Edge-AI compiler

Deploy Edge-AI Models to Bare-Metal Microcontrollers in Seconds

Qovox compiles your trained ONNX or PyTorch model into lean, zero-dependency C++ — built for STM32, ESP32, ARM Cortex-M, RISC-V and automotive ECUs.

3 compiles a month. No card required.

input
# resnet8.onnx  ·  graph summary
input  : float32[1,3,32,32]
Conv   → BatchNorm → Relu
Conv   → BatchNorm → Relu
Add    (residual)
GlobalAveragePool
Gemm   (10 classes)
output : float32[1,10]

params: 78K · fp32 · 312 KB
C++qovox_model.h
// generated by Qovox · no libc++, no heap
#pragma once
#include <stdint.h>

namespace qovox {
  constexpr size_t ARENA = 9216;
  void infer(const int8_t* in,
             int8_t* out);
}
// static arena · fused conv+bn+relu
arena 9 KBflash 78 KBheap 0 Bdeps none

Runs on your chipset · Reads your framework

STM32MCU
ESP32MCU
ARM Cortex-MCPU
RISC-VISA
PyTorchFramework
ONNXFormat
Product

Built for hardware that can't afford waste

Every byte of RAM and flash counts on a microcontroller. Qovox compiles it away.

INT8 Quantization & Pruning

Calibrated post-training quantization and structured pruning shrink models up to 4× with minimal accuracy loss.

fp32 → int8structured pruningper-channel scales

Zero-Dependency C++ Export

Plain C++ header and weights. No runtime, no heap, no vendor lock-in.

CAN-Bus & Automotive Readiness

Deterministic, static-memory output suited to ECU and CAN-bus pipelines.

Sub-millisecond Latency

Fused layers and SIMD-friendly loops keep small vision and signal models in real-time budgets.

< 1 ms

Benchmarks

TFLite Micro vs. Qovox compiled header

Same model, same board, same accuracy.

TFLite MicroQovox
RAM usage
24 KB
9 KB
Flash footprint
260 KB
78 KB
Inference speed
3.0 ms
1.0 ms

Illustrative sample (ResNet-8, Cortex-M4). Shorter bars are better.

Model Zoo

Quantization results on popular models

Size and latency before and after Qovox INT8 compilation. Filter by task or search by name.

ModelCategoryFP32 → INT8 sizeCompression FP32 → Qovox WASM latencySpeedupActions

Illustrative sample figures, not measured results. Latency is single-inference time in the browser (WASM).

Upload, compile, ship

Three steps from a trained model to code you can build.

01

Upload

Drop in an ONNX, PyTorch, TensorFlow or GGUF model file.

02

Compile

Qovox fuses layers, quantizes weights and emits C++.

03

Ship

Download the source and build it anywhere with CMake.

Live demo

Run a quantized model in your browser

Pick a sample model. ONNX models run locally with ONNX Runtime Web (nothing is uploaded); other formats show a simulated Qovox compiler pipeline.

The model loads when this section comes into view.

Live benchmark
Backend—
CPU threads—
Size reduction—
Inference time—

Output inspector
Select a model and press “Run Model”.

Compile your first model free

3 compiles a month, no card required.