Documentation

Qovox Engine SDK

Run Qovox-compiled ONNX models directly in the browser with qovox-engine.js. No server involved: inference runs on your visitor's device with WebGPU or multi-threaded WASM.

Engine v1.2.0Vanilla JS · React · Next.js · Vue 3ONNX Runtime Web

Quickstart

You need three things on the same origin: the engine script, your compiled .onnx model, and (recommended) the manifest.json that sits next to the model. The engine downloads onnxruntime-web on demand, so there is nothing else to install.

1. Add the script

Copy qovox-engine.js to your site and include it with a normal script tag. It exposes a global QovoxEngine.

<script src="./qovox-engine.js"></script>

2. Load a model and run inference

<script src="./qovox-engine.js"></script>
<script>
  (async () => {
    const model  = await QovoxEngine.load('./model_quantized.onnx');
    const result = await model.predict(model.sampleInputs());
    console.log(result.meta);
  })();
</script>

result.meta tells you what actually happened. The values below are an example; yours will differ.

{
  backend: "wasm",        // "webgpu" or "wasm", the backend that really ran
  threads: 1,             // WASM threads (see Server headers)
  inferenceMs: 4.21,      // pure inference time
  fallbackReason: null,   // set if WebGPU was abandoned for WASM
  numericWarning: null,   // set if the output still contains NaN/Inf
  inputs:  [ { name: "input", dims: [1, 3, 224, 224], type: "float32" } ],
  outputs: [ { name: "output", dims: [1, 1000], type: "float32", finite: true } ]
}

3. Feed real data

sampleInputs() generates random test data. For real inputs, pass a tensor-like object keyed by input name. This example sends an already-preprocessed Float32Array (NCHW) to an image model:

const S = 224;
const inputName = model.inputNames[0];

const out = await model.predict({
  [inputName]: { data: pixels, dims: [1, 3, S, S] },   // pixels: Float32Array(3*S*S)
});

const logits = out[model.outputNames[0]].data;           // Float32Array
console.log(out.meta.inferenceMs + ' ms on ' + out.meta.backend);

Privacy: the model and the inputs never leave the visitor's browser. The only network requests are the model download and the one-time onnxruntime-web download from jsDelivr (or your own ortBaseUrl).

The manifest file

If a manifest.json exists next to the model, the engine uses it for input names, dtypes, default shapes and backend hints. Without it, sampleInputs() cannot work and you must always pass { data, dims } yourself.

{
  "inputs":  [ { "name": "input", "dtype": "float32", "shape": [1, 3, 224, 224] } ],
  "outputs": [ "output" ],
  "quantization": { "mode": "dynamic_int8" },
  "recommended_backend": "wasm"
}

Recognised fields: inputs[].name, dtype, shape, default_shape, outputs, quantization.mode, recommended_backend, backend_note, ort_web_version, ort_base_url.

Framework integration guides

The engine is a UMD file, so it works as a plain script tag or through any bundler. In frameworks, load the model once on the client, keep it in a ref (not in reactive state), and dispose() it on unmount.

Use a script tag, create the model once and reuse it for every inference.

<progress id="bar" max="1" value="0"></progress>
<button id="run" disabled>Run</button>
<pre id="out"></pre>

<script src="./qovox-engine.js"></script>
<script>
  let model;

  async function init() {
    model = await QovoxEngine.load('./model_quantized.onnx', {
      onProgress: (loaded, total) => {
        if (total) document.getElementById('bar').value = loaded / total;
      },
    });
    document.getElementById('run').disabled = false;
  }

  document.getElementById('run').addEventListener('click', async () => {
    const result = await model.predict(model.sampleInputs());
    document.getElementById('out').textContent = JSON.stringify(result.meta, null, 2);
  });

  init().catch((err) => { document.getElementById('out').textContent = err.message; });
</script>

A reusable hook. The model lives in a ref; React state only holds progress and errors. The dynamic import() keeps the engine out of server rendering.

'use client';   // Next.js App Router: this file must be a Client Component
import { useEffect, useRef, useState } from 'react';

export function useQovoxModel(url, options) {
  const modelRef = useRef(null);
  const [status, setStatus] = useState({ ready: false, progress: 0, error: null });

  useEffect(() => {
    let cancelled = false;
    (async () => {
      try {
        const mod = await import('./qovox-engine.js');
        const QovoxEngine = mod.default ?? mod;
        const model = await QovoxEngine.load(url, {
          ...options,
          onProgress: (loaded, total) => {
            if (!cancelled && total) setStatus((s) => ({ ...s, progress: loaded / total }));
          },
        });
        if (cancelled) { await model.dispose(); return; }
        modelRef.current = model;
        setStatus({ ready: true, progress: 1, error: null });
      } catch (error) {
        if (!cancelled) setStatus((s) => ({ ...s, error }));
      }
    })();
    return () => {
      cancelled = true;
      if (modelRef.current) { modelRef.current.dispose(); modelRef.current = null; }
    };
  }, [url]);

  return { modelRef, ...status };
}
'use client';
import { useState } from 'react';
import { useQovoxModel } from './useQovoxModel';

export default function Demo() {
  const { modelRef, ready, progress, error } = useQovoxModel('/models/model_quantized.onnx');
  const [meta, setMeta] = useState(null);

  async function run() {
    const model = modelRef.current;
    const result = await model.predict(model.sampleInputs());
    setMeta(result.meta);
  }

  if (error) return <p>Failed to load: {error.message}</p>;
  if (!ready) return <p>Loading model… {Math.round(progress * 100)}%</p>;
  return (
    <div>
      <button onClick={run}>Run model</button>
      {meta && <pre>{JSON.stringify(meta, null, 2)}</pre>}
    </div>
  );
}

Next.js: put the model, manifest.json and (if you use the script-tag route instead of import) qovox-engine.js in public/. COOP/COEP headers are set in next.config.js via headers(); see Server headers below.

Composition API. Use shallowRef for the model: wrapping it in a deep reactive proxy is unnecessary and can interfere with the underlying runtime objects.

<script setup>
import { ref, shallowRef, onMounted, onBeforeUnmount } from 'vue';
import * as engineModule from './qovox-engine.js';

const QovoxEngine = engineModule.default ?? engineModule;

const model = shallowRef(null);
const progress = ref(0);
const error = ref(null);
const meta = ref(null);
let disposed = false;

onMounted(async () => {
  try {
    const m = await QovoxEngine.load('/models/model_quantized.onnx', {
      onProgress: (loaded, total) => { if (total) progress.value = loaded / total; },
    });
    if (disposed) { await m.dispose(); return; }
    model.value = m;
  } catch (e) {
    error.value = e;
  }
});

onBeforeUnmount(() => {
  disposed = true;
  if (model.value) model.value.dispose();
});

async function run() {
  const result = await model.value.predict(model.value.sampleInputs());
  meta.value = result.meta;
}
</script>

<template>
  <p v-if="error">Failed to load: {{ error.message }}</p>
  <p v-else-if="!model">Loading model… {{ Math.round(progress * 100) }}%</p>
  <div v-else>
    <button @click="run">Run model</button>
    <pre v-if="meta">{{ JSON.stringify(meta, null, 2) }}</pre>
  </div>
</template>

API reference

The global QovoxEngine exposes load(), capabilities() and version. load() resolves to a model object with predict(), sampleInputs() and a few helpers.

QovoxEngine.load(url, options)

const model = await QovoxEngine.load(url: string, options?: object): Promise<QovoxModel>

Downloads the model, picks a backend, creates the inference session and, for WebGPU, verifies it with a probe inference.

OptionDefaultDescription
backend'auto''auto', 'webgpu' or 'wasm'. With 'auto', INT8-quantized models go straight to WASM (WebGPU does not reliably run integer operators); other models try WebGPU first, then WASM. Requesting 'webgpu' on a browser without it throws.
tryWebGpuForInt8falseWith backend: 'auto', try WebGPU even for INT8 models. Results are still verified and rolled back to WASM if wrong.
onProgressnone(loaded, total) => void, called while the model downloads. total is 0 when the server sends no Content-Length, so guard against dividing by zero.
threads'auto'WASM thread count (auto uses up to 4). Only effective on a cross-origin isolated page; otherwise 1.
fallbacktrueIf WebGPU throws or returns NaN/Inf, rebuild the session on WASM from the already-downloaded bytes and retry, at load time and on every predict().
verifytrueRun a probe inference after load. true = WebGPU only; 'always' = also for WASM; false = skip.
graphOptimizationLevel'all''all', 'extended', 'basic' or 'disabled'.
ortBaseUrlpinned jsDelivrWhere to load onnxruntime-web from. Set this to self-host it (relative URLs resolve against the manifest URL).
ortVersion'1.22.0'Version used to build the default CDN URL.
manifestUrlmanifest.json next to the modelCustom manifest location.
const model = await QovoxEngine.load('/models/model_quantized.onnx', {
  backend: 'auto',
  tryWebGpuForInt8: false,
  onProgress: (loaded, total) => console.log(total ? Math.round((loaded / total) * 100) + '%' : loaded + ' bytes'),
});

model.predict(inputs)

const result = await model.predict(inputs): Promise<{ [outputName]: { data, dims, type } }>

Calls are queued, so concurrent predict() calls are safe.

Accepted inputs

FormNotes
Object keyed by input name{ input_ids: {data, dims}, attention_mask: {data, dims} }. Required for models with more than one input; every input name must be present.
{ data, dims, type? }Single-input shorthand. type defaults to the manifest dtype, then float32.
Typed array / arraySingle-input shorthand that uses the manifest's default shape, so a manifest is required.
ort.TensorPassed through unchanged.

Data is coerced to the right typed array. For int64 inputs, plain numbers are converted to BigInt64Array for you. The number of values must equal the product of dims, or predict() rejects with a clear message.

Return value

The resolved object maps each output name to { data, dims, type }, with a non-enumerable meta property attached:

const result = await model.predict(inputs);

const logits = result[model.outputNames[0]].data;   // typed array
const { backend, threads, inferenceMs, fallbackReason, numericWarning } = result.meta;

Prefer an explicit shape? predictDetailed() resolves to { outputs, meta } instead (so result.outputs and result.meta).

meta fieldDescription
backend'webgpu' or 'wasm': the backend that produced this result.
threadsThread count in use.
inferenceMsInference time in milliseconds (model run only, excludes pre/post-processing).
fallbackReasonWhy WebGPU was abandoned, or null.
numericWarningSet when the active backend still returns NaN/Inf, otherwise null.
inputs, outputsName, dims and type of every tensor. Each output also has a finite flag.

model.sampleInputs()

const inputs = model.sampleInputs(): { [inputName]: { data, dims, type } }

Builds ready-to-use test inputs from the manifest shapes. Use it for smoke tests, demos and benchmarking without writing preprocessing code. It throws if the manifest has no shape for an input.

  • float inputs: random values in [-1, 1]
  • names containing mask: ones (never zeros, since a fully masked row can produce NaN)
  • names containing position: 0..n-1 along the last axis
  • names containing type or segment: zeros
  • other integer inputs (such as input_ids): small valid ids in 1..97
const inputs = model.sampleInputs();
// { input: { data: Float32Array(150528), dims: [1, 3, 224, 224], type: 'float32' } }

// Quick benchmark on your own device:
const times = [];
for (let i = 0; i < 20; i++) {
  const r = await model.predict(inputs);
  times.push(r.meta.inferenceMs);
}
console.log('median ms:', times.sort((a, b) => a - b)[10]);

Other methods and properties

MemberDescription
model.predictDetailed(inputs)Like predict() but resolves to { outputs, meta }.
model.info()Diagnostics: engine version, backend, threads, crossOriginIsolated, threadsNote (explains single-threading), quantization info, fallback reason, declared inputs.
model.dispose()Releases the session and memory. Call it when you are done (component unmount, page teardown).
model.inputNames, model.outputNamesArrays of tensor names.
model.inputsInput metadata (name, dtype, shape, defaultShape) from the manifest.
model.backend, model.threadsCurrent backend and thread count (can change from webgpu to wasm after a fallback).
model.note, model.fallbackReason, model.numericWarning, model.lastMetaWhy auto chose this backend; why WebGPU was abandoned; NaN/Inf warning; metadata of the latest run.
QovoxEngine.capabilities()Async. Returns { webgpu, crossOriginIsolated, threads } for the current browser.
QovoxEngine.versionEngine version string.

Server configuration: COOP / COEP headers

Multi-threaded WASM needs SharedArrayBuffer, and browsers only expose it on cross-origin isolated pages. Without the two headers below the engine still works, but ONNX Runtime Web is limited to a single thread. With them, the engine uses up to 4 threads.

HeaderValue
Cross-Origin-Opener-Policysame-origin
Cross-Origin-Embedder-Policyrequire-corp

Send both headers on the HTML document that runs the engine.

server {
  listen 443 ssl;
  server_name example.com;
  root /var/www/site;

  add_header Cross-Origin-Opener-Policy   "same-origin"  always;
  add_header Cross-Origin-Embedder-Policy "require-corp"  always;

  location / {
    try_files $uri $uri/ =404;
  }
}
sudo nginx -t && sudo systemctl reload nginx
<IfModule mod_headers.c>
  Header always set Cross-Origin-Opener-Policy   "same-origin"
  Header always set Cross-Origin-Embedder-Policy "require-corp"
</IfModule>
sudo a2enmod headers && sudo systemctl reload apache2
from fastapi import FastAPI
from fastapi.staticfiles import StaticFiles

app = FastAPI()

@app.middleware("http")
async def isolation_headers(request, call_next):
    response = await call_next(request)
    response.headers["Cross-Origin-Opener-Policy"] = "same-origin"
    response.headers["Cross-Origin-Embedder-Policy"] = "require-corp"
    return response

app.mount("/", StaticFiles(directory="public", html=True), name="site")
const express = require('express');
const app = express();

app.use((req, res, next) => {
  res.setHeader('Cross-Origin-Opener-Policy', 'same-origin');
  res.setHeader('Cross-Origin-Embedder-Policy', 'require-corp');
  next();
});

app.use(express.static('public'));
app.listen(3000);
// next.config.js
module.exports = {
  async headers() {
    return [{
      source: '/(.*)',
      headers: [
        { key: 'Cross-Origin-Opener-Policy',   value: 'same-origin' },
        { key: 'Cross-Origin-Embedder-Policy', value: 'require-corp' },
      ],
    }];
  },
};

Verify it worked

console.log(crossOriginIsolated);          // true when isolation is active

const model = await QovoxEngine.load('./model_quantized.onnx');
console.log(model.threads);                // > 1 on a multi-core device
console.log(model.info().threadsNote);     // null, or the reason it is single-threaded

Cross-origin resources: with require-corp, every cross-origin subresource on that page (scripts, images, fonts, the model download) must be served with CORS or a Cross-Origin-Resource-Policy header, or the browser blocks it. If something stops loading after you enable the headers, either self-host that resource (for example set ortBaseUrl to your own copy of onnxruntime-web), apply the headers only to the pages that run the engine, or test Cross-Origin-Embedder-Policy: credentialless, which is supported in Chromium and Firefox but, as far as I know, not in Safari. Always test in your target browsers.

Error handling & fallbacks

The engine already protects you from the most common WebGPU problem: if a WebGPU session throws or returns NaN/Inf, it rebuilds the session on WASM from the bytes it already downloaded and retries, with no second download. Your code should still handle load failures and surface what actually ran.

async function createModel(url, onProgress) {
  try {
    const model = await QovoxEngine.load(url, { onProgress });
    if (model.fallbackReason) console.info('Fell back to WASM:', model.fallbackReason);
    if (model.numericWarning) console.warn('Numerical warning:', model.numericWarning);
    return model;
  } catch (err) {
    // Network failure, CORS, unsupported browser, or ONNX Runtime could not load
    showError('Could not load the model: ' + err.message);
    return null;
  }
}

async function infer(model, inputs) {
  try {
    const result = await model.predict(inputs);
    if (result.meta.numericWarning) {
      // Output still has NaN/Inf: do not trust it. Log it and tell the user.
      console.warn(result.meta.numericWarning);
    }
    return result;
  } catch (err) {
    showError('Inference failed: ' + err.message);   // for example a shape mismatch
    return null;
  }
}

Best practices

  • Keep backend: 'auto' and fallback: true (the defaults) unless you have a measured reason to change them.
  • Load the model once, lazily (when the feature scrolls into view or is first used), and reuse it. Call dispose() on teardown.
  • Show meta.backend and meta.inferenceMs in debug UIs. They tell you immediately whether a fallback happened.
  • Never feed all-zero attention masks. sampleInputs() avoids this, and your own inputs should too.
  • Handle the case where onProgress reports total = 0 (no Content-Length).

Troubleshooting

SymptomLikely cause and fix
threads is always 1The page is not cross-origin isolated. Add the COOP/COEP headers and check crossOriginIsolated. See model.info().threadsNote.
Model download fails or CORS errorHost the model on the same origin, or allow your origin through CORS on the host (Hugging Face and most CDNs do). With COEP enabled, the host must also be CORS or CORP enabled.
Could not load onnxruntime-webThe CDN is blocked (CSP, firewall, offline). Self-host the files and set ortBaseUrl.
NaN or Inf in outputsUsually WebGPU with an INT8 model; the engine rolls back to WASM automatically. If numericWarning persists on WASM, check your input preprocessing and masks.
Input "x" has N values but shape [...] needs MThe data length does not match dims. Check the shape in your manifest or your tensor.
shape unknown, pass {data, dims} / No manifest shapeNo manifest.json was found next to the model. Add it, set manifestUrl, or pass { data, dims } explicitly.
manifest.json 404 in the network tabHarmless: the manifest is optional. Provide one to enable sampleInputs() and backend hints.

Need help? Email qovox222@gmail.com or try the live demo to see the engine in action.