Tiny AI — overview¶
Tiny AI lets you ship small, specialized perceptual-quality models next to the classic VMAF SVM, without a second ML runtime and without giving up libvmaf's C-only deployment. The feature is gated on -Denable_dnn=auto|enabled|disabled (default auto) and consumed through one public header, libvmaf/dnn.h.
New here? Start with the index, which shows the runnable path in one command.
The four capabilities¶
| # | Name | Shape | What you do with it | Shipped model |
|---|---|---|---|---|
| C1 | FR regressor | feature-vector → MOS | Replace or augment the upstream vmaf_v0.6.1 SVM with a lightweight MLP trained on your own reference dataset. | vmaf_tiny_v2, vmaf_tiny_v3, vmaf_tiny_v4, fr_regressor_v1 to v3 |
| C2 | NR metric | frame → MOS | Predict quality without a reference (live encodes, consumer telemetry). | nr_metric_v1 |
| C3 | Learned filter | frame → frame | Denoise, deblock or sharpen before encoding to save VMAF/PSNR budget. | learned_filter_v1 |
| C4 | LLM dev helpers | repo-time only | Review, commit-message drafting, docgen. Never linked into libvmaf. | none (lives in dev-llm/) |
C1, C2 and C3 share one runtime, the ONNX Runtime C API.
How the pieces fit together¶
Text description
- Surfaces: vmaf CLI (--tiny-model), FFmpeg (vf_libvmaf, vf_vmaf_pre).
- Boxes: ai/ (train, export, register (PyTorch, Lightning)), model/tiny/ (.onnx + sidecar .json, registry.json), core/src/dnn/ (vmaf_use_tiny_model(), ONNX Runtime).
- ai/ → model/tiny/ (export_to_onnx) → core/src/dnn/ (load) → vmaf CLI.
- core/src/dnn/ → FFmpeg.
- Scenario 1, Train and ship: A model leaves Python as an ONNX file and is scored by the C runtime. ai/ trains the model and exports it to ONNX with its sidecar. vmaf_use_tiny_model() opens the committed model through ONNX Runtime. The CLI and the FFmpeg filters share that one runtime.
- Training depends on PyTorch and Lightning; the runtime depends only on ONNX Runtime.
- The boundary between them is the .onnx file and its sidecar JSON on disk.
Training lives in Python and depends on PyTorch and Lightning. Runtime lives in C and depends only on ONNX Runtime. The boundary is the .onnx plus sidecar JSON pair on disk. Git LFS is not used, see tiny-blob-storage.md.
Runtime availability¶
The tiny-AI extractors (lpips, dists_sq, fastdvdnet_pre, mobilesal, transnet_v2) need a libvmaf build with ONNX Runtime support:
- On a build compiled with
-Denable_dnn=disabled, extractorinitreturns-ENOSYSbefore probingmodel_pathor the extractor-specific environment variable. - On DNN-enabled builds, a missing model path stays a normal configuration error and returns
-EINVAL.
When to reach for Tiny AI¶
| You want to… | Use | Read |
|---|---|---|
| Beat the upstream SVM's PLCC on your own MOS data | C1 | training.md |
| Score VMAF without a reference | C2 | inference.md |
| Pre-filter frames before encoding | C3 | inference.md |
| Compare a new model's PLCC/SROCC/RMSE to the SVM baseline | none | benchmarks.md |
| Understand the operator allowlist and signature model | none | security.md |
Documentation rule¶
Every PR that adds or changes a tiny-AI surface ships its documentation in the same PR. The project-wide rule is agent hard rules (rule 7); the tiny-AI five-point bar is per-pr-doc-bar.md (ADR-0042).
Related documents¶
- roadmap.md: status of each tiny-AI item and the Wave 1 scope (LPIPS, saliency, per-shot CRF,
vmaf_post, allowlistLoop/If, MCP VLM tool). - training.md:
vmaf-trainCLI, dataset manifests, export flow. - inference.md: CLI, C API and ffmpeg filter surfaces.
- benchmarks.md: accuracy and throughput methodology.
- security.md: operator allowlist, size cap, Sigstore verification.
Per-model reference¶
Every shipped tiny-AI checkpoint has its own usage page under models/, as ADR-0042 requires. Pages by capability:
- Full-reference regressors: vmaf_tiny_v1, vmaf_tiny_v1_medium (legacy baselines), vmaf_tiny_v2, vmaf_tiny_v3, vmaf_tiny_v4 (progressive VMAF-tiny series; the v5 corpus-expansion proposal remains deferred and no
vmaf_tiny_v5.onnxis shipped), fr_regressor_v1, fr_regressor_v2, fr_regressor_v2 codec-aware, fr_regressor_v2 probabilistic, fr_regressor_v3. - Perceptual distance: LPIPS-SqueezeNet (registry card), DISTS-Sq (smoke checkpoint).
- No-reference and MOS heads: nr_metric_v1, KoNViD MOS head v1.
- Filters and pre-filters: learned_filter_v1, fastdvdnet_pre (5-frame temporal pre-filter).
- Saliency and shots: saliency_student_v1, saliency_student_v2, mobilesal, u2netp mirror, transnet_v2 (shot-boundary detector, about 1M parameters).
- CI smoke fixtures, not quality models: smoke_v0, smoke_v0_symbolic_batch, smoke_fp16_v0, smoke_multi_output_v0.