Tiny-AI — roadmap¶
This page shows what is shipped, planned and deferred on the tiny-AI surface. The four capabilities already in-tree are described in overview.md. The roadmap covers the expansion beyond the scope of ADR-0020 to ADR-0023.
Note
Wave 1 is locked by ADR-0107 (supersedes ADR-0036, the original 2026-04-17 popup approval). Later waves are non-binding and document direction only.
Note
The shipped models were trained against the vmaf_v0.6.1 teacher. The one-shot retrain against vmaf_v1.0.16_3d0h and the remaining tiny-AI training are RC9 work, see the project roadmap.
Status at a glance¶
| Item | Status | ADR | Evidence |
|---|---|---|---|
Training (ai/, vmaf-train) | Shipped | ADR-0020 | training.md |
Inference (core/src/dnn/, ONNX Runtime, op allowlist, 50 MB cap) | Shipped | ADR-0023 | inference.md |
Model signing verification (--tiny-model-verify) | Shipped | ADR-0211 | security.md |
| Model registry | Shipped, 26 entries | ADR-0211 | model-registry.md |
FFmpeg vf_libvmaf tiny options, vmaf_pre (8/10/12-bit, optional chroma) | Shipped | ADR-0482 | ffmpeg-patches/0001, ffmpeg-patches/0002 |
FR baselines (fr_regressor_v1, vmaf_tiny_v2) | Shipped | ADR-0249, ADR-0244 | model cards |
nr_metric_v1, learned_filter_v1 | Shipped | ADR-0168 | nr_metric_v1, learned_filter_v1 |
| LPIPS-SqueezeNet FR extractor | Shipped | none | lpips_sq |
| DISTS-Sq extractor | Shipped with smoke checkpoint; production weights pending (T7-DISTS-followup) | none | dists_sq |
| MobileSal scoring extractor, saliency students | Shipped; saliency_student_v2 is the production default | ADR-0218, ADR-0444 | mobilesal |
tools/vmaf-roi encoder ROI sidecar | Shipped | none | vmaf-roi |
TransNet V2 shot boundaries, tools/vmaf-perShot | Shipped | ADR-0261, ADR-0222 | transnet_v2, vmaf-perShot |
| FastDVDnet extractor (real weights) | Shipped; FFmpeg vmaf_pre_temporal filter planned | ADR-0255 | fastdvdnet_pre |
Allowlist Loop / If with bounded-iteration guard | Shipped | ADR-0169, ADR-0171 | security.md |
vmaf-train tune (Optuna) | Shipped | none | training.md |
FFmpeg vmaf_post filter | Planned, no patch number assigned | none | none |
describe_worst_frames MCP tool | Planned | none | none |
| CLIP-IQA pseudo-labeler, KADID-10k pipeline, Ray tuning backend | Deferred | none | none |
| GPU-parity CI for tiny-AI EPs | Not implemented; cross-EP variance is checked manually | none | inference.md |
vmaf_tiny_v5 corpus expansion | Deferred | ADR-0287 | none |
1. Shipped baseline¶
The shipped surface:
- Training.
ai/(PyTorch + Lightning) and thevmaf-trainCLI. - Inference.
core/src/dnn/: the ONNX Runtime C API behind a 74-entry op allowlist, a model cap of at most 50 MB and a path-hardened loader. - C API.
vmaf_use_tiny_model()andVmafDnnSessionopen, run and close. - CLI.
vmaf --tiny-model PATH --tiny-device STR. The 12 device strings are listed in inference.md. - FFmpeg.
ffmpeg-patches/0001adds the tiny-model options tovf_libvmaf, andffmpeg-patches/0002adds thevmaf_prelearned-filter filter. - Checkpoints.
model/tiny/holds 26 registry entries, see model-registry.md. - Signing.
--tiny-model-verifyis wired tocosign verify-blob(ADR-0211 / T6-9).registry.jsoncarries SHA-256 pins and Sigstore bundle paths.
Still outstanding: there is no GPU-parity CI, so cross-execution-provider variance is verified manually.
2. Wave 1¶
All four sub-lists below were approved in the popup behind ADR-0107 (a paraphrased restatement of the original ADR-0036). Shipping baselines was the blocker for everything else.
2.1 Ship baselines¶
| Model | Role | Status | Target / result |
|---|---|---|---|
fr_regressor_v1.onnx | C1 FR | Shipped 2026-04-29 (ADR-0249). The local Netflix Public drop unblocked the original deferral. | Mean LOSO PLCC vs vmaf_v0.6.1 in model/tiny/fr_regressor_v1.json; ship gate is at least 0.95 |
vmaf_tiny_v2.onnx | C1 FR (canonical-6 fusion) | Shipped 2026-04-29 (ADR-0244). The 3-corpus parquet (Netflix, KoNViD, BVI-DVC D+C) closed the gap the Netflix-only regressor had left. | Netflix LOSO PLCC 0.9978 ± 0.0021 (9 folds x 5 seeds); KoNViD 5-fold PLCC 0.9998; about 257-parameter mlp_small with bundled StandardScaler |
nr_metric_v1.onnx | C2 NR | Shipped 2026-04-25 (ADR-0168) | KoNViD-1k val/MSE 0.382 (RMSE about 0.62 on 1 to 5 MOS); MobileNet-tiny, about 19K parameters |
learned_filter_v1.onnx | C3 filter | Shipped 2026-04-25 (ADR-0168) | KoNViD-1k self-supervised val/L1 0.019 on normalised luma; 4-block residual CNN, about 19K parameters |
The C2 and C3 first run exercised the full pipeline end to end: fetch_konvid_1k.py, vmaf-train manifest-scan, extract_konvid_frames.py, train_konvid.py, export_tiny_models.py, model/tiny/registry.json. C1 followed on 2026-04-29 once the Netflix Public Dataset was locally available.
2.2 LPIPS-SqueezeNet as an FR baseline¶
- Why. Industry-standard perceptual FR. It complements the homegrown C1 with an externally validated reference point. The SqueezeNet variant fits under the size cap (about 2.5M parameters plus about 1.25M frozen features).
- Integration. A feature extractor under
core/src/feature/that callsvmaf_dnn_session_*and emitslpips_sqper frame next to VMAF's composite features. - ONNX. Stock convolutions and global pooling, static input shape, opset 17, no custom ops. Upstream reference:
richzhang/PerceptualSimilarity. - Status. Shipped.
2.3 DISTS-Sq as the LPIPS companion¶
- Why. Bristol VI-Lab's NVC audit flags DISTS as the deep-feature FR companion to LPIPS.
- Integration.
core/src/feature/feature_dists.cmirrors LPIPS' two-input DNN session and emitsdists_sqper frame. - Status. The extractor ships with a smoke checkpoint. Production weights remain
T7-DISTS-followup.
2.4 MobileSal: saliency-weighted VMAF and encoder ROI¶
One saliency model (about 2.5M parameters) feeds two surfaces:
- Scoring side. Multiply the saliency map into the per-pixel residual before spatial pooling in existing feature extractors. This is the SVMAF variant published in academic work but never shipped.
- Encoder side. Emit a per-CTU QP-offset map consumed by
x265 --qpfileor the SVT-AV1 ROI API. Large bitrate win at fixed subjective quality.
Integration status:
mobilesalextractor (scoring side, T6-2a). Emits a scalarsaliency_meanper frame. It shipped with the historical smoke checkpoint first. Production use now points at the fork-trainedsaliency_student_v1weights, andmobilesal_placeholder_v0stays as a registry smoke and legacy artefact. Seemodels/mobilesal.mdand ADR-0218.tools/vmaf-roi(encoder side, T6-2b). Shipped. It writes an ASCII grid for x265 (--qpfile-style) and rawint8_tbinary for SVT-AV1 (--roi-map-file), accepts 8/10/12/16-bit planar YUV and handles one frame per invocation. Seedocs/usage/vmaf-roi.md. Wave-2 follow-ups are a multi-frame batch mode and--blend edge-density.- Evaluation.
eval_saliency_per_mb.pyreports IoU after reducing masks to the block grids the encoder ROI paths consume. Use it before promoting a temporal or video-saliency model.
The upstream MobileSal swap is no longer the production path: ADR-0257 records the CC BY-NC-SA, Google-Drive and RGB-D blockers. The production path is the fork-trained DUTS saliency student, with the same input / saliency_map tensor contract as the placeholder. saliency_student_v2 is the production default since 2026-05-15 (IoU 0.7105 against v1's 0.6558, +8.3%; ADR-0444). Use model/tiny/saliency_student_v2.onnx for new encodes. saliency_student_v1 is kept for regression baselines.
2.5 Per-shot CRF predictor and TransNet V2 shot boundaries¶
- Why. Content-adaptive encoding without an ML framework in the encoder. These are the smallest models on this roadmap (under 1M each), with disproportionate bitrate-at-quality savings.
- Pipeline. (1) TransNet V2 (about 1M) produces per-frame shot-change scores, which become a list of shot timestamps. (2) A per-shot CRF predictor takes a downsampled thumbnail plus classical features (motion energy, spatial complexity) and predicts the CRF that hits the target VMAF on that shot.
- Integration. The standalone CLI
tools/vmaf-perShotwrites an encoder-ingestible sidecar. It does not run inside libvmaf: its output is a parameter hint, not a quality score.
Status:
- Shot-boundary extractor shipped (T6-3a, 2026-04-29, real upstream weights in T6-3a-followup, ADR-0261). The
transnet_v2extractor runs upstream's 100-frame windows (ADR-1527) with the[1, 100, 3, 27, 48] -> [1, 100]ONNX contract and the Soucek and Lokoc 2020 MIT checkpoint atmodel/tiny/transnet_v2.onnx(smoke: false), wrapped by the fork's NTCHW adapter. It emits per-frameshot_boundary_probabilityandshot_boundaryflags. Seemodels/transnet_v2.md. tools/vmaf-perShotshipped (T6-3b, 2026-04-29, ADR-0222), seevmaf-perShot. v1 uses a transparent linear-blend predictor and a frame-difference shot detector fallback. TransNet V2 is available as the libvmaf feature extractor for pipelines that consume feature-collector output directly.- v2 deferred. It will swap the linear blend for a small trained MLP under the same CSV / JSON schema, under a separate ADR, once a labelled per-shot CRF corpus exists.
3. FFmpeg and encoder expansion¶
The slots below are not filled by the current ffmpeg-patches/ series.
3.1 vmaf_pre: 10-bit and chroma (shipped)¶
ffmpeg-patches/0002 ("add vmaf_pre filter (8-bit + 10-bit + optional chroma)") accepts 8, 10 and 12-bit planar 4:2:0, 4:2:2 and 4:4:4 input, and filters the U/V planes when chroma=1 (default 0, luma only). Chroma and HDR sources are where classical pre-filters leave the most budget on the table. Filter options are in inference.md.
3.2 New vmaf_post filter (planned)¶
- Why. Today the pair (reference, distorted) is scored. A post filter would score the actually decoded stream inside an ffmpeg pipeline with the C2 NR model, sharing the backbone with the in-tree NR metric.
- Integration. A new patch in the series, with a filter mirroring
vmaf_pre's shape: frame in, score out, no frame out (measurement only). No patch number is assigned:0004is the Vulkan backend selector shim, seeffmpeg-patches/series.txt.
3.3 FastDVDnet temporal pre-filter¶
- Why. A published temporal denoise CNN (about 2.5M parameters, 5-frame window). Denoise-before-encode is a well-validated bitrate lever for noisy or grainy sources.
- Cost. The filter needs a 5-frame buffer, a bigger lift than per-frame filters. Deferred if Wave 1 is too wide.
- Integration. A new
vmaf_pre_temporalfilter, or a mode flag onvmaf_pre. - Status. The extractor is shipped (T6-7, 2026-04-29, real upstream weights in T6-7b, ADR-0255). The
fastdvdnet_preextractor uses a 5-slot ring buffer with the[1, 5, H, W] -> [1, 1, H, W]ONNX contract and the real m-tassano/FastDVDnet checkpoint undermodel/tiny/fastdvdnet_pre.onnx(smoke: false), wrapped by the fork's luma adapter. The FFmpegvmaf_pre_temporalfilter that consumes the denoised frame buffer remains to be written. Seemodels/fastdvdnet_pre.md.
4. Op-allowlist expansion: bounded Loop and If (shipped)¶
Decision: whitelist Loop and If with a bounded-iteration guard. Published transformer and optical-flow architectures that target ONNX export have bounded loops, and unbounded loops are a sandbox risk (infinite compute, adversarial model). The allowlist entries, the 1024 trip-count cap at export time and the 16-node, depth-8 caps at load time are described in security.md (ADR-0169, ADR-0171).
The expansion unlocks:
- MUSIQ (about 27M): NR transformer with multi-scale attention.
- RAFT-Small (about 1M): optical flow with an iterative GRU update.
- Small VLMs (SmolVLM 256M family): transformer decoder.
Non-goal: Scan, whose more expressive iteration semantics would need a much larger analysis pass.
5. MCP and LLM surfaces¶
5.1 describe_worst_frames MCP tool (planned)¶
When VMAF says "frame 847 is bad", the user still has to open the frame to see why. A local VLM closes that loop in plain English, for example "underexposed in the foreground; mild banding on the sky gradient". It is a debugging affordance, not a scoring component.
Implementation: a new method in mcp-server/vmaf-mcp/, taking a VMAF JSON output path and N.
- Pick the N frames with the largest VMAF delta from the per-frame scores.
- Extract those frames as PNGs, reusing ffmpeg.
- Run SmolVLM (about 256M) locally with a prompt template that asks for artifact types and plausible causes.
- Return a JSON list of
{frame_index, vmaf, caption}.
Model choice is the SmolVLM family. If the 256M variant misses, fall back to Moondream2 (1.8B quantized Q4, fits in 4 GB VRAM). The VLM runs through ONNX Runtime under the extended allowlist (section 4). Absolute path resolution and the 50 MB cap still apply. Larger VLMs need the compile-time VMAF_DNN_DEFAULT_MAX_BYTES constant in core/src/dnn/model_loader.h bumped and the library rebuilt (the historical VMAF_MAX_MODEL_BYTES environment override was retired in T7-12).
6. Training-side items¶
Not in Wave 1, listed so they are not forgotten:
- CLIP-IQA pseudo-labeler. Offline bootstrap for NR datasets.
- KADID-10k synthetic distortion pipeline. Classical augmentation.
- Hyperparameter-tuning Ray backend. Once
tunestabilizes.
vmaf-train tune (Optuna) is implemented, see training.md.
7. Infrastructure items¶
- GPU-parity CI (outstanding). CPU against CUDA and CPU against OpenVINO cross-device variance as a required status check (at most 1e-4 FP32 and 1e-2 FP16, per
inference.md). - Sigstore verification (shipped, ADR-0211).
--tiny-model-verifyis wired tocosign verify-blob; production deployments should set it on. - Model registry (shipped, ADR-0211).
model/tiny/registry.jsoncarries SHA-256 pins, Sigstore bundle paths and license metadata for all 26 entries. See model-registry.md.
8. Out of scope¶
- Training inside libvmaf. ML framework dependencies stay in
ai/and Python. - A second inference runtime (TFLite, ggml). ONNX Runtime is the one runtime.
- Cloud-only or API-dependent models. Everything runs locally.
- Models larger than 50 MB. The cap is the compile-time
VMAF_DNN_DEFAULT_MAX_BYTESconstant; bump it incore/src/dnn/model_loader.hand rebuild when a use case genuinely needs it. Scanand arbitrary control flow, see section 4.
9. Related documents¶
- overview.md: the four existing capabilities.
- training.md:
vmaf-trainCLI and dataset flow. - inference.md: CLI, C API and ffmpeg surfaces.
- benchmarks.md: PLCC/SROCC/RMSE methodology.
- security.md: op allowlist and size cap.
- ADR-0107: this roadmap's authority (supersedes ADR-0036).