vmaf_tiny_v2 — feature-fusion VMAF estimator¶
vmaf_tiny_v2 is a tiny multi-layer perceptron that predicts a VMAF score from six classic libvmaf features (adm2, vif_scale0..3, motion2 — the canonical-6 set used by vmaf_v0.6.1). It replaces vmaf_tiny_v1 as the default tiny VMAF fusion model: same input contract, same output range, +0.005–0.018 PLCC across the validation chain (Netflix LOSO + KoNViD 5-fold).
The model is only the regressor. Feature extraction is unchanged —
adm,vif, andmotionare computed by the existing libvmaf CPU/GPU paths. v2 just gives you a smaller, more accurate fusion head than the upstream SVM.
What the output means¶
A single scalar per frame, on the same 0–100 VMAF scale as the classic SVM regressor.
| Value | Interpretation |
|---|---|
| 100 | Perceptually identical to the reference |
| 80–95 | High-quality encode |
| 60–80 | Visible compression artifacts |
| < 60 | Heavy degradation |
Shipped checkpoint¶
| Field | Value |
|---|---|
| Model name | vmaf_tiny_v2 |
| SHA-256 | d7000bf5c5fd1ed39528546dd8464db44a13cb73cfed33fd8f134e5848b009a7 |
| Location | model/tiny/vmaf_tiny_v2.onnx |
| Architecture | mlp_small — Linear(6, 16) → ReLU → Linear(16, 8) → ReLU → Linear(8, 1), ~257 params |
| Input | features — float32 [N, 6], dynamic batch |
| Feature order | adm2, vif_scale0, vif_scale1, vif_scale2, vif_scale3, motion2 |
| Output | vmaf — float32 [N] |
| ONNX opset | 17 |
| Quantisation | fp32 (size already <2 KB; 8-bit has no shipping payoff) |
| License | BSD-2-Clause-Patent |
| Registry entry | vmaf_tiny_v2 in model/tiny/registry.json |
| Sidecar | model/tiny/vmaf_tiny_v2.json |
| Exporter | ai/scripts/export_vmaf_tiny_v2.py |
| Trainer | ai/scripts/train_vmaf_tiny_v2.py |
The graph bakes the StandardScaler (mean, std) from the training set as Constant nodes that run before the MLP — the runtime feeds raw feature values, the trust-root sha256 covers the calibration values too. There is no out-of-band scaler file to ship or distribute.
Fresh exports add a run_provenance block to the sidecar. It records the exporter entrypoint, CLI arguments, checkpoint input, and ONNX / sidecar output paths so refreshed model artifacts can be traced without reading local shell history.
Effective topology:
features [N, 6]
|
Sub <- mean ([6] constant)
|
Div <- std ([6] constant)
|
Linear(6,16) → ReLU → Linear(16,8) → ReLU → Linear(8,1)
|
Squeeze(-1) -> vmaf [N]
Training data¶
- Netflix Public Dataset (9 sources × encodings — local extract).
- KoNViD-1k (5-fold extract; not redistributed; terms in nr_metric_v1).
- BVI-DVC subsets A + B + C + D (full coverage).
All combined into runs/full_features_4corpus.parquet (330 499 frame-rows × the FULL_FEATURES columns of the time (22; the pool now has 26) + vmaf teacher score from vmaf_v0.6.1). The 4-corpus union is what we fit the StandardScaler and the MLP on for the production export. LOSO + 5-fold are the validation methodology, not the deployment recipe.
The shipped weights were retrained on the full 4-corpus union after the 3-corpus sweep validated the canonical-6 + lr=1e-3 + 90ep configuration (Phase-3 chain → ADR-0244). Adding the BVI-DVC A + B subsets brings the row count from 305 795 to 330 499 (+24 704 rows, +8.1 %) and keeps train PLCC at 0.9999 / RMSE 0.153.
Training data terms¶
This model was trained on the Netflix Public Dataset, KoNViD-1k and BVI-DVC subsets A to D. The terms below are quoted as each source states them (read 2026-10-04); the dataset terms list where each comes from and which models it trained.
Netflix Public Dataset: https://github.com/Netflix/vmaf/blob/0fb4152418d0351901e9c5fd2d30668dced89cdb/resource/doc/datasets.md
We provide a dataset publicly available to the community for training, testing and verification of results purposes.
(please request for access and we will grant it)
The stated purpose includes training; the page states no other terms.
KoNViD-1k: http://database.mmsp-kn.de/konvid-1k-database.html
KoNViD-1k is freely available to the research community.
We took YFCC100m as a baseline database, consisting of 793436 Creative Commons (CC) video sequences
The page names no licence and offers the database to the research community; each clip keeps the Creative Commons licence of its YFCC100M upload.
BVI-DVC: https://fan-aaron-zhang.github.io/assets/copyrights/BVI-DVC.txt
The 800 videos can be used to train CNN models for deep video compression or other computer video tasks, such as image/video super-resolution, image/video enhancement, video frame interpolation, etc..
This database has been compiled by the University of Bristol, Bristol, UK, comprising sequences originally generated by various sources. All intellectual property rights remain with the originators of each sequence. The test sequences from source (15) mentioned above shall only be used for academic research (no commercial use). Material from other sources can also be employed for developing future video coding standards and for evaluating performance of test models in JVET and the subsequent standardization project, and for the relational parent body activities. This copyright and permission notice shall be duplicated whenever the data is copied. The University of Bristol makes no warranties with respect to the material and expressly disclaims any warranties regarding its fitness for any purpose. Unless the above conditions are agreed to by the recipient, no permission is granted for any use and copying of the data. By using the database and sequences, the user agrees to the conditions of this copyright and disclaimer.
Source (15), the Ultra Video Group (Tampere University) sequences, is restricted to academic research.
Reading. The fork ships these weights under BSD-2-Clause-Patent: they are fitted parameters that cannot reproduce a clip, an image or a label, and no dataset file is redistributed. That is the fork's reading, not a permission from the dataset's authors. Where a dataset limits its use to research and that limit binds the weights where you use them, treat the model as research-only.
Retrain. RC9 retrains this model on data cleared for redistribution (T-TINY-AI-RETRAIN-CLEARED-DATA-2026-10-04 in state; ADR-1490, ADR-1570).
Validation¶
Per the Phase-3 chain (Research-0027 → 0028 → 0029 → 0030):
| Methodology | PLCC | SROCC | Notes |
|---|---|---|---|
| Netflix LOSO (9 folds × 5 seeds) | 0.9978 ± 0.0021 | 0.9959 ± 0.0027 | +0.005–0.018 over Subset-B baseline |
| KoNViD 5-fold | 0.9998 | 0.9989 | corpus-portability gate (Phase-3b) |
The min-PLCC = 0.97 ship gate runs in ai/scripts/validate_vmaf_tiny_v2.py against runs/full_features_netflix.parquet first 100 rows; refuses to exit-0 below the gate. Pass --out-json when preserving promotion evidence; the JSON report includes ADR-0661 run_provenance for the ONNX, parquet, parsed gate arguments, and report path.
Usage — CLI¶
# Attach vmaf_tiny_v2 alongside the classic regressor.
vmaf -r ref.yuv -d dis.yuv -w 1920 -h 1080 -p 420 -b 8 \
--tiny-model model/tiny/vmaf_tiny_v2.onnx \
--tiny-device auto
--tiny-model loads the tiny model alongside the classic models rather than replacing the SVM: the classic score stays in the output, and the tiny model's score is added under the feature name vmaf_tiny_model (the sidecar has no name).
Input features
Loading the model makes the run compute its input features (adm2, vif_scale0..3, motion2, with default options), and the model scores every frame once the run is flushed. A frame without one of them fails the run with a message naming it; no input is read as 0.0 (ADR-1520).
--tiny-device auto walks CUDA, OpenVINO GPU, ROCm, CoreML, then CPU. The model is so small (<2 KB) that the dispatch overhead dominates wall-clock on every device; CPU is usually the fastest path.
Usage — Python (ONNX Runtime)¶
For research workflows that already have the canonical-6 features in hand (e.g. from runs/full_features_*.parquet):
import numpy as np
import onnxruntime as ort
import pandas as pd
sess = ort.InferenceSession("model/tiny/vmaf_tiny_v2.onnx",
providers=["CPUExecutionProvider"])
df = pd.read_parquet("runs/full_features_netflix.parquet").head(100)
features = df[
["adm2", "vif_scale0", "vif_scale1",
"vif_scale2", "vif_scale3", "motion2"]
].to_numpy(dtype=np.float32)
(vmaf,) = sess.run(None, {"features": features})
print(vmaf[:5]) # -> per-frame VMAF estimates
Reproducer¶
# 1. Train on the 4-corpus parquet (~12 min CPU on a typical dev box).
python3 ai/scripts/train_vmaf_tiny_v2.py \
--parquet runs/full_features_4corpus.parquet \
--out-ckpt /tmp/vmaf_tiny_v2.pt \
--out-stats /tmp/vmaf_tiny_v2_stats.json
# 2. Export to ONNX with bundled scaler stats.
python3 ai/scripts/export_vmaf_tiny_v2.py \
--ckpt /tmp/vmaf_tiny_v2.pt \
--out-onnx model/tiny/vmaf_tiny_v2.onnx \
--out-sidecar model/tiny/vmaf_tiny_v2.json
# 3. Validate (PLCC must be >= 0.97 on the Netflix slice).
python3 ai/scripts/validate_vmaf_tiny_v2.py \
--onnx model/tiny/vmaf_tiny_v2.onnx \
--parquet runs/full_features_netflix.parquet \
--rows 100 --min-plcc 0.97 \
--out-json runs/vmaf_tiny_v2_validate.json
The training stats JSON includes ADR-0661 run_provenance with the trainer entrypoint, argv, parsed hyperparameters, parquet input, checkpoint target, and stats target. Keep that block with any refreshed stats used for export.
Limitations¶
- The model fuses six already-extracted features — it is not a pixel-input quality model. To use it from raw YUV, the feature extraction stage runs first (the regular libvmaf path).
- Trained on SDR content. HDR coverage is out of scope until the upstream HDR feature extractors land.
- The 4-corpus parquet uses
vmaf_v0.6.1as the teacher score; v2 cannot exceedvmaf_v0.6.1in absolute correctness — it approximates the SVM with a much smaller MLP. - Bit-exactness across CPU/GPU execution providers is not guaranteed (ADR-0042 / ADR-0119 — places=4 tolerance applies to tiny-AI models too).