Skip to content

vmaf_tiny_v4 — top-rung VMAF feature-fusion estimator (opt-in only)

vmaf_tiny_v4 is a tiny multi-layer perceptron that predicts a VMAF score from the same six classic libvmaf features as vmaf_tiny_v2 (adm2, vif_scale0..3, motion2 — the canonical-6 set used by vmaf_v0.6.1). It carries about 4x the parameters of v3 (mlp_large = 6 → 64 → 32 → 16 → 1, 3 073 params vs v3's 769) to answer the question raised in PR #294: does the next rung on the architecture ladder buy further headroom over v3's PLCC = 0.9986 ± 0.0015?

Empirical answer: no, the ladder saturates

v4 ships as an opt-in-only model. Its Netflix 9-fold LOSO PLCC is 0.9987 ± 0.0015, statistically indistinguishable from v3 (+0.0001 mean, identical std). The production default stays vmaf_tiny_v2; the higher-tier opt-in stays v3. See ADR-0390 for the "ladder stops here" rationale.

What the output means

A single scalar per frame on the same 0–100 VMAF scale as the classic SVM regressor and as v2 / v3. Identical interpretation table:

Value Interpretation
100 Perceptually identical to the reference
80–95 High-quality encode
60–80 Visible compression artifacts
< 60 Heavy degradation

Shipped checkpoint

Field Value
Model name vmaf_tiny_v4
SHA-256 (fp32) 2bccd339caa08f25433120546e45f2cb435288edeeefeafc56f3668986904e8f
Location model/tiny/vmaf_tiny_v4.onnx
Architecture mlp_large — Linear(6,64) → ReLU → Linear(64,32) → ReLU → Linear(32,16) → ReLU → Linear(16,1), 3 073 params
Input features — float32 [N, 6], dynamic batch
Feature order adm2, vif_scale0, vif_scale1, vif_scale2, vif_scale3, motion2
Output vmaf — float32 [N]
ONNX opset 17
ONNX size 14 046 bytes
Quantisation fp32 + dynamic-PTQ int8 sidecar (ADR-0275)
License BSD-2-Clause-Patent
Registry entry vmaf_tiny_v4 in model/tiny/registry.json
Sidecar model/tiny/vmaf_tiny_v4.json
Exporter ai/scripts/export_vmaf_tiny_v4.py
Trainer ai/scripts/train_vmaf_tiny_v4.py
LOSO harness ai/scripts/eval_loso_vmaf_tiny_v4.py

The graph bakes the StandardScaler (mean, std) from the training set as Constant nodes that run before the MLP — exactly the same runtime contract as v2 / v3. There is no out-of-band scaler file to ship.

Fresh exports add a run_provenance block to the sidecar. It records the exporter entrypoint, CLI arguments, checkpoint input, and ONNX / sidecar output paths so refreshed model artifacts can be traced without reading local shell history.

Effective topology:

features [N, 6]
   |
   Sub  <- mean   ([6] constant)
   |
   Div  <- std    ([6] constant)
   |
   Linear(6, 64)  -> ReLU
   Linear(64, 32) -> ReLU
   Linear(32, 16) -> ReLU
   Linear(16, 1)
   |
   Squeeze(axis=-1) -> vmaf [N]

How to invoke

The CLI attaches v4 with --tiny-model <path>, alongside the classic models:

vmaf --reference ref.yuv --distorted dist.yuv \
     --width 1920 --height 1080 --pixel_format 420 --bitdepth 8 \
     --tiny-model model/tiny/vmaf_tiny_v4.onnx

The score is added under the feature name vmaf_tiny_model (the sidecar has no name).

Input features

Loading the model makes the run compute its input features (adm2, vif_scale0..3, motion2, with default options), and the model scores every frame once the run is flushed. A frame without one of them fails the run with a message naming it; no input is read as 0.0 (ADR-1520).

The MCP vmaf_score tool exposes the flag as the tiny_model argument and passes the string unchanged to --tiny-model, so it must be a path as well; there is no registry-id lookup.

Training recipe (reproducible)

# 1. Train (~40 s wall on 16-thread CPU, ~10 min CPU time).
python3 ai/scripts/train_vmaf_tiny_v4.py \
    --parquet runs/full_features_4corpus.parquet \
    --out-ckpt runs/vmaf_tiny_v4.pt \
    --out-stats runs/vmaf_tiny_v4_scaler.json

# 2. Export to ONNX with bundled scaler stats.
python3 ai/scripts/export_vmaf_tiny_v4.py \
    --ckpt runs/vmaf_tiny_v4.pt \
    --out-onnx model/tiny/vmaf_tiny_v4.onnx \
    --out-sidecar model/tiny/vmaf_tiny_v4.json

# 3. Smoke validate (PLCC >= 0.97 gate; v2 / v3 diff sanity check).
python3 ai/scripts/validate_vmaf_tiny_v4.py \
    --onnx model/tiny/vmaf_tiny_v4.onnx \
    --parquet runs/full_features_netflix.parquet \
    --rows 5000 --min-plcc 0.97 \
    --v2-onnx model/tiny/vmaf_tiny_v2.onnx \
    --v3-onnx model/tiny/vmaf_tiny_v3.onnx \
    --out-json runs/vmaf_tiny_v4_validate.json

# 4. 9-fold LOSO eval (~12 s wall total).
python3 ai/scripts/eval_loso_vmaf_tiny_v4.py \
    --parquet runs/full_features_netflix.parquet \
    --out-json runs/vmaf_tiny_v4_loso_metrics.json

The training stats JSON includes ADR-0661 run_provenance with the trainer entrypoint, argv, parsed hyperparameters, parquet input, checkpoint target, and stats target. Keep that block with any refreshed stats used for export.

The generated LOSO report includes run_provenance with the evaluation entrypoint, original argv, parsed hyperparameters, feature parquet input, and report target path. Treat that block as the run identity when comparing v4 refreshes or multi-run audit output.

The smoke-validation JSON from --out-json also carries ADR-0661 run_provenance, including the validated ONNX, parquet slice, optional v2/v3 comparison models, parsed gate arguments, and report path.

Hyperparameters (identical to v2 / v3 — only the architecture changes):

  • Optimiser: Adam @ lr=1e-3.
  • Loss: MSE.
  • Epochs: 90.
  • Batch size: 256.
  • Seed: 0.
  • Standardisation: corpus-wide mean / std on the training split, baked into ONNX.

Training data terms

This model was trained on the Netflix Public Dataset, KoNViD-1k and BVI-DVC subsets A to D. The terms below are quoted as each source states them (read 2026-10-04); the dataset terms list where each comes from and which models it trained.

Netflix Public Dataset: https://github.com/Netflix/vmaf/blob/0fb4152418d0351901e9c5fd2d30668dced89cdb/resource/doc/datasets.md

We provide a dataset publicly available to the community for training, testing and verification of results purposes.

(please request for access and we will grant it)

The stated purpose includes training; the page states no other terms.

KoNViD-1k: http://database.mmsp-kn.de/konvid-1k-database.html

KoNViD-1k is freely available to the research community.

We took YFCC100m as a baseline database, consisting of 793436 Creative Commons (CC) video sequences

The page names no licence and offers the database to the research community; each clip keeps the Creative Commons licence of its YFCC100M upload.

BVI-DVC: https://fan-aaron-zhang.github.io/assets/copyrights/BVI-DVC.txt

The 800 videos can be used to train CNN models for deep video compression or other computer video tasks, such as image/video super-resolution, image/video enhancement, video frame interpolation, etc..

This database has been compiled by the University of Bristol, Bristol, UK, comprising sequences originally generated by various sources. All intellectual property rights remain with the originators of each sequence. The test sequences from source (15) mentioned above shall only be used for academic research (no commercial use). Material from other sources can also be employed for developing future video coding standards and for evaluating performance of test models in JVET and the subsequent standardization project, and for the relational parent body activities. This copyright and permission notice shall be duplicated whenever the data is copied. The University of Bristol makes no warranties with respect to the material and expressly disclaims any warranties regarding its fitness for any purpose. Unless the above conditions are agreed to by the recipient, no permission is granted for any use and copying of the data. By using the database and sequences, the user agrees to the conditions of this copyright and disclaimer.

Source (15), the Ultra Video Group (Tampere University) sequences, is restricted to academic research.

Reading. The fork ships these weights under BSD-2-Clause-Patent: they are fitted parameters that cannot reproduce a clip, an image or a label, and no dataset file is redistributed. That is the fork's reading, not a permission from the dataset's authors. Where a dataset limits its use to research and that limit binds the weights where you use them, treat the model as research-only.

Retrain. RC9 retrains this model on data cleared for redistribution (T-TINY-AI-RETRAIN-CLEARED-DATA-2026-10-04 in state; ADR-1490, ADR-1570).

Validation results

Netflix 9-fold LOSO (single seed, parity with v3 protocol)

Source n PLCC SROCC RMSE
BigBuckBunny 1500 0.9995 0.9964 0.86
BirdsInCage 1440 0.9999 0.9998 0.30
CrowdRun 1050 0.9999 0.9998 0.76
ElFuente1 1260 0.9993 0.9952 1.13
ElFuente2 1620 0.9989 0.9996 1.22
FoxBird 900 0.9952 0.9955 2.82
OldTownCross 1050 0.9992 0.9999 1.23
Seeking 1500 0.9989 0.9961 1.63
Tennis 720 0.9977 0.9978 1.15
Mean 0.9987 0.9978 1.23
Std (n-1) 0.0015 0.0020 0.70

Comparison vs v2 / v3

Metric v2 (mlp_small, 257) v3 (mlp_medium, 769) v4 (mlp_large, 3 073)
NF LOSO mean PLCC 0.9978 (5-seed) 0.9986 (1-seed) 0.9987 (1-seed)
NF LOSO std PLCC 0.0021 0.0015 0.0015
NF LOSO mean SROCC 0.9959 0.9977 0.9978
5000-row smoke PLCC 0.9998 1.0000 1.0000
Train-set RMSE 0.153 0.112 0.104
ONNX size (bytes) 2 446 4 496 14 046
Params 257 769 3 073

The +0.0001 mean PLCC delta v3 → v4 is well below 1 std of either model. Train-set RMSE keeps improving (v4 over-fits a bit harder thanks to the ~4x parameters), but the held-out PLCC has saturated.

Why pick v4 (or not)

Model Params ONNX bytes LOSO PLCC (mean) Pick it when
vmaf_tiny_v2 257 2 446 0.9978 (5-seed) You want the tightest bundle, lowest dispatch overhead and the cited Phase-3 baseline (production default)
vmaf_tiny_v3 769 4 496 0.9986 (1-seed) You want the higher-tier model with the smallest ONNX that still beats v2 measurably
vmaf_tiny_v4 3 073 14 046 0.9987 (1-seed) You want the top of the measured ladder, run CPU-only inference where ~14 KB is irrelevant, and want maximum train-set fidelity

v4's win over v3 is below noise, so v3 is the usual opt-in choice.

Quantisation (dynamic-PTQ int8 sidecar — ADR-0275)

A dynamic-PTQ int8 sibling lives at model/tiny/vmaf_tiny_v4.int8.onnx, produced by ai/scripts/ptq_dynamic.py. The runtime redirect in vmaf_dnn_session_open (ADR-0174) loads the int8 sidecar when the registry entry's quant_mode != "fp32"; the fp32 file stays on disk as the regression baseline.

Field Value
Quant mode dynamic (weights-only int8; activations quantised at runtime)
Sidecar file model/tiny/vmaf_tiny_v4.int8.onnx
sha256 (int8) 203a25905a3797b1cc3e3f347f4f6f491b491000d4424017beef64219767a9e9
File size 7 769 B int8 vs 14 046 B fp32 (×0.55) — 45 % smaller
PLCC drop (int8 vs fp32) 0.000145 (Netflix features parquet, ~11k rows)
Budget 0.01 (per ADR-0174 / ADR-0173 default)
CI gate ai-quant-accuracy step in Tiny AI job

mlp_large carries enough weight mass (3 073 params) that the int8 weight tensors actually drive the on-disk size. v4 is therefore the first MLP-flavoured tiny model where dynamic PTQ delivers a meaningful size win — v2 (~257 params) and v3 (~769 params) shrink only marginally because their fp32 ONNX is dominated by op metadata and Constant scaler nodes rather than weights.

# Reproduce
python ai/scripts/ptq_dynamic.py model/tiny/vmaf_tiny_v4.onnx
python ai/scripts/measure_quant_drop.py model/tiny/vmaf_tiny_v4.onnx
# -> [PASS] vmaf_tiny_v4   mode=dynamic PLCC=0.999855  drop=0.000145  budget=0.0100

Limitations

  • Same canonical-6 input contract as v2 / v3 — no new features. v4's quality ceiling is the canonical-6 information bottleneck, not its arch.
  • Trained 4-corpus (NF Public + KoNViD + BVI-DVC A+B+C+D, 330 499 rows). Out-of-distribution content (HDR, 8K, screen content, animation outside the corpora) inherits the corpus's coverage limitations.
  • Single-seed LOSO; v3 also single-seed. v2 was 5-seed. A multi-seed v4 LOSO study (5+ seeds) would tighten the variance estimate but is not gating; the saturation evidence is decisive enough.
  • The architecture ladder stops here. Future tiny-VMAF gains require regime change (richer features, larger corpus, ensembles, distillation), not deeper / wider MLPs. See ADR-0390.