FR regressor v2 — codec-aware (vmaf-tune corpus consumer)¶
fr_regressor_v2 — codec-conditioned successor to fr_regressor_v1. Maps a 6-D canonical libvmaf feature vector plus a 14-D codec block to a VMAF teacher score. Trained on the JSONL corpus emitted by vmaf-tune corpus (Phase A, ADR-0237).
Status: production checkpoint.
model/tiny/registry.jsonregistersfr_regressor_v2.onnxwithsmoke: false, SHA-256 pin67934b0b61c73eb852d84ffb34e3333756e8da2530179ecc830336133e63e69e, and an in-sample PLCC of 0.9794 on the vmaf-tune Phase-A JSONL corpus. The old scaffold-only card text is superseded; the follow-up line is now the v3 16-slot vocabulary / LOSO production checkpoint, documented infr_regressor_v3.md.
Inputs¶
Two named tensors, dynamic batch axis:
features, shape(N, 6)— canonical-6 libvmaf features (StandardScaler-normalised at training time using the mean/std baked into the sidecar JSON):
| Index | Feature |
|---|---|
| 0 | adm2 |
| 1 | vif_scale0 |
| 2 | vif_scale1 |
| 3 | vif_scale2 |
| 4 | vif_scale3 |
| 5 | motion2 |
codec, shape(N, 14)— codec block (12 encoder one-hot slots,preset_norm,crf_norm), not normalised (already in[0, 1]):
| Index | Slot |
|---|---|
| 0 | encoder_onehot[libx264] |
| 1 | encoder_onehot[libx265] |
| 2 | encoder_onehot[libsvtav1] |
| 3 | encoder_onehot[libvvenc] |
| 4 | encoder_onehot[libvpx-vp9] |
| 5 | encoder_onehot[h264_nvenc] |
| 6 | encoder_onehot[hevc_nvenc] |
| 7 | encoder_onehot[av1_nvenc] |
| 8 | encoder_onehot[h264_qsv] |
| 9 | encoder_onehot[hevc_qsv] |
| 10 | encoder_onehot[av1_qsv] |
| 11 | encoder_onehot[unknown] |
| 12 | preset_norm (preset ordinal / 9) |
| 13 | crf_norm (CRF / 63) |
Encoder vocabulary (encoder_vocab_version 2) is closed and ordered — index 0..11 is load-bearing; bumping the vocabulary requires a re-train. The unknown bucket lets corpora without codec metadata pass an all-zeros + unknown=1 vector and degrade gracefully.
CRF normalised by 63 — the union upper bound across all supported encoders (libsvtav1 / libvpx-vp9 max). x264 / x265 use CRF up to 51; values above their per-encoder max are clipped at read time.
Preset ordinal table per encoder lives in ai/scripts/train_fr_regressor_v2.py (PRESET_ORDINAL); the canonical 0..9 scale carries the speed-quality direction consistently across encoders. libsvtav1's numeric 0..13 presets are squashed to 0..9.
Architecture¶
The shipped graph (read from the ONNX initialisers) is a GELU MLP over the concatenated 20-D input (6 features + 14 codec values): three hidden layers of 32 units, then a single output unit (about 2 820 parameters), the shape ADR-0291 records. The trainer's defaults (--hidden 32 --depth 3 in ai/scripts/train_fr_regressor_v2.py) reproduce it, and the sidecar's training block records hidden and depth. Before 2026-10-04 the defaults were --hidden 16 --depth 2, a smaller model than the shipped one.
Output¶
score, shape (N,) — a scalar VMAF-aligned quality score per sample, same MOS range as v1 (typically [0, 100]).
Codec-blind fallback¶
For inference paths that don't carry codec metadata, pass an all-zeros codec vector with encoder_onehot[unknown]=1 and preset_norm=0.5, crf_norm=0.5. The model degrades to a v1-like estimate; no graph surgery required.
Training corpus¶
vmaf-tune Phase A JSONL (tools/vmaf-tune/src/vmaftune/corpus.py). One row per (source, encoder, preset, crf) cell with schema_version=1. The trainer reads the JSONL row-by-row; the canonical-6 features come from each row's measured libvmaf feature payload when present, with compatibility aliases for historical corpus runs. --smoke remains available for CI/load-path validation, but the committed fr_regressor_v2.onnx is the production export recorded in the registry.
Training data terms¶
This model was trained on vmaf-tune Phase A encodes of the Netflix Public Dataset's nine references. The terms below are quoted as each source states them (read 2026-10-04); the dataset terms list where each comes from and which models it trained.
Netflix Public Dataset: https://github.com/Netflix/vmaf/blob/0fb4152418d0351901e9c5fd2d30668dced89cdb/resource/doc/datasets.md
We provide a dataset publicly available to the community for training, testing and verification of results purposes.
(please request for access and we will grant it)
The stated purpose includes training; the page states no other terms.
Reading. The fork ships these weights under BSD-2-Clause-Patent: they are fitted parameters that cannot reproduce a clip, an image or a label, and no dataset file is redistributed. That is the fork's reading, not a permission from the dataset's authors. Where a dataset limits its use to research and that limit binds the weights where you use them, treat the model as research-only.
Retrain. RC9 retrains this model on data cleared for redistribution (T-TINY-AI-RETRAIN-CLEARED-DATA-2026-10-04 in state; ADR-1490, ADR-1570).
CLI¶
# Smoke (synthetic corpus, validates the pipeline only)
python ai/scripts/train_fr_regressor_v2.py --smoke
# Production (real Phase A corpus)
python ai/scripts/train_fr_regressor_v2.py \
--corpus runs/vmaf_tune_corpus.jsonl \
--epochs 30
The script bakes the StandardScaler over the canonical-6 dims into the sidecar JSON (feature_mean / feature_std); the codec block is unscaled. Output ONNX is opset 17, dynamic batch axis, op-allowlist checked.
The sidecar and --metrics-out JSON include run_provenance (ai-run-provenance-v1): trainer entrypoint, command arguments, real-corpus path/hash or synthetic-smoke, and output targets. Use this block when comparing refreshed codec-aware runs so stale Phase-A corpora are visible without reverse-engineering shell history.
Checkpoint facts¶
| Field | Value |
|---|---|
| Model id | fr_regressor_v2 |
| Location | model/tiny/fr_regressor_v2.onnx |
| Architecture | GELU MLP, 3 x 32 hidden units, codec conditioning block |
| Input | features [N, 6], codec [N, 14] |
| Output | score [N] |
| ONNX opset | 17 |
| License | BSD-2-Clause-Patent |
| Registry entry | fr_regressor_v2 in model/tiny/registry.json ("smoke": false) |
| SHA-256 | 67934b0b61c73eb852d84ffb34e3333756e8da2530179ecc830336133e63e69e |
Runnable usage example¶
# Score reference and distorted video using the codec-aware model:
vmaf \
--reference python/test/resource/yuv/src01_hrc00_576x324.yuv \
--distorted python/test/resource/yuv/src01_hrc01_576x324.yuv \
--width 576 --height 324 --pixel_format 420 --bitdepth 8 \
--tiny-model model/tiny/fr_regressor_v2.onnx \
--tiny-codec libx264 --tiny-preset medium --tiny-crf 28 \
--json --output /tmp/fr_v2.json
The score is attached under the feature name vmaf_tiny_model (the sidecar has no name). --tiny-codec takes an encoder from the sidecar encoder_vocab (see Inputs); --tiny-preset and --tiny-crf set preset_norm and crf_norm.
Input features and codec context
Loading the model makes the run compute its input features (adm2, vif_scale0..3, motion2, with default options), and the model scores every frame once the run is flushed. A frame without one of them fails the run with a message naming it; no input is read as 0.0. The model also needs --tiny-codec (with the encode's --tiny-preset and --tiny-crf): without it the run stops on the first frame instead of scoring a guessed codec block (ADR-1520).
Known limitations¶
- Feature dependency: requires the canonical-6 feature set (
adm2,vif_scale0..3,motion2) extracted from 8-bit luma planes. - Closed encoder vocabulary: supports
libx264,libx265,libsvtav1,libvvenc,libvpx-vp9,h264_nvenc,hevc_nvenc,av1_nvenc,h264_qsv,hevc_qsv,av1_qsv, andunknown(fallback slot). Thevmaf --tiny-codecoption accepts exactly these names (plus common ffprobe aliases such ash264orhevc). - In-sample evaluation only: the only recorded quality figure is the in-sample PLCC 0.9794 (SROCC 0.9640, RMSE 3.01 on 216 rows) in the sidecar's
trainingblock; there is no held-out or leave-one-source-out number for v2. The LOSO-gated successor isfr_regressor_v3. - Execution providers: validated on CPU (
CPUExecutionProvider) and CUDA (CUDAExecutionProvider). - External data file: uses external data format; companion
fr_regressor_v2.onnx.datamust remain co-located.
See also¶
- ADR-0272 — original scaffold decision; this card now reflects the promoted production checkpoint.
- ADR-0235 — the parent codec-aware decision.
- ADR-0237 — vmaf-tune Phase A (the corpus producer).
- ADR-0249 —
fr_regressor_v1baseline. - Research-0058 — feasibility digest, including the open question on production corpus diversity.
fr_regressor_v2_codec_aware.md— superseded ADR-0235-era design card (canonical-9 / FULL_FEATURES path). The shipped model isfr_regressor_v2(this card), not a separatefr_regressor_v2_codec_aware.onnx.