FR regressor v3 — codec-aware on ENCODER_VOCAB v3 (16-slot)¶
fr_regressor_v3 — codec-aware FR regressor trained on ENCODER_VOCAB_V3 (16 slots). Parallel-shipped successor to fr_regressor_v2. Maps a 6-D canonical libvmaf feature vector plus an 18-D codec block (16 encoder one-hot + preset_norm + crf_norm) to a VMAF teacher score scalar.
Status: production checkpoint (gate-passed)
- Gate: mean LOSO PLCC is 0.9975 across the 9 Netflix Public Dataset sources, above the ADR-0302 ship gate of 0.95 (the gate ADR-0291 cleared on v2).
- Registry: ships under ADR-0323; the registry row
fr_regressor_v3hassmoke: false. - v2 stays authoritative for its slot: the live
ENCODER_VOCAB_VERSION = 2inai/scripts/train_fr_regressor_v2.pyremains authoritative forfr_regressor_v2.onnx. Promoting v3 to "the" canonicalfr_regressor_v2.onnxslot is a separate follow-up PR (see ADR-0302 §Production-flip checklist).
Inputs¶
Two named tensors, dynamic batch axis (matches the vmaf_dnn_session_run two-input contract from ADR-0040 / ADR-0022).
features¶
Shape (N, 6): canonical-6 libvmaf features, StandardScaler-normalised at training time using the mean/std baked into the sidecar JSON (feature_mean, feature_std):
| Index | Feature |
|---|---|
| 0 | adm2 |
| 1 | vif_scale0 |
| 2 | vif_scale1 |
| 3 | vif_scale2 |
| 4 | vif_scale3 |
| 5 | motion2 |
codec_block¶
Shape (N, 18): codec block, not normalised (already in [0, 1]):
| Index | Slot |
|---|---|
| 0 | encoder_onehot[libx264] |
| 1 | encoder_onehot[libaom-av1] |
| 2 | encoder_onehot[libx265] |
| 3 | encoder_onehot[h264_nvenc] |
| 4 | encoder_onehot[hevc_nvenc] |
| 5 | encoder_onehot[av1_nvenc] |
| 6 | encoder_onehot[h264_amf] |
| 7 | encoder_onehot[hevc_amf] |
| 8 | encoder_onehot[av1_amf] |
| 9 | encoder_onehot[h264_qsv] |
| 10 | encoder_onehot[hevc_qsv] |
| 11 | encoder_onehot[av1_qsv] |
| 12 | encoder_onehot[libvvenc] |
| 13 | encoder_onehot[libsvtav1] |
| 14 | encoder_onehot[h264_videotoolbox] |
| 15 | encoder_onehot[hevc_videotoolbox] |
| 16 | preset_norm (0.5 on every row) |
| 17 | crf_norm ((CRF - 19) / 18) |
The two scalar slots are not v2's. The trainer (_build_codec_block() in ai/scripts/train_fr_regressor_v3.py) sets preset_norm to 0.5 on every row and min-max normalises the CRF over its corpus. The shipped checkpoint's corpus (corpus_sha256 58512e6c...) spans CQ 19 to 37: the ensemble trainer records that range for the same file, and both trainers compute it with the same min and max. The sidecar declares the encoding (codec_preset_norm: constant, codec_preset_value: 0.5, codec_crf_norm: minmax, codec_crf_min: 19, codec_crf_max: 37), and libvmaf fills the block that way (ADR-1558): with --tiny-crf 28 the slot is 0.5, --tiny-preset has no effect (the run says so in a warning), and a CRF outside 19..37 gives a value outside [0, 1], as the trainer's formula would. Before 2026-10-04 libvmaf filled v2's ordinal preset / 9 and CRF / 63 here, inputs the model was not trained on.
The encoder vocabulary is closed and order-stable per ADR-0235: the index of each codec is the one-hot column index baked into the trained ONNX.
Differences from v2¶
| Aspect | fr_regressor_v2 | fr_regressor_v3 |
|---|---|---|
| Codec block | (N, 14): 12 encoder slots + 2 | (N, 18): 16 encoder slots + 2 |
| Slot order | ENCODER_VOCAB v2 (libx264, libx265, libsvtav1, ...) | ENCODER_VOCAB_V3 (the 13 slots of the ADR-0291 layout, then libsvtav1, h264_videotoolbox, hevc_videotoolbox appended at 13, 14, 15) |
unknown slot | Yes (slot 11, fallback for novel codecs) | No: the closed 16-slot vocabulary covers every adapter registered under tools/vmaf-tune/src/vmaftune/codec_adapters/ |
| Output name | score | vmaf (matches the teacher-score column of the corpus rows; sidecar output_names: ["vmaf"]) |
| Input name of the codec tensor | codec | codec_block |
The v3 slot order follows the layout documented in ADR-0291 and kept as ENCODER_VOCAB_V3 since PR #401 (ADR-0302 scaffold); it is not the column order of the shipped v2 sidecar (for example libx265 is index 1 in v2 and 2 in v3).
Note
The vmaf CLI validates --tiny-codec against the loaded model's sidecar encoder_vocab, so the v3 names (libaom-av1, h264_amf, hevc_videotoolbox, ...) are accepted with the v3 model even though the --help text lists the v2 names. A name not in the sidecar vocabulary is rejected. Because v3 has no unknown slot, always pass an encoder name that is in its vocabulary. The Python-side convention for callers without codec metadata is slot 0 (libx264), see "Codec-blind fallback".
Output¶
vmaf, shape (N,) — a scalar VMAF-aligned quality score per sample, same MOS range as v1/v2 (typically [0, 100]).
Training corpus¶
Two corpus shapes are accepted, mapped to the same internal feature / codec-block tensors at load time:
vmaf-tunecorpus, schema v3 (preferred, ADR-0366). One row per (source, encoder, preset, crf) encode, canonical-6 means / stddevs computed from libvmaf'spooled_metricsblock:
{"schema_version": 3, "src": "BigBuckBunny_25fps.yuv",
"encoder": "h264_nvenc", "preset": "p4", "crf": 19,
"vmaf_score": 95.86,
"adm2_mean": 0.99, "vif_scale0_mean": 0.88,
"vif_scale1_mean": 0.99, "vif_scale2_mean": 0.996,
"vif_scale3_mean": 0.998, "motion2_mean": 0.0,
"adm2_std": 0.01, "vif_scale0_std": 0.02, ...}
Rows with NaN canonical-6 means (libvmaf did not expose the feature, or the encode failed) are dropped before the StandardScaler is fitted — never imputed to 0.0. Legacy v2 corpora that carry only vmaf_score raise ValueError and point operators at this ADR; they cannot train this regressor.
hw_encoder_corpus.pyper-frame corpus (legacy / NVENC-only).runs/phase_a/full_grid/per_frame_canonical6.jsonl(5,640 rows). One row per frame, bare canonical-6 column names, target columnvmaf, quality knobcq. The training cohort the gate-passing v3 checkpoint was fit on.
NVENC-only corpus caveat¶
Warning
The current Phase A corpus drop is NVENC-only (slot 3, h264_nvenc). The remaining 15 vocabulary slots received zero training examples in this checkpoint.
Consequences for inference at the untrained slots:
- The MLP weights for the 15 unused one-hot columns remain at their Glorot initialisation. The signal for those codecs comes through the canonical-6 features,
preset_normandcrf_norm. - Predictions are degraded but not random: the canonical-6 features alone clear ~0.99 PLCC on the v1 single-input baseline (ADR-0249), so untrained-codec predictions inherit that baseline behaviour modulo the small one-hot column shift.
- The ADR-0235 multi-codec lift floor (at least +0.005 PLCC over v1) is not yet measured: the NVENC-only corpus does not exercise other codecs, so v3's lift over v1 reduces to v1 vs v1 on NVENC.
v3 ships as the production graph regardless: it is forward-compatible with the broader 16-slot schema, and re-using v2 would block multi-codec follow-up corpora. The lift floor will be enforced retroactively when a future Phase A corpus drop covers at least 3 codec families.
This caveat is the dominant reason the live ENCODER_VOCAB_VERSION stays at 2 in train_fr_regressor_v2.py: fr_regressor_v2.onnx remains the production graph for cross-codec inference, and v3 is a parallel checkpoint that wins on NVENC-specific predictions and serves as the schema-flip dry-run.
Training data terms¶
This model was trained on vmaf-tune Phase A encodes of the Netflix Public Dataset's nine references. The terms below are quoted as each source states them (read 2026-10-04); the dataset terms list where each comes from and which models it trained.
Netflix Public Dataset: https://github.com/Netflix/vmaf/blob/0fb4152418d0351901e9c5fd2d30668dced89cdb/resource/doc/datasets.md
We provide a dataset publicly available to the community for training, testing and verification of results purposes.
(please request for access and we will grant it)
The stated purpose includes training; the page states no other terms.
Reading. The fork ships these weights under BSD-2-Clause-Patent: they are fitted parameters that cannot reproduce a clip, an image or a label, and no dataset file is redistributed. That is the fork's reading, not a permission from the dataset's authors. Where a dataset limits its use to research and that limit binds the weights where you use them, treat the model as research-only.
Retrain. RC9 retrains this model on data cleared for redistribution (T-TINY-AI-RETRAIN-CLEARED-DATA-2026-10-04 in state; ADR-1490, ADR-1570).
Codec-blind fallback¶
For inference paths that don't carry codec metadata, pass an all-zeros codec block with encoder_onehot[libx264]=1 (slot 0 is the fork's "default" canonical SW encoder), preset_norm=0.5, crf_norm=0.5. The model degrades to a v1-like estimate; no graph surgery required. Through libvmaf that is --tiny-codec libx264 --tiny-crf 28; the vocabulary has no unknown entry, so libvmaf does not choose an encoder for you.
Training recipe¶
Identical to fr_regressor_v2 and the deep-ensemble LOSO trainer (ADR-0319):
- 9-fold leave-one-source-out (LOSO) over the unique
srcvalues. - Per-fold StandardScaler fit on the training rows only (mirrors
eval_loso_vmaf_tiny_v3.py). FRRegressor(in_features=6, hidden=64, depth=2, dropout=0.1, num_codecs=18).- Adam(
lr=5e-4,weight_decay=1e-5), MSE loss, batch_size=32, 200 epochs. - Final ship checkpoint is fit on the entire corpus (no held-out split) once the LOSO gate passes — the LOSO fold is the gate, not the ship checkpoint.
Headline results¶
Mean LOSO PLCC 0.9975 ± 0.0018 (n = 9 sources). Per-source PLCC:
| Source | PLCC | SROCC | RMSE |
|---|---|---|---|
| BigBuckBunny_25fps | 0.9973 | 0.9878 | 0.787 |
| BirdsInCage_30fps | 0.9988 | 0.9989 | 0.432 |
| CrowdRun_25fps | 0.9996 | 0.9972 | 0.677 |
| ElFuente1_30fps | 0.9987 | 0.8805 | 0.822 |
| ElFuente2_30fps | 0.9950 | 0.9984 | 3.288 |
| FoxBird_25fps | 0.9945 | 0.9329 | 0.904 |
| OldTownCross_25fps | 0.9981 | 0.9951 | 0.810 |
| Seeking_25fps | 0.9989 | 0.9877 | 1.013 |
| Tennis_24fps | 0.9962 | 0.9436 | 1.061 |
Every source clears the relaxed per-source PLCC floor (0.85) from Research-0078 §Retrain ship gate criterion 3, and the mean clears the 0.95 hard floor with ~5 percentage points of margin. The min/max spread (0.9945 → 0.9996) is well under the 0.005 ensemble-spread bound from ADR-0303.
CLI¶
# Production (real Phase A corpus)
python ai/scripts/train_fr_regressor_v3.py \
--corpus runs/phase_a/full_grid/per_frame_canonical6.jsonl
# Smoke (synthetic corpus, validates the pipeline only)
python ai/scripts/train_fr_regressor_v3.py --smoke
The script bakes the full-corpus StandardScaler over the canonical-6 dims into the sidecar JSON (feature_mean / feature_std); the codec block is unscaled. Output ONNX is opset 17, dynamic batch axis, op-allowlist checked. Smoke mode skips the ship gate; real-corpus mode exits non-zero on gate-fail.
The sidecar includes run_provenance (ai-run-provenance-v1) with the trainer entrypoint, parsed arguments, corpus path/hash, and output targets. Smoke runs point at the generated temporary corpus, which makes the sidecar explicit that the output is a pipeline check rather than a real Phase-A training result.
Checkpoint facts¶
| Field | Value |
|---|---|
| Model id | fr_regressor_v3 |
| Location | model/tiny/fr_regressor_v3.onnx |
| Architecture | MLP with 16-slot codec conditioning block |
| Input | features [N, 6], codec_block [N, 18] |
| Output | vmaf [N] |
| ONNX opset | 17 |
| License | BSD-2-Clause-Patent |
| Registry entry | fr_regressor_v3 in model/tiny/registry.json ("smoke": false) |
| SHA-256 | eaa16d23461eda74940b2ed590edfcaf13428aade294e47792a5a15f4d3b999c |
Runnable usage example¶
# Evaluate quality using the vmaf CLI with the v3 codec-aware model:
vmaf \
--reference python/test/resource/yuv/src01_hrc00_576x324.yuv \
--distorted python/test/resource/yuv/src01_hrc01_576x324.yuv \
--width 576 --height 324 --pixel_format 420 --bitdepth 8 \
--tiny-model model/tiny/fr_regressor_v3.onnx \
--tiny-codec libx264 --tiny-preset medium --tiny-crf 28 \
--json --output /tmp/fr_v3.json
The score is attached under the feature name vmaf_tiny_model (the sidecar has no name).
Input features and codec context
Loading the model makes the run compute its input features (adm2, vif_scale0..3, motion2, with default options), and the model scores every frame once the run is flushed. A frame without one of them fails the run with a message naming it; no input is read as 0.0. The model also needs --tiny-codec (with the encode's --tiny-preset and --tiny-crf): without it the run stops on the first frame instead of scoring a guessed codec block (ADR-1520).
Known limitations¶
- Feature dependency: requires the canonical-6 feature set (
adm2,vif_scale0..3,motion2) extracted from 8-bit luma planes. - Closed 16-slot vocabulary: the model has no
unknownslot. ThevmafCLI rejects an encoder name that is not in the sidecar vocabulary; callers that build the codec block themselves use slot 0 (libx264) as their fallback convention. - Execution providers: validated on CPU (
CPUExecutionProvider) and CUDA (CUDAExecutionProvider).
See also¶
- ADR-0323 — this PR's decision record.
- ADR-0302 — the v3 16-slot schema scaffold + ship gate definition.
- ADR-0291 — v2 production-flip; defines the 0.95 LOSO PLCC ship gate v3 reuses.
- ADR-0235 — the parent codec-aware decision; ≥+0.005 PLCC multi-codec lift floor.
- ADR-0319 — LOSO trainer pattern this script reuses.
- Research-0078 — schema expansion plan + retrain checklist.
fr_regressor_v2— v2 model card; v3 is the parallel-shipped successor on the 16-slot vocab.