Skip to content

FR regressor v2 — codec-aware (vmaf-tune corpus consumer)

fr_regressor_v2 — codec-conditioned successor to fr_regressor_v1. Maps a 6-D canonical libvmaf feature vector plus a 14-D codec block to a VMAF teacher score. Trained on the JSONL corpus emitted by vmaf-tune corpus (Phase A, ADR-0237).

Status: production checkpoint. model/tiny/registry.json registers fr_regressor_v2.onnx with smoke: false, SHA-256 pin 67934b0b61c73eb852d84ffb34e3333756e8da2530179ecc830336133e63e69e, and an in-sample PLCC of 0.9794 on the vmaf-tune Phase-A JSONL corpus. The old scaffold-only card text is superseded; the follow-up line is now the v3 16-slot vocabulary / LOSO production checkpoint, documented in fr_regressor_v3.md.

Inputs

Two named tensors, dynamic batch axis:

  • features, shape (N, 6) — canonical-6 libvmaf features (StandardScaler-normalised at training time using the mean/std baked into the sidecar JSON):
Index Feature
0 adm2
1 vif_scale0
2 vif_scale1
3 vif_scale2
4 vif_scale3
5 motion2
  • codec, shape (N, 14) — codec block (12 encoder one-hot slots, preset_norm, crf_norm), not normalised (already in [0, 1]):
Index Slot
0 encoder_onehot[libx264]
1 encoder_onehot[libx265]
2 encoder_onehot[libsvtav1]
3 encoder_onehot[libvvenc]
4 encoder_onehot[libvpx-vp9]
5 encoder_onehot[h264_nvenc]
6 encoder_onehot[hevc_nvenc]
7 encoder_onehot[av1_nvenc]
8 encoder_onehot[h264_qsv]
9 encoder_onehot[hevc_qsv]
10 encoder_onehot[av1_qsv]
11 encoder_onehot[unknown]
12 preset_norm (preset ordinal / 9)
13 crf_norm (CRF / 63)

Encoder vocabulary (encoder_vocab_version 2) is closed and ordered — index 0..11 is load-bearing; bumping the vocabulary requires a re-train. The unknown bucket lets corpora without codec metadata pass an all-zeros + unknown=1 vector and degrade gracefully.

CRF normalised by 63 — the union upper bound across all supported encoders (libsvtav1 / libvpx-vp9 max). x264 / x265 use CRF up to 51; values above their per-encoder max are clipped at read time.

Preset ordinal table per encoder lives in ai/scripts/train_fr_regressor_v2.py (PRESET_ORDINAL); the canonical 0..9 scale carries the speed-quality direction consistently across encoders. libsvtav1's numeric 0..13 presets are squashed to 0..9.

Architecture

The shipped graph (read from the ONNX initialisers) is a GELU MLP over the concatenated 20-D input (6 features + 14 codec values): three hidden layers of 32 units, then a single output unit (about 2 820 parameters), the shape ADR-0291 records. The trainer's defaults (--hidden 32 --depth 3 in ai/scripts/train_fr_regressor_v2.py) reproduce it, and the sidecar's training block records hidden and depth. Before 2026-10-04 the defaults were --hidden 16 --depth 2, a smaller model than the shipped one.

Output

score, shape (N,) — a scalar VMAF-aligned quality score per sample, same MOS range as v1 (typically [0, 100]).

Codec-blind fallback

For inference paths that don't carry codec metadata, pass an all-zeros codec vector with encoder_onehot[unknown]=1 and preset_norm=0.5, crf_norm=0.5. The model degrades to a v1-like estimate; no graph surgery required.

Training corpus

vmaf-tune Phase A JSONL (tools/vmaf-tune/src/vmaftune/corpus.py). One row per (source, encoder, preset, crf) cell with schema_version=1. The trainer reads the JSONL row-by-row; the canonical-6 features come from each row's measured libvmaf feature payload when present, with compatibility aliases for historical corpus runs. --smoke remains available for CI/load-path validation, but the committed fr_regressor_v2.onnx is the production export recorded in the registry.

Training data terms

This model was trained on vmaf-tune Phase A encodes of the Netflix Public Dataset's nine references. The terms below are quoted as each source states them (read 2026-10-04); the dataset terms list where each comes from and which models it trained.

Netflix Public Dataset: https://github.com/Netflix/vmaf/blob/0fb4152418d0351901e9c5fd2d30668dced89cdb/resource/doc/datasets.md

We provide a dataset publicly available to the community for training, testing and verification of results purposes.

(please request for access and we will grant it)

The stated purpose includes training; the page states no other terms.

Reading. The fork ships these weights under BSD-2-Clause-Patent: they are fitted parameters that cannot reproduce a clip, an image or a label, and no dataset file is redistributed. That is the fork's reading, not a permission from the dataset's authors. Where a dataset limits its use to research and that limit binds the weights where you use them, treat the model as research-only.

Retrain. RC9 retrains this model on data cleared for redistribution (T-TINY-AI-RETRAIN-CLEARED-DATA-2026-10-04 in state; ADR-1490, ADR-1570).

CLI

# Smoke (synthetic corpus, validates the pipeline only)
python ai/scripts/train_fr_regressor_v2.py --smoke

# Production (real Phase A corpus)
python ai/scripts/train_fr_regressor_v2.py \
    --corpus runs/vmaf_tune_corpus.jsonl \
    --epochs 30

The script bakes the StandardScaler over the canonical-6 dims into the sidecar JSON (feature_mean / feature_std); the codec block is unscaled. Output ONNX is opset 17, dynamic batch axis, op-allowlist checked.

The sidecar and --metrics-out JSON include run_provenance (ai-run-provenance-v1): trainer entrypoint, command arguments, real-corpus path/hash or synthetic-smoke, and output targets. Use this block when comparing refreshed codec-aware runs so stale Phase-A corpora are visible without reverse-engineering shell history.

Checkpoint facts

Field Value
Model id fr_regressor_v2
Location model/tiny/fr_regressor_v2.onnx
Architecture GELU MLP, 3 x 32 hidden units, codec conditioning block
Input features [N, 6], codec [N, 14]
Output score [N]
ONNX opset 17
License BSD-2-Clause-Patent
Registry entry fr_regressor_v2 in model/tiny/registry.json ("smoke": false)
SHA-256 67934b0b61c73eb852d84ffb34e3333756e8da2530179ecc830336133e63e69e

Runnable usage example

# Score reference and distorted video using the codec-aware model:
vmaf \
    --reference python/test/resource/yuv/src01_hrc00_576x324.yuv \
    --distorted python/test/resource/yuv/src01_hrc01_576x324.yuv \
    --width 576 --height 324 --pixel_format 420 --bitdepth 8 \
    --tiny-model model/tiny/fr_regressor_v2.onnx \
    --tiny-codec libx264 --tiny-preset medium --tiny-crf 28 \
    --json --output /tmp/fr_v2.json

The score is attached under the feature name vmaf_tiny_model (the sidecar has no name). --tiny-codec takes an encoder from the sidecar encoder_vocab (see Inputs); --tiny-preset and --tiny-crf set preset_norm and crf_norm.

Input features and codec context

Loading the model makes the run compute its input features (adm2, vif_scale0..3, motion2, with default options), and the model scores every frame once the run is flushed. A frame without one of them fails the run with a message naming it; no input is read as 0.0. The model also needs --tiny-codec (with the encode's --tiny-preset and --tiny-crf): without it the run stops on the first frame instead of scoring a guessed codec block (ADR-1520).

Known limitations

  • Feature dependency: requires the canonical-6 feature set (adm2, vif_scale0..3, motion2) extracted from 8-bit luma planes.
  • Closed encoder vocabulary: supports libx264, libx265, libsvtav1, libvvenc, libvpx-vp9, h264_nvenc, hevc_nvenc, av1_nvenc, h264_qsv, hevc_qsv, av1_qsv, and unknown (fallback slot). The vmaf --tiny-codec option accepts exactly these names (plus common ffprobe aliases such as h264 or hevc).
  • In-sample evaluation only: the only recorded quality figure is the in-sample PLCC 0.9794 (SROCC 0.9640, RMSE 3.01 on 216 rows) in the sidecar's training block; there is no held-out or leave-one-source-out number for v2. The LOSO-gated successor is fr_regressor_v3.
  • Execution providers: validated on CPU (CPUExecutionProvider) and CUDA (CUDAExecutionProvider).
  • External data file: uses external data format; companion fr_regressor_v2.onnx.data must remain co-located.

See also

  • ADR-0272 — original scaffold decision; this card now reflects the promoted production checkpoint.
  • ADR-0235 — the parent codec-aware decision.
  • ADR-0237 — vmaf-tune Phase A (the corpus producer).
  • ADR-0249 — fr_regressor_v1 baseline.
  • Research-0058 — feasibility digest, including the open question on production corpus diversity.
  • fr_regressor_v2_codec_aware.md — superseded ADR-0235-era design card (canonical-9 / FULL_FEATURES path). The shipped model is fr_regressor_v2 (this card), not a separate fr_regressor_v2_codec_aware.onnx.