vmaf-tune corpus — encoder grid sweep and corpus schema¶
vmaf-tune corpus encodes a reference clip over a grid of presets and CRF values with one encoder, scores each encode with vmaf, and writes one JSONL row per cell. The corpus feeds recommend, benchmark and the predictor training. Overview: vmaf-tune.md.
Run a sweep¶
The grid is the Cartesian product of --preset and --crf. This run writes six rows:
vmaf-tune corpus \
--source ref.yuv \
--width 1920 --height 1080 --pix-fmt yuv420p \
--framerate 24 --duration 10 \
--preset medium --preset slow \
--crf 22 --crf 28 --crf 34 \
--output corpus.jsonl
--source is repeatable: pass one flag per source clip. Encodes are written to --encode-dir and deleted after scoring unless --keep-encodes is set. On success the command prints wrote N rows -> corpus.jsonl on stderr and exits 0. Exit code 2 means a missing required flag or an unavailable --score-backend.
Note
--encoder takes one value per run. To sweep several codecs, run the command once per encoder, or use vmaf-tune compare --encoders a,b,c.
Choose another encoder¶
--encoder accepts any of the 19 registered adapters (see codec adapters). The libsvtav1 adapter takes the same x264-style preset names and translates them to SVT-AV1 integer presets internally. AV1 CRF values span 0..63; the informative window is (20, 50):
vmaf-tune corpus \
--source ref.yuv \
--width 1920 --height 1080 --pix-fmt yuv420p \
--framerate 24 --duration 10 \
--encoder libsvtav1 \
--preset medium --preset slow \
--crf 28 --crf 35 --crf 42 \
--output corpus_av1.jsonl
The corpus row records the preset name ("medium"); the FFmpeg command line carries the integer SVT-AV1 expects (-preset 7). The full mapping is on the AV1 codecs page (ADR-0294).
Flags¶
| Flag | Default | Meaning |
|---|---|---|
--source PATH | required | Reference video. Repeatable. |
--width N / --height N | required | Source resolution. |
--pix-fmt PFMT | yuv420p | Forwarded to ffmpeg -pix_fmt. |
--framerate F | 24.0 | Source framerate. |
--duration S | 0.0 | Source duration in seconds, used for the bitrate calculation. |
--encoder NAME | libx264 | Codec adapter; one of the 19 registered adapters. |
--preset P | required | Preset name. Repeatable. Names are encoder specific; see codec adapters. |
--crf N | required | CRF (quality knob) value. Repeatable. Optional with --coarse-to-fine, which picks the CRF axis itself. |
--output PATH | corpus.jsonl | JSONL destination. |
--encode-dir PATH | .workingdir/cache/vmafx-tune/encodes | Scratch directory for encodes; gitignored by convention. |
--keep-encodes | off | Keep encoded files after scoring. |
--vmaf-model NAME | per encode height | Score every row with this model. Without it the model follows the encode height (resolution-aware). |
--neg | off | Score with the NEG variant of the model that applies. |
--ffmpeg-bin PATH | ffmpeg | ffmpeg binary. |
--ffprobe-bin PATH | ffprobe | ffprobe binary, used for HDR detection. |
--vmaf-bin PATH | vmaf | vmaf binary. |
--score-backend NAME | auto | libvmaf scoring backend: auto, cpu, cuda, sycl, hip or metal. See score backend. |
--no-source-hash | off | Skip src_sha256: faster on large YUV files, but loses provenance. |
--two-pass | off | Two-pass encode for adapters with supports_two_pass (libx264, libx265, libvpx-vp9, libaom-av1, libvvenc); other adapters warn on stderr and run single-pass. Doubles encode time. See multi-pass. |
--sample-clip-seconds N | 0.0 | Encode and score only the centre N seconds of each source; 0 uses the full source. See HDR and sampling. |
--coarse-to-fine | off | Coarse-then-fine CRF search instead of the full grid. See coarse-to-fine. |
--coarse-step N | 10 | CRF step of the coarse pass. |
--fine-radius N | 5 | Radius around the best coarse CRF searched in the fine pass. |
--fine-step N | 1 | CRF step of the fine pass. |
--target-vmaf F | none | Target for --coarse-to-fine. |
--auto-hdr | on | Probe each source with ffprobe; inject HDR encoder flags and the HDR-aware scoring path when PQ or HLG signalling is found. |
--force-sdr | off | Treat every source as SDR and skip detection. |
--force-hdr-pq | off | Treat every source as HDR PQ (SMPTE-2084) without probing. Useful for raw YUV, which carries no colour metadata. |
--force-hdr-hlg | off | Treat every source as HDR HLG (ARIB STD-B67) without probing. |
The four HDR flags are mutually exclusive.
Model selection on the CLI
Without --vmaf-model, corpus selects the VMAF model per encode height (resolution-aware); with it, every row is scored with that model. --neg takes the NEG variant of either. The choice is printed on stderr and recorded in each row's vmaf_model.
Hardware encoders
A hardware encoder (NVENC, QSV, AMF, VideoToolbox) is probed with a one-frame encode before the sweep; when the host cannot run it, corpus exits 2 with the probe's reason. See hardware encoders.
Cache
The encode cache has no corpus flags. It is enabled from Python with CorpusOptions(cache_enabled=True, cache_dir=...); see cache.
Corpus JSONL schema¶
Each row is one JSON object on its own line. The full key list is exported as vmaftune.CORPUS_ROW_KEYS and versioned by vmaftune.SCHEMA_VERSION, currently 3. Changing a row's shape means bumping the version in step with the predictor training code.
| Schema | Added |
|---|---|
| v2 | clip_mode, for sample-clip mode (ADR-0301) |
| v3 | HDR provenance triple hdr_transfer / hdr_primaries / hdr_forced, shot statistics, the canonical-6 aggregates and the enc_internal_* columns |
Row identity and source¶
| Key | Type | Description |
|---|---|---|
schema_version | int | Currently 3. |
run_id | str | Per-row UUID4 hex. |
timestamp | str | UTC ISO-8601, seconds precision. |
src | str | Path to the reference. |
src_sha256 | str | SHA-256 of the reference; empty with --no-source-hash. |
width / height | int | Source dimensions. |
pix_fmt | str | Source pixel format. |
framerate | float | Source framerate. |
duration_s | float | Source duration in seconds. |
Encode and score¶
| Key | Type | Description |
|---|---|---|
encoder | str | Codec adapter name, for example libx264. |
encoder_version | str | Detected encoder version, for example libx264-164. |
preset | str | Encoder preset. |
crf | int | Quality knob value. |
extra_params | list[str] | Additional encoder arguments; [] for a plain grid run. |
encode_path | str | Path to the encoded file; empty when not retained. |
encode_size_bytes | int | Encoded file size. |
bitrate_kbps | float | \(\dfrac{8 \cdot \mathrm{encode\_size\_bytes} / 1000}{\mathrm{duration\_s}}\). |
encode_time_ms | float | Wall-clock encode time. |
vmaf_score | float | Pooled-mean VMAF; NaN if scoring was skipped or failed. |
vmaf_model | str | Model version string that scored the row. |
score_time_ms | float | Wall-clock scoring time. |
ffmpeg_version | str | Detected ffmpeg version. |
vmaf_binary_version | str | Detected vmaf binary version. |
exit_status | int | First non-zero of the encode and score exit codes. |
clip_mode | str | "full" (default) or "sample_<N>s" from --sample-clip-seconds. v2+. |
HDR provenance (v3+)¶
| Key | Type | Description |
|---|---|---|
hdr_transfer | str | "" for SDR, "pq" (SMPTE-2084) or "hlg" (ARIB STD-B67). |
hdr_primaries | str | Raw ffprobe color_primaries, for example bt2020; empty for SDR. |
hdr_forced | bool | true when --force-hdr-* or --force-sdr overrode detection. |
Content features (v3+)¶
| Key | Type | Description |
|---|---|---|
shot_count | int | TransNet-V2 shots in the source; 0 when shot detection is unavailable. |
shot_avg_duration_sec | float | Mean shot length in seconds; 0.0 when unavailable. |
shot_duration_std_sec | float | Population standard deviation of shot lengths, a content-class proxy (animation low, live action high). |
adm2_mean | float | Per-frame ADM2 mean (canonical-6); NaN when scoring was skipped (ADR-0366). |
vif_scale0_mean … motion2_std | float | The remaining canonical-6 mean and standard-deviation aggregates, 12 columns in total. |
The twelve canonical-6 columns (adm2_mean, adm2_std, vif_scale0_mean through vif_scale3_std, motion2_mean, motion2_std) come from libvmaf's pooled metrics. When the scoring model omits VIF, including the default vmaf_v1.0.16_3d0h model (ADR-1168, ADR-1169), both the Python vmaf-tune and the Go vmafx-tune pass --feature vif to libvmaf and parse feature aliases with option suffixes, so every canonical-6 column holds a real value instead of NaN.
Encoder-internal statistics (v3+, ADR-0400)¶
| Key | Type | Description |
|---|---|---|
enc_internal_qp_mean | float | Per-frame QP mean from the pass-1 statistics. 0.0 for opt-out adapters. |
enc_internal_qp_std | float | Per-frame QP standard deviation. |
enc_internal_bits_mean | float | Per-frame bit-cost mean (tex+mv+misc). |
enc_internal_bits_std | float | Per-frame bit-cost standard deviation. |
enc_internal_mv_mean | float | Per-frame motion-vector bit-cost mean. |
enc_internal_mv_std | float | Per-frame motion-vector bit-cost standard deviation. |
enc_internal_itex_mean | float | Mean intra-texture cost across I/i frames. |
enc_internal_ptex_mean | float | Mean predicted-texture cost across P/B/b frames. |
enc_internal_intra_ratio | float | Fraction of macroblocks coded as intra. |
enc_internal_skip_ratio | float | Fraction of macroblocks coded as skip. |
Adapters that declare supports_encoder_stats = True populate the ten enc_internal_* columns. Today these are libx264 and libx265; the parser normalises x264 macroblock counters and x265 icu / pcu / scu CTU counters into the same intra, predicted and skip ratio columns.
Hardware encoders (NVENC, AMF, QSV, VideoToolbox) and the AV1 and VVC software encoders (libaom-av1, libsvtav1, libvvenc) opt out and emit 0.0 in every column, so the schema stays uniform across a corpus. The cost for opt-in adapters is encode time: the harness runs a stats-only -pass 1 invocation before the production CRF encode, which roughly doubles per-encode wall time.
Example row¶
{
"schema_version": 3,
"run_id": "0a3b1c8b...",
"timestamp": "2026-05-03T16:00:00+00:00",
"src": "ref.yuv",
"src_sha256": "",
"width": 1920, "height": 1080, "pix_fmt": "yuv420p",
"framerate": 24.0, "duration_s": 10.0,
"encoder": "libx264", "encoder_version": "libx264-164",
"preset": "medium", "crf": 28,
"extra_params": [],
"encode_path": "",
"encode_size_bytes": 845210,
"bitrate_kbps": 676.168,
"encode_time_ms": 4321.0,
"vmaf_score": 92.41,
"vmaf_model": "vmaf_v1.0.16_3d0h",
"score_time_ms": 1820.5,
"ffmpeg_version": "6.1.1",
"vmaf_binary_version": "3.2.1",
"exit_status": 0,
"clip_mode": "full",
"shot_count": 12,
"shot_avg_duration_sec": 0.83,
"shot_duration_std_sec": 0.41,
"adm2_mean": 9.73, "adm2_std": 0.12,
"enc_internal_qp_mean": 25.23,
"enc_internal_qp_std": 0.12,
"enc_internal_bits_mean": 4975.0,
"enc_internal_bits_std": 1820.5,
"enc_internal_mv_mean": 60.3,
"enc_internal_mv_std": 32.1,
"enc_internal_itex_mean": 8000.0,
"enc_internal_ptex_mean": 1500.0,
"enc_internal_intra_ratio": 0.07,
"enc_internal_skip_ratio": 0.16
}
JSON artifact portability¶
Human-facing CLI JSON, report artifacts, the executor result JSONL (tune_results*.jsonl) and local sidecar state files are strict RFC 8259 JSON. Values that are non-finite in memory (NaN, Infinity, -Infinity) are written as null, so notebooks, dashboards, FFmpeg profile consumers and MCP clients can read them with strict JSON decoders.
Corpus JSONL rows are the exception: they are the training interchange format, and their missing-feature semantics are the ones documented in the schema above.