Models (v1)¶
As of June 2026, this repository ships a new generation of VMAF models, referred to as VMAF v1. These models have demonstrated better accuracy compared to the previous set of models (VMAF v0). Netflix describes them in the tech blog. For the previous generation of models, see overview.md.
The v1 models are ported verbatim from Netflix upstream (commit 4718b4f5f). They live under model/vmaf_v1.0.16/ (and model/vmaf_v1.0.16_hfr/ for the high-frame-rate variants) and are selected via the --model option, in the same way as the v0 models.
vmaf_v1.0.16_3d0h is this fork's default model
Since ADR-1169, scoring without --model uses vmaf_v1.0.16_3d0h, and the 4K ladder uses vmaf_v1.0.16_1d5h_2160. Upstream Netflix still defaults to vmaf_v0.6.1, so the same command gives different numbers on upstream and on this fork: pin --model version=vmaf_v0.6.1 when you need the old values.
There is no NEG model for v1
There is no NEG counterpart to any v1 model. To score with a NEG model, name a v0.6.1 one: --model version=vmaf_v0.6.1neg (or vmaf_4k_v0.6.1neg). The --neg flag belongs to vmaf-tune; see VMAF NEG.
Overview of VMAF v1 models¶
VMAF v1 supports the following models:
- Standard 1080p Model: Calibrated for 1080p video viewed at a standard 3H distance. It uses an operating range of [0, 100].
- Phone Model: Derived by setting the normalized viewing distance to 5H (based on experimental data), this model adjusts the DLM, AIM, and chroma feature calculations to reflect reduced artifact visibility on smaller screens viewed from a greater relative distance. It retains the standard [0, 100] range.
- 4K Model: Two v1 4K models are released:
- A 1.5H variant, based on a discerning 4K@1.5H viewing condition. This variant is conceptually similar to its v0 4K counterpart and operates on a [0, 100] range. For most users, this variant is the default choice.
- A 3H variant, based on a consumer-like 4K@3H viewing condition. This variant operates on a [0, 110] range, which helps to quantify the additional perceptual benefit of 4K resolution over 1080p when both are viewed at 3H.
VMAF v1 should ideally be applied at 10-bit precision for SDR, which helps more accurately capture the presence of banding. Even if the encoded video is 8-bit, VMAF can still be measured at 10 bits by appropriately preprocessing both video inputs.
Which model file should I use?¶
| Scenario | Display | Normalized viewing distance | Model file | Score range |
|---|---|---|---|---|
| Standard 1080p | 1080p | 3H | model/vmaf_v1.0.16/vmaf_v1.0.16_3d0h.json | [0, 100] |
| Phone | 1080p | 5H | model/vmaf_v1.0.16/vmaf_v1.0.16_5d0h.json | [0, 100] |
| 4K default | 2160p | 1.5H | model/vmaf_v1.0.16/vmaf_v1.0.16_1d5h_2160.json | [0, 100] |
| 4K consumer TV | 2160p | 3H | model/vmaf_v1.0.16/vmaf_v1.0.16_3d0h_2160.json | [0, 110] |
Each model is also embedded as a built-in (-Dbuilt_in_models=true, the default), so it can be selected by version= without a file path:
vmaf \
--reference ref.yuv \
--distorted dis.yuv \
--width 1920 --height 1080 --pixel_format 420 --bitdepth 10 \
--model version=vmaf_v1.0.16_3d0h \
--output output.xml
Example invocation using an explicit model path:
vmaf \
--reference ref.yuv \
--distorted dis.yuv \
--width 1920 --height 1080 --pixel_format 420 --bitdepth 10 \
--model path=model/vmaf_v1.0.16/vmaf_v1.0.16_3d0h.json \
--output output.xml
Specifying encode-side parameters¶
If the encode-side width, height, and bit depth of the distorted video are known (i.e., the dimensions and bit depth at which the video was actually encoded, before any rescaling for display), they can be passed to the CAMBI feature used by VMAF as additional --model parameters. For example, if we have a 1280x720 8-bit encoded video that is displayed on a 1080p display, we measure VMAF at a 1920x1080 resolution as follows:
vmaf \
--reference ref.yuv \
--distorted dis.yuv \
--width 1920 --height 1080 --pixel_format 420 --bitdepth 10 \
--model path=model/vmaf_v1.0.16/vmaf_v1.0.16_3d0h.json:cambi.enc_width=1280:cambi.enc_height=720:cambi.enc_bitdepth=8 \
--output output.xml
These overrides are merged into the model's CAMBI feature options, so the CAMBI instance evaluated by VMAF is the one that takes them into account — no separate CAMBI instance is registered. When enc_width/enc_height/enc_bitdepth are not provided, CAMBI falls back to the input width/height and bitdepth.
Model-declared conversion target¶
A model file may carry an optional conversion_target block inside model_dict. It names the colorspace the model's features are defined in, so that libvmaf converts the input pictures to it before scoring (the groundwork for HDR models; no shipped model has one yet):
"conversion_target": {
"colorspace": {"range": "limited", "primaries": "bt2020", "trc": "smpte2084", "matrix": "ictcp"},
"pixel_format": "420",
"bit_depth": 10
}
| Key | Required | Values |
|---|---|---|
colorspace.range | yes | limited, full |
colorspace.primaries | yes | bt709, bt2020, smpte432 |
colorspace.trc | yes | bt709, smpte2084 |
colorspace.matrix | yes | bt709, bt2020nc, ictcp |
pixel_format | no | "420", "422", "444"; absent keeps the source's |
bit_depth | no | integer 8 to 16; absent keeps the source's |
Loading a model rejects (-EINVAL) an unknown name, a missing colour attribute, an unknown value, a bit_depth outside 8 to 16 and a block without colorspace. With a target, vmaf_read_pictures() needs the source colorimetry of both inputs (CLI: input colorimetry; C: vmaf_set_input_colorimetry()), skips the conversion for a picture that already matches, and requires every model of a run to declare the same target or none. See Pictures.
High-frame-rate (HFR) content¶
For high-frame-rate (HFR) content, VMAF v1 provides _hfr model variants under model/vmaf_v1.0.16_hfr/. The model-selection logic above carries over identically — just substitute the _hfr directory and filename.
HFR here means frame rates roughly double the common 24/30 fps cases, such as 60 fps. The _hfr variants are calibrated for the ~50/60 fps regime.
Note
"HFR" can refer to higher frame rates in some applications; for these models it means the ~50/60 fps regime above.
Compared to the standard models, the _hfr variants use a wider, five-frame temporal motion window with moving-average smoothing. The SAD of frame n is taken between frames n-2 and n, and motion2 of frame n is the smaller of the SADs on both sides of it (so it reads frames n-3, n-1 and n+1); see Motion, five-frame window. This reduces the quality under-prediction that v0 showed at higher frame rates without inflating scores from the denser inter-frame signal.
HFR handling remains an area of active improvement: the current variants don't fully capture the perceptual impact of high frame rates, and Netflix expects to refine it in future releases.
Fork status (VMAFX-specific)¶
The four SDR (non-HFR) v1.0.16 models — vmaf_v1.0.16_3d0h, vmaf_v1.0.16_3d0h_2160, vmaf_v1.0.16_5d0h, and vmaf_v1.0.16_1d5h_2160 — are fully supported on the fork's CPU path and score correctly (the 1080p 3H model reproduces the upstream golden VMAF of 82.816060 on the Netflix src01 reference pair; before ADR-1477 the fork printed 82.816059).
The four _hfr variants (vmaf_v1.0.16_hfr_3d0h, vmaf_v1.0.16_hfr_3d0h_2160, vmaf_v1.0.16_hfr_5d0h, vmaf_v1.0.16_hfr_1d5h_2160) score on the CPU since ADR-1478. They set motion_five_frame_window=true and motion_moving_average=true on the motion feature (Motion, five-frame window); the fork returned -ENOTSUP for the first of the two until then. Per frame, every feature and the model score equal upstream Netflix 9e48141b at 17 significant digits on the measured clips, on the scalar, AVX2 and default paths, serial and with worker threads.
build/tools/vmaf \
--reference ref.yuv --distorted dist.yuv \
--width 1920 --height 1080 --pixel_format 420 --bitdepth 8 \
--model version=vmaf_v1.0.16_hfr_3d0h \
--output /dev/stdout
On --backend cuda, sycl and hip an _hfr model runs its motion feature on the device as well: the motion twins compute the five-frame window with the CPU's bits (ADR-1491). On --backend metal the CPU motion extractor computes it and the other features stay on the device. The JSON receipt lists which extractor ran where.
GPU and SIMD acceleration (fork-specific)¶
The v1 models, the _hfr models included, run on the CPU (scalar and AVX2 / AVX-512 / NEON SIMD) and on the CUDA, SYCL, HIP and Metal backends. Only the CPU scalar fixed-point path is the archival reference the Netflix golden-data checkpoints assert against. See overview.md for the backend table and backends for setup.