Features¶
A feature extractor is a per-frame computation that libvmaf runs as part of scoring. Each extractor publishes one or more named metrics into the result report, and VMAF models fuse some of them into the final VMAF score. You can also request an extractor on its own, with no model and no fusion, when all you want is PSNR, SSIM, CIEDE2000 or another single metric. Pass its registered name to --feature.
This page is the index of every extractor. Use it to find the registered name, the backends that accelerate it, whether a device twin is bit-identical to the CPU, and the page that documents it. The option reference of the extractors without a page of their own follows the table.
Per ADR-0100 every user-discoverable extractor ships what / range / invocation / input formats / limitations in the same PR as the code.
Select an extractor¶
Name the extractor after --feature. Options go after the name: the first = ends the name, : separates options, and each option is key=value.
# Single extractor, no model
vmaf --reference ref.yuv --distorted dis.yuv \
--width 1920 --height 1080 --pixel_format 420 --bitdepth 8 \
--no_prediction --feature psnr \
--output score.json --json
# Several extractors
vmaf ... --feature psnr --feature ssim --feature ciede --feature cambi ...
# Per-extractor options
vmaf ... --feature "psnr=enable_mse=true:enable_apsnr=true" ...
vmaf ... --feature "adm=adm_enhn_gain_limit=1.0" ...
- CLI aliases —
integer_motionselectsmotion,integer_motion2selectsmotion_v2,integer_ssimselectsssim,integer_ms_ssimselectsfloat_ms_ssimandinteger_psnrselectspsnr. No otherinteger_*name is an alias: the fixed-point VIF and ADM arevifandadm. - ffmpeg —
libvmaf=feature=name=vif(set throughav_opt_set); options follow the name with:, for examplelibvmaf=feature=name=float_viforlibvmaf=feature=name=motion_v2. - C API —
vmaf_use_feature(ctx, "psnr", opts)orvmaf_use_features_from_model():
VmafFeatureDictionary *opts = NULL;
vmaf_feature_dictionary_set(&opts, "enable_mse", "true");
vmaf_use_feature(ctx, "psnr", opts);
See usage/cli.md for the full CLI grammar and api/index.md for the dictionary ownership rules.
With --backend cuda|sycl|hip|metal, --feature <name> runs that backend's twin of the extractor when one exists. You can also name a twin directly, for example --feature float_adm_cuda. A backend that was not compiled in fails instead of falling back to the CPU (ADR-0498); see Backends for the runtime dispatch rules.
Extractor coverage¶
Names in the first column are the registered names in core/src/feature/feature_extractor.cpp. Twin columns show the registered twin name, or — when none is registered. Exactness comes from the declaration files scripts/ci/exact_twins.d/<feature>.<backend>: exact means the parity gate compares the twin with tolerance 0, bound <value> is a documented tolerance from scripts/ci/cross_backend_calibration.py. The generated list of declarations is cross-backend-exact-twins.md.
| Registered name | Measures | CPU SIMD | CUDA twin | SYCL twin | HIP twin | Metal twin | Exactness of the CUDA / SYCL / HIP twins | Page |
|---|---|---|---|---|---|---|---|---|
vif | Visual information fidelity, four scales (core) | AVX2, AVX-512, NEON | vif_cuda | vif_sycl | vif_hip | integer_vif_metal | cuda exact, sycl exact, hip exact | vif |
float_vif | VIF in float (core) | AVX2 (convolution only) | float_vif_cuda | float_vif_sycl | float_vif_hip | float_vif_metal | cuda exact, sycl exact, hip exact | vif |
motion | Motion2, fixed-point (core) | AVX2, AVX-512, NEON | motion_cuda | motion_sycl | motion_hip | integer_motion_metal | cuda exact, sycl exact, hip exact | motion |
motion_v2 | Pipelined Motion2 | AVX2, AVX-512, NEON | motion_v2_cuda | motion_v2_sycl | motion_v2_hip | motion_v2_metal | cuda exact, sycl exact, hip exact | motion |
float_motion | Motion2 in float (core) | AVX2, AVX-512, NEON | float_motion_cuda | float_motion_sycl | float_motion_hip | float_motion_metal | cuda exact, sycl exact, hip exact | motion |
adm | Additive detail metric, fixed-point (core) | AVX2, AVX-512, NEON | adm_cuda | adm_sycl | adm_hip | integer_adm_metal | cuda exact, sycl exact, hip exact | adm |
float_adm | ADM in float (core) | AVX2, AVX-512, NEON | float_adm_cuda | float_adm_sycl | float_adm_hip | float_adm_metal | cuda exact, sycl exact, hip exact | adm |
cambi | Banding index | AVX2, AVX-512, NEON | cambi_cuda | cambi_sycl | cambi_hip | integer_cambi_metal | cuda exact, sycl exact, hip exact | cambi |
ciede | CIEDE2000 colour difference | AVX2, AVX-512, NEON | ciede_cuda | ciede_sycl | ciede_hip | integer_ciede_metal | cuda bound 1e-9, sycl bound 1e-9, hip bound 1e-9 | ciede |
psnr | PSNR, fixed-point (Y, Cb, Cr) | AVX2, AVX-512, NEON | psnr_cuda | psnr_sycl | psnr_hip | integer_psnr_metal | cuda exact, sycl exact, hip exact | psnr |
float_psnr | PSNR in float (luma) | AVX2, AVX-512, NEON | float_psnr_cuda | float_psnr_sycl | float_psnr_hip | float_psnr_metal | cuda exact, sycl exact, hip exact | psnr |
psnr_hvs | PSNR weighted by a contrast-sensitivity function | AVX2, NEON | psnr_hvs_cuda | psnr_hvs_sycl | psnr_hvs_hip | integer_psnr_hvs_metal | cuda exact, sycl exact, hip exact | psnr-hvs |
ssim | SSIM, fixed-point | AVX2 | integer_ssim_cuda | integer_ssim_sycl | integer_ssim_hip | integer_ssim_metal | cuda exact, sycl exact, hip exact | ssim |
float_ssim | SSIM in float | AVX2, AVX-512, NEON | float_ssim_cuda | float_ssim_sycl | float_ssim_hip | float_ssim_metal | cuda exact, sycl exact, hip exact | ssim |
float_ms_ssim | Multi-scale SSIM | AVX2, AVX-512, NEON | float_ms_ssim_cuda | float_ms_ssim_sycl | integer_ms_ssim_hip | float_ms_ssim_metal | cuda exact, sycl exact, hip exact | ms-ssim |
ssimulacra2 | SSIMULACRA 2 in XYB space | AVX2, AVX-512, NEON, SVE2 | ssimulacra2_cuda | ssimulacra2_sycl | ssimulacra2_hip | ssimulacra2_metal | cuda exact, sycl exact, hip exact | ssimulacra2 |
float_moment | First and second moments | AVX2, AVX-512, NEON, SVE2 | float_moment_cuda | float_moment_sycl | float_moment_hip | float_moment_metal | cuda exact, sycl exact, hip exact | float-moment |
speed_chroma | SpEED chroma (research) | AVX2, AVX-512, NEON | speed_chroma_cuda | speed_chroma_sycl | speed_chroma_hip | — | cuda exact, sycl exact, hip exact | speed_qa |
speed_temporal | SpEED temporal (research) | AVX2, AVX-512, NEON | speed_temporal_cuda | speed_temporal_sycl | speed_temporal_hip | — | cuda exact, sycl exact, hip exact | speed_qa |
speed_qa | SpEED-QA | — | — | — | — | — | — | speed_qa |
niqe | No-reference naturalness (distorted frame only) | — | — | — | — | — | — | niqe |
y_funque_plus | Y-FUNQUE+ atoms (fused SVR deferred) | — | — | — | — | — | — | y-funque-plus |
brisque | No-reference BRISQUE | — | — | — | — | — | — | brisque |
delta_e_itp | Delta E ITP colour difference | — | — | — | — | — | — | delta_e_itp |
pu21 | PU21 perceptually uniform PSNR / SSIM | — | — | — | — | — | — | pu21 |
tad | Temporal absolute difference (Rust pilot, -Denable_rust_features=true) | — | — | — | — | — | — | tad |
lpips | Learned perceptual image patch similarity (ONNX) | — | — | — | — | — | — | tiny-ai-extractors |
dists_sq | DISTS-shaped perceptual distance (ONNX, placeholder weights) | — | — | — | — | — | — | dists |
fastdvdnet_pre | Temporal denoising pre-filter, residual only (ONNX) | — | — | — | — | — | — | tiny-ai-extractors |
transnet_v2 | Shot-boundary detector (ONNX) | — | — | — | — | — | — | tiny-ai-extractors |
mobilesal | Saliency mean (ONNX) | — | — | — | — | — | — | tiny-ai-extractors |
Published metrics¶
| Extractor | Metrics it publishes |
|---|---|
vif | vif_scale0, vif_scale1, vif_scale2, vif_scale3 |
float_vif | float_vif_scale0 … float_vif_scale3 |
motion | motion2 (+ motion with debug=true) |
motion_v2 | VMAF_integer_feature_motion_v2_sad_score, VMAF_integer_feature_motion2_v2_score, VMAF_integer_feature_motion3_v2_score |
float_motion | float_motion2 (+ float_motion with debug=true, the default) |
adm | adm2, adm_scale0 … adm_scale3, aim_score, adm3_score |
float_adm | float_adm2, adm_scale0 … adm_scale3, aim_score, adm3_score |
cambi | cambi |
ciede | ciede2000 |
psnr | psnr_y, psnr_cb, psnr_cr (+ MSE and APSNR when enabled) |
float_psnr | float_psnr (luma only; the CPU extractor emits a single luma score) |
psnr_hvs | psnr_hvs, psnr_hvs_y, psnr_hvs_cb, psnr_hvs_cr |
ssim | ssim |
float_ssim | float_ssim (+ L/C/S when enabled) |
float_ms_ssim | float_ms_ssim (+ per-scale L/C/S when enabled) |
ssimulacra2 | ssimulacra2 |
niqe | niqe (no-reference; scores the distorted frame only) |
y_funque_plus | y_funque_plus_ms_ssim, y_funque_plus_dlm, y_funque_plus_mad (atoms only) |
float_moment | float_moment_ref1st, float_moment_dis1st, float_moment_ref2nd, float_moment_dis2nd |
speed_chroma (--feature speed_chroma) | Speed_chroma_feature_speed_chroma_u_score, ..._v_score, ..._uv_score |
speed_temporal (--feature speed_temporal) | Speed_temporal_feature_speed_temporal_score |
speed_qa | speed_qa |
brisque | brisque (no-reference) |
delta_e_itp | delta_e_itp |
pu21 | pu21_psnr, pu21_ssim |
tad | tad, tad_sad |
lpips | lpips |
dists_sq | dists_sq |
fastdvdnet_pre | fastdvdnet_pre_l1_residual |
transnet_v2 | shot_boundary_probability, shot_boundary |
mobilesal | saliency_mean |
float_ansnr | removed: float_ansnr and float_anpsnr are no longer emitted |
Notes on the table¶
- Core extractors (
vif,float_vif,motion,float_motion,adm,float_adm) are inputs of the shipped VMAF models (see models/overview.md); the others are standalone. - Metal twins are registered when libvmaf is built with Metal. None is declared exact; Metal has no
speed_chroma,speed_temporalorspeed_qatwin (GAP-METAL-MISSING-SPEED-TWINS).ssimulacra2_metalruns a hybrid host/GPU pipeline. Metal parity runs on an Apple device from the macOS tester bundle, not on a hosted runner (ADR-1496). float_psnr,float_adm,float_vif,float_motion,float_moment,speed_chromaandspeed_temporal(and their twins) are registered only when libvmaf is built with-Denable_float=true(VMAF_FLOAT_FEATURES=1), which is the default (core/meson_options.txt).psnrGPU chroma — the twins honourenable_chroma(defaulttrue) and emitpsnr_cb/psnr_crnext topsnr_y. Withenable_chroma=false, onlypsnr_yis emitted, as on the CPU. YUV 4:0:0 sources always produce luma only. GPU chroma parity for CUDA and SYCL came with ADR-0453.float_ssimandfloat_ms_ssim— thessim_accumulate_avx512reduction is vectorised, bit-exact against scalar (ADR-0139, PR #342, about 7 to 11 percent less wall-clock time on the SSIM and MS-SSIM hot path). The fixed-pointssimhas an AVX2 path only.- Tiny-AI extractors run their ONNX graph on the ORT execution provider selected with
--tiny-device(CPU, CUDA, OpenVINO, ROCm); libvmaf has no SIMD or GPU path for their pre- and post-processing. Seedocs/ai/inference.md. - HIP — 19 extractors are registered in
feature_extractor_list[]and resolve throughvmaf_get_feature_extractor_by_name(T7-10b). Seebackends/hip/overview.md. float_vif_hiptakes part in model-driven dispatch through the build optionenable_float_vif_hip_autodispatch, on by default since ADR-2092 (vif).nullis a registered no-op extractor used by tests.float_ansnrwas removed (PR #38, ADR-0865):--feature float_ansnrfails with feature-not-found on a current build. See ANSNR.
Depending on your build configuration not every backend is available; see backends/.
Non-finite result handling¶
A failed extractor computation never becomes a plausible metric. If VIF, ADM, SSIM, MS-SSIM, SSIMULACRA2, TransNet V2 or the final VMAF piecewise mapping produces NaN or an unexpected infinity, libvmaf fails that frame with -EINVAL. It does not substitute a minimum, a maximum, a threshold result or zero.
What this means for you:
- The frame has no score. The CLI treats the negative return as a runtime error and prints no fabricated value; C API callers handle the existing negative-errno contract (API error semantics). The log callback names the extractor, the frame and the offending value where known.
- Hidden values are checked too. CPU, CUDA, HIP, SYCL and Metal host paths validate the complete enabled score set before their first collector write. That covers VIF debug numerators and denominators, ADM reductions before their precision floor, and every MS-SSIM L/C/S atom even when
enable_lcs=false. - All four VIF ratios must be finite before any scale is published (scale 0 stays unclamped, scales 1 to 3 then apply their minimum).
- The same complete-set rule holds for float and integer VIF and for SSIMULACRA2 on scalar, SIMD, CUDA, HIP, SYCL and Metal.
- Finite results are unchanged, including ADM's defined perfect aggregate flat-frame result, a finite ADM per-scale
0/0(now an explicit1.0) and scores that legitimately reach a configured clamp.
One output is non-finite on purpose: with dB output enabled and clip_db=false, a finite perfect SSIM or MS-SSIM raw score reports positive infinity (ADR-1221). Set clip_db=true to cap it at max_db. NaN raw scores and invalid dB ceilings still fail the frame.
Note
Those substitutions used to look legitimate: an invalid SSIMULACRA2 result appeared as the perfect 100.0, an invalid SSIM as max_db, and an undefined ADM AIM ratio as the perfect 1.0.
Per-feature GPU dispatch hints (T7-26 / ADR-0181)¶
Each feature carries a small VmafFeatureCharacteristics descriptor that drives the per-backend dispatch decision: graph replay or direct submit on SYCL, graph capture or streams on CUDA. The per-backend dispatch_strategy modules consume it (core/src/{cuda,sycl}/dispatch_strategy.{c,h}). The defaults match the behaviour before T7-26 byte for byte: graph replay above 720p area, direct submit below.
Two environment variables override the default. Each takes a comma-separated list of feature:strategy pairs and wins over the registry default for the named features:
| Env var | Strategy values | Effect |
|---|---|---|
VMAF_SYCL_DISPATCH | graph / direct | Per-feature SYCL graph-replay override. |
VMAF_CUDA_DISPATCH | graph / direct | Per-extractor CUDA graph-capture override, keyed by the CUDA extractor name (vif_cuda:graph) and read when the extractor initialises. Graph capture is not implemented: graph logs a warning and runs direct (env-var reference). |
Examples:
# Force ADM to direct submit on SYCL (default below 720p, override above):
VMAF_SYCL_DISPATCH=adm:direct vmaf [...] --feature adm_sycl --backend sycl
# Mix per-feature strategies:
VMAF_SYCL_DISPATCH=vif:graph,motion:direct,adm:graph vmaf [...]
The legacy global knobs VMAF_SYCL_USE_GRAPH=1 and VMAF_SYCL_NO_GRAPH=1 remain as aliases that force every feature to graph or to direct; the per-feature VMAF_SYCL_DISPATCH takes precedence.
Option reference¶
Extractors with their own page link to it. The sections below give the options of the rest. Each option is a feature parameter, so a non-default value appends its alias to the published feature name.
VIF — Visual Information Fidelity¶
VIF measures information-fidelity loss between reference and distorted at four Gaussian-pyramid scales. In the original Sheikh/Bovik formulation the scales are combined into a single score; in VMAF each scale is kept as a separate feature so the model can learn per-scale weights. See vif for the twin measurements and the minimum frame size.
- Run —
--feature vif(fixed-point, default) or--feature float_vif; ffmpeglibvmaf=feature=name=vif; C APIvmaf_use_feature(ctx, "vif", opts). - Output metrics —
vif_scale0,vif_scale1,vif_scale2,vif_scale3: per-scale fidelity ratios in[0, 1]. Higher is better (1 = reference-identical). Withdebug=truealsovif,vif_num,vif_den, and per-scale*_num/*_den. - Input formats — YUV 4:2:0 / 4:2:2 / 4:4:4 / 4:0:0, 8 / 10 / 12 / 16 bpc. Y plane only.
- Backends —
vif: AVX2, AVX-512, NEON, CUDA, SYCL, HIP, Metal.float_vif: AVX2 convolution, CUDA, SYCL, HIP, Metal.
| Option | Alias | Declared by | Type | Default | Range | Effect |
|---|---|---|---|---|---|---|
debug | — | both | bool | false | — | Emit vif, vif_num, vif_den and the per-scale numerator and denominator. |
vif_enhn_gain_limit | egl | both | double | 100.0 | 1.0–100.0 | Cap on the enhancement-gain ratio, so over-sharpened output cannot saturate. 1.0 disables the enhancement-gain path (matches pre-v1.3 behaviour). |
vif_skip_scale0 | ssclz | both | bool | false | — | Skip the finest scale: exclude it from the aggregate and report it as zero (below). |
vif_kernelscale | ks | float_vif | double | 1.0 | 0.1–4.0 | Scale the Gaussian kernel standard deviation. |
vif_prescale | ps | float_vif | double | 1.0 | 0.1–4.0 | Resize factor applied to the frame before VIF. |
vif_prescale_method | pm | float_vif | string | nearest | — | Resize method for the prescale. |
vif_scale1_min_val … vif_scale3_min_val | s1miv … s3miv | float_vif | double | 0.0 | 0.0–1.0 | Minimum value of scale 1 … 3; smaller values are set to it. |
vif_sigma_nsq | snsq | float_vif | double | 2.0 | 0.0–5.0 | Neural-noise variance. |
vif_skip_scale0 and what it publishes¶
With vif_skip_scale0=true the finest VIF scale is dropped from the aggregate integer_vif score and its own score is reported as exactly 0.0 rather than as a computed ratio. The two are separate effects, and a backend has to do both:
| Key | Value when skipping |
|---|---|
integer_vif | sum over scales 1–3 only |
VMAF_integer_feature_vif_scale0_score | 0.0 |
integer_vif_num_scale0 / ..._den_scale0 (debug only) | 0.0 / -1.0 |
Because the option is a FEATURE_PARAM, setting it changes the published feature names: the alias ssclz is appended, so the scale-0 score is filed under integer_vif_scale0_ssclz. Read that key, not the default one. Where a GPU twin declares one of these options, its alias must match the CPU extractor so an equivalent configuration publishes the same key (ADR-1312).
The CPU reference simply never computes scale 0. The GPU twins do compute all four scales and apply the skip when they publish, so the zeroing lives at the emission site rather than in the kernel. integer_vif honours this on CPU, CUDA, SYCL, HIP and Metal.
Reference — Sheikh H. R., Bovik A. C., "Image information and visual quality," IEEE TIP 15(2):430–444, 2006.
Motion¶
Motion measures temporal activity: both frames are blurred with a fixed low-pass filter and the mean absolute pixel difference between the current and the previous reference luma is taken. motion2 is the improved version with proper padding and boundary handling; the unfixed motion is kept behind debug=true for back-compat. See the Motion page for output ranges, the five-frame window and the per-variant backend matrix.
- Run —
--feature motion(fixed-point),--feature float_motionor--feature motion_v2; ffmpeglibvmaf=feature=name=motion(orname=motion_v2); C APIvmaf_use_feature(ctx, "motion", opts)(formotion_v2,vmaf_use_feature(ctx, "motion_v2", NULL)). - Output range —
[0, ∞)before themotion_max_valclamp (default 10 000). Zero for a frozen reference; grows with motion content. - Input formats — YUV 4:2:0 / 4:2:2 / 4:4:4, 8 / 10 / 12 / 16 bpc. Y plane only.
- Limitations — temporal. The extractor carries state across frames (two previous blurred references) and has a flush callback that emits the final frame's score after the input stream ends. Single-frame scoring is not supported; Motion2 on frame 0 is defined as
0.0.
Output metrics¶
| Extractor | Metrics |
|---|---|
motion | motion2 (the shipped feature), motion3, the SAD score (CPU, CUDA, SYCL and HIP; see Motion), and the legacy motion with debug=true |
float_motion | float_motion2, float_motion3, and float_motion (the legacy variant, emitted by default because debug defaults to true) |
motion_v2 | VMAF_integer_feature_motion_v2_sad_score, VMAF_integer_feature_motion2_v2_score, VMAF_integer_feature_motion3_v2_score |
Options of motion and float_motion¶
| Option | Alias | motion | float_motion | Type | Default | Range | Effect |
|---|---|---|---|---|---|---|---|
debug | — | yes (default false) | yes (default true) | bool | see column | — | Emit the legacy motion alongside motion2. |
motion_force_zero | force_0 | yes | yes | bool | false | — | Override all scores to 0.0, for deterministic test fixtures. |
motion_fps_weight | mfw | yes | yes | double | 1.0 | 0.0–5.0 | Multiplicative FPS-aware correction applied before clamping. |
motion_blend_factor | mbf | yes | yes | double | 1.0 | 0.0–1.0 | Blend factor of motion3. |
motion_blend_offset | mbo | yes | yes | double | 40.0 | 0.0–1000.0 | Score offset at which blending of motion3 begins. |
motion_max_val | mmxv | yes | yes | double | 10000.0 | 0.0–10000.0 | Upper clamp of the emitted scores. |
motion_five_frame_window | mffw | yes | — | bool | false | — | Take each SAD against the frame two back instead of the previous one. |
motion_moving_average | mma | yes | — | bool | false | — | Two-frame moving average of motion3. |
motion_add_scale1 | mdc | — | yes | bool | false | — | Add a half-resolution bilinear-downsampled SAD term on top of the full-resolution SAD. |
motion_add_uv | mau | — | yes | bool | false | — | Add the U and V plane SADs to the Y SAD. |
motion_filter_size | mfs | — | yes | int | 5 | 0–9 | Blur kernel size; 5 is the original Motion2 filter, 3 a cheaper variant. |
force_0 is the collector-key suffix on the CPU, CUDA, SYCL and HIP motion twins; backend selection does not change the published key.
Options of motion_v2¶
motion_v2 declares seven options with the same aliases, defaults and ranges as the motion rows above: motion_force_zero, motion_blend_factor, motion_blend_offset, motion_fps_weight, motion_max_val, motion_five_frame_window and motion_moving_average. It has no debug option.
Motion v2 — pipelined Motion2¶
motion_v2 exploits the linearity of the blur kernel. Instead of storing the blurred previous reference across frames, it folds the frame difference, blur and absolute sum into one row-at-a-time pipeline that needs a single scratch row. The score equals Motion2 apart from the SAD-against-sum semantics described below. It is a separate extractor so callers can opt into the pipelined arithmetic without touching the legacy motion registry entry.
- Output metrics —
VMAF_integer_feature_motion_v2_sad_scoreis the per-frame sum of absolute blurred differences (frame 0 emits0.0).VMAF_integer_feature_motion2_v2_scoreis the Motion2-equivalent score (current frame plus the next frame's score, divided by 2, matching the legacy temporal smoothing).VMAF_integer_feature_motion3_v2_scoreis the perceptually blended and clipped score, optionally moving-averaged (Motion). - Output range —
[0, ∞), in the units of Motion2. - Limitations — temporal. The extractor caches its own previous reference in a GPU-side ping-pong (the framework's
prev_refslot is unused on the GPU paths). Frame 0 is0.0, and the final frame's smoothed score is emitted by the flush callback, as formotion.
Motion backends¶
- CPU — AVX2, AVX-512 and NEON for
motion,float_motionandmotion_v2. - GPU, three-frame window — all GPU backends emit
motion,motion2andmotion3(the latter since T3-15(c) / ADR-0219). - GPU, five-frame window (
motion_five_frame_window=true, motion.md) — runs on the CPU and on the CUDA, SYCL and HIP twins ofmotionandmotion_v2, bit-identical to the CPU (ADR-1491). On Metal the CPU extractor computes it. motion_add_uv,motion_add_scale1,motion_filter_size— implemented on the CPU and byfloat_motion_hip(ADR-1404, see motion.md); the CUDA, SYCL and Metalfloat_motiontwins do not declare them. See backends/cuda/overview.md §Known gaps and backends/sycl/overview.md §Known gaps.float_motiontwins —float_motion_cudaandfloat_motion_syclreturn the CPU scores bit for bit at--precision max(ADR-1409, ADR-1411), and so doesfloat_motion_hip(ADR-1419). The Metalfloat_motiontwin has not been measured on a device.motion_v2twins — bit-exact against the CPU scalar reference on 8-bit and 10-bit inputs (max_abs_diff = 0.0across the cross-backend gate fixture), forinteger_motion_v2_cuda.c,integer_motion_v2_sycl.cppandinteger_motion_v2_hip.c. They share a design:- one dispatch over
(prev_ref - cur_ref)exploits convolution linearity and skips the per-frame blurred-state buffer; - a raw-pixel ping-pong of two private device buffers caches the previous frame's Y plane;
- per-workgroup
int64SAD partials reduce on the host; - \(\mathrm{motion2\_v2\_score} = \min(\mathrm{score}_i, \mathrm{score}_{i+1})\) is emitted in
flush(). motion_v2padding — mirror padding diverges from the correspondingmotion_*kernels by one pixel at the boundary. CPUinteger_motion_v2.cuses reflect-101 mirror2*size - idx - 2(ADR-0662 corrected stale GPU-side prose that had documented-1).
ADM — Additive Detail Metric¶
ADM measures detail loss and additive impairment at four wavelet sub-band scales. Its options, outputs, backend measurements and edge-case behaviour are on the ADM page: run it with --feature adm (fixed-point) or --feature float_adm.
AIM above 1¶
integer_aim of the fixed-point adm can exceed 1, while float_adm's aim is clipped at 1; both match upstream Netflix/vmaf (ADR-1417). The full explanation and the numbers are in AIM above 1.
CAMBI — Contrast-Aware Multiscale Banding Index¶
CAMBI has parameters enough for its own reference: see cambi.
- Run —
--feature cambi; options such asfull_refgo after the name (--feature cambi=full_ref=true). - Output —
cambiin[0, ∞); 0 = no banding, larger = more visible banding. Typical "bad" content sits in1–10. - Backends — CPU (AVX2, AVX-512, NEON) and CUDA, SYCL, HIP and Metal twins. The SYCL, CUDA and HIP twins run every stage on the device (ADR-1357, ADR-1379, ADR-1378).
CIEDE2000 — colour-difference metric¶
--feature ciede converts both frames to CIELAB and averages the CIEDE2000 ΔE per pixel. Output ciede2000 in [0, ~100] (smaller is better), no options, no 4:0:0 input. See ciede for the twin measurements and the agreement with Netflix's source.
PSNR¶
Peak Signal-to-Noise Ratio on each colour plane. The fixed-point psnr path is the default; the float_psnr path is kept for parity with upstream consumers of the float pipeline. See PSNR for the full comparison.
- Run —
--feature psnror--feature float_psnr; ffmpeglibvmaf=feature=name=psnr. - Output metrics (fixed) —
psnr_y,psnr_cb,psnr_cr. Withenable_mse=truealsomse_y/cb/cr. Withenable_apsnr=truealsoapsnr_y/cb/cr(aggregate across the whole clip, emitted at flush). The clip's squared error is summed in 128 bits, so a clip of any length gives the exact aggregate; a 64-bit sum wrapped at frame 33 of 16K at 16 bits with every sample at the maximum difference.float_psnremitsfloat_psnr, luma only. - Output range — dB, saturated at
6 × bpc + 12: 60 dB for 8 bpc, 72 dB for 10 bpc, 84 dB for 12 bpc, 108 dB for 16 bpc. That value is reported both when the two planes are identical (MSE=0, where the true PSNR is+inf) and, by default, when a computed score exceeds it. Setuncapped=trueto report the true value while keeping the MSE=0 sentinel (ADR-1193), ormin_sseto raise both at once. - Input formats — YUV 4:2:0 / 4:2:2 / 4:4:4 / 4:0:0, 8 / 10 / 12 / 16 bpc.
- Limitations — temporal flag set only because of
apsnraccumulation; per-frame PSNR itself is stateless.
| Option | Type | Default | Effect |
|---|---|---|---|
enable_chroma | bool | true | Include psnr_cb / psnr_cr; set false for luma-only. |
enable_mse | bool | false | Emit mse_y/cb/cr alongside PSNR. |
enable_apsnr | bool | false | Emit clip-aggregate apsnr_y/cb/cr at flush. |
reduced_hbd_peak | bool | false | Scale the high-bit-depth peak to match 8-bit content. |
min_sse | double | 0.0 | Clamp the minimum MSE, so both the PSNR ceiling and the identical-frame sentinel. |
uncapped | bool | false | Report the true PSNR instead of truncating at the ceiling; the MSE=0 sentinel is unaffected. |
float_psnr declares uncapped only.
How the twins treat the options:
- The GPU twins honour
enable_chromaand emitpsnr_cb/psnr_cridentically to the CPU path; passenable_chroma=falsefor luma-only operation on any backend. uncappedis mirrored on every GPU twin under the same name and default.psnr_syclandpsnr_cudaalso implementenable_mse,enable_apsnr,reduced_hbd_peakandmin_sse, bit-exact with the CPU.psnr_hipdoes so through the CPU's own helpers and matched the CPU bit for bit on a gfx1036 (ADR-1382).- On Metal those four options keep
psnron the CPU (see PSNR). float_psnradds CUDA, SYCL, HIP and Metal twins on the float pipeline and acceptsuncappedon all of them.
PSNR-HVS¶
PSNR weighted by a human-visual-system contrast-sensitivity function applied in the DCT domain. It correlates better with subjective quality than plain PSNR on blocking-style distortions. See psnr-hvs.
- Run —
--feature psnr_hvs. - Output metrics —
psnr_hvs,psnr_hvs_y,psnr_hvs_cb,psnr_hvs_cr. - Output range — dB, typically
20–60. - Input formats — 8 to 12 bpc, YUV 4:0:0 / 4:2:0 / 4:2:2 / 4:4:4. Deeper input is rejected at init.
- Options —
enable_chroma(defaulttrue). - Backends — scalar (Xiph reference), AVX2 (ADR-0159) and NEON on aarch64 (ADR-0160). The 8×8 integer DCT block is vectorised eight rows in parallel (butterfly, transpose, butterfly, transpose); the float accumulators stay scalar to keep byte-identity with the reference. Verified bit-identical to scalar on the three Netflix golden pairs, with about 3.58x DCT microbenchmark speedup on AVX2. The GPU twins
psnr_hvs_cuda,psnr_hvs_syclandpsnr_hvs_hipare bit-identical to the CPU (ADR-1397, ADR-1401; see agreement with the CPU extractor).
SSIM / MS-SSIM¶
Structural Similarity Index on luma. MS-SSIM extends SSIM to five Gaussian-pyramid scales and fuses them with the Wang 2003 weights. See ssim and ms-ssim for the precision of each twin.
- Run —
--feature ssim(fixed-point),--feature float_ssimor--feature float_ms_ssim; ffmpeglibvmaf=feature=name=ssim. - Output metrics —
ssim(one scalar in[0, 1]);float_ssim(scalar in[0, 1], withenable_lcs=truealsofloat_ssim_l,float_ssim_c,float_ssim_s);float_ms_ssim(scalar in[0, 1], withenable_lcs=truethe per-scale triplesfloat_ms_ssim_{l,c,s}_scale{0..4}, withenable_chroma=truealsofloat_ms_ssim_cbandfloat_ms_ssim_cr). - Output range —
[0, 1], higher is better. Withenable_db=truethe score is \(-10 \log_{10}(1 - \mathrm{score})\);clip_db=truecaps the infinite value of identical frames. - Input formats — YUV 4:2:0 / 4:2:2 / 4:4:4, 8 / 10 / 12 / 16 bpc.
Options¶
| Option | ssim | float_ssim | float_ms_ssim | Type | Default | Range | Effect |
|---|---|---|---|---|---|---|---|
enable_lcs | — | yes | yes | bool | false | — | Emit the L / C / S components (per scale for MS-SSIM). |
enable_db | yes | yes | yes | bool | false | — | Report \(-10 \log_{10}(1 - \mathrm{score})\) instead of the raw ratio. |
clip_db | yes | yes | yes | bool | false | — | Cap dB values based on the minimum representable MSE. |
scale | — | yes | — | int | 0 | 0–10 | Decimation factor of float_ssim: 0 = auto per Wang 2003, 1 = none, 2–10 explicit. |
enable_chroma | — | — | yes | bool | false | — | Score the chroma planes too. |
Minimum dimensions¶
float_ms_ssimneeds at least 176×176 luma. The five Gaussian-pyramid scales force a2⁴ = 16× downsample on the smallest level; smaller inputs (for example QCIF) make the decimate kernel produce undefined output, so init rejects them with-EINVALand a log message (ADR-0153).ssimandfloat_ssimhave no such constraint, butfloat_ssimscores windows of 11×11 samples. A luma plane with fewer than 11 samples in a direction (after thescaledecimation) has no window, and the score is0(the sum of no terms over(w − 10) · (h − 10), as in Netflix's libvmaf) on the scalar and every SIMD path.- With exactly 10 samples in a direction the divisor is 0, the mean is \(0 / 0\) and the frame fails with a non-finite-score error (ADR-1302).
- The GPU twins do not run such a plane:
--backend <gpu> --feature float_ssimcomputes it on the CPU and says so, and--feature float_ssim_cuda(or_sycl,_hip) fails at init (ADR-1324).
Twins and options¶
| Twin | Options it implements | Arithmetic against the CPU |
|---|---|---|
integer_ssim_cuda | enable_db, clip_db (ADR-1373) | Bit for bit (ADR-1424). |
integer_ssim_sycl | enable_db, clip_db (ADR-1365) | Bit for bit (ADR-1443). |
integer_ssim_hip | enable_db, clip_db (ADR-1382) | Declared exact (exact_twins.d/ssim.hip). |
integer_ssim_metal | see ssim | Precision stated on the SSIM page. |
float_ssim_cuda | enable_lcs, enable_db, clip_db, scale; accepts and ignores the enable_chroma the CPU float_ssim never had | Equal to the CPU on every measured frame (ADR-1399, ADR-1464). |
float_ssim_sycl | enable_lcs, enable_db, clip_db, scale | Bit for bit: decimation on the device (ADR-1370), terms added in the CPU's order (ADR-1463). |
float_ssim_hip | enable_db, clip_db, enable_lcs | CPU's own l * c * s arithmetic (ADR-1405, ADR-1382). |
float_ssim_metal | scale 1 only; larger scales run on the CPU extractor | — |
float_ms_ssim_cuda | all MS-SSIM options | Bit-identical at --precision max, per-scale enable_lcs outputs included (ADR-1403, ADR-1465). |
integer_ms_ssim_hip (registered for float_ms_ssim) | all MS-SSIM options | Same arithmetic as the CUDA twin, bit-identical on a gfx1036 (HIP backend). |
float_ms_ssim_sycl | all MS-SSIM options | Bit-identical without fp64 on the device (ADR-1414); see MS-SSIM. |
float_ms_ssim_metal | all MS-SSIM options | Within the cross-backend tolerance of 5e-5. |
Notes on the twins:
- The
enable_lcsoption ships across all MS-SSIM backends: CPU, CUDA, SYCL, HIP and Metal emit the same 15float_ms_ssim_{l,c,s}_scale{0..4}metrics on top of the combined score (T7-35 / ADR-0243). - On SYCL and CUDA, identical frames report the CPU's
+inforclip_dbceiling on the device too.float_ssim_cudacomputes the CPU's per-pixell * c * sand fp32 frame mean, so identical frames report the CPU's value (+inf, or 72.247 dB for flat frames).float_ssim_hipreports the same finite 72.247 dB the CPU gives some identical frames. float_ssim_sycl,float_ssim_cudaandfloat_ssim_hipdecimate on the device at the automatic scale and at every explicit one, so 1080p and 4Kfloat_ssimrun on those backends.- The
float_ssim_cudahost sum costs about 1 ns per scored window and thefloat_ms_ssim_cudaone about 2.1 ns per scored window. The SYCL host sum costs time only where the scale is 1 (SYCL backend).
MS-SSIM decimate (fork-local)¶
The 9-tap 9/7 biorthogonal wavelet low-pass filter that produces scales 1–4 runs through ms_ssim_decimate in core/src/feature/ms_ssim_decimate.c. SIMD variants live in core/src/feature/x86/ms_ssim_decimate_avx2.c (8-wide), core/src/feature/x86/ms_ssim_decimate_avx512.c (16-wide) and core/src/feature/arm64/ms_ssim_decimate_neon.c (4-wide). Dispatch prefers AVX-512 over AVX2 over scalar on x86 and NEON over scalar on aarch64 at runtime through vmaf_get_cpu_flags(). All four paths are strictly byte-identical (per-lane fmaf / _mm{256,512}_fmadd_ps / vfmaq_n_f32 with broadcast coefficients and scalar-fallback borders); core/test/test_ms_ssim_decimate.c verifies it across 1x1, 8x8, 9x9, border-edge and 1920x1080 cases (ADR-0125).
ANSNR — Adjusted Noise SNR¶
Note
Removed. The CPU implementation and all GPU twins (CUDA, SYCL, HIP, Metal) were removed in commit 70ed8b3ce3 (PR #38); the Vulkan backend and its kernel source were removed in ADR-0726. No float_ansnr source remains in the tree, and --feature float_ansnr returns a feature-not-found error on any current build. The historical GPU kernel design (ADR-0194) is kept for reference only.
ANSNR was SNR after a noise-shaping Wiener filter: historical VMAF input that no shipped model consumed. See ansnr.
SSIMULACRA 2 — perceptual similarity in XYB space¶
--feature ssimulacra2 scores one scalar per frame in [0, 100] (higher is better; identical frames return exactly 100). It is a port of the libjxl reference metric. See ssimulacra2 for the score bands, the options and the twin measurements.
- Input formats — YUV 4:2:0 / 4:2:2 / 4:4:4, 8 / 10 / 12 bpc. Chroma is nearest-neighbour upsampled to luma resolution. Minimum input 8×8.
- Options —
yuv_matrix(int, default0, range0–3):0BT.709 limited,1BT.601 limited,2BT.709 full,3BT.601 full. - CPU backends — AVX2, AVX-512, NEON and SVE2, bit-identical to scalar.
- Twins —
ssimulacra2_cuda,ssimulacra2_syclandssimulacra2_hipare device-resident and return the CPU's score bit for bit (exact_twins.d/ssimulacra2.{cuda,sycl,hip}; ADR-1391, ADR-1433, ADR-1363, ADR-1446, ADR-1445).ssimulacra2_metalruns a hybrid host/GPU pipeline. Name a twin with--feature ssimulacra2_cuda(or_sycl,_hip) and pair it with the matching--backendflag for exclusive GPU dispatch. - Limitations —
create_recursive_gaussianderives its coefficients with Cramer's rule in doubles, which yields the samen2/d1floats as libjxl'sInv3x3Matrixfor σ=1.5 at 10-decimal precision but is not guaranteed bit-exact at every σ; the fork pins σ=1.5, libjxl'skSigma.
The SIMD and determinism history is under History.
Float moment — first and second statistical moments¶
--feature float_moment emits float_moment_ref1st, float_moment_dis1st, float_moment_ref2nd and float_moment_dis2nd, has no options and runs on all four backends. See float-moment.
Tiny-AI extractors¶
lpips, dists_sq, fastdvdnet_pre, mobilesal and transnet_v2 run an ONNX model through ONNX Runtime and take one option, model_path. All five are on tiny-ai-extractors; the summaries below keep the headings that other pages link to.
LPIPS — learned perceptual image patch similarity¶
--feature lpips=model_path=/path/to/lpips.onnx emits lpips (lower is more similar). See tiny-ai-extractors.
DISTS-Sq — deep image structure and texture similarity¶
--feature dists_sq=model_path=/path/to/dists_sq.onnx emits dists_sq; the shipped checkpoint is a smoke placeholder. See dists and tiny-ai-extractors.
FastDVDnet pre — temporal denoising pre-filter¶
--feature fastdvdnet_pre=model_path=... emits the diagnostic fastdvdnet_pre_l1_residual, not a quality score. See tiny-ai-extractors.
mobilesal — MobileSal saliency map (tiny-AI, NR / single-input)¶
--feature mobilesal=model_path=... emits saliency_mean for the distorted frame, 8-bit input only. See tiny-ai-extractors.
transnet_v2 — TransNet V2 shot-boundary detector (tiny-AI, NR / single-input)¶
--feature transnet_v2=model_path=... emits shot_boundary_probability and shot_boundary over a 100-frame window. See tiny-ai-extractors.
Speed (chroma + temporal) — Netflix research extractors¶
speed_chroma and speed_temporal are research-stage extractors ported from Netflix upstream commit d3647c73, with its dependency 4ad6e0ea for the vif_tools helpers. They share a spatial-pooling backbone in the style of SpEED-QA and a per-frame neural-network-shaped weighting. They register when libvmaf is built with -Denable_float=true, the default. See also speed_qa.
Their scores are Netflix's: every value compared with a build of Netflix master is identical, and the CUDA, HIP and SYCL twins return the CPU's scores bit for bit (ADR-1477; the comparison, the options where the fork differs on purpose and what changed on 2026-10-02 are on the SpEED page).
| Extractor | Output metrics | Input formats | Backends |
|---|---|---|---|
speed_chroma | Speed_chroma_feature_speed_chroma_u_score, ..._v_score, ..._uv_score (no y score) | YUV 4:2:0 / 4:2:2 / 4:4:4, 8 / 10 / 12 / 16 bpc | CPU (scalar, AVX2, AVX-512, NEON); CUDA, SYCL, HIP; Metal missing (GAP-METAL-MISSING-SPEED-TWINS) |
speed_temporal | Speed_temporal_feature_speed_temporal_score | same | same |
speed_chroma scores the U and V planes of the input pictures, lifted with picture_copy(..., channel) from the same upstream commit, plus a combined U+V score. speed_temporal scores a small cyclic frame buffer and captures distortions that appear only across consecutive frames, such as flicker and judder.
Speed options¶
| Option | Alias | Declared by | Type | Default | Range | Effect |
|---|---|---|---|---|---|---|
speed_kernelscale | ks | both | double | 1.0 | 0.1–4.0 | Scale of the Gaussian kernel (2.0 doubles the standard deviation and enlarges the kernel). |
speed_prescale | ps | both | double | 1.0 | 0.1–4.0 | Resize factor of each plane before the SpEED filters; values above 1 upsample. |
speed_prescale_method | psm | both | string | nearest | nearest, bilinear, bicubic, lanczos4 | Resize method. |
speed_sigma_nn | snn | both | double | 0.29 | 0.1–2.0 | Standard deviation of the neural noise. |
speed_nn_floor | nnf | both | double | 0.0 | 0.0–1.0 | Neural-noise floor, as a share of sigma_nn. |
speed_max_val | mxv | both | double | 1000.0 | 0.0–1000.0 | Upper clamp of the scores. |
speed_weight_var_mode | wvm | speed_chroma | int | 0 | 0–6 | Approach to variance-based weighting. |
speed_use_ref_diff | urd | speed_temporal | bool | false | — | Debug mode: additional output. |
Defaults match Netflix upstream; core/src/feature/speed.c carries the help strings.
Speed performance and exactness notes¶
- The SIMD paths vectorise the 25x25 covariance matrix each plane needs. They return the scalar path's bits on every input: a kernel keeps one lane per covariance sum instead of splitting one sum over lanes, and multiplies and adds separately (ADR-1459). Scalar, AVX2, AVX-512 and NEON dispatch give the same scores (
core/test/test_speed_simd.ccompares the sums bit for bit). - The kernels upstream Netflix ships for AVX2 and AVX-512 are faster on large planes and differ from scalar in the last bits of a sum. This fork does not use them, and
speed_chroma+speed_temporaltake up to 12 % more CPU time per frame for it (3840x2160, AVX-512). - Each plane goes through a Gaussian anti-alias filter and is then decimated by 16 in both directions. On x86 the whole plane is filtered (AVX2) and one sample in 256 is kept. On every other target the filter is evaluated only at the kept samples (
vif_filter1d_dec16_s()incore/src/feature/vif_tools.c, Netflix/vmaf76ea5f03), which returns the same bits as filtering the plane with the scalar filter and decimating it;core/test/test_speed_filter.ccompares the two withmemcmp. Scores do not change on any target. - In subsampled formats (4:2:0, 4:2:2) an odd luma width or height produces an extra chroma row or column to cover the last luma sample (
vmaf_chroma_extent());speed_chromasizes its buffers fromspeed_chroma_dimensions(), matchingpicture.c.
Stability — research. The option grammar and score scale may shift in future Netflix upstream commits; track upstream releases before pinning these features into a downstream pipeline.
Related¶
- CAMBI — banding-specific extractor.
- Confidence Interval — bootstrapped uncertainty on the final VMAF score.
- Bad cases — how to report content where extractors disagree with subjective ratings.
- Backends — which SIMD / GPU paths get picked at runtime.
- Models — how the fixed-point core extractors feed into the shipped VMAF models.
- ADR-0100 — the per-surface doc bar this page satisfies.
Licensing of the extractors (ADR-1250)¶
Extractor sources carry their terms per file as an SPDX-License-Identifier. An extractor the fork wrote is EUPL-1.2. One that carries Netflix's, libjxl's, Xiph's or IQA's code (every SIMD and GPU kernel of an upstream metric, and the scalar ports of third-party references) keeps that code's terms and its copyright notice. scripts/dev/relicense_fork_files.py --list prints the verdict and the reason for every file. See ADR-1250.
History¶
Newest first.
Dated changes, September and October 2026¶
- 2026-10-03 — ADR-1500 aligned the NEON and SVE2
float_momentsums with the scalar order (float-moment). - 2026-10-02 — SpEED scores follow Netflix's upstream double math (ADR-1477; see the SpEED page).
- 2026-10-01 —
--backend <gpu> --feature float_momentstopped computing on the CPU with a no-twin warning and runs the twin. - 2026-09-30 — The CPU
speed_temporaloverran its frame buffers for anyspeed_prescaleabove 1, so the run crashed or corrupted memory; it now sizes them for the upscaled plane (Netflix/vmaf#1626). Output atspeed_prescale1 and below is unchanged, and so isspeed_chroma, which prescale never affected.core/test/test_speed_frame_buffers.cextracts three frames at 1.0, 1.5, 2.0 and 4.0. Also on this datespeed_chroma(CPU, CUDA and HIP) sized its buffers by floor division for odd subsampled frame sizes, sopicture_copy()wrote past the allocation; it now sizes fromspeed_chroma_dimensions().
Dated changes, April and May 2026¶
- 2026-05-28 — The Vulkan backend and its column were removed (ADR-0726). Earlier ADRs that mention Vulkan or a Vulkan stub describe historical state.
- 2026-04-29 — Port of upstream Netflix/vmaf
b949cebf: themotion_add_scale1,motion_add_uv,motion_filter_sizeandmotion_max_valoptions. With the defaults, output is bit-identical to the pre-port baseline on the Y-plane SIMD fast path; non-default options route through the scalarcompute_motion()path. The same port letsfloat_motionemitmotion3_score(a perceptual blend ofmotion2, controlled bymotion_blend_factor/motion_blend_offset) on the second frame. The trained VMAF models do not consumemotion3_scoreand are unchanged. - 2026-04-29 — SSIMULACRA 2: ARM64 SVE2 ports of the IIR blur and
picture_to_linear_rgb(ADR-0213), moving the SVE2 deferral notes in Research-0016 and Research-0017 from "deferred" to "shipped". The SVE2 path runs alongside NEON on hosts that advertise thesve2HWCAP and falls back to NEON otherwise. - 2026-04-25 — SSIMULACRA 2: three SIMD ports landed, the pointwise and reduction kernels (ADR-0161), the IIR blur (ADR-0162) and
picture_to_linear_rgb(ADR-0163). All SIMD paths build with-ffp-contract=offin dedicated split static libraries to pin cross-host bit-exactness. The source-levelFP_CONTRACT OFFpragmas are wrapped for GCC so warning-clean builds keep the same no-FMA contract; compilers that ignore the pragma rely on the split-library flag.
SSIMULACRA 2 background¶
- Origin — a fork-added scalar port of the libjxl reference metric, with a bit-close C port of libjxl's
FastGaussian3-pole recursive IIR as the pyramid blur (ADR-0130, Research-0007). - SIMD parity — scalar and SIMD outputs are byte-identical on the fork's host matrix (
core/test/test_ssimulacra2_simd.c, 11 unit tests). - Cross-host determinism — libm
cbrtfandpowf(x, 2.4)are replaced with deterministic polynomials:vmaf_ss2_cbrtf(bit-trick init plus two Newton-Raphson iterations, about 7e-7 accuracy) and a 1024-entry sRGB-EOTF LUT (about 5e-7 accuracy) (ADR-0164). - CI snapshot gate —
python/test/ssimulacra2_test.pypins 48-frame mean, min, max, hmean, frame-0 and frame-47 values atplaces=4. - Warning hygiene — the registration structs of the CPU, SIMD and tiny-AI extractors are kept warning-clean across the hosted CI matrix; that does not change feature names, option names, score formulas or backend selection.
- GPU twins — first reported as optional follow-up work (BACKLOG T3-8) and shipped in ADR-0206. On a first GPU iteration GPU
cbrtfdiffered from libm by up to 42 ULP and cascaded to a 1.59e-2 pooled-score drift, which is why every path uses the sharedvmaf_ss2_cbrtf. - Contraction — the CUDA and SYCL IIR kernels are built without contraction (
--fmad=falseon CUDA,-fp-model=preciseor contraction off on SYCL) son2*sum - d1*prev1 - prev2keeps its CPU ordering.
Overview table notes¶
float_ansnrwas removed per ADR-0865.- The CAMBI Vulkan kernel (T7-36 / ADR-0210) was removed with the backend in ADR-0726.