Backends¶
Pick a backend by hardware, build libvmaf with its option, and select it at run time with --backend. The CPU path (scalar C plus SIMD) is always built; GPU backends are opt-in at build time, except Metal, which auto probes on macOS.
Choose a backend¶
Run every meson setup line from the repository root (the Meson source directory is core/). The Run-time column lists the flags of the vmaf CLI; see the CLI reference.
| Hardware | Backend | Build option | Run-time selection | SDK | Status | Guide |
|---|---|---|---|---|---|---|
| x86-64 CPU | AVX2 / AVX-512 SIMD | on by default (enable_asm, enable_avx512) | --backend cpu, --cpumask | nasm 2.14+ for AVX-512 | CPU reference; SIMD kernels are checked against the scalar reference | x86 SIMD |
| aarch64 CPU | NEON / SVE2 | on by default (enable_asm) | --backend cpu, --cpumask | none | scalar bits for the contracted features | ARM |
| NVIDIA GPU | CUDA | meson setup build core -Denable_cuda=true | --backend cuda, --no_cuda | CUDA toolkit 13.4.2 (nvcc) | 24 of 25 gate features bit-identical to the CPU; ciede bounded at 1e-9 | CUDA |
| Intel GPU (Arc, Xe iGPU, Battlemage) | SYCL | meson setup build core -Denable_sycl=true | --backend sycl, --sycl_device N, --no_sycl | oneAPI DPC++ 2026.1 (icpx), Level Zero 1.34.0, ocloc | 24 of 25 gate features bit-identical; ciede bounded at 1e-9 | SYCL |
| AMD GPU | HIP | meson setup build core -Denable_hip=true -Denable_hipcc=true | --backend hip, --hip_device N, --no_hip | ROCm 10.1.0 (hipcc) | 19 registered extractors; 24 of 25 gate features bit-identical; ciede bounded at 1e-9 | HIP |
| Apple Silicon | Metal | -Denable_metal=auto (default; probes on macOS) | --backend metal, --metal_device N, --no_metal | Xcode / Metal.framework | 17 registered, parity-tested kernels; SpEED is the one gap; no exactness declared yet | Metal |
Versions come from build-config.env (CUDA_VERSION, ONEAPI_VERSION, LEVEL_ZERO_VERSION, ROCM_VERSION). "Gate features" are the 25 features of the cross-backend parity gate; "bit-identical" is defined under Numerical agreement.
Note
HIP and Metal are opt-in at run time: without --backend hip / --hip_device (or the Metal equivalents) the backend is never used, even in a binary built with it. The Vulkan backend was removed in ADR-0726.
Source: the fragments in scripts/ci/exact_twins.d/ and LIBM_TWINS in scripts/ci/cross_backend_calibration.py. An exact twin returns the CPU extractor's bits (=) and is compared with tolerance 0; a libm-bound twin differs only through the math library, within the stated bound (≤); a dash means no twin of that backend is declared either way.
Data table
| Feature | CUDA | SYCL | HIP | Metal |
|---|---|---|---|---|
adm | exact | exact | exact | not declared |
cambi | exact | exact | exact | not declared |
ciede | libm bound (≤ 1e-09) | libm bound (≤ 1e-09) | libm bound (≤ 1e-09) | not declared |
float_adm | exact | exact | exact | exact |
float_moment | exact | exact | exact | exact |
float_motion | exact | exact | exact | exact |
float_ms_ssim | exact | exact | exact | exact |
float_ms_ssim_chroma | exact | exact | exact | exact |
float_ms_ssim_lcs | exact | exact | exact | exact |
float_psnr | exact | exact | exact | exact |
float_ssim | exact | exact | exact | exact |
float_ssim_lcs | exact | exact | exact | exact |
float_vif | exact | exact | exact | exact |
motion | exact | exact | exact | exact |
motion_debug | exact | exact | exact | exact |
motion_mffw | exact | exact | exact | exact |
motion_v2 | exact | exact | exact | exact |
motion_v2_mffw | exact | exact | exact | exact |
psnr | exact | exact | exact | exact |
psnr_hvs | exact | exact | exact | not declared |
speed_chroma | exact | exact | exact | not declared |
speed_temporal | exact | exact | exact | not declared |
ssim | exact | exact | exact | exact |
ssimulacra2 | exact | exact | exact | exact |
vif | exact | exact | exact | not declared |
How selection works¶
At build time the options above decide which backends exist in the binary. At run time the CLI chooses among those compiled in.
Auto and explicit selection¶
--backend auto(the default) soft-falls back: if the priority backend fails to initialise, the run silently drops to CPU with a stderr log line.--backend NAME(cpu,cuda,sycl,hip,metal) is exclusive. For a GPU backend, an init failure is a hard error and the process exits with code100.--no_cuda,--no_sycl,--no_hip,--no_metaland--cpumaskremove candidates from the list.
Programmatic selection uses VmafConfiguration (log_level, n_threads, n_subsample, cpumask, gpumask in libvmaf.h) plus the per-backend vmaf_<backend>_import_state() calls. Python tooling (compat/python-vmaf, ExternalProgramCaller) injects --backend <name> into every child vmaf call when the environment variable VMAF_FORCE_BACKEND (or VMAF_BACKEND) is set.
Dispatch precedence¶
Inside libvmaf, the first rule that applies wins:
Text description
- Runs on: GPU twin (CUDA, SYCL, HIP or Metal), CPU extractor (AVX-512, AVX2, NEON, SVE2 or scalar C).
- Boxes: Options (--backend, --model, --feature), Features to compute (named by the model or the CLI), Extractor registry (by feature name and backend flag), Device twin? (provided_features, geometry).
- Options → Features to compute → Extractor registry (vmaf_use_feature) → Device twin? → GPU twin (twin fits).
- Device twin? → CPU extractor (no twin).
- Scenario 1, GPU twin: With a GPU backend compiled in and allowed, a feature with a twin runs on the device. The model names its features; --backend cuda allows CUDA. Each feature is looked up in the registry. The CPU extractor names a CUDA twin through provided_features. The twin accepts the frame geometry and options, so it runs.
- Scenario 2, CPU: Without a fitting twin, the CPU extractor runs its best SIMD path. --backend cpu, or no GPU backend compiled in. The feature is looked up. No twin is eligible. The CPU extractor runs; vmaf_init_cpu picked AVX-512, AVX2, NEON or scalar C.
- Disabled backends (--no_cuda, --no_sycl, ...) leave the candidate list first.
- With --backend NAME a missing or failing backend is an error, never a silent fall back to the CPU.
- Backends the user disabled (
--no_cuda,--no_sycl,--no_hip,--no_metal,--cpumask) are removed from the candidates. - If a surviving GPU backend has a twin of the feature, and the twin honours the options and the input size, the GPU twin runs.
- Otherwise the best available CPU SIMD path runs; scalar C is the fallback.
The --feature name matters. A name is used exactly as given in auto mode: --feature ciede runs the CPU extractor, --feature ciede_sycl runs the SYCL twin. With an explicit device --backend, a CPU extractor name runs on that backend's twin when the twin can honour the options and input size, and falls back to the CPU with a warning otherwise (ADR-1359; details).
Explicit-backend semantics (backend name)¶
What --backend NAME guarantees, per ADR-0498 and ADR-0543:
| Situation | Exit code | stderr | JSON output |
|---|---|---|---|
| Explicit GPU backend not compiled in or fails to init | 100 (VMAF_EXIT_BACKEND_INIT_FAILED) | vmaf: --backend NAME requested but ...; refusing to silently fall back to CPU (ADR-0498) | with --output X.json, the file is overwritten by a one-line descriptor with keys error, backend_requested, errno, adr (always "ADR-0498") and exit_code |
GPU-pinned feature name (*_cuda, *_sycl, *_hip, *_metal) whose backend is not active | 100 | same refusal, no silent CPU twin | same descriptor |
| Success | 0 | silent on stdout | gains backend_used and feature_backends (below) |
| Any other failure | other non-zero codes (for example 101 no frames decoded, 102 input read error, 255 generic) | error text | none |
The 100 code lets a CI gate match [[ $rc -eq 100 ]] to tell backend failures from other errors without parsing stderr. The vulkan token was removed with ADR-0726; passing it is an unsupported-backend error, and a _vulkan feature name has no extractor.
Example:
# Explicit HIP; errors out hard if no AMD GPU is available.
vmaf --reference ref.yuv --distorted dist.yuv \
--width 1920 --height 1080 --pixel_format 420 --bitdepth 8 \
--model version=vmaf_v0.6.1 --backend hip \
--json --output /tmp/s.json
# stdout silent on success; /tmp/s.json carries:
# { ..., "backend_used": "hip" }
# On init failure: exit = 100 (ADR-0543), stderr:
# vmaf: --backend hip requested but init failed; refusing to
# silently fall back to CPU (ADR-0498)
# AND /tmp/s.json is overwritten with a structured error descriptor:
# {"error": "vmaf_hip_state_init failed",
# "backend_requested": "hip", "errno": -19,
# "adr": "ADR-0498", "exit_code": 100}
The backend receipt¶
The JSON output carries a top-level "backend_used": "NAME" key echoing what actually ran (cpu, cuda, sycl, hip, metal), and a "feature_backends" array listing the backend of every extractor. Since ADR-1359 backend_used names a device only when at least one extractor ran on it. It mirrors the MCP-layer echo added by PR #1251. See the CLI receipt.
Numerical agreement¶
"Exact" means: at --precision max, the twin returns the same bits as the CPU extractor on every input, not within a tolerance. Where a twin differs only in the math library, the gate holds it to a stated bound instead (today ciede at 1e-9, LIBM_TWINS in scripts/ci/cross_backend_calibration.py).
Where to find the details:
- Per-twin table: the generated cross-backend exact twins page lists every declared twin. Declarations are one file per twin in
scripts/ci/exact_twins.d/<feature>.<backend>. - Gate and tolerances: the cross-backend gate page gives the tolerance table, workflow coverage, failure output and the local sweep command. The matrix runner compares CPU against CUDA, SYCL, HIP and Metal. CI currently runs CPU to SYCL
float_ssimon the conditionally enabled Arc A380 lane; it is not an all-backend, every-PR matrix. Backend-specific tests cover the wider native surface. - Parity against Netflix: the upstream parity guard checks that inherited CPU code returns what Netflix/vmaf returns, bit for bit, except for recorded deviations.
GPU results are subject to the gate's tolerances (ADR-0214) and must not be run against CPU-golden equality assertions.
Guides¶
- RC1 external tester guide: bounded hardware discovery, four-frame CPU-referenced correctness checks, and one shareable report across CPU, CUDA, SYCL, HIP and Metal.
- x86 SIMD (AVX2 / AVX-512): SIMD optimisation notes.
- ARM NEON / SVE2: aarch64 build, runtime and per-feature coverage.
- CUDA: NVIDIA GPU backend, build and invocation; CUDA twin notes for per-twin detail.
- NVTX profiling: profiling CUDA with NVIDIA Nsight.
- SYCL / oneAPI: Intel GPU backend, build and invocation.
- SYCL bundling: self-contained deployment without oneAPI.
- SYCL on Windows: native MSVC and oneAPI build, run and test.
- HIP / AMD ROCm: 19 registered feature extractors (the full table is on that page).
float_ansnr_hipwas removed in commit70ed8b3ce3(PR #38). - Metal / Apple Silicon: runtime plus 17 wired, registered, parity-tested feature kernels; the SpEED family is the one remaining Metal-twin gap.
- Kernel scaffolding templates and the context-API contract: for backend contributors.
- Vulkan: removed in ADR-0726; historical reference only.
Not every feature has every twin. The coverage matrix is in feature metrics and in each per-backend page.
Parity tests are resolution-sensitive¶
Several extractors change behaviour with picture size, so a parity test that pins one fixture cannot see the whole contract. Rules to keep in mind:
- The shared SSIM / MS-SSIM auto-scale is \(\max(1, \operatorname{round}(\min(w, h) / 256))\), which is always
1belowmin(w, h) = 384. - The ADM border crop is
(int)(dim * 0.1 - 0.5), which is0only for band dimensions<= 14; a zero crop pulls the first and last row and column into the sum. - Every CUDA parity test also runs as a
*_largevariant against a 960x540 fixture (ADR-1206), which crosses the decimation boundary and is not a multiple of the kernel block width. When adding a parity test, add it to the large-fixture list incore/test/meson.buildtoo.
float_ssim decimation by backend¶
| Backend | Decimation on the device | Behaviour at auto scale above 1 |
|---|---|---|
| CPU | reference | reference |
| CUDA | yes, equal to --backend cpu at --precision max on every frame measured (576x324, 1920x1080, 3840x2160, 8 to 16 bits; ADR-1399, CUDA guide) | runs on the device |
| SYCL | yes, bit for bit (ADR-1370, SYCL guide) | runs on the device for 1080p and 4K and every explicit scale; falls back only when the reduced picture is smaller than SSIM's 11x11 window |
| HIP | no (scale=1 only) | libvmaf replaces that context with CPU float_ssim, keeping its options (ADR-1324) |
| Metal | no (scale=1 only) | same CPU replacement (ADR-1324) |
For automatic model dispatch through the host-picture API, libvmaf evaluates the first picture before backend initialisation: auto-scale 1 stays on the selected twin, while a resolved value above 1 replaces only that context with CPU float_ssim and preserves its options (HIP and Metal). This makes common broadcast dimensions complete without pretending those backends implement decimation.
Explicitly naming float_ssim_{hip,metal} keeps the direct capability contract: auto at min(w, h) >= 384 returns -EINVAL, and scale=1 opts into full-resolution GPU SSIM. Pinning scale=1 on the GPU while leaving the CPU on auto compares different quantities, not two backends.
The SYCL device-buffer-only vmaf_read_pictures_sycl() path has no host pictures, so float_ssim_sycl cannot run there and fails its first frame with an error.
Related¶
- CLI reference:
--no_cuda/--no_sycl/--sycl_device/--cpumask/--gpumaskflags. - ADR-0022: tiny-AI runtime, separate from classic VMAF backend dispatch (tiny-AI uses ONNX Runtime execution providers).
- ADR-0027: base-image and toolchain pins for GPU CI.
History¶
Parity-test gaps that hid three defects¶
Three separate defects hid behind the resolution gap above: the speed_chroma 4K launch bug (ADR-1202), the float-ADM edge-indexing bug (ADR-1204) and the float_ssim_cuda scale=1-only limitation, lifted by ADR-1399.
ADR-0543 explicit-backend hardening¶
ADR-0543 extended ADR-0498: the explicit-backend exit code became a dedicated 100 instead of the generic 255 (int -1 truncated to uint8_t), the JSON error descriptor was added, and a GPU-pinned feature name now fails like an explicit backend instead of silently registering the CPU twin.