Cross-backend GPU-parity gate¶
scripts/ci/cross_backend_parity_gate.py runs the same clip through the CPU and one or more GPU backends (CUDA, SYCL, HIP, Metal), compares every per-frame metric, and says whether each backend returns the CPU's numbers within the allowed tolerance. It writes a JSON file and a Markdown summary.
In CI the gate runs in one lane only: the SYCL Parity (Arc A380) job in .github/workflows/sycl-parity.yml, which compares CPU and SYCL float_ssim on the self-hosted Arc runner. For every other backend and feature you run it yourself, as Run it locally shows. Vulkan and its hosted lavapipe lane were removed by ADR-0726.
For users: what "exact" means¶
A GPU twin is exact when it returns the same bits as the CPU extractor for the same input. The gate compares an exact cell with tolerance 0, at --precision max (%.17g) so the default %.6f rounding cannot hide a last-bit difference. Every other cell passes when each per-frame value is within an absolute tolerance, 5e-5 by default (four decimal places).
The generated table of exact twins lists every (feature, backend) pair that is exact today, with the ADR and the measurement behind it. Netflix golden assertions remain untouched; parity thresholds never replace the CPU golden-data gate.
To check your own GPU build against the CPU:
- Build
vmafwith the backend, or use the dev container (Run it locally). -
Run the gate on a clip pair and read
OK,FAIL,ERRORorSKIPper cell (Read the output):
Backends¶
| Backend | Extractor suffix | Device flag passed to vmaf | Gate option | Default device |
|---|---|---|---|---|
cpu | none | none | ||
cuda | _cuda | --gpumask | --cuda-device | 1 (first GPU) |
sycl | _sycl | --sycl_device | --sycl-device | 0 |
hip | _hip | --hip_device | --hip-device | 0 |
metal | _metal | --metal_device | --metal-device | 0 |
The command line default is --backends cpu cuda. No CI job runs the HIP or Metal cells; HIP is for a local run on an AMD host, and Metal runs from the macOS tester bundle (Metal).
Extractors whose twin is not named <feature>_<backend> are listed in BACKEND_EXTRACTOR_ALIASES:
float_ms_ssimisinteger_ms_ssim_hipon HIP;ssimisinteger_ssim_cuda,integer_ssim_sycl,integer_ssim_hipandinteger_ssim_metal;- the Metal twins of the fixed-point extractors carry the CPU file's name (
integer_adm_metal,integer_motion_metal, and so on).
Command-line options¶
| Option | Default | Meaning |
|---|---|---|
--vmaf-binary | required | Path to the vmaf binary under test. |
--reference, --distorted | required | Reference and distorted YUV files. |
--width, --height | required | Frame size in pixels. |
--pixel-format | 420 | Chroma subsampling of the YUV files. |
--bitdepth | 8 | Bit depth of the YUV files. |
--features | every feature in FEATURE_METRICS | Features (cells) to run. |
--backends | cpu cuda | Backends to pair; cpu is the reference. |
--hold-exact | none | Compare every cell of these backends exactly, although no fragment lists them (Metal). |
--fp16-features | none | Features that use the FP16 contract, absolute tolerance 1e-2. |
--cuda-device, --sycl-device, --hip-device, --metal-device | see Backends | Device index per backend. |
--gpu-id | none | Runtime GPU identifier (for example cuda:8.6, sycl:0); picks the most specific matching row in the calibration table. With no match, the feature's built-in tolerance stays authoritative. |
--calibration-table | scripts/ci/gpu_ulp_calibration.yaml | Per-device calibration rows (ADR-0234). |
--workdir | <tmp>/vmaf_parity_gate | Scratch directory for per-run output. |
--json-out, --md-out | none | Write the machine-readable and the Markdown summary. |
Tolerance contract¶
The default is 5e-5 (places=4), the fork's GPU-versus-CPU contract from ADR-0125, ADR-0138 and ADR-0140. Feature-specific relaxations live in FEATURE_TOLERANCE inside cross_backend_parity_gate.py.
| Cell | Tolerance | Source |
|---|---|---|
Any feature, both sides cpu or a listed exact twin | 0 (bit-identical, at --precision max) | The ADR in the twin's fragment under scripts/ci/exact_twins.d/. |
ciede, every pair of CPU, CUDA, SYCL and HIP | 1e-9 at --precision max | ADR-1426, ADR-1436, ADR-1448; see math-library twins. |
vif, adm, psnr (planes psnr_y, psnr_cb, psnr_cr), float_moment, cambi, motion family | 5e-5 | ADR-0125, ADR-0138, ADR-0140, ADR-0360. |
ssim (fixed-point extractor) | 5e-5 | ADR-0564 (int64 moments, one double term per pixel). |
float_ssim, float_ssim_lcs, float_ms_ssim, float_ms_ssim_lcs, float_ms_ssim_chroma, float_psnr, float_motion, float_vif, float_adm | 5e-5 | ADR-0188, ADR-0192, ADR-0215, ADR-1382. |
speed_chroma (the three scores speed_chroma_u, _v, _uv), speed_temporal (the one score speed_temporal), for a twin that is not exact | 5e-5 | Places=4; the CUDA, HIP and SYCL twins are exact (ADR-1477). |
psnr_hvs (a twin that is not exact) | 5e-4 at 576x324 and below, \(5 \times 10^{-4} \sqrt{N / N_{576 \times 324}}\) above | ADR-0191 (DCT plus per-block float reduction); ADR-1361 (area scaling). |
ciede (a twin that is not in the 1e-9 row) | 5e-3 | ADR-0187 (per-pixel pow, sqrt, sin, atan2). |
ssimulacra2 | 5e-3 | ADR-0192 (XYB cube root plus IIR blur). |
Notes on single cells¶
motion_debugismotionwithdebug=trueand addsinteger_motion. Both cells compareVMAF_integer_feature_motion_sad_score, the per-frame score thatmotion2andmotion3derive from, next tointeger_motion2andinteger_motion3(ADR-1418).motion_mffwandmotion_v2_mffwaremotionandmotion_v2withmotion_five_frame_window=true:motion_moving_average=true, the option set of thevmaf_v1.0.16_hfr_*models. They compare the SAD score,motion2andmotion3under their option-suffixed names (ADR-1491).float_motioncomparesmotion,motion2andmotion3;motion3since 2026-10-03, when the SYCL twin started to emit it as the CUDA and HIP twins do (T-GPU-FLOAT-MOTION3-MISSING-2026-09-30).float_ssim_lcsrunsfloat_ssimwithenable_lcs=trueand comparesfloat_ssim_l,float_ssim_candfloat_ssim_snext to the score, at the same5e-5(ADR-1382).-
float_ms_ssim_chromarunsfloat_ms_ssimwithenable_chroma=trueand comparesfloat_ms_ssim_cbandfloat_ms_ssim_crnext to the score; CUDA, SYCL and HIP are listed exact. Every chroma plane must be at least 176 pixels on a side, so on a fixture with smaller chroma, such as the 576x324 4:2:0 Netflix pair (288x162), the cell is reportedSKIPwith that reason and nothing is run (FEATURE_MIN_CHROMA_DIM). Run it on a 1080p 4:2:0, a 4:2:2 or a 4:4:4 pair:python3 scripts/ci/cross_backend_parity_gate.py \ --vmaf-binary build/tools/vmaf \ --reference python/test/resource/yuv/checkerboard_1920_1080_10_3_0_0.yuv \ --distorted python/test/resource/yuv/checkerboard_1920_1080_10_3_10_0.yuv \ --width 1920 --height 1080 --backends cpu cuda \ --features float_ms_ssim_chroma -
Every registered twin is a gate cell. A twin registered in
core/src/feature/feature_extractor.cppthat is no gate feature's extractor is guarded by its own unit test only;speed_temporalwas the one such twin on CUDA, SYCL and HIP.core/test/test_parity_gate_covers_registered_twins.pyfails when a new twin of a gated backend has no gate feature, and when the tables of this gate and ofcross_backend_vif_diff.pydiffer.
Area-scaled psnr_hvs¶
Currently unused by the gate
The CPU psnr_hvs adds every coefficient error of a plane into one float, so its rounding error, and the achievable CPU/GPU agreement, grows with the number of 8x8 blocks.
Both cross_backend_parity_gate.py and cross_backend_vif_diff.py take the psnr_hvs tolerance from the table or calibration row as the contract at 576x324 and multiply it by sqrt(N / N_576x324) for larger fixtures, where N is the luma plane's term count (64 per block, a block every 7 pixels). Frames of 576x324 or smaller keep 5e-4; 1920x1080 gets 1.67e-3 and 3840x2160 3.34e-3. The label in the output shows the factor, for example default+area x6.69 (ADR-1361 derives it from the float-accumulation bound).
This is the contract of a twin that sums each block on the device. The CUDA, SYCL and HIP twins no longer do (they are exact), and the Metal twin does but has no listed psnr_hvs fragment, so no psnr_hvs cell of the gate uses it today. It stays for a caller that names no backends and for a twin added later.
Exact twins¶
A GPU twin that returns the CPU extractor's bits for a feature is listed by adding one file, scripts/ci/exact_twins.d/<feature>.<backend> (for example psnr_hvs.cuda), and nothing else shared (ADR-1428). The file holds two key: value lines:
adr:one or moreADR-NNNNthat exist underdocs/adr/, the ADR that establishes exactness;evidence:one line: fixtures and result.
cross_backend_calibration.py builds EXACT_TWINS from the directory at import and rejects an unknown key, a missing key, an empty file, an unknown feature or backend, and a missing or empty directory. The generated table of exact twins is rendered from the same files by make docs-fragments-write.
A listed CUDA, SYCL or HIP twin also needs a recorded row on the depth and layout matrix: every twin at 8, 10, 12 and 16 bits and in 4:2:0, 4:2:2 and 4:4:4, == against the CPU on generated fixtures. scripts/ci/exact_twin_matrix.py --record writes the row, and test_exact_twin_matrix_contract fails while a listed twin has none.
A cell whose two sides are the CPU extractor or a listed twin is compared with tolerance 0, at every frame size and ahead of any calibration row, and both sides run with --precision max so that a last-bit difference is not rounded away by the default %.6f output. The label in the output is exact:ADR-1397 for every listed twin. An explicit --fp16-features <feature> still selects the FP16 contract. A cell with a backend that is not listed keeps the tolerance of the tolerance contract.
The rule for listing¶
Listing needs a measurement that shows bit-identity (--precision max, the Netflix 576x324 pairs, the 1080p checkerboard pairs and testdata/bbb 4K) and an ADR that records it. A listed twin that drifts is fixed; it is never given a tolerance and never taken off the list to make a gate pass.
Identical on those fixtures is necessary, not sufficient (ADR-1437). A twin is listed when it also reaches the CPU's value by construction: integer sums on the device, the CPU's own helpers on the host, or the CPU's arithmetic type for type.
Natural 8-bit content does not tell the two apart, and the repository's 10-, 12- and 16-bit Netflix clips are the 8-bit clip shifted left. Before listing a twin, also run it on full-range noise at 8, 10, 12 and 16 bits: float_psnr_hip and float_moment_hip matched every real clip measured and differ there.
What the comparison assumes¶
The equality holds between runs of one vmaf binary, which is how the gate runs a cell. The dB value goes through the host's log10.
A binary built with oneAPI icx before ADR-1495 used Intel's libimf and a gcc build glibc, and the two rounded differently by one unit in the last place on a few frames (3 of the 48 Netflix frames), for the CPU extractor and the twins alike. An icx build now links glibc's libm (test_icx_system_libm checks it), so a SYCL build's twins also equal a gcc build's CPU extractor; a binary built with another C library can still differ from either.
An exact cell runs one binary on both sides, so it also needs the CPU extractor of that binary to be the reference arithmetic; for an icx build on an AVX-512 host that needs the x86 SIMD libraries built without FP contraction (#1706).
Twins that differ only in their math library¶
ciede_cuda evaluates ciede.c's expressions in its types and the host adds the per-pixel values in the CPU's order (ADR-1426). It is not bit-identical: the CPU calls glibc's pow, atan2, sin, cos, exp and powf, the device CUDA's, and a few pixels per million round to the neighbouring float.
LIBM_TWINS in scripts/ci/cross_backend_calibration.py gives such a cell its own tolerance (ciede: cuda at 1e-9, label libm:ADR-1426) and runs both sides at --precision max. The bound is for frames of 576x324 and larger; one such pixel weighs more in a smaller frame.
The three twins that carry this bound:
| Twin | How it differs from the CPU | ADR |
|---|---|---|
ciede_cuda | CUDA's math library against glibc's. | ADR-1426 |
ciede_sycl | A SYCL kernel has no fp64 type, so it runs the same statements on fp32 pairs, which decide all but about one pixel in a million the way fp64 does. | ADR-1436 |
ciede_hip | Runs the SYCL twin's pair statements from the same header. | ADR-1448 |
Measured against a GCC build on 180 frames (Netflix 576x324 at 8 to 16 bits and as 10-bit 4:2:2, Sparks, both 1080p checkerboards, BBB 1920x1080 and 3840x2160), since ADR-1467 writes the squares of ciede.c as products:
| Twin | Frames identical | Largest difference | Frames identical before ADR-1467 | Largest before |
|---|---|---|---|---|
| CUDA | 127 | 5.2e-12 | 113 | 2.0e-11 |
| SYCL | 124 | 5.2e-12 | 111 | 2.0e-11 |
| HIP | 127 | 5.2e-12 | 113 | 2.0e-11 |
The Netflix pair is identical on every frame. What is left is glibc's powf(x, 7). The bound stays 1e-9: it is the size of one differing pixel on the smallest gated frame, and there are fewer such pixels, not smaller ones.
Metal¶
The metal backend runs where an Apple device is: the macOS tester bundle carries the gate and runs it on its four fixtures on the tester's Mac (ADR-1496). The SpEED features have no Metal twin and are not run there.
No fragment lists a Metal twin yet, so the bundle runs the gate with --hold-exact metal: every Metal cell is compared with tolerance 0 at --precision max, or at the LIBM_TWINS bound for ciede (1e-9), and the label in the output is held-exact:ADR-1496. That run is the measurement a scripts/ci/exact_twins.d/<feature>.metal fragment cites; the option never replaces the fragment.
The bundle leaves float_ssim and float_ssim_lcs out of the 1080p fixtures, where the Metal twin's automatic scale (4) is one it does not implement (T-METAL-FLOAT-SSIM-SCALE-GT1-2026-09-29); the report lists them under left_out.
The gate must stay importable with the Python standard library alone and read nothing beyond scripts/ci/cross_backend_calibration.py, scripts/lib/safe_subprocess.py, scripts/ci/exact_twins.d/ and the ADR files the fragments cite: those are the files the bundle copies (GATE_FILES in tools/rc1-tester/image/prepare_build.py).
Run it locally¶
Use the dev-MCP container, which carries the backend toolchains and the repository fixture mounts.
-
Build and start the container:
-
Run the gate inside it:
docker exec vmaf-dev-mcp bash -lc ' cd /workspace && python3 scripts/ci/cross_backend_parity_gate.py \ --vmaf-binary /usr/local/bin/vmaf \ --reference testdata/ref_576x324_48f.yuv \ --distorted testdata/dis_576x324_48f.yuv \ --width 576 --height 324 \ --backends cpu cuda \ --features float_ssim vif \ --json-out /tmp/parity.json \ --md-out /tmp/parity.md '
On an AMD host, the same runner compares the HIP twins, including the float_ssim enable_lcs cell:
python3 scripts/ci/cross_backend_parity_gate.py \
--vmaf-binary build-hip/tools/vmaf \
--reference testdata/ref_576x324_48f.yuv \
--distorted testdata/dis_576x324_48f.yuv \
--width 576 --height 324 \
--backends cpu hip --hip-device 0 \
--features float_ssim float_ssim_lcs psnr motion_v2 vif \
--json-out /tmp/parity-hip.json --md-out /tmp/parity-hip.md
Pin each concurrent run to different hardware; do not multiplex one device across parallel parity jobs.
Whole models on one backend¶
The gate compares one extractor at a time. Whether a model runs wholly on a device is a separate question: the option gate (ADR-1183, ADR-1316) sends an extractor to the CPU when its twin cannot honour a model option, and the run still succeeds. core/test/test_gpu_v1_models_no_fallback.py checks the model set end to end. For every built-in vmaf_v1.0.16* model and the default model, it scores these fixtures with --backend cpu and with the device backend at --precision max:
- the 576x324
src01_hrc00/src01_hrc01pair at 8, 10 and 12 bits 4:2:0 and 10 bits 4:2:2; - the first 16 frames of the 3840x2160 pair in
testdata/bbb.
It fails when the device run's feature_backends lists an extractor on any other backend, or when any per-frame, pooled or aggregate value differs. Each run also scores --feature float_adm=adm_csf_mode=1, a known fallback, and fails if the check does not report it.
Meson registers one test per backend the build enables: test_cuda_v1_models_no_fallback, test_sycl_v1_models_no_fallback and test_hip_v1_models_no_fallback (suites slow, gpu and the backend name). Each one skips with the reason when it finds no device or a missing fixture. testdata/bbb and python/test/resource/yuv are not in the repository. When ~/.cache/vmafx-locks exists (or VMAFX_LOCK_DIR), each device run takes that backend's lock file. Run a SYCL build's test with the oneAPI environment loaded:
Read the output¶
The JSON artifact contains one record per cell with status, tolerance_abs, tolerance_source, n_frames, per_metric_max_abs_diff, per_metric_mismatches, and a free-text note for errors. Its top-level schema_version versions the format. The Markdown artifact contains the same cell summary plus a failure-detail section.
A cell's status is one of:
| Status | Meaning |
|---|---|
OK | Every per-frame metric is within tolerance. |
FAIL | At least one per-frame mismatch exceeds tolerance_abs. |
ERROR | Execution failed before diffing, the frame counts differ, or one backend does not emit a metric the cell compares (the note names the backend and the metrics). |
SKIP | The fixture cannot exercise the cell (the chroma of a float_ms_ssim_chroma fixture is below 176 pixels); the note says why. A skipped cell does not fail the run and is not a pass. |
Add a feature or backend¶
- Add the feature-to-metric mapping to
FEATURE_METRICSin both parity scripts (cross_backend_parity_gate.pyandcross_backend_vif_diff.py), with the metrics the CPU extractor emits by default. Metric names must match the keys emitted by the selectedvmaffeature extractor. - Add a
FEATURE_TOLERANCEentry only when the feature differs from the default5e-5, and cite the measurement or ADR that owns the relaxation. - Declare a twin exact, once it is measured bit-identical, by adding
scripts/ci/exact_twins.d/<feature>.<backend>(adr:andevidence:), then runmake docs-fragments-write. Do not edit the calibration module, the tests or this page for it. - Add or update unit tests in
scripts/ci/test_cross_backend_parity_gate.pyand update the tolerance table above where the feature's tolerance changed. The exact-twin tests hold for any set of fragments and need no edit. - For a new backend, extend the suffix and device-selection maps, add executable coverage, and update the consuming workflow explicitly. Adding support to the script alone does not create CI coverage.
Relationship to other gates¶
Text description
- Gates: Unit tests (run_meson_test.py -- -C build), Netflix golden gate (make test-netflix-golden), Cross-backend parity gate (every GPU cell against the CPU).
- Boxes: Change (source, tests, docs), Isolated golden build (gcc or clang, no FP contraction), Netflix assertions (python/test/, never edited), EXACT_TWINS, LIBM_TWINS (tolerance 0 or a derived bound).
- Change → Unit tests → Netflix golden gate → Cross-backend parity gate.
- Netflix golden gate → Isolated golden build (VMAF_BUILD_DIR).
- Netflix golden gate → Netflix assertions (assertAlmostEqual).
- Cross-backend parity gate → EXACT_TWINS, LIBM_TWINS (tolerance).
- Scenario 1, Gates in order: Each gate must pass before the change lands. The unit tests run under the credential-safe Meson wrapper. The golden gate builds the CPU library on its own build directory. Netflix's assertions run against that build. The parity gate runs each GPU backend against the CPU. An exact twin must match to the last bit; a libm twin within its bound.
- The golden assertions are Netflix ground truth and are never edited; a drifting score is fixed in the code.
- An exact twin is compared at --precision max with tolerance 0.
| Gate | Role |
|---|---|
| Netflix golden (ADR-0024) | CPU numerical correctness; required and untouchable. |
| SYCL Parity (Arc A380) | Conditional required lane; CPU and SYCL float_ssim on real Arc hardware. It runs for eligible non-draft in-repository pull requests, pushes to master and manual dispatches when SYCL_ARC_RUNNER_ENABLED=true and the runner is online; when the lane is disabled, the required-check aggregator explicitly accepts its skip. |
| Backend Meson parity tests | Backend-specific correctness, including large-fixture variants where registered. |
test_<backend>_v1_models_no_fallback | Every vmaf_v1.0.16* model and the default model run wholly on the device and return the CPU run's bits (Whole models on one backend). |
| This matrix runner outside CI | Broader CPU, CUDA, SYCL and HIP feature sweeps and calibration evidence. |
| macOS tester bundle (tester page) | Runs this gate's Metal cells, held exact, on an outside tester's Mac (ADR-1496). |
Per-backend snapshots (testdata/scores_cpu_*.json) | Snapshot-based regression checks, not pairwise parity. |
Sources¶
- ADR-0214: original matrix design.
- ADR-0726: removal of Vulkan and its hosted matrix lanes.
- ADR-1177: current Arc runner and lane-switch contract.
- ADR-1428: exact-twin fragments.
History¶
- ADR-1477:
speed_chromaandspeed_temporalwere libm twins (5e-6and4e-5; ADR-1430, ADR-1452, ADR-1460) because the fork'sspeed.ccalledlog2f, each twin roundedlog2on the device, and glibc'slog2fmoved a few scores by one to five steps of the fp32 result.speed.cevaluates Netflix's fp64log2again, and the twins form the entropies and the score on the host with those statements and the same C library, so the six cells are exact:speed_chroma.{cuda,hip,sycl}andspeed_temporal.{cuda,hip,sycl}underscripts/ci/exact_twins.d/. An AdaptiveCpp build of the SYCL twin is outside that (ADR-1358). - ADR-1467: the first ciede measurements predate the change that writes the squares of
ciede.cas products. On an RTX 4090: 1.4e-11 on 200 frames of BBB 3840x2160, 6.9e-13 on the Netflix 576x324 pair; on an Arc A380 the same 1.4e-11 and 6.9e-13; on a gfx1036 1.4e-11 on 48 frames of BBB 3840x2160 and 6.9e-13 on the Netflix pair, where of 437 million pixels compared one by one 2 206 differ through glibc'spowfand 8 through the last bits of a pair. - ADR-1496: the gate gained the
metalbackend,--metal-deviceand--hold-exact.T-GATE-NO-METAL-BACKEND-2026-10-02instate.mdtracks the remaining work: a hardware report that shows every Metal cellOK.