Skip to content

Exact GPU twins

Every row is a (feature, backend) pair whose scores equal the CPU extractor's bits. A parity-gate cell whose two sides are cpu or listed twins of the feature is compared with tolerance 0 at --precision max. The rule for listing a twin is in the cross-backend gate guide.

Feature Backend ADR Evidence
adm cuda ADR-1416 RTX 4090, --precision max: 2034 of 2034 outputs (Netflix 8 to 16 bit, 1080p checkerboards, BBB 4K 50 frames) identical to --backend cpu.
adm hip ADR-1423, ADR-1525 gfx1036, --precision max: 21 fixture pairs (Netflix 8 to 16 bit, 1080p checkerboards, BBB 4K 200 frames) identical to --backend cpu (6192 values with debug=true); with the AIM pass 4141 of 4141 values identical, aim and adm3 included (Netflix 8 to 16 bit with debug=true and the default model's options, 1080p checkerboards, BBB 4K 50 and 200 frames).
adm sycl ADR-1362, ADR-1451 Arc A380 (xe), --precision max: 333 of 333 frames (Netflix 576x324 at 8 to 16 bit and 4:2:2, 1080p checkerboards, full-range noise at 8 to 16 bit, bright 16-bit 1080p, BBB 4K widened to 16 bit, BBB 4K 200 frames) identical to --backend cpu on every gate output.
cambi cuda ADR-1379, ADR-1457 RTX 4090, --precision max: 178 of 178 frames (Netflix 576x324 at 8 to 16 bit and 4:2:2, 1080p checkerboards, Sparks 10 bit, BBB 4K 48 frames, full-range noise at 8 to 16 bit, bright 16-bit 1080p) and 200 frames of BBB 4K identical to --backend cpu on every output; options too. Exact while cambi.c's own top-K double sum is exact (ADR-1379).
cambi hip ADR-1378, ADR-1437 gfx1036, --precision max: 178 of 178 frames (Netflix 576x324 at 8 to 16 bit and 4:2:2, 1080p checkerboards, Sparks 10 bit, BBB 4K 48 frames, full-range noise at 8 to 16 bit) identical to --backend cpu on every output; options too.
cambi sycl ADR-1357, ADR-1451 Arc A380 (xe), --precision max: 333 of 333 frames (Netflix 576x324 at 8 to 16 bit and 4:2:2, 1080p checkerboards, full-range noise at 8 to 16 bit, bright 16-bit 1080p, BBB 4K widened to 16 bit, BBB 4K 200 frames) identical to --backend cpu on every gate output. Exact while cambi.c's own top-K double sum is exact (ADR-1357).
float_adm cuda ADR-1420, ADR-1442 RTX 4090, --precision max, CPU reference dividing (ADR-1442): 791 of 791 scores and 2034 of 2034 outputs with debug=true (Netflix 8 to 16 bit, 1080p checkerboards, BBB 4K 50 frames) identical to --backend cpu; gate 0 on 200 BBB frames.
float_adm hip ADR-1458, ADR-1442 gfx1036, --precision max, CPU reference dividing (ADR-1442): 1246 of 1246 scores (Netflix 576x324 at 8 to 16 bit and as 10-bit 4:2:2, 1080p checkerboards, Sparks 10 bit, full-range noise at four depths, bright 16-bit 1080p, BBB 4K 48 frames) identical to --backend cpu; 3204 of 3204 outputs with debug=true; five option sets identical.
float_adm metal ADR-1496, ADR-1498 Apple M4 Pro (macOS 26.6), macOS tester bundle built from 860050c3f, outside-tester report #2118 (docs/hardware-reports/2026-10-05-apple-m4-pro.json): the parity gate's metal cell, held exact at --precision max, is 0 on the Netflix 576x324 pair at 8 and 10 bit and both 1080p checkerboards (48 + 3 + 3 + 3 frames); test_metal_float_adm_parity == on 19 of 19 cases (noise, 10, 12 and 16 bit, odd, smallest, narrow and 1080p frames, every option case).
float_adm sycl ADR-1434, ADR-1442 Arc A380 (xe), --precision max, CPU reference dividing (ADR-1442): Netflix 576x324 at 8, 10, 12 and 16 bit and 4:2:2 10-bit, both 1080p checkerboards, BBB 3840x2160 200 frames identical to --backend cpu of the same build, debug=true included (3600 of 3600 outputs on BBB).
float_moment cuda ADR-1453, ADR-1497 RTX 4090, --precision max: four outputs of 262 of 262 frames (Netflix 576x324 at 8 to 16 bit and 4:2:2, 1080p checkerboards, Sparks 10 bit, BBB 4K 48 frames, noise at 8 to 16 bit and at 40x40 to 64x64, bright 16-bit 1080p, BBB 1080p and 4K widened to 16 bit) identical to --backend cpu; past 2^53 units of 2^-16 the CPU's rounded sum (ADR-1497): 16 of 16 frames of 16-bit 4K noise, 4 of 4 of 16-bit 8K noise, 96 of 96 of BBB 4K widened to 16 bit three ways identical.
float_moment hip ADR-1447, ADR-1497 gfx1036, --precision max: four outputs of 250 of 250 frames (the fourteen sweep fixtures at 8 to 16 bit up to BBB 4K, BBB 1080p and 4K widened to 16 bit) identical to --backend cpu; past 2^53 units of 2^-16 the CPU's rounded sum (ADR-1497): 16 of 16 frames of 16-bit 4K noise, 4 of 4 of 16-bit 8K noise, 96 of 96 of BBB 4K widened to 16 bit three ways identical.
float_moment metal ADR-1496, ADR-1498 Apple M4 Pro (macOS 26.6), macOS tester bundle built from 860050c3f, outside-tester report #2118 (docs/hardware-reports/2026-10-05-apple-m4-pro.json): the parity gate's metal cell, held exact at --precision max, is 0 on the Netflix 576x324 pair at 8 and 10 bit and both 1080p checkerboards (48 + 3 + 3 + 3 frames); test_metal_float_moment_parity == on 7 of 7 cases (8 to 16 bit, bright 16 bit, a 16-bit sum past 2^53).
float_moment sycl ADR-1449, ADR-1497 Arc A380 (xe), --precision max: four outputs of 288 of 288 frames (Netflix 576x324 at 8 to 16 bit and 4:2:2, 1080p checkerboards, BBB 4K 200 frames, noise at 8 to 16 bit, bright 16-bit 1080p, BBB 4K widened to 16 bit) identical to --backend cpu; past 2^53 units of 2^-16 the CPU's rounded sum (ADR-1497): 16 of 16 frames of 16-bit 4K noise, 4 of 4 of 16-bit 8K noise, 96 of 96 of BBB 4K widened to 16 bit three ways identical.
float_motion cuda ADR-1409 RTX 4090, --precision max: Netflix 576x324, 1080p checkerboards, BBB 3840x2160 motion, motion2, motion3 identical to --backend cpu.
float_motion hip ADR-1419 gfx1036, --precision max: Netflix 576x324 (8 and 10 bit), 1080p checkerboards, BBB 3840x2160 motion/motion2/motion3 identical to --backend cpu (1617/1617 values over seven option sets).
float_motion metal ADR-1496, ADR-1498 Apple M4 Pro (macOS 26.6), macOS tester bundle built from 860050c3f, outside-tester report #2118 (docs/hardware-reports/2026-10-05-apple-m4-pro.json): the parity gate's metal cell, held exact at --precision max, is 0 on the Netflix 576x324 pair at 8 and 10 bit and both 1080p checkerboards (48 + 3 + 3 + 3 frames); test_metal_float_motion_parity == on 16 of 16 cases (noise at 8, 10 and 12 bit, 960x540 noise, 1080p checkerboard, scale-1 and chroma options, motion3, one frame, force zero).
float_motion sycl ADR-1411 Arc A380 (xe): Netflix pair, both 1080p checkerboards, 200 BBB 4K frames identical to --backend cpu; 10, 12 and 16 bit too.
float_ms_ssim cuda ADR-1403, ADR-1465 RTX 4090, --precision max: the kernel stores every window's l, c and s terms and the host adds each scale's sums in the CPU's raster order (ADR-1465). The 176x176 frame of core/test/float_ms_ssim_order_frame.h, on which per-block sums returned 0x3f7c499f for float_ms_ssim_c_scale1, returns the CPU's 0x3f7c49a0, and a second constructed frame its 0x3f7cd999 for float_ms_ssim_l_scale0 (test_cuda_float_ms_ssim_order). 15 750 of 15 750 values identical to --backend cpu on Netflix 576x324 at 8 to 16 bit and 4:2:2, 1080p checkerboards, Sparks 10 bit, BBB 4K 200 frames and at 16 bit, full-range noise at 8 to 16 bit, bright 16-bit 1080p, with enable_lcs, enable_db and clip_db; 960 000 of 960 000 on 60 000 noise frames at 176x176.
float_ms_ssim hip ADR-1403, ADR-1437, ADR-1438 gfx1036, --precision max: 178 of 178 frames (Netflix 576x324 at 8 to 16 bit and 4:2:2, 1080p checkerboards, Sparks 10 bit, BBB 4K 48 frames, full-range noise at 8 to 16 bit) identical to --backend cpu on every output; options too. The per-scale sums are added in the CPU's raster order: the constructed 176x176 pair of core/test/float_ms_ssim_order_frame.h (T-GPU-FLOAT-SSIM-FRAME-SUM-ORDER-2026-10-02) returns the CPU's float bits 0x3f7c49a0 for float_ms_ssim_c_scale1 (0x3f7c499f before) and the CPU's value on all 16 outputs.
float_ms_ssim metal ADR-1496, ADR-1498 Apple M4 Pro (macOS 26.6), macOS tester bundle built from 860050c3f, outside-tester report #2118 (docs/hardware-reports/2026-10-05-apple-m4-pro.json): the parity gate's metal cell, held exact at --precision max, is 0 on the Netflix 576x324 pair at 8 and 10 bit and both 1080p checkerboards (48 + 3 + 3 + 3 frames); test_metal_float_ms_ssim_parity == on 8 of 8 cases (the order frame of float_ms_ssim_order_frame.h, 640x480, minimum and odd frames, enable_db and clip_db, chroma).
float_ms_ssim sycl ADR-1414, ADR-1466 Arc A380 (xe), --precision max: the 176x176 noise pair seed 2437157 (core/test/ssim_order_noise.h) scores the CPU's 0x3f7d0db6 on l_scale0 (0x3f7d0db7 before ADR-1466) and the pair of core/test/float_ms_ssim_order_frame.h the CPU's 0x3f7c49a0 on c_scale1; 333 of 333 frames (Netflix 576x324 at 8 to 16 bit and 4:2:2, 1080p checkerboards, full-range noise at 8 to 16 bit, bright 16-bit 1080p, BBB 4K widened to 16 bit, BBB 4K 200 frames) identical to --backend cpu on every gate output.
float_ms_ssim_chroma cuda ADR-1403, ADR-1465 RTX 4090, --precision max: the twin runs the luma pipeline once per plane (T-MS-SSIM-GPU-CHROMA-OPTION-DRIFT-2026-09-06). float_ms_ssim, float_ms_ssim_cb and float_ms_ssim_cr identical to --backend cpu (GCC build) on every frame, 6804 of 6804 values: 1080p checkerboards (1 px and 10 px), Netflix 576x324 4:2:2 10 bit and 4:4:4 at 8 and 10 bit, BBB 1080p 4:4:4 24 frames and BBB 4K 4:2:0 30 frames, with enable_lcs, enable_db and clip_db too; on the Netflix 4:2:0 pair (288x162 chroma) both refuse the option at init.
float_ms_ssim_chroma hip ADR-1403, ADR-1437, ADR-1438 gfx1036, --precision max: the twin runs the luma pipeline once per plane (T-MS-SSIM-GPU-CHROMA-OPTION-DRIFT-2026-09-06; it used to accept enable_chroma and drop float_ms_ssim_cb and float_ms_ssim_cr). The three scores identical to --backend cpu (GCC build) on every frame, 6804 of 6804 values: 1080p checkerboards (1 px and 10 px), Netflix 576x324 4:2:2 10 bit and 4:4:4 at 8 and 10 bit, BBB 1080p 4:4:4 24 frames and BBB 4K 4:2:0 30 frames, with enable_lcs, enable_db and clip_db too; on the Netflix 4:2:0 pair (288x162 chroma) both refuse the option at init.
float_ms_ssim_chroma metal ADR-1496, ADR-1498 Apple M4 Pro (macOS 26.6), macOS tester bundle built from 860050c3f, outside-tester report #2118 (docs/hardware-reports/2026-10-05-apple-m4-pro.json): the parity gate's metal cell, held exact at --precision max, is 0 on both 1080p checkerboards (3 + 3 frames; the 576x324 pairs skip it, chroma below the 176-pixel minimum); test_metal_float_ms_ssim_parity == on 8 of 8 cases, test_float_ms_ssim_chroma_exact among them.
float_ms_ssim_chroma sycl ADR-1414, ADR-1466 Arc A380 (xe), --precision max, build of master 9aa990455 (the SYCL twin computes chroma since ADR-1299): float_ms_ssim, float_ms_ssim_cb and float_ms_ssim_cr identical to --backend cpu of the same binary on every frame, 5508 of 5508 values (1080p checkerboards, Netflix 576x324 4:2:2 10 bit and 4:4:4 at 8 and 10 bit, BBB 1080p 4:4:4 24 frames, BBB 4K 4:2:0 30 frames, with enable_lcs, enable_db and clip_db); against a GCC build 52 of 6804 values differ by at most 2.1e-14, all from the Intel math library's pow() and log10() in the host combine (T-ICX-LIBIMF-HOST-MATH-2026-10-01).
float_ms_ssim_lcs cuda ADR-1403, ADR-1465 RTX 4090, --precision max: the kernel stores every window's l, c and s terms and the host adds each scale's sums in the CPU's raster order (ADR-1465). The 176x176 frame of core/test/float_ms_ssim_order_frame.h, on which per-block sums returned 0x3f7c499f for float_ms_ssim_c_scale1, returns the CPU's 0x3f7c49a0, and a second constructed frame its 0x3f7cd999 for float_ms_ssim_l_scale0 (test_cuda_float_ms_ssim_order). 15 750 of 15 750 values identical to --backend cpu on Netflix 576x324 at 8 to 16 bit and 4:2:2, 1080p checkerboards, Sparks 10 bit, BBB 4K 200 frames and at 16 bit, full-range noise at 8 to 16 bit, bright 16-bit 1080p, with enable_lcs, enable_db and clip_db; 960 000 of 960 000 on 60 000 noise frames at 176x176.
float_ms_ssim_lcs hip ADR-1403, ADR-1437, ADR-1438 gfx1036, --precision max: 178 of 178 frames (Netflix 576x324 at 8 to 16 bit and 4:2:2, 1080p checkerboards, Sparks 10 bit, BBB 4K 48 frames, full-range noise at 8 to 16 bit) identical to --backend cpu on every output; options too. The per-scale sums are added in the CPU's raster order: the constructed 176x176 pair of core/test/float_ms_ssim_order_frame.h (T-GPU-FLOAT-SSIM-FRAME-SUM-ORDER-2026-10-02) returns the CPU's float bits 0x3f7c49a0 for float_ms_ssim_c_scale1 (0x3f7c499f before) and the CPU's value on all 16 outputs.
float_ms_ssim_lcs metal ADR-1496, ADR-1498 Apple M4 Pro (macOS 26.6), macOS tester bundle built from 860050c3f, outside-tester report #2118 (docs/hardware-reports/2026-10-05-apple-m4-pro.json): the parity gate's metal cell, held exact at --precision max, is 0 on the Netflix 576x324 pair at 8 and 10 bit and both 1080p checkerboards (48 + 3 + 3 + 3 frames); test_metal_float_ms_ssim_parity == on 8 of 8 cases (the order frame of float_ms_ssim_order_frame.h, 640x480, minimum and odd frames, enable_db and clip_db, chroma).
float_ms_ssim_lcs sycl ADR-1414, ADR-1466 Arc A380 (xe), --precision max: the 176x176 noise pair seed 2437157 (core/test/ssim_order_noise.h) scores the CPU's 0x3f7d0db6 on l_scale0 (0x3f7d0db7 before ADR-1466) and the pair of core/test/float_ms_ssim_order_frame.h the CPU's 0x3f7c49a0 on c_scale1; 333 of 333 frames (Netflix 576x324 at 8 to 16 bit and 4:2:2, 1080p checkerboards, full-range noise at 8 to 16 bit, bright 16-bit 1080p, BBB 4K widened to 16 bit, BBB 4K 200 frames) identical to --backend cpu on every gate output.
float_psnr cuda ADR-1455, ADR-1499 RTX 4090, --precision max: 268 of 268 frames (Netflix 576x324 at 8 to 16 bit and 4:2:2, 1080p checkerboards, Sparks 10 bit, BBB 4K 48 frames, full-range noise at 8 to 16 bit and at 40x40 to 64x64, bright 16-bit 1080p, BBB 1080p and 4K widened to 16 bit) identical to --backend cpu; uncapped=true too; past 2^53 units of 2^-16 the CPU's sum of its rows (ADR-1499): 8 of 8 frames of 16-bit 4K half-range noise, 16 of 16 of 16-bit 4K full-range noise, 96 of 96 of BBB 4K widened to 16 bit three ways identical.
float_psnr hip ADR-1440, ADR-1499 gfx1036, --precision max: 178 of 178 frames (Netflix 576x324 at 8 to 16 bit and 4:2:2, 1080p checkerboards, Sparks 10 bit, BBB 4K 48 frames, full-range noise at 8 to 16 bit, bright 16 bit 1080p) identical to --backend cpu; uncapped=true too; past 2^53 units of 2^-16 the CPU's sum of its rows (ADR-1499): 8 of 8 frames of 16-bit 4K half-range noise, 16 of 16 of 16-bit 4K full-range noise, 96 of 96 of BBB 4K widened to 16 bit three ways identical.
float_psnr metal ADR-1496, ADR-1498 Apple M4 Pro (macOS 26.6), macOS tester bundle built from 860050c3f, outside-tester report #2118 (docs/hardware-reports/2026-10-05-apple-m4-pro.json): the parity gate's metal cell, held exact at --precision max, is 0 on the Netflix 576x324 pair at 8 and 10 bit and both 1080p checkerboards (48 + 3 + 3 + 3 frames); test_metal_float_psnr_parity == on 10 of 10 cases (8 to 16 bit, uncapped 10 bit, bright 16 bit, identical frames, a 16-bit sum past 2^53).
float_psnr sycl ADR-1450, ADR-1499 Arc A380 (xe), --precision max: 288 of 288 frames (Netflix 576x324 at 8 to 16 bit and 4:2:2, 1080p checkerboards, BBB 4K 200 frames, full-range noise at 8 to 16 bit, bright 16-bit 1080p, BBB 4K widened to 16 bit) identical to --backend cpu; uncapped=true too; past 2^53 units of 2^-16 the CPU's sum of its rows (ADR-1499): 8 of 8 frames of 16-bit 4K half-range noise, 16 of 16 of 16-bit 4K full-range noise, 96 of 96 of BBB 4K widened to 16 bit three ways identical.
float_ssim cuda ADR-1399, ADR-1464 RTX 4090, --precision max: the kernels store every window's terms and the host adds each frame sum in the CPU's raster order (ADR-1464). The constructed 64x64 frame of core/test/float_ssim_order_frame.h, on which a per-block sum returned 0xb4e2b621, returns the CPU's 0xb4e2b622 (test_cuda_float_ssim_order). 9828 of 9828 values identical to --backend cpu on Netflix 576x324 at 8 to 16 bit and 4:2:2, 1080p checkerboards, Sparks 10 bit, BBB 4K 200 frames and at 16 bit, full-range noise at 8 to 16 bit and at 40x40 to 64x64, bright 16-bit 1080p, with enable_lcs, scale 1 to 3, enable_db and clip_db; 880 000 of 880 000 on 220 000 noise frames at 64x64.
float_ssim hip ADR-1441, ADR-1438 gfx1036, --precision max: 178 of 178 frames (Netflix 576x324 at 8 to 16 bit and 4:2:2, 1080p checkerboards, Sparks 10 bit, BBB 4K 48 frames, full-range noise at 8 to 16 bit) identical to --backend cpu on every output; scale, enable_db and clip_db too. The frame sums are added in the CPU's raster order: the constructed 64x64 pair of core/test/float_ssim_order_frame.h (T-GPU-FLOAT-SSIM-FRAME-SUM-ORDER-2026-10-02) returns the CPU's float bits 0xb4e2b622 (0xb4e2b621 before).
float_ssim metal ADR-1496, ADR-1498 Apple M4 Pro (macOS 26.6), macOS tester bundle built from 860050c3f, outside-tester report #2118 (docs/hardware-reports/2026-10-05-apple-m4-pro.json): the parity gate's metal cell, held exact at --precision max, is 0 on the Netflix 576x324 pair (48 frames; the 1080p fixtures leave it out: float_ssim_metal runs scale 1 only, T-METAL-FLOAT-SSIM-SCALE-GT1-2026-09-29); test_metal_float_ssim_parity == on 8 of 8 cases (the order frame of float_ssim_order_frame.h, order noise, 8-bit and high-bit-depth texture, enable_db, flat frames with and without clip_db).
float_ssim sycl ADR-1370, ADR-1463, ADR-1451 Arc A380 (xe), --precision max: the constructed 64x64 pair of core/test/float_ssim_order_frame.h scores the CPU's 0xb4e2b622 (0xb4e2b621 before ADR-1463); 333 of 333 frames (Netflix 576x324 at 8 to 16 bit and 4:2:2, 1080p checkerboards, full-range noise at 8 to 16 bit, bright 16-bit 1080p, BBB 4K widened to 16 bit, BBB 4K 200 frames) identical to --backend cpu on every gate output. scale=1 and scale=3 on 138 frames too (BBB 4K 50 frames).
float_ssim_lcs cuda ADR-1399, ADR-1464 RTX 4090, --precision max: the kernels store every window's terms and the host adds each frame sum in the CPU's raster order (ADR-1464). The constructed 64x64 frame of core/test/float_ssim_order_frame.h, on which a per-block sum returned 0xb4e2b621, returns the CPU's 0xb4e2b622 (test_cuda_float_ssim_order). 9828 of 9828 values identical to --backend cpu on Netflix 576x324 at 8 to 16 bit and 4:2:2, 1080p checkerboards, Sparks 10 bit, BBB 4K 200 frames and at 16 bit, full-range noise at 8 to 16 bit and at 40x40 to 64x64, bright 16-bit 1080p, with enable_lcs, scale 1 to 3, enable_db and clip_db; 880 000 of 880 000 on 220 000 noise frames at 64x64.
float_ssim_lcs hip ADR-1441, ADR-1438 gfx1036, --precision max: 178 of 178 frames (Netflix 576x324 at 8 to 16 bit and 4:2:2, 1080p checkerboards, Sparks 10 bit, BBB 4K 48 frames, full-range noise at 8 to 16 bit) identical to --backend cpu on every output; scale, enable_db and clip_db too. The frame sums are added in the CPU's raster order: the constructed 64x64 pair of core/test/float_ssim_order_frame.h (T-GPU-FLOAT-SSIM-FRAME-SUM-ORDER-2026-10-02) returns the CPU's float bits 0xb4e2b622 (0xb4e2b621 before).
float_ssim_lcs metal ADR-1496, ADR-1498 Apple M4 Pro (macOS 26.6), macOS tester bundle built from 860050c3f, outside-tester report #2118 (docs/hardware-reports/2026-10-05-apple-m4-pro.json): the parity gate's metal cell, held exact at --precision max, is 0 on the Netflix 576x324 pair (48 frames; the 1080p fixtures leave it out: float_ssim_metal runs scale 1 only, T-METAL-FLOAT-SSIM-SCALE-GT1-2026-09-29); test_metal_float_ssim_parity == on 8 of 8 cases (the order frame of float_ssim_order_frame.h, order noise, 8-bit and high-bit-depth texture, enable_db, flat frames with and without clip_db).
float_ssim_lcs sycl ADR-1370, ADR-1463, ADR-1451 Arc A380 (xe), --precision max: the constructed 64x64 pair of core/test/float_ssim_order_frame.h scores the CPU's 0xb4e2b622 with enable_lcs and equals the CPU on the luminance, contrast and structure means (float_ssim 0xb4e2b621 before ADR-1463); 333 of 333 frames (Netflix 576x324 at 8 to 16 bit and 4:2:2, 1080p checkerboards, full-range noise at 8 to 16 bit, bright 16-bit 1080p, BBB 4K widened to 16 bit, BBB 4K 200 frames) identical to --backend cpu on every gate output. scale=1 and scale=2 on 138 frames too (BBB 4K 50 frames).
float_vif cuda ADR-1412 RTX 4090, --precision max: 452 of 452 outputs (Netflix 8 to 16 bit, 1080p checkerboards, BBB 4K 50 frames) identical; gate 0 on 200 BBB frames.
float_vif hip ADR-1444 gfx1036, --precision max: 712 of 712 scores (Netflix 576x324 at 8 to 16 bit and as 10-bit 4:2:2, 1080p checkerboards, Sparks 10 bit, full-range noise at four depths, bright 16-bit 1080p, BBB 4K 48 frames) identical to --backend cpu; 1922 values with debug=true and the feature options.
float_vif metal ADR-1496, ADR-1498 Apple M4 Pro (macOS 26.6), macOS tester bundle built from 860050c3f, outside-tester report #2118 (docs/hardware-reports/2026-10-05-apple-m4-pro.json): the parity gate's metal cell, held exact at --precision max, is 0 on the Netflix 576x324 pair at 8 and 10 bit and both 1080p checkerboards (48 + 3 + 3 + 3 frames); test_metal_float_vif_parity == on 8 of 8 cases (default, debug, model options, skip_scale0, scale minimums, 10 bit, small odd frame).
float_vif sycl ADR-1422 Arc A380, --precision max: 0 on the Netflix pair, both 1080p checkerboard pairs and 200 frames of BBB 3840x2160.
motion cuda ADR-1372, ADR-1457 RTX 4090, --precision max: 196 of 196 frames (Netflix 576x324 at 8 to 16 bit and 4:2:2, 1080p checkerboards, Sparks 10 bit, BBB 4K 48 frames, full-range noise at 8 to 16 bit and at 40x40 to 64x64, bright 16-bit 1080p) and 200 frames of BBB 4K identical to --backend cpu on every output; options too.
motion hip ADR-1377, ADR-1437 gfx1036, --precision max: 178 of 178 frames (Netflix 576x324 at 8 to 16 bit and 4:2:2, 1080p checkerboards, Sparks 10 bit, BBB 4K 48 frames, full-range noise at 8 to 16 bit) identical to --backend cpu on every output; options too.
motion metal ADR-1496, ADR-1498 Apple M4 Pro (macOS 26.6), macOS tester bundle built from 860050c3f, outside-tester report #2118 (docs/hardware-reports/2026-10-05-apple-m4-pro.json): the parity gate's metal cell, held exact at --precision max, is 0 on the Netflix 576x324 pair at 8 and 10 bit and both 1080p checkerboards (48 + 3 + 3 + 3 frames); test_metal_integer_motion_parity == on 14 of 14 cases (default, debug, force zero, weight and cap, blend, tiny, odd and large frames, the five-frame window at 8 and 10 bit).
motion sycl ADR-1371, ADR-1451 Arc A380 (xe), --precision max: 333 of 333 frames (Netflix 576x324 at 8 to 16 bit and 4:2:2, 1080p checkerboards, full-range noise at 8 to 16 bit, bright 16-bit 1080p, BBB 4K widened to 16 bit, BBB 4K 200 frames) identical to --backend cpu on every gate output, VMAF_integer_feature_motion_sad_score included since the twin emits it.
motion_debug cuda ADR-1372, ADR-1457 RTX 4090, --precision max: 196 of 196 frames (Netflix 576x324 at 8 to 16 bit and 4:2:2, 1080p checkerboards, Sparks 10 bit, BBB 4K 48 frames, full-range noise at 8 to 16 bit and at 40x40 to 64x64, bright 16-bit 1080p) and 200 frames of BBB 4K identical to --backend cpu on every output; options too.
motion_debug hip ADR-1377, ADR-1437 gfx1036, --precision max: 178 of 178 frames (Netflix 576x324 at 8 to 16 bit and 4:2:2, 1080p checkerboards, Sparks 10 bit, BBB 4K 48 frames, full-range noise at 8 to 16 bit) identical to --backend cpu on every output; options too.
motion_debug metal ADR-1496, ADR-1498 Apple M4 Pro (macOS 26.6), macOS tester bundle built from 860050c3f, outside-tester report #2118 (docs/hardware-reports/2026-10-05-apple-m4-pro.json): the parity gate's metal cell, held exact at --precision max, is 0 on the Netflix 576x324 pair at 8 and 10 bit and both 1080p checkerboards (48 + 3 + 3 + 3 frames); test_metal_integer_motion_parity == on 14 of 14 cases (default, debug, force zero, weight and cap, blend, tiny, odd and large frames, the five-frame window at 8 and 10 bit).
motion_debug sycl ADR-1371, ADR-1451 Arc A380 (xe), --precision max: 333 of 333 frames (Netflix 576x324 at 8 to 16 bit and 4:2:2, 1080p checkerboards, full-range noise at 8 to 16 bit, bright 16-bit 1080p, BBB 4K widened to 16 bit, BBB 4K 200 frames) identical to --backend cpu on every gate output, VMAF_integer_feature_motion_sad_score included since the twin emits it.
motion_mffw cuda ADR-1491 RTX 4090, --precision max: 104 of 104 frames (Netflix 576x324 at 8 and 10 bit, 1080p checkerboard, BBB 4K 50 frames) identical to --backend cpu on every gate output, and the gate cell 0 on the Netflix pair (48 frames) and BBB 4K (200 frames); test_cuda_motion_five_frame_window (== on every output, 11 / 1 / 2 / 3 frames, 8 and 10 bit).
motion_mffw hip ADR-1491 gfx1036, --precision max: 104 of 104 frames (Netflix 576x324 at 8 and 10 bit, 1080p checkerboard, BBB 4K 50 frames) identical to --backend cpu on every gate output, and the gate cell 0 on the Netflix pair (48 frames) and BBB 4K (200 frames); test_hip_motion_five_frame_window (== on every output, 11 / 1 / 2 / 3 frames, 8 and 10 bit).
motion_mffw metal ADR-1496, ADR-1498 Apple M4 Pro (macOS 26.6), macOS tester bundle built from 860050c3f, outside-tester report #2118 (docs/hardware-reports/2026-10-05-apple-m4-pro.json): the parity gate's metal cell, held exact at --precision max, is 0 on the Netflix 576x324 pair at 8 and 10 bit and both 1080p checkerboards (48 + 3 + 3 + 3 frames); test_metal_integer_motion_parity == on 14 of 14 cases, the five-frame window at 8 and 10 bit among them.
motion_mffw sycl ADR-1491 Arc A380 (xe), --precision max: 104 of 104 frames (Netflix 576x324 at 8 and 10 bit, 1080p checkerboard, BBB 4K 50 frames) identical to --backend cpu on every gate output, with the combined graph and without, and the gate cell 0 on the Netflix pair (48 frames) and BBB 4K (200 frames); test_sycl_motion_five_frame_window (== on every output, 11 / 1 / 2 / 3 frames, 8 and 10 bit).
motion_v2 cuda ADR-1372, ADR-1457 RTX 4090, --precision max: 196 of 196 frames (Netflix 576x324 at 8 to 16 bit and 4:2:2, 1080p checkerboards, Sparks 10 bit, BBB 4K 48 frames, full-range noise at 8 to 16 bit and at 40x40 to 64x64, bright 16-bit 1080p) and 200 frames of BBB 4K identical to --backend cpu on every output; options too.
motion_v2 hip ADR-1377, ADR-1437 gfx1036, --precision max: 178 of 178 frames (Netflix 576x324 at 8 to 16 bit and 4:2:2, 1080p checkerboards, Sparks 10 bit, BBB 4K 48 frames, full-range noise at 8 to 16 bit) identical to --backend cpu on every output; options too.
motion_v2 metal ADR-1496, ADR-1498 Apple M4 Pro (macOS 26.6), macOS tester bundle built from 860050c3f, outside-tester report #2118 (docs/hardware-reports/2026-10-05-apple-m4-pro.json): the parity gate's metal cell, held exact at --precision max, is 0 on the Netflix 576x324 pair at 8 and 10 bit and both 1080p checkerboards (48 + 3 + 3 + 3 frames); test_metal_motion_v2_parity == on 15 of 15 cases (default, weight, cap, blend, force zero, one frame, tiny, odd and large frames, the five-frame window at 8 and 10 bit).
motion_v2 sycl ADR-1371, ADR-1451 Arc A380 (xe), --precision max: 333 of 333 frames (Netflix 576x324 at 8 to 16 bit and 4:2:2, 1080p checkerboards, full-range noise at 8 to 16 bit, bright 16-bit 1080p, BBB 4K widened to 16 bit, BBB 4K 200 frames) identical to --backend cpu on every gate output.
motion_v2_mffw cuda ADR-1491 RTX 4090, --precision max: 104 of 104 frames (Netflix 576x324 at 8 and 10 bit, 1080p checkerboard, BBB 4K 50 frames) identical to --backend cpu on every gate output, and the gate cell 0 on the Netflix pair (48 frames) and BBB 4K (200 frames); test_cuda_motion_five_frame_window (== on every output, 11 / 1 / 2 / 3 frames, 8 and 10 bit).
motion_v2_mffw hip ADR-1491 gfx1036, --precision max: 104 of 104 frames (Netflix 576x324 at 8 and 10 bit, 1080p checkerboard, BBB 4K 50 frames) identical to --backend cpu on every gate output, and the gate cell 0 on the Netflix pair (48 frames) and BBB 4K (200 frames); test_hip_motion_five_frame_window (== on every output, 11 / 1 / 2 / 3 frames, 8 and 10 bit).
motion_v2_mffw metal ADR-1496, ADR-1498 Apple M4 Pro (macOS 26.6), macOS tester bundle built from 860050c3f, outside-tester report #2118 (docs/hardware-reports/2026-10-05-apple-m4-pro.json): the parity gate's metal cell, held exact at --precision max, is 0 on the Netflix 576x324 pair at 8 and 10 bit and both 1080p checkerboards (48 + 3 + 3 + 3 frames); test_metal_motion_v2_parity == on 15 of 15 cases, the five-frame window at 8 and 10 bit among them.
motion_v2_mffw sycl ADR-1491 Arc A380 (xe), --precision max: 104 of 104 frames (Netflix 576x324 at 8 and 10 bit, 1080p checkerboard, BBB 4K 50 frames) identical to --backend cpu on every gate output, with the combined graph and without, and the gate cell 0 on the Netflix pair (48 frames) and BBB 4K (200 frames); test_sycl_motion_five_frame_window (== on every output, 11 / 1 / 2 / 3 frames, 8 and 10 bit).
psnr cuda ADR-1373, ADR-1457 RTX 4090, --precision max: 196 of 196 frames (Netflix 576x324 at 8 to 16 bit and 4:2:2, 1080p checkerboards, Sparks 10 bit, BBB 4K 48 frames, full-range noise at 8 to 16 bit and at 40x40 to 64x64, bright 16-bit 1080p) and 200 frames of BBB 4K identical to --backend cpu on every output; options too.
psnr hip ADR-1382, ADR-1437 gfx1036, --precision max: 178 of 178 frames (Netflix 576x324 at 8 to 16 bit and 4:2:2, 1080p checkerboards, Sparks 10 bit, BBB 4K 48 frames, full-range noise at 8 to 16 bit) identical to --backend cpu on every output; options too.
psnr metal ADR-1496, ADR-1498 Apple M4 Pro (macOS 26.6), macOS tester bundle built from 860050c3f, outside-tester report #2118 (docs/hardware-reports/2026-10-05-apple-m4-pro.json): the parity gate's metal cell, held exact at --precision max, is 0 on the Netflix 576x324 pair at 8 and 10 bit and both 1080p checkerboards (48 + 3 + 3 + 3 frames); test_metal_integer_psnr_parity == on 11 of 11 cases (8 to 16 bit, odd frame, identical frames, options, min_sse, apsnr with --subsample 2).
psnr sycl ADR-1365, ADR-1451 Arc A380 (xe), --precision max: 333 of 333 frames (Netflix 576x324 at 8 to 16 bit and 4:2:2, 1080p checkerboards, full-range noise at 8 to 16 bit, bright 16-bit 1080p, BBB 4K widened to 16 bit, BBB 4K 200 frames) identical to --backend cpu on every gate output.
psnr_hvs cuda ADR-1397 RTX 4090, --precision max: Netflix 576x324 (8/10/12 bit, 4:2:2), 1080p checkerboards, BBB 1080p and 4K identical to --backend cpu.
psnr_hvs hip ADR-1401 gfx1036: Netflix, 1080p checkerboards, BBB 1080p and 4K identical to --backend cpu.
psnr_hvs sycl ADR-1401 Arc A380 (xe): Netflix, 1080p checkerboards, BBB 1080p and 4K identical to --backend cpu; gate 0 on 200 BBB 4K frames.
speed_chroma cuda ADR-1477 RTX 4090, --precision max, GCC build, glibc 2.44: 759 of 759 values at the default options (the Netflix 576x324 pair at 8, 10, 12 and 16 bit, as 10-bit 4:2:2 and as 8- and 12-bit 4:4:4, both 1080p checkerboards, Sparks, noise at four bit depths, a bright 16-bit 1080p pair, a 1080p and a 10-bit 720p gradient, akiyo 352x288, BBB 1920x1080 and 3840x2160) and 2052 of 2052 over 12 option sets (the vmaf_v1.0.16 options, every weighting mode, kernelscale, sigma_nn, nn_floor, four prescales) identical to --backend cpu of the same build.
speed_chroma hip ADR-1477 gfx1036, --precision max, GCC build, glibc 2.44: 759 of 759 values at the default options (the Netflix 576x324 pair at 8, 10, 12 and 16 bit, as 10-bit 4:2:2 and as 8- and 12-bit 4:4:4, both 1080p checkerboards, Sparks, noise at four bit depths, a bright 16-bit 1080p pair, a 1080p and a 10-bit 720p gradient, akiyo 352x288, BBB 1920x1080 and 3840x2160) and 2052 of 2052 over 12 option sets (the vmaf_v1.0.16 options, every weighting mode, kernelscale, sigma_nn, nn_floor, four prescales) identical to --backend cpu of the same build.
speed_chroma sycl ADR-1477 Arc A380 (xe), --precision max, icx build, libimf: 759 of 759 values at the default options (the Netflix 576x324 pair at 8, 10, 12 and 16 bit, as 10-bit 4:2:2 and as 8- and 12-bit 4:4:4, both 1080p checkerboards, Sparks, noise at four bit depths, a bright 16-bit 1080p pair, a 1080p and a 10-bit 720p gradient, akiyo 352x288, BBB 1920x1080 and 3840x2160) and 2052 of 2052 over 12 option sets (the vmaf_v1.0.16 options, every weighting mode, kernelscale, sigma_nn, nn_floor, four prescales) identical to --backend cpu of the same build.
speed_temporal cuda ADR-1477 RTX 4090, --precision max, GCC build, glibc 2.44: 256 of 256 values at the default options (the Netflix 576x324 pair at 8, 10, 12 and 16 bit, as 10-bit 4:2:2 and as 8- and 12-bit 4:4:4, both 1080p checkerboards, Sparks, noise at four bit depths, a bright 16-bit 1080p pair, a 1080p and a 10-bit 720p gradient, akiyo 352x288, BBB 1920x1080 and 3840x2160 and a 256x144 clip) and 342 of 342 over 6 option sets (speed_use_ref_diff, kernelscale, sigma_nn, nn_floor, two prescales) identical to --backend cpu of the same build.
speed_temporal hip ADR-1477 gfx1036, --precision max, GCC build, glibc 2.44: 256 of 256 values at the default options (the Netflix 576x324 pair at 8, 10, 12 and 16 bit, as 10-bit 4:2:2 and as 8- and 12-bit 4:4:4, both 1080p checkerboards, Sparks, noise at four bit depths, a bright 16-bit 1080p pair, a 1080p and a 10-bit 720p gradient, akiyo 352x288, BBB 1920x1080 and 3840x2160 and a 256x144 clip) and 342 of 342 over 6 option sets (speed_use_ref_diff, kernelscale, sigma_nn, nn_floor, two prescales) identical to --backend cpu of the same build.
speed_temporal sycl ADR-1477 Arc A380 (xe), --precision max, icx build, libimf: 256 of 256 values at the default options (the Netflix 576x324 pair at 8, 10, 12 and 16 bit, as 10-bit 4:2:2 and as 8- and 12-bit 4:4:4, both 1080p checkerboards, Sparks, noise at four bit depths, a bright 16-bit 1080p pair, a 1080p and a 10-bit 720p gradient, akiyo 352x288, BBB 1920x1080 and 3840x2160 and a 256x144 clip) and 342 of 342 over 6 option sets (speed_use_ref_diff, kernelscale, sigma_nn, nn_floor, two prescales) identical to --backend cpu of the same build.
ssim cuda ADR-1424 RTX 4090, --precision max: 113 of 113 frames (Netflix 8 to 16 bit, 1080p checkerboards, BBB 4K 50 frames) identical; gate 0 on 200 BBB frames.
ssim hip ADR-1438 gfx1036, --precision max: 178 of 178 frames (Netflix 576x324 at 8 to 16 bit and 4:2:2, 1080p checkerboards, Sparks 10 bit, BBB 4K 48 frames, full-range noise at 8 to 16 bit) identical to --backend cpu; with enable_db and clip_db too.
ssim metal ADR-1496, ADR-1498 Apple M4 Pro (macOS 26.6), macOS tester bundle built from 860050c3f, outside-tester report #2118 (docs/hardware-reports/2026-10-05-apple-m4-pro.json): the parity gate's metal cell, held exact at --precision max, is 0 on the Netflix 576x324 pair at 8 and 10 bit and both 1080p checkerboards (48 + 3 + 3 + 3 frames); test_metal_integer_ssim_parity == on 16 of 16 cases (8 to 16 bit, odd, tiny and one-pixel frames, 1080p, enable_db, inverted at 8 and 16 bit, identical frames).
ssim sycl ADR-1443 Arc A380 (xe), --precision max: Netflix 576x324 at 8 bit (48 frames), 10, 12 and 16 bit and 4:2:2 10-bit, both 1080p checkerboards, BBB 3840x2160 200 frames identical to --backend cpu, enable_db and clip_db included.
ssimulacra2 cuda ADR-1433 RTX 4090, --precision max: 113 of 113 frames (Netflix 8 to 16 bit, 1080p checkerboards, BBB 4K 50 frames) identical to --backend cpu; gate 0 on 200 BBB frames.
ssimulacra2 hip ADR-1445 gfx1036, --precision max: 178 of 178 frames (Netflix 576x324 at 8 to 16 bit and as 10-bit 4:2:2, 1080p checkerboards, Sparks 10 bit, full-range noise at four depths, bright 16-bit 1080p, BBB 4K 48 frames) identical to --backend cpu; 186 of 186 with yuv_matrix 1, 2 and 3.
ssimulacra2 metal ADR-1496, ADR-1498 Apple M4 Pro (macOS 26.6), macOS tester bundle built from 860050c3f, outside-tester report #2118 (docs/hardware-reports/2026-10-05-apple-m4-pro.json): the parity gate's metal cell, held exact at --precision max, is 0 on the Netflix 576x324 pair at 8 and 10 bit and both 1080p checkerboards (48 + 3 + 3 + 3 frames); test_metal_ssimulacra2_parity == on its 8 score cases (identical frames, lower third, large, odd and 10-bit frames). The executable then crashed in its 4:0:0 refusal case, a validation defect with no score involved (T-METAL-SSIMULACRA2-YUV400-ACCEPTED-2026-10-05).
ssimulacra2 sycl ADR-1446 Arc A380 (xe), --precision max: 266 of 266 frames (Netflix 576x324 at 8 bit, 48 frames, at 10, 12 and 16 bit and as 10-bit 4:2:2, both 1080p checkerboards, BBB 3840x2160 200 frames) identical to --backend cpu; 153 of 153 with yuv_matrix 1, 2 and 3.
vif cuda ADR-1462 RTX 4090, --precision max: the kernels read the CPU's log2 table, uploaded at init (all 32768 entries equal, test_cuda_vif_log2_table); 1392 of 1392 scores on 348 frames (Netflix 576x324 at 8 to 16 bit and 4:2:2, 1080p checkerboards, Sparks 10 bit, BBB 4K 200 frames, noise at 8 to 16 bit and at 40x40 to 64x64, bright 16-bit 1080p) identical to --backend cpu, debug=true, vif_enhn_gain_limit=1.0 and vif_skip_scale0 too.
vif hip ADR-1435 gfx1036, --precision max: 440 of 440 scores (Netflix 576x324 at 8 and 10 bit, 1080p checkerboards, Sparks 10 bit, BBB 4K 48 frames) identical to --backend cpu; 2460 values with debug=true at 8 to 16 bit and 4:2:2.
vif sycl ADR-1432 Arc A380 (xe), --precision max: Netflix 576x324 at 8, 10, 12 and 16 bit and 4:2:2 10-bit, both 1080p checkerboards, BBB 3840x2160 200 frames identical to --backend cpu, debug=true included.