Exact GPU twins¶
Every row is a (feature, backend) pair whose scores equal the CPU extractor's bits. A parity-gate cell whose two sides are cpu or listed twins of the feature is compared with tolerance 0 at --precision max. The rule for listing a twin is in the cross-backend gate guide.
| Feature | Backend | ADR | Evidence |
|---|---|---|---|
adm | cuda | ADR-1416 | RTX 4090, --precision max: 2034 of 2034 outputs (Netflix 8 to 16 bit, 1080p checkerboards, BBB 4K 50 frames) identical to --backend cpu. |
adm | hip | ADR-1423, ADR-1525 | gfx1036, --precision max: 21 fixture pairs (Netflix 8 to 16 bit, 1080p checkerboards, BBB 4K 200 frames) identical to --backend cpu (6192 values with debug=true); with the AIM pass 4141 of 4141 values identical, aim and adm3 included (Netflix 8 to 16 bit with debug=true and the default model's options, 1080p checkerboards, BBB 4K 50 and 200 frames). |
adm | sycl | ADR-1362, ADR-1451 | Arc A380 (xe), --precision max: 333 of 333 frames (Netflix 576x324 at 8 to 16 bit and 4:2:2, 1080p checkerboards, full-range noise at 8 to 16 bit, bright 16-bit 1080p, BBB 4K widened to 16 bit, BBB 4K 200 frames) identical to --backend cpu on every gate output. |
cambi | cuda | ADR-1379, ADR-1457 | RTX 4090, --precision max: 178 of 178 frames (Netflix 576x324 at 8 to 16 bit and 4:2:2, 1080p checkerboards, Sparks 10 bit, BBB 4K 48 frames, full-range noise at 8 to 16 bit, bright 16-bit 1080p) and 200 frames of BBB 4K identical to --backend cpu on every output; options too. Exact while cambi.c's own top-K double sum is exact (ADR-1379). |
cambi | hip | ADR-1378, ADR-1437 | gfx1036, --precision max: 178 of 178 frames (Netflix 576x324 at 8 to 16 bit and 4:2:2, 1080p checkerboards, Sparks 10 bit, BBB 4K 48 frames, full-range noise at 8 to 16 bit) identical to --backend cpu on every output; options too. |
cambi | sycl | ADR-1357, ADR-1451 | Arc A380 (xe), --precision max: 333 of 333 frames (Netflix 576x324 at 8 to 16 bit and 4:2:2, 1080p checkerboards, full-range noise at 8 to 16 bit, bright 16-bit 1080p, BBB 4K widened to 16 bit, BBB 4K 200 frames) identical to --backend cpu on every gate output. Exact while cambi.c's own top-K double sum is exact (ADR-1357). |
float_adm | cuda | ADR-1420, ADR-1442 | RTX 4090, --precision max, CPU reference dividing (ADR-1442): 791 of 791 scores and 2034 of 2034 outputs with debug=true (Netflix 8 to 16 bit, 1080p checkerboards, BBB 4K 50 frames) identical to --backend cpu; gate 0 on 200 BBB frames. |
float_adm | hip | ADR-1458, ADR-1442 | gfx1036, --precision max, CPU reference dividing (ADR-1442): 1246 of 1246 scores (Netflix 576x324 at 8 to 16 bit and as 10-bit 4:2:2, 1080p checkerboards, Sparks 10 bit, full-range noise at four depths, bright 16-bit 1080p, BBB 4K 48 frames) identical to --backend cpu; 3204 of 3204 outputs with debug=true; five option sets identical. |
float_adm | metal | ADR-1496, ADR-1498 | Apple M4 Pro (macOS 26.6), macOS tester bundle built from 860050c3f, outside-tester report #2118 (docs/hardware-reports/2026-10-05-apple-m4-pro.json): the parity gate's metal cell, held exact at --precision max, is 0 on the Netflix 576x324 pair at 8 and 10 bit and both 1080p checkerboards (48 + 3 + 3 + 3 frames); test_metal_float_adm_parity == on 19 of 19 cases (noise, 10, 12 and 16 bit, odd, smallest, narrow and 1080p frames, every option case). |
float_adm | sycl | ADR-1434, ADR-1442 | Arc A380 (xe), --precision max, CPU reference dividing (ADR-1442): Netflix 576x324 at 8, 10, 12 and 16 bit and 4:2:2 10-bit, both 1080p checkerboards, BBB 3840x2160 200 frames identical to --backend cpu of the same build, debug=true included (3600 of 3600 outputs on BBB). |
float_moment | cuda | ADR-1453, ADR-1497 | RTX 4090, --precision max: four outputs of 262 of 262 frames (Netflix 576x324 at 8 to 16 bit and 4:2:2, 1080p checkerboards, Sparks 10 bit, BBB 4K 48 frames, noise at 8 to 16 bit and at 40x40 to 64x64, bright 16-bit 1080p, BBB 1080p and 4K widened to 16 bit) identical to --backend cpu; past 2^53 units of 2^-16 the CPU's rounded sum (ADR-1497): 16 of 16 frames of 16-bit 4K noise, 4 of 4 of 16-bit 8K noise, 96 of 96 of BBB 4K widened to 16 bit three ways identical. |
float_moment | hip | ADR-1447, ADR-1497 | gfx1036, --precision max: four outputs of 250 of 250 frames (the fourteen sweep fixtures at 8 to 16 bit up to BBB 4K, BBB 1080p and 4K widened to 16 bit) identical to --backend cpu; past 2^53 units of 2^-16 the CPU's rounded sum (ADR-1497): 16 of 16 frames of 16-bit 4K noise, 4 of 4 of 16-bit 8K noise, 96 of 96 of BBB 4K widened to 16 bit three ways identical. |
float_moment | metal | ADR-1496, ADR-1498 | Apple M4 Pro (macOS 26.6), macOS tester bundle built from 860050c3f, outside-tester report #2118 (docs/hardware-reports/2026-10-05-apple-m4-pro.json): the parity gate's metal cell, held exact at --precision max, is 0 on the Netflix 576x324 pair at 8 and 10 bit and both 1080p checkerboards (48 + 3 + 3 + 3 frames); test_metal_float_moment_parity == on 7 of 7 cases (8 to 16 bit, bright 16 bit, a 16-bit sum past 2^53). |
float_moment | sycl | ADR-1449, ADR-1497 | Arc A380 (xe), --precision max: four outputs of 288 of 288 frames (Netflix 576x324 at 8 to 16 bit and 4:2:2, 1080p checkerboards, BBB 4K 200 frames, noise at 8 to 16 bit, bright 16-bit 1080p, BBB 4K widened to 16 bit) identical to --backend cpu; past 2^53 units of 2^-16 the CPU's rounded sum (ADR-1497): 16 of 16 frames of 16-bit 4K noise, 4 of 4 of 16-bit 8K noise, 96 of 96 of BBB 4K widened to 16 bit three ways identical. |
float_motion | cuda | ADR-1409 | RTX 4090, --precision max: Netflix 576x324, 1080p checkerboards, BBB 3840x2160 motion, motion2, motion3 identical to --backend cpu. |
float_motion | hip | ADR-1419 | gfx1036, --precision max: Netflix 576x324 (8 and 10 bit), 1080p checkerboards, BBB 3840x2160 motion/motion2/motion3 identical to --backend cpu (1617/1617 values over seven option sets). |
float_motion | metal | ADR-1496, ADR-1498 | Apple M4 Pro (macOS 26.6), macOS tester bundle built from 860050c3f, outside-tester report #2118 (docs/hardware-reports/2026-10-05-apple-m4-pro.json): the parity gate's metal cell, held exact at --precision max, is 0 on the Netflix 576x324 pair at 8 and 10 bit and both 1080p checkerboards (48 + 3 + 3 + 3 frames); test_metal_float_motion_parity == on 16 of 16 cases (noise at 8, 10 and 12 bit, 960x540 noise, 1080p checkerboard, scale-1 and chroma options, motion3, one frame, force zero). |
float_motion | sycl | ADR-1411 | Arc A380 (xe): Netflix pair, both 1080p checkerboards, 200 BBB 4K frames identical to --backend cpu; 10, 12 and 16 bit too. |
float_ms_ssim | cuda | ADR-1403, ADR-1465 | RTX 4090, --precision max: the kernel stores every window's l, c and s terms and the host adds each scale's sums in the CPU's raster order (ADR-1465). The 176x176 frame of core/test/float_ms_ssim_order_frame.h, on which per-block sums returned 0x3f7c499f for float_ms_ssim_c_scale1, returns the CPU's 0x3f7c49a0, and a second constructed frame its 0x3f7cd999 for float_ms_ssim_l_scale0 (test_cuda_float_ms_ssim_order). 15 750 of 15 750 values identical to --backend cpu on Netflix 576x324 at 8 to 16 bit and 4:2:2, 1080p checkerboards, Sparks 10 bit, BBB 4K 200 frames and at 16 bit, full-range noise at 8 to 16 bit, bright 16-bit 1080p, with enable_lcs, enable_db and clip_db; 960 000 of 960 000 on 60 000 noise frames at 176x176. |
float_ms_ssim | hip | ADR-1403, ADR-1437, ADR-1438 | gfx1036, --precision max: 178 of 178 frames (Netflix 576x324 at 8 to 16 bit and 4:2:2, 1080p checkerboards, Sparks 10 bit, BBB 4K 48 frames, full-range noise at 8 to 16 bit) identical to --backend cpu on every output; options too. The per-scale sums are added in the CPU's raster order: the constructed 176x176 pair of core/test/float_ms_ssim_order_frame.h (T-GPU-FLOAT-SSIM-FRAME-SUM-ORDER-2026-10-02) returns the CPU's float bits 0x3f7c49a0 for float_ms_ssim_c_scale1 (0x3f7c499f before) and the CPU's value on all 16 outputs. |
float_ms_ssim | metal | ADR-1496, ADR-1498 | Apple M4 Pro (macOS 26.6), macOS tester bundle built from 860050c3f, outside-tester report #2118 (docs/hardware-reports/2026-10-05-apple-m4-pro.json): the parity gate's metal cell, held exact at --precision max, is 0 on the Netflix 576x324 pair at 8 and 10 bit and both 1080p checkerboards (48 + 3 + 3 + 3 frames); test_metal_float_ms_ssim_parity == on 8 of 8 cases (the order frame of float_ms_ssim_order_frame.h, 640x480, minimum and odd frames, enable_db and clip_db, chroma). |
float_ms_ssim | sycl | ADR-1414, ADR-1466 | Arc A380 (xe), --precision max: the 176x176 noise pair seed 2437157 (core/test/ssim_order_noise.h) scores the CPU's 0x3f7d0db6 on l_scale0 (0x3f7d0db7 before ADR-1466) and the pair of core/test/float_ms_ssim_order_frame.h the CPU's 0x3f7c49a0 on c_scale1; 333 of 333 frames (Netflix 576x324 at 8 to 16 bit and 4:2:2, 1080p checkerboards, full-range noise at 8 to 16 bit, bright 16-bit 1080p, BBB 4K widened to 16 bit, BBB 4K 200 frames) identical to --backend cpu on every gate output. |
float_ms_ssim_chroma | cuda | ADR-1403, ADR-1465 | RTX 4090, --precision max: the twin runs the luma pipeline once per plane (T-MS-SSIM-GPU-CHROMA-OPTION-DRIFT-2026-09-06). float_ms_ssim, float_ms_ssim_cb and float_ms_ssim_cr identical to --backend cpu (GCC build) on every frame, 6804 of 6804 values: 1080p checkerboards (1 px and 10 px), Netflix 576x324 4:2:2 10 bit and 4:4:4 at 8 and 10 bit, BBB 1080p 4:4:4 24 frames and BBB 4K 4:2:0 30 frames, with enable_lcs, enable_db and clip_db too; on the Netflix 4:2:0 pair (288x162 chroma) both refuse the option at init. |
float_ms_ssim_chroma | hip | ADR-1403, ADR-1437, ADR-1438 | gfx1036, --precision max: the twin runs the luma pipeline once per plane (T-MS-SSIM-GPU-CHROMA-OPTION-DRIFT-2026-09-06; it used to accept enable_chroma and drop float_ms_ssim_cb and float_ms_ssim_cr). The three scores identical to --backend cpu (GCC build) on every frame, 6804 of 6804 values: 1080p checkerboards (1 px and 10 px), Netflix 576x324 4:2:2 10 bit and 4:4:4 at 8 and 10 bit, BBB 1080p 4:4:4 24 frames and BBB 4K 4:2:0 30 frames, with enable_lcs, enable_db and clip_db too; on the Netflix 4:2:0 pair (288x162 chroma) both refuse the option at init. |
float_ms_ssim_chroma | metal | ADR-1496, ADR-1498 | Apple M4 Pro (macOS 26.6), macOS tester bundle built from 860050c3f, outside-tester report #2118 (docs/hardware-reports/2026-10-05-apple-m4-pro.json): the parity gate's metal cell, held exact at --precision max, is 0 on both 1080p checkerboards (3 + 3 frames; the 576x324 pairs skip it, chroma below the 176-pixel minimum); test_metal_float_ms_ssim_parity == on 8 of 8 cases, test_float_ms_ssim_chroma_exact among them. |
float_ms_ssim_chroma | sycl | ADR-1414, ADR-1466 | Arc A380 (xe), --precision max, build of master 9aa990455 (the SYCL twin computes chroma since ADR-1299): float_ms_ssim, float_ms_ssim_cb and float_ms_ssim_cr identical to --backend cpu of the same binary on every frame, 5508 of 5508 values (1080p checkerboards, Netflix 576x324 4:2:2 10 bit and 4:4:4 at 8 and 10 bit, BBB 1080p 4:4:4 24 frames, BBB 4K 4:2:0 30 frames, with enable_lcs, enable_db and clip_db); against a GCC build 52 of 6804 values differ by at most 2.1e-14, all from the Intel math library's pow() and log10() in the host combine (T-ICX-LIBIMF-HOST-MATH-2026-10-01). |
float_ms_ssim_lcs | cuda | ADR-1403, ADR-1465 | RTX 4090, --precision max: the kernel stores every window's l, c and s terms and the host adds each scale's sums in the CPU's raster order (ADR-1465). The 176x176 frame of core/test/float_ms_ssim_order_frame.h, on which per-block sums returned 0x3f7c499f for float_ms_ssim_c_scale1, returns the CPU's 0x3f7c49a0, and a second constructed frame its 0x3f7cd999 for float_ms_ssim_l_scale0 (test_cuda_float_ms_ssim_order). 15 750 of 15 750 values identical to --backend cpu on Netflix 576x324 at 8 to 16 bit and 4:2:2, 1080p checkerboards, Sparks 10 bit, BBB 4K 200 frames and at 16 bit, full-range noise at 8 to 16 bit, bright 16-bit 1080p, with enable_lcs, enable_db and clip_db; 960 000 of 960 000 on 60 000 noise frames at 176x176. |
float_ms_ssim_lcs | hip | ADR-1403, ADR-1437, ADR-1438 | gfx1036, --precision max: 178 of 178 frames (Netflix 576x324 at 8 to 16 bit and 4:2:2, 1080p checkerboards, Sparks 10 bit, BBB 4K 48 frames, full-range noise at 8 to 16 bit) identical to --backend cpu on every output; options too. The per-scale sums are added in the CPU's raster order: the constructed 176x176 pair of core/test/float_ms_ssim_order_frame.h (T-GPU-FLOAT-SSIM-FRAME-SUM-ORDER-2026-10-02) returns the CPU's float bits 0x3f7c49a0 for float_ms_ssim_c_scale1 (0x3f7c499f before) and the CPU's value on all 16 outputs. |
float_ms_ssim_lcs | metal | ADR-1496, ADR-1498 | Apple M4 Pro (macOS 26.6), macOS tester bundle built from 860050c3f, outside-tester report #2118 (docs/hardware-reports/2026-10-05-apple-m4-pro.json): the parity gate's metal cell, held exact at --precision max, is 0 on the Netflix 576x324 pair at 8 and 10 bit and both 1080p checkerboards (48 + 3 + 3 + 3 frames); test_metal_float_ms_ssim_parity == on 8 of 8 cases (the order frame of float_ms_ssim_order_frame.h, 640x480, minimum and odd frames, enable_db and clip_db, chroma). |
float_ms_ssim_lcs | sycl | ADR-1414, ADR-1466 | Arc A380 (xe), --precision max: the 176x176 noise pair seed 2437157 (core/test/ssim_order_noise.h) scores the CPU's 0x3f7d0db6 on l_scale0 (0x3f7d0db7 before ADR-1466) and the pair of core/test/float_ms_ssim_order_frame.h the CPU's 0x3f7c49a0 on c_scale1; 333 of 333 frames (Netflix 576x324 at 8 to 16 bit and 4:2:2, 1080p checkerboards, full-range noise at 8 to 16 bit, bright 16-bit 1080p, BBB 4K widened to 16 bit, BBB 4K 200 frames) identical to --backend cpu on every gate output. |
float_psnr | cuda | ADR-1455, ADR-1499 | RTX 4090, --precision max: 268 of 268 frames (Netflix 576x324 at 8 to 16 bit and 4:2:2, 1080p checkerboards, Sparks 10 bit, BBB 4K 48 frames, full-range noise at 8 to 16 bit and at 40x40 to 64x64, bright 16-bit 1080p, BBB 1080p and 4K widened to 16 bit) identical to --backend cpu; uncapped=true too; past 2^53 units of 2^-16 the CPU's sum of its rows (ADR-1499): 8 of 8 frames of 16-bit 4K half-range noise, 16 of 16 of 16-bit 4K full-range noise, 96 of 96 of BBB 4K widened to 16 bit three ways identical. |
float_psnr | hip | ADR-1440, ADR-1499 | gfx1036, --precision max: 178 of 178 frames (Netflix 576x324 at 8 to 16 bit and 4:2:2, 1080p checkerboards, Sparks 10 bit, BBB 4K 48 frames, full-range noise at 8 to 16 bit, bright 16 bit 1080p) identical to --backend cpu; uncapped=true too; past 2^53 units of 2^-16 the CPU's sum of its rows (ADR-1499): 8 of 8 frames of 16-bit 4K half-range noise, 16 of 16 of 16-bit 4K full-range noise, 96 of 96 of BBB 4K widened to 16 bit three ways identical. |
float_psnr | metal | ADR-1496, ADR-1498 | Apple M4 Pro (macOS 26.6), macOS tester bundle built from 860050c3f, outside-tester report #2118 (docs/hardware-reports/2026-10-05-apple-m4-pro.json): the parity gate's metal cell, held exact at --precision max, is 0 on the Netflix 576x324 pair at 8 and 10 bit and both 1080p checkerboards (48 + 3 + 3 + 3 frames); test_metal_float_psnr_parity == on 10 of 10 cases (8 to 16 bit, uncapped 10 bit, bright 16 bit, identical frames, a 16-bit sum past 2^53). |
float_psnr | sycl | ADR-1450, ADR-1499 | Arc A380 (xe), --precision max: 288 of 288 frames (Netflix 576x324 at 8 to 16 bit and 4:2:2, 1080p checkerboards, BBB 4K 200 frames, full-range noise at 8 to 16 bit, bright 16-bit 1080p, BBB 4K widened to 16 bit) identical to --backend cpu; uncapped=true too; past 2^53 units of 2^-16 the CPU's sum of its rows (ADR-1499): 8 of 8 frames of 16-bit 4K half-range noise, 16 of 16 of 16-bit 4K full-range noise, 96 of 96 of BBB 4K widened to 16 bit three ways identical. |
float_ssim | cuda | ADR-1399, ADR-1464 | RTX 4090, --precision max: the kernels store every window's terms and the host adds each frame sum in the CPU's raster order (ADR-1464). The constructed 64x64 frame of core/test/float_ssim_order_frame.h, on which a per-block sum returned 0xb4e2b621, returns the CPU's 0xb4e2b622 (test_cuda_float_ssim_order). 9828 of 9828 values identical to --backend cpu on Netflix 576x324 at 8 to 16 bit and 4:2:2, 1080p checkerboards, Sparks 10 bit, BBB 4K 200 frames and at 16 bit, full-range noise at 8 to 16 bit and at 40x40 to 64x64, bright 16-bit 1080p, with enable_lcs, scale 1 to 3, enable_db and clip_db; 880 000 of 880 000 on 220 000 noise frames at 64x64. |
float_ssim | hip | ADR-1441, ADR-1438 | gfx1036, --precision max: 178 of 178 frames (Netflix 576x324 at 8 to 16 bit and 4:2:2, 1080p checkerboards, Sparks 10 bit, BBB 4K 48 frames, full-range noise at 8 to 16 bit) identical to --backend cpu on every output; scale, enable_db and clip_db too. The frame sums are added in the CPU's raster order: the constructed 64x64 pair of core/test/float_ssim_order_frame.h (T-GPU-FLOAT-SSIM-FRAME-SUM-ORDER-2026-10-02) returns the CPU's float bits 0xb4e2b622 (0xb4e2b621 before). |
float_ssim | metal | ADR-1496, ADR-1498 | Apple M4 Pro (macOS 26.6), macOS tester bundle built from 860050c3f, outside-tester report #2118 (docs/hardware-reports/2026-10-05-apple-m4-pro.json): the parity gate's metal cell, held exact at --precision max, is 0 on the Netflix 576x324 pair (48 frames; the 1080p fixtures leave it out: float_ssim_metal runs scale 1 only, T-METAL-FLOAT-SSIM-SCALE-GT1-2026-09-29); test_metal_float_ssim_parity == on 8 of 8 cases (the order frame of float_ssim_order_frame.h, order noise, 8-bit and high-bit-depth texture, enable_db, flat frames with and without clip_db). |
float_ssim | sycl | ADR-1370, ADR-1463, ADR-1451 | Arc A380 (xe), --precision max: the constructed 64x64 pair of core/test/float_ssim_order_frame.h scores the CPU's 0xb4e2b622 (0xb4e2b621 before ADR-1463); 333 of 333 frames (Netflix 576x324 at 8 to 16 bit and 4:2:2, 1080p checkerboards, full-range noise at 8 to 16 bit, bright 16-bit 1080p, BBB 4K widened to 16 bit, BBB 4K 200 frames) identical to --backend cpu on every gate output. scale=1 and scale=3 on 138 frames too (BBB 4K 50 frames). |
float_ssim_lcs | cuda | ADR-1399, ADR-1464 | RTX 4090, --precision max: the kernels store every window's terms and the host adds each frame sum in the CPU's raster order (ADR-1464). The constructed 64x64 frame of core/test/float_ssim_order_frame.h, on which a per-block sum returned 0xb4e2b621, returns the CPU's 0xb4e2b622 (test_cuda_float_ssim_order). 9828 of 9828 values identical to --backend cpu on Netflix 576x324 at 8 to 16 bit and 4:2:2, 1080p checkerboards, Sparks 10 bit, BBB 4K 200 frames and at 16 bit, full-range noise at 8 to 16 bit and at 40x40 to 64x64, bright 16-bit 1080p, with enable_lcs, scale 1 to 3, enable_db and clip_db; 880 000 of 880 000 on 220 000 noise frames at 64x64. |
float_ssim_lcs | hip | ADR-1441, ADR-1438 | gfx1036, --precision max: 178 of 178 frames (Netflix 576x324 at 8 to 16 bit and 4:2:2, 1080p checkerboards, Sparks 10 bit, BBB 4K 48 frames, full-range noise at 8 to 16 bit) identical to --backend cpu on every output; scale, enable_db and clip_db too. The frame sums are added in the CPU's raster order: the constructed 64x64 pair of core/test/float_ssim_order_frame.h (T-GPU-FLOAT-SSIM-FRAME-SUM-ORDER-2026-10-02) returns the CPU's float bits 0xb4e2b622 (0xb4e2b621 before). |
float_ssim_lcs | metal | ADR-1496, ADR-1498 | Apple M4 Pro (macOS 26.6), macOS tester bundle built from 860050c3f, outside-tester report #2118 (docs/hardware-reports/2026-10-05-apple-m4-pro.json): the parity gate's metal cell, held exact at --precision max, is 0 on the Netflix 576x324 pair (48 frames; the 1080p fixtures leave it out: float_ssim_metal runs scale 1 only, T-METAL-FLOAT-SSIM-SCALE-GT1-2026-09-29); test_metal_float_ssim_parity == on 8 of 8 cases (the order frame of float_ssim_order_frame.h, order noise, 8-bit and high-bit-depth texture, enable_db, flat frames with and without clip_db). |
float_ssim_lcs | sycl | ADR-1370, ADR-1463, ADR-1451 | Arc A380 (xe), --precision max: the constructed 64x64 pair of core/test/float_ssim_order_frame.h scores the CPU's 0xb4e2b622 with enable_lcs and equals the CPU on the luminance, contrast and structure means (float_ssim 0xb4e2b621 before ADR-1463); 333 of 333 frames (Netflix 576x324 at 8 to 16 bit and 4:2:2, 1080p checkerboards, full-range noise at 8 to 16 bit, bright 16-bit 1080p, BBB 4K widened to 16 bit, BBB 4K 200 frames) identical to --backend cpu on every gate output. scale=1 and scale=2 on 138 frames too (BBB 4K 50 frames). |
float_vif | cuda | ADR-1412 | RTX 4090, --precision max: 452 of 452 outputs (Netflix 8 to 16 bit, 1080p checkerboards, BBB 4K 50 frames) identical; gate 0 on 200 BBB frames. |
float_vif | hip | ADR-1444 | gfx1036, --precision max: 712 of 712 scores (Netflix 576x324 at 8 to 16 bit and as 10-bit 4:2:2, 1080p checkerboards, Sparks 10 bit, full-range noise at four depths, bright 16-bit 1080p, BBB 4K 48 frames) identical to --backend cpu; 1922 values with debug=true and the feature options. |
float_vif | metal | ADR-1496, ADR-1498 | Apple M4 Pro (macOS 26.6), macOS tester bundle built from 860050c3f, outside-tester report #2118 (docs/hardware-reports/2026-10-05-apple-m4-pro.json): the parity gate's metal cell, held exact at --precision max, is 0 on the Netflix 576x324 pair at 8 and 10 bit and both 1080p checkerboards (48 + 3 + 3 + 3 frames); test_metal_float_vif_parity == on 8 of 8 cases (default, debug, model options, skip_scale0, scale minimums, 10 bit, small odd frame). |
float_vif | sycl | ADR-1422 | Arc A380, --precision max: 0 on the Netflix pair, both 1080p checkerboard pairs and 200 frames of BBB 3840x2160. |
motion | cuda | ADR-1372, ADR-1457 | RTX 4090, --precision max: 196 of 196 frames (Netflix 576x324 at 8 to 16 bit and 4:2:2, 1080p checkerboards, Sparks 10 bit, BBB 4K 48 frames, full-range noise at 8 to 16 bit and at 40x40 to 64x64, bright 16-bit 1080p) and 200 frames of BBB 4K identical to --backend cpu on every output; options too. |
motion | hip | ADR-1377, ADR-1437 | gfx1036, --precision max: 178 of 178 frames (Netflix 576x324 at 8 to 16 bit and 4:2:2, 1080p checkerboards, Sparks 10 bit, BBB 4K 48 frames, full-range noise at 8 to 16 bit) identical to --backend cpu on every output; options too. |
motion | metal | ADR-1496, ADR-1498 | Apple M4 Pro (macOS 26.6), macOS tester bundle built from 860050c3f, outside-tester report #2118 (docs/hardware-reports/2026-10-05-apple-m4-pro.json): the parity gate's metal cell, held exact at --precision max, is 0 on the Netflix 576x324 pair at 8 and 10 bit and both 1080p checkerboards (48 + 3 + 3 + 3 frames); test_metal_integer_motion_parity == on 14 of 14 cases (default, debug, force zero, weight and cap, blend, tiny, odd and large frames, the five-frame window at 8 and 10 bit). |
motion | sycl | ADR-1371, ADR-1451 | Arc A380 (xe), --precision max: 333 of 333 frames (Netflix 576x324 at 8 to 16 bit and 4:2:2, 1080p checkerboards, full-range noise at 8 to 16 bit, bright 16-bit 1080p, BBB 4K widened to 16 bit, BBB 4K 200 frames) identical to --backend cpu on every gate output, VMAF_integer_feature_motion_sad_score included since the twin emits it. |
motion_debug | cuda | ADR-1372, ADR-1457 | RTX 4090, --precision max: 196 of 196 frames (Netflix 576x324 at 8 to 16 bit and 4:2:2, 1080p checkerboards, Sparks 10 bit, BBB 4K 48 frames, full-range noise at 8 to 16 bit and at 40x40 to 64x64, bright 16-bit 1080p) and 200 frames of BBB 4K identical to --backend cpu on every output; options too. |
motion_debug | hip | ADR-1377, ADR-1437 | gfx1036, --precision max: 178 of 178 frames (Netflix 576x324 at 8 to 16 bit and 4:2:2, 1080p checkerboards, Sparks 10 bit, BBB 4K 48 frames, full-range noise at 8 to 16 bit) identical to --backend cpu on every output; options too. |
motion_debug | metal | ADR-1496, ADR-1498 | Apple M4 Pro (macOS 26.6), macOS tester bundle built from 860050c3f, outside-tester report #2118 (docs/hardware-reports/2026-10-05-apple-m4-pro.json): the parity gate's metal cell, held exact at --precision max, is 0 on the Netflix 576x324 pair at 8 and 10 bit and both 1080p checkerboards (48 + 3 + 3 + 3 frames); test_metal_integer_motion_parity == on 14 of 14 cases (default, debug, force zero, weight and cap, blend, tiny, odd and large frames, the five-frame window at 8 and 10 bit). |
motion_debug | sycl | ADR-1371, ADR-1451 | Arc A380 (xe), --precision max: 333 of 333 frames (Netflix 576x324 at 8 to 16 bit and 4:2:2, 1080p checkerboards, full-range noise at 8 to 16 bit, bright 16-bit 1080p, BBB 4K widened to 16 bit, BBB 4K 200 frames) identical to --backend cpu on every gate output, VMAF_integer_feature_motion_sad_score included since the twin emits it. |
motion_mffw | cuda | ADR-1491 | RTX 4090, --precision max: 104 of 104 frames (Netflix 576x324 at 8 and 10 bit, 1080p checkerboard, BBB 4K 50 frames) identical to --backend cpu on every gate output, and the gate cell 0 on the Netflix pair (48 frames) and BBB 4K (200 frames); test_cuda_motion_five_frame_window (== on every output, 11 / 1 / 2 / 3 frames, 8 and 10 bit). |
motion_mffw | hip | ADR-1491 | gfx1036, --precision max: 104 of 104 frames (Netflix 576x324 at 8 and 10 bit, 1080p checkerboard, BBB 4K 50 frames) identical to --backend cpu on every gate output, and the gate cell 0 on the Netflix pair (48 frames) and BBB 4K (200 frames); test_hip_motion_five_frame_window (== on every output, 11 / 1 / 2 / 3 frames, 8 and 10 bit). |
motion_mffw | metal | ADR-1496, ADR-1498 | Apple M4 Pro (macOS 26.6), macOS tester bundle built from 860050c3f, outside-tester report #2118 (docs/hardware-reports/2026-10-05-apple-m4-pro.json): the parity gate's metal cell, held exact at --precision max, is 0 on the Netflix 576x324 pair at 8 and 10 bit and both 1080p checkerboards (48 + 3 + 3 + 3 frames); test_metal_integer_motion_parity == on 14 of 14 cases, the five-frame window at 8 and 10 bit among them. |
motion_mffw | sycl | ADR-1491 | Arc A380 (xe), --precision max: 104 of 104 frames (Netflix 576x324 at 8 and 10 bit, 1080p checkerboard, BBB 4K 50 frames) identical to --backend cpu on every gate output, with the combined graph and without, and the gate cell 0 on the Netflix pair (48 frames) and BBB 4K (200 frames); test_sycl_motion_five_frame_window (== on every output, 11 / 1 / 2 / 3 frames, 8 and 10 bit). |
motion_v2 | cuda | ADR-1372, ADR-1457 | RTX 4090, --precision max: 196 of 196 frames (Netflix 576x324 at 8 to 16 bit and 4:2:2, 1080p checkerboards, Sparks 10 bit, BBB 4K 48 frames, full-range noise at 8 to 16 bit and at 40x40 to 64x64, bright 16-bit 1080p) and 200 frames of BBB 4K identical to --backend cpu on every output; options too. |
motion_v2 | hip | ADR-1377, ADR-1437 | gfx1036, --precision max: 178 of 178 frames (Netflix 576x324 at 8 to 16 bit and 4:2:2, 1080p checkerboards, Sparks 10 bit, BBB 4K 48 frames, full-range noise at 8 to 16 bit) identical to --backend cpu on every output; options too. |
motion_v2 | metal | ADR-1496, ADR-1498 | Apple M4 Pro (macOS 26.6), macOS tester bundle built from 860050c3f, outside-tester report #2118 (docs/hardware-reports/2026-10-05-apple-m4-pro.json): the parity gate's metal cell, held exact at --precision max, is 0 on the Netflix 576x324 pair at 8 and 10 bit and both 1080p checkerboards (48 + 3 + 3 + 3 frames); test_metal_motion_v2_parity == on 15 of 15 cases (default, weight, cap, blend, force zero, one frame, tiny, odd and large frames, the five-frame window at 8 and 10 bit). |
motion_v2 | sycl | ADR-1371, ADR-1451 | Arc A380 (xe), --precision max: 333 of 333 frames (Netflix 576x324 at 8 to 16 bit and 4:2:2, 1080p checkerboards, full-range noise at 8 to 16 bit, bright 16-bit 1080p, BBB 4K widened to 16 bit, BBB 4K 200 frames) identical to --backend cpu on every gate output. |
motion_v2_mffw | cuda | ADR-1491 | RTX 4090, --precision max: 104 of 104 frames (Netflix 576x324 at 8 and 10 bit, 1080p checkerboard, BBB 4K 50 frames) identical to --backend cpu on every gate output, and the gate cell 0 on the Netflix pair (48 frames) and BBB 4K (200 frames); test_cuda_motion_five_frame_window (== on every output, 11 / 1 / 2 / 3 frames, 8 and 10 bit). |
motion_v2_mffw | hip | ADR-1491 | gfx1036, --precision max: 104 of 104 frames (Netflix 576x324 at 8 and 10 bit, 1080p checkerboard, BBB 4K 50 frames) identical to --backend cpu on every gate output, and the gate cell 0 on the Netflix pair (48 frames) and BBB 4K (200 frames); test_hip_motion_five_frame_window (== on every output, 11 / 1 / 2 / 3 frames, 8 and 10 bit). |
motion_v2_mffw | metal | ADR-1496, ADR-1498 | Apple M4 Pro (macOS 26.6), macOS tester bundle built from 860050c3f, outside-tester report #2118 (docs/hardware-reports/2026-10-05-apple-m4-pro.json): the parity gate's metal cell, held exact at --precision max, is 0 on the Netflix 576x324 pair at 8 and 10 bit and both 1080p checkerboards (48 + 3 + 3 + 3 frames); test_metal_motion_v2_parity == on 15 of 15 cases, the five-frame window at 8 and 10 bit among them. |
motion_v2_mffw | sycl | ADR-1491 | Arc A380 (xe), --precision max: 104 of 104 frames (Netflix 576x324 at 8 and 10 bit, 1080p checkerboard, BBB 4K 50 frames) identical to --backend cpu on every gate output, with the combined graph and without, and the gate cell 0 on the Netflix pair (48 frames) and BBB 4K (200 frames); test_sycl_motion_five_frame_window (== on every output, 11 / 1 / 2 / 3 frames, 8 and 10 bit). |
psnr | cuda | ADR-1373, ADR-1457 | RTX 4090, --precision max: 196 of 196 frames (Netflix 576x324 at 8 to 16 bit and 4:2:2, 1080p checkerboards, Sparks 10 bit, BBB 4K 48 frames, full-range noise at 8 to 16 bit and at 40x40 to 64x64, bright 16-bit 1080p) and 200 frames of BBB 4K identical to --backend cpu on every output; options too. |
psnr | hip | ADR-1382, ADR-1437 | gfx1036, --precision max: 178 of 178 frames (Netflix 576x324 at 8 to 16 bit and 4:2:2, 1080p checkerboards, Sparks 10 bit, BBB 4K 48 frames, full-range noise at 8 to 16 bit) identical to --backend cpu on every output; options too. |
psnr | metal | ADR-1496, ADR-1498 | Apple M4 Pro (macOS 26.6), macOS tester bundle built from 860050c3f, outside-tester report #2118 (docs/hardware-reports/2026-10-05-apple-m4-pro.json): the parity gate's metal cell, held exact at --precision max, is 0 on the Netflix 576x324 pair at 8 and 10 bit and both 1080p checkerboards (48 + 3 + 3 + 3 frames); test_metal_integer_psnr_parity == on 11 of 11 cases (8 to 16 bit, odd frame, identical frames, options, min_sse, apsnr with --subsample 2). |
psnr | sycl | ADR-1365, ADR-1451 | Arc A380 (xe), --precision max: 333 of 333 frames (Netflix 576x324 at 8 to 16 bit and 4:2:2, 1080p checkerboards, full-range noise at 8 to 16 bit, bright 16-bit 1080p, BBB 4K widened to 16 bit, BBB 4K 200 frames) identical to --backend cpu on every gate output. |
psnr_hvs | cuda | ADR-1397 | RTX 4090, --precision max: Netflix 576x324 (8/10/12 bit, 4:2:2), 1080p checkerboards, BBB 1080p and 4K identical to --backend cpu. |
psnr_hvs | hip | ADR-1401 | gfx1036: Netflix, 1080p checkerboards, BBB 1080p and 4K identical to --backend cpu. |
psnr_hvs | sycl | ADR-1401 | Arc A380 (xe): Netflix, 1080p checkerboards, BBB 1080p and 4K identical to --backend cpu; gate 0 on 200 BBB 4K frames. |
speed_chroma | cuda | ADR-1477 | RTX 4090, --precision max, GCC build, glibc 2.44: 759 of 759 values at the default options (the Netflix 576x324 pair at 8, 10, 12 and 16 bit, as 10-bit 4:2:2 and as 8- and 12-bit 4:4:4, both 1080p checkerboards, Sparks, noise at four bit depths, a bright 16-bit 1080p pair, a 1080p and a 10-bit 720p gradient, akiyo 352x288, BBB 1920x1080 and 3840x2160) and 2052 of 2052 over 12 option sets (the vmaf_v1.0.16 options, every weighting mode, kernelscale, sigma_nn, nn_floor, four prescales) identical to --backend cpu of the same build. |
speed_chroma | hip | ADR-1477 | gfx1036, --precision max, GCC build, glibc 2.44: 759 of 759 values at the default options (the Netflix 576x324 pair at 8, 10, 12 and 16 bit, as 10-bit 4:2:2 and as 8- and 12-bit 4:4:4, both 1080p checkerboards, Sparks, noise at four bit depths, a bright 16-bit 1080p pair, a 1080p and a 10-bit 720p gradient, akiyo 352x288, BBB 1920x1080 and 3840x2160) and 2052 of 2052 over 12 option sets (the vmaf_v1.0.16 options, every weighting mode, kernelscale, sigma_nn, nn_floor, four prescales) identical to --backend cpu of the same build. |
speed_chroma | sycl | ADR-1477 | Arc A380 (xe), --precision max, icx build, libimf: 759 of 759 values at the default options (the Netflix 576x324 pair at 8, 10, 12 and 16 bit, as 10-bit 4:2:2 and as 8- and 12-bit 4:4:4, both 1080p checkerboards, Sparks, noise at four bit depths, a bright 16-bit 1080p pair, a 1080p and a 10-bit 720p gradient, akiyo 352x288, BBB 1920x1080 and 3840x2160) and 2052 of 2052 over 12 option sets (the vmaf_v1.0.16 options, every weighting mode, kernelscale, sigma_nn, nn_floor, four prescales) identical to --backend cpu of the same build. |
speed_temporal | cuda | ADR-1477 | RTX 4090, --precision max, GCC build, glibc 2.44: 256 of 256 values at the default options (the Netflix 576x324 pair at 8, 10, 12 and 16 bit, as 10-bit 4:2:2 and as 8- and 12-bit 4:4:4, both 1080p checkerboards, Sparks, noise at four bit depths, a bright 16-bit 1080p pair, a 1080p and a 10-bit 720p gradient, akiyo 352x288, BBB 1920x1080 and 3840x2160 and a 256x144 clip) and 342 of 342 over 6 option sets (speed_use_ref_diff, kernelscale, sigma_nn, nn_floor, two prescales) identical to --backend cpu of the same build. |
speed_temporal | hip | ADR-1477 | gfx1036, --precision max, GCC build, glibc 2.44: 256 of 256 values at the default options (the Netflix 576x324 pair at 8, 10, 12 and 16 bit, as 10-bit 4:2:2 and as 8- and 12-bit 4:4:4, both 1080p checkerboards, Sparks, noise at four bit depths, a bright 16-bit 1080p pair, a 1080p and a 10-bit 720p gradient, akiyo 352x288, BBB 1920x1080 and 3840x2160 and a 256x144 clip) and 342 of 342 over 6 option sets (speed_use_ref_diff, kernelscale, sigma_nn, nn_floor, two prescales) identical to --backend cpu of the same build. |
speed_temporal | sycl | ADR-1477 | Arc A380 (xe), --precision max, icx build, libimf: 256 of 256 values at the default options (the Netflix 576x324 pair at 8, 10, 12 and 16 bit, as 10-bit 4:2:2 and as 8- and 12-bit 4:4:4, both 1080p checkerboards, Sparks, noise at four bit depths, a bright 16-bit 1080p pair, a 1080p and a 10-bit 720p gradient, akiyo 352x288, BBB 1920x1080 and 3840x2160 and a 256x144 clip) and 342 of 342 over 6 option sets (speed_use_ref_diff, kernelscale, sigma_nn, nn_floor, two prescales) identical to --backend cpu of the same build. |
ssim | cuda | ADR-1424 | RTX 4090, --precision max: 113 of 113 frames (Netflix 8 to 16 bit, 1080p checkerboards, BBB 4K 50 frames) identical; gate 0 on 200 BBB frames. |
ssim | hip | ADR-1438 | gfx1036, --precision max: 178 of 178 frames (Netflix 576x324 at 8 to 16 bit and 4:2:2, 1080p checkerboards, Sparks 10 bit, BBB 4K 48 frames, full-range noise at 8 to 16 bit) identical to --backend cpu; with enable_db and clip_db too. |
ssim | metal | ADR-1496, ADR-1498 | Apple M4 Pro (macOS 26.6), macOS tester bundle built from 860050c3f, outside-tester report #2118 (docs/hardware-reports/2026-10-05-apple-m4-pro.json): the parity gate's metal cell, held exact at --precision max, is 0 on the Netflix 576x324 pair at 8 and 10 bit and both 1080p checkerboards (48 + 3 + 3 + 3 frames); test_metal_integer_ssim_parity == on 16 of 16 cases (8 to 16 bit, odd, tiny and one-pixel frames, 1080p, enable_db, inverted at 8 and 16 bit, identical frames). |
ssim | sycl | ADR-1443 | Arc A380 (xe), --precision max: Netflix 576x324 at 8 bit (48 frames), 10, 12 and 16 bit and 4:2:2 10-bit, both 1080p checkerboards, BBB 3840x2160 200 frames identical to --backend cpu, enable_db and clip_db included. |
ssimulacra2 | cuda | ADR-1433 | RTX 4090, --precision max: 113 of 113 frames (Netflix 8 to 16 bit, 1080p checkerboards, BBB 4K 50 frames) identical to --backend cpu; gate 0 on 200 BBB frames. |
ssimulacra2 | hip | ADR-1445 | gfx1036, --precision max: 178 of 178 frames (Netflix 576x324 at 8 to 16 bit and as 10-bit 4:2:2, 1080p checkerboards, Sparks 10 bit, full-range noise at four depths, bright 16-bit 1080p, BBB 4K 48 frames) identical to --backend cpu; 186 of 186 with yuv_matrix 1, 2 and 3. |
ssimulacra2 | metal | ADR-1496, ADR-1498 | Apple M4 Pro (macOS 26.6), macOS tester bundle built from 860050c3f, outside-tester report #2118 (docs/hardware-reports/2026-10-05-apple-m4-pro.json): the parity gate's metal cell, held exact at --precision max, is 0 on the Netflix 576x324 pair at 8 and 10 bit and both 1080p checkerboards (48 + 3 + 3 + 3 frames); test_metal_ssimulacra2_parity == on its 8 score cases (identical frames, lower third, large, odd and 10-bit frames). The executable then crashed in its 4:0:0 refusal case, a validation defect with no score involved (T-METAL-SSIMULACRA2-YUV400-ACCEPTED-2026-10-05). |
ssimulacra2 | sycl | ADR-1446 | Arc A380 (xe), --precision max: 266 of 266 frames (Netflix 576x324 at 8 bit, 48 frames, at 10, 12 and 16 bit and as 10-bit 4:2:2, both 1080p checkerboards, BBB 3840x2160 200 frames) identical to --backend cpu; 153 of 153 with yuv_matrix 1, 2 and 3. |
vif | cuda | ADR-1462 | RTX 4090, --precision max: the kernels read the CPU's log2 table, uploaded at init (all 32768 entries equal, test_cuda_vif_log2_table); 1392 of 1392 scores on 348 frames (Netflix 576x324 at 8 to 16 bit and 4:2:2, 1080p checkerboards, Sparks 10 bit, BBB 4K 200 frames, noise at 8 to 16 bit and at 40x40 to 64x64, bright 16-bit 1080p) identical to --backend cpu, debug=true, vif_enhn_gain_limit=1.0 and vif_skip_scale0 too. |
vif | hip | ADR-1435 | gfx1036, --precision max: 440 of 440 scores (Netflix 576x324 at 8 and 10 bit, 1080p checkerboards, Sparks 10 bit, BBB 4K 48 frames) identical to --backend cpu; 2460 values with debug=true at 8 to 16 bit and 4:2:2. |
vif | sycl | ADR-1432 | Arc A380 (xe), --precision max: Netflix 576x324 at 8, 10, 12 and 16 bit and 4:2:2 10-bit, both 1080p checkerboards, BBB 3840x2160 200 frames identical to --backend cpu, debug=true included. |