ADR-1499: The float_psnr twins add each row's exact sum in the CPU's order, and return the CPU's bits on every frame¶
- Status: Accepted
- Date: 2026-10-03
- Deciders: lusoris
- Tags:
cuda,sycl,hip,gpu-parity,numerics,psnr,testing,rc3,fork-local - Supersedes: the bound past 2^53 units stated in ADR-1440, ADR-1450 and ADR-1455. Their integer block sums stand.
Context¶
float_psnr.c::extract() forms each sample difference's square in float, adds a row's squares into a double (noise_line(), scalar or AVX2, AVX-512, NEON) and adds the rows into one double. The CUDA, SYCL and HIP twins add the same float squares as integers in units of 1 / scaler^2, per 16x16 block, and divided the exact frame total on the host. That is the CPU's sum while it is at most 2^53 units. On a 16-bit frame whose mean squared error times its pixel count passes 2^37 on the 8-bit scale the CPU's adds of the rows round, and the twins, which rounded the exact total once, were within a stated bound: 0 of 8 frames of 16-bit 3840x2160 noise (reference in the upper half of the range, distorted frame in the lower) identical on the RTX 4090 and the gfx1036, 1 of 8 on the Arc A380. No ledger row recorded the gap. After the same decision for float_moment (ADR-1497), the maintainer decided that float_psnr must return the CPU's bits on every input for rc.3 as well.
Decision¶
We will lay each block (CUDA, HIP) or work-group (SYCL) out over 256 pixels of one row instead of 16x16 pixels, and form the CPU's sum on the host with one helper for the three twins, core/src/feature/float_psnr_rows.h (vmaf_float_psnr_row_noise()): each row's segment sums added in 64-bit integers, the rows added into one double in order, then the CPU's two divisions. The HIP kernel puts its two 16-bit halves together into one uint64 per block.
Why it is the CPU's double on every input:
- In units of 1 / scaler^2 every term is an integer below 2^32, the float square of the sample difference (the twins' term since ADR-1440 / 1450 / 1455). A row has at most 2^15 terms, so every partial sum inside a row is an integer below 2^47, exact in a
doublein any order: everynoise_line()variant returns the row's exact integer sum. - The device's segment sums and the host's 64-bit row sums are exact integers of the same terms, so the host holds each row's sum exactly, and converting it to
doubleis exact (below 2^47). - The host adds those values into a
doublerow after row, the CPU's adds of the same operands in the same order: every rounding past 2^53, ties included, is the CPU's. - The CPU adds the rows' true values (the integer sums times 2^-2k for scaler = 2^k); multiplying every operand by a power of two scales every exactly-rounded result by it while nothing underflows or overflows (the smallest nonzero value is 2^-16), so the host's sum in units, divided by scaler^2, is the CPU's sum. The division by the pixel count is the CPU's.
No device arithmetic of the ordered-sum kind is needed: the exact-sum machinery of float_moment_sum.h exists because moment.c adds pixel by pixel; float_psnr.c adds row by row, and its rows are exact.
Alternatives considered¶
| Option | Pros | Cons | Why not chosen |
|---|---|---|---|
| Blocks of one row, the rows' exact sums added in order on the host (this ADR) | The CPU's adds of the CPU's operands, by the argument above; same number of blocks and read-back bytes as before; no new kernel | The block shape changes (256 x 1) | Chosen |
| Keep 16x16 blocks and read back per-row sums with 64-bit atomics | Fewer host adds | An atomic per row segment, a buffer to clear per frame, and a second layout | Not needed: a block of one row already is a segment |
float_moment_sum.h's walk on the device | One exact-sum design for both extractors | Four kernels and a walk to reproduce a sum the host forms with h adds | The CPU's rows are exact; nothing to emulate |
| Keep the bound | No change | Not the CPU's bits | The maintainer decided against it (References) |
Consequences¶
- Positive: at
--precision maxagainst--backend cpu, every frame identical on an RTX 4090, an Arc A380 and a gfx1036: 8 of 8 frames of the 16-bit 3840x2160 half-range noise above (0, 1 and 0 of 8 before), 16 of 16 of 16-bit 3840x2160 full-range noise, 32 of 32 of BBB 3840x2160 widened to 16 bits in each of three ways. The parity gate'sfloat_psnrcell reads 0 at tolerance 0 on the Netflix pair and on the half-range noise on each device (FAIL on master there).test_{cuda,sycl,hip}_float_psnr_parityassert==on that frame and on frames whose exact sum is 2^53 - 1, 2^53, and 2^53 followed by three rows of sum 1; each fails on master.test_float_psnr_rowsholds the helper against the CPU extractor on the host, a 7680x4320 frame included. - Positive: no measurable cost (per 16-bit 3840x2160 frame, before and after: 4.12 and 4.12 ms on the RTX 4090, 7.16 and 7.06 ms on the A380, 10.9 and 9.3 ms on the gfx1036).
- Neutral / follow-ups:
- The HIP parity test now uses
core/test/float_psnr_twin_parity.h, as the CUDA and SYCL tests do (the RC5 follow-up of ADR-1455). float_psnr_metalstill adds fp32 block sums (T-METAL-FLOAT-PSNR-FP32-BLOCK-SUMS-2026-10-02); its fix should take this layout and helper too.- Guards:
test_float_psnr_rows(host), the threetest_*_float_psnr_exact_contract.py(a block that spans rows, a frame total instead of rows, planted), the three device parity tests.
References¶
Q(maintainer popup, 2026-10-03, float_psnr past 2^53 units): "Make it exact for rc.3 (Recommended)".- ADR-1497: the same decision for
float_moment. docs/state.md:T-GPU-FLOAT-PSNR-PAST-2-53-2026-10-03(opened and closed by this ADR).