Appendix of integer accumulator bounds: every integer accumulator, size product and offset of the CUDA and HIP feature twins, read on master 571565a47 (before the fixes the main page lists). Rows marked OVERFLOW or DEPENDS name their state row on the main page.
Source: master 571565a47.
Mode: read only. Nothing was edited, built or run on a device. The only commands were grep, sed and cat, plus a few python3 bound computations. The ADM replay is scripts/dev/adm_cm_row_bound.py.
Scope:
every file under core/src/feature/cuda/ and core/src/feature/hip/ (130 files);
the shared core/src/feature/*.h headers those files include for GPU-side integer math;
core/src/cuda/*.c, core/src/hip/*.c and core/src/cuda/cuda_helper.cuh, for size products only.
Envelope:
16K: W ≤ 15360, H ≤ 8640, N = 132,710,400.
8K DCI: 8192x4320.
Cap: W, H ≤ 32768, N ≤ \(2^{30}\).
Samples up to 16 bit, 4:4:4.
Worst-case content: max-difference frames or constructed extremal patterns.
The audit was split into eight feature groups. Each group read every file of its group in full and derived every bound from the code it read. Each group reports its own constants and derivations. Every row that is not SAFE was then re-checked every row that is not SAFE against the source, and spot-checked a set of SAFE rows (listed at the end of this section).
Feature sections: every feature has CUDA and HIP subsections.
Integer ADM is the exception. Its HIP kernels are line-for-line ports of the CUDA kernels, so each ADM row cites both files and is tagged "CUDA …; HIP …".
Rows on shared headers sit in their own subsection of the feature that uses them.
Two verdicts changed from the first draft, both in integer ADM: - Scale-0 CM row total: DEPENDS → OVERFLOW@16K. The envelope covers every W ≤ 15360, and W = 63–64 wraps at default options. - Scale-0 CM frame accumulator: DEPENDS → SAFE. The frame sum cannot wrap for any int64 row value.
The (int32_t) cast binds to the shifted product only. The multiplication by the double gain yields a double, which is implicitly converted to int32_t *before* min(rst, t_val) / max(rst, t_val).
Bound: |r| ≤ |o| ≤ 1.449e9 at scale 1 and ≤ 7.5e8 at scales 2-3. Those are the DWT band maxima; they agree with adm_csf_fixed_point.h:88-90.
The default gain is 100 (adm_options.h:40). With angle_flag set, |r| > 21,474,837 produces a double outside the int32 range, which is UB under C++ [conv.fpint].
This does not depend on frame size. It needs only a band value above about 1.5 % of the scale-1 maximum and a dis band within 1° of the ref band, which is ordinary content on high-quality encodes.
The CPU (integer_adm_kernels.h:380-385) takes MIN(rst * gain, t) in double first, which is defined.
The GPU scores match the CPU only because the hardware conversions saturate: INT32_MAX gives min(INT32_MAX, t) = t. This is consistent with the existing bit-exact parity results, but it is not guaranteed under LLVM (fptosi out of range is poison).
Fix: rst = (int32_t)fmin((double)r * gain, (double)t) and the mirrored fmax, as the CPU does.
**Integer ADM scale-0 contrast-masking row total, int64 (CUDA, HIP, and the CPU reference).**
Locations:
CUDA adm_cm.cu:561, 894 (accum_row), :515 (warp_reduce), :519.
HIP adm_cm.hip:440, 719, :457, :465.
CPU integer_adm_kernels.h:1017-1019, 1079-1099.
Term: ((x² + 2^28) >> 29) · x >> (ceil(log2 Wb) − 4), where Wb = W/2. The shift uses the band width (integer_adm.c:711 passes w2; integer_adm_kernels.h:989).
Worst case: ref == dis, so a = 0 and thr = 0, and x = |band| · weight. With 4 pixel rows (0, 0, max, 0), the exact column-sign maximum comes from the DP in scripts/dev/adm_cm_row_bound.py. It was re-run for this page.
Row maximum as a fraction of INT64_MAX at the default Watson weights:
W
row max / INT64_MAX
48
0.807
56
0.878
**64**
**1.021 (wraps)**
128
0.973
1080p (1920)
0.857
8K DCI
0.912
15360
0.855
16384
0.912
cap
0.912
The 8-bit maximum at W = 64 is 1.009 and the 10-bit maximum is 1.018.
W ≥ 17 is accepted (ADM_MIN_FRAME_DIM, adm_csf_fixed_point.h:277), so W = 63–64 runs.
With non-default CSF weights the row also overflows at 16K and at the cap: h/v ≥ 38,406 or d ≥ 60,965 at 16K, h/v ≥ 37,588 at the cap. Those weights are reachable through adm_csf_scale / adm_csf_diag_scale / the viewing-geometry options.
The bound is not monotonic in W, because the normalising shift is ceil(log2 Wb).
The CPU has the same signed int64 sum, so all backends share the defect: signed-overflow UB on the CPU, and a two's-complement wrap on the GPUs.
16K 4:4:4: 3 · 2194 · 1234 = 8,122,188 blocks, so 31,728 chunks. That is under the cap, with a 3.2 % margin.
The cap is passed above 8,388,608 blocks, for example:
16384x8640 4:4:4: 8,662,680 blocks.
luma-only frames above about 20.3K x 20.3K.
cap 4:4:4: 256,779 chunks.
Above that, chunk_offsets[c ≥ 32768] are never written. d_scratch is a hipMalloc buffer that is never cleared, so hvs_compact_hip writes packed_terms + chunk_offsets[chunk] + intra from stale or garbage offsets. The results are out-of-bounds device writes, a short total_terms, and wrong scores.
The CUDA twin (psnr_hvs_score.cu:443) scans every chunk.
The uint32_t sum itself is safe: ≤ 4,207,058,112 at the cap.
**CUDA SpEED covariance divisor rounded to fp32 (cuda/speed/speed_score.cu:1219).**
Code: static_cast<float>(g.sub_w * g.sub_h).
This is not a wrap: the uint32_t product is at most 67,010,596. The problem is the fp32 exact-integer range, \(2^{24}\).
The count is exact at 16K for every option value: the largest is 3836 · 2156 = 8,270,416 at speed_prescale 4.
Above 16K with a prescale above about 2, the count can be inexact. For example, W = H = 32740 at prescale 4 gives sub 8181, and 8181² = 66,928,761 is odd and above \(2^{24}\).
The CPU (speed.c:847) divides the double sum by the exact double count. The twin is then no longer bit-exact.
The mean divisor (:1183) rounds the count the same way the CPU's compute_mean (speed.c:775) does, so it is SAFE.
**HIP SpEED covariance divisor (hip/speed/speed_hip_device.h:624).** Same defect as row 4.
The HIP vertical pass stores tmp.mu1 without the (uint16_t) truncation that the CPU (integer_vif.c:450-451) and CUDA (filter1d.cu:603-604) apply.
For bpc 9-15 with samples above \(2^{\mathrm{bpc}} - 1\), the stored value reaches 8,388,480, and the 17-tap sum reaches 5.5e11. That wraps, and HIP then differs from the CPU.
In-range samples are SAFE: ≤ 4,294,901,760.
**HIP integer VIF vif_downsample_storeaccum_ref_rd (hip/integer_vif/vif_statistics.hip:506-507, 329-336).** Same cause as row 9: ref_convol is not truncated.
fm_mirror() (cuda/float_motion/float_motion_score.cu:43-50) reflects once without a clamp.
fm_load_tile() (:85-90) loads the full 20x20 tile for every block. For W or H in {3..9, 17}, an index goes negative (down to −13) and the load reads before the plane.
The HIP twin and the CUDA motion SAD kernels clamp through vmaf_*_tile_index().
A2 (CUDA integer VIF, full-mask shuffle inside a divergent branch).
vif_hori_flush_accums() → warp_reduce() runs __shfl_down_sync(0xffffffff, …) under if (y < h && x_start < w) (cuda/integer_vif/filter1d.cu, the kernel body around :513-535).
Lanes past the plane edge do not take part, which is undefined for a full mask.
The code shape is verified. The effect on hardware is not.
**A3 (core/src/cuda/cuda_helper.cuh:129-130).** (x >> 32) << 32 on a negative long long is UB before C++20.
A4 (stale comments).hip/float_psnr/float_psnr_score.hip:107-111,139-141 describe a partials[2*block] layout, but the code writes one u64 per block.
A5 (outside scope, CPU).
third_party/xiph/psnr_hvs.c:311-317: the int index reaches exactly INT_MAX for a >8-bit 32768² plane.
integer_motion.c:193/243: the uint32_t row_sad wraps for out-of-range samples (bpc 9, W ≥ 513).
sycl/speed_sycl_pipeline.cpp:808: the same fp32 covariance divisor as rows 4-5.
A6 (memory footprint only, no wrap).
HIP integer ADM tmp_accum takes 24 B/pixel: 3.2 GB at 16K and 25.8 GB at the cap.
CUDA vif_cuda allocates 7 planes that no kernel reads: 3.2e9 B at 16K.
The integer SSIM twins need 7 · \(2^{33}\) B at the cap.
SpEED at prescale 4 needs about 34 GB at 16K.
Every one of these is a size_t product, so allocation fails cleanly.
Common root of rows 8-10. No libvmaf entry point checks that samples are ≤ \(2^{\mathrm{bpc}}\) − 1. The CPU truncates or uses wide types in places where HIP does not.
Read adm_decouple_inline.cuh:125-170 and adm_decouple_inline.hip:150-170.
Read the CPU integer_adm_kernels.h:355-385.
Confirmed the gain default 100.0 at adm_options.h:40.
Row 2:
Re-ran scripts/dev/adm_cm_row_bound.py; the output is the table in row 2.
Confirmed the term and shifts: adm_cm.cu:216-220, :500-561; integer_adm_kernels.h:989-1003.
Confirmed the band width w2 (integer_adm.c:711) and the minimum frame size 17 (adm_csf_fixed_point.h:277).
For the frame accumulator: region rows ≤ 0.875 · \(2^{\lceil \log_2 H_b \rceil}\), with the maximum at Hb = 16 or 32. So |Σ (row >> s)| < 0.875 · \(2^{63}\) for any int64 row value.
Row 3:
Read psnr_hvs_score.hip:430-526 and integer_psnr_hvs_hip.c:169-189, :375-394.
Recomputed the 16K block and chunk counts.
Rows 4-5:
Read speed_score.cu:1175-1222 and speed_hip_device.h:545-628.
Read the CPU speed.c:766-776, :835-850, :1285-1317, and the option range speed.c:1518-1523 (speed_prescale 0.1..4.0).
Recomputed sub at 16K at prescale 4: 3836 x 2156.
Found that 32768 at prescale 4 gives 8186² = 67,010,596, which is a multiple of 4 and so exact in fp32. Only some sizes above \(2^{24}\) are inexact, for example 32740.
Rows 6-8:
Read integer_psnr_cuda.c:380-400, the apsnr_sse sites in both twins, and float_psnr_score.hip:1-181.
Recomputed the wrap frames.
Rows 9-10: read vif_statistics.hip:455-540 against filter1d.cu:585-612.
Row 11: read integer_motion_cuda.c:525-550, :822-831, and grepped both trees for other (int)index narrowings (none).
Spot-checked SAFE rows, each read in the source:
CUDA psnr_score.cu (uint64 SSE, int64 diff, one atomic per block).
CUDA motion_v2_score.cu: |h| ≤ 65535; SAD ≤ N · 65535 = \(2^{46}\) at the cap.
CUDA moment_score.cu: u64 sums; ref2 ≤ \(2^{62}\) at the cap.
SSIMULACRA 2 CUDA unsigned index 2u * pixels + i: ≤ 3 · \(2^{30} - 1\) < \(2^{32}\).
HIP SpEED (uint32_t)plane_bytes: \(2^{31}\) at the cap, which fits.
G1: integer ADM, CUDA and HIP twins (integer-overflow audit)¶
Repo: master at 571565a47. All paths are under core/src/feature/. Read-only. Nothing was built or run on a device. The bounds come from the code, the composite-filter bounds per DWT scale of core/src/feature/adm_csf_fixed_point.h (held by test_integer_adm_cm_budget), and scripts/dev/adm_cm_row_bound.py (the exact maximum of a scale-0 CM row over every column sign pattern, found by dynamic programming over the overlapping 4-tap windows).
Constants used by every row:
DWT taps. lo = {15826, 27411, 7345, -4240} and hi = {-4240, -7345, 27411, -15826}. Each has an absolute sum of 54822. The lo taps sum to 46342 (integer_adm.h:187-191; the GPU copies are in adm_fixed_parameters() and adm_hip_fixed_params()).
Band maxima (half the absolute sum of the composite filter; they agree with adm_csf_fixed_point.h:88-90):
scale 0: every band ≤ 22,929.8.
scale 1: band_a ≤ 1,491,390,732 and h/v/d ≤ 1,448,979,042.
scale 2: ≤ 751,510,749.
scale 3: ≤ 748,642,077.
Every maximum is below INT32_MAX. They are per sample and do not depend on frame size.
Defaults (Watson97, 3H, 1080): 36,453 for h/v and 49,417 for d (tabulated).
The options adm_csf_mode and adm_csf_scale / adm_csf_diag_scale (continuous doubles, 0..50, applied multiplicatively in barten_csf_tools.h:159-160) can place a weight anywhere below its budget. So can a non-default adm_norm_view_dist or adm_ref_display_height.
Region.cols = Wb - 2*int(0.1*Wb - 0.5), where Wb is the band width (W/2 at scale 0).
CUDA adm_dwt2.cu:115-141; HIP adm_dwt2.hip:130-156
accum of the s123 horizontal pass, then (int32_t) band stores at :120/127/134/141 (HIP :135/142/149/156)
int64_t → int32_t
4 taps × tmp (≤ 1.0587e9). Shifts: scale 1 h15, scale 2 h16, scale 3 h15, which is the CPU's i4_dwt2_round() (integer_adm_kernels.h:1370-1380)
accum ≤ 4.9e13. Band ≤ 1,491,390,732 (scale-1 band_a)
same
SAFE (accum \(2^{45.5}\) < \(2^{63}\); band 0.6945·INT32_MAX)
CUDA adm_dwt2.cu:182; HIP adm_dwt2.hip:197
d_picture[y_in * src_stride + x]
int
Row index × element stride. CUDA: pitch/2 (16-bit) or pitch (8-bit) from cuMemAllocPitch (integer_adm_cuda.c:1236-1240). HIP: a packed plane, luma_stride = w (integer_adm_hip.c:1113)
≤ 8639·15360 + 15359 = 1.327e8
≤ 32767·32768 + 32767 = 1,073,741,823
SAFE (cap 0.50·INT32_MAX)
CUDA adm_dwt2.cu:68-71, 120-141, 279; HIP adm_dwt2.hip:83-86, 135-156, 294
|r| ≤ 1.449e9 (scale 1), ≤ 7.5e8 (scales 2-3), × gain ≤ 100, converted to int32 *before* min(rst, t). Out of range whenever angle_flag holds and |r| > \(2^{31}\)/gain = 21,474,837 (gain 100): 1.5 % of the scale-1 band maximum, 2.9 % at scales 2-3. Ordinary edge content does this. The CPU (integer_adm_kernels.h:380, rst = MIN(rst*gain, t)) takes the minimum in double first, which is defined
r·gain ≤ 1.449e11 ≫ INT32_MAX (any frame size)
same
**OVERFLOW@16K (defect, independent of frame size).** The conversion is UB under C++ [conv.fpint] (LLVM fptosi gives poison). The scores match the CPU only if the device conversion saturates, in which case min(INT32_MAX, t) = t. The parity tests are bit-exact on RTX 4090 and gfx1036 (state row T-ADM-AIM-BARTEN), which is consistent with saturation. The PTX/ISA wording was not verified. Fix: rst = (int32_t)min((double)r*gain, (double)t), as the CPU does
CUDA cuda/integer_adm/adm_csf.cu:169-170, cuda/integer_adm/adm_cm.cu:154-155, 690-691; HIP hip/integer_adm/adm_csf.hip:191-192, hip/integer_adm/adm_cm.hip:208-209 (HIP widens to int64 there)
dst_val = i_rfactor * (uint32_t)a_val (wraps mod \(2^{32}\), then converted to int32), then dst_val + i_shiftsadd[band]
uint32_t → int32_t; int32_t
|a| ≤ 22,930 × weight ≤ 65,535 (d); add ≤ 65535
true product ≤ 1.5027e9, plus 65535
same
SAFE (exact; < INT32_MAX)
CUDA adm_csf.cu:170, adm_cm.cu:155,691 (int16 return); HIP adm_csf.hip:192, adm_cm.hip:209
i16_dst_val / csf_a / csf_r = (dst + add) >> 15 (h/v) or >> 17 (d), narrowed to int16
CUDA adm_csf.cu:171, AIM adm_cm.cu:844-847; HIP adm_csf.hip:193, AIM adm_cm.hip:658
flt = (4369*\|i16\| + 2048) >> 12, narrowed to int16 (csf_f and the AIM neighbour)
int → int16_t
≤ 4369·32,611/4096 = 34,786. It wraps negative once |i16| ≥ 30,720, i.e. weight·|band| ≥ 1.0066e9, i.e. an h/v weight ≥ 43,900 at |band| 22,930. Default weight 36,453 gives |i16| ≤ 25,509 and flt ≤ 27,212. The CPU narrows identically (integer_adm_kernels.h:488-489)
default: 27,212 fits
same
DEPENDS: on the CSF weight option, not on frame size. At default options it is SAFE. With a non-default h/v weight in [43,900, 46,603) (reachable through adm_csf_scale in Barten modes, or through viewing geometry) the int16 wraps negative, the threshold falls, and the excess can pass the ADR-1472 square budget (next table). Shared bit-for-bit with the CPU
CUDA adm_csf.cu:88, adm_cm.cu:86, 348, 660; HIP adm_csf.hip:129-130, adm_cm.hip:198
(i_rfactor * int64_t(a or r)) + 2^27 >> 28 (s123 CSF), result narrowed to int32
Fits only for x ≤ \(2^{30} - 1\) (h/v, shift 29) or x ≤ 1,518,500,249 (d, shift 30). With thr ≥ 0: x ≤ 22,930·46,603 = 1.0686e9 (h/v) and ≤ 22,930·65,535 = 1.5027e9 (d)
within budget at default and at every weight below the limit, provided thr ≥ 0
same
DEPENDS: SAFE while thr ≥ 0. It wraps when a negative int16 flt (row "flt narrowed to int16" above; h/v weight ≥ 43,900) lowers thr below 0. Not frame-size related
CUDA adm_cm.cu:218; HIP adm_cm.hip:150
x_sq, scales 1-3
int64_t → int32_t
x = |csf| − thr with thr ≥ −27 (flt ≥ −1 per sample, ADR-0155), so x ≤ 1,518,500,249 by the budget
**Scale-0 DLM and AIM row accumulator: per-thread partial, then the int64 row total, then (row_total + 2^(s-1)) >> s**
**int64_t**
**Term:** cm_cube(x) = ((x² + \(2^{28}\)) >> 29)·x >> (ceil(log2 Wb) − 4) for h/v, and the same with shifts 30 and −3 for d.; **Worst case:** thr = 0, which happens when ref == dis (DLM; a = 0 makes every csf_a and flt 0) or when ref is flat (AIM; r = 0, signal = t·w, csf_r = 0). Then x = |band|·weight.; **Terms per row:** cols ≈ 0.8·Wb, so row = R·16·T with R = cols/\(2^{\lceil \log_2 W_b \rceil}\) ∈ [0.47, 0.875] and T = x_sq·x. R is 0.750 at 16K and 0.800 at 8K DCI, 16384 wide and the cap.; **Content:** 4 pixel rows (0, 0, max, 0), so the vertical hi output is 27,411 (16-bit) or 27,304 (8-bit), and the column sign pattern is maximised exactly (scripts/dev/adm_cm_row_bound.py).
**Default weights:** h/v row max = 0.855·INT64_MAX (W = 15360), d 0.533.; **Non-default weight:** h/v ≥ 38,406 or d ≥ 60,965 → overflow (budget limit 46,603 gives 1.79·INT64_MAX).
**Default:** 0.912·INT64_MAX (8K DCI, 16384, 32768).; **Non-default:** h/v weight ≥ 37,588 or d ≥ 59,667 → overflow (1.91x at the limit).
**OVERFLOW@16K** (integrator relabel from the reader's DEPENDS: the envelope covers every W ≤ 15360 and W = 63–64 wraps at default options; the bound is not monotonic in W. Reader's note: depends on the CSF weight option and the band-width class, not on growth in frame size.); **At default options:** SAFE at 1080p (0.857), 16K and cap. It **overflows at W = 63–64 (Wb = 32, R = 0.875): 1.021·INT64_MAX at 16-bit, 1.009 at 8-bit, 1.018 at 10-bit.** W = 128 gives 0.973.; **With non-default CSF weights:** it overflows at every resolution, 16K and cap included.; The int64 wrap makes the band numerator negative or wrong, so powf ends in NaN or a wrong score. The CPU reference has the identical int64 inner (integer_adm_kernels.h:1017-1019, 1079-1099), where it is signed-overflow UB, so all three backends share the defect. The ADR-1472 budget bounds only the square, not the row sum.
CUDA adm_cm.cu:519; HIP adm_cm.hip:465
Scale-0 frame accumulator adm_cm[0][b] / adm_aim_cm[0][b] (atomicAdd on the unsigned long long alias)
int64_t
Σ rows of (row >> ceil(log2 Hb)); rows/\(2^{s}\) ≤ 0.81
≤ 0.81·row max
same
SAFE (integrator relabel from DEPENDS: region rows ≤ 0.875·\(2^{s}\) with s = ceil(log2 Hb) (max at Hb = 16 or 32), each term ≤ \(2^{63-s}\), so the sum stays below 0.875·\(2^{63}\) for any int64 row value; a wrapped row makes the value wrong but cannot make this sum wrap)
CUDA adm_cm.cu:310-329, 704-737; HIP adm_cm.hip:324, 609
CUDA adm_cm.cu:386, 781 (thread_accum), :281, 291 (warp_reduce, warp_sums, row_total); HIP adm_cm.hip:842 (AIM thread_accum), :632, and the DLM two-pass path :772 (per-pixel int32 excess) → :550 (temp_value) → :558
Scales 1-3 DLM and AIM row accumulator
int64_t (HIP scratch int32_t)
Term = ((x²+\(2^{29}\))>>30)·x >> ceil(log2 Wb), with T ≤ 2,147,483,647 · 1,518,500,249 = 3.261e18. Row ≤ cols/\(2^{\lceil \log_2 W_b \rceil}\) · T < T. HIP scratch value = excess ≤ 1.5185e9
≤ 3.27e18
≤ 3.27e18
SAFE at every size and weight (2.8x below \(2^{63}\))
cuda/integer_adm_cuda.c (1816 lines, read in full: 1-560 line by line, the rest by targeted reads and greps)
3 host rows, plus the stride sources of the DWT index row. No integer host sums: conclusion is in double (adm_dlm_terms, adm_aim_num), and tmp_res is memset every frame, so **no clip-level integer accumulator**
in the decouple rows (identical arithmetic, including the double → int32 defect)
hip/integer_adm/adm_dwt2.hip (381, full)
in the DWT rows (identical)
hip/integer_adm/adm_dwt2_rows.h (71)
in the index-helper row
adm_cm_accumulator.h (84, full)
adm_cm_round_row_total and adm_csf_den_round_row_total are folded into the CM and csf_den row rows; adm_cm_excess_s0 has its own row
integer_adm_kernels.h (1463, full read of the GPU-relevant contexts, kernels and shifts)
CPU reference. It is cited for the shared scale-0 row-total overflow (:1017-1019, :1079-1099), the int16 flt narrowing (:488-489) and the defined MIN before conversion (:380). No GPU-only accumulator
adm_csf_fixed_point.h (307, full)
none (double limits; adm_half_shift is uint32 and bounded). Supplies the weight budgets used above
adm_score.h (99)
none (double only)
adm_angle_flag.h (264; fp64 path)
none for G1 (CUDA and HIP pass int64 operands bounded by the decouple rows; the int64 reformulation is SYCL/Metal only)
adm_gain_limit.h (91, full)
1 row (not reachable from CUDA or HIP)
integer_adm.c
read only for the shifts (i4_dwt2_round, adm_cm_ctx_init) and the driver
Row counts by verdict, 42 rows in total:
verdict
rows
SAFE
38 (integrator: frame accumulator relabelled to SAFE)
OVERFLOW@16K
2 (the double → int32 gain conversion, size-independent; the scale-0 CM row total, relabelled by the integrator)
Repo: master at 571565a47. Read-only. Paths below are relative to core/src/feature/.
Envelope: 16K N = 132,710,400 (\(2^{26.98}\)), cap N = \(2^{30}\), W,H <= 32768. Samples <= 65535 (bpc 16) and in range (<= \(2^{\mathrm{bpc}} - 1\)) unless a row says otherwise.
Basis for integer VIF, checked against the CPU reference (integer_vif.h:39-46, integer_vif.c:175-228, 254-341, 406-456):
Every row of vif_filter1d_table sums to exactly 65536. Widths are {17, 9, 5, 3}, and the largest tap is 43,728 (scale 3 centre). So any tap sum is at most 65536 * (max input), whatever W and H are.
Scale-0 vertical shifts at bpc b: shift_VP = b, shift_VP_sq = 2(b-8). Scales 1-3 shift by 16. The horizontal shift is always 16. The GPU twins copy these values: integer_vif_cuda.c:572-589 and integer_vif_hip.c:334-341.
Inputs to scales 1-3 are the uint16 decimated planes: at most 65,535 for 16-bit input, at most 65,280 for 8-bit input.
The int64 accumulators hold the sums of one frame and one scale. They are cleared every frame (integer_vif_cuda.c:811, integer_vif_hip.c:689), so the frame count does not matter.
SAFE for in-range samples. Samples above \(2^{\mathrm{bpc}} - 1\) at bpc 9-15 wrap modulo \(2^{16}\) here, with the same cast as CPU integer_vif.c:205-206, 450-451, so CPU and CUDA agree and the accumulator itself never wraps.
Asides (CUDA integer VIF, no overflow, no verdict):
Possible undefined-behaviour shuffle, unverified.filter1d.cu:513/535: vif_hori_flush_accums() (warp_reduce, __shfl_down_sync(0xffffffff, …)) runs only inside if (y < h && x_start < w). When the scale width is not a multiple of 64, lanes of the edge warp skip the shuffle that the full mask names. The CUDA guide calls this undefined, and the values read from those lanes are undefined. This is not an overflow. test_integer_vif_cpu_cuda_parity (256x144, so a 32-pixel width at scale 3 with 16 active lanes) passes places=4, so current hardware returns benign values. It is not checked on a device here.
Idle horizontal blocks.filter1d_16_grid (integer_vif_cuda.c:620) sets grid_hori_x = ceil(w/128), but a block covers 256 pixels (vpt = 2), so about half the x-blocks do nothing. There is no wrap.
Unused buffer memory.data_sz also carves 2 stride_16 planes and 5 stride_32 planes (mu1, mu2, mu1_32, mu2_32, ref_sq, dis_sq, ref_dis) that no kernel reads. That is about 3.18e9 B at 16K and 2.58e10 B at cap of device memory, which matters for memory, not for overflow.
(accum_mu1 + (uint32_t)add_shift_VP) >> shift_VP, stored to uint32 tmp **without** the (uint16_t) cast that CPU (integer_vif.c:205-206, 450-451) and CUDA (filter1d.cu:603-610) apply
uint32_t
pre-shift <= 4,294,934,528
<= 65,535 for in-range samples
same
SAFE for in-range samples (out-of-range: next two rows)
hip/integer_vif/vif_statistics.hip:499-500, then 530-534
tmp.mu1 is not truncated: up to (65536*65535 + \(2^{b-1}\)) >> b = 8,388,480 (b=9) to 131,070 (b=15); 17 taps with sum 65536
5.5e11 (b=9) to 8.59e9 (b=15), wraps past \(2^{32}\)
same (independent of W and H)
DEPENDS on the input sample range: only malformed samples above \(2^{\mathrm{bpc}}\)-1. CPU and CUDA truncate to 16 bits at the store, so their sum stays <= 4,294,901,760; HIP wraps instead, giving HIP != CPU. In-range input is SAFE.
hip/integer_vif/vif_statistics.hip:506-507, then 329-336
(uint16_t)roundf(log2f((float)(0x8000 + i)) * 2048), unsigned i < 32768
float to uint16_t
entry max round(log2(65535)*2048) = 32,768
32,768
same
SAFE (< 65536)
float VIF: CUDA (float_vif_cuda) and HIP (float_vif_hip)¶
The accumulators are fp32 throughout: fvif_row_sum and fvif_sum_rows, as in the CPU. They are never converted to an integer, so they are out of scope. Only index, size and float-to-int math is listed. The only conversion is (float)((int32_t)exponent - 127).
file:line
variable / expression
type (exact C type)
what it sums (per-term max, term count)
worst-case bound at 16K
at cap
verdict
float_vif_gpu_common.h:136-139
fvif_term_index: ((size_t)x * height + y) * 2
size_t
term index
<= 2N = \(2^{28}\)
\(2^{31}\)
SAFE (size_t)
float_vif_gpu_common.h:186-190
(float)((int32_t)exponent - 127)
uint32_t to int32_t
exponent <= 255
[-127, 128]
same
SAFE
float_vif_gpu_common.h:272-275, 289-291
loop uint32_t x < width; (size_t)y * FVIF_TERM_FLOATS (+1)
uint32_t; size_t
row and column indices
<= 15,360 / 2*8640
<= \(2^{16}\)
SAFE
cuda/float_vif/float_vif_score.cu:48-50
plane[y * stride_bytes + x], plane + y * stride_bytes
float ADM: CUDA (float_adm_cuda) and HIP (float_adm_hip)¶
The accumulators are fp32: the per-row sum fadm_row_sum and the frame fold fadm_fold_rows. They are never converted to an integer. The device code has no float-to-int conversion; fadm_abs is a bit mask. Integer math is index and size arithmetic only. Band planes at scale 0 are half_w0 = (W+1)/2 by half_h0 = (H+1)/2, with buf_stride = (half_w0 + 3) & ~3.
file:line
variable / expression
type (exact C type)
what it sums (per-term max, term count)
worst-case bound at 16K
at cap
verdict
cuda/float_adm/float_adm_score.cu:66-68
plane[y * stride_bytes + x], plane + y * stride_bytes
All 27 files in group G2 were read in full or every grep hit was read in context. Rows per file (111 rows in total: 109 SAFE, 2 DEPENDS, 0 OVERFLOW@16K, 0 OVERFLOW@CAP-ONLY):
cuda/float_adm_cuda.c: 8
cuda/float_adm_cuda.h: none (declares the PTX symbol only)
cuda/float_adm/float_adm_device.h: none (macros and typedefs only)
cuda/float_adm/float_adm_score.cu: 7
cuda/float_vif_cuda.c: 3
cuda/float_vif_cuda.h: none (declares the PTX symbol only)
cuda/float_vif/float_vif_device.h: none (macros and typedefs only)
cuda/float_vif/float_vif_score.cu: 6
cuda/integer_vif_cuda.c: 7
cuda/integer_vif_cuda.h: none (struct types only; vif_accums is 7 int64_t, covered under the .cuh rows)
cuda/integer_vif/filter1d.cu: 20 (includes the warp_reduce/atomicAdd_int64 rows)
adm_float_reference.h: integer parts are adm_border_s and adm_pool_bands_s (whose body, get_noise_constant, uses w * h). Both are covered by the float ADM rows; 0 rows of its own.
Auxiliary files read to bound the terms:
integer_vif.c and integer_vif.h: the CPU reference shifts, casts, filter table and log2_32/64.
cuda/cuda_helper.cuh:125-139: warp_reduce and atomicAdd_int64.
cuda/cuda_tile_index.h: the reflect/clamp index helpers.
mem.h:26-31: ALIGN_CEIL, MAX_ALIGN = 32.
cuda/common.h:54-55: warp size 32, cache line 128.
adm_tools.c:66-99: get_noise_constant and adm_border_s.
adm_options.h:24: ADM_BORDER_FACTOR = 0.1.
src/meson.build:1553: CUDA builds with --std c++20.
test_integer_vif_cpu_cuda_parity.c: the fixture is 256x144.
G3a: motion family, CUDA + HIP twins (integer-overflow audit)¶
Repo: master at 571565a47. Read-only. Paths are relative to core/src/feature/ unless they start with core/.
Envelope: 16K N = 132,710,400 (W 15360, H 8640); 8K DCI N = 35,389,440; 1080p N = 2,073,600; cap W,H <= 32768, N <= \(2^{30}\) (core/src/picture.c:46); bpc 8..16 (core/src/picture.c:37-38).
horizontal h = (sum f_k v_k + 2^15) >> 16: |h| <= |v|max. In range |h| <= 65535 (reached by a uniform max-difference frame: ref 65535 / dis 0). Out-of-range worst |h| <= 8,388,480.
Frame SAD = sum over N pixels of |h|. In range: 16K 8,697,176,064,000 (\(2^{42.98}\)); 8K 2,319,246,950,400 (\(2^{41.08}\)); 1080p 135,893,376,000 (\(2^{36.98}\)); cap 70,367,670,435,840 (\(2^{46.0}\)). Out-of-range worst (bpc 9): 16K 1.113e15 (\(2^{49.98}\)); cap 9,007,061,815,787,520 (< \(2^{53}\) = 9,007,199,254,740,992).
The SAD accumulator is zeroed before every frame (CUDA cuda/integer_motion_sad_cuda.c:177, HIP hip/integer_motion_sad_hip.c:172), so there is no clip-level (cross-frame) integer accumulator in G3a; frames-to-overflow: not applicable.
motion SAD kernel (motion_cuda + motion_v2_cuda) - CUDA¶
file:line
variable / expression
type (exact C type)
what it sums (per-term max, term count)
worst-case bound at 16K
at cap
verdict
cuda/integer_motion_v2/motion_v2_score.cu:65
plane + ((ptrdiff_t)y * stride) then + x on const T *
ptrdiff_t byte offset, int x element index
row offset of a packed plane (pitch = W*bpp, integer_motion_sad_cuda.c:179)
863930720 + 215359 = 265,420,798 B
3276765536 + 232767 = 2,147,483,646 B
SAFE (\(2^{31}\) < \(2^{63}\); would be OVERFLOW@CAP-ONLY in int32 byte math)
frame counter of the last batch-boundary collect (boundary when index % 8 == 7)
n/a (frame count)
n/a (frame count)
DEPENDS (frame count: index \(2^{31} - 1\) is a batch boundary (== 7 mod 8), so last_batch_boundary = INT_MAX and flush's INT_MAX + 1 is signed-overflow UB; index >= \(2^{31}\) makes (int)index an implementation-defined narrowing. Needs >= 2,147,483,647 frames: 414 days at 60 fps, 828 days at 30 fps)
const int tile_ox = blockIdx.x * FM_BX - FM_RADIUS;
unsigned arithmetic (wraps to 4294967294u at block 0) then converted to int
tile origin
-2 .. 15342
-2 .. 32750
SAFE (unsigned wrap is defined; the int conversion is implementation-defined before C++20 and modular on nvcc; the motion SAD kernel and the HIP twin cast to int first)
cuda/float_motion/float_motion_score.cu:43-50
fm_mirror: 2 * (sup - 1) - idx
int
reflected index
<= 30716
<= 65532
SAFE as arithmetic; see Aside A (unclamped negative index)
integer_motion.h:53 (included by cuda/integer_motion_cuda.h:25; edge_16 is not called by any G3a twin)
accum += filter[k] * src[i_tap * stride + j_tap]
uint16_t * uint16_t promotes to int; uint32_t accum; int index
5 taps, term <= 2638665535 = 1,729,206,510 (< INT32_MAX, no UB); sum <= 6553665535
accum 4,294,901,760; index 132,710,399 (stride ~ W)
accum same (< UINT32_MAX); index 1,073,741,823 (stride ~ W)
SAFE (not used by the twins; recorded for completeness; index stays < \(2^{31}\) only while the stride in samples stays < 65536)
Aside A: CUDA float_motion tile index can go negative (out-of-bounds read, not an integer overflow)¶
cuda/float_motion/float_motion_score.cu:43-50 (fm_mirror) reflects once and does not clamp, and fm_load_tile (:83-90) loads a fixed 20x20 tile for every block, padding threads included. The largest tile index on an axis of extent sup is 16 * (ceil(sup / 16) - 1) + 17; it maps below zero when that exceeds 2 * (sup - 1), i.e. for sup in {3..9, 17}. Width or height 3 gives index -13 (13 rows before the ref_in allocation when it is the row axis), 17 gives -1. Those samples feed no valid output (the consumed taps stay in range from 3x3 up), but the load reads outside the buffer. The HIP twin clamps (hip/float_motion/float_motion_score.hip:101-104, ADR-1381), and the CUDA motion SAD kernel clamps (cuda/cuda_tile_index.h). core/test/test_cuda_motion_tiny_frames.c:67 covers 3x3 and 17x17 for motion_cuda and motion_v2_cuda only, not float_motion_cuda. The init guard only refuses below 3x3 (cuda/float_motion_cuda.c:262).
Aside B: CPU reference row accumulator (outside G3a, for the CPU group)¶
integer_motion.c:193 and :243 hold a row's SAD in uint32_t row_sad. In range it is <= 32768 * 65535 = 2,147,450,880 (< \(2^{32}\)), so it is safe up to the cap. If a 16-bit container carries samples above \(2^{\mathrm{bpc}} - 1\) (bpc 9..15; sample range against bpc was not checked in this audit), abs(h) reaches 8,388,480 at bpc 9 and the row sum wraps from W >= 513 (bpc 10: abs(h) <= 4,194,240, W >= 1025). The GPU twins add in 64 bits and would then differ from the CPU.
All paths under core/src/feature/. Every file was read in full, or grepped with the brief's pattern list and every hit read in context (for the 600-900-line host files).
file
rows
cuda/float_motion_cuda.c
3 (+ float_motion_sad.h call)
cuda/float_motion_cuda.h
none (PTX symbol only)
cuda/float_motion/float_motion_score.cu
6 (+ Aside A)
cuda/integer_motion_cuda.c
6
cuda/integer_motion_cuda.h
none (includes integer_motion.h; see Shared headers)
cuda/integer_motion_sad_cuda.c
3
cuda/integer_motion_sad_cuda.h
none (declarations; sad documented as uint64)
cuda/integer_motion_v2_cuda.c
4
cuda/integer_motion_v2_cuda.h
none (PTX symbol only)
cuda/integer_motion_v2/motion_v2_score.cu
13
cuda/cuda_tile_index.h (helper, read for index math)
folded into the motion_v2_score.cu reflect row
hip/float_motion_hip.c
6
hip/float_motion_hip.h
none (HSACO symbol only)
hip/float_motion/float_motion_rows.h
4 (+ float_motion_sad.h call)
hip/float_motion/float_motion_score.hip
3
hip/integer_motion_hip.c
4
hip/integer_motion_sad_hip.c
3
hip/integer_motion_sad_hip.h
none (declarations)
hip/integer_motion_v2_hip.c
3 (2 shared with integer_motion_hip.c rows)
hip/integer_motion_v2_hip.h
none (HSACO symbol only)
hip/integer_motion_v2/motion_v2_score.hip
12
hip/hip_tile_index.h (helper, read for index math)
folded into the reflect rows
float_motion_sad.h
1 (used by both CUDA and HIP float twins)
integer_motion.h
1 (edge_16, not called by the twins; filter[] taps match the twins')
motion_tools.h
none (float filter constants only)
motion_blend_tools.h
none (double blend only)
integer_motion.c (CPU reference, tap values and shifts)
Aside B only
core/src/hip/shared_frame.{h,c} (checked: HIP luma planes are packed, so the packed pitch the SAD kernel gets is right)
Repo: master @ 571565a47. Paths are relative to core/src/feature/. Envelope: 16K N = 132,710,400 (\(2^{26.98}\)); cap N = \(2^{30}\). Max 16-bit term \(65535^{2}\) = 4,294,836,225. Its fp32 square is 4,294,836,224 (the twins' float-square terms). Both values are below \(2^{32}\). Row W <= 32768 = \(2^{15}\).
256 terms; in contract each <= \((2^{\mathrm{bpc}} - 1)^{2}\) <= \(4095^{2}\) = 16,769,025
256*\(4095^{2}\) = 4,292,870,400
same (per block, not size dependent)
DEPENDS: safe only if every sample <= \(2^{\mathrm{bpc}} - 1\). The margin to UINT32_MAX is 2,096,895. No sample-range check exists in picture.c or libvmaf.c (grepped). A 10- or 12-bit picture whose uint16 codewords differ by >= 4096 across all 256 pixels of a row segment wraps the block sum (256*\(4096^{2}\) = \(2^{32}\)) with no error. The CUDA twin (uint64) and the CPU (double) do not wrap
Note: the kernel doc comments (float_psnr_score.hip:107-111, 139-141) describe a partials[2 * block] / [2 * block + 1] layout. The code writes one u64 per block (partials[block_idx]), and the host reads one u64 per block. The comments are out of date. This does not cause an overflow.
one per-plane frame SSE per frame, each <= N*\(65535^{2}\)
per frame 5.70e17 (\(2^{58.98}\)); u64 wraps on frame 33
per frame 4.61e18 (< \(2^{62}\)); wraps on frame 5
DEPENDS (frame count): 16-bit max-diff wraps on frame 2072 at 1080p (\(2^{52.98}\)/frame), frame 122 at 8K DCI (\(2^{57.08}\)), frame 33 at 16K, frame 5 at cap. Max-diff wrap frame at 12 bit: 530,502 / 31,085 / 8,290 / 1,025. At 10 bit: 8.5M / 498,076 / 132,821 / 16,417. At 8 bit: 136.8M / 8.0M / 2.1M / 264,205. The wrap gives no error, and the apsnr_* aggregate comes out too high. The CPU's integer_psnr.c:64,211,249 uses the same uint64_t, so the twin and the CPU wrap identically
cuda/integer_psnr_cuda.c:395
s->apsnr_n_pixels[p] += (uint64_t)height * width
uint64_t
N per frame
\(2^{64}\)/N = 1.39e11 frames
\(2^{34}\) frames (1.7e10)
SAFE (>= \(2^{34}\) frames at cap, about 18 years at 30 fps)
DEPENDS (frame count): same derivation and wrap frames as CUDA integer_psnr_cuda.c:392-394 (1080p 2072, 8K 122, 16K 33, cap 5 at 16-bit max-diff). The CPU's integer_psnr.c uses the same type
grid (frame_h, 2) x 256 lanes for row_totals / row_units
unsigned launch dims
gridDim.x = H
8640
32768
SAFE (x-dim limit \(2^{31} - 1\))
(double)sums_host[k] at integer_moment_cuda.c:329-332 converts uint64 to double. It is not integer math. Above \(2^{53}\), exactness is a parity question that ADR-1497 handles; it does not cause a wrap.
HIP (hip/float_moment/moment_score.hip, hip/float_moment_hip.c)¶
none (integer parts: 1 << (bpc - 8u) <= 256 at :390/:392 and the size_t xyz loop at :410; e.k at :296 is an ldexp exponent from ff_math.h, not an accumulator)
ordered_sum.h (read in part, :97-118, :160-192, :245-262, for the units bounds)
counted under float_moment_sum.h
integer_psnr.c (CPU, grepped for parity only)
not counted (uses the same uint64_t apsnr types at :64-65, :211-212, :249-250)
G4a: integer-overflow audit of SSIM (integer and float) and MS-SSIM, CUDA and HIP twins¶
Repo: master at 571565a47. This audit was read-only. Every path below is relative to core/src/feature/. Envelope: 16K N = 132,710,400; cap N = \(2^{30}\) (W, H <= 32768); 16-bit samples, 4:4:4.
Derived term maxima (shared by every integer-SSIM row)¶
Integer Gaussian gaussian_filter_init(1.5, 5) (integer_ssim.c:60-96) has kernel_len = floor(1.5sqrt(-2 ln(sqrt(pi/2)1.5/256))) = floor(4.70) = 4, so it has 9 taps [2,9,28,55,68,55,28,9,2] with a sum of exactly 256. The CUDA kernel (ISSIM_KERNEL, integer_ssim_score.cu:70) and the HIP kernel (integer_ssim_score.hip:84) copy it.
Horizontal moments (a window of at most 9 taps, at most 65535 per sample):
mux and muy: at most 256*65535 = 16,776,960 (\(2^{24}\)).
x2, xy and y2: at most 256*\(65535^{2}\) = 1,099,478,073,600 (< \(2^{40}\)).
w: at most 256.
Vertical moments (9 taps over the horizontal moments):
mux and muy: at most 65536*65535 = 4,294,901,760 (< \(2^{32}\)).
x2, xy and y2: at most 65536*\(65535^{2}\) = 281,466,386,841,600 (< \(2^{48}\)).
w: at most 65,536 (\(2^{16}\)).
The bounds above are local to one window and do not depend on the frame size. Frame-level weight sums are Σ m.w ≤ 65536*N: 8,697,308,774,400 (\(2^{42.98}\)) at 16K and \(2^{46}\) at the cap.
The float_ssim decimation is a fixed-point sum in units of \(2^{-52}\) over a scale x scale window, with tap = fl(1/\(\mathrm{scale}^{2}\)) and sample ≤ 65535/256 = 255.996. A window therefore sums to at most about 256, which is ≤ 255.996*\(2^{52}\) ≈ \(2^{59.99998}\) < \(2^{60}\). This bound holds for every scale. The cap of 128 on the scale (SSIM_MAX_EXACT_SCALE / VMAF_HIP_SSIM_MAX_EXACT_SCALE) is there for exactness, not for overflow.
No accumulator in G4a spans more than one frame (clip level). Every sum is a local in collect() or a per-frame device buffer, so no frames-to-overflow figure applies.
SAFE (both < \(2^{32}\); \(\mathrm{peak}^{2}\) has a margin of 131,070 below UINT32_MAX; the shared double helper vmaf_metal_ms_ssim_max_db is not used here)
Device memory at the cap: the CUDA and HIP integer SSIM twins allocate 6 int64 moment planes plus 1 double term plane, 7 x \(2^{33}\) B = 56 GiB. Every size is computed in size_t and passed to size_t allocator parameters (cuda/common.h:148, 273; cuda/kernel_template.h:187-188; hip/kernel_template.h:125). At that size the allocation fails cleanly; no size wraps.
**warp_reduce(int64_t) (cuda_helper.cuh:129-130, outside this group):** the helper left-shifts the shuffled signed high half (<< 32), which is UB for negative values before C++20. That is unreachable here: the weights are ≥ 0 and below \(2^{23}\), so the high half is 0.
Grid dimensions: the largest gridDim.y in the group is ceil(32768/8) = 4096, well under 65535. No grid expression wraps.
**picture_copy() (picture_copy.cpp, outside this group):** the CUDA and HIP MS-SSIM hosts call it. It walks rows with ptrdiff_t pointer increments and unsigned loop counters, and has no integer product.
Tightest margin in the group: the signed int index at cuda/integer_ms_ssim/ms_ssim_score.cu:154 (\(2^{30} - 1\) against INT32_MAX). Next is peak * peak in unsigned at cuda/integer_ms_ssim_cuda.c:306 and hip/integer_ms_ssim_hip.c:620 (131,070 below UINT32_MAX). Both are safe under the envelope.
G4b: SSIMULACRA 2, CUDA + HIP twins, ordered-sum header (integer-overflow audit)¶
Repo: master @ 571565a47. Read-only; no build, no device run. All paths below are relative to core/src/feature/.
Envelope inputs taken from the code: - VMAF_PIC_DIM_MAX 32768u (core/src/picture.c:46, checked at :185), VMAF_PIC_BPC_MIN/MAX 8u/16u (picture.c:37-38, checked at :181). - Both twins refuse w or h < 8 (cuda/ssimulacra2_cuda.c:788, hip/ssimulacra2_hip.c:931). That makes pixels >= 64 and chunks >= 1. - Chunk = 256 lanes x 4 px = 1024 px (SS2C_CHUNK_PIXELS, SS2H_CHUNK_PIXELS). A batch is 1024 chunks (SS2C_BATCH, SS2H_BATCH). - Scale pyramid: scale_w[i] = (scale_w[i-1]+1)/2. The pyramid stops before the first scale with a side below 8. Chunks are counted per plane (one channel); ceil applies.
scale
16K W x H
N
chunks
cap W x H
N
chunks
0
15360x8640
132,710,400
129,600
32768x32768
\(2^{30}\)
\(2^{20}\) = 1,048,576
1
7680x4320
33,177,600
32,400
\(16384^{2}\)
\(2^{28}\)
262,144
2
3840x2160
8,294,400
8,100
\(8192^{2}\)
\(2^{26}\)
65,536
3
1920x1080
2,073,600
2,025
\(4096^{2}\)
\(2^{24}\)
16,384
4
960x540
518,400
507
\(2048^{2}\)
\(2^{22}\)
4,096
5
480x270
129,600
127
\(1024^{2}\)
\(2^{20}\)
1,024
Device buffers are sized for scale 0. That scale has the most chunks.
The brief assumed fp64 term planes (6 terms x 3 channels x N x 8 B). These do not exist. Every kernel recomputes the six fp64 terms from the five blurred fp32 planes and the two XYB fp32 planes (ss2c_terms / ss2h_terms). Per scale, the code stores only three per-(channel, sum, chunk) arrays: - chunk_sums: double, 8 B - plan: int16, 2 B - units: 2 x int64, 16 B
Neither twin has a clip-level (cross-frame) integer accumulator. Each frame's score is appended to the collector on its own, so no row is DEPENDS.
ordered_sum (shared, ordered_sum.h, ADR-1433). Used by the CUDA and HIP twins¶
The increment of one term in binade e is the term divided by u = \(2^{e-52}\), rounded. In integers this is mantissa >> (e - ex): - mantissa is in [\(2^{52}\), \(2^{53}\)). - shift = 0 gives an increment of at most \(2^{53}\) - 1. - shift >= 1 gives at most \(2^{52}\) after rounding up. - shift > 53 gives 0. - A term that does not fit (higher binade, negative, inf or NaN) gives VMAF_ORDSUM_UNFIT = \(2^{54}\).
Every increment is >= 0 (round_shifted returns whole + {0,1}).
The composition vmaf_ordsum_then saturates at \(2^{54}\). The int64 range therefore does not depend on the term count, the chunk size or N. For comparison, an uncapped sum would wrap after 512 UNFIT terms (512 x \(2^{54}\) = \(2^{63}\)). With 1024 terms at the maximum increment it would reach 1024 x (\(2^{53} - 1\)) = \(2^{63}\) - 1024, which still fits.
The cap never changes a result. A capped value (>= \(2^{54}\)) always fails total > VMAF_ORDSUM_BINADE_END in add_chunk, and that chunk is then added term by term.
file:line
variable / expression
type (exact C type)
what it sums (per-term max, term count)
worst-case bound at 16K
at cap
verdict
ordered_sum.h:180
whole = (int64_t)(mantissa >> shift)
int64_t (from uint64_t)
mantissa < \(2^{53}\), shift in [1,53] (guards :217, :220)
vmaf_ordsum_then: a.even + (b.odd\|b.even), a.odd + ..., then vmaf_ordsum_cap
int64_t
both operands in [0, 2^54]: an increment (<= \(2^{53} - 1\)) or UNFIT (\(2^{54}\)), or an earlier then result capped to \(2^{54}\). Term count: 4 per lane plus an 8-level lane tree = 1024 per chunk; it does not matter because of the cap
pre-cap <= \(2^{55}\); stored <= \(2^{54}\)
same (N-independent)
SAFE (\(2^{55}\) < \(2^{63}\); saturating)
ordered_sum.h:320-321
total = m + ((m & 1) ? units.odd : units.even)
int64_t
m in [\(2^{52}\), \(2^{53}\)) (running sum's significand), units <= \(2^{54}\)
e in [-900,900] (binade_bits(s) == plan, :318); total in [\(2^{52}\), \(2^{53}\)] (:322)
exponent field in [123, 1924] < 2047; no carry into the sign bit
same
SAFE
ordered_sum.h:125, :144
(bits >> 52) & 0x7ff; (int)field - 1023
unsigned, int
11-bit field
[-1023, 1024]
same
SAFE
ordered_sum.h:92-97 to cuda/.../ssimulacra2_device.cu:394, :401, hip/.../ssimulacra2_device.hip:654, :662
plan int -> (short) -> int16_t
short / int16_t
plan values: [-900, 900], PLAN_ZERO 32767, PLAN_TERMS -32768
all values in int16 range (endpoints exact)
same
SAFE (no narrowing loss)
SSIMULACRA 2: CUDA twin (cuda/ssimulacra2_cuda.c, cuda/ssimulacra2_cuda.h, cuda/ssimulacra2/ssimulacra2_device.cu, cuda/ssimulacra2/ssimulacra2_blur.cu)¶
There are no integer accumulators besides the ordered-sum units. The rows below are size products, index and offset math, counters and grid dims.
Input pictures are cuMemAllocPitch planes. pitch is carried as size_t (ssimulacra2_cuda.h:71, set at ssimulacra2_cuda.c:441 from (size_t)stride) and multiplied in size_t (device.cu:82).
file:line
variable / expression
type (exact C type)
what it sums (per-term max, term count)
worst-case bound at 16K
at cap
verdict
CUDA ssimulacra2_cuda.c:327
peak = (float)((1u << bpc) - 1u)
unsigned
bpc in [8,16] (picture.c:181)
65,535
65,535
SAFE (< \(2^{16}\))
CUDA ssimulacra2_cuda.c:349-350, :359-360
(width + sh) >> sh; (scale_w[i-1] + 1u) / 2u
unsigned
dims <= 32768
<= 15,361
<= 32,769
SAFE
CUDA ssimulacra2_cuda.c:370-373
ss2c_chunks: (unsigned)((pixels + 1023) / 1024)
size_t -> unsigned
N / 1024 per plane
129,600
\(2^{20}\)
SAFE (\(2^{20}\) < \(2^{32}\))
CUDA ssimulacra2_cuda.c:380
(size_t)cw * (size_t)ch
size_t
W x H
132,710,400
\(2^{30}\)
SAFE
CUDA ssimulacra2_cuda.c:454-457
yuv grid gx = (W+15)/16, gy = (H+7)/8, z = 2
unsigned
gridDim.y limit 65535
gx 960, gy 1,080
gx 2,048, gy 4,096
SAFE (gy <= 65535)
CUDA ssimulacra2_cuda.c:469
unsigned pixels = scale_w[scale] * scale_h[scale]
unsigned x unsigned
W x H of the scale; passed as a kernel arg unsigned pixels
132,710,400
\(2^{30}\) = 1,073,741,824
SAFE (\(2^{30}\) < \(2^{32}\); also < INT32_MAX)
CUDA ssimulacra2_cuda.c:471-472
xyb grid gx = (pixels + 255) / 256, gy = 2
unsigned
gridDim.x limit \(2^{31} - 1\)
518,400
4,194,304 = \(2^{22}\)
SAFE
CUDA ssimulacra2_cuda.c:498-503
blur grids gh = ceil(H/32) (x), gv = ceil(W/64) (x); y = 5 jobs, z = 3
w = (int)a.width; (int)(chunk * 32) - lead; n, left, col
int (from unsigned)
chunk <= ceil((W + radius - 1) / 32)
<= ~15,400
<= ~32,800
SAFE (<< \(2^{31}\))
CUDA ssimulacra2_blur.cu:136
(chunk + 1u) * 32 + lane
unsigned
tile column
<= ~15,450
<= ~32,850
SAFE
CUDA ssimulacra2_blur.cu:150
(size_t)(row0 + r) * a.width + (unsigned)col
size_t
pass-buffer offset
< 132.7 M
< \(2^{30}\)
SAFE
CUDA ssimulacra2_blur.cu:163, :167
col = blockIdx.x * 64 + threadIdx.x; (size_t)blockIdx.z * a.width * a.height + col
unsigned; size_t
—
< 3 x 132.7 M
< 3 x \(2^{30}\)
SAFE
CUDA ssimulacra2_blur.cu:170-182
h = (int)a.height; n, left, right; (size_t)left * w, (size_t)n * w
int; size_t
—
n <= 8,644; offsets < 132.7 M
n <= 32,772; offsets < \(2^{30}\)
SAFE
SSIMULACRA 2: HIP twin (hip/ssimulacra2_hip.c, hip/ssimulacra2_hip.h, hip/ssimulacra2/ssimulacra2_device.hip)¶
The HIP twin does not read the picture's device pitch. submit() packs every plane into pinned staging with row_bytes = plane_w x bytes_per_sample (ss2h_stage_plane, :895-907) and copies it into a flat hipMalloc per plane. The kernel therefore indexes sy * plane_w + sx in size_t.
SSIMULACRA 2: shared headers, integer parts (ssimulacra2_math.h, ssimulacra2_eotf_lut.h, ssimulacra2_score.h, ssimulacra2_pixel_format.h)¶
These are compiled into both device modules (CUDA: device.cu:39-43; HIP: device.hip:54-58).
file:line
variable / expression
type (exact C type)
what it sums (per-term max, term count)
worst-case bound at 16K
at cap
verdict
ssimulacra2_math.h:77
u.i = u.i / 3u + 0x2a5137a0u (cube-root seed)
uint32_t
x > 0 (:66). Largest bit pattern: +NaN 0x7fffffff; -NaN passes x <= 0 as false, up to 0xffffffff
0xffffffff / 3 + 0x2a5137a0 = 0x7fa68cf5
same (input-value bound, not N)
SAFE (< \(2^{32}\), unsigned; no wrap)
ssimulacra2_math.h:107-111
idx = x * 1023.0f; i = (int)idx; lut[i], lut[i + 1]
float -> int
x in (0, 1) after the guards at :101, :104. Largest x = 1 - \(2^{-24}\) gives fl(x * 1023) = 1022.99994 (checked with numpy float32), so i <= 1022 and i + 1 <= 1023 = size - 1
i in [0, 1022]
same
SAFE (index in bounds). Aside: NaN passes both guards, and (int)NaN is UB on the host. It is not reachable: the inputs are integer samples times finite coefficients, clamped by ss2c_clampf / ss2h_clampf
ordered_sum.h (full, 344 lines): 10 rows. Covers the binade increments, round_shifted, then composition and cap, add_chunk total, from_units_bits exponent and the int16 plan narrowing.
cuda/ssimulacra2_cuda.c (850 lines): 17 rows. Covers host size products, chunk count, grids, readback and the unsigned pixels kernel arg.
cuda/ssimulacra2_cuda.h (120 lines): none of its own. It holds the types used by the rows above: pitch is size_t, pixels is size_t, chunks is unsigned, units is int64_t, plan is int16_t.
ssimulacra2_eotf_lut.h (1061 lines; integer parts only): none.
ssimulacra2_score.h (55 lines): none.
ssimulacra2_pixel_format.h (51 lines): none.
Cross-checked, not in the group: core/src/picture.c:37-38, :46, :181-185 (bpc 8-16, dims <= 32768); core/src/cuda/common.h:148, :273 (vmaf_cuda_buffer_alloc and vmaf_cuda_buffer_host_alloc take size_t).
Verdict totals: 88 rows (shared ordered_sum 10, CUDA 42, HIP 34, math.h 2). SAFE 88, OVERFLOW@16K 0, OVERFLOW@CAP-ONLY 0, DEPENDS 0. Five more entries are listed as "none" (lut decl, score.h, pixel_format.h, and the two struct headers).
Tightest margin in the group: CUDA ssimulacra2_xyb, which indexes with the 32-bit unsigned2u * pixels + i (device.cu:315, :349). At cap the index reaches 3,221,225,471, which is below \(2^{32}\) = 4,294,967,296 (headroom \(2^{0.42}\)). It is safe only because N <= \(2^{30}\) and the type is unsigned: as int it would overflow at cap, and with a fourth plane it would wrap. The HIP twin uses size_t for the same index.
G5: CAMBI and PSNR-HVS, CUDA and HIP twins: integer-overflow audit¶
Repo: master at 571565a47. Read-only; nothing built or run on a device. Paths are relative to core/src/feature/.
Envelope: 16K = 15360x8640, N = 132,710,400 (\(2^{26.98}\)). Cap = W,H <= 32768, N <= \(2^{30}\).
Window guard.vmaf_cambi_check_window_fits_lut() (cambi.c:1833) runs in the CUDA init (integer_cambi_cuda.c:400) and the HIP init (integer_cambi_hip.c:393). It rejects any window w with \(w^{2}\) >= 4226, for both the encode window and the source window. So w <= 65, pad <= 32, and no window holds more than 65*65 = 4225 pixels, at any resolution.
What the guard means for resolution. At window_size >= 15 (the option minimum), the guard rejects W+H >= 26,400, or W+H >= 52,400 with the high-res speed-up. So W=H=32768 never runs CAMBI.
Largest reachable per-scale n is about 1.742e8 (\(2^{27.38}\)).
Largest reachable preprocess/mask plane is about 6.86e8 (\(2^{29.35}\)), with the speed-up.
16K runs only with window_size 15 or 16, or window_size <= 32 with the speed-up.
The "cap" column still uses n = \(2^{30}\), the conservative value. No verdict changes because of this.
c-value maximum. c = w * p0 * pm / (p0 + pm). Weight w <= 9 (g_contrast_weights, cambi.c:97). p0 and pm count pixels of two different levels in one window, so p0 + pm <= 4225. That gives c <= 9 * 2112 * 2113 / 4225 = 9506.25 (\(2^{13.22}\)). The fixed point is c * \(2^{24}\) < \(2^{37.22}\). Call this F.
No clip-level accumulators. Every integer sum is reset each frame (memset of the select/results/frame state), and the score leaves as a per-frame double. Frames-to-overflow does not apply.
CUDA (cuda/integer_cambi/cambi_score.cu, cuda/integer_cambi_cuda.c, cuda/integer_cambi_cuda.h)¶
Window count of one level for one column. Runs are added and removed one thread at a time, in sequence: remove row y-pad-1, then add row y+pad. The true count stays in 0..4225.
Bit-depth guard. Every implementation rejects bpc > 12: CUDA cuda/integer_psnr_hvs_cuda.c:107, HIP hip/integer_psnr_hvs_hip.c:401, CPU third_party/xiph/psnr_hvs.c:440. Samples are therefore <= 4095, and 16-bit input cannot reach the kernels.
DCT intermediates. od_bin_fdct8 is applied to columns, then to rows. It is bounded with affine arithmetic over the 64 samples in [0, M]: the exact linear part, plus +/-0.5 for every rounding shift.
12-bit:
max |intermediate| = 32,795
max |t*K + r| = 379,769,216 (\(2^{28.50}\); row pass, K = 11585)
max |AC coefficient| = 16,414, so \(\mathrm{AC}^{2}\) = 2.69e8 (\(2^{28.0}\))
max |coefficient| = 32,795
16-bit (rejected input), for reference: max |t*K + r| = 6.07e9 (\(2^{32.5}\)) and \(\mathrm{AC}^{2}\) = 6.87e10 (\(2^{36}\)), so both would overflow int32. The products first exceed \(2^{31}\) at 15-bit samples, which leaves 2.5 bits of margin above the guard.
Block counts. Step 7, 8x8 blocks, so blocks per plane = ((W-8)/7+1) * ((H-8)/7+1).
16K: 2,707,396 per plane; 4:4:4 total over 3 planes 8,122,188.
Cap: 21,911,761 per plane; 4:4:4 total 65,735,283 (\(2^{25.97}\)).
No clip-level accumulators. Per-frame float sums are converted to double scores.
CUDA (cuda/integer_psnr_hvs/psnr_hvs_score.cu, cuda/integer_psnr_hvs_cuda.c, cuda/integer_psnr_hvs_cuda.h)¶
file:line
variable / expression
type (exact C type)
what it sums (per-term max, term count)
worst-case bound at 16K
at cap
verdict
CUDA cuda/integer_psnr_hvs/psnr_hvs_score.cu:144-162
od_bin_fdct8 lifting products (t * K + r) >> s, K <= 21895
int
One product per lifting step, column and row pass; 12-bit samples
<= 379,769,216 (\(2^{28.50}\)); per block, so independent of resolution
same
SAFE (\(2^{28.5}\) < \(2^{31}\); bpc <= 12 guard. 16-bit would reach \(2^{32.5}\), but init rejects it.)
CUDA cuda/integer_psnr_hvs/psnr_hvs_score.cu:129-143,153-156
Butterflies t0 - t1, t4 + t5, t6 + t7, t0 += t6h
int
Intermediates
<= 32,795
same
SAFE
CUDA cuda/integer_psnr_hvs/psnr_hvs_score.cu:110-113
od_dct_rshift ((unsigned)a >> (32 - b)) + (unsigned)a, then (int) and >> b
unsigned int (deliberate modular wrap that adds the sign bit)
abs(a) <= 32,795
Exact truncation toward zero (equals OD_UNBIASED_RSHIFT32)
same
SAFE (intentional wrap; the unsigned-to-int conversion is two's complement on nvcc)
CUDA cuda/integer_psnr_hvs/psnr_hvs_score.cu:311-312
coefficient * coefficient (mask energy; DC skipped)
int
AC coefficient squared
<= 16,\(414^{2}\) = 269,419,396 (\(2^{28.0}\))
same
SAFE (12-bit guard; 16-bit would reach \(2^{36}\), rejected)
CUDA cuda/integer_psnr_hvs/psnr_hvs_score.cu:339
abs(block[ref] - block[dist])
int
Coefficient difference
<= 65,590
same
SAFE
CUDA cuda/integer_psnr_hvs/psnr_hvs_score.cu:357-360
HIP hip/integer_psnr_hvs/psnr_hvs_score.hip:437,482
b = chunk * 256u + tid
unsigned
<= total_blocks + 255
8.1e6
6.6e7
SAFE
HIP hip/integer_psnr_hvs/psnr_hvs_score.hip:462-471
hvs_scan_prefix_hip running += chunk_totals[c] for c < limit, where limit = num_chunks < 32768u ? num_chunks : 32768u -> chunk_offsets[c], header->total_terms
uint32_t sum; unsigned loop cap 32,768 chunks
**The sum is safe (<= \(2^{31.97}\), as CUDA). The cap is not.** num_chunks = ceil(total_blocks / 256), and only the first 32,768 chunks (8,388,608 blocks) get an offset. Chunks at or past 32,768 read whatever d_scratch held: hipMalloc garbage, or the previous frame's value. total_terms counts only the first 32,768 chunks.
4:4:4: 256,779 chunks; luma only: 85,593. Both > 32,768.
**OVERFLOW@CAP-ONLY** (a count cap rather than an arithmetic wrap). It fails once total_blocks > 8,388,608, e.g. 16384x8640 4:4:4 (8,662,680 blocks) or luma/4:0:0 above about 20.3K x 20.3K. hvs_compact_hip then writes packed_terms + chunk_offsets[chunk] + intra from unset offsets: out-of-bounds device writes, a wrong total_terms readback, and NaN or wrong plane scores. The CUDA twin (psnr_hvs_score.cu:443) loops over every chunk and is not affected.
HIP hip/integer_psnr_hvs/psnr_hvs_score.hip:498-519
(size_t)first_block * 64; end - start (uint32_t); (size_t)total_terms * sizeof(float)
size_t / uint32_t
Host plane ranges. If the scan-prefix cap row has fired, end - start can underflow; vmaf_psnr_hvs_plane_score_compacted then returns NaN because n_compact > n_terms.
<= \(2^{28.95}\)
<= \(2^{31.97}\)
SAFE (inside the cap)
Shared host tail (psnr_hvs_score.c, the implementation behind psnr_hvs_score.h; both twins call it)¶
SpEED does not run on the full plane. speed_internal_init_dimensions() (speed_internal.c:65-96) computes the operating geometry, and si_gpu_fill_geometry() (speed_internal.c:833-858) narrows it to uint32_t in SpeedGpuGeometry:
scaled = lround(src * speed_prescale). The option speed_prescale is in [0.1, 4.0] on all four extractors (cuda/speed_chroma_cuda.c:87-88, cuda/speed_temporal_cuda.c:88-89, hip/speed_chroma_hip.c:89-90, hip/speed_temporal_hip.c:92-93). **The scaled plane can therefore be up to 16x the picture plane.**
down = scaled >> 4, trunc = down / 5 * 5, blocks = (trunc_w/5)*(trunc_h/5), sub = trunc - 4.
chroma 4:4:4 uses the luma size (speed_chroma_dimensions(), speed_internal.h:100). The temporal extractor uses luma.
Every bound below was derived at two prescales: p=1 (default) and p=4 (option max, the worst case). Values come from the formulas above, computed in python.
point
scaled WxH
down
trunc
blocks
sub_w*sub_h
16K p=1
15360x8640
960x540
960x540
20,736
956*536 = 512,416
16K p=4
61440x34560
3840x2160
3840x2160
331,776
3836*2156 = 8,270,416
cap p=1
32768x32768
2048x2048
2045x2045
167,281
\(2041^{2}\) = 4,165,681
cap p=4
131072x131072
8192x8192
8190x8190
2,683,044
\(8186^{2}\) = 67,010,596
\(32740^{2}\) p=4 (inside cap)
\(130960^{2}\)
\(8185^{2}\)
\(8185^{2}\)
2,679,769
\(8181^{2}\) = 66,928,761 (odd)
Channels: chroma 4, temporal 2. Raw plane bytes: 16K, 16-bit: 265,420,800; at the cap, \(32768^{2}\)*2 = \(2^{31}\).
There are no integer pixel accumulators in SpEED. Every sum over pixels (means, covariance, filters, solve, variance) is fp32 or an fp32 pair. There are no atomics, __shfl, warp/wave integer reductions or histograms in any G6 file. The integer surface is sizes, indices, grid dims, the uint32 tail layout, int-to-float conversions of counts, float-to-int coordinate conversions and the uint64 singular tally.
means divisor static_cast<float>(g.sub_w * g.sub_h)
uint32_t product -> float
sub product
8,270,416 (exact)
67,010,596 (rounded above \(2^{24}\))
SAFE (no wrap; the CPU compute_mean at speed.c:775 is float / size_t, which also converts the count to float, so the rounding is identical)
cuda/speed/speed_score.cu:1219-1220
covariance divisor static_cast<float>(g.sub_w * g.sub_h) passed to ff_div_to_float
uint32_t product -> float
sub product, must be exact
p4 max 8,270,416 < \(2^{24}\): exact
p > ~2.0 gives count > \(2^{24}\). \(32740^{2}\) at p4: 66,928,761 (odd) is not an fp32 value
OVERFLOW@CAP-ONLY (fp32 exact-integer range, not UB). The CPU compute_covariance_row (speed.c:847) divides the double sum by size_t -> double (exact). The twin divides by the rounded float, so the covariance can differ from the CPU and the twin is no longer bit-exact. At default p=1 the cap max is 4,165,681: exact
cuda/speed/speed_score.cu:1195-1199, 1222-1223
covariance ch = blockIdx.x / 325, triangle entry, x * kN + y
SAFE (same float conversion as the CPU, speed.c:775)
hip/speed/speed_hip_device.h:601-602
ch * 25 + x
uint32_t
< 100
99
99
SAFE
hip/speed/speed_hip_device.h:607-613
total = sub_w * sub_h; pos += group; (size_t)(xr + row) * trunc_w + xc + col
uint32_t / size_t
sub product (+256)
8,270,672
67,010,852
SAFE (\(2^{26}\) < \(2^{32}\))
hip/speed/speed_hip_device.h:624-625
covariance divisor (float)(p->geometry.sub_w * p->geometry.sub_h) passed to speed_hd_ff_div_to_float
uint32_t product -> float
sub product, must be exact
p4 max 8,270,416 < \(2^{24}\): exact
\(32740^{2}\) at p4: 66,928,761 is not an fp32 value
OVERFLOW@CAP-ONLY (fp32 exact-integer range, not UB). Same defect as CUDA :1219: the CPU speed.c:847 divides by the exact double count, so bit parity breaks for p > ~2.0 beyond 16K
Device memory, not an integer issue. The speed_prescale maximum of 4 makes the SCALED buffer 4 ch * 16 * N * 4 B. That is about 34 GB at 16K and 256 GiB at the cap. The allocation fails cleanly (vmaf_cuda_buffer_alloc / hipMalloc return an errno). All size math is size_t.
The same covariance divisor form exists outside G6 in sycl/speed_sycl_pipeline.cpp:808 (static_cast<float>(a.sub_w * a.sub_h)). Not audited here; it most likely shares the OVERFLOW@CAP-ONLY parity issue.
The CPU compute_mean (speed.c:775) also rounds its count to float. The means rows are therefore SAFE for parity even above \(2^{24}\). Only the covariance divisor differs from the CPU.
device pitch, CUDA rounds widthInBytes up to the device pitch alignment
pitch >= 30720 B
pitch >= 65536 B; plane = \(2^{31}\) B
SAFE in host (size_t). Device-side consumers that take this pitch as int and form y * pitch reach \(2^{31}\) at cap -> see per-feature rows. w/h of the cookie are capped at 32768 by libvmaf.c:432 (device_alloc_check_args itself checks only bpc).
x += int64_t(__shfl_down_sync(.., x & 0xffffffff, i)) \| int64_t(__shfl_down_sync(.., x >> 32, i) << 32); atomicAdd((uint64_cu *)address, (uint64_cu)val)
int64_t (shuffled as long long halves); atomic on unsigned long long
32 lanes of the caller's int64 partials (the bound of each sum is the caller's: ADM rows, VIF frame sums, SSIM weights, rows in the per-feature sections)
caller-defined
caller-defined
SAFE as a helper: the halves reconstruct the neighbour's value exactly, and the atomic adds two's-complement bits, so the int64 result is the caller's exact sum while it fits. Aside: (x >> 32) << 32 left-shifts a negative long long when a partial is negative, which is UB before C++20 (nvcc defaults to C++17); the ADM ADR-0155 negative i4 term is the only negative caller.
Coverage addendum (integrator): core/src/cuda/cuda_helper.cuh (1 row above). Shared headers included by the twins and not named by a reader, checked by the integrator: ff_pair.h (none: fp32 pair arithmetic, no integer type), motion_window.h (none: declarations), vif_tools.h (none: float filters and declarations), picture_copy.h (none: declaration; ptrdiff_t dst_stride).
Per-file row counts and "none" verdicts are in each group's own Coverage subsection above. This list maps every in-scope file to the group that read it in full (G = section above).
core/src/feature/cuda and core/src/feature/hip (130 files)¶
CPU references the groups read to derive term maxima or compare arithmetic (no rows of their own unless a group states one): integer_adm.c, integer_vif.c, integer_motion.c, integer_psnr.c, ciede.c, integer_ssim.c, iqa/decimate.c, cambi.c, speed.c, speed_internal.c, third_party/xiph/psnr_hvs.c, psnr_hvs_score.c, core/src/picture.c, core/src/libvmaf.c (dimension guards).