Skip to content

ADRs tagged numerics

Auto-generated by scripts/docs/generate-adr-by-tag.sh. Edit ADR Tags: lines to update.

96 ADR(s) carry this tag.

ID Title
ADR-0594 Per-kernel hip_cu_extra_flags dispatch — disable FMA contraction for ssimulacra2_blur HIP HSACO
ADR-0688 HIP wave32 carry-preserving int64 reduction for VIF and motion kernels
ADR-1326 Use a fixed-point oracle for SYCL motion-add-UV parity
ADR-1357 Run the SYCL CAMBI extractor entirely on the device
ADR-1358 The SYCL SpEED twins are device-resident, with the 25x25 linear algebra on the device and exact fp32 arithmetic
ADR-1361 Scale the psnr_hvs cross-backend tolerance with the CPU's float-sum length
ADR-1362 The SYCL integer ADM twin computes AIM on the device and finalises every ADM output in the CPU's float arithmetic
ADR-1363 The SYCL ssimulacra2 twin is device-resident, and float_ms_ssim_sycl waits once per frame
ADR-1365 SYCL PSNR, SSIM and float-motion twins take the CPU option tables
ADR-1367 Every SYCL feature TU compiles with contraction off and correctly rounded fp32 division and square root
ADR-1370 float_ssim_sycl decimates on the device, bit-identical to the CPU
ADR-1371 SYCL motion differences the frames before the blur, in one shared kernel
ADR-1372 CUDA motion differences the frames before the blur, in the kernel motion_v2 already had
ADR-1373 CUDA PSNR, SSIM and motion twins take the CPU option tables and the CPU's arithmetic
ADR-1377 HIP motion differences the frames before the blur, in one shared kernel, and waits only in collect
ADR-1378 Run the HIP CAMBI extractor entirely on the device
ADR-1379 Run the CUDA CAMBI extractor entirely on the device
ADR-1380 The CUDA SpEED twins are device-resident, with the 25x25 linear algebra on the device and CPU-exact fp32 arithmetic
ADR-1382 HIP PSNR, SSIM and float-motion twins take the CPU option tables
ADR-1384 Run the HIP SpEED twins entirely on the device, in the CPU's fp32 arithmetic
ADR-1390 The HIP ssimulacra2 twin is device-resident with tiled row-pass Gaussian blur
ADR-1391 The CUDA ssimulacra2 twin is device-resident
ADR-1393 CAMBI's c-values walks clip the window to the frame at every edge
ADR-1395 SYCL kernels use no scratch memory on Intel GPUs
ADR-1397 GPU psnr_hvs twins reproduce the CPU's running float sum bit for bit
ADR-1399 float_ssim_cuda decimates and convolves on the device with the CPU's arithmetic
ADR-1400 integer_ssim_hip sums small frames in the CPU's raster order
ADR-1401 psnr_hvs_sycl and psnr_hvs_hip return the CPU's scores bit for bit; the fp64-free masking threshold is an integer square root
ADR-1402 Integer ADM keeps the scale-0 masking centre tap in int32 and clamps the excess in int64
ADR-1403 Every CUDA feature kernel compiles without FMA contraction, and a twin spells the fused operations its reference performs
ADR-1405 float_ssim_hip decimates on the device, bit-identical to the CPU
ADR-1407 Every HIP kernel compiles with contraction off and correctly rounded fp32 division and square root
ADR-1409 float_motion_cuda adds its SAD in the CPU's order and returns the CPU's scores bit for bit
ADR-1411 float_motion_sycl adds its SAD in the CPU's order and returns the CPU's scores bit for bit
ADR-1412 float_vif_cuda computes the CPU's arithmetic, adds in the CPU's order and returns its scores bit for bit
ADR-1413 Every integer ADM implementation bounds the enhancement gain with the scalar's truncated double product
ADR-1414 float_ms_ssim_sycl computes the CPU's arithmetic and returns its per-scale means bit for bit
ADR-1415 every x86 SIMD library is built without FP contraction
ADR-1416 adm_cuda takes its CSF weights, its rounding shifts and its score conclusion from the CPU's routines and folds the denominator once per row
ADR-1417 Integer AIM stays the unclipped ratio upstream defines; float AIM stays clipped at 1
ADR-1419 float_motion_hip stores its differences transposed and adds each row in the CPU's order
ADR-1420 float_adm_cuda computes the CPU's arithmetic, divides through the host's reciprocal estimate and returns the CPU's scores bit for bit
ADR-1422 float_vif_sycl computes the CPU's arithmetic without an fp64 type and returns its scores bit for bit
ADR-1423 adm_hip takes its weights, shifts and score conclusion from the CPU, folds the denominator once per row and clears its accumulators after the upload
ADR-1424 integer_ssim_cuda adds its terms in the CPU's raster order, on the host, and returns the CPU's score bit for bit
ADR-1426 ciede_cuda computes the CPU's arithmetic and adds in the CPU's order; what remains is the math library, and the gate bounds it
ADR-1430 speed_chroma_cuda keeps its correctly rounded log2; the gate bounds what glibc's log2f adds
ADR-1432 vif_sycl computes the gain terms of the integer VIF in exact integer arithmetic and returns the CPU's scores bit for bit
ADR-1433 ssimulacra2_cuda returns the sums of the CPU's loops, formed on the device from integer increments per binade
ADR-1434 float_adm_sycl computes the CPU's arithmetic without an fp64 type, adds in the CPU's order and returns the CPU's scores bit for bit
ADR-1435 vif_hip reads the CPU's log2 table instead of evaluating log2f() on the device, and returns the CPU's scores bit for bit
ADR-1436 ciede_sycl runs the CPU's statements on fp32 pairs and adds in the CPU's order; it lands where the CUDA twin does, 1.4e-11 from the CPU
ADR-1437 motion_hip, motion_v2_hip, psnr_hip, integer_ms_ssim_hip and cambi_hip are declared exact twins; float_psnr_hip and float_moment_hip are not
ADR-1438 integer_ssim_hip adds its terms in the CPU's raster order at every frame size and returns the CPU's score bit for bit
ADR-1440 float_psnr_hip adds its squared differences as integers and returns the CPU's score bit for bit
ADR-1441 float_ssim_hip forms its window sums through the arithmetic float_ms_ssim_hip shares with the CPU and returns the CPU's score bit for bit
ADR-1442 float ADM divides; the processor's reciprocal estimate leaves the reference and the CUDA twin
ADR-1443 integer_ssim_sycl computes the CPU's fp64 term in 64-bit integers and adds in the CPU's order; it returns the CPU's ssim bit for bit
ADR-1444 float_vif_hip runs the arithmetic of the CUDA twin from one shared header and returns the CPU's scores bit for bit
ADR-1445 ssimulacra2_hip evaluates the CPU's fp64 terms and returns the sums of the CPU's loops
ADR-1446 ssimulacra2_sycl forms the CPU's fp64 terms in 64-bit integers and returns the sums of the CPU's loops; it returns the CPU's score bit for bit
ADR-1447 float_moment_hip adds the float squares the CPU adds, and is bit-identical while the CPU's own sum is exact
ADR-1448 ciede_hip runs the CPU's arithmetic in fp32 pairs, from the header the SYCL twin runs; what remains is glibc's powf and the last bits of a pair
ADR-1449 float_moment_sycl adds the float squares the CPU adds, and is bit-identical while the CPU's own sum is exact
ADR-1450 float_psnr_sycl adds its squared differences as integers, and is bit-identical to the CPU
ADR-1451 adm_sycl, motion_sycl, motion_v2_sycl, psnr_sycl, float_ssim_sycl and cambi_sycl are declared exact twins; speed_chroma_sycl is not
ADR-1452 the gate bounds speed_chroma_hip against the CPU by what glibc's log2f adds, as it does for the CUDA twin
ADR-1453 float_moment_cuda adds the float squares the CPU adds, and is bit-identical while the CPU's own sum is exact
ADR-1455 float_psnr_cuda adds its squared differences as integers, and is bit-identical to the CPU
ADR-1456 vif_cuda keeps its device log2f(), proven equal to the CPU's log2 table on every entry, and is declared an exact twin
ADR-1457 motion_cuda, motion_v2_cuda, psnr_cuda, float_ssim_cuda, float_ms_ssim_cuda and cambi_cuda are declared exact twins
ADR-1458 float_adm_hip runs the CPU's arithmetic from the header the CUDA twin runs and returns the CPU's scores bit for bit
ADR-1459 SpEED's covariance kernels return the scalar kernel's bits; they vectorise across sums, not within one
ADR-1460 speed_temporal becomes a parity-gate feature with a derived bound, and every registered twin of a gated backend has to be a gate cell
ADR-1461 no C or C++ translation unit is built with floating-point contraction; the strict policy is a project argument
ADR-1462 vif_cuda reads the CPU's log2 table instead of evaluating log2f() on the device
ADR-1463 float_ssim_sycl forms the CPU's fp64 terms in 64-bit integers and adds them on the host in the CPU's raster order
ADR-1464 float_ssim_cuda adds its frame sums in the CPU's raster order, on the host
ADR-1465 float_ms_ssim_cuda adds the terms of every scale in the CPU's raster order, on the host
ADR-1466 float_ms_ssim_sycl stores every window's l, c and s of every scale and adds them on the host in the CPU's raster order
ADR-1467 ciede.c writes its squares as products, so ciede2000 no longer depends on the compiler or on the C library's powf for them; a GCC build moves by up to 2e-11
ADR-1475 The quantisation step of integer ADM is Netflix's expression again, so vmaf_v0.6.1 returns Netflix master's bits
ADR-1476 ciede2000() forms its two products in float, as Netflix's source does; the double casts of PR #552 are removed and the three GPU twins follow
ADR-1477 SpEED evaluates Netflix's fp64 expressions again, and its GPU twins form the entropies and the score on the host with the same statements, so every twin returns the CPU's scores bit for bit on any C library
ADR-1484 float_ms_ssim takes the magnitude of a scale's terms before pow(); upstream raises a negative structure term to a fractional power and returns NaN on anti-correlated frames
ADR-1485 the aggregate PSNR of a plane without error is the per-frame cap; upstream publishes a ceiling 54 dB above it on the 1080p checkerboard chroma
ADR-1487 code inherited from Netflix/vmaf evaluates as Netflix's source does, a difference needs an ADR, and a guard compares every emitted value against the recorded upstream head
ADR-1488 the psnr_hvs masking threshold is the double root of a float product, as Netflix's source writes it; the double product of PR #552 is removed and the AVX2, NEON, CUDA, HIP and SYCL forms follow, bit for bit
ADR-1489 The CSF weights of float ADM are Netflix's float arithmetic again, so float_adm differs from Netflix by the division alone
ADR-1495 icx and icpx builds link glibc's libm, not Intel's libimf
ADR-1497 The float_moment twins form the CPU's rounded second-moment sum past 2^53 units, and return the CPU's bits on every frame
ADR-1498 Metal twins take the exact designs of their CUDA, HIP and SYCL twins, on a strict FP kernel policy
ADR-1499 The float_psnr twins add each row's exact sum in the CPU's order, and return the CPU's bits on every frame
ADR-1500 The NEON and SVE2 float_moment kernels add in the scalar's order and return its bits on every input and vector length
ADR-1561 integer VIF converts its residual variance through vif_sv_sq(), which returns x86's value without the undefined double to int32_t conversion
ADR-1601 integer VIF forms the denominator log argument sigma_nsq + sigma1_sq in uint32_t