ADRs tagged rc3¶
Auto-generated by scripts/docs/generate-adr-by-tag.sh. Edit ADR Tags: lines to update.
110 ADR(s) carry this tag.
| ID | Title |
|---|---|
| ADR-1357 | Run the SYCL CAMBI extractor entirely on the device |
| ADR-1363 | The SYCL ssimulacra2 twin is device-resident, and float_ms_ssim_sycl waits once per frame |
| ADR-1367 | Every SYCL feature TU compiles with contraction off and correctly rounded fp32 division and square root |
| ADR-1369 | SYCL twins read the planes the state uploads once per frame; opt-in shared chroma planes and a device-side slot fence |
| ADR-1370 | float_ssim_sycl decimates on the device, bit-identical to the CPU |
| ADR-1378 | Run the HIP CAMBI extractor entirely on the device |
| ADR-1379 | Run the CUDA CAMBI extractor entirely on the device |
| ADR-1380 | The CUDA SpEED twins are device-resident, with the 25x25 linear algebra on the device and CPU-exact fp32 arithmetic |
| ADR-1384 | Run the HIP SpEED twins entirely on the device, in the CPU's fp32 arithmetic |
| ADR-1386 | One script runs the home GPU box's RC3 verify commands, row by row, under per-device locks |
| ADR-1390 | The HIP ssimulacra2 twin is device-resident with tiled row-pass Gaussian blur |
| ADR-1391 | The CUDA ssimulacra2 twin is device-resident |
| ADR-1395 | SYCL kernels use no scratch memory on Intel GPUs |
| ADR-1397 | GPU psnr_hvs twins reproduce the CPU's running float sum bit for bit |
| ADR-1399 | float_ssim_cuda decimates and convolves on the device with the CPU's arithmetic |
| ADR-1401 | psnr_hvs_sycl and psnr_hvs_hip return the CPU's scores bit for bit; the fp64-free masking threshold is an integer square root |
| ADR-1403 | Every CUDA feature kernel compiles without FMA contraction, and a twin spells the fused operations its reference performs |
| ADR-1405 | float_ssim_hip decimates on the device, bit-identical to the CPU |
| ADR-1407 | Every HIP kernel compiles with contraction off and correctly rounded fp32 division and square root |
| ADR-1408 | A VmafContext uploads each plane of a frame once and every HIP twin reads that copy |
| ADR-1409 | float_motion_cuda adds its SAD in the CPU's order and returns the CPU's scores bit for bit |
| ADR-1410 | SYCL CLI picture pool allocates pinned host USM to bypass staging upload |
| ADR-1411 | float_motion_sycl adds its SAD in the CPU's order and returns the CPU's scores bit for bit |
| ADR-1412 | float_vif_cuda computes the CPU's arithmetic, adds in the CPU's order and returns its scores bit for bit |
| ADR-1414 | float_ms_ssim_sycl computes the CPU's arithmetic and returns its per-scale means bit for bit |
| ADR-1415 | every x86 SIMD library is built without FP contraction |
| ADR-1416 | adm_cuda takes its CSF weights, its rounding shifts and its score conclusion from the CPU's routines and folds the denominator once per row |
| ADR-1418 | Motion parity cells compare what every twin emits; a missing metric is a cell error |
| ADR-1419 | float_motion_hip stores its differences transposed and adds each row in the CPU's order |
| ADR-1420 | float_adm_cuda computes the CPU's arithmetic, divides through the host's reciprocal estimate and returns the CPU's scores bit for bit |
| ADR-1422 | float_vif_sycl computes the CPU's arithmetic without an fp64 type and returns its scores bit for bit |
| ADR-1423 | adm_hip takes its weights, shifts and score conclusion from the CPU, folds the denominator once per row and clears its accumulators after the upload |
| ADR-1424 | integer_ssim_cuda adds its terms in the CPU's raster order, on the host, and returns the CPU's score bit for bit |
| ADR-1426 | ciede_cuda computes the CPU's arithmetic and adds in the CPU's order; what remains is the math library, and the gate bounds it |
| ADR-1427 | A HIP frame queues its accumulator clears after its upload |
| ADR-1428 | Exact GPU twins are declared by one fragment file each, not by a shared literal |
| ADR-1430 | speed_chroma_cuda keeps its correctly rounded log2; the gate bounds what glibc's log2f adds |
| ADR-1432 | vif_sycl computes the gain terms of the integer VIF in exact integer arithmetic and returns the CPU's scores bit for bit |
| ADR-1433 | ssimulacra2_cuda returns the sums of the CPU's loops, formed on the device from integer increments per binade |
| ADR-1434 | float_adm_sycl computes the CPU's arithmetic without an fp64 type, adds in the CPU's order and returns the CPU's scores bit for bit |
| ADR-1435 | vif_hip reads the CPU's log2 table instead of evaluating log2f() on the device, and returns the CPU's scores bit for bit |
| ADR-1436 | ciede_sycl runs the CPU's statements on fp32 pairs and adds in the CPU's order; it lands where the CUDA twin does, 1.4e-11 from the CPU |
| ADR-1437 | motion_hip, motion_v2_hip, psnr_hip, integer_ms_ssim_hip and cambi_hip are declared exact twins; float_psnr_hip and float_moment_hip are not |
| ADR-1438 | integer_ssim_hip adds its terms in the CPU's raster order at every frame size and returns the CPU's score bit for bit |
| ADR-1440 | float_psnr_hip adds its squared differences as integers and returns the CPU's score bit for bit |
| ADR-1441 | float_ssim_hip forms its window sums through the arithmetic float_ms_ssim_hip shares with the CPU and returns the CPU's score bit for bit |
| ADR-1442 | float ADM divides; the processor's reciprocal estimate leaves the reference and the CUDA twin |
| ADR-1443 | integer_ssim_sycl computes the CPU's fp64 term in 64-bit integers and adds in the CPU's order; it returns the CPU's ssim bit for bit |
| ADR-1444 | float_vif_hip runs the arithmetic of the CUDA twin from one shared header and returns the CPU's scores bit for bit |
| ADR-1445 | ssimulacra2_hip evaluates the CPU's fp64 terms and returns the sums of the CPU's loops |
| ADR-1446 | ssimulacra2_sycl forms the CPU's fp64 terms in 64-bit integers and returns the sums of the CPU's loops; it returns the CPU's score bit for bit |
| ADR-1447 | float_moment_hip adds the float squares the CPU adds, and is bit-identical while the CPU's own sum is exact |
| ADR-1448 | ciede_hip runs the CPU's arithmetic in fp32 pairs, from the header the SYCL twin runs; what remains is glibc's powf and the last bits of a pair |
| ADR-1449 | float_moment_sycl adds the float squares the CPU adds, and is bit-identical while the CPU's own sum is exact |
| ADR-1450 | float_psnr_sycl adds its squared differences as integers, and is bit-identical to the CPU |
| ADR-1451 | adm_sycl, motion_sycl, motion_v2_sycl, psnr_sycl, float_ssim_sycl and cambi_sycl are declared exact twins; speed_chroma_sycl is not |
| ADR-1452 | the gate bounds speed_chroma_hip against the CPU by what glibc's log2f adds, as it does for the CUDA twin |
| ADR-1453 | float_moment_cuda adds the float squares the CPU adds, and is bit-identical while the CPU's own sum is exact |
| ADR-1454 | A large subtree AGENTS.md is a generated index over one page per topic |
| ADR-1455 | float_psnr_cuda adds its squared differences as integers, and is bit-identical to the CPU |
| ADR-1456 | vif_cuda keeps its device log2f(), proven equal to the CPU's log2 table on every entry, and is declared an exact twin |
| ADR-1457 | motion_cuda, motion_v2_cuda, psnr_cuda, float_ssim_cuda, float_ms_ssim_cuda and cambi_cuda are declared exact twins |
| ADR-1458 | float_adm_hip runs the CPU's arithmetic from the header the CUDA twin runs and returns the CPU's scores bit for bit |
| ADR-1459 | SpEED's covariance kernels return the scalar kernel's bits; they vectorise across sums, not within one |
| ADR-1460 | speed_temporal becomes a parity-gate feature with a derived bound, and every registered twin of a gated backend has to be a gate cell |
| ADR-1461 | no C or C++ translation unit is built with floating-point contraction; the strict policy is a project argument |
| ADR-1462 | vif_cuda reads the CPU's log2 table instead of evaluating log2f() on the device |
| ADR-1463 | float_ssim_sycl forms the CPU's fp64 terms in 64-bit integers and adds them on the host in the CPU's raster order |
| ADR-1464 | float_ssim_cuda adds its frame sums in the CPU's raster order, on the host |
| ADR-1465 | float_ms_ssim_cuda adds the terms of every scale in the CPU's raster order, on the host |
| ADR-1466 | float_ms_ssim_sycl stores every window's l, c and s of every scale and adds them on the host in the CPU's raster order |
| ADR-1467 | ciede.c writes its squares as products, so ciede2000 no longer depends on the compiler or on the C library's powf for them; a GCC build moves by up to 2e-11 |
| ADR-1468 | A SYCL kernel requires only a sub-group size every default AOT target accepts (16 or 32); the six kernels that required 8 move to 16 |
| ADR-1469 | the psnr_hvs SIMD butterfly is two functions, cut at the same statement in the AVX2 and the NEON twin |
| ADR-1470 | An enum that C and C++ translation units both see has one size: no C++-only underlying type other than int's width |
| ADR-1471 | The clang-tidy lanes are measured in the dev container, with the device toolchains |
| ADR-1472 | The integer ADM weight limits follow from the contrast-masking cube and the largest wavelet coefficient of a scale |
| ADR-1473 | The x86 float ADM wavelet and CSF kernels return the scalar bits and are dispatched; the two reduction kernels are removed |
| ADR-1474 | A fork-created header that replays upstream arithmetic is EUPL-1.2 AND the licences of exactly that code, and the relicensing check becomes a required CI job |
| ADR-1476 | ciede2000() forms its two products in float, as Netflix's source does; the double casts of PR #552 are removed and the three GPU twins follow |
| ADR-1477 | SpEED evaluates Netflix's fp64 expressions again, and its GPU twins form the entropies and the score on the host with the same statements, so every twin returns the CPU's scores bit for bit on any C library |
| ADR-1478 | Port motion_five_frame_window from Netflix; the deferral of ADR-0337 ends for this option |
| ADR-1479 | ciede upsamples 4:2:2 chroma with the horizontal flag for columns and the vertical flag for rows; upstream has the two swapped, and ciede2000 differs by up to 0.153 on 4:2:2 input |
| ADR-1480 | speed_temporal sizes its frame buffers for the prescaled height; upstream sizes them for the source height and reads past them when speed_prescale is above 1 |
| ADR-1481 | an extractor that fails on a worker thread fails the run; upstream drops the worker's error and returns success without the metric |
| ADR-1482 | integer adm on frames of 17 to 32 pixels reads inside the frame and rounds a zero shift with 0; upstream reads index -1 there, and scale 3 differs by up to 0.23 |
| ADR-1483 | a subsampled chroma plane of an odd-sized picture is rounded up; upstream rounds down, and chroma metrics on odd sizes differ by up to 0.83 dB |
| ADR-1484 | float_ms_ssim takes the magnitude of a scale's terms before pow(); upstream raises a negative structure term to a fractional power and returns NaN on anti-correlated frames |
| ADR-1485 | the aggregate PSNR of a plane without error is the per-frame cap; upstream publishes a ceiling 54 dB above it on the 1080p checkerboard chroma |
| ADR-1486 | float_motion scales its scale-1 planes with the stride it was called with; upstream recomputes the stride from the plane width, and motion_add_scale1 with motion_add_uv differs by up to 25 |
| ADR-1487 | code inherited from Netflix/vmaf evaluates as Netflix's source does, a difference needs an ADR, and a guard compares every emitted value against the recorded upstream head |
| ADR-1488 | the psnr_hvs masking threshold is the double root of a float product, as Netflix's source writes it; the double product of PR #552 is removed and the AVX2, NEON, CUDA, HIP and SYCL forms follow, bit for bit |
| ADR-1491 | The CUDA, SYCL and HIP motion twins compute motion_five_frame_window: the frame two back on the device, the CPU's window function on the host |
| ADR-1494 | adm and float_adm refuse frames below 17x17; upstream's integer ADM crashes there and its float ADM returns values that at 8x8 and 12x9 depend on the heap |
| ADR-1495 | icx and icpx builds link glibc's libm, not Intel's libimf |
| ADR-1497 | The float_moment twins form the CPU's rounded second-moment sum past 2^53 units, and return the CPU's bits on every frame |
| ADR-1499 | The float_psnr twins add each row's exact sum in the CPU's order, and return the CPU's bits on every frame |
| ADR-1500 | The NEON and SVE2 float_moment kernels add in the scalar's order and return its bits on every input and vector length |
| ADR-1501 | The float_adm_sycl term kernel takes the large register file and leaves the sub-group size to the compiler, so it uses no scratch memory on Xe2 |
| ADR-1502 | Tests that assert an exact result compare floats by their bits through one helper, core/test/float_bits.h, which treats a NaN as identical to nothing |
| ADR-1561 | integer VIF converts its residual variance through vif_sv_sq(), which returns x86's value without the undefined double to int32_t conversion |
| ADR-1571 | a GPU dispatch variable is documented only when library code reads it; VMAF_CUDA_DISPATCH is read at extractor init, VMAF_HIP_DISPATCH and its predicate are removed |
| ADR-1601 | integer VIF forms the denominator log argument sigma_nsq + sigma1_sq in uint32_t |
| ADR-1673 | A push to master never cancels the runs of an earlier master commit; the concurrency group carries the SHA there |
| ADR-1686 | a Scorecard master run whose master moved on to a descendant ends cancelled |
| ADR-1707 | Cut v1.0.0-rc.3 without waiting for outside-hardware reports |
| ADR-1830 | vif_sycl runs at SIMD-16 only; the SIMD-32 kernels and VMAF_SYCL_VIF_SUBGROUP_SIZE are removed |
| ADR-1836 | vif_cuda names its scores after enable_chroma when the caller sets it |
| ADR-1917 | The integer ADM scale-0 horizontal and vertical weight limit is 43900, set by the CSF stage's 16-bit magnitude |
| ADR-1918 | Samples above 2^bpc - 1 are invalid input; an opt-in check refuses them |