ADRs tagged cuda¶
Auto-generated by scripts/docs/generate-adr-by-tag.sh. Edit ADR Tags: lines to update.
201 ADR(s) carry this tag.
| ID | Title |
|---|---|
| ADR-0022 | Inference runtime is ONNX Runtime via execution providers |
| ADR-0027 | Non-conservative image pins with experimental toolchain flags |
| ADR-0121 | Windows GPU build-only matrix legs (MSVC + CUDA, MSVC + oneAPI SYCL) |
| ADR-0122 | CUDA gencode coverage + actionable init-failure logging |
| ADR-0123 | CUDA prev_ref null-deref on ffmpeg libvmaf_cuda path |
| ADR-0131 | Port Netflix#1382 — cuMemFreeAsync → cuMemFree in vmaf_cuda_picture_free |
| ADR-0150 | Port Netflix #1472 — CUDA feature extraction on Windows (MSYS2/MinGW) |
| ADR-0156 | CUDA backend: graceful error propagation (Netflix#1420) |
| ADR-0157 | CUDA preallocation memory leak fix + vmaf_cuda_state_free public API (Netflix#1300) |
| ADR-0181 | Global feature-characteristics registry + per-backend dispatch strategy |
| ADR-0182 | GPU long-tail batch 1 — psnr + ciede + moment on CUDA / SYCL / Vulkan |
| ADR-0188 | GPU long-tail batch 2 — psnr_hvs / ssim / ms_ssim across CUDA / SYCL / Vulkan |
| ADR-0192 | GPU long-tail batch 3 — closing every remaining metric gap (motion_v2 / float_ansnr / ssimulacra2 / cambi + float twins) |
| ADR-0194 | float_ansnr GPU kernels — single-dispatch 3x3 + 5x5 filters with per-WG float partials |
| ADR-0195 | float_psnr GPU kernels — single-dispatch diff² with float partials, bit-exact vs CPU |
| ADR-0196 | float_motion GPU kernels — float twin of integer_motion blur+SAD |
| ADR-0197 | float_vif GPU kernels — 4-scale pyramid with mirror-asymmetry fix |
| ADR-0202 | float_adm CUDA + SYCL twins — sixth Group B float kernel finishes |
| ADR-0206 | ssimulacra2 CUDA + SYCL twins |
| ADR-0214 | GPU-parity CI gate (T6-8) — cross-device variance matrix |
| ADR-0219 | motion3 GPU coverage on Vulkan + CUDA + SYCL (3-frame window) |
| ADR-0234 | GPU-generation-aware ULP calibration head |
| ADR-0239 | Backend-agnostic GPU picture pool (gpu_picture_pool.{h,c}) |
| ADR-0243 | enable_lcs MS-SSIM extras on CUDA + Vulkan |
| ADR-0246 | Per-backend GPU kernel scaffolding templates (CUDA + Vulkan) |
| ADR-0271 | Wire integer_ms_ssim_cuda through the CUDA fence-batching helper |
| ADR-0299 | GPU scoring backend for vmaf-tune (--score-backend) |
| ADR-0345 | cambi × {CUDA, SYCL, HIP} GPU port strategy |
| ADR-0351 | CUDA PSNR — chroma extension (psnr_cb / psnr_cr) |
| ADR-0358 | CUDA motion correctness — SAD race, pinned-mem leak, and motion2/motion3 precision parity with CPU |
| ADR-0360 | CAMBI CUDA port (Strategy II hybrid, T3-15a) |
| ADR-0374 | Build-time-optional public APIs return -ENOSYS when disabled |
| ADR-0378 | Per-picture CUDA streams must use CU_STREAM_NON_BLOCKING |
| ADR-0385 | Feature-extractor deduplication by provided-feature names |
| ADR-0408 | FFmpeg libvmaf filter — CUDA backend selector |
| ADR-0410 | ssimulacra2_cuda GPU module leak + per-scale malloc removal |
| ADR-0431 | Split CUDA and CPU Feature Passes for FR-from-NR Extraction |
| ADR-0447 | Motion features under-report on HFR / 50p content |
| ADR-0451 | Local dev-MCP container for live probing |
| ADR-0453 | PSNR enable_chroma option parity across all GPU backends |
| ADR-0454 | VIF CUDA shared-memory staging for horizontal and vertical filter passes |
| ADR-0456 | SSIMULACRA2 CUDA Blur: 3-Channel Kernel Fusion and V-Pass Transpose for Coalesced Access |
| ADR-0464 | CAMBI CUDA spatial-mask shared-memory tile |
| ADR-0483 | Extract shared vmaf_gpu_dispatch_parse_env tokenizer |
| ADR-0485 | Extract VMAF_LIFECYCLE_ZERO macro to eliminate struct-init duplication across HIP and Metal kernel templates |
| ADR-0486 | Codify the three-function GPU backend context-API contract in docs |
| ADR-0487 | Wire adm_min_val option into integer_adm GPU backends |
| ADR-0488 | Shared once-snapshot helper for GPU dispatch env variables |
| ADR-0514 | dev-MCP container exposes every host GPU backend (CUDA + SYCL + Vulkan + HIP) |
| ADR-0529 | Replace /dev/dri/by-path bind with whole /dev/dri bind in dev container |
| ADR-0542 | Full GPU backend plumbing in the dev-mcp container |
| ADR-0564 | Real integer_ssim GPU kernels (CUDA, HIP, SYCL) — replace silent float_ssim substitution |
| ADR-0567 | Real On-Device GPU Kernels for speed_chroma and speed_temporal (4 Backends) |
| ADR-0573 | Dev-mcp container — ubuntu:26.04 + CUDA 13.2 + hipcc + ocloc |
| ADR-0574 | CUDA Twins for HDR-Model Features — Phase 1 (aim, adm3) |
| ADR-0576 | ffmpeg-patches n8.1.1 full-feature-exposure sync |
| ADR-0582 | MS-SSIM enable_db and clip_db option parity on CUDA and SYCL backends |
| ADR-0590 | Wire enable_db / clip_db into the CUDA and SYCL MS-SSIM twins |
| ADR-0591 | Restore rfe_hw_flags per-frame bitmask cache after PR #1067 clobber |
| ADR-0596 | Delete orphan and duplicate HIP/CUDA translation units |
| ADR-0597 | integer_vif is luma-only across every backend; CUDA enable_chroma is a documented no-op |
| ADR-0599 | Cross-Backend Parity Audit — Full Extractor Matrix (2026-05-18) |
| ADR-0603 | Ubuntu 26.04 (Resolute Raccoon) fallout fixes — CUDA 13.2, Python 3.14, apt renames |
| ADR-0605 | Renovate customManagers for all dev/Containerfile pinned dependencies |
| ADR-0662 | Vulkan Motion Lavapipe Parity |
| ADR-0664 | Install Windows CUDA directly in CI |
| ADR-0667 | vmaf-tune score backend native priority |
| ADR-0699 | VMAFX Helm Chart and Kubernetes Manifests with 3-Vendor GPU Device-Plugin Support |
| ADR-0738 | Bump local CUDA toolkit pin to 13.3 + R610 minimum driver (partial — CI deferred) |
| ADR-0743 | CUDA VIF filter1d ncu-driven performance optimizations |
| ADR-0744 | CUDA adm_cm __launch_bounds__(128, 8) register reduction (ms_ssim_decimate smem tiling reverted) |
| ADR-0746 | integer_adm_cuda — emit integer_adm3 + integer_aim (parity with CPU) |
| ADR-0750 | Hardware Measurement Verdict for PR perf/cuda-ms-ssim-decimate-adm-cm-ncu-driven |
| ADR-0753 | Runtime Resolution-Aware CUDA Kernel Variant Dispatch |
| ADR-0754 | 0754-cuda-ssim-vert-combine-ldg-pinned-leak.md |
| ADR-0756 | CUDA F3 struct-by-value kernel audit (scope + dispatch order) |
| ADR-0757 | 0757-cuda-ms-ssim-vert-lcs-horiz-ldg.md |
| ADR-0759 | HIP ADM — AdmBufferHip passed by pointer (F3 fix) |
| ADR-0760 | CUDA motion kernel multi-resolution ncu profiling methodology |
| ADR-0762 | CUDA CIEDE2000 8bpc/16bpc — __ldg() read-only cache routing (F3 fix) |
| ADR-0763 | 0763-cuda-adm-decouple-ldg.md |
| ADR-0764 | psnr_hvs CUDA kernel — __ldg() + __restrict__ + __launch_bounds__(64) |
| ADR-0773 | CUDA ADM decouple-inline — __ldg() F3 fix on active path |
| ADR-0777 | Thread-Safety Audit — CUDA / SYCL / HIP Backends |
| ADR-0780 | NOLINT Cluster Refactor — Slab Allocator, SYCL Stride, ADM Band-Size |
| ADR-0787 | 0787-libvmaf-api-error-path-audit.md |
| ADR-0840 | Fix cu_state leak on import failure and gpu_dispatch_env TOCTOU |
| ADR-0841 | Environment variable reference page and canonical naming |
| ADR-0845 | CUDA motion — multi-frame SAD batching to reduce per-launch overhead |
| ADR-0868 | GPU backend kernel parity-test coverage gap-fill |
| ADR-0874 | Name magic numbers in fork-added C surfaces (CERT INT07-C closeout pass 1) |
| ADR-0886 | CUDA kernel parity test coverage — round 2 gap-fill |
| ADR-0928 | VmafPicture v2 — explicit per-backend GPU state |
| ADR-0947 | CUDA kernel parity coverage — round 3 (float-path twins + ssimulacra2) |
| ADR-0954 | Host-only unit test for shared GPU dispatch runtime |
| ADR-0956 | CUDA kernel parity coverage — round 4 (last 5 uncovered kernels) |
| ADR-0960 | GPU runtime error-path leak fixes — round 25 (A.1 + A.2 + A.3) |
| ADR-0964 | Implement speed_internal.c and wire speed_{chroma,temporal}_{hip,sycl} |
| ADR-0965 | CUDA SpEED TU repair — align with current CudaFunctions table (closes T-CUDA-SPEED-TU-REPAIR-2026-05-31) |
| ADR-0970 | test_gpu_picture_pool.c: remove unused malloc + dead code (Round 27 audit D.3 + D.4) |
| ADR-0982 | GPU runtime bug audit — round 26 (init/teardown leak sweep) |
| ADR-0990 | Restore double-precision L/C/S accumulation in CUDA ms_ssim_vert_lcs |
| ADR-1011 | Add static to TU-internal CUDA helper functions — VIF, ADM, motion |
| ADR-1025 | R6 CUDA/HIP kernel correctness fixes |
| ADR-1053 | Default docker-compose runtime to nvidia and expand GPU capabilities |
| ADR-1078 | ms_ssim option parity across HIP and SYCL backends |
| ADR-1090 | Fix CUDA stream and event leaks on init error paths |
| ADR-1100 | Skip GPU-flagged extractors when flags == 0 in vmaf_get_feature_extractor_by_feature_name |
| ADR-1108 | CUDA motion_v2 twin emits motion3_v2_score |
| ADR-1142 | Whole-codebase standards; lint debt only ratchets down |
| ADR-1143 | CUDA and Intel SYCL Backend Gap Closure |
| ADR-1167 | Row-level rounding accumulator and border row selection for integer ADM GPU kernels |
| ADR-1185 | Per-backend performance baselines are median-of-N, one backend per build dir |
| ADR-1191 | Integer ADM rejects CSF configurations its fixed-point storage cannot represent |
| ADR-1192 | Keep the recorded Netflix benchmark snapshot; do not regenerate it while the GPU paths are broken |
| ADR-1193 | Opt-in uncapped option splits the PSNR infinity sentinel from the truncation |
| ADR-1194 | One integer-ADM angle_flag predicate for every backend |
| ADR-1197 | The threaded flush leaves GPU extractors to their own backend flush |
| ADR-1199 | Order caller-written CUDA pictures once per frame, at the dispatch point |
| ADR-1202 | GPU SpEED-chroma twins report singularity separately from failure |
| ADR-1203 | psnr_hvs_cuda defaults enable_chroma to true, matching every other backend |
| ADR-1204 | GPU ADM contrast-masking twins clamp the far edge instead of mirroring it |
| ADR-1205 | The ssimulacra2 FMA unification extends to the scalar fallback and every GPU host copy |
| ADR-1206 | Every CUDA parity test also runs against a second, larger fixture |
| ADR-1212 | The GPU float_moment twins normalise by the bit-depth scaler on the host |
| ADR-1214 | The float-ADM GPU twins ignore adm_csf_scale in Watson mode and share the CPU's option aliases |
| ADR-1215 | The 16-bpc CUDA PSNR kernel takes the plane index the host has always passed |
| ADR-1216 | The GPU motion3 twins apply motion_fps_weight exactly once |
| ADR-1217 | The GPU float-VIF kernels read vif_sigma_nsq and vif_enhn_gain_limit from their options |
| ADR-1218 | The GPU SpEED twins zero the device solution and report singularity from the temporal path |
| ADR-1220 | The GPU float-ADM kernels honour adm_p_norm, adm_bypass_cm and adm_skip_scale0 |
| ADR-1221 | clip_db is a ceiling on the MS-SSIM dB output, not a clamp on the linear score |
| ADR-1223 | The CUDA backend requires compute capability 8.0 (Ampere), and CI standardises on CUDA 13.3.1 |
| ADR-1224 | CUDA Tile C++ is not adopted; the audit's incidental findings are |
| ADR-1226 | Size the CUDA AIM CM launch by SM count, not by a fixed rows-per-thread |
| ADR-1228 | A recurring "faster than upstream, and still exact" milestone |
| ADR-1285 | CUDA is a coordinated pin — group it in Renovate, gate the rest |
| ADR-1290 | The tidy ratchet owns its per-lane compilation database |
| ADR-1296 | GPU init failure paths are tested device-free, by compiling the backend TU against runtime stubs |
| ADR-1300 | Install CUDA on every CI leg from NVIDIA's own distribution |
| ADR-1306 | Replace nvidia/cuda base images with digest-pinned Ubuntu and version-locked apt installation |
| ADR-1320 | CUDA fatbin and HIP HSACO kernel header dependency tracking |
| ADR-1325 | Normalize integer ADM Barten weights with one exponent per scale |
| ADR-1336 | Tear down CUDA resources in their owning context |
| ADR-1372 | CUDA motion differences the frames before the blur, in the kernel motion_v2 already had |
| ADR-1373 | CUDA PSNR, SSIM and motion twins take the CPU option tables and the CPU's arithmetic |
| ADR-1374 | CUDA integer ADM and VIF guard tiny frames like their SYCL twins |
| ADR-1379 | Run the CUDA CAMBI extractor entirely on the device |
| ADR-1380 | The CUDA SpEED twins are device-resident, with the 25x25 linear algebra on the device and CPU-exact fp32 arithmetic |
| ADR-1386 | One script runs the home GPU box's RC3 verify commands, row by row, under per-device locks |
| ADR-1391 | The CUDA ssimulacra2 twin is device-resident |
| ADR-1392 | CUDA motion, PSNR and moment kernels add one atomic per block, the motion SAD filters separably, and PSNR selects its plane with constant indices |
| ADR-1397 | GPU psnr_hvs twins reproduce the CPU's running float sum bit for bit |
| ADR-1399 | float_ssim_cuda decimates and convolves on the device with the CPU's arithmetic |
| ADR-1402 | Integer ADM keeps the scale-0 masking centre tap in int32 and clamps the excess in int64 |
| ADR-1403 | Every CUDA feature kernel compiles without FMA contraction, and a twin spells the fused operations its reference performs |
| ADR-1406 | Preallocate pinned host pictures for zero-copy 4K CLI CUDA upload |
| ADR-1409 | float_motion_cuda adds its SAD in the CPU's order and returns the CPU's scores bit for bit |
| ADR-1412 | float_vif_cuda computes the CPU's arithmetic, adds in the CPU's order and returns its scores bit for bit |
| ADR-1416 | adm_cuda takes its CSF weights, its rounding shifts and its score conclusion from the CPU's routines and folds the denominator once per row |
| ADR-1418 | Motion parity cells compare what every twin emits; a missing metric is a cell error |
| ADR-1420 | float_adm_cuda computes the CPU's arithmetic, divides through the host's reciprocal estimate and returns the CPU's scores bit for bit |
| ADR-1424 | integer_ssim_cuda adds its terms in the CPU's raster order, on the host, and returns the CPU's score bit for bit |
| ADR-1426 | ciede_cuda computes the CPU's arithmetic and adds in the CPU's order; what remains is the math library, and the gate bounds it |
| ADR-1430 | speed_chroma_cuda keeps its correctly rounded log2; the gate bounds what glibc's log2f adds |
| ADR-1431 | vmaf_read_pictures() owns the pictures it is given on every return |
| ADR-1433 | ssimulacra2_cuda returns the sums of the CPU's loops, formed on the device from integer increments per binade |
| ADR-1442 | float ADM divides; the processor's reciprocal estimate leaves the reference and the CUDA twin |
| ADR-1444 | float_vif_hip runs the arithmetic of the CUDA twin from one shared header and returns the CPU's scores bit for bit |
| ADR-1453 | float_moment_cuda adds the float squares the CPU adds, and is bit-identical while the CPU's own sum is exact |
| ADR-1455 | float_psnr_cuda adds its squared differences as integers, and is bit-identical to the CPU |
| ADR-1456 | vif_cuda keeps its device log2f(), proven equal to the CPU's log2 table on every entry, and is declared an exact twin |
| ADR-1457 | motion_cuda, motion_v2_cuda, psnr_cuda, float_ssim_cuda, float_ms_ssim_cuda and cambi_cuda are declared exact twins |
| ADR-1458 | float_adm_hip runs the CPU's arithmetic from the header the CUDA twin runs and returns the CPU's scores bit for bit |
| ADR-1460 | speed_temporal becomes a parity-gate feature with a derived bound, and every registered twin of a gated backend has to be a gate cell |
| ADR-1462 | vif_cuda reads the CPU's log2 table instead of evaluating log2f() on the device |
| ADR-1464 | float_ssim_cuda adds its frame sums in the CPU's raster order, on the host |
| ADR-1465 | float_ms_ssim_cuda adds the terms of every scale in the CPU's raster order, on the host |
| ADR-1472 | The integer ADM weight limits follow from the contrast-masking cube and the largest wavelet coefficient of a scale |
| ADR-1477 | SpEED evaluates Netflix's fp64 expressions again, and its GPU twins form the entropies and the score on the host with the same statements, so every twin returns the CPU's scores bit for bit on any C library |
| ADR-1491 | The CUDA, SYCL and HIP motion twins compute motion_five_frame_window: the frame two back on the device, the CPU's window function on the host |
| ADR-1497 | The float_moment twins form the CPU's rounded second-moment sum past 2^53 units, and return the CPU's bits on every frame |
| ADR-1499 | The float_psnr twins add each row's exact sum in the CPU's order, and return the CPU's bits on every frame |
| ADR-1509 | An NVIDIA GPU tester image that ships no NVIDIA file and measures every CUDA twin on a tester's GPU |
| ADR-1516 | A Windows CUDA tester zip that ships no NVIDIA file and measures every CUDA twin on a tester's Windows PC |
| ADR-1517 | The production GPU images are built on Debian 13, ship only the vendor files libvmaf loads, and share the tester images' licence records |
| ADR-1561 | integer VIF converts its residual variance through vif_sv_sq(), which returns x86's value without the undefined double to int32_t conversion |
| ADR-1571 | a GPU dispatch variable is documented only when library code reads it; VMAF_CUDA_DISPATCH is read at extractor init, VMAF_HIP_DISPATCH and its predicate are removed |
| ADR-1590 | Every build stores its GPU device code compressed at the toolchain's strongest setting, and the build refuses raw device code |
| ADR-1601 | integer VIF forms the denominator log argument sigma_nsq + sigma1_sq in uint32_t |
| ADR-1685 | A post-1.0 embedding milestone: zero-copy device-frame import with fences, asynchronous window scores, and an unchanged licence |
| ADR-1768 | The FFmpeg libvmaf and libvmaf_cuda filters print no pooled score after a mid-run error |
| ADR-1829 | RC4 owns the whole device-memory import API, not only the first full Rust metric |
| ADR-1836 | vif_cuda names its scores after enable_chroma when the caller sets it |
| ADR-1874 | The vmaf CLI reports its usable backends; the score-backend selectors read that report |
| ADR-1917 | The integer ADM scale-0 horizontal and vertical weight limit is 43900, set by the CSF stage's 16-bit magnitude |
| ADR-1918 | Samples above 2^bpc - 1 are invalid input; an opt-in check refuses them |
| ADR-1929 | VMAFx device frames, fences and the import rule: the shared contract the backend lanes implement |
| ADR-2023 | VMAFx device frames on CUDA: one library stream per device, fences on it, release events recorded where the frame is released |
| ADR-2134 | adm_cm_aim_line_kernel_4 gets its own register budget of 209 for the exact scale-0 angle flag |
| ADR-2796 | Every clang-tidy ratchet lane is a hosted, path-routed required check |