Skip to content

ADRs tagged cuda

Auto-generated by scripts/docs/generate-adr-by-tag.sh. Edit ADR Tags: lines to update.

201 ADR(s) carry this tag.

ID Title
ADR-0022 Inference runtime is ONNX Runtime via execution providers
ADR-0027 Non-conservative image pins with experimental toolchain flags
ADR-0121 Windows GPU build-only matrix legs (MSVC + CUDA, MSVC + oneAPI SYCL)
ADR-0122 CUDA gencode coverage + actionable init-failure logging
ADR-0123 CUDA prev_ref null-deref on ffmpeg libvmaf_cuda path
ADR-0131 Port Netflix#1382 — cuMemFreeAsync → cuMemFree in vmaf_cuda_picture_free
ADR-0150 Port Netflix #1472 — CUDA feature extraction on Windows (MSYS2/MinGW)
ADR-0156 CUDA backend: graceful error propagation (Netflix#1420)
ADR-0157 CUDA preallocation memory leak fix + vmaf_cuda_state_free public API (Netflix#1300)
ADR-0181 Global feature-characteristics registry + per-backend dispatch strategy
ADR-0182 GPU long-tail batch 1 — psnr + ciede + moment on CUDA / SYCL / Vulkan
ADR-0188 GPU long-tail batch 2 — psnr_hvs / ssim / ms_ssim across CUDA / SYCL / Vulkan
ADR-0192 GPU long-tail batch 3 — closing every remaining metric gap (motion_v2 / float_ansnr / ssimulacra2 / cambi + float twins)
ADR-0194 float_ansnr GPU kernels — single-dispatch 3x3 + 5x5 filters with per-WG float partials
ADR-0195 float_psnr GPU kernels — single-dispatch diff² with float partials, bit-exact vs CPU
ADR-0196 float_motion GPU kernels — float twin of integer_motion blur+SAD
ADR-0197 float_vif GPU kernels — 4-scale pyramid with mirror-asymmetry fix
ADR-0202 float_adm CUDA + SYCL twins — sixth Group B float kernel finishes
ADR-0206 ssimulacra2 CUDA + SYCL twins
ADR-0214 GPU-parity CI gate (T6-8) — cross-device variance matrix
ADR-0219 motion3 GPU coverage on Vulkan + CUDA + SYCL (3-frame window)
ADR-0234 GPU-generation-aware ULP calibration head
ADR-0239 Backend-agnostic GPU picture pool (gpu_picture_pool.{h,c})
ADR-0243 enable_lcs MS-SSIM extras on CUDA + Vulkan
ADR-0246 Per-backend GPU kernel scaffolding templates (CUDA + Vulkan)
ADR-0271 Wire integer_ms_ssim_cuda through the CUDA fence-batching helper
ADR-0299 GPU scoring backend for vmaf-tune (--score-backend)
ADR-0345 cambi × {CUDA, SYCL, HIP} GPU port strategy
ADR-0351 CUDA PSNR — chroma extension (psnr_cb / psnr_cr)
ADR-0358 CUDA motion correctness — SAD race, pinned-mem leak, and motion2/motion3 precision parity with CPU
ADR-0360 CAMBI CUDA port (Strategy II hybrid, T3-15a)
ADR-0374 Build-time-optional public APIs return -ENOSYS when disabled
ADR-0378 Per-picture CUDA streams must use CU_STREAM_NON_BLOCKING
ADR-0385 Feature-extractor deduplication by provided-feature names
ADR-0408 FFmpeg libvmaf filter — CUDA backend selector
ADR-0410 ssimulacra2_cuda GPU module leak + per-scale malloc removal
ADR-0431 Split CUDA and CPU Feature Passes for FR-from-NR Extraction
ADR-0447 Motion features under-report on HFR / 50p content
ADR-0451 Local dev-MCP container for live probing
ADR-0453 PSNR enable_chroma option parity across all GPU backends
ADR-0454 VIF CUDA shared-memory staging for horizontal and vertical filter passes
ADR-0456 SSIMULACRA2 CUDA Blur: 3-Channel Kernel Fusion and V-Pass Transpose for Coalesced Access
ADR-0464 CAMBI CUDA spatial-mask shared-memory tile
ADR-0483 Extract shared vmaf_gpu_dispatch_parse_env tokenizer
ADR-0485 Extract VMAF_LIFECYCLE_ZERO macro to eliminate struct-init duplication across HIP and Metal kernel templates
ADR-0486 Codify the three-function GPU backend context-API contract in docs
ADR-0487 Wire adm_min_val option into integer_adm GPU backends
ADR-0488 Shared once-snapshot helper for GPU dispatch env variables
ADR-0514 dev-MCP container exposes every host GPU backend (CUDA + SYCL + Vulkan + HIP)
ADR-0529 Replace /dev/dri/by-path bind with whole /dev/dri bind in dev container
ADR-0542 Full GPU backend plumbing in the dev-mcp container
ADR-0564 Real integer_ssim GPU kernels (CUDA, HIP, SYCL) — replace silent float_ssim substitution
ADR-0567 Real On-Device GPU Kernels for speed_chroma and speed_temporal (4 Backends)
ADR-0573 Dev-mcp container — ubuntu:26.04 + CUDA 13.2 + hipcc + ocloc
ADR-0574 CUDA Twins for HDR-Model Features — Phase 1 (aim, adm3)
ADR-0576 ffmpeg-patches n8.1.1 full-feature-exposure sync
ADR-0582 MS-SSIM enable_db and clip_db option parity on CUDA and SYCL backends
ADR-0590 Wire enable_db / clip_db into the CUDA and SYCL MS-SSIM twins
ADR-0591 Restore rfe_hw_flags per-frame bitmask cache after PR #1067 clobber
ADR-0596 Delete orphan and duplicate HIP/CUDA translation units
ADR-0597 integer_vif is luma-only across every backend; CUDA enable_chroma is a documented no-op
ADR-0599 Cross-Backend Parity Audit — Full Extractor Matrix (2026-05-18)
ADR-0603 Ubuntu 26.04 (Resolute Raccoon) fallout fixes — CUDA 13.2, Python 3.14, apt renames
ADR-0605 Renovate customManagers for all dev/Containerfile pinned dependencies
ADR-0662 Vulkan Motion Lavapipe Parity
ADR-0664 Install Windows CUDA directly in CI
ADR-0667 vmaf-tune score backend native priority
ADR-0699 VMAFX Helm Chart and Kubernetes Manifests with 3-Vendor GPU Device-Plugin Support
ADR-0738 Bump local CUDA toolkit pin to 13.3 + R610 minimum driver (partial — CI deferred)
ADR-0743 CUDA VIF filter1d ncu-driven performance optimizations
ADR-0744 CUDA adm_cm __launch_bounds__(128, 8) register reduction (ms_ssim_decimate smem tiling reverted)
ADR-0746 integer_adm_cuda — emit integer_adm3 + integer_aim (parity with CPU)
ADR-0750 Hardware Measurement Verdict for PR perf/cuda-ms-ssim-decimate-adm-cm-ncu-driven
ADR-0753 Runtime Resolution-Aware CUDA Kernel Variant Dispatch
ADR-0754 0754-cuda-ssim-vert-combine-ldg-pinned-leak.md
ADR-0756 CUDA F3 struct-by-value kernel audit (scope + dispatch order)
ADR-0757 0757-cuda-ms-ssim-vert-lcs-horiz-ldg.md
ADR-0759 HIP ADM — AdmBufferHip passed by pointer (F3 fix)
ADR-0760 CUDA motion kernel multi-resolution ncu profiling methodology
ADR-0762 CUDA CIEDE2000 8bpc/16bpc — __ldg() read-only cache routing (F3 fix)
ADR-0763 0763-cuda-adm-decouple-ldg.md
ADR-0764 psnr_hvs CUDA kernel — __ldg() + __restrict__ + __launch_bounds__(64)
ADR-0773 CUDA ADM decouple-inline — __ldg() F3 fix on active path
ADR-0777 Thread-Safety Audit — CUDA / SYCL / HIP Backends
ADR-0780 NOLINT Cluster Refactor — Slab Allocator, SYCL Stride, ADM Band-Size
ADR-0787 0787-libvmaf-api-error-path-audit.md
ADR-0840 Fix cu_state leak on import failure and gpu_dispatch_env TOCTOU
ADR-0841 Environment variable reference page and canonical naming
ADR-0845 CUDA motion — multi-frame SAD batching to reduce per-launch overhead
ADR-0868 GPU backend kernel parity-test coverage gap-fill
ADR-0874 Name magic numbers in fork-added C surfaces (CERT INT07-C closeout pass 1)
ADR-0886 CUDA kernel parity test coverage — round 2 gap-fill
ADR-0928 VmafPicture v2 — explicit per-backend GPU state
ADR-0947 CUDA kernel parity coverage — round 3 (float-path twins + ssimulacra2)
ADR-0954 Host-only unit test for shared GPU dispatch runtime
ADR-0956 CUDA kernel parity coverage — round 4 (last 5 uncovered kernels)
ADR-0960 GPU runtime error-path leak fixes — round 25 (A.1 + A.2 + A.3)
ADR-0964 Implement speed_internal.c and wire speed_{chroma,temporal}_{hip,sycl}
ADR-0965 CUDA SpEED TU repair — align with current CudaFunctions table (closes T-CUDA-SPEED-TU-REPAIR-2026-05-31)
ADR-0970 test_gpu_picture_pool.c: remove unused malloc + dead code (Round 27 audit D.3 + D.4)
ADR-0982 GPU runtime bug audit — round 26 (init/teardown leak sweep)
ADR-0990 Restore double-precision L/C/S accumulation in CUDA ms_ssim_vert_lcs
ADR-1011 Add static to TU-internal CUDA helper functions — VIF, ADM, motion
ADR-1025 R6 CUDA/HIP kernel correctness fixes
ADR-1053 Default docker-compose runtime to nvidia and expand GPU capabilities
ADR-1078 ms_ssim option parity across HIP and SYCL backends
ADR-1090 Fix CUDA stream and event leaks on init error paths
ADR-1100 Skip GPU-flagged extractors when flags == 0 in vmaf_get_feature_extractor_by_feature_name
ADR-1108 CUDA motion_v2 twin emits motion3_v2_score
ADR-1142 Whole-codebase standards; lint debt only ratchets down
ADR-1143 CUDA and Intel SYCL Backend Gap Closure
ADR-1167 Row-level rounding accumulator and border row selection for integer ADM GPU kernels
ADR-1185 Per-backend performance baselines are median-of-N, one backend per build dir
ADR-1191 Integer ADM rejects CSF configurations its fixed-point storage cannot represent
ADR-1192 Keep the recorded Netflix benchmark snapshot; do not regenerate it while the GPU paths are broken
ADR-1193 Opt-in uncapped option splits the PSNR infinity sentinel from the truncation
ADR-1194 One integer-ADM angle_flag predicate for every backend
ADR-1197 The threaded flush leaves GPU extractors to their own backend flush
ADR-1199 Order caller-written CUDA pictures once per frame, at the dispatch point
ADR-1202 GPU SpEED-chroma twins report singularity separately from failure
ADR-1203 psnr_hvs_cuda defaults enable_chroma to true, matching every other backend
ADR-1204 GPU ADM contrast-masking twins clamp the far edge instead of mirroring it
ADR-1205 The ssimulacra2 FMA unification extends to the scalar fallback and every GPU host copy
ADR-1206 Every CUDA parity test also runs against a second, larger fixture
ADR-1212 The GPU float_moment twins normalise by the bit-depth scaler on the host
ADR-1214 The float-ADM GPU twins ignore adm_csf_scale in Watson mode and share the CPU's option aliases
ADR-1215 The 16-bpc CUDA PSNR kernel takes the plane index the host has always passed
ADR-1216 The GPU motion3 twins apply motion_fps_weight exactly once
ADR-1217 The GPU float-VIF kernels read vif_sigma_nsq and vif_enhn_gain_limit from their options
ADR-1218 The GPU SpEED twins zero the device solution and report singularity from the temporal path
ADR-1220 The GPU float-ADM kernels honour adm_p_norm, adm_bypass_cm and adm_skip_scale0
ADR-1221 clip_db is a ceiling on the MS-SSIM dB output, not a clamp on the linear score
ADR-1223 The CUDA backend requires compute capability 8.0 (Ampere), and CI standardises on CUDA 13.3.1
ADR-1224 CUDA Tile C++ is not adopted; the audit's incidental findings are
ADR-1226 Size the CUDA AIM CM launch by SM count, not by a fixed rows-per-thread
ADR-1228 A recurring "faster than upstream, and still exact" milestone
ADR-1285 CUDA is a coordinated pin — group it in Renovate, gate the rest
ADR-1290 The tidy ratchet owns its per-lane compilation database
ADR-1296 GPU init failure paths are tested device-free, by compiling the backend TU against runtime stubs
ADR-1300 Install CUDA on every CI leg from NVIDIA's own distribution
ADR-1306 Replace nvidia/cuda base images with digest-pinned Ubuntu and version-locked apt installation
ADR-1320 CUDA fatbin and HIP HSACO kernel header dependency tracking
ADR-1325 Normalize integer ADM Barten weights with one exponent per scale
ADR-1336 Tear down CUDA resources in their owning context
ADR-1372 CUDA motion differences the frames before the blur, in the kernel motion_v2 already had
ADR-1373 CUDA PSNR, SSIM and motion twins take the CPU option tables and the CPU's arithmetic
ADR-1374 CUDA integer ADM and VIF guard tiny frames like their SYCL twins
ADR-1379 Run the CUDA CAMBI extractor entirely on the device
ADR-1380 The CUDA SpEED twins are device-resident, with the 25x25 linear algebra on the device and CPU-exact fp32 arithmetic
ADR-1386 One script runs the home GPU box's RC3 verify commands, row by row, under per-device locks
ADR-1391 The CUDA ssimulacra2 twin is device-resident
ADR-1392 CUDA motion, PSNR and moment kernels add one atomic per block, the motion SAD filters separably, and PSNR selects its plane with constant indices
ADR-1397 GPU psnr_hvs twins reproduce the CPU's running float sum bit for bit
ADR-1399 float_ssim_cuda decimates and convolves on the device with the CPU's arithmetic
ADR-1402 Integer ADM keeps the scale-0 masking centre tap in int32 and clamps the excess in int64
ADR-1403 Every CUDA feature kernel compiles without FMA contraction, and a twin spells the fused operations its reference performs
ADR-1406 Preallocate pinned host pictures for zero-copy 4K CLI CUDA upload
ADR-1409 float_motion_cuda adds its SAD in the CPU's order and returns the CPU's scores bit for bit
ADR-1412 float_vif_cuda computes the CPU's arithmetic, adds in the CPU's order and returns its scores bit for bit
ADR-1416 adm_cuda takes its CSF weights, its rounding shifts and its score conclusion from the CPU's routines and folds the denominator once per row
ADR-1418 Motion parity cells compare what every twin emits; a missing metric is a cell error
ADR-1420 float_adm_cuda computes the CPU's arithmetic, divides through the host's reciprocal estimate and returns the CPU's scores bit for bit
ADR-1424 integer_ssim_cuda adds its terms in the CPU's raster order, on the host, and returns the CPU's score bit for bit
ADR-1426 ciede_cuda computes the CPU's arithmetic and adds in the CPU's order; what remains is the math library, and the gate bounds it
ADR-1430 speed_chroma_cuda keeps its correctly rounded log2; the gate bounds what glibc's log2f adds
ADR-1431 vmaf_read_pictures() owns the pictures it is given on every return
ADR-1433 ssimulacra2_cuda returns the sums of the CPU's loops, formed on the device from integer increments per binade
ADR-1442 float ADM divides; the processor's reciprocal estimate leaves the reference and the CUDA twin
ADR-1444 float_vif_hip runs the arithmetic of the CUDA twin from one shared header and returns the CPU's scores bit for bit
ADR-1453 float_moment_cuda adds the float squares the CPU adds, and is bit-identical while the CPU's own sum is exact
ADR-1455 float_psnr_cuda adds its squared differences as integers, and is bit-identical to the CPU
ADR-1456 vif_cuda keeps its device log2f(), proven equal to the CPU's log2 table on every entry, and is declared an exact twin
ADR-1457 motion_cuda, motion_v2_cuda, psnr_cuda, float_ssim_cuda, float_ms_ssim_cuda and cambi_cuda are declared exact twins
ADR-1458 float_adm_hip runs the CPU's arithmetic from the header the CUDA twin runs and returns the CPU's scores bit for bit
ADR-1460 speed_temporal becomes a parity-gate feature with a derived bound, and every registered twin of a gated backend has to be a gate cell
ADR-1462 vif_cuda reads the CPU's log2 table instead of evaluating log2f() on the device
ADR-1464 float_ssim_cuda adds its frame sums in the CPU's raster order, on the host
ADR-1465 float_ms_ssim_cuda adds the terms of every scale in the CPU's raster order, on the host
ADR-1472 The integer ADM weight limits follow from the contrast-masking cube and the largest wavelet coefficient of a scale
ADR-1477 SpEED evaluates Netflix's fp64 expressions again, and its GPU twins form the entropies and the score on the host with the same statements, so every twin returns the CPU's scores bit for bit on any C library
ADR-1491 The CUDA, SYCL and HIP motion twins compute motion_five_frame_window: the frame two back on the device, the CPU's window function on the host
ADR-1497 The float_moment twins form the CPU's rounded second-moment sum past 2^53 units, and return the CPU's bits on every frame
ADR-1499 The float_psnr twins add each row's exact sum in the CPU's order, and return the CPU's bits on every frame
ADR-1509 An NVIDIA GPU tester image that ships no NVIDIA file and measures every CUDA twin on a tester's GPU
ADR-1516 A Windows CUDA tester zip that ships no NVIDIA file and measures every CUDA twin on a tester's Windows PC
ADR-1517 The production GPU images are built on Debian 13, ship only the vendor files libvmaf loads, and share the tester images' licence records
ADR-1561 integer VIF converts its residual variance through vif_sv_sq(), which returns x86's value without the undefined double to int32_t conversion
ADR-1571 a GPU dispatch variable is documented only when library code reads it; VMAF_CUDA_DISPATCH is read at extractor init, VMAF_HIP_DISPATCH and its predicate are removed
ADR-1590 Every build stores its GPU device code compressed at the toolchain's strongest setting, and the build refuses raw device code
ADR-1601 integer VIF forms the denominator log argument sigma_nsq + sigma1_sq in uint32_t
ADR-1685 A post-1.0 embedding milestone: zero-copy device-frame import with fences, asynchronous window scores, and an unchanged licence
ADR-1768 The FFmpeg libvmaf and libvmaf_cuda filters print no pooled score after a mid-run error
ADR-1829 RC4 owns the whole device-memory import API, not only the first full Rust metric
ADR-1836 vif_cuda names its scores after enable_chroma when the caller sets it
ADR-1874 The vmaf CLI reports its usable backends; the score-backend selectors read that report
ADR-1917 The integer ADM scale-0 horizontal and vertical weight limit is 43900, set by the CSF stage's 16-bit magnitude
ADR-1918 Samples above 2^bpc - 1 are invalid input; an opt-in check refuses them
ADR-1929 VMAFx device frames, fences and the import rule: the shared contract the backend lanes implement
ADR-2023 VMAFx device frames on CUDA: one library stream per device, fences on it, release events recorded where the frame is released
ADR-2134 adm_cm_aim_line_kernel_4 gets its own register budget of 209 for the exact scale-0 angle flag
ADR-2796 Every clang-tidy ratchet lane is a hosted, path-routed required check