Skip to content

ADRs tagged gpu

Auto-generated by scripts/docs/generate-adr-by-tag.sh. Edit ADR Tags: lines to update.

175 ADR(s) carry this tag.

ID Title
ADR-0101 SYCL USM-backed picture pre-allocation pool
ADR-0127 Vulkan compute backend — vendor-neutral GPU path alongside CUDA/SYCL/HIP
ADR-0175 Vulkan compute backend — scaffold-only audit-first PR (T5-1)
ADR-0176 Vulkan VIF cross-backend gate (lavapipe + Arc nightly)
ADR-0177 Vulkan motion kernel + motion cross-backend gate
ADR-0178 Vulkan ADM kernel (T5-1c-adm)
ADR-0181 Global feature-characteristics registry + per-backend dispatch strategy
ADR-0182 GPU long-tail batch 1 — psnr + ciede + moment on CUDA / SYCL / Vulkan
ADR-0187 ciede2000 Vulkan kernel — float-precision per-pixel ΔE
ADR-0188 GPU long-tail batch 2 — psnr_hvs / ssim / ms_ssim across CUDA / SYCL / Vulkan
ADR-0189 float_ssim Vulkan kernel — host decimation, 2-dispatch GPU
ADR-0190 float_ms_ssim Vulkan kernel — 5-level pyramid + Wang product on host
ADR-0191 float_psnr_hvs Vulkan kernel — overlapping 8×8 DCT blocks + per-plane log transform
ADR-0192 GPU long-tail batch 3 — closing every remaining metric gap (motion_v2 / float_ansnr / ssimulacra2 / cambi + float twins)
ADR-0193 motion_v2 Vulkan kernel — single-dispatch SAD via convolution linearity
ADR-0194 float_ansnr GPU kernels — single-dispatch 3x3 + 5x5 filters with per-WG float partials
ADR-0195 float_psnr GPU kernels — single-dispatch diff² with float partials, bit-exact vs CPU
ADR-0196 float_motion GPU kernels — float twin of integer_motion blur+SAD
ADR-0197 float_vif GPU kernels — 4-scale pyramid with mirror-asymmetry fix
ADR-0199 float_adm Vulkan kernel — sixth Group B float twin
ADR-0201 ssimulacra2 Vulkan kernel
ADR-0202 float_adm CUDA + SYCL twins — sixth Group B float kernel finishes
ADR-0205 cambi GPU feasibility spike
ADR-0206 ssimulacra2 CUDA + SYCL twins
ADR-0210 cambi Vulkan integration (Strategy II hybrid)
ADR-0212 HIP (AMD ROCm) compute backend — scaffold-only audit-first PR (T7-10)
ADR-0214 GPU-parity CI gate (T6-8) — cross-device variance matrix
ADR-0216 Vulkan PSNR — chroma extension (psnr_cb / psnr_cr)
ADR-0219 motion3 GPU coverage on Vulkan + CUDA + SYCL (3-frame window)
ADR-0220 SYCL feature kernels are unconditionally fp64-free
ADR-0234 GPU-generation-aware ULP calibration head
ADR-0239 Backend-agnostic GPU picture pool (gpu_picture_pool.{h,c})
ADR-0240 GPU backend public-header pattern doc (PR3 of GPU dedup, doc-only)
ADR-0241 HIP first-consumer kernel — integer_psnr_hip via mirrored kernel-template
ADR-0243 enable_lcs MS-SSIM extras on CUDA + Vulkan
ADR-0246 Per-backend GPU kernel scaffolding templates (CUDA + Vulkan)
ADR-0254 HIP second-consumer kernel — float_psnr_hip via mirrored kernel-template
ADR-0259 HIP third-consumer kernel — ciede_hip via mirrored kernel-template
ADR-0260 HIP fourth-consumer kernel — float_moment_hip via mirrored kernel-template
ADR-0266 HIP fifth kernel-template consumer — float_ansnr_hip
ADR-0267 HIP sixth kernel-template consumer — motion_v2_hip
ADR-0271 Wire integer_ms_ssim_cuda through the CUDA fence-batching helper
ADR-0273 HIP seventh kernel-template consumer — float_motion_hip
ADR-0274 HIP eighth kernel-template consumer — float_ssim_hip
ADR-0290 NVENC codec adapters for vmaf-tune (h264 / hevc / av1)
ADR-0314 vmaf-tune --score-backend=vulkan (vendor-neutral GPU scoring)
ADR-0315 Vendor-neutral VVC encode strategy — tiered Tier-1-now / Tier-2-backlog / Tier-3-revisit
ADR-0338 macOS Vulkan-via-MoltenVK CI lane (advisory) for the Vulkan backend
ADR-0345 cambi × {CUDA, SYCL, HIP} GPU port strategy
ADR-0351 CUDA PSNR — chroma extension (psnr_cb / psnr_cr)
ADR-0353 Vulkan submit-pool migration PR-B — six secondary kernels
ADR-0356 Two-level GPU reduction for Vulkan VIF / ADM / motion accumulators
ADR-0360 CAMBI CUDA port (Strategy II hybrid, T3-15a)
ADR-0361 Metal compute backend — scaffold-only audit-first PR (T8-1)
ADR-0372 HIP Batch-1 — integer_psnr_hip and float_ansnr_hip Real Kernels
ADR-0373 HIP Batch-2 — float_motion_hip Real Kernel
ADR-0375 HIP batch-3 — float_moment_hip and float_ssim_hip real kernels
ADR-0376 Fix silent error-swallow in Vulkan buffer-invalidate readback functions
ADR-0377 HIP batch-4 — ciede_hip and integer_motion_v2_hip real kernels
ADR-0378 Per-picture CUDA streams must use CU_STREAM_NON_BLOCKING
ADR-0385 Feature-extractor deduplication by provided-feature names
ADR-0391 ciede2000 Vulkan NVIDIA places=4 fork debt is a structural f32/f64 precision gap
ADR-0410 ssimulacra2_cuda GPU module leak + per-scale malloc removal
ADR-0415 CAMBI SYCL port — closes last CUDA-to-SYCL parity gap
ADR-0420 Metal backend runtime (T8-1b)
ADR-0421 Metal first kernel — integer_motion_v2 (T8-1c)
ADR-0422 CLI HIP and Metal Backend Selectors
ADR-0423 Metal IOSurface zero-copy import (T8-IOS)
ADR-0445 Persistent VkPipelineCache for Vulkan compute backend
ADR-0451 Local dev-MCP container for live probing
ADR-0454 VIF CUDA shared-memory staging for horizontal and vertical filter passes
ADR-0458 SYCL CAMBI queue-sync collapse + SSIM horizontal SLM staging
ADR-0464 CAMBI CUDA spatial-mask shared-memory tile
ADR-0484 Extend kernel-scaffolding.md with HIP and Metal lifecycle contract
ADR-0486 Codify the three-function GPU backend context-API contract in docs
ADR-0488 Shared once-snapshot helper for GPU dispatch env variables
ADR-0489 CAMBI SYCL — Replace GPU-to-GPU q.wait() Calls with Event Chains (SY-1)
ADR-0514 dev-MCP container exposes every host GPU backend (CUDA + SYCL + Vulkan + HIP)
ADR-0519 Implement vmaf_hip_import_state to unblock --backend hip
ADR-0523 Register vmaf_fex_integer_motion_hip in the extractor list
ADR-0530 HIP feature-extractor flag promotion and HIP_DEVICE picture-buffer type
ADR-0533 Full HIP feature-extractor registration sweep
ADR-0537 HIP integer VIF kernel crash fix — filter upload, bounds, HtoD staging
ADR-0539 integer ADM HIP kernels — real implementation replacing weak HSACO stubs
ADR-0541 Pin dev-MCP container Intel NEO + ROCm runtimes to versions matching the host kernel
ADR-0552 Deterministic wavefront reduction for integer_vif_hip horizontal kernels
ADR-0563 HIP extractor audit — verification of 9 remaining scaffold claims
ADR-0564 Real integer_ssim GPU kernels (CUDA, HIP, SYCL) — replace silent float_ssim substitution
ADR-0567 Real On-Device GPU Kernels for speed_chroma and speed_temporal (4 Backends)
ADR-0568 Default sycl_icpx_aot_targets to full Intel arch list
ADR-0587 Real Metal Compute Kernels for CAMBI
ADR-0593 HIP integer_moment kernel — register real HSACO blob alongside psnr / psnr_hvs
ADR-0667 vmaf-tune score backend native priority
ADR-0699 VMAFX Helm Chart and Kubernetes Manifests with 3-Vendor GPU Device-Plugin Support
ADR-0726 Drop Vulkan backend
ADR-0804 Add vmaf_context_get_backend — additive ABI introspection
ADR-0852 Wire speed_chroma_hip and speed_temporal_hip into HIP Build and Dispatch
ADR-0883 HIP kernel parity-test coverage round 2
ADR-0884 SYCL kernel coverage round 2 — five additional CPU-vs-SYCL parity gates
ADR-0886 CUDA kernel parity test coverage — round 2 gap-fill
ADR-0928 VmafPicture v2 — explicit per-backend GPU state
ADR-0945 HIP kernel parity-test coverage round 3
ADR-0946 SYCL kernel coverage round 3 (float family + PSNR-HVS)
ADR-0954 Host-only unit test for shared GPU dispatch runtime
ADR-0957 SYCL kernel coverage round 4 (float_moment + SpEED + SSIMULACRA2)
ADR-0958 HIP kernel parity-test coverage round 4
ADR-0959 Metal kernel parity coverage round 4 — closeout
ADR-0982 GPU runtime bug audit — round 26 (init/teardown leak sweep)
ADR-0985 SYCL parity divergence investigation — float_ssim + ssimulacra2 on Arc A380
ADR-0989 Wire motion_add_uv through integer_motion_sycl; emit warning on motion_five_frame_window
ADR-1001 SYCL parity round 5 — CAMBI CPU vs. SYCL parity gate
ADR-1004 HIP kernel parity-test coverage round 5
ADR-1030 HIP adm_decouple dangling body + VIF wavefront 32-bit carry + Metal motion vertical halo
ADR-1034 Fix SYCL integer_vif rd_stride OOB on odd widths and integer_motion UV queue sync gap
ADR-1106 HIP motion_v2 mirror is reflect-101 (-2), correcting ADR-0377's -1 parity claim
ADR-1121 SYCL QSV zero-copy — P010 pixel normalization and separate-session decode contract
ADR-1129 Align release containers with the published tag and runtime ABI
ADR-1154 AMD ROCm HIP Backend Gap Closure and Extractor Promotion
ADR-1167 Row-level rounding accumulator and border row selection for integer ADM GPU kernels
ADR-1176 Metal motion_v2 mirror closeout and reflect-101 parity contract
ADR-1177 Containerised self-hosted GitHub Actions runner for Intel Arc SYCL parity CI
ADR-1179 Fix Intel Arc SYCL Crashes and Default Model Resolution Divergence
ADR-1264 The HIP scaffold posture reports -ENOSYS, and its tests check both sites
ADR-1296 GPU init failure paths are tested device-free, by compiling the backend TU against runtime stubs
ADR-1299 Compute MS-SSIM chroma on SYCL
ADR-1300 Install CUDA on every CI leg from NVIDIA's own distribution
ADR-1312 GPU option aliases match CPU collector keys
ADR-1316 Mark extractor options whose implementation is default-only
ADR-1319 Admit self-hosted GPU jobs through a live fail-closed probe
ADR-1320 CUDA fatbin and HIP HSACO kernel header dependency tracking
ADR-1324 Resolve GPU float-SSIM auto-scale before backend initialization
ADR-1357 Run the SYCL CAMBI extractor entirely on the device
ADR-1358 The SYCL SpEED twins are device-resident, with the 25x25 linear algebra on the device and exact fp32 arithmetic
ADR-1359 The CLI maps --feature <cpu-name> to the explicit --backend's twin
ADR-1360 Generate SYCL AOT images at compile time and fail the build when they are missing
ADR-1362 The SYCL integer ADM twin computes AIM on the device and finalises every ADM output in the CPU's float arithmetic
ADR-1363 The SYCL ssimulacra2 twin is device-resident, and float_ms_ssim_sycl waits once per frame
ADR-1364 Register the SYCL device images of a Windows MSVC build through one explicit device link
ADR-1367 Every SYCL feature TU compiles with contraction off and correctly rounded fp32 division and square root
ADR-1368 Build and run the oneAPI release image on Debian 13 with pinned Intel packages
ADR-1369 SYCL twins read the planes the state uploads once per frame; opt-in shared chroma planes and a device-side slot fence
ADR-1370 float_ssim_sycl decimates on the device, bit-identical to the CPU
ADR-1378 Run the HIP CAMBI extractor entirely on the device
ADR-1379 Run the CUDA CAMBI extractor entirely on the device
ADR-1380 The CUDA SpEED twins are device-resident, with the 25x25 linear algebra on the device and CPU-exact fp32 arithmetic
ADR-1384 Run the HIP SpEED twins entirely on the device, in the CPU's fp32 arithmetic
ADR-1390 The HIP ssimulacra2 twin is device-resident with tiled row-pass Gaussian blur
ADR-1391 The CUDA ssimulacra2 twin is device-resident
ADR-1395 SYCL kernels use no scratch memory on Intel GPUs
ADR-1399 float_ssim_cuda decimates and convolves on the device with the CPU's arithmetic
ADR-1403 Every CUDA feature kernel compiles without FMA contraction, and a twin spells the fused operations its reference performs
ADR-1405 float_ssim_hip decimates on the device, bit-identical to the CPU
ADR-1407 Every HIP kernel compiles with contraction off and correctly rounded fp32 division and square root
ADR-1408 A VmafContext uploads each plane of a frame once and every HIP twin reads that copy
ADR-1498 Metal twins take the exact designs of their CUDA, HIP and SYCL twins, on a strict FP kernel policy
ADR-1523 HIP twins run on the device of the imported state
ADR-1525 adm_hip computes AIM on the device and is dispatched
ADR-1547 The Helm chart derives the GPU resource name from the vendor and the Intel kernel driver, with an explicit override
ADR-1571 a GPU dispatch variable is documented only when library code reads it; VMAF_CUDA_DISPATCH is read at extractor init, VMAF_HIP_DISPATCH and its predicate are removed
ADR-1590 Every build stores its GPU device code compressed at the toolchain's strongest setting, and the build refuses raw device code
ADR-1685 A post-1.0 embedding milestone: zero-copy device-frame import with fences, asynchronous window scores, and an unchanged licence
ADR-1806 Run Metal kernel files on the host through a Metal Shading Language shim in tests
ADR-1829 RC4 owns the whole device-memory import API, not only the first full Rust metric
ADR-1830 vif_sycl runs at SIMD-16 only; the SIMD-32 kernels and VMAF_SYCL_VIF_SUBGROUP_SIZE are removed
ADR-1874 The vmaf CLI reports its usable backends; the score-backend selectors read that report
ADR-1880 Format envelope and device-targeted scoring in the 1.0.0 candidates
ADR-1900 Deterministic verification and state contract for VPL decode retry ceiling
ADR-1929 VMAFx device frames, fences and the import rule: the shared contract the backend lanes implement
ADR-2023 VMAFx device frames on CUDA: one library stream per device, fences on it, release events recorded where the frame is released
ADR-2056 float_adm files its debug ratio unsuffixed; a second debug instance is refused
ADR-2091 VMAFx device frames on SYCL: readers copy on the device behind the frame's ready event, the release waits on every reader
ADR-2092 VMAFx device frames on HIP: one library stream per device copies every frame for the twins, dma-bufs as external memory, sync_file checked on the host
ADR-2132 HIP imports GL textures through EGL dma-buf export; the runtime's GL interop is not used
ADR-2133 Device import takes 4:2:2 and 4:4:4: semi-planar NV16 / NV24 family and packed Y210 / Y410, one CPU reference, converted on the device
ADR-2478 Rust replaces the host-side C and C++ through 3.0, behind the unchanged C ABI