ADRs tagged gpu¶
Auto-generated by scripts/docs/generate-adr-by-tag.sh. Edit ADR Tags: lines to update.
175 ADR(s) carry this tag.
| ID | Title |
|---|---|
| ADR-0101 | SYCL USM-backed picture pre-allocation pool |
| ADR-0127 | Vulkan compute backend — vendor-neutral GPU path alongside CUDA/SYCL/HIP |
| ADR-0175 | Vulkan compute backend — scaffold-only audit-first PR (T5-1) |
| ADR-0176 | Vulkan VIF cross-backend gate (lavapipe + Arc nightly) |
| ADR-0177 | Vulkan motion kernel + motion cross-backend gate |
| ADR-0178 | Vulkan ADM kernel (T5-1c-adm) |
| ADR-0181 | Global feature-characteristics registry + per-backend dispatch strategy |
| ADR-0182 | GPU long-tail batch 1 — psnr + ciede + moment on CUDA / SYCL / Vulkan |
| ADR-0187 | ciede2000 Vulkan kernel — float-precision per-pixel ΔE |
| ADR-0188 | GPU long-tail batch 2 — psnr_hvs / ssim / ms_ssim across CUDA / SYCL / Vulkan |
| ADR-0189 | float_ssim Vulkan kernel — host decimation, 2-dispatch GPU |
| ADR-0190 | float_ms_ssim Vulkan kernel — 5-level pyramid + Wang product on host |
| ADR-0191 | float_psnr_hvs Vulkan kernel — overlapping 8×8 DCT blocks + per-plane log transform |
| ADR-0192 | GPU long-tail batch 3 — closing every remaining metric gap (motion_v2 / float_ansnr / ssimulacra2 / cambi + float twins) |
| ADR-0193 | motion_v2 Vulkan kernel — single-dispatch SAD via convolution linearity |
| ADR-0194 | float_ansnr GPU kernels — single-dispatch 3x3 + 5x5 filters with per-WG float partials |
| ADR-0195 | float_psnr GPU kernels — single-dispatch diff² with float partials, bit-exact vs CPU |
| ADR-0196 | float_motion GPU kernels — float twin of integer_motion blur+SAD |
| ADR-0197 | float_vif GPU kernels — 4-scale pyramid with mirror-asymmetry fix |
| ADR-0199 | float_adm Vulkan kernel — sixth Group B float twin |
| ADR-0201 | ssimulacra2 Vulkan kernel |
| ADR-0202 | float_adm CUDA + SYCL twins — sixth Group B float kernel finishes |
| ADR-0205 | cambi GPU feasibility spike |
| ADR-0206 | ssimulacra2 CUDA + SYCL twins |
| ADR-0210 | cambi Vulkan integration (Strategy II hybrid) |
| ADR-0212 | HIP (AMD ROCm) compute backend — scaffold-only audit-first PR (T7-10) |
| ADR-0214 | GPU-parity CI gate (T6-8) — cross-device variance matrix |
| ADR-0216 | Vulkan PSNR — chroma extension (psnr_cb / psnr_cr) |
| ADR-0219 | motion3 GPU coverage on Vulkan + CUDA + SYCL (3-frame window) |
| ADR-0220 | SYCL feature kernels are unconditionally fp64-free |
| ADR-0234 | GPU-generation-aware ULP calibration head |
| ADR-0239 | Backend-agnostic GPU picture pool (gpu_picture_pool.{h,c}) |
| ADR-0240 | GPU backend public-header pattern doc (PR3 of GPU dedup, doc-only) |
| ADR-0241 | HIP first-consumer kernel — integer_psnr_hip via mirrored kernel-template |
| ADR-0243 | enable_lcs MS-SSIM extras on CUDA + Vulkan |
| ADR-0246 | Per-backend GPU kernel scaffolding templates (CUDA + Vulkan) |
| ADR-0254 | HIP second-consumer kernel — float_psnr_hip via mirrored kernel-template |
| ADR-0259 | HIP third-consumer kernel — ciede_hip via mirrored kernel-template |
| ADR-0260 | HIP fourth-consumer kernel — float_moment_hip via mirrored kernel-template |
| ADR-0266 | HIP fifth kernel-template consumer — float_ansnr_hip |
| ADR-0267 | HIP sixth kernel-template consumer — motion_v2_hip |
| ADR-0271 | Wire integer_ms_ssim_cuda through the CUDA fence-batching helper |
| ADR-0273 | HIP seventh kernel-template consumer — float_motion_hip |
| ADR-0274 | HIP eighth kernel-template consumer — float_ssim_hip |
| ADR-0290 | NVENC codec adapters for vmaf-tune (h264 / hevc / av1) |
| ADR-0314 | vmaf-tune --score-backend=vulkan (vendor-neutral GPU scoring) |
| ADR-0315 | Vendor-neutral VVC encode strategy — tiered Tier-1-now / Tier-2-backlog / Tier-3-revisit |
| ADR-0338 | macOS Vulkan-via-MoltenVK CI lane (advisory) for the Vulkan backend |
| ADR-0345 | cambi × {CUDA, SYCL, HIP} GPU port strategy |
| ADR-0351 | CUDA PSNR — chroma extension (psnr_cb / psnr_cr) |
| ADR-0353 | Vulkan submit-pool migration PR-B — six secondary kernels |
| ADR-0356 | Two-level GPU reduction for Vulkan VIF / ADM / motion accumulators |
| ADR-0360 | CAMBI CUDA port (Strategy II hybrid, T3-15a) |
| ADR-0361 | Metal compute backend — scaffold-only audit-first PR (T8-1) |
| ADR-0372 | HIP Batch-1 — integer_psnr_hip and float_ansnr_hip Real Kernels |
| ADR-0373 | HIP Batch-2 — float_motion_hip Real Kernel |
| ADR-0375 | HIP batch-3 — float_moment_hip and float_ssim_hip real kernels |
| ADR-0376 | Fix silent error-swallow in Vulkan buffer-invalidate readback functions |
| ADR-0377 | HIP batch-4 — ciede_hip and integer_motion_v2_hip real kernels |
| ADR-0378 | Per-picture CUDA streams must use CU_STREAM_NON_BLOCKING |
| ADR-0385 | Feature-extractor deduplication by provided-feature names |
| ADR-0391 | ciede2000 Vulkan NVIDIA places=4 fork debt is a structural f32/f64 precision gap |
| ADR-0410 | ssimulacra2_cuda GPU module leak + per-scale malloc removal |
| ADR-0415 | CAMBI SYCL port — closes last CUDA-to-SYCL parity gap |
| ADR-0420 | Metal backend runtime (T8-1b) |
| ADR-0421 | Metal first kernel — integer_motion_v2 (T8-1c) |
| ADR-0422 | CLI HIP and Metal Backend Selectors |
| ADR-0423 | Metal IOSurface zero-copy import (T8-IOS) |
| ADR-0445 | Persistent VkPipelineCache for Vulkan compute backend |
| ADR-0451 | Local dev-MCP container for live probing |
| ADR-0454 | VIF CUDA shared-memory staging for horizontal and vertical filter passes |
| ADR-0458 | SYCL CAMBI queue-sync collapse + SSIM horizontal SLM staging |
| ADR-0464 | CAMBI CUDA spatial-mask shared-memory tile |
| ADR-0484 | Extend kernel-scaffolding.md with HIP and Metal lifecycle contract |
| ADR-0486 | Codify the three-function GPU backend context-API contract in docs |
| ADR-0488 | Shared once-snapshot helper for GPU dispatch env variables |
| ADR-0489 | CAMBI SYCL — Replace GPU-to-GPU q.wait() Calls with Event Chains (SY-1) |
| ADR-0514 | dev-MCP container exposes every host GPU backend (CUDA + SYCL + Vulkan + HIP) |
| ADR-0519 | Implement vmaf_hip_import_state to unblock --backend hip |
| ADR-0523 | Register vmaf_fex_integer_motion_hip in the extractor list |
| ADR-0530 | HIP feature-extractor flag promotion and HIP_DEVICE picture-buffer type |
| ADR-0533 | Full HIP feature-extractor registration sweep |
| ADR-0537 | HIP integer VIF kernel crash fix — filter upload, bounds, HtoD staging |
| ADR-0539 | integer ADM HIP kernels — real implementation replacing weak HSACO stubs |
| ADR-0541 | Pin dev-MCP container Intel NEO + ROCm runtimes to versions matching the host kernel |
| ADR-0552 | Deterministic wavefront reduction for integer_vif_hip horizontal kernels |
| ADR-0563 | HIP extractor audit — verification of 9 remaining scaffold claims |
| ADR-0564 | Real integer_ssim GPU kernels (CUDA, HIP, SYCL) — replace silent float_ssim substitution |
| ADR-0567 | Real On-Device GPU Kernels for speed_chroma and speed_temporal (4 Backends) |
| ADR-0568 | Default sycl_icpx_aot_targets to full Intel arch list |
| ADR-0587 | Real Metal Compute Kernels for CAMBI |
| ADR-0593 | HIP integer_moment kernel — register real HSACO blob alongside psnr / psnr_hvs |
| ADR-0667 | vmaf-tune score backend native priority |
| ADR-0699 | VMAFX Helm Chart and Kubernetes Manifests with 3-Vendor GPU Device-Plugin Support |
| ADR-0726 | Drop Vulkan backend |
| ADR-0804 | Add vmaf_context_get_backend — additive ABI introspection |
| ADR-0852 | Wire speed_chroma_hip and speed_temporal_hip into HIP Build and Dispatch |
| ADR-0883 | HIP kernel parity-test coverage round 2 |
| ADR-0884 | SYCL kernel coverage round 2 — five additional CPU-vs-SYCL parity gates |
| ADR-0886 | CUDA kernel parity test coverage — round 2 gap-fill |
| ADR-0928 | VmafPicture v2 — explicit per-backend GPU state |
| ADR-0945 | HIP kernel parity-test coverage round 3 |
| ADR-0946 | SYCL kernel coverage round 3 (float family + PSNR-HVS) |
| ADR-0954 | Host-only unit test for shared GPU dispatch runtime |
| ADR-0957 | SYCL kernel coverage round 4 (float_moment + SpEED + SSIMULACRA2) |
| ADR-0958 | HIP kernel parity-test coverage round 4 |
| ADR-0959 | Metal kernel parity coverage round 4 — closeout |
| ADR-0982 | GPU runtime bug audit — round 26 (init/teardown leak sweep) |
| ADR-0985 | SYCL parity divergence investigation — float_ssim + ssimulacra2 on Arc A380 |
| ADR-0989 | Wire motion_add_uv through integer_motion_sycl; emit warning on motion_five_frame_window |
| ADR-1001 | SYCL parity round 5 — CAMBI CPU vs. SYCL parity gate |
| ADR-1004 | HIP kernel parity-test coverage round 5 |
| ADR-1030 | HIP adm_decouple dangling body + VIF wavefront 32-bit carry + Metal motion vertical halo |
| ADR-1034 | Fix SYCL integer_vif rd_stride OOB on odd widths and integer_motion UV queue sync gap |
| ADR-1106 | HIP motion_v2 mirror is reflect-101 (-2), correcting ADR-0377's -1 parity claim |
| ADR-1121 | SYCL QSV zero-copy — P010 pixel normalization and separate-session decode contract |
| ADR-1129 | Align release containers with the published tag and runtime ABI |
| ADR-1154 | AMD ROCm HIP Backend Gap Closure and Extractor Promotion |
| ADR-1167 | Row-level rounding accumulator and border row selection for integer ADM GPU kernels |
| ADR-1176 | Metal motion_v2 mirror closeout and reflect-101 parity contract |
| ADR-1177 | Containerised self-hosted GitHub Actions runner for Intel Arc SYCL parity CI |
| ADR-1179 | Fix Intel Arc SYCL Crashes and Default Model Resolution Divergence |
| ADR-1264 | The HIP scaffold posture reports -ENOSYS, and its tests check both sites |
| ADR-1296 | GPU init failure paths are tested device-free, by compiling the backend TU against runtime stubs |
| ADR-1299 | Compute MS-SSIM chroma on SYCL |
| ADR-1300 | Install CUDA on every CI leg from NVIDIA's own distribution |
| ADR-1312 | GPU option aliases match CPU collector keys |
| ADR-1316 | Mark extractor options whose implementation is default-only |
| ADR-1319 | Admit self-hosted GPU jobs through a live fail-closed probe |
| ADR-1320 | CUDA fatbin and HIP HSACO kernel header dependency tracking |
| ADR-1324 | Resolve GPU float-SSIM auto-scale before backend initialization |
| ADR-1357 | Run the SYCL CAMBI extractor entirely on the device |
| ADR-1358 | The SYCL SpEED twins are device-resident, with the 25x25 linear algebra on the device and exact fp32 arithmetic |
| ADR-1359 | The CLI maps --feature <cpu-name> to the explicit --backend's twin |
| ADR-1360 | Generate SYCL AOT images at compile time and fail the build when they are missing |
| ADR-1362 | The SYCL integer ADM twin computes AIM on the device and finalises every ADM output in the CPU's float arithmetic |
| ADR-1363 | The SYCL ssimulacra2 twin is device-resident, and float_ms_ssim_sycl waits once per frame |
| ADR-1364 | Register the SYCL device images of a Windows MSVC build through one explicit device link |
| ADR-1367 | Every SYCL feature TU compiles with contraction off and correctly rounded fp32 division and square root |
| ADR-1368 | Build and run the oneAPI release image on Debian 13 with pinned Intel packages |
| ADR-1369 | SYCL twins read the planes the state uploads once per frame; opt-in shared chroma planes and a device-side slot fence |
| ADR-1370 | float_ssim_sycl decimates on the device, bit-identical to the CPU |
| ADR-1378 | Run the HIP CAMBI extractor entirely on the device |
| ADR-1379 | Run the CUDA CAMBI extractor entirely on the device |
| ADR-1380 | The CUDA SpEED twins are device-resident, with the 25x25 linear algebra on the device and CPU-exact fp32 arithmetic |
| ADR-1384 | Run the HIP SpEED twins entirely on the device, in the CPU's fp32 arithmetic |
| ADR-1390 | The HIP ssimulacra2 twin is device-resident with tiled row-pass Gaussian blur |
| ADR-1391 | The CUDA ssimulacra2 twin is device-resident |
| ADR-1395 | SYCL kernels use no scratch memory on Intel GPUs |
| ADR-1399 | float_ssim_cuda decimates and convolves on the device with the CPU's arithmetic |
| ADR-1403 | Every CUDA feature kernel compiles without FMA contraction, and a twin spells the fused operations its reference performs |
| ADR-1405 | float_ssim_hip decimates on the device, bit-identical to the CPU |
| ADR-1407 | Every HIP kernel compiles with contraction off and correctly rounded fp32 division and square root |
| ADR-1408 | A VmafContext uploads each plane of a frame once and every HIP twin reads that copy |
| ADR-1498 | Metal twins take the exact designs of their CUDA, HIP and SYCL twins, on a strict FP kernel policy |
| ADR-1523 | HIP twins run on the device of the imported state |
| ADR-1525 | adm_hip computes AIM on the device and is dispatched |
| ADR-1547 | The Helm chart derives the GPU resource name from the vendor and the Intel kernel driver, with an explicit override |
| ADR-1571 | a GPU dispatch variable is documented only when library code reads it; VMAF_CUDA_DISPATCH is read at extractor init, VMAF_HIP_DISPATCH and its predicate are removed |
| ADR-1590 | Every build stores its GPU device code compressed at the toolchain's strongest setting, and the build refuses raw device code |
| ADR-1685 | A post-1.0 embedding milestone: zero-copy device-frame import with fences, asynchronous window scores, and an unchanged licence |
| ADR-1806 | Run Metal kernel files on the host through a Metal Shading Language shim in tests |
| ADR-1829 | RC4 owns the whole device-memory import API, not only the first full Rust metric |
| ADR-1830 | vif_sycl runs at SIMD-16 only; the SIMD-32 kernels and VMAF_SYCL_VIF_SUBGROUP_SIZE are removed |
| ADR-1874 | The vmaf CLI reports its usable backends; the score-backend selectors read that report |
| ADR-1880 | Format envelope and device-targeted scoring in the 1.0.0 candidates |
| ADR-1900 | Deterministic verification and state contract for VPL decode retry ceiling |
| ADR-1929 | VMAFx device frames, fences and the import rule: the shared contract the backend lanes implement |
| ADR-2023 | VMAFx device frames on CUDA: one library stream per device, fences on it, release events recorded where the frame is released |
| ADR-2056 | float_adm files its debug ratio unsuffixed; a second debug instance is refused |
| ADR-2091 | VMAFx device frames on SYCL: readers copy on the device behind the frame's ready event, the release waits on every reader |
| ADR-2092 | VMAFx device frames on HIP: one library stream per device copies every frame for the twins, dma-bufs as external memory, sync_file checked on the host |
| ADR-2132 | HIP imports GL textures through EGL dma-buf export; the runtime's GL interop is not used |
| ADR-2133 | Device import takes 4:2:2 and 4:4:4: semi-planar NV16 / NV24 family and packed Y210 / Y410, one CPU reference, converted on the device |
| ADR-2478 | Rust replaces the host-side C and C++ through 3.0, behind the unchanged C ABI |