Skip to content

ADRs tagged hip

Auto-generated by scripts/docs/generate-adr-by-tag.sh. Edit ADR Tags: lines to update.

157 ADR(s) carry this tag.

ID Title
ADR-0212 HIP (AMD ROCm) compute backend — scaffold-only audit-first PR (T7-10)
ADR-0241 HIP first-consumer kernel — integer_psnr_hip via mirrored kernel-template
ADR-0254 HIP second-consumer kernel — float_psnr_hip via mirrored kernel-template
ADR-0259 HIP third-consumer kernel — ciede_hip via mirrored kernel-template
ADR-0260 HIP fourth-consumer kernel — float_moment_hip via mirrored kernel-template
ADR-0266 HIP fifth kernel-template consumer — float_ansnr_hip
ADR-0267 HIP sixth kernel-template consumer — motion_v2_hip
ADR-0273 HIP seventh kernel-template consumer — float_motion_hip
ADR-0274 HIP eighth kernel-template consumer — float_ssim_hip
ADR-0315 Vendor-neutral VVC encode strategy — tiered Tier-1-now / Tier-2-backlog / Tier-3-revisit
ADR-0345 cambi × {CUDA, SYCL, HIP} GPU port strategy
ADR-0372 HIP Batch-1 — integer_psnr_hip and float_ansnr_hip Real Kernels
ADR-0373 HIP Batch-2 — float_motion_hip Real Kernel
ADR-0374 Build-time-optional public APIs return -ENOSYS when disabled
ADR-0375 HIP batch-3 — float_moment_hip and float_ssim_hip real kernels
ADR-0377 HIP batch-4 — ciede_hip and integer_motion_v2_hip real kernels
ADR-0380 FFmpeg libvmaf filter — HIP backend selector patch (0011)
ADR-0422 CLI HIP and Metal Backend Selectors
ADR-0451 Local dev-MCP container for live probing
ADR-0460 Dispatch-strategy registry audit 2026-05-15
ADR-0468 HIP float_adm real kernel (ninth HIP consumer)
ADR-0469 float_psnr HIP twin — wire enable_chroma option
ADR-0471 Add enable_chroma to integer_psnr_hip (chroma parity with CUDA/SYCL/Vulkan twins)
ADR-0484 Extend kernel-scaffolding.md with HIP and Metal lifecycle contract
ADR-0486 Codify the three-function GPU backend context-API contract in docs
ADR-0514 dev-MCP container exposes every host GPU backend (CUDA + SYCL + Vulkan + HIP)
ADR-0519 Implement vmaf_hip_import_state to unblock --backend hip
ADR-0523 Register vmaf_fex_integer_motion_hip in the extractor list
ADR-0530 HIP feature-extractor flag promotion and HIP_DEVICE picture-buffer type
ADR-0533 Full HIP feature-extractor registration sweep
ADR-0537 HIP integer VIF kernel crash fix — filter upload, bounds, HtoD staging
ADR-0539 integer ADM HIP kernels — real implementation replacing weak HSACO stubs
ADR-0541 Pin dev-MCP container Intel NEO + ROCm runtimes to versions matching the host kernel
ADR-0542 Full GPU backend plumbing in the dev-mcp container
ADR-0552 Deterministic wavefront reduction for integer_vif_hip horizontal kernels
ADR-0561 0561-hip-gfx-targets-fallback-widening.md
ADR-0563 HIP extractor audit — verification of 9 remaining scaffold claims
ADR-0564 Real integer_ssim GPU kernels (CUDA, HIP, SYCL) — replace silent float_ssim substitution
ADR-0566 0566-hip-vif-per-feature-places4-gate.md
ADR-0567 Real On-Device GPU Kernels for speed_chroma and speed_temporal (4 Backends)
ADR-0576 ffmpeg-patches n8.1.1 full-feature-exposure sync
ADR-0592 Remove float_vif_score weak HSACO stub now that real HIP kernel ships
ADR-0593 HIP integer_moment kernel — register real HSACO blob alongside psnr / psnr_hvs
ADR-0594 Per-kernel hip_cu_extra_flags dispatch — disable FMA contraction for ssimulacra2_blur HIP HSACO
ADR-0596 Delete orphan and duplicate HIP/CUDA translation units
ADR-0599 Cross-Backend Parity Audit — Full Extractor Matrix (2026-05-18)
ADR-0604 Add Renovate customManager for ROCm apt-repo tracking
ADR-0605 Renovate customManagers for all dev/Containerfile pinned dependencies
ADR-0623 Scaffold audit P2 — half-finished implementation fixes
ADR-0639 Scaffold-audit P1 — backend precheck, HIP picture, mobilesal bpc, DNN multi-output
ADR-0667 vmaf-tune score backend native priority
ADR-0688 HIP wave32 carry-preserving int64 reduction for VIF and motion kernels
ADR-0699 VMAFX Helm Chart and Kubernetes Manifests with 3-Vendor GPU Device-Plugin Support
ADR-0759 HIP ADM — AdmBufferHip passed by pointer (F3 fix)
ADR-0777 Thread-Safety Audit — CUDA / SYCL / HIP Backends
ADR-0780 NOLINT Cluster Refactor — Slab Allocator, SYCL Stride, ADM Band-Size
ADR-0787 0787-libvmaf-api-error-path-audit.md
ADR-0852 Wire speed_chroma_hip and speed_temporal_hip into HIP Build and Dispatch
ADR-0868 GPU backend kernel parity-test coverage gap-fill
ADR-0883 HIP kernel parity-test coverage round 2
ADR-0928 VmafPicture v2 — explicit per-backend GPU state
ADR-0945 HIP kernel parity-test coverage round 3
ADR-0949 HIP motion3 parity test skips cleanly when HIPCC kernels are not built
ADR-0950 Fix symmetric "adm" vs "adm_hip" feature-name bug in test_hip_adm_parity and add ENOSYS skip
ADR-0954 Host-only unit test for shared GPU dispatch runtime
ADR-0958 HIP kernel parity-test coverage round 4
ADR-0964 Implement speed_internal.c and wire speed_{chroma,temporal}_{hip,sycl}
ADR-1004 HIP kernel parity-test coverage round 5
ADR-1025 R6 CUDA/HIP kernel correctness fixes
ADR-1030 HIP adm_decouple dangling body + VIF wavefront 32-bit carry + Metal motion vertical halo
ADR-1071 Promote HIP ms_ssim_vert_lcs to double precision (ADR-0990 parity)
ADR-1078 ms_ssim option parity across HIP and SYCL backends
ADR-1100 Skip GPU-flagged extractors when flags == 0 in vmaf_get_feature_extractor_by_feature_name
ADR-1103 1103-hip-vif-mirror2-boundary.md
ADR-1106 HIP motion_v2 mirror is reflect-101 (-2), correcting ADR-0377's -1 parity claim
ADR-1142 Whole-codebase standards; lint debt only ratchets down
ADR-1154 AMD ROCm HIP Backend Gap Closure and Extractor Promotion
ADR-1167 Row-level rounding accumulator and border row selection for integer ADM GPU kernels
ADR-1185 Per-backend performance baselines are median-of-N, one backend per build dir
ADR-1191 Integer ADM rejects CSF configurations its fixed-point storage cannot represent
ADR-1193 Opt-in uncapped option splits the PSNR infinity sentinel from the truncation
ADR-1194 One integer-ADM angle_flag predicate for every backend
ADR-1202 GPU SpEED-chroma twins report singularity separately from failure
ADR-1204 GPU ADM contrast-masking twins clamp the far edge instead of mirroring it
ADR-1205 The ssimulacra2 FMA unification extends to the scalar fallback and every GPU host copy
ADR-1211 integer_adm_hip stages the luma plane onto the device before launching
ADR-1212 The GPU float_moment twins normalise by the bit-depth scaler on the host
ADR-1213 ciede_hip sizes its chroma staging with the picture's ceil dimensions
ADR-1214 The float-ADM GPU twins ignore adm_csf_scale in Watson mode and share the CPU's option aliases
ADR-1216 The GPU motion3 twins apply motion_fps_weight exactly once
ADR-1217 The GPU float-VIF kernels read vif_sigma_nsq and vif_enhn_gain_limit from their options
ADR-1218 The GPU SpEED twins zero the device solution and report singularity from the temporal path
ADR-1219 The HIP and Metal CAMBI twins use the shared TVI bisection and the CPU's border rules
ADR-1220 The GPU float-ADM kernels honour adm_p_norm, adm_bypass_cm and adm_skip_scale0
ADR-1221 clip_db is a ceiling on the MS-SSIM dB output, not a clamp on the linear score
ADR-1225 Migrate the HIP backend to ROCm 10.0.0, installed from digest-pinned container images
ADR-1228 A recurring "faster than upstream, and still exact" milestone
ADR-1237 CAMBI Anti-Dithering AVX2 Vectorization, SpEED SIMD QR Dispatch, and Threaded GPU Flush Alignment
ADR-1263 __HIP_PLATFORM_AMD__ is declared once by the build, not by each source
ADR-1264 The HIP scaffold posture reports -ENOSYS, and its tests check both sites
ADR-1289 The three SSIMULACRA2 HIP host functions may be split; their no-split citations are withdrawn
ADR-1290 The tidy ratchet owns its per-lane compilation database
ADR-1296 GPU init failure paths are tested device-free, by compiling the backend TU against runtime stubs
ADR-1320 CUDA fatbin and HIP HSACO kernel header dependency tracking
ADR-1325 Normalize integer ADM Barten weights with one exponent per scale
ADR-1377 HIP motion differences the frames before the blur, in one shared kernel, and waits only in collect
ADR-1378 Run the HIP CAMBI extractor entirely on the device
ADR-1381 HIP tile loads and the ADM vertical DWT clamp their rows; vif_hip hands frames below 16 pixels to the CPU
ADR-1382 HIP PSNR, SSIM and float-motion twins take the CPU option tables
ADR-1384 Run the HIP SpEED twins entirely on the device, in the CPU's fp32 arithmetic
ADR-1386 One script runs the home GPU box's RC3 verify commands, row by row, under per-device locks
ADR-1390 The HIP ssimulacra2 twin is device-resident with tiled row-pass Gaussian blur
ADR-1400 integer_ssim_hip sums small frames in the CPU's raster order
ADR-1401 psnr_hvs_sycl and psnr_hvs_hip return the CPU's scores bit for bit; the fp64-free masking threshold is an integer square root
ADR-1402 Integer ADM keeps the scale-0 masking centre tap in int32 and clamps the excess in int64
ADR-1404 float_motion_hip emits motion3 and implements every CPU float_motion option on the device
ADR-1405 float_ssim_hip decimates on the device, bit-identical to the CPU
ADR-1407 Every HIP kernel compiles with contraction off and correctly rounded fp32 division and square root
ADR-1408 A VmafContext uploads each plane of a frame once and every HIP twin reads that copy
ADR-1419 float_motion_hip stores its differences transposed and adds each row in the CPU's order
ADR-1423 adm_hip takes its weights, shifts and score conclusion from the CPU, folds the denominator once per row and clears its accumulators after the upload
ADR-1427 A HIP frame queues its accumulator clears after its upload
ADR-1435 vif_hip reads the CPU's log2 table instead of evaluating log2f() on the device, and returns the CPU's scores bit for bit
ADR-1437 motion_hip, motion_v2_hip, psnr_hip, integer_ms_ssim_hip and cambi_hip are declared exact twins; float_psnr_hip and float_moment_hip are not
ADR-1438 integer_ssim_hip adds its terms in the CPU's raster order at every frame size and returns the CPU's score bit for bit
ADR-1440 float_psnr_hip adds its squared differences as integers and returns the CPU's score bit for bit
ADR-1441 float_ssim_hip forms its window sums through the arithmetic float_ms_ssim_hip shares with the CPU and returns the CPU's score bit for bit
ADR-1444 float_vif_hip runs the arithmetic of the CUDA twin from one shared header and returns the CPU's scores bit for bit
ADR-1445 ssimulacra2_hip evaluates the CPU's fp64 terms and returns the sums of the CPU's loops
ADR-1447 float_moment_hip adds the float squares the CPU adds, and is bit-identical while the CPU's own sum is exact
ADR-1448 ciede_hip runs the CPU's arithmetic in fp32 pairs, from the header the SYCL twin runs; what remains is glibc's powf and the last bits of a pair
ADR-1452 the gate bounds speed_chroma_hip against the CPU by what glibc's log2f adds, as it does for the CUDA twin
ADR-1458 float_adm_hip runs the CPU's arithmetic from the header the CUDA twin runs and returns the CPU's scores bit for bit
ADR-1460 speed_temporal becomes a parity-gate feature with a derived bound, and every registered twin of a gated backend has to be a gate cell
ADR-1472 The integer ADM weight limits follow from the contrast-masking cube and the largest wavelet coefficient of a scale
ADR-1477 SpEED evaluates Netflix's fp64 expressions again, and its GPU twins form the entropies and the score on the host with the same statements, so every twin returns the CPU's scores bit for bit on any C library
ADR-1491 The CUDA, SYCL and HIP motion twins compute motion_five_frame_window: the frame two back on the device, the CPU's window function on the host
ADR-1497 The float_moment twins form the CPU's rounded second-moment sum past 2^53 units, and return the CPU's bits on every frame
ADR-1499 The float_psnr twins add each row's exact sum in the CPU's order, and return the CPU's bits on every frame
ADR-1511 An AMD GPU tester image that ships only the ROCm runtime files the HIP build loads, with the source of its LGPL parts
ADR-1517 The production GPU images are built on Debian 13, ship only the vendor files libvmaf loads, and share the tester images' licence records
ADR-1523 HIP twins run on the device of the imported state
ADR-1525 adm_hip computes AIM on the device and is dispatched
ADR-1561 integer VIF converts its residual variance through vif_sv_sq(), which returns x86's value without the undefined double to int32_t conversion
ADR-1571 a GPU dispatch variable is documented only when library code reads it; VMAF_CUDA_DISPATCH is read at extractor init, VMAF_HIP_DISPATCH and its predicate are removed
ADR-1590 Every build stores its GPU device code compressed at the toolchain's strongest setting, and the build refuses raw device code
ADR-1601 integer VIF forms the denominator log argument sigma_nsq + sigma1_sq in uint32_t
ADR-1685 A post-1.0 embedding milestone: zero-copy device-frame import with fences, asynchronous window scores, and an unchanged licence
ADR-1829 RC4 owns the whole device-memory import API, not only the first full Rust metric
ADR-1874 The vmaf CLI reports its usable backends; the score-backend selectors read that report
ADR-1917 The integer ADM scale-0 horizontal and vertical weight limit is 43900, set by the CSF stage's 16-bit magnitude
ADR-1918 Samples above 2^bpc - 1 are invalid input; an opt-in check refuses them
ADR-1929 VMAFx device frames, fences and the import rule: the shared contract the backend lanes implement
ADR-2092 VMAFx device frames on HIP: one library stream per device copies every frame for the twins, dma-bufs as external memory, sync_file checked on the host
ADR-2132 HIP imports GL textures through EGL dma-buf export; the runtime's GL interop is not used
ADR-2134 adm_cm_aim_line_kernel_4 gets its own register budget of 209 for the exact scale-0 angle flag
ADR-2796 Every clang-tidy ratchet lane is a hosted, path-routed required check