ADRs tagged hip¶
Auto-generated by scripts/docs/generate-adr-by-tag.sh. Edit ADR Tags: lines to update.
157 ADR(s) carry this tag.
| ID | Title |
|---|---|
| ADR-0212 | HIP (AMD ROCm) compute backend — scaffold-only audit-first PR (T7-10) |
| ADR-0241 | HIP first-consumer kernel — integer_psnr_hip via mirrored kernel-template |
| ADR-0254 | HIP second-consumer kernel — float_psnr_hip via mirrored kernel-template |
| ADR-0259 | HIP third-consumer kernel — ciede_hip via mirrored kernel-template |
| ADR-0260 | HIP fourth-consumer kernel — float_moment_hip via mirrored kernel-template |
| ADR-0266 | HIP fifth kernel-template consumer — float_ansnr_hip |
| ADR-0267 | HIP sixth kernel-template consumer — motion_v2_hip |
| ADR-0273 | HIP seventh kernel-template consumer — float_motion_hip |
| ADR-0274 | HIP eighth kernel-template consumer — float_ssim_hip |
| ADR-0315 | Vendor-neutral VVC encode strategy — tiered Tier-1-now / Tier-2-backlog / Tier-3-revisit |
| ADR-0345 | cambi × {CUDA, SYCL, HIP} GPU port strategy |
| ADR-0372 | HIP Batch-1 — integer_psnr_hip and float_ansnr_hip Real Kernels |
| ADR-0373 | HIP Batch-2 — float_motion_hip Real Kernel |
| ADR-0374 | Build-time-optional public APIs return -ENOSYS when disabled |
| ADR-0375 | HIP batch-3 — float_moment_hip and float_ssim_hip real kernels |
| ADR-0377 | HIP batch-4 — ciede_hip and integer_motion_v2_hip real kernels |
| ADR-0380 | FFmpeg libvmaf filter — HIP backend selector patch (0011) |
| ADR-0422 | CLI HIP and Metal Backend Selectors |
| ADR-0451 | Local dev-MCP container for live probing |
| ADR-0460 | Dispatch-strategy registry audit 2026-05-15 |
| ADR-0468 | HIP float_adm real kernel (ninth HIP consumer) |
| ADR-0469 | float_psnr HIP twin — wire enable_chroma option |
| ADR-0471 | Add enable_chroma to integer_psnr_hip (chroma parity with CUDA/SYCL/Vulkan twins) |
| ADR-0484 | Extend kernel-scaffolding.md with HIP and Metal lifecycle contract |
| ADR-0486 | Codify the three-function GPU backend context-API contract in docs |
| ADR-0514 | dev-MCP container exposes every host GPU backend (CUDA + SYCL + Vulkan + HIP) |
| ADR-0519 | Implement vmaf_hip_import_state to unblock --backend hip |
| ADR-0523 | Register vmaf_fex_integer_motion_hip in the extractor list |
| ADR-0530 | HIP feature-extractor flag promotion and HIP_DEVICE picture-buffer type |
| ADR-0533 | Full HIP feature-extractor registration sweep |
| ADR-0537 | HIP integer VIF kernel crash fix — filter upload, bounds, HtoD staging |
| ADR-0539 | integer ADM HIP kernels — real implementation replacing weak HSACO stubs |
| ADR-0541 | Pin dev-MCP container Intel NEO + ROCm runtimes to versions matching the host kernel |
| ADR-0542 | Full GPU backend plumbing in the dev-mcp container |
| ADR-0552 | Deterministic wavefront reduction for integer_vif_hip horizontal kernels |
| ADR-0561 | 0561-hip-gfx-targets-fallback-widening.md |
| ADR-0563 | HIP extractor audit — verification of 9 remaining scaffold claims |
| ADR-0564 | Real integer_ssim GPU kernels (CUDA, HIP, SYCL) — replace silent float_ssim substitution |
| ADR-0566 | 0566-hip-vif-per-feature-places4-gate.md |
| ADR-0567 | Real On-Device GPU Kernels for speed_chroma and speed_temporal (4 Backends) |
| ADR-0576 | ffmpeg-patches n8.1.1 full-feature-exposure sync |
| ADR-0592 | Remove float_vif_score weak HSACO stub now that real HIP kernel ships |
| ADR-0593 | HIP integer_moment kernel — register real HSACO blob alongside psnr / psnr_hvs |
| ADR-0594 | Per-kernel hip_cu_extra_flags dispatch — disable FMA contraction for ssimulacra2_blur HIP HSACO |
| ADR-0596 | Delete orphan and duplicate HIP/CUDA translation units |
| ADR-0599 | Cross-Backend Parity Audit — Full Extractor Matrix (2026-05-18) |
| ADR-0604 | Add Renovate customManager for ROCm apt-repo tracking |
| ADR-0605 | Renovate customManagers for all dev/Containerfile pinned dependencies |
| ADR-0623 | Scaffold audit P2 — half-finished implementation fixes |
| ADR-0639 | Scaffold-audit P1 — backend precheck, HIP picture, mobilesal bpc, DNN multi-output |
| ADR-0667 | vmaf-tune score backend native priority |
| ADR-0688 | HIP wave32 carry-preserving int64 reduction for VIF and motion kernels |
| ADR-0699 | VMAFX Helm Chart and Kubernetes Manifests with 3-Vendor GPU Device-Plugin Support |
| ADR-0759 | HIP ADM — AdmBufferHip passed by pointer (F3 fix) |
| ADR-0777 | Thread-Safety Audit — CUDA / SYCL / HIP Backends |
| ADR-0780 | NOLINT Cluster Refactor — Slab Allocator, SYCL Stride, ADM Band-Size |
| ADR-0787 | 0787-libvmaf-api-error-path-audit.md |
| ADR-0852 | Wire speed_chroma_hip and speed_temporal_hip into HIP Build and Dispatch |
| ADR-0868 | GPU backend kernel parity-test coverage gap-fill |
| ADR-0883 | HIP kernel parity-test coverage round 2 |
| ADR-0928 | VmafPicture v2 — explicit per-backend GPU state |
| ADR-0945 | HIP kernel parity-test coverage round 3 |
| ADR-0949 | HIP motion3 parity test skips cleanly when HIPCC kernels are not built |
| ADR-0950 | Fix symmetric "adm" vs "adm_hip" feature-name bug in test_hip_adm_parity and add ENOSYS skip |
| ADR-0954 | Host-only unit test for shared GPU dispatch runtime |
| ADR-0958 | HIP kernel parity-test coverage round 4 |
| ADR-0964 | Implement speed_internal.c and wire speed_{chroma,temporal}_{hip,sycl} |
| ADR-1004 | HIP kernel parity-test coverage round 5 |
| ADR-1025 | R6 CUDA/HIP kernel correctness fixes |
| ADR-1030 | HIP adm_decouple dangling body + VIF wavefront 32-bit carry + Metal motion vertical halo |
| ADR-1071 | Promote HIP ms_ssim_vert_lcs to double precision (ADR-0990 parity) |
| ADR-1078 | ms_ssim option parity across HIP and SYCL backends |
| ADR-1100 | Skip GPU-flagged extractors when flags == 0 in vmaf_get_feature_extractor_by_feature_name |
| ADR-1103 | 1103-hip-vif-mirror2-boundary.md |
| ADR-1106 | HIP motion_v2 mirror is reflect-101 (-2), correcting ADR-0377's -1 parity claim |
| ADR-1142 | Whole-codebase standards; lint debt only ratchets down |
| ADR-1154 | AMD ROCm HIP Backend Gap Closure and Extractor Promotion |
| ADR-1167 | Row-level rounding accumulator and border row selection for integer ADM GPU kernels |
| ADR-1185 | Per-backend performance baselines are median-of-N, one backend per build dir |
| ADR-1191 | Integer ADM rejects CSF configurations its fixed-point storage cannot represent |
| ADR-1193 | Opt-in uncapped option splits the PSNR infinity sentinel from the truncation |
| ADR-1194 | One integer-ADM angle_flag predicate for every backend |
| ADR-1202 | GPU SpEED-chroma twins report singularity separately from failure |
| ADR-1204 | GPU ADM contrast-masking twins clamp the far edge instead of mirroring it |
| ADR-1205 | The ssimulacra2 FMA unification extends to the scalar fallback and every GPU host copy |
| ADR-1211 | integer_adm_hip stages the luma plane onto the device before launching |
| ADR-1212 | The GPU float_moment twins normalise by the bit-depth scaler on the host |
| ADR-1213 | ciede_hip sizes its chroma staging with the picture's ceil dimensions |
| ADR-1214 | The float-ADM GPU twins ignore adm_csf_scale in Watson mode and share the CPU's option aliases |
| ADR-1216 | The GPU motion3 twins apply motion_fps_weight exactly once |
| ADR-1217 | The GPU float-VIF kernels read vif_sigma_nsq and vif_enhn_gain_limit from their options |
| ADR-1218 | The GPU SpEED twins zero the device solution and report singularity from the temporal path |
| ADR-1219 | The HIP and Metal CAMBI twins use the shared TVI bisection and the CPU's border rules |
| ADR-1220 | The GPU float-ADM kernels honour adm_p_norm, adm_bypass_cm and adm_skip_scale0 |
| ADR-1221 | clip_db is a ceiling on the MS-SSIM dB output, not a clamp on the linear score |
| ADR-1225 | Migrate the HIP backend to ROCm 10.0.0, installed from digest-pinned container images |
| ADR-1228 | A recurring "faster than upstream, and still exact" milestone |
| ADR-1237 | CAMBI Anti-Dithering AVX2 Vectorization, SpEED SIMD QR Dispatch, and Threaded GPU Flush Alignment |
| ADR-1263 | __HIP_PLATFORM_AMD__ is declared once by the build, not by each source |
| ADR-1264 | The HIP scaffold posture reports -ENOSYS, and its tests check both sites |
| ADR-1289 | The three SSIMULACRA2 HIP host functions may be split; their no-split citations are withdrawn |
| ADR-1290 | The tidy ratchet owns its per-lane compilation database |
| ADR-1296 | GPU init failure paths are tested device-free, by compiling the backend TU against runtime stubs |
| ADR-1320 | CUDA fatbin and HIP HSACO kernel header dependency tracking |
| ADR-1325 | Normalize integer ADM Barten weights with one exponent per scale |
| ADR-1377 | HIP motion differences the frames before the blur, in one shared kernel, and waits only in collect |
| ADR-1378 | Run the HIP CAMBI extractor entirely on the device |
| ADR-1381 | HIP tile loads and the ADM vertical DWT clamp their rows; vif_hip hands frames below 16 pixels to the CPU |
| ADR-1382 | HIP PSNR, SSIM and float-motion twins take the CPU option tables |
| ADR-1384 | Run the HIP SpEED twins entirely on the device, in the CPU's fp32 arithmetic |
| ADR-1386 | One script runs the home GPU box's RC3 verify commands, row by row, under per-device locks |
| ADR-1390 | The HIP ssimulacra2 twin is device-resident with tiled row-pass Gaussian blur |
| ADR-1400 | integer_ssim_hip sums small frames in the CPU's raster order |
| ADR-1401 | psnr_hvs_sycl and psnr_hvs_hip return the CPU's scores bit for bit; the fp64-free masking threshold is an integer square root |
| ADR-1402 | Integer ADM keeps the scale-0 masking centre tap in int32 and clamps the excess in int64 |
| ADR-1404 | float_motion_hip emits motion3 and implements every CPU float_motion option on the device |
| ADR-1405 | float_ssim_hip decimates on the device, bit-identical to the CPU |
| ADR-1407 | Every HIP kernel compiles with contraction off and correctly rounded fp32 division and square root |
| ADR-1408 | A VmafContext uploads each plane of a frame once and every HIP twin reads that copy |
| ADR-1419 | float_motion_hip stores its differences transposed and adds each row in the CPU's order |
| ADR-1423 | adm_hip takes its weights, shifts and score conclusion from the CPU, folds the denominator once per row and clears its accumulators after the upload |
| ADR-1427 | A HIP frame queues its accumulator clears after its upload |
| ADR-1435 | vif_hip reads the CPU's log2 table instead of evaluating log2f() on the device, and returns the CPU's scores bit for bit |
| ADR-1437 | motion_hip, motion_v2_hip, psnr_hip, integer_ms_ssim_hip and cambi_hip are declared exact twins; float_psnr_hip and float_moment_hip are not |
| ADR-1438 | integer_ssim_hip adds its terms in the CPU's raster order at every frame size and returns the CPU's score bit for bit |
| ADR-1440 | float_psnr_hip adds its squared differences as integers and returns the CPU's score bit for bit |
| ADR-1441 | float_ssim_hip forms its window sums through the arithmetic float_ms_ssim_hip shares with the CPU and returns the CPU's score bit for bit |
| ADR-1444 | float_vif_hip runs the arithmetic of the CUDA twin from one shared header and returns the CPU's scores bit for bit |
| ADR-1445 | ssimulacra2_hip evaluates the CPU's fp64 terms and returns the sums of the CPU's loops |
| ADR-1447 | float_moment_hip adds the float squares the CPU adds, and is bit-identical while the CPU's own sum is exact |
| ADR-1448 | ciede_hip runs the CPU's arithmetic in fp32 pairs, from the header the SYCL twin runs; what remains is glibc's powf and the last bits of a pair |
| ADR-1452 | the gate bounds speed_chroma_hip against the CPU by what glibc's log2f adds, as it does for the CUDA twin |
| ADR-1458 | float_adm_hip runs the CPU's arithmetic from the header the CUDA twin runs and returns the CPU's scores bit for bit |
| ADR-1460 | speed_temporal becomes a parity-gate feature with a derived bound, and every registered twin of a gated backend has to be a gate cell |
| ADR-1472 | The integer ADM weight limits follow from the contrast-masking cube and the largest wavelet coefficient of a scale |
| ADR-1477 | SpEED evaluates Netflix's fp64 expressions again, and its GPU twins form the entropies and the score on the host with the same statements, so every twin returns the CPU's scores bit for bit on any C library |
| ADR-1491 | The CUDA, SYCL and HIP motion twins compute motion_five_frame_window: the frame two back on the device, the CPU's window function on the host |
| ADR-1497 | The float_moment twins form the CPU's rounded second-moment sum past 2^53 units, and return the CPU's bits on every frame |
| ADR-1499 | The float_psnr twins add each row's exact sum in the CPU's order, and return the CPU's bits on every frame |
| ADR-1511 | An AMD GPU tester image that ships only the ROCm runtime files the HIP build loads, with the source of its LGPL parts |
| ADR-1517 | The production GPU images are built on Debian 13, ship only the vendor files libvmaf loads, and share the tester images' licence records |
| ADR-1523 | HIP twins run on the device of the imported state |
| ADR-1525 | adm_hip computes AIM on the device and is dispatched |
| ADR-1561 | integer VIF converts its residual variance through vif_sv_sq(), which returns x86's value without the undefined double to int32_t conversion |
| ADR-1571 | a GPU dispatch variable is documented only when library code reads it; VMAF_CUDA_DISPATCH is read at extractor init, VMAF_HIP_DISPATCH and its predicate are removed |
| ADR-1590 | Every build stores its GPU device code compressed at the toolchain's strongest setting, and the build refuses raw device code |
| ADR-1601 | integer VIF forms the denominator log argument sigma_nsq + sigma1_sq in uint32_t |
| ADR-1685 | A post-1.0 embedding milestone: zero-copy device-frame import with fences, asynchronous window scores, and an unchanged licence |
| ADR-1829 | RC4 owns the whole device-memory import API, not only the first full Rust metric |
| ADR-1874 | The vmaf CLI reports its usable backends; the score-backend selectors read that report |
| ADR-1917 | The integer ADM scale-0 horizontal and vertical weight limit is 43900, set by the CSF stage's 16-bit magnitude |
| ADR-1918 | Samples above 2^bpc - 1 are invalid input; an opt-in check refuses them |
| ADR-1929 | VMAFx device frames, fences and the import rule: the shared contract the backend lanes implement |
| ADR-2092 | VMAFx device frames on HIP: one library stream per device copies every frame for the twins, dma-bufs as external memory, sync_file checked on the host |
| ADR-2132 | HIP imports GL textures through EGL dma-buf export; the runtime's GL interop is not used |
| ADR-2134 | adm_cm_aim_line_kernel_4 gets its own register budget of 209 for the exact scale-0 angle flag |
| ADR-2796 | Every clang-tidy ratchet lane is a hosted, path-routed required check |