HIP (AMD ROCm) compute backend¶
The HIP backend runs VMAFx feature extractors on AMD GPUs through ROCm. Build with -Denable_hip=true -Denable_hipcc=true, then run vmaf --backend hip. This page covers building, running, what is implemented and what is still open. Per-twin detail is on the pages listed under More HIP pages.
Status
All 19 registered HIP extractors run on AMD hardware. 18 of them return the CPU extractor's values bit for bit, and ciede_hip is within 1e-9 of the CPU. The state ledger docs/state.md holds every open item, and the dated change record is on the history page. All measurements on these pages come from one integrated GPU (gfx1036); a discrete AMD GPU has not been measured.
Requirements¶
| Item | Value |
|---|---|
| ROCm | 7.0 or later builds; 10.1.0 is the version tested in CI, in the dev container and in the published GPU images (build-config.env ROCM_VERSION, ADR-1225) |
| Libraries | libamdhip64 and <hip/hip_runtime_api.h>, found through the hip-lang package, HIP_PATH, or /opt/rocm |
| Compiler | hipcc in PATH when enable_hipcc=true |
| Hardware | An AMD GPU visible to ROCm; the default fat binary targets gfx90a, gfx1030, gfx1036 and gfx1100 |
ROCm 10 has no apt channel. Since ROCm 7.14 AMD builds and releases through "TheRock", and repo.radeon.com/rocm/apt/ ends at 7.2.4. VMAFx therefore installs ROCm from the digest-pinned rocm/dev-ubuntu-26.04:10.1.0-full container image. On a workstation, use your distribution's ROCm packages (any 7.0+ release builds) or the dev container. CI runs scripts/ci/install-rocm-from-image.sh, which streams the image's /opt/rocm out of the registry without a 29 GB docker pull.
Build¶
-
Configure from the repository root. The Meson source directory is
core/. -
Compile.
-
Run the test suite. Device tests skip with exit 77 when no AMD device is visible.
Build options¶
| Option | Default | Effect |
|---|---|---|
enable_hip | false | Compiles the HIP host runtime. With it off, every public libvmaf_hip.h entry point returns -ENOSYS. |
enable_hipcc | false | Compiles the device kernels with hipcc and embeds the HSACO code objects. Without it, every extractor returns -ENOSYS at init(); float_ssim_hip, integer_ssim_hip and vmaf_hip_picture_alloc log an error that names -Denable_hipcc=true. Requires enable_hip=true. |
hip_gfx_targets | empty (auto-detect) | Comma-separated --offload-arch list for the HSACO fat binary. See GFX targets. |
enable_float_vif_hip_autodispatch | true | Sets VMAF_FEATURE_EXTRACTOR_HIP on float_vif_hip, so --backend hip and a VMAFx context on a HIP device select it for float_vif (ADR-0623; on by default since ADR-2092). Off: the twin runs only when named. |
compress_device_code | true | Stores each kernel's code object bundle compressed (hipcc --offload-compress --offload-compression-level=22, a zstd CCOB bundle): 1.09 MB instead of 17.9 MB for the 25 tester targets. The HIP runtime decompresses a bundle when hipModuleLoadData() loads it (measured with the ROCm 7.2.4 runtime on a gfx1036: no change in load time); the AMD tester image ships the runtime of the ROCm that compiled its kernels. See compress_device_code. |
Pre-compiled HSACO fat binaries are not bundled without hipcc, because ROCm needs target-specific code objects. The CI compile lane (Ubuntu HIP) builds with -Denable_hipcc=false, and no GitHub-hosted runner has an AMD GPU, so CI does not compile or exercise the kernels. HIP runtime types (hipDevice_t, hipStream_t) cross the public ABI as uintptr_t, which keeps libvmaf_hip.h free of <hip/hip_runtime.h>.
GFX targets¶
hipcc --genco produces one HSACO blob per --offload-arch target. Meson resolves the target list in this order:
- The
-Dhip_gfx_targets=<csv>override. rocm_agent_enumerator(thegfx*lines).hipconfig --amdgpu-target.- The fallback list
gfx90a,gfx1030,gfx1036,gfx1100(CDNA2 server, RDNA2 desktop, the Raphael APU iGPU and RDNA3).
Steps 2 and 3 succeed only when the build host can see a GPU. In a no-GPU sandbox (BuildKit, CI) both return nothing and the build uses step 4. The HIP HSACO targets: line of the Meson configure output shows the resolved list.
To shrink the fat binary, pin one target or a comma-separated list:
rocm_agent_enumerator prints the target of the installed GPU.
Why the fallback list is wide
The fallback was gfx90a only until ADR-0561. That narrow list shipped libvmaf.so binaries that failed at runtime on the fork's own dev host (a Raphael APU, gfx1036) with hip_fatbin.cpp: No compatible code objects found for: gfx1030. Under ROCm 6.x and 7.x that host needed HSA_OVERRIDE_GFX_VERSION=10.3.0 to alias gfx1036 onto the allowlisted gfx1030. ROCm 10 supports gfx1036 natively, so the override is gone (ADR-1225).
Run¶
--backend hip selects the HIP twin of every feature the model or the --feature list needs, and runs the rest on the CPU. --backend hip pins device 0; --hip_device N picks another device by ordinal.
vmaf --reference ref.yuv --distorted dist.yuv \
--width 576 --height 324 --pixel_format 420 --bitdepth 8 \
--backend hip --json --output out.json
The JSON output names the extractor that ran for each feature under feature_backends, and the backend that ran under backend_used.
| Flag | Effect |
|---|---|
--backend hip | Exclusive HIP selection; disables the other backends before dispatch. An explicit request for a backend that was not compiled in fails with exit code 100. |
--hip_device N | Selects the HIP GPU by ordinal; every HIP twin runs on it. An ordinal the runtime does not have fails with exit code 100 under --backend hip. |
--no_hip | Forbids HIP dispatch even when the backend is built in. |
--feature NAME | Runs an extractor. A CPU name runs the HIP twin under --backend hip; the twin's own name (psnr_hip) always runs it. |
CLI reference | backend selection
A twin that carries no HIP flag runs only when you name it. In a default build every twin carries it; a build with -Denable_float_vif_hip_autodispatch=false leaves float_vif_hip without it, so it needs the _hip name there:
vmaf --reference ref.yuv --distorted dist.yuv \
--width 576 --height 324 --pixel_format 420 --bitdepth 8 \
--backend hip --feature float_vif_hip --no_prediction --json --output out.json
A twin named this way runs on the thread that calls vmaf_read_pictures(), with or without --threads.
FFmpeg selects the device with the hip_device=N filter option (patch 0011-libvmaf-wire-hip-backend-selector.patch in ffmpeg-patches/, ADR-0380).
A program that already holds its frames on the GPU (a decoder's dma-bufs, a renderer's GL textures, HIP device memory) hands them to the VMAFx API instead of vmaf_read_pictures(): see HIP devices. On a host with a HIP device, the import is checked against host uploads by the hip suite:
python3 scripts/ci/run_meson_test.py -- -C build --suite hip \
test_vmafx_import_hip test_vmafx_import_hip_bitexact \
test_vmafx_import_hip_fence test_vmafx_import_hip_gl
The bit-exactness run prints the cells, values and imports it compared and the host copies it counted (0), and repeats a cell that differs, printing each repeat (see the gfx1036 loses stream commands).
The HIP backend reads no dispatch environment variable: every HIP twin submits directly, and --backend hip picks a twin by its registration flag. VMAF_HIP_DISPATCH, which earlier pages listed, was read by a function nothing called and has been removed (ADR-1571).
Registered extractors¶
Nineteen extractors are registered in core/src/feature/feature_extractor.cpp. "Named only" means the extractor carries no VMAF_FEATURE_EXTRACTOR_HIP flag, so --backend hip does not pick it for the CPU name.
| Registered name | CPU feature | --backend hip picks it | Emits | Added in |
|---|---|---|---|---|
psnr_hip | psnr | yes | psnr_y, psnr_cb, psnr_cr | ADR-0241 |
float_psnr_hip | float_psnr | yes | float_psnr | ADR-0254 |
ciede_hip | ciede | yes | ciede2000 | ADR-0259, PR #1016, ADR-1448 |
float_moment_hip | float_moment | yes | four float_moment_* | ADR-0260 |
motion_v2_hip | motion_v2 | yes | motion_v2 SAD, motion2_v2, motion3_v2 | ADR-0267 |
motion_hip | motion | yes | motion, motion2, motion3 | ADR-0523, PR #1004 |
float_motion_hip | float_motion | yes | motion, motion2, motion3 | ADR-0373 |
float_ssim_hip | float_ssim | yes | float_ssim (and L, C, S with enable_lcs) | ADR-0375 |
float_vif_hip | float_vif | yes (named only with -Denable_float_vif_hip_autodispatch=false) | float_vif_scale0..3 | ADR-0379, ADR-0592 |
float_adm_hip | float_adm | yes | adm2, adm_scale0..3, aim, adm3 | ADR-0468, PR #1024, ADR-1458 |
psnr_hvs_hip | psnr_hvs | yes | psnr_hvs and per-channel values | PR #995 |
cambi_hip | cambi | yes | cambi | PR #996, ADR-1378 |
ssimulacra2_hip | ssimulacra2 | yes | ssimulacra2 | PR #1000 |
vif_hip | vif | yes | vif_scale0..3 | PR #1001 |
adm_hip | adm | yes | integer_adm2, integer_aim, integer_adm3, integer_adm_scale0..3 | PR #1007, ADR-1423, ADR-1525 |
integer_ms_ssim_hip | float_ms_ssim | yes | float_ms_ssim (and _cb, _cr, L, C, S) | PR #1013 |
integer_ssim_hip | ssim | yes | ssim | PR #999, ADR-0564 |
speed_chroma_hip | speed_chroma | yes | SpEED chroma features | ADR-0567, ADR-0852, ADR-1384 |
speed_temporal_hip | speed_temporal | yes | SpEED temporal features | ADR-0567, ADR-0852, ADR-1384 |
The kernel file, algorithm and per-twin notes for each row are in HIP twins.
Agreement with the CPU¶
18 twins return the CPU extractor's values bit for bit. A twin is declared exact by one file scripts/ci/exact_twins.d/<feature>.hip, and the cross-backend parity gate then compares it with tolerance 0 (generated list of exact twins). The one bounded twin takes its bound from LIBM_TWINS in scripts/ci/cross_backend_calibration.py.
The table is the state measured on a gfx1036 (ROCm 7.2.4, glibc 2.44) at --precision max against --backend cpu: 110 frames of typical content (the Netflix 576x324 pair at 8 and 10 bits, both 1920x1080 checkerboard pairs, Sparks 480x270 at 10 bits, 48 frames of BBB 3840x2160) and 68 frames that stress the arithmetic (12 and 16 bits, 10-bit 4:2:2, full-range noise at four depths, a bright 16-bit 1080p pair).
"Values identical" counts every output of every frame; the speed_* rows come from a wider set of clips (Research-1437 has the sweep per output and per fixture). "Largest difference before" is the difference measured when the sweep first ran (2026-10-01, when eight of these twins were exact), before the twin took the CPU's arithmetic.
| CPU feature | HIP twin | Values identical | Largest difference | Before | Exact twin |
|---|---|---|---|---|---|
motion (also debug=true) | motion_hip | 534 of 534 (712 of 712) | 0 | 0 | yes (ADR-1437) |
motion_v2 | motion_v2_hip | 534 of 534 | 0 | 0 | yes (ADR-1437) |
psnr | psnr_hip | 534 of 534 | 0 | 0 | yes (ADR-1437) |
float_ms_ssim (also enable_lcs) | integer_ms_ssim_hip | 178 of 178 (2848 of 2848) | 0 | 0 | yes (ADR-1437) |
cambi | cambi_hip | 178 of 178 | 0 | 0 | yes (ADR-1437) |
adm (also debug=true) | adm_hip | 890 of 890 (4141 of 4141 with aim and adm3) | 0 | 0 | yes (ADR-1423, ADR-1525) |
float_motion | float_motion_hip | 534 of 534 | 0 | 0 | yes (ADR-1419) |
psnr_hvs | psnr_hvs_hip | 680 of 680 (8 to 12 bits) | 0 | 0 | yes (ADR-1401) |
vif | vif_hip | 712 of 712 | 0 | 5.4e-7 | yes (ADR-1435) |
ssim | integer_ssim_hip | 178 of 178 | 0 | 1.1e-11 | yes (ADR-1438) |
float_psnr | float_psnr_hip | 178 of 178 | 0 | 7.6e-8 dB | yes (ADR-1440, ADR-1499) |
float_ssim (also enable_lcs) | float_ssim_hip | 178 of 178 (712 of 712) | 0 | 5.4e-7 | yes (ADR-1441) |
float_vif | float_vif_hip | 712 of 712 | 0 | 1.1e-4 | yes (ADR-1444) |
ssimulacra2 | ssimulacra2_hip | 178 of 178 | 0 | 7.6e-11 | yes (ADR-1445) |
float_moment | float_moment_hip | 712 of 712 | 0 | 1.0e-4 | yes (ADR-1447, ADR-1497) |
float_adm (also debug=true) | float_adm_hip | 1246 of 1246 (3204 of 3204) | 0 | 1.3e-5 | yes (ADR-1458) |
speed_chroma | speed_chroma_hip | 759 of 759 | 0 | 1.4e-6 | yes (ADR-1477) |
speed_temporal | speed_temporal_hip | 256 of 256 | 0 | 4.8e-7 | yes (ADR-1477) |
ciede | ciede_hip | 115 of 178 | 1.4e-11 | 1.1e-5 | no: the C library's powf and the last bits of an fp32 pair, bounded at 1e-9 (ADR-1448) |
Exact twins cost time. Cost of exactness lists the frame time before and after each twin took the CPU's arithmetic. To re-run the comparison:
python3 scripts/ci/run_meson_test.py -- -C build-hip test_hip_exact_twins
python3 scripts/ci/cross_backend_parity_gate.py --vmaf-binary build-hip/tools/vmaf \
--reference python/test/resource/yuv/src01_hrc00_576x324.yuv \
--distorted python/test/resource/yuv/src01_hrc01_576x324.yuv \
--width 576 --height 324 --backends cpu hip \
--features vif motion motion_debug motion_v2 adm psnr float_moment ssim \
float_ssim float_ssim_lcs float_ms_ssim float_ms_ssim_lcs float_psnr \
float_motion float_vif float_adm psnr_hvs ssimulacra2 cambi ciede \
speed_chroma
Known gaps¶
Open items only; the ledger row ids are in docs/state.md.
adm_hipis slower than the CPU on an integrated GPU. On the gfx1036 of the measuring host (2 compute units) the twin takes about 220 ms per 3840x2160 frame, the CPUadmextractor 15 ms with 16 threads; the AIM pass recomputes csf(r) at all nine taps of every threshold, as the CUDA twin does, and is two thirds of that. The default model under--backend hipgoes from 65 to 273 ms per 3840x2160 frame because its ADM now runs on that device (T-HIP-ADM-AIM-INLINE-COST-2026-10-04, RC8).- Imported frames are copied once more on the device. The VMAFx API imports device pointers, dma-bufs, arrays and GL textures without a host copy, but the twins copy each imported frame into their own buffers on the device's library stream, where the CUDA twins read the picture itself (ADR-2092; RC8 tuning row
T-HIP-IMPORT-TWIN-DEVICE-COPY-2026-10-06). See picture uploads. - Import limits of the runtime. A sync_file is an acquire fence only, checked on the host: ROCm 10.1 imports no external semaphore that could carry one (ROCm 7.2.4 aborts the process trying), so no HIP stream waits on or signals one (
T-HIP-ROCM-NO-SYNC-FILE-SEMAPHORE-2026-10-06). OpenGL textures import only from a GLX context on the device's GPU, and not at all with ROCm 10.1, whose runtime maps a texture but cannot read it: the import is refused naming the runtime (T-HIP-ROCM10-GL-TEXTURE-READ-2026-10-06; ROCm 7.2.4 reads them). A HIP device has no frame pools. - Six twins stage their own copy of the frame instead of reading the shared planes:
integer_ms_ssim_hip,psnr_hvs_hip,cambi_hip,speed_chroma_hip,speed_temporal_hipandssimulacra2_hip(T-HIP-SHARED-FRAME-REMAINING-TWINS-2026-10-01, RC8). - Exact twins are slower than before, and some are slower than the CPU on the gfx1036. Open tuning rows, all RC8 and all with correct scores:
T-HIP-FLOAT-SSIM-EXACT-THROUGHPUT-2026-10-01,T-HIP-FLOAT-VIF-EXACT-THROUGHPUT-2026-10-02,T-HIP-SSIMULACRA2-EXACT-THROUGHPUT-2026-10-02,T-HIP-CIEDE-EXACT-THROUGHPUT-2026-10-02,T-HIP-FLOAT-MS-SSIM-EXACT-THROUGHPUT-2026-10-02,T-SYCL-HIP-PSNR-HVS-EXACT-SUM-THROUGHPUT-2026-10-01andT-GPU-FLOAT-MOMENT-EXACT-SUM-COST-2026-10-03. - No CI device. CI compiles the host code without
hipcc; the kernels are exercised only on a developer's AMD GPU.
The gfx1036 loses stream commands¶
This is a platform defect, deferred as T-HIP-GFX1036-DROPPED-DISPATCHES-2026-10-01. On a gfx1036 (ROCm 7.2.4, Linux 7.2.8) a HIP stream now and then never runs a run of the commands it was given, roughly once per \(10^{4}\) frames, on master as well. A HIP twin then reports a wrong score for that frame: vif_hip reports the sums of two frames when the memset of its accumulators is lost, or fails the run with invalid ratio when a scale's kernel is lost.
Nothing in VMAFx sets it off, and no runtime setting tried stops it.
To check a device or a driver update, run the probe that reproduces it without VMAFx:
hipcc -O2 --offload-arch=gfx1036 scripts/dev/hip_dispatch_drop_probe.hip -o /tmp/probe
for i in 1 2 3 4 5; do /tmp/probe 100000 12 0; done
Each run prints bad_frames and lost (dispatches that never ran); both are 0 on a healthy stack. On the gfx1036 five runs gave 55 bad frames in 500000 and 382 lost dispatches in 6.0 million; on Linux 7.2.9 three runs of 20000 frames gave 0, 17 and 3 bad frames (2026-10-06, load average 24 to 38). Until a driver update clears it, compare HIP scores from this device over repeated runs and treat a single-frame mismatch as suspect, not as a code defect.
A separate, deferred report (T-HIP-GFX1036-SDMA-READ-FAULT-2026-10-01) is one psnr_hvs_hip run killed by a GPU memory access fault raised by the copy engine; 131 further runs of the same command were clean.
More HIP pages¶
| Page | Content |
|---|---|
| HIP twins | Kernel notes, floating-point policy, the cost of exactness and one section per twin with its measurements |
| Picture uploads and device state | Shared frame planes, zero-copy status, accumulator clearing |
| History | Dated status entries and the ADR-0537, ADR-0539 and ADR-1103 bring-up |
| Backend selection | How the backends compete and how --backend resolves |
References¶
- ADR-0212 — the original scaffold.
- ADR-0241 — first consumer (
psnr_hip). - ADR-0254 — second consumer (
float_psnr_hip). - ADR-0259 — third consumer.
- ADR-0260 — fourth consumer (
float_moment_hip). - ADR-0266 — fifth consumer (
float_ansnr_hip), retained for historical traceability. The kernel and its CPU twin were removed in ADR-0709 (PR #38); ANSNR is no longer a registered feature on any backend. - ADR-0267 — sixth consumer (
motion_v2_hip). - ADR-0372 — batch-1 kernels.
- ADR-0373 — batch-2 kernels.
- ADR-0375 — batch-3 kernels.
- ADR-0377 — batch-4 kernels.
docs/adr/0379-hip-float-vif.md— unavailable historical reference forfloat_vif_hip; ADR-0592 records the later removal of its weak stub after the real kernel shipped.- ADR-0380 — FFmpeg selector.
- ADR-0468 —
float_adm_hip. - ADR-0523 — register
vmaf_fex_integer_motion_hip. - ADR-0533 — full HIP-extractor registration sweep (six more TUs wired into
hip_sourcesandfeature_extractor_list[]). - Research-0432 — AMD market-share and ROCm Linux maturity survey.
Former section names¶
ADRs and research digests link to these headings; each points to the section that now holds its content.
A frame clears its accumulators after its upload (ADR-1427)¶
Now under clearing accumulators after the upload.
vif_hip returns the CPU's scores bit for bit (2026-10-01)¶
Now under vif_hip in the twin notes.