ADR-1590: Every build stores its GPU device code compressed at the toolchain's strongest setting, and the build refuses raw device code¶
- Status: Accepted
- Date: 2026-10-04
- Deciders: maintainer, agent
- Tags: build, meson, cuda, hip, sycl, gpu, packaging, fork-local
Context¶
The GPU kernels of all three backends are embedded in libvmaf and, because the test executables link libvmaf.a, in every test program as well. Before this decision each toolchain stored them its own way:
- nvcc compresses only the PTX entries of a fatbin by default (
fatbinary --help: "Compress ptx and debug images", defaulttrue); the six cubins of each kernel went in raw. The 21 fatbins took 10.97 MB, the same 10.97 MB in each of the 267 test executables of a CUDA build, and the Windows CUDA tester zip carried that payload in each of its test programs. - hipcc stores a plain offload bundle: 17.9 MB for the 20 kernels and the 25 targets of the tester kits, again in every one of 282 test executables.
- icpx compressed the ahead-of-time images (
--offload-compress, ADR-1360) at its default zstd level 10 on the compile line, but the SPIR-V fallback image is generated by the final link, which carried no compression flag, so every binary held it raw (1.9 MB inlibvmaf.so).
The maintainer asked for every build and every published artifact to be compressed as well as practical, as the project standard, without changing what is computed. This record covers the device code; the archives and images that carry it are decided separately (ADR-1591, its own pull request).
Decision¶
We will compress the device code of every backend in every build at the strongest setting its toolchain offers, through one list per backend in core/src/meson.build (BEGIN/END VMAF {CUDA,HIP,SYCL} device code compression policy), switched by the Meson option compress_device_code (default true):
| Backend | Flags | Where |
|---|---|---|
| CUDA (nvcc 13.4) | -Xfatbin=-compress-all --compress-mode=size | every fatbin compile |
| HIP (ROCm clang 22 / ROCm 7.2.4 here, ROCm 10 in the container) | --offload-compress --offload-compression-level=22 | every hipcc --genco, test probes included |
| SYCL (icpx 2026.0 here, 2026.1 pinned) | --offload-compress --offload-compression-level=22 | the AOT compile line, the final link (SPIR-V image), the MSVC device link |
Level 22 is zstd's maximum. A compiler that cannot compress (the clang CUDA driver of enable_nvcc=false, AdaptiveCpp, an nvcc without --compress-mode, a clang without --offload-compression-level) stops configure with an error that names -Dcompress_device_code=false; nothing is skipped silently. With the option off, nvcc stores every entry raw (--no-compress) and hipcc and icpx store plain images. The build then checks its own output: core/src/check_device_compression.py fails the build when a fatbin holds a raw cubin or PTX entry, a code object bundle is not a compressed (CCOB) bundle, or a SYCL image section of libvmaf.so is not zstd frames.
nvcc's --concat (new in nvcc 13.4) is not used: NVIDIA documents no driver floor for it, while a CUDA 13 build otherwise promises every R580 driver.
Measurements¶
Host ryzen-4090-arc (RTX 4090 driver 615.71.09, gfx1036 with the ROCm 7.2.4 runtime, Arc A380 on the xe driver), build options of the tester kits (--buildtype=release --strip -Db_lto=false -Denable_float=true -Denable_tests=true), HIP for the 25 tester targets, SYCL for the 19 default AOT targets. "Before" is origin/master 2889f963a, "after" this change.
Device code (the CUDA and HIP variants are the same commands with only the compression flags changed):
| CUDA, 21 fatbins (6 cubins + 2 PTX each) | Bytes |
|---|---|
raw (--no-compress) | 14,606,544 |
| master (nvcc default: PTX only) | 10,965,504 |
--compress-mode=speed, all entries | 4,349,256 |
--compress-mode=balance, all entries | 3,200,528 |
--compress-mode=size, all entries | 2,599,576 |
size + --concat (not adopted) | 1,980,136 |
| HIP, 20 bundles x 25 targets | Bytes |
|---|---|
| plain bundle (master) | 17,866,480 |
| zstd level 3 (clang's default) | 1,290,717 |
| zstd level 10 | 1,254,212 |
| zstd level 19 | 1,088,567 |
| zstd level 22 | 1,088,092 |
SYCL images in libvmaf.so | master | after |
|---|---|---|
spir64_gen (19 targets, 45 images) | 6,421,923 (level 10) | 5,229,295 (level 22) |
spir64 (SPIR-V fallback) | 1,897,800 (raw) | 444,348 (level 22) |
Binaries:
| Build | libvmaf.so before | after | test executables before | after |
|---|---|---|---|---|
| CUDA | 14,225,896 | 5,857,768 | 267: 2.53 GB | 0.92 GB |
| HIP | 21,007,408 | 4,229,136 | 282: 4.09 GB | 0.65 GB |
| SYCL | 12,807,096 | 10,161,080 | 265: 2.39 GB | 1.85 GB |
Tester kit payloads (the test programs of tools/rc1-tester/image/unit-tests.txt and the backend's list, tools/vmaf and libvmaf.so, staged with prepare_build.py stage from the host builds; the images compress layers with gzip, the Windows zip with deflate level 9):
| Payload | On disk before | after | gzip -6 tar before | after | zip level 9 before | after |
|---|---|---|---|---|---|---|
| NVIDIA image list (103 tests) | 1.198 GB | 0.436 GB | 396 MB | 323 MB | 397 MB | 330 MB |
| AMD image list (110 tests) | 2.031 GB | 0.319 GB | 574 MB | 204 MB | 584 MB | 208 MB |
| Intel image list (104 tests) | 1.057 GB | 0.815 GB | 752 MB | 627 MB | 767 MB | 640 MB |
| Windows CUDA zip list, Linux build (117 tests) | 1.278 GB | 0.466 GB | 422 MB | 345 MB | 423 MB | 352 MB |
Load time, unchanged within noise:
| Measure | Before | After |
|---|---|---|
CUDA: cuModuleLoadData() of all 21 fatbins, fresh process, median of 9 | 1.39 ms | 2.99 ms (--concat: 4.50 ms) |
HIP: hipModuleLoadData() of all 20 bundles, median of 5 | 10.9 ms (plain) | 10.2 ms |
SYCL: zstd decode of every image of libvmaf.so (the runtime decodes only the images it uses) | 38 ms (level 10) | 37 ms (level 22) |
vmaf one frame, 15 GPU twins plus the default model, CUDA, median of 11 | 207 ms | 201 ms |
| the same on HIP | 100 ms | 104 ms |
| the same on SYCL | 205 ms | 200 ms |
Build cost: compiling the 21 fatbins took 11.0 s at nvcc's default and 10.8 s in size mode, the 20 HIP bundles 51 s at level 3 and 55 s at level 22 (six jobs on a loaded host). zstd level 22 instead of 10 costs 11.3 s of CPU for all 45 SYCL images (spread over 45 translation units) and 0.36 s per link for the SPIR-V image.
Nothing computed changes¶
- CUDA: the 168 cubins and PTX files
cuobjdump -xelf all -xptx allextracts from the compressed fatbins are byte-identical to those of a-Dcompress_device_code=falsebuild of the same tree. - HIP: the 500 code objects
clang-offload-bundler --unbundleextracts are identical in.text,.rodata,.noteand.dataand in their symbol tables, except for the name of the__hip_cuid_<hash>symbol, which hashes the compile options. - Scores:
vmaf --precision maxwith 15 GPU twins and the default model, on the Netflix 576x324 pair and both 1080p checkerboard pairs, gives identical values before and after on every backend: 2,862 of 2,862 metric cells each on the RTX 4090, the gfx1036 and the A380. - Device suites on the compressed builds: CUDA 66 of 66
gputests andtest_cuda_parity_gate_default_run; HIP 73 of 73; SYCL 67 of 67 with AOT and with-Dsycl_icpx_aot_targets=(SPIR-V only), among themtest_cuda_exact_twins,test_hip_exact_twins,test_hip_adm_exact,test_sycl_exact_twinsandtest_sycl_kernel_scratch.
Driver and runtime minimums¶
- CUDA:
--compress-modeoutput loads on drivers from CUDA 12.4 (R550) on (nvcc 13.4 manual,--compress-mode). A CUDA 13 build already needs R580 (CUDA 13.4 release notes, minor version compatibility table), so the floor does not move. - HIP: the runtime decompresses a
CCOBbundle inhipModuleLoadData(); loaded here with the ROCm 7.2.4 runtime. The AMD tester image ships the runtime of the ROCm that compiled its kernels. - SYCL: the oneAPI runtime decodes the zstd images when it first builds a kernel from them; measured on the A380 for both the AOT and the SPIR-V image.
Alternatives considered¶
| Option | Pros | Cons | Why not chosen |
|---|---|---|---|
| Strongest setting everywhere, option to turn off, configure error when unsupported (chosen) | Smallest device code with documented load floors; one place per backend; never silently uncompressed | Builds with clang CUDA or AdaptiveCpp must pass -Dcompress_device_code=false | |
Add nvcc --concat | Fatbins 24% smaller again (1.98 MB) | New in nvcc 13.4 with no documented driver floor; a tester on an R580 to R610 driver might not load the kernels; module load 4.5 ms instead of 3.0 ms | Revisit when NVIDIA documents its floor |
| Keep each toolchain's default level (nvcc speed mode, zstd 3 or 10) | Slightly faster compiles | 1.7x (CUDA) and 1.2x (HIP, SYCL) larger device code; the maintainer asked for the strongest compression | Contradicts the decision |
| Compress where the compiler can, skip elsewhere | No configure error for clang CUDA or AdaptiveCpp | A build silently ships raw device code (the project's no-silent-fallback rule) | Rejected by the request |
| Compress the binaries after linking (upx and similar) | Backend-independent | Rewrites vendor-linked binaries (the tester licence rules forbid modifying vendor files), defeats page sharing, breaks code signing | Not a device code measure |
Consequences¶
- Positive: device code is 4.2x (CUDA), 16.4x (HIP) and 1.6x (SYCL
libvmaf.soimages) smaller; the HIP test executables drop from 4.09 GB to 0.65 GB per build, and every tester kit and image is smaller on disk and to download. - Negative: builds with the clang CUDA driver or AdaptiveCpp need an extra option; level 22 costs a few seconds of CPU per build.
- Neutral / follow-ups:
core/test/test_device_code_compression.pyholds the policy (each list's flags, every compile site, no flag spelled outside the blocks, the build checks, the option default) and the checker's fixtures.scripts/ci/gen-sycl-compile-commands.pystrips the level flag for clang-tidy. A new device compile site takes its backend's list.
References¶
- req (maintainer, 2026-10-04, as relayed in the coordinator brief): compress all of our builds as well as we can; that should be the standard.
- nvcc 13.4 manual,
--compress-mode,--no-compress,--concat: https://docs.nvidia.com/cuda/cuda-compiler-driver-nvcc/index.html;fatbinary --help(CUDA 13.4.92) for-compress-alland the PTX-only default; nvcc 12.6 and 12.8 manuals (the option first appears in 12.8). - CUDA 13.4 release notes, minor version compatibility (13.x needs >= 580): https://docs.nvidia.com/cuda/cuda-toolkit-release-notes/index.html.
- RAPIDS
rapids_cuda_enable_fatbin_compression(driver 550.54.14+ for size mode under CUDA 12, 580+ under CUDA 13;-Xfatbin=-compress-all): https://github.com/rapidsai/rapids-cmake/blob/main/rapids-cmake/cuda/enable_fatbin_compression.cmake. - LLVM
OffloadBundler.cpp(zstd default level 3 with long-distance matching) andOptions.td(--offload-compression-level=from LLVM 19); ROCm "Code portability and compression": https://rocm.docs.amd.com/projects/llvm-project/en/docs-7.2.4/conceptual/code-portability.html. - intel/llvm SYCL users manual,
--offload-compressand--offload-compression-level(default 10): https://github.com/intel/llvm/blob/sycl/sycl/doc/UsersManual.md. - AdaptiveCpp 25.10
acpp --help(no compression option). - ADR-1360, ADR-1364, ADR-1403, ADR-1407, ADR-1503; research digest Research-1590.