Skip to content

ADR-1590: Every build stores its GPU device code compressed at the toolchain's strongest setting, and the build refuses raw device code

  • Status: Accepted
  • Date: 2026-10-04
  • Deciders: maintainer, agent
  • Tags: build, meson, cuda, hip, sycl, gpu, packaging, fork-local

Context

The GPU kernels of all three backends are embedded in libvmaf and, because the test executables link libvmaf.a, in every test program as well. Before this decision each toolchain stored them its own way:

  • nvcc compresses only the PTX entries of a fatbin by default (fatbinary --help: "Compress ptx and debug images", default true); the six cubins of each kernel went in raw. The 21 fatbins took 10.97 MB, the same 10.97 MB in each of the 267 test executables of a CUDA build, and the Windows CUDA tester zip carried that payload in each of its test programs.
  • hipcc stores a plain offload bundle: 17.9 MB for the 20 kernels and the 25 targets of the tester kits, again in every one of 282 test executables.
  • icpx compressed the ahead-of-time images (--offload-compress, ADR-1360) at its default zstd level 10 on the compile line, but the SPIR-V fallback image is generated by the final link, which carried no compression flag, so every binary held it raw (1.9 MB in libvmaf.so).

The maintainer asked for every build and every published artifact to be compressed as well as practical, as the project standard, without changing what is computed. This record covers the device code; the archives and images that carry it are decided separately (ADR-1591, its own pull request).

Decision

We will compress the device code of every backend in every build at the strongest setting its toolchain offers, through one list per backend in core/src/meson.build (BEGIN/END VMAF {CUDA,HIP,SYCL} device code compression policy), switched by the Meson option compress_device_code (default true):

Backend Flags Where
CUDA (nvcc 13.4) -Xfatbin=-compress-all --compress-mode=size every fatbin compile
HIP (ROCm clang 22 / ROCm 7.2.4 here, ROCm 10 in the container) --offload-compress --offload-compression-level=22 every hipcc --genco, test probes included
SYCL (icpx 2026.0 here, 2026.1 pinned) --offload-compress --offload-compression-level=22 the AOT compile line, the final link (SPIR-V image), the MSVC device link

Level 22 is zstd's maximum. A compiler that cannot compress (the clang CUDA driver of enable_nvcc=false, AdaptiveCpp, an nvcc without --compress-mode, a clang without --offload-compression-level) stops configure with an error that names -Dcompress_device_code=false; nothing is skipped silently. With the option off, nvcc stores every entry raw (--no-compress) and hipcc and icpx store plain images. The build then checks its own output: core/src/check_device_compression.py fails the build when a fatbin holds a raw cubin or PTX entry, a code object bundle is not a compressed (CCOB) bundle, or a SYCL image section of libvmaf.so is not zstd frames.

nvcc's --concat (new in nvcc 13.4) is not used: NVIDIA documents no driver floor for it, while a CUDA 13 build otherwise promises every R580 driver.

Measurements

Host ryzen-4090-arc (RTX 4090 driver 615.71.09, gfx1036 with the ROCm 7.2.4 runtime, Arc A380 on the xe driver), build options of the tester kits (--buildtype=release --strip -Db_lto=false -Denable_float=true -Denable_tests=true), HIP for the 25 tester targets, SYCL for the 19 default AOT targets. "Before" is origin/master 2889f963a, "after" this change.

Device code (the CUDA and HIP variants are the same commands with only the compression flags changed):

CUDA, 21 fatbins (6 cubins + 2 PTX each) Bytes
raw (--no-compress) 14,606,544
master (nvcc default: PTX only) 10,965,504
--compress-mode=speed, all entries 4,349,256
--compress-mode=balance, all entries 3,200,528
--compress-mode=size, all entries 2,599,576
size + --concat (not adopted) 1,980,136
HIP, 20 bundles x 25 targets Bytes
plain bundle (master) 17,866,480
zstd level 3 (clang's default) 1,290,717
zstd level 10 1,254,212
zstd level 19 1,088,567
zstd level 22 1,088,092
SYCL images in libvmaf.so master after
spir64_gen (19 targets, 45 images) 6,421,923 (level 10) 5,229,295 (level 22)
spir64 (SPIR-V fallback) 1,897,800 (raw) 444,348 (level 22)

Binaries:

Build libvmaf.so before after test executables before after
CUDA 14,225,896 5,857,768 267: 2.53 GB 0.92 GB
HIP 21,007,408 4,229,136 282: 4.09 GB 0.65 GB
SYCL 12,807,096 10,161,080 265: 2.39 GB 1.85 GB

Tester kit payloads (the test programs of tools/rc1-tester/image/unit-tests.txt and the backend's list, tools/vmaf and libvmaf.so, staged with prepare_build.py stage from the host builds; the images compress layers with gzip, the Windows zip with deflate level 9):

Payload On disk before after gzip -6 tar before after zip level 9 before after
NVIDIA image list (103 tests) 1.198 GB 0.436 GB 396 MB 323 MB 397 MB 330 MB
AMD image list (110 tests) 2.031 GB 0.319 GB 574 MB 204 MB 584 MB 208 MB
Intel image list (104 tests) 1.057 GB 0.815 GB 752 MB 627 MB 767 MB 640 MB
Windows CUDA zip list, Linux build (117 tests) 1.278 GB 0.466 GB 422 MB 345 MB 423 MB 352 MB

Load time, unchanged within noise:

Measure Before After
CUDA: cuModuleLoadData() of all 21 fatbins, fresh process, median of 9 1.39 ms 2.99 ms (--concat: 4.50 ms)
HIP: hipModuleLoadData() of all 20 bundles, median of 5 10.9 ms (plain) 10.2 ms
SYCL: zstd decode of every image of libvmaf.so (the runtime decodes only the images it uses) 38 ms (level 10) 37 ms (level 22)
vmaf one frame, 15 GPU twins plus the default model, CUDA, median of 11 207 ms 201 ms
the same on HIP 100 ms 104 ms
the same on SYCL 205 ms 200 ms

Build cost: compiling the 21 fatbins took 11.0 s at nvcc's default and 10.8 s in size mode, the 20 HIP bundles 51 s at level 3 and 55 s at level 22 (six jobs on a loaded host). zstd level 22 instead of 10 costs 11.3 s of CPU for all 45 SYCL images (spread over 45 translation units) and 0.36 s per link for the SPIR-V image.

Nothing computed changes

  • CUDA: the 168 cubins and PTX files cuobjdump -xelf all -xptx all extracts from the compressed fatbins are byte-identical to those of a -Dcompress_device_code=false build of the same tree.
  • HIP: the 500 code objects clang-offload-bundler --unbundle extracts are identical in .text, .rodata, .note and .data and in their symbol tables, except for the name of the __hip_cuid_<hash> symbol, which hashes the compile options.
  • Scores: vmaf --precision max with 15 GPU twins and the default model, on the Netflix 576x324 pair and both 1080p checkerboard pairs, gives identical values before and after on every backend: 2,862 of 2,862 metric cells each on the RTX 4090, the gfx1036 and the A380.
  • Device suites on the compressed builds: CUDA 66 of 66 gpu tests and test_cuda_parity_gate_default_run; HIP 73 of 73; SYCL 67 of 67 with AOT and with -Dsycl_icpx_aot_targets= (SPIR-V only), among them test_cuda_exact_twins, test_hip_exact_twins, test_hip_adm_exact, test_sycl_exact_twins and test_sycl_kernel_scratch.

Driver and runtime minimums

  • CUDA: --compress-mode output loads on drivers from CUDA 12.4 (R550) on (nvcc 13.4 manual, --compress-mode). A CUDA 13 build already needs R580 (CUDA 13.4 release notes, minor version compatibility table), so the floor does not move.
  • HIP: the runtime decompresses a CCOB bundle in hipModuleLoadData(); loaded here with the ROCm 7.2.4 runtime. The AMD tester image ships the runtime of the ROCm that compiled its kernels.
  • SYCL: the oneAPI runtime decodes the zstd images when it first builds a kernel from them; measured on the A380 for both the AOT and the SPIR-V image.

Alternatives considered

Option Pros Cons Why not chosen
Strongest setting everywhere, option to turn off, configure error when unsupported (chosen) Smallest device code with documented load floors; one place per backend; never silently uncompressed Builds with clang CUDA or AdaptiveCpp must pass -Dcompress_device_code=false
Add nvcc --concat Fatbins 24% smaller again (1.98 MB) New in nvcc 13.4 with no documented driver floor; a tester on an R580 to R610 driver might not load the kernels; module load 4.5 ms instead of 3.0 ms Revisit when NVIDIA documents its floor
Keep each toolchain's default level (nvcc speed mode, zstd 3 or 10) Slightly faster compiles 1.7x (CUDA) and 1.2x (HIP, SYCL) larger device code; the maintainer asked for the strongest compression Contradicts the decision
Compress where the compiler can, skip elsewhere No configure error for clang CUDA or AdaptiveCpp A build silently ships raw device code (the project's no-silent-fallback rule) Rejected by the request
Compress the binaries after linking (upx and similar) Backend-independent Rewrites vendor-linked binaries (the tester licence rules forbid modifying vendor files), defeats page sharing, breaks code signing Not a device code measure

Consequences

  • Positive: device code is 4.2x (CUDA), 16.4x (HIP) and 1.6x (SYCL libvmaf.so images) smaller; the HIP test executables drop from 4.09 GB to 0.65 GB per build, and every tester kit and image is smaller on disk and to download.
  • Negative: builds with the clang CUDA driver or AdaptiveCpp need an extra option; level 22 costs a few seconds of CPU per build.
  • Neutral / follow-ups: core/test/test_device_code_compression.py holds the policy (each list's flags, every compile site, no flag spelled outside the blocks, the build checks, the option default) and the checker's fixtures. scripts/ci/gen-sycl-compile-commands.py strips the level flag for clang-tidy. A new device compile site takes its backend's list.

References