Skip to content

Research-1590: GPU device code compression

  • Status: Active
  • Workstream: ADR-1590
  • Last updated: 2026-10-04

Question

Which compression does each device compiler offer for the GPU code it embeds, which setting is the strongest, which drivers and runtimes load the result, and does any of it change the code the GPU runs?

Sources

Findings

Measurements are in ADR-1590; the reproduction commands are below.

  1. nvcc stored the cubins raw. Of the 14.6 MB of the 21 uncompressed fatbins, nvcc's default compressed only the two PTX entries per kernel (10.97 MB). --compress-mode=size compresses every entry; with it -Xfatbin=-compress-all produced byte-identical fatbins here, and is kept so that no entry is left out by a size heuristic of another nvcc.
  2. --concat packs the six cubins of a kernel before compressing them and saves another 24%, and the RTX 4090 on driver 615 loads it, but NVIDIA gives no driver floor for this nvcc 13.4 addition.
  3. hipcc's bundle compression is zstd: level 22 gives 1.088 MB for 17.9 MB of code objects (level 3, clang's default: 1.29 MB). Level 19 and 22 differ by a few hundred bytes; 22 is zstd's maximum and costs no measurable compile time on these sizes.
  4. icpx generates the SPIR-V fallback image at the final link, also in an AOT build under -fno-sycl-rdc: an AOT object linked without --offload-compress gives a binary whose __CLANG_OFFLOAD_BUNDLE__sycl-spir64 section starts with the SPIR-V magic, the same object linked with the flag gives a zstd frame. ADR-1360 put the flag on the compile line only, so every binary carried a raw SPIR-V image (1.9 MB in libvmaf.so).
  5. Compressing changes no device code: the decompressed cubins and PTX are byte-identical to an uncompressed build's, and the HIP code objects differ only in the name of __hip_cuid_<hash>, a hash of the compile options.

Reproduction

# CUDA: fatbins of one tree, compressed and not, then the extracted cubins and PTX.
meson setup b-on core -Denable_cuda=true
meson setup b-off core -Denable_cuda=true -Dcompress_device_code=false
ninja -C b-on && ninja -C b-off   # or only the src/*.fatbin targets
for d in b-on b-off; do for f in $d/src/*.fatbin; do
  n=$(basename "$f" .fatbin); mkdir -p "x/$d/$n"
  (cd "x/$d/$n" && cuobjdump -xelf all -xptx all "$OLDPWD/$f"); done; done
diff -r x/b-on x/b-off && echo identical

# HIP: unbundle every target of every bundle and compare the sections.
clang-offload-bundler --type=o --input=b-on/src/psnr_score.hsaco --list
clang-offload-bundler --type=o --input=b-on/src/psnr_score.hsaco \
  --unbundle --targets=hipv4-amdgcn-amd-amdhsa--gfx1036 --output=on.co

# The build-time check, on any build directory.
python3 core/src/check_device_compression.py --stamp /dev/null \
  --fatbin b-on/src/*.fatbin

Alternatives explored

  • nvcc speed and balance modes: 4.35 MB and 3.20 MB against 2.60 MB in size mode; the module load time of all 21 fatbins is 1.4 ms raw and 3.0 ms in size mode, which a process start does not notice.
  • Leaving clang CUDA and AdaptiveCpp builds uncompressed without an error: rejected, a build that cannot honour the option must say so.

Open questions

  • The driver floor of nvcc --concat. Enable it once NVIDIA documents the floor or an R580 driver is measured loading it.
  • The Windows CUDA tester zip itself is measured by its workflow run; the figures here are from the Linux build of the same test lists.