Skip to content

Research-0734: CUDA 13.3 Non-Blog Findings — Bug Fixes and APIs the Developer Blog Skipped

Date: 2026-05-28 Status: Final Author: lusoris / Claude (agent) Reproducer: WebFetch against the five primary sources listed in §Sources


Executive Summary

The CUDA 13.3 developer blog focused on Tile C++, CompileIQ, green contexts, and DLPack interoperability. The full release notes reveal a materially larger set of changes relevant to VMAFX: a silent-data-corruption fix in the compiler's thread-reconvergence pass (present since 12.8, affects all kernel workloads including our CUDA feature kernels), a second silent-data-corruption fix in __mul24() that has existed since 11.1 and affects any kernel performing 24-bit integer multiply on compile-time constants (our ADM and VIF integer-mode kernels use __mul24 patterns), new remote CPU-to-GPU managed memory mapping that could simplify our CUDA picture-import path, a cuStreamBeginRecaptureToGraph() API relevant to graph-capture in the benchmark CLI, C++23 now officially supported by nvcc (directly enabling ADR-0732's modernisation plan), and official C++23 support in NVRTC. Additionally, Maxwell/Pascal/Volta support was removed in CUDA 13.0 (already past), and cuSPARSE performance improvements (CSR SpMV ALG2 +11%) are not directly relevant but inform any future sparse-feature work. The blog mentioned none of the corruption fixes.


Sources

# URL Status
1 https://docs.nvidia.com/cuda/cuda-toolkit-release-notes/index.html Fetched
2 https://docs.nvidia.com/cuda/cuda-c-programming-guide/index.html Fetched (no 13.3 changelog section; deferred to PDF)
3 https://docs.nvidia.com/nsight-compute/ReleaseNotes/index.html Fetched (latest = 2026.2, CUDA 13.3 era)
4 https://docs.nvidia.com/cuda/cuda-runtime-api/index.html Not fetched — release notes source (1) covers runtime API changes
5 https://download.nvidia.com/XFree86/Linux-x86_64/610.43.02/README/changelog.html 404 — driver package for CUDA 13.3 not yet in XFree86 index

Cross-Reference: research-0734 vs Prior CUDA Research

The prior CUDA research files in docs/research/ (see 0047-cuda-graph-capture-feasibility.md, 0091-cambi-cuda-integration.md, 0092-motion-cuda-sub4k-perf-root-cause-2026-05-10.md, etc.) cover workload-specific CUDA correctness and performance analysis. None cover the CUDA toolkit-level bug fixes documented here. This digest is additive; there is no duplication risk with any of the numbered 004x–013x CUDA research files.

Per task instructions, the following topics are intentionally excluded as covered by an earlier research pass (research-0734 predecessor): Tile C++, CompileIQ, green contexts, DLPack.


Findings Table

Category A: Correctness / Bug Fixes

Source Section Verbatim Quote VMAFX Relevance Recommendation
CUDA TK Release Notes §CUDA Compiler Bug Fix — CUDA 13.3 "Fixed a compiler issue, present since CUDA 12.8, that could cause compiler-inserted thread reconvergence to fail and leave stale or corrupted values in registers, resulting in incorrect program execution." CRITICAL. Any kernel compiled with CUDA 12.8–13.2 nvcc is potentially affected. Our CUDA feature kernels (ADM, VIF, SSIM, CAMBI, etc.) all use divergent control flow. Silent register corruption → wrong VMAF scores without assertion failure. DO-NOW: Rebuild all CUDA kernels with CUDA 13.3 nvcc. Add a comment to core/src/cuda/ build instructions and dev/Containerfile noting the minimum safe toolkit version for production builds is 13.3.
CUDA TK Release Notes §CUDA Math Bug Fix — CUDA 13.3 (introduced 11.1) "Fixed an issue where silent data corruption could occur when the CUDA Math API __mul24() intrinsic was called with compile-time constant inputs due to undefined behavior from compiler optimizations." HIGH. The integer VIF and ADM CUDA kernels (integer_vif_cuda.c, integer_adm_score.cu) perform 24-bit integer multiplications. If any such call uses compile-time constants (e.g., __mul24(val, FIXED_SCALE_FACTOR)), scores are silently wrong on CUDA 11.1–13.2. DO-NOW: Audit every __mul24() call in core/src/feature/cuda/ and core/src/cuda/ for compile-time-constant arguments. Replace with __mul24(val, (int)runtime_var) patterns if any are found. File a tracking row in docs/state.md until audit completes.
CUDA TK Release Notes §CUDA Tools Bug Fix — CUDA 13.3 "Fixed Data Race in WGMMA A/B Register Copy Propagation" where "ptxas could incorrectly copy-propagate across wgmma.wait_group.sync.aligned." LOW for current kernels (no WGMMA in VMAFX feature kernels — Hopper tensor core ops). Relevant if tiny-AI inference (ORT) uses Hopper WGMMA. WAIT-FOR-X: Monitor whether ORT upgrades its Hopper matmul path to WGMMA; if so, CUDA 13.3 ptxas is the minimum safe compiler.
CUDA TK Release Notes §nvJPEG Bug Fix "Fixed an issue with boundary handling when decoding a region of interest with NVJPEG_FLAGS_UPSAMPLING_WITH_INTERPOLATION enabled." NOT RELEVANT. VMAFX does not use nvJPEG. Not worth it.

Category B: New APIs / Runtime

Source Section Verbatim Quote VMAFX Relevance Recommendation
CUDA TK Release Notes §CUDA Driver New API — 13.3 "Added cuStreamBeginRecaptureToGraph() API, which allows applications to initiate stream capture into an existing source graph." MEDIUM. The VMAFX benchmark CLI (vmaf_bench.c) and the potential warm-path graph-capture path (see 0047-cuda-graph-capture-feasibility.md) currently must recreate the full graph on each segment change. cuStreamBeginRecaptureToGraph() enables recapture into an existing graph object, eliminating reallocation cost on parameter-only changes. WAIT-FOR-X: Only actionable once graph-capture is adopted (research-0047 recommendation). Note in graph-capture ADR when filed.
CUDA TK Release Notes §Managed Memory New Feature — 13.3 "Added support for remote CPU-to-GPU mapping of managed memory" and "Added System-Allocated Memory (SAM) migration support for CDMM." MEDIUM. Our cuda/picture.c allocates YUV frames via cudaMallocManaged + explicit prefetch hints. Remote CPU-to-GPU mapping of managed memory could eliminate the explicit cudaMemPrefetchAsync calls for YUV input on systems where the GPU is on a remote NUMA node (multi-GPU or NVLink topology in the Phase 4b distributed cluster). SAM migration is relevant for future Confidential Computing or CXL-attached GPU scenarios. WAIT-FOR-X: Relevant for Phase 4b cluster deployment (ADR-0709). Revisit when Phase 4b.3 (GPU scheduler integration) is scoped.
CUDA TK Release Notes §CUDA Driver New NVML API "Added nvmlDeviceGetRemappedRows_v2 NVML API. In addition to the information returned by nvmlDeviceGetRemappedRows, nvmlDeviceGetRemappedRows_v2 returns the number of inactive row remappings." LOW-MEDIUM. Phase 4b health reporting (ADR-0709 §monitoring) uses NVML for GPU health. _v2 adds inactive-row count, useful for early ECC degradation detection in production nodes. WAIT-FOR-X: Add to the Phase 4b monitoring design (health-check probe in the vmafx-node sidecar). Not urgent.
CUDA TK Release Notes §CUDA Python New Feature "Added the cuda.core.checkpoint module for CUDA process checkpointing." LOW for current workloads. Potentially relevant for long-running feature-extraction jobs (K150K sweep, CHUG re-extract) if process migration is needed. NOT-WORTH-IT for current scope. Note for future.

Category C: Compiler / Build

Source Section Verbatim Quote VMAFX Relevance Recommendation
CUDA TK Release Notes §CUDA Compiler New Feature — 13.3 "Added official C++23 support in nvcc and NVRTC." HIGH. ADR-0732 (C++23 internals migration plan, docs/research/0732-vmafx-cpp23-internals-migration-plan.md) identified nvcc C++23 support as a prerequisite blocker for using C++23 in CUDA kernel files. CUDA 13.3 officially unblocks this. Previously only --std=c++20 was supported by nvcc for device code. DO-NOW: Update core/src/meson.build CUDA flags from -std=c++20 (or whatever is currently set) to -std=c++23 once the container baseline is pinned to CUDA 13.3. Cross-reference ADR-0732 deliverable.
CUDA TK Release Notes §CUDA Compiler New Feature — 13.3 "Added nvprune functionality to nvcc" for "streamline deployment artifacts and manage builds for targeted GPU architectures." MEDIUM. Container image size has been a concern (dev/Containerfile ships all sm_ variants). nvprune via nvcc can strip unused .cubin sections from the final .so, reducing libvmaf.so install size when a single --arch=sm_XX is targeted. WAIT-FOR-X: Actionable in the Phase 4b Docker image optimization pass. Add to dev/Containerfile build notes.
CUDA TK Release Notes §CUDA Compiler New Feature — 13.3 "Added support for Advanced Control Files (ACFs) in NVIDIA compiler toolchains through the --apply-controls=<file> option." LOW. ACFs allow per-function or per-loop compiler controls (unrolling, vectorisation) without source annotations. Potentially useful as an alternative to #pragma unroll in hot CUDA loops, but our existing pragmas are well-placed. NOT-WORTH-IT until a specific kernel benchmark shows a pragma-management pain point.

Category D: Deprecations / Removals

Source Section Verbatim Quote VMAFX Relevance Recommendation
CUDA TK Release Notes §cuFFT Deprecation "Using cuFFT link-time optimized (LTO) kernels now requires NVRTC." (13.2 change) LOW. VMAFX does not currently use cuFFT LTO. If added in the future, NVRTC dependency must be accounted for. NOT-WORTH-IT currently.
CUDA TK Release Notes §NPP Removal "All legacy NPP APIs without the _Ctx suffix have been deprecated and are now removed." MEDIUM-HIGH. Any NPP call in core/src/cuda/ (e.g., image conversion, resize ops) that uses the non-_Ctx forms is now a link error on CUDA 13.3. DO-NOW: grep -r 'npp[A-Z]' core/src/ — if any non-_Ctx NPP calls appear, port them to _Ctx variants before pinning CUDA 13.3. (Likely none, as VMAFX does not heavily use NPP, but verify.)
CUDA TK Release Notes §cuSOLVER Deprecation "cuSOLVERMg is deprecated and may be removed in an upcoming major release." "cuSOLVERSp and cuSOLVERRf are fully deprecated and may be removed in an upcoming major release." NOT RELEVANT. VMAFX does not use cuSOLVER. Not worth it.
CUDA TK Release Notes §Nsight Eclipse Removal "Legacy Nsight Eclipse Edition plugins are no longer delivered in CUDA Toolkit packages beginning with CUDA 13.3." Affects developer tooling only — not runtime. Use ncu CLI (already our convention). Not worth it.
CUDA TK Release Notes §cuFFT Note (13.0, already past) "Removed support for Maxwell, Pascal, and Volta GPUs, corresponding to compute capabilities earlier than Turing." Already known / past. VMAFX CI matrix targets sm_75+ (Turing). Confirms Maxwell/Pascal sm_50/sm_60 targets can be dropped from any residual CUDA_ARCH_LIST entries in meson_options.txt. DO-NOW (minor): Audit meson_options.txt and dev/Containerfile for any residual sm_50, sm_52, sm_60 arch entries; remove them since CUDA 13.x will not compile them.

Category E: Performance (not on blog)

Source Section Verbatim Quote VMAFX Relevance Recommendation
CUDA TK Release Notes §cuBLAS Performance — 13.3 "Improved TF32 matrix multiplication performance on Hopper by 11% geometric mean, with up to 40% speedup for some small problems." MEDIUM. The tiny-AI ORT inference path (core/src/dnn/) targets Hopper on our dev machine. cuBLAS TF32 matmul is on the ORT execution provider's hot path. This is a free speedup on ORT upgrade. WAIT-FOR-X: Free gain on ORT upgrade to CUDA 13.3 cuBLAS. No code change needed. Note in ORT upgrade checklist.
CUDA TK Release Notes §cuBLAS Performance — 13.3 "Improved TF32 TN matrix multiplication performance on Blackwell and Blackwell Ultra by 27% geometric mean, with up to 3.5x speedup." LOW-MEDIUM (future). Relevant when the dev machine upgrades to Blackwell. Not worth it right now.
CUDA TK Release Notes §cuSPARSE Performance — 13.3 "Improved CSR SpMV ALG2 performance by an average of 11%." NOT RELEVANT to current feature kernels. Not worth it.

Category F: Nsight Compute 2026.2 (CUDA 13.3 era)

Source Section Verbatim Quote VMAFX Relevance Recommendation
Nsight Compute Release Notes New Metric "metrics for the size of SASS instructions" and "warp-can't issue stall samples per HW warp ID slot" MEDIUM. The warp-can't-issue stall per HW warp ID slot metric directly maps to register-pressure stalls in our VIF and ADM CUDA kernels, which are the historically hot kernels. This is finer-grained than the prior warp-stall metric. DO-NOW (tooling): Update the /profile-hotpath skill's NCU metric set to include warp_cant_issue_stall_per_hw_warp_id when ncu 2026.2+ is available in the container.
Nsight Compute Release Notes Source Page Enhancement "Register Dependencies analysis to identify general purpose register dependencies and occupancy issues due to live register pressure" HIGH for profiling. Our AVX-512 float convolution fix (PR in state.md) was driven by register pressure; this new metric would have surfaced it directly. DO-NOW (tooling): Document in docs/development/profiling.md that NCU 2026.2+ Register Dependencies view is the recommended first diagnostic for occupancy-limited kernels.

Actionable Summary

Do-Now (no external dependency, clear VMAFX benefit)

  1. Rebuild all CUDA kernels with CUDA 13.3 nvcc to pick up the thread-reconvergence silent-corruption fix (present since 12.8). Update dev/Containerfile CUDA pin.
  2. Audit __mul24() calls for compile-time-constant arguments in core/src/feature/cuda/ — the 13.3 Math fix (present since 11.1) covers this exact pattern. Estimated scope: ~15 minutes grep + review.
  3. Drop residual sm_50/sm_52/sm_60 arch entries from meson_options.txt and dev/Containerfile; CUDA 13.0 removed Maxwell/Pascal support (confirmed in release notes; safe to clean up now).
  4. Enable -std=c++23 for CUDA device code in core/src/meson.build once the container is pinned to CUDA 13.3 (unblocks ADR-0732 migration plan).
  5. NPP _Ctx audit: grep -r 'npp[A-Z]' core/src/ — port any non-_Ctx calls before pinning CUDA 13.3 (removal is a link error, not a warning).

Wait-For-X (actionable when a specific future milestone is reached)

  • cuStreamBeginRecaptureToGraph() — revisit when graph-capture ADR is filed (research-0047 follow-up).
  • Managed memory remote mapping / SAM — revisit at Phase 4b.3 GPU scheduler scoping (ADR-0709).
  • TF32 cuBLAS Hopper speedup — free gain on next ORT upgrade, no code change needed.
  • nvprune in nvcc — add to Phase 4b Docker image optimization pass.
  • NCU warp-can't-issue stall per HW warp ID slot metric — add to /profile-hotpath skill when ncu 2026.2 lands in container.

Not Worth It

  • cuSOLVER / cuSPARSE deprecations (VMAFX does not use these libraries).
  • nvJPEG boundary-handling fix (no nvJPEG usage).
  • cufftDebug deprecation (no cuFFT usage in hot path).
  • ACFs compiler controls (existing #pragma unroll coverage is adequate).
  • CUDA process checkpointing (not relevant to current workload model).

What Was Unexpectedly Absent

  1. No cuDNN release notes in the CUDA 13.3 document. The toolkit release notes do not mention cuDNN at all for 13.3 — cuDNN ships on its own cadence and has separate release notes. The CUDA 13.3 blog also skipped this. Implication: ORT's cuDNN dependency must be tracked via the cuDNN-specific release notes, not the CUDA toolkit notes.
  2. No driver-level changelog accessible. The XFree86 driver README at the expected URL (610.43.02, the driver paired with CUDA 13.3) returned 404 — NVIDIA had not published it to the XFree86 index at time of research. Driver-level scheduler and UVM changes therefore cannot be confirmed from primary sources.
  3. No PTX ISA 9.3 detail in the release notes. The release notes reference PTX 9.3 only with "For new features from PTX, refer to PTX ISA version 9.3" — the actual changelog requires a separate PTX ISA fetch. No VMAFX-impacting PTX changes are known to be in 9.3, but this gap is noted.

Six Deep-Dive Deliverables Checklist (ADR-0108)

  1. Research digest: This file IS the research deliverable.
  2. Decision matrix: no decision matrix needed: research-only — findings drive Do-Now action items, not an architectural decision between alternatives.
  3. Rebase-sensitive invariants: no rebase-sensitive invariants — digest is doc-only; no C source changes.
  4. Reproducer: WebFetch the five URLs listed in §Sources above.
  5. Changelog fragment: changelog.d/changed/docs-cuda-13.3-non-blog-findings.md (created in this PR).
  6. Rebase notes entry: docs/rebase-notes.md entry added in this PR.

Per-Surface Docs (CLAUDE.md §12 r10)

no user-discoverable surface change — research digest