The CUDA 13.3 developer blog focused on Tile C++, CompileIQ, green contexts, and DLPack interoperability. The full release notes reveal a materially larger set of changes relevant to VMAFX: a silent-data-corruption fix in the compiler's thread-reconvergence pass (present since 12.8, affects all kernel workloads including our CUDA feature kernels), a second silent-data-corruption fix in __mul24() that has existed since 11.1 and affects any kernel performing 24-bit integer multiply on compile-time constants (our ADM and VIF integer-mode kernels use __mul24 patterns), new remote CPU-to-GPU managed memory mapping that could simplify our CUDA picture-import path, a cuStreamBeginRecaptureToGraph() API relevant to graph-capture in the benchmark CLI, C++23 now officially supported by nvcc (directly enabling ADR-0732's modernisation plan), and official C++23 support in NVRTC. Additionally, Maxwell/Pascal/Volta support was removed in CUDA 13.0 (already past), and cuSPARSE performance improvements (CSR SpMV ALG2 +11%) are not directly relevant but inform any future sparse-feature work. The blog mentioned none of the corruption fixes.
404 — driver package for CUDA 13.3 not yet in XFree86 index
Cross-Reference: research-0734 vs Prior CUDA Research¶
The prior CUDA research files in docs/research/ (see 0047-cuda-graph-capture-feasibility.md, 0091-cambi-cuda-integration.md, 0092-motion-cuda-sub4k-perf-root-cause-2026-05-10.md, etc.) cover workload-specific CUDA correctness and performance analysis. None cover the CUDA toolkit-level bug fixes documented here. This digest is additive; there is no duplication risk with any of the numbered 004x–013x CUDA research files.
Per task instructions, the following topics are intentionally excluded as covered by an earlier research pass (research-0734 predecessor): Tile C++, CompileIQ, green contexts, DLPack.
"Fixed a compiler issue, present since CUDA 12.8, that could cause compiler-inserted thread reconvergence to fail and leave stale or corrupted values in registers, resulting in incorrect program execution."
CRITICAL. Any kernel compiled with CUDA 12.8–13.2 nvcc is potentially affected. Our CUDA feature kernels (ADM, VIF, SSIM, CAMBI, etc.) all use divergent control flow. Silent register corruption → wrong VMAF scores without assertion failure.
DO-NOW: Rebuild all CUDA kernels with CUDA 13.3 nvcc. Add a comment to core/src/cuda/ build instructions and dev/Containerfile noting the minimum safe toolkit version for production builds is 13.3.
CUDA TK Release Notes §CUDA Math
Bug Fix — CUDA 13.3 (introduced 11.1)
"Fixed an issue where silent data corruption could occur when the CUDA Math API __mul24() intrinsic was called with compile-time constant inputs due to undefined behavior from compiler optimizations."
HIGH. The integer VIF and ADM CUDA kernels (integer_vif_cuda.c, integer_adm_score.cu) perform 24-bit integer multiplications. If any such call uses compile-time constants (e.g., __mul24(val, FIXED_SCALE_FACTOR)), scores are silently wrong on CUDA 11.1–13.2.
DO-NOW: Audit every __mul24() call in core/src/feature/cuda/ and core/src/cuda/ for compile-time-constant arguments. Replace with __mul24(val, (int)runtime_var) patterns if any are found. File a tracking row in docs/state.md until audit completes.
CUDA TK Release Notes §CUDA Tools
Bug Fix — CUDA 13.3
"Fixed Data Race in WGMMA A/B Register Copy Propagation" where "ptxas could incorrectly copy-propagate across wgmma.wait_group.sync.aligned."
LOW for current kernels (no WGMMA in VMAFX feature kernels — Hopper tensor core ops). Relevant if tiny-AI inference (ORT) uses Hopper WGMMA.
WAIT-FOR-X: Monitor whether ORT upgrades its Hopper matmul path to WGMMA; if so, CUDA 13.3 ptxas is the minimum safe compiler.
CUDA TK Release Notes §nvJPEG
Bug Fix
"Fixed an issue with boundary handling when decoding a region of interest with NVJPEG_FLAGS_UPSAMPLING_WITH_INTERPOLATION enabled."
"Added cuStreamBeginRecaptureToGraph() API, which allows applications to initiate stream capture into an existing source graph."
MEDIUM. The VMAFX benchmark CLI (vmaf_bench.c) and the potential warm-path graph-capture path (see 0047-cuda-graph-capture-feasibility.md) currently must recreate the full graph on each segment change. cuStreamBeginRecaptureToGraph() enables recapture into an existing graph object, eliminating reallocation cost on parameter-only changes.
WAIT-FOR-X: Only actionable once graph-capture is adopted (research-0047 recommendation). Note in graph-capture ADR when filed.
CUDA TK Release Notes §Managed Memory
New Feature — 13.3
"Added support for remote CPU-to-GPU mapping of managed memory" and "Added System-Allocated Memory (SAM) migration support for CDMM."
MEDIUM. Our cuda/picture.c allocates YUV frames via cudaMallocManaged + explicit prefetch hints. Remote CPU-to-GPU mapping of managed memory could eliminate the explicit cudaMemPrefetchAsync calls for YUV input on systems where the GPU is on a remote NUMA node (multi-GPU or NVLink topology in the Phase 4b distributed cluster). SAM migration is relevant for future Confidential Computing or CXL-attached GPU scenarios.
WAIT-FOR-X: Relevant for Phase 4b cluster deployment (ADR-0709). Revisit when Phase 4b.3 (GPU scheduler integration) is scoped.
CUDA TK Release Notes §CUDA Driver
New NVML API
"Added nvmlDeviceGetRemappedRows_v2 NVML API. In addition to the information returned by nvmlDeviceGetRemappedRows, nvmlDeviceGetRemappedRows_v2 returns the number of inactive row remappings."
LOW-MEDIUM. Phase 4b health reporting (ADR-0709 §monitoring) uses NVML for GPU health. _v2 adds inactive-row count, useful for early ECC degradation detection in production nodes.
WAIT-FOR-X: Add to the Phase 4b monitoring design (health-check probe in the vmafx-node sidecar). Not urgent.
CUDA TK Release Notes §CUDA Python
New Feature
"Added the cuda.core.checkpoint module for CUDA process checkpointing."
LOW for current workloads. Potentially relevant for long-running feature-extraction jobs (K150K sweep, CHUG re-extract) if process migration is needed.
HIGH. ADR-0732 (C++23 internals migration plan, docs/research/0732-vmafx-cpp23-internals-migration-plan.md) identified nvcc C++23 support as a prerequisite blocker for using C++23 in CUDA kernel files. CUDA 13.3 officially unblocks this. Previously only --std=c++20 was supported by nvcc for device code.
DO-NOW: Update core/src/meson.build CUDA flags from -std=c++20 (or whatever is currently set) to -std=c++23 once the container baseline is pinned to CUDA 13.3. Cross-reference ADR-0732 deliverable.
CUDA TK Release Notes §CUDA Compiler
New Feature — 13.3
"Added nvprune functionality to nvcc" for "streamline deployment artifacts and manage builds for targeted GPU architectures."
MEDIUM. Container image size has been a concern (dev/Containerfile ships all sm_ variants). nvprune via nvcc can strip unused .cubin sections from the final .so, reducing libvmaf.so install size when a single --arch=sm_XX is targeted.
WAIT-FOR-X: Actionable in the Phase 4b Docker image optimization pass. Add to dev/Containerfile build notes.
CUDA TK Release Notes §CUDA Compiler
New Feature — 13.3
"Added support for Advanced Control Files (ACFs) in NVIDIA compiler toolchains through the --apply-controls=<file> option."
LOW. ACFs allow per-function or per-loop compiler controls (unrolling, vectorisation) without source annotations. Potentially useful as an alternative to #pragma unroll in hot CUDA loops, but our existing pragmas are well-placed.
NOT-WORTH-IT until a specific kernel benchmark shows a pragma-management pain point.
LOW. VMAFX does not currently use cuFFT LTO. If added in the future, NVRTC dependency must be accounted for.
NOT-WORTH-IT currently.
CUDA TK Release Notes §NPP
Removal
"All legacy NPP APIs without the _Ctx suffix have been deprecated and are now removed."
MEDIUM-HIGH. Any NPP call in core/src/cuda/ (e.g., image conversion, resize ops) that uses the non-_Ctx forms is now a link error on CUDA 13.3.
DO-NOW:grep -r 'npp[A-Z]' core/src/ — if any non-_Ctx NPP calls appear, port them to _Ctx variants before pinning CUDA 13.3. (Likely none, as VMAFX does not heavily use NPP, but verify.)
CUDA TK Release Notes §cuSOLVER
Deprecation
"cuSOLVERMg is deprecated and may be removed in an upcoming major release." "cuSOLVERSp and cuSOLVERRf are fully deprecated and may be removed in an upcoming major release."
NOT RELEVANT. VMAFX does not use cuSOLVER.
Not worth it.
CUDA TK Release Notes §Nsight Eclipse
Removal
"Legacy Nsight Eclipse Edition plugins are no longer delivered in CUDA Toolkit packages beginning with CUDA 13.3."
Affects developer tooling only — not runtime. Use ncu CLI (already our convention).
Not worth it.
CUDA TK Release Notes §cuFFT
Note (13.0, already past)
"Removed support for Maxwell, Pascal, and Volta GPUs, corresponding to compute capabilities earlier than Turing."
Already known / past. VMAFX CI matrix targets sm_75+ (Turing). Confirms Maxwell/Pascal sm_50/sm_60 targets can be dropped from any residual CUDA_ARCH_LIST entries in meson_options.txt.
DO-NOW (minor): Audit meson_options.txt and dev/Containerfile for any residual sm_50, sm_52, sm_60 arch entries; remove them since CUDA 13.x will not compile them.
"Improved TF32 matrix multiplication performance on Hopper by 11% geometric mean, with up to 40% speedup for some small problems."
MEDIUM. The tiny-AI ORT inference path (core/src/dnn/) targets Hopper on our dev machine. cuBLAS TF32 matmul is on the ORT execution provider's hot path. This is a free speedup on ORT upgrade.
WAIT-FOR-X: Free gain on ORT upgrade to CUDA 13.3 cuBLAS. No code change needed. Note in ORT upgrade checklist.
CUDA TK Release Notes §cuBLAS
Performance — 13.3
"Improved TF32 TN matrix multiplication performance on Blackwell and Blackwell Ultra by 27% geometric mean, with up to 3.5x speedup."
LOW-MEDIUM (future). Relevant when the dev machine upgrades to Blackwell.
Not worth it right now.
CUDA TK Release Notes §cuSPARSE
Performance — 13.3
"Improved CSR SpMV ALG2 performance by an average of 11%."
"metrics for the size of SASS instructions" and "warp-can't issue stall samples per HW warp ID slot"
MEDIUM. The warp-can't-issue stall per HW warp ID slot metric directly maps to register-pressure stalls in our VIF and ADM CUDA kernels, which are the historically hot kernels. This is finer-grained than the prior warp-stall metric.
DO-NOW (tooling): Update the /profile-hotpath skill's NCU metric set to include warp_cant_issue_stall_per_hw_warp_id when ncu 2026.2+ is available in the container.
Nsight Compute Release Notes
Source Page Enhancement
"Register Dependencies analysis to identify general purpose register dependencies and occupancy issues due to live register pressure"
HIGH for profiling. Our AVX-512 float convolution fix (PR in state.md) was driven by register pressure; this new metric would have surfaced it directly.
DO-NOW (tooling): Document in docs/development/profiling.md that NCU 2026.2+ Register Dependencies view is the recommended first diagnostic for occupancy-limited kernels.
Do-Now (no external dependency, clear VMAFX benefit)¶
Rebuild all CUDA kernels with CUDA 13.3 nvcc to pick up the thread-reconvergence silent-corruption fix (present since 12.8). Update dev/Containerfile CUDA pin.
Audit __mul24() calls for compile-time-constant arguments in core/src/feature/cuda/ — the 13.3 Math fix (present since 11.1) covers this exact pattern. Estimated scope: ~15 minutes grep + review.
Drop residual sm_50/sm_52/sm_60 arch entries from meson_options.txt and dev/Containerfile; CUDA 13.0 removed Maxwell/Pascal support (confirmed in release notes; safe to clean up now).
Enable -std=c++23 for CUDA device code in core/src/meson.build once the container is pinned to CUDA 13.3 (unblocks ADR-0732 migration plan).
NPP _Ctx audit:grep -r 'npp[A-Z]' core/src/ — port any non-_Ctx calls before pinning CUDA 13.3 (removal is a link error, not a warning).
Wait-For-X (actionable when a specific future milestone is reached)¶
cuStreamBeginRecaptureToGraph() — revisit when graph-capture ADR is filed (research-0047 follow-up).
Managed memory remote mapping / SAM — revisit at Phase 4b.3 GPU scheduler scoping (ADR-0709).
TF32 cuBLAS Hopper speedup — free gain on next ORT upgrade, no code change needed.
nvprune in nvcc — add to Phase 4b Docker image optimization pass.
NCU warp-can't-issue stall per HW warp ID slot metric — add to /profile-hotpath skill when ncu 2026.2 lands in container.
No cuDNN release notes in the CUDA 13.3 document. The toolkit release notes do not mention cuDNN at all for 13.3 — cuDNN ships on its own cadence and has separate release notes. The CUDA 13.3 blog also skipped this. Implication: ORT's cuDNN dependency must be tracked via the cuDNN-specific release notes, not the CUDA toolkit notes.
No driver-level changelog accessible. The XFree86 driver README at the expected URL (610.43.02, the driver paired with CUDA 13.3) returned 404 — NVIDIA had not published it to the XFree86 index at time of research. Driver-level scheduler and UVM changes therefore cannot be confirmed from primary sources.
No PTX ISA 9.3 detail in the release notes. The release notes reference PTX 9.3 only with "For new features from PTX, refer to PTX ISA version 9.3" — the actual changelog requires a separate PTX ISA fetch. No VMAFX-impacting PTX changes are known to be in 9.3, but this gap is noted.