Skip to content

ADRs tagged performance

Auto-generated by scripts/docs/generate-adr-by-tag.sh. Edit ADR Tags: lines to update.

69 ADR(s) carry this tag.

ID Title
ADR-0138 _iqa_convolve AVX2 bit-exact double-precision fast path
ADR-0139 SSIM SIMD accumulate bit-exact to scalar via per-lane scalar double
ADR-0145 motion_v2 NEON SIMD — bit-exact to scalar
ADR-0147 Thread-pool job-object recycling + inline data buffer
ADR-0159 psnr_hvs AVX2 port — bit-exact DCT vectorization (T3-5)
ADR-0160 psnr_hvs NEON port — bit-exact DCT vectorization (T3-5-neon)
ADR-0161 SSIMULACRA 2 SIMD bit-exact ports — AVX2 + AVX-512 + NEON (T3-1 + T3-2)
ADR-0162 SSIMULACRA 2 IIR blur SIMD ports — AVX2 + AVX-512 + NEON (T3-1 phase 2)
ADR-0251 Vulkan VkImage import — v2 async pending-fence model (T7-29 part 4)
ADR-0252 ssimulacra2 Vulkan host-path AVX2 + NEON SIMD (T-GPU-OPT-VK-2)
ADR-0353 Vulkan submit-pool migration PR-B — six secondary kernels
ADR-0378 Per-picture CUDA streams must use CU_STREAM_NON_BLOCKING
ADR-0383 K150K corpus scoring driver — parallel CPU worker redesign
ADR-0445 Persistent VkPipelineCache for Vulkan compute backend
ADR-0454 VIF CUDA shared-memory staging for horizontal and vertical filter passes
ADR-0464 CAMBI CUDA spatial-mask shared-memory tile
ADR-0489 CAMBI SYCL — Replace GPU-to-GPU q.wait() Calls with Event Chains (SY-1)
ADR-0503 vif_subsample_rd_8_avx512 Loop Fission to Reduce ZMM Register Spill
ADR-0504 AVX-512F port of float separable convolution scanlines
ADR-0551 VCQ-223 LocalExplainer CI timeout — root cause and fix path
ADR-0562 VCQ-223 LocalExplainer hang fix — cap neighbor_samples in test runner
ADR-0600 Port upstream USE_DIRECT_READ zero-copy input path (Netflix/vmaf@30a6e2a8d)
ADR-0607 vmaf-tune compare: decode reference YUV once for the entire run
ADR-0743 CUDA VIF filter1d ncu-driven performance optimizations
ADR-0744 CUDA adm_cm __launch_bounds__(128, 8) register reduction (ms_ssim_decimate smem tiling reverted)
ADR-0750 Hardware Measurement Verdict for PR perf/cuda-ms-ssim-decimate-adm-cm-ncu-driven
ADR-0754 0754-cuda-ssim-vert-combine-ldg-pinned-leak.md
ADR-0757 0757-cuda-ms-ssim-vert-lcs-horiz-ldg.md
ADR-0759 HIP ADM — AdmBufferHip passed by pointer (F3 fix)
ADR-0762 CUDA CIEDE2000 8bpc/16bpc — __ldg() read-only cache routing (F3 fix)
ADR-0763 0763-cuda-adm-decouple-ldg.md
ADR-0773 CUDA ADM decouple-inline — __ldg() F3 fix on active path
ADR-0779 eBPF FUSE read-path bypass for vmafx-node rclone mounts
ADR-0784 AVX2 SIMD path for integer SSIM horizontal moment accumulation
ADR-0845 CUDA motion — multi-frame SAD batching to reduce per-launch overhead
ADR-0907 Wall-clock perf regression gate over the multi-resolution baseline
ADR-0923 Adopt BuildKit cache mounts and ccache across the container build matrix
ADR-0931 MCP server — replace subprocess delegation with direct cgo (Phase 1)
ADR-0932 iter.Seq[T] companion APIs for single-pass Go collections
ADR-0987 AVX-512 path for float_moment feature extractor
ADR-0996 eBPF FUSE bypass for rclone zero-copy path in vmafx-node
ADR-1196 Dispatch the SpEED dense matrix product through bit-exact AVX2 / AVX-512 kernels
ADR-1224 CUDA Tile C++ is not adopted; the audit's incidental findings are
ADR-1226 Size the CUDA AIM CM launch by SM count, not by a fixed rows-per-thread
ADR-1228 A recurring "faster than upstream, and still exact" milestone
ADR-1341 Stage correctness, benchmarking, and retraining across RC1 through RC3
ADR-1357 Run the SYCL CAMBI extractor entirely on the device
ADR-1358 The SYCL SpEED twins are device-resident, with the 25x25 linear algebra on the device and exact fp32 arithmetic
ADR-1362 The SYCL integer ADM twin computes AIM on the device and finalises every ADM output in the CPU's float arithmetic
ADR-1363 The SYCL ssimulacra2 twin is device-resident, and float_ms_ssim_sycl waits once per frame
ADR-1366 The vmaf CLI reads each input on its own thread, a bounded number of frames ahead
ADR-1369 SYCL twins read the planes the state uploads once per frame; opt-in shared chroma planes and a device-side slot fence
ADR-1370 float_ssim_sycl decimates on the device, bit-identical to the CPU
ADR-1371 SYCL motion differences the frames before the blur, in one shared kernel
ADR-1377 HIP motion differences the frames before the blur, in one shared kernel, and waits only in collect
ADR-1378 Run the HIP CAMBI extractor entirely on the device
ADR-1379 Run the CUDA CAMBI extractor entirely on the device
ADR-1380 The CUDA SpEED twins are device-resident, with the 25x25 linear algebra on the device and CPU-exact fp32 arithmetic
ADR-1384 Run the HIP SpEED twins entirely on the device, in the CPU's fp32 arithmetic
ADR-1390 The HIP ssimulacra2 twin is device-resident with tiled row-pass Gaussian blur
ADR-1391 The CUDA ssimulacra2 twin is device-resident
ADR-1392 CUDA motion, PSNR and moment kernels add one atomic per block, the motion SAD filters separably, and PSNR selects its plane with constant indices
ADR-1399 float_ssim_cuda decimates and convolves on the device with the CPU's arithmetic
ADR-1405 float_ssim_hip decimates on the device, bit-identical to the CPU
ADR-1408 A VmafContext uploads each plane of a frame once and every HIP twin reads that copy
ADR-1410 SYCL CLI picture pool allocates pinned host USM to bypass staging upload
ADR-1473 The x86 float ADM wavelet and CSF kernels return the scalar bits and are dispatched; the two reduction kernels are removed
ADR-1794 Bilinear column tables without a width limit
ADR-2795 Integer ADM evaluates two viewing distances in one context