ADRs tagged performance¶
Auto-generated by scripts/docs/generate-adr-by-tag.sh. Edit ADR Tags: lines to update.
69 ADR(s) carry this tag.
| ID | Title |
|---|---|
| ADR-0138 | _iqa_convolve AVX2 bit-exact double-precision fast path |
| ADR-0139 | SSIM SIMD accumulate bit-exact to scalar via per-lane scalar double |
| ADR-0145 | motion_v2 NEON SIMD — bit-exact to scalar |
| ADR-0147 | Thread-pool job-object recycling + inline data buffer |
| ADR-0159 | psnr_hvs AVX2 port — bit-exact DCT vectorization (T3-5) |
| ADR-0160 | psnr_hvs NEON port — bit-exact DCT vectorization (T3-5-neon) |
| ADR-0161 | SSIMULACRA 2 SIMD bit-exact ports — AVX2 + AVX-512 + NEON (T3-1 + T3-2) |
| ADR-0162 | SSIMULACRA 2 IIR blur SIMD ports — AVX2 + AVX-512 + NEON (T3-1 phase 2) |
| ADR-0251 | Vulkan VkImage import — v2 async pending-fence model (T7-29 part 4) |
| ADR-0252 | ssimulacra2 Vulkan host-path AVX2 + NEON SIMD (T-GPU-OPT-VK-2) |
| ADR-0353 | Vulkan submit-pool migration PR-B — six secondary kernels |
| ADR-0378 | Per-picture CUDA streams must use CU_STREAM_NON_BLOCKING |
| ADR-0383 | K150K corpus scoring driver — parallel CPU worker redesign |
| ADR-0445 | Persistent VkPipelineCache for Vulkan compute backend |
| ADR-0454 | VIF CUDA shared-memory staging for horizontal and vertical filter passes |
| ADR-0464 | CAMBI CUDA spatial-mask shared-memory tile |
| ADR-0489 | CAMBI SYCL — Replace GPU-to-GPU q.wait() Calls with Event Chains (SY-1) |
| ADR-0503 | vif_subsample_rd_8_avx512 Loop Fission to Reduce ZMM Register Spill |
| ADR-0504 | AVX-512F port of float separable convolution scanlines |
| ADR-0551 | VCQ-223 LocalExplainer CI timeout — root cause and fix path |
| ADR-0562 | VCQ-223 LocalExplainer hang fix — cap neighbor_samples in test runner |
| ADR-0600 | Port upstream USE_DIRECT_READ zero-copy input path (Netflix/vmaf@30a6e2a8d) |
| ADR-0607 | vmaf-tune compare: decode reference YUV once for the entire run |
| ADR-0743 | CUDA VIF filter1d ncu-driven performance optimizations |
| ADR-0744 | CUDA adm_cm __launch_bounds__(128, 8) register reduction (ms_ssim_decimate smem tiling reverted) |
| ADR-0750 | Hardware Measurement Verdict for PR perf/cuda-ms-ssim-decimate-adm-cm-ncu-driven |
| ADR-0754 | 0754-cuda-ssim-vert-combine-ldg-pinned-leak.md |
| ADR-0757 | 0757-cuda-ms-ssim-vert-lcs-horiz-ldg.md |
| ADR-0759 | HIP ADM — AdmBufferHip passed by pointer (F3 fix) |
| ADR-0762 | CUDA CIEDE2000 8bpc/16bpc — __ldg() read-only cache routing (F3 fix) |
| ADR-0763 | 0763-cuda-adm-decouple-ldg.md |
| ADR-0773 | CUDA ADM decouple-inline — __ldg() F3 fix on active path |
| ADR-0779 | eBPF FUSE read-path bypass for vmafx-node rclone mounts |
| ADR-0784 | AVX2 SIMD path for integer SSIM horizontal moment accumulation |
| ADR-0845 | CUDA motion — multi-frame SAD batching to reduce per-launch overhead |
| ADR-0907 | Wall-clock perf regression gate over the multi-resolution baseline |
| ADR-0923 | Adopt BuildKit cache mounts and ccache across the container build matrix |
| ADR-0931 | MCP server — replace subprocess delegation with direct cgo (Phase 1) |
| ADR-0932 | iter.Seq[T] companion APIs for single-pass Go collections |
| ADR-0987 | AVX-512 path for float_moment feature extractor |
| ADR-0996 | eBPF FUSE bypass for rclone zero-copy path in vmafx-node |
| ADR-1196 | Dispatch the SpEED dense matrix product through bit-exact AVX2 / AVX-512 kernels |
| ADR-1224 | CUDA Tile C++ is not adopted; the audit's incidental findings are |
| ADR-1226 | Size the CUDA AIM CM launch by SM count, not by a fixed rows-per-thread |
| ADR-1228 | A recurring "faster than upstream, and still exact" milestone |
| ADR-1341 | Stage correctness, benchmarking, and retraining across RC1 through RC3 |
| ADR-1357 | Run the SYCL CAMBI extractor entirely on the device |
| ADR-1358 | The SYCL SpEED twins are device-resident, with the 25x25 linear algebra on the device and exact fp32 arithmetic |
| ADR-1362 | The SYCL integer ADM twin computes AIM on the device and finalises every ADM output in the CPU's float arithmetic |
| ADR-1363 | The SYCL ssimulacra2 twin is device-resident, and float_ms_ssim_sycl waits once per frame |
| ADR-1366 | The vmaf CLI reads each input on its own thread, a bounded number of frames ahead |
| ADR-1369 | SYCL twins read the planes the state uploads once per frame; opt-in shared chroma planes and a device-side slot fence |
| ADR-1370 | float_ssim_sycl decimates on the device, bit-identical to the CPU |
| ADR-1371 | SYCL motion differences the frames before the blur, in one shared kernel |
| ADR-1377 | HIP motion differences the frames before the blur, in one shared kernel, and waits only in collect |
| ADR-1378 | Run the HIP CAMBI extractor entirely on the device |
| ADR-1379 | Run the CUDA CAMBI extractor entirely on the device |
| ADR-1380 | The CUDA SpEED twins are device-resident, with the 25x25 linear algebra on the device and CPU-exact fp32 arithmetic |
| ADR-1384 | Run the HIP SpEED twins entirely on the device, in the CPU's fp32 arithmetic |
| ADR-1390 | The HIP ssimulacra2 twin is device-resident with tiled row-pass Gaussian blur |
| ADR-1391 | The CUDA ssimulacra2 twin is device-resident |
| ADR-1392 | CUDA motion, PSNR and moment kernels add one atomic per block, the motion SAD filters separably, and PSNR selects its plane with constant indices |
| ADR-1399 | float_ssim_cuda decimates and convolves on the device with the CPU's arithmetic |
| ADR-1405 | float_ssim_hip decimates on the device, bit-identical to the CPU |
| ADR-1408 | A VmafContext uploads each plane of a frame once and every HIP twin reads that copy |
| ADR-1410 | SYCL CLI picture pool allocates pinned host USM to bypass staging upload |
| ADR-1473 | The x86 float ADM wavelet and CSF kernels return the scalar bits and are dispatched; the two reduction kernels are removed |
| ADR-1794 | Bilinear column tables without a width limit |
| ADR-2795 | Integer ADM evaluates two viewing distances in one context |