Skip to content

ADRs tagged simd

Auto-generated by scripts/docs/generate-adr-by-tag.sh. Edit ADR Tags: lines to update.

81 ADR(s) carry this tag.

ID Title
ADR-0110 Coverage gate -fprofile-update=atomic for parallel meson tests
ADR-0125 MS-SSIM decimate SIMD fast paths (AVX2 + AVX-512)
ADR-0138 _iqa_convolve AVX2 bit-exact double-precision fast path
ADR-0139 SSIM SIMD accumulate bit-exact to scalar via per-lane scalar double
ADR-0140 SIMD DX framework — header macros + scaffolding skill
ADR-0142 Port Netflix upstream vif_sigma_nsq feature parameter
ADR-0143 Port Netflix upstream generalised AVX convolve for arbitrary filter widths
ADR-0145 motion_v2 NEON SIMD — bit-exact to scalar
ADR-0159 psnr_hvs AVX2 port — bit-exact DCT vectorization (T3-5)
ADR-0160 psnr_hvs NEON port — bit-exact DCT vectorization (T3-5-neon)
ADR-0161 SSIMULACRA 2 SIMD bit-exact ports — AVX2 + AVX-512 + NEON (T3-1 + T3-2)
ADR-0162 SSIMULACRA 2 IIR blur SIMD ports — AVX2 + AVX-512 + NEON (T3-1 phase 2)
ADR-0163 SSIMULACRA 2 picture_to_linear_rgb SIMD ports (T3-1 phase 3)
ADR-0179 float_moment SIMD parity (AVX2 + NEON)
ADR-0180 CPU coverage matrix audit — close 5 stale gaps
ADR-0213 SSIMULACRA 2 SVE2 SIMD parity
ADR-0245 SIMD bit-exact test harness shared header
ADR-0252 ssimulacra2 Vulkan host-path AVX2 + NEON SIMD (T-GPU-OPT-VK-2)
ADR-0328 Cambi cluster port — skip the shared-header rename
ADR-0350 psnr_hvs AVX-512 — re-bench confirms AVX2 ceiling (T3-9 (a))
ADR-0419 Gate SVE2 build probe to non-Darwin hosts
ADR-0452 Port calculate_c_values_row to AVX-512 and NEON
ADR-0463 ADM p-norm fast-path split and VIF scalar-fallback malloc hoist
ADR-0467 SSIMULACRA2 AVX-512 + NEON IIR Blur / picture_to_linear_rgb ULP Audit — Clean Close
ADR-0500 VIF log2 LUT Shrink and Gaussian Filter Cache
ADR-0502 ADM decouple gather prefetch (Approach B)
ADR-0503 vif_subsample_rd_8_avx512 Loop Fission to Reduce ZMM Register Spill
ADR-0504 AVX-512F port of float separable convolution scanlines
ADR-0521 MSVC portability gating — vif_avx512.c noinline/noclone + yuv_input.c S_ISREG/fstat
ADR-0584 0584-moment-sve2-port.md
ADR-0645 Thread integer ADM p-norm through SIMD callbacks
ADR-0771 SIMD twin coverage inventory and gap prioritisation
ADR-0780 NOLINT Cluster Refactor — Slab Allocator, SYCL Stride, ADM Band-Size
ADR-0784 AVX2 SIMD path for integer SSIM horizontal moment accumulation
ADR-0844 float_adm AVX2/AVX-512 F2+F3 — double-precision and FP-contraction
ADR-0853 Remove dead debug-print macros from motion_avx2.c
ADR-0869 Sanitizer-Pass Cleanup — CAMBI Option-Type Mismatch and AVX{2,512} ADM Signed-Shift UB
ADR-0871 SSIM SIMD dispatch installation must be pthread_once-guarded
ADR-0873 ARM64 NEON bit-exactness audit — -ffp-contract=off carve-out scope
ADR-0877 Error-code consistency audit — fork-added MS-SSIM decimate dispatcher
ADR-0891 SIMD bit-exactness round-2 — unify SSIMULACRA 2 colour-matrix on FMA, extend -fp-model=precise to libvmaf_feature_static_lib
ADR-0918 LLVM IR diff harness for bit-exact SIMD paths
ADR-0973 Master CI fixes — Metal MS-SSIM fixture dim + ssimulacra2 icpx XYB bit-exactness
ADR-0987 AVX-512 path for float_moment feature extractor
ADR-1024 R6 per-metric scoring guards — PSNR/ADM correctness fixes
ADR-1025 R6 CUDA/HIP kernel correctness fixes
ADR-1040 Promote integer_ssim_moments_t to shared header (macOS / Windows arm64 build fix)
ADR-1057 Revert float-ADM SIMD dispatch wiring (PR #685) — NEON FMA divergence unfixable in scope
ADR-1104 Remove AVX-512 dispatch from float VIF convolution to restore Netflix golden scores
ADR-1141 Rework the upstream-mirror integer ADM to the fork lint profile, bit-exact
ADR-1142 Whole-codebase standards; lint debt only ratchets down
ADR-1194 One integer-ADM angle_flag predicate for every backend
ADR-1196 Dispatch the SpEED dense matrix product through bit-exact AVX2 / AVX-512 kernels
ADR-1205 The ssimulacra2 FMA unification extends to the scalar fallback and every GPU host copy
ADR-1207 A test gates every feature's score against the host instruction set
ADR-1208 The ssimulacra2 edge-diff SIMD loops take their difference in double
ADR-1237 CAMBI Anti-Dithering AVX2 Vectorization, SpEED SIMD QR Dispatch, and Threaded GPU Flush Alignment
ADR-1253 Scalar references compute their fused multiply-add themselves
ADR-1254 Wide vector register pressure is a Win64 correctness constraint, not a performance one
ADR-1256 Dispatch CAMBI's spatial-mask row SIMD kernels only where they measurably beat scalar
ADR-1257 Retire the Darwin three-tap integer-ADM DWT2 compatibility dispatch
ADR-1260 Windows on ARM64 CPU build-and-test lane
ADR-1283 The whole-tree clang-tidy ratchet gets an arm64 cross lane
ADR-1325 Normalize integer ADM Barten weights with one exponent per scale
ADR-1393 CAMBI's c-values walks clip the window to the frame at every edge
ADR-1402 Integer ADM keeps the scale-0 masking centre tap in int32 and clamps the excess in int64
ADR-1413 Every integer ADM implementation bounds the enhancement gain with the scalar's truncated double product
ADR-1415 every x86 SIMD library is built without FP contraction
ADR-1459 SpEED's covariance kernels return the scalar kernel's bits; they vectorise across sums, not within one
ADR-1469 the psnr_hvs SIMD butterfly is two functions, cut at the same statement in the AVX2 and the NEON twin
ADR-1473 The x86 float ADM wavelet and CSF kernels return the scalar bits and are dispatched; the two reduction kernels are removed
ADR-1482 integer adm on frames of 17 to 32 pixels reads inside the frame and rounds a zero shift with 0; upstream reads index -1 there, and scale 3 differs by up to 0.23
ADR-1488 the psnr_hvs masking threshold is the double root of a float product, as Netflix's source writes it; the double product of PR #552 is removed and the AVX2, NEON, CUDA, HIP and SYCL forms follow, bit for bit
ADR-1490 Insert RC7 CPU capability source of truth and shift the later candidates
ADR-1500 The NEON and SVE2 float_moment kernels add in the scalar's order and return its bits on every input and vector length
ADR-1561 integer VIF converts its residual variance through vif_sv_sq(), which returns x86's value without the undefined double to int32_t conversion
ADR-1601 integer VIF forms the denominator log argument sigma_nsq + sigma1_sq in uint32_t
ADR-1880 Format envelope and device-targeted scoring in the 1.0.0 candidates
ADR-1917 The integer ADM scale-0 horizontal and vertical weight limit is 43900, set by the CSF stage's 16-bit magnitude
ADR-2055 The x86 AVX2 level requires FMA
ADR-2478 Rust replaces the host-side C and C++ through 3.0, behind the unchanged C ABI