ADRs tagged simd¶
Auto-generated by scripts/docs/generate-adr-by-tag.sh. Edit ADR Tags: lines to update.
81 ADR(s) carry this tag.
| ID | Title |
|---|---|
| ADR-0110 | Coverage gate -fprofile-update=atomic for parallel meson tests |
| ADR-0125 | MS-SSIM decimate SIMD fast paths (AVX2 + AVX-512) |
| ADR-0138 | _iqa_convolve AVX2 bit-exact double-precision fast path |
| ADR-0139 | SSIM SIMD accumulate bit-exact to scalar via per-lane scalar double |
| ADR-0140 | SIMD DX framework — header macros + scaffolding skill |
| ADR-0142 | Port Netflix upstream vif_sigma_nsq feature parameter |
| ADR-0143 | Port Netflix upstream generalised AVX convolve for arbitrary filter widths |
| ADR-0145 | motion_v2 NEON SIMD — bit-exact to scalar |
| ADR-0159 | psnr_hvs AVX2 port — bit-exact DCT vectorization (T3-5) |
| ADR-0160 | psnr_hvs NEON port — bit-exact DCT vectorization (T3-5-neon) |
| ADR-0161 | SSIMULACRA 2 SIMD bit-exact ports — AVX2 + AVX-512 + NEON (T3-1 + T3-2) |
| ADR-0162 | SSIMULACRA 2 IIR blur SIMD ports — AVX2 + AVX-512 + NEON (T3-1 phase 2) |
| ADR-0163 | SSIMULACRA 2 picture_to_linear_rgb SIMD ports (T3-1 phase 3) |
| ADR-0179 | float_moment SIMD parity (AVX2 + NEON) |
| ADR-0180 | CPU coverage matrix audit — close 5 stale gaps |
| ADR-0213 | SSIMULACRA 2 SVE2 SIMD parity |
| ADR-0245 | SIMD bit-exact test harness shared header |
| ADR-0252 | ssimulacra2 Vulkan host-path AVX2 + NEON SIMD (T-GPU-OPT-VK-2) |
| ADR-0328 | Cambi cluster port — skip the shared-header rename |
| ADR-0350 | psnr_hvs AVX-512 — re-bench confirms AVX2 ceiling (T3-9 (a)) |
| ADR-0419 | Gate SVE2 build probe to non-Darwin hosts |
| ADR-0452 | Port calculate_c_values_row to AVX-512 and NEON |
| ADR-0463 | ADM p-norm fast-path split and VIF scalar-fallback malloc hoist |
| ADR-0467 | SSIMULACRA2 AVX-512 + NEON IIR Blur / picture_to_linear_rgb ULP Audit — Clean Close |
| ADR-0500 | VIF log2 LUT Shrink and Gaussian Filter Cache |
| ADR-0502 | ADM decouple gather prefetch (Approach B) |
| ADR-0503 | vif_subsample_rd_8_avx512 Loop Fission to Reduce ZMM Register Spill |
| ADR-0504 | AVX-512F port of float separable convolution scanlines |
| ADR-0521 | MSVC portability gating — vif_avx512.c noinline/noclone + yuv_input.c S_ISREG/fstat |
| ADR-0584 | 0584-moment-sve2-port.md |
| ADR-0645 | Thread integer ADM p-norm through SIMD callbacks |
| ADR-0771 | SIMD twin coverage inventory and gap prioritisation |
| ADR-0780 | NOLINT Cluster Refactor — Slab Allocator, SYCL Stride, ADM Band-Size |
| ADR-0784 | AVX2 SIMD path for integer SSIM horizontal moment accumulation |
| ADR-0844 | float_adm AVX2/AVX-512 F2+F3 — double-precision and FP-contraction |
| ADR-0853 | Remove dead debug-print macros from motion_avx2.c |
| ADR-0869 | Sanitizer-Pass Cleanup — CAMBI Option-Type Mismatch and AVX{2,512} ADM Signed-Shift UB |
| ADR-0871 | SSIM SIMD dispatch installation must be pthread_once-guarded |
| ADR-0873 | ARM64 NEON bit-exactness audit — -ffp-contract=off carve-out scope |
| ADR-0877 | Error-code consistency audit — fork-added MS-SSIM decimate dispatcher |
| ADR-0891 | SIMD bit-exactness round-2 — unify SSIMULACRA 2 colour-matrix on FMA, extend -fp-model=precise to libvmaf_feature_static_lib |
| ADR-0918 | LLVM IR diff harness for bit-exact SIMD paths |
| ADR-0973 | Master CI fixes — Metal MS-SSIM fixture dim + ssimulacra2 icpx XYB bit-exactness |
| ADR-0987 | AVX-512 path for float_moment feature extractor |
| ADR-1024 | R6 per-metric scoring guards — PSNR/ADM correctness fixes |
| ADR-1025 | R6 CUDA/HIP kernel correctness fixes |
| ADR-1040 | Promote integer_ssim_moments_t to shared header (macOS / Windows arm64 build fix) |
| ADR-1057 | Revert float-ADM SIMD dispatch wiring (PR #685) — NEON FMA divergence unfixable in scope |
| ADR-1104 | Remove AVX-512 dispatch from float VIF convolution to restore Netflix golden scores |
| ADR-1141 | Rework the upstream-mirror integer ADM to the fork lint profile, bit-exact |
| ADR-1142 | Whole-codebase standards; lint debt only ratchets down |
| ADR-1194 | One integer-ADM angle_flag predicate for every backend |
| ADR-1196 | Dispatch the SpEED dense matrix product through bit-exact AVX2 / AVX-512 kernels |
| ADR-1205 | The ssimulacra2 FMA unification extends to the scalar fallback and every GPU host copy |
| ADR-1207 | A test gates every feature's score against the host instruction set |
| ADR-1208 | The ssimulacra2 edge-diff SIMD loops take their difference in double |
| ADR-1237 | CAMBI Anti-Dithering AVX2 Vectorization, SpEED SIMD QR Dispatch, and Threaded GPU Flush Alignment |
| ADR-1253 | Scalar references compute their fused multiply-add themselves |
| ADR-1254 | Wide vector register pressure is a Win64 correctness constraint, not a performance one |
| ADR-1256 | Dispatch CAMBI's spatial-mask row SIMD kernels only where they measurably beat scalar |
| ADR-1257 | Retire the Darwin three-tap integer-ADM DWT2 compatibility dispatch |
| ADR-1260 | Windows on ARM64 CPU build-and-test lane |
| ADR-1283 | The whole-tree clang-tidy ratchet gets an arm64 cross lane |
| ADR-1325 | Normalize integer ADM Barten weights with one exponent per scale |
| ADR-1393 | CAMBI's c-values walks clip the window to the frame at every edge |
| ADR-1402 | Integer ADM keeps the scale-0 masking centre tap in int32 and clamps the excess in int64 |
| ADR-1413 | Every integer ADM implementation bounds the enhancement gain with the scalar's truncated double product |
| ADR-1415 | every x86 SIMD library is built without FP contraction |
| ADR-1459 | SpEED's covariance kernels return the scalar kernel's bits; they vectorise across sums, not within one |
| ADR-1469 | the psnr_hvs SIMD butterfly is two functions, cut at the same statement in the AVX2 and the NEON twin |
| ADR-1473 | The x86 float ADM wavelet and CSF kernels return the scalar bits and are dispatched; the two reduction kernels are removed |
| ADR-1482 | integer adm on frames of 17 to 32 pixels reads inside the frame and rounds a zero shift with 0; upstream reads index -1 there, and scale 3 differs by up to 0.23 |
| ADR-1488 | the psnr_hvs masking threshold is the double root of a float product, as Netflix's source writes it; the double product of PR #552 is removed and the AVX2, NEON, CUDA, HIP and SYCL forms follow, bit for bit |
| ADR-1490 | Insert RC7 CPU capability source of truth and shift the later candidates |
| ADR-1500 | The NEON and SVE2 float_moment kernels add in the scalar's order and return its bits on every input and vector length |
| ADR-1561 | integer VIF converts its residual variance through vif_sv_sq(), which returns x86's value without the undefined double to int32_t conversion |
| ADR-1601 | integer VIF forms the denominator log argument sigma_nsq + sigma1_sq in uint32_t |
| ADR-1880 | Format envelope and device-targeted scoring in the 1.0.0 candidates |
| ADR-1917 | The integer ADM scale-0 horizontal and vertical weight limit is 43900, set by the CSF stage's 16-bit magnitude |
| ADR-2055 | The x86 AVX2 level requires FMA |
| ADR-2478 | Rust replaces the host-side C and C++ through 3.0, behind the unchanged C ABI |