Skip to content

Research-2045: integer VIF AVX2 stages — 2026-09-08

Finding and scope

The configured release CPU profile reports 17 clang-tidy warnings in vif_avx2.c, including four oversized statistic/subsample functions, missing own-header declarations, macro parentheses and offset widening. Cppcheck also reports dead initial product stores and read-only aliases. The adjacent integer_adm.h uses a quoted system stdio.h include; its one-line correction changes no numeric content.

This is an internal implementation of the existing touched-file policy (ADR-0141/1142). It introduces no public API, option, dispatch policy or numerical contract. The human-facing surface documentation rule and FFmpeg patch update are exempt because there is no user-visible surface change.

Decision and alternatives

Choice Result
Keep large kernels and add size suppressions Retains avoidable debt; rejected.
Reimplement filters or widen shared callback signatures Changes unrelated arithmetic/dispatch contracts; rejected.
Extract private forced-inline stages Selected; preserve each output's operations, rounding and lane layout.

The original packed 64-bit mean additions over 32-bit products remain. Eight-bit moments retain their shuffled mean products and final 0xD8 shuffle; 16-bit moments retain their separate packing before subtraction. Vertical 8-bit filtering keeps the center followed by symmetric tap pairs. Vertical 16-bit filtering keeps its left-to-right taps, unsigned products, rounding and shifts. Subsampling retains its scale+1 coefficient table, scale-dependent normalization, scalar tails and final copy/padding order. No scalar/SIMD boundary moves.

Two exact Cppcheck constParameterPointer exceptions preserve the statistic callback signatures: integer_vif.c assigns scalar, AVX2, AVX-512 and NEON functions to the same VifState function-pointer types taking VifPublicState *. Read-only implementation aliases are qualified without changing that dispatch contract. The obsolete dead-store NOLINT markers and only the overwritten initial product stores are removed.

Reproducer and numerical checks

With the normal CPU build configured with assembly enabled:

meson test -C build --print-errorlogs test_integer_vif_avx2_stages \
  test_integer_vif_log2 test_vif_skip_scale0
python3 scripts/ci/tidy-ratchet.py --lane cpu --build-dir build \
  --only core/src/feature/x86/vif_avx2.c \
  --only core/test/test_integer_vif_avx2_stages.c --write

The new native test calls the actual scalar and AVX2 statistic functions and compares both numerator/denominator encodings and complete fixture storage, including five final vertical planes and reflected padding. Its 864 cases cover all integer bit depths from 8 through 16, four scales, two patterns and twelve widths spanning short, full-vector and tail paths. Missing AVX2 produces an explicit skip; the retained local run executes AVX2 through the shared ADR-0245 test gate. Private object linkage keeps the test available in shared-library configurations without exporting internal symbols.

An independent retained old/new public-API probe runs the unchanged integer vif extractor with debug output at %.17g: 504 contexts, 1,008 frames and 15,120 finite feature values are bit-identical. It covers 8/10/12/16-bit, 14 geometries from 17×17 through 1920×1080, three input patterns, gain caps 1/2/100 and identity/distorted frames. No golden assertions or thresholds change.

A separate direct old/new probe instruments both VIF translation units and its fixture with ASan/UBSan. All 1,008 statistic and 756 subsample comparisons preserve result and workspace bytes. Linked unchanged library dependencies are not sanitizer-instrumented; this is a scoped check, not full-tree sanitizer acceptance.

Validation limits and evidence

Clang-tidy 22.1.8 measures zero warnings and zero uncited markers in the production TU and new test. Its scoped writer tightens the production count from 17 to zero and preserves unmeasured entries. Exhaustive Cppcheck finds no actionable diagnostics in the touched source, new test and the ADM header's actual consumer. The focused Cppcheck project retains ten unused-function diagnostics from unchanged included headers (including unused shared SIMD-harness helpers); a clean touched file is not a full-tree lint pass. The baseline is generated by the committed scoped writer, with unmeasured entries preserved.

Raw commands, hashes, old/new scores, stage results and the bounded warmed 1080p timing sanity check are retained under .workingdir2/evidence/vif-simd-native-lint-2026-09-08/. The initial timing check identified a repeatable 8-bit regression. Bisecting the extractions isolated GCC's full unrolling of the small inlined 8-bit second-moment loop: the statistic grew from about 5.6 KB/81 YMM stack references to 11 KB/454. Immediate result stores and buffer snapshots alone did not recover it. A GCC-only #pragma GCC unroll 1 keeps that one tap loop compact while all stages remain inline. This is a scheduling constraint, not a diagnostic or numerical exemption; GCC documents that unroll factors 0/1 block unrolling.

Seven alternating warmed 1080p rounds on CPU 24 time only feature extraction; picture generation/allocation are outside the measured interval. Four warmup frames precede twelve measured frames per depth. Final median new/old ratios for 8/10/12/16-bit are 0.9958/0.9922/0.9965/1.0010. Raw variation includes 8-bit old 24.29–25.02 ms versus new 24.37–26.90 ms per frame and a 12-bit new 37.33 ms outlier. This bounded check recovered the repeatable slowdown; it does not establish a speedup or general performance guarantee.

The numeric negative control changes one vertical coefficient and is rejected by the bitwise probe. A one-past-workspace load produces an ASan heap-buffer-overflow, confirming that the bounds control is active. No Windows/GPU execution, full-tree lint, Netflix gate or RC1-ready claim follows from these focused checks.