Research-1475: the fork's integer ADM against Netflix master, before and after the quantisation step went back to float¶
- Date: 2026-10-02
- Companion ADR: ADR-1475
- Rows:
T-UPSTREAM-AB-SCORE-DELTA-2026-09-07,T-ADM-CSF-EXPONENT-NOT-UPSTREAM-2026-10-01
Question¶
Does the fork's CPU return Netflix master's bits for integer ADM and for the models that read it, and if not, which expression is responsible?
Method¶
- Netflix/vmaf
cea2b4d8and the fork, both built with GCC 16.2.1, meson release,-Denable_float=true, x86-64 without-march. Netflix compiles with-O3 -std=c11, the fork with-O3 -std=c23 -ffp-contract=off; both are ISO modes, and neitherlibvmaf.sohas a fused multiply-add outside a function named*_avx2/*_avx512. - One C program per tree, linked against that tree's
libvmaf.a, drives the public API (vmaf_use_feature(),vmaf_model_load_from_path(),vmaf_read_pictures()) and prints every value the feature collector holds with%.17g. Netflix's CLI prints six decimals, which hides this defect. - Dispatch:
cpumask63 (scalar), 48 (AVX2 only) and 0 (the host's AVX-512). - 31 fixtures: the Netflix 576x324 pair at 8, 10, 12 and 16 bits and as 10-bit 4:2:2, both 1920x1080 checkerboard pairs, Big Buck Bunny at 1920x1080 and 3840x2160 (48 frames each), Sparks, synthetic noise at 8 to 16 bits, a bright 16-bit frame, two gradients, 4:4:4 and 4:0:0 clips, and frames from 352x288 down to 8x8. The
admextractor withdebug=trueand 19 option variants; 25 model files, which are byte-identical in the two trees.
Result¶
Values identical to Netflix / values compared, with the largest absolute difference:
Before (39929960f) | After | |
|---|---|---|
integer_adm2, scalar | 52 / 658, 8.5e-8 on decoded pictures | 616 / 658 |
integer_adm2, AVX2 | 52 / 658 | 636 / 658 |
integer_adm2, AVX-512 | 52 / 658 | 632 / 658 |
integer_adm3, scalar | 4 / 658, 5.5e-8 | 616 / 658 |
integer_adm_scale1, integer_adm_scale2 | 40 and 118 / 658, 2.5e-7 | 658 / 658 |
vmaf_v0.6.1, AVX-512 | 32 / 504, 1.83e-5 | 504 / 504 |
vmaf_v0.6.1, AVX2 | 32 / 504, 1.83e-5 | 504 / 504 |
vmaf_v0.6.1, scalar | 32 / 504 | 498 / 504 |
vmaf_v0.6.1neg, AVX-512 | 18 / 504, 1.81e-5 | 504 / 504 |
vmaf_4k_v0.6.1, AVX-512 | 32 / 504, 1.49e-5 | 504 / 504 |
vmaf_b_v0.6.3 (all bootstrap outputs), AVX-512 | 418 / 6048, 2.11e-5 | 6048 / 6048 |
option variants of adm identical at scalar and AVX-512 | 5 of 19 | 8 of 19 |
What still differs after the change is deliberate, or Netflix's own code disagreeing with itself:
- The four noise fixtures (
integer_adm_scale0, and through itadm2andadm3): the fork keeps the scale-0 masking centre tap in int32 (ADR-1402). Netflix's scalar and vector paths differ from each other there by the same amounts (adm2up to 9.4e-5), which is why the scalar column has more differing frames than the vector ones and why six frames ofvmaf_v0.6.1remain at scalar dispatch. - Frames of 17 to 24 pixels (scale 3), where Netflix reads outside the band.
- Of the option variants:
adm_csf_mode=1(the Barten weights, ADR-1472), andadm_enhn_gain_limit=1.2at vector dispatch (ADR-1413); the other eight differ on the 8-bit noise fixture at scalar dispatch only.
The prediction is not involved: with identical features the two trees return identical scores for every model, the bootstrap models included.
The expression¶
dwt_quant_step() returns 2 a 10^(k temp^2) / amplitude. Netflix forms k * temp * temp in float; the fork formed it in double since PR #552. For the default viewing geometry the reciprocal weights compare as follows (the fork's first): scale 0 0x1.1cc772p-6 and 0x1.1cc774p-6, its diagonal 0x1.820d4ep-8 and 0x1.820d54p-8, scale 1 0x1.060508p-5 and 0x1.06050cp-5. core/test/test_integer_adm_quant_step.c holds the step's bits from a Netflix build for five geometries.
Twins and snapshots¶
adm_cuda(RTX 4090),adm_hip(gfx1036),adm_sycl(Arc A380): the parity tests assert equality and pass; the parity gate'sadmcell is compared at tolerance 0 on the Netflix pair, both checkerboard pairs and Big Buck Bunny 1920x1080.testdata/scores_cpu_{576,640,720,1080,4k}.json: a build of the previous master reproduces all five files value for value on clips made withtestdata/generate.sh's commands; with the change 59, 41, 44, 45 and 38 of 720 values move, by at most 2e-5. The committed values are those of the golden build profile (scripts/ci/setup-golden-build.sh) inside the dev container: GCC 15.2.0, glibc 2.43, Ubuntu 26.04, imagesha256:43ef1e32cb32b148a076ed6dff73b72d7a6566ca3882bf90954b8a34a74761fc. A host build of the same profile with GCC 16.2.1 and glibc 2.44 returns the same values at%.17g: 0 of 3600 per-frame values and no pooled value differ over the five files, so for this model the result does not depend on which of the two compilers and C libraries built it.- A build with Intel's compiler: the step has the same bits as under GCC and glibc (the test above passes in an icx build). One of 240 frames at the five snapshot resolutions then differs between an icx and a GCC build in
integer_adm_scale1(7.9e-8), because Intel'spowfrounds that frame's argument differently; with glibc'slibmloaded the icx-built library returns the GCC build's values.