ADR-1478: Port motion_five_frame_window from Netflix; the deferral of ADR-0337 ends for this option¶
- Status: Accepted (Superseded-in-part 2026-10-06 by ADR-1491 for decision 5 (GPU twins leave motion_five_frame_window to the CPU))
- Date: 2026-10-02
- Deciders: lusoris
- Tags:
upstream-port,motion,feature-extractor,models,picture-pool,gpu-parity,golden-gate,rc3
Context¶
Netflix/vmaf has a second temporal window for the integer motion feature (a2b59b77, 2026-05-08, kept through the rewrite a4a1492d): with motion_five_frame_window=true the SAD of frame n is taken against frame n-2 instead of n-1, and motion2 of frame n is the smaller of the SADs of frames n-1 and n+1. The four shipped models model/vmaf_v1.0.16_hfr/vmaf_v1.0.16_hfr_*.json set the option.
The fork declared the option and refused it: ADR-0337 made motion_v2 return -ENOTSUP at init() and ADR-0994 did the same for motion, because the frame n-2 reaches an extractor through a second field on VmafFeatureExtractor (prev_prev_ref) and through vmaf_read_pictures(), which the fork had decomposed (ADR-0152) and given its own picture-ownership rules (ADR-0778, ADR-1431). ADR-0337 deferred that plumbing "to a follow-up PR"; none came. The cost by October 2026: the four _hfr models could not be scored on the fork, and 13 Netflix golden tests carried a fork-added @unittest.skip (nine in python/test/feature_extractor_test.py, four in python/test/vmaf_v1_quality_runner_test.py).
The upstream parity audit of 2026-10-02 listed the option as a missing port, and the maintainer decided that inherited code follows Netflix's source, with deviations only by ADR, and that this option is ported now, on the CPU and on the GPU twins.
Upstream no longer has motion_v2: a4a1492d replaced integer_motion.c with the pipelined implementation and deleted integer_motion_v2.c. The fork keeps both extractors, with the same SAD pipeline.
Upstream also makes every run keep the reference picture of frame n-2, whether or not an extractor reads it, and raises its picture-pool sizes to match. A review of the first revision of this port, which did the same with a warning for small pools, found that a preallocated pool of three pictures, enough for a serial run until then, would stall on the third frame: a hang for an existing API caller that never sets the option.
Decision¶
The fork computes motion_five_frame_window.
motionruns upstream's statements.integer_motion.c::extract()takes the SAD againstfex->prev_prev_refwhen the option is set (min_idx = 2), and the-ENOTSUPguard is gone. The flush is upstream's, as before.motion_v2does the same, as upstream's lastinteger_motion_v2.c(a4a1492d^) did. The two extractors keep their two option tables (ADR-0337, alternative A1), but the derivation ofmotion2andmotion3from the SAD scores exists once:vmaf_motion_window_flush()ininteger_motion.c, declared incore/src/feature/motion_window.h, called by both.- The framework hands a PREV_REF extractor frame
n-1, and framen-2to an extractor that reads it.VmafFeatureExtractorgainsreads_prev_prev_ref(), whichmotionandmotion_v2answer with the option. The context keepsprev_prev_ref, and rotates the two pictures per frame as upstream does, only while a registered extractor answers true (admit_prev_prev_ref()incore/src/libvmaf.c); otherwise it keeps framen-1alone, exactly as before the port. This is a deliberate deviation from Netflix, which keeps framen-2in every run. It changes which pictures the context holds, not what any extractor computes; the proof is under Consequences. The fork's ownership rule stays: an extractor gets counted references of its own (fex_take_prev_refs()/fex_release_prev_ref()), where upstream copies the structs. - A pool that would stall is refused, never left to wait. While frame
n-2is kept, a preallocated pool needs at least four pictures (the two kept reference pictures and the current pair). Whichever comes second,vmaf_preallocate_pictures()or the registration of the extractor (vmaf_use_feature(),vmaf_use_features_from_model(), and the first-frame CPU fallback of ADR-1324), returns-EINVALwith one error line namingpic_cntand the minimum. Without such an extractor the pool behaves as before the port: no minimum, and the default pool staysn_threads * 2(upstream'sn_threads * 2 + 2while framen-2is kept). The CLI sizes its pool before it loads the models, so a serial run takes four pictures, one more than it needs without the window; a threaded run keeps its count, which is also upstream's. - GPU twins leave the option to the CPU until they have the window.
motion_cuda,motion_syclandmotion_hipmark the optionVMAF_OPT_FLAG_DEFAULT_ONLY(ADR-1316), so a model or a--feature motionthat sets it is computed by the CPU extractor on every backend; the Metal twin and the fourmotion_v2twins do not declare the option, with the same effect (ADR-1359). A twin named directly with the option still fails (-ENOTSUP, or "unknown option"). The device implementations follow per backend. - The 13 skip markers are removed. No assertion changes.
Statements of earlier ADRs this replaces: ADR-0337 §Decision, "motion_five_frame_window=true is rejected at init() with -ENOTSUP", and its deferral of the picture-pool plumbing; ADR-0994 decision 1 (the guard in integer_motion.c). The rest of both stands, including ADR-0337's choice of duplicate option tables. ADR-0219's -ENOTSUP for a GPU twin stands until that twin has the window.
Alternatives considered¶
| Option | Pros | Cons | Why not chosen |
|---|---|---|---|
| Keep the deferral | No framework change | Four shipped models unusable, 13 golden tests skipped, a known difference from Netflix with no reason behind it | Maintainer decision: port now |
The extractor keeps its own copy of frame n-2 | No change to vmaf_read_pictures() | With worker threads each worker has a private extractor and sees an arbitrary subset of frames, so no extractor can hold "the frame two back"; the framework is the only place that sees frames in order | Does not work with --threads |
Keep frame n-2 in every run, as upstream does, and warn about a pool below four (this ADR's first revision) | Upstream's code; no mechanism of the fork's own | A preallocated pool of three, which served a serial API caller until now, stalls on the third frame; a warning does not make that additive for an existing caller (HISS-14) | Frame n-2 only for an extractor that reads it, a pool too small for it refused with -EINVAL |
The same, as a breaking change (! and a Migration: footer) | Upstream's code | Every existing caller with a pool of three has to change for a picture it never uses | Not needed: the conditional window moves no score |
Port to motion only, motion_v2 keeps -ENOTSUP | Smaller diff | Two extractors with one option table and different behaviour for the same option | Both, through one window function |
A copy of the five-frame flush in motion_v2 and in every twin | Each file reads on its own | The fork already carries one copy of the three-frame flush per twin (T-GPU-CUDA-HIP-DUPLICATED-KERNELS-2026-10-02); a second window would double them | One definition, vmaf_motion_window_flush() |
| Implement the twins in the same pull request | One step | Three devices to verify before the CPU port, the models and the golden tests can land | CPU first with a correct fallback; twins stacked per backend |
Consequences¶
- Positive:
motionwith the option equals Netflix9e48141bbit for bit: 2232 harness runs (31 fixtures, 8 option sets, scalar / AVX2 / default dispatch, serial and 1 and 4 worker threads), 95 130 values at%.17g, 0 differences.- The four
vmaf_v1.0.16_hfr_*models score. Against Netflix9e48141bon eight clips, three dispatch modes and three thread modes (288 runs), per frame: their motion features are identical (22 032 of 22 032 values),admandcambitoo (58 752 of 58 752), and the model score differs on 5436 of 7344 frames by at most 2.5e-5 becausespeed_chromadiffers (by at most 3.5e-5), which the non-HFR models share and another pull request of the parity work owns. On the audit's converged tree (the fork with every other difference from Netflix undone) with this port applied, all 119 088 values of those 288 runs are identical, the model scores included. - Netflix golden gate (x86): 280 passed, 3 skipped (271 and 12 before); the v1 model file 9 passed, 0 skipped (5 and 4 before). On the aarch64 cross build under qemu-user the 13 tests pass as well.
- With
--backend cuda,syclorhipan_hfrmodel scores, its motion feature on the CPU. - Keeping frame
n-2only for a reader moves no score. With the conditional window, against Netflix9e48141b:motionwith the option and with the same eight option sets without it, NUM_FEATURES; the four_hfrmodels and the four SDRvmaf_v1.0.16models, NUM_MODELS; on the audit's converged tree, NUM_UND5. - Without an extractor that reads frame
n-2, a context holds the pictures it held before the port and a pool of three serves a whole sequence (test_pool_of_three_without_the_window, serial and with a worker). - Negative:
- With the option the context keeps one more reference picture alive (one frame of memory), as upstream does in every run.
- An API caller that registers a five-frame extractor (an
_hfrmodel) next to a preallocated pool below four pictures gets-EINVALand has to make the pool four. Callers that allocate each picture withvmaf_picture_alloc()(the FFmpeg filters) are not affected. The CLI preallocates four pictures in a serial run, one more than before. reads_prev_prev_ref()and the admission check are a mechanism upstream does not have; a sync keeps them (docs/development/rebase-sensitive-invariants.md).- Neutral / follow-ups:
- Device implementations of the window for the CUDA, SYCL and HIP twins of
motionandmotion_v2, each bit-identical to the CPU, with a parity gate cell that sets the option (T-GPU-MOTION-FIVE-FRAME-WINDOW-2026-10-02). The Metal twins keep the CPU fallback: no device to measure on. - Whether the fork keeps
motion_v2as a second extractor at all, now that upstream has one, is a deduplication question (RC5,T-MOTION-V2-SECOND-EXTRACTOR-2026-10-03) and is not decided here. What this ADR settles of ADR-0337's duplicate-surface concern is the behaviour: one window function. The two option tables remain. - Upstream's own model documentation describes the window as "frames i-2, i, i+2"; the code, and the golden assertions that pin it, read frames
n-3,n-1andn+1formotion2of framen. The fork follows the code and documents what it computes (docs/metrics/motion.md).
References¶
Q(popup answer of the maintainer, 2026-10-02): "Port now, CPU and twins (Recommended)".req(popup answer, 2026-10-02): "Netflix's source, deviations only by ADR".- Review of #1887 before merge (2026-10-03): the unconditional window of the first revision made a preallocated pool of three a hang for existing API callers; this revision keeps frame
n-2only for a reader and refuses a pool too small for it. - Netflix/vmaf
a2b59b77(libvmaf/motion_v2: add motion_five_frame_window),4e469601,a4a1492d(libvmaf: replace integer_motion with pipelined v2 variant, rename); compared against master9e48141b. - ADR-0337, ADR-0994, ADR-0219, ADR-0152, ADR-0778, ADR-1431, ADR-1316, ADR-1359, ADR-0024, ADR-1317, ADR-1461.