ADR-1491: The CUDA, SYCL and HIP motion twins compute motion_five_frame_window: the frame two back on the device, the CPU's window function on the host¶
- Status: Accepted
- Date: 2026-10-02
- Deciders: lusoris
- Tags:
motion,cuda,sycl,hip,gpu-parity,exact-twins,upstream-port,rc3
Context¶
ADR-1478 ported Netflix's motion_five_frame_window to the CPU extractors motion and motion_v2 and left the GPU twins as they were: motion_cuda, motion_sycl and motion_hip declared the option default-only, the motion_v2 twins did not declare it, and model dispatch computed the motion feature on the CPU whenever the option was set (T-GPU-MOTION-FIVE-FRAME-WINDOW-2026-10-02). The maintainer's decision was to port the option on the CPU and on the twins, each twin returning the CPU's bits.
The twins of both extractors are declared exact (scripts/ci/exact_twins.d/). What the option changes is small on the device and larger on the host:
- Device. The SAD kernel already takes the two planes it differences as arguments. With the option one of them is the frame two back instead of the previous one, so the twin has to keep one more frame.
- Host.
motion2of framenbecomesmin(SAD[n-1], SAD[n+1]), with special cases at both ends, and frames 0 and 1 have no SAD. The twins ofmotionemitmotion2andmotion3frame by frame as SADs arrive, with their own copy of the three-frame arithmetic; the twins ofmotion_v2carry one copy each of the CPU's flush.
Decision¶
Each twin takes its SAD against the frame two back when the option is set, and derives motion2 and motion3 with the CPU's own function, vmaf_motion_window_flush() (core/src/feature/motion_window.h, ADR-1478), in flush().
- The frame two back. CUDA: the ring of raw planes has three slots instead of two; frame
nwrites slotn % ringand reads slot(n + 1) % ring. HIP: two kept planes instead of one; framenreads planen % depthand the copy behind the SAD overwrites it. SYCLmotion_v2_sycl: a ring of three, as on CUDA. SYCLmotion_sycl: its two planes get fixed roles (framen-2and framen-1) and two device copies after each frame's kernel advance them. No kernel changes on any backend. - The first two frames report a SAD of 0 from
collect()and launch no readback.motion_syclenqueues the kernel on every frame all the same and ignores the result of the first two: its work is recorded once into a command graph and replayed, so what a frame enqueues cannot depend on its index. - The window on the host. With the option, the twins of
motionstore the SAD score per frame and nothing else;flush()callsvmaf_motion_window_flush(). Their three-frame path is unchanged. The twins ofmotion_v2call the function for both windows, which removes their copies of the flush. - Option tables.
motion_cuda,motion_syclandmotion_hipdropVMAF_OPT_FLAG_DEFAULT_ONLYand the-ENOTSUP; the threemotion_v2twins declare the option as the CPU does. - Gate. Two cells,
motion_mffwandmotion_v2_mffw(the option with the moving average, the option set of the HFR models), exact for all three backends. - Metal. The Metal twins are not changed: they do not declare the option, so the CPU extractor keeps computing it there. Not run: no device.
motion_syclrefuses the option together with its ownmotion_add_uv(-ENOTSUP): the CPUmotionhas no chroma mode, so that combination has no reference to equal.
Statements of earlier ADRs this replaces: ADR-1478 decision 5 for the CUDA, SYCL and HIP twins (they no longer leave the option to the CPU; the Metal part stands), and ADR-0219's -ENOTSUP for motion_five_frame_window on those three twins.
Alternatives considered¶
| Option | Pros | Cons | Why not chosen |
|---|---|---|---|
| Keep the CPU fallback (the state after ADR-1478) | No twin code | The motion feature of the HFR models leaves the device on every GPU run; the maintainer asked for the twins | Not what was decided |
Extend each motion twin's frame-by-frame emission to the five-frame window | motion2 / motion3 stay available before the flush, as in the three-frame path | A third copy of the window per backend, with the end cases (SAD[3] for frame 2, the last frame, one- to three-frame inputs) to get right three times; the CPU itself has the scores only at the flush | The CPU's function at the flush: exact by construction |
Move the three-frame path of the motion twins onto the shared function too | One path per twin, less code | Changes when a GPU run's motion2 becomes available (frame by frame today, at the flush then), which callers that read scores before the flush would notice; a deduplication step, not part of this port | Left for RC5 (T-GPU-CUDA-HIP-DUPLICATED-KERNELS-2026-10-02) |
motion_sycl: a ring of three planes, as on CUDA | No extra copy per frame | The combined command graph is recorded for two slots and replayed; a ring of three does not fit two recordings, and tying the planes to the slot parity would silently turn the window into the three-frame one on a path where the slot does not alternate | Two fixed planes and two copies outside the graph |
motion_sycl: read and overwrite one plane per slot inside the recorded graph | No extra plane or copy | The copy would have to run after the kernel that reads the same plane, inside a replayed graph whose replay of copy nodes the runtime notes call unreliable | Copies as plain queue commands behind the replay's barrier |
Consequences¶
- Positive:
- A GPU run of an
_hfrmodel keeps its motion feature on the device, and--feature motion=motion_five_frame_window=trueruns the twin. - Every output equals the CPU's bit for bit. Device tests (
test_<backend>_motion_five_frame_window): six option sets on both extractors, sequences of 11, 1, 2 and 3 frames, 8 and 10 bits,==on every output of every frame, on an RTX 4090, an Arc A380 and a gfx1036. Clips at--precision max(Netflix 576x324 at 8 and 10 bits, a 1080p checkerboard pair, 50 frames of BBB 3840x2160; four option sets including an HFR model): 1352 of 1352 motion values identical on each of the three devices. - The twins of
motion_v2lose their copies of the flush (45 to 62 lines net each on CUDA, SYCL and HIP). - Parity gate:
motion_mffwandmotion_v2_mffwreport 0 against the CPU on CUDA, SYCL and HIP, on the Netflix pair (48 frames) and on 200 frames of BBB 3840x2160; the three-frame cellsmotion,motion_debugandmotion_v2stay at 0.motion_syclwas measured with the combined command graph and without (VMAF_SYCL_USE_GRAPH=1,VMAF_SYCL_NO_GRAPH=1). - Negative:
- One more luma plane on the device while the option is set (CUDA and
motion_v2_sycl: a third ring slot; HIP: a second kept plane).motion_syclkeeps its two planes and pays one more device copy of a luma plane per frame. - With the option the
motiontwins publishmotion2andmotion3at the flush, as the CPU does, not frame by frame. - Neutral / follow-ups:
- The Metal twins keep the CPU fallback until someone can measure a device implementation.
- No SYCL kernel was added or changed, so the scratch-memory ratchet (ADR-1395), the sub-group sizes (ADR-1468) and the ahead-of-time targets are untouched.
References¶
Q(popup answer of the maintainer, 2026-10-02): "Port now, CPU and twins (Recommended)".- ADR-1478 (the CPU port and the window function), ADR-0219 (its
-ENOTSUPfor the three twins ofmotionends here), ADR-1372, ADR-1371, ADR-1377 (the shared SAD kernels), ADR-1108 (themotion_v2twins' flush), ADR-1316, ADR-1428, ADR-1395, ADR-1468. - Netflix/vmaf
a2b59b77,a4a1492d.