Motion¶
Motion measures temporal activity in the reference stream by computing the mean absolute difference (MAD) between consecutive Gaussian-blurred luma frames. It is used as a VMAF model feature to weight distortion scores: scenes with high motion are treated differently from static content. A frozen reference scores 0; larger values mean more temporal activity.
Three extractor variants are registered:
| Extractor name | Algorithm | Temporal? | Output keys in the JSON log |
|---|---|---|---|
motion | Integer fixed-point (Motion2) | Yes | integer_motion2, integer_motion3, VMAF_integer_feature_motion_sad_score (integer_motion with debug=true) |
motion_v2 | Integer pipelined (Motion2 v2) | Yes | VMAF_integer_feature_motion_v2_sad_score, VMAF_integer_feature_motion2_v2_score, VMAF_integer_feature_motion3_v2_score |
float_motion | Floating-point (Motion2) | Yes | motion2, motion3, motion (the last is on by default) |
How to run¶
Options follow the feature name after the first =, separated by :.
# Integer motion (default)
build/tools/vmaf \
--reference ref.yuv --distorted dist.yuv \
--width 1920 --height 1080 --pixel_format 420 --bitdepth 8 \
--no_prediction --feature motion --output /dev/stdout
# Integer motion with an FPS weight
build/tools/vmaf \
--reference ref.yuv --distorted dist.yuv \
--width 1920 --height 1080 --pixel_format 420 --bitdepth 8 \
--no_prediction --feature motion=motion_fps_weight=0.8 \
--output /dev/stdout
# Float motion with the 3-tap filter
build/tools/vmaf \
--reference ref.yuv --distorted dist.yuv \
--width 1920 --height 1080 --pixel_format 420 --bitdepth 8 \
--no_prediction --feature float_motion=motion_filter_size=3 \
--output /dev/stdout
An option with a non-default value is appended to the output key as its alias and value, for example integer_motion2_mfw_0.8 for motion_fps_weight=0.8. The CLI also accepts integer_motion and integer_motion2 as aliases of motion and motion_v2.
motion extractor (integer fixed-point)¶
Registered name: motion (VmafFeatureExtractor vmaf_fex_integer_motion).
The extractor blurs the difference of two reference luma frames with a separable 5-tap Gaussian filter and sums the absolute values (the frame's SAD). motion2 of a frame is the smaller of the SADs on both sides of it. By default it looks one frame back and one frame ahead (the three-frame window); with motion_five_frame_window it reads three frames back and one ahead (Five-frame window).
Output features¶
| Feature name | Description | Condition |
|---|---|---|
VMAF_integer_feature_motion2_score | Motion2 score (shipped VMAF model input) | Always |
VMAF_integer_feature_motion3_score | Perceptually blended motion score | Always |
VMAF_integer_feature_motion_score | Raw (unfixed) motion score for back-compat | debug=true only |
VMAF_integer_feature_motion_sad_score | The frame's SAD score: weighted by motion_fps_weight, capped at motion_max_val; the value motion2 and motion3 are derived from | Always |
Frame 0 always emits motion2_score = 0.0. The output range is [0, motion_max_val] (default 10 000), with no inherent upper bound; larger values indicate more temporal activity.
When motion2 and motion3 are final¶
motion2 of a frame reads the SAD of the frame after it, so neither score of frame i exists when frame i is read. The engine writes both as soon as the SAD scores of frames 0 to max(i + 1, min_idx) are in, where min_idx is 1, or 2 with motion_five_frame_window (ADR-2090):
- without worker threads, by the time
vmaf_read_pictures()of framei + 1returns; - with worker threads, in the first
vmaf_read_pictures()after the workers finished framei + 1, or in a read of the score (vmaf_feature_score_at_index(),vmaf_score_at_index(),vmaf_score_pooled()), which waits for the workers; a VMAFx context with an open window derives them as soon as the worker finishes the frame (window scores); - on a device backend, when the twin's SAD of frame
i + 1is collected: onevmaf_read_pictures()later, and formotion_cudaat the end of its readback batch of eight frames; - the last frame has no frame after it: its scores are written by the flush,
vmaf_read_pictures(vmaf, NULL, NULL, 0).
The values are the ones a flush over the whole stream computes, bit for bit: the derivation runs upstream's statements frame by frame in index order and carries the moving average from one frame to the next. A per-frame model score (a metadata callback, a VMAFx window over a VMAF model) is therefore available one frame after its frame, not only after the flush. motion_v2 and every GPU twin of motion and motion_v2 follow the same rule; float_motion always has (it writes frame i - 1 while it reads frame i).
Options¶
| Option | Alias | Type | Default | Range | Effect |
|---|---|---|---|---|---|
debug | none | bool | false | n/a | Emit motion_score (legacy unfixed variant) alongside motion2_score; motion_cuda, motion_sycl and motion_hip default to false too |
motion_force_zero | force_0 | bool | false | n/a | Override all emitted scores to 0.0; used for deterministic fixtures |
motion_fps_weight | mfw | double | 1.0 | 0.0–5.0 | Multiplicative FPS-aware correction applied before clamping |
motion_blend_factor | mbf | double | 1.0 | 0.0–1.0 | Blend factor for motion3_score |
motion_blend_offset | mbo | double | 40.0 | 0.0–1000.0 | Score offset at which blending begins for motion3_score |
motion_max_val | mmxv | double | 10000.0 | 0.0–10000.0 | Upper clamp applied to emitted scores |
motion_five_frame_window | mffw | bool | false | n/a | Take each SAD against the frame two back instead of the previous one (Five-frame window); also on motion_cuda, motion_sycl and motion_hip |
motion_moving_average | mma | bool | false | n/a | Apply a two-frame moving average to motion3_score |
motion_fps_weight is applied exactly once
The weight scales the SAD-derived score once, in extract(), on the CPU and on every GPU twin. Up to and including v3.2.1 the CUDA, SYCL and HIP twins applied it again, so motion3_score carried the weight squared.
The CPU reference scales the SAD-derived score by motion_fps_weight in extract(), stores the weighted value as motion_sad_score, and blends that already-weighted value into motion2_score / motion3_score without touching the weight again. The old GPU behaviour showed whenever the option was set away from 1.0 (motion2_score was always correct). Fixed per ADR-1216; the test_<backend>_motion3_parity tests now pin motion_fps_weight = 0.6 and assert CPU/GPU parity on the derived integer_motion3_mfw_0.6 key, so the default value can no longer hide a squared weight.
Backends¶
| Backend | Extractor | Source | Arithmetic vs CPU | Evidence (fragments in scripts/ci/exact_twins.d/) |
|---|---|---|---|---|
| Scalar C | motion | integer_motion.c | Reference | n/a |
| AVX2, AVX-512, NEON | motion | x86/motion_avx2.c, x86/motion_avx512.c, arm64/motion_neon.c | Same bits | n/a |
| CUDA | motion_cuda | feature/cuda/integer_motion_cuda.c | Exact (ADR-1372, ADR-1457) | motion.cuda, motion_debug.cuda, motion_mffw.cuda |
| SYCL | motion_sycl | feature/sycl/integer_motion_sycl.cpp | Exact (ADR-1371, ADR-1451) | motion.sycl, motion_debug.sycl, motion_mffw.sycl |
| HIP | motion_hip | feature/hip/integer_motion_hip.c | Exact (ADR-1377, ADR-1437) | motion.hip, motion_debug.hip, motion_mffw.hip |
| Metal | integer_motion_metal | feature/metal/integer_motion_metal.mm | Not exact (see below) | none |
How the twins match the CPU:
- CPU. It computes the SAD of the blurred difference of the two frames, rounding after the vertical and after the horizontal filter pass.
- CUDA, SYCL, HIP. Each does the same. The CUDA twin runs the kernel
motion_v2_cudaalready used, and its debugmotionscore is the CPU's (weighted bymotion_fps_weight, capped atmotion_max_val). - Metal. The twin still blurs each frame and compares the blurred frames, which rounds differently. The SYCL twin did the same until 2026-09-29 and its
motion2_scorewas up to 2.0e-4 off on 17x17 frames and 1.3e-5 on the Netflix 576x324 pair; expect the same from the Metal twin, which has not been measured yet (its row instate.md).
What the twins emit:
| Backend | motion2, motion3 (3-frame window) | motion_sad_score | motion_force_zero | Five-frame window |
|---|---|---|---|---|
| CUDA | Yes | Yes, since 2026-10-02 | Yes | Yes (ADR-1491) |
| SYCL | Yes | Yes, since 2026-10-02 | Yes, since 2026-10-02 | Yes (ADR-1491) |
| HIP | Yes | Yes | Yes | Yes (ADR-1491) |
| Metal | motion and motion2 only (no motion3) | No (T-GPU-MOTION-SAD-SCORE-NOT-EMITTED-2026-10-02 in state.md) | not recorded | No: the CPU extractor computes it |
The parity gate compares the SAD score in its motion and motion_debug cells, so a twin without it fails the cell.
motion_sad_score was missing on CUDA and SYCL before 2026-10-02
A --backend cuda or --backend sycl run lacked VMAF_integer_feature_motion_sad_score before that date (T-CUDA-MOTION-SAD-SCORE-NOT-EMITTED-2026-10-02).
motion_force_zero was ignored on motion_sycl before 2026-10-02
The twin declared the option and ignored it: with motion_force_zero=true it returned the measured motion2 / motion3 under the _force_0 names where the CPU returns 0 (T-SYCL-MOTION-FORCE-ZERO-IGNORED-2026-10-02). That reached a shipped model: model/other_models/vmaf_v0.6.1mfz.json sets the option, and on --backend sycl it scored the Netflix 576x324 pair 76.668 where the CPU scores 72.321. Both give 72.321 now.
motion_v2 extractor (pipelined integer)¶
Registered name: motion_v2 (VmafFeatureExtractor vmaf_fex_integer_motion_v2).
A pipelined re-implementation that exploits the linearity of the blur kernel:
The frame difference, blur, and absolute-sum are fused into a single row-at-a-time pipeline requiring only one scratch row. Per-frame blurred-state storage is eliminated.
Output features¶
| Feature name | Description | Condition |
|---|---|---|
VMAF_integer_feature_motion_v2_sad_score | Per-frame sum of absolute blurred differences | Always |
VMAF_integer_feature_motion2_v2_score | Motion2-equivalent smoothed score | Always |
VMAF_integer_feature_motion3_v2_score | Perceptually blended + clipped score (optional 2-frame moving average) | Always |
motion3_v2_score is the motion_v2 analogue of motion v1's motion3_score: a per-frame motion_blend(motion2, motion_blend_factor, motion_blend_offset) clipped to motion_max_val, optionally smoothed by a two-frame moving average (motion_moving_average=true). It is written host-side as soon as the frame after it is scored, the last frame's at the end-of-stream flush (When motion2 and motion3 are final). The output range is [0, motion_max_val], in the same units as motion.
Options¶
| Option | Alias | Type | Default | Range | Effect |
|---|---|---|---|---|---|
motion_force_zero | force_0 | bool | false | n/a | Override all scores to 0.0 |
motion_fps_weight | mfw | double | 1.0 | 0.0–5.0 | FPS-aware multiplicative correction |
motion_blend_factor | mbf | double | 1.0 | 0.0–1.0 | Blend factor for motion3-style score |
motion_blend_offset | mbo | double | 40.0 | 0.0–1000.0 | Blend offset |
motion_max_val | mmxv | double | 10000.0 | 0.0–10000.0 | Upper clamp |
motion_five_frame_window | mffw | bool | false | n/a | SAD against the frame two back, as on motion (Five-frame window); also on motion_v2_cuda, motion_v2_sycl and motion_v2_hip (ADR-1491), on the CPU for --backend metal |
motion_moving_average | mma | bool | false | n/a | Two-frame moving average |
Backends¶
| Backend | Extractor | Source | vs CPU | Evidence (fragments) |
|---|---|---|---|---|
| Scalar C | motion_v2 | integer_motion_v2.c | Reference | n/a |
| AVX2, AVX-512, NEON | motion_v2 | x86/motion_v2_avx2.c, x86/motion_v2_avx512.c, arm64/motion_v2_neon.c (ADR-0145, bit-exact) | Same bits | n/a |
| CUDA | motion_v2_cuda | feature/cuda/integer_motion_v2_cuda.c | Exact (ADR-1372, ADR-1457) | motion_v2.cuda, motion_v2_mffw.cuda |
| SYCL | motion_v2_sycl | feature/sycl/integer_motion_v2_sycl.cpp | Exact (ADR-1371, ADR-1451) | motion_v2.sycl, motion_v2_mffw.sycl |
| HIP | motion_v2_hip | feature/hip/integer_motion_v2_hip.c | Exact (ADR-1377, ADR-1437) | motion_v2.hip, motion_v2_mffw.hip |
| Metal | motion_v2_metal | feature/metal/integer_motion_v2_metal.mm | Bit-exact on the cross-backend gate fixture | none |
All four GPU twins emit motion3_v2_score and accept the motion_blend_factor, motion_blend_offset, motion_max_val and motion_moving_average options (ADR-1108):
- Where it runs. The score is computed host-side over the kernel's SAD scores via the shared
motion_blend_tools.hhelper. The CUDA twin landed first (#909); the SYCL, HIP and Metal twins mirror itsflush_fexpost-process byte for byte. - Parity. It is bit-exact at default options (
max_abs_diff = 0.0on the Netflixsrc01576x324 pair, 48 frames), verified by the per-backendtest_<backend>_motion_v2_paritytests, which assertmotion_v2_sad,motion2_v2andmotion3_v2atplaces=4and skip cleanly when the backend device is absent. - Mirror rule. The host-side post-process is identical across all GPU twins, so any change to the CPU
motion_v2flush blend, clip, seed or moving-average logic must be mirrored into all four in the same PR. - Edge formula. All GPU kernels use the CPU
integer_motion_v2.c::mirrorhigh-edge formula (\(2\,\mathrm{size} - \mathrm{idx} - 2\)).
motion_fps_weight and motion_max_val on motion_v2
The CPU reference stores MIN(sad * motion_fps_weight, motion_max_val) as motion_v2_sad_score and derives motion2_v2 / motion3_v2 from it; a one-frame input still gets motion2_v2 = motion3_v2 = 0. motion_v2_cuda and motion_v2_hip do the same (ADR-1373, ADR-1382), and so does motion_v2_sycl. The Metal twin stores the raw SAD, applies motion_fps_weight in the host-side flush without the motion_max_val cap on motion2_v2, and emits no motion2_v2 / motion3_v2 for a one-frame input. The paths agree at the default motion_fps_weight = 1.0 and motion_max_val = 10000 (T-GPU-TWIN-PARITY-GAPS-OUTSIDE-CUDA-2026-09-30 in state.md).
float_motion extractor (floating-point)¶
Registered name: float_motion (VmafFeatureExtractor vmaf_fex_float_motion).
Floating-point twin of motion using float arithmetic throughout. It provides additional options for chroma channels and a half-resolution scale-1 SAD term. The output range is [0, motion_max_val], with the same semantics as motion.
Output features¶
| Feature name | Description | Condition |
|---|---|---|
VMAF_feature_motion2_score | Motion2 score | Always |
VMAF_feature_motion3_score | Perceptually blended score | Always |
VMAF_feature_motion_score | Raw (unfixed) motion score | debug=true (the default) |
Options¶
| Option | Alias | Type | Default | Range | Effect |
|---|---|---|---|---|---|
debug | none | bool | true | n/a | Emit motion_score alongside motion2_score |
motion_force_zero | force_0 | bool | false | n/a | Override all scores to 0.0 |
motion_fps_weight | mfw | double | 1.0 | 0.0–5.0 | FPS-aware multiplicative correction |
motion_blend_factor | mbf | double | 1.0 | 0.0–1.0 | Blend factor for motion3_score |
motion_blend_offset | mbo | double | 40.0 | 0.0–1000.0 | Blend offset for motion3_score |
motion_add_scale1 | mdc | bool | false | n/a | Add half-resolution SAD term on top of the full-resolution SAD of each plane (CPU and float_motion_hip) |
motion_add_uv | mau | bool | false | n/a | Sum U and V plane SADs into the score (CPU and float_motion_hip) |
motion_filter_size | mfs | int | 5 | 0–9 | Blur filter: 3 = 3-tap, 1 = no blur, any other value = the 5-tap Motion2 filter (CPU and float_motion_hip) |
motion_max_val | mmxv | double | 10000.0 | 0.0–10000.0 | Upper clamp applied to emitted scores |
motion_add_scale1, motion_add_uv, motion_filter_size and motion_max_val were ported from Netflix/vmaf commit b949cebf. With default settings the output is bit-identical to the pre-port baseline on the Y-plane SIMD fast path. motion_add_uv needs a pixel format with chroma planes; 4:0:0 input is refused at init.
Backends¶
| Backend | Extractor | Source | vs CPU | motion3 | Options accepted | Evidence (fragment) |
|---|---|---|---|---|---|---|
| Scalar C, AVX2, AVX-512, NEON | float_motion | float_motion.c, x86/float_motion_avx2.c, x86/float_motion_avx512.c, arm64/float_motion_neon.c | Reference | Yes | All | n/a |
| CUDA | float_motion_cuda | feature/cuda/float_motion_cuda.c (ADR-0196) | Exact (ADR-1409) | Yes | debug, motion_force_zero, motion_fps_weight, motion_max_val, motion_blend_factor, motion_blend_offset | float_motion.cuda |
| SYCL | float_motion_sycl | feature/sycl/float_motion_sycl.cpp (ADR-0196) | Exact (ADR-1411) | Yes, since 2026-10-03 | As CUDA | float_motion.sycl |
| HIP | float_motion_hip | feature/hip/float_motion_hip.c (ADR-0273) | Exact (ADR-1419) | Yes (ADR-1404) | The whole CPU table, including motion_add_scale1, motion_add_uv, motion_filter_size | float_motion.hip |
| Metal | float_motion_metal | feature/metal/float_motion_metal.mm | Not exact (per-block sum) | No (motion and motion2 only; T-GPU-FLOAT-MOTION3-MISSING-2026-09-30 in state.md) | Not motion_max_val; see below | none |
The CUDA, SYCL and Metal twins do not declare motion_add_scale1, motion_add_uv or motion_filter_size, so a request with one of them is computed on the CPU (ADR-1183). On Metal a motion_max_val setting keeps float_motion on the CPU as well, and the debug motion score is emitted without the fps weight.
Why the float twins can be exact¶
The score depends on the order of the additions: the CPU adds the absolute differences of a row into one float, the row sums into a second one, and divides in float, and those running sums round at every step. A twin that sums per block instead is closer to the exact mean and differs from the CPU in the low digits.
- CUDA (ADR-1409). The twin adds each row on the device in the CPU's order and the rows on the host. It returns
motion,motion2andmotion3bit for bit at--precision max, at every frame size and sample depth. Before the change it was off by 3e-6 on the Netflix 576x324 pair, 2.4e-5 at 3840x2160 and 1.4e-4 on 1920x1080 checkerboards. - SYCL (ADR-1411). One work-item per row adds the row on the device and the host adds the rows.
motionandmotion2equal the CPU's on every frame (measured on an Arc A380: the Netflix pair, both 1080p checkerboard pairs, 200 frames of BBB 3840x2160, and the Netflix pair at 10, 12 and 16 bits). The second pass over the blurred planes costs 0.38 ms per 3840x2160 frame on that GPU (3.85 to 4.23 ms). - HIP (ADR-1419). The twin stores its differences transposed and adds each row in the CPU's order, so it is exact too (
float_motion.hip: 1617 of 1617 values over seven option sets on a gfx1036).
motion3 and the options on the twins¶
- CUDA.
motion3is the CPU's blend ofmotion2(motion_fps_weight, then themotion_blend_factor/motion_blend_offsetblend, then themotion_max_valcap), with frame 0 taken from the first SAD and the last frame from the flush. It takes both blend options (aliasesmbf/mbo). - HIP. The same, and the whole CPU option table (ADR-1404): the filter is a kernel argument, the scale-1 term a second kernel, and
motion_add_uvruns both on the U and V planes. - SYCL.
motion3and both blend options since 2026-10-03. Itsmotion3is the CPU's bit for bit, because its SAD already is (ADR-1411) and the blend is the CPU's host arithmetic. motion_max_valandmotion_fps_weight.float_motion_sycl,float_motion_cudaandfloat_motion_hiptakemotion_max_val(aliasmmxv) and, like the CPU, scale every score they emit (the debugmotiontoo) bymotion_fps_weightbefore capping it atmotion_max_val(ADR-1365, ADR-1373, ADR-1382).motion_force_zero.float_motion_cudaandmotion_cudapublish zeros from the first frame;motion_hipandfloat_motion_hipdo as well.
Measured for the SYCL motion3 on an Arc A380 at --precision max against --backend cpu: every motion, motion2 and motion3 value is identical on the Netflix 576x324 pair (48 frames), both 1920x1080 checkerboard pairs and 200 frames of BBB 3840x2160. That holds at the default options and with motion_blend_factor=0.5:motion_blend_offset=2, and on the small fixtures also with motion_fps_weight=2:motion_max_val=4, with all four score options together, with motion_force_zero and for a one-frame input (motion3 = 0).
# On float_motion_sycl with a motion cap
# (keys become motion_mmxv_4 / motion2_mmxv_4 / motion3_mmxv_4)
vmaf --reference ref.yuv --distorted dist.yuv \
--width 576 --height 324 --pixel_format 420 --bitdepth 8 \
--backend sycl --no_prediction \
--feature float_motion=motion_max_val=4 --output /dev/stdout
# On float_motion_sycl with the motion3 blend
# (keys become motion_mbf_0.5_mbo_2 / motion2_mbf_0.5_mbo_2 / motion3_mbf_0.5_mbo_2)
vmaf --reference ref.yuv --distorted dist.yuv \
--width 576 --height 324 --pixel_format 420 --bitdepth 8 \
--backend sycl --no_prediction \
--feature float_motion=motion_blend_factor=0.5:motion_blend_offset=2 \
--output /dev/stdout
Five-frame window¶
motion_five_frame_window=true is the motion feature of the four vmaf_v1.0.16_hfr_* models (VMAF v1 models): vmaf_v1.0.16_hfr_3d0h, vmaf_v1.0.16_hfr_3d0h_2160, vmaf_v1.0.16_hfr_5d0h and vmaf_v1.0.16_hfr_1d5h_2160. They set it together with motion_moving_average=true. It is Netflix's option (a2b59b77, a4a1492d), ported with its arithmetic unchanged (ADR-1478); the scores equal upstream Netflix 9e48141b bit for bit. No other shipped model sets it.
What it changes¶
With the option, for a sequence of N frames:
| Score | Three-frame window (default) | Five-frame window |
|---|---|---|
SAD of frame n | frames n-1 and n; 0 for n = 0 | frames n-2 and n; 0 for n < 2 |
motion2 of frame n | min(SAD[n], SAD[n+1]); SAD[n] for the last frame; 0 for n = 0 | min(SAD[n-1], SAD[n+1]); SAD[3] for n = 2; SAD[n] for the last frame; 0 for n < 2 |
motion3 of the frames without a SAD | the blended SAD[1] | the blended SAD[2] (frames 0 and 1) |
So motion2 of frame n reads the frames n-3, n-1 and n+1, a span of five frames (it does not read frame n itself). motion3 is derived from motion2 as in the default mode (motion_blend_factor, motion_blend_offset, motion_max_val, motion_moving_average). A sequence of one or two frames has no SAD and every score is 0.
How to run¶
build/tools/vmaf \
--reference ref.yuv --distorted dist.yuv \
--width 1920 --height 1080 --pixel_format 420 --bitdepth 8 \
--no_prediction \
--feature motion=motion_five_frame_window=true:motion_moving_average=true \
--output /dev/stdout
The scores are written under names that carry the option aliases: integer_motion2_mffw_mma, integer_motion3_mffw_mma and VMAF_integer_feature_motion_sad_score_mffw_mma.
What the framework keeps¶
For the extra frame the framework keeps the reference pictures of the two frames before the current one (prev_ref, prev_prev_ref):
- When. Only while an extractor with the option is registered; otherwise it keeps the frame before the current one alone, as without the port. Netflix keeps both in every run; the fork's narrower rule changes which pictures stay allocated, not a score (ADR-1478).
- Pool size. A preallocated picture pool then needs at least four pictures: the current pair and those two.
- Error. A smaller pool is refused with
-EINVALwhen the extractor is registered or the pool is allocated, whichever comes second, instead of stalling on the third frame (C API).
GPU twins¶
motion_cuda, motion_sycl and motion_hip compute the five-frame window (ADR-1491), and so do motion_v2_cuda, motion_v2_sycl and motion_v2_hip:
- They keep the frame two back on the device, take the SAD against it with the kernel of the three-frame window, and derive
motion2/motion3with the CPU's own function when the last frame is in. - With the option, a twin publishes
motion2andmotion3at the end of the run, as the CPU does, not frame by frame. - Their outputs equal the CPU's bit for bit (measured on an RTX 4090, an Arc A380 and a gfx1036): 104 of 104 frames (Netflix 576x324 at 8 and 10 bit, a 1080p checkerboard, BBB 4K 50 frames) on every backend, and the parity gate compares the cell
motion_mffw(andmotion_v2_mffw) with tolerance 0 (fragmentsmotion_mffw.{cuda,hip,sycl}andmotion_v2_mffw.{cuda,hip,sycl}). motion_metalandmotion_v2_metaldo not declare the option: on--backend metalthe CPU extractor computes the motion feature of a model or--feature motionthat sets it (not measured, no device).
Input format constraints¶
All three extractors:
- Accept YUV 4:2:0 / 4:2:2 / 4:4:4, 8 / 10 / 12 / 16 bpc.
- Operate on the Y (luma) plane only by default. Chroma is only included when
motion_add_uv=trueis set onfloat_motion, and only for formats other than YUV 4:0:0. - Require a minimum frame size of 3x3 pixels (5-tap Gaussian minimum dimension = filter_radius + 1 = 3). Smaller frames are rejected with
-EINVALatinit(). Withmotion_filter_size=3the minimum is 2x2. - Are temporal extractors: frame 0 always emits
0.0for all motion scores.
The minimum applies to every plane that is actually convolved, not only to luma. With motion_add_uv=true on float_motion, the U and V planes are blurred at their subsampled dimensions, so a 4x4 YUV 4:2:0 frame, whose luma clears the 3x3 floor, presents a 2x2 chroma plane and is rejected. In practice motion_add_uv=true needs at least 5x5 for 4:2:0 (chroma 3x3), 5x3 for 4:2:2, and 3x3 for 4:4:4. The error message names the plane that failed:
libvmaf ERROR float_motion: chroma plane 2x2 is below the 5-tap filter
minimum 3x3; refusing to avoid out-of-bounds mirror reads
History¶
- 2026-10-03.
float_motion_syclemitsmotion3and takesmotion_blend_factor/motion_blend_offset. Before, a--backend sycl --feature float_motionrun wrote nomotion3and gave no warning, and a request with a blend option ran on the CPU. - 2026-10-02.
motion_cudaandmotion_syclemitVMAF_integer_feature_motion_sad_score;motion_force_zeroworks onmotion_sycl(see the warnings under Backends). - 2026-09-30, ADR-1372, ADR-1377.
motion_cudaandmotion_hiptake the diff-first order. On an RTX 4090 the CUDA twin'smotion2_scoreandmotion3_scoreequal the CPU's bit for bit on the Netflix 576x324 pair and on 50 frames of a 3840x2160 clip, where the previous order was 1.26e-5 and 6.9e-5 off (T-CUDA-MOTION-BLUR-THEN-DIFF-2026-09-29instate.md). On a gfx1036 the HIP twin's scores equal the CPU's on every frame of the Netflix pair, where the previous order was 1.26e-5 off (T-HIP-MOTION-BLUR-THEN-DIFF-2026-09-29). - 2026-09-30,
motion_force_zerocrash. Before it,float_motion_cudaandmotion_cudacrashed onmotion_force_zero(T-GPU-MOTION-FORCE-ZERO-FIRST-FRAME-SEGV-2026-09-30), and so didmotion_hipandfloat_motion_hip(T-HIP-MOTION-FORCE-ZERO-NULL-SUBMIT-2026-09-30). - 2026-09-29, ADR-1371.
motion_sycltakes the diff-first order. - ADR-1166. Before it, the
float_motionguard validated luma only, and a sub-minimum chroma plane reached the convolution and read out of bounds (Netflix/vmaf#1582, Netflix/vmaf#1581). The same guard is now present on the three Metal twins, which had none at all (Netflix/vmaf#1580). - ADR-1216. GPU
motion3_scorecarriedmotion_fps_weightsquared up to and including v3.2.1 (see the warning under Options).
Rust implementation¶
The integer motion extractor has a Rust twin, motion_rust (core/src/rust/feature/motion), a port of integer_motion.c that returns the C extractor's scores bit for bit: VMAF_integer_feature_motion_sad_score, motion2 and motion3 (and motion with debug), with every option of the table, including motion_five_frame_window and motion_moving_average. It is built with -Denable_rust_features=true and selected with VMAF_FEATURE_IMPL=rust (every motion of the run, including the one a model asks for) or by name with --feature motion_rust; without either the C extractor runs. Like the C extractor it appends motion2 and motion3 of a frame as soon as the frame after it is scored (ADR-2090), so a window over a VMAF model completes before the end of the stream with the twin too. The JSON feature_backends entry names the twin that ran. Check the twin against the C extractor on your own clips with:
python3 scripts/ci/rust_twin_diff.py --vmaf build-rs/tools/vmaf \
--feature motion --fixtures netflix