VMAFx window scores¶
A window score is a pooled score over a range of frames that you ask for before the frames are final and read when they are: the per-window statistics of a live encode (n_stats, #2138), or a rolling score a live-production plugin shows without blocking its render thread (#2238). The calls are part of the VMAFx API preview and are declared in vmafx/score.h; the reference page lists every field. The design is recorded in ADR-2074.
Ask for a window¶
VmafxWindowRequest req = VMAFX_WINDOW_REQUEST_INIT;
req.target = VMAFX_WINDOW_TARGET_MODEL; /* or MODEL_SET, FEATURE */
req.model = model; /* the window holds a reference */
req.pool_mask = VMAFX_POOL_MASK_MIN | VMAFX_POOL_MASK_MAX |
VMAFX_POOL_MASK_MEAN | VMAFX_POOL_MASK_HARMONIC_MEAN;
req.first = 48; /* inclusive, as vmafx_score_pooled() */
req.last = 95;
VmafxWindow *window = NULL;
status = vmafx_window_submit(context, &req, &window, &error);
vmafx_window_submit() returns at once. Submit windows before their frames, while they arrive, or after the flush; a window whose frames are all final completes right after the submit. At most 1024 windows of one context are open at a time; one more is refused with VMAFX_E_BUSY naming context.
| Target | Set | Same values as |
|---|---|---|
VMAFX_WINDOW_TARGET_MODEL | model | vmafx_score_pooled() |
VMAFX_WINDOW_TARGET_MODEL_SET | model_set | vmafx_score_pooled_model_set() |
VMAFX_WINDOW_TARGET_FEATURE | feature (copied) | vmafx_feature_score_pooled() |
pool_mask has bit 1 << p for each VmafxPool method p (VMAFX_POOL_MASK_MIN to VMAFX_POOL_MASK_PERC20); one window computes all of them.
When a window completes¶
A window completes when every scored frame of [first, last] has its final score, and then it is never revised. Each context has a completion thread, started by its first window, that finds this out: the worker thread that finishes a frame wakes it, and so do vmafx_submit(), vmafx_flush(), vmafx_context_import_score() and vmafx_window_submit(). It sleeps until woken (no polling), looks only at frames it has not looked at before, and never waits for work still on a worker thread. A window therefore completes when its last frame is final, whether or not the thread that feeds the context calls again.
- Without worker threads (
n_threads0) a window over frames that are final per frame completes right after the submit of its last frame. - With worker threads it completes when the worker finishes its last frame, also while the producer is busy elsewhere or waiting: about 2 ms after the start of that frame's submit on the Netflix 576x324 pair.
- On a device backend a frame's scores are collected when the next frame is submitted (the engine double-buffers device extractors), one submit later.
- Features that read the next frame.
motion2andmotion3of a frame need the motion of the frame after it (and, withmotion_five_frame_window, frames 0 and 1 need frame 2). The engine derives them as soon as that frame is scored (ADR-2090), so a window over them, or over a VMAF model (every model readsmotion2), completes one frame after itslast: without worker threads right after the submit of framelast + 1; with worker threads when the worker finishes framelast + 1(the completion thread lets the engine derive them before it looks), also while the producer makes no further call; on a device backend one submit later, and withmotion_cudaat its readback batch of eight frames. The last frame of the stream has no frame after it: itsmotion2andmotion3become final at the flush, which completes the windows that end there. - At the flush every open window completes over the frames the stream had. A window whose
lastlies past the stream's last frame is flaggedVMAFX_WINDOW_PARTIALand poolsn_framesframes fromfirst; one that no frame of the stream reaches completes withVMAFX_E_RANGE. - When the context is destroyed every open window completes with
VMAFX_E_INVALID, and every callback runs beforevmafx_context_destroy()returns. Window handles stay valid until you release them.
The stream's last frame is the highest index submitted or imported with vmafx_context_import_score(), so a window over scores you import completes when its last score is imported, with or without a flush.
Read the result¶
Three ways, all thread-safe, and you may use them together:
| Call | Returns |
|---|---|
vmafx_window_poll(window, &result, &error) | VMAFX_OK and the result once complete, else VMAFX_PENDING without an error |
vmafx_window_wait(window, timeout_ns, &result, &error) | The same after waiting up to timeout_ns (UINT64_MAX: no limit) |
req.on_complete(window, &result, req.user) | Called once on the context's callback thread |
VmafxWindowResult r = VMAFX_WINDOW_RESULT_INIT;
if (vmafx_window_poll(window, &r, NULL) == VMAFX_OK && r.status == VMAFX_OK)
printf("frames %llu-%llu: mean %.6f, min %.6f%s\n",
(unsigned long long)r.first, (unsigned long long)(r.first + r.n_frames - 1),
r.value[VMAFX_POOL_MEAN], r.value[VMAFX_POOL_MIN],
(r.flags & VMAFX_WINDOW_PARTIAL) ? " (partial)" : "");
vmafx_window_release(window);
| Field | Meaning |
|---|---|
status | VMAFX_OK, or why the window has no values (VMAFX_E_RANGE: no scored frame of it in the stream; VMAFX_E_NOTFOUND: a frame has no score of the target after the flush; VMAFX_E_INVALID: the context was destroyed first). The failure is also logged to the context |
flags | VMAFX_WINDOW_PARTIAL when the stream ended before last |
first, last | As requested |
n_frames | Frames pooled over, from first |
n_scored | Those the context scored: with n_subsample > 1 only every n_subsample-th index |
value[p] | The pooled score of method p for each bit of pool_mask (a model set's bagging score) |
stddev[p], ci95_lo[p], ci95_hi[p] | A model set's other three bootstrap values |
name | The model's, the set's or the feature's name |
Every value is the synchronous call's over the same frames, bit for bit: a window and vmafx_score_pooled() run one pooling implementation with the same arguments. A failure of the window itself is in result.status; the return value of a poll or a wait is about the call.
Threads¶
- The calls that feed a context (
vmafx_submit(),vmafx_flush(),vmafx_context_import_score(),vmafx_window_submit()) are externally synchronised, as every context call is. The completion thread uses the context's engine between them: each call that enters the engine takes the context's engine lock, so the engine sees one caller at a time. vmafx_window_poll(),vmafx_window_wait()andvmafx_window_release()may run on any thread at any time, the feeding thread included: waiting there for a window whose frames are on the workers returns when they finish.- Callbacks run on one callback thread per context, started by the first window that has a callback, in completion order. It is not the completion thread, so a slow callback delays later callbacks but never the completion of a window. A callback may poll, wait on and release windows, its own included, and must not call anything else on the window's context: no
vmafx_submit(), novmafx_flush(), novmafx_context_destroy(). A message the library raises during a callback reaches the context's log callback. vmafx_window_release()cancels an open window: its callback never runs. Released from another thread while its callback runs, the release waits for the callback to return; released from that callback, it does not wait.vmafx_context_destroy()first lets the completion thread finish the changes already signalled, so a window whose frames were final before the destroy completes with its values; then it completes the rest withVMAFX_E_INVALID, runs every callback and stops both threads. A destroy that fails (ADR-1336) leaves the threads running and every open window open.
Cut a stream into windows: the window clock¶
n_stats and n_stats_frames (the FFmpeg filter's window option group) mean the same in every consumer because the library cuts the windows:
VmafxWindowClockConfig cfg = VMAFX_WINDOW_CLOCK_CONFIG_INIT;
cfg.n_stats = 2.0; /* seconds; or cfg.n_stats_frames = 48 */
VmafxWindowClock *clock = NULL;
vmafx_window_clock_create(&cfg, &clock, &error);
/* per frame, before or after submitting it: */
VmafxWindowSpan span = VMAFX_WINDOW_SPAN_INIT;
if (vmafx_window_clock_frame(clock, index, pts_ns, &span, NULL) == VMAFX_OK)
submit_window(span.first, span.last); /* the window that just ended */
/* at the end of the stream: */
if (vmafx_window_clock_finish(clock, &span, NULL) == VMAFX_OK)
submit_window(span.first, span.last); /* flagged VMAFX_WINDOW_PARTIAL */
vmafx_window_clock_destroy(clock);
- By time: window
kholds the frames whose presentation time lies in[t0 + k * n_stats, t0 + (k + 1) * n_stats),t0the first frame's.n_statsis rounded to whole nanoseconds, so membership is exact integer arithmetic.span.start_nsandspan.end_nsare the window's bounds. - By frame count: window
kholds the frames whose index lies in[i0 + k * n_stats_frames, i0 + (k + 1) * n_stats_frames).span.start_nsandspan.end_nsare its first and last frame's times. - A window ends when the first frame past it arrives, so its span comes with that frame. Windows without a frame (a gap in the stream) are skipped:
span.windowcan jump. vmafx_window_clock_finish()returns the window open at the end of the stream, flaggedVMAFX_WINDOW_PARTIALunless it is a full window by frame count, orVMAFX_PENDINGwhen none is open.- Indices must increase and times must not decrease; a refused frame leaves the clock as it was.
Submitting the span right away, before the frame that ended it is submitted, lets the window complete in that submit when its frames are final.
Frames in flight and backpressure¶
vmafx_context_max_in_flight(context) is the most frames of each input the context holds when vmafx_submit() returns, whatever your rate. With R = vmafx_context_frame_retention() and T worker threads it is R + 2 * T * (R + 1), plus 1 on a device backend:
- without worker threads a submit scores its frame before it returns, and the context keeps the
Rreference frames temporal extractors read; - with them, a submit returns once its frame is queued and waits while
Tframes wait for a worker, so at most2 * Tframes are in flight, each with itsRearlier reference frames. A faster producer is held back invmafx_submit(); the queue never grows past the bound.
A producer that recycles its frames (a ring of textures, a VmafxFramePool) needs vmafx_context_max_in_flight() frames plus the one it is filling, and reuses a frame once its release fence (vmafx_frame_release_fence()) is signalled.
A live plugin¶
core/test/test_vmafx_window_live.c imitates a live-production plugin on the CPU device: a render thread paces frames at 60 per second, renders each into a texture of a ring of vmafx_context_max_in_flight() + 1, signals its acquire fence, imports it with vmafx_context_import_frame() and submits it; a window clock cuts 0.2-second windows, and windows of 12 frames are also submitted ahead; a poller thread polls every window. It holds that
- every window equals an offline session on the same frames bit for bit, and
core/test/test_vmafx_window_cli.pyholds the windows against thevmafCLI's per-frame scores of the same pair, pooled with the engine's arithmetic; - a window over frames that are final per frame completes within two frame periods (33.3 ms) of the start of the submit of its last frame, with 0 and 2 worker threads: measured worst about 2.3 ms for windows submitted ahead and about 19 ms for windows the clock cuts (those are submitted one frame later, with the frame that ends them);
- a window completes while the producer stalls: with 2 worker threads every window's frame was still on a worker when its submit returned, and the window completed about 2 ms later with no further call;
- the textures the library holds never exceed the bound, also when the producer is not paced;
- no frame is copied through the host (the host-copy counter stays 0).
Run it with the fixtures in python/test/resource/yuv:
python3 scripts/ci/run_meson_test.py -- -C build test_vmafx_window \
test_vmafx_window_clock test_vmafx_window_live test_vmafx_window_cli
The device lanes run the same harness on CUDA, SYCL, HIP and Metal textures.