Roadmap¶
VMAFx tracks its plan in GitHub, not in a document that drifts. This page is a map of where that plan lives and how the releases are sequenced.
- Board — VMAFx Roadmap (public)
- Milestones — all milestones
- Epics — issues labelled
epic, each with a child task list
Releases¶
| Milestone | Theme |
|---|---|
| 1.0.0 | First release: RC1 correctness and tester reports, RC2 stabilisation, RC3 twin exactness, RC4 first full Rust metric, zero-copy import, the VMAFx API, the FFmpeg and GStreamer integrations and the cloud-native platform, RC5 deduplication, tool consolidation and new metrics, RC6 GPU capability table, RC7 CPU capability table, RC8 benchmarking and tuning, RC9 real model retraining, then final |
| 1.1 | Integrations and live quality: the OBS Studio plugin (#2239), #2148, #2147, #2144, #2146, #2159; no-reference mode in the plugins (#2413) with its first model (#2166), the live P.1204 monitor (#2417), AI-generated video scoring (#2418), WebRTC, streaming outputs, libVLC and cookbook recipes, and the rest of the milestone |
| 1.2 | Encoder feedback, embedding and platforms: the rest of the embedding epic (#2067), #2164, #2156; encoder-side predictors (#2416), the VLC plugin (#2358), the mobile and WebAssembly targets |
| 1.3 | New metrics with exact twins: #2165, #2167 (a decision after the external runner), the artefact detectors (#2272), ST-GREED (#2394), picks from #2168; our own no-reference models (#2415) and foveated scoring (#2419) |
| 1.4 | Metric A/B comparison, the best current mix and more training data: #2240, #2241 |
| 1.5 | The next model generation: #2242 |
| 2.0 | Breaking changes only: the libvmaf.h compatibility library removed (ADR-1852), the rest of #1254 |
| 2.1 | Rust P1a: the default-model extractors Rust-default with SIMD (#2573) |
| 2.2 | Rust P1b: the remaining CPU extractors and SIMD (#2574) |
| 2.3 | Rust P2: engine, model loading, predict and pooling, generated ABI shells (#2575) |
| 2.4 | Rust P3: CLI, tools, MCP server, ONNX Runtime host, host glue (#2576, #2581) |
| 2.5 | Rust P4 and P5: GPU host runtimes and the kernel verdict per backend (#2577, #2578) |
| 3.0 | Host C and C++ removed outside the exception list (#2579) |
ADR-2001 set this layout on 2026-10-06 and declared it the last change of the milestone map until 2.0, apart from bugs and findings. The cloud-native work (scoring API contract, server mode, observability, the cloud-native platform, containers, Helm, operator, GPU pool arbiter) is part of 1.0.0. Every release after 1.0.0 runs its own candidate cycle: features first, then deduplication, tests and bug fixing, then capability tables with benchmarks and tuning, then training if models change, then the release; the same phase rules as for 1.0.0 apply inside each cycle.
Two milestones are deliberately rolling rather than tied to a release:
- Models & benchmarks — retraining cadence, benchmark baselines, corpus work.
- Code health & deduplication — the fork adds and changes a lot, so slimming it is recurring work, not a one-off.
How 1.0.0 is gated¶
The fork's first release candidate, v1.0.0-rc.1, was published on 2026-09-27; older tags are inherited upstream history. ADR-1341 gives each first-release candidate one responsibility. ADR-1421 maps the stages to tags, so that each stage number matches its v1.0.0-rc.N tag (it supersedes the mapping of ADR-1352), and ADR-1490 inserts the CPU capability stage as RC7, which moves benchmarking to RC8 and retraining to RC9. ADR-1868 folds the work that joined 1.0.0 on 2026-10-05 into those candidates without new numbers: the new API and provenance into RC4, tool consolidation, new metrics and the Metal SpEED twins into RC5, training readiness into RC8. ADR-1880 adds the format envelope: an overflow audit at 8K and 16K and 8K exactness in RC3, the supported resolutions, bit depths and layouts per backend and device in the RC6 and RC7 tables, throughput per resolution in RC8; and device-targeted scoring (one decode scored for several displays) in RC5. ADR-2001 moves the cloud-native work into 1.0.0 without new numbers: the versioned scoring API contract, server mode and observability in RC4, containers, Helm chart, operator and GPU pool arbiter in RC5, distributed throughput in RC8; a native GStreamer element, an API ready for OBS Studio and real-time FFmpeg GPU scoring in RC4. The OBS Studio plugin follows in 1.1. ADR-2342 records the scope decisions of 2026-10-06 and 2026-10-07, again without new numbers: RC4 also holds the FFmpeg series redesign, the input-format work, the engineering principles per language, the FFmpeg audit, the observability package and the cloud-native platform; RC5 also holds live alignment, interlaced video, region masks, container input, bits per pixel and BD-rate, HandBrake support, the minimal run-result timeline and the reusable provenance workflow.
What this means for you¶
- Use the newest candidate. Each candidate keeps the Netflix golden scores; later candidates make GPU and SIMD results match the CPU and remove duplicated code.
- No performance claims before RC8. Candidates up to RC7 are about correctness; speed numbers are measured and tuned in RC8.
- Model retraining comes last, in RC9.
- Help test. Results from hardware the project does not own count as evidence: run the tester image.
The stages¶
| Stage | Theme | Tracking issue |
|---|---|---|
| RC1 | correctness and tester readiness | — |
| RC2 | stabilisation and repair | — |
| RC3 | twin exactness, overflow audit at 8K and 16K | #1721 |
| RC4 | first full Rust metric, zero-copy device-frame import, new VMAFx API and its bindings, FFmpeg filters and series redesign, GStreamer element, provenance, input formats, scoring API contract, server mode, observability, cloud-native platform, engineering principles, OBS-ready API | #1723 |
| RC5 | deduplication, tool consolidation (live alignment, interlaced video, region masks, container input, bits per pixel and BD-rate, HandBrake, run-result timeline), new metrics with exact twins, Metal SpEED twins, device-targeted scoring, containers, Helm, operator, GPU pool arbiter | #1724 |
| RC6 | GPU capability source of truth, GPU format envelope, legacy GPU build variants | #1725 |
| RC7 | CPU capability source of truth, CPU format envelope, full SIMD ladder on five architectures | #1885 |
| RC8 | benchmark and tune, throughput per resolution, distributed throughput, training readiness | #1245 |
| RC9 | real retraining | #1246, #1242 |
Final v1.0.0 | — | — |
RC1 — correctness and tester readiness¶
- In scope: Release-blocking correctness, reliability, security, build, packaging, backend usability, and a portable report path for outside hardware
- Exit boundary: The exact candidate head is green; no confirmed RC1 blocker or untriaged
docs/state.mdrow remains; a tester can return artifact/environment identity, device and tool versions, backend availability, correctness/parity results, commands, logs, and failures
RC2 — stabilisation and repair¶
- In scope: The dependency updates and correctness fixes merged since rc.1, delivered to testers through the same report path; no benchmark or training work
- Exit boundary: The RC1 boundary, re-established on the rc.2 head
RC3 — twin exactness¶
- In scope: Every GPU and SIMD twin returns the CPU extractor's scores bit for bit, or carries a measured tolerance recorded in an ADR; no SYCL kernel uses scratch memory. Since 2026-10-05 (ADR-1880): an integer-overflow audit of every extractor and twin at 8K and 16K frame sizes with 16-bit samples
- Exit boundary: Per-twin parity measured at
--precision maxon the Netflix pairs, the 1080p checkerboard pairs, the 4K fixture and 8K cells; no accumulator overflows at 16K with 16-bit samples; Netflix golden assertions unchanged
RC4 — first full Rust metric, zero-copy import and the VMAFx API¶
Also in RC4 since 2026-10-05 (ADR-1868): provenance on every score (#2142) through the new API.
- In scope: The whole
vmaf_v1.0.16_3d0hpath (cambi, speed_chroma, integer adm3, integer motion3, model prediction) in Rust; the C ABI is unchanged and the GPU twins stay CUDA, SYCL and HIP code. The whole device-memory import API (ADR-1829): additive import onVmafPicture2with fences in both directions for CUDA, SYCL, HIP and Metal, NV12 and P010 on the GPU, CUDA without its device-to-device copy, SYCL chroma import and D3D11, Metal IOSurface and MTLTexture bound without the CPU copy, a HIP import path, FFmpeg filters that take hardware frames. The new VMAFx C API (vmafx/*.h,libvmafx.so.1) generated with every other surface fromcore/api/vmafx.toml,libvmaf.has a thin compatibility library on it, and the FFmpeg filters under VMAFx names (vmafx,vmafx_tune,vmafx_pre) (ADR-1852). Since 2026-10-06 (ADR-2001): the versioned scoring API contract (#2155) and server mode with observability (#1251), generated from the same definition; a nativevmafxGStreamer element (#2236) and CI conformance of thevmafelement of upstream's gst-plugins-bad on the compatibilitylibvmaf.so.3(#2237); an API ready for OBS Studio (#2238: texture import including OpenGL interop, asynchronous window scores, bounded queues); real-time FFmpeg GPU scoring withn_stats(#2138). Since 2026-10-07 (ADR-2342): language bindings generated from the definition (Rust, Go, Python); the FFmpeg patch series redesigned in one change from0001, grouped by purpose, after an audit of everything VMAFx uses in FFmpeg; host input of every semi-planar and packed layout, RGB with a stated matrix and every bit depth from 8 to 16; incrementalmotion2/motion3for live windows; Vulkan frame import; engineering principles per language with warnings as errors in every one; the observability package (#2430); Linux arm64 and older-glibc binaries and the package-manager channels (#2437); the cloud-native platform (#2431): state out of the processes (PostgreSQL, a job queue, a two-tier cache, object storage, OCI artifacts), scaling on queue depth, and CRDs, proto, OpenAPI and the Helm schema generated from the definition; conformance of SVT Encore's VMAF options on the compatibility library (#2364) - Exit boundary: The Rust path is bit-identical to the C path on the parity fixtures; an imported device frame scores bit-identically to the same frame uploaded from the host, with no host copy of pixel data and fence-ordering tests that fail when a fence is skipped; the golden-data gate passes through the compatibility library and every generated surface is checked against the definition
RC5 — deduplication¶
- In scope: One implementation per behaviour across GPU twins and host code, the Rust code included;
libgpudispatchextracted, folding in the per-backend import code RC4 wrote (#1455). One implementation per tool (#1249), thetools/surface finished and the known unfinished surfaces closed (#1250, #1270, #1272). The new metrics with their twins written once onlibgpudispatch: ΔE-ITP, PU21, NIQE, BRISQUE, Y-FUNQUE+ (#1247, #1248), HDR-SSIM and HDR-MS-SSIM (#2161), XPSNR (#2158); Metal twins ofspeed_chromaandspeed_temporal(#2160). Device-targeted scoring (ADR-1880): device profiles (phone, tablet, laptop, TV, VR per eye, portrait included), each a target resolution, a scaling and a viewing distance per display height mapped onto the ADM optionsnvdandrdh; one decode scored for many targets; a short research pass first, the profile table generated by the RC6 / RC7 table machinery. Cloud-native deployment (ADR-2001): containers, Helm chart and a kind plus kuttl test setup (#1252); the operator, the controller / node split and the GPU pool arbiter inlibgpudispatch(#1253). Since 2026-10-07 (ADR-2342) the tool consolidation also carries live alignment of two feeds with timecode (#2354), interlaced video (#2361), region masks (#2362), container input (#2363), bits per pixel and BD-rate (#2284), HandBrake support (#2407) and the minimal run-result timeline they report on; one reusable build-and-provenance workflow for every image; and a small Vulkan compute experiment with a written verdict. The Go tools replace the Python MCP server andvmaf-tunehere, not after 1.0.0 - Exit boundary: Scores unchanged against the RC3 reference; duplicated code removed rather than moved; every new twin bit-identical to its CPU extractor or within a measured libm bound recorded in an ADR; a multi-target run scores each target as a separate run with that target's options would
RC6 — GPU capability source of truth¶
- In scope: A per-vendor capability table generated from
nvcc --list-gpu-arch,oclocand ROCmllc -mcpu=help, checked in with a CI drift check, covering the twins RC5 adds; dispatch and kernel parameters read it, with a generic fallback for unknown devices; every kernel compiled and statically audited for every target (scratch, spills, register ceiling, fp64). The table also declares the format envelope per backend and device (ADR-1880): maximum resolution up to 16K with measured memory limits and tiling where needed, bit depths 8 to 16, chroma layouts, odd and portrait sizes, each row backed by a test. Legacy build variants (ADR-2001): CUDA 12.x builds for Maxwell, Pascal and Volta (sm_50 to sm_72; CUDA 13.4 starts at compute_75), the Intel legacy compute runtime for Gen9 to Gen11 iGPUs, and every AMD target the pinned ROCm compiler still emits, each bit-exact and listed in the table - Exit boundary: Drift check green, the envelope included; audit clean for every listed target
RC7 — CPU capability source of truth¶
- In scope: The CPU twin of RC6: a checked-in table, generated by one script, of the CPU features each SIMD kernel needs (from the per-translation-unit compile flags in
core/src/meson.buildand the runtime gates incore/src/x86/cpu.candcore/src/arm/cpu.c), with a CI drift check; a per-function disassembly audit for x86 and aarch64; every dispatch level run bit-exact against scalar under emulation. The full SIMD ladder (ADR-2001): every extractor gets a bit-exact kernel at every useful ISA level with runtime dispatch, on x86-64 (SSE2, SSSE3, SSE4.1, AVX, AVX2, AVX-512, AVX-512 ICL, AVX10), AArch64 (NEON, dotprod/i8mm, SVE, SVE2), RISC-V RVV 1.0, POWER VSX and LoongArch LSX/LASX; qemu-user CI for architectures without hardware, the golden gate on each. The CPU format envelope (resolution up to 16K with measured memory limits, bit depths 8 to 16, chroma layouts, odd and portrait sizes) in the same table, each row test-backed (ADR-1880). A reference conformance column: every extractor with an original implementation is proven against it, the default becomes reference-exact and Netflix's behaviour a named compatibility mode that the golden gate runs in (ADR-2343, #2286) - Exit boundary: Drift check green; audit finds no instruction outside the feature set a gate guarantees; parity green under Intel SDE (AVX2-only model, Skylake-X, Ice Lake, Sapphire Rapids, the AMD AVX-512 set) and qemu (aarch64 NEON, SVE2 at more than one vector length). Reports from real Xeon or Apple Silicon machines are extra evidence, not a requirement; timing is RC8
RC8 — benchmark and tune¶
- In scope: Comparable benchmark baselines, profiling, hardware-generation retuning, and measured performance fixes, including the speed RC3 gave up for exactness; throughput per resolution, 16K included (the envelope itself is RC6 / RC7 evidence), and distributed throughput across nodes (ADR-2001). Training readiness: automatic temporal alignment and the HDR-input guard for SDR models (#2163, #2157), the external-metric runner and estimator calibration (#2162, #2143), the HDR conversion-check workflow (#2145), the mini retrain in CI and the measured resource plan of #1246
- Exit boundary: Results identify the exact artifact, fixtures, host, drivers and runtimes; accepted wins are re-measured and preserve correctness/parity; the mini retrain passes every stage
RC9 — real retraining¶
- In scope: The locked one-shot model retraining programme on the clean, tuned tree, started only when every precondition of #1246 holds (the RC4 to RC8 items above included). The shipped v1 models read compatibility-mode features until this run; the retrain trains on reference-exact features (ADR-2343)
- Exit boundary: Model-quality gates, model cards, registry/signing metadata, and unchanged Netflix golden assertions pass
Final v1.0.0¶
- In scope: Accepted RC9 output plus any required repair candidate
- Exit boundary: Publication preflight passes on the immutable final tag
“Done fixing” is deliberately bounded rather than a promise that no future bug will be found. RC1 and RC2 are ready when there are no confirmed, actionable release blockers and no untriaged rows. Performance-only findings belong to RC8, real training belongs to RC9, and externally blocked work stays explicitly deferred with its trigger and evidence.
If any later candidate exposes a correctness regression, fix it and rerun the affected stage evidence before proceeding. Do not pull general benchmarking into RC1 to RC7, or real training before RC8 evidence is accepted. Speed that RC3 gives up for exactness is recorded as a tuning row and recovered in RC8; it is not traded back for a tolerance. The RC1 and RC2 report envelope may run a short correctness and device-engagement smoke; it does not make a performance claim.
Ordinary Renovate and other version PRs remain mergeable throughout the sequence when normal required checks, review, pinning, and component-specific validation pass. Security updates are prioritised rather than being the only allowed updates. Because evidence is exact-head, any later merge requires the affected candidate checks or measurements to be rerun.
Rust core¶
ADR-2478 decides that Rust replaces all host-side C and C++ by 3.0 behind the unchanged C ABI, exported from Rust and generated from the API definition. The epic is #2567; the phases and their milestones are in the ADR's phase table and the research digest is Research-2479. Native GPU device sources and C-only host glue stay on a named exception list with re-evaluation triggers. The C implementation of a layer is the differential oracle until it is deleted; the libvmaf.h compatibility removal stays in 2.0.
Things that do not change¶
Some guarantees are load-bearing for downstream users and hold across every milestone above, with one exception: 2.0 removes the libvmaf.h compatibility library (ADR-2001, ADR-1852). Until then:
- The Netflix golden values are never edited. They are the numerical ground truth; if scores drift, the code is wrong.
- The
libvmaf.soABI and the FFmpeglibvmaffilter name stay stable, even as the internals move to C++23 and parts of the tooling move to Go and Rust. - The public C API under
core/include/libvmaf/stays source-compatible. - Release artifacts are built in the container, never from a host build.
Contributing against the roadmap¶
Pick an epic, read its task list, and open a PR that closes one line of it. Epics are intentionally coarse — sub-tasks become their own issues when someone starts them, so the tracker reflects work in progress rather than a wish list.