VMAFX Phase 4b — Distributed Platform Architecture¶
This page shows how the controller, node and operator binaries fit together on a Kubernetes cluster, and how each one is configured. Read the diagram and the responsibilities table first; the storage, GPU affinity and implementation-status sections follow.
Status
The controller, node, operator and server binaries are implemented under cmd/ and documented under docs/server/. The umbrella ADR-0709 is Accepted (status update of 2026-10-06). ADR-2350 replaces the controller's SQLite queue with PostgreSQL and moves the platform core (stateless replicas, queue-driven scaling, generated CRDs and protobuf) into RC4; the pages below change as each part ships. The implementation status table lists each sweep step.
The layout replaces the single-binary model from Phase 3 (ADR-0701) and Phase 4a (ADR-0702) with a controller/node/operator split designed for horizontal scale on heterogeneous GPU clusters. vmafx-server is the Phase 4a single-binary scoring service; it remains a separate binary (see gRPC server).
Component diagram¶
Text description
- Thin clients: vmafx CLI (Go, proposed), vmafx-mcp (Go, JSON-RPC), vmafx-tune (Go).
- Control plane: vmafx-operator (CRDs, reconcilers, HPA), vmafx-controller (gRPC + HTTP, job queue, scheduler).
- Worker nodes: vmafx-node (NVIDIA, CUDA EP), vmafx-node (AMD, ROCm EP), vmafx-node (Intel, OpenVINO EP).
- Storage: Object store (S3, GCS, Azure, SFTP via rclone), Model registry (.onnx + registry.json).
- Boxes: Kubernetes API (CRD watch), Training sidecar (Python, PyTorch + Lightning).
- Thin clients → vmafx-controller (gRPC).
- vmafx-operator → Kubernetes API (reconcile).
- vmafx-operator → vmafx-controller (pod lifecycle) → Worker nodes (work items).
- vmafx-node → Training sidecar (triples).
- vmafx-node → Training sidecar (triples).
- Worker nodes → Object store (rclone read).
- Worker nodes → Model registry (ONNX load).
- Training sidecar → Model registry (updated .onnx).
- Scenario 1, Score a job: A client job runs on a node that matches its backend. A client submits a job over gRPC. A node with the required backend pulls it. The node reads the media through rclone and loads the model.
- Scenario 2, Online training: Scored triples fine-tune the model beside the node. Nodes pass (reference, distorted, score, metadata) triples to the sidecar. The sidecar writes the updated .onnx to the registry.
- Each worker node runs libvmaf through cgo, FFmpeg as a subprocess, ONNX Runtime and an rclone mount.
- The training sidecar runs beside NVIDIA and AMD nodes and fine-tunes the model from scored triples.
Component responsibilities¶
| Component | Language | Image | Key responsibilities |
|---|---|---|---|
vmafx-controller | Go | distroless/cc | gRPC + HTTP API, job queue, node registry, scheduler, /healthz /readyz /metrics |
vmafx-operator | Go (controller-runtime) | distroless/cc | Watches VmafxJob / VmafxNode / VmafxModelTraining CRDs (VmafxTenant is read by the controller, not the operator); reconciles pod lifecycle; drives HPA |
vmafx-node | Go | distroless/cc + ffmpeg + rclone | Pulls work, runs ffmpeg subprocess, scores via libvmaf cgo, AI inference via Go ONNX Runtime, captures training triples |
training-sidecar | Python (PyTorch + Lightning) | pytorch base | Consumes (ref, dis, score, metadata) triples from co-located node; continuously fine-tunes ONNX model; writes updated .onnx to model registry |
vmafx-mcp | Go | distroless/cc | MCP JSON-RPC server; 5 of its 24 tools (submit_job, get_job, cancel_job, list_jobs, vmaf_score_remote) call the controller over gRPC, the others run the vmaf CLI directly (see MCP tools) |
vmafx-server | Go | distroless/cc | Phase 4a single-binary scoring service: gRPC VmafxScoring + REST (see gRPC server, REST adapter) |
vmafx-tune | Go | distroless/cc | Encoder-ladder optimizer; submits jobs to controller |
CRD summary¶
All four CRDs belong to the group vmafx.dev, version v1 (deploy/helm/vmafx/crds/).
| CRD | Group / version | Scope | Purpose |
|---|---|---|---|
VmafxJob | vmafx.dev/v1 | Namespaced | Describes a scoring / encoding / QA job (source, models, encoder params, target node pool) |
VmafxNode | vmafx.dev/v1 | Namespaced | Describes a node pool (GPU vendor, count, image, resource limits) |
VmafxModelTraining | vmafx.dev/v1 | Namespaced | Describes a sidecar training run (base model, training config, output target) |
VmafxTenant | vmafx.dev/v1 | Namespaced | Maps an auth tenant to its identity provider, suspension switch and role whitelist; the controller reads and enforces it (see auth) |
Storage flow (zero-copy via rclone)¶
A job's sources are local paths, http(s) URLs or rclone remotes. The node reads remotes in the mode VMAFX_STORAGE_MODE selects (ADR-1526); neither mode writes the clip to the node's disk:
Object store (S3 / GCS / Azure Blob / SFTP)
│
├─ http-serve: rclone serve http ──► HTTP GET ──► pipe (/dev/fd/3, /dev/fd/4)
│ │
└─ mount: rclone mount (FUSE) ──► local path ────┤
▼
vmaf CLI (score)
GPU pool affinity¶
Node pods are scheduled via k8s nodeSelector / nodeAffinity resource keys:
| Vendor | Resource key | Backend |
|---|---|---|
| NVIDIA | nvidia.com/gpu | CUDA EP |
| AMD | amd.com/gpu | ROCm EP + HIP |
| Intel | gpu.intel.com/i915 or gpu.intel.com/xe (Helm gpu.intelDriver) | OpenVINO EP + SYCL |
Each backend runs through whichever GPU device plugin is allocated to the pod (per ADR-0701). (The Vulkan backend was removed in ADR-0726.)
Implementation status¶
The nine sweep steps of ADR-0709 and their state in the tree. A step counts as done when its code or artefact exists on master.
| Phase | Description | Input dependency | State |
|---|---|---|---|
| 4b.1 | vmafx-server → vmafx-controller (job queue, node registry, scheduler) | Phase 4a vmafx-server PR merged | Done (cmd/vmafx-controller) |
| 4b.2 | vmafx-node Go binary (libvmaf cgo, ffmpeg, Go ONNX Runtime) | vmafx-sys Rust bindings (Phase 4a) | Done (cmd/vmafx-node, ADR-0713); controller pull loop per ADR-1524 |
| 4b.3 | vmafx-operator kubebuilder skeleton + CRDs | Phase 4b.1 | Done (cmd/vmafx-operator) |
| 4b.4 | ffmpeg latest + ffmpeg-patches/ bundled in node image | Phase 4b.2 | Done (docker/Dockerfile.node, ADR-0717) |
| 4b.5 | rclone integration (node distroless layer + mount lifecycle) | Phase 4b.2 | Done (rclone stage in docker/Dockerfile.node; executor wiring per ADR-1526) |
| 4b.6 | eBPF research digest + ONE concrete optimization | Phase 4b.2 (baseline measurement) | Research done (Research-0733); descriptor tracker wired, opt-in (ADR-1539); no read path bypassed yet |
| 4b.7 | Sidecar training v1 (Python sidecar + triple-capture API) | Phase 4b.2 + Phase 4b.3 | Python sidecar present; not deployed by the chart (sidecar guide) |
| 4b.8 | C ABI break + ffmpeg-patches update | Phase 4b.4 | Open; tracked by ADR-0709 |
| 4b.9 | Native build sunset (Docker + Helm only release artifacts) | Phase 4b.8 | Open; tracked by ADR-0709 |
Related documents¶
- ADR-0709 — umbrella decision record
- ADR-0702 — Phase 4a foundation
- ADR-0701 — Phase 3 cloud-native redesign
- ADR-0699 — Helm chart + k8s manifests
- c4-container.md — C4 Level 2 container view
- Controller, node, operator — per-binary guides