MCP tool reference¶
Per-tool request / response schemas and error semantics for the MCP servers (overview). The Go server vmafx-mcp serves all 24 tools; the Python server vmaf-mcp serves the first 19 in the table below. Source of truth: the tool registrations in cmd/vmafx-mcp/tools.go and the list_tools() handler in mcp-server/vmaf-mcp/src/vmaf_mcp/server.py. The names, property types and required arguments the two servers share are recorded in mcp-server/vmaf-mcp/tool-contract.json, written from the Python handler and checked against both servers (Tests).
Every tool returns a single TextContent message whose body is a JSON document. On error the body has shape {"error": "<string>"}, so clients can json.loads() unconditionally and branch on the presence of error.
Tool summary¶
| Tool | Arguments | Servers | Execution |
|---|---|---|---|
vmaf_score | ref, dis, width, height, pixfmt, bitdepth, optional model, backend, subsample, precision, scoring extras | Python, Go | CLI subprocess; Go: optional cgo |
list_models | none | Python, Go | filesystem |
list_backends | none | Python, Go | vmaf --list-backends probe |
run_benchmark | none | Python, Go | bench_all.sh |
eval_model_on_split | model, features, split, input_name | Python, Go | native |
compare_models | models, features, split | Python, Go | native |
describe_worst_frames | ref, dis, width, height, pixfmt, bitdepth, optional n | Python, Go | CLI + VLM |
probe_backend | backend | Python, Go | CLI probe |
vmaf_version | none | Python, Go | vmaf --version |
vmaf_score_encoded | reference_encoded, distorted_encoded, optional model, backend, subsample, precision | Python, Go | ffmpeg + CLI |
list_extractors | none | Python, Go | vmaf probe |
describe_model | name | Python, Go | CLI; Go: optional cgo |
run_compare, run_ladder, run_tune_per_shot | see each tool | Python, Go | vmaf-tune |
vmaf_per_shot, vmaf_roi, vmaf_bench, vmaf_vpl | see each tool | Python, Go | sidecar binaries |
submit_job, get_job, cancel_job, list_jobs, vmaf_score_remote | see each tool | Go only | gRPC |
vmaf_score¶
Score one (ref, dis) YUV pair and return the full VMAF JSON report.
Input schema¶
The scoring tools (vmaf_score, vmaf_score_encoded, describe_worst_frames) take the arguments below; the Tools column names the tools that accept each one. Both servers serve these schemas from one generated file (options.gen.json, generated from the option groups of core/api/vmafx.toml), so the Go and Python servers cannot drift, and each argument reaches the same vmaf flag (CLI column of the CLI reference). A model that is not given is the library default model (VMAF_DEFAULT_MODEL_VERSION, vmaf_v1.0.16_3d0h in this release); pass model="version=vmaf_v0.6.1" to reproduce the numbers of releases before 1.0.0-rc.4.
| Argument | Type | Required | Default | Tools | Description |
|---|---|---|---|---|---|
ref | string | yes | vmaf_score, describe_worst_frames | Reference video: a .y4m file, or raw planar .yuv together with the width, height, pixel format and bit depth. | |
dis | string | yes | vmaf_score, describe_worst_frames | Distorted video, in the same form as the reference. | |
width | uint >= 1 | yes | vmaf_score, describe_worst_frames | Width of raw .yuv input in pixels. | |
height | uint >= 1 | yes | vmaf_score, describe_worst_frames | Height of raw .yuv input in pixels. | |
pixfmt | 420 | 422 | 444 | yes | vmaf_score, describe_worst_frames | Chroma subsampling of raw .yuv input. | |
bitdepth | 8 | 10 | 12 | 16 | yes | vmaf_score, describe_worst_frames | Bits per sample of raw .yuv input. | |
reference_encoded | string | yes | vmaf_score_encoded | Path to the reference encoded video (MP4/MKV/Y4M/...). Must be under an allowlisted root (VMAF_MCP_ALLOW). | |
distorted_encoded | string | yes | vmaf_score_encoded | Path to the distorted encoded video. | |
model | string | library default (VMAF_DEFAULT_MODEL_VERSION) | vmaf_score, vmaf_score_encoded, describe_worst_frames | Model, colon-delimited: version= a built-in model, path= a model file, name= the name in the report, disable_clip, enable_transform, <feature>.<option>=<value> overloads. Several models score in one pass (repeat the option; the filter separates them with |). | |
disable_clip | bool | false | vmaf_score, vmaf_score_encoded | Do not clip the model score to [0, 100] (the model's disable_clip). | |
enable_transform | bool | false | vmaf_score, vmaf_score_encoded | Apply the model's score transform (the model's enable_transform). | |
backend | auto | cpu | cuda | sycl | hip | metal | auto | vmaf_score, vmaf_score_encoded, describe_worst_frames | Backend: auto uses the available ones; any other value runs that backend alone and fails when it is not available. | |
feature | string (list) | vmaf_score, vmaf_score_encoded | Additional feature extractor, name[=key=value:...] (for example psnr or cambi=full_ref=true); several may be given (the filter separates them with |). Mutually exclusive with the CTC presets. | ||
aom_ctc | v1.0 | v2.0 | v3.0 | v4.0 | v5.0 | v6.0 | v7.0 | vmaf_score, vmaf_score_encoded | AOM common test conditions preset: a fixed model and feature set. | ||
nflx_ctc | v1.0 | vmaf_score, vmaf_score_encoded | Netflix common test conditions preset: a fixed model and feature set. | ||
tiny_model | string | vmaf_score, vmaf_score_encoded | Tiny ONNX model to load alongside the classic models. | ||
tiny_device | auto | cpu | cuda | openvino | openvino-npu | openvino-cpu | openvino-gpu | coreml | coreml-ane | coreml-gpu | coreml-cpu | rocm | auto | vmaf_score, vmaf_score_encoded | ONNX Runtime execution provider of the tiny model. | |
dnn_ep | auto | cpu | cuda | openvino | openvino-npu | openvino-cpu | openvino-gpu | coreml | coreml-ane | coreml-gpu | coreml-cpu | rocm | vmaf_score, vmaf_score_encoded | Alias of the tiny-model device under the ONNX Runtime name (execution provider); both set the same setting. | ||
tiny_threads | uint | vmaf_score, vmaf_score_encoded | Intra-op threads of the CPU execution provider (0: the runtime's default). | ||
tiny_fp16 | bool | false | vmaf_score, vmaf_score_encoded | Request fp16 input and output where the execution provider supports it. | |
tiny_model_verify | bool | false | vmaf_score, vmaf_score_encoded | Require a Sigstore bundle verification of the tiny model (cosign verify-blob) before it loads; a missing bundle, a missing cosign or a failed verification refuses the model. | |
tiny_codec | string | vmaf_score, vmaf_score_encoded | Encoder of the distorted clip, required by codec-aware tiny models (fr_regressor_v2/v3), which refuse to score without it. Must be in the model sidecar's encoder_vocab; the ffprobe names h264, hevc, av1, vp9 and vvc are accepted. | ||
tiny_preset | string | vmaf_score, vmaf_score_encoded | Encoder preset (medium, slow, p4, 5, ...), read as the encoder defines it. Unset: ordinal 5 (medium). A model trained with one preset (fr_regressor_v3) ignores it and warns. | ||
tiny_crf | uint 0..63 | vmaf_score, vmaf_score_encoded | CRF or QP used for the encode, normalised as the model sidecar declares. Required with the codec and preset. | ||
tiny_resize | bilinear | nearest | bicubic | disabled | disabled | vmaf_score, vmaf_score_encoded | Resize filter for NCHW tiny models whose input size differs from the frame; disabled refuses the mismatch (-ERANGE). The three filters give scores about 2% apart: record the filter with the model. | |
no_reference | bool | false | vmaf_score, vmaf_score_encoded | No-reference mode; needs a no-reference tiny model. The reference becomes a formality: only the distorted picture is scored. | |
threads | uint | vmaf_score, vmaf_score_encoded | Worker threads of the feature extractors, capped to the hardware threads (0: score in the calling thread). | ||
frame_cnt | uint >= 1 | vmaf_score, vmaf_score_encoded | Score at most this many frames. | ||
frame_skip_ref | uint | vmaf_score, vmaf_score_encoded | Skip this many frames at the start of the reference. | ||
frame_skip_dist | uint | vmaf_score, vmaf_score_encoded | Skip this many frames at the start of the distorted video. | ||
no_prediction | bool | false | vmaf_score, vmaf_score_encoded | Extract features only; no model score. | |
cpumask | uint | vmaf_score, vmaf_score_encoded | Bitmask of CPU instruction sets the extractors must not use. | ||
gpumask | uint | vmaf_score, vmaf_score_encoded | Bitmask of GPU operations the extractors must not use. | ||
sycl_device | uint | vmaf_score, vmaf_score_encoded | SYCL GPU by index (unset: selected automatically). | ||
hip_device | uint | vmaf_score, vmaf_score_encoded | HIP GPU by index (opt-in: HIP is off unless this or --backend hip is given). | ||
metal_device | uint | vmaf_score, vmaf_score_encoded | Metal GPU by index (opt-in: Metal is off unless this or --backend metal is given). | ||
subsample | uint >= 1 | 1 | vmaf_score, vmaf_score_encoded | Score every n-th frame (1: every frame). | |
precision | legacy | max | full | 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 | 9 | 10 | 11 | 12 | 13 | 14 | 15 | 16 | 17 | legacy | vmaf_score, vmaf_score_encoded | Score precision: N (1 to 17) writes %.<N>g; max or full write %.17g (round-trip lossless); legacy writes %.6f (Netflix-compatible). The scoring server returns lossless scores unless asked otherwise. | |
output_fmt | json | xml | csv | sub | json | vmaf_score, vmaf_score_encoded | Report format; json and xml carry the backend receipt. | |
csv | bool | false | vmaf_score, vmaf_score_encoded | Write the report as CSV; the same as output_fmt csv. | |
sub | bool | false | vmaf_score, vmaf_score_encoded | Write per-frame scores as subtitles; the same as output_fmt sub. | |
view_distance | float 3..24 | vmaf_score, vmaf_score_encoded | Viewing distance in display heights (ADM adm_norm_view_dist). Unset: the model's value. The ADM extractor's default CSF refuses a distance below 3 (it accepts 0.75 with other CSF modes, which these options do not set). | ||
display_height | uint >= 1 | vmaf_score, vmaf_score_encoded | Height of the reference display in pixels (ADM adm_ref_display_height). Unset: the model's value. | ||
target_width | uint | 0 | vmaf_score, vmaf_score_encoded | Width of the target display the distorted video is scaled to (0: no scaling). Reserved: device-targeted scoring lands in RC5; only the default is accepted. | |
target_height | uint | 0 | vmaf_score, vmaf_score_encoded | Height of the target display the distorted video is scaled to (0: no scaling). Reserved: device-targeted scoring lands in RC5; only the default is accepted. | |
target_scaling | none | bilinear | bicubic | lanczos | none | vmaf_score, vmaf_score_encoded | Scaling filter towards the target display. Reserved: device-targeted scoring lands in RC5; only the default is accepted. | |
n | uint 1..32 | 5 | describe_worst_frames | How many worst-VMAF frames to describe. |
All optional scoring parameters below (feature, aom_ctc/nflx_ctc, the tiny_* / no_reference tiny-AI surface, and the frame-range controls) are also accepted on vmaf_score.
Behaviour¶
The server exec's the local vmaf binary with, effectively:
vmaf -r <ref> -d <dis> --width <w> --height <h> -p <pixfmt> -b <bitdepth> \
-m <model> --precision <precision> -q --json -o <tmp>
# <model> is the caller's model, or version=<library default> when it is
# omitted, with :disable_clip / :enable_transform / :adm.* suffixes appended
# for the model switches; then --subsample when > 1, then the optional
# arguments' flags in the order of the table above;
# plus per-backend flags disabling sibling backends:
# backend=cpu → --no_cuda --no_sycl --no_hip --no_metal
# backend=cuda → --no_sycl --no_hip --no_metal
# backend=sycl → --no_cuda --no_hip --no_metal
# backend=hip → --no_cuda --no_sycl --no_metal
# backend=metal → --no_cuda --no_sycl --no_hip
The JSON written by vmaf is parsed and returned with two backend receipt fields guaranteed by the MCP response (ADR-0495):
backend_requested— verbatim echo of the caller'sbackendarg.backend_used— what actually ran. For an explicitbackendarg this equals the requested value (both wrappers refuse to silently fall back). Forbackend="auto", the Python and Go wrappers preserve thebackend_usedreceipt emitted by the fork'svmafCLI ("cpu","cuda","sycl","hip", or"metal"). Older or external binaries that omit a valid receipt produce"unknown"; neither wrapper guesses from the number of metric keys because that count changes as extractors evolve.
When the local vmaf binary does not advertise the requested backend, the wrapper raises rather than running CPU silently (the bug-1 pattern from the 2026-05-17 probe). Use backend="auto" to opt back into vmaf's own probe.
The wrapper additionally emits a mismatched_model_warning field when the model's intended resolution preset disagrees with the source frame size — e.g. version=vmaf_4k_v0.6.1 on a 576×324 source saturates at 100 on every frame and the warning surfaces the foot-gun. Bespoke ONNX models with no known resolution preset are silent (no false positives). See usage/cli.md for the rest of the report schema; the temp file is always unlinked, even on error.
precisiondefault"legacy". The MCP server passes--precision legacy(%.6f, Netflix-compatible) by default, matching the underlyingvmafCLI default per ADR-0119. Pass"max"(or"17", i.e.%.17g, IEEE-754 round-trip lossless) when a programmatic consumer needs scores that re-parse to the exact same double.
Example call¶
{
"method": "tools/call",
"params": {
"name": "vmaf_score",
"arguments": {
"ref": "python/test/resource/yuv/src01_hrc00_576x324.yuv",
"dis": "python/test/resource/yuv/src01_hrc01_576x324.yuv",
"width": 576,
"height": 324,
"pixfmt": "420",
"bitdepth": 8,
"backend": "cpu"
}
}
}
Response body (abridged):
{
"version": "1.0.0-rc.2",
"pooled_metrics": { "vmaf": { "mean": 76.667831, "...": "..." } },
"frames": [ { "frameNum": 0, "metrics": { "vmaf": 78.8263, "...": "..." } } ]
}
Errors¶
- Path not under an allowlisted root →
{"error": "path ... not under an allowlisted root; set VMAF_MCP_ALLOW to extend."}. - Path does not exist →
{"error": "<abs-path>"}fromFileNotFoundError. - vmaf binary missing →
{"error": "vmaf binary not found at ...; Build first: meson compile -C build."}. - Non-zero vmaf exit →
{"error": "vmaf exited <code>: <stderr>"}. - Caller-requested backend not advertised by the local binary →
{"error": "backend 'cuda' requested but the local vmaf binary does not advertise it (available: ['cpu']); refusing to fall back silently. Pass backend='auto' to let vmaf pick, or rebuild with the requested backend enabled."}(ADR-0495).
Optional scoring parameters¶
The optional arguments of the input schema table (ADR-1117) are accepted on both vmaf_score and vmaf_score_encoded. Each reaches the vmaf flag named in its description and is forwarded only when supplied: omitting them leaves the score identical to a call without them. Both servers read their schemas and the flag of every argument from the same generated file, so they validate the same values and pass the same flags in the same order.
Input validation. MCP inputs are untrusted and are checked against the generated schema before the CLI is spawned:
- An enumerated argument (
tiny_device,dnn_ep,tiny_resize,aom_ctc,nflx_ctc,pixfmt,backend,precision,target_scaling) refuses a value outside its list. - A numeric argument refuses a value outside its range (
tiny_crf0 to 63,frame_cntandsubsampleat least 1,view_distance3 to 24, the device indices and masks at least 0);bitdepthis one of 8, 10, 12, 16.threads0 is accepted and means the CLI's single-threaded scoring. featureis a list of non-empty strings; any other entry refuses the call (it used to be dropped silently).- Conflicting
tiny_deviceanddnn_epvalues, andcsv/subthat disagree withoutput_fmt, are refused. - The reserved device-target arguments (
target_width,target_height,target_scaling) accept only their defaults until device-targeted scoring lands in RC5; another value is refused with that reason.
The error names the argument: invalid tiny_crf 64: must be <= 63.
Model switches. disable_clip, enable_transform, view_distance (ADM adm_norm_view_dist) and display_height (ADM adm_ref_display_height) are appended to the model specification (version=...:disable_clip:adm.adm_norm_view_dist=4.5), each once.
No-reference mode. When no_reference is set, only the distorted picture is scored, so no_reference requires tiny_model (an NR ONNX model — there is no classic NR scorer). The request is rejected with {"error": "no_reference requires tiny_model; no classic NR scorer exists"} otherwise, mirroring the CLI's cli_parse.c gate. In NR mode the ref argument becomes optional on vmaf_score: omit it, or pass any valid YUV of matching geometry (it is not consumed by the scorer). For vmaf_score_encoded the reference video is still decoded as usual.
CTC presets. The aom_ctc / nflx_ctc presets configure a fixed model and feature set and are mutually exclusive with manual feature / model configuration: combining them with feature or a custom model stacks both configurations (the CLI does not reject it, but the result is rarely what you want).
Output format. output_fmt (json by default; csv and sub also as boolean shorthands) selects the report format. Non-JSON formats return a structured text payload ({"format": ..., "output": ...}).
list_models¶
Walk model/ (recursively) and list every .json, .pkl, or .onnx file shipped with the build.
Input schema — no arguments.¶
Response body¶
{
"models": [
{
"name": "vmaf_v0.6.1",
"path": "model/vmaf_v0.6.1.json",
"format": "json",
"size_bytes": 9128
},
{
"name": "lpips_sq_small",
"path": "model/tiny/lpips_sq_small.onnx",
"format": "onnx",
"size_bytes": 4873216
}
]
}
name is the file stem (no extension). Use it with vmaf_score's model field as "version=<name>" for built-in .json models or as a plain path for custom .pkl / .onnx.
Errors — none in the normal case (an empty model/ returns {"models": []}).¶
list_backends¶
Probe the local vmaf binary and report which runtime backends it was built with.
Input schema — no arguments.¶
Response body¶
The server runs vmaf --list-backends and reports a backend true when it was compiled in and its state initialises on this host; cpu is always reported true when the binary exists. A backend that initialises can still fail on a particular score (a driver fault mid-run). Use probe_backend to verify that a backend can actually run a score.
Errors¶
- If the vmaf binary is missing,
cpustaystrueand every GPU flag isfalse— no error is raised. Calllist_backendsbefore other tools to test whether the build is usable.
run_benchmark¶
Run the full multi-fixture benchmark suite (testdata/bench_all.sh) against all available compiled-in backends (CPU, CUDA, SYCL, HIP, Metal — Vulkan removed in ADR-0726) on three canonical YUV fixture pairs built into the harness:
- 576×324, 48 frames, 8-bit — the Netflix golden pair
src01_hrc00 / src01_hrc01 - 1920×1080, 5 frames, 8-bit — the 5-frame 1080p pair
- 3840×2160, 200 frames, 8-bit — the 4K BBB excerpt (
testdata/bbb/)
For each fixture the harness scores all compiled-in backends, prints per-backend VMAF means and wall times, and prints a comparison table showing max per-frame diff between CPU and each GPU backend. See usage/bench.md for more detail.
No per-call arguments
This tool does not accept per-call ref/dis arguments. Per-pair scoring is the job of vmaf_score; bench_all.sh is a fixed-fixture harness (ADR-0517).
Protocol note
run_benchmark runs the full 4K test, which takes 30–60 seconds on a modern GPU. Real MCP clients hold the connection open. The heredoc test pattern (docker exec -i ... vmaf-mcp << EOF ... EOF) makes the server shut down on stdin EOF before the benchmark completes. Use a persistent pipe (sleep 120 |) when testing from the command line. ADR-0517 preserves the diagnosis and repair rationale for the original failure.
Error contract
run_benchmark raises RuntimeError("benchmark failed — no output line containing pooled score / Pearson correlation") on partial or silent pipe failures (ADR-0638). A second legacy implementation that swallowed the failure and returned a partial dict was removed in PR #517 (Layer-5). MCP clients should branch on isError=True; see Cross-tool error conventions.
Input schema¶
Takes no arguments.
Response body¶
{
"exit_code": 0,
"stdout": "=========================================\nTest 1: Official 576x324 (48 frames, 8-bit)\n...",
"stderr": ""
}
The stdout field contains the full human-readable benchmark output. Per-backend JSON result files are written to /tmp/vmaf-bench-<pid>/ (or to VMAF_BENCH_OUTDIR if set).
Errors¶
testdata/bench_all.shmissing → raisesFileNotFoundErrorwith path.- Non-zero
exit_codeis not itself an error field — both stdout and stderr are always returned so the caller can diagnose partial failures. - Non-zero exit + empty stdout + empty stderr →
errorkey is added with a root-cause shortlist and abash -xre-run hint. Common causes: missing vmaf binary or missing fixture YUVs undertestdata/bbb/. - Unavailable backends (for example HIP scaffold-only) produce a
SKIPline in stdout and do not abort the harness.
eval_model_on_split¶
Load an ONNX tiny-AI regressor, run it against a parquet feature cache, filter to a deterministic train / val / test split (keyed by the key column via SHA-256 bucketing — same scheme as vmaf_train), and report correlations against the mos target.
Python server (vmaf-mcp) requires the optional eval extra, which pulls in numpy, pandas, scipy, and onnxruntime:
Go server (vmafx-mcp) needs no Python at all. It reads the parquet with a pure-Go reader, computes the split bucketing and the statistics in pkg/modeleval, and runs the ONNX forward pass through libvmaf's own ONNX Runtime session API. The one requirement is that libvmaf was built with DNN support (meson setup … -Denable_dnn=enabled, or the auto default on a host where ONNX Runtime is discoverable via pkg-config). A libvmaf built without it returns:
{"error": "libvmaf was built without DNN support (ONNX Runtime not found at build time); rebuild with -Denable_dnn=enabled against an installed onnxruntime"}
which is the Go counterpart of the Python server's missing-[eval]-extra error.
Numerical parity between the two servers¶
The Python path hands scipy float32 arrays and scipy does not upcast, so its PLCC and RMSE carry float32 rounding. The Go path computes in float64. Expect agreement to roughly 1e-7 on plcc, 1e-9 on rmse, and exact agreement on srocc (rank-based, so width-independent) — not bit-equality. The Go values are the more accurate of the two.
One deliberate behavioural difference: when a split is degenerate (every prediction identical, so the correlation is mathematically undefined), scipy returns NaN and Python emits a bare NaN, which is not valid JSON. The Go server instead fails with an explicit correlation undefined: input has zero variance error.
Input schema¶
| Field | Type | Required | Default |
|---|---|---|---|
model | string (path to .onnx) | yes | — |
features | string (path to .parquet) | yes | — |
split | "train" \| "val" \| "test" \| "all" | no | "test" |
input_name | string — the ONNX graph's input-tensor name | no | "features" |
Feature-column contract¶
The parquet must contain the column mos (ground-truth subjective score). For the input tensor, the server picks whichever of these columns are present:
adm2vif_scale0,vif_scale1,vif_scale2,vif_scale3motion2
At least one must be present, in that order. The ONNX model must accept a float32 tensor of shape [N, K] where K is the number of columns found.
Response body¶
{
"model": "/home/you/dev/vmaf/model/tiny/lpips_sq_small.onnx",
"features": "/home/you/feature-cache/netflix-public.parquet",
"split": "test",
"n": 137,
"plcc": 0.9743,
"srocc": 0.9612,
"rmse": 3.214,
"columns": ["adm2", "vif_scale0", "vif_scale1", "vif_scale2", "vif_scale3", "motion2"]
}
Errors¶
- Bad split name →
{"error": "split must be one of ('train', 'val', 'test', 'all'); got 'foo'"}. - Missing
moscolumn →{"error": "<path> has no 'mos' column — can't score correlations"}. - Missing all feature columns →
{"error": "... has none of the expected feature columns ..."}. - Fewer than 2 samples in the chosen split →
{"error": "split 'test' has N samples — need ≥2 to compute correlations"}. - Model output shape ≠ target shape →
{"error": "model output shape ... does not match target shape ..."}. evalextra not installed →{"error": "eval_model_on_split requires the 'eval' extra: pip install 'vmaf-mcp[eval]'"}.
compare_models¶
Rank several ONNX models on the same parquet split by descending PLCC. Models that fail to load or score are collected under errors instead of aborting the whole call — so the agent can surface partial results.
Input schema¶
| Field | Type | Required | Default |
|---|---|---|---|
models | array of string (paths to .onnx), minItems: 1 | yes | — |
features | string (path to .parquet) | yes | — |
split | "train" \| "val" \| "test" \| "all" | no | "test" |
input_name | string | no | "features" |
Response body¶
{
"ranked": [
{ "model": "/.../baseline_v3.onnx", "plcc": 0.9743, "srocc": 0.9612, "rmse": 3.21, "n": 137, "split": "test", "columns": [ "..." ] },
{ "model": "/.../baseline_v2.onnx", "plcc": 0.9611, "srocc": 0.9503, "rmse": 3.80, "n": 137, "split": "test", "columns": [ "..." ] }
],
"errors": [
{ "model": "/.../broken.onnx", "error": "model output shape (137, 2) does not match target shape (137,)" }
]
}
ranked is sorted descending by plcc. errors preserves the input order for models that failed, with the raised exception serialised as a string.
Errors¶
- Empty or non-list
models→{"error": "'models' must be a non-empty list of paths"}. - Individual model failures show up under the
errorsarray, not as a top-level error.
describe_worst_frames¶
Score a (ref, dis) pair, pick the N frames with lowest VMAF, extract each as PNG via ffmpeg, and describe the visible artefacts with a local vision-language model through ONNX Runtime GenAI. Without a model it returns the frame metadata with a note. It is a debugging aid for an LLM agent that wants narrative context for low-quality regions. Added in ADR-0172; the model runtime is ADR-1886.
Setting up descriptions¶
The Python server only (the Go server always returns metadata):
pip install 'vmaf-mcp[vlm]'installsonnxruntime-genai.- Put an ONNX Runtime GenAI vision model on the machine, a directory with
genai_config.json, the ONNX graphs and the tokenizer. The tested one is the CPU build of Phi-3.5-vision-instruct-onnx (MIT, 3.2 GB), directorycpu_and_mobile/cpu-int4-rtn-block-32-acc-level-4. Verify the download against the SHA-256 the hub lists for each file. Other families ONNX Runtime GenAI supports (Phi-4 multimodal, the Qwen vision models, LFM2-VL, Mistral 3, Gemma 3) get their image placeholder fromvmaf_mcp/vlm.py. - Start the server with
VMAF_MCP_VLM_MODEL=<that directory>.
The server never downloads a model and runs no model code; it reads the directory once, on the first call. With the tested model, a description took 48 s per frame on four CPU cores and peaked at 8.6 GB of memory.
Input schema¶
The arguments whose Tools column names describe_worst_frames in the vmaf_score input schema table: ref, dis, width, height, pixfmt, bitdepth, model (the library default when omitted), backend, and n (1 to 32, default 5).
Behaviour¶
- Run
vmaf_scoreto populate per-frame VMAF. - Pick the
nframes with smallest VMAF. - For each picked frame, run
ffmpeg -f rawvideowithselect='eq(n,<idx>)'and-fps_mode passthroughto emit a single PNG. This needs FFmpeg 5.1 or newer: FFmpeg 5.0 and older do not know-fps_mode, and FFmpeg 9 removed the older-vsync 0spelling. - Describe the PNG with the configured model (greedy decoding, at most 4096 tokens including the prompt). When the
vlmextra is missing,VMAF_MCP_VLM_MODELis unset, or the directory has nogenai_config.json, every frame'sdescriptionis a note that says which, for example(VLM unavailable: VMAF_MCP_VLM_MODEL is not set. ...). - Return frame metadata + descriptions.
The PNGs are written to a temporary directory of the call, which is removed when the call returns; the png paths in the response name where each frame was, not files that still exist.
Response body¶
{
"model_id": "cpu-int4-rtn-block-32-acc-level-4",
"frames": [
{
"frame_index": 12,
"vmaf": 38.4,
"png": "/tmp/vmaf-mcp-worst-12345/frame_000012.png",
"description": "Heavy DCT blocking on the face and ringing along the chin contour."
}
]
}
model_id is the model directory's name, and null when no model ran.
Errors¶
ffmpegnot on PATH →{"error": "ffmpeg not on PATH; install ffmpeg to use describe_worst_frames"}.ffmpegolder than 5.1 → the Python server fails the call withffmpeg frame-extract failed: ... Unrecognized option 'fps_mode'; the Go server (vmafx-mcp) returns the frame with an emptypngand(frame extraction failed: ...)as itsdescription.- Unsupported
pixfmt/bitdepthcombo →{"error": "unsupported pixfmt/bitdepth combo: ..."}. - VMAF subprocess failure → bubbles up the underlying
vmaf_scoreerror. - A configured model that fails to load or to describe a frame fails the call with that error; nothing falls back to metadata silently.
probe_backend¶
Run a 1-frame VMAF health check to distinguish "compiled in" from "driver present and functional". Uses a tiny 64×64 mid-grey synthetic YUV pair so no fixture files are required. The frame is 64×64 (not smaller) because the CUDA ADM kernel requires at least 36px in each dimension; a sub-36px frame silently returns a null score on a CUDA build, which would otherwise be misreported as runtime_healthy: true. A null or non-finite score now sets runtime_healthy: false with an explanatory error string. Added in ADR-0634.
Input schema¶
| Field | Type | Required | Notes |
|---|---|---|---|
backend | "cpu" \| "cuda" \| "sycl" \| "hip" \| "metal" | yes | Backend to health-check |
Response body¶
{
"backend": "cuda",
"compiled_in": true,
"runtime_healthy": true,
"latency_ms": 312.5,
"score": 100.0,
"error": null
}
compiled_in— whethervmaf --list-backendsreports the backend usable (same aslist_backends).runtime_healthy—trueiff the 1-frame score subprocess exits 0 and returns a finite score;falseon driver errors, ICD missing, KFD ioctl failure, etc.latency_ms— wall time for the subprocess in milliseconds;nullwhen the subprocess could not be launched.score— VMAF mean on the synthetic pair (always near 100 for identical ref/dis);nullon failure.error— human-readable failure reason;nullon success.
Errors¶
All failure conditions are reported as runtime_healthy: false with the reason in the error field — the tool itself does not raise (a non-healthy backend is a valid result, not a protocol error).
vmaf_version¶
Return the local vmaf binary's identity and build flags. Useful for confirming which fork build is running before scoring. Added in ADR-0634.
Input schema — no arguments.¶
Response body¶
{
"binary_path": "/usr/local/bin/vmaf",
"version": "1.0.0-rc.2",
"build_flags": {
"cpu": true,
"cuda": true,
"sycl": false,
"hip": false,
"metal": false
},
"error": null
}
version— extracted fromvmaf --versionoutput;nullif the banner cannot be parsed or the subprocess times out.build_flags— derived fromvmaf --list-backends, same probe aslist_backends.error— set when the binary does not exist;nullon success.
Errors¶
- vmaf binary missing →
errorfield is populated; allbuild_flagsarefalse.
vmaf_score_encoded¶
Score a (reference, distorted) pair of encoded video files (MP4, MKV, Y4M, WebM, etc.) by decoding them to raw YUV via ffmpeg, then scoring with the standard vmaf_score pipeline. Geometry (width, height, pixel format, bit depth) is probed automatically from the reference stream — no manual size entry required. Added in ADR-0634.
Requires ffmpeg and ffprobe on PATH.
Input schema¶
The arguments whose Tools column names vmaf_score_encoded in the vmaf_score input schema table: reference_encoded and distorted_encoded (required paths under an allowlisted root) and every optional scoring parameter (model is the library default when omitted), forwarded to the underlying vmaf run (ADR-1117).
Behaviour¶
ffprobethe reference stream to detect width, height, pixel format, and bit depth.- Decode both reference and distorted in parallel to temp raw YUV files via
ffmpeg -f rawvideo. - Call
_run_vmaf_scorewith the decoded YUV pair and the probed geometry. - Inject
reference_encodedanddistorted_encodedkeys into the response (the original encoded paths) alongside the fullvmaf_scorepayload.
Decoded YUV temp files are automatically cleaned up after scoring via a TemporaryDirectory context manager.
Response body¶
Same shape as vmaf_score, plus two extra keys:
{
"version": "1.0.0-rc.2",
"pooled_metrics": { "vmaf": { "mean": 76.667831, "...": "..." } },
"frames": [ "..." ],
"backend_requested": "auto",
"backend_used": "cpu",
"reference_encoded": "/data/corpus/ref.mp4",
"distorted_encoded": "/data/corpus/dis_crf28.mp4"
}
Errors¶
ffprobenot on PATH → raisesRuntimeError.ffmpegnot on PATH → raisesRuntimeError.- No video stream in input → raises
ValueError. - Unknown pixel format → raises
ValueError(e.g.bgr24is not mappable). ffmpegdecode failure → raisesRuntimeErrorwith the ffmpeg stderr.- Same errors as
vmaf_scorefor the scoring step.
list_extractors¶
Enumerate all VmafFeatureExtractor implementations found in the local core/src/feature/ C source tree. No binary required — the server parses the C source directly. Added in ADR-0638.
Input schema — no arguments.¶
Response body¶
{
"extractors": [
{ "name": "float_adm", "backend": "cpu", "source": "core/src/feature/float_adm.c" },
{ "name": "float_vif", "backend": "cpu", "source": "core/src/feature/float_vif.c" },
{ "name": "float_ssim_hip", "backend": "hip", "source": "core/src/feature/hip/float_ssim_hip.c" }
]
}
name is the string the extractor registers as its canonical identifier. backend is inferred from the symbol-name suffix (_cuda, _sycl, _hip, _metal; everything else is cpu).
Errors — none (returns {"extractors": []} if the source tree is absent).¶
describe_model¶
Return metadata for a VMAF model by name or path. Fixes the Path.stem bug (ADR-0638): vmaf_v0.6.1 is matched against vmaf_v0.6.1.json correctly — not incorrectly trimmed to vmaf_v0.6 as Python's Path.stem would do.
Input schema¶
| Field | Type | Required | Notes |
|---|---|---|---|
name | string | yes | Model stem (vmaf_v0.6.1), full filename (vmaf_v0.6.1.json), or repo-relative path. |
Response body¶
{
"name": "vmaf_v0.6.1",
"path": "model/vmaf_v0.6.1.json",
"format": "json",
"size_bytes": 9128,
"model_type": "LIBSVMNUSVR",
"feature_names": ["VMAF_integer_feature_adm2_score", "..."]
}
model_type and feature_names are populated for JSON models; both are null for .pkl and .onnx files.
Errors¶
- Unknown name →
{"error": "model 'foo' not found; run list_models to see available models."}. - Ambiguous name (two models with the same stem in different subdirs) →
{"error": "model name '...' is ambiguous; matched: [...]. Pass an explicit path instead."}.
run_compare¶
Wrap vmaf-tune compare: compare codec adapters at one or more target VMAF scores and return a ranked report. Requires vmaf-tune to be installed (pip install -e tools/vmaf-tune or set VMAF_TUNE_BIN). Emits MCP progress notifications when params._meta.progressToken is set. ADR-0638.
Input schema¶
| Field | Type | Required | Default | Notes |
|---|---|---|---|---|
src | string | yes | — | Source video (any FFmpeg-readable format or raw YUV). |
target_vmaf | number | no | — | Single VMAF target (legacy single-target schema). |
target_vmafs | string | no | "94,96,97,98" | Comma-separated VMAF targets (multi-target schema). |
encoders | string | no | "libx264,libx265,libsvtav1,libvpx-vp9" | Comma-separated encoder list. |
width | integer | no | — | Source width (raw YUV only). |
height | integer | no | — | Source height (raw YUV only). |
pix_fmt | string | no | "yuv420p" | Source pixel format. |
framerate | number | no | — | Source framerate. |
no_parallel | boolean | no | false | Dispatch encoders sequentially. |
Response body¶
Returns the parsed JSON output of vmaf-tune compare --format json. Shape is the v1 (single-target) or v2 (multi-target) schema from ADR-0513.
Errors¶
- vmaf-tune binary missing →
{"error": "vmaf-tune binary not found at ...; Install with: pip install -e tools/vmaf-tune or set VMAF_TUNE_BIN."}. - Non-zero exit →
{"error": "vmaf-tune compare exited <rc>: <stderr>"}.
run_ladder¶
Wrap vmaf-tune ladder: build a per-title bitrate ladder via convex-hull sweep and emit an HLS / DASH / JSON manifest. Requires vmaf-tune. Emits progress notifications. ADR-0638.
Input schema¶
| Field | Type | Required | Default | Notes |
|---|---|---|---|---|
src | string | yes | — | Source video path. |
resolutions | string | yes | — | Comma-separated WxH list, e.g. "1920x1080,1280x720,854x480". |
target_vmafs | string | yes | — | Comma-separated VMAF targets, e.g. "95,90,85". |
encoder | string | no | "libx264" | Codec adapter. |
quality_tiers | integer | no | 5 | Number of ladder rungs to select. |
format | string | no | "json" | "hls", "dash", or "json". |
spacing | string | no | "log_bitrate" | Knee-spacing strategy. |
framerate | number | no | — | Source framerate. |
Response body¶
For format="hls" or "dash", manifest is a raw string (the M3U8 / MPD text).
Errors — same pattern as run_compare.¶
run_tune_per_shot¶
Wrap vmaf-tune tune-per-shot: detect scene cuts and return per-shot CRF recommendations targeting a VMAF score. Requires vmaf-tune. Emits progress notifications. ADR-0638.
Input schema¶
| Field | Type | Required | Default | Notes |
|---|---|---|---|---|
src | string | yes | — | Source video path. |
target_vmaf | number | no | 92.0 | Target VMAF score. |
encoder | string | no | "libx264" | Codec adapter. |
pix_fmt | string | no | "yuv420p" | Source pixel format. |
framerate | number | no | — | Source framerate. |
scene_threshold | number | no | — | Scene-cut detection threshold (0..1). |
output | string | no | — | Output video path (plan-only if omitted). |
format | string | no | "json" | "json", "shell", or "csv". |
Response body¶
Returns the parsed JSON output of vmaf-tune tune-per-shot --format json (list of per-shot recommendations) or the raw shell/CSV string for other formats.
Errors — same pattern as run_compare.¶
Cross-tool error conventions¶
Since ADR-0634:
- All tool handler exceptions are raised, not caught and returned as
TextContent({"error": ...}). - The
mcplibrary's outer handler (_make_error_result) therefore setsisError=Trueon theCallToolResult, and conformant clients (which branch onresult.isError) treat tool errors as errors. - Before ADR-0634
isErrorstayed implicitlyFalse, so clients misclassified errors as successes.
| Situation | MCP-level behavior |
|---|---|
| Unknown tool name | Raises ValueError; mcp sets isError=True |
| Path outside allowlist | Raises ValueError; mcp sets isError=True |
| Path does not exist | Raises FileNotFoundError; mcp sets isError=True |
Subprocess non-zero (vmaf_score) | Raises RuntimeError; mcp sets isError=True |
| Missing optional extras | Raises RuntimeError; mcp sets isError=True |
probe_backend unhealthy backend | Returns success result with runtime_healthy: false |
Sidecar-binary tools¶
Four tools bridge the CLI binaries meson builds next to vmaf in core/tools/. Both servers implement them and build a byte-identical argv (pinned by cmd/vmafx-mcp/sidecar_parity_test.go).
Every numeric and enum bound below is the bound the corresponding C parser enforces, so an out-of-range value fails with a readable MCP error instead of a usage() dump on stderr.
Binary resolution¶
Each tool resolves its binary in this order, and the first hit wins:
- the tool's own environment override,
- a sibling of the resolved
vmafbinary — soVMAF_BINresolves the whole family, which is what thevmaf-dev-mcpcontainer relies on aftermake install, /usr/local/bin/<name>,<repo>/core/build/tools/<name>,<repo>/build/tools/<name>.
| Tool | Binary | Environment override |
|---|---|---|
vmaf_per_shot | vmaf-perShot | VMAF_PER_SHOT_BIN |
vmaf_roi | vmaf_roi | VMAF_ROI_BIN |
vmaf_bench | vmaf_bench | VMAF_BENCH_BIN |
vmaf_vpl | vmaf_vpl | VMAF_VPL_BIN |
When none of the candidates exists the tool returns isError=true naming both the path it looked for and the meson compile target that builds it.
vmaf_per_shot¶
Wrap vmaf-perShot: scan a raw YUV reference, detect shot boundaries from luma complexity + motion energy, and return a per-shot CRF plan targeting a VMAF score. ADR-0222.
Input schema¶
| Field | Type | Required | Default | Notes |
|---|---|---|---|---|
reference | string (path) | yes | — | Raw planar YUV; must be under an allowed root. |
width | integer 16..65535 | yes | — | Frame width. |
height | integer 16..65535 | yes | — | Frame height. |
pixel_format | "420" \| "422" \| "444" | no | "420" | --pixel_format. |
bitdepth | 8 \| 10 \| 12 \| 16 | no | 8 | --bitdepth. |
target_vmaf | number 0..100 | no | 90 | --target-vmaf. |
crf_min | integer 0..63 | no | 18 | --crf-min; must not exceed crf_max. |
crf_max | integer 0..63 | no | 35 | --crf-max. |
diff_threshold | number 0..255 | no | — | --diff-threshold. Omitted from argv when unset, so the C default (12.0) stays authoritative. |
format | "json" \| "csv" | no | "json" | --format. The MCP tool defaults to json where the C CLI defaults to csv, so the plan comes back structured. |
Response body¶
{
"format": "json",
"exit_code": 0,
"stderr": "vmaf-perShot: wrote 2 shot(s) to -\n",
"plan": {
"target_vmaf": 92.0,
"crf_min": 18,
"crf_max": 35,
"shots": [
{"shot_id": 0, "start_frame": 0, "end_frame": 3, "frames": 4,
"mean_complexity": 0.000051, "mean_motion": 0.020046, "predicted_crf": 25.31}
]
}
}
With format="csv" the plan key is replaced by output, the raw CSV text.
Errors¶
isError=true on: a reference outside the allowlist, any out-of-range argument, crf_min > crf_max, a missing binary, a non-zero exit (vmaf-perShot exited N: <stderr>), or JSON the plan writer produced that does not parse.
vmaf_roi¶
Wrap vmaf_roi: compute a per-CTU saliency grid for one frame of a raw YUV file and emit an encoder ROI sidecar.
Input schema¶
| Field | Type | Required | Default | Notes |
|---|---|---|---|---|
reference | string (path) | yes | — | Raw planar YUV, under an allowed root. |
width | integer 1..16384 | yes | — | Frame width. |
height | integer 1..16384 | yes | — | Frame height. |
frame | integer 0..1000000 | yes | — | 0-based frame index (--frame). |
pixel_format | "420" \| "422" \| "444" | no | "420" | |
bitdepth | 8 \| 10 \| 12 \| 16 | no | 8 | |
ctu_size | integer 8..128 | no | 64 | --ctu-size (x265 max-ctu). |
encoder | "x265" \| "svt-av1" | no | "x265" | Sidecar dialect. |
strength | number 0..64 | no | 6.0 | QP-offset gain (--strength). |
saliency_model | string (path) | no | — | ONNX [1,1,H,W] luma→[0,1] model, under an allowed root. Without it a centre-weighted radial placeholder is used — smoke-test quality only. |
Response body¶
The sidecar is always written to a server-owned temp file, never to stdout: the SVT-AV1 emitter produces a raw int8 grid that is not text-safe. The response carries the bytes in the encoding that matches the emitter.
{
"encoder": "x265",
"sidecar_fmt": "qpfile",
"exit_code": 0,
"stderr": "",
"ctu_size": 64,
"frame": 2,
"bytes": 216,
"grid_cols": 9,
"grid_rows": 6,
"saliency": "placeholder",
"qpfile": "# vmaf-roi qpfile (x265, --qpfile-style)\n# frame=2 ctu=64 cols=9 rows=6 strength=6.000\n4 2 1 -1 -1 -1 1 2 4\n…"
}
For encoder="svt-av1", sidecar_fmt is "roi_map_int8" and qpfile is replaced by roi_map_base64 — the base64 of the raw row-major int8 grid. saliency is "onnx" when a saliency_model was supplied, "placeholder" otherwise.
Errors — same pattern as vmaf_per_shot.¶
vmaf_bench¶
Wrap vmaf_bench: a per-feature micro-benchmark over the built-in synthetic fixtures, or a GPU-vs-CPU correctness comparison.
This is not run_benchmark: that one runs the end-to-end bench_all.sh harness over real YUV fixtures across every backend and takes no arguments (ADR-0513 / ADR-0517). vmaf_bench is the per-feature timing tool.
Input schema¶
| Field | Type | Required | Default | Notes |
|---|---|---|---|---|
frames | integer 2..48 | no | — | --frames; the C default is 10 and 48 is MAX_TEST_FRAMES. Out-of-range values are rejected rather than silently clamped as the C parser does. |
resolution | "576x324" \| "640x480" \| "1280x720" \| "1920x1080" \| "3840x2160" | no | — | --resolution. Omit to test all five. |
bpc | 8 \| 10 \| 12 \| 16 | no | — | --bpc; C default 8. |
data_dir | string (path) | no | — | --data-dir. Must be a directory under an allowed root. Otherwise VMAF_TEST_DATA applies. |
validate | boolean | no | false | --validate: compare GPU vs CPU scores. |
gpu_only | boolean | no | false | --gpu-only: skip the CPU targets. |
device_list | boolean | no | false | --list-devices: list GPU devices and exit. |
Response body¶
{
"mode": "benchmark",
"exit_code": 0,
"stdout": "VMAF Performance Benchmark (…)\nFeature Res Init ms Avg ms …",
"stderr": ""
}
In validate mode the payload additionally carries "validation_failed": true|false and "mode": "validate".
Errors¶
In benchmark mode a non-zero exit is a tool error. In validate mode it is not: vmaf_bench --validate exits 1 to report that the GPU/CPU comparison found deltas, which is a result rather than a failure, so the call succeeds with validation_failed=true. Argument validation errors and a missing binary are always tool errors.
vmaf_vpl¶
Wrap vmaf_vpl: decode an encoded (reference, distorted) pair with Intel oneVPL, import the VA surfaces zero-copy into SYCL over DMA-BUF, and score them.
The binary is only built when the oneVPL + libva + SYCL toolchain is present. On a build without it the tool returns isError=true naming the missing binary — it never silently falls back to another path.
Input schema¶
| Field | Type | Required | Default | Notes |
|---|---|---|---|---|
ref | string (path) | yes | — | Reference encoded video, under an allowed root. |
dis | string (path) | yes | — | Distorted encoded video. |
model | string | no | "vmaf_v0.6.1" | --model. A bare model name; a value containing /, \, a space or a tab is rejected. The default mirrors the sidecar's own. |
frames | integer ≥ 0 | no | 0 | --frames; 0 = all. |
device | integer ≥ 0 | no | 0 | --device (SYCL device index). |
render_node | string | no | "/dev/dri/renderD128" | --render-node. Restricted to /dev/dri/renderD<N> or /dev/dri/card<N> — the value goes straight to open(2), so anything else is rejected. |
fallback | boolean | no | false | --fallback: host upload when the zero-copy import fails. |
Response body¶
{
"exit_code": 0,
"stdout": "Frames: 12\nTime: 0.400 s (30.0 FPS)\nVMAF: 96.123456 (mean)\n…",
"stderr": "",
"model": "vmaf_v0.6.1",
"device": 0,
"render_node": "/dev/dri/renderD128",
"vmaf_score": 96.123456,
"frames_processed": 12
}
vmaf_score and frames_processed are parsed out of the sidecar's summary lines; both are omitted when the sidecar did not print them (for example when the model failed to load).
Control-plane tools (gRPC bridge, Go only)¶
Five tools expose the Phase-4b control plane: the vmafx-controller job API and the vmafx-server scoring API. They exist only in the Go server (vmafx-mcp); the Python server does not ship a gRPC stack. See ADR-1184 for why, and docs/architecture/phase4b-distributed-platform.md for the topology.
Configuration¶
Connection targets are environment-only — a tool argument naming a host would turn the MCP server into an SSRF pivot.
The addresses, the bearer token and the RPC deadline (VMAFX_CONTROLLER_ADDR, VMAFX_SERVER_ADDR, VMAFX_CONTROLLER_TOKEN, VMAFX_GRPC_TIMEOUT) are rows of the Go server's environment table. With auth enabled the token needs vmafx:reader for get_job and list_jobs, and vmafx:writer for submit_job and cancel_job (roles); omit it against a controller started with VMAFX_AUTH_DISABLED=true.
The transport is insecure by default, matching pkg/score.Dial: the control plane is expected to run inside the cluster mesh.
Path arguments are resolved remotely¶
reference / distorted on these tools are not checked against the MCP host's allowlist — the file lives on the worker node's shared mount, so VMAF_MCP_ALLOW does not apply and the file need not exist locally. What is enforced is the shape: the value must be absolute, must contain no .. component, and must contain no NUL / CR / LF.
submit_job¶
Enqueue a scoring job and return its ID.
| Field | Type | Required | Default | Notes |
|---|---|---|---|---|
reference | string (absolute worker-side path) | yes | — | |
distorted | string (absolute worker-side path) | yes | — | |
model | string | no | — | Omit to let the controller apply its own default. |
backend | "auto" \| "cpu" \| "cuda" \| "sycl" \| "hip" \| "metal" | no | "auto" | Node capability the scheduler must match. auto is spelled as the empty string on the wire. |
{
"job_id": "555090fd-1349-471f-8a07-751f37e162cc",
"status": "PENDING",
"controller": "127.0.0.1:9090",
"scoring": {"reference": "/mnt/data/ref.yuv", "distorted": "/mnt/data/dis.yuv",
"model": "", "backend": "cpu"}
}
get_job¶
Fetch a job by ID (job_id, required).
{
"id": "555090fd-…", "status": "CANCELLED", "assigned_node": "", "error": "",
"created_at": 1788622781, "updated_at": 1788622781, "final_score": 0,
"scoring": {"reference": "…", "distorted": "…", "model": "", "backend": "cpu"},
"controller": "127.0.0.1:9090"
}
status is one of PENDING, RUNNING, COMPLETED, FAILED, CANCELLED. partial_results and partial_result_count appear only when the node has reported per-frame partials. An unknown ID is a tool error carrying the controller's NotFound status.
cancel_job¶
Request cancellation of a PENDING or RUNNING job (job_id, required). Returns {"job_id": …, "ok": true, "message": "cancellation requested", "controller": …}. A running job's node stops it within one heartbeat interval (10 s by default, controller guide). ok=false means the controller declined; the message says why.
list_jobs¶
List the controller's current job snapshot: the jobs of the token's tenant (VMAFX_CONTROLLER_TOKEN), or of the dev tenant against a controller with auth disabled.
| Field | Type | Required | Default | Notes |
|---|---|---|---|---|
status_filter | array of status strings | no | — | Any of PENDING, RUNNING, COMPLETED, FAILED, CANCELLED; case-insensitive. Omit for all jobs. |
limit | integer 1..500 | no | 100 | Maximum jobs returned. |
Each entry has the get_job shape. The tool drains the StreamJobs server-streaming RPC, which sends the current snapshot and closes (ADR-0962) — it is not a subscription. truncated=true means the controller had more of the tenant's jobs than limit.
vmaf_score_remote¶
Score a pair on a remote vmafx-server over the unary VmafxScoring.Score RPC. Nothing is read locally.
| Field | Type | Required | Default | Notes |
|---|---|---|---|---|
reference | string (absolute server-side path) | yes | — | |
distorted | string (absolute server-side path) | yes | — | |
model | string | no | — | Omit to let the server apply its own default. |
{"score": 76.6683, "features": {"vif_scale0": 0.79, "adm2": 0.93},
"model": "", "reference": "/mnt/data/ref.yuv", "distorted": "/mnt/data/dis.yuv",
"server": "127.0.0.1:9090"}
Use vmaf_score for files on the MCP host.
Progress notifications¶
All four run_* tools — run_benchmark, run_compare, run_ladder, and run_tune_per_shot — emit notifications/progress when the client supplies a progressToken in the request's _meta object:
{
"method": "tools/call",
"params": {
"name": "run_compare",
"arguments": { "src": "/path/to/video.mp4" },
"_meta": { "progressToken": "my-token-42" }
}
}
The server sends two progress events per tool call:
| Event | progress | total | message |
|---|---|---|---|
| Start | 0.0 | 1.0 | "starting vmaf-tune compare" (tool-specific) |
| Completion | 1.0 | 1.0 | "vmaf-tune compare done" |
No finer-grained progress is available because the tools delegate to a subprocess. Clients without a token receive no progress events (per MCP spec — the server must not send unsolicited progress).
Related¶
- MCP server overview — install, security model, env vars.
- CLI reference — the CLI that
vmaf_scorewraps. vmaf_bench— whatrun_benchmarkdrives.- Tiny-AI inference — what
eval_model_on_split/compare_modelsare scoring. - ADR-0634 — P0 fixes.
- ADR-0100.