DNN session API — libvmaf/dnn.h¶
The DNN surface in core/include/libvmaf/dnn.h lets callers load and run tiny ONNX models from C, either attached to a VmafContext (so DNN scores show up next to SVM scores in the normal VMAF report) or as a standalone session (luma-in / luma-out filter-style inference, no VmafContext required).
This is the runtime half of the tiny-AI surface; training lives in ai/. See ADR-0022 (ORT as the inference runtime) and ADR-0023 (the four user surfaces: CLI, C API, ffmpeg, training).
Availability check¶
Returns 1 if libvmaf was built with DNN support (the enable_dnn Meson feature option, enabled or auto with ONNX Runtime found) and ONNX Runtime is linked, 0 otherwise. It is the cheap way to branch between DNN and classic-only builds at run time. It is safe to call from any thread.
When it returns 0, the other entry points in dnn.h return -ENOSYS (except vmaf_dnn_is_codec_aware(), which returns 0, and vmaf_dnn_session_attached_ep(), which returns NULL). The header is installed only when enable_dnn is enabled or auto.
All entry points are not thread-safe unless stated; see thread-safety.
Device config — VmafDnnConfig¶
typedef enum VmafDnnDevice {
VMAF_DNN_DEVICE_AUTO = 0,
VMAF_DNN_DEVICE_CPU = 1,
VMAF_DNN_DEVICE_CUDA = 2,
VMAF_DNN_DEVICE_OPENVINO = 3, /* OpenVINO GPU with CPU fallback */
VMAF_DNN_DEVICE_ROCM = 4,
VMAF_DNN_DEVICE_COREML = 5,
VMAF_DNN_DEVICE_COREML_ANE = 6,
VMAF_DNN_DEVICE_COREML_GPU = 7,
VMAF_DNN_DEVICE_COREML_CPU = 8,
VMAF_DNN_DEVICE_OPENVINO_NPU = 9,
VMAF_DNN_DEVICE_OPENVINO_CPU = 10,
VMAF_DNN_DEVICE_OPENVINO_GPU = 11,
} VmafDnnDevice;
typedef struct VmafDnnConfig {
VmafDnnDevice device;
int device_index; /* multi-GPU index; 0 for single-GPU/CPU */
int threads; /* CPU EP intra-op threads; 0 = ORT default */
bool fp16_io; /* request fp16 tensors when supported */
} VmafDnnConfig;
device | Execution provider (EP) tried | Fallback |
|---|---|---|
AUTO | CUDA, then OpenVINO GPU, then ROCm, then CoreML | CPU. OpenVINO NPU is explicit-only: small graphs can pay a noticeable NPU power-state latency floor. |
CPU | ORT CPU EP | none |
CUDA | CUDAExecutionProvider, when the linked ORT exports it | CPU if the EP append fails; the session still opens |
OPENVINO | OpenVINO device_type=GPU | OpenVINO device_type=CPU, then CPU |
OPENVINO_NPU, _CPU, _GPU | OpenVINO pinned to one device_type (NPU, CPU, GPU), no OpenVINO fallback | CPU when the EP or the silicon is missing |
ROCM | ROCMExecutionProvider | CPU |
COREML | CoreMLExecutionProvider, CoreML chooses compute units (ALL) | CPU on non-Apple ORT builds |
COREML_ANE | CoreML with MLComputeUnits = CPUAndNeuralEngine | CPU on non-Apple ORT builds |
COREML_GPU | CoreML with CPUAndGPU | CPU on non-Apple ORT builds |
COREML_CPU | CoreML with CPUOnly | CPU on non-Apple ORT builds |
Other fields:
device_index: multi-GPU index; 0 for a single GPU or CPU.threads: CPU EP intra-op threads.0lets ORT pick; set it explicitly when pinning affinity or benchmarking.fp16_io: stages fp32 to fp16 for model slots declaredFLOAT16; OpenVINO also receivesprecision=FP16. It does not change float32 slots, so it is not a speed switch for fp32-only graphs.
EP choice is a preference, not a requirement: when a requested provider is missing from the linked ORT build, session open falls back to CPU. Check which EP bound if your application must fail instead.
Pass NULL for cfg in any function that accepts one to use VMAF_DNN_DEVICE_AUTO with zero device index, zero threads, and no fp16 I/O.
Attached mode — vmaf_use_tiny_model¶
Register a tiny ONNX model on a live VmafContext. The model participates in the per-frame pipeline; its outputs appear in the report alongside SVM scores. Use this when you want "VMAF + tiny AI score" in the same run.
A feature-vector model (rank-2 input) reads libvmaf features (ADR-1520):
- Attaching registers the extractors that write the features its sidecar names (
feature_orderorfeatures; canonical-6 for a six-wide model without a sidecar list), with their default options. - Its scores are computed when the context is flushed (
vmaf_read_pictures(ctx, NULL, NULL, 0)), so read them after the flush. - A frame that lacks one of the inputs makes the flush return
-EINVAL, and the log names the feature and the frame. Frames before it keep their scores; framesn_subsampledrops are not scored.
An image model (rank-4 input) is scored as each frame is read.
Returns:
0— success.-ENOSYS— built without DNN support.-EINVAL— bad args (nullctxoronnx_path).-ENOENT—onnx_pathdoes not exist or is not a regular file.-E2BIG— file exceeds the compile-time 50 MB cap (VMAF_DNN_DEFAULT_MAX_BYTES— defence against adversarial bloat; see ADR-0039). The historicalVMAF_MAX_MODEL_BYTESenv override was retired in T7-12.-ENOMEM— session allocation failed (ORT env, session options, or internal buffer allocation).- Negative
errnofrom the operator-allowlist walk if the model contains a banned op. -ENOTSUP— a feature-vector model whose sidecar names another number of features than the input is wide, or whose second input does not match the sidecar'sencoder_vocab.-EINVAL— a feature-vector model whose sidecar names a feature no extractor writes.
Equivalent CLI flag: --tiny-model <path> (usage/cli.md).
Attached scores are written to the normal feature collector:
- A single-output model preserves the historical key: the sidecar
namefield (orvmaf_tiny_modelwhen no sidecar name exists). - A multi-output model emits one scalar score per graph output. The key is
<sidecar-name>_<output-name>, whereoutput-namecomes from sidecaroutput_names[]when the array length matches the ONNX output count, or from the ONNX graph output name otherwise. The suffix is sanitized to[A-Za-z0-9_]; duplicate sanitized suffixes fall back to deterministicoutput<slot>_<attempt>keys.
Codec-aware tiny-model inputs — vmaf_dnn_set_codec_context¶
Codec-conditioned tiny models (e.g. the v2 ladder regressor) accept a small categorical block alongside the per-frame features: encoder identity, preset ordinal, and CRF / QP. vmaf_dnn_set_codec_context populates that block on the attached tiny model so the loop body does not need to re-supply it per frame.
Warning
Call it after vmaf_use_tiny_model() and before the first vmaf_read_pictures(). A codec-aware model does not score until this call has returned 0: vmaf_read_pictures() returns -EINVAL for it and logs tiny model <name> is codec-aware: ... (ADR-1520).
int vmaf_dnn_set_codec_context(VmafContext *ctx,
const char *codec_name,
const char *preset,
int crf);
Parameters¶
| Parameter | Notes |
|---|---|
ctx | Context with a tiny model already attached via vmaf_use_tiny_model(). |
codec_name | Encoder name (libx264, libx265, libsvtav1, libvpx-vp9, h264_nvenc, ...). NULL or "" selects the vocabulary's "unknown" entry; a vocabulary without one (fr_regressor_v3) returns -ENOENT. ffprobe aliases (h264, hevc, av1, vp9, vvc) map to their canonical encoder names. |
preset | Preset string (medium, slow, p4, 5, ...), looked up in a per-encoder ordinal table. NULL or an unknown preset defaults to ordinal 5 (mid-tier). |
crf | CRF / QP integer; clamped to [0, 63]. |
Returns¶
| Code | Meaning |
|---|---|
0 | Codec block written (a named encoder, or the vocabulary's "unknown" entry for NULL / ""). |
-ENOENT | codec_name is not in the model's encoder_vocab, or is NULL / "" and the vocabulary has no "unknown" entry. The model will not score. |
-ENOSYS | libvmaf was built without DNN support (-Denable_dnn=disabled). |
-EINVAL | ctx is NULL or no tiny model is attached. |
-ENOTSUP | The attached model has no codec block (rank-4 image model or rank-2 single-input model). |
Equivalent CLI flags (the vmaf CLI calls this internally): --tiny-codec, --tiny-preset, --tiny-crf — see usage/cli.md. Per-codec / per-preset vocabularies live in the model's sidecar JSON under encoder_vocab and preset_vocab; the loader bakes them into the runtime descriptor at vmaf_use_tiny_model() time. The block layout is [encoder_onehot(N_VOCAB), preset_norm, crf_norm].
Detect a codec-aware model¶
Returns 1 when the tiny model attached to ctx requires a codec context (its sidecar declares "codec_aware": true and it was loaded with a codec_block second input). Returns 0 when no model is attached, the model has no codec block, ctx is NULL, or libvmaf was built without DNN. Call it after vmaf_use_tiny_model() to emit a clear error before inference when --tiny-codec style conditioning is required. It is safe to call before vmaf_read_pictures().
Tiny-model auto-resize — vmaf_dnn_set_resize_mode¶
NCHW tiny models declare a fixed input shape at training time (e.g. 224×224 for the nr_metric_v1 NR scorer). When the user-supplied frame dims don't match, the per-frame dispatch resamples the luma plane to the model dims using the selected filter before invoking ONNX Runtime. Bit-exact when source dims already equal model dims (the routine forwards to vmaf_tensor_from_luma unchanged). Introduced in ADR-0550.
typedef enum VmafDnnResizeMode {
VMAF_DNN_RESIZE_DISABLED = 0, /* default; mismatch → -ERANGE */
VMAF_DNN_RESIZE_BILINEAR = 1, /* OpenCV INTER_LINEAR / torchvision */
VMAF_DNN_RESIZE_NEAREST = 2, /* nearest, floor coord */
VMAF_DNN_RESIZE_BICUBIC = 3, /* Catmull-Rom (a = -0.5) */
} VmafDnnResizeMode;
int vmaf_dnn_set_resize_mode(VmafContext *ctx, VmafDnnResizeMode mode);
Filter semantics¶
| Mode | Equivalent | When to use |
|---|---|---|
DISABLED | None — size mismatch returns -ERANGE | Parity harnesses; strict-mode pipelines (default). |
BILINEAR | torchvision Resize(..., antialias=False) / OpenCV INTER_LINEAR | Every shipped NR / image-input tiny-AI model was trained against this. |
NEAREST | OpenCV INTER_NEAREST; deterministic floor of source coord | Cheaper; debugging dispatch without a filter parameter. |
BICUBIC | Separable Catmull-Rom (a = -0.5); torchvision BICUBIC | Parity with exporters that used transforms.Resize(interpolation=BICUBIC). |
The three filter modes produce scores that differ by approximately 2% on the same input — treat filter choice as a model hyperparameter and document it alongside the model checkpoint.
Resize-mode returns¶
| Code | Meaning |
|---|---|
0 | Resize mode updated; takes effect on the next vmaf_read_pictures(). |
-EINVAL | ctx is NULL, or mode is outside the enum range. |
-ENOSYS | libvmaf was built without DNN support. |
Equivalent CLI flag: --tiny-resize <bilinear|nearest|bicubic|disabled> — see usage/cli.md. May be called before or after vmaf_use_tiny_model(); the setting is sticky for the lifetime of the context.
Standalone sessions — VmafDnnSession¶
Standalone mode is for filter-style inference that does not need a VmafContext — e.g. a learned de-banding preprocessor that mutates a luma plane before downstream processing.
typedef struct VmafDnnSession VmafDnnSession;
int vmaf_dnn_session_open (VmafDnnSession **out, const char *onnx_path, const VmafDnnConfig *cfg);
void vmaf_dnn_session_close(VmafDnnSession *sess);
vmaf_dnn_session_open() returns 0, -ENOSYS (no DNN support), -EINVAL (NULL out or onnx_path), -ENOENT (file missing), -E2BIG (over the 50 MB cap) or -EIO (ONNX Runtime failure). cfg may be NULL for VMAF_DNN_DEVICE_AUTO. vmaf_dnn_session_close() accepts NULL, tears down the ORT session and its binding caches, and invalidates any pointer from vmaf_dnn_session_attached_ep(). Pair every successful open with one close.
Both vmaf_dnn_session_open and vmaf_use_tiny_model apply the same size-cap + operator-allowlist walk. See ADR-0039 for the allowlist and ADR-0041 for an example of an extractor that uses a session under the hood.
Luma-only convenience call¶
int vmaf_dnn_session_run_luma8(VmafDnnSession *sess,
const uint8_t *in, size_t in_stride,
int w, int h,
uint8_t *out, size_t out_stride);
Runs one luma-in / luma-out pass. Only works when the ONNX graph has:
- exactly one float32 input of static shape
[1, 1, H, W], - exactly one output of the same shape.
The implementation:
- Reads
in(uint8 luma), normalises to[0, 1](applies mean/std from the sidecar JSON if present). - Runs ORT.
- De-normalises, rounds, clamps to
[0, 255], writesout.
Errors:
-EINVAL—sess,in, oroutis NULL.-ENOTSUP— graph shape isn't the supported NCHW[1,1,H,W]luma layout, or ORT returned fewer output elements thanw*h.-ERANGE—w/hdon't match the graph's static input shape. Usevmaf_dnn_session_run()for dynamic shapes.-ENOSYS— libvmaf was built without DNN support (enable_dnndisabled); every DNN entry point returns-ENOSYSin that configuration.
10/12/16-bit convenience call¶
int vmaf_dnn_session_run_plane16(VmafDnnSession *sess,
const uint16_t *in, size_t in_stride,
int w, int h, int bpc,
uint16_t *out, size_t out_stride);
The bit-depth-extended sibling of _luma8. Used by the ffmpeg vmaf_pre filter for yuv420p10le / yuv422p10le / yuv444p10le (and the 12-bit LE counterparts), and — at any supported bit depth — to run the same session on chroma planes at their sub-sampled dimensions. Added in ADR-0170 (T6-4).
in/outare packeduint16little-endian single-plane buffers.in_stride/out_strideare in bytes (not samples) — same convention as_luma8, so a 10-bit 1920×1080 plane hasstride ≥ 1920 * 2.bpcin range 9..16 selects the normalisation divisor(1 << bpc) - 1. Passingbpc=8returns-EINVAL— use_luma8for 8-bit input.
The model must still declare [1, 1, H, W] static shape; the only new freedom is the bit depth of the host-side buffer the loader normalises from. A single learned_filter_v1 session works for both luma and chroma — re-call with chroma W/H (the shape is declared dynamic, see the open() comment).
Errors match _luma8, plus -EINVAL for a bpc outside [9, 16].
General named-binding call¶
For models with multiple inputs / outputs or non-luma shapes, use the general call:
typedef struct VmafDnnInput {
const char *name; /* bind by graph name; NULL = positional */
const float *data; /* row-major float32 */
const int64_t *shape; /* rank dims */
size_t rank;
} VmafDnnInput;
typedef struct VmafDnnOutput {
const char *name; /* bind by graph name; NULL = positional */
float *data; /* caller-owned */
size_t capacity; /* element count allocated */
size_t written; /* OUT: element count produced */
} VmafDnnOutput;
int vmaf_dnn_session_run(VmafDnnSession *sess,
const VmafDnnInput *inputs, size_t n_inputs,
VmafDnnOutput *outputs, size_t n_outputs);
Name-binding (name != NULL) resolves by the ONNX graph's declared input / output names. Positional binding (name == NULL) uses the tensor's array index. Mix is allowed but discouraged — pick one style per session.
Errors:
-ENOSYS— built without DNN support.-EINVAL— mismatched arity, null pointers, rank zero.-ENOMEM— allocation failure (per-input staging buffer, or tensor creation).-ENOSPC— someoutputs[i].capacityis smaller than the produced tensor. On this return,outputs[i].writtenis populated with the required element count (the code setswritten = producedbefore the capacity check), so the caller can resize and retry with the same bindings.-EIO— ORT failure (bad graph, EP crash, OOM on device). The diagnostic is logged via theVmafContextlog callback if one is configured (for sessions opened without aVmafContext, logging goes through the library's global log sink).
See ADR-0040 for the rationale behind multi-input/output + named binding.
Which execution provider bound: vmaf_dnn_session_attached_ep¶
Returns the ONNX Runtime EP that actually bound to the session, for diagnostics and for asserting AUTO-chain behaviour. The strings are "CPU", "CUDA", "ROCm", "CoreML", "CoreML:ANE", "CoreML:GPU", "CoreML:CPU", "OpenVINO:CPU", "OpenVINO:GPU" and "OpenVINO:NPU". It returns NULL when sess is NULL or libvmaf has no DNN support. The string is owned by the session and invalid after vmaf_dnn_session_close(). It is safe from any thread once the session is open and no inference is running on it.
Thread-safety¶
- A single
VmafDnnSessionis not re-entrant. Driving inference from two threads requires either per-thread sessions or external locking. - Opening multiple sessions concurrently is safe; they do not share state beyond process-global ORT singletons.
- Attaching a tiny model via
vmaf_use_tiny_model()is subject to the same single-driver rule as the rest of theVmafContextAPI — see the overview. vmaf_dnn_available()andvmaf_dnn_verify_signature()are safe from any thread.
Runnable example — standalone luma filter¶
#include <errno.h>
#include <stdio.h>
#include <stdint.h>
#include <string.h>
#include <libvmaf/dnn.h>
int main(int argc, char **argv)
{
if (!vmaf_dnn_available()) {
fprintf(stderr, "libvmaf was built without DNN support\n");
return 2;
}
VmafDnnSession *sess = NULL;
VmafDnnConfig cfg = { .device = VMAF_DNN_DEVICE_AUTO };
int err = vmaf_dnn_session_open(&sess, argv[1], &cfg);
if (err < 0) {
fprintf(stderr, "open failed: %d (%s)\n", err, strerror(-err));
return 1;
}
const int W = 1920, H = 1080;
uint8_t *in = malloc((size_t)W * H);
uint8_t *out = malloc((size_t)W * H);
/* ...fill `in` with luma from your pipeline... */
err = vmaf_dnn_session_run_luma8(sess, in, W, W, H, out, W);
if (err < 0) fprintf(stderr, "run failed: %d\n", err);
free(in); free(out);
vmaf_dnn_session_close(sess);
return err < 0 ? 1 : 0;
}
Build:
Needs a libvmaf built with DNN support (-Denable_dnn=enabled).
Sigstore signature verification: vmaf_dnn_verify_signature¶
Verify a tiny model's Sigstore bundle before loading it. The same check is behind the vmaf CLI flag --tiny-model-verify (CLI usage).
| Argument | Meaning |
|---|---|
onnx_path | Path to the model file. Its basename is the registry key. |
registry_path | Explicit registry file, or NULL to use registry.json next to the model. |
The function:
- Reads the registry and finds the entry for the model's basename.
- Reads that entry's
sigstore_bundlefield and resolves the bundle path. - Runs
cosign verify-blobthroughposix_spawnp, pinned to the identity regexphttps://github.com/VMAFx/vmafx/.github/workflows/.+and the issuerhttps://token.actions.githubusercontent.com. - Returns
0only when cosign exits0. It fails closed: any error stops model load.
| Return | Meaning |
|---|---|
0 | Verification passed. |
-ENOENT | Registry missing, bundle missing, or no entry for the model. |
-EACCES | cosign is not on PATH. |
-EPROTO | cosign exited non-zero (bad or non-matching signature). |
-ENOSYS | Windows build: the supply-chain path needs POSIX process spawning. |
-EINVAL | onnx_path is NULL. This check runs before the platform check. |
There is no error-text out parameter; the CLI prints the negated errno. Any cosign diagnostics go to the process's own stderr.
CLI coupling
With --tiny-model-verify the CLI calls this function itself, before vmaf_use_tiny_model(), and refuses to load the model on a nonzero result (--tiny-model-verify: signature verification failed for <path> (errno N)). vmaf_use_tiny_model() does not verify by itself.
Provenance: the bundle comes from the fork's release-please and Sigstore signing pipeline (ADR-0010, and the policy in the model registry). Bundles under model/tiny/ carry a keyless OIDC identity tied to the GitHub Actions workflow that built the model. The function is part of the model-registry work in ADR-0211.
Testing standalone sessions¶
Run the session and ORT-internals tests with:
python3 "$(git rev-parse --show-toplevel)/scripts/ci/run_meson_test.py" -- \
-C build --print-errorlogs test_ort_internals test_dnn_session_api
- Use a build configured with
-Denable_dnn=enabledand a discoverable ONNX Runtime package to exercise inference and session creation. A disabled-DNN build exercises only the stub semantics, so that result alone does not validate inference. - The POSIX fixture helper checks that source read errors and destination close or flush failures cannot be reported as successful copies. The close-error case runs first, before ORT initialisation, and confines its file-size limit and signal disposition to a child process. The read-error case stays last.
Known limitations¶
- Attached mode supports multiple scalar output tensors, not vector or image output tensors.
vmaf_use_tiny_model()records every ONNX output when each output tensor contains exactly one scalar value. If any attached output tensor has more than one element, the frame run returns-ENOTSUP. Use the standalonevmaf_dnn_session_run()API for caller-owned vector/image output buffers. See ADR-0646. -
Attached mode caps routed outputs at eight tensors. This mirrors
VMAF_ORT_MAX_IO, the existing ORT wrapper stack-array limit. Models with more outputs should collapse related values into a standalone session output or ship a future ADR that raises the cap across the ORT wrapper and sidecar parser together. -
Operator allowlist covers the set required by tiny FR / NR / filter models shipped in
model/tiny/; untrusted models with new op types will be rejected at_open. Extend the allowlist via the registry — see ADR-0039. - EP selection is a preference, not a requirement (see which EP bound).
- There is no callback / progress hook; inference is synchronous per call.
- Sessions are heap-only; no stack-allocated variant.
Related¶
- ADR-0022 — choice of ONNX Runtime.
- ADR-0023 — where the CLI / C API / ffmpeg / training surfaces intersect.
- ADR-0036 — Wave 1 scope (LPIPS, saliency, per-shot CRF,
vmaf_post, allowlistLoop/If, MCP VLM). - ADR-0039 — operator allowlist + model registry.
- ADR-0040 — multi-input/output named-binding API.
- ADR-0041 — LPIPS-SqueezeNet extractor (consumer of this API).
- ADR-0042 — tiny-AI doc specialisation.
- ../ai/inference.md — CLI-side tiny-AI walkthrough.