GPU backends C API¶
Use this page to add a GPU backend to a libvmaf context from C. Each backend (CUDA, SYCL, HIP, Metal) adds a small header on top of libvmaf.h: a state object, device selection, and backend-specific picture handling. The Vulkan backend was removed (ADR-0726); no Vulkan entry point exists.
Related pages: core API overview, CLI backend selection, backend dispatch rules.
When these headers apply¶
| Backend | Header | Build options | Without the option |
|---|---|---|---|
| CUDA | libvmaf_cuda.h | -Denable_cuda=true | Header not installed; the symbols do not exist. |
| SYCL | libvmaf_sycl.h | -Denable_sycl=true | Header not installed; the symbols do not exist. |
| HIP | libvmaf_hip.h | -Denable_hip=true for the API; -Denable_hipcc=true for device kernels | Header not installed; the library still exports stubs that return -ENOSYS. |
| Metal | libvmaf_metal.h | -Denable_metal=auto (default) or enabled, on macOS | Header installed when Metal is enabled or auto; entry points return -ENOSYS when the build has no Metal, -ENODEV on an unsupported device. |
Registered feature extractors per backend, from the extractor registry:
| Backend | Extractors | Where the list is |
|---|---|---|
| HIP | 19 | HIP overview |
| Metal | 17 (every extractor except SpEED) | Metal backend |
Note
libvmaf does not export HAVE_CUDA, HAVE_HIP and similar macros through pkg-config. To write code that builds against any libvmaf, probe at build time in your own build system, or call vmaf_hip_available() / vmaf_metal_available() at run time.
Toolchain minimums
HIP needs ROCm 7.0 or newer (core/meson_options.txt); the shipped toolchain is ROCm 10.1.0 (build-config.env).
Ownership at a glance¶
All four backends keep the caller responsible for the state allocation. They differ in what import_state copies.
| Backend | import_state | Who frees the state | When | Free call |
|---|---|---|---|---|
| CUDA | Copies VmafCudaState by value into the context | Caller frees the original allocation | After vmaf_close() returned 0 | vmaf_cuda_state_free(state) (plain pointer, not cleared) |
| SYCL | Borrows the pointer | Caller | After vmaf_close() returned 0 | vmaf_sycl_state_free(&state) |
| HIP | Borrows the pointer | Caller | After vmaf_close() returned 0 | vmaf_hip_state_free(&state) |
| Metal | Borrows the pointer | Caller | After vmaf_close() returned 0 | vmaf_metal_state_free(&state) |
Common rules:
- Call
vmaf_close()first. Only an exact0tears down what the context holds; on a nonzero result retry it and keep every state and model alive (close and retry). - A successful CUDA close destroys the embedded copy (stream and, if libvmaf created it, the context). The original allocation still needs
vmaf_cuda_state_free()afterwards or it leaks. - Import the state before
vmaf_use_features_from_model()and before the firstvmaf_read_pictures(). - A state belongs to one context. CUDA returns
-EBUSYfor a second import of the same state or an import into a context that already has CUDA state. - Every function in these headers is not thread-safe; use one state and one context per driver thread.
- ABI: fork-added, stable. The SYCL zero-copy imports are experimental.
CUDA¶
Header: libvmaf_cuda.h.
Call order¶
vmaf_init()
vmaf_cuda_state_init()
vmaf_cuda_import_state() hands a by-value copy to the context
vmaf_use_features_from_model()
vmaf_cuda_preallocate_pictures()
loop:
vmaf_cuda_fetch_preallocated_picture() x2
fill .data[i]
vmaf_read_pictures()
vmaf_read_pictures(NULL, NULL, 0)
vmaf_score_pooled()
vmaf_close() retry on nonzero
vmaf_cuda_state_free() only after close returned 0
State and configuration¶
typedef struct VmafCudaConfiguration {
void *cu_ctx; /* CUcontext; NULL: libvmaf creates one */
} VmafCudaConfiguration;
int vmaf_cuda_state_init(VmafCudaState **cu_state, VmafCudaConfiguration cfg);
int vmaf_cuda_import_state(VmafContext *vmaf, VmafCudaState *cu_state);
int vmaf_cuda_state_free(VmafCudaState *cu_state);
| Field value | Effect |
|---|---|
cu_ctx = NULL | libvmaf creates a fresh context on the current CUDA device (device 0 by default). Typical for standalone tools. |
cu_ctx != NULL | A driver-API CUcontext the application owns, for example one shared with NVENC or NVDEC. It must outlive the state and every importing context until vmaf_close() returns 0. |
Errors: vmaf_cuda_import_state() returns -EBUSY as described above, or another negative errno. vmaf_cuda_state_free() accepts NULL.
Ownership and free¶
vmaf_cuda_state_free() (ADR-0157) frees the original allocation. When the state was never imported it first tears down the stream and context, which is retry-safe: a nonzero result retains the state and its live handles for another call.
VmafCudaState *cuda = NULL;
int err = vmaf_cuda_state_init(&cuda, (VmafCudaConfiguration){ .cu_ctx = NULL });
if (err)
return err;
err = vmaf_cuda_import_state(ctx, cuda); /* ctx holds a by-value copy */
if (err) {
vmaf_cuda_state_free(cuda); /* never imported: tears down too */
return err;
}
/* ... score ... */
int close_err = vmaf_close(ctx);
if (close_err != 0)
close_err = vmaf_close(ctx); /* retained teardown-only context */
if (close_err == 0)
vmaf_cuda_state_free(cuda); /* original allocation outlives close */
Picture preallocation¶
enum VmafCudaPicturePreallocationMethod {
VMAF_CUDA_PICTURE_PREALLOCATION_METHOD_NONE,
VMAF_CUDA_PICTURE_PREALLOCATION_METHOD_DEVICE,
VMAF_CUDA_PICTURE_PREALLOCATION_METHOD_HOST,
VMAF_CUDA_PICTURE_PREALLOCATION_METHOD_HOST_PINNED,
};
typedef struct VmafCudaPictureConfiguration {
struct { unsigned w, h; unsigned bpc; enum VmafPixelFormat pix_fmt; } pic_params;
enum VmafCudaPicturePreallocationMethod pic_prealloc_method;
} VmafCudaPictureConfiguration;
int vmaf_cuda_preallocate_pictures(VmafContext *vmaf, VmafCudaPictureConfiguration cfg);
int vmaf_cuda_fetch_preallocated_picture(VmafContext *vmaf, VmafPicture *pic);
| Method | .data[i] memory | Behaviour of vmaf_cuda_fetch_preallocated_picture | Use when |
|---|---|---|---|
NONE | none | Returns -EINVAL: no pool exists. Allocate pictures yourself. | Diagnostics only. |
DEVICE | cudaMalloc | Takes a picture from a device pool. | Source data already on the GPU (decoder output); no host-to-device copy in vmaf_read_pictures. |
HOST | malloc | Allocates a fresh pageable host picture each call. | Host source; libvmaf copies through a bounce buffer. |
HOST_PINNED | cudaMallocHost | Allocates a fresh pinned host picture each call. | Host source with asynchronous copy overlap; the usual choice for CPU-decoded feeds. |
Preallocating twice on one context returns -EBUSY. Pictures from a fetch follow the normal ownership rule: after vmaf_read_pictures() the context owns them. The CUDA backend page has measured throughput.
Example¶
#include <libvmaf/libvmaf.h>
#include <libvmaf/libvmaf_cuda.h>
#include <libvmaf/model.h>
static int run(VmafContext *vmaf, VmafModel *model, unsigned nframes, double *score)
{
VmafCudaPictureConfiguration pcfg = {
.pic_params = { .w = 1920, .h = 1080, .bpc = 8, .pix_fmt = VMAF_PIX_FMT_YUV420P },
.pic_prealloc_method = VMAF_CUDA_PICTURE_PREALLOCATION_METHOD_HOST_PINNED,
};
int err = vmaf_use_features_from_model(vmaf, model);
if (err) return err;
err = vmaf_cuda_preallocate_pictures(vmaf, pcfg);
if (err) return err;
for (unsigned i = 0; i < nframes; i++) {
VmafPicture ref, dist;
err = vmaf_cuda_fetch_preallocated_picture(vmaf, &ref);
if (err) return err;
err = vmaf_cuda_fetch_preallocated_picture(vmaf, &dist);
if (err) { vmaf_picture_unref(&ref); return err; }
/* fill ref.data[p] / dist.data[p] row by row with stride[p] */
err = vmaf_read_pictures(vmaf, &ref, &dist, i);
if (err) return err; /* the context owns both pictures */
}
err = vmaf_read_pictures(vmaf, NULL, NULL, 0);
if (err) return err;
return vmaf_score_pooled(vmaf, model, VMAF_POOL_METHOD_MEAN, score, 0, nframes - 1);
}
Surround run() with vmaf_init(), vmaf_cuda_state_init(), vmaf_cuda_import_state() and the close sequence above.
Limitations¶
- One device per state.
VmafCudaConfigurationhas no device index: select device N by making its context current (cuCtxSetCurrentorcudaSetDevice) beforevmaf_cuda_state_init(), or pass itsCUcontext. - No stream parameter. libvmaf runs its own streams; an external stream is not exposed.
SYCL¶
Header: libvmaf_sycl.h.
State and configuration¶
typedef struct VmafSyclConfiguration {
int device_index; /* -1: SYCL default device; 0+: Level Zero GPU ordinal */
int enable_profiling; /* non-zero: create the queue with profiling */
} VmafSyclConfiguration;
int vmaf_sycl_state_init(VmafSyclState **out, VmafSyclConfiguration cfg);
int vmaf_sycl_import_state(VmafContext *ctx, VmafSyclState *state);
void vmaf_sycl_state_free(VmafSyclState **state);
int vmaf_sycl_list_devices(void);
- The state owns the
sycl::queue, device allocations, shared frame buffers and profiling data. SYCL USM allocations are queue-scoped and the queue outlives one scoring session, which is why the caller frees the state aftervmaf_close()returns 0. vmaf_sycl_list_devices()enumeratesdevice_type::gpudevices only and prints one line each: ordinal, platform, vendor, driver version, fp64 support. It returns the count, or-EIOon a SYCL exception.vmaf_bench --list-devicesuses it (bench usage).
Picture preallocation¶
enum VmafSyclPicturePreallocationMethod {
VMAF_SYCL_PICTURE_PREALLOCATION_METHOD_NONE = 0,
VMAF_SYCL_PICTURE_PREALLOCATION_METHOD_DEVICE = 1,
VMAF_SYCL_PICTURE_PREALLOCATION_METHOD_HOST = 2,
};
typedef struct VmafSyclPictureConfiguration {
struct { unsigned w, h; unsigned bpc; enum VmafPixelFormat pix_fmt; } pic_params;
enum VmafSyclPicturePreallocationMethod pic_prealloc_method;
} VmafSyclPictureConfiguration;
int vmaf_sycl_preallocate_pictures(VmafContext *ctx, VmafSyclPictureConfiguration cfg);
int vmaf_sycl_picture_fetch(VmafContext *ctx, VmafPicture *pic);
The enumerator values are stable and append-only; serialised configuration and FFI bindings may rely on them. DEVICE and HOST create a 2-deep pool.
| Method | Backing | Use case |
|---|---|---|
NONE | No pool; vmaf_sycl_picture_fetch falls back to vmaf_picture_alloc() | CPU-fed callers and test harnesses |
DEVICE | sycl::malloc_device USM | A decoder or uploader writes directly into GPU-resident planes |
HOST | sycl::malloc_host USM | CPU-visible pooled planes |
The caller owns each picture reference returned by vmaf_sycl_picture_fetch(). Submit it through vmaf_read_pictures(), then release the caller's reference with vmaf_picture_unref(). The pool keeps its own references until vmaf_close() returns 0; after a nonzero close the context and its pool stay retained for retry.
Zero-copy frame-buffer path¶
For a GPU-resident decode pipeline (Intel VPL, VA-API, D3D11) the frame-buffer API exposes two shared Y-plane buffers (reference and distorted) and the SYCL backend's double-buffered upload path, instead of whole VmafPicture instances from the pool.
int vmaf_sycl_init_frame_buffers(VmafContext *ctx, unsigned w, unsigned h, unsigned bpc);
int vmaf_sycl_get_frame_buffers (VmafContext *ctx, void **ref, void **dis);
int vmaf_read_pictures_sycl (VmafContext *ctx, unsigned index); /* replaces vmaf_read_pictures */
int vmaf_sycl_wait_compute (VmafContext *ctx);
int vmaf_flush_sycl (VmafContext *ctx); /* replaces the NULL flush */
vmaf_sycl_init_frame_buffers(vmaf, W, H, 8);
void *ref_buf = NULL, *dis_buf = NULL;
vmaf_sycl_get_frame_buffers(vmaf, &ref_buf, &dis_buf);
for (unsigned i = 0; i < nframes; i++) {
/* write Y-plane luma into ref_buf / dis_buf: kernels, SYCL events,
* dmabuf imports, whatever the decoder exposes */
vmaf_read_pictures_sycl(vmaf, i);
/* vmaf_sycl_wait_compute() is needed only to reuse the buffers for a
* later frame while compute is still in flight */
}
vmaf_flush_sycl(vmaf);
The path holds the luma plane only, and the extractors get no host picture (ADR-1688). Each call to vmaf_read_pictures_sycl() first checks every registered feature extractor. It returns -ENOTSUP before counting the frame, with one libvmaf error line naming each extractor that cannot run on luma alone:
- a SYCL twin that reads chroma or a host picture. That includes
speed_chroma_sycl, which the default modelvmaf_v1.0.16_3d0hneeds forspeed_chroma_uv;motion_syclwithmotion_add_uv=true; andpsnr_syclorpsnr_hvs_syclwithenable_chroma(their default); - a CPU extractor: a feature with no SYCL twin, or a twin that cannot honour the model's options.
What runs: adm_sycl, vif_sycl, motion_sycl, motion_v2_sycl, cambi_sycl, float_moment_sycl, and psnr_sycl / psnr_hvs_sycl with enable_chroma=false. That covers vmaf_v0.6.1 (adm2, vif, motion2), whose scores equal the CPU's on this path. Score anything else from host frames with vmaf_read_pictures(). Before ADR-1688 such a feature failed with a bare -22, crashed the process (float_psnr_sycl), added the SAD of chroma that was never imported (motion_sycl with motion_add_uv), or was dropped from the result without an error (a CPU extractor).
GPU-resident import paths¶
/* Linux, Level Zero */
int vmaf_sycl_dmabuf_import(VmafSyclState *state, int fd, size_t size, void **ptr);
void vmaf_sycl_dmabuf_free (VmafSyclState *state, void *ptr);
int vmaf_sycl_import_va_surface(VmafSyclState *state, void *va_display,
unsigned int va_surface, int is_ref,
unsigned w, unsigned h, unsigned bpc);
int vmaf_sycl_upload_plane(VmafSyclState *state, const void *src, unsigned pitch,
int is_ref, unsigned w, unsigned h, unsigned bpc);
/* Windows only (_WIN32) */
int vmaf_sycl_import_d3d11_surface(VmafSyclState *state, void *d3d11_device,
void *d3d11_texture, unsigned subresource,
int is_ref, unsigned w, unsigned h, unsigned bpc);
| Function | Role | Notes |
|---|---|---|
vmaf_sycl_dmabuf_import | Primitive: turns a DMA-BUF fd into a SYCL device pointer through Level Zero external memory. | Caller keeps the fd. Free the pointer with vmaf_sycl_dmabuf_free (NULL is a no-op). |
vmaf_sycl_import_va_surface | Convenience wrapper over dmabuf for a VA-API decode feed. Preferred on Linux. | Imports the luma plane only (see the zero-copy path above for what that can score). Falls back to vaGetImage plus a host-to-device copy when the DRM-PRIME export fails (older Mesa, proprietary drivers). bpc is 8 or 10. |
vmaf_sycl_upload_plane | Platform-agnostic escape hatch: copies a Y plane from host memory with a row pitch. | Returns once the copy has completed, so the caller may free or reuse the source and the next vmaf_read_pictures_sycl() reads the uploaded plane. Before 2026-10-05 it returned with the copy still in flight, and a 3840x2160 frame scored the wrong pixels. Use when nothing better works, or as a benchmark baseline. |
vmaf_sycl_import_d3d11_surface | Windows: copies the decoded texture into a staging texture, maps it and uploads the plane through vmaf_sycl_upload_plane. | Implemented in core/src/sycl/d3d11_import.cpp. -EINVAL for NULL arguments, zero size, bpc other than 8 or 10, or a texture smaller than w by h; -EIO when the map fails. |
Every import takes is_ref (non-zero: reference buffer, otherwise distorted) and needs vmaf_sycl_init_frame_buffers() first. See ADR-0016 for how these APIs landed and the SYCL overview for the ingestion-path decision tree.
Profiling helpers¶
int vmaf_sycl_profiling_enable (VmafSyclState *state);
void vmaf_sycl_profiling_disable (VmafSyclState *state);
void vmaf_sycl_profiling_print (VmafSyclState *state);
int vmaf_sycl_profiling_get_string(VmafSyclState *state, char **out);
Request profiling at init: set VmafSyclConfiguration.enable_profiling = 1 (or VMAF_SYCL_PROFILE=1 in the environment). Then use the enable and disable pair to choose which frame ranges are timed. The SYCL environment switches (VMAF_SYCL_PROFILE, VMAF_SYCL_TIMING, VMAF_SYCL_IMPORT_DEBUG, VMAF_SYCL_CHECKSUM) are read once, at first use, like VMAF_SYCL_DISPATCH; set them before the first SYCL state is created.
Warning
vmaf_sycl_profiling_enable() does not re-create the queue. It only sets a flag on the state (core/src/sycl/common.cpp). If the queue was not created with profiling, the call succeeds but reading kernel event timings later throws a sycl::exception.
vmaf_sycl_profiling_get_string() returns a caller-owned buffer; release it with free(). It matches vmaf_bench --gpu-profile (bench usage).
Limitations¶
dmabuf_importandimport_va_surfaceneed the SYCL queue on the Level Zero backend (they callsycl::get_native<ext_oneapi_level_zero>). On an OpenCL-backend SYCL build they throwsycl::exception, which the wrapper converts to-EIO. The log line is generic (SYCL DMA-BUF import exception: <what()>), so detect the backend up front withsycl::queue::get_backend()and fall back tovmaf_sycl_upload_planeinstead of parsing the log.vmaf_sycl_init_frame_buffers()is single-resolution. Changingw,horbpcmid-stream needsvmaf_close()and a new init.
HIP¶
Header: libvmaf_hip.h. HIP runtime types (hipDevice_t, hipStream_t) cross the ABI as uintptr_t, so the header does not include <hip/hip_runtime.h>; cast on your side.
State and configuration¶
typedef struct VmafHipConfiguration {
int device_index; /* -1: first compute-capable HIP device */
int flags; /* reserved; pass 0 */
} VmafHipConfiguration;
int vmaf_hip_available(void);
int vmaf_hip_state_init(VmafHipState **out, VmafHipConfiguration cfg);
int vmaf_hip_import_state(VmafContext *ctx, VmafHipState *state);
void vmaf_hip_state_free(VmafHipState **state);
int vmaf_hip_list_devices(void);
A zero-initialised configuration selects device 0.
| Function | Does | Errors |
|---|---|---|
vmaf_hip_available | Returns 1 when libvmaf was built with -Denable_hip=true, else 0. Touches no HIP runtime. | none |
vmaf_hip_state_init | Allocates a state pinned to one HIP device and its compute stream. | -ENODEV no visible HIP device, -EINVAL bad arguments or an ordinal the runtime does not have, another negative errno when the HIP runtime fails, -ENOSYS built without HIP |
vmaf_hip_import_state | Hands the state to a context; the context borrows the pointer. | -EINVAL NULL ctx or state, -ENOSYS built without HIP |
vmaf_hip_state_free | Releases the state and sets *state to NULL. Accepts NULL or a never-imported state. | none |
vmaf_hip_list_devices | Logs ordinal, name and GFX architecture per device at VMAF_LOG_LEVEL_INFO. Returns the count, 0 when the runtime sees no device. | a negative errno when the HIP runtime fails or a device cannot be described, -ENOSYS built without HIP |
Every HIP twin of a context runs on the state's device: it creates its device resources there, and libvmaf makes that device current before each frame's HIP work and before the flush. The thread that calls vmaf_read_pictures() does not change the device.
Ownership and call order¶
After vmaf_hip_import_state() the caller still owns the state. Call vmaf_hip_state_free(&state) only after vmaf_close() returned 0; a nonzero close keeps the context and its borrowed state alive for retry. The state may outlive one scoring session when you run several passes on the same device.
vmaf_init()
vmaf_hip_state_init(&state, cfg)
vmaf_hip_import_state(vmaf, state)
vmaf_use_features_from_model()
loop: vmaf_read_pictures(vmaf, &ref, &dist, i)
vmaf_read_pictures(NULL, NULL, 0)
vmaf_score_pooled(vmaf, ...)
vmaf_close(vmaf) retry on nonzero
vmaf_hip_state_free(&state) only after close returned 0
libvmaf_hip.h has no picture-preallocation functions: feed ordinary host pictures from vmaf_picture_alloc() or the CPU pool.
Feature coverage¶
The registry holds 19 HIP extractors, verified end to end on AMD gfx hardware: PSNR, float-PSNR, CIEDE, float-moment (which also serves integer-moment), float-SSIM, integer-SSIM, MS-SSIM, PSNR-HVS, CAMBI, SSIMULACRA2, integer-motion, integer-motion-v2, float-motion, float-VIF, integer-VIF, integer-ADM, float-ADM, speed-chroma and speed-temporal. The older adm_hip, vif_hip and motion_hip stubs no longer exist. The float_ansnr extractor was removed (PR #38) and is registered on no backend.
The HIP overview has the full extractor table, HSACO fat-binary target selection, build flags, FFmpeg integration (hip_device=N, ADR-0380) and per-kernel notes.
Metal¶
Header: libvmaf_metal.h. Metal types (id<MTLDevice>, id<MTLCommandQueue>) cross the ABI as uintptr_t. The backend targets Apple Silicon (Apple GPU family 7, M1 and later); Intel Macs and other hosts return -ENODEV instead of falling back to the CPU (ADR-0361). The header is installed whenever Metal is enabled so FFmpeg's --enable-libvmaf-metal configure probe finds it (ADR-0437). The registry holds 17 Metal extractors.
State and configuration¶
| Function | Does | Errors |
|---|---|---|
vmaf_metal_available | 1 when built with Metal, else 0. | none |
vmaf_metal_state_init(&state, cfg) | Allocates a state for device device_index (-1: system default; zero-initialised config picks 0). flags is reserved. | -ENODEV not Apple GPU family 7+, -ENOSYS built without Metal, -EINVAL |
vmaf_metal_state_init_external(&state, handles) | Adopts caller-supplied MTLDevice and MTLCommandQueue (VmafMetalExternalHandles, uintptr_t each; 0 selects the system default device or a new queue). Use when the IOSurface source and libvmaf must share one device. Exclusive with vmaf_metal_state_init in one process context. | -EINVAL, -ENODEV, -ENOMEM |
vmaf_metal_import_state(ctx, state) | Hands the state to a context; the context borrows the pointer. | -EINVAL, -ENOSYS |
vmaf_metal_state_free(&state) | Releases a state from either init function; accepts NULL. | none |
vmaf_metal_list_devices | Enumerates Apple GPU family 7+ devices. Returns the count. | -ENOSYS built without Metal |
External handles are never owned by libvmaf: keep them alive until vmaf_metal_state_free() returns.
IOSurface import¶
For FFmpeg and VideoToolbox callers that hold CVPixelBufferRef frames, take the surface with CVPixelBufferGetIOSurface and pass it to libvmaf (ADR-0423).
vmaf_init()
vmaf_metal_state_init(&state, cfg)
vmaf_metal_import_state(vmaf, state)
loop:
vmaf_metal_picture_import(state, iosurface, plane, w, h, bpc, is_ref, index)
vmaf_metal_wait_compute(state)
vmaf_metal_read_imported_pictures(vmaf, index)
vmaf_score_pooled(vmaf, ...)
vmaf_close(vmaf) retry on nonzero
vmaf_metal_state_free(&state) only after close returned 0
| Function | Does |
|---|---|
vmaf_metal_picture_import | Imports one plane (0 = Y, 1 = U, 2 = V) of an IOSurfaceRef into a planar 4:2:0 VmafPicture. The caller keeps the surface; libvmaf locks it read-only and copies the plane. The surface's pixel format decides the read (ADR-1679). NV12 (420v / 420f, bpc 8) and P010 (x420 / xf20, bpc 10) are what VideoToolbox decodes to: planes 1 and 2 are the Cb and Cr samples of their interleaved second plane, and P010 samples are shifted from the top 10 bits of 16 to the bottom 10. Planar 8-bit 4:2:0 (y420 / f420) is read plane by plane. Errors: -ENOTSUP for any other pixel format, -EINVAL (including a bpc that is not the format's, or a surface plane smaller than the frame), -EIO (lock failure), -ENOMEM. |
vmaf_metal_wait_compute | Blocks until Metal work on the state finished. Currently a synchronous no-op, because the import is a host-side copy; a later asynchronous path will drain an MTLSharedEvent. -EINVAL for a NULL state. |
vmaf_metal_read_imported_pictures | Triggers a score read for the imported reference and distorted surfaces at index. All three planes (Y, U, V) of both must have been imported at that index. |
Note
The v1 import path copies on the CPU. True zero-copy GPU binding is a tracked gap (GAP-METAL-IOSURFACE-NOT-TRUE-ZERO-COPY).
Related¶
- Core API overview, DNN sessions.
- CLI backend selection:
--no_cuda,--no_sycl,--sycl_deviceand the other flags. - Backend pages: CUDA, SYCL, HIP, Metal.
vmaf_benchconsumes these APIs for the performance and validation tables.- Governing decisions: ADR-0016, ADR-0022, ADR-0027.
Licensing of the GPU headers (ADR-1250)¶
libvmaf_cuda.h is Netflix's and carries SPDX-License-Identifier: BSD-2-Clause-Patent. libvmaf_sycl.h, libvmaf_hip.h and libvmaf_metal.h are fork-authored and carry SPDX-License-Identifier: EUPL-1.2.
The European Commission reads linking against an EUPL work as creating no condition on the linking program. Distributing libvmaf itself, modified or not, still obliges you to provide its source or point to a repository that has it (EUPL-1.2, Article 5). A modified libvmaf is distributed under the EUPL-1.2. See Embedding VMAFx in another product, ADR-1250 and the README.