SYCL Backend¶
The SYCL backend runs libvmaf's feature extractors on Intel GPUs (Arc, Xe integrated graphics) through oneAPI DPC++ and Level Zero. Build it with -Denable_sycl=true, select it with --backend sycl, and expect scores that equal the CPU's bit for bit for 24 of the 25 gated features.
This page covers what you need to build and run it, how closely it agrees with the CPU, and what is still open. The details live on five further pages:
| Page | What it holds |
|---|---|
| AOT targets | Ahead-of-time device code, the default target list, sub-group sizes |
| Device sanitizer | -Dsycl_device_asan=true, what the DPC++ device AddressSanitizer reports on an Arc GPU and what it cannot prove |
| Twin notes | How each twin reaches the CPU's bits; compile-line guarantees |
| Zero-copy and pictures | QSV / VA-API import, D3D11, picture pre-allocation |
| Developer notes | fp64-free kernels, scratch-memory rules, source layout |
| History | Dated change log with every before / after measurement |
Platform guides: bundling a self-contained binary and SYCL on Windows.
Hardware¶
The backend targets any SYCL device that Level Zero exposes.
- Intel Arc discrete GPUs: A-series (DG2 / ACM) and B-series (Battlemage).
- Intel integrated GPUs: Tiger Lake and later (Xe-LP, Xe-LPG, Xe2).
- Other devices: AMD through the HIP plugin and NVIDIA through the CUDA plugin work only when the DPC++ compiler is built with them. The shipped oneAPI binaries do not include the HIP plugin.
A device that is not in the ahead-of-time list, such as a Data Center GPU Flex or Max card, still works: the binary also carries portable SPIR-V, which the Level Zero runtime compiles on first use.
Requirements¶
| Component | Version | Source of the pin |
|---|---|---|
oneAPI DPC++ (icpx) | 2026.1 | ONEAPI_VERSION in build-config.env |
| Level Zero loader | 1.34.0 | LEVEL_ZERO_VERSION in build-config.env |
| Intel compute runtime | 26.35.39758.10 | INTEL_NEO_VERSION in build-config.env |
Intel ocloc | any release that knows every target | scripts/ci/install-intel-ocloc.sh |
| Linux render node | /dev/dri/renderD* | none |
build-config.env is the single source for these versions; the release image and CI install exactly what it names. Older oneAPI releases may work, but only the pinned set is tested.
ocloc is needed only for ahead-of-time builds (the default); see AOT targets. AdaptiveCpp is an alternative compiler, chosen with -Dsycl_compiler=acpp (see the table below).
Build¶
-
Configure from the repository root. The Meson source directory is
core/: -
Build:
The options that shape a SYCL build:
| Option | Default | Effect |
|---|---|---|
enable_sycl | false | Compile the SYCL backend and its kernels |
sycl_compiler | icpx | Compiler: Intel icpx, or AdaptiveCpp acpp / syclcc (ADR-0335) |
sycl_icpx_aot_targets | 19 Intel targets | Ahead-of-time device list; empty string selects SPIR-V JIT only. Ignored unless the compiler is icpx |
sycl_acpp_targets | generic | AdaptiveCpp --acpp-targets value. Ignored when the compiler is icpx |
enable_cuda | false | May be set together with enable_sycl; both backends then live in one binary |
Note
A SYCL build is an icx / icpx build for its host code as well. Its CPU extractors take log10, pow, powf and the other math functions from glibc's libm, not from Intel's libimf, because the build passes -no-intel-lib=libimf to every link (ADR-1495, build flags). The CPU half of a SYCL build therefore returns a GCC build's scores, and a SYCL twin that equals its own build's CPU extractor equals a GCC build's too. The option applies to the host link only; the device code and its math libraries do not change.
Run¶
./build/tools/vmaf ... # SYCL used automatically if it initialises
./build/tools/vmaf --backend sycl ... # exit 100 if SYCL cannot initialise
./build/tools/vmaf --no_sycl ... # force the CPU path
./build/tools/vmaf --sycl_device 1 ... # pick device index 1
| Flag | Meaning |
|---|---|
--backend sycl | Use SYCL exclusively; an init failure exits with code 100 |
--no_sycl | Disable the SYCL backend |
--sycl_device N | Pick the device by index; without it SYCL's default selector decides, and --backend sycl uses index 0 |
Index 0 is whichever device SYCL's default selector picks, usually the first discrete GPU. Pin an integrated GPU or a specific Arc card with --sycl_device, or with the oneAPI runtime variable ONEAPI_DEVICE_SELECTOR=level_zero:N.
Automatic selection falls back to the CPU when SYCL cannot initialise, for example when a Level Zero or Unified Runtime library fails to load or a container has no render node. The run still prints a score and exits 0; only a stderr line such as problem during vmaf_sycl_state_init, using CPU shows it. --backend sycl makes the failure explicit instead. See how selection works.
Choosing a twin from --feature¶
With --backend sycl, a --feature that names a CPU extractor runs on that extractor's SYCL twin (ADR-1359):
vmaf ... --backend sycl --feature cambi # runs cambi_sycl
vmaf ... --backend sycl --feature psnr=enable_mse=true
vmaf ... --backend sycl --feature float_motion_sycl=motion_max_val=4
The twin runs only when it can honour the options and the frame size and bit depth you gave. Otherwise the CPU extractor runs and vmaf prints one warning that names the reason; the feature_backends receipt in the JSON output lists which extractor ran where. The rules and the warning texts are in CLI: feature extractors on a GPU backend. Naming the twin (adm_sycl) always selects it, and it fails when it cannot run.
A model's feature list resolves the same way, so the default model vmaf_v1.0.16_3d0h runs its VIF, ADM, motion and CAMBI features on the device.
Environment variables¶
| Variable | Effect |
|---|---|
VMAF_SYCL_DISPATCH | Per-feature dispatch strategy, direct or graph (below) |
VMAF_SYCL_USE_GRAPH | 1 forces graph replay for every feature |
VMAF_SYCL_NO_GRAPH | Deprecated: 1 forces direct submission and prints a warning |
VMAF_SYCL_IMPORT_DEBUG | 1 logs shared frame-buffer addresses and each VA import at INFO |
VMAF_SYCL_PROFILE | 1 enables queue profiling events |
VMAF_SYCL_TIMING | 1 records per-extractor timing with a queue wait |
VMAF_SYCL_CHECKSUM | 1 logs a CRC of the uploaded ref and dis frame buffers at INFO |
VMAF_SYCL_SCRATCH_SELFTEST | 0 skips the start-up scratch-memory probes (below) |
VMAF_SYCL_AOT_JOBS | Parallel compiles of the sycl-aot test suite (default 4) |
ONEAPI_DEVICE_SELECTOR | oneAPI runtime device filter, for example level_zero:0 |
VMAF_SYCL_DISPATCH is also listed in the env-var reference.
Dispatch strategy¶
VMAF_SYCL_DISPATCH chooses how kernels reach the device:
| Value | Behaviour |
|---|---|
direct | Submit kernels to an in-order queue. Lower per-frame overhead at small resolutions |
graph | Replay a recorded SYCL graph (ADR-0483). Cuts kernel-launch overhead at 720p and above |
When nothing is set, an area threshold decides: graph from 1280 x 720 pixels upward, direct below. The zero-copy VA-import path of the libvmaf_sycl FFmpeg filter is the exception and defaults to direct at every resolution: there the graph's output is byte-identical, but the per-frame de-tile import plus the graph's compute barrier serialise decode and compute, which costs 15 to 25 % at 4K (ADR-1121). Set VMAF_SYCL_USE_GRAPH=1 or VMAF_SYCL_DISPATCH=<feature>:graph to force the graph there anyway.
VMAF_SYCL_NO_GRAPH is deprecated (ADR-0841): it still works, prints a one-shot warning, and the warning names v4.0 as the removal release. VMAF_SYCL_IMPORT_DEBUG is read once at init, so changing it mid-run has no effect.
Numerical agreement with the CPU¶
The SYCL twins return the CPU extractors' bits. The cross-backend gate compares 25 features; 24 have an exact-twin declaration in scripts/ci/exact_twins.d/*.sycl and are held to tolerance 0. Only ciede is bounded, at 1e-9 (measured 1.4e-11). The generated list is the exact-twin table and the procedure is in the gate guide.
Measured on an Intel Arc A380 (2026-10-02) with the default model, the VMAF score of every frame equals --backend cpu on the Netflix src01 pair (48 frames) and on 50 frames of BBB 3840x2160.
| CPU extractor (gate feature) | SYCL twin | Agreement | ADR |
|---|---|---|---|
adm | adm_sycl | exact | 1362, 1451 |
float_adm | float_adm_sycl | exact | 1434 |
vif | vif_sycl | exact | 1432 |
float_vif | float_vif_sycl | exact | 1422 |
motion, motion_debug, motion_mffw | motion_sycl | exact | 1371, 1451, 1491 |
motion_v2, motion_v2_mffw | motion_v2_sycl | exact | 1451, 1491 |
float_motion | float_motion_sycl | exact | 1411 |
psnr | psnr_sycl | exact | 1451 |
float_psnr | float_psnr_sycl | exact | 1450 |
psnr_hvs | psnr_hvs_sycl | exact | 1401 |
ssim | integer_ssim_sycl | exact | 1443 |
float_ssim, float_ssim_lcs | float_ssim_sycl | exact | 1463 |
float_ms_ssim, float_ms_ssim_lcs, float_ms_ssim_chroma | float_ms_ssim_sycl | exact | 1414, 1466 |
ssimulacra2 | ssimulacra2_sycl | exact | 1446 |
float_moment | float_moment_sycl | exact | 1449, 1497 |
cambi | cambi_sycl | exact while the CPU's own top-K double sum is exact; otherwise within that sum's rounding (2.2e-15 at most over 50 frames of BBB 4K) | 1357 |
speed_chroma, speed_temporal | speed_chroma_sycl, speed_temporal_sycl | exact | 1358, 1477 |
ciede | ciede_sycl | bounded at 1e-9; measured within 1.4e-11 (the host's powf differs) | 1436 |
float_ansnr has no twin; it was removed (ADR-0865). The extractor-by-extractor coverage across all backends is in metrics/features.md.
Four rules about device code make the table possible. Each is enforced by a test, and Twin notes explains how the twins meet them:
| Rule | Meaning | Where |
|---|---|---|
| No fp64 in any kernel | Arc A-series has no fp64; a single fp64 instruction rejects the whole module | ADR-0220, developer notes |
| Strict FP line | -fp-model=precise -ffp-contract=off plus correctly rounded / and sqrt | ADR-1367, twin notes |
| No scratch memory | Kernels that spill return wrong values on Arc A-series under the xe driver | ADR-1395, below |
| Sub-group size 16 or 32 | Xe2 targets reject a kernel that requires 8 | ADR-1468, AOT targets |
The Netflix golden-data gate is CPU-only (principles). The SYCL backend's numerics are pinned by fork-added tests such as test_sycl_exact_twins, not by the Netflix goldens.
AOT targets (default, ADR-0568)¶
The default build compiles device code ahead of time for 19 Intel targets, so the first launch does not pay a JIT compile. A device outside the list falls back to SPIR-V JIT.
| Family | Targets |
|---|---|
| Arc A-series (DG2, ACM) | dg2-g10, dg2-g11, acm-g10, acm-g11, acm-g12 |
| Xe-LP integrated (Tiger, Alder, Raptor Lake) | tgllp, adl-s, adl-p, adl-n, rpl-s, rpl-p |
| Xe-LPG integrated (Meteor, Arrow Lake) | mtl-h, mtl-u, arl-h, arl-s, arl-u |
| Xe2 (Lunar Lake, Battlemage) | lnl-m, bmg-g21, bmg-g31 |
Override the list with -Dsycl_icpx_aot_targets=dg2-g11,mtl-h, or disable AOT with -Dsycl_icpx_aot_targets=''. AOT targets has the silicon behind each name, the build and image checks, the target-list recipes and the sub-group-size table.
Scratch memory on Intel GPUs (ADR-1395)¶
A kernel uses scratch memory when the Intel graphics compiler puts a private array in memory or spills registers there. On an Arc A-series GPU under the Linux xe kernel driver, such kernels return wrong values with no error. This was measured on an Arc A380 with compute runtime 26.35.39758.10 and IGC 2.41.5, where the same kernels were correct under i915.
The fork's rule is that SYCL kernels use no scratch memory (ADR-1395). Since 2026-10-01 none does: the list of kernels that still did, core/src/sycl/scratch_ratchet.txt, is empty. float_adm_sycl was the last extractor on it; until then it returned NaN on an Arc A380 under xe and the run stopped with problem reading pictures.
What you see¶
At the first SYCL initialisation on each device, libvmaf runs two small probe kernels, one with a private array and one that spills. When either returns wrong values it logs a warning like this:
libvmaf WARNING SYCL: Intel(R) Arc(TM) A380 Graphics returns wrong values from kernels that use
scratch memory (private-array probe: 256 of 256 work-items wrong; register-spill probe: 256 of 256).
Seen on Arc A-series GPUs under the Linux xe kernel driver, not under i915. No libvmaf SYCL
extractor uses scratch memory, so its scores are not affected. See
docs/backends/sycl/overview.md (ADR-1395).
The device is still used and every libvmaf extractor gives correct scores on it. The warning stays because it describes the device: a kernel from outside libvmaf, or one added without the audit in the developer notes, would be affected.
The probes cost 3 to 5 ms with a warm compute-runtime kernel cache and about 0.4 s on the first run. VMAF_SYCL_SCRATCH_SELFTEST=0 skips them.
Note
A build older than 2026-10-01 names, in this message, the extractors that still used scratch memory. On such a build compute those features with their CPU extractors (--feature adm rather than adm_sycl, or --no_sycl), or run the GPU under i915.
Checking a device¶
test_sycl_kernel_scratch (in --suite sycl) builds every kernel libvmaf registers for the default GPU and fails on each one that uses scratch memory:
ONEAPI_DEVICE_SELECTOR=level_zero:0 \
python3 scripts/ci/run_meson_test.py -- -C build -v test_sycl_kernel_scratch
The ratchet list that used to exempt known kernels is empty and stays empty; test_sycl_kernel_source_contract.py rejects a new entry without a device. The test skips without a GPU, so CI, which has none, does not run it.
It audited 128 kernels on an Arc A380 at the revision of ADR-1488 (2026-10-02) and found none. On an Arc B580 and an Arc Pro B60 (Xe2, xe driver) it audited 127 kernels and found none (2026-10-03). Both Xe2 cards return correct values from the two probes, so they log no warning. Until 2026-10-03 the term kernel of float_adm_sycl used 128 bytes of scratch memory there (ADR-1501).
VIF kernels run at SIMD-16¶
vif_sycl runs its kernels at sub-group size 16 on every device (ADR-1830); a SYCL device without SIMD-16 sub-groups (no Intel GPU) gets -ENOTSUP from it. The SIMD-32 kernels and the VMAF_SYCL_VIF_SUBGROUP_SIZE variable that forced them were removed on 2026-10-05: they needed a 256-entry register file that Xe-LP integrated GPUs (UHD 770 and similar) do not have, spilled to scratch memory there, and were never faster than SIMD-16 on an Arc A380 (21.19 against 21.20 ms per frame at 3840x2160, 33.85 against 23.50 ms with vif_fused=true). Setting the variable now has no effect.
Profiling¶
- Intel VTune (
vtune-gui) with the GPU Compute analysis type for kernel occupancy and EU utilization. onetracefrom the pti-gpu project for Level Zero API-level tracing.VMAF_SYCL_PROFILE=1gives the queuesenable_profiling; withVMAF_SYCL_NO_GRAPH=1graph extractors submit directly, so every kernel has its own event (Research-1369 describes the event-timing build).- For end-to-end wall-time comparisons against the CUDA and CPU paths, use
make test-netflix-golden, which records per-backend scores and timings. - Programmatic profiling through
VmafSyclState.enable_profiling; see api/gpu.md for the queue-event query API.
Known gaps¶
Only open items are listed. A closed gap is recorded in History and in docs/state.md, which is the authority on status.
| Gap | Effect | Tracked as |
|---|---|---|
| No CI lane executes a SYCL kernel | Device parity is checked by hand on Arc hardware; the isolated Arc workflow has no registered runner | T-SYCL-NO-CI-KERNEL-EXECUTION-2026-09-04 |
| The six row kernels moved to sub-group 16 are verified on Xe2 and Arc A-series only | Scratch freedom and exactness at 16 are unmeasured on Xe-LP and Xe-LPG integrated GPUs | T-SYCL-ROW-KERNELS-SG16-OTHER-DEVICES-2026-10-02 |
| Exact twins cost throughput | Scores are correct; float_ssim, float_ms_ssim, ssim, ssimulacra2, ciede, float_vif, float_moment and psnr_hvs take longer than before exactness, and float_ms_ssim at 4K is slower than a 16-thread CPU run (--backend cpu is faster there) | T-SYCL-FLOAT-SSIM-RASTER-SUM-THROUGHPUT-2026-10-02, T-SYCL-FLOAT-MS-SSIM-RASTER-SUM-THROUGHPUT-2026-10-02, T-SYCL-SSIM-EXACT-THROUGHPUT-2026-10-02, T-SYCL-SSIMULACRA2-EXACT-THROUGHPUT-2026-10-02, T-SYCL-CIEDE-EXACT-THROUGHPUT-2026-10-01, T-SYCL-FLOAT-VIF-EXACT-THROUGHPUT-2026-10-01, T-SYCL-FLOAT-MOMENT-PER-PIXEL-ATOMICS-2026-10-02, T-SYCL-HIP-PSNR-HVS-EXACT-SUM-THROUGHPUT-2026-10-01 (RC8) |
psnr_hvs_sycl is slow on Xe-LP | About 49 ms of kernel time per 3840x2160 frame on a UHD 770, against about 7 ms for 16 CPU threads | T-SYCL-PSNR-HVS-XE-LP-THROUGHPUT-2026-09-29 (RC8) |
float_motion_sycl lacks three CPU options | motion_add_scale1, motion_add_uv and motion_filter_size are not declared, so a request that sets one runs float_motion on the CPU (ADR-1183). motion_sycl has its own motion_add_uv | none (see the motion page) |
| AdaptiveCpp builds are JIT only | sycl_acpp_targets defaults to generic, AdaptiveCpp's portable JIT path; ahead-of-time intel_gpu_<arch> strings (AdaptiveCpp 23.10 and later) are not wired | follow-up task, no state row |
| HIP and CUDA devices through SYCL need a custom DPC++ | The shipped oneAPI compiler includes the Level Zero, OpenCL CPU and CUDA plugins, not HIP | build limitation |
| dmabuf import is Linux-only | vmaf_sycl_dmabuf_import and vmaf_sycl_import_va_surface return -ENOSYS on Windows, where callers use the D3D11 staging path; DMA-BUF is a Linux kernel interface | by design (T-SYCL-DMABUF-IMPORT-WIN32-ENOSYS) |
| Windows MSVC device link is not measured on a GPU | The one explicit device link of the MSVC build gets the strict FP flags, but that path has not run on a GPU (ADR-1364) | none |
See metrics/features.md for the per-extractor coverage matrix and api/gpu.md for the programmatic surface.
References¶
- SYCL 2020 Specification
- Intel oneAPI DPC++ Compiler
- Level Zero Specification
- Intel oneAPI Programming Guide
Licensing of the SYCL kernels (ADR-1250)¶
As with the other backends, a SYCL kernel implementing an upstream Netflix metric keeps that code's terms and copyright notice, while fork-original SYCL code is EUPL-1.2. Four files in core/src/sycl/ additionally carry an outside contributor's work and stay on their current terms until that contributor agrees to a change. See ADR-1250.
Former section names¶
ADRs and research digests link to these headings; each points to the section that now holds its content.
Numerical tolerance vs the CPU scalar path¶
Now under numerical agreement with the CPU.