Intel oneAPI install: local SYCL toolchain¶
Install the Intel oneAPI DPC++ compiler icpx, the ocloc offline compiler and a Level Zero runtime, then configure a SYCL build. The SYCL backend (-Denable_sycl=true) needs all three. This page covers the local developer machine on Linux; for a native Windows build see SYCL on Windows.
CI installs oneAPI from Intel's apt repository on Linux and from Intel's offline installer on Windows; see CI vs local.
Quick path¶
-
Install
ocloc, the offline GPU compiler that the default ahead-of-time build calls (details): -
Install the oneAPI compiler (see Install paths) and activate it for the shell:
-
Configure and build from the repository root:
For a SPIR-V build without ocloc (for example to run clang-tidy), configure with -Dsycl_icpx_aot_targets=: the binary then compiles its kernels at first launch instead.
Pinned version¶
The pin lives in build-config.env; this page does not repeat the numbers except where a command needs one.
| Component | Pin | Notes |
|---|---|---|
| Intel oneAPI compiler (Linux) | ONEAPI_VERSION (2026.1), exact apt build ONEAPI_APT_VERSION | icpx ships with the compiler package; CI and the release image install it from Intel's apt repository. |
| Intel oneAPI (Windows) | ONEAPI_WINDOWS_VERSION (2025.3.0.372); INTEL_BASEKIT_WINDOWS_URL, _SIZE, _SHA256 (the installer the Windows SYCL tester zip checks and installs) | The offline installer URL carries a per-build GUID, so it lags the Linux pin. |
Compute runtime (level-zero-loader) | distro package | Arch / CachyOS: pacman -S level-zero-loader. |
ocloc (GPU offline compiler) | INTEL_NEO_VERSION | Needed for the default ahead-of-time build. |
| Level Zero loader (container) | LEVEL_ZERO_VERSION | Installed by install-intel-ocloc.sh --components build. |
Building with icx / icpx¶
Any SYCL build uses icpx. You can also build the CPU code with Intel's compilers (icx for C, icpx for C++); the vmaf of the dev container (dev/Containerfile), the Linux Intel LLVM job of build.yml and the Ubuntu SYCL and Ubuntu SYCL+CUDA legs of libvmaf-build-matrix.yml are built that way.
-
Configure the host compilers through the environment, then build:
A SYCL build (
-Denable_sycl=true) takesicpxfromsycl_compiler(defaulticpx) for the device code in addition. -
Check the math library. The build passes
-no-intel-lib=libimfto every C and C++ link, so the CPU extractors call glibc'slibmand an icx build returns the scores of a GCC build (ADR-1495). Nothing is needed on the command line; the block is incore/src/meson.buildbetween theBEGIN / END VMAF host math library link policymarkers. Verify it with:
What an icx build does, and does not, change:
| Topic | Behaviour | Reference |
|---|---|---|
| Floating-point flags | Every C and C++ translation unit gets -fp-model=precise -ffp-contract=off, so nothing fuses a * b + c on its own. | ADR-1461, build flags |
| Math library | Linux links glibc libm, not Intel's libimf. | ADR-1495 |
| SYCL device code | Own strict list (sycl_strict_fp_args) on every kernel TU; device math libraries stay linked. | ADR-1367 |
| Netflix golden gate | Never built with icx: make test-netflix-golden uses gcc or clang in core/build-golden, because icx contraction drift moves float-motion and float-VIF scores. | ADR-1317, scripts/ci/setup-golden-build.sh |
Windows icx-cl is not covered
The host-math policy applies to the Unix icx driver only. Whether a Windows build with icx-cl links Intel's math library is unmeasured and explicitly deferred (state row T-ICX-CL-WINDOWS-HOST-MATH-2026-10-03 in state.md). The Windows lane is build-only, so no score from it is compared with anything today.
The ocloc offline compiler¶
The default build compiles native Intel GPU code for every target in sycl_icpx_aot_targets (ADR-0568), and icpx hands that work to Intel's ocloc. The Linux oneAPI compiler does not include it (the Windows one does), so meson setup -Denable_sycl=true stops with an error naming this page's remedy when ocloc is not on PATH (ADR-1360).
On Debian or Ubuntu x86-64, install the release the rest of the tree pins:
bash scripts/ci/install-intel-ocloc.sh # sudo is used when not root
ocloc query OCL_DRIVER_VERSION # prints INTEL_NEO_VERSION
The script downloads intel-ocloc and the two Intel Graphics Compiler packages it loads from the intel/compute-runtime release named by INTEL_NEO_VERSION in build-config.env. It checks them against the release's published SHA-256 sums, installs them and proves that ocloc can compile a kernel. Set GITHUB_TOKEN if the anonymous GitHub API limit is exhausted.
On Arch / CachyOS the intel-compute-runtime package ships ocloc. The vmaf-dev-mcp container already has it.
--components widens the set to the rest of the pinned Intel GPU stack. The oneAPI release image uses both wider sets (ADR-1368):
--components | Installs | For |
|---|---|---|
ocloc (default) | intel-ocloc and the two IGC packages | a build host that already has a Level Zero loader |
build | the ocloc set plus the Level Zero loader and its headers (libze1, libze-dev) at LEVEL_ZERO_VERSION | a build host with no other Level Zero source |
runtime | the whole compute runtime (Level Zero GPU driver, OpenCL ICD, gmmlib, IGC, ocloc), the set vmaf-dev-mcp installs, plus libze1 | a host or image that runs SYCL kernels on an Intel GPU |
The Level Zero loader release publishes no checksum file, so the script checks those debs against the SHA-256 digests GitHub records for the release assets.
Install paths¶
oneAPI installs under /opt/intel/oneapi/. Pick one way to get it:
| Option | Command | When |
|---|---|---|
| Intel apt repository script | scripts/ci/install-intel-oneapi.sh --mode=builder | Debian or Ubuntu; installs the pinned build, as CI and the release image do |
| Intel offline installer | see below | Side-by-side with another version, for A/B compiler benchmarks |
| Arch / CachyOS package | sudo pacman -S intel-oneapi-basekit | One global install; the repository lags Intel by one or two releases |
| AUR | paru -S intel-oneapi-basekit-2025 (or yay -S) | One global install, newer than the Arch repository |
The apt script installs the compiler (--mode=builder) or its SYCL runtime (--mode=runtime) from Intel's apt repository at the exact build ONEAPI_APT_VERSION in build-config.env. It trusts only the repository key named by INTEL_ONEAPI_APT_SIGNER_FINGERPRINT. The release image runs it on Debian 13; it works on any Debian or Ubuntu host.
The AUR package conflicts with the Arch repository package and installs to /opt/intel/oneapi/ (not version-suffixed), so it replaces any existing install. Check the packaged version before relying on it: both Arch sources can be older than the pin in build-config.env.
Intel offline installer¶
The offline .sh installer is about 2.6 GB. Download it from Intel oneAPI Base Toolkit downloads; its URL carries a per-build identifier, so copy it from that page. Install into a versioned directory so it does not replace an existing install at /opt/intel/oneapi/:
VERSION=2026.1
sudo sh ./intel-oneapi-base-toolkit-offline.sh \
--silent --eula accept --components all \
--install-dir /opt/intel/oneapi-$VERSION
Activate it for a shell:
Multi-version coexistence¶
A common situation: the Arch repository or AUR install at /opt/intel/oneapi/ is held back at an older release whose device images no longer match the system's level-zero-loader, while a side-by-side install at /opt/intel/oneapi-<version>/ ships matching binaries. Run against the older install, vmaf_bench silently falls back to the host or fails ze_init; the symptom is "ran on CPU even though a GPU is plugged in".
The bench and lint cycle has to point at the install whose runtime matches the loader. The helper scripts/ci/sycl-bench-env.sh resolves the right setvars.sh and emits an eval-able environment block:
List the GPU devices the benchmark sees, then run it on one by index:
The helper looks, in order, for:
$ONEAPI_PREFIX(explicit override);/opt/intel/oneapi-<version>/(canonical side-by-side layout);/opt/intel/oneapi/<version>/(Intel modulefile-style layout);/opt/intel/oneapi/(fallback; warns if the requested version does not match the install actually present).
If none resolves, it exits 1 with a pointer back to this page. It forwards only CMPLR_ROOT, LD_LIBRARY_PATH, LIBRARY_PATH and PATH, the four variables vmaf_bench, icpx and clang-tidy consume, so the parent shell does not accrete the roughly 40 variables setvars.sh exports.
Verify a new compiler¶
After installing a new oneAPI version:
-
Activate it and force a clean SYCL build; the
icpxversion is baked intocompile_commands.json: -
Check that the SYCL kernels still link:
-
Check that the host math functions still come from glibc (ADR-1495). A new driver that renamed or dropped
-no-intel-lib=libimf, or linked Intel's math library another way, would bringlibimfback, and this test fails then:python3 scripts/ci/run_meson_test.py -- -C core/build-sycl-lint \ test_icx_system_libm --print-errorlogsThe test reads
libvmaf.soandvmafwithreadelfand runsvmaf --versionunderLD_DEBUG=bindings. The option is documented inicx --help("Restrict linking of Intel specific libraries. Valid arguments are libirc, libimf, libirng, libsvml") and in the installed man page (share/man/man1/icx.1); check both on a bump.
Verify SYCL clang-tidy¶
The icpx-aware wrapper scripts/ci/clang-tidy-sycl.sh (ADR-0217) injects the oneAPI SYCL include path and __SYCL_DEVICE_ONLY__=0, so stock LLVM clang-tidy resolves <sycl/sycl.hpp> and SYCL translation units lint cleanly:
echo "core/src/sycl/picture_sycl.cpp
core/src/feature/sycl/integer_adm_sycl.cpp
core/src/feature/sycl/integer_motion_sycl.cpp
core/src/feature/sycl/integer_vif_sycl.cpp" \
| parallel -j$(nproc) "scripts/ci/clang-tidy-sycl.sh \
-p core/build-sycl-lint --quiet {}" \
| grep -E "warning:|error:" \
| wc -l
Expected: 0 across all four files.
If the wrapper cannot locate <sycl/sycl.hpp> automatically, point it at the install:
ICPX_ROOT=/opt/intel/oneapi-2026.1/compiler/latest/linux \
scripts/ci/clang-tidy-sycl.sh -p core/build-sycl-lint --quiet <file>
The wrapper also adapts an icx compile database to stock clang. It rewrites -fp-model= to -ffp-model= in a copy of compile_commands.json, and passes -Wno-unknown-warning-option, -Wno-unknown-pragmas and -Wno-overriding-option. The last one is needed because a target that names the strict floating-point arguments next to the project-wide ones compiles with -fp-model=precise -ffp-contract=off twice; clang's driver comments on the repeat with a warning that has no source location, and scripts/ci/tidy-ratchet.py counts a warning it cannot place as a failed translation unit. To check the wrapper after a toolchain change:
Upgrading oneAPI¶
After a major-version oneAPI bump, walk through these items before declaring the bump complete. None block; each is a follow-up backlog candidate.
-
atomic_refperformance: run the SYCL paths ofvmaf_bench(motion_sycl,adm_sycl) on the canonical Arc or Battlemage host and compare per-frame timings against the previous version's numbers. -
sub_group::shuffle_*codegen: sample the IR for the VIF reduction loop ininteger_vif_sycl.cppand check whether the new compiler removes the_mm-style fallback written against older Arc generations. - Sub-group sizes: kernels must declare 16 or 32 through
VmafSyclSubGroupSize, never a rawreqd_sub_group_sizeattribute; check that the compiler still honours both sizes on every AOT target (ADR-1468). -
group_load/group_store(2025.2+): a rewrite of the ADM DWT vertical and horizontal passes onsycl::ext::oneapi::experimental::group_loadwas sketched (Research-0086 §A.4 emitted GO) and then deferred under ADR-0406: theWG_SIZE × ElementsPerWorkItemdivisibility constraint and the multi-row source contiguity gap defeat it. It re-opens when a tile-geometry redesign yields integer divisibility and Xe2 or Battlemage hardware is available to confirm the register-pressure delta. - OpenVINO EP version bump: the ONNX Runtime bundled with the basekit exposes the NPU plugin through
device_type=NPUon the existingOpenVINOExecutionProvider. Done 2026-05-08 in ADR-0405, which adds--tiny-device=openvino-npu(plusopenvino-cpuandopenvino-gpufor explicit device-type pinning). End-to-end NPU validation still needs a contributor with Meteor, Lunar or Arrow Lake hardware. - C++23 surface: C++23 features (
std::expected,std::print,if consteval) are usable with the currenticpxbut not adopted in any fork-local TU. Defer until a clear use case appears, likely the tiny-AI dispatch layer when the NPU EP lands.
CI vs local¶
The Linux SYCL lanes install ${ONEAPI_APT_PACKAGE} from Intel's apt repository; its version is set once, as ONEAPI_VERSION in build-config.env, and can differ from a local install. They install ocloc with scripts/ci/install-intel-ocloc.sh. The Windows MSVC+SYCL lane pins its own offline-installer release in the workflow.
Local-versus-CI divergence on the SYCL kernel binaries is acceptable as long as both build cleanly and the Ubuntu SYCL matrix row stays green. Bit-identical SYCL output is not a guaranteed invariant across Intel oneAPI releases.
Related¶
- ADR-0217: SYCL toolchain consolidation;
icpxis the chosen DPC++ implementation. docs/backends/sycl/overview.md: user-facing SYCL backend reference (--sycl_deviceand so on).docs/backends/sycl/bundling.md: shipping a self-contained binary without an end-user oneAPI install.build-flags.md: theenable_syclMeson option.sycl-toolchains.md: the AdaptiveCpp alternative.