Skip to content

ADRs tagged sycl

Auto-generated by scripts/docs/generate-adr-by-tag.sh. Edit ADR Tags: lines to update.

168 ADR(s) carry this tag.

ID Title
ADR-0022 Inference runtime is ONNX Runtime via execution providers
ADR-0027 Non-conservative image pins with experimental toolchain flags
ADR-0101 SYCL USM-backed picture pre-allocation pool
ADR-0103 vmaf_sycl_import_d3d11_surface ships as a staging-texture H2D path, not zero-copy
ADR-0118 FFmpeg patches ship as ordered series.txt, not a single carry
ADR-0121 Windows GPU build-only matrix legs (MSVC + CUDA, MSVC + oneAPI SYCL)
ADR-0181 Global feature-characteristics registry + per-backend dispatch strategy
ADR-0182 GPU long-tail batch 1 — psnr + ciede + moment on CUDA / SYCL / Vulkan
ADR-0183 libvmaf_sycl FFmpeg filter — zero-copy QSV / VAAPI import
ADR-0188 GPU long-tail batch 2 — psnr_hvs / ssim / ms_ssim across CUDA / SYCL / Vulkan
ADR-0192 GPU long-tail batch 3 — closing every remaining metric gap (motion_v2 / float_ansnr / ssimulacra2 / cambi + float twins)
ADR-0194 float_ansnr GPU kernels — single-dispatch 3x3 + 5x5 filters with per-WG float partials
ADR-0195 float_psnr GPU kernels — single-dispatch diff² with float partials, bit-exact vs CPU
ADR-0196 float_motion GPU kernels — float twin of integer_motion blur+SAD
ADR-0197 float_vif GPU kernels — 4-scale pyramid with mirror-asymmetry fix
ADR-0202 float_adm CUDA + SYCL twins — sixth Group B float kernel finishes
ADR-0206 ssimulacra2 CUDA + SYCL twins
ADR-0214 GPU-parity CI gate (T6-8) — cross-device variance matrix
ADR-0217 SYCL toolchain cleanup — multi-version recipe + icpx-aware clang-tidy wrapper
ADR-0219 motion3 GPU coverage on Vulkan + CUDA + SYCL (3-frame window)
ADR-0220 SYCL feature kernels are unconditionally fp64-free
ADR-0234 GPU-generation-aware ULP calibration head
ADR-0239 Backend-agnostic GPU picture pool (gpu_picture_pool.{h,c})
ADR-0299 GPU scoring backend for vmaf-tune (--score-backend)
ADR-0315 Vendor-neutral VVC encode strategy — tiered Tier-1-now / Tier-2-backlog / Tier-3-revisit
ADR-0345 cambi × {CUDA, SYCL, HIP} GPU port strategy
ADR-0374 Build-time-optional public APIs return -ENOSYS when disabled
ADR-0406 Defer SYCL ADM DWT group_load rewrite — divisibility blocker
ADR-0407 AdaptiveCpp as a second SYCL toolchain
ADR-0415 CAMBI SYCL port — closes last CUDA-to-SYCL parity gap
ADR-0447 Motion features under-report on HFR / 50p content
ADR-0451 Local dev-MCP container for live probing
ADR-0453 PSNR enable_chroma option parity across all GPU backends
ADR-0458 SYCL CAMBI queue-sync collapse + SSIM horizontal SLM staging
ADR-0460 Dispatch-strategy registry audit 2026-05-15
ADR-0483 Extract shared vmaf_gpu_dispatch_parse_env tokenizer
ADR-0487 Wire adm_min_val option into integer_adm GPU backends
ADR-0488 Shared once-snapshot helper for GPU dispatch env variables
ADR-0489 CAMBI SYCL — Replace GPU-to-GPU q.wait() Calls with Event Chains (SY-1)
ADR-0514 dev-MCP container exposes every host GPU backend (CUDA + SYCL + Vulkan + HIP)
ADR-0526 Add enable_lcs and enable_chroma to float_ms_ssim SYCL twin
ADR-0529 Replace /dev/dri/by-path bind with whole /dev/dri bind in dev container
ADR-0541 Pin dev-MCP container Intel NEO + ROCm runtimes to versions matching the host kernel
ADR-0542 Full GPU backend plumbing in the dev-mcp container
ADR-0544 deduplicate feature_extractor_list[] registrations
ADR-0564 Real integer_ssim GPU kernels (CUDA, HIP, SYCL) — replace silent float_ssim substitution
ADR-0567 Real On-Device GPU Kernels for speed_chroma and speed_temporal (4 Backends)
ADR-0568 Default sycl_icpx_aot_targets to full Intel arch list
ADR-0576 ffmpeg-patches n8.1.1 full-feature-exposure sync
ADR-0582 MS-SSIM enable_db and clip_db option parity on CUDA and SYCL backends
ADR-0590 Wire enable_db / clip_db into the CUDA and SYCL MS-SSIM twins
ADR-0599 Cross-Backend Parity Audit — Full Extractor Matrix (2026-05-18)
ADR-0605 Renovate customManagers for all dev/Containerfile pinned dependencies
ADR-0662 Vulkan Motion Lavapipe Parity
ADR-0667 vmaf-tune score backend native priority
ADR-0699 VMAFX Helm Chart and Kubernetes Manifests with 3-Vendor GPU Device-Plugin Support
ADR-0777 Thread-Safety Audit — CUDA / SYCL / HIP Backends
ADR-0780 NOLINT Cluster Refactor — Slab Allocator, SYCL Stride, ADM Band-Size
ADR-0787 0787-libvmaf-api-error-path-audit.md
ADR-0839 C++23 wave — shadow-identifier and implicit-cast cleanup
ADR-0841 Environment variable reference page and canonical naming
ADR-0868 GPU backend kernel parity-test coverage gap-fill
ADR-0876 Adopt <inttypes.h> PRI macros for fixed-width integer printf formatting
ADR-0884 SYCL kernel coverage round 2 — five additional CPU-vs-SYCL parity gates
ADR-0928 VmafPicture v2 — explicit per-backend GPU state
ADR-0946 SYCL kernel coverage round 3 (float family + PSNR-HVS)
ADR-0954 Host-only unit test for shared GPU dispatch runtime
ADR-0957 SYCL kernel coverage round 4 (float_moment + SpEED + SSIMULACRA2)
ADR-0964 Implement speed_internal.c and wire speed_{chroma,temporal}_{hip,sycl}
ADR-0982 GPU runtime bug audit — round 26 (init/teardown leak sweep)
ADR-0985 SYCL parity divergence investigation — float_ssim + ssimulacra2 on Arc A380
ADR-0989 Wire motion_add_uv through integer_motion_sycl; emit warning on motion_five_frame_window
ADR-1001 SYCL parity round 5 — CAMBI CPU vs. SYCL parity gate
ADR-1026 R6 SYCL kernel correctness — rd-stride OOB and unchecked graph_wait
ADR-1034 Fix SYCL integer_vif rd_stride OOB on odd widths and integer_motion UV queue sync gap
ADR-1078 ms_ssim option parity across HIP and SYCL backends
ADR-1093 Disable two recurring-failure tests via should_fail while root cause is under investigation
ADR-1100 Skip GPU-flagged extractors when flags == 0 in vmaf_get_feature_extractor_by_feature_name
ADR-1121 SYCL QSV zero-copy — P010 pixel normalization and separate-session decode contract
ADR-1122 Adopt and port VMAF v1 models (opt-in, v0.6.1 stays default)
ADR-1142 Whole-codebase standards; lint debt only ratchets down
ADR-1143 CUDA and Intel SYCL Backend Gap Closure
ADR-1145 Derive the Intel NEO compute stack (gmmlib and IGC) dynamically from pinned compute-runtime release metadata
ADR-1179 Fix Intel Arc SYCL Crashes and Default Model Resolution Divergence
ADR-1185 Per-backend performance baselines are median-of-N, one backend per build dir
ADR-1191 Integer ADM rejects CSF configurations its fixed-point storage cannot represent
ADR-1192 Keep the recorded Netflix benchmark snapshot; do not regenerate it while the GPU paths are broken
ADR-1193 Opt-in uncapped option splits the PSNR infinity sentinel from the truncation
ADR-1194 One integer-ADM angle_flag predicate for every backend
ADR-1197 The threaded flush leaves GPU extractors to their own backend flush
ADR-1202 GPU SpEED-chroma twins report singularity separately from failure
ADR-1204 GPU ADM contrast-masking twins clamp the far edge instead of mirroring it
ADR-1205 The ssimulacra2 FMA unification extends to the scalar fallback and every GPU host copy
ADR-1210 The SYCL integer-ADM contrast-masking kernel mirrors its near edge
ADR-1212 The GPU float_moment twins normalise by the bit-depth scaler on the host
ADR-1214 The float-ADM GPU twins ignore adm_csf_scale in Watson mode and share the CPU's option aliases
ADR-1216 The GPU motion3 twins apply motion_fps_weight exactly once
ADR-1217 The GPU float-VIF kernels read vif_sigma_nsq and vif_enhn_gain_limit from their options
ADR-1218 The GPU SpEED twins zero the device solution and report singularity from the temporal path
ADR-1220 The GPU float-ADM kernels honour adm_p_norm, adm_bypass_cm and adm_skip_scale0
ADR-1221 clip_db is a ceiling on the MS-SSIM dB output, not a clamp on the linear score
ADR-1228 A recurring "faster than upstream, and still exact" milestone
ADR-1290 The tidy ratchet owns its per-lane compilation database
ADR-1296 GPU init failure paths are tested device-free, by compiling the backend TU against runtime stubs
ADR-1299 Compute MS-SSIM chroma on SYCL
ADR-1325 Normalize integer ADM Barten weights with one exponent per scale
ADR-1326 Use a fixed-point oracle for SYCL motion-add-UV parity
ADR-1357 Run the SYCL CAMBI extractor entirely on the device
ADR-1358 The SYCL SpEED twins are device-resident, with the 25x25 linear algebra on the device and exact fp32 arithmetic
ADR-1360 Generate SYCL AOT images at compile time and fail the build when they are missing
ADR-1362 The SYCL integer ADM twin computes AIM on the device and finalises every ADM output in the CPU's float arithmetic
ADR-1363 The SYCL ssimulacra2 twin is device-resident, and float_ms_ssim_sycl waits once per frame
ADR-1364 Register the SYCL device images of a Windows MSVC build through one explicit device link
ADR-1365 SYCL PSNR, SSIM and float-motion twins take the CPU option tables
ADR-1367 Every SYCL feature TU compiles with contraction off and correctly rounded fp32 division and square root
ADR-1368 Build and run the oneAPI release image on Debian 13 with pinned Intel packages
ADR-1369 SYCL twins read the planes the state uploads once per frame; opt-in shared chroma planes and a device-side slot fence
ADR-1370 float_ssim_sycl decimates on the device, bit-identical to the CPU
ADR-1371 SYCL motion differences the frames before the blur, in one shared kernel
ADR-1386 One script runs the home GPU box's RC3 verify commands, row by row, under per-device locks
ADR-1395 SYCL kernels use no scratch memory on Intel GPUs
ADR-1401 psnr_hvs_sycl and psnr_hvs_hip return the CPU's scores bit for bit; the fp64-free masking threshold is an integer square root
ADR-1402 Integer ADM keeps the scale-0 masking centre tap in int32 and clamps the excess in int64
ADR-1410 SYCL CLI picture pool allocates pinned host USM to bypass staging upload
ADR-1411 float_motion_sycl adds its SAD in the CPU's order and returns the CPU's scores bit for bit
ADR-1413 Every integer ADM implementation bounds the enhancement gain with the scalar's truncated double product
ADR-1414 float_ms_ssim_sycl computes the CPU's arithmetic and returns its per-scale means bit for bit
ADR-1418 Motion parity cells compare what every twin emits; a missing metric is a cell error
ADR-1422 float_vif_sycl computes the CPU's arithmetic without an fp64 type and returns its scores bit for bit
ADR-1432 vif_sycl computes the gain terms of the integer VIF in exact integer arithmetic and returns the CPU's scores bit for bit
ADR-1434 float_adm_sycl computes the CPU's arithmetic without an fp64 type, adds in the CPU's order and returns the CPU's scores bit for bit
ADR-1436 ciede_sycl runs the CPU's statements on fp32 pairs and adds in the CPU's order; it lands where the CUDA twin does, 1.4e-11 from the CPU
ADR-1443 integer_ssim_sycl computes the CPU's fp64 term in 64-bit integers and adds in the CPU's order; it returns the CPU's ssim bit for bit
ADR-1446 ssimulacra2_sycl forms the CPU's fp64 terms in 64-bit integers and returns the sums of the CPU's loops; it returns the CPU's score bit for bit
ADR-1448 ciede_hip runs the CPU's arithmetic in fp32 pairs, from the header the SYCL twin runs; what remains is glibc's powf and the last bits of a pair
ADR-1449 float_moment_sycl adds the float squares the CPU adds, and is bit-identical while the CPU's own sum is exact
ADR-1450 float_psnr_sycl adds its squared differences as integers, and is bit-identical to the CPU
ADR-1451 adm_sycl, motion_sycl, motion_v2_sycl, psnr_sycl, float_ssim_sycl and cambi_sycl are declared exact twins; speed_chroma_sycl is not
ADR-1460 speed_temporal becomes a parity-gate feature with a derived bound, and every registered twin of a gated backend has to be a gate cell
ADR-1463 float_ssim_sycl forms the CPU's fp64 terms in 64-bit integers and adds them on the host in the CPU's raster order
ADR-1466 float_ms_ssim_sycl stores every window's l, c and s of every scale and adds them on the host in the CPU's raster order
ADR-1468 A SYCL kernel requires only a sub-group size every default AOT target accepts (16 or 32); the six kernels that required 8 move to 16
ADR-1472 The integer ADM weight limits follow from the contrast-masking cube and the largest wavelet coefficient of a scale
ADR-1475 The quantisation step of integer ADM is Netflix's expression again, so vmaf_v0.6.1 returns Netflix master's bits
ADR-1477 SpEED evaluates Netflix's fp64 expressions again, and its GPU twins form the entropies and the score on the host with the same statements, so every twin returns the CPU's scores bit for bit on any C library
ADR-1489 The CSF weights of float ADM are Netflix's float arithmetic again, so float_adm differs from Netflix by the division alone
ADR-1491 The CUDA, SYCL and HIP motion twins compute motion_five_frame_window: the frame two back on the device, the CPU's window function on the host
ADR-1495 icx and icpx builds link glibc's libm, not Intel's libimf
ADR-1497 The float_moment twins form the CPU's rounded second-moment sum past 2^53 units, and return the CPU's bits on every frame
ADR-1499 The float_psnr twins add each row's exact sum in the CPU's order, and return the CPU's bits on every frame
ADR-1501 The float_adm_sycl term kernel takes the large register file and leaves the sub-group size to the compiler, so it uses no scratch memory on Xe2
ADR-1505 An Intel GPU tester image with a backend-neutral GPU section in the report
ADR-1517 The production GPU images are built on Debian 13, ship only the vendor files libvmaf loads, and share the tester images' licence records
ADR-1566 A Windows SYCL tester zip built with /MD that carries its runtime beside every program and measures every SYCL twin on a tester's Intel GPU
ADR-1571 a GPU dispatch variable is documented only when library code reads it; VMAF_CUDA_DISPATCH is read at extractor init, VMAF_HIP_DISPATCH and its predicate are removed
ADR-1590 Every build stores its GPU device code compressed at the toolchain's strongest setting, and the build refuses raw device code
ADR-1685 A post-1.0 embedding milestone: zero-copy device-frame import with fences, asynchronous window scores, and an unchanged licence
ADR-1688 The SYCL zero-copy path admits only extractors that compute from the shared luma, and names every other one
ADR-1761 The libvmaf_sycl filter retries a failed VA import, then stops naming the frame; each input is imported with its own VA display
ADR-1829 RC4 owns the whole device-memory import API, not only the first full Rust metric
ADR-1830 vif_sycl runs at SIMD-16 only; the SIMD-32 kernels and VMAF_SYCL_VIF_SUBGROUP_SIZE are removed
ADR-1874 The vmaf CLI reports its usable backends; the score-backend selectors read that report
ADR-1900 Deterministic verification and state contract for VPL decode retry ceiling
ADR-1917 The integer ADM scale-0 horizontal and vertical weight limit is 43900, set by the CSF stage's 16-bit magnitude
ADR-1918 Samples above 2^bpc - 1 are invalid input; an opt-in check refuses them
ADR-1929 VMAFx device frames, fences and the import rule: the shared contract the backend lanes implement
ADR-1930 sycl_device_asan puts the device sanitizer on every SYCL compile and the link
ADR-2091 VMAFx device frames on SYCL: readers copy on the device behind the frame's ready event, the release waits on every reader