ADRs tagged sycl¶
Auto-generated by scripts/docs/generate-adr-by-tag.sh. Edit ADR Tags: lines to update.
168 ADR(s) carry this tag.
| ID | Title |
|---|---|
| ADR-0022 | Inference runtime is ONNX Runtime via execution providers |
| ADR-0027 | Non-conservative image pins with experimental toolchain flags |
| ADR-0101 | SYCL USM-backed picture pre-allocation pool |
| ADR-0103 | vmaf_sycl_import_d3d11_surface ships as a staging-texture H2D path, not zero-copy |
| ADR-0118 | FFmpeg patches ship as ordered series.txt, not a single carry |
| ADR-0121 | Windows GPU build-only matrix legs (MSVC + CUDA, MSVC + oneAPI SYCL) |
| ADR-0181 | Global feature-characteristics registry + per-backend dispatch strategy |
| ADR-0182 | GPU long-tail batch 1 — psnr + ciede + moment on CUDA / SYCL / Vulkan |
| ADR-0183 | libvmaf_sycl FFmpeg filter — zero-copy QSV / VAAPI import |
| ADR-0188 | GPU long-tail batch 2 — psnr_hvs / ssim / ms_ssim across CUDA / SYCL / Vulkan |
| ADR-0192 | GPU long-tail batch 3 — closing every remaining metric gap (motion_v2 / float_ansnr / ssimulacra2 / cambi + float twins) |
| ADR-0194 | float_ansnr GPU kernels — single-dispatch 3x3 + 5x5 filters with per-WG float partials |
| ADR-0195 | float_psnr GPU kernels — single-dispatch diff² with float partials, bit-exact vs CPU |
| ADR-0196 | float_motion GPU kernels — float twin of integer_motion blur+SAD |
| ADR-0197 | float_vif GPU kernels — 4-scale pyramid with mirror-asymmetry fix |
| ADR-0202 | float_adm CUDA + SYCL twins — sixth Group B float kernel finishes |
| ADR-0206 | ssimulacra2 CUDA + SYCL twins |
| ADR-0214 | GPU-parity CI gate (T6-8) — cross-device variance matrix |
| ADR-0217 | SYCL toolchain cleanup — multi-version recipe + icpx-aware clang-tidy wrapper |
| ADR-0219 | motion3 GPU coverage on Vulkan + CUDA + SYCL (3-frame window) |
| ADR-0220 | SYCL feature kernels are unconditionally fp64-free |
| ADR-0234 | GPU-generation-aware ULP calibration head |
| ADR-0239 | Backend-agnostic GPU picture pool (gpu_picture_pool.{h,c}) |
| ADR-0299 | GPU scoring backend for vmaf-tune (--score-backend) |
| ADR-0315 | Vendor-neutral VVC encode strategy — tiered Tier-1-now / Tier-2-backlog / Tier-3-revisit |
| ADR-0345 | cambi × {CUDA, SYCL, HIP} GPU port strategy |
| ADR-0374 | Build-time-optional public APIs return -ENOSYS when disabled |
| ADR-0406 | Defer SYCL ADM DWT group_load rewrite — divisibility blocker |
| ADR-0407 | AdaptiveCpp as a second SYCL toolchain |
| ADR-0415 | CAMBI SYCL port — closes last CUDA-to-SYCL parity gap |
| ADR-0447 | Motion features under-report on HFR / 50p content |
| ADR-0451 | Local dev-MCP container for live probing |
| ADR-0453 | PSNR enable_chroma option parity across all GPU backends |
| ADR-0458 | SYCL CAMBI queue-sync collapse + SSIM horizontal SLM staging |
| ADR-0460 | Dispatch-strategy registry audit 2026-05-15 |
| ADR-0483 | Extract shared vmaf_gpu_dispatch_parse_env tokenizer |
| ADR-0487 | Wire adm_min_val option into integer_adm GPU backends |
| ADR-0488 | Shared once-snapshot helper for GPU dispatch env variables |
| ADR-0489 | CAMBI SYCL — Replace GPU-to-GPU q.wait() Calls with Event Chains (SY-1) |
| ADR-0514 | dev-MCP container exposes every host GPU backend (CUDA + SYCL + Vulkan + HIP) |
| ADR-0526 | Add enable_lcs and enable_chroma to float_ms_ssim SYCL twin |
| ADR-0529 | Replace /dev/dri/by-path bind with whole /dev/dri bind in dev container |
| ADR-0541 | Pin dev-MCP container Intel NEO + ROCm runtimes to versions matching the host kernel |
| ADR-0542 | Full GPU backend plumbing in the dev-mcp container |
| ADR-0544 | deduplicate feature_extractor_list[] registrations |
| ADR-0564 | Real integer_ssim GPU kernels (CUDA, HIP, SYCL) — replace silent float_ssim substitution |
| ADR-0567 | Real On-Device GPU Kernels for speed_chroma and speed_temporal (4 Backends) |
| ADR-0568 | Default sycl_icpx_aot_targets to full Intel arch list |
| ADR-0576 | ffmpeg-patches n8.1.1 full-feature-exposure sync |
| ADR-0582 | MS-SSIM enable_db and clip_db option parity on CUDA and SYCL backends |
| ADR-0590 | Wire enable_db / clip_db into the CUDA and SYCL MS-SSIM twins |
| ADR-0599 | Cross-Backend Parity Audit — Full Extractor Matrix (2026-05-18) |
| ADR-0605 | Renovate customManagers for all dev/Containerfile pinned dependencies |
| ADR-0662 | Vulkan Motion Lavapipe Parity |
| ADR-0667 | vmaf-tune score backend native priority |
| ADR-0699 | VMAFX Helm Chart and Kubernetes Manifests with 3-Vendor GPU Device-Plugin Support |
| ADR-0777 | Thread-Safety Audit — CUDA / SYCL / HIP Backends |
| ADR-0780 | NOLINT Cluster Refactor — Slab Allocator, SYCL Stride, ADM Band-Size |
| ADR-0787 | 0787-libvmaf-api-error-path-audit.md |
| ADR-0839 | C++23 wave — shadow-identifier and implicit-cast cleanup |
| ADR-0841 | Environment variable reference page and canonical naming |
| ADR-0868 | GPU backend kernel parity-test coverage gap-fill |
| ADR-0876 | Adopt <inttypes.h> PRI macros for fixed-width integer printf formatting |
| ADR-0884 | SYCL kernel coverage round 2 — five additional CPU-vs-SYCL parity gates |
| ADR-0928 | VmafPicture v2 — explicit per-backend GPU state |
| ADR-0946 | SYCL kernel coverage round 3 (float family + PSNR-HVS) |
| ADR-0954 | Host-only unit test for shared GPU dispatch runtime |
| ADR-0957 | SYCL kernel coverage round 4 (float_moment + SpEED + SSIMULACRA2) |
| ADR-0964 | Implement speed_internal.c and wire speed_{chroma,temporal}_{hip,sycl} |
| ADR-0982 | GPU runtime bug audit — round 26 (init/teardown leak sweep) |
| ADR-0985 | SYCL parity divergence investigation — float_ssim + ssimulacra2 on Arc A380 |
| ADR-0989 | Wire motion_add_uv through integer_motion_sycl; emit warning on motion_five_frame_window |
| ADR-1001 | SYCL parity round 5 — CAMBI CPU vs. SYCL parity gate |
| ADR-1026 | R6 SYCL kernel correctness — rd-stride OOB and unchecked graph_wait |
| ADR-1034 | Fix SYCL integer_vif rd_stride OOB on odd widths and integer_motion UV queue sync gap |
| ADR-1078 | ms_ssim option parity across HIP and SYCL backends |
| ADR-1093 | Disable two recurring-failure tests via should_fail while root cause is under investigation |
| ADR-1100 | Skip GPU-flagged extractors when flags == 0 in vmaf_get_feature_extractor_by_feature_name |
| ADR-1121 | SYCL QSV zero-copy — P010 pixel normalization and separate-session decode contract |
| ADR-1122 | Adopt and port VMAF v1 models (opt-in, v0.6.1 stays default) |
| ADR-1142 | Whole-codebase standards; lint debt only ratchets down |
| ADR-1143 | CUDA and Intel SYCL Backend Gap Closure |
| ADR-1145 | Derive the Intel NEO compute stack (gmmlib and IGC) dynamically from pinned compute-runtime release metadata |
| ADR-1179 | Fix Intel Arc SYCL Crashes and Default Model Resolution Divergence |
| ADR-1185 | Per-backend performance baselines are median-of-N, one backend per build dir |
| ADR-1191 | Integer ADM rejects CSF configurations its fixed-point storage cannot represent |
| ADR-1192 | Keep the recorded Netflix benchmark snapshot; do not regenerate it while the GPU paths are broken |
| ADR-1193 | Opt-in uncapped option splits the PSNR infinity sentinel from the truncation |
| ADR-1194 | One integer-ADM angle_flag predicate for every backend |
| ADR-1197 | The threaded flush leaves GPU extractors to their own backend flush |
| ADR-1202 | GPU SpEED-chroma twins report singularity separately from failure |
| ADR-1204 | GPU ADM contrast-masking twins clamp the far edge instead of mirroring it |
| ADR-1205 | The ssimulacra2 FMA unification extends to the scalar fallback and every GPU host copy |
| ADR-1210 | The SYCL integer-ADM contrast-masking kernel mirrors its near edge |
| ADR-1212 | The GPU float_moment twins normalise by the bit-depth scaler on the host |
| ADR-1214 | The float-ADM GPU twins ignore adm_csf_scale in Watson mode and share the CPU's option aliases |
| ADR-1216 | The GPU motion3 twins apply motion_fps_weight exactly once |
| ADR-1217 | The GPU float-VIF kernels read vif_sigma_nsq and vif_enhn_gain_limit from their options |
| ADR-1218 | The GPU SpEED twins zero the device solution and report singularity from the temporal path |
| ADR-1220 | The GPU float-ADM kernels honour adm_p_norm, adm_bypass_cm and adm_skip_scale0 |
| ADR-1221 | clip_db is a ceiling on the MS-SSIM dB output, not a clamp on the linear score |
| ADR-1228 | A recurring "faster than upstream, and still exact" milestone |
| ADR-1290 | The tidy ratchet owns its per-lane compilation database |
| ADR-1296 | GPU init failure paths are tested device-free, by compiling the backend TU against runtime stubs |
| ADR-1299 | Compute MS-SSIM chroma on SYCL |
| ADR-1325 | Normalize integer ADM Barten weights with one exponent per scale |
| ADR-1326 | Use a fixed-point oracle for SYCL motion-add-UV parity |
| ADR-1357 | Run the SYCL CAMBI extractor entirely on the device |
| ADR-1358 | The SYCL SpEED twins are device-resident, with the 25x25 linear algebra on the device and exact fp32 arithmetic |
| ADR-1360 | Generate SYCL AOT images at compile time and fail the build when they are missing |
| ADR-1362 | The SYCL integer ADM twin computes AIM on the device and finalises every ADM output in the CPU's float arithmetic |
| ADR-1363 | The SYCL ssimulacra2 twin is device-resident, and float_ms_ssim_sycl waits once per frame |
| ADR-1364 | Register the SYCL device images of a Windows MSVC build through one explicit device link |
| ADR-1365 | SYCL PSNR, SSIM and float-motion twins take the CPU option tables |
| ADR-1367 | Every SYCL feature TU compiles with contraction off and correctly rounded fp32 division and square root |
| ADR-1368 | Build and run the oneAPI release image on Debian 13 with pinned Intel packages |
| ADR-1369 | SYCL twins read the planes the state uploads once per frame; opt-in shared chroma planes and a device-side slot fence |
| ADR-1370 | float_ssim_sycl decimates on the device, bit-identical to the CPU |
| ADR-1371 | SYCL motion differences the frames before the blur, in one shared kernel |
| ADR-1386 | One script runs the home GPU box's RC3 verify commands, row by row, under per-device locks |
| ADR-1395 | SYCL kernels use no scratch memory on Intel GPUs |
| ADR-1401 | psnr_hvs_sycl and psnr_hvs_hip return the CPU's scores bit for bit; the fp64-free masking threshold is an integer square root |
| ADR-1402 | Integer ADM keeps the scale-0 masking centre tap in int32 and clamps the excess in int64 |
| ADR-1410 | SYCL CLI picture pool allocates pinned host USM to bypass staging upload |
| ADR-1411 | float_motion_sycl adds its SAD in the CPU's order and returns the CPU's scores bit for bit |
| ADR-1413 | Every integer ADM implementation bounds the enhancement gain with the scalar's truncated double product |
| ADR-1414 | float_ms_ssim_sycl computes the CPU's arithmetic and returns its per-scale means bit for bit |
| ADR-1418 | Motion parity cells compare what every twin emits; a missing metric is a cell error |
| ADR-1422 | float_vif_sycl computes the CPU's arithmetic without an fp64 type and returns its scores bit for bit |
| ADR-1432 | vif_sycl computes the gain terms of the integer VIF in exact integer arithmetic and returns the CPU's scores bit for bit |
| ADR-1434 | float_adm_sycl computes the CPU's arithmetic without an fp64 type, adds in the CPU's order and returns the CPU's scores bit for bit |
| ADR-1436 | ciede_sycl runs the CPU's statements on fp32 pairs and adds in the CPU's order; it lands where the CUDA twin does, 1.4e-11 from the CPU |
| ADR-1443 | integer_ssim_sycl computes the CPU's fp64 term in 64-bit integers and adds in the CPU's order; it returns the CPU's ssim bit for bit |
| ADR-1446 | ssimulacra2_sycl forms the CPU's fp64 terms in 64-bit integers and returns the sums of the CPU's loops; it returns the CPU's score bit for bit |
| ADR-1448 | ciede_hip runs the CPU's arithmetic in fp32 pairs, from the header the SYCL twin runs; what remains is glibc's powf and the last bits of a pair |
| ADR-1449 | float_moment_sycl adds the float squares the CPU adds, and is bit-identical while the CPU's own sum is exact |
| ADR-1450 | float_psnr_sycl adds its squared differences as integers, and is bit-identical to the CPU |
| ADR-1451 | adm_sycl, motion_sycl, motion_v2_sycl, psnr_sycl, float_ssim_sycl and cambi_sycl are declared exact twins; speed_chroma_sycl is not |
| ADR-1460 | speed_temporal becomes a parity-gate feature with a derived bound, and every registered twin of a gated backend has to be a gate cell |
| ADR-1463 | float_ssim_sycl forms the CPU's fp64 terms in 64-bit integers and adds them on the host in the CPU's raster order |
| ADR-1466 | float_ms_ssim_sycl stores every window's l, c and s of every scale and adds them on the host in the CPU's raster order |
| ADR-1468 | A SYCL kernel requires only a sub-group size every default AOT target accepts (16 or 32); the six kernels that required 8 move to 16 |
| ADR-1472 | The integer ADM weight limits follow from the contrast-masking cube and the largest wavelet coefficient of a scale |
| ADR-1475 | The quantisation step of integer ADM is Netflix's expression again, so vmaf_v0.6.1 returns Netflix master's bits |
| ADR-1477 | SpEED evaluates Netflix's fp64 expressions again, and its GPU twins form the entropies and the score on the host with the same statements, so every twin returns the CPU's scores bit for bit on any C library |
| ADR-1489 | The CSF weights of float ADM are Netflix's float arithmetic again, so float_adm differs from Netflix by the division alone |
| ADR-1491 | The CUDA, SYCL and HIP motion twins compute motion_five_frame_window: the frame two back on the device, the CPU's window function on the host |
| ADR-1495 | icx and icpx builds link glibc's libm, not Intel's libimf |
| ADR-1497 | The float_moment twins form the CPU's rounded second-moment sum past 2^53 units, and return the CPU's bits on every frame |
| ADR-1499 | The float_psnr twins add each row's exact sum in the CPU's order, and return the CPU's bits on every frame |
| ADR-1501 | The float_adm_sycl term kernel takes the large register file and leaves the sub-group size to the compiler, so it uses no scratch memory on Xe2 |
| ADR-1505 | An Intel GPU tester image with a backend-neutral GPU section in the report |
| ADR-1517 | The production GPU images are built on Debian 13, ship only the vendor files libvmaf loads, and share the tester images' licence records |
| ADR-1566 | A Windows SYCL tester zip built with /MD that carries its runtime beside every program and measures every SYCL twin on a tester's Intel GPU |
| ADR-1571 | a GPU dispatch variable is documented only when library code reads it; VMAF_CUDA_DISPATCH is read at extractor init, VMAF_HIP_DISPATCH and its predicate are removed |
| ADR-1590 | Every build stores its GPU device code compressed at the toolchain's strongest setting, and the build refuses raw device code |
| ADR-1685 | A post-1.0 embedding milestone: zero-copy device-frame import with fences, asynchronous window scores, and an unchanged licence |
| ADR-1688 | The SYCL zero-copy path admits only extractors that compute from the shared luma, and names every other one |
| ADR-1761 | The libvmaf_sycl filter retries a failed VA import, then stops naming the frame; each input is imported with its own VA display |
| ADR-1829 | RC4 owns the whole device-memory import API, not only the first full Rust metric |
| ADR-1830 | vif_sycl runs at SIMD-16 only; the SIMD-32 kernels and VMAF_SYCL_VIF_SUBGROUP_SIZE are removed |
| ADR-1874 | The vmaf CLI reports its usable backends; the score-backend selectors read that report |
| ADR-1900 | Deterministic verification and state contract for VPL decode retry ceiling |
| ADR-1917 | The integer ADM scale-0 horizontal and vertical weight limit is 43900, set by the CSF stage's 16-bit magnitude |
| ADR-1918 | Samples above 2^bpc - 1 are invalid input; an opt-in check refuses them |
| ADR-1929 | VMAFx device frames, fences and the import rule: the shared contract the backend lanes implement |
| ADR-1930 | sycl_device_asan puts the device sanitizer on every SYCL compile and the link |
| ADR-2091 | VMAFx device frames on SYCL: readers copy on the device behind the frame's ready event, the release waits on every reader |