Rebase-sensitive invariants¶
Cross-package invariants that any upstream-sync or rebase agent must preserve. Referenced from the canonical AGENTS.md harness. Per-subtree detail (the load-bearing reasons and mechanics) lives in the AGENTS.md under each subtree; this page is the index. When a rebase touches a cited translation unit, read that subtree harness before resolving conflicts. A subtree AGENTS.md with an AGENTS.d/ next to it is a generated index: read the pages its table names for the paths you touch, and add an invariant as a page there (agents index and topic pages).
The invariants are grouped by area. A GPU or SIMD twin that returns the CPU extractor's bits has one entry per backend; find the feature's section, then the backend within it.
- Documentation and site: entry points, the site toolchain, navigation and search
- Build, test and CI: test-runner sanitisation, editor settings, coverage and CI pins, the dev container, nox, security evidence
- Upstream sync and provenance: the recorded upstream head, deliberate deviations, upstream ports
- Backends, extractors and the parity gate: which backends and extractors exist, how twins are registered and declared exact
- Floating-point policy and device contracts: contraction, libm, strict-FP lines, SYCL scratch / fp64 / sub-group rules
- psnr_hvs: CUDA, SYCL and HIP twins, the masking threshold
- psnr and float_moment: integer block sums, float squares, the NEON / SVE2 order
- SpEED and CAMBI: fp64 expressions, device-resident pipelines, the parity fixture
- SSIMULACRA 2: CPU sum order and single-readback pipelines
- SSIM and MS-SSIM: decimation, raster-order sums, option parity
- Motion: SAD order and the five-frame window
- VIF: statistic arithmetic, the log2 table
- ADM: integer and float ADM, AIM, the divide, row rounding
- CIEDE2000: CPU arithmetic on each backend
Documentation and site¶
-
Documentation entry points: keep
README.mdconcise and link to the topic guides for changing build requirements, backend coverage and model defaults.docs/index.mdanddocs/backends/index.mdshould link to backend guides rather than repeat kernel counts or maturity summaries. Keep the repository-root build instructions indocs/getting-started/index.mdand include Meson'score/source directory when showing a configure command. -
Documentation site design layer (ADR-1508): the design is
docs/stylesheets/vmafx.cssand nothing else: no template override, no Material colour name and nofont:block inmkdocs.yml, so Zensical'sclassicvariant can render it later. The fonts indocs/assets/fonts/match theirvendor.json(scripts/docs/check_vendored_assets.py, run bymake docs-fragments-check). The landing page keeps itsvx-*wrappers when its text changes. See Documentation site design. -
Documentation charts (ADR-1508):
scripts/docs/generate-charts.pywrites each chart's data, both SVG renders, the page block between itsCHARTsentinels and the vendored Vega bundle;make docs-fragments-checkcompares them with a fresh build and the docs CI jobs re-render with--require-render. Regenerate after a change toscripts/ci/exact_twins.d/,LIBM_TWINS,scripts/ci/upstream_parity.d/or the 576x324 snapshots; keep the sentinels when a page is rewritten. See Documentation site design. -
Documentation diagrams (ADR-1508): diagrams are figure specs in
docs/figures/rendered intodocs/assets/figures/bytools/figures/, shown through the MkDocs hook listed inmkdocs.yml;make docs-figuresholds the renders to their specs and every evidence anchor to the code. No Mermaid fence and no ASCII diagram returns on a page that has a figure. See Documentation site design. -
ADR navigation collapsed (ADR-1510): the
ADRsentry ofmkdocs.ymllists the ADR index, the template and the tag index only; noADR-NAV-GENERATEDblock and noscripts/docs/generate-adr-nav.sh.test_adr_navigation_is_collapsedinscripts/docs/tests/test_generators.pyguards it. See scripts/AGENTS.md. -
Site search covers user pages only (ADR-1512):
.meta.ymlfiles underdocs/adr/,docs/research/anddocs/changelog-archive/setsearch.excludethrough thematerial/metaplugin; the index pages and the generated title lists (scripts/docs/generate-record-titles.py) override it, anddocs/rebase-notes.mdanddocs/state.mdkeep their front matter.scripts/docs/check_search_scope.pyreads the built index. See Documentation site design.
Build, test and CI¶
-
Root licence files (ADR-1699): the root holds
LICENSE(the EUPL-1.2, byte for byteLICENSES/EUPL-1.2.txt) andNOTICE(Netflix'sLICENSE, unchanged). An upstream sync that changes Netflix'sLICENSEapplies it toNOTICE; no sync brings backLICENSE-MITor another root licence file. Package licence fields name the licences of the files the package ships, and fork.tomlfiles carry an EUPL-1.2 header. The rootgo.modkeeps itsretract [v1.0.0-rc.1, v1.0.0-rc.2].scripts/ci/check_licence_metadata.py(requiredLicence Provenancejob and a pre-commit hook) refuses each breach. See scripts/ci/AGENTS.md. -
Meson test secret environment sanitization (ADR-1333):
scripts/ci/run_meson_test.pydeletes sensitive GitHub credential keys before Meson starts and records its raw parent environment intestlog.txt. Every supported Make, workflow, preflight, bisection, setup-guidance, and Zed entry point must remain on that wrapper.core/meson.buildretains a default test setup usingenvironment().unset()for (GITHUB_PERSONAL_ACCESS_TOKEN,GITHUB_TOKEN,GH_TOKEN,GH_ENTERPRISE_TOKEN,GITHUB_ENTERPRISE_TOKEN,GITHUB_PAT,GH_PAT,GITHUB_AUTH_TOKEN,GITHUB_API_TOKEN,HOMEBREW_GITHUB_API_TOKEN,ACTIONS_ID_TOKEN_REQUEST_TOKEN,ACTIONS_RUNTIME_TOKEN) at the child and JSON-log layer. The regression contract rejects raw supported-entry-point bypasses, alternate setups, and explicit forbidden-name reintroduction. Preserve the runner, callers, setup, andcore/test/test_meson_secret_env_sanitization.pytogether. Raw external Meson/Ninja test commands are outside this bounded guarantee. -
Zed project settings are project-scoped:
.zed/settings.jsonis parsed as Zed'sProjectSettingsContent, so it must not regainagent,agent_servers, provider/model pins, or permission policy. Preserve the currentdocker exec -i vmaf-dev-mcp vmafx-mcpcontext-server entry, the threeStandards:tasks, and the contract inscripts/ci/tests/test_zed_project_config.py. The scoped mechanics live in.zed/AGENTS.md. -
-qpfileon libx264 goes throughquant_offsets(ADR-2167): patch0007's libx264 hunks load the file withff_qpfile_load()inX264_init()(refusingaq-mode=0and a grid that is not the video's macroblock grid) and add the deltas of input framento the picture'squant_offsetsinsetup_frame(); libx264 has noqpfileparameter, sox264_param_parse(.., "qpfile", ..)must not return.ffmpeg-patches/test/qpfile_check.pyguards it, and the saliency code ofpkg/saliencyandtools/vmaf-tunepasses-qpfile, not-x264-params qpfile=. - Encoder-parameter options are merged, shot starts are seconds:
ffencode.BuildFFmpegCommandandvmaftune.encode.build_ffmpeg_commandjoin repeated-x264-params/-x265-params/-svtav1-params/-vvenc-paramsinto one (MergeCodecParams/merge_codec_params; FFmpeg keeps only the last),-ssof the per-shot probe and signalstats takes seconds (ShotStartArg/_shot_start_arg), and no HDR argv carries-master_displayor-max_cllforhevc_nvenc(FFmpeg has no such option). - libvmaf is a compat library on libvmafx (ADR-1852 decision D3, RC4 WP6):
libvmafx.so.1(engine + VMAFx API) exportsvmafx_*only (core/src/vmafx.map,hide_unlisted = true), plus the libvmaf functions a built backend still keeps (core/src/vmafx_legacy_<backend>.map, declared exceptions ofcore/api/vmafx.toml).libvmaf.so.3(core/src/compat/libvmaf/) defines every other libvmaf function on exportedvmafx_*symbols and links nothing else. Every engine translation unit compiles with the generatedcore/src/vmafx/engine_names_gen.hforced (vmaf_engine_name_args), which renames the engine's own libvmaf bodies tovmaf_engine_<stem>: an upstream sync ports a change to such a body into the engine source unchanged and never adds a definition of a libvmaf name tolibvmaf.so.3by hand (the definition generates it, or amanualfile ofcore/src/compat/libvmaf/declared there). Generated files are never hand-merged: take either side and runpython3 scripts/codegen/vmafx-api.py --write; the Meson testtest_vmafx_api_generated_currentfails on any difference.check_exported_symbols/check_exported_symbols_libvmafcompare both libraries withcore/src/vmafx_symbols.txtandcore/src/libvmaf_symbols.txt, andtest_compat_conformancecompares every compat function with its engine body (traces,%a). See core/src/AGENTS.md. - VMAFx imported frames never take a host copy (ADR-1929):
vmafx_frame_import()binds the producer's planes or converts NV12 / P010 / P016 on the device (de-interleave, P010 shift 6) and nothing else; every place that copies imported pixels through host memory callsvmafx_count_host_copy()(core/src/vmafx/frame_import_hooks.h), and the import tests assert the count stays 0. A frame's release fence is signalled where its last picture reference is dropped (vmafx_frame_release(),pool_frame_release()), never when a submit returns. The CPU conversion uses the row readers ofcore/src/metal/iosurface_layout.h: a sync that changes them keepstest_vmafx_import_bitexactandtest_metal_iosurface_layoutpassing together. - Provenance record in the library (ADR-2073): the engine's report writers (
core/src/output.cpp) embed the record and the backend receipt; the CLI no longer edits its report afterwards, so a rebase must not bring backamend_cli_backend_receipt(). The feature collector records the producer of every feature vector from the producer the engine installs around each extractor call, prediction and import (vmaf_feature_producer_swap(),vmaf_feature_collector_append_from()); every engine loader of a model ends invmaf_model_stamp_loaded(). The record's canonical form (RFC 8785 subset) and whatdigestcovers are fixed by the ADR.core/src/vmafx/exactness_gen.cis generated fromscripts/ci/exact_twins.dbymake docs-fragments-write. See core/src/AGENTS.d/vmafx-provenance.md. - VMAFx CUDA frames are read on one stream per device (ADR-2023): every frame of a CUDA device of the VMAFx API is a CUDA picture on the device's library stream; its acquire wait, conversions and ready event are there, and its release is enqueued there where its last reference is dropped (
core/src/cuda/import_fence.c).core/src/libvmaf.cskips the ADR-1199 barrier only for a pair oforderedpictures and never downloads avmafxpicture.integer_vif_cudareads each picture with its own pitch (VifBufferCuda.dis_stride). Preservetest_vmafx_import_cuda*together. - VMAFx window values are the synchronous pooled values (ADR-2074):
vmafx_score_pooled(),vmafx_feature_score_pooled(),vmafx_score_pooled_model_set()and the windows all callvmafx_pool_engine()(core/src/vmafx/score.c); keep one pooling path. Each context's completion thread finds completion; it is woken by the engine's frame listener (end ofthreaded_extract_batch_func()) and by the hooks insubmit.c,register.candcontext.c, and calls the engine only throughvmafx_engine_enter()/vmafx_engine_leave(context, ...), which take the context's engine lock. Its probes never fence (vmaf_engine_try_score_at_index(),vmaf_engine_feature_written(),vmaf_predict_inputs_written()).vmaf_engine_max_in_flight()mirrorsbatch_job_take_pictures(), the thread pool's enqueue capacity and the device double buffer; a change to one recomputes it.test_vmafx_window,test_vmafx_window_live,test_vmafx_window_cliandtest_vmafx_lifetimeguard it. - VMAFx HIP frames are read on one stream per device (ADR-2092): every frame a HIP device of the VMAFx API imports is a
VMAF_PICTURE_BUFFER_TYPE_HIP_DEVICEpicture on the device's library stream. The twins copy it there (vmaf_hip_picture_upload(), the shared frame), and the reading twin's stream and the null stream wait for the copies;bind_hip_frame()submits each import withhipStreamQuery(); a GL import checks for a GLX context of the device's GPU before any HIP-GL call; a dma-buf is imported with its own size. Keep these when rebasingcore/src/hip/picture_hip.c,shared_frame.cor the three twins that stage on the host.core/test/test_vmafx_import_hip_contract.py,test_hip_shared_frameandtest_vmafx_import_hip*(on a device) guard them together. -
Coverage Gate ratchet + per-PR delta gate (ADR-0922): ADR-0922. Absolute floors live in
scripts/ci/coverage-check.sh(OVERALL_MIN=70,CRITICAL_MIN=90,PER_FILE_MIN[...]); per-PR drop tolerance lives inscripts/ci/coverage-delta-check.sh(default 0.5pp on overall and per-touched-file). Lowering any floor or loosening the delta tolerance requires a new ADR superseding ADR-0922. The Coverage Gate job in.github/workflows/tests-and-quality-gates.ymlinvokes both scripts; the delta gate needsactions/checkoutwithfetch-depth: 0because it runsgit merge-base. See scripts/ci/AGENTS.md §Coverage Gate ratchet for the full coupling. -
CI action pins — Windows MSVC dev env (ADR-0635):
.github/workflows/libvmaf-build-matrix.ymlusesTheMrMilchmann/setup-msvc-dev@79dac248…(v4.0.0, Node.js 24) for the Windows GPU build legs. If upstream ADR-0121 is re-implemented or the Windows legs are rebased, do not reintroduceilammy/msvc-dev-cmd(Node.js 20, deprecated 2026-06-02). TheTheMrMilchmannaction is a drop-in replacement with identicalvcvarsall.batsemantics. Also: both Windows jobs are pinned towindows-2025; do not revert towindows-latest(redirect towindows-2025-vs2026takes effect 2026-06-15). -
dev-MCP Docker container (ADR-0451):
dev/Containerfileinstalls CUDA through the shared installer's exact--mode=fullcontract (ADR-1306).build-config.envowns the apt series, release lock, and exact toolkit/nvcc/cudart package versions; do not restore a floatingcuda-toolkit-13-4command in the Containerfile. It also pins the unversionedintel-basekitmeta-package (Intel does not publish aintel-basekit-2025.3apt package), and the digest-pinnedrocm/dev-ubuntu-26.04:10.1.0-fullimage in therocm-srcstage (ADR-1225 / ADR-1231). If SDK versions are bumped (routine security maintenance), update their shared pins inbuild-config.envand regenerate the mirrors before merging; a ROCm bump additionally means re-validating therocm-srcprune list against its hipcc smoke check, settingROCM_VERSION(Renovate moves only the image pins), moving the/opt/rocm/core-<major>.<minor>paths of therocm-srcstage anddocker/Dockerfile.node, the LLVM sonames oftools/rc1-tester/image/hip-runtime.jsonand the ROCm source pins oftools/rc1-tester/image/licensing.json, and re-running the HIP device suite (ROCm 10.1.0 moved the compiler from LLVM 23 to LLVM 24).dev/scripts/smoke-probe-loop.shassumes the golden pair lives at${VMAF_TESTDATA_PATH}/ref_576x324_48f.yuv/dis_576x324_48f.yuv— do not rename these files. The probe JSON schema fields (ts,host_id,backend_results,mcp_results) are an internal format; updatedocs/development/dev-mcp.mdif the schema changes. This directory does not affect the libvmaf C build or any CI gate. -
Top-level
noxfile.pyis a local-dev affordance, not a CI gate (ADR-0914): The repo-rootnoxfile.pyexposes one session per Python suite (ai,compat_decorator,mcp,vmaf_tune,dev_llm,roi_score,ensemble_kit,rc1_tester,tooling,python_harness) plusall/lintmeta-sessions. CI does not call nox: each suite runs in its own job (Tiny AI,MCP Smoke,RC1 Tester Report,Tooling Tests, onePython Package Testsleg per remaining package), which installs the suite's hash lock and runspytest -rs..github/test-suites.jsonmaps every tracked test file to one suite and every suite to its required checks;scripts/ci/suite_registry.py checkfails on an unwired test file (ADR-1528). When adding a new Python package, updatenoxfile.py, the CI job and the registry together. See test suites anddocs/development/python-test-orchestrator.md. Thepython_harnesssession intentionally delegates totox -c pythonrather than duplicating the Cython + Netflix golden-data setup that lives inpython/tox.ini; do not collapse them. -
Security support and badge evidence —
SECURITY.mddescribes actual VMAFx release support, not inherited Netflix/libvmaf version strings. Keep the passing worksheet tied to a reviewed source revision and the live project record. Configuration, future releases and agent-authored prose cannot establish historical response times, a human developer's knowledge or a completed external badge. The project website is GitHub Pages; keep the short purpose and participation links indocs/index.md, and verify deployed pages before citing new text.
Upstream sync and provenance¶
-
Recorded upstream head (ADR-1474):
docs/development/known-upstream-bugs.mdcarries exactly one heading## Upstream head the fork is at parity with: `<commit id>` (<date>).scripts/ci/upstream_parity_pin.pyreads it and the required checkLicence Provenancecompares every file's licence header against that Netflix/vmaf commit. An upstream port or sync moves the id in the same pull request and keeps the heading's wording; a second heading of that form, or a reworded one, fails the check. A port that brings a file whose path or name now exists upstream changes that file's verdict: runscripts/dev/relicense_fork_files.py --check --upstream-ref <new id>before pushing. See the guide. -
Deliberate deviations from Netflix's source, by ADR: code inherited from Netflix/vmaf evaluates as Netflix's source does unless an ADR says otherwise. Eight fixes that predate that rule have their ADR since 2026-10-02, each with upstream's lines at Netflix
9e48141b, the measured size and the upstream pull request that would end it: ADR-1479 (ciede4:2:2 chroma flags), ADR-1480 (speed_temporalbuffers atspeed_prescaleabove 1), ADR-1481 (a worker's error fails the run), ADR-1482 (integeradmon frames of 17 to 32 pixels), ADR-1483 (odd-sized chroma planes round up), ADR-1484 (float_ms_ssimmagnitude beforepow()), ADR-1485 (apsnrof a plane without error) and ADR-1486 (float_motionscale-1 stride). A sync keeps the fork's side of these lines until the named upstream pull request is merged; the table is in rebase-notes under "Eight deliberate deviations". -
Upstream port — feature/motion options from b949cebf (T-NEW-1): PR #197 (
b949cebf, MERGED 2026-04-29) ported Netflix's feature/motion several-options commit; PR #213 portedd3647c73feature/speedextractors (speed_chroma+speed_temporal;speed.cis in the tree).
Backends, extractors and the parity gate¶
-
GPU long-tail terminus reached — every registered feature extractor has at least one GPU twin (lpips remains ORT-delegated per ADR-0022). Cross-backend tolerances live in
scripts/ci/cross_backend_parity_gate.py. Governing ADRs: ADR-0182 (batch 1: psnr / ciede / moment), ADR-0188 (batch 2: ssim / ms_ssim / psnr_hvs), ADR-0192 (batch 3: motion_v2 / float-twins / ssimulacra2 / cambi;float_ansnrremoved in commit 70ed8b3ce3 / PR #38). See core/src/feature/AGENTS.md. -
Vulkan backend removed (ADR-0726) — the Vulkan backend, its
libvmaf_vulkan.hsurface, thecore/src/vulkan/tree, the Volk-symbol-hiding machinery, and all*_vulkanGLSL kernels (ssim / ms_ssim / motion_v2 / cambi / psnr chroma) no longer exist in the tree. No rebase invariant survives. Treat any lingering Vulkan reference as stale. -
MCP embedded server (ADR-0128, ADR-0209; runtime v3): ADR-0209. Public header
libvmaf_mcp.h; the runtime is live incore/src/mcp/(mcp.c,dispatcher.c,compute_vmaf.c, thetransport_{stdio,uds,sse}.cbodies, vendored cJSON) behindenable_mcpplus three transport sub-flags. The stdio transport is newline-delimited JSON-RPC;-ENOSYSmeans only "feature or transport not built"; the SPSC command ring is v4 work, soqueue_depth/max_drain_per_frameare validated and stored, nothing more. User page: embedded MCP. See core/src/mcp/AGENTS.md. -
HIP backend (T7-10, ADR-0212, PR #200) — public
libvmaf_hip.h, 19 registered feature extractors (HIP overview),enable_hipmeson option defaultfalse, device kernels behindenable_hipcc. -
SVE2 SIMD ports (T7-38, ADR-0213, PR #201) — SSIMULACRA 2 PTLR + IIR-blur SVE2 ports developed against
qemu-aarch64-static. Same bit-exact contract as the existing NEON ports. -
GPU-parity CI gate (T6-8, ADR-0214): ADR-0214. Single source of truth for cross-backend tolerances:
scripts/ci/cross_backend_parity_gate.py. Adding a new GPU twin requires (1)FEATURE_METRICSentry, (2)FEATURE_TOLERANCEentry if it relaxes places=4, (3) row indocs/development/cross-backend-gate.md. Declaring a twin bit-identical adds one filescripts/ci/exact_twins.d/<feature>.<backend>(ADR-1428) and edits no shared line; on a conflict in the generateddocs/development/cross-backend-exact-twins.mdtake master's side and runmake docs-fragments-write. See core/AGENTS.md. -
FastDVDnet temporal pre-filter (T6-7, ADR-0215, PR #203) — 5-frame window pre-filter feeding ssim/ms_ssim.
-
MobileSal saliency extractor (T6-2a, ADR-0218, PR #208) — first half of T6-2 (encoder-side ROI bundle). Saliency-weighted VMAF, sidecar emit for
tools/vmaf-roi. -
TransNet V2 shot-boundary extractor (T6-3a, PR #210) — ~1M params; feeds
tools/vmaf-perShotCRF predictor. -
Model registry + Sigstore (T6-9, ADR-0211, PR #199):
--tiny-model-verifyflag + registry schema + Sigstore bundle paths. Pairs with ADR-0010 (release signing). -
CPU extractors declare the features they write (ADR-1359): the twin lookup pairs a CPU extractor with a device twin through
provided_features.core/src/feature/float_moment.cis an upstream-mirror file whose list the fork changed from upstream's pseudo-name"float_moment"to the four emittedfloat_moment_*names; an upstream sync must keep the fork's list, or--backend <gpu> --feature float_momentfalls back to the CPU again.vmaf_feature_extractor_twin_audit()andtest_every_device_twin_is_reachable(core/test/test_feature_extractor.c) fail when any registered device twin is unreachable. See core/src/feature/AGENTS.md. -
CUDA twins declared exact as a group (ADR-1457):
scripts/ci/exact_twins.d/{motion,motion_debug,motion_v2,psnr,float_ssim,float_ssim_lcs,float_ms_ssim,float_ms_ssim_lcs,cambi}.cudamake the parity gate compare those cells with tolerance 0, andcore/test/test_cuda_exact_twins.choldsmotion_cuda,motion_v2_cuda,psnr_cuda,float_ssim_cuda,float_ms_ssim_cudaandcambi_cudato==on every output. A rebase that changes one of these twins or its CPU extractor keeps them bit-identical; a twin that drifts is fixed, never given a tolerance or taken off the list. -
SYCL twins declared exact as a group (ADR-1451):
scripts/ci/exact_twins.d/{adm,motion,motion_debug,motion_v2,psnr,float_ssim,float_ssim_lcs,cambi}.syclmake the parity gate compare those cells with tolerance 0, andcore/test/test_sycl_exact_twins.choldsadm_sycl,motion_sycl,motion_v2_sycl,psnr_sycl,float_ssim_syclandcambi_syclto==on every output. A rebase that changes one of these twins or its CPU extractor keeps them bit-identical; a twin that drifts is fixed, never given a tolerance or taken off the list.
Floating-point policy and device contracts¶
-
No C or C++ translation unit is built with FP contraction (ADR-1461):
core/src/meson.builddeclaresvmaf_strict_fp_argsas a project argument for C and C++ directly after theVMAF strict FP compiler-argument policyblock, above the first build target. Keep both there on a rebase (Meson refusesadd_project_arguments()after a target), and never give a targetvmaf_fp_model_argsalone or any flag that turns contraction back on.core/test/test_strict_fp_compiler_args.pyreads the compile database of the build it runs in;make test-netflix-golden-arm64runs the golden gate on an aarch64 cross build, where a clang build and a GCC build used to differ. See core/AGENTS.md. -
icx and icpx builds link glibc's libm, not Intel's libimf (ADR-1495):
core/src/meson.builddeclares theVMAF host math library link policyblock directly after the strict FP policy and passes its lists withadd_project_link_arguments()for C and C++, above the first build target: anintel-llvmcompiler gets-no-intel-lib=libimf, every other compiler nothing. The Intel driver otherwise linkslibimfinto every link (it turns a given-lminto-limf -lm), and an icx-builtvmafexported libimf's copies of the math functionslibvmaf.soimports, so the CPU scores of an icx build differed from a GCC build's. Keep the block and both lines on a rebase, and never link Intel's math library back by name or substitute-shared-intel.core/test/test_icx_system_libm.pyreads the build's ownlibvmaf.soandvmaf(skips on non-icx builds) andcore/test/test_strict_fp_compiler_args.pyexecutes the block per compiler pair. See core/AGENTS.md. -
GPU device code is stored compressed (ADR-1590):
core/src/meson.builddefines one compression list per backend between theBEGIN/END VMAF {CUDA,HIP,SYCL} device code compression policymarkers (cuda_compress_args,hip_compress_args,sycl_compress_args), gated by thecompress_device_codeoption (defaulttrue). Every nvcc fatbin, everyhipcc --gencocommand (the test probes incore/test/meson.buildtoo), the SYCL AOT compile,sycl_link_argsand the MSVC device link take the list; the flags are spelled nowhere else. A rebase that adds a device compile site adds the list, and must not drop theerror()that refuses a compiler without compression. The build checks its own output withcore/src/check_device_compression.py;core/test/test_device_code_compression.pyguards the policy without a device. See core/AGENTS.d/device-code-compression.md. -
SYCL strict FP line on every feature TU (ADR-1367):
core/src/meson.builddefinessycl_strict_fp_argsonce, between theBEGIN/END VMAF SYCL strict FP policymarkers: icpx gets-fp-model=precise -ffp-contract=off -foffload-fp32-prec-div -foffload-fp32-prec-sqrtin that order (precise implies contraction on, so contraction-off must follow it), AdaptiveCpp-ffp-contract=off. Every feature TU takes it throughsycl_feature_tail_args; no TU gets a private FP list.sycl_link_argsalso carriessycl_fp32_prec_argsto every link the icpx driver runs, because the SPIR-V JIT image is generated there; dropping it leaves-Dsycl_icpx_aot_targets=builds with approximate/and sqrt. The MSVC build's explicit device link (ADR-1364) generates every image and takessycl_strict_fp_argswhole.core/test/test_strict_fp_compiler_args.pyexecutes the policy andtest_sycl_fp_arith_contractchecks the device arithmetic. -
CUDA device FP policy (ADR-1403): every CUDA fatbin takes
cuda_device_strict_fp_args(--fmad=falseunder nvcc,-ffp-contract=offunder clang CUDA), defined once between theVMAF CUDA device strict FP policymarkers incore/src/meson.build;cuda_cu_extra_flagscarries no floating-point flag. A kernel whose reference fuses writes__fmaf_rn().float_ms_ssim_cudareproducesms_ssim_decimate.c,iqa_convolve()andssim_accumulate_default_scalar()operation for operation and is bit-identical to the CPU.core/test/test_strict_fp_compiler_args.py,core/test/test_cuda_kernel_source_contract.pyandcore/test/test_cuda_float_ms_ssim_parity.cguard it. See core/src/cuda/AGENTS.md and core/src/feature/cuda/AGENTS.md. -
SYCL kernels use no scratch memory (ADR-1395): on an Arc A-series GPU under the Linux xe driver, kernels with a private array in memory or spilled registers return wrong values.
test_sycl_kernel_scratchfails on a scratch kernel missing fromcore/src/sycl/scratch_ratchet.txt, whose extractors must matchkScratchExtractorsincore/src/sycl/scratch_check.cpp; the list only shrinks.integer_vif_syclruns at SIMD-16 only (ADR-1830): a sync must not bring back its SIMD-32 kernels orVMAF_SYCL_VIF_SUBGROUP_SIZE, which spilled on Xe-LP. See core/src/sycl/AGENTS.md and core/src/feature/sycl/AGENTS.md. -
libvmaf_syclimports each input with its own VA display and never skips a frame (ADR-1761): FFmpeg patch0005reads the VA display of both inputs' QSV sessions and imports each input's surfaces with its own; a failed import is retried in a bounded loop and then stops the filter naming the frame, with no pooled score after it. Patch0013'slibvmaf_metalprints no pooled score after a stop either. A refresh or an upstream rebase must not bring back the single display, the pass-through on a failed import, or a score after a stop.core/test/test_sycl_filter_import_contract.pyandcore/test/test_metal_iosurface_filter_contract.pyguard it without a device,ffmpeg-patches/test/check-sycl-import-retry.shon one. libvmafandlibvmaf_cudaprint no pooled score after an error (ADR-1768): FFmpeg patch0021is a deliberate divergence from upstream'svf_libvmaf.c.stop_on_frame()logs the frame and the error once and setsstopped.do_vmaf()anddo_vmaf_cuda()call it on every failed copy and read, andframe_cntadvances only after a successful read. The shareduninit()prints no score after a stop, a failed flush or a failed model. A series refresh keeps the patch; when upstream rewrites these functions, the rule moves into the new code.core/test/test_ffmpeg_libvmaf_stop_contract.pyguards it without a device,ffmpeg-patches/test/check-libvmaf-no-score-after-error.shon one.- SYCL zero-copy admission (ADR-1688):
vmaf_read_pictures_sycl()incore/src/libvmaf.crefuses, before counting a frame, every registered extractor whosereads_shared_luma_only()hook is absent or false for its options, with an error naming it and-ENOTSUP;vmaf_flush_sycl()skips uninitialized extractors. The hooks sit on eight SYCL twins (core/src/feature/sycl/AGENTS.d/zero-copy-admission.md); a twin that starts reading a host picture narrows its hook in the same PR.test_sycl_zero_copy_admissionandtest_sycl_zero_copy_model_gateguard it. See core/src/sycl/AGENTS.md. -
SYCL fp64-less device contract (T7-17, ADR-0220): ADR-0220. SYCL feature kernels are unconditionally fp64-free; a single fp64 instruction in any lambda blocks the whole TU on Arc A-series. See core/src/sycl/AGENTS.md.
-
SYCL kernels require sub-group size 16 or 32 (ADR-1468): the default build compiles every kernel ahead of time for the 19 targets of
sycl_icpx_aot_targets, and the Xe2 targets do not compile a kernel that requires 8.core/src/feature/sycl/sycl_compat.hrejects another size at compile time (VmafSyclSubGroupSize); a rebase must not bring a raw[[sycl::reqd_sub_group_size(N)]]orsub_group_size<N>into a kernel, nor a size 8.core/test/test_sycl_sub_group_size_contract.py(device-free) andcore/test/test_sycl_aot_default_targets.py(suitesycl-aot, compiles every SYCL translation unit for the full default list) guard it;core/test/sycl_aot_targets.pyholds the measured sizes per target family and needs an entry for a target added to the list.
psnr_hvs¶
-
psnr_hvs_cudareturns the CPU's scores bit for bit (ADR-1397):psnr_hvs_score.custores the 64 termscalc_psnrhvs()sums per block, in the CPU's arithmetic (double masking table, the threshold's float product and double root, integer coefficient difference, fatbin built with--fmad=false), andcore/src/feature/psnr_hvs_score.cadds them into one runningfloatin the CPU's order. A change tocalc_psnrhvs()orextract()inthird_party/xiph/psnr_hvs.cchanges the kernel and that file in the same PR.core/test/test_psnr_hvs_twin_exact_sum_contract.pyandtest_psnr_hvs_scoreguard it without a device,test_cuda_psnr_hvs_parityon one; the parity gate compares the twin with tolerance 0 (EXACT_TWINS). See core/src/feature/cuda/AGENTS.md. -
The
psnr_hvsmasking threshold is upstream's statement (ADR-1488):calc_psnrhvs()writess_mask = sqrt(s_mask * s_gvar) / 32.f(and the same ford_mask), as Netflixlibvmaf/src/feature/third_party/xiph/psnr_hvs.c:316-317does: a float product, its root in double, the result stored as float. A sync takes upstream's side of these two lines, and no(double)goes in front of the product (PR #552 added one).x86/psnr_hvs_avx2.candarm64/psnr_hvs_neon.cwrite the same statement incompute_masks(); the CUDA and HIP kernels form the float product and take the double root, the SYCL kernel takessqrt_rn()of the float product (the same value without fp64). A change to the statement changes all six in the same PR.test_psnr_hvs_dispatch_invariance(recorded blocks scored as Netflix master scores them),test_psnr_hvs_simdandtest_psnr_hvs_twin_exact_sum_contract.pyguard it. -
psnr_hvs_syclandpsnr_hvs_hipreturn the CPU's scores bit for bit (ADR-1401): both store the 64 termscalc_psnrhvs()sums per block and callcore/src/feature/psnr_hvs_score.c, as the CUDA twin does. The HIP kernel (psnr_hvs_score.hip) takes the masking table indouble, the threshold as the float product's double root (ADR-1488), and is built with-ffp-contract=off -fhip-fp32-correctly-rounded-divide-sqrt(hip_cu_extra_flags). The SYCL kernel has no fp64: its masking table is a compile-time constant and its threshold comes fromsqrt_rn()incore/src/feature/sycl/sycl_exact_fp.happlied to the float product (the correctly rounded fp32 root, which is the double root rounded to float); it must stay free of scratch memory. A change tocalc_psnrhvs()orextract()inthird_party/xiph/psnr_hvs.cchanges both kernels in the same PR.test_psnr_hvs_twin_exact_sum_contract.pyguards all three twins without a device;test_sycl_psnr_hvs_parity,test_hip_psnr_hvs_parityandtest_sycl_fp_arith_contracton one. See core/src/feature/sycl/AGENTS.md and core/src/feature/hip/AGENTS.md.
psnr and float_moment¶
-
psnr chroma GPU twins (T3-15(b), PR #204) —
psnr_cb/psnr_crdevice kernels alongside the existingpsnr_yfrom ADR-0182. (The original Vulkan implementation was removed with the backend in ADR-0726.) -
float_psnr_cudaadds integers (ADR-1455):core/src/feature/cuda/float_psnr/float_psnr_score.cuforms the CPU's term (diff * diffinfloat, asfloat_psnr.cdoes) with__fmul_rn()as an integer in units of 1 / \(\mathrm{scaler}^{2}\) and reducesuint64values per warp and per block;float_psnr_cuda.c::float_psnr_noise()adds the blocks inuint64and divides the exact total by \(\mathrm{scaler}^{2}\) and the pixel count. An fp32 block sum is exact only up to 24 bits. A change to howfloat_psnr.cforms or adds its terms changes the kernel in the same PR.core/test/test_cuda_float_psnr_exact_contract.pyguards it without a device,test_cuda_float_psnr_parity(==) on one. Each block / work-group lies in ONE row (256 x 1), and the host adds each row's exact sum into a double in row order withcore/src/feature/float_psnr_rows.h(ADR-1499), asfloat_psnr.cadds its rows, so the twin rounds where the CPU rounds past \(2^{53}\) units; a sync must not bring back 16x16 blocks or a frame total rounded once. The HIP twin (ADR-1440) follows the same layout and helper. -
float_psnr_sycladds integers (ADR-1450):core/src/feature/sycl/float_psnr_sycl.cppforms the CPU's term (diff * diffinfloat, asfloat_psnr.cdoes) as an integer in units of 1 / \(\mathrm{scaler}^{2}\) and reducesuint64values per sub-group, per work-group and on the host; an fp32 group sum is exact only up to 24 bits. The host divides the exact total by \(\mathrm{scaler}^{2}\) and the pixel count. A change to howfloat_psnr.cforms or adds its terms changes the kernel in the same PR.core/test/test_sycl_float_psnr_exact_contract.pyguards it without a device,test_sycl_float_psnr_parity(==) on one. Each block / work-group lies in ONE row (256 x 1), and the host adds each row's exact sum into a double in row order withcore/src/feature/float_psnr_rows.h(ADR-1499), asfloat_psnr.cadds its rows, so the twin rounds where the CPU rounds past \(2^{53}\) units; a sync must not bring back 16x16 blocks or a frame total rounded once. The HIP twin (ADR-1440) follows the same layout and helper. -
float_moment_hipadds the CPU's float squares (ADR-1447): the 16-bit kernel ofcore/src/feature/hip/float_moment/moment_score.hipaddsmoment_float_square(), one fp32 product of the sample with itself converted to an integer, wheremoment.c::compute_2nd_moment()forms the square infloat; an exact integer square is another number at 16 bits. The host recovers the moment with the CPU's two divisions. A change to howmoment.cforms or adds its terms changes the kernel in the same PR.core/test/test_hip_float_moment_exact_contract.pyguards it without a device,test_hip_float_moment_parityon one (==, past \(2^{53}\) units too, ADR-1497 below). -
The NEON and SVE2
float_momentkernels add in the scalar's order (ADR-1500):core/src/feature/arm64/moment_neon.candmoment_sve2.cstore each vector of samples (squared infloatfor the second moment) and add the lanes into onedoubleone after the other, asmoment.candx86/moment_avx2.cdo; the SVE2 kernel adds the firstsvcntp_b32active lanes of asvwhilelt_b32predicate and does not depend on the vector length. A sync must not bring back lane accumulators, per-row vector sums or a vector reduction (vaddvq_f64,svaddv_f64): past \(2^{53}\) units the sum rounds on every add.core/test/test_moment_simd.c(==) guards it; run it underqemu-aarch64withsve=off,sve128,sve256,sve512andsve2048after touching any of the four kernels. -
float_moment_cudaadds the CPU's float squares (ADR-1453): the 16bpc kernel ofcore/src/feature/cuda/integer_moment/moment_score.cuaddsmoment_float_square(), one__fmul_rn()product of the sample with itself converted to an integer, wheremoment.c::compute_2nd_moment()forms the square infloat; an exact integer square is another number at 16 bits. The host recovers the moment with the CPU's two divisions. A change to howmoment.cforms or adds its terms changes the kernel in the same PR.core/test/test_cuda_float_moment_exact_contract.pyguards it without a device,test_cuda_float_moment_parityon one (==, past \(2^{53}\) units too, ADR-1497 below). -
float_moment_sycladds the CPU's float squares (ADR-1449): the kernel ofcore/src/feature/sycl/integer_moment_sycl.cppaddsmoment_float_square(), one fp32 product of the sample with itself converted to an integer, wheremoment.c::compute_2nd_moment()forms the square infloat; an exact integer square is another number at 16 bits. The host recovers the moment with the CPU's two divisions. A change to howmoment.cforms or adds its terms changes the kernel in the same PR.core/test/test_sycl_float_moment_exact_contract.pyguards it without a device,test_sycl_float_moment_parityon one (==, past \(2^{53}\) units too, ADR-1497 below). -
The
float_momenttwins form the CPU's rounded second-moment sum past \(2^{53}\) units (ADR-1497): on a frame whose sum of float squares can pass \(2^{53}\) units (vmaf_moment_sum_may_round()), the CUDA, SYCL and HIP hosts run four more kernels after the frame kernel (row totals, row plans, row units, ordered totals) that replace accumulators 2 and 3 with the CPU's sequentially rounded sums. The arithmetic and every lane's steps arecore/src/feature/float_moment_sum.h(integers only); the CUDA and HIP kernels arecore/src/feature/float_moment_sum_gpu.h, compiled intomoment_score.cu/moment_score.hip; the SYCL kernels are ininteger_moment_sycl.cppand pick planes by value. A sync must not drop the four kernels, add a row from its increments withoutvmaf_moment_sum_add_run()'s check, reorder the tree, or bring back the exact sum rounded once. A change tocompute_2nd_moment()'s order or term changes the header andtest_float_moment_sumin the same PR.test_float_moment_sum(host, the kernels' steps againstpicture_copy()+compute_2nd_moment()up to 7680x4320) andtest_float_moment_sum_contract.pyguard it without a device,test_{cuda,sycl,hip}_float_moment_parityon one.
SpEED and CAMBI¶
-
SpEED evaluates Netflix's fp64 expressions (ADR-1477): three places of
core/src/feature/speed.care fp64 arithmetic rounded tofloatonce, as Netflix master has them:1.0 / sqrt(1 + t * t)increate_givens(), thelog2()sum ofupdate_entropy()and thelog2()weights ofget_speed_score(). A sync takes upstream's side on these lines and never restoressqrtf/log2f/0.75f(the fork's port #213 had them: up to 6.6e-4 from Netflix). The GPU twins runspeed.con the device up to the per-block variances, read one block back per frame (SpeedGpuTailLayout,core/src/feature/speed_gpu_common.h) and form the entropies and the score on the host withspeed_internal_gpu_tail_scores()(core/src/feature/speed_internal.c), which holds those statements; the rotation's statement isspeed_givens_unit()(core/src/feature/speed_givens.h) in the kernels. A change toupdate_entropy(),get_speed_score()orspeed_extract_score()changes the tail in the same PR, and a change tocreate_givens()changessi_create_givens()andspeed_givens.h. No kernel may evaluate a logarithm or form a score, and the six gate cells stay exact (scripts/ci/exact_twins.d/speed_{chroma,temporal}.*).core/test/test_speed_upstream_form.cguardsspeed.c's three functions and the tail against upstream's statements evaluated with the host's ownlog2()on every C library, the rotation on every input, and, on glibc only, the CPU against Netflix's values, without a device (its_foreign_libmvariant runs the other libcs' path); the three source contract tests andtest_{cuda,hip,sycl}_speed_*_parity(==) guard the twins. -
SYCL SpEED device-resident pipeline (ADR-1358): every SpEED kernel lives in
core/src/feature/sycl/speed_sycl_pipeline.cppand reproducesspeed.coperation for operation, up to the variances (the entropies and the score are the host's since ADR-1477); like every SYCL feature TU the SpEED TUs build with contraction off (sycl_strict_fp_args, ADR-1367), divide and take square roots throughdiv_rn()/sqrt_rn(), and never wait on the queue mid-frame.core/test/test_sycl_kernel_source_contract.pyguards the layout;scripts/dev/speed_gpu_parity.py --backend syclre-checks bit parity. See core/src/feature/sycl/AGENTS.md. -
SpEED twin parity fixture (ADR-1430, ADR-1452):
core/test/test_cuda_speed_chroma_parity.candcore/test/test_hip_speed_chroma_parity.ckeep the 960x960 textured fixture ofcore/test/speed_chroma_twin_parity.h: a smaller or ramp fixture has a singular covariance and never reaches the scoring path. The comparison is==since ADR-1477; theLIBM_TWINSbounds those two ADRs introduced (5e-6, and4e-5forspeed_temporal, ADR-1460) are gone. -
CUDA CAMBI and SpEED device-resident (ADR-1379, ADR-1380):
cambi_cuda,speed_chroma_cudaandspeed_temporal_cudaread back one result block and wait once per frame, incollect(); a sync must not bring back the host c-values, host pooling, host SpEED linear algebra or a mid-framecuStreamSynchronize. SpEED's block is the tail of ADR-1477 (status words, eigenvalues, variances), from which the host forms the entropies and the score after the wait. The host constants come fromcambi.c(vmaf_cambi_*helpers incambi_internal.h) andspeed_internal_gpu_configure(), shared with the SYCL twins;speed/speed_score.cukeeps its__f*_rnintrinsics and--fmad=false(every CUDA fatbin's, ADR-1403).core/test/test_cuda_device_resident_contract.pyguards the design. See core/src/feature/cuda/AGENTS.md. -
HIP CAMBI and SpEED device-resident pipelines (ADR-1378, ADR-1384): no host stage of
cambi.c/speed.cbefore the frame's wait and no mid-frame wait; one staged upload, one readback, the wait incollect(). SpEED's readback is the tail block of ADR-1477 andcollect()forms the entropies and the score from it on the host. Per-work-item math lives ininteger_cambi/cambi_hip_device.handspeed/speed_hip_device.h, which the host replay tests compile; the SpEED kernel TU keeps-ffp-contract=off -fhip-fp32-correctly-rounded-divide-sqrt. The init-time helpers arecambi.c's (cambi_internal.h) andspeed_internal_gpu_configure(), shared with SYCL.core/test/test_hip_device_resident_contract.pyguards the layout. See core/src/feature/hip/AGENTS.md.
SSIMULACRA 2¶
-
ssimulacra2_hipreturns the CPU's score bit for bit (ADR-1445):ssimulacra2_device.hipevaluates the six per-pixel terms with the CPU's fp64 expressions (ss2h_terms(), no fp32 pairs) and forms their sums withcore/src/feature/ordered_sum.hin the four kernels of the CUDA twin (ADR-1433): 1024-pixel chunks in raster order, lanes composed in lane order, a checked walk, term-by-term fallback in pixel order. A change tossim_map()/edge_diff_map()inssimulacra2.cchangesss2h_terms()in the same PR.core/test/test_hip_ssimulacra2_exact_contract.pyandcore/test/test_ordered_sum.cguard it without a device,test_hip_ssimulacra2_parity(==) on one. See core/src/feature/hip/AGENTS.md. -
SYCL ssimulacra2 / float_ms_ssim single wait (ADR-1363):
ssimulacra2_sycl.cppruns the whole frame on the device and reads one block of per-scale sums incollect(). The exact-fp helpers live incore/src/feature/sycl/sycl_exact_fp.hand need contraction off, which every SYCL feature TU has (ADR-1367).integer_ms_ssim_sycl.cppenqueues every scale insubmit()into its own partials span and waits once.core/test/test_sycl_kernel_source_contract.pyguards all of it. -
ssimulacra2_syclreturns the CPU's score bit for bit (ADR-1446): a SYCL kernel has no fp64 type, so the six per-sample terms are the CPU's doubles computed in 64-bit integers (core/src/feature/sycl/sycl_ssimulacra2_math.h), and their sums come fromcore/src/feature/ordered_sum.hthrough its_bitsforms (core/src/feature/sycl/sycl_ordered_sum.h): 512-pixel chunks in raster order, lanes composed in lane order, a checked walk, term-by-term fallback in pixel order. The fp32 pair sums are advice for the walk's plan and never a result.ordered_sum.his shared with the CUDA and HIP twins and keeps everydoubleinside#ifndef VMAF_ORDSUM_NO_FP64. A change tossim_map()/edge_diff_map()inssimulacra2.cchanges the math header andreference_terms()ofcore/test/test_sycl_ssimulacra2_math.cin the same PR.core/test/test_sycl_ssimulacra2_exact_contract.pyguards it without a device;test_sycl_ssimulacra2_math,test_sycl_ordered_sum,test_sycl_ssimulacra2_parity(==) andscripts/dev/speed_gpu_parity.py --backend sycl --feature ssimulacra2on one. See core/src/feature/sycl/AGENTS.md. -
CUDA ssimulacra2 single readback (ADR-1391):
ssimulacra2_cuda.cenqueues the whole frame insubmit()on the picture stream and reads one block of per-scale sums incollect(); no host compute or host wait mid-frame. The device TUsssimulacra2_deviceandssimulacra2_blurbuild with--fmad=false(as every CUDA fatbin does, ADR-1403), andssimulacra2_device.cucompiles the sharedfeature/ssimulacra2_math.h,ssimulacra2_score.handssimulacra2_eotf_lut.hinto device code through theVMAF_SS2_FUNC/VMAF_SS2_EOTF_LUT_STORAGEhooks, so an upstream change to those helpers must stay valid CUDA device code. The per-pixel SSIM / edge terms are fp64, and their sums are the sums of the CPU's loops (ADR-1433): chunks of 1024 pixels in raster order, integer increments of the running sum's binade composed in pixel order (feature/ordered_sum.h, compiled into device code through theVMAF_ORDSUM_*hooks and tested on the host bycore/test/test_ordered_sum.c), one walk per sum with a term-by-term fallback. The terms must stay non-negative or NaN, and a change tossim_map()/edge_diff_map()inssimulacra2.cchangesss2c_terms()in the same PR.core/test/test_cuda_ssimulacra2_parity.c(==),core/test/test_cuda_ssimulacra2_exact_contract.pyandscripts/dev/speed_gpu_parity.py --backend cuda --feature ssimulacra2re-check parity. See core/src/feature/cuda/AGENTS.md.
SSIM and MS-SSIM¶
-
float_ms_ssim_cudaandinteger_ms_ssim_hipscore every planeenable_chromaasks for (T-MS-SSIM-GPU-CHROMA-OPTION-DRIFT-2026-09-06): both keep geometry, pyramid and term buffers per plane and run the luma pipeline once per scored plane, asfloat_ms_ssim.cdoes; both declare the CPU's four options and providefloat_ms_ssim_cb/float_ms_ssim_cr. Plane count and plane size come fromcore/src/feature/metal/float_ms_ssim_option_semantics.h. A sync must not bring back a fixedn_planes = 1u, a luma-onlyprovided_featuresor a chroma path with its own arithmetic, and the HIP option stays (HISS-14).test_cuda_float_ms_ssim_parityandtest_hip_ms_ssim_parity(==) on a device,test_cuda_float_ms_ssim_exact_contract.pyandtest_hip_kernel_source_contract.pywithout one; gate cellfloat_ms_ssim_chroma. See core/src/feature/cuda/AGENTS.md and core/src/feature/hip/AGENTS.md. -
float_ms_ssim_cudaper-scale sums are the CPU's, in the CPU's order (ADR-1465):ms_ssim_vert_lcsincore/src/feature/cuda/integer_ms_ssim/ms_ssim_score.custores every window'sl,candsat its raster position andinteger_ms_ssim_cuda.c::ms_ssim_scale_sums()adds the three planes of a scale in index order, asiqa/ssim_tools.c::iqa_ssim()adds them. A sync must not bring back a device reduction of the terms: on the frame ofcore/test/float_ms_ssim_order_frame.hper-block sums return the neighbouringfloatforfloat_ms_ssim_c_scale1. That header is shared with the HIP and SYCL twin tests and its bytes are fixed. Preserve the kernel, the host loop,core/test/test_cuda_float_ms_ssim_order.candcore/test/test_cuda_float_ms_ssim_exact_contract.pytogether. -
float_ms_ssim_syclis the CPU's arithmetic (ADR-1414): the decimate spells each tapsycl::fma()asms_ssim_decimate.cfuses it; the window sums and thel/c/sterms come fromcore/src/feature/sycl/sycl_ssim_terms.h, shared withfloat_ssim_sycl(the window sums as exact fp32 pairs,landcas the CPU's doubles in 64-bit integers); every window'sl,candsof every scale is stored unreduced and the host adds them iniqa_ssim()'s raster order (ADR-1466; no reduction may return to the twin); the host rounds each per-scale mean to fp32 and combines asms_ssim.cdoes. A change toms_ssim_decimate.c,iqa/convolve.c,iqa/ssim_tools.c,iqa/ssim_accumulate_lane.horms_ssim.cchanges the header or the twin in the same PR.core/test/test_sycl_ms_ssim_parity.c(==on 18 outputs of 3 frames, and two order pairs: seeded noise andcore/test/float_ms_ssim_order_frame.h, a file shared byte for byte with the CUDA and HIP tests) andcore/test/test_sycl_kernel_source_contract.pyguard it. See core/src/feature/sycl/AGENTS.md. -
Metal IOSurface import reads NV12 / P010 itself (ADR-1679):
core/src/metal/picture_import.mmplans each plane from the surface's CoreVideo pixel format throughcore/src/metal/iosurface_layout.h(bi-planar chroma de-interleaved, P010 shifted, other layouts-ENOTSUP), and FFmpeg patch0013imports planes 0, 1 and 2 of both frames and fails on an import error.core/test/test_metal_iosurface_filter_contract.py,test_metal_iosurface_layoutandtest_metal_iosurface_import_parityguard it. See core/src/metal/AGENTS.md. -
Metal
float_ms_ssimoption parity (ADR-1334):float_ms_ssim_metalexposesenable_db,clip_db,enable_chroma, andenable_lcsmatching CPU/SYCL/HIP twins. It emitsfloat_ms_ssim,float_ms_ssim_cb, andfloat_ms_ssim_cron the GPU, enforces the >= 176 minimum plane dimension at init, resolves YUV400P to one plane before chroma validation, and uses the exact ceil-subsampled 351x351 YUV420P luma boundary. It wiress->enable_db, s->max_dbintovmaf_ms_ssim_emit_scores/vmaf_ssim_emit_score_named. Device-free contracts incore/test/test_metal_ms_ssim_option_semantics,core/test/test_metal_ms_ssim_options_contract.py, andcore/test/test_nonfinite_collector_wiring.pyprotect this against regression. -
SYCL
float_ssimdecimation mirrors the CPU's (ADR-1370):float_ssim_syclreproducesssim.c's box low-pass andiqa/decimate.c::iqa_decimate()bit for bit (int64 fixed-point window sum,KBND_SYMMETRIC,picture_copy()scaling) and sizes its planes with the sharediqa/decimate_dim.h. A change on the CPU side of that pipeline changescore/src/feature/sycl/integer_ssim_sycl.cppin the same PR. See core/src/feature/sycl/AGENTS.md and core/src/feature/iqa/AGENTS.md. -
HIP
float_ssimdecimation mirrors the CPU's (ADR-1405):core/src/feature/hip/float_ssim/ssim_decimate.his the window sum ofiqa/decimate.c::iqa_decimate()withssim.c's box low-pass (int64 fixed-point sum,KBND_SYMMETRIC,picture_copy()scaling), compiled by the kernel and bycore/test/test_hip_float_ssim_decimate.c, which holds it againstiqa_decimate()byte for byte. A change on the CPU side of that pipeline changes the header in the same PR. See core/src/feature/hip/AGENTS.md. -
CUDA
float_ssimis the CPU pipeline on the device (ADR-1399):core/src/feature/cuda/integer_ssim/ssim_score.cureproducesssim.c's box low-pass andiqa/decimate.c::iqa_decimate()(exact int64 window sum, one rounding,KBND_SYMMETRIC),iqa/convolve.c's fp32 products added to adoublesum in both Gaussian passes, and the ADR-1373 per-pixel combine; the host sizes the planes with the sharediqa/decimate_dim.h. Its score equals the CPU's on every measured frame, andtest_cuda_float_ssim_parityasserts equality. A change on the CPU side of that pipeline changes the kernel in the same PR. See core/src/feature/cuda/AGENTS.md and core/src/feature/iqa/AGENTS.md. -
integer_ssim_cudareturns the CPU'sssimbit for bit (ADR-1424):integer_ssim.c::calc_ssim()adds every pixel's term into one double in raster order, sointeger_ssim_vert_combine(core/src/feature/cuda/integer_ssim/integer_ssim_score.cu) stores the terms unreduced andssim_cuda.c::issim_frame_sum()adds the plane it reads back in index order. Do not reduce the double terms on the device and do not reorder the host loop; the int64 weights may stay a block reduction. A change tossim_reduce_row_range()or to the ordercalc_ssim()visits pixels changes the kernel'sissim_term()or the host sum in the same PR.core/test/test_cuda_ssim_exact_contract.pyguards it without a device,test_cuda_ssim_parityon one; the parity gate compares the twin with tolerance 0 (EXACT_TWINS, featuressim). See core/src/feature/cuda/AGENTS.md. -
float_ssim_cudaframe sums are the CPU's, in the CPU's order (ADR-1464): the pass-2 kernels ofcore/src/feature/cuda/integer_ssim/ssim_score.custore every window's terms at its raster position andinteger_ssim_cuda.c::float_ssim_frame_sum()/float_ssim_frame_sums_lcs()add them in index order, asiqa/ssim_tools.c::iqa_ssim()adds them. A sync must not bring back a device reduction of the terms: on the frame ofcore/test/float_ssim_order_frame.ha per-block sum returns the neighbouringfloat. That header is shared with the HIP and SYCL twin tests and its bytes are fixed. Preserve the kernels, the two host loops,core/test/test_cuda_float_ssim_order.candcore/test/test_cuda_float_ssim_exact_contract.pytogether. -
SYCL
float_ssimadds the CPU's terms in the CPU's order (ADR-1463):core/src/feature/sycl/sycl_ssim_terms.h::ssim_double_terms()formsiqa/ssim_accumulate_lane.h'slvandcvas the CPU's doubles in 64-bit integers;float_ssim_syclstores every window's term unreduced and the host adds them in raster order, asiqa/ssim_tools.c::iqa_ssim()does. No reduction may return to the float twin and no host sum may change its order: either moves thefloatmean by one step on frames whose terms cancel. A change tossim_accumulate_lane.hor to the means ofssim_tools.cchanges the header in the same PR.core/test/test_sycl_float_ssim_exact_contract.py(device-free) andcore/test/test_sycl_float_ssim_parity.c(==, with the constructed pair ofcore/test/float_ssim_order_frame.h, a file shared byte for byte with the CUDA and HIP tests) guard it. See core/src/feature/sycl/AGENTS.md. -
integer_ssim_syclreturns the CPU's score bit for bit (ADR-1443):core/src/feature/sycl/sycl_integer_ssim_math.hruns the fp64 operations ofinteger_ssim.c::ssim_reduce_row_range()'s per-pixel term, one for one and in the reference's order, on values held in 64-bit integers (core/src/feature/sycl/sycl_soft_signed.h, onsycl_soft_double.h). The kernel stores the bit pattern of every term unreduced and the host adds the plane incalc_ssim()'s raster order. A change to that expression ininteger_ssim.cchanges the header in the same PR. The twin stays free offloatin its term, of a device reduction of the terms, and of scratch memory (SIMD-16 with the 256-entry register file).core/test/test_sycl_integer_ssim_math.c(host and device),core/test/test_sycl_ssim_exact_contract.pyandcore/test/test_sycl_ssim_parity.cguard it. See core/src/feature/sycl/AGENTS.md.
Motion¶
-
CUDA RC3 CPU parity (ADR-1372, ADR-1373, ADR-1374): both CUDA motion twins run the diff-first SAD kernel of
integer_motion_v2/motion_v2_score.cuthroughinteger_motion_sad_cuda.c; an upstream sync must not bring back the blur-each-framemotion_score.cu.psnr_cuda,integer_ssim_cuda,float_ssim_cudaandfloat_motion_cudacarry the CPU option tables and call the CPU's helpers (psnr_score.h,vmaf_ssim_max_db(),motion_clip());ssim_score.cu::ssim_terms()mirrors the CPU'sl * c * srounding point for rounding point, andinteger_ssim_scorebuilds with--fmad=false(every CUDA fatbin's, ADR-1403) and the CPU's grouping. The integer ADM DWT row and tap arithmetic lives ininteger_adm/adm_dwt2_rows.h, andvif_cudafalls back to the CPU below 16 pixels.float_motion_cudaemits the CPU'smotion3(motion_blend_clip()). The motion SAD, PSNR and moment kernels add one atomic per block (per accumulator) and PSNR selects its plane with constant indices (ADR-1392). Details: core/src/feature/cuda/AGENTS.md. -
float_motion_cudaadds its SAD in the CPU's order (ADR-1409):float_motion.c::compute_motion_simd()keeps one fp32 running sum per row and one over the rows. The twin'sfloat_motion_row_sadkernel runs one thread per row with a plain left-to-right loop, and the host finishes throughcore/src/feature/float_motion_sad.h; the scores are the CPU's bit for bit and the parity gate compares them with tolerance 0 (EXACT_TWINS). A change to the CPU's SAD order, or toconvolution_f32_c_s()'s tap order, changes the kernel and the helper in the same PR.core/test/test_cuda_float_motion_parity.c,core/test/test_float_motion_sad.candcore/test/test_cuda_kernel_source_contract.pyguard it. -
float_motion_sycladds its SAD in the CPU's order (ADR-1411): the same contract as the CUDA twin above.fm_row_sad()incore/src/feature/sycl/float_motion_sycl.cppis one plain left-to-right loop per work-item, launched oversycl::range<1>(height)at sub-group size 16 (ADR-1468), andcollect()finishes throughcore/src/feature/float_motion_sad.h. No group, sub-group or atomic reduction may return to the TU, and the blur needs the SYCL strict FP line (ADR-1367).core/test/test_sycl_float_motion_parity.c(==) andcore/test/test_sycl_kernel_source_contract.pyguard it; the row kernel must stay free of scratch memory (test_sycl_kernel_scratch, ADR-1395). Since 2026-10-03 the twin also emitsmotion3on the host with the CPU'smotion_blend_clip()and declaresmotion_blend_factor/motion_blend_offsetas the CPU table does (T-GPU-FLOAT-MOTION3-MISSING-2026-09-30). A change to howfloat_motion.cemitsmotion3(index 0 from the first SAD, the flush tail, 0 for one frame) changescollect_fex_sycl()/flush_fex_sycl()in the same PR;test_sycl_twin_option_paritycompares every output with==, and the gate'sfloat_motioncell listsmotion3. -
motion_five_frame_windowis Netflix's, on the fork's picture ownership (ADR-1478):extract()and the window ofcore/src/feature/integer_motion.care upstream's statements (a2b59b77,a4a1492d); a sync takes upstream's side for the arithmetic and puts a change to upstream'sflush()intomotion_flush_one()/vmaf_motion_window_flush()(core/src/feature/motion_window.h), whichinteger_motion_v2.ccalls too. That file is deleted upstream and kept here. Incore/src/libvmaf.cupstream struct-copiesprev_refandprev_prev_refinto the extractor and zeroes them; the fork hands out counted references (fex_take_prev_refs()/fex_release_prev_ref(), ADR-0778) and rotates them in the PREV_REF swap offeature_extractor.cpp: keep the fork's side of those hunks and take only which frames are kept. Upstream keeps frame n-2 for every run; the fork keeps it only while a registered extractor'sreads_prev_prev_ref()answers true (the option is on), so a context without one holds what it held before the port, and a preallocated pool below four pictures is then refused with-EINVALat registration or atvmaf_preallocate_pictures(), never left to stall (a deliberate deviation that moves no score). A sync must not bring back the unconditional window or the unconditionaln_threads * 2 + 2ofcheck_picture_pool(). A sync must not bring back a@unittest.skipon the five-frame or_hfrtests underpython/test/.core/test/test_motion_five_frame_window.c,test_read_pictures_failure_ownershipand the Netflix golden gate guard it. See core/src/feature/AGENTS.md and core/src/AGENTS.md. -
The CUDA, SYCL and HIP motion twins compute
motion_five_frame_window(ADR-1491): with the option each twin ofmotionandmotion_v2takes its SAD against the frame two back (CUDA andmotion_v2_sycl: a ring of three raw planes; HIP: two kept planes;motion_sycl: two planes with fixed roles, advanced by two device copies behind the graph replay, with the kernel enqueued on every frame) and derivesmotion2/motion3with the CPU'svmaf_motion_window_flush()(core/src/feature/motion_window.h). Themotion_v2twins hold no copy of the CPU flush. A sync or a cleanup must not bring back a twin's own window arithmetic, aVMAF_OPT_FLAG_DEFAULT_ONLYor-ENOTSUPfor the option on these six twins, or movemotion_sycl's plane copies into the recorded graph. A change tomin_idxor to the frameextract()differences against ininteger_motion.c/integer_motion_v2.cchanges the twins'ring/depthin the same PR.test_{cuda,sycl,hip}_motion_five_frame_window(==, fixturecore/test/motion_five_frame_twin_parity.h) and the exact gate cellsmotion_mffw/motion_v2_mffwguard it on a device;core/test/test_gpu_option_value_capability_contract.pyandcore/test/test_{cuda,sycl,hip}_kernel_source_contract.py(the twins call the function and read no stored score back) without one. The Metal twins do not declare the option; the CPU extractor computes it there.
VIF¶
-
float_vif_syclreturns the CPU's scores bit for bit (ADR-1422): the same contract as the CUDA twin, without an fp64 type. The host takes each scale's Gaussian fromvif_get_filter()and hands it to the kernels by value.core/src/feature/sycl/sycl_float_vif_math.hisvif_pixel_statistic_s()andlog2f_approx()operation for operation; itsone_plus_ratio()evaluates the reference's two fp64 expressions as exact fp32 pairs and replays the fp64 operations in integers next to a rounding boundary.vif_row_sums()adds the terms of a row in one work-item andsum_vif_rows()adds the rows on the host, both in fp32. A change tovif_get_filter(), toVIF_OPT_FAST_LOG2/log2f_approx(), tovif_pixel_statistic_s()or tovif_statistic_s()invif_tools.cchanges that header in the same PR.core/test/test_sycl_float_vif_math.c(host and device),core/test/test_sycl_float_vif_exact_contract.pyandcore/test/test_sycl_float_vif_parity.cguard it; every kernel must stay free of scratch memory (test_sycl_kernel_scratch, ADR-1395). See core/src/feature/sycl/AGENTS.md. -
vif_syclreturns the CPU's scores bit for bit (ADR-1432):core/src/feature/sycl/sycl_integer_vif_math.hreturns the two integersinteger_vif.c::vif_accumulate_pixel()truncates from its fp64 gain (sigma2_sq - g * sigma12andg * g * sigma1_sq), from one integer division and, for a sample within the fp64 chain's rounding error of an integer, from the reference's fp64 operations replayed in 64-bit integers (core/src/feature/sycl/sycl_soft_double.h, shared withfloat_vif_sycl). The host tail rounds each scale's sums tofloatasvif_store_residuals()does. A change to those lines ofinteger_vif.c(the same lines are inx86/vif_avx2.c,x86/vif_avx512.candarm64/vif_neon.c) changes the header in the same PR. The kernels stay free of fp64, ofsycl::mul_hi()on 64-bit operands (wrong values on an Arc A380) and of scratch memory.core/test/test_sycl_integer_vif_math.c,core/test/test_sycl_vif_exact_gain_contract.pyandcore/test/test_sycl_vif_parity.cguard it. See core/src/feature/sycl/AGENTS.md. -
float_vif_cudareturns the CPU's scores bit for bit (ADR-1412): the host takes each scale's Gaussian fromvif_get_filter(), asfloat_vif.cdoes, and hands it to the kernels; no kernel file holds a tap.core/src/feature/float_vif_gpu_common.h(shared withfloat_vif_hipsince ADR-1444; CUDA compiles it throughcore/src/feature/cuda/float_vif/float_vif_device.h, which maps its operators to the__fmul_rn()family) isvif_pixel_statistic_s()andlog2f_approx()operation for operation (vif_sigma_nsqin fp64),float_vif_row_sumsadds the terms of a row in one thread, andfvif_sum_rows()adds the rows on the host, both in fp32 asvif_statistic_s()does. A change tovif_get_filter(), toVIF_OPT_FAST_LOG2/log2f_approx(), tovif_pixel_statistic_s()or tovif_statistic_s()invif_tools.cchanges that header in the same PR.core/test/test_float_vif_device_math.candcore/test/test_cuda_float_vif_exact_contract.pyguard it without a device,test_cuda_float_vif_parityon one; the parity gate compares the twin with tolerance 0 (EXACT_TWINS). See core/src/feature/cuda/AGENTS.md. -
vif_cudareads the CPU's log2 table (ADR-1462):core/src/feature/cuda/integer_vif/vif_statistics.cuhholds the table as the module globalvif_cuda_log2_table,log2_lookup()reads it with the CPU's mask, and no vif kernel source evaluates a logarithm.integer_vif_cuda.c::init_fex_cuda()fills it withvif_log2_table_generate()'s values throughvmaf_cuda_vif_upload_log2_table()before any frame is submitted. When upstream changesvif_statistics.cuhorfilter1d.cu, keep the lookup and do not bringlog_generate()back.core/test/test_cuda_vif_log2_contract.pyguards it without a device,test_cuda_vif_log2_tableon one. -
Every lane of a warp reaches the CUDA warp reductions (T-CUDA-WARP-REDUCE-UB-2026-10-05):
warp_reduce()/warp_reduce_u64()(core/src/cuda/cuda_helper.cuh) shuffle with the full mask, sovif_hori_kernel()incuda/integer_vif/filter1d.cucallsvif_hori_flush_accums()after the per-laney < h && x_start < wbranch, the lanes past the plane edge adding zeros; upstream keeps the flush inside it. Andwarp_reduce(int64_t)adds throughwarp_reduce_u64()on the unsigned bits: do not bring back the two-halves form, which shifts a negative word left.core/test/test_cuda_warp_reduce_contract.pyguards both without a device. -
float_vif_hipreturns the CPU's scores bit for bit (ADR-1444): the twin compilescore/src/feature/float_vif_gpu_common.hwith its default operators, which round once only because every HIP kernel is built withhip_strict_fp_args;float_vif_score.hipdefines no operator, holds no tap and reduces nothing per block. The host takes the taps fromvif_get_filter()and passesvif_sigma_nsqas adouble. A change to the shared header is a change to both twins:core/test/test_hip_float_vif_exact_contract.pyandcore/test/test_float_vif_device_math.cguard it without a device,test_hip_float_vif_parityon one. See core/src/feature/hip/AGENTS.md.
ADM¶
-
float_adm_syclreturns the CPU's scores bit for bit (ADR-1434):core/src/feature/sycl/sycl_float_adm_math.his the same arithmetic as the CUDA twin's device header, without an fp64 type: the three expressionsadm_tools.cevaluates indouble(the enhancement gain, the 1/30 product and the centre tap's 1/15 product) are exact fp32 pairs, and a result next to an fp32 rounding boundary replays the fp64 operations in 64-bit integers (sycl_soft_double.h). The decouple's quotient is fp32n / d, as the reference'sDIVS()is since ADR-1442; it must not become a product with a reciprocal. The header also holds what one work-item of the decouple, term and row-sum kernels does;float_adm_sycl.cpponly launches them. A row is added by one work-item and the rows by the host, both in fp32. The weights, the region, the pooling and the floor are the reference's own. A change toadm_decouple_s(),adm_csf_s(),adm_cm_thresh3x3_s(),adm_csf_den_scale_s()oradm_cm_s()changes this header and the CUDA one in the same PR.core/test/test_sycl_float_adm_math.candcore/test/test_sycl_float_adm_exact_contract.pyguard it,test_sycl_float_adm_parityon a device; the twin is declared exact byscripts/ci/exact_twins.d/float_adm.sycl. See core/src/feature/sycl/AGENTS.md. -
Float ADM CSF weights are upstream's float arithmetic (ADR-1489):
dwt_quant_step()incore/src/feature/adm_tools.hkeepsr,tempandQinfloatand raises 10 to thefloatproductparams->k * temp * temp, as upstream'sadm_tools.hdoes; a sync takes upstream's side and keeps the suppression comment.barten_csf_tools.hforms each product and quotient infloat, as upstream does, and promotes the result with an explicit cast, because the SYCL and Metal twins of integer ADM compile the header as C++, where upstream's implicit promotion calls thefloatmath functions: keep the casts around the results, never on an operand.core/src/feature/metal/float_adm_metal.mmholds a copy of the step and changes with it. With these the only difference between the fork'sfloat_admand Netflix's on x86 is the division (ADR-1442).core/test/test_float_adm_csf_upstream.c(values, the bits of a Netflix build, C against C++) andcore/test/test_float_adm_csf_upstream_contract.py(source shapes, the Metal copy) guard it; the Netflix golden gate would not notice. -
Integer ADM quantisation step is upstream's (ADR-1475):
dwt_quant_step()incore/src/feature/integer_adm_kernels.hraises 10 toparams->k * temp * temp, afloatproduct, exactly as upstream'sinteger_adm.cdoes; with a(double)on an operand (the form #552 introduced) every integer ADM score andvmaf_v0.6.1leave Netflix's values by up to 1.8e-5. A sync takes upstream's side of the statement and keeps the suppression comment above it. The SYCL twin holds its own copy (sycl/integer_adm_sycl.cpp) and changes with it; the Metal twin takes the CPU'sadm_csf_factors()inmetal/integer_adm_metal_host.cand holds no copy.core/test/test_integer_adm_quant_step.c(values, with the bits of a Netflix build) andcore/test/test_integer_adm_quant_step_contract.py(both copies, and no Metal copy) guard it; the Netflix golden gate would not notice. -
Integer ADM scale-0 masking centre tap (ADR-1402): the fork keeps the 1/15 centre tap of the masking threshold in int32 and clamps \(\lvert x \rvert - \mathrm{thr} \cdot 2^{\mathrm{shift}}\) to [0, INT32_MAX] in int64, where upstream master narrows the tap to int16 and subtracts in 32 bits (the fork's own Netflix/vmaf PR #1602, second revision, is not merged upstream). The scalar definition is
adm_cm_thresh()incore/src/feature/integer_adm_kernels.handadm_cm_excess_s0()incore/src/feature/adm_cm_accumulator.h; the AVX2, AVX-512, CUDA, HIP, SYCL and Metal twins return its value bit for bit and change together with it. A sync must not restore the(int16_t)cast, the 16-bit sign extension in the vector thresholds,adm_i16()on the SYCL centre term orabs(x) - (thr << shift).adm_avx2.candadm_avx512.cno longer carry upstream's macros: their scalar parts are the shared kernels.test_integer_adm_cm_threshold,test_integer_adm_simdandtest_gpu_adm_tiny_framesguard it; the Netflix golden gate must be re-run on any change. See core/src/feature/AGENTS.md and the "fix/adm-cm-centre-tap-wrap" entry of rebase-notes. -
Integer ADM scale-0 contrast-masking rows are unsigned: a scale-0 row of the masking reduction is a sum of non-negative cubes that passes INT64_MAX on a 64-pixel-wide picture at the default CSF weights and on any picture with a weight near its ADR-1472 budget. The CPU sums it, and the frame, in
uint64_t(adm_cm_round_row_total_s0()incore/src/feature/adm_cm_accumulator.h,adm_cm_fold_s0()incore/src/feature/integer_adm_kernels.h), the AVX2 / AVX-512 rows add unsigned, and the CUDA, HIP, SYCL and Metal twins sum unsigned too. Upstream Netflix/vmaf sums int64; a sync keeps the fork's form. Scales 1-3 stay signed.test_integer_adm_cm_row_unsigned,test_gpu_adm_tiny_framesandtest_adm_cm_row_rounding_contract.pyguard it. See core/src/feature/AGENTS.md. -
Integer ADM enhancement gain limit (ADR-1413): the limited sample is the double product
rst * adm_enhn_gain_limittruncated toward zero, as the scalar kernels incore/src/feature/integer_adm_kernels.hstore it. The AVX2 and AVX-512 decouple kernels use the truncating conversions (upstream master rounds), and the SYCL twin forms the same value in integers withadm_gain_limit_product()fromcore/src/feature/adm_gain_limit.h. A sync must not restore_mm256_cvtpd_epi32/_mm512_cvtpd_epi32/_mm512_cvtpd_epi64on the product or a fixed-point limit in the twin.test_integer_adm_simd,test_adm_gain_limitandtest_gpu_adm_tiny_framesguard it. See core/src/feature/AGENTS.md and the "fix/adm-decouple-fractional-gain-truncation" entry of rebase-notes. -
adm_cudareturns the CPU's scores bit for bit (ADR-1416):core/src/feature/cuda/integer_adm_cuda.cincludescore/src/feature/integer_adm_kernels.hand takes its CSF weights (adm_csf_factors()), its denominator border and shifts (adm_csf_den_ctx_init(),i4_adm_csf_den_ctx_init()) and its per-scale scores (adm_cm_result(),adm_csf_den_result()and theiri4_forms) from it; it defines none of them itself.integer_adm/adm_csf_den.cufolds one whole row per block throughadm_csf_den_round_row_total()(adm_cm_accumulator.h), with the shifts as kernel arguments. A change to those CPU routines reaches the twin through the header; a change to how the CPU folds a denominator row changes that kernel in the same PR.core/test/test_cuda_adm_exact_contract.pyandtest_adm_cm_row_roundingguard it without a device,test_cuda_adm_parityon one; the parity gate compares the twin with tolerance 0 (EXACT_TWINS). See core/src/feature/cuda/AGENTS.md. -
SYCL integer ADM AIM pass (ADR-1362):
integer_adm_sycl.cppcomputes aim / adm3 on the device and finalises every ADM output in the CPU's float arithmetic (bit-exact with the CPU). The decouple quotient is clamped in int64 before narrowing. Details and the mirror list for upstreaminteger_adm.cchanges: core/src/feature/sycl/AGENTS.md. -
Integer AIM is not clipped, float AIM is (ADR-1417):
core/src/feature/integer_adm.creportsaim_num / den(vmaf_adm_scale_ratios()),core/src/feature/adm.creportsMIN(aim_num / aim_den, 1)(vmaf_adm_finalize_scores()), each as its upstream file does. The shippedvmaf_v1.0.16models read the integeradm3, which the unclipped AIM takes down toadm_min_val. A sync or a cleanup must not unify the two unless upstream does; the Netflix golden gate has no integer AIM above 1 and would not notice.core/test/test_integer_adm_aim_unclipped.cpins both sides. -
float_adm_hipreturns the CPU's scores bit for bit (ADR-1458): it compilescore/src/feature/float_adm_gpu_common.h, the arithmetic of the CUDA twin (next entry), throughcore/src/feature/hip/float_adm/float_adm_hip_math.h, which keeps the shared header's plain operators: under the strict FP list of the HIP kernels they are the reference's operations, the division included. Do not respell them with the__fmul_rn()family, do not reduce per wave or block, and keep the host on the reference's routines (adm_float_reference.h) withadm_frame_size_check()first ininit.core/test/test_hip_float_adm_exact_contract.pyguards it without a device,test_hip_float_adm_math(device arithmetic against the host, value by value) andtest_hip_float_adm_parityon one. See core/src/feature/hip/AGENTS.md. -
float_adm_cudareturns the CPU's scores bit for bit (ADR-1420):core/src/feature/float_adm_gpu_common.h(shared withfloat_adm_hip;core/src/feature/cuda/float_adm/float_adm_device.hgives it the CUDA device spelling) is the decouple, the CSF, the masking threshold and the reduction terms ofadm_tools.coperation for operation (the gain limit and the 1/30 and 1/15 constants in fp64, the angle threshold as(cos^2 * |o|^2) * |t|^2), and its division is the reference's, the IEEE fp32 quotient (__fdiv_rn(); see the next entry).float_adm_row_sumsadds each row in one thread and the host adds the rows, both in fp32. The weights, the reduced region, the pooling and the angle constant come fromadm_tools.citself throughcore/src/feature/adm_float_reference.h; keep those exports, and keep the four reductions ofadm_tools.conadm_pool_bands_s(). A change toadm_decouple_s(),adm_csf_s(),adm_cm_thresh3x3_s(),adm_csf_den_scale_s()oradm_cm_s()changes the device header in the same PR.core/test/test_float_adm_device_math.candcore/test/test_cuda_float_adm_exact_contract.pyguard it without a device,test_cuda_float_adm_parityon one; the parity gate compares the twin with tolerance 0 (EXACT_TWINS). See core/src/feature/cuda/AGENTS.md. -
Float ADM divides (ADR-1442):
core/src/feature/adm_options.hdoes not defineADM_OPT_RECIP_DIVISIONandcore/src/feature/adm_tools.chas oneDIVS(), the plain quotient, with an#errorif the macro is defined. Upstream Netflix defines the macro and multiplies by a reciprocal refined from the processor'sRCPSSestimate, which made the scores depend on the processor. An upstream sync that touches either file keeps the fork's side of both hunks: no macro, norcp_s(), no<emmintrin.h>inadm_tools.c. No twin may bring a reciprocal estimate, a probe of the host or a table of it back, and no CUDA flag may relax the division (--use_fast_math,-prec-div=false).core/test/test_float_adm_divides_contract.pyscans the reference, everyfloat_admfile of every backend andcore/src/meson.build;core/test/test_float_adm_device_math.cchecks the value on inputs where the estimate and the quotient differ.
CIEDE2000¶
-
ciede_cudaruns the CPU's arithmetic (ADR-1426):core/src/feature/cuda/integer_ciede/ciede_device.hisciede.c'sget_lab_color()andciede2000()statement for statement: fp64 where the reference computes in double, float where it stores in float, every float-to-double promotion of a libm argument written out (the kernel is C++). The reference's two float products,c_prime_1 * c_prime_2andr_sub_t * chroma * hue, are upstream's and are float products in every twin (ADR-1476); a sync takes upstream's side of them and no(double)goes in front of either. The kernel stores one float per pixel andciede_frame_sum()(core/src/feature/ciede_frame_sum.h, one definition for the CUDA, SYCL and HIP hosts) adds the read-back plane in raster order. Do not introduce float math functions, a device reduction or another form of the formula. A change toget_lab_color(),ciede2000(),get_r_sub_t()or the order ofextract()'s sum inciede.cchanges that header in the same PR. The twin is not bit-identical (glibc's math library against CUDA's); the gate bounds it at1e-9throughLIBM_TWINS.core/test/test_ciede_device_math.candcore/test/test_cuda_ciede_exact_contract.pyguard it without a device,test_cuda_ciede_parityon one. See core/src/feature/cuda/AGENTS.md. -
ciede_syclruns the CPU's arithmetic on fp32 pairs (ADR-1436):core/src/feature/ciede_ff_math.his the same statements as the CUDA twin'sciede_device.hfor a device without an fp64 type: every fp64 value is an fp32 pair, every math-library call a function ofcore/src/feature/ff_math.h, everyfloatof the reference a float rounded from the pair at the reference's statement. Both headers are backend-neutral and shared withciede_hip(ADR-1448);core/src/feature/sycl/sycl_ciede_math.handsycl_ff_math.honly name the SYCL primitives they are built on. The kernel stores one float per pixel andciede_frame_sum()adds the read-back plane in raster order. Do not introduce the device's fp32 math functions, a device reduction, another form of the formula, or a call the compiler does not inline:ciede_pixel()is flattened into the kernel because a call frame is scratch memory (ADR-1395). The header's constants and tables come fromscripts/dev/gen_sycl_ff_math.py; the tables are read from device memory. A change toget_lab_color(),ciede2000(),get_r_sub_t()or the order ofextract()'s sum inciede.cchanges this header and the CUDA one in the same PR. The twin is not bit-identical (the host'spowf); the gate bounds it at1e-9throughLIBM_TWINS.core/test/test_sycl_ciede_exact_contract.pyguards it without a device,test_sycl_ciede_mathandtest_sycl_ciede_parityon one. See core/src/feature/sycl/AGENTS.md. -
ciede_hipruns the same fp32-pair statements (ADR-1448):core/src/feature/hip/integer_ciede/ciede_score.hipincludescore/src/feature/ciede_ff_math.hthroughcore/src/feature/hip/integer_ciede/ciede_hip_math.h, which names the HIP primitives (core/src/feature/ff_pair.hon plain fp32 operators under the strict FP list,fmaf(),sqrtf(),cbrtf(),expf(0.2f * logf(x))). The device has fp64, but its fp64 math functions cost 17 times the frame time; do not bring them back. The kernel stores one float per pixel and the host adds the plane withciede_frame_sum(). A change to a shared header changes the SYCL twin too: both are re-measured (A380 and gfx1036) in the same PR. The gate bounds the cell at1e-9(LIBM_TWINS), not 0.core/test/test_hip_ciede_exact_contract.pyandtest_hip_ciede_mathguard it without a device,test_hip_ciede_parityon one.