Rebase notes¶
ADM second viewing distance on HIP (2026-10-09)¶
core/src/feature/hip/integer_adm_hip.c:adm_norm_view_dist_extra(nvde) in the option table;adm_hip_scale0()/adm_hip_scale123()became*_transform()(DWT) and*_weigh()(denominator, CSF, CM, AIM at one distance, throughadm_hip_view_buffer());rfactor/i_rfactorare indexed by distance andadm_hip_fixed_params()takes the distance;tmp_resandresults_hosthold one result block per distance; the context initialisers and the host conclusion take the distance asnvd. The merge callback and the second distance's names are the sharedadm_view_dist.c. Fork-only, nothing upstream to merge. On sync: an upstream change to the HIP ADM driver lands in the matching half; keep the per-distance loop.core/test/test_hip_adm_exact.c:test_adm_two_views_exactandtest_adm_merged_registrations_exacthold both distances to the CPU's bits.test_hip_adm_parity.cdrops its recorded option gap;test_hip_adm_exact_contract.pyandtest_hip_adm_buffer_pointer_contract.pypin thenvdcontexts, theres_bytesclear andadm_hip_init_names()'s failure path.
Netflix/vmaf 9f4bd165f: integer_vif copies each row's samples (2026-10-09)¶
core/src/feature/integer_vif.cextract(): the same change as upstream (row_bytes = w << (bpc > 8)), with a comment. On sync: identical; take either side.core/test/test_integer_vif_row_copy.c(fork-only): wide-stride luma ending at an inaccessible page; scores equal at either stride.4068ee3b5(CAMBI 10-bit same-size copy with stride) is not ported: the fork copies row by row sinceT-CAMBI-10BIT-FULLREF-WIDE-SOURCE-ROWS-2026-10-05(decimate_same_size_16b()incore/src/feature/cambi.c), and the Rust twin (core/src/rust/feature/cambi/src/preprocess.rs) does too. On sync, keep the fork's function.
VMAFx device frames on HIP (RC4 WP3 HIP lane)¶
rc4/api-wp3-hip, ADR-2092, ADR-2023, ADR-1929.
- New files
core/src/hip/vmafx_hip.h,vmafx_hip_internal.h,import_device.c,import_frame.c,import_dmabuf.c,import_fence.c,import_gl.c(inlibvmaf_sourcesunderif is_hip_enabled, not inhip_sources, for the reason the CUDA lane gives) and the kernel moduleimport_convert.hip(hip_kernel_sources,import_convert_hsaco). - Moved out of the CUDA lane into shared files, one copy each:
core/src/vmafx/import_convert_kernels.h(the NV12 / P010 / P016 kernels;cuda/import_convert.cuandhip/import_convert.hiponly include it; it is in bothcuda_kernel_shared_headersandhip_kernel_shared_headers),core/src/vmafx/gl_sync.c(the GL loader andvmafx_gl_sync_acquire(), formerly incuda/import_gl.c),core/src/vmafx/release_events.c(the release-event table, formerly incuda/import_fence.c), and the newcore/src/vmafx/sync_file.c. A rebase of the CUDA lane that touches the old copies applies the change to the shared file. core/src/vmafx/*.cdispatch to the HIP lane next to the CUDA one (device.c,device_context.c,frame_import.c,fence.c,frame_import_admit.c,submit.c);fence.cwaits onSYNC_FILEandGL_SYNCfences without a device;frame_pool.crefuses HIP devices.core/src/picture.h:VmafPicturePrivategains an unconditionalhip.str(one layout in every translation unit;HAVE_HIPreaches only some throughconfig.h).core/src/libvmaf.c:translate_picture()passes a HIP device picture through, andhip_refuse_other_reader()refuses it to a non-HIP extractor.core/src/hip/picture_hip.candshared_frame.c: a device picture is copied on its library stream and the reader's stream and the null stream wait (vmaf_hip_stream_wait_library()); an upstream or fork change to the upload path keeps the device branch first.integer_psnr_hvs_hip.c,ssimulacra2_hip.candinteger_ms_ssim_hip.c(with the new kernelms_ssim_picture_to_floatininteger_ms_ssim/ms_ssim_score.hip) branch onvmaf_hip_picture_device_stream()before their host staging;core/test/test_vmafx_import_hip_contract.pyholds the order.core/meson_options.txt:enable_float_vif_hip_autodispatchdefaults totrue; its description and thecore/src/meson.buildcomment keep theADR-0623citationstest_stale_text_contract.pyreads.- Tests:
core/test/vmafx_cuda_cells.his nowvmafx_device_cells.h(both lanes readscripts/ci/exact_twins.d/); the CUDA tests include the new name. core/src/hip/import_frame.cfill_failed(): a GL texture the runtime maps but refuses to read (hipErrorInvalidValue, every read on ROCm 10.1) isVMAFX_E_NOTSUPnamingdesc.memory, notVMAFX_E_DEVICE;test_vmafx_import_hip_glskips on that refusal. Keep both when the GL path changes.- No
libvmaf.h, ABI (0.1.4 unchanged), golden-data or FFmpeg patch impact;--backend hipnow runsfloat_vif_hipforfloat_vif(the CPU's scores, ADR-1444).
ADM second viewing distance on SYCL (2026-10-09)¶
core/src/feature/sycl/integer_adm_sycl.cpp:adm_norm_view_dist_extra(nvde) in the option table;rfactor,i_rfactorandcsf_normalization_shiftare indexed by viewing distance;enqueue_adm_reductions()runs once per distance after the scale's DWT, into that distance's block ofd_accum;collect_view()concludes one distance. The merge callback and the second distance's names are the sharedadm_view_dist.c. Fork-only, nothing upstream to merge. On sync: an upstream change to the SYCL ADM driver keeps the per-distance loop and the[view]index.core/test/test_sycl_adm_parity.c: the recorded option gap is gone;test_adm_two_views_exactandtest_adm_merged_registrations_exacthold both distances to the CPU's bits.core/test/test_adm_view_merge.candcore/test/test_adm_view_dist_contract.pylistadm_syclamong the merging descriptors.
vif_tools.c: AVX2 row dispatch taken for a filter with a tap (2026-10-09)¶
vif_filter1d_vertical_dispatch_s()incore/src/feature/vif_tools.ctakes the AVX2 row pass only forfwidth >= 1, the condition under which its row table is filled for every entryconvolution_f32_avx_rows_s()reads. cppcheck 2.19.0 reportsuninitvaron the table without it. On sync: keep thefwidth >= 1 &&in front ofvif_use_avx2_convolution()when upstream changes that dispatch; no score changes with it (the Netflix golden gate passes unchanged).
Netflix/vmaf 3e1385bed: FFmpeg built with MSVC in CI (2026-10-08)¶
- Upstream adds a Windows row to
.github/workflows/ffmpeg.ymlthat builds FFmpegmasterthrough a Meson port of FFmpeg. The fork's equivalent is theffmpeg-msvc-work/ffmpeg-msvc-gatejobs of.github/workflows/ffmpeg-integration.yml(required checkFFmpeg Windows MSVC, ADR-2783), which runffmpeg-patches/test/build-and-run.shwithFFMPEG_TOOLCHAIN=msvc. On sync: do not import upstream'sffmpeg.ymlrow; the fork has noffmpeg.yml. scripts/ci/upstream-consumer-lib.shnow holdsuc_ffmpeg_graph(moved fromupstream-ffmpeg-compat.sh'sff_graph); the smoke script's score check and the upstream-consumer check share it.docs/getting-started/building-on-windows.md: the ARM64 toolset notes stay under "Native MSVC on Windows ARM64"; "Threads on MSVC", "Library files of an MSVC build" and "FFmpeg with MSVC" follow as sections of their own.ffmpeg-patches/0002and0008callff_set_pixel_formats_from_list2()for theirenum AVPixelFormatlists (cl.exeC4133with the untypedff_set_common_formats_from_list2()). Keep the typed helper when refreshing the series. The smoke script's warning gate undermsvcreads only the lines the series writes and the linker (msvc_findings).
ADM second viewing distance on CUDA, shared merge helpers (2026-10-08)¶
core/src/feature/adm_view_dist.{c,h}(new): the merge callback and the second distance's names, read through the option table by name, for every ADM descriptor.integer_adm.cdrops its own copies;adm_cudasets the same hooks. Fork-only, nothing upstream to merge.core/src/feature/cuda/integer_adm_cuda.c:adm_scale0_device()/adm_scale123_device()became*_transform()(DWT) and*_weigh()(denominator, CSF, CM, AIM at one distance);tmp_resandresults_hosthold one result block per distance; the host conclusion takes the distance as an argument. On sync: an upstream change to the CUDA ADM driver lands in the matching half; keep the per-distance loop.core/test/test_cuda_adm_parity.c: the recorded option gap is gone;test_adm_two_views_exactandtest_adm_merged_registrations_exacthold both distances to the CPU's bits.
Windows math-constant define on the SYCL command lines (2026-10-09)¶
core/meson.build: the Windows-D_USE_MATH_DEFINESis one list,vmaf_math_constant_args(empty on other hosts), used for the project argument and the header checks. On sync: upstream (4e150067b) spellsadd_project_arguments('-D_USE_MATH_DEFINES', ...)literally; keep the fork's list form, whichcore/src/meson.buildreuses.core/src/meson.build:sycl_common_argsandsycl_feature_tail_argsadd the list, because icpx custom targets never see project arguments. A new icpx compile line must take one of the two lists (core/test/test_sycl_math_constants_contract.py).scripts/dev/preflight.sh: the msvcism stage reads the list form and both SYCL lists.
Netflix/vmaf 33e5f0aca + cffd5b77d: ADM shares two viewing distances (2026-10-08)¶
core/src/feature/feature_extractor.h:mergeas upstream, plus the fork-onlyextend_name_dicthook. On sync: keep both; upstream'smergesits betweencloseandoptions, the fork's afterreads_shared_luma_only(designated initializers make the position irrelevant).core/src/fex_ctx_vector.cpp:offer_merge()runs after the dedup loop and checks the extractor name andis_initialized; upstream offers inside its dedup loop on the callback pointer alone. On sync: keep the fork's pass.core/src/feature/integer_adm.c: the fork's per-frame driver is split into*_transform()and*_weigh()per scale; upstream rewrites its singleinteger_compute_adm()with an inner per-distance loop and anAdmScorepair. Upstream keeps a second dictionary (feature_name_dict_extra); the fork extends the one dictionary with<base>:nvdekeys throughadm_extend_name_dict(), which the Rust twin shim also calls.adm_merge_view_dist()differs from upstream'sadm_try_merge_view_dist(): it absorbs a repeat of the second distance and declines adebugincoming context (docs/development/known-upstream-bugs.md). On sync: take upstream's arithmetic changes into the*_weigh()helpers, keep the fork's merge rules.core/src/rust/feature/adm/:adm_rustmirrors the split and files the second distance underscore.rs::EXTRA_VIEW_NAMES.core/test/test_{cuda,sycl,hip}_adm_parity.c,core/test/test_metal_twin_option_tables_contract.py: the twins' missingadm_norm_view_dist_extrais a recorded gap until each backend's pull request adds the option and deletes the gap.
Netflix/vmaf 3b4dd350e: static MSVC builds install vmaf.lib / vmafx.lib (2026-10-08)¶
core/src/meson.build:vmaf_static_name_kwargs(name_prefix: '',name_suffix: 'lib') onlibvmafx = library('vmafx', ...)and on the compatstatic_library('vmaf', ...), forcc.get_argument_syntax() == 'msvc'withdefault_library=staticonly. Upstream splits itslibrary()intoshared_library()+static_library()and names thebothstatic halfvmaf-static.lib; the fork keeps itslibrary()(ADR-2752). On sync: do not import that split; keep the keywords on both targets..github/workflows/libvmaf-build-matrix.yml: the MSVC CUDA / SYCL and ARM64 legs runscripts/ci/check_msvc_library_names.py --prefix installafterninja install.
Netflix/vmaf 4e150067b (+ a8f536b9a): M_PI / M_E from (2026-10-08)¶
core/meson.build:-D_USE_MATH_DEFINESas a C and C++ project argument and intest_argson Windows hosts, beside_GNU_SOURCE(Linux) and_DARWIN_C_SOURCE(macOS). Upstream adds it forcc.get_id() == 'msvc'only; the fork also needs it for clang-cl, icx-cl and MinGW-w64 (whose<math.h>hides the constants under__STRICT_ANSI__, set by-std=c23).- Removed every
#ifndef M_PI/#ifndef M_Efallback and every in-file#define _USE_MATH_DEFINES:adm_csf_tools.h,adm_tools.c,adm_tools.h,barten_csf_tools.h,ciede.c,integer_adm.h,integer_ssim.c,speed.c,speed_internal.c,speed_qa.c,vif_tools.c,y_funque_plus.c,sycl/float_adm_sycl.cpp,sycl/integer_adm_sycl.cpp, and the teststest_adm_angle_flag.c,test_float_adm_csf_upstream.c,test_integer_adm_quant_step.c,test_speed_upstream_form.c. Upstream'sa8f536b9a(guards) is superseded by this. On sync: delete a fallback a sync brings back (core/src/feature/AGENTS.d/math-constants.md).
Netflix/vmaf 8bc5a5c6a + b41d2340a: integer ADM NEON kernels for every scale (2026-10-08)¶
core/src/feature/arm64/adm_neon.c/.h:adm_cm_neon(),i4_adm_cm_neon(),adm_dwt2_s123_combined_neon()andadm_decouple_s123_neon(), dispatched ininit_dispatch_simd()(integer_adm.c; contrast masking only withoutcsf_requires_normalization, as on x86).- The fork's form differs from upstream's on purpose:
- contrast masking: the kernels are interior-row callbacks of the scalar drivers
adm_cm_rows()/i4_adm_cm_rows()(one fold per row, ADR-1167) and finish throughadm_cm_result()/i4_adm_cm_result(), where upstream repeats the factor, shift and pooling code. The scale-0 centre tap stays int32 and the excess is formed in int64 and clamped (ADR-1402); upstream narrows the tap to int16 (vmovn_s32) and subtracts the threshold in int32. The scale-0 row is summed in uint64 (adm_cm_fold_s0()). Upstream's three-row running sums intmp_refare not taken; each block reads its 3x3 neighbourhood. - scale 1-3 decouple: the angle flag of an overshooting lane comes from the scalar
adm_angle_flag(), not from upstream's vector double test. - scale 1-3 DWT: the int64 taps get the scale's rounding term from
i4_dwt2_round()and an arithmetic shift, asi4_dwt2_tap4(), and both pictures share the scalartmp_reflayout. On sync: do not import upstream'sadm_cm_threshold_neon()(int16 tap),adm_cm_sum_row()or its result code; port a change of the scalar kernels into these callbacks. - Tests:
test_integer_adm_simdruns the contrast-masking, centre-tap and decouple tests on aarch64 too, and gainstest_i4_adm_cm_matches_scalar_kernels(also againsti4_adm_cm_avx2/_avx512);test_adm_dwt2_neongainstest_adm_dwt2_s123_neon_matches_scalar.
Praetor pin 3a766f2d56ad, the REUSE workflow and the HISS-10 entries (2026-10-08)¶
chore/praetor-pin-3a766f2d, ADR-2784. A rebase or sync keeps PRAETOR_REF at 3a766f2d56ad... in .github/workflows/standards-gate.yml and the engine's texts of praetor-api.yml, praetor-docs.yml, tools/apicompat/gate/ and tools/figures/ (regenerate with adopt in a throwaway copy, never hand-edit). .github/workflows/reuse.yml is praetor's rendering with both actions pinned to commits; keep the pins, and keep REUSE lint in the aggregator's required and strictMustReport lists, in always and untiered_jobs of .github/ci-tier.json, and reuse-lint in the Pre-Commit job's SKIP. The exceptions: block of .standards.yaml is rendered from .config/lint-exceptions.d/ (HISS-10.toml and HISS-11.toml through PRAETOR_RULES); on a conflict take either side and run python3 scripts/ci/praetor_tidy_coverage.py --write.
Win32 pthread shim: timed wait; host fences on a condition variable (2026-10-08)¶
core/src/compat/win32/pthread.hgainspthread_cond_timedwait()overSleepConditionVariableSRW(). Its deadline arithmetic iscore/src/compat/win32/pthread_timeout.h(vmaf_w32_timeout_ms()), tested on every host bytest_win32_pthread_timeout. Upstream's bundled pthread-win32 (Netflix/vmafbcd6e6159,-Dbundled_winpthreads, and the pthread parts ofa2660554e,8618ba5cd,5079124ea,e2f8b24ab) is not taken (Q-302). On sync: do not add thelibvmaf/subprojects/pthread-win32submodule, the CMake subproject or the option.core/src/vmafx/fence.c: a host fence carries a mutex and a condition variable (CLOCK_MONOTONICon Linux);vmafx_host_fence_signal()broadcasts,vmafx_host_fence_wait()(now non-const) waits in chunks of at most 1 s against the monotonic deadline. The virtual test clock keeps the poll path. Backend lanes that add fence kinds keepvmafx_fence_poll().core/test/test_thread_pool_backpressure.creadstimespec_get(TIME_UTC)on MSVC; the Meson probehas_cond_timedwaitnow finds the shim's function, so the test builds on the MSVC lanes.core/test/test_win32_pthread_shim_contract.py:PROBEDis empty; Linux-onlypthread_condattr_*calls infence.care allowed inside their guard (PLATFORM_ONLY).
Upstream reconcile: arm64 ADM port hazards, #1551 closed (2026-10-08)¶
docs/upstream-reconcile-2026-10-08, ADR-1402, ADR-1413, ADR-1417, ADR-0155. Documentation only; no fork code changes.
arm64 ADM contrast masking and scales 1-3 (upstream 8bc5a5c6a, b41d2340a): port hazards. Port this code only if it is bit-exact with the fork's scalar kernels (ADR-1402, ADR-1413, ADR-1417), checked under qemu-aarch64 with GCC and clang. Three places to check before taking it:
adm_cm_threshold_neon()(arm64/adm_neon.c:364-380) narrows the scale-0 centre tap withvmovn_s32, andadm_cm_accum_neon()shifts the threshold in 32 bits (:401). The fork's scalar keeps the tap in 32 bits and forms the excess in 64 bits with a clamp (ADR-1402); a port must not copy these two lines.i4_adm_cm_threshold_neon()(:760-761) usesvdupq_n_s64(INT32_MIN)as the rounding term to match upstream's scalar (Netflix/vmaf#955). The fork's scalar keeps that behaviour (ADR-0155), so this one is correct to keep.adm_decouple_neon()falls back to the scalar kernel for non-integral gain limits (:273), which agrees with ADR-1413.
adm_cm_neon() only runs for widths of 32 and above and falls back to the scalar kernel below.
Update to the core/src/feature/compat_builtin.h entry (Netflix/vmaf#1551, retracting #1422) further down this page. Do not adopt Netflix/vmaf#1422's __lzcnt form. Upstream's #1551 retracted it and was closed on 2026-10-02 without merging; its algorithm is now upstream's own: 7388bd6fc (in the MSVC compat header) and 7437f3d9a.
Server workloads without VMAFX_BACKEND, unused chart helpers removed (2026-10-08)¶
rc4/api-wp17-cleanup, ADR-2350 D13. The server's Deployment, StatefulSet and Job render env only from .Values.env; no [[chart_env]] entry targets them, and the unread escape of [[chart_env]] is gone, so the generator refuses an entry the workload's binary does not read. vmafx.podSpec, vmafx.containerSpec, vmafx.volumes and templates/sidecar-trainer.yaml are deleted; cmd/vmafx-mcp/main.go has no LOG_LEVEL / LOG_FORMAT copy. A sync that brings any of them back drops it again. test_helm_config_env.py guards the server containers. no upstream file.
Netflix/vmaf ad42c532 + 9cb9479f: SpEED fused anti-alias filter on x86, AVX2 vertical pass (2026-10-08)¶
core/src/feature/speed.cfilter_and_downscale()and its mirrorspeed_internal_filter_and_downscale()(speed_internal.c): the#if ARCH_X86branch (vif_filter1d_s()+vif_dec16_s()) is gone; every target callsvif_filter1d_dec16_s()and copies the decimated plane back. On sync: keep the two files in lockstep, as before.core/src/feature/vif_tools.c:vif_filter1d_dec16_s()takes its vertical pass fromvif_filter1d_vertical_dispatch_s(), which callsconvolution_f32_avx_rows_s()(common/convolution_avx.c, declared incommon/convolution.h) under the same gate asvif_filter1d_s()'s AVX2 convolution. Upstream's form differs on purpose:- upstream's
convolution_f32_avx_dec16_s()carries both passes and a second scalar copy of the decimated horizontal pass; the fork vectorises only the vertical row and keeps one horizontal pass invif_tools.c; - upstream's
vif_filter1d_dec16_scalar_s()split is not taken: the CPU mask selects the scalar pass; - upstream's
VMAF_NO_FUSEasm barrier (convolution.h, and in the existing AVX scanlines) is not taken: every translation unit builds with contraction off (ADR-1461). On sync: do not import those three; port a tap-order or mirror change intoconvolution_f32_avx_rows_s()andvif_filter1d_vertical_s()together. - GPU twins (
cuda/speed/speed_score.cu,hip/speed/speed_hip_device.h,sycl/speed_sycl_pipeline.cpp): comment lines only (same line counts). The Rust twin already ran the fused filter. core/test/test_speed_filter.c: SIMD and old-x86-path comparisons, an aligned layout (the AVX2 convolution of the old path loads aligned), upstream's checkasm size 79x48, andtest_avx_rowsfor the masked tail the decimated pass never reads.
Go binaries' environment and the chart's VMAFX_* entries generated (2026-10-08)¶
rc4/api-wp17-config, ADR-2350 D13. [[config_binaries]], [[config]], [[chart_workloads]], [[chart_env]] and [[chart_maps]] of api/vmafx-platform.toml generate each binary's config_keys.gen.go (its golusoris CompoundKeys), the environment tables of the binaries' pages and of docs/usage/env-vars.md (between BEGIN/END GENERATED markers), and deploy/helm/vmafx/templates/_config.gen.tpl. A template includes vmafx.env.<workload> where its environment list holds the VMAFX_* entries; _helpers.tpl keeps no environment mapping. A change that adds a variable, edits a CompoundKeys list, an environment table row or a VMAFX_* entry of a template by hand moves into the definition instead; on a conflict in a generated file or region take either side and run scripts/codegen/vmafx-api.py --write. test_vmafx_api_generated_current guards the generated files and regions, scripts/ci/tests/test_helm_config_env.py that no other template writes a VMAFX_* entry, and cmd/vmafx-controller/env_test.go that the controller's grpc.* variables reach their keys. no upstream file.
motion2 / motion3 derived frame by frame¶
rc4/api-motion-incremental, ADR-2090, amends ADR-2074 decision 9.
core/src/feature/integer_motion.c: the flush body is split intomotion_window_stamp(),motion_window_count_sads(),motion_window_derive(),vmaf_motion_window_advance()andvmaf_motion_window_flush();motion_flush_one()is unchanged. An upstream sync offlush()(Netflixinteger_motion.c) ports the per-frame statements intomotion_flush_one()and the stamp intomotion_window_stamp(), never back into one loop inflush(): the advance would then miss them.integer_motion_v2.cfollows.motion_window.h:VmafMotionWindowgainsstate(VmafMotionWindowState); every caller ofvmaf_motion_window_flush()(CPU, CUDA, SYCL, HIP, Metal motion twins) also registers.advance.test_motion_window_advance_contract.pylists the TUs.core/src/feature/feature_extractor.h:VmafFeatureExtractorgains the optionaladvancecallback afterflush. A descriptor copied whole (the Rust twin shim of the RC4 Rust lanes) inherits it and must set or clear it.core/src/libvmaf.c:vmaf_engine_read_pictures()is split intoread_pictures_owned()and a wrapper that callsadvance_extractors();vmaf_read_pictures_sycl()andfence_for_read()call it too, andvmaf_engine_advance()exposes it to the VMAFx completion thread (advance_engine()incore/src/vmafx/window.c, engine lock, before each pass). An upstream sync ofvmaf_read_pictures()ports intoread_pictures_owned().vmaf_engine_feature_score_at_index()fences on-EINVALfor a fed frame.- Tests whose expectations moved:
test_score_pooled_eagain(score_pooled(i - 1, i - 1)afterread_pictures(i)now 0),test_vmafx_window(motion windows complete at step 4),test_gpu_float_ssim_auto_scale_contract.py(readsread_pictures_owned()). - No golden-data impact: every motion value is the flush-time value bit for bit (32 988 per-frame values against master, 0 different; Netflix golden gate green); no FFmpeg patch change (
libvmaf.hunchanged, the scores only arrive earlier). - Rust twins (lane request MI-1, Q-093):
VmafxRsTwingainsadvance(last field;VMAFX_RS_ABI_VERSIONstays 1, no release has shipped the Rust ABI); on a conflict incore/src/rust/include/vmafx_rs.htake either side and runscripts/dev/rust-abi-header.sh. The shim mapsadvancetotwin_advance(), never the inherited C callback, andadvance_one_extractor()(core/src/libvmaf.c) initialises a pooled Rust twin before its first advance.motion_rust'swindow.rsportsmotion_window_stamp(),motion_window_count_sads(),motion_window_derive(),vmaf_motion_window_advance()andvmaf_motion_window_flush()statement by statement: an upstream sync that changes them changeswindow.rsin the same PR (test_rust_motion_window_incremental,rust_twin_diff.py --feature motion).
Helm values and schema generated from the platform definition (2026-10-08)¶
rc4/api-wp17-helm, ADR-2350 D13. deploy/helm/vmafx/values.yaml and values.schema.json are written by scripts/codegen/vmafx-api.py from the [[chart]], [chart_root] and [[chart_defs]] tables of api/vmafx-platform.toml; the Kubernetes types of the schema come from api/kubernetes/openapi-subset.json (scripts/codegen/k8s_openapi.py, release and digests in build-config.env). A change that edits either chart file by hand moves into the definition instead; on a conflict in a generated file take either side and run the generator. A new values key is a new [[chart]] entry at its place in the file. test_vmafx_api_generated_current guards both files, test_k8s_openapi_subset_current the subset. no upstream file.
POSIX-only build parts off Windows, msvcism POSIX-header check (2026-10-08)¶
fix/msvcism-posix-headers, ADR-2646.
core/src/meson.build:enable_mcp=trueon Windows is a configureerror();subdir('mcp')andcompat/libvmaf/mcp.ccarryhost_machine.system() != 'windows'.core/test/meson.build: the MCP tests carry the same gate, andsubdir('fuzz')runs only off Windows (fuzz=trueon Windows is anerror()).core/tools/meson.build: thevmaf_vplblock carries the gate and prints a disabled message on Windows. An upstream sync that touches these blocks keeps the gates: themsvcismscan reads them to decide which sources the Windows build compiles.scripts/dev/preflight.shresolves its scanners next to itself (PREFLIGHT_DIR) and fails the stage whenfind-posix-only-headers.pyorlint_exceptions.py filtercannot run.- Exceptions:
.config/lint-exceptions.d/msvcism-posix-headers.toml(two files, expiry 2027-06-30). - Reproducer:
bash scripts/ci/tests/test-preflight-msvcism.sh;scripts/dev/preflight.sh --full --stage msvcism.
VMAFx window scores and the window clock¶
rc4/api-wp4-windows, ADR-1852, ADR-2074.
core/src/vmafx/gainswindow.c(windows, the completion thread, the callback thread, the hooks, the context's engine lock,vmafx_context_max_in_flight()) andwindow_clock.c.vmafx_engine_enter()takes the context's engine lock andvmafx_engine_leave()now takes the context (vmafx_engine_leave(context, previous)): a rebase that adds an engine call to a WP2 / WP3 function uses the pair and never nests it.submit.c,register.candcontext.ccall the hooksvmafx_windows_note_index(),_note_flush(),_init(),_pause(),_resume()and_close();VmafxContext(internal.h) gainswindows, created with the context. A rebase that reorders those functions keeps each hook after the engine call it follows, and the pause beforevmaf_engine_close().score.c: the three pooled functions callvmafx_pool_engine(), which the windows call too; keep one pooling path.fence.c's timed wait is exported asvmafx_host_fence_wait()forvmafx_window_wait().core/src/libvmaf.c:VmafContextandstruct ThreadDataBatchgain a frame listener (vmaf_engine_set_frame_listener()), called at the end ofthreaded_extract_batch_func()after the pictures are released; no frame count, submit or flush path moved (WP5'srun_note_frame()calls are untouched).vmaf_engine_score_at_index()isengine_score_at_index(fence = true); newvmaf_engine_try_score_at_index()(no fence, inputs checked withvmaf_predict_inputs_written()),vmaf_engine_feature_written(),vmaf_engine_try_score_at_index_model_collection(),vmaf_engine_thread_count(),vmaf_engine_subsample()andvmaf_engine_max_in_flight().core/src/predict.cgainsvmaf_predict_inputs_written()(collector reads only). An upstream sync ofvmaf_score_at_index()ports intoengine_score_at_index(); a change tobatch_job_take_pictures(), to the thread pool's enqueue capacity or to the device double buffering recomputesvmaf_engine_max_in_flight().- The branch carries master's #2206 (
read_predicted_collection_score()) as a cherry-pick, which a rebase onto master drops as already applied. - No score, golden-data or FFmpeg patch impact: a window's values come from the synchronous pooling;
libvmaf.hbehaviour is unchanged.
Controller tenant spec is the generated type (2026-10-08)¶
refactor/controller-tenant-spec-generated, ADR-2350 D13. TenantSpec, TenantOIDC, TenantRBAC and TenantScoring in cmd/vmafx-controller/auth/tenants.go are aliases of the VmafxTenant types generated into api/vmafx/v1 from api/vmafx-platform.toml; rbac and scoring are pointers there (pointer = true), as the controller's structs were. A rebase that brings back a struct declaration in tenants.go, or a new tenant field outside the definition, fails auth/tenants_generated_type_test.go. no upstream file.
Operator events cluster-wide (2026-10-08)¶
fix/operator-events-rbac, ADR-2647. deploy/helm/vmafx/templates/operator-rbac.yaml renders a third operator role, the ClusterRole <fullname>-operator-events with its binding (create, patch on core events only), and the release-namespace Role no longer lists events. A rebase that touches the template keeps events out of the namespace Role and out of the custom-resource ClusterRole; scripts/ci/tests/test_helm_service_accounts.py fails otherwise. no upstream file.
VMAFx device frames on CUDA (RC4 WP3 CUDA lane)¶
rc4/api-wp3-cuda, ADR-2023, ADR-1929.
- New files
core/src/cuda/vmafx_cuda.h,vmafx_cuda_internal.h,import_device.c,import_frame.c,import_fence.c,import_gl.c,import_pool.c(inlibvmaf_sources, notcuda_static_lib, because the CUDA common objects link into test programs without the VMAFx sources) and the kernelimport_convert.cu(cuda_cu_sources,import_convert_ptx). core/src/vmafx/*.cdispatch to the lane under#ifdef HAVE_CUDA:device.c(create, count, info, describe, unref),device_context.c(engine import;vmafx_context_release_device()frees it after a successful close),frame_import.c(import_on_device(), the release fence),fence.c(CUDA_EVENT and GL_SYNC),frame_import_admit.c(the per-extractor answer),frame_pool.c(CUDA pools),submit.c(the planted early release).internal.hgrowslane/lane_releaseonVmafxDevice,VmafxFrameandlane_stateonVmafxContext, and exports the import layout and plane checks.vmafx_frame_release()calls the lane's release before the release callback.register.cpicks the device twin when a feature is registered on a device context.core/src/picture.h:VmafPicturePrivate.cudagainsvmafxandordered.core/src/libvmaf.c:cuda_order_pictures_against_producer()skips the ADR-1199 barrier for a pair of ordered pictures, andtranslate_picture_device()refuses a VMAFx frame and counts every other download as a host copy. A rebase that touches these keeps both.integer_vif_cuda:VifBufferCudagainsdis_stride; both pitches are set per frame from the pictures invif_submit_scales(), andvif_vert_load_tiles()takes both. An upstream sync offilter1d.cumust keep the second pitch.VmafxFrame.lane_persistent(internal.h): a CUDA pool frame's event behind its readers;vmafx_frame_pool_acquire()waits on it.float_ms_ssim_cudaconverts level 0 on the device: see "float_ms_ssim_cudabuilds level 0 on the device" (master #2282, carried by this stack until its restack).- Definition:
VMAFX_MEMORY_GL_TEXTURE,VMAFX_FENCE_GL_SYNC,VmafxFrameImport.release/user(ABI 0.1.4);VMAFX_MIN_FRAME_IMPORTis the 0.1.2 size (offsetof(VmafxFrameImport, release)).
Custom resources generated from the platform definition (2026-10-08)¶
rc4/api-wp17-crd, ADR-2350 D13. Every file under api/vmafx/v1 except deepcopy_test.go is generated: the types from api/vmafx-platform.toml (scripts/codegen/vmafx-api.py --write), and zz_generated.deepcopy.go, deploy/helm/vmafx/crds/*.yaml and config/rbac/role.yaml by controller-gen (scripts/codegen/crd_generate.py --write). config/crd/bases/, the hand-written zz_generated_deepcopy.go and the per-kind config/rbac/role_*.yaml files are removed. A change that edits a type, CRD or role by hand moves into the definition or an RBAC marker instead; on a conflict in a generated file take either side and run both generators. test_crd_generated_current guards drift, test_crd_compat narrowing within v1, scripts/ci/tests/test_helm_operator_rbac.py the chart's operator rules. no upstream file.
Protobuf files generated from the platform definition (2026-10-08)¶
rc4/api-wp17-proto, ADR-2350 D13. Every file under proto/ is generated: proto/vmafx/v1/vmafx_api.proto from core/api/vmafx.toml, proto/vmafx/v1/vmafx.proto and proto/vmafx/controller/v1/controller.proto from api/vmafx-platform.toml (scripts/codegen/vmafx-api.py --write), and every *.pb.go under gen/go by scripts/codegen/proto_generate.py --write (buf BUF_VERSION of build-config.env, plugins pinned as go.mod tools). A change that edits a proto or a binding by hand moves into the definition instead; on a conflict in a generated file take either side and run both generators. test_vmafx_api_generated_current and test_proto_generated_current guard it, proto_generate.py --breaking-against the wire. no upstream file.
Helm controller on PostgreSQL and its failover E2E case (2026-10-07)¶
rc4/api-wp17-chart, ADR-2350. The controller Deployment takes its replica count, update strategy, volume and store environment from the _helpers.tpl helpers vmafx.controllerStoreBackend, vmafx.controllerStoreEnv, vmafx.controllerDatabaseDSN and vmafx.controllerTopologySpread; a sync that touches controller.yaml, pdb.yaml or networkpolicy.yaml keeps them, and keeps the SQLite store at one replica. controller.store.* holds one key per setting so the generator of the values and schema (ADR-2350 work package 4) reproduces them unchanged. The E2E case 02-controller-ha moves together with .github/workflows/e2e-k8s.yml, test/e2e/kind-cluster.sh and scripts/ci/test_e2e_runtime_contract.py. no upstream file.
Observability: SLO report, usage and cost, capacity (2026-10-08)¶
rc4/obs-6-slo-cost-capacity, ADR-2349, #2430. Fork-only. The rule file, the chart's PrometheusRule template and values block, and the three dashboards are generated (go run ./tools/obsgen -write): regenerate, never hand-merge. The settings series (vmafx:slo_objective, vmafx:slo_events:rate5m, vmafx:slo_bad_events:rate5m, vmafx:price_job_second, vmafx:price_job) keep job and instance labels, because dashboard-linter requires both matchers on every query; a price rule keeps only a positive price.
libvmaf compat library on libvmafx: engine names, split library targets¶
rc4/api-wp6-compat, ADR-1852 decision D3, ADR-2094.
- The library target is now
libvmafx(libvmafx.so.1: the engine and the VMAFx API) plus the compat targetslibvmaf_shared_lib/libvmaf_static_lib(libvmaf.so.3, sourcescore/src/compat/libvmaf/). A sync that adds a source to the oldlibvmaf_sourceslist adds it tolibvmafx_sources; a sync that changes the library's link arguments changeslibvmafx's. - Every engine translation unit is compiled with the generated
core/src/vmafx/engine_names_gen.h(add_project_arguments), which renames the engine's own definition of each libvmaf function tovmaf_engine_<stem>. An upstream change to a libvmaf function body ports into the engine source as is (core/src/libvmaf.c,model.c,picture.c,picture_v2.c,picture_convert.c,dict.cpp,dnn/,mcp/, the HIP / Metal stubs); the 24 bodies incore/src/libvmaf.cthat WP2 already spelledvmaf_engine_<stem>(vmaf_engine_use_featureand others) take the change under that name. Never add a definition of a libvmaf name to the engine and never compile an engine source without the header. A libvmaf function upstream adds needs a[[compat]]entry incore/api/vmafx.toml(the export checks fail otherwise) and a conformance call incore/test/test_compat_conformance_*.c(the coverage rule fails otherwise). - Tests:
link_with : get_option('default_library') == 'both' ? libvmaf.get_static_lib() : libvmafis nowvmaf_test_link(white-box, engine names) andlibvmaf_public_link(black-box, needsvmaf_public_name_argsinc_argsand, for C++,cpp_args). Resolve a conflict incore/test/meson.buildper hunk and keep the variable names. - The WP2 forwarders at the end of
core/src/libvmaf.care gone; theVmafModel/VmafModelCollectionstructs gainapi_owner. vmaf_engine_read_pictures()clears both callerVmafPicturestructs once the context owns the pictures, in every build (a CUDA build left them pointing at released host translations;test_compat_conformancecompares the traces). The body sits inread_pictures_owned(); an upstream change tovmaf_read_pictures()goes there and keeps the clearing in the wrapper.core/include/libvmaf/*.h: every exported declaration carriesVMAF_DEPRECATED("use <vmafx successor>")(empty unlessVMAF_ENABLE_DEPRECATION_WARNINGS);test_libvmaf_deprecationchecks each marker against the definition. An upstream header sync keeps the markers.- No score impact: the golden gate passes through the compat library and
test_compat_conformancecompares every compat function with its engine body; libvmaf return values are unchanged except the differences listed indocs/api/vmafx/index.md. No FFmpeg patch impact: unpatched FFmpeg n9.0.2 builds against the split library and scores identically. - The two libvmaf functions master gained after this branch's base are compat functions:
vmaf_set_sample_range_check_enabled()sets the context optioncheck_sample_range,vmaf_set_input_colorimetry()callsvmafx_context_set_default_color().vmafx_submit()hands every pair's colour to the engine (vmaf_engine_set_pair_colorimetry(), which compares withvmaf_conversion_policy_color_equal()beforevmaf_conversion_state_set_input_color()incore/src/conversion_context.c). An upstream change to the conversion state keeps that comparison: the same colour after the first converted pair is 0, another one-EBUSY.
Observability: Compose example and smoke test (2026-10-08)¶
rc4/obs-5b-compose, ADR-2399, #2430. Fork-only. deploy/grafana/provisioning/datasources/vmafx.yaml is generated (go run ./tools/obsgen -write): on a conflict take either side and regenerate. Its URLs are the service names of deploy/compose/observability/compose.yaml; renaming a service there changes obsgen/datasources.go too. tools/obssmoke reads dashboards through obsgen.DashboardQueries, the parser CheckDashboard uses.
vmafx-controller job backends (2026-10-07)¶
rc4/api-wp17-controller, ADR-2350. The gRPC handlers talk to cmd/vmafx-controller/backend (Backend), never to queue, nodes or scheduler directly; the SQLite queue sits behind backend.Legacy and the PostgreSQL store behind backend.Postgres. The store's generated pgdb/ takes either side on a conflict and is regenerated with python3 scripts/codegen/sqlc_generate.py --write. A sync that touches grpc_server.go keeps the handlers on the interface and the tenant contract (another tenant's job is PERMISSION_DENIED, an unknown one NOT_FOUND) on both backends: replicas_test.go and grpc_tenant_test.go guard it. no upstream file.
The provenance record moves into the library (2026-10-06)¶
rc4/api-wp5-provenance (RC4 work package 5, ADR-2073), on top of rc4/api-wp8-options. Upstream-mirror files touched:
core/src/feature/feature_collector.{h,cpp}:FeatureVectorgainsproducer,producer_optionsandsource, recorded from the thread's producer (vmaf_feature_producer_swap()) when a vector is created. An upstream change to vector creation keeps the call tofeature_vector_record_producer().core/src/feature/feature_extractor.cpp:ProducerScopearoundextract,collectandflush;core/src/libvmaf.cinstalls the producer around the directfex->flush()loop and appends imported and tiny-model scores withvmaf_feature_collector_append_from();core/src/predict.cappends model scores the same way. A new direct call of an extractor's callbacks needs the same scope, or its features have no producer.core/src/model.{h,c}/model_lifetime.c:VmafModelgainssource,sha256,load_flags,overrides; every loader ends invmaf_model_stamp_loaded()andvmaf_model_feature_overload()records the override. Keep both on an upstream sync of the loaders.core/src/output.cpp: the JSON writer ends withjson_write_provenance()(score_format, the record, the backend receipt) and the XML writer withxml_write_provenance().aggregate_metricsno longer ends with a newline of its own. Netflix's harness reads neither element.core/tools/vmaf.cpp: the JSON splice (amend_json_with_backend_receipt,amend_cli_backend_receipt) andcli_format_backend_members()are gone; do not bring them back on a rebase that touches the output path.core/src/libvmaf.c:VmafContext.run(atomics: frames, size, format, first / flush times) is what a provenance query reads, from any thread;run_note_frame()follows eachvmaf->pic_cnt++of the submit paths and the flush paths storerun.flush_ns. A rebase that adds a submit path (RC4 WP4's asynchronous windows) callsrun_note_frame()where it counts a frame; never letvmaf_engine_run_info()readpic_cntorpic_params.core/src/meson.build:vmafx_build_info.h(configure_file) andvmafx_build_commit.h(vcs_tag) feedprovenance_build.c; keep them next to the library target.- Generated:
core/src/vmafx/exactness_gen.c(python3 scripts/codegen/vmafx_exactness.py --write, part ofmake docs-fragments-write) followsscripts/ci/exact_twins.dand the parity gate's tables; a PR that adds a fragment regenerates it with the exact-twins page (make docs-fragments-check,test_vmafx_exactness_table_current). On a conflict take either side and regenerate.
Observability: chart monitoring generated from the rule code (2026-10-07)¶
rc4/obs-5-packaging, ADR-2399. Fork-only. deploy/helm/vmafx/templates/prometheusrule.yaml, deploy/helm/vmafx/files/dashboards/*.json and the block between # BEGIN obsgen monitoring settings and # END obsgen monitoring settings in deploy/helm/vmafx/values.yaml are written by go run ./tools/obsgen -write: on a conflict take either side and regenerate, never hand-merge. TestGeneratedFilesAreCurrent, TestValidateAgreesWithTheChartSchema and scripts/ci/tests/test_helm_observability.py guard them. A rule must match a histogram bucket with obsgen.LeMatcher, never le="<whole number>" (Prometheus 3 stores le="30.0").
Observability: generated alert rules and runbooks (2026-10-07)¶
rc4/obs-3-alerts, ADR-2349, #2430.
deploy/prometheus/vmafx-rules.yamland its promtool testvmafx-rules.test.yamlare generated (pkg/observability/obsgen): on a conflict take either side and rungo run ./tools/obsgen -write.- Every alert needs a page
docs/observability/runbooks/<slug>.md(TestEveryAlertHasARunbook); renaming an alert renames its page and the mkdocs nav entry together. scripts/ci/pinned-tool.shfetches both pinned tools (dashboard-linter, promtool);build-config.envpinsPROMETHEUS_VERSIONand its sha256.- No score, FFmpeg patch or C API impact.
DCO sign-off check (2026-10-08)¶
community-dco, ADR-2462. The job dco-sign-off in .github/workflows/rule-enforcement.yml, its entry 'DCO Sign-off' in the required list of required-aggregator.yml, scripts/ci/check-dco.py with its bot list BOT_LOGINS, and :gitSignOff in renovate.json are fork-authored. A sync that rewrites either workflow keeps the job and the aggregator entry together (check-aggregator-names.sh fails when they diverge). scripts/ci/tests/test_check_dco.py reads the wiring.
Small-PR track in the deliverables gate (2026-10-08)¶
community-smallpr, ADR-2461. scripts/ci/deliverables-check.sh gained section 2c and a skip in the six-item loop; the PR template gained a paragraph. Both are fork-authored; a sync keeps the fork's side. scripts/ci/tests/test_deliverables_small_pr.py guards the behaviour and reads the template and the sentinel guide.
Community health documents (2026-10-08)¶
community-docs, ADR-2461 (Proposed). ACCESSIBILITY.md, CODE_OF_CONDUCT.md, GOVERNANCE.md, SUPPORT.md, CONTRIBUTING.md and the accessibility issue form are fork-authored; no upstream Netflix/vmaf file carries them in this form, so a sync keeps the fork's side of every hunk. CONTRIBUTING.md still ends with the inherited Netflix guide, which a sync may update below the fork's part.
Rust core migration decision (2026-10-08, ADR-2478)¶
docs/rust-core-plan, ADR-2478. Documentation only: ADR, research digest, roadmap section and one core/AGENTS.d page. No code path changes. A sync that ports a Netflix change into a layer with a Rust successor lands it in the C oracle and in the Rust layer in the same PR. no upstream file.
Credits page and gate (2026-10-08)¶
community-credits, ADR-2485. docs/credits.yaml, docs/credits.md, scripts/docs/{credits_lib,credits_checks}.py, generate-credits.py and check-credits.py are fork-authored. An upstream sync that adds a vendored directory, notice file, font or foreign-copyright source file must add its entry to docs/credits.yaml in the same change, or make docs-fragments-check fails; the gate reads REUSE.toml, so keep its annotations in step. Never hand-edit the tables of docs/credits.md.
Option groups generate every scoring surface (2026-10-06)¶
rc4/api-wp8-options (RC4 work package 8, ADR-2044). Fork-only files except core/tools/cli_parse.cpp and core/tools/vmaf.cpp:
core/tools/cli_parse.cppno longer holds the short option string, theARG_*enum,long_opts[]or the usage text: it includes the generatedcore/tools/cli_options.gen.inc. An upstream sync that adds a CLI option adds it to the option groups ofcore/api/vmafx.tomland regenerates (python3 scripts/codegen/vmafx-api.py --write); never edit the.inc. On the rebase onto master, master's--check-sample-range/--check_sample_rangeand--list-backendsjoin the definition the same way, andcore/test/test_cli_option_table.cpp's frozen list gains them.core/tools/vmaf.cpp's JSON receipt appends aprovenancemember (cli_format_provenance_member()of the C filecore/tools/cli_provenance.c, so no C++ CLI source includes the generated C headers); keep it when the receipt code moves (WP5 moves it into the library's report writer).- Generated files (
*.gen.*,proto/vmafx_api.proto,gen/go/**, the marked regions ofdocs/usage/cli.md,docs/usage/ffmpeg.md,docs/mcp/tools.md,docs/server/api-contract.mdandapi/openapi/vmafx-server-v1.yaml): on a conflict take either side, regenerate, thenbuf generate protoand the oapi-codegen command ofgen/go/AGENTS.md. - Both MCP servers read
options.gen.json; do not bring back the hand schemas (scoringExtraProperties,_scoring_extra_properties) or the hand flag lists ofscoreExtras.appendArgs/ScoreExtras.to_argv.
Controller PostgreSQL store (2026-10-08)¶
rc4/api-wp17-store, ADR-2350. New fork-only package cmd/vmafx-controller/store/ (migrations, sqlc queries, generated pgdb/), scripts/codegen/sqlc_generate.py, the SQLC_* pins in build-config.env and the Meson test test_sqlc_generated_current. Generated files take either side on a conflict and are regenerated with python3 scripts/codegen/sqlc_generate.py --write. no upstream file.
Rust model prediction (RC4 lane P, #1723)¶
core/src/predict.csplitsvmaf_predict_score_at_index()into the score gather (unchanged),predict_compute_c()(the former body, statement for statement) andpredict_compute_rust(); an upstream sync keeps the C body inpredict_compute_c()and any arithmetic change tonormalize(),transform(),clip(),post_process_feature_from_another(),piecewise_*orsvm_predict()is mirrored incore/src/rust/predict/src/in the same PR (scripts/ci/rust_twin_diff.py --models).struct VmafModelcarries three trailing fields (rust_predict_state,rust_predict,predict_raw).predict.creaches Rust only through thestruct VmafRustPredictOpstable declared at the end ofcore/src/predict.h;vmaf_model_destroy()(core/src/model_lifetime.c) callsvmaf_rust_predict_destroy()and freespredict_raw. Keep both on a sync that touches those files. No score, public API or FFmpeg patch impact whileVMAF_FEATURE_IMPLis unset.core/src/rust/include/*.hare cbindgen 0.29.4 output, committed byte for byte (scripts/dev/rust-abi-header.sh); the clang-format hooks andmake formatskip that directory. Regenerate them, never format or merge them by hand (scripts/ci/tests/test_rust_abi_header_verbatim.py).
Rust integer ADM twin (core/src/rust/feature/adm/, RC4) (2026-10-07)¶
rc4/adm-twin, ADR-1713.
- The crate
vmafx-fex-admports the scalar path ofcore/src/feature/integer_adm.c,integer_adm.h,integer_adm_kernels.h,adm_csf_fixed_point.h,adm_cm_accumulator.h,adm_angle_flag.h,adm_score.handbarten_csf_tools.hstatement by statement and registers it asadm_rust(ADR-1713). An upstream sync or rebase that changes the arithmetic, the option table or the emitted names ofadmchanges the crate in the same change and re-runsscripts/ci/rust_twin_diff.py --feature admon every fixture (core/src/feature/AGENTS.d/adm-rust-twin.md). A change to a float table ofinteger_adm.horbarten_csf_tools.halso regeneratescore/src/rust/feature/adm/src/tables_c.rs. No score, public C API or FFmpeg patch impact.
Rust motion twin mirrors integer_motion.c (2026-10-07)¶
rc4/motion-twin, #2096.
core/src/rust/feature/motion/src/{sad,window,extractor}.rsportmotion_score_pipeline_8/16,motion_flush_one/vmaf_motion_window_flushandextract()ofcore/src/feature/integer_motion.c(andmotion_blend()) statement by statement. A sync that changes any of them changes the twin in the same PR;scripts/ci/rust_twin_diff.py --feature motionand thesadtable test insad.rs(values from the C pipelines) guard it. No score, public API or FFmpeg patch impact.
Observability: dashboards, GPU exporters and scraped read errors (2026-10-07)¶
rc4/obs-2-dashboards, ADR-2349, #2430.
pkg/observability.RegisterScrapedtakes a read-error counter (NewReadErrors) and aScrapeGroup; a failed read counts under its source and never returns an invalid metric (that fails the whole/metricspage).vmafx-serverandvmafx-noderecord ScoreStream sessions throughinternal/app/scoringservice.StreamMetrics; a sync that touches eitherScoreStreamhandler keepsBegin/defer End(retErr)and the per-frameFramecall.- The dashboards under
deploy/grafana/dashboards/(now seven) are generated: on a conflict take either side and rungo run ./tools/obsgen -write.build-config.envpinsDASHBOARD_LINTER_VERSIONand its sha256. - No score, FFmpeg patch or C API impact.
OpenTelemetry: Base builds the providers (2026-10-07)¶
fix/otel-providers-constructed, fork-only. internal/app/bootstrap.Base ends with fx.Invoke(func(*otel.Providers) {}); without it fx never builds golusoris's OTel providers and no binary exports anything. Keep it on a rebase or a golusoris bump, unless golusoris's otel.Module invokes them itself. TestBase_ConstructsProvidersNobodyRequests guards it.
Praetor pin 7458a220e1c9 and the managed workflows (2026-10-07)¶
chore/praetor-pin-7458a220, ADR-2440. A rebase or sync keeps PRAETOR_REF at 7458a220e1c9... in .github/workflows/standards-gate.yml, the engine's praetor-api.yml and praetor-docs.yml (never hand-edit: audit refuses a changed byte), and an empty push_branch_exceptions in .github/ci-tier.json. A conflict in either workflow file takes the engine's text: regenerate with adopt --force in a throwaway copy, copy back only those two files. test_praetor_managed_jobs_stop_on_a_draft_before_any_work in scripts/ci/tests/test_ci_routing_contract.py guards the draft stop.
CodeQL sweep: exact float compares in tests (2026-10-06)¶
fix/codeql-test-float-compare. Test-only: no rebase impact beyond the files named in changelog.d/fixed/codeql-test-float-bits-sweep.md. Upstream-mirror tests keep their assertions; only the comparison spelling moved to core/test/float_bits.h (ADR-1502).
Model JSON checked out with LF, generator format test pinned to the hook's clang-format (2026-10-08)¶
fix/master-red-format-pin-model-lf, no ADR (bug fixes). .gitattributes adds model/**/*.json text eol=lf after upstream's *.pkl / *.model lines: an upstream sync that touches .gitattributes keeps the fork's line, or Windows builds report other model hashes again (test_praetor_hashed_files_lf.py fails). scripts/codegen/tests/support.py::pinned_clang_format() ties the format test to the clang-format hook's major in .pre-commit-config.yaml; bump the hook rev and requirements/locks/tooling-tests.in together. No upstream file besides .gitattributes.
RC4: speed_chroma Rust twin keeps Netflix's double-form statements¶
core/src/rust/feature/speed/src/isspeed.candvif_tools.cported statement by statement (BSD-2-Clause-Patent). A change tocreate_givens(),update_entropy(),get_speed_score()(ADR-1477's doublesqrt()/log2()form),EIGENVALUE_EPS(a double), the prescale methods, the Gaussian taps orpicture_copy()changes the matching Rust function in the same PR;scripts/ci/rust_twin_diff.py --feature speed_chromaand the golden tests incore/src/rust/feature/speed/src/lib.rsfail on a one-ulp drift. An upstream sync that touchesspeed.cre-runs both. No score, public API or FFmpeg patch impact: the twin is opt-in.
Rust cambi twin follows cambi.c (RC4, ADR-1713)¶
core/src/rust/feature/cambi/is a statement-by-statement port of the scalar path ofcore/src/feature/cambi.c,cambi.h(reciprocal_lut, theupdate_histogram_*/uh_slide*helpers) andluminance_tools.cpp, and must return the C extractor's bits. An upstream sync that changes any of them (option table, init post-processing, preprocessing, spatial mask, mode filter, c-values, quick-select pooling, EOTFs) changes the matching Rust module in the same PR;reciprocal_lutis copied as literal text, never recomputed.core/test/test_rust_cambi_kernels.c(suiterust) holds the table, the TVI / visibility tables, the adjusted window, the mask index and the resize walk to the C functions;scripts/ci/rust_twin_diff.py --feature cambire-checks the scores. No public C API or FFmpeg patch impact.
VMAFx device frames and fences: shared contract on the CPU device (2026-10-07)¶
rc4/api-wp3-common, ADR-1852, ADR-1929.
core/src/vmafx/gainsdevice_context.c,fence.c,frame_import.c,frame_import_admit.c,frame_import_hooks.{c,h}andframe_pool.c;device.cgrows enumeration, information and the checks ofVmafxDeviceDesc's newflagsandexternalfields. The backend lanes (CUDA, SYCL, HIP, Metal) add their device creation, memory kinds, fence kinds and per-extractor admission behind these functions.frame_host.c: the release of a non-pool frame is the sharedvmafx_frame_release()and signals the frame's release fence after the last read;vmafx_frame_read_desc(),vmafx_frame_host_device()andvmafx_frame_bind()are shared with the import and the pool.submit.cchecks admission before the engine counts a frame.context.cdrops the context's device after a successful close, and reads and checksVmafxContextConfig.import_retry_wait_ns(ABI 0.1.3; thevmaf_initcompat glue sets it to 0, the default).error.ckeeps 1023 bytes of subject and message.frame_import.cincludescore/src/metal/iosurface_layout.hfor the NV12 / P010 / P016 row readers. A change to those readers changes the CPU import and the Metal import together (test_vmafx_import_bitexactandtest_metal_iosurface_layout).scripts/codegen/vmafx_api/ctext.pyspells anoutstring parameterconst char **(it printedconst char * *).core/src/compat/gcc/stdatomic.h(the fallback for a compiler without<stdatomic.h>, from dav1d) gainsatomic_uintptr_t,atomic_uint_fast64_t,atomic_store_explicit,atomic_exchange,atomic_compare_exchange_strong/_weakand two memory orders, the C11 atomicscore/src/vmafx/uses. A re-sync of that file from dav1d keeps them.core/test/vmafx_fixture_util.hholds the fixture pairs and the reader bothtest_vmafx_bitexact.candtest_vmafx_import_bitexact.cuse.- No
libvmaf.h, score, golden-data or FFmpeg patch impact.
Source ADR citations: live bindings derived, registry keeps retired and fixtures (2026-10-07)¶
ci/citations-derived, ADR-2200. Fork-only gate. scripts/ci/source-adr-citations.json is schema 2 with retired and fixtures only; on a conflict in it take master's side and keep only the hand-governed records your branch changed (a retired or fixture site count). A live key is an error: delete it, never re-add --write. An upstream sync that cites an ADR number needs no registry edit.
icx-cl and the Windows icpx: strict FP without the override warning (2026-10-07)¶
build/icx-cl-strict-fp-spelling, ADR-2170, Research-2170. Fork-only: the intel-llvm-cl branch of the strict FP policy in core/src/meson.build is /fp:precise /clang:-fno-fast-math /clang:-fcomplex-arithmetic=full /clang:-ffp-contract=off (it was /fp:precise /Qfma-), and the SYCL policy gives sycl_msvc_device_link builds the same -fno-fast-math -fcomplex-arithmetic=full reset as Linux. A sync keeps the order (model first, contraction-off last). no upstream file.
Cloud-native platform decision (2026-10-07, ADR-2350)¶
rc4/api-wp17-adr, ADR-2350. Records the decision and marks ADR-1119, ADR-0711, ADR-1589, ADR-0719, ADR-1526 and ADR-2001 as partially superseded; comments and AGENTS notes of cmd/vmafx-controller/, cmd/vmafx-operator/ and deploy/helm/vmafx/ now call the SQLite queue transitional. No code path changes. no upstream file.
Observability: one metric definition and the generated Overview dashboard (2026-10-07)¶
rc4/obs-1-metric-definitions, ADR-2349, #2430.
- Every Prometheus family is defined in
pkg/observability/metricdef; the services register it throughpkg/observability'sNewCounter,NewGauge,NewHistogramandRegisterScraped. A sync that brings back aprometheus.New*Vec, a GaugeFunc orpromautoin a service bypasses the label bounds and fails the per-binary contract tests. pkg/observability.NewMetricsreturns(*Metrics, error)andSetControllerSourcesis gone (the controller's queue families are incmd/vmafx-controller/metrics.go). The controller queue'sCancelandReportResultreturn(bool, error); the bool feeds the job counters.deploy/grafana/vmafx-overview.jsonmoved todeploy/grafana/dashboards/vmafx-overview.jsonand is generated, as isdocs/observability/metrics.md: on a conflict take either side and rungo run ./tools/obsgen -write, never merge by hand.vmafx-nodecomposesbootstrap.HTTP(nodeServerOptions), listens onVMAFX_HTTP_ADDR(default:9090) and its Dockerfile stages expose 9090.- No score, FFmpeg patch or C API impact.
Rendered docs: pull requests carry fragments only (2026-10-07)¶
ci/render-at-release, ADR-2197. Fork-only tooling. docs/rebase-notes.md has a fragment block at its top (between two marker comments) rendered from docs/rebase-notes.d/; the entries below the block are the history and stay as they are. CHANGELOG.md, docs/adr/README.md, docs/adr/by-tag/, docs/adr/titles.md and docs/research/titles.md are outputs of make docs-render: on a conflict in one, take master's side and re-render, and never add them back to a branch (deliverables-check.sh refuses them). _order.txt is frozen; a rebased branch drops any line it added. .gitattributes no longer lists merge=union for these files.
Praetor pin afb739ed81f3 (2026-10-07)¶
chore/praetor-pin-afb739ed, ADR-2321. Fork-only governance files; no upstream file. Engine output (take master's side on a conflict, then regenerate in a throwaway copy with the pinned engine): PRAETOR_REF, tools/markdownlint/, the DevContainer bundle (six praetor-source.*.b64 parts), .config/agent/hooks/block_evasion.py, .paperclip/harness.json, .paperclip/rules.md and the register.sources digest in .standards.yaml, and the register block of AGENTS.md with the six compiled context files (praetorctl compile-context). The exceptions: block of .standards.yaml is generated: run python3 scripts/ci/praetor_tidy_coverage.py --write. A nested AGENTS.md or AGENTS.d/ page brought in by a rebase or an upstream sync must pass praetorctl caveman check --kind=context; the indexes come from make docs-fragments-write.
Windows: _wsopen_s permission mask (2026-10-07)¶
fix/msvc-wsopen-pmode. no rebase impact: fork-only compat/path_utf8.c; in svm.cpp the vmaf_open_bin_crt() helper (fork edit of the vendored libsvm open call) masks pmode; keep the mask when re-syncing.
DNN session test: invalid pointer literal (2026-10-07)¶
fix/msvc-int-to-ptr. no rebase impact: fork test file, two literals.
MSVC zero warnings: residual sites (2026-10-07)¶
fix/msvc-zero-warnings-residuals. no rebase impact beyond the series notes above: the same conversion-only edits in files the earlier notes already list (adm_tools.c, float_vif.c, x86/motion_avx*.c, the CUDA / HIP adm_decouple_inline helpers, cli_parse.cpp); the shifts are ((int64_t)1 << n) with n below 31.
Tests: MSVC zero-warning conversions (2026-10-07)¶
fix/msvc-zero-warnings-tests. Netflix-mirror tests (test_speed_chroma.c, test_vif_tools.c, test_ciede.c, test_cambi.c, test_barten_csf.c, test_adm_csf_tools_coverage.c, test_float_adm_csf_upstream.c) keep upstream's values; the fork adds f suffixes to float tables and explicit (float) / (int) conversions. On a sync conflict keep upstream's numbers and re-apply the suffix/cast.
Feature sources: MSVC zero-warning conversions (2026-10-07)¶
fix/msvc-zero-warnings-feature. Upstream-mirror files (adm_tools.[ch], adm_csf_tools.h, barten_csf_tools.h, integer_adm*.[ch], integer_vif.c, vif*.c, vif_tools.c, speed*.c, ciede.c, cambi.c, motion_tools.h, iqa/ssim_tools.c, third_party/xiph/psnr_hvs.c, the x86/ and arm64/ twins) gain explicit (float) / (int) / (double) conversions where the compiler converted implicitly, and f on float table literals. On a sync conflict keep upstream's expression and re-apply the cast on the statement the conversion belongs to; two forms need care: x += double is x = (float)(x + double) (never x += (float)double, which rounds twice), and speed*.c's entropy update keeps the increment in a double so the load of entropy[i] stays after the log2() calls. The twin contract tests (test_*_exact_contract.py, test_sycl_vif_float_sums_contract.py) quote the new statements. A statement-for-statement mirror in a GPU twin needs no change: the value is the same.
MSVC zero warnings: CRT calls, pragmas, declarations (2026-10-07)¶
fix/msvc-zero-warnings-crt. Upstream-mirror files touched: pdjson.c (push / pop renamed json_push / json_pop, with the matching core/test/meson.build symbol list), adm_tools.c and barten_csf_tools.h (the M_PI fallback now spells UCRT's own literal so a second definition is an identical redefinition), model.c, feature_name.cpp, cli_parse.cpp, y4m_input.c, yuv_input.c (bounded memcpy, VMAF_SSCANF, explicit conversions). pelorus_qp_report_csv.c is a vendored file: the second local edit (_wfsopen) must be in pelorus before the next scripts/sync-pelorus-interop.sh, or the C4996 comes back. A sync that brings upstream's strncpy / sscanf / getenv back keeps the fork's crt_portable.h spelling.
Zero warnings: Metal links and the spill probe (2026-10-07)¶
ci/zero-warnings-metal-and-probe, ADR-2170. Fork-only: core/src/metal/meson.build (project link argument after add_languages('objcpp')), core/src/sycl/run_captured.py and the sycl_quiet_launcher of sycl_common_* in core/src/meson.build. no upstream file.
Warnings are errors on the MSVC legs (2026-10-07)¶
ci/msvc-werror-gate, ADR-2170. Fork-only: the msvc mode of scripts/ci/werror-args.sh, its cases and the MSVC leg contract in scripts/ci/tests/test_werror_args.py, the werror: msvc row key, the Warnings-as-errors arguments bash step (id: werror) and ${{ steps.werror.outputs.args }} on the cmd configure lines of libvmaf-build-matrix.yml (windows-gpu-build, windows-arm64) and build.yml (build-work). no upstream file.
Warnings are errors per leg (2026-10-07)¶
ci/warnings-are-errors-per-leg, ADR-2170. Fork-only: scripts/ci/werror-args.sh, its test, and the werror row key and $(scripts/ci/werror-args.sh ...) call in the fork's workflows (libvmaf-build-matrix.yml, sanitizers.yml, go-ci.yml, rust-ci.yml, ffmpeg-integration.yml). core/src/meson.build gained nvcc_werror_flags and hip_werror_args, both empty unless -Dwerror=true. no upstream file.
Zero warnings: icx and clang-cl driver flags (2026-10-07)¶
build/zero-warnings-driver-flags, ADR-2170. core/src/meson.build is a fork file: the icx strict line gained -fno-fast-math -fcomplex-arithmetic=full before -ffp-contract=off (see the page core/AGENTS.d/strict-fp-compiler-args.md), and -pedantic / -fvisibility=* are offered only to non-MSVC-syntax drivers. no upstream file.
Zero warnings: unused code, attributes, deprecated calls (2026-10-07)¶
fix/zero-warnings-unused-and-attributes, ADR-2170. Upstream-mirror files touched: core/include/libvmaf/macros.h (VMAF_EXPORT is empty when __MINGW32__ is defined and the compiler is not clang; keep the fork's branch order: MSVC, MinGW GCC, GNU / clang, empty) and core/src/dnn/meson.build (include_type: 'system' on the ONNX Runtime dependency). A sync keeps both. The Win32 pthread shim names VMAF_W32_CALLBACK / VMAF_W32_STDCALL instead of CALLBACK / __stdcall.
Zero warnings: initialisers, tags, fallthrough (2026-10-07)¶
fix/zero-warnings-initialisers, ADR-2170. Upstream-mirror files touched: core/src/feature/integer_motion.c (option terminator {0}), core/src/feature/ssimulacra2.c ([[fallthrough]]; in yuv_matrix_coeffs()), vendored core/src/mcp/3rdparty/cJSON/cJSON.c (true / false are not redefined when <stdbool.h> leaves them keywords; hunk 7 of its AGENTS.md). On a sync keep the fork's side of each hunk. The Metal .mm option tables list their designators in the declaration order of VmafOption (name, help, alias, offset, type, default_val, min, max, flags) and of VmafFeatureExtractor; a rebase that brings a table from a branch keeps that order. The device-free Metal contract tests accept {} as the terminator.
CI tiers: one definition, a tier job in every pull-request workflow (2026-10-07)¶
ci/fewer-runs, ADR-2169. Fork-only CI; no upstream file. Every workflow with a pull_request trigger starts with a tier job (a call of .github/workflows/ci-tier.yml) and its other jobs need it. On a conflict in a workflow keep both: master's change to the job and the needs: tier / if: needs.tier.outputs.<light|full> == 'true' pair of this branch (a planner gate is always() && needs.tier.outputs.<tier> == 'true'). A job added since belongs to the light or the full tier: .github/ci-tier.json (full_only, always) and test_ci_routing_contract.py say which. renovate.json: the catch-all group is the first packageRules entry (the later rules win); keep it first.
Metal headers: the host double comparison uses compiler builtins (2026-10-06)¶
fix/metal-f64-equal-no-libimf. no rebase impact: fork-only Metal header (core/src/feature/metal/metal_portable.h); no upstream file.
Praetor pin 04cc813ff054 (2026-10-06)¶
chore/praetor-pin-04cc813, ADR-2153. Fork-only governance files. PRAETOR_REF in .github/workflows/standards-gate.yml, tools/markdownlint/verify.mjs and the DevContainer bundle are engine output: take master's side on a conflict and regenerate with adopt --force --lock-source-root <praetor source at the pin> in a throwaway copy. The exceptions: block of .standards.yaml and .config/clang-tidy/measured-sources.txt are generated: after a conflict run python3 scripts/ci/praetor_tidy_coverage.py --write.
CI: the push aggregator reads its own branch's runs (2026-10-06)¶
fix/aggregator-own-branch-runs. no rebase impact: the fork's aggregator workflow and its test harness; no upstream file.
vmaf-tune: the backend probe skips the PATH lookup with a runner (2026-10-06)¶
fix/vmaftune-probe-runner-no-path. no rebase impact: one fork-only vmaf-tune function (backend_report()) and its test; no upstream file, build or public surface.
ADM twins: exact scale-0 angle flag, per-kernel register budget (2026-10-06)¶
fix/adm-angle-flag-s0-int64, ADR-2134. Keep the unsigned-sum form in decouple_angle_flag_s0() of cuda/integer_adm/adm_decouple_inline.cuh and hip/integer_adm/adm_decouple_inline.hip; Netflix has no GPU twin, so a sync has no counterpart. KERNEL_BUDGETS in core/test/test_cuda_adm_cm_register_pressure.py holds the one kernel above 208.
CI: the hosted cpu tidy lane runs the Makefile targets (2026-10-06)¶
fix/tidy-ratchet-unmeasured-baseline-files. The Tidy Ratchet job (.github/workflows/lint-and-format.yml) and nightly clang-tidy-full run make tidy-ratchet-build LANE=cpu and make tidy-ratchet LANE=cpu and install libvpl-dev; a workflow sync must not bring back a job-local meson setup or direct tidy-ratchet.py call (scripts/ci/tests/test_tidy_lane_container.py). tidy-ratchet.py exits 4 on a baseline translation unit it did not measure. No upstream file.
Scorecard single-maintainer exceptions (2026-10-06)¶
ADR-2126. Two files under .config/lint-exceptions.d/ (scorecard-code-review.toml, scorecard-branch-protection.toml). Fork-only; no upstream counterpart. A rebase keeps both entries and their expiry.
Code-scanning sweep: SBOM wheel unpack, _set_seed (2026-10-06)¶
fix/code-scanning-pip-hash-and-seed. The Prepare the SBOM root steps of release-dry-run.yml and supply-chain.yml unpack the built vmaf_mcp wheel with python -m zipfile -e; keep that, because pip install of an unhashed local wheel is a Scorecard Pinned-Dependencies finding and the repository's own lock check forbids a generated requirements file. _set_seed() in ai/src/vmaf_train/predictor_train.py uses find_spec(). Both are fork-only.
CodeQL sweep: Metal headers, include guards, Pelorus test (2026-10-06)¶
fix/codeql-metal-headers-guards. core/src/feature/ssim.h and ms_ssim.h (Netflix files) gained SSIM_H_ / MS_SSIM_H_ guards in the style of motion.h: an upstream sync that rewrites either file keeps the guard. vmaf_mtl_fm_blur() takes const VMAF_MTL_FM_THR VmafMtlFmWindow * (Metal-only fork code). core/test/meson.build renames two static helpers of the vendored Pelorus parser for test_pelorus_interop with c_args; keep them when the mirror is re-vendored (the names must stay private to that executable).
ADM twins: host test of the device decouple header (2026-10-06)¶
fix/codeql-adm-twin-header-tests, T-GPU-ADM-ANGLE-FLAG-S0-INT32-CORNER-2026-10-06. Keep the int64 sums in iadm_angle_flag_s0() (metal/integer_adm.metal); an upstream sync has no counterpart (Netflix has no GPU twin). core/test/meson.build renames run_tests per twin for test_adm_decouple_recip_*: keep the c_args and cpp_args pair together. The CUDA and HIP decouple_angle_flag_s0() stay int32 until the register budget is settled (see the state row).
FFmpeg patch 0022: input colorimetry from the AVFrame (2026-10-06)¶
port/ffmpeg-input-colorimetry, ADR-2093. Fork-only patch, appended to the series: no upstream FFmpeg counterpart. It adds vmaf_color_from_frame() and vmaf_declare_input_color() before do_vmaf(), the color_set member of LIBVMAFContext, and one call each in do_vmaf() and the software branch of do_vmaf_sycl(). A refresh onto a new FFmpeg release must keep both call sites; a libvmaf without vmaf_set_input_colorimetry() (Netflix's) does not link it. Test: core/test/test_ffmpeg_libvmaf_input_colorimetry_contract.py, ffmpeg-patches/test/check-libvmaf-input-colorimetry.sh.
Go test: the controller version test keeps the loader path (2026-10-06)¶
fix/go-controller-version-test-loader-path. no rebase impact: one fork-only Go test (cmd/vmafx-controller/version_flag_test.go); no upstream file, build or public surface.
Go CI: one gosec definition, G703 on the VMAF_BIN lookup (2026-10-06)¶
fix/go-vmaftest-gosec-g703. The gosec flags live only in the Makefile's lint-go, and the gosec (exclude generated) step of .github/workflows/go-ci.yml runs make lint-go; a workflow edit that inlines the command again forks the gate. internal/vmaftest/vmaftest.go keeps its #nosec G703 with the reason. No upstream file is involved.
Port of Netflix/vmaf ed61076b2, 1ddf81607, a6c0ba6d5, 130569c45, efe90c8b8, 5c3f4fb90, 4f3f71b68: HDR-VMAF groundwork (2026-10-06)¶
port/upstream-hdr-groundwork, ADR-2093. Netflix PRs #1671 to #1675, #1677 and #1678, one batch because each depends on the one before.
- No
VmafPicture::color. Upstream'svmaf.cwritespic->color = colorandlibvmaf.c/conversion_policy.creadpic->color; the fork hasvmaf_set_input_colorimetry()(ADR-2093) and passes the colour as an argument (vmaf_conversion_policy_target(ref, ref_color, dist, dist_color, ...)). A sync of those hunks keeps the fork's side;ed61076b2'sfetch_picture()change is not applicable (init_cli_context()calls the setter). libvmaf.cglue lives incore/src/conversion_context.c. Upstream'sconvertmember,convert_picture(),convert_pictures()andregister_conversion_target()are in that file;libvmaf.chasVmafConversionState convert, the register call invmaf_use_features_from_model(), the convert call at the top ofvmaf_read_pictures()(failure releases the pictures, ADR-1431) and the close invmaf_commit_remaining_owners(). A device picture is refused with-ENOTSUP, and so isvmaf_read_pictures_sycl()when a model declares a target (vmaf_conversion_state_refuse_zero_copy()), since no conversion reaches that path.- Files.
libvmaf/tools/cli_parse.c/vmaf.carecore/tools/cli_parse.cpp/vmaf.cpp(the flag handlers are split per attribute, HISS-04). The model parser has a C twin (read_json_model.c, compiled by the fuzz harness) and the built C++ file (read_json_model.cpp);conversion_targetis in both. Thezimghunk of5c3f4fb90is incore/src/picture_convert.c, notpicture.c. The Python harness iscompat/python-vmaf/(color_ref/color_distfollowbackendincall_vmafexec(), so the existing positional order is kept). - CAMBI twins.
4f3f71b68changescambi.conly; the fork changes the same two minimums in the CUDA, HIP, SYCL and Metal option tables. Keep them equal on a sync. - Tests.
test_cli_parse.c(colour tests asrun_color_tests()),test_model.c(the hunk merges into the fork's tables),test_conversion_policy.c(aTaggedPicholds the colour),test_read_pictures_convert.c(colour set through the setter; no unref after a failed read) are upstream's, adapted as noted. The missing-colour log line names--color_range_ref/_distetc.; upstream's names flags that do not exist. - Fixtures. The two dock clips are in
scripts/test/fetch-test-yuvs.shwith md5 sums (Netflix/vmaf_resource5e7b853ba).
Meson secret-env contract test spells paths with forward slashes (2026-10-06)¶
fix/meson-secret-env-test-posix-paths. ADR-1333 entry preserved: only the path spelling in core/test/test_meson_secret_env_sanitization.py changed; the runner, the setup and the credential inventory are untouched. no other rebase impact.
CLI exit status is the libvmaf code modulo 256 on every platform (2026-10-06)¶
fix/cli-exit-status-modulo-256. A sync or refactor of vmaf_cli_main() in core/tools/vmaf.cpp keeps the return of the run result through vmaf_cli_exit_status() (core/tools/cli_exit_status.h); Netflix's main() returns the raw code. core/test/test_cli_exit_status_contract.py and test_cli_exit_status guard it.
CI: prune-corrupt-fixtures runs on bash 3.2 (2026-10-06)¶
fix/prune-fixtures-bash32. no rebase impact: one CI helper script and its fixture test; no source, build or public surface changes.
ADR audit: status header forms (2026-10-06)¶
docs/adr-audit-mechanical. no rebase impact: ADR status headers and dated status updates, plus one ADR number in docs/state.md; no source, build or test file changes.
ADR audit: Proposed ADR statuses (2026-10-06)¶
docs/adr-audit-status. no rebase impact: ADR status lines, dated status updates, one drift-gate exception list and one docs/state.md Deferred row; no source, build or test file changes.
ADR audit: partial supersession in status lines (2026-10-06)¶
docs/adr-audit-partial. no rebase impact: 45 ADR status lines; no source, build or test file changes.
ADR audit: unfilled template blocks and tombstone records (2026-10-06)¶
docs/adr-audit-backfill. no rebase impact: four ADR files lose a template block, four short ADR files are added; no source, build or test file changes.
ADR audit: errata blocks (2026-10-06)¶
docs/adr-audit-errata. no rebase impact: appended errata sections in ADR files; no source, build or test file changes.
SYCL fused VIF reads and writes different downsampled planes (2026-10-06)¶
fix/sycl-vif-fused-rd-pingpong. no rebase impact: fork-only SYCL twin (core/src/feature/sycl/integer_vif_sycl.cpp) and test (core/test/test_sycl_vif_parity.c). A sync or refactor of the fused path keeps vif_rd_output(): a fused scale never writes the planes it reads; scale 1 writes the second pair (d_rd_ref_alt / d_rd_dis_alt, fused mode only).
Release scope of 1.0.0 and the roadmap to 2.0 (2026-10-06)¶
docs/rc3-adr-roadmap-2026-10-06. no rebase impact: ADR-2001, docs, changelog fragment and the candidate-map paragraph of AGENTS.md section 11 with its six compiled projections (edited by the same substitutions, as ADR-1868 and ADR-1880 did). A sync that touches AGENTS.md keeps the fork's section 11 and recompiles the projections from it.
Lefthook and the pre-commit framework run on Windows hosts (2026-10-06)¶
fix/hooks-windows-host, ADR-2012, T-HOOKS-WINDOWS-LEFTHOOK-QUOTING-2026-09-30, T-LEFTHOOK-UNINSTALL-REWRITES-AGENT-HOOK-FILES-2026-09-30.
lefthook.yml:framework-hooksinpre-commitandpre-pushcallscripts/git-hooks/framework-hooks.shfrom a one-line, quote-freerun:. Keep everyrun:inlefthook.ymlon one line and free of double quotes: lefthook passes them tosh -cunescaped on Windows, so a double quote ends the script. Guarded byLefthookBridgeTestsinscripts/githooks/tests/test_install.py.scripts/githooks/install.py: leaves hooks whose shim callscall_lefthook runin place. Install order:lefthook install, thenmake install-hooks..claude/settings.json,.codex/hooks.json: keep sorted keys and two-space indent (json.dumps(..., indent=2, sort_keys=True)plus newline), the formlefthook uninstallwrites them to.requirements/locks/pre-commit.txt:reuse[charset-normalizer]==6.2.0; reuse skipspython-magicon Windows.cmd/vmafx-node/bpf/://go:build linuxon the tracepoint-dependent loader and test files, sogovulncheck ./...andgo vet ./...load the package on Windows.- Fork-only CI and hook files:
scripts/ci/check-container-image-references.py(uses POSIX path spelling for comparison),scripts/ci/tests/test-dedupe-gate.shandtest_envtest_single_source.py(skip POSIX Makefile/shebang parts on Windows),scripts/ci/tests/test_research_digest_ids.pyandscripts/docs/generate-adr-by-tag.sh(LF newlines). No upstream-mirror file changed.
Licence provenance of the Metal integer ADM host files (2026-10-06)¶
fix/master-red-licence-provenance. scripts/dev/relicense_provenance.toml gains two [ports] entries and one [not_ports] line; integer_adm_metal_host.c / .h take the Netflix notice and the dual tag, .config/lint-exceptions.d/spdx.toml its header. An upstream sync that touches integer_adm.c keeps the entries. No other rebase impact.
Cppcheck on the Metal host tests (2026-10-06)¶
fix/master-red-cppcheck-metal. metal_float_motion_math.h and metal_float_vif_math.h carry a cppcheck-suppress-begin / -end passedByValue block (ADR-1498: the headers are shared with MSL); a sync must keep the block and the {0} initialiser in vmaf_mtl_fvif_statistic_args(). Six test files changed in place. No upstream-mirror file; no other rebase impact.
CI gates read the right inputs (2026-10-06)¶
fix/master-red-ci-gates. Fork-only CI and test files: core/test/test_meson_secret_env_sanitization.py (EXTERNAL_CHECKOUT_ROOTS), three workflow files (docs.yml, lint-and-format.yml, tests-and-quality-gates.yml), scripts/ci/tests/test-default-model-single-source.sh and a new scripts/ci/tests/test_tidy_changed_exclusions.py. No upstream-mirror file changed; no rebase impact beyond keeping the .ci/ skip and the three exclude_untidyable() entries.
go fix and cargo fmt applied (2026-10-06)¶
fix/master-red-go-rust-fmt. Formatting and fixer output only, in Go files under cmd/ and pkg/ and one Rust example; no upstream-mirror file. no rebase impact.
CI fixture cache: tracked fixtures put back after the restore (2026-10-06)¶
fix/fixture-cache-tracked-files. no rebase impact: fork-only CI files. scripts/ci/prune-corrupt-fixtures.sh takes --restore-tracked, and the fixture-restore steps of build.yml, libvmaf-build-matrix.yml and tests-and-quality-gates.yml pass it; a workflow edit that moves or copies the restore keeps the flag on the step after it, or a restore-keys hit brings back an older revision of a tracked python/test/resource file.
test_icx_system_libm runs where os has no confstr (2026-10-06)¶
fix/icx-libm-test-no-confstr. no rebase impact: fork-only test file (core/test/test_icx_system_libm.py); its os.confstr patches keep create=True so the Windows legs can run them.
HIP smoke test follows vmaf_hip_context_new()'s device contract (2026-10-06)¶
fix/hip-smoke-context-no-device. no rebase impact: fork-only test (core/test/test_hip_smoke.c); the context case branches on vmaf_hip_device_count() like the state case.
Metal IOSurface import self-test releases its fixtures (2026-10-06)¶
fix/metal-iosurface-selftest-leak. no rebase impact: fork-only test (core/test/test_metal_iosurface_import_parity.c); both builds of imported_psnr() consume the planar pair they are given.
test_adm_decouple_recip builds and runs on Windows (2026-10-06)¶
fix/adm-decouple-recip-test-windows. no rebase impact: fork-only test files and CI check. A host C or C++ file that names __builtin_clz keeps #include "feature/compat_builtin.h" (scripts/ci/check-msvc-clz-shim.sh enforces it), and a test helper fills div_lookup once: on _WIN32 div_lookup_generator() refills the table on every call.
CPU extractor close callbacks are close_fex (2026-10-06)¶
fix/darwin-lto-static-close. Upstream-mirror files touched: ciede.c, float_adm.c, float_moment.c, float_motion.c, float_ms_ssim.c, float_psnr.c, float_ssim.c, float_vif.c, integer_adm.c, integer_ssim.c, integer_vif.c, speed.c, ssimulacra2.c (and the fork's brisque.c, delta_e_itp.c, niqe.c, pu21.c): the static int close(VmafFeatureExtractor *fex) callback and its .close = initialiser are named close_fex. A sync that brings an upstream change to one of these functions keeps the fork's name; a conflict on the definition line or the initialiser takes the fork's side. Upstream's name collides with the C library's labelled close() in a macOS full-LTO link. core/test/test_libc_named_internal_functions.py fails if a static close comes back. See core/src/feature/AGENTS.d/libc-named-statics.md.
Heavy fast-suite tests sized for the sanitizer jobs (2026-10-06)¶
fix/sanitizer-heavy-test-timeouts. no rebase impact: fork-only test files. test_integer_psnr_coverage carries timeout : 480 for its 5e9-sample APSNR wrap case, and test_metal_psnr_hvs_math forms each masking table's terms once for both summations.
Metal host files balance their anonymous namespaces (2026-10-06)¶
fix/metal-float-motion-anon-namespace. no rebase impact: fork-only Metal host code (core/src/feature/metal/float_motion_metal.mm) and a new device-free test, core/test/test_metal_host_source_balance.py.
SYCL twin option cases proven on a device (2026-10-06)¶
test/rc3-sycl-twin-option-regression. no rebase impact: ledger, changelog fragment and this note only; no source or test file changed.
SYCL device AddressSanitizer option (2026-10-06)¶
test/rc3-sycl-device-sanitizer. Fork-only: core/meson_options.txt gains sycl_device_asan and core/src/meson.build defines sycl_asan_args between the SYCL toolchain selection and the MSVC device-link block, appends it to sycl_toolchain_args and sycl_link_args, and to the per-translation-unit AOT override (tu_toolchain_args). Upstream Netflix/vmaf has no SYCL build, so a sync cannot conflict; a rebase onto master keeps the three appends together with core/test/test_sycl_device_asan_option_contract.py. See ADR-1930.
float_vif and SpEED refuse a prescaled plane past the int index (2026-10-05)¶
fix/prescaled-plane-int-index-limit. The fork adds vif_plane_fits_int_index() to the upstream-mirror core/src/feature/vif_tools.h and calls it in float_vif.c::init_scaled_plane() (the scaled-size checks moved out of init() with it, to keep init() under 60 lines), in speed.c::speed_init_dimensions() (after the too-small check) and in speed_internal.c::speed_internal_init_dimensions(). Upstream Netflix/vmaf has no such check and indexes vif_tools.c with int, so a sync keeps the three calls and the helper; if upstream ever widens the indices of vif_tools.c to a 64-bit type, the check can go. core/test/test_prescaled_plane_int_index.c fails without it. See core/src/feature/AGENTS.d/float-vif.md.
Integer ADM: named refusal when viewing geometry is below fixed-point floor (2026-10-06)¶
fix/adm-viewing-floor-named-refusal. core/src/feature/adm_csf_fixed_point.h adds adm_viewing_geometry_check() which logs a named refusal at ERROR when adm_norm_view_dist * adm_ref_display_height < 3240 (the 3240 floor, 1080p at 3H) naming the extractor, parameter values, product, and pointing to float_adm. The CPU extractor (core/src/feature/integer_adm.c) calls it from extract() with "adm"; GPU twins (adm_cuda, adm_hip, adm_sycl, adm_metal) call it during init(). An upstream sync that touches integer ADM option validation or CSF setup must preserve the named refusal and the contract where CPU fails in extract() and GPU twins fail in init().
Opt-in sample range check of vmaf_read_pictures() (2026-10-06)¶
feat/sample-range-check (ADR-1918). Fork-only files: core/src/picture_sample_range.{c,h}, core/test/test_sample_range_check.c, docs/api/sample-range.md. Fork hunks in upstream-mirror files: VmafContext::check_sample_range, vmaf_set_sample_range_check_enabled() and the call in read_pictures_validate_and_prep() in core/src/libvmaf.c; the contract paragraph and the declaration in core/include/libvmaf/libvmaf.h; --check-sample-range in core/tools/cli_parse.cpp / cli_parse.h and the setter call in init_cli_context() (core/tools/vmaf.cpp). An upstream sync that touches vmaf_read_pictures() keeps the check after validate_pic_params() and before any extractor. See core/AGENTS.d/sample-range-check.md.
CUDA warp reductions reached by every lane, on unsigned words (2026-10-06)¶
fix/cuda-warp-reduce-defined. Upstream Netflix/vmaf calls the integer VIF horizontal flush (warp_reduce() of the seven accumulators) inside if (y < h && x_start < w) in cuda/integer_vif/filter1d.cu, and builds warp_reduce(int64_t) in cuda_helper.cuh from two shuffled halves with (x >> 32) << 32. The fork calls vif_hori_flush_accums() after the branch (every lane of a warp reaches the full-mask shuffles; the lanes past the edge add zeros) and adds the 64-bit words as unsigned values through warp_reduce_u64(). A sync that touches either file keeps the fork's form. Outputs are bit-identical (integer sums). See docs/development/rebase-sensitive-invariants.md and core/src/feature/cuda/AGENTS.d/vif.md.
Integer ADM: scale-0 contrast-masking rows summed unsigned; GPU gain product bounded before narrowing (2026-10-05)¶
fix/adm-cm-row-total-unsigned. Upstream Netflix/vmaf sums every contrast-masking row of integer_adm.c in int64_t; the fork sums the scale-0 rows and frame unsigned (adm_cm_round_row_total_s0() in core/src/feature/adm_cm_accumulator.h, adm_cm_fold_s0(), uint64_t in adm_cm_accum_px(), adm_cm_row(), AdmCmRowFn, adm_cm_rows(), adm_cm_result(), cm_row_avx2() / cm_row_avx512()). An upstream sync that touches the scale-0 masking loops keeps the unsigned row and frame: a scale-0 row passes INT64_MAX (core/test/adm_cm_row_overflow_frame.h). Scales 1-3 keep upstream's signed sums. The CUDA and HIP decouple_r_s123() bound the gain product in double before narrowing it, as the CPU's adm_decouple_band_s123() does; do not restore (int32_t)(...) * adm_enhn_gain_limit. See core/src/feature/AGENTS.d/adm-rounding.md.
APSNR clip squared error summed in 128 bits (2026-10-05)¶
fix/apsnr-clip-sse-128. Upstream Netflix/vmaf's integer_psnr.c keeps apsnr.sse[] as uint64_t and adds s->apsnr.sse[p] += sse;; the fork holds it as VmafPsnrClipSse (core/src/feature/psnr_score.h) and adds through vmaf_psnr_clip_sse_add(), because the clip sum wraps past 2^64 on long 12- and 16-bit clips. An upstream sync that touches psnr() / psnr_hbd() or flush() keeps the fork's form. See core/src/feature/AGENTS.d/psnr.md.
Preserve explicit .int8.onnx paths in DNN session open (2026-10-05)¶
Fork-only: core/src/dnn/dnn_api.c (resolve_load_path) gains the kInt8Suffix early return matching core/src/dnn/dnn_attach_api.c:75. On an upstream sync that touches this file, preserve the kInt8Suffix check to avoid deriving <name>.int8.int8.onnx.
VPL decode retry ceiling contract and warning frame drop repair (2026-10-05)¶
The vmaf_vpl decode retry loop is formally verified under the derived 60,000-attempt bound (VPL_DECODE_MAX_ATTEMPTS). Physical Intel Arc A380 hardware (/dev/dri/renderD129) confirmed zero ceiling exhaustion or hangs across 48-frame baseline and long-GOP streams. Extracted core/tools/vmaf_vpl_core.h and .c to decouple status classification and frame loop execution. Fixed a correctness bug where warning codes with valid synchronization points (sts > 0 && sync != NULL, such as MFX_WRN_VIDEO_PARAM_CHANGED) were previously dropped. Added transient retry for MFX_WRN_ALLOC_TIMEOUT_EXPIRED. An 8-test deterministic device-free unit test suite (test_vmaf_vpl_decode_ceiling.c) in Meson fast verifies finite busy recovery, ceiling sensitivity, exact 60,000 attempt exhaustion, multi-frame ordering, warning publication, and hard error fail-fast without GPU hardware. Added automated hardware smoke test test_vmaf_vpl_hardware_smoke.sh under slow / gpu suites.
- Research digest: Research-1900.
- Decision matrix: ADR-1900.
- AGENTS.md invariant:
core/tools/AGENTS.d/vmaf-vpl.md, "VPL decode retry ceiling contract". - Reproducer / smoke:
meson test -C build test_vmaf_vpl_decode_ceilingandmeson test -C build test_vmaf_vpl_hardware_smoke. - Changelog:
changelog.d/fixed/vpl-decode-ceiling-contract.md. - FFmpeg impact: none; no public C header, exported libvmaf API, CLI flag, or FFmpeg patch touched.
Integer accumulator bounds; exact-twin matrix at 8K and 16K (2026-10-05)¶
fix/accumulator-bounds-audit. Fork-only files: docs/development/accumulator-bounds*.md, scripts/dev/adm_cm_row_bound.py, core/test/test_accumulator_bounds_16k.c and its block in core/test/meson.build, the --grid option of scripts/ci/exact_twin_matrix.py and the 8K / 16K blocks of docs/development/exact-twin-matrix.md.
- An upstream sync that changes an accumulator's type, its term or how many terms reach it changes that row of the accumulator-bounds pages in the same PR.
- A new exact twin needs an 8K row per backend and a 16K CPU row (
exact_twin_matrix.py --grid 8k/--grid 16k --record), ortest_exact_twin_matrix_contractfails.
MCP tool contract shared by both servers (2026-10-05)¶
No rebase impact: both MCP servers are fork-only. Keep mcp-server/vmaf-mcp/tool-contract.json generated (python3 -m vmaf_mcp.tool_contract --write); on a conflict in it, take either side and regenerate. cmd/vmafx-mcp/tool_contract_test.go replaces the hand-copied lists that used to be in server_test.go; do not bring them back.
vmaf --list-backends and the score-backend selectors (2026-10-05)¶
Fork-only: core/tools/cli_backends.cpp is new and cli_parse.cpp / vmaf.cpp / cli_parse.h gain the --list-backends option (ARG_LIST_BACKENDS, CLISettings.list_backends, the early return after the getopt loop, the branch in vmaf_cli_main()). On an upstream sync that touches these files keep all four; upstream has no such option. See core/tools/AGENTS.d/list-backends.md.
v1-model test fixtures named by file (2026-10-05)¶
No rebase impact: wording and fixture labels of a fork-only test.
Known upstream GPU defects recorded (2026-10-05)¶
No rebase impact: docs only.
Whole-model no-fallback test for the v1 models (2026-10-05)¶
test/rc3-v1-models-no-fallback (issue #2144). Fork-only files: core/test/test_gpu_v1_models_no_fallback.py, its block in core/test/meson.build, a note in core/src/feature/AGENTS.d/model-options.md.
- An upstream sync that adds an option to a v1 model's
feature_opts_dictsmust add it to the CUDA, SYCL and HIP twin tables in the same PR, ortest_<backend>_v1_models_no_fallbackfails with the extractor named. - The test's known fallback is
float_adm=adm_csf_mode=1. It moves to another default-only option when a twin implementsadm_csf_mode.
Exact-twin depth and layout matrix (2026-10-05)¶
test/rc3-exact-twin-matrix. Fork-only files: scripts/ci/exact_twin_matrix.py, core/test/test_exact_twin_matrix_contract.py, their blocks in core/test/meson.build, docs/development/exact-twin-matrix.md, a note in scripts/ci/AGENTS.d/parity-exact-twins.md.
- A new fragment in
scripts/ci/exact_twins.d/for CUDA, SYCL or HIP needs a recorded matrix row in the same PR (exact_twin_matrix.py --record), ortest_exact_twin_matrix_contractfails. docs/development/exact-twin-matrix.mdis measurement output: on a conflict in a backend's block, take either side and re-record that backend on its device.- The fixture bytes are pinned in the contract test; a change to the generator re-pins them and re-records every backend.
GPU byte-stride contract (2026-10-05)¶
test/rc3-gpu-byte-stride-contract. Fork-only files: core/test/test_gpu_byte_stride_contract.py, its block in core/test/meson.build, core/src/feature/AGENTS.d/gpu-row-stride.md, a section of docs/development/gpu-backend-template.md.
- An upstream sync that brings a CUDA, HIP, SYCL or Metal kernel addressing a 16-bit row as
reinterpret_cast<const uint16_t *>(base) + y * stridefailstest_gpu_byte_stride_contract. Port it through a byte pointer (reinterpret_cast<const uint16_t *>(base + y * stride)) or convert the stride to elements. integer_adm_sycl.cpp(adm_dev_dwt_src()) andinteger_vif_sycl.cpp(dev_read_pixel()) copy their element stride into a local*_elemsvariable for the scales above 0. Keep the names on a rebase; the scan reads them.
Port of Netflix/vmaf golden-assertion updates 5c7770080, 005988ead, 4679db83c, d93495f5c, e3827e4dd (2026-10-05)¶
port/upstream-golden-updates-2026-05, ADR-1828.
- Future upstream syncs port Netflix's own golden-assertion updates verbatim (value and
placesexactly as upstream, never a fork-chosen value) after a measurement against the fork's CPU build like the one behind this PR (summary in the PR body). Upstream's value that the fork does not reproduce is a code question, not a reason to loosen or edit a value. - Where upstream's tip differs from a commit's own value (a later upstream commit), the tip is taken:
float_vifks360o97VMAF_scoreandinput160x90VMAF_scorekeep places 3 and 2, not the 4 / 3 of5c7770080. vmafexec_test.pyloses its_IS_DARWINconstant and the per-platform values, as upstream's4679db83chas it (sixVMAFEXEC_scoreassertions, three of them places 4 to 3). A macOS lane that differs from Linux by more than 5e-4 on those scores would now fail.- Not portable:
test_run_vmaf_runner_v1_modelis skipped in the fork (ADR-0865), so its upstream rows have no assertion to edit; the source part of5c7770080(chroma_correction_parameter,postprocess_feature_from_another) already landed with #2136.
Research digests describe third-party products generically (2026-10-05)¶
No rebase impact: docs only.
RC4 owns the device-memory import API (2026-10-05)¶
docs/rc4-zerocopy-scope. Docs and ledger only; no source file.
- ADR-1829 moves the import API with fences from the post-1.0 embedding milestone into RC4.
AGENTS.mdsection 11,docs/roadmap.md,docs/development/release.mdand the RC4 rows ofdocs/state.mdcarry the new scope; the vendor context files are compiled fromAGENTS.md. - On a sync or rebase, keep the fork's RC4 text in all of them. A conflict in the
AGENTS.mdcandidate-map line is resolved per hunk and followed bypraetorctl compile-contextin the rebasing worktree only. No upstream file is touched.
SYCL kernels scratch-free on Xe-LP: slot, term, scale-0 vif hori, 16-bit motion SAD (2026-10-05)¶
SYCL kernels scratch-free on Xe-LP: slot, term, scale-0 vif hori, 16-bit motion SAD; vif SIMD-16 only (2026-10-05)¶
fix/sycl-xelp-scratch (T-SYCL-ROW-KERNELS-SG16-OTHER-DEVICES-2026-10-02). Fork-only files.
Ss2SlotKernel(ssimulacra2_sycl.cpp),IssimTermKernel(integer_ssim_sycl.cpp) andIntegerVifHoriKernel<0, 16>(integer_vif_sycl.cpp, throughvif_hori_sg_size()/vif_hori_grf_size()) haveVmafSyclKernelShape<0, 256>(ADR-1501): no required sub-group size, the large register file.MotionSadHbdKernel(integer_motion_pipeline_sycl.cpp) isVmafSyclKernelShape<16, 0>. A required SIMD-16 (or SIMD-32 for the motion kernel) spills on Xe-LP, which has no 256-entry register file. Keep these shapes on a rebase;test_sycl_kernel_source_contract.py,test_sycl_ssim_exact_contract.pyandtest_sycl_ssimulacra2_exact_contract.pyrefuse the old ones.integer_vif_sycl.cppruns at SIMD-16 only (ADR-1830): the SIMD-32 hori and fused instances,use_simd16,launch_vif_hori_v2andVMAF_SYCL_VIF_SUBGROUP_SIZEare removed, and so istest_sycl_vif_parity_sg32. They spilled on Xe-LP. A sync must not bring any of them back; the source contract refuses it.
CUDA path is not zero-copy: documentation wording (2026-10-05)¶
docs/v1-gpu-fallbacks-zero-copy. no rebase impact: docs only. Edits docs/usage/ffmpeg.md, docs/backends/cuda/overview.md, docs/backends/hip/uploads.md and a dated correction note in docs/research/0086-tiny-ai-sota-deep-dive-2026-05-08.md; no code, no FFmpeg patch.
vif_cuda names its features before it clears enable_chroma (2026-10-05)¶
fix/cuda-vif-names-before-option-reset (T-GPU-VIF-NAMES-AFTER-OPTION-RESET-2026-10-05, ADR-1836). Fork-only file.
init_fex_cuda()incore/src/feature/cuda/integer_vif_cuda.cbuildsfeature_name_dictright after the log2 table upload and beforevif_drop_vestigial_chroma_option(), and ends inreturn vif_setup_buffers(...). Keep that order on a rebase:test_gpu_twin_name_order_contract.pyrefuses an option write before the dictionary in any CUDA, SYCL or HIP twin, andtest_cuda_vif_log2_contract.pyholds the init tail.test_integer_vif_cpu_cuda_parityreads the_enable_chromanames.
Port of Netflix/vmaf 0497a0f29: vmaf_picture_convert, additive variant (2026-10-05)¶
port/upstream-picture-convert-additive, ADR-1822. This API deliberately differs from upstream's until upstream releases it.
- Do not restore
VmafPicture::color. Upstream'spicture.hhunk addsVmafColor color;betweendata[3]andref; the fork keepsVmafPictureas it is (HISS-14).core/test/test_picture_convert_api.cfails the build when a member is inserted. Take upstream's enums,VmafColor,VmafResampleFilter,VmafPictureConvertTargetand the context typedef as they are; they match. - Init differs. Upstream
vmaf_picture_convert_context_init(ctx, src, target)readssrc->color; the fork hasvmaf_picture_convert_context_init_with_color(ctx, src, src_color, target).vmaf_picture_convert()does not compare the source colour (the picture has none) and does not setdst->color. - Different file. Upstream's code is appended to
libvmaf/src/picture.c; the fork has it incore/src/picture_convert.c(built once aspicture_convert_lib), so a sync ofpicture.chunks of0497a0f29is not applicable. Upstream'stest_picture.ccolour-default test is not ported (no field); the without-zimg test and the layout guard live incore/test/test_picture_convert_api.c, the zimg cases incore/test/test_colorspace.c(upstream's, init call adapted, the "mismatched colour" case becomes a mismatched format case). enable_zimghas upstream's name and default (false); the dependency iszimg >= 2.7through pkg-config, also inRequires.private.- Upstream's
goto failin init is split intovalidate_init_args(),build_graph()andallocate_tmp()(HISS-01, HISS-04).
Port of Netflix/vmaf 7922f2c04, 10ec73c73, 6a7b1ae34: SpEED Python tests (2026-10-05)¶
port/upstream-speed-python-tests. Test-only.
python/test/speed_chroma_feature_extractor_test.py,speed_chroma_quality_runner_test.py,speed_temporal_feature_extractor_test.pyandspeed_temporal_quality_runner_test.pyare Netflix's files at upstream0497a0f29with the SPDX line andblackformatting. All 156 assertions are Netflix's, unchanged (AST-identical). The extractors and runners came with #34 (9e99c0d8c); these four files were left out then.- On a sync, take upstream's changes to these files as they are. They pass on the fork's CPU build because SpEED computes Netflix's expressions (ADR-1477); a value that stops matching is a defect in
speed.c, not in the test.
Tester report keeps the cause of a failure (2026-10-05)¶
fix/tester-keep-failure-diagnostics (T-TESTER-REPORT-DROPS-FAILURE-CAUSE-2026-10-05). Fork-local tester code only; no upstream file.
tools/rc1-tester/src/vmaf_rc1_tester/hw_equiv.py:failure_message()builds a failed run'sFixtureRunErrortext (head line unchanged except the signal name, then the diagnostic lines). Keeprun_fixture_meta()calling it.hw_suites.py:run_one_test()returnsTIMED_OUT/OUTPUT_LIMIT/NOT_STARTEDwith the partial output instead of-1, "";run_unit_tests()recordsabnormal_end()intoreason,casesandcase_messages. A conflict with a change to either function keeps both.- The report schema is unchanged;
tests/test_hw_report.pyandtests/test_hw_metal.pypin the text.
One 4:0:0 check for every ssimulacra2 extractor (2026-10-05)¶
fix/metal-ssimulacra2-refuse-yuv400 (T-METAL-SSIMULACRA2-YUV400-ACCEPTED-2026-10-05). core/src/feature/ssimulacra2_pixel_format.h is new and fork-local: vmaf_ss2_check_pixel_format() and vmaf_ss2_has_chroma().
core/src/feature/ssimulacra2.c(fork-local) lost its staticcheck_pixel_format();init()calls the header with the name"ssimulacra2", so the log line is unchanged.cuda/ssimulacra2_cuda.c,sycl/ssimulacra2_sycl.cppandhip/ssimulacra2_hip.c:init()calls the check first; their*_configure_planes()no longer refuse and returnvoid; the context checks usevmaf_ss2_has_chroma().metal/ssimulacra2_metal.mm'sinit()no longer discardspix_fmt.- On a conflict in any of these files keep the call to the header and do not bring back a private
VMAF_PIX_FMT_YUV400Pcomparison:core/test/test_ssimulacra2_pixel_format_contract.pyfails on either. - Netflix golden data unaffected: no score changes for 4:2:0, 4:2:2 or 4:4:4.
integer_adm_metal: shared uniforms, host logic in C, host replay (2026-10-05)¶
fix/metal-integer-adm-twin-defects (T-METAL-INTEGER-ADM-TWIN-DEFECTS-2026-10-05, ADR-1806). Fork-only files; no upstream counterpart.
core/src/feature/metal/metal_integer_adm_uniforms.hdefinesIadmDims,IadmCsfand the reduction slot addressvmaf_mtl_iadm_accum_word()once forinteger_adm.metaland the host. Keep the structs out of the.metaland.mmon a conflict; a field added on one side goes into the header.core/src/feature/metal/integer_adm_metal_host.cholds every host step that does not touch the Metal API (geometry, buffer sizes, the stage plan, uniforms from the CPU's contexts, scores from the CPU'sadm_cm_result()/adm_csf_den_result());integer_adm_metal.mmonly allocates, binds, encodes and emits. A change to the CPU's integer ADM contexts or result functions reaches the twin through that file.core/src/metal/meson.buildlists the host file inmetal_sources;core/test/meson.buildcompiles it intotest_metal_integer_adm_host_replay(not on Windows).integer_adm.metal: scale 1 runsinteger_adm_dwt_vert_s1(int16 parent); the unusedinteger_adm_csf_r_s0/_s123kernels are gone. No identifier namedkernelorhalfin a header the kernels or the host shim include.core/test/metal_msl_host_shim.handcore/test/metal_msl_host/metal_stdliblet a test compile an unmodified.metalfile on the host.- Netflix golden data unaffected (Metal only, and no CPU code changed).
Port of Netflix/vmaf 2e6bbb657, 3685aa3c1, 2f2bb601b: SubjectiveDatasetReader and SubjectiveDatasetTester (2026-10-05)¶
port/upstream-subjective-dataset-reader. Python harness only.
compat/python-vmaf/routine.py:SubjectiveDatasetReaderandSubjectiveDatasetTesterwith Netflix's constructor signatures, attributes (dataset,kwargs,test_assets,test_raw_assets,results,stats) and methods (read(),derive_assets(),run());read_dataset()andrun_test_on_dataset()are Netflix's thin wrappers over them. The bodies are not Netflix's 350- and 160-line methods but the fork's helpers (_read_dataset_assets(),_resolve_asset_fields(),_build_asset_dict(),_run_test_runner(),_compute_test_stats(), ...) extended with Netflix's new behaviour: per-sidewidth/height(_resolve_side_dimensions(),_assert_equalizable(); the old "ref and dis width must agree" assert is gone), per-side resampling (_resampling_entries(), Netflix's four cases), dataset- and video-levelfps_cmd, the distorted video'sworkfile_yuv_type, the hfr model paths anddelete_workdir(_tester_optional_dict()), the runner on the raw assets after subjective modeling.- Kept fork deviations:
allow_uncalibrated/CalibrationError(ADR-0620; the tester takesallow_uncalibrated), bootstrap keys read from bootstrap runners only, the classifier stats branch. python/test/routine_test.py: Netflix's 14TestReadDatasetcases of2f2bb601band the three classes of3685aa3c1(localmatplotlibimports, as in that commit), assertions unchanged; the rest of the file stays the fork's (golden stop of005988ead). 14 fixtures added and twoenc_width/enc_heightpairs intest_read_dataset_dataset3.py, each with the SPDX line andblack.- On a sync: take Netflix's changes to the two classes into the helpers, not as a copy of the long methods.
Metal twins declared exact; 2026-10-05 hardware reports (2026-10-05)¶
docs/hardware-reports-lawrence-2026-10-05. Report files, exact-twin fragments, gate tests and documentation; no library change.
scripts/ci/exact_twins.d/*.metal(18 files) declare the Metal twins the M4 Pro report of 2026-10-05 measured exact. A sync must not drop them, and a Metal twin that drifts from the CPU is fixed, never given a tolerance (ADR-1428).scripts/ci/test_cross_backend_parity_gate.pyno longer namesmetalas the backend outside the gate (_OFF_GATE_BACKEND = "off_gate") and picks its held-exact Metal examples from the fragment set (_metal_unlisted()). Keep it free of hard-coded fragment-free features or backends.docs/hardware-reports/2026-10-05-*.jsonare the testers' files as submitted; never reformat them:report_sha256covers their content.
Port of Netflix/vmaf 6046b1926: SpEED without enable_float (2026-10-05)¶
port/6046b1926-speed-without-float. Netflix builds speed.c, vif_tools.c and common/convolution.c outside the enable_float block since 4718b4f5f (a hunk the fork's port of that commit, #1024, left out) and registers speed_chroma / speed_temporal outside #if VMAF_FLOAT_FEATURES since 6046b1926. The fork now does both, with its own speed_internal.c moved alongside speed.c because the CUDA, HIP and SYCL SpEED twins link it.
core/src/meson.build: the four sources are at the end of the unconditionallibvmaf_feature_sourceslist. On a conflict keep them there; the float block keepsoffset.c,adm.c,adm_tools.c,vif.cand the float extractors.core/src/feature/feature_extractor.cpp: the two externs and list entries sit after#endifof the float block. The list order of a float build is the same as before.core/test/meson.build:test_speed,test_speed_frame_buffers,test_speed_lanczos4_weights,test_speed_filter,test_speed_upstream_form*andtest_speed_simdlost theirenable_floatgate.- No score moves in a float build; Netflix golden data unaffected.
Hardware-we-need CPU rows read the CPU checks (2026-10-05)¶
fix/hardware-needs-cpu-row-verdict. Documentation generator and its test; no library change.
scripts/docs/generate-hardware-reports.py::_host_verdict()rates a CPU row from the report's own CPU checks (CPU_CHECKS, plusMETAL_CHECKSfor a row withmetal_rows), read throughCHECK_KEYSandPASSINGoftools/rc1-tester/src/vmaf_rc1_tester/hw_report.py, and fromimage.files_match_build. A sync must not bring backreport["verdict"]for CPU rows, and must not copy the passing statuses into the generator: they have one definition, in the tester.CpuRowVerdictTestsfails on the old form.
Port of Netflix/vmaf ac9467ff4: VMAF v1 tech blog link (2026-10-05)¶
port/ac9467ff4-v1-techblog-url. Docs only. Upstream replaced the "tech blog XXX" placeholder in README.md and resource/doc/models_v1.md; the fork's counterpart is docs/models/v1.md, which had dropped the placeholder sentence. The link is in its introduction now; the fork's README.md has no v1 news line.
Port of Netflix/vmaf 8cdd55a03, 314f14b22, 12aa1cb44: three unit tests (2026-10-05)¶
port/upstream-tests-motion-blend-predict-barten. Test-only; no score moves.
core/test/test_motion_blend.c(new, upstream's file with the fork'sstatichelper and ADR-1138 bracket) and itsexecutable()/test()block next totest_barten_csfincore/test/meson.build, with upstream03981a80e'sstdatomic_dependency.core/test/test_barten_csf.c: upstream's 32 v1.0.17barten_watson_blend_csf_mae()cases as four functions of eight assertions underrun_tests_blend_mae_v1017()(function-size limit).core/test/test_predict.c: upstream includespredict.cinto the test to reach the file-staticpost_process_feature_from_another(). The fork links one predictor implementation (Research-2096), sopredict.cexportsvmaf_predict_post_process_feature_from_another_for_test()(declared incore/src/predict.h), a pass-through. On an upstream change to that function's parameters, change the test entry with it; do not bring back#include "predict.c".
Port of Netflix/vmaf d327ed67b, 3dee96664, 560c4e491, 5c7770080 (Python harness) and the CAMBI tests of 095bb1818 / 83b4f1306 (2026-10-05)¶
port/upstream-python-harness-2026-05. Python harness only; no score of the vmaf executable moves.
compat/python-vmaf/core/feature_extractor.pyandquality_runner.py: Netflix's multi-nickname discovery (d327ed67b). It replaces the fork's "shortest key wins" wildcard inVmafexecFeatureExtractorMixinand the "first key wins" wildcard ofFeatureDiscoveryMixin. Fork deviations: the frame-count check isassert_same_frame_count()(count of the first non-empty nickname, where Netflix's loop leaves it unset when feature 0 is absent),atom_features=Nonefalls back tocls.ATOM_FEATURES, and the result loop is_collect_feature_result()(HISS-04).compat/python-vmaf/core/cambi_feature_extractor.py: the fork'scambi_encbdatom feature ofCambiFullReferenceFeatureExtractorexisted only to dodge the shortest-key rule; Netflix'scambiis back, andpython/test/cambi_test.pyis Netflix's file at0497a0f29(SPDX line,black; 14 tests the fork lacked, the 4 Netflix had renamed in095bb1818gone, every value of the common tests identical). On master that file fails 2 of 29 on theCambi_FR_feature_cambi_scorekey.VMAF_FLOAT_FEATURE_OPTION_TARGETS/VMAF_INTEGER_FEATURE_OPTION_TARGETS: the nine options of3dee96664the tables lacked (float:vif_prescale,vif_prescale_method,adm_bypass_cm,adm_adm3_apply_hm,adm_p_norm,adm_skip_aim_scale,motion_add_scale1,motion_add_uv; integer:adm_skip_aim). The fork's ownmotion_five_frame_windowandmotion_moving_averagestay.compat/python-vmaf/core/asset.py:ORDERED_FILTER_LISTis Netflix's order withselect(560c4e491); the fork had putfpsandformatbeforegblur.compat/python-vmaf/core/train_test_model.py:chroma_correction_parameterandpostprocess_feature_from_another()(5c7770080), split into_find_guiding_and_guided()for HISS-04, returning Python floats; the ResPow guard catchesExceptionwhere Netflix has a bareexcept:. Netflix's assertion changes of5c7770080are not ported (golden stop).
Inherited Netflix tags are deleted (ADR-1805, 2026-10-05)¶
chore/rc3-netflix-tags. Repository refs and one hook; no library change.
- VMAFx/vmafx carries none of Netflix's tags: the 26 inherited ones were deleted and recorded in
scripts/release/inherited-upstream-tags.jsonanddocs/development/inherited-tags.md. An upstream sync never pushes a Netflix tag and never fetches them into a clone: keepgit config remote.upstream.tagOpt --no-tags, and nevergit push --tagsfrom a clone that fetchedupstreamwith tags (it would restore them). lefthook.ymlpre-pushrunsscripts/git-hooks/check-push-tags.py(tag-guard); a rebase oflefthook.ymlkeeps the command withuse_stdin: true. Its contract isscripts/release/tests/test_inherited_tags.py.
integer_vif_metal divides in single precision (2026-10-05)¶
fix/metal-vif-single-precision-ratio. One host line, comments, one contract test; no kernel change.
core/src/feature/metal/integer_vif_metal.mm::collect_fex_metal()builds itsVmafVifScoreSetwith.single_precision_ratio = true, asinteger_vif.c::write_scores()and the CUDA, HIP and SYCL twins do. A sync that rewrites the collect path must keep the flag andscale_num_den()'s twofloatroundings.core/test/test_sycl_vif_float_sums_contract.pynow checks the host tail of every integer VIF twin (CUDA, HIP, SYCL, Metal) from one table (TWINS); a new twin or a renamed tail function is a new row there. The file keeps its name and meson test name.
Port of the rest of Netflix/vmaf c70debb10: test_vif_tools, test_speed_chroma (2026-10-05)¶
port/c70debb10-vif-tools-speed-chroma-tests. Entry "0085" below ported the ADM and Barten halves of c70debb10 and left these two files out because the fork then had no VIF runtime helpers; it has had them since ADR-0416. Test-only.
core/test/test_vif_tools.c: upstream's file and tables unchanged apart fromstatic,(void)prototypes and the ADR-1138 bracket.core/test/test_speed_chroma.c: upstream's test includesspeed.c; the fork's tests may not include sources (check-no-non-header-includes), sospeed.cgains two more test entries next to the ADR-1477 ones,speed_internal_cpu_est_params()(scalar kernels, work buffers allocated per call) andspeed_internal_cpu_compute_eigenvalues(), declared inspeed_internal.hwithSpeedInternalEstGeometry; the score case uses the existingspeed_internal_cpu_speed_score(). Upstream's inputs and expected values are file-scope arrays and one helper runs the threeest_params()cases (function-size limit). A change to the signature ofest_params()orcompute_eigenvalues()inspeed.cchanges the entries in the same PR.- Upstream gates both on
enable_float; the fork does not, becausevif_tools.candspeed.care in every build since 6046b1926. On a sync, do not bring the gate back.
Port of Netflix/vmaf 78e11b52c: bilinear prescale column tables (2026-10-05)¶
port/78e11b52c-speed-bilinear-columns. Bit-identical; a speed-up only.
core/src/feature/vif_tools.c: the per-pixelbilinear_interpolation()is gone.vif_bilinear_columns()/vif_bilinear_rows()compute the same values in the same order (xxformed in double and rounded to float,dxfrom the mirrored floor, the four-term sum in the per-pixel order); upstream's public pairvif_scale_frame_bilinear_precompute_columns_s()/vif_scale_frame_bilinear_precomputed_s()is invif_tools.h.- Deviation from upstream, kept deliberately: upstream's generic path keeps one table of
VIF_BILINEAR_MAX_WIDTH(7680) entries on the stack, asserts on wider outputs and refuses them at init infloat_vifandfloat_motion. The fork walks the columns in chunks of 1024 (VIF_BILINEAR_COLUMN_CHUNK), so no caller has a width limit and neither init check exists here; the fork'sfloat_motionscale-1 path has its own scaler inmotion.c. On a sync, do not bring the macro, the assert or the init checks back. core/src/feature/speed.c:SpeedBufferscarries the instance's table (bilinear_x1a/_x2a/_dxa, filled byspeed_alloc_bilinear_columns()when the prescale is bilinear and resamples, freed byspeed_free_bilinear_columns());filter_and_downscale()takes the buffers and its prescale step isspeed_prescale_frame().speed_internal.c's mirror keeps callingvif_scale_frame_s(), which returns the same bits.core/test/test_vif_bilinear.cholds both paths to the per-pixel scaler (copied into the test) withmemcmp.
psnr_hvs_metal stores every term and takes the CPU's masking table (2026-10-05)¶
fix/metal-psnr-hvs-cpu-sum. Metal twin, its host and the shared host tail.
core/src/feature/metal/integer_psnr_hvs.metalstores the 64 terms of every block (terms[slot * 64 + lid]) and sums nothing on the device; every operation on a value comes fromcore/src/feature/metal/metal_psnr_hvs_math.h(onmetal_portable.h, nodouble). A sync must not bring back the per-blockretpartial, the host's float sum of partials or the kernel's own masking table(csf * 0.3885746225901003f)^2.integer_psnr_hvs_metal.mmbinds the masking table at buffer 6, formed withvmaf_psnr_hvs_mask_value()(new incore/src/feature/psnr_hvs_score.c, the CPU's double product stored as float), and scores throughvmaf_psnr_hvs_plane_score(),vmaf_psnr_hvs_combined_score()andvmaf_psnr_hvs_score_db(). The CSF tables moved from the.mminto the header (vmaf_mtl_hvs_csf).- A change to
calc_psnrhvs()inthird_party/xiph/psnr_hvs.cchanges the header in the same PR, as it changes the CUDA, HIP and SYCL kernels.test_metal_psnr_hvs_mathandtest_psnr_hvs_twin_exact_sum_contract.py(now with the Metal twin) guard it without a device.
CAMBI copies a same-size 10-bit plane row by row (2026-10-05)¶
fix/cambi-fullref-wide-source-rows. One function of cambi.c, one new test, three test extensions; no API change.
decimate_same_size_16b()incore/src/feature/cambi.ccopies a 10-bit plane one row at a time with the input's stride and the working picture's stride. Upstream Netflix/vmaf still copiesstride * heightsamples in onememcpy(libvmaf/src/feature/cambi.c, insidedecimate_generic_uint16_and_convert_to_10b()), which shifts every row whenever the strides differ: underfull_refwith a source larger than the picture, and for a caller's picture with a wider stride. Reported upstream as Netflix/vmaf#1670. An upstream sync that touches that function keeps the fork's row loop and never takes upstream'smemcpyback. Drop the fork's version only when upstream merges an equivalent row-by-row copy that honours both strides, and keepcore/test/test_cambi_full_ref_wide_source.ceither way.core/test/test_cambi_full_ref_wide_source.cfails on the single copy (3006 misplaced samples in each stride direction, and 10-bitcambi2.5075 against 5.1444 withsrc_width=640:src_height=480).core/test/test_{cuda,sycl,hip}_exact_twins.cgainedcpu_opts/cpu_keysinExactCaseand one CAMBI case that gives only the CPU extractorfull_ref=true:src_width=1280:src_height=960; keep the fields and the case when the files are merged or regenerated.
Metal CAMBI names and heatmaps (T-METAL-CAMBI-SCORE-NAME-SUFFIXED-2026-10-05, 2026-10-05)¶
fix/metal-cambi-feature-name-order. CAMBI CPU heatmap writers, the Metal twin, a contract test and test_cambi.
core/src/feature/cambi.cis an upstream-mirror file.open_heatmaps()takes(heatmaps_path, enc_width, enc_height, heatmaps_files)instead ofCambiState *,close_heatmap_files()replaces the close loop ofclose_cambi()(which now returns-EIOwhen a close fails), andinit(),cambi_score()andclose_cambi()call the trampolinesvmaf_cambi_open_heatmaps(),vmaf_cambi_dump_c_values()andvmaf_cambi_close_heatmaps()at the bottom of the file. An upstream change toopen_heatmaps(),dump_c_values()or the close loop is ported into these state-free forms; the Metal twin writes its heatmaps through them.core/test/test_cambi_heatmap_writers.cand the cambi.cREFERENCE_LINESofcore/test/test_metal_twin_option_tables_contract.pyfail on a sync that loses them. The file spells the null pointerNULLunder the ADR-1138NOLINTBEGIN(modernize-use-nullptr)bracket (the same hunks as #2111).core/src/feature/metal/integer_cambi_metal.mm:init_fex_metal()builds the feature-name dictionary beforecambi_metal_resolve_dimensions(). Never move it after any write to an option slot; the contract's_name_order_failureschecks every Metal twin.
libvmaf and libvmaf_cuda print no score after an error (ADR-1768, 2026-10-05)¶
fix/ffmpeg-libvmaf-no-score-after-error. New FFmpeg patch 0021, test helpers under ffmpeg-patches/test/.
0021deliberately diverges from upstream FFmpeg'svf_libvmaf.c; keep it on every series refresh and do not drop it when upstream moves the surrounding lines.stop_on_frame()logs the frame and the error once, setsstopped(declared by0005) and frees the frame;do_vmaf()anddo_vmaf_cuda()call it on every failed copy and read, andframe_cntadvances only after a successful read. The shareduninit()prints no pooled score after a stop or a failed flush, and no score line for a failed model. If upstream rewrites these functions, carry the rule into the new code; do not take upstream's side.ffmpeg-patches/test/fault_inject_sycl_import.cbecamefault_inject_libvmaf.c(it now also injectsvmaf_read_pictures()failures); both checks sourcefilter_check_lib.sh.core/test/test_ffmpeg_libvmaf_stop_contract.pyreads the series;ffmpeg-patches/test/check-libvmaf-no-score-after-error.shis the device run.
libvmaf_sycl import retry, per-input VA display (ADR-1761, 2026-10-05)¶
fix/ffmpeg-sycl-import-retry. FFmpeg patch 0005 and a device test under ffmpeg-patches/test/.
0005:LIBVMAFContextgainsva_display_refandstopped.config_props_sycl()reads each input's display throughqsv_link_va_display().do_vmaf_sycl()imports throughimport_va_surface_retry(), a bounded loop ofLIBVMAF_SYCL_IMPORT_TRIESwith anav_usleep()between tries, and never returns the frame after a failed import.uninit_sycl()prints no pooled score whenstoppedis set. The software path copies(height + 1) / 2chroma rows.stoppedsits outside the SYCL block:0013uses it too.0013:do_vmaf_metal()setsstoppedon a failed wait, import or read and advancesframe_cntonly after a read succeeded;uninit_metal()prints no pooled score when it is set. A refresh keeps both; the other later patches moved by offsets only.core/test/test_sycl_filter_import_contract.pyreads the patch;ffmpeg-patches/test/check-sycl-import-retry.shwithfault_inject_libvmaf.c(LD_PRELOAD) is the device run.
Tester selectors follow their own paths; nightly tester image (ADR-1700, ADR-1701, 2026-10-05)¶
ci/tester-legs-own-paths-and-nightly. The impact planner, its map, one workflow and their tests; no library change.
scripts/ci/plan-ci-impact.pyhonours the selector propertyown_paths_onlyinfull_plan()(true only on a known changed path of its own; true when the change list is unknown) and resolves selector inheritance withinheritance_order()andimpact_selectors()instead of the recursive_selector_value(). A sync must not bring the recursion back or move the exception from.github/ci-impact.jsoninto code..github/ci-impact.jsondeclares"own_paths_only": trueontester_imageandwindows_tester_zipand nowhere else;test_ci_impact.py(OwnPathsOnlyContract) fails on a third.docker-publish-tester.ymlruns nightly (schedule, 00:29 UTC, amd64, no publish, no GPU images, concurrency groupnightly);build-gpuexcludes the schedule event andvalidatenarrows the matrix on it. Keep the slot ahead ofnightly.ymland the weekly Release Dry Run (test_pr_time_verify_workflows.pychecks the timeouts against both).
adm_cuda and adm_hip take the CPU's integer reciprocal (2026-10-05)¶
fix/gpu-adm-decouple-integer-reciprocal. decouple_r_s0() of core/src/feature/cuda/integer_adm/adm_decouple_inline.cuh and core/src/feature/hip/integer_adm/adm_decouple_inline.hip takes 2^30 / o from adm_recip_q30() (in the same headers), which equals div_lookup[o + 32768] of core/src/feature/integer_adm.h for every int16 operand. A sync or rebase must not restore int32_t(div_Q_factor / float(o_val)) (another integer for 343 positive operands) and must not replace the fp32 correction with an integer division: that costs 80 registers in adm_cm_line_kernel_8 and fails test_cuda_adm_cm_register_pressure. core/test/test_adm_decouple_recip.cpp (one executable per twin header) fails on the old form without a device.
The tidy coverage check (RC3 exit, 2026-10-05)¶
rc3-tidy-coverage-2. Tooling only. scripts/ci/check-tidy-coverage.py, its hooks in .pre-commit-config.yaml and .config/lint-exceptions.d/clang-tidy-coverage.toml belong together; a sync that adds a .c / .cpp / .mm / .cu / .hip file lands it in a lane (scripts/dev/tidy-lane.sh --write --only <file> <lane>) or adds one entry there. The Pelorus mirror entries mirror scripts/ci/pelorus-mirror-paths.txt (test_check_tidy_coverage.py compares them). write-compile-commands.py keeps its OPTIONAL_COMPILER_RULES: without them the macOS lane loses every .mm.
Every translation unit is read by a tidy lane (RC3 exit, 2026-10-05)¶
rc3-tidy-coverage. Tooling only; no library change.
MakefileTIDY_RATCHET_SETUP_cpucarries-Denable_mcp=true -Denable_mcp_sse=enabled -Denable_mcp_uds=true -Denable_mcp_stdio=true, and theTidy Ratchetjob oflint-and-format.ymlrepeats them (test_tidy_lane_container.pycompares the two). A rebase that takes either side must keep both lines the same.- Lanes
clangandmetalare new (scripts/dev/tidy-lane.shknowsclang;metalruns in theTidy Metalworkflow,tidy-metal.yml).scripts/ci/tidy-baseline-clang.jsonandtidy-baseline-metal.jsonare generated: take one side at a conflict and regenerate. - The exception entries for units no lane reads land in the follow-up pull request, in the shared lint exception list.
tidy-ratchet.py --selectand the scoped write'smeasured_sourcesupdate are additive. core/test/test_mcp_*.c,core/src/mcp/{mcp,dispatcher}.cand three fuzz harnesses carryNOLINTBEGIN(modernize-use-nullptr)blocks citing ADR-1138 (C builds withoutnullptron MSVC); an upstream-style sync must keep them.
Root licence files follow ADR-1250; package licence fields match their files (ADR-1699, 2026-10-05)¶
fix/root-licence-files-eupl. Root files, manifests, go.mod, licence record, the provenance tool, one CI step and one hook; no C library change.
- The root holds
LICENSE(the EUPL-1.2, byte for byteLICENSES/EUPL-1.2.txt) andNOTICE(Netflix'sLICENSE, moved unchanged;REUSE.tomlannotates it). An upstream sync that touches Netflix'sLICENSElands the change inNOTICEand keeps the fork'sLICENSE; it never brings backLICENSE-MIT, aLICENSE-BSD-2-Clause-Patentor any other root file name licensee reads as a licence file.scripts/ci/check_licence_metadata.pyrefuses each. - The root
go.modcarriesretract [v1.0.0-rc.1, v1.0.0-rc.2]; ago modrewrite or a sync ofgo.modkeeps it. tools/rc1-tester/image/licensing.jsontakes the BSD-2-Clause-Patent text (spdx_textsand every Netflix/vmaf_resourcerepotext) fromNOTICE; every Dockerfile licence stage bind-mounts that file (notLICENSE),docker-publish-tester.ymlchecks it out from the recipe ref, also in the GPU job, and thetester_imageselector of.github/ci-impact.jsonlists it.- Licence fields: workspace
Cargo.tomlandbindings/rust/vmafx/Cargo.tomlEUPL-1.2;core/src/feature/rust/tad/Cargo.tomlits ownEUPL-1.2 AND BSD-2-Clause-Patent; Helmartifacthub.io/license: EUPL-1.2;tools/rc1-testerEUPL-1.2 AND MITwithLICENSES/;deny.tomlallowsEUPL-1.2.REUSE.tomlrecordsdeploy/helm/vmafx/charts/prometheus-pushgateway-*.tgzas Apache-2.0 (newLICENSES/Apache-2.0.txt). A Renovate bump of the subchart keeps that pattern matching. - The Python package model moved from
python/test/setup_metadata_test.pyinto the gate module, which the test imports; change it there. scripts/dev/relicense_fork_files.pyclassifies.tomland treatspyproject.tomlas a name without provenance signal;.codex/agents/is excluded. Every fork.tomlfile carries an EUPL-1.2 header; a new one without it failsLicence Provenance(--writeadds it). Ten lock headers were restamped for thepyproject.tomlheader lines.dev/Containerfile: the licences label sits onlibvmaf-build(derived from the files it copies);build-depsandrelease-buildcarry none.
vmaf_sycl_upload_plane waits for its copy (2026-10-05)¶
fix/sycl-upload-plane-fence. SYCL runtime, one public-header comment.
core/src/sycl/common.cpp::vmaf_sycl_upload_plane()callssycl_fence_slot_readers()and waits for the last copy event before it returns. A sync that restores the fire-and-forget copy brings back wrong 4K scores on the zero-copy and D3D11 paths;test_sycl_zero_copy_model_gate(test_upload_plane_orders_compute) fails on it.- The file's four
getenv()calls read throughvmaf_gpu_dispatch_env_get()(ADR-0488), which keepscommon.cppat zero clang-tidy findings.
Hip-lane clang-tidy debt, batch 3: CAMBI, CIEDE, MS-SSIM, SpEED and the HIP test sources (ADR-1142, 2026-10-05)¶
refactor/tidy-zero-hip-3. Lint refactor, no behaviour change.
cambi_hip_device.h,ms_ssim_arith.h,speed_hip_device.handcore/test/hip_float_adm_math_sample.hare included by C translation units: they keepstruct X {...};with a C-onlytypedefand C++-only includes behind__cplusplus.speed_hd_block_statistics()zeroes its array with= {0.0f}, not a range-for.test_hip_cambi_device_math.csplits the registered-extractor check intoreference_score()andextractor_score(); every assertion is kept.test_sycl_fp_arith_contract.ccomputesprod/prod_rootonly inside#if FP_ARITH_HAS_PROD_ROOT.
Hip-lane clang-tidy debt, batch 2: ADM, moment, motion, PSNR and SSIM kernels (ADR-1142, 2026-10-05)¶
refactor/tidy-zero-hip-2. Lint refactor, no behaviour change.
core/test/test_hip_float_psnr_exact_contract.pymatches>> 16u/<< 16uinfloat_psnr_score.hip; the pinned property is unchanged.adm_dwt2.hipforms every signed arithmetic shift throughadm_dwt_asr()(one cited NOLINT, ADR-1423);float_adm_score.hiptakes device addresses throughfadm_device_ptr<T>()(one cited NOLINT, ADR-1458). A sync that inlines either brings the findings back.float_motion_rows.handssim_decimate.hstay valid C (C++-only includes behind__cplusplus,struct Xplus a C-onlytypedef);motion_v2_score.hipkeeps two cited signed-shift NOLINTs (ADR-0138, ADR-0139).
Hip-lane clang-tidy debt, batch 1: VIF, SSIMULACRA 2, PSNR-HVS and the shared GPU headers (ADR-1142, 2026-10-05)¶
refactor/tidy-zero-hip-1. Lint refactor, no behaviour change.
VifBufferHip.refand.disincore/src/feature/hip/integer_vif_hip.hareuint16_t *, assigned invif_hip_layout_planes(); the kernels no longer cast an integer to a pointer. A sync that restoresuintptr_tfields brings backperformance-no-int-to-ptrinvif_statistics.hip.core/test/test_hip_ssimulacra2_exact_contract.pymatches the kernel'skChunkPixels/state.sumspellings; the properties it pins (raster order within a chunk, lane order, pixel-order fallback) are the same.core/src/feature/ordered_sum.hhasvmaf_ordsum_is_odd()in place ofx & 1on a signed value;adm_angle_flag.h,float_adm_gpu_common.handfloat_vif_gpu_common.huse unsigned shift counts andstruct X {...};with a C-onlytypedef(the CUDA, SYCL, Metal and C hosts include them). Keep both forms on a rebase.
rc.3 is cut without outside-hardware reports (ADR-1707, 2026-10-05)¶
docs/rc3-exit-without-outside-hardware. no rebase impact: docs only (ADR, ledger disposition, changelog fragment).
CUDA lane clang-tidy cleanup (ADR-1142, 2026-10-05)¶
refactor/tidy-zero-cuda-1. Lint refactor of CUDA kernels; no behaviour change (scores identical, sm_89 SASS identical or differing by commutative operand order).
core/src/cuda/cuda_device_ptr.cuhis new:VMAF_CUDA_DPTR(T, address)replacesreinterpret_castof aCUdeviceptr. It is a macro because an inline function losesld.global.ncloads. An upstream sync that brings a kernel with areinterpret_cast<T *>(a.field)converts it.- No designated initializers in
.cu/.cuh(nvcc MSVC host frontend,preflight.sh --stage msvcism): aggregates are filled field by field. integer_vif/vif_statistics.cuh:vif_statistic_calculation()is split intovif_sigmas(),vif_gain(),vif_accumulate_log()andvif_statistic_pixel()and loses its unusedhparameter;filter1d.cufollows. A Netflix change to the statistic is ported into the helpers.- Source-text contract tests (
test_cuda_*_contract.py,test_integer_vif_sv_sq_contract.py) follow the new spelling; their assertions are unchanged.
RC4 Rust extractor framework (ADR-1713, 2026-10-05)¶
rc4/rust-extractor-framework (draft, label rc4, lands after the v1.0.0-rc.3 tag). Build system, registry, libvmaf registration paths, Rust workspace, CI. No score changes; the C extractors stay the default.
core/src/meson.build:is_rust_enabledandcdata.set10('HAVE_RUST_FEATURES', ...)sit directly aboveconfig_h_target; the old TAD block (rust_tad_dep,rust_tad_direct_sources,HAVE_RUST_TAD) is replaced byrust_core_dep/rust_shim_sources. A sync that brings back the TAD block or aHAVE_RUST_TADdefine reintroduces the unregistered-TAD defect.core/src/feature/feature_extractor.cpp: every registry walk goes throughregistry_at(); theHAVE_RUST_TADextern and list entry are gone.vmaf_get_feature_extractor_by_feature_name()is split intofirst_pass_eligible()/fallback_eligible()/provides_feature(); an upstream change to that lookup is re-applied on the helpers, keeping the Rust-twin skip.device_twin_flagsincludesVMAF_FEATURE_EXTRACTOR_RUST.core/src/feature/feature_extractor.h: flagVMAF_FEATURE_EXTRACTOR_RUST = 1 << 8and three functions (vmaf_feature_extractor_install_rust_registry,vmaf_feature_extractor_impl_select,vmaf_feature_impl_rust_requested). An upstream flag at bit 8 must move, not this one.core/src/libvmaf.c:vmaf_rust_twins_install()before the registry audit invmaf_ctx_subsystems_init(), andvmaf_feature_extractor_impl_select()invmaf_use_feature(),vmaf_use_features_from_model()andcreate_context_fallback(). Keep all four on a conflict.core/src/feature/tad_rust.cis gated onHAVE_RUST_FEATURES; the TAD crate is an rlib withoutbuild.rs.- Two Cargo workspaces: the root one (bindings)
excludescore/src/rustandcore/src/feature/rust/tad;core/src/rust/Cargo.tomlholds the libvmaf-linked crates and the TAD crate (package.workspace). A sync that adds those crates back to the root members breaks the offline build.rust-ci.ymllost the "Rebuild vmafx-tad after a source edit" step with TAD'sbuild.rs(the code it guarded is gone) and gained the empty-CARGO_HOME offline build. - Generated files:
core/src/rust/include/vmafx_rs.h(take either side, runscripts/dev/rust-abi-header.sh). Lockfiles: take master's side, thencargo metadata --offline --format-version 1in the affected workspace re-adds its members; no version may move.
Post-1.0 embedding milestone is an ADR and a roadmap row (ADR-1685, 2026-10-05)¶
docs/adr-post-1-0-embedding-milestone. no rebase impact: docs only.
The pull-request release legs are required contexts (ADR-1687, 2026-10-05)¶
ci/require-release-dry-run-legs. CI workflows, the impact map and their tests; no library change.
docker-publish-tester.ymlandwindows-tester-bundle.ymlhave no triggerpaths:any more: jobimpactrunsscripts/ci/plan-ci-impact.py,validateneeds it, and the gatesTester Image(needsbuild) andWindows Tester Zip(needsverify) own the required contexts. A sync or rebase must not restore the trigger filters (test_ci_impact.pyrefuses them) or point a gate at an earlier job. Theirvalidatejobs are namedValidate tester image sourceandValidate Windows zip source, apart from the macOS bundle'sValidate source..github/ci-impact.jsonselectorstester_imageandwindows_tester_zipare the former path lists; both workflows are infull_patterns. A new input of either build goes into its selector and intoFORMER_TRIGGER_PATHSofscripts/ci/tests/test_required_release_legs.pytogether.release-dry-run.ymlgains jobgate(Release Dry Run). Inrequired-aggregator.ymlthe three names sit inrequired,strictMustReportanddelayedStrictDependencies, andRelease Dry RuninpullRequestOnly; a conflict in any of those arrays keeps both sides' names.scripts/ci/required_aggregator_harness.pystrips comment lines before it reads therequirednames and takes aneventargument; keep both.
A superseded Scorecard master run ends cancelled (ADR-1686, 2026-10-05)¶
fix/scorecard-superseded-master-runs. CI workflow and its gate script; no C change.
.github/workflows/scorecard.yml,gatejob:if: ${{ !cancelled() }}(notalways(), which keeps the job running through a cancellation), job-levelactions: writewith its comment,RUN_IDin the policy step, and the step's handling of exit 3 (cancel the run, bounded wait,exit 1). An upstream or rebase conflict there keeps all four; ADR-1673's concurrency form does not apply toscorecard.yml, whosescorecard-${{ github.ref }}group keeps the newest pending run.scripts/ci/scorecard_gate.py:final_master(),descendant_distance(),master_identity(reference, sha, compare),fetch_comparison(),Superseded(not aValueError, somain()never turns it into exit 1) andSUPERSEDED_EXIT = 3, which the workflow step tests literally.- Guards:
scripts/ci/tests/test_scorecard_gate.py(MasterSupersessionTests,ComparisonReadTests,MasterCliOutcomeTests) andscripts/ci/tests/test_scorecard_workflow.py(SupersededMasterRunTestsruns the step's ownrun:script with stubgh,python3andsleep).
Metal IOSurface import reads NV12 / P010; libvmaf_metal imports whole frames (ADR-1679, 2026-10-05)¶
fix/ffmpeg-metal-filter-planes. Metal host code, one public-header comment, FFmpeg patch 0013.
core/src/metal/picture_import.mmreads each plane throughcore/src/metal/iosurface_layout.h: the surface's CoreVideo pixel format picks the layout, a bi-planar surface's planes 1 and 2 are the even and odd samples of its second plane, P010 is shifted by 6, anything outside the table returns-ENOTSUP. A sync or refactor must not bring back a copy of the surface's planenas picture planen(IOSurfaceGetBaseAddressOfPlane(surf, plane)).- Patch
0013:do_vmaf_metal()imports planes 0, 1 and 2 of both frames throughimport_metal_frame()and fails on an import error;config_props_metal()checks both inputs withcheck_metal_input()(NV12 / P010 only, format named in the error). A refresh of the series keeps those functions; patches0014to0020only moved by offsets. core/test/test_metal_iosurface_filter_contract.pyreads the patch and the.mm;test_metal_iosurface_layoutruns the header on every host;test_metal_iosurface_import_parity(a Metal parity test, also a self-test on every host) runs the import on IOSurfaces in the macOS tester bundle.
SYCL zero-copy admission (ADR-1688, 2026-10-05)¶
fix/sycl-zero-copy-chroma. libvmaf C, eight SYCL extractor registrations, one public-header comment, FFmpeg patch 0005.
VmafFeatureExtractorgainsreads_shared_luma_onlyafterreads_prev_prev_ref(core/src/feature/feature_extractor.h). C++ designated initializers follow declaration order: an extractor sets.reads_shared_luma_onlyafter.provided_featuresand before.chars/.context_check.core/src/libvmaf.c:sycl_zero_copy_admit()runs first invmaf_read_pictures_sycl(), beforepic_cnt++andvmaf_sycl_advance_frame();vmaf_flush_sycl()skips an extractor that is not initialized. A sync must keep both.- Hooks:
adm_sycl,vif_sycl,motion_v2_sycl,cambi_sycl,float_moment_sycl(true),motion_sycl(!motion_add_uv),psnr_syclandpsnr_hvs_sycl(!enable_chroma). A change that makes one of them read a host picture removes or narrows its hook in the same PR;test_sycl_zero_copy_admissionlists every SYCL extractor's answer. integer_psnr_hvs_sycl.cpp(touched for its hook, brought to zero clang-tidy findings, sycl baseline 5 → 0): the two kernels' Hillis-Steele scan ishvs_group_inclusive_scan(), the compact kernel's per-block work ishvs_record_plane_offsets()(constant plane indices, ADR-1395) andhvs_pack_block_terms(), andreduce_hvs_planes()bounds its plane loop byPSNR_HVS_NUM_PLANES. Integer arithmetic unchanged:test_sycl_psnr_hvs_parity(==), the scratch audit (131 kernels, 0 in scratch) and the CLI on the Netflix pair (48 of 48 frames identical to the CPU at--precision max) on the A380.- Patch
0005:do_vmaf_sycl()incrementsframe_cntonly aftervmaf_read_pictures_sycl()accepted the frame and maps-ENOTSUPto a message;uninit_sycl()prints no score line after a failed pooled score. Later patches moved by offsets only.
Master push runs are not cancelled by concurrency (ADR-1673, 2026-10-05)¶
ci/master-runs-not-cancelled-by-concurrency. Workflows only; no C change.
- Every push-to-master workflow's concurrency group is
... ${{ github.ref == 'refs/heads/master' && github.sha || github.ref }}withcancel-in-progress: ${{ github.ref != 'refs/heads/master' }}. An upstream or rebase conflict in aconcurrency:block keeps this form. The serialising blocks (dev-container-publish.yml,release-please.yml,scorecard.yml,docs.ymldeploy) stay as they were.scripts/ci/tests/test_master_concurrency_contract.pyandscripts/ci/test_security_workflow_contract.pyguard it; add a new push-to-master workflow in the same form.
The node's eBPF object is generated at build time (ADR-1622, 2026-10-05)¶
build/bpf-object-at-build-time. Build system, Go node, CI; no C library change.
cmd/vmafx-node/bpf/rclonebypass_bpfel.ois deleted from the tree and git-ignored;scripts/dev/gen-node-bpf.shgenerates it. A sync or rebase must not restore it (ScorecardBinary-Artifactsflags the ELF) and must keep every caller of the script:go-ci.yml(through.github/actions/gen-node-bpf),docker/Dockerfile.node(go-builderstage and its clang / llvm / libbpf packages),dev/Containerfile(go-buildstage), thenode-bpfMakefile prerequisites ofgo-buildandgo-test.rclonebypass_bpfel.gostays committed with its embed rewritten toembeddedObject()(object_embed.go,embed_generated_object.sh,rclonebypass_bpfel.o.NOTICE); a sync that restores bpf2go's owngo:embedbreaks the lockedGo API Compatibilitygate, which builds the package without the object.praetor-api.ymlis a locked asset and is untouched.build-config.envownsBPF_CLANG_VERSIONandBPF_OBJECT_SHA256; a change torclone_bypass.bpf.c,vmlinux.hor thego:generateflags re-records the digest.gen.go: thego:generateline runs bpf2go at thego.modmodule version (no@version, no-cc); the compiler comes fromBPF2GO_CC, which the script sets.
Metal helper-header licences and tester signature suffix (2026-10-05)¶
fix/master-red-lint-scorecard-2026-10-05. Data, two workflow lines and one test.
scripts/dev/relicense_provenance.tomlgains two[ports]entries (metal_float_motion_math.h,metal_float_psnr_math.h) and two[not_ports]entries (metal_float_moment_sum.h,metal_integer_vif_math.h); five Metal headers carry the dual tag. A sync keeps the entries and the headers'Netflixnotice lines.macos-tester-bundle.ymlandwindows-tester-bundle.ymlsign to"$f.sigstore.json"; a sync must not bring back.bundle(Scorecard counts.sigstore.json, not.bundle);scripts/ci/tests/test_tester_signature_extension.pyguards it.scripts/ci/run_affected_suites.pyprepends the suite venv'sbin/directory toPATHand setsVIRTUAL_ENVso subshells and hook tests execute inside the isolated suite environment;scripts/githooks/tests/test_install.pyprioritisessys.executableonself.env["PATH"].
Build backends of sdist-only lock pins (2026-10-05)¶
chore/close-rc-companions-row-lock-backends. One script, its record, one test, one suite path.
scripts/ci/sdist_only_pins.jsonis generated byscripts/ci/sdist_only_pins.py write(networked);scripts/ci/tests/test_sdist_build_backends_locked.pyreads it offline. A sync that bumps a lock pin regenerates the record in the same change; a record entry no lock pins fails the test as stale..github/test-suites.json: thetoolingsuite'ssource_pathsnamesrequirements/locks/(it namedtooling-tests.txtonly), so a lock change runs the suite. A sync keeps both.
Build backend lock and vmaf-tune host assumptions (2026-10-05)¶
fix/master-red-tests-2026-10-05. One lock input and two test files.
requirements/locks/package-build.incarriespoetry-core==2.5.0(reuse6.2.0 builds from source on Python 3.14); the six locks that include it were regenerated. A sync that regenerates every lock brings unrelated newer pins: keep this change'spoetry-coreentries.tools/vmaf-tune/tests/test_bisect_concurrency_cap.pypassesscore_runnerto every bisect call that reaches scoring;test_bbb_e2e_v14_bug_cluster.pydetects an NVIDIA GPU before the live NVENC probe. A sync keeps both.
Python lock declarations and the cosign verifier contract (2026-10-04)¶
fix/master-red-locks-cosign. Three declarations and one test.
ai/pyproject.tomlandtools/rc1-tester/pyproject.tomlcarryjsonschema==4.26.0indev(rc1-tester also pinsblack==26.5.1); their hash locks are regenerated withmake python-locks-write. A sync that regenerates other locks keeps this one's pins.test-publication-environment-binding.shlists four verifier targets fordocker-publish-operator-node.ymland requires one perpublish-*job; a sync that adds a published image adds its target toVERIFY_TARGETS.
macOS glibc probe and VIF signed sum (2026-10-04)¶
fix/master-red-macos-ubsan. Two test fixes and one cast.
core/test/test_icx_system_libm.pyprobes glibc through_gnu_libc_version()(never a bareos.confstr()); a sync keeps the guard,LibcDetectionTestfails without it.sigma_nsq + sigma1_sqis formed inuint32_tininteger_vif.c,x86/vif_avx2.c,arm64/vif_neon.c,cuda/integer_vif/vif_statistics.cuh,hip/integer_vif/vif_statistics.hipandmetal/integer_vif.metal(ADR-1601); a sync taking Netflix's text must not restore the signed sum.
The registry validator has no fallback (2026-10-04)¶
fix/model-registry-validator-no-fallback. One script and its tests.
ai/scripts/validate_model_registry.pyhas no structural fallback:_jsonschema_errorsreplaces_try_jsonschema_validateand_structural_fallback_validate, and a missingjsonschemais exit 2. A sync that brings either old name back failstest_no_structural_fallback_validator_remains.
Runner unit path and ADR-0931 status (2026-10-04)¶
fix/runner-unit-path-adr-0931-status. A unit file, one doc line and one ADR status.
- The runner unit's two paths assume the clone at
%h/dev/vmafx/vmafx;test_runner_systemd_unit.pyfollows them into the tree, so a sync that restores%h/dev/vmaffails it. - ADR-0931 is
Accepted(Phase 1 only) through a status-update appendix; its body is untouched. A sync keeps the status line and the index row in step.
make coverage-check reads gcovr, not lcov (2026-10-04)¶
fix/coverage-check-gcovr-target. One Makefile target, one script guard, one page.
- The
coverage/coverage-html/coverage-checkrecipes inMakefileuse gcovr andbuild-coverage/coverage.json;scripts/ci/coverage-check.shexits 2 on a non-gcovr input. A sync that restores lcov failstest_make_coverage_target.py. The CI job's recipe (tests-and-quality-gates.yml) is unchanged.
Two stale tiny-AI tests (2026-10-04)¶
fix/ai-suite-master-red. Tests only.
ai/tests/test_dnn_exporter_run_provenance.pywrites a real ONNX graph (the sidecar records its opset) andai/sidecar/tests/test_quickstart_contract.pyfinds the launch section by content. A sync keeps both; the contract must not go back to pinning a heading. The third failure on master (test_validation_report_provenance.py) is fixed by #2045.
Release and tester workflows verify before they publish (ADR-1595, 2026-10-04)¶
feat/ci-release-workflow-dry-run. Workflow triggers, one new workflow, one extracted script.
docker-publish-tester.ymlandwindows-tester-bundle.ymlcarry apull_requesttrigger whosepaths:equals theirpushlist, and theirvalidatejobs output the build matrix (matrix;build-matrix,verify-matrix). A sync that adds a leg adds it to that JSON, not to a literalmatrix:; one that adds a path adds it to both lists (test_pull_request_trigger_has_the_push_paths).supply-chain.yml(jobsbom) callsscripts/release/verify-mcp-sbom.sh; the inlinejqis gone.release-dry-run.ymlmirrors the release's image targets andvmaf-mcpcommands (test_the_images_it_builds_are_the_images_the_release_builds).
Tiny model cards quote their training data's terms (ADR-1570, 2026-10-04)¶
docs/model-dataset-terms. Docs and one contract test.
docs/ai/training-data.mdgains## Dataset terms(canonical verbatim quotes, one###per dataset). Fourteen cards gain## Training data termswith the same quotes;scripts/ci/tests/test_model_card_dataset_terms.pyfails when they differ. An upstream sync never touches these fork-local cards; a rebase that edits a card keeps the section whole.
Wheel force-include no longer repeats package files (2026-10-04)¶
fix/wheel-force-include-duplicates. Packaging metadata and one test.
ai/pyproject.tomlkeeps onlyconfigsinforce-include;dev-llm/pyproject.tomlhas none. A sync that brings an in-packageforce-includeback failstest_no_wheel_force_includes_a_file_its_packages_already_ship.
GPU dispatch variables: CUDA read at init, HIP removed (ADR-1571, 2026-10-04)¶
fix/hip-dispatch-env. CUDA, HIP, docs.
core/src/hip/dispatch_strategy.{c,h}(vmaf_hip_dispatch_supports(),VMAF_HIP_DISPATCH, theg_hip_features[]table) andcore/src/hip/AGENTS.d/dispatch-allowlist.mdare deleted; a sync must not restore them. HIP twin selection stays onVMAF_FEATURE_EXTRACTOR_HIP.feature_extractor.cpp::consult_cuda_dispatch()callsvmaf_cuda_select_strategy()for each CUDA extractor beforeinit()and fails with-ENOSYSon a strategy other than direct. Keep the call whenvmaf_feature_extractor_context_init()changes.test_gpu_dispatch_env_contract.pytiesdocs/usage/env-vars.md'sVMAF_*_DISPATCHrows to readers with a caller incore/src. No score, public C API or FFmpeg patch impact.
Notices and source for the rc.1 / rc.2 images (ADR-1578, 2026-10-04)¶
fix/published-rc-licence-companions. Licence tooling, a manual workflow, records and docs; no build or runtime change.
tools/rc1-tester/image/licensing.py:fetch_debianfalls back from the apt archive to snapshot.debian.org and then to Launchpad (launchpad_fetch);dpkg-foreigncomponents acceptpackage_patterns. A rebase keeps both, and keeps a failed fetch an error..github/actions/image-licence-artifactsgainssource-context(default.); the production callers do not pass it.tools/rc1-tester/image/published-rc/(data, recorded scans) and thepublished-rc-*records describe immutable images: never regenerate a scan from another commit than the release'ssource_commit.- Dry-run fix (
fix/rc-companions-dry-run):published-rc-licence-companions.ymlsetsWORKto${RUNNER_TEMP}/...in a first step (never a..path: upload-artifact refuses it; never the runner context in a job-levelenv).licensing.pydebian_specs()fetches adpkg-foreignpackage's source only when its copyright file declares a copyleft licence (copyleft_declared()); theintel-gpu-stack-aptcomponent ofpublished-rc-oneapi-imagelists Intel's MIT / BSD apt packages, which no archive holds. A rebase keeps both, and a failed fetch of a copyleft package stays an error.
zstd image layers and zopfli zips (ADR-1594, 2026-10-04)¶
build/compress-packages-2, stacked on ADR-1591. Workflows, the Windows zip builder, one lock and docs.
- Every workflow that pushes an image (
docker-publish-{tester,production,operator-node}.yml,dev-container-publish.yml,published-rc-licence-companions.yml) setsIMAGE_COMPRESSION: compression=zstd,compression-level=22,force-compression=true,oci-mediatypes=trueand pushes throughoutputs:ending in it; the tester'sload:steps are back to the shorthand. Droppingforce-compressionleaves cached and base layers gzip; droppingoci-mediatypesmakes the images unpullable by Docker. scripts/ci/build-windows-tester-bundle.py::pack()no longer useszipfile:write_zip()writes the recordszipfilewrites on Windows around zopfli streams. An upstream-style revert tozipfile.writestr()silently drops zopfli; the testtest_pack_deflates_every_entry_at_the_strongest_levelfails on it.requirements/locks/windows-tester-zip.{in,txt}(new) and the rc1-testerdevextra pin the same zopfli; the Windows workflow's recipe overlay checks the lock out withtools/rc1-testerandscripts/ci.- ADR-1591's exception for the dev container is gone.
Every published archive and image at its strongest compression (ADR-1591, 2026-10-04)¶
build/compress-packages. Packaging, workflows and their tests; no compiled code.
.github/workflows/docker-publish-{tester,production,operator-node}.yml: everydocker/build-push-actionstep exports throughoutputs:ending in,${{ env.IMAGE_COMPRESSION }}(nopush:/load:shorthand), and every call of.github/actions/image-licence-artifactspassescompression:. A rebase that adds an image build keeps both;scripts/ci/tests/test_package_compression.pyfails a workflow that pushes an image outside its tables.scripts/ci/build-macos-tester-bundle.shwrites<name>.tar.xzandmacos-tester-bundle.ymlglobs*.tar.xz; the guide's commands name.tar.xz.scripts/ci/build-windows-tester-bundle.py::pack()andtools/rc1-tester/src/vmaf_rc1_tester/bundle.py::_archive_zip()passcompressleveltowritestr(): aZipFile's level never reaches aZipInfoentry.docker/Dockerfile.nodeandlicensing.py fetch_git_archive()rungit archive --format=tar.gz -9;build-native-release-artifacts.shusesgzip -9n.dev-container-publish.ymlkeeps BuildKit's default level as a recorded exception that expires on 2026-12-31 (the test fails after that date).
Node FUSE tools and the chart's node.fuse / node.ebpf (ADR-1593, 2026-10-04)¶
feat/helm-ebpf-and-fuse. Fork-local files only.
docker/Dockerfile.node: thefuse-toolsstage and its twoCOPYlines inruntime-base. Keepmountandumountnext tofusermount3: libfuse 3.17 runs/bin/mountwhenever/etc/mtabexists, and Docker creates it.tools/rc1-tester/image/licensing.jsoncomponentfuse-toolsandscripts/ci/record-copied-debian-libs.sh(optional per-line destination) go with it.deploy/helm/vmafx/templates/node.yamlrenders its securityContext throughvmafx.nodeSecurityContext;templates/node-validate.yamlrefusesstorage.mode: mountwithoutnode.fuse. A sync that restores the plaintoYaml .Values.securityContextdrops the FUSE and eBPF capabilities.scripts/ci/tests/test_helm_node_contract.pyrenders mount mode withnode.fuse.
Python package licence metadata follows the shipped files (ADR-1560, 2026-10-04)¶
fix/package-licence-metadata-test. Packaging metadata and tests only.
python/test/setup_metadata_test.py:test_every_python_package_declares_the_licences_of_the_files_it_shipsreplacestest_all_python_packages_use_pep639_license_expression;tools/rc1-tester/tests/test_licensing_production.pylosestest_a_python_package_declares_the_licences_of_its_files. A sync that brings either old test back reintroduces the duplicate or the failing hard-codedBSD-2-Clause-Patent.python/pyproject.toml(Netflix-derived): an upstream sync keeps the fork'slicenseexpression andlicense-files = ["LICENSES/*"]; upstream'sBSD-2-Clause-Patentunderstates the compiled extension.python/LICENSES/is fork-added.- A new file under another licence in a package, or a new header the
vmafextension includes, fails the test until that package's expression andLICENSES/gain the identifier.
vmaf-tune splits auto, bisect, executor, per_shot, prefilter and score (ADR-1142, 2026-10-04)¶
refactor/vmaf-tune-baselined-modules. Python only, tools/vmaf-tune.
- Public names and signatures are unchanged except
auto.run_auto(src=...), which now acceptsNonefor a smoke plan (the CLI's--srcis optional with--smoke). Test seams stay module attributes (run_encode,run_score,_encode_and_score,_midpoint_lower_quality,_which); a rebase that resolves a conflict in one of these files keeps the helper split, never the old long body, because the HISS baseline no longer holds those functions. - New private helpers carry the old bodies:
auto._CellCtx/_build_cell/_late_short_circuits,bisect._BisectLoop/_bisect_iteration/_score_encoded,executor._ExecCtx/_execute_cell/_per_shot_cell/_saliency_cell,per_shot._per_shot_command,prefilter._make_objective,score._read_score_payload.plan.metadatakeeps its key order. test_help_texts_and_adr_refs.pynow scans these six modules for ADR citations;recommend.pyis still outside it (its own state row).
Tester image build stages copy tools/rc1-tester/src (2026-10-04)¶
fix/tester-image-prepare-build-src. Dockerfile and one test; no C change.
docker/Dockerfile.tester:vmaf-build,sycl-build,cuda-buildandhip-buildeach copytools/rc1-tester/srcnext totools/rc1-tester/image, becauseimage/prepare_build.pyimportsvmaf_rc1_tester.hw_facts. A rebase keeps all four lines;test_dockerfile_script_imports.pyfails when one is missing.
vmaf-tune returns the lowest-bitrate passing encode everywhere (ADR-1562, 2026-10-04)¶
feat/vmaf-tune-lowest-passing-bitrate-pick. tools/vmaf-tune and the Go pkg/recommend, pkg/fast, cmd/vmafx-tune. Fork-only code, no upstream Netflix files. A rebase must keep the single implementation of the rule (recommend.lowest_passing_row, Go lowestPassing): recommend, the interval-aware search, the live-mode pickers (cli._lowest_bitrate_passing, Go LowestBitratePassing) and the ladder's default sampler all call it, and the fast objective (objective_value / objectiveValue) is pinned in both languages by the same table. Do not restore _smallest_passing_crf, SmallestPassingCRF or the abs(vmaf - target) objective. A conflict in fast.py also moves the three helpers split out of its former over-length functions (_extract_sample, _verify_encode, _fast_production_result).
Integer VIF's residual variance goes through vif_sv_sq() (ADR-1561, 2026-10-04)¶
fix/integer-vif-sv-sq-defined. CPU, AVX2, NEON, CUDA and HIP.
- Netflix's
int32_t sv_sq = sigma2_sq - g * sigma12; sv_sq = (uint32_t)(MAX(sv_sq, 0));isuint32_t sv_sq = vif_sv_sq(sigma2_sq, g, sigma12);ininteger_vif.c::vif_accumulate_pixel(),x86/vif_avx2.c::vif_num_log256()andarm64/vif_neon.c::vif_num_log(), and in the CUDA (vif_statistics.cuh) and HIP (vif_statistics.hip) kernels. An upstream sync that touches those lines keeps the call: the upstream conversion is undefined below INT32_MIN (vif_sv_sq()returns the value x86 computes from it).sv_sqisuint32_t, sosv_sq + sigma_nsqstays an unsigned addition. core/src/feature/integer_vif_sv_sq.his in the CUDA and HIPdepend_fileslists ofcore/src/meson.build(test_device_target_header_dependencies).test_integer_vif_sv_sq_contract.pyfails on any C, C++, CUDA, HIP, Objective-C++ or Metal file undercore/src/featureorcore/testthat converts the raw difference to a signed integer, andtest_sycl_vif_exact_gain_contract.pynow expects the call invif_accumulate_pixel()and intest_sycl_integer_vif_math.c'sreference_terms(). No score, public API or FFmpeg patch impact.
Dev image pushed only into a private package (ADR-1564, 2026-10-04)¶
fix/dev-image-private-guard. CI only.
.github/workflows/dev-container-publish.ymlgains the step "Refuse to push unless the package is private" right after checkout; it runsscripts/ci/require-private-ghcr-package.sh VMAFx vmafx-dev-mcp. A rebase that edits the job keeps the step first, withoutif:orcontinue-on-error, and keeps every pushed tag in that package (scripts/ci/tests/test_require_private_ghcr_package.py).
libx265 two-pass cells are pass 1 at the CRF, then ABR (ADR-1565, 2026-10-04)¶
fix/vmaf-tune-x265-two-pass-crf. tools/vmaf-tune (encode.py, codec_adapters/x265.py) and the Go pkg/ffencode, pkg/corpus, pkg/codecadapter. Fork-only. A rebase keeps the two EncodeRequest fields (abr_bitrate_kbps / pass1_output, Go ABRBitrateKbps / Pass1Output), the argv swap (_with_abr_rate_control, withABRRateControl), the two drivers (_encode_abr_two_pass, runABRTwoPassEncode) and the adapter flag two_pass_abr_at_pass1_bitrate together; the libx265 adapter_version stays 2 in both languages. In Go's codecadapter, libx265Adapter is the single definition (after #2047) carrying TwoPassABRAtPass1Bitrate.
The controller keeps evicting nodes and requeues their jobs (2026-10-04)¶
fix/controller-requeue-evicted-node-jobs. Go controller only.
nodes.Registry:StartDetached,SetEvictionHook,ReaperRunning; the reaper's eviction isevictStale+notifyEvicted.Start(ctx)stays for callers with a long-lived context.queue.QueuegainsRequeueNode;provideNodeRegistrytakes the queue, starts the reaper detached and installsrequeueEvictedNode. A rebase that editsprovideNodeRegistrykeeps both.
vmafx-node starts the eBPF descriptor tracker on request (ADR-1539, 2026-10-04)¶
feat/node-ebpf-loader. Go node and cmd/vmafx-node/bpf; no C library change.
cmd/vmafx-node/bpf:rclone_bypass_stub.gois gone; bpf2go output (rclonebypass_bpfel.{go,o}, little-endian only) and a minimalvmlinux.hare committed;preflight.gois new; the loader uses the generated struct mirrors. A change torclone_bypass.bpf.cregenerates and commits both generated files with it.cmd/vmafx-node:ebpf_config.go,ebpf_linux.go,ebpf_other.go; the lifecycle invoke gains_ *ebpfBypassbetween the executor and the controller client.
vmaf-tune saliency accepts any frame height (2026-10-04)¶
fix/vmaf-tune-saliency-height-pad. Fork-local Python and Go only; no C source or public API change.
tools/vmaf-tune/src/vmaftune/saliency.py::compute_saliency_map()andpkg/saliency/saliency.go::ComputeMap()have noheight % 8guard: the tensor is zero-padded to a multiple of 32 and the map cropped back (_infer_frame_mask(),inferMask()). A sync must not restore the guard. The Python function now delegates to_MaskAccumulator,_open_saliency_session()and_infer_frame_mask()(HISS-04 split); keep that shape if upstream-style edits land in either file.
Go ladder scores each rung with its height's VMAF model (2026-10-04)¶
fix/vmaf-tune-go-ladder-model-per-rung. Fork-local Go and docs; no C source or public API change.
cmd/vmafx-tune/cmd/ladder.go::newLadderSampler()passescorpus.SelectVMAFModelVersion(width, height)tobisect.Y4MScoreParams.Model, whichrunVMAFXML()turns into--model version=...; an emptyModelstill leaves the flag off (compare,bisect).pkg/corpus/resolution.goandvmaftune.resolutionsharetools/vmaf-tune/tests/data/resolution_model_table.jsonas their golden table: a change to the height rule or the model names changes that file, both rules and both tests in the same PR.
mobilesal pads frames to a multiple of 8 (ADR-1540, 2026-10-04)¶
fix/saliency-frame-size. Fork-local tiny-AI code; upstream Netflix has no mobilesal extractor.
core/src/feature/feature_mobilesal.c: buffers sized forpw/ph,mobilesal_pad_plane()andmobilesal_cropped_mean(). Do not size the tensor from the frame alone again or average the padded rows.
vmaf-tune help texts and ADR references (2026-10-04)¶
docs/vmaf-tune-help-texts. Fork-local Python and docs only; no C source change.
ladder --crf-sweep, theautosubcommand help andcorpus --two-passbuild their text fromDEFAULT_SAMPLER_CRF_SWEEP,auto.ShortCircuitand the adapters'supports_two_pass; a sync must not reintroduce literal lists.- Many vmaf-tune ADR numbers were reassigned (for example ADR-0279 is now the libaom adapter, the conformal record is ADR-0393). When porting a comment or help string that cites an ADR, check the record's title;
tests/test_help_texts_and_adr_refs.pyfails on an unrelated record in the help, the usage pages and the AGENTS.d pages. cli._fast_proxy_encoder_slotaddsproxy_encoder_slotto thefastJSON.
Helm GPU resource name follows the Intel driver (ADR-1547, 2026-10-04)¶
fix/helm-gpu-resource-name. Helm chart only.
_helpers.tpl:vmafx.gpuResourceNameis the one resolver;vmafx.gpuResourcegates it ongpu.enabled;vmafx.gpuResourceKeyis gone (node.yamlusesvmafx.gpuResourceName). New valuesgpu.intelDriver,gpu.resourceName.networkpolicy.yaml:allow-controller-to-nodetakescontrollerToNode.nodePort | default node.grpcPort; the values file no longer setsnodePort: 50051.
Codec-block encoding from the sidecar (ADR-1558, 2026-10-04)¶
fix/fr-regressor-v3-codec-block. Fork-local DNN code and model metadata.
core/src/dnn/model_loader.{c,h}:VmafCodecBlockEncoding,parse_codec_preset_norm()/parse_codec_crf_norm(),vmaf_dnn_codec_block_fill_encoded()(the old fill is a wrapper).core/src/libvmaf.c:vmaf_ctx_dnn_set_codec_context()passes the sidecar's encoding anddnn_warn_constant_preset()reports an ignored preset.model/tiny/fr_regressor_v3.jsongains the fivecodec_*keys;ai/scripts/train_fr_regressor_v3.pywrites them (crf_range).
Tiny-model metadata checked against the graphs (ADR-1546, 2026-10-04)¶
fix/tiny-model-metadata. Fork-local AI tooling and model metadata; no libvmaf change.
ai/scripts/validate_model_registry.pygains_check_graph_metadata()and helpers on top of the newai/src/aiutils/onnx_signature.py.model/tiny/:nr_metric_v1opset 18,fr_regressor_v2notes andtraining.hidden/depth, the five ensemble seed sidecars rewritten,transnet_v2.jsonoutput_nameoutput_0, schema texts andrelease_url. A re-export must keep these in step with its graph or the required validator job fails.ai/scripts/build_calibration_set.pyis removed.
vmaf-tune adapter-aware coarse window, ladder workdir, auto geometry, uncertainty note (2026-10-04)¶
fix/vmaf-tune-crashes-dead-flags. Fork-local Python only; no C source change.
corpus.coarse_to_fine_searchis a plain function that resolves and validates the window (_resolve_search_window) before it returns the row generator. Do not turn it back into a generator: theValueErrorwould surface mid-write.crf_min/crf_maxdefault toNone, meaningcoarse_search_window(encoder).ladder.SamplerResourcescarries--workdirand the decode semaphore into the default sampler;CorpusOptions.decode_semaphoregates_decode_job_reference.cli._auto_execute_geometryfeedsrun_plan(**geometry); raw YUV without--width/--heightexits 2.cli._uncertainty_unavailable_noteis the only path that answers--with-uncertaintywithout intervals.
TransNet V2 runs upstream's windows (ADR-1527, 2026-10-04)¶
fix/transnet-v2-load. Fork-local tiny-AI code; upstream Netflix has no transnet_v2 extractor.
core/src/feature/transnet_v2.c: windows ofpredict_frames(), a flush callback andVMAF_FEATURE_EXTRACTOR_TEMPORAL; thumbnails in 0..255; the output bound by position. Do not restore the last-slot readout or the 0..1 scaling.core/src/dnn/dnn_api.c:setup_luma_fast_path()probes ranks up toVMAF_DNN_PROBE_MAX_RANK(8) and treats-ERANGEas "no fast path".model/tiny/transnet_v2.jsonandai/scripts/export_transnet_v2.py:output_nameisoutput_0.
vmaf-tune model overrides, cache key v2, QSV chain on encodes, ladder codec strings (2026-10-04)¶
fix/vmaf-tune-silent-overrides. Fork-local Python and Go only; no C source change.
corpus._sweep_score_modelresolves the score model once per job (resolution_aware, thenneg, then HDR); the CLI setsresolution_aware=Falseonly for an explicit--vmaf-model. Do not bring back a per-cell selector that ignoresneg.cache.CACHE_VERSIONis 2 andcache_keytakespasses, the sample-clip window andsettings;corpus._cell_cache_keyfills them. Every adapter carriesadapter_version.- The QSV chain is
_qsv_common.qsv_device_init_args()(filter deviceqsv_dev) plusQSV_UPLOAD_FILTER, applied byencode.build_ffmpeg_commandandcompare._hw_probe_argv; the Go twin ispkg/hwdevice, used bypkg/ffencodeandpkg/encoder. Change both together (cmd/vmafx-tune/AGENTS.mdinvariant 24). ladder.emit_manifest(..., codec_for=)names codecs only through a resolver (vmaftune.codec_strings); there is no default codec string.- AMF
extra_params()returns();pkg/codecadapter/testdata/python_adapters.jsonrecords an empty AMFextra.
vmafx-node reads job sources through pkg/storage (ADR-1526, 2026-10-04)¶
feat/node-storage-wiring, stacked on feat/node-controller-client. Go node, pkg/storage, pkg/libvmaf, Helm chart and go-ci.yml; no C source change.
cmd/vmafx-node/executor_inputs.go(scoreJob) andstorage_config.goare new;provideExecutorreturns(*Executor, error);executeScoringcallsscoreJob. Keepstorage.Open(notstorage.New) inprovideStorage.pkg/libvmaf:readers_unix.go/readers_other.goaddScoreReaders;ScoreOnBackendnow uses the extractedscoreOutputFileandwithScoreDeadlinehelpers.pkg/storage:ParseMode,Open,IsHTTP,Config.MountRoot;Newis deprecated but unchanged.- Helm:
storage.modeenumhttp-serve | mount | auto,storage.mountRoot,VMAFX_RCLONE_CONFIGonly with the Secret.go-ci.ymlinstallsrcloneandfuse3for the real-rclone tests.
vmafx-node pulls jobs from the controller (ADR-1524, 2026-10-04)¶
feat/node-controller-client. Go node, pkg/libvmaf and Helm chart; no C source change, no upstream file touched.
cmd/vmafx-node/controller_*.goandbackoff.goare new;main.goaddsprovideControllerClientand the*controllerClientargument of the lifecycle invoke (stop order: gRPC, client, feedback, scorer). Keep the argument when the invoke is edited (cmd/vmafx-node/AGENTS.mdinvariant 13).pkg/libvmaf.Scorer.ScoreOnBackendadds--backend;Scoredelegates with an empty backend and is unchanged for its callers.deploy/helm/vmafx: thevmafx.controllerAddrhelper is gone;node.yamlrendersVMAFX_CONTROLLER_ADDRonly fromnode.controllerAddr;networkpolicy.yamlgainsallow-node-to-controller. A rebase that brings the helper back reintroduces a default pointing at a Service the chart does not deploy.
Controller tenant registry and chart auth guards (ADR-1519, 2026-10-04)¶
feat/controller-tenant-config. Go controller and Helm chart; no libvmaf change.
cmd/vmafx-controller/auth/tenants.go(registry, resolution) andcmd/vmafx-controller/tenants/(file and Kubernetes sources, refresher) are new;auth.Config.Tenantsselects them andConfig.validateModerefuses a tenant source next toDisabledor the global provider settings. HTTP and gRPC verify through oneMiddleware.verifyBearer.cmd/vmafx-controller/tenant_config.goreadsVMAFX_AUTH_TENANTS_*and loads the tenants while the fx graph is built (a failure stops startup).- Chart:
templates/auth-validate.yaml,templates/controller-tenant-rbac.yamland theallow-server-to-apiserverNetworkPolicy are new;deployment.yamlpasses the tenant source instead of the global provider in tenant mode;tenant-crd-config.yamltestsenabledwithhasKey.
Controller job reads and node sessions are tenant-scoped (ADR-1522, 2026-10-04)¶
fix/controller-tenant-filter. Go controller only; no libvmaf change.
queue.Queue.ListAllis gone;ListByTenantputstenant_id = ?in the SQL.PullWorktakes the tenant;ReportResulttakes aqueue.Report(node, tenant, job, result, orphan predicate), itsUPDATEcarriesAND assigned_node = ?, and only an orphanedRUNNINGjob of the same tenant may be adopted (ErrNotAssignedotherwise). A change to the jobs schema or to these queries keeps the node, status and tenant guards.nodes.Registry.Register/Heartbeat/ValidateSessionandscheduler.Assigntake the caller's tenant;Node.TenantIDrecords it.auth.AssertTenantOwnsandAssertHTTPTenantOwnsname no tenant in their error.
Controller gRPC calls are authorised per method (ADR-1518, 2026-10-04)¶
fix/controller-grpc-roles. Go controller only; no libvmaf change.
cmd/vmafx-controller/auth/policy.goand the interceptors ingrpc_interceptor.goauthenticate and authorise in one function (admitGRPC);auth.Config.MethodRolesis the policy and a method without an entry is refused.cmd/vmafx-controller/grpc_roles.goholds the controller's table; an RPC added tocontroller.protoorvmafx.protoneeds an entry there in the same change (TestEveryServedRPCHasARolePolicy).cmd/vmafx-controller/auth/authtestmints the RS256 tokens of the auth and controller tests; the auth tests'fakeIssuersigns through it.
The GPU images carry their licences and source (ADR-1517, 2026-10-04)¶
fix/prod-licensing-gpu-images. Fork-added build and packaging files only; no libvmaf source change.
docker/Dockerfile.production-gpubuildsfinal-cuda13,final-rocm10andfinal-oneapi2026onRELEASE_BUILDER_BASE/ONEAPI_*(Debian 13); it no longer declaresCUDA_BUILDER,CUDA_RUNTIMEorROCM_RUNTIME, andfinal-cpuis gone (the CPU image isdocker/Dockerfile.production). Each final target copies the receipt of<variant>-licence-check; keep that line on a rebase (test_every_published_production_target_passes_a_licence_check).- The runtimes hold only the vendor files
vmafloads: none for CUDA (the build fails on an NVIDIA file or NEEDED entry),hip-runtime.jsonin/usr/local/lib/rocm,sycl-runtime.jsonin/usr/local/lib/intel. A change to one of those lists changes the tester and the production image together. licensing.json:production-cuda-image,production-rocm-imageandproduction-oneapi-imagetake components by reference ({"from": ..., "id": ...}with the record'srewrite); edit a vendor component in its tester record, never a copy.licensing.pyexpand_shared()resolves the references inload_manifest().docker-publish-production.yml: the GPU jobs passVMAFX_SOURCE_COMMIT/VMAFX_IMAGE_TAG, run.github/actions/image-licence-artifactsbefore the disk cleanup, and the oneAPI smoke checks/usr/local/lib/intel/libur_adapter_*.scripts/release/tests/test-docker-image-runtime-contract.shholds the oneAPI staging and the UMF entry ofsycl-runtime.json.
HIP twins run on the state's device (ADR-1523, 2026-10-04)¶
fix/hip-device-index. Fork-local HIP runtime and twins; no upstream file.
core/src/hip/common.c:vmaf_hip_context_new()checks its index and callshipSetDevice();vmaf_hip_device_count()returns a negative errno for a runtime failure (0 only forhipErrorNoDevice); newvmaf_hip_state_device_index()/vmaf_hip_state_bind().VmafFeatureExtractorgainsint hip_device_indexunderHAVE_HIP, filled byset_fex_hip_device()inlibvmaf.c'sfex_ctx_bind_backends();read_pictures_hip_frame_begin()andflush_context()callvmaf_hip_state_bind(). An upstream change toflush_context()keeps the HIP block at its top.- Every
vmaf_hip_context_new()call undercore/src/feature/hip/passesfex->hip_device_index(CAMBI throughcambi_hip_setup_device(), SpEED throughSpeedHipConfig.device_index). A new HIP twin does the same;test_hip_device_index_contractfails on a literal index.
Feature-vector tiny models own their inputs (ADR-1520, 2026-10-04)¶
fix/tiny-model-missing-features. Fork-local DNN code; upstream Netflix has no tiny-model path, so a sync does not touch these hunks.
core/src/libvmaf.c:dnn_attach_feature_vector()registers the input extractors (dnn_request_input_features()),flush_context()callsdnn_flush_feature_vector()after the backend flushes, andvmaf_ctx_dnn_run_frame()no longer scores rank-2 models. A rebase that touchesflush_context()must keep that call after the CUDA and SYCL flushes and beforevmaf->flushedis set.core/src/dnn/model_loader.c:vmaf_dnn_codec_block_fill()finds"unknown"by name (codec_block_slot()); do not restore the last-slot default.core/tools/vmaf.cpp:apply_tiny_codec()requires--tiny-crf.
adm_hip computes AIM on the device and is dispatched (ADR-1525, 2026-10-04)¶
fix/adm-hip-aim-dispatch. Fork-local HIP twin; the CPU integer_adm.c is unchanged.
core/src/feature/hip/integer_adm/adm_cm.hipgainsadm_cm_aim_line_kernel_4andi4_adm_cm_aim_line_kernel(ports of the CUDA ADR-0746 kernels) and shared helpers (i4_decouple()/s0_decouple(),i4_csf()/s0_csf(),cm_cube(),adm_asr()); the DLM kernels use the same helpers with unchanged values.integer_adm_hip.c:RES_BUFFER_SIZE24 to 36 (DLM, denominator, AIM),adm_skip_aim, the scale-0 launch sharesadm_cm_s0_launch()(shifts from the CPU'sadm_cm_ctx_init()),aim/adm3claimed and.flags = VMAF_FEATURE_EXTRACTOR_HIP.AdmBufferHipgainsadm_aim_cm[4]andinteger_adm_hip.htheAdmCmShiftsHipkernel argument.- An upstream change to the CPU CM (
adm_cm_ctx_init()/i4_adm_cm_ctx_init()withmeasure_aim,adm_csf_cols()) changes these kernels in the same PR;test_hip_adm_exactholds every output to==.
Release files carry their notices (ADR-1513, 2026-10-04)¶
fix/prod-licensing-release-assets. Release tooling only; no libvmaf change.
scripts/release/build-native-release-artifacts.shwrites and checks the notices ofmodels.tar.gz(which now also holdslicenses/) and of the release files (THIRD_PARTY_NOTICES.txt,licenses.tar.gz) withlicensing.py(kindsrelease-models,release-native), after the provenance stamp and before the verify step. The script runs offline in therelease-buildstage: the two kinds use only texts from the repository.supply-chain.yml:attach-to-releaserequires the two notices files;provenanceandmcp-provenanceattest the SPDX SBOMs withactions/attest.
Model attribution in REUSE.toml (ADR-1513, 2026-10-04)¶
fix/prod-licensing-model-attribution. REUSE.toml and model documentation; no code change.
REUSE.tomlannotatesmodel/tiny/lpips_sq.*(BSD-2-Clause AND BSD-3-Clause),model/tiny/transnet_v2.*(MIT),model/tiny/fastdvdnet_pre.*(MIT, the upstream LICENSE's copyright line) and the fork's root models and cards (model/predictor_*,model/konvid_mos_head_v1*,model/*_card.md, 2026 Lusoris). The annotations sit aftermodel/tiny/**because REUSE applies the last matching table. A new tiny model with upstream weights adds its annotation in the same PR (test_model_annotations_match_the_registry).- An upstream sync that adds a Netflix model at
model/root named likepredictor_*would be mis-attributed by the root annotation; none exists.
vmafx-tune-go scores Y4M, scales ladder rungs, exits 2 on usage errors (2026-10-04)¶
fix/vmafx-tune-go-cli-contract. Go CLI and packages only; no C source change.
pkg/bisect/score_y4m.go(Y4MScorer) is the scorer ofcompareandladder;VMAFScoreFuncwraps it. A sync must not bring back a scorer that handsvmafa non-Y4M file without geometry flags.- A ladder rung's encode filter (
bisect.Params.EncodeExtraArgs) and its reference decode both usebisect.ScaleFilter; QSV's upload joins that chain throughappendVideoFilterinpkg/encoder/hardware.go. probeBitrateKbpsreadsstream=bit_rate:format=bit_rate; the stream entry alone isN/Afor Matroska.newRootinstallsuseUsageExitCodeandmarkCommandFlagsRequiredvalidates fromPreRunE(cmd/vmafx-tune/AGENTS.mdinvariant 30);autochecks--srcinvalidateAutoFlags, not withMarkFlagRequired.
The tester image's SBOMs are attested on the platform manifests (2026-10-04)¶
fix/tester-sbom-attest-platform-digest. CI workflow and docs; no source change.
docker-publish-tester.yml'spublishjob selects each SBOM's subject from the merged index (Select the platform manifests the published index lists) and ends withVerify the attestations on the published digests. A rebase keeps both; taking the per-arch digest from the build job's artifact assubject-digestbrings back an attestation no tester can find.
The macOS tester bundle job installs the Metal compiler (2026-10-04)¶
fix/tester-bundle-metal-toolchain. CI workflow only; no source change.
.github/workflows/macos-tester-bundle.yml'sbuildjob has the stepInstall Metal compiler toolchain(xcodebuild -downloadComponent MetalToolchain, thenxcrun -sdk macosx metal --version) after the brew step, with the same wording asbuild.ymlandlibvmaf-build-matrix.yml. A sync or rebase keeps all three; dropping the step brings back themissing Metal Toolchainfailure on runner images without the component.
The site search covers user pages only (ADR-1512, 2026-10-04)¶
docs/site-search-scope. Documentation, its generator and checks.
mkdocs.ymllistsmaterial/meta;docs/adr/.meta.yml,docs/research/.meta.ymlanddocs/changelog-archive/.meta.ymlexclude their pages from the search,docs/adr/by-tag/.meta.ymland the front matter ofdocs/adr/_index_fragments/_header.md,docs/research/README.mdand the generatedtitles.mdpages bring the indexes back. Keep the front matter ofdocs/rebase-notes.mdanddocs/state.md(search: exclude: true) at the top of those files when resolving a conflict there; new entries go below it.docs/adr/titles.mdanddocs/research/titles.mdare generated byscripts/docs/generate-record-titles.py; on a conflict take either side and runmake docs-fragments-write.scripts/docs/check_search_scope.pyruns after the strict build in both docs CI jobs; a new record directory that must stay out of the index gets a.meta.ymland a pattern in that script.
Go service and node images carry their licences and source (ADR-1514, 2026-10-04)¶
fix/prod-licensing-go-images. Fork-added build and packaging files only; no libvmaf source change.
docker/Dockerfile.operator,Dockerfile.go-server,docker/Dockerfile.node: the published targets (operator,go-server,node-cpu) copy the receipt of their licence-check stage; the go-builder stages runlicensing.py scan-goandgo-licences. A change that adds a Go program, a module with an unusual licence file or a copied library changestools/rc1-tester/image/licensing.jsonin the same PR.docker/Dockerfile.node: never re-add--enable-nonfreeto the FFmpeg configure line (the binary then declares itself unredistributable); theffmpeg-builder-cpustage writes/ffmpeg-source/(the patched tree, the configure line, the patch series) and records the copied libraries withscripts/ci/record-copied-debian-libs.sh. An FFmpeg patch refresh keeps those steps aftermake install.build-config.envRCLONE_VERSIONreplacesRCLONE_IMAGE; therclone-binstage runsgo install github.com/rclone/rclone@${RCLONE_VERSION}and dropsRCLONE_VERSIONfrom the environment before running rclone (rclone readsRCLONE_*variables as flags).- The
libvmaf.so*copies in Dockerfile.go-server and Dockerfile.node usefind -maxdepth 1 \( -type f -o -type l \): the glob took Meson's object directorylibvmaf.so.3.0.0.p/into the images.
The ADR navigation is collapsed behind the indexes (ADR-1510, 2026-10-04)¶
docs/site-adr-nav. Documentation, its generator and tests; no source change.
mkdocs.yml'sADRsentry is three static lines (adr/README.md,adr/0000-template.md,adr/by-tag/index.md). TheADR-NAV-GENERATEDblock andscripts/docs/generate-adr-nav.share gone, with their two lines in theMakefile'sdocs-fragments-checkanddocs-fragments-write. A branch from before this change that regenerates the block, or an upstream sync that touchesmkdocs.ymlaround it, takes this side and drops the block; never re-add the generator.scripts/docs/tests/test_generators.py::test_adr_navigation_is_collapsedfails when an ADR page or the block comes back into the navigation.- ADR branches no longer touch
mkdocs.yml; a conflict there on an ADR branch is a leftover of the old generator and resolves to this side.
Documentation diagrams are figure specs (ADR-1508, 2026-10-04)¶
docs/site-diagrams. Documentation and site configuration only.
mkdocs.ymlliststools/figures/mkdocs_hook.pyunderhooks:and no longer has a Mermaid custom fence;exclude_docsholds/figures/and its patterns avoid**(the figuresourcescheck exits 2 on them). Keep all three on a rebase.docs/figures/<slug>.tsanddocs/assets/figures/<slug>.*move together: after editing a spec, or when a cited symbol is renamed in the code, runnode tools/figures/build.mjs buildand commit the outputs. On a conflict indocs/assets/figures/take either side and rebuild.- The pages that hold a
figurefence (ai/overview.md,backends/index.md,development/cross-backend-gate.md,development/release.md,usage/tester-image.md,architecture/phase4b-distributed-platform.md,development/operator.md,server/controller.md) keep the fence when their text is rewritten; an ASCII diagram must not come back beside it.
Production CPU and server images carry their licences and source (ADR-1513, 2026-10-04)¶
fix/prod-licensing-cpu-images. Fork-added build and packaging files only; no libvmaf source change.
docker/Dockerfile.production:cliandservercopy the receipt ofcli-licence-check/server-licence-check, so neither builds without the licence check (artifact kindsproduction-cli-image,production-server-imageintools/rc1-tester/image/licensing.json);cli-source-export/server-source-exportare published as<tag>-sourceand<tag>-server-source. An upstream sync never touches this file; a change that adds files to either image changes the record in the same PR.python-depsbuilds the two wheels in/build-venvand installs only the runtime lock and the wheels into/venv; do not bring the build lock back into the runtime venv (the licence check would also refuse its dist-infos).mcp-server/vmaf-mcp/pyproject.tomlandtools/vmaf-tune/pyproject.tomldeclare the union of their files' SPDX identifiers and shipLICENSES/*(copies of the repository texts, held byte-identical by a test)..github/actions/image-licence-artifactsis the one implementation of the per-platform SPDX attestation and the source image push for production images.- The CPU and server jobs' recovery dispatch also takes
tools/rc1-tester/image/and.github/actions/from the recipe commit (ADR-1347): the licence record is part of the build recipe.
rule-enforcement.yml swallows no exit status (2026-10-04)¶
ci/rule-enforcement-hiss07. CI and baseline only; no libvmaf change.
- The sixteen baselined
|| true(HISS-07) are gone: thegit fetchof the base and head SHAs fails the step; the ADR-collision steps filter throughkeep()(grep status 1 = no match, the only mapped status) underset -euo pipefailand read their lists from a captured variable, not a process substitution whose failureset -enever sees; the open-PR table comes from a capturedpython3whose failure fails the step. A sync that re-adds|| truere-adds the baseline rows; keep this side. - The
release-script-contractjob runscheck-vcs-version-not-bare-sha.shand its test (until now onlymake lint-shran them). .standards-baseline.jsonand the README count: 495 to 479, re-recorded withpraetorctl baseline --record; on conflict re-record at the tip.
Documentation charts from repository data (ADR-1508, 2026-10-04)¶
docs/site-charts. Documentation, its generator and two CI steps; no source, score or build change.
scripts/docs/generate-charts.pyownsdocs/charts/<slug>/data.json,docs/assets/charts/,docs/javascripts/vendor/vega/vega-bundle.js(and its hash in that directory'svendor.json) and the blocks between<!-- >>> CHART <slug>and<!-- <<< CHART <slug> -->ondocs/index.md,docs/backends/index.md,docs/development/upstream-parity.mdanddocs/development/netflix-benchmark-baselines.md. On a conflict in any of them take either side and runmake docs-fragments-write; never hand-edit a render or a block. Keep the sentinel lines when a page's text is rewritten.- A change to
scripts/ci/exact_twins.d/,LIBM_TWINS,scripts/ci/upstream_parity.d/ortestdata/scores_*_576.jsonchanges a chart: regenerate in the same change, ormake docs-fragments-checkfails. - The renders depend on vl-convert-python's version (
docs/requirements-lock.txt); a bump re-renders every chart and the bundle, andTHIRD-PARTY-LICENSES.txtis collected again for the new Vega releases.
lint-and-format.yml swallows no exit status (2026-10-03)¶
ci/lint-and-format-hiss07. CI and baseline only; no libvmaf change.
- The seven
|| trueof the file (HISS-07) are gone.git difffailures fail the step; the markdown filters run throughkeep(), which maps only grep's status 1 (no line matched) to success, underset -o pipefail. A sync that re-adds|| truere-adds the baseline rows; keep this side. - The
pre-commitjob runsscripts/githooks/tests/test_install_hooks_env.pyaftertest_install.py. .standards-baseline.jsonand the README count: 502 to 495, re-recorded withpraetorctl baseline --record; on conflict re-record at the tip.
Hook environments install outside the commit's git environment (2026-10-03)¶
fix/hooks-install-envs-outside-commit-env. Hooks and tests only; no libvmaf change.
lefthook.yml: bothframework-hooksentries runpre-commit install-hooksunderenv -u GIT_INDEX_FILE -u GIT_DIR -u GIT_WORK_TREE -u GIT_OBJECT_DIRECTORYbeforepre-commit run/pre-commit hook-impl. A sync or a lefthook rewrite keeps that line and its place before the run: without it a node hook install in a linked worktree rewrites the worktree's index.scripts/githooks/tests/test_install_hooks_env.pyreads the block fromlefthook.yml, so it fails when the line goes.
The AMD GPU tester image (ADR-1511, 2026-10-03)¶
feat/tester-kit-hip, stacked on feat/tester-kit-cuda. Fork-only tester tooling; no libvmaf source changes.
docker/Dockerfile.testergains thehip-*stages and targetfinal-hipafter the CUDA stages, andARG ROCM_BUILDER(mirrored frombuild-config.envbyscripts/ci/check-base-image-single-source.sh).HIP_GFX_TARGETSthere is the image's offload-target list and must equal the ROCm image'sshare/therock/dist_info.json(the build checks it; a ROCm bump changes both); the build reads the targets back from theHIP HSACO targets:message ofcore/src/meson.build(prepare_build.py hip-targets), so renaming that message fails the image build.scripts/ci/install-rocm-from-image.shgains--keep-docs(keepsshare/doc); without the flag it behaves as before.prepare_build.py:stage_intel_runtime()becamestage_vendor_runtime()(per componentdest, per speclicence_dir,credistoptional); the commandsintel-runtimeandrocm-runtimeboth call it.licensing.py: components may carryvendored_libraries(bundled libraries; copyleft ones keyed by ELF build ID to source archives), a source archive may be a git tree (git+ fullcommit, fetched by that commit and packed withgit archive; thehip-source-fetchstage installs git), andgenerated_build_filesgains the rule forsrc/*_hsaco.c. A HIP kernel added upstream needs a unique file name undercore/src/feature/hip/.- The workflow's GPU matrix gains the leg
hip.
The NVIDIA GPU tester image (ADR-1509, 2026-10-03)¶
feat/tester-kit-cuda. Fork-only tester tooling; no libvmaf source changes.
docker/Dockerfile.testergains thecuda-*stages and targetfinal-cudabetween the SYCL stages and the CPU image'sassembled;finalstays the last stage and the default target. Thecuda-runtimestage fails on any file named like an NVIDIA library and thecuda-buildstage on aNEEDEDentry naming one: the image ships no NVIDIA file. A sync that makeslibvmaflinklibcudart(or any NVIDIA library) breaks the build until ADR-1509 is revisited.- The image's CUDA targets come from the
CUDA gencode = [...]andFound CUDA version = ...linescore/src/meson.buildprints (prepare_build.py cuda-targets); renaming those messages fails the image build. docker-publish-tester.yml: the Intel jobsbuild-sycl/publish-syclare now the matrix jobsbuild-gpu/publish-gpu(kitssycl,cuda); digests pass through the artifacttester-<kit>-digests.licensing.json: artifactcuda-image, and agenerated_build_filesrule of the new formcompiled_fromforsrc/*.fatbin.c(the licence of the.cusource a kernel object was compiled from). A CUDA kernel added upstream needs a unique file name undercore/src/feature/cuda/, or the scan refuses it.tools/rc1-tester/tests/test_sycl_rows_contract.pyis nowtest_gpu_rows_contract.pyand coverscuda-rows.jsontoo: every CUDA row holds every parity-gate feature, so a feature added upstream changes the row map.
The documentation site's design layer (ADR-1508, 2026-10-03)¶
docs/site-design. Documentation and site configuration only; no source, score or build change.
docs/stylesheets/vmafx.cssis the only design layer.mkdocs.ymllists it underextra_css, setsprimary: customandaccent: customin both palettes andfont: false. A sync or a content edit must not bring back a Material colour name, afont:block (it makes the browser load Google Fonts) or acustom_dirtemplate override: Zensical'sclassicvariant renders the same HTML only while the design stays in CSS.docs/assets/fonts/{inter,jetbrains-mono}/hold the upstream fonts subset to the ranges in eachvendor.json, the licence and the manifest;scripts/docs/vendor_fonts.pywrites them from the release archive andscripts/docs/check_vendored_assets.py(inmake docs-fragments-check) fails on a changed, missing or unlisted file. Update a font with the writer, never by hand.docs/index.mdkeeps thevx-hero,vx-steps,vx-backendsandvx-topicswrappers and itshide:front matter when its wording changes.
Open CodeQL alerts fixed in code: header guards, SpEED products, bit-identity tests (ADR-1502, 2026-10-03)¶
fix/codeql-open-alerts. No score moves; every library and tool object file is byte-identical before and after (GCC 16.2.1 and clang 23.1.1, x86-64 and aarch64, -Db_lto=false).
core/src/feature/adm.handmotion.hhave include guards (ADM_H_,MOTION_H_);adm_tools.handadm_csf_tools.hopen their existing guard before the includes and theM_PIfallback instead of after them. Upstream has no guard in the first two and the late one in the others; a sync keeps the fork's placement (CodeQLcpp/missing-header-guard).adm_options.hnames upstream's commented-outADM_OPT_DEBUG_DUMPswitch in prose; a sync does not bring the/* #define ... */line back. The replacement keeps the header's line count, so the dismissed alert on the enum below it (85) keeps its line.speed.c::get_speed_score()andspeed_internal.c::si_gpu_speed_score()writelog2((double)((1 + nn_floor) * sigma_nn)): the conversion of the float product's result, as upstream's implicit promotion does (ADR-1477). Never cast an operand.core/test/test_speed_upstream_form.c(Netflix statements, fork-held):run_netflix_value_tests()has one body in both meson variants; the glibc / foreign-libm difference is the file-scopeuf_netflix_values_skip_reason. The bit helperuf_bits()isvmaf_test_bits_f32()fromcore/test/float_bits.h.
Tester packages carry their licences, an SBOM and their source (ADR-1503, 2026-10-03)¶
fix/tester-artifact-licensing. Fork-added files only, except two licence tags: core/src/feature/mkdirp.h and mkdirp.cpp now say MIT, the terms of Stephen Mathieson's code that Netflix/vmaf carries with an "MIT licensed" comment and no SPDX tag. An upstream sync that touches them keeps the MIT tag.
tools/rc1-tester/image/licensing.jsonrecords every component of the tester image and the macOS bundle;licensing.py checkfails their builds on any file it does not claim. A change todocker/Dockerfile.tester, the bundle script, the Python lock or a base image that adds files changes the record in the same PR (ADR-1503).docker/Dockerfile.tester:finalcopies the receipt of thelicence-checkstage;stripskips*.libs/(vendor libraries ship unmodified, the check compares them with the wheelRECORD);source-exportis published as<tag>-source.REUSE.tomlrecordsmodel/other_models/brisque_live.modelandNOTICE-brisqueasLicenseRef-LIVE-BRISQUE(text inLICENSES/), and the HACL* notice the tester packages ship as MIT.
The Intel GPU tester image and report schema 3 (ADR-1505, 2026-10-03)¶
feat/tester-kit-sycl. Fork-only tester tooling; no libvmaf source changes.
docker/Dockerfile.testergains the stagessycl-build,sycl-licences,sycl-runtime,sycl-refs-genandfinal-sycl, inserted beforefinalso thatfinalstays the default target. Thesycl-buildstage builds in/opt/vmafx/src:test_sycl_kernel_scratchbakes its ratchet path (core/test/../src/sycl/scratch_ratchet.txt) from the source directory, and the runtime stage provides the file at that path. A change to that test'sVMAF_SYCL_SCRATCH_RATCHETdefine changes the image (the build greps the path).- The report is schema 3 (
gpusection,hw_gpu.py,hw_sycl.py,hw_l0probe.py);hw_gate.run_gate(),hw_suites.run_unit_tests()(args, environment, scratch directories, per-test results) andprepare_build.py(suite:lines,twins,intel-runtime) are generalised, Metal keeps its wrappers. A Meson test of thegpuorsyclsuite added upstream of a sync is picked up by the image automatically; a Python test is listed as left out. tools/rc1-tester/image/sycl-rows.jsonnames the SG16 row's tests; renaming one of them failstools/rc1-tester/tests/test_gpu_rows_contract.py.- The licence record (
tools/rc1-tester/image/licensing.json, ADR-1503) gains the artifactsycl-image, a component kinddpkg-foreign(vendor packages without a Debian copyright file or Debian source) and pinnedfetched_texts;licensing.pyhonours both incheck_dpkg(),debian_specs()andfetch_texts(). The Dockerfile'ssycl-licence-checkandsycl-source-exportstages mirror the CPU image's.
The float_psnr twins add each row's exact sum in the CPU's order (ADR-1499, 2026-10-03)¶
fix/float-psnr-exact-past-2-53. Fork-only device and host code; float_psnr.c and its SIMD rows are untouched.
- The CUDA, HIP and SYCL kernels reduce 256 x 1 blocks / work-groups (were 16 x 16), so each partial sum is a segment of one row; the HIP kernel writes one
uint64per block (was twouint32halves). The CUDA geometry moved intocuda/float_psnr_cuda.h(FPSNR_BX/FPSNR_BY), shared by kernel and host. - New host helper
core/src/feature/float_psnr_rows.h(vmaf_float_psnr_row_noise()), the only place the hosts sum partials. - If upstream changes
float_psnr.c's row loop ornoise_line(), the helper,test_float_psnr_rowsand the three twins change in the same PR. core/test/test_hip_float_psnr_parity.cusescore/test/float_psnr_twin_parity.hnow; the cases past 2^53 compare with==. The header no longer has a bound helper:core/test/test_metal_float_psnr_parity.ccallsfloat_psnr_twin_past_2_53_exact()(casetest_float_psnr_16bit_past_2_53_exact), as the CUDA, SYCL and HIP tests do.
The NEON and SVE2 float_moment kernels add in the scalar's order (ADR-1500, 2026-10-03)¶
fix/arm-moment-scalar-order. Fork-added files only (moment_neon.c, moment_sve2.c, the two tests); Netflix/vmaf has no aarch64 moment kernel and moment.c / float_moment.c are untouched. x86 object code is unchanged.
core/src/feature/arm64/moment_neon.candmoment_sve2.cstore each vector and add the lanes into onedoublein raster order (moment_add4(),moment_add_active()), asx86/moment_avx2.cdoes. A sync that changesmoment.c::compute_1st_moment()/compute_2nd_moment()changes all four SIMD kernels in the same PR; none may go back to lane accumulators or a vector reduction.core/test/test_moment_simd.cis rewritten around one kernel table and asserts==(the 1e-7MOMENT_REL_TOLis gone); its SVE2 case and the NEON cases ofcore/test/test_iqa_convolve.cprobevmaf_get_cpu_flags_arm()(arm/cpu.h). Do not switch them back tovmaf_get_cpu_flags()without avmaf_init_cpu()call: the cases then skip on every processor.
The Metal twins run the exact designs of the other GPU backends (ADR-1498, 2026-10-03)¶
fix/metal-twins-exact. Fork-only Metal files plus four shared headers; no upstream-mirror file changes and no CPU score moves.
- Every
.metalfile builds withmetal_shader_strict_fp_args(-fno-fast-math -ffp-contract=off), defined once between theBEGIN / END VMAF Metal shader strict FP policymarkers incore/src/metal/meson.build;test_metal_shader_build_contract.pyguards it and rejects identifiers named after MSL types (half,device, ...). - Each twin's arithmetic is a header under
core/src/feature/metal/onmetal_portable.h.metal_soft_double.h,metal_soft_signed.h,metal_float_vif_math.h,metal_float_adm_math.h,metal_integer_vif_gain.handmetal_integer_ssim_math.hfollow their SYCL headers statement for statement: a change to one copy changes the other in the same PR (RC5 rowT-METAL-SYCL-FP64-FREE-ARITHMETIC-COPIES-2026-10-03). - Shared headers:
ff_pair.h,ff_math.handciede_ff_math.hgain aVMAF_FF_MSL_SUBSETbranch andadm_gain_limit.ha Metal guard; the preprocessed SYCL, HIP and CUDA translation units are byte-identical, and a rebase that edits these headers keeps them so. - The Metal twins call the CPU's own routines on the host (
vmaf_motion_window_flush(),adm_float_reference.h,psnr_score.h,ciede_frame_sum(),cambi_internal.h,vif_get_filter()); an upstream change to one of them reaches Metal without a Metal edit, and no local copy may come back (test_float_adm_csf_upstream_contract.py,test_metal_twin_option_tables_contract.py). core/src/metal/dispatch_strategy.c'sg_metal_features[]equals the registry names andprovided_features[]of every Metal extractor (test_metal_twin_option_tables_contract.py); a rebase that changes a Metal twin's features changes the table in the same PR.
icx and icpx builds link glibc's libm, not Intel's libimf (ADR-1495, 2026-10-03)¶
fix/icx-system-libm. Fork-only build policy; no upstream file is touched and no score of a GCC or clang build moves.
core/src/meson.buildgains theBEGIN / END VMAF host math library link policyblock directly after the strict FP policy and twoadd_project_link_arguments()lines (C and C++) above the first build target; anintel-llvmcompiler gets-no-intel-lib=libimf. A rebase that reorders the top of the file keeps both above the first target (Meson refuses project link arguments after one) and never re-adds Intel's math library by name.core/test/test_icx_system_libm.py(new, suitefast) and two cases incore/test/test_strict_fp_compiler_args.pyguard it; the SYCL lanes oflibvmaf-build-matrix.ymlrun the new test as a step.
float_motion_sycl emits motion3 (2026-10-03)¶
fix/sycl-float-motion3-and-option-tests. Fork-only files; no upstream file touched. No score of an existing output moves.
core/src/feature/sycl/float_motion_sycl.cppprovidesVMAF_feature_motion3_scoreand declaresmotion_blend_factor/motion_blend_offset(host code only, no kernel). Itscollect()/flush()mirror CPUfloat_motion.c::extract()/flush(); an upstream sync that changes how the CPU emitsmotion3changes this file in the same PR, as forfloat_motion_cuda.candfloat_motion_hip.c.FEATURE_METRICS["float_motion"]inscripts/ci/cross_backend_parity_gate.pyandscripts/ci/cross_backend_vif_diff.pylistsmotion3; keep both lists equal. A twin that dropsmotion3fails the cell.core/test/test_sycl_twin_option_parity.cgains themotion3cases and the #1645 regression cases (flat identical frames, single pixel,apsnrwith--subsample 2,motion_v2weight / cap / one frame), plustest_twin_options_are_cpu_options.
The float_moment twins form the CPU's second-moment sum past 2^53 units (ADR-1497, 2026-10-03)¶
fix/float-moment-exact-past-2-53. Fork-only device and host code; no CPU extractor, no upstream-mirror line changes (moment.c, float_moment.c and picture_copy.cpp untouched).
- New shared headers
core/src/feature/float_moment_sum.h(integer arithmetic and every lane's step) andcore/src/feature/float_moment_sum_gpu.h(the bodies of the four CUDA / HIP kernels;cuda/integer_moment/moment_score.cuandhip/float_moment/moment_score.hipdefine theextern "C" __global__kernels around them). Both are in the CUDA and HIPdepend_fileslists ofcore/src/meson.build. cuda/integer_moment_cuda.c,hip/float_moment_hip.candsycl/integer_moment_sycl.cppallocate per-row buffers whenvmaf_moment_sum_may_round()holds and launch the four kernels after the frame kernel;collect()is unchanged.init_fex_cuda()no longer usesCHECK_CUDA_GOTO/ afail:label: the module and its six kernels load inmoment_cuda_load().- If upstream ever changes
compute_2nd_moment()(order, term), the header,test_float_moment_sumand the three twins change in the same PR. Keep the four kernels; never go back to rounding the exact sum once. core/test/float_moment_twin_parity.h: the cases past 2^53 compare with==(file-scopeFLOAT_MOMENT_TWIN_PAST_CASES, shared withtest_float_moment_sum); the dark samples reach 8191. The HIP parity test now uses the shared header like the CUDA and SYCL tests.
The macOS tester bundle measures every open Metal row (2026-10-03)¶
test/metal-report-full-measurement (ADR-1496). No score impact: tests, the parity gate scripts, the tester report and docs. Fork-only paths.
scripts/ci/cross_backend_parity_gate.pyandcross_backend_vif_diff.pycarry the samemetalentries inBACKEND_SUFFIX,BACKEND_DEVICE_FLAGandBACKEND_EXTRACTOR_ALIASES(integer_*_metal); a rebase that edits one table edits both (test_parity_gate_covers_registered_twins.py).--hold-exactis measurement only and never used in a CI lane.- The bundle copies the gate (
GATE_FILESintools/rc1-tester/image/prepare_build.py): it must stay standard-library only and read only those files,exact_twins.dand the ADRs the fragments cite. core/test/test_metal_*_parity.cusemetal_twin.handmetal_run_case(); their case names are referenced bytools/rc1-tester/image/metal-rows.json(test_metal_report_rows_contract.py).metal_parity_testsincore/test/meson.builddrives both the Metal and the self-test builds;test_metal_float_moment_parity_10bitis gone (the shared header covers 10 bits).ciede_twin_parity.h'sCIEDE_TWIN_TOLis#ifndef-guarded.
The parity allowlist page always has a Pending section (2026-10-03)¶
docs/close-parity-pending-row. No rebase impact: a docs generator and ledger rows. scripts/docs/generate-upstream-parity-allowlist.py emits the "Pending" heading with "None" when no fragment is pending; keep that branch, docs/state.md links to the anchor.
Tester dispatch fixes: bash 3.2 and a universal jsonschema lock (2026-10-03)¶
fix/tester-dispatch-bash32-locks. No rebase impact: fork-only paths (scripts/ci/build-macos-tester-bundle.sh, requirements/locks/jsonschema.txt and its manifest entry, tools/rc1-tester/tests/test_bash32_compat.py).
The dev container entrypoint no longer chmods a root-owned /tmp (2026-10-03)¶
fix/dev-entrypoint-tmp-chmod. No rebase impact on the library: dev/scripts/dev-mcp-entrypoint.sh only.
- The unconditional
mkdir -p /tmp && chmod 1777 /tmp(ADR-0498 follow-up) is now a guarded block that changes the mode only when it is not 1777 and the directory belongs to the entrypoint's user. Keep the guard if the line is touched again; seedev/AGENTS.d/entrypoint-unprivileged.md.
Explicit conversion of upstream's float products (CodeQL, 2026-10-03)¶
fix/codeql-float-product-explicit-conversion. No score moves; objdump -d of every touched object is identical.
ciede.c,third_party/xiph/psnr_hvs.c,x86/psnr_hvs_avx2.c,arm64/psnr_hvs_neon.c,integer_adm_kernels.h,adm_tools.h,iqa/convolve.cwrite the conversion of the float product's result:sqrt((double)(a * b)). Upstream has the implicit form; a sync that brings it back keeps the explicit cast (same bits, and CodeQL'scpp/integer-multiplication-cast-to-longreports only implicit widenings). Never cast an operand: that is PR #552 again.
The Cython extension declares init_dwt_band_d() as adm.c defines it (2026-10-02)¶
fix/ci-cython-adm-dwt-band-cursor. No score impact; adm.c is untouched.
compat/python-vmaf/core/adm_dwt2_cy.pyxtext-includescore/src/feature/adm.cand declares its staticinit_dwt_band_d(). Since #1859 the helper takes adouble *cursor and a length in samples; the.pyxfollows. An upstream sync or a rebase that changes the helper's signature changes the declaration and the call in the same PR:core/test/test_cython_adm_dwt_band_decl_contract.pyfails otherwise, without building the extension.
test_pic_preallocation has an explicit 180 s timeout (2026-10-02)¶
fix/ci-asan-pic-preallocation-timeout. No rebase impact: a timeout : argument of one test() line in core/test/meson.build, no code and no scores.
Port of Netflix/vmaf 9e48141b: NEON scale-zero ADM decouple (2026-10-02)¶
port/upstream-9e48141b-neon-adm-decouple. Upstream PR Netflix/vmaf#1656 (Dan Trapp): adm_neon.c gains adm_decouple_neon(), integer_adm.c binds it in init() on NEON. Ported, with these differences:
- The scalar fallback is the fork's kernel. Upstream calls the now-non-static
adm_decouple(); here it is static, soadm_decouple_neon()loopsadm_decouple_cols()frominteger_adm_kernels.hfor a fractional gain limit or a band narrower than four columns. That kernel stores the double productrst * gaintruncated toward zero (ADR-1413), which the vector path never forms: it takes integral limits only, where the product is the int32 one. A sync that brings a fractional-limit vector path must truncate (vcvtq_s64_f64), not round. - Border and angle test are the scalar's helpers.
adm_border_filt()andadm_cos_1deg_sq()replace upstream's inline copies; the angle test is the fp64 expression ofadm_angle_flag_fp64()(ADR-1194). - Split for HISS-04. Upstream's loop body is
adm_neon_decouple4(), its dot products and angle test are helpers; the operation sequence is unchanged. - The dispatch line is
s->adm_decouple = adm_decouple_neon;ininit_dispatch_simd()(integer_adm.c) underARCH_AARCH64. x86 object files are byte-identical before and after (170 of 170). - No checkasm. Upstream's
check_adm.chunk becomes rows ofcore/test/test_integer_adm_simd.c, which now builds on aarch64 and holds the scale-0 kernel toadm_decouple_cols()at gain limits 1, 1.2, 1.5, 2, 3, 7 and 100 on bands that make the angle test pass and the limit bind. Without gains 1, 2 and 3 a limited sample off by one passes (upstream's checkasm has that gap); with them a+1on the limited sample fails at gain 2.
Measured under qemu-aarch64: 1 202 040 cases of a standalone harness (dimensions 1 to 129, three stride kinds, guard pages, 106 gain limits, nine band patterns including -32768 and the angle boundary) equal the scalar kernel bit for bit; 831 of 831 frames of the whole extractor at --precision max equal --cpumask 1 on the three Netflix pairs and on synthetic frames from 17x17 to 641x359 at 8, 10, 12 and 16 bits and 1 to 8 threads.
ciede2000()'s two products are upstream's float products again (ADR-1476, 2026-10-02)¶
fix/ciede-upstream-expression. Scores move by up to 1.3e-9 (1.0e-8 on frames of 24x24 and smaller).
core/src/feature/ciede.c:sqrt(c_prime_1 * c_prime_2)and+ r_sub_t * chroma * huecarry no cast, as in Netflixlibvmaf/src/feature/ciede.c:224-225and:235-236. These two expressions are upstream's again: a sync takes upstream's side. The fork keeps a citedNOLINTand writes the conversion of each product's result explicitly (sqrt((double)(c_prime_1 * c_prime_2)),+ (double)(r_sub_t * chroma * hue)), which compiles to the same object code, andsquare(x)for upstream'spow(x, 2)in the second one (ADR-1467).- Do not re-add
(double)in front of either product to quiet CodeQL'scpp/integer-multiplication-cast-to-long: that was PR #552, and it moved 265 of 327 measured frames away from upstream. - The twins mirror it:
cuda/integer_ciede/ciede_device.h(chroma_product,rotation),ciede_ff_math.hfor SYCL and HIP (c_prime_1 * c_prime_2intofrom_float(),rotation * chroma * hueintoadd_f()). A change to either expression upstream changes all three files in the same PR. - Guards:
test_ciede_upstream_products(the CUDA header forms thefloatproducts),test_ciede_device_math(the header is the CPU extractor, bit for bit),test_sycl_ciede_exact_contract,test_cuda_ciede_exact_contract; on a devicetest_{cuda,sycl,hip}_ciede_parity.
The AVX-512 warm-up is x86/avx512_warm_up.h, with xmm0 in its clobber list (2026-10-02)¶
fix/cpu-avx512-warmup-clobber. No score impact.
core/src/x86/avx512_warm_up.h: fork-only header withstatic inline vmaf_x86_avx512_warm_up();core/src/cpu.cpp::vmaf_init_cpu()(fork file, upstream hascpu.cwithout a warm-up) calls it. Upstream has neither, so a sync does not touch them.- The clobber list is
"xmm0", "zmm0"."zmm0"alone is dropped by clang in a function not compiled for AVX-512, and a link-time-optimised build then zeroes a caller's value inxmm0(T-CPU-AVX512-WARMUP-CLOBBER-2026-10-02). Do not shorten it and do not move the statement back intocpu.cpp:test_cpuinlines it without link-time optimisation. - Guards:
test_cpu(test_avx512_warm_up_keeps_xmm0, AVX-512 host),test_inline_asm_clobber_contract(device-free).
Float ADM: dwt_quant_step() and the Barten CSF are upstream's float arithmetic again (ADR-1489, 2026-10-02)¶
fix/float-adm-barten-upstream-float, stacked on the entry below. Scores move: every float_adm score and every model score that reads one (@MOVE_SHORT@), and integer adm with adm_csf_mode=1, onto Netflix master's arithmetic. What remains between the fork's float_adm and Netflix's on x86 is the division of ADR-1442.
core/src/feature/adm_tools.h::dwt_quant_step(): the three statementsfloat r = ...,float temp = ...,float Q = ...are upstream's (libvmaf/src/feature/adm_tools.h). A sync takes upstream's side of them; the fork adds only the comment and thecodeql[cpp/integer-multiplication-cast-to-long]line aboveQ. Do not bring backdoublelocals or a(double)on an operand ofparams->k * temp * temp: #552 and #760 did.core/src/feature/barten_csf_tools.h: the arithmetic is upstream's, the text is not quite. Upstream writespow(p_0 * spatial_frequency, p_1),exp(- barten_mtf_params_b[i] * spatial_frequency)and so on, and lets the language promote thefloatargument. The fork writes that promotion out (pow((double)(p_0 * spatial_frequency), (double)p_1)), because the SYCL and Metal twins of integer ADM compile this header as C++, where the implicit form calls thefloatmath functions and returns other weights. When a sync brings a hunk of upstream's here, take upstream's arithmetic and keep the casts around eachfloatresult; never move a cast onto an operand ((double)p_0 * spatial_frequencyis the form #44 introduced).linear_interpolate()is upstream's text unchanged. The 18 locals that are never reassigned areconstin the fork (clang-tidy'smisc-const-correctnesson the C++ translation units; the sycl lane's baseline entry for the header is gone): keep the qualifiers.core/src/feature/metal/float_adm_metal.mm::fadm_dwt_quant_step()is a copy of the step and changes with it;core/test/test_float_adm_csf_upstream_contract.pyreads it.core/test/test_float_adm_csf_upstream.cholds the bits of a Netflixcea2b4d8build for 40 steps and 168 Barten weights (asserted on glibc), and compares the header compiled as C with the header compiled as C++ (core/test/barten_csf_cxx.cpp). If upstream changes the formula, the model constants or a Barten parameter, regenerate both tables from an upstream build in the same PR as the port.ADM_OPT_RECIP_DIVISIONstays undefined (ADR-1442); this entry does not change that one.
dwt_quant_step() of integer ADM is upstream's line again (ADR-1475, 2026-10-02)¶
fix/integer-adm-quant-step-upstream-float. Scores move: every integer ADM score and every model score that reads one (vmaf_v0.6.1 by up to 1.83e-5), onto Netflix master's values.
core/src/feature/integer_adm_kernels.h::dwt_quant_step(): the statementfloat Q = 2.0 * params->a * pow(10.0, params->k * temp * temp) / ...is upstream's (libvmaf/src/feature/integer_adm.c). A sync takes upstream's side of that statement; the fork adds only the comment and the explicit conversion of the product's result,pow(10.0, (double)(params->k * temp * temp))(same object code). Do not re-add a(double)on an operand of the product: #552 did, and it moved the CSF weights of scales 1 to 3 by 1 to 3 units in the last place.- The same statement lives in two twins that cannot include the C header:
core/src/feature/sycl/integer_adm_sycl.cpp::dwt_quant_step()andcore/src/feature/metal/integer_adm_metal.mm::iadm_dwt_quant_step(). Both form the exponent in a namedfloatand promote that. A change upstream makes to the function goes into all three;core/test/test_integer_adm_quant_step_contract.pyreads them. core/test/test_integer_adm_quant_step.cholds the step's bits from a build of Netflixcea2b4d8for five viewing geometries (asserted on glibc). If upstream changes the model constants or the formula, regenerate the table from an upstream build in the same PR as the port.testdata/scores_cpu_{576,640,720,1080,4k}.jsonwere regenerated (38 to 59 of 720 values each, at most 2e-5). A branch that regenerates them from an older base takes master's files and regenerates again.- Float ADM (
core/src/feature/adm_tools.h) still has its own, wider copy of the function; that is a separate change.
The psnr_hvs masking threshold is upstream's float product again (ADR-1488, 2026-10-02)¶
fix/psnr-hvs-upstream-expression. Scores move by up to 9.4e-7 dB on 27 of 319 measured frames.
core/src/feature/third_party/xiph/psnr_hvs.c:s_mask = sqrt(s_mask * s_gvar) / 32.f;and the same ford_mask, as in Netflixlibvmaf/src/feature/third_party/xiph/psnr_hvs.c:316-317. These two lines are upstream's again: a sync takes upstream's side. The fork keeps a comment and writessqrt((double)(s_mask * s_gvar))(same object code).- Do not re-add
(double)in front of the product to quiet CodeQL'scpp/integer-multiplication-cast-to-long: that was PR #552. - The same statement, same bits, in five more places; a change to it changes all of them in the same PR:
x86/psnr_hvs_avx2.candarm64/psnr_hvs_neon.c(compute_masks()),cuda/integer_psnr_hvs/psnr_hvs_score.cuandhip/integer_psnr_hvs/psnr_hvs_score.hip(hvs_threshold():floatproduct,doubleroot),sycl/integer_psnr_hvs_sycl.cpp(hvs_threshold():sqrt_rn()of thefloatproduct). core/src/feature/sycl/sycl_exact_fp.h:sqrt_prod_rn()andisqrt_floor50()(fork-local, ADR-1401) are removed; nothing used them after this change. The removal is declared inscripts/ci/silent-revert-allowlist.json(two ADR-1488 entries); a branch that still callssqrt_prod_rn()takessqrt_rn(a * b)of the float product instead.test_sycl_fp_arith_contract's fourth result is nowsqrt_rn()of thefloatproduct.- Guards:
test_psnr_hvs_dispatch_invariance(recorded blocks scored as Netflix master scores them; x86-64 and aarch64),test_psnr_hvs_simd,test_psnr_hvs_twin_exact_sum_contract; on a devicetest_{cuda,hip,sycl}_psnr_hvs_parity{,_large},test_{cuda,hip,sycl}_exact_twins,test_sycl_fp_arith_contract. - Float ADM (
core/src/feature/adm_tools.h) has its own copy of the function; the entry above (ADR-1489) covers it.
The motion twins compute motion_five_frame_window (ADR-1491, 2026-10-02)¶
port/motion-five-frame-window-gpu-twins, stacked on the CPU port (ADR-1478). Fork-only code; no upstream file changes.
core/src/feature/cuda/integer_motion_cuda.c,integer_motion_v2_cuda.c:raw[3]/pix[3]withring(2 or 3);prev_doneis the previous frame's event from frame 1 on.motion_cudawith the option emits the SAD score only fromcollect()and callsmotion_flush_window()inflush().motion_v2_cudaflushes throughvmaf_motion_window_flush()always.core/src/feature/hip/integer_motion_hip.c,integer_motion_v2_hip.c:prev_luma[2]withdepth(1 or 2); the same host split.core/src/feature/sycl/integer_motion_sycl.cpp: with the optiond_raw_y[0]/d_raw_y[1]hold framesn-2/n-1; the kernel is enqueued on every frame (the combined graph is recorded once), the planes advance inmotion_post_graph().integer_motion_v2_sycl.cpp:d_pix[3]withring, flush through the shared function.- A rebase that brings back a twin's own copy of the motion2 / motion3 arithmetic, a
VMAF_OPT_FLAG_DEFAULT_ONLYon the option of these six twins, or-ENOTSUPfor it, undoes this;test_<backend>_motion_five_frame_window,core/test/test_gpu_option_value_capability_contract.pyand the gate cellsmotion_mffw/motion_v2_mffwfail.
motion_five_frame_window is upstream's again (ADR-1478, 2026-10-02)¶
port/upstream-motion-five-frame-window. Ports Netflix a2b59b77 (the option, prev_prev_ref, pool sizing) on top of the fork's port of a4a1492d. Scores with the option equal Netflix 9e48141b bit for bit.
core/src/feature/integer_motion.c:extract()and the flush statements are upstream's again; the-ENOTSUPguard of ADR-0994 is gone. A sync takes upstream's side for the arithmetic. What stays the fork's: the helpersmotion_select_pipeline(),motion_flush_one(), andvmaf_motion_window_flush(), which is the body of upstream'sflush()below the feature-name dictionary, exported throughcore/src/feature/motion_window.h. An upstream change toflush()goes into those two functions.core/src/feature/integer_motion_v2.c: upstream deleted this file ina4a1492d; the fork keeps the extractor. Itsextract()is upstream's last text (a4a1492d^), its flush callsvmaf_motion_window_flush(). A change to the window is made once, ininteger_motion.c.core/src/feature/feature_extractor.h:prev_prev_refnext toprev_ref, as upstream.core/src/feature/feature_extractor.cpp: the fork's PREV_REF swap rotates the two fields (upstream has no swap).core/src/libvmaf.c:VmafContext::prev_prev_ref, rotated inread_pictures_update_prev_ref(), released invmaf_commit_remaining_owners(). Upstream copies the two pictures into the extractor as structs and zeroes them afterextract(); the fork hands out counted references (fex_take_prev_refs()/fex_release_prev_ref(), ADR-0778) at the three dispatch sites and in the worker job (batch_job_take_pictures()). Keep the fork's side of every such hunk and take only what upstream changes about which frames are kept.- Deliberate deviation (ADR-1478): upstream keeps frame n-2 in every run and sizes
check_picture_pool()asn_threads * 2 + 2. The fork keeps it only while a registered extractor'sreads_prev_prev_ref()answers true (VmafContext::keep_prev_prev_ref, set byadmit_prev_prev_ref()), adds the+ 2only then, and refuses a preallocated pool below four pictures next to such an extractor with-EINVAL. A sync keeps the fork's side ofread_pictures_update_prev_ref(),check_picture_pool(),vmaf_preallocate_pictures()and the PREV_REF swap. core/tools/vmaf.cpp: pool ofthread_cnt > 0 ? (thread_cnt + 1) * 2 + 1 : 4pictures (upstream's expression) plus the fork's read-ahead pictures; the serial four is needed because the tool sizes the pool before the models load.python/test/feature_extractor_test.py,python/test/vmaf_v1_quality_runner_test.py: the 13@unittest.skiplines the fork added for this option are gone; the files differ from upstream by formatting only in those tests. A sync must not bring a skip back.- GPU twins:
motion_five_frame_windowcarriesVMAF_OPT_FLAG_DEFAULT_ONLYonmotion_cuda,motion_syclandmotion_hipuntil the twin has the window (core/test/test_gpu_option_value_capability_contract.pylists them); the-ENOTSUPin theirinit()stays with the flag.
Tester image and macOS bundle are fork-only additions (ADR-1492, ADR-1493, 2026-10-03)¶
feat/tester-image-arm64. No upstream (Netflix/vmaf) file changes. Fork-only paths: docker/Dockerfile.tester, tools/rc1-tester/ (hw_*.py, image/, vmaf-tester-report), scripts/ci/{check-hardware-reports.py,check-macos-bundle-links.sh,build-macos-tester-bundle.sh}, scripts/docs/generate-hardware-reports.py, docs/hardware-reports/, two workflows and an issue form. tools/rc1-tester/src/vmaf_rc1_tester/safe_process.py gained an optional cwd argument. Makefile docs-fragments-check / -write each gained one hardware-reports line: keep them on a conflict. docs/state.md gained T-TESTER-APPLE-SILICON-EVIDENCE-2026-10-03.
Agent pages name the staged CUDA VIF kernels and the HIP handle header (2026-10-02)¶
docs/agents-notes-and-state-rows. No rebase impact on code: agent pages and two docs/state.md rows only.
core/src/feature/cuda/AGENTS.d/vif.md:integer_vif/filter1d.cuis assembled fromvif_*stages since PR #1860. An upstream hunk in a kernel body goes into the matching stage; the long bodies do not come back.core/src/hip/AGENTS.d/kernel-template.md:uintptr_thandles convert throughhip_handle.honly.docs/state.md:T-GPU-FLOAT-SSIM-FRAME-SUM-ORDER-2026-10-02is under "Recently closed". A branch that still edits it under "Open bugs" takes master's side for that row.
The ADM headers lost upstream's ADM_CM_THRESH_S_* macros and #pragma once (ADR-1142, 2026-10-02)¶
refactor/adm-tools-standards. No score impact: every object file of an x86 and an aarch64 build is byte-identical before and after.
core/src/feature/adm_tools.h: upstream's nineADM_CM_THRESH_S_{0_0, 0_W_M_1, 0_J, H_M_1_0, H_M_1_W_M_1, H_M_1_J, I_J, I_0, I_W_M_1}macros are gone. Nothing expanded them since ADR-1141 moved the masking threshold into the closed formadm_tools.c::adm_cm_thresh3x3_s()(integer twins:integer_adm_kernels.h::adm_cm_thresh()/i4_adm_cm_thresh()). An upstream hunk that touches the macros therefore conflicts, on purpose: port its arithmetic into those functions, keeping the nine terms in the macros' order, and do not re-add the macros (a re-added macro is dead text that silently ignores the upstream change).adm_tools.h,adm_csf_tools.h,adm_options.h:#pragma onceis gone; the#ifndefguards upstream also has are the only guard. Keep it so when a sync brings the pragma back (portability-avoid-pragma-once).adm_csf_tools.hkeeps the_USE_MATH_DEFINES/#ifndef M_PItwo-step the Windows lanes need (ADR-1234); the define carries a cited suppression.integer_adm.h: oneNOLINTBEGIN/NOLINTENDpair around the body for the C++-only checks a C header trips when a SYCL translation unit includes it (ADR-1138);recipindiv_lookup_populate()isconst. The positional initialisers ofdwt_7_9_YCbCr_thresholdstay positional.
compute_adm() is split into helpers; its debug-dump blocks are gone (ADR-1142, 2026-10-02)¶
refactor/adm-c-standards. No score impact: every recorded adm / float_adm output is identical before and after on x86 (scalar, AVX2, AVX-512) and aarch64 (scalar, NEON); both golden gates 271 passed, 12 skipped.
core/src/feature/adm.c no longer lines up with upstream's single 260-line compute_adm(). The function keeps its name and signature; port an upstream hunk into the helper that owns the statement:
Upstream statements in compute_adm() | Now in |
|---|---|
buf_stride, buf_sz_one, the SIZE_MAX check, data_buf, the six init_dwt_band*() calls | adm_alloc_bands() |
ind_size_y / ind_size_x, buf_y_orig / buf_x_orig, the four index rows each | adm_frame_alloc() + adm_alloc_indices() |
the fail: label and the three aligned_free() | adm_frame_free(), called once at the end of compute_adm(); every goto fail is a return of the helper |
adm_dwt2_lo / adm_dwt2 of both pictures | adm_scale_dwt2() |
adm_decouple, adm_csf_den_scale, adm_csf, adm_cm, adm_csf, adm_cm | adm_scale_sums(), same order, same arguments (the options travel in AdmScaleOpts) |
the for (scale ...) loop, num / den / aim_num / aim_den, scores[] | adm_accumulate_scales(); the four doubles are sums[0..3] |
numden_limit, vmaf_adm_floor_pair_named(), vmaf_adm_finalize_scores_named() | compute_adm() |
Kept as upstream has them: the types of every temporary (float per-scale sums added into double frame sums), the order aim_den before aim_num, the halving of w and h after the wavelet, the stdout messages byte for byte.
Dropped: the two #ifdef ADM_OPT_DEBUG_DUMP blocks. They called write_image() and PRINTF(), which nothing in the tree defines, so they could not compile; adm_options.h names the macro in a prose comment only (it was a commented-out #define until the CodeQL sweep of 2026-10-03). The (float *)(void *) casts in init_dwt_band*() are direct casts (bugprone-casting-through-void); init_dwt_band_d() keeps its signature because compat/python-vmaf/core/adm_dwt2_cy.pyx declares it.
RC7 CPU capability inserted, benchmarks are RC8, retrain is RC9 (ADR-1490, 2026-10-02)¶
docs/rc-map-rc7-cpu-capability. No rebase impact on code: labels only.
- A stale branch that opens a tuning row as RC7 relabels it RC8; a training row labelled RC8 becomes RC9. The disposition labels of
docs/state.mdare nowRC7 CPU capability source of truth,RC8 benchmarks, profiling and tuningandRC9 training and model validation; a rebase conflict in that table is resolved byscripts/dev/resolve-state-md-conflict.pyand the id goes under the label its row text earns. - The
tools/rc1-testercatalog phases areRC1,RC8andRC9. RootAGENTS.mdsection 11 and its compiled projections carry the new map; never resolve a conflict there by hand, take master's side and runpraetorctl compile-context.
docs/state.md uses the RC3 to RC8 labels (ADR-1421, 2026-10-02)¶
rc3-ledger-relabel, closes T-STATE-LEDGER-RC-RELABEL-2026-10-01. No rebase impact: ledger labels only.
- A rebase conflict in the disposition table of
docs/state.mdis resolved byscripts/dev/resolve-state-md-conflict.py, which keys each row by its bold label. The labels are nowRC3 twin exactness,RC5 deduplication,RC6 GPU capability source of truth,RC7 benchmarks, profiling and tuningandRC8 training and model validation. A branch written before this change that lists an id under "RC3 performance and backend acceleration" or "RC4 training and model validation" conflicts there: put the id under the label its row text earns (inexact scores RC3, duplication RC5, capability table RC6, correct scores with throughput or host residuals RC7, training RC8).
ADR-1459 — SpEED's covariance kernels are the fork's own and return the scalar kernel's bits (2026-10-02)¶
fix/speed-cov-kernel-exact, closes T-SPEED-COV-KERNEL-X86-NOT-BIT-EXACT-2026-10-02.
- Upstream's covariance kernels are not in the tree.
compute_cov_kernel_avx2()/compute_cov_kernel_avx512()(Netflix30f472b14) are deleted andcompute_cov_kernel_neon()(15297286) was never taken: each splits one sum over vector lanes with fused multiply-adds and does not returncompute_cov_kernel_scalar()'s bits. On upstream sync: do not re-import them or their dispatch inspeed_init(). A change upstream makes to those kernels has to be read for what it means for the row kernels below. core/src/feature/x86/speed_avx2.{c,h}andspeed_avx512.{c,h}keep their names and now holdspeed_cov_row_avx2()/speed_cov_row_avx512();core/src/feature/arm64/speed_neon.{c,h}(new, inarm64_fp_lib) holdsspeed_cov_row_neon(). Contract:core/src/feature/speed_cov.h(new). One lane is one covariance sum, a multiply and then an add; do not introduce an FMA or fold lanes into each other.core/src/feature/speed.cdiffers from upstream in four places:compute_cov_kernel_scalar()is notstatic, forms the product in its own statement and carries the function-scoped no-contraction guard (ADR-1057 pattern: GCCoptimizeattribute, clangfp contract(off)pragma), which has to survive a sync because every kernel is compared with this function;speed_cov_row_scalar()is new; upstream'scompute_covariance()is replaced bycompute_covariance_row(), andcompute_covariance_matrix()walks the lower triangle one row ofyblocks at a time (same pairs, same values, samedoubletofloatconversion);SpeedState::cov_rowandspeed_dispatch_cpu_kernel()dispatch the row kernel, on aarch64 too. An upstream change insidecompute_covariance()is ported intocompute_covariance_row()by hand.core/test/test_speed_simd.ccompares every sum with the production reference bit for bit and is built on aarch64 as well (speed_simd_test_archsincore/test/meson.build); the 1e-9 tolerance, the test's copy of the scalar kernel and thecheck_cov_matrix()case that the Netflix/vmaf#1653 port added for upstream's kernels are gone (its sizes are rows of the new matrix).- No Netflix golden-data, public API or FFmpeg patch impact. Scores are unchanged on x86 and on an aarch64 GCC build.
ADR-1461 — the strict FP policy is a project argument; the golden gate runs for aarch64 (2026-10-02)¶
fix/aarch64-clang-fp-contract, closes T-AARCH64-CLANG-FP-CONTRACT-FEATURE-LIB-2026-10-02.
core/src/meson.build: theVMAF strict FP compiler-argument policyblock moved from belowlibvmaf_cpu_static_libto the top of the file, andadd_project_arguments(vmaf_strict_fp_args, language : ['c', 'cpp'])follows it. On rebase: both stay above the first build target (Meson rejects the call after one), and upstream build changes that add a target need nothing: the argument reaches it.libvmaf_feature_static_libtakesvmaf_cflags_commonalone; do not restore+ vmaf_fp_model_argsthere or on any target (under icx it follows the project argument and turns contraction back on).core/src/metal/meson.build:metal_objcpp_argsends withvmaf_strict_fp_args(project arguments cover C and C++, not Obj-C++). Not built on this host.core/test/meson.build: four test executables dropvmaf_fp_model_args.core/test/test_strict_fp_compiler_args.pychecks the placement, scanscore/src,core/testandcore/toolsfor a target-level flag that undoes the policy, and undermeson testreads the build'scompile_commands.json.Makefile:test-netflix-golden-arm64andbuild-golden-arm64(GOLDEN_ARM64_CC,GOLDEN_ARM64_CROSS_FILE,GOLDEN_ARM64_BUILD_DIR,QEMU_LD_PREFIX); both golden targets shareGOLDEN_PYTEST_ARGS.scripts/ci/setup-golden-build.shtakesGOLDEN_CROSS_FILE;scripts/ci/golden-arm64-preflight.shandbuild-aux/aarch64-linux-gnu-clang.iniare new. Fork-only files.- Scores: x86-64 unchanged (same machine code with GCC). aarch64 clang builds move by up to 6.2e-5 (
speed_chroma), 3.5e-5 (float_vif), 2.3e-2 (speed_temporalon a checkerboard); aarch64 GCC builds by 1.2e-12 in the model score. No Netflix golden assertion, public API or FFmpeg patch changes.
The CLI read-ahead asserts its invariants (2026-10-02)¶
fix/cli-restore-frame-reader-asserts, closes T-CLI-FRAME-READER-ASSERTS-REPLACED-2026-10-02.
core/tools/vmaf.cpp:release_fetched_picture()andFrameReader(start,wait_for_free_slot,publish,next,request_stop) hold sevenassert()s. A sync or a lint pass must not turn them into early returns;core/test/test_cli_frame_reader_asserts_contract.pyfails if one goes.- A local clang-tidy run on a glibc 2.44 host reports
misc-static-asserton them (T-TIDY-GLIBC-244-STATIC-ASSERT-FALSE-POSITIVE-2026-10-02). Upstream Netflix/vmaf has noFrameReader; nothing to keep in step there.
Netflix/vmaf#1653 — SpEED fused anti-alias filter ported, NEON covariance kernel not ported (2026-10-02)¶
port/76ea5f03-speed-fused-filter. Upstream master moved from 8e7a1ac4e to cea2b4d83 (PR 1653, three commits). The fork is at parity with cea2b4d83 apart from the kernel named below.
76ea5f03"speed: fuse scalar antialias filtering and decimation" — ported.core/src/feature/vif_tools.c:vif_filter1d_dec16_s()runs the fork'svif_filter1d_vertical_s()at every 16th row and the newvif_filter1d_horizontal_dec16_s()at every 16th column (upstream has one function with the loops inline; the fork had already splitvif_filter1d_s()into those helpers).core/src/feature/speed.cfilter_and_downscale()and its mirrorspeed_internal_filter_and_downscale()incore/src/feature/speed_internal.ccall it under#elseof#if ARCH_X86and copy the decimated rows back; x86 keepsvif_filter1d_s()+vif_dec16_s(). On rebase: keep the two files in step, keep the x86 branch, and keep the decimated helper's taps, mirror and accumulation order equal tovif_filter1d_horizontal_s();core/test/test_speed_filter.c(upstream's test, split into helpers for the function-size limit, with more sizes and a whole-frame case) fails on any bit of difference and has to be run on a non-x86 build (build-aux/aarch64-linux-gnu.ini,qemu-aarch64).cea2b4d8"checkasm: cover fused SpEED filtering and decimation" — no checkasm tree in the fork. Its sizes (64x64, 255x63, 256x64), padded strides and unaligned source are rows oftest_speed_filter.c.15297286"arm64: add NEON SpEED covariance kernel" — upstream commit not ported: SIMD not bit-exact.compute_cov_kernel_neonadds the products into eight partial sums withvfmaq_f64;compute_cov_kernel_scalarkeeps one running sum. Underqemu-aarch6411.1.1 with GCC 16.1 the two differ in the last bits on 4061 of 18480 sums (35 widths x 11 heights x 3 layouts x 2 mean choices x 8 input patterns), by up to 3.5e-12 relative; upstream's own bound is 1e-10. A NEON kernel that keeps the running sum (lane products, scalar adds in order) is bit-identical under GCC, but it is what GCC 16 and clang 22 already emit for the scalar loop, so it gains nothing; under clang the scalar loop's remainder is contracted tofmaddand its vector body is not, so that kernel differs on 481 of 18480 sums there. On rebase: do not takelibvmaf/src/feature/arm64/speed_neon.{c,h}or theARCH_AARCH64dispatch inspeed_init()unless the arm64 bit-exactness rule gets an exception for this reduction (core/src/feature/arm64/AGENTS.md). The commit's widened checkasm case for the covariance kernels (15 sizes, 3 patterns, 3 layouts, bound1e-10 * (|ref| + 1)) is ported for the kernels the fork dispatches:check_cov_matrix()incore/test/test_speed_simd.c.- The CUDA, HIP and SYCL SpEED twins already evaluate the filter at the decimated samples only, in the scalar arithmetic. Only the comments that name the CPU call sequence changed (
cuda/speed/speed_score.cu,hip/speed/speed_hip_device.h,sycl/speed_sycl_pipeline.cpp), without moving a line. - Upstream branch
speed-fused-avx2moves x86 to the fused path as well. Port it only with x86 before/after identity at--precision maxon scalar, AVX2 and AVX-512 dispatch. - No Netflix golden-data, public API or FFmpeg patch impact.
The CPU clang-tidy lane brought back to baseline (2026-10-02)¶
fix/cpu-tidy-regressions, closes T-TIDY-CPU-LANE-ABOVE-BASELINE-2026-10-02.
core/src/picture_pool.cpp:default_picture_free()const-qualifies the error variable frommunmap()(misc-const-correctness).core/src/read_json_model.cpp: usesstd::cmp_greater_equal()to safely compare signed file size against unsigned buffer capacity (modernize-use-integer-sign-comparison).core/test/test_psnr_hvs_score.c: casts multiplication operands tosize_t(bugprone-implicit-widening-of-multiplication-result) and extracts test buffer allocation into helperalloc_test_buffers()to keep function length under 50 LOC (readability-function-size).core/test/test_read_pictures_failure_ownership.c: extracts picture pool ownership verification into helperverify_pictures_returned_to_pool()to keep function length under 60 LOC (readability-function-size).core/tools/vmaf.cpp: replaces runtimeassert()inFrameReader::read_frame()with explicit boundary check returning-EINVAL(cert-dcl03-c,misc-static-assert).- No Netflix golden-data, public API or FFmpeg patch impact.
The SYCL lint database reads both Ninja rule forms (2026-10-02)¶
fix/sycl-tidy-compdb-depfile-rule, closes T-SYCL-TIDY-COMPDB-DEPFILE-RULE-2026-10-02 (follows the entry below).
scripts/ci/gen-sycl-compile-commands.py:SYCL_COMMAND_PATTERNmatchesCUSTOM_COMMANDandCUSTOM_COMMAND_DEP(Ninja's name for a rule with a depfile). On rebase: a custom target that compiles a SYCL source in a third form needs the pattern extended; the generator exits 1 whenSYCL_BUILD_STATEMENTcounts more icpx.cppstatements than were parsed. Do not remove that count.clang_tidy_command()drops-MD,-MMDand-MF <file>.- Tests:
scripts/ci/tests/test_gen_sycl_compile_commands.py(ParseNinjaTests), wired through thetest-sycl-compile-command-generatorhook.
ADR-1443 — integer_ssim_sycl computes the CPU's fp64 term in integers (2026-10-02)¶
fix/sycl-ssim-cpu-arithmetic, ADR-1443 (builds on ADR-1432's sycl_soft_double.h).
core/src/feature/sycl/sycl_integer_ssim_math.h(new) mirrors the term ofinteger_ssim.c::ssim_reduce_row_range()(nine lines fromw_d = m.w;to the quotient). If upstream Netflix changes them, changeterm_bits(),product_sums_rounded()andproduct_sums_exact()in the same change;core/test/test_sycl_ssim_exact_contract.pyfails when the lines move, andreference_term()incore/test/test_sycl_integer_ssim_math.cholds a verbatim copy.core/src/feature/sycl/sycl_soft_signed.h(new): signed fp64 values in integers (SoftSigned), sum and difference with cancellation, conversion fromuint64_t, a radix-2^19 division. It includessycl_soft_double.hand uses itssoft_mul(),soft_round()and 128-bit helpers. On rebase: an edit to those functions insycl_soft_double.hhas to keep round-to-nearest-even per operation;test_sycl_integer_ssim_mathchecks every operation against the host's fp64.core/src/feature/sycl/integer_ssim_sycl.cpp(fixed-point twin only; thefloat_ssim_syclhalf of the file is untouched): the horizontal passes write five planes (no weight plane),IssimTermKernelreplaceslaunch_issim_vert_combine()and its per-group reduction, the state holdsd_terms/h_terms(oneuint64_tper pixel) instead of the partials,collectadds the plane withinteger_ssim_frame_sum(). The fp32 formula andsycl::reduce_over_groupmust not come back into this half of the file. Shape:ISSIM_TERM_SG16,ISSIM_TERM_GRF256; any other measured shape uses scratch memory (ADR-1395).VMAF_SYCL_ALWAYS_INLINEis defined once, incore/src/feature/sycl/sycl_compat.h;sycl_ff_math.h(ADR-1436) andsycl_soft_signed.htake it from there. On rebase: a header whose functions a kernel calls many times uses the macro; a plaininlinecan stay a call, and a call in a kernel is a scratch-memory frame (ADR-1395).ssimis declared exact forsyclbyscripts/ci/exact_twins.d/ssim.sycl; the row indocs/development/cross-backend-exact-twins.mdis generated (make docs-fragments-write).- Tests:
core/test/ssim_twin_parity.h(shared cases),core/test/test_sycl_ssim_parity.c(==),core/test/test_sycl_integer_ssim_math.cwith its probetest_sycl_integer_ssim_math_probe.cpp,core/test/test_sycl_ssim_exact_contract.py.core/test/test_sycl_twin_option_parity.cholds the twin to equality withenable_db/clip_dbtoo. - No Netflix golden-data, public API or FFmpeg patch impact.
SYCL translation units track their headers through compiler depfiles (2026-10-01)¶
fix/sycl-feature-header-deps, closes T-SYCL-TU-HEADER-DEPS-UNTRACKED-2026-10-01 (ADR-1320 applied to SYCL).
core/src/meson.build: thesycl_common_<name>andsycl_feature_<name>custom targets declaredepfileand passsycl_depfile_args(-MD -MF @DEPFILE@, empty on Windows). On rebase: a new custom target that compiles a SYCL source takes both; without them an edit to a header that source includes leaves the old kernels in the library.core/test/test_device_target_header_dependencies.pyguards it.- Fork-local build wiring; upstream Netflix/vmaf has no SYCL backend. No public API, ABI, FFmpeg patch or Netflix golden-data impact.
ADR-1436 — ciede_sycl runs the CPU's statements on fp32 pairs (2026-10-01)¶
fix/sycl-ciede-cpu-arithmetic, ADR-1436 (after ADR-1426 for the CUDA twin).
core/src/feature/sycl/sycl_ciede_math.h(new) mirrorsciede.c:rgb_to_xyz_map(),xyz_to_lab_map(),lab_color()=get_lab_color(),h_prime(),delta_h_prime(),upcase_h_bar_prime(),upcase_t(),r_sub_t(),delta_e()=ciede2000(). It is the same function set ascore/src/feature/cuda/integer_ciede/ciede_device.h. If upstream Netflix changes one of those routines, or the order ofextract()'s sum, change both headers in the same change;core/test/test_sycl_ciede_exact_contract.pyfails when the mirrored lines move.core/src/feature/sycl/sycl_ff_math.h(new): elementary functions on fp32 pairs. Its constants and tables sit betweenBEGIN GENERATEDandEND GENERATED; editscripts/dev/gen_sycl_ff_math.pyand run it with--write, never the block by hand. On rebase: no fp64 type in either header outsidemake_pair()/make_constants()(ADR-0220).core/src/feature/sycl/integer_ciede_sycl.cpp: the fp32 formula (srgb_to_linear()...ciede2000_dev()) and the per-work-group float partials are gone and must not come back. The kernel stores one float per pixel (d_terms), the host adds them (ciede_frame_sum()), the tables are copied tod_tablesat the first submit.ciede_pixel()keeps__attribute__((flatten, always_inline)): without it the kernel uses scratch memory and returns wrong values on Arc A-series under xe. The kernel is the functorCiedeKernelat SIMD-16 with the default register file; SIMD-32 spills to scratch memory.scripts/ci/cross_backend_calibration.py:LIBM_TWINS["ciede"]gains"sycl": 1e-9. On a conflict with another twin's entry keep both.core/test/test_strict_fp_compiler_args.pyno longer pins the number of probe device links incore/test/meson.build; it requires every one to carry the strict FP policy. On rebase: if the other side changes the pinned number, keep this form.core/test/meson.build: thesycl_ciede_format_variantsloop is gone; the variants are cases oftest_sycl_ciede_parity(core/test/ciede_twin_parity.h).- Tests:
core/test/test_sycl_ciede_math.cwith its probetest_sycl_ciede_math_probe.cppandsycl_ciede_math_probe.h(host and device),core/test/test_sycl_ciede_parity.c(1e-8),core/test/test_sycl_ciede_exact_contract.py. - No Netflix golden-data, public API or FFmpeg patch impact.
integer_vif_cuda resets its accumulators on the picture stream (2026-10-01)¶
fix/cuda-vif-accum-reset-order, closes T-UPSTREAM-1305-CUDA-VIF-ACCUM-STREAM-2026-10-01.
core/src/feature/cuda/integer_vif_cuda.c,vif_submit_plane(): thecuMemsetD8Asyncofs->buf.accum_datatakesvmaf_cuda_picture_get_stream(ref_pic)where upstream'sextract_fex_cuda()takess->str. On rebase: keep the picture stream. Upstream still has the private stream (Netflix/vmaf#1305; the patch there is the same one-argument change).
translate_picture_device() downloads every plane (2026-10-01)¶
fix/cuda-device-input-chroma-download, closes T-CUDA-DEVICE-INPUT-CHROMA-NOT-DOWNLOADED-2026-10-01.
core/src/libvmaf.c,translate_picture_device(): the plane mask ofvmaf_cuda_picture_download_async()is0x7(0x1for 4:0:0) where upstream passes0x1. On rebase: keep the fork's mask. Upstream still copies luma only (Netflix/vmaf#1613).- No public API, ABI, FFmpeg patch or Netflix golden-data impact.
ADR-1429 — vmaf_read_pictures() accepts an index gap; the contract is documented (2026-10-01)¶
docs/api-read-pictures-index-and-eagain, ADR-1429.
core/include/libvmaf/libvmaf.h: the Doxygen ofvmaf_read_pictures()(@param index),vmaf_score_at_index(),vmaf_feature_score_at_index(),vmaf_score_pooled()andvmaf_score_pooled_model_collection()gained the index-gap and-EAGAINtext. Upstream'slibvmaf.hhas none of it; on a sync keep the fork's text and take upstream's signatures.- No code, ABI or FFmpeg patch impact; no Netflix golden-data impact.
ADR-1431 — vmaf_read_pictures() owns its pictures on every return (2026-10-01)¶
fix/read-pictures-consume-on-error, ADR-1431.
core/src/libvmaf.c,vmaf_read_pictures(): builds theReadPicturesFramebefore the validation step and returns throughread_pictures_frame_cleanup()on a flushed context, a failedread_pictures_validate_and_prep()and a failed fallback; a failed CUDA translation returns through the newread_pictures_translate_abort().check_ring_buffer()returns the real error. On rebase: upstream returns the bare error from every one of those points and leaks the pictures; keep the fork's returns. A new earlyreturn err;between the argument checks and the extractor loop reintroduces the hang of thevmafCLI after an out-of-memory.docs/api/index.mdand thevmaf_read_pictures()Doxygen state the rule (ownership on every return);test_read_pictures_monotonicandtest_validate_pic_params_bpcdo not unref after a rejection.- No ABI or FFmpeg patch impact (callers already leave the pictures alone); no Netflix golden-data impact.
ADR-1434 — float_adm_sycl computes the CPU's arithmetic without fp64 (2026-10-01)¶
fix/sycl-float-adm-cpu-arithmetic-exact, ADR-1434 (after ADR-1420 for the CUDA twin), closes T-SYCL-FLOAT-ADM-NOT-CPU-ARITHMETIC-2026-10-01.
core/src/feature/sycl/sycl_float_adm_math.h(new) mirrorsadm_tools.c:divs()=DIVS(),angle_flag()=adm_angle_flag_s()(theADM_OPT_AVOID_ATANbranch),decouple_band()=adm_decouple_band_s(),csf_flt()= thefltstore ofadm_csf_s(),thresh_band()/threshold()=adm_cm_thresh3x3_s(),den_term()/cm_term()= the terms ofadm_csf_den_scale_s()/adm_cm_s(),row_sum()/fold_rows()= their two accumulators. It is the same function set ascore/src/feature/cuda/float_adm/float_adm_device.h. If upstream Netflix changes one of those routines, change both headers in the same change;core/test/test_sycl_float_adm_exact_contract.pyfails when the mirrored lines move.- The header has no fp64 type outside
make_gain_limit()(host code).kOneBy30/kOneBy15are the reference'sdoubleliteralsFLOAT_ONE_BY_30/FLOAT_ONE_BY_15as an fp32 pair and as a 53-bit significand. On rebase: do not replacetimes_constant(),add_scaled()orgain_limited()by fp32 arithmetic, and do not add an fp64 type; either breaks the twin (the first by up to 1e-7, the second by rejecting the TU on Arc A-series, ADR-0220). core/src/feature/sycl/sycl_soft_double.hgainedsoft_from_float_any()andsoft_to_float_any()(subnormal inputs and results).core/src/feature/sycl/float_adm_sycl.cpp: the kernels after the DWT arelaunch_decouple_csf(),launch_terms()andlaunch_row_sums(), each a call into the header. Gone and not to come back:fadm_dwt_quant_step()(the weights areadm_csf_rfactor_s()), the per-sub-group reductions,fadm_accumulate_totals()(adoublesum), the1e-2frame floor. The twin declares the CPU'sadm_f1s0..3,adm_f2s0..3,adm_skip_aim_scaleandadm_skip_scale0.- The earlier note below about
fadm_load_cm_pixel()describes code this change removed. Its rule stands in the new place: every helper of the header takes its band as a constant andBandsholds the three CSF weights as named fields (ADR-1395;core/test/test_sycl_kernel_source_contract.py). float_admis declared exact forsyclbyscripts/ci/exact_twins.d/float_adm.sycl.- Tests:
core/test/test_sycl_float_adm_math.cwith its probetest_sycl_float_adm_math_probe.cppandsycl_float_adm_math_probe.h(host and device),core/test/test_sycl_float_adm_parity.covercore/test/float_adm_twin_parity.h(==, every output),core/test/test_sycl_float_adm_exact_contract.py(15 planted regressions). - Depends on ADR-1420's exports (
adm_float_reference.h), on ADR-1442 (the reference divides:divs()isn / d, and must not become a product with a reciprocal) and on ADR-1432'ssycl_soft_double.h. - No Netflix golden-data, public API or FFmpeg patch impact.
float_adm_sycl uses no scratch memory; the scratch ratchet list is empty (2026-10-01)¶
fix/sycl-float-adm-cpu-arithmetic, closes T-SYCL-XE-SCRATCH-WRONG-RESULTS-2026-10-01.
core/src/feature/sycl/float_adm_sycl.cpp:fadm_load_cm_pixel()takes the band and returns oneFadmCmPixel(original,transformed,angle_flag) instead of aFadmDecouplePixelwhose arrays the callers indexed with the run-time band. On rebase: do not reintroducepixel.original[band]/pixel.transformed[band]infadm_aim_cm_term()orfadm_csf_cm_terms(); that array lives in private memory and the kernel returns NaN on Arc A-series GPUs under xe. The decouple kernel (fadm_decouple_item) keepsFadmDecouplePixel: its band loop has a constant trip count and stays in registers.core/src/sycl/scratch_ratchet.txthas no entry andkScratchExtractorsincore/src/sycl/scratch_check.cppis"". Keep both empty on a conflict;core/test/test_sycl_kernel_source_contract.pyrejects an entry or a name.- The self-test warning has a second wording for the empty list.
- No Netflix golden-data, public API or FFmpeg patch impact.
ADR-1415 — every x86 SIMD library is built without FP contraction (2026-10-01)¶
fix/icx-ssim-avx512-fp-contract, ADR-1415.
core/src/meson.build:x86_avx2_static_libandx86_avx512_static_libtakevmaf_strict_fp_argsinstead ofvmaf_fp_model_args. On rebase: if the other side still spellsvmaf_fp_model_argsfor either library, keep this side; a new x86 SIMD library takes the strict list too. The nine carve-out libraries are unchanged and carry the same flags.core/test/test_strict_fp_compiler_args.py:STRICT_TARGETSnames both general libraries.core/test/meson.build:test_integer_adm_simdtakes_simd_strict_fp_args(it compiles the scalar ADM kernels into its own translation unit; without the flag it fails on icx builds). Newtest_ssim_x86_simdblock (executable()+test()).core/test/test_ssim_x86_simd.c(new): AVX2 and AVX-512ssim_precompute,ssim_varianceandssim_accumulateagainst transcriptions of the scalar functions iniqa/ssim_tools.c. If upstream changes those scalar functions, update the transcriptions with the kernels (ADR-0139 twin group).- No source file of a library changes. GCC builds produce the same objects.
ADR-1414 — float_ms_ssim_sycl computes the CPU's arithmetic (2026-10-01)¶
fix/sycl-float-ms-ssim-cpu-arithmetic, ADR-1414.
core/src/feature/sycl/sycl_ssim_terms.h(new): the per-pixel SSIM arithmetic moved out ofinteger_ssim_sycl.cppunchanged (add_horizontal_tap,add_vertical_tap,round_moments,ssim_terms,ssim_term,term_fixed,FixedSum,float_ssim_constants, the three structs). Both SSIM twins include it. On rebase: if the other side edits one of these helpers insideinteger_ssim_sycl.cpp, apply the edit to the header; do not restore a second copy in either TU.core/src/feature/sycl/integer_ms_ssim_sycl.cpp:decimate_pixel()usessycl::fma()per tap;launch_horiz()becamelaunch_ms_ssim_horiz()with anMsHorizArgsstruct and pair sums;vertical_lcs_pixel()returnsLcsFixed(int64) throughssim_terms();d_partials/h_partialsarestd::int64_t;sum_scale_lcs()adds withFixedSumand rounds each mean to fp32;combine_ms_ssim()takesfabs()of l, c and s; the state lostc3. Each of these is what makes the twin match the CPU: keep this side if the other still has the fp32 forms.- If upstream Netflix changes
ms_ssim_decimate.c,iqa/convolve.c,iqa/ssim_tools.c(ssim_variance_scalar,ssim_accumulate_default_scalar) orms_ssim.c::ms_ssim_score_scales(), mirror it in the header or the twin in the same change. The HIP and Metal twins still use the old arithmetic (T-GPU-FLOAT-MS-SSIM-CPU-ARITHMETIC-2026-10-01). scripts/ci/cross_backend_calibration.py:EXACT_TWINSgainsfloat_ms_ssimandfloat_ms_ssim_lcs:sycl.- Tests:
core/test/test_sycl_ms_ssim_parity.c(new bit-for-bit test, run first in the binary), seven planted regressions incore/test/test_sycl_kernel_source_contract.py,scripts/ci/test_cross_backend_parity_gate.py. - No Netflix golden-data, public API or FFmpeg patch impact.
fix/adm-decouple-fractional-gain-truncation — the integer ADM gain limit truncates on every path (ADR-1413, 2026-10-01)¶
core/src/feature/x86/adm_avx2.cdecouple_gain_avx2()andcore/src/feature/x86/adm_avx512.cdecouple_gain_avx512()/decouple_s123_limit_half_avx512()convertrst * gainwith_mm256_cvttpd_epi32,_mm512_cvttpd_epi32and_mm512_cvttpd_epi64. Upstream master uses the roundingcvtpdforms in all three places (adm_avx2.c:854,adm_avx512.c:964,adm_avx512.c:1407at6ec23e8f2), which differ from its own scalar code for a non-integeradm_enhn_gain_limit. On a sync keep the fork's truncating conversions;test_adm_decouple_matches_scalar_for_gains(test_integer_adm_simd) fails on upstream's.- The scale-0 angle masks (
decouple_angle_mask_avx2(),decouple_angle_mask_avx512()) read themadd_epi16squared magnitudes as unsigned and the dot product as 2^31 where it isINT32_MIN. Upstream reads all three as int32. Keep the fork's reading; the same test covers it. core/src/feature/adm_gain_limit.his fork-only (no upstream counterpart): the integer form of(int64_t)((double)rst * gain). The SYCL twin (core/src/feature/sycl/integer_adm_sycl.cpp, also fork-only) calls it fromadm_dev_gain_limit();gain_limit_to_q31andGainLimitQ31are gone and must not come back. If upstream changes how the scalar forms or converts the product inadm_decouple_band()/adm_decouple_band_s123()(integer_adm_kernels.hhere,integer_adm.cupstream), change the header, the two x86 files andtest_adm_gain_limitin the same PR, and re-run the Netflix golden gate: three of its assertions use the integer extractor at a limit of 1.2.- The CUDA and HIP twins (
adm_decouple_inline.cuh/.hip) are untouched; they assign the double product to anint32_t.
fix/sycl-speed-lanczos4-host-weights — the SYCL SpEED twins read the CPU's lanczos4 weights (2026-10-01)¶
core/src/feature/sycl/speed_sycl_pipeline.cpp:Pipeline::lanczos(device USM, allocated only for a lanczos4 resample),upload_lanczos()at init,ScaleArgs::lanczos, andscale_lanczos()reading the nine column and nine row taps from it.lanczos_weight()andsycl::sinpiare gone. The table comes fromspeed_internal_gpu_lanczos_weights()(speed_internal.c), the routine the CUDA twins use; do not give the SYCL pipeline a routine of its own. Fork-only file, no upstream counterpart.core/src/sycl/scratch_ratchet.txtandkScratchExtractorsinscratch_check.cpp: the eight SpEED kernels (launch_scale,launch_decimate, 8-bit and 16-bit, with and withoutRoundedRangeKernel) are off the ADR-1395 ratchet, so a private array or a spill in them failstest_sycl_kernel_scratchagain. On rebase: a conflict in either file with another scratch-free PR is a set difference; keep every removal.core/test/test_cuda_speed_lanczos4_parity.cis renamedtest_gpu_speed_lanczos4_parity.c: one source, built astest_cuda_speed_lanczos4_parityand, with-DLZ_BACKEND_SYCL=1, astest_sycl_speed_lanczos4_parity. A HIP executable needs a third backend block, not a copy of the file.- No CPU code changes. No Netflix golden-data, public API or FFmpeg patch impact.
docs/adm-integer-aim-unclipped — integer AIM is unclipped and float AIM is clipped, as upstream (ADR-1417, 2026-10-01)¶
- No source change.
core/src/feature/integer_adm.cadm_result_finalise()reportsaim_num / denandcore/src/feature/adm.ccompute_adm()reportsMIN(aim_num / aim_den, 1)(throughvmaf_adm_scale_ratios()andvmaf_adm_finalize_scores()inadm_score.h), matching upstreaminteger_adm.c:3007andadm.c:323at6ec23e8f2. The integer value goes above 1 on a reference without detail (3.1756 on the 64x64 patch picture). - On a sync: if upstream adds the clip to
integer_adm.c, or removes it fromadm.c, port the change, update the integer or float expectations ofcore/test/test_integer_adm_aim_unclipped.cin the same PR, re-run the Netflix golden gate and move the "AIM above 1" section ofdocs/metrics/features.md. Do not unify the two extractors on the fork's own initiative: the shippedvmaf_v1.0.16models read the integeradm3.
perf/sycl-cli-pinned-host-picture-pool — SYCL CLI pinned host USM picture pool (2026-10-01)¶
core/src/picture_pool.h,core/src/picture_pool.c,core/src/picture.h:VmafPicturePoolConfiggains custom allocation callbacks (alloc_picture_callback,free_picture_callback,sync_picture_callback,attach_picture_callback,cookie) and buffer typeVMAF_PICTURE_BUFFER_TYPE_SYCL_HOST_PINNED.core/src/libvmaf.c:prepare_picture_pool()configures the picture pool to allocate SYCL host USM memory when a SYCL state is attached to the context.- Upstream sync note: Preserve custom allocation callbacks in
VmafPicturePoolConfigand picture pool fetch/close functions.
fix/dev-image-icx-native-fma-drift — omit -march=native from dev container reference build (2026-10-01)¶
fix/dev-image-icx-native-fma-drift, ADR-1317, T-DEV-IMAGE-ICX-NATIVE-FMA-DRIFT-2026-09-30.
dev/Containerfile: omitted-Dc_args="-march=native"from thelibvmaf-buildstage reference binary compilation (CC=icx CXX=icpx meson setup core/build core). Under Intel oneAPIicx,-march=nativeenables FMA contraction in unvectorized CPU feature extractors (generating 102vfmaddinstructions inspeed.c), leading to SpEED score drift against standard reference builds.- Tests:
scripts/ci/tests/test_dev_container_reference_build_flags.pyguards against re-introducing-march=nativein the reference binary build stage. - No Netflix golden-data, public API or FFmpeg patch impact.
fix/speed-lanczos4-host-weights — the CUDA SpEED twins read the CPU's lanczos4 weights (2026-10-01)¶
core/src/feature/vif_tools.c(upstream-mirror): the weight loop oflanczos4_interpolation()moved into a staticlanczos4_weights(), and a new exportedvif_scale_lanczos4_axis_weights()fills the per-axis table of those weights for every output sample (VIF_LANCZOS4_TAPSinvif_tools.h). The CPU scaler's results are unchanged: same expressions, same order. Upstream-sync note: a conflict inlanczos4_interpolation()will offer the inlined loop as "theirs". Keep the call tolanczos4_weights(), and apply any upstream change tolanczos4_kernel()or to the(i + 0.5) * ratio - 0.5position ofvif_scale_frame_lanczos4_s()tovif_scale_lanczos4_axis_weights()as well.core/test/test_speed_lanczos4_weights.c(new, no device) replays the table againstvif_scale_frame_s()bit for bit and fails when they drift.core/src/feature/speed_internal.{h,c},speed_gpu_common.h:speed_internal_gpu_lanczos_count()/speed_internal_gpu_lanczos_weights()andSPEED_GPU_LANCZOS_TAPS, the table layout for a device pipeline (taps of every scaled column, then of every scaled row). Backend-neutral: the SYCL and HIP twins can read the same table.core/src/feature/cuda/speed_cuda_pipeline.c,speed/speed_cuda_params.h,speed/speed_score.cu: aSPEED_BUF_LANCZOSdevice buffer uploaded at init and alanczospointer inSpeedCudaFrameArgs;scale_lanczos()reads the weights andlanczos_weight()/sinpif()are gone.test_cuda_device_resident_contract.pyrejects a sine inspeed_score.cu.- No Netflix golden-data, public API or FFmpeg patch impact.
fix/sycl-shared-frame-sticky-geometry — re-allocate shared frame buffers on geometry change (2026-10-01)¶
core/src/sycl/common.cpp:vmaf_sycl_shared_frame_init()previously returned 0 immediately whenshared_ref_buf[0]was already allocated, even if the new context requested a differentw,h, orbpc. This retained the old buffer allocations and pitch when reusing aVmafSyclStateacross contexts of varying dimensions.vmaf_sycl_shared_frame_init()now checks if existing buffers match the requested geometry; if not, it invokessycl_shared_frame_reinit_unwind()to wait for queue completion, free existing command graphs, release shared frame and chroma buffers, and allocate new buffers with the updated geometry.core/src/libvmaf.c: removed the!vmaf_sycl_get_shared_ref(...)check inread_pictures_sycl_prep(), callingvmaf_sycl_shared_frame_init()unconditionally so any geometry change is propagated cleanly.core/test/test_sycl_shared_frame_sticky_geometry.c: added unit test running multiple consecutiveVmafContextinstances of different sizes (64x48, 128x96, 32x24) sharing a singleVmafSyclState, asserting identical scores to CPU.
perf/adm-p3-fast-path — restore adm_p_norm==3.0 fast path (ADR-0463) (2026-10-01)¶
core/src/feature/adm_tools.c,adm_tools.h: restoresadm_sum_cube_s_p3,adm_csf_den_scale_s_p3, andadm_cm_s_p3(ADR-0463 / BUG-048 B3) lost to stale merges. Dispatches inadm_sum_cube_s(),adm_csf_den_scale_s(), andadm_cm_s()whenadm_p_norm == 3.0(default for standard VMAF evaluation).- Eliminates per-pixel
powf()and inner-loop branching on the hot path. Float accumulation matches generic path and outputs are 100% bit-identical across all 48 frames at--precision maxon the Netflix 576x324 reference pair and 1080p checkerboard. - Functions kept within HISS-04 limits (<= 60 LOC) via helper refactoring.
- Passes all 255 fast suite tests and full Netflix CPU golden gate.
fix/sycl-motion-chroma-geometry — motion_sycl motion_add_uv sizes chroma from pixel format (2026-10-01)¶
core/src/feature/sycl/integer_motion_sycl.cpp,motion_configure_chroma(): deriveschroma_wandchroma_hfrom the inputVmafPixelFormatusingvmaf_chroma_extent()frompicture_geometry.hforYUV420P,YUV422P, andYUV444P(rejectingYUV400Pand unknown formats). Previously, it hardcoded(w + 1) >> 1and(h + 1) >> 1regardless of format.core/test/test_sycl_motion_add_uv_parity.c: parametrized to verify both non-zero UV contribution and bit-exact match against the fixed-point scalar oracle acrossYUV420P,YUV422P, andYUV444P.- Upstream-sync note:
motion_add_uvis a fork-added option onmotion_sycl. Upstreaminteger_motion.cdoes not havemotion_add_uv.
docs/sycl-float-ssim-residual — float_ssim_sycl combined formula residual closed (2026-10-01)¶
- PR #1645 (
9e9ea0571) already alignedfloat_ssim_syclwith CPU reference arithmetic incore/src/feature/sycl/integer_ssim_sycl.cpp(ssim_termsandssim_termevaluate exact per-pixel $l \cdot c \cdot s$ in fp32 pairs, fixed-point work-group sumsterm_fixed, and double host reduction). - Verified on Intel Arc A380 under Linux
xekernel driver: max absolute difference against--backend cpuis 0.000e+00 on Netflix 576x324 (48 frames) and BBB 3840x2160 (auto scale andscale=1), down from 7.8e-5.test_sycl_twin_option_paritypasses 13/13 with exact match on flat identical frames (72.247199 dB). - Closes
T-SYCL-FLOAT-SSIM-COMBINED-FORMULA-RESIDUAL-2026-09-29.
perf/cuda-adm-cm-register-pressure — adm_cm.fatbin zero-spill and bounded registers (2026-10-01)¶
core/test/test_cuda_adm_cm_register_pressure.py: new Python regression test verifying that all kernel functions inadm_cm.fatbinacross all compiled CUDA architectures (sm_80,sm_86,sm_89,sm_90,sm_100,sm_120) have zero stack spill (STACK:0), zero local memory spill (LOCAL:0), and bounded register usage (REG <= 208, onsm_89REG <= 176).- Closes
T-CUDA-ADM-CM-REGISTER-PRESSURE-2026-09-07indocs/state.md. - No Netflix golden-data, public C API or FFmpeg patch impact.
ADR-1411 — float_motion_sycl adds its SAD in the CPU's order (2026-10-01)¶
fix/sycl-float-motion-cpu-float-sum, ADR-1411 (follows ADR-1409).
core/src/feature/sycl/float_motion_sycl.cpp: the blur kernel (launch_float_motion) writes the blurred plane only;fm_store_sad(), its local accessor and theprev_blur/sad_partials/compute_sad/wg_count_xmembers ofFmKernelArgsare gone. The newlaunch_float_motion_row_sad()runsfm_row_sad()with one work-item per row (sycl::range<1>(height), sub-group size 8), which adds|cur[j] - prev[j]|left to right into one fp32 accumulator. That loop shape is load-bearing: a group, sub-group, strided or atomic reduction gives a different rounding and the twin stops matching the CPU. On rebase: if the other side still hasfm_store_sad()orsycl::reduce_over_groupin this TU, keep this side.collect()readsheightfloats (h_row_sad) and only callsvmaf_float_motion_score_from_row_sads()(core/src/feature/float_motion_sad.h, added by ADR-1409).reduce_sad()and thedoublesum over work-groups are gone.- If upstream Netflix changes
compute_motion_simd(),float_sad_line_c()or the order ofconvolution_f32_c_s(), mirror it infm_row_sad(), the blur helpers and the shared header in the same change. The HIP and Metal twins still sum per block (T-GPU-FLOAT-MOTION-CPU-FLOAT-SUM-2026-10-01). scripts/ci/cross_backend_calibration.py:EXACT_TWINS["float_motion"]is{"cuda", "sycl"}, so the gate compares the CPU, CUDA and SYCL cells with tolerance 0.- Depends on ADR-1367 (the SYCL strict FP line): with contraction on, the blur is no longer the CPU's and the equality tests fail. The row kernel must stay free of scratch memory (ADR-1395).
- Tests:
core/test/test_sycl_float_motion_parity.c(==, every frame, 8 / 10 / 12 bits, device), five planted regressions incore/test/test_sycl_kernel_source_contract.py,scripts/ci/test_cross_backend_parity_gate.py. - No Netflix golden-data, public API or FFmpeg patch impact.
ADR-1409 — float_motion_cuda adds its SAD in the CPU's order (2026-10-01)¶
fix/cuda-float-motion-cpu-float-sum, ADR-1409.
core/src/feature/cuda/float_motion/float_motion_score.cu: the blur kernels (float_motion_kernel_8bpc/_16bpc) write the blurred plane only; their SAD arguments (prev_blur, the partials buffer,compute_sad) are gone. The newfloat_motion_row_sadkernel runs one thread per row and adds|cur[j] - prev[j]|left to right into one fp32 accumulator. That loop shape is load-bearing: a block, warp, strided or atomic reduction gives a different rounding and the twin stops matching the CPU. On rebase: if the other side still passes the eight (8-bit) or nine (16-bit) kernel arguments, keep this side's five and six.core/src/feature/cuda/float_motion_cuda.c: readback isframe_hfloats;reduce_sad()only callsvmaf_float_motion_score_from_row_sads().core/src/feature/float_motion_sad.h(new, backend neutral) is the tail offloat_motion.c::compute_motion_simd(): onefloatover the rows, then afloatdivision by theintpixel count.- If upstream Netflix changes
compute_motion_simd(),float_sad_line_c()or the order ofconvolution_f32_c_s(), mirror it in the kernel and the helper in the same change. The SYCL, HIP and Metal twins still sum per block (T-GPU-FLOAT-MOTION-CPU-FLOAT-SUM-2026-10-01); do not copy that shape back. scripts/ci/cross_backend_calibration.py:EXACT_TWINSgainsfloat_motion:cuda, so the gate compares that cell with tolerance 0.- Depends on ADR-1403 (
--fmad=falseon the fatbin): with contraction on, the blur is no longer the CPU's and the equality tests fail. - Tests:
core/test/test_cuda_float_motion_parity.c(==, 8 / 10 / 12 bits, device),core/test/test_float_motion_sad.c(new, device-free), five planted regressions incore/test/test_cuda_kernel_source_contract.py,scripts/ci/test_cross_backend_parity_gate.py. - No Netflix golden-data, public API or FFmpeg patch impact.
ADR-1403 — one FP flag list for every CUDA kernel; float_ms_ssim_cuda follows the CPU's arithmetic (2026-10-01)¶
core/src/meson.build:cuda_cu_extra_flagsis empty and may not carry a floating-point flag again. Every fatbin's command takescuda_device_strict_fp_args, defined once between the# BEGIN / # END VMAF CUDA device strict FP policymarkers (--fmad=falsewith the host strict args under nvcc,-ffp-contract=offunder clang CUDA). On rebase: a conflict in the fatbincustom_targetor incuda_cu_extra_flagswill offer the per-kernel--fmad=falseentries as "theirs". Keep the shared list and the empty map; a kernel added by the other side needs no entry (ADR-1397'spsnr_hvs_scoreentry was folded in this way, andtest_psnr_hvs_twin_exact_sum_contract.pynow checks the shared list). The clang branch assignsnvcc_ccbin_flagsandnvcc_host_includesempty; without them-Denable_nvcc=falsedoes not configure.core/src/feature/cuda/integer_ms_ssim/ms_ssim_score.cuandinteger_ms_ssim_cuda.c: the kernels reproducems_ssim_decimate.c(__fmaf_rn()per tap),iqa_convolve()(fp32 products summed in theMsPairfp32 pair, which stands for the reference's fp64 sum) andssim_accumulate_default_scalar()(fp32 denominators, fp32 quotient fors,__fsqrt_rn()); the host passes fp32 constants and rounds each per-scale mean to fp32 before the Wang combine. An upstream change to any of those three CPU routines, or toiqa_ssim()/ms_ssim.c's combine, must be mirrored here. The HIP, SYCL and Metal twins still have the older arithmetic (T-GPU-FLOAT-MS-SSIM-CPU-ARITHMETIC-2026-10-01); do not copy it back.- Tests that fail when either comes undone:
core/test/test_strict_fp_compiler_args.py(policy executed for nvcc and clang, planted per-kernel flags),test_cuda_kernel_source_contract.py(six plantedfloat_ms_ssimregressions),test_cuda_device_resident_contract.py, andtest_cuda_float_ms_ssim_parity(bit-exact on a CUDA device). - Output changes for
float_ms_ssim_cuda,ciede_cuda,float_vif_cudaandfloat_motion_cuda; no snapshot undertestdata/is a CUDA output. No Netflix golden-data, public API or FFmpeg patch impact.
perf/vif-scalar-malloc-hoist — reuse tmpbuf in scalar VIF fallbacks (ADR-0463) (2026-10-01)¶
core/src/feature/vif_tools.c: invif_filter1d_s(),vif_filter1d_sq_s(), andvif_filter1d_xy_s(), reuses the caller-suppliedtmpbufscratch buffer (allocated once per frame bycompute_vif()) instead of allocating and freeingtmpon every call.- Upstream Netflix/vmaf allocated
tmpper filter call viaaligned_malloc(). On ARM64 and CPU architectures without AVX2 convolution, this fired 12 heap allocations and deallocations per frame. - Scores remain 100% bit-identical. No change to algorithm or math.
fix/adm-cm-centre-tap-wrap — integer ADM departs from upstream master's masking centre tap (ADR-1402, 2026-10-01)¶
The fork's integer ADM is not upstream master's on one kind of content. Upstream narrows the centre tap of the scale-0 masking threshold to int16_t and subtracts thr << shift in 32 bits. The fork keeps the tap in int32 and clamps |x| - thr * 2^shift to [0, INT32_MAX] in int64, in every implementation. This is the second revision of the fork's own Netflix/vmaf PR #1602, which upstream has not merged. Scores differ from upstream master only where a scale-0 coefficient reaches 15360: isolated impairments on flat content, full-range noise. The Netflix golden pairs are unchanged.
- When upstream merges #1602 as it stands: the scalar code is already the fork's (
adm_cm_thresh(),adm_cm_excess_s0()); keep the fork's side of every conflict. Do not take upstream's vector hunks (threshold_overflow/threshold_fits): they differ from the scalar for a negative threshold, andtest_integer_adm_simdfails on them. Do not take its CUDA hunk either:adm_cm.cualready callsadm_cm_excess_s0(). Then drop this entry. - When upstream changes #1602 again, or fixes the wrap another way: all of these must change together, and the golden gate must be re-run:
core/src/feature/adm_cm_accumulator.h(adm_cm_excess_s0()),core/src/feature/integer_adm_kernels.h(adm_cm_thresh()),x86/adm_avx2.c(cm_thresh_band_avx2(),cm_excess_avx2()),x86/adm_avx512.c(cm_thresh_band_avx512(),cm_excess_avx512()),cuda/integer_adm/adm_cm.cu(DLM and AIM),hip/integer_adm/adm_cm.hip,sycl/integer_adm_sycl.cpp(adm_dev_csf_centre(),adm_dev_cm_excess_s0()),metal/integer_adm.metal(adm_cm_excess_s0()). NEON has no contrast-masking kernel. - A sync must not restore the
(int16_t)cast on the centre tap, the_mm*_srai_epi32(_mm*_slli_epi32(centre, 16), 16)pair in the vector thresholds,adm_i16()around the SYCL centre term, orabs(x) - (thr << shift)in any of the files above.test_integer_adm_cm_threshold(CPU) andtest_gpu_adm_tiny_frames(device twins) fail on the first three, because a patch picture scores above 1 or a twin leaves the scalar; the sanitizer lane stops on the last.
x86/adm_avx2.c, x86/adm_avx512.c and integer_adm.c no longer line up with upstream. To edit the two x86 files at all, their functions had to fit the 60-line limit (ADR-1298). The scalar kernels moved out of integer_adm.c into core/src/feature/integer_adm_kernels.h (same names as in the "refactor/c-rework-adm" entry below, plus column-range variants such as adm_decouple_cols(), adm_csf_cols(), adm_dwt2_hpass()), and the x86 files call them for edge rows and leftover columns instead of expanding their own copies of upstream's macros. Every upstream *_avx256 / *_avx512 macro is now a static function:
ADM_CM_THRESH_S_I_J_*iscm_thresh_band_*()+cm_thresh_*();ADM_CM_ACCUM_ROUND_*iscm_excess_*()+cm_accum_*(); the row loop iscm_block_*(),cm_tail_block_*(),cm_row_pass_*()andcm_row_*(), driven by the scalaradm_cm_rows(), which owns the per-row fold (ADR-1167). TheI4_*macros arei4_cm_thresh_band_*(),i4_cm_cube_*()andi4_cm_row_*(), driven byi4_adm_cm_rows().- The decouple, CSF, denominator and DWT bodies are
decouple_*,csf_*,csf_den_*,dwt2_*andi4_dwt2_*helpers with the upstream arithmetic unchanged. ADR-0502's prefetch isdecouple_prefetch_avx512(). - Re-port an upstream hunk to these files by hand into the helper that owns the expression. A hunk to a scalar tail or edge macro in an x86 file has no counterpart: the code is the shared kernel.
- Two things are deliberately not upstream's: the leftover columns of a contrast-masking row are the top lanes of one more vector block that ends at the last column (upstream runs them in scalar code), and the AVX2 cube shift is arithmetic, by a bias folded into the rounding term (upstream's second revision calls a variable-count
sra_epi64helper). sycl/integer_adm_sycl.cpp: the two DWT launches were split the same way (AdmDwtVertArgs,AdmDwtHoriArgs,adm_dev_dwt_*()helpers); the kernels compute the same values and still use no scratch memory (ADR-1395).
No public API, CLI or FFmpeg patch impact.
port/upstream-1590-model-collection-growth-test — a failed model-collection growth keeps the collection (2026-10-01)¶
core/src/model.c,vmaf_model_collection_append(): when therealloc()that doubles the model array fails, the function returns-ENOMEMdirectly. Upstream Netflix/vmaf hasif (!m) goto fail;there, and itsfaillabel clears*model_collection, which loses the existing collection. The fork's own upstream PR #1590 proposes the direct return; it is open. Upstream-sync note: a conflict in this function will offergoto failas "theirs". Keep the direct return.core/test/test_model_collection_growth.c(new, fork-only) fails if it comes back. Drop this entry when upstream merges #1590 or an equivalent.- The test links with
-Wl,--wrap=realloc, so it is built only on Linux with the static archive and without LTO, liketest_registration_partial_copy. - No library change. No Netflix golden-data, public API or FFmpeg patch impact.
port/upstream-1604-odd-dimension-readback-test — odd-sized frames read back whole (2026-10-01)¶
core/test/test_video_input_odd_dims.c(new, fork-only) is the fork's counterpart of thetest_video_input.cthat Netflix/vmaf PR #1604 adds. It is not a copy: upstream's test expects the picture to carry floor chroma and the reader to skip the rest; the fork'sVmafPicturecarries ceiling chroma, so the test requires the picture to hold every sample the file stores. If upstream merges #1604, do not take itstest_video_input.c, and do not take its skip logic inyuv_input.c/y4m_input.ceither: with ceiling chroma there is nothing to skip, andskip_byteswould always be zero.- What the test depends on:
vmaf_chroma_extent()(core/src/picture_geometry.h) rounding up, and the readers sizing a frame asw * h + 2 * ceil(w / 2) * ceil(h / 2). A rebase that restoresw >> ss_horin the picture geometry fails it withthe picture does not carry the planes the file stores. - The y4m 4:2:2 case is left out on purpose: the y4m reader resamples
C422chroma to the jpeg siting, so its output is not the file's samples. - No reader change. No Netflix golden-data, public API or FFmpeg patch impact.
port/upstream-1602-adm-cm-threshold-shift — the scalar scale-0 ADM masking excess is modular (2026-10-01)¶
core/src/feature/adm_cm_accumulator.hgainsadm_cm_excess_s0(x, thr, shift):|x| - (thr << shift)modulo 2^32, inuint32_t.adm_cm_accum_round()(core/src/feature/integer_adm.c) calls it. Upstream Netflix/vmaf spells the expressionabs(x) - ((int32_t)(thr) << shift_xsub)in itsADM_CM_ACCUM_ROUNDmacro, which is undefined oncethris negative. Upstream-sync note: when re-porting an upstream change to that macro intoadm_cm_accum_round(), keep the helper call. The sanitizer lane stopstest_integer_adm_cm_thresholdon the old expression.- Superseded the same day by ADR-1402 (entry "fix/adm-cm-centre-tap-wrap" above): the helper now clamps in int64, the x86 files and the CUDA, HIP and Metal kernels use the same definition, and the
(int16_t)cast on the centre tap is gone. - No Netflix golden-data, public API or FFmpeg patch impact: scores are bit-identical on every input measured.
docs/upstream-reconcile-2026-10-01 — what an upstream sync can skip (2026-10-01)¶
Checked against the fork's code at master 591d53449, not against ledgers; the evidence is in docs/state.md under "Confirmed not-affected". Upstream head at the time: 6ec23e8f2.
Upstream commits since the September port: take none.
6ec23e8f2(void *arithmetic ininteger_vif.c): the fork'svif_buffers_alloc()already usesuint8_t *. A conflict there is two spellings of one fix; keep ours.3c07efea6(AVX2 casts and lane indexing): the fork already uses_mm256_castps_si256()andextract_epi64_128(). Do not addmm_hadd_epi64()next to it.15f1447c6(no VLAs for MSVC): the fork has none. Itsalloca()calls and theHAVE_MALLOC_H/HAVE_ALLOCA_Hprobes are not wanted. One side change is not in the fork:-EINVALwhen a generated sub-model name is truncated (read_json_model); take it by hand if that loader is touched.295293a76,2f92791c9(bundledya_getopt,getopt_longdetection): the fork hascore/tools/compat/win32/getopt.c. Do not importlibvmaf/src/compat/getopt/.aeaf2877d(-fps_mode passthrough): ported asT-FFMPEG9-VSYNC-REMOVED-2026-09-28. Its__version__change is not taken.
The fork's own upstream pull requests: recognise them if they land. The fork's tree already carries these fixes.
-
1588, #1589, #1599, #1600, #1620, #1621, #1627, #1629: same fix already in¶
the fork. Keep the fork's side of any conflict. One contract is wider here than upstream's and must survive:vmaf_use_feature()consumes its dictionary on a failed copy as well (#1588). -
1590: same fix; the fork's test is
(PR #1663).test_model_collection_growth¶ -
1591: the fork keeps a pool that started at least one worker; upstream's¶
version tears it down. Keeppool_spawn_workers(). -
1601: the fork sums the 16-bit vertical DWT in int64¶
(adm_dwt2_vpass16_tap4()); upstream's version starts an int32 sum from the offset. Either is correct. Do not end up with both. -
1602: do not take piecemeal. The fork has taken its second revision¶
in every path, with its own vector forms (ADR-1402; entry "fix/adm-cm-centre-tap-wrap" above says which hunks to refuse). -
1603: touches
libvmaf/test/checkasm/, which the fork does not carry.¶ -
1604: the fork needs none of its reader changes, and must not take its¶
test_video_input.c; the fork's test istest_video_input_odd_dims(PR #1664). -
1606: patches a VLA the fork replaced with
ModelArrays.¶
port/upstream-15f1447c6-submodel-name-truncation — port sub-model name truncation check (2026-09-30)¶
- Upstream Netflix/vmaf commit
15f1447c6(MSVC: Avoid the use of variable-length arrays (#1428)): the upstream commit avoided VLAs by replacingsprintfwithsnprintfinmodel_collection_parseand returning-EINVALif the generated sub-model name is truncated. - The fork had already eliminated VLAs for MSVC portability in earlier waves, but still cast
snprintfto(void)atcore/src/read_json_model.cpp:760andcore/src/read_json_model.c:771. - Both
read_json_model.cpp(C++23 parser) andread_json_model.c(C parser twin) now checkn < 0 || (size_t)n >= cfg_name_szand return-EINVALon truncation. - Both parser twins cleanly invoke
teardown_models(model, model_collection)on all error paths inmodel_collection_parse_loop, preventing partial model collection leaks. - Regression tests
test_json_model_collection_submodel_name_truncationincore/test/test_model.candtest_model_collection_submodel_name_truncationincore/test/test_model_collection_api.cexercise 10,000 minimal submodels forcing++i == 10000to verify-EINVALand zero leaks. - No Netflix golden-data, score arithmetic, or public API impact.
fix/cli-raw-odd-420 — CLI accepts odd dimensions for raw YUV chroma-subsampled inputs (ADR-1398) (2026-10-01)¶
core/tools/vmaf.cpp:validate_chroma_alignment()formerly rejected odd widths for 4:2:0 and 4:2:2 and odd heights for 4:2:0 withodd width/height %d not allowed...(ADR-0461). Because.y4mpadded dimensions to 16, odd dimensions were already accepted and evaluated using ceiling chroma ((dim + 1) / 2,core/src/picture_geometry.h). Per user decision 2026-10-01 ("Accept both (Recommended)"), the raw YUV reader accepts odd dimensions matching.y4m.validate_chroma_alignment()returns 0.core/tools/test/test_vmaf_option_dict_ownership.sh: Case 2 previously relied on odd height refusal to verify early CLI options cleanup. Case 2 now tests mismatched dimensions (64x64ref vs64x32dist) to test pre-registration failure cleanup without tripping on odd dimensions.core/tools/test/test_vmaf_raw_odd_dims.sh: added positive, negative, and boundary tests (19x19, 1921x1081, 19x20, 20x19, 19x19 422, 1x1 boundary, and truncated file size tests).python/test/vmafx_cli_test.py: addedtest_raw_odd_dimensions_matches_y4m,test_raw_odd_boundary_1x1, andtest_raw_odd_file_size_mismatch_fails_cleanly.- No Netflix golden assertions or C-API ABI impact. Upstream sync notes: keep
validate_chroma_alignment()accepting odd dimensions unless upstream adopts an equivalent or superseding contract.
perf/sycl-cambi-no-scratch — cambi_sycl launch_reset scratch-free on Intel GPUs (2026-10-01)¶
core/src/feature/sycl/integer_cambi_sycl.cpplaunch_reset(): uses an explicit 1Dnd_range<1>(TOTAL_ITEMS = 5 * RADIX_BINS, local size 256) and a scalar select chain fortopkinstead of an indexed array in the lambda closure. Capturingunsigned topk[CAMBI_SYCL_NUM_SCALES]by value caused IGC on the Arc A380 under the Linux xe driver to allocate 1280 B of private stack memory, and the basic 2Drangecaused DPC++ to wrap the launch inRoundedRangeKernelwith 896 B of private memory.- Invariant: both kernels eliminated their private memory (0 B private, 0 B spill in
.zeinfo,RoundedRangeKernelwrapper dropped). Parity is bit-identical to--backend cpuon all tested fixtures (48/48 on Netflix 576x324, 50/50 on BBB 4K, max abs diff 0.0). Throughput at 4K on Arc A380 is 17.05 ms/frame. - When PR #1660 (ADR-1395) lands, the two
cambi_sycllines incore/src/sycl/scratch_ratchet.txtandcambi_syclinkScratchExtractorsincore/src/sycl/scratch_check.cppmust be deleted as nocambi_syclkernel uses scratch memory.
perf/sycl-motion-hbd-no-scratch — scratch-free SYCL motion and motion_v2 kernels above 15 bpc on Arc A380 under xe (ADR-1395) (2026-10-01)¶
core/src/feature/sycl/integer_motion_pipeline_sycl.cpp: fork-only (upstream Netflix/vmaf has no SYCL backend). The 16-bit vertical accumulation pipeline (submit_sad<int64_t>) is specialized withMotionSadHbdKernel, derived fromVmafSyclKernelShape<32, 256>incore/src/feature/sycl/sycl_compat.h. On Intel Arc A380 under the Linuxxedriver, 128-register allocation caused 768 B register spills per thread forint64_tvertical taps, which corrupted calculations without an error (e.g.test_sycl_motion_tiny_frames16-bit 3x3 frame 0 failed). Requesting the 256-entry register file completely eliminates the 768 B spill (spill_size: 0,private_size: 0).core/src/feature/sycl/sycl_compat.h:VmafSyclKernelShape<SG, GRF>(ADR-1395) already landed on master via PR #1660.core/src/sycl/scratch_ratchet.txtandcore/src/sycl/scratch_check.cpp: removedmotion_syclandmotion_v2_syclnow thatsubmit_sad<int64_t>is scratch-free.- Verified on Arc A380 (
ryzen-4090-arc) underxe:test_sycl_motion_tiny_framespasses (8, 10, and 16-bit across all 9 geometries, bit-for-bit exact vs scalar CPU),test_sycl_motion3_parity,test_sycl_motion_add_uv_parity,test_sycl_motion_v2_paritypass, and 50 frames of 16-bit 4K BBB match the CPU reference bit-for-bit (max abs diff 0.0). - No Netflix golden-data, public API or FFmpeg patch impact.
perf/sycl-speed-no-scratch — make SpEED kernels scratch-free on Intel Arc GPUs (ADR-1395) (2026-10-01)¶
core/src/feature/sycl/speed_sycl_pipeline.cpp:RawPlanesreplaces array membersminuend[kMaxChannels]andsubtrahend[kMaxChannels]with scalar pointersminuend_0..3andsubtrahend_0..3. Apick_planeselect chain andRawBound<T>/FloatBoundfunctor classes bind channel pointers once per work-item inlaunch_scaleandlaunch_decimate. Loop unrolling (#pragma unroll) applied to bicubic and lanczos coordinate/weight loops inscale_bicubicandscale_lanczos.- Invariant: all 8 SpEED kernels (
launch_scaleandlaunch_decimateforuint8_tanduint16_twith and withoutRoundedRangeKernel) compile with zero private memory and zero register spill in IGCzeinfo(private_size 0,spill_size 0). - Parity: passes
test_sycl_speed_chroma_parity,test_sycl_speed_singular_parity,test_sycl_speed_chroma_parity_large,test_sycl_speed_temporal_parity, andtest_sycl_speed_temporal_parity_large.speed_gpu_parity.pyyields bit-identical parity (0.000e+00max abs diff) against the CPU reference across all 48 frames of 576x324 and 50 frames of BBB 4K 3840x2160 forspeed_chroma_u,speed_chroma_v,speed_chroma_uv, andspeed_temporal. - No Netflix golden-data, public API or FFmpeg patch impact: GPU kernel implementation only; numerical output is bit-identical to the CPU reference.
fix/cli-pre-registration-opts-leak — the CLI releases its option dictionaries on every exit path (2026-09-30)¶
core/tools/cli_parse.cppcli_free()(upstream-mirror function, fork
body): it frees every feature_cfg[i].opts_dict and every model_config[i].feature_overload[j].opts_dict still in CLISettings, as well as the option buffers. Upstream Netflix/vmaf frees only the buffers and leaks the dictionaries on every early exit; do not take its version back. - core/tools/vmaf.cpp: every hand-off of an opts_dict to libvmaf (use_cli_feature() → vmaf_use_feature(), vmaf_model_feature_overload(), vmaf_model_collection_feature_overload()) clears the settings' pointer with std::exchange first, so cli_free() never frees a dictionary libvmaf took. A new call site that passes an opts_dict to one of those calls must do the same, or the run double-frees. use_cli_feature() puts the options back only when a second vmaf_use_feature() without options also returns -EINVAL (unknown extractor name, the one path on which libvmaf hands them back). - core/test/test_cli_parse.c release_parsed(), core/test/fuzz/fuzz_cli_parse.c and core/tools/test/test_vmaf_option_dict_ownership.sh depend on that contract; the test and the fuzz harness no longer free dictionaries themselves.
fix/code-scanning-include-and-sast — CodeQL include-non-header and universal PR SAST coverage (ADR-1389) (2026-09-30)¶
core/test/test_feature_backend_twin.c: resolves CodeQL alert #1309 (cpp/include-non-header). The test previously unity-includedcore/src/libvmaf.c. It now links againstlibvmafviacore/test/meson.buildand uses narrow internal test accessors declared incore/src/libvmaf_priv.h:vmaf_backend_twin_verdict_for_test,vmaf_context_fake_backend_for_test,vmaf_context_set_gpumask_for_test,vmaf_context_append_registered_feature_extractor_for_test, andvmaf_context_resolve_context_fallbacks_for_test.core/src/libvmaf.c: static test helper functions implementing the above accessors. Internal state and symbols remain private without exposing unwanted ABI surfaces.scripts/ci/check-no-non-header-includes.sh: guardscore/test/against future.c/.cppinclusions. Wired into.pre-commit-config.yamland.github/workflows/rule-enforcement.yml. Unit tests inscripts/ci/tests/test-check-no-non-header-includes.sh..github/workflows/security-scans.yml: resolves Scorecard alert #6 SAST.CodeQL (Actions)now runs unconditionally on all pull requests and pushes, ensuring 100% commit SAST coverage across all PR types (including docs-only PRs) without path filtering or diff-skipping. Documented in ADR-1389..github/workflows/required-aggregator.yml: added'CodeQL (Actions)'torequiredJobNamesso pull requests require green SAST scanning.- No Netflix golden-data, public API or FFmpeg patch impact.
fix/speed-temporal-prescale-overflow — speed_temporal and speed_chroma frame buffers (2026-09-30)¶
core/src/feature/speed.c,speed_temporalinit():frame_sizeisfloat_stride * s->speed_state.dimensions.alloc_height, notfloat_stride * h. Upstream Netflix/vmaf still has thehline (Netflix/vmaf#1626); the same one-line fix is proposed upstream (PR #1627). Upstream-sync note: drop this entry when the upstream fix is ported. Until then, an upstream sync that touches thespeed_temporalinit()must keepalloc_height; restoringhbrings back the heap overrun atspeed_prescaleabove 1, whichtest_speed_frame_buffersreports under ASan (sanitizers.yml).core/src/feature/speed.c,init_chroma(): sizes chroma dimensions viaspeed_chroma_dimensions(), which rounds up odd luma dimensions usingvmaf_chroma_extent(). Fixed inspeed_chroma_cuda.candspeed_chroma_hip.cas well.core/src/feature/cuda/speed_temporal_cuda.c:st_solve_launch_dims()fixes the block thread count overflow whenu_nb > 256(1080p, or 576x324 at prescale 4.0).core/test/test_speed_frame_buffers.c(new, fork-only, float-gated liketest_speed): regression tests forspeed_temporalprescale andspeed_chromaodd sizes under ASan.- No Netflix golden-data, public API or FFmpeg patch impact: output at
speed_prescale <= 1.0is bit-identical, and none of the Netflix reference pairs runs SpEED.
perf/cuda-psnr-hvs-device-convert — psnr_hvs_cuda reads the device pictures directly (CUDA port of ADR-1369) (2026-09-30)¶
core/src/feature/cuda/integer_psnr_hvs/psnr_hvs_score.cu,core/src/feature/cuda/integer_psnr_hvs_cuda.c,core/src/feature/cuda/integer_psnr_hvs_cuda.h: fork-only (upstream Netflix/vmaf has no CUDA twin). The host round trip (issue_d2h_plane,convert_plane,issue_h2d_plane) and the private float planes are gone; the kernel reads the raw 8- to 12-bit samples of the device pictures (hvs_load_block), two threads per 8x8 block, one launch for every plane into oneVmafCudaKernelReadbackbuffer.- Invariant: the per-block arithmetic is the previous CUDA kernel's, and
reduce_hvs_planes()adds each plane's partials in block order infloat(the sum ADR-1361 calibrated). Changing either changes the output:vmaf --feature psnr_hvs_cuda --precision maxbefore and after a rebase must stay bit-identical at 576x324 and 3840x2160. - 4:0:0 input is luma only (as on master and on the CPU);
test_psnr_hvs_yuv400_paritypins it, andtest_psnr_hvs_odd_depth_paritypins 9- and 11-bit parity with the CPU. core/test/test_cuda_module_lifecycle_contract.py:integer_psnr_hvs_cuda.cleftEXPECTED_BUFFER_OWNERS; it owns noVmafCudaBufferof its own any more.
fix/float-moment-gpu-twin-reachable — the CPU float_moment declares the features it writes (2026-10-01)¶
core/src/feature/float_moment.cis an upstream-mirror file. Upstream Netflix/vmaf hasprovided_features[] = {"float_moment", NULL}; the fork has the four namesextract()writes (float_moment_ref1st,float_moment_dis1st,float_moment_ref2nd,float_moment_dis2nd). Upstream-sync note: a conflict here offers the pseudo-name as "theirs". Keep the fork's list: the ADR-1359 twin lookup and model dispatch findfloat_moment_{cuda,sycl,hip,metal}through these names, and with the pseudo-name--backend <gpu> --feature float_momentruns the CPU extractor with a "has no twin" warning. No other line of the file changed.core/src/feature/feature_extractor.{cpp,h}: new internalvmaf_feature_extractor_twin_audit()next tovmaf_feature_extractor_list_audit(). Fork-only file; it is called bycore/test/test_feature_extractor.cand not fromvmaf_init().- Scores do not change on the CPU. No Netflix golden-data, public C API or FFmpeg patch impact. Naming the twin and the CPU extractor together (
--backend cuda --feature float_moment_cuda --feature float_moment) now registers the twin once, like every other CPU and twin pair; it used to run both and fail withfeature "float_moment_ref1st" cannot be overwritten.
chore/rc3-home-gpu-retest — RC3 home GPU retest kit (ADR-1386) (2026-09-30)¶
scripts/dev/rc3-home-gpu-retest.shandscripts/dev/rc3_retest_helpers.py: fork-only harness. Runs verify-and-time commands for open RC3docs/state.mdrows onryzen-4090-arcunder per-device locks (~/.cache/vmafx-locks/{cuda-4090,sycl-a380,hip-gfx1036}.lock). Commands are explicitly encoded per row and backend, never parsed at run time.scripts/dev/tests/test_rc3_home_gpu_retest.pyvalidates argument handling, lock isolation, and row existence contracts againstdocs/state.md.core/test/test_meson_secret_env_sanitization.py:EXPECTED_RUNNER_PATHSregistersscripts/dev/rc3-home-gpu-retest.shas an authorized caller ofscripts/ci/run_meson_test.py.
fix/cuda-pic-prealloc-check — the pinned CUDA picture keeps its allocating state (2026-09-30)¶
core/src/cuda/picture_cuda.cvmaf_cuda_picture_alloc_pinned(): keeppriv->cuda.state = cuda_state;. Upstream Netflix/vmaf master sets onlypriv->cuda.ctxthere, anddefault_release_pinned_picture()then loadsstate->fthrough a NULL state (SIGSEGV in upstream'stest_cuda_pic_preallocationhost-pinned case). The upstream fix is the open Netflix/vmaf#1573, hunk (a). An upstream sync orport-upstream-committhat takes upstream's function body must keep the line.test_pinned_picture_release_uses_the_allocating_stateincore/test/test_cuda_runtime_unwind.cfails without it and runs without a device.core/test/test_cuda_runtime_unwind.c: fork-only. The allocation checks free what they already allocated before they return,test_lifecycle_close_preserves_failed_handles_and_first_errorasserts throughcheck_sync_failed_lifecycle_close()after its frees, and the legacy runner is split into_a,_band_c. With these changes the file is clean under clang-tidy (CUDA lane) and cppcheck.
perf/sycl-adm-no-spill — spill-free integer ADM row reduction on DG2 (ADR-1395) (2026-09-30)¶
core/src/feature/sycl/integer_adm_sycl.cpp: fork-only (upstream Netflix/vmaf has no SYCL). The CSF denominator and contrast measure reductions inlaunch_csf_den_cmrun in two sequential column reduction phases (CSF denominator 3 sums intoden[], then DLM and AIM contrast measures 6 sums intocm[]), staging sub-group partials into local memory before folding. Do not combine them into a single loop: keeping all nine 64-bit accumulators live across the column loop causes IGC to spill 864 B/thread at SIMD16 on DG2 (dg2-g11, Intel Arc A380), and under the Linuxxedriver scratch memory corrupts accumulator reads and produces zero sums, failingtest_sycl_adm_parity(T-SYCL-ADM-CM-SCRATCH-2026-09-30).- Both phases compile with zero private memory and zero spill memory on
dg2-g11,adl-s, andbmg-g21. - NASA JPL Rule 4: helper functions
adm_dev_den_px,adm_dev_cm_px,adm_dev_sg_partials,adm_dev_fold_rows, andlaunch_csf_den_cmeach remain under 25 LOC. core/test/test_adm_cm_row_rounding_contract.pycontinues to validate the fold inadm_dev_fold_row.
perf/sycl-psnr-hvs-no-scratch — scratch-free SYCL psnr_hvs kernel on Arc A380 under xe (ADR-1395) (2026-09-30)¶
core/src/feature/sycl/integer_psnr_hvs_sycl.cpp: fork-only (upstream Netflix/vmaf has no SYCL). The kernel eliminates all private memory and register spills on DG2 at SIMD16 (private_size: 0,spill: 0). Dynamic plane indexing inhvs_locate()was replaced byhvs_pick()to avoid spilling the 152-bytePsnrHvsKernelArgsstruct into private memory. The 4x4 sub-block accumulator arrays (means[4],variances[4]) inhvs_variance_ratio()were restructured into scalar members (HvsQuadrants) with unrolled row-sum walks in the CPU's traversal order. Do not reintroduce dynamic array indexing or thread-private arrays into the kernel lambda; on the Linuxxekernel driver on Intel Arc A380, any private memory access produces corrupted values (~20 dB drift at 4K).- Verified with
oclocinspecting.ze_infoacrossdg2-g11(Arc A380, SIMD16),adl-s(UHD 770, SIMD8), andbmg-g21(Arc B580, SIMD16). - Verified parity on physical Arc A380 (
ryzen-4090-arc) underxe: 576x324 Netflix pair within 8.37e-5 dB (gate 5e-4), 1080p within 1.71e-3 dB, 4K BBB (22 frames) frame 0psnr_hvs_y33.161817 dB (CPU 33.171624 dB, delta 0.0098 dB). 4K BBB(t(22) - t(2)) / 20throughput: 12.55 ms/frame (corrupted) -> 10.90 ms/frame (correct). - No Netflix golden-data, public API or FFmpeg patch impact.
fix/sycl-aot-check-single-target — the AOT image check accepts bare native images (2026-09-30)¶
core/src/sycl/check_aot_image.py: fork-only (upstream Netflix/vmaf has no SYCL).oclocwrites an image per TU in one of two forms: anarfat binary when-devicenames two or more acronyms (even two that share an IP version), a bare zebin (plain ELF) when it names one. The check must accept both; never narrow it back toararchives, which failed every single-target build (-Dsycl_icpx_aot_targets=dg2-g11, the single-target example indocs/backends/sycl/overview.md). A TU thatsycl_icpx_aot_igc_skipleaves with one target is a bare zebin too.- A bare zebin's IP version is the u32 of its
.note.intelgt.compatIntelGT note of type 6 (IntelGTSectionType::productConfigin intel/compute-runtimeshared/source/device_binary_format/zebin/zebin_elf.h), laid out asHardwareIpVersioninshared/source/helpers/hw_ip_version.h(architecture bits 31:22, release 21:14, revision 5:0) and printed asarchitecture.release.revision, the tokenocloc idsprints. The bare zebin is modelled as an image carrying that one IP version, so the completeness and--partialrules, and thecount incomplete <= declared partial TUsrule, apply to both forms unchanged.elf_sections(),parse_images()andcheck()are the three pieces; keep the walker strict (zero padding is the only skippable byte) and keep thearmember alignment relative to the archive start. - Summary text is unchanged for an all-fat-binary library (
31 spir64_gen fat binaries; ...); a bare zebin addsN spir64_gen native images. The stamp content is informational, nothing parses it. core/test/test_sycl_aot_image_check.pygains synthetic bare-zebin fixtures (zebin(),intelgt_notes(),ip_word());elf64()takes section types and, as a list of pairs, repeated names. No device or oneAPI is needed. No Netflix golden-data, public API or FFmpeg patch impact; the Meson wiring (sycl_aot_image_checkincore/src/meson.build) is unchanged.
fix/ssimulacra2-reject-yuv400 — CPU ssimulacra2 refuses 4:0:0 at init (2026-10-01)¶
core/src/feature/ssimulacra2.cis fork-only (upstream Netflix/vmaf has nossimulacra2); no Netflix golden-data, public C API or FFmpeg patch impact.init()now returns-EINVALforVMAF_PIX_FMT_YUV400P/VMAF_PIX_FMT_UNKNOWNbefore it allocates anything; keep that check first ifinit()is restructured, sinceconvert_picture_to_linear_rgband every SIMDpicture_to_linear_rgbread U and V unconditionally.core/test/test_ssimulacra2_coverage.c::test_ssimulacra2_rejects_yuv400guards it.- Touching the file made it subject to the HISS touched-file rule, so four functions were split without changing any arithmetic (the TU builds with
-ffp-contract=off):create_recursive_gaussiancallssolve_cramer_3x3,picture_to_linear_rgbtakes its constants fromyuv_matrix_coeffs,init/closesharealloc_buffers/free_buffers(thegoto failpath is gone), andextractcallsscore_one_scale/downsample_both. The twoNOLINTNEXTLINE(readability-function-size)lines they carried are gone with the long bodies. Output is bit-identical to master (every frame of five fixtures x fouryuv_matrixvalues x AVX-512 / AVX2 / scalar).
fix/cambi-short-frame-oob — CAMBI c-values walks stay inside short and narrow frames (Netflix/vmaf#1628) (2026-09-30)¶
- Partly a port, partly fork-local.
core/src/feature/cambi.c(c_values_first_pass,c_values_top_edge,c_values_bottom_edge) andcore/src/feature/x86/cambi_avx2.c(the_avx2twins) carry the loop bounds of Netflix/vmaf#1629:MIN(pad_size, height),MIN(pad_size + 1, height)andMAX(height - pad_size, 0). The fork split upstream's singlecalculate_c_values()into those helpers, soc_values_first_pass/c_values_first_pass_avx2gained aheightparameter; upstream's diff does not apply verbatim. When #1629 merges upstream, a sync keeps the fork's helpers and checks the three bounds are present in both files. - If upstream instead rejects such frames in
init()(the alternative Netflix/vmaf#1628 raises), keep the fork's clipping: a 1920x160 frame is valid video that CAMBI can score with the window clipped to the rows that exist, the same clipping every frame gets at its top and bottom edge, and rejecting it would makecambifail on inputs that only needed these bounds (Research-2132 has the comparison). - Fork-local, no upstream counterpart yet: the first column loop of every row step in the same helpers (
c_values_first_pass,c_values_top_edge,c_values_middle_slide,c_values_bottom_edgeand their_avx2twins) runs toMIN(pad_size, width)instead ofpad_size. Upstream's loops read the columns past a frame narrower thanpad_size; a sync that takes upstream's loops must keep this bound ortest_calculate_c_values_narrow_framefails. - Fork-local:
core/src/feature/cambi_c_values_frame.h(cambi_calculate_c_values_frame, the walk the AVX2 scan, AVX-512 and NEON drivers share) has the same three row bounds and already visits only the columns belowwidth. The header's comment already says a change to the walk incambi.c/cambi_avx2.cmust be mirrored there; this is such a change. core/test/test_cambi.c:test_calculate_c_values_short_frameis the fork's version of upstream's sentinel test, rewritten for the fork's test style and extended: every driver the host has, heights 1 to 10, and a from-scratch reference for each in-frame c-value.test_calculate_c_values_narrow_frameis its column twin (widths 1 topad + 2, stale content past the width). Keep the SIMD gates of both onvmaf_get_cpu_flags_x86();vmaf_get_cpu_flags()is 0 in this binary. Upstream's version gates its AVX2 leg onvmaf_get_cpu_flags(), so on a sync do not take it over the fork's.check_c_values_avx2_parity()in the same file now readsvmaf_get_cpu_flags_x86(). #1479 made that change on its branch, but when it was rebased onto #1483 (merged earlier the same day) it kept its comment and took #1483's helper with the old gate, so the code change never reached master (T-CAMBI-AVX2-PARITY-GATE-LOST-IN-REBASE-2026-09-30).core/test/test_cambi_stage_simd.c: the frame sweep adds 1-row andpad-row heights for the production configuration.- No public API, FFmpeg patch or Netflix golden-data impact; golden gate and
python/test/cambi_test.pypass unchanged.
fix/vmaf-init-output-only-handle — vmaf_init never reads the incoming handle (ADR-1396) (2026-09-30)¶
core/src/libvmaf.cvmaf_init(): fork body. It sets*vmaf = NULLright after thevmaf == NULLcheck and*vmaf = vonly on success. Upstream assigns*vmaf = malloc(...)up front and leaves the freed pointer there when set-up fails; keep the fork's order. Do not bring back ADR-1032'sif (*vmaf) return -EINVAL;guard.core/test/test_context.c:test_vmaf_init_ignores_the_incoming_handleandtest_vmaf_init_overwrites_an_open_handlereplacetest_vmaf_init_double_init_guard.- Public header documentation only; no symbol, signature or FFmpeg patch change (the FFmpeg filters pass a zeroed
LIBVMAFContextmember). No Netflix golden-data impact.
feat/cuda-float-ssim-scale — float_ssim_cuda decimates and convolves like the CPU (ADR-1399) (2026-10-01)¶
- All touched library files are fork-only (
core/src/feature/cuda/integer_ssim_cuda.c,core/src/feature/cuda/integer_ssim/ssim_score.cu); no upstream-mirror file changes, and no Netflix golden-data, public C API or FFmpeg patch impact. ssim_score.cunow mirrors three CPU files. A sync or rebase that changes any of these on the CPU side must change the kernel in the same PR, ortest_cuda_float_ssim_parity(equality with the CPU) andtest_cuda_float_ssim_decimate(planes byte for byte) fail:core/src/feature/ssim.c: the automatic scale rule,ssim_low_pass_alloc()'s tap1.0f / (float)(scale * scale), and the in-placeiqa_decimate()of both planes;core/src/feature/iqa/decimate.c/convolve.c:iqa_filter_pixel()'s window offsets,KBND_SYMMETRIC, the fp32prodand itsdoublesum; andiqa_convolve_1d_separable()'s fp32 product,doublesum and one(float)rounding per pass;core/src/feature/iqa/ssim_tools.c:ssim_precompute_scalar()'s fp32 products (already mirrored for the combine by ADR-1373).- The plane size comes from
core/src/feature/iqa/decimate_dim.h, which the SYCL twin shares (ADR-1370). Keep that header include-free. - NVCC compiles
ssim_score.cuwith FMA contraction on. Every rounding in it is an intrinsic (__fmul_rn,__dadd_rn,__double2float_rn,__ll2float_rn); a conflict resolution that rewrites one as a plain*or+changes scores.test_cuda_kernel_source_contract.pypins them. - The 8bpc and 16bpc pass-1 kernels lost their unused
widthargument and acalculate_ssim_horiz_planesand twocalculate_ssim_decimate_*kernels were added; the launch argument arrays ininteger_ssim_cuda.cfollow the new signatures (ADR-1215). Resolve a conflict in one file together with the other. float_ssim_hipandfloat_ssim_metalstill implement scale 1 only; the HIP counterpart is a separate branch.core/test/test_gpu_float_ssim_auto_scale_contract.py,core/test/test_feature_backend_twin.candcore/tools/test/test_vmaf_feature_backend.shlist which backends decimate: a rebase across the HIP change keeps both.
perf/cuda-ssimulacra2-device-resident — device-resident ssimulacra2_cuda (ADR-1391) (2026-10-01)¶
- All touched files are fork-only (upstream Netflix/vmaf has no CUDA ssimulacra2 and no
ssimulacra2extractor); no Netflix golden-data, public C API or FFmpeg patch impact. core/src/feature/cuda/ssimulacra2_cuda.cis now submit/collect with no host stage. The host pipeline (ss2c_stage_raw_planes, host YUV / XYB, the per-scale downloads,ss2c_host_combine, the host downsample) is gone, and so aressimulacra2/ssimulacra2_mul.cu, the ADR-0456 kernels (ssimulacra2_blur_h3,ssimulacra2_transpose,ssimulacra2_blur_v3_transposed) and themodule_mul/ssimulacra2_mulfatbin. A sync that brings any of them back reverts ADR-1391. The device kernels aressimulacra2/ssimulacra2_device.cu(YUV, XYB, combine, downsample) andssimulacra2/ssimulacra2_blur.cu(ssimulacra2_blur_h/_v), with the by-value argument structs inssimulacra2_cuda.h.- A change to
ssimulacra2.c'spicture_to_linear_rgb,linear_rgb_to_xyb,fast_gaussian_1d,multiply_3plane,downsample_2x2,ssim_map,edge_diff_maporpool_scoremust be mirrored inssimulacra2_device.cu/ssimulacra2_blur.cu(ssimulacra2_yuv_to_linear,ssimulacra2_xyb,ss2c_iir_step,ss2c_blur_input,ssimulacra2_downsample,ss2c_accumulate) andss2c_pool_score, then re-checked withscripts/dev/speed_gpu_parity.py --backend cuda --feature ssimulacra2 --max-abs-diff 1e-9 --vmaf "$PWD/build/tools/vmaf"andtest_cuda_ssimulacra2_parity(both at 1e-9). core/src/feature/ssimulacra2_math.h/ssimulacra2_score.hgained theVMAF_SS2_FUNCqualifier hook andssimulacra2_eotf_lut.h(generated byscripts/gen_ssimulacra2_eotf_lut.py, which emits the hook) theVMAF_SS2_EOTF_LUT_STORAGEhook; host code keeps the oldstatic inline/static const. The CUDA twin compiles these headers as device code, so keep them free of host-only calls. They are listed incuda_kernel_shared_headersincore/src/meson.build, so every CUDA fatbin rebuilds when one changes; keep them there.core/src/meson.build:cuda_cu_sourcesswapsssimulacra2_mulforssimulacra2_device, which also gets acuda_cu_extra_flagsentry (vmaf_cuda_host_strict_fp_args + ['--fmad=false']);core/test/test_strict_fp_compiler_args.pyasserts it.
fix/sycl-rc3-parity — RC3 SYCL parity follow-ups (2026-09-30)¶
core/src/feature/sycl/integer_ssim_sycl.cpp: fork-only (upstream Netflix/vmaf has no SYCL). Fixed flat/identical-frame handling to reproduce CPU scoring without shortcuts, grouping((w*a)*b)/denand preserving ADR-1370 fp32 frame-mean rounding.core/src/feature/sycl/integer_psnr_sycl.cpp: fork-only. AddedVMAF_FEATURE_EXTRACTOR_TEMPORALflag for correct--subsamplebehavior.core/src/feature/sycl/integer_motion_v2_sycl.cpp: fork-only. Aligned FPS weighting incollect()and motion score clipping with CPU reference.- No Netflix golden-data, public API or FFmpeg patch impact.
ci/mingw-ucrt64 — migrate Windows MinGW CI leg to UCRT64 (ADR-1387) (2026-09-30)¶
.github/workflows/libvmaf-build-matrix.yml: the MSYS2 Windows matrix leg migrated frommsystem: MINGW64withmingw-w64-x86_64-*tomsystem: UCRT64withmingw-w64-ucrt-x86_64-*packages, resolving the deprecation warning frommsys2/setup-msys2v2.33.0 (#1609, ADR-1387). Matrix job display name renamed fromWindows MinGW64toWindows UCRT64..github/workflows/required-aggregator.yml: required status check context updated from'Windows MinGW64'to'Windows UCRT64'. The two must remain identical (scripts/ci/check-aggregator-names.sh).docs/getting-started/building-on-windows.md: manual MSYS2 prerequisite command updated to installmingw-w64-ucrt-x86_64-*packages under UCRT64.- No Netflix golden-data, public C API or FFmpeg patch impact.
ci/release-pat-mode-gate-exemption — PAT-mode release PR authoring gate exemption (ADR-1388) (2026-09-30)¶
scripts/ci/release-pr-exempt.sh: added--diff,--diff-file,--base,--head, and--pat-useroptions and corresponding environment variablesDIFF_FILE,BASE_SHA,HEAD_SHA, andRELEASE_BOT_PAT_USER. Evaluates PR diff against the approved release file set (manifest, config, changelog, changelog.d, and dynamically parsedextra-filesversion markers) when the PR author matches the designated PAT user (lusoris). Non-release diffs or unauthorized authors fail closed..github/workflows/rule-enforcement.yml: exportedBASE_SHA: ${{ github.event.pull_request.base.sha }}andHEAD_SHA: ${{ github.event.pull_request.head.sha }}tosteps.release_pracross all five authoring-discipline jobs (deliverables-check,doc-substance-check,state-md-check,silent-revert-check,ffmpeg-patches-surface-check).scripts/git-hooks/pre-push-pr-body-lint.sh: generates diff againstorigin/masterwhen pushing arelease-please--*branch so local pre-push evaluation mirrors CI.docs/adr/1388-release-pat-mode-gate-exemption.md: documents dual-path exemption protocol (bot author vs PAT author with verified release-only diff).- No impact on Netflix golden data, SIMD/GPU kernels, public API or FFmpeg patches.
perf/hip-psnr-hvs-device-convert — HIP psnr_hvs native sample upload and device conversion (ADR-1369 port) (2026-09-30)¶
core/src/feature/hip/integer_psnr_hvs_hip.c: fork-only (upstream Netflix/vmaf has no HIP backend). Replaces host-side float conversion loops with native sample upload viavmaf_hip_picture_upload(). Device buffersd_ref/d_distare sized to sample width (1 byte for 8 bpc, 2 bytes for 9–12 bpc). Removed unused pinned host buffersh_uint_refandh_uint_dist(eliminating 6 redundant allocations).core/src/feature/hip/integer_psnr_hvs/psnr_hvs_score.hip: fork-only. Replacedconst float *kernel parameters withconst void *andint wideflag (0 for 8 bpc uint8_t, 1 for 16 bpc uint16_t). Kernel reads and converts raw samples directly on the device. Resolves a latent scaling bug on 9-bit and 11-bit depths. Fuses plane dispatches into a single kernel (n_dispatches_per_frame = 1), reducing 4K frame time to 18.90 ms/frame on AMD gfx1036.core/test/test_hip_psnr_hvs_parity.c: addedtest_psnr_hvs_deep_parityasserting exact parity against CPU for 9, 10, 11, and 12-bit inputs.- No Netflix golden-data, public API or FFmpeg patch impact.
fix/sycl-scratch-free-kernels — SYCL kernels use no scratch memory (ADR-1395) (2026-10-01)¶
core/src/sycl/scratch_check.{h,cpp}(new, fork-only): the two probe kernels, the warning-only self-testvmaf_sycl_state_initcalls, and the kernel audit behindtest_sycl_kernel_scratch.kScratchExtractorsthere and the extractor column ofcore/src/sycl/scratch_ratchet.txtmust name the same extractors; the test compares them. A PR that clears a ratchet kernel deletes its line and, when it was an extractor's last, the extractor fromkScratchExtractors. Never add a line to make the test pass.- The ratchet list keys on mangled kernel ids. A launcher whose name or parameter types change produces a new id: the test then reports the kernel as unlisted (fails if it still uses scratch) and the old line as not registered. Update the line with the id the FAIL message prints.
core/src/feature/sycl/sycl_compat.h:VmafSyclKernelShape<SG, GRF>andVMAF_SYCL_FUNCTOR_SG_SIZE. Under icpx the sub-group size of a functor derived from it lives in itsget(properties_tag); do not also put thereqd_sub_group_sizeattribute on its call operator (icpx warns and ignores one of them).core/src/feature/sycl/integer_vif_sycl.cpp: fork-only. The horizontal and fused launchers submit the functorsIntegerVifHoriKernel/IntegerVifFusedKernel; their SIMD-32 instances take the 256-entry register file (vif_grf_size()). Keep it: without it they spill up to 8832 bytes per thread and score 0/0 on an Arc A380 under xe.VMAF_SYCL_VIF_SUBGROUP_SIZEreads throughvmaf_gpu_dispatch_env_get()(ADR-0488).core/test/meson.build:test_sycl_kernel_scratch(ratchet path passed asVMAF_SYCL_SCRATCH_RATCHET) andtest_sycl_vif_parity_sg32(test_sycl_vif_paritywithVMAF_SYCL_VIF_SUBGROUP_SIZE=32).- No Netflix golden-data, public API or FFmpeg patch impact: the new functions are internal, and the two environment variables are documented in
docs/backends/sycl/overview.md.
perf/hip-ssimulacra2-device-resident — device-resident SSIMULACRA2 on HIP (ADR-1390) (2026-09-30)¶
core/src/feature/hip/ssimulacra2_hip.c,core/src/feature/hip/ssimulacra2/ssimulacra2_device.hip: fork-only (upstream Netflix/vmaf has no HIP backend).ssimulacra2_hipruns the complete frame on the device with one raw plane upload insubmit(), on-device YUV-to-linear, XYB, IIR Gaussian blurs with a tiled shared-memory row pass (SS2H_ROW_TILErows, single-wave blocks, two-slot ring, register prefetch), exact fp32-pair per-pixel SSIM and edge sums over a deterministic LDS reduction tree, 2x2 downsampling, and one 864-byte readback incollect().- Built with
-ffp-contract=offincore/src/meson.buildto preserve bit-level agreement. - Numerical contract: within 1e-9 of CPU reference at
--precision max(Netflix 576x324: 1.123e-12, BBB 4K: 5.826e-13). - No Netflix golden-data, public API or FFmpeg patch impact.
fix/state-md-three-way-resolver — three-way docs/state.md conflict resolver (ADR-1383) (2026-09-30)¶
scripts/dev/resolve-state-md-conflict.py: fork-only (upstream Netflix/vmaf has nodocs/state.md). It reads the conflicted path's index stages (:1:base,:2:ours,:3:theirs), never the conflict markers, and merges rows and move tombstones by bug id, disposition rows by label, and every other line three-way by line. Do not restore the old "ours wins, the branch adds only unseen ids" rule: mid-rebase ours already contains the branch's earlier commits, so that rule keeps stale rows.ID_PATTERN,ROW_REandTOMBSTONE_REmirror the id and tombstone shapes inscripts/ci/check-state-md-rows.sh; a change to one belongs in the other in the same PR.DISPOSITION_SECTIONmust match the##heading of the disposition table indocs/state.md.- The tool writes bytes with LF endings (
write_bytes), notwrite_text, which turns every line ending into CRLF on Windows. scripts/dev/test-resolve-state-md-conflict.pyruns in thestate.md row hygiene (ADR-0165)step of.github/workflows/rule-enforcement.ymland inscripts/ci/test_git_fixture_isolation.py. It scrubs everyGIT_*variable before it creates a repository; keep it that way.- No Netflix golden-data, public API or FFmpeg patch impact.
fix/cuda-rc3-parity — CUDA motion order, CPU option tables, tiny-frame guards, one atomic per block (ADR-1372, ADR-1373, ADR-1374, ADR-1392) (2026-09-30)¶
core/src/feature/cuda/integer_motion_sad_cuda.{h,c}(new, fork-only): the one host path to the motion SAD kernel ofinteger_motion_v2/motion_v2_score.cu, used bymotion_cudaandmotion_v2_cuda.integer_motion/motion_score.cu(upstream NVIDIA blur-each-frame kernel) is deleted with itsmotion_score_ptxtarget. An upstream sync that brings backmotion_score.cu,calculate_motion_scoreor the blurred ping-pong reintroduces the blur-then-diff order; keep the diff-first kernel andtest_cuda_motion_tiny_frames(compares with==). If upstream changesmotion_score_pipeline_8/_16ininteger_motion.c, mirror it inmotion_v2_score.cuonce (both CUDA motion twins).integer_motion_cuda.c: raw-luma ping-pongraw[2]replacesblur[2];submit()passes the previous frame'seventasprev_done, so the copy and the kernel are ordered on the device. The ADR-0845 batch readback (motion_readback_slots()) synchronises once; do not restore the extracuStreamSynchronizebefore the copies. The debugVMAF_integer_feature_motion_scoreisMIN(sad * mfw, mmxv), as ininteger_motion.c::extract.integer_adm/adm_dwt2_rows.h(new): row and tap arithmetic ofadm_dwt2.cu;adm_dwt2_load_column()andcalculate_indices()call it, and the scale-0 load clamps throughcuda_tile_index.h. Upstream-mirror NVIDIA code changed here: on a sync, keep the header calls and thestatic_asserts that tie theDWT_8_VERT_HORI(4, 16, 32768, 128, 8, ...)instantiation to the header's geometry.test_cuda_adm_dwt2_rowsreplays the arithmetic device-free.integer_vif_cuda.c:context_check+context_fallback_name = "vif"and aninit()guard belowvif_cuda_min_dim()(16), before any CUDA state is read (ADR-1324 pattern, asvif_sycl).- Options (ADR-1373):
integer_psnr_cuda.ccallspsnr_score.hfor every score and gained aflushforapsnr_*;ssim_cuda.candinteger_ssim_cuda.cemit throughvmaf_ssim_max_db()/ thenonfinite_score.hemitters;integer_ssim/ssim_score.cu::ssim_terms()computes the CPU'sl * c * swith the CPU's types and rounding points (mirror ofiqa/ssim_tools.candiqa/ssim_accumulate_lane.h: when an upstream sync changes either, changessim_terms()in the same PR), its partials are doubles andinteger_ssim_cuda.crounds the frame means to fp32; it gainedcalculate_ssim_vert_combine_lcs.core/src/meson.buildbuildsinteger_ssim_scorewith--fmad=false(cuda_cu_extra_flags, the HIP twin's-ffp-contract=off), andinteger_ssim_score.cugroups each term asinteger_ssim.cdoes.float_ssim_cudakeepsenable_chromaas an ignored option (HISS-14);float_motion_cuda.croutes every score throughmotion_clip().integer_motion_v2_cuda.cpublishes the CPU's weighted, capped SAD and derivesmotion2_v2/motion3_v2from it likeinteger_motion_v2.c::flush.integer_psnr_cuda.cis TEMPORAL and zeroes every plane accumulator on the picture stream. - Tests:
core/test/test_cuda_module_lifecycle_contract.pyinventory listsinteger_motion_sad_cuda.cas the motion module owner;test_device_target_header_dependencies.pycounts 21 CUDA fatbin targets. core/src/libvmaf.c(engine, found by the RTX 4090 run):init_before_dispatch()initialises an extractor that hassubmit()andcollect()beforeread_pictures_cuda_submit_current()andread_pictures_dispatch_one()choose between the asynchronous path andextract(). The motion twins'init()swaps inextract()undermotion_force_zero; with the choice made first, the first frame called the clearedsubmit()and crashed. A sync or refactor of those two functions must keep the init ahead of the decision;test_cuda_kernel_source_contract.pypins the order, andtest_cuda_motion_tiny_frames/test_cuda_twin_option_parityrunmotion_force_zeroon a device.float_motion_cuda.c(found by the RTX 4090 run,T-GPU-FLOAT-MOTION3-MISSING-2026-09-30): providesVMAF_feature_motion3_scoreand declaresmotion_blend_factor/motion_blend_offsetin the CPU table's order;motion_blend_clip()isfloat_motion.c::motion_blend_clip. If upstream changes the CPU motion3 (blend, index-0 or flush emission), change the twin in the same PR;test_cuda_float_motion_paritycompares every frame of all three scores.- ADR-1392 (kernel reductions):
integer_motion_v2/motion_v2_score.cucomputes the vertical pass once per block (vertical_pass()intos_v) and adds one atomic per block (add_block_sad());integer_psnr/psnr_score.cusums eight pixels per thread (geometry ininteger_psnr_cuda.h, shared withpsnr_cuda_dispatch()), adds one atomic per block (add_block_sse()) and selects the plane with constant indices (plane_row());integer_moment/moment_score.cutakes the same layout (integer_moment_cuda.h) and adds one atomic per accumulator per block (add_block_sums()). Keep all of it on a sync of these upstream-derived NVIDIA kernels: per-warp atomics to a single accumulator serialise the kernels, andpic.data[plane]on the by-value kernel parameter puts both pictures on every thread's stack. - clang-tidy clean-up of touched CUDA hosts (HISS-04, no behaviour change):
integer_vif_cuda.cgainedvif_submit_plane()/vif_submit_scales()and a kernel-name table invif_get_filter1d_functions();integer_ssim_cuda.csplitinit_fex_cuda()intofloat_ssim_check_geometry()/float_ssim_load_kernels(). On a conflict keep the helpers and move the upstream statement into them in order.
No public C API, CLI syntax or FFmpeg patch impact. CPU scores are bit-identical (no CPU extractor changed; the engine change only moves the first-frame init() of a submit / collect extractor ahead of the dispatch decision). motion_cuda scores move to the CPU's (measured on an RTX 4090: from 1.26e-5 to 0.0 on the Netflix pair, from 6.9e-5 to 0.0 on 50 frames of 3840x2160); motion_v2_cuda is unchanged; default float_ssim_cuda moves towards the CPU's by an fp32 rounding; default integer_ssim_cuda per-pixel terms move to the CPU's (no fused c1, c2, y2 * w; the CPU's grouping); psnr_cuda, float_motion_cuda and motion_v2_cuda default scores are unchanged (motion_v2_cuda moves only with a non-default motion_fps_weight / motion_max_val or a one-frame input); float_motion_cuda output gains motion3, and the ADR-1392 kernels return the same integers as before.
perf/sycl-adm-aim-device — AIM pass on the SYCL integer ADM twin (ADR-1362) (2026-09-29)¶
core/src/feature/sycl/integer_adm_sycl.cpp: fork-only (upstream Netflix/vmaf has no SYCL). The twin claimsVMAF_integer_feature_aim_score/_adm3_scoreagain. Per scale:launch_decouple_csfwritesd_csf_f(|csf(t - r)| / 30) andd_csf_f_aim(|csf(r)| / 30);launch_csf_den_cmis one work-group per region row with nine sums (CSF denominator, DLM and AIM for h, v, d), folded once per row throughadm_cm_round_row_total(). An upstream change toadm_csf()/i4_adm_csf(),adm_cm()/i4_adm_cm(), themeasure_aimrole swap, the threshold macros or the scale finalisation ininteger_adm.chas to be mirrored inadm_dev_*and inadm_cm_scale_cpu()/adm_den_scale_cpu()in the same sync;test_sycl_adm_parityandtest_sycl_adm_tiny_framesfail on any aim / adm3 bit difference.- Keep the int64 clamp in
adm_dev_decouple_k(); narrowing the quotient to int32 before the clamp is the old defect (integer_adm_scale2up to 1.40e-6 off the CPU at 4K). - Every output (adm2,
integer_adm_scale*, debug num / den, aim, adm3) uses the CPU-float finaliser (adm_scale_cpu/adm_terms/adm_finalise). The twin's old double finaliser is deleted on purpose; a conflict resolution that restoresconclude_adm_cm/conclude_adm_csf_denbreaks the bit-exact tests. core/test/test_adm_cm_row_rounding_contract.pynow pins the SYCL fold inadm_dev_fold_row(was an inline expression inlaunch_csf_den_cm_3band)..standards-baseline.json: re-recorded 190 -> 185; the five over-long functions of the old file are gone and the two DWT launchers moved lines. Re-record from a clean tree after a conflict; never edit it by hand.- No Netflix golden-data, public API or FFmpeg patch impact.
perf/sycl-psnr-hvs-light-twins-4k — SYCL twins share the uploaded planes (ADR-1369) (2026-09-29)¶
core/src/sycl/common.cpp/common.h: fork-only (upstream Netflix/vmaf has no SYCL). NewSyclSharedChromamember ofVmafSyclState,vmaf_sycl_shared_chroma_init/_upload,vmaf_sycl_get_shared_plane,vmaf_sycl_queue_after_upload, andsycl_fence_slot_readers()called first insidevmaf_sycl_shared_frame_upload'stry. Keep the chroma upload out ofvmaf_sycl_shared_frame_upload(luma-only runs must not pay for it) and keep the fence before the ref-plane upload; the ref-before-dis order andsycl_enqueue_plane_upload'sstatic_cast<unsigned>are unchanged.core/src/feature/sycl/integer_psnr_sycl.cpp: fork-local. The chroma staging buffers (d_chroma_*,h_chroma_*,stage_chroma_plane) are gone. Rebased onto ADR-1365 (#1624), which owns the option table,configure_scores,emit_planeandflush_fex_syclof the same TU; this change ownsallocate_chroma,psnr_pre_graph,psnr_post_graph,launch_sseand the chroma upload insubmit_fex_sycl. A conflict keeps both halves.core/src/feature/sycl/integer_psnr_hvs_sycl.cpp: fork-local kernel; per-block float expressions must stay verbatim (bit-identity).integer_motion_v2_sycl.cpp: host path only —curfrom the shared frame, the ADR-1371 pipeline'scur_copyfor the ping-pong; the kernel stays ininteger_motion_pipeline_sycl.cpp.core/test/test_sycl_init_unwind.cpp+core/test/meson.build: the--wrap=vmaf_sycl_shared_chroma_initlink argument and its wrapper go together.
perf/cli-frame-readahead — per-input reader threads in the vmaf CLI (ADR-1366) (2026-09-29)¶
core/tools/vmaf.cpp: upstream Netflixvmaf.ckeeps the inlinefor (picture_index = 0 ;; picture_index++)loop that callsfetch_picture()for the reference and then the distorted frame. The fork's loop isscore_frames()(bounded by--frame_cntorUINT_MAX) fed by twoFrameReaders, andrun_frame_loop()only starts, stops and joins them. A sync conflict there resolves to the fork's side; port an upstream change to how a frame is read intofetch_picture(), which the reader threads and the inline path both call, and an upstream change to the per-frame scoring step intoscore_frames().preallocate_cli_pictures()sizes the pool as2 * (threads + 1) + 1plus2 * kReadaheadDepthwhen read-ahead is on. If upstream changes the base term, keep the read-ahead term added to it.core/tools/meson.build: thevmaf/vmafxtargets takevmaf_cli_deps = vmaf_tool_deps + [thread_lib]forstd::thread.- Keep
request_stop()on both readers before eitherjoin(), and the ring slot reservation before the pool fetch; seecore/tools/AGENTS.md§Frame read-ahead.
fix/sycl-fp-contract-all-tus — one strict FP line for every SYCL feature TU (ADR-1367) (2026-09-29)¶
- Folds ADR-1363's
sycl_exact_fp_args/sycl_exact_fp_sources(#1627) intosycl_strict_fp_args. If a sync brings either name back, or any per-TU FP list (extra_args) into thesycl_feature_sourcesloop, drop it:core/test/test_strict_fp_compiler_args.pyandcore/test/test_sycl_kernel_source_contract.pyfail on both. core/src/meson.build:sycl_fp32_prec_argsandsycl_strict_fp_argslive between# BEGIN/END VMAF SYCL strict FP policy, beforesycl_dependency. icpx order-fp-model=precise -ffp-contract=off -foffload-fp32-prec-div -foffload-fp32-prec-sqrtis load-bearing (precise implies contraction on).sycl_link_argsis['-fsycl']plus the precision pair wherever the icpx driver links; ADR-1360's "link carries-fsyclonly" now means no device targets at the link. The SPIR-V JIT image is still device-linked there, so a sync that returns the link to plain-fsyclmakes-Dsycl_icpx_aot_targets=builds (the SYCL parity lane) approximate again;test_sycl_fp_arith_contractfails on that path. The MSVC build (ADR-1364) links withlink.exeand generates every image in its explicit device link, which takessycl_strict_fp_argswhole; keep it there.core/src/feature/sycl/integer_vif_sycl.cppdev_vif_stats_log_domain():sv_sqissycl::fma(-g, sigma12, sigma2_sq)on purpose. The CPU computes it in fp64; with contraction off a separate fp32 product moves the scores further from the CPU. Keep the explicitfmaif the VIF statistic is re-synced from upstream or the CUDA/HIP twins.core/test/test_sycl_fp_arith_probe.cpp+test_sycl_fp_arith_contract.c: the probe is compiled by acustom_targetwithsycl_toolchain_args + sycl_feature_tail_args, so it follows the feature line automatically; do not give it private flags. Under the MSVC device link it gets its own explicit device link with libvmaf's arguments, becauselink.exenever wraps device code.docs/state.md:T-SYCL-FP-MODEL-PRECISE-CONTRACTS-2026-09-29closed;T-CUDA-FP-CONTRACT-DEFAULT-2026-09-29,T-HIP-FP-CONTRACT-DEFAULT-2026-09-29andT-CLI-FLOAT-MOMENT-NO-TWIN-2026-09-29opened (RC3).- No Netflix golden-data, public C API or FFmpeg patch impact; CPU code is unchanged.
perf/sycl-ssimulacra2-msssim-device-resident — device-resident ssimulacra2_sycl, single-wait float_ms_ssim_sycl (ADR-1363) (2026-09-29)¶
- All touched files are fork-only (upstream Netflix/vmaf has no SYCL, and its
ssimulacra2does not exist); no Netflix golden-data, public C API or FFmpeg patch impact. core/src/feature/sycl/ssimulacra2_sycl.cppis now submit/collect and holds no host stage: a sync that restoresss2s_host_combine,ss2s_host_linear_rgb_to_xyb,ss2s_downsample_2x2orss2s_picture_to_linear_rgb, or adds await()outsidecollect_fex_sycl, failscore/test/test_sycl_kernel_source_contract.py. A change tossimulacra2.c'spicture_to_linear_rgb,linear_rgb_to_xyb,fast_gaussian_1d,downsample_2x2,ssim_map,edge_diff_maporpool_scoremust be mirrored in the device functions of the same names (ss2s_yuv_pixel,ss2s_xyb_pixel,ss2s_iir_step,ss2s_down_pixel,ss2s_ssim_term,ss2s_edge_terms,ss2s_pool_score) and re-checked withscripts/dev/speed_gpu_parity.py --backend sycl --feature ssimulacra2 --max-abs-diff 1e-9.core/src/feature/ssimulacra2_math.h:vmaf_ss2_cbrtfdivides throughVMAF_SS2_FDIV(a, b), default((a) / (b)), so host code is unchanged; keep the hook if the cube root is rewritten.core/src/feature/sycl/sycl_exact_fp.his the one home ofdiv_rn,sqrt_rnand the fp32-pair helpers, moved verbatim out ofspeed_sycl_pipeline.cpp(which now has using-declarations) plusff_negandff_div. Its slow paths callsycl::ext::intel::math::fdiv_rn/fsqrt_rnonly under__SYCL_DEVICE_ONLY__: DPC++ also emits a host copy of every kernel body, and the extension's host fallback inlibsycl-devicelib-host.aneeds libm symbols the static test links do not resolve (test_speed,test_sycl_ssimulacra2_parityfailed to link before). The SYCL feature TUs are custom targets without header dependency tracking: after editing the header, touchssimulacra2_sycl.cppandspeed_sycl_pipeline.cppin an incremental build.core/src/meson.buildrenamessycl_speed_strict_fp_args/sycl_speed_sourcestosycl_exact_fp_args/sycl_exact_fp_sourcesand addsssimulacra2_sycl; a branch still using the old names must follow the rename (the source contract checks the list).core/src/feature/sycl/integer_ms_ssim_sycl.cpp:d_l/c/s_partialsandh_l/c/s_partialsbecame oned_partials/h_partialspair with a per-(plane, scale)partial_offset; the horizontal and vertical passes moved fromcollect()tosubmit()(enqueue_scale_lcs), andcollect()sums (sum_scale_lcs) after one wait. Any change that adds a plane or a scale must extend the offset table inallocate_ms_ssim_buffers.scripts/dev/speed_gpu_parity.pygained--featureand--max-abs-diff; the defaults keep the SpEED behaviour.
fix/sycl-aot-at-link — SYCL AOT images built at compile time, checked after link (ADR-1360) (2026-09-29)¶
core/src/meson.build: fork-only (upstream Netflix/vmaf has no SYCL). The AOT argument listsycl_icpx_aot_base_argscarries-fno-sycl-rdcand--offload-compress; without-fno-sycl-rdcthe link, which passes only-fsyclthroughsycl_dependency, drops everyspir64_genimage. Do not move the target flags tosycl_dependency.link_argsinstead: link-time AOT reruns device codegen for all oflibvmaf.aat each of its 113 test links, and one ocloc crash aborts the whole link. Configure errors when AOT targets are set andoclocis missing; thesycl_aot_image_checkcustom target runscore/src/sycl/check_aot_image.pyonlibvmaf.so.sycl_icpx_aot_igc_skip(the Xe2 targets ofinteger_psnr_hvs_sycl) is the only allowed gap; remove the entry when the reworked kernel offix/sycl-b580-psnr-hvs-adm-tinylands, whichever branch merges second.build-config.envINTEL_NEO_VERSIONreplacesdev/Containerfile'sARG NEO_VER; the Containerfile NEO step sources the copied config and also installsintel-ocloc, and Renovate's compute-runtime manager watchesbuild-config.env. Never reintroduceARG NEO_VER.dev/scripts/fetch-intel-neo.py --components oclocfeedsscripts/ci/install-intel-ocloc.sh, used by every Linux SYCL CI leg and bydocker/Dockerfile.production-gpu's oneAPI builder.scripts/ci/gen-sycl-compile-commands.pystrips--offload-compress; the clang-tidy SYCL job configures-Dsycl_icpx_aot_targets=and needs no ocloc.
fix/release-provenance-attest — GitHub build-provenance attestations replace slsa-github-generator (ADR-1356) (2026-09-29)¶
.github/workflows/supply-chain.yml: fork-only; upstream Netflix/vmaf has no release provenance. Theslsa-provenance/mcp-slsa-provenancereusable-workflow calls are gone;provenance/mcp-provenancerun SHA-pinnedactions/attest-build-provenanceinrelease-publish, andattach-to-releaseattaches*.sigstore.jsonbundles. Never restoreslsa-framework/slsa-github-generatoror any tag-referenced action: the VMAFx organisation enforcessha_pinning_required, including on actions inside called reusable workflows (the v1.0.0-rc.2 failure).- This supersedes item 5 of
fix/scorecard-pins-best-practicesbelow ("SLSA GitHub generator must remain tag-pinned") and the "single permitted exception" in the SHA-pin entries further down: there is no exception any more, and the.github/AGENTS.mdsync-gate grep no longer filters the generator. scripts/release/tests/test-publication-environment-binding.shpins the job names, permissions, environment, subjects, verification step and bundle names; rename them together with the workflow.
perf/sycl-cambi-device-resident — device-resident cambi_sycl (ADR-1357) (2026-09-29)¶
core/src/feature/sycl/integer_cambi_sycl.cpp: fork-only. The twin now reimplementscambi.c's c-values and top-K pooling on the device instead of calling them on the host. An upstream Netflix change toc_value_pixel, thecalculate_c_valueswindow walk (c_values_first_pass,_top_edge,_middle_slide,_bottom_edge),spatial_pooling,cambi_preprocessingorfilter_modemust be mirrored into the device kernels in the same sync;test_sycl_cambi_parityfails on any per-frame difference.core/src/feature/cambi.c/cambi_internal.h: one fork-added trampoline,vmaf_cambi_reciprocal_lut(), in the fork block at the end ofcambi.c; the upstream-mirror body is unchanged. If upstream regeneratesreciprocal_lutincambi.h, the device picks the new table up through the accessor;test_cambi's "differs from1.0f / i" assertion documents the current table and may need updating.- The twin's init repeats
cambi.c's reciprocal-LUT window guard (check_window_fits_lut, same -EINVAL and message, encode and source windows). If upstream changes that guard or the table size, change both. - No other backend changes; the CUDA, HIP and Metal twins keep their host residual (RC3 rows in
docs/state.md).
fix/cli-feature-backend-twin — --feature runs the explicit --backend's twin (ADR-1359) (2026-09-29)¶
core/include/libvmaf/libvmaf.h,core/src/libvmaf.c: fork-only, additive.vmaf_feature_backend_twin()andvmaf_registered_feature_extractor()carryVMAF_EXPORT(ADR-0379). An upstream sync that rewrites the tail oflibvmaf.hor thevmaf_use_features_from_model*()block oflibvmaf.cmust keep both.vmaf_use_feature()keeps upstream's exact-name contract.core/src/feature/feature_extractor.{h,cpp}:vmaf_get_feature_extractor_twin()is the only CPU-to-twin pairing. It reusesvmaf_get_feature_extractor_by_feature_name()and must require the backend flag on the result, because that lookup's ADR-0530 second pass can return an extractor of another backend or the CPU one. Never add a name-mangling table.core/tools/vmaf.cpp:register_cli_feature()keeps the ADR-0543 suffix gate beforecli_feature_extractor();explicit_backend_requested()delegates tocli_backend_is_device()incore/tools/cli_feature_backend.cpp, the singleauto/cputest.write_cli_output()buildsbackend_usedfrom the registered extractors after the final flush; do not restore the oldactive_backend_name()from the initialised states. Upstream Netflix has nobackend_usedkey.ffmpeg-patches/: unaffected; the filters callvmaf_use_feature()by exact name and no patch touches the new symbols.
perf/sycl-ciede-throughput — ciede_sycl stages chroma at native size (2026-09-29)¶
core/src/feature/sycl/integer_ciede_sycl.cpp: fork-only.submit()packs Y, U and V at their native size (stage_plane(), chroma bypicture.c's ceil rule) and the kernel reads chroma at(x >> ss_hor, y >> ss_ver). Do not bring back the hostupscale_planefrom before this change or floor the chroma size. The horizontal index followsss_horand the vertical oness_ver, matching the fork's fixedciede.c::scale_chroma_planes, not upstream's transposed pair.core/test/test_sycl_ciede_parity.ctakesFIXTURE_PIX_FMTandFIXTURE_BPC;core/test/meson.buildbuilds the_oddw,_422_10band_444variants. Keep them with the TU.scripts/ci/tidy-baseline-sycl.json: scoped tightening ofinteger_ciede_sycl.cppfrom 14 to 0 (clang-tidy 22.1.8). On a conflict, reruntidy-ratchet.py --onlyrather than merging the JSON by hand.
perf/sycl-speed-device-resident — SYCL SpEED twins are device-resident (ADR-1358) (2026-09-29)¶
core/src/feature/sycl/speed_sycl_pipeline.cppholds every SpEED kernel;speed_chroma_sycl.cpp,speed_temporal_sycl.cppandspeed_sycl_host.cpphold none and do not wait on the queue outsidepipeline_collect()/pipeline_wait().core/test/test_sycl_kernel_source_contract.pyenforces this, the absence of the fp64 type in the pipeline, and the absence of calls to the host linear algebra (speed_internal_compute_eigenvalues,_qr_factorize,_qt_multiply,_filter_and_downscale,picture_copy).- The pipeline is bit-identical to
speed.conly while every sum keeps the reference order and every product feeding an add stays in a named temporary; division and square root go throughdiv_rn()/sqrt_rn(), log2 throughspeed_log2(), the fp64EIGENVALUE_EPScomparisons throughbelow_eps()/below_eps_scaled(). An upstream change tospeed.c(eigen sweep, QR, scoring,vif_tools.cfiltering) must be mirrored there and re-checked withscripts/dev/speed_gpu_parity.py --backend sycl. core/src/meson.build: the four SpEED TUs build withsycl_speed_strict_fp_args(-ffp-contract=offafter-fp-model=precise). Keep the order; do not move the flags onto the shared SYCL feature line without re-measuring every other twin (T-SYCL-FP-MODEL-PRECISE-CONTRACTS-2026-09-29).speed_internal.cgainsspeed_internal_entropy_constant()andspeed_internal_base_entropy()(same expressions asspeed.c), declared in the newcore/src/feature/speed_constants.h;speed.candspeed_internal.hare untouched. The CUDA/HIP twins still use the host helpers.- The four SpEED SYCL TUs keep their helpers in many short anonymous-namespace blocks and define the
speed_sycl::API with qualified names: the HISS-04 scanner counts a namespace block as one function, so a block over 60 lines fails the touched-file gate. - No Netflix golden-data, public API or FFmpeg patch impact.
fix/ffmpeg9-fps-mode — -fps_mode passthrough replaces -vsync 0 (2026-09-28)¶
compat/python-vmaf/core/executor.py: ports Netflix/vmafaeaf2877d; the decode command now matches upstream. Upstream'spython/vmaf/__init__.py__version__bump to 4.0.0 is deliberately not ported (ADR-1127): resolve a sync conflict there by keeping the fork's release-owned version.mcp-server/vmaf-mcp/src/vmaf_mcp/server.pyandcmd/vmafx-mcp/impl.go: fork-only. Never reintroduce-vsync; FFmpeg 9, whichbuild-config.envpins, rejects it.
fix/cli-unescape-values-svm-swap — libsvm uses std::swap for libc++ 23 (2026-09-28)¶
core/src/svm.cpp: libsvm's globaltemplate <class T> void swap(T &, T &)is replaced byusing std::swap;. libc++ 23__split_buffer::__swap_layoutscallsswapunqualified afterusing std::swap, so forstd::vector<svm_node>argument-dependent lookup found both templates and the call was ambiguous (upstream issue 1616; libc++ 21.1.8 and 22.1.8 qualify the call and were unaffected). Do not restore the template on a libsvm re-vendor.- The upstream fix (upstream PR 1617) also replaces libsvm's
min/maxwithstd::min/std::max. The fork keeps libsvm's pair: on a tie or a NaN operand libsvm returns the second argument and the standard ones the first, which would change training results (the sign of a zero rho, NaN propagation in working-set selection). Scoring never reaches them. A port of that commit should take theswaphunk only, unless the semantic change is decided separately. Solver_NU'sselect_working_set,calculate_rhoanddo_shrinkingcarryoverride, which clears clang's-Winconsistent-missing-overridenow thatSolveis marked. Keep them on a re-vendor.- The check that this stayed correct: the Netflix 576x324 pair scores byte-identically at
--precision maxagainst the pre-change binary.
fix/cli-unescape-values-svm-swap — CLI option values keep their backslashes (ADR-1355) (2026-09-28)¶
core/tools/cli_parse.cpp:cli_unescape()is split intocli_unescape_key()(ADR-1190's\:\=\.\\, for keys, the--featurename and both halves of an overload key) andcli_unescape_value()(values: a backslash is data unless it belongs to a run directly before:/=or at the end of the value, which is read in pairs). Upstream Netflix still splits these strings withstrsepand has neither function. Invariant: never route a value through the key unescaper —..\,\\serverand\.cachepaths lose bytes — and keep the value pairing in step withcli_split(), which treats a:after an odd run of backslashes as literal.pkg/cliopt:EscapeValueandcli_unescape_value()are one grammar in two languages; change both, and the round-trip test'sunescape/splitmirrors, in the same commit.core/test/test_cli_parse.c: the seven ADR-1355 cases run through their ownrun_value_backslash_testsrunner (sevenmu_run_testexpansions is thereadability-function-sizeceiling, ADR-0141).ffmpeg-patches/: unaffected; the filter takes the remainder after the first=and never splits on:.
agent/fix-codex-hook-paths-3139 — keep Codex hooks worktree-relative (2026-09-25)¶
All seven commands in .codex/hooks.json must retain the quoted, repository-local-Git-environment-clearing $(env -u GIT_DIR -u GIT_WORK_TREE ... git rev-parse --show-toplevel)/.codex/hooks/<script>.sh form. Do not restore the retired /home/kilian/dev/vmaf path, substitute a new absolute checkout, or reduce the command to a launch-directory-relative path during a configuration regeneration. Keep the exact event/matcher matrix, script executable modes, and the test-codex-hook-config pre-commit/pre-push caller together.
- Research digest: Codex hook-path portability audit.
- Decision matrix: no ADR needed; only-one-way broken-path correction.
- AGENTS.md invariant:
scripts/ci/AGENTS.md, “Codex repository-hook path contract”. - Reproducer / smoke:
python3 -B scripts/ci/tests/test_codex_hook_config.py. - Changelog:
changelog.d/fixed/codex-hook-path-portability.md. - FFmpeg impact: none; no public C header, CLI flag, Meson option, or patch surface changed.
agent/fix-research-0033-0034-3139 — ratchet research identifier drift (2026-09-25)¶
Commit 5ac5b4167 renamed the HIP-applicability and CI-pipeline-audit digests from 0033/0034 to 0432/0433, but a later collector merge restored the old files and their old links alongside the renamed copies. Preserve only the 0432/0433 files and targets. Preserve check-research-digest-ids.py, its generated exact-debt baseline, planted red-cap tests, path-complete pre-commit hooks, and required Rule Enforcement invocation together. ADR-1335 binds the baseline to the trusted merge base: branch JSON must exactly match its tree and may only reduce trusted debt. Bootstrap remains an explicit one-time operator mode and never belongs in CI or hooks.
- Research digest: Research-2114 records the laundering reproducer, trusted-base delta, and adversarial matrix.
- Decision matrix: ADR-1335.
- AGENTS.md invariant:
scripts/ci/AGENTS.md, “Research-digest identifier ratchet”, plusdocs/research/AGENTS.mdheader-normalization guidance. - Reproducer / smoke:
python3 -B scripts/ci/check-research-digest-ids.pyandpython3 -B scripts/ci/tests/test_research_digest_ids.py. - Changelog:
changelog.d/fixed/research-digest-0033-0034-resurrection.md. - FFmpeg/public-surface impact: none; no public C header, CLI flag, Meson option, score, model, snapshot, FFmpeg patch, benchmark, tuning, or retraining surface changed.
agent/meson-secret-env-sanitize — sanitize secret environment variables in Meson tests (2026-09-25)¶
Meson test execution inherits host environment variables by default and writes the raw parent mapping to build/meson-logs/testlog.txt before applying test setups. Preserve scripts/ci/run_meson_test.py and every inventoried Make, workflow, preflight, bisection, setup-guidance, and Zed caller so those entry points delete sensitive GitHub credential keys (GITHUB_PERSONAL_ACCESS_TOKEN, GITHUB_TOKEN, GH_TOKEN, GH_ENTERPRISE_TOKEN, GITHUB_ENTERPRISE_TOKEN, GITHUB_PAT, GH_PAT, GITHUB_AUTH_TOKEN, GITHUB_API_TOKEN, HOMEBREW_GITHUB_API_TOKEN, ACTIONS_ID_TOKEN_REQUEST_TOKEN, ACTIONS_RUNTIME_TOKEN) before Meson starts. core/meson.build retains the same denylist in the sole default test setup for the child and JSON-log boundary. Meson can select an alternate setup and applies per-test environments after the setup; the regression contract therefore inventories all supported callers, rejects direct test-target bypasses across shell, multiline YAML (plain and quoted keys), and Python implicit list/tuple continuations, requires the default to remain the only add_test_setup under core/, and rejects explicit forbidden-name reintroduction. Its subprocess probes default to a load-tolerant 120-second deadline; the optional override accepts only finite values from 60 through 300 seconds and fails closed otherwise. Probes never copy arbitrary host variables and inspect only disposable synthetic logs. Raw external Meson/Ninja commands remain an explicit unsupported bypass.
- Research digest: Research-1333.
- Decision matrix: ADR-1333.
- AGENTS.md invariant:
core/AGENTS.mdanddocs/development/rebase-sensitive-invariants.md, "Meson test secret environment sanitization". - Reproducer / smoke:
python3 -m unittest core.test.test_meson_secret_env_sanitizationandpython3 scripts/ci/run_meson_test.py -- -C build test_meson_secret_env_sanitization. - Changelog:
changelog.d/security/1333-meson-test-secret-env-sanitization.md. - FFmpeg impact: none; no public C header, CLI flag, Meson option, or patch surface changed.
agent/sycl-motion-uv-tolerance — fixed-point oracle for SYCL motion-add-UV (2026-09-25)¶
test_sycl_motion_add_uv_parity must compare motion_sycl with the scalar fixed-point oracle, not with CPU float_motion. Preserve the five integer coefficients, reflect-101 mapping, vertical and horizontal rounding stages, exact per-plane SAD, YUV420 geometry, and normalized binary64 error bound in the oracle. Both the 256x144 and registered 960x540 variants are required; a rebase must not restore the former large-fixture exclusion or an empirical absolute tolerance. Production SYCL source is unchanged.
- Research digest: Research-2112.
- Decision matrix: ADR-1326.
- AGENTS.md invariant:
core/test/AGENTS.md, “SYCL motion-add-UV fixed-point oracle”; andcore/src/feature/sycl/AGENTS.md, “motion_add_uv fixed-point parity”. - Reproducer / smoke:
ONEAPI_DEVICE_SELECTOR=level_zero:gpu meson test -C build-sycl --no-rebuild --print-errorlogs test_sycl_motion_add_uv_parity test_sycl_motion_add_uv_parity_large. - Changelog:
changelog.d/fixed/sycl-motion-add-uv-fixed-oracle.md. - FFmpeg impact: none; no public C header, CLI flag, Meson option, or patch surface changed.
agent/adm-cm-rounding-observable — preserve raw row-fold observability (2026-09-25)¶
Integer-ADM contrast masking must apply shift_inner_accum exactly once after the complete row reduction in the scalar CPU reference, AVX2/AVX-512, CUDA, HIP, SYCL and Metal. Preserve the private adm_cm_round_row_total() seam, all 72 inline x86 SIMD band folds, the equivalent inline SYCL fold, the MSL-local twin, the ten scalar/GPU call shapes, and the CUDA/HIP embedded-kernel header dependencies when resolving upstream reduction changes. Preserve the seam's signed rounding argument because CUDA i4 passes ADR-0155's negative term. Score parity is not evidence for this invariant because the later float conversion erases one-unit placement errors.
- Research digest: Research-2111.
- Decision matrix: existing ADR-1167; no new decision was required.
- AGENTS.md invariant:
core/src/feature/AGENTS.md, “integer_adm row-level rounding invariant”, plus the HIP exception andcore/test/AGENTS.mdguard. - Reproducer / smoke:
python3 core/test/test_adm_cm_row_rounding_contract.pyandmeson test -C BUILD test_adm_cm_row_rounding. - Changelog:
changelog.d/fixed/adm-cm-row-rounding-observability.md. - FFmpeg impact: none; no public header, exported API, CLI flag or Meson option changed.
audit/rc1-flake-survey-df0b — keep scheduled CI aligned with required lanes (2026-09-25)¶
The scheduled whole-tree CPU ratchet must retain the required PR lane's GCC 15, clang-tidy 22, -Db_lto=false, and explicit analyzer path. The standalone and sanitizer fuzz workflows both compile full libvmaf with Clang 22 and ASan; keep their 30-minute job budgets aligned when resolving workflow conflicts.
- Research digest: Research-2107.
- Decision matrix: ADR-1321.
- AGENTS.md invariant: existing
.github/AGENTS.mdclang-tidy repository rule andscripts/ci/AGENTS.mdconfigured-native-lint rules; no new invariant. - Reproducer / smoke:
python3 -B scripts/ci/test_fail_closed_ci.py. - Changelog:
changelog.d/fixed/rc1-nightly-clang-tidy-fuzz-timeout.md. - FFmpeg impact: none; no public C header, CLI flag, Meson option, or patch surface changed.
agent/merge-train-runtime-closure-df0b — close merge-train runtime migration (2026-09-25)¶
Closes T-MERGE-TRAIN-CONTROL-2026-09-08 after verifying the local merge-train runtime migration under ADR-1244. Live runtime adapters in /home/kilian/dev/vmafx/vmafx/.claude/mergetrain (train.sh, rebase-clean.sh, watchdog.sh, merge_train_operator.py) are hash-bound to committed gateway a7a58dd8f39576dc2b0a86fb5af518003496b147261dc5294706cbce976f39c7 (blob 4ca44da40edfc03fc500174a193b293ca8cfdc3e), backed by immutable receipt migration-xghk30zh. Legacy unrestricted actors remain absent; foreign train processes belong to their actual external cwd; held and non-master PRs fail closed; worktree and branch ownership are protected; release PR #1213 is excluded; required checks and exact-head validation receipts fail closed; and all 26 disposable Git/Make regression tests pass. Corrects stale operator documentation paths in docs/development/merge-train.md.
- Research digest: Merge-train control investigation.
- Decision matrix: ADR-1244.
- AGENTS.md invariant:
scripts/dev/AGENTS.md, "Developer control scripts". - Reproducer / smoke:
python3 -m unittest discover -s scripts/dev/tests -p 'test_*merge_train_guard.py'. - Changelog:
changelog.d/fixed/merge-train-control-closure.md. - FFmpeg impact: none; no C/C++ or library surface touched.
fix/gpu-float-ssim-auto-scale-1e6f — dimension-aware model fallback (2026-09-25)¶
CUDA, SYCL, HIP and Metal float_ssim remain scale-1-only GPU kernels, but a model-selected host-picture context must no longer fail at common dimensions when scale=0 resolves above 1. Preserve each descriptor's context_check and context_fallback_name = "float_ssim", the model-only allow_context_fallback marker, and the resolver call after picture validation but before CUDA translation or extractor initialization. Only -ENOTSUP requests fallback; parser errors and directly named GPU extractors retain their existing errors. A replacement must clone the option dictionary, rebind context-owned backend state and invalidate cached CUDA residency flags.
- Research digest: Research-2108.
- Decision matrix: ADR-1324.
- AGENTS.md invariant:
core/src/feature/AGENTS.md, “Model options gate GPU twin selection”. - Reproducer / smoke:
python3 core/test/test_gpu_float_ssim_auto_scale_contract.pyandmeson test -C BUILD test_feature_collector. - Changelog:
changelog.d/fixed/gpu-float-ssim-auto-scale-fallback.md. - FFmpeg impact: none; no public C header, exported API, CLI flag or Meson option changed.
agent/fix-barten-mode1-9e1b — normalize integer ADM Barten weights (2026-09-25)¶
Integer ADM uses one shared power-of-two normalization exponent per DWT scale when CSF weights exceed the original fixed-point arithmetic budget. Preserve the strict 2^16 scale-0 and 2^30 scale-1..3 limits, apply the same exponent to all three bands, and restore 3k in every CPU/CUDA/SYCL/HIP/Metal contrast-masking finalizer. The denominator must continue to use the original floating-point CSF factors. Already-representable k=0 configurations retain their existing values and CPU SIMD dispatch; normalized CPU configurations use the scalar weighted-CSF/CM stages while keeping unrelated SIMD stages.
Metal integer ADM now implements CSF modes 0..3. Do not restore its former VMAF_OPT_FLAG_DEFAULT_ONLY bit or mode-0 init rejection. The four GPU float_adm twins remain mode-0-only.
- Research digest: Research-2109.
- Decision matrix: ADR-1325.
- AGENTS.md invariant:
core/src/feature/AGENTS.md, “Integer ADM Barten weights use one exponent per scale”. - Reproducer / smoke:
meson test -C BUILD test_adm_csf_representable; then run each builttest_{cuda,sycl,hip}_adm_tiny_framesparity executable on its own idle device. - Changelog:
changelog.d/fixed/integer-adm-barten-fixed-point-normalization.md. - FFmpeg impact: none; no public header, C API, CLI flag, Meson option, or FFmpeg patch changed.
agent/fix-doxygen-public-api-warnings-6ba5 — drive public C API Doxygen warnings to zero and fail closed (2026-09-25)¶
Public C headers in core/include/libvmaf/*.h are now strictly warning-free under core/doc/Doxyfile.public-api (WARN_AS_ERROR = YES) and CI workflow .github/workflows/doxygen-public-api.yml (DOXYGEN_WARNING_CEILING: "0"). Multi-variable declarations (unsigned w, h;) in public structs must remain split into separate lines with individual doc comments so Doxygen attaches docs to all members. Unrecognized @field tags must not be reintroduced (use inline /**< ... */). Standardized @note Thread safety: replaces invalid @thread-safety annotations. The vendored pelorus/ mirror remains excluded from the public C API documentation scope. When rebasing or resolving conflicts in public headers, ensure complete per-member documentation and verify with python3 -B core/test/test_gpu_public_header_docs.py and doxygen core/doc/Doxyfile.public-api.
- Research digest: Research digest.
- Decision matrix: ADR-1315.
- AGENTS.md invariant:
core/include/libvmaf/AGENTS.md, "Doxygen-clean public API". - Reproducer / smoke:
python3 -B core/test/test_gpu_public_header_docs.pyanddoxygen core/doc/Doxyfile.public-api. - Changelog:
changelog.d/fixed/1315-doxygen-public-api-fail-closed.md. - FFmpeg impact: none; no symbol was removed or renamed; ABI structs were not modified.
agent/fix-gpu-runner-label-a7f77d — fail-closed hardware admission (2026-09-25)¶
sycl-arc and gpu-full are deliberately different runner capabilities. Preserve sycl-parity.yml as the sole hardware float_ssim owner; do not restore SYCL float_ssim Parity under tests-and-quality-gates.yml or relabel the isolated Arc runner as gpu-full. Every self-hosted job must remain behind a hosted probe of its complete runs-on label set, and the aggregator must require success whenever that lane's switch is true.
- Research digest: GPU runner admission.
- Decision matrix: ADR-1319.
- AGENTS.md invariant:
scripts/ci/AGENTS.md, “Self-hosted hardware admission invariants (ADR-1319)”. - Reproducer / smoke:
python3 -B scripts/ci/test_self_hosted_runner_workflow_contract.pyandbash scripts/ci/tests/test-runner-available.sh. - Changelog:
changelog.d/fixed/1319-self-hosted-gpu-admission.md. - FFmpeg impact: none; no public C header, CLI flag, Meson option, or FFmpeg-patch surface changed.
agent/gpu-option-value-capability-d5df — value-aware model fallback (2026-09-25)¶
VmafOption distinguishes canonical schema from an extractor's narrower implementation capability with VMAF_OPT_FLAG_DEFAULT_ONLY. Preserve that bit on the entries enumerated by core/test/test_gpu_option_value_capability_contract.py until the corresponding kernel implements the option's non-default values. ADR-1325 implemented modes 0..3 for Metal integer ADM and therefore removed that entry; the remaining eight restrictions are still authoritative. Do not narrow or remove the CPU-mirrored name, alias, default, range, or FEATURE_PARAM bit: those fields remain collector-key authority.
vmaf_use_features_from_model() must call the value-aware helper before creating a GPU context. A valid non-default value falls back to the CPU for that feature; a valid default stays on device; malformed values remain the normal parser's error; explicitly selected GPU extractors retain their direct -EINVAL. No public C surface or FFmpeg patch changes.
The structured context-create failure helper owns the cloned extractor's private allocation. Preserve its unconditional private-state free: the NIQE unknown-option regression is LeakSanitizer-red without it (264 bytes) and green with it.
- Research digest: Research-2105.
- Decision matrix: ADR-1316.
- AGENTS.md invariant:
core/src/feature/AGENTS.md, “Model options gate GPU twin selection”. - Reproducer / smoke:
python3 core/test/test_gpu_option_value_capability_contract.pyandmeson test -C BUILD test_feature_extractor test_feature_collector. - Changelog:
changelog.d/fixed/gpu-option-value-capability-fallback.md. - FFmpeg impact: none; no public header, C API, CLI flag, or Meson option changed.
agent/fix-golden-gate-build-dir-6ba5 — isolate Netflix golden gate build profile to prevent ICX FP drift (2026-09-25)¶
make test-netflix-golden now uses an isolated CPU-only build profile in GOLDEN_BUILD_DIR ?= core/build-golden, compiled via scripts/ci/setup-golden-build.sh enforcing an explicitly supported compiler (gcc or clang). compat/python-vmaf/__init__.py and config.py read VMAF_BUILD_DIR from the environment to decouple the test harness from core/build. When resolving rebase conflicts in Makefile or compat/python-vmaf/__init__.py, preserve GOLDEN_BUILD_DIR, target build-golden, and VMAF_BUILD_DIR injection in test-netflix-golden.
- Research digest: Research-1317.
- Decision matrix: ADR-1317.
- AGENTS.md invariant:
AGENTS.md§8 andcompat/python-vmaf/AGENTS.md, "Isolated build profile (ADR-1317)". - Reproducer / smoke:
make test-netflix-goldenandpytest python/test/golden_gate_isolation_test.py scripts/ci/tests/test_golden_gate_makefile_contract.py. - Changelog:
changelog.d/fixed/golden-gate-icx-drift-build-isolation.md. - FFmpeg impact: none; no public C API, headers, or CLI flags changed.
agent/fix-pershot-input-ceiling-139c — operator frame ceiling for vmaf-perShot (ADR-1318) (2026-09-25)¶
vmaf-perShot now provides -F, --frames <N> (with aliases --frame_cnt and --max-frames) defaulting to 0U (unbounded compatibility contract). Bounded scans on FIFOs, streams, and /dev/zero now terminate promptly and cleanly with exit code 0 instead of reading ~4.29e9 frames or appearing hung. The scan loop tracks frame_idx in uint64_t and probes one additional read at the built-in boundary before indexing it, resolving the ADR-1287 UINT32_MAX off-by-one check so an input of exactly UINT32_MAX complete frames is accepted when EOF is reached. core/tools/vmaf_per_shot_input.c deliberately consumes luma and chroma exactly: a seek beyond regular-file EOF is not evidence that a raw frame exists, and ferror must never be collapsed into clean EOF. Preserve the reduced-boundary test target when resolving Meson conflicts.
- Research digest: Research-1318.
- Decision matrix: ADR-1318.
- AGENTS.md invariant:
core/tools/AGENTS.md, "Scan stops atVMAF_PER_SHOT_MAX_FRAMESor--framesceiling". - Reproducer / smoke:
meson test -C core/build test_vmaf_per_shot. - Changelog:
changelog.d/fixed/pershot-endless-input-ceiling.md. - FFmpeg impact: none; no public libvmaf header, C API, or scoring behavior changed.
agent/fix-gpu-option-aliases-hiss-984823 — preserve CPU collector-key aliases (2026-09-25)¶
CUDA, SYCL, and HIP option tables now use the same force_0, ks, and ssclz aliases as their CPU reference extractors. These spellings are load-bearing: ADR-1183 puts a non-default option's alias into the published feature key. When resolving an upstream or backend-table conflict, preserve the CPU alias and extend core/test/test_gpu_option_alias_contract.py for any new equivalent twin option.
- Research digest: Research-2104.
- Decision matrix: ADR-1312.
- AGENTS.md invariant:
core/src/feature/AGENTS.md, “Twin option tables mirror the CPU's aliases and semantics”. - Reproducer / smoke:
python3 core/test/test_gpu_option_alias_contract.py. - Changelog:
changelog.d/fixed/gpu-option-alias-parity.md. - FFmpeg impact: none; no public header, C API, CLI flag, or Meson option changed.
agent/semgrep-registry-advisory-edbf — isolate moving registry packs from the merge gate (ADR-1314) (2026-09-25)¶
.github/workflows/security-scans.yml has two deliberately different Semgrep authorities. Preserve the repository-owned .semgrep.yml upload through github/codeql-action/upload-sarif; it creates the required Semgrep OSS check. Preserve the unpinned registry-pack output as an ordinary actions/upload-artifact artifact; restoring category: semgrep-registry would again let an upstream pack update block every pull request without a VMAFx diff.
- Research digest:
docs/research/1314-semgrep-registry-sarif-routing.md. - Decision matrix: ADR-1314 compares accepting the shared gate, removing the scan, a second Code Scanning configuration, and artifact-only retention.
- AGENTS.md invariant:
.github/AGENTS.mdrecords the SARIF authority split. - Reproducer:
python3 -m unittest scripts.ci.test_security_workflow_contract scripts.ci.test_fail_closed_ci. - Changelog:
changelog.d/security/semgrep-registry-advisory-boundary.md.
agent/reconcile-hip-scaffold-state-8763 — reconcile duplicate HIP scaffold state (2026-09-25)¶
No rebase impact: this is a documentation-only state-ledger correction. The former Open item T-HIP-SCAFFOLD-TESTS-FAIL-2026-09-16 and the Recently closed item T-HIP-SCAFFOLD-ENOSYS-MASKED-2026-09-19 described the same four remaining default-scaffold failures after PR #1425. PR #1506's verified integration commit 11a47f39b1 carries the signed a5a9ec69e fix and is an ancestor of the current collector. The Open duplicate is removed and its provenance is retained in the closed row.
- Research digest: no digest needed; this is a trivial reconciliation against existing ADR-1264, exact Git history, current source, and executable tests.
- Decision matrix: no alternatives; retaining contradictory Open and closed records is invalid, while deleting the provenance would lose the first-stage PR #1425 history, so the older identifier is folded into the authoritative closed record.
- AGENTS.md invariant note: no new rebase-sensitive invariant. The existing ADR-1264 HIP scaffold and test invariants are unchanged.
- Reproducer / smoke:
meson test -C build-hip-scaffold --print-errorlogs --num-processes 1 test_hip_float_vif_parity test_hip_psnr_hvs_parity test_hip_psnr_hvs_parity_large test_hip_speed_singular_parity. - Changelog:
changelog.d/changed/hip-scaffold-state-ledger-reconciliation.md. - ADR: no ADR needed; no architecture, policy, runtime behavior, or scope decision changed.
agent/fix-sycl-tidy-path-99d79a — resolve SYCL clang-tidy wrapper to absolute path for safe_subprocess (2026-09-25)¶
When invoking make tidy-ratchet LANE=sycl, the lane passed scripts/ci/clang-tidy-sycl.sh as a relative executable path. Under ADR-1270, scripts/lib/safe_subprocess.py validates that all allowlisted executables are either bare commands (single path component resolved via PATH) or absolute paths (candidate.is_absolute()). Relative executable paths containing slashes are rejected with CommandValidationError: allowlisted executable must be bare or absolute.
The fix addresses this at both the caller and runner levels: 1. Makefile: TIDY_RATCHET_EXTRA_sycl defines --clang-tidy $(CURDIR)/scripts/ci/clang-tidy-sycl.sh, anchoring the wrapper path to the active worktree root so that Make invocations from arbitrary directories (make -C <dir>) produce an absolute path. 2. scripts/ci/tidy-ratchet.py: Added resolve_clang_tidy(binary, repo_root) (HISS-04 compliant, <= 60 LOC) to convert multi-component relative binary paths to absolute paths relative to Path.cwd() (or repository root), while leaving bare binary names unchanged for standard PATH lookup. Both clang_tidy_version() and run_one() use resolve_clang_tidy(). 3. scripts/ci/tests/test_tidy_ratchet.py: Added regression tests (test_relative_wrapper_path_in_subdirectory_survives_safe_subprocess, ResolveClangTidy, SyclLaneFlags) verifying safe execution across subdirectories.
Do not resolve rebase conflicts by stripping $(CURDIR) or removing resolve_clang_tidy(), as this will re-introduce the safe_subprocess validation error during make tidy-ratchet LANE=sycl.
fix/source-adr-citation-provenance — preserve exact decision identities (ADR-1311) (2026-09-25)¶
Plain ADR-NNNN references in implementation and build/control files are bound by scripts/ci/source-adr-citations.json to the exact ADR filename and exact source path/counts audited in Research-1311. Preserve the registry, the always-run pre-commit hook, and scripts/ci/tests/test_check_source_adr_citations.py together. After resolving a conflict that adds, removes, or renumbers a source citation, audit the context and run python3 scripts/ci/check-source-adr-citations.py --write; review the registry diff before accepting it. Never resolve drift by repointing a missing number at a plausible current ADR.
ADR-0557/0558 (abandoned split SpEED plans), ADR-0722 (superseded logging attempt), and ADR-0864 (unfiled Markdown-lint cleanup) are reserved historical identities. Do not allocate those numbers or collapse them into their related live ADRs. Synthetic ADR-0099/9997/9998/9999 uses remain exact fixture-only exceptions. mkdocs.yml, Markdown, changelog prose, patches, and model/binary data are deliberately out of scope; broadening the scanner to those prose surfaces recreates false positives.
The unit fixture's Git and checker subprocesses must continue to discard all inherited GIT_* variables, disable caller system/global configuration, hooks, and signing. A commit hook supplies an alternate index; letting a disposable fixture inherit it can replace the real staged index with fixture paths.
fix/gpu-picture-pool-alloc-error — handle pool allocation failure with -ENOMEM (2026-09-24)¶
When malloc(sizeof(*p)) fails in vmaf_gpu_picture_pool_init (core/src/gpu_picture_pool.cpp), the function previously returned err (initialized to 0) while setting *pool = nullptr. This incorrectly signaled success to callers with a null pool handle. The path now returns -ENOMEM directly while preserving *pool = nullptr.
Preserve the standard malloc/free calls in gpu_picture_pool.cpp and the force-included allocator interposer in core/test/test_gpu_picture_pool_alloc_interpose.h. Do not resolve rebase conflicts by switching malloc back to std::malloc, as this breaks preprocessor interposition in the deterministic regression test target test_gpu_picture_pool_alloc_failure without brittle production-only hooks. Re-run test_gpu_picture_pool_alloc_failure (suite: fast) after resolving any rebase conflicts touching core/src/gpu_picture_pool.cpp.
fix/codeql-float-equality-alerts — semantic floating-point comparisons for CodeQL (ADR-1308) (2026-09-24)¶
Resolves all six live GitHub CodeQL cpp/equality-on-floats alerts on current origin/master (alerts 168, 927, 1101, 1201, 1221, 1244) at their semantic root causes without scanner suppression or tolerance loosening.
core/src/feature/feature_name.cpp(option_double_equals): Option deduplication requires consistent matching of float/double values. Do not resolve conflicts by reverting toval == opt_val. NaN is never equal (even to NaN), preserving IEEE-754 semantics; signed zeros+0.0 == -0.0compare equal; same infinities compare equal; other finite values compare via 64-bit representation bit identity.core/src/predict.c(float_values_equal):vmaf_predict_score_at_indextests whetherguided_scorediffers from the sentinel value. Do not revert tost->guided_score != st->sentinel. The helper properly recognizes that NaN is never equal (sentinel check exits early for NaN), signed zeros+0.0 == -0.0compare equal, same infinities compare equal, and finite values compare via exact 64-bit bit identity.core/test/test_svm_api.c(svm_labels_equal): SVM labels are discrete integers stored in double.svm_labels_equaluses 64-bit bit identity with signed-zero equivalence (+0.0 == -0.0) and same-infinity behavior, rejecting NaN (never equal). It does not rely ona - b == 0.0or finiteness checks.core/src/feature/brisque_math.h(span != 0.0 && isfinite(span)): Inbrisque_range_scale, asserts that the normalization interval[lo, hi]is non-degenerate and finite (span != 0.0 && isfinite(span)). Do not revert toassert(hi != lo). In addition,brisque_fit_aggdwas split intobrisque_aggd_accumulateto satisfy HISS-04 (maximum 60 LOC per function in touched files); keep the helper static and intact.core/src/mcp/3rdparty/cJSON/cJSON.c(d - (double)item->valueint == 0.0): Preserves upstream cJSON behavior and the host compiler's floating-point model (e.g., DAZ under Intel icx-fp-model=fastvs subnormals on GCC/Clang) while using a difference-from-zero expression CodeQL accepts.core/test/test_cambi.c(float_bits_equal): Incheck_c_values_avx2_parity, comparesc_scalar[i]andc_avx2[i]usingfloat_bits_equalto assert bit-for-bit identical IEEE single-precision results between AVX2 SIMD and scalar kernel paths. Do not revert toc_scalar[i] == c_avx2[i].
fix/bug048-test-restorations — preserve picture, CLI, and registration seams (2026-09-24)¶
core/test/test_picture.c directly pins the public YUV400P allocation contract at both 1920x1080 and odd 577x323 dimensions: the luma plane is allocated with the requested geometry, while both chroma pointers are NULL and both chroma geometries are zero. Consumer tests that happen to allocate a monochrome picture are not a substitute for this seam.
core/test/test_cli_parse.c keeps direct cases for --precision=max, --precision=legacy, --precision=6, and --sycl_device 3. Preserve those explicit-option tests separately from the vmafx default and --netflix-compat override cases.
Metal registration is already covered more deeply than the historical smoke functions: test_metal_kernel_coverage_audit.c checks all 17 registered kernels, while test_metal_kernel_registration.c pins the relevant lookup and temporal-flag contracts. Do not reintroduce duplicate per-extractor functions into test_metal_smoke.c. New Metal and HIP registrations must instead follow the Registration coverage invariant sections in their backend AGENTS.md files.
No FFmpeg patch update is required: production code, public declarations, Meson options, and Netflix golden assertions are unchanged.
fix/bug048-dev-mcp-resilience — preserve runtime-record and retry controls (2026-09-24)¶
The dev-MCP files are fork-local, but core/src/libvmaf.c follows upstream Netflix and is conflict-prone. Preserve these three controls through any reconciliation:
dev-mcp-entrypoint.shmatches complete runtime records. SYCL accepts a leading bracketedlevel_zero:gpuoropencl:gpurecord; HIP accepts a fullName: gfx...orDevice Type: GPUline. A loose token search reintroduces both missed devices and diagnostic false positives. Runscripts/ci/tests/test-dev-mcp-entrypoint-probe.shafter conflicts.dev-mcp-healthcheck.shis still a stdio-compatible CLI check. It adds a driver query only when/dev/nvidia0exists; do not replace it with a Unix socket check or an unconditional NVIDIA dependency. Keep Compose's 45-second start period and rundev/scripts/test-dev-mcp-healthcheck.sh.output_file_open()retries_open()/open()exactly once only when the first failure isEINTR. The Linux control wrapsopen64, selected by this project's large-file flags, in a static non-LTO test target. If those flags change, inspect the library's undefined symbol before changing the wrapper; the test must continue proving that fault injection actually fired and that exactly two calls occurred.
No public C surface, FFmpeg patch, output schema, or Netflix golden assertion changes. Research and alternatives: Research-2084.
fix/bug048-smoke-probe-contract — current CLI and Go MCP contracts (2026-09-24)¶
No upstream impact: dev/ and the Go MCP service are fork-local. Preserve the probe's evidence contract when resolving a conflict: each backend is selected with --backend, the raw fixture declares --pixel_format 420 --bitdepth 8, the score comes from the JSON output file, and backend_used must match the request. The production MCP binary is vmafx-mcp; a client must initialize the stdio session before calling list_extractors or vmaf_score, and must not close stdin until the response arrives because EOF disconnects the Go SDK session. The legacy probe JSON keys list_features and compute_vmaf remain stable consumer keys, not MCP operation names. Run bash dev/scripts/test-smoke-probe-loop.sh after any resolution touching the probe, CLI options, or Go MCP tool surface.
fix/bug048-feature-correlation — filter non-numeric columns in feature correlation (2026-09-24)¶
No upstream impact: ai/ is fork-only (ai/scripts/feature_correlation.py). Restores BUG-048 item A11 (originally commit 5fc73913b, clobbered in 384d97d03). Parquets containing string/metadata columns (e.g. codec, chug_orientation) are filtered through select_dtypes(include='number') before to_numpy(dtype=np.float64), then all-null and constant numeric features are removed before analysis. The constant check is repeated after complete-case filtering because removing rows for a sibling feature can erase the variance of a previously valid column. All skipped sets are logged and recorded in the JSON report; this prevents NumPy's undefined-correlation warning and prevents a constant feature entering consensus_topk through a zero-score tie. An unavailable optional scikit-learn method emits an empty result map instead of a non-standard JSON NaN value. Selected feature and target rows must also be finite: NaN and both infinities are removed before Pearson or optional scikit-learn analysis. The CLI rejects non-finite --redundancy-threshold values, and the complete report is validated with allow_nan=False before the existing atomic manifest write. Preserve this fail-closed boundary when rebasing shared CLI or provenance helpers. Companion regressions in ai/tests/test_feature_correlation.py cover string/metadata columns, unavailable and constant numeric columns including a retained-row constant, non-finite feature/target rows and thresholds, missing scikit-learn, the one-feature case, and an empty usable schema.
fix/bug048-bootstrap-name-owner — keep ADR-0480's shared suffix owner (2026-09-24)¶
core/src/bootstrap_names.h is the sole owner of bootstrap collection-score suffixes and buffer sizing. Both core/src/libvmaf.c and core/src/predict.c must include it and use all four suffix symbols; do not resolve a layout-sync conflict by restoring local literals. The loops intentionally remain separate because one pools collector values and the other appends per-index scores. Run python3 core/test/test_bootstrap_name_contract.py after either consumer changes. No public API or numerical behavior changes (ADR-0480, Research-0480).
fix/bug048-report-output-restoration — restore the shared complete bundle (2026-09-24)¶
No upstream rebase impact: tools/vmaf-tune/ is fork-only. Preserve the contract that compare --format both and report --format both each emit .json, .html, and .md, in that order. The report regression was hidden because its test copied the intended dispatch logic instead of calling the production writer; keep the regression bound to vmaftune.report.write_report_outputs. Both subcommands must keep calling that single helper; restoring CLI-local copies would reintroduce the B10 drift fixed by historical commit 3a63383af. Research-0714 remains the design authority. No ADR is needed for these one-way restorations of already documented behavior.
fix/bug048-doxygen-contracts — preserve GPU public-header semantics (2026-09-24)¶
core/include/libvmaf/libvmaf_cuda.h must continue to document the live by-value import model: init returns a caller-owned allocation, import copies it without transferring ownership, vmaf_close() precedes the single-pointer vmaf_cuda_state_free(), and that free cannot NULL the caller's handle. Do not restore PR #712's old "borrows the state pointer" sentence; core/src/libvmaf.c copies *cu_state into the context.
VmafSyclPicturePreallocationMethod keeps explicit, append-only values NONE=0, DEVICE=1, and HOST=2. Their implementation mapping remains no-pool/vmaf_picture_alloc, sycl::malloc_device, and sycl::malloc_host, respectively. Run python3 -B core/test/test_gpu_public_header_docs.py after resolving a conflict in either header.
No FFmpeg patch update is required: no declaration, signature, enumerator value, or integration behavior changed. The public header, human API guide, test, research, changelog, state evidence, and this rebase note are fork-local documentation/test changes; Netflix golden assertions are untouched.
fix/bug048-helm-vulkan-docs — keep removed Vulkan out of live chart guidance (2026-09-24)¶
The Helm chart maps NVIDIA, AMD, and Intel device-plugin resources only to the active CUDA, HIP, and SYCL backends. ADR-0726 removed Vulkan; an older chart change reintroduced prose claiming it remained available implicitly through every vendor allocation. Preserve the removal notice in templates/NOTES.txt, templates/_helpers.tpl, and the two Kubernetes operator guides when resolving conflicts with chart history. No runtime template expression or public API is changed.
fix/bug048-vmaf-log — diagnostics are one record with one severity prefix (2026-09-24)¶
- Preserve the trailing
\non the guarded diagnostics infeature/luminance_tools.cpp,feature/speed.c, andfeature/vif.c. - Preserve CUDA initialization message bodies without a leading
Error:;vmaf_log(VMAF_LOG_LEVEL_ERROR, ...)already renders the severity. - Run
python3 core/test/test_vmaf_log_callsite_format.pyafter an upstream sync that touches those files. The guard covers every call site restored from9d57a93bf, plus the second CUDA-init failure path introduced later.
fix/bug048-sycl-usm-init-unwind — failed SYCL init owns its unwind (2026-09-24)¶
The framework does not call close after an extractor init returns an error. Preserve the local close_fex_sycl(fex) (or close_fex_issim_sycl) call on every post-allocation, post-dictionary, and post-graph-registration failure in the twelve touched feature TUs. Do not resolve a conflict by restoring the historical bare returns from 5d070b0b4, or by moving cleanup into the generic framework: four sibling TUs already self-unwind and would then be double-closed.
The executable contract is core/test/test_sycl_init_unwind.cpp. It is a device-free GNU-ld interposer and is intentionally registered only on Linux, static-library builds with b_lto=false; LTO resolves the calls before --wrap can see them. Re-run test_sycl_init_unwind after any sync touching these init/close pairs. Research and the exact historical boundary are in docs/research/2101-bug048-sycl-init-unwind-restoration-2026-09-24.md.
fix/bug048-issue-reference-provenance — archived tracker identities (2026-09-24)¶
The active tracker reused VMAFx/vmafx#239, VMAFx/vmafx#241, VMAFx/vmafx#310, VMAFx/vmafx#857, VMAFx/vmafx#866, and VMAFx/vmafx#870 for pull requests unrelated to historical records carried by this tree. Preserve the explicit lusoris/vmaf spelling in every protected Vulkan async-fence, CUDA CAMBI, ADR-sweep, BVI-DVC, profile, sync-report, source-invariant, changelog-fragment, and generated-changelog context. A conflict resolution that restores a bare or active-repository form silently points readers at a newer object.
scripts/ci/check-issue-reference-provenance.py is intentionally narrower than a generic Markdown link checker: it matches proven historical contexts by stable prose anchor and requires their archived repository namespace. Keep the checker, scripts/ci/tests/test_issue_reference_provenance.py, the always-run pre-commit hook, and the Rule Enforcement self-test together. Do not widen it to reject normal bare references to the active fork or Netflix/vmaf upstream references. See Research-2089.
fix/bug048-zed-restoration — current Zed project contract (2026-09-24)¶
No upstream Netflix/vmaf code is touched. Preserve the selective Zed 1.18.1 restoration if .zed/ or developer documentation conflicts: project settings exclude agent, agent_servers, provider/model pins, and permission policy; the MCP entrypoint is docker exec -i vmaf-dev-mcp vmafx-mcp; the three Standards: tasks remain mandatory. Run python3 -m pytest -q scripts/ci/tests/test_zed_project_config.py after any resolution. Do not copy the archived 1.3.6 configuration back into live files.
fix/codeql-misc-c-alerts-rc1 — CodeQL misc C/C++ alert fixes (2026-09-24)¶
core/src/feature/iqa/convolve.cwiden-after-multiply bit-exact invariant (ADR-0138). Multiplication is kept in single-precision float (const float prod = img[...] * k->kernel_h[...];) and explicitly cast to(double)before accumulation (sum += (double)prod;).- Do NOT pre-widen operands (e.g.
sum += (double)a * b;): this evaluates the multiply indoubleand violates the ADR-0138 bit-exactness contract, desynchronizing the scalar reference from the AVX2 (_mm256_mul_ps), AVX-512, and NEON (vmul_f32) SIMD twins (verified viatest_iqa_convolve). - Do NOT revert to direct
sum += a * b;: this creates compiler-generated widening conversions flagged by CodeQL'sIntMultToLong.ql(Alert 1005). See Research-2031. core/src/feature/moment.csecond-moment reduction contract (ADR-0179 / ADR-0987). Squaring is evaluated in single-precision float (const float term = pic_ * pic_;) and explicitly cast to(double)before accumulation (cum += (double)term;).- Unlike
convolve.c,moment.cis governed by ADR-0179 and ADR-0987 under a tolerance-bounded non-byte-exact reduction contract (MOMENT_REL_TOL = 1e-7) rather than bit-exactness. - Do NOT pre-widen operands to double or revert to implicit widening (
cum += pic_ * pic_;), which restores the source pattern behind historically dismissed CodeQL Alert 707 (cpp/integer-multiplication-cast-to-long). core/src/pdjson.henum json_typesequential numbering. All enumerators (JSON_NONE = 0throughJSON_NULL = 11) are explicitly assigned. This satisfies CodeQLcpp/irregular-enum-init(AV Rule 145 / Alert 1064) while strictly preserving the public ABI. Verified bycore/test/test_pdjson.c::test_enum_json_type_abi_contract.core/tools/cli_parse.cppusage()discrete overloads. Discrete template overloads for 1, 2, and 3 arguments prevent empty parameter pack instantiations forRest &&...rest, closing CodeQLcpp/unused-local-variableandcpp/unused-static-variable(Alerts 1002/1003). Covered by adversarial exit tests incore/test/test_cli_parse_long_only_args.c..github/workflows/security-scans.ymlMeson configure outside extraction. Meson configure runs beforegithub/codeql-action/initwith the build directory in${{ runner.temp }}/build. Running configure before extraction prevents Meson compiler probe files (testfile.c) from polluting the CodeQL database (current hosted probe alert 1279, historical 1278 / pre-merge 1232–1235); keeping the build tree outside$GITHUB_WORKSPACEprevents generated artifacts from being indexed as project source.
docs/release-sequence-rcs — RC responsibilities stay separated (2026-09-26)¶
No upstream source impact: this change is fork-only release governance and documentation. Preserve ADR-1341's phase boundary when rebasing release, roadmap, backlog, or model-training documents: RC1 owns correctness completion plus the reproducible outside-hardware report path; RC2 owns benchmark/profiling/tuning; RC3 owns the one-shot real retrain. Do not resolve a conflict by restoring generic “post-RC” training or by moving performance work back into RC1.
Ordinary Renovate/version PRs remain mergeable under existing required gates. Their merge invalidates affected exact-head candidate evidence and triggers revalidation; it does not restore a blanket version freeze. The Netflix golden assertions remain untouched.
feat/rc1-tester-report-bundle — RC1 explicit-backend evidence bundle (2026-09-26)¶
No upstream impact: tools/rc1-tester/ and the linked usage documentation are fork-only. Preserve the fail-closed correctness invariant: every run includes a CPU reference and PASS requires four consecutive frames (through temporal frame 3), finite model metrics, bounded/consistent VMAF, and backend_used exactly matching each explicit request. CPU must match the pinned snapshot and each accelerator must match that run's CPU result within 5e-5 per metric and frame; never restore version-only, exit-zero-only, finite-only, or auto success. Keep the pinned model/snapshot/fixtures hashed, preserve the bounded process-group/ output contract, and keep status 100 as backend-unavailable evidence. Root backend_used is not per-feature dispatch proof. RC1 collects build/correctness reports, RC2 owns benchmark/tuning work, and RC3 owns real training. The collector must not invoke either later phase. No Netflix CPU golden assertion is modified. See ADR-1342.
fix/mcp-cyclic-imports — Python transports form an import DAG (2026-09-23)¶
No upstream impact: mcp-server/ is fork-only. Preserve vmaf_mcp/http_scoring.py as the seam between the canonical scoring implementation in server.py and the HTTP adapter in http_transport.py. The dependency remains one-way: server.py may import the HTTP startup entry, but http_transport.py depends only on the shared interface and must never import server.py. Canonical registration must not replace an adapter installed by an embedding process; direct run_http_server callers may inject one, which must be bound to that application without mutating the process-wide registry (ADR-1304), and must fail before binding when neither an injected nor canonical adapter exists. Re-run tests/test_import_graph.py after resolving any conflict that touches these three modules; it walks function-local imports as well as module-level imports, matching the dependency edges CodeQL reports.
fix/bug048-strict-json — strict tool report boundaries (2026-09-24)¶
Preserve both halves of BUG048/A9 when resolving changes around the tool entry points. tools/external-bench/compare.py --out-json keeps missing values as nan for aggregation and the text table, but render_json() maps every non-finite aggregate float to JSON null and calls json.dumps with allow_nan=False. vmaf-roi-score keeps the finite check in blend_scores(), maps its ValueError to exit 65 without writing a report, and retains allow_nan=False in _emit().
Commit 384d97d03 once replaced both files wholesale and removed these boundaries while their docs and Research-0722 survived. Run the two package suites, especially the all-wrapper-failure and three non-finite ROI cases, after any conflict involving these paths. No public libvmaf or FFmpeg patch surface is involved.
fix/msvc-strict-fp-flags — compiler-native no-contraction flags (2026-09-23)¶
vmaf_fp_model_args and vmaf_strict_fp_args in core/src/meson.build are a single policy consumed by x86 and AArch64 SIMD carve-outs, strict scalar-reference libraries, and core/test/meson.build. vmaf_cuda_host_strict_fp_args carries the host-side spelling through nvcc. Do not resolve a conflict by restoring per-target ['-ffp-contract=off'] literals or by rebuilding a separate test mapping: the old duplication sent ignored Unix flags to Windows drivers and could place a SIMD kernel and its scalar reference under different arithmetic semantics.
Preserve the compiler distinctions and order. Unix intel-llvm requires -fp-model=precise first and -ffp-contract=off last; Windows intel-llvm-cl requires /fp:precise /Qfma-; MSVC uses /fp:precise; and clang-cl needs /clang:-ffp-contract=off. Windows nvcc must forward /fp:precise to cl.exe instead of -ffp-contract=off. The executable contract is core/test/test_strict_fp_compiler_args.py; run it after any rebase touching these Meson blocks.
fix/codeql-unused-static-alerts — preserve test-local source identities (2026-09-24)¶
-
Do not restore redundant implementation sources.
test_picture*does not compilethread_pool.c;test_predictandtest_model*do not compilepdjson.c. Those targets either do not use the implementation or already link its owning library/object. -
Keep intentional direct copies uniquely identifiable. Each pdjson copy is built as a target-local static library whose private helpers have unique names; its public
json_*API remains unchanged. Each pdjson executable also keeps a target-localrun_testsalias.test_thread_pool_backpressureretains aliases for the fourvmaf_thread_pool_*entry points around the deliberatethread_pool.cinclusion. This prevents CodeQL from coalescing a repeated helper while orphaning one link-target-specific call graph, without manufacturing unused public APIs for cppcheck. The aliases are test-only and must not enter installed headers or the production ABI. -
Keep configuration-only helpers configuration-owned.
vector_unchangedis defined and used only underFEX_VECTOR_ALLOC_TEST. Do not restore[[maybe_unused]]or add unrelated default-build assertions to make it appear reachable. Preserve the non-LTO allocation-failure tests that exercise its actual contract. -
Consolidate orphan picture test identities.
test_picture,test_picture_v2, andtest_picture_pool_error_pathsshare thetest_picture_impltest-local static library forpicture.c,mem.cpp, andref.cpp. The error-path test still compilespicture_pool.cdirectly, so its ADR-0960 access to internal pool entry points is preserved. -
Keep predictor tests out of the implementation translation unit.
test_predict.cincludespredict_internal.hfor the exact-order pure mapping/equality helpers and links the production predictor. Do not restore#include "predict.c"or theVMAF_PREDICT_TEST_NONFINITE_LOGmacro: that unity include creates a second static graph that CodeQL reports as unreachable and replaces the shipped warning with a private test callback. Preserve bothtest_predict_source_authority.pyandtest_predict_nonfinite_log_output.pywhen resolving test-build conflicts. -
Keep one compiled predictor source authority.
predict_c_libincore/src/meson.buildis the only target that compilespredict.c;libvmafextracts its object and private-source tests link it throughpredict_test_dependencies. Do not restorepredict.ctolibvmaf_sources, any test source list, or a shared test-source array. The old graph compiled the TU 53 times and left an orphan static scan identity in whole-build CodeQL extraction.test_predict_source_authority.pylocks both Meson boundaries. -
Keep the collector-only test on the predictor source authority.
test_feature_collectormust retainvmaf_cflags_commonandpredict_test_dependencies. Do not restore../src/predict.c: the linkedpredict_c_libarchive satisfies the unity-includedlibvmaf.creferences without creating another predictor graph.
Research-2096 contains the exact-query evidence. Re-run the complete CodeQL database extraction after rebases that alter these target source lists or dependencies; ordinary runtime tests cannot detect a repeated-compilation identity regression. The historical reviewed receipt is .workingdir/evidence/codeql-unused-static-2026-09-24/unused-static-final.csv; because that ignored evidence path is not present in a clean checkout, it is not an acceptance input. From a clean repository root, the following commands create a fresh CodeQL 2.27.0 database, run codeql/cpp-queries 1.8.3, and validate results fail-closed against the official 9-column CodeQL CSV schema using scripts/ci/check-codeql-unused-static.py. The validator confirms 0 selected rows remain in the lane target paths, including core/src/picture.c. The pre-correction inventory has six repository-wide rows (two in picture.c, three in predict.c, and one in cambi.c); the corrected replay has four (the predict.c and cambi.c rows). CODEQL_BIN may override the documented local installation path. The temporary directory is retained and printed for inspection.
That four-row remainder is historical. A 2026-09-25 database from exact base 71c3c155717f4496c3c572e479d5202984e9c3b5 found five predictor rows created by the unity include and repeated private-source predictor builds, with no CAMBI row. Research-2096 records the entity/call-edge probe, single-authority repair, and follow-up closure evidence. The final clean 1,559-step extraction contains one predictor compile command, and the official 1.8.3 query returns zero repository-wide rows.
set -euo pipefail
codeql_bin="${CODEQL_BIN:-$HOME/.cache/codeql-2.27.0/codeql}"
codeql_run="$(mktemp -d /tmp/vmafx-codeql-unused-static-replay.XXXXXX)"
codeql_spec='codeql/cpp-queries@1.8.3:Best Practices/Unused Entities/UnusedStaticFunctions.ql'
cat >"$codeql_run/build.sh" <<'BUILD_SH'
#!/bin/sh
set -eu
build_dir="${1:?missing CodeQL build directory}"
meson setup "$build_dir" core -Denable_cuda=false -Denable_sycl=false
meson compile -C "$build_dir"
BUILD_SH
"$codeql_bin" version | rg -q 'release 2\.27\.0\.'
"$codeql_bin" database create "$codeql_run/db" \
--language=cpp --source-root=. --threads=4 \
--command="/bin/sh $codeql_run/build.sh $codeql_run/build"
"$codeql_bin" database analyze --rerun --format=csv \
--output="$codeql_run/unused-static.csv" \
"$codeql_run/db" "$codeql_spec"
python3 scripts/ci/check-codeql-unused-static.py "$codeql_run/unused-static.csv"
printf 'CodeQL replay artifacts: %s\n' "$codeql_run"
fix/codeql-include-non-header-alerts — internal header and link seams for test suites (2026-09-24)¶
- Test targets link against libvmaf instead of unity-including source files. Alerts 908, 943, 955, 1043, 1203, 1218, and 1241 resolved
.c/.cppinclusions by establishing proper internal header declarations: core/src/feature/luminance_tools.h: declares narrowvmaf_luminance_test_*trampolines while the implementation helpers remain translation-unit-local.core/src/model.h: declares narrow built-in-model count and iterator-version test accessors whileVmafBuiltInModelandBUILT_IN_MODEL_CNTremain private tomodel.c.core/src/libvmaf_priv.h: declaresvmaf_context_is_flushed,vmaf_context_has_thread_pool, and test flush triggers.core/src/feature/cambi_internal.h: preserves GPU-facingvmaf_cambi_*routines and declaresvmaf_cambi_test_*internal test trampolines for static stage and helper coverage. Do not re-introduce#include "cambi.c",#include "libvmaf.c", or other.cinclusions during rebase conflict resolution.- Meson test target dependencies:
test_featurebuildsfeature_name.cppdirectly;test_model,test_flush_context_ordering,test_cambi, andtest_cambi_stage_simdlink withlibvmaf/libvmaf.get_static_lib(). Preserve these linker configurations incore/test/meson.build.
fix/sycl-tidy-required-rc1 — SYCL tidy strict-reporting and header coverage (2026-09-24)¶
Commit 6475fa9ea had already promoted clang-tidy-sycl (Tidy SYCL) to a required, non-advisory ADR-1297 gate. This branch hardens that existing policy; it does not perform the promotion.
Tidy SYCLbelongs to bothrequiredandstrictMustReport. The job has no workflow or job-level path filter: its detect step skips only the expensive body, while the context itself reports on every eligible PR and master push. Absence is therefore a workflow failure, never a path skip.- Changed-file detection in
lint-and-format.ymlcovers SYCL headers. The file patterns for the job include'core/src/sycl/*.h'and'core/src/feature/sycl/*.h'in addition to.cppand.hppin each pull-request, push-fallback, normal-push, and dispatch command. Do not let one complete branch mask missing coverage in another. - The Go and SYCL contract suites share one real aggregator driver. Keep execution in
scripts/ci/required_aggregator_harness.py; duplicated Node drivers can drift in polling time and result decoding. test_sycl_tidy_workflow_contract.pyis wired to CI and local hooks. The contract is executed bydeep-dive-checklistinrule-enforcement.ymland by thetest-sycl-tidy-workflow-contractlocal hook in.pre-commit-config.yaml. The exact ADR-1297 strict set is also pinned inscripts/ci/tests/test_hiss_replay_contract.py.
fix/codeql-vif-large-parameter-alerts — VIF AVX-512 internal helper pointer convention (2026-09-24)¶
core/src/feature/x86/vif_avx512.cinternal stage helpers takeconst *.vif_horizontal_energy_pack512,vif_vertical_mean8,vif_vertical_energy8,vif_vertical_store8,vif_vertical_store_mean8,vif_vertical_energy16,vif_vertical_store_mean16, andvif_vertical_store_energy16pass aggregate vector structs (VifPair512,VifTaps8,VifEnergy512) byconst *rather than by value. Under System V AMD64 and Windows x64 ABIs, objects > 64 bytes cannot be passed in vector registers; passing them by value forces stack copies and triggers CodeQLcpp/large-parameteralerts 1108–1112. Under GCC 16.2.1-O3, the complete hot.textsection is byte-for-byte identical to an independently builtorigin/masterobject. Do not revert these internal parameters to pass-by-value on rebase.- Public ABI in
core/src/feature/x86/vif_avx512.his unchanged. None of the modified helper functions are declared in headers or exported from the static library. - Parity test harness in
core/test/test_integer_vif_avx512_stages.cincludes a red check. Itstest_integer_vif_avx512_stages_red_checkcase verifies baseline bit-exactness and proves the harness detects 1-bit input and intermediate plane perturbations. Preserve both cases on rebase. - Bounded macros and narrow function-size suppression preserve the codegen invariant. In
core/src/feature/x86/vif_avx512.c,vif_subsample_rd_8_vert_jandvif_subsample_rd_8_horiz_jdecompose repeated unrolled vector operations into bounded macros (VIF_VERT_LOAD10_REF,VIF_VERT_LOAD10_DIS,VIF_VERT_MADD5,VIF_HORIZ_TAP8) to satisfy HISS-04 / NASA Rule 4 function length constraints ($\le 60$ source LOC) without raising baseline debt. Replacing the macros with static forced-inline helper functions was proven to alter GCC SSA register allocation due to address-taken vector pointer arguments (e.g. swapping%zmm3and%zmm13), breaking the byte-identical.textmachine code contract (SHA-256:80b48e27e202ca98c3a124351740fdecba1e19df53adeebbcd237889fe44374c). BecauseVIF_HORIZ_TAP8unrolls 9 taps withinvif_subsample_rd_8_horiz_j, direct clang-tidy 22.1.8'sreadability-function-sizecounts macro-expanded statements (128 statements vs. threshold 120) despite source LOC being 36 ($\le 60$). A narrow inline suppressionNOLINTNEXTLINE(readability-function-size)citing ADR-0138, ADR-0139, ADR-0141, and Research-2098 is required to preserve this bit-exact codegen invariant without relaxing global tidy configuration or adding baseline debt. Macro locals inVIF_VERT_MADD5use compliant non-reserved identifiers (t0lo–t4hi). Do not extract helper functions or remove the suppression on rebase.
fix/silent-revert-restorations-1 — restore ADR-0982 GPU partial-init leak unwinds (BUG-048 Sec A3) (2026-09-23)¶
- ADR-0982 partial-init unwinds and error handling restored in CUDA and SYCL runtimes. PR #503 (
fbde2f91e) originally implemented ADR-0982 to fix resource leaks on error paths in GPU runtime initialization and teardown, but squash merge PR #504 (a12373faa) silently reverted these changes. This branch restores the missing unwinds: core/src/cuda/picture_cuda.c:vmaf_cuda_picture_alloc()zeroesprivstruct immediately upon allocation (memset(priv, 0, sizeof(*priv))) to ensure NULL sentinels for unwinds. On plane allocation failure (device_pic_alloc_planes() != 0),device_pic_unwind()is called withDEV_PIC_UNWIND_DATAso that prior successfully allocated device planes are freed (cuMemFree(pic->data[i])) and their pointers set to NULL, preventing plane memory leaks.core/src/cuda/common.c: invmaf_cuda_release(), failure paths forcuStreamDestroy,cuCtxPopCurrent, andcuDevicePrimaryCtxReleaseroute viafail_release_funcsso thatcuda_free_functions(&f)is executed andcu_statezeroed, rather than leaking the dlopen'd driver table on release errors.core/src/cuda/drain_batch.c: indrain_stream_ensure(), ifcuCtxPopCurrentfails after stream creation,fail_after_streamis taken to destroyg_drain_batch.drain_strwithcuStreamDestroybefore returning-ENOTRECOVERABLE.core/src/sycl/common.cpp: invmaf_sycl_graph_register(),state->combined_queueis allocated before incrementingstate->num_graph_extractorsso a queue allocation failure does not leave a half-registered extractor entry.- Guarded by deterministic mock-driver regression test.
core/test/test_cuda_runtime_unwind.ccompiles against a simulated in-memoryCudaFunctionstable and verifies that plane allocation failures unwind all previously allocated planes without hardware GPU requirements, running in the default fast test suite. - Rebase resolution: Upstream Netflix does not carry these unwinds. When syncing with upstream or resolving conflicts in
picture_cuda.corcommon.c, keep fork'sDEV_PIC_UNWIND_DATAstage on plane allocation failures and thefail_release_funcstable cleanup invmaf_cuda_release().
fix/codeql-svm-lifecycle-alerts — Solver RAII cleanup and parser loop bounds (2026-09-24)¶
core/src/svm.cppSolver lifecycle uses idempotentsolve_cleanup().solve_cleanup()frees and zeroesp,y,alpha,alpha_status,active_set,G, andG_bar. Callingsolve_cleanup()insolve_finish(),~Solver(), and entry ofsolve_setup()eliminates exception leaks without causing double-frees on normal completion. Do not revert to rawdelete[]insolve_finish()or remove pointer nulling on rebase, which reintroduces heap-use-after-free/double-free aborts under ASan.parse_support_vectors()uses a bounded while loop with sentinel check. Replacingfor (size_t i = 0; ...; ++i)withwhile (i < sv_buffer.size())and assertingi < sv_buffer.size()eliminates loop variable mutation inside the body and prevents out-of-bounds reads. Function length must stay <= 60 LOC to comply with HISS-04 (currently 59 LOC).
fix/sycl-a380-snapshots-bug040 — production SYCL DMA host-buffer race fixed (2026-09-24)¶
- Both
core/src/libvmaf.cpicture-ownership paths must wait for the final SYCL upload. The picture pool (tools/vmaf.cpp, 3 slots for 1-thread default) recyclesdistto the reader thread immediately aftervmaf_picture_unref. Withcopy_queue.memcpyDMA still running on the host buffer, the reader'sfreadoverwrites bytes in-flight — corrupting the bottom half of 4Kdistplanes with next-frame data.read_pictures_frame_cleanupcovers the serial path.threaded_read_pictures_batchwaits after enqueue while the caller's original counted refs are still live, then releases them on both success and enqueue-error paths; waiting later inread_pictures_frame_cleanup_after_batchis unsafe because the worker may already have dropped its copies. Both paths wait forlast_upload_event(the final dist-plane DMA) before the caller can refill a pooled picture. In a combined CUDA+SYCL build, the serial wait must precede CUDA's host-cleanup early return. Do not drop or move either call during a rebase; the event does not touchcombined_queue, so it does not serialize GPU compute.core/test/test_sycl_cuda_serial_upload_lifetime.ccompiles only with both backends enabled and poisons released 4K host storage atn_threads=0andn_threads=1to pin both orderings. testdata/test_sycl_4k_repeat_determinism.pyis the production stability harness. It defaults to 20 serial and 20--threads 14K SYCL runs and asserts every frame, pooled, aggregate, and backend value is identical after removing only measuredfps.VMAF_BINbinds to its adjacent Mesonbuild/srclibrary;VMAF_LIB_DIRis the explicit override for another layout. It usespytest.mark.skipifwhen 4K YUV fixtures are absent, so it is always safe to collect. Do not weaken the full-report comparison without rerunning the Arc A380 evidence.- The nondeterminism was not in graph-replay or the compute queue.
VMAF_SYCL_NO_GRAPH=1showed identical drift; the corruption was in host memory before any compute submitted. Do not re-introduce event dependencies betweencombined_queueoperations to "fix" determinism — the root cause was a buffer lifetime issue, not a command-queue ordering. The full investigation and rejected alternatives are recorded in Research-2082.
fix/sycl-a380-snapshots-bug040 — CPU and A380 SYCL snapshots regenerated together (2026-09-23)¶
testdata/scores_cpu_{720,1080,4k}.jsonandtestdata/scores_sycl_a380_*.jsonmoved together on purpose. Onlyref/dis_576x324_48f.yuvandref/dis_640x480_48f.yuvare committed in the repository; the 720p, 1080p, and 4k fixtures are gitignored and derived bytestdata/generate.shfrom the authoritative Big Buck Bunny 4K MP4 source (bbb_sunflower_2160p_30fps_normal.mp4). Because FFmpeg n9.0.1 (x264) encodes the distorted clips with slightly different bitstreams than the historical April 2026 encoder, deriving fresh fixtures naturally shifted the CPU scores at 720p, 1080p, and 4k by ~6-8 points. Both CPU and SYCL snapshots were regenerated from the exact same newly-derived fixtures in the same pass. Do not "restore" the old CPU 720/1080/4k snapshots during a rebase or merge conflict; doing so breaks cross-backend parity against the regenerated SYCL snapshots.testdata/run_sycl_scores.pyrequires--backend sycland device pinning. The script previously selected the backend negatively with--no_cuda, which stopped selecting SYCL once HIP was added. It now explicitly passes--backend sycl. It also configuresONEAPI_DEVICE_SELECTOR=level_zero:gputo select an Intel Level Zero GPU. The regenerated 576p and 4K snapshots were independently reproduced with byte-identical numeric payloads after dropping the measuredfpsand build-derivedversionfields. The diagnosticVMAF_SYCL_CHECKSUMpath was disabled, so the results do not depend on its blocking checksum copy or hide a production queue race.- Committed 576x324 and 640x480 fixture files remain bit-identical.
testdata/ref_576x324_48f.yuv,dis_576x324_48f.yuv,ref_640x480_48f.yuv, anddis_640x480_48f.yuvwere not regenerated and match master bit-for-bit.
fix/drop-nvidia-cuda-base — dropping nvidia/cuda base images in favor of Ubuntu 26.04 + apt (2026-09-24)¶
build-config.envsetsCUDA_BUILDERandCUDA_RUNTIMEtoubuntu:26.04@sha256:.... Do not revert them tonvidia/cuda:...images during a rebase. Upstream OCI images for CUDA point releases (like 13.4.2) lag or are skipped entirely, whereas NVIDIA's official apt repository publishes day 1 (ADR-1306).scripts/ci/check-base-image-single-source.shrule 3 expectsubuntu:${DEV_UBUNTU}@for both CUDA keys. It additionally requiresCUDA_BUILDERandCUDA_RUNTIMEto equalDEV_BASEbyte-for-byte, including the digest, and owns the narrowdocker/dev/ubuntu-26.04-cuda.Dockerfilemirror. Reverting those rules or the ARG default lines inDockerfile,docker/Dockerfile.production-gpu,docker/Dockerfile.node, or the CUDA compatibility Dockerfile will failmake base-images-sync.scripts/ci/check-cuda-pin-lockstep.pyretired theimageshape. The coordinated pin now tracks 7 sites across 2 files (build-config.envanddocker/Dockerfile.production-gpu). The residual regex catches any re-introducednvidia/cudaimage tags. Only the apt series and runtime label are mechanically derived; do not make--writeguess NVIDIA component build numbers.scripts/ci/install-cuda-toolkit.shsupports--mode=builder,--mode=runtime, and--mode=full. Every core apt operand is an exactpackage=versionfrom the live NVIDIA metadata, and every installed version is checked withdpkg-query. Containers invoke it directly as root withoutsudo; host/CI runners invoke it usingsudo.dev/Containerfilemust keep using shared--mode=full, not a second floating apt recipe.- Renovate owns
CUDA_VERSIONthroughcustom.nvidia-cuda-redist, not Docker tags. Keep the HTML datasource, exactredistrib_X.Y.Z.jsonextraction, and scoped timestamp-optional/manual-review rule together. Reintroducing the oldnvidia/cudapackage group silently restores the publication bottleneck this branch removes.CUDA_APT_LOCK_RELEASEmust remain outside Renovate ownership so every release bump fails until the exact toolkit, nvcc, and cudart versions have been reviewed live.
fix/mypy-prepush-config-baseline-rc1 — baseline evaluates branch checker config (2026-09-24)¶
scripts/git-hooks/pre-push-mypy.pyevaluates baseline mypy under branch checker configuration. When resolving rebase conflicts or updating pre-push hooks, preserve the configuration synchronization logic inbaseline_fingerprints()and the scope widening inselected_paths(). The disposable baseline worktree checked out at the merge base must receive copies of the branch's checker configuration files (pyproject.toml,mypy.ini,.mypy.ini,setup.cfg) before executing the baseline checker. This ensures that configuration adjustments (such aspython_versionraises or strictness increases) do not attribute pre-existing debt in merge-base code as new errors introduced by the branch.- Merge-base source files must remain preserved. The baseline worktree must not copy branch
.pyfiles into the baseline worktree. Only checker configuration files are copied, while source code remains checked out atbase. - Resolve
ai/srcthrough one module identity. Preservefollow_imports = "skip"on the legacyai.src.*mypy override. The canonicalvmaf_train.*tree is checked by the dedicatedai/srcinvocation with--explicit-package-bases; traversing both names in the widened configuration-change run makes mypy abort withSource file found twicebefore any finding comparison. - Fail-closed behavior is required. If configuration parsing fails, mypy exits outside its ordinary 0/1 statuses, or exit 1 carries no parseable finding, the hook raises a
RuntimeErrorand terminates with exit code 2. A blocker remains fatal even after partial findings, and one module-identity run's exit 1 must not mask a later blocker.
agent/fix-ci-mypy-no-files-rc1 — required mypy gate is fail closed (2026-09-25)¶
- Hosted and local mypy share one gate implementation. Preserve the
Python Lintworkflow call toscripts/git-hooks/pre-push-mypy.py; do not restore a rawmypy ai/ scripts/directory scan or an advisory shell tail. The workflow must continue to installrequirements/locks/mypy.txtwith--require-hashesand fetch full history so both trees are available. - Event-specific merge-base authority is intentional. Pull requests retain the hook's
origin/masterdefault. Master pushes pass the exactgithub.event.beforecommit throughVMAFX_MYPY_BASE_REF, so the required post-merge run checks what landed instead of comparingmasterwith itself. The hook validates an explicit value withgit rev-parse --verify --end-of-options <ref>^{commit}before computing the merge base and fails closed when it cannot resolve the ref. - Preserve the existing checker semantics. Keep the baseline-configuration synchronization, finding fingerprints,
--no-site-packagesisolation, tracked-file scope, and the dedicatedai/src/invocation with--explicit-package-bases. ADR-1310 changes only CI's base selection and status propagation; it does not redefine the pre-push delta policy. - Typed test decorators are part of the checked surface. In checker-only environments, pytest's marker factory is untyped. Preserve the typed
_parametrize()adapter inai/sidecar/tests/test_online_trainer.py; replacing it with directpytest.mark.parametrizeerases the decorated test signature and fails the required gate.
fix/scorecard-pins-best-practices — hash-locked Python installs and OpenSSF hardening (2026-09-23)¶
requirements/locks/manifest.jsonis the sole compiler authority for Python locks. All lock files underrequirements/locks/,docs/,python/,ai/,mcp-server/,dev/,dev-llm/, andtools/carry input digests and generator version metadata (uv 0.12.18). Do not edit lock files by hand or resolve rebase conflicts by taking one side's hashes; runmake python-locks-writeto regenerate them from the merged inputs.--require-hashesis enforced repository-wide. Every executable pip install in.github/workflows/,Dockerfile*,scripts/setup/,Makefile, and literalsession.install(...)calls innoxfile.pymust specify--require-hashes -r <lockfile>, or--no-depsfor local editables/wheels. Requirement targets are exact manifest outputs or explicitinstall_aliases; never restore basename/suffix matching.scripts/ci/check_python_dependency_locks.py checkenforces this inmake lintand pre-commit.- PEP 517 build-isolation dependencies for pure sdist packages.
docs/requirements.txtpinssetuptools>=77.0.1andwheel>=0.45.1so they are hashed intodocs/requirements-lock.txt. Workflows invoking this lockfile pass--no-build-isolationto prevent pip from attempting unhashed PyPI downloads. dev-linters.txttargets Python 3.12 portability.dev-linters.inis compiled with--python-version 3.12to ensure universal markers include dependencies required on Python 3.12 workstations (e.g.tomli).- SLSA GitHub generator must remain tag-pinned.
slsa-framework/slsa-github-generatorrequires semantic@vX.Y.Ztags for its trusted builder verification (slsa-verifier#12; ADR-1128). Never convert its refs to commit SHAs. - Nox uses manifest-owned locks and package-compatible interpreters. Bootstrap Nox from
requirements/locks/nox.txt; each package session owns a dedicated development lock. Preserve Python 3.12 onroi_scoreandensemble_kit, whose package metadata excludes Python 3.14. - Installer tooling (
pip) excluded from bootstrap build locks.requirements/locks/build.inandbuild.txtpin build dependencies (meson,ninja) only.pipis installer tooling provided by runners/operating systems; pinningpipinsidebuild.txtcaused uninstallation failures on Debian/Ubuntu systems with packaged pip distributions lackingRECORDmetadata. - Truthful package-wide license review for
text-unidecode.actions/dependency-review-actionevaluates SPDX license expressions underdeny-licenses: GPL-3.0, AGPL-3.0.python-slugifybrings intext-unidecode, licensed asArtistic-1.0-Perl OR GPL-1.0-only OR GPL-2.0-or-later, which VMAFx consumes underArtistic-1.0-Perl. GitHub Dependency Review compares PURLs package-wide ignoring versions;allow-dependencies-licenses: pkg:pypi/text-unidecodetruthfully allows the package using exact package-wide purl syntax. - Root
Dockerfileisolates Python tooling in/opt/vmaf-venv. Prevents packaging conflicts (such as Debian's pre-installedpython3-packaging) when installing hash-locked dependencies into the container image, eliminating the need for--break-system-packages. - Fail-closed authority edges are cross-platform and parser-independent. Preserve the classifier's explicit allowlist: generic
*.inandmanifest.jsonbasenames outside the owned requirements subtree are not dependency-only. The workflow fallback accepts simple quoted block keys and must report the same pre-checkout helper ordering failure as PyYAML. Nox annotated or literalgetattr(session, "install")aliases remain scanned. Manifest output, input, and alias-consumer paths must be local and repo-relative under POSIX and Windows semantics.
fix/rust-ci-path-filters — required workflows always emit gates (2026-09-24)¶
- Do not restore workflow-level
paths:orpaths-ignore:to a workflow that hosts an aggregator-required context.build.yml,dev-container-build.yml,docker-image.yml,doxygen-public-api.yml,ffmpeg-integration.yml,helm-chart.yml, andrust-ci.ymlmust always start. Theirimpactjobs select distinctly named heavyworkjobs;if: always()gate jobs alone own the exact required context names and fail closed on planner/work disagreement (BUG-098). Keep those twelve names instrictMustReport; absence is no longer an accepted path-skip outcome. GitHub creates each gate check only after itsneedswork completes, so keep every planner/work display name in the aggregator'sdelayedStrictDependenciesmap and preserve the paginated check-run fetch. - Keep public libvmaf headers in
selectors.rust.patterns.vmafx-sysgenerates FFI bindings fromcore/include/libvmaf/libvmaf.husingbindgen, socore/include/libvmaf/**must select Rust work even though trigger-level filters no longer exist. - Keep each of these workflow files in
full_patterns. A change to routing or gate structure must select every lane, not rely on the selector being edited. Re-runtest_ci_impact.py, the Rust workflow contract,actionlint, andcheck-aggregator-names.shafter resolving conflicts in this block.
fix/codeql-python-alerts-rc1 — active exception semantics and identity comparison for CodeQL Python alerts (2026-09-24)¶
compat/python-vmaf/core/executor.pymaintains active exception diagnostics for FIFO workers._run_fifo_workercatches(BrokenPipeError, EOFError, OSError)during traceback transmission, safely bounds channel cleanup via_safe_close_channelinfinally, and uses_safe_add_exception_noteto attach pipe failure diagnostics without calling a custom exception override or allowing note-storage failure to replace the target exception, and guarantees secondary close or send failures never displace the primary target exception._fifo_worker_failuretreats EOF and OS-level read failures on the diagnostic pipe as immediate failures with synthesized child traceback context (distinguishingEOFErrorfromOSError), ensuring dead child processes triggerRuntimeErrorrather than delaying onNone.compat/python-vmaf/core/train_test_model.pyuses elementwise identity comparisonis None. InRegressorMixin._get_scatter_arrays,ys_label_stddevmasksNoneusing[x is None for x in ys_label_stddev.flat], zeroes those entries, converts the array to float, and zeroes remainingNaNs. Future upstream rebases must not restoreys_label_stddev == Noneor# noqa: E711.compat/python-vmaf/tools/misc.pyandtools/scanf.pypreserve explanatory comments and robust error handling.check_scanf_matchretains fallback from sscanfFormatError/IncompleteCaptureErrortofnmatch, andisFileLikesafely returnsFalsewhenseek()raises(IOError, OSError, ValueError).- Verification:
PYTHONPATH=python:compat pytest python/test/executor_test.py python/test/train_test_model_test.py python/test/python_harness_scanf_locale_bugs_test.pyand CodeQL query evaluation againstEmptyExcept.qlandEqualsNone.ql.
fix/semgrep-python-warning-alerts — Semgrep SHA-1 upgrade and socket permission invariant (2026-09-23)¶
compat/python-vmaf/tools/decorator.pymemoization keys strictly use SHA-256. The upgrade fromhashlib.sha1tohashlib.sha256(..., usedforsecurity=False)is governed by ADR-1307 (partially superseding ADR-1222 for its SHA-1 keep-open disposition) and clears Semgrep alerts 947, 948, and 949 at source (0 findings in SARIF). Ephemeral runtime and per-process memoization caches accept clean cold invalidation on upgrade; no legacy SHA-1 fallback or# nosemgrepsuppression is retained. Concurrency is serialized viathreading.RLock(), cross-process cache updates inpersist_to_fileare synchronized via re-entrant_file_lock(fcntl.flockon POSIX,msvcrt.lockingon Windows) and disk cache merge, including recursive entries, and atomic writes use unique temporary files viatempfile.mkstemp. The reviewed base wrote JSON directly to the destination; it never used PID-based temp files. Do not revert to SHA-1 or introduce non-unique temporary names during upstream merges.ai/sidecar/online_trainer.pyuses owner-only mode0o600and owns paths by identity (ADR-1309). The old0o660rationale relied on a cross-UID/same-GID Helm topology that the current chart does not wire. Do not restore unconditional group access or a Semgrep suppression. A future group-shared mode needs an explicit configuration surface, chart wiring, threat model, and end-to-end tests. Preserve the adjacent owner-only lifetime claim; no-followlstatchecks; non-blocking socket probes where onlyECONNREFUSEDproves staleness;EADDRINUSErefusal for a full accept queue or any pending/unverified result; unchanged device/inode validation before stale removal; bound identity verification before listen; and owned-identity-only cleanup. Symlinks, ordinary files, active listeners, and replacement objects must survive. Keep the_ConnectionRegistryaccept-before-spawn and promptclose_all()shutdown invariants. The POSIX socket regressions run through the AI suite; all 26 decorator regressions run throughcompat_decoratorNox and the hosted Linux/macOS/real-Windows matrix.- The hosted decorator lane exact-pins
pytest==9.1.1. This branch predates the hash-locked Python dependency infrastructure being developed separately, so it must remain independently executable rather than reference a lock file absent from its base. When rebasing after that infrastructure lands, reconcile this direct pin with the manifest-owned lock selected forbuild.yml; do not restore an unversionedpip install pytest. - Standalone
ai/sidecar/online_trainer.pyinvocations requireVMAFX_SIDECAR_CHECKPOINT_DIR. The production default/mnt/vmafx-models/onlineassumes a container volume mount. Standalone quick-starts and doc contracts must explicitly configure a writable checkpoint directory alongsideVMAFX_SIDECAR_SOCKETto avoid failing closed withPermissionErroron root-owned/mnt. - The required sidecar suite is warning-clean on PyTorch 2.14. Preserve equal-length flattened predictions and targets (including batch size one), explicit count-mismatch rejection, and the opset-17 tuple-argument
dynamic_shapes=({0: "batch"},)export. Do not restoresqueeze(-1)plus broadcasting or the legacydynamic_axesexporter argument. Derive the new-sample window from batch size and replay mix, reserve only that oldest window, sample replay without replacement when history is sufficient and with replacement only when history is smaller, and keep one training owner through restore or commit. Only the reserved new-row count advances the checkpoint gate; replay rows never do. A failedRuntimeErrororValueErrorstep restores that window at the front; concurrent arrivals remain queued behind it, cannot train ahead, and cannot be cleared by its eventual retry. Keep the admitted backlog bounded with explicit retry backpressure; do not admit a capacity-rejected sample to the replay buffer. Preserve admission-aware ACKs: an admitted sample restored after a failed step isok: true/retry_queued: trueand must not be resubmitted, while a capacity-deferred sample remainsok: falsewithretryable: trueif the oldest-window step fails. The Go client must retain the latter in flight across reconnects and retry it ahead of the bounded queue without incrementingdelivered, while preserving the former as an accepted delivery. Never put this retry back into a queue slot that a concurrent producer can refill. Preserve the explicit Go send disposition: only transport ambiguity andretryable: trueretain an in-flight sample; permanent local JSON encoding failures incrementdropped, remain on the current connection, and cannot starve later valid samples. Non-retryable sidecar rejection remains terminal without changing either counter. Socket lifecycle regressions signal readiness only after the reallisten()succeeds and surface every server-thread exception to the parent test.
integration/zero-warning-hiss21 — the silent-revert allowlist is a live, expiring file (2026-09-22)¶
scripts/ci/silent-revert-allowlist.jsondescribes the difference between this branch andmaster, not a permanent policy. Each entry exists only because master still carries the state being superseded: the retired numbered workspace root thatc2a3c7e0fwrote in place of.corpus/(ADR-1277), and the by-value HIP ADM kernels that92ea978a4left after silently reverting ADR-0759. Once master carries the superseded state,check-silent-revert.pystops producing those findings and the entries are dead weight — delete them rather than carrying them forward. A rebase that keeps an entry whose finding no longer exists has left a suppression behind, which is the failure ADR-1291'sundoes+evidencefields exist to make visible.- Do not widen an entry to resolve a rebase conflict. Each entry matches on the detector, the exact path, the commit undone and every evidence line. If a rebase moves the reversal onto new paths or new lines, the honest resolution is to re-run
make silent-revert-check, read the new findings, and re-derive the entry from them; droppingundoesor looseningevidenceto make the gate quiet converts a declaration into the path exclusion ADR-1291 refused. - The two detector repairs are behavioural, not cosmetic.
reverse-hunknow skips paths the merge deletes andresurrectednow requires the line to be text the target once held and lost. A rebase that restores either detector's older body — for instance by taking master's side ofcheck-silent-revert.pywholesale — puts back twelve false findings on this branch alone.scripts/ci/tests/test_check_silent_revert.pyis the tell: five of its cases fail against the unrepaired gate.
chore/hiss21-core-tools — governance surfaces a rebase must not undo (2026-09-21)¶
.standards-baseline.jsonwas re-recorded downward, 1411 → 965 → 938, withpraetorctl baseline -recordfrom a clean clone of this branch, after the branch's HISS-21 burn-down cleared the debt. The 965 step was recorded against the praetorctl build pinned on 2026-09-21 (7c0f803d40ee); 938 is the same tree re-recorded against7ec6f6ca5e28, which fixes that build's signature-relative function-length regression. The file is append-only downward (see the rule further down this page): resolve a rebase conflict here by re-recording on the merged tree, never by taking whichever side has the larger count, and never with-allow-increase.README.mdcarries a managed<!-- praetor:readme-governance:start -->block. The audit engine validates its content, not just the markers, so the block is not free prose: a rebase that reflows it, renames the commands back tostandardsctl, or lets theDebt Baselinerow drift from.standards-baseline.json's recorded count will failpraetorctl auditwithmanaged README governance block is stale. Keep the row in step with the baseline whenever the baseline is re-recorded..config/hiss/testdata/HISS-04/c/negative/loc-at-cap-60.cis length-critical.exactly_sixtymust span exactly 60 lines from its opening brace through its closing brace, which is what the brace-tracked Praetor matcher counts; the signature lines above the brace are not counted. Adding or removing one line inside the function turns the negative fixture into a false positive or stops it pinning the boundary, andpraetorctl hiss coverage --verifyfails either way. Measured against praetorctl7ec6f6ca5e28: a brace-to-brace span of 60 is clean, 61 is reported. Note the two enforcers differ by exactly one line — clang-tidy'sreadability-function-size(LineThreshold: 60) counts closing-brace line minus opening-brace line, so it is clean at a span of 61 and reports at 62 (clang-tidy 22.1.8, this repository's.clang-tidy). A function Praetor rejects can still be clean under clang-tidy; Praetor is the stricter of the two and is what gates.
Historical note for anyone re-reading an older commit message on this branch: the praetorctl build pinned on 2026-09-21 (7c0f803d40ee) measured the span from the signature line instead of the opening brace, which inflated every wrapped-signature function by one or more lines. Commit messages and notes written against that build quote signature-relative spans; the numbers above are the corrected, brace-relative ones. Re-measure rather than trusting a quoted span.
chore/hiss21-core-tools — yuv_input_open cleanup path is fork-shaped (2026-09-21)¶
Upstream keeps the goto fail form; the fork splits the pixel-format and buffer-size decision into yuv_input_set_plane_geometry() (ADR-0977's size_t-precision cast lives there now), so resolve a sync conflict in favour of the helper rather than restoring the label.
fix/bug-arm64-tidy — the ratchet grows an arm64 cross lane (2026-09-21)¶
No upstream C-source impact: the change is Makefile (fork-only section below the "Fork-specific targets" marker), .github/workflows/lint-and-format.yml, scripts/ci/AGENTS.md, docs/, and the new scripts/ci/tidy-baseline-arm64.json. Netflix ships none of these.
Rebase impact for anyone touching core/src/feature/arm64/. Those 20 translation units, and the ARCH_AARCH64 bodies of the core/test/ SIMD parity tests, are now measured by the ADR-1283 arm64 ratchet lane and only by it — no x86 build in the project emits a compile command for them. Re-measure after any change there:
meson setup build-arm64 core --cross-file build-aux/aarch64-linux-gnu.ini \
-Denable_cuda=false -Denable_sycl=false -Db_lto=false
make tidy-ratchet LANE=arm64 TIDY_RATCHET_BUILD_DIR=build-arm64
Two invariants the lane depends on. TIDY_RATCHET_EXTRA_arm64 must keep both --extra-arg=--target=… and --extra-arg=--sysroot=…: the compile database is produced by aarch64-linux-gnu-gcc, and clang-tidy will otherwise parse <arm_neon.h> / <arm_sve.h> against the host's x86 headers and fail every NEON translation unit as a compile error (ratchet exit 4, which is a failed measurement, never a clean one). And the generated model-JSON → C sources under build-arm64/src/ must exist on disk before measuring, exactly as the cpu lane builds before it measures.
exclude_untidyable() in lint-and-format.yml keeps its ^core/src/feature/arm64/ entry on purpose — that job's CPU-only build/ really has no command for those files. Do not "fix" it by deleting the line on a conflict; the lane, not the exclusion list, is what bounds this tree.
The baseline in this change was recorded on a workstation cross toolchain, so it is comparable only against the same one. If the lane is ever promoted to a CI context, re-record it from that runner's own measurement (ADR-1230), the way tidy-baseline-cpu.json is taken from the tidy-ratchet-cpu artifact.
fix/bug-silent-revert — a merge may not quietly rewind the target (2026-09-21)¶
Fork-local CI tooling; no upstream C-source impact. Preserve scripts/ci/check-silent-revert.py, its fixture suite scripts/ci/tests/test_check_silent_revert.py, the Silent-Revert Guard job in rule-enforcement.yml and its entry in required-aggregator.yml when resolving workflow conflicts. The gate must keep measuring the merge result (git merge-tree --write-tree base head), not git diff base..head: the branch-tree form reports every file the target changed that the branch never touched and is unusable on any branch that is behind. It must keep resolving the live target tip rather than github.event.pull_request.base.sha, since the defect class is master moving after the branch was cut.
ADR-1291 supersedes this section's original no-allowlist rule with two narrow declaration mechanisms. A one-off whole-PR revert uses a revert: title, reverts: #N, or intentional revert: <reason>; the unedited intentional revert: REASON placeholder must keep failing. An accepted ADR may instead declare a live reversal in silent-revert-allowlist.json, but only with detector, exact path, exact full commit for reverse-hunk, and an evidence regex that matches every line. That file is an expiring declaration of one known reversal, never a bare path or source-tree exclusion. Keep the fail-closed exits (unresolvable ref, no merge base, conflicting merge, git below 2.38) — a case the gate cannot analyse must never print "clean".
GENERATED_PREFIXES covers rendered files only (CHANGELOG.md, docs/adr/README.md, docs/adr/by-tag/, mkdocs.yml, the standards and tidy baselines); do not widen it to source trees to quiet a finding. The conflict-marker exclusion in is_evidence() is load-bearing: 0c494cca0 once committed three markers into core/src/feature/cuda/integer_vif_cuda.c and the PR that deleted them reset the file to its pre-marker blob.
Regression command: python3 scripts/ci/tests/test_check_silent_revert.py (22 tests; test_real_history_replay needs 31a51afb2 and 92ea978a4 in the clone and skips otherwise). See ADR-1284.
fix/bug-gpu-lint — GPU NOLINT citations and the SYCL tidy database (2026-09-22)¶
Fork-local; no upstream Netflix C source is touched (upstream has no SYCL, HIP or Metal backend, and the CUDA edits are comment-only). Two invariants to preserve when rebasing lint tooling:
make tidy-ratchet/tidy-ratchet-writemust keep expanding the per-laneTIDY_RATCHET_COMPDB_<lane>hook betweenwrite-compile-commands.pyandtidy-ratchet.py. Forsyclthat hook isscripts/ci/gen-sycl-compile-commands.py; meson emits the SYCL feature TUs asCUSTOM_COMMANDrules, so dropping the hook silently returns the lane to measuring zero SYCL translation units.scripts/ci/tests/test_tidy_ratchet_sycl_compdb.pyis the regression contract. The GPU lanes additionally need-Db_lto=falseand a build directory outside the repository; both are documented in the Makefile anddocs/development/ci.md.- Every GPU-lane
NOLINTnow carries an inlineADR-NNNNtoken inside the windowcount_uncited_nolints()scans (previous, same or next line, or the enclosing/* ... */comment — note the block scan looks forward only, so a citation placed above the marker in the same comment does not count). Do not reflow those comments in a way that moves the token out of the window, and do not restore the seven suppressions this branch deleted: each was verified to suppress a diagnostic that clang-tidy does not emit.launch_dwt_hori_pairincore/src/feature/sycl/integer_adm_sycl.cpplost its unusedh_addparameter and the matchingDwtShiftsfield; the horizontal pass derives that addend fromh_shift, so do not reintroduce the parameter on a merge.
fix/configured-lint-warning-exit — diagnostics cannot pass as green (2026-09-21)¶
The fork-local configured-lint driver must pass --warnings-as-errors=* to every clang-tidy invocation. Clang-tidy's default exit status is zero for ordinary warnings, so relying on return codes alone silently accepts findings from upstream, vendored, test, GPU, and fork-native translation units. Preserve the all-diagnostic promotion and its zero-exit-warning regression fixture when rebasing CI tooling. Do not replace it with a baseline, touched-file filter, log-text heuristic, or origin exemption. Cppcheck must still run after any clang-tidy failure so both reports remain available.
fix/go-duplicate-cleanup — shared Go service and CLI plumbing (2026-09-21)¶
Fork-local ownership cleanup with no upstream C-source impact. Preserve internal/app/scoringservice as the single implementation of server/controller metrics, scorer lifecycle, legacy probes, and JSON responses. Preserve pkg/model.CLIArgument{,OrDefault} as the formatter for Go subprocess callers, the corpus alias to scorebackend.UnavailableError, standard-library sorted registry keys, and the shared root-resource deep-copy helper. When rebasing a binary or Go port, adapt the shared owner rather than restoring a local copy. praetorctl dedupe scan . is the regression command and must remain explicit in the required Standards workflow, pre-commit, pre-push, and make verify-all; the general audit does not include it.
fix/ci-fail-closed — test and scan exit status is evidence (fork-local, 2026-09-20)¶
No upstream C-source impact. Preserve scripts/ci/test_fail_closed_ci.py and both of its callers when resolving workflow or hook conflicts. Required test, coverage, benchmark, and discovery commands return their real exit status. A step may continue solely to collect diagnostics when a later if: always() step checks steps.<id>.outcome and fails the job. Advisory Semgrep policy is unchanged, but its command must still expose a failed step outcome. Do not restore ignore_outcome, coverage -i, || true, or disabled pipefail on these paths. Coverage GPU is required and listed by the aggregator; do not restore its stale (advisory) name or job-level continue-on-error.
fix/ai-trainer-warning-cleanup — preserve executable trainer helpers (2026-09-21)¶
The five AI corpus/trainer scripts in this cleanup are fork-local. Upstream syncs have no direct overlap, but future trainer refactors must preserve two runtime fixes: train_fr_regressor._standardize() returns new arrays instead of mutating pandas-owned NumPy views, and train_konvid_mos_head._export_onnx() uses dynamic_shapes with a two-row example batch so current PyTorch exports a genuinely dynamic batch axis without warning. Do not restore the old in-place normalisation or dynamic_axes call. Parser and workflow helpers remain below the HISS-04 60-line limit; no public CLI option or model threshold changed.
fix/sycl-strict-clean — strict SYCL diagnostics and AOT command contract (fork-local, 2026-09-21)¶
All touched core/src/feature/sycl/*.cpp files are intentionally warning-free under the fork's oneAPI, clang-tidy, cppcheck, and HISS profiles. Upstream origin is not an exemption: when resolving conflicts, preserve the phase helpers and explicit size-domain arithmetic rather than restoring long kernel lambdas or analyzer suppressions. Device-side arithmetic remains fp32-only and filter/reduction order remains unchanged.
The icpx multi-target command must keep -Xsycl-target-backend=spir64_gen '-device <list>'; unqualified -Xs sends the device selector to the portable spir64 target and produces an unused-argument warning. Keep the paired removal in scripts/ci/gen-sycl-compile-commands.py and its test_sycl_aot_command.py pre-commit contract when either Meson source or the lint projection is rebased. Language standards stay in Meson's built-in fallback lists; do not restore manual -std= probes.
fix/pelorus-interop-sync-v022 — exact v0.2.2 parser safety mirror (2026-09-20)¶
No Netflix-upstream file is involved. The cross-repo conflict surface is the fork-local Pelorus mirror, its sync/lint tooling, and the existing required Pre-Commit workflow. ADR-1276 records the re-pin and fail-closed maintenance contract while preserving ADR-1113's base mirror decision.
PELORUS_VENDOR_SHAis the full released v0.2.2 commit93bef1206d68d9e09024c08a12732fb8e77b9b16. ABI stays 1.3. A future rebase must not infer that an unchanged ABI minor makes a parser fix optional: reviewed correctness/security releases are re-pin triggers too.pelorus_interop.ccopies the wire header and directory entries into aligned locals before access. Do not restore pointer casts from byte addresses; a valid caller buffer may have any base alignment. Keep the rejection of aheader_sizethat is not 8-byte aligned.- From its first vendored include onward,
test_pelorus_interop.cis exact Pelorus source except for the include rewrite. Do not reapply PR #1351's VMAFx-onlyNOLINTband or(void)casts. Formatting/tidy exclusions belong in VMAFx tooling. The drift guard renders the VMAFx prefix canonically from the source pin and ABI version, then compares the complete fixture exactly. - Native-lint exemption uses an explicit manifest-owned path set. The sync guard rejects extra or missing tracked files in the Pelorus header/source namespaces; do not restore prefix-wide exemption without that fail-closed manifest check.
- The guard must keep reading the pinned Git object and failing closed for a plain directory or a checkout that lacks it. The existing required
Pre-Commitjob checks out that exact object and runs the default guard. - The fixture's
fopen(path, "w")remains an upstream-owned CodeQL finding, tracked separately indocs/state.md. Fix it in Pelorus and re-vendor; do not patch only the VMAFx mirror.
fix/bounded-process-execution — repository automation process boundary (fork-local, 2026-09-20)¶
scripts/lib/safe_subprocess.pyis the process-execution boundary for Python automation underscripts/: executable allowlist, bounded argv and captured output, explicit deadline, closed unused stdin, and process-group cleanup on timeout, output overflow, and caller cancellation. Cancellation cleanup must finish beforeCancelledErrorpropagates. When an upstream sync or script port adds a directsubprocesslaunch in this scope, adapt it to the boundary; do not restore anS603annotation.- Consumer tests deliberately preserve each command's prior return and output semantics. Keep the
allowed_executablesset narrow and command-specific; broadening it to whatever happens to be onPATHdefeats the boundary. scripts/__init__.pyand canonicalscripts.lib.safe_subprocessimports are load-bearing. Direct-path scripts first prepend their resolved repository root; do not restore thelib.safe_subprocessfallback, which gives mypy two names for the same file. Keep the two-root regression test.scripts/ci/agent-eligibility-precheck.pyimports tracker and process exceptions throughscripts.*. Its GitHub search/list checks intentionally fail soft after emitting a notice; keep the two CLI regressions wired to the process-boundary hook so aCommandFailedidentity mismatch cannot turn an offline dispatch check into a traceback..github/ci-impact.jsonclassifies every tracked top-level entry. Add new roots toknown_prefixesorknown_filesin the same change that creates them so routing does not silently degrade to the fail-closed full plan.
See ADR-1270 and Research-2071.
chore/ffmpeg-n9.0.2 — stable patch baseline (fork-local, 2026-09-20)¶
build-config.envownsFFMPEG_TAG=n9.0.2;Dockerfile,Dockerfile.ffmpeg,dev/Containerfile, anddocker/Dockerfile.nodeare generated mirrors. A future stable-tag refresh must update them throughscripts/ci/ffmpeg_patch_stack.py --refresh, never as independent pins.- The existing 18 integration patches remain byte-identical; patch 0019 adds the 126-diagnostic GCC 14/16 warning-clean hardening. All 19 entries replay cumulatively on upstream commit
946fcce07b6dcd0331c8cc609192aeff5e1924f8, producing tree1fd76f79179a6a51b5bb773a89c5334c6324c755. Preserve series order and runpython3 scripts/ci/ffmpeg_patch_stack.py --check; independentgit apply --checkcalls are not an equivalent gate. - FFmpeg n9.0.2 has removed libnpp support. Its retained
--enable-libnppswitch only warns that enabling it does nothing, so the dev image deliberately omits the flag to keep configure warning-clean. Re-add it only if a future FFmpeg release restores a real probe and the matching CUDA contract is validated. - Patch 0019 must not be replaced by warning suppressions or component removal. VVC remains enabled; its scaled-prediction scratch belongs to
VVCLocalContext. The maintained root CUDA, compatibility, dev, and node FFmpeg builders, hosted integration lanes, and patch smoke harness all use--fatal-warningsplus a compiler-log warning gate. The ordinary hosted compatibility matrix applies patch 0019 alone so it remains independent of fork integration surfaces while compiling the warning-clean pinned source. Hosted integration and the smoke harness compile all test programs and run every generated, sample-independent FATE target inside that log gate; do not narrow it back to production objects because APV, CABAC, checkasm, and newer compiler versions have their own warning inventory. Capturemake -s fate-listbefore filtering its output tofate-*; on a pristine tree it can also emit a generated-makefile status line, which is not a target. Release-tag checkouts must usescripts/ci/checkout-annotated-tag.sh: direct shallow clones warn on annotated FFmpeg and AMF tags in the container Git version.Dockerfile.ffmpegmust continue replaying the canonical series rather than the deletedpatches/ffmpeg-libvmaf-gpu.patchpath. No patch path may fall back to fuzz-capablepatch -p1. The dev image's encoder inventory is fail-closed because listing compiled encoders needs no device. - The partial libvmaf builders in
docker/Dockerfile.nodeandDockerfile.go-servermust copyscripts/ci/check-msvc-clz-shim.sh, install bothxxdandmake, preservelibvmaf.so*links withcp -a, and stage Meson's generatedlibvmaf.pc. Do not restore the handwritten pkg-config template based onVMAFX_VERSION: that is the product version, while FFmpeg probes the independent libvmaf interface version. core/meson.builduses orderedc_std/cpp_stdpreferences per ADR-1273; do not restore direct standard flags from ADR-1056.core/src/model.ckeeps the built-in table behindVMAF_BUILT_IN_MODELS, andcore/test/test_model.cmust continue passing with that option disabled.- The root Makefile resolves
VIRTUAL_ENV_PATHwith$(abspath $(VENV))and passes that absolute directory to Meson/Ninja recipes. Meson invokes its recorded Ninja path fromcore/buildto createcompile_commands.json; restoring a relative$(VENV)/binprefix makesmake lintfail after the build. Keepcheck_makefile_venv_paths()and its fixture in sync.
See Research-2073 and ADR-1273.
perf/cambi-simd-gaps-2 — AVX-512 and NEON for every CAMBI stage, scanned AVX2 c-values (fork-local, 2026-09-18)¶
Everything here is fork-local: upstream Netflix/vmaf ships AVX2 CAMBI kernels only. What a sync must keep:
core/src/feature/cambi.c,setup_callbacks(): upstream's AVX2 block stays as upstream writes it except for one line:calc_c_values_callbackbinds the fork'scalculate_c_values_scan_avx2, not upstream'scalculate_c_values_avx2. Upstream's driver visits every column of the sliding-histogram walk and measured 0.81–0.83x of scalar in icx builds (icx builds the published container), 0.80–1.11x under Clang depending on the build; the scanned driver is 2.1–5.5x scalar (Research-2065). Keep the fork's binding on a sync.calculate_c_values_avx2itself stays inx86/cambi_avx2.cas upstream writes it, built and checked bytest_cambiandtest_cambi_stage_simd, so upstream changes to it merge cleanly; if upstream reworks it, re-measure against the scanned driver before switching back. After the AVX2 block the fork adds an AVX-512 block (derivative, c-values, mode filter, decimate, dp and mask rows, underHAVE_AVX512) and an aarch64 block (derivative, c-values, decimate, dp row).filter_modeandcompute_mask_rowstay scalar on aarch64 on purpose (ADR-1256).anti_dithering_filter()tries AVX-512, then upstream's AVX2 branch, and NEON on aarch64. When upstream rewrites this block, re-add the fork's branches instead of taking upstream's version wholesale; that is howe3fd1c88adropped the previous AVX-512 and NEON dispatch.core/src/feature/cambi_c_values_frame.his the fork's copy of thecalculate_c_valueswalk (first pass, top edge, middle slide, bottom edge, thev_band_base/v_band_sizederivation), driven by per-ISA column scans and thecambi.hupdate helpers; the AVX2, AVX-512 and NEON drivers all use it. The scans' per-column predicates (cambi_column_in_band,cambi_column_slide_needed) live there too. If upstream changes that walk, those helpers oruh_slide's skip condition, mirror the change here and in the scans (scan_*_avx2at the end ofx86/cambi_avx2.c,scan_*_avx512inx86/cambi_avx512.c,scan_*_neoninarm64/cambi_neon.c). A scan may flag too many columns but never too few.test_cambi_stage_simdcompares both drivers with the scalarcalculate_c_valuesand fails on any mismatch.cambi_increment_range_neon/cambi_decrement_range_neonwere retired: the NEON driver's plain C loops compile to the same adds. Do not bring them back from an old branch.core/test/test_cambi.c:test_calculate_c_values_scalar_avx2_paritygates onvmaf_get_cpu_flags_x86()(CPUID). Keep it if upstream touches that test; the oldvmaf_get_cpu_flags()gate silently skipped the comparison.
See Research-2065.
perf/cambi-spatial-mask-simd — upstream 86da14d03 adapted, not verbatim (2026-09-18)¶
Upstream 86da14d03 ("feature/cambi: AVX2 vectorize spatial-mask dp row and mask row") is ported, with three differences a sync must keep:
core/src/feature/x86/cambi_avx2.c:compute_dp_row_avx2carries the running prefix ascarry += broadcast(block total)instead of re-broadcasting lane 7 of the carried scan, andinclusive_prefix_epi32moves the low half's total withpshufd+ zeroingvperm2i128instead ofpermute2x128+shuffle+blend. Upstream's form is 0.64–0.74x of scalar under Clang and icx.compute_mask_row_avx2biases both compare operands by 2^31, so it is exact for anymask_index, not only box sums below 2^31. If upstream later changes these functions, take their intent and re-bench; do not replace the fork's bodies.core/src/feature/cambi.c: dispatch matches upstream for AVX2 and adds AVX-512 (#if HAVE_AVX512) for both rows and NEON for the dp row only (ADR-1256).vmaf_cambi_get_spatial_mask(the GPU twins' trampoline) passes the scalar row kernels.compute_dp_row/compute_mask_roware non-static with prototypes incambi.h, as upstream made them; the other functions upstream exported for checkasm staystaticin the fork.- No checkasm in the fork: upstream's
check_cambi.ccases are covered bycore/test/test_cambi_spatial_mask_simd.cinstead;test_cambi.ctakes upstream's two extraget_spatial_mask_for_indexarguments.
compute_*_row_avx512 and compute_*_row_neon are fork-local and have no upstream counterpart. The same branch refactors calculate_c_values_row_neon into a per-pixel helper (touched-file lint, bit-exact under test_cambi_simd). See Research-2062.
ci/retire-i686-lane — the fork stays 64-bit only; x86 SIMD sources use no x86-64-only intrinsics (2026-09-18)¶
.github/workflows/libvmaf-build-matrix.yml: there is no i686 row (ADR-0691, ADR-1258). Upstream Netflix/vmaf has its own 32-bit cross build (f6d6dde1); do not port it. A merge that restoresi686: truerows is wrong: that is how the lane came back in384d97d03.core/src/feature/x86/adm_avx2.c,adm_avx512.c: 64-bit lane extraction from an__m128igoes throughextract_epi64_128(), defined besideextract_epi64(). Upstream calls_mm_extract_epi64directly; keep the fork form.core/src/feature/x86/psnr_avx2.c,psnr_sse_line_16_avx2(): the final 64-bit sum is read with_mm_storel_epi64, not_mm_cvtsi128_si64.
ci/retire-i686-lane — the CI build matrix of record (ADR-1259) (2026-09-18)¶
.github/workflows/libvmaf-build-matrix.ymlandbuild.ymlboth run;build.ymldoes not replace the matrix. ADR-0689, ADR-0691, ADR-0710 and ADR-0728 are superseded by ADR-1259. A merge resolution that drops or restores a lane is wrong unless an ADR asks for it:384d97d03undid ADR-0689 and ADR-0691 that way.- The unreleased changelog fragments
changelog.d/removed/native-build-sunset.md,changelog.d/removed/0691-vmafx-drop-legacy-build-paths.mdandchangelog.d/changed/0689-vmafx-ci-matrix-dedupe.mdare deleted on purpose: they announced removals that never happened. Do not restore them.
fix/restore-reverted-security-fixes — two fixes that a stale merge and a re-vendor undid (2026-09-19)¶
Both regressions below came from taking a whole-file "theirs" over a security fix. Neither conflicts with upstream Netflix/vmaf (all files are fork-local); the risk is a future fork-internal rebase, squash-merge or re-vendor. See docs/state.md, T-VENDORED-CJSON-BANNED-FUNCTIONS-REVERTED-2026-09-19 and T-SHELL-INJECTION-ROUND2-REVERTED-2026-09-19.
- Vendored cJSON is upstream 1.7.19 plus a fork delta; a re-vendor must re-apply the delta.
core/src/mcp/3rdparty/cJSON/cJSON.cdiffers from upstream on purpose: boundedsnprintf/memcpyinstead of the bannedsprintf/strcpy(ADR-0683, ADR-1061),cJSON_GetArraySizesaturating atINT_MAX, and nogotoor function over 60 lines (ADR-1142).6ab6a58b1(PR #883) dropped upstream's files in unchanged and reverted the first two. The delta, function by function, and the re-vendor procedure are incore/src/mcp/3rdparty/cJSON/AGENTS.md.cJSON.his unmodified upstream.core/test/test_cjson.cpins the version string, so a re-vendor failstest_versionuntil the delta has been carried over and the version updated. - Never answer a finding in vendored code with an exclusion. The
vmaf-no-strcpy-strcat-sprintfrule in.semgrep.ymlno longer excludes/core/src/mcp/3rdparty/**,.semgrepignoreno longer listscJSON.c, and thesemgrep-localhook no longer skipscore/src/pdjson.{c,h}. If a rebase brings any of those back,scripts/ci/tests/test_semgrep_vendored_scope.pyfails; resolve by keeping the exclusion out, not by editing the test. scripts/ci/sycl-bench-env.shanddev/scripts/dev-mcp-entrypoint.sh: when resolving a conflict, never take the side that hasbash -c "... '$ROOT/setvars.sh' ..."oreval "${cmd}".d9c33dc68(PR #414) did exactly that to45d536962(PR #350), along with PR #350's state row, rebase note andscripts/ci/AGENTS.mdrows. The correct forms arebash -c '... "$1/setvars.sh" ...' _ "$ROOT"and"${prog}" 2>&1 | grep ....- Keep the three new pre-commit hooks wired (
test-semgrep-vendored-scope,test-sycl-bench-env,test-dev-mcp-entrypoint-probe). The sycl test existed when its fix was reverted and would have failed; nothing ran it.test-dev-mcp-entrypoint-probe.shextracts_probe_with_retryfrom the real entrypoint by its opening line_probe_with_retry() {, so keep that function at top level under that name. core/test/test_cjson.cmust stay outsideif get_option('enable_mcp')incore/test/meson.build. It compilescJSON.citself, which is the only reason the CPU-lane Tidy Ratchet and Cppcheck jobs see that file (the lane builds without MCP).cJSON.chas no entry inscripts/ci/tidy-baseline-cpu.jsonbecause it measures zero; any warning a rebase introduces there is a ratchet regression..standards-baseline.jsonwas re-recorded downward for this change from a clean clone with the pinned engine. A rebase that reintroduces a banned call or agotoincJSON.cnow fails the baseline-only CI audit as new debt, not only the local touched-file rule.
fix/hip-pageable-upload-race — HIP extractors wait for their picture uploads (2026-09-19)¶
Rebase impact: none upstream. Netflix ships no HIP backend; every touched source under core/src/hip/, core/src/feature/hip/ and the two HIP tests is fork-local, and the core/test/meson.build hunk sits inside the enable_hip block upstream never edits.
Invariants (see also core/src/feature/hip/AGENTS.md, "Picture uploads"):
- A HIP extractor must not return from
submit()while an upload from a host picture is in flight. Every plane taken fromVmafPicture::datagoes throughvmaf_hip_picture_upload()(core/src/hip/picture_hip.{c,h}), on the extractor's private stream. A port of a CUDA twin brings a barecuMemcpy2DAsyncwith it: CUDA pictures are device memory with a ready event, HIP pictures are pageable host memory, so replace the copy with the helper rather than translating it tohipMemcpy2DAsync. - The helper waits on an event, not on the stream, so that a null-stream caller does not wait for every other stream on the device. Keep the wait when a copy fails part-way: the copies already enqueued still read the pictures.
- The wait costs throughput on a single-queue GPU (T-HIP-UPLOAD-WAIT-THROUGHPUT- 2026-09-19). Replace it with extractor-owned pinned staging, never by removing it.
test_hip_upload_racefails on every run without it. - A new HIP extractor that uploads a picture gets a row in
race_cases[]incore/test/test_hip_upload_race.c. The reference picture is not exempt:vmaf_read_pictures()keeps it alive throughprev_ref, which hides the defect from a pooled test but does not remove it. core/src/hip/hip_handle.his the one place that converts theuintptr_thandles ofkernel_template.hback tohipStream_t/hipEvent_t. The twelve touched extractors no longer define__HIP_PLATFORM_AMD__themselves;hip/meson.buildpasses it.
The twelve extractor files were also brought to zero clang-tidy findings and zero HISS findings, which the touched-file rule requires: init() / close() share one *_release() teardown in place of the goto ladders, and it drains the stream before it frees a buffer. float_psnr_hip, float_moment_hip and float_ssim_hip used to free first. A conflict in those functions should be resolved towards the single teardown.
Touched files: core/src/hip/picture_hip.{c,h}, core/src/hip/hip_handle.h (new), core/src/feature/hip/{ciede,float_adm,float_moment,float_motion,float_psnr,float_ssim,float_vif,integer_motion,integer_motion_v2,integer_psnr,integer_ssim,integer_vif}_hip.c, core/test/hip_pooled_fixture.h (new), core/test/test_hip_upload_race.c (new), core/test/test_hip_ssim_parity.c, core/test/meson.build, core/src/feature/hip/AGENTS.md, core/src/hip/AGENTS.md, docs/backends/hip/overview.md, docs/state.md, changelog.d/fixed/hip-pageable-upload-race.md, scripts/ci/tidy-baseline-hip.json.
fix/hip-integer-ssim-int64-kernel — HIP integer SSIM runs the CPU kernel (2026-09-18)¶
Rebase impact: none upstream. Every touched source is fork-local: Netflix ships no HIP backend, and integer_ssim_hip.c, integer_ssim_score.hip, core/src/hip/dispatch_strategy.c and test_hip_ssim_parity.c exist only in the fork. The core/src/meson.build hunk adds one hip_cu_extra_flags entry and the core/test/meson.build hunk replaces the should_fail registration with a variant loop; both are inside enable_hip blocks upstream never edits.
Invariants (see also core/src/feature/hip/AGENTS.md):
integer_ssim_score.hipmirrorsinteger_ssim.c::calc_ssim(). The 9-tap integer weights, the k_min / k_max truncation formulas,SSIM_K1/SSIM_K2spelled(0.01 * 0.01)/(0.03 * 0.03), and the operand order of the per-pixel term are all load-bearing. If an upstream sync changesinteger_ssim.c, change every GPU twin (CUDA, SYCL, HIP, Metal) in the same PR.hip_cu_extra_flags['integer_ssim_score']keeps-ffp-contract=off. Without it the per-pixel term fuses into FMAs and stops rounding like the CPU's. The 1x1 case shows it: with the flag HIP matches the CPU exactly on 5 of 5 frames.submit()waits for its two host-to-device uploads before returning. The pictures are pageable memory that the pool recycles as soon assubmit()returns. Keep the wait until a HIP picture pool (T7-10c) hands device pictures to the extractor.test_hip_ssim_parityfeeds 8 pooled frames and fails every run without it.- The block reduction in
integer_ssim_vert_combineis sized for the 16x8 launch ininteger_ssim_hip.c. ChangeISSIM_BLOCK_X/ISSIM_BLOCK_YandISSIM_HIP_BLOCK_X/ISSIM_HIP_BLOCK_Ytogether.
Touched files: core/src/feature/hip/integer_ssim/integer_ssim_score.hip (rewritten), core/src/feature/hip/integer_ssim_hip.c (rewritten), core/src/feature/hip/integer_ssim_hip.h (comment), core/src/hip/dispatch_strategy.c (two table entries, ADR-1138 bracket), core/src/meson.build, core/test/meson.build, core/test/test_hip_ssim_parity.c, core/src/feature/hip/AGENTS.md, docs/backends/hip/overview.md, docs/metrics/ssim.md, docs/metrics/features.md, docs/state.md, changelog.d/fixed/hip-integer-ssim-int64-kernel.md, CHANGELOG.md, scripts/ci/tidy-baseline-hip.json (scoped tightening).
perf/hip-adm-buffer-by-pointer — HIP integer ADM buffer by pointer, re-applied (2026-09-18)¶
ADR-0759 (the four HIP ADM kernels that read AdmBufferHip take it by pointer) landed in #101 and was silently reverted by the next merge, #102, whose branch predated it (T-HIP-ADM-ADR0759-REVERTED-2026-09-18). The HIP twin is fork-only, so an upstream sync cannot conflict with it; the risk is another fork branch cut before this change. When merging or rebasing anything that touches these files, keep:
core/src/feature/hip/integer_adm/adm_csf.hipandadm_cm.hip:adm_csf_kernel_1_4,i4_adm_csf_kernel_1_4,i4_adm_cm_line_kernelandadm_cm_line_kernel_8takeconst AdmBufferHip *__restrict__ buf_ptr.adm_csf.hipalso carries its lint restructure (helpers in an anonymous namespace,csf_band_sample()); kernel names and launch layouts are unchanged.core/src/feature/hip/integer_adm_hip.c:AdmStateHip::buf_dev, uploaded byadm_hip_upload_buf()at the end ofadm_hip_init_device()and freed byadm_hip_free_buf_dev()inclose_fex_hip()and the init failure paths; the four launches pass(void *)&s->buf_dev.
Check after any merge that touches them: grep -n 'AdmBufferHip buf' core/src/feature/hip/integer_adm/*.hip must print nothing. Kernel and host must change together: a by-value kernel launched with &s->buf_dev gets 328 bytes copied from that address as the struct and dereferences whatever follows buf_dev in AdmStateHip.
fix/gpu-adm-dwt2-16bit-overflow — GPU integer ADM 16-bit vertical DWT sum (2026-09-18)¶
Upstream Netflix/vmaf's CUDA integer ADM sums the scale-0 vertical DWT response of 16-bit samples in int32_t, which overflows once three samples reach 42456 (T-GPU-ADM-DWT2-16BIT-INT32-OVERFLOW-2026-09-18). When a sync touches these files, keep the fork form:
core/src/feature/cuda/integer_adm/adm_dwt2.cuand its HIP twinadm_dwt2.hip: the fused scale-0 kernel accumulates inDwtVertAccum<T>::type, which is int64 foruint16_tinput. The kernel is also split intoadm_dwt2_load_column(),adm_dwt2_vert_tile()andadm_dwt2_hori_tile(), and the device helpers sit in an anonymous namespace; re-apply upstream's arithmetic intent onto that structure rather than taking its file.core/src/feature/metal/integer_adm.metal: the raw vertical DWT kernel sums inlong.
fix/gpu-adm-tiny-frames — GPU integer ADM on frames 17 to 32 pixels wide (2026-09-18)¶
Upstream Netflix/vmaf ships the CUDA integer ADM this fork mirrors, and it carries both defects fixed here (T-GPU-ADM-TINY-FRAME-SHIFT-2026-09-18). When a sync touches these files, keep the fork form:
core/src/feature/cuda/integer_adm_cuda.c: the scale-0 cube and inner-accum rounding constants areadm_half_shift(x)fromcore/src/feature/adm_csf_fixed_point.h, where upstream writes1 << (x - 1);init_fex_cuda()starts withadm_frame_size_check().core/src/feature/cuda/integer_adm/adm_cm.cu, both scale-0 kernels: the neighbour clamps arepos_x[2] = min(pos_x[2], w - 1)andpos_y = min(pos_y, h - 1). Upstream'spos - max(0, 2 * (x - w) + 1)uses the base index and reads one column and one row past the band.test_cuda_adm_tiny_framesfails on either upstream form.integer_adm_cuda.c, 10/16-bit path:curr_ref_stridecomes fromref_picandcurr_dis_stridefromdis_pic; upstream swaps them. Keep the fork form.- Both files were restructured to zero clang-tidy findings (ADR-1142): helpers in anonymous namespaces, unused kernel parameters unnamed, the host glue split into single-purpose functions. Kernel names and parameter layouts are unchanged. A sync that touches them ports upstream's intent into the fork structure rather than taking upstream's text.
The HIP (integer_adm_hip.c, integer_adm/adm_cm.hip) and SYCL (integer_adm_sycl.cpp) twins are fork-only; the same rules apply to them. integer_adm_sycl.cpp's internals now sit in an anonymous namespace rather than behind C-style static.
SYCL now reproduces the CPU's integer semantics exactly where it used to widen (T-SYCL-ADM-INT16-SEMANTICS-2026-09-18): scale-0 intermediates wrap to 16 bits through adm_i16(), the diagonal csf_a rounds with 65535, and the scale 1-3 filter terms round with I4_FLT_ROUND, the wrapped -2^31 of the Netflix#955 quirk that ADR-0155 keeps for the golden values (entry 0048 below covers the CPU and CUDA/HIP forms; this is the SYCL one). If an upstream change to integer_adm.c alters any of those narrowings or rounding terms, change the SYCL twin with it; test_sycl_adm_tiny_frames (noise geometries) catches a mismatch.
port/upstream-2026-09 — Netflix/vmaf 03b5562c5..86da14d03 (2026-09-18)¶
Reconciles upstream through 86da14d03 (previous mark f85a85369, PR #1456). See Research-2063 for the evidence behind each verdict.
Ported:
03b5562c5:core/src/feature/x86/adm_avx2.cadm_decouple_avx2bounds its tail withright - ((right - left) % 8). The loop starts atleft, so a bound computed from column 0 stores pastright. The scale-0 AVX2 kernel is now consistent withadm_decouple_s123_avx2and every AVX-512 decouple.ea012e387(adapted):core/src/feature/arm64/adm_neon.chorizontal pass. The 8-wide loop stops at the same guarded bound the x86 DWT2 kernels use,half_w >= 2 ? half_w - 1 - ((half_w - 2) % 8) : 1, not at upstream's(w_half - 2) - ((w_half - 3) % 8). Everything from there on, and column 0, goes throughadm_dwt2_8_neon_hpass_column()andind_x. The fork's earlier "redo the last column" block is gone because the tail subsumes it. The kernel is split into row helpers, and the NEON accumulate/store macros are nowadm_neon_macc4()/adm_neon_store_shifted(). On a conflict, keep the fork structure and re-apply only upstream's arithmetic intent.1786bd961(one hunk):integer_adm.cinit_buffers()zeroesdata_buf, mirroring upstream'sadm_buffer_alloc(). The checkasm framework, the un-static refactor andadm_buffer_alloc()itself are not ported.
Fork-only fix upstream still needs:
integer_adm.cadm_dwt2_vpass_16()and the scalar vertical loops ofadm_dwt2_16_avx2()/adm_dwt2_16_avx512(): the 16-bit vertical DWT response is formed byadm_dwt2_vpass16_tap4()ininteger_adm.h, in int64. Upstream sumsfilter[k] * s[k]inint32_t, which overflows once three consecutive 16-bit samples reach 42456 (T-ADM-DWT2-16BIT-INT32-OVERFLOW-2026-09-18). When a sync touches these loops, keep the helper;test_integer_adm_dwt16_rangeaborts on the sanitizer lane with the upstream form. Scores are identical either way.integer_adm.cdwt2_src_indices_1d(): the mirrored tail starts at(n_half > 2u) ? n_half - 2u : 1uand the first loop is bounded withi + 2 < n_half. Upstream'sn_half - 2restarts the tail at 0 whenn_half == 2(any frame dimension from 17 to 32 at scale 3) and reads index -1. Keep the fork form on every sync. The zeroing above does not make the upstream form safe.core/test/test_integer_adm_tiny_frames.cguards it, and on the ASan lane it aborts on the old bound.x86/adm_avx2.candx86/adm_avx512.c: every rounding constant of a right shift isadm_half_shift(x)fromadm_csf_fixed_point.h, where upstream writes(uint32_t)pow(2, (x - 1)). At frame widths 17 to 32 the scale-0 cube shift is 0, and upstream's form converts infinity to an integer; the AVX-512 build turns that into0xFFFFFFFF. When a sync touches these lines, keep the helper.test_integer_adm_tiny_widths_simd_matches_scalarfails at 17x70 with the upstream form on an AVX-512 host, and the UBSan lane flags it on any x86 host.adm_half_shift()moved there frominteger_adm.c, whose frame-size check is nowadm_frame_size_check().x86/adm_avx2.candx86/adm_avx512.c,ADM_CM_THRESH_S_I_J_avx256/_avx512: after each centre tap'ssrai(..., 12)the fork sign-extends the low 16 bits (srai(slli(x, 16), 16)), reproducing the scalar reference's(int16_t)conversion. Upstream's vector macros keep 32 bits and differ from its own scalar on full-range content (T-ADM-CM-SIMD-NOISE-NOT-BIT-EXACT-2026-09-18). Keep the fork form;test_integer_adm_simd_noisefails without it.
Retired (ADR-1257): adm_dwt2_8_neon_apple_legacy() and the #if defined(__APPLE__) NEON dispatch branch in integer_adm.c. Apple AArch64 dispatches adm_dwt2_8_neon() like every other AArch64 host. Do not reintroduce a platform-specific DWT2 path. The three akiyo assertions in python/test/vmafexec_test.py whose Darwin branch recorded the old three-tap result (88.030322 if _IS_DARWIN else 88.030463) are to assert 88.030463 on all platforms. That file is protected by the golden-file edit guard, so the change is applied by a maintainer. The other two _IS_DARWIN branches (ADR-0418 libm) stay.
Skipped:
8f7d50d29: keep the fork's x86 DWT2 bound (0ed57f9f1, PR #1339). It already covers all six kernels. Upstream's form covers only the two s123 kernels and is one column more conservative.cba9343ed,c023bb7cb: already fixed bya013c1410(PR #1134) and6d61106ed(PR #1156).1801915be: checkasm CI workflow. Not applicable without the framework.86da14d03: CAMBI AVX2 dp/mask row. Handled by the separate CAMBI PR.
Tests: core/test/simd_bitexact_test.h gains guard-band helpers (simd_test_guard_fill, simd_test_guard_count_outside, SIMD_GUARD_ASSERT_UNTOUCHED). test_integer_adm_simd, test_adm_dwt2_x86, test_adm_dwt2_neon and test_vif_neon use them with the production band strides, and add small-size sweeps. test_integer_vif_avx2_stages and test_integer_vif_avx512_stages sweep heights. When porting upstream checkasm cases later, keep the guard bands: they are what catches out-of-region stores.
Canonical envtest installer (2026-09-08)¶
Keep Make, Go CI and controller-suite guidance on scripts/ci/setup-envtest.sh. The tool release and Kubernetes default live in build-config.env; preserve exact executable metadata checking, first-GOPATH/GOBIN selection, overrides and failure propagation. setup-envtest-env prints a shell-quoted export. Application Go dependencies, controller behavior and native/FFmpeg APIs are unchanged. See Research-2058.
Fedora Dockerfile parser repair (2026-09-08)¶
Preserve the optional SYCL branch's seven literal oneAPI configuration lines, GPG checks and write/install/cleanup short-circuiting in docker/dev/fedora-40.Dockerfile. Keep the printf form; the previous literal-backslash-n heredoc failed both Docker and Scorecard parsing. Base/SDK resolution, CUDA, public C/CLI and FFmpeg interfaces are unchanged. See Research-2056.
OpenSSF passing evidence and support correction (2026-09-08)¶
Keep VMAFx release support distinct from inherited libvmaf version strings. SECURITY.md gives current reporting/remediation policy; targets are not proof of historical compliance. Preserve the major-new-functionality test requirement in CONTRIBUTING.md. The passing worksheet tracks 67 official criteria at its dated source revision and project 14549; recheck criteria and public-master evidence before attesting. Unknown personal, private-history, crypto and release facts must remain unanswered until verified. No native/public API, numerical or FFmpeg rebase impact.
Repository security enforcement (2026-09-08)¶
Keep the canonical master policy and read-only checker together. REST omission of bypass actors requires a verified GraphQL count compared with the declared actor list, never an assumed match. Preserve strict checks, the GitHub Actions app binding, and the review requirement on the ruleset. Amended 2026-09-15 by ADR-1252: the policy now declares exactly one User bypass actor because a single-maintainer repository cannot satisfy its own independent-review requirement. Do not "restore" the zero-bypass assertion on a rebase without removing that actor from the live ruleset first, or the Scorecard master gate fails. No native or public C API rebase impact. See ADR-1248. Preserve the short purpose and participation links in docs/index.md; the website evidence uses GitHub Pages URLs with source links and a deployment check before new text is attested. Keep existing published topic-page evidence distinct from that pending improvement; the recorded external assessment is a dated in-progress snapshot, not permanent certification. No native/public API, numerical or FFmpeg rebase impact.
Scorecard exact-head gates (2026-09-08)¶
Preserve ADR-1247's separate PR-local and master-full scopes, immutable same-run artifact identity, before/after source binding, complete check sets, upstream risk weights and unrounded 8.5 floor. Keep publisher restrictions and its OIDC permission isolated from the gate jobs. The aggregator must require success from the applicable scope without waiting for the other event. Scanner errors are never exceptions; the exact no-release state is visibly unassessed. No C API, numerical baseline or FFmpeg surface changes.
Configured lint fixture bootstrap isolation (2026-09-08)¶
Preserve the real Make target and recursive build in test_lint_configured.py. Create its failing/recording pip prerequisite before fake Meson/Ninja, retain PIP_NO_INDEX=1, and assert no pip call, real venv or sentinel overwrite. Outer -o flags do not propagate to recursive Make. Production dependency rules, native sources and baselines stay unchanged. See Research-1246.
fix/observation-fixture-const — IQA/motion coverage inputs (2026-09-08)¶
Preserve the twelve const fixture inputs, writable IQA filter/kernel storage and exact assertion/API-call order. run_boundary_tests groups the first five IQA cases without adding a case count; its caller propagates the first failure. Production headers and numerical bodies are untouched. No public C API or FFmpeg rebase impact; see Research-2053.
fix/dnn-tests-native-lint-20260908 — ORT/session test inputs and copying¶
Keep the 16 read-only input/shape array qualifiers in the ORT and session API tests. All original 83 cases, 208 assertions and API call order remain intact; the session driver's private groups must return the first failure without counting helper groups. Keep the added POSIX read-error regression: a short fread must stop the copy loop, and ferror must make copy_file fail. This fixes test-fixture copying only; production APIs, model bytes and Netflix golden assertions are untouched. Preserve the two zero warning entries. See Research-2052. The copier must also check destination fclose after closing both streams. Keep its close-error regression before any ORT initialization, with limits/signals changed only in the child and a distinct setup-failure exit; retain the read-error case last.
fix/metric-coverage-const — preserve metric coverage setup (2026-09-08)¶
Keep the motion-v2, SSIM and PSNR coverage tests' const descriptor views and thirteen private setup helpers. Each caller immediately returns the original failure before its unchanged picture/extract/score/teardown path. Preserve all 133 assertions, 183 API calls, nineteen ordered registrations and literals. No production, public-header, test-registration or FFmpeg rebase impact. See Research-2054.
fix/fex-context-vector-20260908 — option-aware context identity¶
Keep the shared provided-feature base comparison from ADR-0385, followed by canonical keys derived from each context's parsed feature parameters. Base-only matching silently drops option-distinct motion/model registrations. Equivalent CPU/GPU twins, defaults and aliases still deduplicate with the first registration winning; absent provided-feature lists retain the name/options fallback.
Both comparison allocation failures and checked-growth failures return -ENOMEM without consuming the incoming context or changing existing vector storage. Preserve the runtime UINT_MAX and SIZE_MAX bounds, the portable registration and public-score controls, and Linux linker-wrapped allocation failure tests. The vector keeps its existing C-visible layout and manual pointer-array allocator; do not restore the obsolete prologue claiming a std::vector owns its storage. See Research-2047.
fix/svm-tests-native-lint — observation-only parser/API tests (2026-09-08)¶
Preserve the parser's header-size/header-order groups and exact case order, const model/query views, existing assertions and public SVM calls. The driver helpers return the first failure without incrementing the test count. Keep ADR-1138/ADR-1166 C NULL brackets. Vendored library bodies, headers, test registration and FFmpeg integration are unchanged. See Research-2049.
fix/cambi-avx2-native-lint — local reciprocal-table names (2026-09-08)¶
Keep reciprocals for the five local CAMBI AVX2 parameter bindings and retain reciprocal_lut at the outer global-table call sites. All function types, expressions, gather widths and test registrations are unchanged. No algorithm, header/API or FFmpeg rebase impact. See Research-2051.
fix/speed-test-native-lint-20260908 — preserve existing SpEED cases¶
Keep the read-only descriptor views in test_speed.c and test_speed_qa.c. The temporal QA fixture's allocation and extractor setup groups must return failures immediately to their original registered test. Preserve all 67 assertion expressions/messages, 10 registrations, input literals and API call order. Neither group is a new case. Keep the ADR-1138 C NULL brackets and the two measured zero warning entries. No production, API, FFmpeg or upstream algorithm change; see Research-2050.
docs/readme-entrypoint-20260908 — concise documentation entry points¶
Keep the README as a short introduction and guide index. Changing SDK pins, backend coverage and model defaults belong with their canonical topic docs, not repeated badge values, kernel counts or maturity tables. Root build commands use Meson's core/ source directory; the CPU guide explicitly disables optional GPU backends. Preserve the existing build-guide heading anchor for incoming links. Documentation-only correction; no native or FFmpeg surface impact.
refactor/cambi-production-lint-20260908 — read-only views and GPU helpers¶
Keep CAMBI's validation/preprocessing inputs and paired private scale-score wrapper declaration read-only. Preserve every formula, callback type and GPU trampoline, including the three documented scaffolds. Exact declaration annotations distinguish fixed callback types and out-of-profile/scaffold exports; unused checks elsewhere remain enabled. No public API or FFmpeg surface changes. See the equivalence receipt.
fix/adm-simd-native-lint — integer ADM local declarations (2026-09-08)¶
Keep AVX2/AVX-512 read-only band descriptors, threshold aliases and fixed arrays const. Preserve row-local SIMD accumulators and declaration-at-use scalar temporaries without changing expression order, shifts, clipping, p-norm behavior or LUT prefetch. Dispatch signatures remain unchanged; output band storage is writable. Existing cited function-size exceptions remain numerical invariants. No public C API or FFmpeg surface impact.
refactor/test-feature-extractor-lint-20260908 — read-only test views¶
Keep the twelve const qualifications in core/test/test_feature_extractor.c without changing its assertions, case order or fixture lifetime. The called production APIs already accept these read-only inputs. No production API, FFmpeg or numerical rebase impact; see the preservation receipt.
fix/vif-simd-native-lint — integer AVX2 stages (2026-09-08)¶
Preserve the private stages in vif_avx2.c while retaining the original integer tap/reduction order, per-scale shifts, packed mean additions and separate 8/16-bit variance lane layouts. Keep scalar tails, copy/padding and VifState callback signatures unchanged. The new native stage test compares real scalar/AVX2 results and vertical workspace bytes. Preserve the system <stdio.h> include in integer_adm.h; it changes no numeric declarations. See Research-2045 and core/src/feature/x86/AGENTS.md. No public/FFmpeg impact.
fix/cppcheck-c-header-model-20260908 — official pthread type model (2026-09-08)¶
Keep write_cppcheck_posix_model.py in both configured local lint and the required Cppcheck workflow, and load its generated path rather than bare --library=posix. The full installed model supplies pthread's C aggregate types without forcing a platform or language. Versions lacking pthread_cond_init receive its correct contract; the newer defective model loses only argument 2's non-null marker because POSIX permits default attributes as NULL. Preserve actual-header positive/negative controls, install-relative fallback, analyzer validation, atomic publication, diagnostic categories and compile-database variants. Do not replace the correction with call-site suppressions, native header constructors, or a vendored full model. Fork-only analyzer wiring; no libvmaf/FFmpeg API impact.
fix/tensor-io-test-cleanup-20260908 — preserve tensor test coverage (2026-09-08)¶
Keep read-only tensor fixtures const and grouped test drivers below the existing function-size limit. Preserve the explicit unsupported dtype/resize values and individual cited analyzer markers; those calls protect rejection behavior, not an accidental cast. All numerical assertions, test order and one execution per real case remain unchanged. The test compiles real tensor I/O with DNN disabled. See core/test/dnn/AGENTS.md. Fork-only test cleanup; no public API or FFmpeg impact.
fix/rocm-2604-restore-20260908 — released vendor image¶
Keep ROCM_BUILDER and ROCM_RUNTIME on the digest-pinned Ubuntu 26.04 ROCm 10 image selected by ADR-1231. Regenerate Dockerfile mirrors from build-config.env; do not restore the ROCm 24.04 exemptions from the old rollback. The pruned SDK stage must compile/link a HIP kernel and check its host loader, not just report a version. Preserve the node runtime library layout. See verification.
fix/roi-reader-bounds-20260908 — ROI input boundaries (2026-09-08)¶
Keep vmaf_roi_input.h shared by the CLI and its boundary test: validate depth and extent locally, saturate rounded luma before narrowing to 8 bits, and traverse the placeholder using its validated allocation count. Existing CLI dimensions, rounding below saturation, radial arithmetic and encoder sidecar byte layouts remain unchanged. Preserve the ADR-1138 C NULL brackets and the cited single-threaded getopt invariant. Fork-only CLI implementation; no public libvmaf or FFmpeg filter surface changes.
fix/convolution-horizontal-boundary (2026-09-08)¶
Preserve the output-based horizontal split in both common AVX kernels: j_vec_end is the first final scalar output; SIMD source starts stop at j_vec_end - radius. The masked final load/store keeps original AVX2 mul/add and AVX-512 FMA regions without discarded out-of-plane accesses. Keep clamped tiny-width borders and the common horizontal pass per ISA. The regression uses configured private objects, runtime ISA checks and tight final rows. No public API or FFmpeg integration surface changes.
fix/vif-native-lint — scalar VIF decomposition (2026-09-08)¶
Preserve the ten-plane workspace layout, filter/decimation/statistic order, float-to-double promotions, scale reductions and temporal first-frame values in core/src/feature/vif.c. The own-header declarations retain all three legacy external symbols; the C NULL bracket follows ADR-1138. Debug dumps write initialized planes and scalar numerator/denominator outputs. Keep the odd-stride and temporal EOF/error lifecycle test registered. No public C API or FFmpeg surface changes; the scalar arithmetic remains Netflix-compatible.
docs/upstream-pr-links — what the fork has open upstream (2026-09-19)¶
docs/development/known-upstream-bugs.md lists the eight pull requests this fork has open against Netflix/vmaf and the six upstream defects it found and did not report. The next sync should read it before resolving conflicts in integer_adm.c, adm_avx2.c, adm_avx512.c or output.c: if an upstream PR landed, the incoming side may already carry the fork's fix, and in two cases it carries a different fix. Upstream #1601 starts the 16-bit DWT sum from the normalization offset rather than widening the accumulator to int64 as the fork's
1477 does, because the int64 form costs 3.5 to 6 % of throughput upstream. Do¶
not resolve that conflict by keeping both. Upstream #1494, by a maintainer, refactors the same ADM functions and will force a rebase either way.
fix/svm-cppcheck-219 — libsvm is split, and stays split (2026-09-19)¶
core/src/svm.cpp is no longer close to upstream libsvm's layout. The allocation macro goes through svm_checked_malloc, and 14 oversized blocks are split into named helpers: the solver, both working-set selections, the trainer, the sigmoid trainer, the model parser and two class bodies. A re-vendor that drops a fresh libsvm in will undo all of it, exactly as the 1.7.19 cJSON re-vendor undid that file's fixes. Re-apply the delta rather than replacing the file, and check with the two gates that found this: cppcheck 2.19 (which comes from the ubuntu-26.04 runner image, not from a pin) and the size limit, which applies to every block in a file the pull request touches. Every split preserved its expressions in order; the check that it stayed correct is the Netflix pair byte-identical at --precision max, which is the first thing to re-run.
ci/ubuntu-2604-matrix-tail — the lanes Renovate could not see (2026-09-19)¶
Renovate's github-actions manager rewrites runs-on: values only. A lane that names its image as a matrix os: key, or inside an expression, is invisible to it: build.yml's Linux Intel LLVM row, libvmaf-build-matrix.yml's Ubuntu ARM clang row and the Cppcheck job's ARC_RUNNERS_ENABLED fallback all stayed on 24.04 after the bump. When the next image generation arrives, grep for ubuntu- rather than trusting the bot's diff. .github/actionlint.yaml carries ubuntu-26.04 and ubuntu-26.04-arm because actionlint 1.7.12 does not know them; drop an entry once it ships the label.
renovate/ubuntu-26.04 — job artefacts leave /tmp (2026-09-19)¶
The hosted runner label moved from ubuntu-24.04 to ubuntu-26.04 across 21 workflows. On that image /tmp is a RAM-backed tmpfs with a per-user quota, so anything large must live under RUNNER_TEMP: the Tiny AI and MCP virtualenvs, the ONNX Runtime archives and their cache directory, and the Kubernetes end-to-end workflow's buildx layer cache and three-image tar. Do not move them back while resolving a conflict; the failure is [Errno 122] Disk quota exceeded, not a full disk. In run: blocks use ${RUNNER_TEMP}; in action inputs use ${{ runner.temp }}, because with: has no shell. The contract test scripts/ci/test_e2e_runtime_contract.py enforces exactly that split.
fix/bug-mypy-pyver — mypy models the required Python (2026-09-21)¶
[tool.mypy] python_version in pyproject.toml is 3.14 and tracks requires-python; take 3.14 on any conflict and never resolve back towards 3.10. Below 3.12 mypy refuses to parse the PEP 695 type statement in numpy's bundled __init__.pyi, and that blocking [syntax] error aborts the entire ai/src/ pass before a single source file is checked. The comment above the value carries that reason; keep it with the value. The 60 findings the raise makes visible are pre-existing debt — measured against a merge base carrying the same 3.14, the change introduces none — so do not absorb a conflict here by adding type: ignore, widening ignore_missing_imports, or relaxing strict. [tool.black] and [tool.ruff] target-version are deliberately left at their older values in this change. The pairing is now enforced rather than commented: check_mypy_python_version in scripts/ci/check-workflow-versions.py fails the always-run pre-commit gate when python_version, the requires-python floor and PYTHON_CI_VERSION stop agreeing, or when python_version is deleted instead of reverted. A rebase that moves requires-python must move python_version in the same commit, and scripts/ci/tests/test_mypy_python_version_single_source.py (discovered by the test-base-image-single-source hook's test_*single_source.py pattern) is the fixture that says so. CI's own Python Lint job is untouched by any of this: it installs only mypy, so its output is byte-identical at 3.10 and 3.14 and it checks zero source files either way. Fork-only tooling; no native API or FFmpeg patch impact. See ADR-1282.
fix/pre-push-mypy-delta — introduced findings only (2026-09-19)¶
scripts/git-hooks/pre-push-mypy.py keeps the ai//scripts/ merge-base scope from the note below, and adds two invariants (ADR-1261). Paths under ai/src/ run in their own mypy invocation with --explicit-package-bases; that directory is a mypy_path base and without the flag mypy refuses the file for having two module names, which blocked every push touching it. The same files are re-checked at the merge base in a disposable worktree and only new findings fail, because CI's mypy is advisory and inherited findings vary with the checkout's installed stub packages. Do not restore the raw exit-status propagation: an unattributable non-zero exit now fails closed with its own message, which the regression suite pins. python_version moved to 3.14 in its own change (ADR-1282, fix/bug-mypy-pyver); take 3.14 on any conflict here. Fork-only tooling; no native API or FFmpeg patch impact.
fix/pre-push-mypy-scope — merge-base ownership (2026-09-08)¶
Preserve the existing ai//scripts/ Python touched-file policy in scripts/git-hooks/pre-push-mypy.py. Check every branch-owned path against the master merge base, including type changes; do not restore the remote old-tip/new-tip intersection. The always-run, filename-free hook invocation, lexical symlink identity, safe target validation and outgoing-HEAD check are paired with a real Git rebase regression. Fork-only tooling; no native API or FFmpeg patch impact. This fixes implementation of AGENTS.md §12.10.
fix/git-fixture-environment-isolation — fixture caller safety (2026-09-08)¶
Keep inherited GIT_* and caller global/system configuration out of the FFmpeg replay/smoke, dependency-classifier, Level Zero and agent-cleanup test fixtures, including setup and assertions. git -C alone can still mutate a caller's config, refs, object store or index. Preserve the disposable caller matrix in scripts/ci/test_git_fixture_isolation.py, its actual linked-worktree hook control, and its local/CI registration. This implements the existing ADR-1240 isolation contract; no production Git operation or native/FFmpeg API changes.
fix/configured-lint-driver — native lint selection (2026-09-08)¶
Local Make lint regenerates Meson metadata without option overrides before building, then retains its native database and every configured tracked source/command variant. Keep engine roots, C++ tools, tests and tracked vendors included, inactive backends explicitly outside the profile, and numeric GCC LTO adaptation confined to the private analyzer database. Preserve the scratch-Git/Make fixture and pre-commit registration. Existing ADR-1142 ratchet baselines and remote lane policies remain authoritative; no libvmaf/FFmpeg surface changes.
FFmpeg stable-release patch maintenance (2026-09-08)¶
Preserve build-config.env as the FFmpeg remote/tag owner, the ordered ffmpeg-patches/series.txt, and FFmpeg Patch Stack in the required aggregator. Regenerate with scripts/ci/ffmpeg_patch_stack.py --refresh; checks fetch the reviewed tag, while scheduled --latest refreshes select stable releases only. Canonical patch mail metadata changes without changing the applied source tree. Local hooks use disposable Git state and fail closed on network/replay errors. See ADR-1240.
fix/agent-cleanup-preserve-work-20260908 — preserve agent state (2026-09-08)¶
cleanup-agent-state.sh is fork-owned. Preserve ADR-1239's report-only default, explicit worktree selections, non-force removal, file-state guards and stash retention when rebasing developer tooling. Branch existence never proves a stash is redundant. Regression: bash scripts/dev/test-cleanup-agent-state.sh.
fix/docs-guide-references — developer/usage links and push tools (2026-09-08)¶
Preserve canonical ADR IDs alongside link slugs, backend overview paths, and the historical-but-unpublished T7-3 notebook note. The MkDocs push hook selects documentation before checking availability; missing MkDocs blocks selected docs, while direct non-doc invocation still skips. The real-Git fixture tests both cases. No libvmaf/FFmpeg API impact.
fix/generated-adr-freshness — generated metadata (2026-09-08)¶
Regenerate with make docs-fragments-write after combining ADR fragments. Tags must precede navigation. Preserve required Docs/local freshness checks, fragment coverage validation, and generator fixtures (ADR-1242). Accepted ADR bodies and tag taxonomy are unchanged; only mutable fragments and rendered outputs are refreshed.
fix/worktree-hook-dispatch — local hook lifetime (2026-09-08)¶
Preserve regular dispatchers, actual Git argument/stdin forwarding, and independent MkDocs/PR-body push checks (ADR-1241). Keep the disposable lifecycle fixture wired into required Pre-Commit CI. Fork-only tooling; no Netflix C API or FFmpeg patch impact.
fix/base-image-unpinned-reference-guard — Level Zero config consumer (2026-09-08)¶
The development SDK stage reads LEVEL_ZERO_VERSION from its copied build-config.env during the download RUN; no Docker ARG mirror is needed. Keep both URL fields and the Renovate manager attached to that owner. The single-source regression test executes the actual command with stubs. ROCm continues through the image manager; do not restore the obsolete literal workflow manager. No upstream rebase impact: these container and CI surfaces are fork-local.
fix/base-image-unpinned-reference-guard — Renovate file selection (2026-09-08)¶
Custom-manager file patterns use one /regex/ delimiter pair. Preserve positive tracked-file fixtures and the base-manager coverage of every Dockerfile where the built-in manager is disabled. The fixture iterates active managers, so removing an obsolete manager does not reintroduce or require it. No upstream rebase impact: Renovate configuration and these tests are fork-local.
fix/base-image-unpinned-reference-guard — container reference guard (2026-09-08)¶
The ADR-1231 scanner must reject direct external FROM/COPY references regardless of digest presence. Preserve the Python instruction scanner, the exact local consumer exceptions and the fixture suite in scripts/ci/tests/. Shared image ARG defaults remain one per physical line for the shell mirror writer. No upstream rebase impact: these guard scripts and fixtures are fork-local.
fix/renovate-draft-automerge-deadlock — dependency bumps can merge again (2026-09-15)¶
renovate.json gains "draftPR": false in seven places: the six packageRules that carry "automerge": true and the vulnerabilityAlerts block. The global "draftPR": true at the bottom of the file stays.
Rebase-sensitive in two ways:
- Do not "simplify" this by removing the global
draftPR. The per-rule overrides exist precisely so that review-needed bumps keep opening as drafts, which is what PR #1411 added the global flag for. Dropping the global and keeping the overrides inverts the behaviour and re-floods the ready queue. - Keep the overrides paired with
automerge. A rule that gains"automerge": truelater must gain"draftPR": falsewith it, or it deadlocks again: ADR-0679 makes the required aggregator fail on drafts, so a draft can never satisfy branch protection and automerge can never fire. If a rule losesautomerge, thedraftPRoverride should go with it.
PR #1416 also edits dependency declarations (single-sourcing versions); it does not touch renovate.json's packageRules, so the two should not collide. If a conflict does appear here, keep both sides: they are independent keys.
fix/windows-cuda-cl-fallback (2026-09-08)¶
In core/src/meson.build, the PATH fallback must assign cl_path from cl_exe.full_path() before forming nvcc_ccbin_flags. The next PowerShell command derives MSVC includes from cl_path; assigning only the flags leaves an undefined variable when vswhere fails. Netflix PR #1472 at b7b65e64 has the same defect. Preserve the assignment when porting that discovery block.
core/test/test_windows_cuda_compiler_discovery.py executes the current block through Meson using Windows host metadata and stubbed tool responses. It covers successful discovery, empty/error fallback and missing cl, and is registered in fast on POSIX build hosts. This is configure coverage, not a Windows GPU test.
Scoped lint-baseline tightening (ADR-1243)¶
tidy-ratchet.py --only --write updates only successfully measured source entries, preserves all unselected headers/TUs and full-report metadata, and rejects increased allowance. Preserve its failure and atomic-write tests when rebasing the CI tools. A scoped report must never replace the full baseline; the required whole-tree lane and diagnostic-only --only semantics remain.
fix/thread-pool-queue-bound — Netflix queue-capacity fix (2026-09-08)¶
Adapt Netflix 8fc71e3006f0b21e8e31d6e5d1b904332149ad9e from libvmaf/src/thread_pool.c to core/src/thread_pool.c. Keep the fork's inline payload/free-list recycling, two-argument callback, worker private-data cleanup, checked primitive initialization and error OR/reset. Capacity follows the successfully created worker count. Destruction must also wait for admitted producers to leave capacity waits; the upstream broadcast by itself does not protect their lifetime. The isolated pthread-injection test gates these paths.
fix/merge-train-ownership-guard — local control boundary (2026-09-08)¶
The fork-local gateway scripts/dev/merge_train_guard.py and its pre-commit regression hook must move together. Preserve ADR-1244's shared hold/base/owner guard on every mutation, rebase-before-ready ordering, exact-head leases, non-force cleanup, and actual full-gate receipt generation. Do not restore an unrestricted local agent operator alongside it. Runtime migration remains an explicit operator step documented in docs/development/merge-train.md.
fix/renovate-draft-automerge-deadlock — dependency bumps can merge again (2026-09-15)¶
renovate.json gains "draftPR": false in seven places: the six packageRules that carry "automerge": true and the vulnerabilityAlerts block. The global "draftPR": true at the bottom of the file stays.
Rebase-sensitive in two ways:
- Do not "simplify" this by removing the global
draftPR. The per-rule overrides exist precisely so that review-needed bumps keep opening as drafts, which is what PR #1411 added the global flag for. Dropping the global and keeping the overrides inverts the behaviour and re-floods the ready queue. - Keep the overrides paired with
automerge. A rule that gains"automerge": truelater must gain"draftPR": falsewith it, or it deadlocks again: ADR-0679 makes the required aggregator fail on drafts, so a draft can never satisfy branch protection and automerge can never fire. If a rule losesautomerge, thedraftPRoverride should go with it.
PR #1416 also edits dependency declarations (single-sourcing versions); it does not touch renovate.json's packageRules, so the two should not collide. If a conflict does appear here, keep both sides: they are independent keys.
port/upstream-2026-08-sync — August upstream reconciliation (2026-09-15)¶
Three things a future upstream sync must know about this range.
core/src/ext/x86/x86inc.asmnow carries upstream's CET block. The.note.gnu.property/GNU_PROPERTY_X86_FEATURE_1_SHSTKstanza was taken verbatim from upstreamf85a85369, so the next sync should resolve this file in upstream's favour rather than re-applying ours. The file is otherwise a vendored dav1d/x264 header and must stay that way. Note that meson does not track it as a dependency ofcpuid.asm: after touching it, delete the object or the change is silently ignored by an incremental build.- The AVX-512 targets no longer pass
-mavx512vbmi. This matches upstreameb1045795, so the fourc_argslists incore/src/meson.buildconverge with upstream instead of diverging. Do not reintroduce the flag: no source in the tree uses a VBMI intrinsic, and requiring it excludes Skylake-SP and Cascade Lake. If a future kernel does use one, give that kernel its own target. - Do NOT take upstream's
arm64/motion_neon.c8-bit pipeline. Upstream3137d5525addsmotion_score_pipeline_8_neoninarm64/motion_neon.h. This fork already exports that symbol fromarm64/motion_v2_neon.c, alongside a 16-bit twin upstream does not have, andinteger_motion.cdispatches both. Taking upstream's file duplicates the symbol and would lose the 16-bit path. Its test was ported instead ascore/test/test_motion_pipeline_neon.c, which asserts both pipelines are bit-exact against the scalar reference; keep that test when resolving, and keep ours registered inside the arm64 guard.
chore/praetor-governance-adoption — AGENTS.md is now a compiled source (2026-09-15, refreshed 2026-09-18)¶
The canonical AGENTS.md is compiled by praetorctl compile-context into six vendor context files under a hard 300-line-per-target budget: CLAUDE.md, .cursor/rules/hiss-invariants.mdc, .github/copilot-instructions.md, .windsurfrules, .gemini/GEMINI.md and .codex/rules.md. Consequences for any rebase or upstream-sync agent:
- Never resolve a conflict by editing a vendor context file. They are generated byte-for-byte from
AGENTS.mdplus a two-line header. Resolve inAGENTS.md, then runmake compile-context; the requiredStandards & Invariant Verification Gaterunscompile-context --verifyand fails on any hand edit. AGENTS.mdis linted as agent text. The pinned engine failscompile-context --verifyandauditwhenAGENTS.mdreads as prose. Resolve a conflict in the internal register, then runpraetorctl caveman check AGENTS.md; to prove a rewrite dropped nothing, runpraetorctl caveman floor <old> AGENTS.md.- Adding to
AGENTS.mdcan break the build. Every target must stay at or below 300 lines; the largest sits at 263. New long-form rules belong in adocs/development/page that the harness imports, which is why §12 Hard rules and §13 Rebase-sensitive invariants live atdocs/development/agent-hard-rules.mdanddocs/development/rebase-sensitive-invariants.md. - The project-wide invariant index moved. An upstream-sync agent that used to read
AGENTS.md§13 must now readdocs/development/rebase-sensitive-invariants.md; the per-subtreeAGENTS.mdfiles it points at are unchanged. - Personas have one editable copy. Each
.mdpersona under.claude/agents/,.codex/agents/,.github/agents/and.gemini/agents/is a projection of the same name under.agents/agents/. Resolve there and recompile; the gate rejects a projection that differs from, or has no source in,.agents/agents/. .standards-baseline.jsonis append-only downward. A rebase that reintroduces agoto, an unchecked error or an over-long function fails the audit against the baseline rather than the compiler. Regenerate withpraetorctl baseline --recordonly when debt genuinely decreased, from a clean clone and with the engine the CI gate pins.
fix/ai-1270-blockers — DISTS, MobileSal, and predictor stub triage (2026-09-08)¶
Triage and point-of-use guards for issue #1270 blockers:
-
core/src/feature/feature_dists.clogs a point-of-use warning when loading placeholder checkpoint.model/tiny/dists_sq.onnxis a 3-op synthetic MSE smoke placeholder (vmaf_tiny_dists_sq_placeholder_v0), not the learned Ding et al. multi-scale backbone. Point-of-useVMAF_LOG_LEVEL_WARNINGadded todists_sq_initwhen loadingdists_sq.onnxor any placeholder graph. Do not silence this warning until real learned DISTS weights and the feature extractor stack are implemented. Tracked indocs/state.mdunderT-DISTS-PLACEHOLDER-CHECKPOINT-2026-09-08. -
docs/ai/models/mobilesal.mdandcore/src/feature/feature_mobilesal.cupdated for production saliency.mobilesal.onnxis the legacy smoke placeholder (vmaf_tiny_mobilesal_placeholder_v0). Production saliency usesmodel/tiny/saliency_student_v2.onnx(ADR-0444) orsaliency_student_v1.onnx(ADR-0286).mobilesal_initwarns atVMAF_LOG_LEVEL_WARNINGwhen loadingmobilesal.onnx, recommendingsaliency_student_v2.onnx. -
Software and AMF predictor models are synthetic stubs, guarded at point of use.
Predictor(tools/vmaf-tune/src/vmaftune/predictor.py), CLIvmaf-tune predict(tools/vmaf-tune/src/vmaftune/cli.py), and Go session (pkg/predictor/ortsession.go) detect and warn on synthetic-stub predictor models (synthetic-stub-N=100, ADR-0325). Real-corpus models exist only for NVENC and QSV. Tracked indocs/state.mdunderT-PREDICTOR-SOFTWARE-AMF-STUB-MODELS-2026-09-08.
refactor/1241-modernization-finish — JSON model parser twin lockstep and rebrand finish (2026-09-08)¶
core/src/read_json_model.c,core/src/read_json_model.cpp: C and C++23 parser twins brought into full lockstep. Invariant: keepread_json_model.c(compiled intofuzz_json_modelbycore/test/fuzz/meson.build) andread_json_model.cpp(compiled intoread_json_model_cpp23_libforlibvmaf) in lockstep on any parser change.read_json_model.cgains the ADR-1060 defect #5 stream-error check (if (json_get_error(s)) return -EINVAL;after the unknown-key skip loop inmodel_parse).read_json_model.cppgains ADR-0887 cross-key per-feature length mismatch validation (sync_n_featuresin all walkers,validate_feature_arraysinparse_model_dict).tools/vmaf-roi-score/README.md,tools/vmaf-tune/README.md,docs/research/README.md: Residual "lusoris vmaf fork" product names updated to "VMAFx fork".
perf/1245-benchmark-tuning-pass — CAMBI AVX2 anti-dithering, SpEED SIMD QR, and HIP threaded pipeline (2026-09-08)¶
core/src/feature/cambi.c:anti_dithering_filter()now checksVMAF_X86_CPU_FLAG_AVX2and dispatches toanti_dithering_filter_avx2()incore/src/feature/x86/cambi_avx2.c. The fallback remains the upstream scalar loop. When rebasing against upstream Netflix, preserve the AVX2 check and dispatch branch.core/src/feature/x86/cambi_avx2.candcambi_avx2.h: fork-added AVX2 implementation ofanti_dithering_filter_avx2(). Must preserve 32-bit zero-extended accumulation and lane permute (0xD8) to ensure bit-exact parity with scalar arithmetic.core/src/feature/speed_internal.c:si_mat_mul()now dispatches throughspeed_matmul_avx512/speed_matmul_avx2/speed_matmul_scalarper ADR-1237, lifting the ADR-1196 scalar hold.core/src/libvmaf.c:batch_extractor_skip(),read_pictures_should_skip(), andflush_non_temporal_cpu_extractors()includeVMAF_FEATURE_EXTRACTOR_HIPandVMAF_FEATURE_EXTRACTOR_METALin their GPU extractor sets.flush_context_threaded()drainsgpu_pendingfor non-CUDA/SYCL extractors before temporal flushes, matchingflush_context_serial(). Preserve this alignment so HIP works under--threads N.testdata/bench_all.sh: prefers/opt/intel/oneapi/setvars.shover legacy 2025.3 paths, and Test 2 targetscheckerboard_1920_1080_10_3_0_0.yuvandcheckerboard_1920_1080_10_3_1_0.yuv.
feat/1242-tiny-ai-completion — tiny-AI model cards and int8 fallback test coverage (2026-09-08)¶
core/test/dnn/test_dnn_session_api.c,core/test/dnn/test_vmaf_use_tiny_model.c: fork-added C unit tests verifying the ADR-1032 second fallback trigger (int8 session creation failure retrying fp32 baseline without leak).docs/ai/models/: added missing model cardssmoke_multi_output_v0.mdandsmoke_v0_symbolic_batch.md; brought existing cards into compliance with ADR-0042.no rebase impact: fork-added tests and docs; upstream Netflix/vmaf has no dnn test or docs/ai tree.
fix/venv-gate-basename-false-positive — tracked-venv gate pattern (2026-09-05)¶
No rebase impact: scripts/ci/check-no-tracked-venv.sh and its test are fork-added.
perf/hot-path-1245 — SpEED matrix_mul gains a kernel-pointer parameter (2026-09-06)¶
core/src/feature/speed.c: upstream Netflix'smatrix_mul(Matrix *dst, const Matrix *x, const Matrix *y)is nowmatrix_mul(..., speed_matmul_fn matmul), and the pointer is threaded throughmatrix_qr_decomposition()andsolve_linear_system()fromSpeedState::matmul(ADR-1196). The multiply loop itself moved out ofmatrix_mulinto the exportedspeed_matmul_scalar()in the same file. An upstream commit that touches any of those three signatures will conflict: re-thread the parameter rather than reverting to the three-argument form, and keepspeed_matmul_scalarnon-static—core/test/test_speed_simd.clinks against it directly so the SIMD twins are gated against the production reference rather than a copy of it.core/src/feature/x86/speed_matmul_avx2.c,x86/speed_matmul_avx512.c: fork-added. Invariant: both must stay in their own-ffp-contract=offstatic libraries (x86_speed_matmul_avx2/x86_speed_matmul_avx512incore/src/meson.build). Folding them intox86_avx2_sources/x86_avx512_sourcesputs them under-mfmawith contraction enabled, the compiler fuses the explicit_mm*_mul_ps/_mm*_add_pspairs into a single-rounding FMA, and the scores drift from the scalar reference. Same carve-out rationale asx86_ssim_avx2andx86_float_adm_avx2.core/src/feature/speed_internal.c: not touched. Itssi_mat_mul()is the ADR-0964 duplicate of the same loop for the GPU twins' host side and stays scalar on purpose; do not "resync" it tospeed.cas if the difference were drift.
fix/gpu-cambi-parity-drift — align the CUDA/SYCL CAMBI twins with cambi.c (2026-09-05)¶
core/src/feature/cuda/integer_cambi/cambi_score.cu,core/src/feature/sycl/integer_cambi_sycl.cpp: fork-added GPU CAMBI kernels with no upstream Netflix counterpart. Invariant: these kernels mirrorcambi.c's host-side semantics exactly, not "something reasonable". Two places where the natural GPU idiom is wrong and must not be "simplified" back: (1)cambi_spatial_mask_kernel/launch_spatial_maskmust contribute zero for taps outside the image when accumulating the 7x7 zero-derivative box sum —cambi.c's summed-area table zero-pads (compute_dp_rowwithactual_width = 0); clamping to the edge pixel inflates the sum, because edge pixels arezero_derivative = 1by construction. (2) the verticalfilter_modepass must return early fory == 0andy == height - 1:cambi.c::filter_modewrites only output rows1 .. height-2and leaves both border rows at their pre-filter values. Both twins write the V pass back into the buffer the H pass read from, so the early return preserves exactly those pre-filter pixels.core/src/feature/cuda/integer_cambi_cuda.c:cambi_high_res_speedupis no longer "reserved". It must be resolved against the encode pixel count ininit_fex_cuda(mirroringcambi.c:622-640), halve the adjusted window incambi_cuda_adjust_window(cambi.c:471) and trigger one extra decimation before scale 0 insubmit_fex_cuda(cambi.c:1621). The default modelvmaf_v1.0.16_3d0hsetshrs=1080, so dropping any of the three silently changes every>= 1080pCUDA score.core/test/test_cuda_cambi_parity.c,core/test/test_sycl_cambi_parity.c: the second ("textured") fixture is a regression gate, not decoration — the original quantised-gradient fixture is flat along every border and down every column and cannot observe either defect. Keep both fixtures on rebase.core/src/feature/cambi.cis unchanged by this branch (Netflix golden gate).
fix/cambi-cuda-context — CUDA CAMBI context push/pop and model options twin selection gate (2026-09-05)¶
core/src/feature/cuda/integer_cambi_cuda.c: fork-added CUDA CAMBI extractor. Invariant: every device-touching entry point (init_fex_cuda,submit_fex_cuda,close_fex_cuda) must pushfex->cu_state->ctxupon entry and cleanly pop it on all exit paths (balancedfail_after_poplabels). Its option table and TVI initialization must mirrorcambi.c.core/src/feature/feature_extractor.cpp/core/src/libvmaf.c: option validation and GPU twin gating (ADR-1183).vmaf_fex_ctx_parse_optionsrejects unknown option keys with-EINVAL.vmaf_use_features_from_modelchecks GPU twin option support against model requirements and dispatches unsupported twins to the CPU reference. Preserve this gating on rebase to prevent silent option drops.
fix/t-upstream-1494-adm-csf-mode-irfactor-ov — integer-ADM CSF representability guard (2026-09-06)¶
-
core/src/feature/x86/adm_avx2.c,core/src/feature/x86/adm_avx512.c: the horizontal DWT2 tail bound ishalf_w >= 2 ? half_w - 1 - ((half_w - 2) % N) : 1(withhalf_w = (w + 1) / 2) in all six kernels —adm_dwt2_8_avx2,adm_dwt2_16_avx2,adm_dwt2_s123_combined_avx2,adm_dwt2_8_avx512,adm_dwt2_16_avx512andadm_dwt2_s123_combined_avx512. Invariant: the last columnhalf_w - 1must always fall to the scalar tail loop, which is the only one that applies theind_xmirror; the upstream boundhalf_w - ((half_w - 1) % N)gave that column to the vector loop wheneverhalf_w % N == 1. This is the x86 twin of the NEON fix (T-ADM-DWT2-NEON-PARITY-2026-08-30) and the residual (3) of T-UPSTREAM-1564-ADM-CM-GPU-BORDER-AND-ROUNDING-2026-09-03. Upstream Netflix still carries the unguarded bound, so on any upstream sync touching x86 DWT2 keep the guarded expression andcore/test/test_adm_dwt2_x86.c, which pins bit-exactness against the scalar kernel at w in {34, 66, 130, 258} (half_w % N == 1) plus w=576. -
core/src/feature/adm_csf_fixed_point.his fork-added. It owns the fixed-point exponents (2^21 / 2^23 at scale 0, 2^32 at scales 1-3), the storage bounds, the tabulated-fast-path predicate, and the scale-0 narrowing conversion. Invariant: the conversion must staydouble-valued ((double)float_weight * pow(2, N)), because that is exactly what the four(uint16_t)(rfactor1[k] * pow2_N)expressions it replaced evaluated —float * doublepromotes todouble. Changing it to afloatproduct would move scores. core/src/feature/integer_adm.c:adm_csf_rfactor_scale0()now delegates to that header, andadm_csf_config_check()(called once frominit(), cached inAdmState::csf_config_err, returned byextract()) refuses weights the storage cannot hold. On a rebase, keep the verdict inextract()beside the pre-existingnvd * rdh >= 3240guard:core/test/test_adm_coverage.c ::test_adm_invalid_view_dist_returns_einvalpins that an unsupported ADM configuration initialises and then fails at extract time. See ADR-1191.core/src/feature/cuda/integer_adm_cuda.c,core/src/feature/hip/integer_adm_hip.c,core/src/feature/sycl/integer_adm_sycl.cpp: each keeps its own copy ofadm_csf_factors()/adm_csf_rfactor_scale0()(upstream parity, unchanged) and gains anadm_csf_config_check()called from its owninit(). Invariant: all four backends must apply the CPU bounds from the shared header, even where the twin's own storage is wider — the SYCL twin holds scale 0 inuint32_t, but accepting a configuration the CPU rejects would break the ADR-1183 option / feature-name parity contract.core/src/feature/x86/adm_avx2.candadm_avx512.cdeliberately keep their own byte-identical copies of the scale-0 conversion. They are unreachable with an out-of-range weight now thatinit()gates the configuration, and leaving them untouched keeps the SIMD bit-exactness story unchanged. Do not "unify" them into the shared header without re-runningcore/test/test_integer_adm_simd.c.
build/union-merge-append-only-docs — union merge for the bookkeeping files (2026-09-06)¶
No rebase impact on upstream code: .gitattributes is fork-owned. Note the mechanic: a merge driver is read from the merge base, so this only helps branches whose base already contains the attribute — every branch open when it lands still conflicts once more, then stops. docs/state.md is excluded on purpose (moved rows must not be duplicated).
fix/t-upstream-818-pooling-enum-no-percentil — percentile temporal pooling (2026-09-06)¶
core/include/libvmaf/libvmaf.h: upstream-mirrored public header. The fork appendsVMAF_POOL_METHOD_MEDIAN/_PERC5/_PERC10/_PERC20afterVMAF_POOL_METHOD_HARMONIC_MEANand definesVMAF_HAVE_PERCENTILE_POOLING(ADR-1188). Invariant: the growth is append-only — if upstream ever adds its own enumerator, append it after the fork's four rather than renumbering, and never reorder the first five. On a sync that touches this enum, re-check thestatic_assert(VMAF_POOL_METHOD_NB == 9)incore/src/output.cppand the value assertions incore/test/test_pool_percentile.c.core/src/libvmaf.c:pool_reduce()keeps the upstream accumulator arithmetic; the fork addspool_accumulate(),PoolSamples,pool_samples_push()andpool_reduce_percentile()around it. Invariant: percentiles must never be derived from the accumulators, and the accumulator methods must never be derived from the sorted buffer (ADR-1118 golden-gate isolation). Porting an upstream change tovmaf_feature_score_pooledmeans porting it intopool_accumulate()'s loop body, which is where the upstream loop now lives verbatim.core/src/percentile.h: fork-added, header-onlystatic inline.core/src/predict.clost its file-staticscore_compare/percentileto this header (identical expressions). Invariant: keep themstatic inlinein the header — moving them into a.cputs the golden-asserted bootstrapci_p95arithmetic behind a cross-TU call.core/src/output.cpp: writers iterate the fork-addedpool_report_order[]instead of[1, VMAF_POOL_METHOD_NB), sopooled_metricskeeps emitting exactlymin,max,mean,harmonic_mean. Invariant: an upstream diff that reintroduces theNB-bounded loop would silently widen the report schema; keep the explicit table.ffmpeg-patches/0018-libvmaf-map-percentile-pool-methods.patch: new tail patch againstn9.0.1, extends FFmpeg's stockpool_method_map(and mapsmax, which upstream FFmpeg never did). Guarded by#ifdef VMAF_HAVE_PERCENTILE_POOLING, so it still builds against a Netflix libvmaf. Verified: all 18 patches replay onto pristinen9.0.1viagit am --3way.
fix/t-upstream-766-cli-option-string-delimit — escape-aware --model / --feature splitting (2026-09-06)¶
core/tools/cli_parse.cpp: upstream Netflix carries this file (ascli_parse.c) and still splits the option strings withstrsep. The fork replaced all nine split sites withcli_split()/cli_unescape()(ADR-1190) and deleted thevmaf_cli_strsepshim together with the#ifndef HAVE_STRSEPfork. Invariants to preserve on a sync: splitting and unescaping are two passes (cli_splitmust leave backslashes in place so an escape written for the:pass survives into the=pass, andcli_unescapemust run exactly once, on a token that will not be split again — unescaping twice would eat a user's literal backslash); a key/value pair's value is the whole remainder after the first unescaped=, never a secondstrsep(that second split is the silent-truncation bug); and the model-overload key must be split on.before it is unescaped. If an upstream commit reintroduces astrsepcall here, port its intent ontocli_split, do not restore the call.core/test/test_cli_parse.c: the eightT-UPSTREAM-766cases are fork-added and are registered throughrun_model_delimiter_tests/run_feature_delimiter_testsrather than a single runner, because more than about sevenmu_run_testexpansions tripreadability-function-size(ADR-0141). Keep the per-return NULLcitedNOLINTNEXTLINE(modernize-use-nullptr)markers (ADR-1138) — this TU must keep spelling the null pointer constantNULLfor the MSVC C lane.pkg/cliopt: fork-added, no upstream counterpart. Invariant:EscapeValueandcli_unescape()are one grammar in two languages — a change to the C escape set must change the Go escaper (and its round-trip test) in the same commit.ffmpeg-patches/: deliberately unchanged. Thelibvmaf_tunefilter'sload_model()takes the whole remainder after the first=and never splits on:, and ffmpeg's own filtergraph parser owns escaping at that layer, so teaching it the CLI grammar would double-unescape.
ci/release-artifacts-built-in-dev-container — native release artifacts built on self-hosted canonical runner (ADR-1178) (2026-09-05)¶
No rebase impact: all touched files (.github/actionlint.yaml, .github/workflows/dev-container-publish.yml, .github/workflows/supply-chain.yml, scripts/release/verify-native-release-artifacts.sh, scripts/release/tests/test-verify-native-release-artifacts.sh, scripts/ci/check-container-build.sh, scripts/ci/tests/test-check-container-build.sh, docs) are fork-local CI workflows, verification scripts, and documentation with no upstream Netflix/vmaf counterpart. No public C API, header, Meson option, or golden assertion is touched.
fix/tidy-lane-lto-flag — the tidy lane builds without LTO (2026-09-05)¶
No rebase impact: .github/workflows/lint-and-format.yml is fork-added. Invariant: any build whose only purpose is to emit compile_commands.json for clang-tidy must configure with -Db_lto=false, because the project default (b_lto_threads=4, ADR-1172) renders as a GCC-only -flto=<n> that clang rejects outright.
fix/state-md-duplicate-rows — one row per bug id (2026-09-05)¶
No rebase impact: docs/state.md, scripts/ci/check-state-md-rows.sh and its test are fork-added. The gate exists because of rebases: resolving a docs/state.md conflict by keeping both sides is the documented shortcut for the append-only sections, and it silently duplicates a row that a PR was moving between sections. After any such resolution, run bash scripts/ci/check-state-md-rows.sh.
A rebase can also drop the move hunk while keeping the status edit, which leaves one copy of the row under ## Open bugs with closed or fixed in its own status cell — no duplicate, and every id/row check passes. The gate now compares each row's status token against the section heading it sits under, so that resolution fails too. The status cell is the column a header calls Status, or the last non-empty cell, and only the word that opens it is read; the repair is to move the row, never to rewrite the status to match where the rebase left it.
fix/state-md-move-tombstone — reject contradictory move bookkeeping (2026-09-25)¶
No rebase impact: the row-hygiene script, fixture, documentation, and docs/state.md are fork-local. Preserve the fourth gate alongside the existing duplicate and status checks: a moved to Recently closed tombstone under Open bugs may not coexist with a table row carrying the same id in that section. During a state conflict, keep the tombstone plus the authoritative row under Recently closed; never retain the stale Open row merely because it has no parseable Status cell.
fix/sycl-motion2-checkerboard-drift — clip integer_motion2 score to motion_max_val (2026-09-05)¶
no rebase impact: fork-local SYCL feature extractor and tests.
core/src/feature/sycl/integer_motion_sycl.cpp: wholly fork-added (upstream Netflix/vmaf has no SYCL backend). Collector calls now appendmotion2_clipped(lines 841, 848) andlast_motion2inflush_fex_sycl(line 896), matching CPU reference behavior whenmotion_max_valis set.core/test/test_sycl_motion3_parity.c: fork-added test; added 1080p checkerboard test case verifyinginteger_motion2_mmxv_18andinteger_motion3_mmxv_18clipping and CPU/SYCL parity.core/test/test_sycl_motion_add_uv_parity.c: historically adjusted the tolerance to 2e-4 for 3-plane fixed-point integer motion; ADR-1326 later replaced that empirical comparison with a fixed-point oracle.python/test/sycl_motion_parity_test.py: fork-added Python parity tests for checkerboard and src01 pairs.
fix/cambi-cuda-context — CUDA CAMBI context push/pop and model options twin selection gate (2026-09-05)¶
core/src/feature/cuda/integer_cambi_cuda.c: fork-added CUDA CAMBI extractor. Invariant: every device-touching entry point (init_fex_cuda,submit_fex_cuda,close_fex_cuda) must pushfex->cu_state->ctxupon entry and cleanly pop it on all exit paths (balancedfail_after_poplabels). Its option table and TVI initialization must mirrorcambi.c.core/src/feature/feature_extractor.cpp/core/src/libvmaf.c: option validation and GPU twin gating (ADR-1183).vmaf_fex_ctx_parse_optionsrejects unknown option keys with-EINVAL.vmaf_use_features_from_modelchecks GPU twin option support against model requirements and dispatches unsupported twins to the CPU reference. Preserve this gating on rebase to prevent silent option drops.core/test/test_feature_extractor.c: upstream Netflix carries this file with a flatrun_tests()and a handful of cases; the fork's copy is now mostly fork-added regression tests, split one-behaviour-per-function and registered throughrun_registry_tests/run_context_tests/run_option_testsso thereadability-function-sizebranch budget (ADR-0141) holds. On a sync, add any new upstream case to the matching group runner rather than re-flatteningrun_tests(), and keep the file-scopedNOLINTBEGIN(modernize-use-nullptr)bracket (ADR-1138) — the TU must keep spelling the null pointer constantNULLfor the required MSVC C lane.core/src/feature/feature_extractor.cpp: this TU is C++, not the C twin, so ADR-1138'sNULL-for-MSVC exemption does not apply — the file now spells the null pointer constantnullptrthroughout and carries noNOLINT. Its file-local symbols (feature_extractor_list[],vmaf_fex_ctx_parse_options,check_pic_buf_type,find_fex_list_entry/grow_fex_list/init_fex_list_slot/get_fex_list_entry,ctx_pool_ensure_slot_ctx,ctx_pool_claim_slot) live in anonymous namespaces rather than beingstatic(misc-use-anonymous-namespace). When porting an upstream change tovmaf_feature_extractor_context_extract, note the picture buf_type/backend validation now lives incheck_pic_buf_type()and the pool-slot registration is split acrossfind_fex_list_entry/grow_fex_list/init_fex_list_slot, so the two entry points stay inside thereadability-function-sizebudget (ADR-0141); re-inlining them re-opens the warning.grow_fex_listuses a hardif (pool->capacity == 0) return -EINVAL;guard rather thanassert(), because clang-tidy 22'smisc-static-assert/cert-dcl03-cflags everyassert()whose condition contains no non-constexpr call. Fail-closed teardown invariant (ADR-1336 follow-up): CUDA contexts publishclose_requiredbefore feature init;context_destroy()returns-EBUSYwhile either that partial-init obligation or an initialized close obligation remains. Worker-private data,vmaf_fex_ctx_pool, andRegisteredFeatureExtractorsuse a close-only prepare followed by a destroy-only commit. A failed prepare retains the owner container and publicVmafContextfor retry. Preserve prepare order after worker drain: worker-private contexts, pooled contexts, registered contexts, then the CUDA drain stream. Do not move close back into a void free callback or free a vector/pool after a child close fails. GPU picture pools likewise remember committed slots, and CUDA state/function tables are cleared only after release succeeds. Exact zero is the only public close commit; normalize positive pthread-style errno values and retain the context after every nonzero result. Apply the same bounded retry and dependency retention to embedded callers, includingcore/src/mcp/compute_vmaf.c; never destroy its model first. Preserve the CUDA state's internal imported marker: unimported state-free performs retry-safe runtime release, imported wrapper free remains allocation-only after exact-zero close, and duplicate imports return-EBUSYrather than aliasing or overwriting live ownership. Non-CUDA failed-init callbacks remain outsideclose_requireduntil separately audited (see ADR-1336 §Follow-up).core/src/feature/feature_extractor.h: upstream Netflix header. Keeps the upstream__VMAF_FEATURE_EXTRACTOR_H__include guard,<stdint.h>/<stdlib.h>, plain Ctypedef structand untyped flag enums, because roughly a hundred C translation units include it. clang-tidy has no compile command for a header and falls back tofeature_extractor.cpp's, so it analyses the file as C++ and proposes C++-only rewrites; a single file-scopedNOLINTBEGIN(...)/NOLINTEND(...)bracket after the licence block and after the closing#endifsuppresses them. Two constraints on that bracket: theADR-NNNNcitations must sit inside theNOLINTBEGINmarker's own block comment, not in a separate comment above it —scripts/ci/tidy-ratchet.py::count_uncited_nolintsscans only the current line, its neighbours and the marker's own comment, and counts anything else as an uncited NOLINT, which fails the ADR-1142 ratchet; and the justification text must not contain the character pair that closes a block comment (an earlier draft wrote acore/src/feature/wildcard glob, silently truncated the comment mid-file, and turned the header into parse errors). On a sync, keep the bracket balanced and do not "modernise" the header — every rewrite it suppresses breaks the C includers.
feat/dnn-int8-redirect-and-sidecar-fixes — dnn_attach_api.c helper split (2026-09-05)¶
core/src/dnn/dnn_attach_api.c: the int8 redirect is nowresolve_quantised_load_path(); the sidecar load and the post-open attach areload_optional_sidecar()/attach_opened_session().vmaf_use_tiny_model()is back under thereadability-function-sizethreshold and no longer tripsbugprone-redundant-branch-conditionorclang-analyzer-deadcode.DeadStores.- All three helpers live inside the
#if VMAF_HAVE_DNNguard, so a-Denable_dnn=disabledbuild is byte-for-byte the ADR-0374 stub it was before. no rebase impact: fork-added DNN loader TU with no upstream Netflix counterpart.
feat/dnn-int8-redirect-and-sidecar-fixes — docs/ai gap closeout for #1242 (2026-09-05)¶
docs/ai/sidecar-online-training.md,docs/ai/extractor-template.md,docs/ai/inference.md: replaced three "planned" placeholders with the audited state of the tree (sidecar checkpoint quarantine,transnet_v2sliding window, self-hosted GPU runner). OpenedT-GPU-RUNNER-LABEL-MISMATCH-2026-09-05indocs/state.md.no rebase impact: fork-added documentation only; upstream Netflix/vmaf has no docs/ai tree.
feat/dnn-int8-redirect-and-sidecar-fixes — qat_train.py rank-4 image loader (2026-09-05)¶
ai/scripts/qat_train.py:_build_train_loader_factorynow dispatches on the rank ofqat.input_shape; new_build_image_loader_factoryreads an NCHW.npzfor rank-4 models._config_input_rankis the single place the rank is read.ai/tests/test_qat_train_loader.py,docs/ai/quantization.md: coverage + docs.no rebase impact: fork-added tiny-AI training script; upstream Netflix/vmaf has no ai/ tree.
feat/dnn-int8-redirect-and-sidecar-fixes — measure_quant_drop.py path overrides (2026-09-05)¶
ai/scripts/measure_quant_drop.py: added--fp32/--int8/--budget/--idso a pair of ONNX files can be gated without amodel/tiny/registry.jsonentry. Registry-driven--alland positional forms are untouched.ai/tests/test_measure_quant_drop_unit.py,docs/ai/quantization.md: coverage + docs.no rebase impact: fork-added tiny-AI training script; upstream Netflix/vmaf has no ai/ tree.
feat/dnn-int8-redirect-and-sidecar-fixes — declare onnx_has_scaler in vmaf_tiny_v3.int8.json (2026-09-05)¶
model/tiny/vmaf_tiny_v3.int8.json: added"onnx_has_scaler": true, so the C runtime stops normalising the canonical-6 vector that the graph already scales.core/test/dnn/test_registry.sh,python/test/model_registry_schema_test.py,ai/scripts/validate_model_registry.py: added a consistency check asserting that anymodel/tiny/*.int8.onnxbaking scaler ops (Sub/Div) has a companion sidecar declaring"onnx_has_scaler": true. Detection prefers theonnxparser and falls back to a protobuf byte scan on legs without it.no rebase impact: fork-added tiny-AI model artifact and fork-added validation tests; upstream Netflix/vmaf ships no ONNX model registry.
feat/dnn-int8-redirect-and-sidecar-fixes — wire int8 redirect and fp32 fallback into vmaf_use_tiny_model (2026-09-05)¶
core/src/dnn/dnn_attach_api.c: wired.int8.onnxredirect and ADR-1032 debug fallback intovmaf_use_tiny_model()when sidecarquant_mode != VMAF_QUANT_FP32.core/test/dnn/test_vmaf_use_tiny_model.c: unit tests for int8 redirect, fp32 fallback, and missing external data error path.core/src/dnn/dnn_api.c+core/src/dnn/dnn_attach_api.c: both twins retry the fp32 baseline once whenvmaf_ort_open()fails on an int8 graph that already cleared the size cap and the op allowlist. Without it--tiny-model model/tiny/nr_metric_v1.onnxregressed to-EIOon any ONNX Runtime build lacking aConvIntegerkernel, whichcore/test/dnn/test_cli.shcatches. The retry must stay in both twins — the invariant note incore/src/dnn/AGENTS.mdpins that.no rebase impact: fork-added tiny-AI DNN loader surface with no upstream Netflix counterpart.
chore/modernization-leftovers-1241 — redundant cpp_std override_options cleanup (2026-09-05)¶
core/src/meson.build,core/test/meson.build,core/tools/meson.build: upstream-mirror files (Netflix/vmaflibvmaf/src|test|tools/meson.build, renamed under ADR-0700). The hunks removed here are fork-only: thelibvmaf_cpu_cpp_stdtoken block and everyoverride_options : ['cpp_std=...']on the fork's isolated*_cpp23_lib/metadata_handler_cpp20_libstatic libs,libvmaf_cpu_static_lib, thevmaf/vmafxtools and thetest_cli_parse*/test_picture_pool_cpp_error_pathstargets. Upstream has no C++ TUs there, so a sync cannot re-introduce them; if a future upstream hunk lands next to one of these targets, do not resurrect a per-targetcpp_stdoverride — the standard is project-wide viaadd_project_argumentsincore/meson.build(ADR-1003 / ADR-1056). Theb_lto=falseoverrides incore/src/meson.build(AVX-512) andcore/test/meson.build(test_output_lto_override, macOS) are intentional and must survive.core/test/fuzz/meson.build,core/AGENTS.md,docs/development/cpp23-extractor-pattern.md: fork-added; no rebase impact.
chore/modernization-leftovers-1241 — rebrand residual scrub (2026-09-05)¶
no rebase impact: every touched file is fork-added (.claude/, ai/scripts/, dev-llm/, mcp-server/, docs/development/, the root and python/ pyproject.toml, CLAUDE.md); upstream Netflix/vmaf has no counterpart for any of them. Deliberately NOT renamed — keep them on any sync: the libvmaf.so soname, the libvmaf ffmpeg filter name, the version scheme, the lusoris.* ONNX metadata keys asserted by ai/tests/test_export_u2netp_mirror.py, and model/tiny/transnet_v2.onnx (its baked-in producer_name only changes at the next re-export).
fix/metal-cambi-hrs-option — option-table sync for the Metal cambi twin (2026-09-05)¶
core/src/feature/metal/integer_cambi_metal.mm: fork-added (upstream Netflix has no Metal backend); no upstream sync conflict. Preserves the option table parity requirement:cambi_high_res_speedup(aliashrs, int, default 0, min 0, max 2160) must be retained so model dispatch using default modelvmaf_v1.0.16_3d0hselects the Metal twin rather than falling back to CPU. Also preserves the decimation and window size adjustments for resolutions >= 1080p. The >= 1080p / 1440p / 2160p pixel-count thresholds are taken from the sharedCAMBI_HIGH_RES_SPEEDUP_THRESHOLD_*macros incore/src/feature/cambi_internal.h— do not reintroduce a Metal-local copy, that is exactly how the twins drift apart.core/test/test_metal_integer_cambi_parity.c: fork-added unit test. Asserts option table registration and parity between CPU and Metal extractors withcambi_high_res_speedup. Rebase-sensitive invariant:vmaf_use_feature()takes ownership of theVmafFeatureDictionaryon every path except the argument-validation guards, so each runner must build its own dictionary — handing one dictionary to both backends is a use-after-free.
fix/cuda-drain-batch-per-state-lifetime — drain-batch ownership and the read fence (2026-09-05)¶
core/src/cuda/drain_batch.{c,h} and core/test/test_cuda_drain_batch.c are fork-added (ADR-0242); core/src/libvmaf.c is an upstream-mirror file and the two hunks there are fork-only: the fence_for_read() helper above vmaf_feature_score_at_index() and its two call sites. On rebase, keep the invariant that vmaf_cuda_drain_batch_open() takes the owning VmafCudaState — an upstream signature without the owner reintroduces the cross-context bleed.
feat/otel-init-all-go-binaries — finish the OpenTelemetry rollout across every Go binary (ADR-0782, ADR-1119; epic #1241) (2026-09-05)¶
no rebase impact: every touched file is fork-added — the Go tree (internal/app/bootstrap/, internal/oteltest/, cmd/vmafx-*, pkg/observability/otel_instruments.go, pkg/ai/infer.go, go.mod/go.sum) and fork-added docs; upstream Netflix/vmaf has no Go code.
internal/app/bootstrap/bootstrap.go:Basegainsfx.Decorate(withServiceIdentity)(service.version frompkg/version,OTEL_SERVICE_NAMEhonoured behindVMAFX_OTEL_SERVICE_NAME) and the package gainsHTTPTracing/TraceHTTPHandler(otelhttp,<METHOD> <path>, probes filtered). Keep the decorator shape — it must not replace golusoris's ownotel.Optionsprovider.cmd/vmafx-server/main.go,cmd/vmafx-controller/main.go:bootstrap.HTTPTracingnext togolusoris.HTTP; the server'sapp_test.go::productionGraphmirrors it.cmd/vmafx-operator/internal/controller/vmafxjob_controller.go:grpc.DialContext→grpcmod.NewConnFactory().Dial(otelgrpc client handler; drops thestaticchecknolint).cmd/vmafx-mcp/tools.go::addRawTool:vmafx.mcp.toolspan;main.gowraps the HTTP transport handler inbootstrap.TraceHTTPHandleroutermost.cmd/vmafx-tune/cmd/golusoris.go:deps.OTel,vmafx.tune.commandspan ended beforeapp.Stop.pkg/ai/infer.go:Inferhas named results and avmafx.onnx.inferencespan;vmafx-ort-runnerintentionally untouched (ADR-1134 exemption).pkg/observability/otel_instruments.go: additive constantsSpanMCPTool,SpanTuneCommand,AttrMCPTool,AttrTuneCommand;InitOTelkept (ADR-0927) but documented as unused by binaries.go.mod:otelhttppromoted from indirect to direct; no version change.- Docs:
docs/development/observability.mdrewritten for the golusoris reality (OTLP/gRPC :4317,VMAFX_OTEL_*keys, sample ratio 1.0, per-binary span table);docs/observability/otel.mdrefreshed;cmd/AGENTS.mdandinternal/app/bootstrap/AGENTS.mdadded.
fix/gpu-init-leaks-and-hip-mirror — fix CUDA init error path leaks and close verified GPU state issues (2026-09-04)¶
no rebase impact: fork-local CUDA and documentation files.
core/src/feature/cuda/speed_chroma_cuda.c,core/src/feature/cuda/speed_temporal_cuda.c,core/src/feature/cuda/integer_ms_ssim_cuda.c,core/src/feature/cuda/integer_psnr_hvs_cuda.c: wholly fork-added CUDA feature extractors (upstream Netflix/vmaf does not have GPU SpEED, and upstream CUDA extractors lack these fork-specific teardown paths). No upstream rebase conflicts.docs/state.md: closed out resolved audit tasksT-HIP-MOTION-V2-MIRROR-OFF-BY-ONE-2026-06-13,T-SYCL-INIT-LEAKS-EXC-2026-06-19,T-SPEED-GPU-REGISTRY-ORPHAN-2026-06-19, andT-CUDA-INIT-SUBMIT-LEAKS-2026-06-19.
build/bound-lto-link-parallelism — b_lto_threads=4 default (2026-09-04)¶
Rebase impact: core/meson.build default_options is a fork-edited hunk of an upstream file (libvmaf/meson.build upstream has no b_lto); on rebase keep both b_lto=true and b_lto_threads=4 together — the second exists only because the first is on (ADR-1172).
ci/release-please-gate-warning — idle-green release-please without the App (2026-09-04)¶
No rebase impact: .github/workflows/release-please.yml, scripts/release/, ADR-1171 and docs/development/release.md are fork-added with no upstream counterpart. Invariant kept from ADR-1151: no step ever authenticates release-please with GITHUB_TOKEN; the creds step gates every write step, and its severity is warning on push, error on workflow_dispatch.
ci/shorten-job-names — shorten CI job display names and gate aggregator list (2026-09-04)¶
no rebase impact: changes GitHub Actions workflow job display names (.github/workflows/*.yml), branch-protection aggregator list (required-aggregator.yml), adds verification gate (scripts/ci/check-aggregator-names.sh), and updates fork-added documentation (docs/development/ci-job-names.md, docs/development/ci.md, docs/development/release.md, AGENTS.md, CLAUDE.md, .github/AGENTS.md, scripts/ci/AGENTS.md). Upstream Netflix/vmaf uses a completely different CI setup; preserve fork-local workflow files on any sync.
fix/sycl-adm-shift-reachability — verify non-negative shift reachability in integer_adm_sycl.cpp (2026-09-04)¶
core/src/feature/sycl/integer_adm_sycl.cpp: added invariant documentation comments at both normalization shift sites (launch_decouple_csfandlaunch_adm_cm_line). Proved thatks = 17 - clz >= 1is an algebraic invariant guaranteed by the enclosingabs_oh >= 32768(2^15) guard, matching the unclamped structure of CPUget_best15_from32and CUDAadm_decouple_inline.cuh. No code changes or numeric divergence.no rebase impact: SYCL integer ADM is a fork-added backend with no upstream Netflix counterpart.
fix/sycl-ssimulacra2-blur-recurrence-and-arc-calibration — revert pseudo-Kahan recurrence and calibrate Arc A380 (2026-09-05)¶
core/src/feature/sycl/ssimulacra2_sycl.cpp: wholly fork-added (upstream Netflix has no SYCL backend); no upstream sync conflict. Invariant (ADR-0985): the Charalampidis 3-pole autoregressive IIR blur recurrence ($o_k = n2 \cdot \text{sum} - d1 \cdot \text{prev1} - \text{prev2}$) has no running accumulator; do NOT add pseudo-Kahan recurrence or modify the output equation, as any additive feedback shifts the poles outside the unit circle and causes geometric divergence ($> 10^{25}$ / NaN / saturation at 100.0). Must remain bit-exact with the CUDA twinssimulacra2_blur.cu.scripts/ci/gpu_ulp_calibration.yaml: fork-added calibration database.sycl:0x8086:0x56a*andarc:dg2-g10entries are calibrated to5.0e-2(places=1) based on hardware measurement on Intel Arc A380 over 48-framesrc01sequences, capturing fp64-less accumulation across the 6-scale pyramid.docs/adr/0985-sycl-parity-divergence-2026-06-03.md: ADR-0985 marked Accepted with Option C decision matrix.docs/research/0985-sycl-parity-divergence-2026-06-03.md: Research-0985 updated with mathematical derivation of the pseudo-Kahan pole instability and empirical hardware measurements.
fix/cli-metal-define — define HAVE_METAL for the CLI so its Metal paths are compiled at all (2026-09-04)¶
core/tools/meson.build: added Metal branch (if is_metal_enabled) that appends-DHAVE_METAL=1tovmaf_tool_cflagsandmetal_depstovmaf_tool_deps.core/tools/meson.buildis fork-modified; upstream has no Metal backend. On upstream sync, keep this branch intact so that the CLI continues to compile Metal translation units on macOS.
docs/venv-recipe — replace impossible venv recipe with verified one (2026-09-04)¶
docs/development/languages.md: no rebase impact: docs/development/ is fork-added.
fix/vmaf-tune-python-fast-path — Python fast-path probe decoding, feature parsing, and normalisation parity (2026-09-05)¶
No rebase impact: all touched files (tools/vmaf-tune/src/vmaftune/, tools/vmaf-tune/tests/, docs/) are fork-added Python tuning tooling with no upstream Netflix/vmaf counterpart. No public C API, header, Meson option, or golden assertion is touched. - tools/vmaf-tune/src/vmaftune/cli.py & fast.py: probe distorted containers (.mp4) are decoded to temporary raw YUV before running libvmaf feature extraction (with guaranteed cleanup) and non-zero exit codes raise RuntimeError (no zero-fill). - tools/vmaf-tune/src/vmaftune/proxy.py: added load_proxy_sidecar and normalise_features adhering to fr_regressor_v2.json StandardScaler parameters, aligned ENCODER_VOCAB_V2 ordering, and mapped unrecognized encoders to "unknown" (slot 11) when allow_unknown=True. - tools/vmaf-tune/src/vmaftune/score.py: parse_feature_aggregates handles integer_* keys and falls back to per-frame averages when pooled metrics are absent.
fix/ai-ptq-static-pin-qdq — pin ONNX Runtime static-PTQ output format to QDQ (2026-09-05)¶
no rebase impact: fork-only ai/ script
fix/security-cleanup-1243 — widen integer index operands in convolve, moment, psnr (2026-09-04)¶
core/src/feature/iqa/convolve.c: upstream-mirror file. Widen(ptrdiff_t)y * dst_w + xindst[...]vertical pass while preservingfloat * floatsingle-rounded arithmetic for SIMD bit-exactness contract (ADR-0138). On upstream sync, preserve the widening.core/src/feature/moment.c: upstream-mirror file. Widen(ptrdiff_t)i * stride_ + jincompute_1st_momentandcompute_2nd_momentwhile preservingpic_ * pic_float multiplication for SIMD bit-exactness contract (ADR-0179). On upstream sync, preserve the widening.core/src/feature/psnr.c: upstream-mirror file. Widen(ptrdiff_t)i * ref_stride_ + jand(ptrdiff_t)i * dis_stride_ + jincompute_psnrwhile preservingdiff * difffloat multiplication for SIMD bit-exactness. On upstream sync, preserve the widening.ai/scripts/extract_ugc_features.py,ai/tests/test_extract_ugc_features.py: wholly fork-added tooling and test files with no upstream Netflix/vmaf counterpart. No rebase impact.core/tools/spinner.h: upstream-mirror header. Added#ifndef VMAF_SPINNER_H/#define VMAF_SPINNER_Hheader guard. On upstream sync, preserve header guards.core/tools/cli_parse.cpp: fork-added C++ translation unit (replacingcli_parse.c, ADR-0809 / ADR-1155). Added non-variadicusage(app, reason)overload alongside variadic template. No upstream counterpart.core/test/test_model_feature_overload_ownership.c: fork-added test file. Rephrased comment text to avoid CodeQL commented-out code heuristic. No upstream counterpart.core/src/pdjson.c: vendored third-party parser (pdjson). Removed redundant lower bound comparisons in UTF-8 sequence length validation. On upstream sync, preserve bounds cleanup.mcp-server/vmaf-mcp/tests/test_parity_argv.py: fork-added test file. Removed unusedpytestimport. No rebase impact.osv-scanner.toml: fork-added configuration file ignoringGO-2026-5932for unimportedopenpgpsubpackage. No rebase impact.
fix/sycl-adm-tidy-debt — SYCL ADM warning cleanup + tidy-lane scoping (2026-09-04)¶
core/src/feature/sycl/integer_adm_sycl.cpp: wholly fork-added (upstream Netflix has no SYCL backend); no sync conflict. Two things to preserve: the designated initialisers must stay in struct declaration order (ISO C++ requires it; MSVC rejects the reverse), andks = 17 - clzat lines ~705 and ~1032 must NOT be clamped without the CPU-parity analysis tracked indocs/state.md(T-SYCL-ADM-NEGATIVE-SHIFT-REACHABILITY-2026-09-04)..github/workflows/lint-and-format.yml,scripts/ci/clang-tidy-sycl.sh,scripts/ci/gen-sycl-compile-commands.py: fork-added. The SYCL tidy lane deliberately builds onlyinclude/vcs_version.hbefore analysis. Restoring a fullmeson compilethere re-creates the scoping bug where any TU's compiler warning fails the lane regardless of the PR's diff.
feat/ai-teacher-single-source — AI teacher model follows default model single source (ADR-1173) (2026-09-04)¶
ai/: feature extractors (extract_full_features.py,extract_k150k_features.py,bvi_dvc_to_full_features.py,extract_ugc_features.py,konvid_to_full_features.py,konvid_to_vmaf_pairs.py,bvi_dvc_to_corpus_jsonl.py) and scoring helpers (scores.py) now dynamically resolve their teacher model fromai.data.scores.resolve_teacher_model()(backingvmaftune.defaultmodel.DEFAULT_MODELper ADR-1168) instead of hardcodingvmaf_v0.6.1. They stampteacher_modelon every row and manifest. Upstream syncs touching these scripts should preserveresolve_teacher_model()and the row-levelteacher_modelcolumn.ai/data/feature_extractor.pyandai/scripts/extract_k150k_features.py: raw feature extraction lists append"adm3"toFULL_FEATURESandFEATURE_NAMES. The canonical-6 student features (DEFAULT_FEATURES) remain frozen.ai/scripts/combine_full_feature_parquets.py,ai/scripts/train_vmaf_tiny_v5.py,ai/scripts/eval_loso_vmaf_tiny_v5.py: enforce intra-table and cross-table teacher model uniformity, refusing mixed-model datasets and unprovenanced tables without--assume-teacher <name>.scripts/ci/check-default-model-single-source.sh: removed wholesale^ai/exemption from the gate'sallow_re.ai/data/netflix_loader.py(load_or_compute(..., cache_valid=)) andai/train/dataset.py: the per-clip$VMAF_TINY_AI_CACHEentry is revalidated against the resolved teacher; a stale or unstamped entry is a cache miss. Keep the predicate when touching the loader — dropping it silently relabels pre-ADR-1173vmaf_v0.6.1caches as the current teacher.
docs/state-sweep-four-closed-rows — docs/state.md bookkeeping sweep (2026-09-04)¶
no rebase impact: docs-only
fix/hip-motion-v2-parity-test-wiring — register test_hip_motion_v2_parity in meson.build (2026-09-04)¶
no rebase impact: fork-only test wiring in core/test/meson.build and documentation updates in core/src/feature/hip/AGENTS.md, docs/adr/1154-hip-backend-gaps.md, and docs/state.md. Upstream Netflix/vmaf has no HIP backend or HIP parity test suite.
fix/sycl-v1-model-crash — Intel Arc SYCL default model crashes and feature parity (2026-09-05)¶
core/src/feature/cambi.c,core/src/feature/cambi_internal.h:vmaf_cambi_init_tvi_and_vlt()exposed with extern "C" linkage so GPU and CPU cambi extractors share table initialization logic. Upstream sync should preserve this helper.core/src/feature/sycl/integer_cambi_sycl.cpp: Addedcambi_high_res_speedup(aliashrs) tooptions_cambi_syclto maintain feature-name parity with CPU CAMBI under modelvmaf_v1.0.16_3d0h. Sized histogram buffer toMAX(num_bins, v_band_size).core/src/feature/sycl/speed_chroma_sycl.cpp,core/src/feature/sycl/speed_temporal_sycl.cpp: Replaceddoubleaccumulators and workgroup local accessors withfloatto satisfy ADR-0220 on fp64-less Intel Arc devices.core/src/meson.build: Passed_x86_simd_strict_fp_extra(-fp-model=precise) tox86_avx2_static_libandx86_avx512_static_libwhen compiling withicx.python/test/sycl_default_model_test.py: Wholly fork-added regression test gating--backend sycldefault model execution. No upstream rebase conflict.
ci/sycl-arc-self-hosted-runner — containerised self-hosted GitHub Actions runner for Intel Arc SYCL CI (ADR-1177) (2026-09-04)¶
dev/Containerfile.runner: fork-added; derives fromvmaf-dev-mcp:localwith GitHub Actions runner v2.337.0 and non-rootrunneruser (uid 1001). Preserves all oneAPI SYCL tools and Level-Zero runtime. No upstream counterpart.dev/docker-compose.runner.yml: fork-added compose file passing through only the Intel Arc A380 render node (${ARC_RENDER_NODE:-/dev/dri/renderD129}, by-pathpci-0000:03:00.0-render) with NVIDIA and AMD device isolation,seccomp=unconfined, and 8 CPU / 16 GB limits.dev/scripts/runner-entrypoint.sh: fork-added runner entrypoint handling token configuration and ephemeral execution..github/workflows/sycl-parity.yml: fork-added workflow; runs on[self-hosted, linux, x64, sycl-arc]. Strictly prohibits execution on untrusted forks (github.event.pull_request.head.repo.full_name == github.repository)..github/workflows/required-aggregator.yml: addedSYCL Parity (Arc A380)to therequiredarray; the check is switched byvars.SYCL_ARC_RUNNER_ENABLED(absence/skip accepted while disabled; skip = loud failure while enabled). No runner API call in the aggregator.scripts/ci/check-runner-available.sh+scripts/ci/tests/test-runner-available.sh: fork-added hosted probe (lane switch + online check viasecrets.SYCL_RUNNER_PROBE_TOKEN; API errors fail loudly).dev/scripts/arc-render-node.sh: fork-added; resolves the single Intel render node forARC_RENDER_NODE.scripts/ci/gpu_ulp_calibration.yaml: added calibratedfloat_ssim: 5.0e-4entry for Arc A380sycl:0x8086:0x56a*.core/test/meson.build: tagged all 23 SYCL tests withsuite : ['fast', 'gpu', 'sycl']. Upstream sync conflict resolution: preserve thesuiteadditions on any upstream test additions.- Rebase impact: minimal. Upstream Netflix/vmaf has no SYCL backend, no self-hosted runner infrastructure, and no
required-aggregator.yml. If upstream touchescore/test/meson.build, keep the fork's SYCL test declarations and suite tags.
fix/metal-motion-v2-mirror-closeout — Metal motion_v2 mirror closeout and test observability (2026-09-04)¶
no rebase impact: fork-only Metal backend (core/src/feature/metal/integer_motion_v2.metal, core/test/test_metal_motion_v2_parity.c, core/src/feature/metal/AGENTS.md, ADR-1176). All touched files are fork-added surfaces with no upstream Netflix/vmaf counterpart.
fix/vmaf-tune-v1-canonical-features — request VIF explicitly so canonical-6 columns populate under v1 default (2026-09-04)¶
No rebase impact: all touched files (tools/vmaf-tune/, pkg/corpus/, pkg/fast/) are fork-added Python and Go tuning tooling with no upstream Netflix/vmaf counterpart. No public C API, public header, Meson option, or golden assertion is touched.
fix/vmaf-tune-report-audit-and-svtav1-hdr-knob-docs — vmaf-tune report audit findings #2–#10 and SVT-AV1-HDR knob docs (2026-09-04)¶
No rebase impact: fork-only tools/vmaf-tune and documentation surfaces (tools/vmaf-tune/, docs/usage/vmaf-tune.md, docs/usage/vmaf-tune-codec-adapters.md). No upstream Netflix/vmaf counterpart, no C engine files, and no Netflix golden test assertions touched.
chore/drop-ansnr — remove ansnr feature extractor (ADR-0865) (2026-09-04)¶
- Cleanly finalized removal of the legacy ANSNR feature extractor (ADR-0865).
- Removed residual dead configuration entries in CI parity tooling:
scripts/ci/cross_backend_parity_gate.py(float_ansnrmetric tuple and tolerance entry),scripts/ci/cross_backend_vif_diff.py(float_ansnrtuple), andscripts/ci/gpu_ulp_calibration.yaml(float_ansnrULP entry). - Recorded feature deprecation row in
docs/development/deprecations.md. - Added rebase-sensitive invariant in
core/src/feature/AGENTS.md. - Rebase impact: Upstream Netflix/vmaf still carries
ansnr/float_ansnrin its C tree (libvmaf/src/feature/ansnr.c,libvmaf/src/feature/ansnr.h,libvmaf/src/feature/ansnr_options.h,libvmaf/src/feature/ansnr_tools.c,libvmaf/src/feature/ansnr_tools.h,libvmaf/src/feature/float_ansnr.c, and x86/arm64 SIMD pathsansnr_avx2.c,ansnr_avx512.c,ansnr_neon.c). On rebase or upstream sync, re-drop any restored ansnr files and do not allowansnrregistrations back intofeature_extractor.cpp.
ci/flaky-legs-1236 — unblock UDS listener accept on stop and resilient macOS Homebrew (2026-09-04)¶
core/src/mcp/mcp.c:stop_uds()now invokesshutdown(server->uds_listen_fd, SHUT_RDWR)beforeclose(). On Linux, closing a listeningAF_UNIXsocket does not unblockaccept(2)on another thread;shutdown()is required to unblock the thread and returnEINVAL. Preserve this shutdown call on any upstream rebase touchingcore/src/mcp/mcp.c.core/src/mcp/transport_uds.c:vmaf_mcp_uds_thread_mainchecksuds_runningand guardsuds_listen_fddefensively before loop entry to prevent assertions if stopped immediately..github/workflows/build.ymland.github/workflows/libvmaf-build-matrix.yml: Homebrew installation on macOS uses a 3-attempt retry loop with backoff andbrew fetch --retry, plusHOMEBREW_NO_AUTO_UPDATE=1andHOMEBREW_NO_INSTALL_CLEANUP=1. Wholly fork-added workflows.
ci/release-artifacts-built-in-dev-container — native release artifacts built in canonical dev container (ADR-1178) (2026-09-04)¶
ci/release-artifacts-built-in-dev-container — native release artifacts built on self-hosted canonical runner (ADR-1178) (2026-09-05)¶
No rebase impact: all touched files (.github/actionlint.yaml, .github/workflows/dev-container-publish.yml, .github/workflows/supply-chain.yml, scripts/release/verify-native-release-artifacts.sh, scripts/release/tests/test-verify-native-release-artifacts.sh, scripts/ci/check-container-build.sh, scripts/ci/tests/test-check-container-build.sh, docs) are fork-local CI workflows, verification scripts, and documentation with no upstream Netflix/vmaf counterpart. No public C API, header, Meson option, or golden assertion is touched.
fix/vmaftune-state-bugs — libx264 two-pass CRF conflict fix (2026-09-03)¶
No rebase impact: all touched files (pkg/codecadapter/, pkg/ffencode/, pkg/corpus/, tools/vmaf-tune/) are fork-added Go and Python tuning tooling with no upstream Netflix/vmaf counterpart. No public C API, header, Meson option, or golden assertion is touched.
fix/vmafx-tune-go-gaps — resolve vmafx-tune Go parity gaps (#1272) (2026-09-04)¶
cmd/vmafx-tune/cmd/predict.go: wired saliency moments (pkg/saliency.ComputeMapandcomputeSaliencyMoments) intorunPredictvianewPredictSaliencyFunc, replacing the--use-saliencyusage error and allowing--use-saliencyto feed moments intopredictor.ExtractFeaturesand feature vectors (with graceful degradation to 0.0 moments when inference is unavailable).cmd/vmafx-tune/main.go,cmd/vmafx-tune/cmd/root.go,cmd/vmafx-tune/cmd/compare.go,docs/usage/vmafx-tune-go.md: removed stale comments, docstrings, and deadstubSubcommandreferencing the retired Pythonvmaf-tunebinary.- Wholly fork-added:
cmd/vmafx-tune/,pkg/predictor/,pkg/saliency/,pkg/tune/are all fork-local; upstream Netflix/vmaf has no Go rate-quality tuning CLI. No upstream rebase conflict.
feat/vmafx-cli-alias — vmafx CLI alias and --netflix-compat override (ADR-0690/0696) (2026-09-04)¶
core/tools/cli_parse.cpp,core/tools/cli_parse.h: detectsvmafxmode viadetect_vmafx_mode(argv[0]), setting modernized defaults (precision_max = true,precision_fmt = "%.17g", startup bannerVMAFX version <V> (precision=max), andvmafx --versionprintingVMAFX <V> (auto-backend, precision=max)).--netflix-compat/--netflix_compat: final post-parse override incli_parse()forcing CPU backend (settings->backend = VMAF_BACKEND_CPU),%.6fprecision format, and selectingVMAF_NETFLIX_COMPAT_MODEL_VERSIONinvalidate_cli_settings().core/include/libvmaf/model.h: added#define VMAF_NETFLIX_COMPAT_MODEL_VERSION "vmaf_v0.6.1"with the single-source pin comment (/* vmaf-model-pin: ... */). Must be preserved on upstream rebase to satisfyscripts/ci/check-default-model-single-source.sh.core/tools/meson.build: installsvmafxas a symlink tovmafviainstall_symlinkon POSIX systems with build-dir custom targetvmafx_build_symlink, or as a separate Windows executable compiled from the same sources. Keep this block when rebasing tool build configurations.ai/pyproject.toml,tools/vmaf-tune/pyproject.toml,mcp-server/vmaf-mcp/pyproject.toml: companion entrypointsvmafx-train,vmafx-tune,vmafx-mcpadded alongside legacy commands.
feat/default-model-v1-0-16 — loud model-dimension validation in the CLI (2026-09-04)¶
core/src/feature/feature_dimensions.h: wholly fork-added. Single point that turns the extractors' own minimum-dimension rules into a CLI-facing check. Do NOT copy thresholds into it; it must keep readingCAMBI_MIN_WIDTH_HEIGHTfromcambi_internal.handSPEED_INTERNAL_MIN_DIMENSIONfromspeed_internal.h, so the check and the extractors cannot drift apart.core/src/feature/cambi_internal.h,core/src/feature/speed_internal.h: gained small*_validate_dimensions()/speed_chroma_dimensions()helpers so the header above has something to call. Upstream Netflix carriescambi.cbut not these helpers; on a sync, keep the fork's helpers and re-derive the threshold from whatever upstream's cambi asserts internally.core/tools/vmaf.cpp: the validation runs inload_model_entryandload_model_collection_entryafter the load succeeds and before feature overloads. Upstream'svmaf.chas neither function in this shape (ADR-0809 C++ conversion); resolve conflicts by keeping the fork's version and re-applying the twovmaf_validate_model_dimensionscalls. The error is intentionally NOT gated on--quiet.python/test/ssimulacra2_test.py: fork-added test;test_ssimulacra2_small_160x90must keep passing--model version=vmaf_v0.6.1— it measures ssimulacra2 only and must not depend on whether the current default model can run at 160x90.
docs/readme-overhaul — README.md overhaul for clarity and accuracy (2026-09-03)¶
no rebase impact: edits fork documentation (README.md, CHANGELOG.md, changelog.d/changed/readme-overhaul.md, docs/state.md, docs/rebase-notes.md) only. Upstream Netflix/vmaf has a completely separate README; if an upstream sync touches README.md, preserve the fork's overhauled version.
chore/drop-ansnr — scrub residual ansnr references across code and comments (ADR-0865) (2026-09-03)¶
ai/data/feature_extractor.py,core/src/feature/feature_extractor.cpp,core/src/feature/offset.c,core/src/feature/x86/moment_avx2.c,core/src/hip/kernel_template.h,core/test/test_hip_smoke.c,mcp-server/vmaf-mcp/tests/test_p1_tools.py: removed stale comments and docstrings referencing the sunset ANSNR /float_ansnrfeature extractor.docs/metrics/ansnr.md: retained as a concise metric deprecation stub pointing callers topsnr_yandpsnr_hvs; updated ADR citation from ADR-0709 to ADR-0865.compat/python-vmaf/core/quality_runner.pyandcore/test/test_metal_kernel_coverage_audit.c: deliberately preserved load-bearing backward compatibility stubs and negative dispatch tests.- Rebase impact: None. All modifications touch fork-added comments or fork-added test/doc surfaces. If upstream touches
core/src/feature/offset.c, preserve theadm.c / motion.ccomment text.
feat/mcp-tinyai-flags — tiny-AI scoring flags and input validation (2026-09-03)¶
no rebase impact: MCP servers (cmd/vmafx-mcp and mcp-server/vmaf-mcp) are wholly fork-added surfaces with no upstream Netflix/vmaf counterpart.
fix/untrack-venv-symlink — no virtualenv path may be tracked (2026-09-04)¶
.gitignore: the.venv*line (no trailing slash) is load-bearing next to.venv*/. The slash form matches directories only; a symlink or file named.venvneeds the slash-less form. Upstream Netflix ignores nothing venv-related, so a sync will not conflict here, but do not "simplify" the pair back to one line.scripts/ci/check-no-tracked-venv.sh: wholly fork-added backstop. An ignore rule stopsgit addfrom picking a path up; it does nothing about an already-tracked path and nothing againstgit add -f. Keep the gate even though the ignore rule exists.
chore/dedup-sweep-2026-09-03 — collapse dead C twins (gpu_picture_pool, opt) and stale pre-rename paths (2026-09-03)¶
core/src/opt.c: deleted dead C translation unit. Upstream still haslibvmaf/src/opt.c. If an upstream sync touchesopt.c, port any new option keys or types intocore/src/opt.cpp(which is C++23 withstd::optionaland ADR-1080 UBSan fixes, compiled intolibvmafviaopt_cpp23_lib) and keepcore/src/opt.cdeleted.core/src/gpu_picture_pool.c: deleted dead C translation unit. Wholly fork-added; upstream has no GPU picture pool. No upstream rebase conflict.core/test/meson.build:test_integer_ssim_simdandtest_motion_avx512_parityupdated to compile/linkgpu_picture_pool.cppandwave8_opt_only_objects/log_cpp23_test_objects;test_gpu_picture_pool_partial_initwired. Fork-added test wiring; resolve any rebase conflict by preserving the references to.cppand the test object libraries.docs/usage/{bd-rate,matlab,python}.md,core/tools/meson.build,core/tools/compat/win32/getopt.{c,h},core/tools/vmaf_roi_core.h,testdata/bench_all.sh: mechanical path updates fromlibvmaf/->core/andpython/vmaf/->compat/python-vmaf/(ADR-0700). No upstream rebase impact.
fix/picture-pool-twin-drift — port concurrency and lifecycle fixes to picture_pool.cpp (2026-09-04)¶
core/src/picture_pool.cpp: C++ twin ofcore/src/picture_pool.ccompiled intolibvmafviapicture_pool_cpp23_lib(ADR-0768). Ported four fixes frompicture_pool.cthat had drifted: (1) ADR-0778 Fix-E two-pass picture preallocation to avoid leaking buffers whenvmaf_picture_allocfails; (2) ADR-1020 Fix 3 stack-local snapshot ofpool->pictures[idx]under mutex before unlock; (3) ADR-0960 Fix A.3pic->priv = nullptrafter free on fetch failure; (4) ADR-0960 Fix A.2pthread_cond_signal(&pool->available)onreturn_to_poolto wake waiting threads. Preserved C++23 / ADR-1138 idioms (nullptr,std::free). On upstream rebase: Netflix/vmaf has only the C file; resolve any future changes to picture pool by keepingpicture_pool.cppin sync withpicture_pool.c.core/test/meson.build: addedtest_picture_pool_cpp_error_pathscompilingpicture_pool.cppdirectly into an internal error-path test target withoverride_options : ['cpp_std=' + libvmaf_cpu_cpp_std].
fix/fixture-cache-poisoning — fixture caches must not be written by failed runs (2026-09-04)¶
.github/workflows/{build,libvmaf-build-matrix,tests-and-quality-gates}.yml: all three are fork-added and have no upstream counterpart, so no sync conflict is expected. The invariant they now encode is easy to undo by accident: the fixture cache MUST stay split intoactions/cache/restoreplus a separateactions/cache/savegated onsuccess(). Collapsing them back into the combinedactions/cacheaction silently restores the poisoning bug, because that action's post-job save runs even when the job was cancelled or failed.scripts/ci/prune-corrupt-fixtures.sh,scripts/ci/test-prune-corrupt-fixtures.sh: wholly fork-added. The pruner deletes files, so its match rules are deliberately narrow — empty, a Git-LFS pointer, an HTML error page, a JSON API error. Do not widen them to include "file is smaller than expected": several real fixtures are legitimately tiny, and the self-test pins that case.compat/python-vmaf/config.py::download_reactivelyis the reason any of this is needed — it re-fetches only when the local file is ABSENT. If a future sync makes it validate content instead, the pruner becomes redundant and can go.
fix/vcs-version-bare-sha — VMAF_VERSION must never be a bare commit SHA (2026-09-03)¶
core/include/meson.build: upstream Netflix/vmaf carries the samevcs_tag()call with--always. The fork deliberately drops that flag and pins an explicitfallback: meson.project_version(). On an upstream sync this file will conflict; keep the fork's side. Restoring--alwaysreintroduces the defect where a tagless or shallow checkout yields a bare abbreviated object name asVMAF_VERSION.scripts/ci/check-vcs-version-not-bare-sha.shfails the build if the flag comes back, so a careless conflict resolution is caught rather than shipped..github/workflows/build.yml: fork-added workflow; no upstream counterpart. Thefetch-depth: 0on the checkout is load-bearing (git describe needs tags plus the commit distance), as is|| exit /b 1in the Windowsforloop (GitHub runsshell: cmdwith/V:OFF, so without it only the last executable's exit code reaches the step result).scripts/ci/check-vcs-version-not-bare-sha.sh,changelog.d/fixed/*: wholly fork-added.
feat/mcp-score-gaps — MCP scoring surface completeness (epic #1240) (2026-09-03)¶
no rebase impact: MCP servers (cmd/vmafx-mcp and mcp-server/vmaf-mcp), pkg/libvmaf/paths.go, and docs/mcp/ are wholly fork-added surfaces with no upstream Netflix/vmaf counterpart.
gap/hip-bucket-v2 — AMD ROCm HIP backend gap closure (ADR-1154) (2026-09-03)¶
core/src/feature/hip/andcore/src/hip/: all touched files (ciede_hip.c,float_adm_hip.c,float_moment_hip.c,float_motion_hip.c,float_psnr_hip.c,float_ssim_hip.c,integer_cambi_hip.c,integer_motion_v2_hip.c,integer_ms_ssim_hip.c,integer_psnr_hip.c,integer_psnr_hvs_hip.c,integer_ssim_hip.c,dispatch_strategy.c,picture_hip.c) are wholly fork-added (upstream Netflix/vmaf does not have a HIP backend). No upstream rebase conflicts will occur here.core/src/feature/hip/integer_adm/adm_decouple.hip,core/src/feature/hip/integer_moment_hip.h,core/src/feature/hip/integer_moment/moment_score.hip: deleted orphan uncompiled fork files. No upstream counterpart.core/src/libvmaf.c: inflush_context_serial, drainedgpu_pendingfor all non-CUDA/SYCL extractors before flush. This is fork-added GPU pipeline logic; resolve any future upstream merge conflict by preserving the loop.Makefile:test-netflix-goldentarget combinesCUDA_VISIBLE_DEVICES=""andVMAF_FORCE_BACKEND=cpu. Preserve both on rebase.core/test/meson.build: updated deferral comments fortest_hip_adm_parityandtest_hip_ssim_parityto cite ADR-1154.
fix/test-feature-tidy-clean — modernise the assertions ported by #1219 (2026-09-03)¶
core/test/test_feature.cpp: upstream-mirror test (Netflixlibvmaf/test/test_feature.c), and the sole surviving side since #1219 deleted the C twin. This change rewrites the fork-local portions to C++ idiom, so a future upstream sync will conflict here in predictable, mechanical ways. Resolve them as follows:- Upstream spells
NULL; the fork spellsnullptr. This is a C++ TU, so the fork's spelling wins. ADR-1138'sNULLrule is scoped to C translation units (MSVC/std:clatesthas no Cnullptr) and does not apply here. - Upstream uses
typedef struct {...} Name;inside the test bodies; the fork hoistsTestStateto namespace scope as a plainstructbecause the shared option table needsoffsetof(TestState, ...)at namespace scope. Keep the fork's shape. - Upstream's option tables end with a
{0}sentinel; the fork uses{}. Equivalent zero-initialisation, and{}is whatmodernize-use-designated-initializersaccepts. - Upstream has one large
test_feature_name_from_options(); the fork splits it into four named cases sharing a namespace-scopeg_optionstable plus akAllDefaultsbaseline, because the single function exceeded the 60-linereadability-function-sizethreshold. Port new upstream assertions into whichever split case matches, or add a new one and register it inrun_tests(). - The fork frees each heap result before asserting on it.
mu_assertexpands to an earlyreturn, so asserting with a live pointer leaks it on failure. Preserve this ordering when porting upstream assertions, which do not observe it. - The
#include "feature/feature_name.cpp"unity include is fork-local (ADR-0729) and carries a citedNOLINTNEXTLINE(bugprone-suspicious-include); upstream includes the.c. Keep the fork's include and its citation. scripts/ci/tidy-baseline-cpu.json: fork-local ratchet state (ADR-1142), no upstream counterpart. Only this file's entry was removed; every other file keeps CI's measured number. No rebase impact.
fix/json-model-libsvm-dup-key-leak — duplicate-key leaks in the model parser (2026-09-03)¶
core/src/svm.cpp: vendored libsvm. The fork adds fiveexceptAssert(!model-><field>, "duplicate <field> row in model file")guards inSVMModelParser::parse_header()forrho,label,probA,probBandnSV. Upstream libsvm has no such guard and will happilyMallocover the previous pointer. A re-vendor must re-apply them or the 8-byte-per-field leak returns.core/src/read_json_model.candcore/src/read_json_model.cpp: the samesvm_free_and_destroy_model(&model->svm)call must exist inparse_libsvm_modelin both files. They are a twin pair that is not inscripts/ci/twin-drift-allowlist.txt: the library builds the.cpp, whilecore/test/fuzz/meson.buildcompiles the.cdirectly into the fuzz harness. Fixing only one leaves the other leaking, and the symptom depends on which binary you test — the library-side reproducer looks fixed while the fuzz lane stays red, or vice versa.scripts/ci/twin-drift-check.shreports the.cas a "test-only twin side".core/src/model.c: upstream-mirror.vmaf_model_destroywalksmodel->feature_cap, where upstream walksmodel->n_features. The fork's form frees feature slots thatparse_feature_opts_dictspopulated without bumpingn_features; reverting it ton_featuresreintroduces the leak. It is safe becausefeature_capis the allocated element count andensure_feature_capacityzeroes newly grown slots.core/test/fuzz/json_model_corpus/seed_duplicate_model_key.jsonandseed_duplicate_rho_row.json: fork-added corpus seeds, no upstream counterpart.core/test/fuzz/json_model_known_crashes/holds the reproducer for the still-open third leak and is excluded from the nightly seed path.
fix/code-scanning-open-alerts — resolve open code-scanning alerts & re-audit security dismissals (2026-09-03)¶
no rebase impact: fork-local fixes and cleanups.
core/src/feature/feature_name.cpp: replaced trivial single-caseswitchwithif/else.core/src/feature/mkdirp.cpp: convertedforloop modifying loop variable to idiomaticwhile.core/src/pdjson.c: removed redundant lower bound0xC2 <= uinutf8_seq_length(simplified tou <= 0xDF).compat/python-vmaf/tools/decorator.py: addedusedforsecurity=Falseto 3 SHA-1 memoization keys.ai/sidecar/online_trainer.py: this historical branch preserved0o660and a# nosemgreprationale. Do not preserve that disposition in a current rebase: ADR-1309 supersedes it with owner-only0o600, no suppression, and an identity-checked pathname lifecycle after the claimed group-peer Helm topology was found not to be wired.mcp-server/vmaf-mcp: removed unused import intest_smoke_e2e.pyand convertedserver.pyHTTP branch import to dynamicimportlibto break static circular import.go.mod,go.sum: upgradedgolang.org/x/cryptotov0.56.0.
feat/default-model-v1-0-16 — the fork's default model is vmaf_v1.0.16_3d0h (ADR-1169) (2026-09-03)¶
Permanent, user-visible divergence from upstream. Read this before any sync.
core/include/libvmaf/model.h: the fork definesVMAF_DEFAULT_MODEL_VERSIONas"vmaf_v1.0.16_3d0h". Upstream has no such macro and still hardcodes"vmaf_v0.6.1"inlibvmaf/tools/cli_parse.c. An upstream sync will look like it wants to revert the default. It does not — keep the fork's value. Verified against upstream master on 2026-09-03.core/include/libvmaf/model.handcore/src/model.c: public model accessorsvmaf_model_feature_countandvmaf_model_feature_nameare fork-added and upstream Netflix has no counterpart, so on sync keep them.core/tools/cli_parse.cpp: themodel_cnt == 0fallback reads the macro. The AOM CTC preset in the same file keeps the literal"vmaf_v0.6.1"with avmaf-model-pin:comment, because the CTC specification mandates that exact model; that literal must survive a sync too, and the two must not be "unified".python/test/vmafexec_test.py: upstream-mirror golden test. The fork changes exactly one line intest_run_vmafexec_runner_use_default_built_in_model— itsoptional_dictnamesvmaf_v0.6.1instead of settinguse_default_built_in_model: True. No assertion value differs from upstream. On a sync, upstream will restore theuse_default_built_in_modelform; re-apply the fork's explicit model, because with the fork's default that test raisesKeyError('VMAFEXEC_vif_scale0_score')(the v1.0.16 family does not emitvif_scale0..3ormotion2). Never resolve this by editing anassertAlmostEqualvalue.python/test/default_model_test.py: fork-added, no upstream counterpart. HoldsEXPECTED_DEFAULT_MODEL; update it if the default ever changes again.compat/python-vmaf/core/quality_runner.py:VmafQualityRunner. DEFAULT_MODEL_FILEPATHstaysvmaf_v0.6.1.jsonon purpose — that harness exists to reproduce Netflix's published numbers. Do not "fix" it to follow the fork default.- Fork-only, no upstream counterpart:
pkg/model/,pkg/corpus/resolution.go,tools/*/defaultmodel.py,tools/vmaf-tune/src/vmaftune/resolution.py,scripts/ci/check-default-model-single-source.sh, the ADR and the docs. - NEG invariant:
DefaultNEGVersion/DEFAULT_MODEL_NEGmust stay independent constants naming the v0.6.1 family. Never re-derive them asDefaultVersion + "neg"— with a v1 default that synthesisesvmaf_v1.0.16_3d0hneg, which does not exist.
feat/default-model-single-source — one definition of the default model (ADR-1168) (2026-09-03)¶
core/include/libvmaf/model.h: upstream-mirror header. The fork adds#define VMAF_DEFAULT_MODEL_VERSIONand declaresVMAF_EXPORT const char *vmaf_default_model_version(void);. Neither exists upstream. An upstream sync that rewrites this header must keep both; they are the anchor for the whole single-source scheme andscripts/ci/check-default-model-single-source.shhard-fails without the macro.core/src/model.c: upstream-mirror. The fork appends thevmaf_default_model_version()definition at end of file, deliberately aftervmaf_model_version_next(), so an upstream diff to the built-in model table above it does not conflict with it.core/tools/cli_parse.cpp: upstream-mirror (ascli_parse.cupstream). Two divergences. (1) Themodel_cnt == 0fallback readsVMAF_DEFAULT_MODEL_VERSIONwhere upstream writes"vmaf_v0.6.1"— keep the fork's macro. (2) The two AOM CTC preset entries keep the literal"vmaf_v0.6.1"and carry avmaf-model-pin:comment; that is correct and must survive, because the CTC specification mandates that exact model and it must NOT follow the fork's default. If an upstream sync drops the comment the gate will fail until it is restored.core/tools/vmaf_vpl.c,core/src/mcp/compute_vmaf.c: fork-added tools; both now include<libvmaf/model.h>for the macro. No upstream counterpart.core/tools/cli_parse.c: untouched. It is the uncompiled twin (onlycli_parse.cppis incore/tools/meson.build) and is allowlisted in the gate rather than edited, so this change does not collide with the twin decision in PR #1222.- Everything else is fork-only with no upstream counterpart:
pkg/model/,tools/*/defaultmodel.py, the Go and Python call sites,scripts/ci/check-default-model-single-source.shand its test,docs/development/default-model.md, the ADR and the mkdocs nav entry.
fix/adm-cm-gpu-border-and-rounding — integer ADM GPU border indexing and row-level rounding (ADR-1167) (2026-09-03)¶
core/src/feature/cuda/integer_adm/adm_cm.cu&core/src/feature/hip/integer_adm/adm_cm.hip:- Defect 1 (border row selection): Replaced running pointer offsets with explicit absolute indexing
{row_top, row_bot, col_l, col_r}and evaluated centercsf_aati * src_stride + jwheni == 0 && top <= 0. - Defect 2 (distributed rounding shift): Changed kernels to stride across columns with full row accumulation into 64-bit integers and applied
(row_total + add_shift_inner_accum) >> shift_inner_accumonce per row before atomic add intoaccum_global. core/src/feature/cuda/integer_adm_cuda.c&core/src/feature/hip/integer_adm_hip.c:- Launched CM kernels with
gridDim.x = 1to ensure single-block/warp column striding per row. core/test/test_adm_small_border.c,core/test/test_adm_wide_rounding.c,core/test/meson.build:- Added new regression parity tests exercising small border frames and wide row rounding for CUDA and HIP. no rebase impact: all touched GPU kernels, host wrappers, and tests are fork-added (upstream Netflix/vmaf does not have CUDA or HIP ADM kernels).
fix/twin-dead-sides — resolve dead twin sides (T-TWIN-DEAD-SIDES-2026-09-02) (2026-09-03)¶
core/src/model.cpp: fork-added twin deleted.core/src/model.cis the sole authoritative model TU and tracks upstream Netflixlibvmaf/src/model.cdirectly. No rebase conflict onmodel.c.core/test/test_dict.c: deleted. Upstreamlibvmaf/test/test_dict.cchanges should be ported tocore/test/test_dict.cppduring future upstream syncs.core/test/test_feature.c: deleted. Upstreamlibvmaf/test/test_feature.cchanges should be ported tocore/test/test_feature.cppduring future upstream syncs.
refactor/c-rework-adm — integer ADM upstream-mirror rework (2026-09-02)¶
core/src/feature/integer_adm.c and core/src/feature/adm_tools.c (both keep the Netflix header) were restructured under ADR-0141 / ADR-1141 with every kernel expression, integer width, rounding term and float summation order kept verbatim. An upstream Netflix hunk to either file no longer applies textually; re-port it by hand into the function that now owns the code. On conflict keep the fork's version. Function map:
integer_adm.c
- The twenty
ADM_CM_THRESH_S_*/I4_ADM_CM_THRESH_S_*/ADM_CM_ACCUM_ROUND/I4_ADM_CM_ACCUM_ROUNDmacros are gone. The nine corner / edge / interior threshold variants areadm_cm_thresh()/i4_adm_cm_thresh(): the row / column before the first edge mirrors to index 1, the one past the last edge clamps to the last index (i_m1 = i == 0 ? 1 : i - 1,i_p1 = i == h - 1 ? h - 1 : i + 1, same forj), nine terms added in the macro order with the(int32_t)centre-term cast of scales 1..3 kept (the scale-0(int16_t)cast was removed by ADR-1402, 2026-10-01).adm_cm_accum_round()/i4_adm_cm_accum_round()carry the cube rounding over anAdmCmBand(shift_sub, add_shift_sq, shift_sq, add_shift_cub, shift_cub). An upstream change to the neighbourhood or the rounding lands there, once. adm_cm()/i4_adm_cm(): prologue inadm_cm_ctx_init()/i4_adm_cm_ctx_init()(AdmCmCtx/I4AdmCmCtx), per-sampleadm_cm_accum_px()/i4_adm_cm_accum_px()(i4_adm_cm_scale()for the CSF weighting), per-rowadm_cm_row()/i4_adm_cm_row(), sharedadm_cm_fold(),adm_num_scale(). The upstream four-way border branch is the pair of predicatesleft_edge = left <= 0/right_edge = right > w - 1passed to the row helper; do not reintroduce the branch (its third arm indexedrfactor[i * src_stride + w - 1], an unreachable out-of-bounds read). Rounding terms go throughadm_half_shift()(guardedpow(2, shift - 1)).adm_decouple()/adm_decouple_s123(): per-bandadm_decouple_band()/adm_decouple_band_s123(), sharedadm_angle_flag(); parameters renamedgain/lut(positions unchanged — theAdmStateprototypes and the SIMD twins are untouched). Theint32_t tmp_knarrowing and theMIN/MAXdouble-to-int assignments are upstream semantics; keep them.adm_csf()/i4_adm_csf()/adm_csf_den_scale()/adm_csf_den_s123()shareadm_csf_factors(),adm_csf_rfactor_scale0(),adm_border()/adm_border_filt(),adm_den_scale_finalise(),i4_cube_term(). The ADR-0155 rounding terms (Netflix#955,int32_t, sign-negated for scales 1..3) live ini4_adm_round_terms()with the file-scopei4_shift_dst[]/i4_shift_flt[]; do not widen them.- DWT:
adm_dwt2_8()/_8_lo()/_16()/_16_lo()are loops overadm_dwt2_vpass_8()/adm_dwt2_vpass_16()andadm_dwt2_hpass()onadm_dwt2_tap4()(tmphi == NULLselects the low-pass-only variant);adm_dwt2_s123_combined()isi4_dwt2_vpass()+i4_dwt2_hpass()/i4_dwt2_hpass_bands()oni4_dwt2_tap4()(ref and dis still interleaved per row).dwt2_src_indices_filt()callsdwt2_src_indices_1d()twice. integer_compute_adm(s, ref_pic, dis_pic, res)takes the parameters fromAdmStateand fillsAdmResult; per scale it callsinteger_adm_scale0()(including theadm_skip_scale0low-pass path) orinteger_adm_scale_s123();adm_src_stride()andadm_result_finalise()hold the stride selection and thenumden_limit/den == 0finalisation.init()isinit_dispatch_scalar()+init_dispatch_simd()+init_buffers(); failure runsfree_buffers()(shared withclose(), nogoto).extract()delegates thedebug=trueappends toextract_debug_features()overscale_feature_names[]/debug_scale_feature_names[](append order unchanged).- Surviving suppressions, all cited: the file-scoped
NOLINTBEGIN/END(modernize-use-nullptr)bracket (ADR-1138; keep theNOLINTENDline at EOF when appending),readability-non-const-parameter cppcheck constParameterCallbackon the two decouplelutparameters andcppcheck constParameterCallbackonextract()'s pictures (frozen dispatch /VmafFeatureExtractor::extractprototypes), the cross-TUmisc-use-internal-linkagemarker onvmaf_fex_integer_adm.
adm_tools.c
adm_decouple_s():adm_angle_flag_s()(bothADM_OPT_AVOID_ATANarms),adm_decouple_band_s(),adm_border_filt_s().adm_csf_s()/adm_csf_den_scale_s()/adm_cm_s()shareadm_csf_rfactor_s()(+adm_csf_factor_overrides_s()for theadm_f1sN/adm_f2sNoverrides),adm_border_s()andadm_fold3_s().adm_cm_s()mirrors the integer shape:AdmCmCtxS,adm_cm_thresh3x3_s()(the nineadm_tools.hADM_CM_THRESH_S_*float macros, same summation order; the header macros are now unused),adm_cm_accum_px_s(),adm_cm_row_s(). Inner / outer accumulators stayfloat(golden-gated; ADR-0418 widened onlyadm_sum_cube_s).adm_dwt2_s()is deliberately NOT split: it carries the ADR-1057optimize("-ffp-contract=off")attribute /#pragma clang fp contract(off)bracket, and helpers would each need the attribute; thereadability-function-sizemarker cites this.adm_dwt2_lo_s()keeps its default contraction semantics — do not share helpers between the two.adm_dwt2_d()usesadm_dwt2_tap4_d()/adm_dwt2_hpass_d().dwt2_src_indices_filt_s()callsdwt2_src_indices_1d_s()twice.get_noise_constant()isstatic;adm_dwt2_lo_d()andadm_buffer_copy()(no caller in the tree, never declared in the header) are removed.adm_dwt2_d()stays: the Cython extension
gap/cuda-intel-bucket — CUDA and Intel SYCL backend gap closure (2026-09-02)¶
- Dead CUDA source cleanup: Removed
core/src/feature/cuda/integer_adm/adm_decouple.cu(superseded by inline decoupling inadm_csf.cuviaadm_decouple_inline.cuh) and uncompiled/orphanedcore/src/feature/cuda/resolution_dispatch.c/.h. No rebase impact as these files were dead in the fork tree. - Unified GPU dispatch environment wiring:
core/src/cuda/dispatch_strategy.candcore/src/sycl/dispatch_strategy.cppnow route throughvmaf_gpu_dispatch_env_get(defined ingpu_dispatch_env.cpp). Local WindowsINIT_ONCE/ POSIXpthread_onceand rawgetenvcalls withNOLINT(concurrency-mt-unsafe)were removed.gpu_dispatch_env_cpp23_libis linked intolibvmaf_feature_static_libso all backends, test binaries, and shared libraries inherit it without duplicate definitions. - CUDA graph dispatch honest fallback: When
VMAF_CUDA_DISPATCH=graphis requested,vmaf_cuda_select_strategyemits a clear warning and falls back toVMAF_CUDA_DISPATCH_DIRECTsince static graph capture is not implemented for the driver API. - Python harness backend selection:
compat/python-vmaf/__init__.py(ExternalProgramCaller.call_vmafexecandcall_vmafexec_multi_features) readsVMAF_FORCE_BACKEND(andVMAF_BACKEND), mapping it to--backend <name>. - CI GPU test scoping:
.github/workflows/tests-and-quality-gates.ymlscopes pytest withVMAF_FORCE_BACKENDaway from the 5 Netflix CPU golden assertion files (quality_runner_test.py,feature_extractor_test.py,vmafexec_test.py,vmafexec_feature_extractor_test.py,result_test.py) to prevent false failures from ULP-relaxed GPU float differences. - SYCL Win32 stub logging:
core/src/sycl/dmabuf_import.cppemits an informative error log explaining that DMA-BUF is a Linux kernel primitive before returning-ENOSYSon_WIN32.
refactor/c-rework-tools-v2 — Upstream-mirror CLI and tool translation units lint rework (ADR-1155) (2026-09-02)¶
core/tools/vmaf.cpp: Reworked in place to C++23. Implemented RAII resource guards (VmafResourceGuard), encapsulatedModelArraysaccessors, decomposed monolithic functions into modular helpers. Upstream syncs tovmaf.c/vmaf.cppwill need manual conflict resolution; keep the RAII cleanup and helper structure.core/tools/cli_parse.cpp: Reworked to C++23. Modernized argument parsing with typed helpers,std::string_viewcomparisons, explicit bounds checks (--width > 0,--height > 0).core/tools/cli_parse.c: DELETED. Resolved as dead twin under ADR-1153 precedent after verifying 0 unique behaviors or assertions vscli_parse.cpp. Test targets (test_cli_parse,test_cli_parse_long_only_args,fuzz_cli_parse) now compilecli_parse.cpp. If upstream touchescli_parse.c, port changes directly tocli_parse.cpp; do NOT resurrectcli_parse.c.core/tools/y4m_input.candcore/tools/vmaf_bench.c: RetainNULLas the null pointer constant (ADR-1138) to maintain MSVC/std:clatestcompatibility on Windows CI legs. NOLINTBEGIN/NOLINTEND brackets suppressmodernize-use-nullptr. Fixed arithmetic types (size_t/ptrdiff_t) to eliminate overflow warnings.core/tools/cli_parse.h: Header guard renamed from reserved__VMAF_CLI_PARSE_H__toVMAF_CLI_PARSE_H.
gap/metal-bucket — Metal gap bucket closure & dispatch alignment (2026-09-02)¶
core/src/feature/metal/*.mm: set.flags = VMAF_FEATURE_EXTRACTOR_METALacross all 9 previously-unflagged Metal feature descriptors (float_adm_metal,float_vif_metal,integer_adm_metal,integer_cambi_metal,integer_ciede_metal,integer_psnr_hvs_metal,integer_ssim_metal,integer_vif_metal,ssimulacra2_metal).core/src/feature/feature_extractor.cpp: includedVMAF_FEATURE_EXTRACTOR_METALingpu_maskso that CPU-only requests (flags == 0) filter out Metal extractors identically to CUDA/SYCL/HIP.core/src/libvmaf.c: addedHAVE_METALcheck incompute_fex_flagsto ORVMAF_FEATURE_EXTRACTOR_METALwhen a Metal context is active (vmaf->metal.state != NULL).core/src/dnn/ort_backend.c,core/src/dnn/ort_backend_internal.h,core/test/dnn/test_ort_internals.c: probed CoreML under#ifdef __APPLE__first inVMAF_DNN_DEVICE_AUTO, implemented selection order table helpervmaf_ort_internal_auto_ep_order(int is_apple)and unit tests.core/include/libvmaf/libvmaf_metal.h,docs/backends/metal/index.md,docs/metrics/features.md,docs/ai/inference.md: aligned documentation and doc comments with runtime truth.
fix/neo-derive-matched-set — derive Intel NEO matched set at build time (ADR-1145) (2026-09-02)¶
no rebase impact: fork-only container and Renovate configuration.
dev/scripts/fetch-intel-neo.pydynamically resolves the matched set of gmmlib and IGC deb packages from the pinnedNEO_VERrelease assets and verifies their sha256 checksums at container build time.- Preserve its GitHub-only HTTPS boundary, exact-host authorization, bounded metadata reads, atomic downloads, and fail-closed asset/package validation.
- Preserve the optional BuildKit
github_tokensecret transport from ADR-1271. Never restoreARG GITHUB_TOKEN,ENV GITHUB_TOKEN, or token-valued--build-arg; raw builds without a secret and Compose builds with an unset/empty host variable must stay anonymous. dev/ContainerfileremovesGMMLIB_VERandIGC_VERARGs; Renovate regex managers for gmmlib and IGC removed fromrenovate.json.
gap/cpu-ci-bucket — retire dead orphan motion_v2 x86 SIMD duplicate files (GAP-BUILD-ORPHAN-DEAD-SIMD-MOTION-V2)¶
core/src/feature/x86/motion_v2_avx2.{c,h}andmotion_v2_avx512.{c,h}were removed. These files were orphaned duplicates ofcore/src/feature/x86/motion_avx2.{c,h}andmotion_avx512.{c,h}(which were added when upstream pipelined integer motion ina4a1492d3). The library build and tests already linked againstmotion_avx2.candmotion_avx512.c. Consumers (core/src/feature/integer_motion_v2.c,core/test/test_motion_v2_simd.c,core/test/test_motion_avx512_parity.c) now uniformly#include "x86/motion_avx2.h"and#include "x86/motion_avx512.h".
chore/intel-neo-matched-set-bump (2026-09-02)¶
no rebase impact: dev/Containerfile is fork-local.
fix/fuzz-dict-cpp-and-setup-meson — fuzz dict.cpp + setup script meson (2026-08-31)¶
core/test/fuzz/meson.build,scripts/setup/ubuntu.sh— both fork-added (ADR-0270 fuzz harnesses; setup script has no upstream counterpart). no rebase impact: neither file exists upstream.
ci/impact-planner — required CI routed by measured impact (2026-09-02)¶
.github/workflows/*.yml,.github/ci-impact.json,scripts/ci/plan-ci-impact.py,scripts/ci/tests/test_ci_impact.py— all fork-local CI. no rebase impact: upstream Netflix/vmaf has none of these files.- Edit-sensitive pair (not rebase-sensitive): a workflow that hosts a check named in
required-aggregator.ymlmust never regain a workflow-levelpaths:/paths-ignore:filter — the contract testWorkflowContract.test_required_contexts_workflows_have_no_path_filtersfails if one does. Route inside the job via the planner instead.
ci/twin-drift-gate — .c/.cpp twin-drift + stale-source-reference gate (2026-09-02)¶
scripts/ci/twin-drift-check.sh,scripts/ci/twin-drift-allowlist.txt,scripts/ci/tests/test-twin-drift-check.sh, thetwin-drift-checkjob in.github/workflows/lint-and-format.yml, its row inrequired-aggregator.ymland thetwin-drift-checkpre-push hook are fork-local (ADR-1135; no upstream counterpart). Keep the workflowname:and the aggregator row identical — the aggregator matches names exactly.- The gate reads every tracked
meson.build,python/setup.pyand*.pyx, most of which are upstream-mirror files. After an upstream sync that adds, renames or removes a source, runbash scripts/ci/twin-drift-check.shbefore pushing: a stale path in an upstream-mirror build file is fixed in the build file (never allowlisted), and a new same-directory.c/.cpppair must have both sides compiled or the dead side listed in the allowlist with a reason. core/test/fuzz/meson.build(fork-added, ADR-0270):fuzz_json_modelnow lists../../src/dict.cppwithcpp_args : fuzz_flags— the same hunk as #1186; whichever lands second rebases onto an identical line.
refactor/go-dedup-tune-shadow — ADR-1137 shadow-package consolidation (2026-09-02)¶
Go-only; no upstream Netflix/vmaf counterpart, so no rebase conflict surface. Supersedes item 1 of the "vmafx-tune Go port integration" note below: internal/pyjson and internal/pyjsonstrict are gone, and pkg/pyjson is the single CPython-JSON encoder — the two Python entry points they mirrored (json.dumps with bare NaN tokens, jsonio.dumps_strict with null) are one Options.NonFinite field. Invariants a future change must not undo:
- Shared layers live outside
pkg/tune/.pkg/pyjson,pkg/pymath,pkg/hdr,pkg/codecadapter,pkg/predictor(now also home to the one ORT-session adapter,ORTSession/NewWithModel) andpkg/ffencodeare the one implementation each;pkg/tune/{auto,sidecar,executor}consume them.pkg/tune/{codec,predictor,pyjson}exist only as one-file transitional aliases becausepkg/tune/sidecar/andcmd/vmafx-tune/cmd/sidecar.gobelong to the in-flight #1187; when that lands, repoint those imports and delete the aliases. Do not re-createpkg/tune/{hdr,pymath},internal/pyjson*orcmd/vmafx-tune/cmd/ortsession.gowhen rebasing a branch that still uses them — repoint the import. EncodeRequest/BuildFFmpegCommand/ParseVersionsinpkg/corpus,pkg/encodeprofileandpkg/tune/executorare aliases and one-line wrappers overpkg/ffencode. Their argv tables still pin the contract under the local name; a fix belongs inpkg/ffencode, never in a re-grown local copy.pkg/encodeprofile's wrapper keeps the strict preset / quality check on a registered codec;pkg/corpus.DetectHDRkeeps Python's missing-file check in front ofpkg/hdr.Detect.- The Python is the tiebreaker where the duplicates disagreed: content light
int()truncation, the libx264 fallback for a partial coefficient table,repr()float thresholds in argv, and[]/{}for a nil Go slice or map. The AMF argv de-duplication (ADR-1125, pkg/codecadapterAGENTS.mdinvariant 3) now applies topkg/tune/executoras well. - Every Python-derived fixture moved with its winner —
pkg/pyjson/testdata/float_repr.txt,pkg/codecadapter/testdata/python_adapters.json,pkg/predictor/testdata/python_predictor.json,pkg/hdr/testdata/,pkg/pymath/testdata/. Regenerate them only alongside a coordinated change on both sides.
refactor/x86-adm-avx-macro-hygiene — x86 ADM AVX2/AVX-512 macro hygiene (2026-09-02)¶
No rebase impact against upstream Netflix: core/src/feature/x86/adm_avx2.c and core/src/feature/x86/adm_avx512.c are fork-added SIMD implementations and have no counterparts in upstream Netflix/vmaf.
Invariants preserved: - Fully bit-exact outputs against scalar integer_adm.c across all 11 AVX2 and AVX-512 functions (adm_decouple_*, adm_dwt2_*, adm_cm_*, i4_adm_cm_*, adm_csf_*, i4_adm_csf_*, adm_csf_den_*). - All 11 kernel functions maintain // NOLINTNEXTLINE(readability-function-size) with inline ADR-0138/0139 and ADR-0141 citations to preserve register allocation and vector reduction order. - Macro parameters across threshold and accumulation macros (ADM_CM_THRESH_*, I4_ADM_CM_THRESH_*, ADM_CM_ACCUM_ROUND*) are strictly parenthesized. - Pointer arithmetic offsets explicitly cast to (ptrdiff_t) to prevent implicit widening warnings on integer products. - Dead print_* debug macros in adm_avx512.c removed; #include "adm_avx2.h" added to adm_avx2.c for internal linkage consistency with public declarations.
refactor/test-model-tidy-clean — clang-tidy clean test_model.c and test_output.c (2026-09-02)¶
Upstream-mirror files touched: core/test/test_model.c and core/test/test_output.c. Both files were brought to 0 clang-tidy warnings without altering, deleting, or skipping any Netflix test assertion. When rebasing against upstream changes to these tests, note the following structural reorganizations: - core/test/test_output.c: - Assertion groups split into static check helpers: - check_csv_basic_output, check_csv_subsample_output (from test_csv_basic, test_csv_subsample_and_custom_format) - check_sub_basic_output (from test_sub_basic) - check_xml_basic_structure, check_xml_basic_metrics, check_xml_basic_output (from test_xml_basic) - check_json_basic_structure, check_json_basic_pooled, check_json_basic_aggregates, check_json_basic_output (from test_json_basic_and_format) - check_json_nan_inf_output (from test_json_nan_and_inf) - check_json_empty_collector_output (from test_json_empty_collector) - check_write_output_json (from test_write_output_json_path) - check_write_output_format (from test_write_output_with_format_custom) - check_pic_cnt_zero_json, check_pic_cnt_zero_xml (from test_write_output_pic_cnt_zero) - Concurrency/portability: replaced getenv("TMPDIR") with P_tmpdir fallback in make_temp_path(). - Memory leak fixes: ensured allocated file buffers (out) are freed on every return path. - Test runner: split run_tests into run_output_tests_part1 and run_output_tests_part2. - core/test/test_model.c: - NULL modernized to nullptr (C23 standard across fork). - #include "model.c" annotated with ADR-0278 / ADR-0141 NOLINT comment for white-box static model testing. - test_model_feature split into test_model_feature_step1, test_model_feature_step2, and check_model_feature_entry. - test_model_set_flags decomposed into test_model_set_flags_transform_and_clip, test_model_set_flags_default_opts, check_model_neg_feature_opts, and test_model_set_flags_neg_opts. - Buffer allocation lifetime: free(buf) placed immediately after vmaf_read_json_model_from_buffer and vmaf_read_json_model_collection_from_buffer parses. - String formatting: replaced variadic append_fmt with bounded append_str and append_uint, removing <stdarg.h> and avoiding VAList analyzer false positives. - JSON builder complexity: extracted append_65_feature_names, append_65_feature_slopes_intercepts, and check_65_feature_model for test_json_model_allows_more_than_64_features; extracted build_11_knot_json and check_11_knot_model for test_json_model_allows_more_than_10_knots. - Test runner: replaced linear macro expansion in run_tests with table-driven test_cases[] array and run_tests iterator.
feat/go-ort-runner — vmafx-ort-runner built in-tree (2026-09-02, ADR-1134)¶
No upstream code impact: cmd/vmafx-ort-runner/, pkg/ai, pkg/libvmaf, the Go stages of dev/Containerfile and .github/workflows/go-ci.yml are fork-local with no Netflix/vmaf counterpart. Rebase-sensitive contracts:
- The runner's wire format IS
pkg/ai.Registry.Infer's argv (--model <path> --inputs '<JSON array>') and its stdout (one JSON array line).cmd/vmafx-ort-runner/main_test.goandpkg/ai/infer_runner_test.gopin the two halves; change them together. Exit codes 0/1/2/3 are part of the contract (pkg/aiand the usage page both key onexit status 3= libvmaf without ONNX Runtime). pkg/libvmaf.DNNSession.Predictwith an empty input name binds positionally (NULLVmafDnnInput.name). Do not "simplify" it back to an unconditionalC.CString(inputName): that binds to an input literally named""and breaks the runner against every shipped predictor.go-ci.ymlinstalls the ONNX Runtime tarball and builds libvmaf with-Denable_dnn=enabledso the real-ORT branches ofpkg/libvmaf/dnn_test.go,pkg/ai'sTestInfer_RealRunnerand the runner smoke execute; dropping the install step turns them back into silent skips.dev/Containerfile's go-build stage asserts seven binaries (test -x /out/vmafx-ort-runner) and the dev-mcp stage smoke-runs the runner afterCOPY --from=go-build; keep both in step when acmd/is added or removed.renovate.jsontracksORT_VERSIONingo-ci.ymlwith the same regex manager asdev/Containerfile.
renovate/pypi-aiohttp-vulnerability — aiohttp security floor (2026-08-31)¶
No rebase impact: mcp-server/vmaf-mcp/, docs/mcp/, and the changelog fragment are fork-local and have no Netflix upstream counterparts. Preserve the aiohttp>=3.14.3 minimum when resolving future dependency refreshes: it is the first release outside GHSA-cq5v-8q36-5273 / CVE-2026-69244's affected range. No static-file or follow_symlinks compatibility exception is needed; the MCP HTTP transport registers dynamic routes only.
fix/e2e-k8s-runtime-contract — execute the real chart runtime (2026-08-31)¶
.github/workflows/e2e-k8s.ymlmust passtarget: node-cputo the nodedocker/build-push-actionstep.docker/Dockerfile.nodeends with thenode-syclstage, so an omitted target silently builds Intel runtime layers; the formerBACKEND=cpubuild argument was undeclared and had no effect. The same workflow must buildDockerfile.go-server --target go-serverand export/load the operator, node, and servere2e-testtags before Helm runs. The node builder must copymodel/.into/dist/model/and assert/dist/model/vmaf_v0.6.1.json; copying the directory itself creates a nested model root that disagrees withVMAFX_MODEL_DIR.test/e2e/kind-cluster.shapplies CRDs directly. Do not restore the old full Helm “CRD install” plus fallback: it launched the default server before its image was loaded and hid the rollout failure behind a successful CRD apply. Every create/reuse, apply, kuttl invocation, diagnostic read, score, and teardown is coupled to one absolute dedicated kubeconfig. Preserve the exactkind-${KIND_CLUSTER_NAME}current-context and loopback-server guard; never fall back to a process-wide Kubernetes context or suppress teardown failure.test/e2e/kuttl-tests/01-chart-cpu-score/is the executable integration boundary. It installs the chart's default Deployment workload on CPU, mounts validated Y4M fixtures, and requires a finite real/v1/scoreresponse. The chart Service and server Pod templates must shareapp.kubernetes.io/component: serverso an enabled operator's metrics port cannot become a scoring endpoint. ADR-1353 (1.0.0-rc.2) later added the discriminator to the server Deployment/StatefulSet selectors as well, with a documented one-time upgrade step; keep it there. Do not restore the removed Pod-creation, operator-heartbeat, MinIO/rclone, or trainer-sidecar cases unless the production reconcilers first implement and provision every asserted prerequisite.scripts/ci/test_e2e_runtime_contract.pyenforces these couplings in the always-on Rules workflow and again before the gated image build; never leave it only behind the E2E trigger gate.pkg/libvmaf/libvmaf.gomust pass subprocess models with the CLI parameter grammar-m path=/absolute/model.json. A bare path is rejected by the CLI parser and turns every file-backed server score into HTTP 500; preserveTestScore_PassesModelAsCLIPathParameteracross scorer or parser rebases..github/workflows/security-scans.ymlmust keepgithub.event_namein its concurrency group. Bothscheduleandpushuserefs/heads/master; a ref-only group lets the weekly scan cancel master-push CodeQL (or vice versa). Preserve same-eventcancel-in-progress: trueand the always-onscripts/ci/test_security_workflow_contract.pyguard together. Netflix upstream has none of these fork-local files, so there is no upstream conflict; preserve the explicit targets and executable runtime boundary during fork-local CI edits.
fix/sanitizers-meson-c23 — Cython extern follows mem.c -> mem.cpp (2026-08-30)¶
compat/python-vmaf/core/adm_dwt2_cy.pyx— rebase-sensitive. Upstream Netflix still haslibvmaf/src/mem.cand its.pyxstill text-includes it. The fork renamedlibvmaf/tocore/(ADR-0700) and converted that TU to C++23mem.cpp(#1133), so the fork's extern readscdef extern from "mem.h". An upstream sync touching this.pyxwill conflict — keep the fork's header-based extern; reverting to a.ctext-include reintroduces a failure that only shows up in the tox legs, never in a meson build.python/setup.py— fork-local: appends../core/src/mem.cppto the extension sources and setslanguage="c++"for the link driver.
fix/sanitizers-meson-c23 — meson from PyPI + declared c23 floor (2026-08-30)¶
core/meson.build— rebase-sensitive.meson_versionraised from'>= 0.58.0'to'>= 1.4.0'. This is a fork-local edit to the upstream project declaration, so an upstream sync that touches theproject()call will conflict here. Keep the fork's>= 1.4.0: it is load-bearing for the fork'sc_std=c23default option (ADR-0692), which upstream does not set. If a future sync ever dropsc_std=c23, this pin may be relaxed back to upstream's value..github/workflows/*.yml— fork-local CI only, no upstream counterpart. 15apt-get install … mesonsites replaced withsudo pip3 install --break-system-packages --quiet meson. Purely additive against upstream, no rebase impact.
fix/precommit-master-green — isort retired in favour of ruff (2026-08-30)¶
pyproject.toml,.pre-commit-config.yaml— fork-local tooling config; upstream Netflix/vmaf has neither a ruff nor an isort configuration, so an upstream sync cannot conflict here..github/workflows/required-aggregator.yml+lint-and-format.yml— paired invariant, not rebase-sensitive but edit-sensitive. The aggregator matches required checks by exact job name. ThePython Lintjob was renamed toPython Lint (Ruff + Black + mypy); both files must always be changed together. Renaming one alone leaves the aggregator waiting on a check name that never registers, which blocks every PR rather than failing loudly.
feat/vmafx-tune-go-fast — Phase A.5 fast path ported to Go (2026-08-30)¶
All fork-added surfaces; no upstream-mirror files touched, so nothing here conflicts with a Netflix upstream sync.
- New Go packages —
pkg/fast/(fast-path search + pipeline + proxy seam - Python-
reprJSON encoder),pkg/scorebackend/(libvmaf backend detection and strict selection),pkg/conformal/(split-conformal and CV+ prediction intervals). All three are Go ports of fork-local Python modules undertools/vmaf-tune/src/vmaftune/(fast.py,score_backend.py,conformal.py); the Python side is untouched and remains canonical until the ADR-0703 sunset. - New module dependencies —
github.com/c-bata/goptunav0.9.0 (MIT; a Go implementation of Optuna's TPE sampler) and its only transitive need,gonum.org/v1/gonumv0.17.0.go mod tidyadds exactly these two lines; the gorm / mysql / postgres requirements in goptuna's owngo.modare pruned because only the root package andgoptuna/tpeare imported. pkg/encoderadditive fields —EncodeParams.InputArgs,EncodeParams.OutputPathandEncodeResult.OutputSizeBytes. All optional;runEncodekeeps its previous behaviour when they are zero, socompareandladderare unaffected. If a rebase conflicts inrunEncode, keep theInputArgssplice before-iand theOutputPathbranch around theos.CreateTempblock — raw-YUV probing depends on both.cmd/vmafx-tune/cmd/root.go—fastmoved out of the loud-fail stub slice into the ported list, andExecutegrew afastExitCode(err)check so the subcommand's 2 / 3 exit contract survives cobra's blanket exit 1. A rebase that re-adds{"fast", ...}to the stub slice would shadow the real command; drop the stub entry, not thenewFastCmd()registration.- Python-side divergences are deliberate — the Go probe path decodes container encodes to raw YUV, resolves
integer_-prefixed libvmaf pooled keys, reads the encoder vocabulary and the StandardScaler from the model sidecar, and applies that scaler. Each of those corrects a defect invmaf-tune fast(seechangelog.d/added/vmafx-tune-go-fast-subcommand.md). Do not "restore parity" by reverting them; if the Python is fixed later, the two converge on the Go behaviour. - Do not lift the proxy port guard on rebase —
ORTProxy.Scorerefuses a multi-port ONNX graph on purpose. See invariant 15 incmd/vmafx-tune/AGENTS.md.
feat/vmafx-tune-go-encoder-introspection (ADR-0770)¶
No rebase impact: pure Go additions plus one Markdown doc page. No upstream C or Python file is modified — the Python vmaf-tune tree is read as the port reference but left byte-unchanged, and the Netflix golden assertions under python/test/ are untouched.
Files added: pkg/benchmark/benchmark.go + benchmark_test.go + testdata/, pkg/codecadapter/codecadapter.go + codecadapter_test.go, pkg/encodeprofile/{profile,encode,pycompat}.go + {encodeprofile,encode}_test.go + testdata/, internal/pyjson/pyjson.go + pyjson_test.go, cmd/vmafx-tune/cmd/{benchmark,encodeprofile,exitcode}.go + {benchmark,encodeprofile}_test.go, changelog.d/added/vmafx-tune-go-benchmark-encode-profile.md.
Files modified: cmd/vmafx-tune/cmd/root.go (register benchmark + encode-profile, drop their stubs, route the exit status through exitCodeOf), cmd/vmafx-tune/cmd/root_test.go (ported/stub lists), cmd/vmafx-tune/AGENTS.md (invariants 13–15), docs/usage/vmafx-tune-go.md (new subcommand sections), docs/rebase-notes.md (this entry).
Parity invariant to preserve on any future edit. The golden files under pkg/benchmark/testdata/ and pkg/encodeprofile/testdata/ were generated by running the Python vmaftune modules over the committed fixtures, not by recording the Go output. If a future change alters either package's output, regenerate them the same way (drive vmaftune.benchmark / vmaftune.encoder_profile + vmaftune.encode directly) rather than blessing whatever Go now emits — otherwise the tests stop proving parity and start merely asserting self-consistency.
fix/codeql-quality-batch — code-scanning hygiene (2026-06-27)¶
Small behaviour-neutral quality fixes. Upstream-mirror touches to re-apply on the next upstream sync: removed an unused from collections.abc import Hashable in compat/python-vmaf/tools/decorator.py, and unused pytest/tempfile imports in two compat/python-vmaf/tests/ files. Fork files: core/tools/vmaf.cpp (2-label switch→if), and new include guards on core/src/feature/moment.h / alias.h. The bulk of the code-scanning backlog was resolved by dismissal (verified false-positive/intentional via the GitHub code-scanning API), not code change — see docs/state.md T-CODEQL-QUALITY-BATCH.
fix/round4-ffmpeg-patches — libvmaf_sycl filter leak + QSV NULL guard (2026-06-27)¶
ffmpeg-patches/0005-libvmaf-add-libvmaf-sycl-filter.patch gained two fixes in its libavfilter/vf_libvmaf.c hunk (new-count 335→353): uninit_sycl now calls vmaf_sycl_state_free() after vmaf_close(), and do_vmaf_sycl NULL-guards the QSV mfxHDLPair chain. The patch was regenerated surgically (only those two + blocks + the hunk-header recount; the configure/Makefile/allfilters hunks are byte-unchanged). Do NOT let git format-patch/git am --3way regenerate the whole patch — that fuzzed the configure probe >= 3.0.0→2.0.0 / libvmaf/libvmaf_sycl.h→libvmaf_sycl.h. Verified by a full 16-patch git apply --3way series replay against n8.1.1. Keep these two + blocks on re-sync. Finding #21 (a redundant but idempotent check_pkg_config libvmaf_sycl configure probe) was intentionally left in place.
fix/round4-cli-build-go — round-4 audit bug-fix bundle (2026-06-27)¶
All fork-added/fork-modified surfaces (no upstream-mirror conflict risk):
core/src/meson.build— added_x86_simd_strict_fp_extrato thex86_float_adm_avx2/x86_float_adm_avx512carve-outs (icx fp-model parity; no-op on gcc/clang). Keep when re-syncing the meson SIMD carve-out block.core/tools/vmaf.cpp,core/tools/vmaf_bench.c— fork-added CLI timing helpers (wall_time_s/now_ms): zero-init + cachedstaticQPF frequency.core/tools/meson.build— comment-only path fix.pkg/libvmaf/paths.go— fork-added Go MCP path allowlist (AllowedRootsfail-closed viadiscoverRepoRoot).RepoRoot()signature unchanged.
fix/round4-c-bundle — round-4 audit bug-fix bundle (2026-06-27)¶
Audit-derived bug fixes; several touch upstream-mirror files, so the next upstream sync must preserve these hunks (they are not in Netflix/vmaf):
core/src/feature/ciede.c— init returns-EINVAL(not-ENOMEM) for an unsupported bitdepth; close uses two independentvmaf_picture_unrefguards (was a conjunctive guard that leakeds->refon partial alloc).core/src/feature/cambi.c— close guardsvmaf_picture_unrefons->pics[i].refso never-allocated slots don't poisonerr.core/src/feature/integer_ssim.c— comment-only correction of the GPU-twin note (theconst double smfix itself is unchanged).core/src/read_json_model.c— partial-collection teardown on the non-string-key early return inmodel_collection_parse_loop.core/src/model.c—vmaf_model_collection_appendshort-name path nowgoto fail_modelto freemc->model.
Fork-added files (no upstream-sync concern): core/src/feature/cuda/speed_*, core/src/dnn/ort_backend.c.
gorust-rederive — bound GPU/AI probe subprocesses + Rust -sys picture double-free footgun (2026-06-27)¶
Rebase impact: none on upstream Netflix/vmaf — all fork-local. Every touched surface (pkg/gpu/detect.go, pkg/ai/infer.go, bindings/rust/vmafx-sys/src/safe.rs, core/src/meson.build) is fork-added Go/Rust/build code with no upstream counterpart, so a future /sync-upstream sees no conflict here.
Cross-crate invariant: vmafx-sys::safe::VmafContext::read_pictures and the higher-level vmafx::Context::read_pictures must keep aligned picture-ownership semantics — both consume pictures by value (move) and neither manually unrefs on the error path (the libvmaf contract takes ownership for the call's duration; a second unref is a use-after-free against a CUDA-enabled libvmaf). The vmafx crate side was settled by PR #1056 (round-3 R3-2); this change brings the -sys crate to the same contract. Do not revert either to a borrowing signature or re-add an error-path unref. See bindings/rust/vmafx-sys/AGENTS.md.
Note: this change also restores docs/state.md (truncated to 0 bytes by PR
1055, the pelorus ABI re-vendor) and docs/rebase-notes.md itself (truncated¶
to 0 bytes by PR #1060, the FMA-ADM fix) — two unrelated accidental wipes on master that are recovered here from their last-good blobs.
feat/pelorus-abi-minor3-consume — re-pin vendored Pelorus ABI to minor-3 + consume PEL_SEC_COMPLEXITY (2026-06-27)¶
Rebase impact: none on upstream Netflix/vmaf — all fork-local (ADR-1120, builds on ADR-1113 + ADR-1118). Cross-repo ABI parity invariant: the vendored Pelorus interop mirror is single-sourced in VMAFx/pelorus (ADR-0103) and pinned by PELORUS_VENDOR_SHA in scripts/sync-pelorus-interop.sh — now 818d844 (ABI 1.3, was 835e097 / ABI 1.0). The drift-guard CI gate (sync-pelorus-interop.sh without --update) fails on any divergence from the pin, so a future maintainer must NOT hand-edit the vendored files (core/include/libvmaf/pelorus/*.h, core/src/interop/pelorus_*.c, and the body of core/test/test_pelorus_interop.c from its first vendored #include on) — fix defects upstream in pelorus and re-vendor via --update. - Manifest invariant: the script's manifest array, core/src/meson.build, and the test_pelorus_interop target in core/test/meson.build must stay in lockstep with the pelorus source set. Minor-3 added pelorus/denoise.h + pelorus_denoise_params.c + pelorus_qp_report_csv.c; the last is REQUIRED to link the fixture (pel_x265_csv_parse). - --update now re-vendors the fixture body (previously only the six manifest files), preserving the Lusoris-authored header before the first vendored include. The drift check compares the body whitespace-insensitively. - Vendored files are lint/format-excluded by prefix glob (core/src/interop/pelorus_, core/include/libvmaf/pelorus/) in .pre-commit-config.yaml, Makefile, and scripts/ci/assertion-density.sh — new vendored files matching those prefixes are covered automatically; no new exclusion entries are needed. - Complexity modulation (golden-isolation invariant, rebase-sensitive): perceptual_weight.c::complexity_modulation MUST return exactly 1.0 when PEL_SEC_COMPLEXITY is absent or complexity is non-finite — that is what keeps the no-side-data golden path bit-exact (Netflix 576×324 pair = 76.667831). The guard is test_complexity_modulates_weight/_grid_zero.
fix/bughunt-cuda — CUDA pinned-buffer leaks, motion SAD precision, errno fidelity (2026-06-27)¶
Rebase impact: none on upstream — all fork-local. The CUDA backend (core/src/feature/cuda/) is a fork addition with no Netflix/vmaf counterpart. Touches only core/src/feature/cuda/{float_vif_cuda.c,float_adm_cuda.c, integer_motion_cuda.c,speed_temporal_cuda.c,speed_chroma_cuda.c}. No public header, CLI flag, meson-option, ffmpeg-patch, or Netflix golden-gate surface changes (the golden gate is CPU-only; SpEED is not in the golden pairs). The leak fixes fire only in close_fex_cuda / init-error paths (no success-path behaviour change); the errno fixes only change the value returned on an already-failing CUDA error path (-EIO → the mapped errno); the integer_motion_cuda precision change brings GPU SAD output closer to the CPU double-precision reference (GPU-only, not bit-exact with CPU by design). Rebase-sensitive invariant for the next syncer: in speed_temporal_cuda.c / speed_chroma_cuda.c, every fail: label reached from CHECK_CUDA_GOTO must return _cuda_err; (the macro-mapped errno), not a literal -EIO — matching the CHECK_CUDA_RETURN convention in cuda_helper.cuh. The two manual cuMemcpyDtoH / cuCtxPushCurrent boolean checks deliberately keep their literal -EIO.
fix/bughunt-core-engine — core-engine error-path fixes (2026-06-27)¶
Rebase impact: low — three upstream-mirror files touched, all on error/cleanup paths. core/src/libvmaf.c, core/src/feature/feature_collector.c, and core/src/model.c are upstream-mirror files with Netflix counterparts, so a future /sync-upstream may produce small conflicts here. The changes are fork-local divergences confined to failure paths: - threaded_read_pictures_batch (libvmaf.c) is a fork-added threaded-batch helper (not in upstream), so its enqueue-failure unref fix carries no upstream conflict risk. Two adjacent doc comments in the same function were tightened to keep it under the fork's readability-function-size LineThreshold (60) — purely cosmetic, no behaviour change. - aggregate_vector_append (feature_collector.c): one-line -EINVAL→-ENOMEM on the feature-name malloc-failure path. Mirrors the fork's own .cpp twin. If upstream rewrites this allocation, prefer the -ENOMEM semantics. - vmaf_model_collection_append (model.c): grow-path realloc failure no longer takes the shared fail: label (which nulls *model_collection); it returns -ENOMEM inline. Rebase-sensitive invariant: only the fresh-allocation failures may null the caller's out-param; the grow path must leave the still-valid existing collection (and the caller's handle) intact. No public header, CLI flag, meson-option, ffmpeg-patch, or Netflix golden-gate surface changes; all three edits fire only on malloc/realloc/enqueue failure, so success-path scores are unchanged (golden gate verified green).
fix/bughunt-mcp — MCP Go↔Python parity + HTTP hardening (2026-06-27)¶
no rebase impact: edits the fork-only MCP servers (cmd/vmafx-mcp/{main.go,impl.go,impl_direct.go} + new cmd/vmafx-mcp/http_security.go, mcp-server/vmaf-mcp/src/vmaf_mcp/http_transport.py) + tests + docs/state.md + changelog. No libvmaf C-API / CLI / meson_options.txt / public-header change → no ffmpeg-patch impact. Rebase-sensitive invariant — HTTP transport security parity (cmd/vmafx-mcp/AGENTS.md invariant #13): the Go securityMiddleware / bind logic (http_security.go) and the Python _make_security_middleware / _resolve_bind_host (http_transport.py) MUST share the same ADR-0967 env contract (VMAFX_MCP_HTTP_TOKEN constant-time bearer, VMAFX_MCP_HTTP_NO_AUTH=1 opt-out, refuse-all-401-when-neither-set, 4 MiB body limit, VMAFX_MCP_HTTP_BIND default 127.0.0.1). Precision-default parity (ADR-0119 / ADR-1117): both servers default vmaf_score precision to legacy (%.6f) on every transport / dispatch path. A rebase touching either server must keep both in lock-step.
fix/mcp-probe-backend-required — MCP probe_backend required-arg message (2026-06-20)¶
Rebase impact: none on upstream — fork-local. The MCP server (mcp-server/vmaf-mcp/) is a fork addition with no Netflix/vmaf counterpart. One-line change in _call_tool_dispatch's probe_backend branch (removes a redundant explicit guard, relies on the existing KeyError→ValueError wrapper). No public C-API / CLI / header impact.
fix/speed-extractor-oob-deadlock-heap-corruption — GPU SpEED covariance + eigenbasis correctness + safety (2026-06-20)¶
Rebase impact: none on upstream — all fork-local. The SpEED feature (speed_chroma / speed_temporal) and all of its GPU backends are fork-additions with no Netflix/vmaf counterpart. Touches the fork-only GPU extractors (core/src/feature/cuda/{speed_chroma_cuda.c,speed_temporal_cuda.c, speed/speed_score.cu}, core/src/feature/hip/{speed_chroma_hip.c, speed_temporal_hip.c,speed/speed_score.hip}, core/src/feature/sycl/{speed_chroma_sycl.cpp,speed_temporal_sycl.cpp}), the fork-only core/src/feature/speed.c CPU host (init-return propagation only — the global covariance math itself is unchanged), and the GPU parity test fixture. No public header, CLI, meson-option, ffmpeg-patch, or Netflix golden-gate surface changes (SpEED is not in the golden pairs). Rebase-sensitive invariant for the next syncer: the GPU means/cov kernels must stay on the CPU's global covariance formulation (means[25] over the full phase-shifted submatrix, NOT per-tile means[25*num_blocks]), and the ref/dis paths must keep separate covariance + eigenvalue bases — recorded in core/src/feature/cuda/AGENTS.md and verified by test_cuda_speed_{chroma,temporal}_parity at 1e-4. If a future change touches any one backend's kernels, mirror it across all four (CPU + CUDA + HIP + SYCL).
fix/audit-runtime-bugs-batch — 18 audit runtime bugs: SYCL/CUDA init leaks, AI crash-hardening, MCP parity (2026-06-20)¶
Rebase impact: none on upstream. All 18 fixes are fork-local and touch only fork-added files with no upstream Netflix/vmaf counterpart: core/src/feature/sycl/integer_motion_sycl.cpp, core/src/feature/cuda/integer_ssim_cuda.c, core/src/feature/cuda/integer_vif_cuda.c, core/src/feature/ssimulacra2.c (fork-added SSIMULACRA2 extractor), cmd/vmafx-mcp/impl.go, and five ai/ extraction/training scripts (bvi_dvc_to_full_features.py, extract_full_features.py, konvid_to_full_features.py, train_fr_regressor_v2.py, vmaf_train/datamodule.py). No public header, CLI flag, meson-option, ffmpeg-patch, or golden-gate surface changes; all C/SYCL/CUDA edits fire only on already-failing error/OOM paths so success-path behaviour and scores are unchanged. Rebase-sensitive note for the next person syncing: the cmd/vmafx-mcp/impl.go change deletes the last vulkan backend reference in the Go MCP server to keep it byte-compatible with the Python MCP server after the Vulkan removal (ADR-0726); if a sync re-introduces a vulkan keyword in either MCP server, both must move together. No new rebase-sensitive invariants worth a dedicated AGENTS.md entry beyond the existing SYCL/CUDA error-path notes.
fix/sycl-psnr-hvs-chroma-ceiling — SYCL psnr_hvs odd-dimension chroma geometry (2026-06-20)¶
Rebase impact: none on upstream. Fork-local one-line correctness fix in the fork-only SYCL feature extractor core/src/feature/sycl/integer_psnr_hvs_sycl.cpp, which has no upstream Netflix/vmaf counterpart. init_fex_sycl now derives the 4:2:0 / 4:2:2 chroma plane dims with ceiling division ((w + 1U) >> 1) instead of floor (w >> 1), matching picture.c / the CPU reference / the CUDA + HIP twins. No public-header, CLI, meson-option, ffmpeg-patch, or golden-gate surface changes; even-dimension behaviour is byte-identical to before. Rebase-sensitive note for the next person syncing: the picture allocator's ceiling subsample convention ((dim + ss) >> ss) is the single source of truth for chroma plane dims — any new GPU feature extractor that re-derives plane dimensions in its own init must use the ceiling form, not floor; the floor form only agrees on even dimensions and silently drops the last chroma block strip otherwise. This is the same class of bug as the PSNR and Vulkan chroma ceiling fixes already in tree.
fix/metal-drain-motion2 — Metal end-of-stream drain + frame-0 motion2 (2026-06-20)¶
Rebase impact: none on upstream. All changes are fork-local (the Metal backend has no upstream Netflix/vmaf counterpart) plus one additive bit in the shared flush path. Touches: core/src/feature/feature_extractor.h (adds VMAF_FEATURE_EXTRACTOR_METAL = 1 << 7 to the VmafFeatureExtractorFlags enum — a fork-added enum; bit 7 is the next free slot after the fork's HIP bit 6), core/src/libvmaf.c (a new #ifdef HAVE_METAL drain branch in flush_context_serial, gated so non-Metal builds are byte-unchanged), and 8 fork-only core/src/feature/metal/*.mm extractors (set the new flag; + float_motion_metal.mm collect-index fix).
Rebase-sensitive notes for the next person syncing: - flush_context_serial is a fork-local rewrite of the upstream flush. If an upstream sync re-touches the end-of-stream flush, the fork's per-backend drain blocks (CUDA / HIP / SYCL / Metal) must be re-applied — each GPU backend whose extractors carry a VMAF_FEATURE_EXTRACTOR_<BACKEND> flag needs its pending gpu_pending final-frame collect() drained before its flush() runs, or the last frame's score is dropped. Do not drop the Metal branch. - Frame-0 motion2 contract. Every motion-family extractor (CPU + all GPU twins) appends motion2 = 0.0 at index 0 and a no-op at index 1; index ≥ 2 emits min(prev, cur) at index − 1. float_motion_metal now matches this exactly — keep it aligned with integer_motion_metal and the HIP / CUDA twins on any future motion2 change (cross-backend invariant, see core/src/feature/metal/AGENTS.md). - Darwin-only. Not buildable / not exercised on the Linux dev or CI lane; re-validate on Apple Silicon after any upstream flush-path sync.
fix/k150k-training-data-integrity — fail-loud on empty-frame clips + MOS-join key mismatch (2026-06-20)¶
Rebase impact: none on upstream. All changes are fork-local in ai/scripts/extract_k150k_features.py (a fork-added training script with no upstream Netflix counterpart). Two defensive guards: _process_clip raises on an empty frame list instead of writing an all-NaN row + marking the clip done; the MOS-label join gains an mp4.stem fallback + an up-front coverage hard-fail. Invariants recorded in ai/AGENTS.md (do not revert the lookup to a single mos_map.get(clip_name, NaN); keep the staging-first .done ordering). No C-library, ABI, golden-data, or ffmpeg-patch surface touched.
fix/speed-gpu-registry — restore orphaned GPU SpEED registrations + delete dead feature_extractor.c (2026-06-19)¶
Rebase impact: none on upstream. All changes are fork-local. Touches the fork-only registry file core/src/feature/feature_extractor.cpp (adds six externs + array entries under the existing #if HAVE_{CUDA,SYCL,HIP} blocks), deletes the fork-only dead twin core/src/feature/feature_extractor.c (orphaned by PR #875's .c→.cpp split; meson compiled only the .cpp), and adds by-name resolution asserts to core/src/feature/../test/test_feature_extractor.c. Rebase-sensitive note for the next person syncing: there is now exactly ONE registry file (feature_extractor.cpp); if an upstream/Netflix sync re-introduces a feature_extractor.c it must be reconciled into the .cpp, not kept alongside it — the split-brain is what this fix removes. Stale feature_extractor.c references remain in ~40 sibling source comments (Metal .mm, HIP .c) and in historical docs/adr/* / docs/research/* (audit trail — do NOT rewrite); the live comment sweep for the non-ADR consumer files is deferred to the RC LOW doc-hygiene PR and coordinated with the in-flight Metal PR #986.
fix/sycl-init-leaks-exception-safety — SYCL init error-path + exception-boundary hardening (2026-06-19)¶
Rebase impact: none on upstream — fork-local SYCL error-path + exception-boundary hardening. Touches only fork-added SYCL sources (core/src/feature/sycl/integer_adm_sycl.cpp, core/src/feature/sycl/integer_vif_sycl.cpp, core/src/sycl/common.cpp, core/src/sycl/dmabuf_import.cpp), none of which have an upstream Netflix/vmaf counterpart. No public header, CLI, meson-option, ffmpeg-patch, or golden-gate surface changes; success-path behaviour is unchanged (cleanup/exception handling only fires on already-failing paths).
fix/cuda-init-submit-leaks — CUDA error-path resource frees (2026-06-19)¶
Rebase impact: none on upstream — fork-local CUDA error-path hardening. Touches only fork-added CUDA feature extractors under core/src/feature/cuda/ (integer_ms_ssim_cuda.c, integer_psnr_hvs_cuda.c, ssimulacra2_cuda.c, speed_chroma_cuda.c); adds NULL-guarded frees / cleanup-goto routing on init + submit failure paths only. No public-header, meson, CLI, ffmpeg-patch, or golden-gate impact, and success-path behaviour is unchanged.
fix/hip-chroma-mcp-parity — psnr_hip enable_chroma option + MCP Go/Python parity (2026-06-20)¶
Rebase impact: none on upstream — all fork-local. Touches the fork-only HIP extractor (core/src/feature/hip/integer_psnr_hip.c: add an enable_chroma VmafOption), the fork-only Go MCP server (cmd/vmafx-mcp/impl.go + impl_test.go: drop the dead vulkan backend value, drop the unsupported --format from tune-per-shot), and the fork-only Python MCP server (server.py: probe_backend ValueError guard). No upstream Netflix/vmaf file is touched; no public C-API/CLI surface changes (the enable_chroma option already exists on the CPU/CUDA psnr twins).
fix/tox-py314-scipy-118 — tox env py311→py314 (2026-06-20)¶
Rebase impact: none on upstream — fork-local CI config only. Touches python/tox.ini (envlist py311→py314, matching the CI setup-python 3.14.5, since the fork's requirements.txt deps now require ≥3.12) plus a changelog fragment + state.md row. No source or test code changed.
feat/upstream-v1.0.16-models (2026-06-20)¶
Rebase impact: low (additive model data + one C registry block + one meson embed block; no public-header / CLI / ffmpeg-patch / golden-gate change). Verbatim port of Netflix upstream commit 4718b4f5f ("Add VMAF v1.0.16 SDR models, documentation, and tests"). Because it is a pure upstream port, it is exempt from the ADR-0108 six-deliverable rule (CLAUDE §12 r11); the changelog fragment + this rebase note are still provided.
What was ported and how it was adapted to the fork's diverged layout (ADR-0700 libvmaf/ → core/):
libvmaf/src/model.c→core/src/model.c: added the 8externdecls + 8built_in_models[]registry entries, mirroring the existingvmaf_v0.6.1/vmaf_4k_v0.6.1negidiom byte-for-byte (the fork's struct is the sameVmafBuiltInModel {version, data, data_len}).libvmaf/src/meson.build→core/src/meson.build: added twoforeachblocks embedding the v1.0.16 + v1.0.16_hfr JSONs via the samexxd -i -n src_@PLAINNAME@custom_targetthe fork already uses for the v0 models. The v1 models live in their ownmodel/vmaf_v1.0.16{,_hfr}/subdirectories, so a dedicated dir prefix is used (matching upstream).model/vmaf_v1.0.16/*.json+model/vmaf_v1.0.16_hfr/*.json(8 files): copied verbatim viagit checkout 4718b4f5f -- ….python/test/vmaf_v1_quality_runner_test.py: copied verbatim (path NOT renamed). 46 new golden assertions; no pre-existing assertion touched.- Upstream's
resource/doc/models_v1.md+ themodels.md→models_v0.mdrename + the README "News" line do not map: the fork has noresource/doc/model docs (it consolidated them underdocs/models/), so the new doc lives atdocs/models/v1.md(added to the mkdocs nav, with a cross-link fromdocs/models/overview.md). The README change is moot — the fork's README diverged and no longer linksresource/doc/models.md.
Deliberately NOT ported (rebase-sensitive — re-check on the next upstream sync): the upstream commit also bundled an unrelated feature-source reorg in meson.build (moving speed.c, common/convolution.c, vif_tools.c out of the float_enabled block into the always-on list). The fork already wires those sources differently, so applying the upstream hunk would conflict / double-list. If a future sync touches that region, reconcile against the fork's current libvmaf_feature_sources layout, not the upstream diff.
Known fork gap (load-bearing invariant): the 4 _hfr models embed and register but cannot be scored until motion_five_frame_window=true + motion_moving_average=true are implemented (the prev_prev_ref 5-frame plumbing deferred per ADR-0337). The 4 non-HFR models score correctly (1080p 3H == upstream golden VMAF 82.816059). Do not "fix" the HFR runtime error by deleting the option from the model JSONs — the JSONs are verbatim Netflix data; the fix is to land the 5-frame motion plumbing.
feat/golusoris-tune (2026-06-15)¶
Rebase impact: low (Go-only, additive + in-place rewrite of one binary's composition root). Phase-1 of the golusoris adoption (ADR-1119): migrates the cmd/vmafx-tune CLI from a hand-built cobra.Command root onto the golusoris clikit (cobra + fx) framework. Touches only cmd/vmafx-tune/* + its docs / changelog / rebase-notes; no C / meson / public-header / ffmpeg-patch / golden-gate impact, and the Python tools/vmaf-tune harness is untouched. cmd/vmafx-tune is non-cgo (no pkg/libvmaf import), so no libvmaf.so build is needed to compile or test it.
Files: new cmd/vmafx-tune/cmd/golusoris.go (the withGolusoris adapter + configOptions + levelledLogger); cmd/vmafx-tune/cmd/root.go rewritten to clikit.New + clikit.Command; compare.go / ladder.go / report.go subcommand builders re-wired through clikit and their run* functions now take (ctx, deps, flags); new cmd/vmafx-tune/cmd/root_test.go; existing tests updated for the new run* signatures; cmd/vmafx-tune/AGENTS.md invariants extended; docs/usage/vmafx-tune-go.md documents the VMAFX_LOG_LEVEL / VMAFX_LOG_FORMAT surface.
Rebase-sensitive invariants for follow-up PRs and any golusoris bump:
- clikit
WithFxis long-running, not one-shot.clikit.WithFxbuilds anfx.Appand callsapp.Run()(blocks until signal) and never surfaces anfx.Invokeerror as the exit code. One-shot tuning subcommands therefore useclikit.WithRunE(withGolusoris(fn)), wherewithGolusorisbuilds the graph frombootstrap.Base,fx.Populates the deps, runsfn, and returns its error. Do not "simplify" these toclikit.WithFx(golusoris.Core, fx.Invoke(fn))— the CLI would block and lose its exit code. fx.NopLoggeris deliberate. A one-shot CLI must not print fx provide/invoke/lifecycle chatter on every run;bootstrap.FxLogger()(which routes fx events onto the app logger) is for long-running services only. The injected*slog.Loggerstill carries domain diagnostics.levelledLoggercompensates for a golusoris v0.4.0 scoping gap. A root-scopefx.Replace(config.Options{EnvPrefix:"VMAFX_"})reaches root-scope consumers (our domain code reads the right config) but does not penetrate thegolusoris.logsubmodule's ownconfig.Optionsdependency, so the auto-built logger falls back to the defaultAPP_prefix and stays atLevelInfo.withGolusoristherefore addsfx.Decorate(levelledLogger)to rebuild the*slog.Loggerfrom the root config at theVMAFX_-configured level/format. Delete this decorator once golusoris makes the root config override penetrate submodules (track upstream alongside golusoris #234); theTestGolusorisInjection_ConfigDrivesLogLeveltest guards the behavior.VMAFX_env prefix.configOptions()setsEnvPrefix:"VMAFX_"to match the fork-wide env contract (ADR-1119). golusoris splits every underscore into the config delimiter, soVMAFX_LOG_LEVEL→log.level.
feat/golusoris-foundation (2026-06-14)¶
Rebase impact: low (Go-only, additive). Phase 0 of the golusoris fx framework adoption (ADR-1119). Adds github.com/golusoris/golusoris v0.3.1 to go.mod/go.sum (and the widened transitive closure — fx/dig, koanf, chi, river transitives; go build ./... + all test binaries compile clean, no version-skew since both repos already pinned identical shared-dep versions), one new package internal/app/bootstrap/bootstrap.go (Base fx module set + FxLogger), and an Info/Get() addition to pkg/version/version.go (the interim stand-in for golusoris#226). No C/meson/CLI/public-header change → no ffmpeg-patch impact, no golden-gate impact. No binary is migrated in this PR — the six cmd/vmafx-* composition roots are rewritten in the subsequent phased PRs (vmafx-server first; vmafx-controller gated on golusoris#225). Rebase-sensitive note for the follow-up PRs: each binary's fx.New(...) must lead with fx.Replace(config.Options{EnvPrefix:"VMAFX_"}) before golusoris.Core to preserve the VMAFX_ env contract, and the cgo libvmaf.Scorer provider must order its OnStop after the gRPC server's (drain before Close()). Docs: docs/adr/1119-* + fragment + _order.txt + README row, docs/research/1119-*, changelog.d/chore/1119-golusoris-foundation.md.
feat/metric-brisque (2026-06-14)¶
Rebase impact: low. Adds four fork-only files (core/src/feature/brisque.c, brisque_math.h, brisque_model.h, core/test/test_brisque.c) plus the vendored model + provenance (model/other_models/brisque_live.model, NOTICE-brisque, model/brisque_live_card.md) — no upstream twin (BRISQUE is fork-added; the model is the LIVE-lab allmodel bundled under a documented research-use exception, ADR-1115). The model is embedded into the binary at build time via an xxd -i Meson custom_target (the same mechanism libvmaf's JSON models use; brisque_model.h only declares the generated src_brisque_live_model[] / _len externs), so the giant byte array never enters the tree — keeping it under the 1 MB large-file gate. Additive registration: one extern VmafFeatureExtractor vmaf_fex_brisque; + one feature_extractor_list[] entry in core/src/feature/feature_extractor.cpp (the LIVE C++23 registry, NOT the dead feature_extractor.c twin), a model embed custom_target + one source line in core/src/meson.build, one executable() (linking the generated model TU) + one test() in core/test/meson.build. Edits docs/metrics/brisque.md (new), mkdocs.yml (BRISQUE nav row), docs/state.md, docs/rebase-notes.md, changelog.d/, core/src/feature/AGENTS.md (invariant note), testdata/scores_cpu_brisque.json (new), docs/research/1101-brisque-nr-metric.md (new), and the ADR index (docs/adr/1115-brisque-nr-metric.md + fragment + _order.txt + README row). CPU-only scalar extractor; no public C-API / ABI / CLI flag / meson_options.txt change → no ffmpeg-patch impact (reachable via the existing generic --feature path). First feature-extractor consumer of the vendored libsvm (core/src/svm.cpp/svm.h) — if a future upstream sync changes the libsvm parser/predict ABI, brisque.c's svm_parse_model_from_buffer + svm_predict calls must be re-checked alongside predict.c. Load-bearing invariants if the algorithm is ever touched: GGD (not AGGD) for the MSCN field, Gaussian sigma=7/6 (not 1.166), MATLAB antialiased bicubic (not INTER_CUBIC), the inline range arrays (not allrange), no output clamp — all required for parity with the bundled trained model (see core/src/feature/AGENTS.md and ADR-1115).
feat/metric-y-funque-plus (2026-06-14)¶
Rebase impact: low. Fork-only additive metric — no upstream twin. Adds two fork-only files (core/src/feature/y_funque_plus.c, core/test/test_y_funque_plus.c). Additive registration only: one extern VmafFeatureExtractor vmaf_fex_y_funque_plus; + one feature_extractor_list[] entry in core/src/feature/feature_extractor.cpp (the live C++23 registry — NOT the dead feature_extractor.c twin), a dedicated libvmaf_y_funque_plus_static_lib static_library() + one extract_all_objects() line in core/src/meson.build (mirrors the ssimulacra2 -ffp-contract=off carve-out), and one executable() + one test() in core/test/meson.build. Edits docs/metrics/y-funque-plus.md (new), docs/metrics/features.md (one new row), docs/state.md, docs/rebase-notes.md, changelog.d/, and the ADR index (docs/adr/1114-y-funque-plus-atoms.md + fragment + _order.txt + README.md row). CPU-only scalar extractor; no public C-API / ABI / CLI flag / meson_options.txt / public-header change → no ffmpeg-patch impact (CLAUDE §12 r14 N/A; reachable via the generic --feature path). Load-bearing invariants if the algorithm is ever touched: the Haar butterfly uses the pywt 'haar' convention cH=(a+b-c-d)/2, cV=(a-b+c-d)/2 (NOT the H/V-swapped form the design dossier text mistakenly listed — pywt was verified directly); the DLM numerator pools rest^3 WITHOUT abs while the denominator pools the ref detail WITH abs (pyr_features.py:54/61); the 2x downscale is OpenCV INTER_CUBIC (Keys cubic a=-0.75), the dominant cross-host parity risk — keep -ffp-contract=off.
feat/pelorus-sidedata-reader (2026-06-14)¶
Rebase impact: low-to-medium. Fork-only additive feature (ADR-1118), builds on the vendored Pelorus interop ABI (ADR-1113). Adds three fork-only files (core/include/libvmaf/perceptual_weight.h public C-API, core/src/feature/perceptual_weight.{c,h} the weight module + internal contract, core/test/test_perceptual_weight.c the golden-isolation test) — no upstream twin. Edits to shared files, all additive: - core/src/libvmaf.c — the rebase-sensitive one. Adds (1) two #includes, (2) a VmafPerceptualWeightStore perceptual; field at the tail of struct VmafContext (after dnn), (3) a vmaf_perceptual_weight_store_destroy call in vmaf_close after vmaf_ctx_dnn_free, (4) three new public entry points after vmaf_import_feature_score, and (5) the weighting branch inside vmaf_feature_score_pooled plus a new static pool_reduce helper + PoolAccumulators struct just above it. Load-bearing invariant: the no-side-data path through vmaf_feature_score_pooled MUST stay byte-identical to upstream — the weighted accumulators are only summed when vmaf_perceptual_weight_active() is true, and the MEAN/HARMONIC_MEAN reduce runs the literal upstream expression when weighting is inactive. A rebase that refactors this function must preserve that bit-exactness (the golden gate depends on it; test_perceptual_weight.c guards it). - core/src/meson.build — one source line (feature/perceptual_weight.c) in libvmaf_sources (NOT the feature static lib — it is a pooling helper, not a registered extractor). - core/include/libvmaf/meson.build — one install_headers entry. - core/test/meson.build — one executable() + one test(). - ffmpeg: ffmpeg-patches/0017-libvmaf-read-pelorus-sidedata.patch (new public C-API consumed by vf_libvmaf + new perceptual_weight AVOption → ffmpeg-patch impact per CLAUDE r14) appended at the tail of ffmpeg-patches/series.txt. The patch anchors on stable post-0016 context (score_fmt option line, VmafContext *vmaf; struct line, the do_vmaf vmaf_read_pictures call); CI validates it via a full series replay against a clean n8.1 checkout (git am --3way), not standalone git apply. - Docs/index: docs/api/perceptual-weight.md (new), mkdocs.yml (api nav row + ADR-1118 nav row), docs/state.md, docs/research/1102-*.md (new), changelog.d/added/, and the ADR index (docs/adr/1118-perceptual-sidedata-weighting.md + fragment + _order.txt + regenerated README.md).
feat/mcp-tiny-ai-feature-coverage (2026-06-14)¶
no rebase impact: all touched code is fork-local. The MCP servers (cmd/vmafx-mcp/{tools.go,impl.go,impl_direct.go,main.go,score_extras_test.go} and mcp-server/vmaf-mcp/src/vmaf_mcp/server.py + tests/test_score_extras_adr1117.py) do not exist in upstream Netflix/vmaf, so no mechanical merge conflict is possible. The change adds optional scoring parameters that shell out to existing vmaf CLI flags — it does not add, rename, or remove any public C-API entry point, CLI flag, meson_options.txt entry, public header, or LIBVMAFContext field, so per CLAUDE.md §12 r14 there is no ffmpeg-patch impact (the patches under ffmpeg-patches/ consume libvmaf symbols, not the MCP servers). Rebase-sensitive invariant the diff must preserve: the Go (scoringExtraProperties()) and Python (_scoring_extra_properties()) schema generators MUST stay byte-identical (same keys/enums/defaults/descriptions) per cmd/vmafx-mcp/AGENTS.md §1 — the parity tests (server_test.go::TestToolSchemasMatchPython, test_score_extras_adr1117.py) and the source-of-truth flag names in core/tools/cli_parse.c are the backstop. Also edits docs/mcp/tools.md, docs/state.md, changelog.d/, the ADR index (ADR-1117 + fragment + _order.txt + regenerated README.md), and a research digest — all fork-local docs.
feat/metric-niqe (2026-06-14)¶
Rebase impact: low. Adds four fork-only files (core/src/feature/niqe.c, niqe_math.h, niqe_model.h, core/test/test_niqe.c) — no upstream twin (NIQE is fork-trained against model/other_models/niqe_v0.1.pkl). Additive registration: one extern VmafFeatureExtractor vmaf_fex_niqe; + one feature_extractor_list[] entry in core/src/feature/feature_extractor.cpp (the LIVE C++23 registry, NOT the dead feature_extractor.c twin), one source line in core/src/meson.build, one executable() + one test() in core/test/meson.build. Edits docs/metrics/niqe.md (new), mkdocs.yml (NIQE nav row + regenerated ADR-nav block via scripts/docs/generate-adr-nav.sh), docs/state.md, docs/rebase-notes.md, changelog.d/, testdata/scores_cpu_niqe.json (new), and the ADR index (docs/adr/1112-niqe-nr-metric.md + fragment + _order.txt + regenerated README.md). CPU-only scalar extractor; no public C-API / ABI / CLI flag / meson_options.txt change → no ffmpeg-patch impact (the feature is reachable via the existing generic --feature path). Load-bearing invariants if the algorithm is ever touched: the AGGD N keeps the trailing *aggdratio factor and the MSCN maps + PIL bicubic half-res output stay float32-rounded — both are required for parity with the pkl the model was trained against (see core/src/feature/AGENTS.md and ADR-1112).
feat/metal-standalone-batch (2026-06-14)¶
Rebase impact: low. Adds 12 fork-only files (4 kernels x {.metal,_metal.mm,test}) for integer_ciede / integer_psnr_hvs / integer_cambi / ssimulacra2 — no upstream twins. Additive registration (4 externs + 4 list entries in feature_extractor.c
if HAVE_METAL; 4 .mm sources + 4 custom_targets + 4 metal_air_files in¶
core/src/metal/meson.build; a foreach test block in core/test/meson.build). Edits docs/metrics/features.md (+Metal on the 4 rows), state.md, changelog. cambi is a Strategy-II hybrid (GPU kernels + exact-CPU host residual via cambi_internal.h), matching ADR-0205. Metal-only; no public C-API/CLI change -> no ffmpeg-patch impact.
feat/metric-delta-e-itp (2026-06-14)¶
Rebase impact: low. Fork-only additive metric — no upstream twin. Adds three new files (core/src/feature/delta_e_itp.c, core/src/feature/delta_e_itp_math.h, core/test/test_delta_e_itp.c). Additive registration only: one extern + one feature_extractor_list[] entry in core/src/feature/feature_extractor.cpp (the live C++23 file — NOT the stale feature_extractor.c twin, which is dead per ADR-0846 and is a separate cleanup), one source line in core/src/meson.build (next to ciede.c), one executable() + one test() block in core/test/meson.build (next to test_ciede), and one nav entry in mkdocs.yml. No public C-API / ABI / CLI flag / meson_options.txt / public-header change → no ffmpeg-patch impact (CLAUDE §12 r14 N/A). The metric mirrors the CPU ciede.c structure (chroma-upsample helpers, 8/16-bit reads, double-precision frame sum); if the upstream ciede.c chroma-upsampling helpers are ever refactored, the copied-verbatim scale_chroma_planes / scale_chroma_planes_hbd in delta_e_itp.c are independent and need no follow-up. Compiled unconditionally (CPU); no backend flag.
feat/metric-pu21 (2026-06-14)¶
Rebase impact: low (fork-only additive). Adds five fork-only files (core/src/feature/pu21.c, pu21_math.h, pu21_ssim.c, pu21_ssim.h, core/test/test_pu21.c) — no upstream twin. Additive registration: extern + list entry in core/src/feature/feature_extractor.cpp (the active C++ registry — NOT the dead feature_extractor.c, which the build does not compile), pu21.c + pu21_ssim.c added to the unconditional libvmaf_feature_sources in core/src/meson.build (next to ciede.c), test block + test() row in core/test/meson.build. Edits docs/metrics/pu21.md (new), mkdocs.yml, docs/adr/README.md, docs/state.md, docs/rebase-notes.md, changelog.d/. Reuses only the read-only iqa Gaussian-convolve helper (iqa/convolve.c) and the read-only Gaussian window table (iqa/ssim_tools.h); the golden float_ssim/iqa_ssim (L=255) is untouched — PU21 ships its own L=256 SSIM. No public C-API / ABI / CLI flag / meson_options.txt change → no ffmpeg-patch impact. If the iqa convolve layout (output packed at the reduced stride w-kw+1) ever changes, pu21_ssim.c's reduction stride must follow.
feat/metal-integer-adm (2026-06-14)¶
Rebase impact: low. Adds three fork-only files (core/src/feature/metal/integer_adm.metal, integer_adm_metal.mm, core/test/test_metal_integer_adm_parity.c) — no upstream twin. Additive registration: extern + list entry in core/src/feature/feature_extractor.c (#if HAVE_METAL), .mm source + custom_target + metal_air_files entry in core/src/metal/meson.build, test block in core/test/meson.build. Edits docs/metrics/features.md (adm fixed-point GPU column += SYCL/HIP/Metal), docs/state.md, docs/rebase-notes.md, changelog.d/. Metal-only (-Denable_metal=enabled); no public C-API / ABI / CLI / meson_options.txt change → no ffmpeg-patch impact. Mirrors the CPU integer_adm.c fixed-point DWT pipeline — if that algorithm changes, the Metal twin must follow.
fix/mcp-schema-bitdepth-vulkan (2026-06-14)¶
no rebase impact: edits the fork-only MCP servers (cmd/vmafx-mcp/tools.go, mcp-server/vmaf-mcp/src/vmaf_mcp/server.py) + docs/mcp/tools.md + state.md + changelog. No libvmaf C-API/CLI change. The bitdepth enum + backend enum must stay in sync between the Python and Go MCP tool schemas (byte-compatible pair).
feat/metal-integer-vif (2026-06-14)¶
Rebase impact: low. Adds three fork-only files (core/src/feature/metal/integer_vif.metal, integer_vif_metal.mm, core/test/test_metal_integer_vif_parity.c) — no upstream twin. Additive registration: extern + list entry in core/src/feature/feature_extractor.c (#if HAVE_METAL), .mm source + custom_target + metal_air_files entry in core/src/metal/meson.build, test block in core/test/meson.build. Edits docs/metrics/vif.md, docs/state.md, docs/rebase-notes.md, changelog.d/. Metal-only (-Denable_metal=enabled); no public C-API / ABI / CLI / meson_options.txt change → no ffmpeg-patch impact. Mirrors the CPU integer_vif.c fixed-point arithmetic + the float_vif_metal scaffold; if the CPU integer-VIF math changes, the Metal twin must follow.
feat/metal-float-adm (2026-06-14)¶
Rebase impact: low. Adds three fork-only files (core/src/feature/metal/float_adm.metal, float_adm_metal.mm, core/test/test_metal_float_adm_parity.c) — no upstream twin. Additive registration: extern + list entry in core/src/feature/feature_extractor.c (#if HAVE_METAL), .mm source + custom_target + metal_air_files entry in core/src/metal/meson.build, test block in core/test/meson.build. Edits docs/metrics/features.md (float_adm GPU column += Metal), docs/state.md, docs/rebase-notes.md, changelog.d/. Metal-only (-Denable_metal=enabled); no public C-API / ABI / CLI / meson_options.txt change → no ffmpeg-patch impact. Core-VMAF kernel; mirrors the CUDA float_adm/ DWT+CSF+CM pipeline — if a future change alters that algorithm, the Metal twin must follow.
feat/metal-float-vif (2026-06-14)¶
Rebase impact: low. Adds three fork-only files (core/src/feature/metal/float_vif.metal, float_vif_metal.mm, core/test/test_metal_float_vif_parity.c) — no upstream twin. Additive registration: one extern + one list entry in core/src/feature/feature_extractor.c (#if HAVE_METAL), one .mm source + one custom_target + one metal_air_files entry in core/src/metal/meson.build, one test block in core/test/meson.build. Edits docs/metrics/vif.md, docs/state.md, docs/rebase-notes.md, changelog.d/ — keep both additive hunks on concurrent-branch conflict. Metal-only (compiles under -Denable_metal=enabled); no public C-API / ABI / CLI / meson_options.txt change → no ffmpeg-patch impact. Part of the Metal full-parity sweep (9 real kernels); float_vif is core-VMAF.
feat/metal-integer-ssim (2026-06-14)¶
Rebase impact: low. Adds three fork-only files (core/src/feature/metal/integer_ssim.metal, integer_ssim_metal.mm, core/test/test_metal_integer_ssim_parity.c) — no upstream twin. Registration edits are additive: one extern + one list entry in core/src/feature/feature_extractor.c (inside the #if HAVE_METAL block), one .mm source + one custom_target + one metal_air_files entry in core/src/metal/meson.build, one test block in core/test/meson.build. Edits docs/metrics/ssim.md, docs/state.md, docs/rebase-notes.md, changelog.d/ — keep both additive hunks if a concurrent branch also edits them. The Metal kernel only compiles under -Denable_metal=enabled (macOS); no public libvmaf C-API / ABI / CLI / meson_options.txt change, so no ffmpeg-patch (CLAUDE §12 r14) impact. Scope note: Metal full-parity is 9 real kernels (not 11) — integer_moment/integer_ms_ssim are not distinct extractors.
feat/gpu-motion3-v2-twins (2026-06-14)¶
Rebase impact: low. Touches three fork-added GPU wrappers (core/src/feature/{sycl,hip,metal}/integer_motion_v2_{sycl.cpp,hip.c,metal.mm}) — none have an upstream twin, so no upstream-sync conflict — plus their fork-only parity tests (core/test/test_{sycl,hip,metal}_motion_v2_parity.c). Each mirrors the merged CUDA flush_fex_cuda motion3_v2 post-process (fix/cuda-motion-v2-motion3-emission, ADR-1108) byte-for-byte and reuses the shared motion_blend_tools.h helper; no GPU kernel is modified. Edits docs/metrics/motion.md, docs/state.md, docs/rebase-notes.md, changelog.d/ — keep both additive hunks if a concurrent branch also edits them. No public libvmaf C-API, ABI, header, CLI, or meson_options.txt change, so no ffmpeg-patch (CLAUDE §12 r14) impact. If a future change alters the CPU integer_motion_v2.c::flush blend/clip/seed/moving-average logic, all four GPU twins (cuda/sycl/hip/metal) must be updated in the same PR to keep the places=4 parity gate green.
fix/cuda-motion-v2-motion3-emission (2026-06-13)¶
Rebase impact: low. Touches core/src/feature/cuda/integer_motion_v2_cuda.c (fork-added CUDA wrapper — no upstream twin, so no upstream-sync conflict), core/test/test_cuda_motion_v2_parity.c (fork-only test), and core/src/feature/cuda/AGENTS.md (fork doc). Adds docs/adr/1108-*.md, changelog.d/fixed/1108-*.md. Edits docs/metrics/motion.md, docs/adr/README.md, and docs/state.md — these can conflict with a concurrent branch that also edits the same doc; keep both additive hunks (the motion3_v2 rows/paragraph here plus whatever the other branch adds). The motion3_v2 emission reuses the existing motion_blend_tools.h host helper and the established vmaf_feature_collector_append_with_dict API — no public libvmaf C-API, ABI, header, CLI, or meson_options.txt surface change, so no ffmpeg-patch (CLAUDE §12 r14) impact. The CUDA kernel itself is unchanged; only the host-side flush + option table grew.
feat/vmafx-scorestream-phase2 (2026-06-13)¶
no rebase impact: all changes are in fork-local Go files that do not exist in upstream Netflix/vmaf — pkg/libvmaf/stream.go (+ test), pkg/libvmaf/libvmaf.go (adds the exported Scorer.ResolveModel wrapper), cmd/vmafx-server/grpc_server.go, cmd/vmafx-node/server/server.go, cmd/vmafx-node/main.go, and their tests, plus docs/ and changelog.d/. The cgo path links against the public libvmaf C ABI (vmaf_picture_alloc / vmaf_read_pictures / vmaf_score_at_index / vmaf_score_pooled / vmaf_feature_score_at_index) — all stable upstream entry points in core/include/libvmaf/libvmaf.h; no upstream-mirrored C source is modified, so no mechanical conflict is possible. If a future upstream sync renamed any of those public functions, pkg/libvmaf/{direct,stream}.go would need the same one-line follow per the existing cgo-coupling invariant.
fix/json-model-feature-name-leak (2026-06-13)¶
Rebase impact: low. Touches core/src/read_json_model.c (upstream-mirrored, libvmaf/src/read_json_model.c upstream) and its fork-only C++23 twin core/src/read_json_model.cpp (ADR-0761 / ADR-0846 Wave 8) — both gain one free(model->feature[index].name) line plus a comment inside append_feature_name, immediately before the strdup. Also adds one test function + one registration line to core/test/test_model.c. Upstream lacks the duplicate-key overwrite guard, so a future sync that rewrites append_feature_name in the .c file will conflict on that hunk only; keep the fork's free-before-strdup (it fixes a real leak the upstream code shares). The .cpp twin is fork-only and never receives upstream hunks. No public API, ABI, header, or CLI surface changes — no ffmpeg-patch impact.
fix/golden-cpu-regression-restore (2026-06-13)¶
Rebase impact: low. Touches core/src/feature/vif_tools.c (removes #if HAVE_AVX512 dispatch blocks from the three float VIF functions). This file also exists in upstream Netflix/vmaf. Future upstream syncs that modify vif_tools.c will see a clean merge on any hunk that does not overlap with the three removed dispatch blocks. If upstream ever adds AVX-512 float VIF dispatch, the upstream version must be audited for Netflix golden parity before enabling it on this fork.
docs/rc-deferred-closeout (2026-06-13)¶
no rebase impact: changes confined to docs/state.md (move T-DOC-LEGACY-RUNNER from Open to Recently Closed) and docs/metrics/cambi.md (remove stale Vulkan section, doc-only). Conflicts with a concurrent branch that also edits state.md: keep both state-update rows; the row order within Recently Closed does not matter.
test/float-extractor-cpu-coverage — Float extractor CPU-path unit tests (2026-06-13)¶
no rebase impact: test-only addition. New files: core/test/test_float_{psnr,moment,ssim,ms_ssim,vif,adm,motion}_coverage.c. Edit to core/test/meson.build adds 7 new executable targets after line 1705 (test_float_vif_min_dim). No conflict risk unless a concurrent branch also adds tests immediately after that same line; in that case, append both blocks in source order.
fix/hip-meson-speed-tus-dedup — Remove duplicate HIP speed TU entries (2026-06-12)¶
no rebase impact: change confined to core/src/hip/meson.build (build config only). Removes the duplicate speed_chroma_hip.c / speed_temporal_hip.c wiring block that ADR-0852 introduced when ADR-0964 had already included those TUs earlier in the same hip_sources list. If a concurrent branch also edits core/src/hip/meson.build, keep both sets of changes; the resolved file must contain each speed TU exactly once.
fix/docker-ffmpeg-tag-n811-pin — Docker FFMPEG_TAG n8.1 → n8.1.1 + patch 0016 context (2026-06-12)¶
no rebase impact: Dockerfile and Dockerfile.ffmpeg pin changes are build-config only; patch 0016 context line fix is an ffmpeg-patches-internal correction with no effect on C API, ABI, or libvmaf source. Files touched: Dockerfile, Dockerfile.ffmpeg, .pre-commit-config.yaml, ffmpeg-patches/0016-libvmaf-wire-score-fmt-on-all-vmaf-filters.patch.
chore/codeql-cpp-cleanup-bundle — CodeQL C++ note-level cleanup (2026-06-12)¶
no rebase impact: all changes are local variable renames, dead-code removals, and comment additions. Files touched: adm_avx2.c, adm_avx512.c, integer_adm.c, feature_collector.c, feature_collector.cpp, feature_name.c, libvmaf.c, mkdirp.c, speed.c, vif_tools.c, pdjson.c, predict.c, svm.cpp, test_score_pooled_eagain.c, test_tensor_io.c, test_cambi.c, test_integer_adm_simd.c, test_svm_api.c. No API or ABI changes; no semantic changes to score computation. If a concurrent branch modifies any of these files, resolve by keeping both sets of changes — variable renames are local and non-conflicting.
chore/bundle-fable-5-findings — 4 Fable deep-hunt fixes (2026-06-12)¶
core/src/feature/x86/integer_ssim_avx2.c: reorder w*(s*s) to (w*s)*s for the 16-bit accumulation; only affects integer_ssim AVX2 16-bit path. No conflict risk on other branches unless they also modify integer_ssim_avx2.c accumulation order. core/src/libvmaf.c: three separate hunks — bpc &&→|| in validate_pic_params; Phase 2 CUDA PREV_REF vmaf_picture_ref instead of bare copy; dist translate error-propagation in read_pictures_cuda_translate. If a concurrent branch edits libvmaf.c in those functions, resolve by keeping all three fixes; they are independent. core/test/test_validate_pic_params_bpc.c and core/test/meson.build: new test file and meson registration. No conflict risk unless another branch adds a test with the same name. cmd/vmafx-server/concurrency.go, concurrency_test.go, grpc_server.go, http_server.go, main.go: ScoreLimiter addition. If a concurrent branch also modifies main.go flag parsing or grpc_server.go/http_server.go handler signatures, resolve by preserving the WithLimiter constructors and the --max-concurrent-scores flag.
fix/master-855-tip-3-reds — bootstrap-test recal + Dockerfile ldconfig (2026-06-08, no ADR)¶
no rebase impact: python/test/local_explainer_test.py line 276 expected value and places argument changed (fork-local test, not Netflix golden data); Dockerfile gains a single RUN ldconfig line after make install. If a concurrent branch also edits python/test/local_explainer_test.py lines 271-277, resolve by keeping places=3 and the # ADR-0418 macOS-libm Δ relax comments. If a concurrent branch edits Dockerfile around the libvmaf build block, ensure RUN ldconfig is present immediately after the make install line.
fix/containerfile-gid-and-stale-rename — GID/UID 1000 → 2000 (2026-06-08, ADR-1101)¶
no rebase impact: changes confined to dev/Containerfile (GID/UID values), docs/adr/1101-containerfile-gid-uid-2000.md (new ADR), and changelog.d/fixed/1101-containerfile-gid-uid-2000.md (new fragment). No production C source, public header, meson build files, or Python package modified. If a concurrent branch also edits dev/Containerfile, the only conflict will be in the groupadd/useradd lines; resolve by keeping GID/UID 2000.
fix/matrix-5-real-bugs (2026-06-08, no ADR — 5 correctness bug fixes)¶
core/src/feature/hip/integer_vif/vif_statistics.hip: removed #define AMD_WAVEFRONT_SIZE 64; reduction loop and lane guards now use warpSize device variable. Conflicts possible if another branch edits the same wavefront-reduce section; resolve by keeping the warpSize-based version. core/src/feature/hip/float_vif/float_vif_score.hip, float_motion/float_motion_score.hip, float_psnr/float_psnr_score.hip, float_moment/moment_score.hip: similar pattern — shared-memory arrays resized for minimum warp size (32); runtime warpSize used for loops. Conflict risk is low (only these wavefront-size definitions changed); keep the warpSize-based version. core/src/libvmaf.c: ref = &ref_host; dist = &dist_host guarded by if (hw_flags & HW_FLAG_HOST). Conflicts possible if another branch modifies the same #ifdef HAVE_CUDA block; resolve by keeping the HW_FLAG_HOST guard. mcp-server/vmaf-mcp/src/vmaf_mcp/server.py: _PROBE_YUV_WIDTH/HEIGHT bumped from 32 to 64; runtime_healthy set to score is not None. Low conflict risk. ffmpeg-patches/0005-libvmaf-add-libvmaf-sycl-filter.patch: FILTER_SINGLE_PIXFMT replaced by FILTER_PIXFMTS; do_vmaf_sycl and config_props_sycl split on AV_PIX_FMT_QSV. Conflicts possible if another branch edits patch 0005; apply this version first, then rebase the other. dev/Containerfile: RUN bash .../fetch-test-yuvs.sh layer added. Low conflict risk.
test/ai-scripts-coverage-round3 (2026-06-06, no ADR — test-only)¶
no rebase impact: adds two new test files (ai/tests/test_calibrate_phase_f_recipes_unit.py and ai/tests/test_analyze_knob_sweep_unit.py) and one changelog fragment. No existing C source, public API, upstream-mirrored Python, or golden assertion is modified.
docs/r12-c-api-doc-completeness (2026-06-06, no ADR — doc-only)¶
no rebase impact: comment-only changes to core/include/libvmaf/libvmaf_cuda.h, core/include/libvmaf/libvmaf_sycl.h, core/include/libvmaf/dnn.h, core/include/libvmaf/picture_v2.h, core/include/libvmaf/libvmaf.h, and core/include/libvmaf/model.h. No C sources, build files, or public API signatures touched — Doxygen comment additions only.
docs/doxygen-private-headers-r4 (2026-06-07)¶
no rebase impact: purely additive Doxygen comment blocks inserted into 10 internal headers under core/src/. No include paths, struct layouts, or function signatures are changed. Conflicts only if another branch inserts text at the same line positions in these headers.
fix/pic-pool-odr-cuda-gpumask-cov-floor (2026-06-08)¶
core/src/meson.build: adds cpp_args to picture_pool_cpp23_lib. Conflicts possible if another branch modifies the same static_library() block; resolve by keeping both the cpp_args line and the other change. core/tools/test/test_vmaf_cuda_gpumask.sh and core/tools/test/meson.build: shell guard + timeout added; low conflict risk. scripts/ci/coverage-check.sh and .github/workflows/tests-and-quality-gates.yml: per-file floor and pytest timeout changed; low conflict risk (numeric/string values only).
fix/cuda-done-path-double-unref-ort-coverage (2026-06-07)¶
no rebase impact: changes confined to core/src/libvmaf.c (split read_pictures_cuda_cleanup into full and _device_only variants inside the existing #ifdef HAVE_CUDA block — non-CUDA builds are unchanged; the call site at the done=true branch is guarded by the same #ifdef HAVE_CUDA) and core/src/dnn/ort_backend.c (collapse a dead else branch into a single-line ternary in ort_log_and_release_status — no behaviour change on any exercised code path, coverage-only impact).
fix/ci-multi-platform-bundle-838 (2026-06-07)¶
no rebase impact: changes confined to core/tools/cli_parse.cpp (const-qualifier on local strsep parameter — isolated #ifndef HAVE_STRSEP compat block), core/src/opt.cpp (replace static_cast<int> with memcpy in a single switch statement — no surrounding context dependency), core/src/feature/feature_extractor.cpp (add extern "C" wrappers around existing extern declarations — purely syntactic, no semantic change), core/src/libvmaf.c (add #ifdef HAVE_CUDA cleanup call in the done=true branch of vmaf_read_pictures — guarded by HAVE_CUDA; non-CUDA builds are unchanged), and python/test/vmafexec_feature_extractor_test.py (lower places=6 to places=4 on 5 per-frame assertions).
fix/go-rust-ci-red-bundle (2026-06-07)¶
no rebase impact: changes confined to .github/workflows/go-ci.yml (env var addition to go test step), cmd/vmafx-operator/internal/controller/vmafxnode_controller_test.go (timestamp truncation), cmd/vmafx-mcp/impl_direct.go (restore ValidatePath calls), and bindings/rust/vmafx-sys/Cargo.toml (add [lib] doctest = false). No C source, public header, or upstream-mirrored code modified.
fix/build-matrix-macos-windows-fixes (2026-06-07)¶
Rebase-sensitive (meson.build): core/src/meson.build gains dependencies : [pthread_dependency] on both picture_pool_cpp23_lib and gpu_picture_pool_cpp23_lib static library targets (~lines 1768–1788). If a concurrent branch adds other fields to those static_library() calls, merge both sets of fields.
Other changes are not rebase-sensitive: - core/src/feature/arm64/motion_v2_neon.c: rewrite of neon_any_nonzero_s32 (isolated function, no surrounding context). - compat/python-vmaf/__init__.py: two call-sites of --cpumask changed from "-1" to "4294967295". - .github/workflows/libvmaf-build-matrix.yml: two Vulkan matrix rows removed; if a concurrent branch also removes the Vulkan step bodies (Install Vulkan SDK, Cache meson subprojects (Vulkan wraps), Run Vulkan smoke tests (macOS MoltenVK), etc.), take both removals. - python/test/python_harness_coverage_test.py: test expectation update (--cpumask -1 → --cpumask 4294967295; test_run_preserves_user_env expected dict gains LC_ALL/LANG).
fix/nightly-bisect-tracker-issue (2026-06-07)¶
no rebase impact: changes confined to .github/workflows/nightly-bisect.yml, scripts/ci/post-bisect-comment.py, docs/state.md, and changelog.d/fixed/nightly-bisect-tracker-issue.md. No C source, public header, Go source, or test logic modified.
fix/feature-extractor-flags-zero-skip-gpu (2026-06-07)¶
no rebase impact: the change is confined to a single function body in core/src/feature/feature_extractor.c (lines 443–473). No header changes, no meson.build changes, no new files except the ADR and changelog fragment. The only other file touched is core/test/test_picture.c (missing <string.h> include added). Neither file is a high-contention rebase target.
fix/sycl-fsycl-link-propagation (2026-06-07)¶
Rebase-sensitive: modifies core/src/meson.build and core/test/meson.build — two high-contention build files that accumulate edits from most GPU-backend PRs.
In core/src/meson.build: - The sycl_dependency declare_dependency block gains link_args: ['-fsycl']. - The vmaf_link_args += ['-fsycl'] line and its surrounding comment block are replaced with a shorter comment referencing sycl_dependency. If a concurrent branch adds entries to vmaf_link_args, take that branch's additions and keep the updated comment.
In core/test/meson.build: - The test('test_sycl_motion_add_uv_parity', ...) call loses should_fail: true and its accompanying ADR-1093 comment block. If a concurrent branch adds new SYCL test executables nearby, no conflict is expected; should_fail on other tests is unaffected.
In core/test/test_sycl_motion_add_uv_parity.c: - Feature-name queries updated (integer_motion2_mau, float_motion2_mau). Conflicts only if another branch edits the same query lines.
fix/mcp-resource-uri-validation (2026-06-07)¶
no rebase impact: single-function change in cmd/vmafx-mcp/impl_direct.go (resolveModelArgToPath) and one new test in cmd/vmafx-mcp/impl_direct_test.go. Only the Go cmd/vmafx-mcp package is touched; no C sources, no public headers, no test fixtures, no build files. Conflicts only if another branch edits resolveModelArgToPath or adds tests to impl_direct_test.go.
fix/cross-platform-path-list-separator (2026-06-06)¶
no rebase impact: single-line change in pkg/libvmaf/paths.go replacing strings.Split(extra, ":") with filepath.SplitList(extra). Only the Go pkg/libvmaf package is touched; no C sources, no public headers, no test fixtures, no build files. Conflicts only if another branch edits the same AllowedRoots function in that file.
fix/neon-motion-zero-skip (2026-06-06)¶
no rebase impact: single-file change to core/src/feature/arm64/motion_v2_neon.c. Replaces the neon_hadd_s32 (signed horizontal sum) early-exit check with neon_any_nonzero_s32 (bitwise OR-fold) in both motion_score_pipeline_8_neon and motion_score_pipeline_16_neon. No public API, no header, no test data, no upstream-mirrored file is modified. Conflicts only if another branch edits the same static helper region of that file.
fix/helm-values-completeness-adr-1074 (ADR-1074, 2026-06-06)¶
no rebase impact: changes are confined to deploy/helm/vmafx/values.yaml, deploy/helm/vmafx/values.schema.json, and three templates (templates/statefulset.yaml, templates/node.yaml, templates/networkpolicy.yaml). No C source, public header, upstream-mirrored file, Python test, or golden-data assertion is touched. Conflict risk exists only if another branch edits those same Helm files concurrently.
test/coverage-pkg-observability (2026-06-06)¶
no rebase impact: changes are confined to pkg/observability/coverage_gaps_test.go (new test file), pkg/observability/AGENTS.md (invariant notes), and changelog.d/added/observability-coverage-gaps.md (fragment). No production source, public header, or build file is modified. Conflicts only if another branch edits the same lines in AGENTS.md or rebase-notes.md.
fix/sanitizer-deselect-tests-and-quality-gates (2026-06-06)¶
no rebase impact: CI-only change to .github/workflows/tests-and-quality-gates.yml adding test_gpu_picture_pool_uaf, test_integer_motion_v2_coverage, and test_pic_preallocation to the ADR-0347 per-sanitizer EXCLUDE patterns for address, undefined, and thread. No source, header, test, or build file is modified. Conflicts only if another branch edits the same case block in that workflow file.
fix/mcp-score-at-index-eagain-guard (ADR-1073, 2026-06-06)¶
no rebase impact: changes are confined to core/src/libvmaf.c (vmaf_score_at_index guard condition), core/src/mcp/compute_vmaf.c (n_threads restored to 1u, debug code removed), and core/test/test_mcp_smoke.c (fixture dimensions 64→192, debug print removed). No public API surface, no upstream-mirrored file is modified. The guard change is a one-line fix that does not affect the call signature or semantics observable to callers that never encounter multi-frame pools.
fix/skip-motion-five-frame-window-adr-0337 (ADR-0337, 2026-06-06)¶
no rebase impact: only python/test/feature_extractor_test.py is modified — 9 test methods gain @unittest.skip decorators. No C source, public header, upstream-mirrored file, or golden-data assertion is touched. Rebase against Netflix/vmaf master or any feature branch has zero conflict risk.
fix/prev-ref-batch-refcount-and-motion-score (ADR-1072, 2026-06-06)¶
Files touched: core/src/libvmaf.c (two sites in threaded_extract_batch_func and one in threaded_extract_func), core/test/test_hip_ms_ssim_parity.c (FIXTURE_H 144→192), core/test/test_cuda_float_ms_ssim_parity.c (FIXTURE_H 144→192), core/test/test_hip_motion_parity.c (add debug=1 opts, add feature.h include), docs/adr/1072-prev-ref-batch-refcount-leak.md, docs/adr/README.md, docs/state.md, docs/rebase-notes.md, changelog.d/fixed/1072-prev-ref-batch-refcount-leak.md.
Rebase impact: The libvmaf.c hunks add vmaf_picture_unref + memset + memset(f->prev_ref) inside the VMAF_FEATURE_EXTRACTOR_PREV_REF block in threaded_extract_batch_func. If a concurrent branch modifies the same block or the unref: label region, resolve by keeping both the concurrent change and the new unref-before-memset + zero-f->prev_ref logic from this branch. The test fixture changes (144→192) and the debug-flag addition are self-contained with no shared invariants. No public-API, ABI, or upstream-mirrored file changes.
fix/test-failures-macos-dnn (2026-06-06, no ADR — bug fixes)¶
Files touched: core/src/gpu_picture_pool.{c,cpp}, core/src/libvmaf.c, core/src/feature/integer_motion.c, core/src/feature/feature_extractor.cpp, core/test/test_framesync.c, core/test/test_integer_motion_coverage.c, changelog.d/fixed/macos-dnn-test-failures-6-fixes.md, docs/rebase-notes.md.
Rebase impact: All changes are internal bug fixes with no public-API or ABI changes. If a concurrent branch modifies vmaf_score_at_index (libvmaf.c), vmaf_gpu_picture_pool_init (gpu_picture_pool.{c,cpp}), or integer_motion.c init(), resolve conflicts by keeping both the concurrent change and the err != -EAGAIN / *pool = NULL / w < 3 || h < 3 guards from this branch. The test fixes in test_framesync.c and test_integer_motion_coverage.c are self-contained; no invariants span other branches.
no rebase impact on public API, build flags, or upstream-mirrored files.
docs/doxygen-public-header-drift (2026-06-06, no ADR — doc-only fix)¶
no rebase impact: comment-only changes to core/include/libvmaf/libvmaf_cuda.h, core/include/libvmaf/libvmaf_vulkan.h, and core/include/libvmaf/libvmaf_sycl.h. No C sources, build files, or public API signatures touched.
chore/ci-workflow-audit-sha-pin-dead-jobs (2026-06-06, no ADR — workflow hygiene)¶
no rebase impact: changes are entirely in .github/workflows/ (SHA pin, dead-job removal, comment correction). No C sources, public API, build flags, or upstream-mirrored files are touched.
test/go-vmafx-mcp-handler-coverage (2026-06-06, no ADR — test-only)¶
Files touched: cmd/vmafx-mcp/impl_handlers_test.go (new), cmd/vmafx-mcp/AGENTS.md, changelog.d/added/go-vmafx-mcp-handler-coverage.md, docs/rebase-notes.md.
Rebase impact: test-only addition; no production code changed. If a concurrent branch adds a new tool handler to impl.go, add a corresponding error-path test to impl_handlers_test.go following the established pattern (t.Setenv("VMAF_BIN", "/nonexistent/...") for binary-dependent handlers).
fix/r10-cpp23-wave-error-paths (2026-06-06, ADR-1060)¶
Files touched: core/src/feature/feature_extractor.cpp, core/src/read_json_model.cpp
Rebase impact: no rebase impact. All changes are internal to existing functions with no public-API or header changes. Branches that also touch feature_extractor.cpp should verify the free_fex_list label and the context-create parse-options error path merge cleanly.
fix/helm-chart-security-hardening (2026-06-06, ADR-1058)¶
Files touched: deploy/helm/vmafx/templates/pdb.yaml (new), deploy/helm/vmafx/templates/operator-rbac.yaml, deploy/helm/vmafx/templates/networkpolicy.yaml, deploy/helm/vmafx/values.yaml, deploy/helm/vmafx/values.schema.json
Rebase impact: The operator RBAC resource names changed: *-operator-role (ClusterRole) is replaced by *-operator-crds (ClusterRole) + *-operator-ns (Role). Any branch that patches operator-rbac.yaml will conflict on the resource name. Run helm upgrade (not in-place patch) when applying to existing operator installs. The networkPolicy.allow schema is now additionalProperties: false; any branch that adds a new allow.* key must also enumerate it in values.schema.json.
fix/rust-clippy-library-strictness (2026-06-06, ADR-1063)¶
Files touched: bindings/rust/vmafx-sys/src/lib.rs, bindings/rust/vmafx-sys/src/safe.rs, bindings/rust/vmafx/src/lib.rs, bindings/rust/vmafx/src/picture.rs, bindings/rust/vmafx/src/error.rs, core/src/feature/rust/tad/src/lib.rs
Rebase impact: vmafx-sys/src/lib.rs no longer uses crate-level #![allow(clippy::all)]; the generated bindings are now in a private mod bindings with the allow scoped to that module. Any branch that adds new hand-written code to vmafx-sys/src/lib.rs or safe.rs must write clippy-clean code. The VmafContext::default() call is gone — branches that depend on it must use VmafContext::new() instead. The #![deny(unsafe_op_in_unsafe_fn)] in safe.rs and tad/src/lib.rs will cause a compile error on any in-flight branch that adds a bare unsafe operation inside an unsafe fn without an explicit unsafe {} block.
fix/msvc-cpp-std-vc-latest-1056 (2026-06-06, ADR-1056)¶
Files touched: core/meson.build, core/AGENTS.md
Rebase impact: core/meson.build no longer carries cpp_std=c++23 in default_options. Any branch that adds cpp_std=... to default_options will conflict with this change. The add_project_arguments('-std=c++23') block must remain beneath the cxx = meson.get_compiler('cpp') line and above the first cc.check_header call. The get_option('cpp_std') == 'none' guard must be preserved; removing it would cause the SYCL leg to receive both -Dcpp_std=c++14 (from the workflow) and -std=c++23 (from the else branch), which is a compile error.
fix/ci-pin-cuda-132-jimver (2026-06-06, no ADR — CI configuration pin fix)¶
no rebase impact: CI-only change (.github/workflows/build.yml, .github/workflows/libvmaf-build-matrix.yml). No C sources, public API, or upstream-mirrored files are touched.
fix/macos-docker-platform-unblock (2026-06-04, no ADR — build bug fix)¶
no rebase impact: adds <string_view> include to core/tools/vmaf.cpp (no logic change) and replaces VmafCudaFunctions with CudaFunctions in 13 CUDA close callbacks (correct type name, no ABI/API change). Neither modification touches upstream-mirrored code paths or public API signatures.
revert/float-adm-simd-dispatch-neon-fma (2026-06-06, ADR-1057)¶
no rebase impact: removes adm_prime_simd_dispatch() from adm_tools.h and adm_tools.c; removes the call site added to float_adm.c::init() by PR #685; deletes core/test/test_float_adm_simd.c and its meson.build entries. Any in-flight branch that rebases onto a version of adm_tools.h that still contains adm_prime_simd_dispatch() will see a merge conflict at the declaration — resolve by simply not including the declaration (the function no longer exists after this revert). The SIMD kernel files (adm_tools_avx2.c, adm_tools_neon.c, etc.) are untouched; the functions remain compiled and linkable for a future re-dispatch PR.
fix/pr1161-neon-adm-contract (2026-08-31, ADR-1057 follow-up)¶
Rebase impact: preserve the scoped non-contracting arithmetic in core/src/feature/adm_tools.c::adm_dwt2_s and core/src/feature/arm64/float_adm_dwt2_neon.c. The scalar function carries an in-body Clang contract(off) pragma and a GCC optimize("-ffp-contract=off") attribute. Do not widen the Clang pragma to the whole scalar translation unit: that changes unrelated ADM reductions. The NEON twin retains separate vmulq_laneq_f32 plus vaddq_f32 operations, starting every accumulator at +0 before its four taps, split scalar multiply/add, and its dedicated -ffp-contract=off build flag. The initial +0 preserves scalar-identical signed zero. Do not introduce vfmaq, fmaf, or initialize an accumulator directly from tap 0.
Retest test_float_adm_dwt2_neon, including its signed-zero case, under both Clang and GCC AArch64 builds through QEMU.
The integer ADM contract has a separate, intentional platform boundary. Keep adm_dwt2_8_neon() four-tap and scalar-bit-exact everywhere. In integer_adm.c, Apple AArch64 production dispatch must select adm_dwt2_8_neon_apple_legacy(), which runs the universal kernel and replaces only output column j == 0 with the historical three-tap boundary recorded by the immutable Darwin quality tests. Linux AArch64 must continue to select the universal kernel. Retest test_adm_dwt2_neon under AArch64/QEMU and the complete macOS Python quality suite; the former locks both kernel contracts, while only the latter executes the real __APPLE__ dispatch. Never alter Netflix assertions, snapshots, or parity tolerances to resolve a mismatch.
fix/core-test-regressions-pr-train (2026-06-04, no ADR — bug fixes)¶
Files touched: core/src/gpu_picture_pool.cpp, core/src/feature/feature_extractor.cpp, core/src/feature/integer_motion.c, core/src/predict.c, core/test/test_framesync.c, core/test/test_integer_motion_coverage.c, core/test/test_score_pooled_eagain.c
Rebase impact: Any concurrent branch that also edits feature_extractor_list[] must preserve the &vmaf_fex_integer_motion_v2 entry. Any branch that adds a new motion extractor with VMAF_FEATURE_EXTRACTOR_PREV_REF flag benefits from the context_extract prev_ref management added here. Branches modifying predict_load_feature_score must not regress the -EAGAIN vs -EINVAL distinction for unwritten feature vectors (Netflix#755 / ADR-0154).
fix/legacy-runner-import-stub-adr0749 (2026-06-04, no ADR — bug fix)¶
Files touched: compat/python-vmaf/core/quality_runner.py, docs/state.md, changelog.d/fixed/legacy-runner-import-stub-adr0749.md
Rebase impact: The VmafLegacyQualityRunner stub is fork-local and does not conflict with upstream Netflix/vmaf (which never had this class). No upstream sync will touch compat/python-vmaf/core/quality_runner.py in a way that removes the stub; if upstream adds a class with the same name, the stub must be removed rather than overwritten.
fix/arm-motion-v2-re-register-and-test-order¶
Files touched: core/src/meson.build, core/src/feature/feature_extractor.c, core/test/test_integer_motion_coverage.c
Rebase impact: no rebase impact from other branches expected. If a concurrent PR touches feature_extractor_list[] or meson.build's CPU source list, preserve integer_motion_v2.c registration and the &vmaf_fex_integer_motion_v2 list entry — removing them breaks all "motion_v2" lookups on CPU-only builds.
ci/dev-container-gate-adr0819 (2026-06-04, ADR-0819)¶
no rebase impact: adds .github/workflows/dev-container-build.yml and docs/adr/0819-dev-container-ci-gate.md. No C source, public C API, upstream-mirrored Python, Netflix golden-assertion file, or ffmpeg-patches file is touched.
docs/mkdocs-strict-nav-conformance¶
no rebase impact: changes are isolated to mkdocs.yml nav entries and a changelog fragment. No C source, public C API, upstream-mirrored Python, Netflix golden-assertion file, or ffmpeg-patches file is touched.
ci/promote-gpu-coverage-gate-required¶
no rebase impact: changes are isolated to the CI workflow file and docs. No C source, public C API, upstream-mirrored Python, Netflix golden-assertion file, or ffmpeg-patches file is touched.
fix/containerfile-user-hardening-adr1042¶
no rebase impact: container hardening changes only (USER directive and ARG/ENV scoping). No public API or upstream-mirrored C code touched.## fix/r9-helm-vmaftune-grpc-bugs (2026-06-04)
no rebase impact: changes are confined to deploy/helm/vmafx/ (Helm chart config only), tools/vmaf-tune/src/vmaftune/cli.py (Python), and cmd/vmafx-node/online_feedback.go (fork-local Go binary). None of these files has an upstream Netflix/vmaf counterpart.
fix/r6-sycl-kernel-correctness (2026-06-04)¶
Files touched: core/src/feature/sycl/integer_vif_sycl.cpp, core/src/feature/sycl/integer_motion_sycl.cpp, core/src/feature/sycl/integer_adm_sycl.cpp
Rebase impact: no rebase impact — all files are fork-local SYCL paths with no upstream counterparts.
fix/r6-cuda-hip-kernel-correctness (2026-06-04)¶
Files touched: core/src/feature/cuda/integer_vif/filter1d.cu, core/src/feature/cuda/integer_adm/adm_cm.cu, core/src/feature/hip/integer_adm/adm_decouple.hip, core/src/feature/hip/integer_vif/vif_statistics.hip
Rebase impact: no rebase impact — all four files are fork-local GPU paths with no upstream counterparts. The CUDA files are in feature/cuda/ which Netflix upstream does not ship; the HIP files are fully fork-added.
fix/r6-metric-scoring-guards (2026-06-04)¶
Files touched: core/src/feature/integer_psnr.c, core/src/feature/x86/psnr_avx2.c, core/src/feature/x86/psnr_avx512.c, core/src/feature/arm64/psnr_neon.c, core/src/feature/adm.c, core/src/feature/integer_adm.c, core/src/feature/float_adm.c
Rebase impact: no rebase impact — all fixes are in error-path / edge-case branches that upstream has not touched since the fork. The APSNR cap formula change (* 2 removed) only affects scores on nearly-perfect sequences; it is not a Netflix golden-data assertion value.
fix/r7-ci-wf-concurrency-timeout (2026-06-04, ADR-1035)¶
Files touched: .github/workflows/nightly.yml, .github/workflows/nightly-bisect.yml, .github/workflows/supply-chain.yml, .github/workflows/release-please.yml, .github/workflows/scorecard.yml, .github/workflows/rust-ci.yml, .github/workflows/go-ci.yml, .github/workflows/e2e-k8s.yml
No rebase impact: pure CI configuration changes with no code-path dependencies. Upstream Netflix/vmaf does not carry these workflows.
fix/r7-docs-broken-links-mkdocs-nav (2026-06-04)¶
Files touched: docs/development/build-flags.md, docs/metrics/features.md, mkdocs.yml
No rebase impact: documentation-only changes. No C library, public header, or Netflix golden-assertion file is touched.
fix/r7-mcp-precision-subsample-drift (2026-06-04, ADR-1038)¶
Files touched: cmd/vmafx-mcp/impl.go, cmd/vmafx-mcp/tools.go, mcp-server/vmaf-mcp/src/vmaf_mcp/server.py
No rebase impact: pure default-value changes. No C library, public header, upstream Python harness, or Netflix golden-assertion file is touched.
fix/r7-vendored-svm-realloc-oom (2026-06-04, ADR-1039)¶
Files touched: core/src/svm.cpp
no rebase impact: three internal realloc safety patches. No public C API, public header, upstream-mirrored Python, or Netflix golden-assertion file is touched. If an upstream Netflix/vmaf commit also fixes these same three sites, take the upstream version (which is also a MEM04-C fix) and drop this patch at rebase time.
fix/r7-licensing-spdx-svm-copyright (2026-06-04)¶
Files touched: Cargo.toml, bindings/rust/vmafx/Cargo.toml, ai/pyproject.toml, mcp-server/vmaf-mcp/pyproject.toml, dev-llm/pyproject.toml, python/pyproject.toml, tools/ensemble-training-kit/pyproject.toml, tools/vmaf-roi-score/pyproject.toml, tools/vmaf-tune/pyproject.toml, core/src/svm.cpp
No rebase impact: license field corrections and copyright header additions have no effect on build or test outputs. Upstream Netflix/vmaf does not carry Cargo.toml or any of these pyproject.toml files.
fix/sycl-speed-incomplete-type-access (2026-06-04)¶
Files touched: core/src/feature/sycl/speed_chroma_sycl.cpp, core/src/feature/sycl/speed_temporal_sycl.cpp
no rebase impact: internal build-fix replacing direct struct member dereferences with the existing public API call vmaf_sycl_get_queue_ptr(). No public C API, public header, upstream-mirrored Python, or Netflix golden-assertion file is touched. If an upstream commit adds a SYCL speed extractor, ensure it also uses vmaf_sycl_get_queue_ptr() rather than direct struct access.
fix/cli-narrowing-casts-vmaf-cpp (2026-06-04)¶
Files touched: core/tools/vmaf.cpp
no rebase impact: three static_cast<unsigned>(...) wrappers added to the VmafPictureConfiguration initializer at line ~1360. No public C API, public header, upstream-mirrored Python, or Netflix golden-assertion file is touched. If an upstream commit modifies the VmafPictureConfiguration initializer or adds new pic_params fields, verify the cast pattern is preserved.
fix/release-please-config-json-parse-error (2026-06-04)¶
no rebase impact: removes a duplicate array element from release-please-config.json. No C source, public header, Python harness, or Netflix golden-assertion file is touched. Any in-flight branch that modifies release-please-config.json should simply ensure the ai package's changelog-sections array no longer contains two chore entries.
fix/simd-psnr-16bit-scalar-tail-overflow (2026-06-04)¶
Files touched: core/src/feature/x86/psnr_avx2.c, core/src/feature/x86/psnr_avx512.c, core/src/feature/arm64/psnr_neon.c
no rebase impact: internal arithmetic fix in scalar tail loops. No public C API, public header, upstream-mirrored Python, or Netflix golden-assertion file is touched. The change affects only the three SIMD backends' scalar-remainder path for 16-bit PSNR; the SIMD main loop is unchanged. Port of any upstream commit touching these files should verify that the (uint32_t)abs(...) pattern is preserved in the scalar tail if the upstream change modifies it.
fix/r6-cpu-scoring-nan-ub-guards (2026-06-04)¶
Files touched: core/src/feature/integer_psnr.c, core/src/feature/ms_ssim.c, core/src/feature/float_ssim.c, core/src/feature/float_ms_ssim.c, core/src/feature/iqa/ssim_tools.c, core/src/feature/adm.c, core/src/feature/integer_adm.c, core/src/feature/float_adm.c, core/src/feature/motion.c, core/src/feature/cambi.c, docs/adr/1033-cpu-scoring-nan-ub-guards.md, changelog.d/fixed/1033-cpu-scoring-nan-ub-guards.md
no rebase impact: all changes are internal correctness fixes inside CPU-path scoring functions. No public C API headers, no meson_options.txt entries, no ffmpeg-patches/ series entries, and no Netflix golden-assertion files are touched. Rebasing on top of any upstream commit that modifies these same source files may produce minor context conflicts in the guard blocks; resolve by keeping both the upstream change and the NaN guard. ADR-1033.
fix/vmaf-init-double-init-guard-vmaf-close-pointer-contract (2026-06-04, ADR-1032)¶
Files touched: core/src/libvmaf.c, core/src/dnn/dnn_api.c, core/include/libvmaf/libvmaf.h, core/test/test_context.c
no rebase impact: all changes are fork-local bug-fixes with no upstream equivalents. vmaf_init guard is a new branch (no upstream logic removed), vmaf_close header change is documentation-only, and the DNN fallback path touches a fork-added sidecar-loading block that does not exist in Netflix upstream. No Netflix golden assertions or upstream-mirrored Python are touched.
fix/cuda-vif-filter1d-adm-cm-opprec (2026-06-04)¶
Files touched: core/src/feature/cuda/integer_vif/filter1d.cu, core/src/feature/cuda/integer_adm/adm_cm.cu
no rebase impact: pure kernel arithmetic fixes. No public C API header, no meson build option, no FFmpeg patch surface, and no upstream-mirrored Python file is touched. The fixes correct two silent arithmetic defects (a typo in the rd-filter upper-bound guard in filter1d.cu and a missing parenthesis pair in two x_sq reduction loops in adm_cm.cu). Cross-backend SYCL/HIP/ Vulkan ADM and VIF twins do not carry the same expressions and are unaffected.
fix/sycl-vif-rd-stride-motion-uv-sync (2026-06-04)¶
Files touched: core/src/feature/sycl/integer_vif_sycl.cpp, core/src/feature/sycl/integer_motion_sycl.cpp, docs/adr/1034-sycl-vif-rd-stride-motion-uv-sync.md, changelog.d/fixed/sycl-vif-rd-stride-motion-uv-sync.md
If this branch rebounds onto a commit that changes the rd_stride or rd_size allocation in integer_vif_sycl.cpp, re-verify that both the scalar (SIMD-32) and SIMD-16 kernel variants use (e_w + 1U) / 2U as the stride and that the allocation uses ((w + 1U) / 2U) * ((h + 1U) / 2U). If a future PR routes UV H2D copies through copy_queue and updates last_upload_event, the vmaf_sycl_queue_wait(state) added in submit_fex_sycl can be removed in favour of the GPU-side barrier — track this as a follow-up optimization.
fix(hip,metal): HIP adm_decouple dangling body + VIF wavefront carry + Metal motion vertical halo (ADR-1030, 2026-06-04)¶
Files touched: core/src/feature/hip/integer_adm/adm_decouple.hip, core/src/feature/hip/integer_vif/vif_statistics.hip, core/src/feature/metal/float_motion.metal, docs/adr/1030-hip-metal-kernel-correctness.md, changelog.d/fixed/hip-metal-kernel-correctness-1030.md
Rebase impact: low. These are self-contained correctness fixes inside GPU-only kernel files. No public C API, no CPU feature extractor, no CLI flag, and no Netflix golden assertion is touched. Any branch that also modifies adm_decouple.hip will need to re-apply the dangling-body removal; branches touching vif_statistics.hip wavefront_reduce_i64 will need to keep the integer-addition reassembly. Metal float_motion.metal conflicts are straightforward to resolve by preserving TILE_H=20 and the - HALF_FW origin offsets.
docs/vulkan-overview-mark-removed-adr0726 (2026-06-04)¶
Files touched: docs/backends/vulkan/overview.md, docs/api/vulkan-image-import.md, docs/state.md, changelog.d/chore/vulkan-docs-mark-removed.md
no rebase impact: docs-only changes. No C source, public header, upstream-mirrored Python, or Netflix golden-assertion file is touched. The changes add removal notices to two Vulkan documentation files that still described the backend as active after ADR-0726 removed it.
docs(post-rename): scrub residual libvmaf/ paths (ADR-0700)¶
no rebase impact: doc-only path corrections. All changed files are under docs/, AGENTS.md, CONTRIBUTING.md, and one comment in core/include/libvmaf/libvmaf_mcp.h. No C source files changed. No public headers changed (the comment in libvmaf_mcp.h is prose, not an include path). No Netflix golden assertions touched.
docs(usage,api): correct backend auto-priority + Doxygen drift in public headers¶
Files touched: docs/usage/vmafx-cli.md, docs/usage/vmaf-tune-score-backend.md, docs/usage/vmaf-tune.md, docs/usage/bench.md, docs/usage/ffmpeg.md, core/include/libvmaf/libvmaf.h, core/include/libvmaf/libvmaf_hip.h, core/include/libvmaf/AGENTS.md, core/include/libvmaf/model.h, changelog.d/changed/backend-autopriority-doxygen-drift.md, docs/rebase-notes.md.
No rebase impact: doc-only and Doxygen-only edits. No C source, public C symbol, ABI surface, Netflix golden assertion, or upstream-mirrored implementation is affected. The model.h change replaces a @field block with per-member inline comments — comment-only; no struct layout change.
docs/mcp-tools-audit-fixes¶
Files touched: docs/mcp/index.md, docs/mcp/tools.md, docs/mcp/http-transport.md, docs/mcp/release-channel.md, changelog.d/changed/mcp-tools-catalogue-audit-fixes.md, docs/rebase-notes.md.
No rebase impact: doc-only scrub. No C source, public headers, Netflix golden assertions, MCP server Python/Go source, or upstream-mirrored symbols are touched. No branch logic changed.
docs(post-vulkan-drop): residual scrub + fix -Denable_vulkan=true (invalid Meson) → =enabled¶
Branch: docs/post-vulkan-drop-residual-scrub
no rebase impact: docs-only change. Fixes stale Vulkan references in docs/ai/datasets/k150k.md, docs/mcp/tools.md, docs/api/index.md, and docs/api/gpu.md. No C source, public header, upstream-mirrored Python, or Netflix golden-assertion file is touched.
docs(rebrand): scrub residual Lusoris-fork references¶
Files touched: docs/usage/cli.md, docs/ai/mos-corpora.md, docs/ai/konvid-1k-ingestion.md, docs/ai/konvid-150k-ingestion.md, docs/development/release.md, docs/development/automated-rule-enforcement.md, docs/mcp/index.md, docs/architecture/c4-context.md, docs/architecture/c4-container.md, docs/metrics/bad-cases.md, CONTRIBUTING.md, AGENTS.md, changelog.d/changed/scrub-lusoris-fork-refs.md, docs/rebase-notes.md.
No rebase impact: doc-only text substitutions (branding strings, fork issue URL, HIP status text). No C source, public header, Netflix golden assertion, upstream-mirrored symbol, version string, or copyright header was modified.
docs(versions): bump stale Go + required-checks-count + Python pins¶
Files touched: CLAUDE.md, docs/development/languages.md, docs/development/release.md, docs/architecture/c4-context.md, docs/mcp/index.md, docs/getting-started/install/windows.md, docs/ai/training.md, changelog.d/changed/bump-stale-docs-go-checks-python-pins.md, docs/rebase-notes.md.
No rebase impact: docs-only scrub; no C source, public header, Netflix golden assertion, or upstream-mirrored symbol is affected.
docs(copyright): drop "and Claude (Anthropic)" from fork headers — residual sweep¶
Files touched: README.md, dev-llm/src/vmaf_dev_llm/__init__.py, scripts/lib/__init__.py, changelog.d/changed/copyright-drop-anthropic-residuals.md, docs/rebase-notes.md.
No rebase impact: text-only copyright-line change in three files missed by the ADR-0861 / ADR-0776 sweeps. No C source, public header, Netflix golden assertion, upstream-mirrored symbol, or build system touched.
docs(post-ansnr): scrub residual ANSNR references (PR #38 follow-up)¶
Files touched: docs/api/gpu.md, docs/backends/hip/overview.md, docs/backends/index.md, docs/backends/arm/overview.md, docs/backends/metal/index.md, docs/development/build-flags.md, docs/development/cross-backend-gate.md, docs/metrics/features.md, docs/mcp/tools.md, README.md, core/src/feature/metal/AGENTS.md, core/src/hip/AGENTS.md, core/src/feature/cuda/AGENTS.md, AGENTS.md, changelog.d/changed/post-ansnr-doc-scrub.md.
No rebase impact: doc-only changes (no C source, public header, Netflix golden assertions, or upstream-mirrored symbols affected). If an upstream Netflix/vmaf PR adds float_ansnr back, take the upstream side only in the C sources; the fork's doc changes apply only to fork-specific backend docs.
fix(rebrand): correct C++ badge (c++11→c++23) + drop Vulkan from GPU badge¶
Files touched: README.md, changelog.d/fixed/readme-badges-cpp23-drop-vulkan.md, docs/rebase-notes.md.
No rebase impact: doc-only edit to README.md badge lines; no C source, public header, Netflix golden assertion, or upstream-mirrored symbol is affected.
test(hip): parity coverage round 5 — speed_chroma + speed_temporal (2026-06-04, ADR-1004)¶
Files touched: core/test/test_hip_speed_chroma_parity.c, core/test/test_hip_speed_temporal_parity.c, core/test/meson.build, docs/adr/1004-hip-kernel-coverage-round5.md, docs/adr/README.md, docs/state.md, changelog.d/added/1004-hip-kernel-coverage-round5.md
no rebase impact: the two new test TUs are fork-local additions with no upstream analogue. The meson.build additions are append-only within the if hip_enabled block. No C source, public header, Netflix golden assertion, or upstream-mirrored Python file is modified.
chore/build-cpp-std-c23-bump (2026-06-04, ADR-1003)¶
Files touched: core/meson.build, core/AGENTS.md, core/test/meson.build, docs/adr/1003-cpp-std-c23-bump.md, docs/adr/README.md, changelog.d/changed/cpp-std-c23-bump.md
Rebase impact: Low. The cpp_std=c++11 → cpp_std=c++23 change in core/meson.build may conflict with any upstream Netflix/vmaf PR that also touches default_options. Netflix upstream still uses c++11; on conflict, keep c++23 (the fork's stated standard). The core/test/meson.build fix for test_feature_collector_coverage is fork-local; take the fork side on any conflict.
test(mcp-server): coverage push round 4¶
Files touched: mcp-server/vmaf-mcp/tests/test_coverage_round4.py, changelog.d/added/mcp-server-coverage-round4.md, docs/rebase-notes.md.
Rebase impact: None. Fork-local Python test file with no upstream analogue; no C source, public header, or Netflix golden-assertion file is touched.
test(sycl): parity coverage round 5 — CAMBI parity gate¶
Branch: test/sycl-parity-round5-cambi
no rebase impact: adds core/test/test_sycl_cambi_parity.c (new file, no upstream analogue), one meson.build registration block, ADR-1001, and a changelog fragment. No C source, public header, feature extractor implementation, or Netflix golden-assertion file is touched.
chore(rust): bump bindgen 0.69 → 0.72 + workspace edition 2021 → 2024 (ADR-1002)¶
Branch: chore/rust-edition-2024-bindgen-072
Touches: Cargo.toml, Cargo.lock, bindings/rust/vmafx/Cargo.toml, bindings/rust/vmafx-sys/Cargo.toml, core/src/feature/rust/tad/src/lib.rs, docs/adr/1002-rust-edition-2024-bindgen-072.md, changelog.d/chore/rust-edition-2024-bindgen-072.md.
No rebase impact on upstream Netflix/vmaf code. All changed files are fork-local Rust crates (vmafx-sys, vmafx, vmafx-tad) with no upstream analogue. No C source, public header, upstream-mirrored Python, or Netflix golden-assertion file is touched. A future upstream port cannot conflict with Rust workspace settings since Netflix/vmaf has no Rust code. The bindgen-consumed header paths consumed by bindgen remain at core/include/libvmaf/ (ADR-0700 path); any future upstream header change that adds or removes a symbol is handled automatically by re-running cargo build (bindgen regenerates on every build).
fix(cppcheck): resolve Whole-Project warnings¶
Files touched: core/src/feature/integer_ssim.c, core/src/picture_pool.cpp, core/src/read_json_model.cpp, core/tools/vmaf.cpp, core/test/test_ssimulacra2_simd.c, core/test/dnn/test_tensor_io.c, .cppcheck-suppressions.txt, changelog.d/fixed/cppcheck-whole-project-warnings.md
Rebase impact: None for upstream Netflix/vmaf cherry-picks. All changes are either fork-local files (picture_pool.cpp, opt.cpp suppression) or minimal defensive additions (null checks, format-specifier corrections, struct-member initialisation) that do not alter external behaviour. The %d → %u format fixes in vmaf.cpp and read_json_model.cpp are cosmetic; the VmafModel{} initialisation is semantically equivalent to memset(m, 0, …) on any IEEE-754 platform.
docs(coverage): ADR-0922 coverage-gate runbook (2026-06-04)¶
Files touched: docs/development/coverage-gate.md (new), changelog.d/added/coverage-gate-runbook.md (new)
Rebase impact: None. Documentation-only addition; no source, build, or CI files are modified.
fix(cppcheck): motion_avx512 missing sub-kernel functions¶
Files touched: core/src/feature/x86/motion_avx512.c, core/src/feature/x86/motion_avx512.h
Rebase impact: None. Both files are fork-local SIMD additions. The four new public symbols (sad_avx512, y_convolution_8_avx512, y_convolution_16_avx512, x_convolution_16_avx512) are additive and have no upstream Netflix/vmaf equivalents. No existing symbol is renamed, removed, or ABI-changed.
chore/tech-stack-badges-go-pin-bump (2026-06-04, ADR-1000)¶
Files touched: README.md, go.mod, .github/workflows/go-ci.yml, docs/adr/1000-tech-stack-badges-go-rust-pins.md, docs/adr/_index_fragments/1000-tech-stack-badges-go-rust-pins.md, docs/adr/_index_fragments/_order.txt, changelog.d/changed/tech-stack-badges-go-pin-bump.md
Rebase impact: None for C/SYCL/CUDA/HIP/Vulkan/Rust code. go.mod minimum version is bumped 1.25.0 → 1.26.4; this only affects builds that run go build / go test. Upstream Netflix/vmaf has no Go code, so no upstream cherry-pick will conflict with this change. The README badge block change is purely additive; no upstream port touches the README badge section.
fix/tsan-framesync-stdatomic-cxx (2026-06-04, ADR-0999)¶
Files touched: core/src/framesync.h, core/src/ref.h
Rebase impact: None. Both files are upstream-mirror headers touched only in the preprocessor guard section; no function signatures or struct members are changed. Upstream Netflix/vmaf does not compile feature_extractor.cpp as C++ (they use a C-only build), so this guard addition will not conflict with any upstream cherry-pick. ref.h guard widening from _MSC_VER to all C++ is backward-compatible: non-MSVC C compilers are unchanged (#if defined(__cplusplus) is false in C mode).
fix(metal): hoist feature_extractor.h above extern "C" in Metal .mm files¶
Files touched: core/src/feature/metal/float_moment_metal.mm, core/src/feature/metal/float_motion_metal.mm, core/src/feature/metal/float_ms_ssim_metal.mm, core/src/feature/metal/float_psnr_metal.mm, core/src/feature/metal/float_ssim_metal.mm, core/src/feature/metal/integer_motion_metal.mm, core/src/feature/metal/integer_motion_v2_metal.mm, core/src/feature/metal/integer_psnr_metal.mm
Rebase impact: None. All changed files are fork-local Metal backend sources. No upstream Netflix/vmaf files are touched. The change is purely an include-order fix (moves feature_extractor.h above its enclosing extern "C" block); no API, ABI, or algorithm change.
fix(arm64): guard framesync.h stdatomic include for C++ mode¶
Branch: fix/arm64-clang-stdatomic-cxx-conflict
Files touched: changelog.d/fixed/arm64-clang-stdatomic-cxx-framesync.md, docs/state.md, docs/rebase-notes.md.
no rebase impact: The framesync.h guard is already present via ADR-0999 (fix/tsan-framesync-stdatomic-cxx); this PR adds the ARM64-specific changelog fragment and state.md tracking row.
port/upstream-speed-chroma-simd-30f472b14 (2026-06-03, upstream 30f472b14)¶
Files touched: core/src/feature/x86/speed_avx2.c, core/src/feature/x86/speed_avx2.h, core/src/feature/x86/speed_avx512.c, core/src/feature/x86/speed_avx512.h, core/src/feature/speed.c, core/src/meson.build, core/test/test_speed_simd.c, core/test/meson.build
Rebase impact: Reduces delta — this port lands the upstream commit verbatim (new AVX2 + AVX-512 covariance-sum kernels, function-pointer dispatch). Future /sync-upstream passes that touch speed.c will see a smaller diff because the kernel dispatch pattern is now present on both sides. The compute_cov_kernel_fn typedef and SpeedState::compute_cov_kernel field are fork additions; any upstream change to the compute_covariance signature must also update the typedef here.
test/go-coverage-push (2026-06-04)¶
Files touched: cmd/vmafx-controller/{grpc_server.go,grpc_server_test.go,http_cancel_test.go,main_test.go,main_extra_test.go,auth/grpc_interceptor.go,auth/middleware.go,queue/queue_listall_test.go}, cmd/vmafx-mcp/impl.go, cmd/vmafx-node/{executor_test.go,main_test.go,online_feedback_pump_test.go}, cmd/vmafx-operator/internal/controller/{vmafxjob_applystatus_test.go,vmafxmodeltraining_applystatus_test.go,vmafxmodeltraining_controller.go,vmafxnode_controller.go}, cmd/vmafx-server/{grpc_server.go,http_cancel_test.go,main_extra_test.go}, pkg/observability/otel_instruments_test.go, pkg/score/grpc_client_unary_test.go
Rebase impact: Low. All changes are either test files (no rebase conflict possible on pure test additions) or targeted bug fixes in production code (grpc_server.go undefined-var fix, operator int32 type cast, MCP Vulkan backend dispatch). The auth ContextWithClaims export and probeHealthz method are additive. No public header or proto changes.
test/compat-python-vmaf-coverage-push (2026-06-03)¶
Files touched: compat/python-vmaf/tests/ (new directory), pyproject.toml (testpaths + pythonpath additions)
Rebase impact: None. Pure test addition; no production code changed. The pyproject.toml diff only appends to testpaths and pythonpath — if a concurrent branch adds entries in the same section a trivial conflict resolution is required (keep both entries).
vmafx-title-rebrand (2026-06-03, no ADR)¶
Files touched: README.md, mkdocs.yml, pyproject.toml, CONTRIBUTING.md
Rebase impact: None. All four files are fork-local metadata surfaces (project title, site name, package description, contributor heading). Upstream Netflix/vmaf does not touch any of these files; no merge conflict is possible on rebase.
feat(vmaf-tune): ADR-0498 follow-up #7 — encoder stats, x264 detection, backend dispatch, codec-list parser¶
Files touched: tools/vmaf-tune/src/vmaftune/encode.py, tools/vmaf-tune/src/vmaftune/fast.py, tools/vmaf-tune/src/vmaftune/codec_adapters/__init__.py, tools/vmaf-tune/tests/test_encode_dispatcher_per_adapter.py, tools/vmaf-tune/tests/test_adr_0498_followup7.py
Rebase impact: None. All changed files are fork-local to tools/vmaf-tune/; no upstream Netflix/vmaf files are touched. The _VERSION_PROBE_PATTERNS dict is additive (new keys only). The parse_available_codecs function is new; no existing symbol is renamed or removed. The _build_production_sample_extractor signature change (new backend=None kwarg) is backward-compatible. The test_encode_dispatcher_per_adapter.py fix (capture first call only) resolves a test fragility introduced by the probe-cache expansion; no merge conflict expected against Netflix upstream since that test is fork-added.
fix/cuda-duplicate-csf-r-definitions (2026-06-03)¶
Files touched: core/src/feature/cuda/integer_adm/adm_cm.cu
Rebase impact: None. Purely removes a duplicate code block introduced by a merge-order accident (PR #565 admin-merged while master already had the same helpers). No upstream file is touched; no public header changes.
feat/ai-run-manifest-12-scripts (ADR-0668 follow-up)¶
No rebase impact. Pure Python-only change to ai/scripts/train_konvid.py. No C/header files modified. No upstream Netflix/vmaf files touched. The only observable change is the addition of a train_konvid.manifest.json sidecar emitted after training completes.
cuda-adm-decouple-inline-ldg (2026-05-29, ADR-0773)¶
Files touched: core/src/feature/cuda/integer_adm/adm_csf.cu, core/src/feature/cuda/integer_adm/adm_cm.cu
Rebase impact: None. Both files are fork-added CUDA kernel translation units that do not exist in upstream Netflix/vmaf master (ADM CUDA port is fork-local). No rebase conflict is possible.
The change is a pure performance annotation: const T *__restrict__ pointer extraction before hot inner loops and __ldg() on all per-pixel DWT2 band reads. If upstream Netflix ever adds their own ADM CUDA port, these files will need to be re-reviewed against theirs; the F3 pattern should carry forward.
feat/vmafx-tune-go-stage4-report (ADR-0770)¶
No rebase impact: pure Go CLI and pkg/report additions. No upstream C/Python files modified. Files added: cmd/vmafx-tune/cmd/report.go, pkg/report/multi.go, pkg/report/multi_test.go, docs/adr/0770-vmafx-tune-go-stage4-report.md, changelog.d/added/vmafx-tune-go-stage4-report.md. Files modified: cmd/vmafx-tune/cmd/root.go (register report + ladder), cmd/vmafx-tune/AGENTS.md (invariants 8–9), docs/usage/vmafx-tune-go.md (Stage-4 section), docs/adr/README.md (new row), docs/rebase-notes.md (this entry).
doxygen-thread-safety-tags (2026-05-29, ADR-0788)¶
Files touched: core/include/libvmaf/libvmaf.h, core/include/libvmaf/picture.h, core/include/libvmaf/feature.h, core/include/libvmaf/model.h, core/include/libvmaf/dnn.h
Rebase impact: Low. These are comment-only additions. An upstream sync that modifies the same function signatures may create minor merge-fuzz on the Doxygen blocks; resolve by re-applying the @thread-safety tags to whatever the upstream version of the comment looks like.
containerfile-layer-optimization (ADR-0790, 2026-05-29)¶
Files touched: dev/Containerfile
Rebase impact: None. dev/Containerfile is fork-local (not present in upstream Netflix/vmaf). No rebase conflict is possible.
phase-4b8-c-abi-break-scoping (2026-05-29)¶
Files touched: docs/adr/0767-phase-4b8-c-abi-break-scoping.md, docs/research/research-0752-phase-4b8-c-abi-break-scoping.md, docs/adr/README.md, changelog.d/changed/0767-phase-4b8-c-abi-break-scoping.md
Rebase impact: No rebase impact. This is a scoping/design document with no source changes. The implementation PR (when it lands) will touch core/include/libvmaf/*.h and every ffmpeg-patches/ file — that implementation PR will carry its own rebase note cataloguing the specific header and patch changes. When upstream Netflix/vmaf adds symbols to libvmaf.h or model.h between now and the v4 implementation, the ADR-0767 removal list should be checked against the upstream additions to avoid removing a symbol upstream has just added.
docs/hip-picture-stub-comment-closeout (ADR-0299, 2026-06-03)¶
core/src/picture.h — comment on VMAF_PICTURE_BUFFER_TYPE_HIP_DEVICE updated to reflect that picture_hip.{c,h} is fully implemented (ADR-0299); the old text described it as a stub.
Rebase impact: NONE — comment-only change; no logic, no ABI delta.
chore/cambi-drop-vulkan-scaffold — remove CAMBI Vulkan scaffolding per ADR-0726 (2026-06-03)¶
No rebase impact on upstream C/Python code.
Files modified are fork-local: core/src/feature/vulkan/cambi_vulkan.c (deleted), core/src/feature/vulkan/shaders/cambi_{preprocess,derivative,filter_mode,decimate,mask_dp}.comp (deleted), core/test/test_cambi_vulkan.c (deleted), core/src/vulkan/meson.build (CAMBI source + shader entries removed), core/src/feature/cambi_internal.h (comment updated), core/src/feature/cuda/integer_cambi_cuda.c (comments updated), core/src/feature/hip/integer_cambi_hip.c (comment updated), changelog.d/removed/cambi-vulkan-scaffold.md (new).
Rebase impact: None on upstream sync (no Netflix file touched).
CI scaffold-comment refresh (2026-06-03)¶
.github/workflows/fuzz.yml — header comment updated: ADR-0882 citation added alongside ADR-0270/0311. .github/workflows/libvmaf-build-matrix.yml — Metal matrix lane comment and name: field updated from "T8-1 scaffold" to "runtime" (ADR-0420 landed).
no rebase impact: comment-only change; no logic or structure altered.
Single ledger of fork-local changes that need attention when this fork syncs from upstream/master (Netflix/vmaf). Required by ADR-0108: every fork-local
Second-opinion batch smoke scaffold + pytest path fix (ADR-0991, 2026-06-03)¶
Files touched: ai/pyproject.toml (add pythonpath = ["scripts"] to pytest config), ai/testdata/smoke-second-opinion-batch/ (new: batch.json, fixtures/*.jsonl, README.md), docs/adr/0991-second-opinion-batch-runs.md (new), docs/research/research-0991-second-opinion-batch-2026-06-03.md (new), changelog.d/fixed/0991-second-opinion-batch-pytest-path.md (new).
Rebase impact: None on upstream sync (no Netflix/vmaf upstream file touched). The ai/pyproject.toml addition is additive; no conflict risk.
controller-multi-tenant-auth-gateway (2026-05-29, ADR-0794)¶
Files touched: cmd/vmafx-controller/auth/ (new package), cmd/vmafx-controller/main.go, cmd/vmafx-controller/grpc_server.go, cmd/vmafx-controller/http_server.go, cmd/vmafx-controller/queue/queue.go, cmd/vmafx-controller/queue/schema.sql, deploy/helm/vmafx/crds/vmafx.dev_vmafxtenants.yaml (new), deploy/helm/vmafx/templates/tenant-crd-config.yaml (new), deploy/helm/vmafx/templates/deployment.yaml, deploy/helm/vmafx/values.yaml, docs/server/auth.md (new), docs/adr/0794-controller-multi-tenant-auth-gateway.md (new).
Rebase impact: None. All touched files are fork-local additions (vmafx-controller, Helm chart, docs) that do not exist in upstream Netflix/vmaf. The SQLite schema change (tenant_id column) is additive and non-breaking. No upstream rebase conflict is possible.
KoNViD / UGC / BVI-DVC saliency batch manifests (ADR-0993, 2026-06-03)¶
Files touched: ai/batch-manifests/saliency/konvid-150k.json (new), ai/batch-manifests/saliency/ugc.json (new), ai/batch-manifests/saliency/bvi-dvc.json (new), docs/ai/saliency-feature-materializer.md (corpus-specific manifests section), docs/adr/0993-konvid-ugc-bvi-saliency-batch-launch.md (new), docs/adr/README.md (index row), changelog.d/added/konvid-ugc-bvi-saliency-batch-manifests.md (new).
Rebase impact: None on upstream sync (no Netflix file touched). All new files are fork-local; no upstream path conflicts.
ADR-0992 — MOS-label batch-run manifests for KonViD and CHUG¶
Files touched: ai/configs/mos-label-batch-konvid.json (new), ai/configs/mos-label-batch-chug.json (new), ai/tests/test_mos_label_batch_runs_smoke.py (new), ai/tests/test_batch_materialize_mos_labels.py (sys.path bug fix), docs/ai/mos-label-materializer.md, docs/adr/0992-mos-label-batch-runs.md (new), docs/adr/README.md, changelog.d/added/0992-mos-label-batch-runs.md (new), and this file.
Rebase impact: No rebase impact on upstream sync (all touched files are fork-local; no Netflix/vmaf source file is modified). No cross-branch impact: the new ai/configs/*.json files are independent and will not conflict with any in-flight branch.
Changelog-fragment section hygiene (2026-05-30)¶
Files touched: changelog.d/perf/*.md → changelog.d/changed/perf-*.md (27 renames), changelog.d/performance/*.md → changelog.d/changed/perf-*.md (5 renames), changelog.d/README.md, release-please-config.json, docs/adr/0892-conventional-commits-and-changelog-fragment-hygiene.md (new), docs/research/0892-conventional-commits-audit-2026-05-30.md (new), changelog.d/fixed/conventional-commits-audit.md (new).
Rebase impact: None on upstream sync (no Netflix file touched). Cross-branch impact on fork: any in-flight feature branch holding a changelog.d/perf/*.md or changelog.d/performance/*.md file will hit a rename-detection conflict on rebase. git rebase with default -X settings detects the rename cleanly; if a conflict surfaces, the fix is to drop the in-flight branch's copy of the file and re-add the content under changelog.d/changed/perf-<topic>.md. The migrated files had their leading ### Performance / ## perf(…) headings stripped (renderer adds ### Changed itself); in-flight branches that added a new perf/ fragment should follow the same pattern.
See ADR-0892.
fix/ci-docs-pr-trigger — docs.yml PR trigger (2026-06-03, ADR-0986)¶
No rebase impact on upstream C/Python code.
Files modified are fork-local: .github/workflows/docs.yml (trigger + permissions update), docs/adr/0986-ci-docs-pr-trigger.md (new), docs/adr/_index_fragments/0986-ci-docs-pr-trigger.md (new), docs/adr/_index_fragments/_order.txt (appended), changelog.d/fixed/ci-docs-pr-trigger-0986.md (new), docs/rebase-notes.md (this entry).
Netflix upstream ships no GitHub Actions workflows. No rebase conflict is possible.
Research-0760 — Rust crate audit (docs + ADR-0707 correction, 2026-05-29)¶
No rebase impact on upstream C/Python code.
All files modified are fork-local: docs/research/research-0760-rust-crate-audit.md (new), changelog.d/added/rust-crate-audit-0760.md (new), docs/adr/0707-vmafx-rust-pilot-feature.md (corrected enable_rust_features default description from "true" to "false"), docs/rebase-notes.md (this entry).
Neither core/meson_options.txt, core/src/meson.build, nor any C/Rust source is modified. No Netflix upstream file is touched. No rebase conflict is possible.
fix/helm-node-deployment-deduplicate (2026-05-30, ADR-0713 / ADR-0719)¶
Files touched: deploy/helm/vmafx/templates/node.yaml (modified), deploy/helm/vmafx/templates/node-deployment.yaml (deleted).
Rebase impact: None. The deploy/helm/ tree is fork-only — Netflix upstream ships no Helm chart. The duplicate-Deployment collision and its fix live entirely within fork-added templates.
The two templates both rendered a Deployment named {{ include "vmafx.fullname" . }}-node under .Values.node.enabled, which made helm install fail with a duplicate-resource error and left Phase 4b distributed scoring uninstallable. The richer node.yaml (liveness/readiness probes, GPU resource injection, metrics port + Service, VMAFX_NODE_ID per ADR-0713) is kept; the rclone Secret mount + storage-mode / model-dir env vars from the deleted node-deployment.yaml were folded into node.yaml.
libvmaf.Score / ScoreDirect ctx.Context plumbing (2026-05-31, fix/libvmaf-score-ctx)¶
Files touched: pkg/libvmaf/libvmaf.go (Score signature: ctx as first param; exec.CommandContext + WaitDelay = 2s), pkg/libvmaf/direct.go (ScoreDirect signature: ctx as first param; per-frame ctx.Err() check at the top of the read+queue loop; rename of local ctx *C.VmafContext -> vmafCtx to avoid shadowing), pkg/libvmaf/libvmaf_test.go, pkg/libvmaf/direct_test.go (call-site updates + new cancel tests), cmd/vmafx-server/{http_server.go,grpc_server.go,http_cancel_test.go}, cmd/vmafx-controller/{http_server.go,grpc_server.go,http_cancel_test.go}, cmd/vmafx-node/executor.go, cmd/vmafx-mcp/impl_direct.go.
Rebase impact: All fork-local. pkg/libvmaf/ is a fork-only Go wrapper around the public libvmaf C ABI; cmd/vmafx-* are entirely fork-local binaries with no upstream counterparts. No headers in core/include/ were changed and no upstream-mirrored C source was touched, so upstream syncs cannot collide.
Action on next upstream sync: None. The C API surface (vmaf_init / vmaf_read_pictures / vmaf_score_pooled / vmaf_close) the Go layer wraps is unchanged; we only renamed a local C.VmafContext* variable inside Go.
vmafx-tune-go deep bug audit (2026-05-31, fix/vmafx-tune-go-audit-20260531)¶
Files touched: pkg/report/report.go, pkg/report/sanitize_test.go (new), pkg/bisect/bisect.go, pkg/bisect/nan_parse_test.go (new), pkg/bisect/timeout_test.go (new), pkg/encoder/encoder.go, pkg/encoder/discover.go, pkg/encoder/discover_test.go, pkg/encoder/discover_cache_test.go (new), pkg/encoder/timeout_test.go (new), cmd/vmafx-tune/cmd/compare.go, cmd/vmafx-tune/cmd/ladder.go, cmd/vmafx-tune/cmd/ladder_nan_test.go (new), changelog.d/fixed/0979-vmafx-tune-go-deep-bug-audit.md (new).
Rebase impact: Fork-local only. Every file lives under pkg/{report,bisect,encoder} or cmd/vmafx-tune/, which are 100% fork additions (the vmafx-tune-go Stage-1 surface from ADR-0705 / ADR-0713; no Netflix upstream counterpart exists). An upstream sync will not encounter conflicts on any of these files.
On-disk surface changes (relevant to in-tree callers):
- New public helper
report.SanitizeBisectSamples([]bisect.Sample) []any— exported so the schema-v2 sweep emitter incmd/vmafx-tune/cmd.emitSweepJSONcan apply the same nested NaN→null coercion the Python emitter (_nan_to_noneintools/vmaf-tune/src/vmaftune/compare.py) has used since the RFC-8259 hardening of 2026-05-17. - New env-var knobs
VMAFX_TUNE_ENCODE_TIMEOUT(default60m),VMAFX_TUNE_SCORE_TIMEOUT(default30m),VMAFX_TUNE_PROBE_TIMEOUT(default30s) for the ffmpeg / vmaf / ffprobe subprocess upper bounds. Operators can lower these in CI to fail-fast instead of hanging a job. - Codec-discovery cache key is now the binary path, not a one-shot
sync.Once. Callers that depended on the old "first probe wins forever" shape (none in tree as of this PR) will see a re-probe on binary-path change.
Python-surfaces bug-audit bundle (2026-05-31, fix/python-surfaces-bug-audit)¶
no rebase impact: REASON — fork-local Python files only. Touches: ai/src/corpus/base.py (fork-added, ADR-0371), ai/src/vmaf_train/data/{datasets,manifest_scan,feature_dump,frame_dataset,frame_loader}.py (fork-added tiny-AI training surface), and mcp-server/vmaf-mcp/src/vmaf_mcp/server.py (fork-added MCP server, no upstream equivalent). No core/src/ or upstream-mirror file is touched.
Fork-local files: ai/src/corpus/base.py, ai/src/vmaf_train/data/datasets.py, ai/src/vmaf_train/data/manifest_scan.py, ai/src/vmaf_train/data/feature_dump.py, ai/src/vmaf_train/data/frame_dataset.py, ai/src/vmaf_train/data/frame_loader.py, mcp-server/vmaf-mcp/src/vmaf_mcp/server.py, mcp-server/vmaf-mcp/tests/test_server.py, ai/tests/test_python_surfaces_bug_audit.py (new), mcp-server/vmaf-mcp/tests/test_python_surfaces_bug_audit.py (new), changelog.d/fixed/python-surfaces-bug-audit-2026-05-31.md (new), docs/research/0983-python-surfaces-bug-audit-2026-05-31.md (new).
chore/gosec-findings-fix-v2 (2026-06-01, ADR-0983)¶
no rebase impact: the Go surface (cmd/, pkg/, gen/, api/vmafx/v1/) is wholly fork-local. Netflix/vmaf has no Go code. The sweep touches only Go files plus .github/workflows/go-ci.yml, docs/adr/, docs/research/, changelog.d/security/, and the regression test cmd/vmafx-mcp/impl_gosec_test.go. No C, no SIMD, no GPU, no upstream-mirror file is touched.
Re-run of the earlier chore/gosec-findings-fix (PR #509, closed without merge) against the post-#505 / post-#508 master tip. The prior PR conflicted with PR #505's pkg/bisect/bisect.go + pkg/encoder/encoder.go exec.CommandContext + per-stage timeout plumbing; this v2 sweep applies the same security-hardening fixes while preserving the ctx + timeout. Same set of touched files; same single real bug fixed (describeModel path traversal). No markdownlint / formatter regression; both parser CI scripts green.
test_svm_parser link + vmafx-operator audit (2026-05-31, fix/test-svm-parser-link-plus-operator-audit)¶
Files touched: core/test/meson.build (added ../src/thread_locale.c to the test_svm_parser source list), api/vmafx/v1/vmafxjob_types.go, api/vmafx/v1/vmafxnode_types.go, api/vmafx/v1/vmafxmodeltraining_types.go, config/crd/bases/vmafx.dev_vmafxjobs.yaml, config/crd/bases/vmafx.dev_vmafxnodes.yaml, config/crd/bases/vmafx.dev_vmafxmodeltrainings.yaml, deploy/helm/vmafx/crds/*.yaml (synced copies), cmd/vmafx-operator/internal/controller/vmafxnode_controller.go, cmd/vmafx-operator/internal/controller/vmafxnode_probehealthz_test.go (new), cmd/vmafx-operator/internal/controller/vmafxmodeltraining_controller_branch_test.go (int32 casts).
Rebase impact: All changes are fork-local — the operator, vmafx.dev/v1 CRDs, and Helm chart are 100% additions on this fork (no upstream counterparts). The single upstream-mirrored file is core/test/meson.build; the change there is one-line additive (append '../src/thread_locale.c'), no conflict surface. No public C ABI is touched; the libsvm vendor remains observation-only per ADR-0889.
No load-bearing invariants; no AGENTS.md rebase-pin required.
core/src lifecycle + memory audit (2026-05-31, fix/core-lifecycle-memory-audit)¶
Files touched: core/src/picture_pool.c, core/src/model.c, core/src/model.cpp, core/src/predict.c, core/src/output.c, core/src/dict.c, core/src/feature/feature_collector.cpp, core/test/test_predict.c (new test case), core/test/test_model.c (new test case), core/test/test_output.c (new test case).
Rebase impact: Touch points are all upstream-mirrored TUs. Each fix is a narrow correctness patch (NULL guard, errno sign, errno code, missing return-value propagation, free-on-error path) — none of them changes the public C ABI, the entry-point list, or the data layout of any struct.
On upstream sync:
picture_pool.c::pool_preallocate_picturescleanup: trivialvmaf_picture_unref → aligned_freeswap on a fork-only code path (Netflix has nopool_preallocate_picturesin this form).model.{c,cpp}::vmaf_model_load+vmaf_model_collection_loadNULL guards: add theif (!version) return -EINVAL;block at the top of each function. Conflicts only if upstream reorders the body.predict.csign and propagation fixes: small textual deltas on upstream-mirrored functions. If upstream changes the sign convention, fall in line with upstream.output.cCSV/SUB NULL guards: paste the same three-line guard the XML/JSON writers already have (ADR-0602).dict.c::dict_normalize_numeric: one-word changestrtof→strtod. Conflicts only if upstream switches to a different parser entirely.feature_collector.cpp::aggregate_vector_append: one-word change-EINVAL→-ENOMEM.
No load-bearing invariants; no AGENTS.md rebase-pin required.
Markdown-lint full-ruleset discharge (2026-05-31, ADR-0980)¶
Files touched: ~1,400 .md files across docs/, .claude/, core/, ai/, tools/, bindings/, mcp-server/, scripts/, cmd/, top-level README/CONTRIBUTING/CODE_OF_CONDUCT. .markdownlint.json is unchanged.
Rebase impact: None for upstream-mirrored TUs. The added <!-- markdownlint-disable ... --> comments live only in fork-added / fork-modified .md files; upstream-vendored .md files in subprojects/, core/test/data/, python/test/resource/, compat/python-vmaf/resource/, compat/python-vmaf/matlab/, model/, and testdata/ are excluded from the gate (see .pre-commit-config.yaml markdownlint-cli2 exclude: regex) and are not touched by this PR.
On upstream sync, no resolution is required for .md files. If a future upstream PR adds a new fork-mirrored .md file that brings new violations, either fix the content or extend the per-file disable comment for that file; do not modify .markdownlint.json.
vmafx-server + pkg/score bug-audit (2026-05-31, ADR-0978)¶
Files touched: pkg/observability/observability.go, pkg/observability/observability_test.go, pkg/score/grpc_client.go, pkg/score/grpc_client_test.go, cmd/vmafx-server/grpc_server.go, cmd/vmafx-server/http_server.go, cmd/vmafx-server/main_test.go, cmd/vmafx-server/grpc_recovery_test.go (new).
Rebase impact: None. All five surfaces are fork-local Go code:
pkg/observability/is a fork-added package; Netflix/vmaf has no equivalent.pkg/score/is a fork-added wrapper around the fork's vmafx.v1 proto; Netflix/vmaf does not ship a gRPC client.cmd/vmafx-server/is a fork-added binary (ADR-0703); Netflix/vmaf has no equivalent gRPC + HTTP scoring service.
No C ABI, no public header, no upstream-mirrored TU touched. Upstream syncs do not interact with this change.
core/tools input-reader safety (2026-05-31, ADR-0977)¶
Files touched: core/tools/y4m_input.c, core/tools/yuv_input.c, core/tools/vmaf_bench.c, core/test/test_y4m_alloc_failure.c (new), core/test/meson.build.
Rebase impact: PARTIAL. y4m_input.c and yuv_input.c are vendored from Daala via upstream Netflix/vmaf; upstream still carries the unchecked malloc returns and the int-precision dst_buf_sz arithmetic. On any upstream sync that touches the y4m / yuv parsers, keep our (size_t) casts and the explicit if (!_y4m->dst_buf) return -1; block — the diff is localised (the size-arithmetic stanza lines and the return 0; tail of y4m_input_open_impl).
vmaf_bench.c is fork-only (no upstream churn).
Sync action: Mechanical merge if upstream touches the same lines: prefer the fork side at the size-arithmetic stanza and the malloc-failure block in y4m_input_open_impl, prefer the fork bench_cleanup label structure in vmaf_bench::bench_feature.
Test suite: NULL-check malloc sweep (2026-05-31, ADR-0971)¶
Files touched: core/test/test_ssimulacra2_simd.c, core/test/test_framesync.c, core/test/test_pic_preallocation.c, core/test/AGENTS.md.
Rebase impact: None. All changes are purely additive NULL-checks in test-only files. Netflix/vmaf does not carry these test files upstream (test_ssimulacra2_simd.c, test_pic_preallocation.c are fork-added; test_framesync.c has fork-local modifications). No C API or public ABI is touched. Subsequent upstream syncs do not interact with this change.
Public-header ISO-reserved include guards renamed (2026-05-31)¶
Files touched: core/include/libvmaf/libvmaf.h, core/include/libvmaf/picture.h, core/include/libvmaf/feature.h, core/include/libvmaf/model.h, core/include/libvmaf/macros.h, core/include/libvmaf/vmaf_assert.h, core/include/libvmaf/dnn.h, core/include/libvmaf/libvmaf_cuda.h, core/include/libvmaf/libvmaf_sycl.h.
Rebase impact: REAL. Six of the nine renamed headers (libvmaf.h, picture.h, feature.h, model.h, libvmaf_cuda.h, plus arguably dnn.h if upstream ever ports the tiny-AI surface) are upstream-mirrored from Netflix/vmaf. Upstream still ships the ISO-reserved __VMAF_*__ guard pattern that SEI CERT DCL37-C bans (ADR-0972).
Sync action: On any upstream sync that touches these six headers, keep our LIBVMAF_<BASENAME>_H lines and drop the upstream __VMAF_*__ ones. The diff is mechanical (3 lines per header — the #ifndef, the #define, and the closing #endif comment); no semantic merge required. The full guard-rename table is in ADR-0972 §Decision.
The remaining three renamed headers (macros.h, vmaf_assert.h, libvmaf_sycl.h) are fork-only and never receive upstream churn.
If a future upstream sync changes the LIBVMAF_* pattern itself (e.g. Netflix adopts the same fix with a different spelling), reopen ADR-0972 to decide whether to converge.
Rust vmafx safe binding crate scaffold (2026-05-31)¶
Files touched: Cargo.toml (workspace), bindings/rust/vmafx/ (new crate).
Rebase impact: None. Netflix/vmaf has no Rust bindings upstream. The new crate is a pure addition under bindings/rust/, parallel to the existing vmafx-sys crate (ADR-0706). The workspace Cargo.toml gains one members entry; no upstream file is touched. Subsequent upstream syncs do not interact with this code.
If a future upstream PR adds a Rust workspace (extremely unlikely), the fork's bindings/rust/vmafx/ and bindings/rust/vmafx-sys/ paths must not collide with the upstream layout. As of n8.1 there is no precedent.
gRPC ScoreStream Phase 1 (2026-05-31)¶
Files touched: proto/vmafx.proto, gen/go/vmafx.pb.go, gen/go/vmafx_grpc.pb.go, cmd/vmafx-server/grpc_server.go, cmd/vmafx-server/AGENTS.md, pkg/score/grpc_client.go, pkg/score/grpc_client_test.go, pkg/score/AGENTS.md, docs/architecture/grpc-streaming.md, docs/architecture/index.md, docs/adr/0933-grpc-streaming-multi-frame-scoring.md, docs/adr/_index_fragments/0933-grpc-streaming-multi-frame-scoring.md, docs/adr/_index_fragments/_order.txt, docs/adr/README.md, changelog.d/added/0933-grpc-streaming-phase1.md.
Rebase impact: None against upstream — this surface is entirely fork-local (Netflix/vmaf has no Go gRPC service). The proto package stays vmafx.v1; the unary Score / Health RPCs are unchanged. ScoreStream is purely additive. The Phase 1 server handler returns codes.Unimplemented after validating the opening StreamConfig.
If a future upstream port touches core/ in a way that changes the public C API consumed by pkg/libvmaf, the Phase 2 wiring of ScoreStream to libvmaf will need to mirror that change — but Phase 1 is server-stub-only and doesn't reach the C surface yet.
Native bash pre-commit hook (ADR-0924, 2026-05-31)¶
no rebase impact: all paths are fork-local — scripts/githooks/ (new directory), docs/development/pre-commit-hooks.md, docs/adr/0924-*.md, docs/research/0924-*.md, changelog.d/added/native-pre-commit-hooks.md. The Makefile changes rename hooks-install → install-hooks (with the old name kept as a legacy alias), in a fork-only target that upstream Netflix does not define. No upstream-mirrored file is touched.
Metal kernel parity tests round 3 (2026-05-31)¶
Files touched: core/test/meson.build, core/test/test_metal_integer_motion_parity.c (new), core/test/test_metal_float_motion_parity.c (new), core/test/test_metal_float_moment_parity.c (new), core/test/test_metal_float_ms_ssim_parity.c (new)
Rebase impact: None. Closes the per-kernel parity coverage gap for the remaining four Metal extractors after PR #351 (registration audit) and PR #379 (round-2 parity: motion_v2, integer_psnr, float_psnr, float_ssim). All four new files live under the existing fork-local enable_metal block in core/test/meson.build (the entire Metal backend is fork-added — ADR-0361 / ADR-0421 / ADR-0589 / T8-2a — and absent from upstream Netflix/vmaf). The block edit appended four new executable() + test() pairs immediately after the round-2 block (PR #379 has since merged); the surrounding endif boundaries are untouched so upstream syncs cannot conflict here.
If upstream ever ports a Metal backend, the test files would need re-pointing at the upstream kernel names; the synthetic-fixture + -ENODEV skip pattern carries forward unchanged.
vmaf-tune coverage push — lowest-covered modules (2026-05-31)¶
Files touched: tools/vmaf-tune/tests/test_coverage_push_lowcov_modules.py, changelog.d/added/vmaf-tune-coverage-push.md.
Rebase impact: None. tools/vmaf-tune/ is fork-only (no upstream Netflix counterpart); the new test file imports only public + underscore- prefixed seams that already existed in the package. The 92 added tests are pure unit-level (no subprocess / no ffmpeg / no ONNX / no GPU) and exercise documented error paths in uncertainty.py, _gop_common.py, proxy.py, predictor_features.py, benchmark.py, encoder_profile.py, and fast.py. If a future refactor renames any of the targeted internal helpers (_parse_fps, _run_probe_encode, _run_signalstats, _parse_frame_sizes, _mean, _resolve_baseline, _row_encode_fps, _row_score_fps, _resolve_model_path), update the corresponding import in this single test file.
Core MCP transport coverage push (2026-05-31)¶
Files touched: core/test/test_mcp_coverage.c (new), core/test/meson.build, changelog.d/added/core-mcp-coverage-push.md, docs/research/core-mcp-coverage-push-2026-05-31.md.
Rebase impact: None. The embedded MCP server (core/src/mcp/, core/include/libvmaf/libvmaf_mcp.h) is fork-only — upstream Netflix/vmaf has no MCP surface — so this test-only push is fully self-contained and never lands on a Netflix file. If upstream ever adds an MCP-shaped surface, treat the test as canonical fork-side coverage and reconcile by name. Companion: ADR-0108 deliverables in docs/research/core-mcp-coverage-push-2026-05-31.md.
phase3-subset-sweep readonly-view fix (2026-05-31)¶
Files touched: ai/scripts/phase3_subset_sweep.py, ai/tests/test_phase3_subset_sweep_unit.py.
Rebase impact: None — ai/scripts/phase3_subset_sweep.py is fork-original (Research-0027 Phase-3 tooling, no upstream Netflix analogue). The fix tightens an internal contract (_standardize_inplace now refuses read-only inputs and the caller forces a writeable copy via to_numpy(copy=True)); there is no public API change and no coupling to upstream files. Safe to carry through any upstream sync.
GPU runtime error-path leak fixes (ADR-0960, 2026-05-31)¶
no rebase impact: REASON — all changes are in fork-local error paths of core/src/cuda/common.c (new fail_after_stream label between two existing labels) and core/src/picture_pool.c (one pthread_cond_signal call and two pic->priv = NULL assignments). No upstream Netflix/vmaf logic is altered. The new test file core/test/test_picture_pool_error_paths.c is wholly fork-added with no upstream counterpart.
queue PullWork rollback on post-update Get failure (2026-05-31, ADR-0961)¶
no rebase impact: pure Go controller-internal fix. cmd/vmafx-controller/queue/ is entirely fork-added (no upstream Netflix/vmaf equivalent); upstream syncs do not touch this subtree.
ai/src NaN propagation guards — eval.correlations + tune._read_best_metric (2026-05-31, ADR-0963)¶
Files touched: ai/src/vmaf_train/eval.py, ai/src/vmaf_train/tune.py, ai/tests/test_eval_correlations.py, ai/tests/test_tune_objective.py.
Rebase impact: None — ai/src/vmaf_train/ is entirely fork-local with no upstream Netflix/vmaf equivalent. No C surface is touched. No upstream coupling.
Helm chart seccompProfile + node-deployment image helper (2026-05-31, ADR-0969)¶
no rebase impact: REASON — both changes are entirely within deploy/helm/vmafx/ which is fork-added infrastructure with no upstream counterpart in Netflix/vmaf. Netflix upstream does not ship a Helm chart; upstream syncs never touch this directory. PR #439 (ADR-0930) has since merged cleanly on top (it modified values.yaml in a non-conflicting block and did not touch node-deployment.yaml).
MCP HTTP transport security hardening (2026-05-31, ADR-0967)¶
no rebase impact: REASON — changes are confined to the fork-local MCP server subtree (mcp-server/vmaf-mcp/). Netflix upstream has no MCP server; this entire subtree will never merge upstream. The security middleware, auth helpers, and bind-host resolver are fork-invented code with no upstream counterpart.
HIP kernel parity-test coverage round 4 (2026-05-31, ADR-0958)¶
Files touched: core/test/test_hip_ssimulacra2_parity.c, core/test/test_hip_float_ssim_parity.c, core/test/meson.build.
Rebase impact: Low — the 2 new tests are fork-added consumers of fork-added HIP feature extractors (ssimulacra2_hip, float_ssim_hip). Upstream Netflix has no HIP backend, so neither the test sources nor the meson registration block has an upstream-mirror analogue. The skip-on--ENOSYS contract matches the round-1/2/3 template (PR #351 / PR #372 / PR #443) — if upstream ever ships a HIP backend the tests can be kept verbatim; their CPU side calls only public C-API entry points (vmaf_init, vmaf_use_feature, vmaf_read_pictures, vmaf_feature_score_at_index, vmaf_close) that are upstream-stable.
The round-4 plan also covered speed_chroma_hip / speed_temporal_hip parity gates, but those were deferred when the container build surfaced a pre-existing latent link defect — the helpers speed_internal_init_dimensions / speed_internal_float_stride are declared in core/src/feature/speed_internal.h but never defined. The same defect blocks the analogous CUDA / SYCL speed-family TUs from linking (none are currently wired into their respective meson archives). A follow-up PR adding core/src/feature/speed_internal.c will unblock all three GPU backends simultaneously. Tracked as T-HIP-SPEED-INTERNAL-IMPL-MISSING-2026-05-31 in docs/state.md.
Companion: docs/adr/0958-hip-kernel-coverage-round4.md, docs/research/0958-hip-kernel-coverage-round4-2026-05-31.md, changelog.d/added/0958-hip-kernel-coverage-round4.md.
Controller infrastructure fixes — StreamJobs + reaper stop signal (2026-05-31, ADR-0962)¶
No rebase impact: all changes are confined to the fork-local controller package (cmd/vmafx-controller/) and the Queue interface in cmd/vmafx-controller/queue/queue.go. Netflix upstream does not own these paths (the controller is a Phase 4b addition, not a port of Netflix code). The nodes.Registry context-propagation change is entirely within fork-local code and has no interaction with libvmaf C sources.
vmaf_mcp_stop() idempotent (CAS instead of exchange) (2026-05-31)¶
Files touched: core/src/mcp/mcp.c, core/test/test_mcp_stop_idempotent.c, core/test/meson.build.
Rebase impact: None — core/src/mcp/ is fork-only (Netflix has no MCP surface). The fix replaces three atomic_exchange(running, 2) + dual-value-guard pairs with three atomic_compare_exchange_strong(expected=1, desired=2) calls, keeping the existing 3-state state machine semantics intact and matching the CAS pattern already used by vmaf_mcp_start_{stdio,uds,sse}. The new regression test (test_mcp_stop_idempotent.c) is also fork-only. Sync impact: no Netflix file references vmaf_mcp_* symbols.
compat/python-vmaf/ scanf + ProcessRunner locale fixes (2026-05-31, ADR-0955)¶
Files touched: compat/python-vmaf/tools/scanf.py, compat/python-vmaf/__init__.py, python/test/python_harness_scanf_locale_bugs_test.py (new fork-only test).
Rebase impact: Medium. Both fixes live inside the upstream-mirror tree (compat/python-vmaf/), so a future upstream sync may overwrite them.
tools/scanf.py::makeFormattedHandler.applyWidth— the upstream code has an inverted width guard:
def applyWidth(handler):
if width is None:
return makeWidthLimitedHandler(handler, width, ignoreWhitespace=True)
return handler
The fork swaps the branches so implicit-width converters return handler and explicit-width converters return the capped wrapper. When porting an upstream commit that re-touches this function, verify the swapped semantics are preserved. If Netflix has independently fixed the same bug, drop the fork delta and update ADR-0955's status to Superseded by upstream.
__init__.py::ProcessRunner.run— upstream sets the C locale viaenv.setdefault("LC_ALL", "C")/env.setdefault("LANG", "C"). The fork replaces bothsetdefaultcalls with unconditional assignment (env["LC_ALL"] = "C"/env["LANG"] = "C") so a parent shell with non-EnglishLC_ALL/LANGcannot defeat the override. When porting an upstream commit that re-touchesProcessRunner.run, preserve the unconditional assignment pattern.
The regression test python/test/python_harness_scanf_locale_bugs_test.py exercises both code paths and will fail if either fix regresses during an upstream sync.
GPU dispatch-runtime host-only unit test (2026-05-31, ADR-0954)¶
Files touched: core/test/test_gpu_dispatch_runtime.c (new), core/test/meson.build.
Rebase impact: Low. The new test executable is fork-local — upstream Netflix/vmaf does not ship the gpu_dispatch_env, gpu_dispatch_parse, or per-backend dispatch_strategy TUs targeted by the test (those are all ADR-0181 / ADR-0488 / ADR-0483 fork additions). The wiring in core/test/meson.build lives in the fork-added test region near other test_* entries; no upstream collision is possible. If upstream ever adds dispatch-strategy abstractions of its own, the test would coexist by name.
Python harness coverage push round 2 (2026-05-31)¶
Files touched: python/test/python_harness_coverage_test.py (new — 82 cases).
Rebase impact: None. The new test file lives under python/test/, exercises only fork-touched modules under compat/python-vmaf/, and does not modify any Netflix golden assertAlmostEqual value (CLAUDE.md §8). Upstream Netflix has no analogue at the compat/ path (that subtree exists because of ADR-0700). When /sync-upstream runs, this file is fork-only and needs no re-baselining. Companion: PR #412 (test/compat-python-vmaf-coverage) round 1, PR #413 (fix/decorator-persist-encode) — neither overlap.
HIP ADM parity test feature-name + ENOSYS skip (ADR-0950, 2026-05-31)¶
Files touched: core/test/test_hip_adm_parity.c.
Rebase impact: no rebase impact: Netflix/vmaf upstream has no HIP backend at all (HIP is a fork-exclusive backend per ADR-0212); theadm_hipextractor and its parity test only exist on this fork. There is no upstream counterpart to reconcile during sync. Companion fix to ADR-0949 (motion3 sibling); both tests now follow the same two-axis (enable_hip × enable_hipcc) skip predicate. Companion docs: docs/adr/0950-hip-adm-parity-feature-name-and-enosys-skip.md, changelog.d/fixed/0950-test-hip-adm-parity-feature-name-and-enosys-skip.md.
go-services-coverage-round2 (2026-05-31)¶
Files touched: cmd/vmafx-tune/cmd/unit_internal_test.go, cmd/vmafx-tune/cmd/unit_internal_fixtures_test.go, cmd/vmafx-controller/grpc_server_test.go, cmd/vmafx-controller/queue/queue_extra_test.go, pkg/encoder/version_extract_test.go, changelog.d/added/go-services-coverage-round2.md.
Rebase impact: None. The Go cmd/ and pkg/ trees are wholly fork-added — upstream Netflix/vmaf has no Go layer. All new files are test-only and never enter the libvmaf C build, the Python harness, or the FFmpeg patch stack. No production code is touched, so the upstream rebase boundary is unaffected. The cmd/vmafx-controller grpc_server tests carry the //go:build cgo tag mirroring the production source file, so they compile only when cgo is enabled (matching the existing main_test.go invariant).
dev/Containerfile libvmaf → core path fix (2026-05-31, ADR-0966)¶
No rebase impact: pure path fix, no upstream coupling. dev/Containerfile is entirely fork-local and the only change is substituting three occurrences of the old source-directory name libvmaf/ with core/ following the ADR-0700 rename. If a future sync touches dev/Containerfile (unlikely — Netflix does not ship a dev container), re-run grep -n 'libvmaf/' dev/Containerfile to confirm no stale references were re-introduced by the merge. The library output name (libvmaf.so) and stage name (libvmaf-build) are intentionally preserved as references to the product, not the source directory.
SIMD bit-exactness round-2 — SSIMULACRA 2 FMA unification + lib-FP-model extension (2026-05-30, ADR-0891)¶
CUDA kernel parity coverage round 3 (2026-05-31)¶
Files touched: core/test/test_cuda_float_psnr_parity.c, core/test/test_cuda_float_vif_parity.c, core/test/test_cuda_float_ms_ssim_parity.c, core/test/test_cuda_float_moment_parity.c, core/test/test_cuda_ssimulacra2_parity.c, core/test/meson.build (+5 executable() + test() blocks under the existing if get_option('enable_cuda') guard, suite ['fast', 'gpu']), docs/adr/0947-cuda-kernel-coverage-round3.md, docs/adr/README.md (+1 row), docs/adr/_index_fragments/_order.txt (+1 line), docs/research/cuda-kernel-coverage-round3-2026-05-31.md, changelog.d/added/cuda-kernel-coverage-round3.md.
Rebase impact: None. All five test files are fork-local (test_cuda_*_parity.c pattern is fork-only; upstream Netflix/vmaf has no equivalent test scaffold). core/test/meson.build edits are additive blocks inside the existing enable_cuda guard — no upstream file in this region. If upstream Netflix adds new CUDA kernels with matching names (float_psnr_cuda, float_vif_cuda, float_ms_ssim_cuda, float_moment_cuda, ssimulacra2_cuda), the parity tests continue to work unchanged. If upstream adds new test files near test_integer_vif_cpu_cuda_parity (the closest neighbour in meson.build) the additive blocks may need re-anchoring — trivial 3-way merge.
PRs #351 (round 1) and #374 (round 2) both inserted test entries under the same enable_cuda guard in core/test/meson.build and have since merged; the sequential three-way merges resolved cleanly at landing time.
ADR template — optional supply-chain / SBOM / carbon sections (2026-05-31)¶
Files touched: docs/adr/0000-template.md, docs/adr/README.md
Rebase impact: None. Upstream Netflix/vmaf does not maintain an ADR template; the entire docs/adr/ tree is fork-local. The new optional sections (## Supply-chain impact, ## SBOM delta, ## Carbon / footprint) appear between ## Consequences and ## References. No upstream conflict surface.
vmafx-operator zap → slog uniformity (2026-05-31)¶
Files touched: cmd/vmafx-operator/main.go, cmd/vmafx-operator/internal/controller/suite_test.go, cmd/vmafx-operator/AGENTS.md, go.mod, go.sum.
Rebase impact: None against Netflix/vmaf (the operator is a fork-only Go package; upstream ships no Kubernetes operator). Rebase impact does exist against the kubebuilder v4 template itself: future scaffold upgrades will re-introduce sigs.k8s.io/controller-runtime/pkg/log/zap imports in main.go and suite_test.go. When re-running kubebuilder edit / operator-sdk init, re-apply the slog bridge:
main.go: replace the zap.Options block withslog.NewJSONHandler(os.Stderr, &slog.HandlerOptions{Level: ...})passed throughlogr.FromSlogHandler.suite_test.go: replacezap.New(zap.WriteTo(GinkgoWriter), zap.UseDevMode(true))withslog.NewTextHandler(GinkgoWriter, &slog.HandlerOptions{Level: slog.LevelDebug})throughlogr.FromSlogHandler.
The cmd/vmafx-operator/AGENTS.md invariant #6 documents this; check it before merging any upstream-template re-sync PR.
MCP server cgo direct path Phase 1 (2026-05-31, ADR-0931)¶
Files touched: pkg/libvmaf/direct.go, pkg/libvmaf/errors.go, pkg/libvmaf/direct_test.go, pkg/libvmaf/errors_test.go, pkg/libvmaf/AGENTS.md, cmd/vmafx-mcp/impl.go, cmd/vmafx-mcp/impl_direct.go, cmd/vmafx-mcp/impl_direct_test.go, cmd/vmafx-mcp/AGENTS.md.
Rebase impact: None against Netflix upstream. The change is entirely fork-local: it adds a new in-process cgo scoring path (ScoreDirect, ValidateModel) to pkg/libvmaf/ (which does not exist upstream) and wires two MCP tool handlers (vmaf_score, describe_model) in cmd/vmafx-mcp/ (which also does not exist upstream) to take that path when VMAFX_MCP_DIRECT=1. The libvmaf public C ABI used (vmaf_init / vmaf_use_features_from_model / vmaf_read_pictures / vmaf_score_pooled / vmaf_model_load_from_path / vmaf_picture_alloc / vmaf_picture_unref / vmaf_model_destroy / vmaf_close) is the canonical entry-point set documented in core/include/libvmaf/; the upstream signatures change rarely and any rename would already break core/tools/vmaf.c, so this code rides along.
If upstream renames or removes any of those entry points, update pkg/libvmaf/direct.go to match, then run the unit suite (LD_LIBRARY_PATH=$(pwd)/core/build-cpu/src go test ./pkg/libvmaf/ ./cmd/vmafx-mcp/).
OpenTelemetry tracing full roll-out — ADR-0782 (2026-06-03)¶
Files touched: pkg/observability/otel_instruments.go (new), cmd/vmafx-controller/grpc_server.go, cmd/vmafx-node/executor.go, cmd/vmafx-node/main.go, cmd/vmafx-server/main.go, cmd/vmafx-mcp/main.go, cmd/vmafx-controller/queue/queue.go, deploy/grafana/vmafx-overview.json (new), deploy/helm/vmafx/templates/otel-collector-sidecar.yaml (new), docs/observability/otel.md (new), docs/adr/0782-otel-tracing.md (new).
Rebase impact: None on the Netflix/vmaf C tree. Entirely fork-local Go instrumentation. No upstream C surfaces touched. The otel_instruments.go file is a pure addition; span call-sites follow the StartSpan/EndSpan pattern and do not change function signatures. If a future port touches executor.go or grpc_server.go, the span stanzas are additive and do not conflict with upstream semantics.
OpenTelemetry traces + metrics — Phase 1 (2026-05-31)¶
Files touched: pkg/observability/otel.go (new), pkg/observability/otel_test.go (new), pkg/observability/AGENTS.md (new), cmd/vmafx-controller/main.go, cmd/vmafx-controller/grpc_server.go, docs/development/observability.md (new), docs/adr/0927-opentelemetry-traces-metrics-phase1.md (new), go.mod, go.sum.
Rebase impact: None on the Netflix/vmaf C tree. The change is entirely fork-local Go code under pkg/observability and cmd/vmafx-controller. Upstream Netflix/vmaf has no Go services, so there is no cross-repo file to reconcile on sync. The added OTel dependencies (go.opentelemetry.io/otel, otelgrpc, OTLP HTTP exporters) live in go.mod and do not touch the C build.
When Phase 2 wires OTel into vmafx-node / vmafx-server / vmafx-mcp / vmafx-tune, follow the call-site pattern documented in pkg/observability/AGENTS.md (the 5 s bounded shutdown is mandatory). Each subsequent service ships as its own PR with its own ADR.
mkdocs ADR nav restructure + by-tag generator (2026-05-31)¶
Files touched: mkdocs.yml, scripts/docs/generate-adr-nav.sh, scripts/docs/generate-adr-by-tag.sh, docs/adr/by-tag/*.md (auto-generated, 443 files), docs/adr/0937-mkdocs-nav-decade-buckets.md, docs/adr/_index_fragments/0937-*.md, docs/adr/_index_fragments/_order.txt (append).
Rebase impact: None. All files are fork-only:
- Upstream Netflix/vmaf has no
mkdocs.yml, nodocs/adr/tree, and noscripts/docs/directory. - The sentinel-bounded splice region in
mkdocs.yml(# >>> ADR-NAV-GENERATED/# <<< ADR-NAV-GENERATED) is fork-local and unaffected by any upstream doc reorganisation. - The
docs/adr/by-tag/tree is regenerated byscripts/docs/generate-adr-by-tag.sh --write; on every ADR add / edit / tag-edit, re-run the script (or rely on the--checkCI gate once wired into.github/workflows/docs.yml).
If the per-hundred bucket labels in LABELS inside scripts/docs/generate-adr-nav.sh drift away from the actual bucket themes (e.g., the 0800s and 0900s fill out with a clear topic), edit the dict and re-run --write.
BuildKit cache mounts on container build matrix (2026-05-31)¶
Files touched: Dockerfile, docker/Dockerfile.production-gpu, dev/Containerfile, Dockerfile.go-server, docs/adr/0923-buildkit-cache-mounts.md, changelog.d/changed/buildkit-cache-mounts.md.
Rebase impact: None. These four Dockerfiles are fork-local (Netflix's upstream has only the top-level Dockerfile which we already heavily customise; the production-gpu / dev / go-server trio are wholly fork-added). The change introduces three patterns worth preserving across rebases:
# syntax=docker/dockerfile:1.7header at the top of each file.RUN --mount=type=cache,target=/var/cache/apt,sharing=locked --mount=type=cache,target=/var/lib/apt,sharing=locked apt-get ...on every apt invocation, with the matchingrm -rf /var/lib/apt/lists/*cleanup REMOVED.RUN --mount=type=cache,target=$CCACHE_DIR,sharing=locked CCACHE_DIR=... <build command>around every meson/ninja/cmake invocation;ccacheinstalled as a build dependency; FFmpeg gets--cc='ccache gcc' --cxx='ccache g++'; cmake gets-DCMAKE_{C,CXX}_COMPILER_LAUNCHER=ccache.
If upstream Netflix adds new RUN apt-get install lines to the top-level Dockerfile, prepend the apt cache mount pair. If they add new C/C++ compile steps, wrap them with the ccache mount + env var.
The vmaf user uid/gid is now explicitly pinned to 1000 in dev/Containerfile so BuildKit --mount=...,uid=1000,gid=1000 directives resolve to the same identity that runs the build — preserve that pin on rebase.
Pre-existing test failures across ai/, vmaf-tune, mcp-server (2026-05-30)¶
Files touched: ai/tests/conftest.py, ai/tests/test_codec_aware_fr.py, ai/tests/test_dnn_exporter_run_provenance.py, ai/tests/test_export_roundtrip.py, ai/tests/test_qat_smoke.py, ai/tests/test_registry.py, ai/tests/test_train_fr_regressor_v2_ensemble_loso_train.py, ai/tests/test_train_fr_regressor_v3.py, ai/tests/test_tune_cli.py, ai/tests/test_variance_mode.py, ai/tests/test_conftest_pytorch_lightning_guard.py (new), ai/pyproject.toml, tools/vmaf-tune/src/vmaftune/ladder.py, tools/vmaf-tune/tests/test_ladder.py, mcp-server/vmaf-mcp/src/vmaf_mcp/http_transport.py, mcp-server/vmaf-mcp/tests/test_http_transport.py.
Rebase impact: None. All three touched subsystems are fork-local:
ai/— entirely fork-added (tiny-AI training); upstream Netflix/vmaf has no Python training package.tools/vmaf-tune/— fork-added recommendation tool; upstream has no equivalent.mcp-server/vmaf-mcp/— fork-added MCP JSON-RPC server; upstream has no equivalent.
No cross-repo conflict possible. The requires_pytorch_lightning() helper in ai/tests/conftest.py is a generic environment-probe pattern that will keep working unchanged for any future torch / torchvision / torchmetrics ABI drift; the only knob to revisit is whether to widen the broad except Exception if some future failure mode warrants more specific handling.
Unified Python test orchestrator — top-level noxfile.py (2026-05-31, ADR-0914)¶
Files touched: noxfile.py (new), docs/development/python-test-orchestrator.md (new), docs/adr/0914-unified-python-test-orchestrator.md (new), docs/adr/_index_fragments/0914-unified-python-test-orchestrator.md (new), docs/adr/_index_fragments/_order.txt, docs/research/0914-python-test-orchestrator-audit-2026-05-31.md (new), changelog.d/added/0914-unified-python-test-orchestrator.md (new).
Rebase impact: None. The orchestrator is entirely fork-local — upstream Netflix/vmaf ships only the python/ legacy harness and its python/tox.ini, neither of which this change modifies. The new noxfile.py lives at repo root, a path upstream does not occupy. If upstream ever adds its own noxfile.py, treat the conflict as fork-takes-priority: our file delegates to upstream's python/tox.ini via the python_harness session, so behaviour is preserved.
clang-tidy modernize-* family enablement (2026-05-31)¶
Files touched: .clang-tidy, core/src/feature/feature_collector.cpp, core/src/metadata_handler.cpp.
Rebase impact: Low. .clang-tidy is fork-local; upstream Netflix does not ship one. feature_collector.cpp is fork-renamed from upstream .c under ADR-0725-family migrations — if an upstream sync brings a new .c patch that touches feature_collector, the patch likely applies cleanly to the .cpp (extern "C" linkage is preserved) but should be replayed in the C++ idiom (nullptr not NULL, <cstring> not <string.h>). metadata_handler.cpp is wholly fork- local with no upstream counterpart.
When syncing: keep the four -modernize-* opt-outs in .clang-tidy (noise / C-ABI hostility rationale documented in ADR-0915). If upstream ever ships their own clang-tidy config, merge by union — drop our opt-outs only with an explicit ADR.
cargo-deny supply-chain policy (2026-05-31)¶
Files touched: deny.toml (new), .github/workflows/rust-ci.yml (new cargo-deny job + deny.toml / core/src/feature/rust/** path filters), core/src/feature/rust/tad/Cargo.toml (publish = false).
Rebase impact: None against upstream Netflix/vmaf — deny.toml, the cargo-deny CI job, and the Rust workspace itself are all fork-local additions. Upstream does not maintain a Rust workspace, so no merge surface exists. The publish = false change to core/src/feature/rust/tad/Cargo.toml is also fork-local (core/src/feature/rust/ is an ADR-0707 pilot directory that does not exist upstream).
If a future upstream sync starts shipping a Rust workspace of its own, reconcile by extending deny.toml's [graph] members implicit-include behaviour (cargo-deny picks up workspace members automatically) and audit whether upstream's choice of licenses / banned-crate stance differs from ours. See ADR-0917.
Pixel-format edge coverage test (2026-05-31)¶
Files touched: core/test/test_pixel_format_edge_coverage.c (new), core/test/meson.build (one executable + one test() registration).
Rebase impact: Low. The new test file is wholly fork-local and only links against the public extractor / picture / collector C surface (no internal-source #include). If upstream Netflix renames any of the API entry points the test uses (vmaf_get_feature_extractor_by_name, vmaf_feature_extractor_context_create / _extract / _close / _destroy, vmaf_feature_collector_init / _get_score / _destroy, vmaf_picture_alloc / _unref), update the test accordingly. The meson.build additions sit between the existing test_psnr block and test_framesync; no upstream core/test/meson.build reordering should conflict, since the inserted block is immediately adjacent to fork-only neighbours.
ADR-0912.
ADR README drift sweep (2026-05-31)¶
Files touched: docs/adr/README.md, docs/adr/_index_fragments/_order.txt, 35 new + 7 rewritten files under docs/adr/_index_fragments/[0-9]*.md, 3 orphan fragments removed under docs/adr/_index_fragments/, changelog.d/fixed/adr-readme-regen.md.
Rebase impact: None. The fragment tree and README.md are entirely fork-local (upstream Netflix/vmaf has no ADR directory). The sweep only re-aligns three fork-local index sources against the already-authoritative docs/adr/[0-9]*-*.md ADR file set, with no content changes to any ADR body. Future regenerations are mechanical via scripts/docs/concat-adr-index.sh --write.
codespell sweep + .codespellrc (2026-05-31)¶
Files touched: .codespellrc (new), CONTRIBUTING.md, docs/metrics/cambi.md, docs/adr/0910-codespell-sweep-config.md (new), changelog.d/fixed/codespell-sweep.md (new).
Rebase impact: Low. .codespellrc skip-list explicitly excludes every Netflix-author / vendored / upstream-mirrored file enumerated in ADR-0910 §Context (e.g. compat/python-vmaf/*, python/test/*, core/src/feature/{x86,arm64,cuda,hip,common,metal}/*, core/src/svm.cpp, core/src/pdjson.c, core/tools/y4m_input.c, core/tools/cli_parse.c, core/README.md, core/tools/README.md, core/test/test_picture.c), so re-running codespell after a sync surfaces only newly-introduced fork typos. If upstream lands new files under the skipped trees that the fork later adopts as fork-local (e.g. a new feature extractor we then modify), drop the matching skip row and re-run codespell to catch any latent typos.
If upstream changes path layout (rename core/ back to libvmaf/, etc.), update the skip-list paths in .codespellrc to match. ignore-words-list is independent of upstream layout.
Re-run: codespell --config .codespellrc (or just codespell from the repo root — picks up .codespellrc automatically). Expected output: no findings on a clean tree.
.gitignore staleness audit (ADR-0905, 2026-05-30)¶
Files touched: .gitignore, python/.gitignore.
Rebase impact: None. Both files are fork-local (the rules trimmed or rewired all originate from fork additions and the post-ADR-0700 directory rename). Upstream Netflix/vmaf maintains its own .gitignore independently; the matlab MEX block, the Cython adm_dwt2_cy block, and the legacy python/.gitignore scope were fork-only artefacts of the rename and never tracked upstream. On the next /sync-upstream, Netflix's .gitignore will merge cleanly because the trimmed rules (.gradle/, .pypirc) and the rewired matlab paths (compat/python-vmaf/matlab/**/*.mex*) do not overlap any upstream rule.
cpp const/noexcept/nodiscard annotation sweep (2026-05-30)¶
Files touched: core/src/dict.cpp, core/src/feature/feature_collector.cpp, core/src/feature/feature_name.cpp, core/src/fex_ctx_vector.cpp, core/src/opt.cpp.
Rebase impact: None. All annotations are added to fork-local TU-internal static helpers and one TU-local lambda in C++23 files that were introduced by the ADR-0723 / ADR-0727 / ADR-0729 / ADR-0731 C++ migration waves. The extern "C" public-ABI entry points are untouched, so no upstream header rebase is affected. If upstream Netflix introduces new fork-only C++ static helpers, apply the same [[nodiscard]] / noexcept discipline so the lint posture stays uniform.
libvmaf-public-header-doc-gaps-round3 (2026-05-30)¶
Files touched:
core/include/libvmaf/picture.h(doc comments on enum + opaque typedef + 2 entry points, plus NOLINT-cited include guard)core/include/libvmaf/libvmaf.h(doc comments on 2 enums + opaque typedef + 1 struct, plus NOLINT-cited include guard)core/include/libvmaf/libvmaf_cuda.h(doc comments on opaque typedef + config struct + enum + 1 picture-config struct, plus NOLINT-cited include guard)
Rebase impact: Low. The doc-comment additions land above unchanged upstream-mirror declarations; any future Netflix upstream that touches the same function signatures, enum bodies, or struct definitions will produce a tractable 3-way merge — the doc text is fork-local and git merge will preserve our /** ... */ block above whatever upstream rewrites the declaration to. No identifier renames; no ABI/source impact.
The NOLINT annotations on __VMAF_H__ / __VMAF_PICTURE_H__ / __VMAF_CUDA_H__ are inline comments only — they do not alter the include guard symbols themselves, so upstream's preprocessor identity remains intact. Same pattern PR #327 (round 2) used for feature.h / model.h / dnn.h. If a future upstream sync changes the guard form (unlikely — these have been stable for years), the NOLINT cites become redundant and can be removed in a follow-on cleanup.
libvmaf-public-header-doc-gaps-round2 (2026-05-30)¶
Files touched: - core/include/libvmaf/feature.h (doc comments + NOLINT-cited guard) - core/include/libvmaf/model.h (doc comments + NOLINT-cited guard) - core/include/libvmaf/dnn.h (vmaf_dnn_session_close doc + NOLINT-cited guard)
Rebase impact: Low. The doc-comment additions land above unchanged upstream-mirror declarations; any future Netflix upstream that touches the same function signatures will produce a tractable 3-way merge — the doc text is fork-local and git merge will preserve our /** ... */ block above whatever upstream rewrites the signature to. No identifier renames; no ABI/source impact.
The NOLINT annotations on __VMAF_FEATURE_H__ / __VMAF_MODEL_H__ / __VMAF_DNN_H__ are inline comments only — they do not alter the include guard symbols themselves, so upstream's preprocessor identity remains intact. If a future upstream sync changes the guard form (unlikely — these have been stable for years), the NOLINT cites become redundant and can be removed in a follow-on cleanup. /binary symbol renames; consumers of the patch stack (ffmpeg-patches/) and the Go/Rust bindings see identical declarations.
Bash strict-mode + trap-cleanup sweep (2026-05-30, ADR-0899)¶
Files touched: scripts/run_unittests.sh, scripts/ai/fetch-tiny-blobs.sh, dev/scripts/smoke-probe-loop.sh, scripts/ci/check-agent-worktree-drift.sh, scripts/ci/test_check_agent_worktree_drift.sh, scripts/ci/check-adr-numbering.sh, scripts/ci/check-dispatch-registry.sh, scripts/adr/next-free.sh, tools/ensemble-training-kit/_platform_detect.sh.
Rebase impact: None. All 9 files are fork-local (Netflix upstream has neither scripts/adr/, scripts/ci/check-*-drift*, scripts/ai/fetch-tiny-blobs.sh, dev/scripts/smoke-probe-loop.sh, tools/ensemble-training-kit/, nor the in-tree scripts/run_unittests.sh in this form). No conflict risk on sync-upstream.
Conflict watchpoints (none expected): if a future upstream sync introduces a Netflix-side scripts/run_unittests.sh, the strict-mode set -eu block at the top of our version is the only carrier of fork-specific behaviour and trivially survives a 3-way merge.
Metal kernel parity tests round 2 (2026-05-30)¶
Files touched: core/test/meson.build, core/test/test_metal_motion_v2_parity.c (new), core/test/test_metal_integer_psnr_parity.c (new), core/test/test_metal_float_psnr_parity.c (new), core/test/test_metal_float_ssim_parity.c (new)
Rebase impact: None. All four new files live under the existing fork-local enable_metal block in core/test/meson.build (the entire Metal backend is fork-added — ADR-0361 / ADR-0421 / ADR-0589 — and absent from upstream Netflix/vmaf). The block edit appends four new executable() + test() pairs immediately after the test_metal_install_header block; the surrounding endif boundaries are untouched so upstream syncs cannot conflict here.
If upstream ever ports a Metal backend, the test files would need re-pointing at the upstream kernel names; the synthetic-fixture + -ENODEV skip pattern from test_sycl_motion3_parity.c carries forward unchanged.
.claude/skills/ — ADR-0700 path drift cleanup (2026-05-30)¶
Files touched: .claude/skills/add-gpu-backend/scaffold.sh, .claude/skills/build-vmaf/build.sh, .claude/skills/build-vmaf/SKILL.md, .claude/skills/regen-docs/SKILL.md, .claude/skills/add-simd-path/templates/simd_feature.c.template
Rebase impact: None. Files are entirely fork-local (the .claude/ tree does not exist upstream — see ADR-0331 / ADR-0700). The change rewrites four residual libvmaf/ source-tree references to core/ to match the post-ADR-0700 layout. Public install-path references (core/include/libvmaf/..., libvmaf.so) are unchanged.
When syncing from upstream Netflix/vmaf, this file does not need attention; the conflict surface is empty.
ADR-0871 — SSIM SIMD dispatch pthread_once guard — 2026-05-30¶
Low rebase impact. The fix sits in two fork-added zones:
core/src/feature/iqa/ssim_tools.c— the file is a Tom-Distler BSD-2011 import, but the four globals (g_ssim_precompute,g_ssim_variance,g_ssim_accumulate,g_iqa_convolve), the setter functions, and the newiqa_ssim_install_dispatch_oncehelper are fork additions (Distler's 2011 import has no SIMD dispatch). The pthread_once guard and atomic-installer publish are appended to the existing fork-added block. A future re-import of Tom Distler's IQA would not collide because the new code lives in fork-added territory.core/src/feature/iqa/ssim_simd.h— fork-added header (Netflix/vmaf has no equivalent); appends one declaration.core/src/feature/float_ssim.candcore/src/feature/float_ms_ssim.c— the dispatch-install bodies are fork additions; the change factors them into a callback and routes the call through the once-helper. The Netflix-upstream init() bodies are unchanged beyond the dispatch block, so a future upstream change to the init() prologue would merge cleanly.
Fork-local files: core/src/feature/iqa/ssim_tools.c (fork-added dispatch zone), core/src/feature/iqa/ssim_simd.h (fork-added header), core/src/feature/float_ssim.c (fork-added SIMD-install block), core/src/feature/float_ms_ssim.c (fork-added SIMD-install block), docs/adr/0871-ssim-dispatch-pthread-once.md, docs/research/tsan-race-audit-2026-05-30.md, changelog.d/fixed/tsan-race-audit.md.
sanitizer-pass-cleanup (2026-05-30, ADR-0869)¶
Files touched:
core/src/feature/cambi.c— adds twointshadow slots (window_size_opt,max_log_contrast_opt) toCambiState; the options table targets them;init()copies into the existinguint16_truntime fields.core/src/feature/x86/adm_avx2.c— moves theuint32_tcast inside the shift in four DWT2 filter-packing expressions.core/src/feature/x86/adm_avx512.c— same as AVX2.
Rebase impact:
- CAMBI: upstream Netflix's
CambiStatedoes not have the_optshadow slots. On upstream sync, expect a context conflict on the struct definition and on the two option-table entries. Resolution is to keep the fork's shadow slots and the init-bridge assignments; upstream's option entries should be re-pointed at the_optshadows. - ADM AVX2/AVX-512: the four filter-packing expressions are upstream-mirrored code. On upstream sync, a textual conflict is possible at every occurrence; the fork's resolution is the inside-cast (
((uint32_t)filter[k] << 16)). Bit-exact with upstream output; safe to keep.
Verified clean under ASan+UBSan against the full unit-test suite (63 tests OK) and the vmaf CLI on 4:2:0 8-bit, 4:2:2 10-bit, 4:2:0 12-bit. Cambi tuned-options feature-name derivation (cambi_mlc_3_ws_63) works.
SIMD bit-exactness round-2 — SSIMULACRA 2 FMA unification + lib-FP-model extension (2026-05-30, ADR-0891)¶
Files touched: core/src/meson.build, core/src/feature/x86/ssimulacra2_avx2.c, core/src/feature/x86/ssimulacra2_avx512.c, core/test/test_ssimulacra2_simd.c.
Rebase impact: Low — SSIMULACRA 2 is fork-added (no upstream coupling) and the meson helper _libvmaf_feature_icx_args mirrors the existing _x86_simd_strict_fp_extra pattern from ADR-0339 (round-1). If upstream Netflix ever adds an intel-llvm build matrix and ships scalar references inside libvmaf_feature_static_lib that participate in SIMD bit-exactness tests, reuse _libvmaf_feature_icx_args rather than minting a new helper. The FMA-based picture_to_linear_rgb colour matrix is fully self-contained inside the SSIMULACRA 2 TUs; no upstream Netflix file references those symbols. Companion: docs/adr/0891-simd-bit-exact-round2-fmaf-libvmaf-feature-icx.md, changelog.d/fixed/0891-simd-bit-exact-round2.md.
SIMD strict-FP flags for icx (2026-05-30)¶
Files touched: core/src/meson.build, core/test/meson.build, core/src/feature/AGENTS.md
Rebase impact: Low. The changes add an icx-specific compile flag (-fp-model=precise) to x86 SIMD carve-out static libs and to the three SIMD bit-exactness test executables (test_psnr_hvs_simd, test_ms_ssim_decimate, test_ssimulacra2_simd). The flag is added only when cc.get_id() returns 'intel-llvm' or 'intel-llvm-cl', so GCC and vanilla Clang builds are unaffected.
If upstream Netflix adds new SIMD carve-out static libs, apply the same _x86_simd_strict_fp_extra pattern to them so the icx build stays green. If Netflix adds new SIMD test executables that compare a scalar reference against SIMD output, add _simd_strict_fp_args to their c_args.
Coverage Gate ORT accessor coverage (2026-05-30)¶
Files touched: core/test/dnn/test_ort_internals.c, changelog.d/fixed/coverage-gate-ort-backend-accessor.md.
Rebase impact: None. The added test exercises a fork-only public accessor (vmaf_ort_output_name_at) on a fork-only file (core/src/dnn/ort_backend.c); the test TU itself is fork-only under ADR-0112's testability surface. Upstream Netflix/vmaf has no ORT backend, so there is no cross-repo file to reconcile on sync. The ADR-0114 per-file floor override (PER_FILE_MIN["core/src/dnn/ort_backend.c"]=78) stays in place; the coverage delta (409 → 413 / 526 = 78.5 %) is the per-file safety margin restored after PR #129 grew the denominator with unreachable error-handling.
unused-testdata-debug-scripts-cleanup (2026-05-30, ADR-0880)¶
Files touched: testdata/check_borders.py (deleted), testdata/compare_a380.py (deleted), testdata/scores_sycl_b580_576_mq.json (deleted).
Rebase impact: None. All three files were fork-added and not present in upstream Netflix/vmaf. No upstream patch context references them. Future /sync-upstream runs will not surface any conflicts on these paths.
trivy-container-scan-baseline (2026-05-30, ADR-0878)¶
Files touched: docker/Dockerfile.production, docker/Dockerfile.production-gpu
Rebase impact: None. Both files are fork-added (no upstream Netflix/vmaf equivalents — Netflix ships no production Dockerfile). The added USER nonroot:nonroot directive on each final stage will not conflict on any future upstream sync. If upstream ever publishes their own Dockerfile, the fork's containers stay separate (the GHCR namespace is vmafx/).
go-nilness-staticcheck-audit (2026-05-30)¶
Files touched: cmd/vmafx-server/{main.go,http_server.go}, cmd/vmafx-controller/{main.go,http_server.go}, cmd/vmafx-mcp/impl.go, cmd/vmafx-node/main_test.go, pkg/ai/infer_test.go, pkg/bisect/bisect_test.go.
Rebase impact: None. Every modified file is fork-original Go code under cmd/vmafx-* / pkg/*; Netflix/vmaf upstream does not ship Go code in these paths. No upstream conflict possible.
iwyu-audit (2026-05-30) — fork-only files, append-only direct includes¶
Files touched: 16 fork-authored sources under core/src/feature/, core/src/feature/x86/, core/test/, core/tools/.
Rebase impact: None. All modified files carry the Lusoris-only license header (filtered explicitly during scope selection — files with a Netflix header were skipped to preserve upstream-parity per CLAUDE.md §12 r12). The diff consists of removing dead #include directives and adding direct includes for symbols previously reached transitively. Upstream Netflix/vmaf does not contain any of these files in the form modified here, so there is no conflict surface for a future sync-upstream to navigate.
Follow-up: A second-phase IWYU pass on core/src/dnn/*, core/src/{cuda,sycl,hip,vulkan}/, and the DNN-gated feature extractors is owed (the host CPU-only build cannot exercise VMAF_HAVE_DNN because ONNX Runtime is not installed locally). That pass will run inside the vmaf-dev-mcp container per CLAUDE.md §12 r15.
magic-number-audit cert-int07c (2026-05-30, ADR-0874)¶
Files touched: core/src/mcp/{mcp_internal.h,mcp.c,compute_vmaf.c,transport_sse.c}, core/src/picture.c, core/src/cuda/picture_cuda.c, core/src/libvmaf.c.
Rebase impact: Low. All five core/src/mcp/* files and core/src/cuda/picture_cuda.c are fork-added; upstream Netflix/vmaf has neither MCP nor a CUDA picture-allocator with these bounds. core/src/picture.c and core/src/libvmaf.c are fork-mirrored — the renames touch fork-added helpers (dnn_*_output_feature_name) and the fork's VMAF_PIC_BPC_{MIN,MAX} hardening (originally a fork-local guard against bpc < 8 || bpc > 16). A future upstream sync that re-introduces a raw 8/16 predicate on those lines should keep the fork's named constants — they are not bit-exact changes and do not alter behaviour. No new public C-API symbols introduced.
eintr-and-io-error-audit (2026-05-30, ADR-0872)¶
Files touched: core/src/mcp/transport_stdio.c, core/src/mcp/transport_uds.c, core/src/libvmaf.c, core/src/feature/cambi.c, core/src/sycl/dmabuf_import.cpp, core/tools/vmaf_vpl.c.
Rebase impact: Low. The MCP transports are fully fork-local (no upstream peer). libvmaf.c, cambi.c, and vmaf_vpl.c carry fork-local hunks (vmaf_write_output, heatmaps close() fail-path, VPL VA-API init) that are already non-shared with upstream — the new (void) casts sit inside those hunks. dmabuf_import.cpp is wholly fork-added (no upstream file). No upstream conflict expected on the next sync; if Netflix ever adds their own MCP transport, the EINTR retry pattern should be ported there too.
adr-0100-per-surface-doc-audit (2026-05-30)¶
Files touched: docs/development/build-flags.md, docs/api/dnn.md, docs/usage/cli.md, changelog.d/added/adr-0100-per-surface-doc-audit.md.
Rebase impact: None. All four files are fork-added (the upstream Netflix/vmaf tree has no docs/development/build-flags.md, no docs/api/dnn.md, no docs/usage/cli.md at the fork's depth, and no changelog.d/). The audit closes per-surface doc gaps for fork-local surfaces (codec-context DNN API, codec/preset/CRF/resize CLI flags, six Meson options) that originated in fork ADRs (ADR-0335, ADR-0361, ADR-0519, ADR-0550, ADR-0568, ADR-0623, ADR-0707, ADR-0726). No upstream file is touched; no rebase conflict possible.
go-pkg-coverage-push (2026-05-30)¶
Files touched: pkg/observability/observability_test.go, pkg/report/report_test.go, pkg/encoder/discover_test.go, pkg/libvmaf/paths_test.go, pkg/gpu/parsers_test.go, pkg/gpu/probe_shim_test.go, pkg/bisect/parse_test.go, pkg/storage/internals_test.go, changelog.d/added/go-pkg-coverage-push.md.
Rebase impact: None. The Go pkg/ tree is wholly fork-added — upstream Netflix/vmaf has no Go layer. All new files are test-only and never enter the libvmaf C build, the Python harness, or the FFmpeg patch stack. No production code is touched, so the upstream rebase boundary is unaffected.
python-type-annotations-audit (2026-05-30)¶
Files touched: ai/src/aiutils/{__init__,jsonl_utils,parquet_utils}.py, ai/src/corpus/base.py, mcp-server/vmaf-mcp/src/vmaf_mcp/{server,http_transport}.py, tools/vmaf-tune/src/vmaftune/{auto,benchmark,corpus,encoder_profile, fr_from_nr_adapter,hdr,predictor_features,report,saliency,score, score_backend,sidecar}.py, tools/vmaf-tune/src/vmaftune/codec_adapters/_gop_common.py, pyproject.toml.
Rebase impact: None. Every touched file is fork-added (ai/, mcp-server/, tools/vmaf-tune/) or fork-only mypy config (pyproject.toml [tool.mypy.overrides]). Upstream Netflix/vmaf does not ship any of these trees; on a future upstream sync there is no conflict surface.
The change is a pure type-annotation tightening — no runtime semantics change. The one functional change is the removal of a dead-code duplicate _run_benchmark() definition in mcp-server/vmaf-mcp/src/vmaf_mcp/server.py; the deleted copy was silently shadowed at import time by the progress-token-aware implementation 575 lines later, so removal is behaviour-preserving.
openapi-rest-schema (2026-05-29, ADR-0797)¶
Files touched: api/openapi/vmafx-server-v1.yaml, api/openapi/oapi-codegen.yaml, gen/go/oapi/vmafx_server_v1.gen.go, cmd/vmafx-server/rest_adapter.go, cmd/vmafx-server/swagger_ui.go, cmd/vmafx-server/http_server.go, cmd/vmafx-server/grpc_server.go, cmd/vmafx-server/main.go, docs/server/rest.md
Rebase impact: None. All touched files are fork-local additions in the Go server layer (cmd/vmafx-server/, api/, gen/go/) that do not exist in upstream Netflix/vmaf. No rebase conflicts are possible.
The newHTTPServer signature gained a *grpcServer parameter; any fork-local branch that calls newHTTPServer with the old 4-argument form will fail to compile and must add the grpcServer argument.
ADR-0783 — Kubernetes e2e integration test harness (2026-05-29)¶
No rebase impact on upstream C/Python code.
All files are wholly fork-local additions: test/e2e/kind-cluster.sh, test/e2e/fixtures/gen-tiny-yuv.sh, test/e2e/fixtures/ref.yuv, test/e2e/fixtures/dist.yuv, test/e2e/kuttl-tests/ (all test case YAML), .github/workflows/e2e-k8s.yml, docs/k8s/integration-tests.md, docs/adr/0783-k8s-e2e-integration-test-harness.md, changelog.d/added/k8s-e2e-integration-test-harness.md.
Netflix upstream has no Kubernetes test infrastructure; no merge conflict risk. A sync-upstream that adds an upstream e2e directory would not conflict with this harness because Netflix uses libvmaf/ path roots that the fork has renamed to core/ (ADR-0700).
cuda-ms-ssim-vert-lcs-horiz-ldg (2026-05-29, ADR-0757)¶
Files touched: core/src/feature/cuda/integer_ms_ssim/ms_ssim_score.cu
Rebase impact: None. The modified file is a fork-added CUDA kernel TU that does not exist in upstream Netflix/vmaf master (ms_ssim CUDA port is fork-local). No rebase conflict is possible.
The change is a pure performance annotation: __launch_bounds__(128), const float *__restrict__ pointer extraction, and __ldg() on inner-loop loads. If upstream Netflix ever adds their own ms_ssim CUDA port, this file will need to be re-reviewed against theirs; the F3 pattern should carry forward.
cpp23 orphan .c sweep — metadata_handler.c (2026-05-29)¶
Files touched: core/src/metadata_handler.c (deleted)
Rebase impact: None. The file was dead source — never referenced by any meson.build after ADR-0708 renamed it to metadata_handler.cpp. Upstream Netflix/vmaf still uses metadata_handler.c; on future upstream sync, the upstream .c file will reappear in the patch context but meson.build will continue to reference only .cpp. No conflict possible: the deletion only affects the fork-local tree.
Rule for future cpp23 conversions: when renaming foo.c → foo.cpp in meson.build, always git rm core/src/foo.c in the same commit. Leaving both files in tree causes the source tree to diverge from the build definition.
cuda-readback-free-host-pinned-leak sweep (2026-05-29)¶
Files touched: core/src/cuda/kernel_template.h, docs/backends/kernel-scaffolding.md
Rebase impact: None. The fix is entirely in fork-added files (kernel_template.h is a Lusoris-added header; kernel-scaffolding.md is fork-added documentation). No upstream Netflix/vmaf file is modified.
The changed function (vmaf_cuda_kernel_readback_free) did not exist in upstream — it was introduced by the fork's kernel-template ADR. No rebase conflict is possible.
ADR-0753 — CUDA resolution-aware dispatch scaffold (2026-05-29)¶
Files touched (initial + extended scope):
core/src/feature/cuda/resolution_dispatch.{h,c}(new)core/src/feature/cuda/integer_adm/adm_cm.cu(two kernel macros)core/src/feature/cuda/integer_adm_cuda.c(include, struct field, init, dispatch)core/src/feature/cuda/integer_vif/filter1d.cu(FILTER1D_8_HORI_NO_BOUNDS macro + instantiation)core/src/feature/cuda/integer_vif_cuda.c(struct field, init, resolution-aware dispatch in filter1d_8)core/src/feature/cuda/integer_ssim/ssim_score.cu(calculate_ssim_vert_combine_no_bounds)core/src/feature/cuda/integer_ssim_cuda.c(struct field, init, resolution-aware dispatch in submit_fex_cuda)core/src/feature/cuda/AGENTS.md(invariant notes + verified wirings table)docs/adr/0753-cuda-resolution-aware-dispatch.md(new; extended policy table)docs/backends/cuda/overview.md(kernel dispatch table extended)docs/research/0753-cuda-resolution-aware-dispatch-design.md(new)changelog.d/added/cuda-resolution-aware-dispatch.md(new)
Rebase impact: Low on resolution_dispatch.{h,c} — these are wholly new fork-local files; no upstream conflict possible.
adm_cm.cu: The ADM_CM_LINE macro was split into ADM_CM_LINE_BOUNDED and ADM_CM_LINE_NO_BOUNDS. If upstream Netflix modifies adm_cm.cu after the fork diverges, the split needs to be reapplied around the new macro body. The extern "C" wrapping (ADR-0747) must be preserved for both entries.
integer_adm_cuda.c: The AdmStateCuda struct grew one field (func_adm_cm_line_kernel_8_no_bounds). If upstream adds fields to the struct in the same location, resolve the merge conflict by keeping both additions. The new #include "feature/cuda/resolution_dispatch.h" line must survive any upstream shuffle of the include block.
On rebase: verify that both cuModuleGetFunction calls in the init block still reference valid kernel symbol names from adm_cm.cu.
Research-0751 4K baseline + PR #79 adm_cm A/B (2026-05-29)¶
Files touched: docs/research/0751-cross-backend-4k-baseline-and-pr79-adm-cm-4k-measure.md, changelog.d/changed/cross-backend-4k-baseline.md
Rebase impact: None. Research-only digest; no source code changed. No upstream conflict possible — these are fork-added measurement artifacts.
CI round-3 fix — .semgrepignore, .gitleaks.toml, codeql-config.yml, compat/python-vmaf/ (2026-05-28)¶
Files touched: .semgrepignore, .gitleaks.toml, .github/codeql-config.yml, compat/python-vmaf/core/feature_extractor.py, core/test/test_hip_smoke.c, ai/src/aiutils/jsonl_utils.py, ai/src/vmaf_train/registry.py, .github/workflows/libvmaf-build-matrix.yml.
Rebase impact: Low. All changes are either CI config fixes (path corrections post-ADR-0700 rename) or code fixes for missing functions and removed extractors.
On upstream sync:
.semgrepignoreand.gitleaks.tomlare fork-local; no upstream conflict expected.codeql-config.ymlis fork-local; no upstream conflict expected.compat/python-vmaf/core/feature_extractor.py: if Netflix upstream modifiespython/vmaf/core/feature_extractor.py(old path), the rename-shim must preserve the removal offloat_ansnrfromVmafIntegerFeatureExtractor's features list. The legacy path (VmafFeatureExtractor, line 301) may still referencefloat_ansnrif upstream restores it; that's intentional pending the legacy-runner sunset decision.core/test/test_hip_smoke.c: if upstream addsfloat_ansnr_hipback, the removed test function must be restored.
docs/research/0734-r610-driver-changelog-audit-2026-05-28.md — R610 driver audit¶
No rebase impact. This is a documentation-only research digest; it does not touch any C sources, build files, or API surfaces. No upstream sync conflict expected.
docs/research/0734-cudnn-version-audit-20260528.md — cuDNN/ORT audit (doc-only)¶
No rebase impact on upstream C/Python code: this PR adds only doc and changelog files. No C source, header, or Python source is modified.
If a future upstream sync adds cuDNN pinning or onnxruntime-gpu to any Python requirement, re-check dev/Containerfile lines 529–539 (ORT install) and ai/pyproject.toml for compatibility with the then-current cuDNN series.
Fork-local files added: docs/research/0734-cudnn-version-audit-20260528.md (new), changelog.d/changed/docs-cudnn-version-audit.md (new), docs/rebase-notes.md (this entry), docs/state.md (new deferred row).
Periodic drift sweep — upstream syncs may reintroduce libvmaf/ refs¶
After every Netflix/vmaf upstream sync, run the inventory grep from PR chore/post-rename-drift-sweep-20260528 to catch any new libvmaf/[a-z] or python/vmaf/ directory references outside ADR bodies and CHANGELOG.md. Files to recheck: Makefile, Dockerfile, .github/codeql-config.yml, IDE settings, skill scripts, and any newly-added utility under scripts/. See changelog fragment changelog.d/fixed/post-rename-drift-sweep.md for the full inventory commands.## port/upstream-batch-threading-picture-pool (2026-06-04)
Files touched: core/src/libvmaf.c, core/src/meson.build
Rebase impact: if a future upstream commit adds more #ifdef VMAF_BATCH_THREADING blocks, those blocks must be removed in the same port PR — the fork no longer uses the flag. The non-batch threaded_read_pictures path was removed; it is not recoverable from the fork without re-introducing the old per-extractor thread pool enqueue pattern.
.github/workflows/tests-and-quality-gates.yml — coverage job deselects slow vifks360 test¶
The coverage job's --deselect list includes python/test/quality_runner_test.py::QualityRunnerTest::test_run_vmaf_runner_float_vifks360o97 because the test exceeds the 60 s per-test limit on GitHub-hosted runners and truncates the suite. If upstream Netflix/vmaf adds a test with a similar name in a future sync, verify it does not also use a very large vif_kernelscale before removing the deselect. The deselect is CI-only; the test runs in the Netflix golden gate without a per-test timeout.
.github/workflows/ — post-ADR-0700 path rename (libvmaf/ → core/)¶
If an upstream Netflix/vmaf sync or cherry-pick brings new CI references to libvmaf/ (path filters, cd libvmaf, find libvmaf/src), they must be remapped to core/ in the same PR. The fork's source tree is rooted at core/ per ADR-0700; any upstream workflow or Makefile that still hardcodes libvmaf/ as a source directory will silently build from a non-existent path on this fork. Additionally, replace any gitleaks/gitleaks-action usage with the direct gitleaks CLI binary — the action requires a GITLEAKS_LICENSE for org repos even when public.
docker/Dockerfile.node — vmafx-node worker image + ffmpeg n8.2 (ADR-0717)¶
ffmpeg-patches now validated against both n8.1.1 and n8.2. The node Dockerfile pins FFMPEG_TAG=n8.2. When the next upstream sync lands, confirm:
ffmpeg-patches/still applies against the new tag. RunFFMPEG_SHA=<new-tag> bash ffmpeg-patches/test/build-and-run.sh.- If patches fail, rebase the affected patches and update
FFMPEG_TAGin bothdev/Containerfileanddocker/Dockerfile.nodein the same PR (CLAUDE.md §12 r14). pkg/encoder/encoder.goshells out to ffmpeg. If a new FFmpeg version changes a codec's CLI flag, update the encoder package to match.
Touched files: docker/Dockerfile.node, cmd/vmafx-node/main.go, cmd/vmafx-node/probe/probe.go, cmd/vmafx-node/probe/probe_test.go, cmd/vmafx-node/server/server.go, cmd/vmafx-node/server/server_test.go, docs/adr/0717-vmafx-node-ffmpeg-latest.md, docs/development/vmafx-node.md, changelog.d/added/node-ffmpeg-latest.md, docs/state.md (this entry), docs/rebase-notes.md (this entry).
feat/speed-python-compat-extractors (Research-0732, item #2) — low-conflict upstream port¶
No structural rebase impact. This PR adds fork-local content to paths (compat/python-vmaf/core/feature_extractor.py, compat/python-vmaf/core/quality_runner.py, python/test/feature_extractor_test.py, docs/metrics/speed_qa.md) that are already diverged from upstream (python/vmaf/core/… in Netflix/vmaf). When syncing from upstream:
- If Netflix/vmaf updates
SpeedChromaFeatureExtractororSpeedTemporalFeatureExtractor(e.g. bumps VERSION), apply the equivalent change tocompat/python-vmaf/core/feature_extractor.py. - If Netflix/vmaf adds new SpEED QualityRunner subclasses, port them to
compat/python-vmaf/core/quality_runner.py. - The compat harness mirrors Netflix's class hierarchy intentionally; keep the TYPE, VERSION, and ATOM_FEATURES_TO_VMAFEXEC_KEY_DICT in sync.
cmd/vmafx-server — Go gRPC + HTTP server (ADR-0703)¶
no rebase impact on upstream C/Python code: the Go server is entirely fork-local (cmd/, pkg/, gen/, proto/, go.mod, go.sum, Dockerfile.go-server, buf.gen.yaml). None of these paths overlap with Netflix/vmaf upstream.
If a future upstream sync touches model/ (model JSON schema changes) or core/include/libvmaf/libvmaf.h (public ABI), review:
pkg/libvmaf/libvmaf.go— the cgo#includeand JSON parsing inparseOutput.- The
ScoreResponse.featuresmap keys (derived frompooled_metricskeys in the vmaf CLI JSON output; key names are stable but new keys may appear).
Touched files: cmd/vmafx-server/main.go, cmd/vmafx-server/grpc_server.go, cmd/vmafx-server/http_server.go, cmd/vmafx-server/main_test.go, pkg/libvmaf/libvmaf.go, pkg/libvmaf/libvmaf_test.go, pkg/observability/observability.go, proto/vmafx.proto, proto/buf.yaml, buf.gen.yaml, gen/go/vmafx.pb.go, gen/go/vmafx_grpc.pb.go, go.mod, go.sum, Dockerfile.go-server, docs/server/grpc.md, docs/adr/0703-vmafx-server-go-grpc.md, changelog.d/added/vmafx-server-go.md, docs/state.md, deploy/helm/vmafx/values.yaml (image repository update).
PR that touches upstream-shared paths or establishes a rebase-sensitive invariant adds an entry here. PRs with no rebase impact state "no rebase impact" in the PR description and skip the entry.
docs/hw-backend-audit-2026-05-28 — doc-only, no rebase impact¶
No upstream rebase impact: this PR adds a research digest (docs/research/0733-hardware-backend-audit-2026-05-28.md), a changelog fragment, and a docs/state.md update. No C source, build system, or upstream-shared path is touched. Netflix/vmaf upstream syncs are unaffected.
feat/vmafx-phase4-language-modernization-foundation (ADR-0702) — fork-only, no Netflix conflict¶
No upstream rebase impact. The files added in this PR (go.mod, Cargo.toml, pkg/, cmd/, bindings/, .github/workflows/go-ci.yml, .github/workflows/rust-ci.yml) are entirely fork-local. Netflix/vmaf upstream does not have a Go or Rust surface; cherry-picks from upstream are unaffected.
The docs/principles.md, docs/development/languages.md, .gitignore, and Makefile additions are additive; the Makefile targets are named distinctly (go-build, go-test, rust-build, rust-test) and do not conflict with any upstream Makefile target.
feat/vmafx-tune-go-stage1 (ADR-0705) — fork-only, no Netflix conflict¶
No upstream rebase impact: the Go port lives entirely under cmd/vmafx-tune/, pkg/encoder/, pkg/bisect/, and pkg/report/. These directories do not exist in upstream Netflix/vmaf. The Python tools/vmaf-tune/ is unchanged. go.mod and go.sum are fork-local additions that upstream does not carry. Cherry-picks from upstream that touch tools/vmaf-tune/ Python source files are unaffected by this PR.
feat/vmafx-mcp-go-port (ADR-0704) — fork-only, no Netflix conflict¶
No upstream rebase impact: this PR adds cmd/vmafx-mcp/, pkg/libvmaf/, go.mod, and go.sum — all entirely fork-local. The Python MCP server at mcp-server/vmaf-mcp/ is unchanged. Netflix/vmaf upstream does not contain any Go code or an MCP server. Cherry-picks from upstream are unaffected.
chore/post-cutover-url-sweep — fork-only URL change, no Netflix conflict¶
No upstream rebase impact: this change replaces all occurrences of the lusoris/vmaf GitHub repository slug with VMAFx/vmafx following the GitHub org cutover. All affected strings are fork-local (CI workflow URLs, GHCR image paths, ADR cross-references, doc URLs). Netflix/vmaf upstream does not contain any of these references. Cherry-picks from upstream are unaffected.
refactor/vmafx-repo-layout (ADR-0700) — IMPORTANT: breaks all in-flight PRs¶
Upstream sync strategy: upstream Netflix/vmaf patches arrive with libvmaf/ paths. When cherry-picking or porting upstream commits after ADR-0700 merged, rewrite paths in the patch stream:
# Single commit
git format-patch -1 <upstream-sha> --stdout \
| sed 's|libvmaf/|core/|g' \
| git am --3way
# Range of commits
git format-patch <base>..<tip> --stdout \
| sed 's|libvmaf/|core/|g' \
| git am --3way
In-flight PR rebase recipe: after git rebase origin/master, resolve each libvmaf/ path conflict by renaming to core/, and each python/vmaf/ conflict by renaming to compat/python-vmaf/.
Python import compatibility: import vmaf continues to work via the compat/vmaf symlink (→ python-vmaf/) when compat/ is on sys.path, and via the python/vmaf/__init__.py shim when python/ is on sys.path. No from vmaf. import lines need changing.
What stays the same: libvmaf.so, libvmaf.pc, <libvmaf/...> C install-path headers, all public C symbols (VmafContext, vmaf_init, etc.), ffmpeg filter names.
Touched files: all source-tree path references across CI workflows, Makefile, scripts, docs, agent configs, and the libvmaf/ and python/vmaf/ directories themselves.
feat/ai-run-manifest-helper (ADR-0678)¶
No upstream rebase impact: this touches fork-local AI helper code, AI scripts, tests, Claude skills, docs, ADR/research notes, and a changelog fragment. Upstream Netflix/vmaf does not ship these AI provenance helpers or local training utilities.
Invariant: new standalone AI artifact sidecars use aiutils.run_manifest.write_run_manifest() so the shared envelope and run_provenance block stay deduplicated. Existing stable report schemas may continue embedding build_run_provenance() directly.
Smoke: .venv/bin/python -m pytest ai/tests/test_run_manifest.py ai/tests/test_build_bisect_cache.py ai/tests/test_legacy_extractor_manifests.py ai/tests/test_ptq_scripts.py ai/tests/test_qat_smoke.py -q
Touched files: ai/src/aiutils/run_manifest.py, ai/scripts/ptq_dynamic.py, ai/scripts/ptq_static.py, ai/scripts/qat_train.py, ai/scripts/build_bisect_cache.py, ai/scripts/collect_gpu_calibration_data.py, ai/scripts/extract_ugc_features.py, ai/scripts/extract_konvid_frames.py, AI tests, .claude/skills/ai-run-manifest/SKILL.md, AI package/Claude guidance, docs/ai/*.md, docs/adr/0678-*.md, docs/research/0699-*.md, changelog.d/added/0678-*.md, and this file.
feat/ai-dataset-fetch-manifests (ADR-0677)¶
No upstream rebase impact: this touches fork-local AI dataset fetch helpers, tests, docs, package AGENTS notes, ADR/research notes, and a changelog fragment. Upstream Netflix/vmaf does not ship these local downloader scripts.
Invariant: dataset fetch helpers that seed later AI JSONL/parquet builders write deterministic ADR-0661 run-manifest sidecars before conversion. fetch_konvid_1k.py defaults to <root>/fetch_manifest.json; fetch_youtube_ugc_subset.py keeps --manifest as the content manifest and defaults the run sidecar to <manifest>.run-manifest.json.
Smoke: .venv/bin/python -m pytest ai/tests/test_dataset_fetch_manifests.py -q
Touched files: ai/scripts/fetch_konvid_1k.py, ai/scripts/fetch_youtube_ugc_subset.py, ai/tests/test_dataset_fetch_manifests.py, ai/AGENTS.md, docs/ai/training.md, docs/ai/training-data.md, docs/ai/konvid-1k-ingestion.md, docs/ai/youtube-ugc-ingestion.md, docs/ai/mos-corpora.md, docs/adr/0677-*.md, docs/research/0698-*.md, changelog.d/added/0677-*.md, and this file.
feat/mos-corpus-adapter-manifests (ADR-0676)¶
No upstream rebase impact: this touches fork-local AI MOS corpus adapters, tests, docs, package AGENTS notes, ADR/research notes, and a changelog fragment. Upstream Netflix/vmaf does not ship these local MOS-corpus ingestion scripts.
Invariant: CHUG, KoNViD-1k, KoNViD-150k, YouTube-UGC, LSVQ, LIVE-VQC, and Waterloo-IVC source adapters write <output>.manifest.json by default using corpus.base.write_ingest_manifest() and ADR-0661 run_provenance. Keep new MOS adapter CLIs on this sidecar contract before their JSONL rows feed aggregation, model-card refreshes, or signal-mix audits.
Smoke: .venv/bin/python -m pytest ai/tests/test_corpus_base.py ai/tests/test_chug.py ai/tests/test_konvid_1k.py ai/tests/test_konvid_150k.py ai/tests/test_lsvq.py ai/tests/test_live_vqc.py ai/tests/test_waterloo_ivc.py ai/tests/test_youtube_ugc.py -q
Touched files: ai/src/corpus/base.py, ai/scripts/chug_to_corpus_jsonl.py, ai/scripts/konvid_1k_to_corpus_jsonl.py, ai/scripts/konvid_150k_to_corpus_jsonl.py, ai/scripts/youtube_ugc_to_corpus_jsonl.py, ai/scripts/lsvq_to_corpus_jsonl.py, ai/scripts/live_vqc_to_corpus_jsonl.py, ai/scripts/waterloo_ivc_to_corpus_jsonl.py, ai/tests/test_corpus_base.py, ai/tests/test_chug.py, ai/AGENTS.md, docs/ai/*.md ingestion docs, docs/adr/0676-*.md, docs/research/0697-*.md, changelog.d/added/0676-*.md, and this file.
feat/full-feature-exporter-manifests (ADR-0668 follow-up)¶
No upstream rebase impact: this touches fork-local AI corpus exporters, tests, docs, package AGENTS notes, a research digest, and a changelog fragment. Upstream Netflix/vmaf does not ship these KoNViD or BVI-DVC training-table builders.
Invariant: ai/scripts/konvid_to_full_features.py and ai/scripts/bvi_dvc_to_full_features.py write <out>.manifest.json by default using aiutils.run_manifest. Keep the manifest beside refreshed local parquets so later model cards can prove source roots, cache/model inputs, feature order, and row/clip counts.
Smoke: .venv/bin/python -m pytest ai/tests/test_konvid_full_features.py ai/tests/test_bvi_dvc_dir_mode.py -q
Touched files: ai/scripts/konvid_to_full_features.py, ai/scripts/bvi_dvc_to_full_features.py, ai/tests/test_konvid_full_features.py, ai/tests/test_bvi_dvc_dir_mode.py, ai/AGENTS.md, docs/ai/training.md, docs/ai/bvi-dvc-corpus-ingestion.md, docs/research/0696-full-feature-exporter-manifests.md, changelog.d/added/0696-full-feature-exporter-manifests.md, and this file.
feat/u2netp-mirror-exporter (ADR-0671)¶
No upstream rebase impact: this touches fork-local tiny-AI exporter tooling, tests, docs, ADR/research notes, and a changelog fragment. Upstream Netflix/vmaf does not ship the U2NetP mirror workflow.
Invariant: ai/scripts/export_u2netp_mirror.py imports an audited local xuebinqin/U-2-Net checkout and writes a gitignored ONNX plus manifest. Do not vendor upstream U-2-Net source here, do not accept non-Apache license text, and do not commit model/u2netp_mirror.onnx.
Smoke: .venv/bin/python -m pytest ai/tests/test_export_u2netp_mirror.py -q
Touched files: ai/scripts/export_u2netp_mirror.py, ai/tests/test_export_u2netp_mirror.py, ai/AGENTS.md, docs/ai/u2netp-mirror.md, docs/ai/models/u2netp_mirror_card.md, docs/ai/training.md, docs/adr/0671-*.md, docs/adr/_index_fragments/0671-*.md, docs/research/0691-*.md, changelog.d/added/0671-*.md, and this file.
feat/tune-score-backend-native-priority (ADR-0667)¶
No upstream rebase impact: this touches fork-local vmaf-tune backend-selection code, docs, tests, AGENTS notes, and ADR/research notes. Upstream Netflix/vmaf does not ship the fork vmaf-tune automation harness.
Invariant: tools/vmaf-tune/src/vmaftune/score_backend.py keeps DEFAULT_FALLBACKS = ("cuda", "sycl", "hip", "cpu"). The Vulkan entry was removed when ADR-0726 dropped the Vulkan backend; do not re-add it during backend-selector rebases. CPU remains the final fallback.
Smoke: .venv/bin/python -m pytest tools/vmaf-tune/tests/test_score_backend.py -q
Touched files: tools/vmaf-tune/src/vmaftune/score_backend.py, tools/vmaf-tune/tests/test_score_backend.py, tools/vmaf-tune/AGENTS.md, docs/usage/vmaf-tune.md, docs/usage/vmaf-tune-score-backend.md, docs/adr/0667-*.md, docs/adr/_index_fragments/0667-*.md, docs/research/0687-*.md, changelog.d/changed/0667-*.md, and this file.
feat/tune-report-quick-takeaways (ADR-0666)¶
No upstream rebase impact: this touches fork-local vmaf-tune report rendering, tests, user docs, ADR/research notes, AGENTS notes, and a changelog fragment. Upstream Netflix/vmaf does not ship the fork vmaf-tune profile-card renderer.
Smoke: .venv/bin/python -m pytest tools/vmaf-tune/tests/test_report.py -q
Touched files: tools/vmaf-tune/src/vmaftune/report.py, tools/vmaf-tune/tests/test_report.py, tools/vmaf-tune/AGENTS.md, docs/usage/vmaf-tune.md, docs/adr/0666-*.md, docs/adr/_index_fragments/0666-*.md, docs/research/0686-*.md, changelog.d/added/0666-*.md, and this file.
fix/fast-nr-calibration-quality-guard (ADR-0665)¶
No upstream rebase impact: this touches fork-local tiny-AI calibration tooling, vmaf-tune docs, package AGENTS notes, ADR/research notes, and a changelog fragment. Upstream Netflix/vmaf does not ship the fork nr_metric_v1 fast-NR sidecar calibration workflow.
Smoke: .venv/bin/python -m pytest ai/tests/test_calibrate_nr_threshold.py -q
Touched files: ai/scripts/calibrate_nr_threshold.py, ai/tests/test_calibrate_nr_threshold.py, ai/AGENTS.md, tools/vmaf-tune/AGENTS.md, docs/usage/vmaf-tune-fast-nr.md, docs/ai/training.md, docs/adr/0665-*.md, docs/adr/_index_fragments/0665-*.md, docs/research/0685-*.md, changelog.d/fixed/0665-*.md, and this file.
feat/ai-validation-report-provenance (ADR-0661)¶
No upstream rebase impact: this touches fork-local tiny-AI validation tooling, model docs, ADR/research notes, and changelog fragments. Upstream Netflix/vmaf does not ship the fork tiny-model registry or saliency-student validation surfaces.
Smoke: .venv/bin/python -m pytest ai/tests/test_validation_report_provenance.py -q
Touched files: ai/scripts/validate_model_registry.py, ai/scripts/validate_saliency_student.py, ai/tests/test_validation_report_provenance.py, docs/ai/model-registry.md, docs/ai/models/saliency_student_*.md, docs/ai/training.md, docs/research/0683-*.md, docs/adr/0661-ai-run-manifest-provenance.md, changelog.d/added/0683-*.md, and this file.
feat/vmaf-tiny-validator-report-provenance (ADR-0661)¶
No upstream rebase impact: this touches fork-local tiny-AI validator tooling, model docs, ADR/research notes, and changelog fragments. Upstream Netflix/vmaf does not ship the v2/v3/v4 tiny-VMAF validator CLI family.
Smoke: .venv/bin/python -m pytest ai/tests/test_vmaf_tiny_validator_reports.py -q
Touched files: ai/scripts/validate_vmaf_tiny_v*.py, ai/tests/test_vmaf_tiny_validator_reports.py, docs/ai/models/vmaf_tiny_v*.md, docs/ai/training.md, docs/research/0681-*.md, docs/adr/0661-ai-run-manifest-provenance.md, changelog.d/added/0681-*.md, and this file.
feat/saliency-student-metrics-provenance (ADR-0661)¶
No upstream rebase impact: this touches fork-local AI saliency training tooling, model docs, ADR/research notes, and changelog fragments. Upstream Netflix/vmaf does not ship the DUTS-trained saliency student metrics surface.
Smoke: .venv/bin/python -m pytest ai/tests/test_saliency_student_metrics_provenance.py -q
Touched files: ai/scripts/train_saliency_student.py, ai/scripts/train_saliency_student_v2.py, ai/tests/test_saliency_student_metrics_provenance.py, docs/ai/models/saliency_student_*.md, docs/ai/training.md, docs/research/0680-*.md, docs/adr/0661-ai-run-manifest-provenance.md, changelog.d/added/0680-*.md, and this file.
feat/dnn-exporter-manifest-provenance (ADR-0661)¶
No upstream rebase impact: this touches fork-local AI exporter tooling, tiny-model docs, ADR/research notes, and changelog fragments. Upstream Netflix/vmaf does not ship these DNN feature-model exporter sidecars.
Smoke: .venv/bin/python -m pytest ai/tests/test_dnn_exporter_run_provenance.py -q
Touched files: ai/scripts/export_tiny_models.py, ai/scripts/export_fastdvdnet_pre.py, ai/scripts/export_fastdvdnet_pre_placeholder.py, ai/scripts/export_transnet_v2.py, ai/scripts/export_transnet_v2_placeholder.py, ai/tests/test_dnn_exporter_run_provenance.py, docs/ai/models/*.md, docs/ai/training.md, docs/research/0679-*.md, docs/adr/0661-ai-run-manifest-provenance.md, changelog.d/added/0679-*.md, and this file.
feat/ensemble-manifest-provenance (ADR-0661)¶
No upstream rebase impact: this touches fork-local AI ensemble training tooling, docs, ADR/research notes, and changelog fragments. Upstream Netflix/vmaf does not ship the fr_regressor_v2_ensemble_v1 trainer/manifest surface.
Smoke: .venv/bin/python -m pytest ai/tests/test_train_fr_regressor_v2_ensemble.py -q
Touched files: ai/scripts/train_fr_regressor_v2_ensemble.py, ai/tests/test_train_fr_regressor_v2_ensemble.py, docs/ai/models/fr_regressor_v2_probabilistic.md, docs/ai/training.md, docs/research/0678-*.md, docs/adr/0661-ai-run-manifest-provenance.md, changelog.d/added/0678-*.md, and this file.
feat/nr-threshold-calibration-provenance (ADR-0661)¶
No upstream rebase impact: this touches fork-local AI/vmaf-tune calibration tooling, docs, ADR/research notes, and changelog fragments. Upstream Netflix/vmaf does not ship the vmaf-tune --fast-nr NR threshold calibration path.
Smoke: .venv/bin/python -m pytest ai/tests/test_calibrate_nr_threshold.py -q
Touched files: ai/scripts/calibrate_nr_threshold.py, ai/tests/test_calibrate_nr_threshold.py, docs/usage/vmaf-tune-fast-nr.md, docs/ai/training.md, docs/research/0677-*.md, docs/adr/0661-ai-run-manifest-provenance.md, changelog.d/added/0677-*.md, and this file.
feat/phase-f-calibration-provenance (ADR-0661)¶
No upstream rebase impact: this touches fork-local AI/vmaf-tune calibration tooling, docs, ADR/research notes, and changelog fragments. Upstream Netflix/vmaf does not ship the vmaf-tune auto Phase F recipe calibration path.
Smoke: .venv/bin/python -m pytest ai/tests/test_calibrate_phase_f_recipes.py -q
Touched files: ai/scripts/calibrate_phase_f_recipes.py, ai/tests/test_calibrate_phase_f_recipes.py, docs/usage/vmaf-tune.md, docs/ai/training.md, docs/research/0676-*.md, docs/adr/0661-ai-run-manifest-provenance.md, changelog.d/added/0676-*.md, and this file.
feat/quant-ep-report-provenance (ADR-0661)¶
No upstream rebase impact: this touches fork-local AI investigation tooling, AI docs, ADR/research notes, and changelog fragments. Upstream Netflix/vmaf does not ship this per-EP quantisation harness.
Smoke: .venv/bin/python -m pytest ai/tests/test_measure_quant_drop_per_ep.py -q
Touched files: ai/scripts/measure_quant_drop_per_ep.py, ai/tests/test_measure_quant_drop_per_ep.py, docs/ai/quant-eps.md, docs/research/0006-tinyai-ptq-accuracy-targets.md, docs/research/0675-quant-ep-report-provenance.md, docs/adr/0661-ai-run-manifest-provenance.md, changelog.d/added/0675-*.md, and this file.
fix/windows-cuda-toolkit-installer (ADR-0664)¶
High CI rebase impact: this touches the fork-local build matrix workflow. Upstream Netflix/vmaf does not ship these Windows GPU build-only legs, but workflow syncs can silently restore older action-based setup patterns.
Rebase-sensitive fork invariant:
Build — Windows MSVC + CUDA (build only)installs CUDA 13.2.0 directly from NVIDIA's Windows network installer and verifiesnvcc.exe --version. Do not restoreJimver/cuda-toolkiton this Windows leg without a superseding ADR and a green required Windows CUDA run.- Linux CUDA legs remain unchanged and still use
Jimver/cuda-toolkit.
Smoke: gh pr checks <pr> --watch --required
Touched files: .github/workflows/libvmaf-build-matrix.yml, .github/AGENTS.md, docs/development/ci-runners.md, docs/adr/0664-*.md, docs/research/0664-*.md, changelog.d/fixed/0664-*.md, and this file.
fix/external-bench-wrapper-schema (ADR-0656)¶
No upstream rebase impact: all touched implementation paths are fork-local external-bench tooling, docs, tests, ADR/research, and changelog fragments. Upstream Netflix/vmaf does not ship this benchmark harness.
Rebase-sensitive fork invariant:
summary.competitoremitted by everytools/external-bench/*/run.shwrapper must exactly match the registry key incompare.WRAPPERS. Model/version labels belong in optional metadata, not this identity field, orvalidate_wrapper_output()will reject the result before aggregation.
Smoke: .venv/bin/python -m pytest tools/external-bench/tests/ -q
Touched files: tools/external-bench/, docs/ai/external-bench.md, docs/adr/0656-*.md, docs/research/0656-*.md, changelog.d/fixed/0656-*.md, mkdocs.yml, and this file.
fix/tiny-ai-disabled-runtime-gate (ADR-0660)¶
Low upstream rebase impact: the touched C files are fork-local tiny-AI extractors and helper tests. Upstream Netflix/vmaf does not ship these DNN feature extractors, but conflicts are possible if upstream changes the feature registry or libvmaf's optional-DNN surface.
Rebase-sensitive fork invariant:
- Every tiny-AI feature extractor calls
vmaf_tiny_ai_require_runtime(<feature>)after pixel-format / bit-depth validation and beforevmaf_tiny_ai_resolve_model_path(). Disabled-DNN builds must return-ENOSYSbefore path probing; DNN-enabled builds keep missing model paths as-EINVAL.
Smoke: meson test -C build --suite=fast --print-errorlogs test_lpips test_dists test_fastdvdnet_pre test_mobilesal test_transnet_v2
Touched files: core/src/dnn/tiny_extractor_template.h, core/src/feature/feature_{lpips,dists,mobilesal}.c, core/src/feature/{fastdvdnet_pre,transnet_v2}.c, core/test/tiny_ai_test_template.h, core/src/dnn/AGENTS.md, docs/ai/, docs/metrics/features.md, docs/adr/0660-*.md, docs/research/0660-*.md, changelog.d/fixed/0660-*.md, and this file.
feat/saliency-feature-materializer (ADR-0655)¶
No upstream rebase impact: the implementation is fork-local AI tooling (ai/scripts/, ai/tests/) plus fork-local documentation and changelog files. Upstream Netflix/vmaf does not ship the fork's saliency training-table materializer.
Rebase-sensitive fork invariants:
ai/scripts/materialize_saliency_features.pyowns bulk saliency enrichment for existing JSONL/parquet feature tables; trainers consume the resultingsaliency_mean/saliency_varcolumns instead of silently running saliency inference inside training loops.- The status column remains row-local and human-readable (
ok,skipped-existing,missing-source,missing-geometry,decode-failed,model-failed) so large local sweeps can be audited without scraping stderr. SaliencyMaterializeConfig.default_width/default_height(added PR fixing the Netflix refresh materializer): fallback geometry for raw YUV corpora without container headers. Netflix corpus YUVs are always 1920×1080 at rest.- For
.yuvsources, the ffmpeg decode prepends-f rawvideo -video_size WxH -pix_fmt yuv420pbefore-i; do not remove this for raw-YUV support. - In-process per-file saliency cache in
materialize_rows()avoids redundant decodes for per-frame tables; the cache is scoped to onematerialize_rows()call and does not persist across batch table boundaries.
Smoke: PYTHONPATH=. .venv/bin/python -m pytest ai/tests/test_materialize_saliency_features.py -q
feat/signal-mix-audit (ADR-0650)¶
No upstream rebase impact: all implementation paths are fork-local AI tooling, tests, and documentation. Upstream Netflix/vmaf does not ship this training/audit package or the associated docs.
Rebase-sensitive fork invariants:
ai/scripts/signal_mix_audit.pyremains table-only and side-effect free: no feature extraction, checkpoint export, corpus mutation, or default CI gate.- Signal-family regexes and
docs/ai/signal-mix-audit.mdmust be updated together when new metric families or table columns are introduced. - Missing candidate metrics in the Markdown report are advisory work selectors, not proof that a candidate should be promoted without a corpus run.
Smoke: .venv/bin/python -m pytest ai/tests/test_signal_mix_audit.py -q
Touched files: ai/scripts/signal_mix_audit.py, ai/tests/test_signal_mix_audit.py, ai/AGENTS.md, docs/ai/signal-mix-audit.md, docs/adr/0650-*.md, docs/research/0650-*.md, changelog.d/added/0650-*.md, and this file.
fix/dnn-attached-multi-output (ADR-0646)¶
Low upstream rebase impact: the implementation touches fork-local DNN runtime plumbing plus libvmaf's context bridge. Upstream Netflix/vmaf does not ship the fork's ONNX Runtime attached tiny-AI surface, but conflicts are possible if upstream changes core/src/libvmaf.c near the per-frame pipeline.
Rebase-sensitive fork invariants:
- Single-output attached tiny models keep the historical collector key exactly. Do not append
_scoreor an ONNX output suffix for one-output models. - Multi-output attached models route through
vmaf_ort_run(), notvmaf_ort_infer(). The latter is intentionally a single-output helper. - Sidecar
output_names[]wins only when its count matches the ONNX output count; otherwise ONNX output names are used and sanitized. - Attached mode remains scalar-only. Vector or image output tensors must still use
vmaf_dnn_session_run()until a future ADR defines feature-name flattening.
Smoke: docker exec vmaf-dev-mcp bash -lc 'cd /workspace && rm -rf /tmp/vmaf-dnn-multi-output-build && meson setup /tmp/vmaf-dnn-multi-output-build core -Denable_dnn=enabled -Denable_cuda=false -Denable_sycl=false -Denable_vulkan=disabled -Denable_hip=false -Denable_metal=disabled && meson test -C /tmp/vmaf-dnn-multi-output-build --suite=dnn --print-errorlogs'
Touched files: core/src/libvmaf.c, core/src/dnn/model_loader.*, core/src/dnn/ort_backend.*, core/test/dnn/*, model/tiny/smoke_multi_output_v0.*, scripts/gen_multi_output_smoke_onnx.py, docs/api/dnn.md, docs/ai/, docs/adr/0646-*.md, docs/research/0646-*.md, changelog.d/fixed/0646-*.md, and this file.
fix/ai-refresh-defaults-and-konvid-full-features (ADR-0642)¶
No upstream rebase impact: all touched implementation files live under fork-local ai/ tooling. Upstream Netflix/vmaf does not ship these training scripts, model-refresh docs, or local corpus ledgers.
Rebase-sensitive fork invariants:
- AI feature extraction defaults point at
core/build-cpu/tools/vmaf. Do not regress to/usr/local/bin/vmafor ambiguousbuild/tools/vmaf; stale binaries have previously lacked fork-only extractors. ai/scripts/konvid_to_full_features.pyowns regeneration of bothruns/full_features_konvid.parquetandruns/full_features_konvid_with_folds.parquet. The folded output'ssource=fold0..fold4assignment is a deterministic balanced hash over clip keys and feedseval_multiseed_v3_v4.py.- BVI-DVC full-feature dir mode accepts
.mkv,.mp4, and.yuv. The local.mkvlossless bundle is the known-good refresh input after the raw-YUV copy produced all-zero VMAF in a one-clip smoke. ai/scripts/extract_ugc_features.pyemits the currentFULL_FEATURESschema with an explicitvmaf_v0.6.1model path. Do not restore the historical canonical-6-only UGC table when refreshingfull_features_5corpus.- Aggregate full-feature training tables are rebuilt with
ai/scripts/combine_full_feature_parquets.py; the normalized schema iscorpus, source, frame_index, codec, <FULL_FEATURES>, vmaf.
Smoke: .venv/bin/python -m pytest ai/tests/test_konvid_full_features.py ai/tests/test_extract_ugc_features.py ai/tests/test_combine_full_feature_parquets.py ai/tests/test_feature_extractor_defaults.py ai/tests/test_bvi_dvc_dir_mode.py -q
Touched files: ai/data/feature_extractor.py, ai/scripts/*full_features*.py, ai/scripts/konvid_to_full_features.py, ai/src/vmaf_train/, ai/tests/test_*, ai/AGENTS.md, docs/ai/, docs/adr/0642-*.md, docs/research/0642-*.md, changelog.d/added/, and .workingdir2/AI_REFRESH_2026-05-20.md (ignored local ledger).
fix/dev-container-encoder-probes (ADR-0641)¶
Low upstream rebase impact: implementation changes are fork-local dev-container / vmaf-tune files (dev/, tools/vmaf-tune/, docs, ADR, research, changelog) plus one fork-local FFmpeg integration patch. Upstream Netflix/vmaf does not ship vmaf-tune or this dev-MCP compose stack. The only upstream-adjacent file is ffmpeg-patches/0003-*, which targets FFmpeg n8.1.1 rather than Netflix/vmaf.
Rebase-sensitive fork invariants:
dev/Containerfilemust keep the pinnedintel/vpl-gpu-rtsource build and post-install/usr/lib/x86_64-linux-gnu/libmfx-gen.socheck whenever FFmpeg keeps--enable-libvpl.libvpl-devalone exposes QSV encoders but cannot create an Arc/iGPU session, and installing the runtime outside the dispatcher search path revives the sameMFX_ERR_NOT_FOUNDfailure.dev/docker-compose.ymlmust keep thedev-mcphealthcheck aligned with the stdio entrypoint (vmaf --version), not/sockets/vmaf-mcp.sock.vmaf-tune comparedefaults to the production CPU setlibx265,libsvtav1; archival software codecs remain explicit via--encoders.- QSV VA-API device selection defaults to
autoand uses Intel sysfs vendor-ID discovery; explicit--vaapi-devicepaths still override. ffmpeg-patches/0003-*must callvmaf_sycl_state_free(&s->sycl_state). The public SYCL API frees and nulls aVmafSyclState **; using the older single-pointer call breaks the in-container FFmpeg build with-Wincompatible-pointer-types.
Touched files: dev/Containerfile, dev/docker-compose.yml, dev/AGENTS.md, tools/vmaf-tune/src/vmaftune/{bisect.py,cli.py,compare.py,hw_devices.py}, tools/vmaf-tune/src/vmaftune/codec_adapters/_qsv_common.py, tools/vmaf-tune/tests/, ffmpeg-patches/0003-*, docs/usage/vmaf-tune.md, docs/development/dev-mcp.md, docs/state.md, docs/adr/0641-*.md, docs/research/0641-*.md, changelog.d/fixed/0641-*.md, and this file.
chore/ci-warning-omnibus (ADR-0635)¶
No rebase impact: all touched files are fork-local CI workflow YAML (.github/workflows/libvmaf-build-matrix.yml), fork-added docs (docs/mcp/tools.md, docs/adr/, docs/research/, changelog.d/), and this file. Upstream Netflix/vmaf does not use GitHub Actions workflows that overlap with these changes. No C sources, no public headers, and no FFmpeg patch series are involved.
Touched files: .github/workflows/libvmaf-build-matrix.yml (ilammy→TheMrMilchmann action swap; windows-latest→windows-2025; vulkaninfo stderr redirect + debug demotion; ccache-v2 key prefix), docs/mcp/tools.md (run_benchmark heading backtick removal + a-id drop), docs/adr/0635-ci-warning-omnibus-2026-05-19.md, docs/adr/README.md (one index row), docs/research/ci-warning-omnibus-2026-05-19.md, changelog.d/fixed/0635-ci-warning-omnibus.md, docs/rebase-notes.md (this entry).
ADR-0672 — Saliency materializer temporal controls¶
Saliency-table provenance impact. This widens the ADR-0655 materializer from historical mean-only saliency to the same temporal reducer family exposed by vmaf-tune.
Key invariants:
ai/scripts/materialize_saliency_features.pyforwards--temporal-aggregatorand--ema-alphaintovmaftune.saliency.compute_saliency_map().- Newly computed rows record
saliency_model_id,saliency_aggregator, andsaliency_ema_alphaby default. - Rows skipped because they already contain finite saliency columns must not get invented model/reducer metadata; use
--overwritefor intentional replacement.
Touched files: ai/scripts/materialize_saliency_features.py, ai/tests/test_materialize_saliency_features.py, ai/AGENTS.md, docs/ai/saliency-feature-materializer.md, docs/ai/u2netp-mirror.md, docs/adr/0672-saliency-materializer-temporal-controls.md, docs/research/0692-saliency-materializer-temporal-controls.md, changelog.d/added/0672-saliency-materializer-temporal-controls.md, docs/rebase-notes.md (this entry).
ADR-0654 — Predictor saliency signals¶
vmaf-tune predict --use-saliency is a predictor-feature switch, not the ROI/QP sidecar path. Preserve the temporary raw-yuv420p decode in predictor_features._compute_saliency() before calling saliency.compute_saliency_map(raw_path, width, height, ...); the saliency helper remains raw-YUV-only even though the public predict source can be any FFmpeg-readable container.
predictor_train.project_row() must keep the 14-column predictor input layout stable. When real corpora carry probe_*_avg_bytes, saliency_mean, saliency_var, frame_diff_mean, y_avg, or y_var, preserve those finite values. Only legacy rows should fall back to bitrate-derived probe bytes and zero saliency / signalstats values.
fix/ci-test-failures-omnibus (ADR-0637)¶
No rebase impact: all touched files are fork-local CI configuration (.github/workflows/tests-and-quality-gates.yml), MCP server tests (mcp-server/vmaf-mcp/tests/test_smoke_e2e.py), ADR files, and changelog fragments. No upstream C sources, no public headers, no FFmpeg patch series involved. The timeout and coverage-floor edits are fork-CI-specific and have no upstream equivalent.
fix/scaffold-audit-p0-silent-correctness (ADR-0620)¶
No rebase impact: all touched files are fork-local Python harness files and docs. No upstream C sources, no public headers, no FFmpeg patch series involved. The three fixed Python files (routine.py, train_test_model.py, local_explainer.py) are also present upstream, but the specific exception-handling changes are in fork-added call paths (extended-stats bagging, plot_scatter visualisation, local-explainer model dispatch). If upstream lands a conflicting change to these exact lines, the merge resolution is straightforward: keep the raise paths and update context if the upstream change affects surrounding logic.
Touched files: python/vmaf/tools/exceptions.py (3 new exception classes), python/vmaf/routine.py (P0-1 fix + CalibrationError import), python/vmaf/core/train_test_model.py (P0-2 fix + MissingLabelStddevError import), python/vmaf/core/local_explainer.py (P0-3 fix + EnsembleNotSupportedError import), python/test/test_adr0620_scaffold_audit_p0.py (16 regression tests), docs/adr/0620-scaffold-audit-p0-silent-correctness-fixes.md, docs/adr/README.md (one index row), docs/state.md (3 rows moved from Open to Recently closed), changelog.d/fixed/adr0620-scaffold-audit-p0-silent-correctness.md,
fix/scaffold-audit-p1-feature-plumbing (ADR-0299)¶
Touches core/src/hip/picture_hip.c, core/src/feature/feature_mobilesal.c, and core/src/libvmaf.c. Upstream Netflix/vmaf does not have a HIP backend, the mobilesal extractor, or the DNN multi-output guard — so no rebase conflict is expected on any of the C-side changes.
Touches tools/vmaf-tune/src/vmaftune/cli.py — upstream does not have vmaf-tune. No rebase conflict expected.
Doc paths (docs/api/dnn.md, docs/ai/models/mobilesal.md, docs/state.md, docs/adr/README.md) are fork-local only.
Rebase-sensitive invariant (C): picture_hip.c now compiles in two branches: #ifdef HAVE_HIPCC (real hipMalloc) and #else (-ENOSYS). Any upstream change to picture_hip.h's function signatures must be reflected in both branches.
Touched files: core/src/hip/picture_hip.c, core/src/feature/feature_mobilesal.c, core/src/libvmaf.c (comment-only at lines 1115, 1214), tools/vmaf-tune/src/vmaftune/cli.py, docs/api/dnn.md, docs/ai/models/mobilesal.md, docs/state.md, docs/adr/0639-scaffold-audit-p1-feature-plumbing-fixes.md, docs/adr/README.md, changelog.d/fixed/adr-0613-scaffold-audit-p1.md, docs/rebase-notes.md (this entry).
feat/zed-editor-project-config (ADR-0608)¶
No rebase-sensitive invariants — only .zed/ (new directory), .gitignore (.zed/local/ exclusion), docs/development/ide-setup.md (Zed section), docs/adr/0608-zed-editor-project-config.md (ADR), and supporting fragment/ changelog files are touched. None of these paths overlap with upstream Netflix/vmaf. .vscode/ is unchanged.
Touched files: .zed/settings.json, .zed/tasks.json, .zed/debug.json (new), .gitignore (.zed/local/ entry), docs/development/ide-setup.md (Zed section appended), docs/adr/0608-zed-editor-project-config.md, docs/adr/_index_fragments/0608-zed-editor-project-config.md, docs/adr/_index_fragments/_order.txt, docs/adr/README.md (regenerated), changelog.d/added/0608-zed-editor-project-config.md, docs/rebase-notes.md (this entry).
plan/netflix-grade-encoding-roadmap (ADR-0299 – ADR-0618)¶
No rebase-sensitive invariants — all changes are planning documents only: six ADRs, six research digests, one roadmap overview, one changelog fragment, and ADR index rows in docs/adr/README.md. No C sources, headers, build files, or Python implementation files are touched. No upstream-shared paths are modified.
Touched files: docs/adr/0613-dynamic-optimizer.md, docs/adr/0614-per-shot-abr-rendition.md, docs/adr/0615-fast-nr-prescoring.md, docs/adr/0616-vmaf-neg-integration.md, docs/adr/0617-cross-shot-complexity-weighting.md, docs/adr/0618-content-aware-classifier.md, docs/adr/README.md (index rows), docs/research/0609-dynamic-optimizer-research.md, docs/research/0610-per-shot-abr-rendition-research.md, docs/research/0611-fast-nr-prescoring-research.md, docs/research/0612-vmaf-neg-integration-research.md, docs/research/0613-cross-shot-complexity-weighting-research.md, docs/research/0614-content-aware-classifier-research.md, docs/development/netflix-grade-encoding-pipeline-roadmap-2026-05-19.md, changelog.d/added/netflix-grade-encoding-pipeline-roadmap.md.
chore/scaffold-audit-p3-cleanup (ADR-0621)¶
No rebase-sensitive invariants. All changes are in fork-local files (ai/scripts/, scripts/dev/, python/test/, .semgrepignore, docs/ai/model-registry.md, docs/adr/, docs/state.md, changelog.d/). None of the touched Python test files are shared with Netflix upstream (Netflix does not ship asset_test.py or quality_runner_test.py). The python/test/*.py files the PR touches carry fork-added tests or skip-decorator updates; no upstream test assertions are modified.
Touched files: scripts/dev/permutation_importance.py, ai/scripts/*.py (13 files), python/test/result_test.py, python/test/routine_test.py, python/test/asset_test.py, python/test/feature_extractor_test.py, python/test/quality_runner_test.py, .semgrepignore, docs/ai/model-registry.md, docs/adr/0621-scaffold-audit-p3-cleanup.md, docs/adr/README.md, docs/state.md, changelog.d/fixed/0621-scaffold-audit-p3-cleanup.md.
feat/mcp-p1-vmaftune-extractors-models-progress (ADR-0608)¶
No rebase-sensitive invariants. The only changed files are:
mcp-server/vmaf-mcp/src/vmaf_mcp/server.py— fork-local MCP server, never in Netflix upstream.mcp-server/vmaf-mcp/tests/— fork-local tests.mcp-server/vmaf-mcp/tests/test_smoke_e2e.py— updated expected tool-name set.docs/mcp/tools.md,docs/adr/0608-*.md,docs/adr/README.md,docs/rebase-notes.md— docs.changelog.d/added/0608-*.md— changelog fragment.
No C sources, public headers, meson_options.txt, ffmpeg-patches/, or build files are touched.
chore/renovate-customManagers-dev-image (ADR-0605)¶
No rebase-sensitive invariants — the only change is to renovate.json (adding eight new customManagers entries for Containerfile ARG-pinned deps; extending the FFmpeg manager's managerFilePatterns to also scan dev/Containerfile). renovate.json is fork-local and never appears in upstream Netflix/vmaf. No C sources, headers, or build files are touched.
Touched files: renovate.json (customManagers + packageRules), docs/adr/0605-renovate-custommgr-dev-image.md, docs/adr/README.md (one index row), changelog.d/changed/0605-renovate-custommgr-dev-image.md, docs/rebase-notes.md (this entry).
chore/rocm-7-13-bump-and-renovate-manager (ADR-0604)¶
No rebase-sensitive invariants — the only change is to renovate.json (adding a customManagers entry and customDatasources block for ROCm). renovate.json is fork-local and never appears in upstream Netflix/vmaf. dev/Containerfile is unchanged (7.2.3 remains the correct pin).
Touched files: renovate.json (customManagers + customDatasources), docs/adr/0604-rocm-renovate-manager.md, docs/adr/README.md (one index row), docs/research/rocm-version-audit-2026-05-19.md, changelog.d/changed/0604-rocm-renovate-manager.md, docs/rebase-notes.md (this entry).
fix/ubuntu-26-04-fallout (ADR-0603)¶
No rebase-sensitive invariants — all changes are in the build/CI layer (dev/Containerfile, CI workflow YAML, meson.build nvcc flags, pyproject.toml ceiling bumps) and do not touch any upstream-shared C sources, public headers, or Python test assertions.
The one meson.build addition (-D__MATH_NO_INLINES in cuda_flags) is additive and harmless on any glibc version; if upstream Netflix touches the CUDA flags block in core/src/meson.build, preserve the -D__MATH_NO_INLINES entry alongside whatever upstream adds.
Touched files: dev/Containerfile, core/src/meson.build (cuda_flags), tools/vmaf-tune/pyproject.toml (requires-python ceiling), ai/pyproject.toml (requires-python ceiling), .github/workflows/libvmaf-build-matrix.yml (CUDA version pin), docs/adr/0603-ubuntu-26-04-fallout-fixes.md, docs/adr/README.md (index row), changelog.d/fixed/ubuntu-26-04-fallout.md, docs/rebase-notes.md (this entry).
fix/macos-vmaf-write-output-segv (ADR-0602)¶
No rebase-sensitive invariants — the changes are purely defensive guards (NULL checks, pic_cnt > 0 guards) added to existing functions in core/src/libvmaf.c and core/src/output.c, and a new test in core/test/test_output.c. If upstream Netflix merges any change to vmaf_write_output_with_format or vmaf_write_output_json, re-apply the three guards (vmaf-NULL, feature_collector-NULL, output_path-NULL) and the pic_cnt > 0 guards in json_write_pooled_entry / xml_write_one_metric_pools to the merged version.
Touched files: core/src/libvmaf.c (NULL guards at top of vmaf_write_output_with_format), core/src/output.c (pic_cnt > 0 guards, NULL guards in JSON writer, split xml_write_pooled_and_aggregate into three helpers, remove unused n_frames variables), core/test/test_output.c (test_write_output_pic_cnt_zero regression test), docs/adr/0602-macos-vmaf-write-output-segv.md, docs/adr/README.md (index row), docs/state.md (Recently-closed row), docs/rebase-notes.md (this entry), changelog.d/fixed/0602-macos-vmaf-write-output-segv.md.
fix/vmaftune-qsv-amf-hw-init-and-probe-size (ADR-0601)¶
Rebase impact: tools/vmaf-tune/ only — no libvmaf C sources, public headers, or meson_options.txt touched. Zero upstream conflict surface.
Rebase-sensitive invariants:
compare._QSV_ENCODERSmust stay in sync with the set of QSV encoder names registered incodec_adapters/. If a new QSV adapter is added (e.g.vp9_qsv), add its encoder string to_QSV_ENCODERSin the same commit; omitting it silently skips the VA-API init chain for that encoder.BaseQsvAdapter.qsv_hw_init_args()andcompare._hw_init_args_for_encoder()must produce identical flag sequences. If one is updated, update the other. A test intest_bbb_e2e_v14_bug_cluster.pyverifies this invariant.- The default
_DEFAULT_VAAPI_DEVICE = "/dev/dri/renderD128"is also the default inBaseQsvAdapter.qsv_hw_init_args. Keep them in sync.
Touched files: tools/vmaf-tune/src/vmaftune/compare.py, tools/vmaf-tune/src/vmaftune/cli.py, tools/vmaf-tune/src/vmaftune/codec_adapters/_qsv_common.py, tools/vmaf-tune/src/vmaftune/codec_adapters/_amf_common.py, tools/vmaf-tune/tests/test_bbb_e2e_v14_bug_cluster.py, docs/adr/0601-vmaftune-qsv-amf-hw-init-and-probe-fix.md, docs/adr/README.md (one index row), docs/usage/vmaf-tune.md (--vaapi-device flag + QSV init docs), docs/state.md (T-BBB-V14-HW-ENCODER-PROBE-QSV-INIT-2026-05-18 row), changelog.d/fixed/0601-vmaftune-qsv-amf-hw-init-and-probe-fix.md, docs/rebase-notes.md (this entry).
chore/ffmpeg-patches-n811-full-feature-exposure-sync (ADR-0576)¶
Rebase impact: ffmpeg-patches/ only — no libvmaf C sources, public headers, or meson_options.txt touched. Upstream Netflix/vmaf does not ship ffmpeg-patches/; no rebase conflict surface.
Rebase-sensitive invariants:
- Patch 0014 targets the
LIBVMAFContextstruct andVmafConfigurationinit blocks introduced cumulatively by patches 0003–0013. It must remain the final patch in the series (or be rebased against whichever patch last touches those init blocks if the series is reordered). - The
cpumask/gpumaskAVOption names must match the field names inVmafConfigurationfromcore/include/libvmaf/libvmaf.h. If a future libvmaf refactor renames those fields, patch 0014's struct designators (.cpumask =,.gpumask =) must be updated to match. - The
feature=passthrough in the stocklibvmaffilter continues to cover all extractors infeature_extractor_list[]; no patch is needed for new extractor additions unless they require a dedicated C-API init call (e.g., a newvmaf_<backend>_state_init()entry point).
fix/ffmpeg-patches-score-fmt-gap (ADR-1064)¶
Rebase impact: ffmpeg-patches/ only — adds patch 0016 and updates series.txt and README.md. No libvmaf C sources, public headers, or meson_options.txt touched.
Rebase-sensitive invariants:
- Patch 0016 must come after patch 0014 (which adds
cpumask/gpumasktoLIBVMAFContext). Patch 0016 addsscore_fmtimmediately after theint64_t gpumaskfield; if 0014 is reordered or the struct layout changes, the context lines in 0016's struct hunk must be updated. - The
vmaf_write_output_with_formatsymbol must be present in the libvmaf version checked bypkg-config. If a future refactor renames this entry point, all four uninit paths in patch 0016 must be updated. - Patch 0016 requires
git am --3wayreplay against all 15 preceding patches before verifying clean apply against n8.1.1 (the patch series is cumulative).
Re-test on rebase:
git clone --depth 1 --branch n8.1.1 https://git.ffmpeg.org/ffmpeg.git /tmp/ffmpeg-retest
git -C /tmp/ffmpeg-retest config user.email "lusoris@pm.me"
git -C /tmp/ffmpeg-retest config user.name "Lusoris"
for p in ffmpeg-patches/*.patch; do
git -C /tmp/ffmpeg-retest am --3way "$p" || { echo "FAILED: $p"; break; }
done
# Expect 14 commits applied cleanly, no conflicts.
Upstream Netflix/vmaf has no ffmpeg-patches/; no rebase conflict surface against upstream/master. All 14 patches are fork-local.
feat/vmaftune-bisect-concurrency-cap (ADR-0577)¶
Rebase impact: pure Python — touches only tools/vmaf-tune/ and docs/. No C surface, no meson.build change, no public C-API change, no GPU path change.
Rebase-sensitive invariant: none. The _decode_semaphore singleton and set_decode_semaphore setter are new module-level additions in vmaftune/bisect.py; they do not conflict with any existing upstream pattern. The decode_semaphore keyword argument added to bisect_target_vmaf and make_bisect_predicate is backwards-compatible (defaults to None, falling back to the module-level semaphore).
Touched files: tools/vmaf-tune/src/vmaftune/bisect.py, tools/vmaf-tune/src/vmaftune/cli.py, tools/vmaf-tune/tests/test_bisect_concurrency_cap.py (new), tools/vmaf-tune/tests/test_bisect.py (exports check update), tools/vmaf-tune/tests/test_compare.py (semaphore kwarg assertion), docs/adr/0577-vmaftune-bisect-concurrency-cap-and-aggressive-cleanup.md (new), docs/adr/README.md (one index row), docs/usage/vmaf-tune.md (--max-concurrent-decodes docs + disk-mgmt section), changelog.d/fixed/vmaf-tune-bisect-concurrency-cap-enospc.md (new), docs/rebase-notes.md (this entry).
fix/windows-ci-sdk-pin-22621 (ADR-0575)¶
Rebase impact: tools only — touches core/tools/yuv_input.c. No meson.build change, no public C-API change, no GPU path change.
Rebase-sensitive invariant: #include <sys/stat.h> must remain before the #ifdef _MSC_VER macro block in yuv_input.c. If a rebase reorders these lines (e.g. by re-applying a prior ADR-0521 patch that placed the macros before the include), the MinGW64 and MSVC+SDK-26100 redefinition errors will recur.
Touched files: core/tools/yuv_input.c, docs/adr/0575-windows-msvc-stat-compat-include-order.md, docs/adr/README.md (one index row), docs/state.md (Updated note + T-WINDOWS-STAT-COMPAT row in Recently closed), changelog.d/fixed/0575-windows-stat-compat-include-order.md, docs/rebase-notes.md (this entry).
feat/integer-ssim-gpu-real-kernels (ADR-0564)¶
Rebase impact: low. The change touches two upstream-shared files:
core/src/feature/feature_extractor.c: adds threeexterndeclarations and three list entries (vmaf_fex_integer_ssim_cuda,vmaf_fex_integer_ssim_sycl, and a comment update). On rebase, apply after any upstream changes to this file.core/src/meson.build: adds one entry tocuda_cu_sourcesdict and one entry to the C source list. The meson.build is append-only per fork coordination rules.core/src/feature/hip/integer_ssim_hip.c: full rewrite of the host glue. The pre-existing upstream file used float intermediates; this branch rewrites it to int64. If upstream ever ships a real integer_ssim HIP extractor, it will conflict — prefer the upstream version and re-test.core/src/feature/sycl/integer_ssim_sycl.cpp: appends a new extractor after the existing float_ssim_sycl code. On rebase, confirm the append point is still a clean} /* extern "C" */boundary.
All new files (ssim_cuda.c, ssim_cuda.h, integer_ssim_score.cu) are fork-local with no upstream equivalent; no conflict expected.
Invariant: vmaf_fex_integer_ssim_cuda in ssim_cuda.c provides "ssim". The pre-existing vmaf_fex_integer_ssim_cuda in integer_ssim_cuda.c provides "float_ssim" — the naming is a historical misnomer kept for link-compat. Do not merge or rename without updating feature_extractor.c to match.
Touched files: core/src/feature/cuda/integer_ssim/integer_ssim_score.cu (new), core/src/feature/cuda/ssim_cuda.c (new), core/src/feature/cuda/ssim_cuda.h (new), core/src/feature/hip/integer_ssim_hip.c (rewritten), core/src/feature/sycl/integer_ssim_sycl.cpp (appended), core/src/feature/feature_extractor.c (extern + list entries), core/src/meson.build (PTX + C source entries), docs/adr/0564-integer-ssim-gpu-real-kernels.md, docs/adr/README.md (one index row), docs/research/0564-integer-ssim-gpu-real-kernels.md, docs/state.md (Recently-closed row), changelog.d/added/0564-integer-ssim-gpu-real-kernels.md, docs/rebase-notes.md (this entry).
fix/vmaftune-workdir-tmpfs-enospc (ADR-0598)¶
No rebase impact. All changes are confined to fork-local files:
tools/vmaf-tune/src/vmaftune/bisect.py(fork-added tool).tools/vmaf-tune/src/vmaftune/cli.py(fork-added tool).tools/vmaf-tune/tests/test_workdir_enospc.py(new test file).tools/vmaf-tune/tests/test_compare.py(update expected kwargs).dev/Containerfile(fork-local; whole file is fork-added).dev/scripts/dev-mcp-entrypoint.sh(fork-local).docs/adr/0549-vmaftune-workdir-relocation.md,docs/state.md,docs/usage/vmaf-tune.md,docs/adr/README.md,changelog.d/fixed/vmaf-tune-enospc-workdir.md,docs/rebase-notes.md(fork-only doc tree).
No upstream-shared paths touched. VMAFTUNE_WORKDIR is a new fork-local environment variable; it has no upstream counterpart and poses no rebase conflict risk.
docs/vcq-223-local-explainer-hang-diagnosis (ADR-0563)¶
No rebase impact. All changes are confined to fork-local documentation:
docs/adr/0551-local-explainer-hang-diagnosis.md(new file, fork-local).docs/research/0551-local-explainer-hang.md(new file, fork-local).docs/state.md— updated T-VCQ-223-LOCAL-EXPLAINER-HANG row (fork-local).changelog.d/fixed/0551-local-explainer-hang-diagnosis.md(new file, fork-local).docs/adr/README.md— new index row (fork-local).docs/rebase-notes.md— this entry (fork-local).
No upstream-shared C sources, Python sources, or build files are touched. The @unittest.skip decorator in python/test/local_explainer_test.py is explicitly not removed in this PR — that is a follow-up code change.
chore/hip-cuda-orphan-tu-cleanup (ADR-0546)¶
No rebase impact. All deleted files (adm_hip.c, motion_hip.c, vif_hip.c, feature_hip.h, integer_ciede_hip.c, integer_moment_hip.c, float_ssim_cuda.c) are fork-local additions with no upstream analogue. If upstream ever adds a file with the same name to core/src/feature/hip/ or core/src/feature/cuda/, a sync-upstream cherry-pick will restore it; the deletion here does not create a rebase conflict because the upstream tree never had these paths. The core/src/hip/meson.build edit is entirely fork-local. core/src/feature/hip/AGENTS.md and core/src/feature/cuda/AGENTS.md are fork-local files.
chore/hip-extractor-audit-verify-9 (ADR-0563)¶
No rebase impact. This PR is documentation and audit closure only. All changed files are fork-local:
docs/adr/0563-hip-extractor-audit-verification.md(new ADR, fork-local).docs/research/0563-hip-extractor-audit-verification.md(new research digest, fork-local).docs/state.md(fork-local tracking ledger).docs/adr/README.md(fork-local ADR index).docs/rebase-notes.md(this file, fork-local).changelog.d/changed/0551-hip-extractor-audit-close.md(fork-local fragment).
No upstream Netflix/vmaf file is touched. No libvmaf/ source is touched. No ffmpeg-patches/ file is touched. No meson_options.txt key is added. No new rebase-sensitive invariant is introduced.
fix/dev-container-sycl-hip-runtime (ADR-0543)¶
No rebase impact. All changes are confined to:
dev/Containerfile(fork-local; whole file is fork-added).dev/scripts/dev-mcp-entrypoint.sh(fork-local).dev/AGENTS.md(fork-local invariant note).docs/adr/0541-dev-container-sycl-hip-runtime-fix.md,docs/state.md,docs/development/dev-mcp.md,changelog.d/fixed/0541-dev-container-sycl-hip-runtime.md,docs/adr/README.md(fork-only doc tree).
No upstream-shared paths touched. The container's pinned NEO_VER / IGC_VER / GMMLIB_VER / ROCM_VER ARGs become a recurring maintenance item: when a future host kernel revs the i915 / xe / KFD UAPI, bump the relevant ARG. The dev-mcp-entrypoint.sh visibility probe surfaces such regressions in ≤ 30 s of container start so future bumps are easy to identify.
fix/dev-container-full-gpu-plumbing (ADR-0542)¶
No upstream-mirror paths touched. Modifies:
dev/Containerfile(stage 1 apt list: +intel-media-va-driver-non-free,mesa-va-drivers; revised Vulkan ICD selection comment block).dev/docker-compose.yml(common-env: +HSA_OVERRIDE_GFX_VERSION,HSA_ENABLE_SDMA, +ROCR_VISIBLE_DEVICES; expandedNVIDIA_DRIVER_CAPABILITIESdocumentation comment).dev/scripts/dev-mcp-entrypoint.sh(entrypoint-timeVK_DRIVER_FILESrewrite to exclude lavapipe whenever any real ICD is present).dev/AGENTS.md(new GPU-plumbing invariant section).docs/development/dev-mcp.md(backend matrix + env-var contract + HSA override documentation).docs/adr/0541-…md(+ index row),changelog.d/fixed/0541-…md,docs/state.md(one recently-closed row), this file.
Rebase sensitivity (none — container infra fork-local additive plus documentation): Every touched file lives under dev/, docs/, or changelog.d/. No libvmaf C source, no public header, no meson_options.txt, no ffmpeg-patches/ entry. The CLAUDE.md §12 r14 patch-stack rule does not apply (no libvmaf surface touched). Netflix upstream has no container infra under dev/ to conflict with.
fix/integer-vif-cuda-chroma-plane (ADR-0547)¶
Touches upstream-mirror path. Modifies:
core/src/feature/cuda/integer_vif_cuda.c(upstream-mirror — comment option-help-text clarifications plus a one-shot warn-on-true block for the vestigialenable_chromaoption; no kernel changes, no behaviour changes for any caller that doesn't setenable_chroma=true).
Why fork-local. Upstream Netflix/vmaf's CUDA VIF (verified at Netflix/vmaf@32780bd9b6:core/src/feature/cuda/integer_vif_cuda.c) neither carries the enable_chroma option nor has an n_planes field — it hardcodes data[0]-only access. The option was added by the fork-local PR #949 and the abandoned PR #948 attempted to mirror it on CPU; only the CUDA option landed and it was always a no-op.
Sync rule. If upstream ever adds genuine multi-plane VIF (would be a significant departure from the Sheikh & Bovik 2006 definition), revisit this clarification:
- Drop the
vmaf_log VMAF_LOG_LEVEL_WARNINGblock frominit_fex_cuda. - Restore
s->enable_chroma = false;to the active-clamp form OR plumbenable_chromainto the dispatch loop (depending on upstream's shape). - Update
docs/metrics/vif.mdto advertise the per-chroma-plane features that newly exist. - Move the
docs/state.mdrow from "Confirmed not-affected" to a normal closed-bug row.
Until then, sync conflicts on this file should keep both: the fork-local warn-on-true block in init_fex_cuda (search for the ADR-0541 comment anchor) and any incoming upstream changes to the neighbouring kernel-load paths.
Test reference. core/test/test_integer_vif_cpu_cuda_parity.c (suite fast/gpu) is the regression gate; it must continue to pass after any sync.
feat/hip-float-vif-score-kernel-real (ADR-0539)¶
No rebase impact. Touches:
core/src/feature/hip/hip_hsaco_stubs.c— fork-local TU; removes oneVMAF_HSACO_WEAK_STUB(float_vif_score_hsaco)line. Upstream Netflix/vmaf has no HIP backend so no conflict possible.docs/adr/0539-hip-float-vif-stub-removal.md,docs/adr/README.md,docs/state.md,docs/backends/hip/overview.md,core/src/feature/hip/AGENTS.md,changelog.d/fixed/hip-float-vif-stub-removal.md— fork-local docs.
Rebase invariant: the moment another .hip kernel under feature/hip/<extractor>/ becomes standalone-buildable, the same one-line removal must happen in hip_hsaco_stubs.c for its symbol. The AGENTS.md note added by this PR captures the pattern.
fix/hip-integer-vif-kernel-crash (ADR-0538)¶
Touches upstream-mirror paths. Modifies:
core/src/feature/hip/integer_vif/vif_statistics.hip(fork-local — HIP backend addition; no upstream conflict expected).core/src/feature/hip/integer_vif_hip.c(fork-local — added by ADR-0379 / PR #...).core/src/meson.build(upstream-shared — adds entries to the fork-localhip_kernel_sourcesdict, which itself is inside the fork-localif is_hipcc_enabled and is_hip_enabledblock; conflict risk only if upstream lands a totally different HIP build pipeline, which is implausible).core/src/hip/meson.build(fork-local).core/src/feature/hip/AGENTS.md(fork-local).core/src/feature/hip/hip_hsaco_stubs.c(NEW — fork-local).
No verbatim upstream code paths altered. Rebase invariant: if upstream ever adds an integer_vif HIP port, drop the fork's integer_vif/vif_statistics.hip and integer_vif_hip.c and re-evaluate whether the four ADR-0538 defects exist in their port too — three of the four are subtle (filter-half-width parsing, missing rd-write, host-pointer kernel arg) and an upstream re-implementation may well have the same blind spots.
fix/per-shot-bitrate-and-last-shot-chart (ADR-0531)¶
No rebase impact. All changes are confined to tools/vmaf-tune/src/vmaftune/per_shot.py, tools/vmaf-tune/src/vmaftune/cli.py, tools/vmaf-tune/src/vmaftune/report.py, tools/vmaf-tune/tests/test_per_shot.py, tools/vmaf-tune/tests/test_report.py, docs/adr/0531-*.md, docs/adr/README.md, docs/state.md, docs/rebase-notes.md, and changelog.d/fixed/per-shot-bitrate-and-last-shot-chart.md. The tools/vmaf-tune/ tree does not exist in upstream Netflix/vmaf. No conflict risk on sync.
fix/per-shot-segments-readonly-cwd (ADR-0532)¶
fix/per-shot-segments-readonly-cwd (ADR-0532)¶
No rebase impact. All changes are confined to tools/vmaf-tune/src/vmaftune/cli.py, tools/vmaf-tune/tests/test_per_shot.py, docs/usage/vmaf-tune.md, docs/adr/0530-*.md, docs/adr/README.md, docs/state.md, docs/rebase-notes.md, and changelog.d/fixed/0532-per-shot-segments-readonly-cwd.md. The tools/vmaf-tune/ tree does not exist in upstream Netflix/vmaf. No conflict risk on sync.
changelog.d/fixed/0530-per-shot-segments-readonly-cwd.md. The tools/vmaf-tune/ tree does not exist in upstream Netflix/vmaf. No conflict risk on sync.
fix/dev-container-dri-bind (ADR-0528)¶
No rebase impact. The only changed files are dev/docker-compose.yml, dev/AGENTS.md, docs/development/dev-mcp.md, docs/adr/0528-*.md, docs/adr/README.md, docs/rebase-notes.md, docs/state.md, and changelog.d/fixed/dev-container-dri-bind.md. None of these paths exist in upstream Netflix/vmaf. No conflict risk on sync.
fix/compare-rate-quality-chart-from-bisect-samples (ADR-0534)¶
No rebase impact. All changes are confined to fork-local files:
tools/vmaf-tune/src/vmaftune/bisect.py— addedBisectSampledataclass +BisectResult.samplesfield; bisect loop appends a sample per successful probe;to_recommend_resultprojects samples into theRecommendResult.bisect_samplestuple.tools/vmaf-tune/src/vmaftune/compare.py—RecommendResultgained optionalbisect_samplesfield;to_rowemitsbisect_samplesonly when populated (additive v2 schema change); CSV writer pinned toextrasaction="ignore".tools/vmaf-tune/src/vmaftune/report.py—BisectSamplePointadded;CodecSweepPointgained optionalbisect_samples;_sweep_plot_fnrewrites the chart to render from samples when available (legacy connect-the-dots path retained with caveat note when samples absent).tools/vmaf-tune/src/vmaftune/cli.py—--target-vmafsdefault flipped to75,80,85,90,93; both--target-vmafand--target-vmafswrapped with_TrackedDefaultActionso the v1 single-target back-compat path activates only when--target-vmaf NNis explicit and--target-vmafsis at its default;_sweep_point_from_jsonparses the new field.
None of these paths exist in upstream Netflix/vmaf (the entire tools/vmaf-tune/ tree is fork-local). No conflict risk on sync.
ffmpeg-patch stack: no impact (this PR doesn't touch any libvmaf C-API, public header, or meson_options.txt entry).
tooling/adr-atomic-allocator (ADR-0535)¶
No rebase impact. All changes are confined to scripts/adr/next-free.sh, scripts/adr/test-next-free.sh, docs/adr/0535-adr-atomic-allocator.md, docs/adr/README.md, docs/adr/0000-template.md, docs/state.md, docs/rebase-notes.md, docs/development/adr-workflow.md, changelog.d/added/0535-adr-atomic-allocator.md, CLAUDE.md, and AGENTS.md. None of these paths exist in upstream Netflix/vmaf. No conflict risk on sync.
fix/premium-vmaf-target-defaults (ADR-0538)¶
No rebase impact. All changes are confined to fork-local files:
tools/vmaf-tune/src/vmaftune/cli.py— flip the--target-vmafsdefault from75,80,85,90,93(ADR-0534) to94,96,97,98and update the help text + supersession note.tools/vmaf-tune/src/vmaftune/bisect.py— add_ABSOLUTE_CRF_RANGE_BY_NAME+_absolute_crf_range(adapter); default the bisect search window to that absolute range; bypassadapter.validate's CRF gate inside_encode_and_scorein favour of an explicit absolute-range check.tools/vmaf-tune/tests/test_bisect.py— updatetest_crf_range_defaults_to_adapter_quality_range->test_crf_range_defaults_to_encoder_absolute_range; add three regression tests pinning libx264 / libx265 / libsvtav1 premium-archival targets atok=Truewith achieved VMAF within 0.5 of target.tools/vmaf-tune/tests/test_compare_rate_quality_sweep.py— update the default-target-vmafs assertion to94,96,97,98.tools/vmaf-tune/AGENTS.md— rewrite the--target-vmafsdefault rebase-sensitive-invariant note; add a new invariant for the bisect's encoder-absolute-range default.docs/usage/vmaf-tune.md— supersede the rate-quality-sweep section's defaults / rationale; add the High-VMAF bisect contract subsection with the per-codec absolute-range table.docs/adr/0538-premium-vmaf-target-defaults-and-bisect.md,docs/adr/0534-...md(status flip),docs/adr/README.md,docs/research/0537-...md,docs/state.md,docs/rebase-notes.md,changelog.d/fixed/0538-premium-vmaf-target-defaults-and-bisect.md.
None of these paths exist in upstream Netflix/vmaf (the entire tools/vmaf-tune/ tree is fork-local). No conflict risk on sync.
ffmpeg-patch stack: no impact (this PR does not touch any libvmaf C-API, public header, or meson_options.txt entry).
feat/bvi-dvc-pre-extracted-input (ADR-0527)¶
No rebase impact. All changes are confined to ai/scripts/bvi_dvc_to_full_features.py, ai/tests/test_bvi_dvc_dir_mode.py, and doc / AGENTS.md / changelog files. None of these paths exist in upstream Netflix/vmaf. No conflict risk on sync.
fix/hip-motion-extractor-register (ADR-0523)¶
No rebase impact. The only changed file is core/src/feature/feature_extractor.c, which is a fork-local file (the HIP and Metal extractor blocks it contains have no upstream equivalent). Upstream Netflix/vmaf does not ship integer_motion_hip and the #if HAVE_HIP block does not exist in upstream. No conflict risk on sync.
fix/dnn-symbolic-batch-dim (ADR-0524)¶
Rebase-sensitive — core/src/libvmaf.c carries the fork-local tiny-AI loader path (vmaf_ctx_dnn_attach and the two helpers dnn_attach_nchw / dnn_attach_feature_vector); Netflix upstream does not ship a tiny-model surface. Changes:
dnn_attach_nchwacceptsin_shape[0] ∈ {1, -1}(symbolic batch folded to 1). Thein_shape[1] != 1(channels) reject is now separated from the batch check so each surface has its own diagnostic. The H/W reject message was sharpened to call out symbolic dims explicitly.dnn_attach_feature_vectorgained the same batch policy before the feature-width check; the optional rank-2 second-input shape probe (extra_shape) follows the same rule.- Per-frame inference (
vmaf_ctx_dnn_run_frame_nchwand the feature-vector run path) is unchanged — both already emitshape[0] = 1on the ORT Run call, so symbolic batch is purely a load-time concern.
core/src/dnn/AGENTS.md gained an "Invariant — symbolic batch dim acceptance (ADR-0524)" section. Reverting the batch acceptance breaks every shipped NR tiny model (model/tiny/nr_metric_v1*.onnx) plus any future trainer using the PyTorch dynamic_axes default.
ffmpeg-patch stack: no impact. The tiny-AI loader sits behind vmaf_use_tiny_model, which the in-tree FFmpeg patches do not touch.
Test fixture: model/tiny/smoke_v0_symbolic_batch.onnx is a fork-local 166-byte Identity graph with dim_param='batch' on dim 0. The fixture has no sidecar (loader handles -ENOENT gracefully) and is not listed in model/tiny/registry.json (which catalogues shipped models, not test fixtures).
fix/cli-no-reference-wire (ADR-0520)¶
Rebase-sensitive — core/tools/cli_parse.c + core/tools/vmaf.c are upstream-shared paths. Changes:
cli_parse.c: the reference-required gate at the end ofcli_parse()is now conditional on!settings->no_reference; the new branch requirestiny_model_pathand force-enablesno_prediction. If upstream Netflix reintroduces an unconditionalif (!settings->path_ref)(the original shape pre-PR), restore the guard. Theno_referencefield has been in tree since the tiny-AI surface landed, so the merge conflict is a literal hunk replace.vmaf.c: in themain()body thefile_ref = fopen(c.path_ref, ...)call now opensc.path_distwhenc.no_referenceis true. If an upstream sync collapses the open into a helper, propagate the conditional.open_input_videosalso gained ano_reference-aware error message (usesc->path_distwhen ref is being faked).core/tools/AGENTS.md: new ADR-0519 entry under "Governing ADRs" documents the CLI gate invariant + frame-loop invariant. Keep the entry through future merges.
core/src/libvmaf.c is not touched; the public API (vmaf_read_pictures, vmaf_use_tiny_model) is unchanged. The rank-4 DNN dispatch in vmaf_ctx_dnn_run_frame_nchw is upstream- internal and consumes ref argument bytes without caring about the slot semantics — the CLI's open-twice strategy works precisely because that dispatch is slot-agnostic. If an upstream refactor changes the dispatch to consult both ref and dist (e.g. for FR-only dual-input models), the fork-side wiring needs to either pass dist explicitly or expose a public vmaf_read_pictures_nr API.
ffmpeg-patch stack: no impact. The fork's FFmpeg filter does not surface NR-mode wiring today.
Netflix upstream does not ship --no-reference; the flag is a fork-local addition.
fix/msvc-unistd-gating (ADR-0521)¶
Rebase sensitivity: low — targeted portability guards on upstream-shared files.
Two files touched: core/src/feature/x86/vif_avx512.c and core/tools/yuv_input.c.
vif_avx512.c is a fork-local AVX-512 TU (no Netflix/vmaf upstream equivalent). The VMAF_NOINLINE_NOCLONE macro is added at the TU level and does not affect public headers or the ABI.
yuv_input.c has an upstream counterpart in Netflix/vmaf. The _WIN32 shims (fstat → _fstat64, S_ISREG, off_t) are added inside the existing #ifdef _WIN32 block, immediately after the already-present _fileno alias. On upstream sync: check whether Netflix has independently added MSVC portability to yuv_input.c; if so, prefer their solution and drop the fork-local block. The change is a four-line addition inside an existing guarded block — low merge conflict risk.
No ffmpeg-patches file touches either file. No public API change.
fix/per-shot-scene-threshold-and-1-shot-chart (ADR-0513)¶
No rebase impact. Changes confined to fork-local trees: tools/vmaf-tune/src/vmaftune/per_shot.py (new split_long_shots helper + diff_threshold / framerate / max_shot_duration_sec kwargs on detect_shots), tools/vmaf-tune/src/vmaftune/cli.py (--scene-threshold + --max-shot-duration flags on tune-per-shot), tools/vmaf-tune/src/vmaftune/report.py (_shot_plot_fn uses ax.hlines bands instead of a step plot), tools/vmaf-tune/tests/test_per_shot.py + test_report.py (6 new regression tests). The C-side core/tools/vmaf_per_shot.c is untouched — the new --scene-threshold flag passes through to the existing --diff-threshold C option that has been in tree since ADR-0222. Docs: ADR-0513, docs/adr/README.md index row, docs/usage/vmaf-tune.md flag rows + "Tuning scene sensitivity" section, docs/state.md Recently-closed rows, changelog.d/fixed/per-shot-scene-threshold-and-1-shot-chart.md. Netflix upstream does not ship tools/vmaf-tune/.
feat/compare-rate-quality-sweep — ADR-0516¶
No rebase impact. Changes confined to fork-local files: tools/vmaf-tune/src/vmaftune/compare.py (new compare_codecs_sweep, SweepReport, probe_encoder_available, detect_schema_version, v2 emitters, DEFAULT_CPU_ENCODERS, HARDWARE_ENCODERS, SCHEMA_VERSION_V1, SCHEMA_VERSION_V2), tools/vmaf-tune/src/vmaftune/cli.py (the _run_compare runner gains --target-vmafs parsing + sweep dispatch, the _run_report runner ingests v2 JSON into CodecSweepPoint), tools/vmaf-tune/src/vmaftune/report.py (new CodecSweepPoint, compute_pareto_frontier, _sweep_plot_fn per-codec line chart, v2 summary table renderers in both markdown + HTML), tools/vmaf-tune/tests/test_compare_rate_quality_sweep.py (new file, 24 regression tests), tools/vmaf-tune/AGENTS.md (v1 vs v2 schema invariant note + per-target bisect predicate construction rule), docs/usage/vmaf-tune.md (multi-target sweep section + flag table update + schema migration note), docs/adr/0516-vmaf-tune-compare-rate-quality-sweep.md (new), docs/adr/README.md (index row), docs/state.md (Recently closed row), changelog.d/added/compare-rate-quality-sweep.md (new fragment). Netflix upstream does not ship tools/vmaf-tune/; no upstream-shared C sources, public headers, Meson options, or ffmpeg-patches/ patches are touched.
fix/compare-source-is-container-plumbing (ADR-0509)¶
No rebase impact. Changes confined to tools/vmaf-tune/ (fork-local package) — src/vmaftune/cli.py (the _run_compare runner, the new _TrackedDefaultAction argparse action, _stamp_tracked_default_sentinels, and the _resolve_compare_source_geometry helper) and tests/test_compare.py (7 new regression tests). Netflix upstream does not ship tools/vmaf-tune/; no upstream-shared C sources, public headers, or build files are modified. The ADR (0509) and changelog fragment are fork-local docs only.
fix/chug-extract-vmaf-alignment — ADR-0510¶
No rebase impact. Changes confined to fork-local files: ai/scripts/extract_k150k_features.py, ai/scripts/chug_extract_features.py, ai/tests/test_extract_k150k_features.py, ai/tests/test_chug.py, ai/tests/test_chug_extract_features_smoke.py (new), docs/adr/0510-chug-extract-vmaf-alignment-fr-from-nr-guard.md (new), docs/adr/README.md (index row), docs/rebase-notes.md (this entry), docs/state.md (Recently closed row), ai/AGENTS.md (K150K-A invariant update), changelog.d/fixed/0509-*.md (new). The entire ai/ package and the FR-from-NR adapter pattern are fork-local — Netflix upstream has no CHUG ingestion, no K150K-A extractor, and no FR-from-NR adapter. No upstream-shared code, headers, build files, public C-API, or feature extractors are modified; the libvmaf CLI and all backends are unchanged.
fix/vulkan-two-variant-vif-shader (ADR-0512, supersedes ADR-0492)¶
No rebase impact on Netflix upstream — the Vulkan backend and its GLSL compute shaders are entirely fork-local (the core/src/vulkan/ and core/src/feature/vulkan/ trees do not exist in upstream). Fork-internal rebase invariants:
core/src/feature/vulkan/shaders/vif.compwas renamed intovif_fp64.comp+ new siblingvif_fp32.comp(the original file is removed). Any future patch series that targetsvif.compby name must be retargeted onto both variants — kernel changes touch BOTH in lockstep (seecore/src/feature/vulkan/AGENTS.md).VmafVulkanContextgained anint has_float64field (core/src/vulkan/vulkan_internal.h). Wire-compatible: feature TUs read it via the internal header, not the public ABI.VmafVulkanConfigurationgained a publicint require_fp64field (core/include/libvmaf/libvmaf_vulkan.h). Append-only ABI extension — existing zero-initialised callers get the auto-fallback default.- New internal entry point
vmaf_vulkan_context_new_with_opts(out, device_index, require_fp64); the originalvmaf_vulkan_context_newis preserved as a wrapper that passesrequire_fp64 = 0. - New CLI flag
--vulkan-require-fp64(and underscore alias--vulkan_require_fp64); the usage string was split across twofprintfcalls to stay under the C99 4095-char string-literal limit.
fix/dev-container-backend-exposure (ADR-0514)¶
No rebase impact. dev/Containerfile, dev/docker-compose.yml, and dev/AGENTS.md are entirely fork-local — upstream Netflix/vmaf does not ship the vmaf-dev-mcp container stack. If upstream ever ships its own dev container, merge by adopting upstream's image discipline and re-applying the four invariants documented in dev/AGENTS.md (tcm/latest/lib on LD_LIBRARY_PATH, no VK_ICD_FILENAMES pin, /dev/dri/by-path bind-mount, build-time backend probe).
fix/mcp-run-benchmark-repair — no rebase impact¶
All changed files (mcp-server/, testdata/bench_all.sh, docs/adr/0513-*, docs/mcp/, changelog.d/, docs/state.md) are fork-local. bench_all.sh is a fork-local benchmarking helper not present in Netflix/vmaf upstream. No rebase action required on upstream sync.
fix/restore-cuda-kernel-lifecycle-helpers¶
No rebase impact. Investigation confirmed VmafCudaKernelLifecycle, VmafCudaKernelReadback, and helper functions are intact in core/src/cuda/kernel_template.h (fork-local, ADR-0246). Changes confined to docs/state.md and changelog.d/ -- both fork-local, not present in Netflix upstream.
refactor/aiutils-vmaftune-corpus-dedup — no rebase impact¶
tools/vmaf-tune/ is fork-local. ai/src/aiutils/ is fork-local. No upstream-shared files are touched; no rebase action required.
fix/saliency-per-mb-eval-2026-05-15 — integer_vif enable_chroma¶
refactor/gpu-dispatch-env-pthread-once (ADR-0488)¶
No rebase impact: adds core/src/gpu_dispatch_env.{h,c} (new fork-local files) and modifies cuda/dispatch_strategy.c, vulkan/dispatch_strategy.c, sycl/dispatch_strategy.cpp — all fork-local TUs with no Netflix upstream equivalents. If upstream Netflix ever introduces their own dispatch env handling, merge by adopting their approach and dropping this helper.
No rebase impact: doc-only change. docs/state.md is fork-local and not present in Netflix upstream; upstream syncs do not touch it.
test/output-public-api-coverage-2026-05-16¶
No rebase impact. All changes are confined to core/test/test_output.c and changelog.d/. The test file is fork-local; upstream Netflix/vmaf does not ship test_output.c. No upstream-shared C sources, public headers, or build files are modified.
fix/sycl-motion-fps-weight-vulkan-import-status-2026-05-16¶
Sub-task B -- integer_motion_v2_sycl.cpp: adds motion_fps_weight to MotionV2StateSycl struct and options_motion_v2_sycl[]. If upstream Netflix ever adds motion_fps_weight to integer_motion_v2.c (the CPU reference), both the SYCL and CUDA motion_v2 twins should pick it up in the same PR per the invariant added to core/src/feature/sycl/AGENTS.md.
Sub-task A -- libvmaf_vulkan.h: removes stale -ENOSYS until T7-29 part 2 lands from the @return lines of vmaf_vulkan_import_image, vmaf_vulkan_wait_compute, and vmaf_vulkan_read_imported_pictures. No upstream rebase conflict expected -- the public Vulkan header is fork-local.
No rebase impact: fix/dev-mcp-stage3-and-bundled-fixes-2026-05-16 touches only dev/Containerfile, dev/AGENTS.md, docs/research/0135-*, and changelog.d/fixed/dev-mcp-container-stage-3.md. These are all fork-local infra files; no upstream-shared code, headers, build files, or feature extractors are modified. No sync-upstream conflicts expected.
No rebase impact: audit/t3-9b-ssimulacra2-ulp-audit — doc-only PR (ADR-0467, changelog fragment, BACKLOG update). No C files touched. No upstream-shared paths modified.
No rebase impact: feat/tiny-ai-registry-ci-and-saliency-v2-promotion-2026-05-15 touches model/tiny/registry.json (fork-local tiny-AI registry), docs/ai/models/ (fork-local model cards), docs/adr/0444-* (fork-local ADR), and the registry-validate CI job (fork-local CI). No upstream-shared code, headers, build files, or feature extractors are modified; the saliency model change is registry and docs only — the C-side mobilesal extractor is unaffected. Sync-upstream conflicts in this area are not expected.
No rebase impact: fix/mcp-embedded-docs-live-2026-05-14 updates fork-local MCP documentation and tools/vmaf-tune auto-planner code only; it does not touch upstream-shared code, headers, build files, or rebase-sensitive invariants.
The intended reader is whoever runs the next /sync-upstream (see ADR-0002 and .claude/skills/sync-upstream/). Read top-to-bottom before resolving conflicts.
Format¶
Each entry is a ### NNNN — short title heading with three fields:
- Touches: paths likely to conflict on upstream merge.
- Invariant: what the fork relies on that an upstream change could silently drop.
- Re-test: the command(s) to run after the merge to confirm the invariant survived. Reproducer-style — no surrounding prose required.
IDs are assigned in commit order and never reused. A single entry may cover several PRs in one workstream; cross-link from the ID heading.
Entries (backfilled 2026-04-18 per ADR-0108 adoption)¶
perf/vif-cpu-workspace-hoist-2026-05-16 — VifState scratch buffer hoist (ADR-0452)¶
- Touches:
core/src/feature/vif.c,core/src/feature/float_vif.c,core/src/feature/vif.h. - Invariant:
VifStategains afloat *vif_buffield (VIF_SCRATCH_BUF_CNT × scaled_float_stride × scaled_hbytes, allocated ininit, freed inclose).compute_vif's signature gains a trailingfloat *data_bufparameter — callers must pass a buffer of at least10 × ALIGN_CEIL(w * sizeof(float)) × hbytes. If an upstream Netflix commit modifiescompute_vif's signature or adds fields to the implicit scratch layout, the fork's extra parameter must be reconciled with the upstream change. The fork does NOT carry the upstream per-frame allocation; if upstream adds a new scratch sub-plane, extendVIF_SCRATCH_BUF_CNTand theVifState::vif_bufallocation size in the same PR. - Re-test:
```shell ninja -C build meson test -C build 2>&1 | grep -E "Ok|Fail" # Confirm 0 failures
perf/cambi-sycl-event-chain-2026-05-16 — CAMBI SYCL GPU-to-GPU event chains (SY-1)¶
- Touches:
core/src/feature/sycl/integer_cambi_sycl.cpp,core/src/feature/sycl/AGENTS.md,docs/adr/0471-cambi-sycl-event-chain.md. - Invariant:
launch_spatial_mask,launch_decimate, andlaunch_filter_modenow returnsycl::eventand accept asycl::event depparameter (exceptlaunch_spatial_mask, which has no predecessor). If upstream Netflix ever rewrites the CAMBI SYCL port (unlikely — SYCL is fork-only), preserve the event-chain structure and ensure the two semantically-requiredq.wait()points (post-H2D and post-D2H) remain. The CUDA twin (integer_cambi_cuda.c) retains synchronous v1 posture and is not affected by this change. - Re-test:
meson test -C build --suite=fast cambi_sycl
python3 scripts/ci/cross_backend_parity_gate.py --feature cambi --places 4
fix/psnr-enable-chroma-gpu-parity-2026-05-16 — PSNR enable_chroma option GPU parity¶
- Touches:
core/src/feature/cuda/integer_psnr_cuda.c,core/src/feature/sycl/integer_psnr_sycl.cpp,core/src/feature/vulkan/psnr_vulkan.c,docs/metrics/features.md,docs/research/0135-*,docs/adr/0452-*,changelog.d/fixed/psnr-enable-chroma-cross-backend.md,docs/research/0136-psnr-enable-chroma-cross-backend-2026-05-16.md,docs/adr/0453-psnr-enable-chroma-gpu-parity.md. - Invariant: The
enable_chromaoption default istrueon all backends. Then_planesclamp in GPUinit()must stay in the following order: (1)pix_fmt == YUV400Psetsn_planes = 1; (2)!enable_chroma && n_planes > 1also clamps to 1. If upstream Netflix ever adds option-table support to the CUDA/Vulkan twins, port any new options but preserve theenable_chromaentry and itsdefault_val.b = trueexactly — a default flip tofalsewould silently suppress chroma output and break the cross-backend parity gate. - Re-test:
python3 scripts/ci/cross_backend_parity_gate.py \
--backends cpu cuda --features psnr --places 4
python3 scripts/ci/cross_backend_parity_gate.py \
--backends cpu cuda --features psnr --places 4 \
--feature-opts 'psnr=enable_chroma=false' \
--feature-opts 'psnr_cuda=enable_chroma=false'
fix/vmaf-tune-temporal-saliency-2026-05-15 — recommend-saliency temporal aggregation¶
- Touches:
tools/vmaf-tune/src/vmaftune/saliency.py,tools/vmaf-tune/src/vmaftune/cli.py,tools/vmaf-tune/tests/test_saliency.py, anddocs/usage/vmaf-tune.md. - Invariant:
meanremains the default compatibility reducer forrecommend-saliency --saliency-aggregator. Changing the default changes user-visible saliency ROI behaviour and needs an ADR-0396 follow-up plus usage-doc update. - Re-test:
fix/saliency-per-mb-eval-2026-05-15 — saliency per-block IoU evaluator¶
- Touches:
ai/scripts/eval_saliency_per_mb.py,ai/tests/test_eval_saliency_per_mb.py,ai/AGENTS.md,docs/ai/saliency-per-mb-eval.md,docs/ai/index.md,docs/ai/roadmap.md, andmkdocs.yml. - Invariant: video-saliency model promotion should be measured at the encoder ROI block grid, not only full-resolution pixel IoU. Keep the evaluator dependency-light (
numpyplus.npy/ PGM loaders) so training sandboxes can run it without Pillow or OpenCV. - Re-test:
fix/chug-hdr-audit-splits-2026-05-15 — CHUG HDR audit and content-safe splits¶
- Touches:
ai/scripts/chug_extract_features.py,ai/scripts/train_konvid_mos_head.py,ai/scripts/extract_k150k_features.py,ai/tests/test_chug.py,ai/tests/test_train_konvid_mos_head.py,ai/tests/test_extract_k150k_features.py,ai/AGENTS.md,docs/ai/chug-ingestion.md, anddocs/ai/datasets/k150k.md. - Invariant: CHUG train/validation/test partitions are keyed by
chug_content_name, not by individual bitrate-ladder rows. The materialiser writessplit,chug_split_key, andchug_split_policyinto every feature row. Preserve the--audit-outputffprobe HDR metadata audit as a pre-training guard.train_konvid_mos_head.pyconsumes explicit splits when available instead of silently re-shuffling CHUG rows. The FR-from-NR parquet extractor preserves CHUG side metadata when--metadata-jsonlis supplied. - Re-test:
PYTHONPATH=ai/src .venv/bin/python -m pytest ai/tests/test_chug.py ai/tests/test_train_konvid_mos_head.py ai/tests/test_extract_k150k_features.py -q
fix/tiny-ai-rgb-high-bitdepth-2026-05-15 — LPIPS / DISTS high-bit-depth input¶
- Touches:
core/src/dnn/tiny_extractor_template.h,core/src/feature/feature_lpips.c,core/src/feature/feature_dists.c,core/test/test_dists.c,core/src/dnn/AGENTS.md,core/src/feature/AGENTS.md,docs/ai/extractor-template.md,docs/ai/models/lpips_sq.md,docs/ai/models/dists_sq.md,docs/metrics/dists.md, anddocs/metrics/features.md. - Invariant: LPIPS and DISTS-Sq accept planar 8/10/12/16-bit YUV while keeping the ONNX tensor ABI unchanged: ImageNet-normalised RGB8, NCHW
[1,3,H,W], named inputsref/dist, scalar outputscore. High-bit-depth samples are little-endian 16-bit containers rounded into the 8-bit domain before the shared BT.709 limited-range RGB conversion. - Re-test:
meson test -C core/build-tiny-rgb-hbd test_dists test_lpips --print-errorlogs
fix/mcp-runtime-doc-status-2026-05-15 — embedded MCP runtime docs¶
- Touches:
docs/api/mcp.md,docs/development/build-flags.md,core/meson_options.txt, andlibvmaf/AGENTS.md. - Invariant: embedded MCP is no longer an all-entrypoint
-ENOSYSscaffold. Preserve the runtime contract when rebasing: stdio / UDS / loopback-SSE transports are live when their build flags are enabled;compute_vmafuses a per-call ephemeralVmafContext; mutating measurement-thread tools still wait on the future SPSC bridge;enable_mcpremains default-off until that bridge lands. - Re-test:
meson setup /tmp/vmaf-mcp-doc-check -Denable_mcp=true -Denable_mcp_stdio=true -Denable_mcp_uds=true -Denable_mcp_sse=enabled && ninja -C /tmp/vmaf-mcp-doc-check test_mcp_smoke && meson test -C /tmp/vmaf-mcp-doc-check test_mcp_smoke
fix/mcp-compute-vmaf-high-bitdepth-2026-05-15 — MCP compute_vmaf bitdepth¶
- Touches:
core/src/mcp/compute_vmaf.c,core/src/mcp/dispatcher.c,core/test/test_mcp_smoke.c,core/src/mcp/AGENTS.md,docs/api/mcp.md, anddocs/mcp/embedded.md. - Invariant: embedded MCP
compute_vmafaccepts YUV420p at 8/10/12/16 bpc and defaults to 8 whenbitdepthis omitted. High-bit-depth raw samples are little-endian 16-bit words read directly into libvmaf picture storage. Do not silently add YUV422P or YUV444P without extending the tool schema with an explicitpixel_formatargument and matching docs/tests. - Re-test:
meson test -C core/build-mcp-hbd test_mcp_smoke --print-errorlogs
fix/chug-cuda-feature-split-2026-05-15 — FR-from-NR CUDA feature split¶
- Touches:
ai/scripts/extract_k150k_features.py,ai/tests/test_extract_k150k_features.py,ai/AGENTS.md,docs/ai/datasets/k150k.md, anddocs/ai/chug-ingestion.md. - Invariant: CUDA mode in the FR-from-NR extractor uses explicit CUDA feature names for the stable CUDA pass and
--cpu-vmaf-binfor the residual CPU feature pass (float_ssim,cambi). Do not collapse this back into one generic all-feature--backend cudainvocation; local CHUG 10-bit clips reproduced duplicate feature-key writes and CUDA context synchronization failures on that path. - Re-test:
PYTHONPATH=ai/src .venv/bin/python -m pytest ai/tests/test_extract_k150k_features.py -q
fix/vmaf-tune-libvpx-adapter-2026-05-14 — vmaf-tune libvpx-vp9 adapter¶
- Touches:
tools/vmaf-tune/src/vmaftune/codec_adapters/__init__.py,tools/vmaf-tune/src/vmaftune/codec_adapters/libvpx.py,tools/vmaf-tune/src/vmaftune/encode.py, anddocs/usage/vmaf-tune*.md. - Invariant:
libvpx-vp9stays a normal codec-adapter registry entry. Do not add VP9 branches to corpus / encode search loops; the adapter owns-deadline good,-cpu-used,-crf,-b:v 0,-row-mt 1, and FFmpeg-native-pass/-passlogfilewiring.supports_encoder_statsremains false until a binary VP9 first-pass stats parser lands. - Re-test:
PYTHONPATH=tools/vmaf-tune/src .venv/bin/python -m pytest tools/vmaf-tune/tests/test_codec_adapter_libvpx.py tools/vmaf-tune/tests/test_encode_multi_codec.py -q
fix/ai-frame-loader-color-pixfmt-2026-05-14 — packed colour frame loader¶
- Touches:
ai/src/vmaf_train/data/frame_loader.py,ai/tests/test_frame_loader.py,docs/ai/training.md, andai/AGENTS.md. - Invariant: frame-loader support is limited to byte-contiguous formats with unambiguous tensor shape:
gray->HxW, andrgb24/bgr24/rgba/bgra->HxWxC. Planar or subsampled formats such asyuv420pmust keep failing before spawning ffmpeg until a PR adds explicit plane semantics. - Re-test:
PYTHONPATH=ai/src .venv/bin/python -m pytest ai/tests/test_frame_loader.py -q
fix/mkdocs-strict-pre-push-2026-05-15 — mkdocs strict-mode pre-push hook¶
- Touches:
scripts/git-hooks/pre-push-mkdocs-strict.sh(new),scripts/git-hooks/pre-push(delegation call appended),.pre-commit-config.yaml(newmkdocs-strictlocal hook entry),docs/adr/0466-mkdocs-strict-pre-push-hook.md(new ADR). - Invariant: The hook gate mirrors the CI
docs.ymllane (ADR-0403):mkdocs build --strict --quietwith the repo-rootmkdocs.yml. Keeping the hook's config-file flag pointed atmkdocs.ymlin the repo root is load-bearing — ifmkdocs.ymlis ever moved, updatepre-push-mkdocs-strict.shin the same PR. TheSKIP=mkdocs-strictbypass token is the per-hook escape hatch; preserve it across rebases so the CI-gate-mirror contract (which also respectsSKIP) stays coherent. - Re-test:
# Touch a docs file with a known-good anchor, push — hook should pass:
touch docs/index.md && git push
# Touch docs/index.md, add a broken anchor ref, push — hook should block:
echo "[bad](#nonexistent)" >> docs/index.md && git push
fix/dists-extractor-2026-05-14 — DISTS-Sq extractor smoke surface¶
- Touches:
core/src/feature/feature_extractor.c,core/src/feature/feature_dists.c,core/src/meson.build,core/test/meson.build,.gitattributes,model/tiny/registry.json, anddocs/metrics/dists.md. - Invariant:
dists_sqis a registered tiny-AI full-reference extractor that mirrors LPIPS' two-input ABI:model_pathoption,VMAF_DISTS_SQ_MODEL_PATHenvironment fallback, ONNX inputsref/dist, scalar outputscore, and emitted feature keydists_sq.model/tiny/dists_sq.onnxis a smoke placeholder markeddists_sq_placeholder_v0; do not present it as production DISTS weights. - Re-test:
meson test -C build-dists test_dists && .venv/bin/python ai/scripts/validate_model_registry.py model/tiny/registry.json
fix/backlog-gap-pass-10-2026-05-14 — KonViD-150k split score ingestion¶
- Touches:
ai/scripts/konvid_150k_to_corpus_jsonl.py,ai/tests/test_konvid_150k.py,docs/ai/konvid-150k-ingestion.md,ai/AGENTS.md. - Invariant:
konvid_150k_to_corpus_jsonl.pyaccepts both the URL-manifest layout (manifest.csv+clips/) and the staged split score layout (k150ka_scores.csv/k150kb_scores.csvplusk150ka_extracted//k150kb_extracted/). Explicit--manifest-csvstays strict and must not silently fall back. Output JSONL schema remains unchanged. - Re-test:
PYTHONPATH=ai/src .venv/bin/python -m pytest ai/tests/test_konvid_150k.py -q
fix/backlog-gap-pass-11-2026-05-14 — vmaf-tune auto source probe¶
- Touches:
tools/vmaf-tune/src/vmaftune/auto.py,tools/vmaf-tune/src/vmaftune/cli.py,tools/vmaf-tune/tests/test_auto_short_circuits.py,docs/usage/vmaf-tune.md, andtools/vmaf-tune/AGENTS.md. - Invariant:
run_auto(smoke=False, meta_override=None)is not a scaffold. It probes source geometry, duration, and HDR once through_probe_source_meta, using one subprocess runner seam for testability. Probe failure must degrade to conservative defaults rather than raising or reintroducingNotImplementedError. - Re-test:
PYTHONPATH=tools/vmaf-tune/src .venv/bin/python -m pytest \
tools/vmaf-tune/tests/test_auto_short_circuits.py \
tools/vmaf-tune/tests/test_auto_confidence_aware.py \
tools/vmaf-tune/tests/test_auto_recipe_overrides.py \
tools/vmaf-tune/tests/test_auto_phase_f1_f2.py -q
.venv/bin/python -m ruff check \
tools/vmaf-tune/src/vmaftune/auto.py \
tools/vmaf-tune/src/vmaftune/cli.py \
tools/vmaf-tune/tests/test_auto_short_circuits.py
fix/backlog-gap-pass-12-2026-05-14 — MCP docs + SSIMULACRA2 snapshot hardening¶
- Touches:
python/test/ssimulacra2_test.py,docs/mcp/index.md,docs/mcp/embedded.md,docs/mcp/release-channel.md,mcp-server/vmaf-mcp/README.md,mcp-server/AGENTS.md. - Invariant: the SSIMULACRA2 snapshot gate remains fork-local. It pins current extractor output for the 576x324 fixture with explicit x86_64 and arm64/aarch64 baselines, pins the shared 160x90 tail fixture, and must invoke the repo
vmafbinary with an argv list, not a shell string. The external MCP server docs list all seven live tools. The embedded MCP docs describe the v3 runtime accurately: stdio, UDS, and loopback SSE are live;list_featuresandcompute_vmafare live; the SPSC measurement-thread drain and mutating tools remain future work. - Re-test:
PYTHONPATH=python .venv/bin/python -m pytest python/test/ssimulacra2_test.py -q && PYTHONPATH=mcp-server/vmaf-mcp/src .venv/bin/python -m pytest mcp-server/vmaf-mcp/tests/test_server.py -q
fix/read-json-model-dynamic-limits-2026-05-14 — dynamic JSON model arrays¶
- Touches:
core/src/read_json_model.c,core/src/model.h,core/src/model.c, andcore/test/test_model.c. - Invariant: JSON model loading grows
VmafModel.featureandscore_transform.knots.listfrom the payload. Do not restore the old fixedMAX_FEATURE_COUNT/MAX_KNOT_COUNTparser caps; models with 65+ features or 11+ score-transform knots must parse when the JSON is otherwise valid. - Re-test:
meson test -C build test_model --print-errorlogs.
fix/real-scaffold-gap-pass-4-2026-05-14 — vmaf-tune x264 two-pass¶
- Touches:
tools/vmaf-tune/src/vmaftune/codec_adapters/x264.py,tools/vmaf-tune/src/vmaftune/encode.pyconsumers, anddocs/usage/vmaf-tune.md. - Invariant:
libx264opts into the shared Phase F two-pass seam throughsupports_two_pass = Trueandtwo_pass_args() -> ("-pass", N, "-passlogfile", path). The encode driver must stay adapter-driven; do not add an x264 branch inbuild_ffmpeg_command. - Re-test:
PYTHONPATH=tools/vmaf-tune/src .venv/bin/python -m pytest tools/vmaf-tune/tests/test_codec_adapter_x265_two_pass.py tools/vmaf-tune/tests/test_auto_phase_f1_f2.py -q
fix/backlog-gap-pass-8-2026-05-14 — CUDA psnr_hvs drain-batch integration¶
- Touches:
core/src/feature/cuda/integer_psnr_hvs_cuda.c,core/src/feature/cuda/AGENTS.md,docs/backends/cuda/overview.md,docs/development/cuda-profile-2026-05-03.md. - Invariant:
integer_psnr_hvs_cuda.cenqueues all three plane-partial DtoH copies ons->lc.strduring submit, callsvmaf_cuda_kernel_submit_post_record(&s->lc, fex->cu_state), and usesvmaf_cuda_kernel_collect_wait(&s->lc, fex->cu_state)in collect before readingh_partials[]. Do not move the readback + rawcuStreamSynchronize(s->lc.str)back into collect. - Re-test:
meson setup build-cuda-drain libvmaf -Denable_cuda=true -Denable_sycl=false -Denable_vulkan=disabled --buildtype=debug && ninja -C build-cuda-drain src/libvmaf.so.3.0.0 && python3 scripts/ci/cross_backend_vif_diff.py --vmaf-binary "$PWD/build-cuda-drain/tools/vmaf" --reference testdata/ref_576x324_48f.yuv --distorted testdata/dis_576x324_48f.yuv --width 576 --height 324 --feature psnr_hvs --backend cuda --places 3. If the local CUDA fatbin build is blocked by toolkit include-path drift, at minimum compile the touched host TU:ninja -C build-cuda-drain src/liblibvmaf_feature.a.p/feature_cuda_integer_psnr_hvs_cuda.c.o.
fix/tune-scaffold-gap-pass-2-2026-05-14 — vmaf-tune per-shot real bisect CLI¶
- Touches:
tools/vmaf-tune/src/vmaftune/cli.py,tools/vmaf-tune/src/vmaftune/per_shot.py,tools/vmaf-tune/tests/test_per_shot.py,docs/usage/vmaf-tune.md,docs/adr/0392-vmaf-tune-phase-d-per-shot.md,tools/vmaf-tune/AGENTS.md. - Invariant: the CLI default for
vmaf-tune tune-per-shotis the real Phase-B bisect backend. It extracts each detected half-open shot to temporary raw YUV, passes explicit geometry intobisect_target_vmaf, and emits measured per-shot VMAF in the JSON plan.--predicate-module MODULE:CALLABLEis the only CLI path that bypasses real bisect; the adapter-default predicate remains library-only dry-run behaviour. - Re-test:
PYTHONPATH=tools/vmaf-tune/src .venv/bin/python -m pytest tools/vmaf-tune/tests/test_per_shot.py -q.
fix/scaffold-gap-pass-2026-05-14b — vmaf-tune compare real bisect CLI¶
- Touches:
tools/vmaf-tune/src/vmaftune/cli.py,tools/vmaf-tune/src/vmaftune/compare.py,tools/vmaf-tune/tests/test_compare.py,docs/usage/vmaf-tune.md,docs/usage/vmaf-tune-bisect.md,tools/vmaf-tune/AGENTS.md. - Invariant: the CLI default for
vmaf-tune compareis the real Phase-B bisect backend when source geometry is supplied. The programmaticcompare_codecs()default may still returnok=Falsebecause it lacks geometry, but the CLI must not silently rank using a placeholder predicate.--predicate-module MODULE:CALLABLEis the explicit custom/test escape hatch. - Re-test:
PYTHONPATH=tools/vmaf-tune/src .venv/bin/python -m pytest tools/vmaf-tune/tests/test_compare.py -q.
fix/scaffold-gap-pass-2026-05-14 — vmaf-tune hardware predictor real weights¶
- Touches:
tools/vmaf-tune/src/vmaftune/predictor_train.py,model/predictor_{h264,hevc,av1}_{nvenc,qsv}.onnx, matching model cards,docs/ai/predictor.md,tools/vmaf-tune/AGENTS.md. - Invariant: the trainer accepts canonical Phase-A rows and historical hardware-sweep aliases. Do not add an external corpus-conversion script for
runs/phase_a/full_grid/comprehensive.jsonl; the loader is the compatibility seam. - Re-test:
PYTHONPATH=tools/vmaf-tune/src .venv/bin/python -m pytest tools/vmaf-tune/tests/test_predictor_train.py -q.
fix/vmaf-tune-ai-scaffold-state-cleanup — auto HDR dispatch + ensemble seed registry flip (2026-05-14)¶
- Touches:
tools/vmaf-tune/src/vmaftune/auto.py,tools/vmaf-tune/tests/test_auto_short_circuits.py,tools/vmaf-tune/tests/test_auto_recipe_overrides.py,model/tiny/registry.json,python/test/model_registry_schema_test.py,docs/usage/vmaf-tune.md,docs/state.md,docs/research/0100-vmaf-tune-ai-scaffold-audit-2026-05-14.md,changelog.d/fixed/vmaf-tune-ai-scaffold-state-cleanup.md. - Invariant:
vmaf-tune automust usevmaftune.hdr.hdr_codec_args(codec, info)per HDR cell; a single generic PQ tuple is not valid because x265/SVT-AV1/NVENC/VVenC carry HDR signalling through different ffmpeg flag families. Recipe-adjustedeffective_thresholdsfrom_apply_recipe_overridemust be the thresholds used for F.3 decisions and JSON metadata. The fivefr_regressor_v2_ensemble_v1_seed{0..4}registry rows are production entries (smoke: false) only while their sidecars carry matching SHA-256s and a passing PROMOTE gate. - Re-test on rebase:
PYTHONPATH=tools/vmaf-tune/src python -m pytest \
tools/vmaf-tune/tests/test_auto_short_circuits.py \
tools/vmaf-tune/tests/test_auto_recipe_overrides.py \
tools/vmaf-tune/tests/test_auto_confidence_aware.py \
tools/vmaf-tune/tests/test_hdr.py -v
PYTHONPATH=python python -m pytest python/test/model_registry_schema_test.py -v
bash core/test/dnn/test_registry.sh
feat/libvmaf-metal-filter-iosurface — Metal IOSurface zero-copy import (ADR-0423)¶
- Touches:
core/include/libvmaf/libvmaf_metal.h(newVmafMetalExternalHandles+ four entry points appended),core/src/metal/picture_import.mm(new TU implementing the IOSurfaceLock + memcpy ring),core/src/metal/state_priv.h(shared struct defs between common.mm and picture_import.mm),core/src/metal/import.h(internal bridge for libvmaf.c HAVE_METAL block),core/src/metal/common.mm(state-free hook for the import ring),core/src/libvmaf.c(HAVE_METAL block:vmaf_metal_import_state/vmaf_metal_read_imported_pictures),core/src/metal/meson.build(one-line TU registration),core/test/test_metal_smoke.c(input-validation + device-default skip semantics),ffmpeg-patches/0013-libvmaf-add-libvmaf-metal-filter.patch(new),ffmpeg-patches/series.txt,ffmpeg-patches/README.md. - Invariant: the import path is geometry-pinned to the first (w, h, bpc) tuple seen — subsequent imports with a different geometry return
-EINVAL. Ring depth is 2 slots (VMAF_METAL_IMPORT_RING); a slot is identified byindex % VMAF_METAL_IMPORT_RINGand discarded if the caller'sindexno longer matches the stored one. CPU memcpy path is synchronous sovmaf_metal_wait_computeis a no-op (returns 0); do not promote it to aMTLSharedEventdrain without first switching the import body to an asyncMTLCommandBuffersubmission. Apple-Family-7+ gate is enforced insidevmaf_metal_state_init_externalvia[device supportsFamily:MTLGPUFamilyApple7]→-ENODEVon non-Apple hosts; the ffmpeg patch surfaces this asAVERROR(ENODEV)atconfig_props_metaltime. Symbol names are load-bearing for thecheck_pkg_configprobe in patch 0013; do not rename without simultaneously updating the patch. - Re-test on rebase:
meson setup build libvmaf -Denable_metal=enabled \
-Denable_cuda=false -Denable_sycl=false
ninja -C build
nm build/libvmaf/libvmaf.dylib | grep vmaf_metal_picture_import
git -C ffmpeg-8 reset --hard n8.1.1
for p in ffmpeg-patches/000*-*.patch; do
git -C ffmpeg-8 am --3way "$p" || break
done
Upstream Netflix/vmaf has no Metal backend; no rebase conflict surface against upstream/master. The 0013 patch is fork-local.
fix/saliency-per-mb-eval-2026-05-15 (Batch 4) — Metal install + header fix (ADR-0437)¶
- Touches:
core/include/core/meson.build(addsis_metal_enabledguard +libvmaf_metal.htoplatform_specific_headers),core/test/meson.build(addstest_metal_install_headerunderhost_machine.system() == 'darwin'),core/test/test_metal_install_header.c(new compile+link smoke test),docs/api/gpu.md(Metal + HIP symbol corrections, IOSurface sub-API table),docs/adr/0437-*.md,docs/adr/_index_fragments/0437-*.md,changelog.d/fixed/metal-public-header-install-and-import-state.md,docs/state.md. - Invariant: The
is_metal_enabledguard mirrors the Vulkan guard (is_vulkan_enabled): both treatenabledandautoas "install the header". Do not change this to install only onenabled;autoon macOS resolves to a real Metal build and the header must be present for FFmpeg'scheck_pkg_configto succeed (same rationale as ADR-0192 for Vulkan). - No rebase conflict surface: upstream Netflix/vmaf has no Metal backend;
core/include/core/meson.builddiverges from upstream at the firstis_cuda_enabledline. The only conflict risk is a batch that also editsplatform_specific_headers— resolve by keeping both additions. - Re-test on rebase:
meson setup build -Denable_metal=enabled \
-Denable_cuda=false -Denable_sycl=false
ninja -C build
meson install -C build --destdir /tmp/vmaf-test-install
ls /tmp/vmaf-test-install/usr/local/include/libvmaf/libvmaf_metal.h
fix/metal-includes-and-ffmpeg-patch — Metal kernel batch T8-1c–k (ADR-0421)¶
- Touches:
core/src/feature/metal/*.metal(7 new kernel files),core/src/feature/metal/*_metal.mm(7 new dispatch files replacing*_metal.cscaffolds),core/src/metal/meson.build(.aircustom_targets + metallib pipeline),ffmpeg-patches/0012-*. - Invariant: no
atomic_ulongin any.metalfile — Apple MSL silently drops 64-bit atomic updates (CI run 25685703780). All kernels use per-WGfloat/uintpartials array indexed bybid.y * grid_groups.x + bid.x; host reduces indouble. Thefloat_moment_metal.mmcorrectsprovided_features(was wrong in the scaffold:float_moment1/2/std→ correct namesfloat_moment_ref1st/dis1st/ref2nd/dis2nd). - Re-test on rebase (macOS, Apple-Family-7+):
meson setup build libvmaf -Denable_metal=enabled \
-Denable_cuda=false -Denable_sycl=false
ninja -C build
meson test -C build test_metal_smoke
On Linux: same build without -Denable_metal (Metal subdir excluded); no Metal tests registered. No upstream rebase conflict surface.
feat/metal-runtime-t8-1b — Metal backend runtime PR (ADR-0420)¶
- Touches:
core/src/metal/common.{c→mm,h},core/src/metal/picture_metal.{c→mm},core/src/metal/kernel_template.{c→mm},core/src/metal/meson.build,core/test/test_metal_smoke.c,core/src/metal/AGENTS.md. - Invariant: the Metal backend ships three Objective-C++ TUs (
common.mm,picture_metal.mm,kernel_template.mm) instead of the T8-1 pure-C scaffold. Public ABI incore/include/libvmaf/libvmaf_metal.hunchanged. Internalmetal/common.hgained two accessor declarations (vmaf_metal_context_{device,queue}_handle) that consumer TUs call to retrieve bridge-retainedvoid *Metal handles. The Obj-C++ TUs compile with-fobjc-arcviaadd_project_arguments(language: 'objcpp'). ARC +__bridge_retained/__bridge_transfercasts manage the +1 retain that lives on each C-struct slot. Upstream Netflix/vmaf has no Metal backend; there is no rebase conflict surface againstupstream/master. - Re-test:
meson setup build -Denable_metal=enabledon a recent macOS host (Apple Silicon preferred) +meson compile -C build+meson test -C build test_metal_smoke. On Apple-7+ the smoke test exercises real-device paths; on Intel Macs it short-circuits cleanly on-ENODEV. Non-Darwin builds are unaffected —subdir('metal')is already gated to Darwin.
fix/sve2-probe-darwin-gate — SVE2 build probe gated to non-Darwin hosts (ADR-0419)¶
- Touches:
core/src/meson.build(the SVE2cc.compiles()probe block). - Invariant:
is_sve2_supported = falseis forced whenhost_machine.system() == 'darwin', mirroring the runtime__linux__gate incore/src/arm/cpu.c::vmaf_get_cpu_flags_arm(). Apple Silicon (M1–M4) is ARMv8.x without SVE2 hardware, so the build-time and runtime gates must stay in lockstep. - Re-test: if upstream Netflix ever introduces its own SVE2 probe in
core/src/meson.build, drop the fork-local Darwin short-circuit in favour of theirs if it matches thedarwin ⇒ falseinvariant; otherwise layer the Darwin guard on top. Reverse the gate only when (a) Apple ships an arm part with SVE2 — no public roadmap as of 2026-05 — and (b) the runtime probe inarm/cpu.cgrows a Darwin branch (e.g.sysctlbyname("hw.optional.arm.FEAT_SVE2", ...)).
fix/macos-test-recal-post-vif-sync — macOS Python test assertions recalibrated for post-bf9ad333 VMAF/ADM values (ADR-0418)¶
- Touches:
python/test/local_explainer_test.py,python/test/vmafexec_test.py,python/test/vmafexec_feature_extractor_test.py. - Invariant: 9+ assertions in those files were updated to the post-VIF-sync values that the macOS-libm binary actually produces, since Netflix upstream only shipped recalibration fixtures for the
test_run_vmaf_*tests via142c0671/7209110e/d93495f5/fe756c9fand not thelocal_explainer_test::test_explain_vmaf_results,vmafexec_test::test_run_vmafexec_runner_akiyo_*, or the 5×vmafexec_feature_extractor::test_run_float_adm_fextractor_adm_*cases. Each updated line carries an inline# post-VIF-sync (#758) recalcomment so the divergence is greppable. Affected values: local_explainer_test.py:103—76.68425574067017 → 76.66740228116836vmafexec_test.py:871—132.732952 → 132.732323vmafexec_test.py:926, 1032, 1086—88.030463 → 88.030322vmafexec_feature_extractor_test.py:1834—0.9420788125 → 0.9185737499999999vmafexec_feature_extractor_test.py:1897—0.9517253541666667 → 0.8902739375vmafexec_feature_extractor_test.py:1960—0.9554477708333334 → 0.8780868749999998vmafexec_feature_extractor_test.py:2023—0.9662835416666665 → 0.8407157499999999vmafexec_feature_extractor_test.py:3030—0.96851 → 0.962086- Re-test: after the next
/sync-upstream, if upstream has shipped recalibrated fixtures for any of the listed test names, prefer upstream values over the fork-recalibrated ones in this entry. Mechanical:git grep "post-VIF-sync (#758) recal"enumerates every divergence; for each row, diff against the upstream value at the same line. If upstream still hasn't shipped the fixtures, leave the fork values in place — they're verified against the on-the-fly-VIF binary on the macOS-libm precision floor.
fix/vif-upstream-onthefly-filter-sync — VIF synced to Netflix upstream bf9ad333 + 8c645ce3¶
- Touches:
core/src/feature/vif.c,vif.h,vif_tools.c,vif_tools.h,vif_options.h,float_vif.c;python/test/quality_runner_test.py,feature_extractor_test.py,result_test.py,vmafexec_test.py. - Invariant: fork's VIF C-side now matches upstream HEAD verbatim for the listed files. The only fork-local divergence is
float_vif.c::extract()passings->vif_skip_scale0 ? 1 : 0for the newcompute_vif()parameter (instead of the upstream pattern of reading from a flag set ininit). Test cherry-picks took upstream values for VIF score assertions and fork values for VMAF_legacy_score / VMAF_score where the fork's binary diverges from upstream's at places=4 (already pre-loosened). - Re-test: after the next
/sync-upstream, runmeson test -C build(must remain 54/54 OK) andPYTHONPATH=python python -m pytest python/test/feature_extractor_test.py python/test/quality_runner_test.py -q(must show 0 failures excludingniqe_runnerskimage env issue). If upstream reverts or further modifies on-the-fly filter generation, this entry's invariant should re-sync rather than carry a fork-local divergence.
fix/master-build-failures-sycl-vulkan — SYCL macro collision + Vulkan SDK fallback + Cambi FR atom rename¶
- Touches:
core/src/feature/sycl/integer_adm_sycl.cpp,core/src/vulkan/common.c,python/vmaf/core/cambi_feature_extractor.py,python/test/cambi_test.py. - Invariant 1 (SYCL):
adm_options.hdefinesADM_BORDER_FACTORas a C macro; the#undefbefore the constexpr redeclaration must remain if upstream ever changes the macro name or value. If upstream removes the macro entirely, the#ifdef-guarded#undefis a no-op and safe. - Invariant 2 (Vulkan):
#ifndef VK_API_VERSION_1_4guard must remain until Ubuntu 22.04 is retired from CI (or the minimum Vulkan Headers version is bumped past 1.3.280). Track via ADR-0264 (NVIDIA driver regression gate). - Invariant 3 (Cambi FR atom feature):
CambiFullReferenceFeatureExtractoruses"cambi_encbd"(not"cambi") as the atom feature name for the distorted CAMBI score. If upstream changes theenc_bitdepthoption alias from"encbd"to something else, the vmafexec XML key changes and the Python extractor's wildcard prefix must be updated to match. - Re-test:
python3 -m pytest python/test/cambi_test.py -k "full_reference or fullref" -v
fix/precommit-onnx-binary-exclude — ADR collision sweep + pre-commit hook hardening¶
- Touches:
docs/adr/*.md(28 files renumbered to 0388–0415),docs/adr/README.md,docs/adr/_index_fragments/,scripts/ci/check-adr-numbering.sh,.pre-commit-config.yaml,tools/vmaf-tune/tests/test_hdr.py. - Invariant: No rebase impact on libvmaf C sources. The ADR renumbering affects documentation only; no code paths reference ADR numbers at runtime. Any in-flight branches that reference the old ADR numbers (0241-vmaf-tiny-v3, 0279-fr-regressor-v2-probabilistic, etc.) will need their references updated to the new numbers after rebasing onto master.
- Re-test:
bash scripts/ci/check-adr-numbering.shmust print "ADR numbering check passed."pre-commit run end-of-file-fixer ruff-check check-adr-numbering --all-filesmust all pass.
fix/round8-mcp-tmpdir-leak — MCP describe_worst_frames tmp-dir cleanup¶
No rebase impact: this change is MCP-server-only (mcp-server/vmaf-mcp/src/vmaf_mcp/server.py), touches no libvmaf C source, no public C API headers, no Meson build files, and no FFmpeg patch stack entries. Upstream Netflix/vmaf does not have the MCP server. The change adds a shutil.rmtree before the per-invocation PNG generation loop.
- Re-test:
PYTHONPATH=mcp-server/vmaf-mcp/src python -m pytest mcp-server/vmaf-mcp/tests/test_server.py::test_describe_worst_frames_tmpdir_cleared_on_next_call— must report 1 passed.
fix/round8-opt-nan-bypass — NaN rejection in set_option_double¶
- Touches:
core/src/opt.c— adds#include <math.h>and anisnan(n)guard inset_option_double. - Invariant: all callers of
vmaf_option_setwithVMAF_OPT_TYPE_DOUBLEmust receive-EINVALwhen the value string parses to NaN. Upstream Netflix/vmaf'sopt.cdoes not yet have this guard. If Netflix merges a version ofopt.cthat modifiesset_option_double(e.g. to add a new type or change the strtod flow), verify theisnanguard is preserved and still sits before then < min/n > maxchecks. - Re-test:
meson setup core/build-test libvmaf -Denable_cuda=false -Denable_sycl=false -Denable_tests=true && ninja -C core/build-test test/test_opt && core/build-test/test/test_opt— must report 25/25 passed, includingtest_double_nan_is_rejectedandtest_double_inf_rejected_when_max_finite.
fix/fex-dedup-by-provided-feature — feature-extractor dedup by provided-feature names (ADR-0385)¶
- Touches:
core/src/fex_ctx_vector.c(newprovided_features_overlap()helper, updatedfeature_extractor_vector_append()dedup logic);core/test/test_feature_extractor.c(new regression test);core/test/meson.build(addsfex_ctx_vector.cto test target sources). - Rebase impact: Low. The change is entirely internal to
fex_ctx_vector.c; no public C API headers are touched, nocore/include/changes, nomeson_options.txt, no FFmpeg patch stack entries. If upstream Netflix/vmaf rewritesfex_ctx_vector.cin a future sync, port theprovided_features_overlap()helper and its two-stage dedup logic forward; reverting to name-only dedup re-opens T-CUDA-FEATURE-EXTRACTOR-DOUBLE-WRITE on every GPU binary that combines--feature <name>with a default model load. - Re-test:
meson setup build-cpu libvmaf -Denable_cuda=false -Denable_sycl=false
ninja -C build-cpu
meson test -C build-cpu test_feature_extractor
# Expect: 6/6 tests passed, including
# test_fex_vector_dedup_by_provided_feature_name: pass
# Verify no "cannot be overwritten" warnings:
build-cpu/tools/vmaf \
-r python/test/resource/yuv/src01_hrc00_576x324.yuv \
-d python/test/resource/yuv/src01_hrc01_576x324.yuv \
-w 576 -h 324 -p 420 -b 8 --feature adm --threads 1 \
2>&1 | grep "cannot be overwritten" | wc -l
# → 0
fix/pypsnr-ast-eval — JSON log serialization in PyFeatureExtractorMixin¶
No rebase impact: this change is Python-only (python/vmaf/core/feature_extractor.py), touches no C/CUDA/SYCL/HIP/Vulkan/Metal source, no public C API headers, no Meson build files, and no FFmpeg patch stack entries. The log files written by _generate_result are transient per-run scratch files (under workdir/); the format change from Python repr to JSON is invisible to callers. If upstream Netflix/vmaf modifies PyFeatureExtractorMixin._get_feature_scores or _generate_result in a future sync, verify that neither side re-introduces str() / ast.literal_eval — the numpy 2.x incompatibility is the root cause of T-PYPSNR-AST-EVAL.
- Re-test:
PYTHONPATH=$PWD/python python3 -m pytest python/test/feature_extractor_test.py -k pypsnr— must report 8/8 passed.
fix/pypsnr-feature-extractor-import — PyPsnrFeatureExtractor class hierarchy restoration¶
No rebase impact: this change is Python-only (python/vmaf/core/feature_extractor.py), touches no C/CUDA/SYCL/HIP/Vulkan/Metal source, no public C API headers, no Meson build files, and no FFmpeg patch stack entries. If upstream Netflix/vmaf adds or removes PyPsnrFeatureExtractor / PypsnrFeatureExtractor in a future sync, audit feature_extractor.py lines 722–830 to ensure the primary-vs-deprecated alias relationship is preserved (primary = PyPsnr*, deprecated = Pypsnr*).
0086 — TransNet shot-metadata columns + HDR VMAF model port slot (Research-0086, ADR-0300 follow-up)¶
- Touches:
tools/vmaf-tune/src/vmaftune/__init__.py(CORPUS_ROW_KEYS additive trio, noSCHEMA_VERSIONbump),tools/vmaf-tune/src/vmaftune/per_shot.py(summarise_shots,_detect_shots_with_status,ShotMetadata),tools/vmaf-tune/src/vmaftune/corpus.py(_resolve_shot_metadata, row population, newshot_runnerkwarg oniter_rows),tools/vmaf-tune/src/vmaftune/hdr.py(transfer-awareselect_hdr_vmaf_model,hdr_model_name_for,HDR_MODEL_FILENAME, single-shot warning helper), tests + docs. - Invariant:
iter_rowsrunsvmaf-perShotexactly once per source — the cost of TransNet inference is too high to pay per (preset, crf) cell. If a future PR moves shot detection inside the cell loop the corpus-generation wall time roughly doubles. Keep the per-source resolution at the top ofiter_rowsand passShotMetadatadown to_row_for. Additionally:_detect_shots_with_statusis the only call site that distinguishes "real one-shot source" from "fallback because the binary failed" — the publicdetect_shotsshape cannot carry that boolean and downstream consumers depend on the(shots, ok)tuple to emit(0, 0.0, 0.0)sentinel rows. - Upstream conflict probability: zero. Upstream Netflix/vmaf does not carry a
vmaf-tunedirectory, anhdr.py, or a shot-detection harness. The HDR VMAF model port slot (vmaf_hdr_v0.6.1.json) is fork-internal scaffolding — Netflix publishes the canonical artefact outside their publicmodel/tree. No upstream rebase will touch any of these files. - Re-test:
pytest tools/vmaf-tune/tests/test_hdr.py tools/vmaf-tune/tests/test_shot_metadata_columns.py tools/vmaf-tune/tests/test_per_shot.py.
0358 — CUDA motion race + leak + motion2/motion3 precision parity (ADR-0358)¶
- Touches:
core/src/feature/cuda/integer_motion_cuda.c(memset moved froms->strtopic_stream;motion2_scoreemission switched to the CPU'sMIN(score * motion_fps_weight, motion_max_val)post-process in collect + flush;motion3_postprocess_cudaguard relaxed toframe_index > 2for the pre-incremented frame counter;vmaf_cuda_buffer_host_free (s->sad_host)added toclose_fex_cudaand theinit_fex_cudaerror unwind),core/src/feature/cuda/integer_motion/motion_score.cu(shared-tile inner stride paddedTILE_W→ `TILE_PITCH = TILE_W - 1
;launch_bounds(BLOCK_X * BLOCK_Y, 8)added to both bpc kernels),core/src/feature/cuda/integer_motion_v2/motion_v2_score.cu(same padding + launch_bounds for the v2 twin),docs/adr/0358-...md,docs/adr/README.md(index row),docs/backends/cuda/overview.md(motion bit-exact-at-places=4 appendix to "Numerical tolerance vs the CPU scalar path"),docs/state.md(Recently-closed row),changelog.d/fixed/ cuda-motion-race-leak-precision.md. Upstream Netflix/vmaf does not currently ship the motion3-on-CUDA host post-processing surface (motion3_postprocess_cudais fork-local per ADR-0219) so a future rebase touchinginteger_motion_cuda.c` is unlikely to touch the same lines, but a pure-upstream port that resets the post-process to its naive form will silently un-fix bugs 3 + 4. - Invariant: the SAD
cuMemsetD8Asyncruns onpic_stream, NOT on the drain streams->str. The kernel'satomicAddlives onpic_stream; both streams areCU_STREAM_NON_BLOCKINGand there is no event linking them, so co-locating memset + kernel on the same stream is the only thing that orders them. Mirrors the verbatim pattern atinteger_motion_v2_cuda.c:188. Themotion2_scorerow emitted to the feature collector is the weighted-and-clipped valueMIN(score * motion_fps_weight, motion_max_val), NOT the rawmin(prev, cur)SAD score; this matchesinteger_motion.c:563. Themotion3_postprocess_cudamoving-average guard readsframe_index > 2, NOT> 1, becauseframe_indexis pre-incremented before the helper is called. - Re-test: build with
cd libvmaf && meson setup build-cuda -Denable_cuda=true -Denable_sycl=false --buildtype=release && ninja -C build-cuda, then runmeson test -C build-cuda(expect 55/55), then runtools/vmaf -r python/test/resource/yuv/src01_hrc00_576x324.yuv -d python/test/resource/yuv/src01_hrc01_576x324.yuv -w 576 -h 324 -p 420 -b 8 --backend cuda --feature motion_cuda --output /tmp/cuda.json --jsonand the same with--backend cpu --feature motion --output /tmp/cpu.json --json;places=4diff overinteger_motion,integer_motion2,integer_motion3should report0/144mismatches atmax_abs = 0.00e+00.compute-sanitizer --tool memcheck --leak-check full tools/vmaf ... --backend cuda --feature motion_cudareportsLEAK SUMMARY: 0 bytes leaked in 0 allocationspost-fix.compute-sanitizer --tool racecheckreports0 hazards.
0326 — vmaf-tune codec-adapter dispatcher pivot (ADR-0297, HP-1)¶
- Touches:
tools/vmaf-tune/src/vmaftune/encode.py(build_ffmpeg_command+ new_resolve_codec_args/_legacy_codec_argshelpers),tools/vmaf-tune/src/vmaftune/per_shot.py(_segment_commandsignature + body, new_default_segment_preset),tools/vmaf-tune/src/vmaftune/codec_adapters/(11 adapters gainffmpeg_codec_args+extra_params; libaom slice normalised),tools/vmaf-tune/tests/test_encode_dispatcher_per_adapter.py(new),tools/vmaf-tune/tests/test_codec_adapter_libaom.py(slice expectations updated for new contract). Upstream Netflix/vmaf has novmaf-tunesurface, so conflict probability is zero — this entry exists because the dispatcher contract is fork-local and any future adapter PRs need to land bothffmpeg_codec_argsand a matching fixture row. - Invariant: every entry in
tools/vmaf-tune/src/vmaftune/codec_adapters/__init__.py::_REGISTRYships anffmpeg_codec_args(preset, quality) -> list[str]method that returns the codec-correct argv slice (with-c:v <encoder>as the first two tokens). The runtime contract is enforced bytests/test_encode_dispatcher_per_adapter.py::test_fixture_table_covers_every_registered_adapter— adding an adapter without a fixture row fails this meta-test. x264 and x265 argv shapes stay byte-for-byte["-c:v", encoder, "-preset", preset, "-crf", str(quality)](defended by thetest_x{264,265}_argv_byte_for_byte_legacy_shapepinning tests). The legacy fallback inencode._resolve_codec_argsreturns the historic libx264 shape for unregistered encoders so callers that bypass the registry stay invocable. - Re-test:
cd tools/vmaf-tune && pytest tests/test_encode_dispatcher_per_adapter.py tests/test_per_shot.py tests/test_codec_adapter_libaom.py -v(36 + 16 + 12 = 64 tests green on this branch). For a wider check, run the full vmaf-tune suite — pre-existing failures intest_recommend.py,test_resolution.py,test_encode_multi_codec.py(parse_versions(encoder=...)+encoder_runner=), andtest_codec_adapter_{x265,svtav1}.py(parse_versions(encoder=...)) are unrelated and predate HP-1.
0310 — Vulkan VIF int64 reduction race condition Phase 3 fix¶
- Touches:
core/src/feature/vulkan/shaders/vif.comp(replaces all three barebarrier()calls with explicitmemoryBarrierShared(); barrier();pairs covering the Phase-1 cooperative tile load, the Phase-2 vertical-conv shared write, and the Phase-4 cross-subgroup int64 reduction); plus documentation underdocs/research/0089-...md(Phase 3 status appendix),docs/adr/0269-...md(Phase 3 status appendix),docs/state.md(T-VK-VIF-1.4-RESIDUAL closed; new T-VK-VIF-1.4-RESIDUAL-ARC opened),core/src/vulkan/AGENTS.md(Phase 3 update on the existing invariant row),changelog.d/fixed/vif-int64-reduction-race-condition.md. Upstream Netflix/vmaf has no Vulkan backend, so conflict probability for the shader is zero. The entry exists because the fix is rebase-sensitive: any future cherry-pick that touchesvif.compand downgrades amemoryBarrierShared(); barrier();pair back to a barebarrier()will silently re-introduce the NVIDIA Vulkan 1.4 race. - Invariant:
vif.compshared-memory ordering between cooperative-write phases must be release-acquire, not just a bare workgroup-execution barrier. NVIDIA's Vulkan 1.4 default memory model requires the explicit shared-memory release; barebarrier()works at API 1.3 by accident on this driver. SCALE is irrelevant — the fix applies to all four pipeline specialisations because the barrier sites are in the SCALE-shared code. Do NOT remove the explicitmemoryBarrierShared()calls even if a perf review claims they are redundant under the GLSL spec wording: empirical real-hardware evidence in research-0089 2026-05-09 appendix shows otherwise on NVIDIA driver 595.71.05. - Re-test: apply the local API-1.4 bump (
core/src/vulkan/common.c3 sites +vma_impl.cppVMA_VULKAN_VERSION 1004000) on a NVIDIA RTX 4090 + driver 595.71+ machine, build withmeson setup ... -Denable_vulkan=enabled, then runpython3 scripts/ci/cross_backend_vif_diff.py --feature vif --backend vulkan --device 1 --places 4. Expect 0/48 across all four scales. Run the 5-run determinism check from research-0089 §"Reproduction recipe for Phase 3" against--vulkan_device 1; expect 5 identical(integer_vif_num_scale2, integer_vif_den_scale2) = (+2.494358e+04, +2.522523e+04)pairs at frame 5. Note that--vulkan_device 0on this multi-GPU host is the Intel Arc A380 lane and will still fail at API 1.4 (separateT-VK-VIF-1.4-RESIDUAL-ARCrow Open).
0309 — Vulkan VIF API-1.4 Phase 2 dump (T-VK-VIF-1.4-RESIDUAL)¶
- Touches:
docs/research/0089-vulkan-vif-fp-residual-bisect-2026-05-08.md(2026-05-09 status appendix with empirical numbers from the live RTX 4090),docs/state.md(T-VK-VIF-1.4-RESIDUAL row updated with the localisation),core/src/vulkan/AGENTS.md(new invariant row pinning the SCALE = 2 cross-subgroup-reduction memory-model finding),CHANGELOG.md(lusoris fork "Changed" entry). No code touched; the Phase 3 shader memory-model fix lands in a separate PR. Upstream Netflix/vmaf has no Vulkan backend so conflict probability for the AGENTS.md row is zero — entry exists because the empirical localisation flips the open state-row hypothesis from FP-precision to memory-model and retires theplaces=3override path that earlier rebase scaffolding might have suggested. - Invariant:
vif.compSCALE = 2 specialisation's Phase-4 cross-subgroup int64 reduction is non-deterministic on NVIDIA driver 595.71.05 + Vulkan 1.4.341 (lines 547–592,subgroupAdd barrier()+ thread-0 read ofs_lmem). API 1.3 lane is fully deterministic on the same hardware. The fourapiVersionpinning sites incore/src/vulkan/common.c+core/src/vulkan/vma_impl.cppstay at 1.3 until Phase 3 lands the explicit memory-scope barrier and a 5-run determinism gate confirms run-to-run identical(num, den)plusplaces=40/48 on NVIDIA. Theplaces=3override path is eliminated from the unblock options.- Re-test: apply the local API-1.4 bump (
core/src/vulkan/common.c3 sites +vma_impl.cppVMA_VULKAN_VERSION 1004000) on a NVIDIA RTX 4090 + driver 595.71+ machine, build withmeson setup ... -Denable_vulkan=enabled, then run the gate and the 5-run determinism check from research-0089 §"Reproduction recipe for Phase 3". Expect 45/48places=4failures oninteger_vif_scale2(max abs1.527e-02) AND 5 distinct(integer_vif_num_scale2, integer_vif_den_scale2)pairs across 5 runs of--feature 'vif_vulkan=debug=true'. Both observations reproduced bit-for-bit on this session's hardware lane (UUIDe478b41b-5c4f-1ddb-f990-e44916aff4c8).
0309 — vmaf-tune fast CLI surface (ADR-0276 status-update appendix, HP-3)¶
- Touches:
tools/vmaf-tune/src/vmaftune/cli.py(newfastsubparser +_run_fast+_build_fast_sample_extractor+_build_fast_encode_runner+_parse_canonical6_means),tools/vmaf-tune/tests/test_cli_fast.py(new),tools/vmaf-tune/AGENTS.md(new exit-code-contract invariant bullet under the fast-path section),docs/adr/0276-vmaf-tune-fast-path.md(status-update appendix),docs/usage/vmaf-tune.md(new## fastsection). Upstream Netflix/vmaf has no fast-path surface, so conflict probability is zero — entry exists because the canonical-6 parsing path offpooled_metrics.<feature>.meanis sensitive to libvmaf's JSON output shape. - Invariant: The libvmaf JSON layout the
_parse_canonical6_meanshelper consumes is thepooled_metrics.<feature>.meanshape (modern libvmaf 3.x), with a per-frameframes[].metrics.<feature>fallback. Both shapes are covered byparse_vmaf_jsonfor the headline VMAF score inscore.py; canonical-6 means re-use the same surface. If upstream changes the JSON schema (e.g. nestspooled_metricsunder a new key),_parse_canonical6_meansfollows in the same PR — the fast-path proxy depends on the canonical-6 vector being correctly extracted from libvmaf's output. The OOD-gap exit code3from_run_fastis the documented fall-back signal indocs/usage/vmaf-tune.md§ "Fall-back idiom"; do not silently downgrade it to0. The CLI is the only seam that injectssample_extractorandencode_runnerintofast.fast_recommend; the Python API still raisesNotImplementedErrorwhen called without them. - Re-test:
PYTHONPATH=tools/vmaf-tune/src python -m pytest tools/vmaf-tune/tests/test_cli_fast.py tools/vmaf-tune/tests/test_fast.py -v(21 tests). Smoke end-to-end without ffmpeg / ONNX / GPU:vmaf-tune fast --target-vmaf 92 --smoke --n-trials 8should emit a JSON payload whosesmokeistrueandverify_vmafisnull.vmaf-tune fast --helplists every flag in_DOCUMENTED_FAST_FLAGSfromtest_cli_fast.py.
0366 — vmaf-tune corpus schema v3 (ADR-0366)¶
- Touches:
tools/vmaf-tune/src/vmaftune/__init__.py(SCHEMA_VERSION2 → 3, +12 canonical-6 aggregate keys),tools/vmaf-tune/src/vmaftune/score.py(newparse_feature_aggregates,ScoreResult.feature_means/_stds),tools/vmaf-tune/src/vmaftune/corpus.py(writer projects the aggregates into row keys; newread_jsonlwith v2 back-compat),ai/scripts/train_fr_regressor_v[23].py(consume the new columns directly from the corpus DataFrame). All paths are wholly fork-local —tools/vmaf-tune/andai/scripts/are not mirrored upstream — so rebase impact is zero. - Invariant: Phase B/C/D consumers and the FR-regressor trainers rely on the canonical-6
<feature>_meancolumns being present onschema_version >= 3rows and beingNaN(never0.0) when libvmaf does not expose the feature. Keep the writer-sideNaNcontract intact during any future widening; trainers drop NaN rows before fitting StandardScaler. The reader (read_jsonl) preserves the on-diskschema_versionso trainers can filter to>= 3if they need real per-feature data. - Re-test:
cd tools/vmaf-tune && python -m pytest \
tests/test_corpus.py tests/test_corpus_schema_v3.py \
tests/test_corpus_v2_back_compat.py -q
python -m pytest ai/tests/test_train_fr_regressor_v3.py -q
0308 — encoder knob-sweep recipe-regression policy (ADR-0308, docs-only)¶
- Touches:
docs/research/0080-encoder-knob-sweep-findings.md,docs/adr/0308-encoder-knob-sweep-recipe-regression-policy.md,docs/adr/README.md(index row),ai/AGENTS.md(knob-sweep invariant section),changelog.d/changed/encoder-knob-sweep-findings.md. No code touched; companion to PR #400 (ADR-0305 + Research-0077 +ai/scripts/analyze_knob_sweep.py). Upstream Netflix/vmaf has no encoder-knob-sweep surface, so conflict probability is zero — this entry exists only because the policy threshold (7-of-9 structural cut) is rebase-sensitive on the corpus shape. - Invariant: the 7-of-9 source-count threshold from ADR-0308 §Decision point 1 is calibrated against the current 9-source Netflix Public Dataset corpus. If the corpus grows past 9 sources (e.g. UGC expansion per ADR-0287, or HDR additions), re-derive the absolute threshold as a fraction (≥7/9 ≈ 78 %). The structural cluster is sharp on the current corpus (top-15 cells all hit 9-of-9, no observed cells in 4-6 range), so a fractional cut at ~75 % is robust. Do NOT relax
bitrate_tol_pct(default 5.0) orvmaf_tol(default 0.1) inai/scripts/analyze_knob_sweep.pywithout an ADR — those tolerances are calibrated against the per-frame VMAF noise floor and bitrate quantisation in libavformat muxers. - Re-test:
pytest ai/tests/test_knob_sweep_analysis.py -v(script logic; ships in PR #400). Policy gate is offline: regenerateruns/phase_a/full_grid/comprehensive.jsonlviatools/vmaf-tune/src/vmaftune/hw_encoder_corpus.py(3-hour run on a single host with NVENC + QSV) then re-runpython ai/scripts/analyze_knob_sweep.py --jsonl <adapted.jsonl> --out-dir runs/phase_a/full_grid/reports/and diff the resultingsummary.mdagainstdocs/research/0080-encoder-knob-sweep-findings.mdheadline table. Structural cluster (top-15 cells, all 9-of-9) is the invariant to defend.
0228 — Vulkan 1.4 bump deferred (ADR-0264, docs-only)¶
- Touches: none (docs-only PR). Future Step A of T-VK-1.4-BUMP will touch
core/src/feature/vulkan/shaders/vif.compandcore/src/feature/vulkan/shaders/ciede.comp; Step B will touch the threeapiVersionsites incore/src/vulkan/common.c(lines 54, 264, 374) and theVMA_VULKAN_VERSIONdefine incore/src/vulkan/vma_impl.cpp(line 22). - Invariant:
masterstays onVK_API_VERSION_1_3andVMA_VULKAN_VERSION = 1003000. Lifting the constant in any future upstream sync (Netflix doesn't ship a Vulkan backend, so the conflict is improbable) without first auditingprecise/OpDecorate ... NoContractiondecoration onvif.compandciede.compwill reintroduce the NVIDIA-driver regression captured in research-0053. Thepsnr_hvs_strict_shaders-O0list incore/src/vulkan/meson.buildis the existing precedent for shader-side bit-exactness mitigations and should be the place a 1.4-era audit lands its results (potentially expanding to covervif.comp+ciede.compif thepreciseaudit decides the optimizer is the right place to gate). - Re-test: when Step B lands, the gate is
python3 scripts/ci/cross_backend_vif_diff.py --feature vif --backend vulkanand the same with--feature ciedeagainst NVIDIA + RADV + lavapipe; max abs diff must stay ≤5.0e-05(places=4) on all three.
0229 — HIP fifth-consumer kernel float_ansnr_hip (ADR-0266)¶
0228 — y4m_convert_411_422jpeg 1-byte heap-buffer-overflow fix¶
0228 — vmaf-tune resolution-aware model selection (ADR-0289)¶
0282 — vmaf-tune AMD AMF codec adapters (ADR-0282)¶
0228 — tools/vmaf-tune/ codec-agnostic encode dispatcher (ADR-0294)¶
- Touches:
tools/vmaf-tune/src/vmaftune/encode.py— refactored to look up the codec adapter and delegate argv composition. Wholly fork-local.tools/vmaf-tune/src/vmaftune/codec_adapters/__init__.py,codec_adapters/x264.py— adapter contract gainsffmpeg_codec_args(preset, quality)andextra_params(). Both are duck-typed; missing methods fall back to the legacy x264-CRF shape.tools/vmaf-tune/tests/test_encode_multi_codec.py— new 19-test suite pinning the dispatcher contract per codec.docs/usage/vmaf-tune.md— new "Codec adapter contract" section.- Invariant: the harness (
encode.py,corpus.py) must not branch on codec identity. The only codec-aware code is the per-adaptercodec_adapters/*.pyfile. Any future change that adds anif adapter.encoder == "..."to the harness regresses ADR-0294's whole-purpose. The corpus row schema stays at SCHEMA_VERSION=1 —crfis preserved as the row column even when the underlying codec's quality knob is-cq/-qp/ etc.;EncodeRequest.qualityis a request-side property only. Adapters that don't yet exposeffmpeg_codec_argsare intentionally permitted to fall back to the legacy x264-CRF shape; removing that fallback would break in-flight adapter PRs landing one-at-a-time. - Re-test on rebase:
```bash pytest tools/vmaf-tune/tests/ -q # 32 passed (13 existing + 19 multi-codec)
python -c " from pathlib import Path from vmaftune.encode import EncodeRequest, build_ffmpeg_command req = EncodeRequest( source=Path('ref.yuv'), width=1920, height=1080, pix_fmt='yuv420p', framerate=24.0, encoder='libx264', preset='medium', crf=23, output=Path('out.mp4'), ) cmd = build_ffmpeg_command(req) assert cmd[cmd.index('-c:v') + 1] == 'libx264' assert cmd[cmd.index('-preset') + 1] == 'medium' assert cmd[cmd.index('-crf') + 1] == '23' print('x264 dispatcher path OK') "
0260 — vmaf-tune --sample-clip-seconds (ADR-0301)¶
- Touches:
tools/vmaf-tune/src/vmaftune/{cli,corpus,encode,score,__init__}.py— fork-local. No upstream Netflix/vmaf path overlap.tools/vmaf-tune/tests/test_corpus.py,tools/vmaf-tune/AGENTS.md,docs/usage/vmaf-tune.md,docs/adr/0301-vmaf-tune-sample-clip.md,docs/adr/_index_fragments/0301-vmaf-tune-sample-clip.md,docs/adr/_index_fragments/_order.txt,docs/adr/README.md.- Invariant: corpus JSONL
SCHEMA_VERSIONbumped to2— additiveclip_modekey only. Sample-clip windows are mirrored on both sides via FFmpeg input-side-ss/-t(encode) and libvmaf's--frame_skip_ref/--frame_cnt(score). The_resolve_sample_clip()helper is the single source of truth for the centre-anchored slice math; do not duplicate the computation elsewhere. Falls back silently to"full"whenN >= duration_s. - Re-test:
0227 — tools/vmaf-tune/ Phase A scaffold (ADR-0237 Phase A)¶
- Touches:
tools/vmaf-tune/src/vmaftune/codec_adapters/{h264_amf,hevc_amf,av1_amf,_amf_common}.py(new). Wholly fork-local — no upstream Netflix/vmaf path overlap.tools/vmaf-tune/src/vmaftune/codec_adapters/__init__.py— registry extended with three AMF entries.tools/vmaf-tune/tests/test_codec_adapter_amf.py(new).tools/vmaf-tune/tests/test_corpus.py— Phase A test renamed fromtest_known_codecs_phase_a_is_x264_onlytotest_known_codecs_includes_x264_and_amf.tools/vmaf-tune/AGENTS.md— adds AMF preset-compression invariant.docs/usage/vmaf-tune.md— adds Hardware encoders section.- Invariant: the 7-into-3 preset compression table in
_amf_common.py(_PRESET_TO_AMF) is the cross-codec axis Phase B / C consumers depend on. Every AMF adapter accepts the canonical 7 preset names (placebo…ultrafast) and maps them onto the three AMF rungs (quality/balanced/speed). Do not extend the preset vocabulary without amending ADR-0282 — registry uniformity (no codec-identity branching in the harness search loop) rests on every codec accepting the same names. - Re-test:
0227 — tools/vmaf-tune/ Phase A scaffold (ADR-0237 Phase A)¶
- Touches:
tools/vmaf-tune/src/vmaftune/resolution.py(new). Wholly fork-local — no upstream Netflix/vmaf path overlap.tools/vmaf-tune/src/vmaftune/corpus.py— addsCorpusOptions.resolution_aware: bool = Trueand pipes the effective model throughscore_res.request.modelinto the JSONL row.tools/vmaf-tune/src/vmaftune/cli.py— adds--resolution-aware/--no-resolution-aware(BooleanOptionalAction, default on).tools/vmaf-tune/tests/test_resolution.py(new).docs/usage/vmaf-tune.md— new "Resolution-aware mode" section.docs/adr/0289-vmaf-tune-resolution-aware.md(new) +docs/research/0064-vmaf-tune-resolution-aware.md(new).tools/vmaf-tune/AGENTS.md— two new invariant notes.- Invariant: the height-only decision rule (
height >= 2160→vmaf_4k_v0.6.1, elsevmaf_v0.6.1) is the documented contract. The JSONLvmaf_modelfield is now per-row (not per-job) — mixed ladder corpora legitimately contain multiple distinct values across rows. Downstream consumers (Phase B / C / D) must group/filter byvmaf_modelrather than assuming a constant. Width is accepted in the API for symmetry but ignored in the body; do not branch on it without a follow-up ADR. - Re-test:
pytest tools/vmaf-tune/tests/ -q
python tools/vmaf-tune/vmaf-tune corpus --help | grep resolution-aware
0227 — tools/vmaf-tune/ Phase A scaffold (ADR-0237 Phase A)¶
- Touches:
core/tools/y4m_input.c— upstream-mirrored Daala-derived Y4M parser. The fix sits inside the 4:1:1 → 4:2:2-jpeg chroma upsample routiney4m_convert_411_422jpeg, lines ~500–530 in the function's three sub-loops. Upstream Netflix/vmaf carries the same shape; if upstream lands its own fix during a sync, prefer the upstream version and drop ours.core/test/test_y4m_411_oob.c(new, fork-local) — drives the minimal W=2 H=4 4:1:1 stream throughvideo_input_open+video_input_fetch_frame. Wholly fork-added; no upstream collision.core/test/meson.build— addstest_y4m_411_oobexecutable +test()registration.- Invariant: the first two sub-loops of
y4m_convert_411_422jpegmust guard_dst[(x << 1) | 1]writes with(x << 1 | 1) < dst_c_w, matching the third sub-loop's existing guard. Without the guard a 4:1:1 stream of width 2 (dst_c_w == 1) writes one byte past the destination chroma row. - Re-test:
cd libvmaf && meson setup ../build-asan --buildtype=debug -Db_sanitize=address -Db_lundef=false -Denable_cuda=false -Denable_sycl=false -Denable_vulkan=disabledninja -C build-asan test/test_y4m_411_oobASAN_OPTIONS=detect_leaks=0 ./build-asan/test/test_y4m_411_oob— must report1 tests run, 1 passed. Pre-fix the binary aborts withAddressSanitizer: heap-buffer-overflow … WRITE of size 1aty4m_input.c:507.
0270 — saliency_student_v1 fork-trained on DUTS-TR (ADR-0286)¶
- Touches:
model/tiny/registry.json— adds thesaliency_student_v1row. Fork-local registry; no upstream overlap.model/tiny/saliency_student_v1.onnx(+.jsonsidecar) — new weights and metadata. Fork-local.ai/scripts/train_saliency_student.py— new training script. Wholly fork-local underai/, which has no upstream counterpart.docs/ai/models/saliency_student_v1.md,docs/research/0062-saliency-student-from-scratch-on-duts.md,docs/adr/0286-saliency-student-fork-trained-on-duts.md— new docs under fork-local trees.- Invariant: the C-side
feature_mobilesal.cextractor's tensor-name contract —input(NCHW[1, 3, H, W]) andsaliency_map(NCHW[1, 1, H, W]) — must continue to match the ONNX graph for bothsaliency_student_v1.onnxand the legacymobilesal.onnxplaceholder. Future weights swaps can change the graph internals freely but must keep these names + shapes; the smoke test asserts the registration. The op-allowlist constraint (graph uses only ops incore/src/dnn/op_allowlist.c) carries over from ADR-0218 —Resizeis not used;ConvTransposeis the upsample op for v1 to keep the graph load-clean against vanilla origin/master. - Re-test:
.venv/bin/python ai/scripts/validate_model_registry.py
.venv/bin/python -c "
from ai.src.vmaf_train.op_allowlist import check_model
from pathlib import Path
r = check_model(Path('model/tiny/saliency_student_v1.onnx'))
assert r.ok, r.pretty()
print('allowlist OK')
"
meson test -C build --suite=fast mobilesal
0227 — tools/vmaf-tune/ Phase A scaffold (ADR-0237 Phase A)¶
- Touches:
core/src/feature/hip/float_ansnr_hip.{c,h}(new) — fifth consumer ofcore/src/hip/kernel_template.h. Mirrorscore/src/feature/cuda/float_ansnr_cuda.ccall-graph-for-call-graph;init/submit/collect/closeinvoke the kernel-template helpers in the same order; the submit body intentionally bypassesvmaf_hip_kernel_submit_pre_launch(no atomic, kernel writes per-block (sig, noise) interleaved float partials directly).core/src/hip/meson.build— adds the new TU tohip_sources.core/src/feature/feature_extractor.c— adds theextern VmafFeatureExtractor vmaf_fex_float_ansnr_hip;declaration and the registry row under#if HAVE_HIP.core/test/test_hip_smoke.c— addstest_float_ansnr_hip_extractor_registeredsub-test pinning the lookup contract.- Invariant — the
submit_pre_launchbypass is load-bearing. The CUDA twin makes the same choice for the same reason. If a future PR adds asubmit_pre_launchcall tofloat_ansnr_cuda.c's submit path, the HIP twin must follow in the same PR. Likewise the readback shape (wg_count * 2u * sizeof(float)) and the bpc table (peak/psnr_max for 8/10/12/16-bit) mirror the CUDA twin verbatim — keep aligned on rebase. - Re-test on rebase:
cd libvmaf
meson setup build -Denable_hip=true -Denable_cuda=false -Denable_sycl=false
ninja -C build
meson test -C build # 48/48 green (47 CPU + HIP smoke)
0230 — HIP sixth-consumer kernel motion_v2_hip (ADR-0267)¶
- Touches:
core/src/feature/hip/integer_motion_v2_hip.{c,h}(new) — sixth consumer ofcore/src/hip/kernel_template.h. Mirrorscore/src/feature/cuda/integer_motion_v2_cuda.ccall-graph-for-call-graph; carries theVMAF_FEATURE_EXTRACTOR_TEMPORALflag and aflush()callback. The state struct has auintptr_t pix[2]ping-pong slot pair tracked outside the kernel-template (the template models a single device+host pair only).core/src/hip/meson.build— adds the new TU tohip_sources.core/src/feature/feature_extractor.c— adds theextern VmafFeatureExtractor vmaf_fex_integer_motion_v2_hip;declaration and the registry row under#if HAVE_HIP.core/test/test_hip_smoke.c— addstest_motion_v2_hip_extractor_registeredsub-test pinning the lookup contract (extractor name ismotion_v2_hip, matching the CUDA twin'smotion_v2_cudanaming).- Invariant — temporal-extractor + ping-pong shape. The
VMAF_FEATURE_EXTRACTOR_TEMPORALflag bit, theflush()callback registration, and theuintptr_t pix[2]slot pair are load-bearing for the runtime PR (T7-10b). The runtime PR will swapuintptr_t pix[2]for a real device-buffer handle pair matching the CUDA twin'sVmafCudaBuffer *pix[2]. On rebase: if the CUDA twin's flush-pass shape changes (currentlymin(score[i], score[i+1])), update the HIP twin'sflush_fex_hipbody in the same PR. - Re-test on rebase: same as 0229 —
meson test -C buildwithenable_hip=trueexercises the smoke contract.
0227 — ms_ssim_vulkan submit-side migrated to kernel_template (T-GPU-DEDUP-26)¶
- Touches:
core/src/feature/vulkan/ms_ssim_vulkan.c—extract()'s rawVkCommandBuffer/VkFence/vkAllocateCommandBuffers/vkBeginCommandBuffer/vkCreateFence/vkQueueSubmit/vkWaitForFences/vkDestroyFence/vkFreeCommandBuffersblocks becomeVmafVulkanKernelSubmittriples (vmaf_vulkan_kernel_submit_begin/_submit_end_and_wait/_submit_free). One triple covers the decimate-pyramid command buffer; one triple per scale covers the per-scale SSIM submit. The pipeline-side bundles (pl_decimate2-binding 4-variant +pl_ssim10-binding 9-variant) and their_add_variant()chains are unchanged from the prior migration.- Invariant: any future submit-side template change (timeline semaphores, deferred fence release, queue-family parameterisation) must keep the helpers' synchronous-wait + per-frame fence + per-frame command-buffer contract intact, since
ms_ssim_vulkan.cdoes host readback of thel_partials/c_partials/s_partialsbuffers immediately after_submit_end_and_waitreturns. The submit-side contract is the same one already documented incore/src/vulkan/AGENTS.md's "Rebase-sensitive invariants" section forkernel_template.h. - Re-test:
```bash cd libvmaf && meson test -C build python scripts/ci/cross_backend_vif_diff.py \ --vmaf-binary core/build/tools/vmaf \ --reference testdata/ref_576x324_48f.yuv \ --distorted testdata/dis_576x324_48f.yuv \ --width 576 --height 324 \ --feature float_ms_ssim --backend vulkan --places 4
0231 — SHA-pin GitHub Actions (OSSF Pinned-Dependencies)¶
- Touches: every workflow file under
.github/workflows/. All 13 fork workflows (docker-image.yml,docs.yml,ffmpeg-integration.yml,libvmaf-build-matrix.yml,lint-and-format.yml,nightly-bisect.yml,nightly.yml,release-please.yml,rule-enforcement.yml,scorecard.yml,security-scans.yml,supply-chain.yml,tests-and-quality-gates.yml) had theiruses:directives rewritten from<owner>/<repo>@vN[.M.K]to<owner>/<repo>@<40-char-sha> # vN.M.K. 97 references converted; the SLSA reusable-workflow ref insupply-chain.ymlis the single documented holdout (seeInvariantbelow). - Invariant — SHA-pin policy for
uses:. Every action reference in.github/workflows/*.ymlMUST be a 40-char commit SHA with the semver tag preserved as a trailing# vN.M.Kcomment. The OSSF ScorecardPinned-Dependenciescheck parses both forms and a floating tag (@vN) is treated as unpinned and counts against the aggregate score. Single permitted exception: the SLSA generator reusable workflow (slsa-framework/slsa-github-generator/.github/workflows/generator_generic_slsa3.yml) must keep itsvX.Y.Ztag form because GitHub Actions consumers cannot SHA-pin reusable-workflow refs in every code path; the exception is documented inline insupply-chain.ymland survives on each rebase. Why this matters on upstream sync: Netflix upstream does not ship the fork's CI tree, so a/sync-upstreamrun that drags new workflow content (e.g. via repository templates or bot-authored bumps) into.github/workflows/can re-introduce floating-tag references unnoticed. The post-rebase check below is the standing gate — anything that lights up needs to be re-pinned before merging the sync. - Re-test on rebase:
# Anything that prints is a regression — every uses: must be either
# already SHA-pinned (40 hex) or, for the documented SLSA exception,
# the slsa-github-generator reusable-workflow ref.
grep -hnE '^\s*(- )?uses:\s+[^@]+@[^ #]+\s*$' .github/workflows/*.yml \
| grep -vE '@[a-f0-9]{40}' \
| grep -v 'slsa-framework/slsa-github-generator/.github/workflows/'
# SHA-resolution sanity for any new pin (per-action):
gh api repos/<owner>/<repo>/git/ref/tags/<vN.M.K> --jq '.object.sha'
# If the result is a "tag" object (annotated tag), deref:
gh api repos/<owner>/<repo>/git/tags/<sha-from-prev> --jq '.object.sha'
0226 — CUDA drain-batch engine-loop opt (T-GPU-OPT-1)¶
- Touches:
core/src/cuda/drain_batch.{h,c}(new) — TLS drain-batch table + shared drain stream +_open()/register/_flush()/_close()API.core/src/libvmaf.c— engine-side per-frame loop now wraps submit/collect with_open()+_flush()so all CUDA extractorfinishedevents are waited on a single shared drain stream.- All 12 CUDA feature kernels (
core/src/feature/cuda/*.c) register theirfinishedevent +drainedflag with the drain batch on submit; collect skips its privatecuStreamSynchronizewhendrainedis true. - Invariant — drained-flag contract. Every CUDA extractor's collect path must check the per-frame
drainedflag and skip its owncuStreamSynchronizewhen set; otherwise the drain batching is a no-op. The flag is reset tofalseper frame insidevmaf_cuda_drain_batch_register(). - Re-test on rebase:
cd libvmaf
meson setup build -Denable_cuda=true -Denable_sycl=false
ninja -C build
meson test -C build --suite=fast cuda
Expected: all CUDA tests green; bench shows ≥5% wall-clock gain on a 7-extractor VMAF model (model.json with all feature extractors enabled).
0225 — Netflix bench snapshot regen (upstream a44e5e61 motion fix)¶
- Touches:
testdata/netflix_benchmark_results.json— fork-added snapshot. CPU rows now reflect the post-fix motion feature; cuda / sycl rows from the previous regen are preserved unchanged because those backends were not exercised on this rerun (host-environment tooling — wrong renderD path,libvmaf_cudanot enabled in the local FFmpeg build). Future full regens should include cuda / sycl.testdata/bench_all.sh— defaultVMAF=no longer points at/usr/local/bin/vmaf(which on most dev hosts is stuck at the pre-upstream-a44e5e61v3.0.0); now defaults to the in-tree fork build atcore/build/tools/vmaf.testdata/benchmark_netflix.py—FFMPEG,YUVDIRand the hardcodedLD_LIBRARY_PATH=/usr/local/libare now overridable viaVMAF_FFMPEG,VMAF_YUVDIRand any caller-setLD_LIBRARY_PATH.- Invariant: the snapshot's CPU pooled VMAF for
src01_576x324is 76.667828 (post-fix), not 76.668904 (the upstream-buggy mirror). If/sync-upstreamever re-pulls a Netflix change that touchesmotion.cmirror-handling, this number is the reference. - Re-test:
cd libvmaf
meson setup build -Denable_cuda=true -Denable_sycl=false
ninja -C build
LD_LIBRARY_PATH=$(pwd)/build/src python3 \
../testdata/benchmark_netflix.py
Expected CPU pooled rows: 76.667828, 35.068672, 7.985899.
0224 — CUDA graph capture feasibility (research-0047, DEFER)¶
- Touches: none — investigation-only; no code lands. The research digest
docs/research/0047-cuda-graph-capture-feasibility.mddocuments why a CUDA graph capture path on the per-frame submit chain is deferred rather than shipped (realised wall-clock gain capped at ~1-3% vs. the predicted 10-20%, with a 4-slot picture-pool rotation that defeats single-graph capture and forces per-framecuGraphExecKernelNodeSetParamsrebinding for(ref, dis)device pointers). - Invariant: the
kernel_template.hdocstring keeps namingVmafCudaKernelLifecycle.finishedas a graph-capture hook point. Don't prune that comment on rebase — leaving the door open in the template is free, and the digest's "what needs to be true for a future GO" section depends on the hook still being there. - Re-test on rebase:
# Confirm the docstring still references graph capture as the hook
# point — wording change is fine, removal is not.
grep -q "graph capture" core/src/cuda/kernel_template.h
0223 — ADR slug-drift repair in CHANGELOG / rebase-notes (PR #304 follow-up)¶
- Touches:
CHANGELOG.md,docs/rebase-notes.md. No code; no upstream-shared path; no public-API surface. - Invariant: every
[ADR-NNNN](docs/adr/NNNN-slug.md)link in the fork's tracked docs resolves to an actual on-disk file underdocs/adr/. Repaired 4 broken slugs that did not exist on disk (0138-iqa-convolve-avx2-bitexact-double→0138-iqa-convolve-avx2-bitexact-double,0140-simd-dx-framework→0140-simd-dx-framework,0190-ms-ssim-vulkan→0190-ms-ssim-vulkan,0178-vulkan-adm-kernel→0178-vulkan-adm-kernel). All retained their cited NNNN per ADR-0028 (NNNN is immutable once Accepted). - Re-test on rebase: from repo root, the following must print no lines:
for ref in $(grep -ohE 'docs/adr/[0-9]{4}-[a-z0-9-]+\.md' \
CHANGELOG.md docs/rebase-notes.md AGENTS.md docs/state.md \
| sort -u); do
test -f "$ref" || echo "MISSING: $ref"
done
0125 — cambi_vulkan migrated to kernel_template (T-GPU-DEDUP-25, 5-bundle)¶
- Touches:
core/src/feature/vulkan/cambi_vulkan.c— state's quintet (dsl_2bind+ 5×pl_layout_*+shader_modules[CAMBI_PL_COUNT]- shared
desc_pool) collapses to fiveVmafVulkanKernelPipelinebundles (pl_trivial,pl_derivative,pl_filter_mode,pl_decimate,pl_mask_dp), each owning its own descriptor pool. The first slot ofpipelines[]per stage aliases the bundle's base pipeline;CAMBI_PL_FILTER_MODE_V,CAMBI_PL_MASK_SAT_COL, andCAMBI_PL_MASK_THRESHOLDare sibling variants built viavmaf_vulkan_kernel_pipeline_add_variant().
- shared
cambi_vk_alloc_settakes a bundle pointer (->desc_pool/->dsl) — every dispatch site picks the bundle that matches its push-constant struct.- The
cambi_vk_make_dsl/cambi_vk_make_pl/cambi_vk_create_shader/cambi_vk_build_pipelinehelpers are dropped — the template subsumes them. - Invariant — variants destroyed before bundle, base alias must be skipped. Five distinct push-constant struct sizes (
CambiVkPushTrivial/CambiVkPushDerivative/CambiVkPushFilterMode/CambiVkPushDecimate/CambiVkPushMaskDp) force five bundles even though every stage's DSL is 2-binding SSBO;_add_variant()only siblings pipelines under the same layout.close_fexmustvkDestroyPipeline()the variant slots (CAMBI_PL_FILTER_MODE_V,CAMBI_PL_MASK_SAT_COL,CAMBI_PL_MASK_THRESHOLD) before callingvmaf_vulkan_kernel_pipeline_destroy()on each bundle. - Numerical contract: bit-exact preserved. Same shaders + spec-constants + push-constants as before; only the Vulkan pipeline-bundle scaffolding moved to the template. Validated on the Netflix-pair smoke (576×324×8-bit):
cambimean = 0.0, identical to pre-migration (the pair has no banding artifacts). - Rebase impact: low. Builds on top of PR #272's
_add_variant()helper. Upstream Netflix/vmaf has no Vulkan backend, so there is nothing to merge against.
0124 — ssimulacra2_vulkan migrated to kernel_template (T-GPU-DEDUP-24, 4-bundle)¶
- Touches:
core/src/feature/vulkan/ssimulacra2_vulkan.c— state's 16 long-lived pipeline-object fields (4×*_dsl + *_pl + *_shader+ the shareddesc_pool) collapse to fourVmafVulkanKernelPipelinebundles (pl_xyb,pl_mul,pl_blur,pl_ssim), each owning its own descriptor pool. The first slot of each per-bundle pipeline array (xyb_pipelines[0],mul_pipelines[0],blur_pipelines_h[0],ssim_pipelines[0]) aliases the bundle's baseVkPipeline; remaining per-scale / per-pass slots are siblings viavmaf_vulkan_kernel_pipeline_add_variant().ss2v_build_pipeline_int3reroutes through_add_variant()instead of callingvkCreateComputePipelinesdirectly;ss2v_alloc_settakes a bundle pointer (->desc_pool/->dsl) instead of a separate DSL argument; descriptor-set free sites at the tail ofss2v_run_scaleroute to each bundle's pool.- The
ss2v_make_dsl/ss2v_make_pl/ss2v_create_shaderhelpers are dropped — the template subsumes them. - Invariant — variants destroyed before bundle, slot 0 alias must be skipped. Four distinct DSL shapes (XYB = 6 SSBOs, MUL = 3, BLUR = 2, SSIM = 8) prevent collapsing to one bundle:
_add_variant()only siblings pipelines under the same layout.close_fexmustvkDestroyPipeline()the variant slots inxyb_pipelines[1..N-1],mul_pipelines[1..N-1],ssim_pipelines[1..N-1],blur_pipelines_h[1..N-1], and every slot ofblur_pipelines_v[]before callingvmaf_vulkan_kernel_pipeline_destroy()on each bundle, and must skip slot 0 of the first three arrays +blur_pipelines_hto avoid double-freeing the aliased base. - Numerical contract: bit-exact preserved. Same shaders + spec-constants + push-constants as before; only the Vulkan pipeline-bundle scaffolding moved to the template. Validated on the Netflix-pair smoke (576×324×8-bit):
ssimulacra2mean = 24.613842, identical to pre-migration. - Rebase impact: low. Builds on top of PR #272's
_add_variant()helper. Upstream Netflix/vmaf has no ssimulacra2 extractor and no Vulkan backend, so there is nothing to merge against.
0118 — psnr_hvs_vulkan migrated to kernel_template + _add_variant (T-GPU-DEDUP-18)¶
- Touches:
core/src/feature/vulkan/psnr_hvs_vulkan.c— state'sdsl + pipeline_layout + shader + desc_pool + pipeline[3]collapses toVmafVulkanKernelPipeline pl + VkPipeline pipeline_chroma_u + VkPipeline pipeline_chroma_v. Plane 0 is the template's base pipeline; planes 1+2 are siblings viavmaf_vulkan_kernel_pipeline_add_variant().- New
psnr_hvs_plane_pipeline()accessor maps plane index to the rightVkPipelinehandle. - Invariant — variants destroyed before bundle.
close_fexmustvkDestroyPipeline()the chroma U/V variants before callingvmaf_vulkan_kernel_pipeline_destroy(&s->pl)— same rule as ssim_vulkan in T-GPU-DEDUP-7. - Numerical contract: unchanged. Same shaders + spec-constants
- push-constants as before; only the Vulkan pipeline-bundle scaffolding moved to the template.
- Rebase impact: low. Builds on top of PR #272's
_add_variant()helper.
0119 — vif_vulkan migrated to kernel_template + _add_variant (T-GPU-DEDUP-19)¶
- Touches:
core/src/feature/vulkan/vif_vulkan.c— state'sdsl + pipeline_layout + shader + desc_pool + pipelines[4]collapses toVmafVulkanKernelPipeline pl + VkPipeline scale_variants[3]. Scale 0 is the template's base pipeline; scales 1, 2, 3 are siblings viavmaf_vulkan_kernel_pipeline_add_variant().- New
vif_scale_pipeline()accessor maps scale index to the rightVkPipelinehandle (replacess->pipelines[scale]). - Invariant — variants destroyed before bundle.
close_fexmustvkDestroyPipeline()the 3 scale variants before callingvmaf_vulkan_kernel_pipeline_destroy(&s->pl)— same rule as ssim_vulkan in T-GPU-DEDUP-7 and psnr_hvs_vulkan in T-GPU-DEDUP-18. - Numerical contract: unchanged. Same shaders, same spec-constants, same push-constants as before; only the Vulkan pipeline-bundle scaffolding moved to the template.
- Rebase impact: low. Builds on top of PR #272's
_add_variant()helper.
0120 — float_vif_vulkan migrated to kernel_template + _add_variant (T-GPU-DEDUP-20)¶
- Touches:
core/src/feature/vulkan/float_vif_vulkan.c— state collapsesdsl + pipeline_layout + shader + desc_pooltoVmafVulkanKernelPipeline pl; theVkPipeline pipelines[2][4]2-D lookup table is preserved so the existing[mode][scale]dispatch path stays clean, butpipelines[0][0]aliasess->pl.pipeline(the template's base). The other 6 entries are sibling pipelines created viavmaf_vulkan_kernel_pipeline_add_variant().- Invariant — variants destroyed before bundle.
close_fexmustvkDestroyPipeline()the 6 sibling variants (every(mode, scale)except(0, 0)) before callingvmaf_vulkan_kernel_pipeline_destroy(&s->pl)— same rule as ssim_vulkan / psnr_hvs_vulkan / vif_vulkan. - Invariant —
pipelines[0][0]aliasing. The base pipeline handle is owned bys->pl.pipeline; we copy it intopipelines[0][0]after_create()so the dispatch path can use a uniform 2-D lookup. The destroy loop must skip(mode=0, scale=0)to avoid double-freeing the template's pipeline. - Numerical contract: unchanged. Same shaders, spec-constants (
mode+scale), push-constants. Netflix-pair smoke matchesinteger_vifbit-identically to 4 decimals. - Rebase impact: low. Builds on top of PR #272's
_add_variant()helper.
0122 — float_adm_vulkan migrated to kernel_template + _add_variant (T-GPU-DEDUP-22)¶
- Touches:
core/src/feature/vulkan/float_adm_vulkan.c— twin to adm_vulkan (T-GPU-DEDUP-21); 16-pipeline 2-D[stage][scale]array. State collapsesdsl + pipeline_layout + shader + desc_pooltoVmafVulkanKernelPipeline pl.pipelines[0][0]aliasess->pl.pipeline; the other 15 entries are siblings viavmaf_vulkan_kernel_pipeline_add_variant().- Invariants:
- Variants destroyed before bundle.
pipelines[0][0]aliasing — destroy loop must skip(stage=0, scale=0).- Numerical contract: unchanged. Same float (
_ssuffix) primitives fromadm_tools.c; same 5-element spec-constant tuple; same float partial accumulation reduced in double on the host.
0121 — adm_vulkan migrated to kernel_template + _add_variant (T-GPU-DEDUP-21)¶
- Touches:
core/src/feature/vulkan/adm_vulkan.c— state collapsesdsl + pipeline_layout + shader + desc_pooltoVmafVulkanKernelPipeline pl; theVkPipeline pipelines[4][4]2-D lookup is preserved so the per-stage dispatch path stays clean.pipelines[0][0]aliasess->pl.pipeline(the template's base); the other 15 entries are sibling pipelines viavmaf_vulkan_kernel_pipeline_add_variant().- Invariants:
- Variants destroyed before bundle (same rule as ssim_vulkan / psnr_hvs / vif / float_vif).
pipelines[0][0]aliasing — destroy loop must skip(stage=0, scale=0)to avoid double-freeing the template's pipeline.- Numerical contract: unchanged. Same shaders + 5-element spec-constant tuple (width, height, bpc, scale, stage) + push-constants.
- Rebase impact: low. Builds on top of PR #272.
0123 — ms_ssim_vulkan 2-bundle migration (T-GPU-DEDUP-23)¶
- Touches:
core/src/feature/vulkan/ms_ssim_vulkan.c— state collapsesdecimate_dsl + decimate_pl + decimate_shader + ssim_dsl + ssim_pl + ssim_shader + desc_pool(7 fields) to two bundlesVmafVulkanKernelPipeline pl_decimate+pl_ssim. Each bundle owns its own descriptor pool. The kernel has two distinct pipeline shapes (decimate = 2 SSBO bindings, ssim = 10 bindings), so two bundles is the minimum —_add_variant()only siblings pipelines under the same layout.decimate_pipelines[0]aliasespl_decimate.pipeline(the template's base = scale 0). The remainingMS_SSIM_SCALES - 2decimate variants (scales 1..3) are siblings via_add_variant().ssim_pipeline_horiz[0]aliasespl_ssim.pipeline(base = scale 0, pass 0). The other 9 entries (4×ssim_pipeline_horizfor scales 1..4, plus 5×ssim_pipeline_vertfor scales 0..4) are variants.- Invariant — variants destroyed before bundle. Same rule as ADR-0106 entry 0106:
close_fexmust destroydecimate_pipelines[1..3]andssim_pipeline_horiz[1..4]+ssim_pipeline_vert[0..4]before callingvmaf_vulkan_kernel_pipeline_destroy()onpl_decimate/pl_ssim. - Invariant —
[0]aliasing destroy-skip.decimate_pipelines[0]andssim_pipeline_horiz[0]must not be passed tovkDestroyPipelineinclose_fex—_destroy()already releases them viapl_decimate.pipeline/pl_ssim.pipeline. Double-free is UB. The destroy loops inclose_fexstart ati = 1for decimate and skipi == 0for ssim_horiz. - Invariant — per-bundle descriptor pool. The shared
s->desc_poolis gone;alloc_descriptor_setnow takes aconst VmafVulkanKernelPipeline *bundleand usesbundle->desc_pool+bundle->dsl. Per-framevkFreeDescriptorSetscalls must target the matching pool (pl_decimate.desc_poolfor decimate sets,pl_ssim.desc_poolfor ssim sets) — mixing them is undefined behavior. - Numerical contract: unchanged. Same shaders, spec constants, push constants, and dispatch order as before.
float_ms_ssimNetflix-pair smoke (576×324×48f) reports mean 0.963241; ssim pyramid intermediate values bit-identical to pre-migration run. - Rebase impact: low. Upstream Netflix has no Vulkan backend. Conflicts only against the parallel
T-GPU-DEDUP-{18..22}PRs (#284–#288) onCHANGELOG.md/docs/rebase-notes.md— auto-resolve keeps both halves.
0106 — Vulkan kernel template multi-pipeline + ssim/motion migration (T-GPU-DEDUP-7)¶
- Touches:
core/src/vulkan/kernel_template.h— newvmaf_vulkan_kernel_pipeline_add_variant()helper. Takes the base pipeline bundle (DSL / pipeline layout / shader / pool owned byvmaf_vulkan_kernel_pipeline_create) plus a partialVkComputePipelineCreateInfoand produces a siblingVkPipelinere-using the same layout / shader. The base_createand_destroyentry points are unchanged; existing consumers (psnr, moment, ciede) keep working.core/src/feature/vulkan/motion_vulkan.c— state collapsesVkPipeline pipelines[2](kept "for SYCL parity" but functionally identical because COMPUTE_SAD goes through push constants, not spec-constants) to a singleVmafVulkanKernelPipeline pl.create_pipelines/close_fexshrink to template-driven create + destroy.core/src/feature/vulkan/ssim_vulkan.c— state becomesVmafVulkanKernelPipeline pl + VkPipeline pipeline_vert. Pass 0 (horizontal) is the template's base pipeline; pass 1 (vertical) is created via_add_variant().close_fexdestroys the variant first, then callsvmaf_vulkan_kernel_pipeline_destroy()on the bundle.- Invariant — no spec-constant drift between base and variant.
_add_variant()overwritessType/stage.sType/stage.stage/stage.module/layoutof the caller'sVkComputePipelineCreateInfoso the variant is guaranteed to share the base's shader and layout. Callers control the variant's spec-constant viapSpecializationInfo. Reordering these overwrites lets a consumer accidentally bind a different shader module under the same layout — UB at descriptor-set time. - Invariant — variant destroyed before bundle.
close_fexin ssim mustvkDestroyPipeline(s->pipeline_vert)beforevmaf_vulkan_kernel_pipeline_destroy(&s->pl)— the bundle's_destroyreleases the descriptor pool, which thevkAllocateDescriptorSetsissued against the variant pipeline's layout cleanly drops only when the variant pipeline is already gone. - Numerical contract: unchanged. Both kernels run identical shaders + spec-constants + push-constants as before; only the Vulkan boilerplate that creates / destroys the pipeline scaffolding moved to a shared owner. Cross-backend parity gate at
places=4holds — Netflix-pairfloat_ssimsmoke (576×324×48f) reports mean 0.863, identical to pre-migration. - Rebase impact: low. The base pipeline-bundle helpers predate this change (PR #270 / #271); the new
_add_variantis additive. Upstream Netflix has no Vulkan backend to conflict with.
0111 — integer_ciede_cuda migrated to kernel_template (T-GPU-DEDUP-11)¶
- Touches:
core/src/feature/cuda/integer_ciede_cuda.c— state'sCUstream + CUevent + CUevent + VmafCudaBuffer + host-pinned float*quintet collapses toVmafCudaKernelLifecycle lc + VmafCudaKernelReadback rb. init / collect / close call the template'slifecycle_init/readback_alloc/collect_wait/lifecycle_close/readback_freehelpers. submit keeps the pre-launch wait inline (intentional — ciede has no atomic, so the template's pre-launch memset is unnecessary).- Numerical contract: unchanged. Pure CUDA-boilerplate consolidation. The host-side reduction in collect still uses the same
doubleaccumulator over per-block float partials —places=4(ADR-0187) holds.
0112 — integer_moment_cuda migrated to kernel_template (T-GPU-DEDUP-12)¶
- Touches:
core/src/feature/cuda/integer_moment_cuda.c— state's stream/event/device-buffer/host-pinned quintet collapses toVmafCudaKernelLifecycle lc + VmafCudaKernelReadback rb. submit callsvmaf_cuda_kernel_submit_pre_launch(atomic counters require the device-side memset). init / collect / close call the matching template helpers.- Numerical contract: unchanged. Same per-frame atomic accumulators (4× uint64), same
sums_host[i] / n_pixelshost division. - Rebase impact: low. Upstream Netflix has no equivalent template; this consolidation is fork-local.
0113 — integer_motion_v2_cuda migrated to kernel_template (T-GPU-DEDUP-13)¶
- Touches:
core/src/feature/cuda/integer_motion_v2_cuda.c— stream/event pair + sad device+host quintet collapses tolc + rb. Raw-pixel ping-pongpix[2]stays outside the bundle. submit keeps the memset onpic_streaminline rather than callingsubmit_pre_launch(the helper would move the memset tolc.str, which races with the kernel reading the accumulator). init / collect / close call the matching template helpers.- Numerical contract: unchanged. Same D2D copy, same conditional kernel launch on frame ≥ 1, same host-side
min(score[i], score[i+1])flush.
0114 — integer_ssim_cuda migrated to kernel_template (T-GPU-DEDUP-14)¶
- Touches:
core/src/feature/cuda/integer_ssim_cuda.c— stream/event/partials device+host quintet collapses tolc + rb. Five intermediate float buffers (h_ref_mu,h_cmp_mu,h_ref_sq,h_cmp_sq,h_refcmp) stay outside the bundle. submit keeps thecuStreamWaitEvent + horiz + vert + DtoHchain inline — SSIM writes one float per block (no atomic), so the template'ssubmit_pre_launchmemset is unnecessary. init / collect / close use the matching template helpers.- Numerical contract: unchanged. Same horiz-then-vert two-pass pipeline, same per-block float partial reduction in double on the host.
places=4(matching the ciede_cuda precision pattern) holds. - Rebase impact: low. Upstream Netflix has no equivalent; this is fork-added.
0115 — ms_ssim_cuda + psnr_hvs_cuda lifecycle migration (T-GPU-DEDUP-15)¶
- Touches:
core/src/feature/cuda/integer_ms_ssim_cuda.c— stream + 2-event lifecycle replaced withVmafCudaKernelLifecycle lc; multi-level pyramid + SSIM intermediate + 3-partials buffers stay outside the template's single-pair readback bundle.core/src/feature/cuda/integer_psnr_hvs_cuda.c— same shape; 3-plane ref/dist/partials triples remain inline.- Numerical contract: unchanged. The migration only affects init / close boilerplate; submit / collect dispatch and host reduction paths are untouched apart from the
s->str→s->lc.str/s->event→s->lc.submit/s->finished→s->lc.finishedfield renames.
0116 — float_psnr/ansnr/motion cuda → kernel_template (T-GPU-DEDUP-16)¶
- Touches:
core/src/feature/cuda/float_psnr_cuda.c— stream/event/partials quintet →lc + rb; input upload buffersref_in/dis_instay outside the bundle.core/src/feature/cuda/float_ansnr_cuda.c— same shape; rb wraps the (sig, noise) interleaved partials.core/src/feature/cuda/float_motion_cuda.c— same shape; rb wraps the SAD partials,blur[2]ping-pong stays outside.- Numerical contract: unchanged. Same dispatch geometry, same reduction order. Cross-backend parity gate at the kernels' contracted precision (places=3 per ADR-0192) holds.
0117 — float_adm + float_vif cuda lifecycle migration (T-GPU-DEDUP-17)¶
- Touches:
core/src/feature/cuda/float_adm_cuda.c— stream + 2-event lifecycle replaced withVmafCudaKernelLifecycle lc; multi-stage DWT + CSF pipeline state stays outside the template's single-pair readback bundle.core/src/feature/cuda/float_vif_cuda.c— same shape; 4-level pyramid + per-scale (num, den) pairs remain inline.- Numerical contract: unchanged. The migration only affects init / close stream-event boilerplate; submit / collect dispatch and host reduction paths are untouched apart from the field renames.
- Rebase impact: low. Upstream Netflix has no equivalent template; this is fork-added.
0107 — float_psnr_vulkan migrated to kernel_template (T-GPU-DEDUP-8)¶
- Touches:
core/src/feature/vulkan/float_psnr_vulkan.c— state'sdsl + pipeline_layout + shader + pipeline + desc_poolquintet is collapsed into a singleVmafVulkanKernelPipeline pl;create_pipelinesandclose_fexshrink to template-driven create + destroy. No shader changes, no spec-constant changes, no push-constant changes.- Numerical contract: unchanged. The migration is a pure Vulkan-boilerplate consolidation. Cross-backend parity gate at
places=4holds — Netflix-pair smoke reportsfloat_psnrmean 30.755 dB, identical to pre-migration.
0109 — float_ansnr_vulkan + motion_v2_vulkan migrated to kernel_template (T-GPU-DEDUP-9)¶
- Touches:
core/src/feature/vulkan/float_ansnr_vulkan.c— single-pipeline state collapses toVmafVulkanKernelPipeline pl;create_pipelinesandclose_fexshrink to template-driven create + destroy.core/src/feature/vulkan/motion_v2_vulkan.c— same shape.- Numerical contract: unchanged. Pure Vulkan-boilerplate consolidation. Cross-backend parity gate at the kernel's contracted precision holds — Netflix-pair smoke reports
float_ansnrmean 23.51 dB andmotion2_v2_scoremean 3.895, identical to pre-migration.
0110 — float_motion_vulkan migrated to kernel_template (T-GPU-DEDUP-10)¶
- Touches:
core/src/feature/vulkan/float_motion_vulkan.c— single-pipeline state collapses toVmafVulkanKernelPipeline pl;create_pipelinesandclose_fexshrink to template-driven create + destroy.- Numerical contract: unchanged. Pure Vulkan-boilerplate consolidation. Netflix-pair smoke reports
motionmean 4.049 /motion2mean 3.894, identical to pre-migration. - Rebase impact: low. Upstream Netflix has no Vulkan backend.
0108 — Bristol VI-Lab feasibility digest + BVI-CC ingest ADR (Draft)¶
- Touches:
docs/research/0046-bristol-vi-lab-feasibility.md(new) — nine-dataset survey + use-case fit + effort estimate.docs/adr/0241-bristol-bvi-cc-ingest.md(new, Status: Draft) — proposal to ingest BVI-CC as the second tiny-AI corpus.docs/adr/README.md— index row for ADR-0241.CHANGELOG.md— Added entry.- Numerical contract: not applicable (docs-only).
- Rebase impact: none. Pure research deliverables; upstream Netflix has no equivalent surface.
0094 — Vulkan VkImage import v2 async pending-fence (T7-29 part 4 / ADR-0251)¶
- ADR: ADR-0251; predecessor ADR-0186.
- Touches:
core/src/vulkan/import.c— full rewrite of the submission path. Single-fencesubmit_and_waitbecomes per-slotsubmit_to_slot+drain_slot_fence; the newslot_alloc/slot_releasehelpers materialise / tear down a ring slot (staging-pair + cmd buffer + fence).vmaf_vulkan_import_imageindexes into the ring byframe_index % ring_size;vmaf_vulkan_wait_computedrains every outstanding fence.vmaf_vulkan_state_build_pictureswaits the slot's fence before exposing the host pointer. Public-API signatures are unchanged.core/src/vulkan/vulkan_internal.h— newstruct VmafVulkanImportSlot;VmafVulkanImportSlotsbecomes a fixed-capacityVmafVulkanImportSlot ring[VMAF_VULKAN_RING_MAX]plus geometry +ring_size. Two new defines —VMAF_VULKAN_RING_DEFAULT(4) andVMAF_VULKAN_RING_MAX(8).VmafVulkanStategainsrequested_ring_size.core/src/vulkan/common.c—vmaf_vulkan_state_initand_state_init_externalsetrequested_ring_size = VMAF_VULKAN_RING_DEFAULT.core/test/test_vulkan_async_pending_fence.c(new, contract smoke for the v1 → v2 swap).core/test/meson.build— registers the new test under the existingenable_vulkanguard.core/src/vulkan/AGENTS.md(new) — pins the three rebase-sensitive ring invariants.docs/adr/0251-vulkan-async-pending-fence.md(new),docs/research/0042-vulkan-async-pending-fence.md(new),docs/api/gpu.md,docs/backends/vulkan/overview.md,CHANGELOG.md,docs/rebase-notes.md.ffmpeg-patches/0006-libvmaf-add-libvmaf-vulkan-filter.patch— unchanged. The v2 ring is fully internal toVmafVulkanState; the public ABI stays byte-identical so the filter consumes the new path transparently.- Invariant 1 — fixed ring depth at first import.
lazy_alloc_ringis the only place that materialises the ring; once allocated the depth never changes for the lifetime of theVmafVulkanState. Any caller that needs a different depth has to free + re-init. The geometry pinning contract from v1 (ADR-0186) is preserved verbatim. - Invariant 2 —
vkResetFencesonly afterVK_SUCCESSfromvkWaitForFences. Sole reset path lives indrain_slot_fence;fence_in_flightflips back to 0 only after the wait succeeds. A-EIOfrom the wait propagates up without resetting (so a retry would correctly re-wait rather than silently move on). - Invariant 3 —
state_freedrains before destroying.vmaf_vulkan_import_slots_freewalks the ring and callsdrain_slot_fenceon every in-flight slot, then issues onevkQueueWaitIdlebelt-and-braces (any feature kernel that submitted on the same queue may still be running). Reordering this triggers validation-layer "destroying in-use object" errors. - Numerical contract: unchanged. Async submission only changes when the host can read the staging buffer, not which bytes the GPU writes. Cross-backend parity gate (
scripts/ci/cross_backend_parity_gate.py,places=4) holds. - Memory delta: staging arena scales
1 → ring_sizeper direction. At default depth and 1080p 8-bit Y, the per-state host-visible footprint grows from ~4 MiB to ~16 MiB. Documented in ADR-0251 §Consequences.
0090 — cambi_vulkan extractor (T7-36 / ADR-0210)¶
- ADR: ADR-0210; predecessor ADR-0205.
- Touches:
core/src/feature/vulkan/cambi_vulkan.c(replaces the spike scaffold'sinit_stub/extract_stub/close_stubtriple with the full Vulkan-aware lifecycle).core/src/feature/vulkan/shaders/cambi_preprocess.comp(new),cambi_mask_dp.comp(new — unified row-SAT / col-SAT / threshold-compare viaPASS=0/1/2spec const).core/src/feature/cambi.c— appends a small block of public trampolines (vmaf_cambi_*) at the bottom of the file that thinly wrap the file-static helpers. No upstream function-static code is renamed or moved; the entire upstream body of cambi.c above the trampolines stays byte-identical, which keeps Netflix sync straightforward.core/src/feature/cambi_internal.h(new) — internal-only header exposingvmaf_cambi_calculate_c_values,vmaf_cambi_get_spatial_mask, etc., to the GPU twin.core/src/vulkan/meson.build— registers the 5 cambi shaders invulkan_shader_sources[]andcambi_vulkan.cinvulkan_sources.core/src/feature/feature_extractor.c— adds the extern decl + registry entry forvmaf_fex_cambi_vulkanunder#if HAVE_VULKAN.scripts/ci/cross_backend_vif_diff.py—cambirow inFEATURE_METRICSso the cross-backend gate runs atplaces=4against the CPU baseline.docs/adr/0210-cambi-vulkan-integration.md,docs/research/0032-cambi-vulkan-integration.md,docs/backends/vulkan.md,CHANGELOG.md.- Invariant 1 — bit-exactness by construction. Every GPU phase is integer arithmetic (
uint16derivative,int32SAT,>compare, stride-2 gather, 3-elementmode3lookup). The readback into the hostVmafPicturepair is byte-identical to what the CPU would have written; the host residual then runs the unmodified CPUcalculate_c_values+ spatial pooling on those buffers. Any rebase that introduces float arithmetic into one of these GPU phases — e.g., a future Netflix change to the derivative kernel that adds a bilinear interpolation step — will silently breakplaces=4and must be caught at the cross-backend gate. - Invariant 2 —
cambi_internal.hsignatures must stay in lock-step with cambi.c's file-static helpers. The Vulkan twin callsvmaf_cambi_calculate_c_values, which trampolines to the file-staticcalculate_c_values. Any signature change to the latter (extra parameters, type changes) must update the trampoline + header in the same PR or the GPU build breaks. - On upstream sync: cambi.c's file-static helpers are sometimes renamed by upstream (e.g.,
decimate→cambi_decimatewould happen during a Netflix tidy-up). When rebasing, search cambi.c's tail for the trampoline block — its fivestaticcalls (get_spatial_mask,decimate,filter_mode,calculate_c_values,spatial_pooling,weight_scores_per_scale,get_pixels_in_window,increment_range,decrement_range,get_derivative_data_for_row,cambi_preprocessing) need to match the upstream symbol names. Update the trampoline body if upstream renames; signatures should not need to change because the trampoline already takes the function-pointer-typedef form (VmafRangeUpdateretc.). - Re-test on rebase:
python3 scripts/ci/cross_backend_vif_diff.py --backend vulkan --feature cambi --ref testdata/ref_576x324_48f.yuv --dist testdata/dis_576x324_48f.yuv --width 576 --height 324 --pixel-format 420 --bitdepth 8 --frames 48. Should emitplaces=4 PASSwithmax_abs_diff = 0.0. If it diverges, bisect the GPU phases by reading back individual buffers (image_buf/mask_buf/deriv_buf) and comparing against the CPU's in-placepicplane after the equivalent stage.
The pre-ADR-0108 fork-local PRs are summarised by workstream rather than per-PR. Future PRs add entries individually.
0085 — Upstream c70debb1 partial port (adm_csf + barten_csf tests)¶
- No ADR. Pure upstream cherry-pick per ADR-0108 carve-out ("pure upstream syncs and
port-upstream-commitPRs are exempt"). - Upstream source:
c70debb1(Kyle Swanson, 2026-04-28): "libvmaf/test: port new adm/vif/speed tests". The audit row that flagged the gap is T-NEW-2 in the 2026-04-29 quarterly upstream-backlog re-audit (PR #205). - Touches (additive only):
core/src/feature/adm_csf_tools.h— new header (verbatim from upstream); declares the inlineadm_native_csfhelper (DLM-paper CSF) used by the newtest_adm_csfunit.core/test/test_adm_csf.c— new unit (verbatim from upstream); 2mu_assertcases onadm_native_csf(3, 3.0, 1080, {0, 45}).core/test/test_barten_csf.c— new unit (verbatim from upstream); 23mu_assertcases overbarten_rod_cone_sens,barten_mtf,barten_csf,linear_interpolate,barten_watson_blend_csf(all symbols already on the fork).core/test/meson.build— registers the two new executables + addstest('test_adm_csf', ...)andtest('test_barten_csf', ...).CHANGELOG.mdUnreleased § Changed.- Deliberate scope cuts (the upstream commit's other halves are not portable verbatim):
test_vif_tools.c— depends on upstream symbolsNUM_KERNELSCALES, the 21-entryvalid_kernelscalestable,vif_validate_kernelscale,vif_get_filter_size,vif_get_filter,speed_get_antialias_filter, and a[NUM_KERNELSCALES][5][65]filter table that the fork'svif_filter1d_table_s [11][4][65]does not match. Per Research-0024 Strategy E, the fork deliberately diverges from the upstreamvifruntime-helper chain to preserve the ADR-0138 / 0139 / 0142 / 0143 SIMD bit-exactness contract. Porting this test requires porting the runtime helpers first.test_speed_chroma.c—#includesfeature/speed.cdirectly; the fork has no SpEED extractor (feature/speed.cdoes not exist). Pairs with audit row T-NEW-1 (port the SpEED extractor wholesale, or absorb it into the tiny-AI speed metric).- Invariants (rebase-relevant):
- The new
adm_csf_tools.hheader is wholly additive and does not conflict with the existing forkadm_csf_snon-inline helper inadm_tools.h(different signature, different translation units). - The two new tests do not depend on Netflix golden YUVs — they evaluate the closed-form CSF math directly. No golden-data interaction.
- On upstream sync: a future port of the upstream
vifruntime-helper chain (Research-0024 Strategy A reversal) or the SpEED extractor (T-NEW-1) unlocks the deferred halves of this commit. Until then, fork-sidetest_vif_tools.c/test_speed_chroma.cstay absent. - Re-test on rebase:
meson setup build-cpu libvmaf -Denable_cuda=false -Denable_sycl=false
ninja -C build-cpu test_adm_csf test_barten_csf
meson test -C build-cpu test_adm_csf test_barten_csf
0084 — Embedded MCP server scaffold (T5-2, ADR-0209)¶
- ADR: ADR-0209 (audit-first scaffold) on top of the ADR-0128 governance + Research-0005 design.
- Upstream source: fork-local. Netflix/vmaf has no embedded MCP server (and no plans to add one — the workflow is agent-tooling-specific, well outside upstream's library scope).
- Touches:
core/include/libvmaf/libvmaf_mcp.h— new public header.core/include/core/meson.build— newif get_option('enable_mcp')install branch.core/src/mcp/— new directory:mcp.c(stub TU) +meson.build(exposesmcp_sources+mcp_defines).core/src/meson.build— newis_mcp_enabledguard +subdir('mcp')block;mcp_sourcesthreaded into thelibrary('vmaf', ...)source list alongsidednn_sources.core/test/meson.build— newif get_option('enable_mcp')block wiringtest_mcp_smoke.core/test/test_mcp_smoke.c— new 12-sub-test smoke.core/meson_options.txt— newenable_mcpumbrella + three sub-flags (all defaultfalse).- Invariant: every public entry point in
libvmaf_mcp.h(vmaf_mcp_init/_start_sse/_start_uds/_start_stdio/_stop/_close) returns-ENOSYS(or-EINVALon bad arguments) until the T5-2b runtime PR lands. The smoke pins this contract — a runtime PR that flips a return code without flipping the smoke expectation regresses the gate. - On upstream sync: zero interaction with upstream files. Wholly additive directory + boolean build flags. The
subdir('mcp')insertion incore/src/meson.buildlives next to the existingsubdir('dnn')/ Vulkan blocks; an upstream conflict in that area would be confined to those few lines and is mechanical to resolve. - Re-test on rebase:
meson setup build-cpu libvmaf -Denable_cuda=false -Denable_sycl=false -Denable_mcp=false
ninja -C build-cpu && meson test -C build-cpu # baseline still green
meson setup --reconfigure build-cpu libvmaf -Denable_mcp=true \
-Denable_mcp_sse=true -Denable_mcp_uds=true -Denable_mcp_stdio=true
ninja -C build-cpu
meson test -C build-cpu test_mcp_smoke # 12/12 sub-tests pass
0065 — T7-37 Netflix bench rerun + docs/benchmarks.md TBD fill¶
- No ADR. Empirical fill of pre-existing
TBDcells; no new decision. The bench script fixes that this rerun depends on shipped earlier under PR #169 (libvmaf/AGENTS.md backend-engagement foot-guns), PR #170 (--backend cudaactually engages CUDA), and PR #171 (testdata/bench_all.shuses correct flags). Vulkan header install for SDK consumers is PR #175. - Touches (additive only):
docs/benchmarks.md(everyTBDcell replaced with measured numbers; hardware-profile table updated to theryzen-4090-archost the rerun was performed on; "How to reproduce" section now documents fixture acquisition for the gitignored BBB 4K 200-frame pair).CHANGELOG.mdUnreleased § Changed entry. - Invariants (rebase-relevant): none. The numbers are tied to fork commit
41301496and theryzen-4090-arcprofile; an upstream rebase that changes feature pipelines would invalidate the table but not break parsing. - On upstream sync: zero interaction. Pure docs.
- Re-test on rebase:
bash testdata/bench_all.sh(after a fresh fork build) — confirms the bench script drives every live backend and records each row's emitted metrics-key count. A GPU count collapsing to CPU is a fallback warning to corroborate with pool and throughput; never compare against fixed expected counts.
0050 — float_adm_cuda + float_adm_sycl extractors (ADR-0202)¶
- ADR: ADR-0202
- Touches:
core/src/feature/cuda/float_adm/float_adm_score.cu(new)core/src/feature/cuda/float_adm_cuda.{c,h}(new)core/src/feature/sycl/float_adm_sycl.cpp(new)core/src/meson.build— three changes: (1) newfloat_adm_scoreentry incuda_cu_sources, (2) newcuda_cu_extra_flagsdict that threads--fmad=false+-Xcompiler=-ffp-contract=offinto thefloat_adm_scorefatbin only, (3) new SYCL source insycl_feature_sources.core/src/feature/feature_extractor.c(extern decls + list entries forvmaf_fex_float_adm_cuda/vmaf_fex_float_adm_syclunder#if HAVE_CUDA/#if HAVE_SYCL).- Invariant 1 —
--fmad=falsefor the float_adm fatbin only: the angle-flag dot product (ot_dp = oh*th + ov*tv) and the cube reductions (xa*xa*xa,csf_o*csf_o*csf_o) require IEEE-754 add/mul ordering to match the GLSLprecisequalifier infloat_adm.comp. NVCC's default-fmad=truefuses these and drifts pastplaces=4at scale 3 / adm2. The integer ADM kernels sharecuda_flagsbut useint64accumulators where FMA is irrelevant — keep the FMA-on default for them. - Invariant 2 — parent-LL dimension trap: stage 0 at
scale > 0reads the parent's LL band; the mirror/clamp bounds arescale_w/h[scale](= parent's LL output dims = current scale's input dims), NOTscale_w/h[scale - 1](= parent's full image dims). Bothfloat_adm_cuda.candfloat_adm_sycl.cppcite this inline. Do not "simplify" by using the off-by-one neighbour. - Re-test:
CXX=icpx CC=icx meson setup build-cs -Denable_cuda=true \
-Denable_sycl=true -Denable_vulkan=enabled \
-Denable_float=true \
-Dsycl_compiler=/opt/intel/oneapi/compiler/latest/bin/icpx
ninja -C build-cs
python3 scripts/ci/cross_backend_vif_diff.py \
--vmaf-binary build-cs/tools/vmaf \
--reference python/test/resource/yuv/src01_hrc00_576x324.yuv \
--distorted python/test/resource/yuv/src01_hrc01_576x324.yuv \
--width 576 --height 324 --feature float_adm \
--backend cuda --places 4
# Same with --backend sycl on a host with an SYCL device.
# Both must report 0/N mismatches at places=4.
0049 — float_adm_vulkan extractor (ADR-0199)¶
- ADR: ADR-0199
- Touches:
core/src/feature/vulkan/float_adm_vulkan.c(new)core/src/feature/vulkan/shaders/float_adm.comp(new)core/src/vulkan/meson.build(adds the .comp shader and the new .c source)core/src/feature/feature_extractor.c(extern decl + list entry under#if HAVE_VULKAN)scripts/ci/cross_backend_vif_diff.py(float_admentry inFEATURE_METRICS).github/workflows/tests-and-quality-gates.yml(lavapipefloat_admstep atplaces=4)- Invariant: float_adm GPU port uses the
2 * sup - idx - 1mirror form on both axes — matches both the scalaradm_dwt2_sand the AVX2float_adm_dwt2_avx2, which both consume the samedwt2_src_indices_filt_sindex buffer. This is intentionally different from float_vif's GPU mirror (ADR-0197), which uses-2because float_vif's AVX2 path takes a different code branch. Do not "fix" the asymmetry by analogy with float_vif. - Re-test:
meson setup build-vk -Denable_vulkan=enabled -Denable_cuda=false \
-Denable_sycl=false
ninja -C build-vk
meson test -C build-vk
VK_LOADER_DRIVERS_SELECT='*lvp*' python3 \
scripts/ci/cross_backend_vif_diff.py \
--vmaf-binary build-vk/tools/vmaf \
--reference python/test/resource/yuv/src01_hrc00_576x324.yuv \
--distorted python/test/resource/yuv/src01_hrc01_576x324.yuv \
--width 576 --height 324 --feature float_adm --places 4
0083 — SSIMULACRA 2 Vulkan kernel (ADR-0201)¶
- ADR: ADR-0201
- Upstream source: fork-local. No SSIMULACRA 2 extractor in upstream Netflix/vmaf — fully fork-local feature.
- Touches:
core/src/feature/vulkan/ssimulacra2_vulkan.c(new file).core/src/feature/vulkan/shaders/ssimulacra2_xyb.comp,ssimulacra2_blur.comp,ssimulacra2_mul.comp,ssimulacra2_ssim.comp(4 new shader files).core/src/vulkan/meson.build— added 4 shaders tovulkan_shader_sourcesand 1 source tovulkan_sources; added all 4 ssimulacra2 shaders topsnr_hvs_strict_shaders(the-O0strict-mode list, kept its legacy name).core/src/feature/feature_extractor.c— registeredvmaf_fex_ssimulacra2_vulkanin the Vulkan branch of the extractor list (betweenpsnr_hvs_vulkanand the CUDA block).scripts/ci/cross_backend_vif_diff.py— addedssimulacra2toFEATURE_METRICS.- Rebase impact: low — fully additive, no upstream-shared files modified beyond
feature_extractor.c's registry array (which always grows on every new extractor and is not a rebase pain point). - Verification command:
meson setup core/build-vk-ss2 \
-Denable_vulkan=enabled -Denable_cuda=false -Denable_sycl=false \
libvmaf
ninja -C core/build-vk-ss2 tools/vmaf
python3 scripts/ci/cross_backend_vif_diff.py \
--vmaf-binary core/build-vk-ss2/tools/vmaf \
--reference python/test/resource/yuv/src01_hrc00_576x324.yuv \
--distorted python/test/resource/yuv/src01_hrc01_576x324.yuv \
--width 576 --height 324 \
--feature ssimulacra2 --backend vulkan --places 1
# expected: max_abs_diff ≈ 1.59e-2, 0/48 mismatches at places=1
- Follow-ups:
- CUDA + SYCL twins (batch 3 parts 7b + 7c per ADR-0192).
- Performance follow-up: re-bin multiple rows / columns per WG in the IIR blur (currently
local_size = 1, one row/col per WG for correctness). - Optional: rename
psnr_hvs_strict_shaderstostrict_shadersincore/src/vulkan/meson.build(cosmetic — out of scope for this PR).
0001 — SIMD bit-identical reductions for float ADM¶
- Workstream PRs: #18, commits
24c88a32,f082cfd3. - Touches:
core/src/feature/integer_adm.c,core/src/feature/float_adm.c,core/src/feature/x86/adm_avx2.c,core/src/feature/x86/adm_avx512.c,core/src/feature/arm64/adm_neon.c, upstreampython/test/feature_extractor_test.pytest expectations. - Invariant:
sum_cubeandcsf_den_scaleaccumulate cubed values in double precision (via_mm256_cvtps_pd/_mm512_cvtps_pd) in scalar, AVX2, AVX-512, and NEON. Upstream accumulates in float, which produces ~8e-5 drift between scalar and SIMD. Test expectations were tightened to match the double-precision path; an upstream-side accumulator change would re-introduce the drift and break the tightened assertions. - Re-test:
meson test -C build --suite=fast && python -m pytest python/test/feature_extractor_test.py -k adm.
0002 — CUDA ADM decouple-inline buffer elimination¶
- Workstream PRs: commit
787e3382. - Touches:
core/src/feature/cuda/integer_adm_cuda.cu,core/src/feature/cuda/adm_decouple_inline.cuh(new),core/src/feature/cuda/meson.build. Upstream'sadm_decouple.cuis no longer compiled in the fork. - Invariant: CSF and CM CUDA kernels read
ref/disDWT2 buffers directly and computedecouple_r/decouple_ainline via__device__helpers inadm_decouple_inline.cuh. The 6 intermediate buffers (decouple_r,decouple_a,csf_a× {scale-0 int16, scales 1-3 int32}) and the standaloneadm_decouple.cusource are intentionally removed. ~107 MB GPU memory savings at 4K. An upstream change toadm_decouple.cuwill look orphaned and a literal merge would re-introduce the buffer allocations. - Re-test:
meson setup build -Denable_cuda=true && ninja -C build && meson test -C build --suite=cuda.
0003 — SYCL backend (USM pool / D3D11 import / vmaf_sycl_* API)¶
- Workstream PRs: #33, #35, #5 (initial scaffolding), and the picture-pool deadlock fix that landed via #32.
- Touches:
core/include/libvmaf/libvmaf_sycl.h,core/src/sycl/,core/src/feature/sycl/,core/src/libvmaf.c(SYCL public-API entry points),meson_options.txt(enable_sycl). - Invariant:
vmaf_sycl_preallocate_picturesconstructs a realVmafSyclPicturePoolhonoringVmafSyclPicturePreallocationMethod(NONE/DEVICE/HOST);vmaf_sycl_picture_fetchdispatches to the pool when configured. The whole SYCL tree is fork-local and has no upstream counterpart — upstream changes tocore/src/libvmaf.cnear the SYCL entry-point block are likely to conflict. Picture-pool error paths invmaf_read_pictures(libvmaf.c) mustgoto cleanup;rather thanreturn err;to avoid leaking ref/dist pictures into the live-picture set (closes the always-on-pool deadlock fixed in #32 — see ADR-0104). See ADR-0101, ADR-0103, ADR-0104. - Re-test:
meson setup build -Denable_sycl=true && ninja -C build && meson test -C build --suite=sycl(requires oneAPI / icpx).
0004 — DNN runtime + tiny-AI surfaces¶
- Workstream PRs: #5, #8, #21, #22, #23, #31, #34, plus the pre-numbered DNN feat commits (
9b985946,1e5336d3,d122b721). - Touches:
core/include/libvmaf/dnn.h,core/src/dnn/,core/src/feature/feature_lpips.c,model/tiny/,meson_options.txt(enable_onnxruntime). - Invariant: ordered EP selection (CUDA → DML → CPU) with graceful fallback (ADR-0102);
fp16_iodoes host-side fp32↔fp16 cast on the scoring path;VMAF_TINY_MODEL_DIRenforces a path jail on model load (PR #31); the runtime op-allowlist (PR #21) walks the ONNX graph and rejects unknown ops + bounds Loop/Iftrip_countat 1024 (ADR-0036/0107). DNN tree is fork-local; upstream has no DNN code yet, so conflicts here are unlikely but themeson_options.txtandcore/src/meson.buildblocks near the DNN flag may collide. - Re-test:
meson setup build -Denable_onnxruntime=true && ninja -C build && meson test -C build --suite=dnn.
0005 — --precision CLI flag (IEEE-754 round-trip lossless)¶
- Workstream PRs: commit
c989fbd9. - Touches:
core/tools/vmaf.c,core/tools/cli_parse.c,core/include/libvmaf/libvmaf.h(addedvmaf_write_output_with_format),core/src/output.c. - Invariant: default
--precisionis%.17g(round-trip lossless);legacyopts back into upstream's%.6f; the public C API gainedvmaf_write_output_with_formatand the oldvmaf_write_outputroutes through it with the%.17gdefault. ABI-breaking only if upstream adds a same-named function with a different signature. See ADR-0006. - Re-test:
vmaf -r ref.yuv -d dis.yuv ... --precision=fulland diff against--precision=legacy.
0006 — Netflix golden tests preserved verbatim as required gate¶
- Workstream PRs: across the fork's life; codified in ADR-0024.
- Touches:
python/test/quality_runner_test.py,python/test/vmafexec_test.py,python/test/vmafexec_feature_extractor_test.py,python/test/feature_extractor_test.py,python/test/result_test.py,python/test/resource/yuv/. - Invariant:
assertAlmostEqual(...)golden values in the five upstream Python test files are never modified by this fork. Fork-added tests live in separate files (e.g.python/test/test_precision_flag.py). The CI gate "Netflix CPU golden tests (D24)" is required and blocks merge. Upstream changes to these files are accepted unless they relax the assertions. - Re-test:
make test-netflix-golden.
0007 — Build system (CUDA 13.2, oneAPI 2025.3, MkDocs migration)¶
- Workstream PRs: #7, #17, commit
8a995cb0. - Touches:
meson.build,meson_options.txt, top-levelMakefile,docs/(Sphinx → MkDocs Material migration —docs/conf.pyremoved,mkdocs.ymladded),docs/requirements.txt,Dockerfile.*, distro install scripts underscripts/. - Invariant: image pins are non-conservative (ADR-0027) — CUDA 13.2, oneAPI 2025.3, clang-format 22, black 26 — and ship experimental toolchain flags (
--expt-relaxed-constexpr, etc.) deliberately. An upstream sync that pulls in a Dockerfile change targeted at older CUDA or older oneAPI must not relax the pins. - Re-test:
meson setup build -Denable_cuda=true -Denable_sycl=true && ninja -C build && mkdocs build --strict.
0008 — Workspace / docs / MATLAB / resource-tree relocations¶
- Workstream PRs: codified across ADR-0026, ADR-0029, ADR-0030, ADR-0031, ADR-0032, ADR-0033, ADR-0034, ADR-0038.
- Touches: any path-walk in upstream's CI / scripts / docs that assumes the upstream layout (root-level
workspace/,resource/,matlab/, rootunittestscript, rootpatches/). - Invariant: the fork's layout is
python/vmaf/workspace/,python/vmaf/resource/,python/vmaf/matlab/,scripts/unittest,ffmpeg-patches/only,.github/codeql-config.yml. Upstream moves to a different sub-tree (e.g. a hypotheticaltools/workspace/) need to either be applied via a corresponding fork-side relocation or rejected with a rebase note. - Re-test:
python -m pytest python/test/ -k golden(verifies the resource-tree path works);make test-netflix-golden.
0009 — License headers (Lusoris/Claude on wholly-new files¶
2016–2026 on Netflix files)
- Workstream PRs: commits
c159761d,a185f8ef,0e98c949, codified in ADR-0025 / ADR-0105. - Touches: every wholly-new fork file (notably the SYCL tree and
core/src/dnn/) and every Netflix-touched file (year range2016 → 2016–2026). - Invariant: wholly-new fork files carry
Copyright 2026 Lusoris and Claude (Anthropic)under the same BSD-3-Clause-Plus-Patent license; mixed files use a dual-copyright notice. An upstream commit that resets a Netflix file's year range (e.g. back to2016–2020) must be partially rejected — keep the fork's2016–2026. - Re-test: grep that wholly-new fork files retain the Lusoris/Claude header (
grep -L "Copyright 2026 Lusoris" core/src/sycl/*.cpp— expected to match nothing).
0010 — .claude/ agent scaffolding + ADR tree + AGENTS.md / CLAUDE.md¶
- Workstream PRs: #14, #24, #37, plus continuous additions.
- Touches:
.claude/,AGENTS.md,CLAUDE.md,docs/adr/,.github/PULL_REQUEST_TEMPLATE.md. - Invariant: this whole tree is fork-local and has no upstream counterpart. Upstream additions to
.github/(issue templates, workflows) need to merge cleanly with the fork's existing files rather than replacing them. The ADR tree's IDs ≤ 0099 are backfills; new decisions start at 0100 (ADR-0028 / ADR-0106). - Re-test: visual review of
.github/anddocs/adr/README.mdafter the merge.
Pre-ADR-0108 entries above are the result of a one-shot backfill sweep on 2026-04-18; subsequent fork-local PRs add their own entries inline.
0011 — Nightly bisect-model-quality + fixture cache¶
- Workstream PRs: closes #4; sticky tracker issue #40.
- Touches:
.github/workflows/nightly-bisect.yml,ai/scripts/build_bisect_cache.py,ai/testdata/bisect/{features.parquet, models/*.onnx, README.md},scripts/ci/post-bisect-comment.py,docs/ai/bisect-model-quality.md,docs/adr/0109-nightly-bisect-model-quality.md,docs/research/0001-bisect-model-quality-cache.md,mkdocs.yml(nav). - Invariant: the committed parquet + ONNX bytes under
ai/testdata/bisect/must regenerate byte-identically fromai/scripts/build_bisect_cache.pywith seedsFEATURE_SEED=20260418andMODEL_SEED=20260419. The CI--checkstep asserts this before every bisect run, so any upstream pull that bumpspandas/pyarrow/onnxenough to change the serialiser bytes will fail the workflow until the cache is regenerated and committed. - Re-test:
python ai/scripts/build_bisect_cache.py --check
vmaf-train bisect-model-quality \
ai/testdata/bisect/models/model_*.onnx \
--features ai/testdata/bisect/features.parquet \
--min-plcc 0.85 --input-name input
# Expected: "no regression in this range"; first_bad_index None.
Pure upstream code is not touched, so no Netflix-side conflict vector. Only fork-local files; risk is toolchain drift, not merge conflict.
0012 — Upstream ADM port (Netflix 966be8d5)¶
- Workstream PRs: this PR; ports a single upstream commit.
- Touches:
core/src/feature/integer_adm.{c,h},core/src/feature/x86/adm_avx2.{c,h},core/src/feature/x86/adm_avx512.{c,h},core/src/feature/alias.c,core/src/feature/barten_csf_tools.h(new upstream file). - Invariant: the eight ADM files now mirror upstream's content byte-for-byte (modulo our clang-format-22 pass and the Netflix copyright-year bump on the new header). Future
/sync-upstreamruns can take new upstream ADM commits cleanly. Do not revert to a pre-966be8d5ADM kernel without also reverting the call-site signatures ininteger_compute_adm— upstream extendedi4_adm_cmfrom 8 to 13 args. - Re-test:
ninja -C core/build && meson test -C core/build
core/build/tools/vmaf -r python/test/resource/yuv/src01_hrc00_576x324.yuv \
-d python/test/resource/yuv/src01_hrc01_576x324.yuv \
-w 576 -h 324 -p 420 -b 8 \
--model version=vmaf_v0.6.1 -o /tmp/vmaf-port.json
grep '<metric name="vmaf"' /tmp/vmaf-port.json
# Expected: mean ≈ 76.66890 (golden 76.66890519623612, places=4 OK).
0013 — Upstream motion port (Netflix PR #1486 head 2aab9ef1)¶
- Workstream PRs: this PR; ports upstream PR #1486 (4 commits on top of
966be8d5ADM base, head2aab9ef1). Sister to entry 0012. - Touches:
core/src/feature/integer_motion.{c,h},core/src/feature/motion_blend_tools.h(new upstream file),core/src/feature/x86/motion_avx2.c,core/src/feature/x86/motion_avx512.c,core/src/feature/alias.c(additive:integer_motion3row),python/test/{quality_runner,vmafexec,feature_extractor,vmafexec_feature_extractor}_test.py(golden tolerance updates:places=4→places=2on motion-affected asserts; expected values unchanged). - Invariant: motion files mirror upstream byte-for-byte (modulo our clang-format-22 pass). The
alias.crow forinteger_motion3was inserted surgically to avoid clobbering the AVX-512 ADM registration added by entry 0012; new motion3 metric appears in default VMAF model output but is not standalone-loadable via--feature integer_motion3(sub-feature only). Netflix golden VMAF mean shifts76.668904824→76.667830213(well withinplaces=2tolerance the upstream PR loosened to). Do not revertplaces=4on motion-touching assertions without also reverting the motion code. - Re-test:
ninja -C core/build && meson test -C core/build
core/build/tools/vmaf -r python/test/resource/yuv/src01_hrc00_576x324.yuv \
-d python/test/resource/yuv/src01_hrc01_576x324.yuv \
-w 576 -h 324 -p 420 -b 8 \
--model version=vmaf_v0.6.1 -o /tmp/vmaf-motion-port.json
grep -E '<metric name="vmaf"|integer_motion3' /tmp/vmaf-motion-port.json
# Expected: vmaf mean ≈ 76.66783; integer_motion3 mean ≈ 3.98976.
0014 — Coverage gate overhaul + upstream python/test/ reformat¶
- Workstream PRs: this PR (coverage-gate overhaul + in-tree reformat of upstream-mirror Python tests).
- Touches:
.github/workflows/ci.yml(CPU + GPU coverage jobs:-Dc_args=-fprofile-update=atomic/-Dcpp_args=-fprofile-update=atomic,meson test --num-processes 1,-Denable_dnn=enabled, ORT install step on the CPU coverage job,lcov/geninforeplaced bygcovrwith--json-summary/--xml/--txtoutput, artifact renamecoverage-lcov-{cpu,gpu}→coverage-{cpu,gpu}),scripts/ci/coverage-check.sh(rewritten to parse gcovr JSON viapython3 -c— same CLI signature),core/src/dnn/dnn_api.c+ newcore/src/dnn/dnn_attach_api.c(vmaf_use_tiny_modelcarved out into its own TU so the unit-test binaries — which pull indnn_sourcesforfeature_lpips.cbut never linklibvmaf.c— don't end up with an undefined reference tovmaf_ctx_dnn_attachonceenable_dnn=enabledactivates the real bodies),core/src/dnn/meson.build+core/src/meson.build(newdnn_libvmaf_only_sourceslist wired intolibvmaf.soonly),python/test/{feature_extractor,quality_runner,vmafexec,vmafexec_feature_extractor}_test.py(mechanical Black + isort reformat — no assertion values changed, imports regrouped, line wrapping normalised). - Invariant: coverage CI must keep all five pieces in lockstep — (a)
-fprofile-update=atomiccloses the intra-process counter race on SIMD inner loops (vif_avx2.c:673,motion_avx2, etc.) → negative counts →geninfo/gcovr abort; (b)--num-processes 1closes the inter-process race where multiple parallel test binaries merge their counters into the same.gcdafiles for the sharedlibvmaf.soat process exit (per-thread atomicity does not cover this); (c)gcovrdeduplicates.gcnofiles belonging to the same source compiled into multiple targets — without dedup, lcov sums hits across compilation units and yields impossible100% values (
dnn_api.c — 1176%was the smoking gun on the first attempt that had only (a)+(b)); (d) ORT install +enable_dnn=enabledin the coverage job is what makescore/src/dnn/*.cmeasurable in the first place — without ORT, the DNN tree compiles in stub branches and the 85% per-critical-file gate is meaningless; (e)vmaf_use_tiny_modellives indnn_attach_api.cand is added tolibvmaf.soonly viadnn_libvmaf_only_sources— moving it back intodnn_api.creintroduces thevmaf_ctx_dnn_attachundefined-reference link error intest_feature_extractor/test_lpipswheneverenable_dnn=enabled, since those test binaries pull indnn_sourcesforfeature_lpips.cbut never linklibvmaf.c. Lint scope: upstream-mirror Python tests are linted at the same standard as fork-added code; we accept that/sync-upstreamand/port-upstream-commitwill re-trigger Black/isort failures whenever upstream rewrites these files, and the fix is another in-tree reformat pass — never an exclusion. The fork'spyproject.tomland.pre-commit-config.yamlkeeppython/test/resource/(binary fixtures only) excluded;python/test/*.pyis in scope. See ADR-0110 (race fixes, superseded) and ADR-0111 (gcovr + ORT layer). - Re-test:
# Reproduce coverage path locally (requires gcc + python3-pip):
pip install --user 'gcovr>=8.0'
cd libvmaf
meson setup build-cov-test --buildtype=debug -Db_coverage=true \
-Denable_avx512=true -Denable_float=true -Denable_dnn=disabled \
-Dc_args=-fprofile-update=atomic -Dcpp_args=-fprofile-update=atomic
ninja -C build-cov-test
meson test -C build-cov-test --print-errorlogs --num-processes 1
~/.local/bin/gcovr --root .. \
--filter 'src/.*' \
--exclude '.*/test/.*' --exclude '.*/tests/.*' \
--exclude '.*/subprojects/.*' \
--gcov-ignore-parse-errors=negative_hits.warn \
--gcov-ignore-parse-errors=suspicious_hits.warn \
--print-summary --txt build-cov-test/coverage.txt \
--json-summary build-cov-test/coverage.json \
build-cov-test
grep -E 'dnn_api|model_loader' build-cov-test/coverage.txt
# Expected: gcovr completes without "Unexpected negative count" AND no
# per-file percentages exceed 100% (drop --num-processes 1 to reproduce
# the multi-process .gcda merge race; switch back to lcov to reproduce
# the dnn_api.c — 1176% over-count from compilation-unit summation).
# Lint smoke test for upstream-mirror tree:
pre-commit run --files python/test/quality_runner_test.py
# Expected: Black/isort/Ruff all PASS — files are reformatted in-tree
# to fork style and stay clean until the next upstream sync.
0015 — Tox doctest collection skips vmaf/resource/¶
- Workstream PRs: this PR (
fix(ci): skip pytest doctest collection of vmaf/resource/ data files). Surfaced once ADR-0115 consolidated CI triggers tomasterand tox actually started running on PRs. - Touches:
python/tox.ini(single-line--ignore=vmaf/resourceadded to the pytest invocation, plus an explanatory comment block). Pure fork-local; no upstream Python file changes. - Invariant:
pytest --doctest-modulesmust not attempt to import files underpython/vmaf/resource/. Those are parameter / dataset / example-config.pyfiles; several have dots in their stems (e.g.vmaf_v7.2_bootstrap.py) that make them unimportable as Python modules. None carry doctests, so the ignore is correctness rather than a workaround. Do not drop the--ignore=vmaf/resourceflag without first verifying every file under that directory has been renamed to a dot-free stem and is importable. - Re-test:
cd python && tox -e py311 -- --collect-only --doctest-modules \
--ignore=vmaf/resource 2>&1 | grep -c "ERROR collecting vmaf/resource"
# Expected: 0 (was 5 before the fix).
Pure upstream code is not touched, so no Netflix-side conflict vector. Risk is upstream renaming or removing files under python/vmaf/resource/ such that the directory disappears, in which case the --ignore becomes a harmless no-op.
0016 — SYCL -fsycl link-arg gated on icpx CXX¶
- Workstream PRs: this PR (
fix(libvmaf): gate -fsycl link arg on icpx CXX, allow gcc/clang host linker). Surfaced once ADR-0115's CI consolidation added an Ubuntu SYCL job to PR-time CI that usesCXX=g++(host linker) with sidecar icpx for SYCL .cpp compilation. - Touches:
core/src/meson.build(thevmaf_link_argsblock immediately after theis_sycl_enabledflag handling — currently ~lines 696-712). Pure fork-local; no upstream Meson file changes expected. - Invariant:
-fsyclis appended tovmaf_link_argsonly whenmeson.get_compiler('cpp').get_id() == 'intel-llvm'(icpx). Rationale: the documented project mode (see comment nearis_sycl_enabledblock at top ofsrc/meson.build) compiles SYCL.cppfiles viacustom_targetwith icpx, while the project's CXX driver may be gcc / clang / msvc; in that mode the SPIR-V device code is already embedded in the icpx-compiled.ofiles at compile time, and the runtime libraries (libsycl+libsvml+libirc+libze_loader) declared as link dependencies resolve every symbol. Passing-fsyclto a non-icpx linker is a hard error (g++: error: unrecognized command-line option '-fsycl'). Do not remove thecpp.get_id() == 'intel-llvm'guard without first verifying every CI matrix leg uses icpx as the project CXX. - Re-test:
meson setup build -Denable_sycl=true \
-Dcpp_link_args=-Wl,--no-undefined
ninja -C build src/libvmaf.so.3
# Expected: link succeeds; no `-fsycl` errors with gcc/clang host CXX.
Pure fork-local guard; no Netflix-side conflict vector.
0017 — CLI precision default %.6f (Netflix-compat) + frame-skip unref¶
- Workstream PRs: this PR (
fix(cli): revert precision default to %.6f and unref skipped frames). Reverts the default flipped by commitc989fbd9(ADR-0006) per ADR-0119. Companion fix incore/tools/vmaf.cresolves the picture-pool exhaustion in the--frame_skip_ref/distloops surfaced once the always-on picture pool (ADR-0104) made unref'ing skipped pictures mandatory. - Touches:
core/tools/cli_parse.c(VMAF_DEFAULT_PRECISION_FMT+VMAF_LOSSLESS_PRECISION_FMTmacros,resolve_precision_fmt()body,--helptext)core/tools/cli_parse.h(field comments only; struct shape unchanged)core/src/output.c(DEFAULT_SCORE_FORMATmacro)core/tools/vmaf.c(skip loop bodies at thec.frame_skip_ref/c.frame_skip_distfor-loops)python/vmaf/core/result.py(per-frame and aggregate:.6fformatters)python/test/command_line_test.pyis unmodified — Netflix golden assertions stay frozen per CLAUDE.md §8; the binary's output format adapts to them, not the other way around.- Invariant:
vmafCLI default score-output format is%.6f(matches upstream Netflix byte-for-byte).--precision=max|fullselects%.17g(IEEE-754 round-trip lossless).--precision=legacyis a synonym for the default. The library default forvmaf_write_output_with_format(..., score_format=NULL)matches. Skipped frames in the--frame_skip_ref/--frame_skip_distpre-loops arevmaf_picture_unref'd immediately after fetch so the preallocated picture pool is not exhausted before the main scoring loop runs. Do not flip the macros back to%.17gor remove the unrefs without a superseding ADR — both are golden-gate-load-bearing. - Re-test:
ninja -C core/build
python -m pytest python/test/command_line_test.py \
::VmafexecCommandLineTest::test_run_vmafexec \
::VmafexecCommandLineTest::test_run_vmafexec_with_frame_skipping \
::VmafexecCommandLineTest::test_run_vmafexec_with_frame_skipping_unequal \
-v
# Expected: all three PASS in <1 s combined.
Pure fork-local; no Netflix-side conflict vector. If upstream ever changes the default format string, treat their value as the new baseline and reconfirm the golden assertions before adopting.
0018 — FFmpeg patches ship as ordered series.txt¶
- Workstream PRs: this PR (
fix(ci): drop dead sycl trigger + consolidate windows.yml into libvmaf.yml (ADR-0115)). Surfaced once ADR-0115's consolidation routed the docker / FFmpeg-SYCL jobs through the master-targeting CI gate for the first time on this branch — the standalone0003-…sycl…apply broke because it referenced struct fields added by0001-…tiny-model…, the Dockerfile onlyCOPY'd 0003, andffmpeg.ymlreferenced a stale../patches/path. - Touches:
Dockerfile(lines ~86-95 — the FFmpeg patch-apply block),.github/workflows/ffmpeg.yml(theBuild FFmpeg with SYCL patch seriesstep),ffmpeg-patches/000{1,2,3}-*.patch(regenerated via realgit format-patch -3so they carry validindex <sha>..<sha> <mode>lines and committable SHAs). Pure fork-local; no upstream FFmpeg or Netflix file changes. - Invariant: both the Dockerfile and
ffmpeg.ymlwalkffmpeg-patches/series.txtline-by-line and apply each patch viagit applywith apatch -p1fallback. Do not ship a new patch without appending it toseries.txt, and do not reorder existing entries — patch 0003 references LIBVMAFContext fields added by patch 0001, so any out-of-order apply breaks the build at hunk 2 of vf_libvmaf.c. - Two flag-side fixes bundled in the same PR:
--enable-libvmaf-syclis not a valid FFmpeg configure option. Patch 0003 usescheck_pkg_config libvmaf_sycl …auto-detection (matching howlibvmaf_cudais wired) — it never registers the switch. Both Dockerfile and ffmpeg.yml used to pass the flag and configure rejected it withUnknown option "--enable-libvmaf-sycl". SYCL support is now controlled solely by-Denable_sycl=trueat libvmaf build time; FFmpeg picks it up automatically whenlibvmaf-sycl.pcis onPKG_CONFIG_PATH.- The Dockerfile now carries two nvcc-flag ARGs.
NVCC_FLAGS(libvmaf) keeps four-gencodelines plus the experimental--extended-lambda/--expt-relaxed-constexpr/--expt-extended-lambdaflags needed for Thrust/CUB host+device code.FFMPEG_NVCC_FLAGS(FFmpeg) carries a single-gencode arch=compute_75,code=sm_75 -O2— FFmpeg'scheck_nvccrunsnvcc -ptx, which fails withnvcc fatal: Option '--ptx (-ptx)' is not allowed when compiling for multiple GPU architectureson multi-arch input, and--extended-lambdarequires host+device compilation. compute_75 PTX is forward-compatible with all newer GPUs via driver JIT. --enable-libnppis no longer passed to FFmpeg's configure. FFmpeg n8.1's libnpp probe carries an explicitdie "ERROR: libnpp support is deprecated, version 13.0 and up are not supported"(configure:7335-7336) that fires on the base image's CUDA 13.2 libnpp. We don't use scale_npp / transpose_npp / sharpen_npp in any VMAF workflow; cuvid + nvdec + nvenc + libvmaf-cuda is the actual GPU path. Revisit once we move to an FFmpeg release that supports CUDA 13 libnpp upstream.- Patch 0002 (
add-vmaf_pre-filter) gained a missing#include "libavutil/imgutils.h"forav_image_copy_plane(). FFmpeg's libavfilter Makefile builds with-Werror=implicit-function-declarationso this fired during the actual compile (not configure). Caught by a localdocker buildrather than waiting for GitHub Actions — much faster iteration loop. - Re-test:
cd /tmp && rm -rf ffmpeg-test && \
git clone -q --depth 1 -b n8.1 \
https://git.ffmpeg.org/ffmpeg.git ffmpeg-test && \
cd ffmpeg-test && \
while IFS= read -r line; do \
case "$line" in ''|\#*) continue ;; esac; \
git apply "/path/to/vmaf/ffmpeg-patches/$line" \
|| patch -p1 < "/path/to/vmaf/ffmpeg-patches/$line"; \
done < /path/to/vmaf/ffmpeg-patches/series.txt
# Expected: all three patches apply with no rejects; the resulting
# tree compiles with --enable-libvmaf. SYCL is auto-detected via
# check_pkg_config (patch 0003), so no explicit configure flag is
# required when libvmaf-sycl.pc is on PKG_CONFIG_PATH.
Pure fork-local series; no Netflix-side conflict vector. See ADR-0118.
0019 — Coverage Gate annotations: upload-artifact v7 + gcovr filter¶
- Workstream PRs: this PR.
- Touches:
.github/workflows/ci.yml(CPU + GPU coverage steps: gcovr stderr piped throughgrep -vE 'Ignoring (suspicious|negative) hits' ... || true),.github/workflows/{ci,lint,nightly,nightly-bisect,supply-chain,libvmaf}.yml(actions/upload-artifact@v5|@v6 → @v7,actions/download-artifact@v5 → @v7insupply-chain.yml). Note:windows.ymlwas consolidated intolibvmaf.ymlby ADR-0115 / PR #50, so the windows-side bump now lives inlibvmaf.yml'sbuild (MINGW64, …)job. - Invariant: Coverage Gate Annotations panel must finish empty on a clean run. The two pieces are coordinated — (a)
@v7for upload / download artifact actions silences GitHub's Node-20 deprecation banner ahead of the 2026-06-02 forced-Node-24 cutoff; (b) the gcovr stderr filter swallows theIgnoring (suspicious|negative) hitswarnings that gcovr 8 emits for the legitimately-large hit counts in tight ANSNR / VIF / motion inner loops (e.g.ansnr_tools.c:207at ~4.93 G hits across an HD multi-frame coverage suite — real, not gcov bug). The filter is regex-narrow and anchored to gcov's exact warning prefix; any other gcovr warning still surfaces. Upstream (Netflix/vmaf) does not maintain these CI files; rebase impact is limited to the unlikely case that an upstream sync touches the shared.github/workflows/tree, which it currently does not. See ADR-0117. - Re-test:
# Verify gcovr filter locally (after a coverage build per entry 0014):
~/.local/bin/gcovr --root .. \
--filter 'src/.*' \
--exclude '.*/test/.*' --exclude '.*/tests/.*' \
--exclude '.*/subprojects/.*' \
--gcov-ignore-parse-errors=negative_hits.warn \
--gcov-ignore-parse-errors=suspicious_hits.warn \
--print-summary --txt build-cov-test/coverage.txt \
build-cov-test \
2> >(grep -vE 'Ignoring (suspicious|negative) hits' >&2 || true)
# Expected: stderr contains the gcovr summary block but NO
# "Ignoring (suspicious|negative) hits" lines. coverage.txt unchanged.
# Verify all upload/download-artifact instances are on @v7:
grep -rE 'actions/(upload|download)-artifact@v[0-6]' .github/workflows/
# Expected: empty output.
0020 — CI workflow file + display-name renames (Title Case sweep)¶
- Workstream PRs: this PR; renames all six core
.github/workflows/*.ymlfiles to purpose-descriptive kebab-case and normalises every workflowname:and jobname:to Title Case. See ADR-0116. - Touches:
.github/workflows/{ci,lint,security,libvmaf,ffmpeg,docker}.yml(renamed viagit mvtotests-and-quality-gates.yml,lint-and-format.yml,security-scans.yml,libvmaf-build-matrix.yml,ffmpeg-integration.yml,docker-image.yml),README.md(5 badge URLs + labels),docs/principles.md(line 5 workflow-tuple update),.claude/skills/add-gpu-backend/SKILL.md+scaffold.sh(filename refs),docs/adr/0116-*.md(new),docs/adr/README.md(index row),CHANGELOG.md. - Invariant: workflow files are purpose-named; their
name:fields are Title Case sentences with em-dash axis tags; job-levelname:strings are Title Case sentences (Build — / Pre-Commit / Coverage Gate / etc.). Required-status-check contexts inmasterbranch protection are bound to job-level names — when renaming any job, re-pin viagh api --method PUT repos/VMAFx/vmafx/branches/master/protection. The 19 required gates' semantics are unchanged from ADR-0037; only their display strings move. - Re-test:
# Validate every workflow file parses and lists the expected job names.
cd .github/workflows
for f in tests-and-quality-gates.yml lint-and-format.yml security-scans.yml \
libvmaf-build-matrix.yml ffmpeg-integration.yml docker-image.yml; do
yq '.name, .jobs.[].name' "$f" || echo "PARSE FAIL: $f"
done
# Expected: each workflow prints its Title Case workflow name + job names;
# no PARSE FAIL lines.
0021 — DNN-enabled CI matrix legs (gcc + clang + macOS)¶
- Workstream PRs: this PR; adds three new entries to the
libvmaf-buildmatrix in.github/workflows/libvmaf-build-matrix.ymlcovering-Denable_dnn=enabledacross Ubuntu/gcc, Ubuntu/clang, and macOS/clang. See ADR-0120. - Touches:
.github/workflows/libvmaf-build-matrix.yml(3 new matrix entries + ORT install steps + dedicated dnn-suite test step),docs/adr/0120-ai-enabled-ci-matrix-legs.md(new),docs/adr/README.md(index row),CHANGELOG.md(Added entry). - Invariant: the DNN matrix legs install ONNX Runtime via the same pinned source as the dedicated Tiny AI job (tests-and-quality-gates.yml) — Linux: MS tarball at the version pinned by
ORT_VERSION; macOS: Homebrew. When the Tiny AI job's pin changes, the matrix legs'ORT_VERSIONenv in theirInstall ONNX Runtime (linux, DNN leg)step must change to match; otherwise compiler/portability coverage drifts away from the gating leg's actual ABI. - Re-test:
# Local sanity: the matrix file parses and the new job names exist.
yq '.jobs.libvmaf-build.strategy.matrix.include[] | select(.dnn==true) | .name' \
.github/workflows/libvmaf-build-matrix.yml
# Expected output (3 lines):
# Build — Ubuntu gcc (CPU) + DNN
# Build — Ubuntu clang (CPU) + DNN
# Build — macOS clang (CPU) + DNN
# Local DNN build sanity (matches what each leg will run):
meson setup libvmaf core/build --buildtype release \
--prefix $PWD/install -Denable_float=true -Denable_dnn=enabled
ninja -vC core/build install
meson test -C core/build --suite=dnn --print-errorlogs
- Branch protection: the two Linux DNN legs are pinned as required status checks on
masterimmediately after this PR's merge (19 → 21 contexts). The macOS leg stays informational (experimental: true) because Homebrew ORT floats. Re-pin command:
gh api --method PUT repos/VMAFx/vmafx/branches/master/protection \
--input /tmp/protection-update.json
0022 — Windows GPU build-only matrix legs (MSVC + CUDA, MSVC + oneAPI SYCL)¶
- Workstream PRs: this PR; adds a new top-level
windows-gpu-buildjob to.github/workflows/libvmaf-build-matrix.ymlwith two matrix entries (CUDA, SYCL). See ADR-0121. - Touches:
.github/workflows/libvmaf-build-matrix.yml(newwindows-gpu-buildjob),docs/adr/0121-windows-gpu-build-only-legs.md(new),docs/adr/README.md(index row),CHANGELOG.md(Added entry),core/src/compat/win32/pthread.h(new — Win32 pthread shim for MSVC; mirrorscompat/gcc/stdatomic.hpattern),core/src/feature/integer_adm.h(UPSTREAM — converted thedwt_7_9_YCbCr_threshold[3]designated initializer to positional form so MSVC/nvcc-on-Windows accepts the C++ parse; semantically identical, no behavioural change),core/src/ref.handcore/src/feature/feature_extractor.h(UPSTREAM — added#if defined(__cplusplus) && defined(_MSC_VER)branch around#include <stdatomic.h>so MSVC C++ TUs pullatomic_intviausing std::atomic_int;; POSIX paths unchanged),core/src/sycl/d3d11_import.cpp(fix non-existent<libvmaf/log.h>→"log.h"),core/src/sycl/dmabuf_import.cpp(move<unistd.h>inside#if HAVE_SYCL_DMABUFguard for non-VA-API hosts),core/src/sycl/common.cpp(replace POSIXclock_gettime(CLOCK_MONOTONIC)with portablestd::chrono::steady_clock),core/src/feature/x86/motion_avx2.c(UPSTREAM — replace GCC vector-extension__m256i[N]indexing at line 529 with_mm256_extract_epi64; bit-exact),core/src/feature/x86/adm_avx2.c(UPSTREAM — replace 6(__m256i)(_mm256_cmp_ps(...))casts with_mm256_castps_si256(...)and 12__m128i[N]reductions with_mm_extract_epi64; bit-exact),core/src/feature/x86/adm_avx512.c(UPSTREAM — replace 12__m128i[N]reductions with_mm_extract_epi64; bit-exact),core/src/log.c(UPSTREAM — gate<unistd.h>behind!_WIN32, include<io.h>+ redirectisatty/filenoto_isatty/_filenofor MSVC),core/src/feature/integer_vif.c(UPSTREAM — switch thealigned_malloccursor fromvoid *touint8_t *with explicit typed-pointer casts so MSVC accepts the byte-wise pointer arithmetic),core/src/feature/cuda/integer_adm_cuda.c(UPSTREAM — drop unused<unistd.h>include),core/src/dnn/model_loader.c(fork-added — Windows fallback definitions for POSIXS_ISDIR/S_ISREGpath-classification macros),.github/workflows/lint-and-format.yml(fork-added — setlfs: trueon the pre-commit job's checkout so LFS-stored ONNX blobs resolve and don't appear as phantom pre-commit-induced diffs),core/src/feature/x86/motion_avx512.c(UPSTREAM — replace 1__m128i[N]reduction with_mm_extract_epi64; bit-exact),core/src/feature/x86/{vif_statistic_avx2,ansnr_avx2,ansnr_avx512,float_adm_avx2,float_adm_avx512,float_psnr_avx2,float_psnr_avx512,ssim_avx2,ssim_avx512}.c(UPSTREAM — convert 17 sites of trailing__attribute__((aligned(N)))to leading C11_Alignas(N); same alignment, MSVC-portable),core/src/feature/mkdirp.candcore/src/feature/mkdirp.h(UPSTREAM third-party MIT-licensed micro-library — gate<unistd.h>to non-Windows, add<direct.h>+_mkdirfor Windows, addmode_ttypedef for MSVC),core/meson.build(newpthread_dependencygated oncc.check_header('pthread.h')failing),core/src/meson.buildandcore/test/meson.build(threadpthread_dependencyinto every target compiling pthread-using TUs). - Invariant: Windows GPU legs are pinned to the same toolchain versions as the corresponding Linux GPU legs (CUDA 13.0.0, oneAPI BaseKit 2025.3.0.372) so a Linux-vs-Windows divergence implies an MSVC ABI issue, not a tooling-version delta. When either Linux GPU leg bumps its toolchain, the Windows leg must move in lockstep — the Intel installer URL on Windows hard-codes the per-release directory id and the version string, so the bump is two-line edits in the SYCL
Install Intel oneAPI (windows)step (theWINDOWS_BASEKIT_URLenv var). Both legs additionally inject/experimental:c11atomicsintoCFLAGS/CXXFLAGSbecause libvmaf uses C11 atomics that MSVC's<stdatomic.h>rejects without that opt-in flag — when MSVC ships full C11 atomics support, the flag becomes unconditional and can be dropped. Two Windows-only dependency steps round out the parity: the CUDA leg'sJimver/cuda-toolkitsub-package list includes bothcrt(CUDA Runtime Library compile-time headers, shipscrt/host_config.h;cuda_ccclis not a valid Windows sub-package name — installer rejects it) andnvvm(shipsnvvm/bin/cicc.exe+nvvm/libdevice/libdevice.*.bc; without it, nvcc's.cu → PTXstage fails withThe system cannot find the path specified.— on Linux apt pulls NVVM in transitively withcuda-nvcc-XY, Windows requires it explicitly); the SYCL leg builds the Level Zero loader from source (oneapi-src/level-zerov1.18.5 →cmake --build … --target install) because Windows oneAPI BaseKit ships the SYCL runtime but notze_loader.lib, and libvmaf's mesoncc.find_library('ze_loader')needs both the header and the import library. When the Linux aptlevel-zero-devversion moves, bump the L0 git tag to match.core/src/meson.buildguards the explicitsvml/irccc.find_librarycalls behindhost_machine.system() != 'windows'— those calls exist for the gcc/g++ + icpx Linux flow where the host linker is non-Intel; on Windows the host compiler is icx-cl itself and auto-injects the Intel runtime. Round-10 surfaced an additional Windows-only gap: ~14 libvmaf TUs#include <pthread.h>unconditionally, but MSVC and clang-cl ship no pthread (MinGW does, via winpthreads). The fork now ships a header-only Win32 shim atcore/src/compat/win32/pthread.hmapping the in-use pthread subset (mutex / cond / thread create+join+detach) onto SRWLOCK + CONDITION_VARIABLE +_beginthreadex. The shim is wired in viapthread_dependencyincore/meson.build, declared only whencc.check_header('pthread.h')fails — so MinGW and POSIX paths stay untouched. When upstream Netflix/vmaf adds new pthread surface (e.g.,pthread_rwlock_*), extendcompat/win32/pthread.hto cover it. Both nvcc fatbincustom_targets (CUDA) and icpxcustom_targets (SYCLcommon.cpp/picture_sycl.cpp/dmabuf_import.cpp, plus the SYCL feature kernels) bypass meson'sdependencies:plumbing and hand-roll their own-Ilists, so the shim path must be threaded into bothcuda_extra_includesandsycl_inc_flagsexplicitly on Windows. icpx-cl on Windows additionally rejects-fPIC(unsupported option for target 'x86_64-pc-windows-msvc') — sosycl_common_argsandsycl_feature_argsroute their-fPICtoken throughsycl_pic_arg = host_machine.system() != 'windows' ? ['-fPIC'] : []. PIC is the default for Windows DLLs, so dropping the flag is the correct fix rather than a workaround. Round-14 surfaced a third Windows-only blocker:core/src/feature/integer_adm.h(an upstream Netflix file, last touched by upstream port d06dd6cf) initialisesdwt_7_9_YCbCr_threshold[3]with C99 designated initializers ({.a = ..., .k = ..., .f0 = ..., .g = {...}}). The header is included from bothinteger_adm.c(C TU) andcuda/integer_adm/*.cu(C++ TU via nvcc); MSVC's C++ frontend (and nvcc's cudafe++ on Windows) rejects C99 designated initializers without/std:c++20. Converted to positional initialization in the same struct-member order (a / k / f0 / g[4]) — the conversion is provably semantically identical and works in every C/C++ standard, so it costs nothing on the upstream-merge side beyond a trivial conflict marker if upstream Netflix later edits the same lines. Restore designated form post-merge if upstream has it. Round-17 surfaced four more Windows/MSVC-only SYCL blockers, two of which touch upstream-shared headers. (a)core/src/ref.handcore/src/feature/feature_extractor.h(UPSTREAM) unconditionally#include <stdatomic.h>and use theatomic_inttypedef in struct definitions. MSVC's<stdatomic.h>(added in 19.34) only declares the C11 symbols inside the global namespace under C; in C++ compilation (icpx-cl drives the SYCL TUs as C++) MSVC surfaces them only insidenamespace std::. gcc/clang expose both via a GNU extension, so the upstream code works on every other platform. The fork now wraps both headers'#include <stdatomic.h>in#if defined(__cplusplus) && defined(_MSC_VER)→#include <atomic>+using std::atomic_int;, falling through to the original<stdatomic.h>line on every other configuration. ABI is unchanged —atomic_intresolves to the same underlying type. If upstream Netflix adds further C11 atomic typedefs in these headers (e.g.,atomic_uint,atomic_size_t), extend theusing std::lines to cover them. (b)core/src/sycl/d3d11_import.cpp(fork-added) used<libvmaf/log.h>which doesn't exist —log.hlives atcore/src/log.hand is internal. Switched to"log.h"; the icpx invocation already supplies the src-relative-I. (c)core/src/sycl/dmabuf_import.cpp(fork-added) included<unistd.h>at file scope, but POSIXclose()is only used inside the#if HAVE_SYCL_DMABUFVA-API block. Moved the<unistd.h>include inside that guard so non-DMA-BUF builds (Windows MSVC, macOS) compile cleanly. (d)core/src/sycl/common.cpp(fork-added) calledclock_gettime(CLOCK_MONOTONIC), which doesn't exist on Windows. Replaced withstd::chrono::steady_clock(guaranteed monotonic by the C++ standard, portable on every supported host). All four fixes preserve POSIX/Linux behaviour bit-identically and only change the Windows MSVC build path. Round-18 surfaced a fifth Windows blocker on the CUDA leg's CPU SIMD compile path:core/src/feature/x86/motion_avx2.c:529(UPSTREAM, ported in commit 9371a0aa from Netflix PR #1486) computedfinal_accum[0] + final_accum[1] + final_accum[2] + final_accum[3]to extract the four int64 lanes from an__m256i. gcc/clang allow this via the GNU vector-extension treatment of__m256i(it carries__attribute__((vector_size(32)))); MSVC rejects it withC2088: built-in operator '[' cannot be applied to an operand of type '__m256i'. Replaced with_mm256_extract_epi64(final_accum, N)for N ∈ {0..3}, summed — bit-exact lane sum on every compiler. Restore the index form post-merge if upstream Netflix later edits the same lines and your toolchain matrix doesn't include MSVC. Round-19 surfaced the same MSVC pattern at 19 more call sites across the AVX2/AVX-512 ADM and motion files plus six GCC-style vector casts.core/src/feature/x86/adm_avx2.c(UPSTREAM): 6 lines (915-920) used(__m256i)(_mm256_cmp_ps(...))C-style casts that gcc/clang accept via the GNU vector extension; replaced with the dedicated_mm256_castps_si256(...)bit-cast intrinsic. 12 lane-extract sites (r2_h[0]+r2_h[1], etc. at lines 2420 / 2425 / 2430 / 2893 / 2897 / 2901 / 4079 / 4084 / 4089 / 4627 / 4631 / 4635) replaced with_mm_extract_epi64(r2_X, N)summed pair.core/src/feature/x86/adm_avx512.c(UPSTREAM): 6 sister lane-extract sites (lines 4470 / 4477 / 4484 / 4625 / 4631 / 4637) — same fix. The AVX-512 paths reduce a__m512idown to__m128ifirst (via_mm512_extracti64x4_epi64→_mm256_extracti64x2_epi64) before the index, so only the final__m128i[N]step needed changing.core/src/feature/x86/motion_avx512.c(UPSTREAM, ported in 9371a0aa from PR #1486): one finalr2[0]+r2[1]reduction (line 448), same fix. All 19 lane-extract fixes plus the 6 cast fixes are bit-exact rewrites and only change the source-level syntax to MSVC-portable form. Restore the original forms post-merge if upstream Netflix later edits the same lines and your toolchain matrix doesn't include MSVC. Additionallycore/src/sycl/d3d11_import.cpp(fork-added) switched from C-style COBJMACROS helpers (ID3D11Device_CreateTexture2D,…_Release, etc.) to C++ method-call syntax (device->CreateTexture2D,tex->Release) — d3d11.h gates COBJMACROS behind!defined(__cplusplus), so the C-style helpers aren't visible in this.cppTU. The two forms are ABI-equivalent (both dispatch through the COM vtable); the choice is purely lexical and POSIX builds aren't affected (the whole TU is#ifdef _WIN32). Round-20 surfaced two more Windows-only blockers. (a) 17 sites across the x86 SIMD layer used GCC'sfloat tmp[N] __attribute__((aligned(M)));form to align scratch buffers for_mm{256,512}_store_ps. MSVC rejects the trailing-attribute syntax withC2146: syntax error: missing ';' before identifier '__attribute__'. Replaced with the C11-standard_Alignas(M) float tmp[N];(alignment specifier before the type) — works in gcc, clang and MSVC with/std:c11. Files touched (all UPSTREAM):vif_statistic_avx2.c(×2),ansnr_avx2.c(×2),ansnr_avx512.c(×2),float_adm_avx2.c(×2),float_adm_avx512.c(×2),float_psnr_avx2.c(×1),float_psnr_avx512.c(×1),ssim_avx2.c(×4),ssim_avx512.c(×4). The pre-existingvif_avx2.c/vif_avx512.calready define a portableALIGNED(x)macro at file scope and position the attribute before the type, so they compile cleanly under MSVC and were not touched. (b)core/src/feature/mkdirp.c(UPSTREAM, third-party MIT-licensed copy of Stephen Mathieson's micro-library) included<unistd.h>unconditionally but never used POSIXunistdsymbols (onlymkdirvia<sys/stat.h>/<direct.h>). Gated<unistd.h>to non-Windows and added<direct.h>for Windows; switchedmkdir(pathname)→_mkdir(pathname)(the non-deprecated MSVC name).core/src/feature/mkdirp.hadded amode_ttypedef under MSVC since neither<sys/types.h>nor<sys/stat.h>declare it on Windows;modeis ignored on the Windows path anyway. Round-21 surfaced two more blockers (the round-19__m128i[N]sweep missed six sites) plus a pre-commit workflow checkout gap. (a)core/src/feature/x86/adm_avx512.c(UPSTREAM) had six furtherr2_X[0] + r2_X[1]reductions at lines 2128 / 2135 / 2142 / 2589 / 2595 / 2601 that reduce a__m512iaccumulator down to__m128ibefore the lane index. Replaced with the same_mm_extract_epi64(r2_X, N)summed-pair pattern used in round 19 — bit-exact, MSVC-portable. (b)core/src/log.c(UPSTREAM) included<unistd.h>unconditionally to pick up POSIXisatty/fileno. On MSVC both live in<io.h>as_isatty/_fileno; gated the include and macro-redirected the names so the one call site at line 34 compiles on both sides without touching the POSIX path. (c).github/workflows/lint-and-format.yml(fork-added) checks out withoutlfs: true, so themodel/tiny/*.onnxfiles land as LFS pointer stubs. pre-commit's "changes made by hooks" reporter then diffs the stubs against HEAD's real blobs and fails the job even though no hook touched them. Addedlfs: trueto the pre-commit job's checkout. (d)core/src/meson.build—cuda_common_vmaf_libstatic library had nodependencies:list, so the Win32 pthread shim (wired in viapthread_dependencyin core/meson.build) wasn't on its include path;cuda/common.hunconditionally#include <pthread.h>and MSVC failed with C1083. Addeddependencies : [pthread_dependency]— no-op on POSIX (empty list), routes the shim path in on Windows. (e)core/src/feature/integer_vif.c(UPSTREAM) walked one bigaligned_mallocresult asvoid *dataand diddata += pad_size/data += h * stride_16etc. to carve the buffer into typed sub-pointers. gcc/clang accept pointer arithmetic onvoid *as a GNU extension (treatingsizeof(void) == 1); MSVC rejects it withC2036: 'void *': unknown size. Replaced the cursor type withuint8_t *and added explicit casts at assignment sites that take a typed pointer (uint16_t *mu1,uint32_t *mu1_32, etc.). Byte offsets are identical, layout unchanged, bit-exact. If upstream Netflix edits the same loop, reabsorb the walk and re-apply the cursor-type + cast pattern. (f)core/src/feature/cuda/integer_adm_cuda.c(UPSTREAM) included<unistd.h>at line 33 but used no POSIX symbols from it; MSVC failed with C1083. Dropped the unused include outright — simplest fix, no runtime change on any platform. (g)core/src/dnn/model_loader.c(fork-added) usesS_ISDIR/S_ISREGto classify resolved paths. MSVC ships the underlyingS_IFMT/S_IFDIR/S_IFREGbit masks in<sys/stat.h>but not the POSIX classification macros. Added a Windows-only fallback (#ifndef S_ISDIR #define S_ISDIR(m) (((m) & S_IFMT) == S_IFDIR) #endif, same for S_ISREG) guarded by#ifdef _WIN32. Semantically identical to the POSIX macro on Linux/macOS. Round-21e surfaced the final source-portability blockers once the DLL build passed preprocessing. (h)core/src/predict.c,core/src/libvmaf.candcore/src/read_json_model.c(all UPSTREAM) used C99 variable-length arrays —double scores[cnt]at predict.c:385,char name[name_sz]at predict.c:453 and libvmaf.c:1741, pluscfg_name[cfg_name_sz]andgenerated_key[generated_key_sz]in the.jsonmodel-collection parser. gcc/clang accept VLAs as a C11 optional feature; MSVC (even with/std:c11) rejects them outright withC2057: expected constant expression(plus C2466 and C2133 on theconst size_tsized arrays — MSVC treatsconstas runtime-bounded, not a constant expression, even when the initialiser is literal like4 + 1). Replaced each runtime-sized buffer with a smallmalloc+ explicitfreeon every exit path (in predict.c and read_json_model.c agoto out;cleanup arm was introduced because the loops error-exit mid-function). Thegenerated_keybuffer in read_json_model.c uses the narrower fix —char generated_key[5];— since its size (four decimal digits of the bootstrap sub-model index plus NUL) is a true compile-time constant. Buffers are a handful of bytes each (name_szis the model-collection name length plus the fixed_ci_p95_losuffix,scoresholds ~20 doubles,cfg_nameis the name plus_0000suffix), so the heap round-trip is not performance-relevant; the new-ENOMEMfailure mode is handled uniformly by existing callers. The read_json_model.c refactor also plugs a pre-existing leak of thenamebuffer on the earlyreturn -EINVALwhen a JSON object key isn't a string — thegoto out;path freesname+cfg_nameon every exit.core/test/test_feature_extractor.c:56(UPSTREAM) declaredconst unsigned n_threads = 8;and used it as the extent ofVmafFeatureExtractorContext *fex_ctx[n_threads];. Converted toenum { n_threads = 8 };so MSVC sees a constant-expression; every other compiler accepts enum constants identically. Re-absorb if upstream Netflix later edits the same loops and your toolchain matrix omits MSVC. (i) The Windows MSVC build-only legs now build the full tree — CLI tools, unit tests and libvmaf.dll — rather than the previous short cut of disabling-Denable_tools/-Denable_tests. Per user direction ("fix the code ffs"), the tree polyfills the remaining POSIX surfaces on MSVC instead: (core/tools/compat/win32/getopt.h+core/tools/compat/win32/getopt.c) a from-scratch POSIX/GNU-compatiblegetopt_longshim (short / long options,no_argument/required_argument/optional_argument, argv permutation for non-option operands,--explicit stop,=-embedded values). The shim is fork-added (BSD-3-Clause-Plus-Patent, Copyright 2026 Lusoris and Claude) and declared via a singlegetopt_dependencyincore/meson.build, gated oncc.check_header('getopt.h')failing. The dependency auto-propagates the shim.cinto any consuming target via meson'ssources:keyword, so both thevmafCLI (core/tools/meson.build) and thetest_cli_parseunit test (core/test/meson.build) pick it up uniformly. MinGW ships<getopt.h>via mingw-w64-crt, socheck_headersucceeds there and the shim stays out of the TU list. (j) Eleven test executables (test_log,test_dict,test_opt,test_cpu,test_ref,test_feature,test_ciede,test_luminance_tools,test_cli_parse,test_sycl,test_sycl_pic_preallocation) were missingpthread_dependencyin theirdependencies:lists atcore/test/meson.build. On POSIXpthread_dependencyis an empty list so the omission was invisible; on MSVC those TUs transitively includefeature_collector.h→<pthread.h>and fail with C1083. Threaded the dependency through all eleven targets.test_cli_parseadditionally listsgetopt_dependencyto pick up the shim. (k) Three additional VLA sites surfaced once the test harness built on MSVC:test_cambi.c:254hadunsigned w = 5, h = 5; uint16_t buffer[3 * w];; converted toenum { w = 5, h = 5 };so the array extent is a constant expression.test_pic_preallocation.c:382andtest_pic_preallocation.c:506hadconst int num_threads = N; pthread_t threads[num_threads];— MSVC rejectsconst intas non-constant-expression. Converted toenum { num_threads = N, fetches_per_thread = M };. (l)test_ring_buffer.c:23(since removed; the ring-buffer test logic was folded into the CUDA-buffer / pic-preallocation suites) andtest_pic_preallocation.c:26included<unistd.h>forusleep/sleep. Gated behind!_WIN32with a Win32 fallback via<windows.h>+#define usleep(us) Sleep(((us) + 999) / 1000)/#define sleep(s) Sleep((s) * 1000). The conversion rounds sub-millisecondusleepinputs up, which is safe for these test paths (they use 100 µs jitter and 1 s waits). (m)core/tools/vmaf.cincluded<unistd.h>forisatty/fileno. Applied the same gating pattern used inlog.cin round-21(b) — include<io.h>on MSVC and redirectisatty/filenoto_isatty/_filenovia#define. (n)__builtin_clz/__builtin_clzllare GCC intrinsics; MSVC ships__lzcnt/__lzcnt64via<intrin.h>instead. The shim already lived incore/src/feature/integer_vif.hbutinteger_adm.c:939,x86/adm_avx2.c:1425andx86/adm_avx512.c:1217don't include that header. Extracted the shim into a dedicatedcore/src/feature/compat_builtin.h(fork-added) and included it from all four TUs. The guard isdefined(_MSC_VER) && !defined(__clang__), so clang-cl / icx-cl (which provide the GCC intrinsics natively) skip the shim. (o) The SYCL leg's D3D11 import TUcore/src/sycl/d3d11_import.cppis C++ (icpx-cl drives it as C++ on Windows) but included the internal C headerlog.hwithout anextern "C"wrap.log.his an upstream Netflix header with no__cplusplusguard, sovmaf_loggot C++ name-mangled in the .cpp TU and failed to resolve against the C-linkage symbol produced bylog.cat link time (LNK2019from every test target that pulls in the SYCL static lib). Wrapped the#include "log.h"withextern "C" { ... }inside the fork-added .cpp rather than touching the upstream header — keepslog.hidentical to upstream on every/sync-upstream. (p) The Windows MSVC legs build with--default-library=static. libvmaf's public API has no__declspec(dllexport)attributes (upstream Netflix is POSIX-shaped), so a vanilla MSVC shared build producessrc/vmaf-3.dllwith no exported symbols and the toolchain therefore never emits the companionvmaf.libimport library. Downstream tool targets then fail withLNK1181: cannot open input file 'src\vmaf.lib'. The MinGW matrix leg has used--default-library staticsince day one for the same reason (line 387); the MSVC legs now mirror that choice viamatrix.include[].meson_extra. Downstream consumers that want a DLL can either add__declspec(dllexport)decorations to the public API or use a.deffile; that is a separate decision and out of scope for the build-only gate. - Re-test:
# Local sanity: the matrix file parses and the new job names exist.
yq '.jobs.windows-gpu-build.strategy.matrix.include[].name' \
.github/workflows/libvmaf-build-matrix.yml
# Expected output (2 lines):
# Build — Windows MSVC + CUDA (build only)
# Build — Windows MSVC + oneAPI SYCL (build only)
- Branch protection: the two Windows GPU legs are pinned as required status checks on
masterimmediately after this PR's merge. After ADR-0120's two Linux DNN legs the count moves 21 → 23. Re-pin via:
gh api --method PUT repos/VMAFx/vmafx/branches/master/protection \
--input /tmp/protection-update.json
0023 — CUDA gencode coverage (sm_86/sm_89/compute_80 PTX) + init hardening¶
- Workstream PRs: the ADR-0122 PR (gencode + init hardening) and the ADR-0123 follow-up for the
32b115dfpost-cubin-load regression. - Touches:
core/src/meson.build— thegencodearray in theif get_option('enable_nvcc')branch.core/src/cuda/common.c—vmaf_cuda_state_init()error paths (multi-line actionable log,cuda_free_functions()+free(c)+*cu_state = NULLcleanup).docs/backends/cuda/overview.md—## Runtime requirementssection and### GPU architecture coveragetable.- Invariant: the
gencodearray unconditionally emits cubins forsm_75/sm_80/sm_86/sm_89plus acompute_80PTX, independent of hostnvccversion. Upstream Netflix's gencode only ships cubins at Txx major boundaries (sm_75/sm_80/sm_90/sm_100/sm_120); a literal merge that replaces our array with upstream's would re-open the Ampere-sm_86/ Ada-sm_89coverage hole. Thesm_90/sm_100/sm_120entries are still version-gated and should be preserved verbatim if upstream adds new gates. The init-path error messages are fork-local strings; upstream's terse"Error: failed to load CUDA functions"must NOT win a merge. - Re-test:
meson setup build -Denable_cuda=true -Denable_nvcc=true
ninja -C build 2>&1 | grep -E 'compute_(80|86|89)'
# Expect at least -gencode=arch=compute_86,code=sm_86 and
# -gencode=arch=compute_89,code=sm_89 and
# -gencode=arch=compute_80,code=compute_80
# Actionable init message (run without CUDA driver on the loader path):
LD_LIBRARY_PATH= ./build/tools/vmaf --help 2>&1 | grep -qi 'libcuda.so.1' || \
echo "init log regressed"
0024 — vmaf_read_pictures null-guard for CUDA device-only path¶
- Workstream PRs: the ADR-0123 follow-up landed atop the ADR-0122 gencode/init-hardening work.
- Touches:
core/src/libvmaf.c— the non-threaded tail ofvmaf_read_picturesat theprev_refupdate site (line ~1428 in the fork; upstream equivalent is the tail added byf740276a).- Invariant: the
prev_refupdate is guarded byif (ref && ref->ref)so pure-CUDA extractor sets (whereref = &ref_hostbutref_hostwas never populated bytranslate_picture_device) do not deref a NULL refcount. Upstream currently has the same unguarded tail; the bug is masked upstream only because the experimentalVMAF_PICTURE_POOLgate from32b115dfis still in place. A literal upstream merge that removes our null-guard while upstream's experimental gate is still holding would pass tests but re-open thelibvmaf_cudaffmpeg crash the moment the gate flips default-on (which the fork did in65460e3a, ADR-0104). Keep the guard until the upstream null-guard port lands. - Re-test:
# Unit tests cover the non-regression on the library side:
meson test -C build
# End-to-end regression: ffmpeg libvmaf_cuda must exit 0 on a
# CUDA-device-only extractor set (full recipe in ADR-0123).
./ffmpeg -init_hw_device cuda=cu:0 -filter_hw_device cu \
-i /tmp/ref.mp4 -i /tmp/dis.mp4 \
-lavfi "[0:v]format=yuv420p,hwupload_cuda[r];\
[1:v]format=yuv420p,hwupload_cuda[d];\
[r][d]libvmaf_cuda=log_path=/tmp/out.json:log_fmt=json" \
-f null -
0025 — VIF init() fail-path frees advanced byte-cursor¶
- Workstream PRs: PR #47 (rewritten to leak-fix-only after master absorbed the void→uint8_t half via commit
b0a4ac3a, entry 0022 §e). Ports the leak-fix half of upstream Netflix PR #1476. - Touches:
core/src/feature/integer_vif.c(UPSTREAM — 2-line fix in theinit()fail:handler). - Invariant:
init()walksuint8_t *dataforward throughaligned_malloc's one allocation, advancing past each sub-pointer assignment. Ifvmaf_feature_name_dict_from_provided_featuresreturns NULL the fail path must free the base pointers->public.buf.data, never the advanced cursordata. Upstream master still hasaligned_free(data)there — same bug — so this entry is the reminder to not let an upstream sync re-introduce the advanced-cursor form. If upstream lands PR #1476 or an equivalent, the sync can drop this entry. - Re-test:
meson test -C build --suite=fast
# Static check: ripgrep the pattern that must NOT return.
rg -n "aligned_free\(data\)" core/src/feature/integer_vif.c && \
echo 'REGRESSED' || echo 'ok'
0026 — Automated rule-enforcement workflow + copyright pre-commit hook¶
- Workstream PRs: this PR (ADR-0124 adoption). Closes the "rule-without-a-check" gap on ADR-0100 / 0105 / 0106 / 0108.
- Touches (all FORK-ADDED — no upstream overlap):
.github/workflows/rule-enforcement.yml(new),scripts/ci/check-copyright.sh(new),.pre-commit-config.yaml(appended local hook). - Invariant: the
deep-dive-checklistjob is blocking on every PR that is not an upstream port (exempt viaport:title prefix orport/branch). The other three gates (doc-substance-check,adr-backfill-check, copyright pre-commit) are advisory or pre-commit, never CI-blocking; this split is the whole point of ADR-0124 and an upstream sync must not move them into the required-status-check set without a follow-up ADR. The opt-out parser matches/^-?\s*no .* (?:needed|impact|rebase-sensitive)/per ADR-0108 §Opt-out-lines — if upstream ever changes PR-template phrasing (unlikely; this is fork-local), the regex and the template must move together. - Re-test:
# Lint the workflow + hook locally.
pre-commit run --files \
.github/workflows/rule-enforcement.yml \
scripts/ci/check-copyright.sh \
.pre-commit-config.yaml
# Dry-run the copyright hook against a staged source file.
scripts/ci/check-copyright.sh core/src/libvmaf.c && echo ok
# Synthetic PR body that violates ADR-0108 should fail the parser;
# see docs/research/0002-automated-rule-enforcement.md §Verification
# plan for the three test cases.
0027 — SSIMULACRA 2 scalar extractor (libjxl FastGaussian IIR blur)¶
- Workstream PRs: this PR (
feat/ssimulacra2-scalar); proposal ADR in PR #67. - Touches:
core/src/feature/ssimulacra2.c(fork-local, new),core/src/meson.build,core/src/feature/feature_extractor.c. - Invariant: the extractor embeds several tables that must track libjxl upstream — opsin absorbance matrix,
MakePositiveXYBoffsets, 108 pooling weights, polynomial-transform coefficients, and the FastGaussian coefficient-derivation formulas (radius =3.2795·σ + 0.2546, Cramer's 3×3 solve for β, n2/d1 assignment per Charalampidis 2016 (33)). If libjxl ever changes any of these, updatessimulacra2.cin the same PR that syncs upstream. Self-consistency must stay at exactly100.000000for identical ref/dist inputs — this is the cheapest regression check. - Re-test:
meson test -C build --suite=fast
./build/tools/vmaf \
--reference python/test/resource/yuv/src01_hrc00_576x324.yuv \
--distorted python/test/resource/yuv/src01_hrc00_576x324.yuv \
-w 576 -h 324 -p 420 -b 8 --feature ssimulacra2 -o /tmp/self.xml \
&& grep -q 'ssimulacra2="100.000000"' /tmp/self.xml \
&& echo "ok: self-consistency 100.0"
0028 — MS-SSIM separable decimate + AVX2/AVX-512/NEON SIMD¶
- Workstream PRs:
feat/ms-ssim-decimate-simd-v2(supersedes the rebase-incompatiblefeat/ms-ssim-decimate-simd; AVX2/AVX-512, commits7de8cd7fscalar separable,5f93c864AVX2,73436438AVX-512);feat/ms-ssim-decimate-neon-v2(NEON follow-up, stacked). - Touches:
core/src/feature/ms_ssim_decimate.{c,h}(NEW),core/src/feature/x86/ms_ssim_decimate_avx2.{c,h}(NEW),core/src/feature/x86/ms_ssim_decimate_avx512.{c,h}(NEW),core/src/feature/arm64/ms_ssim_decimate_neon.{c,h}(NEW),core/src/feature/ms_ssim.c(call-site change),core/src/meson.build(register new SIMD TUs),core/test/test_ms_ssim_decimate.c(NEW),core/test/meson.build(arm64 gating). - Invariant: the 9-tap 9/7 biorthogonal wavelet LPF coefficients (
ms_ssim_lpf_h/ms_ssim_lpf_v) are duplicated verbatim in five TUs for bit-identity: the scalarms_ssim_decimate.c, the AVX2 variant, the AVX-512 variant, the NEON variant, and upstream'sg_lpf_h/g_lpf_vinms_ssim.c. Any upstream change to the coefficient values or theKBND_SYMMETRICmirror branch iniqa/convolve.cmust be mirrored to all five. If not mirrored, SIMD paths and scalar diverge silently and the bit-equalitymemcmpintest_ms_ssim_decimatecatches it — but only when that test runs, so diff the five files first. - Re-test (on each supported host arch):
# x86_64 host — native build.
meson test -C build
./build/test/test_ms_ssim_decimate
# aarch64 host OR aarch64 cross under qemu — see /tmp/aarch64-cross.txt.
meson setup build-arm64 libvmaf --cross-file /tmp/aarch64-cross.txt \
-Denable_cuda=false -Denable_sycl=false
ninja -C build-arm64
qemu-aarch64-static -L /usr/aarch64-linux-gnu \
build-arm64/test/test_ms_ssim_decimate
# Netflix MS-SSIM golden — places=4 must still pass through SIMD.
.venv/bin/python -m pytest \
python/test/feature_extractor_test.py::FeatureExtractorTest::test_run_ms_ssim_fextractor
0029 — KBND_SYMMETRIC period-based reflection in iqa/convolve.c¶
- Workstream PRs:
feat/ms-ssim-decimate-simd-v2follow-up (CI triage on PR #69, 2026-04-20). - Touches:
core/src/feature/iqa/convolve.c(upstream file, rewrittenKBND_SYMMETRIC). - Invariant:
KBND_SYMMETRIC(img, w, h, x, y, _)must use the period-based form (period = 2*w,period = 2*h) so that offsets with|x| > wor|y| > hstill land in bounds. Upstream's single-reflect form was out-of-bounds wheneverw < kernel_halforh < kernel_half; the latent bug did not reproduce in Netflix golden tests because MS-SSIM pyramids never decimate below ~60×34. Any upstream change that reverts to the single-reflect form must be rejected or re-ported. - Re-test:
./build/test/test_ms_ssim_decimate # test_1x1 border case
.venv/bin/python -m pytest \
python/test/feature_extractor_test.py::FeatureExtractorTest::test_run_ms_ssim_fextractor
0030 — adm_decouple_s123_avx512 stack-array 64-byte alignment¶
- Workstream PRs:
feat/ms-ssim-decimate-simd-v2follow-up (CI triage on PR #69, 2026-04-20). - Touches:
core/src/feature/x86/adm_avx512.c(upstream file, one-line_Alignas(64)onint64_t angle_flag[16]at line 1317).core/test/test_pic_preallocation.c(upstream file, threevmaf_model_destroy(model)calls pairing thevmaf_model_loadintest_picture_pool_basic/_small/_yuv444). - Invariant: the stack slot for
angle_flagmust be 64-byte aligned because two_mm512_loadu_si512(&angle_flag[0/8])loads in the same scope may be promoted to alignedvmovdqa64by LTO. Dropping the_Alignas(64)annotation re-introduces the SEGV under--buildtype=release -Db_lto=true -Db_sanitize=address. Debug / no-LTO builds keepvmovdqu64and cannot flag the regression. Seedocs/development/known-upstream-bugs.md. - Re-test:
meson setup build-asan-lto libvmaf \
-Denable_cuda=false -Denable_sycl=false \
-Db_sanitize=address --buildtype=release -Db_lto=true
ninja -C build-asan-lto test/test_pic_preallocation
ASAN_OPTIONS=detect_leaks=1 \
./build-asan-lto/test/test_pic_preallocation
0031 — Batch-A upstream-port small-fix sweep (ports of unmerged PRs)¶
- Workstream PRs:
feat/batch-a-upstream-small-fix-sweep— commits546a40ee(T0-1),8fed8ad1(T4-4),83a1db46(T4-5),34425dee(T4-6). ADRs 0131, 0132, 0134, 0135. - Touches:
core/src/cuda/picture_cuda.c(one-linecuMemFreeport of Netflix#1382)core/src/feature/feature_collector.c+core/test/test_feature_collector.c(mount/unmount bugfix port of Netflix#1406 + shared-helper test refactor)core/src/meson.build(declare_dependency+override_dependencyport of Netflix#1451)core/include/libvmaf/model.h,core/src/model.c,core/test/test_model.c,docs/api/index.md(built-in model iterator port of Netflix#1424)- Invariant: each of the four upstream PRs is OPEN (unmerged) on the port date; when Netflix merges any of them, the fork's version is correction-bearing (T4-4 test refactor, T4-6 three defect fixes + Doxygen doc expansion), not line-identical. Resolution on upstream merge is always "keep fork version" because the fork's version already satisfies the PR's intent and additionally fixes the defects.
- Netflix#1406 conflict will land in
test_feature_collector.c— fork usesload_three_test_models()helper vs upstream's inline per-modelVmafModel *m0, *m1, *m2;duplication. - Netflix#1424 conflict will land in
core/src/model.candcore/test/test_model.c— fork useselse ifguard +idx + 1 < CNT+ const-qualified test types. - Netflix#1382 and Netflix#1451 are line-identical in substance; merge should be clean aside from trailing-comma style drift.
- Re-test:
meson setup build libvmaf -Denable_cuda=false -Denable_sycl=false
ninja -C build test/test_feature_collector test/test_model
build/test/test_feature_collector
build/test/test_model
# Expected: 6/6 pass in test_feature_collector (mount/unmount
# 3-model sequences); 39/39 pass in test_model (includes
# test_version_next full-iteration invariant).
0032 — Thread-local locale handling for numeric I/O (port of Netflix/vmaf#1430)¶
- Workstream PRs:
port/netflix-1430-thread-locale(T4-3 from the "Batch-A follow-up" sweep, 2026-04-20). - Touches:
core/src/thread_locale.h/core/src/thread_locale.c(new, upstream-authored);core/src/meson.build(twocdata.set('HAVE_USELOCALE'/'HAVE_XLOCALE_H')probes +src_dir + 'thread_locale.c'inlibvmaf_sources);core/src/output.c(four writers gainpush_c()+pop()bracket, preserving fork'sferror(outfile) ? -EIO : 0return contract from ADR-0119);core/src/svm.cpp(drop<locale.h>include; replacesetlocale/strdup/setlocalebracket withvmaf_thread_locale_push_c/pop; addbuffer.imbue(std::locale::classic())to both SVM parser ctors with fork's K&R + 4-space style);core/src/read_json_model.c(bracketmodel_parsewith push/pop);core/test/meson.build(newtest_locale_handlingtarget + test registration);core/test/test_locale_handling.c(new, upstream-authored with three fork corrections for thescore_formatparameter). - Invariant: fork's output writers return
ferror(outfile) ? -EIO : 0— this must survive any upstream refactor of the writer bodies. Thepush_c()call MUST be paired with apop()on every return path (writer bodies have a single tail return, so the pattern is locallypush → body → pop → return ferror-check). Droppingpop()leaks alocale_ton POSIX and leaves the thread locked to "C" on Windows. - Re-test:
meson setup build -Denable_cuda=false -Denable_sycl=false
ninja -C build
meson test -C build test_locale_handling
# Repro the user-visible failure without the fix:
LC_ALL=de_DE.UTF-8 build/tools/vmaf --reference ref.yuv \
--distorted dis.yuv --width 1920 --height 1080 \
--pixel_format 420 --bitdepth 8 --output result.json \
--json
# Assert output contains period decimals, not comma.
python -c "import json; d=json.load(open('result.json')); \
assert all('.' in repr(v) for v in \
[f['metrics']['vmaf'] for f in d['frames']])"
- On upstream sync: when Netflix merges PR #1430, the
(cherry picked from commit 054a97ed…)trailer ingit log port/netflix-1430-thread-localelets the next/sync-upstreamskip this commit. If the upstream diff drifts, redo the three fork corrections listed in ADR-0137 §Decision.
0033 — SSIM / MS-SSIM SIMD bit-exact to scalar via per-lane scalar double¶
- Workstream PRs:
feat/ms-ssim-decimate-neon(this PR — companion to the ADR-0138 convolve fast path). - Touches:
core/src/feature/x86/ssim_avx2.candcore/src/feature/x86/ssim_avx512.c—ssim_accumulate_*rewritten.ssim_precompute_*andssim_variance_*unchanged (they were already bit-exact). Plus the new bit-exactconvolve_avx2.c/convolve_avx512.cand the upstream h-pass OOB fix atiqa/convolve.c:159. - Invariants (see ADR-0139 §Decision):
- Convolve taps — single-rounded
float*float→ widen →doubleadd, NO FMA. Mirrors scalarsum += img[i]*k[j]iniqa/convolve.c. - SSIM accumulate — scalar's
2.0 *literal (2.0 * ref_mu[i] * cmp_mu[i] + C1and2.0 * srsc + C2) is a Cdoubleliteral. Both SIMD accumulators do the2.0 *numerator + division + finall*c*sproduct per-lane in scalar double to match scalar type promotions byte-for-byte. - H-pass outer-loop bound —
y < dst_h + vc - kh_even(noty < dst_h + vc); the- kh_evenis load-bearing because the last cache row on even-tap kernels (e.g. box-8) is never read by the v-pass but was previously written OOB when image height equals kernel height.
Fork-local SSIM SIMD is NOT upstream. If upstream ever adds their own SSIM AVX2/AVX-512, keep the fork's version on conflict — it's the only variant verified bit-exact to scalar at --precision max. - Re-test:
meson setup build -Denable_cuda=false -Denable_sycl=false
ninja -C build
meson test -C build test_iqa_convolve test_ms_ssim_decimate
# Bit-exactness check across dispatch backends:
FIX=python/test/resource/yuv/checkerboard_1920_1080_10_3_0_0.yuv
DIS=python/test/resource/yuv/checkerboard_1920_1080_10_3_1_0.yuv
for m in 255 16 0; do
build/tools/vmaf --cpumask $m --reference $FIX --distorted $DIS \
--width 1920 --height 1080 --pixel_format 420 --bitdepth 8 \
--feature float_ssim --feature float_ms_ssim \
--output /tmp/ssim_$m.xml --precision max
done
diff <(grep -v '<fyi fps' /tmp/ssim_255.xml) \
<(grep -v '<fyi fps' /tmp/ssim_16.xml) # expect empty
diff <(grep -v '<fyi fps' /tmp/ssim_255.xml) \
<(grep -v '<fyi fps' /tmp/ssim_0.xml) # expect empty
- On upstream sync: the AVX2/AVX-512 SSIM surface is entirely fork-local (upstream has VIF/ADM/motion/CAMBI SIMD but no SSIM). If upstream ever introduces SSIM SIMD, their kernel bodies will almost certainly compute
l*c*sin vector float for throughput — do not adopt. The fork's per-lane-scalar-double reduction is required for the bit-exactness claim. Same applies toconvolve_avx2/512— they are fork-only; dispatch sits inssim_tools.cvia_iqa_convolve_set_dispatch.
0034 — SIMD DX framework + NEON SSIM/convolve bit-exact port¶
- Workstream PRs:
feat/simd-dx-framework(this PR, PR #A); ships the two demos on top of which PR #B will consume the framework (ssimulacra2, motion_v2, vif_statistic, ...). - Touches:
core/src/feature/simd_dx.h(new header),core/src/feature/arm64/convolve_neon.c+convolve_neon.h(new NEON port),core/src/feature/arm64/ssim_neon.c(ssim_accumulate_neonrewritten for ADR-0139 bit-exactness;precompute+varianceunchanged),core/src/feature/float_ssim.c+core/src/feature/float_ms_ssim.c(wireiqa_convolve_neoninto the aarch64 dispatch setters),core/src/meson.build(arm64_sources+= convolve_neon.c),core/test/meson.build(test_iqa_convolvearch filter extended toarm64/aarch64),core/test/test_iqa_convolve.c(NEON variant check + aarch64 CPU flag detection),core/test/dnn/meson.build(test_cli.shgated onnot meson.is_cross_build()— bash invokes$VMAF_BINdirectly so meson's exe_wrapper isn't applied), newbuild-aux/aarch64-linux-gnu.inimeson cross-file,.claude/skills/add-simd-path/SKILL.md(upgraded kernel-spec flags). - Invariants (see ADR-0140 §Decision):
simd_dx.his fork-local. Keep the fork's version on upstream conflict. Macro names are ISA-suffixed (_AVX2_4L,_AVX512_8L,_NEON_4L) — do not collapse into a cross-ISA abstraction; the fork's SIMD policy (user-memoryfeedback_simd_dx_scope.md) rules out Highway / simde / xsimd.- The ADR-0138 widen-then-add rule (single-rounded
float * float→ widen →doubleadd, NO FMA) applies to NEON exactly as to AVX2 / AVX-512. The NEON form uses pairedfloat64x2_taccumulators (lo / hi) because NEON has nofloat64x4_t. - The ADR-0139 per-lane scalar-double reduction rule applies to
ssim_accumulate_neonexactly as to the AVX2 / AVX-512 variants. The NEON implementation usesSIMD_ALIGNED_F32_BUF_NEON(_Alignas(16) float name[4]) + a 4-iteration scalar loop. - Re-test (requires
aarch64-linux-gnu-gcc+qemu-user-static+ aarch64 sysroot at/usr/aarch64-linux-gnu):
cd libvmaf
meson setup ../build-aarch64 \
--cross-file ../build-aux/aarch64-linux-gnu.ini \
-Denable_cuda=false -Denable_sycl=false -Denable_dnn=disabled
cd ..
ninja -C build-aarch64
meson test -C build-aarch64 # expect 31/31 OK
# Bit-exactness check scalar vs NEON under QEMU:
REF=python/test/resource/yuv/src01_hrc00_576x324.yuv
DIS=python/test/resource/yuv/src01_hrc01_576x324.yuv
for m in 255 0; do
LD_LIBRARY_PATH=$PWD/build-aarch64/src qemu-aarch64-static \
-L /usr/aarch64-linux-gnu build-aarch64/tools/vmaf \
--cpumask $m --reference $REF --distorted $DIS \
--width 576 --height 324 --pixel_format 420 --bitdepth 8 \
--feature float_ssim --feature float_ms_ssim \
--output /tmp/ssim_$m.xml --precision max
done
diff <(grep -v '<fyi fps' /tmp/ssim_255.xml) \
<(grep -v '<fyi fps' /tmp/ssim_0.xml) # expect empty
- On upstream sync: upstream has no NEON SSIM and no NEON convolve for IQA. If they ever add one, keep the fork's version on conflict — the fork's NEON path is the only variant verified bit-exact to scalar at
--precision max. Thebuild-aux/aarch64-linux-gnu.inicross-file has no upstream equivalent. The/add-simd-pathskill is fork-only; upstream doesn't ship.claude/skills/.
0036 — Port Netflix generalised AVX convolve + ADR-0141 cleanup¶
- Workstream PRs:
port/upstream-f3a628b4-generalized-avx-convolve(this PR). - Upstream commit:
f3a628b4"feature/common: generalize avx convolution for arbitrary filter widths" (Kyle Swanson, 2026-04-21). - Touches:
- convolution.h — upstream-tracking: adds
#define MAX_FWIDTH_AVX_CONV 17. - convolution_avx.c — upstream-tracking (2,500 LoC deletion) plus fork-delta cleanup per ADR-0141: four scanline helpers
convolution_f32_avx_s_1d_*changed from external linkage tostatic(no other TU uses them after the specialised-path removal); stride parameters widened frominttoptrdiff_tin the helpers, with(ptrdiff_t)casts at public-function multiplication sites;#include <stddef.h>added for the type. core/src/feature/vif_tools.c— upstream-tracking: three AVX dispatch sites drop thefwidth == 17 || ... == 3whitelist in favour offwidth <= MAX_FWIDTH_AVX_CONV.python/test/quality_runner_test.py,python/test/vmafexec_test.py— upstream-authored loosening of two full-VMAF-score assertions fromplaces=2(±0.005) toplaces=1(±0.05). Adopted per the ADR-0142 Netflix-authority precedent (project rule #1 addresses fork drift, not upstream-authored test updates the fork must track).- Invariants (see ADR-0143 §Decision):
- Static linkage on scanline helpers — upstream leaves the four
convolution_f32_avx_s_1d_*_scanlinehelpers with external linkage out of habit; the fork narrows them tostatic. On upstream sync: if upstream ever externs them from another TU, that's a flag to re-audit; keep the fork'sstaticunless the reference is real. ptrdiff_tstrides inside helpers — the publicconvolution_f32_avx_*_swrappers keepintstrides (matching the upstream interface +convolution.hdeclarations). Helpers takeptrdiff_tto silencebugprone-implicit-widening-of- multiplication-result. If upstream changes the public interface toptrdiff_t, drop the fork's wrapper-level casts.MAX_FWIDTH_AVX_CONV = 17— the ceiling is upstream's; if upstream bumps it, the fork must rebuild + re-run the VIF golden test pair.- Re-test:
meson setup build -Denable_cuda=false -Denable_sycl=false
ninja -C build
meson test -C build # expect 32/32 OK
clang-tidy -p build core/src/feature/common/convolution_avx.c
# Zero warnings expected on the touched file.
Netflix CPU golden CI leg exercises the two loosened assertions; confirmed locally under meson test. - On upstream sync: upstream is the source of truth for convolution_avx.c, convolution.h, vif_tools.c dispatch, and the two python golden tolerances. On a rebase, prefer upstream for those files except: - Keep the fork's static on the four scanline helpers. - Keep the fork's ptrdiff_t helper signatures + multiplication- site casts (unless upstream adopts them too, in which case converge). - Keep the fork's #include <stddef.h>. If upstream re-introduces a specialised fast path for common widths, evaluate on a per-fwidth perf profile — the fork's /profile-hotpath skill covers this.
0037 — Float convolution AVX-512 port (ADR-0504, fork-local)¶
- Workstream PR:
perf/float-convolution-avx512-port-2026-05-18. - Upstream: no AVX-512 float convolution path in upstream; this is fork-local (ADR-0504).
- Touches (fork-local):
convolution_avx512.c— new TU with four static scanline helpers and three public wrappers, all ported fromconvolution_avx.c(__m256→__m512, FMA added).convolution.h— adds threeconvolution_f32_avx512_*_sdeclarations.vif_tools.c— dispatch updated to testVMAF_X86_CPU_FLAG_AVX512beforeAVX2in all threevif_filter1d_*_sfunctions.core/src/meson.build— addsconvolution_avx512.ctox86_avx512_sources.- Rebase risk: LOW.
convolution_avx512.cis entirely fork-local; upstream changes toconvolution_avx.corconvolution.hmay need to be mirrored here, but the AVX-512 file has no upstream conflict surface. - Gate:
meson test -C build(63/63). Netflix CPU golden tests pass.
0038 — motion_v2 NEON SIMD (fork-local)¶
- Workstream PR:
port/motion-bundle-neon-and-updates(this PR). - Upstream: none — aarch64 NEON for
motion_v2is fork-local. Upstream scalar + AVX2 + AVX-512 variants exist; this PR adds the missing NEON fourth path. Scalar is the bit-exactness ground truth. - Touches (fork-local):
- motion_v2_neon.c — new TU, ~300 LoC. 4-wide int32 SIMD over the 5-tap Gaussian pipeline. Five
static inlinehelpers keep every function under the ADR-0141 60-line budget. - motion_v2_neon.h — new header declaring the two public entry points.
- integer_motion_v2.c — dispatch update: adds an
#if ARCH_AARCH64block ininitthat selects the NEON variant whenVMAF_ARM_CPU_FLAG_NEONis present, mirroring the existing x86 dispatch blocks. core/src/meson.build— addarm64/motion_v2_neon.cto thearm64_sourceslist.- Invariants (see ADR-0145 §Decision):
- Arithmetic right-shift throughout. The fork's AVX2 path uses
_mm256_srlv_epi64(logical) which can diverge from scalar on negative-diff pixels. The NEON port usesvshrq_n_s64(v, 16)for the known Phase-2 shift andvshlq_s64(v, -(int64_t)bpc)for the variable Phase-1 shift — both arithmetic, matching scalar C>>on signed integer. On rebase: keep the arithmetic forms; do NOT adoptvshrq_n_u64or a logical emulation even if it runs faster. - 4-lane stride + mirror tails. SIMD stride = 4; scalar tails cover the remainder. The Phase-2 helper
x_conv_row_sad_neonhands 4 lanes tox_conv_block4_neonand drops to scalar for both left/right edges (j < 2andj + 6 > w). On rebase: preserve the 4-lane stride and the two-sided scalar tail. - Signature parity with AVX2. Both pipeline entry points match the AVX2 + AVX-512 variants'
(const uint8_t *prev, ptrdiff_t, const uint8_t *cur, ptrdiff_t, int32_t *y_row, unsigned w, unsigned h, unsigned bpc)signature. On rebase: if upstream changes the signature, mirror the change here AND in the x86 variants in lockstep. - Re-test:
meson setup build-aarch64 libvmaf \
--cross-file build-aux/aarch64-linux-gnu.ini \
-Denable_cuda=false -Denable_sycl=false
ninja -C build-aarch64
meson test -C build-aarch64 --no-rebuild # expect 31/31 OK
clang-tidy -p build-aarch64 \
core/src/feature/arm64/motion_v2_neon.c
# Zero warnings expected on the touched file.
# NEON-vs-scalar bit-exact diff under QEMU:
YUV=python/test/resource/yuv
for mask in 0 255; do
LD_LIBRARY_PATH=build-aarch64/src \
qemu-aarch64-static -L /usr/aarch64-linux-gnu \
build-aarch64/tools/vmaf \
-r $YUV/src01_hrc00_576x324.yuv \
-d $YUV/src01_hrc01_576x324.yuv \
-w 576 -h 324 -p 420 -b 8 -n --feature motion_v2 \
--cpumask $mask -o /tmp/mv2_$mask.xml --precision max
done
diff <(grep -v 'fps=' /tmp/mv2_0.xml) \
<(grep -v 'fps=' /tmp/mv2_255.xml) # expect empty
- On upstream sync: upstream has no NEON
motion_v2and has not signalled plans to add one. If they ever do, diff their NEON against the fork's: on logical-vs-arithmetic shift, keep the fork's arithmetic form (matches scalar). On the function decomposition (the five helpers), adopt upstream's if it's smaller; the fork's layout is ADR-0141-driven, not a semantic contract. - Follow-up T7-32 (fixed 2026-05-09): The
_mm256_srlv_epi64(logical right shift) inmotion_score_pipeline_16_avx2was replaced withsrav_epi64_imm, an AVX2-safe arithmetic-right-shift emulation: logical shift OR sign-fill mask viasrai_epi32+slli_epi64. Two bugs were closed in the same PR: - AVX2 logical-vs-arithmetic shift:
_mm256_srlv_epi64replaced bysrav_epi64_immincore/src/feature/x86/motion_v2_avx2.c. The emulation is bit-exact with scalar C>> bpcon signedint64_t. - Test scalar reference mirror:
mirror_idxincore/test/test_motion_v2_simd.cused2*size - idx - 1instead of2*size - idx - 2, diverging frominteger_motion_v2.c::mirror(). Fixed to-2. All four adversarial fixtures (neg-diff bpc10/12, mixed-diff bpc10/12) now pass.meson test -C build50/50 OK. On rebase: keepsrav_epi64_imm; do not revert to_mm256_srlv_epi64. The rebase-time invariant is now: AVX2 path uses arithmetic shift (matching NEON and scalar).
0039 — readability-function-size NOLINT sweep (ADR-0146)¶
- ADR: ADR-0146
- Touches:
core/src/dict.ccore/src/picture.ccore/src/picture_pool.ccore/src/predict.ccore/src/libvmaf.ccore/src/output.ccore/src/read_json_model.ccore/src/feature/feature_extractor.ccore/src/feature/feature_collector.ccore/src/feature/iqa/convolve.ccore/src/feature/iqa/ssim_tools.ccore/src/feature/x86/vif_statistic_avx2.c- Invariant: every
readability-function-sizeNOLINT suppression has been replaced by a set of smallstatic(orstatic inline, for the SIMD / IQA files) helpers. The helper names are stable interfaces the surrounding code depends on (e.g.iqa_convolve_1d_separable,iqa_convolve_2d,ssim_compute_stats,ssim_workspace_alloc/_free,vif_stat_simd8_compute/_reduce,struct vif_simd8_lane,read_pictures_extractor_loop,read_pictures_post_extractor,read_pictures_validate_and_prep,read_pictures_update_prev_ref). Upstream Netflix has no equivalent helpers; rebases touching any of these files will conflict against the fork's split shape. - On upstream sync:
- If upstream lands a different decomposition of
_iqa_convolveor_iqa_ssim, prefer upstream's shape only if it keeps the ADR-0138 / ADR-0139 bit-exactness invariants (single-rounded float mul → widen to double → double add; per-lane scalar-float reduction through aligned temp buffer). Otherwise keep the fork's split and re-document the divergence here. - The fork renamed
_calc_scale→iqa_calc_scaleto clear thebugprone-reserved-identifiercheck. If upstream modifies_calc_scale, keep the fork's name and port the behavioural change. model_collection_parse_loopwrites directly tocfg_namerather than throughc->name— if upstream ever rewritesmodel_collection_parse, preserve the direct write (it's what lets the param stay non-const without a NOLINT).- Re-test on rebase (x86, any libsvm-less host):
ninja -C build && meson test -C build
for mask in 0 255; do
VMAF_CPU_MASK=$mask ./build/tools/vmaf \
--reference python/test/resource/yuv/src01_hrc00_576x324.yuv \
--distorted python/test/resource/yuv/src01_hrc01_576x324.yuv \
--width 576 --height 324 --pixel_format 420 --bitdepth 8 \
-m version=vmaf_v0.6.1 -o /tmp/vmaf_$mask.xml
done
diff <(grep -v fyi /tmp/vmaf_0.xml) <(grep -v fyi /tmp/vmaf_255.xml)
# expect exit 0 (Netflix-golden-pair VMAF bit-identical scalar vs SIMD)
Also run clang-tidy -p build on every file in Touches; expect zero warnings. - Follow-up T7-6: decide whether to rename the _iqa_* API surface (convolve / ssim / decimate / img_filter / filter_pixel / get_pixel) across all callers to clear the remaining bugprone-reserved-identifier suppressions in ssim.c, ms_ssim.c, float_ms_ssim.c. Out of scope here.
0040 — Thread-pool job recycling + inline data buffer (ADR-0147)¶
- ADR: ADR-0147
- Touches:
core/src/thread_pool.c - Invariants:
VmafThreadPoolJobcarries a fixed-sizechar inline_data[64]buffer. Payloads ≤ 64 bytes go throughmemcpy(job->inline_data, data, data_sz)+job->data = job->inline_data; payloads > 64 bytes take the legacymallocpath. The cleanup path MUST distinguish the two viajob->data != job->inline_data— a naivefree(job->data)would corrupt the slot. Enforced invmaf_thread_pool_job_clear_data.free_jobslist is protected by the existingqueue.lock; enqueue pops from it beforemallocing, runner recycles onto it after running a job.vmaf_thread_pool_destroywalks the list aftervmaf_thread_pool_waitreturns (all workers have exited → no lock needed). Any reorder that frees the queue lock before thefree_jobswalk is a leak on shutdown.- Fork's
void (*func)(void *data, void **thread_data)signature + per-workerVmafThreadPoolWorkerare fork-local; upstream Netflix #1464 hasfunc(void *data). Keep the fork's signature on any rebase — callers (src/libvmaf.c:threaded_enqueue_oneetc.) depend on the two-arg form. -
On upstream sync: Netflix PR #1464 is CLOSED (not merged) and bundles twelve unrelated optimizations. Only the thread-pool portion is ported here. If upstream ever reopens and merges #1464 (or a successor), cherry-pick only the pool mechanics; reject the payload-signature changes, the ADM / VIF / predict.c pieces (they conflict with ADR-0138 / 0139 / 0142 bit-exactness and with T7-5 predict.c refactor), and the feature-collector capacity bump (fork already capped at 8 for a reason — see
src/feature/feature_collector.c). -
Re-test on rebase (x86, any libsvm-less host):
ninja -C build && meson test -C build
for threads in 1 4; do
for mask in 0 255; do
VMAF_CPU_MASK=$mask ./build/tools/vmaf \
--reference python/test/resource/yuv/src01_hrc00_576x324.yuv \
--distorted python/test/resource/yuv/src01_hrc01_576x324.yuv \
--width 576 --height 324 --pixel_format 420 --bitdepth 8 \
-m version=vmaf_v0.6.1 --threads $threads -o /tmp/vmaf_${threads}_${mask}.xml
done
done
# Expect bit-identical scores (attribute order may differ across
# --threads 1 vs --threads 4 because feature-collector emits in
# insertion order; the numeric values match).
diff <(grep -v fyi /tmp/vmaf_4_0.xml) <(grep -v fyi /tmp/vmaf_4_255.xml)
# expect exit 0 (scalar vs SIMD threaded)
Also run clang-tidy -p build core/src/thread_pool.c — expect zero warnings. Re-run the 500 000-job micro-benchmark from ADR-0147 §Decision if performance is under investigation.
0041 — IQA reserved-identifier rename + cleanup (ADR-0148)¶
- ADR: ADR-0148
- Touches: 21 files across
core/src/feature/(iqa/{convolve,decimate,ssim_tools}.{c,h},iqa/ssim_simd.h,ssim.c,integer_ssim.c,ms_ssim.c,ms_ssim_decimate.h,float_ssim.c,float_ms_ssim.c,x86/convolve_avx2.{c,h},x86/convolve_avx512.{c,h},arm64/convolve_neon.{c,h},AGENTS.md) pluscore/test/test_iqa_convolve.c. - Invariants:
- Every
_iqa_*/_kernel/_ssim_int/_map_reduce/_map/_reduce/_context/_ms_ssim_*/_ssim_*/_alloc_buffers/_free_bufferssymbol and the four underscore-prefixed header guards (_CONVOLVE_H_,_DECIMATE_H_,_SSIM_TOOLS_H_,__VMAF_MS_SSIM_DECIMATE_H__) is renamed to its non-reserved spelling. The fork's IQA surface no longer uses C's reserved-identifier name space. - The
clang-analyzer-security.ArrayBoundNOLINT bracket inssim_accumulate_rowandssim_reduce_row_range(integer_ssim.c) is load-bearing — the inner kernel-loopk_min/k_maxclamping is provably correct (k_min = max(0, hkernel_offs - x),k_max = min(hkernel_sz, hkernel_sz - (x + hkernel_offs - w + 1))) but the analyzer can't follow it across helper boundaries. Do not collapse the bracket. - The
clang-analyzer-unix.MallocNOLINT bracket intest_iqa_convolve.c(check_simd_variant,check_case) is intentional — test exits process on failure path; small allocations leak by design at test end. Do not refactor to free-on-exit. - The cross-TU NOLINT pattern on
compute_ssim(ssim.c) andcompute_ms_ssim(ms_ssim.c) — clang-tidymisc-use-internal-linkageruns per-TU and can't see the header bridge tofloat_ssim.c/float_ms_ssim.c. Keep the inline justification comment. - On upstream sync:
- The Netflix upstream IQA library (
tjdistler/iqa) has been effectively abandoned (last meaningful commit pre-2020). Future rebases will conflict on every renamed symbol; drop the underscore-prefix on each conflict and mirror the fork'siqa_*naming. - If upstream Netflix/vmaf ever reincorporates the IQA naming wholesale, prefer the fork's spellings — this PR is a one-shot mechanical rename with no semantic content.
- Re-test on rebase:
ninja -C build && meson test -C build
for mask in 0 255; do
VMAF_CPU_MASK=$mask ./build/tools/vmaf \
--reference python/test/resource/yuv/src01_hrc00_576x324.yuv \
--distorted python/test/resource/yuv/src01_hrc01_576x324.yuv \
--width 576 --height 324 --pixel_format 420 --bitdepth 8 \
-m version=vmaf_v0.6.1 \
--feature float_ssim --feature float_ms_ssim \
-o /tmp/iqa_$mask.xml
done
diff <(grep -v fyi /tmp/iqa_0.xml) <(grep -v fyi /tmp/iqa_255.xml)
# expect exit 0 (bit-identical scalar vs SIMD on float_ssim/ms_ssim)
Also run clang-tidy -p build on every touched file (excluding arm64/); expect zero warnings.
0042 — Port Netflix #1376 — FIFO-hang fix via Semaphore (ADR-0149)¶
- ADR: ADR-0149
- Upstream commit: Netflix PR #1376, head
1c06ca4f1bb5da38b54db075a27c35ba8ea9d7b7(OPEN upstream as of 2026-04-24). - Touches:
python/vmaf/core/executor.py— baseExecutorclass +ExternalVmafExecutor-style subclass; delete_wait_for_workfiles/_wait_for_procfilespolling loops; rewrite_open_{work,proc}files_in_fifo_modearoundmultiprocessing.Semaphore(0); addopen_sem=Nonekwarg to every_open_{ref,dis}_{work,proc}fileand to the_open_workfilestaticmethod; drop unusedfrom time import sleep.python/vmaf/core/raw_extractor.py—AssetExtractor+DisYUVRawVideoExtractor; addopen_sem=Noneto_open_{ref,dis}_workfileoverrides (release on entry since these are no-ops); delete_wait_for_workfilesoverrides; drop unusedfrom time import sleep.- Fork carve-outs (load-bearing on rebase):
compat/python-vmaf/__init__.py:__version__follows the rootx-release-please-versionmarker — do NOT port upstream's bump to"4.0.0"independently. The fork uses one release stream per ADR-1127.from time import sleepis dropped from both files — upstream leaves the import in place (unused after their patch); the fork removes it because ADR-0141 touched-file rule requires ruff F401 clean.- Upstream typo preserved: the subclass warning message contains "to be created to be created". Comments note the typo inline; do not silently fix on rebase — it's upstream- authored and project policy is verbatim port.
- On upstream sync: upstream PR #1376 is still OPEN. When it merges, re-diff against the merged form; the touched hunks should be conflict-free because the fork now carries the same shape. Re-check whether upstream fixed the "to be created to be created" typo; if so, adopt the fix (it becomes a simple string update).
- Re-test:
python3 -m py_compile python/vmaf/core/executor.py \
python/vmaf/core/raw_extractor.py
ruff check python/vmaf/core/executor.py python/vmaf/core/raw_extractor.py
black --check python/vmaf/core/executor.py python/vmaf/core/raw_extractor.py
# all silent
# No FIFO-mode unit test in the tree; end-to-end harness
# exercise (needs libsvm + ffmpeg + fixtures) goes via
# make test-netflix-golden
# which doesn't exercise fifo_mode path but does verify the
# refactor didn't break executor.py imports.
0043 — Port Netflix #1472 — CUDA on Windows MSYS2/MinGW (ADR-0150)¶
- ADR: ADR-0150
- Upstream commits: Netflix PR #1472 —
15745cdf(portability) +b7b65e64(meson plumbing). Both OPEN upstream as of 2026-04-24. - Touches:
core/src/cuda/common.h— drop<pthread.h>include; rename reserved header guard__VMAF_SRC_CUDA_COMMON_H__→VMAF_SRC_CUDA_COMMON_INCLUDED.core/src/cuda/cuda_helper.cuh—#ifdef DEVICE_CODEguard around<cuda.h>vs<ffnvcodec/dynlink_loader.h>.core/src/picture.h—#ifdef DEVICE_CODEguard around<cuda.h>+ forward-declareVmafCudaStatevs<ffnvcodec/*>+ fulllibvmaf_cuda.h; rename reserved header guard.core/src/feature/integer_adm.h— updated comment abovedwt_7_9_YCbCr_thresholdtable noting the fork's positional-initializer shape vs upstream's#ifndef __CUDACC__shape (see §Fork carve-outs).core/src/feature/cuda/integer_adm/{adm_cm,adm_csf,adm_csf_den,adm_decouple,adm_dwt2}.cu—#ifndef DEVICE_CODEguard around#include "feature_collector.h".core/src/meson.build— Windows nvcc plumbing (+70 LoC underhost_machine.system() == 'windows'):vswhere-basedcl.exediscovery, MSVC + Windows SDK include path injection, CUDA version detection vianvcc --version,nvcc_ccbin_flags+nvcc_host_includesthreaded through everycustom_targetthat invokes nvcc.- Fork carve-outs (load-bearing on rebase):
integer_adm.huses positional initializers, NOT upstream's#ifndef __CUDACC__wrap. Both shapes resolve the MSVC/nvcc C++-designated-initializer issue; the positional form is C++-portable and keeps the table available to future.cuconsumers. Keep the fork's form on rebase.cuda_static_libkeepsdependencies : [pthread_dependency]. Upstream drops it; the fork needs it becausering_buffer.c(built as part ofcuda_static_lib)#includes<pthread.h>directly. On rebase: keep the fork's version.meson.buildgencode coverage block: the fork's ADR-0122 explicit cubin list (sm_75/80/86/89 + compute_80 PTX) sits after the new upstream nvcc-detect block. On rebase, re-assemble the same merged order: nvcc-detect first, then gencode coverage (both host-independent).- Header guards:
_INCLUDEDspellings are fork-local (ADR-0148 precedent). Upstream keeps reserved__VMAF_SRC_*_H__spellings. On rebase, keep_INCLUDED. - On upstream sync: PR #1472 is still OPEN. When merged, re-diff the three conflict-resolved hunks against upstream's final form. Keep fork's version on the four carve-outs above unless upstream meaningfully reshapes those regions.
- Re-test on rebase (Linux host with CUDA toolkit):
meson setup libvmaf core/build-cuda \
-Denable_cuda=true -Denable_nvcc=true -Denable_sycl=false
ninja -C core/build-cuda && meson test -C core/build-cuda
# Expect 6 .fatbin files generated + CLI linked + 35/35 tests pass.
Windows validation is operator-driven — CI does not yet have a Windows + MSYS2 + MinGW + MSVC BuildTools + CUDA runner (tracked as T7-3 in .workingdir2/OPEN.md). - Prerequisites note (Windows only): nv-codec-headers must be built from git master commit 876af32 or later. The release tag n13.0.19.0 is missing cuMemFreeHost, cuStreamCreateWithPriority, cuLaunchHostFunc, and other CudaFunctions members libvmaf uses. Pre-existing issue, not scope of this port.
0058 — libvmaf.pc Cflags leak fix (ADR-0200)¶
- ADR: ADR-0200; bug-fix follow-up to entry 0057.
- Upstream source: fork-local. Netflix has no Vulkan backend.
- Touches:
core/subprojects/packagefiles/volk/meson.build— drops-include volk_priv_remap.hfromvolk_dep.compile_args; keeps-DVK_NO_PROTOTYPES.core/src/vulkan/meson.build— pullsvolk_priv_remap_h_pathfrom the volk subproject and appends['-include', <path>]tovmaf_cflags_common(privatec_args:on libvmaf'slibrary()call).- Invariants (load-bearing):
-includeMUST stay offvolk_dep.compile_args— otherwise it leaks into staticlibvmaf.pcCflags. Test on rebase:meson setup ... -Ddefault_library=static -Denable_vulkan=enabled, thengrep Cflags meson-private/libvmaf.pc— must NOT containvolk_priv_remapor any build-dir absolute path.-includeMUST be applied to libvmaf's compile — every libvmaf TU that calls volk'svk*API needs the rename macros active. Thevmaf_cflags_commoninjection covers this for all libvmaf sub-libraries (libvmaf_feature, libvmaf_cpu, etc.).- The path comes from
subproject('volk').get_variable(...), not from a hardcoded string — survives volk wrap version bumps. - On upstream sync: zero upstream interaction.
- Re-test on rebase / volk wrap bump:
meson setup build-vk-static-test libvmaf -Denable_vulkan=enabled \
-Denable_cuda=false -Denable_sycl=false -Ddefault_library=static
ninja -C build-vk-static-test src/libvmaf.a
grep Cflags build-vk-static-test/meson-private/libvmaf.pc
# Expected: no `volk_priv_remap` substring, no build-dir absolute path
0057 — Volk vk* priv-remap for static-archive builds (ADR-0198)¶
- ADR: ADR-0198; follow-up to ADR-0185.
- Upstream source: fork-local. Netflix/vmaf has no Vulkan backend.
- Touches:
core/subprojects/packagefiles/volk/meson.build— overlay applied on top of the upstream volk wrap. Adds acustom_targetthat runsgen_priv_remap.pyto producevolk_priv_remap.hfrom the upstreamvolk.h, and wires-includeof the generated header intovolk.c'sc_argsandvolk_dep'scompile_args.core/subprojects/packagefiles/volk/gen_priv_remap.py— fork-added generator script (regex againstextern PFN_vkXxx vkXxx;declarations).- Invariants (load-bearing):
- Force-include must propagate to every libvmaf TU pulling in
volk_dep— verified via meson dep graph. Removing the-includefromcompile_argsre-introduces the static-link multi-def cascade. - Generator regex matches every
vk*PFN declaration involk.h— confirmed for volk-1.4.341 (784declarations,784remaps). Bumping the volk wrap version: re-run the generator (it's a configure-time custom target, so it's automatic) and confirm the rename count printed to stdout matches the count of^extern PFN_vklines in the newvolk.h. - The renamed symbols use the
vmaf_priv_prefix — chosen to match no upstream Netflix or Vulkan SDK identifier. Don't rename to_vk*(collides with reserved-identifier C namespace) orvkv_*etc. - On upstream sync: zero upstream interaction. The volk wrap is a libvmaf-managed subproject; Netflix doesn't ship a Vulkan backend.
- Re-test on rebase / after any volk wrap bump:
meson setup build-vk-static libvmaf -Denable_vulkan=enabled \
-Denable_cuda=false -Denable_sycl=false \
-Ddefault_library=static
ninja -C build-vk-static src/libvmaf.a
test "$(nm build-vk-static/src/libvmaf.a 2>/dev/null \
| grep -cE '^[0-9a-f]* (T|D|B|R) vk[A-Z]')" = "0" \
&& echo OK
(Followed by the BtbN-style link reproducer in the ADR References section.)
0056 — SSIMULACRA 2 snapshot gate + fp-contract-off split (ADR-0164)¶
- ADR: ADR-0164
- Upstream source: fork-local. Netflix/vmaf has no SSIMULACRA 2.
- Touches:
- python/test/ssimulacra2_test.py — new fork-added Python test. Uses
subprocess.callagainstExternalProgram.vmafexecwith--feature ssimulacra2; parses the--jsonoutput; asserts pooled + per-frame scores. - Invariants (load-bearing):
- Pinned values are CPU-only — generated on master HEAD after PR #100 merge. Re-generate if the scalar or any SIMD path changes semantically (which per ADR-0161/0162/0163's bit-exactness contract, it shouldn't — any bit-exact refactor leaves pinned values unchanged).
- Tolerance is 4 decimal places (
places=4) — matches 1e-4. The CPU paths are bit-exact so actual drift should be 0; the tolerance is defensive. -ffp-contract=offeverywhere in the ssimulacra2 pipeline:libvmaf_ssimulacra2_static_lib(scalar extractor),x86_ssimulacra2_avx2_lib,x86_ssimulacra2_avx512_lib, andarm64_ssimulacra2_lib(from ADR-0161). All four split out of their umbrella libs so other extractors keep upstream's default FMA policy. Without this the CI GCC/clang hosts drifted ~2e-4 from my AVX-512 authoring host — GCC 10+ defaults-ffp-contract=faston x86 with-mfmaand on aarch64, fusinga*b+cin scalar glue around the SIMD calls. Do NOT remove any of these carve-outs on rebase.- Fixtures are already-checked-in —
src01_hrc00/01_576x324is also the primary Netflix golden fixture; the 160×90 derived one stresses the sub-176 pyramid-termination path. - Do NOT modify the Netflix golden assertions in quality_runner_test.py et al. — those are upstream-pinned. This test is a SEPARATE file that adds fork-specific scores.
- On upstream sync: no upstream interaction. If Netflix adopts SSIMULACRA 2 in the future, cross-reference against their pinning if they add one.
- Re-test on rebase / after any ssimulacra2 change:
- Follow-ups:
- Cross-reference gate against libjxl
tools/ssimulacra2whenssimulacra2_rscargo install is fixed. - Expand fixture coverage if new YUV test assets land.
0055 — SSIMULACRA 2 picture_to_linear_rgb SIMD (ADR-0163)¶
- ADR: ADR-0163
- Upstream source: fork-local. Netflix/vmaf has no SSIMULACRA 2.
- Touches:
- ssimulacra2_avx2.{c,h} — new
ssimulacra2_picture_to_linear_rgb_avx2+ helpers (read_plane_scalar_s2,srgb_to_linear_lane_avx2,compute_matrix_coefs). - ssimulacra2_avx512.{c,h} — 16-wide AVX-512 port.
- ssimulacra2_neon.{c,h} — 4-wide aarch64 port.
- ssimulacra2.c — new
ptlr_fnfield inSsimu2State; dispatch wrapperconvert_picture_to_linear_rgbunpacksVmafPictureintosimd_plane_t[3]; init assigns AVX2/AVX-512/NEON pointers. - ssimulacra2_simd_common.h — new shared header declaring
simd_plane_t. Decouples SIMD TUs fromVmafPicturetype. - test_ssimulacra2_simd.c — new
test_ptlr_420_8,test_ptlr_420_10,test_ptlr_444_8,test_ptlr_444_10,test_ptlr_422_8subtests + scalar referencesref_read_plane,ref_srgb_to_linear,ref_picture_to_linear_rgb. - Invariants (load-bearing):
- Scalar-order matmul —
G = Yn + cb_g * Un + cr_g * Vnchained left-to-right in all three SIMD TUs. Regression test catches reordering drift (~1 ulp). - Per-lane scalar
powf— vector polynomial approximation would drift scalar bit-exactness. Do not replace the lane spill/reload pattern with a vector libm. simd_plane_tlayout —{data, stride, w, h}ordering assumed by all three SIMD TUs. The dispatch wrapper builds this fromVmafPicturefields; layout must match.- Bounds clamping in
read_plane_scalar_*mirrors scalar reference verbatim (if (sx < 0) sx = 0; if (sx >= pw) sx = pw-1;etc.). Do not simplify — removes per-lane safety at plane edges. - Arbitrary chroma ratios fall through to the
int64_tmultiplication branch. Don't remove it — SSIMULACRA 2 is supposed to accept non-standard ratios gracefully. - On upstream sync: no upstream interaction. If Netflix adopts SSIMULACRA 2 in the future and provides a SIMD YUV→RGB path, diff against the fork's — preserve the bit-exactness contract unless ADR-0142 Netflix-authority carve-out opens.
- Re-test on rebase:
ninja -C build && build/test/test_ssimulacra2_simd # 11/11
ninja -C build-aarch64 && \
qemu-aarch64-static -L /usr/aarch64-linux-gnu/ \
build-aarch64/test/test_ssimulacra2_simd # 11/11
- Follow-ups:
- T3-3 SSIMULACRA 2 snapshot-JSON regression test — still pending (gated on
tools/ssimulacra2availability). - SSIMULACRA 2 now has zero scalar hot paths. T3-1 closes in full with phases 1+2+3 (ADR-0161, 0162, 0163).
0054 — SSIMULACRA 2 FastGaussian IIR blur SIMD (ADR-0162)¶
- ADR: ADR-0162
- Upstream source: fork-local. No SSIMULACRA 2 extractor in upstream Netflix/vmaf.
- Touches:
- ssimulacra2_avx2.{c,h} — new
ssimulacra2_blur_plane_avx2+ 2 helpers (hblur_8rows_avx2,vblur_simd_8cols_avx2). - ssimulacra2_avx512.{c,h} — 16-wide port.
- ssimulacra2_neon.{c,h} — 4-wide aarch64 port, uses
vsetq_lane_f32in place of gather. - ssimulacra2.c — adds
blur_fnfunction pointer toSsimu2State, dispatch ininit_simd_dispatch(), call-site inblur_3plane. - test_ssimulacra2_simd.c — new
test_blur+ scalar reference (ref_blur_plane,ref_fast_gaussian_1d). - Invariants (load-bearing):
- Row-batching lane layout — horizontal pass lane
iMUST hold row(y_base + i). Gather index vector entries are(y_base + i) * w(stride-w). Changing this breaks bit-exactness vs scalar. - Scalar left-to-right summation order —
n2_k * sum - d1_k * prev1_k - prev2_kchained sequentially;o0 + o1 + o2at output time is(o0 + o1) + o2. Changing to(o0 + o2) + o1oro0 + (o1 + o2)will drift ~1 ulp and the regression test catches it. col_stateis 6 * w contiguous floats — layout is[prev1_0 | prev1_1 | prev1_2 | prev2_0 | prev2_1 | prev2_2]. SIMD loads assume this layout; changing field order requires updating all three SIMD TUs in lockstep withblur_plane.- NEON lane-set pattern — aarch64 has no gather intrinsic; 4 explicit
vsetq_lane_f32calls per input vector. Do not replace with ald1 {v.s}[lane]-style pseudo-gather without re-verifying bit-exactness. - Scalar tail in vertical pass matches scalar reference body verbatim. Any deviation breaks
memcmpequality on widths that aren't multiples of the SIMD width. - On upstream sync: no upstream interaction. If Netflix adopts SSIMULACRA 2 in the future and provides their own IIR blur SIMD, diff against the fork's and preserve the bit-exactness contract unless an ADR-0142 Netflix-authority carve-out is opened.
- Re-test on rebase:
meson setup build -Denable_cuda=false -Denable_sycl=false
ninja -C build
meson test -C build test_ssimulacra2_simd # 6/6
# aarch64:
ninja -C build-aarch64
qemu-aarch64-static -L /usr/aarch64-linux-gnu/ \
build-aarch64/test/test_ssimulacra2_simd # 6/6
- Follow-ups:
picture_to_linear_rgbSIMD — last scalar hot path in the extractor. 2 calls / frame. Low ROI but mechanical.- T3-3 SSIMULACRA 2 snapshot-JSON regression test — still pending.
0053 — SSIMULACRA 2 SIMD bit-exact ports (ADR-0161)¶
- ADR: ADR-0161
- Upstream source: fork-local. Upstream Netflix/vmaf has no SSIMULACRA 2 extractor at all (fork-added in ADR-0130).
- Touches:
- ssimulacra2_avx2.c / .h — 5 AVX2 kernels + per-lane
cbrtfhelper. - ssimulacra2_avx512.c / .h — 5 AVX-512 kernels; mechanical 16-wide widening of the AVX2 path.
- ssimulacra2_neon.c / .h — 5 NEON kernels; 4-wide aarch64 mirror.
- ssimulacra2.c — adds function-pointer dispatch fields to
Ssimu2State+init_simd_dispatch()helper, calls go through the pointers. - meson.build — registers the three SIMD TUs in
x86_avx2_sources/x86_avx512_sources/arm64_sources. - test_ssimulacra2_simd.c and
test/meson.build— new bit-exact test harness. - Invariants (load-bearing):
- Byte-for-byte bit-exactness to scalar on all 5 vectorised kernels under
FLT_EVAL_METHOD == 0. Regression caught pre- merge: naïve pairing(a+b)+(c+d)vs scalar((a+b)+c)+ddrifts by 1 ULP. Keep sequential scalar-order chains in all three SIMD TUs on rebase. cbrtfis per-lane scalar libm, not a polynomial. Any replacement with a vector cbrt would drift the ssimulacra2 score and break the regression test. Keep the spill/reload pattern.ssim_map/edge_diff_mapreductions use the ADR-0139 per-lanedoublescalar tail. Do NOT SIMD-reduce float lanes then lift to double — summation order changes.downsample_2x2deinterleave uses ISA-appropriate ops: AVX2vshufps+vpermpd, AVX-512vpermt2ps, NEONvuzp1q_f32+vuzp2q_f32. After deinterleave, sum order is((r0e+r0o)+r1e)+r1omatching scalar.#pragma STDC FP_CONTRACT OFFat every TU header. Ignored by aarch64 GCC (non-fatal-Wunknown-pragmas); kept for portability (clang, MSVC).- IIR blur +
picture_to_linear_rgbstay scalar in this PR. Follow-up PRs target these; when they land, re-verify bit-exactness viatest_ssimulacra2_simdexpansion. - Runtime dispatch order: AVX-512 > AVX2 on x86; NEON on aarch64; scalar fallback. Preserve on rebase.
- On upstream sync:
- Upstream has no SSIMULACRA 2 extractor; nothing to merge.
- If Netflix adopts SSIMULACRA 2 in the future, diff their implementation against the fork's scalar + SIMD TUs; keep the fork's bit-exactness contract absent a specific Netflix-authority carve-out ADR.
- Re-test on rebase:
meson setup build -Denable_cuda=false -Denable_sycl=false
ninja -C build
meson test -C build test_ssimulacra2_simd # 5/5
clang-tidy -p build core/src/feature/x86/ssimulacra2_avx2.c \
core/src/feature/x86/ssimulacra2_avx512.c
# aarch64:
ninja -C build-aarch64
qemu-aarch64-static -L /usr/aarch64-linux-gnu/ \
build-aarch64/test/test_ssimulacra2_simd # 5/5
clang-tidy -p build-aarch64 \
core/src/feature/arm64/ssimulacra2_neon.c
- Follow-ups:
- IIR blur vectorisation (
blur_planevertical-pass column batching) — the biggest frame-level wallclock win. picture_to_linear_rgbper-lanepowf— lower ROI but mechanical.- T3-3 SSIMULACRA 2 snapshot-JSON regression test — ADR-0130 deferred; still pending.
0052 — psnr_hvs SIMD bit-exact ports (ADR-0159 AVX2, ADR-0160 NEON)¶
- ADRs: ADR-0159 (AVX2), ADR-0160 (NEON sister port).
- Upstream source: fork-local. Upstream Netflix/vmaf has no psnr_hvs SIMD path.
- Touches:
core/src/feature/x86/psnr_hvs_avx2.c— AVX2 TU.core/src/feature/x86/psnr_hvs_avx2.h— AVX2 header.core/src/feature/arm64/psnr_hvs_neon.c— NEON TU (sister port, ADR-0160).core/src/feature/arm64/psnr_hvs_neon.h— NEON header.core/src/feature/third_party/xiph/psnr_hvs.c— addPsnrHvsState+ runtime dispatch ininit()(AVX2 underARCH_X86, NEON underARCH_AARCH64) + scoped NOLINTBEGIN/END around the upstream Xiph scalar block (kept verbatim as the bit-exact reference).core/src/meson.build— addx86/psnr_hvs_avx2.ctox86_avx2_sourcesandarm64/psnr_hvs_neon.ctoarm64_sources.core/test/test_psnr_hvs_avx2.c,core/test/test_psnr_hvs_neon.c— bit-exact unit tests (x86 and aarch64 respectively).core/test/meson.build— register both tests underenable_asm, arch-gated.- Invariants (load-bearing):
- Bit-exactness to scalar: every
od_coeff(int32) and every finalpsnr_hvs_{y,cb,cr,psnr_hvs}value the AVX2 path emits must be byte-identical to the scalar reference on the Netflix golden pairs. If a rebase introduces any pattern that breaks this (e.g. a floating-point horizontal reduce in the mask accumulator), the unit testtest_psnr_hvs_avx2will fail — don't relax the assertions; fix the SIMD path. - DCT butterfly layout:
butterfly → transpose → butterfly → transpose. The transpose lives insideod_bin_fdct8x8_avx2. Do not move it. - Float accumulators stay scalar: means / variances / mask / error accumulation in
calc_psnrhvs_avx2use the same per-block scalar loop as scalar psnr_hvs — bit-exact by construction. Do not vectorize these with horizontal reductions without replicating ADR-0139's per-lane scalar-float reduction pattern. The cross-block error accumulatorretis threaded throughaccumulate_error()by pointer, not returned-then-summed: each of the 64 per-coefficient contributions per block must hit the outerretdirectly, matching scalar's inlineret += ...atthird_party/xiph/psnr_hvs.cline 355. IEEE-754 float add is non-associative — summing into a local float and then adding the per-block total toretchanges the summation tree and drifts the Netflix golden by ~5.5e-5. #pragma STDC FP_CONTRACT OFFat the TU header disables FMA formation. Required:fmaf(a, b, c)can differ from(a*b)+cby 1 ulp, breaking bit-exactness. Do not remove the pragma; do not add-ffp-contract=fastto the build flags for this TU.- NOLINT suppressions are load-bearing — each cites ADR-0141 inline (bit-exactness scalar-diff auditability for the 30-butterfly function, scalar float→double promotion for
sqrt, extractor-registry extern linkage forvmaf_fex_psnr_hvs, upstream-Xiph scoped block for rebase parity). - On upstream sync:
- Upstream has no psnr_hvs SIMD as of 2026-04-24. Keep fork's version on conflict.
- If upstream ever touches
psnr_hvs.cfor non-SIMD reasons (e.g. a masking-table update), rebase the AVX2 TU to match line-for-line and re-runtest_psnr_hvs_avx2to confirm bit-exactness survives. - NEON follow-up PR is a sister port; its
arm64/psnr_hvs_neon.cwill mirror this ADR's invariants. On rebase, the two SIMD TUs must stay in lock-step with the scalar reference. - Re-test on rebase:
ninja -C build
meson test -C build test_psnr_hvs_avx2
# Expect: 5/5 subtests pass (DCT bit-exact on 3 random seeds +
# delta + constant input).
# CLI-level bit-exactness on Netflix golden (requires the YUV
# fixtures in python/test/resource/yuv/):
# VMAF_CPU_MASK=0 (scalar)
# VMAF_CPU_MASK=255 (AVX2 enabled)
# Diff per-frame psnr_hvs_{y,cb,cr,psnr_hvs} XML fields; expect
# byte-identical across all 3 golden pairs.
0051 — Netflix#1486 motion updates verified present (ADR-0158)¶
- ADR: ADR-0158
- Upstream source: Netflix upstream PR #1486 ("Port motion updates"), MERGED 2026-04-20 as commits
a44e5e6(code) +62f47d5(Netflix golden updates). - Touches: documentation-only; the actual code changes this ADR documents are already in the fork's master via earlier incremental motion3 / blend / five-frame-window commits.
- Invariants (load-bearing for future
/sync-upstream): - The
edge_8mirror fix (i_tap = height - (i_tap - height + 2)) is present atinteger_motion.c:240,x86/motion_avx2.c:147,x86/motion_avx512.c:147. If upstream's mirror line ever diverges again, this is the hunk to watch. - The
motion_max_valfeature option is atinteger_motion.c:57,118-120with default 10000.0 andFEATURE_PARAMflag. Upstream's default = fork's default; don't drift. VMAF_integer_feature_motion3_scoreoutput plumbing is ininteger_motion.c+alias.c.- Fork-local motion extensions (five-frame-window, moving-average, blend, fps_weight) are ADDITIONS on top of Netflix#1486. They are not upstream. Upstream changes to motion extractor internals may conflict with them — diff against
core/src/feature/integer_motion.con every rebase and check that the fork'sMIN(s->score * s->motion_fps_weight, s->motion_max_val)invocations are preserved (lines ~409, ~503). - On upstream sync: nothing to port from Netflix#1486 — it's absorbed. If a future upstream PR touches the same code paths, prefer upstream's version for the scalar/edge handling and the fork's version for the five-frame-window / blend extensions.
- Re-test on rebase:
ninja -C build
meson test -C build
# Expect: 35/35 pass.
# Verify the upstream markers are still in place after rebase:
grep -n "height - (i_tap - height + 2)\|motion_max_val\|VMAF_integer_feature_motion3_score" \
core/src/feature/integer_motion.c \
core/src/feature/alias.c \
core/src/feature/x86/motion_avx2.c \
core/src/feature/x86/motion_avx512.c
# Expect: matches at all 4 files. If any missing, the rebase
# silently dropped the Netflix#1486 content — investigate.
0050 — CUDA preallocation memory leak fix + vmaf_cuda_state_free (ADR-0157)¶
- ADR: ADR-0157
- Upstream source: Netflix upstream issue #1300 (OPEN since 2024; no maintainer fix as of 2026-04-24). User reports GPU memory rises monotonically across init/preallocate/fetch/close cycles.
- Touches:
core/include/libvmaf/libvmaf_cuda.h— new publicvmaf_cuda_state_free()API declaration.core/src/cuda/common.c— newvmaf_cuda_state_free()implementation;vmaf_cuda_release()now callscuda_free_functions();vmaf_cuda_state_init()gets an outer failure unwind;init_with_primary_context()releases the retained primary context onfail_after_pop.core/src/cuda/ring_buffer.c(since folded into the per-stream dispatch + drain machinery; seecore/src/cuda/dispatch_strategy.candcore/src/cuda/drain_batch.c) —vmaf_ring_buffer_close()then unlocked + destroyed the mutex before freeing.core/test/test_cuda_preallocation_leak.c— new GPU-gated reducer (10-cycle loop with full cleanup).core/test/test_cuda_pic_preallocation.c,core/test/test_cuda_buffer_alloc_oom.c— add missingvmaf_cuda_state_free()+vmaf_model_destroy()calls aftervmaf_close()in every test that allocates these.core/test/meson.build— register the new reducer underenable_cudaguard.- Invariants (load-bearing):
- Public contract: every caller of
vmaf_cuda_state_init()MUST callvmaf_cuda_state_free()AFTERvmaf_close()on any VmafContext that imported the state. Informalfree(cu_state)is a silent double-free hazard AFTER close (vmaf_close's vmaf_cuda_release already memset's + frees CudaFunctions internals; vmaf_cuda_state_free only frees the heap allocation itself). vmaf_cuda_release()freesCudaFunctionsvia a saved pointer AFTER thememset. Order matters —memsetfirst socu_state->fis zeroed in the caller's struct, then free via the saved local. Do not re-order.vmaf_ring_buffer_close()unlocks BEFORE destroying the mutex (POSIX requires the mutex be unlocked for destroy).- The cold-start unwind in
init_with_primary_contextreleasescuDevicePrimaryCtxRetain's retained context ifcuStreamCreateWithPriorityfails. - The ADR-0122 / ADR-0123
is_cudastate_empty()null-guards at the top of every publicvmaf_cuda_*entry must continue to compose with the newvmaf_cuda_state_free()(which accepts NULL directly and doesn't call through to the CUDA API). - The new free call order in callers is:
vmaf_close(vmaf)→vmaf_cuda_state_free(cu_state)→vmaf_model_destroy(model). Reversing the first two produces a use-after-free. - On upstream sync:
- Upstream has no
vmaf_cuda_state_free()as of 2026-04-24. Keep the fork's version on any conflict. If upstream eventually lands the same API with a different spelling, prefer upstream's spelling and add a compat alias — but do not break the fork's ABI. vmaf_cuda_release()'scuda_free_functions()call is fork-local. On rebase, keep it.- The ring-buffer
pthread_mutex_unlock+pthread_mutex_destroypair is fork-local. On rebase, keep it. - If upstream refactors
VmafCudaStateownership semantics (unlikely — their pattern has been "leaked state in a long- lived process is acceptable" historically), re-audit this ADR and the new public API. - Re-test on rebase:
ninja -C core/build-cuda
meson test -C core/build-cuda
# Expect: 40/40 pass including test_cuda_preallocation_leak.
# ASan leak-check:
cd libvmaf && meson setup build-asan-cuda \
-Db_sanitize=address -Denable_cuda=true -Denable_sycl=false \
--buildtype=debug
ninja -C build-asan-cuda
ASAN_OPTIONS='detect_leaks=1:leak_check_at_exit=1' \
build-asan-cuda/test/test_cuda_preallocation_leak
# Expect: 0 bytes leaked from core/src/* frames.
# (~180 bytes in libcuda.so.1 is expected — driver's process-
# lifetime cuInit cache, does not grow per cycle.)
0049 — CUDA graceful error propagation (ADR-0156)¶
- ADR: ADR-0156
- Upstream source: Netflix upstream issue #1420 (OPEN as of 2026-04-24). Reports that two concurrent VMAF-CUDA processes crash the second one at
vmaf_cuda_buffer_allocdue toCHECK_CUDA(cuMemAlloc)→assert(0)on OOM. - Touches:
core/src/cuda/cuda_helper.cuh— redefinedCHECK_CUDAfamily. New macrosCHECK_CUDA_GOTO+CHECK_CUDA_RETURN+ helpervmaf_cuda_result_to_errno. Oldassert(0)semantics removed entirely.core/src/cuda/common.c,core/src/cuda/picture_cuda.c,core/src/libvmaf.c— allCHECK_CUDA(...)sites converted; cleanup labels added where contexts / buffers were pushed / allocated.core/src/feature/cuda/integer_motion_cuda.c,integer_vif_cuda.c,integer_adm_cuda.c— same conversion; 12statichelpers promotedvoid → int.core/test/test_cuda_buffer_alloc_oom.c— new GPU-gated reducer.core/test/meson.build— register new test underenable_cudaguard.- Invariants (load-bearing):
CHECK_CUDA_GOTO/CHECK_CUDA_RETURNmust never callassert(0)orabort()on a CUDA error. Any regression back to the upstream abort-on-error semantics re-introduces Netflix#1420 and the NDEBUG footgun.- Every
CHECK_CUDA_GOTOtarget label must pop any previously-pushed CUDA context and free any partially-constructed buffers before returning the errno. The graceful path must not leak resources. vmaf_cuda_result_to_errnouses numericCUresultvalues directly (0 / 1 / 2 / 3 / 4 / 101 / 201 / 400) so host TUs that don't include<cuda.h>can transitively consume the mapping via the inline function. If upstream renumbersCUresultenum values (historically stable — they've been fixed since CUDA 1.0), re-audit the switch.- ADR-0122 / ADR-0123
is_cudastate_empty(...)guards at the top of every publicvmaf_cuda_*entry point must stay — they run before the CUDA API is touched and compose cleanly with the new error propagation. - Twelve
statichelper signatures in the feature extractors areint-returning (wasvoid): any upstream-port that restores thevoidreturn silently regresses the error path. - On upstream sync:
- Upstream Netflix still uses
assert(0)inCHECK_CUDAas of 2026-04-24. Keep the fork's macro definitions incuda_helper.cuhon any upstream conflict — this file is fork-local behaviour. - If upstream eventually lands Netflix#1420 with a similar refactor, prefer the fork's version unless upstream's has identical semantics (no
assert(0)/ noabort()/ translatesCUresultto-errno). Re-verifytest_cuda_buffer_alloc_oomafter rebase. - If upstream adds new
CHECK_CUDA(...)sites in a port, rewrite them toCHECK_CUDA_GOTO/CHECK_CUDA_RETURNas part of the port commit. - If upstream changes any of the 12
statichelper signatures back tovoid, re-promote them tointduring the merge. - Re-test on rebase:
ninja -C core/build-cuda
meson test -C core/build-cuda
# Expect: 39/39 pass including test_cuda_buffer_alloc_oom.
# Reducer check — verify the OOM-to-errno path is live:
meson test -C core/build-cuda test_cuda_buffer_alloc_oom -v
# Expect subtests: request 1 TiB → -ENOMEM; request 0 bytes → 0.
clang-tidy -p core/build-cuda --quiet \
core/src/cuda/common.c \
core/src/cuda/picture_cuda.c \
core/src/feature/cuda/integer_motion_cuda.c \
core/src/feature/cuda/integer_vif_cuda.c \
core/src/feature/cuda/integer_adm_cuda.c \
core/src/libvmaf.c
# Expect exit 0 on every file.
0049 — compute_motion / picture_copy signature changes (b949cebf upstream port)¶
- Upstream commit: Netflix/vmaf b949cebf (feature/motion: port several feature extractor options)
- Prerequisite commit: Netflix/vmaf d3647c73 (picture_copy: add channel parameter)
- PR: upstream/port-b949cebf-motion
Rebase-sensitive invariants:
-
compute_motionsignature change —compute_motion()incore/src/feature/motion.c/motion.hnow takes an extraint motion_decimateparameter (themotion_add_scale1flag). Any new caller added in the fork that callscompute_motion()must pass this parameter. The SIMD integer motion callers (motion_avx2.c,motion_avx512.c) do NOT callcompute_motion()— they use the SAD/convolution dispatch table directly and are unaffected. -
vmaf_image_sad_csignature change — similarly gainsint motion_add_scale1. Any caller in the fork must be updated. Currently only called fromcompute_motion()internally. -
picture_copysignature change — gainsint channelas the last parameter (0=Y, 1=U, 2=V). Every caller in the tree has been updated to pass0(luma). When adding new callers that need UV planes, pass1or2. The fork's CUDA/SYCL/Vulkan callers have been updated in this PR. -
Default behavior preserved — all new options default to no-op values.
motion_add_scale1=false,motion_add_uv=false,motion_blend_factor=1.0,motion_fps_weight=1.0,motion_filter_size=5(= DEFAULT_MOTION_FILTER_SIZE). Integer and float motion2 scores are bit-identical to pre-port baseline. -
vif_scale_frame_sdependency avoided — the upstream b949cebf motion.c importsvif_scale_frame_sfrom vif_tools.h. The fork does not have this function yet (vif options chain is deferred, Research-0024 Strategy E). The bilinear downscaler formotion_add_scale1is implemented as local static functions inmotion.c(motion_scale_bilinear,motion_bilinear_interp,motion_mirror_f). When upstream's vif options chain is eventually ported, reconcile by replacing these local functions withvif_scale_frame_s.
Reproducer:
# verify bit-exactness (default options, scores must be identical):
./core/build/tools/vmaf \
--reference testdata/ref_576x324_48f.yuv \
--distorted testdata/dis_576x324_48f.yuv \
--width 576 --height 324 --pixel_format 420 --bitdepth 8 \
--model path=model/vmaf_v0.6.1.json \
--feature motion --no_prediction --json --output /tmp/motion.json
# integer_motion2 scores must match pre-port baseline at 6 decimal places.
0048 — i4_adm_cm int32 rounding overflow deliberately preserved (ADR-0155)¶
- ADR: ADR-0155
- Upstream source: Netflix upstream issue #955 (OPEN since 2020; no maintainer response as of 2026-04-24). Reports that
add_bef_shift_flt[idx] = (1u << (shift_flt[idx] - 1))incore/src/feature/integer_adm.cscales 1–3 overflowsint32_t(1u << 31 = 0x80000000wraps to-2147483648). Rounding term is sign-negated; ADM scales 1–3 biased low by ≈1 LSB per summed term. - Touches (documentation-only):
docs/adr/0155-adm-i4-rounding-deferred-netflix-955.md— new ADR (this entry's anchor).core/src/feature/integer_adm.c— in-file warning comment above the overflow site (add_bef_shift_flt[]initialiser loop around line 1277). No code change.core/src/feature/AGENTS.md— invariant note under "Rebase-sensitive invariants".- Invariants (load-bearing — do NOT silently "fix"):
integer_adm.ckeepsint32_t add_bef_shift_flt[3]with the overflowing1u << 31assignment. The Netflix golden assertions (python/test/quality_runner_test.py,vmafexec_test.py,feature_extractor_test.py) encode the buggy ADM output. Project hard rule #1 (ADR-0024) prohibits changing those assertions.- Any "fix" that changes ADM numerical output must land together with a coordinated Netflix-authored golden-number update (the ADR-0142 Netflix-authority carve-out). Until Netflix#955 closes upstream, there is no authority to track.
- On upstream sync:
- If Netflix finally lands a fix for #955 (widening the rounding term to
uint32_torint64_t), sync the C-side fix AND the updatedassertAlmostEqualvalues in the same merge. Re-runmake test-netflix-goldenand/cross-backend-diffon the golden pairs to verify the new numbers are consistent across CPU / CUDA / SYCL. - Remove the in-file warning comment above the
add_bef_shift_fltinitialiser loop, flip ADR-0155 toSuperseded by ADR-NNNN, and drop this rebase-notes entry. - If upstream instead closes #955 as wont-fix, keep this entry verbatim and update the ADR status to note upstream's closure.
- Re-test on rebase (gates the invariant by confirming the golden numbers are unchanged):
ninja -C build
make test-netflix-golden
# Expect: VMAF mean 76.66890… on src01_hrc00/01_576x324 golden
# pair — bit-identical to pre-rebase.
0047 — vmaf_score_pooled -EAGAIN for pending features (ADR-0154)¶
- ADR: ADR-0154
- Upstream source: Netflix upstream issue #755 (OPEN as of 2026-04-24). Upstream maintainer closed the door on the streaming use case in 2020 ("you cannot call vmaf_score_pooled() in a loop"); fork reopens it via error-code semantics without changing the retroactive-write design.
- Touches:
core/src/feature/feature_collector.c—vmaf_feature_collector_get_scorereturns-EAGAIN(was-EINVAL) when the requested index is valid but not yet written.core/src/feature/feature_collector.h— inlinevmaf_feature_vector_get_scorenow returns-EINVALfor null/out-of-range and-EAGAINfor not-written (was-1for both). Added#include <errno.h>. Rename reserved__VMAF_FEATURE_COLLECTOR_H__guard toVMAF_FEATURE_COLLECTOR_INCLUDED.core/test/test_score_pooled_eagain.c— new 4-subtest reducer.core/test/meson.build— register the new test.- Invariants (load-bearing, enforced by the reducer):
vmaf_feature_collector_get_score(fc, name, &score, i)returns-EAGAINiff the featurenameis registered andiis in range butscore[i].written == false.- The return stays
-EINVALfor (a) null pointers, (b)i >= feature_vector->capacity, (c) unknown feature name. - The inline fast-path
vmaf_feature_vector_get_scoreuses the same split. - On upstream sync: upstream has not changed the error semantics since 2020. If they do (unlikely), keep the fork's
-EAGAIN— it is strictly more informative and downstream code depending on the split would regress. - Re-test on rebase:
ninja -C build && meson test -C build test_score_pooled_eagain
# Expect: 4/4 subtests pass.
# Reducer check:
git stash push core/src/feature/feature_collector.c core/src/feature/feature_collector.h
ninja -C build && meson test -C build test_score_pooled_eagain
# Expect: Fail: 1 (tests fail without -EAGAIN split).
git stash pop
0046 — float_ms_ssim min-dim guard (ADR-0153)¶
- ADR: ADR-0153
- Upstream source: Netflix upstream issue #1414 (OPEN as of 2026-04-24). No upstream fix has landed; fork adds the guard independently.
- Touches:
core/src/feature/float_ms_ssim.c— add#include "log.h"+#include "iqa/ssim_tools.h"+ amin_dim = GAUSSIAN_LEN << (SCALES - 1)check at the start ofinit; extract SIMD dispatch into a newms_ssim_init_simd_dispatchhelper to keepinitwithin the ADR-0141 60-line budget.core/test/test_float_ms_ssim_min_dim.c— new 3-subtest reducer.core/test/meson.build— register the new test executable.- Invariant (load-bearing, enforced by the reducer):
float_ms_ssim.initreturns-EINVALwhenw < 176 || h < 176, where 176 is computed dynamically from the filter constants. The magic number is not hardcoded — changingSCALESorGAUSSIAN_LENupstream will auto-update the minimum. - On upstream sync: if Netflix upstream lands a similar init-time guard, keep the fork's version — the helper name
ms_ssim_init_simd_dispatchis fork-local (introduced to satisfy ADR-0141) and upstream's patch won't match. Both guards should be compatible; re-verify the reducer after rebase. - Re-test on rebase:
ninja -C build && meson test -C build test_float_ms_ssim_min_dim
# Expect: 3/3 subtests pass.
# Reducer check (confirms the guard is load-bearing):
git stash push core/src/feature/float_ms_ssim.c
ninja -C build && meson test -C build test_float_ms_ssim_min_dim
# Expect: Fail: 1 (tests fail without the guard).
git stash pop
0045 — vmaf_read_pictures monotonic-index guard (ADR-0152)¶
- ADR: ADR-0152
- Upstream source: Netflix upstream issue #910 (OPEN as of 2026-04-24). No upstream fix has landed; the fork adds the guard independently, per the 2021-10-14 maintainer comment that recommended exactly this shape.
- Touches:
core/src/libvmaf.c— addunsigned last_index+bool have_last_indexfields toVmafContext; prepend a monotonic-index check insideread_pictures_validate_and_prep(returns-EINVALon duplicates / regressions); update the two new fields at the tail of the same helper on success.core/test/test_read_pictures_monotonic.c— new 3-subtest reducer covering the Netflix#910 sequence and the two classes of rejection (duplicate, out-of-order).core/test/meson.build— register the new test executable.- Invariant (load-bearing, enforced by the reducer):
vmaf_read_pictures(vmaf, ref, dist, index)returns-EINVALwhenhave_last_index && index <= last_index. Flush (vmaf_read_pictures(vmaf, NULL, NULL, 0)) routes toflush_contextbefore the guard runs — flushing remains always-available independent of the last accepted index. - On upstream sync:
- If Netflix upstream eventually lands a similar guard at the API boundary, keep the fork's version — the helper function name (
read_pictures_validate_and_prep) is fork-local (ADR-0146), upstream's patch will target a different insertion point. Both guards should be compatible; re-verify the reducer after rebase. - If upstream instead lands an internal reordering mechanism (buffer-and-sort frames before dispatch), revisit this decision — the fork's API-level contract is stricter and may need to relax to match. Open a new ADR if so.
- Re-test on rebase:
ninja -C build && meson test -C build test_read_pictures_monotonic
# Expect: 3/3 subtests pass.
# Reducer check (confirms the guard is load-bearing):
git stash push core/src/libvmaf.c
ninja -C build && meson test -C build test_read_pictures_monotonic
# Expect: Fail: 1 (the test rejects the un-guarded behaviour).
git stash pop
0044 — i686 (32-bit x86) build-only CI job (ADR-0151)¶
- ADR: ADR-0151
- Upstream source: Netflix upstream issue #1481 (OPEN as of 2026-04-24). Reports i686 compile failure on
_mm256_extract_epi64. Workaround documented in the issue:-Denable_asm=false. - Touches:
build-aux/i686-linux-gnu.ini— new cross-file; gcc +-m32+cpu_family = 'x86'/cpu = 'i686'. Noexe_wrapper..github/workflows/libvmaf-build-matrix.yml— new matrix row withi686: trueflag + new install-deps step forgcc-multilib+g++-multilib; existing "Run tests" + "Run tox tests (ubuntu)" steps widened with&& !matrix.i686guards.- Invariants:
- The i686 matrix row pins
-Denable_asm=false— this is the upstream-documented workaround for_mm256_extract_epi64's missing declaration on 32-bit x86 targets. Do NOT remove the flag without first gating every_mm256_extract_epi64call site incore/src/feature/x86/adm_avx2.c+motion_avx2.c+adm_avx512.con__x86_64__. Removing the flag naively will re-break the build. - No
exe_wrapperin the cross-file: meson marks tests asSKIP 77even though the host can run i686 binaries natively. Build-only gate by design. - On upstream sync:
- If upstream Netflix fixes #1481 at source (by gating the intrinsic calls on
__x86_64__or by emulating via two_mm256_extract_epi32halves), sync the fix and re-enable ASM on the i686 row (drop-Denable_asm=falsefrommeson_extra). Re-verify bit-exactness via/cross-backend-diffon the x86_64 golden pair. - If upstream marks i686 unsupported in meson (e.g. via a hard error), the fork's i686 row should be removed or downgraded to
continue-on-error: true. - Re-test on rebase (Ubuntu host with
gcc-multilib):
meson setup libvmaf core/build-i686 \
--cross-file=build-aux/i686-linux-gnu.ini \
-Denable_asm=false \
-Denable_cuda=false -Denable_sycl=false
ninja -C core/build-i686
file core/build-i686/tools/vmaf
# Expect: ELF 32-bit LSB pie executable, Intel i386
CI runs this same sequence via the new matrix row.
0058 — Tiny-AI Netflix corpus training scaffold (ADR-0252)¶
- ADR: ADR-0252.
- Upstream source: fork-local. Netflix/vmaf has no tiny-AI training harness or MCP server.
- Touches:
ai/— training harness;NflxLocalDatasetloader reads from--data-root(never from a hardcoded path).docs/ai/training-data.md— corpus path convention and loader API docs; purely additive.mcp-server/vmaf-mcp/tests/test_smoke_e2e.py— new e2e smoke test; references only committed golden fixtures.- Invariants (load-bearing):
- Data path is local-only.
.workingdir2/netflix/is gitignored; no YUV from this corpus is ever committed. The--data-rootCLI flag must remain the sole mechanism for locating the corpus. - Smoke test uses only committed fixtures.
test_smoke_e2e.pyreferencespython/test/resource/yuv/src01_hrc00_576x324.yuv(a committed golden file), never the local corpus path. On upstream sync the golden YUV path must stay stable. - No Netflix golden assertion is modified. The
places=4tolerance intest_smoke_e2e.pyasserts against thevmaf_v0.6.1CPU reference; it is not a golden assertion and may be adjusted by/regen-snapshotswith justification. - On upstream sync: zero interaction with Netflix upstream. The
ai/subtree andmcp-server/are wholly fork-local; upstream merges are conflict-free here. If Netflix ever ships a training harness, reconcile separately. - Re-test on rebase:
cd mcp-server/vmaf-mcp && python -m pytest tests/test_smoke_e2e.py -v
# Requires: meson compile -C build (vmaf binary)
# Skips automatically if binary or golden YUV is absent.
0085 — Research-0030 Phase-3b multi-seed validation (Gate 1 passed)¶
- No ADR. Empirical research digest closing Gate 1 of the 3-gate v2 validation chain. Architecture decision unchanged.
- Upstream source: fork-local. Netflix has no multi-seed validation surface for tiny-AI training.
- Touches (additive only):
docs/research/0030-phase3b-multiseed-validation.md— per-seed PLCC tables + stability analysis + Gate 2/3 plan.ai/scripts/phase3_subset_sweep.py— adds--seedsflag (comma-separated list) + per-seed result aggregation.CHANGELOG.mdUnreleased § Added.- Invariants (rebase-relevant):
- The +0.0175 Δ is multi-seed mean PLCC, not seed-0 PLCC. Don't cite the +0.0106 from Research-0029 once Research-0030 lands; the multi-seed number is more trustworthy.
- Subset B is more stable than canonical-6 across seeds. Don't ship a v2 model citing single-seed numbers — always report multi-seed mean ± seed-mean-std for any tiny-AI metric in a future digest.
- The
--seedsflag aggregates by flattening (seed × fold) pairs. The reportedmean_plccis the mean of alln_seeds × n_foldsmeasurements;seed_mean_plcc_stdis the std across per-seed means, which is the right number for "is the result seed-stable". - On upstream sync: zero interaction. Fork-only research.
- Re-test on rebase: documentation-only PR; the runs/ files reproduce from the canonical command.
0084 — Research-0029 Phase-3b StandardScaler retry (positive result)¶
- No ADR. Empirical research digest; revives the Research-0026 hypothesis after the Research-0028 negative result. The architectural decision (ship
vmaf_tiny_v2) is gated on three validation steps documented in the digest §"Required before shipping". - Upstream source: fork-local. Netflix has no tiny-AI preprocessing-sensitivity analysis surface.
- Touches (additive only):
docs/research/0029-phase3b-standardscaler-results.md— per-fold tables + apples-to-apples comparison + 3-gate pre-shipping checklist.ai/scripts/phase3_subset_sweep.py— adds--standardizeflag +_standardize_inplacehelper.CHANGELOG.mdUnreleased § Added.- Invariants (rebase-relevant):
- StandardScaler statistics MUST be fit per-fold on the train split only. Fitting on the full data would leak held-out information into LOSO; the
_standardize_inplacehelper enforces this by taking only the train slice as input. - A shipped
vmaf_tiny_v2.onnxMUST bundle its scaler(mean, std)in the sidecar JSON — otherwise inference applies different normalisation than training and the win evaporates. Currently UN-implemented; tracked as a §"Caveats" #5 follow-up. - Subset B's feature list is the load-bearing finding:
adm2,adm_scale3,vif_scale2,motion2,ssimulacra2,psnr_hvs,float_ssim. Phase-3c experiments may shift the optimal arch / lr / epochs but should keep this set. - On upstream sync: zero interaction. Fork-only research.
- Re-test on rebase: documentation-only PR; the runs/ files are reproducible from the
--standardizeinvocation in §"Reproducer".
0082 — Research-0028 Phase-3 subset sweep (negative-result digest)¶
- No ADR. Empirical research digest. The architectural decision (no v2 model ships from this Phase) is governed by Research-0027's pre-registered stopping rule.
- Upstream source: fork-local. Netflix has no tiny-AI subset- sweep surface.
- Touches (additive only):
docs/research/0028-phase3-subset-sweep.md— per-fold tables- headline + standardisation caveat + Phase-3b/c/d follow-ups.
CHANGELOG.mdUnreleased § Added.- Invariants (rebase-relevant):
- canonical-6 stays the default until Phase-3b lands a ≥ 0.005 PLCC win (per Research-0027 stopping rule).
- The PLCC drop is most likely a feature-scale issue, not evidence the new features lack signal. Don't cite this digest to retire
ssimulacra2/adm_scale3from the candidate pool; re-test withStandardScalerfirst. - Phase-3 results are seed=0 only. Any v2-shipping decision needs 3-seed mean±std and KoNViD cross-check.
- On upstream sync: zero interaction. Fork-only research.
- Re-test on rebase: documentation-only PR; runs/ files are reproducible from the canonical command in §"Reproducer".
0081 — Research-0027 Phase-2 feature importance results¶
- No ADR. Empirical research digest closing Research-0026 Phase 2; the architectural decision (Subset A / B / C) is deferred to Phase-3 results in a future digest.
- Upstream source: fork-local. Netflix has no cross-metric feature-importance analysis surface.
- Touches (additive only):
docs/research/0027-phase2-feature-importance.md— per-method top-10 + consensus + redundancy + Phase-3 subset recommendations.CHANGELOG.mdUnreleased § Added.- Invariants (rebase-relevant):
- Consensus top-10 is the load-bearing finding:
adm2,adm_scale3,ssimulacra2,vif_scale2. Phase-3 candidate subsets MUST include all four. - The 11-pair redundancy table is corpus-specific — measurements on Netflix Public 9-source. KoNViD-1k cross- check is a Phase-3 prerequisite if Subsets B/C advance.
runs/full_features_netflix.parquetandruns/full_features_correlation.jsonstay gitignored. Reproducer in §"Reproducer" regenerates both.- On upstream sync: zero interaction. Fork-only research.
- Re-test on rebase: documentation-only PR; the
runs/files are reproducible from the canonical commands.
0080 — Phase-2 analysis scripts (Research-0026 Phase 2 prep)¶
- No ADR. Pure analysis scaffolding; the architectural decision (which features to ship in v2) is gated on Phase 2's numerical output via Research-0027.
- Upstream source: fork-local. Netflix has no tiny-AI training nor cross-metric correlation tooling.
- Touches (additive only):
ai/scripts/extract_full_features.py— parquet extractor over Netflix corpus withFULL_FEATURES. Per-clip JSON cache at$XDG_CACHE_HOME/vmaf-tiny-ai-full/<source>/<dis_stem>.json.ai/scripts/feature_correlation.py— Pearson + MI + LASSO- RF + consensus top-K analyser; outputs JSON.
ai/tests/test_feature_correlation.py— 5 pytest cases against synthetic parquet (no libvmaf dependency).CHANGELOG.mdUnreleased § Added.- Invariants (rebase-relevant):
- The per-clip JSON cache and the
FULL_FEATUREStuple must stay in lock-step. If the tuple grows (or shrinks), pre-existing cache files become stale and silently misalign their storedper_framecolumns with the new tuple. The extractor MUST be re-run with a cleared cache whenFULL_FEATURESchanges. Regression hint:test_default_features_unchangedintest_feature_sets.pyalready guards the canonical 6; extend coverage toFULL_FEATURESif rebases touch it. motion3resolves to extractormotion_v2in_METRIC_TO_EXTRACTOR, notmotion3(the upstream-canonical extractor name in the integer_motion_v2 module). The CLI--feature motion3does NOT exist. The JSON output key isinteger_motion3which_lookupfinds via theinteger_fallback.admandvifaggregates are NOT inFULL_FEATURES. The integer extractor emitsinteger_adm2andinteger_vif_scale0..3but no bareadm/vif. Listing them produced all-NaN columns in v1 — fixed in PR #185 amend.- On upstream sync: zero interaction. Pure fork-side analysis tooling.
- Re-test on rebase:
pytest ai/tests/test_feature_correlation.py ai/tests/test_feature_sets.py -v
# Expect: 14 passed in <1 s.
0079 — Tiny-AI feature-set registry (Research-0026 Phase 1)¶
- No ADR. Pure additive extension of an existing module; the architectural decision (which features, which model) lives in Research-0026's go/no-go gate after Phase 2.
- Upstream source: fork-local. Netflix/vmaf has no tiny-AI training pipeline.
- Touches (additive only):
ai/data/feature_extractor.py— addsFULL_FEATURES(21 entries),FEATURE_SETSregistry,resolve_feature_set()helper._METRIC_TO_EXTRACTORgrew 11 → 25 entries.ai/tests/test_feature_sets.py— new 9-test smoke suite.CHANGELOG.mdUnreleased § Added.- Invariants (rebase-relevant — these are load-bearing):
DEFAULT_FEATURESstays the canonical 6-tuple matchingvmaf_v0.6.1's SVR input layout. Testtest_default_features_unchangedis the regression guard; any quiet broadening would invalidate every shipped tiny-AI ONNX (input-dim baked into the model). If a future change must broaden the default, ship a paired model swap under ADR-0049 sidecar policy.FULL_FEATURESexcludeslpipsandfloat_momentper Research-0026 §"Open questions" Q1. Testtest_full_features_excludes_lpips_and_momentenforces. Adding either would re-classify the experiment from "tiny model on classical features" to "ensemble of DNNs".- Every entry in
FULL_FEATURESMUST have an entry in_METRIC_TO_EXTRACTOR. Testtest_every_full_feature_has_extractor_mappingis the guard — without the mapping the libvmaf CLI silently emits NaN columns for the missing metric. - On upstream sync: zero interaction. Fork-only training surface.
- Re-test on rebase:
0078 — Research-0026 cross-metric feature fusion plan¶
- No ADR. Pure research-plan digest; the architectural decision (which features to add) is deferred to Research-0027 follow-up after Phase 2 numbers land.
- Upstream source: fork-local. Netflix/vmaf has no tiny-AI training and no broader-feature-set hypothesis under investigation.
- Touches (additive only):
docs/research/0026-cross-metric-feature-fusion.md— 4-phase experimental plan + cost estimate + go/no-go criteria.CHANGELOG.mdUnreleased § Added.- Invariants (rebase-relevant):
- The 6-feature canonical baseline (
adm2,vif_scale0..3,motion2) stays the default. Any v2 model is opt-in via a newfeature_setfield in the sidecar JSON; existingvmaf_tiny_v1.onnxusers get the same numbers. lpipsis OUT of the candidate pool (Phase 1/2). It's DNN-based and would blur the line between "tiny model on classical features" and "ensemble of DNNs". Revisit only if classical features can't close the gap.- On upstream sync: zero interaction. Pure fork-side research planning.
- Re-test on rebase: documentation-only; no test surface.
0077 — Research-0025 FoxBird outlier resolved via KoNViD combined training¶
- No ADR. Empirical research digest closing the open question in Research-0023 §5; no architecture or policy decision. Pure documentation of an empirical result.
- Upstream source: fork-local. Netflix/vmaf has no tiny-AI training, no KoNViD-1k integration, and no LOSO eval surface.
- Touches (additive only):
docs/research/0025-foxbird-resolved-via-konvid.md— per-clip table + comparison to Netflix-only baselines + interpretation + caveats + next-experiment list.CHANGELOG.mdUnreleased § Added.- Invariants (rebase-relevant):
- The training-fit per-clip numbers in §"Per-clip result" are NOT held-out generalisation metrics — FoxBird is in the training set. The proper validation is the LOSO sweep on the combined corpus (§"Next experiments" #1). Don't cite the 0.9936 FoxBird PLCC as a generalisation number; cite it as "training-fit on combined corpus, 5.4× RMSE improvement vs Netflix-only".
- Combined trainer command line is canonical. The reproduction recipe in §"Setup" includes
--seed 0,--konvid-val-fraction 0.1,--val-source Tennis,--val-mode netflix-source-and-konvid-holdout. Changing any knob invalidates the per-clip numbers. runs/tiny_combined_canonical/stays gitignored. The final ONNX is reproducible from the parquet + Netflix corpus + the canonical CLI; the durable record is the digest's table.- On upstream sync: zero interaction. Research digest is fork-only.
- Re-test on rebase:
python ai/train/train_combined.py \
--netflix-root .workingdir2/netflix \
--konvid-parquet ai/data/konvid_vmaf_pairs.parquet \
--model-arch mlp_small --epochs 30 --batch-size 256 --lr 1e-3 \
--val-mode netflix-source-and-konvid-holdout \
--val-source Tennis --konvid-val-fraction 0.1 --seed 0 \
--out-dir runs/tiny_combined_canonical
# Expect: FoxBird PLCC ≈ 0.9936 ± 1e-3 (numerical-noise floor),
# mean PLCC ≥ 0.9983 across 9 Netflix clips.
0076 — Research-0024 vif/adm upstream-divergence digest (Strategy E doc)¶
- No ADR. Pure documentation digest; the divergence decisions it ratifies are already governed by ADR-0138 / 0139 / 0142 / 0143 (vif SIMD bit-exactness contract) and ADR-0024 (Netflix golden-data immutability). The digest itself fits the per-PR research-digest deliverable bar from ADR-0108.
- Upstream source: forward-looking — pre-emptively documents the fork's non-port of Netflix
4ad6e0ea/41d42c9e/bc744aa3/8c645ce3(vif chain) and4dcc2f7c(float_adm chain). Strategy A onb949cebfmotion chain stays approved. - Touches (additive only):
docs/research/0024-vif-upstream-divergence.md— 5-strategy decision matrix + numerical-risk analysis for each chain.core/src/feature/AGENTS.md— two new "rebase-sensitive invariants" entries pinning the vif and adm divergences.CHANGELOG.mdUnreleased § Changed.- Invariants (rebase-relevant — these are the whole point):
- Do not port
4ad6e0ea(vif runtime helpers) or8c645ce3(vif prescale options) verbatim. They replace the precomputedvif_filter1d_table_stable whose frozenconst floatGaussians make AVX2 == AVX-512 == NEON == scalar bit-for-bit. A future opt-in second-path port (Strategy C, runtime helpers behind--vif-prescale != 1) is allowed but must not touch the default code path. - Do not port
4dcc2f7cfloat_adm options chain. The 12-parametercompute_admsignature change cascades through SIMD (avx2 / avx512 / neon) and 3 GPU backends (vulkan / cuda / sycl). The newaimfeature has no fork- side golden values; defer until concrete user demand. - Mirror bugfix
41d42c9eis a separate decision. Must come paired withplaces=4 → places=3golden loosening per ADR-0142 Netflix-authority precedent. Not part of Strategy E; eligible for a focused single-purpose PR if any shipped model drifts more thanplaces=3because of the missing fix. b949cebfmotion chain port stays APPROVED under Strategy A (verbatim, float_motion-side only). Float_motion has no precomputed-table investment to protect; existing fork integer_motion already has 6/9 of these options; cheap to mirror onto float_motion.- On upstream sync: zero conflict — pure additions to research/ and AGENTS.md.
- Re-test on rebase: documentation-only PR; rendered markdown is the only verification surface.
# Re-run the diff scan that produced the digest (catches new
# upstream commits since 9dac0a59):
git fetch upstream && git log --pretty=format:'%h %s' \
upstream/master ^origin/master --since="2026-01-01" \
-- core/src/feature/{float_,integer_,}{vif,motion,adm,cambi}*.{c,h} \
core/src/feature/{vif,motion,adm,cambi}_options.h \
| head -30
# If new vif / adm option ports appear, update Research-0024 §"Same
# divergence test for motion + float_adm" before deciding to port.
0075 — Upstream 798409e3 + 314db130 ports (CUDA null-deref + remove all.c)¶
- No ADR. Pure upstream cherry-picks per ADR-0108 carve-out ("pure upstream syncs and
port-upstream-commitPRs are exempt"). - Upstream source:
798409e3(Lawrence Curtis, 2026-04-20): "Fix null deref crash on prev_ref update in pure CUDA pipelines"314db130(Kyle Swanson, 2026-04-28): "libvmaf/feature: remove empty translation unit all.c"- Touches (additive / removal only):
core/src/libvmaf.c— addsif (ref && ref->ref)guard beforevmaf_picture_ref(&vmaf->prev_ref, ref)at the two threaded paths (threaded_enqueue_oneline 1057 andthreaded_read_pictures_batchline 1105). Main path at line 1597 already has the guard.core/src/feature/all.c— file deleted.core/src/meson.build— drops thefeature_src_dir + 'all.c'line.core/src/feature/offset.c— updates the// NOLINTNEXTLINEcomment to dropall.cfrom the list of per-feature consumers.CHANGELOG.mdUnreleased § Fixed (798409e3) + § Changed (314db130).- Invariants (rebase-relevant):
- The fork has THREE prev_ref update sites; all need the
if (ref && ref->ref)guard. The mainvmaf_read_picturespath already had it (viaread_pictures_update_prev_refhelper); the threaded paths (#ifdef VMAF_BATCH_THREADING) inherited the unguarded shape from upstream's old code. Future upstream rebases must preserve all three guards even if Netflix refactors the threaded paths. all.cdeletion is symbol-safe. Allcompute_*functions it forward-declared are reached via per-extractor TUs that#includethe relevant<feature>.h. No external linker dependency onall.c's symbols.- On upstream sync: zero conflict expected — fork now matches upstream tip on these two surfaces.
- Re-test on rebase:
meson setup build-cpu libvmaf -Denable_cuda=false -Denable_sycl=false \
-Denable_vulkan=disabled
ninja -C build-cpu
meson test -C build-cpu # 37 tests, all pass.
0074 — Combined Netflix + KoNViD-1k trainer driver¶
- No ADR. Pure engineering follow-up; the architecture rationale is fully covered by ADR-0203 (training-prep architecture) and Research-0023 §5 (FoxBird-class outlier needs broader corpus).
- Upstream source: fork-local. Netflix/vmaf has no tiny-AI trainer.
- Stacks on the KoNViD-1k loader bridge (PR #178 / rebase-note 0073). Rebase order: land 0073 first.
- Touches (additive only):
ai/train/train_combined.py— concatenating trainer that reuses_build_model/_train_loop/export_onnxfromai/train/train.py.ai/tests/test_train_combined_smoke.py— 5 pytest cases (key splitter +--epochs 0paths, no libvmaf or real corpus required).docs/ai/training.md— "Combining KoNViD with the Netflix corpus" subsection rewritten from "follow-up" to runnable.CHANGELOG.mdUnreleased § Added.- Invariants (rebase-relevant):
- Reuse the canonical training-loop helpers. Don't fork
_build_model/_train_loop/export_onnxinto this file. Both trainers must share the model factory so a future change (e.g. addingmlp_large) lands in one place. - KoNViD train/val splits hold out whole clip keys, not random frames. A frame-level split would let frames from the same clip leak across train/val and inflate PLCC by 5-10 pp (well-known VQA pitfall — same reasoning as ADR-0203's Netflix 1-source-out split).
- Missing data falls back, not errors. Missing
--konvid-parquet→ Netflix-only path. Missing--netflix-root→ KoNViD-only path. Both missing → initial- weights ONNX export +rc=0so the smoke command always produces a deterministic artefact. - On upstream sync: zero interaction; pure fork-local trainer.
- Re-test on rebase:
pytest ai/tests/test_train_combined_smoke.py -v
# Expect: 5 passed (under ~3 s, no libvmaf required).
python ai/train/train_combined.py --epochs 0 \
--netflix-root /tmp/missing --konvid-parquet /tmp/missing.parquet \
--out-dir /tmp/combined_smoke
# Expect: <out-dir>/mlp_small_combined_final.onnx written, rc=0.
0073 — KoNViD-1k → VMAF-pair acquisition + loader bridge¶
- No ADR. Acquisition + loader pieces are pure additions; the methodology fits inside ADR-0203 / Research-0019.
- Upstream source: fork-local. KoNViD-1k integration is a fork-only training-data play.
- Touches (additive only):
ai/scripts/konvid_to_vmaf_pairs.py— acquisition pipeline.ai/train/konvid_pair_dataset.py—KoNViDPairDatasetclass mirroringNetflixFrameDataset's interface.ai/tests/test_konvid_pair_dataset.py— 5 pytest cases.docs/ai/training.md— new "C1 (KoNViD-1k corpus)" section.CHANGELOG.mdUnreleased § Added.- Invariants (rebase-relevant):
KoNViDPairDatasetmirrorsNetflixFrameDatasetshape.feature_dim == 6,numpy_arrays() → (X, y)returns(n_frames, 6)+(n_frames,). IfNetflixFrameDataset's feature order changes, mirror it here.- Acquisition parquet schema is fixed. Required columns:
key,frame_index,vif_scale0..3,adm2,motion2,vmaf. Add freely; do NOT rename / drop those. ai/data/konvid_vmaf_pairs.parquetand$VMAF_TINY_AI_CACHE/konvid-1k/stay gitignored. They regenerate from raw KoNViD.mp4sources.- On upstream sync: zero interaction.
- Re-test on rebase:
pytest ai/tests/test_konvid_pair_dataset.py -v
# Expect: 5 passed
python ai/scripts/konvid_to_vmaf_pairs.py --max-clips 5
# Expect: ~7 s wall, ai/data/konvid_vmaf_pairs.parquet with
# 5 unique keys × ~200 frames each.
0072 — Tiny-AI 3-arch LOSO eval harness + Research-0023¶
- No ADR. Methodology fits inside Research-0023; ADR-0203 already covers the training-prep architecture and the three-arch sweep concept.
- Research digest:
docs/research/0023-loso-3arch-results.md. - Upstream source: fork-local. Netflix/vmaf has no LOSO eval surface.
- Touches (additive only):
ai/scripts/eval_loso_3arch.py— new harness; reuses the_load_session+_load_clip+CLIPShelpers fromeval_loso_mlp_small.py(PR #165).docs/research/0023-loso-3arch-results.md— methodology + per-fold tables formlp_small/mlp_medium/linear.CHANGELOG.mdUnreleased § Added.- Invariants (rebase-relevant):
- Reuse the PR #165 helpers. Don't fork the
_load_sessionexternal-data workaround into a copy — both scripts must keep using the same import. If a follow-up re-exports the shipped baselines with correctedexternal_data.location, both scripts deprecate the workaround simultaneously. runs/andmodel/tiny/training_runs/stay gitignored. The harness writesruns/loso_eval/loso_3arch_eval.{json,md}; the durable record is the table in Research-0023 §2 + the per-fold tables in §3. Regenerate via the loop in §6 of the digest.- On upstream sync: zero interaction. Pure fork-local evaluation harness.
- Re-test on rebase:
python ai/scripts/eval_loso_3arch.py
diff <(jq -r '.archs.mlp_small.aggregate.mean_plcc' runs/loso_eval/loso_3arch_eval.json) <(echo 0.9808)
diff <(jq -r '.archs.mlp_medium.aggregate.mean_plcc' runs/loso_eval/loso_3arch_eval.json) <(echo 0.9727)
diff <(jq -r '.archs.linear.aggregate.mean_plcc' runs/loso_eval/loso_3arch_eval.json) <(echo 0.3679)
# Expect: identical lines on a populated cache + identical fold ONNX.
0071 — T7-16 ADM Vulkan/SYCL drift verified-resolved (doc close)¶
- No ADR. Verification-only close, sister of T7-15.
- Upstream source: fork-local. ADM cross-backend gate is a fork-only test surface; Netflix/vmaf has no Vulkan or SYCL backend.
- Touches (additive only):
docs/state.md— new "Recently closed" row for T7-16..workingdir2/BACKLOG.md— T7-16 row marked closed (local- only planning dossier; gitignored).CHANGELOG.mdUnreleased § Fixed.- Invariants (rebase-relevant):
places=4cross-backend ADM contract. Empiricaladm_scale2max_abs_diff is now 1e-6 (print floor; ULP=0) on Vulkan device 0 (NVIDIA), device 1 (Mesa anv on Arc), and SYCL device 0 (Arc); residualadm_scale1 ≈ 3.1e-5andadm2 ≈ 5e-6on 1/48 frames passplaces=4(5e-5 tolerance) but failplaces=5. Hold the gate atplaces=4.- No ADM kernel source change. Fix is environmental (NVCC + driver + SYCL runtime).
- On upstream sync: zero interaction.
- Re-test on rebase:
python3 scripts/ci/cross_backend_vif_diff.py \
--vmaf-binary core/build/tools/vmaf \
--feature adm --backend vulkan --device 0 --places 4 \
--reference python/test/resource/yuv/src01_hrc00_576x324.yuv \
--distorted python/test/resource/yuv/src01_hrc01_576x324.yuv \
--width 576 --height 324
# Expect: 0/48 mismatches across all 5 ADM metrics.
0070 — T7-15 motion CUDA/SYCL drift verified-resolved (doc close)¶
- No ADR. Verification-only close; no code change in PR #172.
- Upstream source: fork-local. Cross-backend gate is a fork-only test surface; not in Netflix/vmaf.
- Touches (additive only):
docs/state.md— "Recently closed" row for T7-15..workingdir2/BACKLOG.md— T7-15 row marked closed (local- only planning dossier; gitignored).CHANGELOG.mdUnreleased § Fixed.- Invariants (rebase-relevant):
- The
places=4cross-backend gate stays atplaces=4. Empirical max_abs_diff is currently 0.0 (CUDA) or 1e-6 (SYCL/ Vulkan, JSON%frounding floor); tightening toplaces=5could be tempting but the 1e-6 print-floor would then make the SYCL + Vulkan rows fail. Hold atplaces=4until--precision=maxis wired into the diff tool. - No motion-kernel source change. PR #172 didn't modify
core/src/feature/cuda/integer_motion/*.cuorcore/src/feature/sycl/integer_motion_sycl.cpp. The fix is environmental (NVCC + driver), so the next CI run on a fresh image needs to be re-verified against the gate. - On upstream sync: zero interaction.
- Re-test on rebase:
python3 scripts/ci/cross_backend_vif_diff.py \
--vmaf-binary core/build/tools/vmaf \
--reference python/test/resource/yuv/src01_hrc00_576x324.yuv \
--distorted python/test/resource/yuv/src01_hrc01_576x324.yuv \
--width 576 --height 324 --feature motion --backend cuda \
--places 4
# Expect: 0/48 mismatches, max_abs_diff = 0.0
0069 — libvmaf_vulkan.h installed under prefix (build bug)¶
- No ADR. Build-system bug fix; matches existing CUDA / SYCL install conditions.
- Upstream source: fork-local. Vulkan backend is fork-only; Netflix/vmaf has no
libvmaf_vulkan.h. - Touches:
core/include/core/meson.build— adds anis_vulkan_enabledgate that handles thefeatureoption'senabled/autostates; appendslibvmaf_vulkan.htoplatform_specific_headerswhen active.CHANGELOG.mdUnreleased § Fixed.- Invariants (rebase-relevant):
- Install rule mirrors the CUDA / SYCL pattern but uses the feature-option API. The
is_cuda_enabled = get_option('enable_cuda') == trueboolean idiom doesn't apply toenable_vulkanbecause that's a feature option, not a boolean. Use.enabled() or .auto(). Don't "simplify" to== true— that would silently drop the install in theautostate. - Pairs with
ffmpeg-patches/0006-libvmaf-add-libvmaf-vulkan-filter.patchwhich probes for the header viacheck_pkg_config libvmaf_vulkan "libvmaf >= 3.0.0" libvmaf/libvmaf_vulkan.h vmaf_vulkan_state_init_external. Removing the install rule re-introduces lawrence's 2026-04-28 symptom: FFmpeg silently drops thelibvmaf_vulkanfilter despite--enable-libvmaf-vulkan. - On upstream sync: zero interaction; Vulkan backend is fork-only.
- Re-test on rebase:
cd libvmaf
CC=icx CXX=icpx meson setup build -Denable_vulkan=enabled \
-Denable_cuda=true -Denable_sycl=true -Db_lto=false
ninja -C build
meson install -C build --destdir /tmp/libvmaf-install
ls /tmp/libvmaf-install/usr/local/include/libvmaf/libvmaf_vulkan.h
# Expect: file exists.
0066 — --backend cuda inverted-gpumask fix (CLI bug)¶
- No ADR. Bug fix; behaviour now matches the public-header
VmafConfiguration::gpumaskcontract. - Upstream source: fork-local. The
--backendCLI selector was added by the fork (Netflix/vmaf has no exclusive-backend selector). - Touches (additive + 1-line behavioural fix):
core/tools/cli_parse.c::parse_cli_args—--backend cudabranch setsgpumask = 0(wasgpumask = 1).core/test/test_cli_parse.c— 5 new regression tests (test_backend_{cpu,cuda_engages_cuda,cuda_preserves_explicit_gpumask,sycl,vulkan}) plusrun_aom_ctc_tests/run_backend_testshelper split to keeprun_testsunder the function-size budget.CHANGELOG.mdUnreleased § Fixed.- Invariants (rebase-relevant):
VmafConfiguration::gpumasksemantics:if gpumask: disable CUDA.compute_fex_flagsinsrc/libvmaf.croutes CUDA only whengpumask == 0. Any code path that sets a non-zerogpumaskto "request CUDA" silently disables it. The CLI's--backend cudabranch must setgpumask = 0and rely onuse_gpumask = trueto triggervmaf_cuda_state_init. Do not "fix" this back togpumask = 1— it's the bug being fixed.- Explicit
--gpumask=N --backend cudapreserves N. A user who passes--gpumask=2already hasuse_gpumask = true, so the--backend cudabranch's defaulting block (gated on!settings->use_gpumask) is skipped. Thetest_backend_cuda_preserves_explicit_gpumaskregression locks this in. - On upstream sync: zero interaction;
--backendis fork-only. - Re-test on rebase:
./build/test/test_cli_parse | grep -E 'backend_'
# Expect: 5 backend tests pass.
build/tools/vmaf -r REF -d DIS -w 576 -h 324 -p 420 -b 8 \
--model "path=model/vmaf_v0.6.1.json" --threads 1 \
--backend cuda --output cuda.json --json -q
python3 -c "import json; d=json.load(open('cuda.json')); \
assert len(d['frames'][0]['metrics']) == 12, 'CUDA not engaged'"
0067 — Tiny-AI PTQ accuracy across Execution Providers (T5-3e)¶
- No ADR. Investigation/measurement PR; ADR-0129 already governs the PTQ workstream. Findings update
docs/research/0006-tinyai-ptq-accuracy-targets.md§"GPU-EP quantisation" — that section was previously a deferred-open-question; it is now the empirical landing spot. - Research digest: same file (Research-0006).
- Upstream source: fork-local. Netflix/vmaf does not ship a PTQ harness or any tiny-AI ONNX path.
- Touches (additive only):
ai/scripts/measure_quant_drop_per_ep.py— new sibling ofmeasure_quant_drop.py. CPU+CUDA via ORT; Arc / OpenVINO-CPU via the nativeopenvinoPython runtime (noonnxruntime-openvinobecause no cp314 wheel exists). Reuses the_load_sessionrename workaround from PR #165 + avalue_info-strip fix so dynamic-PTQ doesn't choke on the shipped MLP ONNX.docs/ai/quant-eps.md— new user doc; linked fromdocs/ai/index.md.docs/research/0006-tinyai-ptq-accuracy-targets.md— refreshed header, replaced "GPU-EP open question" with the measurement table, fixed pre-existing MD040/MD060 lints surfaced on the touched file.docs/ai/index.md— added the quant-eps row, rewrapped to 80 cols.CHANGELOG.mdUnreleased § Changed.- Invariants (rebase-relevant):
measure_quant_drop.py(the CI gate) is unchanged. The new script is purely additive. Any rebase that conflates the two scripts must keep the CI gate CPU-only — Arc int8 is broken, so a per-EP gate would red-light every PR.value_infostrip is required forvmaf_tiny_v1*dynamic PTQ. The shipped MLP ONNX duplicate weight tensors invalue_info, which makesquantize_dynamicraiseInferred shape and existing shape differ. The fix is in_save_inlined. Don't remove it during a refactor unless the underlying ONNX is regenerated.- CUDA-12 ABI shim. ORT-GPU 1.25 wheels link
libcublasLt.so.12even on CUDA-13 hosts. The reproduction recipe pins thenvidia-*-cu12wheels and prepends them toLD_LIBRARY_PATH. If a future ORT wheel drops the cu12 ABI we can cut the shim, but the script tolerates either since it doesn't import any CUDA symbol itself. - On upstream sync: zero interaction; entirely fork-local.
- Re-test on rebase:
SP=$VIRTUAL_ENV/lib/python3.14/site-packages/nvidia
export LD_LIBRARY_PATH="$SP/cublas/lib:$SP/cudnn/lib:$SP/cuda_nvrtc/lib:$SP/cuda_runtime/lib:$SP/cufft/lib:$SP/curand/lib:$SP/cusolver/lib:$SP/cusparse/lib:$SP/cuda_cupti/lib:$SP/nvtx/lib:$SP/nvjitlink/lib"
python ai/scripts/measure_quant_drop_per_ep.py \
--eps cpu cuda openvino \
--extra-fp32 vmaf_tiny_v1.onnx vmaf_tiny_v1_medium.onnx \
--out runs/quant-eps-$(date +%Y-%m-%d)
# Expected: CPU + CUDA PASS (drop ≤ 1.2e-4); OpenVINO Arc ERR
# (compile failure for Conv-int8) or NaN (MatMul-int8) until a
# newer intel_gpu plugin lands.
0065 — testdata/bench_all.sh correct backend-engagement flags¶
- No ADR. Bug fix; no behavioural surface change beyond "the bench actually engages the backends it claims to now."
- Upstream source: fork-local.
testdata/bench_all.shis a fork-only bench harness; not in Netflix/vmaf. - Touches (additive only):
testdata/bench_all.sh— switched per-row flag pattern from the disable-only singletons (--no_syclfor "CUDA", etc.) to the correct engagement form (--gpumask=0 --no_sycl --no_vulkanfor CUDA,--sycl_device=0 --no_cuda --no_vulkanfor SYCL,--vulkan_device=0 --no_cuda --no_syclfor Vulkan, and--no_cuda --no_sycl --no_vulkanfor CPU). Added a 4th column (Vulkan) to the comparator. Honours$VMAF_BINfor the binary path and$VMAF_ONEAPI_SETVARSfor the oneAPI install location.CHANGELOG.mdUnreleased § Fixed.- Invariants (rebase-relevant):
- Disable-only singletons don't engage a backend.
--no_syclalone leaves CUDA available but unrequested.--no_cudaalone leaves SYCL available but unrequested. The CLI inits CUDA only whenc.use_gpumaskis set; SYCL only whenc.sycl_device >= 0orc.use_gpumask; Vulkan only whenc.vulkan_device >= 0. Any change to those gates that drops one of the per-row flags will re-introduce the silent CPU fallback. Verify after a rebase by recording each live row's JSONframes[0].metricskey count. Treat a GPU count equal to CPU as a fallback warning, never as a fixed expected backend count — seelibvmaf/AGENTS.md§"Backend-engagement foot-guns". gpumasksemantics are inverted from intuition.gpumask=0enables CUDA dispatch;gpumask=1disables it. The per-row CUDA flag is--gpumask=0, not--gpumask=1. Don't "fix" it to--gpumask=1for symmetry with sycl_device/vulkan_device — that's the bug being fixed (parallel to PR #170).- On upstream sync: zero interaction;
testdata/bench_all.shis fork-only. - Re-test on rebase:
VMAF_BENCH_OUTDIR=testdata/bbb/results bash testdata/bench_all.sh
# Record actual live-backend counts and compare within this run:
jq '.frames[0].metrics | keys | length' testdata/bbb/results/t1_cpu.json
jq '.frames[0].metrics | keys | length' testdata/bbb/results/t1_cuda.json
jq '.frames[0].metrics | keys | length' testdata/bbb/results/t1_sycl.json
0063 — Tiny-AI LOSO eval harness for mlp_small¶
- No ADR. The methodology fits inside Research Digest 0022; ADR-0203 already covers the training-prep architecture.
- Research digest:
docs/research/0022-loso-mlp-small-results.md. - Upstream source: fork-local. Netflix/vmaf has no LOSO eval surface.
- Touches (additive only):
ai/scripts/eval_loso_mlp_small.py— new evaluation harness.docs/ai/loso-eval.md— usage doc.docs/research/0022-loso-mlp-small-results.md— methodology + results.CHANGELOG.mdUnreleased § Added.- Invariants (rebase-relevant):
_load_sessionworkaround for renamed-baseline ONNX. The shipped baselinesmodel/tiny/vmaf_tiny_v1*.onnxreference their pre-renameexternal_data.locationvalues. The workaround in_load_sessionrewrites the entries before handing the proto to ORT. Removing the workaround breaks the baseline phase. The proper fix (re-export with matching names) is tracked as a follow-up; until then this code path is load-bearing.runs/andmodel/tiny/training_runs/stay gitignored. The harness writes toruns/loso_eval/by default; do NOT promote any of those outputs into the tree. The 9 fold ONNX and the per-clip JSON cache regenerate from the corpus + trainer + libvmaf CLI.- On upstream sync: zero interaction. Pure fork-local evaluation harness.
- Re-test on rebase:
python ai/scripts/eval_loso_mlp_small.py
diff <(jq -r '.loso_aggregate.mean_plcc' runs/loso_eval/loso_mlp_small_eval.json) <(echo 0.9808)
# Expect: identical line on a populated cache + identical fold ONNX.
0064 — Section-A audit: 9 backlog rows + ADR cross-links¶
- No ADR. Process / docs PR; rows trace back to the individually-cited ADRs / research digests in their own References columns.
- Decision dossier:
.workingdir2/decisions/section-a-decisions-2026-04-28.md. - Source audit:
docs/backlog-audit-2026-04-28.md. - Upstream source: fork-local. Pure backlog hygiene PR; no Netflix code touched.
- Touches (additive only):
.workingdir2/BACKLOG.md— 9 new rows: T3-17, T3-18, T5-3e, T5-4, T7-35, T7-36, T7-37, T7-38; T6-1a row extended with the bisect-cache fixture sub-bullet.docs/research/0006-tinyai-ptq-accuracy-targets.md— drops the "defer until first user" framing on the GPU-EP quantisation open question per user direction; cross-links T5-3e.docs/research/0020-cambi-gpu-strategies.md— v2 follow-up section now cites T7-36 as the gate for opening the v2 row.docs/adr/0205-cambi-gpu-feasibility.md— Decision section's "follow-up integration PR" now cites T7-36.CHANGELOG.mdUnreleased § Changed.- Invariants (rebase-relevant): none. Pure backlog text. Rebase-conflict risk is limited to the same
BACKLOG.mdtable rows that any future row addition would touch; trivial to re-resolve. - On upstream sync: zero interaction.
- Re-test on rebase: none — docs-only.
0062 — ssimulacra2 CUDA + SYCL twins (ADR-0206)¶
- ADR: ADR-0206.
- Upstream source: fork-local. Netflix/vmaf has no SSIMULACRA 2 GPU implementation; this PR adds the CUDA + SYCL twins of the fork's ADR-0201 Vulkan kernel.
- Touches (additive + small wiring edits):
docs/adr/0206-ssimulacra2-cuda-sycl.mdand the index row indocs/adr/README.md.core/src/feature/cuda/ssimulacra2_cuda.{c,h}— new CUDA dispatch.core/src/feature/cuda/ssimulacra2/ssimulacra2_blur.cuandssimulacra2_mul.cu— new CUDA fatbins.core/src/feature/sycl/ssimulacra2_sycl.cpp— new SYCL extractor.core/src/feature/feature_extractor.c— two new extern declarations + two new entries infeature_extractor_list[].core/src/meson.build— addsssimulacra2_blur+ssimulacra2_multocuda_cu_sources, introduces (or extends, if PR #157 / ADR-0202 landed first) thecuda_cu_extra_flagsmap with assimulacra2_blurentry, threadsper_kernel_flagsinto the fatbin custom-target, and lists the two new C / CPP TUs.core/src/cuda/AGENTS.mdandcore/src/sycl/AGENTS.md— rebase invariant notes for the per-kernel--fmad=falseflag and the-fp-model=preciseSYCL build flag.docs/backends/cuda/overview.md,docs/backends/sycl/overview.md,docs/metrics/features.md— coverage matrix updates.CHANGELOG.mdUnreleased § Added.- Invariants (load-bearing on rebase):
- Per-kernel
--fmad=falseforssimulacra2_blur. The IIR'so = n2 * sum - d1 * prev1 - prev2must NOT fuse into FMAs — without the flag the recursive Gaussian's per-step rounding compounds across the 6-scale pyramid pastplaces=4. -fp-model=preciseon the SYCL feature build line. Removing it driftsssimulacra2_syclpastplaces=2through the IIR.- Hybrid host/GPU split mirrors Vulkan. Host runs YUV→RGB, XYB, downsample, and SSIM/EdgeDiff combine in double; GPU runs only mul + IIR blur. Any future PR that ports XYB or YUV→RGB onto the GPU MUST land alongside an updated ADR-0206 and re-validate
places=4on every Netflix CPU pair. - CUDA fex uses
.extract(synchronous), not.submit/.collect. Per-frame raw YUV is D2H-copied frompicture_cuda's device-sideVmafPicture.data[]into pinned host scratch viacuMemcpy2DAsync. Skipping the copy segfaults — direct host reads on aCUdeviceptrare the failure mode the prior agent's WIP hit. - On upstream sync: zero interaction with Netflix. The GPU coverage matrix for
ssimulacra2is wholly fork-local. - Re-test on rebase:
meson setup build_cuda libvmaf -Denable_cuda=true -Denable_sycl=false
ninja -C build_cuda
python3 scripts/ci/cross_backend_vif_diff.py \
--vmaf-binary ./build_cuda/tools/vmaf \
--feature ssimulacra2 --backend cuda --places 4 \
--reference python/test/resource/yuv/src01_hrc00_576x324.yuv \
--distorted python/test/resource/yuv/src01_hrc01_576x324.yuv \
--width 576 --height 324 --pixel-format 420 --bitdepth 8
# Expect: 0/48 mismatches, max_abs_diff ~1e-6.
0061 — cambi GPU feasibility spike (ADR-0205)¶
- ADR: ADR-0205.
- Research digest:
docs/research/0020-cambi-gpu-strategies.md. - Upstream source: fork-local. Netflix/vmaf has no Vulkan backend.
- Touches (additive only):
docs/adr/0205-cambi-gpu-feasibility.md,docs/research/0020-cambi-gpu-strategies.md,docs/adr/README.mdindex row.core/src/feature/vulkan/cambi_vulkan.c— new dormant scaffold (not yet invulkan_sources, not yet registered).core/src/feature/vulkan/shaders/cambi_{derivative,decimate,filter_mode}.comp— new reference GLSL shaders, not yet in the build'sshaderslist.core/src/feature/AGENTS.mdinvariants +CHANGELOG.mdbullet.- Invariants (rebase-relevant):
- Hybrid host/GPU port by decision. If Netflix upstream tightens the c-value formula or histogram update protocol, the host residual call site in the eventual
cambi_vulkan.c::cambi_vulkan_extractmust be updated alongsidecambi.c::calculate_c_values— the same code is reused. Do NOT translate the c-values phase to GPU during any upstream-port PR; that optimisation belongs to the v2 strategy-III PR (deferred). - Scaffolds dormant in the spike PR. The
cambi_vulkan.cextractor returns-ENOSYSfromcambi_vulkan_init_stubuntil the integration follow-up wires it in. Do NOT registervmaf_fex_cambi_vulkan_scaffoldinfeature_extractor.c's list. - Shaders not in the build's shader list. Adding them to
core/src/vulkan/meson.build'svulkan_shaderslist before the integration PR produces orphaned*_spv.hheaders. Leave them alone in this spike PR. - On upstream sync: zero interaction. cambi.c itself is upstream-mirrored — Netflix changes flow through
port-upstream-commit; only the integration PR's host residual call site needs paired attention. - Re-test on rebase:
```bash meson setup build -Denable_vulkan=enabled -Denable_cuda=false -Denable_sycl=false ninja -C build meson test -C build
0059 — Tiny-AI Netflix corpus training prep (ADR-0203)¶
- ADR: ADR-0203.
- Upstream source: fork-local. Netflix/vmaf has no equivalent training surface.
- Touches:
ai/data/— Netflix loader, libvmaf-CLI feature extractor, distillation scoring.ai/train/— PyTorch dataset, eval harness, Lightning-style training entry point.ai/scripts/run_training.sh— convenience wrapper.ai/tests/— five new pytest modules (test_netflix_loader.py,test_dataset.py,test_eval.py,test_train_smoke.py, plusconftest.py).docs/ai/training.md— new "C1 (Netflix corpus)" section; existing sections untouched.ai/AGENTS.md— invariants section added.- Invariants (load-bearing):
- Filename ladder regex is fork-specific.
<source>_<quality>_<height>_<bitrate>.yuv(dis) +<source>_<fps>fps.yuv(ref). Upstream may publish a different naming convention later; do NOT merge them — keep this loader scoped to the Netflix corpus, add a sibling loader for any upstream alternative. - Per-clip cache schema is consumed by both dataset and any downstream tooling. Schema is
{features:{feature_names, per_frame, n_frames}, scores:{per_frame, pooled}}. Any change must invalidate$VMAF_TINY_AI_CACHE(delete or version-tag the directory). - Smoke command stays runnable without a built
vmafbinary. The_make_zero_payloadhelper inai.train.datasetinjects a fake payload for--epochs 0so CI gates don't drag a libvmaf build into the Python test surface. - YUV size probe never silently guesses.
probe_yuv_dimseither matches the 1920x1080 default, returns ffprobe's answer, or raises. Tests passassume_dims=(16, 16)explicitly for synthetic fixtures. - On upstream sync: no interaction with upstream. The
ai/subtree is wholly fork-local. - Re-test on rebase:
python -m pytest ai/tests/test_netflix_loader.py \
ai/tests/test_dataset.py ai/tests/test_eval.py \
ai/tests/test_train_smoke.py -v
python ai/train/train.py --epochs 0 --data-root /tmp/mock_corpus \
--assume-dims 16x16 --val-source BetaSrc --out-dir /tmp/out
0073 — Tiny-AI QAT trainer + first per-model QAT pass (T5-4)¶
- ADR: ADR-0207 (design), ADR-0208 (per-model impl).
- Touches:
ai/train/qat.py(new),ai/scripts/qat_train.py(rewrite fromNotImplementedErrorscaffold),ai/configs/learned_filter_v1_qat.yaml(new),ai/tests/test_qat_smoke.py(new),docs/ai/quantization.md(QAT tier added). All paths are wholly fork-local; no upstream Netflix/vmaf interaction. - Invariants:
- Two-step pipeline (PyTorch QAT → fp32 ONNX → ORT static-quantize) is load-bearing. Both the legacy ONNX exporter (
quantized::conv2d) and the new TorchDynamo exporter (Conv2dPackedParamsBase.__obj_flatten__) refuse to consumeconvert_fxoutput on PyTorch 2.11. The bridge (state-dict diff to a fresh fp32 module + ORT static-quantize) is the only path that yields a QDQ ONNX. Do NOT collapse to a single-stepconvert_fx → torch.onnx.exportuntil both PyTorch issues are fixed; re-check both exporters on each PyTorch upgrade. - State-dict transfer matches by submodule name + shape.
_copy_qat_weights_into_fp32walksfp32_statekeys, finds the same key in the FX-prepared module, copies the tensor. Tiny-AI models today have stable submodule names (entry,body.*,exit); a model architecture that uses top-levelnn.Sequentialwould break this becauseprepare_qat_fxrenames Sequential children to numeric indices. TheRuntimeError("0 tensors copied")guard catches the silent failure mode. - FX preparation runs on CPU. PyTorch 2.11's FX symbolic tracer is flaky on CUDA buffers; the trainer migrates the model to CPU before
prepare_qat_fxand back to the accelerator for the fine-tune phase. The smoke test deliberately exercises the CPU path so this stays covered. torch.ao.quantizationdeprecation will hard-fail in PyTorch 2.10. Migration target istorchao.quantization.pt2e(prepare_pt2e/convert_pt2e); the two-step pipeline is mostly pt2e-compatible — only the FX-prep call changes.- On upstream sync: no interaction with upstream. The
ai/subtree is fully fork-local. - Re-test on rebase:
python -m pytest ai/tests/test_qat_smoke.py -v
python ai/scripts/qat_train.py \
--config ai/configs/learned_filter_v1_qat.yaml \
--output /tmp/qat_smoke.int8.onnx --smoke
0074 — GPU-parity matrix CI gate (T6-8 / ADR-0214)¶
- Touched surfaces (fork-local):
scripts/ci/cross_backend_parity_gate.py(new),.github/workflows/tests-and-quality-gates.yml(newvulkan-parity-matrix-gatejob),docs/development/cross-backend-gate.md(new),docs/backends/index.md(cross-backend section),libvmaf/AGENTS.md(rebase-sensitive invariant note). - Why this matters on rebase: the CI lane and the matrix-gate script are entirely fork-local. Upstream Netflix/vmaf has no comparable gate; conflicts on rebase are restricted to the CI workflow file when upstream rearranges its own jobs. The gate's Python script lives outside
core/src/so the upstream-sync path doesn't see it. - Invariants the gate enforces:
- Per-feature absolute tolerance is declared in one place (
FEATURE_TOLERANCEinscripts/ci/cross_backend_parity_gate.py). Tightening a tolerance requires a measurement-driven follow-up ADR; loosening requires a justification ADR (CLAUDE.md §12 r1). - The legacy single-feature gate
scripts/ci/cross_backend_vif_diff.pystays for one release cycle. Sister PRs in this session add to it; the T6-8b cleanup PR deletes it once the matrix gate has soaked. - CUDA / SYCL / hardware-Vulkan are advisory until a self-hosted runner is registered. The script supports them via
--backends; flipping the CI lane to required is a follow-up wiring change, not a code change. - On upstream sync: no interaction with upstream
tests-and-quality-gates.yml(the gate job is fork-added); rebase conflicts limited to insertion-order in the workflow file. - Re-test on rebase:
cd libvmaf && meson setup build \
-Denable_cuda=false -Denable_sycl=false \
-Denable_vulkan=enabled -Denable_float=true \
--buildtype=release && ninja -C build
cd ..
python3 scripts/ci/cross_backend_parity_gate.py \
--vmaf-binary core/build/tools/vmaf \
--reference testdata/ref_576x324_48f.yuv \
--distorted testdata/dis_576x324_48f.yuv \
--width 576 --height 324 --backends cpu vulkan \
--json-out /tmp/parity.json --md-out /tmp/parity.md
0220 — SYCL feature kernels are unconditionally fp64-free (T7-17)¶
- Touches:
core/src/sycl/common.cpp(init log line),core/src/sycl/AGENTS.md(new invariant row), all SYCL feature kernels undercore/src/feature/sycl/(no diff today, but the contract pins their shape going forward). - Invariant: every SYCL feature-kernel lambda captures and operates on
float/ integer types only. Nodoubleoperand inside aparallel_forbody, nosycl::reduction<double>, nosycl::plus<double>. A single fp64 instruction in the TU's SPIR-V module causes the Level Zero runtime to reject the entire module on Intel Arc A-series and other fp64-less devices, even when the offending kernel is never submitted. Host-sidedouble(inextract/flushpost-processing, score aggregation, log10 normalisation) remains fine. Concrete patterns in tree: ADM gain limiting via int64 Q31 (gain_limit_to_q31+launch_decouple_csf<false>ininteger_adm_sycl.cpp); VIF gain limiting via fp32sycl::fmin; CIEDE / SSIM accumulators viasycl::reduction<int64_t>/sycl::plus<int64_t>. - On upstream sync: Netflix/vmaf has no SYCL backend upstream; conflicts cannot enter via
git merge. The risk is a fork-local cherry-pick (e.g. a SYCL twin of a new CUDA kernel) bringing adoubleinto a kernel lambda. Audit the lambda capture list and anysycl::reduce*calls against this invariant before merging. - Re-test on rebase:
# Build SYCL backend
meson setup build-sycl libvmaf -Denable_sycl=true CC=icx CXX=icpx
ninja -C build-sycl
# On an fp64-less device (e.g. Intel Arc A380), confirm the
# init log line is INFO-level and reads "device lacks native
# fp64 — kernels already use fp32 + int64 paths, no emulation
# overhead". The SYCL kernels must launch successfully (no
# SPIR-V module rejection from the Level Zero runtime).
build-sycl/tools/vmaf --reference testdata/ref_576x324_48f.yuv \
--distorted testdata/dis_576x324_48f.yuv \
--width 576 --height 324 --backend sycl \
--feature integer_vif --feature integer_adm \
--output /tmp/sycl-fp64less.json --json
0091 — T6-9 model registry schema + --tiny-model-verify (ADR-0211)¶
- No rebase impact: 100% fork-local surface. The registry (
model/tiny/registry.json), its JSON Schema (model/tiny/registry.schema.json), the--tiny-model-verifyCLI flag, and thevmaf_dnn_verify_signature()C entry point are entirely fork-local — none of these paths exist in upstream Netflix/vmaf. Listed here for completeness so a future/sync-upstreamrun sees the surface area was acknowledged. - Touches (additive only):
model/tiny/registry.json,model/tiny/registry.schema.json,ai/scripts/validate_model_registry.py,core/src/dnn/model_loader.{c,h}(addedvmaf_dnn_verify_signature()),core/include/libvmaf/dnn.h(public declaration),core/tools/cli_parse.{c,h}(ARG_TINY_MODEL_VERIFY+tiny_model_verifyfield),core/tools/vmaf.c(call site),core/test/dnn/test_tiny_model_verify.c,python/test/model_registry_schema_test.py,docs/ai/model-registry.md,docs/ai/inference.md,docs/ai/security.md,docs/adr/0209-...md,docs/adr/README.md(index row),CHANGELOG.md,core/src/dnn/AGENTS.md. - Invariants (rebase-relevant):
- Schema is the contract. New registry fields land in
registry.schema.jsonfirst, then inregistry.json, then in any consumers (the C-side parser, the Python validator, the MCP). Reverse order causes mismatch. schema_versionis bounded. The schema accepts only{0, 1}; bump the enum and the loader's check together when adding2.- Banned-function rule applies. The
cosigninvocation usesposix_spawnp(3p)with an explicit argv array. Do not replace withsystem(3)/popen(3)— both shell-parse the command and would re-introduce injection risk. - Bundle-file absence is fail-closed. When
sigstore_bundlepoints at a not-yet-existing file (pre-release state),vmaf_dnn_verify_signature()returns-ENOENT. The CLI surfaces this as a load failure; do not "soften" to a warning without an explicit ADR. - Re-test on rebase:
python3 ai/scripts/validate_model_registry.py
python3 -m pytest python/test/model_registry_schema_test.py -v
meson test -C build-cpu --suite=dnn
0074 — HIP (AMD ROCm) backend scaffold (T7-10)¶
- ADR: ADR-0212.
- Upstream source: fork-local. HIP backend is fork-only; Netflix/vmaf has no
libvmaf_hip.hand noenable_hipmeson option. - Touches:
core/include/libvmaf/libvmaf_hip.h(new).core/include/core/meson.build— adds theis_hip_enabledinstall gate, mirroringis_cuda_enabled/is_sycl_enabledboolean idioms.core/meson_options.txt— newenable_hipboolean option (default false).core/src/meson.build— newis_hip_enabledflag, conditionalsubdir('hip'),hip_sources+hip_depsthreaded throughlibvmaf_feature_static_lib(alongside the existing CUDA / SYCL / Vulkan aggregations) and the top-levellibrary('vmaf', ...)dependencieslist.core/src/hip/(new directory:common.{c,h},picture_hip.{c,h},dispatch_strategy.{c,h},meson.build).core/src/feature/hip/(new directory:adm_hip.c,vif_hip.c,motion_hip.c).core/test/test_hip_smoke.c(new).core/test/meson.build— registers the smoke test underif get_option('enable_hip') == true..github/workflows/libvmaf-build-matrix.yml— addsBuild — Ubuntu HIP (T7-10 scaffold)row.docs/backends/hip/overview.md(new),docs/backends/index.md(planned → scaffold row),docs/research/0033-hip-applicability.md(new),docs/adr/0212-hip-backend-scaffold.md(new),docs/adr/README.md(new index row).libvmaf/AGENTS.md— new "HIP backend scaffold contract" rebase-sensitive invariant entry.CHANGELOG.md— Unreleased § Added.- Invariants (rebase-relevant):
enable_hipis abooleanoption, not afeature. Mirrorsenable_cuda/enable_sycl; do not "harmonise" withenable_vulkan'sfeature/disabledform without an ADR amendment per ADR-0212 § "Decision".- Public C-API entry points return
-ENOSYSfor the scaffold. The smoke test core/test/test_hip_smoke.c pins this. A rebase that "succeeds" by accidentally enabling a code path (e.g. a refactor that early-returns 0 fromvmaf_hip_state_init) breaks the smoke and the runtime PR's contract baseline. hip_sourcesis added tolibvmaf_feature_static_lib, NOT directly to the top-levellibrary('vmaf', ...). The static lib is extracted into libvmaf viaobjects: [..., libvmaf_feature_static_lib.extract_all_objects(recursive: true), ...]at the bottom ofcore/src/meson.build. Addinghip_sourcesto the top library() too would double-link.hip_depsIS added to the top library()dependencies:list. The runtime PR will populatehip_depswith the realdependency('hip-lang')linkage; threading it through the top library() ensures consumers see the transitive dependency.- Header purity:
libvmaf_hip.hdoes not include<hip/hip_runtime.h>. HIP runtime types cross the public ABI asuintptr_t(matches the CUDA / Vulkan precedent; ADR-0212). Don't add<hip/...>includes to the public header during a rebase / runtime-PR bring-up. - No FFmpeg patch: the fork's
ffmpeg-patches/series does not currently consume the HIP API surface. CLAUDE §12 r14 only requires patch updates when an existing patch consumes the surface; the runtime PR (T7-10b) will add thehip_devicefilter option and the corresponding patch. - On upstream sync: zero interaction; HIP backend is fork-only.
- Re-test on rebase:
cd libvmaf
meson setup build-hip -Denable_cuda=false -Denable_sycl=false \
-Denable_hip=true
ninja -C build-hip
meson test -C build-hip test_hip_smoke
# Expect: 9/9 pass.
# Default no-HIP build still works:
meson setup build-cpu -Denable_cuda=false -Denable_sycl=false
ninja -C build-cpu
meson test -C build-cpu --suite=fast
0074 — SSIMULACRA 2 SVE2 SIMD parity (T7-38)¶
- ADR: ADR-0213.
- Touches:
core/src/feature/arm64/ssimulacra2_sve2.{c,h}(new),core/src/feature/ssimulacra2.c(dispatch table override ininit_simd_dispatch),core/src/arm/cpu.{c,h}(HWCAP2_SVE2 probe + newVMAF_ARM_CPU_FLAG_SVE2enum value),core/src/meson.build(cc.compiles probe + optionalarm64_ssimulacra2_sve2static library),core/test/test_ssimulacra2_simd.c(SVE2 picker overrides on the arm64 path + dispatch diagnostic),build-aux/aarch64-linux-gnu-sve2.ini(new cross-file pinningqemu-aarch64-static -cpu max). All paths are wholly fork-local; no upstream Netflix/vmaf code is modified. - Invariants:
- Fixed 4-lane SVE2 predicate. Every kernel uses
svwhilelt_b32(0, 4)so SIMD arithmetic order is identical to the NEON sibling regardless of the runtime vector length. This keeps the ADR-0138 / ADR-0139 / ADR-0140 byte-exact contract intact. Do NOT widen the predicate tosvptrue_b32()without a separate ADR + snapshot regen — variable-length lane reductions perturb the per-step rounding order. - NEON stays the fallback. SVE2 is purely additive; the dispatch table assigns NEON first and only overrides on
VMAF_ARM_CPU_FLAG_SVE2. A toolchain that fails thecc.compiles(... -march=armv9-a+sve2)probe leavesHAVE_SVE2unset and the legacy NEON-only build is unchanged. -ffp-contract=offmirrors the NEON sibling. Without it GCC fuses the per-lane scalar tail'sa*b+cpatterns intofmla, drifting against the SIMD path by ~1 ulp. Thearm64_ssimulacra2_sve2static library carries the flag like its NEON counterpart.- On upstream sync: no interaction with upstream —
arm64/feature TUs and thearm/cpu.{c,h}flag enum are fork-local. An upstream sync that rewritesinit_simd_dispatchincore/src/feature/ssimulacra2.cwould also need the SVE2 cases preserved. - Re-test on rebase:
meson setup build-arm64-sve2 libvmaf \
--cross-file=build-aux/aarch64-linux-gnu-sve2.ini -Denable_asm=true
ninja -C build-arm64-sve2 test/test_ssimulacra2_simd
meson test -C build-arm64-sve2 test_ssimulacra2_simd
# stderr should report `ssimulacra2 simd dispatch: NEON=1 SVE2=1`
# and 11/11 tests should pass.
0075 — enable_lcs MS-SSIM extras on CUDA + Vulkan (T7-35 / ADR-0243)¶
- Touched surfaces (fork-local):
core/src/feature/cuda/integer_ms_ssim_cuda.c(addedenable_lcstoMsSsimStateCuda+options[]+ 15 host-sidevmaf_feature_collector_appendcalls gated on the bool),core/src/feature/vulkan/ms_ssim_vulkan.c(rewroteenable_lcshelp text + addedemit_lcs_metricshelper + gated 15vmaf_feature_collector_appendcalls),scripts/ci/cross_backend_vif_diff.pyscripts/ci/cross_backend_parity_gate.py(newfloat_ms_ssim_lcspseudo-feature +FEATURE_ALIASESmapplaces=4tolerance row). - Why this matters on rebase: the GPU MS-SSIM extractors are fork-local (Netflix upstream has no Vulkan or CUDA MS-SSIM kernel today). The
enable_lcssemantic and the metric names (float_ms_ssim_{l,c,s}_scale{0..4}) must match the upstream CPU reference atcore/src/feature/float_ms_ssim.c:189-221. If upstream ever renames or reorders those metrics, mirror the change on the GPU side in the same merge — public-API contract. - Invariants the contract enforces:
- Default-path output (
enable_lcs=false) stays bit-identical to the pre-T7-35 binary: only the host-side appends are gated; no kernel / shader / device-buffer changes. - Metric ordering is metric-wise (all
l_scale*first, thenc_*, thens_*) — matches the CPU emission order. places=4cross-backend tolerance per ADR-0190; enforced by the newfloat_ms_ssim_lcscell in the parity matrix gate (ADR-0214).- On upstream sync: zero interaction; the GPU twins do not exist upstream. The CPU
float_ms_ssim.cis shared with upstream butenable_lcsis upstream-stable since v3.0.0. - Re-test on rebase:
cd libvmaf && meson setup build-vulkan \
-Denable_cuda=false -Denable_sycl=false \
-Denable_vulkan=enabled -Denable_float=true \
--buildtype=release && ninja -C build-vulkan
cd ..
python3 scripts/ci/cross_backend_vif_diff.py \
--vmaf-binary core/build-vulkan/tools/vmaf \
--reference testdata/ref_576x324_48f.yuv \
--distorted testdata/dis_576x324_48f.yuv \
--width 576 --height 324 \
--feature float_ms_ssim_lcs --backend vulkan --places 4
0075 — 32-bit ADM/cpu fallbacks port (T-NEW-3)¶
- Touched surfaces (upstream-mirror):
core/src/feature/x86/adm_avx2.c,core/src/feature/x86/adm_avx512.c,core/src/x86/cpu.c. Cherry-picks of upstream8a289703(Christopher Degawa, "adm: add fallback for extract_epi64 for 32-bit") and1b6c3886("x86/cpu: remove limit of avx+ on 32-bit"). - Why this matters on rebase: trivially conflict-free with any future upstream
extract_epi64work because we land upstream's exactextract_epi64macro/inline-fn pair. The conflict surface is the fork's clang-format-100col layout inadm_avx2.c/adm_avx512.cand the_Alignas(64)LTO-correctness slot inadm_avx512.c(docs/development/known-upstream-bugs.md); both are preserved verbatim. - Invariants the port preserves:
_Alignas(64) int64_t angle_flag[16]inadm_decouple_s123_avx512stays — without it, LTO can promote the unaligned load tovmovdqa64and fault under--buildtype=release -Db_lto=true.- The
extract_epi64symbol must remain resolved on both__x86_64__(macro to_mm256_extract_epi64) and 32-bit (fallback inline). If a future upstream change inlines the helper differently, keep the conditional definition. - On upstream sync: if Netflix ships further 32-bit fallbacks (motion / psnr — not in this port), expect a parallel
extract_epi64-style helper at the top of each affected SIMD file. The fork should mirror those verbatim into the same files. - Re-test on rebase:
meson setup build-i686 libvmaf \
--cross-file=build-aux/i686-linux-gnu.ini \
-Denable_asm=false
ninja -C build-i686
meson setup build-cpu libvmaf -Denable_avx512=true
ninja -C build-cpu
meson test -C build-cpu
0076 — codec-aware FR regressor surface (T7-CODEC-AWARE / ADR-0235)¶
- Touches:
ai/src/vmaf_train/codec.py(new),ai/src/vmaf_train/models/fr_regressor.py(extended),ai/scripts/bvi_dvc_to_full_features.py,ai/scripts/extract_full_features.py. No upstream-shared paths. - Invariant:
CODEC_VOCABinai/src/vmaf_train/codec.pyis closed and order-stable — the index of each codec is the one-hot column index baked into trained ONNX. Adding a codec appends to the tuple and bumpsCODEC_VOCAB_VERSION; reordering silently invalidates every shippedfr_regressor_v2_*.onnx.FRRegressor(num_codecs=0)must remain the v1 single-input contract — flipping the default would break every existingmodel/tiny/fr_regressor_v1.onnxconsumer. - Re-test:
pytest ai/tests/test_codec_aware_fr.py -v(8 sub-tests covering vocabulary contract + alias table + back-compat). Pure fork-local addition; no upstream rebase impact for the next/sync-upstream.
0075 — feature/speed extractors (T-NEW-1, upstream port d3647c73)¶
- Touches:
core/src/feature/speed.c(new),core/src/feature/picture_copy.{c,h}(signature change — addedint channelparameter),core/src/feature/float_*.ccall sites updated to passchannel=0,core/src/feature/feature_extractor.cregistry block,core/src/feature/alias.c,core/src/meson.build,core/src/feature/vif_tools.{c,h}(helper-function port from upstream4ad6e0ea). - Upstream source: verbatim cherry-pick of Netflix/vmaf
d3647c73("feature/speed: port speed_chroma and speed_temporal extractors") with its dependency4ad6e0ea("feature/vif: port helper functions"). Both are pre-existing on Netflix master and enter the fork as part of the T7-4 audit catch-up. - Invariant:
picture_copy()now takes achannelargument — every fork-local extractor that calls it (CUDAinteger_ms_ssim, Vulkanssim/ms_ssim) passeschannel=0. If upstream later evolves the signature again (e.g. adds bit-depth or stride validation), update those fork-local call sites in lockstep. Speed extractors only register whenVMAF_FLOAT_FEATURES=1(build with-Denable_float=true). - On upstream sync: future Netflix commits in
core/src/feature/speed.capply cleanly because the file is now a verbatim mirror; conflict potential is limited to the registry block infeature_extractor.c(interleave with the fork's Vulkan / SYCL / CUDA blocks) and to any furtherpicture_copysignature evolution. - Re-test on rebase:
```bash meson setup build-cpu libvmaf -Denable_cuda=false \ -Denable_sycl=false -Denable_float=true ninja -C build-cpu meson test -C build-cpu test_speed meson test -C build-cpu # full meson suite make test-netflix-golden # 3 CPU canonical pairs
0221 — CHANGELOG + ADR-index fragment-file pattern (T7-39 / ADR-0221)¶
- What changed: the fork stopped editing
CHANGELOG.mdanddocs/adr/README.mddirectly. Both files are now rendered from fragment trees: changelog.d/<section>/<topic>.md(Keep-a-Changelog sections), plus the migration archivechangelog.d/_pre_fragment_legacy.md.docs/adr/_index_fragments/<NNNN-slug>.md, plusdocs/adr/_index_fragments/_order.txt(frozen commit-merge order manifest) anddocs/adr/_index_fragments/_header.md(table prelude). Two scripts render the consolidated outputs:scripts/release/concat-changelog-fragments.sh --check|--writescripts/docs/concat-adr-index.sh --check|--write- On upstream sync: zero interaction —
CHANGELOG.mdis a fork-local Markdown surface (Netflix upstream doesn't ship a Keep-a-Changelog file in this format), anddocs/adr/is entirely fork-local. A/sync-upstreamrun will not touch the fragment trees. - Re-test on rebase:
bash scripts/release/concat-changelog-fragments.sh --check
bash scripts/docs/concat-adr-index.sh --check
# both must exit 0; otherwise run --write and re-stage.
0077 — DISTS extractor proposal (T7-DISTS / ADR-0236)¶
- What landed: ADR-0236 (Proposed) + Research-0043 design digest
- ADR README index row + CHANGELOG entry.
- Rebase impact: pure fork-local proposal-stage docs; no code, no Netflix-mirror file touched, no ffmpeg-patches change, no public C-API surface change.
- Reproducer (when implementation lands as T7-DISTS):
```sh vmaf --feature dists_sq=model_path=model/tiny/dists_sq.onnx \ --reference ref.yuv --distorted dist.yuv \ --width 1920 --height 1080 --pix_fmt yuv420p
0076 — GPU-gen ULP calibration head (proposal-stage, T7-GPU-ULP-CAL / ADR-0234)¶
- What landed: ADR-0234 (Proposed), Research-0041, data-collection scaffold at
ai/scripts/collect_gpu_calibration_data.py, forward-pointer indocs/usage/cli.mdfor the future--gpu-calibratedflag. - Rebase impact: pure fork-local (proposal docs + Python script); no upstream Netflix/vmaf code touched, no public C-API changes, no ffmpeg-patches changes.
- Reproducer:
```sh python3 ai/scripts/collect_gpu_calibration_data.py --smoke
0095 — Per-backend GPU kernel scaffolding templates (CUDA + Vulkan, ADR-0246)¶
- ADR: ADR-0246.
- Touches:
core/src/cuda/kernel_template.h(new, header-only).core/src/vulkan/kernel_template.h(new, header-only).core/src/cuda/AGENTS.md(new invariant row + dir listing).core/src/vulkan/AGENTS.md(new file).docs/backends/kernel-scaffolding.md(new).docs/adr/0246-gpu-kernel-template.md(new).CHANGELOG.md,docs/adr/README.md. All paths are wholly fork-local. Upstream Netflix/vmaf has no Vulkan backend at all today and the CUDA backend uses different per-kernel scaffolding shapes; nothing here can collide on a pure upstream sync.- Invariants:
- Templates are unused at PR-merge time.
kernel_template.hin bothcore/src/cuda/andcore/src/vulkan/lands with zero call-sites. Each future kernel migration is its own gated PR (places=4cross-backend-diff per ADR-0214). Do not bulk-port existing kernels onto the templates in a single sync — that would short-circuit the per-kernel gate. - Per-backend, not cross-backend. Resist the urge to merge the two templates into a unified
gpu/kernel_template.h. CUDA async-stream + event vs Vulkan command-buffer + fence + descriptor-pool share no concrete shape; a unified API would be lowest-common-denominator. - Helper functions, not macros. The header bodies are
static inlinefunctions for cuda-gdb / Nsight / RenderDoc step-debugging. TheCHECK_CUDA_GOTO/CHECK_CUDA_RETURNmacros incuda_helper.cuhstay where they pay off (textualgoto label), and the templates use them internally. - On upstream sync: no interaction with upstream paths. An upstream sync that touches
core/src/cuda/common.horpicture_cuda.hmay shift the helper signatures the template consumes (vmaf_cuda_buffer_alloc,vmaf_cuda_picture_get_stream, …); update the template if so. - Re-test on rebase:
```bash # CUDA build (configure inside libvmaf/ — see CLAUDE.md §2 note). meson setup core/build-cuda libvmaf \ -Denable_cuda=true -Denable_nvcc=true \ -Denable_vulkan=disabled -Denable_sycl=false ninja -C core/build-cuda meson test -C core/build-cuda
# Vulkan build. meson setup core/build-vulkan libvmaf \ -Denable_vulkan=enabled -Denable_cuda=false -Denable_sycl=false ninja -C core/build-vulkan meson test -C core/build-vulkan
0222 — vmaf-perShot per-shot CRF predictor sidecar (T6-3b)¶
- Touches:
core/tools/meson.build(new executable + test wiring),core/tools/vmaf_per_shot.c(new file — fork-local, no upstream sibling),core/tools/test/meson.build(test row),core/tools/test/test_vmaf_per_shot.sh(new smoke test),core/tools/AGENTS.md(sidecar invariants),docs/usage/cli.md(cross-link),docs/usage/vmaf-perShot.md(new user doc),docs/ai/roadmap.md(T6-3b row update). - Invariant: the sidecar must stay standalone — it does not link the libvmaf metric path. Any upstream patch that tries to fold per-shot CRF prediction into
vmaf_score_*would collapse the encoder-hint vs. quality-score separation recorded in roadmap §2.4 and ADR-0222 §Decision. The CSV / JSON column set (shot_id,start_frame,end_frame,frames,mean_complexity,mean_motion,predicted_crf) is the public schema; downstream encoders consume it directly. - Conflict expectation on
/sync-upstream: low. Upstream Netflix has no per-shot CRF predictor in tree, so there is no natural collision point —tools/meson.buildis the only mutually-edited file and the newexecutable('vmaf-perShot', …)block is appended aftervmaf_bench_deps, well clear of upstream's likely additions. - Reproducer:
```bash meson setup build libvmaf -Denable_cuda=false -Denable_sycl=false \ -Denable_vulkan=disabled ninja -C build meson test -C build test_vmaf_per_shot --print-errorlogs ./build/tools/vmaf-perShot \ --reference testdata/ref_576x324_48f.yuv \ --width 576 --height 324 --pixel_format 420 --bitdepth 8 \ --output /tmp/plan.csv cat /tmp/plan.csv
0075 — vmaf-roi sidecar binary (T6-2b / ADR-0247)¶
- Touches:
core/tools/meson.build— adds thevmaf_roiexecutable target (after the existingvmaftarget, beforevmaf_bench). Append-only; no upstream-shared lines moved or removed.core/test/meson.build— adds thetest_vmaf_roiexecutable +test()registration. Append-only.core/tools/vmaf_roi.c— wholly new, fork-local.core/tools/vmaf_roi_core.h— wholly new, fork-local.core/test/test_vmaf_roi.c— wholly new, fork-local.- Invariant: the
vmaf-roisidecar emits two byte-exact formats that downstream encoder drivers (x265--qpfile, SVT-AV1--roi-map-file) will hard-depend on: - x265 ASCII grid — two
#-prefixed header lines (# vmaf-roi qpfile (x265, --qpfile-style)and# frame=N ctu=S cols=C rows=R strength=F.FFF), space-separated signed integers, one row per CTU row,\nterminator. - SVT-AV1 raw binary — exactly
cols * rowsbytes ofint8_t, row-major, no header. - QP-offset clamp —
+-12(VMAF_ROI_CORE_QP_OFFSET_MAX). - Reduction — per-CTU mean (not max). Switching to max or a percentile changes every downstream encoder result and requires its own ADR.
- Pure helpers in
vmaf_roi_core.h— the per-CTU mean reducer and saliency-to-QP mapper arestatic inlinein a header sotest_vmaf_roicompiles them without dragging the libvmaf link surface in. Moving them into a.cTU breaks the test wiring. - On upstream sync: no interaction with upstream —
tools/is a fork-local surface from upstream's perspective (upstream shipsvmaf.conly). An upstream sync that rewritescore/tools/meson.buildshould preserve thevmaf_roiexecutable block. - Re-test on rebase:
```bash meson setup build-cpu libvmaf \ -Denable_cuda=false -Denable_sycl=false -Denable_tools=true ninja -C build-cpu tools/vmaf_roi test/test_vmaf_roi meson test -C build-cpu test_vmaf_roi ./build-cpu/tools/vmaf_roi \ --reference testdata/ref_576x324_48f.yuv \ --width 576 --height 324 --frame 0 --output - \ --encoder x265 --ctu-size 64 --strength 6.0 | head -3 # First two lines are the # comment header; row 1 of the grid # should be "4 2 1 -1 -1 -1 1 2 4" (placeholder radial map).
0219 — motion3 GPU coverage on Vulkan + CUDA + SYCL (T3-15(c) / ADR-0219)¶
- What changed: The
motionGPU twins (core/src/feature/vulkan/motion_vulkan.c,core/src/feature/cuda/integer_motion_cuda.c,core/src/feature/sycl/integer_motion_sycl.cpp) now emitVMAF_integer_feature_motion3_scorein 3-frame window mode (default). Cross-backend gates extended (scripts/ci/cross_backend_*.pyFEATURE_METRICS["motion"]). - Invariants:
motion3 = host-side scalar post-process of motion2. No device-side state changes; motion3 is computed on the host inextract()/collect()/flush()after the existing SAD reduction. The post-processing function (motion3_postprocess_*) mirrors CPUinteger_motion.clines 510-560 byte-for-byte:clip(motion_blend(motion2 * fps_weight, blend_factor, blend_offset), max_val)with optional moving-average against the unaveraged prior blended value.motion_five_frame_window=truereturns-ENOTSUPatinit()on all three GPU backends. The 5-deep blur ring + second SAD-pair dispatch remain deferred. Do NOT silently fall back to the 3-frame path when the user enables the flag — fail loud per CERT C / CLAUDE.md §12 r4.- CPU motion3 algorithm is the source of truth. Any port of an upstream Netflix change to
integer_motion.cthat touchesmotion_blend(...), themotion_max_valclip, or the moving-average rule MUST be mirrored inmotion3_postprocess_*across all three GPU files in the same PR. The cross-backend gate atplaces=4will catch drift, but only after a full GPU run. - On upstream sync: Pure fork-local additions to GPU TUs. Upstream Netflix has no GPU motion extractor. The
motion_blend_tools.hheader is upstream-mirrored — if a sync rewrites themotion_blend()formula, regenerate the GPU snapshot and re-run the cross-backend gate. - Re-test on rebase:
```bash # CPU sanity (motion3 emission unchanged) ./core/build/tools/vmaf \ --reference python/test/resource/yuv/src01_hrc00_576x324.yuv \ --distorted python/test/resource/yuv/src01_hrc01_576x324.yuv \ --width 576 --height 324 --pixel_format 420 --bitdepth 8 \ --feature motion --output /tmp/motion.json --json python -c "import json; d=json.load(open('/tmp/motion.json')); \ print('motion3 frames:', sum(1 for f in d['frames'] \ if 'integer_motion3' in f.get('metrics', {})))" # Expect 49 (one motion3 per frame).
# Cross-backend gate (Vulkan/lavapipe lane works on every host): python scripts/ci/cross_backend_vif_diff.py \ --feature motion --backend vulkan \ --ref python/test/resource/yuv/src01_hrc00_576x324.yuv \ --dis python/test/resource/yuv/src01_hrc01_576x324.yuv \ --width 576 --height 324 --bitdepth 8 \ --vmaf-bin core/build/tools/vmaf # Expect: integer_motion / integer_motion2 / integer_motion3 all OK at places=4.
0216 — vmaf_tiny_v2 (Phase-3-validated tiny VMAF MLP)¶
- Touches:
model/tiny/registry.json,model/tiny/vmaf_tiny_v2.{onnx,json},ai/scripts/{train,export,validate}_vmaf_tiny_v2.py,ai/AGENTS.md,core/test/dnn/{test_vmaf_tiny_v2.py,meson.build},docs/ai/{models/vmaf_tiny_v2.md,inference.md,roadmap.md},docs/adr/{0244-vmaf-tiny-v2.md,README.md},CHANGELOG.md. All paths are wholly fork-local; no upstream Netflix/vmaf code is modified. - Invariants:
- Bundled scaler stats are part of the trust root. The shipped ONNX bakes
(input - mean) / stdas ConstantSub+Divnodes that run before the MLP. Re-exporting must go throughai/scripts/export_vmaf_tiny_v2.py, which pullsmean/stdfrom the trainer checkpoint and writes them as graph initialisers. Adding an out-of-band scaler step at runtime (e.g., a sidecar JSON consumed by the loader) is forbidden without a follow-up ADR — it splits the trust root and invalidates the registry sha256 contract. - Feature column order is fixed. The graph reads
(adm2, vif_scale0, vif_scale1, vif_scale2, vif_scale3, motion2)in exactly this order; reordering breaks the bundledmean/stdconstants. Any change to the feature set requires a fresh Phase-3 chain (Research-0027 → 0028 → 0029 → 0030). - opset 17. Matches the sister tiny-AI models (
learned_filter_v1,nr_metric_v1,fastdvdnet_pre) and the ORT op-allowlist baseline. Upgrading requires re-validating theSub/Div/Gemm/Relu/Squeezeops againstop_allowlist.c. - On upstream sync: zero interaction. Netflix/vmaf has no equivalent surface; an upstream sync that touches
core/src/dnn/(op-allowlist or model-loader changes) needs to preserveSub/Div/Gemm/Relu/Squeezein the allowlist for opset 17. - Re-test on rebase:
```bash bash core/test/dnn/test_registry.sh python3 core/test/dnn/test_vmaf_tiny_v2.py python3 ai/scripts/validate_vmaf_tiny_v2.py \ --onnx model/tiny/vmaf_tiny_v2.onnx \ --parquet runs/full_features_netflix.parquet \ --rows 100 --min-plcc 0.97 meson test -C build-cpu --suite=dnn
0094 — Tiny-AI extractor template (ADR-0250)¶
- Touches:
core/src/dnn/tiny_extractor_template.h(new),core/src/feature/feature_lpips.c,core/src/feature/fastdvdnet_pre.c,core/src/dnn/AGENTS.md,docs/ai/extractor-template.md(new),docs/adr/0250-tiny-ai-extractor-template.md(new). - Invariants:
- Helper signatures are wire-format-stable.
vmaf_tiny_ai_resolve_model_path(name, option, env_var)andvmaf_tiny_ai_open_session(name, path, &out)produce the user-facing log lines<name>: no model path …and<name>: vmaf_dnn_session_open(<path>) failed: <rc>— downstream tooling greps these. Don't rename or reorder the parameters without bumping every extractor + the recipe doc. - YUV→RGB is bit-exact. The shared
vmaf_tiny_ai_yuv8_to_rgb8_planesis a literal move of the pre-existingfeature_lpips.cbody (BT.709 limited-range, nearest-neighbour chroma upsample). LPIPS / saliency / future colour-sensitive tiny-AI scores depend on byte-exact equality with the prior ad-hoc copies. Any change to the conversion constants or the rounding rule needs a separate ADR + a coordinated snapshot regen —model/tiny/weights aren't re-trained against new colour math casually. - Option-table macro is plain text substitution. The
VMAF_TINY_AI_MODEL_PATH_OPTION(state_t, help)macro emits a single struct literal — no control flow, no recursion, no variadic shenanigans (Power-of-10 rule 1 / rule 9). Don't extend it into a multi-option emitter without a fresh ADR. - On upstream sync: zero interaction with upstream —
feature_lpips.candfastdvdnet_pre.care fork-only files, and the newdnn/tiny_extractor_template.hlives entirely under fork-introducedcore/src/dnn/. An upstream sync that rewrites unrelatedfeature_*.cfiles won't conflict. - Re-test on rebase:
cd libvmaf
meson setup build-cpu -Denable_cuda=false -Denable_sycl=false
ninja -C build-cpu
meson test -C build-cpu --suite=dnn
meson test -C build-cpu test_lpips test_fastdvdnet_pre
# All 10 dnn-suite + both extractor tests must pass.
0095 — Vulkan ring-depth tunable (ADR-0251 follow-up #3)¶
- PR: feat/t7-29-followup3-ring-tunable.
- What rebases need to know:
VmafVulkanConfigurationgrew an additiveunsigned max_outstanding_framesfield. Existing zero-initialised configs continue to receive the canonical default (0 → VMAF_VULKAN_RING_DEFAULT == 4). The clamp helpervmaf_vulkan_clamp_ring_sizemoved fromimport.c(file-local static) tovulkan_internal.h(static inline) sostate_initandlazy_alloc_ringshare one definition; an upstream sync that re-introduces the static inimport.cwould shadow the header helper — drop the duplicate, keep the inline. - New public symbol:
vmaf_vulkan_state_max_outstanding_frames(const VmafVulkanState *)— read-side accessor for the clamped value. Pure additive surface; no upstream collision. - On upstream sync: zero interaction. The ring is wholly fork-introduced (ADR-0251); upstream Netflix has no Vulkan backend.
- Re-test on rebase:
```bash meson setup build libvmaf -Denable_cuda=false -Denable_sycl=false \ -Denable_vulkan=enabled ninja -C build meson test -C build test_vulkan_async_pending_fence # All 8 cases must pass: 4 v2-contract + 4 ring-tunable.
0096 — tools/vmaf-tune/ automation umbrella spec (ADR-0237 / Research-0044)¶
- PR: feat/vmaf-tune-spec.
- What rebases need to know: this PR ships only an umbrella ADR
- research digest under
docs/. No tracked source code, notools/vmaf-tune/directory yet, no Meson changes. An upstream sync touching ffmpeg-patches orlibvmaf/cannot collide with this PR. - On upstream sync: zero interaction. Spec-only PR.
- Re-test on rebase:
# No build/test impact — verify the docs render and links are alive:
ls docs/adr/0237-quality-aware-encode-automation.md \
docs/research/0044-quality-aware-encode-automation.md
grep -c '\[ADR-0237\]' docs/adr/README.md
0097 — test_speed gated on enable_float (fix default-build failure)¶
- PR: fix/test-speed-chroma-registration.
- What rebases need to know:
core/test/meson.buildnow wraps thetest_speedexecutable +test()registration inif get_option('enable_float'). Thespeed_chroma/speed_temporalextractors live inspeed.c, which is only compiled whenenable_float=true(the entries infeature_extractor.care wrapped in#if VMAF_FLOAT_FEATURES), so the test'svmaf_get_feature_extractor_by_name("speed_chroma")returned NULL on a default build (enable_float=false). - On upstream sync: zero interaction.
test_speed.cwas added fork-side via the Netflix port commitd3647c73. The gating pattern matchestest_vulkan_*(if get_option('enable_vulkan').enabled()). - Re-test on rebase:
# default (enable_float=false): test_speed must NOT be in the suite
meson setup build libvmaf -Denable_cuda=false -Denable_sycl=false --reconfigure
ninja -C build
meson test -C build # expect: NO test_speed in the run
# CI shape (enable_float=true): test_speed must run + pass
meson setup build libvmaf -Denable_float=true --reconfigure
ninja -C build
meson test -C build test_speed # expect: 5/5 pass
0098 — Vulkan picture preallocation surface (ADR-0238)¶
- PR: feat/vulkan-picture-preallocation.
- What rebases need to know: ABI grows additively. New public surface in
core/include/libvmaf/libvmaf_vulkan.h:enum VmafVulkanPicturePreallocationMethod,VmafVulkanPictureConfiguration,vmaf_vulkan_preallocate_pictures,vmaf_vulkan_picture_fetch. New enumeratorVMAF_PICTURE_BUFFER_TYPE_VULKAN_DEVICEincore/src/picture.h::VmafPictureBufferType. New TUcore/src/vulkan/picture_vulkan_pool.c(~180 LOC); registered incore/src/vulkan/meson.build. Fork-internal accessorvmaf_vulkan_state_context()(declared invulkan_internal.h) exposes the imported state's VkInstance/VkDevice to the pool — used only bylibvmaf.c::vmaf_vulkan_preallocate_pictures. VmafContextfield added:vmaf->vulkan.poolnext tovmaf->vulkan.state. Thevmaf_close()teardown closes the pool before clearing the state pointer (matches SYCL).- On upstream sync: zero interaction. Vulkan backend is fork-only; upstream Netflix has no Vulkan integration.
- Re-test on rebase:
```bash meson setup build libvmaf -Denable_cuda=false -Denable_sycl=false \ -Denable_vulkan=enabled ninja -C build meson test -C build test_vulkan_pic_preallocation # All 6 cases must pass under ASan/UBSan: # test_method_none_is_a_no_op # test_method_host_allocates_round_robins # test_method_device_allocates_round_robins # test_fetch_without_preallocate_falls_back # test_unknown_method_rejected # test_null_args_rejected
0099 — feature_mobilesal.c + transnet_v2.c migrated to tiny_extractor_template.h¶
- PR: refactor/migrate-ai-to-template.
- What rebases need to know:
feature_mobilesal.candtransnet_v2.cpreviously open-coded the model-path resolution (getenv+ log block), the YUV→RGB kernel (mobilesal only), thevmaf_dnn_session_open+ log boilerplate, and theVmafOption[].model_pathrow. They now use the helpers fromdnn/tiny_extractor_template.h(PR #251) — the same templatefeature_lpips.candfastdvdnet_pre.calready consume. Net −98 LOC of identical boilerplate. - Behavior preserved: bit-exact YUV→RGB conversion (mobilesal used the literal copy of
feature_lpips.c's body that the template hoisted), identical error-log strings, identical option-table flag/type/offset shape. The migratedmobilesal_optionsmacro expands to the same struct literal the hand-rolled version produced. - On upstream sync: zero interaction. Both files are fork-introduced; upstream Netflix has neither extractor.
0100 — cuda/ring_buffer.{c,h} → gpu_picture_pool.{c,h} (ADR-0239)¶
- PR: refactor/gpu-picture-pool-extract.
- What rebases need to know:
core/src/cuda/ring_buffer.candring_buffer.hare removed. The same callback-based round-robin pool lives atcore/src/gpu_picture_pool.{c,h}under renamed symbols (VmafRingBuffer→VmafGpuPicturePool,vmaf_ring_buffer_*→vmaf_gpu_picture_pool_*,_fetch_next_picture→_fetch). All call sites inlibvmaf.cmigrated.core/test/test_ring_buffer.crenamed totest_gpu_picture_pool.cwith the corresponding meson update. - Netflix-upstream interaction: minimal — Netflix's
cuda/ring_buffer.{c,h}last touched in commitcb1d49c6. An upstream sync that resurrects the old names should be redirected to the new ones; the file move is purely fork-local. Netflix#1300mutex-destroy-order fix preserved (ADR-0157) — moved verbatim to the new file; the fix remains attached tovmaf_gpu_picture_pool_close.- SYCL pool migration:
vmaf_sycl_picture_pool_*keeps its public-internal API but now delegates to the generic pool. The SYCL wrapper struct (VmafSyclPicturePool) just owns theVmafSyclCookiestorage.std::mutexdrops out. - Vulkan pool migration: bundled into this PR after #264 merged.
picture_vulkan_pool.crewrites as a thin wrapper around the generic pool — wrapper struct owns per-pool state for the alloc/free callbacks; the generic pool owns the round-robin slots / mutex / unwind. Same pattern as the SYCL migration above. - Re-test on rebase:
meson setup build libvmaf -Denable_cuda=false -Denable_sycl=false
ninja -C build
meson test -C build --suite=dnn
meson test -C build test_lpips test_mobilesal test_transnet_v2 test_fastdvdnet_pre
# All 11 dnn-suite + 4 extractor smoke tests must pass.
meson test -C build # 47/47 pass under ASan/UBSan
# CUDA build (CI-only; pre-existing local nvcc include-path quirk):
meson setup build-cuda libvmaf -Denable_cuda=true
ninja -C build-cuda
meson test -C build-cuda test_gpu_picture_pool
# SYCL build:
meson setup build-sycl libvmaf -Denable_sycl=true
ninja -C build-sycl
meson test -C build-sycl
0104 — psnr_vulkan.c migrated to vulkan/kernel_template.h¶
- PR: refactor/migrate-psnr-vulkan-to-template.
- What rebases need to know:
vulkan/kernel_template.h(410 LOC, ADR-0246, PR #251) shipped with zero consumers. Its docstring designatedpsnr_vulkan.cas the reference implementation. This PR lands the migration as the first consumer of the Vulkan template — paired with PR #269 (the first CUDA template consumer). The 5 long-lived pipeline objects (descriptor-set layout, pipeline layout, shader module, compute pipeline, descriptor pool) collapse from individual struct fields to oneVmafVulkanKernelPipeline plbundle.create_pipeline()(~104 LOC) collapses to a singlevmaf_vulkan_kernel_pipeline_create()call (~30 LOC) — the template owns the descriptor-set layout creation, pipeline layout, shader module, compute pipeline, and descriptor-pool sizing.close_fex()'svkDeviceWaitIdle+ 5×vkDestroy*sweep collapses to onevmaf_vulkan_kernel_pipeline_destroy()call. - Net LOC delta: −55 LOC on
psnr_vulkan.cdirectly. Unlike the CUDA template (where helper-call boilerplate roughly matches the inline savings), the Vulkan template's pipeline creation is dramatic enough that even the first consumer wins. - Bit-exactness gates: spec-constants, push-constant struct, shader bytecode, dispatch grid math, and host-side reduction are byte-identical to the prior implementation. The template only owns descriptor-set layout / pipeline layout / shader module / compute pipeline creation / descriptor pool sizing — none of which affects the kernel's mathematical behaviour. Cross-backend parity gate (places=4) re-runs unchanged.
- On upstream sync: zero interaction.
psnr_vulkan.cis fork-introduced (T7-23 / ADR-0182 / ADR-0216). - Re-test on rebase:
meson setup build libvmaf -Denable_cuda=false -Denable_sycl=false \
-Denable_vulkan=enabled
ninja -C build
meson test -C build # 50/50 pass on lavapipe
# Cross-backend parity gate (places=4):
python scripts/ci/cross_backend_parity_gate.py --feature psnr_y --places 4
0105 — moment_vulkan.c + ciede_vulkan.c migrated to vulkan/kernel_template.h¶
- PR: refactor/migrate-motion-vulkan-to-template (note: the branch name reflects the original intent; motion's two-pipeline shape didn't fit the template's single-pipeline contract, so this PR migrates moment + ciede instead).
- What rebases need to know: second + third consumers of
vulkan/kernel_template.h(after PR #270 = psnr_vulkan, the first consumer). Both files follow the identical migration pattern: - Replace 5 individual pipeline-object fields (
dsl,pipeline_layout,shader,pipeline,desc_pool) with oneVmafVulkanKernelPipeline plbundle. - Replace ~100 LOC of
create_pipeline()body (descriptor-set layout + pipeline layout + shader module + compute pipeline + descriptor pool boilerplate) with a singlevmaf_vulkan_kernel_pipeline_create()call. - Replace
close_fex()'svkDeviceWaitIdle+ 5×vkDestroy*sweep with onevmaf_vulkan_kernel_pipeline_destroy()call. - Per-file LOC deltas:
moment_vulkan.c: −60 LOC (450 → 390).ciede_vulkan.c: −59 LOC (536 → 477).- Net: −119 LOC.
- Bit-exactness preserved: spec-constants (width/height/bpc/ subgroup_size identical across both), push-constant structs (
MomentPushConsts,CiedePushConsts), shader bytecodes (moment_spv,ciede_spv), dispatch grid math, and host-side reductions are byte-identical to the prior implementation. Cross-backend parity gates (places=4 for moment integer reduce; places=2 for ciede transcendentals per ADR-0187) re-run unchanged. motion_vulkan.cdeferred: motion uses two pipelines (first frame vs subsequent) sharing one DSL + layout + shader + pool. The template's current shape produces one pipeline per descriptor; splitting motion across twoVmafVulkanKernelPipelineinstances would duplicate the shared objects. Tracked as a follow-up template extension (multi-pipeline support).- On upstream sync: zero interaction. Both files are fork-introduced (T7-23 / ADR-0182 / ADR-0187).
- Re-test on rebase:
```bash meson setup build libvmaf -Denable_cuda=false -Denable_sycl=false \ -Denable_vulkan=enabled ninja -C build meson test -C build # 50/50 pass on lavapipe (under ASan/UBSan) python scripts/ci/cross_backend_parity_gate.py --feature float_moment_ref1st --places 4 python scripts/ci/cross_backend_parity_gate.py --feature ciede2000 --places 2
0101 — GPU backend pattern doc (ADR-0240)¶
- PR: docs/gpu-backend-template.
- What rebases need to know: doc-only PR. Adds
docs/development/gpu-backend-template.md(recipe new GPU backends follow) andcore/include/libvmaf/AGENTS.md(public-headers-tree invariant note). No source code, no meson changes, no ABI impact. - On upstream sync: zero interaction. Both files are fork-introduced.
- Re-test on rebase:
```bash # Doc-only — verify links resolve: test -f docs/development/gpu-backend-template.md test -f core/include/libvmaf/AGENTS.md grep -c 'gpu-backend-template' core/include/libvmaf/AGENTS.md
0102 — Tiny-AI test registration macro (tiny_ai_test_template.h)¶
- PR: refactor/test-registration-macro.
- What rebases need to know: new
core/test/tiny_ai_test_template.hemits the four standard registration tests (<name>_is_registered,<name>_provides_primary_feature,<name>_options_table_well_formed,<name>_init_rejects_missing_model) via theVMAF_TINY_AI_DEFINE_REGISTRATION_TESTS(ext, feat, env, prefix)macro. The four per-extractor test files (test_lpips.c,test_mobilesal.c,test_transnet_v2.c,test_fastdvdnet_pre.c) shrank from ~140 LOC each to ~20-50 LOC. Net −286 LOC. Behavior bit-exact preserved (same assertions, same env-var save/restore dance, same setenv shim for MSVCRT). TransNet V2 keeps two extractor-specific extra tests (binary-flag round-trip + provided_features list-termination) that the macro doesn't cover. - On upstream sync: zero interaction. The four test files are fork-introduced (per ADR-0042 / ADR-0168 / ADR-0220 / ADR-0223 / ADR-0215).
- Re-test on rebase:
```bash meson setup build libvmaf -Denable_cuda=false -Denable_sycl=false ninja -C build meson test -C build test_lpips test_mobilesal test_transnet_v2 test_fastdvdnet_pre # 4/4 binaries pass; 18 individual tests total (4x4 standard + 2 # TransNet V2 extras).
0103 — integer_psnr_cuda.c migrated to cuda/kernel_template.h¶
- PR: refactor/migrate-psnr-cuda-to-template.
- What rebases need to know:
cuda/kernel_template.hshipped with no consumers in PR #251 (ADR-0246). This PR migrates the first consumer (integer_psnr_cuda.c) — the file the template's own docstring explicitly designated as the reference. TheCUstream + CUevent + CUeventtriple and the(VmafCudaBuffer device, void *host_pinned, size_t bytes)readback pair are now dispensed by the template helpers (vmaf_cuda_kernel_lifecycle_init/_close,vmaf_cuda_kernel_readback_alloc/_free,vmaf_cuda_kernel_submit_pre_launch,vmaf_cuda_kernel_collect_wait) instead of being open-coded.PsnrStateCudashrinks: replaces three fields (event+finished+str) with oneVmafCudaKernelLifecycle - replaces (
sse+sse_host) with oneVmafCudaKernelReadback. - Net LOC delta: +8 LOC on
integer_psnr_cuda.calone — the helpers add per-call boilerplate. The dedup win materialises as more CUDA feature kernels (motion / moment / ssim / vif / adm) migrate one-at-a-time in follow-up PRs. Each subsequent migration saves ~15 LOC. - Bit-exactness gates: kernel launch + reduction logic unchanged. The migration only touches state-management boilerplate around the kernel; the SSE accumulator math, the per-bpc kernel function lookup, the host-side
log10score formula, and the dispatch grid-dim calculation are byte-identical to the prior implementation. Netflix golden gate + CPU/CUDA cross-backend parity gate (places=4) re-run unchanged. - On upstream sync: zero interaction.
integer_psnr_cuda.cis fork-introduced (T7-23 / ADR-0182). - Re-test on rebase:
meson setup build libvmaf -Denable_cuda=true
ninja -C build
meson test -C build # CUDA test suite must pass
# Cross-backend parity gate:
python scripts/ci/cross_backend_parity_gate.py --feature psnr_y --places 4
0125 — Vulkan submit-side template + fence pool + descriptor pre-alloc bundle (ADR-0256)¶
- Touches:
core/src/vulkan/kernel_template.h— fork-local. Output landing inruns/phase_a/is gitignored — rerun the script to reproduce.VmafVulkanKernelSubmitPoolstruct +_create/_destroy/_acquirehelpers +vmaf_vulkan_kernel_descriptor_sets_allochelper. Upstream has no Vulkan backend — no merge surface.core/src/feature/vulkan/{psnr_hvs,vif,float_vif,float_adm}_vulkan.c— fork-local kernel TUs, also no upstream peer.- Invariant: the four migrated kernels keep all per-frame
VkFence+VkCommandBuffer+VkDescriptorSetresources alive across frames in the pool. Pre-bound descriptor sets rely on the kernel'sVmafVulkanBuffer *handles being init-time stable (allocated ininit(), freed only inclose_fex).vmaf_vulkan_kernel_pipeline_destroydestroys the descriptor pool — pre-allocated sets are released implicitly via the pool; callers must NOT callvkFreeDescriptorSetson them. - Re-test on rebase:
meson setup build libvmaf -Denable_vulkan=enabled
ninja -C build
VK_ICD_FILENAMES=/usr/share/vulkan/icd.d/nvidia_icd.json \
meson test -C build test_vulkan_smoke \
test_vulkan_async_pending_fence \
test_vulkan_pic_preallocation
python scripts/ci/cross_backend_vif_diff.py \
--vmaf-binary build/tools/vmaf \
--reference testdata/ref_576x324_48f.yuv \
--distorted testdata/dis_576x324_48f.yuv \
--width 576 --height 324 --pixel-format 420 --bitdepth 8 \
--feature vif --backend vulkan --places 4
python scripts/ci/cross_backend_vif_diff.py \
--vmaf-binary build/tools/vmaf \
--reference testdata/ref_576x324_48f.yuv \
--distorted testdata/dis_576x324_48f.yuv \
--width 576 --height 324 --pixel-format 420 --bitdepth 8 \
--feature adm --backend vulkan --places 4
0107 — psnr_hvs_cuda async upload + persistent pinned staging (T-GPU-OPT-2/3)¶
- Touches:
core/src/feature/cuda/integer_psnr_hvs_cuda.c— only consumer; fork-local from inception (T7-23 / ADR-0188 / ADR-0191). State addsupload_str(dedicated H2D stream),upload_done(cross-stream completion event), and per-plane persistent pinnedh_uint_ref[3]/h_uint_dist[3]staging buffers allocated once ininit_fex_cuda. The per-call helperupload_plane_cudais split intoissue_d2h_plane(pic-stream D2H),convert_plane(CPU normalise), andissue_h2d_plane(upload-stream H2D).submit_fex_cudaruns the three phases explicitly and recordsupload_doneafter the last H2D, thencuStreamWaitEvents onlc.strbefore kernel launches.core/src/cuda/AGENTS.md— adds a rebase-sensitive invariant entry under §Rebase-sensitive invariants documenting the three-phase flow + persistent staging contract.- Invariant: the pinned
h_uint_*andh_ref/h_distbuffers are never freed and re-allocated mid-stream; the H2Ds must run onupload_str(not onlc.str) so thecuStreamWaitEventcross-stream link is meaningful; theupload_doneevent is recorded after the last H2D for the current frame and waited on once before the first kernel launch of that frame. CUDA graph capture (future T-GPU-OPT-N) depends on the no-per-frame-alloc invariant; collapsing the three-phase split or re-introducing per-framevmaf_cuda_buffer_host_alloccalls breaks that follow-up. Bit-exactness gate isplaces=3forpsnr_hvs_y / cb / crand the combinedpsnr_hvs(matches the existing matrix; notplaces=4). - On upstream sync: zero interaction.
integer_psnr_hvs_cuda.cis fork-introduced (T7-23 / ADR-0188 / ADR-0191). - Re-test on rebase:
meson setup build libvmaf -Denable_cuda=true -Denable_sycl=false
ninja -C build
meson test -C build
python3 scripts/ci/cross_backend_vif_diff.py \
--vmaf-binary core/build/tools/vmaf \
--reference python/test/resource/yuv/src01_hrc00_576x324.yuv \
--distorted python/test/resource/yuv/src01_hrc01_576x324.yuv \
--width 576 --height 324 --pixel-format 420 --bitdepth 8 \
--feature psnr_hvs --backend cuda --places 3
0227 — output.c writer-format unit tests (R3 of coverage-gap-2026-05-02)¶
- Touches:
core/test/test_output.c(new) — exercises the four writers incore/src/output.c(XML / JSON / CSV / SUB) end-to-end viatmpfile()-backed sinks and a syntheticVmafFeatureCollector. Pure test-only; no production code change.core/test/meson.build— registerstest_outputnext totest_feature_collector(mirrors that test's wiring:link_with: libvmaf+ libsvm objects + log/predict/metadata helpers).- Invariant: the test pulls
libvmaf.candoutput.cin via#include "*.c"(mirroring the precedent intest_feature_collector.c) so the per-translation-unit.gcnolands in the test build dir and gcovr aggregates output.c's coverage. The mu-test framework macro (mu_assert) deliberately early-returns from eachstatic char *test_*()body — that's why every test body tripsclang-analyzer-unix.Malloc"potential leak" notes (cleanup runs only on the success-tail path). This pattern is shared across everycore/test/test_*.cfile and is load- bearing (per ADR-0141 NOLINT carve-out): replacing it with goto- cleanup would obscure the per-assertion failure message. - On upstream sync: zero interaction.
output.cis upstream- mirrored, but this PR doesn't touch it. The test only depends on the four public function signatures (vmaf_write_output_{xml, json,csv,sub}); if Netflix renames or reorders those, the test fails to compile and the rebase author updates it then. - Re-test on rebase:
cd libvmaf && meson setup build -Denable_cuda=false -Denable_sycl=false
ninja -C build && ./build/test/test_output
0126 — OSSF Scorecard policy (ADR-0263)¶
- Touches:
.github/workflows/scorecard.yml(line 45 — thegithub/codeql-action/upload-sarif@<sha>pin). The rest of the policy is doc-only (docs/adr/0263-*.md,docs/research/0053-*.md,changelog.d/security/). Upstream Netflix/vmaf does not ship a Scorecard workflow, so the path itself is fork-introduced and won't conflict. - Invariant: the
upload-sarifSHA must point to a commit that currently exists ingithub/codeql-action's git tree. A SHA that was oncev4head but no longer exists in the action repository triggers Scorecard's "imposter commit" defence and breaks the workflow with a 400 error againstapi.scorecard.dev. Verify on every Dependabot bump by spot-checkinggh api /repos/github/codeql-action/commits/<sha>returns 200. - On upstream sync: zero interaction.
- Re-test on rebase:
```bash # Confirm the pin still resolves to a real commit: pin=$(grep -oE 'codeql-action/upload-sarif@[a-f0-9]{40}' \ .github/workflows/scorecard.yml | head -1 | cut -d@ -f2) gh api "/repos/github/codeql-action/commits/$pin" --jq '.sha' # Then watch the next master push for a green Scorecard run: gh run list --workflow scorecard --repo VMAFx/vmafx --limit 1
0228 — U-2-Net u2netp saliency replacement deferred (ADR-0265)¶
- Touches: docs-only.
docs/adr/0265-u2netp-saliency-replacement-blocked.md— new ADR continuing the deferral chain started by ADR-0257.docs/research/0055-u2netp-saliency-replacement-survey.md— new research digest (upstream survey + license + distribution -allowlist audit + alternatives walk).docs/ai/models/mobilesal.md— pointer block updated to reference both ADR-0257 (first blocker) and ADR-0265 (second blocker).model/tiny/registry.json—mobilesal_placeholder_v0notesfield updated to reference ADR-0265 alongside ADR-0257 (no schema / sha256 / file changes).model/tiny/mobilesal.json— sidecarnotesfield updated in lockstep.scripts/gen_mobilesal_placeholder_onnx.py— generator notes string updated so re-running is idempotent against the new sidecar / registry text.CHANGELOG.md— Changed entry viachangelog.d/changed/T6-2a-followup-u2netp-replacement-deferred.md.docs/adr/README.md— index row viadocs/adr/_index_fragments/0265-u2netp-saliency-replacement-blocked.md.- Invariant: zero C-side surface change.
feature_mobilesal.ctensor-name contract (inputinput→ outputsaliency_map, NCHW float32[1, 3, H, W]→[1, 1, H, W]) is unchanged; the on-diskmodel/tiny/mobilesal.onnx(sha256f1226310…) is unchanged;mobilesal_placeholder_v0'ssmoke: trueflag is unchanged. Any future drop-in (U-2-Net viaT6-2a-mirror-u2netp-via-release+T6-2a-widen-allowlist-resize, distilled student, or BASNet / PoolNet survey result) replaces the.onnxand bumps the registry sha256 without touching the C side. - On upstream sync: zero interaction.
feature_mobilesal.c, the registry, the ADR, and the research digest are all fork-local (T6-2a; ADR-0218 / ADR-0257 / ADR-0265; not present in Netflix upstream). - Re-test on rebase:
meson setup build -Denable_cuda=false -Denable_sycl=false
ninja -C build
meson test -C build test_mobilesal
python3 ai/scripts/validate_model_registry.py
bash scripts/docs/concat-adr-index.sh --check
bash scripts/release/concat-changelog-fragments.sh --check
0108 — ssim_accumulate_avx512 per-lane double reduction vectorised¶
- ADR: ADR-0139 (existing; no new ADR — the per-lane reduction order is unchanged).
- Touches:
core/src/feature/x86/ssim_avx512.c— thessim_accumulate_block_avx512body. The per-lane scalarssim_accumulate_lanecalls (16 of them) are replaced by two 8-wide__m512dpasses that computelv,cv,sv, andlv*cv*svlane-wise in vector double. Aligneddouble[16]spill buffers replace the previous_Alignas(64) float[16]×6spill, and the scalar accumulation loop now does 4×16vaddsdinstead of 16 invocations of the per-lane helper.CHANGELOG.md— Changed entry.- This file — this entry.
- Invariant (load-bearing for ADR-0139 bit-exactness):
- Per-lane double computation order is byte-identical:
((2.0 * rm) * cm + C1) / l_den, then(2.0 * srsc + C2) / c_den, then(lv * cv) * sv. No FMA contraction (separate_mm512_mul_pd+_mm512_add_pd—_mm512_fmadd_pdis forbidden because it changes the rounding count and would diverge from scalar's two-stepmul+add). - Float→double widening uses
_mm512_cvtps_pdwhich is IEEE-754-exact for finite floats (52-bit mantissa fits 23-bit float losslessly). - Lane-by-lane left-to-right reduction order preserved:
local_ssim += t_ssim[k]fork = 0..15. Tree reductions (pairwise add, dual-accumulator unroll) are forbidden — they break running-sum associativity against scalar. - AVX2 / NEON twins kept on the per-lane scalar path. Verified bit-identical against the new AVX-512 at
--precision maxon the Netflixsrc01_hrc00/01_576x324and thecheckerboard_1920_1080_10_3_*_0pairs. The bit-exactness contract (ADR-0139) is per-lane, not per-ISA algorithm — so AVX2 / NEON stay scalar-per-lane until a dedicated PR vectorises them with the same care. - Rebase impact: zero conflict with Netflix upstream — the whole SSIM SIMD surface is fork-local (no upstream SSIM SIMD exists). Conflicts only arise if upstream changes
ssim_accumulate_default_scalariniqa/ssim_tools.c; in that case both the AVX2 / NEON per-lane helper and the AVX-512 vector-double block need a coordinated update preserving the three invariants above. - Re-test on rebase:
meson setup build libvmaf -Denable_cuda=false -Denable_sycl=false
ninja -C build
meson test -C build
# Bit-exact at --precision max, scalar vs AVX2 vs AVX-512:
for MASK in 0 16 255; do
core/build/tools/vmaf -r python/test/resource/yuv/src01_hrc00_576x324.yuv \
-d python/test/resource/yuv/src01_hrc01_576x324.yuv \
-w 576 -h 324 -p 420 -b 8 \
--feature float_ms_ssim --feature float_ssim \
--xml -o /tmp/m${MASK}.xml --precision max --cpumask $MASK
done
diff <(grep -v 'fyi fps' /tmp/m0.xml) <(grep -v 'fyi fps' /tmp/m16.xml) # empty
diff <(grep -v 'fyi fps' /tmp/m0.xml) <(grep -v 'fyi fps' /tmp/m255.xml) # empty
- Why this matters on rebase: an upstream commit that touches
core/src/feature/ssimulacra2.ccould prompt a "let's also port the GPU XYB while we're here" follow-up. The ledger entry is the standing answer: don't, the measurement was redone on NVIDIA in May 2026 and the result still failedplaces=4by five decades. See Research-0047.
0126 — FastDVDnet real upstream weights drop (ADR-0253)¶
- What changed: replaces
model/tiny/fastdvdnet_pre.onnxwith the wrapped real upstream FastDVDnet checkpoint (sha256eb9444cf6f07eefdc7f4f68d09131074dbd1dcee6f88a331ba684dd2fb5937d4, ~9.5 MiB), refreshes the sidecarmodel/tiny/fastdvdnet_pre.json, flips the registry row'ssmoke: true → falseand addslicense: "MIT"+ the upstream commit pinc8fdf61. New exporterai/scripts/export_fastdvdnet_pre.py(the older_placeholder.pyexporter is retained for reference). New ADRdocs/adr/0255-fastdvdnet-pre-real-weights.md; user-facing docdocs/ai/models/fastdvdnet_pre.mdrewritten with provenance, license attribution, and reproduce-the-export instructions. - Upstream source: fork-local. Netflix/vmaf does not ship a FastDVDnet temporal pre-filter; the C extractor and ONNX surface are entirely fork-introduced (ADR-0215). The wrapped weights are attribution-only (upstream
m-tassano/fastdvdnetMIT). - On upstream sync: zero interaction. Every file touched (
ai/scripts/export_fastdvdnet_pre*.py,model/tiny/fastdvdnet_pre.*,docs/ai/models/fastdvdnet_pre.md,docs/adr/0253-*.md, CHANGELOG fragment, ADR index fragment) lives in fork-introduced trees. - Re-test on rebase:
# Re-derive the ONNX from the pinned upstream checkpoint.
mkdir -p /tmp/fastdvdnet_upstream && cd /tmp/fastdvdnet_upstream
curl -L -O https://raw.githubusercontent.com/m-tassano/fastdvdnet/c8fdf61/model.pth
curl -L -O https://raw.githubusercontent.com/m-tassano/fastdvdnet/c8fdf61/models.py
cd /path/to/vmaf
python3 ai/scripts/export_fastdvdnet_pre.py \
--upstream-dir /tmp/fastdvdnet_upstream
python3 ai/scripts/validate_model_registry.py
meson test -C build --suite=fast --print-errorlogs test_fastdvdnet_pre
0127 — ONNX op-allowlist gains Resize (ADR-0258)¶
- Touches:
core/src/dnn/op_allowlist.c— fork-local file (no upstream counterpart). One new entry"Resize"under the/* convolutional */block.core/test/dnn/test_op_allowlist.c,core/test/dnn/test_onnx_scan.c— fork-local DNN tests.ai/tests/test_op_allowlist.py— fork-local Python parity test.- Invariant: the C allowlist is the single source of truth; the Python regex parser in
ai/src/vmaf_train/op_allowlist.pywalks the sameop_allowlist.cfile. Any future entry only needs the C edit — Python symmetry is automatic. - Upstream source: fork-local. Netflix/vmaf has no ONNX op- allowlist surface; the entire
core/src/dnn/tree is fork- introduced. - On upstream sync: zero interaction. Every file touched lives in fork-introduced trees.
- Re-test on rebase:
meson test -C build test_op_allowlist test_onnx_scan
PYTHONPATH=ai/src python -m pytest ai/tests/test_op_allowlist.py
0231 — vif.comp + ciede.comp precise decorations (ADR-0269 / Step A of Vulkan 1.4 bump)¶
- Touches:
core/src/feature/vulkan/shaders/vif.comp(3 local-variable type qualifiers:g,sv_sq,gg_sigma_f→precise float),core/src/feature/vulkan/shaders/ciede.comp(yuv_to_rgboutputs,rgb_to_xyzmatmul accumulators,ciede2000chroma magnitudes + half-axes + s_l/c/h + lightness/chroma/hue + final ΔE). - Invariant: Both shaders are fork-local (Vulkan backend is fork-added; upstream Netflix/vmaf has no Vulkan compute kernels). The
precisekeyword is GLSL 4.50 standard syntax; glslc 2026.1 lowers it to per-resultOpDecorate NoContraction. The decorations are load-bearing for the cross-backend gate on NVIDIA driver 595.71+ — removing them would re-introduce the 42/48 ciede regression at API 1.3 documented in research-0054. - On upstream sync: zero interaction. Both shader files are entirely fork-introduced; upstream has no Vulkan compute path.
- Re-test on rebase:
# Re-confirm the cross-backend gate on a Vulkan-capable host.
meson setup core/build -Denable_vulkan=enabled
ninja -C core/build
python3 scripts/ci/cross_backend_vif_diff.py \
--vmaf-binary core/build/tools/vmaf \
--reference python/test/resource/yuv/src01_hrc00_576x324.yuv \
--distorted python/test/resource/yuv/src01_hrc01_576x324.yuv \
--width 576 --height 324 --pixel-format 420 --bitdepth 8 \
--feature vif --backend vulkan --places 4
python3 scripts/ci/cross_backend_vif_diff.py \
--vmaf-binary core/build/tools/vmaf \
--reference python/test/resource/yuv/src01_hrc00_576x324.yuv \
--distorted python/test/resource/yuv/src01_hrc01_576x324.yuv \
--width 576 --height 324 --pixel-format 420 --bitdepth 8 \
--feature ciede --backend vulkan --places 4
# Confirm SPIR-V still emits NoContraction post-rebase.
glslc --target-env=vulkan1.3 -O \
core/src/feature/vulkan/shaders/vif.comp -o /tmp/vif.spv
spirv-dis /tmp/vif.spv | grep -c NoContraction # expect ≥ 60
Expected on NVIDIA 595.71+: vif 0/48 OK, ciede 5/48 FAIL (max abs 8.9e-05 — pre-existing fork debt at API 1.3, see ADR-0269). On RADV / lavapipe: bit-exact (precise is a no-op there).
0229 — fr_regressor_v2 codec-aware scaffold (ADR-0272)¶
- ADR: ADR-0272
- Touches:
ai/scripts/train_fr_regressor_v2.py(new) — Phase A JSONL consumer; trains the codec-aware FRRegressor.model/tiny/fr_regressor_v2.onnx(new, smoke) — placeholder ONNX from--smokemode; re-baked on production training.model/tiny/fr_regressor_v2.json(new) — sidecar.model/tiny/registry.json— new entry withsmoke: true.docs/adr/0272-fr-regressor-v2-codec-aware-scaffold.md(new).docs/adr/README.md— index row.docs/research/0058-fr-regressor-v2-feasibility.md(new).docs/ai/models/fr_regressor_v2.md(new) — model card.ai/AGENTS.md— invariant note (codec block layout + ENCODER_VOCAB ordering).CHANGELOG.md— Added entry.- Invariant: the 8-D codec block layout is
[encoder_onehot(6), preset_norm, crf_norm]withENCODER_VOCAB = (libx264, libx265, libsvtav1, libvvenc, libvpx-vp9, unknown)in load-bearing order. CRF normaliser is/63(union upper bound). Preset normaliser is/9. Bumping the vocabulary requires a re-train; existing checkpoints pin the order they were trained against viaencoder_vocab_versionin the sidecar. The two-input ONNX (features,codec) follows the LPIPS-Sq precedent (ADR-0040 / ADR-0041). - Rebase impact: entirely fork-local; pure additive; no upstream-mirror file is touched. Phase A schema (consumed by this trainer) is itself fork-local (
tools/vmaf-tune/). No conflict expected on/sync-upstream. - Re-test on rebase:
0311 — libFuzzer harness expansion: yuv_input + cli_parse (ADR-0311)¶
- ADR: ADR-0311; parent ADR-0270.
- Touches:
core/test/fuzz/fuzz_yuv_input.c(new)core/test/fuzz/fuzz_cli_parse.c(new)core/test/fuzz/meson.build— two newexecutable(...)blocks for the harnesses, plus a sharedfuzz_vidinput_sourceslist.core/test/fuzz/yuv_input_corpus/*(new — 6 seeds covering 8/10-bit × 4:2:0 / 4:2:2 / 4:4:4 plus a truncated-frame seed).core/test/fuzz/cli_parse_corpus/*(new — 6 seeds covering the--feature,--model,--reference, YUV-flag, and--helpshapes).core/test/fuzz/README.md— Targets table extended..github/workflows/fuzz.yml— matrix gainsfuzz_yuv_input+fuzz_cli_parse; per-harness wall-clock budget reduced from 300 s to 60 s so the 3-target matrix fits the existingtimeout-minutes: 15cap.docs/development/fuzzing.md— runbook table + smoke commands extended.docs/adr/0311-libfuzzer-harness-expansion.md(new)docs/research/0083-libfuzzer-harness-expansion-target-survey.md(new)libvmaf/AGENTS.md— new invariant block for the one-parser-one-harness rule.CHANGELOG.md— Added entry.- Invariant:
- The fuzz scaffold remains opt-in (
-Dfuzz=true) — every defaultmeson setupinvocation must continue to skip it. fuzz_yuv_inputre-includestools/yuv_input.cand the rest of the vidinput trio as build inputs. Upstream Netflix/vmaf splits or renames of those source files need the matchingmeson.buildsource-list update.fuzz_cli_parsere-includestools/cli_parse.cas a build input and links againstlibvmafforvmaf_version()and feature-dictionary symbols. The-Wl,--wrap=exitlink arg is load-bearing — without it,usage()'sexit(1)would terminate the fuzzer process on first bad input.LLVMFuzzerTestOneInputkeeps external linkage; the scaffold-wide// NOLINTNEXTLINE(misc-use-internal-linkage)pattern is correct for libFuzzer's name-resolved entry-point ABI.- Rebase impact: any upstream sync that touches
core/tools/{yuv_input,cli_parse}.cmust re-run the 60 s smoke per harness on the merged tip; record any new-found crash-* artefact under the matching<target>_known_crashes/dir, not in<target>_corpus/. The__wrap_exitshim infuzz_cli_parse.cis GNU-ld / lld-only; do not assume it works on Apple ld without an-undefined,dynamic_lookupfallback. - Re-test on rebase:
CC=clang CXX=clang++ \
meson setup build-fuzz libvmaf \
--buildtype=debug \
-Db_sanitize=address \
-Db_lundef=false \
-Dfuzz=true \
-Denable_cuda=false -Denable_sycl=false -Denable_vulkan=disabled
ninja -C build-fuzz \
test/fuzz/fuzz_y4m_input \
test/fuzz/fuzz_yuv_input \
test/fuzz/fuzz_cli_parse
./build-fuzz/test/fuzz/fuzz_yuv_input \
-seed=0 -runs=1000 \
core/test/fuzz/yuv_input_corpus/
./build-fuzz/test/fuzz/fuzz_cli_parse \
-seed=0 -runs=1000 \
core/test/fuzz/cli_parse_corpus/
0229 — libFuzzer scaffold for the YUV4MPEG2 parser (ADR-0270)¶
- ADR: ADR-0270
- Touches:
core/test/fuzz/fuzz_y4m_input.c(new)core/test/fuzz/meson.build(new)core/test/fuzz/README.md(new)core/test/fuzz/y4m_input_corpus/*(new — six seeds)core/test/fuzz/y4m_input_known_crashes/*(new — one 411-chroma OOB reproducer; excluded from CI corpus)core/test/meson.build—subdir('fuzz')line.core/meson_options.txt— newoption('fuzz', ...)..github/workflows/fuzz.yml(new — nightly 5-minute job).docs/development/fuzzing.md(new — operator runbook).docs/adr/0270-fuzzing-scaffold.md(new)docs/research/0059-libfuzzer-scaffold-y4m.md(new)docs/state.md— new Open-bug row for the 411-chroma OOB write.CHANGELOG.md— Added entry.- Invariant: the fuzz scaffold is opt-in — every default
meson setupinvocation must continue to skip it. The harness links statically againstcore/tools/{y4m_input,yuv_input,vidinput}.crather thanlibvmaf.soso the public C-API surface stays unchanged. - Rebase impact: the harness re-includes
core/tools/y4m_input.cas a build input. Any upstream Netflix/vmaf change that splits or renames the tool sources (e.g. moves the parser intocore/src/) needs the correspondingmeson.buildsource list update and the harness re-test below. They4m_input_known_crashes/y4m_411_w2_h4_oob_dst.y4mreproducer is the regression gate for the parser fix; do not delete it on upstream sync — if upstream lands the same fix, port the reproducer back intoy4m_input_corpus/as a permanent seed. - Re-test on rebase:
CC=clang CXX=clang++ \
meson setup build-fuzz libvmaf \
--buildtype=debug \
-Db_sanitize=address \
-Db_lundef=false \
-Dfuzz=true \
-Denable_cuda=false -Denable_sycl=false -Denable_vulkan=disabled
ninja -C build-fuzz test/fuzz/fuzz_y4m_input
./build-fuzz/test/fuzz/fuzz_y4m_input \
-max_total_time=60 \
core/test/fuzz/y4m_input_corpus/
# Verify the known-crash reproducer still triggers (until the fix lands):
./build-fuzz/test/fuzz/fuzz_y4m_input \
core/test/fuzz/y4m_input_known_crashes/y4m_411_w2_h4_oob_dst.y4m
0231 — HIP seventh-consumer kernel float_motion_hip (ADR-0273)¶
- ADR: ADR-0273
- Touches:
core/src/feature/hip/float_motion_hip.c(new) — seventh consumer ofcore/src/hip/kernel_template.h. Mirrorscore/src/feature/cuda/float_motion_cuda.ccall-graph-for-call-graph;init/submit/collect/closeinvoke the kernel-template helpers in the same order;flush()callback for tail-frame motion2 emission;motion_force_zeroshort-circuit posture (fex->extractswap withsubmit / collect / flush / closenulled). Submit path intentionally bypassesvmaf_hip_kernel_submit_pre_launch(kernel writes per-WG SAD float partials directly, no atomic, no memset).core/src/feature/hip/float_motion_hip.h(new)core/src/hip/meson.build— new entry inhip_sources.core/src/feature/feature_extractor.c— extern declaration plusfeature_extractor_list[]entry under#if HAVE_HIP.core/test/test_hip_smoke.c— new sub-testtest_float_motion_hip_extractor_registered(also asserts theVMAF_FEATURE_EXTRACTOR_TEMPORALflag bit) and a row intest_table[].docs/adr/0273-hip-seventh-consumer-float-motion.md(new)docs/adr/README.md— index row.docs/backends/hip/overview.md— seventh / eighth consumer note.core/src/hip/AGENTS.md— invariant note.CHANGELOG.md— Added entry (joint with ADR-0274).- Invariant — three-buffer ping-pong +
motion_force_zeroshort-circuit are load-bearing. The state struct carries threeuintptr_tbuffer slots (ref_in,blur[2]) that the runtime PR (T7-10b) will swap for real device-buffer handles matching the CUDA twin'sVmafCudaBuffer *ref_in+VmafCudaBuffer *blur[2]field shape. Themotion_force_zeroshort-circuit (fex->extractswap, kernel-template helpers nulled) must stay aligned with the CUDA twin on every refactor — otherwise the runtime PR's helper-body flip diverges between the two backends. Thesubmit_pre_launchbypass mirrors the CUDA twin; if a future PR adds asubmit_pre_launchcall tofloat_motion_cuda.c's submit path, the HIP twin must follow in the same PR. - Rebase impact: entirely fork-local. New files are HIP-specific. The only upstream-touching edit is
feature_extractor.c, but the change sits inside an existing#if HAVE_HIPblock (ADR-0241); upstream has noHAVE_HIPso no conflict is expected. - Re-test on rebase:
meson setup build libvmaf -Denable_hip=true \
-Denable_cuda=false -Denable_sycl=false -Denable_vulkan=disabled
ninja -C build
meson test -C build test_hip_smoke
0232 — HIP eighth-consumer kernel float_ssim_hip (ADR-0274)¶
- ADR: ADR-0274
- Touches:
core/src/feature/hip/float_ssim_hip.c(new) — eighth consumer ofcore/src/hip/kernel_template.h. Mirrorscore/src/feature/cuda/integer_ssim_cuda.ccall-graph-for-call-graph (the CUDA file registersvmaf_fex_float_ssim_cudadespite itsinteger_filename). First multi-dispatch HIP consumer (chars.n_dispatches_per_frame == 2). Submit path intentionally bypassesvmaf_hip_kernel_submit_pre_launch(kernel writes per-block float partials directly). State struct carries fiveuintptr_tintermediate float buffer slots (h_ref_mu,h_cmp_mu,h_ref_sq,h_cmp_sq,h_refcmp) tracked outside the kernel-template's readback bundle.validate_dims_hipandinit_dims_hiphelpers extracted frominit()to fit thereadability-function-sizebudget.core/src/feature/hip/float_ssim_hip.h(new)core/src/hip/meson.build— new entry inhip_sources.core/src/feature/feature_extractor.c— extern declaration plusfeature_extractor_list[]entry under#if HAVE_HIP.core/test/test_hip_smoke.c— new sub-testtest_float_ssim_hip_extractor_registered(also assertschars.n_dispatches_per_frame == 2) and a row intest_table[].docs/adr/0274-hip-eighth-consumer-float-ssim.md(new)docs/adr/README.md— index row.docs/backends/hip/overview.md— seventh / eighth consumer note (joint).core/src/hip/AGENTS.md— invariant note.CHANGELOG.md— Added entry (joint with ADR-0273).- Invariant — multi-dispatch + five-slot buffer pyramid + v1
scale=1validation are load-bearing. The state struct carries fiveuintptr_tintermediate float buffer slots that the runtime PR (T7-10b) will swap for real device-buffer handles matching the CUDA twin'sVmafCudaBuffer *h_*field shape — any drift in the CUDA twin's slot count requires a paired update here. Thechars.n_dispatches_per_frame == 2characteristic is asserted in the smoke test; do not silently lower it. The v1scale=1-EINVALvalidation surface (invalidate_dims_hip) must stay aligned with the CUDA twin'scompute_scale/vmaf_logchain. The HIP twin'svalidate_dims_hip/init_dims_hipextraction is intentional for the function-size budget; do not re-inline without verifying the budget still passes. - Rebase impact: entirely fork-local; same posture as ADR-0273.
- Re-test on rebase:
meson setup build libvmaf -Denable_hip=true \
-Denable_cuda=false -Denable_sycl=false -Denable_vulkan=disabled
ninja -C build
meson test -C build test_hip_smoke
0229 — vmaf_tiny_v3 + vmaf_tiny_v4 dynamic-PTQ int8 sidecars (ADR-0275)¶
0278 — vmaf-tune libaom-av1 codec adapter (2026-05-03)¶
0228 — vmaf-tune libx265 codec adapter (ADR-0288)¶
0280 — vmaf-tune NVENC codec adapters (ADR-0290)¶
- Touches:
tools/vmaf-tune/src/vmaftune/codec_adapters/{h264_nvenc,hevc_nvenc,av1_nvenc,_nvenc_common}.py(new). Wholly fork-local — no upstream Netflix/vmaf overlap.tools/vmaf-tune/src/vmaftune/codec_adapters/__init__.py— registry expanded.tools/vmaf-tune/tests/test_codec_adapter_nvenc.py(new).tools/vmaf-tune/tests/test_corpus.py— Phase-A registry assertion updated.tools/vmaf-tune/AGENTS.md— invariant note expanded.docs/usage/vmaf-tune.md— "Hardware encoders (NVENC)" section.docs/adr/0290-vmaf-tune-nvenc-adapters.md(new) +docs/adr/README.mdindex row.docs/research/0065-vmaf-tune-nvenc-adapters.md(new).CHANGELOG.md— Added entry.- Invariant:
known_codecs()returns the four-codec tuple("av1_nvenc", "h264_nvenc", "hevc_nvenc", "libx264"); the mnemonic preset map (ultrafast/superfast/veryfast→p1,faster→p2,fast→p3,medium→p4,slow→p5,slower→p6,slowest/placebo→p7) is the canonical cross-codec preset alignment that downstream Phase B/C consumers assume. The CQ window is the hardware-permitted[0, 51]; the Phase A informative window is[15, 40]. - Rebase impact: zero —
tools/vmaf-tune/is wholly fork-local and has no upstream Netflix/vmaf path overlap. - Re-test on rebase:
0227 — ffmpeg-patches/ series re-verified against n8.1 (2026-05-03)¶
- Touches:
tools/vmaf-tune/src/vmaftune/codec_adapters/x265.py(new),tools/vmaf-tune/src/vmaftune/codec_adapters/__init__.py(registry add),tools/vmaf-tune/src/vmaftune/encode.py(parse_versions(stderr, encoder=…)gains a per-codec branch),tools/vmaf-tune/src/vmaftune/cli.py(help-text wording only),tools/vmaf-tune/tests/test_codec_adapter_x265.py(new),tools/vmaf-tune/tests/test_corpus.py(membership-based codec list assertion). - Invariant: the codec-adapter contract documented in
tools/vmaf-tune/AGENTS.md(multi-codec from day one; the search loop never branches on codec identity). Theparse_versionssignature is still backward-compatible —encoderdefaults tolibx264so callers from before this PR keep working. - Upstream source: fork-local.
tools/vmaf-tune/is fork-only; upstream Netflix/vmaf does not ship encode automation. - On upstream sync: zero interaction. Confirm the
_index_fragments/_order.txtrow for0288-vmaf-tune-codec-adapter-x265remains present after any cross-merge. - Re-test on rebase:
0227 — ffmpeg-patches/ series re-verified against n8.1 (2026-05-03)¶
- Touches:
tools/vmaf-tune/src/vmaftune/codec_adapters/libaom.py(new),tools/vmaf-tune/src/vmaftune/codec_adapters/__init__.py(registry row + import),tools/vmaf-tune/tests/test_corpus.py(membership assertion relaxed from== ("libx264",)to"libx264" in known_codecs()),tools/vmaf-tune/tests/test_codec_adapter_libaom.py(new),tools/vmaf-tune/AGENTS.md(preset-vocabulary invariant). - Invariant: the cross-codec preset vocabulary (
placebo, slowest, slower, slow, medium, fast, faster, veryfast, superfast, ultrafast) is shared across AV1-family adapters so one--presetaxis covers x264 / x265 / svtav1 / libaom-av1. Each adapter maps the human name onto its codec-specific knob; do not introduce per-adapter preset names. - Upstream source: fork-local.
tools/vmaf-tune/is the fork-introduced quality-aware encode automation harness (ADR-0237); it has no upstream Netflix/vmaf counterpart. - On upstream sync: zero interaction with
upstream/master. Self-contained intools/vmaf-tune/anddocs/. - Re-test on rebase:
0227 — ffmpeg-patches/ series re-verified against n8.1 (2026-05-03)¶
- ADR: ADR-0275
- Touches:
model/tiny/vmaf_tiny_v3.int8.onnx(new, 4 267 B)model/tiny/vmaf_tiny_v4.int8.onnx(new, 7 769 B)model/tiny/registry.json— newvmaf_tiny_v3andvmaf_tiny_v4rows withquant_mode,int8_sha256,quant_accuracy_budget_plccfields.model/tiny/vmaf_tiny_v3.json,model/tiny/vmaf_tiny_v4.json— same fields mirrored into the per-model sidecars.docs/ai/models/vmaf_tiny_v3.md,docs/ai/models/vmaf_tiny_v4.md— new "Quantisation" sections.docs/adr/0275-vmaf-tiny-v3-v4-ptq.md(new) and ADR index row.CHANGELOG.md— Added entry.- Invariant:
python ai/scripts/measure_quant_drop.py --allreports[PASS]for bothvmaf_tiny_v3(drop ≤ 0.001 on Netflix features) andvmaf_tiny_v4(drop ≤ 0.001), inside the 0.01 per-model budget. The runtime redirect from ADR-0174 picks the.int8.onnxsibling when an operator's registry overlay declaresquant_mode: dynamic. - Rebase impact: entirely fork-local — neither v3 nor v4 nor the dynamic-PTQ harness exists upstream. The new int8 ONNX bytes ship as committed binaries (mirroring
learned_filter_v1andnr_metric_v1); they are well below the few-MB external-data threshold and don't require the sigstore +.onnx.datapattern. - Re-test on rebase:
```bash python ai/scripts/validate_model_registry.py python ai/scripts/measure_quant_drop.py --all
0229 — NVIDIA-Vulkan ciede2000 places=4 fork debt root-cause (ADR-0273)¶
- Touched files: docs-only.
docs/adr/0273-...precision-gap.md(new) +_index_fragments/row +_order.txtappend.docs/research/0055-ciede-vulkan-nvidia-f32-f64-root-cause.md(new) +docs/research/README.mdindex row.docs/state.md— Open-bugs rowT-VK-CIEDE-F32-F64.docs/backends/vulkan/overview.md— NVIDIA-hardware caveat.changelog.d/changed/ciede-vulkan-nvidia-f32-f64-precision-gap.md(new).core/src/vulkan/AGENTS.md— invariant cross-link.- Invariant: the ciede.comp shader's f32 precision contract is load-bearing — promoting to f64 would silently change scores on every Vulkan device that supports
shaderFloat64and create a per-device-feature-bit divergence (RTX 4090 has it; many consumer GPUs don't). The CPUciede.c::get_lab_colordoing its colour-space chain indoubleis upstream Netflix behaviour and must not be narrowed to f32 to "fix" the GPU gap (would change Netflix golden ground truth). The 5/48 NVIDIA places=4 mismatch on the highest-ΔE frames is expected and documented; do not attempt to "fix" it without re-reading ADR-0273 first. - Rebase impact: zero — docs-only. The CPU and shader sources this ADR analyses are unchanged by this PR. If a future upstream rebase touches
ciede.c::get_lab_color(thedoublechain) the ADR's reasoning still holds; if upstream changes the CPU reference's precision posture, ADR-0273 needs aStatus: Supersededentry. - Re-test on rebase: a manual NVIDIA-hardware run if available:
```bash cd libvmaf && meson setup build \ -Denable_vulkan=enabled -Denable_cuda=false && ninja -C build cd .. python3 scripts/ci/cross_backend_vif_diff.py \ --vmaf-binary $PWD/core/build/tools/vmaf \ --reference testdata/ref_576x324_48f.yuv \ --distorted testdata/dis_576x324_48f.yuv \ --width 576 --height 324 \ --feature ciede --backend vulkan --device 0 --places 4 # Expected post-PR-346 (when merged): 5/48 mismatches at 1.78× threshold. # Expected pre-PR-346 (current master): 42/48 mismatches at higher ratio. # If the count drops below 5/48 on NVIDIA, ADR-0273 should record the # delta and consider closing T-VK-CIEDE-F32-F64.
0229 — tools/vmaf-tune fast Phase A.5 scaffold (ADR-0276)¶
- Touches:
tools/vmaf-tune/src/vmaftune/fast.py(new),tools/vmaf-tune/src/vmaftune/cli.py(newfastsubcommand branch),tools/vmaf-tune/pyproject.toml(new[fast]extra),tools/vmaf-tune/tests/test_fast.py(new),tools/vmaf-tune/AGENTS.md(new invariants),docs/usage/vmaf-tune.md(new "Phase A.5" section),docs/adr/0276-vmaf-tune-fast-path.md(new ADR),docs/research/0060-vmaf-tune-fast-path.md(new digest). - Invariant: the
fastsubcommand is opt-in and never automatically replaces the Phase A grid path. The slow grid is the ground-truth corpus generator (ADR-0237 contract); fast-path is for the recommendation use case only. Optuna is a lazy-imported optional dep gated behind the[fast]extra — importing it at module scope outsidefast.py(or its tests) breaks the zero-dep core install. - Rebase impact: entirely fork-local; the tool sits under
tools/vmaf-tune/which is fork-added, and no upstream files are touched. Upstream Netflix/vmaf has no analogous surface. - Re-test on rebase:
pip install -e 'tools/vmaf-tune[fast]'
pytest tools/vmaf-tune/tests/test_fast.py -v
vmaf-tune fast --smoke --target-vmaf 92
0229 — vmaf-tune recommend subcommand (ADR-0237 Phase B-lite)¶
- Touches:
tools/vmaf-tune/src/vmaftune/recommend.py(new). Wholly fork-local — no upstream Netflix/vmaf path overlap.tools/vmaf-tune/src/vmaftune/cli.py— addsrecommendsubparser;corpussubcommand untouched.tools/vmaf-tune/tests/test_recommend.py(new). 13-case smoke suite, mocks all binaries; runs in <100 ms.docs/usage/vmaf-tune.md— adds## recommendsection.- Invariant:
recommendconsumes the existingCORPUS_ROW_KEYSschema unchanged —vmaf_score,bitrate_kbps,crf,preset,encoder,exit_status. No schema bump. If a future PR bumpsSCHEMA_VERSION, both thecorpuswriter and therecommendreader must be updated in lockstep; tests assert this viatest_corpus_row_keys_match_init_contract. - Rebase impact: zero —
tools/vmaf-tune/is wholly fork-local; no upstream surface touches it. - Re-test on rebase:
0228 — integer_ms_ssim_cuda.c joins drain_batch (T-GPU-OPT-2 / ADR-0271)¶
- Touches:
core/src/feature/cuda/integer_ms_ssim_cuda.c. No upstream Netflix/vmaf changes expected here — the file is fork-added (CUDA twin of the upstream-portms_ssim_score.cu) and the surface this PR redrew (per-scalel_partials[i]/c_partials[i]/s_partials[i]arrays + the per-scaleh_l_partials[i]/h_c_partials[i]/h_s_partials[i]pinned host shadows + thesubmit()<→collect()work redistribution + thecuEventRecord(s->lc.finished, s->lc.str)+vmaf_cuda_drain_batch_register(&s->lc)tail) is also entirely fork-local. - Invariant: the engine-scope drain-batch contract from ADR-0271 / drain_batch.h. The kernel-launch order on
s->lc.strmust stay stable:decimate (× 4)then for each scalei ∈ 0..4horiz⇒vert_lcs⇒ DtoH(l_partials[i]) ⇒ DtoH(c_partials[i]) ⇒ DtoH(s_partials[i])thencuEventRecord(s->lc.finished, s->lc.str)thenvmaf_cuda_drain_batch_register(&s->lc). Same-stream ordering is what makes the shared SSIM intermediates (h_ref_mu,h_cmp_mu,h_ref_sq,h_cmp_sq,h_refcmp`) safe across scales without explicit sync — any change that parallelises the per-scale work onto multiple streams breaks bit-exactness unless per-scale intermediates are also added. - On upstream sync: zero interaction (the file is fork-added). If a future upstream PR adds an
integer_ms_ssim_cuda.cof its own, the merger must reconcile the per-scale partials topology + the drain_batch tail with whatever the new upstream shape brings. - Re-test on rebase:
cd libvmaf && meson setup build -Denable_cuda=false -Denable_sycl=false
ninja -C build # confirms the CPU build still links cleanly
# If the dev host has a working nvcc / host-compiler pair:
meson setup build_cuda -Denable_cuda=true -Denable_sycl=false
ninja -C build_cuda src/liblibvmaf_feature.a.p/feature_cuda_integer_ms_ssim_cuda.c.o
# Netflix CPU golden gate (CPU is the bit-exactness ground truth):
make test-netflix-golden
# Cross-backend parity (places=4 gate, ADR-0214):
/cross-backend-diff
0277 — ffmpeg-patches refresh against n8.1 — 2026-05-04 (ADR-0277)¶
- Touches:
ffmpeg-patches/is unchanged (no content drift). Doc-only entries land in: docs/adr/0277-ffmpeg-patches-refresh-2026-05-04.md— new ADR.docs/adr/_index_fragments/0277-ffmpeg-patches-refresh-2026-05-04.md— index row.docs/adr/_index_fragments/_order.txt— manifest append.changelog.d/changed/ffmpeg-patches-refresh-2026-05-04.md— Changed entry.- This file — this entry.
- Invariant:
ffmpeg-patches/series.txtorder is load-bearing — patches0002…0006build on each other and only apply cleanly cumulatively. The verification gate is a series replay, not a per-patchgit apply --check(per ADR-0118 + CLAUDE.md §12 r14). - On upstream sync: zero interaction. Netflix/vmaf has no
ffmpeg-patches/tree; this is a fork-local integration surface. - Re-test on rebase (also: re-replay procedure for the next refresh):
# Clone pristine n8.1
git -C /tmp clone --depth 1 --branch n8.1 \
https://github.com/FFmpeg/FFmpeg.git ff-replay-$(date +%F)
cd /tmp/ff-replay-$(date +%F)
git switch -c refresh-$(date +%F)
git config user.email refresh@local && git config user.name "Refresh Bot"
# Replay the series cumulatively
for p in /path/to/vmaf/ffmpeg-patches/000*-*.patch; do
git am --3way "$p" || break
done
# Regenerate and compare to in-tree
mkdir -p /tmp/ff-regen-$(date +%F)
git format-patch n8.1.. -o /tmp/ff-regen-$(date +%F)/
# Diff old vs new excluding pure format-patch noise
for i in 1 2 3 4 5 6; do
orig=$(ls /path/to/vmaf/ffmpeg-patches/000${i}-*.patch)
regen=$(ls /tmp/ff-regen-$(date +%F)/000${i}-*.patch)
diff -u \
<(grep -v "^From [0-9a-f]\|^Date:\|^index " "$orig") \
<(grep -v "^From [0-9a-f]\|^Date:\|^index " "$regen") \
| head -40
done
If only stylistic diffs surface (PATCH N/M numbering, MIME headers, hunk-context counts, hunk offset shifts against cumulative state), keep originals — record a no-drift refresh ADR. If real content drift surfaces, regenerate and ship the refresh PR with the regenerated patches plus a content-summary ADR.
End-to-end vf_libvmaf smoke is best run from CI (ffmpeg-integration.yml) against an installed libvmaf prefix — the meson-uninstalled .pc does not satisfy FFmpeg's #include <libvmaf.h> probe (the headers live under libvmaf/libvmaf.h only; the system-installed .pc carries an extra -I${includedir}/libvmaf shortcut that the uninstalled .pc omits).
0229 — T7-5 NOLINT-sweep closeout (ADR-0278)¶
- Touched files:
core/src/feature/integer_adm.c(1 NOLINT cite, line ~988adm_decouple_s123— upstream-mirror Netflix966be8d5).core/src/feature/cuda/ssimulacra2_cuda.c(3 NOLINT cites:ss2c_picture_to_linear_rgb,ss2c_host_combine,ss2c_run_scale_gpu/extract_fex_cuda).core/src/feature/vulkan/ssimulacra2_vulkan.c(3 NOLINT cites:ss2v_setup_gaussian,ss2v_picture_to_linear_rgb,ss2v_run_scale).core/src/feature/vulkan/cambi_vulkan.c(1 NOLINT cite:cambi_vk_extract).core/src/feature/sycl/integer_adm_sycl.cpp(6 cites, SYCL kernel-launch entries).core/src/feature/sycl/integer_motion_sycl.cpp(2 cites).core/src/feature/sycl/integer_vif_sycl.cpp(4 cites).core/tools/vmaf.c(3 cites:copy_picture_data,init_gpu_backends,main).- Invariant: zero behavioural change. Edits are inside comment blocks — appended
(ADR-0141 §2 ... load-bearing invariant; T7-5 sweep closeout — ADR-0278)to existing prose justifications. No function bodies split. The 12 SYCL sites share an identical justification string verbatim; preserving the byte-for-byte duplicate is the load-bearing documentation pattern (grep-able across the SYCL TUs). - On upstream sync: minimal interaction. The cite-only edits live inside comment blocks above the function signatures; rebases will surface them as touched lines but the function bodies are unchanged. For
integer_adm.c's upstream-mirror block (Netflix966be8d5), the comment edit at line 984–991 is cosmetic — keep the fork's version on conflict (it merely names the ADR; the underlying prose is unchanged). - Re-test on rebase:
```bash # 1. Programmatic audit must report 0 missing citations python3 - <<'PY' import re, os paths = [os.path.join(r, f) for r, _, fs in os.walk('libvmaf/src') for f in fs if f.endswith(('.c','.cpp','.h'))] paths.append('core/tools/vmaf.c') miss = total = 0 for p in paths: with open(p) as fh: ls = fh.readlines() for i, line in enumerate(ls): if 'NOLINT' in line and 'readability-function-size' in line and 'NOLINTEND' not in line: total += 1 ctx = [line]; j = i - 1 while j >= 0 and j > i - 14: s = ls[j].strip() if not s: break if s.startswith(('//','/','')): ctx.insert(0, ls[j]); j -= 1 else: break buf = ''.join(ctx) if 'ADR-' not in buf and not re.search(r'[Rr]esearch-?\d', buf): miss += 1 print(f"sites={total} missing={miss}") PY
# 2. Build + Netflix golden gate meson setup build -Denable_cuda=false -Denable_sycl=false ninja -C build make test-netflix-golden
0231 — vmaf-tune score path decodes mp4 -> raw YUV¶
- Touches:
tools/vmaf-tune/src/vmaftune/score.py(new_decode_to_raw_yuv+_needs_decodehelpers,run_scoreshells out to ffmpeg whenreq.distorted.suffix not in {.yuv, .y4m});tools/vmaf-tune/tests/test_corpus.py(3 new regression tests + the smoke-end-to-end mock now also stubs the ffmpeg decode call). - Invariant: the decode-back is the contract the libvmaf CLI imposes — mp4/webm/etc.
--distortedis silently rejected as raw-yuv with the wrong byte count, surfacing asexit_status=234. Future encoder adapters that emit non-raw containers inherit this decode automatically. Do not "optimise" the temp YUV away without first migrating the corpus pipeline to theffmpeg+libvmaffilter (which can pipe an mp4 stream in directly). - On upstream sync: zero interaction.
vmaf-tuneis fork-only tooling; upstream Netflix/vmaf has no analogue. - Re-test on rebase:
```bash cd tools/vmaf-tune && python3 -m pytest tests/ # plus an end-to-end smoke (needs a real raw YUV + ffmpeg + vmaf): ./vmaf-tune corpus --source /path/to/ref.yuv --width 1920 \ --height 1080 --pix-fmt yuv420p --framerate 25 --duration 6 \ --encoder libx264 --preset medium --crf 23 \ --output /tmp/smoke.jsonl --no-source-hash # expect: vmaf_score is a real number, not NaN.
0232 — CUDA build pins nvcc --std c++20¶
- Touches:
core/src/meson.buildline 686 (cuda_flags = [...]). - Invariant: nvcc 12.x clamps host C++ at C++17 by default; 13.x accepts up to C++20. Bumping the host stdlib past nvcc's default (any gcc >= 16, libstdc++ ships C++23 features) breaks the host-side parse in
<type_traits>/<bits/utility.h>. Forcing--std c++20on CUDA 13+ keeps the host headers parseable. Do not drop this flag without first checking the host gcc version against nvcc's default. - On upstream sync: zero interaction. Netflix/vmaf doesn't ship the
cuda_flagslist shape we use (their CUDA build is the original pre-fork pattern); a sync that touchescore/src/meson.buildaround theis_cuda_enabledbranch should keep the--std c++20injection. - Re-test on rebase:
meson setup core/build-cuda -Denable_cuda=true \
-Denable_sycl=false -Denable_vulkan=disabled
ninja -C core/build-cuda
# smoke
./core/build-cuda/tools/vmaf --gpumask=0 --no_sycl --no_vulkan \
-r .workingdir2/netflix/ref/BigBuckBunny_25fps.yuv \
-d .workingdir2/netflix/ref/BigBuckBunny_25fps.yuv \
-w 1920 -h 1080 -p 420 -b 8
0233 — CUDA motion flush_fex_cuda idempotency guard¶
- Touches:
core/src/feature/cuda/integer_motion_cuda.c— factored anappend_if_unwrittenhelper and routed the two motion2 / motion3 final-frame writes through it. - Invariant: under T-GPU-OPT-1 (PR #312 / ADR-0242), the pending-collect inside
flush_context_cudamay already have writtenmotion2_score[s->index]/motion3_score[s->index]beforeflush_fex_cudaruns. Any future motion-cuda flush logic that emits the same (feature, index) pair must keep this idempotency contract orflush_context_cudawill mis-surface as "context could not be synchronized". - On upstream sync: the bug only exists because the fork's
flush_context_cudaruns the pending-collect before the per-extractor flush. Netflix/vmaf upstream doesn't have the T-GPU-OPT-1 drain pattern, so the pre-#312 code path didn't duplicate-write. If Netflix lands a similar pattern, the fix shape mirrors what's done here. - Re-test on rebase:
ninja -C core/build-cuda
./core/build-cuda/tools/vmaf --gpumask=0 --no_sycl --no_vulkan \
-r python/test/resource/yuv/src01_hrc00_576x324.yuv \
-d python/test/resource/yuv/src01_hrc01_576x324.yuv \
-w 576 -h 324 -p 420 -b 8 \
--model path=model/vmaf_v0.6.1.json --threads 1 -q \
--output /tmp/cuda.json --json
# Expect: clean run, no "cannot be overwritten" warning,
# no "problem flushing context" error.
0234 — hw_encoder_corpus.py Phase A real-corpus runner¶
- Touches: new
scripts/dev/hw_encoder_corpus.py(no existing caller; opt-in tooling). Output landing inruns/phase_a/is gitignored — rerun the script to reproduce.docs/development/intel-arc-vaapi-driver-priority.md. Output landing inruns/phase_a/is gitignored — rerun the script to reproduce. stratified sample, 58 KiB). - Invariant: the script's QSV path forces
env['LIBVA_DRIVER_NAME']='iHD'(set by the calling shell, not inside the script) when targeting/dev/dri/renderD129on a multi-card host that has NVIDIA's libva-driver-nvidia shim installed. Without that, libva picks up NVIDIA's NVDEC-VAAPI translation and the MFX session handshake fails with -9. See the companion doc for the failure mode + fix. - On upstream sync: zero interaction. The script lives under
scripts/dev/(fork-only); upstream Netflix/vmaf has no comparable Phase A corpus tooling. - Re-test on rebase:
python3 scripts/dev/hw_encoder_corpus.py \
--vmaf-bin core/build-cuda/tools/vmaf \
--source .workingdir2/netflix/ref/BigBuckBunny_25fps.yuv \
--width 1920 --height 1080 --pix-fmt yuv420p --framerate 25 \
--encoder h264_nvenc --cq 25 \
--out /tmp/smoke.jsonl
# Expect: 1 cell × ~150 frames, per-frame canonical-6 + vmaf,
# encoder=h264_nvenc, cq=25.
0235 — fr_regressor_v2 ENCODER_VOCAB v2 (hw codec extension)¶
- Touches:
ai/scripts/train_fr_regressor_v2.py—ENCODER_VOCABgains 6 hw-codec entries (3 NVENC + 3 QSV);ENCODER_VOCAB_VERSIONbumps 1 -> 2;PRESET_ORDINALgains 6 sub-tables forp1..p7(NVENC) and the libx264-aligned QSV preset family. - Invariant: vocab order is load-bearing — index of every entry is baked into trained model graphs as a one-hot column position. New entries MUST be appended (never inserted into the middle), and the
unknownsentinel MUST stay last (UNKNOWN_ENCODER_INDEX = N - 1). BumpingENCODER_VOCAB_VERSIONsignals that any v1-graph ONNX needs re-export against v2 before consuming v2 training rows. - On upstream sync: zero interaction.
train_fr_regressor_v2.pyis fork-only (Phase B prereq, ADR-0237 / ADR-0272). - Re-test on rebase:
python3 ai/scripts/train_fr_regressor_v2.py --corpus <jsonl> --epochs 200 --no-export— expect PLCC > 0.95 on a multi-codec corpus.
0276 — vmaf_tiny_v5 corpus-expansion probe (ADR-0287) — defer¶
- What changed: research-only addition. New scripts under
ai/scripts/(fetch_youtube_ugc_subset.py,extract_ugc_features.py,train_vmaf_tiny_v5.py,eval_loso_vmaf_tiny_v5.py), new ADRdocs/adr/0276-*.md, new research digestdocs/research/0057-*.md, and one CHANGELOG entry. No new ONNX artefact undermodel/tiny/, no registry change, no public C-API / CLI / meson_options change. The probe trained an architecturally identical mlp_small on a 5-corpus parquet (4-corpus + 27 000 UGC rows); the 1-σ ship gate did not clear (Δ PLCC = +0.00005), so the exporter that the prior agent had drafted (export_vmaf_tiny_v5.py) was discarded before the commit. - Upstream source: fork-local. Netflix/vmaf has no tiny-AI corpus-expansion surface; nothing on the upstream side touches these files.
- On upstream sync: zero interaction. The v5 surface lives entirely under
ai/scripts/+docs/adr/+docs/research/, all of which are fork-introduced trees. The shipped v2 model (model/tiny/vmaf_tiny_v2.onnx) and its registry row are untouched. - Re-test on rebase:
# No code under test on rebase — purely research artefacts.
# If revisiting the corpus expansion, the reproducer is in the
# research digest:
python3 ai/scripts/fetch_youtube_ugc_subset.py \
--out-dir .workingdir2/ugc/download \
--n-stems 30 \
--manifest .workingdir2/ugc/manifest.json
python3 ai/scripts/extract_ugc_features.py \
--manifest .workingdir2/ugc/manifest.json \
--yuv-dir .workingdir2/ugc/yuv \
--vmaf-bin build-cpu/tools/vmaf \
--out-parquet runs/full_features_ugc.parquet \
--max-height 360 --max-frames 300 --threads 8
python3 ai/scripts/eval_loso_vmaf_tiny_v5.py \
--parquet-base runs/full_features_4corpus.parquet \
--parquet-extra runs/full_features_ugc.parquet \
--out-json runs/vmaf_tiny_v5_loso_metrics.json
0227 — vmaf-tune Intel QSV codec adapters (ADR-0281)¶
- What changed: fork-local additions under
tools/vmaf-tune/src/vmaftune/codec_adapters/—_qsv_common.py,h264_qsv.py,hevc_qsv.py,av1_qsv.py, plus registry rows incodec_adapters/__init__.pyand a new test filetools/vmaf-tune/tests/test_codec_adapter_qsv.py. Doc updates:docs/usage/vmaf-tune.md(Hardware encoders section),docs/adr/0281-vmaf-tune-qsv-adapters.md,docs/research/0066-vmaf-tune-qsv-adapters.md,tools/vmaf-tune/AGENTS.md,CHANGELOG.md. - Upstream source: fork-local.
tools/vmaf-tune/is fork-introduced under ADR-0237; Netflix/vmaf has no corresponding tree. - On upstream sync: zero interaction. Upstream cannot conflict with this PR's paths.
- Invariant: the registry exposes exactly four codecs (
av1_qsv,h264_qsv,hevc_qsv,libx264— alphabetical), each adapter validates its(preset, quality)pair, and the QSV preset vocabulary is the seven x264-style names (veryslow…veryfast, noultrafast/superfast). The encode pipeline (encode.py) remains x264-CRF-tied and will be widened in a separate PR — the QSV adapters are inert until then. Future codec families that share parameter shape (NVENC, AMF) follow the same_<family>_common.py+ N thin adapters pattern. - Re-test on rebase:
0230 — K150K-A corpus extraction script (ADR-0362)¶
- Touches:
ai/scripts/extract_k150k_features.py(new fork-only file),ai/AGENTS.md(K150K invariant note appended),docs/adr/README.md(ADR-0362 index row),CHANGELOG.md,docs/rebase-notes.md. - Invariant:
extract_k150k_features.pyrequiresbuild-cpu/tools/vmaf(fork build withssimulacra2+motion_v2). If upstream Netflix adds these extractors to their own release binary, the--vmaf-bindefault may be updated to the system binary -- but only after verifying that the metric JSON key names match the aliases in_METRIC_ALIASES. TheFEATURE_NAMEStuple is column-order-locked to the parquet schema; any reorder invalidates trained checkpoints that consume the parquet. - Upstream interaction: none. Script is fork-only; the K150K clips and parquet are gitignored. Upstream Netflix/vmaf does not ship a K150K extractor.
- Re-test on rebase:
0229 — vmaf-tune libvvenc + NN-VC codec adapter (ADR-0285)¶
- Touches:
tools/vmaf-tune/src/vmaftune/codec_adapters/vvenc.py(new fork-only file),tools/vmaf-tune/src/vmaftune/codec_adapters/__init__.py(registry edit, fork-only),tools/vmaf-tune/tests/test_codec_adapter_vvenc.py(new),tools/vmaf-tune/tests/test_corpus.py(relaxes theknown_codecs() == ("libx264",)assertion to"libx264" in known_codecs()since the registry now spans multiple codecs). - Invariant: the codec-adapter registry is fork-introduced (Phase A of ADR-0237) and lives entirely outside the upstream Netflix tree, so
tools/vmaf-tune/does not touch upstream paths. The only rebase-sensitive surface is theCORPUS_ROW_KEYSschema insrc/vmaftune/__init__.py(per the Phase A invariant intools/vmaf-tune/AGENTS.md); this PR adds the adapter without changing the schema. - Upstream interaction: none.
tools/vmaf-tune/is not in Netflix/vmaf upstream. - Re-test on rebase:
- Status update 2026-05-09: the original
nnvc_intratoggle was removed (it emitted a fabricatedIntraNNkey that does not exist in any released VVenC). Replaced with a curated 9-knob real-VVenC 1.14.0 tuning surface (PerceptQPA,InternalBitDepth,Tier,Tiles,MaxParallelFrames,RPR,SAO,ALF,CCALF). Defaults preserve the bit-exact Phase A grid baseline.adapter_versionbumped to"2"so cache keys invalidate. See ADR-0285 §"Status update 2026-05-09".no rebase impact: REASON(fork-local file, no upstream-tree touch).
0228 — vmaf-tune Phase D scaffold (ADR-0276)¶
- Touches:
tools/vmaf-tune/src/vmaftune/per_shot.py,tools/vmaf-tune/src/vmaftune/cli.py,tools/vmaf-tune/tests/test_per_shot.py,docs/usage/vmaf-tune.md,docs/adr/0276-vmaf-tune-phase-d-per-shot.md. - Invariant: scaffold-only. The module relies on a stable predicate signature
(shot, target_vmaf, encoder) -> (crf, predicted_vmaf)that Phase B's bisect (PR #347) drops into later.Shotranges are half-open[start_frame, end_frame)even though the C-sidevmaf-perShotJSON/CSV sidecar uses an inclusiveend_frame— normalisation happens at the parse boundary in_parse_per_shot_json/parse_per_shot_csv.vmaf-perShotschema lives indocs/usage/vmaf-perShot.mdand is fork-local (ADR-0222), so upstream cannot drift it; the only rebase risk is fork-internal renames. - Upstream source: entirely fork-local.
tools/vmaf-tune/is fork-introduced (ADR-0237). Netflix/vmaf upstream has no encode-automation surface. - On upstream sync: zero interaction expected. No file in this PR overlaps an upstream-mirrored path.
- Re-test on rebase:
python -m pytest tools/vmaf-tune/tests/test_per_shot.py -q
python tools/vmaf-tune/vmaf-tune tune-per-shot --help
0229 — vmaf-tune SVT-AV1 codec adapter (ADR-0278)¶
- Touches:
tools/vmaf-tune/src/vmaftune/codec_adapters/svtav1.py(new),tools/vmaf-tune/src/vmaftune/codec_adapters/__init__.py(registry),tools/vmaf-tune/src/vmaftune/encode.py(parse_versionsextended for the SVT-AV1 banner pattern),tools/vmaf-tune/src/vmaftune/corpus.py(optionalffmpeg_preset_tokenhook). - Invariant:
PRESET_NAME_TO_INTis closed and order-stable; the integer values are baked into corpus rows that downstreamfr_regressor_v2(ADR-0235) trains on. Reordering or rewriting the table silently changes the integer SVT-AV1 receives. The codec key"libsvtav1"matchesCODEC_VOCAB[2]inai/src/vmaf_train/codec.py— keep them aligned on any rename. - Upstream source: fork-local.
tools/vmaf-tune/is a fork-introduced tree (see entry 0227 — Phase A scaffold). No Netflix/vmaf upstream interaction. - On upstream sync: zero interaction. Lives entirely under the fork-local
tools/vmaf-tune/tree. - Re-test on rebase:
0230 — fr_regressor_v2 PROD ship (ADR-0352)¶
- ADR: ADR-0352
0230 — fr_regressor_v2 PROD ship (ADR-0291)¶
-
ADR: ADR-0291
-
Touches:
model/tiny/fr_regressor_v2.onnx(binary, refreshed),model/tiny/fr_regressor_v2.json(sidecar, sha256 + metrics),model/tiny/registry.json(smoke flag flip, sha256 update),runs/phase_a/full_grid/per_frame_canonical6.jsonl(training corpus — fork-local artefact underruns/), companion docs. - Re-test recipe: see Research-0068 §Reproducer. Ship gate is LOSO PLCC ≥ 0.95 on the per-source folds; current run reports 0.9681 ± 0.0207.
- Rebase invariant: the per-frame canonical-6 corpus must be rebuilt from
runs/phase_a/{nvenc,qsv}_pf.jsonl(PR #392) before any retrain; do not re-train against the cell-onlycomprehensive.jsonl(it lacks the per-frame features and produces PLCC ≈ 0.7 — the smoke baseline). - No upstream interaction:
fr_regressor_v2is fork-local (ADR-0272).
0229 — vmaf-tune Phase E ladder generator (ADR-0295)¶
- ADR: ADR-0295
- Touches: entirely fork-local under
tools/vmaf-tune/. New moduletools/vmaf-tune/src/vmaftune/ladder.py, new test filetools/vmaf-tune/tests/test_ladder.py, two new subcommand blocks intools/vmaf-tune/src/vmaftune/cli.py. No upstream-shared paths touched. - Invariant:
vmaftune.ladder.convex_hullreturns a strictly monotonic Pareto frontier (both bitrate and vmaf monotonically increasing);select_kneesreturns exactlymin(n, len(hull))rungs in ascending bitrate order;emit_manifest("hls")produces one#EXT-X-STREAM-INFper rung with monotonically-increasingBANDWIDTH=values. The default_default_sampleris intentionallyNotImplementedError— production callers must inject a Phase B bisect-driven sampler. Phase B integration PR (gated on PR #347) swaps the default; the test suite continues to inject a synthetic stub. - Rebase impact: none — fork-local Python tool; upstream Netflix/vmaf does not ship a
tools/vmaf-tune/tree. - Re-test on rebase:
0229 — fr_regressor_v2 probabilistic head scaffold (ADR-0279)¶
- Touches:
ai/scripts/train_fr_regressor_v2_ensemble.py(new — fork-local).ai/scripts/eval_probabilistic_proxy.py(new — fork-local).model/tiny/fr_regressor_v2_ensemble_v1*.onnx,fr_regressor_v2_ensemble_v1.json(new artefacts; smoke probes).model/tiny/registry.json— five newkind: "fr"rows (fr_regressor_v2_ensemble_v1_seed{0..4}); existing entries untouched.ai/AGENTS.md— new "fr_regressor_v2_ensemble_v1 — probabilistic head" section pinning the per-member ONNX I/O contract, manifest-as-runtime-entry-point invariant, ensemble-size pin, confidence-rule one-of, codec-vocab parity, and smoke-artefact posture.docs/ai/models/fr_regressor_v2_probabilistic.md(new model card).docs/research/0067-fr-regressor-v2-probabilistic.md(new audit digest).docs/adr/0279-fr-regressor-v2-probabilistic.md(new ADR; Proposed). Index row appended todocs/adr/README.md.CHANGELOG.md—### Addedrow under "Unreleased — lusoris fork".- Invariant: the per-member ONNX I/O contract (two inputs:
features [N, 6]standardised +codec_onehot [N, NUM_CODECS]; one outputscore [N]) and the manifest'sconfidencerule (one-of"ensemble"/"ensemble+conformal") are the C-side adapter's load-bearing contract. Per-member ensembles are stockFRRegressor(num_codecs=NUM_CODECS)calls — flipping to a v1-shaped single-input graph silently invalidates the manifest.CODEC_VOCABparity withai/src/vmaf_train/codec.pyis required. - On upstream sync: zero interaction expected. Wholly fork-local; no upstream Netflix/vmaf path overlap. The
ai/package is fork-introduced (see ADR-0021, ADR-0036) — upstream has no probabilistic-regressor surface. If upstream ever ships its ownfr_regressor_v2variant, do NOT merge — register both ids side-by-side. - Re-test on rebase:
python ai/scripts/train_fr_regressor_v2_ensemble.py --smoke
python ai/scripts/eval_probabilistic_proxy.py --smoke
python ai/scripts/validate_model_registry.py
0287 — vmaf-tune saliency-aware ROI tuning (ADR-0293)¶
- Touches:
tools/vmaf-tune/src/vmaftune/saliency.py,tools/vmaf-tune/src/vmaftune/cli.py(newrecommendsubcommand),tools/vmaf-tune/AGENTS.md(saliency invariant),docs/usage/vmaf-tune.md(saliency section). - Upstream source: fork-local. The
vmaf-tunetree was introduced in PR #329 (ADR-0237 Phase A) and has no upstream Netflix counterpart. - On upstream sync: zero interaction — pure fork-local Python package under
tools/vmaf-tune/. - Invariant: the saliency-to-QP-offset signal blend (
offset = (2*sal − 1) * foreground_offset, clamped to ±12) is bit-for-bit equivalent tovmaf-roi's C-side blend (ADR-0247).tests/test_saliency.pypins the contract; ifvmaf-roi's C blend changes,saliency.pyfollows in the same PR. The test seam contract (session_factory=…,encode_runner=…) lets the suite run withoutonnxruntimeorffmpeg. - Re-test on rebase:
0229 — tools/vmaf-roi-score/ Option C scaffold (ADR-0296)¶
- ADR: ADR-0296
- Touches:
tools/vmaf-roi-score/pyproject.toml(new)tools/vmaf-roi-score/vmaf-roi-score(new console shim)tools/vmaf-roi-score/src/vmafroiscore/__init__.py(new)tools/vmaf-roi-score/src/vmafroiscore/cli.py(new)tools/vmaf-roi-score/src/vmafroiscore/score.py(new)tools/vmaf-roi-score/src/vmafroiscore/mask.py(new)tools/vmaf-roi-score/tests/test_combine.py(new)tools/vmaf-roi-score/README.md(new)tools/vmaf-roi-score/AGENTS.md(new)docs/adr/0296-vmaf-roi-saliency-weighted.md(new)docs/adr/_index_fragments/0296-vmaf-roi-saliency-weighted.md(new)docs/adr/_index_fragments/_order.txt— append-only.docs/research/0069-vmaf-roi-saliency-weighted.md(new)docs/usage/vmaf-roi-score.md(new)changelog.d/added/T6-2c-vmaf-roi-score-scaffold.md(new)- Invariant:
tools/vmaf-roi-score/is wholly fork-local. No upstream Netflix/vmaf surface owns or interacts with this directory. The combine math is a pure linear blend on Pythonfloat; the JSON schema is pinned byROI_RESULT_KEYSandSCHEMA_VERSION = 1. Schema bumps require an ADR-0288 supersession. Naming guard: do not confuse withcore/tools/vmaf_roi.c(ADR-0247) — that's the encoder-steering binary. The scoring tool here isvmaf-roi-score; the names diverge deliberately. - Rebase impact: zero. Pure-Python tool under
tools/; not part of the libvmaf C build, not part of any Netflix-mirrored surface. - Re-test on rebase:
0228 — vmaf-tune compare codec-comparison mode (research-0061 Bucket #7)¶
- Touches:
tools/vmaf-tune/src/vmaftune/compare.py(new). Wholly fork-local; no upstream Netflix/vmaf path overlap.tools/vmaf-tune/src/vmaftune/cli.py— adds thecomparesubparser and_run_comparerouter.tools/vmaf-tune/tests/test_compare.py(new). Mocked predicate; noffmpeg/vmafbinaries required.tools/vmaf-tune/AGENTS.md— invariant note for the predicate seam andCOMPARE_ROW_KEYScontract.docs/usage/vmaf-tune.md— new "Codec comparison" section.- Invariant:
compare.compare_codecsorchestrates per-codec ranking via an injectedpredicate(codec, src, target_vmaf) -> RecommendResultcallable. The orchestration must not branch on codec name; new codecs land as one-file additions undercodec_adapters/and are picked up automatically by the registry.COMPARE_ROW_KEYSis the JSON / CSV column contract — same maintenance discipline asCORPUS_ROW_KEYS. - Rebase impact: entirely fork-local. The Phase A + Phase B recommend backend (ADR-0237) is fork-internal; upstream Netflix/vmaf has no
tools/vmaf-tune/tree. - Re-test on rebase:
```shell pytest tools/vmaf-tune/tests/test_compare.py -v PYTHONPATH=tools/vmaf-tune/src python -m vmaftune.cli compare \ --src /tmp/ref.yuv --target-vmaf 92 --format markdown
0229 — vmaf-tune --score-backend GPU score wiring (ADR-0299)¶
- Touches:
tools/vmaf-tune/src/vmaftune/score_backend.py(new). Wholly fork-local —tools/vmaf-tune/has no upstream Netflix/vmaf overlap.tools/vmaf-tune/src/vmaftune/{score,corpus,cli}.py(additive kwargs, no API removals).tools/vmaf-tune/tests/test_score_backend.py(new).docs/usage/vmaf-tune.md(new GPU section + flag row).docs/adr/0299-vmaf-tune-gpu-score.md(new).docs/research/0071-vmaf-tune-gpu-score-backend.md(new).- Invariant: the libvmaf CLI exposes
--backend NAMEwith valuesauto|cpu|cuda|sycl|vulkanexactly. Help-text parser inscore_backend.parse_supported_backendspins this format. If upstream renames the flag or reformats the help line on merge, the parser silently degrades to "CPU only" — the test fixtures intest_score_backend.pywill catch the format change but only if re-run. - Upstream source: fork-local. Netflix upstream's CLI does not ship a
--backendselector (CPU-only). - On upstream sync: zero interaction.
vmaf-tunelives entirely in fork-introduced paths and consumes only the fork's--backendflag. - Re-test on rebase:
pytest tools/vmaf-tune/tests/test_score_backend.py -v
# If the libvmaf help text reformats, parse_supported_backends
# will return {"cpu"} on test_parse_full_backend_line_yields_all_four
# and the test fails loudly.
0261 — vmaf-tune HDR-aware encode + score path (2026-05-03)¶
- What changed: fork-local addition under
tools/vmaf-tune/src/vmaftune/hdr.pyplus wiring intocorpus.py/cli.py/score.py. Adds ffprobe-driven HDR detection, codec-specific HDR ffmpeg flag dispatch, schema-v2 corpus row keys (hdr_transfer,hdr_primaries,hdr_forced), and four--auto-hdr/--force-*CLI modes. See ADR-0300. - Upstream source: zero.
tools/vmaf-tune/is fork-introduced (Phase A under ADR-0237). - On upstream sync: zero interaction. Upstream Netflix/vmaf ships no encode automation surface; this tree is entirely fork-local and lives outside
libvmaf/andpython/. - Schema migration note:
SCHEMA_VERSIONbumped 1 → 2. The three new keys are additive — Phase B / C loaders treat missing keys as SDR for backward compat with v1 rows. - Re-test on rebase:
python -m pytest tools/vmaf-tune/tests/ -q
python -m vmaftune.cli corpus --help # confirm --auto-hdr surfaces
0298 — vmaf-tune content-addressed cache (ADR-0298)¶
- What changed: fork-local. New module
tools/vmaf-tune/src/vmaftune/cache.py; cache integration intools/vmaf-tune/src/vmaftune/corpus.py(iter_rowsnow consults the cache before encode/score); new CLI flags--no-cache,--cache-dir,--cache-size-gbincli.py. Codec-adapterProtocolgainsadapter_version: str; the lone Phase-A x264 adapter pins"1". - Upstream source: none.
tools/vmaf-tune/is fork-introduced (ADR-0237) and has no upstream counterpart. - On upstream sync: zero interaction with Netflix/vmaf master. The module sits entirely under
tools/vmaf-tune/, which upstream does not ship. - Invariant for future codec adapters: every
CodecAdaptermust declareadapter_version: str. Bump it whenever the adapter's argv shape, preset list, or quality range changes — otherwise the cache returns stale results post-upgrade. The contract is asserted bytest_cache_key_diffs_on_each_fieldintests/test_cache.py. - Re-test on rebase:
```bash pytest tools/vmaf-tune/tests/test_cache.py -v
0283 — vmaf-tune Apple VideoToolbox adapters (2026-05-05)¶
- What changed: fork-local addition under
tools/vmaf-tune/src/vmaftune/codec_adapters/. New files:h264_videotoolbox.py,hevc_videotoolbox.py,_videotoolbox_common.py, plus the registry hook in__init__.py. See ADR-0283. - Update 2026-05-09:
prores_videotoolbox.pyadapter added to the same registry pattern (broadcast / prosumer ProRes intermediate). Quality knob differs — ProRes is a fixed-rate codec, so the harness's--crfslot carries the integer ProRes tier id (0=proxy→ 5=xq) rather than a-q:vvalue._videotoolbox_common.pyextended withPRORES_PROFILE_*constants +validate_prores_videotoolbox()/prores_profile_name()helpers; profile ids verified against FFmpeg n8.1.1libavcodec/videotoolboxenc.c. See the Status update appendix in ADR-0283. - Upstream source: zero.
tools/vmaf-tune/is fork-introduced (Phase A under ADR-0237). - On upstream sync: zero interaction.
- Re-test on rebase:
python -m pytest tools/vmaf-tune/tests/test_codec_adapter_videotoolbox.py -q
python -m pytest tools/vmaf-tune/tests/test_codec_adapter_prores_videotoolbox.py -q
0228 — vmaf-tune coarse-to-fine CRF search (ADR-0306)¶
- What changed: fork-local tooling. Adds
coarse_to_fine_search()totools/vmaf-tune/src/vmaftune/corpus.py, plumbs new CLI flags ontovmaf-tune corpus(--coarse-to-fine,--coarse-step,--fine-radius,--fine-step,--target-vmaf), and ships a newvmaf-tune recommendsubcommand. Widenstools/vmaf-tune/src/vmaftune/codec_adapters/x264.pyquality_rangefrom(15, 40)to(0, 51). JSONL row schema unchanged (SCHEMA_VERSION=1). - Upstream source: fork-local. The whole
tools/vmaf-tune/tree is fork-introduced (ADR-0237); upstream Netflix/vmaf has no encode-automation surface. - On upstream sync: zero interaction.
tools/vmaf-tune/is not mirrored from upstream. - Re-test on rebase:
0314 — vmaf-tune --score-backend=vulkan (ADR-0314)¶
- Touches:
tools/vmaf-tune/src/vmaftune/cli.py(additive argparse flag oncorpus+recommendsubparsers; resolvesselect_backendand catchesBackendUnavailableErrorfor clean exit-2).tools/vmaf-tune/src/vmaftune/score.py(additivebackendkwarg onbuild_vmaf_commandandrun_score;None= no flag emitted).tools/vmaf-tune/src/vmaftune/corpus.py(newCorpusOptions.score_backendfield, defaultNone; forwarded intorun_score).tools/vmaf-tune/tests/test_score_backend.py(additive Vulkan-specific tests; pre-existing tests now pass after thebackend=kwarg lands).docs/adr/0314-vmaf-tune-score-backend-vulkan.md(new).docs/usage/vmaf-tune.md(new "Vulkan score backend" subsection under the existing GPU-scoring section).tools/vmaf-tune/AGENTS.md(invariant note: argparse choices stay in sync with libvmaf--backendvocabulary).changelog.d/added/vmaf-tune-score-backend-vulkan.md(new).- Invariant:
score_backend.ALL_BACKENDS = ("cpu", "cuda", "sycl", "vulkan")is the exact set libvmaf'score/tools/cli_parse.c--backendalternation accepts. Adding a new harness-side value without the libvmaf-side wiring produces silent strict-mode failures on hosts that probe positively for it. - Upstream source: zero. Netflix upstream's CLI does not ship a
--backendselector; bothtools/vmaf-tune/andcore/src/vulkan/are fork-introduced. - On upstream sync: zero interaction. No upstream-mirror file is touched.
- Re-test on rebase:
pytest tools/vmaf-tune/tests/test_score_backend.py -v -k vulkan
pytest tools/vmaf-tune/tests/test_score_backend.py -v
Failures here usually indicate the libvmaf help-text format changed; score_backend.parse_supported_backends test fixtures pin the format and will fail loudly.
0303 — fr_regressor_v2 ensemble prod flip (ADR-0303)¶
- ADR: ADR-0303
- Touches: entirely fork-local.
ai/scripts/train_fr_regressor_v2_ensemble_loso.py(new — 9-fold LOSO trainer over the five ensemble seeds; emitsloso_seed{N}.jsonartefacts).scripts/ci/ensemble_prod_gate.py(new — reads fiveloso_seed{N}.jsonfiles, returns exit 0 iffmean(PLCC_i) ≥ 0.95ANDmax - min ≤ 0.005).ai/AGENTS.md— appended "Ensemble registry invariant" paragraph under the existingfr_regressor_v2_ensemble_v1section.docs/adr/0303-fr-regressor-v2-ensemble-prod-flip.md(new),docs/research/0075-fr-regressor-v2-ensemble-prod-flip.md(new),changelog.d/added/fr-regressor-v2-ensemble-prod-flip.md(new).- Rebase invariant: the production ship gate is two-part —
mean_i(PLCC_i) ≥ 0.95ANDmax_i(PLCC_i) - min_i(PLCC_i) ≤ 0.005over five seeds. The variance bound is load-bearing: removing it silently allows a one-seed-wins-four-seeds-tie configuration that invalidates the ensemble's predictive-distribution semantics. Both thresholds live inscripts/ci/ensemble_prod_gate.py; do not weaken either without superseding ADR-0303. - Rebase invariant (registry): the five
fr_regressor_v2_ensemble_v1_seed{0..4}registry rows aresmoke: trueon master at this commit; flipping them tofalseis the follow-up flip PR's job, gated on a real-corpus LOSO run + the CI gate. Do not flip seed rows during a rebase merge conflict resolution. - Re-test on rebase:
python3 -c "import ast; ast.parse(open('ai/scripts/train_fr_regressor_v2_ensemble_loso.py').read())"
python3 -c "import ast; ast.parse(open('scripts/ci/ensemble_prod_gate.py').read())"
python ai/scripts/train_fr_regressor_v2_ensemble_loso.py --help
python scripts/ci/ensemble_prod_gate.py --help
- Upstream source: zero.
fr_regressor_v2and its ensemble are fork-introduced (parent ADR-0272 / ADR-0279). - On upstream sync: zero interaction.
0313 — CI required-checks aggregator (2026-05-05)¶
- What changed: fork-local CI policy. New
.github/workflows/required-aggregator.yml— single workflow that runs on every non-draft PR and verifies the 23 named required checks reportedsuccess/skipped/neutral(or didn't appear at all, which is the path-filter-rejection semantics). Aggregator becomes the single branch-protection required check, replacing the 23-name list from ADR-0037. - Touches:
.github/workflows/required-aggregator.yml(new),docs/adr/0313-ci-required-checks-aggregator.md(new),changelog.d/added/ci-required-checks-aggregator.md(new),docs/adr/README.md(+1 row),docs/adr/_index_fragments/_order.txt(+1 line + new fragment file). - Upstream source: zero. Branch-protection policy is fork-only.
- On upstream sync: zero interaction with Netflix/vmaf master.
- Manual operator step at adoption (uses PATCH, not PUT — corrected from the original ADR-0313 body which had the wrong verb):
echo '{"strict": false, "contexts": ["Required Checks Aggregator"]}' | \
gh api -X PATCH "repos/VMAFx/vmafx/branches/master/protection/required_status_checks" --input -
- Re-test on rebase:
# YAML lint passes
python3 -c "import yaml; yaml.safe_load(open('.github/workflows/required-aggregator.yml'))"
0305 — encoder knob-space Pareto analysis (2026-05-05)¶
- What changed: fork-local. New analysis scaffold for the 12,636-cell encoder knob sweep that backs
tools/vmaf-tune/codec_adapters/*recipe defaults. New files:ai/scripts/analyze_knob_sweep.py(per-(source, codec, rc_mode)Pareto hull on(bitrate_kbps, vmaf_score),encode_time_mstiebreaker, regression-detection check),ai/tests/test_knob_sweep_analysis.py(synthetic 20-row JSONL fixture). Methodology + scaffolded findings: see ADR-0305 + Research-0077. Companion to Research-0063. - Touches: none upstream-shared. Sits entirely under
ai/(fork-local since the tiny-AI training surface, ADR-0021) anddocs/{adr,research}/(fork ledger). - Upstream source: zero. The 12,636-cell sweep, the Pareto scaffold, and the regression-detection invariant are fork-introduced; Netflix/vmaf master ships no encoder knob-sweep tooling.
- On upstream sync: zero interaction with Netflix/vmaf master.
- Invariant for future codec adapter PRs: per the
ai/AGENTS.mdknob-sweep corpus invariant (ADR-0305), recipes that regress vs the bare encoder at matched bitrate within the same(source, codec, rc_mode)slice MUST NOT ship as adapter defaults. New adapter PRs cite the per-slice hull row fromreports/summary.md(or "no hull entry yet — bare default") in their PR description. Thecomprehensive.jsonlsweep file is generated locally and lives underruns/phase_a/full_grid/(gitignored — never committed). - Re-test on rebase:
0302 — ENCODER_VOCAB v3 schema expansion (ADR-0302)¶
- Touches:
ai/scripts/train_fr_regressor_v2.py(adds anENCODER_VOCAB_V3parallel constant; does not modify the liveENCODER_VOCABorENCODER_VOCAB_VERSION). - Invariant:
ENCODER_VOCABis append-only and order-stable (per ADR-0235). The v3 scaffold preserves the v2 slot ordering verbatim — slots 0..12 are bit-identical to the v2 vocab; slots 13/14/15 appendlibsvtav1,h264_videotoolbox,hevc_videotoolbox. The liveENCODER_VOCAB_VERSION = 2remains the source of truth until the follow-up retrain PR clears the LOSO PLCC ship gate. - Upstream interaction: zero.
ai/scripts/train_fr_regressor_v2.pyis fork-introduced (ADR-0272) and has no upstream counterpart. - Re-test on rebase:
python3 -c "
import importlib.util, pathlib
spec = importlib.util.spec_from_file_location(
't', pathlib.Path('ai/scripts/train_fr_regressor_v2.py')
)
m = importlib.util.module_from_spec(spec)
spec.loader.exec_module(m)
assert len(m.ENCODER_VOCAB_V3) == 16
assert m.ENCODER_VOCAB_VERSION == 2
print('OK')
"
0304 — vmaf-tune fast-path prod wiring (ADR-0304)¶
- Touches:
tools/vmaf-tune/src/vmaftune/fast.py(replaces the ADR-0276 scaffold'sNotImplementedErrorpaths with concrete Optuna TPE + v2 proxy + GPU verify wiring); new moduletools/vmaf-tune/src/vmaftune/proxy.py(centralised seam forfr_regressor_v2ONNX inference); expandedtools/vmaf-tune/tests/test_fast.py. Doc-side: ADR-0304, Research-0076,tools/vmaf-tune/AGENTS.mdinvariant note. - Upstream source: zero.
tools/vmaf-tune/andmodel/tiny/fr_regressor_v2.onnxare both fork-introduced (ADR-0237 / ADR-0352). - Invariant: the production proxy is always
fr_regressor_v2(no smoke models in the production path) and a single GPU verify pass at recommend-end is mandatory — proxy alone never wins. Thevmaftune.proxy.run_proxyhelper is the single seam every fast-path consumer goes through; future probabilistic-head / ensemble migrations land in that one module. ENCODER_VOCAB v2 one-hot ordering is frozen by ADR-0352 and pinned inproxy.ENCODER_VOCAB_V2— keep in sync withai/scripts/train_fr_regressor_v2.py; drift raisesProxyErrorat inference time before bad predictions ship. - On upstream sync: zero interaction with Netflix/vmaf master.
- Re-test on rebase:
0307 — vmaf-tune ladder default sampler wiring (ADR-0307)¶
- What changed: fork-local tooling.
tools/vmaf-tune/src/vmaftune/ladder.py::_default_samplerno longer raisesNotImplementedError; it composescorpus.iter_rows(Phase A encode + score) withrecommend.pick_target_vmaf(smallest CRF clearing target VMAF) overDEFAULT_SAMPLER_CRF_SWEEP = (18, 23, 28, 33, 38)at the adapter's mid-range preset. Module-level docstring + AGENTS.md invariant updated. New tests intools/vmaf-tune/tests/test_ladder.pystubiter_rowsviamonkeypatch.setattrso no live ffmpeg / vmaf binaries are needed. - Upstream source: fork-local. The whole
tools/vmaf-tune/tree is fork-introduced (ADR-0237); upstream Netflix/vmaf has no encode-automation / ladder surface. - On upstream sync: zero interaction.
tools/vmaf-tune/is not mirrored from upstream. - Rebase invariant: the 5-point sweep
(18, 23, 28, 33, 38)is the load-bearing default; downstream Phase E callers size their wall-time budget against five encodes per(resolution, target_vmaf)cell. Do not widen / narrow it without an ADR-0307 follow-up. TheSamplerFnseam stays open — callers needing finer grids pass an explicitsampler=. - Re-test on rebase:
0309 — fr_regressor_v2 ensemble real-corpus retrain harness (ADR-0309)¶
- ADR: ADR-0309
- Touches: entirely fork-local.
ai/scripts/run_ensemble_v2_real_corpus_loso.sh(new — Bash wrapper that loops the five seeds over the existingtrain_fr_regressor_v2_ensemble_loso.pyagainst.workingdir2/netflix/).ai/scripts/validate_ensemble_seeds.py(new — calls the ADR-0303 gate and writesPROMOTE.json/HOLD.jsonwith a corpus sha256 snapshot).ai/tests/test_validate_ensemble_seeds.py(new — 7 tests, synthetic JSON fixtures for both verdict paths).ai/AGENTS.md— appended "Registry-flip is a separate PR (ADR-0309)" paragraph under the existingfr_regressor_v2_ensemble_v1section.docs/adr/0309-fr-regressor-v2-ensemble-real-corpus-retrain.md,docs/research/0081-fr-regressor-v2-ensemble-real-corpus-methodology.md,docs/ai/ensemble-v2-real-corpus-retrain-runbook.md(all new).- Rebase invariant: the harness is decoupled from the registry mutation. Neither the wrapper nor the validator touches
model/tiny/registry.json; the registry flip is a separate follow-up PR gated on a passingPROMOTE.json. Auto-flipping on PROMOTE was rejected in ADR-0309's alternatives matrix specifically because rebase-time mutation of shipped registry rows is the foot-gun this invariant exists to prevent. - Re-test on rebase:
python -m pytest ai/tests/test_validate_ensemble_seeds.py -v
python ai/scripts/validate_ensemble_seeds.py --help
bash -n ai/scripts/run_ensemble_v2_real_corpus_loso.sh
- Upstream source: zero.
- On upstream sync: zero interaction.
0310 — BVI-DVC corpus ingestion for fr_regressor_v2 (ADR-0310)¶
- Touches:
ai/scripts/bvi_dvc_to_corpus_jsonl.py(new fork-only adapter),ai/scripts/merge_corpora.py(new fork-only shard merger),ai/tests/test_merge_corpora.py(new),docs/ai/bvi-dvc-corpus-ingestion.md(new),docs/adr/0310-bvi-dvc-corpus-ingestion.md(new),docs/research/0082-bvi-dvc-corpus-feasibility.md(new),ai/AGENTS.md(BVI-DVC invariant note). - Invariant: the BVI-DVC archive and any extracted artefacts (parquet, cached libvmaf JSON, JSONL corpus shard) are research-only and stay local — only derived
fr_regressor_v2_*.onnxweights ship. The merge utility validates every row against the canonicalvmaftune.CORPUS_ROW_KEYStuple; the schema is the merge contract. Re-shape here is a pure transform on the cached libvmaf JSON; no ffmpeg / vmaf binary is invoked. The(src_sha256, encoder, preset, crf)natural key is load-bearing for de-duplication across mirrors and re-encodes. - Upstream interaction: none.
ai/is fork-introduced; BVI-DVC is not part of Netflix/vmaf upstream. - Re-test on rebase:
ADR-0312 — ffmpeg-patches/ vmaf-tune integration (2026-05-05)¶
- Files:
ffmpeg-patches/0007-libvmaf-tune-qpfile-unified.patch,ffmpeg-patches/0008-add-libvmaf_tune-filter.patch,ffmpeg-patches/0009-pass-autotune-cli-glue.patch,ffmpeg-patches/series.txt,ffmpeg-patches/README.md. - Rebase invariant: patches
0007–0009plug into the cumulative state after patches0001–0006apply against pristinen8.1. Per-patchgit apply --checkin isolation is the wrong gate; use the series-replay command in CLAUDE.md §12 r14 instead. - vmaf-tune patch invariant: the qpfile parser at
libavcodec/qpfile_parser.{c,h}is shared across all three encoder adapters in patch 0007. Future encoders that grow a-qpfileAVOption inherit it; do not fork the parser. Whentools/vmaf-tune/src/vmaftune/saliency.py's qpfile output format changes (new column, different frame-type alphabet, …), patch 0007 must change in the same PR (CLAUDE.md §12 r14). - vf_libvmaf_tune full-scoring promotion (2026-05-06): patch 0008 originally shipped as a scaffold (linear CRF↔VMAF interpolation, no libvmaf scoring) per ADR-0312's deferred-alternatives column. The filter now mirrors
vf_libvmaf.c's CPU framesync pipeline end-to-end (vmaf_init+vmaf_model_load+vmaf_use_features_from_modelin init(); per-framevmaf_picture_alloc+ memcpy +vmaf_read_pictures; flush +vmaf_score_pooled(MEAN)in uninit()). The CRF recommendation remains a piece-wise linear projection from the observed VMAF; per-clip Optuna TPE search stays intools/vmaf-tune/src/vmaftune/recommend.py. Rebase-side: the new filter still depends only on libvmaf's CPU C-API (vmaf_init,vmaf_model_load,vmaf_use_features_from_model,vmaf_read_pictures,vmaf_score_pooled,vmaf_close,vmaf_picture_alloc/unref); zero new symbols beyond whatvf_libvmaf.calready requires, so future libvmaf rebases that pass the existing libvmaf filter pass this one too. ADR-0312 sub-decision retired. - n7+ API migration (2026-05-06): patch 0008 originally referenced the removed
AVFilterLink::frame_ratemember directly (n6-era API); in n7+ that field moved offAVFilterLinkonto a newFilterLinkstruct accessed viaff_filter_link(AVFilterLink *)fromlibavfilter/filters.h. Patch 0008 now usesff_filter_link(outlink)->frame_rate = ff_filter_link(mainlink)->frame_rate;inconfig_output(), mirroring patches 0005/0006 which were already written against the post-n7 API. The bug slipped through CI because the FFmpeg-Vulkan lane only buildsvf_libvmaf.o, notvf_libvmaf_tune.c; the full SYCL lane catches it now that PR #415 addedffmpeg-patches/**to the integration workflow's path filter. Discovery: PR #415 / ADR-0317. - Upstream source: zero. The vmaf-tune integration is fork-introduced; pure upstream syncs are unaffected.
- On upstream sync: zero interaction with libvmaf master. FFmpeg-side rebases when n8.1 → n8.x land in
ffmpeg-patches/test/build-and-run.sh'sFFMPEG_SHAare tracked separately under each refresh ADR (e.g., ADR-0277 for the 2026-05-04 refresh). - Re-test on rebase:
git -C /path/to/ffmpeg-8 reset --hard n8.1
for p in ffmpeg-patches/000*-*.patch; do
git -C /path/to/ffmpeg-8 am --3way "$p" || break
done
# Build smoke (libvmaf-disabled — patches 0001–0006 skipped if libvmaf_dnn
# is not built). With libvmaf_dnn available:
cd /path/to/ffmpeg-8 && ./configure --enable-libvmaf --enable-libx264 --enable-libsvtav1 --enable-libaom --enable-gpl
make -j$(nproc) ffmpeg
./ffmpeg -hide_banner -h encoder=libx264 2>&1 | grep -i qpfile
- 2026-05-06 update — patch 0007 SVT-AV1 ROI bridge promoted from scaffold to full impl: the libsvtav1 hunk now sets
enc_params.enable_roi_map = true, builds oneSvtAv1RoiMapEvtper qpfile frame upfront ineb_enc_init(per-MB qp_offsets averaged into per-64×64-SBb64_seg_mapof up to 8 segment QPs; uniform binning when the value span exceeds the segment budget), and attaches each event as aROI_MAP_EVENTpriv-data node fromeb_send_frame()withnode->size = sizeof(SvtAv1RoiMapEvt*)(the validation contract enforced by SVT-AV1'sresource_coordination_process.c). Lifetime invariant: events + maps live for the entire encode session because SVT-AV1 reads ROI_MAP_EVENT data via shallow-copied pointers on async pipeline threads (perenc_handle.c::copy_private_data_list);eb_enc_closefrees them. Wiring is gated onSVT_AV1_CHECK_VERSION(1, 6, 0); older SVT-AV1 builds keep the log-and-continue fallback. libaom remains scaffold-only — itsAOME_SET_ROI_MAPbridge stays a separate follow-up. No new ADR per CLAUDE.md §12 r8 (executes the existing ADR-0312 decision). - 2026-05-06 update — patch 0007 libaom-av1 ROI bridge promoted from scaffold to full impl: the libaom-av1 hunk now caches the parsed
VmafTuneQpFileinAOMContext, allocates a segment-id map at libaom's mode-info grid (ALIGN_POWER_OF_TWO(dim, 8) >> 2, sinceav1/common/enums.h::MI_SIZE == 4), and on every encoded frame picks up to 8 segment QPs from the per-frame qp_offset value range (uniform linear binning when the span exceedsAOM_MAX_SEGMENTS == 8), paints the per-mi segment map by expanding each per-16×16-MB qp_offset into a 4×4 block of mi cells, and issuesaom_codec_control(&ctx->encoder, AOME_SET_ROI_MAP, &roi_map). Lifetime invariant: libaom deep-copies the segment map anddelta_q[]table on every control call (perav1/encoder/encoder.c::av1_set_roi_map memcpy), so a single buffer is reused across frames and freed inaom_free(). The qpfile is also freed there. Trade-off: the 8-segment cap rounds nearby qp_offsets together when the saliency model emits more than 8 distinct values per frame; finer granularity requiresvmaf-tune corpusinstead. This retires the libaom-av1 deferral noted under ADR-0312 — both AV1 encoder hooks (libsvtav1 and libaom-av1) are now full-impl. No new ADR per CLAUDE.md §12 r8 (executes the existing ADR-0312 decision).
0315 — Vendor-neutral VVC encode strategy (ADR-0315 / Research-0085)¶
- ADR: ADR-0315
- Digest: Research-0085
- Touches: docs-only.
docs/research/0085-vendor-neutral-vvc-encode-landscape.md(new).docs/adr/0315-vendor-neutral-vvc-encode-strategy.md(new).docs/adr/_index_fragments/0315-vendor-neutral-vvc-encode-strategy.md(new).docs/adr/_index_fragments/_order.txt(one-line append).changelog.d/added/research-0085-vendor-neutral-vvc-encode.md(new).docs/rebase-notes.md(this entry).- Rebase invariant: none. The research digest and ADR are pure surveys with no code dependencies; nothing in the fork's source tree references them in a way that breaks on upstream rebase.
- Upstream source: zero. VVC encode strategy is a fork-local decision; upstream Netflix/vmaf has no codec adapter or encode-automation surface.
- On upstream sync: zero interaction. Pure docs.
- Re-test on rebase:
- 2026-05-06 follow-up (Research-0085 verification pass):
docs/research/0085-vendor-neutral-vvc-encode-landscape.mdflipped fromStatus: SKELETONtoStatus: Active. Most[UNVERIFIED]claims are now backed by primary-source URLs (NVIDIA SDK 13.0 docs, AMD AMF GitHub, Intel oneVPL GitHub +mfxstructures.h+CHANGELOG.md, Khronos registry, Phoronix Mesa/RADV coverage, VVenC issue tracker, ZLUDA repo).- ADR-0315's
## Contextand## Alternatives consideredrefreshed with the verified data points. Status staysProposed. [UNVERIFIED]count in the digest dropped 25 → 10; remaining items are legitimate gaps (NN-VC quality lift, vvenc per-kernel profile, HHI's non-public roadmap).- No code touched. No rebase impact beyond the existing docs-only posture.
0316 — cli_parse.c error() long-only-option fix (ADR-0316)¶
- ADR: ADR-0316 (follow-up to ADR-0311).
- Digest: none — bug-fix; fix shape fits in the ADR/commit body.
- Touches:
core/tools/cli_parse.c(3 lines — call-site arg change at theARG_THREADS/ARG_SUBSAMPLE/ARG_CPUMASKhandlers).core/test/fuzz/fuzz_cli_parse.c(removedknown_assert_in_inputearly-reject filter).core/test/fuzz/cli_parse_corpus/cli_threads_abbrev_assert.argv(promoted fromcli_parse_known_crashes/).core/test/test_cli_parse_long_only_args.c(new fork()-based regression test).core/test/meson.build(new test wiring, gated off Windows alongsidetest_y4m_411_oob).core/tools/AGENTS.md(added a long-only-options invariant note next to the existingcli_parse.crules).- Rebase invariant: load-bearing.
cli_parse.cis upstream-mirror with fork additions; the three handlers carry the fork-local shape of passing theARG_*enum value (not't'/'s'/'c') toparse_unsigned(). If an upstream sync re-introduces the original short-option char shape, the assert returns and the parked-then-promoted reproducer (cli_parse_corpus/cli_threads_abbrev_assert.argv) will surface it in the next nightly fuzz run. - Upstream source: the bug shape exists in Netflix/vmaf master too (long-only options were added upstream with the same short-option-char placeholder). When the fork ports an upstream fix that overlaps these handlers, prefer the
parse_unsigned(optarg, ARG_*, argv[0])form already on the fork. - On upstream sync: re-apply the three-line change in
cli_parse.cif upstream resets the call-site args. The unit test is fork-local and stays. - Re-test on rebase:
meson setup core/build libvmaf -Denable_tests=true \
-Denable_cuda=false -Denable_sycl=false
ninja -C core/build test/test_cli_parse_long_only_args
meson test -C core/build test_cli_parse_long_only_args -v
ADR-0317 — CI flake fix: doc-only PR path-filter (2026-05-06)¶
- Touched files:
.github/workflows/docker-image.yml— addedpaths:filter on bothpush:andpull_request:triggers..github/workflows/ffmpeg-integration.yml— addedpaths:filter on bothpush:andpull_request:triggers (covers all four matrix lanes: gcc, clang, SYCL, Vulkan).docs/adr/0317-ci-doc-only-pr-flake-fix.md,docs/adr/README.md(index row),changelog.d/fixed/ci-doc-only-pr-flakes.md.- Rebase invariant: not load-bearing. Workflow-only change. Both files are fork-local CI; upstream Netflix/vmaf does not ship a Docker workflow or an FFmpeg-integration matrix in this shape, so rebase conflicts are unlikely. If a future upstream sync introduces an overlapping
docker-image.ymlor FFmpeg matrix, prefer the fork's path-filtered form — the rationale (ADR-0313 aggregator posture, doc-only-PR runner-time burn) is fork-specific. - Upstream source: none — fork-local CI workflows.
- On upstream sync: no action required. If reviewers later add new build inputs (e.g. a top-level
docker-compose.yml, a newffmpeg-patches/*.txtconfig file), extend thepaths:lists in the same PR that adds the input. - Follow-up not in this ADR: patch
ffmpeg-patches/0008-add-libvmaf_tune-filter.patchline 256 (outlink->frame_rate = mainlink->frame_rate;) needs to migrate to theff_filter_link()accessor introduced in FFmpeg n7+, matching the pattern already in patches 0005 / 0006. Tracked separately; the path-filter does not hide it (any libvmaf/ or ffmpeg-patches/ PR will still trip the SYCL lane). - Re-test on rebase:
python3 -c "import yaml; \
yaml.safe_load(open('.github/workflows/docker-image.yml')); \
yaml.safe_load(open('.github/workflows/ffmpeg-integration.yml')); \
print('OK')"
0319 — fr_regressor_v2 ensemble LOSO trainer — real loader + per-fold training (ADR-0319)¶
- Touches:
ai/scripts/train_fr_regressor_v2_ensemble_loso.py(real_load_corpus+_train_one_seedbodies),ai/scripts/run_ensemble_v2_real_corpus_loso.sh(wrapper argv fix),docs/ai/ensemble-v2-real-corpus-retrain-runbook.md(Step 0 corpus-generation section),ai/AGENTS.md(canonical-6 schema invariant note),ai/tests/test_train_fr_regressor_v2_ensemble_loso_*.py(loader + train schema tests). Closes the deferrals tracked in rebase-notes §0303 + §0309. - Upstream source: none — fork-local ML training infrastructure. Netflix/vmaf upstream has no
fr_regressor_v2surface, no LOSO trainer, and no canonical-6 corpus tooling. - Invariant: the trainer's
_load_corpusaccepts the canonical-6 JSONL schema emitted byscripts/dev/hw_encoder_corpus.pybit-for-bit — required keys per row are(src, encoder, cq, frame_index, vmaf, adm2, vif_scale0..3, motion2). Codec block layout is 12-slotENCODER_VOCABv2 one-hot + constantpreset_norm = 0.5+crf_norm = (cq - cq_min) / (cq_max - cq_min). Schema changes require anENCODER_VOCAB_VERSIONbump and full ensemble retrain per the existing closed-vocabulary rule (ADR-0235 / ADR-0352). Fold-level StandardScaler is fit on the training rows only; leaking the held-out source's distribution into the scaler would silently inflate per-fold PLCC. - On upstream sync: no action required. If upstream Netflix/vmaf ever adds a competing LOSO trainer under
python/vmaf/, do NOT merge them — keep the fork's training stack underai/per the AGENTS.md scope rule. - Re-test on rebase:
pytest ai/tests/test_train_fr_regressor_v2_ensemble_loso_loader.py \
ai/tests/test_train_fr_regressor_v2_ensemble_loso_train.py -v
bash -n ai/scripts/run_ensemble_v2_real_corpus_loso.sh
ADR-0323 — fr_regressor_v3 train + register on ENCODER_VOCAB v3 (2026-05-06)¶
- Scope:
ai/scripts/train_fr_regressor_v3.py(new),ai/tests/test_train_fr_regressor_v3.py(new),model/tiny/fr_regressor_v3.onnx(new, real-weight checkpoint from a 9-fold LOSO gate-pass at mean PLCC 0.9975),model/tiny/fr_regressor_v3.json(new sidecar withencoder_vocab_version: 3and full per-fold trace),model/tiny/registry.json(newfr_regressor_v3row,smoke: false),ai/AGENTS.md(v3 retrain invariant section gains a "Status" subsection recording the gate result),docs/ai/models/fr_regressor_v3.md(new model card),docs/adr/0323-fr-regressor-v3-train-and-register.md+ index row,changelog.d/added/fr-regressor-v3-train-register.md. - Rebase impact: zero. Fork-local feature; no upstream Netflix/vmaf surface is touched. The 16-slot
ENCODER_VOCAB_V3imported fromtrain_fr_regressor_v2.pywas already landed by PR #401 (ADR-0302). - On upstream sync: no action required. The v3 model ships alongside v2 —
fr_regressor_v2.onnxand its sidecar are unchanged; the v3 row is appended to the registry and sorted alphabetically. If a future upstream sync ever lands a competingfr_regressor_v3model underpython/vmaf/, do NOT cross-link them — the fork's training stack lives underai/. - Watch out for: the live
ENCODER_VOCAB_VERSIONinai/scripts/train_fr_regressor_v2.pystays at 2 (per ADR-0302's invariant). Do not bump it to 3 in this PR or in any downstream port; the in-place promotion of v3 over v2 is a separate "promote v3 to authoritative" PR per ADR-0302's production-flip checklist. - Re-test on rebase:
pytest ai/tests/test_train_fr_regressor_v3.py -v
bash core/test/dnn/test_registry.sh # must report OK: 20+
python -c "import onnx; onnx.checker.check_model(onnx.load('model/tiny/fr_regressor_v3.onnx')); print('OK')"
ADR-0321 — fr_regressor_v2_ensemble_v1 full production flip (2026-05-06)¶
- Scope:
ai/scripts/export_ensemble_v2_seeds.py(new),model/tiny/fr_regressor_v2_ensemble_v1_seed{0..4}.onnx(real full-corpus-trained weights replacing the 3025-byte synthetic scaffold bytes),model/tiny/fr_regressor_v2_ensemble_v1_seed{0..4}.json(new per-seed sidecars),model/tiny/registry.json(sha256 +smoke: falseon the five seed rows),ai/AGENTS.md(new invariant: the registry-flip is now done; future re-flips require a fresh PROMOTE.json + re-run of the export driver). - Rebase impact: zero. This is a fork-local production-flip; no upstream Netflix/vmaf surface is touched. The 12-slot
ENCODER_VOCABv2 carried in each sidecar is the same one the LOSO trainer (ADR-0319) bakes into the codec-block layout, so there is no rebase-time vocabulary drift to worry about. - Watch out for: if a future upstream sync ever introduces a competing
fr_regressor_v2_ensemble_*model underpython/vmaf/, do NOT cross-link them — the fork's ensemble weights are gated onruns/ensemble_v2_real/PROMOTE.jsonand are not portable to a different training stack. - Re-test on rebase:
bash core/test/dnn/test_registry.sh # must report OK: 19
python -c "import onnx; \
[onnx.checker.check_model(onnx.load(f'model/tiny/fr_regressor_v2_ensemble_v1_seed{i}.onnx')) \
for i in range(5)]; print('OK')"
ADR-0324 — Ensemble training kit (2026-05-06)¶
- Touches:
tools/ensemble-training-kit/(new),docs/adr/0324-ensemble-training-kit.md(new),docs/adr/README.md(index row),changelog.d/added/0324-ensemble-training-kit.md(new). No engine code touched; no upstream-shared paths. - Invariant: the kit assumes the LOSO wrapper hard-codes seeds
(0 1 2 3 4). The orchestrator surfaces a warning if--seedsdeviates but still hands off to the wrapper. If a future PR parameterises the wrapper's seed list, update both the wrapper and the kit's pass-through logic in lockstep. - On upstream sync: no action required. The kit lives entirely under
tools/ensemble-training-kit/(a fork-local path) and only invokes other fork-local scripts (ai/scripts/,scripts/dev/,scripts/ci/). - Re-test on rebase:
bash -n tools/ensemble-training-kit/*.sh
bash tools/ensemble-training-kit/make-distribution-tarball.sh /tmp/kit-test.tar.gz
tar -tzf /tmp/kit-test.tar.gz | grep -q "tools/ensemble-training-kit/run-full-pipeline.sh"
ADR-0335 — Hardware-capability priors (2026-05-08)¶
- Touches:
ai/data/hardware_caps.csv(new),ai/scripts/hardware_caps_loader.py(new),ai/tests/test_hardware_caps.py(new),ai/AGENTS.md(one new bullet under "Rebase-sensitive invariants"),docs/ai/hardware-capability-priors.md(new),docs/research/0088-hardware-capability-priors-2026-05-08.md(new),docs/adr/0335-hardware-capability-priors.md(new),docs/adr/_index_fragments/0335-hardware-capability-priors.md(new),docs/adr/_index_fragments/_order.txt(one-line append),CHANGELOG.md(Added bullet under[Unreleased] — lusoris fork). No upstream-shared paths. - Invariant: the table is prior-only. The schema check in
hardware_caps_loader.pyrejects benchmark-shaped header columns (fps_*,throughput,mbps,latency,watts,tdp,score_*,vmaf_*), community-wiki source URLs (wikipedia.org,wikichip.org), empty fields, and rows withencoding_blocks=0. Adding throughput / quality columns is forbidden — that pathology was the contributor-pack digest's category-1 NO-GO finding. Schema extensions need a new ADR, not a silent column bump. Thecap_vector_for()return-dict shape is load-bearing: trainers / corpus writers consumehwcap_*columns by name; reordering or renaming silently breaks downstream parquet schemas. - On upstream sync: no action required. The whole surface lives under
ai/anddocs/— Netflix upstream has no equivalent. - Re-test on rebase:
python -m pytest ai/tests/test_hardware_caps.py -v # must report 23 passed
python ai/scripts/hardware_caps_loader.py # JSON dump, 6+ rows
ADR-0332 — External-competitor benchmark harness (2026-05-08)¶
- Touches:
tools/external-bench/(new),docs/adr/0332-external-bench-wrapper-only.md(new),docs/adr/_index_fragments/0332-external-bench-wrapper-only.md(new),docs/adr/_index_fragments/_order.txt(one-line append),docs/adr/README.md(regenerated),changelog.d/added/external-bench-harness.md(new),docs/research/0087-external-bench-competitor-survey-2026-05-08.md(new). No engine code touched; no upstream-shared paths. - Invariant: the harness is wrapper-only — never vendor or link
x264-pVMAF(GPL-2.0) into this fork. Future competitors follow the same pattern (tools/external-bench/<competitor>/run.shinvokes a user-installed binary via env var; output schema-shimmed into the canonical JSON shape). The output schema (frames[].{frame_idx, predicted_vmaf_or_mos, runtime_ms}+summary.{competitor, plcc, srocc, rmse, runtime_total_ms, params, gflops}) is the contract between every wrapper andcompare.py.run_wrapper'srunnerparameter MUST stay resolved at call time (not via default-arg binding) so monkeypatch-based tests work. - On upstream sync: no action required. The harness lives entirely under
tools/external-bench/(a fork-local path) and never touches Netflix-shared code. - Re-test on rebase:
python3 -m pytest tools/external-bench/tests/ -q # must report 7 passed
bash -n tools/external-bench/*/run.sh
0327 — Conformal-VQA prediction surface for vmaf-tune (ADR-0279)¶
- Touches:
tools/vmaf-tune/src/vmaftune/conformal.py(new),tools/vmaf-tune/src/vmaftune/predictor.py(Predictor.predict_vmaf_with_uncertainty),tools/vmaf-tune/src/vmaftune/cli.py(predictsubcommand gains--with-uncertainty/--calibration-sidecar/--alpha),tools/vmaf-tune/tests/test_conformal.py(new),docs/ai/conformal-vqa.md(new). No engine code touched; no upstream-shared paths. - Invariant: the conformal wrapper sits outside the ONNX graph and adds no new runtime dependency —
conformal.pyimports only the standard library (math,statistics,dataclasses,json,warnings). Future calibration-sidecar shapes use themethoddiscriminator string for versioning; do not rename"split-conformal"/"cv-plus"without bumping the loader. ThePredictor.predict_vmaf_with_uncertaintysignature is the Python-API contract consumed byvmaf-tune predict --with-uncertainty; renaming or reordering its keyword args breaks the CLI in lockstep. - On upstream sync: no action required.
vmaf-tuneis a fork-local tool; upstream Netflix/vmaf has no per-shot prediction surface. - Re-test on rebase:
python3 -m pytest tools/vmaf-tune/tests/test_conformal.py -q
python3 -m pytest tools/vmaf-tune/tests/test_predictor.py -q
CI paths-ignore deny-list on heavy workflows (ADR-0341, 2026-05-09)¶
- Touches:
.github/workflows/libvmaf-build-matrix.yml(fork-local —paths-ignore:block underpull_request:),.github/workflows/tests-and-quality-gates.yml(fork-local — same block),docs/adr/0341-ci-paths-ignore-doc-only-prs.md+ index fragment,changelog.d/changed/ci-paths-ignore-doc-only.md. - Invariant: the deny-list must stay strictly documentation-only (
docs/**,**/*.md,changelog.d/**,CHANGELOG.md,.workingdir2/**). Any path that contributes to a build, test, or lint input —libvmaf/**,meson.build,meson_options.txt,subprojects/**,python/**,ai/**,mcp-server/**,model/**,testdata/**,.github/workflows/**— must NEVER appear in the deny-list, otherwise the corresponding required check is silently skipped on a code-touching PR. The Required Checks Aggregator (ADR-0313) catches only the doc-only case (no required check ever ran for any required name); a too-broad deny-list would lose build coverage without anyone noticing. - On upstream sync: Netflix/vmaf upstream does not carry these two workflow files (they are fork-local additions). No sync conflict expected.
- Re-test on rebase:
HDR VMAF model search — Path C documentation only (2026-05-09)¶
- Files added (this fork only; upstream Netflix/vmaf has none of these):
model/vmaf_hdr_model_card.md— discoverable warning that the HDR scoring path falls back to the SDRvmaf_v0.6.1.jsonweights. Filename deliberately uses.md, not.json, so thevmaftune.hdr.select_hdr_vmaf_modelglob (vmaf_hdr_*.json) keeps returningNone.docs/research/0089-hdr-vmaf-model-search.md— verbatim trail of the source-or-train survey (URLs + access dates).changelog.d/added/hdr-vmaf-model-search.md— release-notes fragment per ADR-0221.- ADR-0300 grew an inline
### Status update 2026-05-09: HDR model statussection. - Why no model JSON ships: Path A negative findings (no public Netflix HDR VMAF model exists; HDRMAX is a different algorithm not loadable by libvmaf's JSON path). Path B deferred behind gated subjective HDR corpora + multi-day training compute. No fabricated weights are introduced.
- On upstream sync: if Netflix lands
vmaf_hdr_*.jsoninNetflix/vmaf/model/, port via/port-upstream-commit; the resolver picks it up automatically with novmaftunechange. Then deletemodel/vmaf_hdr_model_card.md(or rewrite it as a normal model card describing the upstream weights). Watch https://github.com/Netflix/vmaf/issues/645 for the upstream release announcement. - Re-test on rebase: no behavioural change — pure docs. Sanity:
python3 -c "from pathlib import Path; \
import sys; sys.path.insert(0,'tools/vmaf-tune/src'); \
from vmaftune.hdr import select_hdr_vmaf_model; \
print(select_hdr_vmaf_model(Path('model')))"
# Expect: None — confirms the .md card does not match the glob
ADR-0349 — fr_regressor_v3 namespace resolution (2026-05-09)¶
- Rebase impact: none. Docs-only change — adds ADR-0349, an append-only status appendix on ADR-0302 per ADR-0028, a
## fr_regressor_* namespace mapblock inai/AGENTS.md, and two changelog fragments. No upstream Netflix/vmaf surface touched; nofr_regressor_*registry rows touched (sha256s for_v1,_v2,_v2_ensemble_v1_seed{0..4},_v3all unchanged); no C / Python / ONNX bytes modified. - What to check after a rebase: nothing automated. The only drift risk is a future agent claiming
fr_regressor_v3plus_featuresfor an unrelated workstream —ai/AGENTS.mdcarries the reservation; reviewers verify the map row exists before approving any newfr_regressor_*registry id. - Reproducer:
```bash # ADR + AGENTS.md namespace map present and consistent: test -f docs/adr/0349-fr-regressor-v3-namespace.md grep -q "fr_regressor_* namespace map" ai/AGENTS.md grep -q "fr_regressor_v3plus_features" ai/AGENTS.md docs/adr/0349-fr-regressor-v3-namespace.md # Status appendix present on ADR-0302: grep -q "Status update 2026-05-09: namespace collision resolved" \ docs/adr/0302-encoder-vocab-v3-schema-expansion.md # Existing v3 production row bit-identical (sha256 unchanged): python3 -c "
import json reg = json.load(open('model/tiny/registry.json')) v3 = next(m for m in reg['models'] if m['id'] == 'fr_regressor_v3') assert v3['sha256'] == 'eaa16d23461eda74940b2ed590edfcaf13428aade294e47792a5a15f4d3b999c', v3 assert v3['smoke'] is False print('OK: fr_regressor_v3 production row unchanged') "
Registry test still passes:¶
bash core/test/dnn/test_registry.sh
0327 — Pre-push PR-body deliverables validator hook¶
- Touches:
scripts/ci/validate-pr-body.sh(new),scripts/git-hooks/pre-push(new),scripts/ci/test-validate-pr-body.sh(new),Makefile(hooks-installtarget adds the pre-push symlink). Re-usesscripts/ci/deliverables-check.shparser verbatim — no upstream-shared file is modified. - Invariant: parser shape parity with
.github/workflows/rule-enforcement.ymldeep-dive-checklist gate (ADR-0108). The validator constructs aPATHshim that interceptsgit diff --name-onlycalls only; every othergitinvocation falls through to the real binary. - On upstream sync: not applicable — these files are entirely fork-local and Netflix has no equivalent. If
scripts/ci/deliverables-check.shis ever rewritten or moved, the validator's exec path (scripts/ci/deliverables-check.sh) and the test harness's expected exit codes must follow. bash scripts/ci/test-validate-pr-body.sh # 8/8 cases pass
0320 — Semgrep # nosemgrep cites on Netflix-upstream Python harness (Research-0090)¶
- Touches:
python/vmaf/core/asset.py,python/vmaf/core/executor.py,python/vmaf/core/feature_extractor.py,python/vmaf/core/quality_runner.py,python/vmaf/core/result_store.py,python/vmaf/tools/decorator.py,python/test/command_line_test.py,python/test/feature_extractor_test.py,python/test/ssimulacra2_test.py,python/vmaf/config.py. - Invariant: every fork-added
# nosemgrep: <rule-id>line is paired with an inline cite toResearch-0090. The cite + rule-id pair is the load-bearing artifact (per memoryfeedback_no_guessing: every "false positive" claim ships its safety proof). If an upstream sync removes the cited line of code, drop the cite-comment block too. If upstream adds adefusedxmlfix at theElementTree.parse()site (feature_extractor.py:115,quality_runner.py:1496), keep upstream's fix and drop our suppressions. config.py:40(the SSL-bypass deletion) is a fork-exclusive security fix; if upstream resurrectsssl._create_unverified_contexton a sync, do not re-merge it — the bypass clobbers the process-global default and is unjustified per Research-0090, F1. semgrep scan --config=p/cwe-top-25 --config=p/c --config=p/python . \ --metrics=off --json | jq '.results | length'
# expect 0 — every legit finding either has a # nosemgrep cite or was fixed
0321 — Security-scans workflow registry-pack list (Research-0090)¶
- Touches:
.github/workflows/security-scans.yml,.github/workflows/lint-and-format.yml. - Invariant: the registry packs the workflow cites (
p/cwe-top-25+p/c+p/python) are validated againsthttps://semgrep.dev/c/p/<pack>— the previously-citedp/cert-c-strict,p/cert-cpp-strict, andp/cpppacks were retired by Semgrep in 2025 and 404. Thelint-and-format.ymlpull of${{ github.* }}intoenv:(clang-tidy + clang-tidy-sycl steps) defusesrun-shell-injection; preserve the pattern on any edit. See Research-0090, F2/F3. for pack in p/cwe-top-25 p/c p/python; do code=\((curl -sIL "https://semgrep.dev/c/\)" | head -1 | awk '{print $2}') [ "$code" = "200" ] && echo "\({pack}: OK" || echo "\): FAIL ($code)"
0320 — CodeQL C bulk sweep (78 deferred alerts → 60 fixed, 14 deferred to T7-5)¶
- Touches:
core/src/feature/{cambi.c,ciede.c,integer_adm.c,integer_psnr.c,adm_tools.h,third_party/xiph/psnr_hvs.c},core/src/feature/x86/{adm_avx2.c,adm_avx512.c,ansnr_avx2.c,ansnr_avx512.c,vif_avx2.c,vif_avx512.c},core/src/{pdjson.c,svm.cpp},core/test/{test_cpu.c,test_model.c},core/tools/{y4m_input.c,yuv_input.c,vmaf_bench.c}. All butvmaf_bench.care upstream-mirror Netflix files. - Invariant: widening casts on integer multiplications (
(size_t),(uint64_t),(double)) are LHS-prefixed before the multiply, never wrapped around the whole expression — the latter is a no-op againstcpp/integer-multiplication-cast-to-long. Deleted commented-out blocks (e.g., the AVX-512 VP-loop dead variant inadm_avx512.c::adm_dwt2_inverse) are gone for good; if upstream brings them back, they reintroduce the alerts.iqa/convolve.cwas deliberately left untouched: prefixing(double)on the float×float multiplications inside the scalar reference path breaks bit-exactness against the AVX2 path enforced bytest_iqa_convolve— CodeQL alert deferred to a follow-up that updates both paths in lockstep. - On upstream sync: any upstream change that re-introduces the deleted comment blocks or rewrites the cast forms will surface the alerts again. The
cambi_scoresignature change (CambiBuffers buffers→const CambiBuffers *buffers) is fork-local and likely to conflict with upstream patches that touch that function. The 14 deferredVifBufferlarge-parameter alerts are tracked under T7-5 (multi-backend coordinated refactor including NEON). - Re-test on rebase: cd libvmaf && meson test -C build # all 50+ C tests make test-netflix-golden # upstream golden gate
# Re-run CodeQL on master afterwards; the 60 fixed alerts must stay closed.
CodeQL cpp/declaration-hides-variable sweep (2026-05-09)¶
- What changed: Mechanical rename / scope-tighten / dedupe sweep closing 64 open
cpp/declaration-hides-variableCodeQL alerts onmaster. Touched files:core/src/feature/cambi.c,core/src/feature/x86/adm_avx2.c,core/src/feature/x86/adm_avx512.c,core/src/feature/x86/vif_avx2.c,core/src/feature/x86/vif_avx512.c. All five are upstream-mirror; the Netflix copyright header is preserved on each. - Renames adopted (semantic over
_2suffix): cambi.c: innerint errshadowing function-scopeerrbecomesmkdir_err(heatmaps init) andsrc_err(full-ref extract path).adm_avx2.c/adm_avx512.c: thej == 0first-column special-case block is wrapped in{ ... }so itsj0..j3ands0..s3stop being visible to the per-jtail loop. The inner duplicate__m256i add_shift_HP_vex = _mm256_set1_epi32(32768)(and 512-bit twin) is removed — bit-identical to the function-scope value already in scope. The__m256i rfactor1that shadowed the function-scopefloat rfactor1[3]becomesrfactor_v0/_v1/_v2(and the AVX-512 twin likewise).vif_avx2.c/vif_avx512.c: tap-loop locals followf_tap,r_top/r_bot,d_top/d_botfor the s0 stage, andf_tap0/f_tap1,r_back0/r_fwd0, etc. for the AVX-512 paired-tap stage. Inner per-fj__m256i fq/__m512i fqshadows of the centre-tap broadcast becomef_tap. Inner-block duplicates of function-scoperef/dis/stride/ii(identical types and initialisers) are simply removed. The two scalarVifResiduals residualsdeclarations that shadowed function-scopeResiduals512 residualsbecometail_residuals. The twoconst uint16_t fcoeffdeclarations that shadowed function-scope__m512i fcoeffbecomefcoeff_scalar.- Invariant: bit-exactness gate — the rename sweep must not change any score. The Netflix CPU golden 3 (
src01_hrc00,checkerboard_1,checkerboard_10) ran clean against this PR. All 76 VMAF-targeted Python tests pass; the 9 unrelated pre-existing failures (NIQE, PyPSNR, FileSystemResultStore) reproduce on a pristineorigin/mastercheckout. - On upstream sync: Netflix has no equivalent renames on upstream
masteras of2026-05-09. When syncing, prefer the fork's renamed identifiers (the CodeQL gate depends on them). If Netflix later renames the same locals differently, reconcile by keeping fork names and updating any imported chunks at port time. - Re-test on rebase: meson test -C build --suite=fast PYTHONPATH=$PWD/python python3 -m pytest \ python/test/quality_runner_test.py -k test_run_vmaf \ python/test/vmafexec_test.py \ python/test/vmafexec_feature_extractor_test.py \ -m "not slow" -q
ADR-0209 v1 stdio runtime (T5-2b) — Embedded MCP server (2026-05-08)¶
- Touches:
core/src/mcp/{mcp.c,dispatcher.c,transport_stdio.c,mcp_internal.h,meson.build,3rdparty/cJSON/{cJSON.c,cJSON.h,LICENSE}},core/test/test_mcp_smoke.c,core/test/meson.build. All paths are fork-local. cJSON is vendored verbatim from upstreamDaveGamble/cJSON@v1.7.18under its MIT license. - Invariant: every TU under
core/src/mcp/(other than the vendored cJSON dir) is fork-local with theCopyright 2026 Lusoris and Claude (Anthropic)header; cJSON keeps its upstream MIT header verbatim. The public ABI incore/include/libvmaf/libvmaf_mcp.his unchanged from T5-2 — only function bodies flipped from-ENOSYSto working implementations. SSE / UDS still return-ENOSYSso the v2 PR can wire them without touching the public surface. - On upstream sync: no action required. Netflix/vmaf upstream has no embedded MCP surface; the entire
core/src/mcp/subtree is fork-local. If upstream ever adds an MCP surface, expect a port-only sync since names will collide. cd libvmaf && meson setup build -Denable_cuda=false -Denable_sycl=false \ -Denable_mcp=true -Denable_mcp_stdio=true ninja -C build && meson test -C build test_mcp_smoke -v
ADR-0334 — state.md-touch-check CI gate (2026-05-08)¶
- Touches:
.github/workflows/rule-enforcement.yml(new top-level jobstate-md-touch-check),scripts/ci/state-md-touch-check.sh(new),scripts/ci/test-state-md-touch-check.sh(new),scripts/ci/AGENTS.md(new rebase-sensitive-surface row),.github/PULL_REQUEST_TEMPLATE.md(already carries the "Bug-status hygiene" section +no state delta: REASONopt-out — coupled to the script's regex). No upstream-shared paths. - Invariant: the gate's trigger predicate (Conventional-Commit
fix:prefix, barebugtoken in title, GitHub close-keywordscloses/fixes/resolves#N, unchecked Bug-status-hygiene checkbox) and opt-out sentinel (no state delta: REASON) match the wording of the## Bug-status hygienesection in.github/PULL_REQUEST_TEMPLATE.md. Reword the template only alongside the script. The job carries thepull_request.draft == false || github.event_name != 'pull_request'gate (ADR-0331 pattern) — keep that on any future hoist into the required-aggregator set. - On upstream sync: Netflix/vmaf has no equivalent rule. No conflict expected; the workflow file is fork-introduced.
- Re-test on rebase: bash scripts/ci/test-state-md-touch-check.sh python3 -c "import yaml; yaml.safe_load(open('.github/workflows/rule-enforcement.yml')); print('YAML OK')" pre-commit run shellcheck --files scripts/ci/state-md-touch-check.sh scripts/ci/test-state-md-touch-check.sh pre-commit run shfmt --files scripts/ci/state-md-touch-check.sh scripts/ci/test-state-md-touch-check.sh
SYCL PSNR chroma extension (T3-15(b), 2026-05-09)¶
- Touches:
core/src/feature/sycl/integer_psnr_sycl.cpp(per-extractor chroma device buffers, per-plane SSE accumulators, and aprovided_featuresextension topsnr_y/psnr_cb/psnr_cr),core/src/sycl/AGENTS.md(per-kernel rebase-sensitive invariant for the chroma-on-per-extractor-buffer arrangement),docs/metrics/features.md(footnote ¹ refresh — all three GPU PSNR extractors now emit chroma),docs/adr/0192-gpu-long-tail-batch-3.mdReferences-section status update,changelog.d/added/sycl-psnr-chroma.md. - Invariant on the chroma upload path: chroma planes ride on per-extractor device buffers populated by host-side staging copies in the combined-graph
pre_fncallback — NOT the SYCL state's shared frame buffer (vmaf_sycl_shared_frame_init), which is luma-only by design. Luma stays graph-recorded; chroma SSE kernels run direct inpost_fnon the same in-order combined queue. The CUDA twin (PR #520 / commit 7f3d58a5) uses the existing CUDA per-plane picture infrastructure and therefore has no equivalent invariant. - On upstream sync: Netflix/vmaf upstream has no SYCL backend at all, so conflict probability is zero on
psnr_sycl. If an upstream port to the fork's SYCL runtime someday extendsvmaf_sycl_shared_frame_initto allocate chroma planes, the PSNR extension can be migrated onto it and the per-extractor chroma buffers retired — but only after a cross-backend gate run confirms bit-exactness against CPU atplaces=4(ADR-0214). source /opt/intel/oneapi/setvars.sh CC=icx CXX=icpx meson setup build-sycl libvmaf \ -Denable_sycl=true -Denable_cuda=false ninja -C build-sycl python3 scripts/ci/cross_backend_vif_diff.py \ --vmaf-binary build-sycl/tools/vmaf \ --reference testdata/ref_576x324_48f.yuv \ --distorted testdata/dis_576x324_48f.yuv \ --width 576 --height 324 --pixel-format 420 --bitdepth 8 \ --feature psnr --backend sycl --device 0
# Expect 0/48 mismatches across psnr_y / psnr_cb / psnr_cr at places=4.
```text
Cppcheck nullPointer false-positive in dict.c (2026-05-09)¶
Files pinned:
core/src/dict.c:121(one-line redundant-condition fix indict_overwrite_existing). Why this rebase-note exists: Master CI'sCppcheck (Whole Project)gate started failing on commit14b5ffba(#537) and blocked every open PR because each PR rebases onto a broken master. The cppcheck finding was likely always present but masked bypaths-ignorefiltering on the prior workflow shape; PR #530 widened cppcheck's trigger surface and exposed it. Deleted the redundant&& valguard sincevalis already checked at the public entry-pointvmaf_dictionary_set(dict.c:137). No behavior change; cppcheck flags the original as "either the val check is redundant or there's a possible null deref" because it can't prove the interprocedural guarantee. Rebase-sensitivity: zero — change is local todict.c. Future upstream sync of this file should keep the fix or re-run cppcheck locally to confirm absence of recurrence.
Aggregator timeout bump (2026-05-09)¶
Files pinned:
.github/workflows/required-aggregator.yml(deadline 30→90 min, job timeout 35→100 min) Why: 41 PRs in flight 2026-05-09 morning hit Aggregator timeouts while real CI eventually passed. Bumping both deadlines unblocks the train without touching the underlying matrix. Rebase-sensitivity: zero — workflow file is wholly fork-local.
ARC self-hosted runner pool — pilot Cppcheck routing (2026-05-09)¶
.github/workflows/lint-and-format.yml(Cppcheckruns-on:ternary). Why: opt-in graceful migration; ADR-0359 + docs/development/ci-runners.md document the flip-the-variable recipe when the cluster is degraded. Rebase-sensitivity: zero — workflow file is fork-local.
ADR-0338 — macOS Vulkan-via-MoltenVK CI lane (2026-05-09)¶
- Touches:
.github/workflows/libvmaf-build-matrix.yml(fork-local — addsBuild — macOS Vulkan via MoltenVK (advisory)lane, addscontinue-on-errorplumbing onmatrix.experimental && matrix.moltenvk, addsInstall MoltenVK + Vulkan loader/headers (macOS)step, addsRun Vulkan smoke tests (macOS MoltenVK)step, gates the existing test/cache/tox steps on!matrix.moltenvk),docs/backends/vulkan/moltenvk.md(new fork-local doc),docs/adr/0127-vulkan-compute-backend.md(status-update appendix per the ADR's Proposed status — body untouched),docs/adr/0338-macos-vulkan-via-moltenvk-lane.md(new),docs/adr/_index_fragments/0338-macos-vulkan-via-moltenvk-lane.mdplus_order.txtappend (new),docs/research/0089-moltenvk-feasibility-on-fork-shaders.md(new),changelog.d/added/macos-vulkan-via-moltenvk-lane.md(new). - Invariant on the upstream-mirror file: none —
libvmaf-build-matrix.ymlis fork-local. The new lane'scontinue-on-errorclause MUST stay scoped tomatrix.experimental == true && matrix.moltenvk == trueso existingexperimental: truematrix entries (e.g. the macOS DNN lane) keep their default fail-fast behaviour.VK_ICD_FILENAMESMUST point at/opt/homebrew/etc/vulkan/icd.d/MoltenVK_icd.json— note theetc/vulkansegment, NOTshare/vulkan(the homebrew formula's install layout usesetc/; verified againstFormula/m/molten-vk.rb). - On upstream sync: Netflix upstream has no macOS Vulkan lane and no MoltenVK awareness; nothing to reconcile. If a future MoltenVK release drops support for
GL_EXT_shader_atomic_int64translation,moment.compwill fail on the lane; the fix path is in ADR-0338 §Decision (lane iscontinue-on-errorso it does not block PRs) — update the known-limitations table indocs/backends/vulkan/moltenvk.mdand either pin a working MoltenVK version in the brew install line or rewrite the shader. - Re-test on rebase:
python3 -c "import yaml; yaml.safe_load(open('.github/workflows/libvmaf-build-matrix.yml'))" && \
echo "YAML parse OK"
# Confirm the lane is still in the matrix:
grep -q "Build — macOS Vulkan via MoltenVK (advisory)" \
.github/workflows/libvmaf-build-matrix.yml
# Confirm the lane is NOT promoted to required-aggregator until one
# green run on master (per ADR-0338):
! grep -q "macOS Vulkan via MoltenVK" \
.github/workflows/required-aggregator.yml
# Confirm the ICD path is the etc/ one, not share/:
grep -q "etc/vulkan/icd.d/MoltenVK_icd.json" \
.github/workflows/libvmaf-build-matrix.yml
ADR-0363 — Mend Renovate replaces Dependabot (2026-05-09)¶
- Touches:
renovate.json(new, repo-root),.github/workflows/renovate.yml(new),.github/dependabot.yml(deleted — renamed to.github/dependabot.yml.disabled),docs/development/dependency-bot.md(new operator playbook),changelog.d/changed/renovate-supersedes-dependabot.md(new),docs/adr/0363-renovate-replaces-dependabot.md(new),docs/adr/_index_fragments/0363-renovate-replaces-dependabot.md(new). - Invariant:
.github/dependabot.ymlno longer exists onmaster; the disabled copy isdependabot.yml.disabled. On upstream sync, if Netflix ever ships their owndependabot.yml, do NOT restore it — the fork intentionally uses Renovate. Merge the upstream file intodependabot.yml.disabledfor reference only. - Upstream interaction: none. Netflix/vmaf upstream has no Renovate config. Conflict risk is zero unless upstream adds
renovate.jsonor restoresdependabot.yml. - Re-test on rebase:
# Verify the workflow SHA-pin is still present and non-floating:
grep -E 'renovatebot/github-action@[a-f0-9]{40}' .github/workflows/renovate.yml
# Verify dependabot.yml is still absent:
test ! -f .github/dependabot.yml && echo "ok: dependabot.yml absent"
# Validate renovate.json syntax (requires Node):
node -e "JSON.parse(require('fs').readFileSync('renovate.json','utf8')); console.log('JSON valid')"
ADR-0355 — Symphony-inspired agent-dispatch infrastructure (2026-05-09)¶
Files added (all fork-introduced, none mirror upstream):
.claude/workflows/_template.md,.claude/workflows/codeql-alert-sweep.md,.claude/workflows/simd-port.md,.claude/workflows/feature-extractor-port.md.scripts/lib/__init__.py,scripts/lib/backlog_tracker.py,scripts/lib/AGENTS.md.scripts/ci/agent-eligibility-precheck.py(new row inscripts/ci/AGENTS.md"Rebase-sensitive surfaces" table).docs/development/agent-dispatch.md. Why this rebase-note exists: pure additive, all paths are fork-only (.claude/,scripts/lib/, fork-only docs). Upstream Netflix/vmaf has no.claude/, noscripts/lib/, and nodocs/development/agent-dispatch.md, so the merge surface is zero on/sync-upstream. The only coupling is internal betweenscripts/ci/agent-eligibility-precheck.pyandscripts/lib/backlog_tracker.py(sys.path import). Both files move together; documented inscripts/lib/AGENTS.mdand a new row inscripts/ci/AGENTS.md. Rebase-sensitivity: zero w.r.t. upstream. Internal-only: renamingBacklogItemfield names or theBacklogTracker/GitHubTrackerpublic method signatures is a breaking change for the precheck and any future state-audit script — guard via the smoke listed in Research-0091 §"Smoke results" before any rename PR. Format-coupling note: the BACKLOG.md row regex (scripts/lib/backlog_tracker.py:_ID_PATTERN) is brittle against table-shape edits. If a future BACKLOG.md edit adds a column or renames a status word, the parser will silently mis-classify rows — the smoke parses 101 rows on master at 2026-05-09; expect ≥ 100 after any structural edit.
0350 — psnr_hvs AVX-512 ceiling re-bench (ADR-0350, T3-9 (a))¶
docs/adr/0350-psnr-hvs-avx512-ceiling.md— closure ADR.docs/adr/0160-psnr-hvs-neon-bitexact.md— appended### Status update 2026-05-09appendix.docs/research/0091-psnr-hvs-avx512-bench-2026-05-09.md— empirical companion (cycle share, Amdahl ceiling, reproducer). Why this rebase-note exists: T3-9 (a) closes as AVX2 ceiling. The result has zero rebase-sensitivity by itself — no engine code changes — but the bit-exactness invariants that lock it to a ceiling do. The 78.42 % scalar tail incalc_psnrhvs_avx2/calc_psnrhvs_neonis locked by ADR-0138 / ADR-0139's "per-lane-scalar float reduction" rule (carried by ADR-0159 / ADR-0160). If a future upstream sync ofcore/src/feature/third_party/xiph/psnr_hvs.c(the Xiph/Daala DCT) changes the per-block summation tree — e.g. partial folding, re-ordered means, vectorised mask reductions — the AVX2 + NEON TUs incore/src/feature/x86/psnr_hvs_avx2.candcore/src/feature/arm64/psnr_hvs_neon.cMUST be re-audited against the new scalar reference, and the ceiling argument in ADR-0350 must be re-run (because the 78 / 15 cycle-share split would shift). Rebase-sensitivity: low for the ceiling decision itself (empirical re-bench on a current host is cheap — 30 seconds via the reproducer in Research-0091 §7); high for the underlying bit-exactness invariants the decision rests on (Netflix golden trips on ≥ 5.5e-5 drift per ADR-0160 §Context). The ADR-0350 §Verification reproducer is the gate — re-run it if the cycle share shifts, the Netflix normal-pair fixture changes, or a new host class (e.g. wide-issue Granite Rapids) goes into CI.
0320 — FFmpeg n8.1 → n8.1.1 base bump (2026-05-09)¶
- Touches:
ffmpeg-patches/series.txt(header comment),ffmpeg-patches/README.md(apply / verify / smoke sections),ffmpeg-patches/test/build-and-run.sh(FFMPEG_SHAdefault),scripts/ci/ffmpeg-patches-check.sh(header comment;FFMPEG_BRANCHenv default unchanged atrelease/8.1since the branch tracks point releases),docs/development/automated-rule-enforcement.md(gate description). The 9.patchfiles themselves are unchanged — every patch in the series applied cleanly, cumulatively, against pristinen8.1.1viagit am --3way. - Upstream source: FFmpeg upstream point release n8.1.1 (commit
239f2c7"Bump micro for 8.1.1") — bug-fix-only on top of n8.1, no API or AVOption breakage that the patch stack consumes. - Invariant: the patch stack continues to apply against the current tip of FFmpeg's
release/8.1branch. Per ADR-0118 and ADR-0186 §FFmpeg patch coupling, the verification gate is cumulativegit am --3wayagainst a pristine checkout, not per-patch standalone apply. The scripts/ci/ffmpeg-patches-check.sh local gate usesgit apply(no commit) but accumulates state in the same way. - On upstream sync: no action required. If a future FFmpeg point release (n8.1.2 or n8.2) lands new hunks that conflict with one of the patches, regenerate the affected patches via
git format-patchon the resolved state, bump the references in the five files listed under "Touches", and add a fresh rebase-notes entry citing the conflict file(s). - Re-test on rebase:
cd /tmp && rm -rf ffmpeg-n811 && \
git clone --depth 1 --branch n8.1.1 \
https://git.ffmpeg.org/ffmpeg.git ffmpeg-n811
git -C /tmp/ffmpeg-n811 config user.email agent@local
git -C /tmp/ffmpeg-n811 config user.name agent
for p in ffmpeg-patches/000*-*.patch; do
git -C /tmp/ffmpeg-n811 am --3way "$p" || break
done
bash scripts/ci/ffmpeg-patches-check.sh
ADR-0281 follow-up — QSV install-matrix discoverability backfill (2026-05-08)¶
- Touches:
docs/getting-started/install/{arch,fedora,ubuntu,macos,windows}.md(new## Intel QSVsection per page),docs/adr/0281-vmaf-tune-qsv-adapters.md(status-update appendix per ADR-0028),changelog.d/changed/qsv-install-matrix-docs.md(new fragment). No code, no engine, no upstream-shared C / Python source touched. Pure documentation backfill closing the SYCL-audit research-0086 Topic C gap (issue #464). - Invariant: each per-OS QSV section pins the package names against verified upstream URLs with a
Verified 2026-05-08access date. The hardware-generation matrix is sourced from the public Wikipedia "Intel Quick Sync Video — Hardware decoding and encoding" table; if Intel revises which generation supports AV1 encode (e.g. backports the encoder to Lunar Lake / Meteor Lake silicon currently absent from the table), the matrix in all five pages must move in lockstep — the Arch / Fedora / Ubuntu / Windows pages all carry the same matrix verbatim. The macOS page deliberately omits the matrix (QSV unsupported on macOS). - On upstream sync: no action required — Netflix/vmaf upstream does not ship per-OS install pages under
docs/getting-started/install/; that tree is fork-only.
# Lint the install pages (markdownlint via pre-commit):
pre-commit run --files docs/getting-started/install/*.md
# Verify each page (except alpine + macos) still carries the matrix:
for f in arch fedora ubuntu windows; do grep -q 'Arc Battlemage' "docs/getting-started/install/${f}.md" || echo "MISSING: ${f}"
# Confirm the macOS page documents QSV as unsupported:
grep -q 'Intel QSV. is unsupported on macOS' docs/getting-started/install/macos.md
0333 — vmaf-tune Phase F multi-pass encoding (ADR-0333)¶
Touches:
tools/vmaf-tune/src/vmaftune/codec_adapters/__init__.py(CodecAdapter Protocol gainssupports_two_pass: bool+two_pass_args(...))tools/vmaf-tune/src/vmaftune/codec_adapters/x265.py(overrides both)tools/vmaf-tune/src/vmaftune/encode.py(EncodeRequestgainspass_number/stats_path;build_ffmpeg_commandadds the 2-pass argv splice + pass-1 null-muxer redirect; newrun_two_pass_encode)tools/vmaf-tune/src/vmaftune/corpus.py(CorpusOptions.two_pass, routing initer_rows)tools/vmaf-tune/src/vmaftune/cli.py(--two-passflag oncorpus/recommendsubparsers) Invariant: 2-pass encoding routes through the codec adapter viasupports_two_pass+two_pass_args(pass_number, stats_path). The encode driver never branches on codec name. Adapters withsupports_two_pass = Falseare honoured silently (single-pass fallback with stderr warning); the seam is open for sibling codec adapters (libx264, libsvtav1, libvvenc, libaom-av1) to opt in by overriding the two methods on their adapter file alone. This is the fork-local extension to the ADR-0237 Phase A multi-codec contract; upstream Netflix/vmaf has no equivalent and does not own this code path. Re-test:
(Optional, requires ffmpeg + libx265 in the runner's PATH:)
VMAF_TUNE_INTEGRATION=1 python -m pytest \
tests/test_codec_adapter_x265_two_pass.py::test_real_x265_two_pass_smoke -q
Rebase-sensitivity: zero from upstream — tools/vmaf-tune/ is fork-local. The only concern is the codec_adapters Protocol shape: a future upstream commit that adds a sibling codec adapter SHOULD inherit the supports_two_pass = False default and either explicitly opt in or leave the flag off. Downstream sibling-codec PRs in this fork should follow the ADR-0288 / ADR-0333 pattern: one adapter file, override the two methods, add a test file mirroring test_codec_adapter_x265_two_pass.py.
ADR-0360 — CAMBI CUDA port (T3-15a, 2026-05-09)¶
Files pinned:
core/src/feature/cuda/integer_cambi_cuda.c(new)core/src/feature/cuda/integer_cambi_cuda.h(new)core/src/feature/cuda/integer_cambi/cambi_score.cu(new)core/src/feature/feature_extractor.c(addedvmaf_fex_cambi_cudato list)core/src/meson.build(addedcambi_scoretocuda_cu_sources, addedinteger_cambi_cuda.cto CUDA feature sources)
Why: The CUDA twin of vmaf_fex_cambi (Strategy II hybrid — three GPU kernels for the embarrassingly parallel stages; calculate_c_values + topK on CPU). Registers vmaf_fex_cambi_cuda under #if HAVE_CUDA guard.
Rebase-sensitivity: low. The three new files are wholly fork-local and will not conflict. The two upstream-shared files have small, self-contained hunks:
feature_extractor.c: theextern vmaf_fex_cambi_cudadeclaration and the&vmaf_fex_cambi_cudaarray entry are inside a#if HAVE_CUDAblock. Upstream's additions to this file (new feature extractors, new dispatch flags) will not conflict unless Netflix adds their own CUDA twin for CAMBI (unlikely — they don't ship a CUDA backend).meson.build: thecambi_scoreentry in thecuda_cu_sourcesdict and theinteger_cambi_cuda.cline in the CUDA sources list. Any upstream changes tomeson.buildthat restructure thecuda_cu_sourcesdict would require a manual merge; the dict entries are sorted alphabetically by key, socambi_scorelands betweenadm_scoreandmotion_score.
If upstream adds cambi_cuda themselves: drop the fork copy and check for API divergence. Strategy II hybrid is the natural choice; the upstream implementation may differ if they choose Strategy III (fully-on-GPU calculate_c_values).
cambi_internal.h dependency: integer_cambi_cuda.c includes core/src/feature/cambi_internal.h (fork-added trampoline exposing cambi.c's static helpers). If upstream significantly refactors cambi.c (renames vmaf_cambi_preprocessing, vmaf_cambi_calculate_c_values, etc.), cambi_internal.h must be updated alongside. This is the same dependency the Vulkan twin (cambi_vulkan.c) has — see ADR-0210's rebase note for the full list of exposed functions.
Vulkan submit-pool PR-B: six secondary kernels (2026-05-09, ADR-0353)¶
Files changed:
core/src/feature/vulkan/ssim_vulkan.ccore/src/feature/vulkan/ciede_vulkan.ccore/src/feature/vulkan/ms_ssim_vulkan.ccore/src/feature/vulkan/motion_v2_vulkan.ccore/src/feature/vulkan/float_psnr_vulkan.ccore/src/feature/vulkan/float_motion_vulkan.ccore/src/feature/vulkan/AGENTS.mddocs/adr/0353-vulkan-submit-pool-pr-b-six-kernels.md
Why this rebase-note exists: six Vulkan host-glue TUs were migrated from per-frame command-buffer and descriptor-set allocation to the VmafVulkanKernelSubmitPool abstraction (ADR-0256). Any Netflix upstream sync that touches these same files (unlikely — they are fork-local) must preserve the VmafVulkanKernelSubmitPool fields in the state struct and the pool-destroy-before-pipeline-destroy ordering in close_fex().
Rebase-sensitivity: low. All six files are entirely fork-local; Netflix upstream does not have a Vulkan backend. The submit-pool API is defined in core/src/vulkan/kernel.h (also fork-local). No public header or C-API surface was changed; the FFmpeg patch series is unaffected.
Key invariant to preserve on rebase: vmaf_vulkan_kernel_submit_pool_destroy MUST be called before vmaf_vulkan_kernel_pipeline_destroy in every migrated kernel's close_fex(). See core/src/feature/vulkan/AGENTS.md §"Submit-pool ordering invariant".
0354 — Vulkan submit-pool PR-C: submit_pool_destroy-before-pipeline ordering¶
- Touches:
core/src/feature/vulkan/cambi_vulkan.c,core/src/feature/vulkan/ssimulacra2_vulkan.c,core/src/feature/vulkan/float_ansnr_vulkan.c,core/src/feature/vulkan/moment_vulkan.c. - Invariant: In every migrated extractor,
vmaf_vulkan_kernel_submit_pool_destroy()MUST precede everyvmaf_vulkan_kernel_pipeline_destroy()call inclose_fex(). Reversing the order frees the pool's command buffers after the pipeline's command pool is destroyed — undefined behaviour per Vulkan spec §6.2. - Re-test:
meson test -C build --suite=vulkanpasses.scripts/ci/cross_backend_vif_diff.pyshowsplaces=4for all four extractors on all three target devices (RTX 4090, Arc A380, RADV iGPU).
0231 — Vulkan submit-pool migration PR A: adm + motion + psnr (ADR-0291)¶
0231 — Vulkan submit-pool migration PR A: adm + motion + psnr (ADR-0352)¶
- Touches:
core/src/feature/vulkan/adm_vulkan.c,core/src/feature/vulkan/motion_vulkan.c,core/src/feature/vulkan/psnr_vulkan.c(all fork-local Vulkan kernels; no upstream C paths touched),changelog.d/changed/vulkan-submit-pool-pr-a-adm-motion-psnr.md,docs/adr/0291-vulkan-submit-pool-pr-a-adm-motion-psnr.md. - Invariant: Each migrated TU adds
VmafVulkanKernelSubmitPool sub_pooland pre-allocatedVkDescriptorSetfield(s) to its state struct. The pool must be destroyed (vmaf_vulkan_kernel_submit_pool_destroy) beforevmaf_vulkan_kernel_pipeline_destroyinclose_fex(); reversing the order would destroy the descriptor pool while the submit pool still holds live command buffer + fence references. Descriptor sets allocated viavmaf_vulkan_kernel_descriptor_sets_allocare freed implicitly by the descriptor pool tear-down — do NOT callvkFreeDescriptorSetson them inclose_fex(). Formotion_vulkan, the pre-allocated set is rebound once per frame viavkUpdateDescriptorSetsbecause the blur ping-pong changes whichblur[]slot is "current"; foradm_vulkanandpsnr_vulkanthe sets are stable afterinit()and require no per-frame update. - Upstream interaction: none. All three files are fork-local Vulkan kernel TUs not present in Netflix/vmaf upstream.
- On upstream sync: zero interaction. Upstream cannot conflict with this PR's paths. The Vulkan backend is entirely fork-introduced.
- Re-test on rebase:
meson test -C build --suite=fast
# Cross-backend parity gate (places=4):
python python/test/cross_backend_diff.py \
--features adm motion psnr \
--backend vulkan cpu \
--places 4 \
--yuv testdata/yuv/src01_hrc00_576x324.yuv \
testdata/yuv/src01_hrc01_576x324.yuv
ADR-0350 — FFmpeg libvmaf filter CUDA backend selector (0010 patch)¶
Patch: ffmpeg-patches/0010-libvmaf-wire-cuda-backend-selector.patch.
libavfilter/vf_libvmaf.c— addscudaAVOption + state field + init / cleanup / picture-pool wiring underCONFIG_LIBVMAF_CUDA && !CONFIG_LIBVMAF_CUDA_FILTER.configure— adds--enable-libvmaf-cuda(EXTERNAL_LIBRARY_LISTentry + help text), promoteslibvmaf_cudafrom blanket-autodetect to gatedenabled libvmaf_cuda && require_pkg_config + check, preserves theenabled libvmaf && check_pkg_config libvmaf_cudain-filter probe so the new selector still works without the explicit flag when libvmaf ships CUDA. Why this rebase-note exists: Patch0010extends the SYCL (0003) / Vulkan (0004) per-context backend selectors to CUDA on the regularlibvmaffilter. The patch coexists with the upstream dedicatedlibvmaf_cudafilter (CONFIG_LIBVMAF_CUDA_FILTER) by gating its struct field and code paths on!CONFIG_LIBVMAF_CUDA_FILTER— the dedicated filter keeps owning its owncu_statefield. CLAUDE.md §12 r14 makes the patch update mandatory because the change touches a filter consumer of thevmaf_cuda_state_init/_import_state/_state_free/_preallocate_pictures/_fetch_preallocated_pictureC-API surface inlibvmaf_cuda.h. Rebase-sensitivity: low. The patch'svf_libvmaf.chunks are context-anchored on the SYCL/Vulkan selector blocks; if upstream FFmpeg renamesCONFIG_LIBVMAF_CUDA_FILTERor moves thelibvmaf_cuda.hinclude, the include guard at the top of the file needs the corresponding update. The configure hunks are context-anchored on the existing--enable-libvmaf-sycl/--enable-libvmaf-vulkanlines — those have proven stable across n8.0 → n8.1 → n8.1.1, so drift risk is low. WhenVmafCudaConfigurationever grows adevice_indexfield upstream, swap thecudaboolean for anint cuda_devicemirroring SYCL's shape (separate ADR + patch refresh). Verification gate: cumulativegit am --3wayreplay offfmpeg-patches/000{1..9}-*.patch+0010-*against pristine FFmpegn8.1.1PASS (2026-05-09). Build oflibavfilter/vf_libvmaf.oPASS under bothCONFIG_LIBVMAF_CUDA=0(selector errors at filter- init time per#elsebranch) andCONFIG_LIBVMAF_CUDA=1 && !CONFIG_LIBVMAF_CUDA_FILTER(selector active, picture-pool wiring compiles).
0320 — Vulkan instance / VMA apiVersion bump to 1.4 (Step B)¶
- Touches:
core/src/vulkan/common.c,core/src/vulkan/vma_impl.cpp,core/src/vulkan/AGENTS.md. - Invariant: the four
apiVersionsites (lines 54, 264, 374 ofcommon.c; line 22 ofvma_impl.cpp) request Vulkan 1.4, not 1.3. Together with the Step-Aprecisedecorations invif.comp/ciede.comp(PR #346) and the Phase-3 cross-subgroup release-acquire fix (PR #511), this gates the cross-backend places=4 contract on Arc + RADV. NVIDIA closure depends on Phase 3c (PR #512; block-on-merge until that lands). Netflix upstream does not carry a VMA dependency or a Vulkan backend; no upstream merge conflict expected on these files. - Re-test on rebase:
meson setup build -Denable_vulkan=enabled -Denable_cuda=false \
-Denable_sycl=false --buildtype=release
ninja -C build
for D in 0 1 2; do
python3 scripts/ci/cross_backend_parity_gate.py \
--vmaf-binary build/tools/vmaf \
--reference python/test/resource/yuv/src01_hrc00_576x324.yuv \
--distorted python/test/resource/yuv/src01_hrc01_576x324.yuv \
--width 576 --height 324 --pixel-format 420 --bitdepth 8 \
--backends cpu vulkan --vulkan-device "$D" \
--features vif ciede adm motion psnr
done
# All 0/N mismatches at places=4 once Phase 3c (PR #512) has landed.
ADR-0332 v2 runtime (T5-2c) — Embedded MCP server UDS + real compute_vmaf (2026-05-09)¶
- Touches:
core/src/mcp/{mcp.c,dispatcher.c,mcp_internal.h,meson.build,compute_vmaf.c,transport_uds.c},core/test/test_mcp_smoke.c. All paths are fork-local. No new third-party vendor drop in v2 — mongoose vendoring stays deferred to v3 with the SSE transport. - Invariant: same as ADR-0209 v1 — the entire
core/src/mcp/subtree is fork-local; the public ABI incore/include/libvmaf/libvmaf_mcp.his unchanged (only function bodies flipped —vmaf_mcp_start_udsfrom-ENOSYSto a working AF_UNIX listener;compute_vmaffrom a{"status":"deferred_to_v2"}placeholder to a realvmaf_score_pooledbinding). Per ADR-0128 § operational guardrails the UDS socket file is created mode 0700; thatchmodhappens invmaf_mcp_start_udsafterbindand is a load-bearing security invariant — do NOT relax it on rebase.compute_vmafruns on a per-call ephemeralVmafContextso the host's main scoring run is unperturbed; do NOT rewire it to reuseserver->ctxbecausevmaf_score_pooledcommits the model destructively to the context. - On upstream sync: no action required. Netflix/vmaf upstream has no embedded MCP surface. If upstream adds one, expect a port-only sync since names will collide.
- Re-test on rebase:
cd libvmaf && meson setup build -Denable_cuda=false -Denable_sycl=false \
-Denable_mcp=true -Denable_mcp_stdio=true \
-Denable_mcp_uds=true
ninja -C build && meson test -C build test_mcp_smoke -v
# Real-score smoke (single 576x324 pair):
build/test/test_mcp_smoke 2>&1 | tail -3 # expects "16 tests run, 16 passed"
ADR-0332 v3 runtime (T5-2d) — Embedded MCP server SSE transport (2026-05-09)¶
- Touches:
core/src/mcp/{mcp.c,mcp_internal.h,meson.build,transport_sse.c},core/meson_options.txt,core/test/test_mcp_smoke.c,docs/mcp/embedded.md,docs/adr/0332-mcp-runtime-v2.md(status-update appendix). All paths are fork-local. No third-party vendor drop in v3 — the originally-planned mongoose vendor was reversed because cesanta/mongoose 7.18 is GPL-2.0-only OR commercial, incompatible with the fork's BSD-3-Clause-Plus-Patent license (verified at upstream LICENSE 2026-05-09). The SSE transport is plain POSIX sockets in fork-owned C (~500 LOC). - Invariant: same as ADR-0209 / ADR-0332 v2 — the entire
core/src/mcp/subtree is fork-local; the public ABI incore/include/libvmaf/libvmaf_mcp.his unchanged (onlyvmaf_mcp_start_sse's body flipped from-ENOSYSto a working AF_INET listener). The SSE listener bindsINADDR_LOOPBACKonly; do NOT switch toINADDR_ANYwithout a separate ADR + auth design (v3 ships intentionally without CORS/Bearer/per-session auth on the assumption of a same-host trust boundary). The SSE stop path usesshutdown(SHUT_RDWR)beforeclose()— plainclose()of an AF_INET listening fd from another thread does NOT unblockaccept()on Linux; do NOT remove theshutdowncall.enable_mcp_sseis now afeatureoption (defaultauto), notboolean false. - On upstream sync: no action required. Netflix/vmaf upstream has no embedded MCP surface. Do NOT re-introduce mongoose (or any GPL-licensed HTTP library) on a future rebase without first amending CLAUDE §1 and adding a separate license-compatibility ADR.
- Re-test on rebase:
cd libvmaf && meson setup build -Denable_cuda=false -Denable_sycl=false \
-Denable_mcp=true -Denable_mcp_stdio=true \
-Denable_mcp_uds=true \
-Denable_mcp_sse=enabled
ninja -C build && meson test -C build test_mcp_smoke -v
build/test/test_mcp_smoke 2>&1 | tail -3 # expects "17 tests run, 17 passed"
Status update 2026-05-09 — placeholder-ref hardening¶
- Additional touches: same set as the 2026-05-08 ADR-0334 entry, no new files. The hardening adds a
git diff -U0 ... -- docs/state.mdcall insidescripts/ci/state-md-touch-check.sh(case 4a) plus 10 additional fixture cases inscripts/ci/test-state-md-touch-check.sh. - New invariant: inserted lines in
docs/state.md(lines starting with+, excluding the+++ b/...header) must not containthis PR/this commit/ bareTBD/<PR>/#NNN. Canonical accept forms arePR #Nandcommit `<sha>`. The placeholder vocabulary is coupled to PR #541's audit findings — reword in lockstep with the ADR-0334 status-update appendix if the fork's row template changes. - Re-test on rebase: same
bash scripts/ci/test-state-md-touch-check.shrun as the 2026-05-08 entry; the harness now reports18/18 passed(was8/8 passed).
0347 — Sanitizer matrix test-set scope (ADR-0347)¶
- Touches:
.github/workflows/tests-and-quality-gates.ymljobsanitizers(build + test step),core/test/meson.build(no edits — the absence of anysuite: 'unit'tag is the upstream state we now work with rather than against). - Invariant: the sanitizer job runs the full C unit-test set per sanitizer with a per-sanitizer deselect list driven by a
caseblock on${{ matrix.sanitizer }}. The deselect lists are load-bearing — each entry corresponds to a real bug tracked indocs/state.md. Under UBSan the build adds-Dc_args=-fno-sanitize=function -Dcpp_args=-fno-sanitize=functionto suppress the K&R-prototype harness UB; the mesoncasebranch must keep this build flag in sync with the test deselect entries. An upstream rebase that adds new test files viacore/test/meson.buildinherits full sanitizer coverage automatically (the workflow enumerates tests viameson test --list). - On upstream sync: if upstream Netflix lands a
suite: 'unit'tagging convention, the workflow is robust to it (we already enumerate frommeson test --list, not from--suite=unit). If upstream rewrites the harness to declarestatic char *test_X(void)with a(void)parameter, the-fno-sanitize=functionflag becomes redundant — leave it in place (zero cost) until a deliberate cleanup PR reverts the suppression. If upstream lands a fix for any of the surfaced defects (SVMModelParservalidation,feature_collectormetadata leak,integer_adm::div_lookuprace,framesyncmutex mismatch), drop the corresponding deselect row from the workflow'scaseblock in the same PR that pulls the upstream fix. cd libvmaf for SAN in address undefined thread; do EXTRA=() [ "$SAN" = undefined ] && EXTRA=( "-Dc_args=-fno-sanitize=function" "-Dcpp_args=-fno-sanitize=function" ) rm -rf "build-$SAN" CC=clang CXX=clang++ LDFLAGS=-fuse-ld=lld \ meson setup "build-$SAN" -Db_sanitize="$SAN" \ -Denable_cuda=false -Denable_sycl=false --buildtype=debug \ -Db_lto=false -Db_lundef=false "${EXTRA[@]}" meson compile -C "build-$SAN" case "$SAN" in address) EXCLUDE='test_model$|test_predict$|test_float_ms_ssim_min_dim$' ;; undefined) EXCLUDE='test_model$' ;; thread) EXCLUDE='test_model$|test_pic_preallocation$|test_framesync$' ;; esac TESTS=$(meson test -C "build-$SAN" --list \ | grep '^libvmaf:' \ | grep -vE "$EXCLUDE" \ | sed 's/^libvmaf://') meson test -C "build-$SAN" --print-errorlogs $TESTS
CodeQL bulk mechanical sweep — Python tree (2026-05-09)¶
- Why this matters on rebase: no rebase impact. The diff lives entirely in
python/vmaf/and one fork-local helper (core/src/vulkan/spv_embed.py). None of the touched Python modules have been changed by Netflix upstream in over four years; the closest churn is unrelated additions topython/vmaf/script/run_*.pydriver flags. A future/sync-upstreamwill land on a clean tree. - What changed: dead imports removed;
exit()→sys.exit()in seven CLI driver scripts;open(...)→with open(...)inpython/vmaf/tools/decorator.pyandcore/src/vulkan/spv_embed.py; typedexcept KeyError: passbodies got an explanatory one-line comment to satisfypy/empty-except;passremoved where it was a no-op tail statement; one commented-out debug block deleted fromtools/misc.py. - Re-test on rebase:
python3 -c "import ast; [ast.parse(open(f).read()) for f in (...)]"over the touched files;ruff checkover the same set must produce no NEW errors versus master baseline.
0345 — cambi × {CUDA, SYCL, HIP} GPU port planning (ADR-0345, docs-only)¶
- Touches:
docs/research/0091-cambi-gpu-port-planning-2026-05-09.md(new),docs/adr/0345-cambi-gpu-port-strategy.md(new),docs/adr/_index_fragments/0345-cambi-gpu-port-strategy.md(new fragment),docs/adr/_index_fragments/_order.txt(append slot),changelog.d/changed/cambi-gpu-planning-digest.md(new). No code. Companion to the per-port PRs that follow per the digest's §6 ordered plan (CUDA → SYCL → HIP). - Upstream source: none — fork-local planning artefact. Netflix/vmaf upstream has no CUDA / SYCL / HIP cambi twin and no plans to add one on those backends.
- Invariant: the planning round locks Strategy II host-staged hybrid for the three pending backends, inheriting verbatim from ADR-0205 §Decision and ADR-0210 §Decision. The cross-backend gate contract for cambi is
places=4from day one on all backends — by construction (integer-only GPU pre-passes; byte-identical readback; unmodified host residual). If any per-port PR sees empirical drift from CPU, fix the kernel — never relax the gate (memoryfeedback_no_test_weakening). The sharedcambi_internal.hhost residual surface (shipped with PR #196 for the Vulkan port) is the load-bearing reuse point — all four GPU twins (Vulkan, CUDA, SYCL, HIP) link against it and inherit any future CPU-side c-value formula change automatically. - On upstream sync: no action required. If a future upstream sync introduces a Netflix/vmaf cambi GPU twin (extremely unlikely — Netflix has no public CUDA / SYCL / HIP cambi work), evaluate whether to drop the fork's twin in favour of upstream's per the standard prefer-upstream rule; otherwise no action.
- Re-test on rebase: docs-only — no compile / runtime gate. The Strategy III v2 follow-up (parked per ADR-0205 §Out of scope) gets its own ADR + rebase-notes entry when profile data lands.
0320 — Vulkan VIF API-1.4 NVIDIA residual Phase 3b (deferral)¶
- Touches:
core/src/feature/vulkan/shaders/vif.comp(comment-only update at the Phase-4 reduction site — documents the Phase-3b candidate-fix experiments and the driver-side hypothesis; no code logic change vs. PR #511);docs/adr/0269-vif-ciede-precise-step-a.md(appended Phase-3b status update appendix; ADR body remains frozen per ADR-0028);docs/research/0090-...md(new);docs/state.md(rowT-VK-VIF-1.4-RESIDUAL-ARCretired in favour ofT-VK-VIF-1.4-RESIDUAL-NVIDIA-DEFERREDafter the hardware-mapping correction);core/src/vulkan/AGENTS.md(Phase 3b update + rebase invariant for cross-backend gate device-name selection);changelog.d/fixed/vif-arc-mesa-anv-int64-reduction.md(new fragment). - Invariant: the workgroup-scope
memoryBarrierShared(); barrier();pair PR #511 introduced is load-bearing for the Arc + RADV lanes at API 1.4 and stays. Phase 3b confirmed it cannot be downgraded back to a barebarrier()even if the NVIDIA residual ever closes — Arc's clean state is contingent on the workgroup-scope pair. - Cross-backend gate device-selection invariant (NEW): scripts that target a specific Vulkan vendor must select by
deviceNamesubstring, not by--vulkan_device <index>.vmaf_vulkan_context_new's device sort is stable inside the samedevtype_scorebucket and thevkEnumeratePhysicalDevicesenumeration order is host-policy-dependent (driver registration order in/etc/vulkan/icd.d/, Mesa device-select layer,VK_LOADER_*env vars). PR #511's commit message inverted the device map on this fork's CI workstation; the empirical numbers it cited as "NVIDIA" actually came from Arc and vice versa. New cross-backend lanes targeting a specific vendor should not inherit the off-by-one. - On upstream sync:
vif.compis fork-local; no upstream Netflix/vmaf has a Vulkan path. Cherry-picks from upstream cannot reach this file. - Re-test on rebase (assumes a multi-GPU CI workstation with NVIDIA + Arc + RADV; lavapipe-only CI lanes are a no-op for the API-1.4 residual since lavapipe never reproduced the bug):
# Local API-1.4 bump (off-master reproducer; do NOT commit).
sed -i 's/VK_API_VERSION_1_3/VK_API_VERSION_1_4/g' \ core/src/vulkan/common.c sed -i 's/VMA_VULKAN_VERSION 1003000/VMA_VULKAN_VERSION 1004000/' \ core/src/vulkan/vma_impl.cpp cd libvmaf && meson setup build -Denable_vulkan=enabled \ -Denable_cuda=false -Denable_sycl=false && ninja -C build cd ..
# NVIDIA lane — expected 45/48 FAIL scale 2 until either the
# manual int64 subgroup-reduction patch lands or NVIDIA fixes
# the driver. Arc + RADV expected 0/48.
python3 scripts/ci/cross_backend_vif_diff.py \ --vmaf-binary core/build/tools/vmaf \ --reference testdata/ref_576x324_48f.yuv \ --distorted testdata/dis_576x324_48f.yuv \ --width 576 --height 324 \ --feature vif --backend vulkan --device
# Revert local bump after testing.
sed -i 's/VK_API_VERSION_1_4/VK_API_VERSION_1_3/g' \ core/src/vulkan/common.c sed -i 's/VMA_VULKAN_VERSION 1004000/VMA_VULKAN_VERSION 1003000/' \ core/src/vulkan/vma_impl.cpp
Upstream-port-later batch — Research-0090 18-commit triage close-out (2026-05-09)¶
- Touches:
docs/state.md(one row in "Deferred (waiting on external trigger)"), this file,changelog.d/changed/upstream-port-later-batch-2026-05-09.md. No code touched. Companion to PR #446 (Research-0090) and the in-flight PRs #497 (MyTestCase super-PR), #443 / #444 (cambi-docs duplicate pair). - Per-commit classification (input set: 18 PORT_LATER SHAs from Research-0090):
| # | Upstream SHA | Subject (truncated) | Verdict | Reopen / forward path |
|---|---|---|---|---|
| 1 | 38e905d1 | adopt MyTestCase + reformat BD-rate test data | PORT_DEFERRED | Subsumed by PR #497 commit e1dbdc09; close out when #497 merges |
| 2 | 005988ea | adopt MyTestCase + port new tests + align fifo_mode | PORT_DEFERRED | Subsumed by PR #497 commit 6c05afe2; close out when #497 merges |
| 3 | 4679db83 | fix VMAFEXEC_score tolerances for macOS FP precision | PORT_DEFERRED w/ Netflix-golden guard | PR #497 commit 0004d2cf — must preserve fork's golden places= values byte-for-byte (CLAUDE §8 / ADR-0024) |
| 4 | 3e075107 | adopt MyTestCase + update score values in vmafexec tests | PORT_DEFERRED | Subsumed by PR #497 commit 0004d2cf; close out when #497 merges |
| 5 | e3827e4d | adopt MyTestCase + port new tests in asset/bootstrap/local_explainer | PORT_DEFERRED | Subsumed by PR #497 commit 6c05afe2; close out when #497 merges |
| 6 | 25ff9f18 | remove empty VmafossexecCommandLineTest stub | PORT_DEFERRED → CHERRY-PICK after #497 | Pure 13-line deletion. PR #497 currently RE-EMITS the stub; once #497 lands, cherry-pick this commit standalone (zero-conflict against post-#497 tip). |
| 7 | 3a041a97 | adopt MyTestCase + update score values | PORT_DEFERRED | Subsumed by PR #497 commit d52d9221; close out when #497 merges |
| 8 | ead2d12b | fix vif_scale3 + adm3_egl_1 tolerances for macOS FP precision | PORT_DEFERRED w/ Netflix-golden guard | PR #497 commit b5a3f61b — Netflix-golden tolerance guard same as row 3 |
| 9 | 6c097fc4 | reduce ADM/VIF tolerances for macOS FP precision | PORT_DEFERRED w/ Netflix-golden guard | PR #497 commit f3881d5c — Netflix-golden tolerance guard same as row 3 |
| 10 | 7df50f3a | align testutil with full set of fixture functions | PORT_DEFERRED | Subsumed by PR #497 commit f1ae0495; close out when #497 merges |
| 11 | 322ca041 | replace temporal slicing with pre-sliced YUV fixtures | PORT_DEFERRED | Subsumed by PR #497 commit 7d9d9a10; close out when #497 merges. Sequencing matters: this commit must land before rows 12, 14, 15, 17 (the YUV-fixture consumers); #497 already orders them correctly. |
| 12 | 74bdce1b | align vmafexec_feature_extractor_test (aim/adm3/motion3) | PORT_DEFERRED | Subsumed by PR #497 commit 07e7cb48; close out when #497 merges |
| 13 | a3776335 | align feature_extractor_test (aim/adm3/motion3) | PORT_DEFERRED | Subsumed by PR #497 commit 15a6874d; close out when #497 merges |
| 14 | 0341f730 | remove duplicate test_run_vmaf_integer_fextractor | PORT_DEFERRED → CHERRY-PICK after #497 | Pure 76-line deletion. Same disposition as row 6 — #497 currently re-emits the duplicate; cherry-pick standalone after #497. |
| 15 | 9fa593eb | port feature_extractor tests for aim/adm3/motion3 + new options | PORT_DEFERRED | Subsumed by PR #497 commit ab21b694; close out when #497 merges |
| 16 | d93495f5 | reduce tolerance for VMAF scores in quality_runner tests | PORT_DEFERRED w/ Netflix-golden guard | PR #497 — Netflix-golden tolerance guard same as row 3 |
| 17 | 7d1ad54b | port feature extractor tests for aim/adm3/motion3 | PORT_DEFERRED | Subsumed by PR #497 commit 44b9e626; close out when #497 merges |
| 18 | 721569bc | resource/doc: cambi_high_res_speedup + motion2 score | PORT_DEFERRED → DEDUP | Already in flight on TWO branches (PR #443 + PR #444). Maintainer picks one and abandons the other per Research-0090 §Recommended action #4. No third port-PR opened. |
- Invariant: after PR #497 merges, the Research-0090 PORT_LATER bucket reduces to exactly two follow-up cherry-picks against post-#497 master:
git cherry-pick 25ff9f18(delete emptyVmafossexecCommandLineTest).git cherry-pick 0341f730(delete duplicatetest_run_vmaf_integer_fextractor). Both are pure deletions onpython/test/command_line_test.pyandpython/test/feature_extractor_test.pyrespectively; no score change, no Netflix-golden interaction. They were excluded from PR #497 because the v2 super-PR's diff state currently RE-EMITS those identifiers (likely because #497 cherry-picked from an earlier upstream tip than25ff9f18/0341f730).- Netflix-golden guard (binding): per CLAUDE §8 / ADR-0024, the three Netflix CPU golden pairs in
python/test/quality_runner_test.py,vmafexec_test.py,vmafexec_feature_extractor_test.py,feature_extractor_test.py,result_test.py(1 normalsrc01_hrc00↔hrc01+ 2 checkerboard) carry hard-codedassertAlmostEqualrows that are NEVER modified by a fork PR. Upstream commits4679db83,ead2d12b,6c097fc4,d93495f5explicitly LOWERplaces=on a subset of those rows (their stated motivation is macOS FP precision drift, not a true score change). Reviewer of PR #497 must verify that the 3 golden pairs retain fork tolerances byte-for-byte; only non-golden rows may adopt the relaxations. - On upstream sync: future
/sync-upstreamruns that re-detect these 18 SHAs should match this entry via the SHA list and short-circuit Pass-2 classification (skip re-triage). - Re-test on rebase: none required at the time of this commit (no code touched); after the two follow-up cherry-picks (
25ff9f18+0341f730) eventually land, run meson test -C build --suite=fast make test-netflix-golden # 3/3 CPU goldens still pass
0356 — Vulkan two-level GPU reduction for VIF / ADM / motion¶
- Touches:
core/src/feature/vulkan/vif_vulkan.c,adm_vulkan.c,motion_vulkan.c,core/src/vulkan/picture_vulkan.{h,c},core/src/vulkan/meson.build,core/src/feature/vulkan/shaders/vif_reduce.comp,adm_reduce.comp,motion_reduce.comp. - Invariant: The
vif_reduce.comp/adm_reduce.comp/motion_reduce.compshaders areACCUM_FIELDS=7/ACCUM_SLOTS=6/ single-field. If an upstream sync adds or removes fields from the per-WG accumulator layout invif.comp/adm.comp/motion.comp, the corresponding reducer shader and the host-sideVIF_ACCUM_FIELDS/ADM_ACCUM_SLOTS_PER_WGconstants must be updated in lockstep. Mismatch = silent miscompute. - Re-test:
meson test -C build --suite=fast+ runscripts/ci/cross-backend-diff.sh --backend=vulkan --places=4against the Netflix normal pair on any available Vulkan device. e notes
Single ledger of fork-local changes that need attention when this fork syncs from upstream/master (Netflix/vmaf). Required by ADR-0108: every fork-local PR that touches upstream-shared paths or establishes a rebase-sensitive invariant adds an entry here. PRs with no rebase impact state "no rebase impact" in the PR description and skip the entry.
The intended reader is whoever runs the next /sync-upstream (see ADR-0002 and .claude/skills/sync-upstream/). Read top-to-bottom before resolving conflicts.
Format¶
Each entry is a ### NNNN — short title heading with three fields:
- Touches: paths likely to conflict on upstream merge.
- Invariant: what the fork relies on that an upstream change could silently drop.
- Re-test: the command(s) to run after the merge to confirm the invariant survived. Reproducer-style — no surrounding prose required.
IDs are assigned in commit order and never reused. A single entry may cover several PRs in one workstream; cross-link from the ID heading.
Entries (backfilled 2026-04-18 per ADR-0108 adoption)¶
0332 — Agent worktree-drift hard guard (ADR-0332)¶
- Touches:
.pre-commit-config.yaml(one newlocalhook id),Makefile(hooks-installtarget — comment-only edit),scripts/ci/check-agent-worktree-drift.sh(new),scripts/ci/test_check_agent_worktree_drift.sh(new),AGENTS.md(new section §12a),docs/development/agent-worktree-discipline.md(new),docs/adr/0332-*.md(new),docs/adr/_index_fragments/(new fragment +_order.txtappend),changelog.d/added/agent-worktree-drift-guard.md(new). - Invariant on upstream-mirror files: none — every touched path is fork-local. The pre-commit hook ID
agent-worktree-drift-guardis unique to this fork; upstream Netflix/vmaf has nolocalhook block in its (also non-existent).pre-commit-config.yaml. - On upstream sync: no expected conflict. The
.pre-commit-config.yamlblock is fork-only; if Netflix ever introduces its own pre-commit config we'll rebase theagent-worktree-drift-guardlocalhook on top of upstream's blocks but the YAML structure is independent. - Re-test on rebase:
```bash bash scripts/ci/test_check_agent_worktree_drift.sh # End-to-end: refused commit from main with active agent. cd "$(git -C . rev-parse --show-toplevel)" && \ bash scripts/ci/check-agent-worktree-drift.sh ; echo "exit=$?" # Allowed commit from inside an agent worktree. cd "$(git -C . rev-parse --show-toplevel)/.claude/worktrees/agent-
0320 — psnr_cuda chroma extension (ADR-0351)¶
- Touches:
core/src/feature/cuda/integer_psnr/psnr_score.cu(kernel — fork-only file, BSD-3-Clause-Plus-Patent / Lusoris+Claude header) andcore/src/feature/cuda/integer_psnr_cuda.c(host glue — fork-only file, same header). Neither path is in upstream Netflix/vmaf. The PTX module key (psnr_score) andnvccextra flags table incore/src/meson.buildare unchanged by this PR. - Invariant: the kernel signature now takes a
planeparameter (unsigned) appended after(width, height). The host file'spsnr_cuda_dispatchpackskernelParams[]in the exact(ref, dis, sse, &width, &height, &plane)order — any refactor that reorders or drops the trailing argument silently breaks the chroma path becausecuLaunchKernelcannot validate argument types. The host also relies on thepicture_cudaupload path having uploaded all 3 planes for non-YUV400Pinputs (seelibvmaf.c::translate_picture_host'supload_mask); a future "minimise upload" optimisation must consult the extractor'scharsto decide which planes can be skipped, not assume luma-only. - Upstream interaction: none — CUDA backend is not in Netflix/vmaf upstream. Cherry-picks from upstream that touch
core/src/feature/integer_psnr.c(the CPU twin) need attention only if they changepsnr_name[],mse_name[], or theenable_chromasemantics; the CUDA path mirrors those conventions byte-for-byte to keep the cross-backend gate clean. - Re-test on rebase:
```bash cd libvmaf && meson setup build -Denable_cuda=true && ninja -C build python ../scripts/ci/cross_backend_vif_diff.py \ --vmaf-binary build/tools/vmaf \ --reference ../testdata/ref_576x324_48f.yuv \ --distorted ../testdata/dis_576x324_48f.yuv \ --width 576 --height 324 --pixel-format 420 --bitdepth 8 \ --feature psnr --backend cuda --places 4
ADR-0994 — coverage build-break fix: remove vmaf_fex_integer_motion_v2 from feature_extractor.cpp + guard motion_five_frame_window in integer_motion.c (2026-06-03)¶
- Touches:
core/src/feature/integer_motion.c,core/src/feature/feature_extractor.cpp,docs/adr/0994-coverage-build-fix-motion-v2-ref.md(new),docs/adr/README.md,docs/state.md,docs/rebase-notes.md,changelog.d/fixed/0994-coverage-build-fix-motion-v2-ref.md(new). - No rebase-sensitive invariants: bug fix only. Both files touched are production C/C++ sources; no option table or ABI surface changed.
- Relation to ADR-0337: when the
prev_prev_refpicture-pool refactor (ADR-0337 deferred hunk) lands, flip the-ENOTSUPguard ininteger_motion.c::init()and restore thefex->prev_prev_refreference inextract(), then remove themotion_five_frame_windowcomment block added here.
ADR-0337 — motion_v2 public option surface duplication (2026-05-09)¶
- Touches:
core/src/feature/integer_motion_v2.c,core/src/feature/x86/motion_v2_avx2.c,core/src/feature/x86/motion_v2_avx512.c,core/src/feature/arm64/motion_v2_neon.c,docs/adr/0337-motion-v2-public-api-options.md(new),docs/adr/_index_fragments/0337-motion-v2-public-api-options.md(new),docs/adr/_index_fragments/_order.txt(one-line append),docs/adr/README.md(regenerated),changelog.d/added/motion-v2-public-api-options.md(new),changelog.d/fixed/motion-v2-mirror-off-by-one.md(new),docs/state.md(deferral row update),core/src/feature/AGENTS.md(invariant note). - Upstream cluster ported: Netflix/vmaf
856d3835(mirror off-by-one fix, propagated to scalar + AVX2 + AVX-512 + fork-local NEON),c17dd898(motion_max_valoption),a2b59b77(motion_five_frame_windowoption, partial — see Deferred-hunks below),4e469601(remaining options +motion3_v2_scoreprovided feature, manual port adapted to 3-frame mode only). - Architectural decision: ADR-0337 picks A1 — duplicate option surfaces between motion v1 and motion_v2. v1 (
integer_motion.c) and v2 (integer_motion_v2.c) each register their ownVmafOption[]table; the seven option names match upstream byte-for-byte so future/sync-upstreamruns find no behavioural delta. The duplication is purely textual (~80 LOC of option-table rows + 7 struct fields). Touching one extractor's help string requires touching the other; ADR-0141 catches drift on the next edit. - Invariants:
motion_five_frame_window=truereturns-ENOTSUPatinit()on motion_v2. Mirrors ADR-0219 §Decision's GPU motion3 precedent. The 3-frame default mode is fully supported. When the picture-pool plumbing follow-up lands, the-ENOTSUPguard flips to aprev_prev_reflookup; until then any caller passing=truesees a hard error.- motion v1's option surface is the source of truth for the seven shared option names. ADR-0158 carries v1's history; ADR-0337 carries v2's. Both extractors emit independently into the feature collector under
VMAF_integer_feature_motion*_score(v1) andVMAF_integer_feature_motion*_v2_score(v2) — there is no shared output namespace. - GPU twins (CUDA / SYCL / HIP / Vulkan) of
motion_v2do NOT yet register the option surface in this PR. Theirmotion3_v2_scoreemission is out of scope per ADR-0337 §Consequences; whether the GPU twins gain the same options follows when a model needs the score there. The mirror off-by-one fix in856d3835is propagated only to scalar + AVX2 + AVX-512 + NEON in this PR; CUDA / SYCL / HIP / Vulkan mirror formulae stay on the pre-fix2*size - idx - 1form and document the divergence. Refresh tracked as a follow-up. - Deferred upstream hunks (kept in upstream commits but not ported in this PR; tracked here so the next
/sync-upstreamfinds them): a2b59b77hunks incore/src/feature/feature_extractor.h(addsVmafPicture prev_prev_refto the per-extractor framework struct),core/src/libvmaf.c(picture-pool sizingn_threads * 2 + 2,prev_prev_refplumbing throughthreaded_extract_func/threaded_extract_batch_func/threaded_read_pictures/threaded_read_pictures_batch,vmaf_closecleanup),core/tools/vmaf.c(CLI picture-pool sizing). Conflicts in 8 regions on the fork'sread_pictures*decomposition (ADR-0152 monotonic-index gate) make an in-PR port unsafe; the picture-pool refactor will land as its own PR with a five-frame-window fixture underpython/test/once the framework changes are reviewed in isolation.a2b59b77's 5-frame branch inextract()(usesfex->prev_prev_refdirectly) lands together with the framework hunks above.4e469601's 5-frameflush()branch (variablestride/min_idx = 2,lo_idx/hi_idxwindow) is collapsed in this PR to themin_idx = 1constant for 3-frame mode only. The branch is structurally oneif (s->motion_five_frame_window)away from full upstream parity; reinstate when theprev_prev_refplumbing lands.- On upstream sync: future
/sync-upstreamagainstNetflix/vmaf masterwill report all four commits as already ported (cherry-picked or hand-ported with(cherry picked from …)trailers). When the picture-pool refactor PR lands, this ledger row gains a "5-frame mode now wired" sub-bullet and the-ENOTSUPguard flips. Avoid re-discovering the four commits as pending — the deferred hunks are tracked above, not the whole commits.
0310 — Vulkan VIF int64 reduction race condition Phase 3 fix¶
- Touches:
core/src/feature/vulkan/shaders/vif.comp(replaces all three barebarrier()calls with explicitmemoryBarrierShared(); barrier();pairs covering the Phase-1 cooperative tile load, the Phase-2 vertical-conv shared write, and the Phase-4 cross-subgroup int64 reduction); plus documentation underdocs/research/0089-...md(Phase 3 status appendix),docs/adr/0269-...md(Phase 3 status appendix),docs/state.md(T-VK-VIF-1.4-RESIDUAL closed; new T-VK-VIF-1.4-RESIDUAL-ARC opened),core/src/vulkan/AGENTS.md(Phase 3 update on the existing invariant row),changelog.d/fixed/vif-int64-reduction-race-condition.md. Upstream Netflix/vmaf has no Vulkan backend, so conflict probability for the shader is zero. The entry exists because the fix is rebase-sensitive: any future cherry-pick that touchesvif.compand downgrades amemoryBarrierShared(); barrier();pair back to a barebarrier()will silently re-introduce the NVIDIA Vulkan 1.4 race. - Invariant:
vif.compshared-memory ordering between cooperative-write phases must be release-acquire, not just a bare workgroup-execution barrier. NVIDIA's Vulkan 1.4 default memory model requires the explicit shared-memory release; barebarrier()works at API 1.3 by accident on this driver. SCALE is irrelevant — the fix applies to all four pipeline specialisations because the barrier sites are in the SCALE-shared code. Do NOT remove the explicitmemoryBarrierShared()calls even if a perf review claims they are redundant under the GLSL spec wording: empirical real-hardware evidence in research-0089 2026-05-09 appendix shows otherwise on NVIDIA driver 595.71.05. - Re-test: apply the local API-1.4 bump (
core/src/vulkan/common.c3 sites +vma_impl.cppVMA_VULKAN_VERSION 1004000) on a NVIDIA RTX 4090 + driver 595.71+ machine, build withmeson setup ... -Denable_vulkan=enabled, then runpython3 scripts/ci/cross_backend_vif_diff.py --feature vif --backend vulkan --device 1 --places 4. Expect 0/48 across all four scales. Run the 5-run determinism check from research-0089 §"Reproduction recipe for Phase 3" against--vulkan_device 1; expect 5 identical(integer_vif_num_scale2, integer_vif_den_scale2) = (+2.494358e+04, +2.522523e+04)pairs at frame 5. Note that--vulkan_device 0on this multi-GPU host is the Intel Arc A380 lane and will still fail at API 1.4 (separateT-VK-VIF-1.4-RESIDUAL-ARCrow Open).
0309 — Vulkan VIF API-1.4 Phase 2 dump (T-VK-VIF-1.4-RESIDUAL)¶
- Touches:
docs/research/0089-vulkan-vif-fp-residual-bisect-2026-05-08.md(2026-05-09 status appendix with empirical numbers from the live RTX 4090),docs/state.md(T-VK-VIF-1.4-RESIDUAL row updated with the localisation),core/src/vulkan/AGENTS.md(new invariant row pinning the SCALE = 2 cross-subgroup-reduction memory-model finding),CHANGELOG.md(lusoris fork "Changed" entry). No code touched; the Phase 3 shader memory-model fix lands in a separate PR. Upstream Netflix/vmaf has no Vulkan backend so conflict probability for the AGENTS.md row is zero — entry exists because the empirical localisation flips the open state-row hypothesis from FP-precision to memory-model and retires theplaces=3override path that earlier rebase scaffolding might have suggested. - Invariant:
vif.compSCALE = 2 specialisation's Phase-4 cross-subgroup int64 reduction is non-deterministic on NVIDIA driver 595.71.05 + Vulkan 1.4.341 (lines 547–592,subgroupAddbarrier()+ thread-0 read ofs_lmem). API 1.3 lane is fully deterministic on the same hardware. The fourapiVersionpinning sites incore/src/vulkan/common.c+core/src/vulkan/vma_impl.cppstay at 1.3 until Phase 3 lands the explicit memory-scope barrier and a 5-run determinism gate confirms run-to-run identical(num, den)plusplaces=40/48 on NVIDIA. Theplaces=3override path is eliminated from the unblock options. - Re-test: apply the local API-1.4 bump (
core/src/vulkan/common.c3 sites +vma_impl.cppVMA_VULKAN_VERSION 1004000) on a NVIDIA RTX 4090 + driver 595.71+ machine, build withmeson setup ... -Denable_vulkan=enabled, then run the gate and the 5-run determinism check from research-0089 §"Reproduction recipe for Phase 3". Expect 45/48places=4failures oninteger_vif_scale2(max abs1.527e-02) AND 5 distinct(integer_vif_num_scale2, integer_vif_den_scale2)pairs across 5 runs of--feature 'vif_vulkan=debug=true'. Both observations reproduced bit-for-bit on this session's hardware lane (UUIDe478b41b-5c4f-1ddb-f990-e44916aff4c8).
0308 — encoder knob-sweep recipe-regression policy (ADR-0308, docs-only)¶
- Touches:
docs/research/0080-encoder-knob-sweep-findings.md,docs/adr/0308-encoder-knob-sweep-recipe-regression-policy.md,docs/adr/README.md(index row),ai/AGENTS.md(knob-sweep invariant section),changelog.d/changed/encoder-knob-sweep-findings.md. No code touched; companion to PR #400 (ADR-0305 + Research-0077 +ai/scripts/analyze_knob_sweep.py). Upstream Netflix/vmaf has no encoder-knob-sweep surface, so conflict probability is zero — this entry exists only because the policy threshold (7-of-9 structural cut) is rebase-sensitive on the corpus shape. - Invariant: the 7-of-9 source-count threshold from ADR-0308 §Decision point 1 is calibrated against the current 9-source Netflix Public Dataset corpus. If the corpus grows past 9 sources (e.g. UGC expansion per ADR-0287, or HDR additions), re-derive the absolute threshold as a fraction (≥7/9 ≈ 78 %). The structural cluster is sharp on the current corpus (top-15 cells all hit 9-of-9, no observed cells in 4-6 range), so a fractional cut at ~75 % is robust. Do NOT relax
bitrate_tol_pct(default 5.0) orvmaf_tol(default 0.1) inai/scripts/analyze_knob_sweep.pywithout an ADR — those tolerances are calibrated against the per-frame VMAF noise floor and bitrate quantisation in libavformat muxers. - Re-test:
pytest ai/tests/test_knob_sweep_analysis.py -v(script logic; ships in PR #400). Policy gate is offline: regenerateruns/phase_a/full_grid/comprehensive.jsonlviatools/vmaf-tune/src/vmaftune/hw_encoder_corpus.py(3-hour run on a single host with NVENC + QSV) then re-runpython ai/scripts/analyze_knob_sweep.py --jsonl <adapted.jsonl> --out-dir runs/phase_a/full_grid/reports/and diff the resultingsummary.mdagainstdocs/research/0080-encoder-knob-sweep-findings.mdheadline table. Structural cluster (top-15 cells, all 9-of-9) is the invariant to defend.
0228 — Vulkan 1.4 bump deferred (ADR-0264, docs-only)¶
- Touches: none (docs-only PR). Future Step A of T-VK-1.4-BUMP will touch
core/src/feature/vulkan/shaders/vif.compandcore/src/feature/vulkan/shaders/ciede.comp; Step B will touch the threeapiVersionsites incore/src/vulkan/common.c(lines 54, 264, 374) and theVMA_VULKAN_VERSIONdefine incore/src/vulkan/vma_impl.cpp(line 22). - Invariant:
masterstays onVK_API_VERSION_1_3andVMA_VULKAN_VERSION = 1003000. Lifting the constant in any future upstream sync (Netflix doesn't ship a Vulkan backend, so the conflict is improbable) without first auditingprecise/OpDecorate ... NoContractiondecoration onvif.compandciede.compwill reintroduce the NVIDIA-driver regression captured in research-0053. Thepsnr_hvs_strict_shaders-O0list incore/src/vulkan/meson.buildis the existing precedent for shader-side bit-exactness mitigations and should be the place a 1.4-era audit lands its results (potentially expanding to covervif.comp+ciede.compif thepreciseaudit decides the optimizer is the right place to gate). - Re-test: when Step B lands, the gate is
python3 scripts/ci/cross_backend_vif_diff.py --feature vif --backend vulkanand the same with--feature ciedeagainst NVIDIA + RADV + lavapipe; max abs diff must stay ≤5.0e-05(places=4) on all three.
0229 — HIP fifth-consumer kernel float_ansnr_hip (ADR-0266)¶
0228 — y4m_convert_411_422jpeg 1-byte heap-buffer-overflow fix¶
0228 — vmaf-tune resolution-aware model selection (ADR-0289)¶
0282 — vmaf-tune AMD AMF codec adapters (ADR-0282)¶
0228 — tools/vmaf-tune/ codec-agnostic encode dispatcher (ADR-0294)¶
- Touches:
tools/vmaf-tune/src/vmaftune/encode.py— refactored to look up the codec adapter and delegate argv composition. Wholly fork-local.tools/vmaf-tune/src/vmaftune/codec_adapters/__init__.py,codec_adapters/x264.py— adapter contract gainsffmpeg_codec_args(preset, quality)andextra_params(). Both are duck-typed; missing methods fall back to the legacy x264-CRF shape.tools/vmaf-tune/tests/test_encode_multi_codec.py— new 19-test suite pinning the dispatcher contract per codec.docs/usage/vmaf-tune.md— new "Codec adapter contract" section.- Invariant: the harness (
encode.py,corpus.py) must not branch on codec identity. The only codec-aware code is the per-adaptercodec_adapters/*.pyfile. Any future change that adds anif adapter.encoder == "..."to the harness regresses ADR-0294's whole-purpose. The corpus row schema stays at SCHEMA_VERSION=1 —crfis preserved as the row column even when the underlying codec's quality knob is-cq/-qp/ etc.;EncodeRequest.qualityis a request-side property only. Adapters that don't yet exposeffmpeg_codec_argsare intentionally permitted to fall back to the legacy x264-CRF shape; removing that fallback would break in-flight adapter PRs landing one-at-a-time. - Re-test on rebase:
```bash pytest tools/vmaf-tune/tests/ -q # 32 passed (13 existing + 19 multi-codec)
python -c " from pathlib import Path from vmaftune.encode import EncodeRequest, build_ffmpeg_command req = EncodeRequest( source=Path('ref.yuv'), width=1920, height=1080, pix_fmt='yuv420p', framerate=24.0, encoder='libx264', preset='medium', crf=23, output=Path('out.mp4'), ) cmd = build_ffmpeg_command(req) assert cmd[cmd.index('-c:v') + 1] == 'libx264' assert cmd[cmd.index('-preset') + 1] == 'medium' assert cmd[cmd.index('-crf') + 1] == '23' print('x264 dispatcher path OK') "
0260 — vmaf-tune --sample-clip-seconds (ADR-0301)¶
- Touches:
tools/vmaf-tune/src/vmaftune/{cli,corpus,encode,score,__init__}.py— fork-local. No upstream Netflix/vmaf path overlap.tools/vmaf-tune/tests/test_corpus.py,tools/vmaf-tune/AGENTS.md,docs/usage/vmaf-tune.md,docs/adr/0301-vmaf-tune-sample-clip.md,docs/adr/_index_fragments/0301-vmaf-tune-sample-clip.md,docs/adr/_index_fragments/_order.txt,docs/adr/README.md.- Invariant: corpus JSONL
SCHEMA_VERSIONbumped to2— additiveclip_modekey only. Sample-clip windows are mirrored on both sides via FFmpeg input-side-ss/-t(encode) and libvmaf's--frame_skip_ref/--frame_cnt(score). The_resolve_sample_clip()helper is the single source of truth for the centre-anchored slice math; do not duplicate the computation elsewhere. Falls back silently to"full"whenN >= duration_s. - Re-test:
0227 — tools/vmaf-tune/ Phase A scaffold (ADR-0237 Phase A)¶
- Touches:
tools/vmaf-tune/src/vmaftune/codec_adapters/{h264_amf,hevc_amf,av1_amf,_amf_common}.py(new). Wholly fork-local — no upstream Netflix/vmaf path overlap.tools/vmaf-tune/src/vmaftune/codec_adapters/__init__.py— registry extended with three AMF entries.tools/vmaf-tune/tests/test_codec_adapter_amf.py(new).tools/vmaf-tune/tests/test_corpus.py— Phase A test renamed fromtest_known_codecs_phase_a_is_x264_onlytotest_known_codecs_includes_x264_and_amf.tools/vmaf-tune/AGENTS.md— adds AMF preset-compression invariant.docs/usage/vmaf-tune.md— adds Hardware encoders section.- Invariant: the 7-into-3 preset compression table in
_amf_common.py(_PRESET_TO_AMF) is the cross-codec axis Phase B / C consumers depend on. Every AMF adapter accepts the canonical 7 preset names (placebo…ultrafast) and maps them onto the three AMF rungs (quality/balanced/speed). Do not extend the preset vocabulary without amending ADR-0282 — registry uniformity (no codec-identity branching in the harness search loop) rests on every codec accepting the same names. - Re-test:
0227 — tools/vmaf-tune/ Phase A scaffold (ADR-0237 Phase A)¶
- Touches:
tools/vmaf-tune/src/vmaftune/resolution.py(new). Wholly fork-local — no upstream Netflix/vmaf path overlap.tools/vmaf-tune/src/vmaftune/corpus.py— addsCorpusOptions.resolution_aware: bool = Trueand pipes the effective model throughscore_res.request.modelinto the JSONL row.tools/vmaf-tune/src/vmaftune/cli.py— adds--resolution-aware/--no-resolution-aware(BooleanOptionalAction, default on).tools/vmaf-tune/tests/test_resolution.py(new).docs/usage/vmaf-tune.md— new "Resolution-aware mode" section.docs/adr/0289-vmaf-tune-resolution-aware.md(new) +docs/research/0064-vmaf-tune-resolution-aware.md(new).tools/vmaf-tune/AGENTS.md— two new invariant notes.- Invariant: the height-only decision rule (
height >= 2160→vmaf_4k_v0.6.1, elsevmaf_v0.6.1) is the documented contract. The JSONLvmaf_modelfield is now per-row (not per-job) — mixed ladder corpora legitimately contain multiple distinct values across rows. Downstream consumers (Phase B / C / D) must group/filter byvmaf_modelrather than assuming a constant. Width is accepted in the API for symmetry but ignored in the body; do not branch on it without a follow-up ADR. - Re-test:
pytest tools/vmaf-tune/tests/ -q
python tools/vmaf-tune/vmaf-tune corpus --help | grep resolution-aware
0227 — tools/vmaf-tune/ Phase A scaffold (ADR-0237 Phase A)¶
- Touches:
core/tools/y4m_input.c— upstream-mirrored Daala-derived Y4M parser. The fix sits inside the 4:1:1 → 4:2:2-jpeg chroma upsample routiney4m_convert_411_422jpeg, lines ~500–530 in the function's three sub-loops. Upstream Netflix/vmaf carries the same shape; if upstream lands its own fix during a sync, prefer the upstream version and drop ours.core/test/test_y4m_411_oob.c(new, fork-local) — drives the minimal W=2 H=4 4:1:1 stream throughvideo_input_open+video_input_fetch_frame. Wholly fork-added; no upstream collision.core/test/meson.build— addstest_y4m_411_oobexecutable +test()registration.- Invariant: the first two sub-loops of
y4m_convert_411_422jpegmust guard_dst[(x << 1) | 1]writes with(x << 1 | 1) < dst_c_w, matching the third sub-loop's existing guard. Without the guard a 4:1:1 stream of width 2 (dst_c_w == 1) writes one byte past the destination chroma row. - Re-test:
cd libvmaf && meson setup ../build-asan --buildtype=debug -Db_sanitize=address -Db_lundef=false -Denable_cuda=false -Denable_sycl=false -Denable_vulkan=disabledninja -C build-asan test/test_y4m_411_oobASAN_OPTIONS=detect_leaks=0 ./build-asan/test/test_y4m_411_oob— must report1 tests run, 1 passed. Pre-fix the binary aborts withAddressSanitizer: heap-buffer-overflow … WRITE of size 1aty4m_input.c:507.
0270 — saliency_student_v1 fork-trained on DUTS-TR (ADR-0286)¶
- Touches:
model/tiny/registry.json— adds thesaliency_student_v1row. Fork-local registry; no upstream overlap.model/tiny/saliency_student_v1.onnx(+.jsonsidecar) — new weights and metadata. Fork-local.ai/scripts/train_saliency_student.py— new training script. Wholly fork-local underai/, which has no upstream counterpart.docs/ai/models/saliency_student_v1.md,docs/research/0062-saliency-student-from-scratch-on-duts.md,docs/adr/0286-saliency-student-fork-trained-on-duts.md— new docs under fork-local trees.- Invariant: the C-side
feature_mobilesal.cextractor's tensor-name contract —input(NCHW[1, 3, H, W]) andsaliency_map(NCHW[1, 1, H, W]) — must continue to match the ONNX graph for bothsaliency_student_v1.onnxand the legacymobilesal.onnxplaceholder. Future weights swaps can change the graph internals freely but must keep these names + shapes; the smoke test asserts the registration. The op-allowlist constraint (graph uses only ops incore/src/dnn/op_allowlist.c) carries over from ADR-0218 —Resizeis not used;ConvTransposeis the upsample op for v1 to keep the graph load-clean against vanilla origin/master. - Re-test:
.venv/bin/python ai/scripts/validate_model_registry.py
.venv/bin/python -c "
from ai.src.vmaf_train.op_allowlist import check_model
from pathlib import Path
r = check_model(Path('model/tiny/saliency_student_v1.onnx'))
assert r.ok, r.pretty()
print('allowlist OK')
"
meson test -C build --suite=fast mobilesal
0227 — tools/vmaf-tune/ Phase A scaffold (ADR-0237 Phase A)¶
- Touches:
core/src/feature/hip/float_ansnr_hip.{c,h}(new) — fifth consumer ofcore/src/hip/kernel_template.h. Mirrorscore/src/feature/cuda/float_ansnr_cuda.ccall-graph-for-call-graph;init/submit/collect/closeinvoke the kernel-template helpers in the same order; the submit body intentionally bypassesvmaf_hip_kernel_submit_pre_launch(no atomic, kernel writes per-block (sig, noise) interleaved float partials directly).core/src/hip/meson.build— adds the new TU tohip_sources.core/src/feature/feature_extractor.c— adds theextern VmafFeatureExtractor vmaf_fex_float_ansnr_hip;declaration and the registry row under#if HAVE_HIP.core/test/test_hip_smoke.c— addstest_float_ansnr_hip_extractor_registeredsub-test pinning the lookup contract.- Invariant — the
submit_pre_launchbypass is load-bearing. The CUDA twin makes the same choice for the same reason. If a future PR adds asubmit_pre_launchcall tofloat_ansnr_cuda.c's submit path, the HIP twin must follow in the same PR. Likewise the readback shape (wg_count * 2u * sizeof(float)) and the bpc table (peak/psnr_max for 8/10/12/16-bit) mirror the CUDA twin verbatim — keep aligned on rebase. - Re-test on rebase:
cd libvmaf
meson setup build -Denable_hip=true -Denable_cuda=false -Denable_sycl=false
ninja -C build
meson test -C build # 48/48 green (47 CPU + HIP smoke)
0230 — HIP sixth-consumer kernel motion_v2_hip (ADR-0267)¶
- Touches:
core/src/feature/hip/integer_motion_v2_hip.{c,h}(new) — sixth consumer ofcore/src/hip/kernel_template.h. Mirrorscore/src/feature/cuda/integer_motion_v2_cuda.ccall-graph-for-call-graph; carries theVMAF_FEATURE_EXTRACTOR_TEMPORALflag and aflush()callback. The state struct has auintptr_t pix[2]ping-pong slot pair tracked outside the kernel-template (the template models a single device+host pair only).core/src/hip/meson.build— adds the new TU tohip_sources.core/src/feature/feature_extractor.c— adds theextern VmafFeatureExtractor vmaf_fex_integer_motion_v2_hip;declaration and the registry row under#if HAVE_HIP.core/test/test_hip_smoke.c— addstest_motion_v2_hip_extractor_registeredsub-test pinning the lookup contract (extractor name ismotion_v2_hip, matching the CUDA twin'smotion_v2_cudanaming).- Invariant — temporal-extractor + ping-pong shape. The
VMAF_FEATURE_EXTRACTOR_TEMPORALflag bit, theflush()callback registration, and theuintptr_t pix[2]slot pair are load-bearing for the runtime PR (T7-10b). The runtime PR will swapuintptr_t pix[2]for a real device-buffer handle pair matching the CUDA twin'sVmafCudaBuffer *pix[2]. On rebase: if the CUDA twin's flush-pass shape changes (currentlymin(score[i], score[i+1])), update the HIP twin'sflush_fex_hipbody in the same PR. - Re-test on rebase: same as 0229 —
meson test -C buildwithenable_hip=trueexercises the smoke contract.
0227 — ms_ssim_vulkan submit-side migrated to kernel_template (T-GPU-DEDUP-26)¶
- Touches:
core/src/feature/vulkan/ms_ssim_vulkan.c—extract()'s rawVkCommandBuffer/VkFence/vkAllocateCommandBuffers/vkBeginCommandBuffer/vkCreateFence/vkQueueSubmit/vkWaitForFences/vkDestroyFence/vkFreeCommandBuffersblocks becomeVmafVulkanKernelSubmittriples (vmaf_vulkan_kernel_submit_begin/_submit_end_and_wait/_submit_free). One triple covers the decimate-pyramid command buffer; one triple per scale covers the per-scale SSIM submit. The pipeline-side bundles (pl_decimate2-binding 4-variant +pl_ssim10-binding 9-variant) and their_add_variant()chains are unchanged from the prior migration.- Invariant: any future submit-side template change (timeline semaphores, deferred fence release, queue-family parameterisation) must keep the helpers' synchronous-wait + per-frame fence + per-frame command-buffer contract intact, since
ms_ssim_vulkan.cdoes host readback of thel_partials/c_partials/s_partialsbuffers immediately after_submit_end_and_waitreturns. The submit-side contract is the same one already documented incore/src/vulkan/AGENTS.md's "Rebase-sensitive invariants" section forkernel_template.h. - Re-test:
```bash cd libvmaf && meson test -C build python scripts/ci/cross_backend_vif_diff.py \ --vmaf-binary core/build/tools/vmaf \ --reference testdata/ref_576x324_48f.yuv \ --distorted testdata/dis_576x324_48f.yuv \ --width 576 --height 324 \ --feature float_ms_ssim --backend vulkan --places 4
0231 — SHA-pin GitHub Actions (OSSF Pinned-Dependencies)¶
- Touches: every workflow file under
.github/workflows/. All 13 fork workflows (docker-image.yml,docs.yml,ffmpeg-integration.yml,libvmaf-build-matrix.yml,lint-and-format.yml,nightly-bisect.yml,nightly.yml,release-please.yml,rule-enforcement.yml,scorecard.yml,security-scans.yml,supply-chain.yml,tests-and-quality-gates.yml) had theiruses:directives rewritten from<owner>/<repo>@vN[.M.K]to<owner>/<repo>@<40-char-sha> # vN.M.K. 97 references converted; the SLSA reusable-workflow ref insupply-chain.ymlis the single documented holdout (seeInvariantbelow). - Invariant — SHA-pin policy for
uses:. Every action reference in.github/workflows/*.ymlMUST be a 40-char commit SHA with the semver tag preserved as a trailing# vN.M.Kcomment. The OSSF ScorecardPinned-Dependenciescheck parses both forms and a floating tag (@vN) is treated as unpinned and counts against the aggregate score. Single permitted exception: the SLSA generator reusable workflow (slsa-framework/slsa-github-generator/.github/workflows/generator_generic_slsa3.yml) must keep itsvX.Y.Ztag form because GitHub Actions consumers cannot SHA-pin reusable-workflow refs in every code path; the exception is documented inline insupply-chain.ymland survives on each rebase. Why this matters on upstream sync: Netflix upstream does not ship the fork's CI tree, so a/sync-upstreamrun that drags new workflow content (e.g. via repository templates or bot-authored bumps) into.github/workflows/can re-introduce floating-tag references unnoticed. The post-rebase check below is the standing gate — anything that lights up needs to be re-pinned before merging the sync. - Re-test on rebase:
# Anything that prints is a regression — every uses: must be either
# already SHA-pinned (40 hex) or, for the documented SLSA exception,
# the slsa-github-generator reusable-workflow ref.
grep -hnE '^\s*(- )?uses:\s+[^@]+@[^ #]+\s*$' .github/workflows/*.yml \
| grep -vE '@[a-f0-9]{40}' \
| grep -v 'slsa-framework/slsa-github-generator/.github/workflows/'
# SHA-resolution sanity for any new pin (per-action):
gh api repos/<owner>/<repo>/git/ref/tags/<vN.M.K> --jq '.object.sha'
# If the result is a "tag" object (annotated tag), deref:
gh api repos/<owner>/<repo>/git/tags/<sha-from-prev> --jq '.object.sha'
0226 — CUDA drain-batch engine-loop opt (T-GPU-OPT-1)¶
- Touches:
core/src/cuda/drain_batch.{h,c}(new) — TLS drain-batch table + shared drain stream +_open()/register/_flush()/_close()API.core/src/libvmaf.c— engine-side per-frame loop now wraps submit/collect with_open()+_flush()so all CUDA extractorfinishedevents are waited on a single shared drain stream.- All 12 CUDA feature kernels (
core/src/feature/cuda/*.c) register theirfinishedevent +drainedflag with the drain batch on submit; collect skips its privatecuStreamSynchronizewhendrainedis true. - Invariant — drained-flag contract. Every CUDA extractor's collect path must check the per-frame
drainedflag and skip its owncuStreamSynchronizewhen set; otherwise the drain batching is a no-op. The flag is reset tofalseper frame insidevmaf_cuda_drain_batch_register(). - Re-test on rebase:
cd libvmaf
meson setup build -Denable_cuda=true -Denable_sycl=false
ninja -C build
meson test -C build --suite=fast cuda
Expected: all CUDA tests green; bench shows ≥5% wall-clock gain on a 7-extractor VMAF model (model.json with all feature extractors enabled).
0225 — Netflix bench snapshot regen (upstream a44e5e61 motion fix)¶
- Touches:
testdata/netflix_benchmark_results.json— fork-added snapshot. CPU rows now reflect the post-fix motion feature; cuda / sycl rows from the previous regen are preserved unchanged because those backends were not exercised on this rerun (host-environment tooling — wrong renderD path,libvmaf_cudanot enabled in the local FFmpeg build). Future full regens should include cuda / sycl.testdata/bench_all.sh— defaultVMAF=no longer points at/usr/local/bin/vmaf(which on most dev hosts is stuck at the pre-upstream-a44e5e61v3.0.0); now defaults to the in-tree fork build atcore/build/tools/vmaf.testdata/benchmark_netflix.py—FFMPEG,YUVDIRand the hardcodedLD_LIBRARY_PATH=/usr/local/libare now overridable viaVMAF_FFMPEG,VMAF_YUVDIRand any caller-setLD_LIBRARY_PATH.- Invariant: the snapshot's CPU pooled VMAF for
src01_576x324is 76.667828 (post-fix), not 76.668904 (the upstream-buggy mirror). If/sync-upstreamever re-pulls a Netflix change that touchesmotion.cmirror-handling, this number is the reference. - Re-test:
cd libvmaf
meson setup build -Denable_cuda=true -Denable_sycl=false
ninja -C build
LD_LIBRARY_PATH=$(pwd)/build/src python3 \
../testdata/benchmark_netflix.py
Expected CPU pooled rows: 76.667828, 35.068672, 7.985899.
0224 — CUDA graph capture feasibility (research-0047, DEFER)¶
- Touches: none — investigation-only; no code lands. The research digest
docs/research/0047-cuda-graph-capture-feasibility.mddocuments why a CUDA graph capture path on the per-frame submit chain is deferred rather than shipped (realised wall-clock gain capped at ~1-3% vs. the predicted 10-20%, with a 4-slot picture-pool rotation that defeats single-graph capture and forces per-framecuGraphExecKernelNodeSetParamsrebinding for(ref, dis)device pointers). - Invariant: the
kernel_template.hdocstring keeps namingVmafCudaKernelLifecycle.finishedas a graph-capture hook point. Don't prune that comment on rebase — leaving the door open in the template is free, and the digest's "what needs to be true for a future GO" section depends on the hook still being there. - Re-test on rebase:
# Confirm the docstring still references graph capture as the hook
# point — wording change is fine, removal is not.
grep -q "graph capture" core/src/cuda/kernel_template.h
0223 — ADR slug-drift repair in CHANGELOG / rebase-notes (PR #304 follow-up)¶
- Touches:
CHANGELOG.md,docs/rebase-notes.md. No code; no upstream-shared path; no public-API surface. - Invariant: every
[ADR-NNNN](docs/adr/NNNN-slug.md)link in the fork's tracked docs resolves to an actual on-disk file underdocs/adr/. Repaired 4 broken slugs that did not exist on disk (0138-iqa-convolve-avx2-bitexact-double→0138-iqa-convolve-avx2-bitexact-double,0140-simd-dx-framework→0140-simd-dx-framework,0190-ms-ssim-vulkan→0190-ms-ssim-vulkan,0178-vulkan-adm-kernel→0178-vulkan-adm-kernel). All retained their cited NNNN per ADR-0028 (NNNN is immutable once Accepted). - Re-test on rebase: from repo root, the following must print no lines:
for ref in $(grep -ohE 'docs/adr/[0-9]{4}-[a-z0-9-]+\.md' \
CHANGELOG.md docs/rebase-notes.md AGENTS.md docs/state.md \
| sort -u); do
test -f "$ref" || echo "MISSING: $ref"
done
0125 — cambi_vulkan migrated to kernel_template (T-GPU-DEDUP-25, 5-bundle)¶
- Touches:
core/src/feature/vulkan/cambi_vulkan.c— state's quintet (dsl_2bind+ 5×pl_layout_*+shader_modules[CAMBI_PL_COUNT]areddesc_pool) collapses to fiveVmafVulkanKernelPipelinebundles (pl_trivial,pl_derivative,pl_filter_mode,pl_decimate,pl_mask_dp), each owning its own descriptor pool. The first slot ofpipelines[]per stage aliases the bundle's base pipeline;CAMBI_PL_FILTER_MODE_V,CAMBI_PL_MASK_SAT_COL, andCAMBI_PL_MASK_THRESHOLDare sibling variants built viavmaf_vulkan_kernel_pipeline_add_variant().cambi_vk_alloc_settakes a bundle pointer (->desc_pool/->dsl) — every dispatch site picks the bundle that matches its push-constant struct.- The
cambi_vk_make_dsl/cambi_vk_make_pl/cambi_vk_create_shader/cambi_vk_build_pipelinehelpers are dropped — the template subsumes them. - Invariant — variants destroyed before bundle, base alias must be skipped. Five distinct push-constant struct sizes (
CambiVkPushTrivial/CambiVkPushDerivative/CambiVkPushFilterMode/CambiVkPushDecimate/CambiVkPushMaskDp) force five bundles even though every stage's DSL is 2-binding SSBO;_add_variant()only siblings pipelines under the same layout.close_fexmustvkDestroyPipeline()the variant slots (CAMBI_PL_FILTER_MODE_V,CAMBI_PL_MASK_SAT_COL,CAMBI_PL_MASK_THRESHOLD) before callingvmaf_vulkan_kernel_pipeline_destroy()on each bundle. - Numerical contract: bit-exact preserved. Same shaders + spec-constants + push-constants as before; only the Vulkan pipeline-bundle scaffolding moved to the template. Validated on the Netflix-pair smoke (576×324×8-bit):
cambimean = 0.0, identical to pre-migration (the pair has no banding artifacts). - Rebase impact: low. Builds on top of PR #272's
_add_variant()helper. Upstream Netflix/vmaf has no Vulkan backend, so there is nothing to merge against.
0124 — ssimulacra2_vulkan migrated to kernel_template (T-GPU-DEDUP-24, 4-bundle)¶
- Touches:
core/src/feature/vulkan/ssimulacra2_vulkan.c— state's 16 long-lived pipeline-object fields (4×*_dsl + *_pl + *_shader+ the shareddesc_pool) collapse to fourVmafVulkanKernelPipelinebundles (pl_xyb,pl_mul,pl_blur,pl_ssim), each owning its own descriptor pool. The first slot of each per-bundle pipeline array (xyb_pipelines[0],mul_pipelines[0],blur_pipelines_h[0],ssim_pipelines[0]) aliases the bundle's baseVkPipeline; remaining per-scale / per-pass slots are siblings viavmaf_vulkan_kernel_pipeline_add_variant().ss2v_build_pipeline_int3reroutes through_add_variant()instead of callingvkCreateComputePipelinesdirectly;ss2v_alloc_settakes a bundle pointer (->desc_pool/->dsl) instead of a separate DSL argument; descriptor-set free sites at the tail ofss2v_run_scaleroute to each bundle's pool.- The
ss2v_make_dsl/ss2v_make_pl/ss2v_create_shaderhelpers are dropped — the template subsumes them. - Invariant — variants destroyed before bundle, slot 0 alias must be skipped. Four distinct DSL shapes (XYB = 6 SSBOs, MUL = 3, BLUR = 2, SSIM = 8) prevent collapsing to one bundle:
_add_variant()only siblings pipelines under the same layout.close_fexmustvkDestroyPipeline()the variant slots inxyb_pipelines[1..N-1],mul_pipelines[1..N-1],ssim_pipelines[1..N-1],blur_pipelines_h[1..N-1], and every slot ofblur_pipelines_v[]before callingvmaf_vulkan_kernel_pipeline_destroy()on each bundle, and must skip slot 0 of the first three arrays +blur_pipelines_hto avoid double-freeing the aliased base. - Numerical contract: bit-exact preserved. Same shaders + spec-constants + push-constants as before; only the Vulkan pipeline-bundle scaffolding moved to the template. Validated on the Netflix-pair smoke (576×324×8-bit):
ssimulacra2mean = 24.613842, identical to pre-migration. - Rebase impact: low. Builds on top of PR #272's
_add_variant()helper. Upstream Netflix/vmaf has no ssimulacra2 extractor and no Vulkan backend, so there is nothing to merge against.
0118 — psnr_hvs_vulkan migrated to kernel_template + _add_variant (T-GPU-DEDUP-18)¶
- Touches:
core/src/feature/vulkan/psnr_hvs_vulkan.c— state'sdsl + pipeline_layout + shader + desc_pool + pipeline[3]collapses toVmafVulkanKernelPipeline pl + VkPipeline pipeline_chroma_u + VkPipeline pipeline_chroma_v. Plane 0 is the template's base pipeline; planes 1+2 are siblings viavmaf_vulkan_kernel_pipeline_add_variant().- New
psnr_hvs_plane_pipeline()accessor maps plane index to the rightVkPipelinehandle. - Invariant — variants destroyed before bundle.
close_fexmustvkDestroyPipeline()the chroma U/V variants before callingvmaf_vulkan_kernel_pipeline_destroy(&s->pl)— same rule as ssim_vulkan in T-GPU-DEDUP-7. - Numerical contract: unchanged. Same shaders + spec-constants push-constants as before; only the Vulkan pipeline-bundle scaffolding moved to the template.
- Rebase impact: low. Builds on top of PR #272's
_add_variant()helper.
0119 — vif_vulkan migrated to kernel_template + _add_variant (T-GPU-DEDUP-19)¶
- Touches:
core/src/feature/vulkan/vif_vulkan.c— state'sdsl + pipeline_layout + shader + desc_pool + pipelines[4]collapses toVmafVulkanKernelPipeline pl + VkPipeline scale_variants[3]. Scale 0 is the template's base pipeline; scales 1, 2, 3 are siblings viavmaf_vulkan_kernel_pipeline_add_variant().- New
vif_scale_pipeline()accessor maps scale index to the rightVkPipelinehandle (replacess->pipelines[scale]). - Invariant — variants destroyed before bundle.
close_fexmustvkDestroyPipeline()the 3 scale variants before callingvmaf_vulkan_kernel_pipeline_destroy(&s->pl)— same rule as ssim_vulkan in T-GPU-DEDUP-7 and psnr_hvs_vulkan in T-GPU-DEDUP-18. - Numerical contract: unchanged. Same shaders, same spec-constants, same push-constants as before; only the Vulkan pipeline-bundle scaffolding moved to the template.
- Rebase impact: low. Builds on top of PR #272's
_add_variant()helper.
0120 — float_vif_vulkan migrated to kernel_template + _add_variant (T-GPU-DEDUP-20)¶
- Touches:
core/src/feature/vulkan/float_vif_vulkan.c— state collapsesdsl + pipeline_layout + shader + desc_pooltoVmafVulkanKernelPipeline pl; theVkPipeline pipelines[2][4]2-D lookup table is preserved so the existing[mode][scale]dispatch path stays clean, butpipelines[0][0]aliasess->pl.pipeline(the template's base). The other 6 entries are sibling pipelines created viavmaf_vulkan_kernel_pipeline_add_variant().- Invariant — variants destroyed before bundle.
close_fexmustvkDestroyPipeline()the 6 sibling variants (every(mode, scale)except(0, 0)) before callingvmaf_vulkan_kernel_pipeline_destroy(&s->pl)— same rule as ssim_vulkan / psnr_hvs_vulkan / vif_vulkan. - Invariant —
pipelines[0][0]aliasing. The base pipeline handle is owned bys->pl.pipeline; we copy it intopipelines[0][0]after_create()so the dispatch path can use a uniform 2-D lookup. The destroy loop must skip(mode=0, scale=0)to avoid double-freeing the template's pipeline. - Numerical contract: unchanged. Same shaders, spec-constants (
mode+scale), push-constants. Netflix-pair smoke matchesinteger_vifbit-identically to 4 decimals. - Rebase impact: low. Builds on top of PR #272's
_add_variant()helper.
0122 — float_adm_vulkan migrated to kernel_template + _add_variant (T-GPU-DEDUP-22)¶
- Touches:
core/src/feature/vulkan/float_adm_vulkan.c— twin to adm_vulkan (T-GPU-DEDUP-21); 16-pipeline 2-D[stage][scale]array. State collapsesdsl + pipeline_layout + shader + desc_pooltoVmafVulkanKernelPipeline pl.pipelines[0][0]aliasess->pl.pipeline; the other 15 entries are siblings viavmaf_vulkan_kernel_pipeline_add_variant().- Invariants:
- Variants destroyed before bundle.
pipelines[0][0]aliasing — destroy loop must skip(stage=0, scale=0).- Numerical contract: unchanged. Same float (
_ssuffix) primitives fromadm_tools.c; same 5-element spec-constant tuple; same float partial accumulation reduced in double on the host.
0121 — adm_vulkan migrated to kernel_template + _add_variant (T-GPU-DEDUP-21)¶
- Touches:
core/src/feature/vulkan/adm_vulkan.c— state collapsesdsl + pipeline_layout + shader + desc_pooltoVmafVulkanKernelPipeline pl; theVkPipeline pipelines[4][4]2-D lookup is preserved so the per-stage dispatch path stays clean.pipelines[0][0]aliasess->pl.pipeline(the template's base); the other 15 entries are sibling pipelines viavmaf_vulkan_kernel_pipeline_add_variant().- Invariants:
- Variants destroyed before bundle (same rule as ssim_vulkan / psnr_hvs / vif / float_vif).
pipelines[0][0]aliasing — destroy loop must skip(stage=0, scale=0)to avoid double-freeing the template's pipeline.- Numerical contract: unchanged. Same shaders + 5-element spec-constant tuple (width, height, bpc, scale, stage) + push-constants.
- Rebase impact: low. Builds on top of PR #272.
0123 — ms_ssim_vulkan 2-bundle migration (T-GPU-DEDUP-23)¶
- Touches:
core/src/feature/vulkan/ms_ssim_vulkan.c— state collapsesdecimate_dsl + decimate_pl + decimate_shader + ssim_dsl + ssim_pl + ssim_shader + desc_pool(7 fields) to two bundlesVmafVulkanKernelPipeline pl_decimate+pl_ssim. Each bundle owns its own descriptor pool. The kernel has two distinct pipeline shapes (decimate = 2 SSBO bindings, ssim = 10 bindings), so two bundles is the minimum —_add_variant()only siblings pipelines under the same layout.decimate_pipelines[0]aliasespl_decimate.pipeline(the template's base = scale 0). The remainingMS_SSIM_SCALES - 2decimate variants (scales 1..3) are siblings via_add_variant().ssim_pipeline_horiz[0]aliasespl_ssim.pipeline(base = scale 0, pass 0). The other 9 entries (4×ssim_pipeline_horizfor scales 1..4, plus 5×ssim_pipeline_vertfor scales 0..4) are variants.- Invariant — variants destroyed before bundle. Same rule as ADR-0106 entry 0106:
close_fexmust destroydecimate_pipelines[1..3]andssim_pipeline_horiz[1..4]+ssim_pipeline_vert[0..4]before callingvmaf_vulkan_kernel_pipeline_destroy()onpl_decimate/pl_ssim. - Invariant —
[0]aliasing destroy-skip.decimate_pipelines[0]andssim_pipeline_horiz[0]must not be passed tovkDestroyPipelineinclose_fex—_destroy()already releases them viapl_decimate.pipeline/pl_ssim.pipeline. Double-free is UB. The destroy loops inclose_fexstart ati = 1for decimate and skipi == 0for ssim_horiz. - Invariant — per-bundle descriptor pool. The shared
s->desc_poolis gone;alloc_descriptor_setnow takes aconst VmafVulkanKernelPipeline *bundleand usesbundle->desc_pool+bundle->dsl. Per-framevkFreeDescriptorSetscalls must target the matching pool (pl_decimate.desc_poolfor decimate sets,pl_ssim.desc_poolfor ssim sets) — mixing them is undefined behavior. - Numerical contract: unchanged. Same shaders, spec constants, push constants, and dispatch order as before.
float_ms_ssimNetflix-pair smoke (576×324×48f) reports mean 0.963241; ssim pyramid intermediate values bit-identical to pre-migration run. - Rebase impact: low. Upstream Netflix has no Vulkan backend. Conflicts only against the parallel
T-GPU-DEDUP-{18..22}PRs (#284–#288) onCHANGELOG.md/docs/rebase-notes.md— auto-resolve keeps both halves.
0106 — Vulkan kernel template multi-pipeline + ssim/motion migration (T-GPU-DEDUP-7)¶
- Touches:
core/src/vulkan/kernel_template.h— newvmaf_vulkan_kernel_pipeline_add_variant()helper. Takes the base pipeline bundle (DSL / pipeline layout / shader / pool owned byvmaf_vulkan_kernel_pipeline_create) plus a partialVkComputePipelineCreateInfoand produces a siblingVkPipelinere-using the same layout / shader. The base_createand_destroyentry points are unchanged; existing consumers (psnr, moment, ciede) keep working.core/src/feature/vulkan/motion_vulkan.c— state collapsesVkPipeline pipelines[2](kept "for SYCL parity" but functionally identical because COMPUTE_SAD goes through push constants, not spec-constants) to a singleVmafVulkanKernelPipeline pl.create_pipelines/close_fexshrink to template-driven create + destroy.core/src/feature/vulkan/ssim_vulkan.c— state becomesVmafVulkanKernelPipeline pl + VkPipeline pipeline_vert. Pass 0 (horizontal) is the template's base pipeline; pass 1 (vertical) is created via_add_variant().close_fexdestroys the variant first, then callsvmaf_vulkan_kernel_pipeline_destroy()on the bundle.- Invariant — no spec-constant drift between base and variant.
_add_variant()overwritessType/stage.sType/stage.stage/stage.module/layoutof the caller'sVkComputePipelineCreateInfoso the variant is guaranteed to share the base's shader and layout. Callers control the variant's spec-constant viapSpecializationInfo. Reordering these overwrites lets a consumer accidentally bind a different shader module under the same layout — UB at descriptor-set time. - Invariant — variant destroyed before bundle.
close_fexin ssim mustvkDestroyPipeline(s->pipeline_vert)beforevmaf_vulkan_kernel_pipeline_destroy(&s->pl)— the bundle's_destroyreleases the descriptor pool, which thevkAllocateDescriptorSetsissued against the variant pipeline's layout cleanly drops only when the variant pipeline is already gone. - Numerical contract: unchanged. Both kernels run identical shaders + spec-constants + push-constants as before; only the Vulkan boilerplate that creates / destroys the pipeline scaffolding moved to a shared owner. Cross-backend parity gate at
places=4holds — Netflix-pairfloat_ssimsmoke (576×324×48f) reports mean 0.863, identical to pre-migration. - Rebase impact: low. The base pipeline-bundle helpers predate this change (PR #270 / #271); the new
_add_variantis additive. Upstream Netflix has no Vulkan backend to conflict with.
0111 — integer_ciede_cuda migrated to kernel_template (T-GPU-DEDUP-11)¶
- Touches:
core/src/feature/cuda/integer_ciede_cuda.c— state'sCUstream + CUevent + CUevent + VmafCudaBuffer + host-pinned float*quintet collapses toVmafCudaKernelLifecycle lc + VmafCudaKernelReadback rb. init / collect / close call the template'slifecycle_init/readback_alloc/collect_wait/lifecycle_close/readback_freehelpers. submit keeps the pre-launch wait inline (intentional — ciede has no atomic, so the template's pre-launch memset is unnecessary).- Numerical contract: unchanged. Pure CUDA-boilerplate consolidation. The host-side reduction in collect still uses the same
doubleaccumulator over per-block float partials —places=4(ADR-0187) holds.
0112 — integer_moment_cuda migrated to kernel_template (T-GPU-DEDUP-12)¶
- Touches:
core/src/feature/cuda/integer_moment_cuda.c— state's stream/event/device-buffer/host-pinned quintet collapses toVmafCudaKernelLifecycle lc + VmafCudaKernelReadback rb. submit callsvmaf_cuda_kernel_submit_pre_launch(atomic counters require the device-side memset). init / collect / close call the matching template helpers.- Numerical contract: unchanged. Same per-frame atomic accumulators (4× uint64), same
sums_host[i] / n_pixelshost division. - Rebase impact: low. Upstream Netflix has no equivalent template; this consolidation is fork-local.
0113 — integer_motion_v2_cuda migrated to kernel_template (T-GPU-DEDUP-13)¶
- Touches:
core/src/feature/cuda/integer_motion_v2_cuda.c— stream/event pair + sad device+host quintet collapses tolc + rb. Raw-pixel ping-pongpix[2]stays outside the bundle. submit keeps the memset onpic_streaminline rather than callingsubmit_pre_launch(the helper would move the memset tolc.str, which races with the kernel reading the accumulator). init / collect / close call the matching template helpers.- Numerical contract: unchanged. Same D2D copy, same conditional kernel launch on frame ≥ 1, same host-side
min(score[i], score[i+1])flush.
0114 — integer_ssim_cuda migrated to kernel_template (T-GPU-DEDUP-14)¶
- Touches:
core/src/feature/cuda/integer_ssim_cuda.c— stream/event/partials device+host quintet collapses tolc + rb. Five intermediate float buffers (h_ref_mu,h_cmp_mu,h_ref_sq,h_cmp_sq,h_refcmp) stay outside the bundle. submit keeps thecuStreamWaitEvent + horiz + vert + DtoHchain inline — SSIM writes one float per block (no atomic), so the template'ssubmit_pre_launchmemset is unnecessary. init / collect / close use the matching template helpers.- Numerical contract: unchanged. Same horiz-then-vert two-pass pipeline, same per-block float partial reduction in double on the host.
places=4(matching the ciede_cuda precision pattern) holds. - Rebase impact: low. Upstream Netflix has no equivalent; this is fork-added.
0115 — ms_ssim_cuda + psnr_hvs_cuda lifecycle migration (T-GPU-DEDUP-15)¶
- Touches:
core/src/feature/cuda/integer_ms_ssim_cuda.c— stream + 2-event lifecycle replaced withVmafCudaKernelLifecycle lc; multi-level pyramid + SSIM intermediate + 3-partials buffers stay outside the template's single-pair readback bundle.core/src/feature/cuda/integer_psnr_hvs_cuda.c— same shape; 3-plane ref/dist/partials triples remain inline.- Numerical contract: unchanged. The migration only affects init / close boilerplate; submit / collect dispatch and host reduction paths are untouched apart from the
s->str→s->lc.str/s->event→s->lc.submit/s->finished→s->lc.finishedfield renames.
0116 — float_psnr/ansnr/motion cuda → kernel_template (T-GPU-DEDUP-16)¶
- Touches:
core/src/feature/cuda/float_psnr_cuda.c— stream/event/partials quintet →lc + rb; input upload buffersref_in/dis_instay outside the bundle.core/src/feature/cuda/float_ansnr_cuda.c— same shape; rb wraps the (sig, noise) interleaved partials.core/src/feature/cuda/float_motion_cuda.c— same shape; rb wraps the SAD partials,blur[2]ping-pong stays outside.- Numerical contract: unchanged. Same dispatch geometry, same reduction order. Cross-backend parity gate at the kernels' contracted precision (places=3 per ADR-0192) holds.
0117 — float_adm + float_vif cuda lifecycle migration (T-GPU-DEDUP-17)¶
- Touches:
core/src/feature/cuda/float_adm_cuda.c— stream + 2-event lifecycle replaced withVmafCudaKernelLifecycle lc; multi-stage DWT + CSF pipeline state stays outside the template's single-pair readback bundle.core/src/feature/cuda/float_vif_cuda.c— same shape; 4-level pyramid + per-scale (num, den) pairs remain inline.- Numerical contract: unchanged. The migration only affects init / close stream-event boilerplate; submit / collect dispatch and host reduction paths are untouched apart from the field renames.
- Rebase impact: low. Upstream Netflix has no equivalent template; this is fork-added.
0107 — float_psnr_vulkan migrated to kernel_template (T-GPU-DEDUP-8)¶
- Touches:
core/src/feature/vulkan/float_psnr_vulkan.c— state'sdsl + pipeline_layout + shader + pipeline + desc_poolquintet is collapsed into a singleVmafVulkanKernelPipeline pl;create_pipelinesandclose_fexshrink to template-driven create + destroy. No shader changes, no spec-constant changes, no push-constant changes.- Numerical contract: unchanged. The migration is a pure Vulkan-boilerplate consolidation. Cross-backend parity gate at
places=4holds — Netflix-pair smoke reportsfloat_psnrmean 30.755 dB, identical to pre-migration.
0109 — float_ansnr_vulkan + motion_v2_vulkan migrated to kernel_template (T-GPU-DEDUP-9)¶
- Touches:
core/src/feature/vulkan/float_ansnr_vulkan.c— single-pipeline state collapses toVmafVulkanKernelPipeline pl;create_pipelinesandclose_fexshrink to template-driven create + destroy.core/src/feature/vulkan/motion_v2_vulkan.c— same shape.- Numerical contract: unchanged. Pure Vulkan-boilerplate consolidation. Cross-backend parity gate at the kernel's contracted precision holds — Netflix-pair smoke reports
float_ansnrmean 23.51 dB andmotion2_v2_scoremean 3.895, identical to pre-migration.
0110 — float_motion_vulkan migrated to kernel_template (T-GPU-DEDUP-10)¶
- Touches:
core/src/feature/vulkan/float_motion_vulkan.c— single-pipeline state collapses toVmafVulkanKernelPipeline pl;create_pipelinesandclose_fexshrink to template-driven create + destroy.- Numerical contract: unchanged. Pure Vulkan-boilerplate consolidation. Netflix-pair smoke reports
motionmean 4.049 /motion2mean 3.894, identical to pre-migration. - Rebase impact: low. Upstream Netflix has no Vulkan backend.
0108 — Bristol VI-Lab feasibility digest + BVI-CC ingest ADR (Draft)¶
- Touches:
docs/research/0046-bristol-vi-lab-feasibility.md(new) — nine-dataset survey + use-case fit + effort estimate.docs/adr/0241-bristol-bvi-cc-ingest.md(new, Status: Draft) — proposal to ingest BVI-CC as the second tiny-AI corpus.docs/adr/README.md— index row for ADR-0241.CHANGELOG.md— Added entry.- Numerical contract: not applicable (docs-only).
- Rebase impact: none. Pure research deliverables; upstream Netflix has no equivalent surface.
0094 — Vulkan VkImage import v2 async pending-fence (T7-29 part 4 / ADR-0251)¶
- ADR: ADR-0251; predecessor ADR-0186.
- Touches:
core/src/vulkan/import.c— full rewrite of the submission path. Single-fencesubmit_and_waitbecomes per-slotsubmit_to_slot+drain_slot_fence; the newslot_alloc/slot_releasehelpers materialise / tear down a ring slot (staging-pair + cmd buffer + fence).vmaf_vulkan_import_imageindexes into the ring byframe_index % ring_size;vmaf_vulkan_wait_computedrains every outstanding fence.vmaf_vulkan_state_build_pictureswaits the slot's fence before exposing the host pointer. Public-API signatures are unchanged.core/src/vulkan/vulkan_internal.h— newstruct VmafVulkanImportSlot;VmafVulkanImportSlotsbecomes a fixed-capacityVmafVulkanImportSlot ring[VMAF_VULKAN_RING_MAX]plus geometry +ring_size. Two new defines —VMAF_VULKAN_RING_DEFAULT(4) andVMAF_VULKAN_RING_MAX(8).VmafVulkanStategainsrequested_ring_size.core/src/vulkan/common.c—vmaf_vulkan_state_initand_state_init_externalsetrequested_ring_size = VMAF_VULKAN_RING_DEFAULT.core/test/test_vulkan_async_pending_fence.c(new, contract smoke for the v1 → v2 swap).core/test/meson.build— registers the new test under the existingenable_vulkanguard.core/src/vulkan/AGENTS.md(new) — pins the three rebase-sensitive ring invariants.docs/adr/0251-vulkan-async-pending-fence.md(new),docs/research/0042-vulkan-async-pending-fence.md(new),docs/api/gpu.md,docs/backends/vulkan/overview.md,CHANGELOG.md,docs/rebase-notes.md.ffmpeg-patches/0006-libvmaf-add-libvmaf-vulkan-filter.patch— unchanged. The v2 ring is fully internal toVmafVulkanState; the public ABI stays byte-identical so the filter consumes the new path transparently.- Invariant 1 — fixed ring depth at first import.
lazy_alloc_ringis the only place that materialises the ring; once allocated the depth never changes for the lifetime of theVmafVulkanState. Any caller that needs a different depth has to free + re-init. The geometry pinning contract from v1 (ADR-0186) is preserved verbatim. - Invariant 2 —
vkResetFencesonly afterVK_SUCCESSfromvkWaitForFences. Sole reset path lives indrain_slot_fence;fence_in_flightflips back to 0 only after the wait succeeds. A-EIOfrom the wait propagates up without resetting (so a retry would correctly re-wait rather than silently move on). - Invariant 3 —
state_freedrains before destroying.vmaf_vulkan_import_slots_freewalks the ring and callsdrain_slot_fenceon every in-flight slot, then issues onevkQueueWaitIdlebelt-and-braces (any feature kernel that submitted on the same queue may still be running). Reordering this triggers validation-layer "destroying in-use object" errors. - Numerical contract: unchanged. Async submission only changes when the host can read the staging buffer, not which bytes the GPU writes. Cross-backend parity gate (
scripts/ci/cross_backend_parity_gate.py,places=4) holds. - Memory delta: staging arena scales
1 → ring_sizeper direction. At default depth and 1080p 8-bit Y, the per-state host-visible footprint grows from ~4 MiB to ~16 MiB. Documented in ADR-0251 §Consequences.
0090 — cambi_vulkan extractor (T7-36 / ADR-0210)¶
- ADR: ADR-0210; predecessor ADR-0205.
- Touches:
core/src/feature/vulkan/cambi_vulkan.c(replaces the spike scaffold'sinit_stub/extract_stub/close_stubtriple with the full Vulkan-aware lifecycle).core/src/feature/vulkan/shaders/cambi_preprocess.comp(new),cambi_mask_dp.comp(new — unified row-SAT / col-SAT / threshold-compare viaPASS=0/1/2spec const).core/src/feature/cambi.c— appends a small block of public trampolines (vmaf_cambi_*) at the bottom of the file that thinly wrap the file-static helpers. No upstream function-static code is renamed or moved; the entire upstream body of cambi.c above the trampolines stays byte-identical, which keeps Netflix sync straightforward.core/src/feature/cambi_internal.h(new) — internal-only header exposingvmaf_cambi_calculate_c_values,vmaf_cambi_get_spatial_mask, etc., to the GPU twin.core/src/vulkan/meson.build— registers the 5 cambi shaders invulkan_shader_sources[]andcambi_vulkan.cinvulkan_sources.core/src/feature/feature_extractor.c— adds the extern decl + registry entry forvmaf_fex_cambi_vulkanunder#if HAVE_VULKAN.scripts/ci/cross_backend_vif_diff.py—cambirow inFEATURE_METRICSso the cross-backend gate runs atplaces=4against the CPU baseline.docs/adr/0210-cambi-vulkan-integration.md,docs/research/0032-cambi-vulkan-integration.md,docs/backends/vulkan.md,CHANGELOG.md.- Invariant 1 — bit-exactness by construction. Every GPU phase is integer arithmetic (
uint16derivative,int32SAT,>compare, stride-2 gather, 3-elementmode3lookup). The readback into the hostVmafPicturepair is byte-identical to what the CPU would have written; the host residual then runs the unmodified CPUcalculate_c_values+ spatial pooling on those buffers. Any rebase that introduces float arithmetic into one of these GPU phases — e.g., a future Netflix change to the derivative kernel that adds a bilinear interpolation step — will silently breakplaces=4and must be caught at the cross-backend gate. - Invariant 2 —
cambi_internal.hsignatures must stay in lock-step with cambi.c's file-static helpers. The Vulkan twin callsvmaf_cambi_calculate_c_values, which trampolines to the file-staticcalculate_c_values. Any signature change to the latter (extra parameters, type changes) must update the trampoline + header in the same PR or the GPU build breaks. - On upstream sync: cambi.c's file-static helpers are sometimes renamed by upstream (e.g.,
decimate→cambi_decimatewould happen during a Netflix tidy-up). When rebasing, search cambi.c's tail for the trampoline block — its fivestaticcalls (get_spatial_mask,decimate,filter_mode,calculate_c_values,spatial_pooling,weight_scores_per_scale,get_pixels_in_window,increment_range,decrement_range,get_derivative_data_for_row,cambi_preprocessing) need to match the upstream symbol names. Update the trampoline body if upstream renames; signatures should not need to change because the trampoline already takes the function-pointer-typedef form (VmafRangeUpdateretc.). - Re-test on rebase:
python3 scripts/ci/cross_backend_vif_diff.py --backend vulkan --feature cambi --ref testdata/ref_576x324_48f.yuv --dist testdata/dis_576x324_48f.yuv --width 576 --height 324 --pixel-format 420 --bitdepth 8 --frames 48. Should emitplaces=4 PASSwithmax_abs_diff = 0.0. If it diverges, bisect the GPU phases by reading back individual buffers (image_buf/mask_buf/deriv_buf) and comparing against the CPU's in-placepicplane after the equivalent stage.
The pre-ADR-0108 fork-local PRs are summarised by workstream rather than per-PR. Future PRs add entries individually.
0085 — Upstream c70debb1 partial port (adm_csf + barten_csf tests)¶
- No ADR. Pure upstream cherry-pick per ADR-0108 carve-out ("pure upstream syncs and
port-upstream-commitPRs are exempt"). - Upstream source:
c70debb1(Kyle Swanson, 2026-04-28): "libvmaf/test: port new adm/vif/speed tests". The audit row that flagged the gap is T-NEW-2 in the 2026-04-29 quarterly upstream-backlog re-audit (PR #205). - Touches (additive only):
core/src/feature/adm_csf_tools.h— new header (verbatim from upstream); declares the inlineadm_native_csfhelper (DLM-paper CSF) used by the newtest_adm_csfunit.core/test/test_adm_csf.c— new unit (verbatim from upstream); 2mu_assertcases onadm_native_csf(3, 3.0, 1080, {0, 45}).core/test/test_barten_csf.c— new unit (verbatim from upstream); 23mu_assertcases overbarten_rod_cone_sens,barten_mtf,barten_csf,linear_interpolate,barten_watson_blend_csf(all symbols already on the fork).core/test/meson.build— registers the two new executables + addstest('test_adm_csf', ...)andtest('test_barten_csf', ...).CHANGELOG.mdUnreleased § Changed.- Deliberate scope cuts (the upstream commit's other halves are not portable verbatim):
test_vif_tools.c— depends on upstream symbolsNUM_KERNELSCALES, the 21-entryvalid_kernelscalestable,vif_validate_kernelscale,vif_get_filter_size,vif_get_filter,speed_get_antialias_filter, and a[NUM_KERNELSCALES][5][65]filter table that the fork'svif_filter1d_table_s [11][4][65]does not match. Per Research-0024 Strategy E, the fork deliberately diverges from the upstreamvifruntime-helper chain to preserve the ADR-0138 / 0139 / 0142 / 0143 SIMD bit-exactness contract. Porting this test requires porting the runtime helpers first.test_speed_chroma.c—#includesfeature/speed.cdirectly; the fork has no SpEED extractor (feature/speed.cdoes not exist). Pairs with audit row T-NEW-1 (port the SpEED extractor wholesale, or absorb it into the tiny-AI speed metric).- Invariants (rebase-relevant):
- The new
adm_csf_tools.hheader is wholly additive and does not conflict with the existing forkadm_csf_snon-inline helper inadm_tools.h(different signature, different translation units). - The two new tests do not depend on Netflix golden YUVs — they evaluate the closed-form CSF math directly. No golden-data interaction.
- On upstream sync: a future port of the upstream
vifruntime-helper chain (Research-0024 Strategy A reversal) or the SpEED extractor (T-NEW-1) unlocks the deferred halves of this commit. Until then, fork-sidetest_vif_tools.c/test_speed_chroma.cstay absent. - Re-test on rebase:
meson setup build-cpu libvmaf -Denable_cuda=false -Denable_sycl=false
ninja -C build-cpu test_adm_csf test_barten_csf
meson test -C build-cpu test_adm_csf test_barten_csf
0084 — Embedded MCP server scaffold (T5-2, ADR-0209)¶
- ADR: ADR-0209 (audit-first scaffold) on top of the ADR-0128 governance + Research-0005 design.
- Upstream source: fork-local. Netflix/vmaf has no embedded MCP server (and no plans to add one — the workflow is agent-tooling-specific, well outside upstream's library scope).
- Touches:
core/include/libvmaf/libvmaf_mcp.h— new public header.core/include/core/meson.build— newif get_option('enable_mcp')install branch.core/src/mcp/— new directory:mcp.c(stub TU) +meson.build(exposesmcp_sources+mcp_defines).core/src/meson.build— newis_mcp_enabledguard +subdir('mcp')block;mcp_sourcesthreaded into thelibrary('vmaf', ...)source list alongsidednn_sources.core/test/meson.build— newif get_option('enable_mcp')block wiringtest_mcp_smoke.core/test/test_mcp_smoke.c— new 12-sub-test smoke.core/meson_options.txt— newenable_mcpumbrella + three sub-flags (all defaultfalse).- Invariant: every public entry point in
libvmaf_mcp.h(vmaf_mcp_init/_start_sse/_start_uds/_start_stdio/_stop/_close) returns-ENOSYS(or-EINVALon bad arguments) until the T5-2b runtime PR lands. The smoke pins this contract — a runtime PR that flips a return code without flipping the smoke expectation regresses the gate. - On upstream sync: zero interaction with upstream files. Wholly additive directory + boolean build flags. The
subdir('mcp')insertion incore/src/meson.buildlives next to the existingsubdir('dnn')/ Vulkan blocks; an upstream conflict in that area would be confined to those few lines and is mechanical to resolve. - Re-test on rebase:
meson setup build-cpu libvmaf -Denable_cuda=false -Denable_sycl=false -Denable_mcp=false
ninja -C build-cpu && meson test -C build-cpu # baseline still green
meson setup --reconfigure build-cpu libvmaf -Denable_mcp=true \
-Denable_mcp_sse=true -Denable_mcp_uds=true -Denable_mcp_stdio=true
ninja -C build-cpu
meson test -C build-cpu test_mcp_smoke # 12/12 sub-tests pass
0065 — T7-37 Netflix bench rerun + docs/benchmarks.md TBD fill¶
- No ADR. Empirical fill of pre-existing
TBDcells; no new decision. The bench script fixes that this rerun depends on shipped earlier under PR #169 (libvmaf/AGENTS.md backend-engagement foot-guns), PR #170 (--backend cudaactually engages CUDA), and PR #171 (testdata/bench_all.shuses correct flags). Vulkan header install for SDK consumers is PR #175. - Touches (additive only):
docs/benchmarks.md(everyTBDcell replaced with measured numbers; hardware-profile table updated to theryzen-4090-archost the rerun was performed on; "How to reproduce" section now documents fixture acquisition for the gitignored BBB 4K 200-frame pair).CHANGELOG.mdUnreleased § Changed entry. - Invariants (rebase-relevant): none. The numbers are tied to fork commit
41301496and theryzen-4090-arcprofile; an upstream rebase that changes feature pipelines would invalidate the table but not break parsing. - On upstream sync: zero interaction. Pure docs.
- Re-test on rebase:
bash testdata/bench_all.sh(after a fresh fork build) — confirms the bench script drives every live backend and records each row's emitted metrics-key count. A GPU count collapsing to CPU is a fallback warning to corroborate with pool and throughput; never compare against fixed expected counts.
0050 — float_adm_cuda + float_adm_sycl extractors (ADR-0202)¶
- ADR: ADR-0202
- Touches:
core/src/feature/cuda/float_adm/float_adm_score.cu(new)core/src/feature/cuda/float_adm_cuda.{c,h}(new)core/src/feature/sycl/float_adm_sycl.cpp(new)core/src/meson.build— three changes: (1) newfloat_adm_scoreentry incuda_cu_sources, (2) newcuda_cu_extra_flagsdict that threads--fmad=false+-Xcompiler=-ffp-contract=offinto thefloat_adm_scorefatbin only, (3) new SYCL source insycl_feature_sources.core/src/feature/feature_extractor.c(extern decls + list entries forvmaf_fex_float_adm_cuda/vmaf_fex_float_adm_syclunder#if HAVE_CUDA/#if HAVE_SYCL).- Invariant 1 —
--fmad=falsefor the float_adm fatbin only: the angle-flag dot product (ot_dp = oh*th + ov*tv) and the cube reductions (xa*xa*xa,csf_o*csf_o*csf_o) require IEEE-754 add/mul ordering to match the GLSLprecisequalifier infloat_adm.comp. NVCC's default-fmad=truefuses these and drifts pastplaces=4at scale 3 / adm2. The integer ADM kernels sharecuda_flagsbut useint64accumulators where FMA is irrelevant — keep the FMA-on default for them. - Invariant 2 — parent-LL dimension trap: stage 0 at
scale > 0reads the parent's LL band; the mirror/clamp bounds arescale_w/h[scale](= parent's LL output dims = current scale's input dims), NOTscale_w/h[scale - 1](= parent's full image dims). Bothfloat_adm_cuda.candfloat_adm_sycl.cppcite this inline. Do not "simplify" by using the off-by-one neighbour. - Re-test:
CXX=icpx CC=icx meson setup build-cs -Denable_cuda=true \
-Denable_sycl=true -Denable_vulkan=enabled \
-Denable_float=true \
-Dsycl_compiler=/opt/intel/oneapi/compiler/latest/bin/icpx
ninja -C build-cs
python3 scripts/ci/cross_backend_vif_diff.py \
--vmaf-binary build-cs/tools/vmaf \
--reference python/test/resource/yuv/src01_hrc00_576x324.yuv \
--distorted python/test/resource/yuv/src01_hrc01_576x324.yuv \
--width 576 --height 324 --feature float_adm \
--backend cuda --places 4
# Same with --backend sycl on a host with an SYCL device.
# Both must report 0/N mismatches at places=4.
0049 — float_adm_vulkan extractor (ADR-0199)¶
- ADR: ADR-0199
- Touches:
core/src/feature/vulkan/float_adm_vulkan.c(new)core/src/feature/vulkan/shaders/float_adm.comp(new)core/src/vulkan/meson.build(adds the .comp shader and the new .c source)core/src/feature/feature_extractor.c(extern decl + list entry under#if HAVE_VULKAN)scripts/ci/cross_backend_vif_diff.py(float_admentry inFEATURE_METRICS).github/workflows/tests-and-quality-gates.yml(lavapipefloat_admstep atplaces=4)- Invariant: float_adm GPU port uses the
2 * sup - idx - 1mirror form on both axes — matches both the scalaradm_dwt2_sand the AVX2float_adm_dwt2_avx2, which both consume the samedwt2_src_indices_filt_sindex buffer. This is intentionally different from float_vif's GPU mirror (ADR-0197), which uses-2because float_vif's AVX2 path takes a different code branch. Do not "fix" the asymmetry by analogy with float_vif. - Re-test:
meson setup build-vk -Denable_vulkan=enabled -Denable_cuda=false \
-Denable_sycl=false
ninja -C build-vk
meson test -C build-vk
VK_LOADER_DRIVERS_SELECT='*lvp*' python3 \
scripts/ci/cross_backend_vif_diff.py \
--vmaf-binary build-vk/tools/vmaf \
--reference python/test/resource/yuv/src01_hrc00_576x324.yuv \
--distorted python/test/resource/yuv/src01_hrc01_576x324.yuv \
--width 576 --height 324 --feature float_adm --places 4
0083 — SSIMULACRA 2 Vulkan kernel (ADR-0201)¶
- ADR: ADR-0201
- Upstream source: fork-local. No SSIMULACRA 2 extractor in upstream Netflix/vmaf — fully fork-local feature.
- Touches:
core/src/feature/vulkan/ssimulacra2_vulkan.c(new file).core/src/feature/vulkan/shaders/ssimulacra2_xyb.comp,ssimulacra2_blur.comp,ssimulacra2_mul.comp,ssimulacra2_ssim.comp(4 new shader files).core/src/vulkan/meson.build— added 4 shaders tovulkan_shader_sourcesand 1 source tovulkan_sources; added all 4 ssimulacra2 shaders topsnr_hvs_strict_shaders(the-O0strict-mode list, kept its legacy name).core/src/feature/feature_extractor.c— registeredvmaf_fex_ssimulacra2_vulkanin the Vulkan branch of the extractor list (betweenpsnr_hvs_vulkanand the CUDA block).scripts/ci/cross_backend_vif_diff.py— addedssimulacra2toFEATURE_METRICS.- Rebase impact: low — fully additive, no upstream-shared files modified beyond
feature_extractor.c's registry array (which always grows on every new extractor and is not a rebase pain point). - Verification command:
meson setup core/build-vk-ss2 \
-Denable_vulkan=enabled -Denable_cuda=false -Denable_sycl=false \
libvmaf
ninja -C core/build-vk-ss2 tools/vmaf
python3 scripts/ci/cross_backend_vif_diff.py \
--vmaf-binary core/build-vk-ss2/tools/vmaf \
--reference python/test/resource/yuv/src01_hrc00_576x324.yuv \
--distorted python/test/resource/yuv/src01_hrc01_576x324.yuv \
--width 576 --height 324 \
--feature ssimulacra2 --backend vulkan --places 1
# expected: max_abs_diff ≈ 1.59e-2, 0/48 mismatches at places=1
- Follow-ups:
- CUDA + SYCL twins (batch 3 parts 7b + 7c per ADR-0192).
- Performance follow-up: re-bin multiple rows / columns per WG in the IIR blur (currently
local_size = 1, one row/col per WG for correctness). - Optional: rename
psnr_hvs_strict_shaderstostrict_shadersincore/src/vulkan/meson.build(cosmetic — out of scope for this PR).
0001 — SIMD bit-identical reductions for float ADM¶
- Workstream PRs: #18, commits
24c88a32,f082cfd3. - Touches:
core/src/feature/integer_adm.c,core/src/feature/float_adm.c,core/src/feature/x86/adm_avx2.c,core/src/feature/x86/adm_avx512.c,core/src/feature/arm64/adm_neon.c, upstreampython/test/feature_extractor_test.pytest expectations. - Invariant:
sum_cubeandcsf_den_scaleaccumulate cubed values in double precision (via_mm256_cvtps_pd/_mm512_cvtps_pd) in scalar, AVX2, AVX-512, and NEON. Upstream accumulates in float, which produces ~8e-5 drift between scalar and SIMD. Test expectations were tightened to match the double-precision path; an upstream-side accumulator change would re-introduce the drift and break the tightened assertions. - Re-test:
meson test -C build --suite=fast && python -m pytest python/test/feature_extractor_test.py -k adm.
0002 — CUDA ADM decouple-inline buffer elimination¶
- Workstream PRs: commit
787e3382. - Touches:
core/src/feature/cuda/integer_adm_cuda.cu,core/src/feature/cuda/adm_decouple_inline.cuh(new),core/src/feature/cuda/meson.build. Upstream'sadm_decouple.cuis no longer compiled in the fork. - Invariant: CSF and CM CUDA kernels read
ref/disDWT2 buffers directly and computedecouple_r/decouple_ainline via__device__helpers inadm_decouple_inline.cuh. The 6 intermediate buffers (decouple_r,decouple_a,csf_a× {scale-0 int16, scales 1-3 int32}) and the standaloneadm_decouple.cusource are intentionally removed. ~107 MB GPU memory savings at 4K. An upstream change toadm_decouple.cuwill look orphaned and a literal merge would re-introduce the buffer allocations. - Re-test:
meson setup build -Denable_cuda=true && ninja -C build && meson test -C build --suite=cuda.
0003 — SYCL backend (USM pool / D3D11 import / vmaf_sycl_* API)¶
- Workstream PRs: #33, #35, #5 (initial scaffolding), and the picture-pool deadlock fix that landed via #32.
- Touches:
core/include/libvmaf/libvmaf_sycl.h,core/src/sycl/,core/src/feature/sycl/,core/src/libvmaf.c(SYCL public-API entry points),meson_options.txt(enable_sycl). - Invariant:
vmaf_sycl_preallocate_picturesconstructs a realVmafSyclPicturePoolhonoringVmafSyclPicturePreallocationMethod(NONE/DEVICE/HOST);vmaf_sycl_picture_fetchdispatches to the pool when configured. The whole SYCL tree is fork-local and has no upstream counterpart — upstream changes tocore/src/libvmaf.cnear the SYCL entry-point block are likely to conflict. Picture-pool error paths invmaf_read_pictures(libvmaf.c) mustgoto cleanup;rather thanreturn err;to avoid leaking ref/dist pictures into the live-picture set (closes the always-on-pool deadlock fixed in #32 — see ADR-0104). See ADR-0101, ADR-0103, ADR-0104. - Re-test:
meson setup build -Denable_sycl=true && ninja -C build && meson test -C build --suite=sycl(requires oneAPI / icpx).
0004 — DNN runtime + tiny-AI surfaces¶
- Workstream PRs: #5, #8, #21, #22, #23, #31, #34, plus the pre-numbered DNN feat commits (
9b985946,1e5336d3,d122b721). - Touches:
core/include/libvmaf/dnn.h,core/src/dnn/,core/src/feature/feature_lpips.c,model/tiny/,meson_options.txt(enable_onnxruntime). - Invariant: ordered EP selection (CUDA → DML → CPU) with graceful fallback (ADR-0102);
fp16_iodoes host-side fp32↔fp16 cast on the scoring path;VMAF_TINY_MODEL_DIRenforces a path jail on model load (PR #31); the runtime op-allowlist (PR #21) walks the ONNX graph and rejects unknown ops + bounds Loop/Iftrip_countat 1024 (ADR-0036/0107). DNN tree is fork-local; upstream has no DNN code yet, so conflicts here are unlikely but themeson_options.txtandcore/src/meson.buildblocks near the DNN flag may collide. - Re-test:
meson setup build -Denable_onnxruntime=true && ninja -C build && meson test -C build --suite=dnn.
0005 — --precision CLI flag (IEEE-754 round-trip lossless)¶
- Workstream PRs: commit
c989fbd9. - Touches:
core/tools/vmaf.c,core/tools/cli_parse.c,core/include/libvmaf/libvmaf.h(addedvmaf_write_output_with_format),core/src/output.c. - Invariant: default
--precisionis%.17g(round-trip lossless);legacyopts back into upstream's%.6f; the public C API gainedvmaf_write_output_with_formatand the oldvmaf_write_outputroutes through it with the%.17gdefault. ABI-breaking only if upstream adds a same-named function with a different signature. See ADR-0006. - Re-test:
vmaf -r ref.yuv -d dis.yuv ... --precision=fulland diff against--precision=legacy.
0006 — Netflix golden tests preserved verbatim as required gate¶
- Workstream PRs: across the fork's life; codified in ADR-0024.
- Touches:
python/test/quality_runner_test.py,python/test/vmafexec_test.py,python/test/vmafexec_feature_extractor_test.py,python/test/feature_extractor_test.py,python/test/result_test.py,python/test/resource/yuv/. - Invariant:
assertAlmostEqual(...)golden values in the five upstream Python test files are never modified by this fork. Fork-added tests live in separate files (e.g.python/test/test_precision_flag.py). The CI gate "Netflix CPU golden tests (D24)" is required and blocks merge. Upstream changes to these files are accepted unless they relax the assertions. - Re-test:
make test-netflix-golden.
0007 — Build system (CUDA 13.2, oneAPI 2025.3, MkDocs migration)¶
- Workstream PRs: #7, #17, commit
8a995cb0. - Touches:
meson.build,meson_options.txt, top-levelMakefile,docs/(Sphinx → MkDocs Material migration —docs/conf.pyremoved,mkdocs.ymladded),docs/requirements.txt,Dockerfile.*, distro install scripts underscripts/. - Invariant: image pins are non-conservative (ADR-0027) — CUDA 13.2, oneAPI 2025.3, clang-format 22, black 26 — and ship experimental toolchain flags (
--expt-relaxed-constexpr, etc.) deliberately. An upstream sync that pulls in a Dockerfile change targeted at older CUDA or older oneAPI must not relax the pins. - Re-test:
meson setup build -Denable_cuda=true -Denable_sycl=true && ninja -C build && mkdocs build --strict.
0008 — Workspace / docs / MATLAB / resource-tree relocations¶
- Workstream PRs: codified across ADR-0026, ADR-0029, ADR-0030, ADR-0031, ADR-0032, ADR-0033, ADR-0034, ADR-0038.
- Touches: any path-walk in upstream's CI / scripts / docs that assumes the upstream layout (root-level
workspace/,resource/,matlab/, rootunittestscript, rootpatches/). - Invariant: the fork's layout is
python/vmaf/workspace/,python/vmaf/resource/,python/vmaf/matlab/,scripts/unittest,ffmpeg-patches/only,.github/codeql-config.yml. Upstream moves to a different sub-tree (e.g. a hypotheticaltools/workspace/) need to either be applied via a corresponding fork-side relocation or rejected with a rebase note. - Re-test:
python -m pytest python/test/ -k golden(verifies the resource-tree path works);make test-netflix-golden.
0009 — License headers (Lusoris/Claude on wholly-new files¶
2016–2026 on Netflix files)
- Workstream PRs: commits
c159761d,a185f8ef,0e98c949, codified in ADR-0025 / ADR-0105. - Touches: every wholly-new fork file (notably the SYCL tree and
core/src/dnn/) and every Netflix-touched file (year range2016 → 2016–2026). - Invariant: wholly-new fork files carry
Copyright 2026 Lusoris and Claude (Anthropic)under the same BSD-3-Clause-Plus-Patent license; mixed files use a dual-copyright notice. An upstream commit that resets a Netflix file's year range (e.g. back to2016–2020) must be partially rejected — keep the fork's2016–2026. - Re-test: grep that wholly-new fork files retain the Lusoris/Claude header (
grep -L "Copyright 2026 Lusoris" core/src/sycl/*.cpp— expected to match nothing).
0010 — .claude/ agent scaffolding + ADR tree + AGENTS.md / CLAUDE.md¶
- Workstream PRs: #14, #24, #37, plus continuous additions.
- Touches:
.claude/,AGENTS.md,CLAUDE.md,docs/adr/,.github/PULL_REQUEST_TEMPLATE.md. - Invariant: this whole tree is fork-local and has no upstream counterpart. Upstream additions to
.github/(issue templates, workflows) need to merge cleanly with the fork's existing files rather than replacing them. The ADR tree's IDs ≤ 0099 are backfills; new decisions start at 0100 (ADR-0028 / ADR-0106). - Re-test: visual review of
.github/anddocs/adr/README.mdafter the merge.
Pre-ADR-0108 entries above are the result of a one-shot backfill sweep on 2026-04-18; subsequent fork-local PRs add their own entries inline.
0011 — Nightly bisect-model-quality + fixture cache¶
- Workstream PRs: closes #4; sticky tracker issue #40.
- Touches:
.github/workflows/nightly-bisect.yml,ai/scripts/build_bisect_cache.py,ai/testdata/bisect/{features.parquet, models/*.onnx, README.md},scripts/ci/post-bisect-comment.py,docs/ai/bisect-model-quality.md,docs/adr/0109-nightly-bisect-model-quality.md,docs/research/0001-bisect-model-quality-cache.md,mkdocs.yml(nav). - Invariant: the committed parquet + ONNX bytes under
ai/testdata/bisect/must regenerate byte-identically fromai/scripts/build_bisect_cache.pywith seedsFEATURE_SEED=20260418andMODEL_SEED=20260419. The CI--checkstep asserts this before every bisect run, so any upstream pull that bumpspandas/pyarrow/onnxenough to change the serialiser bytes will fail the workflow until the cache is regenerated and committed. - Re-test:
python ai/scripts/build_bisect_cache.py --check
vmaf-train bisect-model-quality \
ai/testdata/bisect/models/model_*.onnx \
--features ai/testdata/bisect/features.parquet \
--min-plcc 0.85 --input-name input
# Expected: "no regression in this range"; first_bad_index None.
Pure upstream code is not touched, so no Netflix-side conflict vector. Only fork-local files; risk is toolchain drift, not merge conflict.
0012 — Upstream ADM port (Netflix 966be8d5)¶
- Workstream PRs: this PR; ports a single upstream commit.
- Touches:
core/src/feature/integer_adm.{c,h},core/src/feature/x86/adm_avx2.{c,h},core/src/feature/x86/adm_avx512.{c,h},core/src/feature/alias.c,core/src/feature/barten_csf_tools.h(new upstream file). - Invariant: the eight ADM files now mirror upstream's content byte-for-byte (modulo our clang-format-22 pass and the Netflix copyright-year bump on the new header). Future
/sync-upstreamruns can take new upstream ADM commits cleanly. Do not revert to a pre-966be8d5ADM kernel without also reverting the call-site signatures ininteger_compute_adm— upstream extendedi4_adm_cmfrom 8 to 13 args. - Re-test:
ninja -C core/build && meson test -C core/build
core/build/tools/vmaf -r python/test/resource/yuv/src01_hrc00_576x324.yuv \
-d python/test/resource/yuv/src01_hrc01_576x324.yuv \
-w 576 -h 324 -p 420 -b 8 \
--model version=vmaf_v0.6.1 -o /tmp/vmaf-port.json
grep '<metric name="vmaf"' /tmp/vmaf-port.json
# Expected: mean ≈ 76.66890 (golden 76.66890519623612, places=4 OK).
0013 — Upstream motion port (Netflix PR #1486 head 2aab9ef1)¶
- Workstream PRs: this PR; ports upstream PR #1486 (4 commits on top of
966be8d5ADM base, head2aab9ef1). Sister to entry 0012. - Touches:
core/src/feature/integer_motion.{c,h},core/src/feature/motion_blend_tools.h(new upstream file),core/src/feature/x86/motion_avx2.c,core/src/feature/x86/motion_avx512.c,core/src/feature/alias.c(additive:integer_motion3row),python/test/{quality_runner,vmafexec,feature_extractor,vmafexec_feature_extractor}_test.py(golden tolerance updates:places=4→places=2on motion-affected asserts; expected values unchanged). - Invariant: motion files mirror upstream byte-for-byte (modulo our clang-format-22 pass). The
alias.crow forinteger_motion3was inserted surgically to avoid clobbering the AVX-512 ADM registration added by entry 0012; new motion3 metric appears in default VMAF model output but is not standalone-loadable via--feature integer_motion3(sub-feature only). Netflix golden VMAF mean shifts76.668904824→76.667830213(well withinplaces=2tolerance the upstream PR loosened to). Do not revertplaces=4on motion-touching assertions without also reverting the motion code. - Re-test:
ninja -C core/build && meson test -C core/build
core/build/tools/vmaf -r python/test/resource/yuv/src01_hrc00_576x324.yuv \
-d python/test/resource/yuv/src01_hrc01_576x324.yuv \
-w 576 -h 324 -p 420 -b 8 \
--model version=vmaf_v0.6.1 -o /tmp/vmaf-motion-port.json
grep -E '<metric name="vmaf"|integer_motion3' /tmp/vmaf-motion-port.json
# Expected: vmaf mean ≈ 76.66783; integer_motion3 mean ≈ 3.98976.
0014 — Coverage gate overhaul + upstream python/test/ reformat¶
- Workstream PRs: this PR (coverage-gate overhaul + in-tree reformat of upstream-mirror Python tests).
- Touches:
.github/workflows/ci.yml(CPU + GPU coverage jobs:-Dc_args=-fprofile-update=atomic/-Dcpp_args=-fprofile-update=atomic,meson test --num-processes 1,-Denable_dnn=enabled, ORT install step on the CPU coverage job,lcov/geninforeplaced bygcovrwith--json-summary/--xml/--txtoutput, artifact renamecoverage-lcov-{cpu,gpu}→coverage-{cpu,gpu}),scripts/ci/coverage-check.sh(rewritten to parse gcovr JSON viapython3 -c— same CLI signature),core/src/dnn/dnn_api.c+ newcore/src/dnn/dnn_attach_api.c(vmaf_use_tiny_modelcarved out into its own TU so the unit-test binaries — which pull indnn_sourcesforfeature_lpips.cbut never linklibvmaf.c— don't end up with an undefined reference tovmaf_ctx_dnn_attachonceenable_dnn=enabledactivates the real bodies),core/src/dnn/meson.build+core/src/meson.build(newdnn_libvmaf_only_sourceslist wired intolibvmaf.soonly),python/test/{feature_extractor,quality_runner,vmafexec,vmafexec_feature_extractor}_test.py(mechanical Black + isort reformat — no assertion values changed, imports regrouped, line wrapping normalised). - Invariant: coverage CI must keep all five pieces in lockstep — (a)
-fprofile-update=atomiccloses the intra-process counter race on SIMD inner loops (vif_avx2.c:673,motion_avx2, etc.) → negative counts →geninfo/gcovr abort; (b)--num-processes 1closes the inter-process race where multiple parallel test binaries merge their counters into the same.gcdafiles for the sharedlibvmaf.soat process exit (per-thread atomicity does not cover this); (c)gcovrdeduplicates.gcnofiles belonging to the same source compiled into multiple targets — without dedup, lcov sums hits across compilation units and yields impossible100% values (
dnn_api.c — 1176%was the smoking gun on the first attempt that had only (a)+(b)); (d) ORT install +enable_dnn=enabledin the coverage job is what makescore/src/dnn/*.cmeasurable in the first place — without ORT, the DNN tree compiles in stub branches and the 85% per-critical-file gate is meaningless; (e)vmaf_use_tiny_modellives indnn_attach_api.cand is added tolibvmaf.soonly viadnn_libvmaf_only_sources— moving it back intodnn_api.creintroduces thevmaf_ctx_dnn_attachundefined-reference link error intest_feature_extractor/test_lpipswheneverenable_dnn=enabled, since those test binaries pull indnn_sourcesforfeature_lpips.cbut never linklibvmaf.c. Lint scope: upstream-mirror Python tests are linted at the same standard as fork-added code; we accept that/sync-upstreamand/port-upstream-commitwill re-trigger Black/isort failures whenever upstream rewrites these files, and the fix is another in-tree reformat pass — never an exclusion. The fork'spyproject.tomland.pre-commit-config.yamlkeeppython/test/resource/(binary fixtures only) excluded;python/test/*.pyis in scope. See ADR-0110 (race fixes, superseded) and ADR-0111 (gcovr + ORT layer). - Re-test:
# Reproduce coverage path locally (requires gcc + python3-pip):
pip install --user 'gcovr>=8.0'
cd libvmaf
meson setup build-cov-test --buildtype=debug -Db_coverage=true \
-Denable_avx512=true -Denable_float=true -Denable_dnn=disabled \
-Dc_args=-fprofile-update=atomic -Dcpp_args=-fprofile-update=atomic
ninja -C build-cov-test
meson test -C build-cov-test --print-errorlogs --num-processes 1
~/.local/bin/gcovr --root .. \
--filter 'src/.*' \
--exclude '.*/test/.*' --exclude '.*/tests/.*' \
--exclude '.*/subprojects/.*' \
--gcov-ignore-parse-errors=negative_hits.warn \
--gcov-ignore-parse-errors=suspicious_hits.warn \
--print-summary --txt build-cov-test/coverage.txt \
--json-summary build-cov-test/coverage.json \
build-cov-test
grep -E 'dnn_api|model_loader' build-cov-test/coverage.txt
# Expected: gcovr completes without "Unexpected negative count" AND no
# per-file percentages exceed 100% (drop --num-processes 1 to reproduce
# the multi-process .gcda merge race; switch back to lcov to reproduce
# the dnn_api.c — 1176% over-count from compilation-unit summation).
# Lint smoke test for upstream-mirror tree:
pre-commit run --files python/test/quality_runner_test.py
# Expected: Black/isort/Ruff all PASS — files are reformatted in-tree
# to fork style and stay clean until the next upstream sync.
0015 — Tox doctest collection skips vmaf/resource/¶
- Workstream PRs: this PR (
fix(ci): skip pytest doctest collection of vmaf/resource/ data files). Surfaced once ADR-0115 consolidated CI triggers tomasterand tox actually started running on PRs. - Touches:
python/tox.ini(single-line--ignore=vmaf/resourceadded to the pytest invocation, plus an explanatory comment block). Pure fork-local; no upstream Python file changes. - Invariant:
pytest --doctest-modulesmust not attempt to import files underpython/vmaf/resource/. Those are parameter / dataset / example-config.pyfiles; several have dots in their stems (e.g.vmaf_v7.2_bootstrap.py) that make them unimportable as Python modules. None carry doctests, so the ignore is correctness rather than a workaround. Do not drop the--ignore=vmaf/resourceflag without first verifying every file under that directory has been renamed to a dot-free stem and is importable. - Re-test:
cd python && tox -e py311 -- --collect-only --doctest-modules \
--ignore=vmaf/resource 2>&1 | grep -c "ERROR collecting vmaf/resource"
# Expected: 0 (was 5 before the fix).
Pure upstream code is not touched, so no Netflix-side conflict vector. Risk is upstream renaming or removing files under python/vmaf/resource/ such that the directory disappears, in which case the --ignore becomes a harmless no-op.
0016 — SYCL -fsycl link-arg gated on icpx CXX¶
- Workstream PRs: this PR (
fix(libvmaf): gate -fsycl link arg on icpx CXX, allow gcc/clang host linker). Surfaced once ADR-0115's CI consolidation added an Ubuntu SYCL job to PR-time CI that usesCXX=g++(host linker) with sidecar icpx for SYCL .cpp compilation. - Touches:
core/src/meson.build(thevmaf_link_argsblock immediately after theis_sycl_enabledflag handling — currently ~lines 696-712). Pure fork-local; no upstream Meson file changes expected. - Invariant:
-fsyclis appended tovmaf_link_argsonly whenmeson.get_compiler('cpp').get_id() == 'intel-llvm'(icpx). Rationale: the documented project mode (see comment nearis_sycl_enabledblock at top ofsrc/meson.build) compiles SYCL.cppfiles viacustom_targetwith icpx, while the project's CXX driver may be gcc / clang / msvc; in that mode the SPIR-V device code is already embedded in the icpx-compiled.ofiles at compile time, and the runtime libraries (libsycl+libsvml+libirc+libze_loader) declared as link dependencies resolve every symbol. Passing-fsyclto a non-icpx linker is a hard error (g++: error: unrecognized command-line option '-fsycl'). Do not remove thecpp.get_id() == 'intel-llvm'guard without first verifying every CI matrix leg uses icpx as the project CXX. - Re-test:
meson setup build -Denable_sycl=true \
-Dcpp_link_args=-Wl,--no-undefined
ninja -C build src/libvmaf.so.3
# Expected: link succeeds; no `-fsycl` errors with gcc/clang host CXX.
Pure fork-local guard; no Netflix-side conflict vector.
0017 — CLI precision default %.6f (Netflix-compat) + frame-skip unref¶
- Workstream PRs: this PR (
fix(cli): revert precision default to %.6f and unref skipped frames). Reverts the default flipped by commitc989fbd9(ADR-0006) per ADR-0119. Companion fix incore/tools/vmaf.cresolves the picture-pool exhaustion in the--frame_skip_ref/distloops surfaced once the always-on picture pool (ADR-0104) made unref'ing skipped pictures mandatory. - Touches:
core/tools/cli_parse.c(VMAF_DEFAULT_PRECISION_FMT+VMAF_LOSSLESS_PRECISION_FMTmacros,resolve_precision_fmt()body,--helptext)core/tools/cli_parse.h(field comments only; struct shape unchanged)core/src/output.c(DEFAULT_SCORE_FORMATmacro)core/tools/vmaf.c(skip loop bodies at thec.frame_skip_ref/c.frame_skip_distfor-loops)python/vmaf/core/result.py(per-frame and aggregate:.6fformatters)python/test/command_line_test.pyis unmodified — Netflix golden assertions stay frozen per CLAUDE.md §8; the binary's output format adapts to them, not the other way around.- Invariant:
vmafCLI default score-output format is%.6f(matches upstream Netflix byte-for-byte).--precision=max|fullselects%.17g(IEEE-754 round-trip lossless).--precision=legacyis a synonym for the default. The library default forvmaf_write_output_with_format(..., score_format=NULL)matches. Skipped frames in the--frame_skip_ref/--frame_skip_distpre-loops arevmaf_picture_unref'd immediately after fetch so the preallocated picture pool is not exhausted before the main scoring loop runs. Do not flip the macros back to%.17gor remove the unrefs without a superseding ADR — both are golden-gate-load-bearing. - Re-test:
ninja -C core/build
python -m pytest python/test/command_line_test.py \
::VmafexecCommandLineTest::test_run_vmafexec \
::VmafexecCommandLineTest::test_run_vmafexec_with_frame_skipping \
::VmafexecCommandLineTest::test_run_vmafexec_with_frame_skipping_unequal \
-v
# Expected: all three PASS in <1 s combined.
Pure fork-local; no Netflix-side conflict vector. If upstream ever changes the default format string, treat their value as the new baseline and reconfirm the golden assertions before adopting.
0018 — FFmpeg patches ship as ordered series.txt¶
- Workstream PRs: this PR (
fix(ci): drop dead sycl trigger + consolidate windows.yml into libvmaf.yml (ADR-0115)). Surfaced once ADR-0115's consolidation routed the docker / FFmpeg-SYCL jobs through the master-targeting CI gate for the first time on this branch — the standalone0003-…sycl…apply broke because it referenced struct fields added by0001-…tiny-model…, the Dockerfile onlyCOPY'd 0003, andffmpeg.ymlreferenced a stale../patches/path. - Touches:
Dockerfile(lines ~86-95 — the FFmpeg patch-apply block),.github/workflows/ffmpeg.yml(theBuild FFmpeg with SYCL patch seriesstep),ffmpeg-patches/000{1,2,3}-*.patch(regenerated via realgit format-patch -3so they carry validindex <sha>..<sha> <mode>lines and committable SHAs). Pure fork-local; no upstream FFmpeg or Netflix file changes. - Invariant: both the Dockerfile and
ffmpeg.ymlwalkffmpeg-patches/series.txtline-by-line and apply each patch viagit applywith apatch -p1fallback. Do not ship a new patch without appending it toseries.txt, and do not reorder existing entries — patch 0003 references LIBVMAFContext fields added by patch 0001, so any out-of-order apply breaks the build at hunk 2 of vf_libvmaf.c. - Two flag-side fixes bundled in the same PR:
--enable-libvmaf-syclis not a valid FFmpeg configure option. Patch 0003 usescheck_pkg_config libvmaf_sycl …auto-detection (matching howlibvmaf_cudais wired) — it never registers the switch. Both Dockerfile and ffmpeg.yml used to pass the flag and configure rejected it withUnknown option "--enable-libvmaf-sycl". SYCL support is now controlled solely by-Denable_sycl=trueat libvmaf build time; FFmpeg picks it up automatically whenlibvmaf-sycl.pcis onPKG_CONFIG_PATH.- The Dockerfile now carries two nvcc-flag ARGs.
NVCC_FLAGS(libvmaf) keeps four-gencodelines plus the experimental--extended-lambda/--expt-relaxed-constexpr/--expt-extended-lambdaflags needed for Thrust/CUB host+device code.FFMPEG_NVCC_FLAGS(FFmpeg) carries a single-gencode arch=compute_75,code=sm_75 -O2— FFmpeg'scheck_nvccrunsnvcc -ptx, which fails withnvcc fatal: Option '--ptx (-ptx)' is not allowed when compiling for multiple GPU architectureson multi-arch input, and--extended-lambdarequires host+device compilation. compute_75 PTX is forward-compatible with all newer GPUs via driver JIT. --enable-libnppis no longer passed to FFmpeg's configure. FFmpeg n8.1's libnpp probe carries an explicitdie "ERROR: libnpp support is deprecated, version 13.0 and up are not supported"(configure:7335-7336) that fires on the base image's CUDA 13.2 libnpp. We don't use scale_npp / transpose_npp / sharpen_npp in any VMAF workflow; cuvid + nvdec + nvenc + libvmaf-cuda is the actual GPU path. Revisit once we move to an FFmpeg release that supports CUDA 13 libnpp upstream.- Patch 0002 (
add-vmaf_pre-filter) gained a missing#include "libavutil/imgutils.h"forav_image_copy_plane(). FFmpeg's libavfilter Makefile builds with-Werror=implicit-function-declarationso this fired during the actual compile (not configure). Caught by a localdocker buildrather than waiting for GitHub Actions — much faster iteration loop. - Re-test:
cd /tmp && rm -rf ffmpeg-test && \
git clone -q --depth 1 -b n8.1 \
https://git.ffmpeg.org/ffmpeg.git ffmpeg-test && \
cd ffmpeg-test && \
while IFS= read -r line; do \
case "$line" in ''|\#*) continue ;; esac; \
git apply "/path/to/vmaf/ffmpeg-patches/$line" \
|| patch -p1 < "/path/to/vmaf/ffmpeg-patches/$line"; \
done < /path/to/vmaf/ffmpeg-patches/series.txt
# Expected: all three patches apply with no rejects; the resulting
# tree compiles with --enable-libvmaf. SYCL is auto-detected via
# check_pkg_config (patch 0003), so no explicit configure flag is
# required when libvmaf-sycl.pc is on PKG_CONFIG_PATH.
Pure fork-local series; no Netflix-side conflict vector. See ADR-0118.
0019 — Coverage Gate annotations: upload-artifact v7 + gcovr filter¶
- Workstream PRs: this PR.
- Touches:
.github/workflows/ci.yml(CPU + GPU coverage steps: gcovr stderr piped throughgrep -vE 'Ignoring (suspicious|negative) hits' ... || true),.github/workflows/{ci,lint,nightly,nightly-bisect,supply-chain,libvmaf}.yml(actions/upload-artifact@v5|@v6 → @v7,actions/download-artifact@v5 → @v7insupply-chain.yml). Note:windows.ymlwas consolidated intolibvmaf.ymlby ADR-0115 / PR #50, so the windows-side bump now lives inlibvmaf.yml'sbuild (MINGW64, …)job. - Invariant: Coverage Gate Annotations panel must finish empty on a clean run. The two pieces are coordinated — (a)
@v7for upload / download artifact actions silences GitHub's Node-20 deprecation banner ahead of the 2026-06-02 forced-Node-24 cutoff; (b) the gcovr stderr filter swallows theIgnoring (suspicious|negative) hitswarnings that gcovr 8 emits for the legitimately-large hit counts in tight ANSNR / VIF / motion inner loops (e.g.ansnr_tools.c:207at ~4.93 G hits across an HD multi-frame coverage suite — real, not gcov bug). The filter is regex-narrow and anchored to gcov's exact warning prefix; any other gcovr warning still surfaces. Upstream (Netflix/vmaf) does not maintain these CI files; rebase impact is limited to the unlikely case that an upstream sync touches the shared.github/workflows/tree, which it currently does not. See ADR-0117. - Re-test:
# Verify gcovr filter locally (after a coverage build per entry 0014):
~/.local/bin/gcovr --root .. \
--filter 'src/.*' \
--exclude '.*/test/.*' --exclude '.*/tests/.*' \
--exclude '.*/subprojects/.*' \
--gcov-ignore-parse-errors=negative_hits.warn \
--gcov-ignore-parse-errors=suspicious_hits.warn \
--print-summary --txt build-cov-test/coverage.txt \
build-cov-test \
2> >(grep -vE 'Ignoring (suspicious|negative) hits' >&2 || true)
# Expected: stderr contains the gcovr summary block but NO
# "Ignoring (suspicious|negative) hits" lines. coverage.txt unchanged.
# Verify all upload/download-artifact instances are on @v7:
grep -rE 'actions/(upload|download)-artifact@v[0-6]' .github/workflows/
# Expected: empty output.
0020 — CI workflow file + display-name renames (Title Case sweep)¶
- Workstream PRs: this PR; renames all six core
.github/workflows/*.ymlfiles to purpose-descriptive kebab-case and normalises every workflowname:and jobname:to Title Case. See ADR-0116. - Touches:
.github/workflows/{ci,lint,security,libvmaf,ffmpeg,docker}.yml(renamed viagit mvtotests-and-quality-gates.yml,lint-and-format.yml,security-scans.yml,libvmaf-build-matrix.yml,ffmpeg-integration.yml,docker-image.yml),README.md(5 badge URLs + labels),docs/principles.md(line 5 workflow-tuple update),.claude/skills/add-gpu-backend/SKILL.md+scaffold.sh(filename refs),docs/adr/0116-*.md(new),docs/adr/README.md(index row),CHANGELOG.md. - Invariant: workflow files are purpose-named; their
name:fields are Title Case sentences with em-dash axis tags; job-levelname:strings are Title Case sentences (Build — / Pre-Commit / Coverage Gate / etc.). Required-status-check contexts inmasterbranch protection are bound to job-level names — when renaming any job, re-pin viagh api --method PUT repos/VMAFx/vmafx/branches/master/protection. The 19 required gates' semantics are unchanged from ADR-0037; only their display strings move. - Re-test:
# Validate every workflow file parses and lists the expected job names.
cd .github/workflows
for f in tests-and-quality-gates.yml lint-and-format.yml security-scans.yml \
libvmaf-build-matrix.yml ffmpeg-integration.yml docker-image.yml; do
yq '.name, .jobs.[].name' "$f" || echo "PARSE FAIL: $f"
done
# Expected: each workflow prints its Title Case workflow name + job names;
# no PARSE FAIL lines.
0021 — DNN-enabled CI matrix legs (gcc + clang + macOS)¶
- Workstream PRs: this PR; adds three new entries to the
libvmaf-buildmatrix in.github/workflows/libvmaf-build-matrix.ymlcovering-Denable_dnn=enabledacross Ubuntu/gcc, Ubuntu/clang, and macOS/clang. See ADR-0120. - Touches:
.github/workflows/libvmaf-build-matrix.yml(3 new matrix entries + ORT install steps + dedicated dnn-suite test step),docs/adr/0120-ai-enabled-ci-matrix-legs.md(new),docs/adr/README.md(index row),CHANGELOG.md(Added entry). - Invariant: the DNN matrix legs install ONNX Runtime via the same pinned source as the dedicated Tiny AI job (tests-and-quality-gates.yml) — Linux: MS tarball at the version pinned by
ORT_VERSION; macOS: Homebrew. When the Tiny AI job's pin changes, the matrix legs'ORT_VERSIONenv in theirInstall ONNX Runtime (linux, DNN leg)step must change to match; otherwise compiler/portability coverage drifts away from the gating leg's actual ABI. - Re-test:
# Local sanity: the matrix file parses and the new job names exist.
yq '.jobs.libvmaf-build.strategy.matrix.include[] | select(.dnn==true) | .name' \
.github/workflows/libvmaf-build-matrix.yml
# Expected output (3 lines):
# Build — Ubuntu gcc (CPU) + DNN
# Build — Ubuntu clang (CPU) + DNN
# Build — macOS clang (CPU) + DNN
# Local DNN build sanity (matches what each leg will run):
meson setup libvmaf core/build --buildtype release \
--prefix $PWD/install -Denable_float=true -Denable_dnn=enabled
ninja -vC core/build install
meson test -C core/build --suite=dnn --print-errorlogs
- Branch protection: the two Linux DNN legs are pinned as required status checks on
masterimmediately after this PR's merge (19 → 21 contexts). The macOS leg stays informational (experimental: true) because Homebrew ORT floats. Re-pin command:
gh api --method PUT repos/VMAFx/vmafx/branches/master/protection \
--input /tmp/protection-update.json
0022 — Windows GPU build-only matrix legs (MSVC + CUDA, MSVC + oneAPI SYCL)¶
- Workstream PRs: this PR; adds a new top-level
windows-gpu-buildjob to.github/workflows/libvmaf-build-matrix.ymlwith two matrix entries (CUDA, SYCL). See ADR-0121. - Touches:
.github/workflows/libvmaf-build-matrix.yml(newwindows-gpu-buildjob),docs/adr/0121-windows-gpu-build-only-legs.md(new),docs/adr/README.md(index row),CHANGELOG.md(Added entry),core/src/compat/win32/pthread.h(new — Win32 pthread shim for MSVC; mirrorscompat/gcc/stdatomic.hpattern),core/src/feature/integer_adm.h(UPSTREAM — converted thedwt_7_9_YCbCr_threshold[3]designated initializer to positional form so MSVC/nvcc-on-Windows accepts the C++ parse; semantically identical, no behavioural change),core/src/ref.handcore/src/feature/feature_extractor.h(UPSTREAM — added#if defined(__cplusplus) && defined(_MSC_VER)branch around#include <stdatomic.h>so MSVC C++ TUs pullatomic_intviausing std::atomic_int;; POSIX paths unchanged),core/src/sycl/d3d11_import.cpp(fix non-existent<libvmaf/log.h>→"log.h"),core/src/sycl/dmabuf_import.cpp(move<unistd.h>inside#if HAVE_SYCL_DMABUFguard for non-VA-API hosts),core/src/sycl/common.cpp(replace POSIXclock_gettime(CLOCK_MONOTONIC)with portablestd::chrono::steady_clock),core/src/feature/x86/motion_avx2.c(UPSTREAM — replace GCC vector-extension__m256i[N]indexing at line 529 with_mm256_extract_epi64; bit-exact),core/src/feature/x86/adm_avx2.c(UPSTREAM — replace 6(__m256i)(_mm256_cmp_ps(...))casts with_mm256_castps_si256(...)and 12__m128i[N]reductions with_mm_extract_epi64; bit-exact),core/src/feature/x86/adm_avx512.c(UPSTREAM — replace 12__m128i[N]reductions with_mm_extract_epi64; bit-exact),core/src/log.c(UPSTREAM — gate<unistd.h>behind!_WIN32, include<io.h>+ redirectisatty/filenoto_isatty/_filenofor MSVC),core/src/feature/integer_vif.c(UPSTREAM — switch thealigned_malloccursor fromvoid *touint8_t *with explicit typed-pointer casts so MSVC accepts the byte-wise pointer arithmetic),core/src/feature/cuda/integer_adm_cuda.c(UPSTREAM — drop unused<unistd.h>include),core/src/dnn/model_loader.c(fork-added — Windows fallback definitions for POSIXS_ISDIR/S_ISREGpath-classification macros),.github/workflows/lint-and-format.yml(fork-added — setlfs: trueon the pre-commit job's checkout so LFS-stored ONNX blobs resolve and don't appear as phantom pre-commit-induced diffs),core/src/feature/x86/motion_avx512.c(UPSTREAM — replace 1__m128i[N]reduction with_mm_extract_epi64; bit-exact),core/src/feature/x86/{vif_statistic_avx2,ansnr_avx2,ansnr_avx512,float_adm_avx2,float_adm_avx512,float_psnr_avx2,float_psnr_avx512,ssim_avx2,ssim_avx512}.c(UPSTREAM — convert 17 sites of trailing__attribute__((aligned(N)))to leading C11_Alignas(N); same alignment, MSVC-portable),core/src/feature/mkdirp.candcore/src/feature/mkdirp.h(UPSTREAM third-party MIT-licensed micro-library — gate<unistd.h>to non-Windows, add<direct.h>+_mkdirfor Windows, addmode_ttypedef for MSVC),core/meson.build(newpthread_dependencygated oncc.check_header('pthread.h')failing),core/src/meson.buildandcore/test/meson.build(threadpthread_dependencyinto every target compiling pthread-using TUs). - Invariant: Windows GPU legs are pinned to the same toolchain versions as the corresponding Linux GPU legs (CUDA 13.0.0, oneAPI BaseKit 2025.3.0.372) so a Linux-vs-Windows divergence implies an MSVC ABI issue, not a tooling-version delta. When either Linux GPU leg bumps its toolchain, the Windows leg must move in lockstep — the Intel installer URL on Windows hard-codes the per-release directory id and the version string, so the bump is two-line edits in the SYCL
Install Intel oneAPI (windows)step (theWINDOWS_BASEKIT_URLenv var). Both legs additionally inject/experimental:c11atomicsintoCFLAGS/CXXFLAGSbecause libvmaf uses C11 atomics that MSVC's<stdatomic.h>rejects without that opt-in flag — when MSVC ships full C11 atomics support, the flag becomes unconditional and can be dropped. Two Windows-only dependency steps round out the parity: the CUDA leg'sJimver/cuda-toolkitsub-package list includes bothcrt(CUDA Runtime Library compile-time headers, shipscrt/host_config.h;cuda_ccclis not a valid Windows sub-package name — installer rejects it) andnvvm(shipsnvvm/bin/cicc.exe+nvvm/libdevice/libdevice.*.bc; without it, nvcc's.cu → PTXstage fails withThe system cannot find the path specified.— on Linux apt pulls NVVM in transitively withcuda-nvcc-XY, Windows requires it explicitly); the SYCL leg builds the Level Zero loader from source (oneapi-src/level-zerov1.18.5 →cmake --build … --target install) because Windows oneAPI BaseKit ships the SYCL runtime but notze_loader.lib, and libvmaf's mesoncc.find_library('ze_loader')needs both the header and the import library. When the Linux aptlevel-zero-devversion moves, bump the L0 git tag to match.core/src/meson.buildguards the explicitsvml/irccc.find_librarycalls behindhost_machine.system() != 'windows'— those calls exist for the gcc/g++ + icpx Linux flow where the host linker is non-Intel; on Windows the host compiler is icx-cl itself and auto-injects the Intel runtime. Round-10 surfaced an additional Windows-only gap: ~14 libvmaf TUs#include <pthread.h>unconditionally, but MSVC and clang-cl ship no pthread (MinGW does, via winpthreads). The fork now ships a header-only Win32 shim atcore/src/compat/win32/pthread.hmapping the in-use pthread subset (mutex / cond / thread create+join+detach) onto SRWLOCK + CONDITION_VARIABLE +_beginthreadex. The shim is wired in viapthread_dependencyincore/meson.build, declared only whencc.check_header('pthread.h')fails — so MinGW and POSIX paths stay untouched. When upstream Netflix/vmaf adds new pthread surface (e.g.,pthread_rwlock_*), extendcompat/win32/pthread.hto cover it. Both nvcc fatbincustom_targets (CUDA) and icpxcustom_targets (SYCLcommon.cpp/picture_sycl.cpp/dmabuf_import.cpp, plus the SYCL feature kernels) bypass meson'sdependencies:plumbing and hand-roll their own-Ilists, so the shim path must be threaded into bothcuda_extra_includesandsycl_inc_flagsexplicitly on Windows. icpx-cl on Windows additionally rejects-fPIC(unsupported option for target 'x86_64-pc-windows-msvc') — sosycl_common_argsandsycl_feature_argsroute their-fPICtoken throughsycl_pic_arg = host_machine.system() != 'windows' ? ['-fPIC'] : []. PIC is the default for Windows DLLs, so dropping the flag is the correct fix rather than a workaround. Round-14 surfaced a third Windows-only blocker:core/src/feature/integer_adm.h(an upstream Netflix file, last touched by upstream port d06dd6cf) initialisesdwt_7_9_YCbCr_threshold[3]with C99 designated initializers ({.a = ..., .k = ..., .f0 = ..., .g = {...}}). The header is included from bothinteger_adm.c(C TU) andcuda/integer_adm/*.cu(C++ TU via nvcc); MSVC's C++ frontend (and nvcc's cudafe++ on Windows) rejects C99 designated initializers without/std:c++20. Converted to positional initialization in the same struct-member order (a / k / f0 / g[4]) — the conversion is provably semantically identical and works in every C/C++ standard, so it costs nothing on the upstream-merge side beyond a trivial conflict marker if upstream Netflix later edits the same lines. Restore designated form post-merge if upstream has it. Round-17 surfaced four more Windows/MSVC-only SYCL blockers, two of which touch upstream-shared headers. (a)core/src/ref.handcore/src/feature/feature_extractor.h(UPSTREAM) unconditionally#include <stdatomic.h>and use theatomic_inttypedef in struct definitions. MSVC's<stdatomic.h>(added in 19.34) only declares the C11 symbols inside the global namespace under C; in C++ compilation (icpx-cl drives the SYCL TUs as C++) MSVC surfaces them only insidenamespace std::. gcc/clang expose both via a GNU extension, so the upstream code works on every other platform. The fork now wraps both headers'#include <stdatomic.h>in#if defined(__cplusplus) && defined(_MSC_VER)→#include <atomic>+using std::atomic_int;, falling through to the original<stdatomic.h>line on every other configuration. ABI is unchanged —atomic_intresolves to the same underlying type. If upstream Netflix adds further C11 atomic typedefs in these headers (e.g.,atomic_uint,atomic_size_t), extend theusing std::lines to cover them. (b)core/src/sycl/d3d11_import.cpp(fork-added) used<libvmaf/log.h>which doesn't exist —log.hlives atcore/src/log.hand is internal. Switched to"log.h"; the icpx invocation already supplies the src-relative-I. (c)core/src/sycl/dmabuf_import.cpp(fork-added) included<unistd.h>at file scope, but POSIXclose()is only used inside the#if HAVE_SYCL_DMABUFVA-API block. Moved the<unistd.h>include inside that guard so non-DMA-BUF builds (Windows MSVC, macOS) compile cleanly. (d)core/src/sycl/common.cpp(fork-added) calledclock_gettime(CLOCK_MONOTONIC), which doesn't exist on Windows. Replaced withstd::chrono::steady_clock(guaranteed monotonic by the C++ standard, portable on every supported host). All four fixes preserve POSIX/Linux behaviour bit-identically and only change the Windows MSVC build path. Round-18 surfaced a fifth Windows blocker on the CUDA leg's CPU SIMD compile path:core/src/feature/x86/motion_avx2.c:529(UPSTREAM, ported in commit 9371a0aa from Netflix PR #1486) computedfinal_accum[0] + final_accum[1] + final_accum[2] + final_accum[3]to extract the four int64 lanes from an__m256i. gcc/clang allow this via the GNU vector-extension treatment of__m256i(it carries__attribute__((vector_size(32)))); MSVC rejects it withC2088: built-in operator '[' cannot be applied to an operand of type '__m256i'. Replaced with_mm256_extract_epi64(final_accum, N)for N ∈ {0..3}, summed — bit-exact lane sum on every compiler. Restore the index form post-merge if upstream Netflix later edits the same lines and your toolchain matrix doesn't include MSVC. Round-19 surfaced the same MSVC pattern at 19 more call sites across the AVX2/AVX-512 ADM and motion files plus six GCC-style vector casts.core/src/feature/x86/adm_avx2.c(UPSTREAM): 6 lines (915-920) used(__m256i)(_mm256_cmp_ps(...))C-style casts that gcc/clang accept via the GNU vector extension; replaced with the dedicated_mm256_castps_si256(...)bit-cast intrinsic. 12 lane-extract sites (r2_h[0]+r2_h[1], etc. at lines 2420 / 2425 / 2430 / 2893 / 2897 / 2901 / 4079 / 4084 / 4089 / 4627 / 4631 / 4635) replaced with_mm_extract_epi64(r2_X, N)summed pair.core/src/feature/x86/adm_avx512.c(UPSTREAM): 6 sister lane-extract sites (lines 4470 / 4477 / 4484 / 4625 / 4631 / 4637) — same fix. The AVX-512 paths reduce a__m512idown to__m128ifirst (via_mm512_extracti64x4_epi64→_mm256_extracti64x2_epi64) before the index, so only the final__m128i[N]step needed changing.core/src/feature/x86/motion_avx512.c(UPSTREAM, ported in 9371a0aa from PR #1486): one finalr2[0]+r2[1]reduction (line 448), same fix. All 19 lane-extract fixes plus the 6 cast fixes are bit-exact rewrites and only change the source-level syntax to MSVC-portable form. Restore the original forms post-merge if upstream Netflix later edits the same lines and your toolchain matrix doesn't include MSVC. Additionallycore/src/sycl/d3d11_import.cpp(fork-added) switched from C-style COBJMACROS helpers (ID3D11Device_CreateTexture2D,…_Release, etc.) to C++ method-call syntax (device->CreateTexture2D,tex->Release) — d3d11.h gates COBJMACROS behind!defined(__cplusplus), so the C-style helpers aren't visible in this.cppTU. The two forms are ABI-equivalent (both dispatch through the COM vtable); the choice is purely lexical and POSIX builds aren't affected (the whole TU is#ifdef _WIN32). Round-20 surfaced two more Windows-only blockers. (a) 17 sites across the x86 SIMD layer used GCC'sfloat tmp[N] __attribute__((aligned(M)));form to align scratch buffers for_mm{256,512}_store_ps. MSVC rejects the trailing-attribute syntax withC2146: syntax error: missing ';' before identifier '__attribute__'. Replaced with the C11-standard_Alignas(M) float tmp[N];(alignment specifier before the type) — works in gcc, clang and MSVC with/std:c11. Files touched (all UPSTREAM):vif_statistic_avx2.c(×2),ansnr_avx2.c(×2),ansnr_avx512.c(×2),float_adm_avx2.c(×2),float_adm_avx512.c(×2),float_psnr_avx2.c(×1),float_psnr_avx512.c(×1),ssim_avx2.c(×4),ssim_avx512.c(×4). The pre-existingvif_avx2.c/vif_avx512.calready define a portableALIGNED(x)macro at file scope and position the attribute before the type, so they compile cleanly under MSVC and were not touched. (b)core/src/feature/mkdirp.c(UPSTREAM, third-party MIT-licensed copy of Stephen Mathieson's micro-library) included<unistd.h>unconditionally but never used POSIXunistdsymbols (onlymkdirvia<sys/stat.h>/<direct.h>). Gated<unistd.h>to non-Windows and added<direct.h>for Windows; switchedmkdir(pathname)→_mkdir(pathname)(the non-deprecated MSVC name).core/src/feature/mkdirp.hadded amode_ttypedef under MSVC since neither<sys/types.h>nor<sys/stat.h>declare it on Windows;modeis ignored on the Windows path anyway. Round-21 surfaced two more blockers (the round-19__m128i[N]sweep missed six sites) plus a pre-commit workflow checkout gap. (a)core/src/feature/x86/adm_avx512.c(UPSTREAM) had six furtherr2_X[0] + r2_X[1]reductions at lines 2128 / 2135 / 2142 / 2589 / 2595 / 2601 that reduce a__m512iaccumulator down to__m128ibefore the lane index. Replaced with the same_mm_extract_epi64(r2_X, N)summed-pair pattern used in round 19 — bit-exact, MSVC-portable. (b)core/src/log.c(UPSTREAM) included<unistd.h>unconditionally to pick up POSIXisatty/fileno. On MSVC both live in<io.h>as_isatty/_fileno; gated the include and macro-redirected the names so the one call site at line 34 compiles on both sides without touching the POSIX path. (c).github/workflows/lint-and-format.yml(fork-added) checks out withoutlfs: true, so themodel/tiny/*.onnxfiles land as LFS pointer stubs. pre-commit's "changes made by hooks" reporter then diffs the stubs against HEAD's real blobs and fails the job even though no hook touched them. Addedlfs: trueto the pre-commit job's checkout. (d)core/src/meson.build—cuda_common_vmaf_libstatic library had nodependencies:list, so the Win32 pthread shim (wired in viapthread_dependencyin core/meson.build) wasn't on its include path;cuda/common.hunconditionally#include <pthread.h>and MSVC failed with C1083. Addeddependencies : [pthread_dependency]— no-op on POSIX (empty list), routes the shim path in on Windows. (e)core/src/feature/integer_vif.c(UPSTREAM) walked one bigaligned_mallocresult asvoid *dataand diddata += pad_size/data += h * stride_16etc. to carve the buffer into typed sub-pointers. gcc/clang accept pointer arithmetic onvoid *as a GNU extension (treatingsizeof(void) == 1); MSVC rejects it withC2036: 'void *': unknown size. Replaced the cursor type withuint8_t *and added explicit casts at assignment sites that take a typed pointer (uint16_t *mu1,uint32_t *mu1_32, etc.). Byte offsets are identical, layout unchanged, bit-exact. If upstream Netflix edits the same loop, reabsorb the walk and re-apply the cursor-type + cast pattern. (f)core/src/feature/cuda/integer_adm_cuda.c(UPSTREAM) included<unistd.h>at line 33 but used no POSIX symbols from it; MSVC failed with C1083. Dropped the unused include outright — simplest fix, no runtime change on any platform. (g)core/src/dnn/model_loader.c(fork-added) usesS_ISDIR/S_ISREGto classify resolved paths. MSVC ships the underlyingS_IFMT/S_IFDIR/S_IFREGbit masks in<sys/stat.h>but not the POSIX classification macros. Added a Windows-only fallback (#ifndef S_ISDIR #define S_ISDIR(m) (((m) & S_IFMT) == S_IFDIR) #endif, same for S_ISREG) guarded by#ifdef _WIN32. Semantically identical to the POSIX macro on Linux/macOS. Round-21e surfaced the final source-portability blockers once the DLL build passed preprocessing. (h)core/src/predict.c,core/src/libvmaf.candcore/src/read_json_model.c(all UPSTREAM) used C99 variable-length arrays —double scores[cnt]at predict.c:385,char name[name_sz]at predict.c:453 and libvmaf.c:1741, pluscfg_name[cfg_name_sz]andgenerated_key[generated_key_sz]in the.jsonmodel-collection parser. gcc/clang accept VLAs as a C11 optional feature; MSVC (even with/std:c11) rejects them outright withC2057: expected constant expression(plus C2466 and C2133 on theconst size_tsized arrays — MSVC treatsconstas runtime-bounded, not a constant expression, even when the initialiser is literal like4 + 1). Replaced each runtime-sized buffer with a smallmalloc+ explicitfreeon every exit path (in predict.c and read_json_model.c agoto out;cleanup arm was introduced because the loops error-exit mid-function). Thegenerated_keybuffer in read_json_model.c uses the narrower fix —char generated_key[5];— since its size (four decimal digits of the bootstrap sub-model index plus NUL) is a true compile-time constant. Buffers are a handful of bytes each (name_szis the model-collection name length plus the fixed_ci_p95_losuffix,scoresholds ~20 doubles,cfg_nameis the name plus_0000suffix), so the heap round-trip is not performance-relevant; the new-ENOMEMfailure mode is handled uniformly by existing callers. The read_json_model.c refactor also plugs a pre-existing leak of thenamebuffer on the earlyreturn -EINVALwhen a JSON object key isn't a string — thegoto out;path freesname+cfg_nameon every exit.core/test/test_feature_extractor.c:56(UPSTREAM) declaredconst unsigned n_threads = 8;and used it as the extent ofVmafFeatureExtractorContext *fex_ctx[n_threads];. Converted toenum { n_threads = 8 };so MSVC sees a constant-expression; every other compiler accepts enum constants identically. Re-absorb if upstream Netflix later edits the same loops and your toolchain matrix omits MSVC. (i) The Windows MSVC build-only legs now build the full tree — CLI tools, unit tests and libvmaf.dll — rather than the previous short cut of disabling-Denable_tools/-Denable_tests. Per user direction ("fix the code ffs"), the tree polyfills the remaining POSIX surfaces on MSVC instead: (core/tools/compat/win32/getopt.h+core/tools/compat/win32/getopt.c) a from-scratch POSIX/GNU-compatiblegetopt_longshim (short / long options,no_argument/required_argument/optional_argument, argv permutation for non-option operands,--explicit stop,=-embedded values). The shim is fork-added (BSD-3-Clause-Plus-Patent, Copyright 2026 Lusoris and Claude) and declared via a singlegetopt_dependencyincore/meson.build, gated oncc.check_header('getopt.h')failing. The dependency auto-propagates the shim.cinto any consuming target via meson'ssources:keyword, so both thevmafCLI (core/tools/meson.build) and thetest_cli_parseunit test (core/test/meson.build) pick it up uniformly. MinGW ships<getopt.h>via mingw-w64-crt, socheck_headersucceeds there and the shim stays out of the TU list. (j) Eleven test executables (test_log,test_dict,test_opt,test_cpu,test_ref,test_feature,test_ciede,test_luminance_tools,test_cli_parse,test_sycl,test_sycl_pic_preallocation) were missingpthread_dependencyin theirdependencies:lists atcore/test/meson.build. On POSIXpthread_dependencyis an empty list so the omission was invisible; on MSVC those TUs transitively includefeature_collector.h→<pthread.h>and fail with C1083. Threaded the dependency through all eleven targets.test_cli_parseadditionally listsgetopt_dependencyto pick up the shim. (k) Three additional VLA sites surfaced once the test harness built on MSVC:test_cambi.c:254hadunsigned w = 5, h = 5; uint16_t buffer[3 * w];; converted toenum { w = 5, h = 5 };so the array extent is a constant expression.test_pic_preallocation.c:382andtest_pic_preallocation.c:506hadconst int num_threads = N; pthread_t threads[num_threads];— MSVC rejectsconst intas non-constant-expression. Converted toenum { num_threads = N, fetches_per_thread = M };. (l)test_ring_buffer.c:23andtest_pic_preallocation.c:26included<unistd.h>forusleep/sleep. Gated behind!_WIN32with a Win32 fallback via<windows.h>+#define usleep(us) Sleep(((us) + 999) / 1000)/#define sleep(s) Sleep((s) * 1000). The conversion rounds sub-millisecondusleepinputs up, which is safe for these test paths (they use 100 µs jitter and 1 s waits). (m)core/tools/vmaf.cincluded<unistd.h>forisatty/fileno. Applied the same gating pattern used inlog.cin round-21(b) — include<io.h>on MSVC and redirectisatty/filenoto_isatty/_filenovia#define. (n)__builtin_clz/__builtin_clzllare GCC intrinsics; MSVC ships__lzcnt/__lzcnt64via<intrin.h>instead. The shim already lived incore/src/feature/integer_vif.hbutinteger_adm.c:939,x86/adm_avx2.c:1425andx86/adm_avx512.c:1217don't include that header. Extracted the shim into a dedicatedcore/src/feature/compat_builtin.h(fork-added) and included it from all four TUs. The guard isdefined(_MSC_VER) && !defined(__clang__), so clang-cl / icx-cl (which provide the GCC intrinsics natively) skip the shim. (o) The SYCL leg's D3D11 import TUcore/src/sycl/d3d11_import.cppis C++ (icpx-cl drives it as C++ on Windows) but included the internal C headerlog.hwithout anextern "C"wrap.log.his an upstream Netflix header with no__cplusplusguard, sovmaf_loggot C++ name-mangled in the .cpp TU and failed to resolve against the C-linkage symbol produced bylog.cat link time (LNK2019from every test target that pulls in the SYCL static lib). Wrapped the#include "log.h"withextern "C" { ... }inside the fork-added .cpp rather than touching the upstream header — keepslog.hidentical to upstream on every/sync-upstream. (p) The Windows MSVC legs build with--default-library=static. libvmaf's public API has no__declspec(dllexport)attributes (upstream Netflix is POSIX-shaped), so a vanilla MSVC shared build producessrc/vmaf-3.dllwith no exported symbols and the toolchain therefore never emits the companionvmaf.libimport library. Downstream tool targets then fail withLNK1181: cannot open input file 'src\vmaf.lib'. The MinGW matrix leg has used--default-library staticsince day one for the same reason (line 387); the MSVC legs now mirror that choice viamatrix.include[].meson_extra. Downstream consumers that want a DLL can either add__declspec(dllexport)decorations to the public API or use a.deffile; that is a separate decision and out of scope for the build-only gate. - Re-test:
# Local sanity: the matrix file parses and the new job names exist.
yq '.jobs.windows-gpu-build.strategy.matrix.include[].name' \
.github/workflows/libvmaf-build-matrix.yml
# Expected output (2 lines):
# Build — Windows MSVC + CUDA (build only)
# Build — Windows MSVC + oneAPI SYCL (build only)
- Branch protection: the two Windows GPU legs are pinned as required status checks on
masterimmediately after this PR's merge. After ADR-0120's two Linux DNN legs the count moves 21 → 23. Re-pin via:
gh api --method PUT repos/VMAFx/vmafx/branches/master/protection \
--input /tmp/protection-update.json
0023 — CUDA gencode coverage (sm_86/sm_89/compute_80 PTX) + init hardening¶
- Workstream PRs: the ADR-0122 PR (gencode + init hardening) and the ADR-0123 follow-up for the
32b115dfpost-cubin-load regression. - Touches:
core/src/meson.build— thegencodearray in theif get_option('enable_nvcc')branch.core/src/cuda/common.c—vmaf_cuda_state_init()error paths (multi-line actionable log,cuda_free_functions()+free(c)+*cu_state = NULLcleanup).docs/backends/cuda/overview.md—## Runtime requirementssection and### GPU architecture coveragetable.- Invariant: the
gencodearray unconditionally emits cubins forsm_75/sm_80/sm_86/sm_89plus acompute_80PTX, independent of hostnvccversion. Upstream Netflix's gencode only ships cubins at Txx major boundaries (sm_75/sm_80/sm_90/sm_100/sm_120); a literal merge that replaces our array with upstream's would re-open the Ampere-sm_86/ Ada-sm_89coverage hole. Thesm_90/sm_100/sm_120entries are still version-gated and should be preserved verbatim if upstream adds new gates. The init-path error messages are fork-local strings; upstream's terse"Error: failed to load CUDA functions"must NOT win a merge. - Re-test:
meson setup build -Denable_cuda=true -Denable_nvcc=true
ninja -C build 2>&1 | grep -E 'compute_(80|86|89)'
# Expect at least -gencode=arch=compute_86,code=sm_86 and
# -gencode=arch=compute_89,code=sm_89 and
# -gencode=arch=compute_80,code=compute_80
# Actionable init message (run without CUDA driver on the loader path):
LD_LIBRARY_PATH= ./build/tools/vmaf --help 2>&1 | grep -qi 'libcuda.so.1' || \
echo "init log regressed"
0024 — vmaf_read_pictures null-guard for CUDA device-only path¶
- Workstream PRs: the ADR-0123 follow-up landed atop the ADR-0122 gencode/init-hardening work.
- Touches:
core/src/libvmaf.c— the non-threaded tail ofvmaf_read_picturesat theprev_refupdate site (line ~1428 in the fork; upstream equivalent is the tail added byf740276a).- Invariant: the
prev_refupdate is guarded byif (ref && ref->ref)so pure-CUDA extractor sets (whereref = &ref_hostbutref_hostwas never populated bytranslate_picture_device) do not deref a NULL refcount. Upstream currently has the same unguarded tail; the bug is masked upstream only because the experimentalVMAF_PICTURE_POOLgate from32b115dfis still in place. A literal upstream merge that removes our null-guard while upstream's experimental gate is still holding would pass tests but re-open thelibvmaf_cudaffmpeg crash the moment the gate flips default-on (which the fork did in65460e3a, ADR-0104). Keep the guard until the upstream null-guard port lands. - Re-test:
# Unit tests cover the non-regression on the library side:
meson test -C build
# End-to-end regression: ffmpeg libvmaf_cuda must exit 0 on a
# CUDA-device-only extractor set (full recipe in ADR-0123).
./ffmpeg -init_hw_device cuda=cu:0 -filter_hw_device cu \
-i /tmp/ref.mp4 -i /tmp/dis.mp4 \
-lavfi "[0:v]format=yuv420p,hwupload_cuda[r];\
[1:v]format=yuv420p,hwupload_cuda[d];\
[r][d]libvmaf_cuda=log_path=/tmp/out.json:log_fmt=json" \
-f null -
0025 — VIF init() fail-path frees advanced byte-cursor¶
- Workstream PRs: PR #47 (rewritten to leak-fix-only after master absorbed the void→uint8_t half via commit
b0a4ac3a, entry 0022 §e). Ports the leak-fix half of upstream Netflix PR #1476. - Touches:
core/src/feature/integer_vif.c(UPSTREAM — 2-line fix in theinit()fail:handler). - Invariant:
init()walksuint8_t *dataforward throughaligned_malloc's one allocation, advancing past each sub-pointer assignment. Ifvmaf_feature_name_dict_from_provided_featuresreturns NULL the fail path must free the base pointers->public.buf.data, never the advanced cursordata. Upstream master still hasaligned_free(data)there — same bug — so this entry is the reminder to not let an upstream sync re-introduce the advanced-cursor form. If upstream lands PR #1476 or an equivalent, the sync can drop this entry. - Re-test:
meson test -C build --suite=fast
# Static check: ripgrep the pattern that must NOT return.
rg -n "aligned_free\(data\)" core/src/feature/integer_vif.c && \
echo 'REGRESSED' || echo 'ok'
0026 — Automated rule-enforcement workflow + copyright pre-commit hook¶
- Workstream PRs: this PR (ADR-0124 adoption). Closes the "rule-without-a-check" gap on ADR-0100 / 0105 / 0106 / 0108.
- Touches (all FORK-ADDED — no upstream overlap):
.github/workflows/rule-enforcement.yml(new),scripts/ci/check-copyright.sh(new),.pre-commit-config.yaml(appended local hook). - Invariant: the
deep-dive-checklistjob is blocking on every PR that is not an upstream port (exempt viaport:title prefix orport/branch). The other three gates (doc-substance-check,adr-backfill-check, copyright pre-commit) are advisory or pre-commit, never CI-blocking; this split is the whole point of ADR-0124 and an upstream sync must not move them into the required-status-check set without a follow-up ADR. The opt-out parser matches/^-?\s*no .* (?:needed|impact|rebase-sensitive)/per ADR-0108 §Opt-out-lines — if upstream ever changes PR-template phrasing (unlikely; this is fork-local), the regex and the template must move together. - Re-test:
# Lint the workflow + hook locally.
pre-commit run --files \
.github/workflows/rule-enforcement.yml \
scripts/ci/check-copyright.sh \
.pre-commit-config.yaml
# Dry-run the copyright hook against a staged source file.
scripts/ci/check-copyright.sh core/src/libvmaf.c && echo ok
# Synthetic PR body that violates ADR-0108 should fail the parser;
# see docs/research/0002-automated-rule-enforcement.md §Verification
# plan for the three test cases.
0027 — SSIMULACRA 2 scalar extractor (libjxl FastGaussian IIR blur)¶
- Workstream PRs: this PR (
feat/ssimulacra2-scalar); proposal ADR in PR #67. - Touches:
core/src/feature/ssimulacra2.c(fork-local, new),core/src/meson.build,core/src/feature/feature_extractor.c. - Invariant: the extractor embeds several tables that must track libjxl upstream — opsin absorbance matrix,
MakePositiveXYBoffsets, 108 pooling weights, polynomial-transform coefficients, and the FastGaussian coefficient-derivation formulas (radius =3.2795·σ + 0.2546, Cramer's 3×3 solve for β, n2/d1 assignment per Charalampidis 2016 (33)). If libjxl ever changes any of these, updatessimulacra2.cin the same PR that syncs upstream. Self-consistency must stay at exactly100.000000for identical ref/dist inputs — this is the cheapest regression check. - Re-test:
meson test -C build --suite=fast
./build/tools/vmaf \
--reference python/test/resource/yuv/src01_hrc00_576x324.yuv \
--distorted python/test/resource/yuv/src01_hrc00_576x324.yuv \
-w 576 -h 324 -p 420 -b 8 --feature ssimulacra2 -o /tmp/self.xml \
&& grep -q 'ssimulacra2="100.000000"' /tmp/self.xml \
&& echo "ok: self-consistency 100.0"
0028 — MS-SSIM separable decimate + AVX2/AVX-512/NEON SIMD¶
- Workstream PRs:
feat/ms-ssim-decimate-simd-v2(supersedes the rebase-incompatiblefeat/ms-ssim-decimate-simd; AVX2/AVX-512, commits7de8cd7fscalar separable,5f93c864AVX2,73436438AVX-512);feat/ms-ssim-decimate-neon-v2(NEON follow-up, stacked). - Touches:
core/src/feature/ms_ssim_decimate.{c,h}(NEW),core/src/feature/x86/ms_ssim_decimate_avx2.{c,h}(NEW),core/src/feature/x86/ms_ssim_decimate_avx512.{c,h}(NEW),core/src/feature/arm64/ms_ssim_decimate_neon.{c,h}(NEW),core/src/feature/ms_ssim.c(call-site change),core/src/meson.build(register new SIMD TUs),core/test/test_ms_ssim_decimate.c(NEW),core/test/meson.build(arm64 gating). - Invariant: the 9-tap 9/7 biorthogonal wavelet LPF coefficients (
ms_ssim_lpf_h/ms_ssim_lpf_v) are duplicated verbatim in five TUs for bit-identity: the scalarms_ssim_decimate.c, the AVX2 variant, the AVX-512 variant, the NEON variant, and upstream'sg_lpf_h/g_lpf_vinms_ssim.c. Any upstream change to the coefficient values or theKBND_SYMMETRICmirror branch iniqa/convolve.cmust be mirrored to all five. If not mirrored, SIMD paths and scalar diverge silently and the bit-equalitymemcmpintest_ms_ssim_decimatecatches it — but only when that test runs, so diff the five files first. - Re-test (on each supported host arch):
# x86_64 host — native build.
meson test -C build
./build/test/test_ms_ssim_decimate
# aarch64 host OR aarch64 cross under qemu — see /tmp/aarch64-cross.txt.
meson setup build-arm64 libvmaf --cross-file /tmp/aarch64-cross.txt \
-Denable_cuda=false -Denable_sycl=false
ninja -C build-arm64
qemu-aarch64-static -L /usr/aarch64-linux-gnu \
build-arm64/test/test_ms_ssim_decimate
# Netflix MS-SSIM golden — places=4 must still pass through SIMD.
.venv/bin/python -m pytest \
python/test/feature_extractor_test.py::FeatureExtractorTest::test_run_ms_ssim_fextractor
0029 — KBND_SYMMETRIC period-based reflection in iqa/convolve.c¶
- Workstream PRs:
feat/ms-ssim-decimate-simd-v2follow-up (CI triage on PR #69, 2026-04-20). - Touches:
core/src/feature/iqa/convolve.c(upstream file, rewrittenKBND_SYMMETRIC). - Invariant:
KBND_SYMMETRIC(img, w, h, x, y, _)must use the period-based form (period = 2*w,period = 2*h) so that offsets with|x| > wor|y| > hstill land in bounds. Upstream's single-reflect form was out-of-bounds wheneverw < kernel_halforh < kernel_half; the latent bug did not reproduce in Netflix golden tests because MS-SSIM pyramids never decimate below ~60×34. Any upstream change that reverts to the single-reflect form must be rejected or re-ported. - Re-test:
./build/test/test_ms_ssim_decimate # test_1x1 border case
.venv/bin/python -m pytest \
python/test/feature_extractor_test.py::FeatureExtractorTest::test_run_ms_ssim_fextractor
0030 — adm_decouple_s123_avx512 stack-array 64-byte alignment¶
- Workstream PRs:
feat/ms-ssim-decimate-simd-v2follow-up (CI triage on PR #69, 2026-04-20). - Touches:
core/src/feature/x86/adm_avx512.c(upstream file, one-line_Alignas(64)onint64_t angle_flag[16]at line 1317).core/test/test_pic_preallocation.c(upstream file, threevmaf_model_destroy(model)calls pairing thevmaf_model_loadintest_picture_pool_basic/_small/_yuv444). - Invariant: the stack slot for
angle_flagmust be 64-byte aligned because two_mm512_loadu_si512(&angle_flag[0/8])loads in the same scope may be promoted to alignedvmovdqa64by LTO. Dropping the_Alignas(64)annotation re-introduces the SEGV under--buildtype=release -Db_lto=true -Db_sanitize=address. Debug / no-LTO builds keepvmovdqu64and cannot flag the regression. Seedocs/development/known-upstream-bugs.md. - Re-test:
meson setup build-asan-lto libvmaf \
-Denable_cuda=false -Denable_sycl=false \
-Db_sanitize=address --buildtype=release -Db_lto=true
ninja -C build-asan-lto test/test_pic_preallocation
ASAN_OPTIONS=detect_leaks=1 \
./build-asan-lto/test/test_pic_preallocation
0031 — Batch-A upstream-port small-fix sweep (ports of unmerged PRs)¶
- Workstream PRs:
feat/batch-a-upstream-small-fix-sweep— commits546a40ee(T0-1),8fed8ad1(T4-4),83a1db46(T4-5),34425dee(T4-6). ADRs 0131, 0132, 0134, 0135. - Touches:
core/src/cuda/picture_cuda.c(one-linecuMemFreeport of Netflix#1382)core/src/feature/feature_collector.c+core/test/test_feature_collector.c(mount/unmount bugfix port of Netflix#1406 + shared-helper test refactor)core/src/meson.build(declare_dependency+override_dependencyport of Netflix#1451)core/include/libvmaf/model.h,core/src/model.c,core/test/test_model.c,docs/api/index.md(built-in model iterator port of Netflix#1424)- Invariant: each of the four upstream PRs is OPEN (unmerged) on the port date; when Netflix merges any of them, the fork's version is correction-bearing (T4-4 test refactor, T4-6 three defect fixes + Doxygen doc expansion), not line-identical. Resolution on upstream merge is always "keep fork version" because the fork's version already satisfies the PR's intent and additionally fixes the defects.
- Netflix#1406 conflict will land in
test_feature_collector.c— fork usesload_three_test_models()helper vs upstream's inline per-modelVmafModel *m0, *m1, *m2;duplication. - Netflix#1424 conflict will land in
core/src/model.candcore/test/test_model.c— fork useselse ifguard +idx + 1 < CNT+ const-qualified test types. - Netflix#1382 and Netflix#1451 are line-identical in substance; merge should be clean aside from trailing-comma style drift.
- Re-test:
meson setup build libvmaf -Denable_cuda=false -Denable_sycl=false
ninja -C build test/test_feature_collector test/test_model
build/test/test_feature_collector
build/test/test_model
# Expected: 6/6 pass in test_feature_collector (mount/unmount
# 3-model sequences); 39/39 pass in test_model (includes
# test_version_next full-iteration invariant).
0032 — Thread-local locale handling for numeric I/O (port of Netflix/vmaf#1430)¶
- Workstream PRs:
port/netflix-1430-thread-locale(T4-3 from the "Batch-A follow-up" sweep, 2026-04-20). - Touches:
core/src/thread_locale.h/core/src/thread_locale.c(new, upstream-authored);core/src/meson.build(twocdata.set('HAVE_USELOCALE'/'HAVE_XLOCALE_H')probes +src_dir + 'thread_locale.c'inlibvmaf_sources);core/src/output.c(four writers gainpush_c()+pop()bracket, preserving fork'sferror(outfile) ? -EIO : 0return contract from ADR-0119);core/src/svm.cpp(drop<locale.h>include; replacesetlocale/strdup/setlocalebracket withvmaf_thread_locale_push_c/pop; addbuffer.imbue(std::locale::classic())to both SVM parser ctors with fork's K&R + 4-space style);core/src/read_json_model.c(bracketmodel_parsewith push/pop);core/test/meson.build(newtest_locale_handlingtarget + test registration);core/test/test_locale_handling.c(new, upstream-authored with three fork corrections for thescore_formatparameter). - Invariant: fork's output writers return
ferror(outfile) ? -EIO : 0— this must survive any upstream refactor of the writer bodies. Thepush_c()call MUST be paired with apop()on every return path (writer bodies have a single tail return, so the pattern is locallypush → body → pop → return ferror-check). Droppingpop()leaks alocale_ton POSIX and leaves the thread locked to "C" on Windows. - Re-test:
meson setup build -Denable_cuda=false -Denable_sycl=false
ninja -C build
meson test -C build test_locale_handling
# Repro the user-visible failure without the fix:
LC_ALL=de_DE.UTF-8 build/tools/vmaf --reference ref.yuv \
--distorted dis.yuv --width 1920 --height 1080 \
--pixel_format 420 --bitdepth 8 --output result.json \
--json
# Assert output contains period decimals, not comma.
python -c "import json; d=json.load(open('result.json')); \
assert all('.' in repr(v) for v in \
[f['metrics']['vmaf'] for f in d['frames']])"
- On upstream sync: when Netflix merges PR #1430, the
(cherry picked from commit 054a97ed…)trailer ingit log port/netflix-1430-thread-localelets the next/sync-upstreamskip this commit. If the upstream diff drifts, redo the three fork corrections listed in ADR-0137 §Decision.
0033 — SSIM / MS-SSIM SIMD bit-exact to scalar via per-lane scalar double¶
- Workstream PRs:
feat/ms-ssim-decimate-neon(this PR — companion to the ADR-0138 convolve fast path). - Touches:
core/src/feature/x86/ssim_avx2.candcore/src/feature/x86/ssim_avx512.c—ssim_accumulate_*rewritten.ssim_precompute_*andssim_variance_*unchanged (they were already bit-exact). Plus the new bit-exactconvolve_avx2.c/convolve_avx512.cand the upstream h-pass OOB fix atiqa/convolve.c:159. - Invariants (see ADR-0139 §Decision):
- Convolve taps — single-rounded
float*float→ widen →doubleadd, NO FMA. Mirrors scalarsum += img[i]*k[j]iniqa/convolve.c. - SSIM accumulate — scalar's
2.0 *literal (2.0 * ref_mu[i] * cmp_mu[i] + C1and2.0 * srsc + C2) is a Cdoubleliteral. Both SIMD accumulators do the2.0 *numerator + division + finall*c*sproduct per-lane in scalar double to match scalar type promotions byte-for-byte. - H-pass outer-loop bound —
y < dst_h + vc - kh_even(noty < dst_h + vc); the- kh_evenis load-bearing because the last cache row on even-tap kernels (e.g. box-8) is never read by the v-pass but was previously written OOB when image height equals kernel height.
Fork-local SSIM SIMD is NOT upstream. If upstream ever adds their own SSIM AVX2/AVX-512, keep the fork's version on conflict — it's the only variant verified bit-exact to scalar at --precision max. - Re-test:
meson setup build -Denable_cuda=false -Denable_sycl=false
ninja -C build
meson test -C build test_iqa_convolve test_ms_ssim_decimate
# Bit-exactness check across dispatch backends:
FIX=python/test/resource/yuv/checkerboard_1920_1080_10_3_0_0.yuv
DIS=python/test/resource/yuv/checkerboard_1920_1080_10_3_1_0.yuv
for m in 255 16 0; do
build/tools/vmaf --cpumask $m --reference $FIX --distorted $DIS \
--width 1920 --height 1080 --pixel_format 420 --bitdepth 8 \
--feature float_ssim --feature float_ms_ssim \
--output /tmp/ssim_$m.xml --precision max
done
diff <(grep -v '<fyi fps' /tmp/ssim_255.xml) \
<(grep -v '<fyi fps' /tmp/ssim_16.xml) # expect empty
diff <(grep -v '<fyi fps' /tmp/ssim_255.xml) \
<(grep -v '<fyi fps' /tmp/ssim_0.xml) # expect empty
- On upstream sync: the AVX2/AVX-512 SSIM surface is entirely fork-local (upstream has VIF/ADM/motion/CAMBI SIMD but no SSIM). If upstream ever introduces SSIM SIMD, their kernel bodies will almost certainly compute
l*c*sin vector float for throughput — do not adopt. The fork's per-lane-scalar-double reduction is required for the bit-exactness claim. Same applies toconvolve_avx2/512— they are fork-only; dispatch sits inssim_tools.cvia_iqa_convolve_set_dispatch.
0034 — SIMD DX framework + NEON SSIM/convolve bit-exact port¶
- Workstream PRs:
feat/simd-dx-framework(this PR, PR #A); ships the two demos on top of which PR #B will consume the framework (ssimulacra2, motion_v2, vif_statistic, ...). - Touches:
core/src/feature/simd_dx.h(new header),core/src/feature/arm64/convolve_neon.c+convolve_neon.h(new NEON port),core/src/feature/arm64/ssim_neon.c(ssim_accumulate_neonrewritten for ADR-0139 bit-exactness;precompute+varianceunchanged),core/src/feature/float_ssim.c+core/src/feature/float_ms_ssim.c(wireiqa_convolve_neoninto the aarch64 dispatch setters),core/src/meson.build(arm64_sources+= convolve_neon.c),core/test/meson.build(test_iqa_convolvearch filter extended toarm64/aarch64),core/test/test_iqa_convolve.c(NEON variant check + aarch64 CPU flag detection),core/test/dnn/meson.build(test_cli.shgated onnot meson.is_cross_build()— bash invokes$VMAF_BINdirectly so meson's exe_wrapper isn't applied), newbuild-aux/aarch64-linux-gnu.inimeson cross-file,.claude/skills/add-simd-path/SKILL.md(upgraded kernel-spec flags). - Invariants (see ADR-0140 §Decision):
simd_dx.his fork-local. Keep the fork's version on upstream conflict. Macro names are ISA-suffixed (_AVX2_4L,_AVX512_8L,_NEON_4L) — do not collapse into a cross-ISA abstraction; the fork's SIMD policy (user-memoryfeedback_simd_dx_scope.md) rules out Highway / simde / xsimd.- The ADR-0138 widen-then-add rule (single-rounded
float * float→ widen →doubleadd, NO FMA) applies to NEON exactly as to AVX2 / AVX-512. The NEON form uses pairedfloat64x2_taccumulators (lo / hi) because NEON has nofloat64x4_t. - The ADR-0139 per-lane scalar-double reduction rule applies to
ssim_accumulate_neonexactly as to the AVX2 / AVX-512 variants. The NEON implementation usesSIMD_ALIGNED_F32_BUF_NEON(_Alignas(16) float name[4]) + a 4-iteration scalar loop. - Re-test (requires
aarch64-linux-gnu-gcc+qemu-user-static+ aarch64 sysroot at/usr/aarch64-linux-gnu):
cd libvmaf
meson setup ../build-aarch64 \
--cross-file ../build-aux/aarch64-linux-gnu.ini \
-Denable_cuda=false -Denable_sycl=false -Denable_dnn=disabled
cd ..
ninja -C build-aarch64
meson test -C build-aarch64 # expect 31/31 OK
# Bit-exactness check scalar vs NEON under QEMU:
REF=python/test/resource/yuv/src01_hrc00_576x324.yuv
DIS=python/test/resource/yuv/src01_hrc01_576x324.yuv
for m in 255 0; do
LD_LIBRARY_PATH=$PWD/build-aarch64/src qemu-aarch64-static \
-L /usr/aarch64-linux-gnu build-aarch64/tools/vmaf \
--cpumask $m --reference $REF --distorted $DIS \
--width 576 --height 324 --pixel_format 420 --bitdepth 8 \
--feature float_ssim --feature float_ms_ssim \
--output /tmp/ssim_$m.xml --precision max
done
diff <(grep -v '<fyi fps' /tmp/ssim_255.xml) \
<(grep -v '<fyi fps' /tmp/ssim_0.xml) # expect empty
- On upstream sync: upstream has no NEON SSIM and no NEON convolve for IQA. If they ever add one, keep the fork's version on conflict — the fork's NEON path is the only variant verified bit-exact to scalar at
--precision max. Thebuild-aux/aarch64-linux-gnu.inicross-file has no upstream equivalent. The/add-simd-pathskill is fork-only; upstream doesn't ship.claude/skills/.
0036 — Port Netflix generalised AVX convolve + ADR-0141 cleanup¶
- Workstream PRs:
port/upstream-f3a628b4-generalized-avx-convolve(this PR). - Upstream commit:
f3a628b4"feature/common: generalize avx convolution for arbitrary filter widths" (Kyle Swanson, 2026-04-21). - Touches:
- convolution.h — upstream-tracking: adds
#define MAX_FWIDTH_AVX_CONV 17. - convolution_avx.c — upstream-tracking (2,500 LoC deletion) plus fork-delta cleanup per ADR-0141: four scanline helpers
convolution_f32_avx_s_1d_*changed from external linkage tostatic(no other TU uses them after the specialised-path removal); stride parameters widened frominttoptrdiff_tin the helpers, with(ptrdiff_t)casts at public-function multiplication sites;#include <stddef.h>added for the type. core/src/feature/vif_tools.c— upstream-tracking: three AVX dispatch sites drop thefwidth == 17 || ... == 3whitelist in favour offwidth <= MAX_FWIDTH_AVX_CONV.python/test/quality_runner_test.py,python/test/vmafexec_test.py— upstream-authored loosening of two full-VMAF-score assertions fromplaces=2(±0.005) toplaces=1(±0.05). Adopted per the ADR-0142 Netflix-authority precedent (project rule #1 addresses fork drift, not upstream-authored test updates the fork must track).- Invariants (see ADR-0143 §Decision):
- Static linkage on scanline helpers — upstream leaves the four
convolution_f32_avx_s_1d_*_scanlinehelpers with external linkage out of habit; the fork narrows them tostatic. On upstream sync: if upstream ever externs them from another TU, that's a flag to re-audit; keep the fork'sstaticunless the reference is real. ptrdiff_tstrides inside helpers — the publicconvolution_f32_avx_*_swrappers keepintstrides (matching the upstream interface +convolution.hdeclarations). Helpers takeptrdiff_tto silencebugprone-implicit-widening-of- multiplication-result. If upstream changes the public interface toptrdiff_t, drop the fork's wrapper-level casts.MAX_FWIDTH_AVX_CONV = 17— the ceiling is upstream's; if upstream bumps it, the fork must rebuild + re-run the VIF golden test pair.- Re-test:
meson setup build -Denable_cuda=false -Denable_sycl=false
ninja -C build
meson test -C build # expect 32/32 OK
clang-tidy -p build core/src/feature/common/convolution_avx.c
# Zero warnings expected on the touched file.
Netflix CPU golden CI leg exercises the two loosened assertions; confirmed locally under meson test. - On upstream sync: upstream is the source of truth for convolution_avx.c, convolution.h, vif_tools.c dispatch, and the two python golden tolerances. On a rebase, prefer upstream for those files except: - Keep the fork's static on the four scanline helpers. - Keep the fork's ptrdiff_t helper signatures + multiplication- site casts (unless upstream adopts them too, in which case converge). - Keep the fork's #include <stddef.h>. If upstream re-introduces a specialised fast path for common widths, evaluate on a per-fwidth perf profile — the fork's /profile-hotpath skill covers this.
0038 — motion_v2 NEON SIMD (fork-local)¶
- Workstream PR:
port/motion-bundle-neon-and-updates(this PR). - Upstream: none — aarch64 NEON for
motion_v2is fork-local. Upstream scalar + AVX2 + AVX-512 variants exist; this PR adds the missing NEON fourth path. Scalar is the bit-exactness ground truth. - Touches (fork-local):
- motion_v2_neon.c — new TU, ~300 LoC. 4-wide int32 SIMD over the 5-tap Gaussian pipeline. Five
static inlinehelpers keep every function under the ADR-0141 60-line budget. - motion_v2_neon.h — new header declaring the two public entry points.
- integer_motion_v2.c — dispatch update: adds an
#if ARCH_AARCH64block ininitthat selects the NEON variant whenVMAF_ARM_CPU_FLAG_NEONis present, mirroring the existing x86 dispatch blocks. core/src/meson.build— addarm64/motion_v2_neon.cto thearm64_sourceslist.- Invariants (see ADR-0145 §Decision):
- Arithmetic right-shift throughout. The fork's AVX2 path uses
_mm256_srlv_epi64(logical) which can diverge from scalar on negative-diff pixels. The NEON port usesvshrq_n_s64(v, 16)for the known Phase-2 shift andvshlq_s64(v, -(int64_t)bpc)for the variable Phase-1 shift — both arithmetic, matching scalar C>>on signed integer. On rebase: keep the arithmetic forms; do NOT adoptvshrq_n_u64or a logical emulation even if it runs faster. - 4-lane stride + mirror tails. SIMD stride = 4; scalar tails cover the remainder. The Phase-2 helper
x_conv_row_sad_neonhands 4 lanes tox_conv_block4_neonand drops to scalar for both left/right edges (j < 2andj + 6 > w). On rebase: preserve the 4-lane stride and the two-sided scalar tail. - Signature parity with AVX2. Both pipeline entry points match the AVX2 + AVX-512 variants'
(const uint8_t *prev, ptrdiff_t, const uint8_t *cur, ptrdiff_t, int32_t *y_row, unsigned w, unsigned h, unsigned bpc)signature. On rebase: if upstream changes the signature, mirror the change here AND in the x86 variants in lockstep. - Re-test:
meson setup build-aarch64 libvmaf \
--cross-file build-aux/aarch64-linux-gnu.ini \
-Denable_cuda=false -Denable_sycl=false
ninja -C build-aarch64
meson test -C build-aarch64 --no-rebuild # expect 31/31 OK
clang-tidy -p build-aarch64 \
core/src/feature/arm64/motion_v2_neon.c
# Zero warnings expected on the touched file.
# NEON-vs-scalar bit-exact diff under QEMU:
YUV=python/test/resource/yuv
for mask in 0 255; do
LD_LIBRARY_PATH=build-aarch64/src \
qemu-aarch64-static -L /usr/aarch64-linux-gnu \
build-aarch64/tools/vmaf \
-r $YUV/src01_hrc00_576x324.yuv \
-d $YUV/src01_hrc01_576x324.yuv \
-w 576 -h 324 -p 420 -b 8 -n --feature motion_v2 \
--cpumask $mask -o /tmp/mv2_$mask.xml --precision max
done
diff <(grep -v 'fps=' /tmp/mv2_0.xml) \
<(grep -v 'fps=' /tmp/mv2_255.xml) # expect empty
- On upstream sync: upstream has no NEON
motion_v2and has not signalled plans to add one. If they ever do, diff their NEON against the fork's: on logical-vs-arithmetic shift, keep the fork's arithmetic form (matches scalar). On the function decomposition (the five helpers), adopt upstream's if it's smaller; the fork's layout is ADR-0141-driven, not a semantic contract. - Follow-up T7-32 (fixed 2026-05-09): The
_mm256_srlv_epi64(logical right shift) inmotion_score_pipeline_16_avx2was replaced withsrav_epi64_imm, an AVX2-safe arithmetic-right-shift emulation: logical shift OR sign-fill mask viasrai_epi32+slli_epi64. Two bugs were closed in the same PR: - AVX2 logical-vs-arithmetic shift:
_mm256_srlv_epi64replaced bysrav_epi64_immincore/src/feature/x86/motion_v2_avx2.c. The emulation is bit-exact with scalar C>> bpcon signedint64_t. - Test scalar reference mirror:
mirror_idxincore/test/test_motion_v2_simd.cused2*size - idx - 1instead of2*size - idx - 2, diverging frominteger_motion_v2.c::mirror(). Fixed to-2. All four adversarial fixtures (neg-diff bpc10/12, mixed-diff bpc10/12) now pass.meson test -C build50/50 OK. On rebase: keepsrav_epi64_imm; do not revert to_mm256_srlv_epi64. The rebase-time invariant is now: AVX2 path uses arithmetic shift (matching NEON and scalar).
0039 — readability-function-size NOLINT sweep (ADR-0146)¶
- ADR: ADR-0146
- Touches:
core/src/dict.ccore/src/picture.ccore/src/picture_pool.ccore/src/predict.ccore/src/libvmaf.ccore/src/output.ccore/src/read_json_model.ccore/src/feature/feature_extractor.ccore/src/feature/feature_collector.ccore/src/feature/iqa/convolve.ccore/src/feature/iqa/ssim_tools.ccore/src/feature/x86/vif_statistic_avx2.c- Invariant: every
readability-function-sizeNOLINT suppression has been replaced by a set of smallstatic(orstatic inline, for the SIMD / IQA files) helpers. The helper names are stable interfaces the surrounding code depends on (e.g.iqa_convolve_1d_separable,iqa_convolve_2d,ssim_compute_stats,ssim_workspace_alloc/_free,vif_stat_simd8_compute/_reduce,struct vif_simd8_lane,read_pictures_extractor_loop,read_pictures_post_extractor,read_pictures_validate_and_prep,read_pictures_update_prev_ref). Upstream Netflix has no equivalent helpers; rebases touching any of these files will conflict against the fork's split shape. - On upstream sync:
- If upstream lands a different decomposition of
_iqa_convolveor_iqa_ssim, prefer upstream's shape only if it keeps the ADR-0138 / ADR-0139 bit-exactness invariants (single-rounded float mul → widen to double → double add; per-lane scalar-float reduction through aligned temp buffer). Otherwise keep the fork's split and re-document the divergence here. - The fork renamed
_calc_scale→iqa_calc_scaleto clear thebugprone-reserved-identifiercheck. If upstream modifies_calc_scale, keep the fork's name and port the behavioural change. model_collection_parse_loopwrites directly tocfg_namerather than throughc->name— if upstream ever rewritesmodel_collection_parse, preserve the direct write (it's what lets the param stay non-const without a NOLINT).- Re-test on rebase (x86, any libsvm-less host):
ninja -C build && meson test -C build
for mask in 0 255; do
VMAF_CPU_MASK=$mask ./build/tools/vmaf \
--reference python/test/resource/yuv/src01_hrc00_576x324.yuv \
--distorted python/test/resource/yuv/src01_hrc01_576x324.yuv \
--width 576 --height 324 --pixel_format 420 --bitdepth 8 \
-m version=vmaf_v0.6.1 -o /tmp/vmaf_$mask.xml
done
diff <(grep -v fyi /tmp/vmaf_0.xml) <(grep -v fyi /tmp/vmaf_255.xml)
# expect exit 0 (Netflix-golden-pair VMAF bit-identical scalar vs SIMD)
Also run clang-tidy -p build on every file in Touches; expect zero warnings. - Follow-up T7-6: decide whether to rename the _iqa_* API surface (convolve / ssim / decimate / img_filter / filter_pixel / get_pixel) across all callers to clear the remaining bugprone-reserved-identifier suppressions in ssim.c, ms_ssim.c, float_ms_ssim.c. Out of scope here.
0040 — Thread-pool job recycling + inline data buffer (ADR-0147)¶
- ADR: ADR-0147
- Touches:
core/src/thread_pool.c - Invariants:
VmafThreadPoolJobcarries a fixed-sizechar inline_data[64]buffer. Payloads ≤ 64 bytes go throughmemcpy(job->inline_data, data, data_sz)+job->data = job->inline_data; payloads > 64 bytes take the legacymallocpath. The cleanup path MUST distinguish the two viajob->data != job->inline_data— a naivefree(job->data)would corrupt the slot. Enforced invmaf_thread_pool_job_clear_data.free_jobslist is protected by the existingqueue.lock; enqueue pops from it beforemallocing, runner recycles onto it after running a job.vmaf_thread_pool_destroywalks the list aftervmaf_thread_pool_waitreturns (all workers have exited → no lock needed). Any reorder that frees the queue lock before thefree_jobswalk is a leak on shutdown.- Fork's
void (*func)(void *data, void **thread_data)signature + per-workerVmafThreadPoolWorkerare fork-local; upstream Netflix #1464 hasfunc(void *data). Keep the fork's signature on any rebase — callers (src/libvmaf.c:threaded_enqueue_oneetc.) depend on the two-arg form. -
On upstream sync: Netflix PR #1464 is CLOSED (not merged) and bundles twelve unrelated optimizations. Only the thread-pool portion is ported here. If upstream ever reopens and merges #1464 (or a successor), cherry-pick only the pool mechanics; reject the payload-signature changes, the ADM / VIF / predict.c pieces (they conflict with ADR-0138 / 0139 / 0142 bit-exactness and with T7-5 predict.c refactor), and the feature-collector capacity bump (fork already capped at 8 for a reason — see
src/feature/feature_collector.c). -
Re-test on rebase (x86, any libsvm-less host):
ninja -C build && meson test -C build
for threads in 1 4; do
for mask in 0 255; do
VMAF_CPU_MASK=$mask ./build/tools/vmaf \
--reference python/test/resource/yuv/src01_hrc00_576x324.yuv \
--distorted python/test/resource/yuv/src01_hrc01_576x324.yuv \
--width 576 --height 324 --pixel_format 420 --bitdepth 8 \
-m version=vmaf_v0.6.1 --threads $threads -o /tmp/vmaf_${threads}_${mask}.xml
done
done
# Expect bit-identical scores (attribute order may differ across
# --threads 1 vs --threads 4 because feature-collector emits in
# insertion order; the numeric values match).
diff <(grep -v fyi /tmp/vmaf_4_0.xml) <(grep -v fyi /tmp/vmaf_4_255.xml)
# expect exit 0 (scalar vs SIMD threaded)
Also run clang-tidy -p build core/src/thread_pool.c — expect zero warnings. Re-run the 500 000-job micro-benchmark from ADR-0147 §Decision if performance is under investigation.
0041 — IQA reserved-identifier rename + cleanup (ADR-0148)¶
- ADR: ADR-0148
- Touches: 21 files across
core/src/feature/(iqa/{convolve,decimate,ssim_tools}.{c,h},iqa/ssim_simd.h,ssim.c,integer_ssim.c,ms_ssim.c,ms_ssim_decimate.h,float_ssim.c,float_ms_ssim.c,x86/convolve_avx2.{c,h},x86/convolve_avx512.{c,h},arm64/convolve_neon.{c,h},AGENTS.md) pluscore/test/test_iqa_convolve.c. - Invariants:
- Every
_iqa_*/_kernel/_ssim_int/_map_reduce/_map/_reduce/_context/_ms_ssim_*/_ssim_*/_alloc_buffers/_free_bufferssymbol and the four underscore-prefixed header guards (_CONVOLVE_H_,_DECIMATE_H_,_SSIM_TOOLS_H_,__VMAF_MS_SSIM_DECIMATE_H__) is renamed to its non-reserved spelling. The fork's IQA surface no longer uses C's reserved-identifier name space. - The
clang-analyzer-security.ArrayBoundNOLINT bracket inssim_accumulate_rowandssim_reduce_row_range(integer_ssim.c) is load-bearing — the inner kernel-loopk_min/k_maxclamping is provably correct (k_min = max(0, hkernel_offs - x),k_max = min(hkernel_sz, hkernel_sz - (x + hkernel_offs - w + 1))) but the analyzer can't follow it across helper boundaries. Do not collapse the bracket. - The
clang-analyzer-unix.MallocNOLINT bracket intest_iqa_convolve.c(check_simd_variant,check_case) is intentional — test exits process on failure path; small allocations leak by design at test end. Do not refactor to free-on-exit. - The cross-TU NOLINT pattern on
compute_ssim(ssim.c) andcompute_ms_ssim(ms_ssim.c) — clang-tidymisc-use-internal-linkageruns per-TU and can't see the header bridge tofloat_ssim.c/float_ms_ssim.c. Keep the inline justification comment. - On upstream sync:
- The Netflix upstream IQA library (
tjdistler/iqa) has been effectively abandoned (last meaningful commit pre-2020). Future rebases will conflict on every renamed symbol; drop the underscore-prefix on each conflict and mirror the fork'siqa_*naming. - If upstream Netflix/vmaf ever reincorporates the IQA naming wholesale, prefer the fork's spellings — this PR is a one-shot mechanical rename with no semantic content.
- Re-test on rebase:
ninja -C build && meson test -C build
for mask in 0 255; do
VMAF_CPU_MASK=$mask ./build/tools/vmaf \
--reference python/test/resource/yuv/src01_hrc00_576x324.yuv \
--distorted python/test/resource/yuv/src01_hrc01_576x324.yuv \
--width 576 --height 324 --pixel_format 420 --bitdepth 8 \
-m version=vmaf_v0.6.1 \
--feature float_ssim --feature float_ms_ssim \
-o /tmp/iqa_$mask.xml
done
diff <(grep -v fyi /tmp/iqa_0.xml) <(grep -v fyi /tmp/iqa_255.xml)
# expect exit 0 (bit-identical scalar vs SIMD on float_ssim/ms_ssim)
Also run clang-tidy -p build on every touched file (excluding arm64/); expect zero warnings.
0042 — Port Netflix #1376 — FIFO-hang fix via Semaphore (ADR-0149)¶
- ADR: ADR-0149
- Upstream commit: Netflix PR #1376, head
1c06ca4f1bb5da38b54db075a27c35ba8ea9d7b7(OPEN upstream as of 2026-04-24). - Touches:
python/vmaf/core/executor.py— baseExecutorclass +ExternalVmafExecutor-style subclass; delete_wait_for_workfiles/_wait_for_procfilespolling loops; rewrite_open_{work,proc}files_in_fifo_modearoundmultiprocessing.Semaphore(0); addopen_sem=Nonekwarg to every_open_{ref,dis}_{work,proc}fileand to the_open_workfilestaticmethod; drop unusedfrom time import sleep.python/vmaf/core/raw_extractor.py—AssetExtractor+DisYUVRawVideoExtractor; addopen_sem=Noneto_open_{ref,dis}_workfileoverrides (release on entry since these are no-ops); delete_wait_for_workfilesoverrides; drop unusedfrom time import sleep.- Fork carve-outs (load-bearing on rebase):
compat/python-vmaf/__init__.py:__version__follows the rootx-release-please-versionmarker — do NOT port upstream's bump to"4.0.0"independently. The fork uses one release stream per ADR-1127.from time import sleepis dropped from both files — upstream leaves the import in place (unused after their patch); the fork removes it because ADR-0141 touched-file rule requires ruff F401 clean.- Upstream typo preserved: the subclass warning message contains "to be created to be created". Comments note the typo inline; do not silently fix on rebase — it's upstream- authored and project policy is verbatim port.
- On upstream sync: upstream PR #1376 is still OPEN. When it merges, re-diff against the merged form; the touched hunks should be conflict-free because the fork now carries the same shape. Re-check whether upstream fixed the "to be created to be created" typo; if so, adopt the fix (it becomes a simple string update).
- Re-test:
python3 -m py_compile python/vmaf/core/executor.py \
python/vmaf/core/raw_extractor.py
ruff check python/vmaf/core/executor.py python/vmaf/core/raw_extractor.py
black --check python/vmaf/core/executor.py python/vmaf/core/raw_extractor.py
# all silent
# No FIFO-mode unit test in the tree; end-to-end harness
# exercise (needs libsvm + ffmpeg + fixtures) goes via
# make test-netflix-golden
# which doesn't exercise fifo_mode path but does verify the
# refactor didn't break executor.py imports.
0043 — Port Netflix #1472 — CUDA on Windows MSYS2/MinGW (ADR-0150)¶
- ADR: ADR-0150
- Upstream commits: Netflix PR #1472 —
15745cdf(portability) +b7b65e64(meson plumbing). Both OPEN upstream as of 2026-04-24. - Touches:
core/src/cuda/common.h— drop<pthread.h>include; rename reserved header guard__VMAF_SRC_CUDA_COMMON_H__→VMAF_SRC_CUDA_COMMON_INCLUDED.core/src/cuda/cuda_helper.cuh—#ifdef DEVICE_CODEguard around<cuda.h>vs<ffnvcodec/dynlink_loader.h>.core/src/picture.h—#ifdef DEVICE_CODEguard around<cuda.h>+ forward-declareVmafCudaStatevs<ffnvcodec/*>+ fulllibvmaf_cuda.h; rename reserved header guard.core/src/feature/integer_adm.h— updated comment abovedwt_7_9_YCbCr_thresholdtable noting the fork's positional-initializer shape vs upstream's#ifndef __CUDACC__shape (see §Fork carve-outs).core/src/feature/cuda/integer_adm/{adm_cm,adm_csf,adm_csf_den,adm_decouple,adm_dwt2}.cu—#ifndef DEVICE_CODEguard around#include "feature_collector.h".core/src/meson.build— Windows nvcc plumbing (+70 LoC underhost_machine.system() == 'windows'):vswhere-basedcl.exediscovery, MSVC + Windows SDK include path injection, CUDA version detection vianvcc --version,nvcc_ccbin_flags+nvcc_host_includesthreaded through everycustom_targetthat invokes nvcc.- Fork carve-outs (load-bearing on rebase):
integer_adm.huses positional initializers, NOT upstream's#ifndef __CUDACC__wrap. Both shapes resolve the MSVC/nvcc C++-designated-initializer issue; the positional form is C++-portable and keeps the table available to future.cuconsumers. Keep the fork's form on rebase.cuda_static_libkeepsdependencies : [pthread_dependency]. Upstream drops it; the fork needs it becausering_buffer.c(built as part ofcuda_static_lib)#includes<pthread.h>directly. On rebase: keep the fork's version.meson.buildgencode coverage block: the fork's ADR-0122 explicit cubin list (sm_75/80/86/89 + compute_80 PTX) sits after the new upstream nvcc-detect block. On rebase, re-assemble the same merged order: nvcc-detect first, then gencode coverage (both host-independent).- Header guards:
_INCLUDEDspellings are fork-local (ADR-0148 precedent). Upstream keeps reserved__VMAF_SRC_*_H__spellings. On rebase, keep_INCLUDED. - On upstream sync: PR #1472 is still OPEN. When merged, re-diff the three conflict-resolved hunks against upstream's final form. Keep fork's version on the four carve-outs above unless upstream meaningfully reshapes those regions.
- Re-test on rebase (Linux host with CUDA toolkit):
meson setup libvmaf core/build-cuda \
-Denable_cuda=true -Denable_nvcc=true -Denable_sycl=false
ninja -C core/build-cuda && meson test -C core/build-cuda
# Expect 6 .fatbin files generated + CLI linked + 35/35 tests pass.
Windows validation is operator-driven — CI does not yet have a Windows + MSYS2 + MinGW + MSVC BuildTools + CUDA runner (tracked as T7-3 in .workingdir2/OPEN.md). - Prerequisites note (Windows only): nv-codec-headers must be built from git master commit 876af32 or later. The release tag n13.0.19.0 is missing cuMemFreeHost, cuStreamCreateWithPriority, cuLaunchHostFunc, and other CudaFunctions members libvmaf uses. Pre-existing issue, not scope of this port.
0058 — libvmaf.pc Cflags leak fix (ADR-0200)¶
- ADR: ADR-0200; bug-fix follow-up to entry 0057.
- Upstream source: fork-local. Netflix has no Vulkan backend.
- Touches:
core/subprojects/packagefiles/volk/meson.build— drops-include volk_priv_remap.hfromvolk_dep.compile_args; keeps-DVK_NO_PROTOTYPES.core/src/vulkan/meson.build— pullsvolk_priv_remap_h_pathfrom the volk subproject and appends['-include', <path>]tovmaf_cflags_common(privatec_args:on libvmaf'slibrary()call).- Invariants (load-bearing):
-includeMUST stay offvolk_dep.compile_args— otherwise it leaks into staticlibvmaf.pcCflags. Test on rebase:meson setup ... -Ddefault_library=static -Denable_vulkan=enabled, thengrep Cflags meson-private/libvmaf.pc— must NOT containvolk_priv_remapor any build-dir absolute path.-includeMUST be applied to libvmaf's compile — every libvmaf TU that calls volk'svk*API needs the rename macros active. Thevmaf_cflags_commoninjection covers this for all libvmaf sub-libraries (libvmaf_feature, libvmaf_cpu, etc.).- The path comes from
subproject('volk').get_variable(...), not from a hardcoded string — survives volk wrap version bumps. - On upstream sync: zero upstream interaction.
- Re-test on rebase / volk wrap bump:
meson setup build-vk-static-test libvmaf -Denable_vulkan=enabled \
-Denable_cuda=false -Denable_sycl=false -Ddefault_library=static
ninja -C build-vk-static-test src/libvmaf.a
grep Cflags build-vk-static-test/meson-private/libvmaf.pc
# Expected: no `volk_priv_remap` substring, no build-dir absolute path
0057 — Volk vk* priv-remap for static-archive builds (ADR-0198)¶
- ADR: ADR-0198; follow-up to ADR-0185.
- Upstream source: fork-local. Netflix/vmaf has no Vulkan backend.
- Touches:
core/subprojects/packagefiles/volk/meson.build— overlay applied on top of the upstream volk wrap. Adds acustom_targetthat runsgen_priv_remap.pyto producevolk_priv_remap.hfrom the upstreamvolk.h, and wires-includeof the generated header intovolk.c'sc_argsandvolk_dep'scompile_args.core/subprojects/packagefiles/volk/gen_priv_remap.py— fork-added generator script (regex againstextern PFN_vkXxx vkXxx;declarations).- Invariants (load-bearing):
- Force-include must propagate to every libvmaf TU pulling in
volk_dep— verified via meson dep graph. Removing the-includefromcompile_argsre-introduces the static-link multi-def cascade. - Generator regex matches every
vk*PFN declaration involk.h— confirmed for volk-1.4.341 (784declarations,784remaps). Bumping the volk wrap version: re-run the generator (it's a configure-time custom target, so it's automatic) and confirm the rename count printed to stdout matches the count of^extern PFN_vklines in the newvolk.h. - The renamed symbols use the
vmaf_priv_prefix — chosen to match no upstream Netflix or Vulkan SDK identifier. Don't rename to_vk*(collides with reserved-identifier C namespace) orvkv_*etc. - On upstream sync: zero upstream interaction. The volk wrap is a libvmaf-managed subproject; Netflix doesn't ship a Vulkan backend.
- Re-test on rebase / after any volk wrap bump:
meson setup build-vk-static libvmaf -Denable_vulkan=enabled \
-Denable_cuda=false -Denable_sycl=false \
-Ddefault_library=static
ninja -C build-vk-static src/libvmaf.a
test "$(nm build-vk-static/src/libvmaf.a 2>/dev/null \
| grep -cE '^[0-9a-f]* (T|D|B|R) vk[A-Z]')" = "0" \
&& echo OK
(Followed by the BtbN-style link reproducer in the ADR References section.)
0056 — SSIMULACRA 2 snapshot gate + fp-contract-off split (ADR-0164)¶
- ADR: ADR-0164
- Upstream source: fork-local. Netflix/vmaf has no SSIMULACRA 2.
- Touches:
- python/test/ssimulacra2_test.py — new fork-added Python test. Uses
subprocess.callagainstExternalProgram.vmafexecwith--feature ssimulacra2; parses the--jsonoutput; asserts pooled + per-frame scores. - Invariants (load-bearing):
- Pinned values are CPU-only — generated on master HEAD after PR #100 merge. Re-generate if the scalar or any SIMD path changes semantically (which per ADR-0161/0162/0163's bit-exactness contract, it shouldn't — any bit-exact refactor leaves pinned values unchanged).
- Tolerance is 4 decimal places (
places=4) — matches 1e-4. The CPU paths are bit-exact so actual drift should be 0; the tolerance is defensive. -ffp-contract=offeverywhere in the ssimulacra2 pipeline:libvmaf_ssimulacra2_static_lib(scalar extractor),x86_ssimulacra2_avx2_lib,x86_ssimulacra2_avx512_lib, andarm64_ssimulacra2_lib(from ADR-0161). All four split out of their umbrella libs so other extractors keep upstream's default FMA policy. Without this the CI GCC/clang hosts drifted ~2e-4 from my AVX-512 authoring host — GCC 10+ defaults-ffp-contract=faston x86 with-mfmaand on aarch64, fusinga*b+cin scalar glue around the SIMD calls. Do NOT remove any of these carve-outs on rebase.- Fixtures are already-checked-in —
src01_hrc00/01_576x324is also the primary Netflix golden fixture; the 160×90 derived one stresses the sub-176 pyramid-termination path. - Do NOT modify the Netflix golden assertions in quality_runner_test.py et al. — those are upstream-pinned. This test is a SEPARATE file that adds fork-specific scores.
- On upstream sync: no upstream interaction. If Netflix adopts SSIMULACRA 2 in the future, cross-reference against their pinning if they add one.
- Re-test on rebase / after any ssimulacra2 change:
- Follow-ups:
- Cross-reference gate against libjxl
tools/ssimulacra2whenssimulacra2_rscargo install is fixed. - Expand fixture coverage if new YUV test assets land.
0055 — SSIMULACRA 2 picture_to_linear_rgb SIMD (ADR-0163)¶
- ADR: ADR-0163
- Upstream source: fork-local. Netflix/vmaf has no SSIMULACRA 2.
- Touches:
- ssimulacra2_avx2.{c,h} — new
ssimulacra2_picture_to_linear_rgb_avx2+ helpers (read_plane_scalar_s2,srgb_to_linear_lane_avx2,compute_matrix_coefs). - ssimulacra2_avx512.{c,h} — 16-wide AVX-512 port.
- ssimulacra2_neon.{c,h} — 4-wide aarch64 port.
- ssimulacra2.c — new
ptlr_fnfield inSsimu2State; dispatch wrapperconvert_picture_to_linear_rgbunpacksVmafPictureintosimd_plane_t[3]; init assigns AVX2/AVX-512/NEON pointers. - ssimulacra2_simd_common.h — new shared header declaring
simd_plane_t. Decouples SIMD TUs fromVmafPicturetype. - test_ssimulacra2_simd.c — new
test_ptlr_420_8,test_ptlr_420_10,test_ptlr_444_8,test_ptlr_444_10,test_ptlr_422_8subtests + scalar referencesref_read_plane,ref_srgb_to_linear,ref_picture_to_linear_rgb. - Invariants (load-bearing):
- Scalar-order matmul —
G = Yn + cb_g * Un + cr_g * Vnchained left-to-right in all three SIMD TUs. Regression test catches reordering drift (~1 ulp). - Per-lane scalar
powf— vector polynomial approximation would drift scalar bit-exactness. Do not replace the lane spill/reload pattern with a vector libm. simd_plane_tlayout —{data, stride, w, h}ordering assumed by all three SIMD TUs. The dispatch wrapper builds this fromVmafPicturefields; layout must match.- Bounds clamping in
read_plane_scalar_*mirrors scalar reference verbatim (if (sx < 0) sx = 0; if (sx >= pw) sx = pw-1;etc.). Do not simplify — removes per-lane safety at plane edges. - Arbitrary chroma ratios fall through to the
int64_tmultiplication branch. Don't remove it — SSIMULACRA 2 is supposed to accept non-standard ratios gracefully. - On upstream sync: no upstream interaction. If Netflix adopts SSIMULACRA 2 in the future and provides a SIMD YUV→RGB path, diff against the fork's — preserve the bit-exactness contract unless ADR-0142 Netflix-authority carve-out opens.
- Re-test on rebase:
ninja -C build && build/test/test_ssimulacra2_simd # 11/11
ninja -C build-aarch64 && \
qemu-aarch64-static -L /usr/aarch64-linux-gnu/ \
build-aarch64/test/test_ssimulacra2_simd # 11/11
- Follow-ups:
- T3-3 SSIMULACRA 2 snapshot-JSON regression test — still pending (gated on
tools/ssimulacra2availability). - SSIMULACRA 2 now has zero scalar hot paths. T3-1 closes in full with phases 1+2+3 (ADR-0161, 0162, 0163).
0054 — SSIMULACRA 2 FastGaussian IIR blur SIMD (ADR-0162)¶
- ADR: ADR-0162
- Upstream source: fork-local. No SSIMULACRA 2 extractor in upstream Netflix/vmaf.
- Touches:
- ssimulacra2_avx2.{c,h} — new
ssimulacra2_blur_plane_avx2+ 2 helpers (hblur_8rows_avx2,vblur_simd_8cols_avx2). - ssimulacra2_avx512.{c,h} — 16-wide port.
- ssimulacra2_neon.{c,h} — 4-wide aarch64 port, uses
vsetq_lane_f32in place of gather. - ssimulacra2.c — adds
blur_fnfunction pointer toSsimu2State, dispatch ininit_simd_dispatch(), call-site inblur_3plane. - test_ssimulacra2_simd.c — new
test_blur+ scalar reference (ref_blur_plane,ref_fast_gaussian_1d). - Invariants (load-bearing):
- Row-batching lane layout — horizontal pass lane
iMUST hold row(y_base + i). Gather index vector entries are(y_base + i) * w(stride-w). Changing this breaks bit-exactness vs scalar. - Scalar left-to-right summation order —
n2_k * sum - d1_k * prev1_k - prev2_kchained sequentially;o0 + o1 + o2at output time is(o0 + o1) + o2. Changing to(o0 + o2) + o1oro0 + (o1 + o2)will drift ~1 ulp and the regression test catches it. col_stateis 6 * w contiguous floats — layout is[prev1_0 | prev1_1 | prev1_2 | prev2_0 | prev2_1 | prev2_2]. SIMD loads assume this layout; changing field order requires updating all three SIMD TUs in lockstep withblur_plane.- NEON lane-set pattern — aarch64 has no gather intrinsic; 4 explicit
vsetq_lane_f32calls per input vector. Do not replace with ald1 {v.s}[lane]-style pseudo-gather without re-verifying bit-exactness. - Scalar tail in vertical pass matches scalar reference body verbatim. Any deviation breaks
memcmpequality on widths that aren't multiples of the SIMD width. - On upstream sync: no upstream interaction. If Netflix adopts SSIMULACRA 2 in the future and provides their own IIR blur SIMD, diff against the fork's and preserve the bit-exactness contract unless an ADR-0142 Netflix-authority carve-out is opened.
- Re-test on rebase:
meson setup build -Denable_cuda=false -Denable_sycl=false
ninja -C build
meson test -C build test_ssimulacra2_simd # 6/6
# aarch64:
ninja -C build-aarch64
qemu-aarch64-static -L /usr/aarch64-linux-gnu/ \
build-aarch64/test/test_ssimulacra2_simd # 6/6
- Follow-ups:
picture_to_linear_rgbSIMD — last scalar hot path in the extractor. 2 calls / frame. Low ROI but mechanical.- T3-3 SSIMULACRA 2 snapshot-JSON regression test — still pending.
0053 — SSIMULACRA 2 SIMD bit-exact ports (ADR-0161)¶
- ADR: ADR-0161
- Upstream source: fork-local. Upstream Netflix/vmaf has no SSIMULACRA 2 extractor at all (fork-added in ADR-0130).
- Touches:
- ssimulacra2_avx2.c / .h — 5 AVX2 kernels + per-lane
cbrtfhelper. - ssimulacra2_avx512.c / .h — 5 AVX-512 kernels; mechanical 16-wide widening of the AVX2 path.
- ssimulacra2_neon.c / .h — 5 NEON kernels; 4-wide aarch64 mirror.
- ssimulacra2.c — adds function-pointer dispatch fields to
Ssimu2State+init_simd_dispatch()helper, calls go through the pointers. - meson.build — registers the three SIMD TUs in
x86_avx2_sources/x86_avx512_sources/arm64_sources. - test_ssimulacra2_simd.c and
test/meson.build— new bit-exact test harness. - Invariants (load-bearing):
- Byte-for-byte bit-exactness to scalar on all 5 vectorised kernels under
FLT_EVAL_METHOD == 0. Regression caught pre- merge: naïve pairing(a+b)+(c+d)vs scalar((a+b)+c)+ddrifts by 1 ULP. Keep sequential scalar-order chains in all three SIMD TUs on rebase. cbrtfis per-lane scalar libm, not a polynomial. Any replacement with a vector cbrt would drift the ssimulacra2 score and break the regression test. Keep the spill/reload pattern.ssim_map/edge_diff_mapreductions use the ADR-0139 per-lanedoublescalar tail. Do NOT SIMD-reduce float lanes then lift to double — summation order changes.downsample_2x2deinterleave uses ISA-appropriate ops: AVX2vshufps+vpermpd, AVX-512vpermt2ps, NEONvuzp1q_f32+vuzp2q_f32. After deinterleave, sum order is((r0e+r0o)+r1e)+r1omatching scalar.#pragma STDC FP_CONTRACT OFFat every TU header. Ignored by aarch64 GCC (non-fatal-Wunknown-pragmas); kept for portability (clang, MSVC).- IIR blur +
picture_to_linear_rgbstay scalar in this PR. Follow-up PRs target these; when they land, re-verify bit-exactness viatest_ssimulacra2_simdexpansion. - Runtime dispatch order: AVX-512 > AVX2 on x86; NEON on aarch64; scalar fallback. Preserve on rebase.
- On upstream sync:
- Upstream has no SSIMULACRA 2 extractor; nothing to merge.
- If Netflix adopts SSIMULACRA 2 in the future, diff their implementation against the fork's scalar + SIMD TUs; keep the fork's bit-exactness contract absent a specific Netflix-authority carve-out ADR.
- Re-test on rebase:
meson setup build -Denable_cuda=false -Denable_sycl=false
ninja -C build
meson test -C build test_ssimulacra2_simd # 5/5
clang-tidy -p build core/src/feature/x86/ssimulacra2_avx2.c \
core/src/feature/x86/ssimulacra2_avx512.c
# aarch64:
ninja -C build-aarch64
qemu-aarch64-static -L /usr/aarch64-linux-gnu/ \
build-aarch64/test/test_ssimulacra2_simd # 5/5
clang-tidy -p build-aarch64 \
core/src/feature/arm64/ssimulacra2_neon.c
- Follow-ups:
- IIR blur vectorisation (
blur_planevertical-pass column batching) — the biggest frame-level wallclock win. picture_to_linear_rgbper-lanepowf— lower ROI but mechanical.- T3-3 SSIMULACRA 2 snapshot-JSON regression test — ADR-0130 deferred; still pending.
0052 — psnr_hvs SIMD bit-exact ports (ADR-0159 AVX2, ADR-0160 NEON)¶
- ADRs: ADR-0159 (AVX2), ADR-0160 (NEON sister port).
- Upstream source: fork-local. Upstream Netflix/vmaf has no psnr_hvs SIMD path.
- Touches:
core/src/feature/x86/psnr_hvs_avx2.c— AVX2 TU.core/src/feature/x86/psnr_hvs_avx2.h— AVX2 header.core/src/feature/arm64/psnr_hvs_neon.c— NEON TU (sister port, ADR-0160).core/src/feature/arm64/psnr_hvs_neon.h— NEON header.core/src/feature/third_party/xiph/psnr_hvs.c— addPsnrHvsState+ runtime dispatch ininit()(AVX2 underARCH_X86, NEON underARCH_AARCH64) + scoped NOLINTBEGIN/END around the upstream Xiph scalar block (kept verbatim as the bit-exact reference).core/src/meson.build— addx86/psnr_hvs_avx2.ctox86_avx2_sourcesandarm64/psnr_hvs_neon.ctoarm64_sources.core/test/test_psnr_hvs_avx2.c,core/test/test_psnr_hvs_neon.c— bit-exact unit tests (x86 and aarch64 respectively).core/test/meson.build— register both tests underenable_asm, arch-gated.- Invariants (load-bearing):
- Bit-exactness to scalar: every
od_coeff(int32) and every finalpsnr_hvs_{y,cb,cr,psnr_hvs}value the AVX2 path emits must be byte-identical to the scalar reference on the Netflix golden pairs. If a rebase introduces any pattern that breaks this (e.g. a floating-point horizontal reduce in the mask accumulator), the unit testtest_psnr_hvs_avx2will fail — don't relax the assertions; fix the SIMD path. - DCT butterfly layout:
butterfly → transpose → butterfly → transpose. The transpose lives insideod_bin_fdct8x8_avx2. Do not move it. - Float accumulators stay scalar: means / variances / mask / error accumulation in
calc_psnrhvs_avx2use the same per-block scalar loop as scalar psnr_hvs — bit-exact by construction. Do not vectorize these with horizontal reductions without replicating ADR-0139's per-lane scalar-float reduction pattern. The cross-block error accumulatorretis threaded throughaccumulate_error()by pointer, not returned-then-summed: each of the 64 per-coefficient contributions per block must hit the outerretdirectly, matching scalar's inlineret += ...atthird_party/xiph/psnr_hvs.cline 355. IEEE-754 float add is non-associative — summing into a local float and then adding the per-block total toretchanges the summation tree and drifts the Netflix golden by ~5.5e-5. #pragma STDC FP_CONTRACT OFFat the TU header disables FMA formation. Required:fmaf(a, b, c)can differ from(a*b)+cby 1 ulp, breaking bit-exactness. Do not remove the pragma; do not add-ffp-contract=fastto the build flags for this TU.- NOLINT suppressions are load-bearing — each cites ADR-0141 inline (bit-exactness scalar-diff auditability for the 30-butterfly function, scalar float→double promotion for
sqrt, extractor-registry extern linkage forvmaf_fex_psnr_hvs, upstream-Xiph scoped block for rebase parity). - On upstream sync:
- Upstream has no psnr_hvs SIMD as of 2026-04-24. Keep fork's version on conflict.
- If upstream ever touches
psnr_hvs.cfor non-SIMD reasons (e.g. a masking-table update), rebase the AVX2 TU to match line-for-line and re-runtest_psnr_hvs_avx2to confirm bit-exactness survives. - NEON follow-up PR is a sister port; its
arm64/psnr_hvs_neon.cwill mirror this ADR's invariants. On rebase, the two SIMD TUs must stay in lock-step with the scalar reference. - Re-test on rebase:
ninja -C build
meson test -C build test_psnr_hvs_avx2
# Expect: 5/5 subtests pass (DCT bit-exact on 3 random seeds +
# delta + constant input).
# CLI-level bit-exactness on Netflix golden (requires the YUV
# fixtures in python/test/resource/yuv/):
# VMAF_CPU_MASK=0 (scalar)
# VMAF_CPU_MASK=255 (AVX2 enabled)
# Diff per-frame psnr_hvs_{y,cb,cr,psnr_hvs} XML fields; expect
# byte-identical across all 3 golden pairs.
0051 — Netflix#1486 motion updates verified present (ADR-0158)¶
- ADR: ADR-0158
- Upstream source: Netflix upstream PR #1486 ("Port motion updates"), MERGED 2026-04-20 as commits
a44e5e6(code) +62f47d5(Netflix golden updates). - Touches: documentation-only; the actual code changes this ADR documents are already in the fork's master via earlier incremental motion3 / blend / five-frame-window commits.
- Invariants (load-bearing for future
/sync-upstream): - The
edge_8mirror fix (i_tap = height - (i_tap - height + 2)) is present atinteger_motion.c:240,x86/motion_avx2.c:147,x86/motion_avx512.c:147. If upstream's mirror line ever diverges again, this is the hunk to watch. - The
motion_max_valfeature option is atinteger_motion.c:57,118-120with default 10000.0 andFEATURE_PARAMflag. Upstream's default = fork's default; don't drift. VMAF_integer_feature_motion3_scoreoutput plumbing is ininteger_motion.c+alias.c.- Fork-local motion extensions (five-frame-window, moving-average, blend, fps_weight) are ADDITIONS on top of Netflix#1486. They are not upstream. Upstream changes to motion extractor internals may conflict with them — diff against
core/src/feature/integer_motion.con every rebase and check that the fork'sMIN(s->score * s->motion_fps_weight, s->motion_max_val)invocations are preserved (lines ~409, ~503). - On upstream sync: nothing to port from Netflix#1486 — it's absorbed. If a future upstream PR touches the same code paths, prefer upstream's version for the scalar/edge handling and the fork's version for the five-frame-window / blend extensions.
- Re-test on rebase:
ninja -C build
meson test -C build
# Expect: 35/35 pass.
# Verify the upstream markers are still in place after rebase:
grep -n "height - (i_tap - height + 2)\|motion_max_val\|VMAF_integer_feature_motion3_score" \
core/src/feature/integer_motion.c \
core/src/feature/alias.c \
core/src/feature/x86/motion_avx2.c \
core/src/feature/x86/motion_avx512.c
# Expect: matches at all 4 files. If any missing, the rebase
# silently dropped the Netflix#1486 content — investigate.
0050 — CUDA preallocation memory leak fix + vmaf_cuda_state_free (ADR-0157)¶
- ADR: ADR-0157
- Upstream source: Netflix upstream issue #1300 (OPEN since 2024; no maintainer fix as of 2026-04-24). User reports GPU memory rises monotonically across init/preallocate/fetch/close cycles.
- Touches:
core/include/libvmaf/libvmaf_cuda.h— new publicvmaf_cuda_state_free()API declaration.core/src/cuda/common.c— newvmaf_cuda_state_free()implementation;vmaf_cuda_release()now callscuda_free_functions();vmaf_cuda_state_init()gets an outer failure unwind;init_with_primary_context()releases the retained primary context onfail_after_pop.core/src/cuda/ring_buffer.c—vmaf_ring_buffer_close()now unlocks + destroys the mutex before freeing.core/test/test_cuda_preallocation_leak.c— new GPU-gated reducer (10-cycle loop with full cleanup).core/test/test_cuda_pic_preallocation.c,core/test/test_cuda_buffer_alloc_oom.c— add missingvmaf_cuda_state_free()+vmaf_model_destroy()calls aftervmaf_close()in every test that allocates these.core/test/meson.build— register the new reducer underenable_cudaguard.- Invariants (load-bearing):
- Public contract: every caller of
vmaf_cuda_state_init()MUST callvmaf_cuda_state_free()AFTERvmaf_close()on any VmafContext that imported the state. Informalfree(cu_state)is a silent double-free hazard AFTER close (vmaf_close's vmaf_cuda_release already memset's + frees CudaFunctions internals; vmaf_cuda_state_free only frees the heap allocation itself). vmaf_cuda_release()freesCudaFunctionsvia a saved pointer AFTER thememset. Order matters —memsetfirst socu_state->fis zeroed in the caller's struct, then free via the saved local. Do not re-order.vmaf_ring_buffer_close()unlocks BEFORE destroying the mutex (POSIX requires the mutex be unlocked for destroy).- The cold-start unwind in
init_with_primary_contextreleasescuDevicePrimaryCtxRetain's retained context ifcuStreamCreateWithPriorityfails. - The ADR-0122 / ADR-0123
is_cudastate_empty()null-guards at the top of every publicvmaf_cuda_*entry must continue to compose with the newvmaf_cuda_state_free()(which accepts NULL directly and doesn't call through to the CUDA API). - The new free call order in callers is:
vmaf_close(vmaf)→vmaf_cuda_state_free(cu_state)→vmaf_model_destroy(model). Reversing the first two produces a use-after-free. - On upstream sync:
- Upstream has no
vmaf_cuda_state_free()as of 2026-04-24. Keep the fork's version on any conflict. If upstream eventually lands the same API with a different spelling, prefer upstream's spelling and add a compat alias — but do not break the fork's ABI. vmaf_cuda_release()'scuda_free_functions()call is fork-local. On rebase, keep it.- The ring-buffer
pthread_mutex_unlock+pthread_mutex_destroypair is fork-local. On rebase, keep it. - If upstream refactors
VmafCudaStateownership semantics (unlikely — their pattern has been "leaked state in a long- lived process is acceptable" historically), re-audit this ADR and the new public API. - Re-test on rebase:
ninja -C core/build-cuda
meson test -C core/build-cuda
# Expect: 40/40 pass including test_cuda_preallocation_leak.
# ASan leak-check:
cd libvmaf && meson setup build-asan-cuda \
-Db_sanitize=address -Denable_cuda=true -Denable_sycl=false \
--buildtype=debug
ninja -C build-asan-cuda
ASAN_OPTIONS='detect_leaks=1:leak_check_at_exit=1' \
build-asan-cuda/test/test_cuda_preallocation_leak
# Expect: 0 bytes leaked from core/src/* frames.
# (~180 bytes in libcuda.so.1 is expected — driver's process-
# lifetime cuInit cache, does not grow per cycle.)
0049 — CUDA graceful error propagation (ADR-0156)¶
- ADR: ADR-0156
- Upstream source: Netflix upstream issue #1420 (OPEN as of 2026-04-24). Reports that two concurrent VMAF-CUDA processes crash the second one at
vmaf_cuda_buffer_allocdue toCHECK_CUDA(cuMemAlloc)→assert(0)on OOM. - Touches:
core/src/cuda/cuda_helper.cuh— redefinedCHECK_CUDAfamily. New macrosCHECK_CUDA_GOTO+CHECK_CUDA_RETURN+ helpervmaf_cuda_result_to_errno. Oldassert(0)semantics removed entirely.core/src/cuda/common.c,core/src/cuda/picture_cuda.c,core/src/libvmaf.c— allCHECK_CUDA(...)sites converted; cleanup labels added where contexts / buffers were pushed / allocated.core/src/feature/cuda/integer_motion_cuda.c,integer_vif_cuda.c,integer_adm_cuda.c— same conversion; 12statichelpers promotedvoid → int.core/test/test_cuda_buffer_alloc_oom.c— new GPU-gated reducer.core/test/meson.build— register new test underenable_cudaguard.- Invariants (load-bearing):
CHECK_CUDA_GOTO/CHECK_CUDA_RETURNmust never callassert(0)orabort()on a CUDA error. Any regression back to the upstream abort-on-error semantics re-introduces Netflix#1420 and the NDEBUG footgun.- Every
CHECK_CUDA_GOTOtarget label must pop any previously-pushed CUDA context and free any partially-constructed buffers before returning the errno. The graceful path must not leak resources. vmaf_cuda_result_to_errnouses numericCUresultvalues directly (0 / 1 / 2 / 3 / 4 / 101 / 201 / 400) so host TUs that don't include<cuda.h>can transitively consume the mapping via the inline function. If upstream renumbersCUresultenum values (historically stable — they've been fixed since CUDA 1.0), re-audit the switch.- ADR-0122 / ADR-0123
is_cudastate_empty(...)guards at the top of every publicvmaf_cuda_*entry point must stay — they run before the CUDA API is touched and compose cleanly with the new error propagation. - Twelve
statichelper signatures in the feature extractors areint-returning (wasvoid): any upstream-port that restores thevoidreturn silently regresses the error path. - On upstream sync:
- Upstream Netflix still uses
assert(0)inCHECK_CUDAas of 2026-04-24. Keep the fork's macro definitions incuda_helper.cuhon any upstream conflict — this file is fork-local behaviour. - If upstream eventually lands Netflix#1420 with a similar refactor, prefer the fork's version unless upstream's has identical semantics (no
assert(0)/ noabort()/ translatesCUresultto-errno). Re-verifytest_cuda_buffer_alloc_oomafter rebase. - If upstream adds new
CHECK_CUDA(...)sites in a port, rewrite them toCHECK_CUDA_GOTO/CHECK_CUDA_RETURNas part of the port commit. - If upstream changes any of the 12
statichelper signatures back tovoid, re-promote them tointduring the merge. - Re-test on rebase:
ninja -C core/build-cuda
meson test -C core/build-cuda
# Expect: 39/39 pass including test_cuda_buffer_alloc_oom.
# Reducer check — verify the OOM-to-errno path is live:
meson test -C core/build-cuda test_cuda_buffer_alloc_oom -v
# Expect subtests: request 1 TiB → -ENOMEM; request 0 bytes → 0.
clang-tidy -p core/build-cuda --quiet \
core/src/cuda/common.c \
core/src/cuda/picture_cuda.c \
core/src/feature/cuda/integer_motion_cuda.c \
core/src/feature/cuda/integer_vif_cuda.c \
core/src/feature/cuda/integer_adm_cuda.c \
core/src/libvmaf.c
# Expect exit 0 on every file.
0049 — compute_motion / picture_copy signature changes (b949cebf upstream port)¶
- Upstream commit: Netflix/vmaf b949cebf (feature/motion: port several feature extractor options)
- Prerequisite commit: Netflix/vmaf d3647c73 (picture_copy: add channel parameter)
- PR: upstream/port-b949cebf-motion
Rebase-sensitive invariants:
-
compute_motionsignature change —compute_motion()incore/src/feature/motion.c/motion.hnow takes an extraint motion_decimateparameter (themotion_add_scale1flag). Any new caller added in the fork that callscompute_motion()must pass this parameter. The SIMD integer motion callers (motion_avx2.c,motion_avx512.c) do NOT callcompute_motion()— they use the SAD/convolution dispatch table directly and are unaffected. -
vmaf_image_sad_csignature change — similarly gainsint motion_add_scale1. Any caller in the fork must be updated. Currently only called fromcompute_motion()internally. -
picture_copysignature change — gainsint channelas the last parameter (0=Y, 1=U, 2=V). Every caller in the tree has been updated to pass0(luma). When adding new callers that need UV planes, pass1or2. The fork's CUDA/SYCL/Vulkan callers have been updated in this PR. -
Default behavior preserved — all new options default to no-op values.
motion_add_scale1=false,motion_add_uv=false,motion_blend_factor=1.0,motion_fps_weight=1.0,motion_filter_size=5(= DEFAULT_MOTION_FILTER_SIZE). Integer and float motion2 scores are bit-identical to pre-port baseline. -
vif_scale_frame_sdependency avoided — the upstream b949cebf motion.c importsvif_scale_frame_sfrom vif_tools.h. The fork does not have this function yet (vif options chain is deferred, Research-0024 Strategy E). The bilinear downscaler formotion_add_scale1is implemented as local static functions inmotion.c(motion_scale_bilinear,motion_bilinear_interp,motion_mirror_f). When upstream's vif options chain is eventually ported, reconcile by replacing these local functions withvif_scale_frame_s.
Reproducer:
# verify bit-exactness (default options, scores must be identical):
./core/build/tools/vmaf \
--reference testdata/ref_576x324_48f.yuv \
--distorted testdata/dis_576x324_48f.yuv \
--width 576 --height 324 --pixel_format 420 --bitdepth 8 \
--model path=model/vmaf_v0.6.1.json \
--feature motion --no_prediction --json --output /tmp/motion.json
# integer_motion2 scores must match pre-port baseline at 6 decimal places.
0048 — i4_adm_cm int32 rounding overflow deliberately preserved (ADR-0155)¶
- ADR: ADR-0155
- Upstream source: Netflix upstream issue #955 (OPEN since 2020; no maintainer response as of 2026-04-24). Reports that
add_bef_shift_flt[idx] = (1u << (shift_flt[idx] - 1))incore/src/feature/integer_adm.cscales 1–3 overflowsint32_t(1u << 31 = 0x80000000wraps to-2147483648). Rounding term is sign-negated; ADM scales 1–3 biased low by ≈1 LSB per summed term. - Touches (documentation-only):
docs/adr/0155-adm-i4-rounding-deferred-netflix-955.md— new ADR (this entry's anchor).core/src/feature/integer_adm.c— in-file warning comment above the overflow site (add_bef_shift_flt[]initialiser loop around line 1277). No code change.core/src/feature/AGENTS.md— invariant note under "Rebase-sensitive invariants".- Invariants (load-bearing — do NOT silently "fix"):
integer_adm.ckeepsint32_t add_bef_shift_flt[3]with the overflowing1u << 31assignment. The Netflix golden assertions (python/test/quality_runner_test.py,vmafexec_test.py,feature_extractor_test.py) encode the buggy ADM output. Project hard rule #1 (ADR-0024) prohibits changing those assertions.- Any "fix" that changes ADM numerical output must land together with a coordinated Netflix-authored golden-number update (the ADR-0142 Netflix-authority carve-out). Until Netflix#955 closes upstream, there is no authority to track.
- On upstream sync:
- If Netflix finally lands a fix for #955 (widening the rounding term to
uint32_torint64_t), sync the C-side fix AND the updatedassertAlmostEqualvalues in the same merge. Re-runmake test-netflix-goldenand/cross-backend-diffon the golden pairs to verify the new numbers are consistent across CPU / CUDA / SYCL. - Remove the in-file warning comment above the
add_bef_shift_fltinitialiser loop, flip ADR-0155 toSuperseded by ADR-NNNN, and drop this rebase-notes entry. - If upstream instead closes #955 as wont-fix, keep this entry verbatim and update the ADR status to note upstream's closure.
- Re-test on rebase (gates the invariant by confirming the golden numbers are unchanged):
ninja -C build
make test-netflix-golden
# Expect: VMAF mean 76.66890… on src01_hrc00/01_576x324 golden
# pair — bit-identical to pre-rebase.
0047 — vmaf_score_pooled -EAGAIN for pending features (ADR-0154)¶
- ADR: ADR-0154
- Upstream source: Netflix upstream issue #755 (OPEN as of 2026-04-24). Upstream maintainer closed the door on the streaming use case in 2020 ("you cannot call vmaf_score_pooled() in a loop"); fork reopens it via error-code semantics without changing the retroactive-write design.
- Touches:
core/src/feature/feature_collector.c—vmaf_feature_collector_get_scorereturns-EAGAIN(was-EINVAL) when the requested index is valid but not yet written.core/src/feature/feature_collector.h— inlinevmaf_feature_vector_get_scorenow returns-EINVALfor null/out-of-range and-EAGAINfor not-written (was-1for both). Added#include <errno.h>. Rename reserved__VMAF_FEATURE_COLLECTOR_H__guard toVMAF_FEATURE_COLLECTOR_INCLUDED.core/test/test_score_pooled_eagain.c— new 4-subtest reducer.core/test/meson.build— register the new test.- Invariants (load-bearing, enforced by the reducer):
vmaf_feature_collector_get_score(fc, name, &score, i)returns-EAGAINiff the featurenameis registered andiis in range butscore[i].written == false.- The return stays
-EINVALfor (a) null pointers, (b)i >= feature_vector->capacity, (c) unknown feature name. - The inline fast-path
vmaf_feature_vector_get_scoreuses the same split. - On upstream sync: upstream has not changed the error semantics since 2020. If they do (unlikely), keep the fork's
-EAGAIN— it is strictly more informative and downstream code depending on the split would regress. - Re-test on rebase:
ninja -C build && meson test -C build test_score_pooled_eagain
# Expect: 4/4 subtests pass.
# Reducer check:
git stash push core/src/feature/feature_collector.c core/src/feature/feature_collector.h
ninja -C build && meson test -C build test_score_pooled_eagain
# Expect: Fail: 1 (tests fail without -EAGAIN split).
git stash pop
0046 — float_ms_ssim min-dim guard (ADR-0153)¶
- ADR: ADR-0153
- Upstream source: Netflix upstream issue #1414 (OPEN as of 2026-04-24). No upstream fix has landed; fork adds the guard independently.
- Touches:
core/src/feature/float_ms_ssim.c— add#include "log.h"+#include "iqa/ssim_tools.h"+ amin_dim = GAUSSIAN_LEN << (SCALES - 1)check at the start ofinit; extract SIMD dispatch into a newms_ssim_init_simd_dispatchhelper to keepinitwithin the ADR-0141 60-line budget.core/test/test_float_ms_ssim_min_dim.c— new 3-subtest reducer.core/test/meson.build— register the new test executable.- Invariant (load-bearing, enforced by the reducer):
float_ms_ssim.initreturns-EINVALwhenw < 176 || h < 176, where 176 is computed dynamically from the filter constants. The magic number is not hardcoded — changingSCALESorGAUSSIAN_LENupstream will auto-update the minimum. - On upstream sync: if Netflix upstream lands a similar init-time guard, keep the fork's version — the helper name
ms_ssim_init_simd_dispatchis fork-local (introduced to satisfy ADR-0141) and upstream's patch won't match. Both guards should be compatible; re-verify the reducer after rebase. - Re-test on rebase:
ninja -C build && meson test -C build test_float_ms_ssim_min_dim
# Expect: 3/3 subtests pass.
# Reducer check (confirms the guard is load-bearing):
git stash push core/src/feature/float_ms_ssim.c
ninja -C build && meson test -C build test_float_ms_ssim_min_dim
# Expect: Fail: 1 (tests fail without the guard).
git stash pop
0045 — vmaf_read_pictures monotonic-index guard (ADR-0152)¶
- ADR: ADR-0152
- Upstream source: Netflix upstream issue #910 (OPEN as of 2026-04-24). No upstream fix has landed; the fork adds the guard independently, per the 2021-10-14 maintainer comment that recommended exactly this shape.
- Touches:
core/src/libvmaf.c— addunsigned last_index+bool have_last_indexfields toVmafContext; prepend a monotonic-index check insideread_pictures_validate_and_prep(returns-EINVALon duplicates / regressions); update the two new fields at the tail of the same helper on success.core/test/test_read_pictures_monotonic.c— new 3-subtest reducer covering the Netflix#910 sequence and the two classes of rejection (duplicate, out-of-order).core/test/meson.build— register the new test executable.- Invariant (load-bearing, enforced by the reducer):
vmaf_read_pictures(vmaf, ref, dist, index)returns-EINVALwhenhave_last_index && index <= last_index. Flush (vmaf_read_pictures(vmaf, NULL, NULL, 0)) routes toflush_contextbefore the guard runs — flushing remains always-available independent of the last accepted index. - On upstream sync:
- If Netflix upstream eventually lands a similar guard at the API boundary, keep the fork's version — the helper function name (
read_pictures_validate_and_prep) is fork-local (ADR-0146), upstream's patch will target a different insertion point. Both guards should be compatible; re-verify the reducer after rebase. - If upstream instead lands an internal reordering mechanism (buffer-and-sort frames before dispatch), revisit this decision — the fork's API-level contract is stricter and may need to relax to match. Open a new ADR if so.
- Re-test on rebase:
ninja -C build && meson test -C build test_read_pictures_monotonic
# Expect: 3/3 subtests pass.
# Reducer check (confirms the guard is load-bearing):
git stash push core/src/libvmaf.c
ninja -C build && meson test -C build test_read_pictures_monotonic
# Expect: Fail: 1 (the test rejects the un-guarded behaviour).
git stash pop
0044 — i686 (32-bit x86) build-only CI job (ADR-0151)¶
- ADR: ADR-0151
- Upstream source: Netflix upstream issue #1481 (OPEN as of 2026-04-24). Reports i686 compile failure on
_mm256_extract_epi64. Workaround documented in the issue:-Denable_asm=false. - Touches:
build-aux/i686-linux-gnu.ini— new cross-file; gcc +-m32+cpu_family = 'x86'/cpu = 'i686'. Noexe_wrapper..github/workflows/libvmaf-build-matrix.yml— new matrix row withi686: trueflag + new install-deps step forgcc-multilib+g++-multilib; existing "Run tests" + "Run tox tests (ubuntu)" steps widened with&& !matrix.i686guards.- Invariants:
- The i686 matrix row pins
-Denable_asm=false— this is the upstream-documented workaround for_mm256_extract_epi64's missing declaration on 32-bit x86 targets. Do NOT remove the flag without first gating every_mm256_extract_epi64call site incore/src/feature/x86/adm_avx2.c+motion_avx2.c+adm_avx512.con__x86_64__. Removing the flag naively will re-break the build. - No
exe_wrapperin the cross-file: meson marks tests asSKIP 77even though the host can run i686 binaries natively. Build-only gate by design. - On upstream sync:
- If upstream Netflix fixes #1481 at source (by gating the intrinsic calls on
__x86_64__or by emulating via two_mm256_extract_epi32halves), sync the fix and re-enable ASM on the i686 row (drop-Denable_asm=falsefrommeson_extra). Re-verify bit-exactness via/cross-backend-diffon the x86_64 golden pair. - If upstream marks i686 unsupported in meson (e.g. via a hard error), the fork's i686 row should be removed or downgraded to
continue-on-error: true. - Re-test on rebase (Ubuntu host with
gcc-multilib):
meson setup libvmaf core/build-i686 \
--cross-file=build-aux/i686-linux-gnu.ini \
-Denable_asm=false \
-Denable_cuda=false -Denable_sycl=false
ninja -C core/build-i686
file core/build-i686/tools/vmaf
# Expect: ELF 32-bit LSB pie executable, Intel i386
CI runs this same sequence via the new matrix row.
0058 — Tiny-AI Netflix corpus training scaffold (ADR-0252)¶
- ADR: ADR-0252.
- Upstream source: fork-local. Netflix/vmaf has no tiny-AI training harness or MCP server.
- Touches:
ai/— training harness;NflxLocalDatasetloader reads from--data-root(never from a hardcoded path).docs/ai/training-data.md— corpus path convention and loader API docs; purely additive.mcp-server/vmaf-mcp/tests/test_smoke_e2e.py— new e2e smoke test; references only committed golden fixtures.- Invariants (load-bearing):
- Data path is local-only.
.workingdir2/netflix/is gitignored; no YUV from this corpus is ever committed. The--data-rootCLI flag must remain the sole mechanism for locating the corpus. - Smoke test uses only committed fixtures.
test_smoke_e2e.pyreferencespython/test/resource/yuv/src01_hrc00_576x324.yuv(a committed golden file), never the local corpus path. On upstream sync the golden YUV path must stay stable. - No Netflix golden assertion is modified. The
places=4tolerance intest_smoke_e2e.pyasserts against thevmaf_v0.6.1CPU reference; it is not a golden assertion and may be adjusted by/regen-snapshotswith justification. - On upstream sync: zero interaction with Netflix upstream. The
ai/subtree andmcp-server/are wholly fork-local; upstream merges are conflict-free here. If Netflix ever ships a training harness, reconcile separately. - Re-test on rebase:
cd mcp-server/vmaf-mcp && python -m pytest tests/test_smoke_e2e.py -v
# Requires: meson compile -C build (vmaf binary)
# Skips automatically if binary or golden YUV is absent.
0085 — Research-0030 Phase-3b multi-seed validation (Gate 1 passed)¶
- No ADR. Empirical research digest closing Gate 1 of the 3-gate v2 validation chain. Architecture decision unchanged.
- Upstream source: fork-local. Netflix has no multi-seed validation surface for tiny-AI training.
- Touches (additive only):
docs/research/0030-phase3b-multiseed-validation.md— per-seed PLCC tables + stability analysis + Gate 2/3 plan.ai/scripts/phase3_subset_sweep.py— adds--seedsflag (comma-separated list) + per-seed result aggregation.CHANGELOG.mdUnreleased § Added.- Invariants (rebase-relevant):
- The +0.0175 Δ is multi-seed mean PLCC, not seed-0 PLCC. Don't cite the +0.0106 from Research-0029 once Research-0030 lands; the multi-seed number is more trustworthy.
- Subset B is more stable than canonical-6 across seeds. Don't ship a v2 model citing single-seed numbers — always report multi-seed mean ± seed-mean-std for any tiny-AI metric in a future digest.
- The
--seedsflag aggregates by flattening (seed × fold) pairs. The reportedmean_plccis the mean of alln_seeds × n_foldsmeasurements;seed_mean_plcc_stdis the std across per-seed means, which is the right number for "is the result seed-stable". - On upstream sync: zero interaction. Fork-only research.
- Re-test on rebase: documentation-only PR; the runs/ files reproduce from the canonical command.
0084 — Research-0029 Phase-3b StandardScaler retry (positive result)¶
- No ADR. Empirical research digest; revives the Research-0026 hypothesis after the Research-0028 negative result. The architectural decision (ship
vmaf_tiny_v2) is gated on three validation steps documented in the digest §"Required before shipping". - Upstream source: fork-local. Netflix has no tiny-AI preprocessing-sensitivity analysis surface.
- Touches (additive only):
docs/research/0029-phase3b-standardscaler-results.md— per-fold tables + apples-to-apples comparison + 3-gate pre-shipping checklist.ai/scripts/phase3_subset_sweep.py— adds--standardizeflag +_standardize_inplacehelper.CHANGELOG.mdUnreleased § Added.- Invariants (rebase-relevant):
- StandardScaler statistics MUST be fit per-fold on the train split only. Fitting on the full data would leak held-out information into LOSO; the
_standardize_inplacehelper enforces this by taking only the train slice as input. - A shipped
vmaf_tiny_v2.onnxMUST bundle its scaler(mean, std)in the sidecar JSON per ADR-0049 — otherwise inference applies different normalisation than training and the win evaporates. Currently UN-implemented; tracked as a §"Caveats" #5 follow-up. - Subset B's feature list is the load-bearing finding:
adm2,adm_scale3,vif_scale2,motion2,ssimulacra2,psnr_hvs,float_ssim. Phase-3c experiments may shift the optimal arch / lr / epochs but should keep this set. - On upstream sync: zero interaction. Fork-only research.
- Re-test on rebase: documentation-only PR; the runs/ files are reproducible from the
--standardizeinvocation in §"Reproducer".
0082 — Research-0028 Phase-3 subset sweep (negative-result digest)¶
- No ADR. Empirical research digest. The architectural decision (no v2 model ships from this Phase) is governed by Research-0027's pre-registered stopping rule.
- Upstream source: fork-local. Netflix has no tiny-AI subset- sweep surface.
- Touches (additive only):
docs/research/0028-phase3-subset-sweep.md— per-fold tables adline + standardisation caveat + Phase-3b/c/d follow-ups.CHANGELOG.mdUnreleased § Added.- Invariants (rebase-relevant):
- canonical-6 stays the default until Phase-3b lands a ≥ 0.005 PLCC win (per Research-0027 stopping rule).
- The PLCC drop is most likely a feature-scale issue, not evidence the new features lack signal. Don't cite this digest to retire
ssimulacra2/adm_scale3from the candidate pool; re-test withStandardScalerfirst. - Phase-3 results are seed=0 only. Any v2-shipping decision needs 3-seed mean±std and KoNViD cross-check.
- On upstream sync: zero interaction. Fork-only research.
- Re-test on rebase: documentation-only PR; runs/ files are reproducible from the canonical command in §"Reproducer".
0081 — Research-0027 Phase-2 feature importance results¶
- No ADR. Empirical research digest closing Research-0026 Phase 2; the architectural decision (Subset A / B / C) is deferred to Phase-3 results in a future digest.
- Upstream source: fork-local. Netflix has no cross-metric feature-importance analysis surface.
- Touches (additive only):
docs/research/0027-phase2-feature-importance.md— per-method top-10 + consensus + redundancy + Phase-3 subset recommendations.CHANGELOG.mdUnreleased § Added.- Invariants (rebase-relevant):
- Consensus top-10 is the load-bearing finding:
adm2,adm_scale3,ssimulacra2,vif_scale2. Phase-3 candidate subsets MUST include all four. - The 11-pair redundancy table is corpus-specific — measurements on Netflix Public 9-source. KoNViD-1k cross- check is a Phase-3 prerequisite if Subsets B/C advance.
runs/full_features_netflix.parquetandruns/full_features_correlation.jsonstay gitignored. Reproducer in §"Reproducer" regenerates both.- On upstream sync: zero interaction. Fork-only research.
- Re-test on rebase: documentation-only PR; the
runs/files are reproducible from the canonical commands.
0080 — Phase-2 analysis scripts (Research-0026 Phase 2 prep)¶
- No ADR. Pure analysis scaffolding; the architectural decision (which features to ship in v2) is gated on Phase 2's numerical output via Research-0027.
- Upstream source: fork-local. Netflix has no tiny-AI training nor cross-metric correlation tooling.
- Touches (additive only):
ai/scripts/extract_full_features.py— parquet extractor over Netflix corpus withFULL_FEATURES. Per-clip JSON cache at$XDG_CACHE_HOME/vmaf-tiny-ai-full/<source>/<dis_stem>.json.ai/scripts/feature_correlation.py— Pearson + MI + LASSO- consensus top-K analyser; outputs JSON.
ai/tests/test_feature_correlation.py— 5 pytest cases against synthetic parquet (no libvmaf dependency).CHANGELOG.mdUnreleased § Added.- Invariants (rebase-relevant):
- The per-clip JSON cache and the
FULL_FEATUREStuple must stay in lock-step. If the tuple grows (or shrinks), pre-existing cache files become stale and silently misalign their storedper_framecolumns with the new tuple. The extractor MUST be re-run with a cleared cache whenFULL_FEATURESchanges. Regression hint:test_default_features_unchangedintest_feature_sets.pyalready guards the canonical 6; extend coverage toFULL_FEATURESif rebases touch it. motion3resolves to extractormotion_v2in_METRIC_TO_EXTRACTOR, notmotion3(the upstream-canonical extractor name in the integer_motion_v2 module). The CLI--feature motion3does NOT exist. The JSON output key isinteger_motion3which_lookupfinds via theinteger_fallback.admandvifaggregates are NOT inFULL_FEATURES. The integer extractor emitsinteger_adm2andinteger_vif_scale0..3but no bareadm/vif. Listing them produced all-NaN columns in v1 — fixed in PR #185 amend.- On upstream sync: zero interaction. Pure fork-side analysis tooling.
- Re-test on rebase:
pytest ai/tests/test_feature_correlation.py ai/tests/test_feature_sets.py -v
# Expect: 14 passed in <1 s.
0079 — Tiny-AI feature-set registry (Research-0026 Phase 1)¶
- No ADR. Pure additive extension of an existing module; the architectural decision (which features, which model) lives in Research-0026's go/no-go gate after Phase 2.
- Upstream source: fork-local. Netflix/vmaf has no tiny-AI training pipeline.
- Touches (additive only):
ai/data/feature_extractor.py— addsFULL_FEATURES(21 entries),FEATURE_SETSregistry,resolve_feature_set()helper._METRIC_TO_EXTRACTORgrew 11 → 25 entries.ai/tests/test_feature_sets.py— new 9-test smoke suite.CHANGELOG.mdUnreleased § Added.- Invariants (rebase-relevant — these are load-bearing):
DEFAULT_FEATURESstays the canonical 6-tuple matchingvmaf_v0.6.1's SVR input layout. Testtest_default_features_unchangedis the regression guard; any quiet broadening would invalidate every shipped tiny-AI ONNX (input-dim baked into the model). If a future change must broaden the default, ship a paired model swap under ADR-0049 sidecar policy.FULL_FEATURESexcludeslpipsandfloat_momentper Research-0026 §"Open questions" Q1. Testtest_full_features_excludes_lpips_and_momentenforces. Adding either would re-classify the experiment from "tiny model on classical features" to "ensemble of DNNs".- Every entry in
FULL_FEATURESMUST have an entry in_METRIC_TO_EXTRACTOR. Testtest_every_full_feature_has_extractor_mappingis the guard — without the mapping the libvmaf CLI silently emits NaN columns for the missing metric. - On upstream sync: zero interaction. Fork-only training surface.
- Re-test on rebase:
0078 — Research-0026 cross-metric feature fusion plan¶
- No ADR. Pure research-plan digest; the architectural decision (which features to add) is deferred to Research-0027 follow-up after Phase 2 numbers land.
- Upstream source: fork-local. Netflix/vmaf has no tiny-AI training and no broader-feature-set hypothesis under investigation.
- Touches (additive only):
docs/research/0026-cross-metric-feature-fusion.md— 4-phase experimental plan + cost estimate + go/no-go criteria.CHANGELOG.mdUnreleased § Added.- Invariants (rebase-relevant):
- The 6-feature canonical baseline (
adm2,vif_scale0..3,motion2) stays the default. Any v2 model is opt-in via a newfeature_setfield in the sidecar JSON; existingvmaf_tiny_v1.onnxusers get the same numbers. lpipsis OUT of the candidate pool (Phase 1/2). It's DNN-based and would blur the line between "tiny model on classical features" and "ensemble of DNNs". Revisit only if classical features can't close the gap.- On upstream sync: zero interaction. Pure fork-side research planning.
- Re-test on rebase: documentation-only; no test surface.
0077 — Research-0025 FoxBird outlier resolved via KoNViD combined training¶
- No ADR. Empirical research digest closing the open question in Research-0023 §5; no architecture or policy decision. Pure documentation of an empirical result.
- Upstream source: fork-local. Netflix/vmaf has no tiny-AI training, no KoNViD-1k integration, and no LOSO eval surface.
- Touches (additive only):
docs/research/0025-foxbird-resolved-via-konvid.md— per-clip table + comparison to Netflix-only baselines + interpretation + caveats + next-experiment list.CHANGELOG.mdUnreleased § Added.- Invariants (rebase-relevant):
- The training-fit per-clip numbers in §"Per-clip result" are NOT held-out generalisation metrics — FoxBird is in the training set. The proper validation is the LOSO sweep on the combined corpus (§"Next experiments" #1). Don't cite the 0.9936 FoxBird PLCC as a generalisation number; cite it as "training-fit on combined corpus, 5.4× RMSE improvement vs Netflix-only".
- Combined trainer command line is canonical. The reproduction recipe in §"Setup" includes
--seed 0,--konvid-val-fraction 0.1,--val-source Tennis,--val-mode netflix-source-and-konvid-holdout. Changing any knob invalidates the per-clip numbers. runs/tiny_combined_canonical/stays gitignored. The final ONNX is reproducible from the parquet + Netflix corpus + the canonical CLI; the durable record is the digest's table.- On upstream sync: zero interaction. Research digest is fork-only.
- Re-test on rebase:
python ai/train/train_combined.py \
--netflix-root .workingdir2/netflix \
--konvid-parquet ai/data/konvid_vmaf_pairs.parquet \
--model-arch mlp_small --epochs 30 --batch-size 256 --lr 1e-3 \
--val-mode netflix-source-and-konvid-holdout \
--val-source Tennis --konvid-val-fraction 0.1 --seed 0 \
--out-dir runs/tiny_combined_canonical
# Expect: FoxBird PLCC ≈ 0.9936 ± 1e-3 (numerical-noise floor),
# mean PLCC ≥ 0.9983 across 9 Netflix clips.
0076 — Research-0024 vif/adm upstream-divergence digest (Strategy E doc)¶
- No ADR. Pure documentation digest; the divergence decisions it ratifies are already governed by ADR-0138 / 0139 / 0142 / 0143 (vif SIMD bit-exactness contract) and ADR-0024 (Netflix golden-data immutability). The digest itself fits the per-PR research-digest deliverable bar from ADR-0108.
- Upstream source: forward-looking — pre-emptively documents the fork's non-port of Netflix
4ad6e0ea/41d42c9e/bc744aa3/8c645ce3(vif chain) and4dcc2f7c(float_adm chain). Strategy A onb949cebfmotion chain stays approved. - Touches (additive only):
docs/research/0024-vif-upstream-divergence.md— 5-strategy decision matrix + numerical-risk analysis for each chain.core/src/feature/AGENTS.md— two new "rebase-sensitive invariants" entries pinning the vif and adm divergences.CHANGELOG.mdUnreleased § Changed.- Invariants (rebase-relevant — these are the whole point):
- Do not port
4ad6e0ea(vif runtime helpers) or8c645ce3(vif prescale options) verbatim. They replace the precomputedvif_filter1d_table_stable whose frozenconst floatGaussians make AVX2 == AVX-512 == NEON == scalar bit-for-bit. A future opt-in second-path port (Strategy C, runtime helpers behind--vif-prescale != 1) is allowed but must not touch the default code path. - Do not port
4dcc2f7cfloat_adm options chain. The 12-parametercompute_admsignature change cascades through SIMD (avx2 / avx512 / neon) and 3 GPU backends (vulkan / cuda / sycl). The newaimfeature has no fork- side golden values; defer until concrete user demand. - Mirror bugfix
41d42c9eis a separate decision. Must come paired withplaces=4 → places=3golden loosening per ADR-0142 Netflix-authority precedent. Not part of Strategy E; eligible for a focused single-purpose PR if any shipped model drifts more thanplaces=3because of the missing fix. b949cebfmotion chain port stays APPROVED under Strategy A (verbatim, float_motion-side only). Float_motion has no precomputed-table investment to protect; existing fork integer_motion already has 6/9 of these options; cheap to mirror onto float_motion.- On upstream sync: zero conflict — pure additions to research/ and AGENTS.md.
- Re-test on rebase: documentation-only PR; rendered markdown is the only verification surface.
# Re-run the diff scan that produced the digest (catches new
# upstream commits since 9dac0a59):
git fetch upstream && git log --pretty=format:'%h %s' \
upstream/master ^origin/master --since="2026-01-01" \
-- core/src/feature/{float_,integer_,}{vif,motion,adm,cambi}*.{c,h} \
core/src/feature/{vif,motion,adm,cambi}_options.h \
| head -30
# If new vif / adm option ports appear, update Research-0024 §"Same
# divergence test for motion + float_adm" before deciding to port.
0075 — Upstream 798409e3 + 314db130 ports (CUDA null-deref + remove all.c)¶
- No ADR. Pure upstream cherry-picks per ADR-0108 carve-out ("pure upstream syncs and
port-upstream-commitPRs are exempt"). - Upstream source:
798409e3(Lawrence Curtis, 2026-04-20): "Fix null deref crash on prev_ref update in pure CUDA pipelines"314db130(Kyle Swanson, 2026-04-28): "libvmaf/feature: remove empty translation unit all.c"- Touches (additive / removal only):
core/src/libvmaf.c— addsif (ref && ref->ref)guard beforevmaf_picture_ref(&vmaf->prev_ref, ref)at the two threaded paths (threaded_enqueue_oneline 1057 andthreaded_read_pictures_batchline 1105). Main path at line 1597 already has the guard.core/src/feature/all.c— file deleted.core/src/meson.build— drops thefeature_src_dir + 'all.c'line.core/src/feature/offset.c— updates the// NOLINTNEXTLINEcomment to dropall.cfrom the list of per-feature consumers.CHANGELOG.mdUnreleased § Fixed (798409e3) + § Changed (314db130).- Invariants (rebase-relevant):
- The fork has THREE prev_ref update sites; all need the
if (ref && ref->ref)guard. The mainvmaf_read_picturespath already had it (viaread_pictures_update_prev_refhelper); the threaded paths (#ifdef VMAF_BATCH_THREADING) inherited the unguarded shape from upstream's old code. Future upstream rebases must preserve all three guards even if Netflix refactors the threaded paths. all.cdeletion is symbol-safe. Allcompute_*functions it forward-declared are reached via per-extractor TUs that#includethe relevant<feature>.h. No external linker dependency onall.c's symbols.- On upstream sync: zero conflict expected — fork now matches upstream tip on these two surfaces.
- Re-test on rebase:
meson setup build-cpu libvmaf -Denable_cuda=false -Denable_sycl=false \
-Denable_vulkan=disabled
ninja -C build-cpu
meson test -C build-cpu # 37 tests, all pass.
0074 — Combined Netflix + KoNViD-1k trainer driver¶
- No ADR. Pure engineering follow-up; the architecture rationale is fully covered by ADR-0203 (training-prep architecture) and Research-0023 §5 (FoxBird-class outlier needs broader corpus).
- Upstream source: fork-local. Netflix/vmaf has no tiny-AI trainer.
- Stacks on the KoNViD-1k loader bridge (PR #178 / rebase-note 0073). Rebase order: land 0073 first.
- Touches (additive only):
ai/train/train_combined.py— concatenating trainer that reuses_build_model/_train_loop/export_onnxfromai/train/train.py.ai/tests/test_train_combined_smoke.py— 5 pytest cases (key splitter +--epochs 0paths, no libvmaf or real corpus required).docs/ai/training.md— "Combining KoNViD with the Netflix corpus" subsection rewritten from "follow-up" to runnable.CHANGELOG.mdUnreleased § Added.- Invariants (rebase-relevant):
- Reuse the canonical training-loop helpers. Don't fork
_build_model/_train_loop/export_onnxinto this file. Both trainers must share the model factory so a future change (e.g. addingmlp_large) lands in one place. - KoNViD train/val splits hold out whole clip keys, not random frames. A frame-level split would let frames from the same clip leak across train/val and inflate PLCC by 5-10 pp (well-known VQA pitfall — same reasoning as ADR-0203's Netflix 1-source-out split).
- Missing data falls back, not errors. Missing
--konvid-parquet→ Netflix-only path. Missing--netflix-root→ KoNViD-only path. Both missing → initial- weights ONNX export +rc=0so the smoke command always produces a deterministic artefact. - On upstream sync: zero interaction; pure fork-local trainer.
- Re-test on rebase:
pytest ai/tests/test_train_combined_smoke.py -v
# Expect: 5 passed (under ~3 s, no libvmaf required).
python ai/train/train_combined.py --epochs 0 \
--netflix-root /tmp/missing --konvid-parquet /tmp/missing.parquet \
--out-dir /tmp/combined_smoke
# Expect: <out-dir>/mlp_small_combined_final.onnx written, rc=0.
0073 — KoNViD-1k → VMAF-pair acquisition + loader bridge¶
- No ADR. Acquisition + loader pieces are pure additions; the methodology fits inside ADR-0203 / Research-0019.
- Upstream source: fork-local. KoNViD-1k integration is a fork-only training-data play.
- Touches (additive only):
ai/scripts/konvid_to_vmaf_pairs.py— acquisition pipeline.ai/train/konvid_pair_dataset.py—KoNViDPairDatasetclass mirroringNetflixFrameDataset's interface.ai/tests/test_konvid_pair_dataset.py— 5 pytest cases.docs/ai/training.md— new "C1 (KoNViD-1k corpus)" section.CHANGELOG.mdUnreleased § Added.- Invariants (rebase-relevant):
KoNViDPairDatasetmirrorsNetflixFrameDatasetshape.feature_dim == 6,numpy_arrays() → (X, y)returns(n_frames, 6)+(n_frames,). IfNetflixFrameDataset's feature order changes, mirror it here.- Acquisition parquet schema is fixed. Required columns:
key,frame_index,vif_scale0..3,adm2,motion2,vmaf. Add freely; do NOT rename / drop those. ai/data/konvid_vmaf_pairs.parquetand$VMAF_TINY_AI_CACHE/konvid-1k/stay gitignored. They regenerate from raw KoNViD.mp4sources.- On upstream sync: zero interaction.
- Re-test on rebase:
pytest ai/tests/test_konvid_pair_dataset.py -v
# Expect: 5 passed
python ai/scripts/konvid_to_vmaf_pairs.py --max-clips 5
# Expect: ~7 s wall, ai/data/konvid_vmaf_pairs.parquet with
# 5 unique keys × ~200 frames each.
0072 — Tiny-AI 3-arch LOSO eval harness + Research-0023¶
- No ADR. Methodology fits inside Research-0023; ADR-0203 already covers the training-prep architecture and the three-arch sweep concept.
- Research digest:
docs/research/0023-loso-3arch-results.md. - Upstream source: fork-local. Netflix/vmaf has no LOSO eval surface.
- Touches (additive only):
ai/scripts/eval_loso_3arch.py— new harness; reuses the_load_session+_load_clip+CLIPShelpers fromeval_loso_mlp_small.py(PR #165).docs/research/0023-loso-3arch-results.md— methodology + per-fold tables formlp_small/mlp_medium/linear.CHANGELOG.mdUnreleased § Added.- Invariants (rebase-relevant):
- Reuse the PR #165 helpers. Don't fork the
_load_sessionexternal-data workaround into a copy — both scripts must keep using the same import. If a follow-up re-exports the shipped baselines with correctedexternal_data.location, both scripts deprecate the workaround simultaneously. runs/andmodel/tiny/training_runs/stay gitignored. The harness writesruns/loso_eval/loso_3arch_eval.{json,md}; the durable record is the table in Research-0023 §2 + the per-fold tables in §3. Regenerate via the loop in §6 of the digest.- On upstream sync: zero interaction. Pure fork-local evaluation harness.
- Re-test on rebase:
python ai/scripts/eval_loso_3arch.py
diff <(jq -r '.archs.mlp_small.aggregate.mean_plcc' runs/loso_eval/loso_3arch_eval.json) <(echo 0.9808)
diff <(jq -r '.archs.mlp_medium.aggregate.mean_plcc' runs/loso_eval/loso_3arch_eval.json) <(echo 0.9727)
diff <(jq -r '.archs.linear.aggregate.mean_plcc' runs/loso_eval/loso_3arch_eval.json) <(echo 0.3679)
# Expect: identical lines on a populated cache + identical fold ONNX.
0071 — T7-16 ADM Vulkan/SYCL drift verified-resolved (doc close)¶
- No ADR. Verification-only close, sister of T7-15.
- Upstream source: fork-local. ADM cross-backend gate is a fork-only test surface; Netflix/vmaf has no Vulkan or SYCL backend.
- Touches (additive only):
docs/state.md— new "Recently closed" row for T7-16..workingdir2/BACKLOG.md— T7-16 row marked closed (local- only planning dossier; gitignored).CHANGELOG.mdUnreleased § Fixed.- Invariants (rebase-relevant):
places=4cross-backend ADM contract. Empiricaladm_scale2max_abs_diff is now 1e-6 (print floor; ULP=0) on Vulkan device 0 (NVIDIA), device 1 (Mesa anv on Arc), and SYCL device 0 (Arc); residualadm_scale1 ≈ 3.1e-5andadm2 ≈ 5e-6on 1/48 frames passplaces=4(5e-5 tolerance) but failplaces=5. Hold the gate atplaces=4.- No ADM kernel source change. Fix is environmental (NVCC + driver + SYCL runtime).
- On upstream sync: zero interaction.
- Re-test on rebase:
python3 scripts/ci/cross_backend_vif_diff.py \
--vmaf-binary core/build/tools/vmaf \
--feature adm --backend vulkan --device 0 --places 4 \
--reference python/test/resource/yuv/src01_hrc00_576x324.yuv \
--distorted python/test/resource/yuv/src01_hrc01_576x324.yuv \
--width 576 --height 324
# Expect: 0/48 mismatches across all 5 ADM metrics.
0070 — T7-15 motion CUDA/SYCL drift verified-resolved (doc close)¶
- No ADR. Verification-only close; no code change in PR #172.
- Upstream source: fork-local. Cross-backend gate is a fork-only test surface; not in Netflix/vmaf.
- Touches (additive only):
docs/state.md— "Recently closed" row for T7-15..workingdir2/BACKLOG.md— T7-15 row marked closed (local- only planning dossier; gitignored).CHANGELOG.mdUnreleased § Fixed.- Invariants (rebase-relevant):
- The
places=4cross-backend gate stays atplaces=4. Empirical max_abs_diff is currently 0.0 (CUDA) or 1e-6 (SYCL/ Vulkan, JSON%frounding floor); tightening toplaces=5could be tempting but the 1e-6 print-floor would then make the SYCL + Vulkan rows fail. Hold atplaces=4until--precision=maxis wired into the diff tool. - No motion-kernel source change. PR #172 didn't modify
core/src/feature/cuda/integer_motion/*.cuorcore/src/feature/sycl/integer_motion_sycl.cpp. The fix is environmental (NVCC + driver), so the next CI run on a fresh image needs to be re-verified against the gate. - On upstream sync: zero interaction.
- Re-test on rebase:
python3 scripts/ci/cross_backend_vif_diff.py \
--vmaf-binary core/build/tools/vmaf \
--reference python/test/resource/yuv/src01_hrc00_576x324.yuv \
--distorted python/test/resource/yuv/src01_hrc01_576x324.yuv \
--width 576 --height 324 --feature motion --backend cuda \
--places 4
# Expect: 0/48 mismatches, max_abs_diff = 0.0
0069 — libvmaf_vulkan.h installed under prefix (build bug)¶
- No ADR. Build-system bug fix; matches existing CUDA / SYCL install conditions.
- Upstream source: fork-local. Vulkan backend is fork-only; Netflix/vmaf has no
libvmaf_vulkan.h. - Touches:
core/include/core/meson.build— adds anis_vulkan_enabledgate that handles thefeatureoption'senabled/autostates; appendslibvmaf_vulkan.htoplatform_specific_headerswhen active.CHANGELOG.mdUnreleased § Fixed.- Invariants (rebase-relevant):
- Install rule mirrors the CUDA / SYCL pattern but uses the feature-option API. The
is_cuda_enabled = get_option('enable_cuda') == trueboolean idiom doesn't apply toenable_vulkanbecause that's a feature option, not a boolean. Use.enabled() or .auto(). Don't "simplify" to== true— that would silently drop the install in theautostate. - Pairs with
ffmpeg-patches/0006-libvmaf-add-libvmaf-vulkan-filter.patchwhich probes for the header viacheck_pkg_config libvmaf_vulkan "libvmaf >= 3.0.0" libvmaf/libvmaf_vulkan.h vmaf_vulkan_state_init_external. Removing the install rule re-introduces lawrence's 2026-04-28 symptom: FFmpeg silently drops thelibvmaf_vulkanfilter despite--enable-libvmaf-vulkan. - On upstream sync: zero interaction; Vulkan backend is fork-only.
- Re-test on rebase:
cd libvmaf
CC=icx CXX=icpx meson setup build -Denable_vulkan=enabled \
-Denable_cuda=true -Denable_sycl=true -Db_lto=false
ninja -C build
meson install -C build --destdir /tmp/libvmaf-install
ls /tmp/libvmaf-install/usr/local/include/libvmaf/libvmaf_vulkan.h
# Expect: file exists.
0066 — --backend cuda inverted-gpumask fix (CLI bug)¶
- No ADR. Bug fix; behaviour now matches the public-header
VmafConfiguration::gpumaskcontract. - Upstream source: fork-local. The
--backendCLI selector was added by the fork (Netflix/vmaf has no exclusive-backend selector). - Touches (additive + 1-line behavioural fix):
core/tools/cli_parse.c::parse_cli_args—--backend cudabranch setsgpumask = 0(wasgpumask = 1).core/test/test_cli_parse.c— 5 new regression tests (test_backend_{cpu,cuda_engages_cuda,cuda_preserves_explicit_gpumask,sycl,vulkan}) plusrun_aom_ctc_tests/run_backend_testshelper split to keeprun_testsunder the function-size budget.CHANGELOG.mdUnreleased § Fixed.- Invariants (rebase-relevant):
VmafConfiguration::gpumasksemantics:if gpumask: disable CUDA.compute_fex_flagsinsrc/libvmaf.croutes CUDA only whengpumask == 0. Any code path that sets a non-zerogpumaskto "request CUDA" silently disables it. The CLI's--backend cudabranch must setgpumask = 0and rely onuse_gpumask = trueto triggervmaf_cuda_state_init. Do not "fix" this back togpumask = 1— it's the bug being fixed.- Explicit
--gpumask=N --backend cudapreserves N. A user who passes--gpumask=2already hasuse_gpumask = true, so the--backend cudabranch's defaulting block (gated on!settings->use_gpumask) is skipped. Thetest_backend_cuda_preserves_explicit_gpumaskregression locks this in. - On upstream sync: zero interaction;
--backendis fork-only. - Re-test on rebase:
./build/test/test_cli_parse | grep -E 'backend_'
# Expect: 5 backend tests pass.
build/tools/vmaf -r REF -d DIS -w 576 -h 324 -p 420 -b 8 \
--model "path=model/vmaf_v0.6.1.json" --threads 1 \
--backend cuda --output cuda.json --json -q
python3 -c "import json; d=json.load(open('cuda.json')); \
assert len(d['frames'][0]['metrics']) == 12, 'CUDA not engaged'"
0067 — Tiny-AI PTQ accuracy across Execution Providers (T5-3e)¶
- No ADR. Investigation/measurement PR; ADR-0129 already governs the PTQ workstream. Findings update
docs/research/0006-tinyai-ptq-accuracy-targets.md§"GPU-EP quantisation" — that section was previously a deferred-open-question; it is now the empirical landing spot. - Research digest: same file (Research-0006).
- Upstream source: fork-local. Netflix/vmaf does not ship a PTQ harness or any tiny-AI ONNX path.
- Touches (additive only):
ai/scripts/measure_quant_drop_per_ep.py— new sibling ofmeasure_quant_drop.py. CPU+CUDA via ORT; Arc / OpenVINO-CPU via the nativeopenvinoPython runtime (noonnxruntime-openvinobecause no cp314 wheel exists). Reuses the_load_sessionrename workaround from PR #165 + avalue_info-strip fix so dynamic-PTQ doesn't choke on the shipped MLP ONNX.docs/ai/quant-eps.md— new user doc; linked fromdocs/ai/index.md.docs/research/0006-tinyai-ptq-accuracy-targets.md— refreshed header, replaced "GPU-EP open question" with the measurement table, fixed pre-existing MD040/MD060 lints surfaced on the touched file.docs/ai/index.md— added the quant-eps row, rewrapped to 80 cols.CHANGELOG.mdUnreleased § Changed.- Invariants (rebase-relevant):
measure_quant_drop.py(the CI gate) is unchanged. The new script is purely additive. Any rebase that conflates the two scripts must keep the CI gate CPU-only — Arc int8 is broken, so a per-EP gate would red-light every PR.value_infostrip is required forvmaf_tiny_v1*dynamic PTQ. The shipped MLP ONNX duplicate weight tensors invalue_info, which makesquantize_dynamicraiseInferred shape and existing shape differ. The fix is in_save_inlined. Don't remove it during a refactor unless the underlying ONNX is regenerated.- CUDA-12 ABI shim. ORT-GPU 1.25 wheels link
libcublasLt.so.12even on CUDA-13 hosts. The reproduction recipe pins thenvidia-*-cu12wheels and prepends them toLD_LIBRARY_PATH. If a future ORT wheel drops the cu12 ABI we can cut the shim, but the script tolerates either since it doesn't import any CUDA symbol itself. - On upstream sync: zero interaction; entirely fork-local.
- Re-test on rebase:
SP=$VIRTUAL_ENV/lib/python3.14/site-packages/nvidia
export LD_LIBRARY_PATH="$SP/cublas/lib:$SP/cudnn/lib:$SP/cuda_nvrtc/lib:$SP/cuda_runtime/lib:$SP/cufft/lib:$SP/curand/lib:$SP/cusolver/lib:$SP/cusparse/lib:$SP/cuda_cupti/lib:$SP/nvtx/lib:$SP/nvjitlink/lib"
python ai/scripts/measure_quant_drop_per_ep.py \
--eps cpu cuda openvino \
--extra-fp32 vmaf_tiny_v1.onnx vmaf_tiny_v1_medium.onnx \
--out runs/quant-eps-$(date +%Y-%m-%d)
# Expected: CPU + CUDA PASS (drop ≤ 1.2e-4); OpenVINO Arc ERR
# (compile failure for Conv-int8) or NaN (MatMul-int8) until a
# newer intel_gpu plugin lands.
0065 — testdata/bench_all.sh correct backend-engagement flags¶
- No ADR. Bug fix; no behavioural surface change beyond "the bench actually engages the backends it claims to now."
- Upstream source: fork-local.
testdata/bench_all.shis a fork-only bench harness; not in Netflix/vmaf. - Touches (additive only):
testdata/bench_all.sh— switched per-row flag pattern from the disable-only singletons (--no_syclfor "CUDA", etc.) to the correct engagement form (--gpumask=0 --no_sycl --no_vulkanfor CUDA,--sycl_device=0 --no_cuda --no_vulkanfor SYCL,--vulkan_device=0 --no_cuda --no_syclfor Vulkan, and--no_cuda --no_sycl --no_vulkanfor CPU). Added a 4th column (Vulkan) to the comparator. Honours$VMAF_BINfor the binary path and$VMAF_ONEAPI_SETVARSfor the oneAPI install location.CHANGELOG.mdUnreleased § Fixed.- Invariants (rebase-relevant):
- Disable-only singletons don't engage a backend.
--no_syclalone leaves CUDA available but unrequested.--no_cudaalone leaves SYCL available but unrequested. The CLI inits CUDA only whenc.use_gpumaskis set; SYCL only whenc.sycl_device >= 0orc.use_gpumask; Vulkan only whenc.vulkan_device >= 0. Any change to those gates that drops one of the per-row flags will re-introduce the silent CPU fallback. Verify after a rebase by recording each live row's JSONframes[0].metricskey count. Treat a GPU count equal to CPU as a fallback warning, never as a fixed expected backend count — seelibvmaf/AGENTS.md§"Backend-engagement foot-guns". gpumasksemantics are inverted from intuition.gpumask=0enables CUDA dispatch;gpumask=1disables it. The per-row CUDA flag is--gpumask=0, not--gpumask=1. Don't "fix" it to--gpumask=1for symmetry with sycl_device/vulkan_device — that's the bug being fixed (parallel to PR #170).- On upstream sync: zero interaction;
testdata/bench_all.shis fork-only. - Re-test on rebase:
VMAF_BENCH_OUTDIR=testdata/bbb/results bash testdata/bench_all.sh
# Record actual live-backend counts and compare within this run:
jq '.frames[0].metrics | keys | length' testdata/bbb/results/t1_cpu.json
jq '.frames[0].metrics | keys | length' testdata/bbb/results/t1_cuda.json
jq '.frames[0].metrics | keys | length' testdata/bbb/results/t1_sycl.json
0063 — Tiny-AI LOSO eval harness for mlp_small¶
- No ADR. The methodology fits inside Research Digest 0022; ADR-0203 already covers the training-prep architecture.
- Research digest:
docs/research/0022-loso-mlp-small-results.md. - Upstream source: fork-local. Netflix/vmaf has no LOSO eval surface.
- Touches (additive only):
ai/scripts/eval_loso_mlp_small.py— new evaluation harness.docs/ai/loso-eval.md— usage doc.docs/research/0022-loso-mlp-small-results.md— methodology + results.CHANGELOG.mdUnreleased § Added.- Invariants (rebase-relevant):
_load_sessionworkaround for renamed-baseline ONNX. The shipped baselinesmodel/tiny/vmaf_tiny_v1*.onnxreference their pre-renameexternal_data.locationvalues. The workaround in_load_sessionrewrites the entries before handing the proto to ORT. Removing the workaround breaks the baseline phase. The proper fix (re-export with matching names) is tracked as a follow-up; until then this code path is load-bearing.runs/andmodel/tiny/training_runs/stay gitignored. The harness writes toruns/loso_eval/by default; do NOT promote any of those outputs into the tree. The 9 fold ONNX and the per-clip JSON cache regenerate from the corpus + trainer + libvmaf CLI.- On upstream sync: zero interaction. Pure fork-local evaluation harness.
- Re-test on rebase:
python ai/scripts/eval_loso_mlp_small.py
diff <(jq -r '.loso_aggregate.mean_plcc' runs/loso_eval/loso_mlp_small_eval.json) <(echo 0.9808)
# Expect: identical line on a populated cache + identical fold ONNX.
0064 — Section-A audit: 9 backlog rows + ADR cross-links¶
- No ADR. Process / docs PR; rows trace back to the individually-cited ADRs / research digests in their own References columns.
- Decision dossier:
.workingdir2/decisions/section-a-decisions-2026-04-28.md. - Source audit:
docs/backlog-audit-2026-04-28.md. - Upstream source: fork-local. Pure backlog hygiene PR; no Netflix code touched.
- Touches (additive only):
.workingdir2/BACKLOG.md— 9 new rows: T3-17, T3-18, T5-3e, T5-4, T7-35, T7-36, T7-37, T7-38; T6-1a row extended with the bisect-cache fixture sub-bullet.docs/research/0006-tinyai-ptq-accuracy-targets.md— drops the "defer until first user" framing on the GPU-EP quantisation open question per user direction; cross-links T5-3e.docs/research/0020-cambi-gpu-strategies.md— v2 follow-up section now cites T7-36 as the gate for opening the v2 row.docs/adr/0205-cambi-gpu-feasibility.md— Decision section's "follow-up integration PR" now cites T7-36.CHANGELOG.mdUnreleased § Changed.- Invariants (rebase-relevant): none. Pure backlog text. Rebase-conflict risk is limited to the same
BACKLOG.mdtable rows that any future row addition would touch; trivial to re-resolve. - On upstream sync: zero interaction.
- Re-test on rebase: none — docs-only.
0062 — ssimulacra2 CUDA + SYCL twins (ADR-0206)¶
- ADR: ADR-0206.
- Upstream source: fork-local. Netflix/vmaf has no SSIMULACRA 2 GPU implementation; this PR adds the CUDA + SYCL twins of the fork's ADR-0201 Vulkan kernel.
- Touches (additive + small wiring edits):
docs/adr/0206-ssimulacra2-cuda-sycl.mdand the index row indocs/adr/README.md.core/src/feature/cuda/ssimulacra2_cuda.{c,h}— new CUDA dispatch.core/src/feature/cuda/ssimulacra2/ssimulacra2_blur.cuandssimulacra2_mul.cu— new CUDA fatbins.core/src/feature/sycl/ssimulacra2_sycl.cpp— new SYCL extractor.core/src/feature/feature_extractor.c— two new extern declarations + two new entries infeature_extractor_list[].core/src/meson.build— addsssimulacra2_blur+ssimulacra2_multocuda_cu_sources, introduces (or extends, if PR #157 / ADR-0202 landed first) thecuda_cu_extra_flagsmap with assimulacra2_blurentry, threadsper_kernel_flagsinto the fatbin custom-target, and lists the two new C / CPP TUs.core/src/cuda/AGENTS.mdandcore/src/sycl/AGENTS.md— rebase invariant notes for the per-kernel--fmad=falseflag and the-fp-model=preciseSYCL build flag.docs/backends/cuda/overview.md,docs/backends/sycl/overview.md,docs/metrics/features.md— coverage matrix updates.CHANGELOG.mdUnreleased § Added.- Invariants (load-bearing on rebase):
- Per-kernel
--fmad=falseforssimulacra2_blur. The IIR'so = n2 * sum - d1 * prev1 - prev2must NOT fuse into FMAs — without the flag the recursive Gaussian's per-step rounding compounds across the 6-scale pyramid pastplaces=4. -fp-model=preciseon the SYCL feature build line. Removing it driftsssimulacra2_syclpastplaces=2through the IIR.- Hybrid host/GPU split mirrors Vulkan. Host runs YUV→RGB, XYB, downsample, and SSIM/EdgeDiff combine in double; GPU runs only mul + IIR blur. Any future PR that ports XYB or YUV→RGB onto the GPU MUST land alongside an updated ADR-0206 and re-validate
places=4on every Netflix CPU pair. - CUDA fex uses
.extract(synchronous), not.submit/.collect. Per-frame raw YUV is D2H-copied frompicture_cuda's device-sideVmafPicture.data[]into pinned host scratch viacuMemcpy2DAsync. Skipping the copy segfaults — direct host reads on aCUdeviceptrare the failure mode the prior agent's WIP hit. - On upstream sync: zero interaction with Netflix. The GPU coverage matrix for
ssimulacra2is wholly fork-local. - Re-test on rebase:
meson setup build_cuda libvmaf -Denable_cuda=true -Denable_sycl=false
ninja -C build_cuda
python3 scripts/ci/cross_backend_vif_diff.py \
--vmaf-binary ./build_cuda/tools/vmaf \
--feature ssimulacra2 --backend cuda --places 4 \
--reference python/test/resource/yuv/src01_hrc00_576x324.yuv \
--distorted python/test/resource/yuv/src01_hrc01_576x324.yuv \
--width 576 --height 324 --pixel-format 420 --bitdepth 8
# Expect: 0/48 mismatches, max_abs_diff ~1e-6.
0061 — cambi GPU feasibility spike (ADR-0205)¶
- ADR: ADR-0205.
- Research digest:
docs/research/0020-cambi-gpu-strategies.md. - Upstream source: fork-local. Netflix/vmaf has no Vulkan backend.
- Touches (additive only):
docs/adr/0205-cambi-gpu-feasibility.md,docs/research/0020-cambi-gpu-strategies.md,docs/adr/README.mdindex row.core/src/feature/vulkan/cambi_vulkan.c— new dormant scaffold (not yet invulkan_sources, not yet registered).core/src/feature/vulkan/shaders/cambi_{derivative,decimate,filter_mode}.comp— new reference GLSL shaders, not yet in the build'sshaderslist.core/src/feature/AGENTS.mdinvariants +CHANGELOG.mdbullet.- Invariants (rebase-relevant):
- Hybrid host/GPU port by decision. If Netflix upstream tightens the c-value formula or histogram update protocol, the host residual call site in the eventual
cambi_vulkan.c::cambi_vulkan_extractmust be updated alongsidecambi.c::calculate_c_values— the same code is reused. Do NOT translate the c-values phase to GPU during any upstream-port PR; that optimisation belongs to the v2 strategy-III PR (deferred). - Scaffolds dormant in the spike PR. The
cambi_vulkan.cextractor returns-ENOSYSfromcambi_vulkan_init_stubuntil the integration follow-up wires it in. Do NOT registervmaf_fex_cambi_vulkan_scaffoldinfeature_extractor.c's list. - Shaders not in the build's shader list. Adding them to
core/src/vulkan/meson.build'svulkan_shaderslist before the integration PR produces orphaned*_spv.hheaders. Leave them alone in this spike PR. - On upstream sync: zero interaction. cambi.c itself is upstream-mirrored — Netflix changes flow through
port-upstream-commit; only the integration PR's host residual call site needs paired attention. - Re-test on rebase:
```bash meson setup build -Denable_vulkan=enabled -Denable_cuda=false -Denable_sycl=false ninja -C build meson test -C build
0059 — Tiny-AI Netflix corpus training prep (ADR-0203)¶
- ADR: ADR-0203.
- Upstream source: fork-local. Netflix/vmaf has no equivalent training surface.
- Touches:
ai/data/— Netflix loader, libvmaf-CLI feature extractor, distillation scoring.ai/train/— PyTorch dataset, eval harness, Lightning-style training entry point.ai/scripts/run_training.sh— convenience wrapper.ai/tests/— five new pytest modules (test_netflix_loader.py,test_dataset.py,test_eval.py,test_train_smoke.py, plusconftest.py).docs/ai/training.md— new "C1 (Netflix corpus)" section; existing sections untouched.ai/AGENTS.md— invariants section added.- Invariants (load-bearing):
- Filename ladder regex is fork-specific.
<source>_<quality>_<height>_<bitrate>.yuv(dis) +<source>_<fps>fps.yuv(ref). Upstream may publish a different naming convention later; do NOT merge them — keep this loader scoped to the Netflix corpus, add a sibling loader for any upstream alternative. - Per-clip cache schema is consumed by both dataset and any downstream tooling. Schema is
{features:{feature_names, per_frame, n_frames}, scores:{per_frame, pooled}}. Any change must invalidate$VMAF_TINY_AI_CACHE(delete or version-tag the directory). - Smoke command stays runnable without a built
vmafbinary. The_make_zero_payloadhelper inai.train.datasetinjects a fake payload for--epochs 0so CI gates don't drag a libvmaf build into the Python test surface. - YUV size probe never silently guesses.
probe_yuv_dimseither matches the 1920x1080 default, returns ffprobe's answer, or raises. Tests passassume_dims=(16, 16)explicitly for synthetic fixtures. - On upstream sync: no interaction with upstream. The
ai/subtree is wholly fork-local. - Re-test on rebase:
python -m pytest ai/tests/test_netflix_loader.py \
ai/tests/test_dataset.py ai/tests/test_eval.py \
ai/tests/test_train_smoke.py -v
python ai/train/train.py --epochs 0 --data-root /tmp/mock_corpus \
--assume-dims 16x16 --val-source BetaSrc --out-dir /tmp/out
0073 — Tiny-AI QAT trainer + first per-model QAT pass (T5-4)¶
- ADR: ADR-0207 (design), ADR-0208 (per-model impl).
- Touches:
ai/train/qat.py(new),ai/scripts/qat_train.py(rewrite fromNotImplementedErrorscaffold),ai/configs/learned_filter_v1_qat.yaml(new),ai/tests/test_qat_smoke.py(new),docs/ai/quantization.md(QAT tier added). All paths are wholly fork-local; no upstream Netflix/vmaf interaction. - Invariants:
- Two-step pipeline (PyTorch QAT → fp32 ONNX → ORT static-quantize) is load-bearing. Both the legacy ONNX exporter (
quantized::conv2d) and the new TorchDynamo exporter (Conv2dPackedParamsBase.__obj_flatten__) refuse to consumeconvert_fxoutput on PyTorch 2.11. The bridge (state-dict diff to a fresh fp32 module + ORT static-quantize) is the only path that yields a QDQ ONNX. Do NOT collapse to a single-stepconvert_fx → torch.onnx.exportuntil both PyTorch issues are fixed; re-check both exporters on each PyTorch upgrade. - State-dict transfer matches by submodule name + shape.
_copy_qat_weights_into_fp32walksfp32_statekeys, finds the same key in the FX-prepared module, copies the tensor. Tiny-AI models today have stable submodule names (entry,body.*,exit); a model architecture that uses top-levelnn.Sequentialwould break this becauseprepare_qat_fxrenames Sequential children to numeric indices. TheRuntimeError("0 tensors copied")guard catches the silent failure mode. - FX preparation runs on CPU. PyTorch 2.11's FX symbolic tracer is flaky on CUDA buffers; the trainer migrates the model to CPU before
prepare_qat_fxand back to the accelerator for the fine-tune phase. The smoke test deliberately exercises the CPU path so this stays covered. torch.ao.quantizationdeprecation will hard-fail in PyTorch 2.10. Migration target istorchao.quantization.pt2e(prepare_pt2e/convert_pt2e); the two-step pipeline is mostly pt2e-compatible — only the FX-prep call changes.- On upstream sync: no interaction with upstream. The
ai/subtree is fully fork-local. - Re-test on rebase:
python -m pytest ai/tests/test_qat_smoke.py -v
python ai/scripts/qat_train.py \
--config ai/configs/learned_filter_v1_qat.yaml \
--output /tmp/qat_smoke.int8.onnx --smoke
0074 — GPU-parity matrix CI gate (T6-8 / ADR-0214)¶
- Touched surfaces (fork-local):
scripts/ci/cross_backend_parity_gate.py(new),.github/workflows/tests-and-quality-gates.yml(newvulkan-parity-matrix-gatejob),docs/development/cross-backend-gate.md(new),docs/backends/index.md(cross-backend section),libvmaf/AGENTS.md(rebase-sensitive invariant note). - Why this matters on rebase: the CI lane and the matrix-gate script are entirely fork-local. Upstream Netflix/vmaf has no comparable gate; conflicts on rebase are restricted to the CI workflow file when upstream rearranges its own jobs. The gate's Python script lives outside
core/src/so the upstream-sync path doesn't see it. - Invariants the gate enforces:
- Per-feature absolute tolerance is declared in one place (
FEATURE_TOLERANCEinscripts/ci/cross_backend_parity_gate.py). Tightening a tolerance requires a measurement-driven follow-up ADR; loosening requires a justification ADR (CLAUDE.md §12 r1). - The legacy single-feature gate
scripts/ci/cross_backend_vif_diff.pystays for one release cycle. Sister PRs in this session add to it; the T6-8b cleanup PR deletes it once the matrix gate has soaked. - CUDA / SYCL / hardware-Vulkan are advisory until a self-hosted runner is registered. The script supports them via
--backends; flipping the CI lane to required is a follow-up wiring change, not a code change. - On upstream sync: no interaction with upstream
tests-and-quality-gates.yml(the gate job is fork-added); rebase conflicts limited to insertion-order in the workflow file. - Re-test on rebase:
cd libvmaf && meson setup build \
-Denable_cuda=false -Denable_sycl=false \
-Denable_vulkan=enabled -Denable_float=true \
--buildtype=release && ninja -C build
cd ..
python3 scripts/ci/cross_backend_parity_gate.py \
--vmaf-binary core/build/tools/vmaf \
--reference testdata/ref_576x324_48f.yuv \
--distorted testdata/dis_576x324_48f.yuv \
--width 576 --height 324 --backends cpu vulkan \
--json-out /tmp/parity.json --md-out /tmp/parity.md
0220 — SYCL feature kernels are unconditionally fp64-free (T7-17)¶
- Touches:
core/src/sycl/common.cpp(init log line),core/src/sycl/AGENTS.md(new invariant row), all SYCL feature kernels undercore/src/feature/sycl/(no diff today, but the contract pins their shape going forward). - Invariant: every SYCL feature-kernel lambda captures and operates on
float/ integer types only. Nodoubleoperand inside aparallel_forbody, nosycl::reduction<double>, nosycl::plus<double>. A single fp64 instruction in the TU's SPIR-V module causes the Level Zero runtime to reject the entire module on Intel Arc A-series and other fp64-less devices, even when the offending kernel is never submitted. Host-sidedouble(inextract/flushpost-processing, score aggregation, log10 normalisation) remains fine. Concrete patterns in tree: ADM gain limiting via int64 Q31 (gain_limit_to_q31+launch_decouple_csf<false>ininteger_adm_sycl.cpp); VIF gain limiting via fp32sycl::fmin; CIEDE / SSIM accumulators viasycl::reduction<int64_t>/sycl::plus<int64_t>. - On upstream sync: Netflix/vmaf has no SYCL backend upstream; conflicts cannot enter via
git merge. The risk is a fork-local cherry-pick (e.g. a SYCL twin of a new CUDA kernel) bringing adoubleinto a kernel lambda. Audit the lambda capture list and anysycl::reduce*calls against this invariant before merging. - Re-test on rebase:
# Build SYCL backend
meson setup build-sycl libvmaf -Denable_sycl=true CC=icx CXX=icpx
ninja -C build-sycl
# On an fp64-less device (e.g. Intel Arc A380), confirm the
# init log line is INFO-level and reads "device lacks native
# fp64 — kernels already use fp32 + int64 paths, no emulation
# overhead". The SYCL kernels must launch successfully (no
# SPIR-V module rejection from the Level Zero runtime).
build-sycl/tools/vmaf --reference testdata/ref_576x324_48f.yuv \
--distorted testdata/dis_576x324_48f.yuv \
--width 576 --height 324 --backend sycl \
--feature integer_vif --feature integer_adm \
--output /tmp/sycl-fp64less.json --json
0091 — T6-9 model registry schema + --tiny-model-verify (ADR-0211)¶
- No rebase impact: 100% fork-local surface. The registry (
model/tiny/registry.json), its JSON Schema (model/tiny/registry.schema.json), the--tiny-model-verifyCLI flag, and thevmaf_dnn_verify_signature()C entry point are entirely fork-local — none of these paths exist in upstream Netflix/vmaf. Listed here for completeness so a future/sync-upstreamrun sees the surface area was acknowledged. - Touches (additive only):
model/tiny/registry.json,model/tiny/registry.schema.json,ai/scripts/validate_model_registry.py,core/src/dnn/model_loader.{c,h}(addedvmaf_dnn_verify_signature()),core/include/libvmaf/dnn.h(public declaration),core/tools/cli_parse.{c,h}(ARG_TINY_MODEL_VERIFY+tiny_model_verifyfield),core/tools/vmaf.c(call site),core/test/dnn/test_tiny_model_verify.c,python/test/model_registry_schema_test.py,docs/ai/model-registry.md,docs/ai/inference.md,docs/ai/security.md,docs/adr/0209-...md,docs/adr/README.md(index row),CHANGELOG.md,core/src/dnn/AGENTS.md. - Invariants (rebase-relevant):
- Schema is the contract. New registry fields land in
registry.schema.jsonfirst, then inregistry.json, then in any consumers (the C-side parser, the Python validator, the MCP). Reverse order causes mismatch. schema_versionis bounded. The schema accepts only{0, 1}; bump the enum and the loader's check together when adding2.- Banned-function rule applies. The
cosigninvocation usesposix_spawnp(3p)with an explicit argv array. Do not replace withsystem(3)/popen(3)— both shell-parse the command and would re-introduce injection risk. - Bundle-file absence is fail-closed. When
sigstore_bundlepoints at a not-yet-existing file (pre-release state),vmaf_dnn_verify_signature()returns-ENOENT. The CLI surfaces this as a load failure; do not "soften" to a warning without an explicit ADR. - Re-test on rebase:
python3 ai/scripts/validate_model_registry.py
python3 -m pytest python/test/model_registry_schema_test.py -v
meson test -C build-cpu --suite=dnn
0074 — HIP (AMD ROCm) backend scaffold (T7-10)¶
- ADR: ADR-0212.
- Upstream source: fork-local. HIP backend is fork-only; Netflix/vmaf has no
libvmaf_hip.hand noenable_hipmeson option. - Touches:
core/include/libvmaf/libvmaf_hip.h(new).core/include/core/meson.build— adds theis_hip_enabledinstall gate, mirroringis_cuda_enabled/is_sycl_enabledboolean idioms.core/meson_options.txt— newenable_hipboolean option (default false).core/src/meson.build— newis_hip_enabledflag, conditionalsubdir('hip'),hip_sources+hip_depsthreaded throughlibvmaf_feature_static_lib(alongside the existing CUDA / SYCL / Vulkan aggregations) and the top-levellibrary('vmaf', ...)dependencieslist.core/src/hip/(new directory:common.{c,h},picture_hip.{c,h},dispatch_strategy.{c,h},meson.build).core/src/feature/hip/(new directory:adm_hip.c,vif_hip.c,motion_hip.c).core/test/test_hip_smoke.c(new).core/test/meson.build— registers the smoke test underif get_option('enable_hip') == true..github/workflows/libvmaf-build-matrix.yml— addsBuild — Ubuntu HIP (T7-10 scaffold)row.docs/backends/hip/overview.md(new),docs/backends/index.md(planned → scaffold row),docs/research/0033-hip-applicability.md(new),docs/adr/0212-hip-backend-scaffold.md(new),docs/adr/README.md(new index row).libvmaf/AGENTS.md— new "HIP backend scaffold contract" rebase-sensitive invariant entry.CHANGELOG.md— Unreleased § Added.- Invariants (rebase-relevant):
enable_hipis abooleanoption, not afeature. Mirrorsenable_cuda/enable_sycl; do not "harmonise" withenable_vulkan'sfeature/disabledform without an ADR amendment per ADR-0212 § "Decision".- Public C-API entry points return
-ENOSYSfor the scaffold. The smoke test core/test/test_hip_smoke.c pins this. A rebase that "succeeds" by accidentally enabling a code path (e.g. a refactor that early-returns 0 fromvmaf_hip_state_init) breaks the smoke and the runtime PR's contract baseline. hip_sourcesis added tolibvmaf_feature_static_lib, NOT directly to the top-levellibrary('vmaf', ...). The static lib is extracted into libvmaf viaobjects: [..., libvmaf_feature_static_lib.extract_all_objects(recursive: true), ...]at the bottom ofcore/src/meson.build. Addinghip_sourcesto the top library() too would double-link.hip_depsIS added to the top library()dependencies:list. The runtime PR will populatehip_depswith the realdependency('hip-lang')linkage; threading it through the top library() ensures consumers see the transitive dependency.- Header purity:
libvmaf_hip.hdoes not include<hip/hip_runtime.h>. HIP runtime types cross the public ABI asuintptr_t(matches the CUDA / Vulkan precedent; ADR-0212). Don't add<hip/...>includes to the public header during a rebase / runtime-PR bring-up. - No FFmpeg patch: the fork's
ffmpeg-patches/series does not currently consume the HIP API surface. CLAUDE §12 r14 only requires patch updates when an existing patch consumes the surface; the runtime PR (T7-10b) will add thehip_devicefilter option and the corresponding patch. - On upstream sync: zero interaction; HIP backend is fork-only.
- Re-test on rebase:
cd libvmaf
meson setup build-hip -Denable_cuda=false -Denable_sycl=false \
-Denable_hip=true
ninja -C build-hip
meson test -C build-hip test_hip_smoke
# Expect: 9/9 pass.
# Default no-HIP build still works:
meson setup build-cpu -Denable_cuda=false -Denable_sycl=false
ninja -C build-cpu
meson test -C build-cpu --suite=fast
0074 — SSIMULACRA 2 SVE2 SIMD parity (T7-38)¶
- ADR: ADR-0213.
- Touches:
core/src/feature/arm64/ssimulacra2_sve2.{c,h}(new),core/src/feature/ssimulacra2.c(dispatch table override ininit_simd_dispatch),core/src/arm/cpu.{c,h}(HWCAP2_SVE2 probe + newVMAF_ARM_CPU_FLAG_SVE2enum value),core/src/meson.build(cc.compiles probe + optionalarm64_ssimulacra2_sve2static library),core/test/test_ssimulacra2_simd.c(SVE2 picker overrides on the arm64 path + dispatch diagnostic),build-aux/aarch64-linux-gnu-sve2.ini(new cross-file pinningqemu-aarch64-static -cpu max). All paths are wholly fork-local; no upstream Netflix/vmaf code is modified. - Invariants:
- Fixed 4-lane SVE2 predicate. Every kernel uses
svwhilelt_b32(0, 4)so SIMD arithmetic order is identical to the NEON sibling regardless of the runtime vector length. This keeps the ADR-0138 / ADR-0139 / ADR-0140 byte-exact contract intact. Do NOT widen the predicate tosvptrue_b32()without a separate ADR + snapshot regen — variable-length lane reductions perturb the per-step rounding order. - NEON stays the fallback. SVE2 is purely additive; the dispatch table assigns NEON first and only overrides on
VMAF_ARM_CPU_FLAG_SVE2. A toolchain that fails thecc.compiles(... -march=armv9-a+sve2)probe leavesHAVE_SVE2unset and the legacy NEON-only build is unchanged. -ffp-contract=offmirrors the NEON sibling. Without it GCC fuses the per-lane scalar tail'sa*b+cpatterns intofmla, drifting against the SIMD path by ~1 ulp. Thearm64_ssimulacra2_sve2static library carries the flag like its NEON counterpart.- On upstream sync: no interaction with upstream —
arm64/feature TUs and thearm/cpu.{c,h}flag enum are fork-local. An upstream sync that rewritesinit_simd_dispatchincore/src/feature/ssimulacra2.cwould also need the SVE2 cases preserved. - Re-test on rebase:
meson setup build-arm64-sve2 libvmaf \
--cross-file=build-aux/aarch64-linux-gnu-sve2.ini -Denable_asm=true
ninja -C build-arm64-sve2 test/test_ssimulacra2_simd
meson test -C build-arm64-sve2 test_ssimulacra2_simd
# stderr should report `ssimulacra2 simd dispatch: NEON=1 SVE2=1`
# and 11/11 tests should pass.
0075 — enable_lcs MS-SSIM extras on CUDA + Vulkan (T7-35 / ADR-0243)¶
- Touched surfaces (fork-local):
core/src/feature/cuda/integer_ms_ssim_cuda.c(addedenable_lcstoMsSsimStateCuda+options[]+ 15 host-sidevmaf_feature_collector_appendcalls gated on the bool),core/src/feature/vulkan/ms_ssim_vulkan.c(rewroteenable_lcshelp text + addedemit_lcs_metricshelper + gated 15vmaf_feature_collector_appendcalls),scripts/ci/cross_backend_vif_diff.py scripts/ci/cross_backend_parity_gate.py(newfloat_ms_ssim_lcspseudo-feature +FEATURE_ALIASESmapplaces=4tolerance row).- Why this matters on rebase: the GPU MS-SSIM extractors are fork-local (Netflix upstream has no Vulkan or CUDA MS-SSIM kernel today). The
enable_lcssemantic and the metric names (float_ms_ssim_{l,c,s}_scale{0..4}) must match the upstream CPU reference atcore/src/feature/float_ms_ssim.c:189-221. If upstream ever renames or reorders those metrics, mirror the change on the GPU side in the same merge — public-API contract. - Invariants the contract enforces:
- Default-path output (
enable_lcs=false) stays bit-identical to the pre-T7-35 binary: only the host-side appends are gated; no kernel / shader / device-buffer changes. - Metric ordering is metric-wise (all
l_scale*first, thenc_*, thens_*) — matches the CPU emission order. places=4cross-backend tolerance per ADR-0190; enforced by the newfloat_ms_ssim_lcscell in the parity matrix gate (ADR-0214).- On upstream sync: zero interaction; the GPU twins do not exist upstream. The CPU
float_ms_ssim.cis shared with upstream butenable_lcsis upstream-stable since v3.0.0. - Re-test on rebase:
cd libvmaf && meson setup build-vulkan \
-Denable_cuda=false -Denable_sycl=false \
-Denable_vulkan=enabled -Denable_float=true \
--buildtype=release && ninja -C build-vulkan
cd ..
python3 scripts/ci/cross_backend_vif_diff.py \
--vmaf-binary core/build-vulkan/tools/vmaf \
--reference testdata/ref_576x324_48f.yuv \
--distorted testdata/dis_576x324_48f.yuv \
--width 576 --height 324 \
--feature float_ms_ssim_lcs --backend vulkan --places 4
0075 — 32-bit ADM/cpu fallbacks port (T-NEW-3)¶
- Touched surfaces (upstream-mirror):
core/src/feature/x86/adm_avx2.c,core/src/feature/x86/adm_avx512.c,core/src/x86/cpu.c. Cherry-picks of upstream8a289703(Christopher Degawa, "adm: add fallback for extract_epi64 for 32-bit") and1b6c3886("x86/cpu: remove limit of avx+ on 32-bit"). - Why this matters on rebase: trivially conflict-free with any future upstream
extract_epi64work because we land upstream's exactextract_epi64macro/inline-fn pair. The conflict surface is the fork's clang-format-100col layout inadm_avx2.c/adm_avx512.cand the_Alignas(64)LTO-correctness slot inadm_avx512.c(docs/development/known-upstream-bugs.md); both are preserved verbatim. - Invariants the port preserves:
_Alignas(64) int64_t angle_flag[16]inadm_decouple_s123_avx512stays — without it, LTO can promote the unaligned load tovmovdqa64and fault under--buildtype=release -Db_lto=true.- The
extract_epi64symbol must remain resolved on both__x86_64__(macro to_mm256_extract_epi64) and 32-bit (fallback inline). If a future upstream change inlines the helper differently, keep the conditional definition. - On upstream sync: if Netflix ships further 32-bit fallbacks (motion / psnr — not in this port), expect a parallel
extract_epi64-style helper at the top of each affected SIMD file. The fork should mirror those verbatim into the same files. - Re-test on rebase:
meson setup build-i686 libvmaf \
--cross-file=build-aux/i686-linux-gnu.ini \
-Denable_asm=false
ninja -C build-i686
meson setup build-cpu libvmaf -Denable_avx512=true
ninja -C build-cpu
meson test -C build-cpu
0076 — codec-aware FR regressor surface (T7-CODEC-AWARE / ADR-0235)¶
- Touches:
ai/src/vmaf_train/codec.py(new),ai/src/vmaf_train/models/fr_regressor.py(extended),ai/scripts/bvi_dvc_to_full_features.py,ai/scripts/extract_full_features.py. No upstream-shared paths. - Invariant:
CODEC_VOCABinai/src/vmaf_train/codec.pyis closed and order-stable — the index of each codec is the one-hot column index baked into trained ONNX. Adding a codec appends to the tuple and bumpsCODEC_VOCAB_VERSION; reordering silently invalidates every shippedfr_regressor_v2_*.onnx.FRRegressor(num_codecs=0)must remain the v1 single-input contract — flipping the default would break every existingmodel/tiny/fr_regressor_v1.onnxconsumer. - Re-test:
pytest ai/tests/test_codec_aware_fr.py -v(8 sub-tests covering vocabulary contract + alias table + back-compat). Pure fork-local addition; no upstream rebase impact for the next/sync-upstream.
0075 — feature/speed extractors (T-NEW-1, upstream port d3647c73)¶
- Touches:
core/src/feature/speed.c(new),core/src/feature/picture_copy.{c,h}(signature change — addedint channelparameter),core/src/feature/float_*.ccall sites updated to passchannel=0,core/src/feature/feature_extractor.cregistry block,core/src/feature/alias.c,core/src/meson.build,core/src/feature/vif_tools.{c,h}(helper-function port from upstream4ad6e0ea). - Upstream source: verbatim cherry-pick of Netflix/vmaf
d3647c73("feature/speed: port speed_chroma and speed_temporal extractors") with its dependency4ad6e0ea("feature/vif: port helper functions"). Both are pre-existing on Netflix master and enter the fork as part of the T7-4 audit catch-up. - Invariant:
picture_copy()now takes achannelargument — every fork-local extractor that calls it (CUDAinteger_ms_ssim, Vulkanssim/ms_ssim) passeschannel=0. If upstream later evolves the signature again (e.g. adds bit-depth or stride validation), update those fork-local call sites in lockstep. Speed extractors only register whenVMAF_FLOAT_FEATURES=1(build with-Denable_float=true). - On upstream sync: future Netflix commits in
core/src/feature/speed.capply cleanly because the file is now a verbatim mirror; conflict potential is limited to the registry block infeature_extractor.c(interleave with the fork's Vulkan / SYCL / CUDA blocks) and to any furtherpicture_copysignature evolution. - Re-test on rebase:
```bash meson setup build-cpu libvmaf -Denable_cuda=false \ -Denable_sycl=false -Denable_float=true ninja -C build-cpu meson test -C build-cpu test_speed meson test -C build-cpu # full meson suite make test-netflix-golden # 3 CPU canonical pairs
0221 — CHANGELOG + ADR-index fragment-file pattern (T7-39 / ADR-0221)¶
- What changed: the fork stopped editing
CHANGELOG.mdanddocs/adr/README.mddirectly. Both files are now rendered from fragment trees: changelog.d/<section>/<topic>.md(Keep-a-Changelog sections), plus the migration archivechangelog.d/_pre_fragment_legacy.md.docs/adr/_index_fragments/<NNNN-slug>.md, plusdocs/adr/_index_fragments/_order.txt(frozen commit-merge order manifest) anddocs/adr/_index_fragments/_header.md(table prelude). Two scripts render the consolidated outputs:scripts/release/concat-changelog-fragments.sh --check|--writescripts/docs/concat-adr-index.sh --check|--write- On upstream sync: zero interaction —
CHANGELOG.mdis a fork-local Markdown surface (Netflix upstream doesn't ship a Keep-a-Changelog file in this format), anddocs/adr/is entirely fork-local. A/sync-upstreamrun will not touch the fragment trees. - Re-test on rebase:
bash scripts/release/concat-changelog-fragments.sh --check
bash scripts/docs/concat-adr-index.sh --check
# both must exit 0; otherwise run --write and re-stage.
0077 — DISTS extractor proposal (T7-DISTS / ADR-0236)¶
- What landed: ADR-0236 (Proposed) + Research-0043 design digest ADR README index row + CHANGELOG entry.
- Rebase impact: pure fork-local proposal-stage docs; no code, no Netflix-mirror file touched, no ffmpeg-patches change, no public C-API surface change.
- Reproducer (when implementation lands as T7-DISTS):
```sh vmaf --feature dists_sq=model_path=model/tiny/dists_sq.onnx \ --reference ref.yuv --distorted dist.yuv \ --width 1920 --height 1080 --pix_fmt yuv420p
0076 — GPU-gen ULP calibration head (proposal-stage, T7-GPU-ULP-CAL / ADR-0234)¶
- What landed: ADR-0234 (Proposed), Research-0041, data-collection scaffold at
ai/scripts/collect_gpu_calibration_data.py, forward-pointer indocs/usage/cli.mdfor the future--gpu-calibratedflag. - Rebase impact: pure fork-local (proposal docs + Python script); no upstream Netflix/vmaf code touched, no public C-API changes, no ffmpeg-patches changes.
- Reproducer:
```sh python3 ai/scripts/collect_gpu_calibration_data.py --smoke
0095 — Per-backend GPU kernel scaffolding templates (CUDA + Vulkan, ADR-0246)¶
- ADR: ADR-0246.
- Touches:
core/src/cuda/kernel_template.h(new, header-only).core/src/vulkan/kernel_template.h(new, header-only).core/src/cuda/AGENTS.md(new invariant row + dir listing).core/src/vulkan/AGENTS.md(new file).docs/backends/kernel-scaffolding.md(new).docs/adr/0246-gpu-kernel-template.md(new).CHANGELOG.md,docs/adr/README.md. All paths are wholly fork-local. Upstream Netflix/vmaf has no Vulkan backend at all today and the CUDA backend uses different per-kernel scaffolding shapes; nothing here can collide on a pure upstream sync.- Invariants:
- Templates are unused at PR-merge time.
kernel_template.hin bothcore/src/cuda/andcore/src/vulkan/lands with zero call-sites. Each future kernel migration is its own gated PR (places=4cross-backend-diff per ADR-0214). Do not bulk-port existing kernels onto the templates in a single sync — that would short-circuit the per-kernel gate. - Per-backend, not cross-backend. Resist the urge to merge the two templates into a unified
gpu/kernel_template.h. CUDA async-stream + event vs Vulkan command-buffer + fence + descriptor-pool share no concrete shape; a unified API would be lowest-common-denominator. - Helper functions, not macros. The header bodies are
static inlinefunctions for cuda-gdb / Nsight / RenderDoc step-debugging. TheCHECK_CUDA_GOTO/CHECK_CUDA_RETURNmacros incuda_helper.cuhstay where they pay off (textualgoto label), and the templates use them internally. - On upstream sync: no interaction with upstream paths. An upstream sync that touches
core/src/cuda/common.horpicture_cuda.hmay shift the helper signatures the template consumes (vmaf_cuda_buffer_alloc,vmaf_cuda_picture_get_stream, …); update the template if so. - Re-test on rebase:
```bash # CUDA build (configure inside libvmaf/ — see CLAUDE.md §2 note). meson setup core/build-cuda libvmaf \ -Denable_cuda=true -Denable_nvcc=true \ -Denable_vulkan=disabled -Denable_sycl=false ninja -C core/build-cuda meson test -C core/build-cuda
# Vulkan build. meson setup core/build-vulkan libvmaf \ -Denable_vulkan=enabled -Denable_cuda=false -Denable_sycl=false ninja -C core/build-vulkan meson test -C core/build-vulkan
0222 — vmaf-perShot per-shot CRF predictor sidecar (T6-3b)¶
- Touches:
core/tools/meson.build(new executable + test wiring),core/tools/vmaf_per_shot.c(new file — fork-local, no upstream sibling),core/tools/test/meson.build(test row),core/tools/test/test_vmaf_per_shot.sh(new smoke test),core/tools/AGENTS.md(sidecar invariants),docs/usage/cli.md(cross-link),docs/usage/vmaf-perShot.md(new user doc),docs/ai/roadmap.md(T6-3b row update). - Invariant: the sidecar must stay standalone — it does not link the libvmaf metric path. Any upstream patch that tries to fold per-shot CRF prediction into
vmaf_score_*would collapse the encoder-hint vs. quality-score separation recorded in roadmap §2.4 and ADR-0222 §Decision. The CSV / JSON column set (shot_id,start_frame,end_frame,frames,mean_complexity,mean_motion,predicted_crf) is the public schema; downstream encoders consume it directly. - Conflict expectation on
/sync-upstream: low. Upstream Netflix has no per-shot CRF predictor in tree, so there is no natural collision point —tools/meson.buildis the only mutually-edited file and the newexecutable('vmaf-perShot', …)block is appended aftervmaf_bench_deps, well clear of upstream's likely additions. - Reproducer:
```bash meson setup build libvmaf -Denable_cuda=false -Denable_sycl=false \ -Denable_vulkan=disabled ninja -C build meson test -C build test_vmaf_per_shot --print-errorlogs ./build/tools/vmaf-perShot \ --reference testdata/ref_576x324_48f.yuv \ --width 576 --height 324 --pixel_format 420 --bitdepth 8 \ --output /tmp/plan.csv cat /tmp/plan.csv
0075 — vmaf-roi sidecar binary (T6-2b / ADR-0247)¶
- Touches:
core/tools/meson.build— adds thevmaf_roiexecutable target (after the existingvmaftarget, beforevmaf_bench). Append-only; no upstream-shared lines moved or removed.core/test/meson.build— adds thetest_vmaf_roiexecutable +test()registration. Append-only.core/tools/vmaf_roi.c— wholly new, fork-local.core/tools/vmaf_roi_core.h— wholly new, fork-local.core/test/test_vmaf_roi.c— wholly new, fork-local.- Invariant: the
vmaf-roisidecar emits two byte-exact formats that downstream encoder drivers (x265--qpfile, SVT-AV1--roi-map-file) will hard-depend on: - x265 ASCII grid — two
#-prefixed header lines (# vmaf-roi qpfile (x265, --qpfile-style)and# frame=N ctu=S cols=C rows=R strength=F.FFF), space-separated signed integers, one row per CTU row,\nterminator. - SVT-AV1 raw binary — exactly
cols * rowsbytes ofint8_t, row-major, no header. - QP-offset clamp —
+-12(VMAF_ROI_CORE_QP_OFFSET_MAX). - Reduction — per-CTU mean (not max). Switching to max or a percentile changes every downstream encoder result and requires its own ADR.
- Pure helpers in
vmaf_roi_core.h— the per-CTU mean reducer and saliency-to-QP mapper arestatic inlinein a header sotest_vmaf_roicompiles them without dragging the libvmaf link surface in. Moving them into a.cTU breaks the test wiring. - On upstream sync: no interaction with upstream —
tools/is a fork-local surface from upstream's perspective (upstream shipsvmaf.conly). An upstream sync that rewritescore/tools/meson.buildshould preserve thevmaf_roiexecutable block. - Re-test on rebase:
```bash meson setup build-cpu libvmaf \ -Denable_cuda=false -Denable_sycl=false -Denable_tools=true ninja -C build-cpu tools/vmaf_roi test/test_vmaf_roi meson test -C build-cpu test_vmaf_roi ./build-cpu/tools/vmaf_roi \ --reference testdata/ref_576x324_48f.yuv \ --width 576 --height 324 --frame 0 --output - \ --encoder x265 --ctu-size 64 --strength 6.0 | head -3 # First two lines are the # comment header; row 1 of the grid # should be "4 2 1 -1 -1 -1 1 2 4" (placeholder radial map).
0219 — motion3 GPU coverage on Vulkan + CUDA + SYCL (T3-15(c) / ADR-0219)¶
- What changed: The
motionGPU twins (core/src/feature/vulkan/motion_vulkan.c,core/src/feature/cuda/integer_motion_cuda.c,core/src/feature/sycl/integer_motion_sycl.cpp) now emitVMAF_integer_feature_motion3_scorein 3-frame window mode (default). Cross-backend gates extended (scripts/ci/cross_backend_*.pyFEATURE_METRICS["motion"]). - Invariants:
motion3 = host-side scalar post-process of motion2. No device-side state changes; motion3 is computed on the host inextract()/collect()/flush()after the existing SAD reduction. The post-processing function (motion3_postprocess_*) mirrors CPUinteger_motion.clines 510-560 byte-for-byte:clip(motion_blend(motion2 * fps_weight, blend_factor, blend_offset), max_val)with optional moving-average against the unaveraged prior blended value.motion_five_frame_window=truereturns-ENOTSUPatinit()on all three GPU backends. The 5-deep blur ring + second SAD-pair dispatch remain deferred. Do NOT silently fall back to the 3-frame path when the user enables the flag — fail loud per CERT C / CLAUDE.md §12 r4.- CPU motion3 algorithm is the source of truth. Any port of an upstream Netflix change to
integer_motion.cthat touchesmotion_blend(...), themotion_max_valclip, or the moving-average rule MUST be mirrored inmotion3_postprocess_*across all three GPU files in the same PR. The cross-backend gate atplaces=4will catch drift, but only after a full GPU run. - On upstream sync: Pure fork-local additions to GPU TUs. Upstream Netflix has no GPU motion extractor. The
motion_blend_tools.hheader is upstream-mirrored — if a sync rewrites themotion_blend()formula, regenerate the GPU snapshot and re-run the cross-backend gate. - Re-test on rebase:
```bash # CPU sanity (motion3 emission unchanged) ./core/build/tools/vmaf \ --reference python/test/resource/yuv/src01_hrc00_576x324.yuv \ --distorted python/test/resource/yuv/src01_hrc01_576x324.yuv \ --width 576 --height 324 --pixel_format 420 --bitdepth 8 \ --feature motion --output /tmp/motion.json --json python -c "import json; d=json.load(open('/tmp/motion.json')); \ print('motion3 frames:', sum(1 for f in d['frames'] \ if 'integer_motion3' in f.get('metrics', {})))" # Expect 49 (one motion3 per frame).
# Cross-backend gate (Vulkan/lavapipe lane works on every host): python scripts/ci/cross_backend_vif_diff.py \ --feature motion --backend vulkan \ --ref python/test/resource/yuv/src01_hrc00_576x324.yuv \ --dis python/test/resource/yuv/src01_hrc01_576x324.yuv \ --width 576 --height 324 --bitdepth 8 \ --vmaf-bin core/build/tools/vmaf # Expect: integer_motion / integer_motion2 / integer_motion3 all OK at places=4.
0216 — vmaf_tiny_v2 (Phase-3-validated tiny VMAF MLP)¶
- Touches:
model/tiny/registry.json,model/tiny/vmaf_tiny_v2.{onnx,json},ai/scripts/{train,export,validate}_vmaf_tiny_v2.py,ai/AGENTS.md,core/test/dnn/{test_vmaf_tiny_v2.py,meson.build},docs/ai/{models/vmaf_tiny_v2.md,inference.md,roadmap.md},docs/adr/{0244-vmaf-tiny-v2.md,README.md},CHANGELOG.md. All paths are wholly fork-local; no upstream Netflix/vmaf code is modified. - Invariants:
- Bundled scaler stats are part of the trust root. The shipped ONNX bakes
(input - mean) / stdas ConstantSub+Divnodes that run before the MLP. Re-exporting must go throughai/scripts/export_vmaf_tiny_v2.py, which pullsmean/stdfrom the trainer checkpoint and writes them as graph initialisers. Adding an out-of-band scaler step at runtime (e.g., a sidecar JSON consumed by the loader) is forbidden without a follow-up ADR — it splits the trust root and invalidates the registry sha256 contract. - Feature column order is fixed. The graph reads
(adm2, vif_scale0, vif_scale1, vif_scale2, vif_scale3, motion2)in exactly this order; reordering breaks the bundledmean/stdconstants. Any change to the feature set requires a fresh Phase-3 chain (Research-0027 → 0028 → 0029 → 0030). - opset 17. Matches the sister tiny-AI models (
learned_filter_v1,nr_metric_v1,fastdvdnet_pre) and the ORT op-allowlist baseline. Upgrading requires re-validating theSub/Div/Gemm/Relu/Squeezeops againstop_allowlist.c. - On upstream sync: zero interaction. Netflix/vmaf has no equivalent surface; an upstream sync that touches
core/src/dnn/(op-allowlist or model-loader changes) needs to preserveSub/Div/Gemm/Relu/Squeezein the allowlist for opset 17. - Re-test on rebase:
```bash bash core/test/dnn/test_registry.sh python3 core/test/dnn/test_vmaf_tiny_v2.py python3 ai/scripts/validate_vmaf_tiny_v2.py \ --onnx model/tiny/vmaf_tiny_v2.onnx \ --parquet runs/full_features_netflix.parquet \ --rows 100 --min-plcc 0.97 meson test -C build-cpu --suite=dnn
0094 — Tiny-AI extractor template (ADR-0250)¶
- Touches:
core/src/dnn/tiny_extractor_template.h(new),core/src/feature/feature_lpips.c,core/src/feature/fastdvdnet_pre.c,core/src/dnn/AGENTS.md,docs/ai/extractor-template.md(new),docs/adr/0250-tiny-ai-extractor-template.md(new). - Invariants:
- Helper signatures are wire-format-stable.
vmaf_tiny_ai_resolve_model_path(name, option, env_var)andvmaf_tiny_ai_open_session(name, path, &out)produce the user-facing log lines<name>: no model path …and<name>: vmaf_dnn_session_open(<path>) failed: <rc>— downstream tooling greps these. Don't rename or reorder the parameters without bumping every extractor + the recipe doc. - YUV→RGB is bit-exact. The shared
vmaf_tiny_ai_yuv8_to_rgb8_planesis a literal move of the pre-existingfeature_lpips.cbody (BT.709 limited-range, nearest-neighbour chroma upsample). LPIPS / saliency / future colour-sensitive tiny-AI scores depend on byte-exact equality with the prior ad-hoc copies. Any change to the conversion constants or the rounding rule needs a separate ADR + a coordinated snapshot regen —model/tiny/weights aren't re-trained against new colour math casually. - Option-table macro is plain text substitution. The
VMAF_TINY_AI_MODEL_PATH_OPTION(state_t, help)macro emits a single struct literal — no control flow, no recursion, no variadic shenanigans (Power-of-10 rule 1 / rule 9). Don't extend it into a multi-option emitter without a fresh ADR. - On upstream sync: zero interaction with upstream —
feature_lpips.candfastdvdnet_pre.care fork-only files, and the newdnn/tiny_extractor_template.hlives entirely under fork-introducedcore/src/dnn/. An upstream sync that rewrites unrelatedfeature_*.cfiles won't conflict. - Re-test on rebase:
cd libvmaf
meson setup build-cpu -Denable_cuda=false -Denable_sycl=false
ninja -C build-cpu
meson test -C build-cpu --suite=dnn
meson test -C build-cpu test_lpips test_fastdvdnet_pre
# All 10 dnn-suite + both extractor tests must pass.
0095 — Vulkan ring-depth tunable (ADR-0251 follow-up #3)¶
- PR: feat/t7-29-followup3-ring-tunable.
- What rebases need to know:
VmafVulkanConfigurationgrew an additiveunsigned max_outstanding_framesfield. Existing zero-initialised configs continue to receive the canonical default (0 → VMAF_VULKAN_RING_DEFAULT == 4). The clamp helpervmaf_vulkan_clamp_ring_sizemoved fromimport.c(file-local static) tovulkan_internal.h(static inline) sostate_initandlazy_alloc_ringshare one definition; an upstream sync that re-introduces the static inimport.cwould shadow the header helper — drop the duplicate, keep the inline. - New public symbol:
vmaf_vulkan_state_max_outstanding_frames(const VmafVulkanState *)— read-side accessor for the clamped value. Pure additive surface; no upstream collision. - On upstream sync: zero interaction. The ring is wholly fork-introduced (ADR-0251); upstream Netflix has no Vulkan backend.
- Re-test on rebase:
```bash meson setup build libvmaf -Denable_cuda=false -Denable_sycl=false \ -Denable_hip=false -Denable_vulkan=disabled \ -Denable_float=true ninja -C build && meson test -C build # 51/52 OK; 1 pre-existing # T7-32 fail on # test_motion_v2_simd # (ADR-0038 follow-up)
# Smoke the new options + ENOTSUP guard: build/tools/vmaf --reference python/test/resource/yuv/src01_hrc00_576x324.yuv \ --distorted python/test/resource/yuv/src01_hrc01_576x324.yuv \ --width 576 --height 324 --pixel_format 420 --bitdepth 8 \ --feature 'motion_v2=motion_blend_factor=0.5' \ --xml -o /tmp/r.xml --no_prediction grep motion3_v2 /tmp/r.xml | head -3 # → 49 frames with VMAF_integer_feature_motion3_v2_score_mbf_0.5
# ENOTSUP guard: build/tools/vmaf --reference python/test/resource/yuv/src01_hrc00_576x324.yuv \ --distorted python/test/resource/yuv/src01_hrc01_576x324.yuv \ --width 576 --height 324 --pixel_format 420 --bitdepth 8 \ --feature 'motion_v2=motion_five_frame_window=1' \ --xml -o /tmp/r2.xml --no_prediction 2>&1 # → "problem loading feature extractor: motion_v2" # → stderr: "motion_v2: motion_five_frame_window=true is not supported …"
ADR-index backfill 2026-05-08 (this PR)¶
- Touches:
docs/adr/_index_fragments/0235-codec-aware-fr-regressor.md(new),docs/adr/_index_fragments/0236-dists-extractor.md(new),docs/adr/_index_fragments/0238-vulkan-picture-preallocation.md(new),docs/adr/_index_fragments/0239-gpu-picture-pool-dedup.md(new),docs/adr/_index_fragments/0251-vulkan-async-pending-fence.md(new),docs/adr/_index_fragments/0279-fr-regressor-v2-probabilistic.md(new),docs/adr/_index_fragments/_order.txt(six slugs appended),docs/adr/README.md(eight rows appended; one duplicate ADR-0279 row deduplicated). - Invariant: no engine code touched; no upstream-shared paths. Pure fork-local index maintenance.
- On upstream sync: no action required.
docs/adr/is a fork-local tree. - Coordination with #468 (27-ADR status sweep): both PRs touch ADR metadata. They do not conflict at the file level (#468 edits ADR bodies; this PR adds index fragments + appends README rows for the eight previously-unindexed ADRs). At merge time the README append-tail may overlap if #468 lands later index rows for its swept ADRs; whichever lands first, the second rebases by re-running
scripts/docs/concat-adr-index.sh --checkand inserting any newly-stale rows in commit-merge order. - Known finding (out of scope):
scripts/docs/concat-adr-index.sh --checkcurrently reports a much larger fragment-vs-README drift than this PR introduces — many ADRs have rows inREADME.mdwithout corresponding_index_fragments/files, and several_order.txtslugs have no fragment yet. Running--writeblindly would drop ~37 README rows for ADRs unrelated to this PR. The ADR-0221 fragment-driven contract therefore could not be enforced via a clean--writehere; eight new rows were appended directly to keep the change scoped. A separate sweep PR is needed to flush the residual drift. - Re-test on rebase:
for n in 0235 0236 0238 0239 0251 0276 0279 0315; do
grep -cE "^\| \[ADR-$n\]" docs/adr/README.md # must be ≥ 1
done
bash scripts/docs/concat-adr-index.sh >/dev/null # must succeed
-Denable_vulkan=enabled
ninja -C build meson test -C build test_vulkan_async_pending_fence
# All 8 cases must pass: 4 v2-contract + 4 ring-tunable.
0096 — tools/vmaf-tune/ automation umbrella spec (ADR-0237 / Research-0044)¶
- PR: feat/vmaf-tune-spec.
- What rebases need to know: this PR ships only an umbrella ADR research digest under
docs/. No tracked source code, notools/vmaf-tune/directory yet, no Meson changes. An upstream sync touching ffmpeg-patches orlibvmaf/cannot collide with this PR. - On upstream sync: zero interaction. Spec-only PR.
- Re-test on rebase:
# No build/test impact — verify the docs render and links are alive:
ls docs/adr/0237-quality-aware-encode-automation.md \
docs/research/0044-quality-aware-encode-automation.md
grep -c '\[ADR-0237\]' docs/adr/README.md
0097 — test_speed gated on enable_float (fix default-build failure)¶
- PR: fix/test-speed-chroma-registration.
- What rebases need to know:
core/test/meson.buildnow wraps thetest_speedexecutable +test()registration inif get_option('enable_float'). Thespeed_chroma/speed_temporalextractors live inspeed.c, which is only compiled whenenable_float=true(the entries infeature_extractor.care wrapped in#if VMAF_FLOAT_FEATURES), so the test'svmaf_get_feature_extractor_by_name("speed_chroma")returned NULL on a default build (enable_float=false). - On upstream sync: zero interaction.
test_speed.cwas added fork-side via the Netflix port commitd3647c73. The gating pattern matchestest_vulkan_*(if get_option('enable_vulkan').enabled()). - Re-test on rebase:
# default (enable_float=false): test_speed must NOT be in the suite
meson setup build libvmaf -Denable_cuda=false -Denable_sycl=false --reconfigure
ninja -C build
meson test -C build # expect: NO test_speed in the run
# CI shape (enable_float=true): test_speed must run + pass
meson setup build libvmaf -Denable_float=true --reconfigure
ninja -C build
meson test -C build test_speed # expect: 5/5 pass
0098 — Vulkan picture preallocation surface (ADR-0238)¶
- PR: feat/vulkan-picture-preallocation.
- What rebases need to know: ABI grows additively. New public surface in
core/include/libvmaf/libvmaf_vulkan.h:enum VmafVulkanPicturePreallocationMethod,VmafVulkanPictureConfiguration,vmaf_vulkan_preallocate_pictures,vmaf_vulkan_picture_fetch. New enumeratorVMAF_PICTURE_BUFFER_TYPE_VULKAN_DEVICEincore/src/picture.h::VmafPictureBufferType. New TUcore/src/vulkan/picture_vulkan_pool.c(~180 LOC); registered incore/src/vulkan/meson.build. Fork-internal accessorvmaf_vulkan_state_context()(declared invulkan_internal.h) exposes the imported state's VkInstance/VkDevice to the pool — used only bylibvmaf.c::vmaf_vulkan_preallocate_pictures. VmafContextfield added:vmaf->vulkan.poolnext tovmaf->vulkan.state. Thevmaf_close()teardown closes the pool before clearing the state pointer (matches SYCL).- On upstream sync: zero interaction. Vulkan backend is fork-only; upstream Netflix has no Vulkan integration.
- Re-test on rebase:
```bash meson setup build libvmaf -Denable_cuda=false -Denable_sycl=false \ -Denable_vulkan=enabled ninja -C build meson test -C build test_vulkan_pic_preallocation # All 6 cases must pass under ASan/UBSan: # test_method_none_is_a_no_op # test_method_host_allocates_round_robins # test_method_device_allocates_round_robins # test_fetch_without_preallocate_falls_back # test_unknown_method_rejected # test_null_args_rejected
0099 — feature_mobilesal.c + transnet_v2.c migrated to tiny_extractor_template.h¶
- PR: refactor/migrate-ai-to-template.
- What rebases need to know:
feature_mobilesal.candtransnet_v2.cpreviously open-coded the model-path resolution (getenv+ log block), the YUV→RGB kernel (mobilesal only), thevmaf_dnn_session_open+ log boilerplate, and theVmafOption[].model_pathrow. They now use the helpers fromdnn/tiny_extractor_template.h(PR #251) — the same templatefeature_lpips.candfastdvdnet_pre.calready consume. Net −98 LOC of identical boilerplate. - Behavior preserved: bit-exact YUV→RGB conversion (mobilesal used the literal copy of
feature_lpips.c's body that the template hoisted), identical error-log strings, identical option-table flag/type/offset shape. The migratedmobilesal_optionsmacro expands to the same struct literal the hand-rolled version produced. - On upstream sync: zero interaction. Both files are fork-introduced; upstream Netflix has neither extractor.
0100 — cuda/ring_buffer.{c,h} → gpu_picture_pool.{c,h} (ADR-0239)¶
- PR: refactor/gpu-picture-pool-extract.
- What rebases need to know:
core/src/cuda/ring_buffer.candring_buffer.hare removed. The same callback-based round-robin pool lives atcore/src/gpu_picture_pool.{c,h}under renamed symbols (VmafRingBuffer→VmafGpuPicturePool,vmaf_ring_buffer_*→vmaf_gpu_picture_pool_*,_fetch_next_picture→_fetch). All call sites inlibvmaf.cmigrated.core/test/test_ring_buffer.crenamed totest_gpu_picture_pool.cwith the corresponding meson update. - Netflix-upstream interaction: minimal — Netflix's
cuda/ring_buffer.{c,h}last touched in commitcb1d49c6. An upstream sync that resurrects the old names should be redirected to the new ones; the file move is purely fork-local. Netflix#1300mutex-destroy-order fix preserved (ADR-0157) — moved verbatim to the new file; the fix remains attached tovmaf_gpu_picture_pool_close.- SYCL pool migration:
vmaf_sycl_picture_pool_*keeps its public-internal API but now delegates to the generic pool. The SYCL wrapper struct (VmafSyclPicturePool) just owns theVmafSyclCookiestorage.std::mutexdrops out. - Vulkan pool migration: bundled into this PR after #264 merged.
picture_vulkan_pool.crewrites as a thin wrapper around the generic pool — wrapper struct owns per-pool state for the alloc/free callbacks; the generic pool owns the round-robin slots / mutex / unwind. Same pattern as the SYCL migration above. - Re-test on rebase:
meson setup build libvmaf -Denable_cuda=false -Denable_sycl=false
ninja -C build
meson test -C build --suite=dnn
meson test -C build test_lpips test_mobilesal test_transnet_v2 test_fastdvdnet_pre
# All 11 dnn-suite + 4 extractor smoke tests must pass.
meson test -C build # 47/47 pass under ASan/UBSan
# CUDA build (CI-only; pre-existing local nvcc include-path quirk):
meson setup build-cuda libvmaf -Denable_cuda=true
ninja -C build-cuda
meson test -C build-cuda test_gpu_picture_pool
# SYCL build:
meson setup build-sycl libvmaf -Denable_sycl=true
ninja -C build-sycl
meson test -C build-sycl
0104 — psnr_vulkan.c migrated to vulkan/kernel_template.h¶
- PR: refactor/migrate-psnr-vulkan-to-template.
- What rebases need to know:
vulkan/kernel_template.h(410 LOC, ADR-0246, PR #251) shipped with zero consumers. Its docstring designatedpsnr_vulkan.cas the reference implementation. This PR lands the migration as the first consumer of the Vulkan template — paired with PR #269 (the first CUDA template consumer). The 5 long-lived pipeline objects (descriptor-set layout, pipeline layout, shader module, compute pipeline, descriptor pool) collapse from individual struct fields to oneVmafVulkanKernelPipeline plbundle.create_pipeline()(~104 LOC) collapses to a singlevmaf_vulkan_kernel_pipeline_create()call (~30 LOC) — the template owns the descriptor-set layout creation, pipeline layout, shader module, compute pipeline, and descriptor-pool sizing.close_fex()'svkDeviceWaitIdle+ 5×vkDestroy*sweep collapses to onevmaf_vulkan_kernel_pipeline_destroy()call. - Net LOC delta: −55 LOC on
psnr_vulkan.cdirectly. Unlike the CUDA template (where helper-call boilerplate roughly matches the inline savings), the Vulkan template's pipeline creation is dramatic enough that even the first consumer wins. - Bit-exactness gates: spec-constants, push-constant struct, shader bytecode, dispatch grid math, and host-side reduction are byte-identical to the prior implementation. The template only owns descriptor-set layout / pipeline layout / shader module / compute pipeline creation / descriptor pool sizing — none of which affects the kernel's mathematical behaviour. Cross-backend parity gate (places=4) re-runs unchanged.
- On upstream sync: zero interaction.
psnr_vulkan.cis fork-introduced (T7-23 / ADR-0182 / ADR-0216). - Re-test on rebase:
meson setup build libvmaf -Denable_cuda=false -Denable_sycl=false \
-Denable_vulkan=enabled
ninja -C build
meson test -C build # 50/50 pass on lavapipe
# Cross-backend parity gate (places=4):
python scripts/ci/cross_backend_parity_gate.py --feature psnr_y --places 4
0105 — moment_vulkan.c + ciede_vulkan.c migrated to vulkan/kernel_template.h¶
- PR: refactor/migrate-motion-vulkan-to-template (note: the branch name reflects the original intent; motion's two-pipeline shape didn't fit the template's single-pipeline contract, so this PR migrates moment + ciede instead).
- What rebases need to know: second + third consumers of
vulkan/kernel_template.h(after PR #270 = psnr_vulkan, the first consumer). Both files follow the identical migration pattern: - Replace 5 individual pipeline-object fields (
dsl,pipeline_layout,shader,pipeline,desc_pool) with oneVmafVulkanKernelPipeline plbundle. - Replace ~100 LOC of
create_pipeline()body (descriptor-set layout + pipeline layout + shader module + compute pipeline + descriptor pool boilerplate) with a singlevmaf_vulkan_kernel_pipeline_create()call. - Replace
close_fex()'svkDeviceWaitIdle+ 5×vkDestroy*sweep with onevmaf_vulkan_kernel_pipeline_destroy()call. - Per-file LOC deltas:
moment_vulkan.c: −60 LOC (450 → 390).ciede_vulkan.c: −59 LOC (536 → 477).- Net: −119 LOC.
- Bit-exactness preserved: spec-constants (width/height/bpc/ subgroup_size identical across both), push-constant structs (
MomentPushConsts,CiedePushConsts), shader bytecodes (moment_spv,ciede_spv), dispatch grid math, and host-side reductions are byte-identical to the prior implementation. Cross-backend parity gates (places=4 for moment integer reduce; places=2 for ciede transcendentals per ADR-0187) re-run unchanged. motion_vulkan.cdeferred: motion uses two pipelines (first frame vs subsequent) sharing one DSL + layout + shader + pool. The template's current shape produces one pipeline per descriptor; splitting motion across twoVmafVulkanKernelPipelineinstances would duplicate the shared objects. Tracked as a follow-up template extension (multi-pipeline support).- On upstream sync: zero interaction. Both files are fork-introduced (T7-23 / ADR-0182 / ADR-0187).
- Re-test on rebase:
```bash meson setup build libvmaf -Denable_cuda=false -Denable_sycl=false \ -Denable_vulkan=enabled ninja -C build meson test -C build # 50/50 pass on lavapipe (under ASan/UBSan) python scripts/ci/cross_backend_parity_gate.py --feature float_moment_ref1st --places 4 python scripts/ci/cross_backend_parity_gate.py --feature ciede2000 --places 2
0101 — GPU backend pattern doc (ADR-0240)¶
- PR: docs/gpu-backend-template.
- What rebases need to know: doc-only PR. Adds
docs/development/gpu-backend-template.md(recipe new GPU backends follow) andcore/include/libvmaf/AGENTS.md(public-headers-tree invariant note). No source code, no meson changes, no ABI impact. - On upstream sync: zero interaction. Both files are fork-introduced.
- Re-test on rebase:
```bash # Doc-only — verify links resolve: test -f docs/development/gpu-backend-template.md test -f core/include/libvmaf/AGENTS.md grep -c 'gpu-backend-template' core/include/libvmaf/AGENTS.md
0102 — Tiny-AI test registration macro (tiny_ai_test_template.h)¶
- PR: refactor/test-registration-macro.
- What rebases need to know: new
core/test/tiny_ai_test_template.hemits the four standard registration tests (<name>_is_registered,<name>_provides_primary_feature,<name>_options_table_well_formed,<name>_init_rejects_missing_model) via theVMAF_TINY_AI_DEFINE_REGISTRATION_TESTS(ext, feat, env, prefix)macro. The four per-extractor test files (test_lpips.c,test_mobilesal.c,test_transnet_v2.c,test_fastdvdnet_pre.c) shrank from ~140 LOC each to ~20-50 LOC. Net −286 LOC. Behavior bit-exact preserved (same assertions, same env-var save/restore dance, same setenv shim for MSVCRT). TransNet V2 keeps two extractor-specific extra tests (binary-flag round-trip + provided_features list-termination) that the macro doesn't cover. - On upstream sync: zero interaction. The four test files are fork-introduced (per ADR-0042 / ADR-0168 / ADR-0220 / ADR-0223 / ADR-0215).
- Re-test on rebase:
```bash meson setup build libvmaf -Denable_cuda=false -Denable_sycl=false ninja -C build meson test -C build test_lpips test_mobilesal test_transnet_v2 test_fastdvdnet_pre # 4/4 binaries pass; 18 individual tests total (4x4 standard + 2 # TransNet V2 extras).
0103 — integer_psnr_cuda.c migrated to cuda/kernel_template.h¶
- PR: refactor/migrate-psnr-cuda-to-template.
- What rebases need to know:
cuda/kernel_template.hshipped with no consumers in PR #251 (ADR-0246). This PR migrates the first consumer (integer_psnr_cuda.c) — the file the template's own docstring explicitly designated as the reference. TheCUstream + CUevent + CUeventtriple and the(VmafCudaBuffer device, void *host_pinned, size_t bytes)readback pair are now dispensed by the template helpers (vmaf_cuda_kernel_lifecycle_init/_close,vmaf_cuda_kernel_readback_alloc/_free,vmaf_cuda_kernel_submit_pre_launch,vmaf_cuda_kernel_collect_wait) instead of being open-coded.PsnrStateCudashrinks: replaces three fields (event+finished+str) with oneVmafCudaKernelLifecyclereplaces (sse+sse_host) with oneVmafCudaKernelReadback. - Net LOC delta: +8 LOC on
integer_psnr_cuda.calone — the helpers add per-call boilerplate. The dedup win materialises as more CUDA feature kernels (motion / moment / ssim / vif / adm) migrate one-at-a-time in follow-up PRs. Each subsequent migration saves ~15 LOC. - Bit-exactness gates: kernel launch + reduction logic unchanged. The migration only touches state-management boilerplate around the kernel; the SSE accumulator math, the per-bpc kernel function lookup, the host-side
log10score formula, and the dispatch grid-dim calculation are byte-identical to the prior implementation. Netflix golden gate + CPU/CUDA cross-backend parity gate (places=4) re-run unchanged. - On upstream sync: zero interaction.
integer_psnr_cuda.cis fork-introduced (T7-23 / ADR-0182). - Re-test on rebase:
meson setup build libvmaf -Denable_cuda=true
ninja -C build
meson test -C build # CUDA test suite must pass
# Cross-backend parity gate:
python scripts/ci/cross_backend_parity_gate.py --feature psnr_y --places 4
0125 — Vulkan submit-side template + fence pool + descriptor pre-alloc bundle (ADR-0256)¶
- Touches:
core/src/vulkan/kernel_template.h— fork-local. Output landing inruns/phase_a/is gitignored — rerun the script to reproduce.VmafVulkanKernelSubmitPoolstruct +_create/_destroy/_acquirehelpers +vmaf_vulkan_kernel_descriptor_sets_allochelper. Upstream has no Vulkan backend — no merge surface.core/src/feature/vulkan/{psnr_hvs,vif,float_vif,float_adm}_vulkan.c— fork-local kernel TUs, also no upstream peer.- Invariant: the four migrated kernels keep all per-frame
VkFence+VkCommandBuffer+VkDescriptorSetresources alive across frames in the pool. Pre-bound descriptor sets rely on the kernel'sVmafVulkanBuffer *handles being init-time stable (allocated ininit(), freed only inclose_fex).vmaf_vulkan_kernel_pipeline_destroydestroys the descriptor pool — pre-allocated sets are released implicitly via the pool; callers must NOT callvkFreeDescriptorSetson them. - Re-test on rebase:
meson setup build libvmaf -Denable_vulkan=enabled
ninja -C build
VK_ICD_FILENAMES=/usr/share/vulkan/icd.d/nvidia_icd.json \
meson test -C build test_vulkan_smoke \
test_vulkan_async_pending_fence \
test_vulkan_pic_preallocation
python scripts/ci/cross_backend_vif_diff.py \
--vmaf-binary build/tools/vmaf \
--reference testdata/ref_576x324_48f.yuv \
--distorted testdata/dis_576x324_48f.yuv \
--width 576 --height 324 --pixel-format 420 --bitdepth 8 \
--feature vif --backend vulkan --places 4
python scripts/ci/cross_backend_vif_diff.py \
--vmaf-binary build/tools/vmaf \
--reference testdata/ref_576x324_48f.yuv \
--distorted testdata/dis_576x324_48f.yuv \
--width 576 --height 324 --pixel-format 420 --bitdepth 8 \
--feature adm --backend vulkan --places 4
0107 — psnr_hvs_cuda async upload + persistent pinned staging (T-GPU-OPT-2/3)¶
- Touches:
core/src/feature/cuda/integer_psnr_hvs_cuda.c— only consumer; fork-local from inception (T7-23 / ADR-0188 / ADR-0191). State addsupload_str(dedicated H2D stream),upload_done(cross-stream completion event), and per-plane persistent pinnedh_uint_ref[3]/h_uint_dist[3]staging buffers allocated once ininit_fex_cuda. The per-call helperupload_plane_cudais split intoissue_d2h_plane(pic-stream D2H),convert_plane(CPU normalise), andissue_h2d_plane(upload-stream H2D).submit_fex_cudaruns the three phases explicitly and recordsupload_doneafter the last H2D, thencuStreamWaitEvents onlc.strbefore kernel launches.core/src/cuda/AGENTS.md— adds a rebase-sensitive invariant entry under §Rebase-sensitive invariants documenting the three-phase flow + persistent staging contract.- Invariant: the pinned
h_uint_*andh_ref/h_distbuffers are never freed and re-allocated mid-stream; the H2Ds must run onupload_str(not onlc.str) so thecuStreamWaitEventcross-stream link is meaningful; theupload_doneevent is recorded after the last H2D for the current frame and waited on once before the first kernel launch of that frame. CUDA graph capture (future T-GPU-OPT-N) depends on the no-per-frame-alloc invariant; collapsing the three-phase split or re-introducing per-framevmaf_cuda_buffer_host_alloccalls breaks that follow-up. Bit-exactness gate isplaces=3forpsnr_hvs_y / cb / crand the combinedpsnr_hvs(matches the existing matrix; notplaces=4). - On upstream sync: zero interaction.
integer_psnr_hvs_cuda.cis fork-introduced (T7-23 / ADR-0188 / ADR-0191). - Re-test on rebase:
meson setup build libvmaf -Denable_cuda=true -Denable_sycl=false
ninja -C build
meson test -C build
python3 scripts/ci/cross_backend_vif_diff.py \
--vmaf-binary core/build/tools/vmaf \
--reference python/test/resource/yuv/src01_hrc00_576x324.yuv \
--distorted python/test/resource/yuv/src01_hrc01_576x324.yuv \
--width 576 --height 324 --pixel-format 420 --bitdepth 8 \
--feature psnr_hvs --backend cuda --places 3
0227 — output.c writer-format unit tests (R3 of coverage-gap-2026-05-02)¶
- Touches:
core/test/test_output.c(new) — exercises the four writers incore/src/output.c(XML / JSON / CSV / SUB) end-to-end viatmpfile()-backed sinks and a syntheticVmafFeatureCollector. Pure test-only; no production code change.core/test/meson.build— registerstest_outputnext totest_feature_collector(mirrors that test's wiring:link_with: libvmaf+ libsvm objects + log/predict/metadata helpers).- Invariant: the test pulls
libvmaf.candoutput.cin via#include "*.c"(mirroring the precedent intest_feature_collector.c) so the per-translation-unit.gcnolands in the test build dir and gcovr aggregates output.c's coverage. The mu-test framework macro (mu_assert) deliberately early-returns from eachstatic char *test_*()body — that's why every test body tripsclang-analyzer-unix.Malloc"potential leak" notes (cleanup runs only on the success-tail path). This pattern is shared across everycore/test/test_*.cfile and is load- bearing (per ADR-0141 NOLINT carve-out): replacing it with goto- cleanup would obscure the per-assertion failure message. - On upstream sync: zero interaction.
output.cis upstream- mirrored, but this PR doesn't touch it. The test only depends on the four public function signatures (vmaf_write_output_{xml, json,csv,sub}); if Netflix renames or reorders those, the test fails to compile and the rebase author updates it then. - Re-test on rebase:
cd libvmaf && meson setup build -Denable_cuda=false -Denable_sycl=false
ninja -C build && ./build/test/test_output
0126 — OSSF Scorecard policy (ADR-0263)¶
- Touches:
.github/workflows/scorecard.yml(line 45 — thegithub/codeql-action/upload-sarif@<sha>pin). The rest of the policy is doc-only (docs/adr/0263-*.md,docs/research/0053-*.md,changelog.d/security/). Upstream Netflix/vmaf does not ship a Scorecard workflow, so the path itself is fork-introduced and won't conflict. - Invariant: the
upload-sarifSHA must point to a commit that currently exists ingithub/codeql-action's git tree. A SHA that was oncev4head but no longer exists in the action repository triggers Scorecard's "imposter commit" defence and breaks the workflow with a 400 error againstapi.scorecard.dev. Verify on every Dependabot bump by spot-checkinggh api /repos/github/codeql-action/commits/<sha>returns 200. - On upstream sync: zero interaction.
- Re-test on rebase:
```bash # Confirm the pin still resolves to a real commit: pin=$(grep -oE 'codeql-action/upload-sarif@[a-f0-9]{40}' \ .github/workflows/scorecard.yml | head -1 | cut -d@ -f2) gh api "/repos/github/codeql-action/commits/$pin" --jq '.sha' # Then watch the next master push for a green Scorecard run: gh run list --workflow scorecard --repo VMAFx/vmafx --limit 1
0228 — U-2-Net u2netp saliency replacement deferred (ADR-0265)¶
- Touches: docs-only.
docs/adr/0265-u2netp-saliency-replacement-blocked.md— new ADR continuing the deferral chain started by ADR-0257.docs/research/0055-u2netp-saliency-replacement-survey.md— new research digest (upstream survey + license + distribution- op-allowlist audit + alternatives walk).
docs/ai/models/mobilesal.md— pointer block updated to reference both ADR-0257 (first blocker) and ADR-0265 (second blocker).model/tiny/registry.json—mobilesal_placeholder_v0notesfield updated to reference ADR-0265 alongside ADR-0257 (no schema / sha256 / file changes).model/tiny/mobilesal.json— sidecarnotesfield updated in lockstep.scripts/gen_mobilesal_placeholder_onnx.py— generator notes string updated so re-running is idempotent against the new sidecar / registry text.CHANGELOG.md— Changed entry viachangelog.d/changed/T6-2a-followup-u2netp-replacement-deferred.md.docs/adr/README.md— index row viadocs/adr/_index_fragments/0265-u2netp-saliency-replacement-blocked.md.- Invariant: zero C-side surface change.
feature_mobilesal.ctensor-name contract (inputinput→ outputsaliency_map, NCHW float32[1, 3, H, W]→[1, 1, H, W]) is unchanged; the on-diskmodel/tiny/mobilesal.onnx(sha256f1226310…) is unchanged;mobilesal_placeholder_v0'ssmoke: trueflag is unchanged. Any future drop-in (U-2-Net viaT6-2a-mirror-u2netp-via-release+T6-2a-widen-allowlist-resize, distilled student, or BASNet / PoolNet survey result) replaces the.onnxand bumps the registry sha256 without touching the C side. - On upstream sync: zero interaction.
feature_mobilesal.c, the registry, the ADR, and the research digest are all fork-local (T6-2a; ADR-0218 / ADR-0257 / ADR-0265; not present in Netflix upstream). - Re-test on rebase:
meson setup build -Denable_cuda=false -Denable_sycl=false
ninja -C build
meson test -C build test_mobilesal
python3 ai/scripts/validate_model_registry.py
bash scripts/docs/concat-adr-index.sh --check
bash scripts/release/concat-changelog-fragments.sh --check
0108 — ssim_accumulate_avx512 per-lane double reduction vectorised¶
- ADR: ADR-0139 (existing; no new ADR — the per-lane reduction order is unchanged).
- Touches:
core/src/feature/x86/ssim_avx512.c— thessim_accumulate_block_avx512body. The per-lane scalarssim_accumulate_lanecalls (16 of them) are replaced by two 8-wide__m512dpasses that computelv,cv,sv, andlv*cv*svlane-wise in vector double. Aligneddouble[16]spill buffers replace the previous_Alignas(64) float[16]×6spill, and the scalar accumulation loop now does 4×16vaddsdinstead of 16 invocations of the per-lane helper.CHANGELOG.md— Changed entry.- This file — this entry.
- Invariant (load-bearing for ADR-0139 bit-exactness):
- Per-lane double computation order is byte-identical:
((2.0 * rm) * cm + C1) / l_den, then(2.0 * srsc + C2) / c_den, then(lv * cv) * sv. No FMA contraction (separate_mm512_mul_pd+_mm512_add_pd—_mm512_fmadd_pdis forbidden because it changes the rounding count and would diverge from scalar's two-stepmul+add). - Float→double widening uses
_mm512_cvtps_pdwhich is IEEE-754-exact for finite floats (52-bit mantissa fits 23-bit float losslessly). - Lane-by-lane left-to-right reduction order preserved:
local_ssim += t_ssim[k]fork = 0..15. Tree reductions (pairwise add, dual-accumulator unroll) are forbidden — they break running-sum associativity against scalar. - AVX2 / NEON twins kept on the per-lane scalar path. Verified bit-identical against the new AVX-512 at
--precision maxon the Netflixsrc01_hrc00/01_576x324and thecheckerboard_1920_1080_10_3_*_0pairs. The bit-exactness contract (ADR-0139) is per-lane, not per-ISA algorithm — so AVX2 / NEON stay scalar-per-lane until a dedicated PR vectorises them with the same care. - Rebase impact: zero conflict with Netflix upstream — the whole SSIM SIMD surface is fork-local (no upstream SSIM SIMD exists). Conflicts only arise if upstream changes
ssim_accumulate_default_scalariniqa/ssim_tools.c; in that case both the AVX2 / NEON per-lane helper and the AVX-512 vector-double block need a coordinated update preserving the three invariants above. - Re-test on rebase:
meson setup build libvmaf -Denable_cuda=false -Denable_sycl=false
ninja -C build
meson test -C build
# Bit-exact at --precision max, scalar vs AVX2 vs AVX-512:
for MASK in 0 16 255; do
core/build/tools/vmaf -r python/test/resource/yuv/src01_hrc00_576x324.yuv \
-d python/test/resource/yuv/src01_hrc01_576x324.yuv \
-w 576 -h 324 -p 420 -b 8 \
--feature float_ms_ssim --feature float_ssim \
--xml -o /tmp/m${MASK}.xml --precision max --cpumask $MASK
done
diff <(grep -v 'fyi fps' /tmp/m0.xml) <(grep -v 'fyi fps' /tmp/m16.xml) # empty
diff <(grep -v 'fyi fps' /tmp/m0.xml) <(grep -v 'fyi fps' /tmp/m255.xml) # empty
- Why this matters on rebase: an upstream commit that touches
core/src/feature/ssimulacra2.ccould prompt a "let's also port the GPU XYB while we're here" follow-up. The ledger entry is the standing answer: don't, the measurement was redone on NVIDIA in May 2026 and the result still failedplaces=4by five decades. See Research-0047.
0126 — FastDVDnet real upstream weights drop (ADR-0253)¶
- What changed: replaces
model/tiny/fastdvdnet_pre.onnxwith the wrapped real upstream FastDVDnet checkpoint (sha256eb9444cf6f07eefdc7f4f68d09131074dbd1dcee6f88a331ba684dd2fb5937d4, ~9.5 MiB), refreshes the sidecarmodel/tiny/fastdvdnet_pre.json, flips the registry row'ssmoke: true → falseand addslicense: "MIT"+ the upstream commit pinc8fdf61. New exporterai/scripts/export_fastdvdnet_pre.py(the older_placeholder.pyexporter is retained for reference). New ADRdocs/adr/0255-fastdvdnet-pre-real-weights.md; user-facing docdocs/ai/models/fastdvdnet_pre.mdrewritten with provenance, license attribution, and reproduce-the-export instructions. - Upstream source: fork-local. Netflix/vmaf does not ship a FastDVDnet temporal pre-filter; the C extractor and ONNX surface are entirely fork-introduced (ADR-0215). The wrapped weights are attribution-only (upstream
m-tassano/fastdvdnetMIT). - On upstream sync: zero interaction. Every file touched (
ai/scripts/export_fastdvdnet_pre*.py,model/tiny/fastdvdnet_pre.*,docs/ai/models/fastdvdnet_pre.md,docs/adr/0253-*.md, CHANGELOG fragment, ADR index fragment) lives in fork-introduced trees. - Re-test on rebase:
# Re-derive the ONNX from the pinned upstream checkpoint.
mkdir -p /tmp/fastdvdnet_upstream && cd /tmp/fastdvdnet_upstream
curl -L -O https://raw.githubusercontent.com/m-tassano/fastdvdnet/c8fdf61/model.pth
curl -L -O https://raw.githubusercontent.com/m-tassano/fastdvdnet/c8fdf61/models.py
cd /path/to/vmaf
python3 ai/scripts/export_fastdvdnet_pre.py \
--upstream-dir /tmp/fastdvdnet_upstream
python3 ai/scripts/validate_model_registry.py
meson test -C build --suite=fast --print-errorlogs test_fastdvdnet_pre
0127 — ONNX op-allowlist gains Resize (ADR-0258)¶
- Touches:
core/src/dnn/op_allowlist.c— fork-local file (no upstream counterpart). One new entry"Resize"under the/* convolutional */block.core/test/dnn/test_op_allowlist.c,core/test/dnn/test_onnx_scan.c— fork-local DNN tests.ai/tests/test_op_allowlist.py— fork-local Python parity test.- Invariant: the C allowlist is the single source of truth; the Python regex parser in
ai/src/vmaf_train/op_allowlist.pywalks the sameop_allowlist.cfile. Any future entry only needs the C edit — Python symmetry is automatic. - Upstream source: fork-local. Netflix/vmaf has no ONNX op- allowlist surface; the entire
core/src/dnn/tree is fork- introduced. - On upstream sync: zero interaction. Every file touched lives in fork-introduced trees.
- Re-test on rebase:
meson test -C build test_op_allowlist test_onnx_scan
PYTHONPATH=ai/src python -m pytest ai/tests/test_op_allowlist.py
0231 — vif.comp + ciede.comp precise decorations (ADR-0269 / Step A of Vulkan 1.4 bump)¶
- Touches:
core/src/feature/vulkan/shaders/vif.comp(3 local-variable type qualifiers:g,sv_sq,gg_sigma_f→precise float),core/src/feature/vulkan/shaders/ciede.comp(yuv_to_rgboutputs,rgb_to_xyzmatmul accumulators,ciede2000chroma magnitudes + half-axes + s_l/c/h + lightness/chroma/hue + final ΔE). - Invariant: Both shaders are fork-local (Vulkan backend is fork-added; upstream Netflix/vmaf has no Vulkan compute kernels). The
precisekeyword is GLSL 4.50 standard syntax; glslc 2026.1 lowers it to per-resultOpDecorate NoContraction. The decorations are load-bearing for the cross-backend gate on NVIDIA driver 595.71+ — removing them would re-introduce the 42/48 ciede regression at API 1.3 documented in research-0054. - On upstream sync: zero interaction. Both shader files are entirely fork-introduced; upstream has no Vulkan compute path.
- Re-test on rebase:
# Re-confirm the cross-backend gate on a Vulkan-capable host.
meson setup core/build -Denable_vulkan=enabled
ninja -C core/build
python3 scripts/ci/cross_backend_vif_diff.py \
--vmaf-binary core/build/tools/vmaf \
--reference python/test/resource/yuv/src01_hrc00_576x324.yuv \
--distorted python/test/resource/yuv/src01_hrc01_576x324.yuv \
--width 576 --height 324 --pixel-format 420 --bitdepth 8 \
--feature vif --backend vulkan --places 4
python3 scripts/ci/cross_backend_vif_diff.py \
--vmaf-binary core/build/tools/vmaf \
--reference python/test/resource/yuv/src01_hrc00_576x324.yuv \
--distorted python/test/resource/yuv/src01_hrc01_576x324.yuv \
--width 576 --height 324 --pixel-format 420 --bitdepth 8 \
--feature ciede --backend vulkan --places 4
# Confirm SPIR-V still emits NoContraction post-rebase.
glslc --target-env=vulkan1.3 -O \
core/src/feature/vulkan/shaders/vif.comp -o /tmp/vif.spv
spirv-dis /tmp/vif.spv | grep -c NoContraction # expect ≥ 60
Expected on NVIDIA 595.71+: vif 0/48 OK, ciede 5/48 FAIL (max abs 8.9e-05 — pre-existing fork debt at API 1.3, see ADR-0269). On RADV / lavapipe: bit-exact (precise is a no-op there).
0229 — fr_regressor_v2 codec-aware scaffold (ADR-0272)¶
- ADR: ADR-0272
- Touches:
ai/scripts/train_fr_regressor_v2.py(new) — Phase A JSONL consumer; trains the codec-aware FRRegressor.model/tiny/fr_regressor_v2.onnx(new, smoke) — placeholder ONNX from--smokemode; re-baked on production training.model/tiny/fr_regressor_v2.json(new) — sidecar.model/tiny/registry.json— new entry withsmoke: true.docs/adr/0272-fr-regressor-v2-codec-aware-scaffold.md(new).docs/adr/README.md— index row.docs/research/0058-fr-regressor-v2-feasibility.md(new).docs/ai/models/fr_regressor_v2.md(new) — model card.ai/AGENTS.md— invariant note (codec block layout + ENCODER_VOCAB ordering).CHANGELOG.md— Added entry.- Invariant: the 8-D codec block layout is
[encoder_onehot(6), preset_norm, crf_norm]withENCODER_VOCAB = (libx264, libx265, libsvtav1, libvvenc, libvpx-vp9, unknown)in load-bearing order. CRF normaliser is/63(union upper bound). Preset normaliser is/9. Bumping the vocabulary requires a re-train; existing checkpoints pin the order they were trained against viaencoder_vocab_versionin the sidecar. The two-input ONNX (features,codec) follows the LPIPS-Sq precedent (ADR-0040 / ADR-0041). - Rebase impact: entirely fork-local; pure additive; no upstream-mirror file is touched. Phase A schema (consumed by this trainer) is itself fork-local (
tools/vmaf-tune/). No conflict expected on/sync-upstream. - Re-test on rebase:
0311 — libFuzzer harness expansion: yuv_input + cli_parse (ADR-0311)¶
- ADR: ADR-0311; parent ADR-0270.
- Touches:
core/test/fuzz/fuzz_yuv_input.c(new)core/test/fuzz/fuzz_cli_parse.c(new)core/test/fuzz/meson.build— two newexecutable(...)blocks for the harnesses, plus a sharedfuzz_vidinput_sourceslist.core/test/fuzz/yuv_input_corpus/*(new — 6 seeds covering 8/10-bit × 4:2:0 / 4:2:2 / 4:4:4 plus a truncated-frame seed).core/test/fuzz/cli_parse_corpus/*(new — 6 seeds covering the--feature,--model,--reference, YUV-flag, and--helpshapes).core/test/fuzz/README.md— Targets table extended..github/workflows/fuzz.yml— matrix gainsfuzz_yuv_input+fuzz_cli_parse; per-harness wall-clock budget reduced from 300 s to 60 s so the 3-target matrix fits the existingtimeout-minutes: 15cap.docs/development/fuzzing.md— runbook table + smoke commands extended.docs/adr/0311-libfuzzer-harness-expansion.md(new)docs/research/0083-libfuzzer-harness-expansion-target-survey.md(new)libvmaf/AGENTS.md— new invariant block for the one-parser-one-harness rule.CHANGELOG.md— Added entry.- Invariant:
- The fuzz scaffold remains opt-in (
-Dfuzz=true) — every defaultmeson setupinvocation must continue to skip it. fuzz_yuv_inputre-includestools/yuv_input.cand the rest of the vidinput trio as build inputs. Upstream Netflix/vmaf splits or renames of those source files need the matchingmeson.buildsource-list update.fuzz_cli_parsere-includestools/cli_parse.cas a build input and links againstlibvmafforvmaf_version()and feature-dictionary symbols. The-Wl,--wrap=exitlink arg is load-bearing — without it,usage()'sexit(1)would terminate the fuzzer process on first bad input.LLVMFuzzerTestOneInputkeeps external linkage; the scaffold-wide// NOLINTNEXTLINE(misc-use-internal-linkage)pattern is correct for libFuzzer's name-resolved entry-point ABI.- Rebase impact: any upstream sync that touches
core/tools/{yuv_input,cli_parse}.cmust re-run the 60 s smoke per harness on the merged tip; record any new-found crash-* artefact under the matching<target>_known_crashes/dir, not in<target>_corpus/. The__wrap_exitshim infuzz_cli_parse.cis GNU-ld / lld-only; do not assume it works on Apple ld without an-undefined,dynamic_lookupfallback. - Re-test on rebase:
CC=clang CXX=clang++ \
meson setup build-fuzz libvmaf \
--buildtype=debug \
-Db_sanitize=address \
-Db_lundef=false \
-Dfuzz=true \
-Denable_cuda=false -Denable_sycl=false -Denable_vulkan=disabled
ninja -C build-fuzz \
test/fuzz/fuzz_y4m_input \
test/fuzz/fuzz_yuv_input \
test/fuzz/fuzz_cli_parse
./build-fuzz/test/fuzz/fuzz_yuv_input \
-seed=0 -runs=1000 \
core/test/fuzz/yuv_input_corpus/
./build-fuzz/test/fuzz/fuzz_cli_parse \
-seed=0 -runs=1000 \
core/test/fuzz/cli_parse_corpus/
0229 — libFuzzer scaffold for the YUV4MPEG2 parser (ADR-0270)¶
- ADR: ADR-0270
- Touches:
core/test/fuzz/fuzz_y4m_input.c(new)core/test/fuzz/meson.build(new)core/test/fuzz/README.md(new)core/test/fuzz/y4m_input_corpus/*(new — six seeds)core/test/fuzz/y4m_input_known_crashes/*(new — one 411-chroma OOB reproducer; excluded from CI corpus)core/test/meson.build—subdir('fuzz')line.core/meson_options.txt— newoption('fuzz', ...)..github/workflows/fuzz.yml(new — nightly 5-minute job).docs/development/fuzzing.md(new — operator runbook).docs/adr/0270-fuzzing-scaffold.md(new)docs/research/0059-libfuzzer-scaffold-y4m.md(new)docs/state.md— new Open-bug row for the 411-chroma OOB write.CHANGELOG.md— Added entry.- Invariant: the fuzz scaffold is opt-in — every default
meson setupinvocation must continue to skip it. The harness links statically againstcore/tools/{y4m_input,yuv_input,vidinput}.crather thanlibvmaf.soso the public C-API surface stays unchanged. - Rebase impact: the harness re-includes
core/tools/y4m_input.cas a build input. Any upstream Netflix/vmaf change that splits or renames the tool sources (e.g. moves the parser intocore/src/) needs the correspondingmeson.buildsource list update and the harness re-test below. They4m_input_known_crashes/y4m_411_w2_h4_oob_dst.y4mreproducer is the regression gate for the parser fix; do not delete it on upstream sync — if upstream lands the same fix, port the reproducer back intoy4m_input_corpus/as a permanent seed. - Re-test on rebase:
CC=clang CXX=clang++ \
meson setup build-fuzz libvmaf \
--buildtype=debug \
-Db_sanitize=address \
-Db_lundef=false \
-Dfuzz=true \
-Denable_cuda=false -Denable_sycl=false -Denable_vulkan=disabled
ninja -C build-fuzz test/fuzz/fuzz_y4m_input
./build-fuzz/test/fuzz/fuzz_y4m_input \
-max_total_time=60 \
core/test/fuzz/y4m_input_corpus/
# Verify the known-crash reproducer still triggers (until the fix lands):
./build-fuzz/test/fuzz/fuzz_y4m_input \
core/test/fuzz/y4m_input_known_crashes/y4m_411_w2_h4_oob_dst.y4m
0231 — HIP seventh-consumer kernel float_motion_hip (ADR-0273)¶
- ADR: ADR-0273
- Touches:
core/src/feature/hip/float_motion_hip.c(new) — seventh consumer ofcore/src/hip/kernel_template.h. Mirrorscore/src/feature/cuda/float_motion_cuda.ccall-graph-for-call-graph;init/submit/collect/closeinvoke the kernel-template helpers in the same order;flush()callback for tail-frame motion2 emission;motion_force_zeroshort-circuit posture (fex->extractswap withsubmit / collect / flush / closenulled). Submit path intentionally bypassesvmaf_hip_kernel_submit_pre_launch(kernel writes per-WG SAD float partials directly, no atomic, no memset).core/src/feature/hip/float_motion_hip.h(new)core/src/hip/meson.build— new entry inhip_sources.core/src/feature/feature_extractor.c— extern declaration plusfeature_extractor_list[]entry under#if HAVE_HIP.core/test/test_hip_smoke.c— new sub-testtest_float_motion_hip_extractor_registered(also asserts theVMAF_FEATURE_EXTRACTOR_TEMPORALflag bit) and a row intest_table[].docs/adr/0273-hip-seventh-consumer-float-motion.md(new)docs/adr/README.md— index row.docs/backends/hip/overview.md— seventh / eighth consumer note.core/src/hip/AGENTS.md— invariant note.CHANGELOG.md— Added entry (joint with ADR-0274).- Invariant — three-buffer ping-pong +
motion_force_zeroshort-circuit are load-bearing. The state struct carries threeuintptr_tbuffer slots (ref_in,blur[2]) that the runtime PR (T7-10b) will swap for real device-buffer handles matching the CUDA twin'sVmafCudaBuffer *ref_in+VmafCudaBuffer *blur[2]field shape. Themotion_force_zeroshort-circuit (fex->extractswap, kernel-template helpers nulled) must stay aligned with the CUDA twin on every refactor — otherwise the runtime PR's helper-body flip diverges between the two backends. Thesubmit_pre_launchbypass mirrors the CUDA twin; if a future PR adds asubmit_pre_launchcall tofloat_motion_cuda.c's submit path, the HIP twin must follow in the same PR. - Rebase impact: entirely fork-local. New files are HIP-specific. The only upstream-touching edit is
feature_extractor.c, but the change sits inside an existing#if HAVE_HIPblock (ADR-0241); upstream has noHAVE_HIPso no conflict is expected. - Re-test on rebase:
meson setup build libvmaf -Denable_hip=true \
-Denable_cuda=false -Denable_sycl=false -Denable_vulkan=disabled
ninja -C build
meson test -C build test_hip_smoke
0232 — HIP eighth-consumer kernel float_ssim_hip (ADR-0274)¶
- ADR: ADR-0274
- Touches:
core/src/feature/hip/float_ssim_hip.c(new) — eighth consumer ofcore/src/hip/kernel_template.h. Mirrorscore/src/feature/cuda/integer_ssim_cuda.ccall-graph-for-call-graph (the CUDA file registersvmaf_fex_float_ssim_cudadespite itsinteger_filename). First multi-dispatch HIP consumer (chars.n_dispatches_per_frame == 2). Submit path intentionally bypassesvmaf_hip_kernel_submit_pre_launch(kernel writes per-block float partials directly). State struct carries fiveuintptr_tintermediate float buffer slots (h_ref_mu,h_cmp_mu,h_ref_sq,h_cmp_sq,h_refcmp) tracked outside the kernel-template's readback bundle.validate_dims_hipandinit_dims_hiphelpers extracted frominit()to fit thereadability-function-sizebudget.core/src/feature/hip/float_ssim_hip.h(new)core/src/hip/meson.build— new entry inhip_sources.core/src/feature/feature_extractor.c— extern declaration plusfeature_extractor_list[]entry under#if HAVE_HIP.core/test/test_hip_smoke.c— new sub-testtest_float_ssim_hip_extractor_registered(also assertschars.n_dispatches_per_frame == 2) and a row intest_table[].docs/adr/0274-hip-eighth-consumer-float-ssim.md(new)docs/adr/README.md— index row.docs/backends/hip/overview.md— seventh / eighth consumer note (joint).core/src/hip/AGENTS.md— invariant note.CHANGELOG.md— Added entry (joint with ADR-0273).- Invariant — multi-dispatch + five-slot buffer pyramid + v1
scale=1validation are load-bearing. The state struct carries fiveuintptr_tintermediate float buffer slots that the runtime PR (T7-10b) will swap for real device-buffer handles matching the CUDA twin'sVmafCudaBuffer *h_*field shape — any drift in the CUDA twin's slot count requires a paired update here. Thechars.n_dispatches_per_frame == 2characteristic is asserted in the smoke test; do not silently lower it. The v1scale=1-EINVALvalidation surface (invalidate_dims_hip) must stay aligned with the CUDA twin'scompute_scale/vmaf_logchain. The HIP twin'svalidate_dims_hip/init_dims_hipextraction is intentional for the function-size budget; do not re-inline without verifying the budget still passes. - Rebase impact: entirely fork-local; same posture as ADR-0273.
- Re-test on rebase:
meson setup build libvmaf -Denable_hip=true \
-Denable_cuda=false -Denable_sycl=false -Denable_vulkan=disabled
ninja -C build
meson test -C build test_hip_smoke
0229 — vmaf_tiny_v3 + vmaf_tiny_v4 dynamic-PTQ int8 sidecars (ADR-0275)¶
0278 — vmaf-tune libaom-av1 codec adapter (2026-05-03)¶
0228 — vmaf-tune libx265 codec adapter (ADR-0288)¶
0280 — vmaf-tune NVENC codec adapters (ADR-0290)¶
- Touches:
tools/vmaf-tune/src/vmaftune/codec_adapters/{h264_nvenc,hevc_nvenc,av1_nvenc,_nvenc_common}.py(new). Wholly fork-local — no upstream Netflix/vmaf overlap.tools/vmaf-tune/src/vmaftune/codec_adapters/__init__.py— registry expanded.tools/vmaf-tune/tests/test_codec_adapter_nvenc.py(new).tools/vmaf-tune/tests/test_corpus.py— Phase-A registry assertion updated.tools/vmaf-tune/AGENTS.md— invariant note expanded.docs/usage/vmaf-tune.md— "Hardware encoders (NVENC)" section.docs/adr/0290-vmaf-tune-nvenc-adapters.md(new) +docs/adr/README.mdindex row.docs/research/0065-vmaf-tune-nvenc-adapters.md(new).CHANGELOG.md— Added entry.- Invariant:
known_codecs()returns the four-codec tuple("av1_nvenc", "h264_nvenc", "hevc_nvenc", "libx264"); the mnemonic preset map (ultrafast/superfast/veryfast→p1,faster→p2,fast→p3,medium→p4,slow→p5,slower→p6,slowest/placebo→p7) is the canonical cross-codec preset alignment that downstream Phase B/C consumers assume. The CQ window is the hardware-permitted[0, 51]; the Phase A informative window is[15, 40]. - Rebase impact: zero —
tools/vmaf-tune/is wholly fork-local and has no upstream Netflix/vmaf path overlap. - Re-test on rebase:
0227 — ffmpeg-patches/ series re-verified against n8.1 (2026-05-03)¶
- Touches:
tools/vmaf-tune/src/vmaftune/codec_adapters/x265.py(new),tools/vmaf-tune/src/vmaftune/codec_adapters/__init__.py(registry add),tools/vmaf-tune/src/vmaftune/encode.py(parse_versions(stderr, encoder=…)gains a per-codec branch),tools/vmaf-tune/src/vmaftune/cli.py(help-text wording only),tools/vmaf-tune/tests/test_codec_adapter_x265.py(new),tools/vmaf-tune/tests/test_corpus.py(membership-based codec list assertion). - Invariant: the codec-adapter contract documented in
tools/vmaf-tune/AGENTS.md(multi-codec from day one; the search loop never branches on codec identity). Theparse_versionssignature is still backward-compatible —encoderdefaults tolibx264so callers from before this PR keep working. - Upstream source: fork-local.
tools/vmaf-tune/is fork-only; upstream Netflix/vmaf does not ship encode automation. - On upstream sync: zero interaction. Confirm the
_index_fragments/_order.txtrow for0288-vmaf-tune-codec-adapter-x265remains present after any cross-merge. - Re-test on rebase:
0227 — ffmpeg-patches/ series re-verified against n8.1 (2026-05-03)¶
- Touches:
tools/vmaf-tune/src/vmaftune/codec_adapters/libaom.py(new),tools/vmaf-tune/src/vmaftune/codec_adapters/__init__.py(registry row + import),tools/vmaf-tune/tests/test_corpus.py(membership assertion relaxed from== ("libx264",)to"libx264" in known_codecs()),tools/vmaf-tune/tests/test_codec_adapter_libaom.py(new),tools/vmaf-tune/AGENTS.md(preset-vocabulary invariant). - Invariant: the cross-codec preset vocabulary (
placebo, slowest, slower, slow, medium, fast, faster, veryfast, superfast, ultrafast) is shared across AV1-family adapters so one--presetaxis covers x264 / x265 / svtav1 / libaom-av1. Each adapter maps the human name onto its codec-specific knob; do not introduce per-adapter preset names. - Upstream source: fork-local.
tools/vmaf-tune/is the fork-introduced quality-aware encode automation harness (ADR-0237); it has no upstream Netflix/vmaf counterpart. - On upstream sync: zero interaction with
upstream/master. Self-contained intools/vmaf-tune/anddocs/. - Re-test on rebase:
0227 — ffmpeg-patches/ series re-verified against n8.1 (2026-05-03)¶
- ADR: ADR-0275
- Touches:
model/tiny/vmaf_tiny_v3.int8.onnx(new, 4 267 B)model/tiny/vmaf_tiny_v4.int8.onnx(new, 7 769 B)model/tiny/registry.json— newvmaf_tiny_v3andvmaf_tiny_v4rows withquant_mode,int8_sha256,quant_accuracy_budget_plccfields.model/tiny/vmaf_tiny_v3.json,model/tiny/vmaf_tiny_v4.json— same fields mirrored into the per-model sidecars.docs/ai/models/vmaf_tiny_v3.md,docs/ai/models/vmaf_tiny_v4.md— new "Quantisation" sections.docs/adr/0275-vmaf-tiny-v3-v4-ptq.md(new) and ADR index row.CHANGELOG.md— Added entry.- Invariant:
python ai/scripts/measure_quant_drop.py --allreports[PASS]for bothvmaf_tiny_v3(drop ≤ 0.001 on Netflix features) andvmaf_tiny_v4(drop ≤ 0.001), inside the 0.01 per-model budget. The runtime redirect from ADR-0174 picks the.int8.onnxsibling when an operator's registry overlay declaresquant_mode: dynamic. - Rebase impact: entirely fork-local — neither v3 nor v4 nor the dynamic-PTQ harness exists upstream. The new int8 ONNX bytes ship as committed binaries (mirroring
learned_filter_v1andnr_metric_v1); they are well below the few-MB external-data threshold and don't require the sigstore +.onnx.datapattern. - Re-test on rebase:
```bash python ai/scripts/validate_model_registry.py python ai/scripts/measure_quant_drop.py --all
0229 — NVIDIA-Vulkan ciede2000 places=4 fork debt root-cause (ADR-0273)¶
- Touched files: docs-only.
docs/adr/0273-...precision-gap.md(new) +_index_fragments/row +_order.txtappend.docs/research/0055-ciede-vulkan-nvidia-f32-f64-root-cause.md(new) +docs/research/README.mdindex row.docs/state.md— Open-bugs rowT-VK-CIEDE-F32-F64.docs/backends/vulkan/overview.md— NVIDIA-hardware caveat.changelog.d/changed/ciede-vulkan-nvidia-f32-f64-precision-gap.md(new).core/src/vulkan/AGENTS.md— invariant cross-link.- Invariant: the ciede.comp shader's f32 precision contract is load-bearing — promoting to f64 would silently change scores on every Vulkan device that supports
shaderFloat64and create a per-device-feature-bit divergence (RTX 4090 has it; many consumer GPUs don't). The CPUciede.c::get_lab_colordoing its colour-space chain indoubleis upstream Netflix behaviour and must not be narrowed to f32 to "fix" the GPU gap (would change Netflix golden ground truth). The 5/48 NVIDIA places=4 mismatch on the highest-ΔE frames is expected and documented; do not attempt to "fix" it without re-reading ADR-0273 first. - Rebase impact: zero — docs-only. The CPU and shader sources this ADR analyses are unchanged by this PR. If a future upstream rebase touches
ciede.c::get_lab_color(thedoublechain) the ADR's reasoning still holds; if upstream changes the CPU reference's precision posture, ADR-0273 needs aStatus: Supersededentry. - Re-test on rebase: a manual NVIDIA-hardware run if available:
```bash cd libvmaf && meson setup build \ -Denable_vulkan=enabled -Denable_cuda=false && ninja -C build cd .. python3 scripts/ci/cross_backend_vif_diff.py \ --vmaf-binary $PWD/core/build/tools/vmaf \ --reference testdata/ref_576x324_48f.yuv \ --distorted testdata/dis_576x324_48f.yuv \ --width 576 --height 324 \ --feature ciede --backend vulkan --device 0 --places 4 # Expected post-PR-346 (when merged): 5/48 mismatches at 1.78× threshold. # Expected pre-PR-346 (current master): 42/48 mismatches at higher ratio. # If the count drops below 5/48 on NVIDIA, ADR-0273 should record the # delta and consider closing T-VK-CIEDE-F32-F64.
0229 — tools/vmaf-tune fast Phase A.5 scaffold (ADR-0276)¶
- Touches:
tools/vmaf-tune/src/vmaftune/fast.py(new),tools/vmaf-tune/src/vmaftune/cli.py(newfastsubcommand branch),tools/vmaf-tune/pyproject.toml(new[fast]extra),tools/vmaf-tune/tests/test_fast.py(new),tools/vmaf-tune/AGENTS.md(new invariants),docs/usage/vmaf-tune.md(new "Phase A.5" section),docs/adr/0276-vmaf-tune-fast-path.md(new ADR),docs/research/0060-vmaf-tune-fast-path.md(new digest). - Invariant: the
fastsubcommand is opt-in and never automatically replaces the Phase A grid path. The slow grid is the ground-truth corpus generator (ADR-0237 contract); fast-path is for the recommendation use case only. Optuna is a lazy-imported optional dep gated behind the[fast]extra — importing it at module scope outsidefast.py(or its tests) breaks the zero-dep core install. - Rebase impact: entirely fork-local; the tool sits under
tools/vmaf-tune/which is fork-added, and no upstream files are touched. Upstream Netflix/vmaf has no analogous surface. - Re-test on rebase:
pip install -e 'tools/vmaf-tune[fast]'
pytest tools/vmaf-tune/tests/test_fast.py -v
vmaf-tune fast --smoke --target-vmaf 92
0229 — vmaf-tune recommend subcommand (ADR-0237 Phase B-lite)¶
- Touches:
tools/vmaf-tune/src/vmaftune/recommend.py(new). Wholly fork-local — no upstream Netflix/vmaf path overlap.tools/vmaf-tune/src/vmaftune/cli.py— addsrecommendsubparser;corpussubcommand untouched.tools/vmaf-tune/tests/test_recommend.py(new). 13-case smoke suite, mocks all binaries; runs in <100 ms.docs/usage/vmaf-tune.md— adds## recommendsection.- Invariant:
recommendconsumes the existingCORPUS_ROW_KEYSschema unchanged —vmaf_score,bitrate_kbps,crf,preset,encoder,exit_status. No schema bump. If a future PR bumpsSCHEMA_VERSION, both thecorpuswriter and therecommendreader must be updated in lockstep; tests assert this viatest_corpus_row_keys_match_init_contract. - Rebase impact: zero —
tools/vmaf-tune/is wholly fork-local; no upstream surface touches it. - Re-test on rebase:
0228 — integer_ms_ssim_cuda.c joins drain_batch (T-GPU-OPT-2 / ADR-0271)¶
- Touches:
core/src/feature/cuda/integer_ms_ssim_cuda.c. No upstream Netflix/vmaf changes expected here — the file is fork-added (CUDA twin of the upstream-portms_ssim_score.cu) and the surface this PR redrew (per-scalel_partials[i]/c_partials[i]/s_partials[i]arrays + the per-scaleh_l_partials[i]/h_c_partials[i]/h_s_partials[i]pinned host shadows + thesubmit()<→collect()work redistribution + thecuEventRecord(s->lc.finished, s->lc.str)+vmaf_cuda_drain_batch_register(&s->lc)tail) is also entirely fork-local. - Invariant: the engine-scope drain-batch contract from ADR-0271 / drain_batch.h. The kernel-launch order on
s->lc.strmust stay stable:decimate (× 4)then for each scalei ∈ 0..4horiz⇒vert_lcs⇒ DtoH(l_partials[i]) ⇒ DtoH(c_partials[i]) ⇒ DtoH(s_partials[i])thencuEventRecord(s->lc.finished, s->lc.str)thenvmaf_cuda_drain_batch_register(&s->lc). Same-stream ordering is what makes the shared SSIM intermediates (h_ref_mu,h_cmp_mu,h_ref_sq,h_cmp_sq,h_refcmp`) safe across scales without explicit sync — any change that parallelises the per-scale work onto multiple streams breaks bit-exactness unless per-scale intermediates are also added. - On upstream sync: zero interaction (the file is fork-added). If a future upstream PR adds an
integer_ms_ssim_cuda.cof its own, the merger must reconcile the per-scale partials topology + the drain_batch tail with whatever the new upstream shape brings. - Re-test on rebase:
cd libvmaf && meson setup build -Denable_cuda=false -Denable_sycl=false
ninja -C build # confirms the CPU build still links cleanly
# If the dev host has a working nvcc / host-compiler pair:
meson setup build_cuda -Denable_cuda=true -Denable_sycl=false
ninja -C build_cuda src/liblibvmaf_feature.a.p/feature_cuda_integer_ms_ssim_cuda.c.o
# Netflix CPU golden gate (CPU is the bit-exactness ground truth):
make test-netflix-golden
# Cross-backend parity (places=4 gate, ADR-0214):
/cross-backend-diff
0277 — ffmpeg-patches refresh against n8.1 — 2026-05-04 (ADR-0277)¶
- Touches:
ffmpeg-patches/is unchanged (no content drift). Doc-only entries land in: docs/adr/0277-ffmpeg-patches-refresh-2026-05-04.md— new ADR.docs/adr/_index_fragments/0277-ffmpeg-patches-refresh-2026-05-04.md— index row.docs/adr/_index_fragments/_order.txt— manifest append.changelog.d/changed/ffmpeg-patches-refresh-2026-05-04.md— Changed entry.- This file — this entry.
- Invariant:
ffmpeg-patches/series.txtorder is load-bearing — patches0002…0006build on each other and only apply cleanly cumulatively. The verification gate is a series replay, not a per-patchgit apply --check(per ADR-0118 + CLAUDE.md §12 r14). - On upstream sync: zero interaction. Netflix/vmaf has no
ffmpeg-patches/tree; this is a fork-local integration surface. - Re-test on rebase (also: re-replay procedure for the next refresh):
# Clone pristine n8.1
git -C /tmp clone --depth 1 --branch n8.1 \
https://github.com/FFmpeg/FFmpeg.git ff-replay-$(date +%F)
cd /tmp/ff-replay-$(date +%F)
git switch -c refresh-$(date +%F)
git config user.email refresh@local && git config user.name "Refresh Bot"
# Replay the series cumulatively
for p in /path/to/vmaf/ffmpeg-patches/000*-*.patch; do
git am --3way "$p" || break
done
# Regenerate and compare to in-tree
mkdir -p /tmp/ff-regen-$(date +%F)
git format-patch n8.1.. -o /tmp/ff-regen-$(date +%F)/
# Diff old vs new excluding pure format-patch noise
for i in 1 2 3 4 5 6; do
orig=$(ls /path/to/vmaf/ffmpeg-patches/000${i}-*.patch)
regen=$(ls /tmp/ff-regen-$(date +%F)/000${i}-*.patch)
diff -u \
<(grep -v "^From [0-9a-f]\|^Date:\|^index " "$orig") \
<(grep -v "^From [0-9a-f]\|^Date:\|^index " "$regen") \
| head -40
done
If only stylistic diffs surface (PATCH N/M numbering, MIME headers, hunk-context counts, hunk offset shifts against cumulative state), keep originals — record a no-drift refresh ADR. If real content drift surfaces, regenerate and ship the refresh PR with the regenerated patches plus a content-summary ADR.
End-to-end vf_libvmaf smoke is best run from CI (ffmpeg-integration.yml) against an installed libvmaf prefix — the meson-uninstalled .pc does not satisfy FFmpeg's #include <libvmaf.h> probe (the headers live under libvmaf/libvmaf.h only; the system-installed .pc carries an extra -I${includedir}/libvmaf shortcut that the uninstalled .pc omits).
0229 — T7-5 NOLINT-sweep closeout (ADR-0278)¶
- Touched files:
core/src/feature/integer_adm.c(1 NOLINT cite, line ~988adm_decouple_s123— upstream-mirror Netflix966be8d5).core/src/feature/cuda/ssimulacra2_cuda.c(3 NOLINT cites:ss2c_picture_to_linear_rgb,ss2c_host_combine,ss2c_run_scale_gpu/extract_fex_cuda).core/src/feature/vulkan/ssimulacra2_vulkan.c(3 NOLINT cites:ss2v_setup_gaussian,ss2v_picture_to_linear_rgb,ss2v_run_scale).core/src/feature/vulkan/cambi_vulkan.c(1 NOLINT cite:cambi_vk_extract).core/src/feature/sycl/integer_adm_sycl.cpp(6 cites, SYCL kernel-launch entries).core/src/feature/sycl/integer_motion_sycl.cpp(2 cites).core/src/feature/sycl/integer_vif_sycl.cpp(4 cites).core/tools/vmaf.c(3 cites:copy_picture_data,init_gpu_backends,main).- Invariant: zero behavioural change. Edits are inside comment blocks — appended
(ADR-0141 §2 ... load-bearing invariant; T7-5 sweep closeout — ADR-0278)to existing prose justifications. No function bodies split. The 12 SYCL sites share an identical justification string verbatim; preserving the byte-for-byte duplicate is the load-bearing documentation pattern (grep-able across the SYCL TUs). - On upstream sync: minimal interaction. The cite-only edits live inside comment blocks above the function signatures; rebases will surface them as touched lines but the function bodies are unchanged. For
integer_adm.c's upstream-mirror block (Netflix966be8d5), the comment edit at line 984–991 is cosmetic — keep the fork's version on conflict (it merely names the ADR; the underlying prose is unchanged). - Re-test on rebase:
```bash # 1. Programmatic audit must report 0 missing citations python3 - <<'PY' import re, os paths = [os.path.join(r, f) for r, _, fs in os.walk('libvmaf/src') for f in fs if f.endswith(('.c','.cpp','.h'))] paths.append('core/tools/vmaf.c') miss = total = 0 for p in paths: with open(p) as fh: ls = fh.readlines() for i, line in enumerate(ls): if 'NOLINT' in line and 'readability-function-size' in line and 'NOLINTEND' not in line: total += 1 ctx = [line]; j = i - 1 while j >= 0 and j > i - 14: s = ls[j].strip() if not s: break if s.startswith(('//','/','')): ctx.insert(0, ls[j]); j -= 1 else: break buf = ''.join(ctx) if 'ADR-' not in buf and not re.search(r'[Rr]esearch-?\d', buf): miss += 1 print(f"sites={total} missing={miss}") PY
# 2. Build + Netflix golden gate meson setup build -Denable_cuda=false -Denable_sycl=false ninja -C build make test-netflix-golden
0231 — vmaf-tune score path decodes mp4 -> raw YUV¶
- Touches:
tools/vmaf-tune/src/vmaftune/score.py(new_decode_to_raw_yuv+_needs_decodehelpers,run_scoreshells out to ffmpeg whenreq.distorted.suffix not in {.yuv, .y4m});tools/vmaf-tune/tests/test_corpus.py(3 new regression tests + the smoke-end-to-end mock now also stubs the ffmpeg decode call). - Invariant: the decode-back is the contract the libvmaf CLI imposes — mp4/webm/etc.
--distortedis silently rejected as raw-yuv with the wrong byte count, surfacing asexit_status=234. Future encoder adapters that emit non-raw containers inherit this decode automatically. Do not "optimise" the temp YUV away without first migrating the corpus pipeline to theffmpeg+libvmaffilter (which can pipe an mp4 stream in directly). - On upstream sync: zero interaction.
vmaf-tuneis fork-only tooling; upstream Netflix/vmaf has no analogue. - Re-test on rebase:
```bash cd tools/vmaf-tune && python3 -m pytest tests/ # plus an end-to-end smoke (needs a real raw YUV + ffmpeg + vmaf): ./vmaf-tune corpus --source /path/to/ref.yuv --width 1920 \ --height 1080 --pix-fmt yuv420p --framerate 25 --duration 6 \ --encoder libx264 --preset medium --crf 23 \ --output /tmp/smoke.jsonl --no-source-hash # expect: vmaf_score is a real number, not NaN.
0232 — CUDA build pins nvcc --std c++20¶
- Touches:
core/src/meson.buildline 686 (cuda_flags = [...]). - Invariant: nvcc 12.x clamps host C++ at C++17 by default; 13.x accepts up to C++20. Bumping the host stdlib past nvcc's default (any gcc >= 16, libstdc++ ships C++23 features) breaks the host-side parse in
<type_traits>/<bits/utility.h>. Forcing--std c++20on CUDA 13+ keeps the host headers parseable. Do not drop this flag without first checking the host gcc version against nvcc's default. - On upstream sync: zero interaction. Netflix/vmaf doesn't ship the
cuda_flagslist shape we use (their CUDA build is the original pre-fork pattern); a sync that touchescore/src/meson.buildaround theis_cuda_enabledbranch should keep the--std c++20injection. - Re-test on rebase:
meson setup core/build-cuda -Denable_cuda=true \
-Denable_sycl=false -Denable_vulkan=disabled
ninja -C core/build-cuda
# smoke
./core/build-cuda/tools/vmaf --gpumask=0 --no_sycl --no_vulkan \
-r .workingdir2/netflix/ref/BigBuckBunny_25fps.yuv \
-d .workingdir2/netflix/ref/BigBuckBunny_25fps.yuv \
-w 1920 -h 1080 -p 420 -b 8
0233 — CUDA motion flush_fex_cuda idempotency guard¶
- Touches:
core/src/feature/cuda/integer_motion_cuda.c— factored anappend_if_unwrittenhelper and routed the two motion2 / motion3 final-frame writes through it. - Invariant: under T-GPU-OPT-1 (PR #312 / ADR-0242), the pending-collect inside
flush_context_cudamay already have writtenmotion2_score[s->index]/motion3_score[s->index]beforeflush_fex_cudaruns. Any future motion-cuda flush logic that emits the same (feature, index) pair must keep this idempotency contract orflush_context_cudawill mis-surface as "context could not be synchronized". - On upstream sync: the bug only exists because the fork's
flush_context_cudaruns the pending-collect before the per-extractor flush. Netflix/vmaf upstream doesn't have the T-GPU-OPT-1 drain pattern, so the pre-#312 code path didn't duplicate-write. If Netflix lands a similar pattern, the fix shape mirrors what's done here. - Re-test on rebase:
ninja -C core/build-cuda
./core/build-cuda/tools/vmaf --gpumask=0 --no_sycl --no_vulkan \
-r python/test/resource/yuv/src01_hrc00_576x324.yuv \
-d python/test/resource/yuv/src01_hrc01_576x324.yuv \
-w 576 -h 324 -p 420 -b 8 \
--model path=model/vmaf_v0.6.1.json --threads 1 -q \
--output /tmp/cuda.json --json
# Expect: clean run, no "cannot be overwritten" warning,
# no "problem flushing context" error.
0234 — hw_encoder_corpus.py Phase A real-corpus runner¶
- Touches: new
scripts/dev/hw_encoder_corpus.py(no existing caller; opt-in tooling). Output landing inruns/phase_a/is gitignored — rerun the script to reproduce.docs/development/intel-arc-vaapi-driver-priority.md. Output landing inruns/phase_a/is gitignored — rerun the script to reproduce. stratified sample, 58 KiB). - Invariant: the script's QSV path forces
env['LIBVA_DRIVER_NAME']='iHD'(set by the calling shell, not inside the script) when targeting/dev/dri/renderD129on a multi-card host that has NVIDIA's libva-driver-nvidia shim installed. Without that, libva picks up NVIDIA's NVDEC-VAAPI translation and the MFX session handshake fails with -9. See the companion doc for the failure mode + fix. - On upstream sync: zero interaction. The script lives under
scripts/dev/(fork-only); upstream Netflix/vmaf has no comparable Phase A corpus tooling. - Re-test on rebase:
python3 scripts/dev/hw_encoder_corpus.py \
--vmaf-bin core/build-cuda/tools/vmaf \
--source .workingdir2/netflix/ref/BigBuckBunny_25fps.yuv \
--width 1920 --height 1080 --pix-fmt yuv420p --framerate 25 \
--encoder h264_nvenc --cq 25 \
--out /tmp/smoke.jsonl
# Expect: 1 cell × ~150 frames, per-frame canonical-6 + vmaf,
# encoder=h264_nvenc, cq=25.
0235 — fr_regressor_v2 ENCODER_VOCAB v2 (hw codec extension)¶
- Touches:
ai/scripts/train_fr_regressor_v2.py—ENCODER_VOCABgains 6 hw-codec entries (3 NVENC + 3 QSV);ENCODER_VOCAB_VERSIONbumps 1 -> 2;PRESET_ORDINALgains 6 sub-tables forp1..p7(NVENC) and the libx264-aligned QSV preset family. - Invariant: vocab order is load-bearing — index of every entry is baked into trained model graphs as a one-hot column position. New entries MUST be appended (never inserted into the middle), and the
unknownsentinel MUST stay last (UNKNOWN_ENCODER_INDEX = N - 1). BumpingENCODER_VOCAB_VERSIONsignals that any v1-graph ONNX needs re-export against v2 before consuming v2 training rows. - On upstream sync: zero interaction.
train_fr_regressor_v2.pyis fork-only (Phase B prereq, ADR-0237 / ADR-0272). - Re-test on rebase:
python3 ai/scripts/train_fr_regressor_v2.py --corpus <jsonl> --epochs 200 --no-export— expect PLCC > 0.95 on a multi-codec corpus.
0276 — vmaf_tiny_v5 corpus-expansion probe (ADR-0287) — defer¶
- What changed: research-only addition. New scripts under
ai/scripts/(fetch_youtube_ugc_subset.py,extract_ugc_features.py,train_vmaf_tiny_v5.py,eval_loso_vmaf_tiny_v5.py), new ADRdocs/adr/0276-*.md, new research digestdocs/research/0057-*.md, and one CHANGELOG entry. No new ONNX artefact undermodel/tiny/, no registry change, no public C-API / CLI / meson_options change. The probe trained an architecturally identical mlp_small on a 5-corpus parquet (4-corpus + 27 000 UGC rows); the 1-σ ship gate did not clear (Δ PLCC = +0.00005), so the exporter that the prior agent had drafted (export_vmaf_tiny_v5.py) was discarded before the commit. - Upstream source: fork-local. Netflix/vmaf has no tiny-AI corpus-expansion surface; nothing on the upstream side touches these files.
- On upstream sync: zero interaction. The v5 surface lives entirely under
ai/scripts/+docs/adr/+docs/research/, all of which are fork-introduced trees. The shipped v2 model (model/tiny/vmaf_tiny_v2.onnx) and its registry row are untouched. - Re-test on rebase:
# No code under test on rebase — purely research artefacts.
# If revisiting the corpus expansion, the reproducer is in the
# research digest:
python3 ai/scripts/fetch_youtube_ugc_subset.py \
--out-dir .workingdir2/ugc/download \
--n-stems 30 \
--manifest .workingdir2/ugc/manifest.json
python3 ai/scripts/extract_ugc_features.py \
--manifest .workingdir2/ugc/manifest.json \
--yuv-dir .workingdir2/ugc/yuv \
--vmaf-bin build-cpu/tools/vmaf \
--out-parquet runs/full_features_ugc.parquet \
--max-height 360 --max-frames 300 --threads 8
python3 ai/scripts/eval_loso_vmaf_tiny_v5.py \
--parquet-base runs/full_features_4corpus.parquet \
--parquet-extra runs/full_features_ugc.parquet \
--out-json runs/vmaf_tiny_v5_loso_metrics.json
0227 — vmaf-tune Intel QSV codec adapters (ADR-0281)¶
- What changed: fork-local additions under
tools/vmaf-tune/src/vmaftune/codec_adapters/—_qsv_common.py,h264_qsv.py,hevc_qsv.py,av1_qsv.py, plus registry rows incodec_adapters/__init__.pyand a new test filetools/vmaf-tune/tests/test_codec_adapter_qsv.py. Doc updates:docs/usage/vmaf-tune.md(Hardware encoders section),docs/adr/0281-vmaf-tune-qsv-adapters.md,docs/research/0066-vmaf-tune-qsv-adapters.md,tools/vmaf-tune/AGENTS.md,CHANGELOG.md. - Upstream source: fork-local.
tools/vmaf-tune/is fork-introduced under ADR-0237; Netflix/vmaf has no corresponding tree. - On upstream sync: zero interaction. Upstream cannot conflict with this PR's paths.
- Invariant: the registry exposes exactly four codecs (
av1_qsv,h264_qsv,hevc_qsv,libx264— alphabetical), each adapter validates its(preset, quality)pair, and the QSV preset vocabulary is the seven x264-style names (veryslow…veryfast, noultrafast/superfast). The encode pipeline (encode.py) remains x264-CRF-tied and will be widened in a separate PR — the QSV adapters are inert until then. Future codec families that share parameter shape (NVENC, AMF) follow the same_<family>_common.py+ N thin adapters pattern. - Re-test on rebase:
0229 — vmaf-tune libvvenc + NN-VC codec adapter (ADR-0285)¶
- Touches:
tools/vmaf-tune/src/vmaftune/codec_adapters/vvenc.py(new fork-only file),tools/vmaf-tune/src/vmaftune/codec_adapters/__init__.py(registry edit, fork-only),tools/vmaf-tune/tests/test_codec_adapter_vvenc.py(new),tools/vmaf-tune/tests/test_corpus.py(relaxes theknown_codecs() == ("libx264",)assertion to"libx264" in known_codecs()since the registry now spans multiple codecs). - Invariant: the codec-adapter registry is fork-introduced (Phase A of ADR-0237) and lives entirely outside the upstream Netflix tree, so
tools/vmaf-tune/does not touch upstream paths. The only rebase-sensitive surface is theCORPUS_ROW_KEYSschema insrc/vmaftune/__init__.py(per the Phase A invariant intools/vmaf-tune/AGENTS.md); this PR adds the adapter without changing the schema. - Upstream interaction: none.
tools/vmaf-tune/is not in Netflix/vmaf upstream. - Re-test on rebase:
- Status update 2026-05-09: the original
nnvc_intratoggle was removed (it emitted a fabricatedIntraNNkey that does not exist in any released VVenC). Replaced with a curated 9-knob real-VVenC 1.14.0 tuning surface (PerceptQPA,InternalBitDepth,Tier,Tiles,MaxParallelFrames,RPR,SAO,ALF,CCALF). Defaults preserve the bit-exact Phase A grid baseline.adapter_versionbumped to"2"so cache keys invalidate. See ADR-0285 §"Status update 2026-05-09".no rebase impact: REASON(fork-local file, no upstream-tree touch).
0228 — vmaf-tune Phase D scaffold (ADR-0276)¶
- Touches:
tools/vmaf-tune/src/vmaftune/per_shot.py,tools/vmaf-tune/src/vmaftune/cli.py,tools/vmaf-tune/tests/test_per_shot.py,docs/usage/vmaf-tune.md,docs/adr/0276-vmaf-tune-phase-d-per-shot.md. - Invariant: scaffold-only. The module relies on a stable predicate signature
(shot, target_vmaf, encoder) -> (crf, predicted_vmaf)that Phase B's bisect (PR #347) drops into later.Shotranges are half-open[start_frame, end_frame)even though the C-sidevmaf-perShotJSON/CSV sidecar uses an inclusiveend_frame— normalisation happens at the parse boundary in_parse_per_shot_json/parse_per_shot_csv.vmaf-perShotschema lives indocs/usage/vmaf-perShot.mdand is fork-local (ADR-0222), so upstream cannot drift it; the only rebase risk is fork-internal renames. - Upstream source: entirely fork-local.
tools/vmaf-tune/is fork-introduced (ADR-0237). Netflix/vmaf upstream has no encode-automation surface. - On upstream sync: zero interaction expected. No file in this PR overlaps an upstream-mirrored path.
- Re-test on rebase:
python -m pytest tools/vmaf-tune/tests/test_per_shot.py -q
python tools/vmaf-tune/vmaf-tune tune-per-shot --help
0229 — vmaf-tune SVT-AV1 codec adapter (ADR-0278)¶
- Touches:
tools/vmaf-tune/src/vmaftune/codec_adapters/svtav1.py(new),tools/vmaf-tune/src/vmaftune/codec_adapters/__init__.py(registry),tools/vmaf-tune/src/vmaftune/encode.py(parse_versionsextended for the SVT-AV1 banner pattern),tools/vmaf-tune/src/vmaftune/corpus.py(optionalffmpeg_preset_tokenhook). - Invariant:
PRESET_NAME_TO_INTis closed and order-stable; the integer values are baked into corpus rows that downstreamfr_regressor_v2(ADR-0235) trains on. Reordering or rewriting the table silently changes the integer SVT-AV1 receives. The codec key"libsvtav1"matchesCODEC_VOCAB[2]inai/src/vmaf_train/codec.py— keep them aligned on any rename. - Upstream source: fork-local.
tools/vmaf-tune/is a fork-introduced tree (see entry 0227 — Phase A scaffold). No Netflix/vmaf upstream interaction. - On upstream sync: zero interaction. Lives entirely under the fork-local
tools/vmaf-tune/tree. - Re-test on rebase:
0230 — fr_regressor_v2 PROD ship (ADR-0352)¶
- ADR: ADR-0352
- Touches:
model/tiny/fr_regressor_v2.onnx(binary, refreshed),model/tiny/fr_regressor_v2.json(sidecar, sha256 + metrics),model/tiny/registry.json(smoke flag flip, sha256 update),runs/phase_a/full_grid/per_frame_canonical6.jsonl(training corpus — fork-local artefact underruns/), companion docs. - Re-test recipe: see Research-0068 §Reproducer. Ship gate is LOSO PLCC ≥ 0.95 on the per-source folds; current run reports 0.9681 ± 0.0207.
- Rebase invariant: the per-frame canonical-6 corpus must be rebuilt from
runs/phase_a/{nvenc,qsv}_pf.jsonl(PR #392) before any retrain; do not re-train against the cell-onlycomprehensive.jsonl(it lacks the per-frame features and produces PLCC ≈ 0.7 — the smoke baseline). - No upstream interaction:
fr_regressor_v2is fork-local (ADR-0272).
0229 — vmaf-tune Phase E ladder generator (ADR-0295)¶
- ADR: ADR-0295
- Touches: entirely fork-local under
tools/vmaf-tune/. New moduletools/vmaf-tune/src/vmaftune/ladder.py, new test filetools/vmaf-tune/tests/test_ladder.py, two new subcommand blocks intools/vmaf-tune/src/vmaftune/cli.py. No upstream-shared paths touched. - Invariant:
vmaftune.ladder.convex_hullreturns a strictly monotonic Pareto frontier (both bitrate and vmaf monotonically increasing);select_kneesreturns exactlymin(n, len(hull))rungs in ascending bitrate order;emit_manifest("hls")produces one#EXT-X-STREAM-INFper rung with monotonically-increasingBANDWIDTH=values. The default_default_sampleris intentionallyNotImplementedError— production callers must inject a Phase B bisect-driven sampler. Phase B integration PR (gated on PR #347) swaps the default; the test suite continues to inject a synthetic stub. - Rebase impact: none — fork-local Python tool; upstream Netflix/vmaf does not ship a
tools/vmaf-tune/tree. - Re-test on rebase:
0229 — fr_regressor_v2 probabilistic head scaffold (ADR-0279)¶
- Touches:
ai/scripts/train_fr_regressor_v2_ensemble.py(new — fork-local).ai/scripts/eval_probabilistic_proxy.py(new — fork-local).model/tiny/fr_regressor_v2_ensemble_v1*.onnx,fr_regressor_v2_ensemble_v1.json(new artefacts; smoke probes).model/tiny/registry.json— five newkind: "fr"rows (fr_regressor_v2_ensemble_v1_seed{0..4}); existing entries untouched.ai/AGENTS.md— new "fr_regressor_v2_ensemble_v1 — probabilistic head" section pinning the per-member ONNX I/O contract, manifest-as-runtime-entry-point invariant, ensemble-size pin, confidence-rule one-of, codec-vocab parity, and smoke-artefact posture.docs/ai/models/fr_regressor_v2_probabilistic.md(new model card).docs/research/0067-fr-regressor-v2-probabilistic.md(new audit digest).docs/adr/0279-fr-regressor-v2-probabilistic.md(new ADR; Proposed). Index row appended todocs/adr/README.md.CHANGELOG.md—### Addedrow under "Unreleased — lusoris fork".- Invariant: the per-member ONNX I/O contract (two inputs:
features [N, 6]standardised +codec_onehot [N, NUM_CODECS]; one outputscore [N]) and the manifest'sconfidencerule (one-of"ensemble"/"ensemble+conformal") are the C-side adapter's load-bearing contract. Per-member ensembles are stockFRRegressor(num_codecs=NUM_CODECS)calls — flipping to a v1-shaped single-input graph silently invalidates the manifest.CODEC_VOCABparity withai/src/vmaf_train/codec.pyis required. - On upstream sync: zero interaction expected. Wholly fork-local; no upstream Netflix/vmaf path overlap. The
ai/package is fork-introduced (see ADR-0021, ADR-0036) — upstream has no probabilistic-regressor surface. If upstream ever ships its ownfr_regressor_v2variant, do NOT merge — register both ids side-by-side. - Re-test on rebase:
python ai/scripts/train_fr_regressor_v2_ensemble.py --smoke
python ai/scripts/eval_probabilistic_proxy.py --smoke
python ai/scripts/validate_model_registry.py
0287 — vmaf-tune saliency-aware ROI tuning (ADR-0293)¶
- Touches:
tools/vmaf-tune/src/vmaftune/saliency.py,tools/vmaf-tune/src/vmaftune/cli.py(newrecommendsubcommand),tools/vmaf-tune/AGENTS.md(saliency invariant),docs/usage/vmaf-tune.md(saliency section). - Upstream source: fork-local. The
vmaf-tunetree was introduced in PR #329 (ADR-0237 Phase A) and has no upstream Netflix counterpart. - On upstream sync: zero interaction — pure fork-local Python package under
tools/vmaf-tune/. - Invariant: the saliency-to-QP-offset signal blend (
offset = (2*sal − 1) * foreground_offset, clamped to ±12) is bit-for-bit equivalent tovmaf-roi's C-side blend (ADR-0247).tests/test_saliency.pypins the contract; ifvmaf-roi's C blend changes,saliency.pyfollows in the same PR. The test seam contract (session_factory=…,encode_runner=…) lets the suite run withoutonnxruntimeorffmpeg. - Re-test on rebase:
0229 — tools/vmaf-roi-score/ Option C scaffold (ADR-0296)¶
- ADR: ADR-0296
- Touches:
tools/vmaf-roi-score/pyproject.toml(new)tools/vmaf-roi-score/vmaf-roi-score(new console shim)tools/vmaf-roi-score/src/vmafroiscore/__init__.py(new)tools/vmaf-roi-score/src/vmafroiscore/cli.py(new)tools/vmaf-roi-score/src/vmafroiscore/score.py(new)tools/vmaf-roi-score/src/vmafroiscore/mask.py(new)tools/vmaf-roi-score/tests/test_combine.py(new)tools/vmaf-roi-score/README.md(new)tools/vmaf-roi-score/AGENTS.md(new)docs/adr/0296-vmaf-roi-saliency-weighted.md(new)docs/adr/_index_fragments/0296-vmaf-roi-saliency-weighted.md(new)docs/adr/_index_fragments/_order.txt— append-only.docs/research/0069-vmaf-roi-saliency-weighted.md(new)docs/usage/vmaf-roi-score.md(new)changelog.d/added/T6-2c-vmaf-roi-score-scaffold.md(new)- Invariant:
tools/vmaf-roi-score/is wholly fork-local. No upstream Netflix/vmaf surface owns or interacts with this directory. The combine math is a pure linear blend on Pythonfloat; the JSON schema is pinned byROI_RESULT_KEYSandSCHEMA_VERSION = 1. Schema bumps require an ADR-0288 supersession. Naming guard: do not confuse withcore/tools/vmaf_roi.c(ADR-0247) — that's the encoder-steering binary. The scoring tool here isvmaf-roi-score; the names diverge deliberately. - Rebase impact: zero. Pure-Python tool under
tools/; not part of the libvmaf C build, not part of any Netflix-mirrored surface. - Re-test on rebase:
0228 — vmaf-tune compare codec-comparison mode (research-0061 Bucket #7)¶
- Touches:
tools/vmaf-tune/src/vmaftune/compare.py(new). Wholly fork-local; no upstream Netflix/vmaf path overlap.tools/vmaf-tune/src/vmaftune/cli.py— adds thecomparesubparser and_run_comparerouter.tools/vmaf-tune/tests/test_compare.py(new). Mocked predicate; noffmpeg/vmafbinaries required.tools/vmaf-tune/AGENTS.md— invariant note for the predicate seam andCOMPARE_ROW_KEYScontract.docs/usage/vmaf-tune.md— new "Codec comparison" section.- Invariant:
compare.compare_codecsorchestrates per-codec ranking via an injectedpredicate(codec, src, target_vmaf) -> RecommendResultcallable. The orchestration must not branch on codec name; new codecs land as one-file additions undercodec_adapters/and are picked up automatically by the registry.COMPARE_ROW_KEYSis the JSON / CSV column contract — same maintenance discipline asCORPUS_ROW_KEYS. - Rebase impact: entirely fork-local. The Phase A + Phase B recommend backend (ADR-0237) is fork-internal; upstream Netflix/vmaf has no
tools/vmaf-tune/tree. - Re-test on rebase:
```shell pytest tools/vmaf-tune/tests/test_compare.py -v PYTHONPATH=tools/vmaf-tune/src python -m vmaftune.cli compare \ --src /tmp/ref.yuv --target-vmaf 92 --format markdown
0229 — vmaf-tune --score-backend GPU score wiring (ADR-0299)¶
- Touches:
tools/vmaf-tune/src/vmaftune/score_backend.py(new). Wholly fork-local —tools/vmaf-tune/has no upstream Netflix/vmaf overlap.tools/vmaf-tune/src/vmaftune/{score,corpus,cli}.py(additive kwargs, no API removals).tools/vmaf-tune/tests/test_score_backend.py(new).docs/usage/vmaf-tune.md(new GPU section + flag row).docs/adr/0299-vmaf-tune-gpu-score.md(new).docs/research/0071-vmaf-tune-gpu-score-backend.md(new).- Invariant: the libvmaf CLI exposes
--backend NAMEwith valuesauto|cpu|cuda|sycl|vulkanexactly. Help-text parser inscore_backend.parse_supported_backendspins this format. If upstream renames the flag or reformats the help line on merge, the parser silently degrades to "CPU only" — the test fixtures intest_score_backend.pywill catch the format change but only if re-run. - Upstream source: fork-local. Netflix upstream's CLI does not ship a
--backendselector (CPU-only). - On upstream sync: zero interaction.
vmaf-tunelives entirely in fork-introduced paths and consumes only the fork's--backendflag. - Re-test on rebase:
pytest tools/vmaf-tune/tests/test_score_backend.py -v
# If the libvmaf help text reformats, parse_supported_backends
# will return {"cpu"} on test_parse_full_backend_line_yields_all_four
# and the test fails loudly.
0261 — vmaf-tune HDR-aware encode + score path (2026-05-03)¶
- What changed: fork-local addition under
tools/vmaf-tune/src/vmaftune/hdr.pyplus wiring intocorpus.py/cli.py/score.py. Adds ffprobe-driven HDR detection, codec-specific HDR ffmpeg flag dispatch, schema-v2 corpus row keys (hdr_transfer,hdr_primaries,hdr_forced), and four--auto-hdr/--force-*CLI modes. See ADR-0300. - Upstream source: zero.
tools/vmaf-tune/is fork-introduced (Phase A under ADR-0237). - On upstream sync: zero interaction. Upstream Netflix/vmaf ships no encode automation surface; this tree is entirely fork-local and lives outside
libvmaf/andpython/. - Schema migration note:
SCHEMA_VERSIONbumped 1 → 2. The three new keys are additive — Phase B / C loaders treat missing keys as SDR for backward compat with v1 rows. - Re-test on rebase:
python -m pytest tools/vmaf-tune/tests/ -q
python -m vmaftune.cli corpus --help # confirm --auto-hdr surfaces
0298 — vmaf-tune content-addressed cache (ADR-0298)¶
- What changed: fork-local. New module
tools/vmaf-tune/src/vmaftune/cache.py; cache integration intools/vmaf-tune/src/vmaftune/corpus.py(iter_rowsnow consults the cache before encode/score); new CLI flags--no-cache,--cache-dir,--cache-size-gbincli.py. Codec-adapterProtocolgainsadapter_version: str; the lone Phase-A x264 adapter pins"1". - Upstream source: none.
tools/vmaf-tune/is fork-introduced (ADR-0237) and has no upstream counterpart. - On upstream sync: zero interaction with Netflix/vmaf master. The module sits entirely under
tools/vmaf-tune/, which upstream does not ship. - Invariant for future codec adapters: every
CodecAdaptermust declareadapter_version: str. Bump it whenever the adapter's argv shape, preset list, or quality range changes — otherwise the cache returns stale results post-upgrade. The contract is asserted bytest_cache_key_diffs_on_each_fieldintests/test_cache.py. - Re-test on rebase:
```bash pytest tools/vmaf-tune/tests/test_cache.py -v
0283 — vmaf-tune Apple VideoToolbox adapters (2026-05-05)¶
- What changed: fork-local addition under
tools/vmaf-tune/src/vmaftune/codec_adapters/. New files:h264_videotoolbox.py,hevc_videotoolbox.py,_videotoolbox_common.py, plus the registry hook in__init__.py. See ADR-0283. - Update 2026-05-09:
prores_videotoolbox.pyadapter added to the same registry pattern (broadcast / prosumer ProRes intermediate). Quality knob differs — ProRes is a fixed-rate codec, so the harness's--crfslot carries the integer ProRes tier id (0=proxy→ 5=xq) rather than a-q:vvalue._videotoolbox_common.pyextended withPRORES_PROFILE_*constants +validate_prores_videotoolbox()/prores_profile_name()helpers; profile ids verified against FFmpeg n8.1.1libavcodec/videotoolboxenc.c. See the Status update appendix in ADR-0283. - Upstream source: zero.
tools/vmaf-tune/is fork-introduced (Phase A under ADR-0237). - On upstream sync: zero interaction.
- Re-test on rebase:
python -m pytest tools/vmaf-tune/tests/test_codec_adapter_videotoolbox.py -q
python -m pytest tools/vmaf-tune/tests/test_codec_adapter_prores_videotoolbox.py -q
0228 — vmaf-tune coarse-to-fine CRF search (ADR-0306)¶
- What changed: fork-local tooling. Adds
coarse_to_fine_search()totools/vmaf-tune/src/vmaftune/corpus.py, plumbs new CLI flags ontovmaf-tune corpus(--coarse-to-fine,--coarse-step,--fine-radius,--fine-step,--target-vmaf), and ships a newvmaf-tune recommendsubcommand. Widenstools/vmaf-tune/src/vmaftune/codec_adapters/x264.pyquality_rangefrom(15, 40)to(0, 51). JSONL row schema unchanged (SCHEMA_VERSION=1). - Upstream source: fork-local. The whole
tools/vmaf-tune/tree is fork-introduced (ADR-0237); upstream Netflix/vmaf has no encode-automation surface. - On upstream sync: zero interaction.
tools/vmaf-tune/is not mirrored from upstream. - Re-test on rebase:
0314 — vmaf-tune --score-backend=vulkan (ADR-0314)¶
- Touches:
tools/vmaf-tune/src/vmaftune/cli.py(additive argparse flag oncorpus+recommendsubparsers; resolvesselect_backendand catchesBackendUnavailableErrorfor clean exit-2).tools/vmaf-tune/src/vmaftune/score.py(additivebackendkwarg onbuild_vmaf_commandandrun_score;None= no flag emitted).tools/vmaf-tune/src/vmaftune/corpus.py(newCorpusOptions.score_backendfield, defaultNone; forwarded intorun_score).tools/vmaf-tune/tests/test_score_backend.py(additive Vulkan-specific tests; pre-existing tests now pass after thebackend=kwarg lands).docs/adr/0314-vmaf-tune-score-backend-vulkan.md(new).docs/usage/vmaf-tune.md(new "Vulkan score backend" subsection under the existing GPU-scoring section).tools/vmaf-tune/AGENTS.md(invariant note: argparse choices stay in sync with libvmaf--backendvocabulary).changelog.d/added/vmaf-tune-score-backend-vulkan.md(new).- Invariant:
score_backend.ALL_BACKENDS = ("cpu", "cuda", "sycl", "vulkan")is the exact set libvmaf'score/tools/cli_parse.c--backendalternation accepts. Adding a new harness-side value without the libvmaf-side wiring produces silent strict-mode failures on hosts that probe positively for it. - Upstream source: zero. Netflix upstream's CLI does not ship a
--backendselector; bothtools/vmaf-tune/andcore/src/vulkan/are fork-introduced. - On upstream sync: zero interaction. No upstream-mirror file is touched.
- Re-test on rebase:
pytest tools/vmaf-tune/tests/test_score_backend.py -v -k vulkan
pytest tools/vmaf-tune/tests/test_score_backend.py -v
Failures here usually indicate the libvmaf help-text format changed; score_backend.parse_supported_backends test fixtures pin the format and will fail loudly.
0303 — fr_regressor_v2 ensemble prod flip (ADR-0303)¶
- ADR: ADR-0303
- Touches: entirely fork-local.
ai/scripts/train_fr_regressor_v2_ensemble_loso.py(new — 9-fold LOSO trainer over the five ensemble seeds; emitsloso_seed{N}.jsonartefacts).scripts/ci/ensemble_prod_gate.py(new — reads fiveloso_seed{N}.jsonfiles, returns exit 0 iffmean(PLCC_i) ≥ 0.95ANDmax - min ≤ 0.005).ai/AGENTS.md— appended "Ensemble registry invariant" paragraph under the existingfr_regressor_v2_ensemble_v1section.docs/adr/0303-fr-regressor-v2-ensemble-prod-flip.md(new),docs/research/0075-fr-regressor-v2-ensemble-prod-flip.md(new),changelog.d/added/fr-regressor-v2-ensemble-prod-flip.md(new).- Rebase invariant: the production ship gate is two-part —
mean_i(PLCC_i) ≥ 0.95ANDmax_i(PLCC_i) - min_i(PLCC_i) ≤ 0.005over five seeds. The variance bound is load-bearing: removing it silently allows a one-seed-wins-four-seeds-tie configuration that invalidates the ensemble's predictive-distribution semantics. Both thresholds live inscripts/ci/ensemble_prod_gate.py; do not weaken either without superseding ADR-0303. - Rebase invariant (registry): the five
fr_regressor_v2_ensemble_v1_seed{0..4}registry rows aresmoke: trueon master at this commit; flipping them tofalseis the follow-up flip PR's job, gated on a real-corpus LOSO run + the CI gate. Do not flip seed rows during a rebase merge conflict resolution. - Re-test on rebase:
python3 -c "import ast; ast.parse(open('ai/scripts/train_fr_regressor_v2_ensemble_loso.py').read())"
python3 -c "import ast; ast.parse(open('scripts/ci/ensemble_prod_gate.py').read())"
python ai/scripts/train_fr_regressor_v2_ensemble_loso.py --help
python scripts/ci/ensemble_prod_gate.py --help
- Upstream source: zero.
fr_regressor_v2and its ensemble are fork-introduced (parent ADR-0272 / ADR-0279). - On upstream sync: zero interaction.
0313 — CI required-checks aggregator (2026-05-05)¶
- What changed: fork-local CI policy. New
.github/workflows/required-aggregator.yml— single workflow that runs on every non-draft PR and verifies the 23 named required checks reportedsuccess/skipped/neutral(or didn't appear at all, which is the path-filter-rejection semantics). Aggregator becomes the single branch-protection required check, replacing the 23-name list from ADR-0037. - Touches:
.github/workflows/required-aggregator.yml(new),docs/adr/0313-ci-required-checks-aggregator.md(new),changelog.d/added/ci-required-checks-aggregator.md(new),docs/adr/README.md(+1 row),docs/adr/_index_fragments/_order.txt(+1 line + new fragment file). - Upstream source: zero. Branch-protection policy is fork-only.
- On upstream sync: zero interaction with Netflix/vmaf master.
- Manual operator step at adoption (uses PATCH, not PUT — corrected from the original ADR-0313 body which had the wrong verb):
echo '{"strict": false, "contexts": ["Required Checks Aggregator"]}' | \
gh api -X PATCH "repos/VMAFx/vmafx/branches/master/protection/required_status_checks" --input -
- Re-test on rebase:
# YAML lint passes
python3 -c "import yaml; yaml.safe_load(open('.github/workflows/required-aggregator.yml'))"
0305 — encoder knob-space Pareto analysis (2026-05-05)¶
- What changed: fork-local. New analysis scaffold for the 12,636-cell encoder knob sweep that backs
tools/vmaf-tune/codec_adapters/*recipe defaults. New files:ai/scripts/analyze_knob_sweep.py(per-(source, codec, rc_mode)Pareto hull on(bitrate_kbps, vmaf_score),encode_time_mstiebreaker, regression-detection check),ai/tests/test_knob_sweep_analysis.py(synthetic 20-row JSONL fixture). Methodology + scaffolded findings: see ADR-0305 + Research-0077. Companion to Research-0063. - Touches: none upstream-shared. Sits entirely under
ai/(fork-local since the tiny-AI training surface, ADR-0021) anddocs/{adr,research}/(fork ledger). - Upstream source: zero. The 12,636-cell sweep, the Pareto scaffold, and the regression-detection invariant are fork-introduced; Netflix/vmaf master ships no encoder knob-sweep tooling.
- On upstream sync: zero interaction with Netflix/vmaf master.
- Invariant for future codec adapter PRs: per the
ai/AGENTS.mdknob-sweep corpus invariant (ADR-0305), recipes that regress vs the bare encoder at matched bitrate within the same(source, codec, rc_mode)slice MUST NOT ship as adapter defaults. New adapter PRs cite the per-slice hull row fromreports/summary.md(or "no hull entry yet — bare default") in their PR description. Thecomprehensive.jsonlsweep file is generated locally and lives underruns/phase_a/full_grid/(gitignored — never committed). - Re-test on rebase:
0302 — ENCODER_VOCAB v3 schema expansion (ADR-0302)¶
- Touches:
ai/scripts/train_fr_regressor_v2.py(adds anENCODER_VOCAB_V3parallel constant; does not modify the liveENCODER_VOCABorENCODER_VOCAB_VERSION). - Invariant:
ENCODER_VOCABis append-only and order-stable (per ADR-0235). The v3 scaffold preserves the v2 slot ordering verbatim — slots 0..12 are bit-identical to the v2 vocab; slots 13/14/15 appendlibsvtav1,h264_videotoolbox,hevc_videotoolbox. The liveENCODER_VOCAB_VERSION = 2remains the source of truth until the follow-up retrain PR clears the LOSO PLCC ship gate. - Upstream interaction: zero.
ai/scripts/train_fr_regressor_v2.pyis fork-introduced (ADR-0272) and has no upstream counterpart. - Re-test on rebase:
python3 -c "
import importlib.util, pathlib
spec = importlib.util.spec_from_file_location(
't', pathlib.Path('ai/scripts/train_fr_regressor_v2.py')
)
m = importlib.util.module_from_spec(spec)
spec.loader.exec_module(m)
assert len(m.ENCODER_VOCAB_V3) == 16
assert m.ENCODER_VOCAB_VERSION == 2
print('OK')
"
0304 — vmaf-tune fast-path prod wiring (ADR-0304)¶
- Touches:
tools/vmaf-tune/src/vmaftune/fast.py(replaces the ADR-0276 scaffold'sNotImplementedErrorpaths with concrete Optuna TPE + v2 proxy + GPU verify wiring); new moduletools/vmaf-tune/src/vmaftune/proxy.py(centralised seam forfr_regressor_v2ONNX inference); expandedtools/vmaf-tune/tests/test_fast.py. Doc-side: ADR-0304, Research-0076,tools/vmaf-tune/AGENTS.mdinvariant note. - Upstream source: zero.
tools/vmaf-tune/andmodel/tiny/fr_regressor_v2.onnxare both fork-introduced (ADR-0237 / ADR-0352). - Invariant: the production proxy is always
fr_regressor_v2(no smoke models in the production path) and a single GPU verify pass at recommend-end is mandatory — proxy alone never wins. Thevmaftune.proxy.run_proxyhelper is the single seam every fast-path consumer goes through; future probabilistic-head / ensemble migrations land in that one module. ENCODER_VOCAB v2 one-hot ordering is frozen by ADR-0352 and pinned inproxy.ENCODER_VOCAB_V2— keep in sync withai/scripts/train_fr_regressor_v2.py; drift raisesProxyErrorat inference time before bad predictions ship. - On upstream sync: zero interaction with Netflix/vmaf master.
- Re-test on rebase:
0307 — vmaf-tune ladder default sampler wiring (ADR-0307)¶
- What changed: fork-local tooling.
tools/vmaf-tune/src/vmaftune/ladder.py::_default_samplerno longer raisesNotImplementedError; it composescorpus.iter_rows(Phase A encode + score) withrecommend.pick_target_vmaf(smallest CRF clearing target VMAF) overDEFAULT_SAMPLER_CRF_SWEEP = (18, 23, 28, 33, 38)at the adapter's mid-range preset. Module-level docstring + AGENTS.md invariant updated. New tests intools/vmaf-tune/tests/test_ladder.pystubiter_rowsviamonkeypatch.setattrso no live ffmpeg / vmaf binaries are needed. - Upstream source: fork-local. The whole
tools/vmaf-tune/tree is fork-introduced (ADR-0237); upstream Netflix/vmaf has no encode-automation / ladder surface. - On upstream sync: zero interaction.
tools/vmaf-tune/is not mirrored from upstream. - Rebase invariant: the 5-point sweep
(18, 23, 28, 33, 38)is the load-bearing default; downstream Phase E callers size their wall-time budget against five encodes per(resolution, target_vmaf)cell. Do not widen / narrow it without an ADR-0307 follow-up. TheSamplerFnseam stays open — callers needing finer grids pass an explicitsampler=. - Re-test on rebase:
0309 — fr_regressor_v2 ensemble real-corpus retrain harness (ADR-0309)¶
- ADR: ADR-0309
- Touches: entirely fork-local.
ai/scripts/run_ensemble_v2_real_corpus_loso.sh(new — Bash wrapper that loops the five seeds over the existingtrain_fr_regressor_v2_ensemble_loso.pyagainst.workingdir2/netflix/).ai/scripts/validate_ensemble_seeds.py(new — calls the ADR-0303 gate and writesPROMOTE.json/HOLD.jsonwith a corpus sha256 snapshot).ai/tests/test_validate_ensemble_seeds.py(new — 7 tests, synthetic JSON fixtures for both verdict paths).ai/AGENTS.md— appended "Registry-flip is a separate PR (ADR-0309)" paragraph under the existingfr_regressor_v2_ensemble_v1section.docs/adr/0309-fr-regressor-v2-ensemble-real-corpus-retrain.md,docs/research/0081-fr-regressor-v2-ensemble-real-corpus-methodology.md,docs/ai/ensemble-v2-real-corpus-retrain-runbook.md(all new).- Rebase invariant: the harness is decoupled from the registry mutation. Neither the wrapper nor the validator touches
model/tiny/registry.json; the registry flip is a separate follow-up PR gated on a passingPROMOTE.json. Auto-flipping on PROMOTE was rejected in ADR-0309's alternatives matrix specifically because rebase-time mutation of shipped registry rows is the foot-gun this invariant exists to prevent. - Re-test on rebase:
python -m pytest ai/tests/test_validate_ensemble_seeds.py -v
python ai/scripts/validate_ensemble_seeds.py --help
bash -n ai/scripts/run_ensemble_v2_real_corpus_loso.sh
- Upstream source: zero.
- On upstream sync: zero interaction.
0310 — BVI-DVC corpus ingestion for fr_regressor_v2 (ADR-0310)¶
- Touches:
ai/scripts/bvi_dvc_to_corpus_jsonl.py(new fork-only adapter),ai/scripts/merge_corpora.py(new fork-only shard merger),ai/tests/test_merge_corpora.py(new),docs/ai/bvi-dvc-corpus-ingestion.md(new),docs/adr/0310-bvi-dvc-corpus-ingestion.md(new),docs/research/0082-bvi-dvc-corpus-feasibility.md(new),ai/AGENTS.md(BVI-DVC invariant note). - Invariant: the BVI-DVC archive and any extracted artefacts (parquet, cached libvmaf JSON, JSONL corpus shard) are research-only and stay local — only derived
fr_regressor_v2_*.onnxweights ship. The merge utility validates every row against the canonicalvmaftune.CORPUS_ROW_KEYStuple; the schema is the merge contract. Re-shape here is a pure transform on the cached libvmaf JSON; no ffmpeg / vmaf binary is invoked. The(src_sha256, encoder, preset, crf)natural key is load-bearing for de-duplication across mirrors and re-encodes. - Upstream interaction: none.
ai/is fork-introduced; BVI-DVC is not part of Netflix/vmaf upstream. - Re-test on rebase:
ADR-0312 — ffmpeg-patches/ vmaf-tune integration (2026-05-05)¶
- Files:
ffmpeg-patches/0007-libvmaf-tune-qpfile-unified.patch,ffmpeg-patches/0008-add-libvmaf_tune-filter.patch,ffmpeg-patches/0009-pass-autotune-cli-glue.patch,ffmpeg-patches/series.txt,ffmpeg-patches/README.md. - Rebase invariant: patches
0007–0009plug into the cumulative state after patches0001–0006apply against pristinen8.1. Per-patchgit apply --checkin isolation is the wrong gate; use the series-replay command in CLAUDE.md §12 r14 instead. - vmaf-tune patch invariant: the qpfile parser at
libavcodec/qpfile_parser.{c,h}is shared across all three encoder adapters in patch 0007. Future encoders that grow a-qpfileAVOption inherit it; do not fork the parser. Whentools/vmaf-tune/src/vmaftune/saliency.py's qpfile output format changes (new column, different frame-type alphabet, …), patch 0007 must change in the same PR (CLAUDE.md §12 r14). - vf_libvmaf_tune full-scoring promotion (2026-05-06): patch 0008 originally shipped as a scaffold (linear CRF↔VMAF interpolation, no libvmaf scoring) per ADR-0312's deferred-alternatives column. The filter now mirrors
vf_libvmaf.c's CPU framesync pipeline end-to-end (vmaf_init+vmaf_model_load+vmaf_use_features_from_modelin init(); per-framevmaf_picture_alloc+ memcpy +vmaf_read_pictures; flush +vmaf_score_pooled(MEAN)in uninit()). The CRF recommendation remains a piece-wise linear projection from the observed VMAF; per-clip Optuna TPE search stays intools/vmaf-tune/src/vmaftune/recommend.py. Rebase-side: the new filter still depends only on libvmaf's CPU C-API (vmaf_init,vmaf_model_load,vmaf_use_features_from_model,vmaf_read_pictures,vmaf_score_pooled,vmaf_close,vmaf_picture_alloc/unref); zero new symbols beyond whatvf_libvmaf.calready requires, so future libvmaf rebases that pass the existing libvmaf filter pass this one too. ADR-0312 sub-decision retired. - n7+ API migration (2026-05-06): patch 0008 originally referenced the removed
AVFilterLink::frame_ratemember directly (n6-era API); in n7+ that field moved offAVFilterLinkonto a newFilterLinkstruct accessed viaff_filter_link(AVFilterLink *)fromlibavfilter/filters.h. Patch 0008 now usesff_filter_link(outlink)->frame_rate = ff_filter_link(mainlink)->frame_rate;inconfig_output(), mirroring patches 0005/0006 which were already written against the post-n7 API. The bug slipped through CI because the FFmpeg-Vulkan lane only buildsvf_libvmaf.o, notvf_libvmaf_tune.c; the full SYCL lane catches it now that PR #415 addedffmpeg-patches/**to the integration workflow's path filter. Discovery: PR #415 / ADR-0317. - Upstream source: zero. The vmaf-tune integration is fork-introduced; pure upstream syncs are unaffected.
- On upstream sync: zero interaction with libvmaf master. FFmpeg-side rebases when n8.1 → n8.x land in
ffmpeg-patches/test/build-and-run.sh'sFFMPEG_SHAare tracked separately under each refresh ADR (e.g., ADR-0277 for the 2026-05-04 refresh). - Re-test on rebase:
git -C /path/to/ffmpeg-8 reset --hard n8.1
for p in ffmpeg-patches/000*-*.patch; do
git -C /path/to/ffmpeg-8 am --3way "$p" || break
done
# Build smoke (libvmaf-disabled — patches 0001–0006 skipped if libvmaf_dnn
# is not built). With libvmaf_dnn available:
cd /path/to/ffmpeg-8 && ./configure --enable-libvmaf --enable-libx264 --enable-libsvtav1 --enable-libaom --enable-gpl
make -j$(nproc) ffmpeg
./ffmpeg -hide_banner -h encoder=libx264 2>&1 | grep -i qpfile
-
2026-05-06 update — patch 0007 SVT-AV1 ROI bridge promoted from scaffold to full impl: the libsvtav1 hunk now sets
enc_params.enable_roi_map = true, builds oneSvtAv1RoiMapEvtper qpfile frame upfront ineb_enc_init(per-MB qp_offsets averaged into per-64×64-SBb64_seg_mapof up to 8 segment QPs; uniform binning when the value span exceeds the segment budget), and attaches each event as aROI_MAP_EVENTpriv-data node fromeb_send_frame()withnode->size = sizeof(SvtAv1RoiMapEvt*)(the validation contract enforced by SVT-AV1'sresource_coordination_process.c). Lifetime invariant: events + maps live for the entire encode session because SVT-AV1 reads ROI_MAP_EVENT data via shallow-copied pointers on async pipeline threads (perenc_handle.c::copy_private_data_list);eb_enc_closefrees them. Wiring is gated onSVT_AV1_CHECK_VERSION(1, 6, 0); older SVT-AV1 builds keep the log-and-continue fallback. libaom remains scaffold-only — itsAOME_SET_ROI_MAPbridge stays a separate follow-up. No new ADR per CLAUDE.md §12 r8 (executes the existing ADR-0312 decision). -
2026-05-06 update — patch 0007 libaom-av1 ROI bridge promoted from scaffold to full impl: the libaom-av1 hunk now caches the parsed
VmafTuneQpFileinAOMContext, allocates a segment-id map at libaom's mode-info grid (ALIGN_POWER_OF_TWO(dim, 8) >> 2, sinceav1/common/enums.h::MI_SIZE == 4), and on every encoded frame picks up to 8 segment QPs from the per-frame qp_offset value range (uniform linear binning when the span exceedsAOM_MAX_SEGMENTS == 8), paints the per-mi segment map by expanding each per-16×16-MB qp_offset into a 4×4 block of mi cells, and issuesaom_codec_control(&ctx->encoder, AOME_SET_ROI_MAP, &roi_map). Lifetime invariant: libaom deep-copies the segment map anddelta_q[]table on every control call (perav1/encoder/encoder.c::av1_set_roi_map memcpy), so a single buffer is reused across frames and freed inaom_free(). The qpfile is also freed there. Trade-off: the 8-segment cap rounds nearby qp_offsets together when the saliency model emits more than 8 distinct values per frame; finer granularity requiresvmaf-tune corpusinstead. This retires the libaom-av1 deferral noted under ADR-0312 — both AV1 encoder hooks (libsvtav1 and libaom-av1) are now full-impl. No new ADR per CLAUDE.md §12 r8 (executes the existing ADR-0312 decision).
0315 — Vendor-neutral VVC encode strategy (ADR-0315 / Research-0085)¶
- ADR: ADR-0315
- Digest: Research-0085
- Touches: docs-only.
docs/research/0085-vendor-neutral-vvc-encode-landscape.md(new).docs/adr/0315-vendor-neutral-vvc-encode-strategy.md(new).docs/adr/_index_fragments/0315-vendor-neutral-vvc-encode-strategy.md(new).docs/adr/_index_fragments/_order.txt(one-line append).changelog.d/added/research-0085-vendor-neutral-vvc-encode.md(new).docs/rebase-notes.md(this entry).- Rebase invariant: none. The research digest and ADR are pure surveys with no code dependencies; nothing in the fork's source tree references them in a way that breaks on upstream rebase.
- Upstream source: zero. VVC encode strategy is a fork-local decision; upstream Netflix/vmaf has no codec adapter or encode-automation surface.
- On upstream sync: zero interaction. Pure docs.
- Re-test on rebase:
- 2026-05-06 follow-up (Research-0085 verification pass):
docs/research/0085-vendor-neutral-vvc-encode-landscape.mdflipped fromStatus: SKELETONtoStatus: Active. Most[UNVERIFIED]claims are now backed by primary-source URLs (NVIDIA SDK 13.0 docs, AMD AMF GitHub, Intel oneVPL GitHub +mfxstructures.h+CHANGELOG.md, Khronos registry, Phoronix Mesa/RADV coverage, VVenC issue tracker, ZLUDA repo).- ADR-0315's
## Contextand## Alternatives consideredrefreshed with the verified data points. Status staysProposed. [UNVERIFIED]count in the digest dropped 25 → 10; remaining items are legitimate gaps (NN-VC quality lift, vvenc per-kernel profile, HHI's non-public roadmap).- No code touched. No rebase impact beyond the existing docs-only posture.
0316 — cli_parse.c error() long-only-option fix (ADR-0316)¶
- ADR: ADR-0316 (follow-up to ADR-0311).
- Digest: none — bug-fix; fix shape fits in the ADR/commit body.
- Touches:
core/tools/cli_parse.c(3 lines — call-site arg change at theARG_THREADS/ARG_SUBSAMPLE/ARG_CPUMASKhandlers).core/test/fuzz/fuzz_cli_parse.c(removedknown_assert_in_inputearly-reject filter).core/test/fuzz/cli_parse_corpus/cli_threads_abbrev_assert.argv(promoted fromcli_parse_known_crashes/).core/test/test_cli_parse_long_only_args.c(new fork()-based regression test).core/test/meson.build(new test wiring, gated off Windows alongsidetest_y4m_411_oob).core/tools/AGENTS.md(added a long-only-options invariant note next to the existingcli_parse.crules).- Rebase invariant: load-bearing.
cli_parse.cis upstream-mirror with fork additions; the three handlers carry the fork-local shape of passing theARG_*enum value (not't'/'s'/'c') toparse_unsigned(). If an upstream sync re-introduces the original short-option char shape, the assert returns and the parked-then-promoted reproducer (cli_parse_corpus/cli_threads_abbrev_assert.argv) will surface it in the next nightly fuzz run. - Upstream source: the bug shape exists in Netflix/vmaf master too (long-only options were added upstream with the same short-option-char placeholder). When the fork ports an upstream fix that overlaps these handlers, prefer the
parse_unsigned(optarg, ARG_*, argv[0])form already on the fork. - On upstream sync: re-apply the three-line change in
cli_parse.cif upstream resets the call-site args. The unit test is fork-local and stays. - Re-test on rebase:
meson setup core/build libvmaf -Denable_tests=true \
-Denable_cuda=false -Denable_sycl=false
ninja -C core/build test/test_cli_parse_long_only_args
meson test -C core/build test_cli_parse_long_only_args -v
ADR-0317 — CI flake fix: doc-only PR path-filter (2026-05-06)¶
- Touched files:
.github/workflows/docker-image.yml— addedpaths:filter on bothpush:andpull_request:triggers..github/workflows/ffmpeg-integration.yml— addedpaths:filter on bothpush:andpull_request:triggers (covers all four matrix lanes: gcc, clang, SYCL, Vulkan).docs/adr/0317-ci-doc-only-pr-flake-fix.md,docs/adr/README.md(index row),changelog.d/fixed/ci-doc-only-pr-flakes.md.- Rebase invariant: not load-bearing. Workflow-only change. Both files are fork-local CI; upstream Netflix/vmaf does not ship a Docker workflow or an FFmpeg-integration matrix in this shape, so rebase conflicts are unlikely. If a future upstream sync introduces an overlapping
docker-image.ymlor FFmpeg matrix, prefer the fork's path-filtered form — the rationale (ADR-0313 aggregator posture, doc-only-PR runner-time burn) is fork-specific. - Upstream source: none — fork-local CI workflows.
- On upstream sync: no action required. If reviewers later add new build inputs (e.g. a top-level
docker-compose.yml, a newffmpeg-patches/*.txtconfig file), extend thepaths:lists in the same PR that adds the input. - Follow-up not in this ADR: patch
ffmpeg-patches/0008-add-libvmaf_tune-filter.patchline 256 (outlink->frame_rate = mainlink->frame_rate;) needs to migrate to theff_filter_link()accessor introduced in FFmpeg n7+, matching the pattern already in patches 0005 / 0006. Tracked separately; the path-filter does not hide it (any libvmaf/ or ffmpeg-patches/ PR will still trip the SYCL lane). - Re-test on rebase:
python3 -c "import yaml; \
yaml.safe_load(open('.github/workflows/docker-image.yml')); \
yaml.safe_load(open('.github/workflows/ffmpeg-integration.yml')); \
print('OK')"
0319 — fr_regressor_v2 ensemble LOSO trainer — real loader + per-fold training (ADR-0319)¶
- Touches:
ai/scripts/train_fr_regressor_v2_ensemble_loso.py(real_load_corpus+_train_one_seedbodies),ai/scripts/run_ensemble_v2_real_corpus_loso.sh(wrapper argv fix),docs/ai/ensemble-v2-real-corpus-retrain-runbook.md(Step 0 corpus-generation section),ai/AGENTS.md(canonical-6 schema invariant note),ai/tests/test_train_fr_regressor_v2_ensemble_loso_*.py(loader + train schema tests). Closes the deferrals tracked in rebase-notes §0303 + §0309. - Upstream source: none — fork-local ML training infrastructure. Netflix/vmaf upstream has no
fr_regressor_v2surface, no LOSO trainer, and no canonical-6 corpus tooling. - Invariant: the trainer's
_load_corpusaccepts the canonical-6 JSONL schema emitted byscripts/dev/hw_encoder_corpus.pybit-for-bit — required keys per row are(src, encoder, cq, frame_index, vmaf, adm2, vif_scale0..3, motion2). Codec block layout is 12-slotENCODER_VOCABv2 one-hot + constantpreset_norm = 0.5+crf_norm = (cq - cq_min) / (cq_max - cq_min). Schema changes require anENCODER_VOCAB_VERSIONbump and full ensemble retrain per the existing closed-vocabulary rule (ADR-0235 / ADR-0352). Fold-level StandardScaler is fit on the training rows only; leaking the held-out source's distribution into the scaler would silently inflate per-fold PLCC. - On upstream sync: no action required. If upstream Netflix/vmaf ever adds a competing LOSO trainer under
python/vmaf/, do NOT merge them — keep the fork's training stack underai/per the AGENTS.md scope rule. - Re-test on rebase:
pytest ai/tests/test_train_fr_regressor_v2_ensemble_loso_loader.py \
ai/tests/test_train_fr_regressor_v2_ensemble_loso_train.py -v
bash -n ai/scripts/run_ensemble_v2_real_corpus_loso.sh
ADR-0323 — fr_regressor_v3 train + register on ENCODER_VOCAB v3 (2026-05-06)¶
- Scope:
ai/scripts/train_fr_regressor_v3.py(new),ai/tests/test_train_fr_regressor_v3.py(new),model/tiny/fr_regressor_v3.onnx(new, real-weight checkpoint from a 9-fold LOSO gate-pass at mean PLCC 0.9975),model/tiny/fr_regressor_v3.json(new sidecar withencoder_vocab_version: 3and full per-fold trace),model/tiny/registry.json(newfr_regressor_v3row,smoke: false),ai/AGENTS.md(v3 retrain invariant section gains a "Status" subsection recording the gate result),docs/ai/models/fr_regressor_v3.md(new model card),docs/adr/0323-fr-regressor-v3-train-and-register.md+ index row,changelog.d/added/fr-regressor-v3-train-register.md. - Rebase impact: zero. Fork-local feature; no upstream Netflix/vmaf surface is touched. The 16-slot
ENCODER_VOCAB_V3imported fromtrain_fr_regressor_v2.pywas already landed by PR #401 (ADR-0302). - On upstream sync: no action required. The v3 model ships alongside v2 —
fr_regressor_v2.onnxand its sidecar are unchanged; the v3 row is appended to the registry and sorted alphabetically. If a future upstream sync ever lands a competingfr_regressor_v3model underpython/vmaf/, do NOT cross-link them — the fork's training stack lives underai/. - Watch out for: the live
ENCODER_VOCAB_VERSIONinai/scripts/train_fr_regressor_v2.pystays at 2 (per ADR-0302's invariant). Do not bump it to 3 in this PR or in any downstream port; the in-place promotion of v3 over v2 is a separate "promote v3 to authoritative" PR per ADR-0302's production-flip checklist. - Re-test on rebase:
pytest ai/tests/test_train_fr_regressor_v3.py -v
bash core/test/dnn/test_registry.sh # must report OK: 20+
python -c "import onnx; onnx.checker.check_model(onnx.load('model/tiny/fr_regressor_v3.onnx')); print('OK')"
ADR-0321 — fr_regressor_v2_ensemble_v1 full production flip (2026-05-06)¶
- Scope:
ai/scripts/export_ensemble_v2_seeds.py(new),model/tiny/fr_regressor_v2_ensemble_v1_seed{0..4}.onnx(real full-corpus-trained weights replacing the 3025-byte synthetic scaffold bytes),model/tiny/fr_regressor_v2_ensemble_v1_seed{0..4}.json(new per-seed sidecars),model/tiny/registry.json(sha256 +smoke: falseon the five seed rows),ai/AGENTS.md(new invariant: the registry-flip is now done; future re-flips require a fresh PROMOTE.json + re-run of the export driver). - Rebase impact: zero. This is a fork-local production-flip; no upstream Netflix/vmaf surface is touched. The 12-slot
ENCODER_VOCABv2 carried in each sidecar is the same one the LOSO trainer (ADR-0319) bakes into the codec-block layout, so there is no rebase-time vocabulary drift to worry about. - Watch out for: if a future upstream sync ever introduces a competing
fr_regressor_v2_ensemble_*model underpython/vmaf/, do NOT cross-link them — the fork's ensemble weights are gated onruns/ensemble_v2_real/PROMOTE.jsonand are not portable to a different training stack. - Re-test on rebase:
bash core/test/dnn/test_registry.sh # must report OK: 19
python -c "import onnx; \
[onnx.checker.check_model(onnx.load(f'model/tiny/fr_regressor_v2_ensemble_v1_seed{i}.onnx')) \
for i in range(5)]; print('OK')"
ADR-0324 — Ensemble training kit (2026-05-06)¶
- Touches:
tools/ensemble-training-kit/(new),docs/adr/0324-ensemble-training-kit.md(new),docs/adr/README.md(index row),changelog.d/added/0324-ensemble-training-kit.md(new). No engine code touched; no upstream-shared paths. - Invariant: the kit assumes the LOSO wrapper hard-codes seeds
(0 1 2 3 4). The orchestrator surfaces a warning if--seedsdeviates but still hands off to the wrapper. If a future PR parameterises the wrapper's seed list, update both the wrapper and the kit's pass-through logic in lockstep. - On upstream sync: no action required. The kit lives entirely under
tools/ensemble-training-kit/(a fork-local path) and only invokes other fork-local scripts (ai/scripts/,scripts/dev/,scripts/ci/). - Re-test on rebase:
bash -n tools/ensemble-training-kit/*.sh
bash tools/ensemble-training-kit/make-distribution-tarball.sh /tmp/kit-test.tar.gz
tar -tzf /tmp/kit-test.tar.gz | grep -q "tools/ensemble-training-kit/run-full-pipeline.sh"
ADR-0332 — External-competitor benchmark harness (2026-05-08)¶
- Touches:
tools/external-bench/(new),docs/adr/0332-external-bench-wrapper-only.md(new),docs/adr/_index_fragments/0332-external-bench-wrapper-only.md(new),docs/adr/_index_fragments/_order.txt(one-line append),docs/adr/README.md(regenerated),changelog.d/added/external-bench-harness.md(new),docs/research/0087-external-bench-competitor-survey-2026-05-08.md(new). No engine code touched; no upstream-shared paths. - Invariant: the harness is wrapper-only — never vendor or link
x264-pVMAF(GPL-2.0) into this fork. Future competitors follow the same pattern (tools/external-bench/<competitor>/run.shinvokes a user-installed binary via env var; output schema-shimmed into the canonical JSON shape). The output schema (frames[].{frame_idx, predicted_vmaf_or_mos, runtime_ms}+summary.{competitor, plcc, srocc, rmse, runtime_total_ms, params, gflops}) is the contract between every wrapper andcompare.py.run_wrapper'srunnerparameter MUST stay resolved at call time (not via default-arg binding) so monkeypatch-based tests work. - On upstream sync: no action required. The harness lives entirely under
tools/external-bench/(a fork-local path) and never touches Netflix-shared code.
ADR-0331 — Skip CI on draft pull requests (2026-05-08)¶
- Touches:
.github/workflows/{docker-image,security-scans,lint-and-format,ffmpeg-integration,libvmaf-build-matrix,rule-enforcement,tests-and-quality-gates}.yml(per-jobif:clause +pull_request.typeslist).required-aggregator.ymlis unchanged — it already adopted the pattern under ADR-0313. No upstream-shared paths. - Invariant: every top-level job in the eight fork workflows that trigger on
pull_requestcarriesif: github.event_name != 'pull_request' || github.event.pull_request.draft == false. Thepull_request:block listsready_for_reviewintypes:so promotion of a draft fires CI exactly once. The second clause keepspush:triggers (no PR object) intact. If an upstream merge introduces a new top-level job, that job MUST inherit the gate; otherwise drafts will silently consume one matrix slot per push. -
On upstream sync: Netflix/vmaf upstream does not gate on draft state; if a sync brings in new
pull_requestworkflow content, replay the gate on every newly-introduced top-level job. Composing with an existingif:follows thecoverage-gpupattern — wrap both predicates in${{ ... && ( ... ) }}. -
Re-test on rebase:
```bash python3 -c "import yaml; names=['docker-image','security-scans','lint-and-format','required-aggregator','ffmpeg-integration','libvmaf-build-matrix','rule-enforcement','tests-and-quality-gates']; [yaml.safe_load(open(f'.github/workflows/{n}.yml')) for n in names]; print('OK')" # Spot-check the gate is present on every top-level job: for f in docker-image security-scans lint-and-format ffmpeg-integration \ libvmaf-build-matrix rule-enforcement tests-and-quality-gates \ required-aggregator; do grep -c "pull_request.draft == false" ".github/workflows/${f}.yml" done # Each must report >= 1.
SSIM extractor registration fix (2026-05-08)¶
- Touches:
core/src/feature/feature_extractor.c(upstream-mirror — adds one extern + one registry-array entry near the existing SSIM rows),core/src/feature/integer_ssim.c(upstream-mirror — adds#include "config.h"and refreshes the file-scope comment abovevmaf_fex_ssim),core/src/meson.build(addsinteger_ssim.cto the source list — fork-local diff),core/test/test_feature_extractor.c(adds one regression test alongside the existing tests),docs/metrics/features.md(table row + footnote ²),docs/state.md,changelog.d/fixed/ssim-extractor-registration.md. - Invariant on the upstream-mirror files: the registry-array entry must remain inside the unconditional CPU block (the same block as
&vmaf_fex_float_ssim/&vmaf_fex_float_ms_ssim) —vmaf_fex_ssimis CPU-only with no SIMD or GPU twin. Theconfig.hinclude ininteger_ssim.cis load-bearing on Vulkan-enabled LTO builds becausefeature_extractor.candinteger_ssim.cmust agree onHAVE_VULKAN/HAVE_CUDA/HAVE_SYCLfor theVmafFeatureExtractorstruct layout to match across TUs. - On upstream sync: if Netflix ever lands its own integer-SSIM registry row, drop the fork's row in favour of upstream's; the file structure is identical. If upstream removes
integer_ssim.centirely (the file has been dormant on master for years), revert the meson.build addition. Otherwise no action. - Re-test on rebase:
meson setup build -Denable_cuda=false -Denable_sycl=false && ninja -C build
./build/test/test_feature_extractor # 5/5 pass, includes new ssim row
./build/tools/vmaf --reference testdata/ref_576x324_48f.yuv \
--distorted testdata/dis_576x324_48f.yuv \
--width 576 --height 324 --pixel_format 420 --bitdepth 8 \
--feature ssim --output /tmp/ssim_smoke.json && \
grep -q '<metric name="ssim"' /tmp/ssim_smoke.json
# Vulkan-enabled LTO build (-Wlto-type-mismatch must stay clean)
meson setup build-vulkan -Denable_vulkan=enabled --reconfigure && \
ninja -C build-vulkan tools/vmaf
CI paths-ignore deny-list on heavy workflows (ADR-0341, 2026-05-09)¶
- Touches:
.github/workflows/libvmaf-build-matrix.yml(fork-local —paths-ignore:block underpull_request:),.github/workflows/tests-and-quality-gates.yml(fork-local — same block),docs/adr/0341-ci-paths-ignore-doc-only-prs.md+ index fragment,changelog.d/changed/ci-paths-ignore-doc-only.md. - Invariant: the deny-list must stay strictly documentation-only (
docs/**,**/*.md,changelog.d/**,CHANGELOG.md,.workingdir2/**). Any path that contributes to a build, test, or lint input —libvmaf/**,meson.build,meson_options.txt,subprojects/**,python/**,ai/**,mcp-server/**,model/**,testdata/**,.github/workflows/**— must NEVER appear in the deny-list, otherwise the corresponding required check is silently skipped on a code-touching PR. The Required Checks Aggregator (ADR-0313) catches only the doc-only case (no required check ever ran for any required name); a too-broad deny-list would lose build coverage without anyone noticing. - On upstream sync: Netflix/vmaf upstream does not carry these two workflow files (they are fork-local additions). No sync conflict expected.
- Re-test on rebase:
HDR VMAF model search — Path C documentation only (2026-05-09)¶
- Files added (this fork only; upstream Netflix/vmaf has none of these):
model/vmaf_hdr_model_card.md— discoverable warning that the HDR scoring path falls back to the SDRvmaf_v0.6.1.jsonweights. Filename deliberately uses.md, not.json, so thevmaftune.hdr.select_hdr_vmaf_modelglob (vmaf_hdr_*.json) keeps returningNone.docs/research/0089-hdr-vmaf-model-search.md— verbatim trail of the source-or-train survey (URLs + access dates).changelog.d/added/hdr-vmaf-model-search.md— release-notes fragment per ADR-0221.- ADR-0300 grew an inline
### Status update 2026-05-09: HDR model statussection. - Why no model JSON ships: Path A negative findings (no public Netflix HDR VMAF model exists; HDRMAX is a different algorithm not loadable by libvmaf's JSON path). Path B deferred behind gated subjective HDR corpora + multi-day training compute. No fabricated weights are introduced.
- On upstream sync: if Netflix lands
vmaf_hdr_*.jsoninNetflix/vmaf/model/, port via/port-upstream-commit; the resolver picks it up automatically with novmaftunechange. Then deletemodel/vmaf_hdr_model_card.md(or rewrite it as a normal model card describing the upstream weights). Watch https://github.com/Netflix/vmaf/issues/645 for the upstream release announcement. - Re-test on rebase: no behavioural change — pure docs. Sanity:
python3 -c "from pathlib import Path; \
import sys; sys.path.insert(0,'tools/vmaf-tune/src'); \
from vmaftune.hdr import select_hdr_vmaf_model; \
print(select_hdr_vmaf_model(Path('model')))"
# Expect: None — confirms the .md card does not match the glob
ADR-0349 — fr_regressor_v3 namespace resolution (2026-05-09)¶
- Rebase impact: none. Docs-only change — adds ADR-0349, an append-only status appendix on ADR-0302 per ADR-0028, a
## fr_regressor_* namespace mapblock inai/AGENTS.md, and two changelog fragments. No upstream Netflix/vmaf surface touched; nofr_regressor_*registry rows touched (sha256s for_v1,_v2,_v2_ensemble_v1_seed{0..4},_v3all unchanged); no C / Python / ONNX bytes modified. - What to check after a rebase: nothing automated. The only drift risk is a future agent claiming
fr_regressor_v3plus_featuresfor an unrelated workstream —ai/AGENTS.mdcarries the reservation; reviewers verify the map row exists before approving any newfr_regressor_*registry id. - Reproducer:
```bash # ADR + AGENTS.md namespace map present and consistent: test -f docs/adr/0349-fr-regressor-v3-namespace.md grep -q "fr_regressor_* namespace map" ai/AGENTS.md grep -q "fr_regressor_v3plus_features" ai/AGENTS.md docs/adr/0349-fr-regressor-v3-namespace.md # Status appendix present on ADR-0302: grep -q "Status update 2026-05-09: namespace collision resolved" \ docs/adr/0302-encoder-vocab-v3-schema-expansion.md # Existing v3 production row bit-identical (sha256 unchanged): python3 -c "
import json reg = json.load(open('model/tiny/registry.json')) v3 = next(m for m in reg['models'] if m['id'] == 'fr_regressor_v3') assert v3['sha256'] == 'eaa16d23461eda74940b2ed590edfcaf13428aade294e47792a5a15f4d3b999c', v3 assert v3['smoke'] is False print('OK: fr_regressor_v3 production row unchanged') "
Registry test still passes:¶
bash core/test/dnn/test_registry.sh
0327 — Pre-push PR-body deliverables validator hook¶
- Touches:
scripts/ci/validate-pr-body.sh(new),scripts/git-hooks/pre-push(new),scripts/ci/test-validate-pr-body.sh(new),Makefile(hooks-installtarget adds the pre-push symlink). Re-usesscripts/ci/deliverables-check.shparser verbatim — no upstream-shared file is modified. - Invariant: parser shape parity with
.github/workflows/rule-enforcement.ymldeep-dive-checklist gate (ADR-0108). The validator constructs aPATHshim that interceptsgit diff --name-onlycalls only; every othergitinvocation falls through to the real binary. - On upstream sync: not applicable — these files are entirely fork-local and Netflix has no equivalent. If
scripts/ci/deliverables-check.shis ever rewritten or moved, the validator's exec path (scripts/ci/deliverables-check.sh) and the test harness's expected exit codes must follow. bash scripts/ci/test-validate-pr-body.sh # 8/8 cases pass
0320 — Semgrep # nosemgrep cites on Netflix-upstream Python harness (Research-0090)¶
- Touches:
python/vmaf/core/asset.py,python/vmaf/core/executor.py,python/vmaf/core/feature_extractor.py,python/vmaf/core/quality_runner.py,python/vmaf/core/result_store.py,python/vmaf/tools/decorator.py,python/test/command_line_test.py,python/test/feature_extractor_test.py,python/test/ssimulacra2_test.py,python/vmaf/config.py. - Invariant: every fork-added
# nosemgrep: <rule-id>line is paired with an inline cite toResearch-0090. The cite + rule-id pair is the load-bearing artifact (per memoryfeedback_no_guessing: every "false positive" claim ships its safety proof). If an upstream sync removes the cited line of code, drop the cite-comment block too. If upstream adds adefusedxmlfix at theElementTree.parse()site (feature_extractor.py:115,quality_runner.py:1496), keep upstream's fix and drop our suppressions. config.py:40(the SSL-bypass deletion) is a fork-exclusive security fix; if upstream resurrectsssl._create_unverified_contexton a sync, do not re-merge it — the bypass clobbers the process-global default and is unjustified per Research-0090, F1. semgrep scan --config=p/cwe-top-25 --config=p/c --config=p/python . \ --metrics=off --json | jq '.results | length'
# expect 0 — every legit finding either has a # nosemgrep cite or was fixed
0321 — Security-scans workflow registry-pack list (Research-0090)¶
- Touches:
.github/workflows/security-scans.yml,.github/workflows/lint-and-format.yml. - Invariant: the registry packs the workflow cites (
p/cwe-top-25+p/c+p/python) are validated againsthttps://semgrep.dev/c/p/<pack>— the previously-citedp/cert-c-strict,p/cert-cpp-strict, andp/cpppacks were retired by Semgrep in 2025 and 404. Thelint-and-format.ymlpull of${{ github.* }}intoenv:(clang-tidy + clang-tidy-sycl steps) defusesrun-shell-injection; preserve the pattern on any edit. See Research-0090, F2/F3. for pack in p/cwe-top-25 p/c p/python; do code=\((curl -sIL "https://semgrep.dev/c/\)" | head -1 | awk '{print $2}') [ "$code" = "200" ] && echo "\({pack}: OK" || echo "\): FAIL ($code)"
0320 — CodeQL C bulk sweep (78 deferred alerts → 60 fixed, 14 deferred to T7-5)¶
- Touches:
core/src/feature/{cambi.c,ciede.c,integer_adm.c,integer_psnr.c,adm_tools.h,third_party/xiph/psnr_hvs.c},core/src/feature/x86/{adm_avx2.c,adm_avx512.c,ansnr_avx2.c,ansnr_avx512.c,vif_avx2.c,vif_avx512.c},core/src/{pdjson.c,svm.cpp},core/test/{test_cpu.c,test_model.c},core/tools/{y4m_input.c,yuv_input.c,vmaf_bench.c}. All butvmaf_bench.care upstream-mirror Netflix files. - Invariant: widening casts on integer multiplications (
(size_t),(uint64_t),(double)) are LHS-prefixed before the multiply, never wrapped around the whole expression — the latter is a no-op againstcpp/integer-multiplication-cast-to-long. Deleted commented-out blocks (e.g., the AVX-512 VP-loop dead variant inadm_avx512.c::adm_dwt2_inverse) are gone for good; if upstream brings them back, they reintroduce the alerts.iqa/convolve.cwas deliberately left untouched: prefixing(double)on the float×float multiplications inside the scalar reference path breaks bit-exactness against the AVX2 path enforced bytest_iqa_convolve— CodeQL alert deferred to a follow-up that updates both paths in lockstep. - On upstream sync: any upstream change that re-introduces the deleted comment blocks or rewrites the cast forms will surface the alerts again. The
cambi_scoresignature change (CambiBuffers buffers→const CambiBuffers *buffers) is fork-local and likely to conflict with upstream patches that touch that function. The 14 deferredVifBufferlarge-parameter alerts are tracked under T7-5 (multi-backend coordinated refactor including NEON). - Re-test on rebase: cd libvmaf && meson test -C build # all 50+ C tests make test-netflix-golden # upstream golden gate
# Re-run CodeQL on master afterwards; the 60 fixed alerts must stay closed.
CodeQL cpp/declaration-hides-variable sweep (2026-05-09)¶
- What changed: Mechanical rename / scope-tighten / dedupe sweep closing 64 open
cpp/declaration-hides-variableCodeQL alerts onmaster. Touched files:core/src/feature/cambi.c,core/src/feature/x86/adm_avx2.c,core/src/feature/x86/adm_avx512.c,core/src/feature/x86/vif_avx2.c,core/src/feature/x86/vif_avx512.c. All five are upstream-mirror; the Netflix copyright header is preserved on each. - Renames adopted (semantic over
_2suffix): cambi.c: innerint errshadowing function-scopeerrbecomesmkdir_err(heatmaps init) andsrc_err(full-ref extract path).adm_avx2.c/adm_avx512.c: thej == 0first-column special-case block is wrapped in{ ... }so itsj0..j3ands0..s3stop being visible to the per-jtail loop. The inner duplicate__m256i add_shift_HP_vex = _mm256_set1_epi32(32768)(and 512-bit twin) is removed — bit-identical to the function-scope value already in scope. The__m256i rfactor1that shadowed the function-scopefloat rfactor1[3]becomesrfactor_v0/_v1/_v2(and the AVX-512 twin likewise).vif_avx2.c/vif_avx512.c: tap-loop locals followf_tap,r_top/r_bot,d_top/d_botfor the s0 stage, andf_tap0/f_tap1,r_back0/r_fwd0, etc. for the AVX-512 paired-tap stage. Inner per-fj__m256i fq/__m512i fqshadows of the centre-tap broadcast becomef_tap. Inner-block duplicates of function-scoperef/dis/stride/ii(identical types and initialisers) are simply removed. The two scalarVifResiduals residualsdeclarations that shadowed function-scopeResiduals512 residualsbecometail_residuals. The twoconst uint16_t fcoeffdeclarations that shadowed function-scope__m512i fcoeffbecomefcoeff_scalar.- Invariant: bit-exactness gate — the rename sweep must not change any score. The Netflix CPU golden 3 (
src01_hrc00,checkerboard_1,checkerboard_10) ran clean against this PR. All 76 VMAF-targeted Python tests pass; the 9 unrelated pre-existing failures (NIQE, PyPSNR, FileSystemResultStore) reproduce on a pristineorigin/mastercheckout. - On upstream sync: Netflix has no equivalent renames on upstream
masteras of2026-05-09. When syncing, prefer the fork's renamed identifiers (the CodeQL gate depends on them). If Netflix later renames the same locals differently, reconcile by keeping fork names and updating any imported chunks at port time. - Re-test on rebase: meson test -C build --suite=fast PYTHONPATH=$PWD/python python3 -m pytest \ python/test/quality_runner_test.py -k test_run_vmaf \ python/test/vmafexec_test.py \ python/test/vmafexec_feature_extractor_test.py \ -m "not slow" -q
ADR-0209 v1 stdio runtime (T5-2b) — Embedded MCP server (2026-05-08)¶
- Touches:
core/src/mcp/{mcp.c,dispatcher.c,transport_stdio.c,mcp_internal.h,meson.build,3rdparty/cJSON/{cJSON.c,cJSON.h,LICENSE}},core/test/test_mcp_smoke.c,core/test/meson.build. All paths are fork-local. cJSON is vendored verbatim from upstreamDaveGamble/cJSON@v1.7.18under its MIT license. - Invariant: every TU under
core/src/mcp/(other than the vendored cJSON dir) is fork-local with theCopyright 2026 Lusoris and Claude (Anthropic)header; cJSON keeps its upstream MIT header verbatim. The public ABI incore/include/libvmaf/libvmaf_mcp.his unchanged from T5-2 — only function bodies flipped from-ENOSYSto working implementations. SSE / UDS still return-ENOSYSso the v2 PR can wire them without touching the public surface. - On upstream sync: no action required. Netflix/vmaf upstream has no embedded MCP surface; the entire
core/src/mcp/subtree is fork-local. If upstream ever adds an MCP surface, expect a port-only sync since names will collide. cd libvmaf && meson setup build -Denable_cuda=false -Denable_sycl=false \ -Denable_mcp=true -Denable_mcp_stdio=true ninja -C build && meson test -C build test_mcp_smoke -v
ADR-0334 — state.md-touch-check CI gate (2026-05-08)¶
- Touches:
.github/workflows/rule-enforcement.yml(new top-level jobstate-md-touch-check),scripts/ci/state-md-touch-check.sh(new),scripts/ci/test-state-md-touch-check.sh(new),scripts/ci/AGENTS.md(new rebase-sensitive-surface row),.github/PULL_REQUEST_TEMPLATE.md(already carries the "Bug-status hygiene" section +no state delta: REASONopt-out — coupled to the script's regex). No upstream-shared paths. - Invariant: the gate's trigger predicate (Conventional-Commit
fix:prefix, barebugtoken in title, GitHub close-keywordscloses/fixes/resolves#N, unchecked Bug-status-hygiene checkbox) and opt-out sentinel (no state delta: REASON) match the wording of the## Bug-status hygienesection in.github/PULL_REQUEST_TEMPLATE.md. Reword the template only alongside the script. The job carries thepull_request.draft == false || github.event_name != 'pull_request'gate (ADR-0331 pattern) — keep that on any future hoist into the required-aggregator set. - On upstream sync: Netflix/vmaf has no equivalent rule. No conflict expected; the workflow file is fork-introduced.
- Re-test on rebase: bash scripts/ci/test-state-md-touch-check.sh python3 -c "import yaml; yaml.safe_load(open('.github/workflows/rule-enforcement.yml')); print('YAML OK')" pre-commit run shellcheck --files scripts/ci/state-md-touch-check.sh scripts/ci/test-state-md-touch-check.sh pre-commit run shfmt --files scripts/ci/state-md-touch-check.sh scripts/ci/test-state-md-touch-check.sh
SYCL PSNR chroma extension (T3-15(b), 2026-05-09)¶
- Touches:
core/src/feature/sycl/integer_psnr_sycl.cpp(per-extractor chroma device buffers, per-plane SSE accumulators, and aprovided_featuresextension topsnr_y/psnr_cb/psnr_cr),core/src/sycl/AGENTS.md(per-kernel rebase-sensitive invariant for the chroma-on-per-extractor-buffer arrangement),docs/metrics/features.md(footnote ¹ refresh — all three GPU PSNR extractors now emit chroma),docs/adr/0192-gpu-long-tail-batch-3.mdReferences-section status update,changelog.d/added/sycl-psnr-chroma.md. - Invariant on the chroma upload path: chroma planes ride on per-extractor device buffers populated by host-side staging copies in the combined-graph
pre_fncallback — NOT the SYCL state's shared frame buffer (vmaf_sycl_shared_frame_init), which is luma-only by design. Luma stays graph-recorded; chroma SSE kernels run direct inpost_fnon the same in-order combined queue. The CUDA twin (PR #520 / commit 7f3d58a5) uses the existing CUDA per-plane picture infrastructure and therefore has no equivalent invariant. - On upstream sync: Netflix/vmaf upstream has no SYCL backend at all, so conflict probability is zero on
psnr_sycl. If an upstream port to the fork's SYCL runtime someday extendsvmaf_sycl_shared_frame_initto allocate chroma planes, the PSNR extension can be migrated onto it and the per-extractor chroma buffers retired — but only after a cross-backend gate run confirms bit-exactness against CPU atplaces=4(ADR-0214). source /opt/intel/oneapi/setvars.sh CC=icx CXX=icpx meson setup build-sycl libvmaf \ -Denable_sycl=true -Denable_cuda=false ninja -C build-sycl python3 scripts/ci/cross_backend_vif_diff.py \ --vmaf-binary build-sycl/tools/vmaf \ --reference testdata/ref_576x324_48f.yuv \ --distorted testdata/dis_576x324_48f.yuv \ --width 576 --height 324 --pixel-format 420 --bitdepth 8 \ --feature psnr --backend sycl --device 0
# Expect 0/48 mismatches across psnr_y / psnr_cb / psnr_cr at places=4.
```text
Cppcheck nullPointer false-positive in dict.c (2026-05-09)¶
Files pinned:
core/src/dict.c:121(one-line redundant-condition fix indict_overwrite_existing). Why this rebase-note exists: Master CI'sCppcheck (Whole Project)gate started failing on commit14b5ffba(#537) and blocked every open PR because each PR rebases onto a broken master. The cppcheck finding was likely always present but masked bypaths-ignorefiltering on the prior workflow shape; PR #530 widened cppcheck's trigger surface and exposed it. Deleted the redundant&& valguard sincevalis already checked at the public entry-pointvmaf_dictionary_set(dict.c:137). No behavior change; cppcheck flags the original as "either the val check is redundant or there's a possible null deref" because it can't prove the interprocedural guarantee. Rebase-sensitivity: zero — change is local todict.c. Future upstream sync of this file should keep the fix or re-run cppcheck locally to confirm absence of recurrence.
Aggregator timeout bump (2026-05-09)¶
Files pinned:
.github/workflows/required-aggregator.yml(deadline 30→90 min, job timeout 35→100 min) Why: 41 PRs in flight 2026-05-09 morning hit Aggregator timeouts while real CI eventually passed. Bumping both deadlines unblocks the train without touching the underlying matrix. Rebase-sensitivity: zero — workflow file is wholly fork-local.
ARC self-hosted runner pool — pilot Cppcheck routing (2026-05-09)¶
.github/workflows/lint-and-format.yml(Cppcheckruns-on:ternary). Why: opt-in graceful migration; ADR-0359 + docs/development/ci-runners.md document the flip-the-variable recipe when the cluster is degraded. Rebase-sensitivity: zero — workflow file is fork-local.
ADR-0338 — macOS Vulkan-via-MoltenVK CI lane (2026-05-09)¶
- Touches:
.github/workflows/libvmaf-build-matrix.yml(fork-local — addsBuild — macOS Vulkan via MoltenVK (advisory)lane, addscontinue-on-errorplumbing onmatrix.experimental && matrix.moltenvk, addsInstall MoltenVK + Vulkan loader/headers (macOS)step, addsRun Vulkan smoke tests (macOS MoltenVK)step, gates the existing test/cache/tox steps on!matrix.moltenvk),docs/backends/vulkan/moltenvk.md(new fork-local doc),docs/adr/0127-vulkan-compute-backend.md(status-update appendix per the ADR's Proposed status — body untouched),docs/adr/0338-macos-vulkan-via-moltenvk-lane.md(new),docs/adr/_index_fragments/0338-macos-vulkan-via-moltenvk-lane.mdplus_order.txtappend (new),docs/research/0089-moltenvk-feasibility-on-fork-shaders.md(new),changelog.d/added/macos-vulkan-via-moltenvk-lane.md(new). - Invariant on the upstream-mirror file: none —
libvmaf-build-matrix.ymlis fork-local. The new lane'scontinue-on-errorclause MUST stay scoped tomatrix.experimental == true && matrix.moltenvk == trueso existingexperimental: truematrix entries (e.g. the macOS DNN lane) keep their default fail-fast behaviour.VK_ICD_FILENAMESMUST point at/opt/homebrew/etc/vulkan/icd.d/MoltenVK_icd.json— note theetc/vulkansegment, NOTshare/vulkan(the homebrew formula's install layout usesetc/; verified againstFormula/m/molten-vk.rb). - On upstream sync: Netflix upstream has no macOS Vulkan lane and no MoltenVK awareness; nothing to reconcile. If a future MoltenVK release drops support for
GL_EXT_shader_atomic_int64translation,moment.compwill fail on the lane; the fix path is in ADR-0338 §Decision (lane iscontinue-on-errorso it does not block PRs) — update the known-limitations table indocs/backends/vulkan/moltenvk.mdand either pin a working MoltenVK version in the brew install line or rewrite the shader. - Re-test on rebase:
python3 -c "import yaml; yaml.safe_load(open('.github/workflows/libvmaf-build-matrix.yml'))" && \
echo "YAML parse OK"
# Confirm the lane is still in the matrix:
grep -q "Build — macOS Vulkan via MoltenVK (advisory)" \
.github/workflows/libvmaf-build-matrix.yml
# Confirm the lane is NOT promoted to required-aggregator until one
# green run on master (per ADR-0338):
! grep -q "macOS Vulkan via MoltenVK" \
.github/workflows/required-aggregator.yml
# Confirm the ICD path is the etc/ one, not share/:
grep -q "etc/vulkan/icd.d/MoltenVK_icd.json" \
.github/workflows/libvmaf-build-matrix.yml
ADR-0363 — Mend Renovate replaces Dependabot (2026-05-09)¶
- Touches:
renovate.json(new, repo-root),.github/workflows/renovate.yml(new),.github/dependabot.yml(deleted — renamed to.github/dependabot.yml.disabled),docs/development/dependency-bot.md(new operator playbook),changelog.d/changed/renovate-supersedes-dependabot.md(new),docs/adr/0363-renovate-replaces-dependabot.md(new),docs/adr/_index_fragments/0363-renovate-replaces-dependabot.md(new). - Invariant:
.github/dependabot.ymlno longer exists onmaster; the disabled copy isdependabot.yml.disabled. On upstream sync, if Netflix ever ships their owndependabot.yml, do NOT restore it — the fork intentionally uses Renovate. Merge the upstream file intodependabot.yml.disabledfor reference only. - Upstream interaction: none. Netflix/vmaf upstream has no Renovate config. Conflict risk is zero unless upstream adds
renovate.jsonor restoresdependabot.yml. - Re-test on rebase:
# Verify the workflow SHA-pin is still present and non-floating:
grep -E 'renovatebot/github-action@[a-f0-9]{40}' .github/workflows/renovate.yml
# Verify dependabot.yml is still absent:
test ! -f .github/dependabot.yml && echo "ok: dependabot.yml absent"
# Validate renovate.json syntax (requires Node):
node -e "JSON.parse(require('fs').readFileSync('renovate.json','utf8')); console.log('JSON valid')"
ADR-0355 — Symphony-inspired agent-dispatch infrastructure (2026-05-09)¶
Files added (all fork-introduced, none mirror upstream):
.claude/workflows/_template.md,.claude/workflows/codeql-alert-sweep.md,.claude/workflows/simd-port.md,.claude/workflows/feature-extractor-port.md.scripts/lib/__init__.py,scripts/lib/backlog_tracker.py,scripts/lib/AGENTS.md.scripts/ci/agent-eligibility-precheck.py(new row inscripts/ci/AGENTS.md"Rebase-sensitive surfaces" table).docs/development/agent-dispatch.md. Why this rebase-note exists: pure additive, all paths are fork-only (.claude/,scripts/lib/, fork-only docs). Upstream Netflix/vmaf has no.claude/, noscripts/lib/, and nodocs/development/agent-dispatch.md, so the merge surface is zero on/sync-upstream. The only coupling is internal betweenscripts/ci/agent-eligibility-precheck.pyandscripts/lib/backlog_tracker.py(sys.path import). Both files move together; documented inscripts/lib/AGENTS.mdand a new row inscripts/ci/AGENTS.md. Rebase-sensitivity: zero w.r.t. upstream. Internal-only: renamingBacklogItemfield names or theBacklogTracker/GitHubTrackerpublic method signatures is a breaking change for the precheck and any future state-audit script — guard via the smoke listed in Research-0091 §"Smoke results" before any rename PR. Format-coupling note: the BACKLOG.md row regex (scripts/lib/backlog_tracker.py:_ID_PATTERN) is brittle against table-shape edits. If a future BACKLOG.md edit adds a column or renames a status word, the parser will silently mis-classify rows — the smoke parses 101 rows on master at 2026-05-09; expect ≥ 100 after any structural edit.
0350 — psnr_hvs AVX-512 ceiling re-bench (ADR-0350, T3-9 (a))¶
docs/adr/0350-psnr-hvs-avx512-ceiling.md— closure ADR.docs/adr/0160-psnr-hvs-neon-bitexact.md— appended### Status update 2026-05-09appendix.docs/research/0091-psnr-hvs-avx512-bench-2026-05-09.md— empirical companion (cycle share, Amdahl ceiling, reproducer). Why this rebase-note exists: T3-9 (a) closes as AVX2 ceiling. The result has zero rebase-sensitivity by itself — no engine code changes — but the bit-exactness invariants that lock it to a ceiling do. The 78.42 % scalar tail incalc_psnrhvs_avx2/calc_psnrhvs_neonis locked by ADR-0138 / ADR-0139's "per-lane-scalar float reduction" rule (carried by ADR-0159 / ADR-0160). If a future upstream sync ofcore/src/feature/third_party/xiph/psnr_hvs.c(the Xiph/Daala DCT) changes the per-block summation tree — e.g. partial folding, re-ordered means, vectorised mask reductions — the AVX2 + NEON TUs incore/src/feature/x86/psnr_hvs_avx2.candcore/src/feature/arm64/psnr_hvs_neon.cMUST be re-audited against the new scalar reference, and the ceiling argument in ADR-0350 must be re-run (because the 78 / 15 cycle-share split would shift). Rebase-sensitivity: low for the ceiling decision itself (empirical re-bench on a current host is cheap — 30 seconds via the reproducer in Research-0091 §7); high for the underlying bit-exactness invariants the decision rests on (Netflix golden trips on ≥ 5.5e-5 drift per ADR-0160 §Context). The ADR-0350 §Verification reproducer is the gate — re-run it if the cycle share shifts, the Netflix normal-pair fixture changes, or a new host class (e.g. wide-issue Granite Rapids) goes into CI.
0320 — FFmpeg n8.1 → n8.1.1 base bump (2026-05-09)¶
- Touches:
ffmpeg-patches/series.txt(header comment),ffmpeg-patches/README.md(apply / verify / smoke sections),ffmpeg-patches/test/build-and-run.sh(FFMPEG_SHAdefault),scripts/ci/ffmpeg-patches-check.sh(header comment;FFMPEG_BRANCHenv default unchanged atrelease/8.1since the branch tracks point releases),docs/development/automated-rule-enforcement.md(gate description). The 9.patchfiles themselves are unchanged — every patch in the series applied cleanly, cumulatively, against pristinen8.1.1viagit am --3way. - Upstream source: FFmpeg upstream point release n8.1.1 (commit
239f2c7"Bump micro for 8.1.1") — bug-fix-only on top of n8.1, no API or AVOption breakage that the patch stack consumes. - Invariant: the patch stack continues to apply against the current tip of FFmpeg's
release/8.1branch. Per ADR-0118 and ADR-0186 §FFmpeg patch coupling, the verification gate is cumulativegit am --3wayagainst a pristine checkout, not per-patch standalone apply. The scripts/ci/ffmpeg-patches-check.sh local gate usesgit apply(no commit) but accumulates state in the same way. - On upstream sync: no action required. If a future FFmpeg point release (n8.1.2 or n8.2) lands new hunks that conflict with one of the patches, regenerate the affected patches via
git format-patchon the resolved state, bump the references in the five files listed under "Touches", and add a fresh rebase-notes entry citing the conflict file(s). - Re-test on rebase:
cd /tmp && rm -rf ffmpeg-n811 && \
git clone --depth 1 --branch n8.1.1 \
https://git.ffmpeg.org/ffmpeg.git ffmpeg-n811
git -C /tmp/ffmpeg-n811 config user.email agent@local
git -C /tmp/ffmpeg-n811 config user.name agent
for p in ffmpeg-patches/000*-*.patch; do
git -C /tmp/ffmpeg-n811 am --3way "$p" || break
done
bash scripts/ci/ffmpeg-patches-check.sh
ADR-0281 follow-up — QSV install-matrix discoverability backfill (2026-05-08)¶
- Touches:
docs/getting-started/install/{arch,fedora,ubuntu,macos,windows}.md(new## Intel QSVsection per page),docs/adr/0281-vmaf-tune-qsv-adapters.md(status-update appendix per ADR-0028),changelog.d/changed/qsv-install-matrix-docs.md(new fragment). No code, no engine, no upstream-shared C / Python source touched. Pure documentation backfill closing the SYCL-audit research-0086 Topic C gap (issue #464). - Invariant: each per-OS QSV section pins the package names against verified upstream URLs with a
Verified 2026-05-08access date. The hardware-generation matrix is sourced from the public Wikipedia "Intel Quick Sync Video — Hardware decoding and encoding" table; if Intel revises which generation supports AV1 encode (e.g. backports the encoder to Lunar Lake / Meteor Lake silicon currently absent from the table), the matrix in all five pages must move in lockstep — the Arch / Fedora / Ubuntu / Windows pages all carry the same matrix verbatim. The macOS page deliberately omits the matrix (QSV unsupported on macOS). - On upstream sync: no action required — Netflix/vmaf upstream does not ship per-OS install pages under
docs/getting-started/install/; that tree is fork-only.
# Lint the install pages (markdownlint via pre-commit):
pre-commit run --files docs/getting-started/install/*.md
# Verify each page (except alpine + macos) still carries the matrix:
for f in arch fedora ubuntu windows; do grep -q 'Arc Battlemage' "docs/getting-started/install/${f}.md" || echo "MISSING: ${f}"
# Confirm the macOS page documents QSV as unsupported:
grep -q 'Intel QSV. is unsupported on macOS' docs/getting-started/install/macos.md
0333 — vmaf-tune Phase F multi-pass encoding (ADR-0333)¶
Touches:
tools/vmaf-tune/src/vmaftune/codec_adapters/__init__.py(CodecAdapter Protocol gainssupports_two_pass: bool+two_pass_args(...))tools/vmaf-tune/src/vmaftune/codec_adapters/x265.py(overrides both)tools/vmaf-tune/src/vmaftune/encode.py(EncodeRequestgainspass_number/stats_path;build_ffmpeg_commandadds the 2-pass argv splice + pass-1 null-muxer redirect; newrun_two_pass_encode)tools/vmaf-tune/src/vmaftune/corpus.py(CorpusOptions.two_pass, routing initer_rows)tools/vmaf-tune/src/vmaftune/cli.py(--two-passflag oncorpus/recommendsubparsers) Invariant: 2-pass encoding routes through the codec adapter viasupports_two_pass+two_pass_args(pass_number, stats_path). The encode driver never branches on codec name. Adapters withsupports_two_pass = Falseare honoured silently (single-pass fallback with stderr warning); the seam is open for sibling codec adapters (libx264, libsvtav1, libvvenc, libaom-av1) to opt in by overriding the two methods on their adapter file alone. This is the fork-local extension to the ADR-0237 Phase A multi-codec contract; upstream Netflix/vmaf has no equivalent and does not own this code path. Re-test:
(Optional, requires ffmpeg + libx265 in the runner's PATH:)
VMAF_TUNE_INTEGRATION=1 python -m pytest \
tests/test_codec_adapter_x265_two_pass.py::test_real_x265_two_pass_smoke -q
Rebase-sensitivity: zero from upstream — tools/vmaf-tune/ is fork-local. The only concern is the codec_adapters Protocol shape: a future upstream commit that adds a sibling codec adapter SHOULD inherit the supports_two_pass = False default and either explicitly opt in or leave the flag off. Downstream sibling-codec PRs in this fork should follow the ADR-0288 / ADR-0333 pattern: one adapter file, override the two methods, add a test file mirroring test_codec_adapter_x265_two_pass.py.
ADR-0360 — CAMBI CUDA port (T3-15a, 2026-05-09)¶
Files pinned:
core/src/feature/cuda/integer_cambi_cuda.c(new)core/src/feature/cuda/integer_cambi_cuda.h(new)core/src/feature/cuda/integer_cambi/cambi_score.cu(new)core/src/feature/feature_extractor.c(addedvmaf_fex_cambi_cudato list)core/src/meson.build(addedcambi_scoretocuda_cu_sources, addedinteger_cambi_cuda.cto CUDA feature sources)
Why: The CUDA twin of vmaf_fex_cambi (Strategy II hybrid — three GPU kernels for the embarrassingly parallel stages; calculate_c_values + topK on CPU). Registers vmaf_fex_cambi_cuda under #if HAVE_CUDA guard.
Rebase-sensitivity: low. The three new files are wholly fork-local and will not conflict. The two upstream-shared files have small, self-contained hunks:
feature_extractor.c: theextern vmaf_fex_cambi_cudadeclaration and the&vmaf_fex_cambi_cudaarray entry are inside a#if HAVE_CUDAblock. Upstream's additions to this file (new feature extractors, new dispatch flags) will not conflict unless Netflix adds their own CUDA twin for CAMBI (unlikely — they don't ship a CUDA backend).meson.build: thecambi_scoreentry in thecuda_cu_sourcesdict and theinteger_cambi_cuda.cline in the CUDA sources list. Any upstream changes tomeson.buildthat restructure thecuda_cu_sourcesdict would require a manual merge; the dict entries are sorted alphabetically by key, socambi_scorelands betweenadm_scoreandmotion_score.
If upstream adds cambi_cuda themselves: drop the fork copy and check for API divergence. Strategy II hybrid is the natural choice; the upstream implementation may differ if they choose Strategy III (fully-on-GPU calculate_c_values).
cambi_internal.h dependency: integer_cambi_cuda.c includes core/src/feature/cambi_internal.h (fork-added trampoline exposing cambi.c's static helpers). If upstream significantly refactors cambi.c (renames vmaf_cambi_preprocessing, vmaf_cambi_calculate_c_values, etc.), cambi_internal.h must be updated alongside. This is the same dependency the Vulkan twin (cambi_vulkan.c) has — see ADR-0210's rebase note for the full list of exposed functions.
Vulkan submit-pool PR-B: six secondary kernels (2026-05-09, ADR-0353)¶
Files changed:
core/src/feature/vulkan/ssim_vulkan.ccore/src/feature/vulkan/ciede_vulkan.ccore/src/feature/vulkan/ms_ssim_vulkan.ccore/src/feature/vulkan/motion_v2_vulkan.ccore/src/feature/vulkan/float_psnr_vulkan.ccore/src/feature/vulkan/float_motion_vulkan.ccore/src/feature/vulkan/AGENTS.mddocs/adr/0353-vulkan-submit-pool-pr-b-six-kernels.md
Why this rebase-note exists: six Vulkan host-glue TUs were migrated from per-frame command-buffer and descriptor-set allocation to the VmafVulkanKernelSubmitPool abstraction (ADR-0256). Any Netflix upstream sync that touches these same files (unlikely — they are fork-local) must preserve the VmafVulkanKernelSubmitPool fields in the state struct and the pool-destroy-before-pipeline-destroy ordering in close_fex().
Rebase-sensitivity: low. All six files are entirely fork-local; Netflix upstream does not have a Vulkan backend. The submit-pool API is defined in core/src/vulkan/kernel.h (also fork-local). No public header or C-API surface was changed; the FFmpeg patch series is unaffected.
Key invariant to preserve on rebase: vmaf_vulkan_kernel_submit_pool_destroy MUST be called before vmaf_vulkan_kernel_pipeline_destroy in every migrated kernel's close_fex(). See core/src/feature/vulkan/AGENTS.md §"Submit-pool ordering invariant".
0354 — Vulkan submit-pool PR-C: submit_pool_destroy-before-pipeline ordering¶
- Touches:
core/src/feature/vulkan/cambi_vulkan.c,core/src/feature/vulkan/ssimulacra2_vulkan.c,core/src/feature/vulkan/float_ansnr_vulkan.c,core/src/feature/vulkan/moment_vulkan.c. - Invariant: In every migrated extractor,
vmaf_vulkan_kernel_submit_pool_destroy()MUST precede everyvmaf_vulkan_kernel_pipeline_destroy()call inclose_fex(). Reversing the order frees the pool's command buffers after the pipeline's command pool is destroyed — undefined behaviour per Vulkan spec §6.2. - Re-test:
meson test -C build --suite=vulkanpasses.scripts/ci/cross_backend_vif_diff.pyshowsplaces=4for all four extractors on all three target devices (RTX 4090, Arc A380, RADV iGPU).
0231 — Vulkan submit-pool migration PR A: adm + motion + psnr (ADR-0291)¶
0231 — Vulkan submit-pool migration PR A: adm + motion + psnr (ADR-0352)¶
- Touches:
core/src/feature/vulkan/adm_vulkan.c,core/src/feature/vulkan/motion_vulkan.c,core/src/feature/vulkan/psnr_vulkan.c(all fork-local Vulkan kernels; no upstream C paths touched),changelog.d/changed/vulkan-submit-pool-pr-a-adm-motion-psnr.md,docs/adr/0291-vulkan-submit-pool-pr-a-adm-motion-psnr.md. - Invariant: Each migrated TU adds
VmafVulkanKernelSubmitPool sub_pooland pre-allocatedVkDescriptorSetfield(s) to its state struct. The pool must be destroyed (vmaf_vulkan_kernel_submit_pool_destroy) beforevmaf_vulkan_kernel_pipeline_destroyinclose_fex(); reversing the order would destroy the descriptor pool while the submit pool still holds live command buffer + fence references. Descriptor sets allocated viavmaf_vulkan_kernel_descriptor_sets_allocare freed implicitly by the descriptor pool tear-down — do NOT callvkFreeDescriptorSetson them inclose_fex(). Formotion_vulkan, the pre-allocated set is rebound once per frame viavkUpdateDescriptorSetsbecause the blur ping-pong changes whichblur[]slot is "current"; foradm_vulkanandpsnr_vulkanthe sets are stable afterinit()and require no per-frame update. - Upstream interaction: none. All three files are fork-local Vulkan kernel TUs not present in Netflix/vmaf upstream.
- On upstream sync: zero interaction. Upstream cannot conflict with this PR's paths. The Vulkan backend is entirely fork-introduced.
- Re-test on rebase:
meson test -C build --suite=fast
# Cross-backend parity gate (places=4):
python python/test/cross_backend_diff.py \
--features adm motion psnr \
--backend vulkan cpu \
--places 4 \
--yuv testdata/yuv/src01_hrc00_576x324.yuv \
testdata/yuv/src01_hrc01_576x324.yuv
ADR-0350 — FFmpeg libvmaf filter CUDA backend selector (0010 patch)¶
Patch: ffmpeg-patches/0010-libvmaf-wire-cuda-backend-selector.patch.
libavfilter/vf_libvmaf.c— addscudaAVOption + state field + init / cleanup / picture-pool wiring underCONFIG_LIBVMAF_CUDA && !CONFIG_LIBVMAF_CUDA_FILTER.configure— adds--enable-libvmaf-cuda(EXTERNAL_LIBRARY_LISTentry + help text), promoteslibvmaf_cudafrom blanket-autodetect to gatedenabled libvmaf_cuda && require_pkg_config + check, preserves theenabled libvmaf && check_pkg_config libvmaf_cudain-filter probe so the new selector still works without the explicit flag when libvmaf ships CUDA. Why this rebase-note exists: Patch0010extends the SYCL (0003) / Vulkan (0004) per-context backend selectors to CUDA on the regularlibvmaffilter. The patch coexists with the upstream dedicatedlibvmaf_cudafilter (CONFIG_LIBVMAF_CUDA_FILTER) by gating its struct field and code paths on!CONFIG_LIBVMAF_CUDA_FILTER— the dedicated filter keeps owning its owncu_statefield. CLAUDE.md §12 r14 makes the patch update mandatory because the change touches a filter consumer of thevmaf_cuda_state_init/_import_state/_state_free/_preallocate_pictures/_fetch_preallocated_pictureC-API surface inlibvmaf_cuda.h. Rebase-sensitivity: low. The patch'svf_libvmaf.chunks are context-anchored on the SYCL/Vulkan selector blocks; if upstream FFmpeg renamesCONFIG_LIBVMAF_CUDA_FILTERor moves thelibvmaf_cuda.hinclude, the include guard at the top of the file needs the corresponding update. The configure hunks are context-anchored on the existing--enable-libvmaf-sycl/--enable-libvmaf-vulkanlines — those have proven stable across n8.0 → n8.1 → n8.1.1, so drift risk is low. WhenVmafCudaConfigurationever grows adevice_indexfield upstream, swap thecudaboolean for anint cuda_devicemirroring SYCL's shape (separate ADR + patch refresh). Verification gate: cumulativegit am --3wayreplay offfmpeg-patches/000{1..9}-*.patch+0010-*against pristine FFmpegn8.1.1PASS (2026-05-09). Build oflibavfilter/vf_libvmaf.oPASS under bothCONFIG_LIBVMAF_CUDA=0(selector errors at filter- init time per#elsebranch) andCONFIG_LIBVMAF_CUDA=1 && !CONFIG_LIBVMAF_CUDA_FILTER(selector active, picture-pool wiring compiles).
0320 — Vulkan instance / VMA apiVersion bump to 1.4 (Step B)¶
- Touches:
core/src/vulkan/common.c,core/src/vulkan/vma_impl.cpp,core/src/vulkan/AGENTS.md. - Invariant: the four
apiVersionsites (lines 54, 264, 374 ofcommon.c; line 22 ofvma_impl.cpp) request Vulkan 1.4, not 1.3. Together with the Step-Aprecisedecorations invif.comp/ciede.comp(PR #346) and the Phase-3 cross-subgroup release-acquire fix (PR #511), this gates the cross-backend places=4 contract on Arc + RADV. NVIDIA closure depends on Phase 3c (PR #512; block-on-merge until that lands). Netflix upstream does not carry a VMA dependency or a Vulkan backend; no upstream merge conflict expected on these files. - Re-test on rebase:
meson setup build -Denable_vulkan=enabled -Denable_cuda=false \
-Denable_sycl=false --buildtype=release
ninja -C build
for D in 0 1 2; do
python3 scripts/ci/cross_backend_parity_gate.py \
--vmaf-binary build/tools/vmaf \
--reference python/test/resource/yuv/src01_hrc00_576x324.yuv \
--distorted python/test/resource/yuv/src01_hrc01_576x324.yuv \
--width 576 --height 324 --pixel-format 420 --bitdepth 8 \
--backends cpu vulkan --vulkan-device "$D" \
--features vif ciede adm motion psnr
done
# All 0/N mismatches at places=4 once Phase 3c (PR #512) has landed.
ADR-0332 v2 runtime (T5-2c) — Embedded MCP server UDS + real compute_vmaf (2026-05-09)¶
- Touches:
core/src/mcp/{mcp.c,dispatcher.c,mcp_internal.h,meson.build,compute_vmaf.c,transport_uds.c},core/test/test_mcp_smoke.c. All paths are fork-local. No new third-party vendor drop in v2 — mongoose vendoring stays deferred to v3 with the SSE transport. - Invariant: same as ADR-0209 v1 — the entire
core/src/mcp/subtree is fork-local; the public ABI incore/include/libvmaf/libvmaf_mcp.his unchanged (only function bodies flipped —vmaf_mcp_start_udsfrom-ENOSYSto a working AF_UNIX listener;compute_vmaffrom a{"status":"deferred_to_v2"}placeholder to a realvmaf_score_pooledbinding). Per ADR-0128 § operational guardrails the UDS socket file is created mode 0700; thatchmodhappens invmaf_mcp_start_udsafterbindand is a load-bearing security invariant — do NOT relax it on rebase.compute_vmafruns on a per-call ephemeralVmafContextso the host's main scoring run is unperturbed; do NOT rewire it to reuseserver->ctxbecausevmaf_score_pooledcommits the model destructively to the context. - On upstream sync: no action required. Netflix/vmaf upstream has no embedded MCP surface. If upstream adds one, expect a port-only sync since names will collide.
- Re-test on rebase:
cd libvmaf && meson setup build -Denable_cuda=false -Denable_sycl=false \
-Denable_mcp=true -Denable_mcp_stdio=true \
-Denable_mcp_uds=true
ninja -C build && meson test -C build test_mcp_smoke -v
# Real-score smoke (single 576x324 pair):
build/test/test_mcp_smoke 2>&1 | tail -3 # expects "16 tests run, 16 passed"
ADR-0332 v3 runtime (T5-2d) — Embedded MCP server SSE transport (2026-05-09)¶
- Touches:
core/src/mcp/{mcp.c,mcp_internal.h,meson.build,transport_sse.c},core/meson_options.txt,core/test/test_mcp_smoke.c,docs/mcp/embedded.md,docs/adr/0332-mcp-runtime-v2.md(status-update appendix). All paths are fork-local. No third-party vendor drop in v3 — the originally-planned mongoose vendor was reversed because cesanta/mongoose 7.18 is GPL-2.0-only OR commercial, incompatible with the fork's BSD-3-Clause-Plus-Patent license (verified at upstream LICENSE 2026-05-09). The SSE transport is plain POSIX sockets in fork-owned C (~500 LOC). - Invariant: same as ADR-0209 / ADR-0332 v2 — the entire
core/src/mcp/subtree is fork-local; the public ABI incore/include/libvmaf/libvmaf_mcp.his unchanged (onlyvmaf_mcp_start_sse's body flipped from-ENOSYSto a working AF_INET listener). The SSE listener bindsINADDR_LOOPBACKonly; do NOT switch toINADDR_ANYwithout a separate ADR + auth design (v3 ships intentionally without CORS/Bearer/per-session auth on the assumption of a same-host trust boundary). The SSE stop path usesshutdown(SHUT_RDWR)beforeclose()— plainclose()of an AF_INET listening fd from another thread does NOT unblockaccept()on Linux; do NOT remove theshutdowncall.enable_mcp_sseis now afeatureoption (defaultauto), notboolean false. - On upstream sync: no action required. Netflix/vmaf upstream has no embedded MCP surface. Do NOT re-introduce mongoose (or any GPL-licensed HTTP library) on a future rebase without first amending CLAUDE §1 and adding a separate license-compatibility ADR.
- Re-test on rebase:
cd libvmaf && meson setup build -Denable_cuda=false -Denable_sycl=false \
-Denable_mcp=true -Denable_mcp_stdio=true \
-Denable_mcp_uds=true \
-Denable_mcp_sse=enabled
ninja -C build && meson test -C build test_mcp_smoke -v
build/test/test_mcp_smoke 2>&1 | tail -3 # expects "17 tests run, 17 passed"
Status update 2026-05-09 — placeholder-ref hardening¶
- Additional touches: same set as the 2026-05-08 ADR-0334 entry, no new files. The hardening adds a
git diff -U0 ... -- docs/state.mdcall insidescripts/ci/state-md-touch-check.sh(case 4a) plus 10 additional fixture cases inscripts/ci/test-state-md-touch-check.sh. - New invariant: inserted lines in
docs/state.md(lines starting with+, excluding the+++ b/...header) must not containthis PR/this commit/ bareTBD/<PR>/#NNN. Canonical accept forms arePR #Nandcommit `<sha>`. The placeholder vocabulary is coupled to PR #541's audit findings — reword in lockstep with the ADR-0334 status-update appendix if the fork's row template changes. - Re-test on rebase: same
bash scripts/ci/test-state-md-touch-check.shrun as the 2026-05-08 entry; the harness now reports18/18 passed(was8/8 passed).
0347 — Sanitizer matrix test-set scope (ADR-0347)¶
- Touches:
.github/workflows/tests-and-quality-gates.ymljobsanitizers(build + test step),core/test/meson.build(no edits — the absence of anysuite: 'unit'tag is the upstream state we now work with rather than against). - Invariant: the sanitizer job runs the full C unit-test set per sanitizer with a per-sanitizer deselect list driven by a
caseblock on${{ matrix.sanitizer }}. The deselect lists are load-bearing — each entry corresponds to a real bug tracked indocs/state.md. Under UBSan the build adds-Dc_args=-fno-sanitize=function -Dcpp_args=-fno-sanitize=functionto suppress the K&R-prototype harness UB; the mesoncasebranch must keep this build flag in sync with the test deselect entries. An upstream rebase that adds new test files viacore/test/meson.buildinherits full sanitizer coverage automatically (the workflow enumerates tests viameson test --list). - On upstream sync: if upstream Netflix lands a
suite: 'unit'tagging convention, the workflow is robust to it (we already enumerate frommeson test --list, not from--suite=unit). If upstream rewrites the harness to declarestatic char *test_X(void)with a(void)parameter, the-fno-sanitize=functionflag becomes redundant — leave it in place (zero cost) until a deliberate cleanup PR reverts the suppression. If upstream lands a fix for any of the surfaced defects (SVMModelParservalidation,feature_collectormetadata leak,integer_adm::div_lookuprace,framesyncmutex mismatch), drop the corresponding deselect row from the workflow'scaseblock in the same PR that pulls the upstream fix. cd libvmaf for SAN in address undefined thread; do EXTRA=() [ "$SAN" = undefined ] && EXTRA=( "-Dc_args=-fno-sanitize=function" "-Dcpp_args=-fno-sanitize=function" ) rm -rf "build-$SAN" CC=clang CXX=clang++ LDFLAGS=-fuse-ld=lld \ meson setup "build-$SAN" -Db_sanitize="$SAN" \ -Denable_cuda=false -Denable_sycl=false --buildtype=debug \ -Db_lto=false -Db_lundef=false "${EXTRA[@]}" meson compile -C "build-$SAN" case "$SAN" in address) EXCLUDE='test_model$|test_predict$|test_float_ms_ssim_min_dim$' ;; undefined) EXCLUDE='test_model$' ;; thread) EXCLUDE='test_model$|test_pic_preallocation$|test_framesync$' ;; esac TESTS=$(meson test -C "build-$SAN" --list \ | grep '^libvmaf:' \ | grep -vE "$EXCLUDE" \ | sed 's/^libvmaf://') meson test -C "build-$SAN" --print-errorlogs $TESTS
CodeQL bulk mechanical sweep — Python tree (2026-05-09)¶
- Why this matters on rebase: no rebase impact. The diff lives entirely in
python/vmaf/and one fork-local helper (core/src/vulkan/spv_embed.py). None of the touched Python modules have been changed by Netflix upstream in over four years; the closest churn is unrelated additions topython/vmaf/script/run_*.pydriver flags. A future/sync-upstreamwill land on a clean tree. - What changed: dead imports removed;
exit()→sys.exit()in seven CLI driver scripts;open(...)→with open(...)inpython/vmaf/tools/decorator.pyandcore/src/vulkan/spv_embed.py; typedexcept KeyError: passbodies got an explanatory one-line comment to satisfypy/empty-except;passremoved where it was a no-op tail statement; one commented-out debug block deleted fromtools/misc.py. - Re-test on rebase:
python3 -c "import ast; [ast.parse(open(f).read()) for f in (...)]"over the touched files;ruff checkover the same set must produce no NEW errors versus master baseline.
0345 — cambi × {CUDA, SYCL, HIP} GPU port planning (ADR-0345, docs-only)¶
- Touches:
docs/research/0091-cambi-gpu-port-planning-2026-05-09.md(new),docs/adr/0345-cambi-gpu-port-strategy.md(new),docs/adr/_index_fragments/0345-cambi-gpu-port-strategy.md(new fragment),docs/adr/_index_fragments/_order.txt(append slot),changelog.d/changed/cambi-gpu-planning-digest.md(new). No code. Companion to the per-port PRs that follow per the digest's §6 ordered plan (CUDA → SYCL → HIP). - Upstream source: none — fork-local planning artefact. Netflix/vmaf upstream has no CUDA / SYCL / HIP cambi twin and no plans to add one on those backends.
- Invariant: the planning round locks Strategy II host-staged hybrid for the three pending backends, inheriting verbatim from ADR-0205 §Decision and ADR-0210 §Decision. The cross-backend gate contract for cambi is
places=4from day one on all backends — by construction (integer-only GPU pre-passes; byte-identical readback; unmodified host residual). If any per-port PR sees empirical drift from CPU, fix the kernel — never relax the gate (memoryfeedback_no_test_weakening). The sharedcambi_internal.hhost residual surface (shipped with PR #196 for the Vulkan port) is the load-bearing reuse point — all four GPU twins (Vulkan, CUDA, SYCL, HIP) link against it and inherit any future CPU-side c-value formula change automatically. - On upstream sync: no action required. If a future upstream sync introduces a Netflix/vmaf cambi GPU twin (extremely unlikely — Netflix has no public CUDA / SYCL / HIP cambi work), evaluate whether to drop the fork's twin in favour of upstream's per the standard prefer-upstream rule; otherwise no action.
- Re-test on rebase: docs-only — no compile / runtime gate. The Strategy III v2 follow-up (parked per ADR-0205 §Out of scope) gets its own ADR + rebase-notes entry when profile data lands.
0320 — Vulkan VIF API-1.4 NVIDIA residual Phase 3b (deferral)¶
- Touches:
core/src/feature/vulkan/shaders/vif.comp(comment-only update at the Phase-4 reduction site — documents the Phase-3b candidate-fix experiments and the driver-side hypothesis; no code logic change vs. PR #511);docs/adr/0269-vif-ciede-precise-step-a.md(appended Phase-3b status update appendix; ADR body remains frozen per ADR-0028);docs/research/0090-...md(new);docs/state.md(rowT-VK-VIF-1.4-RESIDUAL-ARCretired in favour ofT-VK-VIF-1.4-RESIDUAL-NVIDIA-DEFERREDafter the hardware-mapping correction);core/src/vulkan/AGENTS.md(Phase 3b update + rebase invariant for cross-backend gate device-name selection);changelog.d/fixed/vif-arc-mesa-anv-int64-reduction.md(new fragment). - Invariant: the workgroup-scope
memoryBarrierShared(); barrier();pair PR #511 introduced is load-bearing for the Arc + RADV lanes at API 1.4 and stays. Phase 3b confirmed it cannot be downgraded back to a barebarrier()even if the NVIDIA residual ever closes — Arc's clean state is contingent on the workgroup-scope pair. - Cross-backend gate device-selection invariant (NEW): scripts that target a specific Vulkan vendor must select by
deviceNamesubstring, not by--vulkan_device <index>.vmaf_vulkan_context_new's device sort is stable inside the samedevtype_scorebucket and thevkEnumeratePhysicalDevicesenumeration order is host-policy-dependent (driver registration order in/etc/vulkan/icd.d/, Mesa device-select layer,VK_LOADER_*env vars). PR #511's commit message inverted the device map on this fork's CI workstation; the empirical numbers it cited as "NVIDIA" actually came from Arc and vice versa. New cross-backend lanes targeting a specific vendor should not inherit the off-by-one. - On upstream sync:
vif.compis fork-local; no upstream Netflix/vmaf has a Vulkan path. Cherry-picks from upstream cannot reach this file. - Re-test on rebase (assumes a multi-GPU CI workstation with NVIDIA + Arc + RADV; lavapipe-only CI lanes are a no-op for the API-1.4 residual since lavapipe never reproduced the bug):
# Local API-1.4 bump (off-master reproducer; do NOT commit).
sed -i 's/VK_API_VERSION_1_3/VK_API_VERSION_1_4/g' \ core/src/vulkan/common.c sed -i 's/VMA_VULKAN_VERSION 1003000/VMA_VULKAN_VERSION 1004000/' \ core/src/vulkan/vma_impl.cpp cd libvmaf && meson setup build -Denable_vulkan=enabled \ -Denable_cuda=false -Denable_sycl=false && ninja -C build cd ..
# NVIDIA lane — expected 45/48 FAIL scale 2 until either the
# manual int64 subgroup-reduction patch lands or NVIDIA fixes
# the driver. Arc + RADV expected 0/48.
0230 — ssimulacra2_cuda GPU module unload + per-scale malloc removal (ADR-0356)¶
- Touches:
core/src/feature/cuda/ssimulacra2_cuda.c(fork-only — fork-added CUDA extractor),core/src/feature/cuda/ssimulacra2/ssimulacra2_blur.cu(fork-only kernel),core/src/cuda/AGENTS.md(fork-local package guidance). - Invariant: every
cuModuleLoadDatain the fork's CUDA extractors must be paired with a guardedcuModuleUnloadin the matchingclose_fex_cuda, betweencuStreamSynchronizeandcuStreamDestroy. The leak is invisible tocompute-sanitizer --tool memcheck(the tool's leak-checker is scoped tocuMem*Alloconly). The XYB H2D / D2H byte counts shrink to the valid sub-region per scale; the device-sideplane_full_pixelsstride contract (kernels assume each plane starts at full-resolution offsets) stays unchanged. Pinned scratch reservationsh_ref_lin_ds/h_dis_lin_dsare owned byss2c_alloc_buffersand freed byclose_fex_cudavia the existingSS2C_FREE_HOSTmacro. - Upstream interaction: none.
ssimulacra2_cudais fork-added per ADR-0206 and has no upstream Netflix/vmaf twin. meson test -C core/build test_ssimulacra2_simd
python3 scripts/ci/cross_backend_vif_diff.py \ --vmaf-binary core/build/tools/vmaf \ --reference testdata/ref_576x324_48f.yuv \ --distorted testdata/dis_576x324_48f.yuv \ --width 576 --height 324 \
--feature vif --backend vulkan --device <NVIDIA-index>
# Revert local bump after testing.
sed -i 's/VK_API_VERSION_1_4/VK_API_VERSION_1_3/g' \ core/src/vulkan/common.c sed -i 's/VMA_VULKAN_VERSION 1004000/VMA_VULKAN_VERSION 1003000/' \ core/src/vulkan/vma_impl.cpp
Upstream-port-later batch — Research-0090 18-commit triage close-out (2026-05-09)¶
- Touches:
docs/state.md(one row in "Deferred (waiting on external trigger)"), this file,changelog.d/changed/upstream-port-later-batch-2026-05-09.md. No code touched. Companion to PR #446 (Research-0090) and the in-flight PRs #497 (MyTestCase super-PR), #443 / #444 (cambi-docs duplicate pair). - Per-commit classification (input set: 18 PORT_LATER SHAs from Research-0090):
| # | Upstream SHA | Subject (truncated) | Verdict | Reopen / forward path |
|---|---|---|---|---|
| 1 | 38e905d1 | adopt MyTestCase + reformat BD-rate test data | PORT_DEFERRED | Subsumed by PR #497 commit e1dbdc09; close out when #497 merges |
| 2 | 005988ea | adopt MyTestCase + port new tests + align fifo_mode | PORT_DEFERRED | Subsumed by PR #497 commit 6c05afe2; close out when #497 merges |
| 3 | 4679db83 | fix VMAFEXEC_score tolerances for macOS FP precision | PORT_DEFERRED w/ Netflix-golden guard | PR #497 commit 0004d2cf — must preserve fork's golden places= values byte-for-byte (CLAUDE §8 / ADR-0024) |
| 4 | 3e075107 | adopt MyTestCase + update score values in vmafexec tests | PORT_DEFERRED | Subsumed by PR #497 commit 0004d2cf; close out when #497 merges |
| 5 | e3827e4d | adopt MyTestCase + port new tests in asset/bootstrap/local_explainer | PORT_DEFERRED | Subsumed by PR #497 commit 6c05afe2; close out when #497 merges |
| 6 | 25ff9f18 | remove empty VmafossexecCommandLineTest stub | PORT_DEFERRED → CHERRY-PICK after #497 | Pure 13-line deletion. PR #497 currently RE-EMITS the stub; once #497 lands, cherry-pick this commit standalone (zero-conflict against post-#497 tip). |
| 7 | 3a041a97 | adopt MyTestCase + update score values | PORT_DEFERRED | Subsumed by PR #497 commit d52d9221; close out when #497 merges |
| 8 | ead2d12b | fix vif_scale3 + adm3_egl_1 tolerances for macOS FP precision | PORT_DEFERRED w/ Netflix-golden guard | PR #497 commit b5a3f61b — Netflix-golden tolerance guard same as row 3 |
| 9 | 6c097fc4 | reduce ADM/VIF tolerances for macOS FP precision | PORT_DEFERRED w/ Netflix-golden guard | PR #497 commit f3881d5c — Netflix-golden tolerance guard same as row 3 |
| 10 | 7df50f3a | align testutil with full set of fixture functions | PORT_DEFERRED | Subsumed by PR #497 commit f1ae0495; close out when #497 merges |
| 11 | 322ca041 | replace temporal slicing with pre-sliced YUV fixtures | PORT_DEFERRED | Subsumed by PR #497 commit 7d9d9a10; close out when #497 merges. Sequencing matters: this commit must land before rows 12, 14, 15, 17 (the YUV-fixture consumers); #497 already orders them correctly. |
| 12 | 74bdce1b | align vmafexec_feature_extractor_test (aim/adm3/motion3) | PORT_DEFERRED | Subsumed by PR #497 commit 07e7cb48; close out when #497 merges |
| 13 | a3776335 | align feature_extractor_test (aim/adm3/motion3) | PORT_DEFERRED | Subsumed by PR #497 commit 15a6874d; close out when #497 merges |
| 14 | 0341f730 | remove duplicate test_run_vmaf_integer_fextractor | PORT_DEFERRED → CHERRY-PICK after #497 | Pure 76-line deletion. Same disposition as row 6 — #497 currently re-emits the duplicate; cherry-pick standalone after #497. |
| 15 | 9fa593eb | port feature_extractor tests for aim/adm3/motion3 + new options | PORT_DEFERRED | Subsumed by PR #497 commit ab21b694; close out when #497 merges |
| 16 | d93495f5 | reduce tolerance for VMAF scores in quality_runner tests | PORT_DEFERRED w/ Netflix-golden guard | PR #497 — Netflix-golden tolerance guard same as row 3 |
| 17 | 7d1ad54b | port feature extractor tests for aim/adm3/motion3 | PORT_DEFERRED | Subsumed by PR #497 commit 44b9e626; close out when #497 merges |
| 18 | 721569bc | resource/doc: cambi_high_res_speedup + motion2 score | PORT_DEFERRED → DEDUP | Already in flight on TWO branches (PR #443 + PR #444). Maintainer picks one and abandons the other per Research-0090 §Recommended action #4. No third port-PR opened. |
- Invariant: after PR #497 merges, the Research-0090 PORT_LATER bucket reduces to exactly two follow-up cherry-picks against post-#497 master:
git cherry-pick 25ff9f18(delete emptyVmafossexecCommandLineTest).git cherry-pick 0341f730(delete duplicatetest_run_vmaf_integer_fextractor). Both are pure deletions onpython/test/command_line_test.pyandpython/test/feature_extractor_test.pyrespectively; no score change, no Netflix-golden interaction. They were excluded from PR #497 because the v2 super-PR's diff state currently RE-EMITS those identifiers (likely because #497 cherry-picked from an earlier upstream tip than25ff9f18/0341f730).- Netflix-golden guard (binding): per CLAUDE §8 / ADR-0024, the three Netflix CPU golden pairs in
python/test/quality_runner_test.py,vmafexec_test.py,vmafexec_feature_extractor_test.py,feature_extractor_test.py,result_test.py(1 normalsrc01_hrc00↔hrc01+ 2 checkerboard) carry hard-codedassertAlmostEqualrows that are NEVER modified by a fork PR. Upstream commits4679db83,ead2d12b,6c097fc4,d93495f5explicitly LOWERplaces=on a subset of those rows (their stated motivation is macOS FP precision drift, not a true score change). Reviewer of PR #497 must verify that the 3 golden pairs retain fork tolerances byte-for-byte; only non-golden rows may adopt the relaxations. - On upstream sync: future
/sync-upstreamruns that re-detect these 18 SHAs should match this entry via the SHA list and short-circuit Pass-2 classification (skip re-triage). - Re-test on rebase: none required at the time of this commit (no code touched); after the two follow-up cherry-picks (
25ff9f18+0341f730) eventually land, run meson test -C build --suite=fast make test-netflix-golden # 3/3 CPU goldens still pass ADR-0108: every fork-local PR that touches upstream-shared paths or establishes a rebase-sensitive invariant adds an entry here. PRs with no rebase impact state "no rebase impact" in the PR description and skip the entry.
The intended reader is whoever runs the next /sync-upstream (see ADR-0002 and .claude/skills/sync-upstream/). Read top-to-bottom before resolving conflicts.
Format¶
Each entry is a ### NNNN — short title heading with three fields:
- Touches: paths likely to conflict on upstream merge.
- Invariant: what the fork relies on that an upstream change could silently drop.
- Re-test: the command(s) to run after the merge to confirm the invariant survived. Reproducer-style — no surrounding prose required.
IDs are assigned in commit order and never reused. A single entry may cover several PRs in one workstream; cross-link from the ID heading.
Entries (backfilled 2026-04-18 per ADR-0108 adoption)¶
0310 — Vulkan VIF int64 reduction race condition Phase 3 fix¶
- Touches:
core/src/feature/vulkan/shaders/vif.comp(replaces all three barebarrier()calls with explicitmemoryBarrierShared(); barrier();pairs covering the Phase-1 cooperative tile load, the Phase-2 vertical-conv shared write, and the Phase-4 cross-subgroup int64 reduction); plus documentation underdocs/research/0089-...md(Phase 3 status appendix),docs/adr/0269-...md(Phase 3 status appendix),docs/state.md(T-VK-VIF-1.4-RESIDUAL closed; new T-VK-VIF-1.4-RESIDUAL-ARC opened),core/src/vulkan/AGENTS.md(Phase 3 update on the existing invariant row),changelog.d/fixed/vif-int64-reduction-race-condition.md. Upstream Netflix/vmaf has no Vulkan backend, so conflict probability for the shader is zero. The entry exists because the fix is rebase-sensitive: any future cherry-pick that touchesvif.compand downgrades amemoryBarrierShared(); barrier();pair back to a barebarrier()will silently re-introduce the NVIDIA Vulkan 1.4 race. - Invariant:
vif.compshared-memory ordering between cooperative-write phases must be release-acquire, not just a bare workgroup-execution barrier. NVIDIA's Vulkan 1.4 default memory model requires the explicit shared-memory release; barebarrier()works at API 1.3 by accident on this driver. SCALE is irrelevant — the fix applies to all four pipeline specialisations because the barrier sites are in the SCALE-shared code. Do NOT remove the explicitmemoryBarrierShared()calls even if a perf review claims they are redundant under the GLSL spec wording: empirical real-hardware evidence in research-0089 2026-05-09 appendix shows otherwise on NVIDIA driver 595.71.05. - Re-test: apply the local API-1.4 bump (
core/src/vulkan/common.c3 sites +vma_impl.cppVMA_VULKAN_VERSION 1004000) on a NVIDIA RTX 4090 + driver 595.71+ machine, build withmeson setup ... -Denable_vulkan=enabled, then runpython3 scripts/ci/cross_backend_vif_diff.py --feature vif --backend vulkan --device 1 --places 4. Expect 0/48 across all four scales. Run the 5-run determinism check from research-0089 §"Reproduction recipe for Phase 3" against--vulkan_device 1; expect 5 identical(integer_vif_num_scale2, integer_vif_den_scale2) = (+2.494358e+04, +2.522523e+04)pairs at frame 5. Note that--vulkan_device 0on this multi-GPU host is the Intel Arc A380 lane and will still fail at API 1.4 (separateT-VK-VIF-1.4-RESIDUAL-ARCrow Open).
0309 — Vulkan VIF API-1.4 Phase 2 dump (T-VK-VIF-1.4-RESIDUAL)¶
- Touches:
docs/research/0089-vulkan-vif-fp-residual-bisect-2026-05-08.md(2026-05-09 status appendix with empirical numbers from the live RTX 4090),docs/state.md(T-VK-VIF-1.4-RESIDUAL row updated with the localisation),core/src/vulkan/AGENTS.md(new invariant row pinning the SCALE = 2 cross-subgroup-reduction memory-model finding),CHANGELOG.md(lusoris fork "Changed" entry). No code touched; the Phase 3 shader memory-model fix lands in a separate PR. Upstream Netflix/vmaf has no Vulkan backend so conflict probability for the AGENTS.md row is zero — entry exists because the empirical localisation flips the open state-row hypothesis from FP-precision to memory-model and retires theplaces=3override path that earlier rebase scaffolding might have suggested. - Invariant:
vif.compSCALE = 2 specialisation's Phase-4 cross-subgroup int64 reduction is non-deterministic on NVIDIA driver 595.71.05 + Vulkan 1.4.341 (lines 547–592,subgroupAddbarrier()+ thread-0 read ofs_lmem). API 1.3 lane is fully deterministic on the same hardware. The fourapiVersionpinning sites incore/src/vulkan/common.c+core/src/vulkan/vma_impl.cppstay at 1.3 until Phase 3 lands the explicit memory-scope barrier and a 5-run determinism gate confirms run-to-run identical(num, den)plusplaces=40/48 on NVIDIA. Theplaces=3override path is eliminated from the unblock options. - Re-test: apply the local API-1.4 bump (
core/src/vulkan/common.c3 sites +vma_impl.cppVMA_VULKAN_VERSION 1004000) on a NVIDIA RTX 4090 + driver 595.71+ machine, build withmeson setup ... -Denable_vulkan=enabled, then run the gate and the 5-run determinism check from research-0089 §"Reproduction recipe for Phase 3". Expect 45/48places=4failures oninteger_vif_scale2(max abs1.527e-02) AND 5 distinct(integer_vif_num_scale2, integer_vif_den_scale2)pairs across 5 runs of--feature 'vif_vulkan=debug=true'. Both observations reproduced bit-for-bit on this session's hardware lane (UUIDe478b41b-5c4f-1ddb-f990-e44916aff4c8).
0308 — encoder knob-sweep recipe-regression policy (ADR-0308, docs-only)¶
- Touches:
docs/research/0080-encoder-knob-sweep-findings.md,docs/adr/0308-encoder-knob-sweep-recipe-regression-policy.md,docs/adr/README.md(index row),ai/AGENTS.md(knob-sweep invariant section),changelog.d/changed/encoder-knob-sweep-findings.md. No code touched; companion to PR #400 (ADR-0305 + Research-0077 +ai/scripts/analyze_knob_sweep.py). Upstream Netflix/vmaf has no encoder-knob-sweep surface, so conflict probability is zero — this entry exists only because the policy threshold (7-of-9 structural cut) is rebase-sensitive on the corpus shape. - Invariant: the 7-of-9 source-count threshold from ADR-0308 §Decision point 1 is calibrated against the current 9-source Netflix Public Dataset corpus. If the corpus grows past 9 sources (e.g. UGC expansion per ADR-0287, or HDR additions), re-derive the absolute threshold as a fraction (≥7/9 ≈ 78 %). The structural cluster is sharp on the current corpus (top-15 cells all hit 9-of-9, no observed cells in 4-6 range), so a fractional cut at ~75 % is robust. Do NOT relax
bitrate_tol_pct(default 5.0) orvmaf_tol(default 0.1) inai/scripts/analyze_knob_sweep.pywithout an ADR — those tolerances are calibrated against the per-frame VMAF noise floor and bitrate quantisation in libavformat muxers. - Re-test:
pytest ai/tests/test_knob_sweep_analysis.py -v(script logic; ships in PR #400). Policy gate is offline: regenerateruns/phase_a/full_grid/comprehensive.jsonlviatools/vmaf-tune/src/vmaftune/hw_encoder_corpus.py(3-hour run on a single host with NVENC + QSV) then re-runpython ai/scripts/analyze_knob_sweep.py --jsonl <adapted.jsonl> --out-dir runs/phase_a/full_grid/reports/and diff the resultingsummary.mdagainstdocs/research/0080-encoder-knob-sweep-findings.mdheadline table. Structural cluster (top-15 cells, all 9-of-9) is the invariant to defend.
0228 — Vulkan 1.4 bump deferred (ADR-0264, docs-only)¶
- Touches: none (docs-only PR). Future Step A of T-VK-1.4-BUMP will touch
core/src/feature/vulkan/shaders/vif.compandcore/src/feature/vulkan/shaders/ciede.comp; Step B will touch the threeapiVersionsites incore/src/vulkan/common.c(lines 54, 264, 374) and theVMA_VULKAN_VERSIONdefine incore/src/vulkan/vma_impl.cpp(line 22). - Invariant:
masterstays onVK_API_VERSION_1_3andVMA_VULKAN_VERSION = 1003000. Lifting the constant in any future upstream sync (Netflix doesn't ship a Vulkan backend, so the conflict is improbable) without first auditingprecise/OpDecorate ... NoContractiondecoration onvif.compandciede.compwill reintroduce the NVIDIA-driver regression captured in research-0053. Thepsnr_hvs_strict_shaders-O0list incore/src/vulkan/meson.buildis the existing precedent for shader-side bit-exactness mitigations and should be the place a 1.4-era audit lands its results (potentially expanding to covervif.comp+ciede.compif thepreciseaudit decides the optimizer is the right place to gate). - Re-test: when Step B lands, the gate is
python3 scripts/ci/cross_backend_vif_diff.py --feature vif --backend vulkanand the same with--feature ciedeagainst NVIDIA + RADV + lavapipe; max abs diff must stay ≤5.0e-05(places=4) on all three.
0229 — HIP fifth-consumer kernel float_ansnr_hip (ADR-0266)¶
0228 — y4m_convert_411_422jpeg 1-byte heap-buffer-overflow fix¶
0228 — vmaf-tune resolution-aware model selection (ADR-0289)¶
0282 — vmaf-tune AMD AMF codec adapters (ADR-0282)¶
0228 — tools/vmaf-tune/ codec-agnostic encode dispatcher (ADR-0294)¶
- Touches:
tools/vmaf-tune/src/vmaftune/encode.py— refactored to look up the codec adapter and delegate argv composition. Wholly fork-local.tools/vmaf-tune/src/vmaftune/codec_adapters/__init__.py,codec_adapters/x264.py— adapter contract gainsffmpeg_codec_args(preset, quality)andextra_params(). Both are duck-typed; missing methods fall back to the legacy x264-CRF shape.tools/vmaf-tune/tests/test_encode_multi_codec.py— new 19-test suite pinning the dispatcher contract per codec.docs/usage/vmaf-tune.md— new "Codec adapter contract" section.- Invariant: the harness (
encode.py,corpus.py) must not branch on codec identity. The only codec-aware code is the per-adaptercodec_adapters/*.pyfile. Any future change that adds anif adapter.encoder == "..."to the harness regresses ADR-0294's whole-purpose. The corpus row schema stays at SCHEMA_VERSION=1 —crfis preserved as the row column even when the underlying codec's quality knob is-cq/-qp/ etc.;EncodeRequest.qualityis a request-side property only. Adapters that don't yet exposeffmpeg_codec_argsare intentionally permitted to fall back to the legacy x264-CRF shape; removing that fallback would break in-flight adapter PRs landing one-at-a-time. - Re-test on rebase:
```bash pytest tools/vmaf-tune/tests/ -q # 32 passed (13 existing + 19 multi-codec)
python -c " from pathlib import Path from vmaftune.encode import EncodeRequest, build_ffmpeg_command req = EncodeRequest( source=Path('ref.yuv'), width=1920, height=1080, pix_fmt='yuv420p', framerate=24.0, encoder='libx264', preset='medium', crf=23, output=Path('out.mp4'), ) cmd = build_ffmpeg_command(req) assert cmd[cmd.index('-c:v') + 1] == 'libx264' assert cmd[cmd.index('-preset') + 1] == 'medium' assert cmd[cmd.index('-crf') + 1] == '23' print('x264 dispatcher path OK') "
0260 — vmaf-tune --sample-clip-seconds (ADR-0301)¶
- Touches:
tools/vmaf-tune/src/vmaftune/{cli,corpus,encode,score,__init__}.py— fork-local. No upstream Netflix/vmaf path overlap.tools/vmaf-tune/tests/test_corpus.py,tools/vmaf-tune/AGENTS.md,docs/usage/vmaf-tune.md,docs/adr/0301-vmaf-tune-sample-clip.md,docs/adr/_index_fragments/0301-vmaf-tune-sample-clip.md,docs/adr/_index_fragments/_order.txt,docs/adr/README.md.- Invariant: corpus JSONL
SCHEMA_VERSIONbumped to2— additiveclip_modekey only. Sample-clip windows are mirrored on both sides via FFmpeg input-side-ss/-t(encode) and libvmaf's--frame_skip_ref/--frame_cnt(score). The_resolve_sample_clip()helper is the single source of truth for the centre-anchored slice math; do not duplicate the computation elsewhere. Falls back silently to"full"whenN >= duration_s. - Re-test:
0227 — tools/vmaf-tune/ Phase A scaffold (ADR-0237 Phase A)¶
- Touches:
tools/vmaf-tune/src/vmaftune/codec_adapters/{h264_amf,hevc_amf,av1_amf,_amf_common}.py(new). Wholly fork-local — no upstream Netflix/vmaf path overlap.tools/vmaf-tune/src/vmaftune/codec_adapters/__init__.py— registry extended with three AMF entries.tools/vmaf-tune/tests/test_codec_adapter_amf.py(new).tools/vmaf-tune/tests/test_corpus.py— Phase A test renamed fromtest_known_codecs_phase_a_is_x264_onlytotest_known_codecs_includes_x264_and_amf.tools/vmaf-tune/AGENTS.md— adds AMF preset-compression invariant.docs/usage/vmaf-tune.md— adds Hardware encoders section.- Invariant: the 7-into-3 preset compression table in
_amf_common.py(_PRESET_TO_AMF) is the cross-codec axis Phase B / C consumers depend on. Every AMF adapter accepts the canonical 7 preset names (placebo…ultrafast) and maps them onto the three AMF rungs (quality/balanced/speed). Do not extend the preset vocabulary without amending ADR-0282 — registry uniformity (no codec-identity branching in the harness search loop) rests on every codec accepting the same names. - Re-test:
0227 — tools/vmaf-tune/ Phase A scaffold (ADR-0237 Phase A)¶
- Touches:
tools/vmaf-tune/src/vmaftune/resolution.py(new). Wholly fork-local — no upstream Netflix/vmaf path overlap.tools/vmaf-tune/src/vmaftune/corpus.py— addsCorpusOptions.resolution_aware: bool = Trueand pipes the effective model throughscore_res.request.modelinto the JSONL row.tools/vmaf-tune/src/vmaftune/cli.py— adds--resolution-aware/--no-resolution-aware(BooleanOptionalAction, default on).tools/vmaf-tune/tests/test_resolution.py(new).docs/usage/vmaf-tune.md— new "Resolution-aware mode" section.docs/adr/0289-vmaf-tune-resolution-aware.md(new) +docs/research/0064-vmaf-tune-resolution-aware.md(new).tools/vmaf-tune/AGENTS.md— two new invariant notes.- Invariant: the height-only decision rule (
height >= 2160→vmaf_4k_v0.6.1, elsevmaf_v0.6.1) is the documented contract. The JSONLvmaf_modelfield is now per-row (not per-job) — mixed ladder corpora legitimately contain multiple distinct values across rows. Downstream consumers (Phase B / C / D) must group/filter byvmaf_modelrather than assuming a constant. Width is accepted in the API for symmetry but ignored in the body; do not branch on it without a follow-up ADR. - Re-test:
pytest tools/vmaf-tune/tests/ -q
python tools/vmaf-tune/vmaf-tune corpus --help | grep resolution-aware
0227 — tools/vmaf-tune/ Phase A scaffold (ADR-0237 Phase A)¶
- Touches:
core/tools/y4m_input.c— upstream-mirrored Daala-derived Y4M parser. The fix sits inside the 4:1:1 → 4:2:2-jpeg chroma upsample routiney4m_convert_411_422jpeg, lines ~500–530 in the function's three sub-loops. Upstream Netflix/vmaf carries the same shape; if upstream lands its own fix during a sync, prefer the upstream version and drop ours.core/test/test_y4m_411_oob.c(new, fork-local) — drives the minimal W=2 H=4 4:1:1 stream throughvideo_input_open+video_input_fetch_frame. Wholly fork-added; no upstream collision.core/test/meson.build— addstest_y4m_411_oobexecutable +test()registration.- Invariant: the first two sub-loops of
y4m_convert_411_422jpegmust guard_dst[(x << 1) | 1]writes with(x << 1 | 1) < dst_c_w, matching the third sub-loop's existing guard. Without the guard a 4:1:1 stream of width 2 (dst_c_w == 1) writes one byte past the destination chroma row. - Re-test:
cd libvmaf && meson setup ../build-asan --buildtype=debug -Db_sanitize=address -Db_lundef=false -Denable_cuda=false -Denable_sycl=false -Denable_vulkan=disabledninja -C build-asan test/test_y4m_411_oobASAN_OPTIONS=detect_leaks=0 ./build-asan/test/test_y4m_411_oob— must report1 tests run, 1 passed. Pre-fix the binary aborts withAddressSanitizer: heap-buffer-overflow … WRITE of size 1aty4m_input.c:507.
0270 — saliency_student_v1 fork-trained on DUTS-TR (ADR-0286)¶
- Touches:
model/tiny/registry.json— adds thesaliency_student_v1row. Fork-local registry; no upstream overlap.model/tiny/saliency_student_v1.onnx(+.jsonsidecar) — new weights and metadata. Fork-local.ai/scripts/train_saliency_student.py— new training script. Wholly fork-local underai/, which has no upstream counterpart.docs/ai/models/saliency_student_v1.md,docs/research/0062-saliency-student-from-scratch-on-duts.md,docs/adr/0286-saliency-student-fork-trained-on-duts.md— new docs under fork-local trees.- Invariant: the C-side
feature_mobilesal.cextractor's tensor-name contract —input(NCHW[1, 3, H, W]) andsaliency_map(NCHW[1, 1, H, W]) — must continue to match the ONNX graph for bothsaliency_student_v1.onnxand the legacymobilesal.onnxplaceholder. Future weights swaps can change the graph internals freely but must keep these names + shapes; the smoke test asserts the registration. The op-allowlist constraint (graph uses only ops incore/src/dnn/op_allowlist.c) carries over from ADR-0218 —Resizeis not used;ConvTransposeis the upsample op for v1 to keep the graph load-clean against vanilla origin/master. - Re-test:
.venv/bin/python ai/scripts/validate_model_registry.py
.venv/bin/python -c "
from ai.src.vmaf_train.op_allowlist import check_model
from pathlib import Path
r = check_model(Path('model/tiny/saliency_student_v1.onnx'))
assert r.ok, r.pretty()
print('allowlist OK')
"
meson test -C build --suite=fast mobilesal
0227 — tools/vmaf-tune/ Phase A scaffold (ADR-0237 Phase A)¶
- Touches:
core/src/feature/hip/float_ansnr_hip.{c,h}(new) — fifth consumer ofcore/src/hip/kernel_template.h. Mirrorscore/src/feature/cuda/float_ansnr_cuda.ccall-graph-for-call-graph;init/submit/collect/closeinvoke the kernel-template helpers in the same order; the submit body intentionally bypassesvmaf_hip_kernel_submit_pre_launch(no atomic, kernel writes per-block (sig, noise) interleaved float partials directly).core/src/hip/meson.build— adds the new TU tohip_sources.core/src/feature/feature_extractor.c— adds theextern VmafFeatureExtractor vmaf_fex_float_ansnr_hip;declaration and the registry row under#if HAVE_HIP.core/test/test_hip_smoke.c— addstest_float_ansnr_hip_extractor_registeredsub-test pinning the lookup contract.- Invariant — the
submit_pre_launchbypass is load-bearing. The CUDA twin makes the same choice for the same reason. If a future PR adds asubmit_pre_launchcall tofloat_ansnr_cuda.c's submit path, the HIP twin must follow in the same PR. Likewise the readback shape (wg_count * 2u * sizeof(float)) and the bpc table (peak/psnr_max for 8/10/12/16-bit) mirror the CUDA twin verbatim — keep aligned on rebase. - Re-test on rebase:
cd libvmaf
meson setup build -Denable_hip=true -Denable_cuda=false -Denable_sycl=false
ninja -C build
meson test -C build # 48/48 green (47 CPU + HIP smoke)
0230 — HIP sixth-consumer kernel motion_v2_hip (ADR-0267)¶
- Touches:
core/src/feature/hip/integer_motion_v2_hip.{c,h}(new) — sixth consumer ofcore/src/hip/kernel_template.h. Mirrorscore/src/feature/cuda/integer_motion_v2_cuda.ccall-graph-for-call-graph; carries theVMAF_FEATURE_EXTRACTOR_TEMPORALflag and aflush()callback. The state struct has auintptr_t pix[2]ping-pong slot pair tracked outside the kernel-template (the template models a single device+host pair only).core/src/hip/meson.build— adds the new TU tohip_sources.core/src/feature/feature_extractor.c— adds theextern VmafFeatureExtractor vmaf_fex_integer_motion_v2_hip;declaration and the registry row under#if HAVE_HIP.core/test/test_hip_smoke.c— addstest_motion_v2_hip_extractor_registeredsub-test pinning the lookup contract (extractor name ismotion_v2_hip, matching the CUDA twin'smotion_v2_cudanaming).- Invariant — temporal-extractor + ping-pong shape. The
VMAF_FEATURE_EXTRACTOR_TEMPORALflag bit, theflush()callback registration, and theuintptr_t pix[2]slot pair are load-bearing for the runtime PR (T7-10b). The runtime PR will swapuintptr_t pix[2]for a real device-buffer handle pair matching the CUDA twin'sVmafCudaBuffer *pix[2]. On rebase: if the CUDA twin's flush-pass shape changes (currentlymin(score[i], score[i+1])), update the HIP twin'sflush_fex_hipbody in the same PR. - Re-test on rebase: same as 0229 —
meson test -C buildwithenable_hip=trueexercises the smoke contract.
0227 — ms_ssim_vulkan submit-side migrated to kernel_template (T-GPU-DEDUP-26)¶
- Touches:
core/src/feature/vulkan/ms_ssim_vulkan.c—extract()'s rawVkCommandBuffer/VkFence/vkAllocateCommandBuffers/vkBeginCommandBuffer/vkCreateFence/vkQueueSubmit/vkWaitForFences/vkDestroyFence/vkFreeCommandBuffersblocks becomeVmafVulkanKernelSubmittriples (vmaf_vulkan_kernel_submit_begin/_submit_end_and_wait/_submit_free). One triple covers the decimate-pyramid command buffer; one triple per scale covers the per-scale SSIM submit. The pipeline-side bundles (pl_decimate2-binding 4-variant +pl_ssim10-binding 9-variant) and their_add_variant()chains are unchanged from the prior migration.- Invariant: any future submit-side template change (timeline semaphores, deferred fence release, queue-family parameterisation) must keep the helpers' synchronous-wait + per-frame fence + per-frame command-buffer contract intact, since
ms_ssim_vulkan.cdoes host readback of thel_partials/c_partials/s_partialsbuffers immediately after_submit_end_and_waitreturns. The submit-side contract is the same one already documented incore/src/vulkan/AGENTS.md's "Rebase-sensitive invariants" section forkernel_template.h. - Re-test:
```bash cd libvmaf && meson test -C build python scripts/ci/cross_backend_vif_diff.py \ --vmaf-binary core/build/tools/vmaf \ --reference testdata/ref_576x324_48f.yuv \ --distorted testdata/dis_576x324_48f.yuv \ --width 576 --height 324 \ --feature float_ms_ssim --backend vulkan --places 4
0231 — SHA-pin GitHub Actions (OSSF Pinned-Dependencies)¶
- Touches: every workflow file under
.github/workflows/. All 13 fork workflows (docker-image.yml,docs.yml,ffmpeg-integration.yml,libvmaf-build-matrix.yml,lint-and-format.yml,nightly-bisect.yml,nightly.yml,release-please.yml,rule-enforcement.yml,scorecard.yml,security-scans.yml,supply-chain.yml,tests-and-quality-gates.yml) had theiruses:directives rewritten from<owner>/<repo>@vN[.M.K]to<owner>/<repo>@<40-char-sha> # vN.M.K. 97 references converted; the SLSA reusable-workflow ref insupply-chain.ymlis the single documented holdout (seeInvariantbelow). - Invariant — SHA-pin policy for
uses:. Every action reference in.github/workflows/*.ymlMUST be a 40-char commit SHA with the semver tag preserved as a trailing# vN.M.Kcomment. The OSSF ScorecardPinned-Dependenciescheck parses both forms and a floating tag (@vN) is treated as unpinned and counts against the aggregate score. Single permitted exception: the SLSA generator reusable workflow (slsa-framework/slsa-github-generator/.github/workflows/generator_generic_slsa3.yml) must keep itsvX.Y.Ztag form because GitHub Actions consumers cannot SHA-pin reusable-workflow refs in every code path; the exception is documented inline insupply-chain.ymland survives on each rebase. Why this matters on upstream sync: Netflix upstream does not ship the fork's CI tree, so a/sync-upstreamrun that drags new workflow content (e.g. via repository templates or bot-authored bumps) into.github/workflows/can re-introduce floating-tag references unnoticed. The post-rebase check below is the standing gate — anything that lights up needs to be re-pinned before merging the sync. - Re-test on rebase:
# Anything that prints is a regression — every uses: must be either
# already SHA-pinned (40 hex) or, for the documented SLSA exception,
# the slsa-github-generator reusable-workflow ref.
grep -hnE '^\s*(- )?uses:\s+[^@]+@[^ #]+\s*$' .github/workflows/*.yml \
| grep -vE '@[a-f0-9]{40}' \
| grep -v 'slsa-framework/slsa-github-generator/.github/workflows/'
# SHA-resolution sanity for any new pin (per-action):
gh api repos/<owner>/<repo>/git/ref/tags/<vN.M.K> --jq '.object.sha'
# If the result is a "tag" object (annotated tag), deref:
gh api repos/<owner>/<repo>/git/tags/<sha-from-prev> --jq '.object.sha'
0226 — CUDA drain-batch engine-loop opt (T-GPU-OPT-1)¶
- Touches:
core/src/cuda/drain_batch.{h,c}(new) — TLS drain-batch table + shared drain stream +_open()/register/_flush()/_close()API.core/src/libvmaf.c— engine-side per-frame loop now wraps submit/collect with_open()+_flush()so all CUDA extractorfinishedevents are waited on a single shared drain stream.- All 12 CUDA feature kernels (
core/src/feature/cuda/*.c) register theirfinishedevent +drainedflag with the drain batch on submit; collect skips its privatecuStreamSynchronizewhendrainedis true. - Invariant — drained-flag contract. Every CUDA extractor's collect path must check the per-frame
drainedflag and skip its owncuStreamSynchronizewhen set; otherwise the drain batching is a no-op. The flag is reset tofalseper frame insidevmaf_cuda_drain_batch_register(). - Re-test on rebase:
cd libvmaf
meson setup build -Denable_cuda=true -Denable_sycl=false
ninja -C build
meson test -C build --suite=fast cuda
Expected: all CUDA tests green; bench shows ≥5% wall-clock gain on a 7-extractor VMAF model (model.json with all feature extractors enabled).
0225 — Netflix bench snapshot regen (upstream a44e5e61 motion fix)¶
- Touches:
testdata/netflix_benchmark_results.json— fork-added snapshot. CPU rows now reflect the post-fix motion feature; cuda / sycl rows from the previous regen are preserved unchanged because those backends were not exercised on this rerun (host-environment tooling — wrong renderD path,libvmaf_cudanot enabled in the local FFmpeg build). Future full regens should include cuda / sycl.testdata/bench_all.sh— defaultVMAF=no longer points at/usr/local/bin/vmaf(which on most dev hosts is stuck at the pre-upstream-a44e5e61v3.0.0); now defaults to the in-tree fork build atcore/build/tools/vmaf.testdata/benchmark_netflix.py—FFMPEG,YUVDIRand the hardcodedLD_LIBRARY_PATH=/usr/local/libare now overridable viaVMAF_FFMPEG,VMAF_YUVDIRand any caller-setLD_LIBRARY_PATH.- Invariant: the snapshot's CPU pooled VMAF for
src01_576x324is 76.667828 (post-fix), not 76.668904 (the upstream-buggy mirror). If/sync-upstreamever re-pulls a Netflix change that touchesmotion.cmirror-handling, this number is the reference. - Re-test:
cd libvmaf
meson setup build -Denable_cuda=true -Denable_sycl=false
ninja -C build
LD_LIBRARY_PATH=$(pwd)/build/src python3 \
../testdata/benchmark_netflix.py
Expected CPU pooled rows: 76.667828, 35.068672, 7.985899.
0224 — CUDA graph capture feasibility (research-0047, DEFER)¶
- Touches: none — investigation-only; no code lands. The research digest
docs/research/0047-cuda-graph-capture-feasibility.mddocuments why a CUDA graph capture path on the per-frame submit chain is deferred rather than shipped (realised wall-clock gain capped at ~1-3% vs. the predicted 10-20%, with a 4-slot picture-pool rotation that defeats single-graph capture and forces per-framecuGraphExecKernelNodeSetParamsrebinding for(ref, dis)device pointers). - Invariant: the
kernel_template.hdocstring keeps namingVmafCudaKernelLifecycle.finishedas a graph-capture hook point. Don't prune that comment on rebase — leaving the door open in the template is free, and the digest's "what needs to be true for a future GO" section depends on the hook still being there. - Re-test on rebase:
# Confirm the docstring still references graph capture as the hook
# point — wording change is fine, removal is not.
grep -q "graph capture" core/src/cuda/kernel_template.h
0223 — ADR slug-drift repair in CHANGELOG / rebase-notes (PR #304 follow-up)¶
- Touches:
CHANGELOG.md,docs/rebase-notes.md. No code; no upstream-shared path; no public-API surface. - Invariant: every
[ADR-NNNN](docs/adr/NNNN-slug.md)link in the fork's tracked docs resolves to an actual on-disk file underdocs/adr/. Repaired 4 broken slugs that did not exist on disk (0138-iqa-convolve-avx2-bitexact-double→0138-iqa-convolve-avx2-bitexact-double,0140-simd-dx-framework→0140-simd-dx-framework,0190-ms-ssim-vulkan→0190-ms-ssim-vulkan,0178-vulkan-adm-kernel→0178-vulkan-adm-kernel). All retained their cited NNNN per ADR-0028 (NNNN is immutable once Accepted). - Re-test on rebase: from repo root, the following must print no lines:
for ref in $(grep -ohE 'docs/adr/[0-9]{4}-[a-z0-9-]+\.md' \
CHANGELOG.md docs/rebase-notes.md AGENTS.md docs/state.md \
| sort -u); do
test -f "$ref" || echo "MISSING: $ref"
done
0125 — cambi_vulkan migrated to kernel_template (T-GPU-DEDUP-25, 5-bundle)¶
- Touches:
core/src/feature/vulkan/cambi_vulkan.c— state's quintet (dsl_2bind+ 5×pl_layout_*+shader_modules[CAMBI_PL_COUNT]areddesc_pool) collapses to fiveVmafVulkanKernelPipelinebundles (pl_trivial,pl_derivative,pl_filter_mode,pl_decimate,pl_mask_dp), each owning its own descriptor pool. The first slot ofpipelines[]per stage aliases the bundle's base pipeline;CAMBI_PL_FILTER_MODE_V,CAMBI_PL_MASK_SAT_COL, andCAMBI_PL_MASK_THRESHOLDare sibling variants built viavmaf_vulkan_kernel_pipeline_add_variant().cambi_vk_alloc_settakes a bundle pointer (->desc_pool/->dsl) — every dispatch site picks the bundle that matches its push-constant struct.- The
cambi_vk_make_dsl/cambi_vk_make_pl/cambi_vk_create_shader/cambi_vk_build_pipelinehelpers are dropped — the template subsumes them. - Invariant — variants destroyed before bundle, base alias must be skipped. Five distinct push-constant struct sizes (
CambiVkPushTrivial/CambiVkPushDerivative/CambiVkPushFilterMode/CambiVkPushDecimate/CambiVkPushMaskDp) force five bundles even though every stage's DSL is 2-binding SSBO;_add_variant()only siblings pipelines under the same layout.close_fexmustvkDestroyPipeline()the variant slots (CAMBI_PL_FILTER_MODE_V,CAMBI_PL_MASK_SAT_COL,CAMBI_PL_MASK_THRESHOLD) before callingvmaf_vulkan_kernel_pipeline_destroy()on each bundle. - Numerical contract: bit-exact preserved. Same shaders + spec-constants + push-constants as before; only the Vulkan pipeline-bundle scaffolding moved to the template. Validated on the Netflix-pair smoke (576×324×8-bit):
cambimean = 0.0, identical to pre-migration (the pair has no banding artifacts). - Rebase impact: low. Builds on top of PR #272's
_add_variant()helper. Upstream Netflix/vmaf has no Vulkan backend, so there is nothing to merge against.
0124 — ssimulacra2_vulkan migrated to kernel_template (T-GPU-DEDUP-24, 4-bundle)¶
- Touches:
core/src/feature/vulkan/ssimulacra2_vulkan.c— state's 16 long-lived pipeline-object fields (4×*_dsl + *_pl + *_shader+ the shareddesc_pool) collapse to fourVmafVulkanKernelPipelinebundles (pl_xyb,pl_mul,pl_blur,pl_ssim), each owning its own descriptor pool. The first slot of each per-bundle pipeline array (xyb_pipelines[0],mul_pipelines[0],blur_pipelines_h[0],ssim_pipelines[0]) aliases the bundle's baseVkPipeline; remaining per-scale / per-pass slots are siblings viavmaf_vulkan_kernel_pipeline_add_variant().ss2v_build_pipeline_int3reroutes through_add_variant()instead of callingvkCreateComputePipelinesdirectly;ss2v_alloc_settakes a bundle pointer (->desc_pool/->dsl) instead of a separate DSL argument; descriptor-set free sites at the tail ofss2v_run_scaleroute to each bundle's pool.- The
ss2v_make_dsl/ss2v_make_pl/ss2v_create_shaderhelpers are dropped — the template subsumes them. - Invariant — variants destroyed before bundle, slot 0 alias must be skipped. Four distinct DSL shapes (XYB = 6 SSBOs, MUL = 3, BLUR = 2, SSIM = 8) prevent collapsing to one bundle:
_add_variant()only siblings pipelines under the same layout.close_fexmustvkDestroyPipeline()the variant slots inxyb_pipelines[1..N-1],mul_pipelines[1..N-1],ssim_pipelines[1..N-1],blur_pipelines_h[1..N-1], and every slot ofblur_pipelines_v[]before callingvmaf_vulkan_kernel_pipeline_destroy()on each bundle, and must skip slot 0 of the first three arrays +blur_pipelines_hto avoid double-freeing the aliased base. - Numerical contract: bit-exact preserved. Same shaders + spec-constants + push-constants as before; only the Vulkan pipeline-bundle scaffolding moved to the template. Validated on the Netflix-pair smoke (576×324×8-bit):
ssimulacra2mean = 24.613842, identical to pre-migration. - Rebase impact: low. Builds on top of PR #272's
_add_variant()helper. Upstream Netflix/vmaf has no ssimulacra2 extractor and no Vulkan backend, so there is nothing to merge against.
0118 — psnr_hvs_vulkan migrated to kernel_template + _add_variant (T-GPU-DEDUP-18)¶
- Touches:
core/src/feature/vulkan/psnr_hvs_vulkan.c— state'sdsl + pipeline_layout + shader + desc_pool + pipeline[3]collapses toVmafVulkanKernelPipeline pl + VkPipeline pipeline_chroma_u + VkPipeline pipeline_chroma_v. Plane 0 is the template's base pipeline; planes 1+2 are siblings viavmaf_vulkan_kernel_pipeline_add_variant().- New
psnr_hvs_plane_pipeline()accessor maps plane index to the rightVkPipelinehandle. - Invariant — variants destroyed before bundle.
close_fexmustvkDestroyPipeline()the chroma U/V variants before callingvmaf_vulkan_kernel_pipeline_destroy(&s->pl)— same rule as ssim_vulkan in T-GPU-DEDUP-7. - Numerical contract: unchanged. Same shaders + spec-constants push-constants as before; only the Vulkan pipeline-bundle scaffolding moved to the template.
- Rebase impact: low. Builds on top of PR #272's
_add_variant()helper.
0119 — vif_vulkan migrated to kernel_template + _add_variant (T-GPU-DEDUP-19)¶
- Touches:
core/src/feature/vulkan/vif_vulkan.c— state'sdsl + pipeline_layout + shader + desc_pool + pipelines[4]collapses toVmafVulkanKernelPipeline pl + VkPipeline scale_variants[3]. Scale 0 is the template's base pipeline; scales 1, 2, 3 are siblings viavmaf_vulkan_kernel_pipeline_add_variant().- New
vif_scale_pipeline()accessor maps scale index to the rightVkPipelinehandle (replacess->pipelines[scale]). - Invariant — variants destroyed before bundle.
close_fexmustvkDestroyPipeline()the 3 scale variants before callingvmaf_vulkan_kernel_pipeline_destroy(&s->pl)— same rule as ssim_vulkan in T-GPU-DEDUP-7 and psnr_hvs_vulkan in T-GPU-DEDUP-18. - Numerical contract: unchanged. Same shaders, same spec-constants, same push-constants as before; only the Vulkan pipeline-bundle scaffolding moved to the template.
- Rebase impact: low. Builds on top of PR #272's
_add_variant()helper.
0120 — float_vif_vulkan migrated to kernel_template + _add_variant (T-GPU-DEDUP-20)¶
- Touches:
core/src/feature/vulkan/float_vif_vulkan.c— state collapsesdsl + pipeline_layout + shader + desc_pooltoVmafVulkanKernelPipeline pl; theVkPipeline pipelines[2][4]2-D lookup table is preserved so the existing[mode][scale]dispatch path stays clean, butpipelines[0][0]aliasess->pl.pipeline(the template's base). The other 6 entries are sibling pipelines created viavmaf_vulkan_kernel_pipeline_add_variant().- Invariant — variants destroyed before bundle.
close_fexmustvkDestroyPipeline()the 6 sibling variants (every(mode, scale)except(0, 0)) before callingvmaf_vulkan_kernel_pipeline_destroy(&s->pl)— same rule as ssim_vulkan / psnr_hvs_vulkan / vif_vulkan. - Invariant —
pipelines[0][0]aliasing. The base pipeline handle is owned bys->pl.pipeline; we copy it intopipelines[0][0]after_create()so the dispatch path can use a uniform 2-D lookup. The destroy loop must skip(mode=0, scale=0)to avoid double-freeing the template's pipeline. - Numerical contract: unchanged. Same shaders, spec-constants (
mode+scale), push-constants. Netflix-pair smoke matchesinteger_vifbit-identically to 4 decimals. - Rebase impact: low. Builds on top of PR #272's
_add_variant()helper.
0122 — float_adm_vulkan migrated to kernel_template + _add_variant (T-GPU-DEDUP-22)¶
- Touches:
core/src/feature/vulkan/float_adm_vulkan.c— twin to adm_vulkan (T-GPU-DEDUP-21); 16-pipeline 2-D[stage][scale]array. State collapsesdsl + pipeline_layout + shader + desc_pooltoVmafVulkanKernelPipeline pl.pipelines[0][0]aliasess->pl.pipeline; the other 15 entries are siblings viavmaf_vulkan_kernel_pipeline_add_variant().- Invariants:
- Variants destroyed before bundle.
pipelines[0][0]aliasing — destroy loop must skip(stage=0, scale=0).- Numerical contract: unchanged. Same float (
_ssuffix) primitives fromadm_tools.c; same 5-element spec-constant tuple; same float partial accumulation reduced in double on the host.
0121 — adm_vulkan migrated to kernel_template + _add_variant (T-GPU-DEDUP-21)¶
- Touches:
core/src/feature/vulkan/adm_vulkan.c— state collapsesdsl + pipeline_layout + shader + desc_pooltoVmafVulkanKernelPipeline pl; theVkPipeline pipelines[4][4]2-D lookup is preserved so the per-stage dispatch path stays clean.pipelines[0][0]aliasess->pl.pipeline(the template's base); the other 15 entries are sibling pipelines viavmaf_vulkan_kernel_pipeline_add_variant().- Invariants:
- Variants destroyed before bundle (same rule as ssim_vulkan / psnr_hvs / vif / float_vif).
pipelines[0][0]aliasing — destroy loop must skip(stage=0, scale=0)to avoid double-freeing the template's pipeline.- Numerical contract: unchanged. Same shaders + 5-element spec-constant tuple (width, height, bpc, scale, stage) + push-constants.
- Rebase impact: low. Builds on top of PR #272.
0123 — ms_ssim_vulkan 2-bundle migration (T-GPU-DEDUP-23)¶
- Touches:
core/src/feature/vulkan/ms_ssim_vulkan.c— state collapsesdecimate_dsl + decimate_pl + decimate_shader + ssim_dsl + ssim_pl + ssim_shader + desc_pool(7 fields) to two bundlesVmafVulkanKernelPipeline pl_decimate+pl_ssim. Each bundle owns its own descriptor pool. The kernel has two distinct pipeline shapes (decimate = 2 SSBO bindings, ssim = 10 bindings), so two bundles is the minimum —_add_variant()only siblings pipelines under the same layout.decimate_pipelines[0]aliasespl_decimate.pipeline(the template's base = scale 0). The remainingMS_SSIM_SCALES - 2decimate variants (scales 1..3) are siblings via_add_variant().ssim_pipeline_horiz[0]aliasespl_ssim.pipeline(base = scale 0, pass 0). The other 9 entries (4×ssim_pipeline_horizfor scales 1..4, plus 5×ssim_pipeline_vertfor scales 0..4) are variants.- Invariant — variants destroyed before bundle. Same rule as ADR-0106 entry 0106:
close_fexmust destroydecimate_pipelines[1..3]andssim_pipeline_horiz[1..4]+ssim_pipeline_vert[0..4]before callingvmaf_vulkan_kernel_pipeline_destroy()onpl_decimate/pl_ssim. - Invariant —
[0]aliasing destroy-skip.decimate_pipelines[0]andssim_pipeline_horiz[0]must not be passed tovkDestroyPipelineinclose_fex—_destroy()already releases them viapl_decimate.pipeline/pl_ssim.pipeline. Double-free is UB. The destroy loops inclose_fexstart ati = 1for decimate and skipi == 0for ssim_horiz. - Invariant — per-bundle descriptor pool. The shared
s->desc_poolis gone;alloc_descriptor_setnow takes aconst VmafVulkanKernelPipeline *bundleand usesbundle->desc_pool+bundle->dsl. Per-framevkFreeDescriptorSetscalls must target the matching pool (pl_decimate.desc_poolfor decimate sets,pl_ssim.desc_poolfor ssim sets) — mixing them is undefined behavior. - Numerical contract: unchanged. Same shaders, spec constants, push constants, and dispatch order as before.
float_ms_ssimNetflix-pair smoke (576×324×48f) reports mean 0.963241; ssim pyramid intermediate values bit-identical to pre-migration run. - Rebase impact: low. Upstream Netflix has no Vulkan backend. Conflicts only against the parallel
T-GPU-DEDUP-{18..22}PRs (#284–#288) onCHANGELOG.md/docs/rebase-notes.md— auto-resolve keeps both halves.
0106 — Vulkan kernel template multi-pipeline + ssim/motion migration (T-GPU-DEDUP-7)¶
- Touches:
core/src/vulkan/kernel_template.h— newvmaf_vulkan_kernel_pipeline_add_variant()helper. Takes the base pipeline bundle (DSL / pipeline layout / shader / pool owned byvmaf_vulkan_kernel_pipeline_create) plus a partialVkComputePipelineCreateInfoand produces a siblingVkPipelinere-using the same layout / shader. The base_createand_destroyentry points are unchanged; existing consumers (psnr, moment, ciede) keep working.core/src/feature/vulkan/motion_vulkan.c— state collapsesVkPipeline pipelines[2](kept "for SYCL parity" but functionally identical because COMPUTE_SAD goes through push constants, not spec-constants) to a singleVmafVulkanKernelPipeline pl.create_pipelines/close_fexshrink to template-driven create + destroy.core/src/feature/vulkan/ssim_vulkan.c— state becomesVmafVulkanKernelPipeline pl + VkPipeline pipeline_vert. Pass 0 (horizontal) is the template's base pipeline; pass 1 (vertical) is created via_add_variant().close_fexdestroys the variant first, then callsvmaf_vulkan_kernel_pipeline_destroy()on the bundle.- Invariant — no spec-constant drift between base and variant.
_add_variant()overwritessType/stage.sType/stage.stage/stage.module/layoutof the caller'sVkComputePipelineCreateInfoso the variant is guaranteed to share the base's shader and layout. Callers control the variant's spec-constant viapSpecializationInfo. Reordering these overwrites lets a consumer accidentally bind a different shader module under the same layout — UB at descriptor-set time. - Invariant — variant destroyed before bundle.
close_fexin ssim mustvkDestroyPipeline(s->pipeline_vert)beforevmaf_vulkan_kernel_pipeline_destroy(&s->pl)— the bundle's_destroyreleases the descriptor pool, which thevkAllocateDescriptorSetsissued against the variant pipeline's layout cleanly drops only when the variant pipeline is already gone. - Numerical contract: unchanged. Both kernels run identical shaders + spec-constants + push-constants as before; only the Vulkan boilerplate that creates / destroys the pipeline scaffolding moved to a shared owner. Cross-backend parity gate at
places=4holds — Netflix-pairfloat_ssimsmoke (576×324×48f) reports mean 0.863, identical to pre-migration. - Rebase impact: low. The base pipeline-bundle helpers predate this change (PR #270 / #271); the new
_add_variantis additive. Upstream Netflix has no Vulkan backend to conflict with.
0111 — integer_ciede_cuda migrated to kernel_template (T-GPU-DEDUP-11)¶
- Touches:
core/src/feature/cuda/integer_ciede_cuda.c— state'sCUstream + CUevent + CUevent + VmafCudaBuffer + host-pinned float*quintet collapses toVmafCudaKernelLifecycle lc + VmafCudaKernelReadback rb. init / collect / close call the template'slifecycle_init/readback_alloc/collect_wait/lifecycle_close/readback_freehelpers. submit keeps the pre-launch wait inline (intentional — ciede has no atomic, so the template's pre-launch memset is unnecessary).- Numerical contract: unchanged. Pure CUDA-boilerplate consolidation. The host-side reduction in collect still uses the same
doubleaccumulator over per-block float partials —places=4(ADR-0187) holds.
0112 — integer_moment_cuda migrated to kernel_template (T-GPU-DEDUP-12)¶
- Touches:
core/src/feature/cuda/integer_moment_cuda.c— state's stream/event/device-buffer/host-pinned quintet collapses toVmafCudaKernelLifecycle lc + VmafCudaKernelReadback rb. submit callsvmaf_cuda_kernel_submit_pre_launch(atomic counters require the device-side memset). init / collect / close call the matching template helpers.- Numerical contract: unchanged. Same per-frame atomic accumulators (4× uint64), same
sums_host[i] / n_pixelshost division. - Rebase impact: low. Upstream Netflix has no equivalent template; this consolidation is fork-local.
0113 — integer_motion_v2_cuda migrated to kernel_template (T-GPU-DEDUP-13)¶
- Touches:
core/src/feature/cuda/integer_motion_v2_cuda.c— stream/event pair + sad device+host quintet collapses tolc + rb. Raw-pixel ping-pongpix[2]stays outside the bundle. submit keeps the memset onpic_streaminline rather than callingsubmit_pre_launch(the helper would move the memset tolc.str, which races with the kernel reading the accumulator). init / collect / close call the matching template helpers.- Numerical contract: unchanged. Same D2D copy, same conditional kernel launch on frame ≥ 1, same host-side
min(score[i], score[i+1])flush.
0114 — integer_ssim_cuda migrated to kernel_template (T-GPU-DEDUP-14)¶
- Touches:
core/src/feature/cuda/integer_ssim_cuda.c— stream/event/partials device+host quintet collapses tolc + rb. Five intermediate float buffers (h_ref_mu,h_cmp_mu,h_ref_sq,h_cmp_sq,h_refcmp) stay outside the bundle. submit keeps thecuStreamWaitEvent + horiz + vert + DtoHchain inline — SSIM writes one float per block (no atomic), so the template'ssubmit_pre_launchmemset is unnecessary. init / collect / close use the matching template helpers.- Numerical contract: unchanged. Same horiz-then-vert two-pass pipeline, same per-block float partial reduction in double on the host.
places=4(matching the ciede_cuda precision pattern) holds. - Rebase impact: low. Upstream Netflix has no equivalent; this is fork-added.
0115 — ms_ssim_cuda + psnr_hvs_cuda lifecycle migration (T-GPU-DEDUP-15)¶
- Touches:
core/src/feature/cuda/integer_ms_ssim_cuda.c— stream + 2-event lifecycle replaced withVmafCudaKernelLifecycle lc; multi-level pyramid + SSIM intermediate + 3-partials buffers stay outside the template's single-pair readback bundle.core/src/feature/cuda/integer_psnr_hvs_cuda.c— same shape; 3-plane ref/dist/partials triples remain inline.- Numerical contract: unchanged. The migration only affects init / close boilerplate; submit / collect dispatch and host reduction paths are untouched apart from the
s->str→s->lc.str/s->event→s->lc.submit/s->finished→s->lc.finishedfield renames.
0116 — float_psnr/ansnr/motion cuda → kernel_template (T-GPU-DEDUP-16)¶
- Touches:
core/src/feature/cuda/float_psnr_cuda.c— stream/event/partials quintet →lc + rb; input upload buffersref_in/dis_instay outside the bundle.core/src/feature/cuda/float_ansnr_cuda.c— same shape; rb wraps the (sig, noise) interleaved partials.core/src/feature/cuda/float_motion_cuda.c— same shape; rb wraps the SAD partials,blur[2]ping-pong stays outside.- Numerical contract: unchanged. Same dispatch geometry, same reduction order. Cross-backend parity gate at the kernels' contracted precision (places=3 per ADR-0192) holds.
0117 — float_adm + float_vif cuda lifecycle migration (T-GPU-DEDUP-17)¶
- Touches:
core/src/feature/cuda/float_adm_cuda.c— stream + 2-event lifecycle replaced withVmafCudaKernelLifecycle lc; multi-stage DWT + CSF pipeline state stays outside the template's single-pair readback bundle.core/src/feature/cuda/float_vif_cuda.c— same shape; 4-level pyramid + per-scale (num, den) pairs remain inline.- Numerical contract: unchanged. The migration only affects init / close stream-event boilerplate; submit / collect dispatch and host reduction paths are untouched apart from the field renames.
- Rebase impact: low. Upstream Netflix has no equivalent template; this is fork-added.
0107 — float_psnr_vulkan migrated to kernel_template (T-GPU-DEDUP-8)¶
- Touches:
core/src/feature/vulkan/float_psnr_vulkan.c— state'sdsl + pipeline_layout + shader + pipeline + desc_poolquintet is collapsed into a singleVmafVulkanKernelPipeline pl;create_pipelinesandclose_fexshrink to template-driven create + destroy. No shader changes, no spec-constant changes, no push-constant changes.- Numerical contract: unchanged. The migration is a pure Vulkan-boilerplate consolidation. Cross-backend parity gate at
places=4holds — Netflix-pair smoke reportsfloat_psnrmean 30.755 dB, identical to pre-migration.
0109 — float_ansnr_vulkan + motion_v2_vulkan migrated to kernel_template (T-GPU-DEDUP-9)¶
- Touches:
core/src/feature/vulkan/float_ansnr_vulkan.c— single-pipeline state collapses toVmafVulkanKernelPipeline pl;create_pipelinesandclose_fexshrink to template-driven create + destroy.core/src/feature/vulkan/motion_v2_vulkan.c— same shape.- Numerical contract: unchanged. Pure Vulkan-boilerplate consolidation. Cross-backend parity gate at the kernel's contracted precision holds — Netflix-pair smoke reports
float_ansnrmean 23.51 dB andmotion2_v2_scoremean 3.895, identical to pre-migration.
0110 — float_motion_vulkan migrated to kernel_template (T-GPU-DEDUP-10)¶
- Touches:
core/src/feature/vulkan/float_motion_vulkan.c— single-pipeline state collapses toVmafVulkanKernelPipeline pl;create_pipelinesandclose_fexshrink to template-driven create + destroy.- Numerical contract: unchanged. Pure Vulkan-boilerplate consolidation. Netflix-pair smoke reports
motionmean 4.049 /motion2mean 3.894, identical to pre-migration. - Rebase impact: low. Upstream Netflix has no Vulkan backend.
0108 — Bristol VI-Lab feasibility digest + BVI-CC ingest ADR (Draft)¶
- Touches:
docs/research/0046-bristol-vi-lab-feasibility.md(new) — nine-dataset survey + use-case fit + effort estimate.docs/adr/0241-bristol-bvi-cc-ingest.md(new, Status: Draft) — proposal to ingest BVI-CC as the second tiny-AI corpus.docs/adr/README.md— index row for ADR-0241.CHANGELOG.md— Added entry.- Numerical contract: not applicable (docs-only).
- Rebase impact: none. Pure research deliverables; upstream Netflix has no equivalent surface.
0094 — Vulkan VkImage import v2 async pending-fence (T7-29 part 4 / ADR-0251)¶
- ADR: ADR-0251; predecessor ADR-0186.
- Touches:
core/src/vulkan/import.c— full rewrite of the submission path. Single-fencesubmit_and_waitbecomes per-slotsubmit_to_slot+drain_slot_fence; the newslot_alloc/slot_releasehelpers materialise / tear down a ring slot (staging-pair + cmd buffer + fence).vmaf_vulkan_import_imageindexes into the ring byframe_index % ring_size;vmaf_vulkan_wait_computedrains every outstanding fence.vmaf_vulkan_state_build_pictureswaits the slot's fence before exposing the host pointer. Public-API signatures are unchanged.core/src/vulkan/vulkan_internal.h— newstruct VmafVulkanImportSlot;VmafVulkanImportSlotsbecomes a fixed-capacityVmafVulkanImportSlot ring[VMAF_VULKAN_RING_MAX]plus geometry +ring_size. Two new defines —VMAF_VULKAN_RING_DEFAULT(4) andVMAF_VULKAN_RING_MAX(8).VmafVulkanStategainsrequested_ring_size.core/src/vulkan/common.c—vmaf_vulkan_state_initand_state_init_externalsetrequested_ring_size = VMAF_VULKAN_RING_DEFAULT.core/test/test_vulkan_async_pending_fence.c(new, contract smoke for the v1 → v2 swap).core/test/meson.build— registers the new test under the existingenable_vulkanguard.core/src/vulkan/AGENTS.md(new) — pins the three rebase-sensitive ring invariants.docs/adr/0251-vulkan-async-pending-fence.md(new),docs/research/0042-vulkan-async-pending-fence.md(new),docs/api/gpu.md,docs/backends/vulkan/overview.md,CHANGELOG.md,docs/rebase-notes.md.ffmpeg-patches/0006-libvmaf-add-libvmaf-vulkan-filter.patch— unchanged. The v2 ring is fully internal toVmafVulkanState; the public ABI stays byte-identical so the filter consumes the new path transparently.- Invariant 1 — fixed ring depth at first import.
lazy_alloc_ringis the only place that materialises the ring; once allocated the depth never changes for the lifetime of theVmafVulkanState. Any caller that needs a different depth has to free + re-init. The geometry pinning contract from v1 (ADR-0186) is preserved verbatim. - Invariant 2 —
vkResetFencesonly afterVK_SUCCESSfromvkWaitForFences. Sole reset path lives indrain_slot_fence;fence_in_flightflips back to 0 only after the wait succeeds. A-EIOfrom the wait propagates up without resetting (so a retry would correctly re-wait rather than silently move on). - Invariant 3 —
state_freedrains before destroying.vmaf_vulkan_import_slots_freewalks the ring and callsdrain_slot_fenceon every in-flight slot, then issues onevkQueueWaitIdlebelt-and-braces (any feature kernel that submitted on the same queue may still be running). Reordering this triggers validation-layer "destroying in-use object" errors. - Numerical contract: unchanged. Async submission only changes when the host can read the staging buffer, not which bytes the GPU writes. Cross-backend parity gate (
scripts/ci/cross_backend_parity_gate.py,places=4) holds. - Memory delta: staging arena scales
1 → ring_sizeper direction. At default depth and 1080p 8-bit Y, the per-state host-visible footprint grows from ~4 MiB to ~16 MiB. Documented in ADR-0251 §Consequences.
0090 — cambi_vulkan extractor (T7-36 / ADR-0210)¶
- ADR: ADR-0210; predecessor ADR-0205.
- Touches:
core/src/feature/vulkan/cambi_vulkan.c(replaces the spike scaffold'sinit_stub/extract_stub/close_stubtriple with the full Vulkan-aware lifecycle).core/src/feature/vulkan/shaders/cambi_preprocess.comp(new),cambi_mask_dp.comp(new — unified row-SAT / col-SAT / threshold-compare viaPASS=0/1/2spec const).core/src/feature/cambi.c— appends a small block of public trampolines (vmaf_cambi_*) at the bottom of the file that thinly wrap the file-static helpers. No upstream function-static code is renamed or moved; the entire upstream body of cambi.c above the trampolines stays byte-identical, which keeps Netflix sync straightforward.core/src/feature/cambi_internal.h(new) — internal-only header exposingvmaf_cambi_calculate_c_values,vmaf_cambi_get_spatial_mask, etc., to the GPU twin.core/src/vulkan/meson.build— registers the 5 cambi shaders invulkan_shader_sources[]andcambi_vulkan.cinvulkan_sources.core/src/feature/feature_extractor.c— adds the extern decl + registry entry forvmaf_fex_cambi_vulkanunder#if HAVE_VULKAN.scripts/ci/cross_backend_vif_diff.py—cambirow inFEATURE_METRICSso the cross-backend gate runs atplaces=4against the CPU baseline.docs/adr/0210-cambi-vulkan-integration.md,docs/research/0032-cambi-vulkan-integration.md,docs/backends/vulkan.md,CHANGELOG.md.- Invariant 1 — bit-exactness by construction. Every GPU phase is integer arithmetic (
uint16derivative,int32SAT,>compare, stride-2 gather, 3-elementmode3lookup). The readback into the hostVmafPicturepair is byte-identical to what the CPU would have written; the host residual then runs the unmodified CPUcalculate_c_values+ spatial pooling on those buffers. Any rebase that introduces float arithmetic into one of these GPU phases — e.g., a future Netflix change to the derivative kernel that adds a bilinear interpolation step — will silently breakplaces=4and must be caught at the cross-backend gate. - Invariant 2 —
cambi_internal.hsignatures must stay in lock-step with cambi.c's file-static helpers. The Vulkan twin callsvmaf_cambi_calculate_c_values, which trampolines to the file-staticcalculate_c_values. Any signature change to the latter (extra parameters, type changes) must update the trampoline + header in the same PR or the GPU build breaks. - On upstream sync: cambi.c's file-static helpers are sometimes renamed by upstream (e.g.,
decimate→cambi_decimatewould happen during a Netflix tidy-up). When rebasing, search cambi.c's tail for the trampoline block — its fivestaticcalls (get_spatial_mask,decimate,filter_mode,calculate_c_values,spatial_pooling,weight_scores_per_scale,get_pixels_in_window,increment_range,decrement_range,get_derivative_data_for_row,cambi_preprocessing) need to match the upstream symbol names. Update the trampoline body if upstream renames; signatures should not need to change because the trampoline already takes the function-pointer-typedef form (VmafRangeUpdateretc.). - Re-test on rebase:
python3 scripts/ci/cross_backend_vif_diff.py --backend vulkan --feature cambi --ref testdata/ref_576x324_48f.yuv --dist testdata/dis_576x324_48f.yuv --width 576 --height 324 --pixel-format 420 --bitdepth 8 --frames 48. Should emitplaces=4 PASSwithmax_abs_diff = 0.0. If it diverges, bisect the GPU phases by reading back individual buffers (image_buf/mask_buf/deriv_buf) and comparing against the CPU's in-placepicplane after the equivalent stage.
The pre-ADR-0108 fork-local PRs are summarised by workstream rather than per-PR. Future PRs add entries individually.
0085 — Upstream c70debb1 partial port (adm_csf + barten_csf tests)¶
- No ADR. Pure upstream cherry-pick per ADR-0108 carve-out ("pure upstream syncs and
port-upstream-commitPRs are exempt"). - Upstream source:
c70debb1(Kyle Swanson, 2026-04-28): "libvmaf/test: port new adm/vif/speed tests". The audit row that flagged the gap is T-NEW-2 in the 2026-04-29 quarterly upstream-backlog re-audit (PR #205). - Touches (additive only):
core/src/feature/adm_csf_tools.h— new header (verbatim from upstream); declares the inlineadm_native_csfhelper (DLM-paper CSF) used by the newtest_adm_csfunit.core/test/test_adm_csf.c— new unit (verbatim from upstream); 2mu_assertcases onadm_native_csf(3, 3.0, 1080, {0, 45}).core/test/test_barten_csf.c— new unit (verbatim from upstream); 23mu_assertcases overbarten_rod_cone_sens,barten_mtf,barten_csf,linear_interpolate,barten_watson_blend_csf(all symbols already on the fork).core/test/meson.build— registers the two new executables + addstest('test_adm_csf', ...)andtest('test_barten_csf', ...).CHANGELOG.mdUnreleased § Changed.- Deliberate scope cuts (the upstream commit's other halves are not portable verbatim):
test_vif_tools.c— depends on upstream symbolsNUM_KERNELSCALES, the 21-entryvalid_kernelscalestable,vif_validate_kernelscale,vif_get_filter_size,vif_get_filter,speed_get_antialias_filter, and a[NUM_KERNELSCALES][5][65]filter table that the fork'svif_filter1d_table_s [11][4][65]does not match. Per Research-0024 Strategy E, the fork deliberately diverges from the upstreamvifruntime-helper chain to preserve the ADR-0138 / 0139 / 0142 / 0143 SIMD bit-exactness contract. Porting this test requires porting the runtime helpers first.test_speed_chroma.c—#includesfeature/speed.cdirectly; the fork has no SpEED extractor (feature/speed.cdoes not exist). Pairs with audit row T-NEW-1 (port the SpEED extractor wholesale, or absorb it into the tiny-AI speed metric).- Invariants (rebase-relevant):
- The new
adm_csf_tools.hheader is wholly additive and does not conflict with the existing forkadm_csf_snon-inline helper inadm_tools.h(different signature, different translation units). - The two new tests do not depend on Netflix golden YUVs — they evaluate the closed-form CSF math directly. No golden-data interaction.
- On upstream sync: a future port of the upstream
vifruntime-helper chain (Research-0024 Strategy A reversal) or the SpEED extractor (T-NEW-1) unlocks the deferred halves of this commit. Until then, fork-sidetest_vif_tools.c/test_speed_chroma.cstay absent. - Re-test on rebase:
meson setup build-cpu libvmaf -Denable_cuda=false -Denable_sycl=false
ninja -C build-cpu test_adm_csf test_barten_csf
meson test -C build-cpu test_adm_csf test_barten_csf
0084 — Embedded MCP server scaffold (T5-2, ADR-0209)¶
- ADR: ADR-0209 (audit-first scaffold) on top of the ADR-0128 governance + Research-0005 design.
- Upstream source: fork-local. Netflix/vmaf has no embedded MCP server (and no plans to add one — the workflow is agent-tooling-specific, well outside upstream's library scope).
- Touches:
core/include/libvmaf/libvmaf_mcp.h— new public header.core/include/core/meson.build— newif get_option('enable_mcp')install branch.core/src/mcp/— new directory:mcp.c(stub TU) +meson.build(exposesmcp_sources+mcp_defines).core/src/meson.build— newis_mcp_enabledguard +subdir('mcp')block;mcp_sourcesthreaded into thelibrary('vmaf', ...)source list alongsidednn_sources.core/test/meson.build— newif get_option('enable_mcp')block wiringtest_mcp_smoke.core/test/test_mcp_smoke.c— new 12-sub-test smoke.core/meson_options.txt— newenable_mcpumbrella + three sub-flags (all defaultfalse).- Invariant: every public entry point in
libvmaf_mcp.h(vmaf_mcp_init/_start_sse/_start_uds/_start_stdio/_stop/_close) returns-ENOSYS(or-EINVALon bad arguments) until the T5-2b runtime PR lands. The smoke pins this contract — a runtime PR that flips a return code without flipping the smoke expectation regresses the gate. - On upstream sync: zero interaction with upstream files. Wholly additive directory + boolean build flags. The
subdir('mcp')insertion incore/src/meson.buildlives next to the existingsubdir('dnn')/ Vulkan blocks; an upstream conflict in that area would be confined to those few lines and is mechanical to resolve. - Re-test on rebase:
meson setup build-cpu libvmaf -Denable_cuda=false -Denable_sycl=false -Denable_mcp=false
ninja -C build-cpu && meson test -C build-cpu # baseline still green
meson setup --reconfigure build-cpu libvmaf -Denable_mcp=true \
-Denable_mcp_sse=true -Denable_mcp_uds=true -Denable_mcp_stdio=true
ninja -C build-cpu
meson test -C build-cpu test_mcp_smoke # 12/12 sub-tests pass
0065 — T7-37 Netflix bench rerun + docs/benchmarks.md TBD fill¶
- No ADR. Empirical fill of pre-existing
TBDcells; no new decision. The bench script fixes that this rerun depends on shipped earlier under PR #169 (libvmaf/AGENTS.md backend-engagement foot-guns), PR #170 (--backend cudaactually engages CUDA), and PR #171 (testdata/bench_all.shuses correct flags). Vulkan header install for SDK consumers is PR #175. - Touches (additive only):
docs/benchmarks.md(everyTBDcell replaced with measured numbers; hardware-profile table updated to theryzen-4090-archost the rerun was performed on; "How to reproduce" section now documents fixture acquisition for the gitignored BBB 4K 200-frame pair).CHANGELOG.mdUnreleased § Changed entry. - Invariants (rebase-relevant): none. The numbers are tied to fork commit
41301496and theryzen-4090-arcprofile; an upstream rebase that changes feature pipelines would invalidate the table but not break parsing. - On upstream sync: zero interaction. Pure docs.
- Re-test on rebase:
bash testdata/bench_all.sh(after a fresh fork build) — confirms the bench script drives every live backend and records each row's emitted metrics-key count. A GPU count collapsing to CPU is a fallback warning to corroborate with pool and throughput; never compare against fixed expected counts.
0050 — float_adm_cuda + float_adm_sycl extractors (ADR-0202)¶
- ADR: ADR-0202
- Touches:
core/src/feature/cuda/float_adm/float_adm_score.cu(new)core/src/feature/cuda/float_adm_cuda.{c,h}(new)core/src/feature/sycl/float_adm_sycl.cpp(new)core/src/meson.build— three changes: (1) newfloat_adm_scoreentry incuda_cu_sources, (2) newcuda_cu_extra_flagsdict that threads--fmad=false+-Xcompiler=-ffp-contract=offinto thefloat_adm_scorefatbin only, (3) new SYCL source insycl_feature_sources.core/src/feature/feature_extractor.c(extern decls + list entries forvmaf_fex_float_adm_cuda/vmaf_fex_float_adm_syclunder#if HAVE_CUDA/#if HAVE_SYCL).- Invariant 1 —
--fmad=falsefor the float_adm fatbin only: the angle-flag dot product (ot_dp = oh*th + ov*tv) and the cube reductions (xa*xa*xa,csf_o*csf_o*csf_o) require IEEE-754 add/mul ordering to match the GLSLprecisequalifier infloat_adm.comp. NVCC's default-fmad=truefuses these and drifts pastplaces=4at scale 3 / adm2. The integer ADM kernels sharecuda_flagsbut useint64accumulators where FMA is irrelevant — keep the FMA-on default for them. - Invariant 2 — parent-LL dimension trap: stage 0 at
scale > 0reads the parent's LL band; the mirror/clamp bounds arescale_w/h[scale](= parent's LL output dims = current scale's input dims), NOTscale_w/h[scale - 1](= parent's full image dims). Bothfloat_adm_cuda.candfloat_adm_sycl.cppcite this inline. Do not "simplify" by using the off-by-one neighbour. - Re-test:
CXX=icpx CC=icx meson setup build-cs -Denable_cuda=true \
-Denable_sycl=true -Denable_vulkan=enabled \
-Denable_float=true \
-Dsycl_compiler=/opt/intel/oneapi/compiler/latest/bin/icpx
ninja -C build-cs
python3 scripts/ci/cross_backend_vif_diff.py \
--vmaf-binary build-cs/tools/vmaf \
--reference python/test/resource/yuv/src01_hrc00_576x324.yuv \
--distorted python/test/resource/yuv/src01_hrc01_576x324.yuv \
--width 576 --height 324 --feature float_adm \
--backend cuda --places 4
# Same with --backend sycl on a host with an SYCL device.
# Both must report 0/N mismatches at places=4.
0049 — float_adm_vulkan extractor (ADR-0199)¶
- ADR: ADR-0199
- Touches:
core/src/feature/vulkan/float_adm_vulkan.c(new)core/src/feature/vulkan/shaders/float_adm.comp(new)core/src/vulkan/meson.build(adds the .comp shader and the new .c source)core/src/feature/feature_extractor.c(extern decl + list entry under#if HAVE_VULKAN)scripts/ci/cross_backend_vif_diff.py(float_admentry inFEATURE_METRICS).github/workflows/tests-and-quality-gates.yml(lavapipefloat_admstep atplaces=4)- Invariant: float_adm GPU port uses the
2 * sup - idx - 1mirror form on both axes — matches both the scalaradm_dwt2_sand the AVX2float_adm_dwt2_avx2, which both consume the samedwt2_src_indices_filt_sindex buffer. This is intentionally different from float_vif's GPU mirror (ADR-0197), which uses-2because float_vif's AVX2 path takes a different code branch. Do not "fix" the asymmetry by analogy with float_vif. - Re-test:
meson setup build-vk -Denable_vulkan=enabled -Denable_cuda=false \
-Denable_sycl=false
ninja -C build-vk
meson test -C build-vk
VK_LOADER_DRIVERS_SELECT='*lvp*' python3 \
scripts/ci/cross_backend_vif_diff.py \
--vmaf-binary build-vk/tools/vmaf \
--reference python/test/resource/yuv/src01_hrc00_576x324.yuv \
--distorted python/test/resource/yuv/src01_hrc01_576x324.yuv \
--width 576 --height 324 --feature float_adm --places 4
0083 — SSIMULACRA 2 Vulkan kernel (ADR-0201)¶
- ADR: ADR-0201
- Upstream source: fork-local. No SSIMULACRA 2 extractor in upstream Netflix/vmaf — fully fork-local feature.
- Touches:
core/src/feature/vulkan/ssimulacra2_vulkan.c(new file).core/src/feature/vulkan/shaders/ssimulacra2_xyb.comp,ssimulacra2_blur.comp,ssimulacra2_mul.comp,ssimulacra2_ssim.comp(4 new shader files).core/src/vulkan/meson.build— added 4 shaders tovulkan_shader_sourcesand 1 source tovulkan_sources; added all 4 ssimulacra2 shaders topsnr_hvs_strict_shaders(the-O0strict-mode list, kept its legacy name).core/src/feature/feature_extractor.c— registeredvmaf_fex_ssimulacra2_vulkanin the Vulkan branch of the extractor list (betweenpsnr_hvs_vulkanand the CUDA block).scripts/ci/cross_backend_vif_diff.py— addedssimulacra2toFEATURE_METRICS.- Rebase impact: low — fully additive, no upstream-shared files modified beyond
feature_extractor.c's registry array (which always grows on every new extractor and is not a rebase pain point). - Verification command:
meson setup core/build-vk-ss2 \
-Denable_vulkan=enabled -Denable_cuda=false -Denable_sycl=false \
libvmaf
ninja -C core/build-vk-ss2 tools/vmaf
python3 scripts/ci/cross_backend_vif_diff.py \
--vmaf-binary core/build-vk-ss2/tools/vmaf \
--reference python/test/resource/yuv/src01_hrc00_576x324.yuv \
--distorted python/test/resource/yuv/src01_hrc01_576x324.yuv \
--width 576 --height 324 \
--feature ssimulacra2 --backend vulkan --places 1
# expected: max_abs_diff ≈ 1.59e-2, 0/48 mismatches at places=1
- Follow-ups:
- CUDA + SYCL twins (batch 3 parts 7b + 7c per ADR-0192).
- Performance follow-up: re-bin multiple rows / columns per WG in the IIR blur (currently
local_size = 1, one row/col per WG for correctness). - Optional: rename
psnr_hvs_strict_shaderstostrict_shadersincore/src/vulkan/meson.build(cosmetic — out of scope for this PR).
0001 — SIMD bit-identical reductions for float ADM¶
- Workstream PRs: #18, commits
24c88a32,f082cfd3. - Touches:
core/src/feature/integer_adm.c,core/src/feature/float_adm.c,core/src/feature/x86/adm_avx2.c,core/src/feature/x86/adm_avx512.c,core/src/feature/arm64/adm_neon.c, upstreampython/test/feature_extractor_test.pytest expectations. - Invariant:
sum_cubeandcsf_den_scaleaccumulate cubed values in double precision (via_mm256_cvtps_pd/_mm512_cvtps_pd) in scalar, AVX2, AVX-512, and NEON. Upstream accumulates in float, which produces ~8e-5 drift between scalar and SIMD. Test expectations were tightened to match the double-precision path; an upstream-side accumulator change would re-introduce the drift and break the tightened assertions. - Re-test:
meson test -C build --suite=fast && python -m pytest python/test/feature_extractor_test.py -k adm.
0002 — CUDA ADM decouple-inline buffer elimination¶
- Workstream PRs: commit
787e3382. - Touches:
core/src/feature/cuda/integer_adm_cuda.cu,core/src/feature/cuda/adm_decouple_inline.cuh(new),core/src/feature/cuda/meson.build. Upstream'sadm_decouple.cuis no longer compiled in the fork. - Invariant: CSF and CM CUDA kernels read
ref/disDWT2 buffers directly and computedecouple_r/decouple_ainline via__device__helpers inadm_decouple_inline.cuh. The 6 intermediate buffers (decouple_r,decouple_a,csf_a× {scale-0 int16, scales 1-3 int32}) and the standaloneadm_decouple.cusource are intentionally removed. ~107 MB GPU memory savings at 4K. An upstream change toadm_decouple.cuwill look orphaned and a literal merge would re-introduce the buffer allocations. - Re-test:
meson setup build -Denable_cuda=true && ninja -C build && meson test -C build --suite=cuda.
0003 — SYCL backend (USM pool / D3D11 import / vmaf_sycl_* API)¶
- Workstream PRs: #33, #35, #5 (initial scaffolding), and the picture-pool deadlock fix that landed via #32.
- Touches:
core/include/libvmaf/libvmaf_sycl.h,core/src/sycl/,core/src/feature/sycl/,core/src/libvmaf.c(SYCL public-API entry points),meson_options.txt(enable_sycl). - Invariant:
vmaf_sycl_preallocate_picturesconstructs a realVmafSyclPicturePoolhonoringVmafSyclPicturePreallocationMethod(NONE/DEVICE/HOST);vmaf_sycl_picture_fetchdispatches to the pool when configured. The whole SYCL tree is fork-local and has no upstream counterpart — upstream changes tocore/src/libvmaf.cnear the SYCL entry-point block are likely to conflict. Picture-pool error paths invmaf_read_pictures(libvmaf.c) mustgoto cleanup;rather thanreturn err;to avoid leaking ref/dist pictures into the live-picture set (closes the always-on-pool deadlock fixed in #32 — see ADR-0104). See ADR-0101, ADR-0103, ADR-0104. - Re-test:
meson setup build -Denable_sycl=true && ninja -C build && meson test -C build --suite=sycl(requires oneAPI / icpx).
0004 — DNN runtime + tiny-AI surfaces¶
- Workstream PRs: #5, #8, #21, #22, #23, #31, #34, plus the pre-numbered DNN feat commits (
9b985946,1e5336d3,d122b721). - Touches:
core/include/libvmaf/dnn.h,core/src/dnn/,core/src/feature/feature_lpips.c,model/tiny/,meson_options.txt(enable_onnxruntime). - Invariant: ordered EP selection (CUDA → DML → CPU) with graceful fallback (ADR-0102);
fp16_iodoes host-side fp32↔fp16 cast on the scoring path;VMAF_TINY_MODEL_DIRenforces a path jail on model load (PR #31); the runtime op-allowlist (PR #21) walks the ONNX graph and rejects unknown ops + bounds Loop/Iftrip_countat 1024 (ADR-0036/0107). DNN tree is fork-local; upstream has no DNN code yet, so conflicts here are unlikely but themeson_options.txtandcore/src/meson.buildblocks near the DNN flag may collide. - Re-test:
meson setup build -Denable_onnxruntime=true && ninja -C build && meson test -C build --suite=dnn.
0005 — --precision CLI flag (IEEE-754 round-trip lossless)¶
- Workstream PRs: commit
c989fbd9. - Touches:
core/tools/vmaf.c,core/tools/cli_parse.c,core/include/libvmaf/libvmaf.h(addedvmaf_write_output_with_format),core/src/output.c. - Invariant: default
--precisionis%.17g(round-trip lossless);legacyopts back into upstream's%.6f; the public C API gainedvmaf_write_output_with_formatand the oldvmaf_write_outputroutes through it with the%.17gdefault. ABI-breaking only if upstream adds a same-named function with a different signature. See ADR-0006. - Re-test:
vmaf -r ref.yuv -d dis.yuv ... --precision=fulland diff against--precision=legacy.
0006 — Netflix golden tests preserved verbatim as required gate¶
- Workstream PRs: across the fork's life; codified in ADR-0024.
- Touches:
python/test/quality_runner_test.py,python/test/vmafexec_test.py,python/test/vmafexec_feature_extractor_test.py,python/test/feature_extractor_test.py,python/test/result_test.py,python/test/resource/yuv/. - Invariant:
assertAlmostEqual(...)golden values in the five upstream Python test files are never modified by this fork. Fork-added tests live in separate files (e.g.python/test/test_precision_flag.py). The CI gate "Netflix CPU golden tests (D24)" is required and blocks merge. Upstream changes to these files are accepted unless they relax the assertions. - Re-test:
make test-netflix-golden.
0007 — Build system (CUDA 13.2, oneAPI 2025.3, MkDocs migration)¶
- Workstream PRs: #7, #17, commit
8a995cb0. - Touches:
meson.build,meson_options.txt, top-levelMakefile,docs/(Sphinx → MkDocs Material migration —docs/conf.pyremoved,mkdocs.ymladded),docs/requirements.txt,Dockerfile.*, distro install scripts underscripts/. - Invariant: image pins are non-conservative (ADR-0027) — CUDA 13.2, oneAPI 2025.3, clang-format 22, black 26 — and ship experimental toolchain flags (
--expt-relaxed-constexpr, etc.) deliberately. An upstream sync that pulls in a Dockerfile change targeted at older CUDA or older oneAPI must not relax the pins. - Re-test:
meson setup build -Denable_cuda=true -Denable_sycl=true && ninja -C build && mkdocs build --strict.
0008 — Workspace / docs / MATLAB / resource-tree relocations¶
- Workstream PRs: codified across ADR-0026, ADR-0029, ADR-0030, ADR-0031, ADR-0032, ADR-0033, ADR-0034, ADR-0038.
- Touches: any path-walk in upstream's CI / scripts / docs that assumes the upstream layout (root-level
workspace/,resource/,matlab/, rootunittestscript, rootpatches/). - Invariant: the fork's layout is
python/vmaf/workspace/,python/vmaf/resource/,python/vmaf/matlab/,scripts/unittest,ffmpeg-patches/only,.github/codeql-config.yml. Upstream moves to a different sub-tree (e.g. a hypotheticaltools/workspace/) need to either be applied via a corresponding fork-side relocation or rejected with a rebase note. - Re-test:
python -m pytest python/test/ -k golden(verifies the resource-tree path works);make test-netflix-golden.
0009 — License headers (Lusoris/Claude on wholly-new files¶
2016–2026 on Netflix files)
- Workstream PRs: commits
c159761d,a185f8ef,0e98c949, codified in ADR-0025 / ADR-0105. - Touches: every wholly-new fork file (notably the SYCL tree and
core/src/dnn/) and every Netflix-touched file (year range2016 → 2016–2026). - Invariant: wholly-new fork files carry
Copyright 2026 Lusoris and Claude (Anthropic)under the same BSD-3-Clause-Plus-Patent license; mixed files use a dual-copyright notice. An upstream commit that resets a Netflix file's year range (e.g. back to2016–2020) must be partially rejected — keep the fork's2016–2026. - Re-test: grep that wholly-new fork files retain the Lusoris/Claude header (
grep -L "Copyright 2026 Lusoris" core/src/sycl/*.cpp— expected to match nothing).
0010 — .claude/ agent scaffolding + ADR tree + AGENTS.md / CLAUDE.md¶
- Workstream PRs: #14, #24, #37, plus continuous additions.
- Touches:
.claude/,AGENTS.md,CLAUDE.md,docs/adr/,.github/PULL_REQUEST_TEMPLATE.md. - Invariant: this whole tree is fork-local and has no upstream counterpart. Upstream additions to
.github/(issue templates, workflows) need to merge cleanly with the fork's existing files rather than replacing them. The ADR tree's IDs ≤ 0099 are backfills; new decisions start at 0100 (ADR-0028 / ADR-0106). - Re-test: visual review of
.github/anddocs/adr/README.mdafter the merge.
Pre-ADR-0108 entries above are the result of a one-shot backfill sweep on 2026-04-18; subsequent fork-local PRs add their own entries inline.
0011 — Nightly bisect-model-quality + fixture cache¶
- Workstream PRs: closes #4; sticky tracker issue #40.
- Touches:
.github/workflows/nightly-bisect.yml,ai/scripts/build_bisect_cache.py,ai/testdata/bisect/{features.parquet, models/*.onnx, README.md},scripts/ci/post-bisect-comment.py,docs/ai/bisect-model-quality.md,docs/adr/0109-nightly-bisect-model-quality.md,docs/research/0001-bisect-model-quality-cache.md,mkdocs.yml(nav). - Invariant: the committed parquet + ONNX bytes under
ai/testdata/bisect/must regenerate byte-identically fromai/scripts/build_bisect_cache.pywith seedsFEATURE_SEED=20260418andMODEL_SEED=20260419. The CI--checkstep asserts this before every bisect run, so any upstream pull that bumpspandas/pyarrow/onnxenough to change the serialiser bytes will fail the workflow until the cache is regenerated and committed. - Re-test:
python ai/scripts/build_bisect_cache.py --check
vmaf-train bisect-model-quality \
ai/testdata/bisect/models/model_*.onnx \
--features ai/testdata/bisect/features.parquet \
--min-plcc 0.85 --input-name input
# Expected: "no regression in this range"; first_bad_index None.
Pure upstream code is not touched, so no Netflix-side conflict vector. Only fork-local files; risk is toolchain drift, not merge conflict.
0012 — Upstream ADM port (Netflix 966be8d5)¶
- Workstream PRs: this PR; ports a single upstream commit.
- Touches:
core/src/feature/integer_adm.{c,h},core/src/feature/x86/adm_avx2.{c,h},core/src/feature/x86/adm_avx512.{c,h},core/src/feature/alias.c,core/src/feature/barten_csf_tools.h(new upstream file). - Invariant: the eight ADM files now mirror upstream's content byte-for-byte (modulo our clang-format-22 pass and the Netflix copyright-year bump on the new header). Future
/sync-upstreamruns can take new upstream ADM commits cleanly. Do not revert to a pre-966be8d5ADM kernel without also reverting the call-site signatures ininteger_compute_adm— upstream extendedi4_adm_cmfrom 8 to 13 args. - Re-test:
ninja -C core/build && meson test -C core/build
core/build/tools/vmaf -r python/test/resource/yuv/src01_hrc00_576x324.yuv \
-d python/test/resource/yuv/src01_hrc01_576x324.yuv \
-w 576 -h 324 -p 420 -b 8 \
--model version=vmaf_v0.6.1 -o /tmp/vmaf-port.json
grep '<metric name="vmaf"' /tmp/vmaf-port.json
# Expected: mean ≈ 76.66890 (golden 76.66890519623612, places=4 OK).
0013 — Upstream motion port (Netflix PR #1486 head 2aab9ef1)¶
- Workstream PRs: this PR; ports upstream PR #1486 (4 commits on top of
966be8d5ADM base, head2aab9ef1). Sister to entry 0012. - Touches:
core/src/feature/integer_motion.{c,h},core/src/feature/motion_blend_tools.h(new upstream file),core/src/feature/x86/motion_avx2.c,core/src/feature/x86/motion_avx512.c,core/src/feature/alias.c(additive:integer_motion3row),python/test/{quality_runner,vmafexec,feature_extractor,vmafexec_feature_extractor}_test.py(golden tolerance updates:places=4→places=2on motion-affected asserts; expected values unchanged). - Invariant: motion files mirror upstream byte-for-byte (modulo our clang-format-22 pass). The
alias.crow forinteger_motion3was inserted surgically to avoid clobbering the AVX-512 ADM registration added by entry 0012; new motion3 metric appears in default VMAF model output but is not standalone-loadable via--feature integer_motion3(sub-feature only). Netflix golden VMAF mean shifts76.668904824→76.667830213(well withinplaces=2tolerance the upstream PR loosened to). Do not revertplaces=4on motion-touching assertions without also reverting the motion code. - Re-test:
ninja -C core/build && meson test -C core/build
core/build/tools/vmaf -r python/test/resource/yuv/src01_hrc00_576x324.yuv \
-d python/test/resource/yuv/src01_hrc01_576x324.yuv \
-w 576 -h 324 -p 420 -b 8 \
--model version=vmaf_v0.6.1 -o /tmp/vmaf-motion-port.json
grep -E '<metric name="vmaf"|integer_motion3' /tmp/vmaf-motion-port.json
# Expected: vmaf mean ≈ 76.66783; integer_motion3 mean ≈ 3.98976.
0014 — Coverage gate overhaul + upstream python/test/ reformat¶
- Workstream PRs: this PR (coverage-gate overhaul + in-tree reformat of upstream-mirror Python tests).
- Touches:
.github/workflows/ci.yml(CPU + GPU coverage jobs:-Dc_args=-fprofile-update=atomic/-Dcpp_args=-fprofile-update=atomic,meson test --num-processes 1,-Denable_dnn=enabled, ORT install step on the CPU coverage job,lcov/geninforeplaced bygcovrwith--json-summary/--xml/--txtoutput, artifact renamecoverage-lcov-{cpu,gpu}→coverage-{cpu,gpu}),scripts/ci/coverage-check.sh(rewritten to parse gcovr JSON viapython3 -c— same CLI signature),core/src/dnn/dnn_api.c+ newcore/src/dnn/dnn_attach_api.c(vmaf_use_tiny_modelcarved out into its own TU so the unit-test binaries — which pull indnn_sourcesforfeature_lpips.cbut never linklibvmaf.c— don't end up with an undefined reference tovmaf_ctx_dnn_attachonceenable_dnn=enabledactivates the real bodies),core/src/dnn/meson.build+core/src/meson.build(newdnn_libvmaf_only_sourceslist wired intolibvmaf.soonly),python/test/{feature_extractor,quality_runner,vmafexec,vmafexec_feature_extractor}_test.py(mechanical Black + isort reformat — no assertion values changed, imports regrouped, line wrapping normalised). - Invariant: coverage CI must keep all five pieces in lockstep — (a)
-fprofile-update=atomiccloses the intra-process counter race on SIMD inner loops (vif_avx2.c:673,motion_avx2, etc.) → negative counts →geninfo/gcovr abort; (b)--num-processes 1closes the inter-process race where multiple parallel test binaries merge their counters into the same.gcdafiles for the sharedlibvmaf.soat process exit (per-thread atomicity does not cover this); (c)gcovrdeduplicates.gcnofiles belonging to the same source compiled into multiple targets — without dedup, lcov sums hits across compilation units and yields impossible100% values (
dnn_api.c — 1176%was the smoking gun on the first attempt that had only (a)+(b)); (d) ORT install +enable_dnn=enabledin the coverage job is what makescore/src/dnn/*.cmeasurable in the first place — without ORT, the DNN tree compiles in stub branches and the 85% per-critical-file gate is meaningless; (e)vmaf_use_tiny_modellives indnn_attach_api.cand is added tolibvmaf.soonly viadnn_libvmaf_only_sources— moving it back intodnn_api.creintroduces thevmaf_ctx_dnn_attachundefined-reference link error intest_feature_extractor/test_lpipswheneverenable_dnn=enabled, since those test binaries pull indnn_sourcesforfeature_lpips.cbut never linklibvmaf.c. Lint scope: upstream-mirror Python tests are linted at the same standard as fork-added code; we accept that/sync-upstreamand/port-upstream-commitwill re-trigger Black/isort failures whenever upstream rewrites these files, and the fix is another in-tree reformat pass — never an exclusion. The fork'spyproject.tomland.pre-commit-config.yamlkeeppython/test/resource/(binary fixtures only) excluded;python/test/*.pyis in scope. See ADR-0110 (race fixes, superseded) and ADR-0111 (gcovr + ORT layer). - Re-test:
# Reproduce coverage path locally (requires gcc + python3-pip):
pip install --user 'gcovr>=8.0'
cd libvmaf
meson setup build-cov-test --buildtype=debug -Db_coverage=true \
-Denable_avx512=true -Denable_float=true -Denable_dnn=disabled \
-Dc_args=-fprofile-update=atomic -Dcpp_args=-fprofile-update=atomic
ninja -C build-cov-test
meson test -C build-cov-test --print-errorlogs --num-processes 1
~/.local/bin/gcovr --root .. \
--filter 'src/.*' \
--exclude '.*/test/.*' --exclude '.*/tests/.*' \
--exclude '.*/subprojects/.*' \
--gcov-ignore-parse-errors=negative_hits.warn \
--gcov-ignore-parse-errors=suspicious_hits.warn \
--print-summary --txt build-cov-test/coverage.txt \
--json-summary build-cov-test/coverage.json \
build-cov-test
grep -E 'dnn_api|model_loader' build-cov-test/coverage.txt
# Expected: gcovr completes without "Unexpected negative count" AND no
# per-file percentages exceed 100% (drop --num-processes 1 to reproduce
# the multi-process .gcda merge race; switch back to lcov to reproduce
# the dnn_api.c — 1176% over-count from compilation-unit summation).
# Lint smoke test for upstream-mirror tree:
pre-commit run --files python/test/quality_runner_test.py
# Expected: Black/isort/Ruff all PASS — files are reformatted in-tree
# to fork style and stay clean until the next upstream sync.
0015 — Tox doctest collection skips vmaf/resource/¶
- Workstream PRs: this PR (
fix(ci): skip pytest doctest collection of vmaf/resource/ data files). Surfaced once ADR-0115 consolidated CI triggers tomasterand tox actually started running on PRs. - Touches:
python/tox.ini(single-line--ignore=vmaf/resourceadded to the pytest invocation, plus an explanatory comment block). Pure fork-local; no upstream Python file changes. - Invariant:
pytest --doctest-modulesmust not attempt to import files underpython/vmaf/resource/. Those are parameter / dataset / example-config.pyfiles; several have dots in their stems (e.g.vmaf_v7.2_bootstrap.py) that make them unimportable as Python modules. None carry doctests, so the ignore is correctness rather than a workaround. Do not drop the--ignore=vmaf/resourceflag without first verifying every file under that directory has been renamed to a dot-free stem and is importable. - Re-test:
cd python && tox -e py311 -- --collect-only --doctest-modules \
--ignore=vmaf/resource 2>&1 | grep -c "ERROR collecting vmaf/resource"
# Expected: 0 (was 5 before the fix).
Pure upstream code is not touched, so no Netflix-side conflict vector. Risk is upstream renaming or removing files under python/vmaf/resource/ such that the directory disappears, in which case the --ignore becomes a harmless no-op.
0016 — SYCL -fsycl link-arg gated on icpx CXX¶
- Workstream PRs: this PR (
fix(libvmaf): gate -fsycl link arg on icpx CXX, allow gcc/clang host linker). Surfaced once ADR-0115's CI consolidation added an Ubuntu SYCL job to PR-time CI that usesCXX=g++(host linker) with sidecar icpx for SYCL .cpp compilation. - Touches:
core/src/meson.build(thevmaf_link_argsblock immediately after theis_sycl_enabledflag handling — currently ~lines 696-712). Pure fork-local; no upstream Meson file changes expected. - Invariant:
-fsyclis appended tovmaf_link_argsonly whenmeson.get_compiler('cpp').get_id() == 'intel-llvm'(icpx). Rationale: the documented project mode (see comment nearis_sycl_enabledblock at top ofsrc/meson.build) compiles SYCL.cppfiles viacustom_targetwith icpx, while the project's CXX driver may be gcc / clang / msvc; in that mode the SPIR-V device code is already embedded in the icpx-compiled.ofiles at compile time, and the runtime libraries (libsycl+libsvml+libirc+libze_loader) declared as link dependencies resolve every symbol. Passing-fsyclto a non-icpx linker is a hard error (g++: error: unrecognized command-line option '-fsycl'). Do not remove thecpp.get_id() == 'intel-llvm'guard without first verifying every CI matrix leg uses icpx as the project CXX. - Re-test:
meson setup build -Denable_sycl=true \
-Dcpp_link_args=-Wl,--no-undefined
ninja -C build src/libvmaf.so.3
# Expected: link succeeds; no `-fsycl` errors with gcc/clang host CXX.
Pure fork-local guard; no Netflix-side conflict vector.
0017 — CLI precision default %.6f (Netflix-compat) + frame-skip unref¶
- Workstream PRs: this PR (
fix(cli): revert precision default to %.6f and unref skipped frames). Reverts the default flipped by commitc989fbd9(ADR-0006) per ADR-0119. Companion fix incore/tools/vmaf.cresolves the picture-pool exhaustion in the--frame_skip_ref/distloops surfaced once the always-on picture pool (ADR-0104) made unref'ing skipped pictures mandatory. - Touches:
core/tools/cli_parse.c(VMAF_DEFAULT_PRECISION_FMT+VMAF_LOSSLESS_PRECISION_FMTmacros,resolve_precision_fmt()body,--helptext)core/tools/cli_parse.h(field comments only; struct shape unchanged)core/src/output.c(DEFAULT_SCORE_FORMATmacro)core/tools/vmaf.c(skip loop bodies at thec.frame_skip_ref/c.frame_skip_distfor-loops)python/vmaf/core/result.py(per-frame and aggregate:.6fformatters)python/test/command_line_test.pyis unmodified — Netflix golden assertions stay frozen per CLAUDE.md §8; the binary's output format adapts to them, not the other way around.- Invariant:
vmafCLI default score-output format is%.6f(matches upstream Netflix byte-for-byte).--precision=max|fullselects%.17g(IEEE-754 round-trip lossless).--precision=legacyis a synonym for the default. The library default forvmaf_write_output_with_format(..., score_format=NULL)matches. Skipped frames in the--frame_skip_ref/--frame_skip_distpre-loops arevmaf_picture_unref'd immediately after fetch so the preallocated picture pool is not exhausted before the main scoring loop runs. Do not flip the macros back to%.17gor remove the unrefs without a superseding ADR — both are golden-gate-load-bearing. - Re-test:
ninja -C core/build
python -m pytest python/test/command_line_test.py \
::VmafexecCommandLineTest::test_run_vmafexec \
::VmafexecCommandLineTest::test_run_vmafexec_with_frame_skipping \
::VmafexecCommandLineTest::test_run_vmafexec_with_frame_skipping_unequal \
-v
# Expected: all three PASS in <1 s combined.
Pure fork-local; no Netflix-side conflict vector. If upstream ever changes the default format string, treat their value as the new baseline and reconfirm the golden assertions before adopting.
0018 — FFmpeg patches ship as ordered series.txt¶
- Workstream PRs: this PR (
fix(ci): drop dead sycl trigger + consolidate windows.yml into libvmaf.yml (ADR-0115)). Surfaced once ADR-0115's consolidation routed the docker / FFmpeg-SYCL jobs through the master-targeting CI gate for the first time on this branch — the standalone0003-…sycl…apply broke because it referenced struct fields added by0001-…tiny-model…, the Dockerfile onlyCOPY'd 0003, andffmpeg.ymlreferenced a stale../patches/path. - Touches:
Dockerfile(lines ~86-95 — the FFmpeg patch-apply block),.github/workflows/ffmpeg.yml(theBuild FFmpeg with SYCL patch seriesstep),ffmpeg-patches/000{1,2,3}-*.patch(regenerated via realgit format-patch -3so they carry validindex <sha>..<sha> <mode>lines and committable SHAs). Pure fork-local; no upstream FFmpeg or Netflix file changes. - Invariant: both the Dockerfile and
ffmpeg.ymlwalkffmpeg-patches/series.txtline-by-line and apply each patch viagit applywith apatch -p1fallback. Do not ship a new patch without appending it toseries.txt, and do not reorder existing entries — patch 0003 references LIBVMAFContext fields added by patch 0001, so any out-of-order apply breaks the build at hunk 2 of vf_libvmaf.c. - Two flag-side fixes bundled in the same PR:
--enable-libvmaf-syclis not a valid FFmpeg configure option. Patch 0003 usescheck_pkg_config libvmaf_sycl …auto-detection (matching howlibvmaf_cudais wired) — it never registers the switch. Both Dockerfile and ffmpeg.yml used to pass the flag and configure rejected it withUnknown option "--enable-libvmaf-sycl". SYCL support is now controlled solely by-Denable_sycl=trueat libvmaf build time; FFmpeg picks it up automatically whenlibvmaf-sycl.pcis onPKG_CONFIG_PATH.- The Dockerfile now carries two nvcc-flag ARGs.
NVCC_FLAGS(libvmaf) keeps four-gencodelines plus the experimental--extended-lambda/--expt-relaxed-constexpr/--expt-extended-lambdaflags needed for Thrust/CUB host+device code.FFMPEG_NVCC_FLAGS(FFmpeg) carries a single-gencode arch=compute_75,code=sm_75 -O2— FFmpeg'scheck_nvccrunsnvcc -ptx, which fails withnvcc fatal: Option '--ptx (-ptx)' is not allowed when compiling for multiple GPU architectureson multi-arch input, and--extended-lambdarequires host+device compilation. compute_75 PTX is forward-compatible with all newer GPUs via driver JIT. --enable-libnppis no longer passed to FFmpeg's configure. FFmpeg n8.1's libnpp probe carries an explicitdie "ERROR: libnpp support is deprecated, version 13.0 and up are not supported"(configure:7335-7336) that fires on the base image's CUDA 13.2 libnpp. We don't use scale_npp / transpose_npp / sharpen_npp in any VMAF workflow; cuvid + nvdec + nvenc + libvmaf-cuda is the actual GPU path. Revisit once we move to an FFmpeg release that supports CUDA 13 libnpp upstream.- Patch 0002 (
add-vmaf_pre-filter) gained a missing#include "libavutil/imgutils.h"forav_image_copy_plane(). FFmpeg's libavfilter Makefile builds with-Werror=implicit-function-declarationso this fired during the actual compile (not configure). Caught by a localdocker buildrather than waiting for GitHub Actions — much faster iteration loop. - Re-test:
cd /tmp && rm -rf ffmpeg-test && \
git clone -q --depth 1 -b n8.1 \
https://git.ffmpeg.org/ffmpeg.git ffmpeg-test && \
cd ffmpeg-test && \
while IFS= read -r line; do \
case "$line" in ''|\#*) continue ;; esac; \
git apply "/path/to/vmaf/ffmpeg-patches/$line" \
|| patch -p1 < "/path/to/vmaf/ffmpeg-patches/$line"; \
done < /path/to/vmaf/ffmpeg-patches/series.txt
# Expected: all three patches apply with no rejects; the resulting
# tree compiles with --enable-libvmaf. SYCL is auto-detected via
# check_pkg_config (patch 0003), so no explicit configure flag is
# required when libvmaf-sycl.pc is on PKG_CONFIG_PATH.
Pure fork-local series; no Netflix-side conflict vector. See ADR-0118.
0019 — Coverage Gate annotations: upload-artifact v7 + gcovr filter¶
- Workstream PRs: this PR.
- Touches:
.github/workflows/ci.yml(CPU + GPU coverage steps: gcovr stderr piped throughgrep -vE 'Ignoring (suspicious|negative) hits' ... || true),.github/workflows/{ci,lint,nightly,nightly-bisect,supply-chain,libvmaf}.yml(actions/upload-artifact@v5|@v6 → @v7,actions/download-artifact@v5 → @v7insupply-chain.yml). Note:windows.ymlwas consolidated intolibvmaf.ymlby ADR-0115 / PR #50, so the windows-side bump now lives inlibvmaf.yml'sbuild (MINGW64, …)job. - Invariant: Coverage Gate Annotations panel must finish empty on a clean run. The two pieces are coordinated — (a)
@v7for upload / download artifact actions silences GitHub's Node-20 deprecation banner ahead of the 2026-06-02 forced-Node-24 cutoff; (b) the gcovr stderr filter swallows theIgnoring (suspicious|negative) hitswarnings that gcovr 8 emits for the legitimately-large hit counts in tight ANSNR / VIF / motion inner loops (e.g.ansnr_tools.c:207at ~4.93 G hits across an HD multi-frame coverage suite — real, not gcov bug). The filter is regex-narrow and anchored to gcov's exact warning prefix; any other gcovr warning still surfaces. Upstream (Netflix/vmaf) does not maintain these CI files; rebase impact is limited to the unlikely case that an upstream sync touches the shared.github/workflows/tree, which it currently does not. See ADR-0117. - Re-test:
# Verify gcovr filter locally (after a coverage build per entry 0014):
~/.local/bin/gcovr --root .. \
--filter 'src/.*' \
--exclude '.*/test/.*' --exclude '.*/tests/.*' \
--exclude '.*/subprojects/.*' \
--gcov-ignore-parse-errors=negative_hits.warn \
--gcov-ignore-parse-errors=suspicious_hits.warn \
--print-summary --txt build-cov-test/coverage.txt \
build-cov-test \
2> >(grep -vE 'Ignoring (suspicious|negative) hits' >&2 || true)
# Expected: stderr contains the gcovr summary block but NO
# "Ignoring (suspicious|negative) hits" lines. coverage.txt unchanged.
# Verify all upload/download-artifact instances are on @v7:
grep -rE 'actions/(upload|download)-artifact@v[0-6]' .github/workflows/
# Expected: empty output.
0020 — CI workflow file + display-name renames (Title Case sweep)¶
- Workstream PRs: this PR; renames all six core
.github/workflows/*.ymlfiles to purpose-descriptive kebab-case and normalises every workflowname:and jobname:to Title Case. See ADR-0116. - Touches:
.github/workflows/{ci,lint,security,libvmaf,ffmpeg,docker}.yml(renamed viagit mvtotests-and-quality-gates.yml,lint-and-format.yml,security-scans.yml,libvmaf-build-matrix.yml,ffmpeg-integration.yml,docker-image.yml),README.md(5 badge URLs + labels),docs/principles.md(line 5 workflow-tuple update),.claude/skills/add-gpu-backend/SKILL.md+scaffold.sh(filename refs),docs/adr/0116-*.md(new),docs/adr/README.md(index row),CHANGELOG.md. - Invariant: workflow files are purpose-named; their
name:fields are Title Case sentences with em-dash axis tags; job-levelname:strings are Title Case sentences (Build — / Pre-Commit / Coverage Gate / etc.). Required-status-check contexts inmasterbranch protection are bound to job-level names — when renaming any job, re-pin viagh api --method PUT repos/VMAFx/vmafx/branches/master/protection. The 19 required gates' semantics are unchanged from ADR-0037; only their display strings move. - Re-test:
# Validate every workflow file parses and lists the expected job names.
cd .github/workflows
for f in tests-and-quality-gates.yml lint-and-format.yml security-scans.yml \
libvmaf-build-matrix.yml ffmpeg-integration.yml docker-image.yml; do
yq '.name, .jobs.[].name' "$f" || echo "PARSE FAIL: $f"
done
# Expected: each workflow prints its Title Case workflow name + job names;
# no PARSE FAIL lines.
0021 — DNN-enabled CI matrix legs (gcc + clang + macOS)¶
- Workstream PRs: this PR; adds three new entries to the
libvmaf-buildmatrix in.github/workflows/libvmaf-build-matrix.ymlcovering-Denable_dnn=enabledacross Ubuntu/gcc, Ubuntu/clang, and macOS/clang. See ADR-0120. - Touches:
.github/workflows/libvmaf-build-matrix.yml(3 new matrix entries + ORT install steps + dedicated dnn-suite test step),docs/adr/0120-ai-enabled-ci-matrix-legs.md(new),docs/adr/README.md(index row),CHANGELOG.md(Added entry). - Invariant: the DNN matrix legs install ONNX Runtime via the same pinned source as the dedicated Tiny AI job (tests-and-quality-gates.yml) — Linux: MS tarball at the version pinned by
ORT_VERSION; macOS: Homebrew. When the Tiny AI job's pin changes, the matrix legs'ORT_VERSIONenv in theirInstall ONNX Runtime (linux, DNN leg)step must change to match; otherwise compiler/portability coverage drifts away from the gating leg's actual ABI. - Re-test:
# Local sanity: the matrix file parses and the new job names exist.
yq '.jobs.libvmaf-build.strategy.matrix.include[] | select(.dnn==true) | .name' \
.github/workflows/libvmaf-build-matrix.yml
# Expected output (3 lines):
# Build — Ubuntu gcc (CPU) + DNN
# Build — Ubuntu clang (CPU) + DNN
# Build — macOS clang (CPU) + DNN
# Local DNN build sanity (matches what each leg will run):
meson setup libvmaf core/build --buildtype release \
--prefix $PWD/install -Denable_float=true -Denable_dnn=enabled
ninja -vC core/build install
meson test -C core/build --suite=dnn --print-errorlogs
- Branch protection: the two Linux DNN legs are pinned as required status checks on
masterimmediately after this PR's merge (19 → 21 contexts). The macOS leg stays informational (experimental: true) because Homebrew ORT floats. Re-pin command:
gh api --method PUT repos/VMAFx/vmafx/branches/master/protection \
--input /tmp/protection-update.json
0022 — Windows GPU build-only matrix legs (MSVC + CUDA, MSVC + oneAPI SYCL)¶
- Workstream PRs: this PR; adds a new top-level
windows-gpu-buildjob to.github/workflows/libvmaf-build-matrix.ymlwith two matrix entries (CUDA, SYCL). See ADR-0121. - Touches:
.github/workflows/libvmaf-build-matrix.yml(newwindows-gpu-buildjob),docs/adr/0121-windows-gpu-build-only-legs.md(new),docs/adr/README.md(index row),CHANGELOG.md(Added entry),core/src/compat/win32/pthread.h(new — Win32 pthread shim for MSVC; mirrorscompat/gcc/stdatomic.hpattern),core/src/feature/integer_adm.h(UPSTREAM — converted thedwt_7_9_YCbCr_threshold[3]designated initializer to positional form so MSVC/nvcc-on-Windows accepts the C++ parse; semantically identical, no behavioural change),core/src/ref.handcore/src/feature/feature_extractor.h(UPSTREAM — added#if defined(__cplusplus) && defined(_MSC_VER)branch around#include <stdatomic.h>so MSVC C++ TUs pullatomic_intviausing std::atomic_int;; POSIX paths unchanged),core/src/sycl/d3d11_import.cpp(fix non-existent<libvmaf/log.h>→"log.h"),core/src/sycl/dmabuf_import.cpp(move<unistd.h>inside#if HAVE_SYCL_DMABUFguard for non-VA-API hosts),core/src/sycl/common.cpp(replace POSIXclock_gettime(CLOCK_MONOTONIC)with portablestd::chrono::steady_clock),core/src/feature/x86/motion_avx2.c(UPSTREAM — replace GCC vector-extension__m256i[N]indexing at line 529 with_mm256_extract_epi64; bit-exact),core/src/feature/x86/adm_avx2.c(UPSTREAM — replace 6(__m256i)(_mm256_cmp_ps(...))casts with_mm256_castps_si256(...)and 12__m128i[N]reductions with_mm_extract_epi64; bit-exact),core/src/feature/x86/adm_avx512.c(UPSTREAM — replace 12__m128i[N]reductions with_mm_extract_epi64; bit-exact),core/src/log.c(UPSTREAM — gate<unistd.h>behind!_WIN32, include<io.h>+ redirectisatty/filenoto_isatty/_filenofor MSVC),core/src/feature/integer_vif.c(UPSTREAM — switch thealigned_malloccursor fromvoid *touint8_t *with explicit typed-pointer casts so MSVC accepts the byte-wise pointer arithmetic),core/src/feature/cuda/integer_adm_cuda.c(UPSTREAM — drop unused<unistd.h>include),core/src/dnn/model_loader.c(fork-added — Windows fallback definitions for POSIXS_ISDIR/S_ISREGpath-classification macros),.github/workflows/lint-and-format.yml(fork-added — setlfs: trueon the pre-commit job's checkout so LFS-stored ONNX blobs resolve and don't appear as phantom pre-commit-induced diffs),core/src/feature/x86/motion_avx512.c(UPSTREAM — replace 1__m128i[N]reduction with_mm_extract_epi64; bit-exact),core/src/feature/x86/{vif_statistic_avx2,ansnr_avx2,ansnr_avx512,float_adm_avx2,float_adm_avx512,float_psnr_avx2,float_psnr_avx512,ssim_avx2,ssim_avx512}.c(UPSTREAM — convert 17 sites of trailing__attribute__((aligned(N)))to leading C11_Alignas(N); same alignment, MSVC-portable),core/src/feature/mkdirp.candcore/src/feature/mkdirp.h(UPSTREAM third-party MIT-licensed micro-library — gate<unistd.h>to non-Windows, add<direct.h>+_mkdirfor Windows, addmode_ttypedef for MSVC),core/meson.build(newpthread_dependencygated oncc.check_header('pthread.h')failing),core/src/meson.buildandcore/test/meson.build(threadpthread_dependencyinto every target compiling pthread-using TUs). - Invariant: Windows GPU legs are pinned to the same toolchain versions as the corresponding Linux GPU legs (CUDA 13.0.0, oneAPI BaseKit 2025.3.0.372) so a Linux-vs-Windows divergence implies an MSVC ABI issue, not a tooling-version delta. When either Linux GPU leg bumps its toolchain, the Windows leg must move in lockstep — the Intel installer URL on Windows hard-codes the per-release directory id and the version string, so the bump is two-line edits in the SYCL
Install Intel oneAPI (windows)step (theWINDOWS_BASEKIT_URLenv var). Both legs additionally inject/experimental:c11atomicsintoCFLAGS/CXXFLAGSbecause libvmaf uses C11 atomics that MSVC's<stdatomic.h>rejects without that opt-in flag — when MSVC ships full C11 atomics support, the flag becomes unconditional and can be dropped. Two Windows-only dependency steps round out the parity: the CUDA leg'sJimver/cuda-toolkitsub-package list includes bothcrt(CUDA Runtime Library compile-time headers, shipscrt/host_config.h;cuda_ccclis not a valid Windows sub-package name — installer rejects it) andnvvm(shipsnvvm/bin/cicc.exe+nvvm/libdevice/libdevice.*.bc; without it, nvcc's.cu → PTXstage fails withThe system cannot find the path specified.— on Linux apt pulls NVVM in transitively withcuda-nvcc-XY, Windows requires it explicitly); the SYCL leg builds the Level Zero loader from source (oneapi-src/level-zerov1.18.5 →cmake --build … --target install) because Windows oneAPI BaseKit ships the SYCL runtime but notze_loader.lib, and libvmaf's mesoncc.find_library('ze_loader')needs both the header and the import library. When the Linux aptlevel-zero-devversion moves, bump the L0 git tag to match.core/src/meson.buildguards the explicitsvml/irccc.find_librarycalls behindhost_machine.system() != 'windows'— those calls exist for the gcc/g++ + icpx Linux flow where the host linker is non-Intel; on Windows the host compiler is icx-cl itself and auto-injects the Intel runtime. Round-10 surfaced an additional Windows-only gap: ~14 libvmaf TUs#include <pthread.h>unconditionally, but MSVC and clang-cl ship no pthread (MinGW does, via winpthreads). The fork now ships a header-only Win32 shim atcore/src/compat/win32/pthread.hmapping the in-use pthread subset (mutex / cond / thread create+join+detach) onto SRWLOCK + CONDITION_VARIABLE +_beginthreadex. The shim is wired in viapthread_dependencyincore/meson.build, declared only whencc.check_header('pthread.h')fails — so MinGW and POSIX paths stay untouched. When upstream Netflix/vmaf adds new pthread surface (e.g.,pthread_rwlock_*), extendcompat/win32/pthread.hto cover it. Both nvcc fatbincustom_targets (CUDA) and icpxcustom_targets (SYCLcommon.cpp/picture_sycl.cpp/dmabuf_import.cpp, plus the SYCL feature kernels) bypass meson'sdependencies:plumbing and hand-roll their own-Ilists, so the shim path must be threaded into bothcuda_extra_includesandsycl_inc_flagsexplicitly on Windows. icpx-cl on Windows additionally rejects-fPIC(unsupported option for target 'x86_64-pc-windows-msvc') — sosycl_common_argsandsycl_feature_argsroute their-fPICtoken throughsycl_pic_arg = host_machine.system() != 'windows' ? ['-fPIC'] : []. PIC is the default for Windows DLLs, so dropping the flag is the correct fix rather than a workaround. Round-14 surfaced a third Windows-only blocker:core/src/feature/integer_adm.h(an upstream Netflix file, last touched by upstream port d06dd6cf) initialisesdwt_7_9_YCbCr_threshold[3]with C99 designated initializers ({.a = ..., .k = ..., .f0 = ..., .g = {...}}). The header is included from bothinteger_adm.c(C TU) andcuda/integer_adm/*.cu(C++ TU via nvcc); MSVC's C++ frontend (and nvcc's cudafe++ on Windows) rejects C99 designated initializers without/std:c++20. Converted to positional initialization in the same struct-member order (a / k / f0 / g[4]) — the conversion is provably semantically identical and works in every C/C++ standard, so it costs nothing on the upstream-merge side beyond a trivial conflict marker if upstream Netflix later edits the same lines. Restore designated form post-merge if upstream has it. Round-17 surfaced four more Windows/MSVC-only SYCL blockers, two of which touch upstream-shared headers. (a)core/src/ref.handcore/src/feature/feature_extractor.h(UPSTREAM) unconditionally#include <stdatomic.h>and use theatomic_inttypedef in struct definitions. MSVC's<stdatomic.h>(added in 19.34) only declares the C11 symbols inside the global namespace under C; in C++ compilation (icpx-cl drives the SYCL TUs as C++) MSVC surfaces them only insidenamespace std::. gcc/clang expose both via a GNU extension, so the upstream code works on every other platform. The fork now wraps both headers'#include <stdatomic.h>in#if defined(__cplusplus) && defined(_MSC_VER)→#include <atomic>+using std::atomic_int;, falling through to the original<stdatomic.h>line on every other configuration. ABI is unchanged —atomic_intresolves to the same underlying type. If upstream Netflix adds further C11 atomic typedefs in these headers (e.g.,atomic_uint,atomic_size_t), extend theusing std::lines to cover them. (b)core/src/sycl/d3d11_import.cpp(fork-added) used<libvmaf/log.h>which doesn't exist —log.hlives atcore/src/log.hand is internal. Switched to"log.h"; the icpx invocation already supplies the src-relative-I. (c)core/src/sycl/dmabuf_import.cpp(fork-added) included<unistd.h>at file scope, but POSIXclose()is only used inside the#if HAVE_SYCL_DMABUFVA-API block. Moved the<unistd.h>include inside that guard so non-DMA-BUF builds (Windows MSVC, macOS) compile cleanly. (d)core/src/sycl/common.cpp(fork-added) calledclock_gettime(CLOCK_MONOTONIC), which doesn't exist on Windows. Replaced withstd::chrono::steady_clock(guaranteed monotonic by the C++ standard, portable on every supported host). All four fixes preserve POSIX/Linux behaviour bit-identically and only change the Windows MSVC build path. Round-18 surfaced a fifth Windows blocker on the CUDA leg's CPU SIMD compile path:core/src/feature/x86/motion_avx2.c:529(UPSTREAM, ported in commit 9371a0aa from Netflix PR #1486) computedfinal_accum[0] + final_accum[1] + final_accum[2] + final_accum[3]to extract the four int64 lanes from an__m256i. gcc/clang allow this via the GNU vector-extension treatment of__m256i(it carries__attribute__((vector_size(32)))); MSVC rejects it withC2088: built-in operator '[' cannot be applied to an operand of type '__m256i'. Replaced with_mm256_extract_epi64(final_accum, N)for N ∈ {0..3}, summed — bit-exact lane sum on every compiler. Restore the index form post-merge if upstream Netflix later edits the same lines and your toolchain matrix doesn't include MSVC. Round-19 surfaced the same MSVC pattern at 19 more call sites across the AVX2/AVX-512 ADM and motion files plus six GCC-style vector casts.core/src/feature/x86/adm_avx2.c(UPSTREAM): 6 lines (915-920) used(__m256i)(_mm256_cmp_ps(...))C-style casts that gcc/clang accept via the GNU vector extension; replaced with the dedicated_mm256_castps_si256(...)bit-cast intrinsic. 12 lane-extract sites (r2_h[0]+r2_h[1], etc. at lines 2420 / 2425 / 2430 / 2893 / 2897 / 2901 / 4079 / 4084 / 4089 / 4627 / 4631 / 4635) replaced with_mm_extract_epi64(r2_X, N)summed pair.core/src/feature/x86/adm_avx512.c(UPSTREAM): 6 sister lane-extract sites (lines 4470 / 4477 / 4484 / 4625 / 4631 / 4637) — same fix. The AVX-512 paths reduce a__m512idown to__m128ifirst (via_mm512_extracti64x4_epi64→_mm256_extracti64x2_epi64) before the index, so only the final__m128i[N]step needed changing.core/src/feature/x86/motion_avx512.c(UPSTREAM, ported in 9371a0aa from PR #1486): one finalr2[0]+r2[1]reduction (line 448), same fix. All 19 lane-extract fixes plus the 6 cast fixes are bit-exact rewrites and only change the source-level syntax to MSVC-portable form. Restore the original forms post-merge if upstream Netflix later edits the same lines and your toolchain matrix doesn't include MSVC. Additionallycore/src/sycl/d3d11_import.cpp(fork-added) switched from C-style COBJMACROS helpers (ID3D11Device_CreateTexture2D,…_Release, etc.) to C++ method-call syntax (device->CreateTexture2D,tex->Release) — d3d11.h gates COBJMACROS behind!defined(__cplusplus), so the C-style helpers aren't visible in this.cppTU. The two forms are ABI-equivalent (both dispatch through the COM vtable); the choice is purely lexical and POSIX builds aren't affected (the whole TU is#ifdef _WIN32). Round-20 surfaced two more Windows-only blockers. (a) 17 sites across the x86 SIMD layer used GCC'sfloat tmp[N] __attribute__((aligned(M)));form to align scratch buffers for_mm{256,512}_store_ps. MSVC rejects the trailing-attribute syntax withC2146: syntax error: missing ';' before identifier '__attribute__'. Replaced with the C11-standard_Alignas(M) float tmp[N];(alignment specifier before the type) — works in gcc, clang and MSVC with/std:c11. Files touched (all UPSTREAM):vif_statistic_avx2.c(×2),ansnr_avx2.c(×2),ansnr_avx512.c(×2),float_adm_avx2.c(×2),float_adm_avx512.c(×2),float_psnr_avx2.c(×1),float_psnr_avx512.c(×1),ssim_avx2.c(×4),ssim_avx512.c(×4). The pre-existingvif_avx2.c/vif_avx512.calready define a portableALIGNED(x)macro at file scope and position the attribute before the type, so they compile cleanly under MSVC and were not touched. (b)core/src/feature/mkdirp.c(UPSTREAM, third-party MIT-licensed copy of Stephen Mathieson's micro-library) included<unistd.h>unconditionally but never used POSIXunistdsymbols (onlymkdirvia<sys/stat.h>/<direct.h>). Gated<unistd.h>to non-Windows and added<direct.h>for Windows; switchedmkdir(pathname)→_mkdir(pathname)(the non-deprecated MSVC name).core/src/feature/mkdirp.hadded amode_ttypedef under MSVC since neither<sys/types.h>nor<sys/stat.h>declare it on Windows;modeis ignored on the Windows path anyway. Round-21 surfaced two more blockers (the round-19__m128i[N]sweep missed six sites) plus a pre-commit workflow checkout gap. (a)core/src/feature/x86/adm_avx512.c(UPSTREAM) had six furtherr2_X[0] + r2_X[1]reductions at lines 2128 / 2135 / 2142 / 2589 / 2595 / 2601 that reduce a__m512iaccumulator down to__m128ibefore the lane index. Replaced with the same_mm_extract_epi64(r2_X, N)summed-pair pattern used in round 19 — bit-exact, MSVC-portable. (b)core/src/log.c(UPSTREAM) included<unistd.h>unconditionally to pick up POSIXisatty/fileno. On MSVC both live in<io.h>as_isatty/_fileno; gated the include and macro-redirected the names so the one call site at line 34 compiles on both sides without touching the POSIX path. (c).github/workflows/lint-and-format.yml(fork-added) checks out withoutlfs: true, so themodel/tiny/*.onnxfiles land as LFS pointer stubs. pre-commit's "changes made by hooks" reporter then diffs the stubs against HEAD's real blobs and fails the job even though no hook touched them. Addedlfs: trueto the pre-commit job's checkout. (d)core/src/meson.build—cuda_common_vmaf_libstatic library had nodependencies:list, so the Win32 pthread shim (wired in viapthread_dependencyin core/meson.build) wasn't on its include path;cuda/common.hunconditionally#include <pthread.h>and MSVC failed with C1083. Addeddependencies : [pthread_dependency]— no-op on POSIX (empty list), routes the shim path in on Windows. (e)core/src/feature/integer_vif.c(UPSTREAM) walked one bigaligned_mallocresult asvoid *dataand diddata += pad_size/data += h * stride_16etc. to carve the buffer into typed sub-pointers. gcc/clang accept pointer arithmetic onvoid *as a GNU extension (treatingsizeof(void) == 1); MSVC rejects it withC2036: 'void *': unknown size. Replaced the cursor type withuint8_t *and added explicit casts at assignment sites that take a typed pointer (uint16_t *mu1,uint32_t *mu1_32, etc.). Byte offsets are identical, layout unchanged, bit-exact. If upstream Netflix edits the same loop, reabsorb the walk and re-apply the cursor-type + cast pattern. (f)core/src/feature/cuda/integer_adm_cuda.c(UPSTREAM) included<unistd.h>at line 33 but used no POSIX symbols from it; MSVC failed with C1083. Dropped the unused include outright — simplest fix, no runtime change on any platform. (g)core/src/dnn/model_loader.c(fork-added) usesS_ISDIR/S_ISREGto classify resolved paths. MSVC ships the underlyingS_IFMT/S_IFDIR/S_IFREGbit masks in<sys/stat.h>but not the POSIX classification macros. Added a Windows-only fallback (#ifndef S_ISDIR #define S_ISDIR(m) (((m) & S_IFMT) == S_IFDIR) #endif, same for S_ISREG) guarded by#ifdef _WIN32. Semantically identical to the POSIX macro on Linux/macOS. Round-21e surfaced the final source-portability blockers once the DLL build passed preprocessing. (h)core/src/predict.c,core/src/libvmaf.candcore/src/read_json_model.c(all UPSTREAM) used C99 variable-length arrays —double scores[cnt]at predict.c:385,char name[name_sz]at predict.c:453 and libvmaf.c:1741, pluscfg_name[cfg_name_sz]andgenerated_key[generated_key_sz]in the.jsonmodel-collection parser. gcc/clang accept VLAs as a C11 optional feature; MSVC (even with/std:c11) rejects them outright withC2057: expected constant expression(plus C2466 and C2133 on theconst size_tsized arrays — MSVC treatsconstas runtime-bounded, not a constant expression, even when the initialiser is literal like4 + 1). Replaced each runtime-sized buffer with a smallmalloc+ explicitfreeon every exit path (in predict.c and read_json_model.c agoto out;cleanup arm was introduced because the loops error-exit mid-function). Thegenerated_keybuffer in read_json_model.c uses the narrower fix —char generated_key[5];— since its size (four decimal digits of the bootstrap sub-model index plus NUL) is a true compile-time constant. Buffers are a handful of bytes each (name_szis the model-collection name length plus the fixed_ci_p95_losuffix,scoresholds ~20 doubles,cfg_nameis the name plus_0000suffix), so the heap round-trip is not performance-relevant; the new-ENOMEMfailure mode is handled uniformly by existing callers. The read_json_model.c refactor also plugs a pre-existing leak of thenamebuffer on the earlyreturn -EINVALwhen a JSON object key isn't a string — thegoto out;path freesname+cfg_nameon every exit.core/test/test_feature_extractor.c:56(UPSTREAM) declaredconst unsigned n_threads = 8;and used it as the extent ofVmafFeatureExtractorContext *fex_ctx[n_threads];. Converted toenum { n_threads = 8 };so MSVC sees a constant-expression; every other compiler accepts enum constants identically. Re-absorb if upstream Netflix later edits the same loops and your toolchain matrix omits MSVC. (i) The Windows MSVC build-only legs now build the full tree — CLI tools, unit tests and libvmaf.dll — rather than the previous short cut of disabling-Denable_tools/-Denable_tests. Per user direction ("fix the code ffs"), the tree polyfills the remaining POSIX surfaces on MSVC instead: (core/tools/compat/win32/getopt.h+core/tools/compat/win32/getopt.c) a from-scratch POSIX/GNU-compatiblegetopt_longshim (short / long options,no_argument/required_argument/optional_argument, argv permutation for non-option operands,--explicit stop,=-embedded values). The shim is fork-added (BSD-3-Clause-Plus-Patent, Copyright 2026 Lusoris and Claude) and declared via a singlegetopt_dependencyincore/meson.build, gated oncc.check_header('getopt.h')failing. The dependency auto-propagates the shim.cinto any consuming target via meson'ssources:keyword, so both thevmafCLI (core/tools/meson.build) and thetest_cli_parseunit test (core/test/meson.build) pick it up uniformly. MinGW ships<getopt.h>via mingw-w64-crt, socheck_headersucceeds there and the shim stays out of the TU list. (j) Eleven test executables (test_log,test_dict,test_opt,test_cpu,test_ref,test_feature,test_ciede,test_luminance_tools,test_cli_parse,test_sycl,test_sycl_pic_preallocation) were missingpthread_dependencyin theirdependencies:lists atcore/test/meson.build. On POSIXpthread_dependencyis an empty list so the omission was invisible; on MSVC those TUs transitively includefeature_collector.h→<pthread.h>and fail with C1083. Threaded the dependency through all eleven targets.test_cli_parseadditionally listsgetopt_dependencyto pick up the shim. (k) Three additional VLA sites surfaced once the test harness built on MSVC:test_cambi.c:254hadunsigned w = 5, h = 5; uint16_t buffer[3 * w];; converted toenum { w = 5, h = 5 };so the array extent is a constant expression.test_pic_preallocation.c:382andtest_pic_preallocation.c:506hadconst int num_threads = N; pthread_t threads[num_threads];— MSVC rejectsconst intas non-constant-expression. Converted toenum { num_threads = N, fetches_per_thread = M };. (l)test_ring_buffer.c:23andtest_pic_preallocation.c:26included<unistd.h>forusleep/sleep. Gated behind!_WIN32with a Win32 fallback via<windows.h>+#define usleep(us) Sleep(((us) + 999) / 1000)/#define sleep(s) Sleep((s) * 1000). The conversion rounds sub-millisecondusleepinputs up, which is safe for these test paths (they use 100 µs jitter and 1 s waits). (m)core/tools/vmaf.cincluded<unistd.h>forisatty/fileno. Applied the same gating pattern used inlog.cin round-21(b) — include<io.h>on MSVC and redirectisatty/filenoto_isatty/_filenovia#define. (n)__builtin_clz/__builtin_clzllare GCC intrinsics; MSVC ships__lzcnt/__lzcnt64via<intrin.h>instead. The shim already lived incore/src/feature/integer_vif.hbutinteger_adm.c:939,x86/adm_avx2.c:1425andx86/adm_avx512.c:1217don't include that header. Extracted the shim into a dedicatedcore/src/feature/compat_builtin.h(fork-added) and included it from all four TUs. The guard isdefined(_MSC_VER) && !defined(__clang__), so clang-cl / icx-cl (which provide the GCC intrinsics natively) skip the shim. (o) The SYCL leg's D3D11 import TUcore/src/sycl/d3d11_import.cppis C++ (icpx-cl drives it as C++ on Windows) but included the internal C headerlog.hwithout anextern "C"wrap.log.his an upstream Netflix header with no__cplusplusguard, sovmaf_loggot C++ name-mangled in the .cpp TU and failed to resolve against the C-linkage symbol produced bylog.cat link time (LNK2019from every test target that pulls in the SYCL static lib). Wrapped the#include "log.h"withextern "C" { ... }inside the fork-added .cpp rather than touching the upstream header — keepslog.hidentical to upstream on every/sync-upstream. (p) The Windows MSVC legs build with--default-library=static. libvmaf's public API has no__declspec(dllexport)attributes (upstream Netflix is POSIX-shaped), so a vanilla MSVC shared build producessrc/vmaf-3.dllwith no exported symbols and the toolchain therefore never emits the companionvmaf.libimport library. Downstream tool targets then fail withLNK1181: cannot open input file 'src\vmaf.lib'. The MinGW matrix leg has used--default-library staticsince day one for the same reason (line 387); the MSVC legs now mirror that choice viamatrix.include[].meson_extra. Downstream consumers that want a DLL can either add__declspec(dllexport)decorations to the public API or use a.deffile; that is a separate decision and out of scope for the build-only gate. - Re-test:
# Local sanity: the matrix file parses and the new job names exist.
yq '.jobs.windows-gpu-build.strategy.matrix.include[].name' \
.github/workflows/libvmaf-build-matrix.yml
# Expected output (2 lines):
# Build — Windows MSVC + CUDA (build only)
# Build — Windows MSVC + oneAPI SYCL (build only)
- Branch protection: the two Windows GPU legs are pinned as required status checks on
masterimmediately after this PR's merge. After ADR-0120's two Linux DNN legs the count moves 21 → 23. Re-pin via:
gh api --method PUT repos/VMAFx/vmafx/branches/master/protection \
--input /tmp/protection-update.json
0023 — CUDA gencode coverage (sm_86/sm_89/compute_80 PTX) + init hardening¶
- Workstream PRs: the ADR-0122 PR (gencode + init hardening) and the ADR-0123 follow-up for the
32b115dfpost-cubin-load regression. - Touches:
core/src/meson.build— thegencodearray in theif get_option('enable_nvcc')branch.core/src/cuda/common.c—vmaf_cuda_state_init()error paths (multi-line actionable log,cuda_free_functions()+free(c)+*cu_state = NULLcleanup).docs/backends/cuda/overview.md—## Runtime requirementssection and### GPU architecture coveragetable.- Invariant: the
gencodearray unconditionally emits cubins forsm_75/sm_80/sm_86/sm_89plus acompute_80PTX, independent of hostnvccversion. Upstream Netflix's gencode only ships cubins at Txx major boundaries (sm_75/sm_80/sm_90/sm_100/sm_120); a literal merge that replaces our array with upstream's would re-open the Ampere-sm_86/ Ada-sm_89coverage hole. Thesm_90/sm_100/sm_120entries are still version-gated and should be preserved verbatim if upstream adds new gates. The init-path error messages are fork-local strings; upstream's terse"Error: failed to load CUDA functions"must NOT win a merge. - Re-test:
meson setup build -Denable_cuda=true -Denable_nvcc=true
ninja -C build 2>&1 | grep -E 'compute_(80|86|89)'
# Expect at least -gencode=arch=compute_86,code=sm_86 and
# -gencode=arch=compute_89,code=sm_89 and
# -gencode=arch=compute_80,code=compute_80
# Actionable init message (run without CUDA driver on the loader path):
LD_LIBRARY_PATH= ./build/tools/vmaf --help 2>&1 | grep -qi 'libcuda.so.1' || \
echo "init log regressed"
0024 — vmaf_read_pictures null-guard for CUDA device-only path¶
- Workstream PRs: the ADR-0123 follow-up landed atop the ADR-0122 gencode/init-hardening work.
- Touches:
core/src/libvmaf.c— the non-threaded tail ofvmaf_read_picturesat theprev_refupdate site (line ~1428 in the fork; upstream equivalent is the tail added byf740276a).- Invariant: the
prev_refupdate is guarded byif (ref && ref->ref)so pure-CUDA extractor sets (whereref = &ref_hostbutref_hostwas never populated bytranslate_picture_device) do not deref a NULL refcount. Upstream currently has the same unguarded tail; the bug is masked upstream only because the experimentalVMAF_PICTURE_POOLgate from32b115dfis still in place. A literal upstream merge that removes our null-guard while upstream's experimental gate is still holding would pass tests but re-open thelibvmaf_cudaffmpeg crash the moment the gate flips default-on (which the fork did in65460e3a, ADR-0104). Keep the guard until the upstream null-guard port lands. - Re-test:
# Unit tests cover the non-regression on the library side:
meson test -C build
# End-to-end regression: ffmpeg libvmaf_cuda must exit 0 on a
# CUDA-device-only extractor set (full recipe in ADR-0123).
./ffmpeg -init_hw_device cuda=cu:0 -filter_hw_device cu \
-i /tmp/ref.mp4 -i /tmp/dis.mp4 \
-lavfi "[0:v]format=yuv420p,hwupload_cuda[r];\
[1:v]format=yuv420p,hwupload_cuda[d];\
[r][d]libvmaf_cuda=log_path=/tmp/out.json:log_fmt=json" \
-f null -
0025 — VIF init() fail-path frees advanced byte-cursor¶
- Workstream PRs: PR #47 (rewritten to leak-fix-only after master absorbed the void→uint8_t half via commit
b0a4ac3a, entry 0022 §e). Ports the leak-fix half of upstream Netflix PR #1476. - Touches:
core/src/feature/integer_vif.c(UPSTREAM — 2-line fix in theinit()fail:handler). - Invariant:
init()walksuint8_t *dataforward throughaligned_malloc's one allocation, advancing past each sub-pointer assignment. Ifvmaf_feature_name_dict_from_provided_featuresreturns NULL the fail path must free the base pointers->public.buf.data, never the advanced cursordata. Upstream master still hasaligned_free(data)there — same bug — so this entry is the reminder to not let an upstream sync re-introduce the advanced-cursor form. If upstream lands PR #1476 or an equivalent, the sync can drop this entry. - Re-test:
meson test -C build --suite=fast
# Static check: ripgrep the pattern that must NOT return.
rg -n "aligned_free\(data\)" core/src/feature/integer_vif.c && \
echo 'REGRESSED' || echo 'ok'
0026 — Automated rule-enforcement workflow + copyright pre-commit hook¶
- Workstream PRs: this PR (ADR-0124 adoption). Closes the "rule-without-a-check" gap on ADR-0100 / 0105 / 0106 / 0108.
- Touches (all FORK-ADDED — no upstream overlap):
.github/workflows/rule-enforcement.yml(new),scripts/ci/check-copyright.sh(new),.pre-commit-config.yaml(appended local hook). - Invariant: the
deep-dive-checklistjob is blocking on every PR that is not an upstream port (exempt viaport:title prefix orport/branch). The other three gates (doc-substance-check,adr-backfill-check, copyright pre-commit) are advisory or pre-commit, never CI-blocking; this split is the whole point of ADR-0124 and an upstream sync must not move them into the required-status-check set without a follow-up ADR. The opt-out parser matches/^-?\s*no .* (?:needed|impact|rebase-sensitive)/per ADR-0108 §Opt-out-lines — if upstream ever changes PR-template phrasing (unlikely; this is fork-local), the regex and the template must move together. - Re-test:
# Lint the workflow + hook locally.
pre-commit run --files \
.github/workflows/rule-enforcement.yml \
scripts/ci/check-copyright.sh \
.pre-commit-config.yaml
# Dry-run the copyright hook against a staged source file.
scripts/ci/check-copyright.sh core/src/libvmaf.c && echo ok
# Synthetic PR body that violates ADR-0108 should fail the parser;
# see docs/research/0002-automated-rule-enforcement.md §Verification
# plan for the three test cases.
0027 — SSIMULACRA 2 scalar extractor (libjxl FastGaussian IIR blur)¶
- Workstream PRs: this PR (
feat/ssimulacra2-scalar); proposal ADR in PR #67. - Touches:
core/src/feature/ssimulacra2.c(fork-local, new),core/src/meson.build,core/src/feature/feature_extractor.c. - Invariant: the extractor embeds several tables that must track libjxl upstream — opsin absorbance matrix,
MakePositiveXYBoffsets, 108 pooling weights, polynomial-transform coefficients, and the FastGaussian coefficient-derivation formulas (radius =3.2795·σ + 0.2546, Cramer's 3×3 solve for β, n2/d1 assignment per Charalampidis 2016 (33)). If libjxl ever changes any of these, updatessimulacra2.cin the same PR that syncs upstream. Self-consistency must stay at exactly100.000000for identical ref/dist inputs — this is the cheapest regression check. - Re-test:
meson test -C build --suite=fast
./build/tools/vmaf \
--reference python/test/resource/yuv/src01_hrc00_576x324.yuv \
--distorted python/test/resource/yuv/src01_hrc00_576x324.yuv \
-w 576 -h 324 -p 420 -b 8 --feature ssimulacra2 -o /tmp/self.xml \
&& grep -q 'ssimulacra2="100.000000"' /tmp/self.xml \
&& echo "ok: self-consistency 100.0"
0028 — MS-SSIM separable decimate + AVX2/AVX-512/NEON SIMD¶
- Workstream PRs:
feat/ms-ssim-decimate-simd-v2(supersedes the rebase-incompatiblefeat/ms-ssim-decimate-simd; AVX2/AVX-512, commits7de8cd7fscalar separable,5f93c864AVX2,73436438AVX-512);feat/ms-ssim-decimate-neon-v2(NEON follow-up, stacked). - Touches:
core/src/feature/ms_ssim_decimate.{c,h}(NEW),core/src/feature/x86/ms_ssim_decimate_avx2.{c,h}(NEW),core/src/feature/x86/ms_ssim_decimate_avx512.{c,h}(NEW),core/src/feature/arm64/ms_ssim_decimate_neon.{c,h}(NEW),core/src/feature/ms_ssim.c(call-site change),core/src/meson.build(register new SIMD TUs),core/test/test_ms_ssim_decimate.c(NEW),core/test/meson.build(arm64 gating). - Invariant: the 9-tap 9/7 biorthogonal wavelet LPF coefficients (
ms_ssim_lpf_h/ms_ssim_lpf_v) are duplicated verbatim in five TUs for bit-identity: the scalarms_ssim_decimate.c, the AVX2 variant, the AVX-512 variant, the NEON variant, and upstream'sg_lpf_h/g_lpf_vinms_ssim.c. Any upstream change to the coefficient values or theKBND_SYMMETRICmirror branch iniqa/convolve.cmust be mirrored to all five. If not mirrored, SIMD paths and scalar diverge silently and the bit-equalitymemcmpintest_ms_ssim_decimatecatches it — but only when that test runs, so diff the five files first. - Re-test (on each supported host arch):
# x86_64 host — native build.
meson test -C build
./build/test/test_ms_ssim_decimate
# aarch64 host OR aarch64 cross under qemu — see /tmp/aarch64-cross.txt.
meson setup build-arm64 libvmaf --cross-file /tmp/aarch64-cross.txt \
-Denable_cuda=false -Denable_sycl=false
ninja -C build-arm64
qemu-aarch64-static -L /usr/aarch64-linux-gnu \
build-arm64/test/test_ms_ssim_decimate
# Netflix MS-SSIM golden — places=4 must still pass through SIMD.
.venv/bin/python -m pytest \
python/test/feature_extractor_test.py::FeatureExtractorTest::test_run_ms_ssim_fextractor
0029 — KBND_SYMMETRIC period-based reflection in iqa/convolve.c¶
- Workstream PRs:
feat/ms-ssim-decimate-simd-v2follow-up (CI triage on PR #69, 2026-04-20). - Touches:
core/src/feature/iqa/convolve.c(upstream file, rewrittenKBND_SYMMETRIC). - Invariant:
KBND_SYMMETRIC(img, w, h, x, y, _)must use the period-based form (period = 2*w,period = 2*h) so that offsets with|x| > wor|y| > hstill land in bounds. Upstream's single-reflect form was out-of-bounds wheneverw < kernel_halforh < kernel_half; the latent bug did not reproduce in Netflix golden tests because MS-SSIM pyramids never decimate below ~60×34. Any upstream change that reverts to the single-reflect form must be rejected or re-ported. - Re-test:
./build/test/test_ms_ssim_decimate # test_1x1 border case
.venv/bin/python -m pytest \
python/test/feature_extractor_test.py::FeatureExtractorTest::test_run_ms_ssim_fextractor
0030 — adm_decouple_s123_avx512 stack-array 64-byte alignment¶
- Workstream PRs:
feat/ms-ssim-decimate-simd-v2follow-up (CI triage on PR #69, 2026-04-20). - Touches:
core/src/feature/x86/adm_avx512.c(upstream file, one-line_Alignas(64)onint64_t angle_flag[16]at line 1317).core/test/test_pic_preallocation.c(upstream file, threevmaf_model_destroy(model)calls pairing thevmaf_model_loadintest_picture_pool_basic/_small/_yuv444). - Invariant: the stack slot for
angle_flagmust be 64-byte aligned because two_mm512_loadu_si512(&angle_flag[0/8])loads in the same scope may be promoted to alignedvmovdqa64by LTO. Dropping the_Alignas(64)annotation re-introduces the SEGV under--buildtype=release -Db_lto=true -Db_sanitize=address. Debug / no-LTO builds keepvmovdqu64and cannot flag the regression. Seedocs/development/known-upstream-bugs.md. - Re-test:
meson setup build-asan-lto libvmaf \
-Denable_cuda=false -Denable_sycl=false \
-Db_sanitize=address --buildtype=release -Db_lto=true
ninja -C build-asan-lto test/test_pic_preallocation
ASAN_OPTIONS=detect_leaks=1 \
./build-asan-lto/test/test_pic_preallocation
0031 — Batch-A upstream-port small-fix sweep (ports of unmerged PRs)¶
- Workstream PRs:
feat/batch-a-upstream-small-fix-sweep— commits546a40ee(T0-1),8fed8ad1(T4-4),83a1db46(T4-5),34425dee(T4-6). ADRs 0131, 0132, 0134, 0135. - Touches:
core/src/cuda/picture_cuda.c(one-linecuMemFreeport of Netflix#1382)core/src/feature/feature_collector.c+core/test/test_feature_collector.c(mount/unmount bugfix port of Netflix#1406 + shared-helper test refactor)core/src/meson.build(declare_dependency+override_dependencyport of Netflix#1451)core/include/libvmaf/model.h,core/src/model.c,core/test/test_model.c,docs/api/index.md(built-in model iterator port of Netflix#1424)- Invariant: each of the four upstream PRs is OPEN (unmerged) on the port date; when Netflix merges any of them, the fork's version is correction-bearing (T4-4 test refactor, T4-6 three defect fixes + Doxygen doc expansion), not line-identical. Resolution on upstream merge is always "keep fork version" because the fork's version already satisfies the PR's intent and additionally fixes the defects.
- Netflix#1406 conflict will land in
test_feature_collector.c— fork usesload_three_test_models()helper vs upstream's inline per-modelVmafModel *m0, *m1, *m2;duplication. - Netflix#1424 conflict will land in
core/src/model.candcore/test/test_model.c— fork useselse ifguard +idx + 1 < CNT+ const-qualified test types. - Netflix#1382 and Netflix#1451 are line-identical in substance; merge should be clean aside from trailing-comma style drift.
- Re-test:
meson setup build libvmaf -Denable_cuda=false -Denable_sycl=false
ninja -C build test/test_feature_collector test/test_model
build/test/test_feature_collector
build/test/test_model
# Expected: 6/6 pass in test_feature_collector (mount/unmount
# 3-model sequences); 39/39 pass in test_model (includes
# test_version_next full-iteration invariant).
0032 — Thread-local locale handling for numeric I/O (port of Netflix/vmaf#1430)¶
- Workstream PRs:
port/netflix-1430-thread-locale(T4-3 from the "Batch-A follow-up" sweep, 2026-04-20). - Touches:
core/src/thread_locale.h/core/src/thread_locale.c(new, upstream-authored);core/src/meson.build(twocdata.set('HAVE_USELOCALE'/'HAVE_XLOCALE_H')probes +src_dir + 'thread_locale.c'inlibvmaf_sources);core/src/output.c(four writers gainpush_c()+pop()bracket, preserving fork'sferror(outfile) ? -EIO : 0return contract from ADR-0119);core/src/svm.cpp(drop<locale.h>include; replacesetlocale/strdup/setlocalebracket withvmaf_thread_locale_push_c/pop; addbuffer.imbue(std::locale::classic())to both SVM parser ctors with fork's K&R + 4-space style);core/src/read_json_model.c(bracketmodel_parsewith push/pop);core/test/meson.build(newtest_locale_handlingtarget + test registration);core/test/test_locale_handling.c(new, upstream-authored with three fork corrections for thescore_formatparameter). - Invariant: fork's output writers return
ferror(outfile) ? -EIO : 0— this must survive any upstream refactor of the writer bodies. Thepush_c()call MUST be paired with apop()on every return path (writer bodies have a single tail return, so the pattern is locallypush → body → pop → return ferror-check). Droppingpop()leaks alocale_ton POSIX and leaves the thread locked to "C" on Windows. - Re-test:
meson setup build -Denable_cuda=false -Denable_sycl=false
ninja -C build
meson test -C build test_locale_handling
# Repro the user-visible failure without the fix:
LC_ALL=de_DE.UTF-8 build/tools/vmaf --reference ref.yuv \
--distorted dis.yuv --width 1920 --height 1080 \
--pixel_format 420 --bitdepth 8 --output result.json \
--json
# Assert output contains period decimals, not comma.
python -c "import json; d=json.load(open('result.json')); \
assert all('.' in repr(v) for v in \
[f['metrics']['vmaf'] for f in d['frames']])"
- On upstream sync: when Netflix merges PR #1430, the
(cherry picked from commit 054a97ed…)trailer ingit log port/netflix-1430-thread-localelets the next/sync-upstreamskip this commit. If the upstream diff drifts, redo the three fork corrections listed in ADR-0137 §Decision.
0033 — SSIM / MS-SSIM SIMD bit-exact to scalar via per-lane scalar double¶
- Workstream PRs:
feat/ms-ssim-decimate-neon(this PR — companion to the ADR-0138 convolve fast path). - Touches:
core/src/feature/x86/ssim_avx2.candcore/src/feature/x86/ssim_avx512.c—ssim_accumulate_*rewritten.ssim_precompute_*andssim_variance_*unchanged (they were already bit-exact). Plus the new bit-exactconvolve_avx2.c/convolve_avx512.cand the upstream h-pass OOB fix atiqa/convolve.c:159. - Invariants (see ADR-0139 §Decision):
- Convolve taps — single-rounded
float*float→ widen →doubleadd, NO FMA. Mirrors scalarsum += img[i]*k[j]iniqa/convolve.c. - SSIM accumulate — scalar's
2.0 *literal (2.0 * ref_mu[i] * cmp_mu[i] + C1and2.0 * srsc + C2) is a Cdoubleliteral. Both SIMD accumulators do the2.0 *numerator + division + finall*c*sproduct per-lane in scalar double to match scalar type promotions byte-for-byte. - H-pass outer-loop bound —
y < dst_h + vc - kh_even(noty < dst_h + vc); the- kh_evenis load-bearing because the last cache row on even-tap kernels (e.g. box-8) is never read by the v-pass but was previously written OOB when image height equals kernel height.
Fork-local SSIM SIMD is NOT upstream. If upstream ever adds their own SSIM AVX2/AVX-512, keep the fork's version on conflict — it's the only variant verified bit-exact to scalar at --precision max. - Re-test:
meson setup build -Denable_cuda=false -Denable_sycl=false
ninja -C build
meson test -C build test_iqa_convolve test_ms_ssim_decimate
# Bit-exactness check across dispatch backends:
FIX=python/test/resource/yuv/checkerboard_1920_1080_10_3_0_0.yuv
DIS=python/test/resource/yuv/checkerboard_1920_1080_10_3_1_0.yuv
for m in 255 16 0; do
build/tools/vmaf --cpumask $m --reference $FIX --distorted $DIS \
--width 1920 --height 1080 --pixel_format 420 --bitdepth 8 \
--feature float_ssim --feature float_ms_ssim \
--output /tmp/ssim_$m.xml --precision max
done
diff <(grep -v '<fyi fps' /tmp/ssim_255.xml) \
<(grep -v '<fyi fps' /tmp/ssim_16.xml) # expect empty
diff <(grep -v '<fyi fps' /tmp/ssim_255.xml) \
<(grep -v '<fyi fps' /tmp/ssim_0.xml) # expect empty
- On upstream sync: the AVX2/AVX-512 SSIM surface is entirely fork-local (upstream has VIF/ADM/motion/CAMBI SIMD but no SSIM). If upstream ever introduces SSIM SIMD, their kernel bodies will almost certainly compute
l*c*sin vector float for throughput — do not adopt. The fork's per-lane-scalar-double reduction is required for the bit-exactness claim. Same applies toconvolve_avx2/512— they are fork-only; dispatch sits inssim_tools.cvia_iqa_convolve_set_dispatch.
0034 — SIMD DX framework + NEON SSIM/convolve bit-exact port¶
- Workstream PRs:
feat/simd-dx-framework(this PR, PR #A); ships the two demos on top of which PR #B will consume the framework (ssimulacra2, motion_v2, vif_statistic, ...). - Touches:
core/src/feature/simd_dx.h(new header),core/src/feature/arm64/convolve_neon.c+convolve_neon.h(new NEON port),core/src/feature/arm64/ssim_neon.c(ssim_accumulate_neonrewritten for ADR-0139 bit-exactness;precompute+varianceunchanged),core/src/feature/float_ssim.c+core/src/feature/float_ms_ssim.c(wireiqa_convolve_neoninto the aarch64 dispatch setters),core/src/meson.build(arm64_sources+= convolve_neon.c),core/test/meson.build(test_iqa_convolvearch filter extended toarm64/aarch64),core/test/test_iqa_convolve.c(NEON variant check + aarch64 CPU flag detection),core/test/dnn/meson.build(test_cli.shgated onnot meson.is_cross_build()— bash invokes$VMAF_BINdirectly so meson's exe_wrapper isn't applied), newbuild-aux/aarch64-linux-gnu.inimeson cross-file,.claude/skills/add-simd-path/SKILL.md(upgraded kernel-spec flags). - Invariants (see ADR-0140 §Decision):
simd_dx.his fork-local. Keep the fork's version on upstream conflict. Macro names are ISA-suffixed (_AVX2_4L,_AVX512_8L,_NEON_4L) — do not collapse into a cross-ISA abstraction; the fork's SIMD policy (user-memoryfeedback_simd_dx_scope.md) rules out Highway / simde / xsimd.- The ADR-0138 widen-then-add rule (single-rounded
float * float→ widen →doubleadd, NO FMA) applies to NEON exactly as to AVX2 / AVX-512. The NEON form uses pairedfloat64x2_taccumulators (lo / hi) because NEON has nofloat64x4_t. - The ADR-0139 per-lane scalar-double reduction rule applies to
ssim_accumulate_neonexactly as to the AVX2 / AVX-512 variants. The NEON implementation usesSIMD_ALIGNED_F32_BUF_NEON(_Alignas(16) float name[4]) + a 4-iteration scalar loop. - Re-test (requires
aarch64-linux-gnu-gcc+qemu-user-static+ aarch64 sysroot at/usr/aarch64-linux-gnu):
cd libvmaf
meson setup ../build-aarch64 \
--cross-file ../build-aux/aarch64-linux-gnu.ini \
-Denable_cuda=false -Denable_sycl=false -Denable_dnn=disabled
cd ..
ninja -C build-aarch64
meson test -C build-aarch64 # expect 31/31 OK
# Bit-exactness check scalar vs NEON under QEMU:
REF=python/test/resource/yuv/src01_hrc00_576x324.yuv
DIS=python/test/resource/yuv/src01_hrc01_576x324.yuv
for m in 255 0; do
LD_LIBRARY_PATH=$PWD/build-aarch64/src qemu-aarch64-static \
-L /usr/aarch64-linux-gnu build-aarch64/tools/vmaf \
--cpumask $m --reference $REF --distorted $DIS \
--width 576 --height 324 --pixel_format 420 --bitdepth 8 \
--feature float_ssim --feature float_ms_ssim \
--output /tmp/ssim_$m.xml --precision max
done
diff <(grep -v '<fyi fps' /tmp/ssim_255.xml) \
<(grep -v '<fyi fps' /tmp/ssim_0.xml) # expect empty
- On upstream sync: upstream has no NEON SSIM and no NEON convolve for IQA. If they ever add one, keep the fork's version on conflict — the fork's NEON path is the only variant verified bit-exact to scalar at
--precision max. Thebuild-aux/aarch64-linux-gnu.inicross-file has no upstream equivalent. The/add-simd-pathskill is fork-only; upstream doesn't ship.claude/skills/.
0036 — Port Netflix generalised AVX convolve + ADR-0141 cleanup¶
- Workstream PRs:
port/upstream-f3a628b4-generalized-avx-convolve(this PR). - Upstream commit:
f3a628b4"feature/common: generalize avx convolution for arbitrary filter widths" (Kyle Swanson, 2026-04-21). - Touches:
- convolution.h — upstream-tracking: adds
#define MAX_FWIDTH_AVX_CONV 17. - convolution_avx.c — upstream-tracking (2,500 LoC deletion) plus fork-delta cleanup per ADR-0141: four scanline helpers
convolution_f32_avx_s_1d_*changed from external linkage tostatic(no other TU uses them after the specialised-path removal); stride parameters widened frominttoptrdiff_tin the helpers, with(ptrdiff_t)casts at public-function multiplication sites;#include <stddef.h>added for the type. core/src/feature/vif_tools.c— upstream-tracking: three AVX dispatch sites drop thefwidth == 17 || ... == 3whitelist in favour offwidth <= MAX_FWIDTH_AVX_CONV.python/test/quality_runner_test.py,python/test/vmafexec_test.py— upstream-authored loosening of two full-VMAF-score assertions fromplaces=2(±0.005) toplaces=1(±0.05). Adopted per the ADR-0142 Netflix-authority precedent (project rule #1 addresses fork drift, not upstream-authored test updates the fork must track).- Invariants (see ADR-0143 §Decision):
- Static linkage on scanline helpers — upstream leaves the four
convolution_f32_avx_s_1d_*_scanlinehelpers with external linkage out of habit; the fork narrows them tostatic. On upstream sync: if upstream ever externs them from another TU, that's a flag to re-audit; keep the fork'sstaticunless the reference is real. ptrdiff_tstrides inside helpers — the publicconvolution_f32_avx_*_swrappers keepintstrides (matching the upstream interface +convolution.hdeclarations). Helpers takeptrdiff_tto silencebugprone-implicit-widening-of- multiplication-result. If upstream changes the public interface toptrdiff_t, drop the fork's wrapper-level casts.MAX_FWIDTH_AVX_CONV = 17— the ceiling is upstream's; if upstream bumps it, the fork must rebuild + re-run the VIF golden test pair.- Re-test:
meson setup build -Denable_cuda=false -Denable_sycl=false
ninja -C build
meson test -C build # expect 32/32 OK
clang-tidy -p build core/src/feature/common/convolution_avx.c
# Zero warnings expected on the touched file.
Netflix CPU golden CI leg exercises the two loosened assertions; confirmed locally under meson test. - On upstream sync: upstream is the source of truth for convolution_avx.c, convolution.h, vif_tools.c dispatch, and the two python golden tolerances. On a rebase, prefer upstream for those files except: - Keep the fork's static on the four scanline helpers. - Keep the fork's ptrdiff_t helper signatures + multiplication- site casts (unless upstream adopts them too, in which case converge). - Keep the fork's #include <stddef.h>. If upstream re-introduces a specialised fast path for common widths, evaluate on a per-fwidth perf profile — the fork's /profile-hotpath skill covers this.
0038 — motion_v2 NEON SIMD (fork-local)¶
- Workstream PR:
port/motion-bundle-neon-and-updates(this PR). - Upstream: none — aarch64 NEON for
motion_v2is fork-local. Upstream scalar + AVX2 + AVX-512 variants exist; this PR adds the missing NEON fourth path. Scalar is the bit-exactness ground truth. - Touches (fork-local):
- motion_v2_neon.c — new TU, ~300 LoC. 4-wide int32 SIMD over the 5-tap Gaussian pipeline. Five
static inlinehelpers keep every function under the ADR-0141 60-line budget. - motion_v2_neon.h — new header declaring the two public entry points.
- integer_motion_v2.c — dispatch update: adds an
#if ARCH_AARCH64block ininitthat selects the NEON variant whenVMAF_ARM_CPU_FLAG_NEONis present, mirroring the existing x86 dispatch blocks. core/src/meson.build— addarm64/motion_v2_neon.cto thearm64_sourceslist.- Invariants (see ADR-0145 §Decision):
- Arithmetic right-shift throughout. The fork's AVX2 path uses
_mm256_srlv_epi64(logical) which can diverge from scalar on negative-diff pixels. The NEON port usesvshrq_n_s64(v, 16)for the known Phase-2 shift andvshlq_s64(v, -(int64_t)bpc)for the variable Phase-1 shift — both arithmetic, matching scalar C>>on signed integer. On rebase: keep the arithmetic forms; do NOT adoptvshrq_n_u64or a logical emulation even if it runs faster. - 4-lane stride + mirror tails. SIMD stride = 4; scalar tails cover the remainder. The Phase-2 helper
x_conv_row_sad_neonhands 4 lanes tox_conv_block4_neonand drops to scalar for both left/right edges (j < 2andj + 6 > w). On rebase: preserve the 4-lane stride and the two-sided scalar tail. - Signature parity with AVX2. Both pipeline entry points match the AVX2 + AVX-512 variants'
(const uint8_t *prev, ptrdiff_t, const uint8_t *cur, ptrdiff_t, int32_t *y_row, unsigned w, unsigned h, unsigned bpc)signature. On rebase: if upstream changes the signature, mirror the change here AND in the x86 variants in lockstep. - Re-test:
meson setup build-aarch64 libvmaf \
--cross-file build-aux/aarch64-linux-gnu.ini \
-Denable_cuda=false -Denable_sycl=false
ninja -C build-aarch64
meson test -C build-aarch64 --no-rebuild # expect 31/31 OK
clang-tidy -p build-aarch64 \
core/src/feature/arm64/motion_v2_neon.c
# Zero warnings expected on the touched file.
# NEON-vs-scalar bit-exact diff under QEMU:
YUV=python/test/resource/yuv
for mask in 0 255; do
LD_LIBRARY_PATH=build-aarch64/src \
qemu-aarch64-static -L /usr/aarch64-linux-gnu \
build-aarch64/tools/vmaf \
-r $YUV/src01_hrc00_576x324.yuv \
-d $YUV/src01_hrc01_576x324.yuv \
-w 576 -h 324 -p 420 -b 8 -n --feature motion_v2 \
--cpumask $mask -o /tmp/mv2_$mask.xml --precision max
done
diff <(grep -v 'fps=' /tmp/mv2_0.xml) \
<(grep -v 'fps=' /tmp/mv2_255.xml) # expect empty
- On upstream sync: upstream has no NEON
motion_v2and has not signalled plans to add one. If they ever do, diff their NEON against the fork's: on logical-vs-arithmetic shift, keep the fork's arithmetic form (matches scalar). On the function decomposition (the five helpers), adopt upstream's if it's smaller; the fork's layout is ADR-0141-driven, not a semantic contract. - Follow-up T7-32 (fixed 2026-05-09): The
_mm256_srlv_epi64(logical right shift) inmotion_score_pipeline_16_avx2was replaced withsrav_epi64_imm, an AVX2-safe arithmetic-right-shift emulation: logical shift OR sign-fill mask viasrai_epi32+slli_epi64. Two bugs were closed in the same PR: - AVX2 logical-vs-arithmetic shift:
_mm256_srlv_epi64replaced bysrav_epi64_immincore/src/feature/x86/motion_v2_avx2.c. The emulation is bit-exact with scalar C>> bpcon signedint64_t. - Test scalar reference mirror:
mirror_idxincore/test/test_motion_v2_simd.cused2*size - idx - 1instead of2*size - idx - 2, diverging frominteger_motion_v2.c::mirror(). Fixed to-2. All four adversarial fixtures (neg-diff bpc10/12, mixed-diff bpc10/12) now pass.meson test -C build50/50 OK. On rebase: keepsrav_epi64_imm; do not revert to_mm256_srlv_epi64. The rebase-time invariant is now: AVX2 path uses arithmetic shift (matching NEON and scalar).
0039 — readability-function-size NOLINT sweep (ADR-0146)¶
- ADR: ADR-0146
- Touches:
core/src/dict.ccore/src/picture.ccore/src/picture_pool.ccore/src/predict.ccore/src/libvmaf.ccore/src/output.ccore/src/read_json_model.ccore/src/feature/feature_extractor.ccore/src/feature/feature_collector.ccore/src/feature/iqa/convolve.ccore/src/feature/iqa/ssim_tools.ccore/src/feature/x86/vif_statistic_avx2.c- Invariant: every
readability-function-sizeNOLINT suppression has been replaced by a set of smallstatic(orstatic inline, for the SIMD / IQA files) helpers. The helper names are stable interfaces the surrounding code depends on (e.g.iqa_convolve_1d_separable,iqa_convolve_2d,ssim_compute_stats,ssim_workspace_alloc/_free,vif_stat_simd8_compute/_reduce,struct vif_simd8_lane,read_pictures_extractor_loop,read_pictures_post_extractor,read_pictures_validate_and_prep,read_pictures_update_prev_ref). Upstream Netflix has no equivalent helpers; rebases touching any of these files will conflict against the fork's split shape. - On upstream sync:
- If upstream lands a different decomposition of
_iqa_convolveor_iqa_ssim, prefer upstream's shape only if it keeps the ADR-0138 / ADR-0139 bit-exactness invariants (single-rounded float mul → widen to double → double add; per-lane scalar-float reduction through aligned temp buffer). Otherwise keep the fork's split and re-document the divergence here. - The fork renamed
_calc_scale→iqa_calc_scaleto clear thebugprone-reserved-identifiercheck. If upstream modifies_calc_scale, keep the fork's name and port the behavioural change. model_collection_parse_loopwrites directly tocfg_namerather than throughc->name— if upstream ever rewritesmodel_collection_parse, preserve the direct write (it's what lets the param stay non-const without a NOLINT).- Re-test on rebase (x86, any libsvm-less host):
ninja -C build && meson test -C build
for mask in 0 255; do
VMAF_CPU_MASK=$mask ./build/tools/vmaf \
--reference python/test/resource/yuv/src01_hrc00_576x324.yuv \
--distorted python/test/resource/yuv/src01_hrc01_576x324.yuv \
--width 576 --height 324 --pixel_format 420 --bitdepth 8 \
-m version=vmaf_v0.6.1 -o /tmp/vmaf_$mask.xml
done
diff <(grep -v fyi /tmp/vmaf_0.xml) <(grep -v fyi /tmp/vmaf_255.xml)
# expect exit 0 (Netflix-golden-pair VMAF bit-identical scalar vs SIMD)
Also run clang-tidy -p build on every file in Touches; expect zero warnings. - Follow-up T7-6: decide whether to rename the _iqa_* API surface (convolve / ssim / decimate / img_filter / filter_pixel / get_pixel) across all callers to clear the remaining bugprone-reserved-identifier suppressions in ssim.c, ms_ssim.c, float_ms_ssim.c. Out of scope here.
0040 — Thread-pool job recycling + inline data buffer (ADR-0147)¶
- ADR: ADR-0147
- Touches:
core/src/thread_pool.c - Invariants:
VmafThreadPoolJobcarries a fixed-sizechar inline_data[64]buffer. Payloads ≤ 64 bytes go throughmemcpy(job->inline_data, data, data_sz)+job->data = job->inline_data; payloads > 64 bytes take the legacymallocpath. The cleanup path MUST distinguish the two viajob->data != job->inline_data— a naivefree(job->data)would corrupt the slot. Enforced invmaf_thread_pool_job_clear_data.free_jobslist is protected by the existingqueue.lock; enqueue pops from it beforemallocing, runner recycles onto it after running a job.vmaf_thread_pool_destroywalks the list aftervmaf_thread_pool_waitreturns (all workers have exited → no lock needed). Any reorder that frees the queue lock before thefree_jobswalk is a leak on shutdown.- Fork's
void (*func)(void *data, void **thread_data)signature + per-workerVmafThreadPoolWorkerare fork-local; upstream Netflix #1464 hasfunc(void *data). Keep the fork's signature on any rebase — callers (src/libvmaf.c:threaded_enqueue_oneetc.) depend on the two-arg form. -
On upstream sync: Netflix PR #1464 is CLOSED (not merged) and bundles twelve unrelated optimizations. Only the thread-pool portion is ported here. If upstream ever reopens and merges #1464 (or a successor), cherry-pick only the pool mechanics; reject the payload-signature changes, the ADM / VIF / predict.c pieces (they conflict with ADR-0138 / 0139 / 0142 bit-exactness and with T7-5 predict.c refactor), and the feature-collector capacity bump (fork already capped at 8 for a reason — see
src/feature/feature_collector.c). -
Re-test on rebase (x86, any libsvm-less host):
ninja -C build && meson test -C build
for threads in 1 4; do
for mask in 0 255; do
VMAF_CPU_MASK=$mask ./build/tools/vmaf \
--reference python/test/resource/yuv/src01_hrc00_576x324.yuv \
--distorted python/test/resource/yuv/src01_hrc01_576x324.yuv \
--width 576 --height 324 --pixel_format 420 --bitdepth 8 \
-m version=vmaf_v0.6.1 --threads $threads -o /tmp/vmaf_${threads}_${mask}.xml
done
done
# Expect bit-identical scores (attribute order may differ across
# --threads 1 vs --threads 4 because feature-collector emits in
# insertion order; the numeric values match).
diff <(grep -v fyi /tmp/vmaf_4_0.xml) <(grep -v fyi /tmp/vmaf_4_255.xml)
# expect exit 0 (scalar vs SIMD threaded)
Also run clang-tidy -p build core/src/thread_pool.c — expect zero warnings. Re-run the 500 000-job micro-benchmark from ADR-0147 §Decision if performance is under investigation.
0041 — IQA reserved-identifier rename + cleanup (ADR-0148)¶
- ADR: ADR-0148
- Touches: 21 files across
core/src/feature/(iqa/{convolve,decimate,ssim_tools}.{c,h},iqa/ssim_simd.h,ssim.c,integer_ssim.c,ms_ssim.c,ms_ssim_decimate.h,float_ssim.c,float_ms_ssim.c,x86/convolve_avx2.{c,h},x86/convolve_avx512.{c,h},arm64/convolve_neon.{c,h},AGENTS.md) pluscore/test/test_iqa_convolve.c. - Invariants:
- Every
_iqa_*/_kernel/_ssim_int/_map_reduce/_map/_reduce/_context/_ms_ssim_*/_ssim_*/_alloc_buffers/_free_bufferssymbol and the four underscore-prefixed header guards (_CONVOLVE_H_,_DECIMATE_H_,_SSIM_TOOLS_H_,__VMAF_MS_SSIM_DECIMATE_H__) is renamed to its non-reserved spelling. The fork's IQA surface no longer uses C's reserved-identifier name space. - The
clang-analyzer-security.ArrayBoundNOLINT bracket inssim_accumulate_rowandssim_reduce_row_range(integer_ssim.c) is load-bearing — the inner kernel-loopk_min/k_maxclamping is provably correct (k_min = max(0, hkernel_offs - x),k_max = min(hkernel_sz, hkernel_sz - (x + hkernel_offs - w + 1))) but the analyzer can't follow it across helper boundaries. Do not collapse the bracket. - The
clang-analyzer-unix.MallocNOLINT bracket intest_iqa_convolve.c(check_simd_variant,check_case) is intentional — test exits process on failure path; small allocations leak by design at test end. Do not refactor to free-on-exit. - The cross-TU NOLINT pattern on
compute_ssim(ssim.c) andcompute_ms_ssim(ms_ssim.c) — clang-tidymisc-use-internal-linkageruns per-TU and can't see the header bridge tofloat_ssim.c/float_ms_ssim.c. Keep the inline justification comment. - On upstream sync:
- The Netflix upstream IQA library (
tjdistler/iqa) has been effectively abandoned (last meaningful commit pre-2020). Future rebases will conflict on every renamed symbol; drop the underscore-prefix on each conflict and mirror the fork'siqa_*naming. - If upstream Netflix/vmaf ever reincorporates the IQA naming wholesale, prefer the fork's spellings — this PR is a one-shot mechanical rename with no semantic content.
- Re-test on rebase:
ninja -C build && meson test -C build
for mask in 0 255; do
VMAF_CPU_MASK=$mask ./build/tools/vmaf \
--reference python/test/resource/yuv/src01_hrc00_576x324.yuv \
--distorted python/test/resource/yuv/src01_hrc01_576x324.yuv \
--width 576 --height 324 --pixel_format 420 --bitdepth 8 \
-m version=vmaf_v0.6.1 \
--feature float_ssim --feature float_ms_ssim \
-o /tmp/iqa_$mask.xml
done
diff <(grep -v fyi /tmp/iqa_0.xml) <(grep -v fyi /tmp/iqa_255.xml)
# expect exit 0 (bit-identical scalar vs SIMD on float_ssim/ms_ssim)
Also run clang-tidy -p build on every touched file (excluding arm64/); expect zero warnings.
0042 — Port Netflix #1376 — FIFO-hang fix via Semaphore (ADR-0149)¶
- ADR: ADR-0149
- Upstream commit: Netflix PR #1376, head
1c06ca4f1bb5da38b54db075a27c35ba8ea9d7b7(OPEN upstream as of 2026-04-24). - Touches:
python/vmaf/core/executor.py— baseExecutorclass +ExternalVmafExecutor-style subclass; delete_wait_for_workfiles/_wait_for_procfilespolling loops; rewrite_open_{work,proc}files_in_fifo_modearoundmultiprocessing.Semaphore(0); addopen_sem=Nonekwarg to every_open_{ref,dis}_{work,proc}fileand to the_open_workfilestaticmethod; drop unusedfrom time import sleep.python/vmaf/core/raw_extractor.py—AssetExtractor+DisYUVRawVideoExtractor; addopen_sem=Noneto_open_{ref,dis}_workfileoverrides (release on entry since these are no-ops); delete_wait_for_workfilesoverrides; drop unusedfrom time import sleep.- Fork carve-outs (load-bearing on rebase):
compat/python-vmaf/__init__.py:__version__follows the rootx-release-please-versionmarker — do NOT port upstream's bump to"4.0.0"independently. The fork uses one release stream per ADR-1127.from time import sleepis dropped from both files — upstream leaves the import in place (unused after their patch); the fork removes it because ADR-0141 touched-file rule requires ruff F401 clean.- Upstream typo preserved: the subclass warning message contains "to be created to be created". Comments note the typo inline; do not silently fix on rebase — it's upstream- authored and project policy is verbatim port.
- On upstream sync: upstream PR #1376 is still OPEN. When it merges, re-diff against the merged form; the touched hunks should be conflict-free because the fork now carries the same shape. Re-check whether upstream fixed the "to be created to be created" typo; if so, adopt the fix (it becomes a simple string update).
- Re-test:
python3 -m py_compile python/vmaf/core/executor.py \
python/vmaf/core/raw_extractor.py
ruff check python/vmaf/core/executor.py python/vmaf/core/raw_extractor.py
black --check python/vmaf/core/executor.py python/vmaf/core/raw_extractor.py
# all silent
# No FIFO-mode unit test in the tree; end-to-end harness
# exercise (needs libsvm + ffmpeg + fixtures) goes via
# make test-netflix-golden
# which doesn't exercise fifo_mode path but does verify the
# refactor didn't break executor.py imports.
0043 — Port Netflix #1472 — CUDA on Windows MSYS2/MinGW (ADR-0150)¶
- ADR: ADR-0150
- Upstream commits: Netflix PR #1472 —
15745cdf(portability) +b7b65e64(meson plumbing). Both OPEN upstream as of 2026-04-24. - Touches:
core/src/cuda/common.h— drop<pthread.h>include; rename reserved header guard__VMAF_SRC_CUDA_COMMON_H__→VMAF_SRC_CUDA_COMMON_INCLUDED.core/src/cuda/cuda_helper.cuh—#ifdef DEVICE_CODEguard around<cuda.h>vs<ffnvcodec/dynlink_loader.h>.core/src/picture.h—#ifdef DEVICE_CODEguard around<cuda.h>+ forward-declareVmafCudaStatevs<ffnvcodec/*>+ fulllibvmaf_cuda.h; rename reserved header guard.core/src/feature/integer_adm.h— updated comment abovedwt_7_9_YCbCr_thresholdtable noting the fork's positional-initializer shape vs upstream's#ifndef __CUDACC__shape (see §Fork carve-outs).core/src/feature/cuda/integer_adm/{adm_cm,adm_csf,adm_csf_den,adm_decouple,adm_dwt2}.cu—#ifndef DEVICE_CODEguard around#include "feature_collector.h".core/src/meson.build— Windows nvcc plumbing (+70 LoC underhost_machine.system() == 'windows'):vswhere-basedcl.exediscovery, MSVC + Windows SDK include path injection, CUDA version detection vianvcc --version,nvcc_ccbin_flags+nvcc_host_includesthreaded through everycustom_targetthat invokes nvcc.- Fork carve-outs (load-bearing on rebase):
integer_adm.huses positional initializers, NOT upstream's#ifndef __CUDACC__wrap. Both shapes resolve the MSVC/nvcc C++-designated-initializer issue; the positional form is C++-portable and keeps the table available to future.cuconsumers. Keep the fork's form on rebase.cuda_static_libkeepsdependencies : [pthread_dependency]. Upstream drops it; the fork needs it becausering_buffer.c(built as part ofcuda_static_lib)#includes<pthread.h>directly. On rebase: keep the fork's version.meson.buildgencode coverage block: the fork's ADR-0122 explicit cubin list (sm_75/80/86/89 + compute_80 PTX) sits after the new upstream nvcc-detect block. On rebase, re-assemble the same merged order: nvcc-detect first, then gencode coverage (both host-independent).- Header guards:
_INCLUDEDspellings are fork-local (ADR-0148 precedent). Upstream keeps reserved__VMAF_SRC_*_H__spellings. On rebase, keep_INCLUDED. - On upstream sync: PR #1472 is still OPEN. When merged, re-diff the three conflict-resolved hunks against upstream's final form. Keep fork's version on the four carve-outs above unless upstream meaningfully reshapes those regions.
- Re-test on rebase (Linux host with CUDA toolkit):
meson setup libvmaf core/build-cuda \
-Denable_cuda=true -Denable_nvcc=true -Denable_sycl=false
ninja -C core/build-cuda && meson test -C core/build-cuda
# Expect 6 .fatbin files generated + CLI linked + 35/35 tests pass.
Windows validation is operator-driven — CI does not yet have a Windows + MSYS2 + MinGW + MSVC BuildTools + CUDA runner (tracked as T7-3 in .workingdir2/OPEN.md). - Prerequisites note (Windows only): nv-codec-headers must be built from git master commit 876af32 or later. The release tag n13.0.19.0 is missing cuMemFreeHost, cuStreamCreateWithPriority, cuLaunchHostFunc, and other CudaFunctions members libvmaf uses. Pre-existing issue, not scope of this port.
0058 — libvmaf.pc Cflags leak fix (ADR-0200)¶
- ADR: ADR-0200; bug-fix follow-up to entry 0057.
- Upstream source: fork-local. Netflix has no Vulkan backend.
- Touches:
core/subprojects/packagefiles/volk/meson.build— drops-include volk_priv_remap.hfromvolk_dep.compile_args; keeps-DVK_NO_PROTOTYPES.core/src/vulkan/meson.build— pullsvolk_priv_remap_h_pathfrom the volk subproject and appends['-include', <path>]tovmaf_cflags_common(privatec_args:on libvmaf'slibrary()call).- Invariants (load-bearing):
-includeMUST stay offvolk_dep.compile_args— otherwise it leaks into staticlibvmaf.pcCflags. Test on rebase:meson setup ... -Ddefault_library=static -Denable_vulkan=enabled, thengrep Cflags meson-private/libvmaf.pc— must NOT containvolk_priv_remapor any build-dir absolute path.-includeMUST be applied to libvmaf's compile — every libvmaf TU that calls volk'svk*API needs the rename macros active. Thevmaf_cflags_commoninjection covers this for all libvmaf sub-libraries (libvmaf_feature, libvmaf_cpu, etc.).- The path comes from
subproject('volk').get_variable(...), not from a hardcoded string — survives volk wrap version bumps. - On upstream sync: zero upstream interaction.
- Re-test on rebase / volk wrap bump:
meson setup build-vk-static-test libvmaf -Denable_vulkan=enabled \
-Denable_cuda=false -Denable_sycl=false -Ddefault_library=static
ninja -C build-vk-static-test src/libvmaf.a
grep Cflags build-vk-static-test/meson-private/libvmaf.pc
# Expected: no `volk_priv_remap` substring, no build-dir absolute path
0057 — Volk vk* priv-remap for static-archive builds (ADR-0198)¶
- ADR: ADR-0198; follow-up to ADR-0185.
- Upstream source: fork-local. Netflix/vmaf has no Vulkan backend.
- Touches:
core/subprojects/packagefiles/volk/meson.build— overlay applied on top of the upstream volk wrap. Adds acustom_targetthat runsgen_priv_remap.pyto producevolk_priv_remap.hfrom the upstreamvolk.h, and wires-includeof the generated header intovolk.c'sc_argsandvolk_dep'scompile_args.core/subprojects/packagefiles/volk/gen_priv_remap.py— fork-added generator script (regex againstextern PFN_vkXxx vkXxx;declarations).- Invariants (load-bearing):
- Force-include must propagate to every libvmaf TU pulling in
volk_dep— verified via meson dep graph. Removing the-includefromcompile_argsre-introduces the static-link multi-def cascade. - Generator regex matches every
vk*PFN declaration involk.h— confirmed for volk-1.4.341 (784declarations,784remaps). Bumping the volk wrap version: re-run the generator (it's a configure-time custom target, so it's automatic) and confirm the rename count printed to stdout matches the count of^extern PFN_vklines in the newvolk.h. - The renamed symbols use the
vmaf_priv_prefix — chosen to match no upstream Netflix or Vulkan SDK identifier. Don't rename to_vk*(collides with reserved-identifier C namespace) orvkv_*etc. - On upstream sync: zero upstream interaction. The volk wrap is a libvmaf-managed subproject; Netflix doesn't ship a Vulkan backend.
- Re-test on rebase / after any volk wrap bump:
meson setup build-vk-static libvmaf -Denable_vulkan=enabled \
-Denable_cuda=false -Denable_sycl=false \
-Ddefault_library=static
ninja -C build-vk-static src/libvmaf.a
test "$(nm build-vk-static/src/libvmaf.a 2>/dev/null \
| grep -cE '^[0-9a-f]* (T|D|B|R) vk[A-Z]')" = "0" \
&& echo OK
(Followed by the BtbN-style link reproducer in the ADR References section.)
0056 — SSIMULACRA 2 snapshot gate + fp-contract-off split (ADR-0164)¶
- ADR: ADR-0164
- Upstream source: fork-local. Netflix/vmaf has no SSIMULACRA 2.
- Touches:
- python/test/ssimulacra2_test.py — new fork-added Python test. Uses
subprocess.callagainstExternalProgram.vmafexecwith--feature ssimulacra2; parses the--jsonoutput; asserts pooled + per-frame scores. - Invariants (load-bearing):
- Pinned values are CPU-only — generated on master HEAD after PR #100 merge. Re-generate if the scalar or any SIMD path changes semantically (which per ADR-0161/0162/0163's bit-exactness contract, it shouldn't — any bit-exact refactor leaves pinned values unchanged).
- Tolerance is 4 decimal places (
places=4) — matches 1e-4. The CPU paths are bit-exact so actual drift should be 0; the tolerance is defensive. -ffp-contract=offeverywhere in the ssimulacra2 pipeline:libvmaf_ssimulacra2_static_lib(scalar extractor),x86_ssimulacra2_avx2_lib,x86_ssimulacra2_avx512_lib, andarm64_ssimulacra2_lib(from ADR-0161). All four split out of their umbrella libs so other extractors keep upstream's default FMA policy. Without this the CI GCC/clang hosts drifted ~2e-4 from my AVX-512 authoring host — GCC 10+ defaults-ffp-contract=faston x86 with-mfmaand on aarch64, fusinga*b+cin scalar glue around the SIMD calls. Do NOT remove any of these carve-outs on rebase.- Fixtures are already-checked-in —
src01_hrc00/01_576x324is also the primary Netflix golden fixture; the 160×90 derived one stresses the sub-176 pyramid-termination path. - Do NOT modify the Netflix golden assertions in quality_runner_test.py et al. — those are upstream-pinned. This test is a SEPARATE file that adds fork-specific scores.
- On upstream sync: no upstream interaction. If Netflix adopts SSIMULACRA 2 in the future, cross-reference against their pinning if they add one.
- Re-test on rebase / after any ssimulacra2 change:
- Follow-ups:
- Cross-reference gate against libjxl
tools/ssimulacra2whenssimulacra2_rscargo install is fixed. - Expand fixture coverage if new YUV test assets land.
0055 — SSIMULACRA 2 picture_to_linear_rgb SIMD (ADR-0163)¶
- ADR: ADR-0163
- Upstream source: fork-local. Netflix/vmaf has no SSIMULACRA 2.
- Touches:
- ssimulacra2_avx2.{c,h} — new
ssimulacra2_picture_to_linear_rgb_avx2+ helpers (read_plane_scalar_s2,srgb_to_linear_lane_avx2,compute_matrix_coefs). - ssimulacra2_avx512.{c,h} — 16-wide AVX-512 port.
- ssimulacra2_neon.{c,h} — 4-wide aarch64 port.
- ssimulacra2.c — new
ptlr_fnfield inSsimu2State; dispatch wrapperconvert_picture_to_linear_rgbunpacksVmafPictureintosimd_plane_t[3]; init assigns AVX2/AVX-512/NEON pointers. - ssimulacra2_simd_common.h — new shared header declaring
simd_plane_t. Decouples SIMD TUs fromVmafPicturetype. - test_ssimulacra2_simd.c — new
test_ptlr_420_8,test_ptlr_420_10,test_ptlr_444_8,test_ptlr_444_10,test_ptlr_422_8subtests + scalar referencesref_read_plane,ref_srgb_to_linear,ref_picture_to_linear_rgb. - Invariants (load-bearing):
- Scalar-order matmul —
G = Yn + cb_g * Un + cr_g * Vnchained left-to-right in all three SIMD TUs. Regression test catches reordering drift (~1 ulp). - Per-lane scalar
powf— vector polynomial approximation would drift scalar bit-exactness. Do not replace the lane spill/reload pattern with a vector libm. simd_plane_tlayout —{data, stride, w, h}ordering assumed by all three SIMD TUs. The dispatch wrapper builds this fromVmafPicturefields; layout must match.- Bounds clamping in
read_plane_scalar_*mirrors scalar reference verbatim (if (sx < 0) sx = 0; if (sx >= pw) sx = pw-1;etc.). Do not simplify — removes per-lane safety at plane edges. - Arbitrary chroma ratios fall through to the
int64_tmultiplication branch. Don't remove it — SSIMULACRA 2 is supposed to accept non-standard ratios gracefully. - On upstream sync: no upstream interaction. If Netflix adopts SSIMULACRA 2 in the future and provides a SIMD YUV→RGB path, diff against the fork's — preserve the bit-exactness contract unless ADR-0142 Netflix-authority carve-out opens.
- Re-test on rebase:
ninja -C build && build/test/test_ssimulacra2_simd # 11/11
ninja -C build-aarch64 && \
qemu-aarch64-static -L /usr/aarch64-linux-gnu/ \
build-aarch64/test/test_ssimulacra2_simd # 11/11
- Follow-ups:
- T3-3 SSIMULACRA 2 snapshot-JSON regression test — still pending (gated on
tools/ssimulacra2availability). - SSIMULACRA 2 now has zero scalar hot paths. T3-1 closes in full with phases 1+2+3 (ADR-0161, 0162, 0163).
0054 — SSIMULACRA 2 FastGaussian IIR blur SIMD (ADR-0162)¶
- ADR: ADR-0162
- Upstream source: fork-local. No SSIMULACRA 2 extractor in upstream Netflix/vmaf.
- Touches:
- ssimulacra2_avx2.{c,h} — new
ssimulacra2_blur_plane_avx2+ 2 helpers (hblur_8rows_avx2,vblur_simd_8cols_avx2). - ssimulacra2_avx512.{c,h} — 16-wide port.
- ssimulacra2_neon.{c,h} — 4-wide aarch64 port, uses
vsetq_lane_f32in place of gather. - ssimulacra2.c — adds
blur_fnfunction pointer toSsimu2State, dispatch ininit_simd_dispatch(), call-site inblur_3plane. - test_ssimulacra2_simd.c — new
test_blur+ scalar reference (ref_blur_plane,ref_fast_gaussian_1d). - Invariants (load-bearing):
- Row-batching lane layout — horizontal pass lane
iMUST hold row(y_base + i). Gather index vector entries are(y_base + i) * w(stride-w). Changing this breaks bit-exactness vs scalar. - Scalar left-to-right summation order —
n2_k * sum - d1_k * prev1_k - prev2_kchained sequentially;o0 + o1 + o2at output time is(o0 + o1) + o2. Changing to(o0 + o2) + o1oro0 + (o1 + o2)will drift ~1 ulp and the regression test catches it. col_stateis 6 * w contiguous floats — layout is[prev1_0 | prev1_1 | prev1_2 | prev2_0 | prev2_1 | prev2_2]. SIMD loads assume this layout; changing field order requires updating all three SIMD TUs in lockstep withblur_plane.- NEON lane-set pattern — aarch64 has no gather intrinsic; 4 explicit
vsetq_lane_f32calls per input vector. Do not replace with ald1 {v.s}[lane]-style pseudo-gather without re-verifying bit-exactness. - Scalar tail in vertical pass matches scalar reference body verbatim. Any deviation breaks
memcmpequality on widths that aren't multiples of the SIMD width. - On upstream sync: no upstream interaction. If Netflix adopts SSIMULACRA 2 in the future and provides their own IIR blur SIMD, diff against the fork's and preserve the bit-exactness contract unless an ADR-0142 Netflix-authority carve-out is opened.
- Re-test on rebase:
meson setup build -Denable_cuda=false -Denable_sycl=false
ninja -C build
meson test -C build test_ssimulacra2_simd # 6/6
# aarch64:
ninja -C build-aarch64
qemu-aarch64-static -L /usr/aarch64-linux-gnu/ \
build-aarch64/test/test_ssimulacra2_simd # 6/6
- Follow-ups:
picture_to_linear_rgbSIMD — last scalar hot path in the extractor. 2 calls / frame. Low ROI but mechanical.- T3-3 SSIMULACRA 2 snapshot-JSON regression test — still pending.
0053 — SSIMULACRA 2 SIMD bit-exact ports (ADR-0161)¶
- ADR: ADR-0161
- Upstream source: fork-local. Upstream Netflix/vmaf has no SSIMULACRA 2 extractor at all (fork-added in ADR-0130).
- Touches:
- ssimulacra2_avx2.c / .h — 5 AVX2 kernels + per-lane
cbrtfhelper. - ssimulacra2_avx512.c / .h — 5 AVX-512 kernels; mechanical 16-wide widening of the AVX2 path.
- ssimulacra2_neon.c / .h — 5 NEON kernels; 4-wide aarch64 mirror.
- ssimulacra2.c — adds function-pointer dispatch fields to
Ssimu2State+init_simd_dispatch()helper, calls go through the pointers. - meson.build — registers the three SIMD TUs in
x86_avx2_sources/x86_avx512_sources/arm64_sources. - test_ssimulacra2_simd.c and
test/meson.build— new bit-exact test harness. - Invariants (load-bearing):
- Byte-for-byte bit-exactness to scalar on all 5 vectorised kernels under
FLT_EVAL_METHOD == 0. Regression caught pre- merge: naïve pairing(a+b)+(c+d)vs scalar((a+b)+c)+ddrifts by 1 ULP. Keep sequential scalar-order chains in all three SIMD TUs on rebase. cbrtfis per-lane scalar libm, not a polynomial. Any replacement with a vector cbrt would drift the ssimulacra2 score and break the regression test. Keep the spill/reload pattern.ssim_map/edge_diff_mapreductions use the ADR-0139 per-lanedoublescalar tail. Do NOT SIMD-reduce float lanes then lift to double — summation order changes.downsample_2x2deinterleave uses ISA-appropriate ops: AVX2vshufps+vpermpd, AVX-512vpermt2ps, NEONvuzp1q_f32+vuzp2q_f32. After deinterleave, sum order is((r0e+r0o)+r1e)+r1omatching scalar.#pragma STDC FP_CONTRACT OFFat every TU header. Ignored by aarch64 GCC (non-fatal-Wunknown-pragmas); kept for portability (clang, MSVC).- IIR blur +
picture_to_linear_rgbstay scalar in this PR. Follow-up PRs target these; when they land, re-verify bit-exactness viatest_ssimulacra2_simdexpansion. - Runtime dispatch order: AVX-512 > AVX2 on x86; NEON on aarch64; scalar fallback. Preserve on rebase.
- On upstream sync:
- Upstream has no SSIMULACRA 2 extractor; nothing to merge.
- If Netflix adopts SSIMULACRA 2 in the future, diff their implementation against the fork's scalar + SIMD TUs; keep the fork's bit-exactness contract absent a specific Netflix-authority carve-out ADR.
- Re-test on rebase:
meson setup build -Denable_cuda=false -Denable_sycl=false
ninja -C build
meson test -C build test_ssimulacra2_simd # 5/5
clang-tidy -p build core/src/feature/x86/ssimulacra2_avx2.c \
core/src/feature/x86/ssimulacra2_avx512.c
# aarch64:
ninja -C build-aarch64
qemu-aarch64-static -L /usr/aarch64-linux-gnu/ \
build-aarch64/test/test_ssimulacra2_simd # 5/5
clang-tidy -p build-aarch64 \
core/src/feature/arm64/ssimulacra2_neon.c
- Follow-ups:
- IIR blur vectorisation (
blur_planevertical-pass column batching) — the biggest frame-level wallclock win. picture_to_linear_rgbper-lanepowf— lower ROI but mechanical.- T3-3 SSIMULACRA 2 snapshot-JSON regression test — ADR-0130 deferred; still pending.
0052 — psnr_hvs SIMD bit-exact ports (ADR-0159 AVX2, ADR-0160 NEON)¶
- ADRs: ADR-0159 (AVX2), ADR-0160 (NEON sister port).
- Upstream source: fork-local. Upstream Netflix/vmaf has no psnr_hvs SIMD path.
- Touches:
core/src/feature/x86/psnr_hvs_avx2.c— AVX2 TU.core/src/feature/x86/psnr_hvs_avx2.h— AVX2 header.core/src/feature/arm64/psnr_hvs_neon.c— NEON TU (sister port, ADR-0160).core/src/feature/arm64/psnr_hvs_neon.h— NEON header.core/src/feature/third_party/xiph/psnr_hvs.c— addPsnrHvsState+ runtime dispatch ininit()(AVX2 underARCH_X86, NEON underARCH_AARCH64) + scoped NOLINTBEGIN/END around the upstream Xiph scalar block (kept verbatim as the bit-exact reference).core/src/meson.build— addx86/psnr_hvs_avx2.ctox86_avx2_sourcesandarm64/psnr_hvs_neon.ctoarm64_sources.core/test/test_psnr_hvs_avx2.c,core/test/test_psnr_hvs_neon.c— bit-exact unit tests (x86 and aarch64 respectively).core/test/meson.build— register both tests underenable_asm, arch-gated.- Invariants (load-bearing):
- Bit-exactness to scalar: every
od_coeff(int32) and every finalpsnr_hvs_{y,cb,cr,psnr_hvs}value the AVX2 path emits must be byte-identical to the scalar reference on the Netflix golden pairs. If a rebase introduces any pattern that breaks this (e.g. a floating-point horizontal reduce in the mask accumulator), the unit testtest_psnr_hvs_avx2will fail — don't relax the assertions; fix the SIMD path. - DCT butterfly layout:
butterfly → transpose → butterfly → transpose. The transpose lives insideod_bin_fdct8x8_avx2. Do not move it. - Float accumulators stay scalar: means / variances / mask / error accumulation in
calc_psnrhvs_avx2use the same per-block scalar loop as scalar psnr_hvs — bit-exact by construction. Do not vectorize these with horizontal reductions without replicating ADR-0139's per-lane scalar-float reduction pattern. The cross-block error accumulatorretis threaded throughaccumulate_error()by pointer, not returned-then-summed: each of the 64 per-coefficient contributions per block must hit the outerretdirectly, matching scalar's inlineret += ...atthird_party/xiph/psnr_hvs.cline 355. IEEE-754 float add is non-associative — summing into a local float and then adding the per-block total toretchanges the summation tree and drifts the Netflix golden by ~5.5e-5. #pragma STDC FP_CONTRACT OFFat the TU header disables FMA formation. Required:fmaf(a, b, c)can differ from(a*b)+cby 1 ulp, breaking bit-exactness. Do not remove the pragma; do not add-ffp-contract=fastto the build flags for this TU.- NOLINT suppressions are load-bearing — each cites ADR-0141 inline (bit-exactness scalar-diff auditability for the 30-butterfly function, scalar float→double promotion for
sqrt, extractor-registry extern linkage forvmaf_fex_psnr_hvs, upstream-Xiph scoped block for rebase parity). - On upstream sync:
- Upstream has no psnr_hvs SIMD as of 2026-04-24. Keep fork's version on conflict.
- If upstream ever touches
psnr_hvs.cfor non-SIMD reasons (e.g. a masking-table update), rebase the AVX2 TU to match line-for-line and re-runtest_psnr_hvs_avx2to confirm bit-exactness survives. - NEON follow-up PR is a sister port; its
arm64/psnr_hvs_neon.cwill mirror this ADR's invariants. On rebase, the two SIMD TUs must stay in lock-step with the scalar reference. - Re-test on rebase:
ninja -C build
meson test -C build test_psnr_hvs_avx2
# Expect: 5/5 subtests pass (DCT bit-exact on 3 random seeds +
# delta + constant input).
# CLI-level bit-exactness on Netflix golden (requires the YUV
# fixtures in python/test/resource/yuv/):
# VMAF_CPU_MASK=0 (scalar)
# VMAF_CPU_MASK=255 (AVX2 enabled)
# Diff per-frame psnr_hvs_{y,cb,cr,psnr_hvs} XML fields; expect
# byte-identical across all 3 golden pairs.
0051 — Netflix#1486 motion updates verified present (ADR-0158)¶
- ADR: ADR-0158
- Upstream source: Netflix upstream PR #1486 ("Port motion updates"), MERGED 2026-04-20 as commits
a44e5e6(code) +62f47d5(Netflix golden updates). - Touches: documentation-only; the actual code changes this ADR documents are already in the fork's master via earlier incremental motion3 / blend / five-frame-window commits.
- Invariants (load-bearing for future
/sync-upstream): - The
edge_8mirror fix (i_tap = height - (i_tap - height + 2)) is present atinteger_motion.c:240,x86/motion_avx2.c:147,x86/motion_avx512.c:147. If upstream's mirror line ever diverges again, this is the hunk to watch. - The
motion_max_valfeature option is atinteger_motion.c:57,118-120with default 10000.0 andFEATURE_PARAMflag. Upstream's default = fork's default; don't drift. VMAF_integer_feature_motion3_scoreoutput plumbing is ininteger_motion.c+alias.c.- Fork-local motion extensions (five-frame-window, moving-average, blend, fps_weight) are ADDITIONS on top of Netflix#1486. They are not upstream. Upstream changes to motion extractor internals may conflict with them — diff against
core/src/feature/integer_motion.con every rebase and check that the fork'sMIN(s->score * s->motion_fps_weight, s->motion_max_val)invocations are preserved (lines ~409, ~503). - On upstream sync: nothing to port from Netflix#1486 — it's absorbed. If a future upstream PR touches the same code paths, prefer upstream's version for the scalar/edge handling and the fork's version for the five-frame-window / blend extensions.
- Re-test on rebase:
ninja -C build
meson test -C build
# Expect: 35/35 pass.
# Verify the upstream markers are still in place after rebase:
grep -n "height - (i_tap - height + 2)\|motion_max_val\|VMAF_integer_feature_motion3_score" \
core/src/feature/integer_motion.c \
core/src/feature/alias.c \
core/src/feature/x86/motion_avx2.c \
core/src/feature/x86/motion_avx512.c
# Expect: matches at all 4 files. If any missing, the rebase
# silently dropped the Netflix#1486 content — investigate.
0050 — CUDA preallocation memory leak fix + vmaf_cuda_state_free (ADR-0157)¶
- ADR: ADR-0157
- Upstream source: Netflix upstream issue #1300 (OPEN since 2024; no maintainer fix as of 2026-04-24). User reports GPU memory rises monotonically across init/preallocate/fetch/close cycles.
- Touches:
core/include/libvmaf/libvmaf_cuda.h— new publicvmaf_cuda_state_free()API declaration.core/src/cuda/common.c— newvmaf_cuda_state_free()implementation;vmaf_cuda_release()now callscuda_free_functions();vmaf_cuda_state_init()gets an outer failure unwind;init_with_primary_context()releases the retained primary context onfail_after_pop.core/src/cuda/ring_buffer.c—vmaf_ring_buffer_close()now unlocks + destroys the mutex before freeing.core/test/test_cuda_preallocation_leak.c— new GPU-gated reducer (10-cycle loop with full cleanup).core/test/test_cuda_pic_preallocation.c,core/test/test_cuda_buffer_alloc_oom.c— add missingvmaf_cuda_state_free()+vmaf_model_destroy()calls aftervmaf_close()in every test that allocates these.core/test/meson.build— register the new reducer underenable_cudaguard.- Invariants (load-bearing):
- Public contract: every caller of
vmaf_cuda_state_init()MUST callvmaf_cuda_state_free()AFTERvmaf_close()on any VmafContext that imported the state. Informalfree(cu_state)is a silent double-free hazard AFTER close (vmaf_close's vmaf_cuda_release already memset's + frees CudaFunctions internals; vmaf_cuda_state_free only frees the heap allocation itself). vmaf_cuda_release()freesCudaFunctionsvia a saved pointer AFTER thememset. Order matters —memsetfirst socu_state->fis zeroed in the caller's struct, then free via the saved local. Do not re-order.vmaf_ring_buffer_close()unlocks BEFORE destroying the mutex (POSIX requires the mutex be unlocked for destroy).- The cold-start unwind in
init_with_primary_contextreleasescuDevicePrimaryCtxRetain's retained context ifcuStreamCreateWithPriorityfails. - The ADR-0122 / ADR-0123
is_cudastate_empty()null-guards at the top of every publicvmaf_cuda_*entry must continue to compose with the newvmaf_cuda_state_free()(which accepts NULL directly and doesn't call through to the CUDA API). - The new free call order in callers is:
vmaf_close(vmaf)→vmaf_cuda_state_free(cu_state)→vmaf_model_destroy(model). Reversing the first two produces a use-after-free. - On upstream sync:
- Upstream has no
vmaf_cuda_state_free()as of 2026-04-24. Keep the fork's version on any conflict. If upstream eventually lands the same API with a different spelling, prefer upstream's spelling and add a compat alias — but do not break the fork's ABI. vmaf_cuda_release()'scuda_free_functions()call is fork-local. On rebase, keep it.- The ring-buffer
pthread_mutex_unlock+pthread_mutex_destroypair is fork-local. On rebase, keep it. - If upstream refactors
VmafCudaStateownership semantics (unlikely — their pattern has been "leaked state in a long- lived process is acceptable" historically), re-audit this ADR and the new public API. - Re-test on rebase:
ninja -C core/build-cuda
meson test -C core/build-cuda
# Expect: 40/40 pass including test_cuda_preallocation_leak.
# ASan leak-check:
cd libvmaf && meson setup build-asan-cuda \
-Db_sanitize=address -Denable_cuda=true -Denable_sycl=false \
--buildtype=debug
ninja -C build-asan-cuda
ASAN_OPTIONS='detect_leaks=1:leak_check_at_exit=1' \
build-asan-cuda/test/test_cuda_preallocation_leak
# Expect: 0 bytes leaked from core/src/* frames.
# (~180 bytes in libcuda.so.1 is expected — driver's process-
# lifetime cuInit cache, does not grow per cycle.)
0049 — CUDA graceful error propagation (ADR-0156)¶
- ADR: ADR-0156
- Upstream source: Netflix upstream issue #1420 (OPEN as of 2026-04-24). Reports that two concurrent VMAF-CUDA processes crash the second one at
vmaf_cuda_buffer_allocdue toCHECK_CUDA(cuMemAlloc)→assert(0)on OOM. - Touches:
core/src/cuda/cuda_helper.cuh— redefinedCHECK_CUDAfamily. New macrosCHECK_CUDA_GOTO+CHECK_CUDA_RETURN+ helpervmaf_cuda_result_to_errno. Oldassert(0)semantics removed entirely.core/src/cuda/common.c,core/src/cuda/picture_cuda.c,core/src/libvmaf.c— allCHECK_CUDA(...)sites converted; cleanup labels added where contexts / buffers were pushed / allocated.core/src/feature/cuda/integer_motion_cuda.c,integer_vif_cuda.c,integer_adm_cuda.c— same conversion; 12statichelpers promotedvoid → int.core/test/test_cuda_buffer_alloc_oom.c— new GPU-gated reducer.core/test/meson.build— register new test underenable_cudaguard.- Invariants (load-bearing):
CHECK_CUDA_GOTO/CHECK_CUDA_RETURNmust never callassert(0)orabort()on a CUDA error. Any regression back to the upstream abort-on-error semantics re-introduces Netflix#1420 and the NDEBUG footgun.- Every
CHECK_CUDA_GOTOtarget label must pop any previously-pushed CUDA context and free any partially-constructed buffers before returning the errno. The graceful path must not leak resources. vmaf_cuda_result_to_errnouses numericCUresultvalues directly (0 / 1 / 2 / 3 / 4 / 101 / 201 / 400) so host TUs that don't include<cuda.h>can transitively consume the mapping via the inline function. If upstream renumbersCUresultenum values (historically stable — they've been fixed since CUDA 1.0), re-audit the switch.- ADR-0122 / ADR-0123
is_cudastate_empty(...)guards at the top of every publicvmaf_cuda_*entry point must stay — they run before the CUDA API is touched and compose cleanly with the new error propagation. - Twelve
statichelper signatures in the feature extractors areint-returning (wasvoid): any upstream-port that restores thevoidreturn silently regresses the error path. - On upstream sync:
- Upstream Netflix still uses
assert(0)inCHECK_CUDAas of 2026-04-24. Keep the fork's macro definitions incuda_helper.cuhon any upstream conflict — this file is fork-local behaviour. - If upstream eventually lands Netflix#1420 with a similar refactor, prefer the fork's version unless upstream's has identical semantics (no
assert(0)/ noabort()/ translatesCUresultto-errno). Re-verifytest_cuda_buffer_alloc_oomafter rebase. - If upstream adds new
CHECK_CUDA(...)sites in a port, rewrite them toCHECK_CUDA_GOTO/CHECK_CUDA_RETURNas part of the port commit. - If upstream changes any of the 12
statichelper signatures back tovoid, re-promote them tointduring the merge. - Re-test on rebase:
ninja -C core/build-cuda
meson test -C core/build-cuda
# Expect: 39/39 pass including test_cuda_buffer_alloc_oom.
# Reducer check — verify the OOM-to-errno path is live:
meson test -C core/build-cuda test_cuda_buffer_alloc_oom -v
# Expect subtests: request 1 TiB → -ENOMEM; request 0 bytes → 0.
clang-tidy -p core/build-cuda --quiet \
core/src/cuda/common.c \
core/src/cuda/picture_cuda.c \
core/src/feature/cuda/integer_motion_cuda.c \
core/src/feature/cuda/integer_vif_cuda.c \
core/src/feature/cuda/integer_adm_cuda.c \
core/src/libvmaf.c
# Expect exit 0 on every file.
0049 — compute_motion / picture_copy signature changes (b949cebf upstream port)¶
- Upstream commit: Netflix/vmaf b949cebf (feature/motion: port several feature extractor options)
- Prerequisite commit: Netflix/vmaf d3647c73 (picture_copy: add channel parameter)
- PR: upstream/port-b949cebf-motion
Rebase-sensitive invariants:
-
compute_motionsignature change —compute_motion()incore/src/feature/motion.c/motion.hnow takes an extraint motion_decimateparameter (themotion_add_scale1flag). Any new caller added in the fork that callscompute_motion()must pass this parameter. The SIMD integer motion callers (motion_avx2.c,motion_avx512.c) do NOT callcompute_motion()— they use the SAD/convolution dispatch table directly and are unaffected. -
vmaf_image_sad_csignature change — similarly gainsint motion_add_scale1. Any caller in the fork must be updated. Currently only called fromcompute_motion()internally. -
picture_copysignature change — gainsint channelas the last parameter (0=Y, 1=U, 2=V). Every caller in the tree has been updated to pass0(luma). When adding new callers that need UV planes, pass1or2. The fork's CUDA/SYCL/Vulkan callers have been updated in this PR. -
Default behavior preserved — all new options default to no-op values.
motion_add_scale1=false,motion_add_uv=false,motion_blend_factor=1.0,motion_fps_weight=1.0,motion_filter_size=5(= DEFAULT_MOTION_FILTER_SIZE). Integer and float motion2 scores are bit-identical to pre-port baseline. -
vif_scale_frame_sdependency avoided — the upstream b949cebf motion.c importsvif_scale_frame_sfrom vif_tools.h. The fork does not have this function yet (vif options chain is deferred, Research-0024 Strategy E). The bilinear downscaler formotion_add_scale1is implemented as local static functions inmotion.c(motion_scale_bilinear,motion_bilinear_interp,motion_mirror_f). When upstream's vif options chain is eventually ported, reconcile by replacing these local functions withvif_scale_frame_s.
Reproducer:
# verify bit-exactness (default options, scores must be identical):
./core/build/tools/vmaf \
--reference testdata/ref_576x324_48f.yuv \
--distorted testdata/dis_576x324_48f.yuv \
--width 576 --height 324 --pixel_format 420 --bitdepth 8 \
--model path=model/vmaf_v0.6.1.json \
--feature motion --no_prediction --json --output /tmp/motion.json
# integer_motion2 scores must match pre-port baseline at 6 decimal places.
0048 — i4_adm_cm int32 rounding overflow deliberately preserved (ADR-0155)¶
- ADR: ADR-0155
- Upstream source: Netflix upstream issue #955 (OPEN since 2020; no maintainer response as of 2026-04-24). Reports that
add_bef_shift_flt[idx] = (1u << (shift_flt[idx] - 1))incore/src/feature/integer_adm.cscales 1–3 overflowsint32_t(1u << 31 = 0x80000000wraps to-2147483648). Rounding term is sign-negated; ADM scales 1–3 biased low by ≈1 LSB per summed term. - Touches (documentation-only):
docs/adr/0155-adm-i4-rounding-deferred-netflix-955.md— new ADR (this entry's anchor).core/src/feature/integer_adm.c— in-file warning comment above the overflow site (add_bef_shift_flt[]initialiser loop around line 1277). No code change.core/src/feature/AGENTS.md— invariant note under "Rebase-sensitive invariants".- Invariants (load-bearing — do NOT silently "fix"):
integer_adm.ckeepsint32_t add_bef_shift_flt[3]with the overflowing1u << 31assignment. The Netflix golden assertions (python/test/quality_runner_test.py,vmafexec_test.py,feature_extractor_test.py) encode the buggy ADM output. Project hard rule #1 (ADR-0024) prohibits changing those assertions.- Any "fix" that changes ADM numerical output must land together with a coordinated Netflix-authored golden-number update (the ADR-0142 Netflix-authority carve-out). Until Netflix#955 closes upstream, there is no authority to track.
- On upstream sync:
- If Netflix finally lands a fix for #955 (widening the rounding term to
uint32_torint64_t), sync the C-side fix AND the updatedassertAlmostEqualvalues in the same merge. Re-runmake test-netflix-goldenand/cross-backend-diffon the golden pairs to verify the new numbers are consistent across CPU / CUDA / SYCL. - Remove the in-file warning comment above the
add_bef_shift_fltinitialiser loop, flip ADR-0155 toSuperseded by ADR-NNNN, and drop this rebase-notes entry. - If upstream instead closes #955 as wont-fix, keep this entry verbatim and update the ADR status to note upstream's closure.
- Re-test on rebase (gates the invariant by confirming the golden numbers are unchanged):
ninja -C build
make test-netflix-golden
# Expect: VMAF mean 76.66890… on src01_hrc00/01_576x324 golden
# pair — bit-identical to pre-rebase.
0047 — vmaf_score_pooled -EAGAIN for pending features (ADR-0154)¶
- ADR: ADR-0154
- Upstream source: Netflix upstream issue #755 (OPEN as of 2026-04-24). Upstream maintainer closed the door on the streaming use case in 2020 ("you cannot call vmaf_score_pooled() in a loop"); fork reopens it via error-code semantics without changing the retroactive-write design.
- Touches:
core/src/feature/feature_collector.c—vmaf_feature_collector_get_scorereturns-EAGAIN(was-EINVAL) when the requested index is valid but not yet written.core/src/feature/feature_collector.h— inlinevmaf_feature_vector_get_scorenow returns-EINVALfor null/out-of-range and-EAGAINfor not-written (was-1for both). Added#include <errno.h>. Rename reserved__VMAF_FEATURE_COLLECTOR_H__guard toVMAF_FEATURE_COLLECTOR_INCLUDED.core/test/test_score_pooled_eagain.c— new 4-subtest reducer.core/test/meson.build— register the new test.- Invariants (load-bearing, enforced by the reducer):
vmaf_feature_collector_get_score(fc, name, &score, i)returns-EAGAINiff the featurenameis registered andiis in range butscore[i].written == false.- The return stays
-EINVALfor (a) null pointers, (b)i >= feature_vector->capacity, (c) unknown feature name. - The inline fast-path
vmaf_feature_vector_get_scoreuses the same split. - On upstream sync: upstream has not changed the error semantics since 2020. If they do (unlikely), keep the fork's
-EAGAIN— it is strictly more informative and downstream code depending on the split would regress. - Re-test on rebase:
ninja -C build && meson test -C build test_score_pooled_eagain
# Expect: 4/4 subtests pass.
# Reducer check:
git stash push core/src/feature/feature_collector.c core/src/feature/feature_collector.h
ninja -C build && meson test -C build test_score_pooled_eagain
# Expect: Fail: 1 (tests fail without -EAGAIN split).
git stash pop
0046 — float_ms_ssim min-dim guard (ADR-0153)¶
- ADR: ADR-0153
- Upstream source: Netflix upstream issue #1414 (OPEN as of 2026-04-24). No upstream fix has landed; fork adds the guard independently.
- Touches:
core/src/feature/float_ms_ssim.c— add#include "log.h"+#include "iqa/ssim_tools.h"+ amin_dim = GAUSSIAN_LEN << (SCALES - 1)check at the start ofinit; extract SIMD dispatch into a newms_ssim_init_simd_dispatchhelper to keepinitwithin the ADR-0141 60-line budget.core/test/test_float_ms_ssim_min_dim.c— new 3-subtest reducer.core/test/meson.build— register the new test executable.- Invariant (load-bearing, enforced by the reducer):
float_ms_ssim.initreturns-EINVALwhenw < 176 || h < 176, where 176 is computed dynamically from the filter constants. The magic number is not hardcoded — changingSCALESorGAUSSIAN_LENupstream will auto-update the minimum. - On upstream sync: if Netflix upstream lands a similar init-time guard, keep the fork's version — the helper name
ms_ssim_init_simd_dispatchis fork-local (introduced to satisfy ADR-0141) and upstream's patch won't match. Both guards should be compatible; re-verify the reducer after rebase. - Re-test on rebase:
ninja -C build && meson test -C build test_float_ms_ssim_min_dim
# Expect: 3/3 subtests pass.
# Reducer check (confirms the guard is load-bearing):
git stash push core/src/feature/float_ms_ssim.c
ninja -C build && meson test -C build test_float_ms_ssim_min_dim
# Expect: Fail: 1 (tests fail without the guard).
git stash pop
0045 — vmaf_read_pictures monotonic-index guard (ADR-0152)¶
- ADR: ADR-0152
- Upstream source: Netflix upstream issue #910 (OPEN as of 2026-04-24). No upstream fix has landed; the fork adds the guard independently, per the 2021-10-14 maintainer comment that recommended exactly this shape.
- Touches:
core/src/libvmaf.c— addunsigned last_index+bool have_last_indexfields toVmafContext; prepend a monotonic-index check insideread_pictures_validate_and_prep(returns-EINVALon duplicates / regressions); update the two new fields at the tail of the same helper on success.core/test/test_read_pictures_monotonic.c— new 3-subtest reducer covering the Netflix#910 sequence and the two classes of rejection (duplicate, out-of-order).core/test/meson.build— register the new test executable.- Invariant (load-bearing, enforced by the reducer):
vmaf_read_pictures(vmaf, ref, dist, index)returns-EINVALwhenhave_last_index && index <= last_index. Flush (vmaf_read_pictures(vmaf, NULL, NULL, 0)) routes toflush_contextbefore the guard runs — flushing remains always-available independent of the last accepted index. - On upstream sync:
- If Netflix upstream eventually lands a similar guard at the API boundary, keep the fork's version — the helper function name (
read_pictures_validate_and_prep) is fork-local (ADR-0146), upstream's patch will target a different insertion point. Both guards should be compatible; re-verify the reducer after rebase. - If upstream instead lands an internal reordering mechanism (buffer-and-sort frames before dispatch), revisit this decision — the fork's API-level contract is stricter and may need to relax to match. Open a new ADR if so.
- Re-test on rebase:
ninja -C build && meson test -C build test_read_pictures_monotonic
# Expect: 3/3 subtests pass.
# Reducer check (confirms the guard is load-bearing):
git stash push core/src/libvmaf.c
ninja -C build && meson test -C build test_read_pictures_monotonic
# Expect: Fail: 1 (the test rejects the un-guarded behaviour).
git stash pop
0044 — i686 (32-bit x86) build-only CI job (ADR-0151)¶
- ADR: ADR-0151
- Upstream source: Netflix upstream issue #1481 (OPEN as of 2026-04-24). Reports i686 compile failure on
_mm256_extract_epi64. Workaround documented in the issue:-Denable_asm=false. - Touches:
build-aux/i686-linux-gnu.ini— new cross-file; gcc +-m32+cpu_family = 'x86'/cpu = 'i686'. Noexe_wrapper..github/workflows/libvmaf-build-matrix.yml— new matrix row withi686: trueflag + new install-deps step forgcc-multilib+g++-multilib; existing "Run tests" + "Run tox tests (ubuntu)" steps widened with&& !matrix.i686guards.- Invariants:
- The i686 matrix row pins
-Denable_asm=false— this is the upstream-documented workaround for_mm256_extract_epi64's missing declaration on 32-bit x86 targets. Do NOT remove the flag without first gating every_mm256_extract_epi64call site incore/src/feature/x86/adm_avx2.c+motion_avx2.c+adm_avx512.con__x86_64__. Removing the flag naively will re-break the build. - No
exe_wrapperin the cross-file: meson marks tests asSKIP 77even though the host can run i686 binaries natively. Build-only gate by design. - On upstream sync:
- If upstream Netflix fixes #1481 at source (by gating the intrinsic calls on
__x86_64__or by emulating via two_mm256_extract_epi32halves), sync the fix and re-enable ASM on the i686 row (drop-Denable_asm=falsefrommeson_extra). Re-verify bit-exactness via/cross-backend-diffon the x86_64 golden pair. - If upstream marks i686 unsupported in meson (e.g. via a hard error), the fork's i686 row should be removed or downgraded to
continue-on-error: true. - Re-test on rebase (Ubuntu host with
gcc-multilib):
meson setup libvmaf core/build-i686 \
--cross-file=build-aux/i686-linux-gnu.ini \
-Denable_asm=false \
-Denable_cuda=false -Denable_sycl=false
ninja -C core/build-i686
file core/build-i686/tools/vmaf
# Expect: ELF 32-bit LSB pie executable, Intel i386
CI runs this same sequence via the new matrix row.
0058 — Tiny-AI Netflix corpus training scaffold (ADR-0252)¶
- ADR: ADR-0252.
- Upstream source: fork-local. Netflix/vmaf has no tiny-AI training harness or MCP server.
- Touches:
ai/— training harness;NflxLocalDatasetloader reads from--data-root(never from a hardcoded path).docs/ai/training-data.md— corpus path convention and loader API docs; purely additive.mcp-server/vmaf-mcp/tests/test_smoke_e2e.py— new e2e smoke test; references only committed golden fixtures.- Invariants (load-bearing):
- Data path is local-only.
.workingdir2/netflix/is gitignored; no YUV from this corpus is ever committed. The--data-rootCLI flag must remain the sole mechanism for locating the corpus. - Smoke test uses only committed fixtures.
test_smoke_e2e.pyreferencespython/test/resource/yuv/src01_hrc00_576x324.yuv(a committed golden file), never the local corpus path. On upstream sync the golden YUV path must stay stable. - No Netflix golden assertion is modified. The
places=4tolerance intest_smoke_e2e.pyasserts against thevmaf_v0.6.1CPU reference; it is not a golden assertion and may be adjusted by/regen-snapshotswith justification. - On upstream sync: zero interaction with Netflix upstream. The
ai/subtree andmcp-server/are wholly fork-local; upstream merges are conflict-free here. If Netflix ever ships a training harness, reconcile separately. - Re-test on rebase:
cd mcp-server/vmaf-mcp && python -m pytest tests/test_smoke_e2e.py -v
# Requires: meson compile -C build (vmaf binary)
# Skips automatically if binary or golden YUV is absent.
0085 — Research-0030 Phase-3b multi-seed validation (Gate 1 passed)¶
- No ADR. Empirical research digest closing Gate 1 of the 3-gate v2 validation chain. Architecture decision unchanged.
- Upstream source: fork-local. Netflix has no multi-seed validation surface for tiny-AI training.
- Touches (additive only):
docs/research/0030-phase3b-multiseed-validation.md— per-seed PLCC tables + stability analysis + Gate 2/3 plan.ai/scripts/phase3_subset_sweep.py— adds--seedsflag (comma-separated list) + per-seed result aggregation.CHANGELOG.mdUnreleased § Added.- Invariants (rebase-relevant):
- The +0.0175 Δ is multi-seed mean PLCC, not seed-0 PLCC. Don't cite the +0.0106 from Research-0029 once Research-0030 lands; the multi-seed number is more trustworthy.
- Subset B is more stable than canonical-6 across seeds. Don't ship a v2 model citing single-seed numbers — always report multi-seed mean ± seed-mean-std for any tiny-AI metric in a future digest.
- The
--seedsflag aggregates by flattening (seed × fold) pairs. The reportedmean_plccis the mean of alln_seeds × n_foldsmeasurements;seed_mean_plcc_stdis the std across per-seed means, which is the right number for "is the result seed-stable". - On upstream sync: zero interaction. Fork-only research.
- Re-test on rebase: documentation-only PR; the runs/ files reproduce from the canonical command.
0084 — Research-0029 Phase-3b StandardScaler retry (positive result)¶
- No ADR. Empirical research digest; revives the Research-0026 hypothesis after the Research-0028 negative result. The architectural decision (ship
vmaf_tiny_v2) is gated on three validation steps documented in the digest §"Required before shipping". - Upstream source: fork-local. Netflix has no tiny-AI preprocessing-sensitivity analysis surface.
- Touches (additive only):
docs/research/0029-phase3b-standardscaler-results.md— per-fold tables + apples-to-apples comparison + 3-gate pre-shipping checklist.ai/scripts/phase3_subset_sweep.py— adds--standardizeflag +_standardize_inplacehelper.CHANGELOG.mdUnreleased § Added.- Invariants (rebase-relevant):
- StandardScaler statistics MUST be fit per-fold on the train split only. Fitting on the full data would leak held-out information into LOSO; the
_standardize_inplacehelper enforces this by taking only the train slice as input. - A shipped
vmaf_tiny_v2.onnxMUST bundle its scaler(mean, std)in the sidecar JSON per ADR-0049 — otherwise inference applies different normalisation than training and the win evaporates. Currently UN-implemented; tracked as a §"Caveats" #5 follow-up. - Subset B's feature list is the load-bearing finding:
adm2,adm_scale3,vif_scale2,motion2,ssimulacra2,psnr_hvs,float_ssim. Phase-3c experiments may shift the optimal arch / lr / epochs but should keep this set. - On upstream sync: zero interaction. Fork-only research.
- Re-test on rebase: documentation-only PR; the runs/ files are reproducible from the
--standardizeinvocation in §"Reproducer".
0082 — Research-0028 Phase-3 subset sweep (negative-result digest)¶
- No ADR. Empirical research digest. The architectural decision (no v2 model ships from this Phase) is governed by Research-0027's pre-registered stopping rule.
- Upstream source: fork-local. Netflix has no tiny-AI subset- sweep surface.
- Touches (additive only):
docs/research/0028-phase3-subset-sweep.md— per-fold tables adline + standardisation caveat + Phase-3b/c/d follow-ups.CHANGELOG.mdUnreleased § Added.- Invariants (rebase-relevant):
- canonical-6 stays the default until Phase-3b lands a ≥ 0.005 PLCC win (per Research-0027 stopping rule).
- The PLCC drop is most likely a feature-scale issue, not evidence the new features lack signal. Don't cite this digest to retire
ssimulacra2/adm_scale3from the candidate pool; re-test withStandardScalerfirst. - Phase-3 results are seed=0 only. Any v2-shipping decision needs 3-seed mean±std and KoNViD cross-check.
- On upstream sync: zero interaction. Fork-only research.
- Re-test on rebase: documentation-only PR; runs/ files are reproducible from the canonical command in §"Reproducer".
0081 — Research-0027 Phase-2 feature importance results¶
- No ADR. Empirical research digest closing Research-0026 Phase 2; the architectural decision (Subset A / B / C) is deferred to Phase-3 results in a future digest.
- Upstream source: fork-local. Netflix has no cross-metric feature-importance analysis surface.
- Touches (additive only):
docs/research/0027-phase2-feature-importance.md— per-method top-10 + consensus + redundancy + Phase-3 subset recommendations.CHANGELOG.mdUnreleased § Added.- Invariants (rebase-relevant):
- Consensus top-10 is the load-bearing finding:
adm2,adm_scale3,ssimulacra2,vif_scale2. Phase-3 candidate subsets MUST include all four. - The 11-pair redundancy table is corpus-specific — measurements on Netflix Public 9-source. KoNViD-1k cross- check is a Phase-3 prerequisite if Subsets B/C advance.
runs/full_features_netflix.parquetandruns/full_features_correlation.jsonstay gitignored. Reproducer in §"Reproducer" regenerates both.- On upstream sync: zero interaction. Fork-only research.
- Re-test on rebase: documentation-only PR; the
runs/files are reproducible from the canonical commands.
0080 — Phase-2 analysis scripts (Research-0026 Phase 2 prep)¶
- No ADR. Pure analysis scaffolding; the architectural decision (which features to ship in v2) is gated on Phase 2's numerical output via Research-0027.
- Upstream source: fork-local. Netflix has no tiny-AI training nor cross-metric correlation tooling.
- Touches (additive only):
ai/scripts/extract_full_features.py— parquet extractor over Netflix corpus withFULL_FEATURES. Per-clip JSON cache at$XDG_CACHE_HOME/vmaf-tiny-ai-full/<source>/<dis_stem>.json.ai/scripts/feature_correlation.py— Pearson + MI + LASSO- consensus top-K analyser; outputs JSON.
ai/tests/test_feature_correlation.py— 5 pytest cases against synthetic parquet (no libvmaf dependency).CHANGELOG.mdUnreleased § Added.- Invariants (rebase-relevant):
- The per-clip JSON cache and the
FULL_FEATUREStuple must stay in lock-step. If the tuple grows (or shrinks), pre-existing cache files become stale and silently misalign their storedper_framecolumns with the new tuple. The extractor MUST be re-run with a cleared cache whenFULL_FEATURESchanges. Regression hint:test_default_features_unchangedintest_feature_sets.pyalready guards the canonical 6; extend coverage toFULL_FEATURESif rebases touch it. motion3resolves to extractormotion_v2in_METRIC_TO_EXTRACTOR, notmotion3(the upstream-canonical extractor name in the integer_motion_v2 module). The CLI--feature motion3does NOT exist. The JSON output key isinteger_motion3which_lookupfinds via theinteger_fallback.admandvifaggregates are NOT inFULL_FEATURES. The integer extractor emitsinteger_adm2andinteger_vif_scale0..3but no bareadm/vif. Listing them produced all-NaN columns in v1 — fixed in PR #185 amend.- On upstream sync: zero interaction. Pure fork-side analysis tooling.
- Re-test on rebase:
pytest ai/tests/test_feature_correlation.py ai/tests/test_feature_sets.py -v
# Expect: 14 passed in <1 s.
0079 — Tiny-AI feature-set registry (Research-0026 Phase 1)¶
- No ADR. Pure additive extension of an existing module; the architectural decision (which features, which model) lives in Research-0026's go/no-go gate after Phase 2.
- Upstream source: fork-local. Netflix/vmaf has no tiny-AI training pipeline.
- Touches (additive only):
ai/data/feature_extractor.py— addsFULL_FEATURES(21 entries),FEATURE_SETSregistry,resolve_feature_set()helper._METRIC_TO_EXTRACTORgrew 11 → 25 entries.ai/tests/test_feature_sets.py— new 9-test smoke suite.CHANGELOG.mdUnreleased § Added.- Invariants (rebase-relevant — these are load-bearing):
DEFAULT_FEATURESstays the canonical 6-tuple matchingvmaf_v0.6.1's SVR input layout. Testtest_default_features_unchangedis the regression guard; any quiet broadening would invalidate every shipped tiny-AI ONNX (input-dim baked into the model). If a future change must broaden the default, ship a paired model swap under ADR-0049 sidecar policy.FULL_FEATURESexcludeslpipsandfloat_momentper Research-0026 §"Open questions" Q1. Testtest_full_features_excludes_lpips_and_momentenforces. Adding either would re-classify the experiment from "tiny model on classical features" to "ensemble of DNNs".- Every entry in
FULL_FEATURESMUST have an entry in_METRIC_TO_EXTRACTOR. Testtest_every_full_feature_has_extractor_mappingis the guard — without the mapping the libvmaf CLI silently emits NaN columns for the missing metric. - On upstream sync: zero interaction. Fork-only training surface.
- Re-test on rebase:
0078 — Research-0026 cross-metric feature fusion plan¶
- No ADR. Pure research-plan digest; the architectural decision (which features to add) is deferred to Research-0027 follow-up after Phase 2 numbers land.
- Upstream source: fork-local. Netflix/vmaf has no tiny-AI training and no broader-feature-set hypothesis under investigation.
- Touches (additive only):
docs/research/0026-cross-metric-feature-fusion.md— 4-phase experimental plan + cost estimate + go/no-go criteria.CHANGELOG.mdUnreleased § Added.- Invariants (rebase-relevant):
- The 6-feature canonical baseline (
adm2,vif_scale0..3,motion2) stays the default. Any v2 model is opt-in via a newfeature_setfield in the sidecar JSON; existingvmaf_tiny_v1.onnxusers get the same numbers. lpipsis OUT of the candidate pool (Phase 1/2). It's DNN-based and would blur the line between "tiny model on classical features" and "ensemble of DNNs". Revisit only if classical features can't close the gap.- On upstream sync: zero interaction. Pure fork-side research planning.
- Re-test on rebase: documentation-only; no test surface.
0077 — Research-0025 FoxBird outlier resolved via KoNViD combined training¶
- No ADR. Empirical research digest closing the open question in Research-0023 §5; no architecture or policy decision. Pure documentation of an empirical result.
- Upstream source: fork-local. Netflix/vmaf has no tiny-AI training, no KoNViD-1k integration, and no LOSO eval surface.
- Touches (additive only):
docs/research/0025-foxbird-resolved-via-konvid.md— per-clip table + comparison to Netflix-only baselines + interpretation + caveats + next-experiment list.CHANGELOG.mdUnreleased § Added.- Invariants (rebase-relevant):
- The training-fit per-clip numbers in §"Per-clip result" are NOT held-out generalisation metrics — FoxBird is in the training set. The proper validation is the LOSO sweep on the combined corpus (§"Next experiments" #1). Don't cite the 0.9936 FoxBird PLCC as a generalisation number; cite it as "training-fit on combined corpus, 5.4× RMSE improvement vs Netflix-only".
- Combined trainer command line is canonical. The reproduction recipe in §"Setup" includes
--seed 0,--konvid-val-fraction 0.1,--val-source Tennis,--val-mode netflix-source-and-konvid-holdout. Changing any knob invalidates the per-clip numbers. runs/tiny_combined_canonical/stays gitignored. The final ONNX is reproducible from the parquet + Netflix corpus + the canonical CLI; the durable record is the digest's table.- On upstream sync: zero interaction. Research digest is fork-only.
- Re-test on rebase:
python ai/train/train_combined.py \
--netflix-root .workingdir2/netflix \
--konvid-parquet ai/data/konvid_vmaf_pairs.parquet \
--model-arch mlp_small --epochs 30 --batch-size 256 --lr 1e-3 \
--val-mode netflix-source-and-konvid-holdout \
--val-source Tennis --konvid-val-fraction 0.1 --seed 0 \
--out-dir runs/tiny_combined_canonical
# Expect: FoxBird PLCC ≈ 0.9936 ± 1e-3 (numerical-noise floor),
# mean PLCC ≥ 0.9983 across 9 Netflix clips.
0076 — Research-0024 vif/adm upstream-divergence digest (Strategy E doc)¶
- No ADR. Pure documentation digest; the divergence decisions it ratifies are already governed by ADR-0138 / 0139 / 0142 / 0143 (vif SIMD bit-exactness contract) and ADR-0024 (Netflix golden-data immutability). The digest itself fits the per-PR research-digest deliverable bar from ADR-0108.
- Upstream source: forward-looking — pre-emptively documents the fork's non-port of Netflix
4ad6e0ea/41d42c9e/bc744aa3/8c645ce3(vif chain) and4dcc2f7c(float_adm chain). Strategy A onb949cebfmotion chain stays approved. - Touches (additive only):
docs/research/0024-vif-upstream-divergence.md— 5-strategy decision matrix + numerical-risk analysis for each chain.core/src/feature/AGENTS.md— two new "rebase-sensitive invariants" entries pinning the vif and adm divergences.CHANGELOG.mdUnreleased § Changed.- Invariants (rebase-relevant — these are the whole point):
- Do not port
4ad6e0ea(vif runtime helpers) or8c645ce3(vif prescale options) verbatim. They replace the precomputedvif_filter1d_table_stable whose frozenconst floatGaussians make AVX2 == AVX-512 == NEON == scalar bit-for-bit. A future opt-in second-path port (Strategy C, runtime helpers behind--vif-prescale != 1) is allowed but must not touch the default code path. - Do not port
4dcc2f7cfloat_adm options chain. The 12-parametercompute_admsignature change cascades through SIMD (avx2 / avx512 / neon) and 3 GPU backends (vulkan / cuda / sycl). The newaimfeature has no fork- side golden values; defer until concrete user demand. - Mirror bugfix
41d42c9eis a separate decision. Must come paired withplaces=4 → places=3golden loosening per ADR-0142 Netflix-authority precedent. Not part of Strategy E; eligible for a focused single-purpose PR if any shipped model drifts more thanplaces=3because of the missing fix. b949cebfmotion chain port stays APPROVED under Strategy A (verbatim, float_motion-side only). Float_motion has no precomputed-table investment to protect; existing fork integer_motion already has 6/9 of these options; cheap to mirror onto float_motion.- On upstream sync: zero conflict — pure additions to research/ and AGENTS.md.
- Re-test on rebase: documentation-only PR; rendered markdown is the only verification surface.
# Re-run the diff scan that produced the digest (catches new
# upstream commits since 9dac0a59):
git fetch upstream && git log --pretty=format:'%h %s' \
upstream/master ^origin/master --since="2026-01-01" \
-- core/src/feature/{float_,integer_,}{vif,motion,adm,cambi}*.{c,h} \
core/src/feature/{vif,motion,adm,cambi}_options.h \
| head -30
# If new vif / adm option ports appear, update Research-0024 §"Same
# divergence test for motion + float_adm" before deciding to port.
0075 — Upstream 798409e3 + 314db130 ports (CUDA null-deref + remove all.c)¶
- No ADR. Pure upstream cherry-picks per ADR-0108 carve-out ("pure upstream syncs and
port-upstream-commitPRs are exempt"). - Upstream source:
798409e3(Lawrence Curtis, 2026-04-20): "Fix null deref crash on prev_ref update in pure CUDA pipelines"314db130(Kyle Swanson, 2026-04-28): "libvmaf/feature: remove empty translation unit all.c"- Touches (additive / removal only):
core/src/libvmaf.c— addsif (ref && ref->ref)guard beforevmaf_picture_ref(&vmaf->prev_ref, ref)at the two threaded paths (threaded_enqueue_oneline 1057 andthreaded_read_pictures_batchline 1105). Main path at line 1597 already has the guard.core/src/feature/all.c— file deleted.core/src/meson.build— drops thefeature_src_dir + 'all.c'line.core/src/feature/offset.c— updates the// NOLINTNEXTLINEcomment to dropall.cfrom the list of per-feature consumers.CHANGELOG.mdUnreleased § Fixed (798409e3) + § Changed (314db130).- Invariants (rebase-relevant):
- The fork has THREE prev_ref update sites; all need the
if (ref && ref->ref)guard. The mainvmaf_read_picturespath already had it (viaread_pictures_update_prev_refhelper); the threaded paths (#ifdef VMAF_BATCH_THREADING) inherited the unguarded shape from upstream's old code. Future upstream rebases must preserve all three guards even if Netflix refactors the threaded paths. all.cdeletion is symbol-safe. Allcompute_*functions it forward-declared are reached via per-extractor TUs that#includethe relevant<feature>.h. No external linker dependency onall.c's symbols.- On upstream sync: zero conflict expected — fork now matches upstream tip on these two surfaces.
- Re-test on rebase:
meson setup build-cpu libvmaf -Denable_cuda=false -Denable_sycl=false \
-Denable_vulkan=disabled
ninja -C build-cpu
meson test -C build-cpu # 37 tests, all pass.
0074 — Combined Netflix + KoNViD-1k trainer driver¶
- No ADR. Pure engineering follow-up; the architecture rationale is fully covered by ADR-0203 (training-prep architecture) and Research-0023 §5 (FoxBird-class outlier needs broader corpus).
- Upstream source: fork-local. Netflix/vmaf has no tiny-AI trainer.
- Stacks on the KoNViD-1k loader bridge (PR #178 / rebase-note 0073). Rebase order: land 0073 first.
- Touches (additive only):
ai/train/train_combined.py— concatenating trainer that reuses_build_model/_train_loop/export_onnxfromai/train/train.py.ai/tests/test_train_combined_smoke.py— 5 pytest cases (key splitter +--epochs 0paths, no libvmaf or real corpus required).docs/ai/training.md— "Combining KoNViD with the Netflix corpus" subsection rewritten from "follow-up" to runnable.CHANGELOG.mdUnreleased § Added.- Invariants (rebase-relevant):
- Reuse the canonical training-loop helpers. Don't fork
_build_model/_train_loop/export_onnxinto this file. Both trainers must share the model factory so a future change (e.g. addingmlp_large) lands in one place. - KoNViD train/val splits hold out whole clip keys, not random frames. A frame-level split would let frames from the same clip leak across train/val and inflate PLCC by 5-10 pp (well-known VQA pitfall — same reasoning as ADR-0203's Netflix 1-source-out split).
- Missing data falls back, not errors. Missing
--konvid-parquet→ Netflix-only path. Missing--netflix-root→ KoNViD-only path. Both missing → initial- weights ONNX export +rc=0so the smoke command always produces a deterministic artefact. - On upstream sync: zero interaction; pure fork-local trainer.
- Re-test on rebase:
pytest ai/tests/test_train_combined_smoke.py -v
# Expect: 5 passed (under ~3 s, no libvmaf required).
python ai/train/train_combined.py --epochs 0 \
--netflix-root /tmp/missing --konvid-parquet /tmp/missing.parquet \
--out-dir /tmp/combined_smoke
# Expect: <out-dir>/mlp_small_combined_final.onnx written, rc=0.
0073 — KoNViD-1k → VMAF-pair acquisition + loader bridge¶
- No ADR. Acquisition + loader pieces are pure additions; the methodology fits inside ADR-0203 / Research-0019.
- Upstream source: fork-local. KoNViD-1k integration is a fork-only training-data play.
- Touches (additive only):
ai/scripts/konvid_to_vmaf_pairs.py— acquisition pipeline.ai/train/konvid_pair_dataset.py—KoNViDPairDatasetclass mirroringNetflixFrameDataset's interface.ai/tests/test_konvid_pair_dataset.py— 5 pytest cases.docs/ai/training.md— new "C1 (KoNViD-1k corpus)" section.CHANGELOG.mdUnreleased § Added.- Invariants (rebase-relevant):
KoNViDPairDatasetmirrorsNetflixFrameDatasetshape.feature_dim == 6,numpy_arrays() → (X, y)returns(n_frames, 6)+(n_frames,). IfNetflixFrameDataset's feature order changes, mirror it here.- Acquisition parquet schema is fixed. Required columns:
key,frame_index,vif_scale0..3,adm2,motion2,vmaf. Add freely; do NOT rename / drop those. ai/data/konvid_vmaf_pairs.parquetand$VMAF_TINY_AI_CACHE/konvid-1k/stay gitignored. They regenerate from raw KoNViD.mp4sources.- On upstream sync: zero interaction.
- Re-test on rebase:
pytest ai/tests/test_konvid_pair_dataset.py -v
# Expect: 5 passed
python ai/scripts/konvid_to_vmaf_pairs.py --max-clips 5
# Expect: ~7 s wall, ai/data/konvid_vmaf_pairs.parquet with
# 5 unique keys × ~200 frames each.
0072 — Tiny-AI 3-arch LOSO eval harness + Research-0023¶
- No ADR. Methodology fits inside Research-0023; ADR-0203 already covers the training-prep architecture and the three-arch sweep concept.
- Research digest:
docs/research/0023-loso-3arch-results.md. - Upstream source: fork-local. Netflix/vmaf has no LOSO eval surface.
- Touches (additive only):
ai/scripts/eval_loso_3arch.py— new harness; reuses the_load_session+_load_clip+CLIPShelpers fromeval_loso_mlp_small.py(PR #165).docs/research/0023-loso-3arch-results.md— methodology + per-fold tables formlp_small/mlp_medium/linear.CHANGELOG.mdUnreleased § Added.- Invariants (rebase-relevant):
- Reuse the PR #165 helpers. Don't fork the
_load_sessionexternal-data workaround into a copy — both scripts must keep using the same import. If a follow-up re-exports the shipped baselines with correctedexternal_data.location, both scripts deprecate the workaround simultaneously. runs/andmodel/tiny/training_runs/stay gitignored. The harness writesruns/loso_eval/loso_3arch_eval.{json,md}; the durable record is the table in Research-0023 §2 + the per-fold tables in §3. Regenerate via the loop in §6 of the digest.- On upstream sync: zero interaction. Pure fork-local evaluation harness.
- Re-test on rebase:
python ai/scripts/eval_loso_3arch.py
diff <(jq -r '.archs.mlp_small.aggregate.mean_plcc' runs/loso_eval/loso_3arch_eval.json) <(echo 0.9808)
diff <(jq -r '.archs.mlp_medium.aggregate.mean_plcc' runs/loso_eval/loso_3arch_eval.json) <(echo 0.9727)
diff <(jq -r '.archs.linear.aggregate.mean_plcc' runs/loso_eval/loso_3arch_eval.json) <(echo 0.3679)
# Expect: identical lines on a populated cache + identical fold ONNX.
0071 — T7-16 ADM Vulkan/SYCL drift verified-resolved (doc close)¶
- No ADR. Verification-only close, sister of T7-15.
- Upstream source: fork-local. ADM cross-backend gate is a fork-only test surface; Netflix/vmaf has no Vulkan or SYCL backend.
- Touches (additive only):
docs/state.md— new "Recently closed" row for T7-16..workingdir2/BACKLOG.md— T7-16 row marked closed (local- only planning dossier; gitignored).CHANGELOG.mdUnreleased § Fixed.- Invariants (rebase-relevant):
places=4cross-backend ADM contract. Empiricaladm_scale2max_abs_diff is now 1e-6 (print floor; ULP=0) on Vulkan device 0 (NVIDIA), device 1 (Mesa anv on Arc), and SYCL device 0 (Arc); residualadm_scale1 ≈ 3.1e-5andadm2 ≈ 5e-6on 1/48 frames passplaces=4(5e-5 tolerance) but failplaces=5. Hold the gate atplaces=4.- No ADM kernel source change. Fix is environmental (NVCC + driver + SYCL runtime).
- On upstream sync: zero interaction.
- Re-test on rebase:
python3 scripts/ci/cross_backend_vif_diff.py \
--vmaf-binary core/build/tools/vmaf \
--feature adm --backend vulkan --device 0 --places 4 \
--reference python/test/resource/yuv/src01_hrc00_576x324.yuv \
--distorted python/test/resource/yuv/src01_hrc01_576x324.yuv \
--width 576 --height 324
# Expect: 0/48 mismatches across all 5 ADM metrics.
0070 — T7-15 motion CUDA/SYCL drift verified-resolved (doc close)¶
- No ADR. Verification-only close; no code change in PR #172.
- Upstream source: fork-local. Cross-backend gate is a fork-only test surface; not in Netflix/vmaf.
- Touches (additive only):
docs/state.md— "Recently closed" row for T7-15..workingdir2/BACKLOG.md— T7-15 row marked closed (local- only planning dossier; gitignored).CHANGELOG.mdUnreleased § Fixed.- Invariants (rebase-relevant):
- The
places=4cross-backend gate stays atplaces=4. Empirical max_abs_diff is currently 0.0 (CUDA) or 1e-6 (SYCL/ Vulkan, JSON%frounding floor); tightening toplaces=5could be tempting but the 1e-6 print-floor would then make the SYCL + Vulkan rows fail. Hold atplaces=4until--precision=maxis wired into the diff tool. - No motion-kernel source change. PR #172 didn't modify
core/src/feature/cuda/integer_motion/*.cuorcore/src/feature/sycl/integer_motion_sycl.cpp. The fix is environmental (NVCC + driver), so the next CI run on a fresh image needs to be re-verified against the gate. - On upstream sync: zero interaction.
- Re-test on rebase:
python3 scripts/ci/cross_backend_vif_diff.py \
--vmaf-binary core/build/tools/vmaf \
--reference python/test/resource/yuv/src01_hrc00_576x324.yuv \
--distorted python/test/resource/yuv/src01_hrc01_576x324.yuv \
--width 576 --height 324 --feature motion --backend cuda \
--places 4
# Expect: 0/48 mismatches, max_abs_diff = 0.0
0069 — libvmaf_vulkan.h installed under prefix (build bug)¶
- No ADR. Build-system bug fix; matches existing CUDA / SYCL install conditions.
- Upstream source: fork-local. Vulkan backend is fork-only; Netflix/vmaf has no
libvmaf_vulkan.h. - Touches:
core/include/core/meson.build— adds anis_vulkan_enabledgate that handles thefeatureoption'senabled/autostates; appendslibvmaf_vulkan.htoplatform_specific_headerswhen active.CHANGELOG.mdUnreleased § Fixed.- Invariants (rebase-relevant):
- Install rule mirrors the CUDA / SYCL pattern but uses the feature-option API. The
is_cuda_enabled = get_option('enable_cuda') == trueboolean idiom doesn't apply toenable_vulkanbecause that's a feature option, not a boolean. Use.enabled() or .auto(). Don't "simplify" to== true— that would silently drop the install in theautostate. - Pairs with
ffmpeg-patches/0006-libvmaf-add-libvmaf-vulkan-filter.patchwhich probes for the header viacheck_pkg_config libvmaf_vulkan "libvmaf >= 3.0.0" libvmaf/libvmaf_vulkan.h vmaf_vulkan_state_init_external. Removing the install rule re-introduces lawrence's 2026-04-28 symptom: FFmpeg silently drops thelibvmaf_vulkanfilter despite--enable-libvmaf-vulkan. - On upstream sync: zero interaction; Vulkan backend is fork-only.
- Re-test on rebase:
cd libvmaf
CC=icx CXX=icpx meson setup build -Denable_vulkan=enabled \
-Denable_cuda=true -Denable_sycl=true -Db_lto=false
ninja -C build
meson install -C build --destdir /tmp/libvmaf-install
ls /tmp/libvmaf-install/usr/local/include/libvmaf/libvmaf_vulkan.h
# Expect: file exists.
0066 — --backend cuda inverted-gpumask fix (CLI bug)¶
- No ADR. Bug fix; behaviour now matches the public-header
VmafConfiguration::gpumaskcontract. - Upstream source: fork-local. The
--backendCLI selector was added by the fork (Netflix/vmaf has no exclusive-backend selector). - Touches (additive + 1-line behavioural fix):
core/tools/cli_parse.c::parse_cli_args—--backend cudabranch setsgpumask = 0(wasgpumask = 1).core/test/test_cli_parse.c— 5 new regression tests (test_backend_{cpu,cuda_engages_cuda,cuda_preserves_explicit_gpumask,sycl,vulkan}) plusrun_aom_ctc_tests/run_backend_testshelper split to keeprun_testsunder the function-size budget.CHANGELOG.mdUnreleased § Fixed.- Invariants (rebase-relevant):
VmafConfiguration::gpumasksemantics:if gpumask: disable CUDA.compute_fex_flagsinsrc/libvmaf.croutes CUDA only whengpumask == 0. Any code path that sets a non-zerogpumaskto "request CUDA" silently disables it. The CLI's--backend cudabranch must setgpumask = 0and rely onuse_gpumask = trueto triggervmaf_cuda_state_init. Do not "fix" this back togpumask = 1— it's the bug being fixed.- Explicit
--gpumask=N --backend cudapreserves N. A user who passes--gpumask=2already hasuse_gpumask = true, so the--backend cudabranch's defaulting block (gated on!settings->use_gpumask) is skipped. Thetest_backend_cuda_preserves_explicit_gpumaskregression locks this in. - On upstream sync: zero interaction;
--backendis fork-only. - Re-test on rebase:
./build/test/test_cli_parse | grep -E 'backend_'
# Expect: 5 backend tests pass.
build/tools/vmaf -r REF -d DIS -w 576 -h 324 -p 420 -b 8 \
--model "path=model/vmaf_v0.6.1.json" --threads 1 \
--backend cuda --output cuda.json --json -q
python3 -c "import json; d=json.load(open('cuda.json')); \
assert len(d['frames'][0]['metrics']) == 12, 'CUDA not engaged'"
0067 — Tiny-AI PTQ accuracy across Execution Providers (T5-3e)¶
- No ADR. Investigation/measurement PR; ADR-0129 already governs the PTQ workstream. Findings update
docs/research/0006-tinyai-ptq-accuracy-targets.md§"GPU-EP quantisation" — that section was previously a deferred-open-question; it is now the empirical landing spot. - Research digest: same file (Research-0006).
- Upstream source: fork-local. Netflix/vmaf does not ship a PTQ harness or any tiny-AI ONNX path.
- Touches (additive only):
ai/scripts/measure_quant_drop_per_ep.py— new sibling ofmeasure_quant_drop.py. CPU+CUDA via ORT; Arc / OpenVINO-CPU via the nativeopenvinoPython runtime (noonnxruntime-openvinobecause no cp314 wheel exists). Reuses the_load_sessionrename workaround from PR #165 + avalue_info-strip fix so dynamic-PTQ doesn't choke on the shipped MLP ONNX.docs/ai/quant-eps.md— new user doc; linked fromdocs/ai/index.md.docs/research/0006-tinyai-ptq-accuracy-targets.md— refreshed header, replaced "GPU-EP open question" with the measurement table, fixed pre-existing MD040/MD060 lints surfaced on the touched file.docs/ai/index.md— added the quant-eps row, rewrapped to 80 cols.CHANGELOG.mdUnreleased § Changed.- Invariants (rebase-relevant):
measure_quant_drop.py(the CI gate) is unchanged. The new script is purely additive. Any rebase that conflates the two scripts must keep the CI gate CPU-only — Arc int8 is broken, so a per-EP gate would red-light every PR.value_infostrip is required forvmaf_tiny_v1*dynamic PTQ. The shipped MLP ONNX duplicate weight tensors invalue_info, which makesquantize_dynamicraiseInferred shape and existing shape differ. The fix is in_save_inlined. Don't remove it during a refactor unless the underlying ONNX is regenerated.- CUDA-12 ABI shim. ORT-GPU 1.25 wheels link
libcublasLt.so.12even on CUDA-13 hosts. The reproduction recipe pins thenvidia-*-cu12wheels and prepends them toLD_LIBRARY_PATH. If a future ORT wheel drops the cu12 ABI we can cut the shim, but the script tolerates either since it doesn't import any CUDA symbol itself. - On upstream sync: zero interaction; entirely fork-local.
- Re-test on rebase:
SP=$VIRTUAL_ENV/lib/python3.14/site-packages/nvidia
export LD_LIBRARY_PATH="$SP/cublas/lib:$SP/cudnn/lib:$SP/cuda_nvrtc/lib:$SP/cuda_runtime/lib:$SP/cufft/lib:$SP/curand/lib:$SP/cusolver/lib:$SP/cusparse/lib:$SP/cuda_cupti/lib:$SP/nvtx/lib:$SP/nvjitlink/lib"
python ai/scripts/measure_quant_drop_per_ep.py \
--eps cpu cuda openvino \
--extra-fp32 vmaf_tiny_v1.onnx vmaf_tiny_v1_medium.onnx \
--out runs/quant-eps-$(date +%Y-%m-%d)
# Expected: CPU + CUDA PASS (drop ≤ 1.2e-4); OpenVINO Arc ERR
# (compile failure for Conv-int8) or NaN (MatMul-int8) until a
# newer intel_gpu plugin lands.
0065 — testdata/bench_all.sh correct backend-engagement flags¶
- No ADR. Bug fix; no behavioural surface change beyond "the bench actually engages the backends it claims to now."
- Upstream source: fork-local.
testdata/bench_all.shis a fork-only bench harness; not in Netflix/vmaf. - Touches (additive only):
testdata/bench_all.sh— switched per-row flag pattern from the disable-only singletons (--no_syclfor "CUDA", etc.) to the correct engagement form (--gpumask=0 --no_sycl --no_vulkanfor CUDA,--sycl_device=0 --no_cuda --no_vulkanfor SYCL,--vulkan_device=0 --no_cuda --no_syclfor Vulkan, and--no_cuda --no_sycl --no_vulkanfor CPU). Added a 4th column (Vulkan) to the comparator. Honours$VMAF_BINfor the binary path and$VMAF_ONEAPI_SETVARSfor the oneAPI install location.CHANGELOG.mdUnreleased § Fixed.- Invariants (rebase-relevant):
- Disable-only singletons don't engage a backend.
--no_syclalone leaves CUDA available but unrequested.--no_cudaalone leaves SYCL available but unrequested. The CLI inits CUDA only whenc.use_gpumaskis set; SYCL only whenc.sycl_device >= 0orc.use_gpumask; Vulkan only whenc.vulkan_device >= 0. Any change to those gates that drops one of the per-row flags will re-introduce the silent CPU fallback. Verify after a rebase by recording each live row's JSONframes[0].metricskey count. Treat a GPU count equal to CPU as a fallback warning, never as a fixed expected backend count — seelibvmaf/AGENTS.md§"Backend-engagement foot-guns". gpumasksemantics are inverted from intuition.gpumask=0enables CUDA dispatch;gpumask=1disables it. The per-row CUDA flag is--gpumask=0, not--gpumask=1. Don't "fix" it to--gpumask=1for symmetry with sycl_device/vulkan_device — that's the bug being fixed (parallel to PR #170).- On upstream sync: zero interaction;
testdata/bench_all.shis fork-only. - Re-test on rebase:
VMAF_BENCH_OUTDIR=testdata/bbb/results bash testdata/bench_all.sh
# Record actual live-backend counts and compare within this run:
jq '.frames[0].metrics | keys | length' testdata/bbb/results/t1_cpu.json
jq '.frames[0].metrics | keys | length' testdata/bbb/results/t1_cuda.json
jq '.frames[0].metrics | keys | length' testdata/bbb/results/t1_sycl.json
0063 — Tiny-AI LOSO eval harness for mlp_small¶
- No ADR. The methodology fits inside Research Digest 0022; ADR-0203 already covers the training-prep architecture.
- Research digest:
docs/research/0022-loso-mlp-small-results.md. - Upstream source: fork-local. Netflix/vmaf has no LOSO eval surface.
- Touches (additive only):
ai/scripts/eval_loso_mlp_small.py— new evaluation harness.docs/ai/loso-eval.md— usage doc.docs/research/0022-loso-mlp-small-results.md— methodology + results.CHANGELOG.mdUnreleased § Added.- Invariants (rebase-relevant):
_load_sessionworkaround for renamed-baseline ONNX. The shipped baselinesmodel/tiny/vmaf_tiny_v1*.onnxreference their pre-renameexternal_data.locationvalues. The workaround in_load_sessionrewrites the entries before handing the proto to ORT. Removing the workaround breaks the baseline phase. The proper fix (re-export with matching names) is tracked as a follow-up; until then this code path is load-bearing.runs/andmodel/tiny/training_runs/stay gitignored. The harness writes toruns/loso_eval/by default; do NOT promote any of those outputs into the tree. The 9 fold ONNX and the per-clip JSON cache regenerate from the corpus + trainer + libvmaf CLI.- On upstream sync: zero interaction. Pure fork-local evaluation harness.
- Re-test on rebase:
python ai/scripts/eval_loso_mlp_small.py
diff <(jq -r '.loso_aggregate.mean_plcc' runs/loso_eval/loso_mlp_small_eval.json) <(echo 0.9808)
# Expect: identical line on a populated cache + identical fold ONNX.
0064 — Section-A audit: 9 backlog rows + ADR cross-links¶
- No ADR. Process / docs PR; rows trace back to the individually-cited ADRs / research digests in their own References columns.
- Decision dossier:
.workingdir2/decisions/section-a-decisions-2026-04-28.md. - Source audit:
docs/backlog-audit-2026-04-28.md. - Upstream source: fork-local. Pure backlog hygiene PR; no Netflix code touched.
- Touches (additive only):
.workingdir2/BACKLOG.md— 9 new rows: T3-17, T3-18, T5-3e, T5-4, T7-35, T7-36, T7-37, T7-38; T6-1a row extended with the bisect-cache fixture sub-bullet.docs/research/0006-tinyai-ptq-accuracy-targets.md— drops the "defer until first user" framing on the GPU-EP quantisation open question per user direction; cross-links T5-3e.docs/research/0020-cambi-gpu-strategies.md— v2 follow-up section now cites T7-36 as the gate for opening the v2 row.docs/adr/0205-cambi-gpu-feasibility.md— Decision section's "follow-up integration PR" now cites T7-36.CHANGELOG.mdUnreleased § Changed.- Invariants (rebase-relevant): none. Pure backlog text. Rebase-conflict risk is limited to the same
BACKLOG.mdtable rows that any future row addition would touch; trivial to re-resolve. - On upstream sync: zero interaction.
- Re-test on rebase: none — docs-only.
0062 — ssimulacra2 CUDA + SYCL twins (ADR-0206)¶
- ADR: ADR-0206.
- Upstream source: fork-local. Netflix/vmaf has no SSIMULACRA 2 GPU implementation; this PR adds the CUDA + SYCL twins of the fork's ADR-0201 Vulkan kernel.
- Touches (additive + small wiring edits):
docs/adr/0206-ssimulacra2-cuda-sycl.mdand the index row indocs/adr/README.md.core/src/feature/cuda/ssimulacra2_cuda.{c,h}— new CUDA dispatch.core/src/feature/cuda/ssimulacra2/ssimulacra2_blur.cuandssimulacra2_mul.cu— new CUDA fatbins.core/src/feature/sycl/ssimulacra2_sycl.cpp— new SYCL extractor.core/src/feature/feature_extractor.c— two new extern declarations + two new entries infeature_extractor_list[].core/src/meson.build— addsssimulacra2_blur+ssimulacra2_multocuda_cu_sources, introduces (or extends, if PR #157 / ADR-0202 landed first) thecuda_cu_extra_flagsmap with assimulacra2_blurentry, threadsper_kernel_flagsinto the fatbin custom-target, and lists the two new C / CPP TUs.core/src/cuda/AGENTS.mdandcore/src/sycl/AGENTS.md— rebase invariant notes for the per-kernel--fmad=falseflag and the-fp-model=preciseSYCL build flag.docs/backends/cuda/overview.md,docs/backends/sycl/overview.md,docs/metrics/features.md— coverage matrix updates.CHANGELOG.mdUnreleased § Added.- Invariants (load-bearing on rebase):
- Per-kernel
--fmad=falseforssimulacra2_blur. The IIR'so = n2 * sum - d1 * prev1 - prev2must NOT fuse into FMAs — without the flag the recursive Gaussian's per-step rounding compounds across the 6-scale pyramid pastplaces=4. -fp-model=preciseon the SYCL feature build line. Removing it driftsssimulacra2_syclpastplaces=2through the IIR.- Hybrid host/GPU split mirrors Vulkan. Host runs YUV→RGB, XYB, downsample, and SSIM/EdgeDiff combine in double; GPU runs only mul + IIR blur. Any future PR that ports XYB or YUV→RGB onto the GPU MUST land alongside an updated ADR-0206 and re-validate
places=4on every Netflix CPU pair. - CUDA fex uses
.extract(synchronous), not.submit/.collect. Per-frame raw YUV is D2H-copied frompicture_cuda's device-sideVmafPicture.data[]into pinned host scratch viacuMemcpy2DAsync. Skipping the copy segfaults — direct host reads on aCUdeviceptrare the failure mode the prior agent's WIP hit. - On upstream sync: zero interaction with Netflix. The GPU coverage matrix for
ssimulacra2is wholly fork-local. - Re-test on rebase:
meson setup build_cuda libvmaf -Denable_cuda=true -Denable_sycl=false
ninja -C build_cuda
python3 scripts/ci/cross_backend_vif_diff.py \
--vmaf-binary ./build_cuda/tools/vmaf \
--feature ssimulacra2 --backend cuda --places 4 \
--reference python/test/resource/yuv/src01_hrc00_576x324.yuv \
--distorted python/test/resource/yuv/src01_hrc01_576x324.yuv \
--width 576 --height 324 --pixel-format 420 --bitdepth 8
# Expect: 0/48 mismatches, max_abs_diff ~1e-6.
0061 — cambi GPU feasibility spike (ADR-0205)¶
- ADR: ADR-0205.
- Research digest:
docs/research/0020-cambi-gpu-strategies.md. - Upstream source: fork-local. Netflix/vmaf has no Vulkan backend.
- Touches (additive only):
docs/adr/0205-cambi-gpu-feasibility.md,docs/research/0020-cambi-gpu-strategies.md,docs/adr/README.mdindex row.core/src/feature/vulkan/cambi_vulkan.c— new dormant scaffold (not yet invulkan_sources, not yet registered).core/src/feature/vulkan/shaders/cambi_{derivative,decimate,filter_mode}.comp— new reference GLSL shaders, not yet in the build'sshaderslist.core/src/feature/AGENTS.mdinvariants +CHANGELOG.mdbullet.- Invariants (rebase-relevant):
- Hybrid host/GPU port by decision. If Netflix upstream tightens the c-value formula or histogram update protocol, the host residual call site in the eventual
cambi_vulkan.c::cambi_vulkan_extractmust be updated alongsidecambi.c::calculate_c_values— the same code is reused. Do NOT translate the c-values phase to GPU during any upstream-port PR; that optimisation belongs to the v2 strategy-III PR (deferred). - Scaffolds dormant in the spike PR. The
cambi_vulkan.cextractor returns-ENOSYSfromcambi_vulkan_init_stubuntil the integration follow-up wires it in. Do NOT registervmaf_fex_cambi_vulkan_scaffoldinfeature_extractor.c's list. - Shaders not in the build's shader list. Adding them to
core/src/vulkan/meson.build'svulkan_shaderslist before the integration PR produces orphaned*_spv.hheaders. Leave them alone in this spike PR. - On upstream sync: zero interaction. cambi.c itself is upstream-mirrored — Netflix changes flow through
port-upstream-commit; only the integration PR's host residual call site needs paired attention. - Re-test on rebase:
```bash meson setup build -Denable_vulkan=enabled -Denable_cuda=false -Denable_sycl=false ninja -C build meson test -C build
0059 — Tiny-AI Netflix corpus training prep (ADR-0203)¶
- ADR: ADR-0203.
- Upstream source: fork-local. Netflix/vmaf has no equivalent training surface.
- Touches:
ai/data/— Netflix loader, libvmaf-CLI feature extractor, distillation scoring.ai/train/— PyTorch dataset, eval harness, Lightning-style training entry point.ai/scripts/run_training.sh— convenience wrapper.ai/tests/— five new pytest modules (test_netflix_loader.py,test_dataset.py,test_eval.py,test_train_smoke.py, plusconftest.py).docs/ai/training.md— new "C1 (Netflix corpus)" section; existing sections untouched.ai/AGENTS.md— invariants section added.- Invariants (load-bearing):
- Filename ladder regex is fork-specific.
<source>_<quality>_<height>_<bitrate>.yuv(dis) +<source>_<fps>fps.yuv(ref). Upstream may publish a different naming convention later; do NOT merge them — keep this loader scoped to the Netflix corpus, add a sibling loader for any upstream alternative. - Per-clip cache schema is consumed by both dataset and any downstream tooling. Schema is
{features:{feature_names, per_frame, n_frames}, scores:{per_frame, pooled}}. Any change must invalidate$VMAF_TINY_AI_CACHE(delete or version-tag the directory). - Smoke command stays runnable without a built
vmafbinary. The_make_zero_payloadhelper inai.train.datasetinjects a fake payload for--epochs 0so CI gates don't drag a libvmaf build into the Python test surface. - YUV size probe never silently guesses.
probe_yuv_dimseither matches the 1920x1080 default, returns ffprobe's answer, or raises. Tests passassume_dims=(16, 16)explicitly for synthetic fixtures. - On upstream sync: no interaction with upstream. The
ai/subtree is wholly fork-local. - Re-test on rebase:
python -m pytest ai/tests/test_netflix_loader.py \
ai/tests/test_dataset.py ai/tests/test_eval.py \
ai/tests/test_train_smoke.py -v
python ai/train/train.py --epochs 0 --data-root /tmp/mock_corpus \
--assume-dims 16x16 --val-source BetaSrc --out-dir /tmp/out
0073 — Tiny-AI QAT trainer + first per-model QAT pass (T5-4)¶
- ADR: ADR-0207 (design), ADR-0208 (per-model impl).
- Touches:
ai/train/qat.py(new),ai/scripts/qat_train.py(rewrite fromNotImplementedErrorscaffold),ai/configs/learned_filter_v1_qat.yaml(new),ai/tests/test_qat_smoke.py(new),docs/ai/quantization.md(QAT tier added). All paths are wholly fork-local; no upstream Netflix/vmaf interaction. - Invariants:
- Two-step pipeline (PyTorch QAT → fp32 ONNX → ORT static-quantize) is load-bearing. Both the legacy ONNX exporter (
quantized::conv2d) and the new TorchDynamo exporter (Conv2dPackedParamsBase.__obj_flatten__) refuse to consumeconvert_fxoutput on PyTorch 2.11. The bridge (state-dict diff to a fresh fp32 module + ORT static-quantize) is the only path that yields a QDQ ONNX. Do NOT collapse to a single-stepconvert_fx → torch.onnx.exportuntil both PyTorch issues are fixed; re-check both exporters on each PyTorch upgrade. - State-dict transfer matches by submodule name + shape.
_copy_qat_weights_into_fp32walksfp32_statekeys, finds the same key in the FX-prepared module, copies the tensor. Tiny-AI models today have stable submodule names (entry,body.*,exit); a model architecture that uses top-levelnn.Sequentialwould break this becauseprepare_qat_fxrenames Sequential children to numeric indices. TheRuntimeError("0 tensors copied")guard catches the silent failure mode. - FX preparation runs on CPU. PyTorch 2.11's FX symbolic tracer is flaky on CUDA buffers; the trainer migrates the model to CPU before
prepare_qat_fxand back to the accelerator for the fine-tune phase. The smoke test deliberately exercises the CPU path so this stays covered. torch.ao.quantizationdeprecation will hard-fail in PyTorch 2.10. Migration target istorchao.quantization.pt2e(prepare_pt2e/convert_pt2e); the two-step pipeline is mostly pt2e-compatible — only the FX-prep call changes.- On upstream sync: no interaction with upstream. The
ai/subtree is fully fork-local. - Re-test on rebase:
python -m pytest ai/tests/test_qat_smoke.py -v
python ai/scripts/qat_train.py \
--config ai/configs/learned_filter_v1_qat.yaml \
--output /tmp/qat_smoke.int8.onnx --smoke
0074 — GPU-parity matrix CI gate (T6-8 / ADR-0214)¶
- Touched surfaces (fork-local):
scripts/ci/cross_backend_parity_gate.py(new),.github/workflows/tests-and-quality-gates.yml(newvulkan-parity-matrix-gatejob),docs/development/cross-backend-gate.md(new),docs/backends/index.md(cross-backend section),libvmaf/AGENTS.md(rebase-sensitive invariant note). - Why this matters on rebase: the CI lane and the matrix-gate script are entirely fork-local. Upstream Netflix/vmaf has no comparable gate; conflicts on rebase are restricted to the CI workflow file when upstream rearranges its own jobs. The gate's Python script lives outside
core/src/so the upstream-sync path doesn't see it. - Invariants the gate enforces:
- Per-feature absolute tolerance is declared in one place (
FEATURE_TOLERANCEinscripts/ci/cross_backend_parity_gate.py). Tightening a tolerance requires a measurement-driven follow-up ADR; loosening requires a justification ADR (CLAUDE.md §12 r1). - The legacy single-feature gate
scripts/ci/cross_backend_vif_diff.pystays for one release cycle. Sister PRs in this session add to it; the T6-8b cleanup PR deletes it once the matrix gate has soaked. - CUDA / SYCL / hardware-Vulkan are advisory until a self-hosted runner is registered. The script supports them via
--backends; flipping the CI lane to required is a follow-up wiring change, not a code change. - On upstream sync: no interaction with upstream
tests-and-quality-gates.yml(the gate job is fork-added); rebase conflicts limited to insertion-order in the workflow file. - Re-test on rebase:
cd libvmaf && meson setup build \
-Denable_cuda=false -Denable_sycl=false \
-Denable_vulkan=enabled -Denable_float=true \
--buildtype=release && ninja -C build
cd ..
python3 scripts/ci/cross_backend_parity_gate.py \
--vmaf-binary core/build/tools/vmaf \
--reference testdata/ref_576x324_48f.yuv \
--distorted testdata/dis_576x324_48f.yuv \
--width 576 --height 324 --backends cpu vulkan \
--json-out /tmp/parity.json --md-out /tmp/parity.md
0220 — SYCL feature kernels are unconditionally fp64-free (T7-17)¶
- Touches:
core/src/sycl/common.cpp(init log line),core/src/sycl/AGENTS.md(new invariant row), all SYCL feature kernels undercore/src/feature/sycl/(no diff today, but the contract pins their shape going forward). - Invariant: every SYCL feature-kernel lambda captures and operates on
float/ integer types only. Nodoubleoperand inside aparallel_forbody, nosycl::reduction<double>, nosycl::plus<double>. A single fp64 instruction in the TU's SPIR-V module causes the Level Zero runtime to reject the entire module on Intel Arc A-series and other fp64-less devices, even when the offending kernel is never submitted. Host-sidedouble(inextract/flushpost-processing, score aggregation, log10 normalisation) remains fine. Concrete patterns in tree: ADM gain limiting via int64 Q31 (gain_limit_to_q31+launch_decouple_csf<false>ininteger_adm_sycl.cpp); VIF gain limiting via fp32sycl::fmin; CIEDE / SSIM accumulators viasycl::reduction<int64_t>/sycl::plus<int64_t>. - On upstream sync: Netflix/vmaf has no SYCL backend upstream; conflicts cannot enter via
git merge. The risk is a fork-local cherry-pick (e.g. a SYCL twin of a new CUDA kernel) bringing adoubleinto a kernel lambda. Audit the lambda capture list and anysycl::reduce*calls against this invariant before merging. - Re-test on rebase:
# Build SYCL backend
meson setup build-sycl libvmaf -Denable_sycl=true CC=icx CXX=icpx
ninja -C build-sycl
# On an fp64-less device (e.g. Intel Arc A380), confirm the
# init log line is INFO-level and reads "device lacks native
# fp64 — kernels already use fp32 + int64 paths, no emulation
# overhead". The SYCL kernels must launch successfully (no
# SPIR-V module rejection from the Level Zero runtime).
build-sycl/tools/vmaf --reference testdata/ref_576x324_48f.yuv \
--distorted testdata/dis_576x324_48f.yuv \
--width 576 --height 324 --backend sycl \
--feature integer_vif --feature integer_adm \
--output /tmp/sycl-fp64less.json --json
0091 — T6-9 model registry schema + --tiny-model-verify (ADR-0211)¶
- No rebase impact: 100% fork-local surface. The registry (
model/tiny/registry.json), its JSON Schema (model/tiny/registry.schema.json), the--tiny-model-verifyCLI flag, and thevmaf_dnn_verify_signature()C entry point are entirely fork-local — none of these paths exist in upstream Netflix/vmaf. Listed here for completeness so a future/sync-upstreamrun sees the surface area was acknowledged. - Touches (additive only):
model/tiny/registry.json,model/tiny/registry.schema.json,ai/scripts/validate_model_registry.py,core/src/dnn/model_loader.{c,h}(addedvmaf_dnn_verify_signature()),core/include/libvmaf/dnn.h(public declaration),core/tools/cli_parse.{c,h}(ARG_TINY_MODEL_VERIFY+tiny_model_verifyfield),core/tools/vmaf.c(call site),core/test/dnn/test_tiny_model_verify.c,python/test/model_registry_schema_test.py,docs/ai/model-registry.md,docs/ai/inference.md,docs/ai/security.md,docs/adr/0209-...md,docs/adr/README.md(index row),CHANGELOG.md,core/src/dnn/AGENTS.md. - Invariants (rebase-relevant):
- Schema is the contract. New registry fields land in
registry.schema.jsonfirst, then inregistry.json, then in any consumers (the C-side parser, the Python validator, the MCP). Reverse order causes mismatch. schema_versionis bounded. The schema accepts only{0, 1}; bump the enum and the loader's check together when adding2.- Banned-function rule applies. The
cosigninvocation usesposix_spawnp(3p)with an explicit argv array. Do not replace withsystem(3)/popen(3)— both shell-parse the command and would re-introduce injection risk. - Bundle-file absence is fail-closed. When
sigstore_bundlepoints at a not-yet-existing file (pre-release state),vmaf_dnn_verify_signature()returns-ENOENT. The CLI surfaces this as a load failure; do not "soften" to a warning without an explicit ADR. - Re-test on rebase:
python3 ai/scripts/validate_model_registry.py
python3 -m pytest python/test/model_registry_schema_test.py -v
meson test -C build-cpu --suite=dnn
0074 — HIP (AMD ROCm) backend scaffold (T7-10)¶
- ADR: ADR-0212.
- Upstream source: fork-local. HIP backend is fork-only; Netflix/vmaf has no
libvmaf_hip.hand noenable_hipmeson option. - Touches:
core/include/libvmaf/libvmaf_hip.h(new).core/include/core/meson.build— adds theis_hip_enabledinstall gate, mirroringis_cuda_enabled/is_sycl_enabledboolean idioms.core/meson_options.txt— newenable_hipboolean option (default false).core/src/meson.build— newis_hip_enabledflag, conditionalsubdir('hip'),hip_sources+hip_depsthreaded throughlibvmaf_feature_static_lib(alongside the existing CUDA / SYCL / Vulkan aggregations) and the top-levellibrary('vmaf', ...)dependencieslist.core/src/hip/(new directory:common.{c,h},picture_hip.{c,h},dispatch_strategy.{c,h},meson.build).core/src/feature/hip/(new directory:adm_hip.c,vif_hip.c,motion_hip.c).core/test/test_hip_smoke.c(new).core/test/meson.build— registers the smoke test underif get_option('enable_hip') == true..github/workflows/libvmaf-build-matrix.yml— addsBuild — Ubuntu HIP (T7-10 scaffold)row.docs/backends/hip/overview.md(new),docs/backends/index.md(planned → scaffold row),docs/research/0033-hip-applicability.md(new),docs/adr/0212-hip-backend-scaffold.md(new),docs/adr/README.md(new index row).libvmaf/AGENTS.md— new "HIP backend scaffold contract" rebase-sensitive invariant entry.CHANGELOG.md— Unreleased § Added.- Invariants (rebase-relevant):
enable_hipis abooleanoption, not afeature. Mirrorsenable_cuda/enable_sycl; do not "harmonise" withenable_vulkan'sfeature/disabledform without an ADR amendment per ADR-0212 § "Decision".- Public C-API entry points return
-ENOSYSfor the scaffold. The smoke test core/test/test_hip_smoke.c pins this. A rebase that "succeeds" by accidentally enabling a code path (e.g. a refactor that early-returns 0 fromvmaf_hip_state_init) breaks the smoke and the runtime PR's contract baseline. hip_sourcesis added tolibvmaf_feature_static_lib, NOT directly to the top-levellibrary('vmaf', ...). The static lib is extracted into libvmaf viaobjects: [..., libvmaf_feature_static_lib.extract_all_objects(recursive: true), ...]at the bottom ofcore/src/meson.build. Addinghip_sourcesto the top library() too would double-link.hip_depsIS added to the top library()dependencies:list. The runtime PR will populatehip_depswith the realdependency('hip-lang')linkage; threading it through the top library() ensures consumers see the transitive dependency.- Header purity:
libvmaf_hip.hdoes not include<hip/hip_runtime.h>. HIP runtime types cross the public ABI asuintptr_t(matches the CUDA / Vulkan precedent; ADR-0212). Don't add<hip/...>includes to the public header during a rebase / runtime-PR bring-up. - No FFmpeg patch: the fork's
ffmpeg-patches/series does not currently consume the HIP API surface. CLAUDE §12 r14 only requires patch updates when an existing patch consumes the surface; the runtime PR (T7-10b) will add thehip_devicefilter option and the corresponding patch. - On upstream sync: zero interaction; HIP backend is fork-only.
- Re-test on rebase:
cd libvmaf
meson setup build-hip -Denable_cuda=false -Denable_sycl=false \
-Denable_hip=true
ninja -C build-hip
meson test -C build-hip test_hip_smoke
# Expect: 9/9 pass.
# Default no-HIP build still works:
meson setup build-cpu -Denable_cuda=false -Denable_sycl=false
ninja -C build-cpu
meson test -C build-cpu --suite=fast
0074 — SSIMULACRA 2 SVE2 SIMD parity (T7-38)¶
- ADR: ADR-0213.
- Touches:
core/src/feature/arm64/ssimulacra2_sve2.{c,h}(new),core/src/feature/ssimulacra2.c(dispatch table override ininit_simd_dispatch),core/src/arm/cpu.{c,h}(HWCAP2_SVE2 probe + newVMAF_ARM_CPU_FLAG_SVE2enum value),core/src/meson.build(cc.compiles probe + optionalarm64_ssimulacra2_sve2static library),core/test/test_ssimulacra2_simd.c(SVE2 picker overrides on the arm64 path + dispatch diagnostic),build-aux/aarch64-linux-gnu-sve2.ini(new cross-file pinningqemu-aarch64-static -cpu max). All paths are wholly fork-local; no upstream Netflix/vmaf code is modified. - Invariants:
- Fixed 4-lane SVE2 predicate. Every kernel uses
svwhilelt_b32(0, 4)so SIMD arithmetic order is identical to the NEON sibling regardless of the runtime vector length. This keeps the ADR-0138 / ADR-0139 / ADR-0140 byte-exact contract intact. Do NOT widen the predicate tosvptrue_b32()without a separate ADR + snapshot regen — variable-length lane reductions perturb the per-step rounding order. - NEON stays the fallback. SVE2 is purely additive; the dispatch table assigns NEON first and only overrides on
VMAF_ARM_CPU_FLAG_SVE2. A toolchain that fails thecc.compiles(... -march=armv9-a+sve2)probe leavesHAVE_SVE2unset and the legacy NEON-only build is unchanged. -ffp-contract=offmirrors the NEON sibling. Without it GCC fuses the per-lane scalar tail'sa*b+cpatterns intofmla, drifting against the SIMD path by ~1 ulp. Thearm64_ssimulacra2_sve2static library carries the flag like its NEON counterpart.- On upstream sync: no interaction with upstream —
arm64/feature TUs and thearm/cpu.{c,h}flag enum are fork-local. An upstream sync that rewritesinit_simd_dispatchincore/src/feature/ssimulacra2.cwould also need the SVE2 cases preserved. - Re-test on rebase:
meson setup build-arm64-sve2 libvmaf \
--cross-file=build-aux/aarch64-linux-gnu-sve2.ini -Denable_asm=true
ninja -C build-arm64-sve2 test/test_ssimulacra2_simd
meson test -C build-arm64-sve2 test_ssimulacra2_simd
# stderr should report `ssimulacra2 simd dispatch: NEON=1 SVE2=1`
# and 11/11 tests should pass.
0075 — enable_lcs MS-SSIM extras on CUDA + Vulkan (T7-35 / ADR-0243)¶
- Touched surfaces (fork-local):
core/src/feature/cuda/integer_ms_ssim_cuda.c(addedenable_lcstoMsSsimStateCuda+options[]+ 15 host-sidevmaf_feature_collector_appendcalls gated on the bool),core/src/feature/vulkan/ms_ssim_vulkan.c(rewroteenable_lcshelp text + addedemit_lcs_metricshelper + gated 15vmaf_feature_collector_appendcalls),scripts/ci/cross_backend_vif_diff.py scripts/ci/cross_backend_parity_gate.py(newfloat_ms_ssim_lcspseudo-feature +FEATURE_ALIASESmapplaces=4tolerance row).- Why this matters on rebase: the GPU MS-SSIM extractors are fork-local (Netflix upstream has no Vulkan or CUDA MS-SSIM kernel today). The
enable_lcssemantic and the metric names (float_ms_ssim_{l,c,s}_scale{0..4}) must match the upstream CPU reference atcore/src/feature/float_ms_ssim.c:189-221. If upstream ever renames or reorders those metrics, mirror the change on the GPU side in the same merge — public-API contract. - Invariants the contract enforces:
- Default-path output (
enable_lcs=false) stays bit-identical to the pre-T7-35 binary: only the host-side appends are gated; no kernel / shader / device-buffer changes. - Metric ordering is metric-wise (all
l_scale*first, thenc_*, thens_*) — matches the CPU emission order. places=4cross-backend tolerance per ADR-0190; enforced by the newfloat_ms_ssim_lcscell in the parity matrix gate (ADR-0214).- On upstream sync: zero interaction; the GPU twins do not exist upstream. The CPU
float_ms_ssim.cis shared with upstream butenable_lcsis upstream-stable since v3.0.0. - Re-test on rebase:
cd libvmaf && meson setup build-vulkan \
-Denable_cuda=false -Denable_sycl=false \
-Denable_vulkan=enabled -Denable_float=true \
--buildtype=release && ninja -C build-vulkan
cd ..
python3 scripts/ci/cross_backend_vif_diff.py \
--vmaf-binary core/build-vulkan/tools/vmaf \
--reference testdata/ref_576x324_48f.yuv \
--distorted testdata/dis_576x324_48f.yuv \
--width 576 --height 324 \
--feature float_ms_ssim_lcs --backend vulkan --places 4
0075 — 32-bit ADM/cpu fallbacks port (T-NEW-3)¶
- Touched surfaces (upstream-mirror):
core/src/feature/x86/adm_avx2.c,core/src/feature/x86/adm_avx512.c,core/src/x86/cpu.c. Cherry-picks of upstream8a289703(Christopher Degawa, "adm: add fallback for extract_epi64 for 32-bit") and1b6c3886("x86/cpu: remove limit of avx+ on 32-bit"). - Why this matters on rebase: trivially conflict-free with any future upstream
extract_epi64work because we land upstream's exactextract_epi64macro/inline-fn pair. The conflict surface is the fork's clang-format-100col layout inadm_avx2.c/adm_avx512.cand the_Alignas(64)LTO-correctness slot inadm_avx512.c(docs/development/known-upstream-bugs.md); both are preserved verbatim. - Invariants the port preserves:
_Alignas(64) int64_t angle_flag[16]inadm_decouple_s123_avx512stays — without it, LTO can promote the unaligned load tovmovdqa64and fault under--buildtype=release -Db_lto=true.- The
extract_epi64symbol must remain resolved on both__x86_64__(macro to_mm256_extract_epi64) and 32-bit (fallback inline). If a future upstream change inlines the helper differently, keep the conditional definition. - On upstream sync: if Netflix ships further 32-bit fallbacks (motion / psnr — not in this port), expect a parallel
extract_epi64-style helper at the top of each affected SIMD file. The fork should mirror those verbatim into the same files. - Re-test on rebase:
meson setup build-i686 libvmaf \
--cross-file=build-aux/i686-linux-gnu.ini \
-Denable_asm=false
ninja -C build-i686
meson setup build-cpu libvmaf -Denable_avx512=true
ninja -C build-cpu
meson test -C build-cpu
0076 — codec-aware FR regressor surface (T7-CODEC-AWARE / ADR-0235)¶
- Touches:
ai/src/vmaf_train/codec.py(new),ai/src/vmaf_train/models/fr_regressor.py(extended),ai/scripts/bvi_dvc_to_full_features.py,ai/scripts/extract_full_features.py. No upstream-shared paths. - Invariant:
CODEC_VOCABinai/src/vmaf_train/codec.pyis closed and order-stable — the index of each codec is the one-hot column index baked into trained ONNX. Adding a codec appends to the tuple and bumpsCODEC_VOCAB_VERSION; reordering silently invalidates every shippedfr_regressor_v2_*.onnx.FRRegressor(num_codecs=0)must remain the v1 single-input contract — flipping the default would break every existingmodel/tiny/fr_regressor_v1.onnxconsumer. - Re-test:
pytest ai/tests/test_codec_aware_fr.py -v(8 sub-tests covering vocabulary contract + alias table + back-compat). Pure fork-local addition; no upstream rebase impact for the next/sync-upstream.
0075 — feature/speed extractors (T-NEW-1, upstream port d3647c73)¶
- Touches:
core/src/feature/speed.c(new),core/src/feature/picture_copy.{c,h}(signature change — addedint channelparameter),core/src/feature/float_*.ccall sites updated to passchannel=0,core/src/feature/feature_extractor.cregistry block,core/src/feature/alias.c,core/src/meson.build,core/src/feature/vif_tools.{c,h}(helper-function port from upstream4ad6e0ea). - Upstream source: verbatim cherry-pick of Netflix/vmaf
d3647c73("feature/speed: port speed_chroma and speed_temporal extractors") with its dependency4ad6e0ea("feature/vif: port helper functions"). Both are pre-existing on Netflix master and enter the fork as part of the T7-4 audit catch-up. - Invariant:
picture_copy()now takes achannelargument — every fork-local extractor that calls it (CUDAinteger_ms_ssim, Vulkanssim/ms_ssim) passeschannel=0. If upstream later evolves the signature again (e.g. adds bit-depth or stride validation), update those fork-local call sites in lockstep. Speed extractors only register whenVMAF_FLOAT_FEATURES=1(build with-Denable_float=true). - On upstream sync: future Netflix commits in
core/src/feature/speed.capply cleanly because the file is now a verbatim mirror; conflict potential is limited to the registry block infeature_extractor.c(interleave with the fork's Vulkan / SYCL / CUDA blocks) and to any furtherpicture_copysignature evolution. - Re-test on rebase:
```bash meson setup build-cpu libvmaf -Denable_cuda=false \ -Denable_sycl=false -Denable_float=true ninja -C build-cpu meson test -C build-cpu test_speed meson test -C build-cpu # full meson suite make test-netflix-golden # 3 CPU canonical pairs
0221 — CHANGELOG + ADR-index fragment-file pattern (T7-39 / ADR-0221)¶
- What changed: the fork stopped editing
CHANGELOG.mdanddocs/adr/README.mddirectly. Both files are now rendered from fragment trees: changelog.d/<section>/<topic>.md(Keep-a-Changelog sections), plus the migration archivechangelog.d/_pre_fragment_legacy.md.docs/adr/_index_fragments/<NNNN-slug>.md, plusdocs/adr/_index_fragments/_order.txt(frozen commit-merge order manifest) anddocs/adr/_index_fragments/_header.md(table prelude). Two scripts render the consolidated outputs:scripts/release/concat-changelog-fragments.sh --check|--writescripts/docs/concat-adr-index.sh --check|--write- On upstream sync: zero interaction —
CHANGELOG.mdis a fork-local Markdown surface (Netflix upstream doesn't ship a Keep-a-Changelog file in this format), anddocs/adr/is entirely fork-local. A/sync-upstreamrun will not touch the fragment trees. - Re-test on rebase:
bash scripts/release/concat-changelog-fragments.sh --check
bash scripts/docs/concat-adr-index.sh --check
# both must exit 0; otherwise run --write and re-stage.
0077 — DISTS extractor proposal (T7-DISTS / ADR-0236)¶
- What landed: ADR-0236 (Proposed) + Research-0043 design digest ADR README index row + CHANGELOG entry.
- Rebase impact: pure fork-local proposal-stage docs; no code, no Netflix-mirror file touched, no ffmpeg-patches change, no public C-API surface change.
- Reproducer (when implementation lands as T7-DISTS):
```sh vmaf --feature dists_sq=model_path=model/tiny/dists_sq.onnx \ --reference ref.yuv --distorted dist.yuv \ --width 1920 --height 1080 --pix_fmt yuv420p
0076 — GPU-gen ULP calibration head (proposal-stage, T7-GPU-ULP-CAL / ADR-0234)¶
- What landed: ADR-0234 (Proposed), Research-0041, data-collection scaffold at
ai/scripts/collect_gpu_calibration_data.py, forward-pointer indocs/usage/cli.mdfor the future--gpu-calibratedflag. - Rebase impact: pure fork-local (proposal docs + Python script); no upstream Netflix/vmaf code touched, no public C-API changes, no ffmpeg-patches changes.
- Reproducer:
```sh python3 ai/scripts/collect_gpu_calibration_data.py --smoke
0095 — Per-backend GPU kernel scaffolding templates (CUDA + Vulkan, ADR-0246)¶
- ADR: ADR-0246.
- Touches:
core/src/cuda/kernel_template.h(new, header-only).core/src/vulkan/kernel_template.h(new, header-only).core/src/cuda/AGENTS.md(new invariant row + dir listing).core/src/vulkan/AGENTS.md(new file).docs/backends/kernel-scaffolding.md(new).docs/adr/0246-gpu-kernel-template.md(new).CHANGELOG.md,docs/adr/README.md. All paths are wholly fork-local. Upstream Netflix/vmaf has no Vulkan backend at all today and the CUDA backend uses different per-kernel scaffolding shapes; nothing here can collide on a pure upstream sync.- Invariants:
- Templates are unused at PR-merge time.
kernel_template.hin bothcore/src/cuda/andcore/src/vulkan/lands with zero call-sites. Each future kernel migration is its own gated PR (places=4cross-backend-diff per ADR-0214). Do not bulk-port existing kernels onto the templates in a single sync — that would short-circuit the per-kernel gate. - Per-backend, not cross-backend. Resist the urge to merge the two templates into a unified
gpu/kernel_template.h. CUDA async-stream + event vs Vulkan command-buffer + fence + descriptor-pool share no concrete shape; a unified API would be lowest-common-denominator. - Helper functions, not macros. The header bodies are
static inlinefunctions for cuda-gdb / Nsight / RenderDoc step-debugging. TheCHECK_CUDA_GOTO/CHECK_CUDA_RETURNmacros incuda_helper.cuhstay where they pay off (textualgoto label), and the templates use them internally. - On upstream sync: no interaction with upstream paths. An upstream sync that touches
core/src/cuda/common.horpicture_cuda.hmay shift the helper signatures the template consumes (vmaf_cuda_buffer_alloc,vmaf_cuda_picture_get_stream, …); update the template if so. - Re-test on rebase:
```bash # CUDA build (configure inside libvmaf/ — see CLAUDE.md §2 note). meson setup core/build-cuda libvmaf \ -Denable_cuda=true -Denable_nvcc=true \ -Denable_vulkan=disabled -Denable_sycl=false ninja -C core/build-cuda meson test -C core/build-cuda
# Vulkan build. meson setup core/build-vulkan libvmaf \ -Denable_vulkan=enabled -Denable_cuda=false -Denable_sycl=false ninja -C core/build-vulkan meson test -C core/build-vulkan
0222 — vmaf-perShot per-shot CRF predictor sidecar (T6-3b)¶
- Touches:
core/tools/meson.build(new executable + test wiring),core/tools/vmaf_per_shot.c(new file — fork-local, no upstream sibling),core/tools/test/meson.build(test row),core/tools/test/test_vmaf_per_shot.sh(new smoke test),core/tools/AGENTS.md(sidecar invariants),docs/usage/cli.md(cross-link),docs/usage/vmaf-perShot.md(new user doc),docs/ai/roadmap.md(T6-3b row update). - Invariant: the sidecar must stay standalone — it does not link the libvmaf metric path. Any upstream patch that tries to fold per-shot CRF prediction into
vmaf_score_*would collapse the encoder-hint vs. quality-score separation recorded in roadmap §2.4 and ADR-0222 §Decision. The CSV / JSON column set (shot_id,start_frame,end_frame,frames,mean_complexity,mean_motion,predicted_crf) is the public schema; downstream encoders consume it directly. - Conflict expectation on
/sync-upstream: low. Upstream Netflix has no per-shot CRF predictor in tree, so there is no natural collision point —tools/meson.buildis the only mutually-edited file and the newexecutable('vmaf-perShot', …)block is appended aftervmaf_bench_deps, well clear of upstream's likely additions. - Reproducer:
```bash meson setup build libvmaf -Denable_cuda=false -Denable_sycl=false \ -Denable_vulkan=disabled ninja -C build meson test -C build test_vmaf_per_shot --print-errorlogs ./build/tools/vmaf-perShot \ --reference testdata/ref_576x324_48f.yuv \ --width 576 --height 324 --pixel_format 420 --bitdepth 8 \ --output /tmp/plan.csv cat /tmp/plan.csv
0075 — vmaf-roi sidecar binary (T6-2b / ADR-0247)¶
- Touches:
core/tools/meson.build— adds thevmaf_roiexecutable target (after the existingvmaftarget, beforevmaf_bench). Append-only; no upstream-shared lines moved or removed.core/test/meson.build— adds thetest_vmaf_roiexecutable +test()registration. Append-only.core/tools/vmaf_roi.c— wholly new, fork-local.core/tools/vmaf_roi_core.h— wholly new, fork-local.core/test/test_vmaf_roi.c— wholly new, fork-local.- Invariant: the
vmaf-roisidecar emits two byte-exact formats that downstream encoder drivers (x265--qpfile, SVT-AV1--roi-map-file) will hard-depend on: - x265 ASCII grid — two
#-prefixed header lines (# vmaf-roi qpfile (x265, --qpfile-style)and# frame=N ctu=S cols=C rows=R strength=F.FFF), space-separated signed integers, one row per CTU row,\nterminator. - SVT-AV1 raw binary — exactly
cols * rowsbytes ofint8_t, row-major, no header. - QP-offset clamp —
+-12(VMAF_ROI_CORE_QP_OFFSET_MAX). - Reduction — per-CTU mean (not max). Switching to max or a percentile changes every downstream encoder result and requires its own ADR.
- Pure helpers in
vmaf_roi_core.h— the per-CTU mean reducer and saliency-to-QP mapper arestatic inlinein a header sotest_vmaf_roicompiles them without dragging the libvmaf link surface in. Moving them into a.cTU breaks the test wiring. - On upstream sync: no interaction with upstream —
tools/is a fork-local surface from upstream's perspective (upstream shipsvmaf.conly). An upstream sync that rewritescore/tools/meson.buildshould preserve thevmaf_roiexecutable block. - Re-test on rebase:
```bash meson setup build-cpu libvmaf \ -Denable_cuda=false -Denable_sycl=false -Denable_tools=true ninja -C build-cpu tools/vmaf_roi test/test_vmaf_roi meson test -C build-cpu test_vmaf_roi ./build-cpu/tools/vmaf_roi \ --reference testdata/ref_576x324_48f.yuv \ --width 576 --height 324 --frame 0 --output - \ --encoder x265 --ctu-size 64 --strength 6.0 | head -3 # First two lines are the # comment header; row 1 of the grid # should be "4 2 1 -1 -1 -1 1 2 4" (placeholder radial map).
0219 — motion3 GPU coverage on Vulkan + CUDA + SYCL (T3-15(c) / ADR-0219)¶
- What changed: The
motionGPU twins (core/src/feature/vulkan/motion_vulkan.c,core/src/feature/cuda/integer_motion_cuda.c,core/src/feature/sycl/integer_motion_sycl.cpp) now emitVMAF_integer_feature_motion3_scorein 3-frame window mode (default). Cross-backend gates extended (scripts/ci/cross_backend_*.pyFEATURE_METRICS["motion"]). - Invariants:
motion3 = host-side scalar post-process of motion2. No device-side state changes; motion3 is computed on the host inextract()/collect()/flush()after the existing SAD reduction. The post-processing function (motion3_postprocess_*) mirrors CPUinteger_motion.clines 510-560 byte-for-byte:clip(motion_blend(motion2 * fps_weight, blend_factor, blend_offset), max_val)with optional moving-average against the unaveraged prior blended value.motion_five_frame_window=truereturns-ENOTSUPatinit()on all three GPU backends. The 5-deep blur ring + second SAD-pair dispatch remain deferred. Do NOT silently fall back to the 3-frame path when the user enables the flag — fail loud per CERT C / CLAUDE.md §12 r4.- CPU motion3 algorithm is the source of truth. Any port of an upstream Netflix change to
integer_motion.cthat touchesmotion_blend(...), themotion_max_valclip, or the moving-average rule MUST be mirrored inmotion3_postprocess_*across all three GPU files in the same PR. The cross-backend gate atplaces=4will catch drift, but only after a full GPU run. - On upstream sync: Pure fork-local additions to GPU TUs. Upstream Netflix has no GPU motion extractor. The
motion_blend_tools.hheader is upstream-mirrored — if a sync rewrites themotion_blend()formula, regenerate the GPU snapshot and re-run the cross-backend gate. - Re-test on rebase:
```bash # CPU sanity (motion3 emission unchanged) ./core/build/tools/vmaf \ --reference python/test/resource/yuv/src01_hrc00_576x324.yuv \ --distorted python/test/resource/yuv/src01_hrc01_576x324.yuv \ --width 576 --height 324 --pixel_format 420 --bitdepth 8 \ --feature motion --output /tmp/motion.json --json python -c "import json; d=json.load(open('/tmp/motion.json')); \ print('motion3 frames:', sum(1 for f in d['frames'] \ if 'integer_motion3' in f.get('metrics', {})))" # Expect 49 (one motion3 per frame).
# Cross-backend gate (Vulkan/lavapipe lane works on every host): python scripts/ci/cross_backend_vif_diff.py \ --feature motion --backend vulkan \ --ref python/test/resource/yuv/src01_hrc00_576x324.yuv \ --dis python/test/resource/yuv/src01_hrc01_576x324.yuv \ --width 576 --height 324 --bitdepth 8 \ --vmaf-bin core/build/tools/vmaf # Expect: integer_motion / integer_motion2 / integer_motion3 all OK at places=4.
0216 — vmaf_tiny_v2 (Phase-3-validated tiny VMAF MLP)¶
- Touches:
model/tiny/registry.json,model/tiny/vmaf_tiny_v2.{onnx,json},ai/scripts/{train,export,validate}_vmaf_tiny_v2.py,ai/AGENTS.md,core/test/dnn/{test_vmaf_tiny_v2.py,meson.build},docs/ai/{models/vmaf_tiny_v2.md,inference.md,roadmap.md},docs/adr/{0244-vmaf-tiny-v2.md,README.md},CHANGELOG.md. All paths are wholly fork-local; no upstream Netflix/vmaf code is modified. - Invariants:
- Bundled scaler stats are part of the trust root. The shipped ONNX bakes
(input - mean) / stdas ConstantSub+Divnodes that run before the MLP. Re-exporting must go throughai/scripts/export_vmaf_tiny_v2.py, which pullsmean/stdfrom the trainer checkpoint and writes them as graph initialisers. Adding an out-of-band scaler step at runtime (e.g., a sidecar JSON consumed by the loader) is forbidden without a follow-up ADR — it splits the trust root and invalidates the registry sha256 contract. - Feature column order is fixed. The graph reads
(adm2, vif_scale0, vif_scale1, vif_scale2, vif_scale3, motion2)in exactly this order; reordering breaks the bundledmean/stdconstants. Any change to the feature set requires a fresh Phase-3 chain (Research-0027 → 0028 → 0029 → 0030). - opset 17. Matches the sister tiny-AI models (
learned_filter_v1,nr_metric_v1,fastdvdnet_pre) and the ORT op-allowlist baseline. Upgrading requires re-validating theSub/Div/Gemm/Relu/Squeezeops againstop_allowlist.c. - On upstream sync: zero interaction. Netflix/vmaf has no equivalent surface; an upstream sync that touches
core/src/dnn/(op-allowlist or model-loader changes) needs to preserveSub/Div/Gemm/Relu/Squeezein the allowlist for opset 17. - Re-test on rebase:
```bash bash core/test/dnn/test_registry.sh python3 core/test/dnn/test_vmaf_tiny_v2.py python3 ai/scripts/validate_vmaf_tiny_v2.py \ --onnx model/tiny/vmaf_tiny_v2.onnx \ --parquet runs/full_features_netflix.parquet \ --rows 100 --min-plcc 0.97 meson test -C build-cpu --suite=dnn
0094 — Tiny-AI extractor template (ADR-0250)¶
- Touches:
core/src/dnn/tiny_extractor_template.h(new),core/src/feature/feature_lpips.c,core/src/feature/fastdvdnet_pre.c,core/src/dnn/AGENTS.md,docs/ai/extractor-template.md(new),docs/adr/0250-tiny-ai-extractor-template.md(new). - Invariants:
- Helper signatures are wire-format-stable.
vmaf_tiny_ai_resolve_model_path(name, option, env_var)andvmaf_tiny_ai_open_session(name, path, &out)produce the user-facing log lines<name>: no model path …and<name>: vmaf_dnn_session_open(<path>) failed: <rc>— downstream tooling greps these. Don't rename or reorder the parameters without bumping every extractor + the recipe doc. - YUV→RGB is bit-exact. The shared
vmaf_tiny_ai_yuv8_to_rgb8_planesis a literal move of the pre-existingfeature_lpips.cbody (BT.709 limited-range, nearest-neighbour chroma upsample). LPIPS / saliency / future colour-sensitive tiny-AI scores depend on byte-exact equality with the prior ad-hoc copies. Any change to the conversion constants or the rounding rule needs a separate ADR + a coordinated snapshot regen —model/tiny/weights aren't re-trained against new colour math casually. - Option-table macro is plain text substitution. The
VMAF_TINY_AI_MODEL_PATH_OPTION(state_t, help)macro emits a single struct literal — no control flow, no recursion, no variadic shenanigans (Power-of-10 rule 1 / rule 9). Don't extend it into a multi-option emitter without a fresh ADR. - On upstream sync: zero interaction with upstream —
feature_lpips.candfastdvdnet_pre.care fork-only files, and the newdnn/tiny_extractor_template.hlives entirely under fork-introducedcore/src/dnn/. An upstream sync that rewrites unrelatedfeature_*.cfiles won't conflict. - Re-test on rebase:
cd libvmaf
meson setup build-cpu -Denable_cuda=false -Denable_sycl=false
ninja -C build-cpu
meson test -C build-cpu --suite=dnn
meson test -C build-cpu test_lpips test_fastdvdnet_pre
# All 10 dnn-suite + both extractor tests must pass.
0095 — Vulkan ring-depth tunable (ADR-0251 follow-up #3)¶
- PR: feat/t7-29-followup3-ring-tunable.
- What rebases need to know:
VmafVulkanConfigurationgrew an additiveunsigned max_outstanding_framesfield. Existing zero-initialised configs continue to receive the canonical default (0 → VMAF_VULKAN_RING_DEFAULT == 4). The clamp helpervmaf_vulkan_clamp_ring_sizemoved fromimport.c(file-local static) tovulkan_internal.h(static inline) sostate_initandlazy_alloc_ringshare one definition; an upstream sync that re-introduces the static inimport.cwould shadow the header helper — drop the duplicate, keep the inline. - New public symbol:
vmaf_vulkan_state_max_outstanding_frames(const VmafVulkanState *)— read-side accessor for the clamped value. Pure additive surface; no upstream collision. - On upstream sync: zero interaction. The ring is wholly fork-introduced (ADR-0251); upstream Netflix has no Vulkan backend.
- Re-test on rebase:
```bash meson setup build libvmaf -Denable_cuda=false -Denable_sycl=false \ -Denable_vulkan=enabled ninja -C build meson test -C build test_vulkan_async_pending_fence # All 8 cases must pass: 4 v2-contract + 4 ring-tunable.
0096 — tools/vmaf-tune/ automation umbrella spec (ADR-0237 / Research-0044)¶
- PR: feat/vmaf-tune-spec.
- What rebases need to know: this PR ships only an umbrella ADR research digest under
docs/. No tracked source code, notools/vmaf-tune/directory yet, no Meson changes. An upstream sync touching ffmpeg-patches orlibvmaf/cannot collide with this PR. - On upstream sync: zero interaction. Spec-only PR.
- Re-test on rebase:
# No build/test impact — verify the docs render and links are alive:
ls docs/adr/0237-quality-aware-encode-automation.md \
docs/research/0044-quality-aware-encode-automation.md
grep -c '\[ADR-0237\]' docs/adr/README.md
0097 — test_speed gated on enable_float (fix default-build failure)¶
- PR: fix/test-speed-chroma-registration.
- What rebases need to know:
core/test/meson.buildnow wraps thetest_speedexecutable +test()registration inif get_option('enable_float'). Thespeed_chroma/speed_temporalextractors live inspeed.c, which is only compiled whenenable_float=true(the entries infeature_extractor.care wrapped in#if VMAF_FLOAT_FEATURES), so the test'svmaf_get_feature_extractor_by_name("speed_chroma")returned NULL on a default build (enable_float=false). - On upstream sync: zero interaction.
test_speed.cwas added fork-side via the Netflix port commitd3647c73. The gating pattern matchestest_vulkan_*(if get_option('enable_vulkan').enabled()). - Re-test on rebase:
# default (enable_float=false): test_speed must NOT be in the suite
meson setup build libvmaf -Denable_cuda=false -Denable_sycl=false --reconfigure
ninja -C build
meson test -C build # expect: NO test_speed in the run
# CI shape (enable_float=true): test_speed must run + pass
meson setup build libvmaf -Denable_float=true --reconfigure
ninja -C build
meson test -C build test_speed # expect: 5/5 pass
0098 — Vulkan picture preallocation surface (ADR-0238)¶
- PR: feat/vulkan-picture-preallocation.
- What rebases need to know: ABI grows additively. New public surface in
core/include/libvmaf/libvmaf_vulkan.h:enum VmafVulkanPicturePreallocationMethod,VmafVulkanPictureConfiguration,vmaf_vulkan_preallocate_pictures,vmaf_vulkan_picture_fetch. New enumeratorVMAF_PICTURE_BUFFER_TYPE_VULKAN_DEVICEincore/src/picture.h::VmafPictureBufferType. New TUcore/src/vulkan/picture_vulkan_pool.c(~180 LOC); registered incore/src/vulkan/meson.build. Fork-internal accessorvmaf_vulkan_state_context()(declared invulkan_internal.h) exposes the imported state's VkInstance/VkDevice to the pool — used only bylibvmaf.c::vmaf_vulkan_preallocate_pictures. VmafContextfield added:vmaf->vulkan.poolnext tovmaf->vulkan.state. Thevmaf_close()teardown closes the pool before clearing the state pointer (matches SYCL).- On upstream sync: zero interaction. Vulkan backend is fork-only; upstream Netflix has no Vulkan integration.
- Re-test on rebase:
```bash meson setup build libvmaf -Denable_cuda=false -Denable_sycl=false \ -Denable_vulkan=enabled ninja -C build meson test -C build test_vulkan_pic_preallocation # All 6 cases must pass under ASan/UBSan: # test_method_none_is_a_no_op # test_method_host_allocates_round_robins # test_method_device_allocates_round_robins # test_fetch_without_preallocate_falls_back # test_unknown_method_rejected # test_null_args_rejected
0099 — feature_mobilesal.c + transnet_v2.c migrated to tiny_extractor_template.h¶
- PR: refactor/migrate-ai-to-template.
- What rebases need to know:
feature_mobilesal.candtransnet_v2.cpreviously open-coded the model-path resolution (getenv+ log block), the YUV→RGB kernel (mobilesal only), thevmaf_dnn_session_open+ log boilerplate, and theVmafOption[].model_pathrow. They now use the helpers fromdnn/tiny_extractor_template.h(PR #251) — the same templatefeature_lpips.candfastdvdnet_pre.calready consume. Net −98 LOC of identical boilerplate. - Behavior preserved: bit-exact YUV→RGB conversion (mobilesal used the literal copy of
feature_lpips.c's body that the template hoisted), identical error-log strings, identical option-table flag/type/offset shape. The migratedmobilesal_optionsmacro expands to the same struct literal the hand-rolled version produced. - On upstream sync: zero interaction. Both files are fork-introduced; upstream Netflix has neither extractor.
0100 — cuda/ring_buffer.{c,h} → gpu_picture_pool.{c,h} (ADR-0239)¶
- PR: refactor/gpu-picture-pool-extract.
- What rebases need to know:
core/src/cuda/ring_buffer.candring_buffer.hare removed. The same callback-based round-robin pool lives atcore/src/gpu_picture_pool.{c,h}under renamed symbols (VmafRingBuffer→VmafGpuPicturePool,vmaf_ring_buffer_*→vmaf_gpu_picture_pool_*,_fetch_next_picture→_fetch). All call sites inlibvmaf.cmigrated.core/test/test_ring_buffer.crenamed totest_gpu_picture_pool.cwith the corresponding meson update. - Netflix-upstream interaction: minimal — Netflix's
cuda/ring_buffer.{c,h}last touched in commitcb1d49c6. An upstream sync that resurrects the old names should be redirected to the new ones; the file move is purely fork-local. Netflix#1300mutex-destroy-order fix preserved (ADR-0157) — moved verbatim to the new file; the fix remains attached tovmaf_gpu_picture_pool_close.- SYCL pool migration:
vmaf_sycl_picture_pool_*keeps its public-internal API but now delegates to the generic pool. The SYCL wrapper struct (VmafSyclPicturePool) just owns theVmafSyclCookiestorage.std::mutexdrops out. - Vulkan pool migration: bundled into this PR after #264 merged.
picture_vulkan_pool.crewrites as a thin wrapper around the generic pool — wrapper struct owns per-pool state for the alloc/free callbacks; the generic pool owns the round-robin slots / mutex / unwind. Same pattern as the SYCL migration above. - Re-test on rebase:
meson setup build libvmaf -Denable_cuda=false -Denable_sycl=false
ninja -C build
meson test -C build --suite=dnn
meson test -C build test_lpips test_mobilesal test_transnet_v2 test_fastdvdnet_pre
# All 11 dnn-suite + 4 extractor smoke tests must pass.
meson test -C build # 47/47 pass under ASan/UBSan
# CUDA build (CI-only; pre-existing local nvcc include-path quirk):
meson setup build-cuda libvmaf -Denable_cuda=true
ninja -C build-cuda
meson test -C build-cuda test_gpu_picture_pool
# SYCL build:
meson setup build-sycl libvmaf -Denable_sycl=true
ninja -C build-sycl
meson test -C build-sycl
0104 — psnr_vulkan.c migrated to vulkan/kernel_template.h¶
- PR: refactor/migrate-psnr-vulkan-to-template.
- What rebases need to know:
vulkan/kernel_template.h(410 LOC, ADR-0246, PR #251) shipped with zero consumers. Its docstring designatedpsnr_vulkan.cas the reference implementation. This PR lands the migration as the first consumer of the Vulkan template — paired with PR #269 (the first CUDA template consumer). The 5 long-lived pipeline objects (descriptor-set layout, pipeline layout, shader module, compute pipeline, descriptor pool) collapse from individual struct fields to oneVmafVulkanKernelPipeline plbundle.create_pipeline()(~104 LOC) collapses to a singlevmaf_vulkan_kernel_pipeline_create()call (~30 LOC) — the template owns the descriptor-set layout creation, pipeline layout, shader module, compute pipeline, and descriptor-pool sizing.close_fex()'svkDeviceWaitIdle+ 5×vkDestroy*sweep collapses to onevmaf_vulkan_kernel_pipeline_destroy()call. - Net LOC delta: −55 LOC on
psnr_vulkan.cdirectly. Unlike the CUDA template (where helper-call boilerplate roughly matches the inline savings), the Vulkan template's pipeline creation is dramatic enough that even the first consumer wins. - Bit-exactness gates: spec-constants, push-constant struct, shader bytecode, dispatch grid math, and host-side reduction are byte-identical to the prior implementation. The template only owns descriptor-set layout / pipeline layout / shader module / compute pipeline creation / descriptor pool sizing — none of which affects the kernel's mathematical behaviour. Cross-backend parity gate (places=4) re-runs unchanged.
- On upstream sync: zero interaction.
psnr_vulkan.cis fork-introduced (T7-23 / ADR-0182 / ADR-0216). - Re-test on rebase:
meson setup build libvmaf -Denable_cuda=false -Denable_sycl=false \
-Denable_vulkan=enabled
ninja -C build
meson test -C build # 50/50 pass on lavapipe
# Cross-backend parity gate (places=4):
python scripts/ci/cross_backend_parity_gate.py --feature psnr_y --places 4
0105 — moment_vulkan.c + ciede_vulkan.c migrated to vulkan/kernel_template.h¶
- PR: refactor/migrate-motion-vulkan-to-template (note: the branch name reflects the original intent; motion's two-pipeline shape didn't fit the template's single-pipeline contract, so this PR migrates moment + ciede instead).
- What rebases need to know: second + third consumers of
vulkan/kernel_template.h(after PR #270 = psnr_vulkan, the first consumer). Both files follow the identical migration pattern: - Replace 5 individual pipeline-object fields (
dsl,pipeline_layout,shader,pipeline,desc_pool) with oneVmafVulkanKernelPipeline plbundle. - Replace ~100 LOC of
create_pipeline()body (descriptor-set layout + pipeline layout + shader module + compute pipeline + descriptor pool boilerplate) with a singlevmaf_vulkan_kernel_pipeline_create()call. - Replace
close_fex()'svkDeviceWaitIdle+ 5×vkDestroy*sweep with onevmaf_vulkan_kernel_pipeline_destroy()call. - Per-file LOC deltas:
moment_vulkan.c: −60 LOC (450 → 390).ciede_vulkan.c: −59 LOC (536 → 477).- Net: −119 LOC.
- Bit-exactness preserved: spec-constants (width/height/bpc/ subgroup_size identical across both), push-constant structs (
MomentPushConsts,CiedePushConsts), shader bytecodes (moment_spv,ciede_spv), dispatch grid math, and host-side reductions are byte-identical to the prior implementation. Cross-backend parity gates (places=4 for moment integer reduce; places=2 for ciede transcendentals per ADR-0187) re-run unchanged. motion_vulkan.cdeferred: motion uses two pipelines (first frame vs subsequent) sharing one DSL + layout + shader + pool. The template's current shape produces one pipeline per descriptor; splitting motion across twoVmafVulkanKernelPipelineinstances would duplicate the shared objects. Tracked as a follow-up template extension (multi-pipeline support).- On upstream sync: zero interaction. Both files are fork-introduced (T7-23 / ADR-0182 / ADR-0187).
- Re-test on rebase:
```bash meson setup build libvmaf -Denable_cuda=false -Denable_sycl=false \ -Denable_vulkan=enabled ninja -C build meson test -C build # 50/50 pass on lavapipe (under ASan/UBSan) python scripts/ci/cross_backend_parity_gate.py --feature float_moment_ref1st --places 4 python scripts/ci/cross_backend_parity_gate.py --feature ciede2000 --places 2
0101 — GPU backend pattern doc (ADR-0240)¶
- PR: docs/gpu-backend-template.
- What rebases need to know: doc-only PR. Adds
docs/development/gpu-backend-template.md(recipe new GPU backends follow) andcore/include/libvmaf/AGENTS.md(public-headers-tree invariant note). No source code, no meson changes, no ABI impact. - On upstream sync: zero interaction. Both files are fork-introduced.
- Re-test on rebase:
```bash # Doc-only — verify links resolve: test -f docs/development/gpu-backend-template.md test -f core/include/libvmaf/AGENTS.md grep -c 'gpu-backend-template' core/include/libvmaf/AGENTS.md
0102 — Tiny-AI test registration macro (tiny_ai_test_template.h)¶
- PR: refactor/test-registration-macro.
- What rebases need to know: new
core/test/tiny_ai_test_template.hemits the four standard registration tests (<name>_is_registered,<name>_provides_primary_feature,<name>_options_table_well_formed,<name>_init_rejects_missing_model) via theVMAF_TINY_AI_DEFINE_REGISTRATION_TESTS(ext, feat, env, prefix)macro. The four per-extractor test files (test_lpips.c,test_mobilesal.c,test_transnet_v2.c,test_fastdvdnet_pre.c) shrank from ~140 LOC each to ~20-50 LOC. Net −286 LOC. Behavior bit-exact preserved (same assertions, same env-var save/restore dance, same setenv shim for MSVCRT). TransNet V2 keeps two extractor-specific extra tests (binary-flag round-trip + provided_features list-termination) that the macro doesn't cover. - On upstream sync: zero interaction. The four test files are fork-introduced (per ADR-0042 / ADR-0168 / ADR-0220 / ADR-0223 / ADR-0215).
- Re-test on rebase:
```bash meson setup build libvmaf -Denable_cuda=false -Denable_sycl=false ninja -C build meson test -C build test_lpips test_mobilesal test_transnet_v2 test_fastdvdnet_pre # 4/4 binaries pass; 18 individual tests total (4x4 standard + 2 # TransNet V2 extras).
0103 — integer_psnr_cuda.c migrated to cuda/kernel_template.h¶
- PR: refactor/migrate-psnr-cuda-to-template.
- What rebases need to know:
cuda/kernel_template.hshipped with no consumers in PR #251 (ADR-0246). This PR migrates the first consumer (integer_psnr_cuda.c) — the file the template's own docstring explicitly designated as the reference. TheCUstream + CUevent + CUeventtriple and the(VmafCudaBuffer device, void *host_pinned, size_t bytes)readback pair are now dispensed by the template helpers (vmaf_cuda_kernel_lifecycle_init/_close,vmaf_cuda_kernel_readback_alloc/_free,vmaf_cuda_kernel_submit_pre_launch,vmaf_cuda_kernel_collect_wait) instead of being open-coded.PsnrStateCudashrinks: replaces three fields (event+finished+str) with oneVmafCudaKernelLifecyclereplaces (sse+sse_host) with oneVmafCudaKernelReadback. - Net LOC delta: +8 LOC on
integer_psnr_cuda.calone — the helpers add per-call boilerplate. The dedup win materialises as more CUDA feature kernels (motion / moment / ssim / vif / adm) migrate one-at-a-time in follow-up PRs. Each subsequent migration saves ~15 LOC. - Bit-exactness gates: kernel launch + reduction logic unchanged. The migration only touches state-management boilerplate around the kernel; the SSE accumulator math, the per-bpc kernel function lookup, the host-side
log10score formula, and the dispatch grid-dim calculation are byte-identical to the prior implementation. Netflix golden gate + CPU/CUDA cross-backend parity gate (places=4) re-run unchanged. - On upstream sync: zero interaction.
integer_psnr_cuda.cis fork-introduced (T7-23 / ADR-0182). - Re-test on rebase:
meson setup build libvmaf -Denable_cuda=true
ninja -C build
meson test -C build # CUDA test suite must pass
# Cross-backend parity gate:
python scripts/ci/cross_backend_parity_gate.py --feature psnr_y --places 4
0125 — Vulkan submit-side template + fence pool + descriptor pre-alloc bundle (ADR-0256)¶
- Touches:
core/src/vulkan/kernel_template.h— fork-local. Output landing inruns/phase_a/is gitignored — rerun the script to reproduce.VmafVulkanKernelSubmitPoolstruct +_create/_destroy/_acquirehelpers +vmaf_vulkan_kernel_descriptor_sets_allochelper. Upstream has no Vulkan backend — no merge surface.core/src/feature/vulkan/{psnr_hvs,vif,float_vif,float_adm}_vulkan.c— fork-local kernel TUs, also no upstream peer.- Invariant: the four migrated kernels keep all per-frame
VkFence+VkCommandBuffer+VkDescriptorSetresources alive across frames in the pool. Pre-bound descriptor sets rely on the kernel'sVmafVulkanBuffer *handles being init-time stable (allocated ininit(), freed only inclose_fex).vmaf_vulkan_kernel_pipeline_destroydestroys the descriptor pool — pre-allocated sets are released implicitly via the pool; callers must NOT callvkFreeDescriptorSetson them. - Re-test on rebase:
meson setup build libvmaf -Denable_vulkan=enabled
ninja -C build
VK_ICD_FILENAMES=/usr/share/vulkan/icd.d/nvidia_icd.json \
meson test -C build test_vulkan_smoke \
test_vulkan_async_pending_fence \
test_vulkan_pic_preallocation
python scripts/ci/cross_backend_vif_diff.py \
--vmaf-binary build/tools/vmaf \
--reference testdata/ref_576x324_48f.yuv \
--distorted testdata/dis_576x324_48f.yuv \
--width 576 --height 324 --pixel-format 420 --bitdepth 8 \
--feature vif --backend vulkan --places 4
python scripts/ci/cross_backend_vif_diff.py \
--vmaf-binary build/tools/vmaf \
--reference testdata/ref_576x324_48f.yuv \
--distorted testdata/dis_576x324_48f.yuv \
--width 576 --height 324 --pixel-format 420 --bitdepth 8 \
--feature adm --backend vulkan --places 4
0107 — psnr_hvs_cuda async upload + persistent pinned staging (T-GPU-OPT-2/3)¶
- Touches:
core/src/feature/cuda/integer_psnr_hvs_cuda.c— only consumer; fork-local from inception (T7-23 / ADR-0188 / ADR-0191). State addsupload_str(dedicated H2D stream),upload_done(cross-stream completion event), and per-plane persistent pinnedh_uint_ref[3]/h_uint_dist[3]staging buffers allocated once ininit_fex_cuda. The per-call helperupload_plane_cudais split intoissue_d2h_plane(pic-stream D2H),convert_plane(CPU normalise), andissue_h2d_plane(upload-stream H2D).submit_fex_cudaruns the three phases explicitly and recordsupload_doneafter the last H2D, thencuStreamWaitEvents onlc.strbefore kernel launches.core/src/cuda/AGENTS.md— adds a rebase-sensitive invariant entry under §Rebase-sensitive invariants documenting the three-phase flow + persistent staging contract.- Invariant: the pinned
h_uint_*andh_ref/h_distbuffers are never freed and re-allocated mid-stream; the H2Ds must run onupload_str(not onlc.str) so thecuStreamWaitEventcross-stream link is meaningful; theupload_doneevent is recorded after the last H2D for the current frame and waited on once before the first kernel launch of that frame. CUDA graph capture (future T-GPU-OPT-N) depends on the no-per-frame-alloc invariant; collapsing the three-phase split or re-introducing per-framevmaf_cuda_buffer_host_alloccalls breaks that follow-up. Bit-exactness gate isplaces=3forpsnr_hvs_y / cb / crand the combinedpsnr_hvs(matches the existing matrix; notplaces=4). - On upstream sync: zero interaction.
integer_psnr_hvs_cuda.cis fork-introduced (T7-23 / ADR-0188 / ADR-0191). - Re-test on rebase:
meson setup build libvmaf -Denable_cuda=true -Denable_sycl=false
ninja -C build
meson test -C build
python3 scripts/ci/cross_backend_vif_diff.py \
--vmaf-binary core/build/tools/vmaf \
--reference python/test/resource/yuv/src01_hrc00_576x324.yuv \
--distorted python/test/resource/yuv/src01_hrc01_576x324.yuv \
--width 576 --height 324 --pixel-format 420 --bitdepth 8 \
--feature psnr_hvs --backend cuda --places 3
0227 — output.c writer-format unit tests (R3 of coverage-gap-2026-05-02)¶
- Touches:
core/test/test_output.c(new) — exercises the four writers incore/src/output.c(XML / JSON / CSV / SUB) end-to-end viatmpfile()-backed sinks and a syntheticVmafFeatureCollector. Pure test-only; no production code change.core/test/meson.build— registerstest_outputnext totest_feature_collector(mirrors that test's wiring:link_with: libvmaf+ libsvm objects + log/predict/metadata helpers).- Invariant: the test pulls
libvmaf.candoutput.cin via#include "*.c"(mirroring the precedent intest_feature_collector.c) so the per-translation-unit.gcnolands in the test build dir and gcovr aggregates output.c's coverage. The mu-test framework macro (mu_assert) deliberately early-returns from eachstatic char *test_*()body — that's why every test body tripsclang-analyzer-unix.Malloc"potential leak" notes (cleanup runs only on the success-tail path). This pattern is shared across everycore/test/test_*.cfile and is load- bearing (per ADR-0141 NOLINT carve-out): replacing it with goto- cleanup would obscure the per-assertion failure message. - On upstream sync: zero interaction.
output.cis upstream- mirrored, but this PR doesn't touch it. The test only depends on the four public function signatures (vmaf_write_output_{xml, json,csv,sub}); if Netflix renames or reorders those, the test fails to compile and the rebase author updates it then. - Re-test on rebase:
cd libvmaf && meson setup build -Denable_cuda=false -Denable_sycl=false
ninja -C build && ./build/test/test_output
0126 — OSSF Scorecard policy (ADR-0263)¶
- Touches:
.github/workflows/scorecard.yml(line 45 — thegithub/codeql-action/upload-sarif@<sha>pin). The rest of the policy is doc-only (docs/adr/0263-*.md,docs/research/0053-*.md,changelog.d/security/). Upstream Netflix/vmaf does not ship a Scorecard workflow, so the path itself is fork-introduced and won't conflict. - Invariant: the
upload-sarifSHA must point to a commit that currently exists ingithub/codeql-action's git tree. A SHA that was oncev4head but no longer exists in the action repository triggers Scorecard's "imposter commit" defence and breaks the workflow with a 400 error againstapi.scorecard.dev. Verify on every Dependabot bump by spot-checkinggh api /repos/github/codeql-action/commits/<sha>returns 200. - On upstream sync: zero interaction.
- Re-test on rebase:
```bash # Confirm the pin still resolves to a real commit: pin=$(grep -oE 'codeql-action/upload-sarif@[a-f0-9]{40}' \ .github/workflows/scorecard.yml | head -1 | cut -d@ -f2) gh api "/repos/github/codeql-action/commits/$pin" --jq '.sha' # Then watch the next master push for a green Scorecard run: gh run list --workflow scorecard --repo VMAFx/vmafx --limit 1
0228 — U-2-Net u2netp saliency replacement deferred (ADR-0265)¶
- Touches: docs-only.
docs/adr/0265-u2netp-saliency-replacement-blocked.md— new ADR continuing the deferral chain started by ADR-0257.docs/research/0055-u2netp-saliency-replacement-survey.md— new research digest (upstream survey + license + distribution- op-allowlist audit + alternatives walk).
docs/ai/models/mobilesal.md— pointer block updated to reference both ADR-0257 (first blocker) and ADR-0265 (second blocker).model/tiny/registry.json—mobilesal_placeholder_v0notesfield updated to reference ADR-0265 alongside ADR-0257 (no schema / sha256 / file changes).model/tiny/mobilesal.json— sidecarnotesfield updated in lockstep.scripts/gen_mobilesal_placeholder_onnx.py— generator notes string updated so re-running is idempotent against the new sidecar / registry text.CHANGELOG.md— Changed entry viachangelog.d/changed/T6-2a-followup-u2netp-replacement-deferred.md.docs/adr/README.md— index row viadocs/adr/_index_fragments/0265-u2netp-saliency-replacement-blocked.md.- Invariant: zero C-side surface change.
feature_mobilesal.ctensor-name contract (inputinput→ outputsaliency_map, NCHW float32[1, 3, H, W]→[1, 1, H, W]) is unchanged; the on-diskmodel/tiny/mobilesal.onnx(sha256f1226310…) is unchanged;mobilesal_placeholder_v0'ssmoke: trueflag is unchanged. Any future drop-in (U-2-Net viaT6-2a-mirror-u2netp-via-release+T6-2a-widen-allowlist-resize, distilled student, or BASNet / PoolNet survey result) replaces the.onnxand bumps the registry sha256 without touching the C side. - On upstream sync: zero interaction.
feature_mobilesal.c, the registry, the ADR, and the research digest are all fork-local (T6-2a; ADR-0218 / ADR-0257 / ADR-0265; not present in Netflix upstream). - Re-test on rebase:
meson setup build -Denable_cuda=false -Denable_sycl=false
ninja -C build
meson test -C build test_mobilesal
python3 ai/scripts/validate_model_registry.py
bash scripts/docs/concat-adr-index.sh --check
bash scripts/release/concat-changelog-fragments.sh --check
0108 — ssim_accumulate_avx512 per-lane double reduction vectorised¶
- ADR: ADR-0139 (existing; no new ADR — the per-lane reduction order is unchanged).
- Touches:
core/src/feature/x86/ssim_avx512.c— thessim_accumulate_block_avx512body. The per-lane scalarssim_accumulate_lanecalls (16 of them) are replaced by two 8-wide__m512dpasses that computelv,cv,sv, andlv*cv*svlane-wise in vector double. Aligneddouble[16]spill buffers replace the previous_Alignas(64) float[16]×6spill, and the scalar accumulation loop now does 4×16vaddsdinstead of 16 invocations of the per-lane helper.CHANGELOG.md— Changed entry.- This file — this entry.
- Invariant (load-bearing for ADR-0139 bit-exactness):
- Per-lane double computation order is byte-identical:
((2.0 * rm) * cm + C1) / l_den, then(2.0 * srsc + C2) / c_den, then(lv * cv) * sv. No FMA contraction (separate_mm512_mul_pd+_mm512_add_pd—_mm512_fmadd_pdis forbidden because it changes the rounding count and would diverge from scalar's two-stepmul+add). - Float→double widening uses
_mm512_cvtps_pdwhich is IEEE-754-exact for finite floats (52-bit mantissa fits 23-bit float losslessly). - Lane-by-lane left-to-right reduction order preserved:
local_ssim += t_ssim[k]fork = 0..15. Tree reductions (pairwise add, dual-accumulator unroll) are forbidden — they break running-sum associativity against scalar. - AVX2 / NEON twins kept on the per-lane scalar path. Verified bit-identical against the new AVX-512 at
--precision maxon the Netflixsrc01_hrc00/01_576x324and thecheckerboard_1920_1080_10_3_*_0pairs. The bit-exactness contract (ADR-0139) is per-lane, not per-ISA algorithm — so AVX2 / NEON stay scalar-per-lane until a dedicated PR vectorises them with the same care. - Rebase impact: zero conflict with Netflix upstream — the whole SSIM SIMD surface is fork-local (no upstream SSIM SIMD exists). Conflicts only arise if upstream changes
ssim_accumulate_default_scalariniqa/ssim_tools.c; in that case both the AVX2 / NEON per-lane helper and the AVX-512 vector-double block need a coordinated update preserving the three invariants above. - Re-test on rebase:
meson setup build libvmaf -Denable_cuda=false -Denable_sycl=false
ninja -C build
meson test -C build
# Bit-exact at --precision max, scalar vs AVX2 vs AVX-512:
for MASK in 0 16 255; do
core/build/tools/vmaf -r python/test/resource/yuv/src01_hrc00_576x324.yuv \
-d python/test/resource/yuv/src01_hrc01_576x324.yuv \
-w 576 -h 324 -p 420 -b 8 \
--feature float_ms_ssim --feature float_ssim \
--xml -o /tmp/m${MASK}.xml --precision max --cpumask $MASK
done
diff <(grep -v 'fyi fps' /tmp/m0.xml) <(grep -v 'fyi fps' /tmp/m16.xml) # empty
diff <(grep -v 'fyi fps' /tmp/m0.xml) <(grep -v 'fyi fps' /tmp/m255.xml) # empty
- Why this matters on rebase: an upstream commit that touches
core/src/feature/ssimulacra2.ccould prompt a "let's also port the GPU XYB while we're here" follow-up. The ledger entry is the standing answer: don't, the measurement was redone on NVIDIA in May 2026 and the result still failedplaces=4by five decades. See Research-0047.
0126 — FastDVDnet real upstream weights drop (ADR-0253)¶
- What changed: replaces
model/tiny/fastdvdnet_pre.onnxwith the wrapped real upstream FastDVDnet checkpoint (sha256eb9444cf6f07eefdc7f4f68d09131074dbd1dcee6f88a331ba684dd2fb5937d4, ~9.5 MiB), refreshes the sidecarmodel/tiny/fastdvdnet_pre.json, flips the registry row'ssmoke: true → falseand addslicense: "MIT"+ the upstream commit pinc8fdf61. New exporterai/scripts/export_fastdvdnet_pre.py(the older_placeholder.pyexporter is retained for reference). New ADRdocs/adr/0255-fastdvdnet-pre-real-weights.md; user-facing docdocs/ai/models/fastdvdnet_pre.mdrewritten with provenance, license attribution, and reproduce-the-export instructions. - Upstream source: fork-local. Netflix/vmaf does not ship a FastDVDnet temporal pre-filter; the C extractor and ONNX surface are entirely fork-introduced (ADR-0215). The wrapped weights are attribution-only (upstream
m-tassano/fastdvdnetMIT). - On upstream sync: zero interaction. Every file touched (
ai/scripts/export_fastdvdnet_pre*.py,model/tiny/fastdvdnet_pre.*,docs/ai/models/fastdvdnet_pre.md,docs/adr/0253-*.md, CHANGELOG fragment, ADR index fragment) lives in fork-introduced trees. - Re-test on rebase:
# Re-derive the ONNX from the pinned upstream checkpoint.
mkdir -p /tmp/fastdvdnet_upstream && cd /tmp/fastdvdnet_upstream
curl -L -O https://raw.githubusercontent.com/m-tassano/fastdvdnet/c8fdf61/model.pth
curl -L -O https://raw.githubusercontent.com/m-tassano/fastdvdnet/c8fdf61/models.py
cd /path/to/vmaf
python3 ai/scripts/export_fastdvdnet_pre.py \
--upstream-dir /tmp/fastdvdnet_upstream
python3 ai/scripts/validate_model_registry.py
meson test -C build --suite=fast --print-errorlogs test_fastdvdnet_pre
0127 — ONNX op-allowlist gains Resize (ADR-0258)¶
- Touches:
core/src/dnn/op_allowlist.c— fork-local file (no upstream counterpart). One new entry"Resize"under the/* convolutional */block.core/test/dnn/test_op_allowlist.c,core/test/dnn/test_onnx_scan.c— fork-local DNN tests.ai/tests/test_op_allowlist.py— fork-local Python parity test.- Invariant: the C allowlist is the single source of truth; the Python regex parser in
ai/src/vmaf_train/op_allowlist.pywalks the sameop_allowlist.cfile. Any future entry only needs the C edit — Python symmetry is automatic. - Upstream source: fork-local. Netflix/vmaf has no ONNX op- allowlist surface; the entire
core/src/dnn/tree is fork- introduced. - On upstream sync: zero interaction. Every file touched lives in fork-introduced trees.
- Re-test on rebase:
meson test -C build test_op_allowlist test_onnx_scan
PYTHONPATH=ai/src python -m pytest ai/tests/test_op_allowlist.py
0231 — vif.comp + ciede.comp precise decorations (ADR-0269 / Step A of Vulkan 1.4 bump)¶
- Touches:
core/src/feature/vulkan/shaders/vif.comp(3 local-variable type qualifiers:g,sv_sq,gg_sigma_f→precise float),core/src/feature/vulkan/shaders/ciede.comp(yuv_to_rgboutputs,rgb_to_xyzmatmul accumulators,ciede2000chroma magnitudes + half-axes + s_l/c/h + lightness/chroma/hue + final ΔE). - Invariant: Both shaders are fork-local (Vulkan backend is fork-added; upstream Netflix/vmaf has no Vulkan compute kernels). The
precisekeyword is GLSL 4.50 standard syntax; glslc 2026.1 lowers it to per-resultOpDecorate NoContraction. The decorations are load-bearing for the cross-backend gate on NVIDIA driver 595.71+ — removing them would re-introduce the 42/48 ciede regression at API 1.3 documented in research-0054. - On upstream sync: zero interaction. Both shader files are entirely fork-introduced; upstream has no Vulkan compute path.
- Re-test on rebase:
# Re-confirm the cross-backend gate on a Vulkan-capable host.
meson setup core/build -Denable_vulkan=enabled
ninja -C core/build
python3 scripts/ci/cross_backend_vif_diff.py \
--vmaf-binary core/build/tools/vmaf \
--reference python/test/resource/yuv/src01_hrc00_576x324.yuv \
--distorted python/test/resource/yuv/src01_hrc01_576x324.yuv \
--width 576 --height 324 --pixel-format 420 --bitdepth 8 \
--feature vif --backend vulkan --places 4
python3 scripts/ci/cross_backend_vif_diff.py \
--vmaf-binary core/build/tools/vmaf \
--reference python/test/resource/yuv/src01_hrc00_576x324.yuv \
--distorted python/test/resource/yuv/src01_hrc01_576x324.yuv \
--width 576 --height 324 --pixel-format 420 --bitdepth 8 \
--feature ciede --backend vulkan --places 4
# Confirm SPIR-V still emits NoContraction post-rebase.
glslc --target-env=vulkan1.3 -O \
core/src/feature/vulkan/shaders/vif.comp -o /tmp/vif.spv
spirv-dis /tmp/vif.spv | grep -c NoContraction # expect ≥ 60
Expected on NVIDIA 595.71+: vif 0/48 OK, ciede 5/48 FAIL (max abs 8.9e-05 — pre-existing fork debt at API 1.3, see ADR-0269). On RADV / lavapipe: bit-exact (precise is a no-op there).
0229 — fr_regressor_v2 codec-aware scaffold (ADR-0272)¶
- ADR: ADR-0272
- Touches:
ai/scripts/train_fr_regressor_v2.py(new) — Phase A JSONL consumer; trains the codec-aware FRRegressor.model/tiny/fr_regressor_v2.onnx(new, smoke) — placeholder ONNX from--smokemode; re-baked on production training.model/tiny/fr_regressor_v2.json(new) — sidecar.model/tiny/registry.json— new entry withsmoke: true.docs/adr/0272-fr-regressor-v2-codec-aware-scaffold.md(new).docs/adr/README.md— index row.docs/research/0058-fr-regressor-v2-feasibility.md(new).docs/ai/models/fr_regressor_v2.md(new) — model card.ai/AGENTS.md— invariant note (codec block layout + ENCODER_VOCAB ordering).CHANGELOG.md— Added entry.- Invariant: the 8-D codec block layout is
[encoder_onehot(6), preset_norm, crf_norm]withENCODER_VOCAB = (libx264, libx265, libsvtav1, libvvenc, libvpx-vp9, unknown)in load-bearing order. CRF normaliser is/63(union upper bound). Preset normaliser is/9. Bumping the vocabulary requires a re-train; existing checkpoints pin the order they were trained against viaencoder_vocab_versionin the sidecar. The two-input ONNX (features,codec) follows the LPIPS-Sq precedent (ADR-0040 / ADR-0041). - Rebase impact: entirely fork-local; pure additive; no upstream-mirror file is touched. Phase A schema (consumed by this trainer) is itself fork-local (
tools/vmaf-tune/). No conflict expected on/sync-upstream. - Re-test on rebase:
0311 — libFuzzer harness expansion: yuv_input + cli_parse (ADR-0311)¶
- ADR: ADR-0311; parent ADR-0270.
- Touches:
core/test/fuzz/fuzz_yuv_input.c(new)core/test/fuzz/fuzz_cli_parse.c(new)core/test/fuzz/meson.build— two newexecutable(...)blocks for the harnesses, plus a sharedfuzz_vidinput_sourceslist.core/test/fuzz/yuv_input_corpus/*(new — 6 seeds covering 8/10-bit × 4:2:0 / 4:2:2 / 4:4:4 plus a truncated-frame seed).core/test/fuzz/cli_parse_corpus/*(new — 6 seeds covering the--feature,--model,--reference, YUV-flag, and--helpshapes).core/test/fuzz/README.md— Targets table extended..github/workflows/fuzz.yml— matrix gainsfuzz_yuv_input+fuzz_cli_parse; per-harness wall-clock budget reduced from 300 s to 60 s so the 3-target matrix fits the existingtimeout-minutes: 15cap.docs/development/fuzzing.md— runbook table + smoke commands extended.docs/adr/0311-libfuzzer-harness-expansion.md(new)docs/research/0083-libfuzzer-harness-expansion-target-survey.md(new)libvmaf/AGENTS.md— new invariant block for the one-parser-one-harness rule.CHANGELOG.md— Added entry.- Invariant:
- The fuzz scaffold remains opt-in (
-Dfuzz=true) — every defaultmeson setupinvocation must continue to skip it. fuzz_yuv_inputre-includestools/yuv_input.cand the rest of the vidinput trio as build inputs. Upstream Netflix/vmaf splits or renames of those source files need the matchingmeson.buildsource-list update.fuzz_cli_parsere-includestools/cli_parse.cas a build input and links againstlibvmafforvmaf_version()and feature-dictionary symbols. The-Wl,--wrap=exitlink arg is load-bearing — without it,usage()'sexit(1)would terminate the fuzzer process on first bad input.LLVMFuzzerTestOneInputkeeps external linkage; the scaffold-wide// NOLINTNEXTLINE(misc-use-internal-linkage)pattern is correct for libFuzzer's name-resolved entry-point ABI.- Rebase impact: any upstream sync that touches
core/tools/{yuv_input,cli_parse}.cmust re-run the 60 s smoke per harness on the merged tip; record any new-found crash-* artefact under the matching<target>_known_crashes/dir, not in<target>_corpus/. The__wrap_exitshim infuzz_cli_parse.cis GNU-ld / lld-only; do not assume it works on Apple ld without an-undefined,dynamic_lookupfallback. - Re-test on rebase:
CC=clang CXX=clang++ \
meson setup build-fuzz libvmaf \
--buildtype=debug \
-Db_sanitize=address \
-Db_lundef=false \
-Dfuzz=true \
-Denable_cuda=false -Denable_sycl=false -Denable_vulkan=disabled
ninja -C build-fuzz \
test/fuzz/fuzz_y4m_input \
test/fuzz/fuzz_yuv_input \
test/fuzz/fuzz_cli_parse
./build-fuzz/test/fuzz/fuzz_yuv_input \
-seed=0 -runs=1000 \
core/test/fuzz/yuv_input_corpus/
./build-fuzz/test/fuzz/fuzz_cli_parse \
-seed=0 -runs=1000 \
core/test/fuzz/cli_parse_corpus/
0229 — libFuzzer scaffold for the YUV4MPEG2 parser (ADR-0270)¶
- ADR: ADR-0270
- Touches:
core/test/fuzz/fuzz_y4m_input.c(new)core/test/fuzz/meson.build(new)core/test/fuzz/README.md(new)core/test/fuzz/y4m_input_corpus/*(new — six seeds)core/test/fuzz/y4m_input_known_crashes/*(new — one 411-chroma OOB reproducer; excluded from CI corpus)core/test/meson.build—subdir('fuzz')line.core/meson_options.txt— newoption('fuzz', ...)..github/workflows/fuzz.yml(new — nightly 5-minute job).docs/development/fuzzing.md(new — operator runbook).docs/adr/0270-fuzzing-scaffold.md(new)docs/research/0059-libfuzzer-scaffold-y4m.md(new)docs/state.md— new Open-bug row for the 411-chroma OOB write.CHANGELOG.md— Added entry.- Invariant: the fuzz scaffold is opt-in — every default
meson setupinvocation must continue to skip it. The harness links statically againstcore/tools/{y4m_input,yuv_input,vidinput}.crather thanlibvmaf.soso the public C-API surface stays unchanged. - Rebase impact: the harness re-includes
core/tools/y4m_input.cas a build input. Any upstream Netflix/vmaf change that splits or renames the tool sources (e.g. moves the parser intocore/src/) needs the correspondingmeson.buildsource list update and the harness re-test below. They4m_input_known_crashes/y4m_411_w2_h4_oob_dst.y4mreproducer is the regression gate for the parser fix; do not delete it on upstream sync — if upstream lands the same fix, port the reproducer back intoy4m_input_corpus/as a permanent seed. - Re-test on rebase:
CC=clang CXX=clang++ \
meson setup build-fuzz libvmaf \
--buildtype=debug \
-Db_sanitize=address \
-Db_lundef=false \
-Dfuzz=true \
-Denable_cuda=false -Denable_sycl=false -Denable_vulkan=disabled
ninja -C build-fuzz test/fuzz/fuzz_y4m_input
./build-fuzz/test/fuzz/fuzz_y4m_input \
-max_total_time=60 \
core/test/fuzz/y4m_input_corpus/
# Verify the known-crash reproducer still triggers (until the fix lands):
./build-fuzz/test/fuzz/fuzz_y4m_input \
core/test/fuzz/y4m_input_known_crashes/y4m_411_w2_h4_oob_dst.y4m
0231 — HIP seventh-consumer kernel float_motion_hip (ADR-0273)¶
- ADR: ADR-0273
- Touches:
core/src/feature/hip/float_motion_hip.c(new) — seventh consumer ofcore/src/hip/kernel_template.h. Mirrorscore/src/feature/cuda/float_motion_cuda.ccall-graph-for-call-graph;init/submit/collect/closeinvoke the kernel-template helpers in the same order;flush()callback for tail-frame motion2 emission;motion_force_zeroshort-circuit posture (fex->extractswap withsubmit / collect / flush / closenulled). Submit path intentionally bypassesvmaf_hip_kernel_submit_pre_launch(kernel writes per-WG SAD float partials directly, no atomic, no memset).core/src/feature/hip/float_motion_hip.h(new)core/src/hip/meson.build— new entry inhip_sources.core/src/feature/feature_extractor.c— extern declaration plusfeature_extractor_list[]entry under#if HAVE_HIP.core/test/test_hip_smoke.c— new sub-testtest_float_motion_hip_extractor_registered(also asserts theVMAF_FEATURE_EXTRACTOR_TEMPORALflag bit) and a row intest_table[].docs/adr/0273-hip-seventh-consumer-float-motion.md(new)docs/adr/README.md— index row.docs/backends/hip/overview.md— seventh / eighth consumer note.core/src/hip/AGENTS.md— invariant note.CHANGELOG.md— Added entry (joint with ADR-0274).- Invariant — three-buffer ping-pong +
motion_force_zeroshort-circuit are load-bearing. The state struct carries threeuintptr_tbuffer slots (ref_in,blur[2]) that the runtime PR (T7-10b) will swap for real device-buffer handles matching the CUDA twin'sVmafCudaBuffer *ref_in+VmafCudaBuffer *blur[2]field shape. Themotion_force_zeroshort-circuit (fex->extractswap, kernel-template helpers nulled) must stay aligned with the CUDA twin on every refactor — otherwise the runtime PR's helper-body flip diverges between the two backends. Thesubmit_pre_launchbypass mirrors the CUDA twin; if a future PR adds asubmit_pre_launchcall tofloat_motion_cuda.c's submit path, the HIP twin must follow in the same PR. - Rebase impact: entirely fork-local. New files are HIP-specific. The only upstream-touching edit is
feature_extractor.c, but the change sits inside an existing#if HAVE_HIPblock (ADR-0241); upstream has noHAVE_HIPso no conflict is expected. - Re-test on rebase:
meson setup build libvmaf -Denable_hip=true \
-Denable_cuda=false -Denable_sycl=false -Denable_vulkan=disabled
ninja -C build
meson test -C build test_hip_smoke
0232 — HIP eighth-consumer kernel float_ssim_hip (ADR-0274)¶
- ADR: ADR-0274
- Touches:
core/src/feature/hip/float_ssim_hip.c(new) — eighth consumer ofcore/src/hip/kernel_template.h. Mirrorscore/src/feature/cuda/integer_ssim_cuda.ccall-graph-for-call-graph (the CUDA file registersvmaf_fex_float_ssim_cudadespite itsinteger_filename). First multi-dispatch HIP consumer (chars.n_dispatches_per_frame == 2). Submit path intentionally bypassesvmaf_hip_kernel_submit_pre_launch(kernel writes per-block float partials directly). State struct carries fiveuintptr_tintermediate float buffer slots (h_ref_mu,h_cmp_mu,h_ref_sq,h_cmp_sq,h_refcmp) tracked outside the kernel-template's readback bundle.validate_dims_hipandinit_dims_hiphelpers extracted frominit()to fit thereadability-function-sizebudget.core/src/feature/hip/float_ssim_hip.h(new)core/src/hip/meson.build— new entry inhip_sources.core/src/feature/feature_extractor.c— extern declaration plusfeature_extractor_list[]entry under#if HAVE_HIP.core/test/test_hip_smoke.c— new sub-testtest_float_ssim_hip_extractor_registered(also assertschars.n_dispatches_per_frame == 2) and a row intest_table[].docs/adr/0274-hip-eighth-consumer-float-ssim.md(new)docs/adr/README.md— index row.docs/backends/hip/overview.md— seventh / eighth consumer note (joint).core/src/hip/AGENTS.md— invariant note.CHANGELOG.md— Added entry (joint with ADR-0273).- Invariant — multi-dispatch + five-slot buffer pyramid + v1
scale=1validation are load-bearing. The state struct carries fiveuintptr_tintermediate float buffer slots that the runtime PR (T7-10b) will swap for real device-buffer handles matching the CUDA twin'sVmafCudaBuffer *h_*field shape — any drift in the CUDA twin's slot count requires a paired update here. Thechars.n_dispatches_per_frame == 2characteristic is asserted in the smoke test; do not silently lower it. The v1scale=1-EINVALvalidation surface (invalidate_dims_hip) must stay aligned with the CUDA twin'scompute_scale/vmaf_logchain. The HIP twin'svalidate_dims_hip/init_dims_hipextraction is intentional for the function-size budget; do not re-inline without verifying the budget still passes. - Rebase impact: entirely fork-local; same posture as ADR-0273.
- Re-test on rebase:
meson setup build libvmaf -Denable_hip=true \
-Denable_cuda=false -Denable_sycl=false -Denable_vulkan=disabled
ninja -C build
meson test -C build test_hip_smoke
0229 — vmaf_tiny_v3 + vmaf_tiny_v4 dynamic-PTQ int8 sidecars (ADR-0275)¶
0278 — vmaf-tune libaom-av1 codec adapter (2026-05-03)¶
0228 — vmaf-tune libx265 codec adapter (ADR-0288)¶
0280 — vmaf-tune NVENC codec adapters (ADR-0290)¶
- Touches:
tools/vmaf-tune/src/vmaftune/codec_adapters/{h264_nvenc,hevc_nvenc,av1_nvenc,_nvenc_common}.py(new). Wholly fork-local — no upstream Netflix/vmaf overlap.tools/vmaf-tune/src/vmaftune/codec_adapters/__init__.py— registry expanded.tools/vmaf-tune/tests/test_codec_adapter_nvenc.py(new).tools/vmaf-tune/tests/test_corpus.py— Phase-A registry assertion updated.tools/vmaf-tune/AGENTS.md— invariant note expanded.docs/usage/vmaf-tune.md— "Hardware encoders (NVENC)" section.docs/adr/0290-vmaf-tune-nvenc-adapters.md(new) +docs/adr/README.mdindex row.docs/research/0065-vmaf-tune-nvenc-adapters.md(new).CHANGELOG.md— Added entry.- Invariant:
known_codecs()returns the four-codec tuple("av1_nvenc", "h264_nvenc", "hevc_nvenc", "libx264"); the mnemonic preset map (ultrafast/superfast/veryfast→p1,faster→p2,fast→p3,medium→p4,slow→p5,slower→p6,slowest/placebo→p7) is the canonical cross-codec preset alignment that downstream Phase B/C consumers assume. The CQ window is the hardware-permitted[0, 51]; the Phase A informative window is[15, 40]. - Rebase impact: zero —
tools/vmaf-tune/is wholly fork-local and has no upstream Netflix/vmaf path overlap. - Re-test on rebase:
0227 — ffmpeg-patches/ series re-verified against n8.1 (2026-05-03)¶
- Touches:
tools/vmaf-tune/src/vmaftune/codec_adapters/x265.py(new),tools/vmaf-tune/src/vmaftune/codec_adapters/__init__.py(registry add),tools/vmaf-tune/src/vmaftune/encode.py(parse_versions(stderr, encoder=…)gains a per-codec branch),tools/vmaf-tune/src/vmaftune/cli.py(help-text wording only),tools/vmaf-tune/tests/test_codec_adapter_x265.py(new),tools/vmaf-tune/tests/test_corpus.py(membership-based codec list assertion). - Invariant: the codec-adapter contract documented in
tools/vmaf-tune/AGENTS.md(multi-codec from day one; the search loop never branches on codec identity). Theparse_versionssignature is still backward-compatible —encoderdefaults tolibx264so callers from before this PR keep working. - Upstream source: fork-local.
tools/vmaf-tune/is fork-only; upstream Netflix/vmaf does not ship encode automation. - On upstream sync: zero interaction. Confirm the
_index_fragments/_order.txtrow for0288-vmaf-tune-codec-adapter-x265remains present after any cross-merge. - Re-test on rebase:
0227 — ffmpeg-patches/ series re-verified against n8.1 (2026-05-03)¶
- Touches:
tools/vmaf-tune/src/vmaftune/codec_adapters/libaom.py(new),tools/vmaf-tune/src/vmaftune/codec_adapters/__init__.py(registry row + import),tools/vmaf-tune/tests/test_corpus.py(membership assertion relaxed from== ("libx264",)to"libx264" in known_codecs()),tools/vmaf-tune/tests/test_codec_adapter_libaom.py(new),tools/vmaf-tune/AGENTS.md(preset-vocabulary invariant). - Invariant: the cross-codec preset vocabulary (
placebo, slowest, slower, slow, medium, fast, faster, veryfast, superfast, ultrafast) is shared across AV1-family adapters so one--presetaxis covers x264 / x265 / svtav1 / libaom-av1. Each adapter maps the human name onto its codec-specific knob; do not introduce per-adapter preset names. - Upstream source: fork-local.
tools/vmaf-tune/is the fork-introduced quality-aware encode automation harness (ADR-0237); it has no upstream Netflix/vmaf counterpart. - On upstream sync: zero interaction with
upstream/master. Self-contained intools/vmaf-tune/anddocs/. - Re-test on rebase:
0227 — ffmpeg-patches/ series re-verified against n8.1 (2026-05-03)¶
- ADR: ADR-0275
- Touches:
model/tiny/vmaf_tiny_v3.int8.onnx(new, 4 267 B)model/tiny/vmaf_tiny_v4.int8.onnx(new, 7 769 B)model/tiny/registry.json— newvmaf_tiny_v3andvmaf_tiny_v4rows withquant_mode,int8_sha256,quant_accuracy_budget_plccfields.model/tiny/vmaf_tiny_v3.json,model/tiny/vmaf_tiny_v4.json— same fields mirrored into the per-model sidecars.docs/ai/models/vmaf_tiny_v3.md,docs/ai/models/vmaf_tiny_v4.md— new "Quantisation" sections.docs/adr/0275-vmaf-tiny-v3-v4-ptq.md(new) and ADR index row.CHANGELOG.md— Added entry.- Invariant:
python ai/scripts/measure_quant_drop.py --allreports[PASS]for bothvmaf_tiny_v3(drop ≤ 0.001 on Netflix features) andvmaf_tiny_v4(drop ≤ 0.001), inside the 0.01 per-model budget. The runtime redirect from ADR-0174 picks the.int8.onnxsibling when an operator's registry overlay declaresquant_mode: dynamic. - Rebase impact: entirely fork-local — neither v3 nor v4 nor the dynamic-PTQ harness exists upstream. The new int8 ONNX bytes ship as committed binaries (mirroring
learned_filter_v1andnr_metric_v1); they are well below the few-MB external-data threshold and don't require the sigstore +.onnx.datapattern. - Re-test on rebase:
```bash python ai/scripts/validate_model_registry.py python ai/scripts/measure_quant_drop.py --all
0229 — NVIDIA-Vulkan ciede2000 places=4 fork debt root-cause (ADR-0273)¶
- Touched files: docs-only.
docs/adr/0273-...precision-gap.md(new) +_index_fragments/row +_order.txtappend.docs/research/0055-ciede-vulkan-nvidia-f32-f64-root-cause.md(new) +docs/research/README.mdindex row.docs/state.md— Open-bugs rowT-VK-CIEDE-F32-F64.docs/backends/vulkan/overview.md— NVIDIA-hardware caveat.changelog.d/changed/ciede-vulkan-nvidia-f32-f64-precision-gap.md(new).core/src/vulkan/AGENTS.md— invariant cross-link.- Invariant: the ciede.comp shader's f32 precision contract is load-bearing — promoting to f64 would silently change scores on every Vulkan device that supports
shaderFloat64and create a per-device-feature-bit divergence (RTX 4090 has it; many consumer GPUs don't). The CPUciede.c::get_lab_colordoing its colour-space chain indoubleis upstream Netflix behaviour and must not be narrowed to f32 to "fix" the GPU gap (would change Netflix golden ground truth). The 5/48 NVIDIA places=4 mismatch on the highest-ΔE frames is expected and documented; do not attempt to "fix" it without re-reading ADR-0273 first. - Rebase impact: zero — docs-only. The CPU and shader sources this ADR analyses are unchanged by this PR. If a future upstream rebase touches
ciede.c::get_lab_color(thedoublechain) the ADR's reasoning still holds; if upstream changes the CPU reference's precision posture, ADR-0273 needs aStatus: Supersededentry. - Re-test on rebase: a manual NVIDIA-hardware run if available:
```bash cd libvmaf && meson setup build \ -Denable_vulkan=enabled -Denable_cuda=false && ninja -C build cd .. python3 scripts/ci/cross_backend_vif_diff.py \ --vmaf-binary $PWD/core/build/tools/vmaf \ --reference testdata/ref_576x324_48f.yuv \ --distorted testdata/dis_576x324_48f.yuv \ --width 576 --height 324 \ --feature ciede --backend vulkan --device 0 --places 4 # Expected post-PR-346 (when merged): 5/48 mismatches at 1.78× threshold. # Expected pre-PR-346 (current master): 42/48 mismatches at higher ratio. # If the count drops below 5/48 on NVIDIA, ADR-0273 should record the # delta and consider closing T-VK-CIEDE-F32-F64.
0229 — tools/vmaf-tune fast Phase A.5 scaffold (ADR-0276)¶
- Touches:
tools/vmaf-tune/src/vmaftune/fast.py(new),tools/vmaf-tune/src/vmaftune/cli.py(newfastsubcommand branch),tools/vmaf-tune/pyproject.toml(new[fast]extra),tools/vmaf-tune/tests/test_fast.py(new),tools/vmaf-tune/AGENTS.md(new invariants),docs/usage/vmaf-tune.md(new "Phase A.5" section),docs/adr/0276-vmaf-tune-fast-path.md(new ADR),docs/research/0060-vmaf-tune-fast-path.md(new digest). - Invariant: the
fastsubcommand is opt-in and never automatically replaces the Phase A grid path. The slow grid is the ground-truth corpus generator (ADR-0237 contract); fast-path is for the recommendation use case only. Optuna is a lazy-imported optional dep gated behind the[fast]extra — importing it at module scope outsidefast.py(or its tests) breaks the zero-dep core install. - Rebase impact: entirely fork-local; the tool sits under
tools/vmaf-tune/which is fork-added, and no upstream files are touched. Upstream Netflix/vmaf has no analogous surface. - Re-test on rebase:
pip install -e 'tools/vmaf-tune[fast]'
pytest tools/vmaf-tune/tests/test_fast.py -v
vmaf-tune fast --smoke --target-vmaf 92
0229 — vmaf-tune recommend subcommand (ADR-0237 Phase B-lite)¶
- Touches:
tools/vmaf-tune/src/vmaftune/recommend.py(new). Wholly fork-local — no upstream Netflix/vmaf path overlap.tools/vmaf-tune/src/vmaftune/cli.py— addsrecommendsubparser;corpussubcommand untouched.tools/vmaf-tune/tests/test_recommend.py(new). 13-case smoke suite, mocks all binaries; runs in <100 ms.docs/usage/vmaf-tune.md— adds## recommendsection.- Invariant:
recommendconsumes the existingCORPUS_ROW_KEYSschema unchanged —vmaf_score,bitrate_kbps,crf,preset,encoder,exit_status. No schema bump. If a future PR bumpsSCHEMA_VERSION, both thecorpuswriter and therecommendreader must be updated in lockstep; tests assert this viatest_corpus_row_keys_match_init_contract. - Rebase impact: zero —
tools/vmaf-tune/is wholly fork-local; no upstream surface touches it. - Re-test on rebase:
0228 — integer_ms_ssim_cuda.c joins drain_batch (T-GPU-OPT-2 / ADR-0271)¶
- Touches:
core/src/feature/cuda/integer_ms_ssim_cuda.c. No upstream Netflix/vmaf changes expected here — the file is fork-added (CUDA twin of the upstream-portms_ssim_score.cu) and the surface this PR redrew (per-scalel_partials[i]/c_partials[i]/s_partials[i]arrays + the per-scaleh_l_partials[i]/h_c_partials[i]/h_s_partials[i]pinned host shadows + thesubmit()<→collect()work redistribution + thecuEventRecord(s->lc.finished, s->lc.str)+vmaf_cuda_drain_batch_register(&s->lc)tail) is also entirely fork-local. - Invariant: the engine-scope drain-batch contract from ADR-0271 / drain_batch.h. The kernel-launch order on
s->lc.strmust stay stable:decimate (× 4)then for each scalei ∈ 0..4horiz⇒vert_lcs⇒ DtoH(l_partials[i]) ⇒ DtoH(c_partials[i]) ⇒ DtoH(s_partials[i])thencuEventRecord(s->lc.finished, s->lc.str)thenvmaf_cuda_drain_batch_register(&s->lc). Same-stream ordering is what makes the shared SSIM intermediates (h_ref_mu,h_cmp_mu,h_ref_sq,h_cmp_sq,h_refcmp`) safe across scales without explicit sync — any change that parallelises the per-scale work onto multiple streams breaks bit-exactness unless per-scale intermediates are also added. - On upstream sync: zero interaction (the file is fork-added). If a future upstream PR adds an
integer_ms_ssim_cuda.cof its own, the merger must reconcile the per-scale partials topology + the drain_batch tail with whatever the new upstream shape brings. - Re-test on rebase:
cd libvmaf && meson setup build -Denable_cuda=false -Denable_sycl=false
ninja -C build # confirms the CPU build still links cleanly
# If the dev host has a working nvcc / host-compiler pair:
meson setup build_cuda -Denable_cuda=true -Denable_sycl=false
ninja -C build_cuda src/liblibvmaf_feature.a.p/feature_cuda_integer_ms_ssim_cuda.c.o
# Netflix CPU golden gate (CPU is the bit-exactness ground truth):
make test-netflix-golden
# Cross-backend parity (places=4 gate, ADR-0214):
/cross-backend-diff
0277 — ffmpeg-patches refresh against n8.1 — 2026-05-04 (ADR-0277)¶
- Touches:
ffmpeg-patches/is unchanged (no content drift). Doc-only entries land in: docs/adr/0277-ffmpeg-patches-refresh-2026-05-04.md— new ADR.docs/adr/_index_fragments/0277-ffmpeg-patches-refresh-2026-05-04.md— index row.docs/adr/_index_fragments/_order.txt— manifest append.changelog.d/changed/ffmpeg-patches-refresh-2026-05-04.md— Changed entry.- This file — this entry.
- Invariant:
ffmpeg-patches/series.txtorder is load-bearing — patches0002…0006build on each other and only apply cleanly cumulatively. The verification gate is a series replay, not a per-patchgit apply --check(per ADR-0118 + CLAUDE.md §12 r14). - On upstream sync: zero interaction. Netflix/vmaf has no
ffmpeg-patches/tree; this is a fork-local integration surface. - Re-test on rebase (also: re-replay procedure for the next refresh):
# Clone pristine n8.1
git -C /tmp clone --depth 1 --branch n8.1 \
https://github.com/FFmpeg/FFmpeg.git ff-replay-$(date +%F)
cd /tmp/ff-replay-$(date +%F)
git switch -c refresh-$(date +%F)
git config user.email refresh@local && git config user.name "Refresh Bot"
# Replay the series cumulatively
for p in /path/to/vmaf/ffmpeg-patches/000*-*.patch; do
git am --3way "$p" || break
done
# Regenerate and compare to in-tree
mkdir -p /tmp/ff-regen-$(date +%F)
git format-patch n8.1.. -o /tmp/ff-regen-$(date +%F)/
# Diff old vs new excluding pure format-patch noise
for i in 1 2 3 4 5 6; do
orig=$(ls /path/to/vmaf/ffmpeg-patches/000${i}-*.patch)
regen=$(ls /tmp/ff-regen-$(date +%F)/000${i}-*.patch)
diff -u \
<(grep -v "^From [0-9a-f]\|^Date:\|^index " "$orig") \
<(grep -v "^From [0-9a-f]\|^Date:\|^index " "$regen") \
| head -40
done
If only stylistic diffs surface (PATCH N/M numbering, MIME headers, hunk-context counts, hunk offset shifts against cumulative state), keep originals — record a no-drift refresh ADR. If real content drift surfaces, regenerate and ship the refresh PR with the regenerated patches plus a content-summary ADR.
End-to-end vf_libvmaf smoke is best run from CI (ffmpeg-integration.yml) against an installed libvmaf prefix — the meson-uninstalled .pc does not satisfy FFmpeg's #include <libvmaf.h> probe (the headers live under libvmaf/libvmaf.h only; the system-installed .pc carries an extra -I${includedir}/libvmaf shortcut that the uninstalled .pc omits).
0229 — T7-5 NOLINT-sweep closeout (ADR-0278)¶
- Touched files:
core/src/feature/integer_adm.c(1 NOLINT cite, line ~988adm_decouple_s123— upstream-mirror Netflix966be8d5).core/src/feature/cuda/ssimulacra2_cuda.c(3 NOLINT cites:ss2c_picture_to_linear_rgb,ss2c_host_combine,ss2c_run_scale_gpu/extract_fex_cuda).core/src/feature/vulkan/ssimulacra2_vulkan.c(3 NOLINT cites:ss2v_setup_gaussian,ss2v_picture_to_linear_rgb,ss2v_run_scale).core/src/feature/vulkan/cambi_vulkan.c(1 NOLINT cite:cambi_vk_extract).core/src/feature/sycl/integer_adm_sycl.cpp(6 cites, SYCL kernel-launch entries).core/src/feature/sycl/integer_motion_sycl.cpp(2 cites).core/src/feature/sycl/integer_vif_sycl.cpp(4 cites).core/tools/vmaf.c(3 cites:copy_picture_data,init_gpu_backends,main).- Invariant: zero behavioural change. Edits are inside comment blocks — appended
(ADR-0141 §2 ... load-bearing invariant; T7-5 sweep closeout — ADR-0278)to existing prose justifications. No function bodies split. The 12 SYCL sites share an identical justification string verbatim; preserving the byte-for-byte duplicate is the load-bearing documentation pattern (grep-able across the SYCL TUs). - On upstream sync: minimal interaction. The cite-only edits live inside comment blocks above the function signatures; rebases will surface them as touched lines but the function bodies are unchanged. For
integer_adm.c's upstream-mirror block (Netflix966be8d5), the comment edit at line 984–991 is cosmetic — keep the fork's version on conflict (it merely names the ADR; the underlying prose is unchanged). - Re-test on rebase:
```bash # 1. Programmatic audit must report 0 missing citations python3 - <<'PY' import re, os paths = [os.path.join(r, f) for r, _, fs in os.walk('libvmaf/src') for f in fs if f.endswith(('.c','.cpp','.h'))] paths.append('core/tools/vmaf.c') miss = total = 0 for p in paths: with open(p) as fh: ls = fh.readlines() for i, line in enumerate(ls): if 'NOLINT' in line and 'readability-function-size' in line and 'NOLINTEND' not in line: total += 1 ctx = [line]; j = i - 1 while j >= 0 and j > i - 14: s = ls[j].strip() if not s: break if s.startswith(('//','/','')): ctx.insert(0, ls[j]); j -= 1 else: break buf = ''.join(ctx) if 'ADR-' not in buf and not re.search(r'[Rr]esearch-?\d', buf): miss += 1 print(f"sites={total} missing={miss}") PY
# 2. Build + Netflix golden gate meson setup build -Denable_cuda=false -Denable_sycl=false ninja -C build make test-netflix-golden
0231 — vmaf-tune score path decodes mp4 -> raw YUV¶
- Touches:
tools/vmaf-tune/src/vmaftune/score.py(new_decode_to_raw_yuv+_needs_decodehelpers,run_scoreshells out to ffmpeg whenreq.distorted.suffix not in {.yuv, .y4m});tools/vmaf-tune/tests/test_corpus.py(3 new regression tests + the smoke-end-to-end mock now also stubs the ffmpeg decode call). - Invariant: the decode-back is the contract the libvmaf CLI imposes — mp4/webm/etc.
--distortedis silently rejected as raw-yuv with the wrong byte count, surfacing asexit_status=234. Future encoder adapters that emit non-raw containers inherit this decode automatically. Do not "optimise" the temp YUV away without first migrating the corpus pipeline to theffmpeg+libvmaffilter (which can pipe an mp4 stream in directly). - On upstream sync: zero interaction.
vmaf-tuneis fork-only tooling; upstream Netflix/vmaf has no analogue. - Re-test on rebase:
```bash cd tools/vmaf-tune && python3 -m pytest tests/ # plus an end-to-end smoke (needs a real raw YUV + ffmpeg + vmaf): ./vmaf-tune corpus --source /path/to/ref.yuv --width 1920 \ --height 1080 --pix-fmt yuv420p --framerate 25 --duration 6 \ --encoder libx264 --preset medium --crf 23 \ --output /tmp/smoke.jsonl --no-source-hash # expect: vmaf_score is a real number, not NaN.
0232 — CUDA build pins nvcc --std c++20¶
- Touches:
core/src/meson.buildline 686 (cuda_flags = [...]). - Invariant: nvcc 12.x clamps host C++ at C++17 by default; 13.x accepts up to C++20. Bumping the host stdlib past nvcc's default (any gcc >= 16, libstdc++ ships C++23 features) breaks the host-side parse in
<type_traits>/<bits/utility.h>. Forcing--std c++20on CUDA 13+ keeps the host headers parseable. Do not drop this flag without first checking the host gcc version against nvcc's default. - On upstream sync: zero interaction. Netflix/vmaf doesn't ship the
cuda_flagslist shape we use (their CUDA build is the original pre-fork pattern); a sync that touchescore/src/meson.buildaround theis_cuda_enabledbranch should keep the--std c++20injection. - Re-test on rebase:
meson setup core/build-cuda -Denable_cuda=true \
-Denable_sycl=false -Denable_vulkan=disabled
ninja -C core/build-cuda
# smoke
./core/build-cuda/tools/vmaf --gpumask=0 --no_sycl --no_vulkan \
-r .workingdir2/netflix/ref/BigBuckBunny_25fps.yuv \
-d .workingdir2/netflix/ref/BigBuckBunny_25fps.yuv \
-w 1920 -h 1080 -p 420 -b 8
0233 — CUDA motion flush_fex_cuda idempotency guard¶
- Touches:
core/src/feature/cuda/integer_motion_cuda.c— factored anappend_if_unwrittenhelper and routed the two motion2 / motion3 final-frame writes through it. - Invariant: under T-GPU-OPT-1 (PR #312 / ADR-0242), the pending-collect inside
flush_context_cudamay already have writtenmotion2_score[s->index]/motion3_score[s->index]beforeflush_fex_cudaruns. Any future motion-cuda flush logic that emits the same (feature, index) pair must keep this idempotency contract orflush_context_cudawill mis-surface as "context could not be synchronized". - On upstream sync: the bug only exists because the fork's
flush_context_cudaruns the pending-collect before the per-extractor flush. Netflix/vmaf upstream doesn't have the T-GPU-OPT-1 drain pattern, so the pre-#312 code path didn't duplicate-write. If Netflix lands a similar pattern, the fix shape mirrors what's done here. - Re-test on rebase:
ninja -C core/build-cuda
./core/build-cuda/tools/vmaf --gpumask=0 --no_sycl --no_vulkan \
-r python/test/resource/yuv/src01_hrc00_576x324.yuv \
-d python/test/resource/yuv/src01_hrc01_576x324.yuv \
-w 576 -h 324 -p 420 -b 8 \
--model path=model/vmaf_v0.6.1.json --threads 1 -q \
--output /tmp/cuda.json --json
# Expect: clean run, no "cannot be overwritten" warning,
# no "problem flushing context" error.
0234 — hw_encoder_corpus.py Phase A real-corpus runner¶
- Touches: new
scripts/dev/hw_encoder_corpus.py(no existing caller; opt-in tooling). Output landing inruns/phase_a/is gitignored — rerun the script to reproduce.docs/development/intel-arc-vaapi-driver-priority.md. Output landing inruns/phase_a/is gitignored — rerun the script to reproduce. stratified sample, 58 KiB). - Invariant: the script's QSV path forces
env['LIBVA_DRIVER_NAME']='iHD'(set by the calling shell, not inside the script) when targeting/dev/dri/renderD129on a multi-card host that has NVIDIA's libva-driver-nvidia shim installed. Without that, libva picks up NVIDIA's NVDEC-VAAPI translation and the MFX session handshake fails with -9. See the companion doc for the failure mode + fix. - On upstream sync: zero interaction. The script lives under
scripts/dev/(fork-only); upstream Netflix/vmaf has no comparable Phase A corpus tooling. - Re-test on rebase:
python3 scripts/dev/hw_encoder_corpus.py \
--vmaf-bin core/build-cuda/tools/vmaf \
--source .workingdir2/netflix/ref/BigBuckBunny_25fps.yuv \
--width 1920 --height 1080 --pix-fmt yuv420p --framerate 25 \
--encoder h264_nvenc --cq 25 \
--out /tmp/smoke.jsonl
# Expect: 1 cell × ~150 frames, per-frame canonical-6 + vmaf,
# encoder=h264_nvenc, cq=25.
0235 — fr_regressor_v2 ENCODER_VOCAB v2 (hw codec extension)¶
- Touches:
ai/scripts/train_fr_regressor_v2.py—ENCODER_VOCABgains 6 hw-codec entries (3 NVENC + 3 QSV);ENCODER_VOCAB_VERSIONbumps 1 -> 2;PRESET_ORDINALgains 6 sub-tables forp1..p7(NVENC) and the libx264-aligned QSV preset family. - Invariant: vocab order is load-bearing — index of every entry is baked into trained model graphs as a one-hot column position. New entries MUST be appended (never inserted into the middle), and the
unknownsentinel MUST stay last (UNKNOWN_ENCODER_INDEX = N - 1). BumpingENCODER_VOCAB_VERSIONsignals that any v1-graph ONNX needs re-export against v2 before consuming v2 training rows. - On upstream sync: zero interaction.
train_fr_regressor_v2.pyis fork-only (Phase B prereq, ADR-0237 / ADR-0272). - Re-test on rebase:
python3 ai/scripts/train_fr_regressor_v2.py --corpus <jsonl> --epochs 200 --no-export— expect PLCC > 0.95 on a multi-codec corpus.
0276 — vmaf_tiny_v5 corpus-expansion probe (ADR-0287) — defer¶
- What changed: research-only addition. New scripts under
ai/scripts/(fetch_youtube_ugc_subset.py,extract_ugc_features.py,train_vmaf_tiny_v5.py,eval_loso_vmaf_tiny_v5.py), new ADRdocs/adr/0276-*.md, new research digestdocs/research/0057-*.md, and one CHANGELOG entry. No new ONNX artefact undermodel/tiny/, no registry change, no public C-API / CLI / meson_options change. The probe trained an architecturally identical mlp_small on a 5-corpus parquet (4-corpus + 27 000 UGC rows); the 1-σ ship gate did not clear (Δ PLCC = +0.00005), so the exporter that the prior agent had drafted (export_vmaf_tiny_v5.py) was discarded before the commit. - Upstream source: fork-local. Netflix/vmaf has no tiny-AI corpus-expansion surface; nothing on the upstream side touches these files.
- On upstream sync: zero interaction. The v5 surface lives entirely under
ai/scripts/+docs/adr/+docs/research/, all of which are fork-introduced trees. The shipped v2 model (model/tiny/vmaf_tiny_v2.onnx) and its registry row are untouched. - Re-test on rebase:
# No code under test on rebase — purely research artefacts.
# If revisiting the corpus expansion, the reproducer is in the
# research digest:
python3 ai/scripts/fetch_youtube_ugc_subset.py \
--out-dir .workingdir2/ugc/download \
--n-stems 30 \
--manifest .workingdir2/ugc/manifest.json
python3 ai/scripts/extract_ugc_features.py \
--manifest .workingdir2/ugc/manifest.json \
--yuv-dir .workingdir2/ugc/yuv \
--vmaf-bin build-cpu/tools/vmaf \
--out-parquet runs/full_features_ugc.parquet \
--max-height 360 --max-frames 300 --threads 8
python3 ai/scripts/eval_loso_vmaf_tiny_v5.py \
--parquet-base runs/full_features_4corpus.parquet \
--parquet-extra runs/full_features_ugc.parquet \
--out-json runs/vmaf_tiny_v5_loso_metrics.json
0227 — vmaf-tune Intel QSV codec adapters (ADR-0281)¶
- What changed: fork-local additions under
tools/vmaf-tune/src/vmaftune/codec_adapters/—_qsv_common.py,h264_qsv.py,hevc_qsv.py,av1_qsv.py, plus registry rows incodec_adapters/__init__.pyand a new test filetools/vmaf-tune/tests/test_codec_adapter_qsv.py. Doc updates:docs/usage/vmaf-tune.md(Hardware encoders section),docs/adr/0281-vmaf-tune-qsv-adapters.md,docs/research/0066-vmaf-tune-qsv-adapters.md,tools/vmaf-tune/AGENTS.md,CHANGELOG.md. - Upstream source: fork-local.
tools/vmaf-tune/is fork-introduced under ADR-0237; Netflix/vmaf has no corresponding tree. - On upstream sync: zero interaction. Upstream cannot conflict with this PR's paths.
- Invariant: the registry exposes exactly four codecs (
av1_qsv,h264_qsv,hevc_qsv,libx264— alphabetical), each adapter validates its(preset, quality)pair, and the QSV preset vocabulary is the seven x264-style names (veryslow…veryfast, noultrafast/superfast). The encode pipeline (encode.py) remains x264-CRF-tied and will be widened in a separate PR — the QSV adapters are inert until then. Future codec families that share parameter shape (NVENC, AMF) follow the same_<family>_common.py+ N thin adapters pattern. - Re-test on rebase:
0229 — vmaf-tune libvvenc + NN-VC codec adapter (ADR-0285)¶
- Touches:
tools/vmaf-tune/src/vmaftune/codec_adapters/vvenc.py(new fork-only file),tools/vmaf-tune/src/vmaftune/codec_adapters/__init__.py(registry edit, fork-only),tools/vmaf-tune/tests/test_codec_adapter_vvenc.py(new),tools/vmaf-tune/tests/test_corpus.py(relaxes theknown_codecs() == ("libx264",)assertion to"libx264" in known_codecs()since the registry now spans multiple codecs). - Invariant: the codec-adapter registry is fork-introduced (Phase A of ADR-0237) and lives entirely outside the upstream Netflix tree, so
tools/vmaf-tune/does not touch upstream paths. The only rebase-sensitive surface is theCORPUS_ROW_KEYSschema insrc/vmaftune/__init__.py(per the Phase A invariant intools/vmaf-tune/AGENTS.md); this PR adds the adapter without changing the schema. - Upstream interaction: none.
tools/vmaf-tune/is not in Netflix/vmaf upstream. - Re-test on rebase:
- Status update 2026-05-09: the original
nnvc_intratoggle was removed (it emitted a fabricatedIntraNNkey that does not exist in any released VVenC). Replaced with a curated 9-knob real-VVenC 1.14.0 tuning surface (PerceptQPA,InternalBitDepth,Tier,Tiles,MaxParallelFrames,RPR,SAO,ALF,CCALF). Defaults preserve the bit-exact Phase A grid baseline.adapter_versionbumped to"2"so cache keys invalidate. See ADR-0285 §"Status update 2026-05-09".no rebase impact: REASON(fork-local file, no upstream-tree touch).
0228 — vmaf-tune Phase D scaffold (ADR-0276)¶
- Touches:
tools/vmaf-tune/src/vmaftune/per_shot.py,tools/vmaf-tune/src/vmaftune/cli.py,tools/vmaf-tune/tests/test_per_shot.py,docs/usage/vmaf-tune.md,docs/adr/0276-vmaf-tune-phase-d-per-shot.md. - Invariant: scaffold-only. The module relies on a stable predicate signature
(shot, target_vmaf, encoder) -> (crf, predicted_vmaf)that Phase B's bisect (PR #347) drops into later.Shotranges are half-open[start_frame, end_frame)even though the C-sidevmaf-perShotJSON/CSV sidecar uses an inclusiveend_frame— normalisation happens at the parse boundary in_parse_per_shot_json/parse_per_shot_csv.vmaf-perShotschema lives indocs/usage/vmaf-perShot.mdand is fork-local (ADR-0222), so upstream cannot drift it; the only rebase risk is fork-internal renames. - Upstream source: entirely fork-local.
tools/vmaf-tune/is fork-introduced (ADR-0237). Netflix/vmaf upstream has no encode-automation surface. - On upstream sync: zero interaction expected. No file in this PR overlaps an upstream-mirrored path.
- Re-test on rebase:
python -m pytest tools/vmaf-tune/tests/test_per_shot.py -q
python tools/vmaf-tune/vmaf-tune tune-per-shot --help
0229 — vmaf-tune SVT-AV1 codec adapter (ADR-0278)¶
- Touches:
tools/vmaf-tune/src/vmaftune/codec_adapters/svtav1.py(new),tools/vmaf-tune/src/vmaftune/codec_adapters/__init__.py(registry),tools/vmaf-tune/src/vmaftune/encode.py(parse_versionsextended for the SVT-AV1 banner pattern),tools/vmaf-tune/src/vmaftune/corpus.py(optionalffmpeg_preset_tokenhook). - Invariant:
PRESET_NAME_TO_INTis closed and order-stable; the integer values are baked into corpus rows that downstreamfr_regressor_v2(ADR-0235) trains on. Reordering or rewriting the table silently changes the integer SVT-AV1 receives. The codec key"libsvtav1"matchesCODEC_VOCAB[2]inai/src/vmaf_train/codec.py— keep them aligned on any rename. - Upstream source: fork-local.
tools/vmaf-tune/is a fork-introduced tree (see entry 0227 — Phase A scaffold). No Netflix/vmaf upstream interaction. - On upstream sync: zero interaction. Lives entirely under the fork-local
tools/vmaf-tune/tree. - Re-test on rebase:
0230 — fr_regressor_v2 PROD ship (ADR-0352)¶
- ADR: ADR-0352
- Touches:
model/tiny/fr_regressor_v2.onnx(binary, refreshed),model/tiny/fr_regressor_v2.json(sidecar, sha256 + metrics),model/tiny/registry.json(smoke flag flip, sha256 update),runs/phase_a/full_grid/per_frame_canonical6.jsonl(training corpus — fork-local artefact underruns/), companion docs. - Re-test recipe: see Research-0068 §Reproducer. Ship gate is LOSO PLCC ≥ 0.95 on the per-source folds; current run reports 0.9681 ± 0.0207.
- Rebase invariant: the per-frame canonical-6 corpus must be rebuilt from
runs/phase_a/{nvenc,qsv}_pf.jsonl(PR #392) before any retrain; do not re-train against the cell-onlycomprehensive.jsonl(it lacks the per-frame features and produces PLCC ≈ 0.7 — the smoke baseline). - No upstream interaction:
fr_regressor_v2is fork-local (ADR-0272).
0229 — vmaf-tune Phase E ladder generator (ADR-0295)¶
- ADR: ADR-0295
- Touches: entirely fork-local under
tools/vmaf-tune/. New moduletools/vmaf-tune/src/vmaftune/ladder.py, new test filetools/vmaf-tune/tests/test_ladder.py, two new subcommand blocks intools/vmaf-tune/src/vmaftune/cli.py. No upstream-shared paths touched. - Invariant:
vmaftune.ladder.convex_hullreturns a strictly monotonic Pareto frontier (both bitrate and vmaf monotonically increasing);select_kneesreturns exactlymin(n, len(hull))rungs in ascending bitrate order;emit_manifest("hls")produces one#EXT-X-STREAM-INFper rung with monotonically-increasingBANDWIDTH=values. The default_default_sampleris intentionallyNotImplementedError— production callers must inject a Phase B bisect-driven sampler. Phase B integration PR (gated on PR #347) swaps the default; the test suite continues to inject a synthetic stub. - Rebase impact: none — fork-local Python tool; upstream Netflix/vmaf does not ship a
tools/vmaf-tune/tree. - Re-test on rebase:
0229 — fr_regressor_v2 probabilistic head scaffold (ADR-0279)¶
- Touches:
ai/scripts/train_fr_regressor_v2_ensemble.py(new — fork-local).ai/scripts/eval_probabilistic_proxy.py(new — fork-local).model/tiny/fr_regressor_v2_ensemble_v1*.onnx,fr_regressor_v2_ensemble_v1.json(new artefacts; smoke probes).model/tiny/registry.json— five newkind: "fr"rows (fr_regressor_v2_ensemble_v1_seed{0..4}); existing entries untouched.ai/AGENTS.md— new "fr_regressor_v2_ensemble_v1 — probabilistic head" section pinning the per-member ONNX I/O contract, manifest-as-runtime-entry-point invariant, ensemble-size pin, confidence-rule one-of, codec-vocab parity, and smoke-artefact posture.docs/ai/models/fr_regressor_v2_probabilistic.md(new model card).docs/research/0067-fr-regressor-v2-probabilistic.md(new audit digest).docs/adr/0279-fr-regressor-v2-probabilistic.md(new ADR; Proposed). Index row appended todocs/adr/README.md.CHANGELOG.md—### Addedrow under "Unreleased — lusoris fork".- Invariant: the per-member ONNX I/O contract (two inputs:
features [N, 6]standardised +codec_onehot [N, NUM_CODECS]; one outputscore [N]) and the manifest'sconfidencerule (one-of"ensemble"/"ensemble+conformal") are the C-side adapter's load-bearing contract. Per-member ensembles are stockFRRegressor(num_codecs=NUM_CODECS)calls — flipping to a v1-shaped single-input graph silently invalidates the manifest.CODEC_VOCABparity withai/src/vmaf_train/codec.pyis required. - On upstream sync: zero interaction expected. Wholly fork-local; no upstream Netflix/vmaf path overlap. The
ai/package is fork-introduced (see ADR-0021, ADR-0036) — upstream has no probabilistic-regressor surface. If upstream ever ships its ownfr_regressor_v2variant, do NOT merge — register both ids side-by-side. - Re-test on rebase:
python ai/scripts/train_fr_regressor_v2_ensemble.py --smoke
python ai/scripts/eval_probabilistic_proxy.py --smoke
python ai/scripts/validate_model_registry.py
0287 — vmaf-tune saliency-aware ROI tuning (ADR-0293)¶
- Touches:
tools/vmaf-tune/src/vmaftune/saliency.py,tools/vmaf-tune/src/vmaftune/cli.py(newrecommendsubcommand),tools/vmaf-tune/AGENTS.md(saliency invariant),docs/usage/vmaf-tune.md(saliency section). - Upstream source: fork-local. The
vmaf-tunetree was introduced in PR #329 (ADR-0237 Phase A) and has no upstream Netflix counterpart. - On upstream sync: zero interaction — pure fork-local Python package under
tools/vmaf-tune/. - Invariant: the saliency-to-QP-offset signal blend (
offset = (2*sal − 1) * foreground_offset, clamped to ±12) is bit-for-bit equivalent tovmaf-roi's C-side blend (ADR-0247).tests/test_saliency.pypins the contract; ifvmaf-roi's C blend changes,saliency.pyfollows in the same PR. The test seam contract (session_factory=…,encode_runner=…) lets the suite run withoutonnxruntimeorffmpeg. - Re-test on rebase:
0229 — tools/vmaf-roi-score/ Option C scaffold (ADR-0296)¶
- ADR: ADR-0296
- Touches:
tools/vmaf-roi-score/pyproject.toml(new)tools/vmaf-roi-score/vmaf-roi-score(new console shim)tools/vmaf-roi-score/src/vmafroiscore/__init__.py(new)tools/vmaf-roi-score/src/vmafroiscore/cli.py(new)tools/vmaf-roi-score/src/vmafroiscore/score.py(new)tools/vmaf-roi-score/src/vmafroiscore/mask.py(new)tools/vmaf-roi-score/tests/test_combine.py(new)tools/vmaf-roi-score/README.md(new)tools/vmaf-roi-score/AGENTS.md(new)docs/adr/0296-vmaf-roi-saliency-weighted.md(new)docs/adr/_index_fragments/0296-vmaf-roi-saliency-weighted.md(new)docs/adr/_index_fragments/_order.txt— append-only.docs/research/0069-vmaf-roi-saliency-weighted.md(new)docs/usage/vmaf-roi-score.md(new)changelog.d/added/T6-2c-vmaf-roi-score-scaffold.md(new)- Invariant:
tools/vmaf-roi-score/is wholly fork-local. No upstream Netflix/vmaf surface owns or interacts with this directory. The combine math is a pure linear blend on Pythonfloat; the JSON schema is pinned byROI_RESULT_KEYSandSCHEMA_VERSION = 1. Schema bumps require an ADR-0288 supersession. Naming guard: do not confuse withcore/tools/vmaf_roi.c(ADR-0247) — that's the encoder-steering binary. The scoring tool here isvmaf-roi-score; the names diverge deliberately. - Rebase impact: zero. Pure-Python tool under
tools/; not part of the libvmaf C build, not part of any Netflix-mirrored surface. - Re-test on rebase:
0228 — vmaf-tune compare codec-comparison mode (research-0061 Bucket #7)¶
- Touches:
tools/vmaf-tune/src/vmaftune/compare.py(new). Wholly fork-local; no upstream Netflix/vmaf path overlap.tools/vmaf-tune/src/vmaftune/cli.py— adds thecomparesubparser and_run_comparerouter.tools/vmaf-tune/tests/test_compare.py(new). Mocked predicate; noffmpeg/vmafbinaries required.tools/vmaf-tune/AGENTS.md— invariant note for the predicate seam andCOMPARE_ROW_KEYScontract.docs/usage/vmaf-tune.md— new "Codec comparison" section.- Invariant:
compare.compare_codecsorchestrates per-codec ranking via an injectedpredicate(codec, src, target_vmaf) -> RecommendResultcallable. The orchestration must not branch on codec name; new codecs land as one-file additions undercodec_adapters/and are picked up automatically by the registry.COMPARE_ROW_KEYSis the JSON / CSV column contract — same maintenance discipline asCORPUS_ROW_KEYS. - Rebase impact: entirely fork-local. The Phase A + Phase B recommend backend (ADR-0237) is fork-internal; upstream Netflix/vmaf has no
tools/vmaf-tune/tree. - Re-test on rebase:
```shell pytest tools/vmaf-tune/tests/test_compare.py -v PYTHONPATH=tools/vmaf-tune/src python -m vmaftune.cli compare \ --src /tmp/ref.yuv --target-vmaf 92 --format markdown
0229 — vmaf-tune --score-backend GPU score wiring (ADR-0299)¶
- Touches:
tools/vmaf-tune/src/vmaftune/score_backend.py(new). Wholly fork-local —tools/vmaf-tune/has no upstream Netflix/vmaf overlap.tools/vmaf-tune/src/vmaftune/{score,corpus,cli}.py(additive kwargs, no API removals).tools/vmaf-tune/tests/test_score_backend.py(new).docs/usage/vmaf-tune.md(new GPU section + flag row).docs/adr/0299-vmaf-tune-gpu-score.md(new).docs/research/0071-vmaf-tune-gpu-score-backend.md(new).- Invariant: the libvmaf CLI exposes
--backend NAMEwith valuesauto|cpu|cuda|sycl|vulkanexactly. Help-text parser inscore_backend.parse_supported_backendspins this format. If upstream renames the flag or reformats the help line on merge, the parser silently degrades to "CPU only" — the test fixtures intest_score_backend.pywill catch the format change but only if re-run. - Upstream source: fork-local. Netflix upstream's CLI does not ship a
--backendselector (CPU-only). - On upstream sync: zero interaction.
vmaf-tunelives entirely in fork-introduced paths and consumes only the fork's--backendflag. - Re-test on rebase:
pytest tools/vmaf-tune/tests/test_score_backend.py -v
# If the libvmaf help text reformats, parse_supported_backends
# will return {"cpu"} on test_parse_full_backend_line_yields_all_four
# and the test fails loudly.
0261 — vmaf-tune HDR-aware encode + score path (2026-05-03)¶
- What changed: fork-local addition under
tools/vmaf-tune/src/vmaftune/hdr.pyplus wiring intocorpus.py/cli.py/score.py. Adds ffprobe-driven HDR detection, codec-specific HDR ffmpeg flag dispatch, schema-v2 corpus row keys (hdr_transfer,hdr_primaries,hdr_forced), and four--auto-hdr/--force-*CLI modes. See ADR-0300. - Upstream source: zero.
tools/vmaf-tune/is fork-introduced (Phase A under ADR-0237). - On upstream sync: zero interaction. Upstream Netflix/vmaf ships no encode automation surface; this tree is entirely fork-local and lives outside
libvmaf/andpython/. - Schema migration note:
SCHEMA_VERSIONbumped 1 → 2. The three new keys are additive — Phase B / C loaders treat missing keys as SDR for backward compat with v1 rows. - Re-test on rebase:
python -m pytest tools/vmaf-tune/tests/ -q
python -m vmaftune.cli corpus --help # confirm --auto-hdr surfaces
HP-2 — vmaf-tune HDR iter_rows integration (2026-05-08)¶
- What changed: fork-local.
tools/vmaf-tune/src/vmaftune/corpus.pynow importsvmaftune.hdrand wiresdetect_hdr/hdr_codec_args/select_hdr_vmaf_modelinto the per-source encode + score loop. The 0300 PR landedhdr.pyand the four CLI flags but never imported the module — PQ sources silently encoded as SDR. Schema bumps v2 → v3 because the originally-promisedhdr_transfer/hdr_primaries/hdr_forcedrow columns finally land. See ADR-0300 § Status update 2026-05-08. - Upstream source: zero. Fork-only.
- On upstream sync: zero interaction.
tools/vmaf-tune/is fork-introduced. - Schema migration note:
SCHEMA_VERSION2 → 3 (additive). The three HDR keys default to""/""/Falsefor SDR rows; Phase B / C loaders that ignore unknown keys keep working against v3 rows. - Re-test on rebase:
python -m pytest tools/vmaf-tune/tests/test_hdr.py -q
python -m pytest tools/vmaf-tune/tests/test_corpus.py::test_corpus_row_keys_match_init_contract -q
0298 — vmaf-tune content-addressed cache (ADR-0298)¶
- What changed: fork-local. New module
tools/vmaf-tune/src/vmaftune/cache.py; cache integration intools/vmaf-tune/src/vmaftune/corpus.py(iter_rowsnow consults the cache before encode/score); new CLI flags--no-cache,--cache-dir,--cache-size-gbincli.py. Codec-adapterProtocolgainsadapter_version: str; the lone Phase-A x264 adapter pins"1". - Upstream source: none.
tools/vmaf-tune/is fork-introduced (ADR-0237) and has no upstream counterpart. - On upstream sync: zero interaction with Netflix/vmaf master. The module sits entirely under
tools/vmaf-tune/, which upstream does not ship. - Invariant for future codec adapters: every
CodecAdaptermust declareadapter_version: str. Bump it whenever the adapter's argv shape, preset list, or quality range changes — otherwise the cache returns stale results post-upgrade. The contract is asserted bytest_cache_key_diffs_on_each_fieldintests/test_cache.py. - Re-test on rebase:
```bash pytest tools/vmaf-tune/tests/test_cache.py -v
0283 — vmaf-tune Apple VideoToolbox adapters (2026-05-05)¶
- What changed: fork-local addition under
tools/vmaf-tune/src/vmaftune/codec_adapters/. New files:h264_videotoolbox.py,hevc_videotoolbox.py,_videotoolbox_common.py, plus the registry hook in__init__.py. See ADR-0283. - Update 2026-05-09:
prores_videotoolbox.pyadapter added to the same registry pattern (broadcast / prosumer ProRes intermediate). Quality knob differs — ProRes is a fixed-rate codec, so the harness's--crfslot carries the integer ProRes tier id (0=proxy→ 5=xq) rather than a-q:vvalue._videotoolbox_common.pyextended withPRORES_PROFILE_*constants +validate_prores_videotoolbox()/prores_profile_name()helpers; profile ids verified against FFmpeg n8.1.1libavcodec/videotoolboxenc.c. See the Status update appendix in ADR-0283. - Upstream source: zero.
tools/vmaf-tune/is fork-introduced (Phase A under ADR-0237). - On upstream sync: zero interaction.
- Re-test on rebase:
python -m pytest tools/vmaf-tune/tests/test_codec_adapter_videotoolbox.py -q
python -m pytest tools/vmaf-tune/tests/test_codec_adapter_prores_videotoolbox.py -q
0228 — vmaf-tune coarse-to-fine CRF search (ADR-0306)¶
- What changed: fork-local tooling. Adds
coarse_to_fine_search()totools/vmaf-tune/src/vmaftune/corpus.py, plumbs new CLI flags ontovmaf-tune corpus(--coarse-to-fine,--coarse-step,--fine-radius,--fine-step,--target-vmaf), and ships a newvmaf-tune recommendsubcommand. Widenstools/vmaf-tune/src/vmaftune/codec_adapters/x264.pyquality_rangefrom(15, 40)to(0, 51). JSONL row schema unchanged (SCHEMA_VERSION=1). - Upstream source: fork-local. The whole
tools/vmaf-tune/tree is fork-introduced (ADR-0237); upstream Netflix/vmaf has no encode-automation surface. - On upstream sync: zero interaction.
tools/vmaf-tune/is not mirrored from upstream. - Re-test on rebase:
0314 — vmaf-tune --score-backend=vulkan (ADR-0314)¶
- Touches:
tools/vmaf-tune/src/vmaftune/cli.py(additive argparse flag oncorpus+recommendsubparsers; resolvesselect_backendand catchesBackendUnavailableErrorfor clean exit-2).tools/vmaf-tune/src/vmaftune/score.py(additivebackendkwarg onbuild_vmaf_commandandrun_score;None= no flag emitted).tools/vmaf-tune/src/vmaftune/corpus.py(newCorpusOptions.score_backendfield, defaultNone; forwarded intorun_score).tools/vmaf-tune/tests/test_score_backend.py(additive Vulkan-specific tests; pre-existing tests now pass after thebackend=kwarg lands).docs/adr/0314-vmaf-tune-score-backend-vulkan.md(new).docs/usage/vmaf-tune.md(new "Vulkan score backend" subsection under the existing GPU-scoring section).tools/vmaf-tune/AGENTS.md(invariant note: argparse choices stay in sync with libvmaf--backendvocabulary).changelog.d/added/vmaf-tune-score-backend-vulkan.md(new).- Invariant:
score_backend.ALL_BACKENDS = ("cpu", "cuda", "sycl", "vulkan")is the exact set libvmaf'score/tools/cli_parse.c--backendalternation accepts. Adding a new harness-side value without the libvmaf-side wiring produces silent strict-mode failures on hosts that probe positively for it. - Upstream source: zero. Netflix upstream's CLI does not ship a
--backendselector; bothtools/vmaf-tune/andcore/src/vulkan/are fork-introduced. - On upstream sync: zero interaction. No upstream-mirror file is touched.
- Re-test on rebase:
pytest tools/vmaf-tune/tests/test_score_backend.py -v -k vulkan
pytest tools/vmaf-tune/tests/test_score_backend.py -v
Failures here usually indicate the libvmaf help-text format changed; score_backend.parse_supported_backends test fixtures pin the format and will fail loudly.
0303 — fr_regressor_v2 ensemble prod flip (ADR-0303)¶
- ADR: ADR-0303
- Touches: entirely fork-local.
ai/scripts/train_fr_regressor_v2_ensemble_loso.py(new — 9-fold LOSO trainer over the five ensemble seeds; emitsloso_seed{N}.jsonartefacts).scripts/ci/ensemble_prod_gate.py(new — reads fiveloso_seed{N}.jsonfiles, returns exit 0 iffmean(PLCC_i) ≥ 0.95ANDmax - min ≤ 0.005).ai/AGENTS.md— appended "Ensemble registry invariant" paragraph under the existingfr_regressor_v2_ensemble_v1section.docs/adr/0303-fr-regressor-v2-ensemble-prod-flip.md(new),docs/research/0075-fr-regressor-v2-ensemble-prod-flip.md(new),changelog.d/added/fr-regressor-v2-ensemble-prod-flip.md(new).- Rebase invariant: the production ship gate is two-part —
mean_i(PLCC_i) ≥ 0.95ANDmax_i(PLCC_i) - min_i(PLCC_i) ≤ 0.005over five seeds. The variance bound is load-bearing: removing it silently allows a one-seed-wins-four-seeds-tie configuration that invalidates the ensemble's predictive-distribution semantics. Both thresholds live inscripts/ci/ensemble_prod_gate.py; do not weaken either without superseding ADR-0303. - Rebase invariant (registry): the five
fr_regressor_v2_ensemble_v1_seed{0..4}registry rows aresmoke: trueon master at this commit; flipping them tofalseis the follow-up flip PR's job, gated on a real-corpus LOSO run + the CI gate. Do not flip seed rows during a rebase merge conflict resolution. - Re-test on rebase:
python3 -c "import ast; ast.parse(open('ai/scripts/train_fr_regressor_v2_ensemble_loso.py').read())"
python3 -c "import ast; ast.parse(open('scripts/ci/ensemble_prod_gate.py').read())"
python ai/scripts/train_fr_regressor_v2_ensemble_loso.py --help
python scripts/ci/ensemble_prod_gate.py --help
- Upstream source: zero.
fr_regressor_v2and its ensemble are fork-introduced (parent ADR-0272 / ADR-0279). - On upstream sync: zero interaction.
0313 — CI required-checks aggregator (2026-05-05)¶
- What changed: fork-local CI policy. New
.github/workflows/required-aggregator.yml— single workflow that runs on every non-draft PR and verifies the 23 named required checks reportedsuccess/skipped/neutral(or didn't appear at all, which is the path-filter-rejection semantics). Aggregator becomes the single branch-protection required check, replacing the 23-name list from ADR-0037. - Touches:
.github/workflows/required-aggregator.yml(new),docs/adr/0313-ci-required-checks-aggregator.md(new),changelog.d/added/ci-required-checks-aggregator.md(new),docs/adr/README.md(+1 row),docs/adr/_index_fragments/_order.txt(+1 line + new fragment file). - Upstream source: zero. Branch-protection policy is fork-only.
- On upstream sync: zero interaction with Netflix/vmaf master.
- Manual operator step at adoption (uses PATCH, not PUT — corrected from the original ADR-0313 body which had the wrong verb):
echo '{"strict": false, "contexts": ["Required Checks Aggregator"]}' | \
gh api -X PATCH "repos/VMAFx/vmafx/branches/master/protection/required_status_checks" --input -
- Re-test on rebase:
# YAML lint passes
python3 -c "import yaml; yaml.safe_load(open('.github/workflows/required-aggregator.yml'))"
0305 — encoder knob-space Pareto analysis (2026-05-05)¶
- What changed: fork-local. New analysis scaffold for the 12,636-cell encoder knob sweep that backs
tools/vmaf-tune/codec_adapters/*recipe defaults. New files:ai/scripts/analyze_knob_sweep.py(per-(source, codec, rc_mode)Pareto hull on(bitrate_kbps, vmaf_score),encode_time_mstiebreaker, regression-detection check),ai/tests/test_knob_sweep_analysis.py(synthetic 20-row JSONL fixture). Methodology + scaffolded findings: see ADR-0305 + Research-0077. Companion to Research-0063. - Touches: none upstream-shared. Sits entirely under
ai/(fork-local since the tiny-AI training surface, ADR-0021) anddocs/{adr,research}/(fork ledger). - Upstream source: zero. The 12,636-cell sweep, the Pareto scaffold, and the regression-detection invariant are fork-introduced; Netflix/vmaf master ships no encoder knob-sweep tooling.
- On upstream sync: zero interaction with Netflix/vmaf master.
- Invariant for future codec adapter PRs: per the
ai/AGENTS.mdknob-sweep corpus invariant (ADR-0305), recipes that regress vs the bare encoder at matched bitrate within the same(source, codec, rc_mode)slice MUST NOT ship as adapter defaults. New adapter PRs cite the per-slice hull row fromreports/summary.md(or "no hull entry yet — bare default") in their PR description. Thecomprehensive.jsonlsweep file is generated locally and lives underruns/phase_a/full_grid/(gitignored — never committed). - Re-test on rebase:
0302 — ENCODER_VOCAB v3 schema expansion (ADR-0302)¶
- Touches:
ai/scripts/train_fr_regressor_v2.py(adds anENCODER_VOCAB_V3parallel constant; does not modify the liveENCODER_VOCABorENCODER_VOCAB_VERSION). - Invariant:
ENCODER_VOCABis append-only and order-stable (per ADR-0235). The v3 scaffold preserves the v2 slot ordering verbatim — slots 0..12 are bit-identical to the v2 vocab; slots 13/14/15 appendlibsvtav1,h264_videotoolbox,hevc_videotoolbox. The liveENCODER_VOCAB_VERSION = 2remains the source of truth until the follow-up retrain PR clears the LOSO PLCC ship gate. - Upstream interaction: zero.
ai/scripts/train_fr_regressor_v2.pyis fork-introduced (ADR-0272) and has no upstream counterpart. - Re-test on rebase:
python3 -c "
import importlib.util, pathlib
spec = importlib.util.spec_from_file_location(
't', pathlib.Path('ai/scripts/train_fr_regressor_v2.py')
)
m = importlib.util.module_from_spec(spec)
spec.loader.exec_module(m)
assert len(m.ENCODER_VOCAB_V3) == 16
assert m.ENCODER_VOCAB_VERSION == 2
print('OK')
"
0304 — vmaf-tune fast-path prod wiring (ADR-0304)¶
- Touches:
tools/vmaf-tune/src/vmaftune/fast.py(replaces the ADR-0276 scaffold'sNotImplementedErrorpaths with concrete Optuna TPE + v2 proxy + GPU verify wiring); new moduletools/vmaf-tune/src/vmaftune/proxy.py(centralised seam forfr_regressor_v2ONNX inference); expandedtools/vmaf-tune/tests/test_fast.py. Doc-side: ADR-0304, Research-0076,tools/vmaf-tune/AGENTS.mdinvariant note. - Upstream source: zero.
tools/vmaf-tune/andmodel/tiny/fr_regressor_v2.onnxare both fork-introduced (ADR-0237 / ADR-0352). - Invariant: the production proxy is always
fr_regressor_v2(no smoke models in the production path) and a single GPU verify pass at recommend-end is mandatory — proxy alone never wins. Thevmaftune.proxy.run_proxyhelper is the single seam every fast-path consumer goes through; future probabilistic-head / ensemble migrations land in that one module. ENCODER_VOCAB v2 one-hot ordering is frozen by ADR-0352 and pinned inproxy.ENCODER_VOCAB_V2— keep in sync withai/scripts/train_fr_regressor_v2.py; drift raisesProxyErrorat inference time before bad predictions ship. - On upstream sync: zero interaction with Netflix/vmaf master.
- Re-test on rebase:
0307 — vmaf-tune ladder default sampler wiring (ADR-0307)¶
- What changed: fork-local tooling.
tools/vmaf-tune/src/vmaftune/ladder.py::_default_samplerno longer raisesNotImplementedError; it composescorpus.iter_rows(Phase A encode + score) withrecommend.pick_target_vmaf(smallest CRF clearing target VMAF) overDEFAULT_SAMPLER_CRF_SWEEP = (18, 23, 28, 33, 38)at the adapter's mid-range preset. Module-level docstring + AGENTS.md invariant updated. New tests intools/vmaf-tune/tests/test_ladder.pystubiter_rowsviamonkeypatch.setattrso no live ffmpeg / vmaf binaries are needed. - Upstream source: fork-local. The whole
tools/vmaf-tune/tree is fork-introduced (ADR-0237); upstream Netflix/vmaf has no encode-automation / ladder surface. - On upstream sync: zero interaction.
tools/vmaf-tune/is not mirrored from upstream. - Rebase invariant: the 5-point sweep
(18, 23, 28, 33, 38)is the load-bearing default; downstream Phase E callers size their wall-time budget against five encodes per(resolution, target_vmaf)cell. Do not widen / narrow it without an ADR-0307 follow-up. TheSamplerFnseam stays open — callers needing finer grids pass an explicitsampler=. - Re-test on rebase:
0309 — fr_regressor_v2 ensemble real-corpus retrain harness (ADR-0309)¶
- ADR: ADR-0309
- Touches: entirely fork-local.
ai/scripts/run_ensemble_v2_real_corpus_loso.sh(new — Bash wrapper that loops the five seeds over the existingtrain_fr_regressor_v2_ensemble_loso.pyagainst.workingdir2/netflix/).ai/scripts/validate_ensemble_seeds.py(new — calls the ADR-0303 gate and writesPROMOTE.json/HOLD.jsonwith a corpus sha256 snapshot).ai/tests/test_validate_ensemble_seeds.py(new — 7 tests, synthetic JSON fixtures for both verdict paths).ai/AGENTS.md— appended "Registry-flip is a separate PR (ADR-0309)" paragraph under the existingfr_regressor_v2_ensemble_v1section.docs/adr/0309-fr-regressor-v2-ensemble-real-corpus-retrain.md,docs/research/0081-fr-regressor-v2-ensemble-real-corpus-methodology.md,docs/ai/ensemble-v2-real-corpus-retrain-runbook.md(all new).- Rebase invariant: the harness is decoupled from the registry mutation. Neither the wrapper nor the validator touches
model/tiny/registry.json; the registry flip is a separate follow-up PR gated on a passingPROMOTE.json. Auto-flipping on PROMOTE was rejected in ADR-0309's alternatives matrix specifically because rebase-time mutation of shipped registry rows is the foot-gun this invariant exists to prevent. - Re-test on rebase:
python -m pytest ai/tests/test_validate_ensemble_seeds.py -v
python ai/scripts/validate_ensemble_seeds.py --help
bash -n ai/scripts/run_ensemble_v2_real_corpus_loso.sh
- Upstream source: zero.
- On upstream sync: zero interaction.
0310 — BVI-DVC corpus ingestion for fr_regressor_v2 (ADR-0310)¶
- Touches:
ai/scripts/bvi_dvc_to_corpus_jsonl.py(new fork-only adapter),ai/scripts/merge_corpora.py(new fork-only shard merger),ai/tests/test_merge_corpora.py(new),docs/ai/bvi-dvc-corpus-ingestion.md(new),docs/adr/0310-bvi-dvc-corpus-ingestion.md(new),docs/research/0082-bvi-dvc-corpus-feasibility.md(new),ai/AGENTS.md(BVI-DVC invariant note). - Invariant: the BVI-DVC archive and any extracted artefacts (parquet, cached libvmaf JSON, JSONL corpus shard) are research-only and stay local — only derived
fr_regressor_v2_*.onnxweights ship. The merge utility validates every row against the canonicalvmaftune.CORPUS_ROW_KEYStuple; the schema is the merge contract. Re-shape here is a pure transform on the cached libvmaf JSON; no ffmpeg / vmaf binary is invoked. The(src_sha256, encoder, preset, crf)natural key is load-bearing for de-duplication across mirrors and re-encodes. - Upstream interaction: none.
ai/is fork-introduced; BVI-DVC is not part of Netflix/vmaf upstream. - Re-test on rebase:
ADR-0312 — ffmpeg-patches/ vmaf-tune integration (2026-05-05)¶
- Files:
ffmpeg-patches/0007-libvmaf-tune-qpfile-unified.patch,ffmpeg-patches/0008-add-libvmaf_tune-filter.patch,ffmpeg-patches/0009-pass-autotune-cli-glue.patch,ffmpeg-patches/series.txt,ffmpeg-patches/README.md. - Rebase invariant: patches
0007–0009plug into the cumulative state after patches0001–0006apply against pristinen8.1. Per-patchgit apply --checkin isolation is the wrong gate; use the series-replay command in CLAUDE.md §12 r14 instead. - vmaf-tune patch invariant: the qpfile parser at
libavcodec/qpfile_parser.{c,h}is shared across all three encoder adapters in patch 0007. Future encoders that grow a-qpfileAVOption inherit it; do not fork the parser. Whentools/vmaf-tune/src/vmaftune/saliency.py's qpfile output format changes (new column, different frame-type alphabet, …), patch 0007 must change in the same PR (CLAUDE.md §12 r14). - vf_libvmaf_tune full-scoring promotion (2026-05-06): patch 0008 originally shipped as a scaffold (linear CRF↔VMAF interpolation, no libvmaf scoring) per ADR-0312's deferred-alternatives column. The filter now mirrors
vf_libvmaf.c's CPU framesync pipeline end-to-end (vmaf_init+vmaf_model_load+vmaf_use_features_from_modelin init(); per-framevmaf_picture_alloc+ memcpy +vmaf_read_pictures; flush +vmaf_score_pooled(MEAN)in uninit()). The CRF recommendation remains a piece-wise linear projection from the observed VMAF; per-clip Optuna TPE search stays intools/vmaf-tune/src/vmaftune/recommend.py. Rebase-side: the new filter still depends only on libvmaf's CPU C-API (vmaf_init,vmaf_model_load,vmaf_use_features_from_model,vmaf_read_pictures,vmaf_score_pooled,vmaf_close,vmaf_picture_alloc/unref); zero new symbols beyond whatvf_libvmaf.calready requires, so future libvmaf rebases that pass the existing libvmaf filter pass this one too. ADR-0312 sub-decision retired. - n7+ API migration (2026-05-06): patch 0008 originally referenced the removed
AVFilterLink::frame_ratemember directly (n6-era API); in n7+ that field moved offAVFilterLinkonto a newFilterLinkstruct accessed viaff_filter_link(AVFilterLink *)fromlibavfilter/filters.h. Patch 0008 now usesff_filter_link(outlink)->frame_rate = ff_filter_link(mainlink)->frame_rate;inconfig_output(), mirroring patches 0005/0006 which were already written against the post-n7 API. The bug slipped through CI because the FFmpeg-Vulkan lane only buildsvf_libvmaf.o, notvf_libvmaf_tune.c; the full SYCL lane catches it now that PR #415 addedffmpeg-patches/**to the integration workflow's path filter. Discovery: PR #415 / ADR-0317. - Upstream source: zero. The vmaf-tune integration is fork-introduced; pure upstream syncs are unaffected.
- On upstream sync: zero interaction with libvmaf master. FFmpeg-side rebases when n8.1 → n8.x land in
ffmpeg-patches/test/build-and-run.sh'sFFMPEG_SHAare tracked separately under each refresh ADR (e.g., ADR-0277 for the 2026-05-04 refresh). - Re-test on rebase:
git -C /path/to/ffmpeg-8 reset --hard n8.1
for p in ffmpeg-patches/000*-*.patch; do
git -C /path/to/ffmpeg-8 am --3way "$p" || break
done
# Build smoke (libvmaf-disabled — patches 0001–0006 skipped if libvmaf_dnn
# is not built). With libvmaf_dnn available:
cd /path/to/ffmpeg-8 && ./configure --enable-libvmaf --enable-libx264 --enable-libsvtav1 --enable-libaom --enable-gpl
make -j$(nproc) ffmpeg
./ffmpeg -hide_banner -h encoder=libx264 2>&1 | grep -i qpfile
-
2026-05-06 update — patch 0007 SVT-AV1 ROI bridge promoted from scaffold to full impl: the libsvtav1 hunk now sets
enc_params.enable_roi_map = true, builds oneSvtAv1RoiMapEvtper qpfile frame upfront ineb_enc_init(per-MB qp_offsets averaged into per-64×64-SBb64_seg_mapof up to 8 segment QPs; uniform binning when the value span exceeds the segment budget), and attaches each event as aROI_MAP_EVENTpriv-data node fromeb_send_frame()withnode->size = sizeof(SvtAv1RoiMapEvt*)(the validation contract enforced by SVT-AV1'sresource_coordination_process.c). Lifetime invariant: events + maps live for the entire encode session because SVT-AV1 reads ROI_MAP_EVENT data via shallow-copied pointers on async pipeline threads (perenc_handle.c::copy_private_data_list);eb_enc_closefrees them. Wiring is gated onSVT_AV1_CHECK_VERSION(1, 6, 0); older SVT-AV1 builds keep the log-and-continue fallback. libaom remains scaffold-only — itsAOME_SET_ROI_MAPbridge stays a separate follow-up. No new ADR per CLAUDE.md §12 r8 (executes the existing ADR-0312 decision). -
2026-05-06 update — patch 0007 libaom-av1 ROI bridge promoted from scaffold to full impl: the libaom-av1 hunk now caches the parsed
VmafTuneQpFileinAOMContext, allocates a segment-id map at libaom's mode-info grid (ALIGN_POWER_OF_TWO(dim, 8) >> 2, sinceav1/common/enums.h::MI_SIZE == 4), and on every encoded frame picks up to 8 segment QPs from the per-frame qp_offset value range (uniform linear binning when the span exceedsAOM_MAX_SEGMENTS == 8), paints the per-mi segment map by expanding each per-16×16-MB qp_offset into a 4×4 block of mi cells, and issuesaom_codec_control(&ctx->encoder, AOME_SET_ROI_MAP, &roi_map). Lifetime invariant: libaom deep-copies the segment map anddelta_q[]table on every control call (perav1/encoder/encoder.c::av1_set_roi_map memcpy), so a single buffer is reused across frames and freed inaom_free(). The qpfile is also freed there. Trade-off: the 8-segment cap rounds nearby qp_offsets together when the saliency model emits more than 8 distinct values per frame; finer granularity requiresvmaf-tune corpusinstead. This retires the libaom-av1 deferral noted under ADR-0312 — both AV1 encoder hooks (libsvtav1 and libaom-av1) are now full-impl. No new ADR per CLAUDE.md §12 r8 (executes the existing ADR-0312 decision).
0315 — Vendor-neutral VVC encode strategy (ADR-0315 / Research-0085)¶
- ADR: ADR-0315
- Digest: Research-0085
- Touches: docs-only.
docs/research/0085-vendor-neutral-vvc-encode-landscape.md(new).docs/adr/0315-vendor-neutral-vvc-encode-strategy.md(new).docs/adr/_index_fragments/0315-vendor-neutral-vvc-encode-strategy.md(new).docs/adr/_index_fragments/_order.txt(one-line append).changelog.d/added/research-0085-vendor-neutral-vvc-encode.md(new).docs/rebase-notes.md(this entry).- Rebase invariant: none. The research digest and ADR are pure surveys with no code dependencies; nothing in the fork's source tree references them in a way that breaks on upstream rebase.
- Upstream source: zero. VVC encode strategy is a fork-local decision; upstream Netflix/vmaf has no codec adapter or encode-automation surface.
- On upstream sync: zero interaction. Pure docs.
- Re-test on rebase:
- 2026-05-06 follow-up (Research-0085 verification pass):
docs/research/0085-vendor-neutral-vvc-encode-landscape.mdflipped fromStatus: SKELETONtoStatus: Active. Most[UNVERIFIED]claims are now backed by primary-source URLs (NVIDIA SDK 13.0 docs, AMD AMF GitHub, Intel oneVPL GitHub +mfxstructures.h+CHANGELOG.md, Khronos registry, Phoronix Mesa/RADV coverage, VVenC issue tracker, ZLUDA repo).- ADR-0315's
## Contextand## Alternatives consideredrefreshed with the verified data points. Status staysProposed. [UNVERIFIED]count in the digest dropped 25 → 10; remaining items are legitimate gaps (NN-VC quality lift, vvenc per-kernel profile, HHI's non-public roadmap).- No code touched. No rebase impact beyond the existing docs-only posture.
0316 — cli_parse.c error() long-only-option fix (ADR-0316)¶
- ADR: ADR-0316 (follow-up to ADR-0311).
- Digest: none — bug-fix; fix shape fits in the ADR/commit body.
- Touches:
core/tools/cli_parse.c(3 lines — call-site arg change at theARG_THREADS/ARG_SUBSAMPLE/ARG_CPUMASKhandlers).core/test/fuzz/fuzz_cli_parse.c(removedknown_assert_in_inputearly-reject filter).core/test/fuzz/cli_parse_corpus/cli_threads_abbrev_assert.argv(promoted fromcli_parse_known_crashes/).core/test/test_cli_parse_long_only_args.c(new fork()-based regression test).core/test/meson.build(new test wiring, gated off Windows alongsidetest_y4m_411_oob).core/tools/AGENTS.md(added a long-only-options invariant note next to the existingcli_parse.crules).- Rebase invariant: load-bearing.
cli_parse.cis upstream-mirror with fork additions; the three handlers carry the fork-local shape of passing theARG_*enum value (not't'/'s'/'c') toparse_unsigned(). If an upstream sync re-introduces the original short-option char shape, the assert returns and the parked-then-promoted reproducer (cli_parse_corpus/cli_threads_abbrev_assert.argv) will surface it in the next nightly fuzz run. - Upstream source: the bug shape exists in Netflix/vmaf master too (long-only options were added upstream with the same short-option-char placeholder). When the fork ports an upstream fix that overlaps these handlers, prefer the
parse_unsigned(optarg, ARG_*, argv[0])form already on the fork. - On upstream sync: re-apply the three-line change in
cli_parse.cif upstream resets the call-site args. The unit test is fork-local and stays. - Re-test on rebase:
meson setup core/build libvmaf -Denable_tests=true \
-Denable_cuda=false -Denable_sycl=false
ninja -C core/build test/test_cli_parse_long_only_args
meson test -C core/build test_cli_parse_long_only_args -v
ADR-0317 — CI flake fix: doc-only PR path-filter (2026-05-06)¶
- Touched files:
.github/workflows/docker-image.yml— addedpaths:filter on bothpush:andpull_request:triggers..github/workflows/ffmpeg-integration.yml— addedpaths:filter on bothpush:andpull_request:triggers (covers all four matrix lanes: gcc, clang, SYCL, Vulkan).docs/adr/0317-ci-doc-only-pr-flake-fix.md,docs/adr/README.md(index row),changelog.d/fixed/ci-doc-only-pr-flakes.md.- Rebase invariant: not load-bearing. Workflow-only change. Both files are fork-local CI; upstream Netflix/vmaf does not ship a Docker workflow or an FFmpeg-integration matrix in this shape, so rebase conflicts are unlikely. If a future upstream sync introduces an overlapping
docker-image.ymlor FFmpeg matrix, prefer the fork's path-filtered form — the rationale (ADR-0313 aggregator posture, doc-only-PR runner-time burn) is fork-specific. - Upstream source: none — fork-local CI workflows.
- On upstream sync: no action required. If reviewers later add new build inputs (e.g. a top-level
docker-compose.yml, a newffmpeg-patches/*.txtconfig file), extend thepaths:lists in the same PR that adds the input. - Follow-up not in this ADR: patch
ffmpeg-patches/0008-add-libvmaf_tune-filter.patchline 256 (outlink->frame_rate = mainlink->frame_rate;) needs to migrate to theff_filter_link()accessor introduced in FFmpeg n7+, matching the pattern already in patches 0005 / 0006. Tracked separately; the path-filter does not hide it (any libvmaf/ or ffmpeg-patches/ PR will still trip the SYCL lane). - Re-test on rebase:
python3 -c "import yaml; \
yaml.safe_load(open('.github/workflows/docker-image.yml')); \
yaml.safe_load(open('.github/workflows/ffmpeg-integration.yml')); \
print('OK')"
0319 — fr_regressor_v2 ensemble LOSO trainer — real loader + per-fold training (ADR-0319)¶
- Touches:
ai/scripts/train_fr_regressor_v2_ensemble_loso.py(real_load_corpus+_train_one_seedbodies),ai/scripts/run_ensemble_v2_real_corpus_loso.sh(wrapper argv fix),docs/ai/ensemble-v2-real-corpus-retrain-runbook.md(Step 0 corpus-generation section),ai/AGENTS.md(canonical-6 schema invariant note),ai/tests/test_train_fr_regressor_v2_ensemble_loso_*.py(loader + train schema tests). Closes the deferrals tracked in rebase-notes §0303 + §0309. - Upstream source: none — fork-local ML training infrastructure. Netflix/vmaf upstream has no
fr_regressor_v2surface, no LOSO trainer, and no canonical-6 corpus tooling. - Invariant: the trainer's
_load_corpusaccepts the canonical-6 JSONL schema emitted byscripts/dev/hw_encoder_corpus.pybit-for-bit — required keys per row are(src, encoder, cq, frame_index, vmaf, adm2, vif_scale0..3, motion2). Codec block layout is 12-slotENCODER_VOCABv2 one-hot + constantpreset_norm = 0.5+crf_norm = (cq - cq_min) / (cq_max - cq_min). Schema changes require anENCODER_VOCAB_VERSIONbump and full ensemble retrain per the existing closed-vocabulary rule (ADR-0235 / ADR-0352). Fold-level StandardScaler is fit on the training rows only; leaking the held-out source's distribution into the scaler would silently inflate per-fold PLCC. - On upstream sync: no action required. If upstream Netflix/vmaf ever adds a competing LOSO trainer under
python/vmaf/, do NOT merge them — keep the fork's training stack underai/per the AGENTS.md scope rule. - Re-test on rebase:
pytest ai/tests/test_train_fr_regressor_v2_ensemble_loso_loader.py \
ai/tests/test_train_fr_regressor_v2_ensemble_loso_train.py -v
bash -n ai/scripts/run_ensemble_v2_real_corpus_loso.sh
ADR-0323 — fr_regressor_v3 train + register on ENCODER_VOCAB v3 (2026-05-06)¶
- Scope:
ai/scripts/train_fr_regressor_v3.py(new),ai/tests/test_train_fr_regressor_v3.py(new),model/tiny/fr_regressor_v3.onnx(new, real-weight checkpoint from a 9-fold LOSO gate-pass at mean PLCC 0.9975),model/tiny/fr_regressor_v3.json(new sidecar withencoder_vocab_version: 3and full per-fold trace),model/tiny/registry.json(newfr_regressor_v3row,smoke: false),ai/AGENTS.md(v3 retrain invariant section gains a "Status" subsection recording the gate result),docs/ai/models/fr_regressor_v3.md(new model card),docs/adr/0323-fr-regressor-v3-train-and-register.md+ index row,changelog.d/added/fr-regressor-v3-train-register.md. - Rebase impact: zero. Fork-local feature; no upstream Netflix/vmaf surface is touched. The 16-slot
ENCODER_VOCAB_V3imported fromtrain_fr_regressor_v2.pywas already landed by PR #401 (ADR-0302). - On upstream sync: no action required. The v3 model ships alongside v2 —
fr_regressor_v2.onnxand its sidecar are unchanged; the v3 row is appended to the registry and sorted alphabetically. If a future upstream sync ever lands a competingfr_regressor_v3model underpython/vmaf/, do NOT cross-link them — the fork's training stack lives underai/. - Watch out for: the live
ENCODER_VOCAB_VERSIONinai/scripts/train_fr_regressor_v2.pystays at 2 (per ADR-0302's invariant). Do not bump it to 3 in this PR or in any downstream port; the in-place promotion of v3 over v2 is a separate "promote v3 to authoritative" PR per ADR-0302's production-flip checklist. - Re-test on rebase:
pytest ai/tests/test_train_fr_regressor_v3.py -v
bash core/test/dnn/test_registry.sh # must report OK: 20+
python -c "import onnx; onnx.checker.check_model(onnx.load('model/tiny/fr_regressor_v3.onnx')); print('OK')"
ADR-0321 — fr_regressor_v2_ensemble_v1 full production flip (2026-05-06)¶
- Scope:
ai/scripts/export_ensemble_v2_seeds.py(new),model/tiny/fr_regressor_v2_ensemble_v1_seed{0..4}.onnx(real full-corpus-trained weights replacing the 3025-byte synthetic scaffold bytes),model/tiny/fr_regressor_v2_ensemble_v1_seed{0..4}.json(new per-seed sidecars),model/tiny/registry.json(sha256 +smoke: falseon the five seed rows),ai/AGENTS.md(new invariant: the registry-flip is now done; future re-flips require a fresh PROMOTE.json + re-run of the export driver). - Rebase impact: zero. This is a fork-local production-flip; no upstream Netflix/vmaf surface is touched. The 12-slot
ENCODER_VOCABv2 carried in each sidecar is the same one the LOSO trainer (ADR-0319) bakes into the codec-block layout, so there is no rebase-time vocabulary drift to worry about. - Watch out for: if a future upstream sync ever introduces a competing
fr_regressor_v2_ensemble_*model underpython/vmaf/, do NOT cross-link them — the fork's ensemble weights are gated onruns/ensemble_v2_real/PROMOTE.jsonand are not portable to a different training stack. - Re-test on rebase:
bash core/test/dnn/test_registry.sh # must report OK: 19
python -c "import onnx; \
[onnx.checker.check_model(onnx.load(f'model/tiny/fr_regressor_v2_ensemble_v1_seed{i}.onnx')) \
for i in range(5)]; print('OK')"
ADR-0324 — Ensemble training kit (2026-05-06)¶
- Touches:
tools/ensemble-training-kit/(new),docs/adr/0324-ensemble-training-kit.md(new),docs/adr/README.md(index row),changelog.d/added/0324-ensemble-training-kit.md(new). No engine code touched; no upstream-shared paths. - Invariant: the kit assumes the LOSO wrapper hard-codes seeds
(0 1 2 3 4). The orchestrator surfaces a warning if--seedsdeviates but still hands off to the wrapper. If a future PR parameterises the wrapper's seed list, update both the wrapper and the kit's pass-through logic in lockstep. - On upstream sync: no action required. The kit lives entirely under
tools/ensemble-training-kit/(a fork-local path) and only invokes other fork-local scripts (ai/scripts/,scripts/dev/,scripts/ci/). - Re-test on rebase:
bash -n tools/ensemble-training-kit/*.sh
bash tools/ensemble-training-kit/make-distribution-tarball.sh /tmp/kit-test.tar.gz
tar -tzf /tmp/kit-test.tar.gz | grep -q "tools/ensemble-training-kit/run-full-pipeline.sh"
ADR-0332 — External-competitor benchmark harness (2026-05-08)¶
- Touches:
tools/external-bench/(new),docs/adr/0332-external-bench-wrapper-only.md(new),docs/adr/_index_fragments/0332-external-bench-wrapper-only.md(new),docs/adr/_index_fragments/_order.txt(one-line append),docs/adr/README.md(regenerated),changelog.d/added/external-bench-harness.md(new),docs/research/0087-external-bench-competitor-survey-2026-05-08.md(new). No engine code touched; no upstream-shared paths. - Invariant: the harness is wrapper-only — never vendor or link
x264-pVMAF(GPL-2.0) into this fork. Future competitors follow the same pattern (tools/external-bench/<competitor>/run.shinvokes a user-installed binary via env var; output schema-shimmed into the canonical JSON shape). The output schema (frames[].{frame_idx, predicted_vmaf_or_mos, runtime_ms}+summary.{competitor, plcc, srocc, rmse, runtime_total_ms, params, gflops}) is the contract between every wrapper andcompare.py.run_wrapper'srunnerparameter MUST stay resolved at call time (not via default-arg binding) so monkeypatch-based tests work. - On upstream sync: no action required. The harness lives entirely under
tools/external-bench/(a fork-local path) and never touches Netflix-shared code. - Re-test on rebase:
python3 -c "import yaml; names=['docker-image','security-scans','lint-and-format','required-aggregator','ffmpeg-integration','libvmaf-build-matrix','rule-enforcement','tests-and-quality-gates']; [yaml.safe_load(open(f'.github/workflows/{n}.yml')) for n in names]; print('OK')"
# Spot-check the gate is present on every top-level job:
for f in docker-image security-scans lint-and-format ffmpeg-integration \
libvmaf-build-matrix rule-enforcement tests-and-quality-gates \
required-aggregator; do
grep -c "pull_request.draft == false" ".github/workflows/${f}.yml"
done # Each must report >= 1.
SSIM extractor registration fix (2026-05-08)¶
- Touches:
core/src/feature/feature_extractor.c(upstream-mirror — adds one extern + one registry-array entry near the existing SSIM rows),core/src/feature/integer_ssim.c(upstream-mirror — adds#include "config.h"and refreshes the file-scope comment abovevmaf_fex_ssim),core/src/meson.build(addsinteger_ssim.cto the source list — fork-local diff),core/test/test_feature_extractor.c(adds one regression test alongside the existing tests),docs/metrics/features.md(table row + footnote ²),docs/state.md,changelog.d/fixed/ssim-extractor-registration.md. - Invariant on the upstream-mirror files: the registry-array entry must remain inside the unconditional CPU block (the same block as
&vmaf_fex_float_ssim/&vmaf_fex_float_ms_ssim) —vmaf_fex_ssimis CPU-only with no SIMD or GPU twin. Theconfig.hinclude ininteger_ssim.cis load-bearing on Vulkan-enabled LTO builds becausefeature_extractor.candinteger_ssim.cmust agree onHAVE_VULKAN/HAVE_CUDA/HAVE_SYCLfor theVmafFeatureExtractorstruct layout to match across TUs. - On upstream sync: if Netflix ever lands its own integer-SSIM registry row, drop the fork's row in favour of upstream's; the file structure is identical. If upstream removes
integer_ssim.centirely (the file has been dormant on master for years), revert the meson.build addition. Otherwise no action. - Re-test on rebase:
meson setup build -Denable_cuda=false -Denable_sycl=false && ninja -C build
./build/test/test_feature_extractor # 5/5 pass, includes new ssim row
./build/tools/vmaf --reference testdata/ref_576x324_48f.yuv \
--distorted testdata/dis_576x324_48f.yuv \
--width 576 --height 324 --pixel_format 420 --bitdepth 8 \
--feature ssim --output /tmp/ssim_smoke.json && \
grep -q '<metric name="ssim"' /tmp/ssim_smoke.json
# Vulkan-enabled LTO build (-Wlto-type-mismatch must stay clean)
meson setup build-vulkan -Denable_vulkan=enabled --reconfigure && \
ninja -C build-vulkan tools/vmaf
python3 -m pytest tools/external-bench/tests/ -q # must report 7 passed
bash -n tools/external-bench/*/run.sh
0327 — Conformal-VQA prediction surface for vmaf-tune (ADR-0279)¶
- Touches:
tools/vmaf-tune/src/vmaftune/conformal.py(new),tools/vmaf-tune/src/vmaftune/predictor.py(Predictor.predict_vmaf_with_uncertainty),tools/vmaf-tune/src/vmaftune/cli.py(predictsubcommand gains--with-uncertainty/--calibration-sidecar/--alpha),tools/vmaf-tune/tests/test_conformal.py(new),docs/ai/conformal-vqa.md(new). No engine code touched; no upstream-shared paths. - Invariant: the conformal wrapper sits outside the ONNX graph and adds no new runtime dependency —
conformal.pyimports only the standard library (math,statistics,dataclasses,json,warnings). Future calibration-sidecar shapes use themethoddiscriminator string for versioning; do not rename"split-conformal"/"cv-plus"without bumping the loader. ThePredictor.predict_vmaf_with_uncertaintysignature is the Python-API contract consumed byvmaf-tune predict --with-uncertainty; renaming or reordering its keyword args breaks the CLI in lockstep. - On upstream sync: no action required.
vmaf-tuneis a fork-local tool; upstream Netflix/vmaf has no per-shot prediction surface. - Re-test on rebase:
python3 -m pytest tools/vmaf-tune/tests/test_conformal.py -q
python3 -m pytest tools/vmaf-tune/tests/test_predictor.py -q
CI paths-ignore deny-list on heavy workflows (ADR-0341, 2026-05-09)¶
- Touches:
.github/workflows/libvmaf-build-matrix.yml(fork-local —paths-ignore:block underpull_request:),.github/workflows/tests-and-quality-gates.yml(fork-local — same block),docs/adr/0341-ci-paths-ignore-doc-only-prs.md+ index fragment,changelog.d/changed/ci-paths-ignore-doc-only.md. - Invariant: the deny-list must stay strictly documentation-only (
docs/**,**/*.md,changelog.d/**,CHANGELOG.md,.workingdir2/**). Any path that contributes to a build, test, or lint input —libvmaf/**,meson.build,meson_options.txt,subprojects/**,python/**,ai/**,mcp-server/**,model/**,testdata/**,.github/workflows/**— must NEVER appear in the deny-list, otherwise the corresponding required check is silently skipped on a code-touching PR. The Required Checks Aggregator (ADR-0313) catches only the doc-only case (no required check ever ran for any required name); a too-broad deny-list would lose build coverage without anyone noticing. - On upstream sync: Netflix/vmaf upstream does not carry these two workflow files (they are fork-local additions). No sync conflict expected.
0332 — mkdocs --strict validation policy (ADR-0332)¶
- Touches:
mkdocs.yml(validation block +exclude_docs:),docs/mcp/embedded.md(one anchor fix),docs/research/0055-...md(one anchor fix),docs/{index,state,rebase-notes}.md(small bare-relative-dir-link sweep). All fork-local — no upstream-shared paths touched. - Upstream source: none — Netflix/vmaf upstream uses Sphinx / GitHub-rendered Markdown, not mkdocs. The
mkdocs.ymlconfig is wholly fork-local. - Invariant:
mkdocs.yml validation:must keeplinks.{not_found,unrecognized_links}: infountil either (a) ADR-0028 / ADR-0106 are superseded by a less-strict immutability rule that allows refreshing renamed-ADR cross-refs in frozen ADR bodies, or (b) the ~820 cross-tree-pointer links from docs into source-tree files (../../core/src/...,../../scripts/ci/...,../../.github/workflows/...) are migrated to absolute GitHub URLs or moved intodocs_dir-resident generated content. Promoting either category towarnwhile those conditions hold turns the docs lane permanently red. - On upstream sync: no action — the lane is fork-local.
- Re-test on rebase:
HDR VMAF model search — Path C documentation only (2026-05-09)¶
- Files added (this fork only; upstream Netflix/vmaf has none of these):
model/vmaf_hdr_model_card.md— discoverable warning that the HDR scoring path falls back to the SDRvmaf_v0.6.1.jsonweights. Filename deliberately uses.md, not.json, so thevmaftune.hdr.select_hdr_vmaf_modelglob (vmaf_hdr_*.json) keeps returningNone.docs/research/0089-hdr-vmaf-model-search.md— verbatim trail of the source-or-train survey (URLs + access dates).changelog.d/added/hdr-vmaf-model-search.md— release-notes fragment per ADR-0221.- ADR-0300 grew an inline
### Status update 2026-05-09: HDR model statussection. - Why no model JSON ships: Path A negative findings (no public Netflix HDR VMAF model exists; HDRMAX is a different algorithm not loadable by libvmaf's JSON path). Path B deferred behind gated subjective HDR corpora + multi-day training compute. No fabricated weights are introduced.
- On upstream sync: if Netflix lands
vmaf_hdr_*.jsoninNetflix/vmaf/model/, port via/port-upstream-commit; the resolver picks it up automatically with novmaftunechange. Then deletemodel/vmaf_hdr_model_card.md(or rewrite it as a normal model card describing the upstream weights). Watch https://github.com/Netflix/vmaf/issues/645 for the upstream release announcement. - Re-test on rebase: no behavioural change — pure docs. Sanity:
python3 -c "from pathlib import Path; \
import sys; sys.path.insert(0,'tools/vmaf-tune/src'); \
from vmaftune.hdr import select_hdr_vmaf_model; \
print(select_hdr_vmaf_model(Path('model')))"
# Expect: None — confirms the .md card does not match the glob
mkdocs build --strict # must EXIT=0 with no WARNING lines
ADR-0349 — fr_regressor_v3 namespace resolution (2026-05-09)¶
- Rebase impact: none. Docs-only change — adds ADR-0349, an append-only status appendix on ADR-0302 per ADR-0028, a
## fr_regressor_* namespace mapblock inai/AGENTS.md, and two changelog fragments. No upstream Netflix/vmaf surface touched; nofr_regressor_*registry rows touched (sha256s for_v1,_v2,_v2_ensemble_v1_seed{0..4},_v3all unchanged); no C / Python / ONNX bytes modified. - What to check after a rebase: nothing automated. The only drift risk is a future agent claiming
fr_regressor_v3plus_featuresfor an unrelated workstream —ai/AGENTS.mdcarries the reservation; reviewers verify the map row exists before approving any newfr_regressor_*registry id. - Reproducer:
```bash # ADR + AGENTS.md namespace map present and consistent: test -f docs/adr/0349-fr-regressor-v3-namespace.md grep -q "fr_regressor_* namespace map" ai/AGENTS.md grep -q "fr_regressor_v3plus_features" ai/AGENTS.md docs/adr/0349-fr-regressor-v3-namespace.md # Status appendix present on ADR-0302: grep -q "Status update 2026-05-09: namespace collision resolved" \ docs/adr/0302-encoder-vocab-v3-schema-expansion.md # Existing v3 production row bit-identical (sha256 unchanged): python3 -c "
import json reg = json.load(open('model/tiny/registry.json')) v3 = next(m for m in reg['models'] if m['id'] == 'fr_regressor_v3') assert v3['sha256'] == 'eaa16d23461eda74940b2ed590edfcaf13428aade294e47792a5a15f4d3b999c', v3 assert v3['smoke'] is False print('OK: fr_regressor_v3 production row unchanged') "
Registry test still passes:¶
bash core/test/dnn/test_registry.sh
0327 — Pre-push PR-body deliverables validator hook¶
- Touches:
scripts/ci/validate-pr-body.sh(new),scripts/git-hooks/pre-push(new),scripts/ci/test-validate-pr-body.sh(new),Makefile(hooks-installtarget adds the pre-push symlink). Re-usesscripts/ci/deliverables-check.shparser verbatim — no upstream-shared file is modified. - Invariant: parser shape parity with
.github/workflows/rule-enforcement.ymldeep-dive-checklist gate (ADR-0108). The validator constructs aPATHshim that interceptsgit diff --name-onlycalls only; every othergitinvocation falls through to the real binary. - On upstream sync: not applicable — these files are entirely fork-local and Netflix has no equivalent. If
scripts/ci/deliverables-check.shis ever rewritten or moved, the validator's exec path (scripts/ci/deliverables-check.sh) and the test harness's expected exit codes must follow. bash scripts/ci/test-validate-pr-body.sh # 8/8 cases pass
0320 — Semgrep # nosemgrep cites on Netflix-upstream Python harness (Research-0090)¶
- Touches:
python/vmaf/core/asset.py,python/vmaf/core/executor.py,python/vmaf/core/feature_extractor.py,python/vmaf/core/quality_runner.py,python/vmaf/core/result_store.py,python/vmaf/tools/decorator.py,python/test/command_line_test.py,python/test/feature_extractor_test.py,python/test/ssimulacra2_test.py,python/vmaf/config.py. - Invariant: every fork-added
# nosemgrep: <rule-id>line is paired with an inline cite toResearch-0090. The cite + rule-id pair is the load-bearing artifact (per memoryfeedback_no_guessing: every "false positive" claim ships its safety proof). If an upstream sync removes the cited line of code, drop the cite-comment block too. If upstream adds adefusedxmlfix at theElementTree.parse()site (feature_extractor.py:115,quality_runner.py:1496), keep upstream's fix and drop our suppressions. config.py:40(the SSL-bypass deletion) is a fork-exclusive security fix; if upstream resurrectsssl._create_unverified_contexton a sync, do not re-merge it — the bypass clobbers the process-global default and is unjustified per Research-0090, F1. semgrep scan --config=p/cwe-top-25 --config=p/c --config=p/python . \ --metrics=off --json | jq '.results | length'
# expect 0 — every legit finding either has a # nosemgrep cite or was fixed
0321 — Security-scans workflow registry-pack list (Research-0090)¶
- Touches:
.github/workflows/security-scans.yml,.github/workflows/lint-and-format.yml. - Invariant: the registry packs the workflow cites (
p/cwe-top-25+p/c+p/python) are validated againsthttps://semgrep.dev/c/p/<pack>— the previously-citedp/cert-c-strict,p/cert-cpp-strict, andp/cpppacks were retired by Semgrep in 2025 and 404. Thelint-and-format.ymlpull of${{ github.* }}intoenv:(clang-tidy + clang-tidy-sycl steps) defusesrun-shell-injection; preserve the pattern on any edit. See Research-0090, F2/F3. for pack in p/cwe-top-25 p/c p/python; do code=\((curl -sIL "https://semgrep.dev/c/\)" | head -1 | awk '{print $2}') [ "$code" = "200" ] && echo "\({pack}: OK" || echo "\): FAIL ($code)"
0320 — CodeQL C bulk sweep (78 deferred alerts → 60 fixed, 14 deferred to T7-5)¶
- Touches:
core/src/feature/{cambi.c,ciede.c,integer_adm.c,integer_psnr.c,adm_tools.h,third_party/xiph/psnr_hvs.c},core/src/feature/x86/{adm_avx2.c,adm_avx512.c,ansnr_avx2.c,ansnr_avx512.c,vif_avx2.c,vif_avx512.c},core/src/{pdjson.c,svm.cpp},core/test/{test_cpu.c,test_model.c},core/tools/{y4m_input.c,yuv_input.c,vmaf_bench.c}. All butvmaf_bench.care upstream-mirror Netflix files. - Invariant: widening casts on integer multiplications (
(size_t),(uint64_t),(double)) are LHS-prefixed before the multiply, never wrapped around the whole expression — the latter is a no-op againstcpp/integer-multiplication-cast-to-long. Deleted commented-out blocks (e.g., the AVX-512 VP-loop dead variant inadm_avx512.c::adm_dwt2_inverse) are gone for good; if upstream brings them back, they reintroduce the alerts.iqa/convolve.cwas deliberately left untouched: prefixing(double)on the float×float multiplications inside the scalar reference path breaks bit-exactness against the AVX2 path enforced bytest_iqa_convolve— CodeQL alert deferred to a follow-up that updates both paths in lockstep. - On upstream sync: any upstream change that re-introduces the deleted comment blocks or rewrites the cast forms will surface the alerts again. The
cambi_scoresignature change (CambiBuffers buffers→const CambiBuffers *buffers) is fork-local and likely to conflict with upstream patches that touch that function. The 14 deferredVifBufferlarge-parameter alerts are tracked under T7-5 (multi-backend coordinated refactor including NEON). - Re-test on rebase: cd libvmaf && meson test -C build # all 50+ C tests make test-netflix-golden # upstream golden gate
# Re-run CodeQL on master afterwards; the 60 fixed alerts must stay closed.
CodeQL cpp/declaration-hides-variable sweep (2026-05-09)¶
- What changed: Mechanical rename / scope-tighten / dedupe sweep closing 64 open
cpp/declaration-hides-variableCodeQL alerts onmaster. Touched files:core/src/feature/cambi.c,core/src/feature/x86/adm_avx2.c,core/src/feature/x86/adm_avx512.c,core/src/feature/x86/vif_avx2.c,core/src/feature/x86/vif_avx512.c. All five are upstream-mirror; the Netflix copyright header is preserved on each. - Renames adopted (semantic over
_2suffix): cambi.c: innerint errshadowing function-scopeerrbecomesmkdir_err(heatmaps init) andsrc_err(full-ref extract path).adm_avx2.c/adm_avx512.c: thej == 0first-column special-case block is wrapped in{ ... }so itsj0..j3ands0..s3stop being visible to the per-jtail loop. The inner duplicate__m256i add_shift_HP_vex = _mm256_set1_epi32(32768)(and 512-bit twin) is removed — bit-identical to the function-scope value already in scope. The__m256i rfactor1that shadowed the function-scopefloat rfactor1[3]becomesrfactor_v0/_v1/_v2(and the AVX-512 twin likewise).vif_avx2.c/vif_avx512.c: tap-loop locals followf_tap,r_top/r_bot,d_top/d_botfor the s0 stage, andf_tap0/f_tap1,r_back0/r_fwd0, etc. for the AVX-512 paired-tap stage. Inner per-fj__m256i fq/__m512i fqshadows of the centre-tap broadcast becomef_tap. Inner-block duplicates of function-scoperef/dis/stride/ii(identical types and initialisers) are simply removed. The two scalarVifResiduals residualsdeclarations that shadowed function-scopeResiduals512 residualsbecometail_residuals. The twoconst uint16_t fcoeffdeclarations that shadowed function-scope__m512i fcoeffbecomefcoeff_scalar.- Invariant: bit-exactness gate — the rename sweep must not change any score. The Netflix CPU golden 3 (
src01_hrc00,checkerboard_1,checkerboard_10) ran clean against this PR. All 76 VMAF-targeted Python tests pass; the 9 unrelated pre-existing failures (NIQE, PyPSNR, FileSystemResultStore) reproduce on a pristineorigin/mastercheckout. - On upstream sync: Netflix has no equivalent renames on upstream
masteras of2026-05-09. When syncing, prefer the fork's renamed identifiers (the CodeQL gate depends on them). If Netflix later renames the same locals differently, reconcile by keeping fork names and updating any imported chunks at port time. - Re-test on rebase: meson test -C build --suite=fast PYTHONPATH=$PWD/python python3 -m pytest \ python/test/quality_runner_test.py -k test_run_vmaf \ python/test/vmafexec_test.py \ python/test/vmafexec_feature_extractor_test.py \ -m "not slow" -q
ADR-0209 v1 stdio runtime (T5-2b) — Embedded MCP server (2026-05-08)¶
- Touches:
core/src/mcp/{mcp.c,dispatcher.c,transport_stdio.c,mcp_internal.h,meson.build,3rdparty/cJSON/{cJSON.c,cJSON.h,LICENSE}},core/test/test_mcp_smoke.c,core/test/meson.build. All paths are fork-local. cJSON is vendored verbatim from upstreamDaveGamble/cJSON@v1.7.18under its MIT license. - Invariant: every TU under
core/src/mcp/(other than the vendored cJSON dir) is fork-local with theCopyright 2026 Lusoris and Claude (Anthropic)header; cJSON keeps its upstream MIT header verbatim. The public ABI incore/include/libvmaf/libvmaf_mcp.his unchanged from T5-2 — only function bodies flipped from-ENOSYSto working implementations. SSE / UDS still return-ENOSYSso the v2 PR can wire them without touching the public surface. - On upstream sync: no action required. Netflix/vmaf upstream has no embedded MCP surface; the entire
core/src/mcp/subtree is fork-local. If upstream ever adds an MCP surface, expect a port-only sync since names will collide. cd libvmaf && meson setup build -Denable_cuda=false -Denable_sycl=false \ -Denable_mcp=true -Denable_mcp_stdio=true ninja -C build && meson test -C build test_mcp_smoke -v
ADR-0334 — state.md-touch-check CI gate (2026-05-08)¶
- Touches:
.github/workflows/rule-enforcement.yml(new top-level jobstate-md-touch-check),scripts/ci/state-md-touch-check.sh(new),scripts/ci/test-state-md-touch-check.sh(new),scripts/ci/AGENTS.md(new rebase-sensitive-surface row),.github/PULL_REQUEST_TEMPLATE.md(already carries the "Bug-status hygiene" section +no state delta: REASONopt-out — coupled to the script's regex). No upstream-shared paths. - Invariant: the gate's trigger predicate (Conventional-Commit
fix:prefix, barebugtoken in title, GitHub close-keywordscloses/fixes/resolves#N, unchecked Bug-status-hygiene checkbox) and opt-out sentinel (no state delta: REASON) match the wording of the## Bug-status hygienesection in.github/PULL_REQUEST_TEMPLATE.md. Reword the template only alongside the script. The job carries thepull_request.draft == false || github.event_name != 'pull_request'gate (ADR-0331 pattern) — keep that on any future hoist into the required-aggregator set. - On upstream sync: Netflix/vmaf has no equivalent rule. No conflict expected; the workflow file is fork-introduced.
- Re-test on rebase: bash scripts/ci/test-state-md-touch-check.sh python3 -c "import yaml; yaml.safe_load(open('.github/workflows/rule-enforcement.yml')); print('YAML OK')" pre-commit run shellcheck --files scripts/ci/state-md-touch-check.sh scripts/ci/test-state-md-touch-check.sh pre-commit run shfmt --files scripts/ci/state-md-touch-check.sh scripts/ci/test-state-md-touch-check.sh
SYCL PSNR chroma extension (T3-15(b), 2026-05-09)¶
- Touches:
core/src/feature/sycl/integer_psnr_sycl.cpp(per-extractor chroma device buffers, per-plane SSE accumulators, and aprovided_featuresextension topsnr_y/psnr_cb/psnr_cr),core/src/sycl/AGENTS.md(per-kernel rebase-sensitive invariant for the chroma-on-per-extractor-buffer arrangement),docs/metrics/features.md(footnote ¹ refresh — all three GPU PSNR extractors now emit chroma),docs/adr/0192-gpu-long-tail-batch-3.mdReferences-section status update,changelog.d/added/sycl-psnr-chroma.md. - Invariant on the chroma upload path: chroma planes ride on per-extractor device buffers populated by host-side staging copies in the combined-graph
pre_fncallback — NOT the SYCL state's shared frame buffer (vmaf_sycl_shared_frame_init), which is luma-only by design. Luma stays graph-recorded; chroma SSE kernels run direct inpost_fnon the same in-order combined queue. The CUDA twin (PR #520 / commit 7f3d58a5) uses the existing CUDA per-plane picture infrastructure and therefore has no equivalent invariant. - On upstream sync: Netflix/vmaf upstream has no SYCL backend at all, so conflict probability is zero on
psnr_sycl. If an upstream port to the fork's SYCL runtime someday extendsvmaf_sycl_shared_frame_initto allocate chroma planes, the PSNR extension can be migrated onto it and the per-extractor chroma buffers retired — but only after a cross-backend gate run confirms bit-exactness against CPU atplaces=4(ADR-0214). source /opt/intel/oneapi/setvars.sh CC=icx CXX=icpx meson setup build-sycl libvmaf \ -Denable_sycl=true -Denable_cuda=false ninja -C build-sycl python3 scripts/ci/cross_backend_vif_diff.py \ --vmaf-binary build-sycl/tools/vmaf \ --reference testdata/ref_576x324_48f.yuv \ --distorted testdata/dis_576x324_48f.yuv \ --width 576 --height 324 --pixel-format 420 --bitdepth 8 \ --feature psnr --backend sycl --device 0
# Expect 0/48 mismatches across psnr_y / psnr_cb / psnr_cr at places=4.
```text
Cppcheck nullPointer false-positive in dict.c (2026-05-09)¶
Files pinned:
core/src/dict.c:121(one-line redundant-condition fix indict_overwrite_existing). Why this rebase-note exists: Master CI'sCppcheck (Whole Project)gate started failing on commit14b5ffba(#537) and blocked every open PR because each PR rebases onto a broken master. The cppcheck finding was likely always present but masked bypaths-ignorefiltering on the prior workflow shape; PR #530 widened cppcheck's trigger surface and exposed it. Deleted the redundant&& valguard sincevalis already checked at the public entry-pointvmaf_dictionary_set(dict.c:137). No behavior change; cppcheck flags the original as "either the val check is redundant or there's a possible null deref" because it can't prove the interprocedural guarantee. Rebase-sensitivity: zero — change is local todict.c. Future upstream sync of this file should keep the fix or re-run cppcheck locally to confirm absence of recurrence.
Aggregator timeout bump (2026-05-09)¶
Files pinned:
.github/workflows/required-aggregator.yml(deadline 30→90 min, job timeout 35→100 min) Why: 41 PRs in flight 2026-05-09 morning hit Aggregator timeouts while real CI eventually passed. Bumping both deadlines unblocks the train without touching the underlying matrix. Rebase-sensitivity: zero — workflow file is wholly fork-local.
ARC self-hosted runner pool — pilot Cppcheck routing (2026-05-09)¶
.github/workflows/lint-and-format.yml(Cppcheckruns-on:ternary). Why: opt-in graceful migration; ADR-0359 + docs/development/ci-runners.md document the flip-the-variable recipe when the cluster is degraded. Rebase-sensitivity: zero — workflow file is fork-local.
ADR-0338 — macOS Vulkan-via-MoltenVK CI lane (2026-05-09)¶
- Touches:
.github/workflows/libvmaf-build-matrix.yml(fork-local — addsBuild — macOS Vulkan via MoltenVK (advisory)lane, addscontinue-on-errorplumbing onmatrix.experimental && matrix.moltenvk, addsInstall MoltenVK + Vulkan loader/headers (macOS)step, addsRun Vulkan smoke tests (macOS MoltenVK)step, gates the existing test/cache/tox steps on!matrix.moltenvk),docs/backends/vulkan/moltenvk.md(new fork-local doc),docs/adr/0127-vulkan-compute-backend.md(status-update appendix per the ADR's Proposed status — body untouched),docs/adr/0338-macos-vulkan-via-moltenvk-lane.md(new),docs/adr/_index_fragments/0338-macos-vulkan-via-moltenvk-lane.mdplus_order.txtappend (new),docs/research/0089-moltenvk-feasibility-on-fork-shaders.md(new),changelog.d/added/macos-vulkan-via-moltenvk-lane.md(new). - Invariant on the upstream-mirror file: none —
libvmaf-build-matrix.ymlis fork-local. The new lane'scontinue-on-errorclause MUST stay scoped tomatrix.experimental == true && matrix.moltenvk == trueso existingexperimental: truematrix entries (e.g. the macOS DNN lane) keep their default fail-fast behaviour.VK_ICD_FILENAMESMUST point at/opt/homebrew/etc/vulkan/icd.d/MoltenVK_icd.json— note theetc/vulkansegment, NOTshare/vulkan(the homebrew formula's install layout usesetc/; verified againstFormula/m/molten-vk.rb). - On upstream sync: Netflix upstream has no macOS Vulkan lane and no MoltenVK awareness; nothing to reconcile. If a future MoltenVK release drops support for
GL_EXT_shader_atomic_int64translation,moment.compwill fail on the lane; the fix path is in ADR-0338 §Decision (lane iscontinue-on-errorso it does not block PRs) — update the known-limitations table indocs/backends/vulkan/moltenvk.mdand either pin a working MoltenVK version in the brew install line or rewrite the shader. - Re-test on rebase:
python3 -c "import yaml; yaml.safe_load(open('.github/workflows/libvmaf-build-matrix.yml'))" && \
echo "YAML parse OK"
# Confirm the lane is still in the matrix:
grep -q "Build — macOS Vulkan via MoltenVK (advisory)" \
.github/workflows/libvmaf-build-matrix.yml
# Confirm the lane is NOT promoted to required-aggregator until one
# green run on master (per ADR-0338):
! grep -q "macOS Vulkan via MoltenVK" \
.github/workflows/required-aggregator.yml
# Confirm the ICD path is the etc/ one, not share/:
grep -q "etc/vulkan/icd.d/MoltenVK_icd.json" \
.github/workflows/libvmaf-build-matrix.yml
ADR-0363 — Mend Renovate replaces Dependabot (2026-05-09)¶
- Touches:
renovate.json(new, repo-root),.github/workflows/renovate.yml(new),.github/dependabot.yml(deleted — renamed to.github/dependabot.yml.disabled),docs/development/dependency-bot.md(new operator playbook),changelog.d/changed/renovate-supersedes-dependabot.md(new),docs/adr/0363-renovate-replaces-dependabot.md(new),docs/adr/_index_fragments/0363-renovate-replaces-dependabot.md(new). - Invariant:
.github/dependabot.ymlno longer exists onmaster; the disabled copy isdependabot.yml.disabled. On upstream sync, if Netflix ever ships their owndependabot.yml, do NOT restore it — the fork intentionally uses Renovate. Merge the upstream file intodependabot.yml.disabledfor reference only. - Upstream interaction: none. Netflix/vmaf upstream has no Renovate config. Conflict risk is zero unless upstream adds
renovate.jsonor restoresdependabot.yml. - Re-test on rebase:
# Verify the workflow SHA-pin is still present and non-floating:
grep -E 'renovatebot/github-action@[a-f0-9]{40}' .github/workflows/renovate.yml
# Verify dependabot.yml is still absent:
test ! -f .github/dependabot.yml && echo "ok: dependabot.yml absent"
# Validate renovate.json syntax (requires Node):
node -e "JSON.parse(require('fs').readFileSync('renovate.json','utf8')); console.log('JSON valid')"
ADR-0355 — Symphony-inspired agent-dispatch infrastructure (2026-05-09)¶
Files added (all fork-introduced, none mirror upstream):
.claude/workflows/_template.md,.claude/workflows/codeql-alert-sweep.md,.claude/workflows/simd-port.md,.claude/workflows/feature-extractor-port.md.scripts/lib/__init__.py,scripts/lib/backlog_tracker.py,scripts/lib/AGENTS.md.scripts/ci/agent-eligibility-precheck.py(new row inscripts/ci/AGENTS.md"Rebase-sensitive surfaces" table).docs/development/agent-dispatch.md. Why this rebase-note exists: pure additive, all paths are fork-only (.claude/,scripts/lib/, fork-only docs). Upstream Netflix/vmaf has no.claude/, noscripts/lib/, and nodocs/development/agent-dispatch.md, so the merge surface is zero on/sync-upstream. The only coupling is internal betweenscripts/ci/agent-eligibility-precheck.pyandscripts/lib/backlog_tracker.py(sys.path import). Both files move together; documented inscripts/lib/AGENTS.mdand a new row inscripts/ci/AGENTS.md. Rebase-sensitivity: zero w.r.t. upstream. Internal-only: renamingBacklogItemfield names or theBacklogTracker/GitHubTrackerpublic method signatures is a breaking change for the precheck and any future state-audit script — guard via the smoke listed in Research-0091 §"Smoke results" before any rename PR. Format-coupling note: the BACKLOG.md row regex (scripts/lib/backlog_tracker.py:_ID_PATTERN) is brittle against table-shape edits. If a future BACKLOG.md edit adds a column or renames a status word, the parser will silently mis-classify rows — the smoke parses 101 rows on master at 2026-05-09; expect ≥ 100 after any structural edit.
0350 — psnr_hvs AVX-512 ceiling re-bench (ADR-0350, T3-9 (a))¶
docs/adr/0350-psnr-hvs-avx512-ceiling.md— closure ADR.docs/adr/0160-psnr-hvs-neon-bitexact.md— appended### Status update 2026-05-09appendix.docs/research/0091-psnr-hvs-avx512-bench-2026-05-09.md— empirical companion (cycle share, Amdahl ceiling, reproducer). Why this rebase-note exists: T3-9 (a) closes as AVX2 ceiling. The result has zero rebase-sensitivity by itself — no engine code changes — but the bit-exactness invariants that lock it to a ceiling do. The 78.42 % scalar tail incalc_psnrhvs_avx2/calc_psnrhvs_neonis locked by ADR-0138 / ADR-0139's "per-lane-scalar float reduction" rule (carried by ADR-0159 / ADR-0160). If a future upstream sync ofcore/src/feature/third_party/xiph/psnr_hvs.c(the Xiph/Daala DCT) changes the per-block summation tree — e.g. partial folding, re-ordered means, vectorised mask reductions — the AVX2 + NEON TUs incore/src/feature/x86/psnr_hvs_avx2.candcore/src/feature/arm64/psnr_hvs_neon.cMUST be re-audited against the new scalar reference, and the ceiling argument in ADR-0350 must be re-run (because the 78 / 15 cycle-share split would shift). Rebase-sensitivity: low for the ceiling decision itself (empirical re-bench on a current host is cheap — 30 seconds via the reproducer in Research-0091 §7); high for the underlying bit-exactness invariants the decision rests on (Netflix golden trips on ≥ 5.5e-5 drift per ADR-0160 §Context). The ADR-0350 §Verification reproducer is the gate — re-run it if the cycle share shifts, the Netflix normal-pair fixture changes, or a new host class (e.g. wide-issue Granite Rapids) goes into CI.
0320 — FFmpeg n8.1 → n8.1.1 base bump (2026-05-09)¶
- Touches:
ffmpeg-patches/series.txt(header comment),ffmpeg-patches/README.md(apply / verify / smoke sections),ffmpeg-patches/test/build-and-run.sh(FFMPEG_SHAdefault),scripts/ci/ffmpeg-patches-check.sh(header comment;FFMPEG_BRANCHenv default unchanged atrelease/8.1since the branch tracks point releases),docs/development/automated-rule-enforcement.md(gate description). The 9.patchfiles themselves are unchanged — every patch in the series applied cleanly, cumulatively, against pristinen8.1.1viagit am --3way. - Upstream source: FFmpeg upstream point release n8.1.1 (commit
239f2c7"Bump micro for 8.1.1") — bug-fix-only on top of n8.1, no API or AVOption breakage that the patch stack consumes. - Invariant: the patch stack continues to apply against the current tip of FFmpeg's
release/8.1branch. Per ADR-0118 and ADR-0186 §FFmpeg patch coupling, the verification gate is cumulativegit am --3wayagainst a pristine checkout, not per-patch standalone apply. The scripts/ci/ffmpeg-patches-check.sh local gate usesgit apply(no commit) but accumulates state in the same way. - On upstream sync: no action required. If a future FFmpeg point release (n8.1.2 or n8.2) lands new hunks that conflict with one of the patches, regenerate the affected patches via
git format-patchon the resolved state, bump the references in the five files listed under "Touches", and add a fresh rebase-notes entry citing the conflict file(s). - Re-test on rebase:
cd /tmp && rm -rf ffmpeg-n811 && \
git clone --depth 1 --branch n8.1.1 \
https://git.ffmpeg.org/ffmpeg.git ffmpeg-n811
git -C /tmp/ffmpeg-n811 config user.email agent@local
git -C /tmp/ffmpeg-n811 config user.name agent
for p in ffmpeg-patches/000*-*.patch; do
git -C /tmp/ffmpeg-n811 am --3way "$p" || break
done
bash scripts/ci/ffmpeg-patches-check.sh
ADR-0281 follow-up — QSV install-matrix discoverability backfill (2026-05-08)¶
- Touches:
docs/getting-started/install/{arch,fedora,ubuntu,macos,windows}.md(new## Intel QSVsection per page),docs/adr/0281-vmaf-tune-qsv-adapters.md(status-update appendix per ADR-0028),changelog.d/changed/qsv-install-matrix-docs.md(new fragment). No code, no engine, no upstream-shared C / Python source touched. Pure documentation backfill closing the SYCL-audit research-0086 Topic C gap (issue #464). - Invariant: each per-OS QSV section pins the package names against verified upstream URLs with a
Verified 2026-05-08access date. The hardware-generation matrix is sourced from the public Wikipedia "Intel Quick Sync Video — Hardware decoding and encoding" table; if Intel revises which generation supports AV1 encode (e.g. backports the encoder to Lunar Lake / Meteor Lake silicon currently absent from the table), the matrix in all five pages must move in lockstep — the Arch / Fedora / Ubuntu / Windows pages all carry the same matrix verbatim. The macOS page deliberately omits the matrix (QSV unsupported on macOS). - On upstream sync: no action required — Netflix/vmaf upstream does not ship per-OS install pages under
docs/getting-started/install/; that tree is fork-only.
# Lint the install pages (markdownlint via pre-commit):
pre-commit run --files docs/getting-started/install/*.md
# Verify each page (except alpine + macos) still carries the matrix:
for f in arch fedora ubuntu windows; do grep -q 'Arc Battlemage' "docs/getting-started/install/${f}.md" || echo "MISSING: ${f}"
# Confirm the macOS page documents QSV as unsupported:
grep -q 'Intel QSV. is unsupported on macOS' docs/getting-started/install/macos.md
0333 — vmaf-tune Phase F multi-pass encoding (ADR-0333)¶
Touches:
tools/vmaf-tune/src/vmaftune/codec_adapters/__init__.py(CodecAdapter Protocol gainssupports_two_pass: bool+two_pass_args(...))tools/vmaf-tune/src/vmaftune/codec_adapters/x265.py(overrides both)tools/vmaf-tune/src/vmaftune/encode.py(EncodeRequestgainspass_number/stats_path;build_ffmpeg_commandadds the 2-pass argv splice + pass-1 null-muxer redirect; newrun_two_pass_encode)tools/vmaf-tune/src/vmaftune/corpus.py(CorpusOptions.two_pass, routing initer_rows)tools/vmaf-tune/src/vmaftune/cli.py(--two-passflag oncorpus/recommendsubparsers) Invariant: 2-pass encoding routes through the codec adapter viasupports_two_pass+two_pass_args(pass_number, stats_path). The encode driver never branches on codec name. Adapters withsupports_two_pass = Falseare honoured silently (single-pass fallback with stderr warning); the seam is open for sibling codec adapters (libx264, libsvtav1, libvvenc, libaom-av1) to opt in by overriding the two methods on their adapter file alone. This is the fork-local extension to the ADR-0237 Phase A multi-codec contract; upstream Netflix/vmaf has no equivalent and does not own this code path. Re-test:
(Optional, requires ffmpeg + libx265 in the runner's PATH:)
VMAF_TUNE_INTEGRATION=1 python -m pytest \
tests/test_codec_adapter_x265_two_pass.py::test_real_x265_two_pass_smoke -q
Rebase-sensitivity: zero from upstream — tools/vmaf-tune/ is fork-local. The only concern is the codec_adapters Protocol shape: a future upstream commit that adds a sibling codec adapter SHOULD inherit the supports_two_pass = False default and either explicitly opt in or leave the flag off. Downstream sibling-codec PRs in this fork should follow the ADR-0288 / ADR-0333 pattern: one adapter file, override the two methods, add a test file mirroring test_codec_adapter_x265_two_pass.py.
ADR-0360 — CAMBI CUDA port (T3-15a, 2026-05-09)¶
Files pinned:
core/src/feature/cuda/integer_cambi_cuda.c(new)core/src/feature/cuda/integer_cambi_cuda.h(new)core/src/feature/cuda/integer_cambi/cambi_score.cu(new)core/src/feature/feature_extractor.c(addedvmaf_fex_cambi_cudato list)core/src/meson.build(addedcambi_scoretocuda_cu_sources, addedinteger_cambi_cuda.cto CUDA feature sources)
Why: The CUDA twin of vmaf_fex_cambi (Strategy II hybrid — three GPU kernels for the embarrassingly parallel stages; calculate_c_values + topK on CPU). Registers vmaf_fex_cambi_cuda under #if HAVE_CUDA guard.
Rebase-sensitivity: low. The three new files are wholly fork-local and will not conflict. The two upstream-shared files have small, self-contained hunks:
feature_extractor.c: theextern vmaf_fex_cambi_cudadeclaration and the&vmaf_fex_cambi_cudaarray entry are inside a#if HAVE_CUDAblock. Upstream's additions to this file (new feature extractors, new dispatch flags) will not conflict unless Netflix adds their own CUDA twin for CAMBI (unlikely — they don't ship a CUDA backend).meson.build: thecambi_scoreentry in thecuda_cu_sourcesdict and theinteger_cambi_cuda.cline in the CUDA sources list. Any upstream changes tomeson.buildthat restructure thecuda_cu_sourcesdict would require a manual merge; the dict entries are sorted alphabetically by key, socambi_scorelands betweenadm_scoreandmotion_score.
If upstream adds cambi_cuda themselves: drop the fork copy and check for API divergence. Strategy II hybrid is the natural choice; the upstream implementation may differ if they choose Strategy III (fully-on-GPU calculate_c_values).
cambi_internal.h dependency: integer_cambi_cuda.c includes core/src/feature/cambi_internal.h (fork-added trampoline exposing cambi.c's static helpers). If upstream significantly refactors cambi.c (renames vmaf_cambi_preprocessing, vmaf_cambi_calculate_c_values, etc.), cambi_internal.h must be updated alongside. This is the same dependency the Vulkan twin (cambi_vulkan.c) has — see ADR-0210's rebase note for the full list of exposed functions.
Vulkan submit-pool PR-B: six secondary kernels (2026-05-09, ADR-0353)¶
Files changed:
core/src/feature/vulkan/ssim_vulkan.ccore/src/feature/vulkan/ciede_vulkan.ccore/src/feature/vulkan/ms_ssim_vulkan.ccore/src/feature/vulkan/motion_v2_vulkan.ccore/src/feature/vulkan/float_psnr_vulkan.ccore/src/feature/vulkan/float_motion_vulkan.ccore/src/feature/vulkan/AGENTS.mddocs/adr/0353-vulkan-submit-pool-pr-b-six-kernels.md
Why this rebase-note exists: six Vulkan host-glue TUs were migrated from per-frame command-buffer and descriptor-set allocation to the VmafVulkanKernelSubmitPool abstraction (ADR-0256). Any Netflix upstream sync that touches these same files (unlikely — they are fork-local) must preserve the VmafVulkanKernelSubmitPool fields in the state struct and the pool-destroy-before-pipeline-destroy ordering in close_fex().
Rebase-sensitivity: low. All six files are entirely fork-local; Netflix upstream does not have a Vulkan backend. The submit-pool API is defined in core/src/vulkan/kernel.h (also fork-local). No public header or C-API surface was changed; the FFmpeg patch series is unaffected.
Key invariant to preserve on rebase: vmaf_vulkan_kernel_submit_pool_destroy MUST be called before vmaf_vulkan_kernel_pipeline_destroy in every migrated kernel's close_fex(). See core/src/feature/vulkan/AGENTS.md §"Submit-pool ordering invariant".
0354 — Vulkan submit-pool PR-C: submit_pool_destroy-before-pipeline ordering¶
- Touches:
core/src/feature/vulkan/cambi_vulkan.c,core/src/feature/vulkan/ssimulacra2_vulkan.c,core/src/feature/vulkan/float_ansnr_vulkan.c,core/src/feature/vulkan/moment_vulkan.c. - Invariant: In every migrated extractor,
vmaf_vulkan_kernel_submit_pool_destroy()MUST precede everyvmaf_vulkan_kernel_pipeline_destroy()call inclose_fex(). Reversing the order frees the pool's command buffers after the pipeline's command pool is destroyed — undefined behaviour per Vulkan spec §6.2. - Re-test:
meson test -C build --suite=vulkanpasses.scripts/ci/cross_backend_vif_diff.pyshowsplaces=4for all four extractors on all three target devices (RTX 4090, Arc A380, RADV iGPU).
0231 — Vulkan submit-pool migration PR A: adm + motion + psnr (ADR-0291)¶
0231 — Vulkan submit-pool migration PR A: adm + motion + psnr (ADR-0352)¶
- Touches:
core/src/feature/vulkan/adm_vulkan.c,core/src/feature/vulkan/motion_vulkan.c,core/src/feature/vulkan/psnr_vulkan.c(all fork-local Vulkan kernels; no upstream C paths touched),changelog.d/changed/vulkan-submit-pool-pr-a-adm-motion-psnr.md,docs/adr/0291-vulkan-submit-pool-pr-a-adm-motion-psnr.md. - Invariant: Each migrated TU adds
VmafVulkanKernelSubmitPool sub_pooland pre-allocatedVkDescriptorSetfield(s) to its state struct. The pool must be destroyed (vmaf_vulkan_kernel_submit_pool_destroy) beforevmaf_vulkan_kernel_pipeline_destroyinclose_fex(); reversing the order would destroy the descriptor pool while the submit pool still holds live command buffer + fence references. Descriptor sets allocated viavmaf_vulkan_kernel_descriptor_sets_allocare freed implicitly by the descriptor pool tear-down — do NOT callvkFreeDescriptorSetson them inclose_fex(). Formotion_vulkan, the pre-allocated set is rebound once per frame viavkUpdateDescriptorSetsbecause the blur ping-pong changes whichblur[]slot is "current"; foradm_vulkanandpsnr_vulkanthe sets are stable afterinit()and require no per-frame update. - Upstream interaction: none. All three files are fork-local Vulkan kernel TUs not present in Netflix/vmaf upstream.
- On upstream sync: zero interaction. Upstream cannot conflict with this PR's paths. The Vulkan backend is entirely fork-introduced.
- Re-test on rebase:
meson test -C build --suite=fast
# Cross-backend parity gate (places=4):
python python/test/cross_backend_diff.py \
--features adm motion psnr \
--backend vulkan cpu \
--places 4 \
--yuv testdata/yuv/src01_hrc00_576x324.yuv \
testdata/yuv/src01_hrc01_576x324.yuv
ADR-0350 — FFmpeg libvmaf filter CUDA backend selector (0010 patch)¶
Patch: ffmpeg-patches/0010-libvmaf-wire-cuda-backend-selector.patch.
libavfilter/vf_libvmaf.c— addscudaAVOption + state field + init / cleanup / picture-pool wiring underCONFIG_LIBVMAF_CUDA && !CONFIG_LIBVMAF_CUDA_FILTER.configure— adds--enable-libvmaf-cuda(EXTERNAL_LIBRARY_LISTentry + help text), promoteslibvmaf_cudafrom blanket-autodetect to gatedenabled libvmaf_cuda && require_pkg_config + check, preserves theenabled libvmaf && check_pkg_config libvmaf_cudain-filter probe so the new selector still works without the explicit flag when libvmaf ships CUDA. Why this rebase-note exists: Patch0010extends the SYCL (0003) / Vulkan (0004) per-context backend selectors to CUDA on the regularlibvmaffilter. The patch coexists with the upstream dedicatedlibvmaf_cudafilter (CONFIG_LIBVMAF_CUDA_FILTER) by gating its struct field and code paths on!CONFIG_LIBVMAF_CUDA_FILTER— the dedicated filter keeps owning its owncu_statefield. CLAUDE.md §12 r14 makes the patch update mandatory because the change touches a filter consumer of thevmaf_cuda_state_init/_import_state/_state_free/_preallocate_pictures/_fetch_preallocated_pictureC-API surface inlibvmaf_cuda.h. Rebase-sensitivity: low. The patch'svf_libvmaf.chunks are context-anchored on the SYCL/Vulkan selector blocks; if upstream FFmpeg renamesCONFIG_LIBVMAF_CUDA_FILTERor moves thelibvmaf_cuda.hinclude, the include guard at the top of the file needs the corresponding update. The configure hunks are context-anchored on the existing--enable-libvmaf-sycl/--enable-libvmaf-vulkanlines — those have proven stable across n8.0 → n8.1 → n8.1.1, so drift risk is low. WhenVmafCudaConfigurationever grows adevice_indexfield upstream, swap thecudaboolean for anint cuda_devicemirroring SYCL's shape (separate ADR + patch refresh). Verification gate: cumulativegit am --3wayreplay offfmpeg-patches/000{1..9}-*.patch+0010-*against pristine FFmpegn8.1.1PASS (2026-05-09). Build oflibavfilter/vf_libvmaf.oPASS under bothCONFIG_LIBVMAF_CUDA=0(selector errors at filter- init time per#elsebranch) andCONFIG_LIBVMAF_CUDA=1 && !CONFIG_LIBVMAF_CUDA_FILTER(selector active, picture-pool wiring compiles).
0320 — Vulkan instance / VMA apiVersion bump to 1.4 (Step B)¶
- Touches:
core/src/vulkan/common.c,core/src/vulkan/vma_impl.cpp,core/src/vulkan/AGENTS.md. - Invariant: the four
apiVersionsites (lines 54, 264, 374 ofcommon.c; line 22 ofvma_impl.cpp) request Vulkan 1.4, not 1.3. Together with the Step-Aprecisedecorations invif.comp/ciede.comp(PR #346) and the Phase-3 cross-subgroup release-acquire fix (PR #511), this gates the cross-backend places=4 contract on Arc + RADV. NVIDIA closure depends on Phase 3c (PR #512; block-on-merge until that lands). Netflix upstream does not carry a VMA dependency or a Vulkan backend; no upstream merge conflict expected on these files. - Re-test on rebase:
meson setup build -Denable_vulkan=enabled -Denable_cuda=false \
-Denable_sycl=false --buildtype=release
ninja -C build
for D in 0 1 2; do
python3 scripts/ci/cross_backend_parity_gate.py \
--vmaf-binary build/tools/vmaf \
--reference python/test/resource/yuv/src01_hrc00_576x324.yuv \
--distorted python/test/resource/yuv/src01_hrc01_576x324.yuv \
--width 576 --height 324 --pixel-format 420 --bitdepth 8 \
--backends cpu vulkan --vulkan-device "$D" \
--features vif ciede adm motion psnr
done
# All 0/N mismatches at places=4 once Phase 3c (PR #512) has landed.
ADR-0332 v2 runtime (T5-2c) — Embedded MCP server UDS + real compute_vmaf (2026-05-09)¶
- Touches:
core/src/mcp/{mcp.c,dispatcher.c,mcp_internal.h,meson.build,compute_vmaf.c,transport_uds.c},core/test/test_mcp_smoke.c. All paths are fork-local. No new third-party vendor drop in v2 — mongoose vendoring stays deferred to v3 with the SSE transport. - Invariant: same as ADR-0209 v1 — the entire
core/src/mcp/subtree is fork-local; the public ABI incore/include/libvmaf/libvmaf_mcp.his unchanged (only function bodies flipped —vmaf_mcp_start_udsfrom-ENOSYSto a working AF_UNIX listener;compute_vmaffrom a{"status":"deferred_to_v2"}placeholder to a realvmaf_score_pooledbinding). Per ADR-0128 § operational guardrails the UDS socket file is created mode 0700; thatchmodhappens invmaf_mcp_start_udsafterbindand is a load-bearing security invariant — do NOT relax it on rebase.compute_vmafruns on a per-call ephemeralVmafContextso the host's main scoring run is unperturbed; do NOT rewire it to reuseserver->ctxbecausevmaf_score_pooledcommits the model destructively to the context. - On upstream sync: no action required. Netflix/vmaf upstream has no embedded MCP surface. If upstream adds one, expect a port-only sync since names will collide.
- Re-test on rebase:
cd libvmaf && meson setup build -Denable_cuda=false -Denable_sycl=false \
-Denable_mcp=true -Denable_mcp_stdio=true \
-Denable_mcp_uds=true
ninja -C build && meson test -C build test_mcp_smoke -v
# Real-score smoke (single 576x324 pair):
build/test/test_mcp_smoke 2>&1 | tail -3 # expects "16 tests run, 16 passed"
ADR-0332 v3 runtime (T5-2d) — Embedded MCP server SSE transport (2026-05-09)¶
- Touches:
core/src/mcp/{mcp.c,mcp_internal.h,meson.build,transport_sse.c},core/meson_options.txt,core/test/test_mcp_smoke.c,docs/mcp/embedded.md,docs/adr/0332-mcp-runtime-v2.md(status-update appendix). All paths are fork-local. No third-party vendor drop in v3 — the originally-planned mongoose vendor was reversed because cesanta/mongoose 7.18 is GPL-2.0-only OR commercial, incompatible with the fork's BSD-3-Clause-Plus-Patent license (verified at upstream LICENSE 2026-05-09). The SSE transport is plain POSIX sockets in fork-owned C (~500 LOC). - Invariant: same as ADR-0209 / ADR-0332 v2 — the entire
core/src/mcp/subtree is fork-local; the public ABI incore/include/libvmaf/libvmaf_mcp.his unchanged (onlyvmaf_mcp_start_sse's body flipped from-ENOSYSto a working AF_INET listener). The SSE listener bindsINADDR_LOOPBACKonly; do NOT switch toINADDR_ANYwithout a separate ADR + auth design (v3 ships intentionally without CORS/Bearer/per-session auth on the assumption of a same-host trust boundary). The SSE stop path usesshutdown(SHUT_RDWR)beforeclose()— plainclose()of an AF_INET listening fd from another thread does NOT unblockaccept()on Linux; do NOT remove theshutdowncall.enable_mcp_sseis now afeatureoption (defaultauto), notboolean false. - On upstream sync: no action required. Netflix/vmaf upstream has no embedded MCP surface. Do NOT re-introduce mongoose (or any GPL-licensed HTTP library) on a future rebase without first amending CLAUDE §1 and adding a separate license-compatibility ADR.
- Re-test on rebase:
cd libvmaf && meson setup build -Denable_cuda=false -Denable_sycl=false \
-Denable_mcp=true -Denable_mcp_stdio=true \
-Denable_mcp_uds=true \
-Denable_mcp_sse=enabled
ninja -C build && meson test -C build test_mcp_smoke -v
build/test/test_mcp_smoke 2>&1 | tail -3 # expects "17 tests run, 17 passed"
Status update 2026-05-09 — placeholder-ref hardening¶
- Additional touches: same set as the 2026-05-08 ADR-0334 entry, no new files. The hardening adds a
git diff -U0 ... -- docs/state.mdcall insidescripts/ci/state-md-touch-check.sh(case 4a) plus 10 additional fixture cases inscripts/ci/test-state-md-touch-check.sh. - New invariant: inserted lines in
docs/state.md(lines starting with+, excluding the+++ b/...header) must not containthis PR/this commit/ bareTBD/<PR>/#NNN. Canonical accept forms arePR #Nandcommit `<sha>`. The placeholder vocabulary is coupled to PR #541's audit findings — reword in lockstep with the ADR-0334 status-update appendix if the fork's row template changes. - Re-test on rebase: same
bash scripts/ci/test-state-md-touch-check.shrun as the 2026-05-08 entry; the harness now reports18/18 passed(was8/8 passed).
0347 — Sanitizer matrix test-set scope (ADR-0347)¶
- Touches:
.github/workflows/tests-and-quality-gates.ymljobsanitizers(build + test step),core/test/meson.build(no edits — the absence of anysuite: 'unit'tag is the upstream state we now work with rather than against). - Invariant: the sanitizer job runs the full C unit-test set per sanitizer with a per-sanitizer deselect list driven by a
caseblock on${{ matrix.sanitizer }}. The deselect lists are load-bearing — each entry corresponds to a real bug tracked indocs/state.md. Under UBSan the build adds-Dc_args=-fno-sanitize=function -Dcpp_args=-fno-sanitize=functionto suppress the K&R-prototype harness UB; the mesoncasebranch must keep this build flag in sync with the test deselect entries. An upstream rebase that adds new test files viacore/test/meson.buildinherits full sanitizer coverage automatically (the workflow enumerates tests viameson test --list). - On upstream sync: if upstream Netflix lands a
suite: 'unit'tagging convention, the workflow is robust to it (we already enumerate frommeson test --list, not from--suite=unit). If upstream rewrites the harness to declarestatic char *test_X(void)with a(void)parameter, the-fno-sanitize=functionflag becomes redundant — leave it in place (zero cost) until a deliberate cleanup PR reverts the suppression. If upstream lands a fix for any of the surfaced defects (SVMModelParservalidation,feature_collectormetadata leak,integer_adm::div_lookuprace,framesyncmutex mismatch), drop the corresponding deselect row from the workflow'scaseblock in the same PR that pulls the upstream fix. cd libvmaf for SAN in address undefined thread; do EXTRA=() [ "$SAN" = undefined ] && EXTRA=( "-Dc_args=-fno-sanitize=function" "-Dcpp_args=-fno-sanitize=function" ) rm -rf "build-$SAN" CC=clang CXX=clang++ LDFLAGS=-fuse-ld=lld \ meson setup "build-$SAN" -Db_sanitize="$SAN" \ -Denable_cuda=false -Denable_sycl=false --buildtype=debug \ -Db_lto=false -Db_lundef=false "${EXTRA[@]}" meson compile -C "build-$SAN" case "$SAN" in address) EXCLUDE='test_model$|test_predict$|test_float_ms_ssim_min_dim$' ;; undefined) EXCLUDE='test_model$' ;; thread) EXCLUDE='test_model$|test_pic_preallocation$|test_framesync$' ;; esac TESTS=$(meson test -C "build-$SAN" --list \ | grep '^libvmaf:' \ | grep -vE "$EXCLUDE" \ | sed 's/^libvmaf://') meson test -C "build-$SAN" --print-errorlogs $TESTS
CodeQL bulk mechanical sweep — Python tree (2026-05-09)¶
- Why this matters on rebase: no rebase impact. The diff lives entirely in
python/vmaf/and one fork-local helper (core/src/vulkan/spv_embed.py). None of the touched Python modules have been changed by Netflix upstream in over four years; the closest churn is unrelated additions topython/vmaf/script/run_*.pydriver flags. A future/sync-upstreamwill land on a clean tree. - What changed: dead imports removed;
exit()→sys.exit()in seven CLI driver scripts;open(...)→with open(...)inpython/vmaf/tools/decorator.pyandcore/src/vulkan/spv_embed.py; typedexcept KeyError: passbodies got an explanatory one-line comment to satisfypy/empty-except;passremoved where it was a no-op tail statement; one commented-out debug block deleted fromtools/misc.py. - Re-test on rebase:
python3 -c "import ast; [ast.parse(open(f).read()) for f in (...)]"over the touched files;ruff checkover the same set must produce no NEW errors versus master baseline.
0345 — cambi × {CUDA, SYCL, HIP} GPU port planning (ADR-0345, docs-only)¶
- Touches:
docs/research/0091-cambi-gpu-port-planning-2026-05-09.md(new),docs/adr/0345-cambi-gpu-port-strategy.md(new),docs/adr/_index_fragments/0345-cambi-gpu-port-strategy.md(new fragment),docs/adr/_index_fragments/_order.txt(append slot),changelog.d/changed/cambi-gpu-planning-digest.md(new). No code. Companion to the per-port PRs that follow per the digest's §6 ordered plan (CUDA → SYCL → HIP). - Upstream source: none — fork-local planning artefact. Netflix/vmaf upstream has no CUDA / SYCL / HIP cambi twin and no plans to add one on those backends.
- Invariant: the planning round locks Strategy II host-staged hybrid for the three pending backends, inheriting verbatim from ADR-0205 §Decision and ADR-0210 §Decision. The cross-backend gate contract for cambi is
places=4from day one on all backends — by construction (integer-only GPU pre-passes; byte-identical readback; unmodified host residual). If any per-port PR sees empirical drift from CPU, fix the kernel — never relax the gate (memoryfeedback_no_test_weakening). The sharedcambi_internal.hhost residual surface (shipped with PR #196 for the Vulkan port) is the load-bearing reuse point — all four GPU twins (Vulkan, CUDA, SYCL, HIP) link against it and inherit any future CPU-side c-value formula change automatically. - On upstream sync: no action required. If a future upstream sync introduces a Netflix/vmaf cambi GPU twin (extremely unlikely — Netflix has no public CUDA / SYCL / HIP cambi work), evaluate whether to drop the fork's twin in favour of upstream's per the standard prefer-upstream rule; otherwise no action.
- Re-test on rebase: docs-only — no compile / runtime gate. The Strategy III v2 follow-up (parked per ADR-0205 §Out of scope) gets its own ADR + rebase-notes entry when profile data lands.
0320 — Vulkan VIF API-1.4 NVIDIA residual Phase 3b (deferral)¶
- Touches:
core/src/feature/vulkan/shaders/vif.comp(comment-only update at the Phase-4 reduction site — documents the Phase-3b candidate-fix experiments and the driver-side hypothesis; no code logic change vs. PR #511);docs/adr/0269-vif-ciede-precise-step-a.md(appended Phase-3b status update appendix; ADR body remains frozen per ADR-0028);docs/research/0090-...md(new);docs/state.md(rowT-VK-VIF-1.4-RESIDUAL-ARCretired in favour ofT-VK-VIF-1.4-RESIDUAL-NVIDIA-DEFERREDafter the hardware-mapping correction);core/src/vulkan/AGENTS.md(Phase 3b update + rebase invariant for cross-backend gate device-name selection);changelog.d/fixed/vif-arc-mesa-anv-int64-reduction.md(new fragment). - Invariant: the workgroup-scope
memoryBarrierShared(); barrier();pair PR #511 introduced is load-bearing for the Arc + RADV lanes at API 1.4 and stays. Phase 3b confirmed it cannot be downgraded back to a barebarrier()even if the NVIDIA residual ever closes — Arc's clean state is contingent on the workgroup-scope pair. - Cross-backend gate device-selection invariant (NEW): scripts that target a specific Vulkan vendor must select by
deviceNamesubstring, not by--vulkan_device <index>.vmaf_vulkan_context_new's device sort is stable inside the samedevtype_scorebucket and thevkEnumeratePhysicalDevicesenumeration order is host-policy-dependent (driver registration order in/etc/vulkan/icd.d/, Mesa device-select layer,VK_LOADER_*env vars). PR #511's commit message inverted the device map on this fork's CI workstation; the empirical numbers it cited as "NVIDIA" actually came from Arc and vice versa. New cross-backend lanes targeting a specific vendor should not inherit the off-by-one. - On upstream sync:
vif.compis fork-local; no upstream Netflix/vmaf has a Vulkan path. Cherry-picks from upstream cannot reach this file. - Re-test on rebase (assumes a multi-GPU CI workstation with NVIDIA + Arc + RADV; lavapipe-only CI lanes are a no-op for the API-1.4 residual since lavapipe never reproduced the bug):
# Local API-1.4 bump (off-master reproducer; do NOT commit).
sed -i 's/VK_API_VERSION_1_3/VK_API_VERSION_1_4/g' \ core/src/vulkan/common.c sed -i 's/VMA_VULKAN_VERSION 1003000/VMA_VULKAN_VERSION 1004000/' \ core/src/vulkan/vma_impl.cpp cd libvmaf && meson setup build -Denable_vulkan=enabled \ -Denable_cuda=false -Denable_sycl=false && ninja -C build cd ..
# NVIDIA lane — expected 45/48 FAIL scale 2 until either the
# manual int64 subgroup-reduction patch lands or NVIDIA fixes
# the driver. Arc + RADV expected 0/48.
python3 scripts/ci/cross_backend_vif_diff.py \ --vmaf-binary core/build/tools/vmaf \ --reference testdata/ref_576x324_48f.yuv \ --distorted testdata/dis_576x324_48f.yuv \ --width 576 --height 324 \ --feature vif --backend vulkan --device
# Revert local bump after testing.
sed -i 's/VK_API_VERSION_1_4/VK_API_VERSION_1_3/g' \ core/src/vulkan/common.c sed -i 's/VMA_VULKAN_VERSION 1004000/VMA_VULKAN_VERSION 1003000/' \ core/src/vulkan/vma_impl.cpp
Upstream-port-later batch — Research-0090 18-commit triage close-out (2026-05-09)¶
- Touches:
docs/state.md(one row in "Deferred (waiting on external trigger)"), this file,changelog.d/changed/upstream-port-later-batch-2026-05-09.md. No code touched. Companion to PR #446 (Research-0090) and the in-flight PRs #497 (MyTestCase super-PR), #443 / #444 (cambi-docs duplicate pair). - Per-commit classification (input set: 18 PORT_LATER SHAs from Research-0090):
| # | Upstream SHA | Subject (truncated) | Verdict | Reopen / forward path |
|---|---|---|---|---|
| 1 | 38e905d1 | adopt MyTestCase + reformat BD-rate test data | PORT_DEFERRED | Subsumed by PR #497 commit e1dbdc09; close out when #497 merges |
| 2 | 005988ea | adopt MyTestCase + port new tests + align fifo_mode | PORT_DEFERRED | Subsumed by PR #497 commit 6c05afe2; close out when #497 merges |
| 3 | 4679db83 | fix VMAFEXEC_score tolerances for macOS FP precision | PORT_DEFERRED w/ Netflix-golden guard | PR #497 commit 0004d2cf — must preserve fork's golden places= values byte-for-byte (CLAUDE §8 / ADR-0024) |
| 4 | 3e075107 | adopt MyTestCase + update score values in vmafexec tests | PORT_DEFERRED | Subsumed by PR #497 commit 0004d2cf; close out when #497 merges |
| 5 | e3827e4d | adopt MyTestCase + port new tests in asset/bootstrap/local_explainer | PORT_DEFERRED | Subsumed by PR #497 commit 6c05afe2; close out when #497 merges |
| 6 | 25ff9f18 | remove empty VmafossexecCommandLineTest stub | PORT_DEFERRED → CHERRY-PICK after #497 | Pure 13-line deletion. PR #497 currently RE-EMITS the stub; once #497 lands, cherry-pick this commit standalone (zero-conflict against post-#497 tip). |
| 7 | 3a041a97 | adopt MyTestCase + update score values | PORT_DEFERRED | Subsumed by PR #497 commit d52d9221; close out when #497 merges |
| 8 | ead2d12b | fix vif_scale3 + adm3_egl_1 tolerances for macOS FP precision | PORT_DEFERRED w/ Netflix-golden guard | PR #497 commit b5a3f61b — Netflix-golden tolerance guard same as row 3 |
| 9 | 6c097fc4 | reduce ADM/VIF tolerances for macOS FP precision | PORT_DEFERRED w/ Netflix-golden guard | PR #497 commit f3881d5c — Netflix-golden tolerance guard same as row 3 |
| 10 | 7df50f3a | align testutil with full set of fixture functions | PORT_DEFERRED | Subsumed by PR #497 commit f1ae0495; close out when #497 merges |
| 11 | 322ca041 | replace temporal slicing with pre-sliced YUV fixtures | PORT_DEFERRED | Subsumed by PR #497 commit 7d9d9a10; close out when #497 merges. Sequencing matters: this commit must land before rows 12, 14, 15, 17 (the YUV-fixture consumers); #497 already orders them correctly. |
| 12 | 74bdce1b | align vmafexec_feature_extractor_test (aim/adm3/motion3) | PORT_DEFERRED | Subsumed by PR #497 commit 07e7cb48; close out when #497 merges |
| 13 | a3776335 | align feature_extractor_test (aim/adm3/motion3) | PORT_DEFERRED | Subsumed by PR #497 commit 15a6874d; close out when #497 merges |
| 14 | 0341f730 | remove duplicate test_run_vmaf_integer_fextractor | PORT_DEFERRED → CHERRY-PICK after #497 | Pure 76-line deletion. Same disposition as row 6 — #497 currently re-emits the duplicate; cherry-pick standalone after #497. |
| 15 | 9fa593eb | port feature_extractor tests for aim/adm3/motion3 + new options | PORT_DEFERRED | Subsumed by PR #497 commit ab21b694; close out when #497 merges |
| 16 | d93495f5 | reduce tolerance for VMAF scores in quality_runner tests | PORT_DEFERRED w/ Netflix-golden guard | PR #497 — Netflix-golden tolerance guard same as row 3 |
| 17 | 7d1ad54b | port feature extractor tests for aim/adm3/motion3 | PORT_DEFERRED | Subsumed by PR #497 commit 44b9e626; close out when #497 merges |
| 18 | 721569bc | resource/doc: cambi_high_res_speedup + motion2 score | PORT_DEFERRED → DEDUP | Already in flight on TWO branches (PR #443 + PR #444). Maintainer picks one and abandons the other per Research-0090 §Recommended action #4. No third port-PR opened. |
- Invariant: after PR #497 merges, the Research-0090 PORT_LATER bucket reduces to exactly two follow-up cherry-picks against post-#497 master:
git cherry-pick 25ff9f18(delete emptyVmafossexecCommandLineTest).git cherry-pick 0341f730(delete duplicatetest_run_vmaf_integer_fextractor). Both are pure deletions onpython/test/command_line_test.pyandpython/test/feature_extractor_test.pyrespectively; no score change, no Netflix-golden interaction. They were excluded from PR #497 because the v2 super-PR's diff state currently RE-EMITS those identifiers (likely because #497 cherry-picked from an earlier upstream tip than25ff9f18/0341f730).- Netflix-golden guard (binding): per CLAUDE §8 / ADR-0024, the three Netflix CPU golden pairs in
python/test/quality_runner_test.py,vmafexec_test.py,vmafexec_feature_extractor_test.py,feature_extractor_test.py,result_test.py(1 normalsrc01_hrc00↔hrc01+ 2 checkerboard) carry hard-codedassertAlmostEqualrows that are NEVER modified by a fork PR. Upstream commits4679db83,ead2d12b,6c097fc4,d93495f5explicitly LOWERplaces=on a subset of those rows (their stated motivation is macOS FP precision drift, not a true score change). Reviewer of PR #497 must verify that the 3 golden pairs retain fork tolerances byte-for-byte; only non-golden rows may adopt the relaxations. - On upstream sync: future
/sync-upstreamruns that re-detect these 18 SHAs should match this entry via the SHA list and short-circuit Pass-2 classification (skip re-triage). - Re-test on rebase: none required at the time of this commit (no code touched); after the two follow-up cherry-picks (
25ff9f18+0341f730) eventually land, run meson test -C build --suite=fast make test-netflix-golden # 3/3 CPU goldens still pass
ADR-0357 — Vulkan readback buffer VMA flag separation (PR pending)¶
What changed: picture_vulkan.{c,h} now exposes two sibling allocation functions: vmaf_vulkan_buffer_alloc (UPLOAD, unchanged) and vmaf_vulkan_buffer_alloc_readback (READBACK, HOST_ACCESS_RANDOM). A new vmaf_vulkan_buffer_invalidate wraps vmaInvalidateAllocation. All 17 feature kernel files under core/src/feature/vulkan/ are updated to use the readback variant for accumulator and partial-sum buffers.
core/src/vulkan/picture_vulkan.c— two new functions + shared helper.core/src/vulkan/picture_vulkan.h— two new declarations.- All 17
core/src/feature/vulkan/*.cfiles — alloc and invalidate call sites. Rebase-sensitivity: low — entirely fork-local Vulkan backend code with no upstream Netflix counterpart. If an upstream sync adds new files tocore/src/vulkan/orcore/src/feature/vulkan/, new readback buffers in those files must be classified (UPLOAD vs READBACK) and use the correct allocator per the table in ADR-0350. Conflict risk on the 17 feature files is zero (upstream doesn't touch them).
ADR-0356 — ffmpeg-patches surface-sync CI gate (2026-05-09)¶
Files added:
scripts/ci/ffmpeg-patches-surface-check.sh(new gate script)..github/workflows/rule-enforcement.yml(newffmpeg-patches-surface-checkjob).docs/adr/0356-ffmpeg-patches-surface-gate.md(decision record).docs/development/automated-rule-enforcement.md(user-facing doc update).
Why this rebase-note exists: the gate is fork-local CI; it does not touch any upstream-shared file, so an upstream merge cannot drop its enforcement. However, whoever runs the next /sync-upstream should be aware that ffmpeg-patches/ integrity is now machine-checked on every PR — if a future libvmaf header rename slips through during conflict resolution and breaks the patch stack, the gate will fire on the post-sync PR and surface the omission immediately rather than at the next sync.
Rebase-sensitivity: zero on the upstream-merge path. Indirect benefit: the gate hardens ffmpeg-patches/ against silent drift, so the patch-stack invariants tracked elsewhere in this file (entries referencing ffmpeg-patches/0001…0009) are now machine-defended.
0320 — HIP CI lane apt-installs ROCm runtime (ADR-0212 status update)¶
- Touches:
.github/workflows/libvmaf-build-matrix.yml(HIP laneif: matrix.hipinstall step + base-deps gate),.github/workflows/required-aggregator.yml(HIP lane added to required-check allow-list). Upstream Netflix/vmaf has no HIP backend and no equivalent CI matrix; conflict probability againstupstream/masteris zero. Entry exists to flag the rebase-sensitive ROCm-version pin for future maintainers. - Invariant: the ROCm version pin (
ROCM_VERSION: "7.2.3") in theInstall ROCm / HIP runtimestep must match the version the maintainer's local box runs against. The apt URL ishttps://repo.radeon.com/rocm/apt/<ver>— the version is part of the path, so AMD effectively snapshots each ROCm release as its own apt repo. Bumping the pin is a one-line change but requires re-validating thatrocm-hip-runtime-devstill pulls the same symbol set; in particular,amdhip64major-version changes have historically brokendlopenconsumers.nobleis the codename forubuntu-24.04, which is whatubuntu-latestresolves to on GitHub-hosted runners as of 2024-04. Ifubuntu-latestrolls forward to a newer LTS, the apt repo path component (https://repo.radeon.com/rocm/apt/<ver> <codename> main) needs to be re-checked against https://rocm.docs.amd.com/projects/install-on-linux/en/latest/install/install-methods/package-manager/package-manager-ubuntu.html for the current AMD-supported codename list. - Re-test on rebase:
# Locally, mirror what CI does (assumes ROCm /opt/rocm install on dev box):
meson setup build -Denable_hip=true -Denable_cuda=false -Denable_sycl=false
ninja -C build
./build/test/test_hip_smoke # passes with device_count == 0
# Apt-side: verify the URL still resolves (versioned path)
curl -sfI https://repo.radeon.com/rocm/apt/7.2.3/dists/noble/Release \
&& echo OK || echo "ROCm apt URL drifted — bump ROCM_VERSION"
RN-2026-05-08-cambi-cluster — port 9 of 10 upstream cambi commits¶
- Tracked by: ADR-0328, PR
feat/upstream-port-cambi-cluster-2026-05-08. - Cluster: Netflix upstream commits
d655cefe,9fad7317,767a6780,8c60dc9e,bd278ea6,1091b0c1,77474251,933cccb4,984f281fported verbatim.41bacc83("move shared code to cambi.h") explicitly skipped. - Touches:
core/src/feature/cambi.c,core/src/feature/x86/cambi_avx2.c,core/src/feature/x86/cambi_avx2.h,core/test/test_cambi.c.cambi_reciprocal_lut.hstays (fork commitef6d33e6already added it before upstream). - Invariant: the fork uses a
CAMBI_CALC_C_VALUES_BODYmacro incambi.cto share the calculate_c_values loop nest acrosscalculate_c_values(scalar),calculate_c_values_avx2, andcalculate_c_values_neon. Upstream keeps the three variants as separate function definitions incambi.c(scalar) andcambi_avx2.c(AVX-2) with the helpers exposed viacambi.h. The fork's macro keeps the three drivers in lockstep without externalising the helpers. - Twin-update gaps:
- AVX-512: no
calculate_c_values_row_avx512exists; the AVX-512 dispatch path falls through tocalculate_c_values_avx2. Tracked as a perf follow-up — bit-exactness preserved, only throughput affected. - NEON:
calculate_c_values_neonuses scalarcalculate_c_values_row(no NEONcalculate_c_values_row_neonexists yet). Tracked as a perf follow-up. - CUDA / SYCL: cambi has no GPU twin in those backends (the only existing twin is Vulkan, ADR-0205 Strategy II). The Vulkan twin's host-residual shim
vmaf_cambi_calculate_c_valueswas updated in port 933cccb4 to drop the inc/dec range-updater parameters (now(void)-cast sincecalculate_c_valuesself-dispatches its updaters); ABI-compatible withcambi_internal.hcallers. - On upstream sync: when re-syncing cambi, expect conflicts on the
calculate_c_values_avx2body — upstream keeps it as a function incambi_avx2.c, the fork keeps it insidecambi.cvia the macro. The translation is mechanical: take any inner-loop change from upstream's body, apply it once insideCAMBI_CALC_C_VALUES_BODY. The fork'scalculate_c_values_neonhas no upstream counterpart and stays fork-local. - Re-test on rebase:
cd libvmaf && meson setup build -Denable_cuda=false -Denable_sycl=false
ninja -C build && build/test/test_cambi
# Optional GPU-parity gate when available:
# ./scripts/cross-backend-diff.sh --feature cambi
ADR-0336 — KonViD MOS head v1 (2026-05-08)¶
- Touches:
ai/scripts/train_konvid_mos_head.py(new),ai/tests/test_train_konvid_mos_head.py(new),tools/vmaf-tune/src/vmaftune/predictor.py(addsPredictor.predict_mos+ the optionalkonvid_mos_head_v1.onnxloader;_DEFAULT_COEFFSand_predict_analyticalare unchanged),tools/vmaf-tune/tests/test_predict_mos.py(new),model/konvid_mos_head_v1.onnx(new),model/konvid_mos_head_v1_card.md(new),model/konvid_mos_head_v1.json(new manifest sidecar),docs/adr/0336-konvid-mos-head-v1.md(new),docs/research/0090-konvid-mos-head-design.md(new),docs/state.md(T-MOS-HEAD-PRODFLIP row),changelog.d/added/0336-konvid-mos-head-v1.md(new). All paths are fork-local; upstream Netflix/vmaf has no MOS-head surface and the predictor lives entirely undertools/vmaf-tune/. - Invariant: the MOS-head ONNX I/O contract is two-input named tensors (
featuresshape(N, 11);encoder_onehotshape(N, 1)) -> one output tensor (mosshape(N,)) with the range[1.0, 5.0]baked into the graph via1 + 4 * sigmoid(raw). The 11 feature columns are(adm2, vif_scale0..3, motion2, saliency_mean, saliency_var, shot_count_norm, shot_mean_len_norm, shot_cut_density)in that exact order — they line up withtrain_konvid_mos_head.FEATURE_COLUMNSand the predictor's_predict_mos_via_headzero-fills layout. ENCODER_VOCAB v4 ships a single"ugc-mixed"slot; multi-slot expansion is append-only.Predictor.predict_mosfalls back tomos = (predicted_vmaf - 30) / 14clamped to[1, 5]whenever the ONNX is missing oronnxruntimeis unavailable — that fallback is the documented behaviour, not a bug. - On upstream sync: no action required. The trainer + predictor + MOS head + tests live entirely under fork-local paths (
ai/,tools/vmaf-tune/,model/); upstream syncs cannot touch them.tools/vmaf-tune/src/vmaftune/predictor.pyis fork-local but co-evolves with vmaf-tune; if a future ADR re-shapesShotFeatures, replay the MOS-head feature-column map in lockstep. - Re-test on rebase:
```bash python3 -m pytest ai/tests/test_train_konvid_mos_head.py tools/vmaf-tune/tests/test_predict_mos.py -v python3 ai/scripts/train_konvid_mos_head.py --smoke --no-export # gate must report PASS
ADR-0335 — AdaptiveCpp as a second SYCL toolchain (2026-05-08)¶
- Touches:
core/src/feature/sycl/sycl_compat.h(new),core/src/feature/sycl/*.cpp(10 attribute call sites in 9 files switched from[[intel::reqd_sub_group_size(N)]]toVMAF_SYCL_REQD_SG_SIZE(N)),core/src/meson.build(toolchain branch in the SYCL block + the feature-kernel block),core/meson_options.txt(description bump onsycl_compiler+ newsycl_acpp_targetsoption),docs/development/sycl-toolchains.md(new),docs/adr/0335-adaptivecpp-second-sycl-toolchain.md(new),docs/adr/_index_fragments/0335-adaptivecpp-second-sycl-toolchain.md(new),docs/adr/_index_fragments/_order.txt(append),docs/adr/README.md(regenerated byconcat-adr-index.sh --write),docs/adr/0217-sycl-toolchain-cleanup.md(status-update appendix per ADR-0028),core/src/sycl/AGENTS.md(invariant row),changelog.d/added/0335-adaptivecpp-second-sycl-toolchain.md(new). No upstream-shared paths incore/src/feature/sycl/*.cppare touched onupstream/master(those TUs are fork-local SYCL twins). - Invariant: Intel
icpxstays the primary toolchain. AdaptiveCpp is opt-in via-Dsycl_compiler=acpp. Any new Intel-specific SYCL kernel attribute (e.g. a future[[intel::*]]decoration,sycl::ext::oneapi::experimental::*use) must land behind a new macro incore/src/feature/sycl/sycl_compat.hrather than appear inline. AdaptiveCpp output is not bit-identical to icpx and not bit-identical to scalar CPU (consistent with the existing CPU-only golden gate). The canonical AdaptiveCpp identification macros areSYCL_IMPLEMENTATION_ACPPand the legacySYCL_IMPLEMENTATION_HIPSYCL, both auto-defined by<sycl/sycl.hpp>. - On upstream sync: if a Netflix upstream cherry-pick lands a bare
[[intel::reqd_sub_group_size(N)]](or any Intel-specific SYCL attribute) on a kernel lambda, wrap the attribute in the appropriateVMAF_SYCL_*compat macro before merging. Upstream has no SYCL backend today, so the conflict surface is small. - Re-test on rebase:
# Plumbing parses cleanly with the icpx default still selected:
meson setup /tmp/build-sycl-icpx libvmaf -Denable_sycl=false
# And the macro count is consistent (10 sites under acpp guard):
grep -rl 'VMAF_SYCL_REQD_SG_SIZE' core/src/feature/sycl | wc -l
# → 9 files (the compat header itself defines the macro;
# 9 kernel TUs consume it.)
ADR-0212 §Status update — HIP runtime (T7-10b, 2026-05-08)¶
- Touches:
core/src/hip/common.c,core/src/hip/kernel_template.c,core/src/hip/meson.build,core/test/test_hip_smoke.c,core/test/meson.build(addedhip_depseverywherevulkan_depsalready appears so test executables that statically pull the feature lib resolvehipMemsetAsync/hipFree). - Invariant: the
kernel_template.chelpers andcommon.cpublic API both store HIP runtime handles (hipStream_t,hipEvent_t) asuintptr_tin the structs that cross the public ABI. The header-purity contract documented incore/src/hip/kernel_template.his load-bearing — moving the cast site (or replacinguintptr_twithvoid *) breaks every consumer TU and the publiclibvmaf_hip.hno-<hip/...>guarantee. The fallbackfind_library('amdhip64', dirs: hip_search_paths)exists because ROCm 7.x publishes nohip-lang.pcand the cmake config breaks under meson's CMake probe — the fallback is the supported path on ROCm 7.x. - Re-test on rebase:
PATH=/opt/rocm/bin:$PATH meson setup build --reconfigure \
-Denable_hip=true -Denable_cuda=false -Denable_sycl=false
ninja -C build
meson test -C build test_hip_smoke
The smoke test self-skips the device-resident assertions when vmaf_hip_device_count() == 0, so it stays portable across CI runners that don't expose an AMD GPU.
saliency_student_v2 — Resize-decoder ablation (ADR-0364, 2026-05-09)¶
- Touches:
ai/scripts/train_saliency_student_v2.py(new),model/tiny/saliency_student_v2.{onnx,json}(new),model/tiny/saliency_student_v2_card.md(new),model/tiny/registry.json(new row),docs/ai/models/saliency_student_v2.md(new),docs/adr/0364-saliency-student-v2-resize-decoder.md(new),docs/research/0089-saliency-student-v2-resize-decoder.md(new),changelog.d/added/saliency-student-v2.md(new). All paths are fork-only — no upstream-mirrored files touched. - Invariant: v1 (
saliency_student_v1.onnx, registry idsaliency_student_v1,smoke: false) stays as the production weights for the C-sidemobilesalextractor. v2 is a parallel artefact undermodel/tiny/; promotion to production is a separate PR. The trainer's_ResizeConvmodule produces an ONNX graph withResize(mode=linear,coordinate_transformation_mode=half_pixel) — every op stays oncore/src/dnn/op_allowlist.cpost-ADR-0258. - On upstream sync: no rebase impact — Netflix has no parallel saliency-student model, no consumer of
Resizein the upstream ONNX surface, and nomodel/tiny/registry in the upstream tree. If Netflix ever lands a saliency model, the fork'ssaliency_student_v{1,2}rows stay independent. - Re-test on rebase:
.venv/bin/python ai/scripts/validate_model_registry.py
.venv/bin/python - <<'EOF'
import onnx
g = onnx.load('model/tiny/saliency_student_v2.onnx')
ops = sorted({n.op_type for n in g.graph.node})
assert 'Resize' in ops and 'ConvTranspose' not in ops, ops
print('v2 ONNX op-set:', ops)
EOF
Predictor v2 — real-corpus LOSO trainer + ADR-0303 gate (2026-05-08)¶
- Touches:
ai/scripts/train_predictor_v2_realcorpus.py(new),ai/scripts/run_predictor_v2_training.sh(new),ai/tests/test_train_predictor_v2_realcorpus.py(new),docs/adr/0303-fr-regressor-v2-ensemble-prod-flip.md(Status-update appendix only — body frozen per ADR-0028),changelog.d/added/predictor-v2-realcorpus-trainer.md(new). No upstream-shared paths; the trainer lives entirely under fork-localai/scripts/. - Invariant: the gate constants
SHIP_GATE_MEAN_PLCC = 0.95,SHIP_GATE_PLCC_SPREAD_MAX = 0.005,SHIP_GATE_PER_FOLD_MIN = 0.95,LOSO_FOLD_COUNT = 5mirror ADR-0303 §Decision and the constants inscripts/ci/ensemble_prod_gate.py. They MUST stay in lockstep; if a future ADR changes the gate, update both files (the predictor trainer + the ensemble CI gate) and re-runtest_gate_constants_match_adr_0303. The 14-codec list in_resolve_codecs()is sourced fromvmaftune.predictor._DEFAULT_COEFFSwhen PR #450 is on the path; the hard-coded fallback exists for the bootstrap case where this script lands before #450 merges. Drift between the two is asserted at runtime — adding a 15th codec means updating the mirror. - On upstream sync: no action required. The trainer + tests live entirely under fork-local paths (
ai/scripts/,ai/tests/); upstream Netflix/vmaf has no equivalent surface. PR #450 (the predictor train pipeline) is itself fork-local; an upstream sync that reorganisesai/scripts/would invalidate the relative imports — re-run the test suite if that happens. - Re-test on rebase:
```bash python -m pytest ai/tests/test_train_predictor_v2_realcorpus.py -q bash -n ai/scripts/run_predictor_v2_training.sh python ai/scripts/train_predictor_v2_realcorpus.py --synthetic-smoke --report-out /tmp/p2.json
ADR-0332 — OpenVINO NPU EP wired into tiny-AI dispatch (2026-05-08)¶
- Touches:
core/include/libvmaf/dnn.h,core/src/dnn/ort_backend.{c,h},core/tools/vmaf.c,core/tools/cli_parse.{c,h},core/test/dnn/test_ep_fp16.c,core/test/dnn/test_cli.sh,docs/ai/inference.md,docs/usage/cli.md,docs/development/oneapi-install.md,docs/adr/0332-openvino-npu-ep-wiring.md(new),docs/adr/_index_fragments/0332-openvino-npu-ep-wiring.md(new),changelog.d/added/openvino-npu-ep.md(new). The libvmafdnn/and tools surfaces are fork-local additions; upstream Netflix/vmaf has no tiny-AI / ONNX Runtime dispatch layer, so conflict probability ondnn/is zero. - Invariant:
VmafDnnDeviceenum values9..11(OPENVINO_NPU/OPENVINO_CPU/OPENVINO_GPU) are appended after CoreML5..8. ABI requires these values stay stable across releases — append-only; never renumber. The--tiny-devicevalidator incli_parse.c::ARG_TINY_DEVICEenumerates the keyword set; new keywords append to the validator AND to the help string AND toresolve_tiny_device()invmaf.ctogether. Thevmaf_dnn_session_attached_ep()stable-string list (docs/ai/inference.md+dnn.hdoxygen) gains"OpenVINO:NPU"— consumers asserting on the returned string MUST update. - On upstream sync: no action required for upstream Netflix/vmaf. If a future Netflix sync introduces an unrelated tiny-AI surface (unlikely), reconcile the EP-name list at the merge.
- Re-test on rebase:
cd libvmaf && \
CC=icx CXX=icpx meson setup build -Denable_sycl=true -Denable_cuda=false && \
ninja -C build && \
./build/test/dnn/test_ep_fp16 && \
./build/tools/vmaf --tiny-device=openvino-npu --tiny-device=openvino-cpu \
--tiny-device=openvino-gpu # validator must accept all three keywords
ADR-0365 — CoreML execution provider wiring (2026-05-09)¶
- Touches:
core/include/libvmaf/dnn.h,core/src/dnn/ort_backend.{c,h},core/tools/cli_parse.{c,h},core/tools/vmaf.c,core/test/dnn/test_ep_fp16.c,core/test/dnn/test_cli.sh,docs/ai/inference.md,docs/usage/cli.md. Coordinates with ADR-0332 (OpenVINO NPU EP, PR #496) — both touch the same files; conflicts are mechanical (adjacent enum values, adjacent switch cases, adjacent CLI keyword strings). OpenVINO NPU/CPU/GPU values are 9..11 (after CoreML 5..8). - Invariant:
VmafDnnDeviceenum is append-only. CoreML values are 5..8; OpenVINO pinned variants are 9..11. TheSessionOptionsAppendExecutionProvider("CoreMLExecutionProvider", …)generic form is deliberate so the Linux build needs nocoreml_provider_factory.hinclude. TheMLComputeUnitskey string values (CPUAndNeuralEngine/CPUAndGPU/CPUOnly) are part of the CoreML EP public contract — upstream renames would break the wiring. The AUTO chain inserts CoreML at the last position (after CUDA / OpenVINO / ROCm); reordering changes the Apple-silicon AUTO outcome. - Re-test on rebase:
cd libvmaf && meson setup build -Denable_dnn=auto \
-Denable_cuda=false -Denable_sycl=false \
-Dbuilt_in_models=false && \
ninja -C build && \
./build/test/dnn/test_ep_fp16 && \
VMAF_BIN=$PWD/build/tools/vmaf bash test/dnn/test_cli.sh && \
./build/tools/vmaf --tiny-device coreml-ane 2>&1 | \
grep -q 'Reference' && \
./build/tools/vmaf --tiny-device bogus 2>&1 | \
grep -q 'coreml'
python3 -m pytest tools/external-bench/tests/ -q # must report 7 passed
bash -n tools/external-bench/*/run.sh
0361 — Metal (Apple Silicon) backend scaffold (ADR-0361)¶
- Touches:
core/include/libvmaf/libvmaf_metal.h(new, fork-local) — public C-API for the Metal backend (vmaf_metal_state_init/_import_state/_state_free/vmaf_metal_list_devices/vmaf_metal_available). Mirrors the HIP / Vulkan / SYCL / CUDA public-header convention; opaque runtime types cross the ABI asuintptr_tper ADR-0361 / ADR-0212 / ADR-0184.core/src/metal/{common,picture_metal,dispatch_strategy,kernel_template}.{c,h}AGENTS.md+meson.build(new, fork-local) — backend tree. Every entry point returns-ENOSYS. Thekernel_templatefield shape mirrors the HIP twin modulo the unified-memory buffer collapse (oneMTLBufferwithMTLResourceStorageModeSharedinstead of the (device, pinned-host) readback pair).
core/src/feature/metal/integer_motion_v2_metal.c(new, fork-local) — first kernel-template consumer. Mirrorsfeature/hip/integer_motion_v2_hip.ccall-graph-for-call-graph modulo the single-buffer prev-ref slot (vs the HIP twin'spix[2]ping-pong).core/test/test_metal_smoke.c(new, fork-local) — 14-sub-test smoke pinning the-ENOSYScontract. Mirrorstest_hip_smoke.c.core/meson_options.txt— newenable_metalfeature option (defaultauto). Onautothe parent meson resolves tohost_machine.system() == 'darwin'so non-macOS hosts compile cleanly without the frameworks;enabledforces linkage and fails on non-macOS. Type-featurematchesenable_dnn's auto-resolve shape (Metal on macOS is always available, like DNN on a host with ONNX Runtime); the GPU-vendor-pair boolean-default-off triad (enable_cuda/enable_sycl/enable_hip) does not fit because Metal has no comparable "wrong-host silent flip" risk.core/src/meson.build—is_metal_enabledresolution +subdir('metal')+metal_sources/metal_depsthreaded throughlibvmaf_feature_static_libandlibvmaflibrary() calls alongside CUDA / SYCL / Vulkan / HIP / DNN aggregations.core/test/meson.build—test_metal_smokeexecutable wired under the same auto-on-macOS / explicit-enabled gate.core/src/feature/feature_extractor.c— addsextern VmafFeatureExtractor vmaf_fex_integer_motion_v2_metal;- registry entry under
#if HAVE_METAL.
- registry entry under
.github/workflows/libvmaf-build-matrix.yml— new laneBuild — macOS Metal (T8-1 scaffold)onmacos-latestwith-Denable_metal=enabled. Themacos-latestrunner ships the Metal SDK as part of the system framework set; no extra install step is needed.docs/backends/metal/index.md(new, fork-local) +docs/backends/index.md(row added) + ADR-0361 + index fragment +changelog.d/added/metal-backend-scaffold.md+docs/state.mdrow T8-1b.- Upstream-port footprint: zero — Netflix/vmaf does not ship a Metal backend; this is a wholly fork-local addition. No upstream file is touched. Same posture as the HIP scaffold (T7-10) and the Vulkan scaffold (T5-1).
- Rebase invariants (mirror the HIP scaffold's invariant set):
metal/kernel_template.hmirrorship/kernel_template.hmodulo the unified-memory buffer collapse (singleMTLBufferslot vs the HIP(device, pinned-host)pair). On rebase, if the HIP twin's lifecycle struct gains a third event slot, the Metal twin must follow in the same PR.feature/metal/integer_motion_v2_metal.cmirrorsfeature/hip/integer_motion_v2_hip.ccall-graph-for-call-graph modulo the single-prev_ref-slot collapse (vs the HIP twin'spix[2]ping-pong). On rebase, drift in the HIP twin's submit body (e.g. an addedsubmit_pre_launchcall) requires a paired update here.vmaf_fex_integer_motion_v2_metalregisters without theVMAF_FEATURE_EXTRACTOR_METALflag bit set. The flag bit is reserved for the runtime PR (T8-1b) which adds theVMAF_PICTURE_BUFFER_TYPE_METAL_DEVICEtag and then sets the flag. Same posture as the HIP twin'sVMAF_FEATURE_EXTRACTOR_HIP-deferral; on rebase, leave the flags atVMAF_FEATURE_EXTRACTOR_TEMPORALonly until T8-1b.- Re-test (on macOS only — Linux dev sessions cannot run this lane locally):
And on every host (Linux / Windows included): the default-build gate must stay green — the auto-probe resolves to disabled on non-macOS hosts so meson setup build && ninja -C build runs unchanged.
ADR-0325 — vmaf-tune auto Phase F.1 + F.2 short-circuits (2026-05-08)¶
0327 — Conformal-VQA prediction surface for vmaf-tune (ADR-0279)¶
- Touches:
tools/vmaf-tune/src/vmaftune/conformal.py(new),tools/vmaf-tune/src/vmaftune/predictor.py(Predictor.predict_vmaf_with_uncertainty),tools/vmaf-tune/src/vmaftune/cli.py(predictsubcommand gains--with-uncertainty/--calibration-sidecar/--alpha),tools/vmaf-tune/tests/test_conformal.py(new),docs/ai/conformal-vqa.md(new). No engine code touched; no upstream-shared paths. - Invariant: the conformal wrapper sits outside the ONNX graph and adds no new runtime dependency —
conformal.pyimports only the standard library (math,statistics,dataclasses,json,warnings). Future calibration-sidecar shapes use themethoddiscriminator string for versioning; do not rename"split-conformal"/"cv-plus"without bumping the loader. ThePredictor.predict_vmaf_with_uncertaintysignature is the Python-API contract consumed byvmaf-tune predict --with-uncertainty; renaming or reordering its keyword args breaks the CLI in lockstep. - On upstream sync: no action required.
vmaf-tuneis a fork-local tool; upstream Netflix/vmaf has no per-shot prediction surface. - Re-test on rebase:
python3 -m pytest tools/vmaf-tune/tests/test_conformal.py -q
python3 -m pytest tools/vmaf-tune/tests/test_predictor.py -q
ADR-0364 — vmaf-tune auto Phase F.1 + F.2 short-circuits (2026-05-08)¶
- Touches:
tools/vmaf-tune/src/vmaftune/auto.py(new),tools/vmaf-tune/src/vmaftune/cli.py(addedautosubparser + dispatcher),tools/vmaf-tune/tests/test_auto_short_circuits.py(new),tools/vmaf-tune/AGENTS.md(invariant row),docs/usage/vmaf-tune.md(## autosection),docs/adr/0364-vmaf-tune-phase-f-auto.md(status update — already-accepted body untouched per ADR-0028; appended a### Status updateblock under## References). No upstream-shared paths.
ADR-0325 — vmaf-tune auto Phase F.1 + F.2 short-circuits (2026-05-08)¶
ADR-0371 — Shared CorpusIngestBase (2026-05-10)¶
No rebase impact: pure Python refactor under ai/ — no C/header/patch changes, no upstream-shared paths touched. All six MOS-corpus adapter scripts now import from corpus.base import CorpusIngestBase (PYTHONPATH=ai/src); if a future upstream sync adds a corpus/ directory under ai/ the import path may collide but the risk is negligible (Netflix/vmaf does not carry an ai/ subtree).
- Touches:
tools/vmaf-tune/src/vmaftune/auto.py(new),tools/vmaf-tune/src/vmaftune/cli.py(addedautosubparser + dispatcher),tools/vmaf-tune/tests/test_auto_short_circuits.py(new),tools/vmaf-tune/AGENTS.md(invariant row),docs/usage/vmaf-tune.md(## autosection),docs/adr/0325-vmaf-tune-phase-f-auto.md(status update — already-accepted body untouched per ADR-0028; appended a### Status updateblock under## References). No upstream-shared paths. - Invariant:
SHORT_CIRCUIT_PREDICATESinauto.pyis an ordered tuple, not a set. The seven entries appear in the canonical orderLADDER_SINGLE_RUNG,CODEC_PINNED,PREDICTOR_GOSPEL,SKIP_SALIENCY,SDR_SKIP,SAMPLE_CLIP_PROPAGATE,SKIP_PER_SHOT. The JSON schema records short-circuits in this order underplan.metadata.short_circuits; downstream consumers (CI corpus collector, post-hoc speedup analysis) parse the canonical-order list. Adding an eighth short-circuit (F.3+) appends; never reorder. The Phase D thresholds (PHASE_D_DURATION_GATE_S = 300.0,PHASE_D_SHOT_VARIANCE_GATE = 0.15) are placeholders pending F.3 empirical fit. - On upstream sync: no action required. Module is fork-local (
tools/vmaf-tune/is fork-only). Thevmaf-tuneumbrella ADR-0237 explicitly carves Phases B–F out of upstream scope. - Re-test on rebase:
cd tools/vmaf-tune && python -m pytest tests/test_auto_short_circuits.py -v
PYTHONPATH=tools/vmaf-tune/src python -m vmaftune.cli auto \
--src /dev/null --target-vmaf 93 --max-budget-bitrate 5000 \
--allow-codecs libx264 --sample-clip-seconds 10 --smoke
ADR-0325 — vmaf-tune auto Phase F.3 confidence-aware fallbacks (2026-05-08)¶
- Touches:
tools/vmaf-tune/src/vmaftune/auto.py(F.3 helpers,_confidence_aware_escalation,ConfidenceThresholds,ConfidenceDecision,load_confidence_thresholds, per-cell wiring inrun_auto),tools/vmaf-tune/tests/test_auto_confidence_aware.py(new, 28 tests),tools/vmaf-tune/AGENTS.md(invariant note),docs/usage/vmaf-tune.md(new### Confidence-aware fallbacks (F.3)subsection under## auto),docs/adr/0325-vmaf-tune-phase-f-auto.md(status update appended per ADR-0028; already-Accepted body untouched),changelog.d/added/phase-f3-confidence-aware-fallbacks.md(new). No upstream-shared paths. - Invariant:
DEFAULT_TIGHT_INTERVAL_MAX_WIDTH = 2.0andDEFAULT_WIDE_INTERVAL_MIN_WIDTH = 5.0are an emergency floor (Research-0067), not a target. The production values come from a JSON calibration sidecar produced by the conformal-VQA pipeline (ADR-0279) with the canonical keystight_interval_max_widthandwide_interval_min_width.load_confidence_thresholdsfalls back to the defaults with a one-line WARNING when no sidecar is found; do not silence the warning._confidence_aware_escalationis a pure function of its three inputs and is exposed in__all__so downstream tools (the MCP server'sautoproxy, the CI corpus collector) can embed it directly. The JSON schema records per-cell decisions inplan.metadata.confidence_aware_escalations[](one entry per(rung, codec)cell with keysrung,codec,verdict,interval_width,decision); each cell inplan.cells[]also carriesconfidence_decision+interval_widthso consumers don't need to cross-reference the metadata array index. Adding a fourthConfidenceDecisionvalue is a schema bump — coordinate with downstream JSON consumers. - On upstream sync: no action required.
tools/vmaf-tune/is fork-only; the conformal-VQA prediction surface (ADR-0279) and the F.1 + F.2 scaffold (ADR-0325) are both fork-local. - Re-test on rebase:
cd tools/vmaf-tune && python -m pytest \
tests/test_auto_confidence_aware.py \
tests/test_auto_short_circuits.py \
tests/test_conformal.py -v
ADR-0325 — vmaf-tune auto Phase F.4 per-content-type recipe overrides (2026-05-09)¶
- Touches:
tools/vmaf-tune/src/vmaftune/auto.py(added_apply_recipe_override,_CONTENT_RECIPE_TABLE,get_recipe_for_class, the four_<class>_recipefactories, and theRECIPE_CLASS_*constants; integrated the override intorun_autoand addedrecipe_applied/effective_predictor_target_vmafto the JSON metadata),tools/vmaf-tune/tests/test_auto_recipe_overrides.py(new — 37 assertions),tools/vmaf-tune/tests/test_auto_short_circuits.py(one test updated for the F.4 force-single-rung semantics on animation sources),tools/vmaf-tune/AGENTS.md(invariant row),docs/usage/vmaf-tune.md(### Per-content-type recipes (F.4)subsection),docs/adr/0325-vmaf-tune-phase-f-auto.md(status update appended; already-accepted body untouched per ADR-0028),changelog.d/added/phase-f4-content-recipes.md. No upstream-shared paths. - Invariant:
_CONTENT_RECIPE_TABLEstores factory callables, not literal dicts. Everyget_recipe_for_class/_apply_recipe_overridecall returns a fresh override dict so caller mutations cannot leak between runs. The four override keys honoured by the driver aretight_interval_max_width,force_single_rung,saliency_intensity,target_vmaf_offset; the_RECIPE_KEYSallowlist filters anything else as defence-in-depth. Thetarget_vmaf_offsetshifts onlyeffective_predictor_target_vmaf; the inputtarget_vmaf(production-flip gate) is preserved verbatim. Every threshold value at F.4 is provisional pending F.5 calibration — do not promote a placeholder to "calibrated" in a drive-by edit. - On upstream sync: no action required.
tools/vmaf-tune/is fork-local; ADR-0237 explicitly carves Phases B–F out of upstream scope. - Re-test on rebase:
PYTHONPATH=tools/vmaf-tune/src python -m pytest \
tools/vmaf-tune/tests/test_auto_recipe_overrides.py \
tools/vmaf-tune/tests/test_auto_short_circuits.py \
tools/vmaf-tune/tests/test_auto_confidence_aware.py -v
PYTHONPATH=tools/vmaf-tune/src python -c \
"from pathlib import Path; from vmaftune.auto import run_auto, SourceMeta; \
m = SourceMeta(height=1080, width=1920, content_class='animation', duration_s=120, shot_variance=0.05); \
p = run_auto(src=Path('/dev/null'), target_vmaf=93.0, max_budget_kbps=5000.0, \
allow_codecs=('libx264',), smoke=True, meta_override=m); \
assert p.metadata['recipe_applied'] == 'animation'; \
assert p.metadata['target_vmaf'] == 93.0; \
assert p.metadata['effective_predictor_target_vmaf'] == 95.0; \
print('F.4 smoke OK')"
ADR-0325 — vmaf-tune auto Phase F.5 calibrated recipe overrides (2026-05-09)¶
- Touches:
ai/scripts/calibrate_phase_f_recipes.py(new),ai/data/phase_f_recipes_calibrated.json(new — tracked via the.gitignore!ai/data/phase_f_recipes_calibrated.jsonallow rule),tools/vmaf-tune/src/vmaftune/auto.py(added_F4_PLACEHOLDER_RECIPES,_CALIBRATED_RECIPES_FILENAME,_find_calibrated_recipes_path,_load_calibrated_recipes,_CALIBRATED_RECIPES; the four_<class>_recipefactories now read from_CALIBRATED_RECIPES),tools/vmaf-tune/tests/test_calibrated_recipes.py(new — 14 assertions),docs/usage/vmaf-tune.md(calibrated table replaces the F.4 placeholder table in the### Per-content-type recipes (F.4)subsection),docs/adr/0325-vmaf-tune-phase-f-auto.md(### Status update 2026-05-09: F.5 calibratedappended; already-accepted body untouched per ADR-0028),changelog.d/changed/phase-f5-calibrated-recipes.md,.gitignore(one allow rule for the JSON file). No upstream-shared paths. - Invariant: the
_CONTENT_RECIPE_TABLEfactories now consume_CALIBRATED_RECIPESsnapshotted at module import. The runtime load is a single read; reloading at runtime requiresimportlib.reload(vmaftune.auto). Everyget_recipe_for_class/_apply_recipe_overridecall still returns a fresh dict — the read-only invariant from F.4 is preserved bydict(_CALIBRATED_ RECIPES[<cls>]). The_load_calibrated_recipesloader strips every_provenancesub-dict and filters every key against_RECIPE_KEYSso a malicious or malformed JSON cannot inject unknown keys into a recipe. Per memoryfeedback_no_test_weakening, the calibration cannot widen the production-flip gate beyond the ConfidenceThresholds wide-interval ceiling — the regression testtest_calibrated_ugc_width_below_wide_gate_ceilinglocks this in. - On upstream sync: no action required.
tools/vmaf-tune/,ai/scripts/,ai/data/are all fork-local; ADR-0237 explicitly carves Phases B–F out of upstream scope. - Re-test on rebase:
PYTHONPATH=tools/vmaf-tune/src python -m pytest \
tools/vmaf-tune/tests/test_calibrated_recipes.py \
tools/vmaf-tune/tests/test_auto_recipe_overrides.py -v
python ai/scripts/calibrate_phase_f_recipes.py \
--corpus .workingdir2/konvid-150k/konvid_150k.jsonl \
--out /tmp/recipes_smoke.json \
--max-rows 10000
ADR-0335 — Hardware-capability priors (2026-05-08)¶
- Touches:
ai/data/hardware_caps.csv(new),ai/scripts/hardware_caps_loader.py(new),ai/tests/test_hardware_caps.py(new),ai/AGENTS.md(one new bullet under "Rebase-sensitive invariants"),docs/ai/hardware-capability-priors.md(new),docs/research/0088-hardware-capability-priors-2026-05-08.md(new),docs/adr/0335-hardware-capability-priors.md(new),docs/adr/_index_fragments/0335-hardware-capability-priors.md(new),docs/adr/_index_fragments/_order.txt(one-line append),CHANGELOG.md(Added bullet under[Unreleased] — lusoris fork). No upstream-shared paths. - Invariant: the table is prior-only. The schema check in
hardware_caps_loader.pyrejects benchmark-shaped header columns (fps_*,throughput,mbps,latency,watts,tdp,score_*,vmaf_*), community-wiki source URLs (wikipedia.org,wikichip.org), empty fields, and rows withencoding_blocks=0. Adding throughput / quality columns is forbidden — that pathology was the contributor-pack digest's category-1 NO-GO finding. Schema extensions need a new ADR, not a silent column bump. Thecap_vector_for()return-dict shape is load-bearing: trainers / corpus writers consumehwcap_*columns by name; reordering or renaming silently breaks downstream parquet schemas. - On upstream sync: no action required. The whole surface lives under
ai/anddocs/— Netflix upstream has no equivalent. - Re-test on rebase:
```bash python -m pytest ai/tests/test_hardware_caps.py -v # must report 23 passed python ai/scripts/hardware_caps_loader.py # JSON dump, 6+ rows
ADR-0367 — LSVQ corpus ingestion (2026-05-08)¶
- Touches:
ai/scripts/lsvq_to_corpus_jsonl.py(new),ai/tests/test_lsvq.py(new),docs/adr/0367-lsvq-corpus-ingestion.md(new),docs/adr/README.md(regenerated index),docs/ai/lsvq-ingestion.md(new),docs/research/0090-lsvq-corpus-feasibility.md(new),changelog.d/added/0367-lsvq-ingestion.md(new). No engine code touched; no upstream-shared paths. - Invariant: the JSONL row schema emitted by this adapter is byte-identical to the KonViD-150k Phase 2 adapter (
ai/scripts/konvid_150k_to_corpus_jsonl.py) modulo thecorpusandcorpus_versionliterals. If a future PR widens the row contract (new column, type change), the LSVQ adapter must follow in lockstep — the trainer-side data loader consumes both shards through one schema. - On upstream sync: no action required. The adapter lives entirely under fork-local paths (
ai/scripts/,ai/tests/) and only consumes a fork-local CSV manifest. - Re-test on rebase:
ADR-0325 — Local sidecar training scaffold (2026-05-08)¶
- Touches:
tools/vmaf-tune/src/vmaftune/sidecar.py(new),tools/vmaf-tune/tests/test_sidecar.py(new),docs/adr/0325-local-sidecar-training.md(new),docs/adr/_index_fragments/0325-local-sidecar-training.md(new),docs/adr/_index_fragments/_order.txt(append),docs/adr/README.md(index row),docs/research/0086-local-sidecar-feasibility.md(new),docs/ai/local-sidecar-training.md(new),changelog.d/added/local-sidecar-training-scaffold.md(new),tools/vmaf-tune/AGENTS.md(sidecar invariant note). No engine code touched; no upstream-shared paths. - Invariant: the sidecar's on-disk state schema (
SIDECAR_SCHEMA_VERSION = 1,FEATURE_DIM = 14, the column order in_feature_vector) is the load-bearing pin. Adding columns or reordering them must bumpSIDECAR_SCHEMA_VERSION; otherwise saved state from older harness versions silently aligns mismatched columns to the wrong feature. TheSidecarConfig.predictor_versiontag is the load-bearing pin against shipped-predictor upgrades — bumping it is the contract that invalidates stale corrections without operator intervention. - On upstream sync: no action required. The sidecar lives entirely under
tools/vmaf-tune/(fork-local) and only consumes the existingPredictor/ShotFeaturessurface. Upstream Netflix/vmaf does not ship avmaf-tuneanalogue; conflict probability is zero. - Re-test on rebase:
```bash cd tools/vmaf-tune && python -m pytest tests/test_sidecar.py -v
ADR-0368 — YouTube UGC corpus ingestion (2026-05-08)¶
- Touches:
ai/scripts/youtube_ugc_to_corpus_jsonl.py(new),ai/tests/test_youtube_ugc.py(new),docs/adr/0368-youtube-ugc-corpus-ingestion.md(new),docs/adr/_index_fragments/0368-youtube-ugc-corpus-ingestion.md(new),docs/adr/_index_fragments/_order.txt(one-line append),docs/adr/README.md(regenerated index),docs/ai/youtube-ugc-ingestion.md(new),docs/research/0091-youtube-ugc-corpus-feasibility.md(new),changelog.d/added/0368-youtube-ugc-ingestion.md(new),ai/AGENTS.md(one-paragraph invariant). No engine code touched; no upstream-shared paths. - Invariant: the JSONL row schema emitted by this adapter is byte-identical to the LSVQ adapter (
ai/scripts/lsvq_to_corpus_jsonl.py, ADR-0367) and the KonViD-150k Phase 2 adapter modulo thecorpusandcorpus_versionliterals. If a future PR widens the row contract (new column, type change), all adapters must follow in lockstep.
ADR-0369 — Waterloo IVC 4K-VQA corpus ingestion (2026-05-08)¶
- Touches:
ai/scripts/waterloo_ivc_to_corpus_jsonl.py(new),ai/tests/test_waterloo_ivc.py(new),docs/adr/0369-waterloo-ivc-4k-corpus-ingestion.md(new),docs/adr/_index_fragments/0369-waterloo-ivc-4k-corpus-ingestion.md(new),docs/adr/_index_fragments/_order.txt(one-line append),docs/adr/README.md(regenerated index),docs/ai/waterloo-ivc-4k-ingestion.md(new),docs/research/0091-waterloo-ivc-4k-corpus-feasibility.md(new),changelog.d/added/0369-waterloo-ivc-4k-ingestion.md(new),ai/AGENTS.md(one-paragraph invariant). No engine code touched; no upstream-shared paths. -
Invariant: JSONL row schema is byte-identical to the LSVQ (ADR-0367) and YouTube-UGC (ADR-0368) adapters modulo the
corpusandcorpus_versionliterals. All adapters must change in lockstep on schema widening. -
On upstream sync: no action required.
- Re-test on rebase:
```bash
pytest ai/tests/test_youtube_ugc.py -v
pytest ai/tests/test_waterloo_ivc.py -v
ADR-0325 — predictor stub-models policy (2026-05-08)¶
- Touches:
tools/vmaf-tune/src/vmaftune/predictor_train.py(new),model/predictor_<codec>.onnx× 14 (new),model/predictor_<codec>_card.md× 14 (new),tools/vmaf-tune/tests/test_predictor_train.py(new),docs/ai/predictor.md(new),docs/adr/0325-predictor-stub-models-policy.md(new),docs/adr/README.md+_index_fragments/0325-*.md+_order.txt(index rows),changelog.d/added/predictor-train-pipeline.md(new). No engine code; no upstream-shared paths. - Invariant: the trainer's
CODECStuple is sourced frompredictor._DEFAULT_COEFFSso the two stay in lockstep. Any new codec adapter that lands inpredictor._DEFAULT_COEFFSmust (a) ship a matching synthetic-stub model + card undermodel/predictor_<codec>.{onnx,_card.md}in the same PR, and (b) re-run the trainer to refresh the artefact set. The shipped-model smoke test (test_predictor_loads_each_shipped_model) parameterises overCODECSand will fail if either condition is missed. - On upstream sync: no action required. The predictor + trainer live entirely under
tools/vmaf-tune/(a fork-local path); the model artefacts live undermodel/but use apredictor_<codec>.onnxnaming scheme that does not collide with any upstreammodel/vmaf_*.{json,pkl}ormodel/tiny/*.onnxpath. - Re-test on rebase:
```bash python3 -m pytest tools/vmaf-tune/tests/test_predictor_train.py -q python3 -c " import sys sys.path.insert(0, 'tools/vmaf-tune/src') from vmaftune.predictor_train import main sys.exit(main(['--output-dir', '/tmp/predictor-rebase', '--epochs', '20'])) "
ADR-0325 — vmaf-tune Phase B target-VMAF bisect (2026-05-08)¶
ADR-0297 — vmaf-tune Phase B target-VMAF bisect (2026-05-08)¶
- Touches:
tools/vmaf-tune/src/vmaftune/bisect.py(new),tools/vmaf-tune/src/vmaftune/compare.py(default-predicate error string),tools/vmaf-tune/tests/test_bisect.py(new),tools/vmaf-tune/tests/test_compare.py(renamed default-predicate assertion),tools/vmaf-tune/AGENTS.md(Phase B invariant),docs/adr/0326-vmaf-tune-phase-b-bisect.md(new),docs/adr/_index_fragments/0326-vmaf-tune-phase-b-bisect.md(new),docs/adr/_index_fragments/_order.txt(append),docs/research/0090-vmaf-tune-phase-b-bisect-feasibility.md(new),docs/usage/vmaf-tune-bisect.md(new),changelog.d/added/vmaf-tune-phase-b-bisect.md(new). No upstream Netflix/vmaf surface is touched. - Invariant: the bisect assumes monotone-decreasing VMAF in CRF. Two non-adjacent samples that violate this contract abort the call with a clear error rather than falling back to a different search strategy. Do NOT add a fallback path on rebase — the AGENTS.md Phase B note is load-bearing.
- Companion seam:
compare._default_predicateno longer raisesNotImplementedError("Phase B pending"); it returns a well-formedRecommendResult(ok=False, error=...)pointing callers atmake_bisect_predicate. Any downstream tests that asserted "Phase B pending" verbatim need updating. - On upstream sync: no action required. The module lives entirely under
tools/vmaf-tune/(a fork-local path). - Re-test on rebase:
python3 -m pytest tools/vmaf-tune/tests/test_bisect.py -v
python3 -m pytest tools/vmaf-tune/tests/test_compare.py -v
feat/sycl-integer-cambi-port — CAMBI SYCL twin (T3-15 / ADR-0371, 2026-05-10)¶
- Touches:
core/src/feature/sycl/integer_cambi_sycl.cpp(new file),core/src/feature/feature_extractor.c(extern declaration + list entry under#if HAVE_SYCL),core/src/meson.build(source addition to the SYCL feature list),core/test/test_integer_cambi_sycl.c(new smoke test),core/test/meson.build(test target +gpu_all_depsrefactor),docs/backends/sycl/overview.md(Known gaps update),docs/adr/0371-cambi-sycl-port.md(new ADR). - Invariant:
vmaf_fex_cambi_syclmust remain registered before any Vulkan or CUDA CAMBI extractor infeature_extractor_list[]so SYCL is preferred when the runtime selects a GPU backend. The ordering#if HAVE_SYCL … &vmaf_fex_cambi_syclbefore#if HAVE_VULKAN/#if HAVE_CUDAis load-bearing. Additionally: the host residual callsvmaf_cambi_calculate_c_valuesandvmaf_cambi_spatial_poolingviacambi_internal.htrampoline — if upstream Netflix ever renames or removes those symbols the SYCL twin will silently stop compiling. - Upstream conflict probability: low. Netflix upstream does not carry a
core/src/feature/sycl/directory. The only upstream-shared paths touched arefeature_extractor.c(extern + list entry) andcambi_internal.h(consumed, not modified). A conflict onfeature_extractor.cwould be an upstream addition of a new extractor; resolve by re-inserting thevmaf_fex_cambi_syclentry under#if HAVE_SYCL. - Re-test on rebase:
meson setup build -Denable_sycl=true -Denable_cuda=false && ninja -C build
meson test -C build --suite=fast
fix/float-adm-extractor-loading — enable_float default flip (2026-05-09)¶
No rebase-sensitive invariants. The change is a single default-value flip in core/meson_options.txt (enable_float: false → enable_float: true) and a prose update to docs/development/build-flags.md. No C source was modified; no build-system paths changed; no new symbols were added.
- On upstream sync: if Netflix upstream ever adds their own
enable_floatdefault change, prefer theirs and drop this entry. - Re-test on rebase: run the reproducer —
./build/tools/vmaf --feature float_adm --no_prediction ...— and confirm it no longer prints "problem loading feature extractor".
ADR-0297 — MyTestCase upstream migration (partial port, Batch E, 2026-05-08)¶
- Touches:
python/test/testutil.py,python/test/bd_rate_calculator_test.py,python/test/asset_test.py,python/test/bootstrap_train_test_model_test.py,python/test/local_explainer_test.py,python/test/cy_test.py,python/test/executor_test.py,python/test/raw_extractor_test.py,python/test/cross_validation_test.py,python/test/niqe_train_test_model_test.py,python/vmaf/script/run_testing.py,python/vmaf/tools/misc.py,python/vmaf/tools/testutils.py. Five Netflix golden-pinned files (quality_runner_test.py,vmafexec_test.py,vmafexec_feature_extractor_test.py,feature_extractor_test.py,result_test.py) are deliberately untouched. - Invariant: every
assertAlmostEqual(key, value)pair in the five golden-pinned files remains byte-identical to the fork's pre-port state per ADR-0024. Verified via/tmp/mytestcase-port/verify_golden.pyagainst the multiset baseline/tmp/mytestcase-port/baseline-pairs.json: all 310 + 183 + 37 + 113 + 17 = 660 pairs PASS post-port. CLAUDE.md §1 / §8 forbid altering them. - Deferred upstream commits (still need porting in a future session, in chronological order):
7d1ad54b(port aim/adm3/motion3 fextractor tests),9fa593eb(more aim/adm3/motion3 + new options),0341f730(remove duplicate test_run_vmaf_integer_fextractor),a3776335+74bdce1b(align fork tests with upstream layout for aim/adm3/motion3),322ca041(replace temporal slicing with pre-sliced YUV fixtures),6c097fc4+ead2d12b+4679db83(macOS FP tolerance widenings — many of these are no-ops for our fork because the affected lines do not exist in the fork's current state),005988ea(routine_test MyTestCase + fifo_mode),3a041a97+3e075107(per the user's 2026-05-08 instruction list, these "update score values" upstream commits are PERMANENTLY skipped — porting them would violate ADR-0024). The3cbf352d+eb3374d0anda333ba4c+403dafedrevert-pairs are no-ops upstream and require no port. TheMyTestCasemixin itself is already inpython/vmaf/tools/misc.pyfrom a prior fork-local sync. - Watch out for: when retrying the deferred port, the fork's
feature_extractor_test.pytest method order (psnr -> ansnr -> ssim -> ssim_flat -> ms_ssim) differs from upstream's post-cluster layout (psnr -> ssim -> ms_ssim -> ansnr); the cherry-pick conflicts cluster around this reordering. The aim/adm3/motion3 additive blocks should be transplanted as new test methods rather than merged into existing ones. The verifier script will catch any (key, value) pair drop. - Re-test on rebase:
```bash python3 /tmp/mytestcase-port/verify_golden.py # OVERALL: PASS required pytest python/test/bd_rate_calculator_test.py -v pytest python/test/asset_test.py -v pytest python/test/quality_runner_test.py python/test/vmafexec_test.py \ python/test/vmafexec_feature_extractor_test.py \ python/test/feature_extractor_test.py python/test/result_test.py \ --collect-only -q # 173 tests collected, no errors
ADR-0318 — fr_regressor_v2 ensemble retrain harness fix (2026-05-06)¶
- Touched files:
ai/scripts/run_ensemble_v2_real_corpus_loso.sh— wrapper passes--corpus "$CORPUS_JSONL"+--out-dir "$out_dir", drops--corpus-root/--output. JSONL-existence check replaces the YUV-directory hard-fail (YUV check is informational).docs/ai/ensemble-v2-real-corpus-retrain-runbook.md— adds step 0. Generate the Phase A canonical-6 corpus, expands prereqs table with the JSONL row + Phase A wall-time estimate.docs/adr/0318-ensemble-retrain-harness-fix.md,docs/adr/README.md(index row),changelog.d/fixed/ensemble-retrain-harness-interface.md.- Rebase invariant: not load-bearing. Wrapper-script + doc-only change. The trainer
ai/scripts/train_fr_regressor_v2_ensemble_loso.pyCLI is the authoritative interface (frozen here as part of the decision); any future change to it must update this wrapper in the same PR. - Upstream source: none — fork-local AI training harness.
- On upstream sync: no action required. Path is entirely under
ai/scripts/+docs/ai/+docs/adr/; upstream Netflix/vmaf does not ship these directories. - Re-test on rebase:
```bash bash -n ai/scripts/run_ensemble_v2_real_corpus_loso.sh python3 ai/scripts/train_fr_regressor_v2_ensemble_loso.py --help mkdir -p runs/phase_a/full_grid && \ touch runs/phase_a/full_grid/per_frame_canonical6.jsonl && \ bash ai/scripts/run_ensemble_v2_real_corpus_loso.sh 2>&1 | \ grep -q "unrecognized arguments" && echo "REGRESSED" || echo "OK" rm -rf runs/phase_a runs/ensemble_v2_real
0320 — fr_regressor_v2 ensemble seeds — production flip (ADR-0320)¶
- Touches:
model/tiny/registry.json(fivefr_regressor_v2_ensemble_v1_seed{0..4}rows flipped fromsmoke: truetosmoke: false),model/tiny/fr_regressor_v2_ensemble_v1_seed_flip_PROMOTE.json(new — committed verdict from the ADR-0319 harness run),ai/AGENTS.md(registry-flip invariant updated to record the flip - the going-forward "fresh PROMOTE.json required" rule),
docs/state.md(Recently closed row),docs/adr/0320-fr-regressor-v2-ensemble-seed-flip.md(new ADR). Closes the deferral tracked in rebase-notes §0303 / §0309 / §0319. - Upstream source: none — fork-local registry mutation honouring ADR-0303's flip contract. Netflix/vmaf upstream has no
fr_regressor_v2ensemble surface. - Invariant: any future change to the
fr_regressor_v2_ensemble_v1_seed{0..4}registry rows (sha256 bump after retraining, smoke-flag mutation, ONNX path change) requires a freshruns/ensemble_v2_real/PROMOTE.jsonverdict with mean per-seed LOSO PLCC ≥ 0.95 ANDmax - minspread ≤ 0.005 — the same two-part gate ADR-0303 defined and ADR-0320 honoured. Never mutate these rows during a/sync-upstreamrebase or as a side-effect of any other PR; the harness emits the verdict file but does not mutate the registry. The committed verdict atmodel/tiny/fr_regressor_v2_ensemble_v1_seed_flip_PROMOTE.jsonis the audit-trail anchor for the 2026-05-06 flip. - On upstream sync: no action required. The five rows live in
model/tiny/registry.jsonwhich is fork-local; upstream has no competing entries. If upstream ever ships its ownfr_regressor_v2_ensemble_v1_*registry rows, stop and consult the rebase reviewer — naming collision implies an architectural divergence that needs a Supersedes-ADR, not a mechanical merge. - Re-test on rebase:
python3 -c "import json; \
d = json.load(open('model/tiny/registry.json')); \
seeds = [m for m in d['models'] \
if m['id'].startswith('fr_regressor_v2_ensemble_v1_seed')]; \
assert len(seeds) == 5, seeds; \
assert all(m['smoke'] is False for m in seeds), seeds; \
print('OK: 5 ensemble seeds at smoke=false')"
python3 -c "import json; \
v = json.load(open('model/tiny/fr_regressor_v2_ensemble_v1_seed_flip_PROMOTE.json')); \
assert v['verdict'] == 'PROMOTE'; \
assert v['gate']['passed'] is True; \
print('OK: verdict still PROMOTE')"
--feature ssimulacra2 --backend cuda --places 4
ADR-0372 — HIP batch-1: integer_psnr_hip + float_ansnr_hip real kernels (2026-05-10)¶
- Touches:
core/src/feature/hip/integer_psnr_hip.c(full rewrite),core/src/feature/hip/float_ansnr_hip.c(full rewrite),core/src/feature/hip/integer_psnr_hip.h(HSACO symbol decl underHAVE_HIPCC),core/src/feature/hip/float_ansnr_hip.h(HSACO symbol decl underHAVE_HIPCC),core/src/feature/hip/integer_psnr/psnr_score.hip(new device kernel),core/src/feature/hip/float_ansnr/float_ansnr_score.hip(new device kernel),core/src/hip/kernel_template.{h,c}(vmaf_hip_kernel_submit_post_record— also in PR #612; on merge conflict keep one copy and drop the duplicate),core/src/meson.build(hip_hsaco_sourcesHSACO build pipeline — also in PR #612),docs/backends/hip/overview.md(status table update),docs/adr/0372-hip-batch1-integer-psnr-float-ansnr.md(new). - Invariant (HAVE_HIPCC dual-path): all device-state fields in
PsnrStateHip/AnsnrStateHipand allhipModule_t/hipFunction_tmember declarations live under#ifdef HAVE_HIPCC. Without the flag the scaffold-ENOSYScontract is preserved and the host TU compiles without ROCm SDK headers. On rebase or refactor, never move device-state fields outside the#ifdef HAVE_HIPCCguard — it breaks the CPU-only build. - Invariant (float_ansnr no-memset bypass):
float_ansnr_hip'ssubmit()does not callvmaf_hip_kernel_submit_pre_launch. The device kernel writes per-block(sig, noise)float partials directly into an output buffer (partials[2*block_idx+0]and[+1]); no atomic accumulation means no memset is needed. The partials buffer is sizedwg_count * 2u * sizeof(float)at init. On rebase: if a future PR adds asubmit_pre_launchcall tofloat_ansnr_cuda.c, the HIP twin must follow in the same PR. - Invariant (integer_psnr uint64 split shuffle): the PSNR kernel splits a uint64 warp-reduction into two uint32
__shfl_downcalls (HIP warp size = 64, no native uint64 shuffle). On rebase: if ROCm adds native uint64 shuffle primitives in a future release, the kernel can be simplified — but verify the cross-backend numeric gatemeson test -C build --suite=hip-paritystill passes before landing. - Merge-conflict risk with PR #612:
vmaf_hip_kernel_submit_post_recordinkernel_template.{h,c}and thehip_hsaco_sourcesmeson pipeline are also being added by PR #612 (float_psnr_hip). When the two PRs merge, keep one copy of each and discard the duplicate. The bodies are identical so either direction is safe. - On upstream sync: no action required for PSNR or ANSNR logic (fork-local kernels). If upstream adds its own HIP backend with conflicting
meson.buildvariables, resolve manually against thehip_hsaco_sourcespattern documented in this PR. - Re-test on rebase:
# CPU-only build (no ROCm required): must compile clean
meson setup build_hip_cpu -Denable_hip=true -Denable_hipcc=false \
-Denable_cuda=false -Denable_sycl=false libvmaf
ninja -C build_hip_cpu
# With ROCm + hipcc: HSACO pipeline must produce .hsaco + _hsaco.c
meson setup build_hip_full -Denable_hip=true -Denable_hipcc=true \
-Denable_cuda=false -Denable_sycl=false libvmaf
ninja -C build_hip_full
# Cross-backend numeric gate (requires AMD GPU)
meson test -C build_hip_full --suite=hip-parity
ADR-0373 — HIP batch-2: float_motion_hip real kernel (2026-05-10)¶
- Touches:
core/src/feature/hip/float_motion_hip.c(full rewrite to#ifdef HAVE_HIPCCdual-path;uintptr_topaque slots replaced with realvoid *device pointers),core/src/feature/hip/float_motion/float_motion_score.hip(device kernel — already present, no change in this PR),core/src/meson.build(float_motion_scoreadded tohip_kernel_sources),docs/backends/hip/overview.md(status update to 4/11),docs/adr/0373-hip-batch2-float-motion.md(new). - Invariant (HAVE_HIPCC dual-path):
hipModule_t module,hipFunction_t funcbpc8/funcbpc16, andvoid *ref_in,void *blur[2]live under#ifdef HAVE_HIPCC. Without the flag the scaffold-ENOSYScontract is preserved. Never move these fields outside the guard. - Invariant (temporal blur ping-pong):
cur_bluralternates 0/1 in bothsubmit()andcollect(). The kernel readsblur[1 - s->cur_blur]as "prev" and writesblur[s->cur_blur]as "cur".collect()flipscur_blurafter consuming the partials. On rebase: if the CUDA twin changes the ping-pong direction, the HIP twin must follow in the same PR. - Invariant (first-frame compute_sad=0):
submit()passescompute_sad=0whenindex == 0. The kernel still writescur_blurbut sets all partials to 0.0.collect()emits motion_score=0 and motion2_score=0 forindex==0without SAD accumulation. - Invariant (flush tail):
flush()emitsVMAF_feature_motion2_score = prev_motion_scoreats->index(the last frame) and returns 1. Ifs->index == 0it returns 1 immediately. Mirrorsflush_fex_cudashape exactly. - On upstream sync: no action required (fork-local kernel). If Netflix adds a HIP backend with conflicting
float_motionlogic, resolve against this invariant set. - Re-test on rebase:
# CPU-only build (no ROCm required): must compile clean
meson setup build_hip_cpu -Denable_hip=true -Denable_hipcc=false \
-Denable_cuda=false -Denable_sycl=false libvmaf
ninja -C build_hip_cpu
# With ROCm + hipcc: HSACO pipeline must produce float_motion_score.hsaco
meson setup build_hip_full -Denable_hip=true -Denable_hipcc=true \
-Denable_cuda=false -Denable_sycl=false libvmaf
ninja -C build_hip_full
# Cross-backend numeric gate (requires AMD GPU)
meson test -C build_hip_full --suite=hip-parity
ADR-0375 — HIP batch-3: float_moment_hip + float_ssim_hip real kernels (2026-05-10)¶
- Touches:
core/src/feature/hip/float_moment_hip.c(full rewrite to#ifdef HAVE_HIPCCdual-path),core/src/feature/hip/float_moment_hip.h(HSACO symbol decl underHAVE_HIPCC),core/src/feature/hip/float_moment/moment_score.hip(new device kernel),core/src/feature/hip/float_ssim_hip.c(full rewrite to#ifdef HAVE_HIPCCdual-path),core/src/feature/hip/float_ssim_hip.h(HSACO symbol decl underHAVE_HIPCC),core/src/feature/hip/float_ssim/ssim_score.hip(new device kernel),core/src/meson.build(moment_scoreandssim_scoreadded tohip_kernel_sources),docs/backends/hip/overview.md(status update to 6/11),docs/adr/0375-hip-batch3-float-moment-float-ssim.md(new). - Invariant (HAVE_HIPCC dual-path):
hipModule_t module,hipFunction_t funcbpc8/16,void *ref_in/dis_in(moment) and all fivevoid *d_*intermediate buffers +void *ref_in/cmp_in(SSIM) live under#ifdef HAVE_HIPCCin the respective structs. Free helpers (moment_hip_module_free,ssim_hip_bufs_free) are defined outside the guard with internal#ifdef HAVE_HIPCCbodies (mirrorsfloat_psnr_hip_module_free). On rebase: never move device-state fields outside the guard — it breaks the CPU-only build. - Invariant (moment 7-arg kernel):
calculate_moment_hip_kernel_8bpcand_16bpcboth take 7 arguments (ref, dis, ref_stride, dis_stride, sums, width, height) — the 16bpc kernel does NOT take abpcarg. The host launch must pass 7 args to both functions. If the CUDA twin adds abpcarg tomoment_score_16bpc, the HIP twin must follow. - Invariant (SSIM two-pass stream ordering): both
calculate_ssim_hip_horiz_*andcalculate_ssim_hip_vert_combinerun on the sames->lc.strstream. Implicit stream ordering provides the happens-before between Pass 1 writes and Pass 2 reads — no explicit event is needed between the two launches. On rebase: if the CUDA twin adds an explicit inter-pass sync event, evaluate whether GCN/RDNA's stream ordering guarantees are equivalent before mirroring. - Invariant (SSIM WARPS_PER_BLOCK=2):
SSIM_WARPS_PER_BLOCK = SSIM_BLOCK_SIZE / SSIM_WARP_SIZE = 128 / 64 = 2(vs CUDA's 4 = 128/32). The shared-memory arrays_warp_sums[SSIM_WARPS_PER_BLOCK]must be sized 2. On rebase: if the block size or warp size changes, updatessim_score.hipaccordingly. - Invariant (SSIM scale=1 only):
init_fex_hiprejectsscale != 1with-EINVAL. Mirrors the CUDA and Vulkan SSIM twins. Lifting this constraint is a future batch item and requires a new ADR. - On upstream sync: no action required (fork-local kernels).
- Re-test on rebase:
# CPU-only build (no ROCm required): must compile clean
meson setup build_hip_cpu -Denable_hip=true -Denable_hipcc=false \
-Denable_cuda=false -Denable_sycl=false libvmaf
ninja -C build_hip_cpu
# With ROCm + hipcc: HSACO pipeline must produce moment_score.hsaco + ssim_score.hsaco
meson setup build_hip_full -Denable_hip=true -Denable_hipcc=true \
-Denable_cuda=false -Denable_sycl=false libvmaf
ninja -C build_hip_full
# Cross-backend numeric gate (requires AMD GPU)
meson test -C build_hip_full --suite=hip-parity
feat/hip-float-psnr-first-real — T7-10b: float_psnr_hip first real kernel (ADR-0254)¶
Touches:
core/src/feature/hip/float_psnr_hip.c— complete rewrite from scaffold stub to functional kernel consumer (HIP Module API pattern:hipModuleLoadData+hipModuleLaunchKernel).core/src/feature/hip/float_psnr_hip.h— HSACO symbol extern declarations (float_psnr_score_hsaco[],float_psnr_score_hsaco_len).core/src/feature/hip/float_psnr/float_psnr_score.hip— new HIP device kernel file (warp-64 reduction, GCN/RDNA-specific__shfl_down).core/src/hip/kernel_template.{h,c}— newvmaf_hip_kernel_submit_post_recordhelper for the post-launch event.core/src/hip/meson.build—float_psnr_hip.cadded tohip_sources.core/src/meson.build—hip_hsaco_sourceslist +enable_hipccguard +hipcc/xxdcustom-target pipeline.core/meson_options.txt— newenable_hipccboolean option.core/src/feature/feature_extractor.c—vmaf_fex_float_psnr_hipregistration under#if HAVE_HIP.core/test/test_hip_smoke.c—test_float_psnr_hip_extractor_registered.
Invariants:
-
vmaf_hip_kernel_submit_post_recordcall ordering — must be called afterhipMemcpyAsyncDtoH and before the collect-sidevmaf_hip_kernel_collect_wait. Any reorder breaks the event-fencing contract (the finished event records the end of the readback copy, not the end of the kernel launch). If a future PR refactorskernel_template.c, preserve this ordering constraint. -
#ifdef HAVE_HIPCCwraps all device-dependent state —hipModule_t,hipFunction_t, staging-buffer allocs, and the kernel launch chain are all inside#ifdef HAVE_HIPCCguards so the file compiles cleanly without a ROCm SDK. The#ifndef HAVE_HIPCCpaths return-ENOSYS. Any future real-kernel port must preserve this dual-path pattern. -
Warp size 64 (GCN/RDNA) —
FPSNR_WARPS_PER_BLOCK = 4for block size 256 (vs CUDA's 8 warps at warp size 32). The.hipkernel uses__shfl_down(v, off)without a mask (HIP convention; CUDA uses__shfl_down_sync(0xffffffff, v, off)). Any future port of the CUDA twin's warp-reduction that changes these constants must update both backends. -
enable_hipcc=falsedefault — the meson option defaults tofalseso CI / downstream builds without a ROCm toolchain still compile cleanly. Theenable_hipcc=truepath requireshipcc+xxdinPATHand a ROCm 6+ SDK. Gate any build-system change on both codepaths.
Re-test on rebase:
# CPU-only (no ROCm needed):
meson setup build -Denable_hip=true -Denable_cuda=false -Denable_sycl=false libvmaf
ninja -C build
meson test -C build --suite=fast # test_float_psnr_hip_extractor_registered passes
# With hipcc (ROCm 6+):
meson setup build -Denable_hip=true -Denable_hipcc=true -Denable_cuda=false libvmaf
ninja -C build
# Confirm kernel launches on device by running vmaf with --feature float_psnr_hip
HIP batch-4 -- ciede_hip and integer_motion_v2_hip real kernels (ADR-0377)¶
Rebase-sensitive invariants:
-
Arithmetic shift on int32/int64 in
motion_v2_score.hip— the inner filter right-shifts (>> shift_y,>> shift_x) operate on signed types (int32_t,int64_t). They MUST remain arithmetic (signed) shifts. Converting to logical shifts (e.g.,>> (unsigned), or using bitwise ops) diverges from the CPU reference for negative values. This was the root cause of the AVX2srlv_epi64regression in PR #587. The CUDA twin documents the same constraint. -
Mirror padding diverges from
motion_hip—motion_v2_score.hipuses reflective mirror (2 * size - idx - 1) whilemotion_hip's kernel uses skip-boundary mirror (2 * size - idx - 2). Both match their respective CPU references. Do not unify them on rebase. -
Six YUV staging buffers for
ciede_hip—ciede_hip_bufs_allocallocates ref_y/u/v + dis_y/u/v separately. The chroma buffers are sized atchroma_w * chroma_h, notluma_w * luma_h. If the HIP picture-buffer API changes on rebase (e.g.,VMAF_FEATURE_EXTRACTOR_HIPflag lands and pictures arrive on-device), these staging copies and theirhipMalloc/hipFreecalls must be removed or made conditional. -
#ifdef HAVE_HIPCCdual-path preserved — same invariant as float_psnr_hip (see entry above). All device-dependent state and kernel launches are inside#ifdef HAVE_HIPCCguards.
Re-test on rebase:
# CPU-only (no ROCm needed):
meson setup build -Denable_hip=true -Denable_cuda=false -Denable_sycl=false libvmaf
ninja -C build
meson test -C build # 54/54 pass including test_hip_smoke
speed_qa -- real SpEED-QA implementation (ADR-0253)¶
core/src/feature/speed_qa.c went from a 71-line placeholder scaffold to a ~380-line real implementation. The extractor now sets
speed_qa — real SpEED-QA implementation (ADR-0253)¶
core/src/feature/speed_qa.c went from a 71-line placeholder scaffold to a ~380-line real implementation. The extractor now sets
VMAF_FEATURE_EXTRACTOR_TEMPORAL and carries priv_size = sizeof(SpeedQaState); the registration entry in feature_extractor_list[] is unchanged (always unconditional, outside the #if VMAF_FLOAT_FEATURES block).
No upstream rebase conflict expected. The scaffold was fork-local; Netflix
upstream has no speed_qa.c. The upstream speed.c is unmodified.
Rebase invariant: vmaf_fex_speed_qa must stay outside the VMAF_FLOAT_FEATURES guard in feature_extractor.c -- speed_qa.c is compiled unconditionally (no float dependency). If a future Netflix commit lands a speed_qa.c, audit for algorithm conflicts before merging.
upstream has no speed_qa.c. The speed.c file (upstream port) is unmodified.
Rebase invariant: vmaf_fex_speed_qa must stay outside the VMAF_FLOAT_FEATURES guard in feature_extractor.c — speed_qa.c is compiled unconditionally (no float dependency). If a rebase lands a Netflix speed_qa.c, audit for algorithm conflicts before merging.
Re-test on rebase:
meson setup build_test libvmaf -Denable_cuda=false -Denable_sycl=false
ninja -C build_test test/test_speed_qa
meson test -C build_test test_speed_qa --verbose
# Expected: 5 tests run, 5 passed
0378 — picture-upload stream CU_STREAM_NON_BLOCKING (PR #702, ADR-0378)¶
Touches: core/src/cuda/picture_cuda.c (one line in vmaf_cuda_picture_alloc).
Invariant: priv->cuda.str must be created with CU_STREAM_NON_BLOCKING (via cuStreamCreateWithPriority). The CUDA implicit null-stream serialisation rule makes CU_STREAM_DEFAULT a per-frame context barrier; at sub-4K this reduces CUDA motion throughput to ~0.55x CPU. If a future upstream commit touches vmaf_cuda_picture_alloc and reverts the stream flag, the performance regression returns silently.
Re-test:
meson setup build-cuda libvmaf -Denable_cuda=true -Denable_sycl=false
ninja -C build-cuda
./build-cuda/tools/vmaf_bench --resolution 576x324 --gpu-only --frames 20
# motion (CUDA) @ 576x324 must be >= 30 fps (>= 1x CPU baseline)
ADR-0376 — Vulkan buffer-invalidate void → int fix (GCC 16, 2026-05-10)¶
- Touches:
core/src/feature/vulkan/float_ansnr_vulkan.c(functionreduce_partials, call site inextract),core/src/feature/vulkan/cambi_vulkan.c(functionscambi_vk_readback_image,cambi_vk_readback_mask, call sites incambi_vk_extract). - Invariant:
reduce_partials,cambi_vk_readback_image, andcambi_vk_readback_maskare nowstatic intwith error-propagating call sites. If an upstream sync brings a competing refactor of these functions (e.g., a signature change or a different coherency-flush strategy), thestatic intcontract and call-site error checks must be preserved. - On upstream sync: no risk from Netflix/vmaf upstream — these files are 100% fork-local (Vulkan feature extractors do not exist upstream). The rebase risk is internal: if a fork-local PR changes the Vulkan buffer-API surface (e.g., a new
vmaf_vulkan_buffer_invalidatevariant), verify the call-site error propagation pattern still applies. - Re-test on rebase:
meson setup build-vk-retest -Denable_cuda=false -Denable_sycl=false -Denable_vulkan=true
ninja -C build-vk-retest
# Must compile cleanly under GCC 16 with no -Wreturn-mismatch error
PR-fix-cuda-picture-widening — CUDA picture_cuda.c integer-precision fixes (round-5 clang-tidy)¶
- Touches:
core/src/cuda/picture_cuda.c— upstream-shared CUDA picture allocation and transfer path. - Invariant: Three
WidthInBytes/cuMemAllocPitchwidth arguments now use(size_t)casts to prevent silent 32-bit multiplication overflow before widening tosize_t.aligned_y/aligned_careunsignedwith an explicit1umask literal.vmaf_ref_load()result is stored aslong. If an upstream sync modifies these expressions, ensure the(size_t)casts andunsignedtypes are preserved. - Re-test on rebase:
clang-tidy \
-checks='-*,bugprone-narrowing-conversions,bugprone-implicit-widening-of-multiplication-result' \
-p core/build-cuda \
core/src/cuda/picture_cuda.c
# Must produce zero warnings for the named checks.
PR-fix-cuda-dispatch-getenv — CUDA dispatch_strategy.c getenv() thread-safety fix¶
- Touches:
core/src/cuda/dispatch_strategy.c— fork-local TU (no upstream equivalent). - Invariant:
g_env_once/cache_env_dispatch/g_env_dispmust remain as the single canonical read path forVMAF_CUDA_DISPATCH. If a future PR needs to re-read the variable (e.g., for unit-test reset), it must resetg_env_onceviapthread_once_t g_env_once = PTHREAD_ONCE_INIT;in a test fixture, not callgetenv()directly fromvmaf_cuda_select_strategy. - Re-test on rebase:
clang-tidy \
-checks='-*,concurrency-*' \
-p core/build-cuda \
core/src/cuda/dispatch_strategy.c
# Must produce zero concurrency-mt-unsafe warnings.
0103 — -fvisibility=hidden + VMAF_EXPORT public-API annotation (ADR-0379, Research-0092)¶
- Touches:
core/src/meson.build(vmaf_cflags_common),core/include/libvmaf/*.h(all public headers),core/include/libvmaf/macros.h(new file),core/include/core/meson.build(header install list),core/src/dnn/model_loader.h(vmaf_dnn_verify_signaturedeclaration). - Invariant:
-fvisibility=hiddenis invmaf_cflags_common. Any new publicvmaf_*function added by an upstream sync — whether inlibvmaf.c,picture.c,dict.c, or any other source — must also haveVMAF_EXPORTon its declaration in the matching public header, otherwise it will be hidden inlibvmaf.soand downstream callers will get a link error. Gate:nm -D --defined-only build/src/libvmaf.so.* | grep ' [TW] ' | grep -v ' vmaf_' | wc -lmust be 0. - On upstream sync: upstream Netflix/vmaf does NOT use
-fvisibility=hidden. Any new public entry point in an upstream commit (typically added tocore/src/libvmaf.c+core/include/libvmaf/libvmaf.h) will compile to a hidden symbol on the fork withoutVMAF_EXPORT. The merge author must: - Add
VMAF_EXPORTto the new declaration in the public header. - Run the
nm -Dgate (above) — it must return 0. - Run
meson test -C build— all tests must pass. - Re-test on rebase:
meson setup build-vis libvmaf -Denable_cuda=false -Denable_sycl=false --wipe
ninja -C build-vis
nm -D --defined-only build-vis/src/libvmaf.so.* | grep ' [TW] ' | grep -v ' vmaf_' | wc -l
# Must print 0
meson test -C build-vis
# All tests must pass
PR-fix-cuda-switch-defaults — CUDA feature extractor defensive fixes (round-5 clang-tidy)¶
- Touches:
core/src/feature/cuda/integer_adm_cuda.c,core/src/feature/cuda/integer_vif_cuda.c. - Invariant:
default: break;clauses added to threeswitch(scale)statements. If an upstream sync adds new scale cases to ADM (scales 1–3) or VIF (scales 0–3), the default clause remains valid but no longer exhausts all cases — update the comment accordingly. TheRES_BUFFER_SIZEmacro now has parentheses; any fork-local addition to that macro must preserve them. - Re-test on rebase:
clang-tidy \
-checks='-*,bugprone-macro-parentheses,bugprone-switch-missing-default-case' \
-p core/build-cuda \
core/src/feature/cuda/integer_adm_cuda.c \
core/src/feature/cuda/integer_vif_cuda.c
# Must produce zero warnings for those checks.
PR-fix-picture-align-unsigned-narrowing — integer-sanitizer narrowing/overflow fixes in picture.c, libvmaf.c, tensor_io.c (round-5 -fsanitize=integer sweep)¶
- Touches:
core/src/picture.c(upstream-shared picture geometry),core/src/libvmaf.c(upstream-shared vmaf_init), andcore/src/dnn/tensor_io.c(fork-added f16 ↔ f32 converter). - Invariant: Three narrowing/overflow defects corrected: (1)
aligned_y/aligned_care nowunsignedwithDATA_ALIGN - 1umask — if an upstream sync touchespicture_compute_geometry, ensure theunsignedtype and1uliteral are preserved. (2)vmaf_set_cpu_flags_maskcall site invmaf_inituses(unsigned)(~cfg.cpumask)— if cpumask type changes upstream, revisit the cast. (3)f16_to_f32_onesubnormal path uses a signedint32_t exp_adjcounter — if the f16 converter is reworked, verify no unsigned wrap is reintroduced. - Re-test on rebase:
CC=clang CXX=clang++ meson setup /tmp/build-isan-retest libvmaf \
-Denable_cuda=false -Denable_sycl=false \
--buildtype=debugoptimized -Db_sanitize=integer -Db_lundef=false -Db_lto=false
ninja -C /tmp/build-isan-retest
UBSAN_OPTIONS="halt_on_error=0:abort_on_error=0" \
/tmp/build-isan-retest/test/test_picture 2>&1 | grep "runtime error"
UBSAN_OPTIONS="halt_on_error=0:abort_on_error=0" \
/tmp/build-isan-retest/test/test_read_pictures_monotonic 2>&1 | grep "runtime error"
UBSAN_OPTIONS="halt_on_error=0:abort_on_error=0" \
/tmp/build-isan-retest/test/dnn/test_tensor_io 2>&1 | grep "runtime error"
# All three must produce zero "runtime error" lines.
PR-fix-cuda-pinned-alloc-null-deref — CWE-476 null-deref in vmaf_cuda_picture_alloc_pinned (round-6 cross-PR audit)¶
- Touches:
core/src/cuda/picture_cuda.c— CUDA host TU; no upstream equivalent. - Invariant: The sequential check pattern (
err = vmaf_picture_priv_init(pic); if (err) goto free_data;) must be preserved on any rebase or future modification. The|=idiom evaluates the right-hand side unconditionally regardless of prior failure — PR #700 fixed the identical pattern inpicture.c(CWE-476); this fix closes the same class in the CUDA path. If upstream ever adds a similar pinned-picture allocation function, apply the same sequential-check discipline. Secondary:DATA_ALIGN_PINNED - 1u(with theusuffix) must be preserved on both sides of the alignment mask expression to match thepicture.cpattern fixed by PR #708. - Re-test on rebase:
gcc -fanalyzer -Wno-analyzer-too-complex \
-Ilibvmaf/src -Ilibvmaf/include \
core/src/cuda/picture_cuda.c 2>&1 | grep "CWE-476"
# Must produce zero CWE-476 warnings for vmaf_cuda_picture_alloc_pinned.
meson test -C build --suite=fast
# Must be green.
0380 — FFmpeg HIP backend selector patch (ADR-0380, ffmpeg-patches 0011)¶
- Touches:
ffmpeg-patches/0011-libvmaf-wire-hip-backend-selector.patch,ffmpeg-patches/series.txt,ffmpeg-patches/README.md. - Invariant: The patch is authored against FFmpeg
n8.1.1as the cumulative base (0001..0010 stack applied). Context lines in the patch referenceVmafCudaState *cu_state/cuda_pool_initialised(added by patch 0010) and#include <vulkan/vulkan.h>/#endif(added by patch 0006). If the series is ever rebased to a newer FFmpeg tag (n8.2,n9.x, etc.) the surrounding context invf_libvmaf.cmay shift; run the full series replay against the new tag and regenerate conflicting patches. The HIP cleanup path (vmaf_hip_state_free(&s->hip_state)) uses a double pointer, unlike the CUDA path (vmaf_cuda_state_free(s->cu_state)) which uses a single pointer — this asymmetry is intentional and matcheslibvmaf_hip.h; preserve it. - Error-code invariant (fix/hip-averror-propagation-0011, 2026-05-10): Both HIP error sites in
init()must usereturn AVERROR(-err), notreturn AVERROR(EINVAL).AVERROR(-err)maps the libvmaf-supplied errno (e.g.-ENODEV= -19,-ENOSYS= -38) to the correct FFmpeg error string ("No such device", "Function not implemented"). TheAVERROR(EINVAL)form was the original patch text; it was corrected during the full 0001–0011 e2e test run. If the patch context is regenerated or the hunk is split, verify the fix is preserved. - Re-test on rebase:
git -C /tmp/ffmpeg-8 reset --hard n8.1.1
for p in ffmpeg-patches/000*-*.patch; do
git -C /tmp/ffmpeg-8 am --3way "$p" || break
done
# All 11 patches must apply without conflict.
Research-0094 — integer_motion_v2 flush() dict-leak fix¶
- Touches:
core/src/feature/integer_motion_v2.c. - Invariant: The
dict_locally_ownedflag inflush()(introduced in this fix) relies on the invariant thats->feature_name_dictisNULLatflush()entry only in the registered-context (threaded dispatch) path, and non-NULL only whenextract()has already run on this context (serial / pool-instance path). If a future upstream change causesextract()to clears->feature_name_dictmid-run (e.g., per-scene re-init), the flag will incorrectly take the locally-owned path and free the dict prematurely. The companion unit test (test in test_feature_extractor) guards this via the existing motion_v2 code path. - Re-test on rebase:
meson test -C build --suite=fast
# 53/53 must pass, including test_feature_extractor and test_motion_v2_simd.
ASAN_OPTIONS='detect_leaks=1' ./build-leak/tools/vmaf \
-r python/test/resource/yuv/src01_hrc00_576x324.yuv \
-d python/test/resource/yuv/src01_hrc01_576x324.yuv \
-w 576 -h 324 -p 420 -b 8 --feature motion_v2 \
--output /dev/null --threads 4 2>&1 | grep -E 'leak|SUMMARY'
# Must produce no output (clean).
0382 — Y4M negative-dimension rejection (ADR-0382, T-FUZZ-Y4M-NEG-WIDTH-SEGV)¶
- Touches:
core/tools/y4m_input.c(internal staticy4m_input_open_impl),core/test/fuzz/y4m_input_known_crashes/(new corpus seed). - Invariant: The guard
if (_y4m->pic_w <= 0 || _y4m->pic_h <= 0)must stay between they4m_parse_tags()call and the chroma-type dispatch block. If upstream restructuresy4m_input_open_implor moves the tag parser, the guard must migrate with it so no allocation occurs before the check. They4m_neg_width_null_deref.y4mseed must be replayed on every rebase to confirm the parser returns clean-1rather than SEGV. - No rebase impact on public API or ffmpeg-patches: the fix is internal to
y4m_input_open_impl(astaticfunction); no public header is changed; no ffmpeg-patches patch is affected. - Re-test on rebase:
CC=clang meson setup build-fuzz libvmaf \
-Dfuzz=true -Db_sanitize=address \
-Denable_cuda=false -Denable_sycl=false \
-Denable_vulkan=disabled --buildtype=debug
ninja -C build-fuzz core/test/fuzz/fuzz_y4m_input
./build-fuzz/core/test/fuzz/fuzz_y4m_input \
core/test/fuzz/y4m_input_known_crashes/y4m_neg_width_null_deref.y4m
# Pre-fix: SEGV on address 0x000000000000 inside fread.
# Post-fix: exits 0; stderr prints
# "Invalid YUV4MPEG2 dimensions: W=-8 H=4 (must be > 0)."
fix/picture-odd-dim-chroma-ceiling — picture_compute_geometry ceiling division for odd luma dims¶
- Touches:
core/src/picture.c,core/src/cuda/picture_cuda.c,core/src/feature/cuda/integer_psnr_cuda.c,core/src/feature/cuda/integer_psnr_hvs_cuda.c,core/src/feature/integer_psnr.c,core/test/test_picture.c. - Invariant: All geometry computations for chroma plane dimensions use ceiling division
(dim + ss) >> ss(wheressis 0 or 1). If upstream adds a new allocator or copies the geometry pattern, it must use the same ceiling form. The regression testtest_picture_odd_dim_chroma_ceilingpins this: 577 × 323 YUV420 must producepic.w[1]==289,pic.h[1]==162. - Re-test on rebase:
meson test -C build --suite=fast
# test_picture must pass; it includes test_picture_odd_dim_chroma_ceiling.
# Additionally, the ASan smoke:
python3 -c "
W, H = 577, 323
luma = bytes([128] * W * H)
cw, ch = (W+1)>>1, (H+1)>>1
chroma = bytes([128] * cw * ch)
open('/tmp/odd.yuv','wb').write((luma+chroma+chroma)*3)
"
ASAN_OPTIONS=halt_on_error=1 ./build/tools/vmaf \
--reference /tmp/odd.yuv --distorted /tmp/odd.yuv \
--width 577 --height 323 --pixel_format 420 --bitdepth 8 \
--feature ciede --threads 4
# Must exit 0 with no ASan reports.
fix/motion-mirror-padding-min-dim — 5-tap filter minimum-dimension guard in all motion extractors¶
- Touches:
core/src/feature/integer_motion.c,core/src/feature/integer_motion_v2.c,core/src/feature/float_motion.c,core/src/feature/cuda/integer_motion_cuda.c,core/src/feature/cuda/float_motion_cuda.c,core/src/feature/cuda/integer_motion_v2_cuda.c,core/src/feature/sycl/integer_motion_sycl.cpp,core/src/feature/sycl/float_motion_sycl.cpp,core/src/feature/sycl/integer_motion_v2_sycl.cpp,core/src/feature/vulkan/motion_vulkan.c,core/src/feature/vulkan/motion_v2_vulkan.c,core/src/feature/vulkan/float_motion_vulkan.c,core/src/feature/hip/integer_motion_v2_hip.c,core/src/feature/hip/float_motion_hip.c,core/test/test_motion_min_dim.c,core/test/meson.build. - Invariant: Every motion
init()rejectsw < 3 || h < 3with-EINVALbefore any buffer allocation. The reflect-101 mirror formulaheight - (i_tap - height + 2)requiresheight ≥ filter_width/2 + 1 = 3. If upstream Netflix/vmaf modifies the convolution core to support smaller frames (e.g. by switching to a clamp-to-edge formula), the guard should be re-evaluated. If upstream adds a new motion extractor that also uses the 5-tap kernel, add the same guard to itsinit(). - Re-test on rebase:
meson test -C build --suite=fast
# test_motion_min_dim must pass (13/13 cases).
# Reproducer:
python3 -c "
plane = bytes([128]*1*1)
chroma = bytes([128]*1*1)
frame = plane + chroma + chroma
with open('/tmp/1x1.yuv','wb') as f: f.write(frame*3)
"
./build/tools/vmaf --reference /tmp/1x1.yuv --distorted /tmp/1x1.yuv \
--width 1 --height 1 --pixel_format 420 --bitdepth 8 \
--feature motion --threads 1 2>&1 | grep -E 'EINVAL|minimum|below'
# Must print the "frame 1x1 is below the 5-tap filter minimum" message.
0381 — Vulkan VIF scale 2/3 numerical saturation fix (ADR-0381, PR #718)¶
- Touches:
core/src/feature/vulkan/shaders/float_vif.comp,core/src/feature/vulkan/vif_vulkan.c,core/src/vulkan/meson.build. - Invariant 1 —
float_vif.compmust remain inpsnr_hvs_strict_shaders.meson.build'spsnr_hvs_strict_shaderslist controls which shaders compile withglslc -O0(strict) vs-O(optimised).float_vif.compbelongs to this list because the SPIR-V optimizer's FMA-contraction and reassociation ofsigma1_sq = xx - mu1*mu1triggers catastrophic cancellation at scales 2 and 3 (small local variance), saturating the per-scale score to 1.0 and inflating VMAF by ~+1.07. If the list is re-ordered orfloat_vif.compis accidentally removed, restore it. - Invariant 2 —
precisequalifiers onfloat_vif.compaccumulators. The vertical-pass accumulators (a_mu1,a_mu2,a_xx,a_yy,a_xy), the horizontal-pass accumulators (mu1,mu2,xx,yy,xy), and the sigma expressions (sigma1_sq,sigma2_sq,sigma12) carryprecisequalifiers. These map toOpDecorate NoContractionin SPIR-V and defend against driver-side FMA contraction (Vulkan 1.4 NVIDIA / newer MoltenVK). Do not remove theprecisequalifiers without re-runningplaces=4on every CI hardware lane. - Invariant 3 — integer VIF rd buffer ceiling division.
vif_vulkan.c::alloc_buffers()allocates the per-scale rd buffers with ceiling division:((w + 1u) / 2u) * ((h + 1u) / 2u). For odd input dimensions (e.g. h=81 at scale 2 of a 576×324 input), the shader writesrd_yindices up toh/2 = 40(inclusive), requiring 41 rows. Floor divisionh/2 = 40under-allocates by one row (72 uint32 slots), corrupting the adjacent per-WG int64 accumulator buffer. The ceiling form must be preserved on any refactor or upstream merge touchingalloc_buffers. - Re-test on rebase:
# Build with Vulkan enabled
meson setup build-vk-vif libvmaf -Denable_vulkan=true -Denable_cuda=false -Denable_sycl=false --buildtype=release
ninja -C build-vk-vif
# Run per-scale VIF parity check on the Netflix 576x324 golden pair
build-vk-vif/tools/vmaf \
-r python/test/resource/yuv/src01_hrc00_576x324.yuv \
-d python/test/resource/yuv/src01_hrc01_576x324.yuv \
-w 576 -h 324 -p 420 -b 8 --feature float_vif_vulkan --backend vulkan \
--output /tmp/vk_vif_out.json
# All per-scale VIF scores must be < 1.0; VMAF must be within ±0.5 of CPU.
# Per-scale delta must be < 1e-3 from CPU reference.
fix/recal-adm-f1f2-post-pr731 — Recalibrate fork-local adm_f1f2 assertion after PR #731 AIM port¶
No rebase impact: this change touches only python/test/feature_extractor_test.py (a single assertAlmostEqual value for the fork-local adm_f1s/f2s feature), with a documentation comment explaining the recalibration. No C sources, no public headers, no Meson options, no FFmpeg patch stack entries were modified. If upstream Netflix/vmaf adds its own adm_f1s/f2s noise-weight test in a future sync, verify that the expected value (0.8872294166666667) still matches the post-PR-#731 CPU scalar path output for the src01_hrc00_576x324.yuv ↔ src01_hrc01_576x324.yuv pair with the f1s/f2s parameters listed in test_run_vmaf_fextractor_adm_f1f2.
- Re-test:
PYTHONPATH=$PWD/python python3 -m pytest python/test/feature_extractor_test.py::FeatureExtractorTest::test_run_vmaf_fextractor_adm_f1f2 -v— must report 1 passed.
feat/adm-gpu-param-sync — ADM noise_weight/csf_scale/csf_diag_scale GPU extension¶
- Touches:
core/src/feature/cuda/float_adm_cuda.c,core/src/feature/cuda/integer_adm_cuda.c,core/src/feature/sycl/float_adm_sycl.cpp,core/src/feature/sycl/integer_adm_sycl.cpp,core/src/feature/vulkan/adm_vulkan.c,core/src/feature/vulkan/float_adm_vulkan.c. - Invariant 1 — three-param parity with CPU. Every GPU ADM backend (
float_adm_cuda,integer_adm_cuda,float_adm_sycl,integer_adm_sycl,adm_vulkan,float_adm_vulkan) exposesadm_csf_scale,adm_csf_diag_scale, andnoise_weightwith the same defaults (1.0,1.0,0.03125) as the CPU scalar path added by PR #731. If upstream Netflix ever adds or renames these parameters ininteger_adm.c/float_adm.c, the corresponding GPU files must be updated in the same PR. - Invariant 2 — integer CUDA must NOT include
adm_options.hdirectly.core/src/feature/cuda/integer_adm_cuda.cmust NOT includefeature/adm_options.hdirectly.DEFAULT_ADM_NOISE_WEIGHT,DEFAULT_ADM_CSF_SCALE,DEFAULT_ADM_CSF_DIAG_SCALE, and the full 4-memberenum ADM_CSF_MODEarrive transitively viacuda/integer_adm_cuda.h→feature/integer_adm.h. A direct include reintroduces the 2-memberenum ADM_CSF_MODEfromadm_options.hand produces a redeclaration error. - Invariant 3 — Vulkan integer fast-path gated on CSF-scale defaults.
adm_vulkan.ccontains a hard-codedi_rfactorfast-path for the3.0 * 1080default viewing geometry. It is gated by:bool csf_default = (fabs(s->adm_csf_scale - 1.0) < 1e-9) && (fabs(s->adm_csf_diag_scale - 1.0) < 1e-9). If the fast-path is ever updated, the CSF-default guard must be updated to match; removing or loosening the guard will produce wrong rfactors when non-default CSF scales are in use. - Re-test on rebase:
# CPU-only build + golden test
meson setup build-cpu libvmaf -Denable_cuda=false -Denable_sycl=false \
-Denable_vulkan=disabled
ninja -C build-cpu
make test-netflix-golden
# Verify default params produce unchanged scores
build-cpu/tools/vmaf \
-r python/test/resource/yuv/src01_hrc00_576x324.yuv \
-d python/test/resource/yuv/src01_hrc01_576x324.yuv \
-w 576 -h 324 -p 420 -b 8 \
--feature adm=noise_weight=0.03125:adm_csf_scale=1.0:adm_csf_diag_scale=1.0 \
--output /tmp/adm_param_default.json
# adm2 must match the no-param baseline (places=4).
0383 — K150K parallel CPU driver + feature_extractor_list dedup fix (ADR-0383)¶
- Touches:
ai/scripts/extract_k150k_features.py— driver redesign.ai/AGENTS.md— K150K-A invariant note updated.core/src/feature/feature_extractor.c— duplicate CUDA extractor registration removed (lines 239–240 deduplicated: six CUDA extractors that were registered twice).docs/ai/datasets/k150k.md— user-facing docs updated.docs/adr/0383-k150k-parallel-cpu-driver.md— new ADR.docs/research/0096-k150k-gpu-driver-investigation-2026-05-10.md— new digest.- Invariant 1 —
feature_extractor_list[]must have no duplicate entries. The dedup infeature_extractor_vector_append()is by extractor name, not by provided-feature name. Duplicate entries infeature_extractor_list[]result in both extractors being registered and both writing the same feature-collector slots. If upstream Netflix/vmaf modifiescore/src/feature/feature_extractor.cto add new backend entries, verify that no extractor is registered more than once. - Invariant 2 — CUDA binary double-write via default model auto-load. When
--modelis not specified and--no-predictionis absent, the CLI auto-loadsvmaf_v0.6.1, which registers CUDA twins viavmaf_use_features_from_model(). A subsequent--feature admcall registers the CPU "adm" extractor in addition; both run and double-write. This is a latent bug in the CLI model-auto-load / explicit-feature interaction path. The K150K pipeline works around it by using the CPU binary (no CUDA context). See Research-0096 for full root-cause analysis. - Re-test on rebase:
# Verify no duplicate entries exist in feature_extractor_list[]
grep -c "vmaf_fex_integer_adm_cuda" core/src/feature/feature_extractor.c
# Must print 2 (one extern declaration + one list entry)
# Verify CPU driver produces correct output
python ai/scripts/extract_k150k_features.py --limit 5 --threads-cuda 2 \
--out /tmp/smoke5.parquet && python3 -c \
"import pandas; df=pandas.read_parquet('/tmp/smoke5.parquet'); \
print(df.shape, df.columns.tolist()[:5])"
# Must print (5, 48) and the first five column names.
fix/ci-master-shfmt-cppcheck-semgrep — CI gate fixes (ADR-0384)¶
- Touches:
.pre-commit-config.yaml,.github/workflows/lint-and-format.yml,core/src/feature/adm.c,core/src/feature/ansnr.c,core/src/feature/offset.c,core/src/feature/vif.c,scripts/ci/check-agent-worktree-drift.sh. - Invariant: No rebase-sensitive invariants. The
.pre-commit-config.yamland workflow changes are fork-infrastructure. The(void *)cast fix in the four feature files is a portable C idiom; upstream may or may not carry their own version of these files. The semgrep-comment reword incheck-agent-worktree-drift.shis fork-only. - Re-test on rebase:
# Verify semgrep is clean
semgrep scan --config=.semgrep.yml --error
# Verify cppcheck finds no invalidPointerCast in the four files
cppcheck --enable=portability core/src/feature/adm.c \
core/src/feature/ansnr.c core/src/feature/offset.c \
core/src/feature/vif.c 2>&1 | grep invalidPointerCast
# Must produce no output.
no rebase impact: the CI infrastructure files are fork-local; the C source changes are minimal (cast through void*) and will trivially survive any upstream rebase that doesn't rewrite these specific functions.
fix/thread-pool-pthread-create-unchecked — thread pool pthread_create error handling + n_workers_created race fix¶
- Touches:
core/src/thread_pool.c. - Invariant:
VmafThreadPoolnow has an_workers_createdfield (written once at creation, never decremented) alongside the existingn_threadscounter (decremented by each exiting worker). Any upstream change tothread_pool.cthat adds or renames struct fields or changes thepthread_createcall site must be reconciled against the fork's error-handling block (lines ~170–192) and then_workers_createdfield initialisation. - Re-test on rebase:
meson setup /tmp/build-tp-rebase libvmaf \
-Denable_cuda=false -Denable_sycl=false --buildtype=debugoptimized
meson test -C /tmp/build-tp-rebase
# Must report 54/54 (or more) OK.
ai/tiny-netflix-training-scaffold — tiny-AI Netflix corpus training scaffold draft PR (ADR-0417)¶
- Touches:
docs/adr/0417-tiny-ai-netflix-training-scaffold-pr.md,docs/research/0099-tiny-ai-netflix-training-update.md,docs/adr/_index_fragments/0417-tiny-ai-netflix-training-scaffold-pr.md,changelog.d/added/0417-tiny-ai-netflix-training-scaffold-pr.md,docs/ai/training-data.md(See-also links only). - Invariant: the corpus path
.workingdir2/netflix/is gitignored and must never be committed. The--data-rootCLI flag andVMAF_DATA_ROOTenvironment variable are the only sanctioned ways to point the training scripts at the corpus. Any rebase or upstream sync that modifiesai/ormcp-server/vmaf-mcp/must preserve this invariant; verify withgit check-ignore -v .workingdir2/netflix/ref/(must return the root.gitignoreentry). - Re-test on rebase:
cd mcp-server/vmaf-mcp && python -m pytest tests/test_smoke_e2e.py -v
# Requires: meson compile -C build (for the vmaf binary).
# test_list_tools_returns_expected_names PASSED
# test_list_tools_each_has_input_schema PASSED
# test_call_tool_list_models_returns_list PASSED
# test_call_tool_list_backends_includes_cpu PASSED
# test_call_tool_unknown_name_returns_error_json PASSED
# test_call_tool_vmaf_score_golden_pair PASSED (requires build/tools/vmaf)
no rebase impact on libvmaf C sources: this branch is doc-only (ADR-0417, Research Digest 0099, changelog fragment, ADR index fragment). The MCP smoke test and training-data.md are already in master and untouched by this branch.
0420 — Metal (Apple Silicon) backend runtime (T8-1b / ADR-0420)¶
- Touches:
core/src/metal/common.mm(new, replacescommon.c) — MTLDevice + MTLCommandQueue lifecycle;MTLCreateSystemDefaultDevicefor auto-pick;MTLCopyAllDevicesfor explicit indexing; Apple-Family-7 gate.core/src/metal/picture_metal.mm(new, replacespicture_metal.c) — MTLBuffer allocator withMTLResourceStorageModeShared(zero-copy).core/src/metal/kernel_template.mm(new, replaceskernel_template.c) — private MTLCommandQueue + two MTLSharedEvent handles; blit-fill accumulator zero; cross-queueencodeWaitForEvent;waitUntilCompleteddrain.core/src/metal/common.h— two new internal accessors:vmaf_metal_context_device_handle()+vmaf_metal_context_queue_handle().core/src/metal/meson.build—dependency('Foundation'/'Metal', required: true);-fobjc-arcproject arg forobjcpp; source list flipped from.cto.mm.core/test/test_metal_smoke.c— smoke expectations flipped from-ENOSYSpin to runtime (0on Apple7+,-ENODEVelsewhere).docs/adr/0420-metal-backend-runtime-t8-1b.md+docs/adr/_index_fragments/0420-metal-backend-runtime-t8-1b.md+changelog.d/changed/metal-backend-runtime.md.- Upstream-port footprint: zero — Netflix/vmaf has no Metal backend.
- Rebase invariants:
- Header purity: no
<Metal/Metal.h>in any header or pure-C consumer. Metal handles cross the boundary asvoid */uintptr_t. Do not promote a Metal type into a header on rebase. - ARC bridge-cast discipline:
__bridge_retainedto stash (+1 retain),__bridge_transferto release (−1),__bridgeto borrow (no refcount). A missing_retainedleaks; a missing_transferdouble-frees. - Struct privacy:
struct VmafMetalContextis defined only incommon.mm. Consumers use the accessor pair — never struct-layout introspection. - HIP twin parity for kernel_template: any PR that grows the HIP
kernel_template.clifecycle must propagate the same change tokernel_template.mmin the same PR. - Re-test on rebase (macOS, Apple-Family-7+):
meson setup build libvmaf -Denable_metal=enabled \
-Denable_cuda=false -Denable_sycl=false
ninja -C build
meson test -C build test_metal_smoke # must PASS
On Linux: meson setup build libvmaf -Denable_cuda=false -Denable_sycl=false && ninja -C build (Metal subdir not entered; no Metal test registered).
ADR-0422 — CLI HIP and Metal backend selectors (2026-05-11)¶
Files touched: core/tools/cli_parse.h, core/tools/cli_parse.c, core/tools/vmaf.c, core/test/test_cli_parse.c, core/include/libvmaf/libvmaf_metal.h.
Rebase impact: none — this PR only adds new CLI flags and branches to the standalone vmaf tool. No libvmaf public C API symbols changed; no meson_options.txt entries added; the ffmpeg-patches stack is unaffected. The libvmaf_metal.h change is docstring-only (no API surface delta).
Invariants to preserve on rebase:
CLISettingsincli_parse.hhasno_hip,hip_device,no_metal,metal_devicefields. If upstream adds its own HIP/Metal CLI flags in the same struct, resolve the merge by keeping the fork's field names (they match our header convention) and dropping any upstream stub.--backend cpudisables all five GPU backends (no_cuda,no_sycl,no_vulkan,no_hip,no_metal). If upstream extends the backend enum, ensure the cpu branch stays exhaustive.init_gpu_backends()signature invmaf.cpasseship_state/hip_activeandmetal_state/metal_activeby reference under#ifdef HAVE_HIP/#ifdef HAVE_METALguards. Preserve both the guards and the by-reference convention on rebase.
Smoke-test after rebase:
meson setup build libvmaf -Denable_cuda=false -Denable_sycl=false --buildtype=debug
ninja -C build
meson test -C build test_cli_parse # all 5 new tests must pass
2026-05-14 — vmaf-tune recommend --from-corpus Row Filtering¶
Files touched: tools/vmaf-tune/src/vmaftune/cli.py, tools/vmaf-tune/tests/test_recommend.py, docs/usage/vmaf-tune.md, docs/state.md.
Rebase impact: low. This only aligns the CLI --from-corpus path with the existing vmaftune.recommend.recommend() filtering contract. No corpus schema, encode path, score path, or model artefact changes.
Invariant to preserve on rebase: both CLI and library corpus recommendation must filter through RecommendRequest / recommend() so failed rows, NaN rows, and non-matching encoder / preset rows cannot win from the CLI.
Smoke-test after rebase:
PYTHONPATH=tools/vmaf-tune/src .venv/bin/python -m pytest \
tools/vmaf-tune/tests/test_recommend.py -q
2026-05-14 — vmaf-tune Usage-Doc Scaffold Label Cleanup¶
Files touched: docs/usage/vmaf-tune.md, docs/usage/vmaf-tune-coarse-to-fine.md, docs/usage/vmaf-tune-bitrate-ladder.md, docs/usage/vmaf-tune-ladder-default-sampler.md, docs/usage/vmaf-tune-saliency-aware.md.
Rebase impact: documentation-only. The implementation already lives in tools/vmaf-tune/src/vmaftune/corpus.py, ladder.py, saliency.py, and fast.py; this PR removes stale user-facing stub labels that survived after those surfaces were wired.
Invariant to preserve on rebase: usage docs are implementation-status contracts, not backlog labels. If a command is wired and tested, do not call it a stub/scaffold in docs/usage/; describe the shipped path and name any remaining production limit precisely.
Smoke-test after rebase:
rg -n 'scaffold-only|Status: scaffold only|\(stub\)|\*\*Stub\*\*|recommend --saliency-aware|advisory in scaffold' \
docs/usage/vmaf-tune.md \
docs/usage/vmaf-tune-coarse-to-fine.md \
docs/usage/vmaf-tune-bitrate-ladder.md \
docs/usage/vmaf-tune-ladder-default-sampler.md \
docs/usage/vmaf-tune-saliency-aware.md
2026-05-14 — vmaf-tune fast --time-budget-s Timeout Wiring¶
Files touched: tools/vmaf-tune/src/vmaftune/fast.py, tools/vmaf-tune/src/vmaftune/cli.py, tools/vmaf-tune/tests/test_fast.py, docs/usage/vmaf-tune.md, docs/usage/vmaf-tune-fast-path.md.
Rebase impact: low. This is a fast-path user-surface fix only; no libvmaf public C API, model schema, or FFmpeg patch stack is touched.
Invariant to preserve on rebase: time_budget_s is a soft Optuna timeout. Do not revert it to metadata-only. The JSON n_trials field reports completed trials because it may be lower than the requested --n-trials when the timeout fires.
Smoke-test after rebase:
PYTHONPATH=tools/vmaf-tune/src .venv/bin/python -m pytest \
tools/vmaf-tune/tests/test_fast.py \
tools/vmaf-tune/tests/test_cli_fast.py -q
2026-05-14 — vmaf-tune Public Doc Stub-Label Sweep¶
Files touched: docs/usage/vmaf-tune-resolution-aware.md, docs/ai/ensemble-training-kit.md, docs/ai/models/vmaf_tiny_v5.md, docs/ai/per-pr-doc-bar.md, docs/ai/predictor.md, docs/development/ffmpeg-patches-refresh.md, docs/development/ossf-scorecard.md, tools/vmaf-tune/README.md, tools/vmaf-tune/src/vmaftune/codec_adapters/libaom.py, tools/vmaf-tune/src/vmaftune/per_shot.py.
Rebase impact: low. This PR updates stale public wording and docstrings after already-shipped implementations. It does not change the vmaf-tune row schema, CLI arguments, model defaults, or libvmaf public API.
Invariant to preserve on rebase: user-facing docs describe shipped implementation status, not old backlog labels. Keep intentional scaffold warnings only where the backing implementation or required external artefact is still genuinely missing.
Smoke-test after rebase:
rg -n '^# .*\(stub\)|^# .*stub|> \*\*Stub\*\*|0276-vmaf-tune-phase-d|full prose follows|later PR' \
docs/usage docs/ai docs/development tools/vmaf-tune/README.md -g '*.md'
PYTHONPATH=tools/vmaf-tune/src .venv/bin/python -m pytest \
tools/vmaf-tune/tests/test_resolution.py \
tools/vmaf-tune/tests/test_per_shot.py \
tools/vmaf-tune/tests/test_encode_dispatcher_per_adapter.py -q
.venv/bin/python -m ruff check \
tools/vmaf-tune/src/vmaftune/codec_adapters/libaom.py \
tools/vmaf-tune/src/vmaftune/per_shot.py
2026-05-14 — vmaf-tune Predictor Directory-Corpus Training¶
Files touched: tools/vmaf-tune/src/vmaftune/predictor_train.py, tools/vmaf-tune/tests/test_predictor_train.py, docs/usage/vmaf-tune.md, tools/vmaf-tune/README.md.
Rebase impact: low. This only broadens the trainer's corpus input resolver from a single JSONL file to a file-or-directory source. The corpus row schema, predictor input vector, shipped model defaults, and libvmaf public surface are unchanged.
Invariant to preserve on rebase: directory corpus traversal is recursive and sorted. Keep that determinism so repeated training over .workingdir2/corpus_run/ sees the same row order across filesystems.
Smoke-test after rebase:
PYTHONPATH=tools/vmaf-tune/src .venv/bin/python -m pytest \
tools/vmaf-tune/tests/test_predictor_train.py \
-q
2026-05-14 — vmaf-tune benchmark Phase-G Corpus Report¶
Files touched: tools/vmaf-tune/src/vmaftune/benchmark.py, tools/vmaf-tune/src/vmaftune/cli.py, tools/vmaf-tune/tests/test_benchmark.py, tools/vmaf-tune/AGENTS.md, docs/usage/vmaf-tune.md, docs/adr/0424-vmaf-tune-corpus-benchmark.md, docs/research/0106-vmaf-tune-corpus-benchmark.md.
Rebase impact: low. The new command is a read-only consumer of the existing Phase-A JSONL row schema. It does not change CORPUS_ROW_KEYS, libvmaf public API, FFmpeg patches, or encode/scoring behaviour.
Invariant to preserve on rebase: vmaf-tune benchmark must stay offline. It reads corpus rows and reports matched-quality encoder summaries; live encode comparisons remain owned by vmaf-tune compare.
Smoke-test after rebase:
PYTHONPATH=tools/vmaf-tune/src .venv/bin/python -m pytest \
tools/vmaf-tune/tests/test_benchmark.py -q
2026-05-14 — vmaf-tune auto winner selection¶
Files touched: tools/vmaf-tune/src/vmaftune/auto.py, tools/vmaf-tune/tests/test_auto_short_circuits.py, docs/usage/vmaf-tune.md, tools/vmaf-tune/AGENTS.md.
Rebase impact: low. The Phase F JSON schema now includes metadata.winner and a per-cell selected boolean, but corpus rows and libvmaf public APIs are unchanged.
Invariant to preserve on rebase: keep metadata.winner aligned with exactly one cells[].selected == true row. The selector remains quality/budget ordered per ADR-0428.
Smoke-test after rebase:
PYTHONPATH=tools/vmaf-tune/src .venv/bin/python -m pytest \
tools/vmaf-tune/tests/test_auto_short_circuits.py -q
2026-05-14 — testdata bench_perf portability¶
Files touched: testdata/bench_perf.py, testdata/test_bench_perf.py, docs/benchmarks.md.
Rebase impact: low. The performance JSON snapshots are unchanged; only the FFmpeg lavfi benchmark harness gains configuration and hardware-free smoke surfaces.
Invariant to preserve on rebase: bench_perf.py must not reintroduce mandatory machine-local paths. The MP4 decode test remains opt-in through --bbb-mp4-ref / VMAF_BBB_MP4_REF, while --require-all is the strict mode.
Smoke-test after rebase:
PYTHONPATH=. .venv/bin/python -m pytest testdata/test_bench_perf.py -q
.venv/bin/python testdata/bench_perf.py --list-tests
.venv/bin/python testdata/bench_perf.py --backend cpu --dry-run
2026-05-14 — CHUG HDR Corpus Ingestion + Feature Materialisation¶
Files touched: ai/scripts/chug_to_corpus_jsonl.py, ai/scripts/chug_extract_features.py, ai/tests/test_chug.py, scripts/dev/training_discovery_report.py, docs/ai/chug-ingestion.md, docs/ai/mos-corpora.md, docs/research/0101-training-discovery-synthesis-2026-05-14.md, docs/adr/0426-chug-hdr-corpus-ingestion.md, docs/adr/0427-chug-hdr-feature-materialisation.md, and ai/AGENTS.md.
Rebase impact: low to medium. The new CHUG adapter is fork-local and local-only, but it intentionally widens the MOS-corpus family with an HDR dataset and optional chug_* JSONL metadata fields.
Invariant to preserve on rebase: CHUG media and labels stay out of git. The adapter stores CHUG's raw mos_j as mos_raw_0_100 and maps the trainer-facing mos to [1, 5]; do not silently change that scale. The feature materialiser pairs distorted rows to the matching chug_content_name reference and scales distorted clips to reference geometry before extraction. Keep the license posture non-commercial/share-alike until the README/license mismatch is clarified upstream.
Smoke-test after rebase:
PYTHONPATH=ai/src .venv/bin/python -m pytest ai/tests/test_chug.py -q
python3 scripts/dev/training_discovery_report.py --output /tmp/training_discovery_report.md
2026-05-14 — vmaf-tune ladder Uncertainty CLI Wiring¶
Files touched: tools/vmaf-tune/src/vmaftune/ladder.py, tools/vmaf-tune/src/vmaftune/cli.py, tools/vmaf-tune/tests/test_ladder.py, tools/vmaf-tune/AGENTS.md, docs/usage/vmaf-tune.md, and docs/usage/vmaf-tune-ladder.md.
Rebase impact: low. The normal point-estimate ladder path is unchanged. When --with-uncertainty is set, corpus rows that contain a vmaf_interval object now flow through apply_uncertainty_recipe() before select_knees(). Rows without intervals use the active wide_interval_min_width as a conservative centred fallback interval so point-only corpora still participate in midpoint insertion.
Invariant to preserve on rebase: the uncertainty transform stays post-hull and pre-knee-selection. Do not run it before convex_hull(), or synthetic midpoint rungs can distort the Pareto filtering stage.
Smoke-test after rebase:
PYTHONPATH=tools/vmaf-tune/src .venv/bin/python -m pytest \
tools/vmaf-tune/tests/test_ladder.py \
tools/vmaf-tune/tests/test_ladder_uncertainty.py -q
2026-05-14 — vmaf-tune libaom-av1 saliency ROI Dispatch¶
Files touched: tools/vmaf-tune/src/vmaftune/saliency.py, tools/vmaf-tune/src/vmaftune/codec_adapters/libaom.py, tools/vmaf-tune/src/vmaftune/cli.py, vmaf-tune saliency tests, and the matching usage docs/state/changelog notes.
Rebase impact: low. The change only adds libaom-av1 to the existing saliency ROI dispatch table and uses the FFmpeg patch stack's top-level -qpfile <path> option. It does not alter scoring, predictor inputs, model files, or libvmaf public ABI.
Invariant to preserve on rebase: libaom-av1 saliency uses the shared x264-style 16x16 qpfile writer, but passes it as separate argv tokens ("-qpfile", path). Keep ephemeral cleanup aware of both key=path params and -qpfile path pairs.
Smoke-test after rebase:
PYTHONPATH=tools/vmaf-tune/src .venv/bin/python -m pytest \
tools/vmaf-tune/tests/test_saliency.py \
tools/vmaf-tune/tests/test_saliency_roi_adapters.py \
tools/vmaf-tune/tests/test_saliency_roi_codec.py \
-q
2026-05-14 — Metal Dispatch Support Table¶
Files touched: core/src/metal/dispatch_strategy.c, core/src/metal/dispatch_strategy.h, core/test/test_metal_smoke.c, core/src/metal/AGENTS.md, docs/backends/metal/index.md.
Rebase impact: low. The dispatch predicate now reflects the Metal kernels already compiled into the backend; it does not change kernel math, picture layout, metallib embedding, or public libvmaf_metal.h symbols.
Invariant to preserve on rebase: every newly-landed Metal extractor must append both its extractor name and its provided feature keys to g_metal_features. Unknown features, NULL contexts, and NULL names must keep returning 0.
Smoke-test after rebase:
meson setup build-metal -Denable_metal=enabled
ninja -C build-metal test_metal_smoke
meson test -C build-metal test_metal_smoke
2026-05-14 — Tiny-AI Bisect Cache Real-Feature Bridge¶
Files touched: ai/scripts/build_bisect_cache.py, ai/tests/test_build_bisect_cache.py, ai/testdata/bisect/README.md, docs/ai/bisect-model-quality.md, ai/AGENTS.md.
Rebase impact: low. The committed nightly cache remains generated from the existing deterministic synthetic seeds unless callers pass --source-features. The real-feature path only broadens the generator to materialise an operator-provided parquet into the same features.parquet + linear-ONNX timeline layout.
Invariant to preserve on rebase: the output feature order stays adm2, vif_scale0, vif_scale1, vif_scale2, vif_scale3, motion2, and the output target column stays named mos even when the source uses dmos, target, or score.
Smoke-test after rebase:
PYTHONPATH=ai/src .venv/bin/python -m pytest \
ai/tests/test_build_bisect_cache.py \
ai/tests/test_bisect_model_quality.py -q
PYTHONPATH=ai/src .venv/bin/python ai/scripts/build_bisect_cache.py --check
2026-05-14 — Vulkan VIF Manual Int64 Subgroup Reduction¶
Files touched: core/src/feature/vulkan/shaders/vif.comp, core/src/vulkan/AGENTS.md, docs/adr/0269-vif-ciede-precise-step-a.md, docs/research/0108-vulkan-vif-int64-subgroup-reduction-2026-05-14.md, docs/state.md.
Rebase impact: medium. The shader semantic change is intentionally small, but it is load-bearing for Vulkan API-1.4 parity on NVIDIA. Do not simplify the Phase-4 VIF accumulator path back to subgroupAdd(int64_t) when resolving upstream shader conflicts.
Invariant to preserve on rebase: vif.comp must keep GL_KHR_shader_subgroup_shuffle and the manual reduce_i64_subgroup(...) helper for all seven int64 accumulator fields. The helper exists because NVIDIA RTX 4090 + driver 595.71.05 produced non-deterministic integer_vif_scale2 output through subgroupAdd(int64_t) at Vulkan API 1.4.
Smoke-test after rebase:
glslc --target-env=vulkan1.3 -O \
core/src/feature/vulkan/shaders/vif.comp -o /tmp/vif.spv
ninja -C build-vulkan-int64 tools/vmaf
python3 scripts/ci/cross_backend_vif_diff.py \
--vmaf-binary "$PWD/build-vulkan-int64/tools/vmaf" \
--reference testdata/ref_576x324_48f.yuv \
--distorted testdata/dis_576x324_48f.yuv \
--width 576 --height 324 --feature vif --backend vulkan \
--device 0 --places 4
2026-05-14 — Saliency RGB ingest + SSIMULACRA2 public docs¶
Files touched: tools/vmaf-tune/src/vmaftune/saliency.py, tools/vmaf-tune/tests/test_saliency.py, tools/vmaf-tune/AGENTS.md, docs/usage/vmaf-tune.md, docs/metrics/ssimulacra2.md, docs/metrics/features.md, docs/adr/0430-saliency-rgb-ingest-and-ssimulacra2-docs.md, docs/research/0112-public-doc-gap-batch-2026-05-14.md.
Rebase impact: low. The changed saliency preprocessing is fork-local and keeps the same ONNX model input shape ([1, 3, H, W]).
Invariant to preserve on rebase: compute_saliency_map() must keep Y/U/V yuv420p ingest, BT.709 limited-range YUV-to-RGB conversion, and ImageNet normalisation before invoking saliency_student_v1. The old luma-replicated RGB path is no longer the user-facing contract.
Smoke-test after rebase:
PYTHONPATH=tools/vmaf-tune/src .venv/bin/python -m pytest \
tools/vmaf-tune/tests/test_saliency.py -q
scripts/docs/concat-adr-index.sh --check
2026-05-14 — test_score_pooled_eagain Sanitizer Deselect Retired¶
Files touched: .github/workflows/tests-and-quality-gates.yml, core/src/feature/x86/adm_avx2.c, docs/state.md.
Rebase impact: low. The sanitizer workflow now dispatches test_score_pooled_eagain again in ASan, UBSan, and TSan lanes. The AVX2 ADM helper keeps scalar-path parity for direct-LUT-range values: temp < 32768 returns temp and shift 0, while larger values still use the rounded 15-bit reduction. The remaining T-SANITIZER-DEFECTS-REVEALED-758 exclusions stay in place.
Invariant to preserve on rebase: sanitizer deselect regexes should contain only tests with an active state row. Do not re-add test_score_pooled_eagain unless a fresh sanitizer report is captured and tracked. Do not call __builtin_clz() for ADM direct-LUT values below 32768.
Smoke-test after rebase:
ASAN_OPTIONS=detect_leaks=1:halt_on_error=1 \
./core/build-asan-score/test/test_score_pooled_eagain
UBSAN_OPTIONS=halt_on_error=1:print_stacktrace=1 \
./core/build-ubsan-score/test/test_score_pooled_eagain
TSAN_OPTIONS=halt_on_error=1 \
./core/build-tsan-score/test/test_score_pooled_eagain
2026-05-14 — test_feature_collector Sanitizer Deselect Retired¶
Files touched: .github/workflows/tests-and-quality-gates.yml, docs/state.md.
Rebase impact: low. The sanitizer workflow now dispatches test_feature_collector again in ASan, UBSan, and TSan lanes. The remaining T-SANITIZER-DEFECTS-REVEALED-758 exclusions stay in place.
Invariant to preserve on rebase: sanitizer deselect regexes should contain only tests with an active state row. Do not re-add test_feature_collector unless a fresh sanitizer report is captured and tracked.
Smoke-test after rebase:
ASAN_OPTIONS=detect_leaks=1:halt_on_error=1 \
./core/build-asan-score/test/test_feature_collector
UBSAN_OPTIONS=halt_on_error=1:print_stacktrace=1 \
./core/build-ubsan-score/test/test_feature_collector
TSAN_OPTIONS=halt_on_error=1 \
./core/build-tsan-score/test/test_feature_collector
2026-05-14 — test_pic_preallocation Sanitizer Deselect Retired¶
Files touched: .github/workflows/tests-and-quality-gates.yml, docs/state.md.
Rebase impact: low. The sanitizer workflow now dispatches test_pic_preallocation again in ASan, UBSan, and TSan lanes. The remaining T-SANITIZER-DEFECTS-REVEALED-758 exclusions stay in place.
Invariant to preserve on rebase: sanitizer deselect regexes should contain only tests with an active state row. Do not re-add test_pic_preallocation unless a fresh sanitizer report is captured and tracked.
Smoke-test after rebase:
ASAN_OPTIONS=detect_leaks=1:halt_on_error=1 \
./core/build-asan-score/test/test_pic_preallocation
UBSAN_OPTIONS=halt_on_error=1:print_stacktrace=1 \
./core/build-ubsan-score/test/test_pic_preallocation
TSAN_OPTIONS=halt_on_error=1 \
./core/build-tsan-score/test/test_pic_preallocation
2026-05-14 — vmaf-tune libx265 encoder-stats parser¶
Files touched: tools/vmaf-tune/src/vmaftune/encoder_stats.py, tools/vmaf-tune/src/vmaftune/codec_adapters/x265.py, tools/vmaf-tune/tests/test_encoder_stats_parser_x264.py, tools/vmaf-tune/AGENTS.md, docs/usage/vmaf-tune.md.
Rebase impact: low. The corpus row schema stays at the existing v3 ten-column enc_internal_* contract; this only teaches the parser x265's pass-1 aliases (q-aq, icu, pcu, scu) and fractional CTU counts.
Invariant to preserve on rebase: x264 imb / pmb / smb and x265 icu / pcu / scu must continue to feed the same intra / predicted / skip ratio columns. Do not split the public corpus schema per codec.
Smoke-test after rebase:
PYTHONPATH=tools/vmaf-tune/src .venv/bin/python -m pytest \
tools/vmaf-tune/tests/test_encoder_stats_parser_x264.py -q
2026-05-14 — vmaf-tune predictor directory-corpus orchestration¶
Files touched: tools/vmaf-tune/src/vmaftune/predictor_train.py, tools/vmaf-tune/tests/test_predictor_train.py, docs/usage/vmaf-tune.md.
Rebase impact: low. The loader already supported recursive JSONL directories; this change removes stale is_file() gates in the trainer orchestration so CLI/API callers get the documented real-corpus path. Model format, feature order, corpus row schema, and shipped model bytes are unchanged.
Invariant to preserve on rebase: --corpus <directory> and train_all_codecs(corpus_path=<directory>) must call the same load_corpus() path as single-file inputs. Do not reintroduce file-only guards above the loader.
Smoke-test after rebase:
PYTHONPATH=tools/vmaf-tune/src .venv/bin/python -m pytest \
tools/vmaf-tune/tests/test_predictor_train.py -q
2026-05-15 — vmaf-tune sidecar CLI wiring¶
Files touched: tools/vmaf-tune/src/vmaftune/cli.py, tools/vmaf-tune/tests/test_cli_sidecar.py, tools/vmaf-tune/tests/test_sidecar.py, tools/vmaf-tune/AGENTS.md, docs/ai/local-sidecar-training.md, docs/usage/vmaf-tune.md, docs/research/0122-vmaf-tune-sidecar-cli-2026-05-15.md.
Rebase impact: low. This adds one top-level vmaf-tune sidecar subcommand group and does not change corpus row schemas, predictor ONNX schemas, codec adapters, libvmaf public APIs, or FFmpeg patches.
Invariant to preserve on rebase: the CLI must remain a thin wrapper over vmaftune.sidecar.SidecarPredictor. It must keep the same cache layout (<cache>/<predictor-version>/<codec>/state.json), random host UUID posture, and ShotFeatures feature names as the Python API. Do not add upload, hostname-derived IDs, or predictor mutation to this surface.
Smoke-test after rebase:
cd tools/vmaf-tune && ../../.venv/bin/python -m pytest \
tests/test_cli_sidecar.py tests/test_sidecar.py -q
2026-05-14 — vmaf-tune Phase-B Bisect Sample Clips¶
Files touched: tools/vmaf-tune/src/vmaftune/bisect.py, tools/vmaf-tune/src/vmaftune/cli.py, tools/vmaf-tune/tests/test_bisect.py, tools/vmaf-tune/tests/test_compare.py, tools/vmaf-tune/AGENTS.md, docs/usage/vmaf-tune.md, docs/usage/vmaf-tune-bisect.md, docs/research/0109-vmaf-tune-bisect-sample-clip-2026-05-14.md.
Rebase impact: low. The public addition is one vmaf-tune compare flag and one Python bisect argument. It does not change the compare report schema, codec adapter registry, libvmaf public API, or FFmpeg patch stack.
Invariant to preserve on rebase: sample_clip_seconds in bisect_target_vmaf, make_bisect_predicate, and vmaf-tune compare must compute one centre-anchored sample window and thread it into both EncodeRequest (sample_clip_start_s / sample_clip_seconds) and ScoreRequest (frame_skip_ref / frame_cnt). Bitrate must be normalised against the sample duration when sample-clip mode is active. Unknown duration, non-positive framerate, or samples not shorter than the source remain full-source mode.
Smoke-test after rebase:
PYTHONPATH=tools/vmaf-tune/src .venv/bin/python -m pytest \
tools/vmaf-tune/tests/test_bisect.py \
tools/vmaf-tune/tests/test_compare.py -q
2026-05-14 — vmaf-tune ladder spacing alias fix¶
Files touched: tools/vmaf-tune/src/vmaftune/cli.py, tools/vmaf-tune/src/vmaftune/ladder.py, tools/vmaf-tune/tests/test_ladder.py, tools/vmaf-tune/AGENTS.md, docs/usage/vmaf-tune.md.
Rebase impact: low. This only keeps the Phase-E CLI choices and library spacing modes aligned. The ladder hull math, default sampler, manifest schema, and encode/scoring behaviour are unchanged.
Invariant to preserve on rebase: argparse choices for vmaf-tune ladder --spacing must stay in lockstep with ladder.select_knees(). vmaf is the documented perceptual spacing mode; uniform remains a backwards-compatible alias for that same mode.
Smoke-test after rebase:
2026-05-15 — Tiny-AI real-weight limitation docs¶
Files touched: docs/ai/roadmap.md, docs/ai/models/fastdvdnet_pre.md, docs/metrics/features.md, core/src/dnn/AGENTS.md, core/src/feature/AGENTS.md.
Rebase impact: low. This is a documentation / invariant-note cleanup that aligns user-facing docs with the already-shipped smoke: false FastDVDnet and TransNet V2 registry entries. Model bytes, registry schema, extractor I/O names, and runtime behaviour are unchanged.
Invariant to preserve on rebase: do not reintroduce placeholder-only wording for fastdvdnet_pre or transnet_v2. The remaining follow-ups are the FFmpeg temporal-filter consumer, luma-native FastDVDnet retrain, per-shot CRF aggregation, and true RGB / bilinear TransNet thumbnails.
Smoke-test after rebase:
rg -n "real upstream weights are tracked|ADR-0246|0253-fastdvdnet" \
docs/ai docs/metrics core/src/dnn/AGENTS.md core/src/feature/AGENTS.md
2026-05-15 — vmaf-roi High-Bit-Depth Input¶
Files touched: core/tools/vmaf_roi.c, core/tools/test/meson.build, core/tools/test/test_vmaf_roi_high_bitdepth.sh, core/tools/AGENTS.md, docs/usage/vmaf-roi.md, docs/research/0123-vmaf-roi-high-bitdepth-2026-05-15.md.
Rebase impact: low. This extends an existing CLI flag and does not change libvmaf public APIs, encoder sidecar schemas, or FFmpeg patches.
Invariant to preserve on rebase: vmaf-roi --bitdepth 10|12|16 must seek using full planar YUV frame bytes, including chroma planes and 16-bit sample containers, then downscale luma to the existing luma8 saliency-model contract. Unsupported depths such as 9-bit remain rejected.
Smoke-test after rebase:
2026-05-15 — vmaf-perShot 4:2:2 / 4:4:4 Input¶
Files touched: core/tools/vmaf_per_shot.c, core/tools/test/test_vmaf_per_shot.sh, core/tools/AGENTS.md, docs/usage/vmaf-perShot.md, docs/research/0124-vmaf-pershot-422-444-2026-05-15.md.
Rebase impact: low. This extends one existing CLI option and does not change the CSV / JSON plan schema, libvmaf public APIs, or FFmpeg patch stack.
Invariant to preserve on rebase: vmaf-perShot remains luma-only for detection and CRF prediction, but --pixel_format 420|422|444 must count the selected planar chroma layout when skipping to the next frame. --bitdepth remains limited to 8|10|12|16.
Smoke-test after rebase:
2026-05-15 — CUDA psnr_hvs DCT Parallelisation¶
Files touched: core/src/feature/cuda/integer_psnr_hvs/psnr_hvs_score.cu, core/src/feature/cuda/AGENTS.md, docs/backends/cuda/overview.md, docs/research/0130-cuda-psnr-hvs-dct-parallel-2026-05-15.md.
Rebase impact: low. The host lifecycle, feature names, CLI surface, and public APIs are unchanged. This is a CUDA-kernel scheduling optimisation for an existing extractor.
Invariant to preserve on rebase: only the integer 8x8 DCT passes run across the first eight CUDA threads. Float means, variances, masking, and masked-error accumulation stay thread-0 serial in CPU scan order; do not convert them to warp/block reductions without a separate numeric-contract ADR and cross-backend tolerance update.
Smoke-test after rebase:
python3 scripts/ci/cross_backend_vif_diff.py \
--vmaf-binary "$PWD/core/build-cuda/tools/vmaf" \
--reference testdata/ref_576x324_48f.yuv \
--distorted testdata/dis_576x324_48f.yuv \
--width 576 --height 324 --feature psnr_hvs --backend cuda --places 3
2026-05-15 — test_cli_parse Sanitizer Deselect Retired¶
Files touched: .github/workflows/tests-and-quality-gates.yml, docs/state.md, changelog.d/fixed/sanitizer-cli-parse.md.
Rebase impact: low. This only narrows the ADR-0347 sanitizer deselect regexes after re-verifying test_cli_parse on current master; the CLI parser behavior and public options are unchanged.
Invariant to preserve on rebase: keep test_cli_parse out of the ASan / UBSan / TSan EXCLUDE regexes unless a new sanitizer report is captured and tracked in docs/state.md.
Smoke-test after rebase:
ASAN_OPTIONS=detect_leaks=1:halt_on_error=1:abort_on_error=1:print_summary=1 ./core/build-asan-cli/test/test_cli_parse
UBSAN_OPTIONS=halt_on_error=1:abort_on_error=1:print_summary=1:print_stacktrace=1 ./core/build-ubsan-cli/test/test_cli_parse
TSAN_OPTIONS=halt_on_error=1 ./core/build-tsan-cli/test/test_cli_parse
2026-05-15 — test_predict Sanitizer Deselect Retired¶
Files touched: .github/workflows/tests-and-quality-gates.yml, docs/state.md, changelog.d/fixed/sanitizer-predict.md.
Rebase impact: low. This only narrows the ADR-0347 sanitizer deselect regexes after re-verifying test_predict on current master; prediction logic, model loading, and output scores are unchanged.
Invariant to preserve on rebase: keep test_predict out of the ASan / UBSan / TSan EXCLUDE regexes unless a new sanitizer report is captured and tracked in docs/state.md.
Smoke-test after rebase:
ASAN_OPTIONS=detect_leaks=1:halt_on_error=1:abort_on_error=1:print_summary=1 ./core/build-asan-predict/test/test_predict
UBSAN_OPTIONS=halt_on_error=1:abort_on_error=1:print_summary=1:print_stacktrace=1 ./core/build-ubsan-predict/test/test_predict
TSAN_OPTIONS=halt_on_error=1 ./core/build-tsan-predict/test/test_predict
2026-05-15 — vmaf-tune HDR Dispatch Coverage¶
Files touched: tools/vmaf-tune/src/vmaftune/hdr.py, tools/vmaf-tune/tests/test_hdr.py, tools/vmaf-tune/tests/test_auto_short_circuits.py, tools/vmaf-tune/AGENTS.md, docs/usage/vmaf-tune.md, docs/usage/vmaf-tune-hdr-and-sampling.md, docs/research/0126-vmaf-tune-hdr-dispatch-coverage-2026-05-15.md.
Rebase impact: low. This extends the existing ADR-0300 dispatch table only; it does not change corpus schema, codec-adapter quality knobs, or HDR model lookup.
Invariant to preserve on rebase: hdr_codec_args() remains the single HDR argv contract. Hardware HEVC rows should emit p010le + main10 plus global color tags; hardware AV1 rows should emit p010le plus global color tags; codec-private mastering-display / MaxCLL flags stay limited to verified families.
Smoke-test after rebase:
PYTHONPATH=tools/vmaf-tune/src .venv/bin/python -m pytest \
tools/vmaf-tune/tests/test_hdr.py \
tools/vmaf-tune/tests/test_auto_short_circuits.py -q
2026-05-15 — Docs Pages Strict-Anchor Repair¶
Files touched: docs/ai/quantization.md, docs/api/gpu.md, docs/metrics/ssimulacra2.md, docs/usage/vmaf-tune.md, changelog.d/fixed/docs-pages-anchor-strict.md.
Rebase impact: low. This only corrects MkDocs-rendered internal anchors and adds the missing docs/api/gpu.md HIP / Metal section targets consumed by docs/api/index.md.
Invariant to preserve on rebase: mkdocs build --strict must stay green on master before a docs-affecting PR is merged; do not weaken validation.links.anchors: warn to hide anchor drift.
Smoke-test after rebase:
2026-05-15 — CHUG FULL_FEATURES Parquet Metadata Enrichment¶
Files touched: ai/scripts/enrich_k150k_parquet_metadata.py, ai/tests/test_enrich_k150k_parquet_metadata.py, ai/AGENTS.md, docs/ai/chug-ingestion.md, and docs/ai/datasets/k150k.md.
Rebase impact: low. This adds a recovery utility for local FULL_FEATURES parquet jobs that predate --metadata-jsonl; it does not change the extraction schema or feature column order.
Invariant to preserve on rebase: the enrichment utility matches rows by clip_name, fills missing metadata cells by default, writes parquet atomically, and leaves feature/MOS columns unchanged unless the operator passes --overwrite-metadata.
Smoke-test after rebase:
PYTHONPATH=ai/src .venv/bin/python -m pytest \
ai/tests/test_enrich_k150k_parquet_metadata.py \
ai/tests/test_extract_k150k_features.py -q
fix/saliency-per-mb-eval-2026-05-15 — CLI short-opt + bench atoi fix (Batch 5)¶
Branch: fix/saliency-per-mb-eval-2026-05-15
Files touched: core/tools/cli_parse.c, core/tools/vmaf_bench.c, core/test/test_cli_parse.c, docs/usage/cli.md.
Rebase impact: low. cli_parse.c and vmaf_bench.c are upstream-shared files; if Netflix/vmaf ever adds a new short option or touches the same switch block, the case 'c': fall-through arm may need to be re-applied. The invariant comment (INVARIANT (ADR-0438)) marks the intent clearly. vmaf_bench.c changes are in the SYCL-gated #if defined(HAVE_SYCL) block; upstream is unlikely to add atoi back.
Invariant to preserve on rebase: every entry in short_opts[] in core/tools/cli_parse.c must have a matching case arm in the switch (o) block inside cli_parse(). If Netflix adds a new short option upstream without a case, the same silent-drop bug recurs.
Smoke-test after rebase:
meson setup build -Denable_cuda=false -Denable_sycl=false
ninja -C build test/test_cli_parse
./build/test/test_cli_parse # expect: 18 tests run, 18 passed
fix/motion-fps-weight-all-gpu-backends — motion_fps_weight parity across all GPU twins¶
Branch: fix/saliency-per-mb-eval-2026-05-15 (squash PR #863)
Files touched: core/src/feature/cuda/integer_motion_v2_cuda.c, core/src/feature/sycl/integer_motion_v2_sycl.cpp, core/src/feature/vulkan/motion_v2_vulkan.c, core/src/feature/hip/integer_motion_v2_hip.c, core/src/feature/metal/integer_motion_v2_metal.mm, core/src/feature/cuda/float_motion_cuda.c, core/src/feature/sycl/float_motion_sycl.cpp, core/src/feature/vulkan/float_motion_vulkan.c, core/src/feature/hip/float_motion_hip.c, core/src/feature/metal/float_motion_metal.mm.
Rebase impact: low. All touched files are fork-local or fork-added GPU twins; upstream Netflix/vmaf does not maintain any GPU motion extractor files. No upstream-shared path is modified.
Invariant to preserve on rebase: motion_fps_weight must remain present in every motion-family GPU twin's VmafOption options[] table and applied identically (see canonical note in core/src/feature/cuda/AGENTS.md). If a future PR introduces a new motion GPU backend or a new motion-related option, the same option table and application math must be replicated across all twins in the same PR.
Smoke-test after rebase:
meson setup build -Denable_cuda=false -Denable_sycl=false
ninja -C build
meson test -C build --suite=fast
core/test/meson.build — suite-tagging invariant (fix/meson-suite-fast)¶
Files touched: core/test/meson.build, core/test/AGENTS.md.
Rebase impact: moderate. Upstream Netflix/vmaf periodically adds new test() calls to core/test/meson.build without suite: arguments (that is the upstream convention). Every upstream sync or port-upstream-commit cherry-pick that touches this file must be followed by:
Any output is a missing tag — add the appropriate suite: before merging. Failure to do so silently breaks meson test -C build --suite=fast (the pre-push gate) because Meson's --suite filter matches only tests that declare the named suite; untagged tests are invisible to the filter and the command exits 0 with zero tests run.
Invariant to preserve on rebase: every test(...) call in core/test/meson.build carries a suite: keyword argument. The fast suite is the pre-push gate; simd and gpu are secondary selectors for CI matrix jobs. See core/test/AGENTS.md for the full tag matrix.
Smoke-test after rebase:
meson setup build -Denable_cuda=false -Denable_sycl=false
ninja -C build
meson test -C build --suite=fast --list # must print >20 tests, not 0
perf/cambi-calculate-c-values-avx512-neon-2026-05-16 (ADR-0452)¶
What changed: Added calculate_c_values_row_avx512 and calculate_c_values_row_neon as siblings of the existing calculate_c_values_row_avx2. Updated cambi.c dispatch to assign calculate_c_values_avx512 on AVX-512 hosts and corrected the NEON wrapper to call calculate_c_values_row_neon instead of the scalar fallback.
Rebase impact: low. All modified files are fork-local additions to cambi SIMD infrastructure; upstream Netflix/vmaf does not maintain AVX-512 or NEON CAMBI kernels. No public API surface is changed.
Invariant to preserve on rebase: The twin-update rule (x86/AGENTS.md, arm64/AGENTS.md) now requires that every cambi inner-loop function ported to AVX2 ships with AVX-512 + NEON siblings in the same PR. Do not merge a cambi AVX2 kernel without the matching AVX-512 + NEON files and a dispatch update in cambi.c.
refactor/gpu-dispatch-parse-dedup — shared GPU dispatch env tokenizer (ADR-0483)¶
Branch: refactor/gpu-dispatch-parse-dedup
Files touched: core/src/gpu_dispatch_parse.h (new), core/src/cuda/dispatch_strategy.c, core/src/sycl/dispatch_strategy.cpp, core/src/vulkan/dispatch_strategy.c.
Rebase impact: low. The three dispatch_strategy TUs are fork-local; upstream Netflix/vmaf does not have dispatch_strategy.c files. No public headers, no meson sources, and no link-time symbols change — the new gpu_dispatch_parse.h is a header-only static inline and is not added to any meson source list.
Invariant to preserve on rebase: k_<backend>_strategy_names[] index 0 must equal the backend enum's default value (e.g. VMAF_CUDA_DISPATCH_DIRECT = 0). When adding new strategy enum values, append to both the enum and the table; never reorder either.
perf/chug-drop-ssimulacra2-cuda-self-vs-self-2026-05-16 — K150K/CHUG self-vs-self extraction schema v2¶
Branch: perf/chug-drop-ssimulacra2-cuda-self-vs-self-2026-05-16
Files touched: ai/scripts/extract_k150k_features.py, ai/AGENTS.md, ai/tests/test_extract_k150k_no_ssimulacra2.py.
Rebase impact: low. All touched files are fork-local tiny-AI infrastructure; upstream Netflix/vmaf does not maintain ai/ or K150K extraction pipelines. No upstream-shared C/C++/headers are modified.
Invariant to preserve on rebase: the K150K extraction script (extract_k150k_features.py) is a fork-only feature. If upstream adds its own extract_features.py or similar, keep them separate under different package names; do not merge them. The parquet schema v2 (21-feature, no ssimulacra2) is now authoritative for new K150K/CHUG extraction runs. Existing v1 parquets (22-feature, with ssimulacra2) are grandfathered in; loaders must handle both by detecting feature count at runtime or reading a schema-version sidecar (future work).
Smoke-test after rebase:
feat/psnr-hvs-vulkan-enable-chroma-2026-05-16 — enable_chroma option for psnr_hvs_vulkan (ADR-0585)¶
- Touches:
core/src/feature/vulkan/psnr_hvs_vulkan.c - Invariant:
enable_chromadefaults totrue; do not flip. Whenfalse,n_planes=1, chroma pipelines are not created, and the combinedpsnr_hvsscore is suppressed.close_fex()relies onVK_NULL_HANDLEguards for the chroma pipeline variants. - No rebase impact on upstream files (fork-local Vulkan extractor).
Smoke-test after rebase:
# Default (enable_chroma=true): expect psnr_hvs_y + psnr_hvs_cb + psnr_hvs_cr + psnr_hvs
./build/tools/vmaf --reference src01_hrc00_576x324.yuv \
--distorted src01_hrc01_576x324.yuv \
--width 576 --height 324 --feature psnr_hvs_vulkan
# Luma-only: expect only psnr_hvs_y
./build/tools/vmaf --reference src01_hrc00_576x324.yuv \
--distorted src01_hrc01_576x324.yuv \
--width 576 --height 324 \
--feature-opts 'psnr_hvs_vulkan=enable_chroma=false'
feat/hip-float-adm-real-2026-05-16 — HIP float_adm ninth consumer (ADR-0468)¶
Branch: feat/hip-float-adm-real-2026-05-16
Files touched: core/src/feature/hip/float_adm_hip.c (new), core/src/feature/hip/float_adm_hip.h (new), core/src/feature/hip/float_adm/float_adm_score.hip (new), core/src/hip/meson.build (add TU to hip_sources), core/src/meson.build (add float_adm_score to hip_kernel_sources), core/src/feature/feature_extractor.c (extern decl + #if HAVE_HIP list row), docs/adr/0468-hip-float-adm-real-kernel.md (new), docs/adr/README.md (index row), changelog.d/added/hip-float-adm-real-kernel.md (new).
Rebase impact: low. All new files are fork-local HIP infrastructure; upstream Netflix/vmaf does not maintain a HIP backend. The only upstream-shared file touched is feature_extractor.c, where the change is limited to adding an extern declaration and a single list entry inside #if HAVE_HIP — a block upstream does not have.
Invariant to preserve on rebase: float_adm_hip.c must track float_adm_cuda.c semantically. Any change to the four pipeline stages (DWT coefficients, decouple angle flag parenthesisation, CM threshold 8-neighbour sum, border factor) must be mirrored in both the CUDA and HIP TUs. The warp-size difference (CUDA=32 vs HIP=64) means the shared-memory partial arrays differ in size (FADM_WARPS_PER_BLOCK = 8 vs 4); this is correct and must not be unified.
perf/adm-p-norm-fast-path-vif-arm64-malloc-2026-05-16 (ADR-0463)¶
What changed: Added adm_cm_s_p3, adm_csf_den_scale_s_p3, and adm_sum_cube_s_p3 fast-path variants in adm_tools.c; dispatch added in adm.c:compute_adm. Removed per-call aligned_malloc from the scalar fallback paths of vif_filter1d_s, vif_filter1d_sq_s, and vif_filter1d_xy_s in vif_tools.c — the caller-supplied tmpbuf is used instead.
Rebase impact: low. All modified files (adm_tools.c, adm_tools.h, adm.c, vif_tools.c) are shared with upstream Netflix/vmaf. The ADM changes add new symbols (no existing signatures altered). The VIF changes only remove local malloc/free; the function signatures and caller-supplied tmpbuf contract are unchanged.
Invariant to preserve on rebase: When upstream Netflix/vmaf modifies adm_cm_s, adm_csf_den_scale_s, or adm_sum_cube_s, the corresponding _p3 variants in the fork must receive the same logic change (minus the powf path). When upstream modifies vif_filter1d_* scalar fallbacks, ensure they do not reintroduce aligned_malloc in the fallback body. See core/src/feature/AGENTS.md performance-invariant section.
fix/dispatch-strategy-registry-audit-2026-05-15 — dispatch registry deduplication + HIP/Metal fixes¶
Touches: core/src/feature/feature_extractor.c (SYCL/Vulkan sections of feature_extractor_list[]), core/src/hip/dispatch_strategy.c, core/src/metal/dispatch_strategy.c.
Rebase impact: low for the SYCL/Vulkan deduplication (purely cosmetic — first-match semantics mean behaviour is unchanged). Medium for HIP and Metal dispatch-supports: if an upstream sync adds new feature_extractor_list[] entries for HIP or Metal extractors, they must also be added to g_hip_features[] / g_metal_features[] in the same commit.
Invariant to preserve on rebase: every vmaf_fex_*_hip extractor registered in feature_extractor_list[] must appear in g_hip_features[] in core/src/hip/dispatch_strategy.c. Every vmaf_fex_*_metal extractor must appear (by extractor .name and all provided_features[] keys) in g_metal_features[] in core/src/metal/dispatch_strategy.c. The build does not enforce this — run scripts/ci/check-dispatch-registry.sh after any kernel addition.
Smoke-test after rebase:
meson setup build -Denable_hip=true -Denable_hipcc=false
ninja -C build
# Must compile without errors; vmaf_fex_float_adm_hip must be registered.
# With a ROCm 6+ toolchain:
meson setup build -Denable_hip=true -Denable_hipcc=true
ninja -C build
meson test -C build --suite=fast
perf/vif-cuda-smem-staging-2026-05-16 (ADR-0454)¶
Files touched: core/src/feature/cuda/integer_vif/filter1d.cu, core/src/feature/cuda/AGENTS.md.
Rebase impact: low. filter1d.cu is a fork-local CUDA kernel file (Netflix/vmaf does not maintain CUDA implementations). AGENTS.md is fork-local. No upstream-shared path, public header, or build file is modified.
Invariant to preserve on rebase: the __shared__ smem staging in all four filter template functions must be preserved verbatim (see canonical note added to AGENTS.md). If a future upstream commit adds a filter1d.cu-equivalent (unlikely — Netflix does not ship GPU VIF), reconcile by keeping the smem staging on our side. Do not remove the __syncthreads() between the cooperative load and the compute phase — that barrier is the only thing ordering the smem writes from all threads before any thread reads.
perf/ssimulacra2-cuda-blur-fusion-transpose — 3-channel kernel fusion + V-pass transpose (ADR-0456)¶
- Touches:
core/src/feature/cuda/ssimulacra2/ssimulacra2_blur.cu,core/src/feature/cuda/ssimulacra2_cuda.c,core/src/feature/cuda/AGENTS.md. - Invariant:
ssimulacra2_blur.cumust export exactly 5 kernel symbols:ssimulacra2_transpose,ssimulacra2_blur_h,ssimulacra2_blur_h3,ssimulacra2_blur_v,ssimulacra2_blur_v3_transposed. The host dispatch inssimulacra2_cuda.clooks up all 5 viacuModuleGetFunctionat init time; removing or renaming any symbol causes a hard init failure. The fused kernels usegridDim.z = 3withplane_stride = width * height(full-resolution constant stride, NOT scale-adjusted stride); any change to the stride contract must propagate to both the kernel and the three host dispatch helpers (ss2c_launch_blur_h3,ss2c_launch_transpose,ss2c_launch_blur_v3_transposed). The--fmad=falseflag in thecuda_cu_extra_flagsmap incore/src/meson.buildis load-bearing for theplaces=4cross-backend parity gate; do not remove it. - Re-test:
meson setup core/build_cuda libvmaf -Denable_cuda=true -Denable_sycl=false
ninja -C core/build_cuda tools/vmaf
# Correctness gate:
./core/build_cuda/tools/vmaf \
--reference testdata/ref_576x324_48f.yuv \
--distorted testdata/dis_576x324_48f.yuv \
--width 576 --height 324 --pixel_format 420 --bitdepth 8 \
--feature ssimulacra2_cuda --output /tmp/cuda_out.xml --backend cuda
# Cross-backend parity:
python3 scripts/ci/cross_backend_parity_gate.py \
--features ssimulacra2 --backends cpu cuda --places 4
perf/adm-cm-cuda-warp-reduce-fusion — ADM CM i4 warp-reduce fusion (2026-05-16)¶
What changed: integer_adm/adm_cm.cu — i4_adm_cm_line_kernel (writes INT32 per-thread values to accum_per_thread global scratch) + adm_cm_reduce_line_kernel_4 (reads scratch, cubic-accumulates, warp-reduces) replaced by a single i4_adm_cm_line_kernel_fused kernel that does all three steps internally and writes via atomicAdd_int64. integer_adm_cuda.c updated: one cuLaunchKernel per scale (was two), func_adm_cm_reduce_line_kernel_4 removed from AdmStateCuda.
Rebase impact: low. All touched files are fork-local CUDA kernels and their host glue; Netflix upstream does not maintain GPU ADM CM kernels.
Invariant to preserve on rebase: the fused kernel's shift constants (shift_sq=30, add_shift_sq=1<<29, shift_cub=ceil(log2(w)), shift_inner_accum=ceil(log2(h))) must match those used by adm_cm_reduce_line_kernel in the same file for scale != 0. If the reduce kernel's constants are ever changed, the fused kernel's constants must be updated in lockstep.
Smoke-test after rebase:
meson setup build -Denable_cuda=true -Denable_sycl=false
ninja -C build
meson test -C build --suite=fast
python3 scripts/ci/cross_backend_parity_gate.py \
--features vif --backends cpu cuda --places 4
python3 scripts/ci/cross_backend_parity_gate.py --features adm --backends cpu cuda --places 4
---
### Ghost `moment_vulkan.c` removed (fix/drop-ghost-moment-vulkan-c)
PR #1067 re-introduced the pre-rename `moment_vulkan.c` alongside the new
`float_moment_vulkan.c` that PR #1046 had established. The ghost file was
deleted and `core/src/vulkan/meson.build` updated to reference
`float_moment_vulkan.c` exclusively.
**Rebase impact**: any branch that modified `moment_vulkan.c` must be
re-targeted to `float_moment_vulkan.c` instead.
**Smoke-test after rebase**:
```bash
ninja -C build && echo "no duplicate symbol error"
perf/cache-rfe-hw-flags — cache rfe_hw_flags bitmask (F2-B)¶
File changed: core/src/libvmaf.c — VmafContext struct + vmaf_init + vmaf_use_feature + vmaf_read_pictures.
No rebase impact: the change is entirely internal to libvmaf.c; no public header touched, no FFmpeg-patch surface changed.
Invariant: rfe_hw_flags_dirty must be set to true in vmaf_init (after the memset zeroes it to false). If a future refactor moves the memset or adds a second init path, the dirty flag must be set at every init site.
Smoke-test after rebase:
meson setup build -Denable_cuda=true -Denable_sycl=false
ninja -C build src/liblibvmaf.a.p/libvmaf_src_libvmaf.c.o
# Expected: compiles without error or warning
PR #1067 clobbered four GPU feature options (fix/enable-chroma-pr1067-regression)¶
PR #1067 (bootstrap name-builder refactor) merged a stale base that pre-dated four option additions and overwrote them:
integer_psnr_metal.mm: lostenable_chromafield + option entry +n_planesguard (PR #986; restored in BUG-048 / ADR-1322 and pinned bytest_gpu_psnr_option_parity_contract.py)float_psnr_metal.mm: lostenable_chroma+ per-plane dispatch loop +n_planes(PR #978)psnr_vulkan.c: ceiling division reverted to floor division for chroma geometry (PR #878)vif_vulkan.c: lostvif_skip_scale0field + option entry + score-suppression guards (PR #1057)
Rebase impact: any branch that adds options to these four files and was branched before PR #1067 merged must be rebased onto master (post-fix) to avoid re-clobbering these options.
Smoke-test after rebase:
meson setup build -Denable_cuda=false -Denable_sycl=false --wipe
ninja -C build
meson test -C build --suite=fast
perf/chug-sidecar-bit-depth-key-f6b (2026-05-17)¶
Files touched: ai/scripts/extract_k150k_features.py
What changed: Added "chug_bit_depth" to the keep allowlist in _load_jsonl_metadata. Without this field, _geometry_from_sidecar always returned the default yuv420p pix_fmt even for 10-bit CHUG clips (F6-B / Research-0135). Corrected module and _process_clip docstrings that overstated the ffprobe-skip extent.
Rebase impact: no rebase impact. ai/scripts/extract_k150k_features.py is fork-local; there is no upstream-Netflix equivalent.
Smoke-test after rebase:
feat/hip-psnr-enable-chroma (2026-05-16)¶
File: core/src/feature/hip/integer_psnr_hip.c
No rebase impact: the change is additive (new option + plane-loop). If upstream later changes the HIP PSNR submit/collect call-graph, re-check that the per-plane loop in submit_fex_hip and collect_fex_hip matches whatever new structure upstream introduces. The kernel (psnr_score.hip) is unchanged.
fix/float-vif-skip-scale0-hip-metal¶
Files: core/src/feature/sycl/float_vif_sycl.cpp, core/src/feature/hip/float_vif_hip.c, core/src/feature/metal/float_vif_metal.mm
No rebase impact: the changes are additive (new field + option + host-side guard in the collect path). GPU kernels are unchanged. If upstream later changes the float_vif collect path or adds vif_skip_scale0 natively, re-check that scale-0 suppression in all three backends matches the CPU implementation in float_vif.c.
fix/adm-metal-missing-options¶
File: core/src/feature/metal/integer_adm_metal.mm
No rebase impact: the change moves three implicit defaults from init_fex_metal into the options table. The struct fields and kernel dispatch are unchanged. If upstream adds its own Metal ADM options or renames the default macros, re-check that DEFAULT_ADM_CSF_SCALE, DEFAULT_ADM_CSF_DIAG_SCALE, and DEFAULT_ADM_NOISE_WEIGHT still resolve correctly.
fix/docs-pr-strict-check-batch18¶
Files: .github/workflows/lint-and-format.yml, .github/workflows/required-aggregator.yml
No rebase impact: the change adds a new CI job (docs-lint) and a new entry in the required-aggregator check list. Both are purely additive and contain no fork-local logic that upstream could change. If upstream adds its own docs-lint CI, dedup by dropping our job or merging the two.
chore/svm-h-remove-orphaned-xxx-marker¶
File: core/src/svm.h
No rebase impact: removes an empty /* XXX */ comment from the vendored libsvm header and folds two trailing comment lines on free_sv into one two-line block. If upstream libsvm updates svm.h, re-apply by re-removing the marker (it originates in libsvm upstream and may reappear).
model/tiny/vmaf_tiny_v1_medium.onnx¶
Files: model/tiny/vmaf_tiny_v1_medium.onnx (binary, inline repack), model/tiny/vmaf_tiny_v1_medium.onnx.data (deleted), changelog.d/fixed/model-tiny-v1-medium-external-data.md.
No rebase impact: model-only binary change, no C-API or registry schema change. If upstream Netflix ever ships a file with this name, treat as a conflict and keep the fork's version (it is a fork-local model, not an upstream artifact).
docs/motion-dedicated-page¶
Files: docs/metrics/motion.md (new), docs/metrics/features.md, docs/adr/0491-motion-dedicated-doc-page.md, docs/adr/README.md, changelog.d/added/motion-dedicated-doc-page.md.
No rebase impact: doc-only addition. If upstream Netflix adds a motion extractor or renames existing ones, update docs/metrics/motion.md to match — no code change required.
fix/vulkan-vif-shader-fp64-for-bit-exact¶
Files: core/src/feature/vulkan/shaders/vif.comp, core/src/vulkan/common.c, docs/adr/0492-vulkan-vif-shader-fp64-g-computation.md, docs/backends/vulkan/overview.md, changelog.d/fixed/vulkan-vif-fp64-g-computation.md.
Rebase sensitivity (medium): vif.comp carries the fp64 extension declaration at line 68 and the revised g/sv_sq block at ~line 540. If upstream Netflix modifies the VIF computation path in integer_vif.c, re-verify that the double-precision GLSL block still mirrors the CPU reference exactly (especially the eps constant and the int32 truncation order for sv_sq). The common.c shaderFloat64 probe must stay in sync with any new device-feature guards added to the same function.
fix/test-output-portable-tempfile¶
Files: core/test/test_output.c, changelog.d/fixed/test-output-portable-tempfile.md.
No rebase impact: core/test/test_output.c is fork-local (added by PR #963, fork-only coverage gap follow-up; Netflix upstream has no equivalent file). The Windows-portable make_temp_path() helper sits inside that file and is not exported. Upstream syncs do not touch it.
fix/mcp-probe-findings-2026-05-17¶
Files: mcp-server/vmaf-mcp/src/vmaf_mcp/server.py, mcp-server/vmaf-mcp/tests/test_probe_findings_2026_05_17.py, mcp-server/vmaf-mcp/tests/test_backend_dispatch.py, testdata/bench_all.sh, docs/adr/0495-mcp-probe-bug-fixes.md, docs/mcp/tools.md, docs/state.md, changelog.d/fixed/mcp-probe-bug-cluster-2026-05-17.md.
Rebase sensitivity (low): all changes live in the fork-only mcp-server/ tree and the fork-added testdata/bench_all.sh (Netflix upstream has neither). The new _BACKEND_DISABLE / _BACKEND_PROBE_CACHE helpers and the --no_<backend> flag plumbing in _run_vmaf_score assume the libvmaf CLI continues to advertise --no_<backend> switches in --help and to accept them on the command line. If upstream ever removes them (the new --backend $NAME exclusive selector landed in the fork on 2026-04-28 — see bench_all.sh header comments), swap _probe_backends to parse the selector grammar instead. No upstream- mirrored file is touched.
fix/bbb-e2e-v2-bug-cluster-2026-05-18¶
Files: core/tools/vmaf.c (init_gpu_backends explicit-backend gating + amend_json_with_backend_used helper), tools/vmaf-tune/src/vmaftune/{score,bisect,corpus,ladder,report,encode,cli}.py, tools/vmaf-tune/tests/test_bbb_e2e_v2_bug_cluster.py, dev/Containerfile (matplotlib), dev/scripts/dev-mcp-entrypoint.sh (mkdir -p /tmp), docs/adr/0498-vmaf-tune-bbb-e2e-v2-bug-cluster.md, docs/adr/README.md, docs/state.md, docs/usage/vmaf-tune.md, docs/backends/index.md, docs/development/dev-mcp.md, changelog.d/fixed/vmaf-tune-bbb-e2e-v2-bug-cluster.md.
Rebase sensitivity (medium for core/tools/vmaf.c, low for the rest): the C-side change is bolted on at the end of each backend's state_init failure stanza inside init_gpu_backends; an upstream refactor that restructures that helper (Netflix has no equivalent function — the fork extracted it as ADR-0141 §2 with NOLINTNEXTLINE) would need the explicit-backend if (...) return -1; gates re-applied per backend. The amend_json_with_backend_used helper is fork-local (operates on the file libvmaf wrote — no API change) and survives upstream syncs verbatim. The vmaf-tune fixes live entirely in tools/vmaf-tune/ which is fork-added; no upstream-mirror file is touched. The ScoreRequest.duration_s and CorpusJob.{src_width, src_height} field additions are optional with safe defaults so older test fixtures still compile. ffmpeg-patches are unaffected — no public C API surface changed.
fix/vmaf-tune-ladder-reference-decode-v3¶
Files: tools/vmaf-tune/src/vmaftune/{corpus,score}.py, tools/vmaf-tune/tests/test_bbb_e2e_v3_bug_cluster.py, docs/adr/0499-vmaf-tune-ladder-reference-decode-v3.md, docs/adr/README.md, docs/usage/vmaf-tune.md, changelog.d/fixed/vmaf-tune-ladder-reference-decode-v3.md.
Rebase sensitivity (none): all changes live in tools/vmaf-tune/ which is fork-added — no upstream Netflix file is touched. The new _maybe_decode_reference helper and _decode_source_to_yuv shared building block are private module functions; no public API was added or renamed. Dropping .y4m from _VMAF_RAW_SUFFIXES / VMAF_RAW_SUFFIXES is a behaviour change inside the wrapper that matches what the libvmaf CLI has always done (raw_input_open rejects Y4M files when use_yuv=true — see core/tools/cli_parse.c and the regression test test_vmaf_raw_suffixes_matches_libvmaf_cli_source which cross-checks the table against the CLI source). ffmpeg-patches unaffected. No effect on bisect.py (already decodes the reference per ADR-0498) — the regression test test_bisect_decodes_reference_too pins the existing invariant.
perf/vif-lut-shrink-and-filter-cache (ADR-0500)¶
Touches upstream-mirrored files: core/src/feature/integer_vif.h, integer_vif.c, vif.c, vif.h, float_vif.c, x86/vif_avx512.c.
Rebase note: when pulling upstream changes to any of these files, verify that:
VifPublicState.log2_table(now 32768 entries) is not reverted to 65537 entries.- The
compute_vifsignature addition (precomputed_filters,precomputed_filter_widths) does not conflict with upstream signature changes. - The three
_mm512_i32gather_epi64gather sites invif_avx512.cretain the_mm256_and_si256index mask withVIF_LOG2_TABLE_SIZE - 1. log_generatefills indices[0..32767]withlog2f(32768+i)*2048(not the originallog2f(i)*2048foriin[32767..65535]).
If upstream changes any of the above, a new reconciliation pass is needed.
fix/bbb-e2e-v4-bug-cluster-2026-05-18 (ADR-0501)¶
Files: tools/vmaf-tune/src/vmaftune/{corpus,ladder,cli}.py, tools/vmaf-tune/tests/{test_bbb_e2e_v2_bug_cluster,test_bbb_e2e_v4_bug_cluster}.py, docs/adr/0501-vmaf-tune-bbb-e2e-v4-bug-cluster.md, docs/adr/README.md, docs/usage/vmaf-tune.md, changelog.d/fixed/0501-vmaf-tune-bbb-e2e-v4-bug-cluster.md.
Rebase sensitivity (none): all changes live in tools/vmaf-tune/ (fork-added) and fork-added docs/changelog files — no upstream Netflix file is touched. The corpus / ladder / cli modules are not present in Netflix master. The optional target_width / target_height kwargs added to _decode_source_to_yuv and _maybe_decode_reference default to None so older test fixtures and the ADR-0499 single-resolution code path round-trip unchanged. The samples= kwarg added to emit_manifest / _emit_json is keyword-only with a None default; HLS / DASH emitters silently ignore it. _run_report's stdout JSON grows two new fields (degraded, codec_rows_unavailable) — a strict schema consumer that asserts on absence would need updating, but the existing fields stay populated. ffmpeg-patches are unaffected; no public C API surface changed.
perf/adm-decouple-gather-locality-2026-05-18 (ADR-0502)¶
Files touched: core/src/feature/x86/adm_avx512.c (upstream-mirror), core/src/feature/x86/AGENTS.md, docs/adr/0502-adm-decouple-gather-prefetch.md, docs/adr/README.md, docs/research/0435-adm-decouple-gather-locality.md, changelog.d/performance/adm-decouple-gather-prefetch.md.
Rebase sensitivity (low): the only upstream-shared file touched is adm_avx512.c. The change is a self-contained block (16 lines, guarded by if (j + 32 < right_mod16)) inserted before the three vpgatherdd lines. Conflicts arise only if Netflix upstream modifies adm_decouple_avx512 — resolution: apply the prefetch block to the updated gather cluster in the upstream version. The adm_div_lookup LUT signature is unchanged; the adm_div_lookup[val + 32768] access pattern is identical to scalar. No public C-API surface, no header, no meson build files touched.
ADR-0503: vif_subsample_rd_8_avx512 loop fission (2026-05-18)¶
File: core/src/feature/x86/vif_avx512.c
If Netflix upstream modifies vif_subsample_rd_8_avx512, the two noinline helpers (vif_subsample_rd_8_vert_j, vif_subsample_rd_8_horiz_j) and their parameter structs (VifVertCoeffs8, VifHorizCoeffs8) must be kept in sync with any upstream changes to the accumulation order or filter constant initialisation. The struct fields map 1:1 to the local variables in the original monolithic body (f0-f4, mask2/3/x for vertical; fcoeff-fcoeff4, addnum, mask1 for horizontal). No public C-API surface, no header, no meson build files touched.
fix/bbb-e2e-v5-bug-cluster-2026-05-18 (ADR-0505)¶
Files: tools/vmaf-tune/src/vmaftune/{corpus,ladder,cli}.py, tools/vmaf-tune/tests/{test_bbb_e2e_v4_bug_cluster,test_bbb_e2e_v5_bug_cluster}.py, docs/adr/0505-vmaf-tune-bbb-e2e-v5-bug-cluster.md, docs/adr/README.md, docs/state.md, docs/rebase-notes.md, changelog.d/fixed/0505-vmaf-tune-bbb-e2e-v5-bug-cluster.md.
Rebase sensitivity (none): all changes live in fork-added tools/vmaf-tune/ plus fork-added docs/changelog files — no upstream Netflix file is touched. The new source_is_container wiring in corpus.iter_rows consumes an existing EncodeRequest field (added in earlier fork work); container-detection logic is a suffix-set membership test against the fork-local _VMAF_RAW_SUFFIXES. The new cloud_sink kwarg on make_default_sampler and _default_sampler defaults to None so every existing caller round-trips unchanged.
fix/ladder-duration-clip-ffmpeg-t-flag (ADR-0508)¶
Files: tools/vmaf-tune/src/vmaftune/encode.py, tools/vmaf-tune/tests/test_bbb_e2e_v8_bug_cluster.py, docs/adr/0508-vmaf-tune-ladder-pass1-stats-duration-clip.md, docs/adr/README.md, docs/state.md, docs/rebase-notes.md, changelog.d/fixed/0508-vmaf-tune-ladder-pass1-stats-duration-clip.md.
Rebase sensitivity (none): all changes live in fork-added tools/vmaf-tune/ plus fork-added docs/changelog files — no upstream Netflix file is touched. The fix adds a six-line fallback to build_pass1_stats_command that reads the existing EncodeRequest.duration_s field (introduced by ADR-0506 V6-1) and emits an input-side -t duration_s when the caller did not opt into sample-clip mode. Sample-clip precedence is preserved so existing tests that pin the sample-clip argv shape continue to pass unchanged. No public surface of the libvmaf C API changes; no ffmpeg-patches file consumes tools/vmaf-tune/ Python helpers.
fix/bbb-e2e-v6-bug-cluster-2026-05-18 (ADR-0506)¶
Files: tools/vmaf-tune/src/vmaftune/{corpus,encode,cli}.py, tools/vmaf-tune/tests/test_bbb_e2e_v6_bug_cluster.py, docs/adr/0506-vmaf-tune-bbb-e2e-v6-bug-cluster.md, docs/adr/README.md, docs/state.md, docs/rebase-notes.md, changelog.d/fixed/0506-vmaf-tune-bbb-e2e-v6-bug-cluster.md.
Rebase sensitivity (none): all changes live in fork-added tools/vmaf-tune/ plus fork-added docs/changelog files — no upstream Netflix file is touched. EncodeRequest gains a new duration_s: float = 0.0 field with a back-compatible default so every existing caller round-trips unchanged. _decode_source_to_yuv gains four new kwargs (source_is_raw, source_width, source_height, source_framerate) all defaulting to None/False; the container-source path (which is what every v3/v4/v5 test exercises) takes the legacy branch and emits an identical argv. _maybe_decode_reference and iter_rows wire the new kwargs through. No public surface of the libvmaf C API changes; no ffmpeg-patches file consumes the modified tools/vmaf-tune/ Python helpers.
fix/mcp-backend-probe-allowlist-ladder-score-backend (ADR-0511)¶
Files: mcp-server/vmaf-mcp/src/vmaf_mcp/server.py, mcp-server/vmaf-mcp/tests/test_backend_probe_and_allowlist_0509.py, tools/vmaf-tune/src/vmaftune/{cli,ladder}.py, tools/vmaf-tune/tests/test_ladder_score_backend_0509.py, docs/adr/0511-mcp-backend-probe-allowlist-and-ladder-backend.md, docs/adr/README.md, docs/state.md, docs/mcp/backends.md, docs/usage/vmaf-tune.md, docs/rebase-notes.md, mcp-server/AGENTS.md, tools/vmaf-tune/AGENTS.md, changelog.d/fixed/mcp-and-ladder-backend.md.
Rebase sensitivity (none): every touched file is fork-local — the MCP server (mcp-server/) and the vmaf-tune CLI (tools/vmaf-tune/) are wholly fork-added trees with no upstream counterpart. The libvmaf C surface is untouched (the probe shells out to the existing vmaf --help flag table, no new CLI surface added there), and no ffmpeg-patches/ patch consumes tools/vmaf-tune/ or mcp-server/ Python helpers. make_default_sampler and _default_sampler gain a new score_backend: str | None = None kwarg with a back-compatible default, so every existing caller round-trips unchanged. _run_tune_per_shot is deliberately not touched — the auto → None → libvmaf-picks predicate contract is preserved as documented inline.
fix/windows-mingw64-build-repair (ADR-0515)¶
Rebase sensitivity (none): test-only change confined to a fork-added file (core/test/test_public_api_score.c, added 2026-05-16 from the C-API coverage audit). The Win32 #ifdef branch mirrors the pre-existing pattern in core/test/dnn/test_model_loader.c::test_sidecar_parses; no public C-API or ffmpeg-patches/ surface touched. No upstream-Netflix counterpart for this test file, so upstream rebases cannot collide with the helper.
fix/tiny-model-loader-external-data-and-feature-rank (ADR-0518)¶
Files: core/src/dnn/{model_loader.h,model_loader.c,dnn_ctx.h,ort_backend.h,ort_backend.c,AGENTS.md}, core/src/libvmaf.c, core/test/dnn/{meson.build,test_cli.sh,test_model_loader.c}, docs/adr/0518-tiny-model-loader-external-data-and-feature-rank.md, docs/adr/README.md, docs/ai/inference.md, docs/state.md, docs/research/0518-tiny-model-loader-feature-rank.md, docs/rebase-notes.md, changelog.d/fixed/tiny-model-loader.md.
Rebase sensitivity (low — fork-local dnn surface, additive): All touched files are fork-local except core/src/libvmaf.c, where the changes are confined to the tiny-AI bridge (the VmafContext::dnn struct in the file-private definitions block, vmaf_ctx_dnn_free, vmaf_ctx_dnn_attach, and vmaf_ctx_dnn_run_frame — all fork-added per ADR-0040 / ADR-0042). The struct grows by four fields (in_rank, n_features, extra_in_width, extra_in_buf); no existing offset shifts since the new fields are appended inside the dnn substruct, which is private to libvmaf.c. VmafModelSidecar (declared in core/src/dnn/model_loader.h) grows by n_features / feature_names[VMAF_DNN_MAX_FEATURE_NAMES] / feature_mean[] / feature_std[] / has_feature_scaler — this is a header consumed only by the dnn TU set and the DNN tests; consumers outside core/src/dnn/ or core/test/dnn/ should not depend on the struct layout (it's an internal sidecar contract, not a public API). vmaf_ort_input_shape_at() is a new public symbol on ort_backend.h; the existing vmaf_ort_input_shape() remains as the slot == 0 shortcut. No ffmpeg-patches file consumes any of the changed symbols.
fix/hip-import-state-implementation (ADR-0519)¶
Files: core/src/hip/common.c (delete vmaf_hip_import_state stub body — 9 lines removed), core/src/libvmaf.c (add HAVE_HIP block: header include, hip field on VmafContext, real implementation of vmaf_hip_import_state, cleanup in vmaf_close), core/test/test_hip_smoke.c (replace test_import_state_returns_enosys with test_import_state_validates_arguments + test_import_state_succeeds_with_real_state), docs/adr/0519-hip-import-state-implementation.md, docs/adr/README.md, docs/backends/hip/overview.md, docs/state.md, docs/research/0519-hip-import-state.md, docs/rebase-notes.md, changelog.d/fixed/hip-import-state.md.
Rebase sensitivity (low — fork-local HIP surface, additive): The VmafContext struct in core/src/libvmaf.c grows by one field — a hip substruct holding a single VmafHipState * pointer — gated by #ifdef HAVE_HIP. The field is appended after the existing metal substruct (the last existing GPU substruct), so no offsets shift in CPU-only / CUDA-only / etc. builds. The new vmaf_hip_import_state definition matches the existing public declaration in core/include/libvmaf/libvmaf_hip.h field-for-field; the header is unchanged, so consumers (the fork's core/tools/vmaf.c and any future ffmpeg-patches/ HIP consumer) recompile against the same ABI. core/src/hip/common.c loses the 9-line vmaf_hip_import_state stub; no other in-file functions are touched. On upstream rebase the patch is trivially applicable because Netflix/vmaf master ships no HIP backend; the entire core/src/hip/ tree is fork-local (ADR-0212). No ffmpeg-patches file consumes vmaf_hip_import_state directly today — vf_libvmaf reaches HIP only through the CPU-path fallback for now.
2026-05-18 — --tiny-codec / --tiny-preset / --tiny-crf populate codec block (ADR-0522, PR #TBD)¶
Fork-local. Adds three CLI flags + one public C-API (vmaf_dnn_set_codec_context) that override the ADR-0518 "unknown" codec pre-seed for codec-aware tiny models (fr_regressor_v2). Touched: core/include/libvmaf/dnn.h (public header — new export), core/src/dnn/dnn_attach_api.c (public symbol), core/src/dnn/dnn_ctx.h (bridge), core/src/libvmaf.c (bridge implementation — vmaf_ctx_dnn_set_codec_context), core/src/dnn/model_loader.{c,h} (sidecar encoder_vocab[] parsing + vmaf_dnn_codec_block_fill helper + VmafModelSidecar grows by n_encoder_vocab / encoder_vocab[VMAF_DNN_MAX_ENCODER_VOCAB] / codec_aware), core/tools/cli_parse.{c,h} (three new flags + three new CLISettings fields: tiny_codec, tiny_preset, tiny_crf), core/tools/vmaf.c (call site after vmaf_use_tiny_model), core/test/dnn/test_model_loader.c (8 new tests).
Rebase sensitivity (low — fork-local additive): All touched files are fork-local. VmafModelSidecar and the VmafContext::dnn substruct grow by additive fields only — no existing offset shifts. vmaf_dnn_set_codec_context() is a new VMAF_EXPORT symbol on the public libvmaf/dnn.h surface; the ffmpeg-patches stack does NOT currently consume tiny-model inference (the vf_libvmaf filter wires through the classic metric collector, not the tiny-AI surface), so no patch update is required for this PR (per CLAUDE.md §12 r14: "Does NOT apply to … kernel implementations behind an existing public surface"). The C-side codec_block_preset_ordinal table is a duplicate of train_fr_regressor_v2.py::PRESET_ORDINAL; the core/src/dnn/AGENTS.md invariant note flags both files as a co-edit pair. The sidecar encoder_vocab array is the single source of truth for the vocabulary; vocab bumps (e.g. ADR-0302 v3) only require a new sidecar JSON, no C recompile.
fix/cli-threads-parse-safety-v2 (ADR-0528)¶
Files touched: core/test/test_cli_parse_long_only_args.c, core/tools/cli_parse.c.
Rebase sensitivity (low — fork-local additive): Both files are fork-local. The test (test_cli_parse_long_only_args.c) has no upstream twin — it was added in PR #408 / ADR-0316 to lock down the long-only short-option synthesis bug. cli_parse.c is upstream- adjacent (Netflix maintains its own error() with the same assert(long_opts[n].name) shape and the same sprintf(optname, …) calls). On a future upstream sync, expect a merge conflict on error() if Netflix changes the assert / sprintf lines independently: keep the fork's if (!found) usage(…); return; + snprintf shape and drop the <assert.h> include. No public API surface changes; the ffmpeg-patches stack is untouched.
2026-05-18 — HIP integer_motion flag promotion + HIP_DEVICE buffer enum (ADR-0532, PR #TBD)¶
Extends ADR-0519. Promotes VMAF_FEATURE_EXTRACTOR_HIP on vmaf_fex_integer_motion_hip so the model-driven dispatch picks the HIP kernel instead of the CPU twin when a HIP state is imported. Adds VMAF_PICTURE_BUFFER_TYPE_HIP_DEVICE to the picture-buffer enum (reserved for the future HIP picture pool; HIP TUs still accept HOST and do their own HtoD copy). Wires compute_fex_flags() for HIP, adds a CPU-twin fallback in vmaf_get_feature_extractor_by_feature_name(), drains HIP-flagged extractors' gpu_pending final-frame collect in flush_context_serial(), and routes the HIP integer_motion collect/flush writes through feature_name_dict so the encoded option-aware key matches the predict-side lookup.
Touched: core/src/picture.h (new enum entry), core/src/feature/feature_extractor.c (dispatch buffer-type check + _by_feature_name fallback + new extern + registry row), core/src/libvmaf.c (compute_fex_flags HIP slot + flush_context_serial HIP drain), core/src/feature/hip/integer_motion_hip.c (flag bit set + dict-aware writes), core/src/feature/hip/integer_vif_hip.c (flag bit cleared with citation — un-promotes pending kernel-level fix), core/src/hip/meson.build (compile integer_motion_hip.c), core/src/meson.build (motion_score.hip HSACO), core/src/hip/AGENTS.md (invariant rewrite), core/test/test_hip_smoke.c (registration + flag-dispatch tests), docs/backends/hip/overview.md, docs/rebase-notes.md, changelog.d/added/0530-hip-integer-motion-flag-promotion.md, docs/state.md.
Rebase sensitivity (medium — touches upstream-mirror feature_extractor.c dispatch site):
The dispatch-time HIP buffer-type check is a NEW symmetric block right after the existing CUDA buffer-type check. Any upstream port that touches the CUDA block needs a paired update to the HIP block to keep them symmetric. The CPU-twin fallback pass in _by_feature_name is a documented contract going forward (ADR-0532) — future GPU backend work cannot assume "flag set ⇒ full coverage"; treat the fallback as the established behaviour, not as a bug to fix.
The compute_fex_flags() HIP slot mirrors the existing Vulkan / SYCL slots field-for-field; the flush_context_serial() HIP drain mirrors the SYCL flush_context_sycl drain. Any upstream refactor that relocates either function needs to move all three GPU slots / drains together.
vmaf_fex_integer_vif_hip had VMAF_FEATURE_EXTRACTOR_HIP set speculatively in its batch-1 commit; this PR clears it with an inline citation. Do NOT re-enable on a future rebase without a kernel-level GPU-memory-access-fault fix and an ADR-0532-style per-extractor reproducer.
No public-header change → no ffmpeg-patches/ update required (per CLAUDE.md §12 r14: the new picture-buffer-type enum lives in the libvmaf-private src/picture.h, not the public include/libvmaf/picture.h; the ffmpeg vf_libvmaf filter hands HOST buffers to libvmaf and is unaffected).
feat/hip-register-all-extractors (ADR-0533)¶
Files touched: core/src/hip/meson.build, core/src/feature/feature_extractor.c, core/test/test_hip_smoke.c, core/src/hip/AGENTS.md, docs/backends/hip/overview.md, docs/state.md, docs/adr/0533-hip-all-extractors-registration-sweep.md, docs/adr/README.md, changelog.d/fixed/hip-register-all-extractors.md.
Rebase sensitivity (low — fork-local only): All edits sit in fork-additive HIP plumbing. Upstream Netflix/vmaf ships no HIP backend, so the #if HAVE_HIP blocks in feature_extractor.c are entirely fork-local; the extern + registry entries land inside the same #if HAVE_HIP regions ADR-0523 already extended. core/src/hip/meson.build is a fork-added file (the subdir('hip') invocation is gated on enable_hip). No public C-API surface changes — vmaf_get_feature_extractor_by_name already existed; the sweep only adds rows to the table it reads. The ffmpeg-patches stack is untouched (no new LIBVMAFContext field, no new CLI flag, no new meson_options.txt entry). On a future upstream sync, expect zero conflicts — Netflix never touches HIP files. If a future PR adds a new HIP feature TU, the invariant pinned in core/src/hip/AGENTS.md (every vmaf_fex_*_hip symbol must appear in hip_sources + the extern/registry blocks) must be honoured or the registration drops out silently.
ADR-0538 — per-shot predicate bitrate sidecar (PR #1290 follow-up)¶
No rebase impact: the change is entirely internal to tools/vmaf-tune/src/vmaftune/cli.py (_build_per_shot_bisect_predicate return type change + call-site unpack + dataclasses.replace patch loop) and the corresponding test file. No public API surface, no C code, no meson_options.txt entry, no ffmpeg-patches entry, no new public Python symbol. Netflix upstream never touches vmaf-tune; upstream syncs will not conflict. On a future upstream sync, expect zero conflicts.
ADR-0539 — HIP integer_moment HSACO blob registration¶
No rebase impact: change is one new row in hip_kernel_sources inside core/src/meson.build, gated by the fork-only enable_hip flag. Netflix upstream has no HIP backend and never touches hip_kernel_sources, the four feature/hip/integer_moment/* paths, or the hip_hsaco_stubs.c TU. The ffmpeg-patches stack is untouched (no new LIBVMAFContext field, no new CLI flag, no new meson_options.txt entry). On a future upstream sync, expect zero conflicts.
If a future PR adds yet another HIP host TU that consumes a <name>_hsaco symbol distinct from any existing meson key, the invariant pinned in core/src/feature/hip/AGENTS.md (HSACO symbol naming) must be honoured to avoid the same class of link error this ADR closed.
feat/dev-container-ffmpeg-av1-hwaccel (ADR-0543)¶
Files touched: dev/Containerfile (stage 3.5 apt list + SVT-AV1 / libaom / vvenc / AMF source builds + FFmpeg configure flags + build- time encoder probe), dev/AGENTS.md (four new "FFmpeg encoder exposure invariants"), docs/development/dev-mcp.md (encoder matrix + runtime failure modes + full-sweep reproducer), core/src/meson.build (one-line follow-up to ADR-0523: add 'motion_score' to the hip_kernel_sources dict; surfaced as a stage-3 link blocker during ADR-0543's container rebuild verification), ffmpeg-patches/0007- libvmaf-tune-qpfile-unified.patch (one-line addition: #include <stdbool.h> to libavcodec/libsvtav1.c so enable_roi_map = true compiles — surfaced when the in-image FFmpeg was first built with --enable-libsvtav1; the patched code was never previously exercised because no prior dev-image enabled libsvtav1), docs/adr/0529-…md (+ index fragment), changelog.d/added/0529-…md, docs/state.md (two state rows), this file.
Rebase sensitivity (none — container-only fork-local additive plus one fork-local libvmaf wiring fix): Every touched file lives under dev/, docs/, changelog.d/, or core/src/meson.build. The libvmaf hunk adds one dict entry referencing a fork-local .hip source (ADR-0523 lineage); no upstream conflict possible because Netflix has no HIP backend. No libvmaf C source, no public header, no meson_options.txt, no ffmpeg-patches/ entry. The CLAUDE.md §12 r14 patch-stack rule does not apply — the FFmpeg configure-line change happens in the in-image build only and is orthogonal to the host-side patch series under ffmpeg-patches/. Pin bumps (VVenC v1.12.0, AMF v1.4.36, FFmpeg n8.1.1 via FFMPEG_TAG build-arg) are visible in the ARG lines of dev/Containerfile; bumping them is a local container change.
ADR-0543 — ADR-0498 enforcement hardening (exit code 100 + JSON error + per-feature gate)¶
Summary: Hardens the explicit-backend gate that ADR-0498 introduced in core/tools/vmaf.c. Adds three orthogonal contracts: dedicated exit code 100 (VMAF_EXIT_BACKEND_INIT_FAILED) for --backend NAME init failures, structured JSON error descriptor at the --output path when format is JSON, and a per-feature gate that hard-fails GPU-pinned feature names (*_cuda / *_sycl / *_vulkan / *_hip / *_metal) when the matching backend isn't active.
Files touched: core/tools/vmaf.c (new constants, new helpers write_backend_error_json / feature_backend_suffix / backend_active, new bool *cuda_active_out parameter on init_gpu_backends, per-feature gate in the feature-loading loop, simplified backend_used echo), docs/adr/0543-adr-0498-enforcement- hardening.md (+ index row in docs/adr/README.md), changelog.d/fixed/0543-adr-0498-enforcement-hardening.md, tools/vmaf-tune/tests/test_adr_0543_backend_enforcement.py (13 integration + source-level tests), docs/state.md (Recently closed row), this file.
Rebase sensitivity (none — fork-local additive against an already-fork-local helper): The only C source touched is core/tools/vmaf.c, and only inside the init_gpu_backends helper + its caller — both of which are fork-local additions that do not exist in Netflix/vmaf upstream (Netflix has no SYCL / HIP / Vulkan / Metal backends and no --backend selector). The new bool *cuda_active_out parameter on init_gpu_backends is guarded by #ifdef HAVE_CUDA and only affects the in-tree caller. No upstream conflict possible.
ffmpeg-patches/ impact (none): No public libvmaf C-API entry points added, renamed, or removed. No meson_options.txt flag added. No LIBVMAFContext field added. No vf_libvmaf.c filter variant added. The new exit code is a CLI-level contract observed by wrappers (vmaf-tune, MCP) — FFmpeg's libvmaf filter consumes libvmaf via the C API and is not impacted. CLAUDE.md §12 r14 does not apply.
fix/feature-extractor-list-dedup (ADR-0544)¶
Removes 61 duplicate &vmaf_fex_* entries from core/src/feature/feature_extractor.c's static feature_extractor_list[] (55 Vulkan + 6 SYCL) and adds vmaf_feature_extractor_list_audit(), called from vmaf_init(), that returns -EINVAL if any extractor name or pointer is seen twice.
Rebase sensitivity (low — fork-local hunks only): The duplicated rows lived in fork-local #if HAVE_VULKAN / #if HAVE_SYCL blocks — both backends are absent upstream. The deduped arrangement keeps the same row ordering Netflix would expect for the CPU + CUDA paths (untouched) so the inevitable next sync-upstream sees no diff there. The new public header line in core/src/feature/feature_extractor.h (the vmaf_feature_extractor_list_audit() declaration) is appended after the existing fork-local symbols and before the VmafFeatureExtractorContextFlags block, isolating it from upstream hunks. The vmaf_init() call site lives in a fork-local block (right after vmaf_set_log_level) that already differs from upstream because of the HIP/Vulkan plumbing — a conflict is only possible if Netflix adds a new init-time call there, in which case the resolution is trivial (preserve both calls; the audit is order-independent w.r.t. other init steps).
Touched files: core/src/feature/feature_extractor.{c,h}, core/src/libvmaf.c, core/test/test_feature_extractor.c, docs/adr/0541-*.md, docs/adr/_index_fragments/0541-*.md, docs/adr/_index_fragments/_order.txt, docs/adr/README.md (regenerated), docs/state.md, docs/rebase-notes.md, changelog.d/fixed/0541-*.md. No ffmpeg-patches/, meson_options.txt, or meson.build change (test is exercised by an existing test_feature_extractor target).
chore/wire-or-delete-dead-extractor-files (ADR-0545)¶
Deletes 18 dead Vulkan / Metal feature-extractor source files plus 14 paired orphan shaders (.comp / .metal) from core/src/feature/{vulkan,metal}/ that were never wired into their backend's meson.build. Wires one previously-unwired Metal TU (float_ms_ssim_metal.mm + float_ms_ssim.metal, ADR-0490) into core/src/metal/meson.build. Removes one dead extern in core/src/feature/feature_extractor.c (vmaf_fex_integer_adm_metal) and refreshes the core/src/feature/{vulkan,metal}/AGENTS.md rebase-sensitive invariants section to forbid re-introducing the deleted scaffolds.
Rebase sensitivity (none — pure fork-local housekeeping): All deleted files were fork-local scaffolds added in commit 302bd1673 (2026-05-18, "docs(rules): default to vmaf-dev-mcp container"). Netflix upstream has no Vulkan or Metal backend, so no upstream conflict is possible on the deletes. The lone wired file (float_ms_ssim_metal.mm) is fork-original, references no upstream identifier, and lives under fork-only core/src/metal/. The retained adm_vulkan.c legacy shim is out of scope per ADR-0468. No CPU-path C source, no public header, no meson_options.txt, no ffmpeg-patches/ entry, no Python-binding change — CLAUDE.md §12 r14 (FFmpeg patch-stack sync) does not apply.
feat/vmaf-tune-full-file-and-no-bisect (ADR-0548)¶
Files touched: tools/vmaf-tune/src/vmaftune/cli.py (auto-probe block in _run_tune_per_shot; new _run_compare_crf_sweep function; --no-bisect / --crf-sweep argparse flags in the compare subparser; --width / --height / --framerate made optional in the tune-per-shot subparser), tools/vmaf-tune/tests/test_tune_per_shot_container_src.py (new — Fix A smoke tests), tools/vmaf-tune/tests/test_compare_no_bisect.py (new — Fix B smoke tests), tools/vmaf-tune/AGENTS.md (two new invariant notes), docs/adr/0548-vmaf-tune-full-file-and-no-bisect.md (+ index row in docs/adr/README.md), changelog.d/added/0548-…md, docs/usage/vmaf-tune.md (Fix A and Fix B documentation), this file.
Rebase sensitivity (none — vmaf-tune Python only): All touched files live under tools/vmaf-tune/, docs/, or changelog.d/. No libvmaf C source, no public header, no meson_options.txt, no ffmpeg-patches/ entry is touched. The cli.py changes are purely additive: new function _run_compare_crf_sweep, new optional args in existing subparsers, and an early-return dispatch at the top of _run_compare. No existing API surface is renamed or removed. The probe block at the top of _run_tune_per_shot only executes when args.width is None or args.height is None or args.framerate is None — callers that pass explicit geometry are unaffected. No upstream Netflix/vmaf path is touched.
ADR-0539 — HIP hip_cu_extra_flags dispatch + ssimulacra2_blur -ffp-contract=off¶
No rebase impact: the change is entirely additive in core/src/meson.build inside the if get_option('enable_hipcc') block — a new hip_cu_extra_flags dict and one extra per_kernel_flags list interpolated into the existing hipcc custom_target command. The fall-through (.get(name, [])) keeps the command line byte-identical for every kernel not listed. Netflix upstream ships no HIP backend, so the entire enable_hipcc block is fork-local; upstream syncs will not conflict. The dict mechanism extends naturally — when porting a future CUDA kernel that lists flags in cuda_cu_extra_flags, mirror the entry in hip_cu_extra_flags per core/src/feature/hip/AGENTS.md. No public API surface, no meson_options.txt entry, no ffmpeg-patches entry. On a future upstream sync, expect zero conflicts.
ADR-0539 — integer ADM HIP kernels (real impl, removes ADR-0536 weak stubs)¶
No rebase impact: every touched file is fork-local — the four .hip kernel sources under core/src/feature/hip/integer_adm/ are fork-additive (Netflix ships no HIP backend), the hip_kernel_sources meson dict additions live inside the if is_hip_enabled and is_hipcc_enabled block (also fork-local), and the hip_hsaco_stubs.c weak-fallback file is wholly fork-added under ADR-0536. No public C-API surface changes — kernel symbol names match the GET_FN calls in integer_adm_hip.c exactly, host TU is untouched. No meson_options.txt flag added or renamed (re-uses enable_hip + enable_hipcc). No ffmpeg-patches entry needs an update (no new LIBVMAFContext field, no new CLI flag). On a future upstream sync, expect zero conflicts. If a future PR re-introduces a CUDA-only helper into one of the four kernels (re-breaking the standalone build), do NOT re-add a weak HSACO stub — fix the kernel (invariant pinned in core/src/feature/hip/AGENTS.md).
ADR-0598 — Codec-adapter two_pass_args real implementations¶
Rebase impact: none. The change is entirely fork-local — all modified files live under tools/vmaf-tune/src/vmaftune/codec_adapters/ (fork-added Phase A/F vmaf-tune package) and docs/. No upstream Netflix/vmaf file is touched, no ffmpeg-patches file is touched, no public core/include/ header is touched, no meson_options.txt key is added.
Touched files: tools/vmaf-tune/src/vmaftune/codec_adapters/{svtav1,libaom,vvenc,_nvenc_common,_qsv_common,_amf_common,_videotoolbox_common,h264_videotoolbox,hevc_videotoolbox,av1_videotoolbox,prores_videotoolbox}.py, tools/vmaf-tune/tests/test_codec_adapter_two_pass_real.py, docs/adr/0546-codec-adapter-two-pass-real.md, docs/adr/README.md (one index row), docs/research/0546-codec-adapter-two-pass-real.md, docs/usage/vmaf-tune.md (codec support matrix refresh), docs/state.md (Recently-closed row), changelog.d/added/0546-codec-adapter-two-pass-real.md, docs/rebase-notes.md (this entry).
chore/ai-tooling-env-overrides-split (ADR-0547)¶
Rebase sensitivity (low — fork-local only):
ai/scripts/*.py: every file is fork-local. The edits add anos.environ.get(...)wrap around the existing default-path literal. No upstream conflict possible..gitignore: appends*.bak/*.orig. Trivial to resolve if upstream ever touches the same lines (unlikely — these are universal editor-backup patterns).docs/ai/scripts-env-vars.md(new file),mkdocs.ymlnav entry: fork-local docs tree. No upstream conflict possible.tools/vmaf-tune/src/vmaftune/cli.py.bak: untracked file deletion; no git history impact.
Touched files: .gitignore, fifteen ai/scripts/*.py files, docs/ai/scripts-env-vars.md (new), mkdocs.yml, docs/adr/0547-ai-script-env-vars.md (new), docs/adr/README.md, docs/state.md, docs/rebase-notes.md, changelog.d/changed/0547-ai-script-env-vars.md (new). No ffmpeg-patches/ change (no C-API, CLI flag, or meson_options.txt consumed by a patch — CLAUDE.md §12 r14 exempt).
chore/audit-cleanup-bundle-2 (ADR-0549)¶
No rebase impact. All changes are confined to fork-local files:
core/src/feature/cuda/integer_{ssim,ms_ssim,psnr,moment}_cuda.c(comment addition only — no functional change; upstream parity intact).core/src/feature/sycl/integer_{ssim,ms_ssim,psnr,moment}_sycl.cpp(comment addition only).dev/Containerfile(fork-local; whole file is fork-added)..gitignore(adds.claude/worktrees/line; no upstream conflict).docs/state.md(fork-only doc tree).python/test/vmafexec_test.py(comment deletion only; assertion value andplacesargument are unchanged — no golden-gate impact).docs/adr/0549-audit-cleanup-bundle-2.md,docs/adr/README.md,changelog.d/changed/0549-audit-cleanup-bundle-2.md,docs/rebase-notes.md(this entry).
Touched files: core/src/feature/cuda/integer_{ssim,ms_ssim,psnr,moment}_cuda.c, core/src/feature/sycl/integer_{ssim,ms_ssim,psnr,moment}_sycl.cpp, dev/Containerfile, .gitignore, docs/state.md, python/test/vmafexec_test.py, docs/adr/0549-audit-cleanup-bundle-2.md, docs/adr/README.md (one index row), changelog.d/changed/0549-audit-cleanup-bundle-2.md, docs/rebase-notes.md (this entry).
ADR-0598 — audit bundle (Vulkan-01 / saliency-tune-01 / ai-01)¶
No rebase impact for Vulkan-01: adding &vmaf_fex_integer_motion_vulkan_impl to feature_extractor_list[] in feature_extractor.c is a purely additive change under the existing #if HAVE_VULKAN guard. Netflix has no Vulkan backend, so upstream syncs produce zero conflicts.
No rebase impact for saliency-tune-01: all touched files (tools/vmaf-tune/src/vmaftune/saliency.py, tools/vmaf-tune/src/vmaftune/cli.py) are fork-local. No public C-API, no meson_options.txt entry, no ffmpeg-patches entry.
No rebase impact for ai-01: tools/vmaf-tune/src/vmaftune/predictor_train.py is fork-local. The --emit-stub-card-only flag is additive.
ADR-0550 -- Cross-backend parity matrix 2026-05-18¶
Touches: docs/adr/0550-cross-backend-parity-matrix-2026-05-18.md, docs/research/0550-cross-backend-parity-matrix-2026-05-18.md, docs/adr/README.md (index row), docs/state.md (audit-closed row), changelog.d/added/0550-cross-backend-parity-matrix.md, this file.
Rebase sensitivity (none -- docs-only, no C source): All touched files are under docs/ and changelog.d/. No libvmaf C source, no public header, no meson_options.txt, no ffmpeg-patches/ entry. No numerical-correctness risk: this is a read-only audit that produces only documentation artefacts.
ADR-0561 — HIP gfx_targets fallback widening¶
Branch: fix/hip-gfx-targets-fallback-widening Rebase impact: low — touches only core/src/meson.build, core/meson_options.txt, and docs/backends/hip/overview.md. No kernel code, no public API change.
Rebase-sensitive invariant: The fallback string 'gfx90a,gfx1030,gfx1036,gfx1100' at the end of the four-step probe chain in core/src/meson.build must not regress to 'gfx90a' alone. If a meson.build rebase conflict arises in that region, prefer the wider fallback. The comment block above the fallback explains the rationale.
Touched files: core/src/meson.build (fallback string + comment), core/meson_options.txt (description update), docs/backends/hip/overview.md (-Dhip_gfx_targets section), docs/adr/0561-hip-gfx-targets-fallback-widening.md, docs/adr/README.md (one index row), docs/research/0561-hip-gfx-targets-fallback-widening.md, docs/state.md (Recently-closed row T-HIP-GFX-TARGETS-FALLBACK-2026-05-18), changelog.d/fixed/0561-hip-gfx-targets-fallback-widening.md, docs/rebase-notes.md (this entry).
ADR-0562 — VCQ-223 local-explainer hang fix¶
Rebase impact: none. The change is entirely fork-local — all modified files live under python/ (Python wrapper test harness) and docs/. No upstream Netflix/vmaf C file is touched, no ffmpeg-patches file is touched, no public core/include/ header is touched, no meson_options.txt key is added.
Touched files: python/vmaf/core/quality_runner_extra.py, python/test/local_explainer_test.py, docs/adr/0562-local-explainer-hang-fix.md, docs/adr/README.md (one index row), docs/state.md (Recently-closed row), changelog.d/fixed/vcq-223-local-explainer-hang.md, docs/rebase-notes.md (this entry).
ADR-0559 — Feature coverage audit: speed_chroma + speed_temporal in extraction scripts¶
Rebase impact: minimal. Changes are entirely fork-local — all modified files live under ai/data/feature_extractor.py, ai/scripts/bvi_dvc_to_full_features.py, ai/scripts/extract_full_features.py (docstring only), and docs/. No upstream Netflix/vmaf file is touched, no ffmpeg-patches file is touched, no public core/include/ header is touched, no meson_options.txt key is added.
If a future upstream sync adds speed_chroma or speed_temporal to the upstream FULL_FEATURES equivalent, this fork's tuple will have them already; check for duplicates on merge.
Touched files: ai/data/feature_extractor.py (FULL_FEATURES + _METRIC_TO_EXTRACTOR), ai/scripts/bvi_dvc_to_full_features.py (local FULL_FEATURES + EXTRACTORS), ai/scripts/extract_full_features.py (docstring only), docs/ai/models/konvid_mos_head_v1.md (coverage-gap note), docs/adr/0559-feature-coverage-audit.md, docs/adr/README.md (one index row), docs/research/feature-coverage-audit-2026-05-18.md, changelog.d/added/0559-feature-coverage-audit-speed-features.md, docs/rebase-notes.md (this entry).
ADR-0566 — HIP VIF per-feature places=4 gate (supersedes ADR-0537 follow-up)¶
Branch: fix/hip-vif-svm-amplification-places4-gate Rebase impact: documentation-only — touches only docs/adr/, docs/state.md, docs/rebase-notes.md, and changelog.d/. No kernel code, no meson.build change, no public API change.
Rebase-sensitive invariant: If ADR-0537 is amended in a future PR, ensure the "places=3 is acceptable" follow-up clause is not reintroduced. The supersession is recorded in ADR-0566 and in the Recently-closed row T-HIP-VIF-PLACES3-GATE-INCORRECT-2026-05-18.
Touched files: docs/adr/0566-hip-vif-per-feature-places4-gate.md, docs/adr/README.md (one index row), docs/state.md (Recently-closed row T-HIP-VIF-PLACES3-GATE-INCORRECT-2026-05-18), changelog.d/fixed/0566-hip-vif-per-feature-places4-gate.md, docs/rebase-notes.md (this entry).
ADR-0552 — HIP VIF deterministic wavefront reduction¶
Branch: fix/hip-vif-deterministic-reduce Rebase impact: low — only touches core/src/feature/hip/integer_vif/vif_statistics.hip and documentation files. No public API change. No meson.build change.
Rebase-sensitive invariant: The wavefront_reduce_i64 helper uses __shfl_xor with strides 32, 16, 8, 4, 2, 1 (for AMD 64-lane wavefronts). Do NOT merge with a CUDA-style __shfl_down_sync port that uses strides 16, 8, 4, 2, 1 (32-lane) — the stride list is wrong for AMD and will under-reduce, leaving 32-thread partial sums in the accumulator.
Conflict scenario: If a rebase brings in a change to vif_statistics.hip from the CUDA parity sweep or a vif_hori_16_body template refactor, verify that:
- The outer
if (x < w && y < h)guard is preserved (not replaced by early return). wavefront_reduce_accums(thr)is called before theatomicAddblock.- The
atomicAddblock is insideif ((threadIdx.x % AMD_WAVEFRONT_SIZE) == 0).
Touched files: core/src/feature/hip/integer_vif/vif_statistics.hip, docs/adr/0552-hip-integer-vif-deterministic-reduce.md, docs/adr/README.md (one index row), docs/research/0552-hip-vif-deterministic-reduce.md, docs/state.md (Recently-closed row T-HIP-VIF-PARITY-PLACES4-2026-05-18), changelog.d/fixed/0552-hip-vif-deterministic-reduce.md, docs/rebase-notes.md (this entry).
fix/python-mcp-ai-audit-p0-p1-2026-05-18 (ADR-0556)¶
Python / MCP / AI silent-fallback audit. No upstream-shared paths modified; all fixes are in fork-local Python harness files (tools/vmaf-tune/, mcp-server/, ai/scripts/) or documentation. No rebase-sensitive invariants introduced — the score.py JSONDecodeError guard is a pure additive safety wrapper around an existing json.load call, and the bvi_dvc_to_full_features.py empty-entries guards are early-exits before any loop body runs. Files touched: tools/vmaf-tune/src/vmaftune/score.py, tools/vmaf-tune/src/vmaftune/auto.py, mcp-server/vmaf-mcp/src/vmaf_mcp/server.py, ai/scripts/bvi_dvc_to_full_features.py, ai/scripts/validate_model_registry.py, docs/adr/0556-python-mcp-ai-audit-2026-05-18.md, docs/adr/README.md (one index row), docs/research/python-mcp-ai-audit-2026-05-18.md, changelog.d/fixed/0556-python-mcp-ai-audit.md, docs/state.md (5 T-rows added to Open section),
ADR-0568 — upstream port USE_DIRECT_READ zero-copy input path (Netflix/vmaf@30a6e2a8d)¶
Branch: chore/upstream-port-direct-read-and-speed-wrappers Rebase impact: low — all touched files are upstream-shared paths that the upstream commit also modifies. On the next sync-upstream, these changes should merge cleanly because the fork's version is a strict superset of the upstream diff (same logic, plus fork style conventions: (void)fprintf, explicit (int) casts on fread returns, memcmp(…) != 0).
Rebase-sensitive invariant: The new fetch_into_vmaf_picture vtable field is the last member of video_input_vtbl. Any future vtable extension must append after it or update both the vtable struct and all initialisers (YUV_INPUT_VTBL, Y4M_INPUT_VTBL) simultaneously. The upstream port order is: open_raw, open, get_info, fetch_frame, close, fetch_into_vmaf_picture.
Touched files: core/tools/vidinput.h, core/tools/vidinput.c, core/tools/yuv_input.c, core/tools/y4m_input.c, core/tools/vmaf.c, docs/adr/0567-upstream-port-direct-read.md, docs/adr/README.md (one index row), changelog.d/perf/0567-upstream-port-direct-read.md,
ADR-0568 — sycl_icpx_aot_targets default¶
Rebase impact: low. Adds a new sycl_icpx_aot_targets string option to core/meson_options.txt and wires the corresponding AOT flags in core/src/meson.build. Any upstream Netflix/vmaf change that also touches those two files will produce a trivial two-hunk conflict. Resolution: preserve both the upstream hunk and the fork-added sycl_icpx_aot_targets option + toolchain-flag block. No public API header is touched; no ffmpeg-patches file is touched.
Touched files: core/meson_options.txt (new option), core/src/meson.build (AOT flag wiring, icpx branch), docs/backends/sycl/overview.md (new AOT section), dev/Containerfile (doc comment near meson invocation), docs/adr/0568-sycl-icpx-aot-targets-default.md (new ADR), docs/adr/README.md (one index row), changelog.d/added/0568-sycl-icpx-aot-targets-default.md, docs/state.md (no-bug note),
chore/sdk-version-bumps-may-2026 (ADR-0569)¶
No rebase-sensitive invariants. All changes are version-string edits in dev/Containerfile ARG lines, .pre-commit-config.yaml rev fields, .github/workflows/supply-chain.yml action SHA pins, and python/requirements.txt ceiling. No API changes, no C/Python logic changes.
Conflict scenarios:
dev/Containerfile: The Ubuntu 26.04 base-image PR (in-flight) touches different ARG blocks. If a rebase conflict occurs, keep both sets of version edits — they are in disjoint sections of the file..pre-commit-config.yaml: If a concurrent PR bumps the same tools, prefer the higher version.python/requirements.txt: If the Ubuntu 26.04 PR widens thenumpyceiling in the same PR, both edits are independent; apply both.
Touched files: dev/Containerfile, .pre-commit-config.yaml, .github/workflows/supply-chain.yml, python/requirements.txt, ai/pyproject.toml (comment only), docs/adr/0569-sdk-version-bumps-2026-05-18.md, docs/adr/README.md (one index row), changelog.d/changed/0569-sdk-version-bumps-2026-05-18.md, dev/AGENTS.md (invariant notes), docs/development/dev-mcp.md (version table),
ADR-0574 — CUDA twins for HDR-model aim and adm3 sub-features (Phase 1)¶
Branch: feat/hdr-features-cuda-twins Rebase impact: CUDA kernel and host files only — touches core/src/feature/cuda/float_adm/float_adm_score.cu and core/src/feature/cuda/float_adm_cuda.c. No meson.build change, no public C-API change, no Python/model change.
Rebase-sensitive invariant: FADM_ACCUM_SLOTS = 9 must remain identical in both files. The .cu unit defines the per-WG slot layout ([0..2]=csf_den, [3..5]=cm_num, [6..8]=aim_cm); the .c host uses it for buffer allocation, D2H copy size, and accumulator reads. If a rebase replaces the .cu with a pre-ADR-0574 version (FADM_ACCUM_SLOTS = 6), update float_adm_cuda.c to match in the same commit — a mismatch silently corrupts host memory. The --fmad=false nvcc flag covers all six kernels; do not remove it.
Touched files: core/src/feature/cuda/float_adm/float_adm_score.cu, core/src/feature/cuda/float_adm_cuda.c, docs/adr/0574-hdr-features-cuda-twins-phase-1.md, docs/adr/README.md (one index row), core/src/feature/cuda/AGENTS.md (slot-sync invariant note), docs/research/netflix-upstream-feature-additions-since-sync-2026-05-18.md, docs/metrics/features.md (aim/adm3 sub-feature docs + footnote ⁶), docs/state.md (Recently-closed row T-CUDA-AIM-ADM3-2026-05-18), changelog.d/added/0574-hdr-features-cuda-twins-aim-adm3.md, docs/rebase-notes.md (this entry).
ADR-0606 — macOS SIGSEGV deep-fix in output.c writers (PR #1403 follow-up)¶
Rebase-sensitive invariant: i >= fc->feature_vector[j]->capacity (not >) in all seven frame-iteration bounds checks in core/src/output.c. If upstream Netflix ever backports a fix to the same comparison sites and uses > (the old, buggy form), the rebase must preserve the >= — the > form is a heap buffer overread UB that surfaces on macOS with MALLOC_PERTURB_=198.
The fps computation guard (if (vmaf->pic_cnt == 0 || timer_elapsed == 0)) in core/src/libvmaf.c is similarly rebase-sensitive: if upstream modifies the fps block, preserve the guard before dividing so import-only callers (those that use vmaf_import_feature_score without vmaf_read_pictures) do not produce 0.0/0.0 which may SIGFPE on Apple platforms.
Touched files: core/src/output.c (7 bounds-check sites, json pool-score + frames comma fixes), core/src/libvmaf.c (fps defensive computation), docs/adr/0606-macos-vmaf-write-output-segv-deep-fix.md, docs/adr/README.md (one index row), docs/state.md (Recently-closed row), changelog.d/fixed/0606-macos-vmaf-write-output-segv-deep-fix.md, docs/rebase-notes.md (this entry).
ADR-0612 — vmaf-tune compare: decode reference YUV once (shared-ref fix)¶
ADR-0607 — vmaf-tune compare: decode reference YUV once (shared-ref fix)¶
No rebase impact: all touched files are fork-local Python harness files. No upstream C sources, no public headers, no FFmpeg patch series involved.
Touched files: tools/vmaf-tune/src/vmaftune/compare.py (pre_decoded_ref param on compare_codecs and compare_codecs_sweep), tools/vmaf-tune/src/vmaftune/cli.py (decode-once block + try/finally in _run_compare; imports _decode_to_raw_yuv from .score), tools/vmaf-tune/tests/test_bbb_e2e_v15_shared_ref.py (7 acceptance tests), docs/adr/0607-vmaftune-shared-ref-yuv-decode-once.md, docs/adr/README.md (one index row), changelog.d/fixed/0607-vmaftune-shared-ref-yuv-decode-once.md,
ADR-0612 — Tiny-AI Netflix corpus training scaffold (2026-05-19 iteration)¶
No rebase-sensitive invariants introduced by this PR — all changes are documentation, research digest, and CHANGELOG fragment. No C/CUDA/SIMD paths modified; no loader or test code changed.
The one invariant worth noting for future rebases: the .workingdir2/netflix/ corpus path is local-only and gitignored. If a future rebase touches .gitignore, confirm that the *.yuv and .workingdir2/ entries remain in place. Training scripts must continue to accept --data-root as an explicit CLI flag rather than hard-coding the path.
Touched files: docs/adr/0612-tiny-ai-netflix-training-scaffold-2026-05-19.md, docs/adr/_index_fragments/0612-tiny-ai-netflix-training-scaffold-2026-05-19.md, docs/adr/_index_fragments/_order.txt, docs/research/0612-tiny-ai-netflix-training-scaffold-2026-05-19.md, docs/ai/training-data.md (cross-reference links), changelog.d/added/0612-tiny-ai-netflix-training-scaffold-2026-05-19.md, docs/rebase-notes.md (this entry).
ADR-0626 — SSH debug session on macOS CI failure (tmate)¶
No rebase-sensitive invariants — the change is limited to .github/workflows/libvmaf-build-matrix.yml (one new step) and docs/. No C, CUDA, SYCL, HIP, or Python paths touched.
If an upstream sync touches libvmaf-build-matrix.yml, confirm that the SSH debug session on test failure step is preserved after the merge and that its if: condition still references runner.os == 'macOS', failure(), and github.event_name == 'workflow_dispatch'. The action SHA pin (c0afd6f790e3a5564914980036ebf83216678101) will be bumped automatically by Renovate when a new mxschmitt/action-tmate release is tagged.
Touched files: .github/workflows/libvmaf-build-matrix.yml (one new step after "Run tests"), docs/development/ci-tmate-debug.md (new operator guide), docs/adr/0626-macos-ci-tmate-debug-on-failure.md, docs/adr/README.md (one index row), changelog.d/added/0626-macos-ci-tmate-debug-on-failure.md, docs/rebase-notes.md (this entry).
ADR-0628 — Remote-aware ADR number allocator¶
No rebase impact: all touched files are fork-local tooling and CI configuration. No upstream C sources, public headers, or FFmpeg patch series involved.
scripts/adr/next-free.sh is a fork-added script with no upstream analogue; it will never conflict on a Netflix upstream sync. The .github/workflows/ rule-enforcement.yml change adds a new step to an existing job — this file does not exist upstream, so no conflict is expected. The CLAUDE.md update extends §12 r8 prose only.
Touched files: scripts/adr/next-free.sh (remote-aware allocator + .git/adr-claims/ side-pointer), scripts/adr/tests/test-next-free-remote-aware.sh (new acceptance tests), .github/workflows/rule-enforcement.yml (phase-2 open-PR collision check), CLAUDE.md (§12 r8 extended allocator description), docs/adr/0628-adr-allocator-remote-aware.md, docs/adr/README.md (one index row), changelog.d/fixed/0628-adr-allocator-remote-aware.md,
ADR-0608 — MCP P0 fixes: isError, probe_backend, vmaf_version, vmaf_score_encoded¶
No rebase impact: all touched files are fork-local MCP server Python files and docs. No upstream C sources, no public headers, no FFmpeg patch series involved.
Touched files: mcp-server/vmaf-mcp/src/vmaf_mcp/server.py (isError fix in _call_tool; new _probe_backend, _vmaf_version, _run_vmaf_score_encoded, _ffprobe_geometry, _decode_to_yuv functions; three new Tool registrations), mcp-server/vmaf-mcp/tests/test_mcp_p0_adr0608.py (11 new regression tests), mcp-server/vmaf-mcp/tests/test_smoke_e2e.py (2 test updates for new behavior), mcp-server/vmaf-mcp/README.md (tools table updated to 10 tools, 6 backends), docs/mcp/tools.md (new tool sections, corrected list_backends description and response body, updated error conventions table), docs/adr/0608-mcp-p0-iserror-and-probe-version-encoded.md, docs/adr/README.md (one index row), changelog.d/fixed/adr0608-mcp-p0-iserror-probe-version-encoded.md, changelog.d/added/adr0608-mcp-probe-backend-vmaf-version-encoded.md, docs/rebase-notes.md (this entry).
Master CI repair — DNN coverage, MCP smoke, and formatter drift (2026-05-19)¶
No rebase-sensitive invariants introduced — this PR is CI/test/doc hygiene only. No public C API, backend implementation, model artifact, FFmpeg patch, or Netflix golden-data assertion changed.
If a future rebase touches the DNN tiny-model smoke tests, preserve these contracts:
core/test/dnn/test_cli.shuses--tiny-resize bilinearfor thenr_metric_v1.onnxno-reference smoke because the shipped NR model is 224x224 and strict resize mode intentionally rejects the 576x324 fixture.- The same CLI smoke caps tiny-model inference at
--frame_cnt 1; it is a load/run smoke, not a full-clip numerical benchmark. mcp-server/vmaf-mcp/tests/test_smoke_e2e.pyexpects unknown tool names to raise, matching the ADR-0299 isError contract.- The Netflix golden and lavapipe parity gates keep per-command
timeoutwrappers plus step-leveltimeout-minutes; if a runner/backend hangs, CI must fail diagnostically instead of waiting for the full job-level timeout. - The Netflix golden CI lane intentionally invokes only
QualityRunnerTest::test_run_vmaf_runnerandQualityRunnerTest::test_run_vmaf_runner_checkerboard: together they cover the D24 normal pair plus the checkerboard 10-px and 1-px distorted pairs. Do not expand this lane into the broad Python quality/feature suites; those are separate test coverage, not the golden-data gate. - Keep the D24 normal and checkerboard invocations as separate workflow steps; this preserves a clear failure surface when the 1080p checkerboard pair is slow or stuck. The normal-pair step still needs a multi-minute budget on cold GitHub-hosted runners because the Python runner invokes the feature binary with ADM/VIF/motion debug output. The normal-pair budget is 21 minutes inside a 22-minute step; lowering it back to 7 or 11 minutes has timed out on cold 2026-05-19 GitHub-hosted runners before assertions completed.
- The lavapipe Vulkan VIF cross-backend lane has the same cold-runner constraint. Keep the required VIF step at a 15-minute command wrapper inside a 16-minute step; the old 8-minute wrapper timed out exactly on GitHub-hosted Ubuntu before the diff script could report a result.
python/vmaf/routine.py::run_test_on_dataset()only passes bootstrap stats kwargs when the runner exposes the full bootstrap score-key getter set. Normal VMAF and PSNR runners do not haveget_bagging_score_key()/ CI95 / all-model prediction fields; macOS tox exercises those normal runners throughrun_testing.py, so do not reintroduce unconditional bootstrap-key access.- Python doctests under
python/vmaf/tools/must not rely on platform/version scalar reprs or assertion traceback formatting. NumPy 2 can display scalar values asnp.float64(...), and Python 3.14 can append assert-expression details; keep examples explicit withfloat(...), string formatting, or first-line exception-message printing. core/src/thread_locale.cusesduplocale(LC_GLOBAL_LOCALE)as the base fornewlocale(LC_NUMERIC_MASK, "C", base). Do not restorenewlocale(LC_ALL_MASK, "C", NULL): macOS allocator poisoning can expose poisoned internal locale pointers as SIGSEGV in the output-writer tests.core/test/test_output.cmust not includelibvmaf.coroutput.cdirectly while also linking libvmaf. Usecore/src/libvmaf_priv.h::vmaf_feature_collector_get()for the internal collector access instead; duplicate implementation TUs have crashed Apple ld64 + LTO macOS jobs under allocator poisoning.core/src/dnn/model_loader.c::vmaf_dnn_sidecar_load()rejects oversized sidecars withstat()before opening them. Preserve this cheap metadata-only guard so thetest_vmaf_use_tiny_modeloversized-sidecar case does not enter platform stdio on the expected-EFBIGpath. The regression test copiesmodel/tiny/smoke_v0.onnxrather than synthesising an invalid ONNX blob; that keeps a missed sidecar gate as an assertion failure instead of an ORT invalid-model crash on macOS.core/src/output.cflushes each writer's stream before callingvmaf_thread_locale_pop(). Keep the flush inside the locale lifetime: path-basedvmaf_write_output()usesfdopen()and may otherwise defer the stream flush tofclose()after the temporary C numeric locale has been restored/freed, which is the macOS-only writer SIGSEGV shape.core/test/meson.builddefineslibvmaf_public_linkso public ABI tests link the shared library whendefault_library=both. Do not routetest_public_api_scoreortest_vmaf_use_tiny_modelback throughlibvmaf.get_static_lib()on macOS: Apple ld64 + LTO folds the public call into the test executable and reproduces the writer/DNN SIGSEGV shape. The internaltest_outputtarget keeps its static link forvmaf_feature_collector_get()but disables LTO at the target on Darwin only; Linux clang static-archive links still need-fltobecausesrc/libvmaf.acontains LLVM bitcode there.python/vmaf/core/asset.py::ORDERED_FILTER_LISTincludesfpsandformatbetweenpadandgblur. Keep that order stable: it controls both FFmpeg preprocessing command composition and the slugifiedAssetstring identity. The corresponding properties arefps_cmd,ref_fps_cmd, anddis_fps_cmd;formatis intentionally accessed viaget_filter_cmd("format", target)like the generic filter-only keys.core/src/feature/feature_extractor.hincludes generatedconfig.hbefore definingstruct VmafFeatureExtractor. Keep that include in the header, not just in selected consumer TUs: backend-enabled LTO builds need every extractor definition and the registry to see identicalHAVE_CUDA/HAVE_SYCL/HAVE_VULKANmacro state, otherwise GCC emits-Wlto-type-mismatchand may misoptimise extractor globals.core/src/feature/common/macros.h::FORCE_INLINEalready expands to an inline specifier on GCC/Clang. Do not re-add a second literalinlineto CSF / CAMBI / motion helper declarations; Clang reports the duplicate specifier throughout the build matrix.core/tools/vmaf.c::fetch_picture()owns a preallocated picture slot as soon asvmaf_fetch_preallocated_picture()succeeds. Preserve the EOF/read-error cleanup that unrefs that slot before returning1or-1, and preserve therun_frame_loop()cleanup for the opposite side when only one input read succeeded. Without both, the CLI can score and write output successfully, then hang forever invmaf_close()while the picture pool waits for unread slots to return.scripts/ci/cross_backend_{parity_gate,vif_diff}.pymust keep the backend-specific extractor alias("adm", "vulkan") -> "integer_adm_vulkan". ADR-0586 renamed Vulkan integer ADM to the canonical extractor name while CPU/CUDA/SYCL retained the historicaladm,adm_cuda, andadm_syclnames. Dropping the alias makes the lavapipe parity gate invoke the retiredadm_vulkancompatibility name and fail before comparing scores.core/src/feature/common/convolution_avx512.cvertical scanlines must use_mm512_loadu_ps/_mm512_storeu_ps.MAX_ALIGNis 32 bytes, not 64 bytes; the stride can be a 64-byte multiple while the row base is still only 32-byte aligned. Reintroducing aligned AVX-512 memory ops can crashfloat_vifon AVX-512-capable CPU runners.
Touched files: .github/workflows/tests-and-quality-gates.yml, docs/development/zed-migration-plan-2026-05-19.md, docs/metrics/features.md, docs/usage/python.md, core/src/dnn/AGENTS.md, core/src/dnn/model_loader.c, core/src/dnn/ort_backend.c, core/src/AGENTS.md, core/src/feature/AGENTS.md, core/src/feature/adm_csf_tools.h, core/src/feature/arm64/moment_sve2.c, core/src/feature/arm64/psnr_hvs_neon.c, core/src/feature/arm64/ssimulacra2_host_neon.c, core/src/feature/arm64/ssimulacra2_neon.c, core/src/feature/arm64/ssimulacra2_sve2.c, core/src/feature/barten_csf_tools.h, core/src/feature/cambi.c, core/src/feature/feature_dists.c, core/src/feature/feature_extractor.h, core/src/feature/feature_lpips.c, core/src/feature/feature_mobilesal.c, core/src/feature/fastdvdnet_pre.c, core/src/feature/integer_motion.c, core/src/feature/motion_blend_tools.h, core/src/feature/ssimulacra2.c, core/src/feature/transnet_v2.c, core/src/feature/vulkan/adm_vulkan.c, core/src/feature/vulkan/cambi_vulkan.c, core/src/feature/vulkan/float_vif_vulkan.c, core/src/feature/vulkan/integer_adm_vulkan.c, core/src/feature/vulkan/ssimulacra2_vulkan.c, core/src/feature/vif_tools.c, core/src/feature/x86/psnr_hvs_avx2.c, core/src/feature/x86/ssimulacra2_avx2.c, core/src/feature/x86/ssimulacra2_avx512.c, core/src/feature/x86/ssimulacra2_host_avx2.c, core/src/feature/x86/vif_avx512.c, core/src/libvmaf.c, core/src/libvmaf_priv.h, core/src/framesync.c, core/src/model.c, core/src/output.c, core/src/picture.c, core/src/thread_locale.c, core/src/vulkan/vma_impl.cpp, core/tools/AGENTS.md, core/tools/vmaf.c, core/test/AGENTS.md, core/test/dnn/test_cli.sh, core/test/dnn/test_dnn_session_api.c, core/test/dnn/test_model_loader.c, core/test/dnn/test_ort_internals.c, core/test/dnn/test_tensor_io.c, core/test/dnn/test_vmaf_use_tiny_model.c, core/test/dnn/meson.build, core/test/meson.build, core/test/test_feature_extractor.c, core/test/test_framesync.c, core/test/test_model.c, core/test/test_output.c, core/test/test_predict.c, core/test/test_psnr_hvs_simd.c, core/test/test.c, mcp-server/vmaf-mcp/tests/test_smoke_e2e.py, python/test/asset_test.py, python/vmaf/core/asset.py, core/tools/vmaf_bench.c, core/src/feature/common/AGENTS.md, core/src/feature/common/convolution.h, core/src/feature/common/convolution_avx512.c, core/test/test_vif_simd.c, scripts/ci/AGENTS.md, scripts/ci/cross_backend_parity_gate.py, scripts/ci/cross_backend_vif_diff.py, scripts/ci/test_cross_backend_feature_names.py, docs/state.md, changelog.d/fixed/master-ci-dnn-mcp-coverage-2026-05-19.md, docs/rebase-notes.md (this entry).
ADR-0640 — Tiny-AI Netflix corpus training scaffold (2026-05-20 iteration)¶
No rebase impact on upstream C sources or FFmpeg patches. All touched files are fork-local docs, Python test infrastructure, and the changelog fragment tree.
Key invariants (track when upstream Netflix/vmaf adds its own training surface):
.workingdir2/netflix/is gitignored and the 37 GB corpus is never committed. The--data-rootflag (orVMAF_DATA_ROOTenv var) is the mandatory CLI interface; any training script that hard-codes the corpus path violates this invariant.mcp-server/vmaf-mcp/tests/test_smoke_e2e.pyruns against committed fixtures only (python/test/resource/yuv/src01_hrc00_576x324.yuv). Do not change the smoke test to reference.workingdir2/netflix/.- Architecture selection and the actual training run are deferred to a follow-up PR; do not trigger training from the scaffold branch.
Touched files: docs/adr/0640-tiny-ai-netflix-training-scaffold-2026-05-20.md, docs/adr/_index_fragments/0640-tiny-ai-netflix-training-scaffold-2026-05-20.md, docs/adr/_index_fragments/_order.txt, docs/research/0615-tiny-ai-netflix-training-2026-05-20.md, docs/ai/training-data.md (See also section extended), changelog.d/added/0640-tiny-ai-netflix-training-scaffold-2026-05-20.md, docs/rebase-notes.md (this entry).
ADR-0643 — vmaf-tune encoder-profile report contract¶
No rebase impact on upstream libvmaf C sources. This change touches fork-local vmaf-tune Python code, docs, tests, and the FFmpeg patch stack. The FFmpeg integration is advisory CLI glue only.
Key invariants:
ReportData.to_dict()embedsencoder_profile.schema == "vmaftune.encoder_profile.v1". Future report-shape changes should be additive or should bump the profile schema.vmaf-tune encode-profilemust read raw JSON, HTML, and Markdown reports. HTML raw JSON is escaped in<pre>and intentionally unescaped before parsing.- The profile reader selects one recommendation by
--codec,--target-vmaf, and/or--recommendation-index; it must not implicitly encode every codec or ladder rung. - FFmpeg patch
0015-vmaf-tune-profile-cli-glue.patchstays advisory. Do not duplicate vmaf-tune's JSON/profile selection logic in FFmpeg. - FFmpeg 8.x.x base: upstream tags were fetched on 2026-05-20 and the latest released 8.x.x tag was
n8.1.1(n8.2-devis a dev tag). The fullffmpeg-patches/000*-*.patchseries replayed cleanly against a temporary pristinen8.1.1worktree.
Touched files: tools/vmaf-tune/src/vmaftune/report.py, tools/vmaf-tune/src/vmaftune/encoder_profile.py, tools/vmaf-tune/src/vmaftune/cli.py, tools/vmaf-tune/tests/test_report.py, tools/vmaf-tune/tests/test_encoder_profile.py, tools/vmaf-tune/AGENTS.md, docs/usage/vmaf-tune.md, docs/usage/vmaf-tune-ffmpeg.md, ffmpeg-patches/0015-vmaf-tune-profile-cli-glue.patch, ffmpeg-patches/series.txt, ffmpeg-patches/README.md, docs/adr/0643-vmaf-tune-encoder-profile-contract.md, docs/adr/_index_fragments/0643-vmaf-tune-encoder-profile-contract.md, docs/adr/_index_fragments/_order.txt, docs/research/0643-vmaf-tune-encoder-profile-contract.md, changelog.d/added/vmaf-tune-encoder-profile.md, docs/rebase-notes.md (this entry).
ADR-0644 — vmaf-tune codec runtime variants¶
No upstream Netflix C-source rebase impact. The change is confined to the fork-local tools/vmaf-tune Python CLI/report schema, usage docs, and ADR metadata.
Key invariants:
ADAPTER@VARIANTis a compare display token. The baseADAPTERstill routes through the codec-adapter registry and FFmpeg-c:vencoder name.--encoder-ffmpeg-bin TOKEN=PATHis an exact-token binding. Unknown binding keys are rejected rather than silently falling back to the global--ffmpeg-bin.- Compare JSON/CSV rows now include
adapter,runtime_variant, andffmpeg_bin. Keep these fields together if future schema work touchesCOMPARE_ROW_KEYS.
Touched files: tools/vmaf-tune/src/vmaftune/encoder_runtime.py, tools/vmaf-tune/src/vmaftune/compare.py, tools/vmaf-tune/src/vmaftune/cli.py, tools/vmaf-tune/tests/test_encoder_runtime.py, tools/vmaf-tune/tests/test_compare.py, tools/vmaf-tune/tests/test_compare_no_bisect.py, tools/vmaf-tune/tests/test_compare_rate_quality_sweep.py, tools/vmaf-tune/AGENTS.md, docs/usage/vmaf-tune.md, docs/usage/vmaf-tune-codec-adapters.md, docs/adr/0644-vmaf-tune-codec-runtime-variants.md, docs/adr/_index_fragments/0644-vmaf-tune-codec-runtime-variants.md, docs/adr/_index_fragments/_order.txt, docs/research/0644-vmaf-tune-codec-runtime-variants.md, changelog.d/added/0644-vmaf-tune-codec-runtime-variants.md, docs/rebase-notes.md (this entry).
ADR-0645 — Integer ADM p-norm SIMD callback ABI¶
When rebasing any upstream change that touches integer ADM contrast-measure callbacks, keep adm_p_norm threaded through the scalar and x86 SIMD twins.
Touched ABI group: core/src/feature/integer_adm.c, core/src/feature/x86/adm_avx2.c, core/src/feature/x86/adm_avx512.c, core/src/feature/x86/adm_avx2.h, core/src/feature/x86/adm_avx512.h.
Invariant: adm_cm and i4_adm_cm must accept the p-norm parameter and the final powf exponent must be 1.0f / (float)adm_p_norm in every twin. The default 3.0 path is the Netflix-compatible path; do not split SIMD dispatch back to a hard-coded exponent when resolving conflicts.
ADR-0648 — CHUG HDR MOS trainer entry point¶
CHUG HDR subjective-MOS experiments use ai/scripts/train_chug_hdr_mos_head.py and local chug_hdr_mos_head_v1 manifests. Keep CHUG operator docs on that entry point; do not reintroduce instructions that pass CHUG shards through train_konvid_mos_head.py's KonViD-named flags. The wrapper may reuse the shared MOS-head implementation, but it must pass explicit non-existent KonViD paths so CHUG runs cannot accidentally mix local KonViD rows with HDR MOS shards.
ADR-0649 — CHUG HDR wide MOS feature schema¶
train_chug_hdr_mos_head.py defaults to --feature-schema chug-hdr-wide-v1. That schema is CHUG-local and currently 34 columns: canonical-6 means, p10/p90 / std temporal aggregates, and HDR ladder / geometry metadata. Do not edit the KonViD FEATURE_COLUMNS order to implement CHUG experiments; keep the shipped konvid_mos_head_v1 ONNX on the konvid-v1 11-column schema. Downstream CHUG experiment scripts must read feature_schema and feature_order from the manifest instead of assuming 11 inputs.
ADR-0331 — rule-enforcement ready-for-review trigger repair¶
.github/workflows/rule-enforcement.yml must include ready_for_review in its pull_request.types list, alongside edited. Without ready_for_review, draft-to-ready promotion leaves the ADR-0108, ADR-0100, FFmpeg-surface, ADR-number, backfill, and docs/state.md gates stuck on their draft-time skipped check runs while the heavier workflows rerun correctly. Keep edited as well so PR-body fixes can rerun only the rule-enforcement workflow without burning the full matrix again.
Test build graph — generated vcs_version.h dependency¶
core/test/test_feature_collector.c directly includes core/src/libvmaf.c, and libvmaf.c includes the generated vcs_version.h header. Keep rev_target listed in the test_feature_collector executable sources in core/test/meson.build; otherwise fresh parallel Ninja builds can compile the test before include/vcs_version.h exists and fail nondeterministically.
Vulkan lavapipe CI — motion probes stay out of the VIF job¶
The Vulkan VIF Cross-Backend (lavapipe, places=4) job should not run the known-broken motion / motion_v2 lavapipe probes with continue-on-error: true. GitHub still emits ##[error] annotations for those advisory failures, which makes a passing PR look broken. Keep the documented T-VULKAN-MOTION-LAVAPIPE-INIT debt in docs/state.md and keep the required GPU-Parity Matrix Gate skip list until the Vulkan motion lavapipe bug is actually fixed; do not reintroduce advisory failing steps inside the named VIF gate.
ADR-0647 — fr_regressor_v1 Netflix refresh¶
No upstream Netflix C-source rebase impact. This is a fork-local model artifact refresh: model/tiny/fr_regressor_v1.onnx, its sidecar, registry row, model card, ADR/research docs, and state/changelog metadata.
Key invariants:
- The ADR-0249 model recipe and PLCC ship gate stay unchanged. Do not use this refresh as precedent for changing architecture, feature order, or gate threshold.
- Refreshes must train from a dated current full-feature table, not from stale
runs/full_features_netflix.parquet. fr_regressor_v1.onnxis inline after export. If the exporter rewrites the stale sibling.onnx.datafile while the ONNX has no external initializers, restore the orphan sidecar rather than expanding the model diff.
Touched files: model/tiny/fr_regressor_v1.onnx, model/tiny/fr_regressor_v1.json, model/tiny/registry.json, docs/ai/models/fr_regressor_v1.md, docs/adr/0647-ai-fr-regressor-v1-refresh-20260520.md, docs/adr/_index_fragments/0647-ai-fr-regressor-v1-refresh-20260520.md, docs/adr/_index_fragments/_order.txt, docs/research/0647-ai-fr-regressor-v1-refresh-20260520.md, ai/AGENTS.md, docs/state.md, changelog.d/changed/0647-ai-fr-regressor-v1-refresh-20260520.md, docs/rebase-notes.md (this entry).
ADR-0651 — CHUG HDR row metadata¶
No upstream Netflix C-source rebase impact. This is a fork-local AI corpus-materialisation schema extension in ai/scripts/chug_extract_features.py.
Key invariants:
- CHUG feature rows now preserve
feature_ref_*andfeature_dis_*ffprobe HDR/display metadata for the matched reference and distorted clip. - Unknown ffprobe fields remain explicit as
unknownornull; do not infer panel/display capability in the materialiser. - The existing
--audit-outputcorpus preflight remains the aggregate health check; the row fields are the model-facing copy.
Touched files: ai/scripts/chug_extract_features.py, ai/tests/test_chug.py, docs/ai/chug-ingestion.md, docs/adr/0651-chug-hdr-row-metadata.md, docs/adr/_index_fragments/0651-chug-hdr-row-metadata.md, docs/adr/_index_fragments/_order.txt, docs/adr/README.md, docs/research/0651-chug-hdr-row-metadata.md, ai/AGENTS.md, changelog.d/added/0651-chug-hdr-row-metadata.md, docs/rebase-notes.md (this entry).
ADR-0652 — CHUG visual-signal primitives¶
No upstream Netflix C-source rebase impact. This is a fork-local AI feature-row schema extension in ai/scripts/chug_extract_features.py.
Key invariants:
- CHUG feature rows now include
feature_ref_*,feature_dis_*, andfeature_delta_*luma-domain visual-signal primitives forluma_std,sharpness_laplacian_var,highfreq_abs_mean, andnoise_lap_mad. - These are deterministic diagnostic blur/noise/grain proxies computed from sampled decoded YUV10 luma frames. Do not treat them as a trained no-reference VQA model.
- The visual-signal cache lives beside the existing CHUG feature cache and must be regenerated if the primitive definitions change.
Touched files: ai/scripts/chug_extract_features.py, ai/tests/test_chug.py, docs/ai/chug-ingestion.md, docs/adr/0652-chug-visual-signal-primitives.md, docs/adr/_index_fragments/0652-chug-visual-signal-primitives.md, docs/adr/_index_fragments/_order.txt, docs/adr/README.md, docs/research/0652-chug-visual-signal-primitives.md, ai/AGENTS.md, changelog.d/added/0652-chug-visual-signal-primitives.md, docs/rebase-notes.md (this entry).
ADR-0653 — CHUG display-profile training¶
No upstream Netflix C-source rebase impact. This is a fork-local CHUG HDR MOS training-schema extension under ai/ plus docs and DDD material.
Key invariants:
chug-hdr-wide-v1remains the no-profile CHUG default.--display-profile-jsonselectschug-hdr-display-v1only when the caller did not explicitly pass--feature-schema.- Row-local display fields override the target profile so future multi-display HDR corpora keep their panel axis.
- The display profile is recorded in the emitted manifest with normalized feature values and source sha256.
Touched files: ai/scripts/train_konvid_mos_head.py, ai/scripts/train_chug_hdr_mos_head.py, ai/tests/test_train_konvid_mos_head.py, docs/ai/chug-ingestion.md, docs/ai/mos-corpora.md, docs/ai/models/konvid_mos_head_v1.md, docs/adr/0653-chug-display-profile-training.md, docs/adr/_index_fragments/0653-chug-display-profile-training.md, docs/adr/_index_fragments/_order.txt, docs/research/0653-chug-display-profile-training.md, ai/AGENTS.md, changelog.d/added/0653-chug-display-profile-training.md, docs/rebase-notes.md (this entry).
ADR-0657 — Second-opinion feature materializer¶
No upstream Netflix C-source rebase impact. This is a fork-local AI feature-table enrichment utility under ai/scripts/.
Key invariants:
ai/scripts/materialize_second_opinion_features.pystays table-side: it joins already-generated scorer JSON/JSONL and must not invoke or vendor third-party VQA projects.- Output columns remain namespaced as
second_opinion_<scorer>_*so downstream audits and trainers can detect NR/MOS evidence without colliding with native corpus columns. - Duplicate
(scorer, key)rows are rejected; they usually indicate stale reruns or mismatched row keys and must not be averaged silently.
Touched files: ai/scripts/materialize_second_opinion_features.py, ai/scripts/signal_mix_audit.py, ai/tests/test_second_opinion_features.py, docs/ai/second-opinion-features.md, docs/ai/signal-mix-audit.md, docs/ai/index.md, docs/adr/0657-second-opinion-feature-materializer.md, docs/adr/_index_fragments/0657-second-opinion-feature-materializer.md, docs/adr/_index_fragments/_order.txt, docs/research/0657-second-opinion-feature-materializer.md, ai/AGENTS.md, mkdocs.yml, changelog.d/added/0657-second-opinion-feature-materializer.md, docs/rebase-notes.md (this entry).
ADR-0658 — Project modernization audit¶
No upstream Netflix C-source rebase impact. This is a fork-local developer-tooling audit under scripts/dev/.
Key invariants:
scripts/dev/project_modernization_audit.pyis read-only. It may emit JSON and Markdown, but it must not rewrite.workingdir2/OPEN.md,.workingdir2/BACKLOG.md,docs/state.md, changelog fragments, or PR bodies.- The scanner is advisory queue shaping, not a required CI gate. Its marker matches are intentionally text-based and need human triage.
- Archived scratch remains skipped by default; include it only with
--include-archivesduring deliberate archaeology.
Touched files: scripts/dev/project_modernization_audit.py, scripts/dev/test_project_modernization_audit.py, docs/development/project-modernization-audit.md, docs/adr/0658-project-modernization-audit.md, docs/adr/_index_fragments/0658-project-modernization-audit.md, docs/adr/_index_fragments/_order.txt, docs/adr/README.md, docs/research/0658-project-modernization-audit.md, scripts/AGENTS.md, mkdocs.yml, changelog.d/added/0658-project-modernization-audit.md, docs/rebase-notes.md (this entry).
ADR-0659 — Modernization audit false-positive filter¶
No upstream Netflix C-source rebase impact. This is a fork-local developer-tooling precision fix under scripts/dev/.
Key invariants:
- Live Python
raise NotImplementedError(...)rows remain high-severity audit findings. - Historical closeout prose such as "replaced the NotImplementedError scaffold", Python
except NotImplementedErrorhandlers, and customNotImplementedErrorexception subclasses are not modernization gaps. - Documented
-ENOSYSoptional-build contracts are not modernization gaps; barereturn -ENOSYS;rows outside such context still are. - Add future suppressions as narrow line-context tests; avoid file-level suppressions that could hide new real debt.
Touched files: scripts/dev/project_modernization_audit.py, scripts/dev/test_project_modernization_audit.py, docs/development/project-modernization-audit.md, docs/adr/0659-modernization-audit-false-positive-filter.md, docs/adr/_index_fragments/0659-modernization-audit-false-positive-filter.md, docs/adr/_index_fragments/_order.txt, docs/adr/README.md, docs/research/0659-modernization-audit-false-positive-filter.md, scripts/AGENTS.md, changelog.d/fixed/0659-modernization-audit-false-positive-filter.md, docs/rebase-notes.md (this entry).
ADR-0661 — AI run manifest provenance¶
No upstream Netflix C-source rebase impact. This is fork-local AI tooling under ai/ plus human-facing docs.
Key invariants:
- New AI training/export sidecars should use
aiutils.run_manifest.build_run_provenance()instead of hand-rolled path hashing or argument JSON. - CHUG MOS wrapper runs record
train_chug_hdr_mos_head.pyas the user-facingentrypoint, even though they delegate into the shared KonViD training loop. - Add
shared_trainerwhen wrapper identity and implementation script differ.
Touched files: ai/src/aiutils/run_manifest.py, ai/src/aiutils/__init__.py, ai/scripts/train_konvid_mos_head.py, ai/scripts/train_chug_hdr_mos_head.py, ai/tests/test_run_manifest.py, ai/tests/test_train_konvid_mos_head.py, ai/AGENTS.md, ai/src/aiutils/AGENTS.md, docs/ai/training.md, docs/ai/models/konvid_mos_head_v1.md, docs/ai/chug-ingestion.md, docs/adr/0661-ai-run-manifest-provenance.md, docs/adr/_index_fragments/0661-ai-run-manifest-provenance.md, docs/adr/_index_fragments/_order.txt, docs/adr/README.md, docs/research/0661-ai-run-manifest-provenance.md, changelog.d/added/0661-ai-run-manifest-provenance.md, docs/rebase-notes.md (this entry).
ADR-0662 — Vulkan motion lavapipe parity¶
Rebase-sensitive feature-extractor impact. This changes fork-local GPU motion twins and CI parity routing; keep these invariants when resolving any upstream sync that touches motion, feature registration, or parity scripts.
Key invariants:
integer_motion_vulkanstays before legacymotion_vulkanin the Vulkan registry block so model feature-name dispatch chooses the lavapipe-stable canonical twin.- Both parity scripts keep
BACKEND_EXTRACTOR_ALIASES[("motion", "vulkan")] = "integer_motion_vulkan". - CUDA, SYCL, and Vulkan
motion_v2kernels use the CPUinteger_motion_v2.c::mirrorhigh-edge literal2 * size - idx - 2; do not restore the stale-1formula from old ADR-0193 prose. integer_motion_vulkandefaultsdebug=true, matching CPU, CUDA, and the legacy Vulkan motion extractor, so the rawinteger_motionmetric is emitted for parity.
Touched files: .github/workflows/tests-and-quality-gates.yml, core/src/feature/feature_extractor.c, core/src/feature/vulkan/integer_motion_vulkan.c, core/src/feature/vulkan/shaders/motion_v2.comp, core/src/feature/cuda/integer_motion_v2/motion_v2_score.cu, core/src/feature/sycl/integer_motion_v2_sycl.cpp, scripts/ci/cross_backend_vif_diff.py, scripts/ci/cross_backend_parity_gate.py, docs/metrics/motion.md, docs/metrics/features.md, docs/backends/vulkan/overview.md, docs/api/gpu.md, docs/development/cross-backend-gate.md, docs/adr/0193-motion-v2-vulkan.md, docs/adr/0662-vulkan-motion-lavapipe-parity.md, docs/adr/_index_fragments/0193-motion-v2-vulkan.md, docs/adr/_index_fragments/0662-vulkan-motion-lavapipe-parity.md, docs/adr/_index_fragments/_order.txt, docs/research/0662-vulkan-motion-lavapipe-parity.md, core/src/feature/AGENTS.md, core/src/feature/vulkan/AGENTS.md, core/src/feature/cuda/AGENTS.md, scripts/ci/AGENTS.md, changelog.d/fixed/0662-vulkan-motion-lavapipe-parity.md, docs/rebase-notes.md (this entry).
ADR-0663 — MOS label materializer¶
No upstream Netflix C-source rebase impact. This is fork-local AI training/data-prep plumbing under ai/scripts/.
Key invariants:
ai/scripts/materialize_mos_labels.pystays table-side: it joins subjective MOS labels onto already-extracted feature tables and must not extract features, download corpora, or train models.- Real MOS-head training must not silently synthesize data when explicit real-corpus paths produce zero labelled rows.
--smokeis the documented synthetic path. - Conflicting duplicate label keys are rejected; low unique-key coverage fails by default so stale key joins do not become training inputs.
Touched files: ai/scripts/materialize_mos_labels.py, ai/scripts/train_konvid_mos_head.py, ai/tests/test_materialize_mos_labels.py, ai/tests/test_train_konvid_mos_head.py, docs/ai/mos-label-materializer.md, docs/ai/mos-corpora.md, docs/ai/models/konvid_mos_head_v1.md, docs/ai/index.md, docs/adr/0663-mos-label-materializer.md, docs/adr/_index_fragments/0663-mos-label-materializer.md, docs/adr/_index_fragments/_order.txt, docs/research/0663-mos-label-materializer.md, ai/AGENTS.md, mkdocs.yml, changelog.d/added/0663-mos-label-materializer.md, docs/rebase-notes.md (this entry).
ADR-0661 follow-up — AI model sidecar run provenance¶
AI-sidecar provenance impact. This widens ADR-0661 from MOS-head trainers to the FR regressor training family and the vmaf_tiny exporter family.
Key invariants:
train_fr_regressor.py,train_fr_regressor_v2.py, andtrain_fr_regressor_v3.pysidecars carryrun_provenancebuilt byaiutils.run_manifest.build_run_provenance().- v1/v2 metrics JSON carries the same block, including gate-failed runs where no ONNX export is written.
export_vmaf_tiny_v2.py,export_vmaf_tiny_v3.py, andexport_vmaf_tiny_v4.pysidecars carry the same block with the checkpoint input and ONNX/sidecar output targets.- Do not replace this with per-script argument/path JSON when rebasing AI trainer changes; extend the shared helper instead.
Touched files: ai/scripts/export_vmaf_tiny_v2.py, ai/scripts/export_vmaf_tiny_v3.py, ai/scripts/export_vmaf_tiny_v4.py, ai/scripts/train_fr_regressor.py, ai/scripts/train_fr_regressor_v2.py, ai/scripts/train_fr_regressor_v3.py, ai/tests/test_fr_regressor_run_provenance.py, ai/tests/test_vmaf_tiny_export_run_provenance.py, docs/ai/training.md, docs/ai/models/fr_regressor_v1.md, docs/ai/models/fr_regressor_v2.md, docs/ai/models/fr_regressor_v3.md, docs/ai/models/vmaf_tiny_v2.md, docs/ai/models/vmaf_tiny_v3.md, docs/ai/models/vmaf_tiny_v4.md, docs/adr/0661-ai-run-manifest-provenance.md, docs/research/0664-ai-fr-regressor-run-provenance.md, changelog.d/added/0664-ai-fr-regressor-run-provenance.md, docs/rebase-notes.md (this entry).
ADR-0661 follow-up — AI eval report run provenance¶
AI-eval/validate provenance impact. This widens ADR-0661 from model-producing sidecars to the tiny-VMAF evaluation and validation report families.
Key invariants:
eval_loso_vmaf_tiny_v3.py,eval_loso_vmaf_tiny_v4.py,eval_loso_vmaf_tiny_v5.py, andeval_multiseed_v3_v4.pyreport JSON files carryrun_provenancebuilt byaiutils.run_manifest.build_run_provenance().- Evaluation reports record the feature parquet input(s), parsed eval hyperparameters, original argv, and
report_targetoutput path. validate_ensemble_seeds.pyverdict JSON files carry the same schema and recordloso_dir,corpus_root, seed list, gate thresholds, and thePROMOTE.json/HOLD.jsonoutput path.- Do not restore per-script
json.dumps(...).write_text(...)report writers when rebasing eval-script changes; usewrite_manifest_json()so the JSON shape and newline handling stay shared.
Touched files: ai/scripts/eval_loso_vmaf_tiny_v3.py, ai/scripts/eval_loso_vmaf_tiny_v4.py, ai/scripts/eval_loso_vmaf_tiny_v5.py, ai/scripts/eval_multiseed_v3_v4.py, ai/scripts/validate_ensemble_seeds.py, ai/tests/test_eval_report_run_provenance.py, ai/tests/test_validate_ensemble_seeds.py, docs/ai/training.md, docs/ai/ensemble-v2-real-corpus-retrain-runbook.md, docs/ai/models/vmaf_tiny_v3.md, docs/ai/models/vmaf_tiny_v4.md, docs/ai/models/vmaf_tiny_v5.md, docs/adr/0661-ai-run-manifest-provenance.md, docs/research/0665-ai-eval-report-run-provenance.md, ai/src/aiutils/AGENTS.md, changelog.d/added/0665-ai-eval-report-run-provenance.md, docs/rebase-notes.md (this entry).
ADR-0661 follow-up — Legacy AI eval report run provenance¶
Legacy eval/report provenance impact. This widens ADR-0661 adoption from the refreshed v3/v4/v5 eval family to older durable AI evaluation reports.
Key invariants:
eval_loso_mlp_small.pyandeval_loso_3arch.pyJSON reports carryrun_provenancebuilt byaiutils.run_manifest.build_run_provenance().eval_probabilistic_proxy.py --metrics-outwrites the same block with the ensemble manifest input, optional held-out parquet, and metrics output path.eval_saliency_per_mb.pywrites the same block for CLI output, recording the predicted and ground-truth mask directories plus block settings.- Do not restore direct
json.dump()writers for these durable reports when rebasing old eval-script changes; usewrite_manifest_json()for stable sorting and newline handling.
Touched files: ai/scripts/eval_loso_mlp_small.py, ai/scripts/eval_loso_3arch.py, ai/scripts/eval_probabilistic_proxy.py, ai/scripts/eval_saliency_per_mb.py, ai/tests/test_legacy_eval_report_run_provenance.py, ai/tests/test_eval_saliency_per_mb.py, docs/ai/training.md, docs/ai/loso-eval.md, docs/ai/saliency-per-mb-eval.md, docs/ai/models/fr_regressor_v2_probabilistic.md, docs/adr/0661-ai-run-manifest-provenance.md, docs/research/0666-ai-legacy-eval-report-provenance.md, changelog.d/added/0666-ai-legacy-eval-report-provenance.md, docs/rebase-notes.md (this entry).
ADR-0661 follow-up — Predictor v2 real-corpus report provenance¶
Predictor-v2 report provenance impact. This widens ADR-0661 adoption to the per-codec real-corpus gate report used before predictor-v2 model-card updates.
Key invariants:
ai/scripts/train_predictor_v2_realcorpus.pywritesruns/predictor_v2_realcorpus/report.jsonwith arun_provenanceblock built byaiutils.run_manifest.build_run_provenance().- The report records the trainer entrypoint, original argv, parsed arguments, explicit corpus files, corpus roots, resolved JSONL files, and report target.
- Keep ADR-0303 gate constants untouched; provenance makes failed or insufficient reports reproducible, but it does not change pass/fail logic.
- Do not restore direct
Path.write_text(json.dumps(...))report output here; usewrite_manifest_json()so JSON sorting and trailing-newline behavior stay shared with the other AI provenance reports.
Touched files: ai/scripts/train_predictor_v2_realcorpus.py, ai/tests/test_train_predictor_v2_realcorpus.py, docs/ai/predictor-v2-realcorpus-training.md, docs/ai/training.md, docs/adr/0661-ai-run-manifest-provenance.md, docs/research/0667-predictor-v2-report-provenance.md, changelog.d/added/0667-predictor-v2-report-provenance.md, docs/rebase-notes.md (this entry).
ADR-0661 follow-up — vmaf_tiny train stats provenance¶
vmaf_tiny training provenance impact. This widens ADR-0661 adoption to the pre-export stats JSON files emitted by the vmaf_tiny trainer family.
Key invariants:
train_vmaf_tiny_v2.py,train_vmaf_tiny_v3.py,train_vmaf_tiny_v4.py, andtrain_vmaf_tiny_v5.pywrite their--out-statsJSON with arun_provenanceblock built byaiutils.run_manifest.build_run_provenance().- v2/v3/v4 stats record the parquet input, checkpoint target, stats target, argv, and parsed hyperparameters.
- v5 stats record both
parquet_baseandparquet_extra, plus checkpoint and stats output targets. - Do not restore direct
Path.write_text(json.dumps(...))stats output here; usewrite_manifest_json()so JSON sorting and trailing-newline behavior stay shared with the other AI provenance reports.
Touched files: ai/scripts/train_vmaf_tiny_v2.py, ai/scripts/train_vmaf_tiny_v3.py, ai/scripts/train_vmaf_tiny_v4.py, ai/scripts/train_vmaf_tiny_v5.py, ai/tests/test_vmaf_tiny_train_run_provenance.py, docs/ai/training.md, docs/ai/models/vmaf_tiny_v2.md, docs/ai/models/vmaf_tiny_v3.md, docs/ai/models/vmaf_tiny_v4.md, docs/ai/models/vmaf_tiny_v5.md, docs/adr/0661-ai-run-manifest-provenance.md, docs/research/0668-vmaf-tiny-train-stats-provenance.md, changelog.d/added/0668-vmaf-tiny-train-stats-provenance.md, docs/rebase-notes.md (this entry).
ADR-0661 follow-up — AI materializer audit provenance¶
Materializer audit provenance impact. This widens ADR-0661 adoption to feature-table materializers and signal-mix audit reports that feed retraining or model-mix decisions.
Key invariants:
materialize_mos_labels.py --audit-jsonandmaterialize_second_opinion_features.py --audit-jsonincluderun_provenancein their audit JSON outputs.materialize_saliency_features.py --audit-jsonwrites row counters, the effective config, andrun_provenance; use it for saliency-enriched tables that feed retraining.signal_mix_audit.py --out-jsonincludesrun_provenancefor audited table paths, thresholds, argv, JSON output, and Markdown output.- Do not reintroduce bespoke path hashing or direct audit JSON writers on these surfaces; use
aiutils.run_manifest.build_run_provenance()andwrite_manifest_json().
Touched files: ai/scripts/materialize_mos_labels.py, ai/scripts/materialize_second_opinion_features.py, ai/scripts/materialize_saliency_features.py, ai/scripts/signal_mix_audit.py, ai/tests/test_materialize_mos_labels.py, ai/tests/test_second_opinion_features.py, ai/tests/test_materialize_saliency_features.py, ai/tests/test_signal_mix_audit.py, docs/ai/training.md, docs/ai/mos-label-materializer.md, docs/ai/second-opinion-features.md, docs/ai/saliency-feature-materializer.md, docs/ai/signal-mix-audit.md, docs/adr/0661-ai-run-manifest-provenance.md, docs/research/0669-ai-materializer-audit-provenance.md, changelog.d/added/0669-ai-materializer-audit-provenance.md, docs/rebase-notes.md (this entry).
ADR-0661 follow-up — Ensemble seed export provenance¶
Ensemble export provenance impact. This widens ADR-0661 adoption to the production seed exporter for fr_regressor_v2_ensemble_v1_seed* sidecars.
Key invariants:
ai/scripts/export_ensemble_v2_seeds.pybuilds onerun_provenanceblock per invocation with the corpus, PROMOTE verdict, parsed export args, argv, per-seed ONNX/sidecar targets, and optional registry target.- Each fresh
fr_regressor_v2_ensemble_v1_seed{N}.jsonsidecar receives that block. - Sidecar and optional registry writes use
write_manifest_json()so the JSON formatting contract matches the other ADR-0661 adopters.
Touched files: ai/scripts/export_ensemble_v2_seeds.py, ai/tests/test_export_ensemble_v2_seeds_provenance.py, docs/ai/training.md, docs/ai/models/fr_regressor_v2_probabilistic.md, docs/ai/ensemble-training-kit.md, docs/adr/0661-ai-run-manifest-provenance.md, docs/research/0670-ensemble-seed-export-provenance.md, changelog.d/added/0670-ensemble-seed-export-provenance.md, docs/rebase-notes.md (this entry).
ADR-0661 follow-up — Ensemble LOSO report provenance¶
Ensemble LOSO provenance impact. This widens ADR-0661 adoption to the per-seed loso_seed{N}.json reports emitted by the production ensemble LOSO trainer.
Key invariants:
ai/scripts/train_fr_regressor_v2_ensemble_loso.pywrites eachloso_seed{N}.jsonthroughwrite_manifest_json().- Each report includes
run_provenancewith the trainer entrypoint, original argv, parsed training args, corpus JSONL input, and per-seed report target. - The existing gate keys (
mean_plcc,min_plcc,max_plcc,folds, and seed metadata) remain unchanged forscripts/ci/ensemble_prod_gate.pyandai/scripts/validate_ensemble_seeds.py.
Touched files: ai/scripts/train_fr_regressor_v2_ensemble_loso.py, ai/tests/test_train_fr_regressor_v2_ensemble_loso_train.py, docs/ai/training.md, docs/ai/ensemble-v2-real-corpus-retrain-runbook.md, docs/ai/ensemble-training-kit.md, docs/adr/0661-ai-run-manifest-provenance.md, docs/research/0671-ensemble-loso-report-provenance.md, changelog.d/added/0671-ensemble-loso-report-provenance.md, docs/rebase-notes.md (this entry).
ADR-0661 follow-up — vmaf-train CLI report provenance¶
vmaf-train report provenance impact. This widens ADR-0661 adoption to the user-facing vmaf-train --json report surfaces.
Key invariants:
ai/src/vmaf_train/cli.pyuses_write_cli_report_json()for durable report commands that accept--json.- Covered subcommands:
validate-norm,profile,audit-learned-filter,quantize-int8,cross-backend, andbisect-model-quality. - The provenance block records the CLI entrypoint, argv, parsed options, model/feature/calibration/frame inputs, JSON report output, and generated model output where a command writes one.
Touched files: ai/src/vmaf_train/cli.py, ai/tests/test_tune_cli.py, docs/usage/vmaf-train.md, docs/ai/training.md, docs/adr/0661-ai-run-manifest-provenance.md, docs/research/0672-vmaf-train-cli-report-provenance.md, changelog.d/added/0672-vmaf-train-cli-report-provenance.md, docs/rebase-notes.md (this entry).
ADR-0661 follow-up — Feature-correlation report provenance¶
Feature-correlation provenance impact. This widens ADR-0661 adoption to the feature-ranking report emitted by ai/scripts/feature_correlation.py.
Key invariants:
feature_correlation.py --outwrites the JSON report throughwrite_manifest_json().- The report includes
run_provenancewith the analyzer entrypoint, original argv, parsed target / redundancy / top-K arguments, source parquet input, and JSON report target. - The analytic payload keys (
pearson,redundant_pairs,importances,per_method_topk, andconsensus_topk) remain unchanged for downstream research/audit readers.
Touched files: ai/scripts/feature_correlation.py, ai/tests/test_feature_correlation.py, docs/research/0027-phase2-feature-importance.md, docs/ai/training.md, docs/adr/0661-ai-run-manifest-provenance.md, docs/research/0673-feature-correlation-report-provenance.md, changelog.d/added/0673-feature-correlation-report-provenance.md, docs/rebase-notes.md (this entry).
ADR-0661 follow-up — Phase-3 subset-sweep report provenance¶
Phase-3 sweep provenance impact. This widens ADR-0661 adoption to the model-selection JSON emitted by ai/scripts/phase3_subset_sweep.py.
Key invariants:
phase3_subset_sweep.py --outwrites the JSON report throughwrite_manifest_json().- The report keeps existing subset result keys and adds top-level
run_provenancewith the analyzer entrypoint, original argv, parsed subset / seed / standardization arguments, source parquet input, and JSON report target. - The subset result payload (
features,per_seed,summary) remains unchanged for each requested subset.
Touched files: ai/scripts/phase3_subset_sweep.py, ai/tests/test_phase3_subset_sweep.py, docs/research/0028-phase3-subset-sweep.md, docs/ai/training.md, docs/adr/0661-ai-run-manifest-provenance.md, docs/research/0674-phase3-subset-report-provenance.md, changelog.d/added/0674-phase3-subset-report-provenance.md, docs/rebase-notes.md (this entry).
ADR-0661 follow-up — Quantisation report provenance¶
Quantisation provenance impact. This widens ADR-0661 adoption to the int8 producer/gate scripts used for model-card promotion evidence.
Key invariants:
ai/scripts/ptq_dynamic.py --report-outandai/scripts/ptq_static.py --report-outwrite JSON reports with fp32/int8 sizes, selected quantisation settings, andrun_provenance.ai/scripts/qat_train.py --report-outwrites QAT output/report metadata for the fp32 bridge and final int8 ONNX artifact.ai/scripts/measure_quant_drop.py --out-jsonpreserves per-model gate rows andrun_provenancewithout changing stdout or exit codes.
Touched files: ai/scripts/ptq_dynamic.py, ai/scripts/ptq_static.py, ai/scripts/qat_train.py, ai/scripts/measure_quant_drop.py, ai/tests/test_ptq_scripts.py, ai/tests/test_qat_smoke.py, docs/ai/quantization.md, docs/ai/training.md, docs/adr/0661-ai-run-manifest-provenance.md, docs/research/0682-quantization-report-provenance.md, changelog.d/added/0682-quantization-report-provenance.md, docs/rebase-notes.md (this entry).
ADR-0661 follow-up — CHUG extraction report provenance¶
CHUG extraction provenance impact. This widens ADR-0661 adoption to the local CHUG split manifest and HDR metadata audit JSON emitted before HDR MOS training.
Key invariants:
ai/scripts/chug_extract_features.py --split-manifestkeeps the existing content-level split payload and adds top-levelrun_provenance.ai/scripts/chug_extract_features.py --audit-outputkeeps the existing HDR audit counters/malformed-row payload and adds top-levelrun_provenance.- Feature JSONL rows are unchanged; this PR only stamps the durable split/audit JSON evidence with extractor command and input/output context.
Touched files: ai/scripts/chug_extract_features.py, ai/tests/test_chug.py, docs/ai/chug-ingestion.md, docs/adr/0661-ai-run-manifest-provenance.md, docs/research/0684-chug-extraction-report-provenance.md, changelog.d/added/0684-chug-extraction-report-provenance.md, docs/rebase-notes.md (this entry).
ADR-0668 — AI derived table provenance¶
Derived-table provenance impact. This extends the ADR-0661 manifest pattern from trainer/report JSONs down to the local FULL_FEATURES parquet builders that feed refreshed AI models.
Key invariants:
ai/scripts/extract_k150k_features.pywrites<out>.manifest.jsonby default with feature order, CPU/CUDA extractor split, restart counters, backend worker counts, parquet row count, and sharedrun_provenance.ai/scripts/combine_full_feature_parquets.pywrites<out>.manifest.jsonby default with input labels, per-input row counts, missing-feature fill lists, corpus distribution, output column order, and sharedrun_provenance.ai/scripts/enrich_k150k_parquet_metadata.pywrites<out>.manifest.jsonby default with metadata match/update counters, available metadata keys, overwrite policy, and sharedrun_provenance.- Existing parquet row schemas are unchanged; the manifest is a sibling local evidence artifact.
Touched files: ai/scripts/extract_k150k_features.py, ai/scripts/combine_full_feature_parquets.py, ai/scripts/enrich_k150k_parquet_metadata.py, ai/tests/test_extract_k150k_features.py, ai/tests/test_combine_full_feature_parquets.py, ai/tests/test_enrich_k150k_parquet_metadata.py, ai/AGENTS.md, docs/ai/training.md, docs/ai/chug-ingestion.md, docs/adr/0668-ai-derived-table-provenance.md, docs/research/0688-ai-derived-table-provenance.md, changelog.d/added/0668-ai-derived-table-provenance.md, docs/rebase-notes.md (this entry).
ADR-0669 — AI corpus JSONL provenance¶
Corpus-JSONL provenance impact. This extends the ADR-0661 manifest pattern to the corpus JSONL boundary before trainers consume merged or aggregated row streams.
Key invariants:
ai/scripts/aggregate_corpora.pywrites<output>.manifest.jsonby default with MOS scale conversions, optional corpus-source overrides, aggregate counters, and sharedrun_provenance.ai/scripts/merge_corpora.pywrites<output>.manifest.jsonby default with required vmaf-tune corpus keys, natural dedup key, merge counters, and sharedrun_provenance.- JSONL row schemas are unchanged; run-level evidence belongs in the sidecar.
Touched files: ai/scripts/aggregate_corpora.py, ai/scripts/merge_corpora.py, ai/tests/test_aggregate_corpora.py, ai/tests/test_merge_corpora.py, ai/AGENTS.md, docs/ai/mos-corpora.md, docs/ai/multi-corpus-aggregation.md, docs/ai/training.md, docs/adr/0669-ai-corpus-jsonl-provenance.md, docs/research/0689-ai-corpus-jsonl-provenance.md, changelog.d/added/0669-ai-corpus-jsonl-provenance.md, docs/rebase-notes.md (this entry).
ADR-0670 — AI legacy corpus extraction manifests¶
Legacy trainer-input provenance impact. This extends the ADR-0661 manifest pattern to older corpus/extraction scripts that directly create local trainer-input parquets or vmaf-tune JSONL.
Key invariants:
ai/scripts/extract_full_features.pywrites<out>.manifest.jsonby default with Netflix corpus/cache inputs, VMAF binary evidence, feature list, pair count, row count, and sharedrun_provenance.ai/scripts/konvid_to_vmaf_pairs.pywrites<out>.manifest.jsonby default with KoNViD root, VMAF/model inputs, cache policy, CRF, feature list, clip / frame counters, failed clip IDs, and sharedrun_provenance.ai/scripts/bvi_dvc_to_corpus_jsonl.pywrites<output>.manifest.jsonby default with cache inputs, row schema version, adapter labels, row/cache counters, and sharedrun_provenance.- BVI-DVC JSONL rows must include the current vmaf-tune v3 additive keys; unavailable HDR, shot, canonical-feature aggregate, and encoder-internal values are explicit defaults, not missing columns.
Touched files: ai/scripts/extract_full_features.py, ai/scripts/konvid_to_vmaf_pairs.py, ai/scripts/bvi_dvc_to_corpus_jsonl.py, ai/tests/test_legacy_corpus_extraction_manifests.py, ai/AGENTS.md, docs/ai/training.md, docs/ai/mos-corpora.md, docs/adr/0670-ai-legacy-corpus-extraction-manifests.md, docs/research/0690-ai-legacy-corpus-extraction-manifests.md, changelog.d/added/0670-ai-legacy-corpus-extraction-manifests.md, docs/rebase-notes.md (this entry).
ADR-0673 — Saliency materializer batch manifest¶
Saliency table refresh impact. This adds a batch orchestration layer over ai/scripts/materialize_saliency_features.py. Rebase work that changes SaliencyMaterializeConfig, saliency status values, or table read/write semantics must update both the single-table script and the batch manifest runner/tests together.
Key invariants:
ai/scripts/batch_materialize_saliency_features.pymust import and reuse the single-table materializer functions; it must not duplicate FFmpeg decode, ffprobe fallback, saliency inference, or row status semantics.- Batch manifests carry shared defaults plus per-table overrides. Relative paths resolve from the manifest directory unless
--base-diris supplied. - Batch reports use schema
saliency-materializer-batch-v1and include ADR-0661run_provenance.
Touched files: ai/scripts/batch_materialize_saliency_features.py, ai/tests/test_batch_materialize_saliency_features.py, ai/AGENTS.md, docs/ai/saliency-feature-materializer.md, docs/adr/0673-saliency-materializer-batch-manifest.md, docs/research/0693-saliency-materializer-batch-manifest.md, changelog.d/added/0673-saliency-materializer-batch-manifest.md, docs/rebase-notes.md (this entry).
ADR-0674 — Second-opinion materializer batch manifest¶
Second-opinion table refresh impact. This adds a batch orchestration layer over ai/scripts/materialize_second_opinion_features.py. Rebase work that changes score-sidecar parsing, join-key policy, missing-score semantics, or run-provenance fields must update both the single-table joiner and the batch manifest runner/tests together.
Key invariants:
ai/scripts/batch_materialize_second_opinion_features.pymust import and reusematerialize_second_opinion_features.materialize(); external scorer execution remains outside this repo.- Batch manifests carry shared defaults plus per-table overrides. Relative paths resolve from the manifest directory unless
--base-diris supplied. - Batch reports use schema
second-opinion-materializer-batch-v1and include ADR-0661run_provenance.
Touched files: ai/scripts/batch_materialize_second_opinion_features.py, ai/tests/test_batch_materialize_second_opinion_features.py, ai/AGENTS.md, docs/ai/second-opinion-features.md, docs/adr/0674-second-opinion-materializer-batch-manifest.md, docs/research/0694-second-opinion-materializer-batch-manifest.md, changelog.d/added/0674-second-opinion-materializer-batch-manifest.md, docs/rebase-notes.md (this entry).
ADR-0675 — MOS label materializer batch manifest¶
MOS-labelled table refresh impact. This adds a batch orchestration layer over ai/scripts/materialize_mos_labels.py. Rebase work that changes MOS column inference, key-normalisation, match-rate enforcement, overwrite policy, or run-provenance fields must update both the single-table materializer and the batch manifest runner/tests together.
Key invariants:
ai/scripts/batch_materialize_mos_labels.pymust import and reusematerialize_mos_labels.materialize(); it must not parse MOS rows, extract features, or train models.- Batch manifests carry shared defaults plus per-table overrides. Relative paths resolve from the manifest directory unless
--base-diris supplied. - Batch reports use schema
mos-label-materializer-batch-v1and include ADR-0661run_provenance.
Touched files: ai/scripts/batch_materialize_mos_labels.py, ai/tests/test_batch_materialize_mos_labels.py, ai/AGENTS.md, docs/ai/mos-label-materializer.md, docs/adr/0675-mos-label-materializer-batch-manifest.md, docs/research/0695-mos-label-materializer-batch-manifest.md, changelog.d/added/0675-mos-label-materializer-batch-manifest.md, docs/rebase-notes.md (this entry).
ADR-0679 — CI draft auto-merge gate¶
Merge-train safety impact. The single required branch-protection context, Required Checks Aggregator, must not be skipped on draft PRs. It now runs on drafts and fails intentionally so GitHub cannot treat a draft-era skipped check as sufficient for auto-merge after the PR is marked ready.
Key invariants:
- Keep
required-aggregator.ymlfree of a job-level draft skip. Expensive sibling workflows may still skip draft PRs, but the required aggregate status must fail drafts and rerun onready_for_review. - The aggregator ignores sibling check runs older than the current workflow registration window and chooses the newest run per check name. This prevents stale draft-era skipped checks on the same SHA from masking ready-run checks.
- The ADR collision guard phase 1 compares against
BASE_SHA, not liveorigin/master, so a fast post-merge workflow cannot self-collide against the PR's own ADR.
Touched files: .github/workflows/required-aggregator.yml, .github/workflows/rule-enforcement.yml, .github/AGENTS.md, docs/adr/0679-ci-draft-automerge-gate.md, docs/research/0699-ci-draft-automerge-gate.md, changelog.d/fixed/0679-ci-draft-automerge-gate.md, docs/rebase-notes.md (this entry).
ADR-0680 — Shared AI CLI helper pattern¶
AI script helper impact. Batch manifest runners now share parser and raw argv boilerplate through aiutils.cli_helpers. Rebase work that changes the standard manifest/report/fail-fast flags should update the helper and all batch runner tests together instead of editing each runner independently.
Key invariants:
collect_cli_argv()is the canonical raw-argument capture for ADR-0661 provenance in scripts that accept an injectableargv.add_batch_manifest_arguments()owns--manifest,--base-dir,--report-json,--report-md,--fail-fast, and optional--allow-row-failuresfor batch manifest runners.- Table-specific manifest schemas and materializer semantics stay in the individual runner modules.
Touched files: ai/src/aiutils/cli_helpers.py, ai/scripts/batch_materialize_saliency_features.py, ai/scripts/batch_materialize_second_opinion_features.py, ai/scripts/batch_materialize_mos_labels.py, ai/tests/test_cli_helpers.py, ai/AGENTS.md, ai/src/aiutils/AGENTS.md, .claude/skills/ai-run-manifest/SKILL.md, docs/ai/training.md, docs/adr/0680-ai-cli-helper-pattern.md, docs/research/0700-ai-cli-helper-pattern.md, changelog.d/added/0680-ai-cli-helper-pattern.md, docs/rebase-notes.md (this entry).
ADR-0681 — AI script bootstrap helper¶
AI script import impact. Directly executable ai/scripts/*.py files now use ai/scripts/_script_bootstrap.py::bootstrap_ai_script(__file__) for repo-local imports before they import aiutils, sibling materializers, or vmaf-tune helpers.
Key invariants:
aiutilsmust remain free of startup path mutation; the bootstrap lives inai/scriptsbecause it has to run beforeai/srcis importable.- New ad hoc
sys.path.insert(...)blocks in AI scripts should be avoided. If a script needs a new repo-local root, extend_script_bootstrap.pyandai/tests/test_script_bootstrap.py. - The helper only owns import roots; artifact schemas, materializer rules, and report contents stay in the individual scripts.
Touched files: ai/scripts/_script_bootstrap.py, ai/scripts/batch_materialize_saliency_features.py, ai/scripts/batch_materialize_second_opinion_features.py, ai/scripts/batch_materialize_mos_labels.py, ai/scripts/enrich_k150k_parquet_metadata.py, ai/scripts/combine_full_feature_parquets.py, ai/scripts/extract_k150k_features.py, ai/tests/test_script_bootstrap.py, ai/AGENTS.md, ai/src/aiutils/AGENTS.md, .claude/skills/ai-run-manifest/SKILL.md, docs/ai/training.md, docs/adr/0681-ai-script-bootstrap-helper.md, docs/research/0701-ai-script-bootstrap-helper.md, changelog.d/changed/0681-ai-script-bootstrap-helper.md, docs/rebase-notes.md (this entry).
fix/mcp-cjson-banned-functions (ADR-0683)¶
No upstream rebase impact: core/src/mcp/3rdparty/cJSON/ is fork-local; upstream Netflix/vmaf does not vendor cJSON. There is no rebase conflict risk from the Netflix side.
Invariant: if this directory is synced to a newer cJSON upstream release, verify that no banned functions (sprintf, strcpy) have been re-introduced, and re-apply the fixes documented in ADR-0683. The AGENTS.md in this directory carries the exact grep command to check.
Smoke: ninja -C build && meson test -C build --suite=fast (no dedicated cJSON unit test; the MCP smoke covers the JSON paths).
Touched files: core/src/mcp/3rdparty/cJSON/cJSON.c, core/src/mcp/3rdparty/cJSON/AGENTS.md, docs/adr/0683-cjson-banned-function-remediation.md, docs/adr/README.md, changelog.d/fixed/0683-mcp-cjson-banned-functions.md, docs/rebase-notes.md (this entry).
2026-05-21 follow-up — contract-noise filter widening¶
No upstream Netflix C-source rebase impact. This stays within ADR-0659's scanner-precision policy.
Key invariant: suppress only context-bound false positives: optional-backend contracts that name HAVE_*, enable_*=false, missing loader/runtime, or CPU fallback; unit-test stub prose; and ADR allocator .md.stub reservation wording. Non-implementation "stub" uses such as Python type-stub packages, driver-stub diagnostics, and ABI-pinning disabled-build stub comments are also filtered. Do not add broad file-level allowlists.
Touched files: scripts/dev/project_modernization_audit.py, scripts/dev/test_project_modernization_audit.py, docs/development/project-modernization-audit.md, docs/research/0685-modernization-audit-contract-noise.md, scripts/AGENTS.md, changelog.d/fixed/0685-modernization-audit-contract-noise.md, docs/rebase-notes.md (this entry).
ADR-0682 — Tiny-AI Netflix corpus training scaffold — 2026-05-22 prep scope¶
- ADR: ADR-0682.
- Upstream source: fork-local. Netflix/vmaf has no tiny-AI training surface.
-
Branch:
ai/tiny-netflix-training-scaffold. Key invariants: -
Data path is local-only.
.workingdir2/netflix/is gitignored; YUV files are never committed. Every training script must accept--data-root(or theVMAF_DATA_ROOTenvironment variable) as the sole corpus entry point. - Branch name is the routine's idempotency key. Once
ai/tiny-netflix-training-scaffoldexists on origin, the daily prep-scaffolding routine exits silently. Do not rename or delete the branch until the follow-up architecture-selection PR has merged. - Netflix golden pairs are held-out only. The 3 pairs in
python/test/resource/yuv/(seeCLAUDE.md §8) are correctness gates; they are never used as training data. - Architecture selection is deferred. ADR-0682 and ADR-0242 document the alternatives table but do not pick an architecture. The follow-up PR must resolve questions (A), (B), (C) from ADR-0242 before any training run. Touched files:
docs/adr/0682-tiny-ai-netflix-training-scaffold-2026-05-22.md,docs/adr/_index_fragments/0682-tiny-ai-netflix-training-scaffold-2026-05-22.md,docs/research/0706-tiny-ai-netflix-training-prep-2026-05-22.md,changelog.d/added/0682-tiny-ai-netflix-training-scaffold-2026-05-22.md,docs/rebase-notes.md(this entry).
feat/bindings-rust-vmafx-sys (ADR-0706) — fork-only Rust crate, no Netflix upstream impact¶
No upstream rebase impact: bindings/rust/vmafx-sys, the root Cargo.toml, .github/workflows/rust-ci.yml, and docs/development/rust.md are wholly fork-local. Netflix/vmaf upstream has no Rust surface; upstream cherry-picks and port-upstream-commit syncs are unaffected. The libvmaf C public headers consumed by bindgen remain at core/include/libvmaf/ (ADR-0700 path); any future upstream header change that adds or removes a symbol is handled automatically by re-running cargo build (bindgen regenerates on every build).
feat/vmafx-phase4b-distributed-platform-adr-0709 — fork-only architectural decision, no Netflix upstream impact¶
No upstream rebase impact: ADR-0709 and the Phase 4b architecture diagram (docs/architecture/phase4b-distributed-platform.md) are wholly fork-local documents. Netflix/vmaf upstream has no controller/node/operator architecture, no Go or Rust binaries, and no rclone/eBPF integration. Upstream cherry-picks and port-upstream-commit syncs are unaffected.
The C ABI break decision (Phase 4b.8) will require updating ffmpeg-patches/ when the implementation PR lands; that PR's docs/rebase-notes.md entry will detail the specific patch files affected. This umbrella ADR does not touch any C source files.
Touched files: docs/adr/0709-vmafx-phase4b-distributed-platform.md, docs/architecture/phase4b-distributed-platform.md, changelog.d/added/vmafx-phase4b-umbrella-adr.md, docs/state.md, docs/rebase-notes.md (this entry), docs/adr/README.md.
docs/research-netflix-pipeline-backlog-audit (Research-0732) — research digest only, no Netflix upstream impact¶
no rebase impact: this PR adds only docs/research/0732-netflix-pipeline-backlog-audit.md and a changelog fragment. No C sources, headers, build files, or test fixtures are touched. Netflix/vmaf upstream cherry-picks and port-upstream-commit syncs are unaffected.
Touched files: docs/research/0732-netflix-pipeline-backlog-audit.md, changelog.d/added/0732-netflix-pipeline-backlog-audit.md, docs/state.md, docs/rebase-notes.md (this entry).
refactor/cpp23-pilot-metadata-handler — no upstream Netflix conflict¶
No rebase impact. metadata_handler.c is a fork-local refactor: Netflix/vmaf upstream also has a libvmaf/src/metadata_handler.c at the same path (pre-rename). The rename to .cpp is fork-local (upstream stays .c). If an upstream commit touches libvmaf/src/metadata_handler.c, the port must:
- Apply the upstream diff content to
core/src/metadata_handler.cppmanually (the C code is still valid C++ after the conversion). - Verify the
extern "C"guards inmetadata_handler.hare not disturbed. - Rebuild and re-run
make test-netflix-goldento confirm scores unchanged.
The meson.build change (replacing the src_dir + 'metadata_handler.c' entry with the metadata_handler_cpp20_lib static lib) is entirely fork-local and has no upstream equivalent.
Touched files: core/src/metadata_handler.cpp (was metadata_handler.c), core/src/metadata_handler.h (added extern "C" guards), core/src/meson.build (isolated static lib for C++20), core/test/meson.build (updated .c -> .cpp references), docs/adr/0708-vmafx-cpp23-internals-pilot.md, docs/research/0732-vmafx-cpp23-internals-migration-plan.md, changelog.d/changed/0708-cpp23-internals-pilot.md, docs/state.md (this entry), docs/rebase-notes.md (this entry).
ADR-0707 — TAD Rust pilot (cbindgen integration) — 2026-05-28¶
- ADR: ADR-0707.
- Upstream source: fork-local. Netflix/vmaf has no Rust feature extractors.
- Branch:
feat/tad-rust-pilot
Key rebase invariants:
core/src/feature/feature_extractor.cgains#if HAVE_RUST_TADguards around thevmaf_fex_tadextern and list entry. On upstream sync, ensure these guards are preserved; do not merge the upstream version of this file without re-applying the guards.core/src/meson.buildhas acargo build --releasecustom_target and adeclare_dependencyfor the Rust archive. These are entirely fork-local additions; upstream's meson.build will not have them. The additions appear after thelibvmaf_feature_sourceslist and before thelibvmaf = library()call.tad_rust.cis compiled as a DIRECT source of thelibvmaflibrary target (not intolibvmaf_feature.a). This is an intentional architectural choice; do not move it intolibvmaf_feature_sourceson rebase.Cargo.tomlat the repo root is the workspace manifest. Upstream will never have this file; no merge conflict expected.- The
enable_rust_featuresmeson option incore/meson_options.txtis fork-local; preserve on upstream merges.Cargo.toml(repo root, new),core/meson_options.txt,core/src/meson.build,core/src/feature/feature_extractor.c,core/src/feature/tad_rust.c(new),core/src/feature/rust/tad/(new crate directory),core/test/meson.build,core/test/test_tad_rust.c(new),docs/adr/0707-vmafx-rust-pilot-feature.md,docs/metrics/tad.md,changelog.d/added/tad-rust-pilot.md,
CAMBI Python compat-layer sync v0.5 → v0.8 — 2026-05-28¶
- ADR: no ADR required — 1:1 upstream port with no fork-local divergence.
- Upstream source: Netflix/vmaf
CambiFeatureExtractorversion history through v0.8 (Research-0732 item #4). - Branch:
chore/cambi-python-v0.8-sync
Rebase notes: The fork is now at parity with upstream Netflix/vmaf for the Python CAMBI wrappers as of 2026-05-28. Future Netflix syncs of compat/python-vmaf/core/cambi_feature_extractor.py, compat/python-vmaf/core/cambi_quality_runner.py, and python/test/cambi_test.py should merge cleanly. No fork-local divergence was introduced; this was a pure upstream port.
compat/python-vmaf/core/cambi_feature_extractor.py, compat/python-vmaf/core/cambi_quality_runner.py, python/test/cambi_test.py, changelog.d/changed/cambi-python-v0.8-sync.md,
vmafx-node Go worker binary (ADR-0713)¶
no rebase impact: fork-only addition — all new files under cmd/vmafx-node/, pkg/gpu/, pkg/ai/, gen/go/controller/, docker/Dockerfile.node*, deploy/helm/vmafx/templates/node.yaml. No C sources, no upstream-mirror files touched. pkg/encoder/discover.go and pkg/encoder/hardware.go are new fork-local files; pkg/encoder/encoder.go (already fork-local from ADR-0705) is not modified.
Files added: cmd/vmafx-node/main.go (new), cmd/vmafx-node/executor.go (new), cmd/vmafx-node/main_test.go (new), gen/go/controller/controller.pb.go (new), gen/go/controller/controller_grpc.pb.go (new), pkg/gpu/detect.go (new), pkg/gpu/detect_test.go (new), pkg/ai/infer.go (new), pkg/ai/infer_test.go (new), pkg/encoder/discover.go (new), pkg/encoder/hardware.go (new), docker/Dockerfile.node (new), docker/Dockerfile.node-cpu (new), docker/Dockerfile.node-cuda12 (new), docker/Dockerfile.node-rocm6 (new), docker/Dockerfile.node-sycl-oneapi2026 (new), deploy/helm/vmafx/templates/node.yaml (new), deploy/helm/vmafx/templates/_helpers.tpl (extended), deploy/helm/vmafx/values.yaml (extended — .Values.node section added), docs/server/node.md (new), docs/adr/0713-vmafx-node-impl.md (new), changelog.d/added/vmafx-node.md (new).
Research-0733 — VMAFX eBPF optimization target — 2026-05-28¶
No rebase impact: docs-only PR. All touched files (docs/research/, changelog.d/, docs/state.md, docs/rebase-notes.md) are fork-local with no upstream Netflix/vmaf equivalent. No C source, no build system, no test assertions changed.
Touched files: docs/research/0733-vmafx-ebpf-optimization-target.md (new), changelog.d/changed/ebpf-research.md (new), docs/state.md (new row), docs/rebase-notes.md (this entry).
cmd/vmafx-operator — Kubernetes Operator kubebuilder skeleton (ADR-0714)¶
No rebase impact on upstream C/Python code: the operator is entirely fork-local (api/vmafx/v1/, cmd/vmafx-operator/, config/crd/, config/rbac/, deploy/helm/vmafx/crds/, deploy/helm/vmafx/templates/operator-*.yaml, go.mod, go.sum). None of these paths overlap with Netflix/vmaf upstream.
If a future upstream sync adds a Go module or touches go.mod, merge the dependency lists in go.mod and regenerate go.sum.
Fork-local files: api/vmafx/v1/ (new), cmd/vmafx-operator/ (new), config/crd/bases/ (new), config/rbac/role.yaml (new), deploy/helm/vmafx/crds/ (new), deploy/helm/vmafx/templates/operator-deployment.yaml (new), deploy/helm/vmafx/templates/operator-rbac.yaml (new), deploy/helm/vmafx/values.yaml (operator.* section added), docs/adr/0714-vmafx-operator-skeleton.md, docs/development/operator.md, changelog.d/added/vmafx-operator-skeleton.md,
core/src/feature/cuda/AGENTS.md — __mul24 prohibition invariant (Research-0734, 2026-05-28)¶
The 2026-05-28 audit confirmed zero __mul24 / __umul24 / __mul24hi usages in the fork's CUDA kernel tree. A prohibition invariant was added to core/src/feature/cuda/AGENTS.md. On upstream sync: if Netflix/vmaf ever adds a CUDA kernel that uses these intrinsics, the prohibiton invariant requires the caller to either remove the intrinsic (replace with *) or obtain CODEOWNERS sign-off documenting the minimum-CUDA-13.3 constraint (see the AGENTS.md note for the full acceptance criteria).
No upstream file is currently in conflict; this note exists to alert future sync agents that the invariant file was intentionally added by the fork and should be preserved through rebases.
Research-0734 — CUDA 13.3 fix-list deep audit¶
No rebase impact on upstream C/Python code: this PR is docs-only (research digest, changelog fragment, state.md row, rebase-notes entry). No C source, .cu kernel, or build file is modified.
If a future upstream sync changes dev/Containerfile or Dockerfile CUDA base-image pins, verify that the new pin is >= 13.3 to ensure the NVCC thread-reconvergence fix [6156910] is included.
Fork-local files: docs/research/0734-cuda-13.3-fix-list-deep-audit.md (new), changelog.d/changed/cuda-13.3-fix-list-deep-audit.md (new), docs/state.md (new row), docs/rebase-notes.md (this entry).
scripts/dev/cleanup-agent-state.sh — agent-state cleanup utility¶
No rebase impact on upstream C/Python code: the script is entirely fork-local developer tooling that touches no compiled sources, tests, or public API.
If a future upstream sync adds a scripts/dev/ directory, merge manually (name collision is the only risk; no logic conflict).
Fork-local files added: scripts/dev/cleanup-agent-state.sh (new), docs/development/agent-worktree-discipline.md (cleanup section added), changelog.d/added/dev-cleanup-script.md (new).
Research-0734 — CUDA VIF filter1d ncu hotpath (no rebase impact)¶
no rebase impact: pure research digest; no source files modified.
Research-0744 cross-backend baseline (2026-05-28) — no rebase impact¶
This PR adds only docs/research/0744-cuda-cross-backend-baseline-pre-ncu-perf.md, changelog.d/perf/cuda-cross-backend-baseline.md, and a docs/state.md row. No C, header, Python, or build files are modified. No upstream sync action is required.
docs/research/0734–0738 — CUDA ADM/motion/SSIM/MS-SSIM ncu hotpath profiles (2026-05-28)¶
No rebase impact: research-only documents, no source code changes. The profiling findings (Research-0734 through 0738) are advisory; no kernel modifications were made in this PR. When a follow-up PR implements the integer_ssim_score.cu extern "C" fix (Research-0736 recommendation 1), that PR must also update ssim_cuda.c host glue and verify bit-exact parity against the CPU integer_ssim extractor on the Netflix golden fixture.
C++23 wave adversarial review (2026-05-28)¶
Read-only review of PRs #41, #43, #44, #45, #48, #51, #54, #56, #58. No files were modified by this review. The review digest is in docs/research/cpp23-wave-adversarial-review-20260528.md.
Critical issues that must be fixed before merge:
- PR #43
opt.cpp:strtol/strtodon potentially non-NUL-terminatedstring_view::data() - PR #48
dict.cpp:strtof(float) assigned todouble— precision loss on option values - PR #54
model.cpp:strlen(model->name) - 5Uunsigned underflow → heap overflow - PR #58
ref.cpp:make_unique/ C-callerfree()allocator mismatch
No rebase impact from the review itself; all findings are fixes required in those PRs.
core/src/feature/cuda/integer_ssim/ — extern "C" on new kernels (ADR-0747)¶
Any upstream or fork PR that adds a new __global__ kernel to a .cu file under core/src/feature/cuda/ or core/src/cuda/ must wrap the entry point in extern "C" { } if it is also referenced by cuModuleGetFunction in the host .c glue.
The invariant is enforced by scripts/dev/check-cuda-extern-c.sh. Run it locally before pushing. On upstream sync, if Netflix adds new CUDA kernels to their libvmaf/src/feature/cuda/ tree, check whether those kernels use extern "C" in the upstream source and mirror the pattern here.
This invariant was formalised after the audit that found integer_ssim/integer_ssim_score.cu missing extern "C", silently breaking --feature ssim --backend cuda since introduction (PR #77 fixed the analogous break in ssim_score.cu; ADR-0747 fixes integer_ssim_score.cu).
core/src/feature/cuda/integer_vif/filter1d.cu — ADR-0743 launch_bounds + __ldg¶
No rebase-sensitive invariants for downstream callers — the changes are confined to the device-side kernel body and the FILTER1D_8_HORI macro. The symbol name filter1d_8_horizontal_kernel_2_17_9 is unchanged; the C host file integer_vif_cuda.c continues to load and dispatch it by name.
If an upstream Netflix/vmaf sync introduces changes to filter1d.cu:
- The
__launch_bounds__(128, 10)annotation onFILTER1D_8_HORImust be preserved (or re-applied) — upstream does not carry this hint. - The
__ldg()calls on the 7buf.tmp.*loads must be preserved. - If upstream changes
val_per_threadorHORI_TILE_W, recheck the smem budget constraint (14812 B/block at vpt=4 is smem-limited on sm_89 — see ADR-0743 for derivation). - The ptxas advisory "minnctapersm out of range, ignored" for sm_75/sm_80/ sm_86 is expected and benign; do not treat it as a gate failure.
research-0748 / PR #76 1080p re-measurement — no new rebase invariants¶
The 1080p re-measurement (research-0748) validates PR #76 at production resolution. No new rebase-sensitive invariants beyond those already documented in the ADR-0743 __launch_bounds__ + __ldg entry above. The register budget (48 regs/thread) and __ldg annotations must be preserved on any upstream sync that touches filter1d.cu per the existing note.
One-off container SYCL device-access pattern (--device /dev/dri --group-add 988)¶
No rebase impact on upstream C/Python code.
When running vmaf-dev-mcp:cuda13.3 as a one-off docker run with SYCL needed:
--device /dev/driis not sufficient. The Level Zero GPU ICD requires/dev/dri/by-path/pci-XXXX:YY:ZZ.W-rendersymlinks to enumerate Intel devices. These symlinks are not passed by--device /dev/dri; they require an explicit-v /dev/dri/by-path:/dev/dri/by-path:robind-mount.--group-add renderfails becauserenderis not a group name inside the container. Use--group-add 988(the host render GID, confirmed on this machine).- Source
setvars.shinside the container before invokingsycl-lsorvmaf --backend sycl.
The docker compose deployment (dev/docker-compose.yml) already carries the by-path bind-mount per ADR-0514; this note covers one-off docker run usage.
Fork-local files: docs/research/0734-cross-backend-baseline-with-sycl-20260528.md (new), changelog.d/changed/cross-backend-baseline-with-sycl.md (new), docs/rebase-notes.md (this entry), docs/state.md (new row).
ADR-0752 — Multi-resolution perf benchmark baseline¶
No rebase impact on upstream C/Python code.
New fork-local files only:
scripts/perf/bench-multi-resolution.sh— benchmark harnesstestdata/perf_multi_resolution.json— baseline snapshot (schema_version=1)docs/development/perf.md— usage docsdocs/research/research-0752-perf-bench-multi-resolution-baseline.mddocs/adr/0752-perf-bench-multi-resolution.mdchangelog.d/added/perf-bench-multi-resolution.mdcore/AGENTS.md(new invariant appended)
The upscaled fixture cache files (testdata/ref_1920x1080_48f.yuv, etc.) are generated on first run and should be .gitignored (they are reproducible from the 576×324 native fixture via ffmpeg -vf scale=W:H:flags=bilinear).
perf/cuda-ssim-vert-combine-ldg-launch-bounds-leak-20260529 (ADR-0754)¶
No rebase impact on upstream C/Python code.
core/src/feature/cuda/integer_ssim/ssim_score.cu and core/src/feature/cuda/integer_ssim_cuda.c are wholly fork-local files with no upstream Netflix equivalents. The VmafCudaBuffer struct and the vmaf_cuda_kernel_readback_free / vmaf_cuda_buffer_host_free helpers are fork-local CUDA infrastructure. No Netflix upstream commit will collide with these changes on sync-upstream.
Fork-local files modified: core/src/feature/cuda/integer_ssim/ssim_score.cu (F2 + F4 — ldg() + __launch_bounds), core/src/feature/cuda/integer_ssim_cuda.c (F6 per-caller save+free DROPPED — superseded by helper fix in PR #94), core/src/feature/cuda/AGENTS.md (invariant notes), docs/adr/0754-cuda-ssim-vert-combine-ldg-pinned-leak.md (new), docs/adr/README.md (new row), docs/research/0754-cuda-ssim-vert-combine-ldg-launch-bounds-2026-05-29.md (new), changelog.d/perf/cuda-ssim-vert-combine.md (new), docs/rebase-notes.md (this entry), docs/state.md (new row).
Research-0755 — HIP backend audit (2026-05-29)¶
No rebase impact on upstream C/Python code.
All files modified are fork-local: core/src/feature/hip/AGENTS.md (invariant notes), docs/research/0755-hip-backend-audit-20260529.md (new), changelog.d/changed/hip-backend-audit.md (new), docs/rebase-notes.md (this entry), docs/state.md (new row).
No source files were modified (audit-only). No Netflix upstream commit will collide with these additions on sync-upstream.
research/cuda-f3-struct-by-value-audit-20260529 (2026-05-29)¶
No rebase impact: this PR adds documentation-only files (research digest, ADR, changelog fragment, state.md row). No CUDA source files are modified. VmafCudaBuffer, VmafPicture, and AdmBufferCuda definitions are unchanged; no upstream Netflix/vmaf commit will collide with this PR's diff.
Fork-local files added/modified: docs/research/research-0756-cuda-f3-struct-by-value-audit.md (new), docs/adr/0756-cuda-f3-struct-by-value-audit.md (new), docs/adr/README.md (new row), changelog.d/perf/cuda-f3-struct-by-value-audit.md (new), docs/state.md (new row),
ADR-0755: C++23 Wave 7 — activate cpu.cpp (PR on 2026-05-29)¶
No rebase impact on upstream C/Python code.
core/src/cpu.c was deleted and core/src/meson.build updated to compile cpu.cpp. The file cpu.cpp is wholly fork-local (no upstream Netflix equivalent). No Netflix upstream commit will collide with this deletion.
Fork-local files modified: core/src/cpu.c (deleted), core/src/meson.build (cpu.c → cpu.cpp in libvmaf_cpu_sources), docs/adr/0755-cpp23-wave7-single-file.md (new), docs/adr/README.md (new row), changelog.d/changed/0755-cpp23-wave7-cpu-cpp.md (new), docs/rebase-notes.md (this entry).
research/cuda-motion-ncu-profile-20260529¶
No rebase impact: research-only commit. No source files modified. Files added: docs/research/0760-cuda-motion-ncu-multi-resolution-20260529.md, changelog.d/perf/cuda-motion-ncu-multi-resolution.md, docs/adr/0760-cuda-motion-ncu-multi-resolution.md (research ADR). No upstream collision risk.
HIP ADM buffer-by-pointer refactor (ADR-0759, 2026-05-29)¶
Files touched: core/src/feature/hip/integer_adm/adm_csf.hip, core/src/feature/hip/integer_adm/adm_cm.hip, core/src/feature/hip/integer_adm_hip.c, core/src/feature/hip/AGENTS.md
Rebase impact: None. All touched files are fork-added; no upstream Netflix/vmaf file is modified. The HIP backend does not exist in upstream. No rebase conflict is possible with upstream syncs.
The changed kernel signatures are internal to the HIP dispatch path and are not part of any public API.
ADR-0762 — CUDA CIEDE2000 __ldg() F3 fix (2026-05-29)¶
No rebase impact on upstream C/Python code.
All files modified are fork-local: core/src/feature/cuda/integer_ciede/ciede_score.cu (F3 fix — __ldg + __launch_bounds), core/src/feature/cuda/integer_vif_cuda.c (resolve pre-existing merge-conflict stub from 24bb5daf89), docs/adr/0762-cuda-ciede-ldg.md (new), changelog.d/perf/cuda-ciede-ldg.md (new), docs/rebase-notes.md (this entry), docs/state.md (new row).
ciede_score.cu is entirely fork-local (Netflix upstream has no CUDA ciede kernel). A sync-upstream that adds a CUDA ciede kernel upstream would need to incorporate this __ldg() pattern. The integer_vif_cuda.c conflict resolution keeps the HEAD side (ADR-0743 comment block); no Netflix upstream content was discarded.
ADR-0764 — psnr_hvs CUDA kernel F3 ldg() + __launch_bounds(64) (2026-05-29)¶
No rebase impact on upstream C/Python code: psnr_hvs_score.cu is entirely fork-local. integer_psnr_hvs_cuda.c is unchanged.
If an upstream sync changes the psnr_hvs CPU reference in core/src/feature/third_party/xiph/psnr_hvs.c, verify the CUDA kernel's cooperative tile load and reduction order in psnr_hvs_score.cu are still byte-for-byte equivalent to the CPU's calc_psnrhvs computation pattern.
All files modified are fork-local: core/src/feature/cuda/integer_psnr_hvs/psnr_hvs_score.cu (pointer extraction + __ldg() + __launch_bounds__(64)), docs/adr/0764-psnr-hvs-ldg-launch-bounds.md (new), docs/research/0764-cuda-psnr-hvs-ldg-launch-bounds-2026-05-29.md (new), changelog.d/perf/cuda-psnr-hvs-ldg-launch-bounds.md (new), docs/rebase-notes.md (this entry), docs/state.md (new row).
ADR-0787 — libvmaf API error-path audit (2026-05-29)¶
No rebase impact: this PR adds only documentation files (research digest, ADR, changelog fragment) and no C/Python source changes.
All files modified are fork-local: docs/research/research-0787-libvmaf-api-error-path-audit.md (new), docs/adr/0787-libvmaf-api-error-path-audit.md (new), docs/adr/README.md (new row), changelog.d/fixed/0787-libvmaf-api-error-path-audit.md (new), docs/rebase-notes.md (this entry).
The six implementation fixes recommended by the audit (vmaf_write_output_with_format errno, vmaf_cuda_state_init error codes, vmaf_close unchecked returns, CUDA EBUSY guard, vmaf_init error propagation) will land in a separate fix PR that will carry its own rebase-notes entry. The vmaf_cuda_state_free ABI-normalisation is deferred to a major-version PR.
ADR-0815 — vmafx-operator + vmafx-node distroless Dockerfiles (2026-05-29)¶
No rebase impact on upstream C/Python code.
All files added are fork-local: docker/Dockerfile.operator (new), .github/workflows/docker-publish-operator-node.yml (new), docs/adr/0815-operator-node-distroless-dockerfiles.md (new), changelog.d/added/0815-operator-node-distroless-dockerfiles.md (new), docs/backends/operator.md (new), docs/rebase-notes.md (this entry).
No upstream Netflix/vmaf files are touched. A sync-upstream cannot conflict with these additions. The docker/Dockerfile.node file was already in-tree (ADR-0717); this PR only adds the CI workflow that publishes it.
no rebase impact: REASON — changes are confined to config files (.clang-tidy, .pre-commit-config.yaml, pyproject.toml), fork-owned Python sources in ai/ and scripts/ (UP auto-fixes), and docs. No upstream Netflix/vmaf C source is touched; the HeaderFilterRegex fix has no effect on any upstream file.
ADR-0795 — prev_ref thread-safety hardening — 2026-05-29¶
No rebase impact: all changes are in core/src/libvmaf.c (comments, a rename from fex to shared_fex, and a defensive assert). No logic change; no new symbols; no API change. The modified functions (threaded_extract_func, threaded_extract_batch_func) are fork-local dispatch paths not present in upstream Netflix/vmaf.
Fork-local files: core/src/libvmaf.c (comments + assert), docs/adr/0795-prev-ref-thread-safety.md, changelog.d/fixed/prev-ref-batch-thread-safety.md.
ADR-0882 — fuzz target audit (json_model + dnn_sidecar) — 2026-05-30¶
no rebase impact: REASON — all new files (core/test/fuzz/fuzz_json_model.c, core/test/fuzz/fuzz_dnn_sidecar.c, seed corpora under core/test/fuzz/json_model_corpus/ + core/test/fuzz/dnn_sidecar_corpus/, the known-crash reproducer under core/test/fuzz/json_model_known_crashes/, and ADR-0882 + changelog fragment) are fork-local. Upstream Netflix/vmaf has no libFuzzer harnesses at all (the entire core/test/fuzz/ subtree is fork-added per ADR-0270 + ADR-0311). The core/test/fuzz/meson.build edits sit in a if not get_option('fuzz') guarded subdir that upstream does not descend into. The .github/workflows/fuzz.yml matrix addition extends a fork-only workflow file. The only files touching shared upstream-mirror code are doc edits (docs/state.md, docs/rebase-notes.md, docs/adr/README.md) that always paint the fork-local row pattern.
ADR-0887 — vmaf_model_destroy slopes-OOB fix — 2026-05-30¶
Low rebase impact, but not zero. Touches two upstream-mirrored files:
core/src/read_json_model.c— addssync_n_featureshelper, replacesparse_feature_names' unconditionaln_features++with a per-iterationsync_n_features(model, i)call, adds the same call toparse_slopes/parse_intercepts/parse_feature_opts_dicts, and insertsvalidate_feature_arraysbeforeparse_model_dictreturns.core/src/model.c::vmaf_model_destroy— flips the destroy walk bound frommax(feature_cap, n_features)tomin(feature_cap, n_features).
On upstream sync, if Netflix has independently changed parse_feature_names or vmaf_model_destroy, take the upstream changes for unrelated lines and re-apply this fork's hunks (the new sync_n_features helper, the validate_feature_arrays call, and the min bound in destroy). Upstream Netflix does not currently have this validation pass, so a conflict means upstream changed an adjacent surface — re-applying the fork's hunks post-upstream is mechanical.
Fork-local files (no rebase impact): docs/adr/0887-*.md, docs/research/0887-*.md, core/test/test_model.c regression tests, changelog.d/fixed/vmaf-model-destroy-slopes-oob.md, docs/state.md row.
Feature-extractor coverage round 2 (ADR-0938, 2026-05-31)¶
no rebase impact: REASON — all seven new files (core/test/test_integer_psnr_coverage.c, core/test/test_integer_motion_coverage.c, core/test/test_integer_motion_v2_coverage.c, core/test/test_integer_vif_log2.c, core/test/test_iqa_convolve_coverage.c, core/test/test_barten_csf_coverage.c, core/test/test_ms_ssim_decimate_coverage.c) are fork-local additions under core/test/ and seven additive blocks in core/test/meson.build that do not touch any upstream-mirrored test file. The only contact surface with upstream is the consumed public C-API and the public feature/integer_vif.h / feature/barten_csf_tools.h / feature/iqa/convolve.h headers, which are upstream-mirrored but read-only from these tests. On upstream sync, conflicts are restricted to the meson.build insertion points; reapply the seven test_*_coverage = executable(...) blocks and the matching test('test_*_coverage', ...) rows post-rebase.
Feature-extractor coverage round 3 (ADR-0948, 2026-05-31)¶
no rebase impact: REASON — additions are confined to fork-local test binaries under core/test/ (test_integer_motion_edge16_coverage, test_adm_csf_tools_coverage, test_feature_collector_coverage) and three append-only entries in core/test/meson.build. No upstream-mirrored source touched; no public API delta. On upstream sync the new tests apply cleanly regardless of what Netflix does to the underlying production files because the tests link against the existing libvmaf static target and import public + internal headers that already existed before round 3.
SYCL kernel coverage round 2 (ADR-0884, 2026-05-30)¶
no rebase impact: REASON — all changes are confined to fork-added test files (core/test/test_sycl_adm_parity.c, core/test/test_sycl_ciede_parity.c, core/test/test_sycl_ssim_parity.c, core/test/test_sycl_ms_ssim_parity.c, core/test/test_sycl_motion_v2_parity.c), the meson wiring for those files in core/test/meson.build, and docs / changelog / core/src/feature/sycl/AGENTS.md companion notes. No upstream Netflix/vmaf C source is touched. The SYCL backend itself is fork-original (Netflix/vmaf has no SYCL path), so there is no upstream rebase surface for these tests at all.
CUDA kernel parity tests — round 2 (ADR-0886, 2026-05-30)¶
no rebase impact: REASON — adds five new fork-local test files under core/test/ (test_cuda_adm_parity.c, test_cuda_motion_v2_parity.c, test_cuda_cambi_parity.c, test_cuda_psnr_hvs_parity.c, test_cuda_ssim_parity.c) and wires them through core/test/meson.build inside the existing if cuda_dependency.found() guard. The tests exercise the public C API (vmaf_init / vmaf_use_feature / vmaf_cuda_state_init / vmaf_feature_score_at_index); upstream Netflix/vmaf does not own any of the touched files. Conflict surface on sync is limited to the core/test/meson.build stanza ordering, which is mechanical.
Fork-local files: core/test/test_cuda_adm_parity.c, core/test/test_cuda_motion_v2_parity.c, core/test/test_cuda_cambi_parity.c, core/test/test_cuda_psnr_hvs_parity.c, core/test/test_cuda_ssim_parity.c, core/test/meson.build (new stanzas only), docs/adr/0886-cuda-kernel-coverage-round2.md, docs/research/cuda-kernel-coverage-round2-2026-05-30.md, changelog.d/added/0886-cuda-kernel-coverage-round2.md.
macOS CI ansnr-residual cleanup (ADR-0749 follow-up, 2026-05-30)¶
no rebase impact: REASON — changes are confined to fork-mirrored upstream test files (python/test/feature_extractor_test.py, python/test/quality_runner_test.py, python/test/routine_test.py) where assertions referencing the legacy ansnr / anpsnr keys are dropped or the tests are skipped per ADR-0749 (ansnr feature sunset). The @unittest.skip reasons cite ADR-0749, so on upstream sync the conflict resolution is mechanical: if Netflix upstream still has the legacy assertions they were calibrated against float_ansnr output that this fork no longer produces — keep the skips. If Netflix upstream removes the legacy assertions themselves (matching this fork's direction), drop the local skips.
CI scripts: rebrand-proof assertion-density + tempfile trap (ADR-0968, 2026-05-31)¶
no rebase impact: fork-local — scripts/ci/assertion-density.sh and scripts/release/concat-changelog-fragments.sh are entirely fork-introduced; Netflix upstream has no equivalent files in either path. The only rebase risk is a new upstream scripts/ entry shadowing the directory, which would surface as an explicit conflict rather than a silent behaviour change.
compat/python-vmaf leaf-utility coverage (2026-05-31)¶
no rebase impact: REASON — the new test file lives entirely under python/test/compat_python_vmaf_coverage_test.py (fork-local test directory that Netflix upstream never touches) and imports leaf utilities by their existing public names. No production module under compat/python-vmaf/ is modified; only compat/python-vmaf/AGENTS.md gains one paragraph documenting which leaves carry coverage tests and warning about the latent sha1 bug in tools/decorator.py's persist helpers. Upstream syncs do not own compat/python-vmaf/AGENTS.md (fork-only file).
Master CI regressions — Metal MS-SSIM fixture + ssimulacra2 icpx XYB (ADR-0973, 2026-05-31)¶
no rebase impact: REASON — all touched files are fork-additions with no upstream conflict surface:
core/test/test_metal_float_ms_ssim_parity.c— fork-added in T8-2a; Netflix upstream has no Metal backend.core/test/test_ssimulacra2_simd.c— fork-added SIMD bit-exactness test; Netflix upstream has no SSIMULACRA 2 SIMD paths.docs/adr/0973-*.md,docs/research/0973-*.md,changelog.d/fixed/0973-*.md,core/test/AGENTS.md— fork-only governance / docs.
The fix adds a file-scope #pragma clang fp contract(off) block to test_ssimulacra2_simd.c. If a future contributor refactors the file's scalar reference functions out into a helper header, the pragma block must move with them or the icx FMA contraction returns and test_xyb fails under the all-backends matrix leg.
test_gpu_picture_pool.c Round 27 D.3 + D.4 cleanup (ADR-0970, 2026-05-31)¶
no rebase impact: REASON — core/test/test_gpu_picture_pool.c is a fork-local test file (it was introduced in this fork's PR #266 / ADR-0239; Netflix upstream has no equivalent file). The two changes (remove unused .state malloc, delete dead /* ... */ block) affect only lines that Netflix upstream never touches. core/test/AGENTS.md is also fork-only.
ADR-0922 — coverage ratchet + per-PR delta gate — 2026-05-31¶
No rebase impact. All touched files are fork-local CI / docs infrastructure:
scripts/ci/coverage-check.sh(raisedOVERALL_MIN37 → 60,CRITICAL_MIN85 → 90, tightened eachPER_FILE_MINentry by +5pp).scripts/ci/coverage-delta-check.sh(new — per-PR delta gate)..github/workflows/tests-and-quality-gates.yml(Coverage Gate job: new floor numbers + two new steps that compute base-branch coverage and run the delta gate on pull-request events).docs/adr/0922-coverage-ratchet-aggressive.md,docs/adr/_index_fragments/0922-coverage-ratchet-aggressive.md,docs/adr/_index_fragments/_order.txt,docs/adr/README.md(regenerated byscripts/docs/concat-adr-index.sh).changelog.d/changed/0922-coverage-ratchet-aggressive.md.
Upstream Netflix/vmaf has no coverage gate, so on sync there is nothing to reconcile. The per-PR delta gate's fetch-depth: 0 checkout requirement is worth flagging if the workflow ever gets restructured: a shallow checkout breaks git merge-base HEAD "$BASE_REF".
Metal kernel coverage round 4 — closeout (2026-05-31, ADR-0959)¶
no rebase impact: REASON — every new file path is fork-local Metal-only (Netflix upstream has no Metal backend at all per rebase-notes.md §"feat/libvmaf-metal-filter-iosurface" lineage). The single existing-file edit, core/test/meson.build, appends one executable() + test() block inside the existing enable_metal guard introduced by ADR-0361 (no boundary change, no upstream-mirrored line touched). Upstream sync resolution is trivially "keep theirs" everywhere except inside the if metal_test_opt.enabled() … block, which is fork-only by construction.
Fork-local additions (no rebase impact): core/test/test_metal_kernel_coverage_audit.c, docs/adr/0959-metal-kernel-coverage-round4-closeout.md, docs/research/0959-metal-kernel-coverage-round4-closeout.md, changelog.d/added/metal-kernel-coverage-round4.md, the new audit row in docs/adr/README.md, the T-METAL-KERNEL-PARITY-ROUND4-2026-05-31 row in docs/state.md.
CUDA kernel parity coverage — round 4 (ADR-0956, 2026-05-31)¶
no rebase impact: REASON — all five new files (core/test/test_cuda_float_adm_parity.c, core/test/test_cuda_float_motion_parity.c, core/test/test_cuda_float_ssim_parity.c, core/test/test_cuda_speed_chroma_smoke.c, core/test/test_cuda_speed_temporal_smoke.c) live entirely under fork-local test directories that Netflix upstream never touches. The only modified shared file is core/test/meson.build, where the round 4 block is appended after the existing round 3 / ADR-0541 / motion3 parity blocks inside the existing if get_option('enable_cuda') guard. On upstream sync, if Netflix has independently added test binaries in the same enable_cuda block the conflict is a trivial append-vs-append three-way merge (no shared lines change). Fork-local documentation files (ADR-0956, the round 4 research digest, the changelog fragment, this rebase-notes row) are never authored upstream.
speed_internal.c + SpEED GPU twin wiring (ADR-0964, 2026-05-31)¶
Will bite a rebase. This PR adds core/src/feature/speed_internal.c (a fork-local TU that duplicates ~600 LOC of pure math — eigendecomposition, QR factorisation, matrix helpers — from speed.c). When /sync-upstream ports any change to speed.c's static helpers (compute_eigenvalues, matrix_qr_decomposition, solve_triangular_system, convert_to_tridiagonal, compute_eigenvalues_tridiagonal, compute_covariance_matrix, filter_and_downscale, ...), the same change must be mirrored into speed_internal.c. Symptom of drift: test_sycl_speed_chroma_parity / test_sycl_speed_temporal_parity flag a places=4 violation between CPU and SYCL on Intel Arc.
Also fork-local:
core/src/feature/hip/speed_{chroma,temporal}_hip.c(already in tree, newly wired intocore/src/hip/meson.build).core/src/feature/sycl/speed_{chroma,temporal}_sycl.cpp(already in tree, newly wired intocore/src/meson.buildsycl_feature_sources).core/src/feature/feature_extractor.cexterns + registry rows for the four new GPU extractor symbols, gated on#if HAVE_HIP/#if HAVE_SYCL.core/test/test_sycl_speed_chroma_parity.c+core/test/test_sycl_speed_temporal_parity.c.
If Netflix upstream ever ships its own SpEED GPU implementation that takes a different code-sharing approach (e.g. exposing speed.c helpers via a non-static-prefix), the fork should consider migrating to upstream's pattern; until then speed_internal.c is the canonical location for the shared helpers and the GPU TUs depend on its function names.
CUDA twins (speed_chroma_cuda, speed_temporal_cuda) are NOT wired in this PR — the TUs reference symbols (CHECK_CUDA, CudaFunctions->cuMemAllocHost) that do not exist; they need a repair pass. Tracked as T-CUDA-SPEED-TU-REPAIR-2026-05-31 in docs/state.md.
CUDA SpEED TU repair + wiring (ADR-0965, 2026-05-31)¶
No new rebase risk beyond ADR-0964. This repair PR fixes the latent bugs in speed_chroma_cuda.c and speed_temporal_cuda.c and wires them into meson. The changes are:
CHECK_CUDA(cu_f, CALL)replaced withCHECK_CUDA_GOTO(cu_f, CALL, fail)throughout both TUs.cuMemAllocHost(ptr, sz)replaced withcuMemHostAlloc(ptr, sz, 0x01u)in theALLOC_HOSTmacros (both TUs).
Both changes are mechanical; no algorithmic content was altered. The rebase-note from ADR-0964 above covers the speed_internal.c drift risk (mirror fixes between speed.c and speed_internal.c).
New additions:
core/src/feature/feature_extractor.cexterns + registry rows forvmaf_fex_speed_chroma_cuda/vmaf_fex_speed_temporal_cudaunder#if HAVE_CUDA.core/test/test_cuda_speed_chroma_parity.c+core/test/test_cuda_speed_temporal_parity.c.
Go errors.Join cleanup paths + slog key standardisation (ADR-0935, 2026-05-31)¶
no rebase impact: REASON — every file touched lives in fork-original Phase 4b Go subtree (pkg/bisect/, pkg/encoder/, pkg/storage/, cmd/vmafx-controller/queue/, cmd/vmafx-node/). Netflix upstream ships no Go code under these paths, so a future upstream/master sync cannot conflict here. The cmd/vmafx-tune/AGENTS.md invariant addition is also fork-original. If a follow-up port-PR introduces upstream Go code, the errors.Join discipline documented in cmd/vmafx-tune/AGENTS.md §7 applies on entry.
Generic registry for vmafx-controller (ADR-0925, 2026-05-31)¶
no rebase impact: REASON — touched files are 100 % fork-only Go sources (pkg/registry/registry.go, pkg/registry/registry_test.go, cmd/vmafx-controller/nodes/registry.go, pkg/observability/observability.go). Netflix upstream is a pure C / Python tree; the cmd/ and pkg/ Go trees do not exist there.
VmafPicture v2 design scaffold (ADR-0928, 2026-05-31)¶
Files touched: core/include/libvmaf/picture_v2.h (new), docs/adr/0928-vmaf-picture-v2-explicit-backend-state.md (new), docs/architecture/vmaf-picture-v2-migration.md (new), docs/adr/README.md + docs/adr/_index_fragments/ (index row), changelog.d/added/vmaf-picture-v2-design.md (fragment).
Rebase impact: None for this PR. The new header is declared but not yet wired into meson.build, and v1 (core/include/libvmaf/picture.h) is preserved bit-for-bit — every existing consumer (FFmpeg patches 0002–0006, MCP server, Rust binding scaffold, Python wheels) still sees the v1 surface unchanged. Upstream Netflix/vmaf has no v2 counterpart on the deprecation horizon, so no sync conflict is expected.
Lifecycle (per ADR-0928):
- Cycle N (this PR): header declared, design + scaffold only.
- Cycle N+1: header wired into meson, converters implemented in
core/src/picture.c, v1 marked__attribute__((deprecated)). - Cycle N+2: in-tree backends +
ffmpeg-patches/0002-0006switched to v2 (coordinated per CLAUDE.md §12 r14). - Cycle N+3 (≈ 12 months, target VMAFX v4.0.0): v1 removed, SONAME bump
libvmaf.so.3 → .4.
If upstream Netflix independently adds a VmafPicture v2 of their own before cycle N+3, reconcile by adopting upstream's naming (VmafPicture2 is intentionally generic) and remap our converters; otherwise the cycle-N+3 v1-removal commit is the natural ABI break window.
pathlib sweep + ruff PTH guard (ADR-0936, 2026-05-31)¶
no rebase impact: changes are confined to fork-owned Python — the two console-shim files under tools/vmaf-*/ (fork-added, no upstream twin), fork-owned ai/scripts/, ai/src/corpus/, mcp-server/, scripts/ci/, and tools/vmaf-tune/src/ modules. The pyproject.toml ruff config delta adds PTH to select and lists it in the existing per-file ignores for the upstream-mirror trees (python/**, compat/python-vmaf/**, testdata/**). Upstream Netflix Python is covered by those ignores; an upstream sync will not see the PTH rule applied to their files.
iter.Seq[T] companion APIs for Go packages (ADR-0932, 2026-05-31)¶
no rebase impact: REASON — every touched file is fork-original Go code under pkg/bisect/, pkg/ladder/, pkg/ai/, and cmd/vmafx-controller/nodes/. None of these paths exist upstream (Netflix/vmaf has no Go module), so an upstream sync cannot conflict with the new IterSamples / IterCloud / IterHull / AllSeq / ListModelsSeq surfaces. The deprecated Registry.All / Registry.ListModels shims are likewise fork-local. If a future Netflix upstream adds Go bindings, the conflict is resolution-only at the package-tree level (different directory layout, no symbol overlap).
Skills library expansion — /add-mcp-tool, /add-k8s-resource, /audit-modernization, bisect-common (ADR-0939, 2026-05-31)¶
no rebase impact: all new files land under .claude/skills/, which is fork-local infrastructure (the upstream Netflix/vmaf repo does not ship a .claude/ directory). The accompanying ADR, index fragment, changelog fragment, and research digest are likewise fork-local. The two existing bisect skills (bisect-regression, bisect-model-quality) gain scaffold.sh driver scripts that source .claude/skills/lib/bisect-common.sh — still all fork-local. No upstream files touched.
If upstream Netflix ever adopts .claude/ skills (unlikely — different agent tooling), revisit whether the three new scaffolds should be promoted or stay fork-only. The bisect-common library has no upstream analogue either, so the merge surface is zero.
ai/ dataclass → pydantic v2 migration (ADR-0934, 2026-05-31)¶
no rebase impact: REASON — touched files are entirely fork-local. Upstream Netflix/vmaf does not ship ai/src/vmaf_train/ at all (the package is fork-added — Tiny-AI surface, ADR-0042). TrainConfig (train.py), ModelMetadata (registry.py), and ManifestEntry (data/datasets.py) become pydantic.BaseModels; pydantic>=2.13.4 added to ai/pyproject.toml (already in tree via mcp-server/vmaf-mcp). Sidecar JSON layout byte-identical (ModelMetadata.to_json() uses model_dump(mode="json") + json.dumps(indent=2, sort_keys=True)). On upstream sync the diff cannot conflict — Netflix has no equivalent file to merge into.
Vendored libsvm + IQA test-coverage uplift (2026-05-31, ADR-0952)¶
core/test/test_svm_api.c and core/test/test_iqa_helpers.c are pure fork-local test additions. They link against the vendored libsvm_static_lib (for svm) and libvmaf_feature_static_lib + libvmaf_cpu_static_lib (for iqa) via their extract_all_objects recipes — the same pattern used by test_iqa_convolve.c, test_feature_extractor.c, and PR #381's test_svm_parser.c. No vendored source is touched.
Rebase impact: when Netflix upstream re-pins libsvm (3.24 → 3.36 or later) or when the IQA helpers gain new public functions:
test_svm_api.cassertions on inspector outputs and on the C-SVC / EPSILON-SVR predict round-trip are functional invariants of the libsvm public API; a major-version bump that changes them is a semantic break and should land its own ADR.test_iqa_helpers.c_round()and_cmp_float()assertions document the current asymmetric rounding rule ("trunc toward zero, add sign when |frac| >= 0.5"). If upstream tdistler.com (or Netflix's 2016 update) ever rewrites those helpers to IEEE-754 round-half-to-even, the tests will fail — that is by design; the failure surfaces the unintended numerical change at the rebase diff, not at the integration SSIM result.
The meson wiring in core/test/meson.build inserts two new executables above test_feature_extractor and registers them in the fast suite. Both fragments are isolated; the only adjacency to upstream code is the alphabetical position in the test list.
PR companion to ADR-0889 (PR #381, libsvm parser audit) — the two PRs can land in either order without conflict.
external-bench test coverage backfill (ADR-0332 follow-up, 2026-05-31)¶
no rebase impact: REASON — changes are confined to fork-only files (tools/external-bench/tests/test_compare.py, changelog.d/added/*, docs/research/*). The tools/external-bench/ tree is fork-only per ADR-0332 (no upstream counterpart); coverage backfill (14 new tests for BVI-DVC discovery edge cases, Netflix discovery edge cases, validator rejection paths, run_wrapper missing-output guard, and main() --limit + per-item skip flow) cannot conflict on upstream sync.
HIP motion3 parity test ENOSYS-skip (ADR-0949, 2026-05-31)¶
no rebase impact: REASON — the only file touched in the libvmaf source tree is core/test/test_hip_motion3_parity.c, which is wholly fork-added (no Netflix upstream counterpart — Netflix/vmaf does not ship a HIP backend). The skip-on--ENOSYS change is self-contained inside the test's HIP-path helper; CPU baseline, tolerance, fixture geometry, and end-of-stream handling are unchanged. Upstream sync cannot conflict because no upstream file touches this test path or the motion_hip extractor's scaffold-vs-runtime split.
GitHub Actions custom-action + reusable-workflow audit (ADR-0951, 2026-05-31)¶
no rebase impact: REASON — audit-only PR. No code under core/, python/, ai/, mcp-server/, or tools/ is touched. The only edited files are: docs/adr/0951-github-actions-custom-audit.md (new ADR), docs/adr/README.md + docs/adr/_index_fragments/_order.txt (index rows), docs/research/0951-github-actions-custom-audit.md (digest), changelog.d/changed/github-actions-custom-audit.md (fragment), and this rebase-notes row. The fork-wide SHA-pin invariant in .github/AGENTS.md lines 100–144 is unchanged. On upstream sync the audit conclusions remain valid until Netflix introduces its own .github/actions/ tree or workflow_call: workflow; re-run the three reproducer commands in the research digest to confirm.
fix(mcp-server): NamedTemporaryFile (ADR-0975) — no rebase impact¶
Replaces a local variable assignment in _run_vmaf_score. No C surface, no public API, no upstream-mirrored file touched. On rebase against upstream Netflix/vmaf, this change applies cleanly to the MCP server layer which is entirely fork-local.
ADR-0945 — HIP kernel parity coverage round 3 — 2026-05-31¶
no rebase impact: REASON — the 4 new test files (core/test/test_hip_cambi_parity.c, core/test/test_hip_float_adm_parity.c, core/test/test_hip_float_motion_parity.c, core/test/test_hip_float_psnr_parity.c) live entirely under the fork-only HIP backend tree. Upstream Netflix/vmaf has no HIP backend, no parity tests, and no test_hip_* files; the if get_option('enable_hip') == true block in core/test/meson.build is fork-local (added by the HIP scaffold landing in ADR-0212). Wiring lives strictly inside that block. The only non-test files touched are docs/adr/README.md (index row), docs/adr/_index_fragments/_order.txt, the changelog.d/added/hip-kernel-coverage-round3.md fragment, and the companion docs/research/hip-kernel-coverage-round3-2026-05-31.md audit — all fork-only.
ADR-0918 — LLVM IR diff harness — 2026-05-31¶
no rebase impact: harness is fork-local tooling (scripts/perf/check-ir-diff.sh, scripts/perf/ir-diff-config.yaml, testdata/ir-snapshots/, make ir-diff / make ir-diff-update targets). It snapshots LLVM IR for fork-added SIMD sources only; Netflix upstream never touches these paths. The only upstream coupling is the SIMD source files themselves (core/src/feature/x86/*.c) — if a future upstream sync changes the scalar reference for psnr_hvs / ms_ssim_decimate / ssimulacra2 and the AVX2 twin must change in lockstep, the snapshot regen step (make ir-diff-update) is a normal part of the port — same discipline as the score JSON snapshots under /regen-snapshots. The new core/src/feature/x86/AGENTS.md invariant note flags this for the next sync agent.
vmafx-operator functional test coverage uplift (2026-05-31)¶
no rebase impact: REASON — all four new test files live under cmd/vmafx-operator/internal/controller/ which is fork-added per ADR-0714 (vmafx-operator kubebuilder skeleton). Upstream Netflix/vmaf ships no Go sources and no Kubernetes operator surface; there is nothing to merge against.
Fork-local files: cmd/vmafx-operator/internal/controller/vmafxnode_controller_test.go (new), cmd/vmafx-operator/internal/controller/vmafxjob_controller_branch_test.go (new), cmd/vmafx-operator/internal/controller/vmafxmodeltraining_controller_branch_test.go (new), cmd/vmafx-operator/internal/controller/setup_with_manager_test.go (new), changelog.d/added/operator-functional-coverage.md (new).
ADR-0913 — CHANGELOG.md renderer splice contract + 44 k-line drift sweep — 2026-05-31¶
no rebase impact (upstream): REASON — fork-local infrastructure only. The renderer (scripts/release/concat-changelog-fragments.sh), the rendered file (CHANGELOG.md), and the fragment tree (changelog.d/) are all fork-added; Netflix/vmaf upstream uses a hand-edited CHANGELOG.md with no fragment system.
In-flight fork-branch impact (medium): every in-flight fork branch that added a fragment under the old ## Section / ### Section shape will conflict on its fragment file at rebase time. Resolution is mechanical — keep the bullet content, drop the redundant first-line section header (the renderer emits ### Section itself). Branches that added perf entries under changelog.d/perf/ or changelog.d/performance/ need to rename to changelog.d/changed/perf-<topic>.md (the same convention PR #384 / ADR-0892 introduces). On rebase the renderer's new stderr WARNING surfaces the wrong directory immediately; bash scripts/release/concat-changelog-fragments.sh --check then verifies the fix.
__init__.py export-completeness audit (ADR-0911, 2026-05-31)¶
no rebase impact: REASON — all eight modified __init__.py files are fork-added (ai/__init__.py, ai/data/__init__.py, ai/train/__init__.py, ai/src/vmaf_train/__init__.py, ai/src/vmaf_train/data/__init__.py, dev-llm/src/vmaf_dev_llm/__init__.py, mcp-server/vmaf-mcp/src/vmaf_mcp/__init__.py, scripts/lib/__init__.py). Upstream-mirror packages (compat/python-vmaf/**, python/test/__init__.py) were deliberately left byte-identical per the upstream-mirror rebase-hygiene rule. No upstream Netflix/vmaf file is touched.
ADR-0907 — Wall-clock perf regression gate (2026-05-30)¶
No rebase impact on upstream C/Python code.
New fork-local files only:
scripts/perf/check-regression.py— gate script (stdlib-only)scripts/perf/test_check_regression.py— smoke testsdocs/adr/0907-perf-regression-gate-wall-clock.md(new)docs/adr/_index_fragments/0907-perf-regression-gate-wall-clock.md(new)changelog.d/added/perf-regression-gate.md(new).github/workflows/tests-and-quality-gates.yml(newperf-regressionjob; the disabledcross-backendjob's brokenbench_all.sh --backend=cpu --snapshot-only --tolerance-ulp=2invocation is replaced with a no-op placeholder echo sincebench_all.shdoes not parse those flags)
No upstream Netflix collision risk — the gate consumes only the fork-added testdata/perf_multi_resolution.json baseline (ADR-0752, fork-local).
Slow-test audit (ADR-0908, 2026-05-30)¶
no rebase impact: REASON — all touched files are fork-local. A new ADR (docs/adr/0908-slow-test-audit-2026-05-30.md), a new research digest (docs/research/slow-test-audit-2026-05-30.md), fork-added pytest configuration in three pyproject.toml files registering the slow marker (tools/vmaf-tune/pyproject.toml, ai/pyproject.toml, mcp-server/vmaf-mcp/pyproject.toml), and fork-added test files (tools/vmaf-tune/tests/test_bbb_e2e_v5_bug_cluster.py, tools/vmaf-tune/tests/test_bbb_e2e_v14_bug_cluster.py). None are mirrored from upstream Netflix/vmaf.
ADR status-field drift sweep (2026-05-30)¶
no rebase impact: changes are confined to fork-local ADR markdown files under docs/adr/. Status-field flips on ADR-0573 (→ Superseded by ADR-0738) and Status normalisation on ADR-0105 / ADR-0106 / ADR-0107 (Supersedes-in-Status → Accepted + explicit Supersedes line). Netflix upstream has no docs/adr/ tree; nothing to reconcile on sync. Audit methodology and the full decision matrix live in docs/research/adr-status-drift-audit-2026-05-30.md.
ADR-0903 — Codecov upload wiring (2026-05-30)¶
no rebase impact: REASON — all changes are confined to fork-only files: .github/workflows/tests-and-quality-gates.yml is a fork-added CI workflow (upstream Netflix/vmaf has no equivalent gcovr-based Coverage Gate), docs/adr/0903-wire-codecov-upload.md is fork-only documentation, and changelog.d/added/wire-codecov-upload.md is a fork-only changelog fragment per ADR-0221. The added codecov/codecov-action steps depend only on the Cobertura XML the existing gcovr step already produces; upstream sync cannot break this wiring because the gcovr job itself is fork-only.
ADR-0904 — cargo-machete build-dep ignores (2026-05-30)¶
no rebase impact: REASON — Netflix/vmaf upstream has no Rust workspace. Both touched Cargo.toml files (bindings/rust/vmafx-sys/Cargo.toml, core/src/feature/rust/tad/Cargo.toml) live entirely in fork-local trees (ADR-0702, ADR-0707). The [package.metadata.cargo-machete] blocks add no-op metadata (cargo ignores keys it doesn't know) and cannot conflict with anything upstream might add later.
Signing and attestation audit (ADR-0902, 2026-05-30)¶
no rebase impact: REASON — changes are confined to fork-local CI infrastructure (.github/workflows/docker-publish-production.yml, docs/development/release.md, docs/adr/0902-*.md, docs/research/signing-and-attestation-audit-2026-05-30.md, changelog.d/security/signing-and-attestation-audit.md). The supply-chain workflow (supply-chain.yml) is itself fork-additive (Netflix upstream does not ship a Sigstore + SLSA + SBOM release channel); upstream syncs never touch any of these files.
Doxygen public-API clean (ADR-0953, 2026-05-31)¶
Low rebase impact: every edit lands as additional doxygen comments or per-member /**< desc */ annotations inside the public headers under core/include/libvmaf/. Two of the headers touched are Netflix-upstream-mirrored — picture.h and model.h — but the edits are pure documentation; no struct layout, function signature, or symbol name moves. An upstream sync that re-touches either file should accept its hunks unchanged and let the fork's doc comments remain in place. The four fork-added headers (libvmaf_mcp.h, dnn.h, libvmaf_metal.h, libvmaf_hip.h) are fork-local and have no upstream counterpart. The new core/doc/Doxyfile.public-api, .github/workflows/doxygen-public-api.yml, ADR-0953, research digest, changelog fragment, and AGENTS.md invariant note are fork-local — zero rebase exposure.: warning-clean doxygen build for libvmaf public C API (recovery of #457)): warning-clean doxygen build for libvmaf public C API (recovery of #457))
governance-audit (2026-05-30, ADR-0901)¶
No rebase impact — all changes are fork-local governance files that upstream Netflix/vmaf does not ship:
GOVERNANCE.md(new),MAINTAINERS.md(new) — top-level fork-only..github/CODEOWNERS— append-only additions below the existing rows (the rename of the existing/libvmaf/...rows to/core/...is owned by in-flight PR #321, not this PR).CONTRIBUTING.md— fork-specific block extended with branch-naming, ADR-0108 deliverables, ADR-allocator pointer, governance pointer. The inherited Netflix upstream contribution-guide block at the bottom is unchanged.docs/adr/0901-governance-audit.md,docs/adr/_index_fragments/0901-governance-audit.md,docs/adr/_index_fragments/_order.txt(one-line append),docs/research/governance-audit-2026-05-30.md,changelog.d/added/governance-audit.md— all fork-only paths.
On upstream sync, no conflict is expected. If CODEOWNERS shows a textual conflict because PR #321 landed in-between, the resolution is trivial: keep PR #321's renamed /core/... rows AND keep this PR's new append-only rows. Both edits are non-overlapping at the line level.
ADR-0893 — Pre-commit config audit — 2026-05-30¶
no rebase impact: REASON — .pre-commit-config.yaml is a fork-local config file. Upstream Netflix/vmaf does not ship pre-commit configuration; all revisions and hooks listed are fork-owned. Touches one fork-owned Python file via isort 6.0.1 auto-fix (tools/vmaf-tune/tests/test_codec_adapter_av1_videotoolbox.py), which is itself outside the upstream tree.
libsvm vendored audit — extend SAN-MODEL-MALLOC-OOB to row-ordering (ADR-0889, 2026-05-30)¶
Touches the vendored libsvm parser core/src/svm.cpp, which is wrapped in a file-level NOLINTBEGIN / NOLINTEND cordon. On Netflix-vmaf upstream sync the file is part of the fork-mirrored set: Netflix upstream has not refreshed its vendored libsvm copy since 2020-11 either, so a Netflix-only sync has near-zero conflict risk on this file.
On an upstream libsvm (Chih-Chung Chang / Chih-Jen Lin) sync — deliberately deferred per ADR-0889 — the fork carries three patch families that must be re-applied:
- Thread-locale isolation (ADR-0137) —
buffer.imbue(std::locale::classic())in bothSVMModelParserFileSourceandSVMModelParserBufferSourceconstructors. - JSON in-memory entry point —
svm_parse_model_from_bufferplus theSVMModelParserBufferSourcetemplate instantiation. Consumed byread_json_model.c; removing it breaks JSON-embedded SVM model loading. - SAN-MODEL-MALLOC-OOB hardening —
VMAF_SVM_MAX_AXIS_COUNT(1 << 24) bound,nr_class/total_svaxis-size asserts inparse_header()andparse_support_vectors(),sv_buffer.empty()post-parse guard, plus the row-ordering preconditions (exceptAssert(model->nr_class > 0, ...)) onrho,label,probA,probB,nr_svadded by ADR-0889.
Regression coverage at core/test/test_svm_parser.c (suite fast). On sync, re-run that test plus test_predict and test_model before merging. See core/src/AGENTS.md §10 for the full invariant list.
CI concurrency + cost audit (ADR-0890, 2026-05-30)¶
no rebase impact: REASON — CI-only changes to .github/workflows/ files that are wholly fork-local. Netflix upstream's CI is one .github/workflows/ file with a different name and structure; the five files modified here (ffmpeg-integration.yml, sanitizers.yml, security-scans.yml, lint-and-format.yml, plus the ADR / changelog / state.md surface) have no upstream counterpart. No source / header / patch surface touched; the ffmpeg-patches/ series is unaffected.
ADR-0883 — HIP kernel parity coverage round 2 — 2026-05-30¶
no rebase impact: REASON — the 5 new test files (core/test/test_hip_ciede_parity.c, test_hip_psnr_hvs_parity.c, test_hip_motion_parity.c, test_hip_ssim_parity.c, test_hip_ms_ssim_parity.c) live entirely under the fork-only HIP backend tree. Upstream Netflix/vmaf has no HIP backend, no parity tests, and no test_hip_* files; the if get_option('enable_hip') == true block in core/test/meson.build is fork-local (enable_hip was added by the HIP scaffold landing in ADR-0212). Wiring lives strictly inside that block. The only non-test file touched is docs/adr/README.md (index row) and docs/adr/_index_fragments/_order.txt — both fork-only.
ADR-0876 — printf-format portability sweep (CERT FIO47-C) — 2026-05-30¶
Low rebase impact, scoped to fork-added log / debug call sites. The four touched source files (core/src/libvmaf.c, core/src/sycl/common.cpp, core/src/sycl/dmabuf_import.cpp, core/test/test_motion_v2_simd.c) either are fork-added (the SYCL TUs + the AVX2 test) or contain a fork-added block inside an upstream-mirror file (the tiny-model loader in libvmaf.c, which is post-ADR-0700 fork-edited per git blame). The format-string changes are mechanical: (unsigned long)x + %lu → x + %" PRIu64 " for uint64_t; (long long)x + %lld → x + %" PRId64 " for int64_t; (unsigned long long)x + %llx → x + %" PRIx64 " for uint64_t hex prints. Three call sites in upstream-mirror code (core/src/feature/x86/adm_avx512.c print_128_64 debug macro) and POSIX-off_t / Windows-DWORD sites were intentionally not changed — see docs/research/0876-printf-format-portability-audit.md §2 Class C for the rationale. Future upstream syncs that touch the same lines will conflict trivially; resolve in favour of the PRI-macro form for fixed-width types.
ADR-0877 — error-code consistency audit (MS-SSIM decimate) — 2026-05-30¶
no rebase impact: the four touched TUs (core/src/feature/ms_ssim_decimate.{c,h}, core/src/feature/x86/ms_ssim_decimate_{avx2,avx512}.c, core/src/feature/arm64/ms_ssim_decimate_neon.c) are fork-added 2026-04-20; they have no upstream Netflix/vmaf counterpart. The change converts the malloc-failure branch from bare return -1 to return -ENOMEM and tightens the header docstring to match — no logic change on the hot path. Bit-exactness across scalar / AVX2 / AVX-512 / NEON is preserved (only the cold malloc-failure branch is touched).
ADR-0875 — GitHub Actions hardening audit — 2026-05-30¶
no rebase impact: REASON — all changes are confined to fork-local CI workflows under .github/workflows/ (go-ci.yml, rust-ci.yml, sanitizers.yml, supply-chain.yml). Upstream Netflix/vmaf has a completely different CI pipeline; none of these files exist upstream. Adds top-level permissions: contents: read to the two Go/Rust workflows and persist-credentials: false to five actions/checkout steps. No source code touched.
Fork-local files: .github/workflows/go-ci.yml, .github/workflows/rust-ci.yml, .github/workflows/sanitizers.yml, .github/workflows/supply-chain.yml, docs/adr/0875-github-actions-audit-2026-05-30.md, docs/research/github-actions-audit-2026-05-30.md, changelog.d/security/github-actions-audit-2026-05-30.md.
ADR-0873 — ARM64 NEON bit-exactness audit — 2026-05-30¶
Rebase impact: low, limited to build system and one test file.
core/src/meson.build lines 581–643: the arm64_v8 static lib is split into arm64_v8 (integer-only TUs, unchanged compile flags) and arm64_v8_fp (float-arithmetic TUs, new -ffp-contract=off flag). If upstream Netflix/vmaf adds new NEON TUs to this region, they must be classified as integer or float and placed in the correct lib.
core/src/feature/arm64/float_adm_neon.c: float_adm_sum_cube_neon and float_adm_csf_den_scale_neon now accumulate into float64x2_t instead of float32x4_t. This is a numeric change — if upstream modifies these functions, the double-accumulation pattern must be preserved.
core/src/feature/adm.c: comment-only change (ADR-0873 follow-up note).
core/test/test_motion_v2_simd.c: fill_adversarial_neg and fill_adversarial_mixed moved outside #if ARCH_X86; NEON test arm added for motion_score_pipeline_16_neon. On upstream sync, ensure the x86 test body still compiles.
Fork-local files: core/src/meson.build (lib split), core/src/feature/arm64/float_adm_neon.c (reduction stability), core/src/feature/arm64/AGENTS.md (invariant note), core/src/feature/adm.c (comment), core/test/test_motion_v2_simd.c (NEON test arm), docs/adr/0873-arm64-neon-bit-exactness-audit.md, changelog.d/fixed/arm64-neon-bit-exactness-audit.md.
Logging consistency audit — 2026-05-30¶
No rebase impact on upstream. All routed sites are fork-local: core/src/libvmaf.c (vmaf_write_output — fork-added entry point added by the --precision/output overhaul), core/src/sycl/dispatch_strategy.cpp (fork-only file, ADR-0181), core/src/sycl/common.cpp (fork-only file). Vendored core/src/svm.cpp and upstream-mirror feature extractors (vif.c, adm.c, ms_ssim.c, motion.c, ssim.c) are explicitly deferred precisely because they carry upstream-sync invariants — leaving them untouched preserves the rebase story.
Fork-local files: core/src/libvmaf.c, core/src/sycl/dispatch_strategy.cpp, core/src/sycl/common.cpp, docs/research/logging-consistency-audit-2026-05-30.md, changelog.d/changed/logging-consistency-audit.md.
ADR-0870 — Helm values.schema.json + dev-MCP path drift — 2026-05-30¶
no rebase impact: all touched files are fork-additions (deploy/helm/, dev/Containerfile, dev/docker-compose.yml, .dockerignore, docs/adr/0870-*.md, docs/adr/README.md, docs/adr/_index_fragments/_order.txt, docs/development/k8s-deployment.md, changelog.d/added/0870-*.md, changelog.d/fixed/0870-*.md, docs/state.md). None of these have upstream Netflix/vmaf counterparts. The Containerfile path fixes (libvmaf/ → core/) are the downstream of ADR-0700's repo rename; future rebases against a hypothetical upstream that re-introduced a libvmaf/ directory at the repo root would need their own audit, but no such state exists or is planned.
ADR-0868 — GPU backend kernel coverage gap-fill — 2026-05-30¶
No rebase impact: all changes are net-new fork-local test files under core/test/test_{cuda,hip,sycl,metal}_*_parity*.c plus their meson wiring. No upstream Netflix/vmaf source is touched. The tests target fork-added GPU extractor names (psnr_cuda, ciede_cuda, psnr_hip, vif_hip, psnr_sycl, vif_sycl, the 8 *_metal extractors) which do not exist upstream. Mirrors the existing test_{cuda,hip,sycl}_motion3_parity.c pattern (already fork-local).
Fork-local files: core/test/test_cuda_psnr_parity.c, core/test/test_cuda_ciede_parity.c, core/test/test_hip_psnr_parity.c, core/test/test_hip_vif_parity.c, core/test/test_sycl_psnr_parity.c, core/test/test_sycl_vif_parity.c, core/test/test_metal_kernel_registration.c, core/test/meson.build (additive blocks only, no upstream-touching hunks), docs/adr/0868-gpu-backend-kernel-coverage.md, docs/research/gpu-backend-kernel-coverage-audit-2026-05-30.md, changelog.d/added/0868-gpu-backend-kernel-coverage.md.
test/feature-extractor-coverage-push — 2026-05-30¶
no rebase impact: REASON — test-only changes confined to fork-local files under core/test/. New test file core/test/test_mkdirp.c is wholly fork-added (no upstream equivalent); the four touched files (core/test/test_luminance_tools.c, core/test/test_feature.c, core/test/test_feature_extractor.c, core/test/meson.build) gain only new static char *test_… functions and registrations — no existing logic edited. No production source under core/src/ is touched. The test_mkdirp binary compiles core/src/feature/mkdirp.c directly into its TU; this is the same pattern other test binaries already use (test_ref compiles ../src/ref.c, test_thread_pool compiles ../src/thread_pool.c, etc.), so the linkage introduces no new precedent.
Fork-local files: core/test/test_mkdirp.c (new), core/test/test_luminance_tools.c, core/test/test_feature.c, core/test/test_feature_extractor.c, core/test/meson.build, changelog.d/added/feature-extractor-coverage-push.md.
fix/simd-bug-audit-20260531 — 2026-05-31¶
no rebase impact: fork-local SIMD entry points only. The two patched files (core/src/feature/x86/float_adm_avx2.c, core/src/feature/arm64/float_adm_neon.c) are fork-added SIMD ports of upstream adm_dwt2_s; they are not yet wired through compute_adm (ADR-0873 follow-up). Upstream Netflix/vmaf has neither file. The change harmonises the NULL-allocation guard with the already-shipped AVX-512 sibling (float_adm_avx512.c) which has been the de-facto reference since master tip; no upstream merge can collide. The third file (core/src/feature/arm64/ssimulacra2_host_neon.c) is wholly fork-added (SSIMULACRA 2 is a fork extractor) and the edit is comment- only.
Fork-local files: core/src/feature/x86/float_adm_avx2.c, core/src/feature/arm64/float_adm_neon.c, core/src/feature/arm64/ssimulacra2_host_neon.c, changelog.d/fixed/simd-float-adm-dwt2-unchecked-aligned-malloc.md.
ai/ tempfile + path-safety bandit sweep — 2026-05-30¶
no rebase impact: REASON — every touched file lives under ai/scripts/ or ai/tests/, all of which are wholly fork-local (Netflix upstream ships no tiny-AI training, dataset acquisition, or ONNX export pipeline). No upstream Netflix/vmaf file is touched. Fork-local files: ai/scripts/bvi_dvc_to_full_features.py, ai/scripts/konvid_to_full_features.py, ai/scripts/export_tiny_models.py, ai/scripts/export_u2netp_mirror.py, ai/scripts/export_vmaf_tiny_v{2,3,4}.py, ai/scripts/fetch_konvid_1k.py, ai/scripts/fetch_youtube_ugc_subset.py, ai/tests/test_corpus_base.py, ai/tests/test_feature_extractor_defaults.py, ai/tests/test_merge_corpora.py, ai/tests/test_train_predictor_v2_realcorpus.py, changelog.d/security/ai-tempfile-and-path-safety.md.
Go controller / server / MCP test coverage expansion (2026-05-30)¶
No rebase impact: all touched files are fork-added Go tests under cmd/ — none have an upstream Netflix/vmaf counterpart (upstream ships no Go sources). The PR adds:
cmd/vmafx-controller/main_extra_test.go(new)cmd/vmafx-controller/nodes/registry_edge_test.go(new)cmd/vmafx-server/main_extra_test.go(new)cmd/vmafx-mcp/impl_test.go(new)changelog.d/added/go-controller-mcp-coverage.md(new)docs/state.md(one_Updated:annotation line; no row change)docs/rebase-notes.md(this entry)
No upstream-mirror file is touched.
ADR-0848 — Per-surface doc compliance audit (2026-05-29)¶
Rebase impact: none. This PR adds only docs/research/, docs/adr/, changelog.d/, and docs/state.md changes. No code, no meson, no public headers.
Future rebases: If PRs that fix the three gaps (Issue A / B / C from Research-0848) are in flight, ensure: - Issue A (Vulkan removal docs): no conflict expected — docs/backends/vulkan/, docs/metrics/features.md, docs/development/build-flags.md are rarely touched. - Issue B (deprecations.md): docs/development/deprecations.md is append-only.
Changelog fragment consolidation (2026-05-29)¶
no rebase impact: changelog-only — scripts/release/concat-changelog-fragments.sh awk fix + changelog.d/ fragment moves do not touch any upstream Netflix/vmaf source file. No C, Python, or test changes.
float_adm AVX2/AVX-512 F2+F3 precision fix (ADR-0844, 2026-05-29)¶
Rebase invariant: if an upstream Netflix/vmaf commit changes float_adm_csf_den_scale_s, float_adm_sum_cube_s, or any other reduction function in core/src/feature/float_adm.c, the corresponding AVX2 and AVX-512 variants in core/src/feature/x86/float_adm_avx2.c and float_adm_avx512.c must be updated to preserve the double-precision widening contract (ADR-0844 / ADR-0139). The hadd_pd4 and hsum_ps_to_double helpers are static inline and duplicated across TUs intentionally — do not merge them into a shared header. The -ffp-contract=off per-TU static library carve-out in core/src/meson.build (the x86_float_adm_avx2_lib and x86_float_adm_avx512_lib targets) must be preserved on any rebase that touches the meson.build AVX2/AVX-512 build block; they mirror the ssimulacra2 carve-out already in tree.
AVX-512 motion parity tests (ADR-0854, 2026-05-29)¶
no rebase impact: REASON — changes are confined to new test files (core/test/test_motion_avx512_parity.c, changelog.d/added/motion-avx512-parity-tests.md, docs/adr/0854-motion-avx512-parity-tests.md) and additive changes to core/test/simd_bitexact_test.h (new helper function) and core/test/meson.build (new test registration). No upstream Netflix/vmaf production source is modified; no existing test is changed; no golden assertions are touched.
ADR-0852 — HIP speed extractor wiring (2026-05-29)¶
no rebase impact: the three changed files (core/src/meson.build, core/src/hip/meson.build, core/src/feature/feature_extractor.c) are fork-owned; no upstream Netflix/vmaf C source is touched. The only upstream- adjacent file is feature_extractor.c whose #if HAVE_HIP block is a fork-added section; conflicts are only possible with other HIP-wiring PRs.
Dependency audit 2026-05-30 — golang.org/x/net + x/sys bump¶
No rebase impact: the only changed files are go.mod / go.sum, plus a changelog fragment and a research digest. The Go workspace is a fork-only addition (Netflix/vmaf upstream does not ship Go modules); there is no upstream baseline to rebase against. Versions: golang.org/x/net v0.53.0 -> v0.55.0, golang.org/x/sys v0.43.0 -> v0.45.0, golang.org/x/term v0.42.0 -> v0.43.0, golang.org/x/text v0.36.0 -> v0.37.0 (minimum-version selection).
Fork-local files: go.mod, go.sum, changelog.d/security/dependency-audit-2026-05-30.md, docs/research/dependency-audit-2026-05-30.md.
CodeQL Go coverage + config conflict resolution (ADR-0811, 2026-05-29)¶
no rebase impact: CI-config-only change; no public API surface affected. All changes are confined to .github/codeql-config.yml (Go paths addition + gen/go exclusion), .github/workflows/security-scans.yml (new codeql-go job), docs/adr/0811-security-codeql-go-pvr.md, and the changelog fragment. No upstream Netflix/vmaf files are touched; no C/Python/Go production code is modified. On upstream sync, the CodeQL workflow additions apply cleanly regardless of upstream changes.
Fork-local files: .github/codeql-config.yml, .github/workflows/security-scans.yml, docs/adr/0811-security-codeql-go-pvr.md, changelog.d/security/0811-codeql-go-config-fix.md.
release-please draft mode¶
no rebase impact: release-tooling-only change (release-please-config.json "draft": true). No C sources, headers, or test logic modified.
Coverage-overrides audit — tighten tiny_extractor_template.h (ADR-0881, 2026-05-30)¶
no rebase impact: REASON — changes are confined to fork-only files: scripts/ci/coverage-check.sh (fork-only CI gate), the new docs/adr/0881-*.md ADR, the new docs/research/0881-*.md digest, the ADR index fragment under docs/adr/_index_fragments/, and the changelog.d/changed/ fragment. The threshold ratchet only tightens an existing override (10 → 75) — does not introduce a new path Netflix upstream might also override. Future audits per the codified rule (see ADR-0881 §Decision) are also fork-only since coverage-check.sh itself is fork-only (Netflix upstream has no equivalent gate).
vmafx-operator envtest etcd setup (2026-05-30)¶
no rebase impact: REASON — all changes are in fork-added paths only. Files touched: Makefile (new setup-envtest + setup-envtest-env targets in the Go workspace section, ADR-0702 scope), .github/workflows/go-ci.yml (new pre-test step installing sigs.k8s.io/controller-runtime/tools/setup-envtest@latest + exporting KUBEBUILDER_ASSETS), cmd/vmafx-operator/internal/controller/suite_test.go (top-of-TestControllers t.Skip() guard + nil-testEnv bailout in AfterSuite), cmd/vmafx-operator/AGENTS.md (new invariant #6 documenting the skip-safe envtest pattern), and the changelog.d/fixed/ fragment. cmd/vmafx-operator/ is fork-added per ADR-0714 — upstream Netflix/vmaf ships no Go sources, so no upstream merge can reach these files.
log.c → log.cpp C++23 pilot (ADR-0708 Wave 1, 2026-05-30)¶
Upstream Netflix libvmaf/src/log.c is fork-renamed to core/src/log.cpp. Future port-upstream-commit runs that touch libvmaf/src/log.c must apply changes to core/src/log.cpp instead — the fork-rename mapping is recorded here.
Public C ABI is preserved: core/src/log.h retains the same two function prototypes (vmaf_log, vmaf_set_log_level) and now carries extern "C" guards so it is includable from both C and C++ TUs. The C-mangled exported symbols are unchanged (nm libvmaf.so shows vmaf_log and vmaf_set_log_level with the same C-mangling as the prior log.c build).
Meson wiring: log.cpp compiles in an isolated log_cpp23_lib static_library with override_options: ['cpp_std=' + libvmaf_cpu_cpp_std], mirroring the metadata_handler_cpp20_lib pattern (ADR-0708 metadata_handler pilot). Test executables that previously direct-compiled ../src/log.c (test_lpips, test_dists, test_feature_extractor, test_speed, ...) now pick up log symbols via the shared log_cpp23_test_objects aggregate in core/test/meson.build.
Fork-local files: core/src/log.cpp (was: core/src/log.c, removed), core/src/log.h (added extern "C" guards), core/src/meson.build (replaced log.c source entry with log_cpp23_lib), core/test/meson.build (added log_cpp23_test_objects, removed inline '../src/log.c' source entries from ~20 test execs, wired test_log into the fast suite), docs/adr/0708-vmafx-cpp23-internals-pilot.md (consequences cross-link), changelog.d/changed/log-c-to-cpp23.md.extern "C" guards added: log.h, model.h, read_json_model.h, opt.h. Any upstream commit that adds new declarations to these headers must include the guard-wrapped declaration for correctness. Flag in the port if upstream adds a declaration outside the guard block.
port/upstream-netflix-may-jun-2026 — 2026-06-01¶
Five Netflix upstream commits ported. Each reduces the diff against upstream and therefore reduces future rebase friction.
-
e4b93c6ed (
fetch_picturedirect-read):core/tools/vmaf.cno longer has a#ifdef USE_DIRECT_READbranch. Future upstreams that touchvmaf.cwill now merge cleanly without the compile-guard conflict. -
a4a1492d3 (integer_motion rename):
core/src/feature/integer_motion.candcore/src/feature/x86/motion_avx2.{c,h}/motion_avx512.{c,h}are now at upstream parity.integer_motion_v2.candmotion_v2_avx2/512are fork-local (GPU build paths); any future upstream touch of those names should check whether the GPU backends have been updated to the renamed API. -
c2155d6cd (2160p CSF):
core/src/feature/barten_csf_tools.his now at upstream parity.core/test/test_barten_csf.chas new upstream tests. -
9a078011c (ADM SIMD fix):
core/src/feature/integer_adm.candcore/src/feature/x86/adm_avx2.c+adm_avx512.cat upstream parity. -
30f472b14 (Speed_chroma AVX):
core/src/feature/speed.c,core/src/feature/x86/speed_avx2.{c,h},core/src/feature/x86/speed_avx512.{c,h}are new upstream-mirror files. Future upstream touches to speed.c may need to propagatecompute_cov_kernelinto the GPU speed_chroma extractors.
Fork-local files touched: core/tools/vmaf.c (commit #1 — call-site updates for signature change), core/src/feature/feature_extractor.c (commit #2 — remove CPU v2), core/src/meson.build (commits #2, #5 — add speed_avx2/512, remove motion_v2 CPU build).
ADR-0700 Dockerfile path residuals — 2026-05-30¶
no rebase impact: REASON — touches only fork-added Dockerfiles (docker/Dockerfile.production, docker/Dockerfile.production-gpu, docker/dev/{alpine-3.20,arch,fedora-40}.Dockerfile) and the fork-added dev/Containerfile. None of these files have an upstream Netflix/vmaf counterpart. The change is a literal libvmaf/ → core/ substitution at source-tree positions (meson setup … core, COPY core/, cd core); install-path / package / filter-name occurrences (/usr/local/include/libvmaf/, libvmaf.so, libvmaf-dev, --enable-libvmaf*) are deliberately preserved because they describe the shipped library / package / ffmpeg-filter surface, not the source layout.
ADR-0709 residual ANSNR references in docs + ai/data — 2026-05-30¶
no rebase impact: REASON — all changes are fork-local. Touched files are ai/data/feature_extractor.py (fork-added Python helper, no upstream counterpart), docs/metrics/ansnr.md, docs/backends/index.md, docs/backends/cuda/overview.md, docs/backends/hip/overview.md. The HIP and CUDA overviews and the metric page are fork-only docs; the backends index page is also fork-only. No upstream Netflix/vmaf source is touched. The cleanup closes residual references left over after PR #38 (ADR-0709) removed the float_ansnr extractor from every backend.
fix/post-rename-post-vulkan-sweep — 2026-06-01¶
no rebase impact: post-rename cleanup only. The Containerfile change adds a pkg-install line that cannot conflict with upstream (upstream has no Containerfile). The score_backend.py change is fork-local code with no upstream counterpart. The test and doc updates are purely fork-local.
ADR-0777 — Thread-Safety Audit: CUDA / SYCL / HIP Backends (2026-05-29)¶
no rebase impact: docs/research + docs/adr only; no source files were changed.
Python dep freshness sweep (ADR-0879, 2026-05-30)¶
no rebase impact: all touched files are fork-local — ai/pyproject.toml, mcp-server/vmaf-mcp/pyproject.toml, dev-llm/pyproject.toml, tools/vmaf-tune/pyproject.toml, tools/vmaf-roi-score/pyproject.toml, python/test/requirements.txt. Netflix upstream does not ship the ai/, mcp-server/, dev-llm/, or tools/vmaf-* trees; the only file shared with upstream (python/test/requirements.txt) gained a >=7.1.0 floor on pytest-cov which is purely additive over upstream's bare pytest-cov line. On rebase, keep the bumped floor; if upstream introduces its own ceiling on pytest-cov, intersect rather than overwrite.
pyright-strict-audit (2026-05-30, ADR-0888)¶
no rebase impact: REASON — all touched files are fork-added Python sources under ai/src/ and tools/vmaf-tune/src/ (and the CodecAdapter Protocol in tools/vmaf-tune/src/vmaftune/codec_adapters/__init__.py, also fork-added). No upstream Netflix/vmaf file is touched. The annotation tightening (TYPE_CHECKING torch import, assert-based Optional narrowing, cast through stub gaps, dropped dead Optional comparisons) does not change runtime behaviour — all 12 fixes are pure type-checker compliance. The audit's companion file pyrightconfig.audit.json is intentionally gitignored so this PR doesn't introduce a CI gate before the long-tail cleanup is done.
cuda-ms-ssim-double-precision-lcs (2026-06-03, ADR-0990)¶
no rebase impact: REASON — all touched files are fork-added CUDA sources (core/src/feature/cuda/integer_ms_ssim/ms_ssim_score.cu, core/src/feature/cuda/integer_ms_ssim_cuda.c) and docs. Netflix upstream does not ship a CUDA ms_ssim kernel. The invariant to preserve on any future port of ssim_accumulate_default_scalar changes from upstream: the ms_ssim_vert_lcs kernel must keep double for my_l/my_c/my_s, the shared-memory warp partial arrays, and the c1/c2/c3 parameters — these are load-bearing for the places=4 parity gate (ADR-0990 / ADR-0139). If upstream changes the scalar 2.0 * to 2.0f * (regressing to float), do NOT mirror that change into the CUDA kernel without also updating the parity test tolerance.
sycl-float-ssim-ssimulacra2-parity-research (2026-06-03, ADR-0985)¶
no rebase impact: REASON — changes are: (1) a new test file core/test/test_sycl_float_ssim_parity.c (fork-added, no upstream equivalent), (2) a meson.build test entry, (3) a research document, (4) an ADR, (5) a clarifying comment in integer_ssim_sycl.cpp (fork-added GPU kernel), and (6) state.md / changelog.d fragment updates. No CPU scalar, no public API, no Netflix upstream file is touched.
perf/arm64-float-moment-sve2 (2026-06-03, ADR-0584)¶
Rebase note: core/src/feature/float_moment.c gains an #if HAVE_SVE2 block that selects compute_1st_moment_sve2 / compute_2nd_moment_sve2 over the NEON fallback when VMAF_ARM_CPU_FLAG_SVE2 is set. The scalar default and the NEON path are unchanged; SVE2 is purely additive. core/src/meson.build gains a new arm64_moment_sve2_lib static library inside the existing if is_sve2_supported block. core/test/test_moment_simd.c gains four SVE2 test functions guarded by #if HAVE_SVE2. docs/backends/arm/overview.md updates the per-feature coverage table. No upstream Netflix/vmaf file is touched; no public C API or CLI flag changes.
shared-strict-json-helpers (2026-06-03, ADR-0988)¶
no rebase impact: REASON — all touched files are fork-added Python modules (tools/vmaf-tune/src/vmaftune/compare.py, report.py, benchmark.py; mcp-server/vmaf-mcp/src/vmaf_mcp/server.py) with no upstream Netflix/vmaf equivalent. The changes are import additions and private-function removals; no public API, no C sources, no Netflix golden-data files are touched.
sycl-motion-add-uv (2026-06-03, ADR-0989)¶
no rebase impact: REASON — all changed files are fork-added GPU backends (integer_motion_sycl.cpp, integer_motion_cuda.c, motion_vulkan.c, integer_motion_hip.c, integer_motion_metal.mm). The upstream Netflix integer_motion.c is not modified. If upstream adds motion_add_uv to integer_motion.c in a future sync, check whether the SYCL per-plane normalization formula remains consistent.
avx512-float-moment (2026-06-03, ADR-0987)¶
no rebase impact: REASON — all touched files are fork-added x86 SIMD sources (core/src/feature/x86/moment_avx512.c, moment_avx512.h) and the dispatch addition inside float_moment.c is guarded by HAVE_AVX512 / VMAF_X86_CPU_FLAG_AVX512 ifdefs that are invisible on any non-AVX-512 build path. The four new parity test cases in test_moment_simd.c are also guarded by HAVE_AVX512 and do not touch any upstream Netflix file. No public C API, no meson_options.txt entry, no CLI flag, no ffmpeg patch is changed.
CUDA motion 8-frame SAD batching (ADR-0845, 2026-05-29)¶
core/src/feature/cuda/integer_motion_cuda.c — structural change to MotionStateCuda (sad ring, score_ring, last_batch_boundary fields) and rewrite of submit/collect/flush.
Rebase impact: MEDIUM. If an upstream Netflix commit touches integer_motion_cuda.c, expect a conflict in submit_fex_cuda and collect_fex_cuda. Resolution rules: 1. Keep the batching structure (sad[] ring, batch-boundary sync in collect). 2. Apply upstream logic changes (e.g., score normalization formula, new options) to the batch-emit paths rather than the old per-frame paths. 3. Verify the ADR-0358 invariant: cuMemsetD8Async is always on pic_stream, NOT s->str. 4. The emit_batch_scores() frame_index save/restore must survive; dropping it breaks motion3 moving-average correctness.
core/src/feature/cuda/AGENTS.md — new section "Motion SAD batch fencing": keep verbatim on rebase.
ADR-0930 — Helm NetworkPolicy + PSS baseline — 2026-05-31¶
no rebase impact: REASON — every touched file lives entirely under deploy/helm/vmafx/ (a fork-local directory; Netflix upstream ships no Helm chart), plus fork-local documentation under docs/development/, docs/adr/, docs/research/, docs/state.md, docs/rebase-notes.md (this file), and changelog.d/added/. None of the production C, Go, Python, FFmpeg-patch, or Meson surfaces are touched. Future upstream syncs cannot conflict with this change.
Fork-local files: - deploy/helm/vmafx/values.yaml (UID + seccomp + networkPolicy block) - deploy/helm/vmafx/templates/networkpolicy.yaml (new) - deploy/helm/vmafx/templates/operator-deployment.yaml (inherit from .Values) - deploy/helm/vmafx/templates/tests/test-connection.yaml (inherit from .Values) - deploy/helm/vmafx/templates/NOTES.txt (PSS / NP hints) - docs/development/k8s-deployment.md (Pod security + NetworkPolicy sections) - docs/adr/0930-helm-networkpolicy-pss.md and matching _index_fragments/ entry - docs/research/0930-helm-networkpolicy-pss.md - changelog.d/added/helm-networkpolicy-pss.md - docs/state.md (closed row)
fix(hip): integer_ms_ssim_hip picture_copy normalization — 2026-06-03¶
no rebase impact: REASON — changes touch only core/src/feature/hip/integer_ms_ssim_hip.c (fork-only HIP backend, no upstream equivalent in Netflix/vmaf), docs/state.md (fork-local bug tracker), changelog.d/fixed/ (fragment), and docs/rebase-notes.md (this entry). core/test/test_hip_ms_ssim_parity.c and the corresponding meson.build entry were already present on master (merged from gpu-runtime-bug-audit). No CPU scalar path, no public header, no Netflix upstream file is touched.
feat(simd): integer-ssim-avx2 (ADR-0784)¶
Branch: feat/integer-ssim-avx2 Touches: core/src/feature/integer_ssim.c, core/src/feature/x86/integer_ssim_avx2.{c,h}, core/src/meson.build (x86_avx2_sources), core/test/test_integer_ssim_simd.c, docs/adr/0784-integer-ssim-avx2.md, docs/backends/x86/integer-ssim-avx2.md.
Adds AVX2 dispatch for the horizontal moment accumulation pass in integer_ssim.c. The dispatch uses function pointers in IntegerSsimState; the integer_ssim_moments_t struct in integer_ssim_avx2.h must stay layout-identical to ssim_moments in integer_ssim.c. Any upstream refactor of ssim_moments field order requires a matching update in the AVX2 header.
chore(ci): ci-workflow-name-shortening (ADR-0995)¶
Branch: chore/ci-shorten-workflow-names
no rebase impact: pure CI display-name rename; no C/C++/Python source touched.
fix(dnn): add missing vmaf_ort_internal_input/output_elem_type accessors¶
Branch: fix/dnn-ort-internals-missing-elem-type-accessors Touches: core/src/dnn/ort_backend_internal.h, core/src/dnn/ort_backend.c, changelog.d/fixed/dnn-ort-internals-elem-type-accessors.md.
No rebase impact: the added symbols (VmafOrtElemType, vmaf_ort_internal_input_elem_type, vmaf_ort_internal_output_elem_type) are fork-local internal-test helpers with no upstream Netflix/vmaf analogue. The VmafOrtSession.input_elem_types / .output_elem_types fields and the VMAF_HAVE_DNN guard structure they read from are also fork-local. Conflict probability on these files with upstream is zero.
chore(scripts): modernization-audit scanner — reduce false-positive noise¶
Branch: chore/modernization-audit-false-positive-filter Touches: scripts/dev/project_modernization_audit.py, scripts/dev/test_project_modernization_audit.py, changelog.d/fixed/modernization-audit-calibration-and-closed-row-noise.md.
No rebase impact: changes are confined to the developer-tools scanner and its test file. No C source, public header, upstream-mirrored Python, or Netflix golden-assertion file is touched. The new module-level constants (CALIBRATION_PLACEHOLDER_PATHS, CLOSED_SECTION_HEADINGS_RE, CLOSED_ROW_RE) and the updated scan_state_files / _marker_suppressed functions have no upstream analogue; there is no merge conflict possible.
fix(rust): vmafx-sys Default trait + Rust CI re-trigger¶
Branch: fix/rust-ci-vmafx-sys-build-dep
no rebase impact: adds impl Default for VmafContext in bindings/rust/vmafx-sys/src/safe.rs. Fork-local Rust crate with no upstream analogue; no C source, public header, or Python file is touched.
fix(perf): scaffold perf gate baseline + advisory threshold (ADR-1005)¶
Branch: fix/perf-gate-advisory-threshold-adr1005
no rebase impact: adds --advisory and --skip-if-no-baseline flags to scripts/perf/check-regression.py, updates the CI workflow step comment and flags, and adds docs/development/perf-gate.md. No C source, public header, Netflix golden assertion, or upstream-mirrored Python file is touched.
test(c): CPU feature extractor coverage push — round 3¶
Branch: test/cpu-extractor-coverage-push
no rebase impact: adds four new test-only .c files under core/test/ and wires them into core/test/meson.build. No C source, public header, upstream-mirrored Python, or Netflix golden-assertion file is touched. The new tests exercise existing extractor paths; no new symbols are introduced.
chore(cppcheck): audit + cite all cppcheck-suppress comments¶
Files touched: core/src/feature/vif.c, changelog.d/chore/cppcheck-suppress-cite-audit.md, docs/rebase-notes.md.
No rebase impact: comment-only edit to vif.c; adds [MISRA-C:2012-11.3/EXP36-C] citations to 10 bare cppcheck-suppress invalidPointerCast annotations. No logic changed; no public header, Netflix golden assertion, or upstream-mirrored symbol is affected.
test(compat-python-vmaf): coverage push round 2 — Asset + ResultStore + crossval¶
Branch: test/compat-python-vmaf-coverage-push (or equivalent worktree branch)
no rebase impact: all four new test files (test_asset.py, test_result_store.py, test_cross_validation.py, test_tools_misc.py) live exclusively under compat/python-vmaf/tests/. No C source, public header, upstream-mirrored Python, or Netflix golden-assertion file is touched. The tests exercise existing public APIs only and add no new production code paths.
chore(adr-0726): final Vulkan residual scrub — config flags + Docker + comments¶
Branch: chore/adr-0726-vulkan-residual-scrub
no rebase impact: removes dead Vulkan build-matrix rows and updates stale Vulkan references in docs, CLAUDE.md, and AGENTS.md files to past tense. No C source files changed. No public headers changed. No upstream-mirrored Python files changed. No Netflix golden assertions touched. The only structural change is removing two dead CI matrix rows that would fail anyway (meson rejects the unknown enable_vulkan option).
ADR-1011 — CUDA symbol visibility (2026-06-04)¶
No rebase impact. Adding static to TU-internal functions has no ABI or behaviour effect — all call sites are function-pointer assignments within the same translation unit. No public headers changed.
ADR-1010 — MCP server JSON parse guards (2026-06-04)¶
No rebase impact. Error-handling only — wraps two json.loads calls. No protocol, API, or tool-schema changes. Output format on the success path is unchanged.
ADR-1008 — C lifecycle + test correctness fixes (2026-06-04)¶
No rebase impact. pic_cnt increment timing change only affects error-retry callers (extremely uncommon path). Div-by-zero guard only fires when n_subsample covers all frames in range (degenerate caller). Test fixes in test_feature_collector.c and test_framesync.c are test-only with no production code change.
ADR-1007 — C string/numeric UB fixes (2026-06-04)¶
No rebase impact. All changes are guarded code paths that only fire for unusual caller-supplied values (NULL string defaults, model names shorter than 5 chars, tiny ADM frame dimensions). No public API, golden assertions, or ABI touched.
ADR-1012 — Go queue state-machine guards (2026-06-04)¶
No rebase impact. Both changes affect only the internal SQLite write path of the controller queue. No public proto/gRPC API change. Callers that receive the new 'job was cancelled before assignment' error from PullWork should retry — the controller's own retry loop already does this.
ADR-1009 — Go shutdown goroutine fixes (2026-06-04)¶
No rebase impact. WaitForShutdown drain-window change only affects shutdown timing (returns up to 30s earlier on clean shutdown). GracefulStop hard-stop fallback only fires on stuck streaming RPCs. No public API, ABI, or golden assertion touched.
fix(observability): Prometheus registry isolation + timer leak (ADR-1014)¶
Branch: fix/r5-prometheus-registry
no rebase impact: all changes are in pkg/observability/observability.go and its test file. No C source, public header, upstream-mirrored Python, or Netflix golden-assertion file is touched. The struct gains two private fields (reg prometheus.Registerer, sourcesOnce sync.Once) which are zero-valued before NewMetrics is called — no caller needs updating. The WaitForShutdown time.After → time.NewTimer + defer Stop() change is behaviour-equivalent; the only observable difference is the timer being released promptly on early return rather than at GC time.
fix(operator): Go operator resource-allocation — http.Client + gRPC dial (ADR-1017)¶
Branch: fix/r5-go-timer-ctx
no rebase impact: all changes are in cmd/vmafx-operator/internal/controller/. No public Go API, CRD schema, RBAC manifest, Helm template, or C source is touched. The SetupWithManager change is additive (adds an if r.HTTPClient == nil guard). No Netflix golden assertions touched.
fix(mcp,controller): exec.CommandContext + gRPC panic recovery (ADR-1018)¶
Branch: fix/r5-mcp-exec-ctx
no rebase impact: changes are in cmd/vmafx-mcp/impl.go and cmd/vmafx-controller/grpc_server.go. The runVmafScore Go signature change is internal to cmd/vmafx-mcp (all three callers are in the same file). No public MCP tool schema, JSON-RPC protocol, or gRPC proto file changes. No C source, public header, or Netflix golden assertions touched.
fix/y4m-dst-buf-read-sz-overflow (2026-06-04)¶
Files touched: core/tools/y4m_input.c, docs/adr/1022-y4m-dst-buf-read-sz-overflow.md, changelog.d/fixed/1022-y4m-dst-buf-read-sz-overflow.md
no rebase impact: the fix adds (size_t) casts to five arithmetic expressions in y4m_input_open_impl(). No public header is changed. No API surface changes. The upstream y4m_input.c source differs from this fork's copy (earlier ADR-0977 fixes are already in tree); cherry-picks of upstream Y4M changes will need to re-apply the same cast pattern to any newly introduced chroma branches.
fix(auth,nodes): constant-time session-token compare + JWT nbf validation (ADR-1021, 2026-06-04)¶
Files touched: cmd/vmafx-controller/nodes/registry.go, cmd/vmafx-controller/auth/middleware.go, mcp-server/vmaf-mcp/src/vmaf_mcp/http_transport.py, cmd/vmafx-controller/nodes/registry_test.go, cmd/vmafx-controller/auth/middleware_test.go, docs/adr/1021-session-token-const-time-compare.md, changelog.d/fixed/r5-crypto-const-time-session-token.md
no rebase impact: security-only bug-fix with no public API changes. All modified symbols are internal (non-exported comparison logic, JWT payload struct field). No C source, public C API header, upstream-mirrored Python, Netflix golden assertion, or ffmpeg-patches file is touched.
fix/mcp-asyncio-adr1023 (2026-06-04)¶
Files touched: mcp-server/vmaf-mcp/src/vmaf_mcp/server.py, docs/adr/1023-mcp-asyncio-correctness.md, docs/adr/README.md, changelog.d/fixed/mcp-asyncio-correctness.md
no rebase impact: changes are isolated to the MCP server Python module and docs. No C source, public C header, Netflix golden-assertion file, or upstream-mirrored Python harness is touched. The changes are pure async-safety fixes inside coroutines that already existed; no public API or tool schema changes.
fix/r5-memory-ordering (2026-06-04, ADR-1020)¶
Files touched: core/src/ref.c, core/src/ref.cpp, core/src/ref.h, core/src/feature/feature_collector.h, core/src/feature/feature_collector.cpp, core/src/picture_pool.c
If an upstream Netflix commit touches any of these files, review the following invariants before accepting:
ref.c/ref.cpp: the decrement must remainmemory_order_acq_rel; any upstream change that reverts to bareatomic_fetch_submust be re-annotated.feature_collector.h: thedestroyedfield must survive struct layout changes; all new public entry points that lockfeature_collector->lockmust add the destroyed-guard pattern immediately after the lock call.picture_pool.cfetch path: thepool->pictures[idx]copy must happen before the unlock; if upstream refactors the fetch function, preserve that ordering.
fix/integer-ssim-moments-type-non-x86 (2026-06-04, ADR-1040)¶
Files touched: core/src/feature/integer_ssim.h (new), core/src/feature/integer_ssim.c, core/src/feature/x86/integer_ssim_avx2.h
If an upstream Netflix commit adds or renames fields in the SSIM accumulation buffer, the shared header integer_ssim.h must be updated to match. The layout invariant (six consecutive int64_t fields, identical to the private ssim_moments struct) is documented in ADR-0784 and ADR-1040; any upstream layout change that breaks the direct cast in accum_row_scalar_8 / accum_row_scalar_16 requires a coordinated update to both the typedef and the cast sites.
fix/ci-go-rust-red-adr1041¶
no rebase impact: CI configuration change (go-ci.yml) and test/meson.build build guard. Neither touches public API or upstream-mirrored C code.
feat/0804-vmaf-context-get-backend¶
no rebase impact: purely additive public C API addition (new enum VmafBackend and new function vmaf_context_get_backend()). No upstream-mirrored code is modified; no existing entry points are changed. ADR-0804.
fix/dev-cuda-gpu-passthrough¶
no rebase impact: dev/docker-compose.yml only; changes default runtime: and expands capabilities for the NVIDIA passthrough. No C sources, public API, or upstream-mirrored files are touched. ADR-1053.
fix/cppcheck-vif-suppression-syntax¶
no rebase impact: comment-only change in core/src/feature/vif.c. Corrects cppcheck suppression delimiter from [...] to ; ... in 10 inline comments. No logic, no public API, no upstream-mirrored code changed.
chore/state-md-stale-open-rows-sweep-20260606¶
no rebase impact: docs/state.md only — removes 2 stale Open rows already in Recently Closed, updates 1 Open row's owner reference, and fixes a duplicate Recently Closed row. No C sources, public API, or upstream-mirrored files are touched.
test/go-vmafx-node-coverage-r6¶
no rebase impact: test-only additions to cmd/vmafx-node/executor_extra_test.go and cmd/vmafx-node/bpf/bypass_unit_test.go. No C sources, public API, upstream-mirrored files, or production Go code is changed.
fix/vendored-cjson-pdjson-depth-overflow (ADR-1061, 2026-06-06)¶
Touches two vendored sources: core/src/pdjson.c and core/src/mcp/3rdparty/cJSON/cJSON.c. Neither file exists in the Netflix upstream tree (pdjson and cJSON are fork-local additions). Rebase against Netflix/vmaf master has zero conflict risk from this change.
The PDJSON_STACK_MAX constant added to pdjson.c is a #define at the top of the file; any future vendor sync that replaces the file will need to re-apply the same define or find a better integration point.
test/svm-multiclass-realloc-and-compose-dri-lint (PR #TBD, 2026-06-06, ADR-1066)¶
no rebase impact: adds new test file core/test/test_svm_multiclass.c, new lint script scripts/ci/check-compose-dri-writable.sh, and a step in .github/workflows/dev-container-build.yml. No existing C source, public API, or upstream-mirrored file is modified.
fix/go-staticcheck-r10-timer-body (ADR-1065)¶
no rebase impact: all changes are fork-local Go files (pkg/storage/, cmd/vmafx-controller/, cmd/vmafx-server/) with no upstream-mirrored C or Python code affected. The time.NewTicker refactor is a semantic no-op for rebases; the MaxBytesReader + ReadTimeout additions are internal to the HTTP handler and do not touch any public API surface.
fix/sanitizer-exclusions-huge-alloc-tests¶
no rebase impact: only .github/workflows/sanitizers.yml and a changelog fragment are modified. No C source, public API, or upstream-mirrored file is touched.
fix/pic-prealloc-asan-leak (2026-06-06)¶
no rebase impact: single-line change in core/src/libvmaf.c setting fex_ctx->is_initialized = true before the batch-flush loop in flush_context_threaded. The change is additive — it enables the existing vmaf_feature_extractor_context_close teardown path to run correctly on shared (never-initialized) contexts. No upstream-mirrored file is modified, no public API is affected, no test fixtures change.
fix/win32-pthread-once-redefinition (2026-06-06)¶
no rebase impact: removes a duplicate block from core/src/compat/win32/pthread.h (typedef, macro, BOOL CALLBACK, and inline function). The surviving first definition is unchanged. No upstream-mirrored file is modified, no public API is affected, no test fixtures change.
fix/ubsan-enum-invalid-value-log-opt (2026-06-06)¶
no rebase impact: both changes are one-liner casts (static_cast<int>) in core/src/log.cpp and core/src/opt.cpp. Neither file is upstream-mirrored (both are C++23 rewrites of upstream C originals — ADR-0708 and ADR-0772), no public API is affected, and no test fixtures change.
test/operator-controller-coverage (2026-06-06)¶
no rebase impact: all changes are confined to cmd/vmafx-operator/internal/controller/ test files and the fix to vmafxnode_controller.go (removes the status.lastHeartbeat write — additive correctness only). No public API is affected, no upstream-mirrored file is touched, no test fixtures change. Depends on PR #759 (ADR-1069 CRD schema fix) for the envtest assertions to pass end-to-end with a real API server.
fix/disable-recurring-flaky-tests (2026-06-07)¶
no rebase impact: the change is two should_fail: true additions to core/test/meson.build and a new ADR file. No source files, no public API, no upstream-mirrored files, and no test fixtures are modified. The test binaries remain compiled; only their expected-failure polarity flips in the Meson test registry.
fix/dnn-onnx-domain-bypass (2026-06-07, ADR-1089)¶
no rebase impact: all changes are confined to fork-local files (core/src/dnn/onnx_scan.c, core/src/dnn/onnx_scan.h, core/test/dnn/test_onnx_scan.c). Netflix/vmaf has no ONNX DNN surface; no upstream-mirrored file is touched. The scanner's internal enum gains NODE_DOMAIN_FIELD = 7; the scan_node loop gains a new branch; a 50-line read_domain() helper is added. No public API or CLI surface changes.
Re-test: meson test -C build test_onnx_scan (26/26 pass).
fix/r13-gpu-dispatch-env-fast-path-data-race (2026-06-06)¶
no rebase impact: changes confined to core/src/gpu_dispatch_env.cpp. Adds std::atomic<bool> ready per-slot publication flag. No public header or ABI change; the AtomicBool type is internal to the translation unit.
worktree-wf_b08e0c22-717-2 / fix/framesync-producer-death-deadlock (2026-06-07)¶
no rebase impact: changes confined to core/src/framesync.c and core/src/framesync.h. Adds vmaf_framesync_abort() and aborted flag. The new function is internal to libvmaf; no public header change.
fix/cuda-stream-event-leak-paths (2026-06-07)¶
no rebase impact: changes confined to core/src/cuda/picture_cuda.c and four CUDA feature extractors. Graduated cleanup labels only; no new public API or ABI change.
fix/metal-buffer-ownership-leaks (2026-06-07)¶
no rebase impact: changes confined to core/src/feature/metal/float_ms_ssim_metal.mm and core/src/metal/picture_import.mm. MTLBuffer retain-count fixes only; no public header or ABI change.
fix/mcp-http-edge-cases (2026-06-07)¶
no rebase impact: changes confined to mcp-server/vmaf-mcp/src/vmaf_mcp/http_transport.py and a new test file. No C API or public header change.
fix/vmaftune-corner-cases-r14 (2026-06-07)¶
no rebase impact: changes confined to tools/vmaf-tune/src/vmaftune/cli.py and encode.py. No C API or public header change.
fix/ci-yaml-concurrency-timeout (2026-06-07)¶
no rebase impact: changes confined to 18 .github/workflows/*.yml files. No source, header, or build file is modified.
fix/ci-workflow-permissions-least-privilege (2026-06-07)¶
no rebase impact: changes confined to two .github/workflows/ files. No source, header, or build file is modified.
cov/vmafx-controller-queue-nodes-auth (2026-06-07)¶
no rebase impact: changes confined to cmd/vmafx-controller/queue/queue.go and test files. No public API or ABI change.
fix/operator-crd-status-schema-gaps (2026-06-07)¶
no rebase impact: changes confined to cmd/vmafx-operator/ and CRD YAML files. No C API or libvmaf surface changed.
fix/ms-ssim-hip-adr0990-precision-parity (2026-06-07)¶
no rebase impact: changes confined to core/src/feature/hip/integer_ms_ssim_hip.c and ms_ssim_score.hip. No public header or ABI change.
fix/ms-ssim-option-parity-hip-sycl (2026-06-07)¶
no rebase impact: changes confined to CUDA/HIP/SYCL ms-ssim extractors. No public header or ABI change.
fix/cargo-deny-bsd2-patent-allowlist (2026-06-07)¶
no rebase impact: changes confined to deny.toml. No source, header, or build file is modified.
fix/rust-pilot-clippy (2026-06-07)¶
no rebase impact: changes confined to bindings/rust/ and core/src/feature/rust/tad/. No C API or public header change.
fix/coverage-pkg-storage (2026-06-07)¶
no rebase impact: new test files only (pkg/storage/coverage_test.go, cmd/vmafx-node/bpf/coverage_test.go) + changelog fragments. No production source or public header change.
fix/mcp-streaming-backpressure-disconnect (2026-06-07)¶
no rebase impact: changes confined to cmd/vmafx-mcp/impl*.go and mcp-server/vmaf-mcp/src/vmaf_mcp/server.py. No C API or public header change.
test/r12-thread-safety-batch-tsan (2026-06-07)¶
no rebase impact: adds new test file core/test/test_thread_safety_batch.c and updates core/test/meson.build. No production source or public header change.
fix/r12-picture-ref-unref-error-path-coverage (2026-06-07)¶
no rebase impact: adds test coverage to core/test/test_picture.c only. No production source or public header change.
fix/r14-yuv-input-edge-cases (2026-06-07)¶
no rebase impact: changes confined to core/tools/y4m_input.c. No public header or ABI change.
fix/r14-cli-flag-parsing (2026-06-07)¶
no rebase impact: changes confined to core/tools/cli_parse.c and cli_parse.cpp. No public header or ABI change.
fix/test-malloc-leak-r12 (2026-06-07)¶
no rebase impact: test-only changes to free malloc'd buffers on early exit in test_framesync.c and test_pic_preallocation.c. No production source touched.
fix/test-framework-mu-assert-stderr-output (2026-06-07)¶
no rebase impact: changes confined to core/test/test.c and core/test/test.h. Fix mu_report writing to stdout instead of stderr; add missing include guard.
worktree-wf_392e91a3-897-12 / fix/ci-action-sha-consistency (2026-06-07)¶
no rebase impact: corrects inconsistent action SHAs in e2e-k8s, go-ci, and rust-ci workflows. No source, header, or build file is modified.
fix/ort-error-message-logging (2026-06-07)¶
no rebase impact: changes confined to core/src/dnn/ort_backend.c. No public header or ABI change.
fix/bench-clock-unchecked-returns (2026-06-07)¶
no rebase impact: changes confined to core/tools/vmaf.c and core/tools/vmaf_bench.c. No public header or ABI change.
fix/msvc-windows-portability-hygiene (2026-06-07)¶
no rebase impact: dead code removal in core/src/dnn/model_loader.c and core/src/feature/x86/vif_avx2.c / vif_avx512.c. No public header, ABI, or numeric change.
fix/roi-frame-bytes-odd-dims (2026-06-07)¶
no rebase impact: changes confined to core/tools/vmaf_roi.c and core/test/test_vmaf_roi.c. No public header or ABI change.
fix/vmaf-per-shot-correctness (2026-06-07)¶
no rebase impact: changes confined to core/tools/vmaf_per_shot.c and core/tools/test/test_vmaf_per_shot.sh. No public header or ABI change.
fix/compat-python-vmaf-mode-shim (2026-06-07)¶
no rebase impact: changes confined to compat/python-vmaf/__init__.py, compat/python-vmaf/config.py, and compat/python-vmaf/core/matlab_feature_extractor.py. No C source, public header, or ABI change.
fix/ffmpeg-vmaf-pre-device-full-enum (2026-06-07)¶
no rebase impact: changes confined to ffmpeg-patches/0002-*.patch and ffmpeg-patches/0014-*.patch. No C source or public header in-tree is modified.
fix/observability-otel-trace-context (2026-06-07)¶
no rebase impact: changes confined to cmd/vmafx-server/grpc_server.go, pkg/observability/otel_instruments.go, and pkg/score/grpc_client.go. No C source, public header, or ABI change.
worktree-wf_392e91a3-897-1 / fix/ai-atomic-writes (2026-06-07)¶
no rebase impact: changes confined to ai/src/aiutils/file_utils.py, ai/src/aiutils/run_manifest.py, and five AI scripts. No C source, public header, or build file is modified.
fix/helm-rolling-update-correctness (2026-06-07)¶
no rebase impact: changes confined to deploy/helm/vmafx/ Helm chart templates and values. No C source, public header, or Go source is modified.
fix/r12-dead-code-and-unused-var-after-pr-train (2026-06-07)¶
no rebase impact: removes dead code and unused variables from core/src/feature/integer_motion.c and core/test/test_framesync.c. No public header or ABI change.
fix/motion-coverage-picture-ref-include (2026-06-07)¶
no rebase impact: changes confined to core/test/test_integer_motion_coverage.c. Test-only change. No production source or public header modified.
fix/state-sweep-fix — CI build-matrix + ASan + motion_v2 coverage (2026-06-07)¶
libvmaf-build-matrix.yml: four meson setup / ninja invocations updated from libvmaf to core — ADR-0700 rename follow-up. Conflicts possible if another branch also edits those four lines; resolve by keeping core as the source dir. tests-and-quality-gates.yml: ASAN_OPTIONS: allocator_may_return_null=1 added to the sanitizer step env; no conflict risk. core/src/picture.c and core/src/picture.h: vmaf_picture_pool_flush() added; no conflict risk (new symbol). core/test/test_integer_motion_v2_coverage.c: manual prev_ref assignments and memset calls removed; tests now call extract in a plain loop. Conflicts possible if another branch edits the same test functions; resolve by keeping the wrapper-managed approach (no manual prev_ref assignment).
fix/ffmpeg-vulkan-ci-job-removal (2026-06-07)¶
no rebase impact: only .github/workflows/ffmpeg-integration.yml modified (dead job removed) and docs/state.md + changelog.d/fixed/ffmpeg-vulkan-ci-job-removal.md added. No production source, public header, or meson build files touched.
docs/phase4b9-container-only-publishing (2026-06-08)¶
no rebase impact: docs-only change. CLAUDE.md §15 updated with a new publishing bullet; docs/development/publishing.md and docs/adr/1102-*.md added; docs/adr/README.md gets one new index row; changelog.d/added/1102-*.md added. No production source, public header, meson build files, or test files touched.
fix/codeql-large-parameter-const-pointer (2026-06-12)¶
Rebase-sensitive: function signatures changed across VIF, SpEED, and CAMBI. - VifBuffer: PADDING_SQ_DATA, PADDING_SQ_DATA_2 (integer_vif.h), pad_top_and_bottom, decimate_and_pad, subsample_rd_8/16 (integer_vif.c), vif_subsample_rd_8/16_avx2 + vif_filter1d_8/16_avx2 (x86/vif_avx2.h/.c), vif_subsample_rd_8/16_avx512 (x86/vif_avx512.h/.c), vif_subsample_rd_8/16_neon (arm64/vif_neon.h/.c): all now take const VifBuffer * instead of VifBuffer by value. The function-pointer typedef in VifState was updated accordingly. Call sites pass &buf or &s->public.buf as appropriate. - SpeedDimensions (speed.c): 11 static functions now take const SpeedDimensions *. Call sites in speed_extract_score pass &s->dimensions. - SpeedInternalDimensions (speed_internal.h/.c): speed_internal_filter_and_downscale and speed_internal_compute_cov_matrix now take const SpeedInternalDimensions *. All GPU backend call sites (cuda/hip/sycl speed twins) pass &s->dim. - CambiBuffers (cambi.c): cambi_score now takes const CambiBuffers *. Call site passes &s->buffers. Conflicts possible if upstream or another branch edits these function signatures or adds new call sites; resolve by carrying the const * form forward.
docs/index.md — Vulkan image-import list item removed (docs/remove-stale-vulkan-image-import-ref)¶
no rebase impact: docs-only removal of a stale list item; no code or nav structure changed.
fix/functional-matrix-broken-17 (2026-06-12)¶
no rebase impact: all changes are bug fixes in independent files (bench_all.sh, bisect.py, server.py, cli.py, op_allowlist.c, float_adm_cuda.c, dnn_api.c, Containerfile) with no shared function-signature changes, no renamed symbols, and no upstream-mirrored path modifications.
rc/scaffold-stub-completion — picture_v2 implementation + ai/scripts exit-code fix¶
Files touched: core/src/picture_v2.c (new), core/test/test_picture_v2.c (new), core/include/libvmaf/picture_v2.h, core/src/meson.build, core/include/libvmaf/meson.build, core/test/meson.build, docs/architecture/vmaf-picture-v2-migration.md, ai/scripts/{gen_calibration,quantize_int8,build_calibration_set,eval_loso_fr_regressor_v2, external_benchmark_pvmaf,fetch_lsvq,gen_dists_sq_placeholder_onnx, gen_mobilesal_placeholder_onnx,gen_ssimulacra2_eotf_lut,hdrsdr_vqa_to_corpus_jsonl, my_corpus_to_corpus_jsonl,train_fr_regressor_v4,train_video_saliency_student}.py.
Rebase impact: low. picture_v2.c is a new file; no upstream file was modified. If upstream ever adds its own picture_v2.c (unlikely given it is a fork-local concept), resolve by keeping the fork's implementation. The ai/scripts exit-code fix is purely script-internal; no C or build system conflict possible with upstream.
fix/rc-gate-three-infra-test-bugs (2026-06-13)¶
no rebase impact: all changes are test-only. cmd/vmafx-mcp/server_test.go drops t.Parallel() from one test function (no C/API change). core/test/meson.build adds a TSAN_OPTIONS env entry alongside an existing ASAN_OPTIONS entry for one test. core/test/dnn/test_cli.sh adds a DNN-availability probe near the top of the script. None of these files exist in upstream Netflix/vmaf; no upstream merge conflict is possible.
chore/remove-vulkan-moltenvk-dead-leftovers (2026-06-13)¶
no rebase impact: deletions only (subproject wraps, Docker stages, CI job bodies, moltenvk.md). No shared function signatures changed, no symbols renamed, no upstream-mirrored C paths modified. ABI-reserved enum gaps preserved.
fix/hip-vif-mirror2-boundary (2026-06-13)¶
no rebase impact for upstream syncs: touches only fork-local HIP files (core/src/feature/hip/integer_vif/vif_statistics.hip) and the fork-local HIP VIF parity test (core/test/test_hip_vif_parity.c). Neither file exists in upstream Netflix/vmaf. The docs/adr/ and docs/backends/hip/overview.md changes are also fork-local. Any future upstream sync that adds upstream files under core/src/feature/hip/ would require manual review of boundary semantics, but no mechanical conflict is possible.
feat/pelorus-vendor-interop-abi (2026-06-14)¶
no rebase impact for upstream Netflix/vmaf syncs: every file is fork-local and has no upstream counterpart. New vendored mirror under core/src/interop/pelorus_*.c + core/include/libvmaf/pelorus/*.h (sourced from VMAFx/pelorus@835e097, NOT Netflix upstream), the conformance fixture core/test/test_pelorus_interop.c, scripts/sync-pelorus-interop.sh, and the docs/ADR/changelog/state deliverables. The core/src/meson.build and core/test/meson.build edits append to fork-local lists (the libvmaf source list and the test registrations) and do not touch upstream-mirrored build logic.
Rebase-sensitive invariant (cross-repo, NOT upstream): the vendored files are a byte-identical mirror pinned to VMAFx/pelorus@835e097. Never hand-edit them — clang-format/clang-tidy are deliberately excluded for these paths (dir-local .clang-tidy, .cppcheck-suppressions.txt, the make format path filters, the assertion-density skip, and the auto-format-on-edit.sh PostToolUse hook skip). A pelorus ABI bump is re-synced via scripts/sync-pelorus-interop.sh --update (which also bumps the pin), never by editing the mirror in place. If a future change rewrites these files, re-run the sync guard + the conformance fixture before merging.
feat/pelorus-autotune-control-plane (2026-06-14)¶
no rebase impact for upstream syncs: touches only fork-local files under tools/vmaf-tune/ — a new src/vmaftune/filter_adapters/ package (__init__.py, pelorus_deband.py), a new src/vmaftune/prefilter.py module, three new test files under tests/, plus additive edits to src/vmaftune/cli.py (new imports, a prefilter subparser, a _run_prefilter handler, and one dispatch line). None of these exist in upstream Netflix/vmaf — tools/vmaf-tune/ is entirely fork-added. No shared function signatures changed and no upstream-mirrored paths were touched. The cli.py edits are append-only at well-separated sites (import block, subparser-registration block, handler block, dispatch block), so even a fork-internal rebase against a newer cli.py resolves cleanly. External coupling note: the 10 deband knobs are a verbatim copy of the Pelorus ADR-0110 control-plane contract; a contract change on the Pelorus side requires a matching edit to filter_adapters/pelorus_deband.py in a coordinated two-repo PR (the conformance test fails on drift).
feat/golusoris-server (2026-06-14)¶
no rebase impact for upstream Netflix/vmaf syncs: every file touched is fork-local Go and has no upstream counterpart. The change rewrites cmd/vmafx-server/*.go (the Go gRPC + HTTP scoring service) onto the golusoris fx framework (ADR-1119), updates Dockerfile.go-server, docs/server/grpc.md, and docs/usage/env-vars.md for the env-var rename, and adds an app_test.go fxtest lifecycle test. None of these exist in upstream; libvmaf's C sources, public headers, and the Netflix golden gate are untouched.
Rebase-sensitive invariants (fork-internal, NOT upstream): - R1 cgo-lifetime stop order. The composition root forces the *libvmaf.Scorer to be constructed BEFORE the golusoris *grpc.Server (an fx.Invoke(func(_ *libvmaf.Scorer) {}) registered ahead of the gRPC service-registration invoke, plus scorer-first arg order in that invoke). fx runs OnStop hooks in reverse construction order, so this guarantees the gRPC server's GracefulStop drains in-flight Score calls before the scorer's Close() releases C resources. TestStopOrderScorerAfterGRPC pins this; do not reorder those invokes or flip the arg order without re-deriving the ordering (see the empirical probe rationale in the PR). - go.mod pin. github.com/golusoris/golusoris stays at v0.3.1. The fx migration only adds transitive // indirect deps (go-grpc-middleware/v2) and promotes go-chi/chi/v5 from indirect to direct via go mod tidy; it does not bump the golusoris pin or touch internal/app/bootstrap. - Env-var contract. VMAFX_HTTP_ADDR / VMAFX_GRPC_LISTEN map to the golusoris http.addr / grpc.listen keys under the VMAFX_ prefix. If golusoris renames those keys, the server's documented env contract must follow.
feat/golusoris-node (2026-06-15)¶
no rebase impact for upstream Netflix/vmaf syncs: every file touched is fork-local Go and has no upstream counterpart. The change rewrites cmd/vmafx-node/main.go (the Go gRPC worker root) onto the golusoris fx framework (ADR-1119, Phase-1 PR-3), adds cmd/vmafx-node/providers.go (the fx domain providers) and cmd/vmafx-node/scoring_handler.go (the VmafxScoring impl moved out of the now-removed cmd/vmafx-node/server package), refactors cmd/vmafx-node/online_feedback.go (Start/Close lifecycle), updates docs/usage/env-vars.md for the env-var rename, and adds app_test.go + app_scorestream_test.go (fxtest lifecycle + end-to-end ScoreStream). None of these exist in upstream; libvmaf's C sources, public headers, and the Netflix golden gate are untouched. The eBPF loader under cmd/vmafx-node/bpf/ is unrelated to golusoris and was not touched.
Rebase-sensitive invariants (fork-internal, NOT upstream): - R-node lifecycle stop order. The composition root forces the *libvmaf.Scorer to be constructed first, then the *FeedbackClient + *Executor (a lazy-provider guard fx.Invoke(func(_ *FeedbackClient, _ *Executor) {})), then the golusoris *grpc.Server (a standalone fx.Invoke(func(_ *grpc.Server) {}) lazy-provider guard). fx runs OnStop hooks in reverse construction order, so this guarantees: gRPC GracefulStop drains in-flight Score / ScoreStream calls → FeedbackClient drainer stops → scorer Close(). TestStopOrderNode (app_test.go) pins this against the REAL hook firing order; do not reorder those invokes or flip arg order. - Lazy-provider listener guard. grpc.Module's listener binds in an OnStart hook that only runs if *grpc.Server is consumed. The standalone fx.Invoke(func(_ *grpc.Server) {}) is load-bearing — remove it and the node serves nothing. TestAppStartsAndBinds dials the bound addr to prove it. - FeedbackClient drainer lifetime. NewFeedbackClient(log) no longer takes a context or spawns a goroutine; Start() launches the drainer (bound to an internal, Close-owned context) and Close() stops + awaits it. Both are idempotent. Wired to fx OnStart/OnStop in provideFeedbackClient. - go.mod pin. github.com/golusoris/golusoris stays at v0.4.0. The fx migration only adds the transitive // indirect dep go-grpc-middleware/v2 v2.3.3 via go mod tidy; it does not bump the golusoris pin or touch internal/app/bootstrap. - Env-var contract. VMAFX_GRPC_LISTEN maps to the golusoris grpc.listen key under the VMAFX_ prefix (replaces VMAFX_NODE_ADDR). If golusoris renames that key, the node's documented env contract must follow.
feat/golusoris-operator (2026-06-15)¶
no rebase impact for upstream Netflix/vmaf syncs: every file is fork-local and has no upstream counterpart. cmd/vmafx-operator/main.go is rewritten from a hand-rolled controller-runtime entry point onto the golusoris fx framework (ADR-1119 Phase 1), and cmd/vmafx-operator/main_test.go adds fx-graph validation. The reconcilers under cmd/vmafx-operator/internal/controller/ and the webhooks under cmd/vmafx-operator/internal/webhook/ are unchanged; only their wiring (Setup-against-manager) moved into fx.Invoke hooks. None of these files exist upstream.
Rebase-sensitive invariant (cross-repo, NOT Netflix upstream): this migration requires github.com/golusoris/golusoris >= v0.4.0, because the github.com/golusoris/golusoris/k8s/operator module (introduced by golusoris commit 3df9f1a / PR #224) first appears in tag v0.4.0 and is ABSENT in v0.3.1. The foundation commit (afd66c7ef) pins v0.3.1, which predates k8s/operator — so cmd/vmafx-operator/main.go does not compile until the go.mod golusoris pin is bumped to v0.4.0+. The pin bump is intentionally NOT part of this branch (it is a shared go.mod change owned by the migration orchestrator).
golusoris#227 note: the in-tree main.go does NOT add an app-level ctrl.SetLogger shim — golusoris v0.4.0's operator.Module already calls ctrl.SetLogger(loggerFromSlog(logger)) inside newManager, so a second SetLogger from the app would be redundant. If a future golusoris release reverts that (regressing #227), re-add the shim as an fx.Invoke(func(l *slog.Logger){ ctrl.SetLogger(logr.FromSlogHandler(l.Handler())) }). Likewise webhooks are wired via operator.Options.WebhookPort (also added post-v0.3.1); if that field disappears upstream, the app must stand up its own webhook.NewServer and add it to the manager.
feat/golusoris-mcp (2026-06-15)¶
no rebase impact for upstream Netflix/vmaf syncs: this PR rewrites only cmd/vmafx-mcp/main.go (the Go MCP server composition root) onto the golusoris fx framework (ADR-1119, Phase-1 PR-5), plus the fork-local docs/changelog/AGENTS deliverables. cmd/vmafx-mcp/ is entirely fork-added and has no upstream counterpart. The MCP tool surface (tools.go, impl.go, impl_direct.go, server.go) is byte-unchanged, and no test file changed.
- Composition root. The hand-rolled
flag.Parse+signal.NotifyContext - bespoke stdio/HTTP transport loops + custom
observability.InitOTelare replaced byfx.New(bootstrap.Base, fx.Replace(config.Options{...}), bootstrap.FxLogger(), fx.Provide(buildMCPServer), fx.Invoke(runMCPTransport)).Run(). Mirrorscmd/vmafx-server/main.goandcmd/vmafx-node/main.go. Because the MCP server is NOT a golusoris server module (golusoris ships no MCP module), the transport is owned in therunMCPTransportlifecycle hook rather than by a framework module — if golusoris later adds an MCP module, fold the hook into it. - bootstrap dependency. This PR consumes
internal/app/bootstrap.Baseandbootstrap.FxLogger()but does NOT modify them; it shares the bootstrap stanza with the sibling fx migrations (#932/#934/#935/#936). A rebase that reshapesbootstrap.Base(e.g. when golusoris#226 ships a version module, or golusoris#234's LOG_LEVEL prefix-read lands and the env bridge can be deleted) must re-checkmain()here too. - Env bridge (interim).
main()bridgesVMAFX_LOG_LEVEL → LOG_LEVELandVMAFX_LOG_FORMAT → LOG_FORMATbeforefx.New, identical to the sibling binaries (golusoris#234). Delete all four bridges across the cmd/ tree in one sweep once the carrying golusoris tag lands. - Env-var / flag contract change.
--transport/--portflags removed; replaced byVMAFX_MCP_TRANSPORT(mcp.transport, defaultstdio) andVMAFX_MCP_HTTP_ADDR(mcp.http.addr, default:3000).VMAF_BINandVMAFX_MCP_DIRECTare read directly by the tool handlers (not via koanf) and are unchanged. - Rebase-sensitive invariant — stdio-stdout purity. Nothing in the fx graph may write to stdout in stdio mode (the JSON-RPC framing owns it). golusoris log → stderr,
otel.Moduleis OTLP-gRPC (no stdout),bootstrap.FxLogger()→ slog → stderr. A future rebase that adds anfx.Print-style logger, a stdout OTel exporter, or anyfmt.Printlnto the composition root MUST gate it off in stdio mode. See cmd/vmafx-mcp/AGENTS.md invariant #11.
feat/golusoris-controller (2026-06-15)¶
no rebase impact for upstream Netflix/vmaf syncs: every file touched is fork-local Go and has no upstream counterpart. The change rewrites cmd/vmafx-controller/*.go (the Go gRPC + HTTP controller: SQLite job queue + node registry + FIFO scheduler + JWT auth) onto the golusoris fx framework (ADR-1119, Phase-1 PR-2), refactors cmd/vmafx-controller/nodes/registry.go, adds an app_test.go fxtest lifecycle suite, and updates docs/usage/env-vars.md for the env-var rename. libvmaf's C sources, public headers, and the Netflix golden gate are untouched; no ffmpeg-patch impact.
Rebase-sensitive invariants (fork-internal, NOT upstream): - golusoris pin → v0.4.1. The PR depends on golusoris#225 (grpc.ProvideServerOption, used to chain the JWT auth interceptors), which is NOT in the v0.4.0 tag the shared go.mod currently pins. The committed go.mod keeps the v0.4.0 pin (no replace directive); the binary + app_test.go will not compile until the orchestrator bumps the pin to the tag carrying #225 (expected v0.4.1). go mod tidy against v0.4.0 is otherwise clean — the migration only adds the transitive // indirect go-grpc-middleware/v2 (golusoris HEAD's grpc.Module recovery/logging interceptors). - R1 stop order. The composition root forces the *libvmaf.Scorer, the SQLite queue.Queue, and the *nodes.Registry to be constructed BEFORE the golusoris *grpc.Server (an fx.Invoke(func(_ *libvmaf.Scorer, _ queue.Queue, _ *nodes.Registry) {}) registered ahead of the gRPC service-registration invoke). fx runs OnStop hooks in reverse construction order, so this guarantees the gRPC GracefulStop drains in-flight RPCs before the queue Close, the node-registry reaper stop, and the scorer Close. TestStopOrder pins this; do not reorder those invokes without re-deriving the ordering. - nodes.Registry lifecycle. NewRegistry(log) no longer takes a context or spawns the reaper at construction; the reaper is launched by Start(ctx) (fx OnStart) and stopped + awaited by Close() (fx OnStop). Every call site (production + tests) must drive Start/Close via the lifecycle rather than passing a caller context. - gen/go/controller proto types are hand-written. Unlike the protoc-generated gen/go (scoring) types, gen/go/controller/*.pb.go are hand-maintained and do NOT implement the protobuf-v2 reflection interface, so VmafxController messages cannot be marshaled by the standard gRPC wire codec. In-process handler tests are unaffected; over-the-wire fxtests therefore use the VmafxScoring service. See the orchestrator note below — regenerating gen/go/controller with buf/protoc is the proper fix and is a candidate follow-up. - Env-var contract. VMAFX_HTTP_ADDR / VMAFX_GRPC_LISTEN map to the golusoris http.addr / grpc.listen keys; auth.tenant_claim / auth.roles_claim are golusoris CompoundKeys (preserve the underscore). If golusoris renames those keys, the controller's documented env contract must follow.
fix/mcp-probe-parity (2026-06-15)¶
no rebase impact: edits the fork-only MCP servers (cmd/vmafx-mcp/{impl.go,tools.go,impl_test.go,AGENTS.md}, mcp-server/vmaf-mcp/src/vmaf_mcp/server.py) + docs/mcp/tools.md + docs/state.md + changelog. No libvmaf C-API / CLI / meson_options.txt / public-header change → no ffmpeg-patch impact. Rebase-sensitive invariant — probe_backend parity (cmd/vmafx-mcp/AGENTS.md invariant #12): the Go handleProbeBackend and the Python _probe_backend MUST share the same 64×64 (≥36px/dim, CUDA-ADM minimum) synthetic probe frame AND the same runtime_healthy predicate (null/non-finite vmaf.mean → runtime_healthy=false, error "vmaf returned exit 0 but score was null"). A rebase that touches either probe handler must keep the two in lock-step.
fix/bughunt-feature-cpu — CIEDE 4:2:2 chroma-upsample flag swap + cambi init leak (2026-06-27)¶
Rebase impact: DIVERGES from upstream on ciede.c — read carefully on the next sync. core/src/feature/ciede.c scale_chroma_planes / scale_chroma_planes_hbd is a near-verbatim upstream Netflix file, and upstream carries the identical transposition bug: the horizontal sample index keys off ss_ver and the vertical row advance keys off ss_hor. This fix swaps them to the correct chroma-subsample math (horizontal → ss_hor, vertical → ss_ver). Rebase-sensitive invariant for the next syncer: do NOT let an upstream sync silently revert this — if Netflix re-pulls the transposed lines, keep the fork's corrected flags. The bug only manifests on YUV422P (heap OOB + wrong ciede2000 scores); YUV420P is a no-op (both flags set) and YUV444P never calls the function, so the Netflix golden CIEDE2000 pair (420P) is bit-identical either way. Guarded by test_ciede_scale_chroma_422_8b / _16b in core/test/test_ciede.c, which fail against the buggy/upstream form. The core/src/feature/cambi.c change is a fork-internal error-path unwind (route init() failures through close_cambi()); cambi is upstream-mirrored but the change is confined to the -ENOMEM / -EINVAL error paths with no success-path or scoring delta — a sync that re-pulls cambi init() should re-apply the goto fail unwind. No public header, CLI, meson-option, or ffmpeg-patch surface changes.
fix/bughunt-simd (2026-06-27)¶
no rebase impact: edits fork-added SIMD float_moment paths only (core/src/feature/arm64/moment_sve2.c, core/src/feature/x86/moment_avx2.c, core/src/feature/x86/moment_avx512.c, core/src/feature/arm64/moment_neon.c) + the fork test core/test/test_moment_simd.c. No libvmaf C-API / CLI / meson_options.txt / public-header change -> no ffmpeg-patch impact. No Netflix golden assertion touched. Rebase-sensitive invariant — moment SVE2 lane mapping (core/src/feature/arm64/AGENTS.md): the SVE FCVT .s->.d (svcvt_f64_f32) widens the EVEN-indexed f32 lanes (source element 2i), NOT the lower contiguous lanes; the odd lanes must be widened with the SVE2 FCVTLT (svcvtlt_f64_f32, source element 2i+1). Any future edit to moment_sve2.c must keep the even+odd dual-convert (stepping a full svcntw() register) or it will silently double-count even lanes and drop odd lanes on >128-bit SVE.
fix/bughunt-dnn (2026-06-27)¶
no rebase impact: edits the fork-only DNN tiny-AI / ONNX Runtime path (core/src/dnn/tensor_io.c, core/src/dnn/ort_backend.c) and its fork-only tests (core/test/dnn/test_tensor_io.c, core/test/dnn/test_ort_internals.c) + docs/state.md + changelog. No libvmaf public-header / CLI / meson_options.txt change, so no ffmpeg-patch impact. The whole core/src/dnn/ tree is fork-added (not present upstream Netflix/vmaf), so there is no upstream-parity conflict surface to track on a future sync.
fix/bughunt-ai (2026-06-27)¶
no rebase impact: edits training-harness Python only (ai/scripts/{aggregate_corpora,extract_k150k_features,materialize_saliency_features}.py, ai/train/konvid_pair_dataset.py + ai/tests/). No libvmaf C-API / CLI / meson_options.txt / public-header change → no ffmpeg-patch impact. No rebase-sensitive invariants.
fix/bughunt-cli (2026-06-27)¶
no ffmpeg-patch impact: edits the CLI (core/tools/vmaf.cpp, cli_parse.cpp) only. Deleted dead core/tools/vmaf.c (unreferenced; superseded by vmaf.cpp) + re-pointed 8 stale config/doc refs. Invariant: cli_parse.c is NOT dead — it is the TU compiled into test_cli_parse / test_cli_parse_long_only_args / fuzz_cli_parse; do not delete it on rebase. --help→stdout/exit0, no-frames→exit 101 (VMAF_EXIT_NO_FRAMES_DECODED).
chore/version-3.2.0 (2026-06-27)¶
no rebase impact: bumps core/meson.build version (x-release-please-version) + .release-please-manifest.json . to 3.2.0 / 3.2.0-lusoris.0 to track upstream libvmaf 3.2.0 SONAME. On an upstream sync, keep the fork's libvmaf version aligned with Netflix's (<upstream-X.Y.Z>-lusoris.N).
fix/round3-build-gpu-batch (2026-06-27)¶
no ffmpeg-patch impact. R3-6 HIP integer_vif uninit-err (init err=0). R3-9 NVTX libdl → cc.find_library('dl'). R3-10 ssim AVX2 carve-out + _x86_simd_strict_fp_extra (icx -fp-model=precise; no-op on gcc/clang). Invariant: every x86 SIMD carve-out lib that needs bit-exactness under icx must carry _x86_simd_strict_fp_extra; keep the ssim carve-out aligned with its psnr_hvs/ms_ssim/ssimulacra2 siblings.
feat/vmafx-tune-go-stage5-per-shot (2026-08-30)¶
no ffmpeg-patch impact: Go-only change under cmd/vmafx-tune/ and pkg/. No libvmaf C-API, public header, CLI flag, or meson_options.txt surface is touched, so nothing in ffmpeg-patches/ consumes it. No Netflix golden assertion touched; no Python removed (ADR-0703 / ADR-0704 sunset stays gated on full Go parity).
Rebase-sensitive invariants (full text in cmd/vmafx-tune/AGENTS.md #13-17):
pkg/pershot/plan_json.gowire-struct field order is alphabetical by JSON key — that is what reproduces Python'sjson.dumps(..., sort_keys=True). Reordering the fields ofplanWire/shotWirefor readability silently breaks byte-parity with the Python emitter.pyFloat(Pythonrepr()form:24.0, not Go's24),ensureASCII(ensure_ascii=True) andSetEscapeHTML(false)are part of the same contract.TestRenderPlanJSON_GoldenMatchesPythonis the guard.pkg/encoder/adapter.gocarries two quality windows per codec and they are not interchangeable —AbsoluteLo/Hiis the bisect search domain (ADR-0538),QualityLo/Hithe informative window the per-shot tuner clamps into. They differ forlibx265andlibsvtav1.AdapterEncoder.CRFRange()returns the absolute pair on purpose.EncodeParams.InputArgs(pre--i) vsExtraArgs(post--c:v) is a placement contract — demuxer options and-init_hw_deviceare rejected by ffmpeg anywhere but the pre-input position (ADR-0601);-vfmust stay post-input. Do not merge the two fields..y4mis deliberately absent frompkg/bisectrawYUVSuffixes— vmaf-tune always passes explicit geometry, which flips libvmaf'suse_yuvbranch, and a Y4M header then trips the file-size guard inraw_input_open(ADR-0499).--predicate-module/--fast-nrfail fast rather than being ignored — they exist for CLI-surface parity with the Python parser but have no Go implementation. If an ONNX Go binding lands,--fast-nrgraduates first.
feat/vmafx-tune-go-corpus-sidecar (2026-08-30)¶
no ffmpeg-patch impact: adds Go-only packages (pkg/{codecadapter,corpus,predictor,sidecar,pyjson}) plus two cmd/vmafx-tune subcommands. No libvmaf C-API / public-header / CLI-flag / meson_options.txt change, so nothing the ffmpeg-patches/ series consumes moves. Invariants (full list in pkg/corpus/AGENTS.md): (1) pkg/corpus.RowKeys mirrors vmaftune.CORPUS_ROW_KEYS in order — the canonical-6 columns are indexed positionally downstream; (2) corpus rows render through pkg/pyjson, never encoding/json, because a row carries bare NaN tokens and CPython repr()-style floats; (3) the float aggregates go through pkg/corpus/pysum.go (Neumaier sum() + exact-rational statistics.pstdev()), not naive Go loops — a plain accumulator drifts by a ULP, which is a visible byte difference in the JSONL; (4) .y4m must stay out of vmafRawSuffixes (ADR-0499 Bug #V3-B). The Python tools/vmaf-tune/ tree is untouched and remains the shipped implementation until the ADR-0703 / ADR-0704 sunset.
feat/vmafx-tune-go-auto-sidecar (2026-08-30)¶
no ffmpeg-patch impact: adds Go-only packages under pkg/tune/ plus two new cmd/vmafx-tune/cmd/ subcommands. No libvmaf C-API, CLI, public-header, or meson_options.txt change, and no Netflix golden assertion touched. The Python tools/vmaf-tune/src/vmaftune/ tree is untouched — this work makes the ADR-0703 / ADR-0704 sunset possible, it is not the sunset.
Rebase-sensitive invariants (full list in pkg/tune/AGENTS.md):
-
The
autoplan JSON is deliberately not strict RFC 8259. The Python emitter uses plainjson.dumps(..., sort_keys=True), whose defaultallow_nan=Truewrites a bareNaNfor the uncalibrated conformalinterval_widthevery non-smoke run produces.pkg/tune/pyjsonreproduces that, plus CPython'srepr()float spelling (mandatory.0, the fixed/exponential switch atdecpt <= -4 || decpt > 16) andensure_ascii=Trueescaping. Do not "fix" it withencoding/jsonorMarshalStrict; the--executeJSONL rows are the strict surface, and that asymmetry is intentional on both sides. -
pkg/tune/pymathis the float-parity layer, not an optimisation. Go'smath.Powandmath.Log10land one ULP from the libm CPython calls, and both feed user-visible JSON (estimated_bitrate_kbps,estimated_vmaf). Reverting either to the stdlib fails the parity fixtures. -
Short-circuit evaluation order is part of the output contract (
plan.metadata.short_circuits). Append predicates; never reorder the ten. -
The content recipe must fire before rung selection so
force_single_rungcan collapse a 4K ladder. -
The sidecar feature-vector column order pins every persisted weight (
FeatureDim = 14, stops atWidth;Heightis deliberately absent). Changing it requires aSchemaVersionbump or oldstate.jsonfiles load mis-aligned. -
A predictor-version mismatch must cold-start the sidecar — that is what makes a shipped-model upgrade safe.
-
The host UUID is CSPRNG-random, never machine-derived.
-
The subprocess seam (
hdr.Runner,executor.Runner) is load-bearing: the whole suite runs without ffprobe / ffmpeg / vmaf installed. A non-zero exit is reported in the result; theerrorreturn means a spawn failure.
Parity fixtures under pkg/tune/*/testdata/ were dumped from the in-tree Python modules. Regenerate them only alongside a deliberate coordinated change on both sides — a silent regeneration turns the gate into a tautology.
vmafx-tune Go port integration (landed 2026-08-30)¶
Go-only; no upstream Netflix/vmaf counterpart, so no rebase conflict surface. Two invariants a future change must not undo, both recorded in ADR-1125:
- (Superseded by ADR-1137, 2026-09-02 — both packages are deleted and
pkg/pyjson.Options.NonFiniteselects between the two spellings.)internal/pyjsonandinternal/pyjsonstrictare deliberately two packages, not an accident of the parallel port. They mirror two different Python entry points:json.dumps(bareNaN/Infinitytokens) andvmaftune.jsonio.dumps_strict(non-finite →null, valid RFC 8259). "Deduplicating" them makes one package answer to two output contracts and silently changes the payload of whichever subcommands lose their encoder. codecadapter.ResolveCodecArgs(package-level) validates the preset;(*Adapter).ResolveCodecArgs(method) does not. That asymmetry is load-bearing — the package-level form carries the Python contract where an out-of-vocabulary preset is an error, while the method is the low-level token builder used on already-validated input. Making the method validate breaks the group-6 call paths; making the function skip validation breaksTestBuildFFmpegCommandRejectsBadPreset.
Also: .gitattributes now exempts pkg/benchmark/testdata/*.csv from text=auto. Those goldens assert CRLF (Python's csv default). Re-normalising them makes the benchmark suite fail on fresh checkouts only, which is a slow failure to diagnose.
C++23 twin wiring (Waves 1–5, landed 2026-08-30)¶
Twelve core/src translation units moved from .c to .cpp: cpu, dict, mem, output, ref, thread_locale, fex_ctx_vector, feature_name, luminance_tools, mkdirp, picture_copy, psnr_tools. The .c twins are deleted, so an upstream Netflix patch touching any of those paths will not apply directly — port the hunk into the .cpp file instead of restoring the .c.
Several internal headers gained #ifdef __cplusplus / extern "C" guards (feature/alias.h, output.h, fex_ctx_vector.h, feature/mkdirp.h, feature/psnr_tools.h, test/test.h). The guards are inert for C consumers.
Lesson worth keeping: an unreferenced twin diverges silently. output.c got the ADR-0602 NULL-guard fix while output.cpp did not, and nothing caught it because nothing built output.cpp. A CI check that fails on any .c/.cpp pair where one side is unreferenced would prevent a recurrence.
C23 + C++26 standard bump (ADR-0692, landed 2026-08-30)¶
core/meson.build sets c_std=c23 (was c11) and emits -std=c++26 (was c++23) on every non-MSVC compiler. When rebasing upstream Netflix/vmaf C sources, note that C23 changes the meaning of an empty parameter list: void f() declares void f(void) rather than an unprototyped function. Upstream code carrying K&R-style empty parameter lists will produce type errors here that it does not produce upstream — give such functions their real prototypes rather than reverting the standard. -Wimplicit-fallthrough is also enabled fork-wide, so an unannotated switch fallthrough in ported code needs an explicit [[fallthrough]].
FFmpeg n8.1.1 → n9.0.1 (landed 2026-08-30)¶
The patch stack now targets n9.0.1. Verified by full series replay: all 17 patches apply at full context, cumulatively, against a clean n9.0.1 checkout.
Two patches were regenerated for line drift only — 0002-add-vmaf_pre-filter and 0008-add-libvmaf_tune-filter. The latter drifted because FFmpeg 9 inserted OBJS-$(CONFIG_FRC_AMF_FILTER) between VPP_AMF_FILTER and VPP_QSV_FILTER, which sat inside that hunk's trailing context. Neither regeneration changed a single line of added code.
Note when replaying the series yourself: it must be applied cumulatively. Patches 0002–0006 and 0008 depend on state introduced by earlier patches (0008's context includes CONFIG_VMAF_PRE_FILTER, which patch 0002 adds), so a per-patch git apply --check against pristine upstream reports false failures. Use ffmpeg-patches/test/build-and-run.sh or a sequential git am --3way chain. A shallow (--depth 1) clone also breaks git am --3way, which needs pre-image blobs — clone with full history when replaying.
Lint / format gate repair (landed 2026-08-30)¶
Makefile is shared with upstream Netflix/vmaf, so this is rebase-sensitive.
The fork's lint-* and format-check targets are fork-added (upstream has no equivalent), but they live in the same file upstream edits. Three fork-local constructs to preserve when rebasing:
export PATH := $(CURDIR)/$(VIRTUAL_ENV_PATH):$(PATH), immediately after theNINJA :=line. Upstream has no venv-on-PATH line. Without it the lint tools resolve from the system PATH only and the gates silently self-skip.- The
define require-tool ... endefblock. Upstream has no equivalent. - The absence of
|| trueon everyformat-checkandlint-pystep. If a rebase reintroduces the upstream-eracommand -v X && X ... || trueidiom, the gate silently becomes incapable of failing again — this is exactly the regression this change fixed, and it is invisible because the target still prints "all lints passed".
Point 3 is the one to watch: a conflict resolved in upstream's favour restores a green-but-dead gate with no test failure to signal it.
fix/sycl-qsv-zerocopy-p010-normalize — SYCL QSV zero-copy P010 normalization + separate-session contract (2026-06-30)¶
no ffmpeg-patch impact: edits the fork-added SYCL zero-copy path only (core/src/sycl/dmabuf_import.cpp, core/src/sycl/common.cpp/.h, core/src/sycl/dispatch_strategy.cpp/.h). The new vmaf_sycl_import_debug_enabled() accessor and the va_import_path parameter on vmaf_sycl_select_strategy() live in the SYCL-internal core/src/sycl/ headers, not the public core/include/libvmaf/ surface, and no CLI flag / meson_options.txt / LIBVMAFContext field changed → nothing the ffmpeg-patches/ stack consumes. The separate--init_hw_device qsv=… requirement (FIX-03, ADR-1121) is an ffmpeg invocation pattern documented in docs/backends/sycl/overview.md, not a change to vf_libvmaf.c. Invariants (keep on any upstream sync — the whole SYCL VA-import path is fork-added, so there is no upstream-parity conflict surface): (1) the >> (16 − bpc) MSB→LSB shift (guarded if (bpc > 8)) must stay on every import path — fused into the Tile4 / Y-tiled de-tile store (each sample shifted as written), and via the standalone launch_p010_normalize() kernel on the LINEAR / readback fallbacks (event threaded into vmaf_sycl_set_detile_event() / subsumed by q->wait_and_throw()). Dropping it re-introduces the 64× integer_motion / NaN bug; do NOT re-add a standalone full-plane normalize pass on the tiled paths (it cost ~15% throughput at 4K — keep it fused). (2) Do NOT re-add a DMA_BUF_IOCTL_SYNC flush — it was tried and removed: insufficient for the contamination (the fix is the separate-session contract) and its SYNC_START blocking fence-wait serialised decode→compute. (3) the zero-copy path defaults to DIRECT dispatch — keep va_import_path threaded from state->has_imported into vmaf_sycl_select_strategy() (checked after the env overrides). The combined graph is a net throughput loss on VA-import (byte-identical output, ~15–25% slower at 4K); do NOT let the host-upload-tuned area-threshold re-select graph for it. (4) D-03 verification targets are the de-contaminated oracle (VMAF 97.2350, integer_motion max 26.6935), not the old shared-session 96.133894 baseline.
fix/adm-dwt2-neon-parity (2026-08-30)¶
core/src/feature/arm64/adm_neon.c is fork-added (upstream Netflix/vmaf has no aarch64 ADM DWT2 kernel), so there is no upstream counterpart to conflict with.
One invariant to preserve: the kernel's vertical pass vectorises 16 columns at a time while integer_adm.c dispatches it on !(w % 8). The scalar tail added here covers the gap. If the vector loop is ever widened, the tail's start index (w / 16) * 16 has to widen with it, or widths that are a multiple of the dispatch granularity but not the vector width will again read buf->tmp_ref entries that were never written in that iteration.
2026-08-31 — ADR-1127 single SemVer release stream¶
No upstream rebase impact: preserve VMAFx's one independent vX.Y.Z root release stream, coordinated version-file list, and draft-publication fan-out when importing upstream release metadata. Do not restore the historical -lusoris.N suffix or component release-please packages.
The release fan-out is fail-closed: every write/OIDC job follows the exact-tag version preflight; native and vmaf-mcp hashes use distinct SLSA jobs and asset names; Anchore's implicit uploads stay disabled so SBOMs pass through signing; and attachment refuses missing globs. Manual recovery may overwrite GitHub release assets, but an existing PyPI filename must have the exact reproducible SHA-256 before skip-existing is allowed.
The Linux native release stage must preserve every Meson shared-library chain name (libvmaf.so, its SONAME, and its real name) as identical regular-file assets. The pre-signing artifact round trip must keep proving that the exact downloaded CLI resolves the staged SONAME under env -i and reports the release version; filename-only checks do not protect runtime linkage.
2026-08-31 — ADR-1128 fragment-owned release cuts¶
No upstream code impact: preserve skip-changelog: true and the pre-merge fragment rollover when rebasing release automation. Re-enabling release-please's changelog updater without also replacing the fragment renderer would republish every consumed release entry under Unreleased.
2026-08-31 — ADR-1129 release container runtime alignment¶
The production Dockerfiles and their two publication workflows are fork-local; an upstream sync has no direct file counterpart to prefer. Preserve these coupled invariants:
- Debian 13-compiled CPU, Go-server, and node binaries must not return to a Debian 12 runtime.
Dockerfile.go-servermust build the fork's libvmaf, its CGO server, and the distroless runtime on that same ABI for both amd64 and arm64. The Python MCP server must keep a real Python 3.14 interpreter matching the venv builder. - CUDA 13.3.1, ROCm 7.2.4, and oneAPI 2025.3.1 builders come from their digest-pinned vendor devel images; their final stages use the matching vendor runtime/application families.
- Node FFmpeg dependencies are collected from the native
lddclosure. Do not restore/usr/lib/x86_64-linux-gnucopies; they break the arm64 publish leg. - Both Docker workflows trigger on
release.publishedand gate every image build behindvalidate-release. Release events and manual recovery must identify the same existing published ordinary tag through the input,GITHUB_REF,GITHUB_SHA, checkout, and coordinated version files; manual recovery runs with--ref vX.Y.Z -f tag=vX.Y.Z. Each build keeps cosign signing, CycloneDX SBOM handling, and GitHub-native provenance; smoke jobs verify signatures before pulling and contain no success-masking fallback. - The Python server uses MCP 2.x constructor handlers, not the removed decorator/request-context API. Its production image installs
[eval,http], keeps the documented HTTP CLI dispatch, and explicitly opts the container intoVMAFX_MCP_HTTP_BIND=0.0.0.0without weakening fail-closed auth. - Go server, operator, and node container builds inject the published tag into
github.com/VMAFx/vmafx/pkg/version.version; all three binaries handle--versionbefore starting their long-running fx graphs, and release smokes assert the exact tag. The Go-server lane must keep the exact amd64+arm64 manifest assertion and digest-addressed/healthzplus/readyzprobes. Do not restore the removedmain.buildVersionldflag target or success-masked--helpprobes. - The node's hand-written
libvmaf.pcmust expose both${includedir}and${includedir}/libvmaf: FFmpeg includes the legacy<libvmaf.h>spelling while fork patches also include namespaced<libvmaf/...>headers. - Ubuntu 24.04 vendor builders spell the C23 Meson option
c2x, and GPU custom targets require an in-source-treecore/builddirectory for their relative include paths. CUDA also installs the exact minimum compatiblenv-codec-headerscommit declared inDockerfile.production-gpu; distro headers are older than libvmaf's loader API. Preserve all three mechanics. - The operator's post-ADR-1119 runtime contract is env-only: the Helm template exports
VMAFX_OPERATOR_METRICS_ADDR=:8080,VMAFX_OPERATOR_HEALTH_PROBE_ADDR=:8081,VMAFX_OPERATOR_LEADER_ELECTION, andVMAFX_LOG_LEVEL. Its named ports and probes use8080/8081. Do not restore the ignored pre-fx CLI arguments or the stale8082probe. The Go server, operator, and node helpers default to the repositories and canonicalv<Chart.AppVersion>image tag published by the release workflows; explicit user-supplied image tags remain verbatim. - The node decorates golusoris's shared
grpc.Configonly whengrpc.listenis empty, preserving the historical standalone:50052default without clobbering an explicit:9090file or environment value. Its runtime image exportsVMAFX_MODEL_DIR, not the unusedVMAF_MODEL_PATH; keep that name coupled to the Go config contract. -
Helm validation keeps
HELM_VERSIONand the official archiveHELM_SHA256identical inhelm-chart.ymlande2e-k8s.yml. Download to a file, verify, then extract; never restore a moving remote installer piped into a shell. Keep kuttl's rawsteps.kuttl.outcomefinal assertion when retainingcontinue-on-errorfor diagnostic uploads. -
GHCR, GitHub Release, and Sigstore publication jobs use the protected
release-publishenvironment; PyPI usespypi-publish. Reusable SLSA jobs keepcontents: readplusupload-assets: false, and the protected attach job publishes both provenance artifacts. Container signature checks require the exact@refs/tags/${PUBLISH_TAG}certificate identity, never@.*.
2026-08-31 — oneAPI 2025.3.2 production builder patch¶
No upstream code impact: preserve the production -oneapi2025 container's explicit split patch levels until Intel publishes a matching Ubuntu 24.04 runtime tag. The builder uses digest-pinned oneAPI Base Toolkit 2025.3.2; the final stage uses Intel's latest published oneAPI Runtime 2025.3.1 image. Do not describe the final runtime as 2025.3.2, replace it with the development-heavy basekit, or assemble a hand-picked runtime-library subset. Build and execute the final-oneapi2025 entrypoint after either pin changes. Ubuntu 24.04's Meson 1.3.2 cannot configure this source tree; preserve the single checksum-pinned Meson 1.12.0 wheel stage and its exact-version assertions in the CUDA, ROCm, and oneAPI builders. Preserve cp -r model/. /dist/model/ in the standard production builder and all four GPU-Dockerfile builders: copying model/ itself creates a second model/ directory and breaks the documented VMAF_MODEL_PATH layout.
refactor/c-rework-core — library-core plumbing split into helpers (2026-09-02)¶
Upstream-mirror files reworked under ADR-0141: core/src/libvmaf.c, core/src/predict.c, core/src/feature/feature_collector.c (all keep the Netflix header) plus the fork's C++ twin core/src/read_json_model.cpp. The fuzz-only C twin core/src/read_json_model.c was deliberately not touched. Rebase-sensitive points:
vmaf_init()is nowvmaf_ctx_subsystems_init()+vmaf_ctx_thread_pools_init(). An upstream hunk that adds a subsystem tovmaf_initbelongs invmaf_ctx_subsystems_initwith a matching label in its reverse-order teardown chain;vmaf_inititself only mallocs, seeds the CPU/log state, and frees on failure.*vmafis assigned on success only.vmaf_read_pictures()carries a single#ifdef HAVE_CUDA(around theread_pictures_frame_translatecall, which exists only in CUDA builds) instead of the former six islands. The caller's pictures and their CUDA host/device translations live inReadPicturesFrame;read_pictures_frame_translate/_select_host/_cleanup/_cleanup_after_batchhold the backend branches. Upstream changes to the picture-ownership rules (which unref runs on which path — the PR #838 double-unref regression) must be merged into those four helpers, not re-inlined.- Three
cppcheck-suppress constParameterPointermarkers carry inline citations (vmaf_context_get_backend: frozen public prototype;read_pictures_validate_and_prep: SYCL upload takes mutable pictures;vmaf_feature_collector_unmount_model: prototype shared with the C++ twinfeature_collector.cpp). An upstream hunk touching those signatures must keep the marker on the line directly above the definition.vmaf_feature_collector_get()inlibvmaf_priv.h/libvmaf.cnow takesconst VmafContext *. threaded_extract_batch_func()is split intobatch_thread_data_ensure,batch_extractor_skip,batch_ensure_fex_ctx,batch_extract_one. The ADR-0795 per-thread deep-copy assertion and the ADR-1051 PREV_REF balance (onevmaf_picture_refper PREV_REF extractor, released byfex_release_prev_ref, snapshot released once at the end) are insidebatch_extract_one— keep them there on conflict.batch_extractor_skipandread_pictures_should_skipmust stay in sync (they sharefex_subsample_skip).- The six-site
if (fex->prev_ref.ref) { unref; memset }idiom isfex_release_prev_ref(); an upstream change to the PREV_REF swap infeature_extractor.cppneeds exactly one review of that helper. set_fex_{cuda,sycl}_state/set_fex_framesyncarevoidand are called throughfex_ctx_bind_backends()from bothvmaf_use_featureandvmaf_use_features_from_model; they never failed upstream either, theerr |=chain was dead.vmaf_write_output_with_format()delegates tooutput_file_open(0644 + errno capture, ADR-0602),output_fps(ADR-0606) andoutput_write; theferror/-EIOtail contract from ADR-0119 lives inoutput.cand is untouched.feature_collector.c: the ADR-0154-EAGAIN("written yet?") contract moved verbatim intofeature_vector_read(); the public lock/destroyedhandshake in every entry point is unchanged.predict.c'spredict_load_feature_scoreEAGAIN/EINVAL split (AGENTS.md invariant) is untouched.read_json_model.cpp: helpers are in two anonymous-namespace blocks that bracket the fourextern "C"entry points;vmaf_read_json_modelruns the parser throughmodel_parse_c_locale()(ADR-0137 bracket). A future twin-drift gate comparing it withread_json_model.cmust tolerate the namespace andconstdifferences.- All three C files carry a file-scoped
NOLINTBEGIN/END(modernize-use-nullptr)bracket per ADR-1138 — keep the closing marker at EOF when appending.
refactor/c-rework-vif-motion — scalar VIF and float-motion split into helpers (2026-09-02)¶
Upstream-mirror files reworked under ADR-0141: core/src/feature/integer_vif.c, core/src/feature/vif_tools.c, core/src/feature/float_motion.c (all keep the Netflix header) plus a const on vif_compute_line_residuals's state parameter in core/src/feature/integer_vif.h. Every numeric path is byte-identical to the pre-rework binary (31-case --precision max matrix across the scalar, AVX2 and AVX-512 lanes; see docs/research/2026-09-02-c-rework-vif-motion-bit-exact.md). Rebase-sensitive points:
- The integer VIF statistic lives once.
vif_statistic_8,vif_statistic_16andvif_compute_line_residuals(the tail helpervif_avx2.c/vif_avx512.c/vif_neon.ccall for the columns their block width does not cover) all runvif_horizontal_pixel→vif_accumulate_pixel→vif_store_residuals. An upstream hunk that changes the horizontal pass, thesigma*/g/sv_sq/numer1arithmetic or the finalnum/denformula must be applied to those helpers once, verbatim (same operand types and order), and then mirrored in the three SIMD kernels — do not re-inline a per-function copy. The 8-bit and 16-bit vertical passes arevif_vertical_line_8/vif_vertical_line_16; the 16-bit rounding / shift constants arevif_shift_for_scale(member names unchanged:add_shift_round_VP,shift_VP,add_shift_round_VP_sq,shift_VP_sq). vif_compute_line_residualstakesconst VifPublicState *; upstream's non-const prototype converts implicitly at every SIMD call site.log_generateusesroundf(proven bit-identical toroundover all 32768 entries); do not "restore"roundfor parity — the LUT is the same.init(integer VIF) isvif_init_dispatch+vif_buffers_alloc; the byte-cursor layout of the single allocation (MSVC C2036 workaround) is insidevif_buffers_allocand its offsets are unchanged. There is nofail:label: the dictionary-failure path freesbuf.dataand NULLs it.write_scoresiswrite_scale_scores+write_debug_scoresover thevif_scale_{score,num,den}_names[4]tables; the append order (four scale scores,integer_vif/_num/_den, then num / den per scale) is unchanged and thedoublesums stay explicit left-to-right expressions.decimate_and_padindexes with(ptrdiff_t)i * 2/(ptrdiff_t)j * 2(src_row/src_col) instead of the implicitly widenedunsignedproducts.vif_tools.c: the AVX2-only dispatch (ADR-0504 rationale comment) isvif_use_avx2_convolution; the reflect-101 index isvif_mirror_index; the three publicvif_filter1d_*_sfallbacks arevif_filter1d_vertical_s/_vertical_sq_s/_vertical_xy_s+ the sharedvif_filter1d_horizontal_s. The per-pixel float statistic isvif_pixel_statistic_s; the upstreammatching_creference block is a file-scope comment above it.ceil/flooronfloatoperands areceilf/floorf. A scratch-rowaligned_mallocfailure now logs and returns instead of dereferencing NULL.float_motion.c:MotionStateholdsMotionPlane plane[3](Y, U, V —ref,tmp,blur[3]) instead of the flatref/ref_u/ref_v/tmp*/blur*fields; U and V are allocated only withmotion_add_uv, andmotion_free_planesis the only teardown (initfailure andclose).motion_chroma_heightsruns before any allocation and has adefault:(upstream leaked the Y buffers on the YUV400P-EINVAL). Blur / copy / score aremotion_copy_and_blur→motion_blur_planeandmotion_score_pair(Y, then U, then V — thedoubleadd order is load-bearing). Score clips aremotion_clip/motion_blend_clip; every collector append goes throughmotion_append. The threereadability-function-sizeNOLINTs from the b949cebf port are gone; do not bring them back with an upstream hunk.integer_vif.candfloat_motion.ccarry a file-scopedNOLINTBEGIN/END(modernize-use-nullptr)bracket per ADR-1138 — keep the closing marker at EOF when appending.vif_tools.chas no null-pointer constants and therefore no bracket. The registry symbols keep the citedNOLINTNEXTLINE(misc-use-internal-linkage).flush()carries a citedcppcheck-suppress constParameterCallback.
fix/release-please-setup — 1.0.0 release line + pipeline repair (2026-09-03)¶
core/meson.build— rebase-sensitive, one comment block. Five comment lines were added directly abovevmaf_soname_version = '3.0.0'recording that the ABI SONAME is deliberately independent of the release-please-owned product version on line 2, is hand-bumped only on an ABI break, and is NOT reset by the fork's 1.0.0 first release (ADR-1151). Upstream has no such comment, so a sync that rewrites the region aroundvmaf_soname_versionwill conflict here. Keep the comment; the value itself ('3.0.0') is upstream's and should follow upstream on a bump.core/meson.buildline 2 —version : '3.2.1', # x-release-please-versionis a coordinated release marker owned by release-please, not by upstream. Whatever an upstream sync brings, the fork's value wins and the trailing# x-release-please-versionmarker must survive verbatim: it is the anchor the generic updater rewrites, andscripts/release/verify-release-version.shrequires exactly one such marker per listed file.- Everything else in this change is fork-only release tooling with no upstream counterpart:
release-please-config.json,.release-please-manifest.json,.github/workflows/{release-please,supply-chain,docker-publish-*, required-aggregator,rule-enforcement}.yml,.github/ci-impact.json,scripts/release/,docker/Dockerfile.node,deploy/helm/,bindings/rust/*/Cargo.toml,core/src/feature/rust/tad/Cargo.toml,pkg/version/version.go,scripts/ci/release-pr-exempt.sh,scripts/ci/tests/test-release-pr-exempt.sh, and the docs. No rebase impact.
docs/ai-quantization-wire-format — int8 wire-format reconciliation (2026-09-03)¶
- No rebase impact: docs-only.
docs/ai/quantization.md,docs/state.mdandchangelog.d/are fork-local files with no upstream Netflix counterpart, so an upstream sync cannot conflict here. - Invariant the new text depends on (worth re-checking after any DNN sync, though the whole
core/src/dnn/tree is fork-local today): the documented behaviour is pinned to two code facts — the quantisation entries incore/src/dnn/op_allowlist.c(QuantizeLinear,DequantizeLinear,DynamicQuantizeLinear,MatMulInteger,ConvInteger, and deliberately noQLinear*), and the fp32 fallback branch invmaf_dnn_session_open(core/src/dnn/dnn_api.c) that logs atVMAF_LOG_LEVEL_DEBUGand keepsload_pathon the fp32 baseline. If either changes, the "What the fork loads" and "Loader behaviour and fp32 fallback" sections go stale.
fix/publishing-container-enforcement — make container-only publishing enforceable (2026-09-03)¶
dev/Containerfile— the only rebase-sensitive file in this change, and only mildly so. ARUN printf ... > /etc/vmafx-dev-containerlayer is inserted in thebuild-depsstage, between theLABELblock andARG DEBIAN_FRONTEND. Upstream Netflix/vmaf has nodev/Containerfile, so there is no upstream counterpart and no sync conflict. What matters on a fork-local rewrite of the stage: the marker must stay in the first stage, because every downstream stage (gpu-sdks,libvmaf-build,go-build,dev-mcp) inherits it from there, andscripts/ci/check-container-build.shreads it in all of them. The four key names (vmafx_dev_container,image_title,containerfile,source) are a contract with the gate script and withscripts/ci/tests/test-check-container-build.sh, which parses the marker lines back out ofdev/Containerfileso the fixture cannot drift from the image.scripts/ci/check-container-build.sh,scripts/ci/tests/test-check-container-build.sh,.github/workflows/dev-container-build.yml,.github/workflows/rule-enforcement.yml,docs/development/publishing.md,docs/state.md,changelog.d/— fork-only CI, policy and documentation surfaces with no upstream counterpart. No rebase impact.
Upstream-issue harvest 2026-09-03 (ADR-1166, branch fix/upstream-harvest-2026-09-03)¶
Nine stale Netflix/vmaf reports were verified against this tree and the confirmed subset fixed. The entries below are the ones a future /sync-upstream needs, because each touches an upstream-mirrored file where the two trees now diverge. Full triage table, including the ALREADY-FIXED and NOT-APPLICABLE verdicts, in docs/research/1166-upstream-issue-harvest-2026-09-03.md.
core/src/feature/common/convolution_internal.h — Netflix/vmaf#1582 / #1581¶
The three edge helpers no longer open-code the single-bounce reflect-101 fold. There is one convolution_reflect101(idx, size) FORCE_INLINE helper at the top of the header, and convolution_edge_s / _sq_s / _xy_s each call it once per tap. The fold is iterative (while (idx < 0 || idx >= size)) with a size <= 1 short circuit, because a single bounce only lands in range for size >= radius + 1; below that it falls out the opposite side and the caller dereferences out of bounds.
For every size >= radius + 1 the loop exits on the first iteration and yields the identical index, so this is a pure safety change with no score movement — pinned by core/test/test_convolution_edge_small.c::test_large_plane_bit_identical, which compares a 24x24 run against an explicit single-bounce reference and asserts bit equality.
An upstream hunk that re-introduces the open-coded width - (j_tap - width + 2) form at any of the three sites must be dropped, not merged. Upstream's own #1582 patch introduces a convolution_mirror() helper of the same shape; prefer keeping the fork's name and the header comment that cites both issue numbers.
core/src/feature/common/convolution.c — Netflix/vmaf#1582¶
convolution_x_c_s and convolution_y_c_s now call convolution_clamp_borders(dim, &borders_lo, &borders_hi) immediately after deriving the two bounds. Upstream leaves borders_right / borders_bottom negative for a plane narrower/shorter than the filter, which makes the trailing loop start at a negative index and write dst[i * dst_stride - 1] / dst[-dst_stride + j] — a heap underflow write. The clamp is a no-op for every dim >= filter_width, so no in-contract behaviour changes; it also removes the duplicate recompute when the two border bands would otherwise overlap.
The file also now #include "alignment.h" instead of re-declaring vmaf_floorn / vmaf_ceiln as local externs. core/src/feature/common/convolution.h gained prototypes for convolution_x_c_s / convolution_y_c_s, which already had external linkage; this silences -Wmissing-prototypes and lets the regression test drive the scalar passes without going through the SIMD dispatch.
core/src/feature/integer_motion.c, integer_motion_v2.c, x86/motion_avx2.c, x86/motion_avx512.c, arm64/motion_v2_neon.c — deliberate divergence from Netflix/vmaf#1581¶
These files keep their single-bounce mirror() bodies on purpose. They sit downstream of an init() guard that has rejected w < 3 || h < 3 since Research-0094, so the defective sizes never reach them. Upstream #1581 goes the other way — it fixes mirror() so tiny frames can be scored; the fork errors out instead. A sync that pulls upstream's mirror() change here is a behaviour decision, not a mechanical merge: it would make the guards unnecessary and start producing scores for 1x1 and 2x2 frames, which the fork has deliberately refused since Research-0094.
The same applies to the CUDA / HIP / Metal mirror twins.
core/src/feature/float_vif.c — Netflix/vmaf#1582¶
The min-dimension guard is no longer a hard-coded 9. It is vif_get_min_dim((float)s->vif_kernelscale) — the largest ((filter_width_s / 2) + 1) << s over the four-scale ladder, which is 16 at the default kernelscale. The old floor covered scale 0 only, so 9..15 px input reached the scale-3 convolution with a sub-minimum plane. vif_get_min_dim is new in core/src/feature/vif_tools.{c,h}; upstream has no counterpart, so an upstream hunk that touches the guard will conflict.
core/src/feature/float_motion.c — Netflix/vmaf#1582 / #1581¶
motion_check_min_dim gained a const char *plane argument (for the log message) and is now driven by motion_check_min_dim_all_planes, which also validates the chroma dimensions when motion_add_uv is set, deriving them with picture.c's own (dim + ss) >> ss geometry via the new motion_chroma_shifts helper. Upstream validates nothing here; the fork's own prior guard validated luma only, which is what left the live out-of-bounds read on the chroma blur.
core/src/model.c (and the unbuilt core/src/model.cpp twin) — Netflix/vmaf#1242¶
vmaf_model_feature_overload no longer has an exit: label: the -ENOMEM and dictionary-free failure paths break out of the loop and fall through to the single unconditional vmaf_dictionary_free(&opts_dict). That is exactly the shape the unbuilt C++ twin already had, so the two files are now convergent — keep them that way (T-TWIN-DEAD-SIDES-2026-09-02 tracks the twin's build wiring). vmaf_model_collection_feature_overload gained argument guards (!model || !feature_name || !opts_dict, plus !*model_collection) and now propagates and cleans up after a failed vmaf_dictionary_copy.
Do not adopt upstream's proposed VmafFeatureDictionary ** signature change: it is an API/ABI break that would need its own ADR and soname handling.
core/include/libvmaf/feature.h, model.h, libvmaf.h — Netflix/vmaf#1242¶
The VmafFeatureDictionary ownership contract is now written identically in all three headers: consumed on every path except the argument-validation guards, where the caller still owns it. feature.h and model.h previously documented opposite rules. These are fork-authored doc comments (upstream's headers are much sparser), so an upstream sync will not conflict, but any edit must keep the three copies in step.
core/src/feature/compat_builtin.h — Netflix/vmaf#1551, retracting Netflix/vmaf#1422¶
This file is fork-added (there is no upstream counterpart), but round-21 item (n) recorded the __lzcnt choice as settled, and it is not: __lzcnt emits LZCNT unconditionally, which silently decodes as BSR on any x86-64 without ABM/LZCNT and returns the MSB index instead of the leading-zero count. The shim now uses _BitScanReverse / _BitScanReverse64 and carries a _M_X64 || _M_IX86 architecture guard.
Do not adopt Netflix/vmaf#1422's __lzcnt form — upstream's own #1551 retracts it. scripts/ci/check-msvc-clz-shim.sh enforces this and fails the fast suite if the intrinsic returns anywhere under core/src.
core/tools/spinner.h and core/tools/vmaf.cpp — Netflix/vmaf#743¶
spinner.h is upstream-mirrored and upstream still has the bug open. The braille table itself is byte-for-byte unchanged (56 entries, verified in core/test/test_spinner.cpp); what is new is the spinner_ascii fallback table, spinner_table_for_codepage(), spinner_erase_eol() and the SPINNER_CODEPAGE_UTF8 constant, and the array is now static const char *const. vmaf.cpp gained an #ifdef _WIN32 WindowsConsoleGuard RAII class plus console_output_code_page() / console_vt_enabled() / console_progress_style() / emit_progress_line() in an anonymous namespace; the progress fprintf moved into emit_progress_line(). On POSIX every selector returns the pre-existing value, so the emitted bytes are unchanged.
core/src/meson.build, core/tools/test/meson.build — Netflix/vmaf#1573¶
The nvcc fatbin include list is now built from absolute meson.current_source_dir() / meson.current_build_dir() paths (cuda_inc_flags), matching what the SYCL block below it already did. The relative form only resolved when the build directory was a direct child of core/, which stopped being the documented layout at ADR-0700. libvmaf_private_libs gained the C++ runtime for Netflix/vmaf#1178, detected via _LIBCPP_VERSION rather than the compiler id. The three shell-driven tool tests now declare depends and workdir.
core/src/feature/psnr_tools.cpp, psnr.c, integer_psnr.c, float_psnr.c — ADR-1142 tidy ratchet¶
These four files carry the Netflix upstream copyright header and still track upstream's PSNR implementation. The wave-2 lint pass on the psnr bucket touches them in ways a future port-upstream-commit has to be aware of:
psnr_tools.cpp— thekFormatTablerows are now written with designated initialisers ({.fmt = "yuv420p", .params = {.peak = 255.0, .psnr_max = 60.0}}). The table already diverged from upstream'sstrcmpladder at ADR-0731; this is a syntax-only change on top of that divergence. Peak / psnr_max values are byte-identical to upstream's.psnr.c— now includes its ownpsnr.h. Upstream does not; the include is what tells clang-tidy thatcompute_psnr()legitimately has external linkage. Keep it when replaying an upstream hunk that rewrites the include block.integer_psnr.c/float_psnr.c— both keep upstream'sNULLspelling (ADR-1138: C translation units never use the C23nullptrkeyword, because the requiredBuild — Windows MSVC + CUDAlane compiles them with cl.exe and MSVC's documented/std:clatestfeature set does not include it). Each file therefore carries a file-scoped/* NOLINTBEGIN(modernize-use-nullptr) ... ADR-1138. */…NOLINTENDbracket instead, exactly likecore/src/feature/integer_adm.c. An upstream hunk that adds a pointer initialiser inside the bracket needs no adaptation; a hunk that lands outside it (before theNOLINTBEGINor after theNOLINTEND) does — keep the bracket spanning the whole file. TheVmafFeatureExtractordefinitions additionally carry the ADR-0278NOLINTNEXTLINE(misc-use-internal-linkage)citation used by every other extractor in the fork; an upstream hunk that rewrites those definitions must keep the citation line.
core/src/dnn/*.c are fork-local (no upstream counterpart). They carry the same ADR-1138 bracket for the same MSVC reason, and model_loader.c's function split in this change carries no rebase risk.
feat/gpu-adm-csf-mode-parity — GPU integer-ADM option-table parity (2026-09-05)¶
Rebase-sensitive files: core/src/feature/cuda/integer_adm_cuda.{c,h}, core/src/feature/sycl/integer_adm_sycl.cpp, core/src/feature/hip/integer_adm_hip.{c,h}, core/test/test_{cuda,sycl,hip}_adm_parity.c.
All five sources are fork-local — upstream Netflix/vmaf has no GPU ADM twin beyond CUDA, and even the CUDA one diverged long ago (ADR-0746 added the AIM device pass, ADR-0487 the adm_min_val option). An upstream rebase that touches core/src/feature/integer_adm.c (the reference) can still invalidate this work, because these three twins are now defined as mirrors of it:
-
The option table is a mirror, and the mirror is load-bearing.
vmaf_feature_name_from_options()builds the emitted feature key from the extractor's ownoptions[]. If an upstream sync adds, renames or re-aliases an entry ininteger_adm.c's table, the same edit must land in all three twin tables in the same commit or the twins start emitting a different key than the CPU for the same opts dict and every model lookup that names that feature misses — silently.core/test/test_{cuda,sycl,hip}_adm_parity.ceach carry a..._option_table_mirrors_cputest that walks the CPU table and fails on the first name / alias / type / feature-param-flag mismatch; that test is the tripwire. -
adm_csf_factors()andadm_csf_rfactor_scale0()are copied, not shared. Each twin has a private copy of the two helpers frominteger_adm.c(a shared header would have to be includable from.cppunder icpx and from.cunder nvcc/hipcc; the ADM enum already exists in two conflicting forms —adm_options.hhas{WATSON97, BARTEN, ADM}whileinteger_adm.hhas{WATSON97, BARTEN, BARTEN_WATSON_BLEND, BARTEN_WATSON_BLEND_MAE}). An upstream change to the CSF weights, to the{36453, 36453, 49417}scale-0 constants, or to thenvd * rdhcanonical test must be replicated into all three copies. -
AdmFixedParametersCuda/AdmFixedParametersHipheader dependency tracking (ADR-1320, Research-2106). Historically,core/src/meson.builddeclared device targets with onlyinput : _cuand no header dependencies, meaning editing structs in headers likecore/src/feature/cuda/integer_adm_cuda.hpaired new host layouts with stale device layouts without triggering fatbin/HSACO rebuilds. Filed asT-CUDA-FATBIN-NO-HEADER-DEP-2026-09-05indocs/state.mdand resolved by ADR-1320.core/src/meson.buildnow binds explicitdepend_fileslists (cuda_kernel_shared_headers,hip_kernel_shared_headers) covering the complete repo-local quoted include closure across all 22 CUDA fatbin targets and 22 HIP HSACO targets (plus generatedconfig_h_targetfor CUDA), alongside compiler depfiles (-MD -MF @DEPFILE@on POSIX nvcc and-Xclang -dependency-file -Xclang @DEPFILE@on hipcc; depfile omitted on Windows MSVC). Header edits now reliably trigger incremental Ninja rebuilds without manualtouchworkarounds. -
adm_min_valfloorsadm3only.integer_adm.c::extract()wraps only the adm3 expression inMAX(..., s->adm_min_val);adm2is emitted raw. The Netflix goldenadm_min_val=0.98case pinsVMAF_integer_feature_adm2_min_0.98_scoreat0.9345148541666667, below the floor — that assertion is the contract. All three twins used to clampadm2; they no longer do. -
SYCL and HIP do not provide
aim_score/adm3_score. They have no AIM device pass, so both features are left out ofprovided_features[]and the ADR-0530 name-based fallback routes them to the CPU twin. Do not "fix" a rebase conflict by re-adding them to the array unless the AIM kernels land with it — an earlier draft of this branch emitted them from a hard-codedaim_num = 0.0, which is a fabricated score, not a fallback.
fix/gpu-threads-ctx-sync — threaded flush leaves GPU extractors alone (2026-09-06)¶
Touches core/src/libvmaf.c, which is upstream-mirrored, so a rebase can plausibly reintroduce this. Two invariants:
-
flush_context_threaded()'s first loop must skipVMAF_FEATURE_EXTRACTOR_CUDAandVMAF_FEATURE_EXTRACTOR_SYCL. Its second loop already skipped CUDA; the first did not, and that asymmetry madevmaf --threads Nfail on every GPU backend for everyN. Restoring the plainTEMPORAL-only condition brings the bug straight back. The backend flush paths own GPU extractors in both threaded and serial mode, so they run collect-then-flush in the one order that yields correctmotion2/motion3at a batch boundary. See ADR-1197. -
Do not re-merge the extractor error and the CUDA driver error in
flush_context_cuda(). They are deliberately separate variables (extractor_err,cuda_err). Folding them back into oneerris what made an extractor's-EINVALannounce itself as "context could not be synchronized" while all four driver calls were returning success — the single most misleading symptom in this bug, and the reason it went unfixed.
The guard that used to sit in flush_context_cuda() (if (vmaf->thread_pool && TEMPORAL) continue;) is intentionally deleted, not moved. A rebase that resurrects it alongside invariant 1 will skip the flush entirely for temporal GPU extractors.
RN-2026-09-06 — Netflix benchmark harness paths and flags are host-coupled¶
testdata/benchmark_netflix.py and testdata/bench_all.sh are fork-added and have no upstream counterpart, so a rebase never conflicts them — but three values inside them silently rot and are worth re-checking after any sync:
bench_all.shmust not pass a flag the CLI has removed. It carried--no_vulkanin all three backend flag sets long after ADR-0726 deleted the Vulkan backend; current builds printunrecognized option '--no_vulkan'and the run continues, so the staleness is invisible until something else fails. If a future sync removes another negative selector (--no_cuda,--no_sycl), update the flag sets in the same change.bench_all.shhard-codes--threads 1. That is not cosmetic: oncd52f2670every GPU backend aborts withproblem flushing contextwhen a thread pool is present (T-GPU-CLI-THREADS-CTX-SYNC-2026-09-06). Do not "fix" a red bench row by dropping the flag — that hides the defect the row is now correctly reporting.- The VA-API render node is not stable.
benchmark_netflix.pyused to pin/dev/dri/renderD130for the SYCL/QSV import; on the bench host that is now the AMD iGPU and the Arc A380 isrenderD129. The node is an environment override (VMAF_SYCL_RENDER_NODE), never a literal.
testdata/netflix_benchmark_results.json is deliberately stale as of 2026-09-06 — see ADR-1192. Do not regenerate it as part of a rebase.
ci/container-source-guard — record the container's source revision (2026-09-06)¶
Fork-only tooling (dev/, scripts/dev/, scripts/ci/tests/). One invariant:
/etc/vmafx-dev-sourcemust stay in the LAST stage ofdev/Containerfile. It sits besideENV PATH=/opt/vmaf-venv/bin:...indev-mcp, deliberately far from the ADR-1102/etc/vmafx-dev-containermarker written in the first stage. Moving it up to keep the two markers together looks tidy and breaks it: the first stage is reused by every rebuild, so the file would record the revision of whichever build first populated the layer cache. A marker that reports a stale revision authoritatively is worse than no marker. See ADR-1195.
fix/t-upstream-1109-psnr-cap-truncates — PSNR uncapped option (2026-09-06)¶
Rebase-sensitive files: core/src/feature/integer_psnr.c, core/src/feature/float_psnr.c, core/src/feature/psnr.{c,h}, core/src/feature/{cuda,sycl,hip,metal}/{integer,float}_psnr_*, core/src/feature/metal/{integer,float}_psnr.metal, core/test/test_psnr_uncapped.c.
integer_psnr.c, float_psnr.c and psnr.{c,h} are upstream-mirror files; the eight GPU twins are fork-local. ADR-1193 changed the same expression in all of them, so an upstream sync that touches PSNR needs the following invariants held:
-
psnr_maxhas exactly two roles and they are now separate. Role (a), themse == 0infinity sentinel, is unconditional. Role (b), the truncation of computed values, applies only whenuncapped == false. Upstream's expression conflates them (MIN(10*log10(peak^2 / MAX(mse, 1e-16)), psnr_max)), so a verbatim upstream hunk landing oninteger_psnr.c::psnr_from_mse(),float_psnr.c::extract()orpsnr.c::compute_psnr()silently reintroduces the bug. Resolve such a conflict by keeping the fork's three-arm form and folding any upstream numeric change into both the!uncappedarm and theuncappedcomputed arm. The!uncappedarm is deliberately upstream's expression character-for-character, including theMAX(mse, 1e-16)floor: with amin_ssebelow ~1.9e-11 the ceiling rises past the ~208 dB a floored zero MSE produces, so a re-derivedmse == 0 -> psnr_maxdefault would not be bit-identical there. Do not "simplify" the two computed arms into one. -
The default must stay bit-identical.
core/test/test_psnr_uncapped.ccarries no-change guards (test_psnr_default_still_truncates,test_float_psnr_default_still_truncates) next to the fix assertions. Both directions have to keep passing; a rebase that moves the default 60 dB value is wrong even if the uncapped value is right. -
The option name, type and default are mirrored across ten extractors.
uncapped/VMAF_OPT_TYPE_BOOL/falseappears ininteger_psnr.c,float_psnr.cand each of the eight GPU twins. Adding it to one backend only produces a cross-backend divergence that no CPU test catches. The option is deliberately notVMAF_OPT_FLAG_FEATURE_PARAM: setting it must not renamepsnr_y/float_psnr, because the CPU extractor appends without a name dict while the GPU twins append with one — flagging it would make the two backends emit different keys for the same request. -
compute_psnr()inpsnr.chas no in-tree caller. It is part of the upstream float "tools" layer and is kept in sync deliberately. Its signature grew a trailingbool uncapped; an upstream rebase that reintroduces the three-argument form will compile (nothing calls it), so the mismatch has to be caught by review rather than by the build. -
Metal was not executed. The two
.mmtwins and their.metalcomment blocks were changed by inspection only — no Apple GPU is available on the fork's dev hardware. Treat the Metal hunks as unverified against silicon when reconciling them.
fix/t-upstream-930-adm-angle-flag — one angle_flag predicate for every backend (2026-09-06)¶
Branch: fix/t-upstream-930-adm-angle-flag-predicate-. ADR: ADR-1194. Digest: docs/research/2030-adm-angle-flag-fp64-free.md.
core/src/feature/adm_angle_flag.h (new) — the frozen predicate, once¶
Upstream spells the 1-degree angle_flag test inline in every ADM implementation, and the spellings had drifted apart (T-UPSTREAM-930). The expression now lives in one fork-added header with two entry points:
adm_angle_flag_fp64()holds the upstream expression verbatim. It is golden-frozen (CLAUDE.md rule 1): if an upstream rebase changes the expression, change it here and nowhere else. Do not "simplify" the(float)x / 4096.0narrowing away — the lossy narrowing is the contract.adm_angle_flag_i64()is fork-local: a bit-identical evaluation in 64-bit integers for backends that cannot execute binary64. It hard-codes the significand ofcos(1deg)^2asADM_ANGLE_FLAG_MC/ADM_ANGLE_FLAG_D; if the constant ever moves, both must move with it, andcore/test/test_adm_angle_flag.cfails loudly if they do not.
core/src/feature/integer_adm.c keeps its adm_angle_flag() wrapper so the two call sites read as before; the wrapper is a one-line forward. A rebase conflict inside that wrapper should be resolved toward the header, not by re-inlining the expression.
core/src/feature/cuda/integer_adm/adm_decouple_inline.cuh, hip/integer_adm/adm_decouple_inline.hip¶
Both decouple_angle_flag_s0 and decouple_angle_flag_s123 now forward to adm_angle_flag_fp64(). Upstream's s0 compares the exact int64 products — that is a more accurate angle test than the CPU's, and therefore the wrong one. If an upstream cherry-pick reintroduces the exact-product form, keep the fork's forwarding call; test_adm_angle_flag documents which quadruples the two forms disagree on. The .hip file remains a byte-for-byte port of the .cuh for this helper: edit both.
core/src/feature/sycl/integer_adm_sycl.cpp¶
Both angle-flag sites call adm_angle_flag_i64(). The #pragma clang fp contract(off) blocks that used to guard the float form are gone with it — there is no floating-point arithmetic left to contract. The fp64-free property of this translation unit is load-bearing: one binary64 instruction anywhere in it makes the SYCL runtime reject the whole SPIR-V module on non-fp64 devices (Arc A-series, most iGPUs), so never resolve a conflict here toward adm_angle_flag_fp64().
core/src/feature/metal/integer_adm.metal¶
iadm_angle_flag() is a hand-written MSL mirror of adm_angle_flag_i64() (MSL cannot #include the C header, and has no double type). The C header is the source of truth — any edit to adm_angle_flag_i64() must be copied across in the same commit. IADM_COS_1DEG_SQ is gone; the MSL side now needs only IADM_AF_D.
cmd/vmafx-mcp/, mcp-server/vmaf-mcp/, pkg/libvmaf/paths.go — MCP sidecar + gRPC bridge (#1240)¶
All fork-added; no upstream counterpart, so an upstream sync cannot conflict here. Two invariants a rebase must not quietly break:
- The sidecar argv builders are twins.
buildPerShotArgv/buildRoiArgv/buildBenchArgv/buildVplArgv(cmd/vmafx-mcp/impl_sidecar.go) and_build_per_shot_argv/_build_roi_argv/_build_bench_argv/_build_vpl_argv(mcp-server/vmaf-mcp/src/vmaf_mcp/server.py) must emit the same bytes;cmd/vmafx-mcp/sidecar_parity_test.goruns both and compares. Resolving a conflict on one side without the other silently breaks the gate. The same applies to float formatting: Go'sstrconv.FormatFloat(v, 'f', -1, 64)is mirrored byserver.py::_fmt_float, never byrepr. - The five gRPC bridge tools are Go-only on purpose (ADR-1184). A future "restore parity" sweep must not add them to the Python server or delete them from Go.
The argument bounds in both servers are copied from the C parsers in core/tools/vmaf_per_shot.c, vmaf_roi.c, vmaf_bench.c and vmaf_vpl.c. If a sidecar's CLI grammar changes, the two MCP servers and docs/mcp/tools.md change with it — the bounds are duplicated by design (the MCP layer rejects early so the caller gets a structured error), so the duplication has to be maintained.
docs/ai/retrain-runbook-1246.md¶
no rebase impact: fork-added operator runbook for the one-shot v1.0.16 teacher retrain (epic #1246); no upstream-mirrored files touched.
docs/ms-ssim-gpu-chroma-accuracy — per-backend option-table reality (2026-09-06)¶
Documentation only; no code is touched. One invariant for whoever closes T-MS-SSIM-GPU-CHROMA-OPTION-DRIFT-2026-09-06:
- HIP's
enable_chromais a dead branch, not a working option.init_fex_hipincore/src/feature/hip/integer_ms_ssim_hip.cassignsn_planes = 1uon both arms of itsif (pix_fmt == VMAF_PIX_FMT_YUV400P || !s->enable_chroma). A rebase that "tidies" that into a single assignment loses the marker for the unimplemented path, and one that assumes the else-arm already computes chroma will ship luma-only numbers under a chroma-enabled model. The safety net today isprovided_features, which deliberately lists onlyfloat_ms_ssim, so_cb/_crroute to the CPU twin (ADR-0530). Do not add those names to the array without implementing the planes.
perf/backend-baselines-1245 — per-backend baseline harness (2026-09-06)¶
Fork-local only; nothing here touches an upstream-mirrored file, so a Netflix rebase cannot conflict with the harness or the docs page. Two invariants are worth carrying forward anyway:
-
testdata/bench_all.sh's stdout shape is a consumed interface, not a convenience format. The MCPrun_benchmarktool (ADR-0517) andmake benchboth parse it. The newtestdata/bench_backends.pywas added beside it rather than folded into it for exactly that reason. If a future change wants repetition insidebench_all.sh, the MCP wrapper and its schema tests have to move in the same PR. -
The benchmark fixture directories are gitignored, so a worktree does not have them.
python/test/resource/yuv/(.gitignoreline 199) andtestdata/bbb/*.yuv(line 51) exist only in a full checkout. A benchmark run from a freshgit worktreefails withcould not open file: …that looks like a build problem and is not. Link the directory in before running; never "fix" it by un-ignoring the YUVs — they are hundreds of MB and the 4K pair is ~2.5 GB per file.
fix/changelog-unknown-section-gate — unknown fragment dirs fail (2026-09-06)¶
Fork-only release tooling. One invariant:
warn_unknown_subdirs()must return non-zero and its caller must propagate it. The function is named "warn" for history; since ADR-1198 it is an error path, andrender()calls it aswarn_unknown_subdirs || return 1. Dropping either half restores the silent-loss bug: a fragment under an unknown directory renders nothing, and--checkstill passes because it compares rendered output againstCHANGELOG.mdand both sides agree the entry does not exist. That is not hypothetical — it hid PR #1313's runbook entry on master. Also keep thefind ... >&2 2>/dev/nullredirect order in that block; the reverse swallows the list of lost files.
fix/cuda-adm-picture-ready-race — reproducer for the CUDA/FFmpeg nondeterminism (2026-09-06)¶
Adds scripts/test/repro-cuda-ffmpeg-nondeterminism.sh; no library code is touched. Two things worth knowing before anyone tries to fix the underlying defect:
-
Measure by interleaving, never sequentially. The corruption rate tracks host load (0/60 idle, 14/60 at load ~33, 36/80 at load ~16, 50/50 at load ~69 where it saturates and stops discriminating). Two 80-run samples taken one after another produced an apparent 36→9 "improvement" from a change that an interleaved A/B then showed to be 14/60 vs 14/60 — no effect at all. Run the two arms alternately.
-
Two plausible fixes are already ruled out, by measurement rather than reasoning: waiting on the pictures'
readyevents before the scale-0 DWT2 ininteger_adm_cuda.c, and fencing the shareds->bufagainst the previous frame'ss->strwork. Both are theoretically sound gaps; neither moves the rate. Do not re-propose them without an interleaved measurement. The live lead iscollect_fex_cuda()skippingcuStreamSynchronizeon the ADR-0242drainedpath whilesubmit(N+1)is already overwriting the sharedresults_host.
fix/cuda-adm-picture-ready-race — order caller-written CUDA pictures (2026-09-06)¶
Touches core/src/libvmaf.c, which is upstream-mirrored. Three invariants:
-
The
cuCtxSynchronize()at the top ofread_pictures_extractor_loop()is load-bearing, not defensive. It orders this frame's device data against whoever produced it. With..._PREALLOCATION_METHOD_DEVICEthe caller copies into a libvmaf-owned picture on a stream we never see, and libvmaf records a picture'sreadyevent only insidevmaf_cuda_picture_upload_async()— so in that path everycuStreamWaitEvent(..., ready)in every extractor is vacuous. Removing this barrier as "redundant with the per-extractor ready waits" restores a silent wrong-score bug: 56 of 60 runs corrupted, measured. See ADR-1199. -
It belongs at the dispatch point, not inside an extractor. The corruption was only ever observed in ADM because ADM reads the raw planes first. Moving the barrier into
integer_adm_cuda.cleaves every other CUDA extractor relying on queue position;test_cuda_float_moment_paritywas seen failing under the same GPU contention. -
Do not re-propose the three fixes already ruled out without an interleaved measurement: waiting on the pictures'
readyevents before the scale-0 DWT2, fencing ADM's shareds->bufagainst the previous frame'ss->str, and dropping thedrainedshortcut incollect_fex_cuda(). Each measured 14/60 against 14/60 for control. Reproduce withscripts/test/repro-cuda-ffmpeg-nondeterminism.shunder concurrent CUDA load — CPU load is not a stressor for this race (1/80 at load 22 versus 56/60 with three concurrent CUDA processes), and two builds must be compared by interleaving runs, never sequentially.
fix/container-nv-codec-mirror-fallback — second source for nv-codec-headers (2026-09-06)¶
Touches dev/Containerfile only. Two invariants:
-
Do not "simplify" the fallback back to a single
curl. The original comment justified the single source with "GitHub mirror lags so use code.ffmpeg.org", which is true for an unreleased commit and false for the tag actually pinned —n13.1.15.0is published on both, and the GitHub tarball carries thecuStreamCreateWithPrioritydeclaration the pin exists for. That host was unreachable for over six hours on 2026-09-06 and made the container unbuildable. See ADR-1200. -
Keep the content assertion and the
find-basedcd. The build requiresinclude/ffnvcodec/dynlink_cuda.hand grepsdynlink_loader.hforcuStreamCreateWithPrioritybeforemake install; without it a fallback could install the wrong headers silently, which is worse than the outage. And the two archives unpack to DIFFERENT top-level directories (nv-codec-headersvsnv-codec-headers-<tag>), so the hard-codedcd nv-codec-headersthat used to be here breaks on the mirror.
docs/retrain-gate-status-1246 — measured retrain gate status (2026-09-06)¶
Documentation only. One thing worth knowing:
- The gate table is a measurement log, not a plan. Every cell states how it was checked and on what date. Do not carry a status forward across a rebase without re-running its verification command — the table this replaced had G3 marked FAIL against "PR #1307 & fix/cambi-cuda-context unmerged" when both had already merged, which is exactly the drift the format is meant to prevent.
fix/cuda-speed-chroma-4k-launch — GPU SpEED-chroma singularity contract (2026-09-06)¶
-
A non-zero return from the GPU twins' linalg helpers means hard failure, not singular matrix. The CPU reference overloads one integer for both (
solve_covariance_system()returnscannot_invert, andextract_fex()reads it to impute theuvscore). The GPU twins handle singularity internally and reserve the return value for device errors, so they carry an explicitbool *singular_out. A rebase that "simplifies" that parameter away by re-reading the return value restores a silent wrong-score bug: both channels failing then averages(0 + 0) * 0.5and the run exits 0 with three0.0scores. See ADR-1202. -
SC_SOLVE_WARPS_PER_BLOCKbounds the block size; the block count is what scales with the picture. The pre-fix code had the two inverted, which put every launch above 256 linear systems past CUDA's 1024-thread block limit — i.e. every 4K frame. The SYCL and HIP twins already compute this correctly (local = SOLVE_WG * 8,hipModuleLaunchKernel(..., u_nb, ..., solve_warp)); keep all three consistent. -
The existing GPU parity tests cannot catch this class. They all run below the 256-system threshold, so the launch bug was invisible to
meson test --suite=fast. Verify 4K parity by hand against the CPU backend when touching these files.
fix/codeql-float-widen-mult — float-widening in the vendored PSNR path (2026-09-06)¶
-
compute_psnr()'s(double)diff * diffis a deliberate deviation from upstream. Upstream computes the product infloat. The cast satisfies CodeQL alert 1009 and is safe only because the function is unreachable; an upstream sync that reverts it re-opens the alert but changes no score. -
Do NOT apply the same cast to
core/src/feature/iqa/convolve.c. Its four accumulation sites must keep thefloatmultiply: the AVX2 / AVX-512 / NEON twins widen after multiplying to stay bit-identical (ADR-0138), andtest_iqa_convolvefails the moment the scalar side is widened. CodeQL alert 1005's exact vertical-pass expression carries a narrow source suppression backed by executable SSIM/MS-SSIM/PU21 domain bounds (the product is at most2^28). Preserve the standalone directive immediately before that expression and the domain tests; never broaden the suppression. See research digest 2031.
ADR-1204 / ADR-1205 — ADM CM edge policy and the ssimulacra2 FMA contract¶
-
The ADM contrast-masking edge policy is asymmetric on purpose. The CPU closed form in
core/src/feature/adm_tools.c::adm_cm_thresh3x3_smirrors the near edge to index 1 (i_m1 = (i == 0) ? 1 : i - 1) and clamps the far edge to the last index (i_p1 = (i == h - 1) ? h - 1 : i + 1). Every GPU twin must reproduce both halves. A symmetric mirror (2 * half_w - x - 2) looks tidier and is wrong; it only diverges when the border crop(int)(dim * 0.1 - 0.5)is 0, i.e. band dimensions ≤ 14, so it survives casual testing. If upstream ever rewrites the macro family, re-derive the twins from the closed form rather than from the macros. -
ssimulacra2's YCbCr → linear-RGB conversion is a single-rounded FMA everywhere. ADR-0891 fixed the SIMD kernels; ADR-1205 extended the same contract to the shipped scalar fallback and the four GPU host copies. There are now six copies of these three lines (scalar, CUDA, HIP, Metal, SYCL, plus the SIMD kernels and their tails) and they must stayfmaf()-based and in the same order. The pipeline is ill-conditioned downstream — the edge-diff term is|img - blur(img)|and pooling is a 4-norm — so a 1 ULP change here surfaces as a ~1e-3 score change, not a ~1e-7 one. -
core/test/test_ssimulacra2_simd.cvalidates against its own private scalar reference, not against the shipped functions. That is why the drift in item 2 passed its bit-exactness assertion for as long as it did. When touching the conversion, change the shipped copies and the test reference together, or the test will keep agreeing with itself.
ADR-1216 — motion_fps_weight is applied exactly once (2026-09-07)¶
Branch: fix/gpu-motion3-fps-weight-double.
-
The v1 integer-motion weighting point is
extract(), not the blend.core/src/feature/integer_motion.cscales the SAD-derived score bymotion_fps_weightonce, inextract(), and stores the weighted value asmotion_sad_score.flush()then reads those weighted values back, takes the neighbour min intomotion2, and blendsmotion2intomotion3withmotion_blend()— no second weighting. The CUDA / SYCL / HIP twins each had a host-sidemotion3_postprocess_*()that opened withscore2 * motion_fps_weight, while every caller already passed a weighted, clipped value: the weight came out squared inmotion3_score. If upstream ever moves the weighting point ininteger_motion.c, move it in all three twins together — the twins deliberately have no weight multiplication left in the post-process. -
A
1.0default hides a squared factor.motion_fps_weightdefaults to1.0and1.0² = 1.0, so every parity test that ran the extractors withNULLoptions agreed on a wrong value. The guard is thetest_motion3_fps_weight_applied_oncevariant in each oftest_{cuda,sycl,hip}_motion3_parity.c, which pinsmotion_fps_weight = 0.6and reads the ADR-1183-derivedinteger_motion3_mfw_0.6key. Any option whose default is an arithmetic identity needs a non-default parity variant or it is not actually covered. -
motion_v2has a different, documented divergence — do not "fix" it here. The GPUmotion_v2twins store the raw SAD and apply the weight in the host-side flush, which shifts the seed frame relative to the CPU. That is pre-existing, consistent across all four backends, and described indocs/metrics/motion.md.float_motionon every backend already applies the weight once, at the emission site. Neither was touched by ADR-1216.
ADR-1217 — GPU float-VIF options must reach the kernel (2026-09-07)¶
Branch: fix/gpu-float-vif-options.
-
Kernel-local constants that shadow an option are invisible to review. The CUDA / SYCL / HIP
float_vifcompute kernels each declaredconst float vif_sigma_nsq = 2.0f; const float vif_egl = 100.0f;inside the per-pixel block. The names matched the options exactly, so the arithmetic below them read as correct; nothing at the declaration site said "this is a hardcoded default". When touching these kernels, keep the values as kernel arguments: a missing argument is a compile error, a shadowing local is not. -
sigma_max_invmust be derived on the host, the CPU's way.vif_tools.c::vif_statistic_scomputespowf(vif_sigma_nsq, 2.0f) / (255.0 * 255.0)—powfinfloat, the division indouble, narrowed tofloaton assignment. The host copies reproduce that expression exactly. Do not "simplify" it tonsq * nsqor move it into device code: devicepowfis not guaranteed to round like the host's, and this value feeds thesigma1_sq < vif_sigma_nsqbranch that setsnum_valdirectly. -
cuLaunchKernel/hipModuleLaunchKernelsilently ignore surpluskernelParamsentries. Adding an argument to the kernel signature without adding it to every launch site reads uninitialised parameter memory instead of failing.float_vif_cuda.chas twofunc_computelaunch sites (scale 0 and the scale 1..3 loop) andfloat_vif_hip.cfunnels all four through one helper; both were updated together. -
float_vif_hiphad no parity test.test_hip_vif_parity.ccovers the integervif_hiptwin — a name close enough to look like coverage. The newtest_hip_float_vif_parity.cfills the gap. When adding a HIP twin, check that the parity test named after it actually targets it.
ADR-1218 — SpEED singular-covariance contract on the GPU twins (2026-09-07)¶
Branch: fix/gpu-speed-singular-solution.
-
A singular covariance matrix is a routine condition, not an error.
speed.czeroes the solution and reports the singularity separately from the return code, because the caller still emits a score. The GPU twins have to keep those two channels apart: the function return stays reserved for hard CUDA/HIP/SYCL failures, andsingular_outcarries the numerical condition. ADR-1202 established this for the three chroma twins; ADR-1218 extends it to the three temporal twins. If upstream changesspeed_extract_score()'s one-sided rule, all six twins move together. -
Zero the DEVICE solution, never the host staging buffer. The score kernel reads
d_sol;h_indtermis re-downloaded fromd_indtermat the top of every pipeline run, so a hostmemseton the singular path is dead code that reads as if it did something. UsecuMemsetD8Async/hipMemsetAsync/q.memseton the stream or queue that the score kernel will use. -
The existing SpEED parity fixtures cannot reach the regular path. They are 768x432, whose chroma planes give 4x2 = 8 blocks for a 25x25 covariance — rank-deficient by construction, so
is_matrix_regular()is false on every frame. Any new SpEED test that needs a regular frame must be at least 960x960 (36 chroma blocks, 144 luma). This is easy to get wrong: the test looks like it is exercising the normal path and is not. -
A flat plane is the wrong singular fixture. SpEED subtracts 128 in
picture_copy, so a plane flat at the neutral level zeroes the independent term as well, andsum(sol * indterm)vanishes whatever the solution holds. A plane flat at any other level gives an all-zero covariance, every entropy collapses belowbase_entropy, andget_speed_score()returns exactly 0. Use a COLUMN-CONSTANT plane: singular (20 zero eigenvalues) with five large ones and a non-zero independent term. Better still, make exactly one side singular — that is the case with an observable score difference.
ci/release-bot-pat-fallback — second release-bot identity (2026-09-06)¶
Fork-only CI. One invariant:
- Never fall back to
GITHUB_TOKEN. Themode=nonebranch must stay: it leaves the pipeline idle (warning onpush, error onworkflow_dispatch, ADR-1171) rather than opening a release PR that can never be merged. A PR opened byGITHUB_TOKENreceives zero check runs — that is a GitHub loop-breaker, not a misconfiguration — so "just use the default token" is the one resolution that recreates the original bug (ADR-1151). The two acceptable identities are the App (preferred) andRELEASE_BOT_TOKEN; both are masked through the singleResolve the release-bot tokenstep so no downstream step has to know which is active.
feat/release-candidates — rc.N prereleases before 1.0.0 (2026-09-06)¶
Fork-only release tooling and CI. Three invariants:
-
The prerelease pattern is narrow on purpose.
^v<major>.<minor>.<patch>(-rc\.<n>)?$—rconly, dotted integer, no leading zero. Widening it to full SemVer accepts-beta,-rcbare,-rc.01and-alpha.1+build, none of which this project ships; each is a way to mis-tag a release. See ADR-1201. -
Publishing workflows check tag/flag CONSISTENCY, not absence. Three workflows used to
exit 1onprerelease == true. They now require an-rc.Ntag to be marked prerelease and a final tag not to be. Restoring the blanket rejection blocks the whole RC line; dropping the check entirely allows an RC published as stable, which takes thelatestimage tag. -
The ADR-1151 contract gate must test the prerelease suffix BEFORE its
sort -Vcomparison.printf '1.0.0\n1.0.0-rc.1\n' | sort -V | head -1returns1.0.0, so a naive comparison concludes the RC line has already reached 1.0.0 and fails every RC build whilerelease-asis legitimately still present. The*-*early return is load-bearing, not cosmetic.
feat/rocm-10-migration (ADR-1225)¶
Files touched: dev/Containerfile, dev/docker-compose.yml, dev/AGENTS.md, docker/Dockerfile.production-gpu, docker/Dockerfile.node, .github/workflows/build.yml, .github/workflows/libvmaf-build-matrix.yml, .github/workflows/docker-publish-production.yml, .github/workflows/docker-publish-operator-node.yml, core/src/meson.build (comments only), renovate.json, scripts/ci/install-rocm-from-image.sh (new), AGENTS.md, docs/backends/hip/overview.md, docs/development/dev-mcp.md.
Rebase impact: none against upstream Netflix — every file here is fork-local. The conflict risk is against other fork branches that touch the GPU-SDK layer of dev/Containerfile or either CI HIP lane.
Invariants a future rebase must not undo:
-
Do not restore the
repo.radeon.comapt path.ARG ROCM_VER=and therocm/apt/<ver>source line are gone on purpose: AMD froze that channel at 7.2.4 when ROCm moved to TheRock at 7.14, so re-adding it silently pins the toolchain back to ROCm 7. A rebase that resurrects the apt block from an older branch will still build — which is exactly why it needs saying. -
librocprofiler-registeris not a profiler component. It is a hard link-time dependency oflibamdhip64.soandlibhsa-runtime64.so. Any widening of therocm-srcprune list — for instance back to a tidy-lookinglibrocprof*glob — makes every HIP binary fail at load withlibrocprofiler-register.so.0: cannot open shared object file. The stage's hipcc smoke compile is the guard; keep it. -
/opt/rocm/{bin,lib,include,llvm,…}are/etc/alternativessymlinks in the ROCm 10 image. Copying or extracting/opt/rocmalone yields dangling links, so both therocm-srcstage andscripts/ci/install-rocm-from-image.shrepoint them atcore-10.0/<name>relative to/opt/rocm. Do not "simplify" either loop away, and do not solve it by copying/etc/alternatives— that directory is shared with the consuming image's own packages. -
node-rocmneeds the whole closure, laid out as it was. ROCm 10'slibamdhip64.sopullslibLLVM/libclang-cpp/libamd_comgr/librocm_kpack/librocprofiler-registerplus therocm_sysdepsbundle, wired by$ORIGIN-relative RPATHs. Reverting to the old two-library flat copy into/usr/local/libproduces an image whose every HIP binary dies at load. Verify withlddin a scratch stage carrying only the copied files. -
HSA_OVERRIDE_GFX_VERSIONstays out ofdev/docker-compose.yml. ROCm 10 supportsgfx1036natively; re-adding the10.3.0alias would map the agent togfx1030while meson compilesgfx1036code objects.
fix/pkgconfig-advertises-abi-version — libvmaf.pc Version: is the ABI version (ADR-1235)¶
- Touches:
core/src/meson.build(thepkg_mod.generate()call),core/meson.build(the SONAME comment only). - Invariant:
pkg_mod.generate(version:)takesvmaf_soname_version, nevermeson.project_version(). The two numbers are deliberately independent (ADR-1151): release-please owns the product line, which starts at 1.0.0, while the C API stays on the 3.x line that shipslibvmaf.so.3.libvmaf.pcmust advertise the C API number, because that is what consumers gate on — unpatched upstream FFmpeg requireslibvmaf >= 2.0.0and this fork's ownffmpeg-patches/0004/0005requirelibvmaf >= 3.0.0. Restoringmeson.project_version()here re-breaks every FFmpeg consumer the moment the product version goes below 2.0.0, which it already has. - Rebase impact: Low, but sharp. Upstream Netflix's
core/src/meson.buildspells this lineversion: meson.project_version()and their product version is their ABI version, so the two agree upstream and the line looks like an ordinary upstream hunk. A sync that takes theirs silently reintroduces the bug, and it will not show up in any unit test — only in the FFmpeg lanes andDocker Image Build, and only once the product version is below 2.0.0. Always keep ours. The guard is the comment block immediately above the call; if a merge drops the comment, the resolution was wrong. - Verify after any sync: configure with a 1.x product version and check the generated
.pcagainst the bounds downstream actually uses —
meson setup build core -Denable_cuda=false -Denable_sycl=false
PC=$(find build -name libvmaf.pc | head -1)
grep '^Version:' "$PC" # expect 3.0.0
PKG_CONFIG_PATH=$(dirname "$PC") pkg-config --exists 'libvmaf >= 2.0.0' # must succeed
PKG_CONFIG_PATH=$(dirname "$PC") pkg-config --exists 'libvmaf >= 3.0.0' # must succeed
ADR-1223 — CUDA floor at compute capability 8.0, CI on 13.3.1 (2026-09-07)¶
Branch: feat/cuda-ampere-floor-and-133.
-
The gencode list has a floor now, and it is enforced twice.
sm_75and the CUDA-12-onlycompute_50PTX are gone fromcore/src/meson.build, and theclang-CUDA fallback targetssm_80. Dropping a gencode entry alone is not enough: without a runtime check the failure surfaces asCUDA_ERROR_NO_BINARY_FOR_GPU(222) fromcuModuleLoadData, inside whichever feature extractor loaded first.check_device_arch()incore/src/cuda/common.crejects a sub-8.0 device at init on BOTH paths —init_with_primary_context(before retaining a context, so the unwind is free) andinit_with_provided_context(the caller's context still has to clear the floor). If a future ADR moves the floor again, moveVMAF_CUDA_MIN_COMPUTE_MAJOR/_MINORincore/src/cuda/common.hand the gencode list together. -
The capability predicate is lexicographic, not
major*10 + minor.vmaf_cuda_arch_supported()compares major first and only falls through to minor on a tie. A naivemajor >= 8would admit a hypothetical 8.x below the floor; a naive minor comparison would admit 7.9. Both edges are pinned incore/test/test_cuda_arch_floor.c, which is a pure-function test precisely because no runner in the fleet has a Turing GPU to test the rejection on. -
Upstream Netflix still ships
sm_75. A rebase that takes upstream'smeson.buildgencode block wholesale will silently reintroduce it. Keep the fork's block. -
The
Jimver/cuda-toolkit13.3 blocker is closed and its comment was stale.T-CI-JIMVER-CUDA-133-NOT-AVAILABLEwas real at action v0.2.35; v0.2.36 serves 13.3.1, whichbuild.yml's Linux leg already used while the build matrix stayed pinned at 13.2.0 citing the old ticket. The Windows legs never used Jimver at all — they fetch NVIDIA's network installer directly (cuda_<version>_windows_network.exe) and were never blocked. Verify an installer URL resolves before bumping it: 13.3.0 and 13.3.1 return HTTP 200, 13.3.2 does not exist. -
Windows lanes export a versioned env var.
CUDA_PATH_V13_2becameCUDA_PATH_V13_3; it is derived from$cudaMajorMinorby hand in two separate workflow files, so both move together with the version.
ADR-1229 — the MCP server is the Go binary (2026-09-07)¶
Files touched: dev/Containerfile, dev/scripts/dev-mcp-entrypoint.sh, dev/AGENTS.md, docs/development/dev-mcp.md, mcp-server/vmaf-mcp/README.md (deprecation banner only).
Rebase impact: none against upstream Netflix — the MCP surface is entirely fork-added.
-
vmafx-mcpneeds no install step. Thego-buildstage already doesCOPY --from=go-build /out/ /usr/local/bin/, which includes every./cmd/...binary. A rebase that "restores" a missing install line for the MCP server is adding a second, conflicting copy. -
Do not reinstate
pip install -e /build/vmaf/mcp-server/vmaf-mcp. It was removed deliberately. The Python package is deprecated and its tools are fully covered by the Go binary; reinstating the install quietly returns 15k lines of Python to the image without returning any capability. -
The entrypoint still must not daemonise the server. The reasoning in its header predates the Go swap and still holds: the container stays alive and clients attach with
docker exec -i. The compose healthcheck must therefore remain a CLI check (vmaf --version), nottest -S /sockets/vmaf-mcp.sock— see the ADR-0641 invariant indev/AGENTS.md. -
stdout belongs to JSON-RPC. The Go server logs to stderr. Any change that sends log output to stdout corrupts the protocol stream and shows up as
mcp stdio returned empty responserather than as a logging bug.
ADR-1231 — container base images come from build-config.env (2026-09-07)¶
Rebase-sensitive invariants introduced by this change:
-
A Dockerfile must never name a base image directly again. Every base is an
ARGwhose default mirrorsbuild-config.env, andscripts/ci/check-base-image-single-source.shfails on any digest-pinned literal. An upstream merge that reintroduces a literalFROM debian:…@sha256:…will fail the gate rather than silently forking the pin. Resolve by moving the value intobuild-config.envand referencing theARG. -
COPY --from=<digest-pinned image>is a base-image pin and is rejected. This is the non-obvious half. Four such pins existed and were the most stale in the tree, because noFROM-oriented search finds them. Use a named stage (FROM ${CUDA_RUNTIME} AS cuda-runtime-libs, thenCOPY --from=cuda-runtime-libs …); BuildKit prunes the stage when the selected target does not use it, so it is free. -
Do not hand-edit an
ARGdefault to fix a drift failure. Editbuild-config.envand runmake base-images-sync. Hand-editing puts the two copies back out of agreement in the other direction, which the gate will then report against the file you just "fixed". -
The Ubuntu 24.04 exemptions are deliberate and self-closing.
ROCM_BUILDER,ROCM_RUNTIME,ONEAPI_BUILDERandONEAPI_RUNTIMEare listed indistro_exemptin the gate. Do not extend that list to silence a new failure — it exists only because those two migrations need matching source changes (PR #1386 for ROCm; the oneAPI 2026.1 restructure documented indocs/research/1231-base-image-single-source.md). Each entry is deleted when its migration lands. -
docker/dev/*.Dockerfileis intentionally outside the gate. Those files pin Alpine / Arch / Fedora precisely because they are not the release distro. Unifying their bases would defeat the portability matrix they exist to run.
FFmpeg percentile patch replay (2026-09-08)¶
Patch 0005 already introduces max and percentile mappings in the shared pool_method_map. Patch 0018 must apply to that cumulative state and guard those existing percentile entries with VMAF_HAVE_PERCENTILE_POOLING. Do not re-add the mappings in 0018. Replay every entry in series.txt against a fresh upstream checkout; the old 000*-*.patch glob missed patches 0010–0018.
ADR-1238 — Go validation is a required impact-routed check (2026-09-08)¶
Keep go vet + go test synchronized between go-ci.yml, its # required-aggregator marker, and the aggregator's required list. The workflow starts without path filters and gates heavy steps on go_checks (go plus c_core); its own changes force full impact. Preserve ready-for-review coverage and the documentation-only no-work success. Run python3 scripts/ci/test_go_workflow_contract.py and python3 -m unittest scripts/ci/tests/test_ci_impact.py after workflow rebases. Predictor model-card reads use os.Root; do not restore an unconfined os.ReadFile or symlink escape while reconciling stub warnings.
ADR-1236 — version single-sourcing and Python dependency unification (2026-09-08)¶
Rebase-sensitive invariants introduced by this change:
-
python/pyproject.tomlis the single owner of Python runtime dependencies.python/setup.pyintentionally removes duplicateinstall_requires=[...]. Setuptools natively loads dependencies frompython/pyproject.toml. If an upstream merge reintroducesinstall_requiresinpython/setup.py, delete the block so dependencies remain single-sourced. -
python/requirements.txtis mechanically generated — never hand-edit.python/requirements.txtis derived frompython/pyproject.tomlusingscripts/ci/check-python-requirements-single-source.sh --write(ormake python-deps-sync). Any manual edits will failscripts/ci/check-python-requirements-single-source.shin CI and pre-commit hooks. -
Renovate must ignore
python/requirements.txt.renovate.jsonincludes"python/requirements.txt"inignorePaths. Renovate should only managepython/pyproject.tomlto prevent competing PRs. -
Package version disagreements follow the newest-version policy. Divergent version pins across submodules or dev requirements (e.g.
numpy,scipy,matplotlib,pyarrow) must never be downgraded to resolve a merge conflict. The currentvmafbuild/runtime floor is>=2.5.3for numpy,>=1.18.1for scipy,>=3.11.1for matplotlib, and>=25.0.1for pyarrow. Package version definitions acrosspyproject.tomland__init__.pyfiles must remain synchronized.
Preserve the Level Zero container consumer check when rebasing the ADR-1236 workflow-checker refactor. Do not reintroduce scientific-stack globals without consumers or drift checks; native ORT archive roles and Python dependency floors are separate contracts.
Feature-option sentinel cleanup (2026-09-08)¶
feature_extractor.cpp and feature_name.cpp iterate VmafOption tables through their existing null-name sentinel. Keep the early empty-dictionary return before reporting the first missing option, and preserve aliases, default-value omission and dictionary sorting when rebasing the private feature-name helpers. Public C signatures, emitted keys and GPU fallback behavior are unchanged. Recheck test_feature, test_feature_extractor and test_opt; see the option-sentinel research digest for the focused lint scope and factory-specific cppcheck model corrections.
fix/fex-pool-stable-entries — internal pool growth (2026-09-08)¶
Keep the pool's pointer table separate from stable fex_list_entry allocations. An acquisition retains its entry across pthread_cond_wait(); relocating live entries during geometric growth loses the condition-variable identity. Preserve construction before cnt publication, allocation-overflow checks and one-time options/condition-variable cleanup. test_fex_pool_growth covers synchronized ninth-entry growth with forced relocation and native allocation. Current scoring does not use these acquire/release operations; this fixes the compiled internal API, not a demonstrated scoring hang. No public header or FFmpeg surface changes.
fix/fex-pool-consumer-model — factory visibility (2026-09-08)¶
Cppcheck's POSIX model resolves pthread types, but fex_ctx_vector.cpp cannot see the separately compiled pool factories. Preserve only the four documented uninitMemberVarNoCtor member markers for fex, opts_dict, ctx_list and full; do not suppress atomic fields, the outer pool, or uninitialized reads. The real-header negative control must continue reporting an uninitialized member read. Pool entry construction and runtime behavior are unchanged.
refactor/test-cambi-lint — preserved test coverage (2026-09-08)¶
Keep the named assertion groups and dispatcher groups in test_cambi.c small without dropping cases or changing fixture/expected values. The cleanup retains all 144 assertion expressions/messages, 30 array initializers and 23 original registrations in order. Its direct feature/cambi.c include intentionally tests private static helpers; the narrow include exception does not export a new API. No production code, golden assertions or FFmpeg integration changed.
PSNR format-table cleanup — 2026-09-08¶
No rebase impact: the private C++ format table has an explicit member default and uses a projected standard lookup. Preserve all twelve format/constant rows and psnr_constants() return/output behavior. The C header is unchanged; see the differential checks.
pdjson nesting boundary and lint cleanup (2026-09-08)¶
fix/pdjson-lint-20260908 keeps the pdjson Unlicense provenance and all existing private parser entry points. Preserve ADR-1061's 512-container count, checked capacity arithmetic and publication of stack_top only after successful admission/allocation. Reject zero and oversized stack increments before allocation. The old > PDJSON_STACK_MAX comparison admitted 513 levels. Keep the named zero lookahead sentinel (existing nonzero event values unchanged), read-only getter const qualifiers, first-error formatter guard and split parser phases. The source has only the ADR-1138 C NULL exception, not a blanket NOLINT. Run test_pdjson plus the model/ownership tests after an upstream parser refresh; its streaming, skip, Unicode, invalid-input, allocation and depth assertions are behavioral contracts. See the digest.
Cppcheck exhaustive configured analysis (2026-09-08)¶
Preserve --check-level=exhaustive in both scripts/ci/lint-configured.py and the required Cppcheck workflow, with every existing diagnostic selection and configured command variant. Do not restore normal's branch budget or suppress its coverage notices. The actual-tool suite must retain its branch-heavy positive control and real uninitialized-member/constructor negative controls. ADR-1245 records the measured runtime tradeoff; backend/runner differences still require validation.
Integer VIF AVX-512 native lint cleanup (2026-09-08)¶
Preserve the private forced-inline stages in feature/x86/vif_avx512.c, the fixed mutable-state callback ABI, and the two ADR-0503 noinline/noclone block helpers. Each accumulator retains tap order, lane packing and shifts. Keep the fused two-channel horizontal mean and three-channel energy loops: independent channel loops caused a measured 8-bit slowdown despite bit-exact results. Retain the 16-sample vertical extent / 32-sample vector step and scalar overwrite for 8-bit statistics. The dedicated test_integer_vif_avx512_stages links private configured objects and compares the actual scalar implementation, including reflected temporary-row padding. Keep its Meson registration conditional on is_avx512_enabled, independent of float features. No public/FFmpeg surface change; no golden assertions changed. See Research-2046.
Registration option-copy ownership (2026-09-08)¶
Keep both private cleanup helpers in libvmaf.c: dictionary-copy failure can leave a partial destination, and context-create failure does not consume its input options. Explicit registration still consumes its original dictionary after the existing argument/name guards; model/worker sources remain borrowed. Preserve prior registered features on later failure and successful ownership transfer. Keep the public rejection test independent of the Linux-only partial copy interposer. No public signatures, feature calculations, FFmpeg filter contract or Netflix assertions change. Six exact retained DNN/metadata exports and the weak glibc ABI marker remain documented in Research-2048.
Cppcheck public entrypoint model (2026-09-08)¶
Preserve the one scripts/ci/cppcheck-public-entrypoints.cfg input in local lint-configured.py and the Cppcheck workflow. ADR-1246 permits only reviewed public roots with VMAF_EXPORT declarations in installed headers; private or vendored helper retention is a separate decision. Keep the configured-driver hook's cfg/header/workflow/control triggers and the required real-tool missing-model, unused-private and body-defect negatives. Names are scope- and linkage-blind in Cppcheck; do not reuse public names for private/static code. Both existing severity selections, POSIX model, exhaustive depth, command variants and failure handling remain unchanged. No native/public API, FFmpeg patch, numerical assertion or baseline change. See Research-1246.
scripts/dev/gc_workingdir.py — local state-tree GC (2026-09-15)¶
Two invariants are load-bearing and were both found by the tests rather than by inspection, so do not "simplify" them on a rebase:
- Citations are stored relative to the state root.
git grepreports them with the.workingdir2/prefix, while the walk compares paths relative to the state root. The first version compared the two directly, so no citation ever matched and the protection silently did nothing — the count printed fine. The prefix is stripped on read; keep it that way or add the prefix on both sides. - A citation protects the named path, not the subtree beneath it. A run directory is cited precisely because a document links to it, so the directory must survive; its Go cache, meson build tree and object files must still be reclaimable from underneath it. An earlier version pruned the walk at any cited directory, which protected the entire run tree and would have reclaimed almost nothing, since the 14.6 GB was inside cited run directories.
PROTECTED_NAMES (rescue, evidence, archive, netflix) is a separate, absolute guard: those roots are never walked, citation or not. rescue/ holds the 134-branch recovery bundle.
ADR-1219 — CAMBI TVI bisection and border rules on the HIP/Metal twins (2026-09-07)¶
Branch: fix/gpu-cambi-tvi-shared-bisection.
-
Never re-derive the TVI table in a twin — call
vmaf_cambi_init_tvi_and_vlt(). It is host-side scalar code that runs once ininit(), so there is no performance argument for a per-backend copy, and the CPU's bisection is subtle enough that two independent hand-ports both got it wrong in the same way: they searched the negatedtvi_hard_threshold_conditionand seeded from luma 0 instead ofluma_range.foot. The helper also computesvlt_lumaand validates the derived band, so a twin that calls it cannot drift on any of the three. -
cambi.c::filter_modeleaves output rows 0 andheight-1UNFILTERED. Its vertical writeback sits underif (i > 1)and covers rows1 .. height-2; the horizontal results for the two border rows live only in the 3-row ring buffer and are never written back. A GPU twin that does a clean separable H-then-V pass over all rows is therefore wrong at the top and bottom edge. The guard isif (axis == 1 && (y == 0 || y >= height - 1)) return;— the vertical pass reads the H result from a scratch buffer and writes into the buffer that still holds the pre-filter image, so returning early preserves it exactly. -
get_spatial_mask_for_index()ZERO-PADS its 7x7 box sum; it does not clamp. The summed-area table ismemsetto zero,dp_widthcarries2 * pad_size + 1extra columns, andderiv_valid = (i < height)gates the row dimension, so an out-of-frame tap adds nothing. Clamping each tap to the border pixel counts that pixel's zero-derivative flag up to three extra times per axis and flipsbox_sum > mask_indexon a band of border pixels. -
A CAMBI fixture must actually band. CAMBI counts neighbour differences of
1 .. num_diffs(4 at the defaultmax_log_contrast = 2). The HIP and Metal parity fixtures were 8-bit gradients stepping 32 code levels every 32 columns — an edge, not banding — and scored0.0on the CPU as well, so the parity assertion was0 == 0and hid a total score collapse. Use a 10-bit gradient of one code level every two columns held inside the TVI band (200..900), and assert the CPU score is non-degenerate before comparing. The CUDA and SYCL gradient fixtures are still degenerate the same way; only CUDA's second textured fixture asserts anything.
ADR-1220 — float-ADM options must reach the GPU kernels (2026-09-07)¶
Branch: fix/gpu-float-adm-options.
-
adm_p_normhas FOUR application points, not one.adm_tools.capplies it in the DLM numerator sum and the CSF denominator sum (each special-casingp == 3to a literal cube), in the pooling rootpowf(accum, 1.0f / adm_p_norm), and insideget_noise_constant(w, h, weight, p) = powf(w * h * weight, 1.0f / p). A twin that honours only the last of these — as all four did, and only for the AIM term — produces a hybrid quantity: a sum of cubes raised to1/p. When touching the ADM pooling, change all four together. -
Keep the CPU's
p == 3fast path in the kernel. Devicepowf(x, 3.0f)is not guaranteed to equalx * x * x, so replacing the cube unconditionally would move the DEFAULT path — every shipped model — for no benefit. The kernels carryfadm_pnorm_term(x, p) = (p == 3.0f) ? x*x*x : powf(x, p), mirroringadm_tools.cexactly. -
adm_bypass_cmapplies to BOTHadm_cm()calls.adm.cpasses it to the DLM CM and to the AIM CM. A twin that gates only the DLM kernel is still wrong. -
adm_skip_scale0is a POOLING rule, not a reporting rule.adm.csetsnum_scale = 0andden_scale = 1e-10for scale 0, so scale 0 drops out of the pooledadm2andaim. Zeroing only the reportedadm_scale0sub-score — what the Metal twin did — leaves the pooled score wrong on every frame. No kernel change is needed for it:adm_dwt2_lo_swrites onlyband_a, whichadm_dwt2computes identically, so scales 1..3 are unaffected. -
On Metal, the CM uniform's two padding slots now carry the options.
FadmCsfinfloat_adm.metalandFadmCsfHostinfloat_adm_metal.mmmust stay byte-identical;_pad0/_pad1becamep_norm(float) andbypass_cm(uint), so the size and alignment are unchanged. If you add another option, add it to both structs in the same commit. -
The SYCL float-ADM parity test compared only
adm2. On its fixture the aggregate alone could not see the p-norm defect at all — the per-scale sub-scores could (adm_scale0drifts by3.09e-03whereadm2stays inside the gate). An ADM parity test must read all five features.
ADR-1221 — MS-SSIM clip_db is a dB ceiling (2026-09-07)¶
Branch: fix/gpu-ms-ssim-max-db.
-
clip_dbdoes NOT clamp the linear score.float_ms_ssim.cderivesmax_db = ceil(10 * log10(peak * peak / mse))withmse = 0.5 / (w * h)atinit(), andconvert_to_db()returnsMIN(-10*log10(1 - score), max_db)withscore >= 1.0short-circuiting tomax_db. A twin that clamps the linear score into[0, 1]and then converts without a ceiling returns+Inffor an identical reference/distorted pair and an uncapped value for every other high-similarity pair. Keep themax_dbfield and thems_ssim_convert_to_db()helper in sync with the CPU on all three twins. -
max_dbbelongs ininit(), not inextract().w,handbpcare fixed for the extractor's lifetime; the CPU computes it once and so do the twins. Its derivation uses the CPU's exact expression and integer types —peak * peakinunsigned,0.5 / (w * h)indouble— so do not "simplify" it. -
A dB-option parity test needs an IDENTICAL pair. On a merely high-similarity fixture
-10*log10(1 - score)stays well belowmax_dband both paths agree, so the variant passes against the unfixed twin. Feeding the same picture as reference and distorted drives the score to1.0, which is where the ceiling binds — and is an ordinary thing for a user to do. -
Metal is out of scope here.
float_ms_ssim_metal.mmexposes onlyenable_lcsand rejectsenable_db/clip_db/enable_chroma, so it cannot produce a dB score at all. That is a feature gap rather than a wrong answer; it is tracked indocs/state.md.
ADR-1226 — the CUDA AIM CM launch is sized by SM count (2026-09-07)¶
Files touched: core/src/feature/cuda/integer_adm/adm_cm.cu, core/src/feature/cuda/integer_adm_cuda.c.
Rebase impact: none against upstream Netflix — the AIM CM GPU kernels are fork-local (ADR-0746). Conflict risk is against other fork branches touching integer_adm_cuda.c.
-
adm_cm_aim_line_kernel_8no longer exists. The macro now instantiates_2and_4, and the host picks between them per launch from the device's SM count. A rebase that reinstates acuModuleGetFunction(..., "adm_cm_aim_line_kernel_8")will fail at init withCUDA_ERROR_NOT_FOUND, which is the good outcome; a rebase that reinstates the fixedrows_per_thread = 8while keeping the new module lookups will silently launch the wrong grid for the instantiation it got. -
__launch_bounds__is absent on purpose. The kernel still reports 255 registers with spill, so it looks like an obvious oversight. It was measured:(128, 5)cuts registers to 96 and spill to 8 B — and changes runtime by 0.6%, because the kernel is grid-limited, not occupancy-limited. On top of therows_per_threadfix it is a 3.4% regression. The comment aboveADM_CM_AIM_LINErecords this; do not delete it and do not add the attribute without re-running the sweep in ADR-1226. -
rows_per_threadmay not be raised back to 8 "for efficiency". It is what decides how much of the GPU the launch uses. There is no x-decomposition to compensate, and adding one would change where the>> shift_inner_accumrounding happens and break CPU bit-exactness. -
sm_count == 0is a supported state. IfcuDeviceGetAttributefails, the launch picks the wider instantiation rather than failing. Keep that fallback: it reproduces the pre-ADR-1226 behaviour on a device whose attributes cannot be read.
ADR-1209 — --gpumask and upstream's negative-value accident¶
-
Do not "restore"
--gpumask -1during an upstream sync. Upstream'sparse_unsignedcallsstrtouldirectly, and POSIXstrtoulsilently converts"-1"toULONG_MAXwithout settingerrno. This fork rejects a leading'-'before callingstrtoul(core/tools/cli_parse.cpp::parse_unsigned) precisely to stop that. A sync that pulls upstream's parser back in re-opens the hole for every unsigned option, not just this one. -
core/tools/test/test_vmaf_cuda_gpumask.shdiverges from upstream on purpose. It uses--gpumask 1where upstream writes--gpumask -1. Both mean "disable the GPU feature extractors"; only the fork's spelling survives the fork's argument validation. If a sync reverts those two lines, the test goes back to failing on any host with a GPU while still passing CI on GPU-less runners. -
--gpumaskis not a per-op bitmask. Any non-zero value disables GPU feature-extractor selection wholesale for CUDA and SYCL. The$bitmaskplaceholder in the usage string is inherited and inaccurate; the reference table indocs/usage/cli.mdcarries the real contract.
ADR-1211 — HIP is host-pic; device kernels need staged input¶
-
VmafPicture::data[]is HOST memory under the HIP backend (ADR-0530 host-pic). Any HIP kernel that reads picture planes must be handed a device pointer that the extractor staged itself — passingpic->data[i]straight through faults the GPU with "Page not present or supervisor privilege" and kills the process. -
Do not copy the CUDA call shape verbatim when porting an extractor to HIP. The CUDA twins receive a device picture from
vmaf_cuda_picture_*, so theirdwt2_*_device(...)-style helpers take a device pointer that the caller never had to produce. That is exactly howinteger_adm_hipacquired this bug.integer_psnr_hipshows the correct shape: per-plane device buffers allocated in init,hipMemcpy2DAsynchost-to-device in extract. -
Staged rows are tightly packed, so the element stride passed to the kernel is the plane width, NOT
pic->stride[i]. Reusing the picture's stride against a staged buffer reads past the end of each row.
ADR-1215 — per-plane CUDA kernels must take the plane index¶
-
psnr_cuda_dispatchpassesplanefor both bit depths; both kernels must declare it.cuLaunchKernelignores a surplus trailing argument, so a kernel that omits the parameter compiles, launches and silently readsdata[0]. When adding or syncing a per-plane CUDA kernel, check the kernel signature against thekernelParamsarray by count, not by whether it runs. -
A flat-chroma fixture cannot see a wrong-plane chroma read. Both sides report the
psnr_maxsentinel for identical chroma. The 10-bit variant oftest_cuda_psnr_paritytherefore carries non-flat, ref/dist-different chroma; keep it that way.
ADR-1212 / ADR-1213 — bit-depth normalisation in GPU moment twins, HIP chroma geometry¶
-
The CPU
float_momentnormalises before it accumulates.float_moment.ccallspicture_copy(), which divides high-bit-depth samples by 4 / 16 / 256, and only then doesmoment.csum them. A GPU twin that accumulates the raw codeword must divide its sums by the scaler (first moment) and scaler squared (second moment) on the host — see themoment_scalerblock incuda/integer_moment_cuda.c,sycl/integer_moment_sycl.cppandhip/float_moment_hip.c. Dropping that block reintroduces a 4x–256x error that no 8-bit test can see. -
Parity fixtures must include a >8 bpc case.
FIXTURE_BPCis#ifndef-guarded in thefloat_momentparity TUs and meson registers a_10bitvariant per backend. When adding a twin for any extractor that consumes samples, register a 10-bit variant too; an 8-bit-only fixture has a scaler of 1 and proves nothing about bit-depth handling. -
Chroma geometry comes from
picture.c, not fromw >> 1. Subsampled planes are(w + ss) >> sswide (ceil). Any per-backend staging code that re-derives the size must use that formula;ciede_hipis the example of what floor does on odd widths.
ADR-1206 (SYCL) — which parity tests get a large-fixture variant¶
-
test_sycl_motion_add_uv_parityis registered insycl_parity_large_fixture_tests. ADR-1326 replaced the CPU-float comparison and empirical 2e-4 tolerance with an independent scalar oracle for the fixed coefficients, reflect-101 borders, two rounding stages, exact SAD and YUV420 normalization. Preserve both the 256x144 and 960x540 registrations and the derived binary64 roundoff bound; do not restore the former exclusion. -
HIP and Metal have no large-fixture variants yet, on purpose. The
#ifndef FIXTURE_Wguards were deliberately not applied to their parity TUs either, so there is nothing half-wired to trip over. Register them only together with a run on real hardware. -
The
float_ssimskip is a contract assertion, not a workaround. The GPU twins are v1 scale=1-only and rejectmin(w, h) >= 384with-EINVAL, while the CPU decimates. The large variant stays registered and treats that specific refusal as a skip; any other failure is real.
ADR-1207 / ADR-1208 — ISA invariance and the edge-diff subtraction¶
-
edge_diff_map's per-pixel difference is taken indouble, in every implementation.ed = fabs((double)a - (double)am)— both operands arefloat, so the double subtraction is exact, whereas subtracting in float rounds first. All four SIMD kernels (AVX2, AVX-512, NEON, SVE2) originally vectorised the subtract in float and promoted afterwards, while their own scalar tails used the correct form. Do not "optimise" the subtraction back into vector float: the per-lane loop is scalar anyway, so the vector subtract buys nothing and costs bit-exactness. -
test_ssimulacra2_simd::test_edgecannot catch item 1. It compares the kernel against a scalar reference defined inside the test TU, at 33x21, where the float subtraction happens to be exact. It passed both before and after the fix. The gate that covers this iscore/test/test_feature_isa_invariance.c. -
test_feature_isa_invarianceasserts bit-identity, deliberately. It runs each feature twice through the public API, once with the host ISA and once withcpumaskdisabling every SIMD flag, andmemcmps the twodoublescores. If an upstream sync adds a feature with a SIMD path, add it to theCASEStable. If it makes a feature reject the 256x192 fixture, fix the fixture rather than removing the feature — the size is chosen sofloat_ms_ssim(>= 176 px) andfloat_ssim(auto-scale 1 below 384 px) both run.
ADR-1222 — code-scanning scope, not code-scanning annotations (2026-09-07)¶
Branch: fix/code-scanning-alert-sweep.
-
# nosemgrepcannot close a GitHub code-scanning alert. Semgrep's docs: anosemgrepcomment "still generates findings records that are automatically set to Ignored triage state, rather than excluding code from scanning entirely." That state lives on the Semgrep platform; this workflow writes SARIF and uploads it, so GitHub keeps the alert. Do not add morenosemgrepdirectives expecting alerts to clear — they are documentation for humans only. -
paths/paths-ignoreare INERT for the built C/C++ analysis. GitHub limits them to interpreted languages and to compiled languages analysed without building. The CodeQL job builds with meson + ninja, so every TU it compiles is extracted.core/testis inpaths-ignoreand still produces alerts. Scope for C/C++ is controlled by what the build compiles, or byquery-filterson rule ids. -
A bare directory name in
paths-ignorematches only at the top level.builddoes not matchcore/build;**/builddoes.**must be its own path segment (**foois invalid), and?,+,[,],!are matched literally. -
Follow an "unused variable" note upstream before deleting the name. The
py/unused-global-variablealert on_VALID_AOM_CTCSturned out to be one visible symptom of a duplicated constant block: five further_VALID_*names were defined twice inmcp-server/.../server.py, the second silently shadowing the first. The query cannot flag those because the name is used — just not the first binding.
ADR-1224 — CUDA Tile not adopted; audit findings banked (2026-09-07)¶
Branch: fix/cuda-audit-followups.
-
int64warp reductions must reassemble before adding. CUDA has no 64-bit shuffle, so each step shuffles two 32-bit halves — but summing the halves as independentint32accumulators and recombining at the end drops the carry out of the low half and can overflow a signedint32(UB).warp_reduce(int64_t)incore/src/cuda/cuda_helper.cuhdoes it correctly; use it rather than hand-rolling a second copy, which is howinteger_ssim_score.cuacquired the bug. -
nvcc --threadsis output-neutral,--split-compileis not. Measured:-t 4vs-t 1gives 22/22 byte-identical fatbins.--split-compileproduced three different fatbin hashes across three identical invocations — never adopt it while the release story depends on reproducible builds for keyless Sigstore/SLSA signing. -
Do not re-litigate CUDA Tile from the weak arguments. The compute-capability floor and the CUDA CI pin are NOT reasons against it (ADR-1223 removes both),
ct::mma()does accept float/double, and NVIDIA never said scheduling is "not user-controllable". The surviving objections are: no contraction to accelerate (and a rounding shift inside the ADM accumulation, which is not a sum of products at any dtype), power-of-two tile extents vs 576/1920-wide rows, no reduction-order guarantee, and--fatbin/--tilefatbinnot being composable with--tilefatbinemitting no PTX.
ADR-1228 — upstream A/B performance milestone (2026-09-07)¶
Files touched: testdata/bench_upstream_ab.py (new), docs/benchmarks.md, docs/state.md, docs/adr/1228-*.
Rebase impact: none against upstream Netflix — all fork-local. The harness builds upstream but does not vendor any of it.
-
The upstream ref is pinned on purpose.
DEFAULT_UPSTREAM_REFtracks a release tag, notmaster. Pointing it at a moving branch makes the speedup column measure upstream's churn as much as the fork's work, and a recorded table stops being comparable to itself. Bump the pin deliberately, and re-measure the table in the same PR. -
--max-score-deltais a ratchet, not a tolerance. Its default1e-5sits above the1e-6floor the%.6foutput format imposes and above the known ~5e-6 divergence tracked asT-UPSTREAM-AB-SCORE-DELTA-2026-09-07. Do not raise it to make a run pass; lower it as the delta is localised. -
The comparison is CPU-only by design. Upstream has no SYCL/HIP/Metal backend and a different CUDA feature set, so adding a GPU cell here would measure hardware rather than work. Per-backend numbers belong in
testdata/bench_backends.py. -
The tracked fixtures cannot produce a meaningful speedup. They all run in well under
MIN_USEFUL_SECONDS, so the ratio is startup-dominated and sits near 1.00x whatever the kernels do. The harness warns; do not silence the warning by lowering the threshold.
ADR-1234 — local preflight gate (2026-09-07)¶
-
scripts/dev/preflight.shmust stay in step with the required-check list in.github/workflows/required-aggregator.yml. A new required compiler context without a matching stage recreates the gap this closed: green locally, red in CI, one round-trip per discovery on a single-active-PR queue. -
The
m32stage's-I build/srcis load-bearing, not decoration. It supplies the generatedconfig.h. Drop it and every file aborts at its first#include, the stage reports success, and it silently checks nothing. That is not hypothetical — the first version of the stage passed against a deliberately planted 32-bit break for exactly this reason. -
changed_sources()deliberately includes the working tree, not onlyorigin/master...HEAD. Narrowing it to committed changes makes the script useless for its main job. -
The sanitizer stage's
-Db_lundef=falseis required, not optional. Clang puts the sanitizer runtime in executables, not shared libraries, so without itlibvmaf.sofails to link on undefined__asan_report_*and the stage reports failure on every branch — including ones with no code changes.fuzz.ymlpairs the same two options. -
A missing toolchain skips, it does not fail. Do not "fix" that into a hard error: the script is meant to run on partially provisioned machines, and a stage that fails for want of
gcc-multilibtrains people to ignore the output.
refactor/nolint-citations-cpu-lane — CPU-lane NOLINT citation sweep (2026-09-05)¶
Upstream-mirror files carrying inline NOLINT citations — ADR-0141 §2 / ADR-0278¶
The CPU-lane NOLINT-citation sweep (epic #1237) touched the following upstream-mirror or vendored files. Every hunk is comment-only — the appended ADR-NNNN reference lives inside the existing // NOLINT… / /* … */ comment, so a git am of an upstream commit conflicts only when upstream itself rewrites the same comment line:
core/src/feature/integer_ssim.c— twoNOLINTBEGIN(clang-analyzer-security.ArrayBound)brackets around the upstream kernel-offset clamps.core/src/feature/iqa/ssim_tools.c—readability-non-const-parameteron thepthread_onceguard.core/src/feature/offset.c—misc-use-internal-linkageonoffset_image_s.core/src/feature/third_party/xiph/psnr_hvs.c— the XiphNOLINTBEGINblock gained one preceding citation line; theNOLINTBEGIN(...)list itself is byte-for-byte unchanged.core/src/feature/x86/vif_avx2.c,core/src/feature/x86/vif_avx512.c— theclang-analyzer-deadcode.DeadStorescomments on the upstream-verbatim accumulator zero-init chains.core/src/libvmaf.c— the glibc__libc_single_threadedweak-symbol comment.core/src/picture.c,core/src/read_json_model.c,core/src/output.cpp,core/src/mem.cpp,core/src/pdjson.c(vendored, Unlicense) — trailing comment text only.core/tools/yuv_input.c— trailing comment text only.
Two files changed code layout rather than only comment text, because the lengthened trailing comment would otherwise push clang-format past the 100-column budget: core/src/dnn/model_loader.c (buf[sz] = '\0'; reverts to one line with a preceding NOLINTNEXTLINE) and core/src/ref.cpp / core/src/opt.cpp (same conversion). All three are fork-local files, so no upstream rebase is affected.
Four fork-local SIMD kernels — core/src/feature/x86/psnr_hvs_avx2.c, core/src/feature/x86/ssimulacra2_host_avx2.c and their core/src/feature/arm64/ twins — keep the NOLINTNEXTLINE(...) directive on a single line with a short — ADR-0141 suffix, with the full citation in the block comment above it. Do not "tidy" that by wrapping the justification onto a second // line: the directive then applies to the comment instead of the function, readability-function-size comes back, and the ratchet's citation scan still reports the marker as cited, so nothing but a clang-tidy run against the merge base catches it.
No upstream identifier, kernel body, or numeric expression was modified.
convolution_internal.h — the named-float products are load-bearing (2026-09-16)¶
convolution_edge_s, convolution_edge_sq_s and convolution_edge_xy_s each compute their accumulation as
Do not collapse that back into accum += filter[k] * src[...] on a rebase, and do not "tidy" the temporary away. It is not style: these helpers compute the border pixels of the AVX and AVX-512 convolutions, whose interiors use explicit _mm*_mul_ps + _mm*_add_ps — two roundings. As one expression the C is a contraction candidate, clang fuses it into a single-rounding FMA at every optimisation level, and a SIMD run's border pixels then disagree with a scalar run's. That is T-FLOAT-VIF-ISA-DIVERGENCE-2026-09-16: 1.789e-08 on VMAF_feature_vif_scale0_score, a breach of ADR-0891 caught by ADR-1207's test_feature_isa_invariance.
Two practical consequences for future work:
- Run the ISA-invariance test under clang as well as gcc. gcc does not contract at this site, so a gcc-only run cannot observe the defect. That is exactly how it survived until this branch.
- A
-ffp-contract=offcarve-out inmeson.buildis not an equivalent fix. The pinning belongs in the source, where it survives a change of compiler, optimisation level or flag default.vif_tools.chas no per-filec_argshook anyway — it is compiled as part of a plain source list.
If this test ever fails again, bisect the same way rather than guessing: build with -ffp-contract=off globally to confirm contraction is the cause, then re-enable it one translation unit at a time with #pragma clang fp contract(fast) until the failure returns.
core/test/meson.build — the fault-injection harnesses are gated on non-LTO (2026-09-16)¶
fex_vector_alloc_harness and the test_registration_partial_copy block are both conditioned on not get_option('b_lto'). Do not drop that condition on a rebase to "restore coverage in release": the coverage is not there to restore.
GNU --wrap and ELF interposition both redirect references the linker resolves. With b_lto=true the compiler resolves those calls across translation units first, so the interceptor is never reached, the injected allocation succeeds and test_fex_ctx_vector ends in a double free rather than an assertion — SIGSEGV on the runner, free(): double free detected in tcache 2 locally.
Two things that do not work, both tried:
override_options : ['b_lto=false']on the test executables. They consumeextract_all_objects()from LTO-built libraries, so the non-LTO link fails withfile format not recognized.__attribute__((noinline))onvmaf_feature_name_from_options. It keeps the call but LTO still resolves the symbol internally, so--wrapnever applies.
A real fix needs an interception mechanism that survives LTO. Until then the harness runs in debug, in any -Db_lto=false build and in the sanitizer lane.
psnr_hvs scalar reference carries -ffp-contract=off (2026-09-16)¶
third_party/xiph/psnr_hvs.c is compiled in its own libvmaf_psnr_hvs_scalar_static_lib purely so it can carry -ffp-contract=off. Do not fold it back into libvmaf_feature_sources on a rebase, and do not "simplify" the extra static library away.
Both of its SIMD twins already carry that flag — x86_psnr_hvs_avx2_lib and the arm64 arm64_fp_lib. The scalar reference did not, so clang was free to contract a + b * c in it while the kernels it is compared against were not. On x86-64 nothing contracted, because there is no FMA without -mfma, and the paths agreed by accident. On aarch64 FMA is baseline, clang contracted, and ADR-1207's test_feature_isa_invariance reported psnr_hvs host-isa 14.191670308986598 against scalar 14.191669969203765.
The policy lives in meson.build because it is a build policy: the same source has to round the same way whatever compiles it, and a #pragma in the file would be honoured by some compilers and silently ignored by others (icx ignores #pragma STDC FP_CONTRACT OFF at -O3 without -fp-model=precise). That the file happens to be vendored from xiph is not the reason and must not be used as one: ADR-1142 §1 removed the upstream-mirror tier, so this file is held to the same standards as any other. x86 codegen is unaffected: compiled both ways the instruction stream is identical at 889 instructions with zero FMA.
To reproduce the aarch64 side without ARM hardware, cross-build with clang and run under qemu-aarch64-static; a gcc cross-build will not show it, for the same reason gcc does not show the x86 convolution case.
Strict-FP flag order: precise first, contract last (2026-09-16)¶
Every list that turns FP contraction off for the Intel compiler has to be spelled in this order, and a rebase must not "tidy" it:
c_args : ['-mavx', '-mavx2', '-mfma'] + _x86_simd_strict_fp_extra +
['-ffp-contract=off'] + vmaf_cflags_common,
-fp-model=precise implies -ffp-contract=on. Put it after -ffp-contract=off and it silently re-enables contraction, which is the opposite of why the list exists. Measured on the speed_matmul_avx2 scalar tail with icx 2026.0: -mfma -ffp-contract=off gives zero vfmadd, -mfma -ffp-contract=off -fp-model=precise gives nine, -mfma -fp-model=precise -ffp-contract=off gives zero again.
Twelve carve-outs in core/src/meson.build use this pair, and core/test/meson.build builds _simd_strict_fp_args from the same two flags. They have to move together. The SIMD tests compile their own copies of the scalar references with _simd_strict_fp_args; reordering only the kernels puts the two sides of every comparison on different contraction settings, which is what broke test_ssimulacra2_simd the first time the reorder was attempted and is why T-ICX-FP-CONTRACT-FLAG-ORDER-2026-09-07 stood open with a warning against fixing it.
On GCC and Clang both lists are empty, so the reorder is a no-op and the gcc suite is unchanged. Reproduce with CC=icx CXX=icpx meson setup ... -Db_lto=false and meson test; test_feature_isa_invariance is the test that fails before it.
cppcheck's POSIX model treats fdopen as an unconditional dealloc (2026-09-16)¶
Eight close() calls on the failure path of fdopen() carry a cited cppcheck-suppress doubleFree. Do not delete them, and do not "fix" the code they sit on: POSIX leaves the descriptor open when fdopen() fails, so closing it is required. cppcheck's posix.cfg declares <dealloc>fdopen</dealloc> with no notion of the call failing; 2.13.0 reports the close as a second free, 2.21.1 does not, and CI installs the Ubuntu 24.04 archive's 2.13.0.
Three of the eight sites had an unbraced if. The braces are load-bearing for a different gate: a comment between an unbraced if and its statement makes clang-tidy's readability-braces-around-statements fire, so removing them trades a cppcheck finding for a clang-tidy one.
The clang-tidy baseline belongs to the lane's compiler (2026-09-16)¶
scripts/ci/tidy-baseline-cpu.json is measured with gcc-15 and clang-tidy-22, because gcc supplies the system headers clang-tidy parses. Re-measuring it on a developer box with a different gcc produces a baseline CI will reject: core/src/dict.cpp and core/src/feature/feature_collector.cpp each report one extra cert-dcl03-c,misc-static-assert under gcc-16 that gcc-15 does not.
Rebuild the lane's environment instead — gcc-15 from ppa:ubuntu-toolchain-r/test, clang-tidy-22 from apt.llvm.org, meson from PyPI, on ubuntu:24.04 — and run tidy-ratchet.py --write there. xxd must be installed in that image or the eighteen embedded-model translation units are never generated and the measurement comes out 18 TUs short of the lane's 306.
Scalar references must not call libm fmaf() (2026-09-16)¶
core/src/feature/common/fmaf_exact.h exists because fmaf() is not a fused multiply-add everywhere. On glibc, musl and the UCRT it is. On the legacy msvcrt.dll that MSYS2's MINGW64 links — the environment the required Windows MinGW64 lane builds in — it is not, so a scalar reference that spells its single rounding fmaf() rounds twice there while its SIMD twin rounds once. That breaks the ADR-0891 contract on exactly one lane and nowhere else, which is why it survived until ADR-1207's gate drove the public API both ways.
Do not "simplify" vmaf_fmaf_exact() back to fmaf(), and do not replace it with a * b + c: the first is not fused on legacy msvcrt, the second is not fused anywhere unless the compiler contracts it, which -ffp-contract=off exists to prevent. The two call sites are picture_to_linear_rgb in ssimulacra2.c and both passes of ms_ssim_decimate.c; a new scalar reference whose SIMD twin uses _mm256_fmadd_ps or vfmla needs the same treatment.
Reproducing it on Linux needs an msvcrt-based mingw, not just any mingw. Arch's mingw-w64-gcc is UCRT-based and shows no divergence at all; Fedora's mingw64-gcc is msvcrt-based and reproduces CI's numbers to the last digit. Cross-build in a container and run the .exe under wine — --cpumask 48 gives the AVX2 path, --cpumask 56 the scalar one, and the two should agree bit for bit. Leave AVX-512 out of it (--cpumask 48, not 0): under wine ssim_accumulate_avx512 faults, which is a separate matter.
The cppcheck suppressions file has no vendored tier (2026-09-16)¶
.cppcheck-suppressions.txt is down to what ADR-1142 section 5 allows: generated files, third-party test fixtures, and two entries about cppcheck's own mechanics. Do not re-add a path-wide entry for core/src/svm.cpp, the pelorus interop mirror, or any invalidPrintfArgType_sint / invalidPointerCast / duplicateAssignExpression / shiftNegativeLHS line that a rebase drags back in. There is no "upstream code" tier any more; the findings behind those 31 entries are fixed in the source.
unusedFunction is covered by a mechanism, not a suppression. The six unusedFunction: entries are gone because ADR-1246's scripts/ci/cppcheck-public-entrypoints.cfg models the exported roots. If a new VMAF_EXPORT entry point appears, add it there.
Never write a bare # line in that file. cppcheck 2.13 — the version CI installs from the Ubuntu 24.04 archive — strips the leading # and then fails to parse the empty remainder with Failed to add suppression. No id., which aborts the entire run before a single file is checked. Blank lines are fine; comment lines must carry text. 2.21 accepts both, so this only reproduces against the CI version.
libsvm is held to the tree's standards (2026-09-16)¶
core/src/svm.cpp is vendored but not exempt. It now carries default member initialisers on Solver, deleted copy operations on Kernel / SVC_Q / ONE_CLASS_Q / SVR_Q, a virtual Solver::Solve that Solver_NU overrides, and a destructor on SVMModelParser. Re-syncing from upstream libsvm will drop all of that — reapply it, and re-run the Netflix golden gate afterwards, which is the hard invariant ADR-1142 keeps above every lint rule.
Making Solve virtual is safe because every call site constructs a concrete Solver or Solver_NU and calls through the object, so dispatch is static either way. If a future change ever calls Solve through a Solver *, that equivalence stops holding and the nu-SVC path would start running Solver_NU::Solve where it used to run the base version.
AVX-512 vector register pressure is load-bearing on Windows (2026-09-16)¶
ssim_accumulate_avx512 rebuilds its seven __m512 / __m512d broadcast constants inside ssim_accumulate_block_avx512 rather than hoisting them into the enclosing function and passing them down. That reads like a pessimisation and an upstream sync, a refactor, or a reviewer optimising for the obvious will be tempted to hoist them back out. Do not.
Hoisting them puts the function over 32 live vector values, gcc spills five zmm registers, and on Windows that spill is a crash: the MS x64 unwind contract prevents the and $-64, %rsp realignment gcc uses on SysV, so it emits vmovaps %zmm28,0x1a0(%rsp) against a stack the ABI only guarantees to 16 bytes. Three call sites in four take a general-protection fault, which Windows reports as an access violation on a read of 0xFFFFFFFFFFFFFFFF.
Neither the unit tests nor the ADR-1207 ISA-invariance gate will catch a regression here, because GitHub's Windows runners have no AVX-512 and never execute the path. scripts/ci/check-win64-stack-alignment.py on the Windows MinGW64 lane is what catches it — it disassembles the built objects rather than running them. If that gate fires on a function you just touched, the fix is to cut vector register pressure, not to suppress the finding.
See ADR-1254 and Research-2061.
Two memory-safety fixes in code an upstream sync will overwrite (2026-09-16)¶
Both were found by libFuzzer + ASan and both live in files a sync touches.
core/tools/y4m_input.c — y4m_convert_42xpaldv_42xjpeg(). The scratch base is now captured once and each chroma plane restarts from it:
unsigned char *const scratch = (unsigned char *)aux + 2U * (size_t)c_sz;
for (int pli = 1; pli < 3; pli++) {
unsigned char *tmp = scratch;
Upstream does not have this shape, because upstream never advances tmp — it indexes a fixed base as tmp[y * c_w]. The fork extracted y4m_horizontal_filter_row() and made tmp a running pointer, which is what lost the per-plane reuse and overran aux_buf by a whole plane. If a sync re-inlines the upstream loop, this fix becomes unnecessary; if it keeps the fork's helpers, the reset is load-bearing and must survive. Either way, re-run fuzz_y4m_input against core/test/fuzz/y4m_input_corpus/y4m_420paldv_scratch_overrun.y4m afterwards.
core/src/svm.cpp — parse_support_vectors(). Two additions upstream libsvm does not have: a rejection of non-positive feature indices (libsvm indices are 1-based; -1 is the reserved terminator, and accepting it from the file lets a model forge a sentinel and overrun the total_sv-sized pointer array), and a bound tying the number of parsed runs to total_sv. Re-syncing libsvm drops both. This is the same file that already carries the ADR-1142 rework noted above, so treat the whole of svm.cpp as fork-modified and reapply — then re-run fuzz_json_model against core/test/fuzz/json_model_corpus/svm_forged_sv_sentinel.bin and the Netflix golden gate.
ADR-1250 — EUPL-1.2 relicensing of fork-authored code¶
BUG-003 extends ADR-1250 with a root REUSE.toml default and exact provenance overrides. Upstream syncs are therefore not automatically unaffected: a new or renamed inherited file is covered mechanically by the default but would be misclassified as Lusoris EUPL-1.2 work until an override is added. The same risk applies when an outside contributor edits a previously fork-only file or an append-only aggregate. reuse lint can remain 100% green through that error. Re-run the rename-aware audit in docs/research/bug-003-reuse-provenance-audit-2026-09-24.md, update REUSE.toml and scripts/ci/tests/test_reuse_compliance.py together, and require zero provenance residuals in addition to zero missing metadata.
New ffmpeg-patches/ files need a separate source-unit audit against the exact configured FFmpeg release. Preserve every licence and copyright represented by the touched upstream units; never let the root EUPL-1.2 default classify the patch by repository location alone.
Two things to know when replaying upstream changes:
core/src/feature/speed.cand the other upstream mirrors are unchanged by the relicensing. If a future sync adds a new upstream file, it arrives with Netflix's header and the classifier will leave it alone.- A new fork-authored file should carry
SPDX-License-Identifier: EUPL-1.2.scripts/dev/relicense_fork_files.py --checksays so mechanically; it needs theupstream/masterref present locally (git fetch upstream). - The vendored Pelorus files are not relicensed. The tool's
vendored-mirrorveto ([mirrors.pelorus]inrelicense_provenance.toml) leaves all ten byte-identical to their origin, soscripts/sync-pelorus-interop.sh --updatekeeps working; their terms change only if Pelorus changes them.
The same PR brought every file it relicensed inside the lint profile (ADR-1142). Three results outlive it:
vmaf_ort_open_with_fallback()incore/src/dnn/ort_backend.cis now the only home of the int8 → fp32 session retry.vmaf_dnn_session_open()andvmaf_use_tiny_model()both call it; neither may callvmaf_ort_open()on an int8 path directly (core/src/dnn/AGENTS.md). EP selection invmaf_ort_open()reads the AUTO order fromvmaf_ort_internal_auto_ep_order(), so the tabletest_ort_internals.cpins is the one the code uses.core/test/mu_table.his the table runner forrun_testsbodies over seven tests. It is fork-only; an upstream test that arrives with a longrun_testscan keep itsmu_run_testlist and take the function-size finding into the ratchet, or be converted — either is a clean rebase.- Tidy Changed in
.github/workflows/lint-and-format.ymldefines its exclusion list once, asexclude_untidyable(). Add a family there, not in the four trigger branches that used to carry copies of it. Two of its entries are not backend lanes and are easy to mistake for dead weight on a conflict:^\.config/hiss/testdata/is the HISS rule engine's own fixture tree, where the planted defect is the fixture (HISS-08/c/positive/gets.ccallsgets), and^compat/python-vmaf/matlab/is the upstream MATLAB MEX harness, which dies in the preprocessor onmex.hbecause the MATLAB SDK is not on any runner. Neither family has an entry inbuild/compile_commands.json, so the whole-tree ratchet does not measure them either; deleting either line puts hardclang-diagnostic-erroroutput back into the gate. The same reasoning puts-exclude-dir=.config/hiss/testdataon thegosecstep ingo-ci.yml—go build/go vet/go testskip that tree twice over (dot-directory andtestdata), and gosec's own walker honours neither rule.
libvmaf C++ flags and the exported-symbol gate (ADR-0379 follow-up)¶
core/src/svm.cpp(upstream libsvm mirror) changes one line:Solver::SolutionInfo si = {};insvm_train_one(), so GCC can seesiis initialised when no solver case runs. An upstream sync that touchessvm_train_one()keeps the initialiser.core/src/meson.build: every C++ target takescpp_args : vmaf_cppflags_common. A sync that adds a C++ source to the library inherits it; a new C++ target has to pass it, orcheck_exported_symbolsfails on the symbols it leaks.core/test/test_registration_partial_copy.cppinjects its fault through-Wl,--wrap=vmaf_dictionary_copyand links the static archive; it no longer builds wheredefault_libraryisshared.
SIMD test-scaffolding definition guards mirror their run_tests() call sites (T-NO-ASM-SIMD-TEST-WARNINGS-2026-09-18)¶
No rebase impact: every file touched — core/test/test_vif_simd.c, test_ssimulacra2_simd.c, test_speed_simd.c, test_psnr_hvs_simd.c, test_ms_ssim_decimate.c, test_motion_v2_simd.c, test_iqa_convolve.c, test_integer_ssim_simd.c, test_cambi_simd.c, test_cambi.c — is fork-added SIMD parity-test scaffolding (ADR-0125, ADR-0138, ADR-0161, ADR-0245) with no upstream Netflix/vmaf counterpart; an upstream sync never touches these paths.
Worth knowing for the next SIMD path added to any of these files: several scalar-reference / fixture-builder helpers were defined unconditionally while every caller sat under an ISA guard (#if ARCH_X86, #if HAVE_AVX512, #if ARCH_AARCH64, or the #if ARCH_X86 || ARCH_AARCH64 union run_tests() uses to gate the mu_run_test() list) — a -Denable_asm=false build compiled the helpers in with zero callers left, and -Wunused-function / -Wunused-const-variable fired. The fix wraps each helper in the exact union of ISA conditions its callers use, walking the whole dependency chain: a pick_* dispatcher or ref_* scalar reference that is itself only called from one now-guarded test_* function has to move under the same guard, or the warning just relocates one level down (test_ssimulacra2_simd.c needed one guard spanning its full pick/ref/test helper section for exactly this reason). Adding a new SIMD variant under a narrower guard than its test file's existing union will reopen this class of warning on whichever configuration drops out of the new, narrower condition.
core/src/feature/arm64/vif_neon.c is split into helpers (T-NO-ASM-SIMD-TEST-WARNINGS-2026-09-18)¶
Rebase impact. The same branch brings vif_neon.c from 36 clang-tidy findings to zero (ADR-1142, ADR-0141), without a single NOLINT. Upstream's four macro-expanded kernels are now row loops over static FORCE_INLINE helpers: per-plane vertical and horizontal passes that operate on small lane structs (VifU32x4Pair, VifU64x2Quad, ...), with the 16-bit rounding and shifts held in a VifFilterPlan. The NEON_FILTER_* macros are gone. Each helper issues the same intrinsics as the macro code, in the same lane order, and the first filter tap keeps its own *_init helper wherever upstream seeds an accumulator with a plain product instead of a multiply-accumulate. The output is byte-identical to the pre-split kernels and to scalar dispatch. The two statistic kernels carry the same Research-2045 constParameterPointer exception as their x86 twins. The dead i_dst_stride counters that upstream still has were dropped on this branch too.
Do not take upstream's vif_neon.c wholesale on a sync. Re-apply each upstream change onto the helper structure, then re-run test_vif_neon on an AArch64 build and scripts/ci/tidy-ratchet.py --lane cpu --build-dir <aarch64-clang-build> --only core/src/feature/arm64/vif_neon.c, which must stay at zero. No CI lane measures the arm64 tree, so nothing else catches a regression here.
test_ciede_neon.c platform layer and arm64_strict_fp_args (ADR-1260)¶
Rebase impact. The Windows ARM64 MSVC lane is the first to compile the AArch64 tree with cl.exe, and two things in the tree exist only for it.
core/test/test_ciede_neon.c keeps its guard-page over-read probe on POSIX and on Windows through five entry points: probe_page_size, guarded_row_alloc, guarded_row_free, fault_trap_install / fault_trap_restore and run_kernel_guarded. POSIX is mmap + PROT_NONE + sigsetjmp; Windows is VirtualAlloc + PAGE_NOACCESS + SEH __try / __except. Upstream has no such test, so a sync never touches it; a port of another guard-page test must go through the same five entry points rather than reintroduce <sys/mman.h> under #if ARCH_AARCH64. The probe loop is split into probe_outputs_alloc, probe_slack and probe_overread with no goto; keep it that way (HISS-01, HISS-04).
core/src/meson.build builds the float NEON and SVE2 carve-outs with arm64_strict_fp_args: /fp:precise on msvc, -ffp-contract=off elsewhere. Upstream's meson.build has neither the carve-outs nor the variable; when re-applying upstream changes to the AArch64 block, keep the variable and do not put a literal -ffp-contract=off back into those six c_args. The SVE2 cc.compiles() probe is skipped on msvc for the same reason. See Research-2066.
HIP ADM parity tests: no should_fail, textured fixture (T-HIP-ADM-TESTS-STALE-SHOULD-FAIL-2026-09-18)¶
Rebase impact: fork-local tests and core/test/meson.build only; no upstream file. test_hip_adm_parity, test_hip_adm_small_border and test_hip_adm_wide_rounding are registered without should_fail; the ADR-1154 staging deferral they cited ended with ADR-1211. Do not bring the marker back on a conflict: meson counts an unexpected pass as a failure. test_adm_small_border.c and test_adm_wide_rounding.c (built for CUDA and HIP from one source) skip VMAF_integer_feature_adm3_score under HAVE_HIP (the HIP twin has no AIM pass), print the failing feature and every compared score, and fill their pictures from luma_sample(), a stateless lowbias32 texture. Keep the geometry (160x96 and 1920x144: scale-3 top <= 0 and 60 warps per row) and keep the texture: on the previous smooth ramp the planted pre-ADR-1167 border defect stayed under the 1e-4 gate. PR #1476 carries the same adm3 guard and marker removal on the port/upstream-2026-09 stack; those hunks are identical and merge clean, the fixture change is separate. The CUDA arms were not run here; run test_cuda_adm_small_border and test_cuda_adm_wide_rounding after any rebase that touches adm_cm.cu.
The HIP scaffold posture reports -ENOSYS, at one of two sites (ADR-1264)¶
enable_hipcc defaults to false, so the ordinary -Denable_hip=true build has no device kernels and every HIP extractor must report -ENOSYS. Two things to preserve:
- A scaffold path returns
-ENOSYSdirectly. It does not call a kernel-submit helper with placeholder arguments first.float_vif_hip.candinteger_psnr_hvs_hip.cused to callvmaf_hip_kernel_submit_pre_launch(&s->lc, s->ctx, NULL, …)and return its error; the NULLrbcheck is that helper's first statement, so they always returned-EINVALand theirreturn -ENOSYSwas dead. Do not reintroduce the call, and do not relax the helper's NULL guard to accommodate it. - A HIP parity test checks
-ENOSYSat BOTH observation points — thevmaf_use_feature()return and thevmaf_read_pictures()return. Which one fires depends on whether the extractor gives up at registration or insideextract();speed_temporal_hipdoes the latter, so a registration-only check fails the test instead of skipping it.core/test/test_hip_speed_temporal_parity.cis the reference shape.
__HIP_PLATFORM_AMD__ comes from hip_deps, not from each source (ADR-1263)¶
core/src/hip/meson.build appends declare_dependency(compile_args: ['-D__HIP_PLATFORM_AMD__=1']) to hip_deps outside the if not hip_runtime_dep.found() block. Two ways to break it:
- Moving it back inside the
if. It used to live there, which left thedependency('hip-lang')branch with no definition — invisible, because the eight host sources each carried their own#defineas well. Those defines are gone now, so a pkg-config ROCm would fail outright. - Re-adding
#define __HIP_PLATFORM_AMD__ 1to a HIP host source. It is a reserved identifier (cert-dcl37-c) and makes the next PR that touches that file responsible for removing it again. A new HIP host file needs nothing:hip_depsalready supplies it.
Globs inside block comments break every GPU build (-Wcomment)¶
Writing a path glob such as core/src/feature/hip/*.c inside a /* ... */ comment opens a nested comment; GCC and clang both warn, and the zero-warning gate fails. Fifteen files had it (14 HIP parity tests, one CUDA ADM test) plus core/src/metal/state_priv.h. When a rebase reintroduces one of these comments, spell the set out in prose ("the .c files under core/src/feature/hip/") instead of restoring the glob. core/src/metal/state_priv.h carries an inline note to that effect.
core/tools/vmaf.cpp — the frame loop reports a status, not just a count (ADR-1262)¶
run_frame_loop() returns FrameLoopResult { frames, exit_code }, and the pair of fetch_picture() results is classified by classify_frame_fetch(). Two things there are load-bearing and an upstream sync will try to undo both, because upstream still has the original shape:
- The error test runs before the end-of-stream test.
fetch_picture()returns1at EOF and-1on a read error, soret1 && ret2is true when both sides fail. Upstream's ordering tests that first and therefore reports two corrupt inputs as a clean end of stream — no diagnostic, exit 0. Netflix/vmaf#1604 records this as known and unfixed upstream, so a conflict here will present the buggy order as "theirs". Keep ours. - The loop's status reaches
main(). Upstream returns only the frame count, which is why every read failure exits 0 there. If a rebase collapsesFrameLoopResultback to anunsigned, exit code 102 stops being reachable andcore/tools/test/test_vmaf_read_error_exit.shfails on case 2.
A legitimately shorter stream must stay exit 0 with its ended before warning; that is a deliberate line, not an oversight. See docs/usage/cli.md §Exit codes.
Upstream reports of 2026-09-19 — recognise them when they land¶
Four pull requests and two issues were sent to Netflix/vmaf on 2026-09-19 (docs/development/known-upstream-bugs.md has the table). When a sync brings any of them in:
- #1603 touches
libvmaf/test/checkasm/, which this fork does not carry — nothing to port. - #1604 changes the direct YUV/y4m readers and
fetch_picture(). The fork needs none of it: chroma geometry is already ceiling-based, the direct-read path is compiled out, and reader errors already map to-1. Do not take upstream's newtest_video_input.c: the fork's counterpart iscore/test/test_video_input_odd_dims.c, written for ceiling chroma. The CLI refuses odd 4:2:0 dimensions for raw input; odd-sized y4m input is read and scored. - #1605 adds the early return the fork's
adm_avx2.chas had since PR #792. A conflict there is two spellings of the same guard; keep the fork's. - #1606 patches a VLA the fork replaced with
ModelArrays(ADR-0809) — nothing to port. - #1607 / #1608 are issues. If upstream chooses to support frames below 17 px rather than refuse them, that is a behaviour decision for the fork too, not a mechanical port: the fork currently refuses with
integer_adm requires width >= 17 and height >= 17.
.clang-tidy HeaderFilterRegex must accept absolute paths (ADR-1265)¶
The filter is (^|/)(core/(include|src|tools|test)|python|ai)/.*\.(h|hpp|hxx|cuh)$. The (^|/) is load-bearing: clang-tidy matches the regex against the absolute path from compile_commands.json, and the previous ^core/… form matched nothing, which is how the ratchet ran for months without a single header finding. If a sync or a tidy-config refresh restores the ^-only anchor, Tidy Ratchet will start reporting every header file's count as N -> 0 and ask to tighten — that is the filter breaking again, not a cleanup.
ADR-1214 — Watson-mode CSF rfactors and the float-ADM option aliases¶
-
adm_csf_scale/adm_csf_diag_scaleare Barten-mode options. Inadm_tools.c::adm_csf_rfactor_sthey are consulted only whenadm_csf_mode == ADM_CSF_MODE_BARTEN; the Watson path is1 / quant_step. Every float-ADM GPU twin must compute its Watson rfactor the same way. Do not reintroduceadm_csf_scale / f1— it looks like a harmless default-1.0 multiplier and silently diverges from the CPU for any other value. -
Option aliases are part of the cross-backend contract. Feature names are derived from
alias+ value (ADR-1183), so a twin whose option table spells an alias differently from the CPU emits a different feature key for the same request.float_adm's aliases arescf/scfd; a sync that brings backcs/cdsreintroduces the split key.
HISS-21 claims require replay evidence (ADR-1274)¶
The HISS catalog under .config/hiss/ is executable evidence, not generated decoration. Preserve the catalog and its positive, negative, and gap fixtures when syncing governance files. make hiss-coverage must pass with Praetor's pinned engine on Linux, macOS, and Windows. A new scanner rule or newly closed gap requires a catalog update and a fixture in the same change; never make the matrix green by dropping the contradictory fixture or removing a required context from .github/workflows/required-aggregator.yml. The four governance contexts live in strictMustReport: absence, skip, or neutral is a failure, not an ADR-0313 path-filter exemption.
Hosted replay validation also proved that the draft-only Scorecard guard runs before its artifacts exist. Preserve the non-draft predicate on the artifact upload in .github/workflows/scorecard-policy.yml; if-no-files-found: error remains mandatory once a real scan starts. Keep that workflow in .github/ci-impact.json's full_patterns so changes exercise its contracts.
Meson 1.12 does not materialise compile_commands.json for the configured Ninja builds used by the native lint gates. Preserve the explicit scripts/ci/write-compile-commands.py call between each build and its analyzer, including before the SYCL custom-command augmentation. The exporter must request exactly c_COMPILER and cpp_COMPILER; an unfiltered ninja -t compdb brings link, generator, and phony commands back into analyzer scope. Failed or partial exports must leave the last valid database intact and fail the lane.
The top-level Makefile must also prepend VIRTUAL_ENV_ABS, never relative .venv/bin, when invoking Meson and Ninja. Meson persists the resolved Ninja name and launches it from the build directory during reconfiguration; restoring the relative recipe prefix makes that launch target core/build/.venv/bin/ninja and prevents the native lint gate from starting.
The test-only C build of core/src/log.c preserves a Clang+C23 branch using __builtin_va_start(args, fmt). Clang 21 does not model the __builtin_c23_va_start emitted by its standard macro and otherwise reports a false uninitialized va_list; GCC and MSVC still use va_start. Preserve the non-reserved VMAF_SRC_LOG_H_ include guard in log.h when porting upstream logging changes.
Cppcheck's installed POSIX model is now corrected by scripts/ci/write_cppcheck_posix_model.py before both local and hosted runs. Pre-2.22 models lack pthread_cond_init and receive its correct schema; 2.22's invalid attributes-pointer non-null marker is removed while the condition object remains non-null. Preserve real-tool positive/negative controls and the validation-before-atomic-replace sequence. An upstream model fix is accepted unchanged; duplicate or malformed shapes still fail closed.
vmaf_framesync_init publishes no partial context. Preserve its staged acquire-mutex, retrieve-mutex, condition, and queue-node initialization; every failure returns the negated pthread error and frees only initialized state. test_framesync_init_failure_impl is a separately compiled object whose four pthread init/destroy symbols are mapped to test wrappers. Do not replace it with a production injection hook, textual .c include, NOLINT, or platform-specific linker interposition.
The PR-body pre-push guard must bound gh pr view because a locked desktop keyring can otherwise hang every push. Preserve the public-page fallback, raw Markdown body extraction, schema validation, confirmed-no-PR-only skip, and fail-closed behavior when both metadata sources are unavailable. Do not map an authentication, network, or markup failure to “no open PR.”
CAMBI strict-clean bounded searches and live helpers (2026-09-21)¶
core/src/feature/cambi.c has no file-local clang-tidy or Cppcheck suppression. Preserve that state during upstream syncs. In particular, retain the 16-step TVI bisection, the UINT16_MAX VLT terminus, and the n/partition span bounds in quick-select; their comparison, pivot, swap, and accumulation order is score-sensitive. Keep read_only_picture_view() as the adapter from the mutable extractor callback ABI to CAMBI's const reads.
All ten cambi_internal.h helper exports are deliberately called by the CPU/reference implementation as well as by optional GPU twins. Do not restore unusedFunction annotations or bypass the wrappers when resolving an upstream conflict. The compact CAMBI_OPTION descriptors are the unchanged public option table and keep the declaration below HISS-04's 60-line boundary.
The 48-frame 576x324 regression fixture produced identical scalar and dispatched CPU JSON before and after this cleanup: mean 0.51441210777008473, normalized SHA-256 2fed4234f9f8c018ac9af0810dbf43a0c7a30765bee00e4e55c23de983a1a523. Re-run core/test/test_cambi.c after any conflict; its unreachable VLT, threshold-extreme TVI, and duplicate/descending quick-select cases pin the termination behavior directly.
CUDA integer ADM negative rounding constant (Research-2076)¶
ADR-0155 still requires the scales 1-3 integer-ADM rounding term to be INT32_MIN; changing it to a positive 2^31 value moves Netflix golden scores. Preserve the direct INT32_MIN spelling in integer_adm/adm_csf.cu and both fused paths in integer_adm/adm_cm.cu. Do not restore the prior 1u << 31 unsigned-to-signed conversion, which emits NVCC diagnostic #68-D, and do not replace it with a warning suppression or a widened type. CUDA 13.4 generated byte-identical ADM-CSF and ADM-CM fatbins for the isolated constant-spelling change. The final touched-file cleanup also decomposes oversized kernels into forced-inline helpers; preserve those boundaries even though they change binary layout, because all four CUDA ADM regression executables remain exactly base-identical at runtime.
Core extractor control-flow cleanup (2026-09-21)¶
y_funque_plus.c::init() and libvmaf.c no longer carry their seven historical HISS baseline findings. Preserve structured reverse-order subsystem cleanup, the CUDA collect-before-submit batch boundary, and the SYCL wait/checksum/collect/submit order when resolving upstream conflicts. The helper boundaries are structural only: extractor selection, pending indices, error propagation, and score output remain unchanged. No new public surface or rebase-sensitive policy was introduced.
SYCL integer VIF warning-clean phases (2026-09-21)¶
Rebase impact: core/src/feature/sycl/integer_vif_sycl.cpp only; no public surface, option, metric, tolerance, golden assertion, or twin algorithm changes. The source is decomposed into bounded vertical-accumulation, horizontal-work-item, and initialization phases so strict clang-tidy and HISS-04 pass without suppressions. Private declarations live in short anonymous-namespace blocks; keep those blocks short because HISS measures the block scope too. Do not reintroduce the five forced-unroll pragmas: oneAPI 2026 cannot honour every request across the supported target list and diagnoses the failure. Preserve fp32 device gain, integer operation order, ceiling downsample stride, and the existing init cleanup points. After conflict resolution, compare both vif_fused=false and vif_fused=true against the pre-rebase object on 8-bit and 10-bit inputs; this change measured zero full-precision delta in all four comparisons.
MCP timeout-drain warning regression (2026-09-21)¶
No rebase impact: this changes only the fork-local Python MCP timeout test. The timeout fake closes the first coroutine when injecting TimeoutError, then actually awaits the second communicate() call so the production post-kill drain lifecycle is exercised. Do not restore the former RuntimeWarning suppression or replace the drain await with a fabricated return value.
Python feature-extractor test HISS cleanup (T-HISS-PYTHON-TESTS-2026-09-21)¶
Python test HISS cleanup (T-HISS-PYTHON-TESTS-2026-09-21)¶
No rebase impact on product behavior: twenty Python and MCP test/harness files only split existing setup, fixture data, CLI argument registration, matrix execution, and assertion blocks into class constants or private helpers. The CLI PTY reader and manual YUV-reader tests also replace while True with bounded loops that preserve the same EOF/error checks. Test names, execution order, fixtures, assertion expressions, numeric constants, expected values, tolerances, and parity-gate CLI/output behavior stay unchanged. An upstream textual conflict may take the upstream test body, then reapply the helper boundaries needed by HISS.
Python process execution and 5PL fitting are warning-free (ADR-1278)¶
compat/python-vmaf/tools/misc.py must not restore upstream's forced global fork start method or its fork-only globals. parallel_map() uses loky, executor callers group equal str(asset) keys for serial evaluation, and FIFO helpers use an explicit spawn context. Preserve ordered results and the duplicate-asset serialization regression when resolving upstream conflicts.
compat/python-vmaf/core/train_test_model.py implements the published 5PL curve with b1 multiplying the sigmoid and uses scipy.special.expit. Upstream currently carries the additive-b1/raw-exp form; accepting it would restore an unidentifiable parameter pair, SciPy covariance warnings, and overflow warnings. Netflix golden assertions remain unchanged.
The warning gate is load-bearing too: root pyproject.toml and python/tox.ini promote warnings to errors, and tox no longer disables the warnings plugin. Do not restore -p no:warnings or add ignore filters when an upstream sync starts warning; fix the emitting code or dependency usage.
Go recommend and corpus orchestrator warning cleanup (2026-09-21)¶
no rebase impact: cmd/vmafx-tune/cmd/recommend.go and pkg/corpus/corpus.go are fork-only Go surfaces. The helper boundaries only enforce the 60-line limit; preserve existing CLI flags/output, corpus row order/schema, the distorted-decode fallback, and explicit source-hash/encode-cleanup errors.
Worktree-safe private-state synchronization (ADR-1280)¶
No Netflix rebase impact: scripts/githooks/state-sync.sh, lefthook.yml, and the hook fixture are fork-only governance tooling. Preserve the regular linked-worktree mirror, common-Git exclusive lock, active-worktree Git identity, cache preservation, symlink refusal, and the rule that only derived STATE.md is copied back to canonical private state.
Local data roots are separated by lifecycle (ADR-1277)¶
Do not restore the retired numbered workspace path during an upstream sync. Local state, bounded cache, evidence, and recovery material use .workingdir/; datasets, extracted media, reusable encodes, and derived feature tables use .corpus/. A mechanical substitution of every legacy path with .workingdir/ is incorrect because it recreates the former mixed-lifecycle tree.
Public Markdown may show these paths in operator commands but must not link into either ignored directory. Durable claims must cite tracked docs, ADRs, research, or manifests. Preserve historically accurate prose in old ADRs and changelog entries, while keeping it non-clickable and non-authoritative.
The same cleanup makes scripts/ci/agent-eligibility-precheck.py and scripts/dev/hw_encoder_corpus.py fail closed. These are fork-only tools with no Netflix merge-conflict surface. Preserve non-zero outcomes for unavailable eligibility evidence and for every failed corpus quality point; explicit offline --skip-* flags remain deliberate operator choices.
Non-CUDA high-signal Cppcheck and touched-HISS cleanup (2026-09-21)¶
Score arithmetic and registration order do not change. The native sources narrow error-variable lifetimes, remove two SpEED QR aliases, make two Xiph loop initializers explicit, and simplify a high-bit-depth rounding branch after its 8-bit early return. The picture-pool test replaces obsolete usleep with the same 100 microsecond POSIX nanosleep (and existing one-millisecond Windows delay).
vmaf.cpp no longer has a file-wide anonymous namespace or cleanup goto spine. Keep its short, reopened anonymous-namespace blocks separate: they provide internal linkage without exceeding the scanner's 60-line block limit. CliRunState plus CliRunGuard own the one teardown order documented in core/tools/AGENTS.md. vmaf_bench.c likewise keeps one post-stage cleanup call in bench_feature() / run_feature_collect() and a dedicated SYCL-profile cleanup owner. A rebase must not restore early returns after those owners acquire resources. The compact SpEED option rows preserve name, alias, default, range and array order. Reapply these ownership/helper boundaries on conflict, then rerun the exhaustive Cppcheck command and touched-file HISS audit recorded in Research-2075. - Praetor engine pin moved from 846da590 to f41e74d8f and .standards-baseline.json re-recorded (1411 -> 959). The pin, the baseline and the managed README block's recorded count are one unit: a rebase that reintroduces the old pin must re-record the baseline with the old engine and restore the old count in README.md (ADR-1249). - Praetor engine pin moved from f41e74d8f to 6c772713a133 and .standards-baseline.json re-recorded with --allow-increase (134 with the old engine -> 378 with the new one on the same tree, 182 recorded on master; every added entry is newly measured by praetor d7a3778, 025bbc6 or 53e7594, see ADR-1351). The same unit covers the regenerated praetor-managed files: the compiled agent context, the README governance block, .paperclip/harness.json with the register.sources digest in .standards.yaml, .config/hiss/coverage.yaml titles and the HISS-09 Go fixture, the devcontainer bootstrap (base image vmafx-dev-mcp kept), .config/agent/hooks/block_evasion.py, the documentation gate (praetor-docs.yml, tools/markdownlint/, tools/figures/, the Makefile docs-lint/docs-figures block, the managed .gitattributes and .gitignore tails), .github/rulesets/main.json (rendered for master from repository.default_branch) and the Renovate packageRules entry that freezes those files. Hand-maintained parts of the unit: the documentation block, repository.default_branch and register.surfaces in .standards.yaml, the --offline probe in scripts/git-hooks/hiss-audit.sh (and the pre-push audit job in lefthook.yml that now calls it, with scripts/git-hooks/test-hiss-audit-offline.py), the !/tools/figures/dist/ re-include in .gitignore, the two figure-engine tables at the end of REUSE.toml, the ^tools/figures/ exclusions of the black, ruff and markdownlint hooks in .pre-commit-config.yaml (praetor#578), the praetor-owned .workingdir2 exception in scripts/ci/check-local-data-contract.sh and its test cases (praetor#641, amends ADR-1277), .claude/settings.json, .codex/hooks.json and .gemini/settings.json deliberately left without the praetorctl hook pre-tool entries adopt proposes, and the AGENTS.md invariant table, which adopt --force would replace with a generic advisory one (take compile-context output only). The README count must equal total_infractions in the baseline (378). On conflict, take this branch's copies and re-run the six standardsctl gate steps plus make docs-lint; never hand-edit the locked files. A rebase that returns to the old pin must restore the old engine on every hook PATH as well: the engines do not read each other's .standards.yaml.
vmafx-mcp tool schemas fail closed (2026-09-21)¶
cmd/vmafx-mcp/tools.go registers every tool through toolRegistrar.add, which is the only place a tool's InputSchema is set. add marshals the schemaObj itself; a marshal failure records the tool name, registers nothing, and makes every later add a no-op, so registerTools returns an error, buildServer discards the half-built server, and the fx provider buildMCPServer fails the graph.
Preserve that chain on conflict. A rebase that restores InputSchema: inside the &mcp.Tool{...} literal — or that collapses buildServer / buildMCPServer back to a single return value — has nowhere to put a marshal failure but a log line, and the only thing left to register is a schema that validates nothing. The permissive {"type":"object"} is indistinguishable over the wire from a healthy tool while accepting every argument map, and the Python-parity tests would not catch it because they only compare the tools that are registered. cmd/vmafx-mcp/tool_schema_test.go pins each link; cmd/vmafx-mcp/AGENTS.md invariants #19 and the buildServer seam carry the same rule.
Upstream MATLAB MEX helpers are extracted, not re-inlined (T-HISS-PY-COMPAT-2026-09-21)¶
On conflict in compat/python-vmaf/matlab/, reapply the static band/parse helpers (reduce_* / expand_* / wrap_* sections, the Extend() reduce/expand halves, the corrDn / upConv / histo / pointOp argument parsers, and the STMAD block-statistics helpers) instead of restoring the inline INPROD macros — every index expression and accumulation order is unchanged, which a stubbed-MEX differential harness confirmed bit-identical over ~16k recorded outputs, and the reflect1 default now uses a bounded copy because HISS-08 bans strcpy().
chore/hiss21-core-test — C test bodies split into helpers (2026-09-21)¶
Upstream-mirrored tests under core/test/ (test.h, test_dict.cpp, test_predict.c, test_model.c, test_feature.cpp, test_cuda_pic_preallocation.c) keep every assertion string, expected value and registered test name; on conflict reapply the helper split rather than restoring the single bodies, and keep the added mu_assert_msg in test.h, the short reopened anonymous-namespace blocks in the C++ tests, and the shared core/test/hip_parity_skip.h.
core/test/test_barten_csf.c is the explicit exception and is not split. Its body is the upstream mu_assert(almost_equal(...)) sequence carried verbatim from Netflix c70debb1, and upstream keeps appending cases to it (c2155d6cd added the 2160p CSF rows). Any reshaping of that sequence turns every later upstream sync of this file into a hand-merge, which is the load-bearing invariant its cited // NOLINTNEXTLINE(readability-function-size) protects (ADR-0141 §2, ADR-0278). On conflict, take upstream's case list verbatim and keep the suppression.
chore/hiss21-core-src-simd¶
core/src/sycl/common.cpp gained five same-TU static helpers (sycl_resolve_device, sycl_log_fp64_note, sycl_profiling_enabled, sycl_queue_props, sycl_enqueue_plane_upload, sycl_shared_frame_release, sycl_any_extractor_wants_graph, sycl_run_compute_phase, sycl_apply_input_barriers, sycl_enqueue_all_phases) to clear HISS-01/HISS-04; sycl_shared_frame_release() is now the single cleanup owner that replaced the fail: label, so a rebase must not reintroduce goto fail or an early return that skips it. The SIMD kernels under core/src/feature/{x86,arm64} were left unsplit on purpose — their per-TU -ffp-contract=off carve-outs and inline horizontal reductions are ADR-0138/ADR-0139 bit-exactness invariants.
chore/hiss21-core-src-root: HISS-21 burn-down removed everygotofromcore/src/picture.c,picture_pool.c,picture_pool.cpp,gpu_picture_pool.cpp,predict.c,read_json_model.cand splitvmaf_picture_pool_fetch,vmaf_mcp_start_udsand the threeinterop/pelorus_interop.centry points intostaticteardown/compute owners; on conflict reapply the owner boundaries documented incore/src/AGENTS.md(free order is the contract, arithmetic was moved statement-for-statement only) rather than restoring the upstream label chains.chore/hiss21-core-src-root(vendored mirror):core/src/interop/pelorus_interop.candcore/src/interop/pelorus_qp_report_csv.care ADR-1113 verbatim mirrors oflibpelorus/src/interop.candsrc/qp_report_csv.catPELORUS_VENDOR_SHA. This branch edited both, soscripts/sync-pelorus-interop.sh <pelorus checkout>reports DRIFT (tracked asT-PELORUS-MIRROR-SOURCE-DRIFT-2026-09-22indocs/state.md). A re-vendor (--update) will overwrite these hunks: carry the splits —validate_pack_args,pack_total_size,pack_write_sections,blob_validate_framing(which publishes the header so the blob is cast once per constness),qp_report_copy_frame_stats,qp_cell_average,qp_fold_blocks_to_cells,csv_parse_finishand the boundedsplit_fields— upstream intoVMAFx/pelorusfirst, then re-vendor and bump the pin; do not re-apply them onto a freshly vendored file.
HISS-21 core/src/feature/ top level (2026-09-21): ciede, feature_collector, feature_dists, feature_lpips, float_moment, float_ms_ssim, float_psnr, float_ssim, motion and pu21 lost their cleanup goto ladders to *_init_unwind / *_append_* static helpers in the same TU — on conflict reapply the helper boundaries rather than restoring the label ladders, and keep every arithmetic expression whole across them (ADR-1253).
chore/hiss21-core-src-hip— HIP host code (core/src/feature/hip/**) replaced itsgotocleanup ladders with cascadingstaticunwind helpers and split oversized init/submit/collect/close functions; on conflict keep the helper boundaries and re-check that each tier still frees the same set in the same order as the upstream-twin CUDA ladder it mirrors. TheVmafOptiontables,g_weights[108]andg_hip_features[]keep the upstream-twin one-entry-per-line layout: an earlier revision of this branch packed them behind// clang-format offto shrink a HISS-04 block finding that the current praetor engine no longer raises for file-scope initialiser tables, and the packing was reverted. Inssimulacra2_hip.c,ss2h_picture_to_linear_rgb(),ss2h_run_scale_gpu()andextract_fex_hip()are split intoss2h_yuv_primaries(),ss2h_upload_xyb(),ss2h_download_blurred()andss2h_downsample_for_next_scale()under ADR-1289, which withdraws the ADR-0141 §2 no-split citations those three used to carry; on conflict keep the split side and keep the invariants the comments now name (the ADR-1205 / ADR-0891fmaf()chain, the eight-launch order insidess2h_run_scale_gpu(), the per-scale order inextract_fex_hip()).ssimulacra2_cuda.candcore/src/feature/ssimulacra2.cstill carry their own no-split citations and must not be split along with it.
core/src/cuda/ and core/src/feature/cuda/ no longer contain any explicit goto statement, and the long init/submit/flush functions are split into static helpers (HISS-01 / HISS-04): each old cleanup label is now a *_unwind helper holding that label's statements verbatim, the three fall-through cascades take an explicit stage argument, and CHECK_CUDA_GOTO plus the labels it targets are unchanged. On conflict, reapply the helper boundaries rather than restoring the labels, and keep every moved arithmetic statement whole — splitting one would let FMA contraction change the score.
The CUDA pin is sixteen literals, not one tag (2026-09-21)¶
build-config.env's CUDA_VERSION is the authority for a release that is spelled out in sixteen places across seven files. A merge that carries one of them forward and not the rest now fails scripts/ci/check-cuda-pin-lockstep.py, which is the point — but it also means a conflict resolution that keeps the incoming nvidia/cuda tag has to move CUDA_VERSION, both Jimver/cuda-toolkit inputs, both $cudaVersion and $cudaMajorMinor literals, both cuda-toolkit-NN-N apt names and the org.opencontainers.image.description label on the CUDA runtime image with it. make cuda-pin-sync derives the last five from CUDA_VERSION; the rest are edits. The gate also fails on a CUDA release literal in a spelling it does not recognise, so a rebase that introduces a new one has to teach both the gate and renovate.json's CUDA manager in the same change (ADR-1285).
Do not fold the CUDA manager into the base-image custom manager on conflict: scripts/ci/tests/test_renovate_file_patterns.py asserts there is exactly one docker-datasource manager without a depNameTemplate, and the CUDA one carries nvidia/cuda precisely so the group rule can name it.
fix/bug-hip-adr0759 — HIP ADM buffer stays by pointer (2026-09-21)¶
Fork-local; no upstream Netflix surface. The four HIP integer ADM __global__ kernels that read AdmBufferHip take const AdmBufferHip *__restrict__ buf_ptr, and AdmStateHip owns a device copy (buf_dev) uploaded once near the end of adm_hip_init_device(), passed as args[0] by address, and freed in close_fex_hip() or during failed init.
The current collector does not carry the old tail-calling adm_hip_unwind_* chain. PR #1507 reconciled ADR-0759 with BUG-092 as straight-line paired acquire/release helpers: a failed upload frees its local allocation before publishing s->buf_dev; the later dictionary failure calls adm_hip_free_buf_dev, adm_hip_free_luma, adm_hip_free_buffers, adm_hip_unload_modules, then adm_hip_destroy_stream, and returns -ENOMEM. Normal close releases modules, buf_dev, luma and backing buffers after stream close. A resolution that resurrects either adm_hip_unwind_* or a fail_buf_dev: label is stale; preserve the straight-line BUG-092 release set and order.
This has already been lost once: ADR-0759 landed it in 31a51afb2 (#101) and 92ea978a4 (#102), a squash of a branch cut from an older base, put the by-value signatures back the same day while leaving the AGENTS.md invariant note in place. If a rebase, a squash of a stale branch, or an upstream-shaped conflict resolution reintroduces AdmBufferHip buf in any of adm_csf_kernel_1_4, i4_adm_csf_kernel_1_4, i4_adm_cm_line_kernel or adm_cm_line_kernel_8, take the pointer side — it is the decided design, not a stylistic preference, and it is worth 320 bytes of kernel arguments per launch plus roughly 310 bytes of per-thread scratch on the scale-0 CM kernel.
The single upload is only correct while nothing writes s->buf after init. Any change that starts mutating a field of s->buf per frame has to re-upload the device copy before the next launch, or add a per-frame refresh; the invariant is stated in core/src/feature/hip/AGENTS.md.
AdmFixedParametersHip (248 bytes) is deliberately still by value — ADR-0759's deferred follow-up. Do not "finish the job" on a rebase without an ADR.
Run python3 core/test/test_hip_adm_buffer_pointer_contract.py after every conflict touching these three sources. The test binds the four kernel signatures to the four &s->buf_dev launch arrays, the one-time HtoD upload, the buffer-free reduce kernel, and both teardown paths. It fails against the known stale collector 92ea978a4, so passing it is evidence rather than a documentation-only grep.
python/vmaf/__init__.py — resolve the compat shim by file location (ADR-1292)¶
The shim must not redirect by deleting itself from sys.modules and re-importing its own name. That strategy is correct only while compat/ precedes python/ on sys.path, and a multiprocessing spawn child inherits a sys.path where it does not: the shim then resolves back to itself and recurses until RecursionError, killing the child before it releases the semaphore its parent is blocked on. The parent's sem.acquire() in Executor._open_workfiles_in_fifo_mode has no timeout, so the run hangs until the CI job is cancelled — 62 minutes of silence on the Ubuntu legs.
Keep the importlib.util.spec_from_file_location(__name__, compat/vmaf/__init__.py, submodule_search_locations=[compat/vmaf]) load and the sys.modules[__name__] = _module assignment before exec_module. An upstream-shaped or "simplify this shim" resolution that restores the four-line re-import reintroduces the hang, and it reintroduces it invisibly: every in-process test still passes, because the parent interpreter only takes the working path.
The paired invariant lives in compat/python-vmaf/core/executor.py: ADR-1278's _MULTIPROCESSING_CONTEXT = multiprocessing.get_context("spawn") is what makes a fresh interpreter import vmaf at all. Reverting it to the interpreter default would also mask the shim bug, by going back to forkserver; do not treat that as a fix for anything.
ai/train/qat.py — QAT prepares through torchao pt2e (ADR-1293)¶
Do not restore torch.ao.quantization.quantize_fx.prepare_qat_fx or get_default_qat_qconfig_mapping("x86"). PyTorch deprecated that API wholesale; with ai/pyproject.toml's filterwarnings = ["error"] the call is a test failure, not a log line, and the API has a published removal. The hook captures with torch.export.export(module, example_inputs, strict=True).module() and prepares with torchao.quantization.pt2e.quantize_pt2e.prepare_qat_pt2e under X86InductorQuantizer.
Three invariants a rebase can quietly undo:
_set_mode()must stay. An exported graph module raisesNotImplementedErroron.train()/.eval();_qat_fine_tuneis called with both a raw Lightning module and the prepared graph, so the dispatch onisinstance(module, torch.fx.GraphModule)is load-bearing. A resolution that putsqat_model.cpu().eval()back fails the smoke test immediately.- Phase 4's ONNX export must not go back to
dynamo=False. Its export target is a fresh fp32 module carrying transferred weights, with no observers, so the quantisation buffers that originally justified the legacy TorchScript exporter are not there — and that exporter now warns on its own. The caller'sdynamic_axesis translated into positionaldynamic_shapesordered byinput_names; restoringdynamic_axesunder the dynamo exporter draws a different warning and fails the same test. torch.export.export_for_trainingdoes not exist in torch 2.14. Guides written against 2.5–2.9 still name it; it folded back intotorch.export.export.
The two-step pipeline ADR-0207 made load-bearing is unchanged: QAT conditions weights, ORT quantize_static emits the QDQ graph, convert_pt2e is never called — for the same reason convert_fx never was.
core/src/feature/hip/integer_adm_hip.c + ssimulacra2_hip.c — unwind ladders (T-HIP-INIT-UNWIND-REPORTS-SUCCESS-2026-09-22, ADR-1296)¶
A conflict resolution on either file's init can silently reinstate a use-after-free. Three things must not come back:
- No unwind tier may be handed
hipSuccess. Both ladders terminate inhip_rc(rc)/ss2h_hip_rc(rc), which maphipSuccessto0, so a tier givenhipSuccessmakes the whole unwind report a successfulinitover a state it has just released.git log -S hipSuccesson these two files finds the three sites that did this. If a resolution reintroducesreturn ss2h_init_unwind_mod_mul(s, hip_rc);ininit_fex_hip, orss2h_load_modules'shipError_t *rc_outout-parameter — whose only purpose was ferrying thathipSuccess— the bug is back. adm_hip_unwind_buf_dev_to_host()must stay deleted. It entered the ladder at the host tier, jumping overd_dis_lumaandd_ref_luma. That skip came from a pre-HISS-01goto fail_hostwhose label sat belowfail_ref_luma:, and ADR-0759 threadedbuf_devthrough it without revisiting it. Becausevmaf_feature_extractor_context_closerejects an uninitialised context,close_fex_hipnever runs after a failedinit, so a skipped tier is a permanent leak. The dictionary path now enters atadm_hip_unwind_buf_dev(), which is the exact reverse of the allocation order.- The
core/src/feature/hip/AGENTS.mdnote that called this intentional must not be restored. An earlier revision recorded "two of them deliberately report thehipSuccessleft by the last successful HIP call … not bugs to fix inside a structural refactor". That sentence is why three separate passes over these lines preserved the defect. It is replaced by two invariants with the opposite sense.
HISS-01 still holds: the ladders stay goto-free, one static helper per former label, each tail-calling the next-earlier tier. The fix changes which tier a failure enters at and what value it returns, not the chain's shape.
core/test/test_hip_adm_init_unwind.c and core/test/test_hip_ssimulacra2_init_unwind.c (fast suite, no GPU needed) fail on any of the three regressions. They compile the extractor TU against a complete stub set, so they also break — loudly, at link time — if either TU gains a HIP runtime call; regenerate the stub list with nm -u on the object.
.github/workflows/required-aggregator.yml — the required array is the whole gate (ADR-1297)¶
Branch protection on this repository requires exactly one status context, Required Checks Aggregator (ruleset 22587111). Everything else is decided by const required = [...] inside that workflow. A rebase or conflict resolution that shortens the array silently shortens the merge gate, and a shorter gate is invisible in a diff review: nothing goes red, no job disappears, no test fails. The 44 checks that reported and blocked nothing before ADR-1297 had been in that state for months precisely because the absence of a name looks like nothing.
Rules for any resolution that touches this file:
- Never drop a name to resolve a conflict. Take the union of both sides and reconcile afterwards. Removing a required context is a decision that needs its own ADR superseding ADR-1297, not a merge artefact.
scripts/ci/check-aggregator-names.shmust print OK and is the only mechanical check on this. It enforces set equality in both directions between the array and the# required-aggregatormarkers in the other workflow files, plus one job per required name. It does not and cannot tell you whether a name that belongs in the gate is missing from both sides at once — that is exactly the failure it read OK through. Re-derive the set from the live check list (gh pr checks <n>) when the workflow set changes, not from this file.- Comments inside the array must not contain apostrophes or single-quoted phrases. The checker extracts names with
'([^']+)'over the whole array block, so a'in prose mis-pairs the quotes and silently corrupts the parsed name set. Write "the ADR-0313 rule", not "ADR-0313's rule". - Three check names are load-bearing renames.
Docs Site Build(docs.yml) andDoxygen Public API(doxygen-public-api.yml) exist so those jobs stop reporting under their bare job idsbuildanddoxygen; a required context calledbuildwould be shared by any future unnamedbuildjob, and the aggregator keeps only the newest run per name (T-CI-MSVC-CUDA-SHARED-CHECK-NAME-2026-09-18).Tidy SYCLlost its "(advisory)" suffix together with itscontinue-on-error. Restoring any of the three old names reopens a masking hole or re-asserts an advisory status that no longer exists. experimental: truemust not come back on themacOS clangormacOS clang+DNNrows oflibvmaf-build-matrix.yml, andcontinue-on-error: ${{ matrix.experimental == true }}must not come back on that job. Measured:continue-on-errordoes not neutralise the check-run conclusion — both legs reportedfailureon PR #1518 with the flag set — so the flag never made them advisory, it only made the workflow claim they were.strictMustReportmembership has a precondition. A name belongs there only if its job has no trigger path filter, no conditional skip, and reports on bothpull_requestandpushto master, because this workflow runs on both. That is whySanitizers ASan+UBSanis required but not strict: itsif:excludespush.
refactor/tuning-domain-dedupe — canonical Go tuning helpers (2026-09-20)¶
No rebase impact on upstream Netflix/vmaf: every touched Go path under cmd/ and pkg/ is fork-local, and the change touches no public C API, header, Meson option, FFmpeg patch, or Netflix golden assertion. Within the fork, pkg/scorebackend remains the sole backend vocabulary, probe implementation, and selector. pkg/corpus/backend.go is compatibility-only; do not regrow backend probing there during a branch rebase. Model CLI selector formatting similarly belongs in pkg/model, not in individual tuning drivers.
fix/speed-nonfinite-score — the SpEED score guard (2026-09-23)¶
core/src/feature/speed.c is an upstream mirror and now diverges from Netflix in three places. All three are in ADR-1301; keep them through any rebase.
-
matrix_qr_decomposition()gained a zero-norm guard. Upstream readsvector_div(vec, vector_norm(vec, size), vec, size);with no check. When a deflated minor leaves that column exactly zero in float the divisor is0.0fand every lane computes0.0f / 0.0f. The fork computes the norm once intovnandcontinues when it is zero, which is whatspeed_internal.c'ssi_householder_qrhas always done. A rebase that takes upstream's line back reintroduces a reachable NaN in the CPU reference only — the GPU twins go throughspeed_internal.cand are unaffected, so the symptom appears as a CPU/GPU parity failure rather than as an obvious crash. -
The chroma emit no longer uses
MIN. Upstream appendsMIN(score_u, s->speed_chroma_max_val)and two more like it. The fork callsspeed_internal_clamp_score(), which refuses a non-finite score before comparing. Taking upstream'sMINback restores the masking: a NaN becomesspeed_chroma_max_val. -
The temporal emit clamps. Upstream appends
scoreraw and never appliesspeed_temporal_max_val, although it declares the option and documents it as clipping. The fork clamps through the same helper, matching every GPU twin. This one does change numbers relative to upstream, but only for a score above the bound.
core/src/feature/speed_internal.c gains speed_internal_clamp_score(). That file is fork-authored (ADR-0964) and has no upstream counterpart, so it carries no rebase risk; the CUDA, HIP and SYCL emit paths that call it are fork-local too.
fix/fifo-bounded-wait — bounded FIFO producer readiness (BUG-090, 2026-09-23)¶
compat/python-vmaf/core/executor.py keeps ADR-1278's explicit spawn context, but no FIFO path may restore the post-warning unconditional sem.acquire() inherited from Netflix #1376. Base Executor and NorefExecutorMixin workfile/procfile paths share these invariants:
- one readiness semaphore and diagnostic pipe per producer, so readiness and failure are attributable to the same child;
- a five-second slow-start warning followed by a 60-second hard ceiling;
- child exit status plus available target traceback propagation when readiness was never signaled (bootstrap errors remain on inherited child stderr); and
- sibling producer termination on startup failure.
ExecutorTest.test_fifo_helpers_surface_child_failure uses real spawn processes and covers all four base/no-reference workfile and procfile variants; test_fifo_helpers_bound_live_child_wait proves the hard ceiling terminates live producers. An upstream resolution that restores a shared semaphore or an unbounded acquire reopens BUG-090. See Research-1292 for the failure model and selected supervisor tradeoffs.
fix/mcp-large-body-stream — aiohttp large-body test payload (2026-09-23)¶
No rebase impact: the behavior change is test-only and confined to the fork-local Python MCP server. Preserve the io.BytesIO payload if the HTTP coverage files are reconciled: aiohttp 3.14.3 warns on raw byte bodies above its large-body threshold, and this repository promotes that ResourceWarning to an exception.
- Research digest: no digest needed: trivial compatibility fix confirmed against the installed aiohttp 3.14.3 payload implementation.
- Decision matrix: no alternatives: aiohttp identifies
io.BytesIOas the streaming payload for this exact case, and filtering the warning would weaken the warnings-as-errors contract. - AGENTS.md invariant note: no rebase-sensitive invariants; no production code or public surface changed.
- Reproducer:
cd mcp-server/vmaf-mcp && .venv/bin/python -m pytest -q tests/test_coverage_round4.py::test_auth_413_on_large_body. - Changelog:
changelog.d/fixed/mcp-aiohttp-large-body-stream.md.
fix/bug048-float-motion-hip-lifecycle — force-zero ownership and flush idempotency (2026-09-24)¶
No Netflix upstream counterpart exists: core/src/feature/hip/float_motion_hip.c and its HIP parity test are fork-local. Preserve two coupled invariants when a fork branch or backend twin is reconciled: the force-zero clone retains a close callback that frees feature_name_dict, and the tail-flush duplicate probe uses the dictionary-resolved feature name rather than the base-name literal. The latter matters whenever a feature parameter changes the collector key.
Re-test on a HIP device with meson test -C <hip-build> test_hip_float_motion_parity --print-errorlogs. This changes no public C surface, Meson option, FFmpeg integration patch, or Netflix golden assertion. See Research-2115.
fix/codeql-cjson-warning-alert — test_cjson denormal runtime probe (2026-09-23)¶
core/test/test_cjson.c is a fork-added test file that exercises vendored cJSON and its fork delta (ADR-0683, ADR-1061). Upstream Netflix/vmaf does not vendor cJSON and has no test_cjson.c.
Rebase impact: None. The file has no upstream counterpart; no rebase conflict with upstream Netflix/vmaf is possible. The change replaces the floating-point zero comparison probe in test_print_number_precision with a semantic assertion that the formatted string output from cJSON_CreateNumber(DBL_TRUE_MIN) matches either valid representation ("0" under DAZ or "4.94065645841247e-324" under IEEE-754), eliminating CodeQL alert 1216 while preserving full precision coverage across compilers.
fix/nonfinite-emit-guards — non-finite scores fail the frame (2026-09-23)¶
Netflix-mirror files float_vif.c, integer_adm.c, float_ssim.c, float_ms_ssim.c, adm.c, float_adm.c and predict.c now reject a non-finite score before any MIN, MAX, dB conversion, ratio or default value can turn it into a finite result. Preserve the finiteness check ahead of the comparison on rebase; moving it after publication restores Issue #1526.
The new adm_score.h seam is load-bearing: its outputs are written only after ADM/AIM and per-scale ratios or the ADM3 blend are finite, while a finite flat-frame 0/0 ratio remains a perfect 1.0; nonzero-over-zero still fails. Keep the raw ADM aggregate check before its precision floor: moving it after the ordered comparison lets negative infinity become zero. piecewise_linear_mapping likewise checks before assigning its old 0.0 default. The production prediction path keeps predict_validate_finite after denormalization and after the polynomial and piecewise stages so it warns once with the frame and value before collector publication; do not move that check into the per-segment loop.
The new nonfinite_score.h seam is also load-bearing. CPU, CUDA, HIP, SYCL and Metal VIF, ADM, SSIM and MS-SSIM hosts validate all enabled values before their first collector write; keep backend twins routed through the shared ratio, SSIM-conversion and finite-set emitters. All four VIF ratios must be finite before scale 0 is published; scale 0 remains unclamped, while scales 1-3 then apply their configured minimum. This includes integer VIF and enabled debug atoms. Preserve validation of every raw MS-SSIM L/C/S atom even when it is not emitted; a zero exponent can otherwise hide NaN. Preserve ADR-1221's deliberate exception: finite perfect SSIM/MS-SSIM in unclipped dB mode reports positive infinity; NaN raw scores, invalid ceilings and non-finite conversions still fail.
SSIMULACRA2 is fork-local, but its host calculation is duplicated across every backend. ssimulacra2_score.h owns the edge sign split and final polynomial mapping for scalar, AVX2, AVX-512, NEON, SVE2, CUDA, HIP, SYCL and Metal. Do not re-inline one backend's old ordered comparisons: both are false for NaN and the old final else returned the perfect 100.0.
The helper extractions reduce the generated HISS baseline from 276 to 267 infractions. Preserve the downward .standards-baseline.json ratchet and the matching managed count in README.md; regenerate with the pinned praetorctl instead of restoring stale line fingerprints during a rebase.
fix/windows-utf8-path-contract-rc1 — Windows UTF-8 path contract (ADR-1182, 2026-09-24)¶
- Internal UTF-8 shims (
core/src/compat/path_utf8.{h,c}) must remain unexported. The UTF-8 opener, canonicalization, metadata, mkdir, and remove functions are internal compatibility helpers; they must not carryVMAF_EXPORTor be declared incore/include/libvmaf/public headers, preserving ADR-0379 ABI stability and satisfyingcheck_exported_symbols.py. output_file_openincore/src/libvmaf.cand fork tool openers. Upstream'soutput_file_openused_open()on Windows, decoding path strings with the ANSI code page. The fork routesoutput_file_open()throughvmaf_open_utf8(), which converts UTF-8 strings to UTF-16 withMultiByteToWideCharand calls_wopen(). Fork-added model loaders (core/src/dnn/model_loader.c,core/src/read_json_model.{c,cpp}) and tools (vmaf.cpp,vmaf_bench.c,vmaf_per_shot.c,vmaf_roi.c,vmaf_vpl.c) similarly route filesystem operations through the compatibility layer. Preserve the DNN_wfullpath/_wstat64preflight and CAMBIheatmaps_pathmkdir/open wiring; widening only the final opener recreates the original failure.- Pelorus interop mirror invariant (ADR-1113).
core/src/interop/pelorus_qp_report_csv.cmust NOT be edited directly to usevmaf_fopen_utf8— it is a verbatim mirror oflibpelorus. Any upstream changes to Pelorus must originate inVMAFx/pelorusand be re-vendored viascripts/sync-pelorus-interop.sh. The original state item remains open until that happens; do not describe the contract as coveringpel_x265_csv_parse()meanwhile.
fix/bug048-sycl-residuals — fp32 SpEED and explicit output captures (2026-09-24)¶
The affected SYCL extractors are fork-local, so there is no upstream Netflix/vmaf hunk to adopt. Preserve two implementation contracts when resolving a stale branch:
- the device-kernel regions before
SpeedChromaSyclStateandSpeedTemporalSyclStatecontain nodouble; the chroma path's compensated two-float covariance is the current successor to the original fp32 patch; - the twins retain role-prefixed
launch_{chroma,temporal}_{indterm,score}names. The generic names generated identical unnamed-kernel symbols across the two translation units, so final linking paired one host capture layout with the other device image and produced NaN entropy on an Intel Arc A380; - float PSNR and integer PSNR retain their
FpsnrOutput/PsnrKernelArgscapture structs, while integer moment aliasesd_sumstoe_sumsbefore submitting the kernel and uses the alias for all four atomics.
python3 core/test/test_sycl_kernel_source_contract.py -v is the portable red-cap: its planted mutations prove the old fp64, ambiguous-kernel-name, and raw-pointer forms fail. test_sycl_speed_singular_parity is the device red-cap and passes on the Arc at the unchanged 1e-4 tolerance. No public surface, FFmpeg patch, score snapshot, or Netflix golden assertion changes. See Research-2090.
fix/bug048-ai-cli-helpers — restore shared AI CLI setup (2026-09-24)¶
Twelve fork-local scripts were migrated to ADR-0680/0681 in d02922fc2 and cc4ea5014, then silently returned to their older bodies in d170ef86a. Preserve the current versions' shared setup through rebases:
- bootstrap with
ai/scripts/_script_bootstrap.py, never a localsys.path.insertblock; - use
make_argument_parserandcollect_cli_argv, and pass the normalized vector to both parsing and run provenance; - pass
include_repo_root=Truewhen the script importsai.*; and - keep
ai/tests/test_ai_cli_helper_restoration.pycovering the exact twelve restored files with AST assertions that are insensitive to formatting.
No rebase impact outside fork-authored tiny-AI scripts. No alternatives: only-one-way restoration of accepted ADR-0680/0681 behavior. The smoke command is python -m pytest ai/tests/test_ai_cli_helper_restoration.py ai/tests/test_eval_report_run_provenance.py ai/tests/test_legacy_eval_report_run_provenance.py ai/tests/test_qat_smoke.py ai/tests/test_ptq_scripts.py ai/tests/test_measure_quant_drop_per_ep.py ai/tests/test_dnn_exporter_run_provenance.py ai/tests/test_ptq_cli_contracts.py -q.
The ADR-1291 reverse-hunk declaration for full commit d170ef86affc8e29bf3d486f36018e129230ae97 is deliberately limited to ai/scripts/measure_quant_drop.py, ai/scripts/ptq_dynamic.py, and ai/scripts/ptq_static.py. Retain it only while rebasing this restoration onto a target that still contains the reverted state. Never widen its paths, commit, or evidence regex. Once the restoration is on master and the finding disappears, remove the entry rather than carrying a dormant suppression.
BUG-048 A10 strict JSON emitters (2026-09-24)¶
The AI and vmaf-tune strict-JSON contracts are fork-local and must survive repository-layout or stale-branch conflict resolution. Preserve these coupled invariants:
aiutils.run_manifest.dumps_manifest_json()andwrite_manifest_json()accept any JSON-like root, recursively map all non-finite floats tonull, and serialize withallow_nan=False;write_manifest_json()remains layered onwrite_text_atomic()so strict serialization does not regress the newer crash-safe write guarantee;- AI evaluation reports, legacy corpus/cache manifests, and report-style stdout emitters use the shared strict helpers rather than bare
json.dumps()/json.dump(); and vmaftune.conformal,vmaftune.auto, andvmaftune.ladderroute artifact output throughjsonio.dumps_strict().
Taking the repository-layout side of a conflict reopens the exact clobber: producer commits d8eaf643c, cce8274bc, 48a7c3e1d, and f04bf0e78 were all present, but 384d97d03 / fedce9889 replaced the live blobs. The red-cap tests inject NaN, positive infinity, and negative infinity and parse with a rejecting parse_constant hook. See Research-2086.
No FFmpeg patch impact: this changes fork-local Python serialization only and does not touch libvmaf public headers, C API, CLI options, or Meson surfaces.
fix/bug048-float-moment-metal — exact high-bit-depth reduction (2026-09-24)¶
float_moment.metal and float_moment_metal.mm are fork-local Metal twins whose exact reduction was added by 0fce64b47 / PR #1029 and silently clobbered by c2a3c7e0f / PR #1067. Preserve this pipeline when reconciling the files: four raw ulong moments → four threadgroup ulong[256] scratch arrays → lane-0 exact sum → eight uint32 lo/hi output planes → host uint64 reconstruction → first/second-moment bit-depth scaling. Do not replay
1029's separate lo/hi simd_sum(uint) calls; they lose the carry out of the¶
low half for 16-bit squares.
The 8-bit and 10-bit parity executables are both load-bearing, while test_metal_float_moment_contract.py keeps the production source contract visible on non-Apple builders. No public header, C API, CLI flag, Meson option, Netflix golden assertion, or FFmpeg integration surface changes.
fix/bug048-test-hardening — restore BUG-048 item A12 test hardening and pythonpath (2026-09-24)¶
No rebase impact: test configuration, test helpers, and test files only; no upstream-shared paths or public C library headers touched.
Commit 384d97d03 clobbered commit 993c0ef81 (#1559). This restoration recovers what current authority still requires while documenting what is now moot: 1. mcp-server/vmaf-mcp/pyproject.toml: restore pythonpath = ["src"] under [tool.pytest.ini_options] so vmaf_mcp is discoverable during test collection without requiring an editable pip install, eliminating ModuleNotFoundError: No module named 'vmaf_mcp'. 2. tools/vmaf-tune/tests/test_adr_0543_backend_enforcement.py: restore _binary_supports_backend_flag() helper and wire it into _resolve_vmaf_binary() so pre-fork system binaries (e.g. /usr/local/bin/vmaf) that lack --backend are skipped rather than causing spurious test failures (exit 255 != 100); update source path check to inspect core/tools/vmaf.cpp. Add dedicated unit tests for flag probing and resolver filtering. 3. PyTorch 2.10 deprecation filters: documented as MOOT. The deprecation warning from torch.onnx.export was fixed at the root cause by PR #1518 (T-TINYAI-TORCH-214-WARNINGS-2026-09-22) by migrating export calls to dynamic_shapes = ({0: "batch"},) with the modern TorchDynamo exporter. Under the repo's strict filterwarnings = ["error"] policy, re-introducing blanket suppression filters or legacy dynamo=False would be an anti-pattern.
- Research digest: no new architectural design required; restoration of tested behaviors and reconciliation with PR #1518 zero-warning policy.
- Decision matrix: straightforward restoration of lost test infrastructure; PyTorch filter suppression rejected in favor of the already-landed root-cause fix (
dynamic_shapes). - AGENTS.md invariant note: no rebase-sensitive invariants impacted. Netflix golden assertions preserved untouched.
- Reproducers:
- Without pythonpath:
pytest -c /dev/null -o testpaths=tests mcp-server/vmaf-mcpfails withModuleNotFoundError: No module named 'vmaf_mcp'. - Without binary flag check:
pytest tools/vmaf-tune/tests/test_adr_0543_backend_enforcement.py -k test_adr_0543_per_feature_pinned_to_inactive_backend_failsfails on systems with upstream/usr/local/bin/vmaf(exit 255 != 100). - Changelog:
changelog.d/fixed/restore-bug048-a12-test-hardening.md.
fix/bug048-feature-collector-duplicate — one collector source (2026-09-24)¶
core/src/feature/feature_collector.cpp is the sole implementation authority. Do not restore feature_collector.c when resolving an upstream or branch conflict. The C++ TU intentionally retains the later C-side mutex coverage for model mount/unmount and metadata registration, the complete mounted-model pointer snapshot used across lock drops, the unlocked destroy traversal, the HISS-01 unwind helpers, and the -EAGAIN not-yet-written contract. The public C ABI is unchanged through feature_collector.h.
test_feature_collector_source_authority is the mechanical guard. The deeper reasoning and the exact-master compile/object evidence are in Research-2100.
agent/sycl-parity-branch-count — SYCL motion parity test branch budget (T-SYCL-RATCHET-TEST-BRANCH-COUNT-2026-09-22, 2026-09-24)¶
No rebase impact on production code: changes are strictly test-only in core/test/test_sycl_motion_add_uv_parity.c and core/test/test_sycl_motion3_parity.c, along with a contract test in scripts/ci/tests/test_tidy_ratchet.py and test invariant docs in core/test/AGENTS.md.
Invariants preserved: - Verbatim assertion messages: every mu_assert message string is preserved byte-for-byte across all phase helpers. - Test execution semantics: exact order of feature setup, frame feeding, EOS handling, score extraction, and context cleanup is preserved. - Branch budget: clang-tidy's readability-function-size BranchThreshold of 15 is strictly respected; each helper and caller stays at <= 12 branches (and <= 9 branches in test_sycl_motion_add_uv_parity.c). - Line budget: all functions remain under the 60-line HISS-04 cap. - Zero baseline debt: scripts/ci/tidy-baseline-sycl.json zero baseline is unmodified; 0 warnings generated.
- Research digest: trivial rationale per ADR-0108 / ADR-0141: bounded test-harness refactoring using existing
mu_assert_msgphase helper idiom; no algorithmic, mathematical, kernel, or production API changes. - Decision matrix: helper-based phase decomposition vs ADR-0141 NOLINT suppression citations: phase decomposition chosen because it cures the root cause without suppressions or baseline expansion, keeping the test clean under whole-codebase standards (ADR-1142).
- AGENTS.md invariant note: documented phase helper pattern for linear test pipelines in
core/test/AGENTS.md. - Reproducer:
python3 scripts/ci/tidy-ratchet.py --build-dir /tmp/build-sycl --only core/test/test_sycl_motion_add_uv_parity.c --only core/test/test_sycl_motion3_parity.c --lane sycl --baseline scripts/ci/tidy-baseline-sycl.jsonand running compiled tests on SYCL device:/tmp/build-sycl/test/test_sycl_motion_add_uv_parity/tmp/build-sycl/test/test_sycl_motion3_parity. - Changelog:
changelog.d/fixed/sycl-motion-parity-branch-count.md.
agent/pre-rc1-next-edbf — Windows CLI Unicode argv closure (2026-09-25)¶
The Windows vmaf and vmafx entry point is deliberately wmain; both targets convert each UTF-16 argument to strict UTF-8 before invoking the platform-neutral parser. Preserve -municode for GNU-style Windows links and do not move conversion after cli_parse(): doing either restores the active ANSI-code-page corruption that ADR-1182's internal wide-path helpers cannot repair. POSIX retains its original narrow main unchanged.
The red-cap is core/tools/test/test_vmaf_windows_utf8_argv.cpp. It launches the built CLI with CreateProcessW, uses accented+CJK reference, distorted, and output filenames, and verifies the exact output path. No public header, C API, CLI option, Netflix golden assertion, or FFmpeg patch surface changes. Research and alternatives are recorded in the 2026-09-25 follow-up section of Research-1182.
fix/gpu-test-serialization — shared-device Meson tests stay exclusive (2026-09-23)¶
core/test/meson.build marks every test in the gpu suite with Meson's is_parallel : false scheduling flag. Meson gives that flag global exclusive semantics: it drains already-running tests before starting the GPU test and starts no other test until the GPU test completes. This is intentional because all backend tests configured on a runner share the same finite accelerator queues and memory; do not replace it with longer timeouts or a caller-side -j1 workaround.
An upstream port or conflict resolution that adds or rewrites a GPU test must preserve both its gpu suite tag and is_parallel : false. Reconfigure each affected backend and run meson test -C <build-dir> --no-rebuild test_gpu_serialization_contract check_gpu_test_serialization; the source guard covers dormant backend registrations, while the checker reads Meson's public meson-info/intro-tests.json metadata and fails with every non-exclusive GPU registration.
agent/bug-ledger-float-vif-cuda — keep vif_skip_scale0 mapped in float_vif_cuda.c (2026-09-25)¶
The float_vif_cuda feature extractor correctly parses and plumbs vif_skip_scale0 to match the CPU twin (float_vif). If upstream adds other missing parameters, they must be ported to the CUDA twin to avoid runtime init() failures when models supply them.
- Research digest: no digest needed: trivial.
- Decision matrix: no alternatives: only-one-way fix.
- AGENTS.md invariant: no rebase-sensitive invariants.
- Reproducer / smoke:
meson test -C build-cuda test_cuda_float_vif_parity. - Changelog:
changelog.d/fixed/float-vif-cuda-skip-scale0.md. - FFmpeg impact: none.
agent/float-adm-bypass-cm-gap — SYCL and HIP float-ADM adm_bypass_cm parity (2026-09-25)¶
Closes T-GAP-FLOAT-ADM-BYPASS-CM-SYCL-HIP-2026-09-07. SYCL (float_adm_sycl.cpp) and HIP (float_adm_hip.c, float_adm_score.hip) float_adm twins now expose adm_bypass_cm (alias bcm, int 0..1, default 0) in their option tables and pass bypass_cm down into their respective contrast masking kernels, achieving full parity with CPU, CUDA, and Metal per ADR-1220. When bypass_cm != 0, the 3x3 contrast-masking threshold calculation in both DLM and AIM CM kernels is bypassed (returning 0.0f). Preserved invariants: - SYCL device code remains strictly fp64-free (float32-only). - Tests test_sycl_float_adm_parity (_large) and test_hip_float_adm_parity (_large) assert adm_bypass_cm=1 parity within 1e-4 against CPU float_adm. - Netflix golden assertions untouched.
agent/fix-metal-ms-ssim-review — Metal MS-SSIM option parity (2026-09-25)¶
float_ms_ssim_metal exposes enable_db, clip_db, and enable_chroma under ADR-1334. Preserve its framework-free option-semantics helper, three-plane provided_features/dispatch entries, per-plane 176x176 pyramid minimum, and validation of every L/C/S atom before the weighted product. The Apple parity test creates one option dictionary per vmaf_use_feature() consumer; sharing one between CPU and Metal is a use-after-free because the API consumes it. YUV400P resolves to one active plane before chroma validation. Ceil subsampling makes 351x351 the exact YUV420P luma minimum for 176x176 chroma; do not replace that boundary or its runtime suggestion with floor division or a claimed 352 minimum.
- Research: Research-2110.
- Reproducer:
meson test -C build --no-rebuild test_metal_ms_ssim_option_semantics test_metal_ms_ssim_options_contract. - Apple device gate:
meson test -C build-metal --no-rebuild test_metal_float_ms_ssim_parity. - No public C ABI, CLI, FFmpeg patch, model, snapshot, dependency, benchmark, tuning, training, or Netflix golden assertion changes.
agent/fix-ffmpeg-input-order-3139 — restore exact AV_LOG_INFO and docs for libvmaf input convention (2026-09-25)¶
The FFmpeg libvmaf, libvmaf_*, and libvmaf_tune filters take distorted (main) on pad 0 and reference on pad 1, opposite the Python runner and standalone vmaf CLI. Passing reference on pad 0 changes the direction of the comparison and silently inflated the measured Netflix-pair score from 76.667830 to 83.782079.
This restores the exact executable AV_LOG_INFO reminder in ffmpeg-patches/0001-libvmaf-add-tiny-model-option.patch, corrects every current user-facing two-input VMAF filter command found under docs/ (including direct numeric pads, hardware options, and named filter graphs), and installs a fail-closed checker plus mutation suite. The Make target is called by lint-sh, pre-commit/pre-push, and both FFmpeg patch-stack workflow jobs. The checker intentionally excludes historical ADR, research, state, and rebase records; one narrowly marked wrong example remains in the operator guide to show the hazard.
- Research digest:
docs/research/0730-ffmpeg-libvmaf-smoke-20260527.md. - Decision matrix: no alternatives: only-one-way fix.
- AGENTS.md invariant: no rebase-sensitive invariants.
- Reproducer / smoke:
make ffmpeg-input-contract. - Changelog:
changelog.d/fixed/ffmpeg-input-order-contract.md. - FFmpeg impact: patch 0001 retains the reminder; verify the complete ordered series with
python3 scripts/ci/ffmpeg_patch_stack.py --check.
C++ placement new and delete symbols (ADR-1337)¶
- Any future upstream C++ targets must continue inheriting
vmaf_cppflags_commonto ensure-fvisibility-inlines-hiddenis applied, preventing_ZnwmPvand_ZdlPvS_from leaking into the public ABI.
agent/fix-ai-script-hygiene-3139 — current script environments (2026-09-25)¶
testdata/bench_all.sh derives its default repository root from its own tracked location and honours VMAF_ROOT as the explicit override. Do not restore a developer-specific checkout fallback when resolving benchmark harness conflicts. The harness covers CPU, CUDA, and SYCL only; do not restore retired backend rows, flags, or operator-facing claims in either the script or the linked core/AGENTS.md invocation table. Keep the harness header aligned with the 1080p checkerboard pair used by Test 2. Record each run's emitted metrics-key counts and warn when a GPU count collapses to CPU, but never freeze backend-specific counts as permanent expectations; they move with extractors. Regenerate scripts/ci/source-adr-citations.json when those source comments change; the removed ADR-0726 references no longer own a harness site.
ai/scripts/collect_gpu_calibration_data.py and its manifest fixtures name only the live CUDA/SYCL backend set after ADR-0726 removed Vulkan. The bounded contract in ai/tests/test_legacy_extractor_manifests.py guards both conditions. Preserve load_frames() validation of the top-level object, frames array, and object-shaped frame entries. No hardware benchmark or training run is part of this correction.
The Python and Go MCP auto-dispatch responses must preserve the concrete top-level backend_used receipt written by the fork's vmaf CLI. Do not restore metric-count backend guessing: CPU/CUDA/SYCL counts move as extractors change. When a legacy or external binary provides no valid receipt, return unknown rather than inventing an identity. Keep each top-level JSON-object guard before annotating a score response. The red caps use dated CPU 15 / CUDA 14 / SYCL 24 payloads only to prove identity is independent of count. The Go direct-cgo path is separate and retains its explicit cpu (direct cgo) receipt.
ADR-0482 — vmaf_pre device-string parity (2026-09-25 restoration)¶
ffmpeg-patches/0002-add-vmaf_pre-filter.patch maps all twelve current VmafDnnDevice values from core/include/libvmaf/dnn.h in parse_device() and lists the same strings in the filter option help. Whenever the enum gains a device value, update both this parser/help pair and the corresponding tiny_device mapping in patch 0001 in the same PR. Unknown strings remain a hard AVERROR(EINVAL); there is no auto-detection fallback for a misspelled or unmapped explicit device.
The patch files are cumulative. Verify this contract by replaying ffmpeg-patches/series.txt with git am --3way against the release pinned by build-config.env (currently n9.0.2), never by applying patch 0002 alone. Patch 0002 depends on source state introduced by earlier entries in that ordered series. This restoration changes documentation only; the twelve-entry implementation from commit 4db777126 remains present.
BUG-048 backend invariant guidance restoration (2026-09-25)¶
The backend-local guidance again records two surviving contracts: CUDA module loads own matching unloads on close/unwind, and HIP/Metal dispatch allowlists use exact extractor and provided-feature strings. Preserve the implementation and its local AGENTS.md rule together on conflict resolution. The global scripts/ci/check-dispatch-registry.sh gate covers extractor-symbol registration; it does not replace the HIP/Metal allowlist review and runtime assertions described by those backend guides.
ADR-1336 — CUDA resources are destroyed in their owning context (2026-09-25)¶
Every feature-owned CUmodule, custom CUstream, and CUevent is released through the owner-context helpers in core/src/cuda/common.c; raw destroy calls must not return to core/src/feature/cuda/*.c. Preserve the exact 19-file, 23-handle module inventory in core/test/test_cuda_module_lifecycle_contract.py, including ADM's four modules and SSIMULACRA2's two. New cuModuleLoadData owners must add their state handle, helper-based close and init-unwind, and inventory entry together.
vmaf_cuda_kernel_lifecycle_init() rolls back every earlier successful create when a later create or final context pop fails. Close paths preserve the first error, continue the remaining destroys, clear only successfully destroyed handles, and restore a foreign caller context. Rebase conflicts must keep that ownership behavior even if upstream changes extractor cleanup structure. Run meson test -C build-cuda-unwind test_cuda_runtime_unwind test_cuda_module_lifecycle_contract --print-errorlogs after resolution.
The public close contract is also consumed by the cumulative FFmpeg patch stack. Patch 0020-libvmaf-honor-retryable-close-ownership.patch treats only exact zero as ownership transfer, retries once, and releases models/backend state only after success. Its dedicated CUDA filter retains an AVHWFramesContext reference through close so the borrowed CUcontext cannot expire during a failed retry. Preserve that reference and the early-return ordering when rebasing vf_libvmaf.c. The full 20-patch series replays to tree 0918464997239e1ed03a47f334fdae2b1221c71e on n9.0.1 and 18f1772438491879abc28028d10bda7ae32cc796 on n9.0.2.
ADR-1125 — Go cgo callers select the fork library explicitly (2026-09-26)¶
pkg/libvmaf/libvmaf.go intentionally contains no #cgo LDFLAGS directive. Do not restore an implicit -L... -lvmaf: if the named in-tree directory is absent, the platform linker continues into system paths and may bind an unrelated or stale libvmaf. Local Make targets and go-ci.yml explicitly use core/build-cpu/src; the server, controller, node, and dev-container builders explicitly select the fork library staged inside their image.
When adding a cgo build caller, add its explicit CGO_LDFLAGS selection and extend scripts/ci/test_go_workflow_contract.py in the same change. A direct developer invocation must set both CGO_LDFLAGS (link time) and the relevant runtime loader path, or use the Make targets. This is fork-only build plumbing with no Netflix source or golden-data impact.
ADR-1338 — Go fix clean-tree gate (2026-09-26)¶
The required go vet + go test workflow keeps its published check name, but its first selected source gate after actions/setup-go is go fix -diff ./.... Preserve that exact non-mutating command, the shared go_checks impact predicate, and its placement before Meson, ONNX Runtime, and libvmaf setup so modernization drift fails cheaply. Keep the local go-fix / go-fix-check Make targets and scripts/ci/test_go_workflow_contract.py synchronized with the workflow.
The baseline sweep intentionally includes the hand-maintained api/vmafx/v1/zz_generated_deepcopy.go. It resolved the overlapping stringsseq / slicescontains suggestion in cmd/vmafx-mcp by selecting strings.SplitSeq, and exitCoder embeds error so Go 1.27's errors.AsType preserves the CLI's negative-control test. Do not weaken the gate by permanently excluding either analyzer when rebasing; resolve any new overlap in reviewed source, then require the full default fixer set to be clean.
CodeQL exact-head ownership and arithmetic contracts (2026-09-26)¶
fex_list_entry owns a by-value VmafFeatureExtractor snapshot so a caller's stack descriptor cannot escape into lazily-created contexts. Preserve that ownership on upstream conflict; refresh only framework-managed CUDA, SYCL, and frame-sync pointers under the pool lock. test_fex_pool_growth mutates the original descriptor after registration as the regression guard.
Do not widen float * float operands in iqa_convolve_vertical_pass before their result is accumulated into double. CodeQL alert 1005 is a false positive against the ADR-0138 bit-exact SIMD contract; the adjacent suppression and bitwise kernel test are intentional.
OpenSSF Best Practices passing badge (2026-09-26)¶
Documentation only. The badge record mirrors the answers stored for project 14549; when either side changes, keep its Submitted column equal to the public readback of https://www.bestpractices.dev/projects/14549.json. The README badge points at https://www.bestpractices.dev/projects/14549/badge. No native/public API, numerical or FFmpeg rebase impact.
Release-candidate changelog rollover (2026-09-26)¶
scripts/release/rollover-changelog-fragments.sh and scripts/release/verify-release-version.sh must accept the same version shape, X.Y.Z or X.Y.Z-rc.N, and extract version markers with the same optional -rc.N group. If either changes alone, a release candidate either cannot be cut or cannot be published. test-rollover-changelog-fragments.sh runs the verifier against an RC cut to hold them together. Checks that pin historical prose to a changelog.d/ fragment or to CHANGELOG.md must resolve it through contract_text() in scripts/ci/check-issue-reference-provenance.py, which follows a cut into CHANGELOG.md and docs/changelog-archive/; a raw read of the fragment breaks the next release cut. No native/public API, numerical or FFmpeg rebase impact.
ADR-1345 — changelog archive large-file exemption (2026-09-27)¶
.pre-commit-config.yaml exempts ^docs/changelog-archive/[^/]+\.md$ from check-added-large-files. The pattern and the rollover's archive path (docs/changelog-archive/X.Y.Z.md in scripts/release/rollover-changelog-fragments.sh) move together; test-rollover-changelog-fragments.sh T18 pins the exact pattern and fails if the two diverge. No native/public API, numerical or FFmpeg rebase impact.
1.0.0-rc.1 release-path fixes (2026-09-27)¶
Release-candidate handling must stay consistent across every place that reads a version. verify-release-version.sh, rollover-changelog-fragments.sh, verify-native-release-artifacts.sh, pep440-version.sh and the draft check in release-please.yml all accept exactly X.Y.Z or X.Y.Z-rc.N, and test-release-please-draft-gate.sh proves parity for the workflow's inline jq regex. Python distribution names use the PEP 440 spelling from pep440-version.sh. The pre-push PR-body hook must keep calling scripts/ci/release-pr-exempt.sh rather than re-implementing it. No test may read a changelog.d/ fragment's contents: every release cut deletes them. No native/public API, numerical or FFmpeg rebase impact.
ADR-1346 — hosted-runner release build in the build-deps stage (2026-09-27)¶
build-artifacts must keep building the tag's build-deps stage through scripts/ci/build-dev-container-stage.sh build-deps and compiling through scripts/release/build-native-release-artifacts.sh under docker run --pull never --network none, with no GITHUB_TOKEN on that step and no job-level concurrency group. Never restore the sycl-arc label, a host compile, a GHCR pull or a libvmaf-build release build. Keep the stage-build script's target allowlist, its per-target secret forwarding, and its freedom from --cache-from/--cache-to and --build-arg; check-dev-container-build-secret.py derives the secret-consuming stages from dev/Containerfile and binds both workflow callers. The Dev Container PR gate must keep its release rehearsal (same script, same docker run, local v<manifest version> tag). Keep the gcc-ar/gcc-nm/gcc-ranlib update-alternatives slaves in build-deps (without them LTO links against libvmaf.a fail there), -Denable_dnn=disabled, CCACHE_DISABLE=1 and the GITHUB_SHA == HEAD check in the release script, and keep verify-native-artifacts on ubuntu-26.04 while the bundle needs glibc 2.43 (T-RELEASE-NATIVE-BUNDLE-RELEASE-TRACK-2026-09-27). check-container-build.sh accepts exactly vmaf-dev-mcp and rejects a symlinked stamp. No native/public API, numerical or FFmpeg rebase impact.
docs/rc-phase-shift — first-release candidate mapping shifted by one (2026-09-28)¶
No upstream source impact: this is fork-only release governance and documentation. ADR-1352 amends the tag mapping in ADR-1341 and supersedes the phase names in the docs/release-sequence-rcs entry above. RC2 (v1.0.0-rc.2) is a stabilisation candidate with the RC1 exit bar; RC3 owns benchmarks, profiling and tuning; RC4 owns the one-shot real retrain. When rebasing release, roadmap, runbook, tester or ledger documents, keep phase numbers equal to tag numbers and never restore "RC2 = benchmarks" or "RC3 = retrain" in forward-looking text. tools/rc1-tester/ keeps its name and its catalog phases (RC1, RC3, RC4); the backlog IDs T-RC2-BENCH-TUNE and T-RC3-MODEL-RETRAIN stay stable (ADR-1303). The Netflix golden assertions are untouched.
Native release CLI RUNPATH $ORIGIN (2026-09-28)¶
scripts/release/build-native-release-artifacts.sh must keep running patchelf --set-rpath '$ORIGIN' on the staged copy of the CLI, never on build/tools/vmaf, and the stage the release compiles in must keep installing the pinned patchelf=${PATCHELF_VERSION}: the release compile runs with --network none and cannot install it. Since ADR-1354 that stage is release-build (Debian 13 0.18.0-1.4), not build-deps; a RELEASE_BUILDER_BASE move to another Debian release moves that pin in the same pull request. verify-native-release-artifacts.sh must keep requiring exactly one DT_RUNPATH entry, $ORIGIN, and no DT_RPATH, and must keep running ldd and vmaf --version without LD_LIBRARY_PATH; setting it hid the v1.0.0-rc.1 build-tree RUNPATH $ORIGIN/../src (T-RELEASE-NATIVE-RUNPATH-BUILD-TREE-2026-09-27). No native/public API, numerical or FFmpeg rebase impact.
ADR-1353 — Helm server workload component selector (2026-09-28)¶
deploy/helm/vmafx/templates/deployment.yaml and statefulset.yaml select app.kubernetes.io/component: server in addition to vmafx.selectorLabels, matching the operator and node templates' per-component selectors. Do not drop it back to the release labels: that selector matched the operator, node and helm test Pods. A workload's spec.selector is immutable, so any further selector change needs its own ADR and a documented delete-and-upgrade step like the one in docs/development/k8s-deployment.md#upgrading-from-100-rc1. Keep scripts/ci/check-helm-selector-isolation.py, its test suite and the Workload selector isolation step in .github/workflows/helm-chart.yml together; the step renders Deployment, StatefulSet and Job workloads with the operator, node and PDBs enabled. No native/public API, numerical or FFmpeg rebase impact.
- Research digest: no digest needed; the decision and alternatives are in ADR-1353.
- Decision matrix: ADR-1353.
- AGENTS.md invariant:
deploy/helm/vmafx/AGENTS.md, "Server component selector". - Reproducer / smoke:
python3 -B -m unittest discover -s scripts/ci/tests -p 'test_check_helm_selector_isolation.py', thenhelm template vmafx deploy/helm/vmafx --set operator.enabled=true --set node.enabled=true > r.yamlandpython3 scripts/ci/check-helm-selector-isolation.py r.yaml. - Changelog:
changelog.d/fixed/helm-server-selector.md.
ADR-1354 — native Linux bundle on the Debian 13 release track (2026-09-28)¶
The native release compile runs in the release-build stage of dev/Containerfile, a separate root FROM ${RELEASE_BUILDER_BASE} (Debian 13), not in build-deps. Keep that stage free of any FROM/COPY --from link to the Ubuntu stages, keep its marker write byte-identical to the one in build-deps (scripts/ci/tests/test-check-container-build.sh requires exactly two writes), keep xxd (built-in models) and patchelf pinned to Debian 13's version, and place the stage before gpu-sdks so the file's last stage stays dev-mcp. build-dev-container-stage.sh allowlists release-build and libvmaf-build only; check-dev-container-build-secret.py STAGE_CALLERS requires supply-chain.yml to build release-build and the Dev Container gate to build libvmaf-build and release-build. The release script adds -Denable_tests=false because Debian 13's GCC 14.2 crashes at random while LTO-linking the unit tests. verify-native-artifacts stays pinned to ubuntu-24.04 (never ubuntu-latest) and keeps the RELEASE_RUNTIME_CC start check, which runs the CLI with no LD_LIBRARY_PATH so it also proves the RUNPATH $ORIGIN from the note above. patchelf for that RUNPATH lives in release-build only; do not reinstall it in build-deps. No native/public API, numerical or FFmpeg rebase impact.
SYCL B580 psnr_hvs crash, tile-halo faults, graph-wait result, VIF size and stride, psnr_hvs gate scaling (2026-09-29)¶
fix/sycl-b580-psnr-hvs-adm-tiny, Research-2123, ADR-1361.
core/src/feature/sycl/integer_psnr_hvs_sycl.cpp: the 8x8 DCT runs in local memory (hvs_fdct8_pass(),hvs_ratios_and_first_pass(),hvs_second_pass()), one 1-D transform per work-item and pass; work-item 0 reads coefficients from local memory and recomputesmask[i][j]per coefficient. Do not bring back per-work-itemint[64]/float[64]arrays orod_bin_fdct8x8()on a private block: IGC 2.41.5 crashes the host compiling that at SIMD32 for Xe2, andtest_sycl_psnr_hvs_parity_simd32forces SIMD32 to catch it on any Intel GPU.core/src/feature/sycl/sycl_tile_index.h: every SYCL tile loader that reflects once (integer_admvertical DWT,integer_vifvertical and fused,integer_motion,integer_motion_v2,float_motion,float_vif) passes the reflected index throughvmaf_sycl_tile_index(). An upstream or twin change to a tile geometry or mirror must keep it; dropping it brings backUR_RESULT_ERROR_DEVICE_LOSTon small frames. It is the identity for every consumed sample, so it never explains a score change.-
core/src/sycl/common.cpp:vmaf_sycl_graph_wait()setsgraph_waited_frameonly afterwait_and_throw()succeeds, so after a device fault every collector's call waits again and fails; the five graph collectors return its result before reading host buffers. Keep both halves when rebasing a collector or the wait. -
core/src/feature/sycl/integer_vif_sycl.cpp:VIF_MIN_DIM(16, from the filter tables,static_asserted) is declared through the ADR-1324context_check/context_fallback_name = "vif"pair and guarded again ininit(); the next scale reads the rd buffer at the ceiling stride it was written with. An upstream VIF sync that touches the scale loop or the filter tables must keep both. scripts/ci/cross_backend_calibration.py:area_tolerance_factor()andmetric_delta()are shared by both gates (ADR-1361). A new psnr_hvs tolerance inFEATURE_TOLERANCEor the calibration table is the 576x324 contract; do not pre-scale it..standards-baseline.jsonwas re-recorded downward, 201 -> 190, from a clean clone with the pinned engine (f41e74d8f), and the README debt line follows. The baseline is keyedfile:line(cordanaLLM/praetor#29), so the include and comment lines added tointeger_adm_sycl.cppmoved its seven pre-existing HISS-04 functions without changing them;collect_fex_syclgrew 125 -> 126 lines for the fail-closed graph wait. The other eleven removed entries were debtmasterhad already paid down. A rebase that moves lines in that file re-records again the same way; never with--allow-increase.
No public API, CLI syntax or FFmpeg patch impact. Scores are bit-identical to the previous kernels wherever those ran, except vif_sycl scales 1-3 on odd-width ladders, which now match the CPU, and frames below 16 pixels that model dispatch now scores with the CPU vif.
ADR-1365 — SYCL PSNR, SSIM and float-motion twins take the CPU option tables (2026-09-29)¶
fix/sycl-twin-option-parity, Research-2127, ADR-1365.
core/src/feature/psnr_score.h(new) holds the integerpsnroption math thatinteger_psnr.ccarried inline:vmaf_psnr_peak(),vmaf_psnr_max(),vmaf_psnr_from_mse()(the ADR-1193uncappedsplit, upstreamMIN/MAXmacro-expanded) andvmaf_psnr_aggregate().integer_psnr.c(upstream-mirror) now calls them; its localMIN/MAXmacros andpsnr_from_mse()are gone. An upstream Netflix hunk tointeger_psnr.c::init,psnr()/psnr_hbd()orflush()that changes the math lands in the header, once, for the CPU and every twin that calls it; a hunk that only touches the SSE loops applies as before.core/src/feature/sycl/integer_psnr_sycl.cpp: full CPU option table, options applied on the host throughpsnr_score.h,flush_fex_sycl()publishesapsnr_*. Do not reintroduce a local copy of the math.core/src/feature/sycl/integer_ssim_sycl.cpp: both SSIM twins keep every product in a named temporary, sum the two variances as a pair and return 1 when numerator and denominator are equal, so identical windows score exactly 1 (enable_db+inf /clip_dbceiling like the CPU). Folding a product back into an expression or restoring the four-term left-to-right variance sum breaks the identical-frame cases oftest_sycl_twin_option_parity.enable_lcsuseslaunch_vert_combine_lcs(); keep it separate from the default kernel.core/src/feature/nonfinite_score.h:vmaf_ssim_max_db()is the SSIMclip_dbceiling for GPU twins; the CPUinteger_ssim.c/float_ssim.ckeep their inline copies for upstream parity.core/src/feature/sycl/float_motion_sycl.cpp: every emitted score goes throughmotion_clip();motion_force_zeroshort-circuitssubmit()/collect().
No public C API, CLI syntax or FFmpeg patch impact. CPU scores are bit-identical; psnr_sycl and default float_motion_sycl output are bit-identical to the previous twins; default integer_ssim_sycl / float_ssim_sycl output moves by at most 1.1e-8 (2.2e-8 on identical frames, now exactly 1).
ADR-1371 — SYCL motion differences the frames before the blur, in one shared kernel (2026-09-29)¶
fix/sycl-motion-tiny-frame-parity, Research-1371, ADR-1371.
core/src/feature/sycl/integer_motion_pipeline_sycl.{h,cpp}(new, fork-only): the one SYCL motion SAD kernel,sum |blur(prev - cur)|with the CPUinteger_motion.crounding (>> bpc, then>> 16), reflect-101 borders, int32 vertical sum up to 15 bits per sample and int64 at 16 (host-selectedsubmit_sad<Acc>). Bothmotion_syclandmotion_v2_syclcallmotion_sycl_pipeline::enqueue_sad(); neither extractor TU defines a kernel (Research-2090). Do not restore the per-frame blur (blur(cur) - blur(prev)) or swap the operands tocur - prev: both change the rounding and failtest_sycl_motion_tiny_frames, which compares with==.core/src/feature/sycl/integer_motion_sycl.cpp: raw luma ping-pongd_raw_y[2](filled by a device copy after the kernel) replaces the int32 blurred ping-pong and the unusedd_blur_tmp;cur_bluris nowcur_slot. Withmotion_add_uv,submit()stages U / V into pinnedh_stage_u/h_stage_vandmotion_pre_graphuploads them on the combined queue; the ADR-1034 primary-queue upload andvmaf_sycl_queue_wait()are gone. An upstream-sync or rebase that reintroducesvmaf_sycl_memcpy_h2d_asyncfor the chroma planes also needs the host wait back; keep the staging.core/src/meson.build:integer_motion_pipeline_sycl.cppjoinssycl_feature_sources.core/test/meson.build:test_sycl_motion_tiny_frames.core/test/test_sycl_motion_add_uv_parity.c: the fixed-point oracle differences the frames before the blur, like the kernel and the CPU.- If upstream Netflix changes
motion_score_pipeline_8/_16ininteger_motion.c, mirror the arithmetic in the pipeline TU (one place for both SYCL motion twins).
No public C API, CLI syntax or FFmpeg patch impact. CPU scores are unchanged; motion_sycl scores move to the CPU's (at most 2.0e-4 on 17x17 frames, 1.3e-5 on the Netflix pair), and its motion_add_uv scores move the same way and match the updated fixed-point oracle exactly; the staging change alone is bit-identical. motion_v2_sycl scores are unchanged.
ADR-1364 — Windows MSVC SYCL device link, Windows test runner exit status (2026-09-29)¶
fix/windows-sycl-native-run, Research-2125, ADR-1364.
core/src/meson.build: undersycl_msvc_device_link(MSVC-syntax toolchain, icpx) the SYCL TUs compile as relocatable device code, andsycl_device_link(icpx -fsycl -fsycl-link, AOT device list,--offload-compress,-fp-model=precise,-fsycl-max-parallel-link-jobs=8) plussycl_device_link_anchor(core/src/sycl/coff_add_anchor.py) producesycl_device_link.obj, appended tosycl_feature_objects.sycl_link_argsis/IGNORE:4078there and-fsycleverywhere else. A sync that touches the SYCL toolchain arguments,sycl_dependencyor the feature object list must keep the MSVC branch whole; the Linux ADR-1360 per-TU codegen is unchanged.core/src/sycl/common.cpp: the/include:vmaf_sycl_device_imagespragma andvmaf_sycl_registered_kernel_count()belong together with the anchor symbol name incore/src/meson.build;test_sycl_kernel_registrationandscripts/ci/tests/test_sycl_aot_command.pycheck all three.scripts/ci/run_meson_test.py: POSIX keeps the ADR-1333 same-processexec; Windows runs Meson as a child and returns its status. Do not collapse the two paths:os.execvpon Windows ends the runner with 0..github/workflows/libvmaf-build-matrix.yml: theWindows MSVC+SYCLleg gained a device-free registration step, so the ADR-1333 runner inventory incore/test/test_meson_secret_env_sanitization.pylists five runner calls for that workflow.scripts/ci/cross_backend_parity_gate.py,scripts/ci/cross_backend_vif_diff.py:FEATURE_METRICSnames the keysvmaf --jsonwrites (cambi, notCambi_feature_cambi_score);core/test/test_parity_gate_metric_names.pychecks every entry againstcore/src/feature/alias.c.- Windows test harness:
core/tools/test/test_vmaf_per_shot_input.cinstalls a no-op CRT invalid-parameter handler around its injected read error,core/test/test_device_target_header_dependencies.pydrives Ninja through a Python stand-in compiler,core/tools/test/test_vmaf_per_shot.shskips its/dev/zeroand FIFO cases for native Windows binaries, andcore/test/dnn/meson.buildrefuses the WSLbash.exelauncher. Keep them platform-neutral when rebasing those tests.
ADR-1368 — oneAPI release image on Debian 13 with pinned Intel packages (2026-09-29)¶
docker/Dockerfile.production-gpu builds builder-oneapi2026 and final-oneapi2026 from ONEAPI_BUILDER / ONEAPI_RUNTIME, which build-config.env sets equal to RELEASE_BUILDER_BASE; the base-image gate now enforces that equality and its distro exemption list is empty. Keep the three installers in both stages: scripts/ci/install-intel-oneapi.sh (compiler or runtime at ONEAPI_APT_VERSION, UMF at ONEAPI_UMF_APT_VERSION, repository key pinned by INTEL_ONEAPI_APT_SIGNER_FINGERPRINT) and scripts/ci/install-intel-ocloc.sh --components build|runtime (NEO at INTEL_NEO_VERSION, Level Zero loader at LEVEL_ZERO_VERSION). Dropping the runtime NEO set brings back the Arc B580 segfault; dropping UMF brings back "No device of requested type available". final-oneapi2025 stays as an alias stage and the publish workflow tags the digest both -oneapi2026 and -oneapi2025 (HISS-14); test-docker-image-runtime-contract.sh rejects losing either. INTEL_UMF_RUNTIME_PACKAGE and the Intel 2025 image pins are gone; do not restore them from an older branch. docker/Dockerfile.node's oneapi-runtime-libs stage runs the same runtime installer and copies from /opt/intel/oneapi/redist/lib. No public API, CLI, numerical or FFmpeg patch impact; the image's scores match its CPU backend within the parity gate.
ADR-1370 — float_ssim_sycl decimates on the device (2026-09-29)¶
feat/sycl-float-ssim-scale, Research-2130, ADR-1370.
core/src/feature/iqa/decimate_dim.h(new, include-free) holdsiqa_decimate_dim(), thew / factor + (w & 1)size rule;iqa/decimate.hincludes it andiqa/decimate.c(upstream-mirror) calls it forsw/sh. An upstream change to that size rule lands in the header, once, for the CPU and the SYCL twin.core/src/feature/sycl/integer_ssim_sycl.cpp:float_ssim_sycluploads raw luma (stage_raw_luma()+ one DMA per plane), andlaunch_decimate()reproducespicture_copy()scaling,ssim.c's1.0f / (scale * scale)box andiqa_filter_pixel()(window offsets,KBND_SYMMETRIC, fp32 product, exact sum as int64 units of 2^-52, one RTE conversion). An upstream Netflix hunk toiqa_filter_pixel(),KBND_SYMMETRIC,ssim_low_pass_alloc(),compute_ssim()'s scale rule orpicture_copy()needs the matching device change in the same PR.check_context_sycl()andconfigure_float_ssim()sharefloat_ssim_geometry_supported(); do not restore thescale == 1test. Frame means go throughfloat_ssim_frame_mean()(fp32 rounding likeiqa_ssim()).pack_integer_plane()moved above the float twin and serves both twins.- Tests:
test_gpu_float_ssim_auto_scale_contract.pypins SYCL separately from the scale-1 CUDA / HIP / Metal twins;test_feature_backend_twin.cexpects SYCL to serve 960x540 and every twin to refuse 100x100 atscale=10;test_vmaf_feature_backend.shcase 5 runsfloat_ssim=scale=2on the SYCL twin.
No public C API, CLI syntax, option table or FFmpeg patch impact (the option help string now matches the CPU's). CPU scores are bit-identical. At scale 1 float_ssim_sycl output is byte-identical to the previous twin apart from the fp32 rounding of the frame means (at most 6e-8); at scale > 1 it now runs on the device instead of the CPU fallback.
ADR-1379 / ADR-1380 — CUDA CAMBI and SpEED run entirely on the device (2026-09-30)¶
perf/cuda-rc3-device-resident, Research-1379, ADR-1379, ADR-1380.
core/src/feature/cuda/integer_cambi_cuda.{c,h}andinteger_cambi/cambi_score.cu(fork-local): the ADR-1357 design on CUDA, twelve kernels, one argument struct each, one 88-byte readback and one wait per frame. The Strategy II host residual (cambi_download_and_preprocess,cambi_upload_and_mask,cambi_filter_and_readback,cambi_submit_scale) is gone; do not restore any of it.core/src/feature/cambi.c/cambi_internal.h(upstream-mirror, additive helpers only): the init window guard moved intovmaf_cambi_check_window_fits_lut()(same code and message), andvmaf_cambi_adjust_window(),vmaf_cambi_mask_index(),vmaf_cambi_resize_source_indices(),vmaf_cambi_contrast_weights()andvmaf_cambi_fixed_topk_mean()expose what both device twins need.integer_cambi_sycl.cppdropped its private copies and calls them. An upstream Netflix change toadjust_window_size(),get_mask_index(), thedecimate_generic_*_and_convert_to_10b()walk,g_contrast_weightsoraverage_topk_elements()lands incambi.conce and reaches both twins; keep the helpers when resolving.core/src/feature/cuda/speed/speed_score.cu,speed/speed_cuda_params.h,speed_cuda_pipeline.{c,h}(new, fork-local): the ADR-1358 chain on CUDA, shared byspeed_chroma_cuda.candspeed_temporal_cuda.c, which now only stage planes and read the result.speed_scorebuilds with--fmad=false(cuda_cu_extra_flagsincore/src/meson.build) and spells every rounding with__f*_rnintrinsics. If upstream changesspeed.c's arithmetic, mirror it inspeed_score.cuandspeed_sycl_pipeline.cppin the same PR.core/src/feature/speed_gpu_common.hnow holds the host/device contract (SpeedGpuGeometry,SpeedGpuFilters,SpeedGpuScoring,SpeedGpuChannelBinding,SpeedGpuFrameResult,SpeedGpuConfig);speed_sycl_pipeline.haliases them.speed_internal_gpu_configure()(speed_internal.c) replacedspeed_sycl_host.cpp's configuration code and serves both backends.speed_internal.hincludes the header, so it reachescore/tools/vmaf.cppthroughfeature_dimensions.h; itstypedef structs sit in a citedNOLINTBEGIN(modernize-use-using)bracket (the ADR-1138 shape), which must stay balanced.core/src/feature/speed_log2_hard_cases.h(new, fork-local): the 48 inputs the fp32-pairspeed_log2()of both device twins rounds the wrong way, with their correctly rounded results;speed_score.cuandspeed_sycl_pipeline.cppboth read it. A change to either twin'sspeed_log2()series invalidates the table: rerun the exhaustive replay of Research-1379 before resolving.- Tests:
core/test/test_cuda_device_resident_contract.py(new, fast suite);test_cuda_cambi_paritycompares every frame with==; the CUDA SpEED parity, singular and smoke tests exit 77 without a device;test_cuda_speed_temporal_parity_1080p(1920x1080) guards thespeed_temporal_cudasolve launch that failed above 256 SpEED blocks.test_cuda_module_lifecycle_contract.pynamesspeed_cuda_pipeline.cas the SpEED module and buffer owner.scripts/ci/tidy-baseline-cuda.jsontightened for the rewritten TUs (scoped write).
No public C API, CLI syntax, option table or FFmpeg patch impact. CPU scores are bit-identical (the cambi.c changes move code into functions). CUDA CAMBI scores move to the CPU's (bit-identical except where the CPU's double top-K sum rounds); CUDA SpEED scores move from within 1e-4 of the CPU to bit-identical with a CPU build that rounds log2f correctly and does not fuse multiply-adds. Measured on an RTX 4090 (ryzen-4090-arc) against an icx build: every CAMBI and SpEED frame of the Netflix 576x324 pair and of BBB 3840x2160 identical to --backend cpu; compute-sanitizer memcheck, racecheck and synccheck clean on the parity tests; one readback and one stream synchronisation per frame (CUPTI count). Re-verify after a rebase that touches these files with the commands of Research-1379 finding 8.
ADR-1377 / ADR-1381 / ADR-1382 — HIP RC3 CPU parity: diff-first motion, tiny-frame guards, CPU options (2026-09-30)¶
fix/hip-rc3-parity, Research-1377, ADR-1377, ADR-1381, ADR-1382.
core/src/feature/hip/integer_motion_sad_hip.{h,c}(new): the only host code that loads and launchesinteger_motion_v2/motion_v2_score.hip, the one HIP motion SAD kernel (sum |blur(prev - cur)|, rounded>> bpcthen>> 16, int32 vertical sum at 8 bits, int64 above).motion_hipandmotion_v2_hipboth callvmaf_hip_motion_sad_submit().integer_motion/motion_score.hip, itsmotion_scoreHSACO target andinteger_motion_hip.hare deleted. Do not restore a per-frame blur or swap the operands tocur - prev: both change the rounding and failtest_hip_motion_tiny_frames(==against the scalar CPU) andtest_hip_kernel_source_contract.py. If upstream Netflix changesmotion_score_pipeline_8/_16ininteger_motion.c, mirror it inmotion_v2_score.hip, once, for both HIP twins.core/src/hip/picture_hip.{h,c}:vmaf_hip_picture_upload_staged()plusvmaf_hip_picture_staging_alloc()/_free(). The motion twins copy the luma into an extractor-owned pinned buffer on the host and enqueue the device copy without a host wait;collect()is their one wait. Keepvmaf_hip_picture_upload()(waiting) for every extractor without a staging buffer: the T-HIP-PAGEABLE-UPLOAD-RACE-2026-09-18 invariant still holds.core/src/hip/kernel_template.cnow definesvmaf_hip_rc_to_errno(), whichcore/src/hip/common.halready declared; the motion, PSNR, SSIM and float-motion twins call it instead of private copies.core/src/feature/hip/hip_tile_index.h(new) andcore/src/feature/hip/integer_adm/adm_dwt2_rows.h(new): the motion tile loads and the integer ADM scale-0 vertical DWT clamp their reflected index into the plane;adm_dwt2.hipandinteger_adm_hip.ctake the scale-0 DWT launch geometry fromadm_dwt2_rows.hand the kernelstatic_asserts it. An upstream or CUDA-mirror hunk toadm_dwt2_load_column()must keepadm_dwt2_source_row(). TheAdmBufferHipby-pointer convention (ADR-0759) is untouched.core/src/feature/hip/integer_vif_hip.c:vif_hip_min_dim()(16),check_context_hip()withcontext_fallback_name = "vif", and aninit()guard before any device work.core/src/feature/hip/integer_psnr_hip.c: the CPU option table, options applied on the host throughpsnr_score.h, newflush_fex_hip()forapsnr_*, andVMAF_FEATURE_EXTRACTOR_TEMPORALlike the CPU.integer_ssim_hip.c/float_ssim_hip.c:enable_db/clip_dbthroughvmaf_ssim_max_db();float_ssim_hipenable_lcsselectscalculate_ssim_hip_vert_combine_lcs.integer_ssim_hipscores an identical window exactly its weight (issim_pixel_term()).float_ssim/ssim_score.hippass 2 is the CPU's per-pixell * c * sin double (ssim_lcs()under#pragma clang fp contract(off),ssim_pixel()), one double partial per block, andfssim_hip_cpu_mean()rounds every frame mean to fp32 likeiqa_ssim(); do not restore the combined formula or an identical-window shortcut.float_motion_hip.c:motion_max_valandfm_hip_motion_clip()on every emitted score;float_motion_score.hiptile loads throughfm_tile_index()(hip_tile_index.h).integer_motion_v2_hip.c: the stored SAD isMIN(score * motion_fps_weight, motion_max_val),motion2_v2folds it unweighted, a one-frame run emits both folds.integer_motion_hip.c:debugdefaults to false andVMAF_integer_feature_motion_sad_scoreis emitted every frame, like the CPUmotion.integer_vif_hip.c: the scaffold-ENOSYScomes before the minimum-size check (ADR-1264).integer_motion_hip.c/float_motion_hip.c: undermotion_force_zeroinit()installssubmit_force_zero()/collect_force_zero()instead of clearingsubmit/collect;motion_hipreleases its device objects there (msh_release_device()) and keepsclose(). An upstream or CUDA-mirror hunk that restoresfex->submit = NULLin eitherinit()takes the twin off the asynchronous path; only the engine'sinit_before_dispatch()(core/src/libvmaf.c, #1637) then stands between it and the frame-0 SIGSEGV of T-HIP-MOTION-FORCE-ZERO-NULL-SUBMIT-2026-09-30.scripts/dev/hip_dispatch_drop_probe.hip(new, not built by Meson): standalone probe for T-HIP-GFX1036-DROPPED-DISPATCHES-2026-10-01. No libvmaf dependency, so no rebase interaction; keep it standalone.scripts/ci/cross_backend_parity_gate.py/cross_backend_vif_diff.py:hipbackend (--hip_device),float_ssim_lcscell, andBACKEND_EXTRACTOR_ALIASESkeyed by the base extractor (float_ms_ssimisinteger_ms_ssim_hip).core/src/meson.build:motion_scoreHSACO target gone, the two new headers join the HIP kernel depfile list;core/src/hip/meson.build:integer_motion_sad_hip.c.core/test/meson.build:test_hip_kernel_source_contract,test_hip_adm_dwt2_rows,test_hip_motion_tiny_frames,test_hip_vif_min_dim,test_hip_twin_option_parity.test_device_target_header_dependencies.pyexpects 21 HIP kernel targets.
No public C API, CLI syntax or FFmpeg patch impact. CPU scores are unchanged. motion_hip scores move to the CPU's (about 1.3e-5 on the Netflix pair); motion_v2_hip, psnr_hip with default options, vif_hip from 16x16 up and integer_adm_hip are unchanged by construction; default float_ssim_hip output moves from the combined formula to the CPU's product form (towards the CPU), motion_v2_hip output changes only with non-default motion_fps_weight / motion_max_val, and motion_hip no longer emits the debug score unless debug=true. Measured on a gfx1036 (2026-10-01): motion_hip identical to the CPU motion, every HIP device test OK; see the HIP backend guide, "Measured on a gfx1036". debug score unless debug=true. None of this was measured on an AMD device in this change. output can move at the fp32 rounding level and identical frames now score exactly 1. None of this was measured on an AMD device in this change.
ADR-1378 / ADR-1384 — HIP CAMBI and SpEED run entirely on the device (2026-09-30)¶
perf/hip-rc3-device-resident, Research-1378, ADR-1378, ADR-1384. Built on fix/hip-rc3-parity (ADR-1377's vmaf_hip_picture_upload_staged() and vmaf_hip_rc_to_errno()).
core/src/feature/hip/integer_cambi_hip.c,integer_cambi_hip.h,integer_cambi/cambi_hip_device.h,integer_cambi/cambi_score.hip: the ADR-1357 pipeline on HIP. No host CAMBI stage and no mid-frame wait; one staged upload and one 88-byte readback per frame. The per-work-item math is the header, whichcore/test/test_hip_cambi_device_math.creplays on the host; the parameter block iscambi_hip_plan(). If upstream Netflix changescambi.c's preprocessing, mask, mode filter,c_value_pixel()or pooling, mirror it in the header (and in the SYCL twin) and rerun the replay.core/src/feature/cambi.c/cambi_internal.h: new shared helpersvmaf_cambi_check_window_fits_lut()(cambi.c's own init guard now calls it),vmaf_cambi_resize_source_indices(),vmaf_cambi_adjust_window(),vmaf_cambi_mask_index(),vmaf_cambi_fixed_topk_mean()withVMAF_CAMBI_TOPK_FIXED_SHIFT, andvmaf_cambi_contrast_weights(). The SYCL twin calls them too. An upstream change toadjust_window_size(),get_mask_index(), the decimate walk org_contrast_weightsmust keep these in step. The CUDA device-resident port adds the same helpers with the same signatures; on a rebase between the two keep one copy.core/src/feature/hip/speed_hip_pipeline.{h,c},speed/speed_hip_device.h,speed/speed_pipeline.hip(new; the oldspeed/speed_score.hipsplit is deleted),speed_chroma_hip.c,speed_temporal_hip.c: the ADR-1358 chain on HIP. The kernel TU must keep-ffp-contract=off -fhip-fp32-correctly-rounded-divide-sqrt(hip_cu_extra_flagsincore/src/meson.build); do not replace plain operators with HIP's__f*_rnintrinsics, which contract or approximate.core/src/feature/speed_gpu_common.h(rewritten: the backend-neutralSpeedGpu*contract) andspeed_internal_gpu_configure()inspeed_internal.c: the one init-time configure the SYCL and HIP twins call;speed_sycl_host.cppdelegates to it andspeed_sycl_pipeline.haliases the types. The CUDA device-resident port carries the identical routine; keep one copy on rebase.- The three
init()s return-ENOSYSfirst in a build without hipcc (ADR-1264), then run the host configure (CAMBI's window guard included) before any device call; the contract test pins that order. - Tests:
test_hip_cambi_device_math,test_hip_speed_device_math(fast suite, no device) andtest_hip_device_resident_contract.py(imports the helpers oftest_hip_kernel_source_contract.py).
ADR-1397 — psnr_hvs_cuda returns the CPU's scores bit for bit (2026-10-01)¶
fix/psnr-hvs-twins-cpu-float-sum, Research-1397, ADR-1397 (amends ADR-1361).
core/src/feature/psnr_hvs_score.{c,h}(new, fork-local): the host tail ofcalc_psnrhvs()for GPU twins.vmaf_psnr_hvs_plane_score()adds a plane's terms into one runningfloatin index order and normalises infloat;vmaf_psnr_hvs_combined_score()andvmaf_psnr_hvs_score_db()are the CPUextract()expressions. Built inlibvmaf_psnr_hvs_scalar_static_lib(core/src/meson.build), so it gets the strict floating-point arguments of the scalar reference. An upstream change to the end ofcalc_psnrhvs()(ret /= pixels,ret /= samplemax * samplemax) or to the combination inextract()ofthird_party/xiph/psnr_hvs.cneeds the same change here.core/src/feature/cuda/integer_psnr_hvs/psnr_hvs_score.cu: the kernel stores the 64 terms of every block instead of their sum (hvs_store_terms). The masking table is built at compile time from the CPU's double product (hvs_mask_value,HVS_TABLES), the threshold takes a double product and square root (hvs_threshold), and the coefficient error is an integerabs(). An upstream change to the per-block arithmetic ofcalc_psnrhvs()(means, variances, masking, the error term or its order) needs the matching kernel change in the same PR, ortest_cuda_psnr_hvs_parityfails.core/src/meson.build:'psnr_hvs_score'joinscuda_cu_extra_flagswith--fmad=false. Keep it when resolving a conflict in that dictionary; without it single frames differ in the last bit.core/src/feature/cuda/integer_psnr_hvs_cuda.{c,h}: the readback isPSNR_HVS_TERMS(64) floats per block (hvs_terms_bytes),args.termsreplacesargs.partials, andreduce_hvs_planes()/append_hvs_scores()call the helpers above. Do not restore a host or kernel sum of partials.scripts/ci/cross_backend_calibration.py:EXACT_TWINS/is_exact_pair();cross_backend_parity_gate.pyandcross_backend_vif_diff.pycompare a CPU-and-listed-twin cell with tolerance 0 at--precision max.FEATURE_TOLERANCE["psnr_hvs"], the calibration rows andarea_tolerance_factor()are unchanged and still governpsnr_hvs_sycl.- Tests:
test_cuda_psnr_hvs_paritycompares all four outputs with==(ramp and noise fixtures, 9 to 12 bits, 4:0:0 / 4:2:2 / 4:4:4, 3840x2160 at 8 and 10 bits);test_psnr_hvs_scoreandtest_psnr_hvs_twin_exact_sum_contract.pyare new and device-free.
No public C API, CLI syntax, option table or FFmpeg patch impact. CPU scores are unchanged (no CPU extractor code is touched). psnr_hvs_cuda scores move to the CPU's: by up to 8.4e-5 dB on the Netflix 576x324 pair and 1.7e-2 dB at 3840x2160. psnr_hvs_hip, psnr_hvs_sycl and the Metal twin are untouched and still sum per block. Measured on an RTX 4090 (zeus): every frame of the Netflix 576x324 pairs (8, 10, 12 bits, 4:2:2), the 1920x1080 checkerboard pairs and BBB 1920x1080 / 3840x2160 (8 and 10 bits) equals --backend cpu at --precision max. Re-verify after a rebase that touches these files with the reproducer of Research-1397.
ADR-1400 — integer_ssim_hip sums small frames in the CPU's raster order (2026-10-01)¶
fix/hip-integer-ssim-tiny-identical, Research-1400, ADR-1400.
core/src/feature/hip/integer_ssim/integer_ssim_score.hip: the per-pixel arithmetic is split intoissim_vert_moments(),issim_factors()andissim_cpu_term();issim_pixel_term()(the ADR-1382 identical-window rule) is now only what the per-block kernelinteger_ssim_vert_combineadds. New kernelinteger_ssim_vert_termswritesissim_cpu_term()and the weight of every pixel aty * width + x.core/src/feature/hip/integer_ssim_hip.c: frames of at mostISSIM_HIP_RASTER_MAX_PIXELS(4096,integer_ssim_hip.h) launchinteger_ssim_vert_termsand read back one pair per pixel;collect()adds the pairs in ascending index order, which is the raster order ofinteger_ssim.c::calc_ssim(). That loop order is load-bearing. Do not move the small-frame sum onto the device (one device thread costs 9 ms a frame at 64x64 on a gfx1036) and do not apply the identical-window rule to it: the CPU is not exactly 1 on every identical small frame.- If upstream Netflix changes the expression in
integer_ssim.c::ssim_reduce_row_range(), mirror it inissim_factors()/issim_cpu_term()(HIP) as well as in the CUDA and SYCL twins. - Tests:
core/test/test_hip_ssim_tiny_frames.c(new, device,==withenable_db), four planted regressions incore/test/test_hip_kernel_source_contract.py.
ADR-1407 — every HIP kernel compiles with hip_strict_fp_args (2026-10-01)¶
fix/hip-fp-contract-off, Research-1407, ADR-1407.
core/src/meson.build:hip_cu_extra_flagsand theper_kernel_flagslookup in the HSACO loop are gone.hip_strict_fp_args, defined once between theVMAF HIP strict FP policymarkers, is on every hipcc kernel command. Do not reintroduce a per-kernel table or a second definition on a rebase:core/test/test_hip_strict_fp_policy.pyfails. A new HIP kernel needs no flag entry.core/test/test_sycl_fp_arith_contract.cis now built twice: astest_sycl_fp_arith_contract(default macros) and astest_hip_fp_arith_contract(-DFP_ARITH_PROBE=vmaf_test_hip_fp_arith -DFP_ARITH_DEVICE="HIP") withcore/test/test_hip_fp_arith_probe.{hip,c}. A change to the operands or the host references changes both tests.core/test/test_hip_device_resident_contract.pyreads the SpEED kernel's two flags fromhip_strict_fp_argsinstead of the removed table entry.- Seven HIP twins' outputs change inside their tolerances (
float_adm,float_vif,float_motion,float_ssim,float_ms_ssim,ciede,psnr_hvs); no fork snapshot undertestdata/is a HIP output. - No Netflix golden-data, public API or FFmpeg patch impact.
perf/sycl-float-vif-no-scratch — scratch-free float_vif SYCL kernels (2026-10-01)¶
On the Arc A380 (dg2-g11, PCI 56a5) under the Linux xe driver, SYCL kernels that use scratch memory (private array allocations or register spills) return wrong values without error (ADR-1395, PR #1660).
core/src/feature/sycl/float_vif_sycl.cpp had two kernels with private memory: - launch_compute<0>: spilled 14080 B of private memory per thread at SIMD-32. - launch_decimate<1>: allocated 2432 B of private memory due to dynamic array indexing into the coefficients array inside unrolled loops, defeating IGC SROA.
Changes: - core/src/feature/sycl/sycl_compat.h: added VmafSyclKernelShape<SG, GRF> using oneAPI 2026 <sycl/ext/intel/experimental/grf_size_properties.hpp>. When GRF is 256, it supplies grf_size<256> via the functor's get(properties_tag), granting 256 registers per hardware thread under icpx. - core/src/feature/sycl/float_vif_sycl.cpp: - Replaced runtime coefficient copying with compile-time VifFilterConstants<SCALE>. - Added #pragma unroll to convolution loops in vertical_vif_moments, horizontal_vif_moments, and decimate_vif_pixel. - Encapsulated kernel bodies into FloatVifComputeKernel<SCALE> and FloatVifDecimateKernel<SCALE>. - Configured FloatVifComputeKernel<0> with VmafSyclKernelShape<32, 256>. Scales 1-3 maintain default 128 GRF for maximum EU thread occupancy.
Verification: - Zero scratch: both JIT and dg2-g11 AOT zeinfo show private_size: 0 and spill_size: 0 for all 7 kernels in the translation unit. - Parity restored on Arc A380 under xe: test_sycl_float_vif_parity and test_sycl_float_vif_parity_large pass. speed_gpu_parity.py --backend sycl --feature float_vif confirms max abs diff vs CPU < 4e-5 on Netflix 576x324 and < 8e-6 on BBB 4K (was up to 0.3540 on master). - 4K runtime improved from 25.61 ms/frame to 19.98 ms/frame (22% speedup, median of 3). - Removes float_vif_sycl from core/src/sycl/scratch_ratchet.txt and core/src/sycl/scratch_check.cpp following PR #1660 merge. - scripts/ci/cross_backend_parity_gate.py and scripts/ci/cross_backend_vif_diff.py: FEATURE_METRICS["motion"] reads default emitted keys integer_motion2 and integer_motion3 (excluding debug-only integer_motion).
ADR-1404 — float_motion_hip emits motion3 and implements every CPU float_motion option (2026-10-01)¶
feat/hip-float-motion-motion3-options, ADR-1404.
core/src/feature/hip/float_motion/float_motion_score.hip: the 8-bit and 16-bit kernels are one template (fm_blur_sad<Sample>()overfm_load_tile(),fm_blur_pixel(),fm_block_reduce()); both entry points takefilter_sizeahead ofcompute_sad. New kernelfloat_motion_hip_scale1_sad(fm_bilinear()mirrorsmotion.c::motion_bilinear_interp()with#pragma clang fp contract(off)). If upstream Netflix changesmotion_blur_plane(),vmaf_image_sad_c()ormotion_scale_bilinear(), mirror it here.core/src/feature/hip/float_motion_hip.c: the option table is the CPU's (float_motion.c), in its order; keep it in step, the order spells the feature names. Per-plane stateFmPlaneHip plane[3];motion3throughfm_hip_motion_blend_clip(), which callsmotion_blend()ofmotion_blend_tools.h. Do not add a second copy of the blend.core/tools/test/test_vmaf_feature_backend.sh: on HIP,float_motion=motion_filter_size=3now runs the twin; the HIP fallback case isfloat_adm=adm_skip_scale0=true. The other backends keep themotion_filter_sizefallback case until their twins implement it.- Tests:
test_hip_twin_option_parity(test_float_motion_motion3,_one_frame,_filter_size,_scale1_and_uv,_refusals), six planted regressions intest_hip_kernel_source_contract.py.
ADR-1405 — float_ssim_hip decimates on the device (2026-10-01)¶
feat/hip-float-ssim-scale, Research-1405, ADR-1405.
core/src/feature/hip/float_ssim/ssim_decimate.h(new): one output ofiqa/decimate.c::iqa_decimate()withssim.c's low-pass kernel, in plain C and HIP C++ (vmaf_hip_ssim_decimate_sample(): fp32sample * tap, int64 sum in units of 2^-52, one rounding;vmaf_hip_ssim_symmetric_index()=KBND_SYMMETRIC). If upstream Netflix changesiqa_filter_pixel(),KBND_SYMMETRIC,iqa_decimate(),iqa_decimate_dim()orssim_low_pass_alloc(), change this header in the same PR;core/test/test_hip_float_ssim_decimate.c(device-free, byte compare againstiqa_decimate()) fails until it follows.core/src/feature/hip/float_ssim/ssim_score.hip: new kernelscalculate_ssim_hip_decimate_{8,16}bpcandcalculate_ssim_hip_horiz_f32; the three pass-1 entry points sharessim_horiz<Sample>().core/src/feature/hip/float_ssim_hip.c:in_width/in_heightare the picture,width/heightthe planes the SSIM passes read (iqa_decimate_dim()above scale 1).check_context_hip()refuses only a decimated plane below 11x11 or a scale above 128.core/test/test_hip_float_ssim_parity.cis rewritten around the SYCL case table (tolerance 5e-5);core/test/test_gpu_float_ssim_auto_scale_contract.pytreats HIP like SYCL;core/tools/test/test_vmaf_feature_backend.shexpects the HIP twin forfloat_ssim=scale=2.- No Netflix golden-data, public API or FFmpeg patch impact.
perf/vmaf-tune-cache-batch — batch vmaf-tune TuneCache index writes via dirty flush¶
- Files changed:
tools/vmaf-tune/src/vmaftune/cache.py,tools/vmaf-tune/src/vmaftune/corpus.py,tools/vmaf-tune/tests/test_cache.py - Rebase impact: pure Python tooling changes in
tools/vmaf-tune/. No upstream-shared C headers or core library impacted.
ADR-1408 — a VmafContext uploads each frame plane once for all HIP twins (2026-10-01)¶
perf/hip-shared-plane-uploads, Research-1408, ADR-1408.
core/src/hip/shared_frame.{c,h}(new):VmafHipSharedFrame, owned by theVmafContext(vmaf->hip.frameincore/src/libvmaf.c), andVmafHipPlaneSource, a twin's handle.VmafFeatureExtractor::hip_frame(underHAVE_HIP, next tocu_state/sycl_state) is set byset_fex_hip_frame()and copied inrefresh_fex_runtime_state().core/src/libvmaf.c: the oldread_pictures_extractor_loop()body is nowread_pictures_dispatch_extractors(); the newread_pictures_extractor_loop()wraps it invmaf_hip_shared_frame_begin()/_end()underHAVE_HIP. An upstream change to the dispatch loop goes intoread_pictures_dispatch_extractors().vmaf_close_backends()destroys the shared frame; it must stay after the extractors are closed.- Thirteen HIP twins (
psnr,float_psnr,float_moment,ciede,integer_ssim,float_ssim,vif,float_vif,adm,float_adm,motion,motion_v2,float_motion) no longer allocate picture staging or callvmaf_hip_picture_upload():submit()callsvmaf_hip_plane_source_acquire[_luma]()andclose()callsvmaf_hip_plane_source_close(). A twin ported or rebased from the CUDA side must keep that shape;core/test/test_hip_shared_frame_contract.pylists the adopted twins (ADOPTED) and fails on an own upload. integer_motion_sad_hip.{c,h}:VmafHipMotionSadFrameis now{cur, keep, have_prev, sad, width, height, bpc}. The motion twins keep oneprev_lumaplane and copy the frame's luma into it device-to-device behind the SAD; the pinned staging plane and thepix[2]ping-pong are gone.integer_vif_hip.c:buf.stride(the raw planes' byte stride) is packed,w * bytes-per-sample, no longer rounded up to 64.core/test/test_hip_adm_init_unwind.cstubs the two plane-source entry pointsinteger_adm_hip.ccalls; a new call from that TU intoshared_frame.cneeds a stub there.- No output changes: every metric of every adopted twin is bit-identical before and after.
perf/ai-k150k-tmpfs-scratch — YUV scratch auto-selected to /dev/shm (Win 3)¶
- Files changed:
ai/scripts/extract_k150k_features.py,ai/tests/test_extract_k150k_perf.py,ai/AGENTS.md - Rebase impact: no rebase impact — fork-local Python script and tests only; no upstream-shared C, headers, or public API touched.
Unflagged submit/collect extractors run on the caller thread (2026-10-01)¶
fix/async-extractor-thread-pool, T-ASYNC-EXTRACTOR-THREAD-POOL-EINVAL-2026-10-01.
core/src/libvmaf.c:read_pictures_should_skip()andbatch_extractor_skip()share one predicate,fex_ctx_runs_on_caller_thread()(backend flags, TEMPORAL, orVmafFeatureExtractorContext::caller_thread_dispatch).batch_extractor_skip()takes the registered context, not the extractor. An upstream change to either skip function has to keep both on the predicate; a flag-only test hands a flagless GPU twin (adm_hip,float_vif_hip) to the worker pool again, where every frame fails with-EINVAL.core/src/feature/feature_extractor.{h,cpp}: the new context field is set invmaf_feature_extractor_context_create()from the descriptor'ssubmit/collect, beforeinit()can swap callbacks.- Guard:
core/test/test_async_extractor_thread_pool.c(mock extractors, no device).
ADR-1401 — psnr_hvs_sycl and psnr_hvs_hip return the CPU's scores bit for bit (2026-10-01)¶
fix/psnr-hvs-sycl-hip-exact-sum, Research-1401, ADR-1401 (implements ADR-1397 for SYCL and HIP).
core/src/feature/sycl/sycl_exact_fp.h:isqrt_floor50()andsqrt_prod_rn(a, b)(new, fork-local): the fp32 rounding of the square root of the exact product, which is what a host returns for(float)sqrt((double)a * (double)b). Integer arithmetic only; the header still must not name the fp64 type.core/src/feature/sycl/integer_psnr_hvs_sycl.cpp: the kernel stores the 64 terms of every block (hvs_store_terms) instead of their sum. The masking table is a compile-time constant built from the CPU's double product (hvs_mask_value,MASK_TABLES), each work-item forms its own threshold withsqrt_prod_rn()(hvs_threshold) and the pair exchanges that value, and the coefficient error is an integersycl::abs(). The readback isHVS_TERMS(64) floats per block (hvs_terms_bytes),args.termsreplacesargs.partials, andreduce_hvs_planes()/append_hvs_scores()callpsnr_hvs_score.h. Keep the kernel free of private arrays and of a sum of terms;hvs_mask_value()is the only place the TU may usedoublefor device data, and only at compile time.core/src/feature/hip/integer_psnr_hvs/psnr_hvs_score.hip,integer_psnr_hvs_hip.{c,h}: the same change in the CUDA kernel's form (HVS_TABLES,hvs_thresholdwith a double product and root,hvs_store_terms,PSNR_HVS_HIP_TERMS,psnr_hvs_terms_bytes).core/src/meson.build:'psnr_hvs_score'joinship_cu_extra_flagswith-ffp-contract=off -fhip-fp32-correctly-rounded-divide-sqrt. Keep it when resolving a conflict in that dictionary. The SYCL TU needs no entry: every SYCL feature TU has the strict line of ADR-1367.scripts/ci/cross_backend_calibration.py:EXACT_TWINS["psnr_hvs"]iscuda,sycl,hip.FEATURE_TOLERANCE["psnr_hvs"], the calibration rows andarea_tolerance_factor()are unchanged; no gate backend reads them forpsnr_hvsany more.- Tests:
core/test/psnr_hvs_twin_parity.h(new) holds the fixtures and the bit comparison;test_cuda_psnr_hvs_parity,test_sycl_psnr_hvs_parityandtest_hip_psnr_hvs_parityare thin wrappers that describe their backend. A test without a device now exits 77 (skipped) instead of 0.test_sycl_fp_arith_contractand its probe gain a fourth result (sqrt_prod_rn) and a midpoint operand set;test_psnr_hvs_twin_exact_sum_contract.pycovers the three twins and the helper.
No public C API, CLI syntax, option table or FFmpeg patch impact. CPU scores are unchanged (no CPU extractor code is touched). psnr_hvs_sycl and psnr_hvs_hip scores move to the CPU's: by up to 8.4e-5 dB on the Netflix 576x324 pair and 1.7e-2 dB at 3840x2160. The Metal twin is untouched and still sums per block. Measured on an Arc A380 (xe driver) and a gfx1036 (zeus): every frame of the Netflix 576x324 pairs (8, 10, 12 bits, 4:2:2), the 1920x1080 checkerboard pairs and BBB 1920x1080 / 3840x2160 (8 and 10 bits) equals --backend cpu of the same binary at --precision max. Re-verify after a rebase that touches these files with the reproducer of Research-1401.
HIP SpEED twins read the host's lanczos4 weight table (2026-10-01)¶
fix/hip-speed-lanczos4-host-weights, T-GPU-SPEED-LANCZOS4-PRESCALE-DRIFT-2026-09-30 (HIP half; CUDA and SYCL landed before).
core/src/feature/hip/speed/speed_hip_device.h:SpeedHipParamshas a seventeenth pointer,lanczos(the layout assert follows);speed_hd_lanczos_weight()andspeed_hd_sinpi()are gone, andspeed_hd_scale_lanczos()takes the nine column and nine row taps.core/src/feature/hip/speed_hip_pipeline.c: the arena has alanczosblock, filled at init byspeed_hip_upload_lanczos()fromspeed_internal_gpu_lanczos_weights(). An upstream change tolanczos4_kernel()or to the sample position ofvif_scale_frame_lanczos4_s()reaches the three device backends throughvif_scale_lanczos4_axis_weights(); nothing HIP-specific has to follow.core/test/test_gpu_speed_lanczos4_parity.cis built a third time (-DLZ_BACKEND_HIP=1,test_hip_speed_lanczos4_parity);test_hip_speed_device_mathhas two lanczos4 cases;test_hip_device_resident_contract.pyforbids a device sine and a table that is not the host's (four planted regressions).T-CI-PARITY-GATE-MOTION-DEBUG-DEFAULT-2026-09-29(ADR-1418):core/src/feature/sycl/integer_motion_sycl.cppdeclaresdebugwith defaultfalse, like the CPU, CUDA and HIP motion extractors; keep the four declarations equal (core/test/test_sycl_twin_option_parity.c).scripts/ci/cross_backend_parity_gate.pyandscripts/ci/cross_backend_vif_diff.pycarry amotion_debugcell (motionwithdebug=true) and fail a cell whose runs emit different metric sets; do not restore a comparison over the common subset.
psnr_hvs_sycl and psnr_hvs_hip score 4:0:0; psnr_hvs_hip takes enable_chroma (2026-10-01)¶
fix/psnr-hvs-sycl-hip-yuv400, closes T-SYCL-HIP-PSNR-HVS-YUV400-REFUSED-2026-10-01. No ADR: the behaviour is the CPU extractor's (third_party/xiph/psnr_hvs.c::init) and the CUDA twin's.
core/src/feature/sycl/integer_psnr_hvs_sycl.cpp:validate_hvs_input()no longer rejectsVMAF_PIX_FMT_YUV400P;configure_hvs_geometry()setsn_active_planesto 1 for 4:0:0 orenable_chroma=falseand gives the format a case in its switch.core/src/feature/hip/integer_psnr_hvs_hip.c: newenable_chromaoption andn_planesstate member.psnr_hvs_set_plane_dims()setsn_planes(and accepts 4:0:0); buffer allocation, staging, uploads, the kernel'sargs.n_planes, the plane scores and the emitted features loop topsnr_hvs_plane_count(s)(n_planes, clamped to the three planes the state holds) instead ofPSNR_HVS_NUM_PLANES. Keep it that way when resolving conflicts: a fixed three-plane loop readsdata[1]of a luma-only picture.core/test/psnr_hvs_twin_parity.h:HvsTwin.scores_yuv400is gone (every twin scores 4:0:0),HvsFixture.luma_onlyruns both sides withenable_chroma=false, andhvs_twin_luma_only_identical()is a new shared case that the three twin tests call.
No public C API or CLI syntax change. One new extractor option (psnr_hvs_hip: enable_chroma), documented in docs/metrics/psnr-hvs.md. No FFmpeg patch impact. CPU scores are unchanged.
ADR-1412 — float_vif_cuda computes the CPU's arithmetic and is bit-identical to it (2026-10-01)¶
fix/cuda-float-vif-cpu-arithmetic, Research-1412, ADR-1412.
core/src/feature/cuda/float_vif/float_vif_device.h(new): the kernel argument blocks and the per-pixel arithmetic, in plain C and CUDA C++.fvif_log2()=vif_tools.c::log2f_approx()withhorner_s()andlog2_poly_s[];fvif_pixel_statistic()=vif_pixel_statistic_s()(vif_sigma_nsqin fp64);fvif_row_sum()/fvif_sum_rows()= the two loops ofvif_statistic_s(). If upstream Netflix changes any of those, orvif_get_filter(), or dropsVIF_OPT_FAST_LOG2fromvif_options.h, change this header in the same PR;core/test/test_float_vif_device_math.c(device-free, bit compare againstvif_statistic_s()andcompute_vif()) andcore/test/test_cuda_float_vif_exact_contract.pyfail until it follows.core/src/feature/cuda/float_vif/float_vif_score.cu: no tap table (FVIF_COEFF_S0..S3are gone; the taps arrive inFloatVifCudaTaps), no warp or block reduction. Three kernels, each with one by-value argument block:float_vif_compute(stores two terms per pixel, column by column),float_vif_row_sums(one thread per row),float_vif_decimate. Tile loads go throughcuda_tile_index.h.core/src/feature/cuda/float_vif_cuda.c:float_vif_init_taps()callsvif_get_filter_size()/vif_get_filter()asfloat_vif.c::init()does; the readback is two floats per row per scale (rows_host). New optionsvif_scale1_min_val/vif_scale2_min_val/vif_scale3_min_val, same table entries as the CPU's.scripts/ci/cross_backend_calibration.py:EXACT_TWINS["float_vif"] = {"cuda"}. On a conflict with another twin's entry keep both.core/test/test_cuda_float_vif_parity.cis rewritten around a case table and asserts equality.- No Netflix golden-data, public C API or FFmpeg patch impact. The
float_vif_sycl,float_vif_hipandfloat_vif_metaltwins are untouched and still hold the old tap table (T-GPU-FLOAT-VIF-CPU-ARITHMETIC-2026-10-01).
float_ms_ssim on HIP follows the CPU's arithmetic through one header (2026-10-01)¶
fix/hip-float-ms-ssim-cpu-arithmetic, T-GPU-FLOAT-MS-SSIM-CPU-ARITHMETIC-2026-10-01 (HIP part), ADR-1403.
- New
core/src/feature/hip/integer_ms_ssim/ms_ssim_arith.h(plain C and HIP C++, listed inhip_kernel_shared_headers): the decimate sample, both window passes, the l / c / s terms, and on the host the constants, the per-scale mean and the Wang combine.ms_ssim_score.hipandinteger_ms_ssim_hip.ccall it and keep no arithmetic of their own. - It mirrors
ms_ssim_decimate.c(fused taps),iqa/convolve.c(fp32 products, fp64 sum, one rounding per pass; carried as an fp32 pair),iqa/ssim_tools.c::ssim_variance_scalar()/ssim_accumulate_default_scalar()and the product inms_ssim.c::ms_ssim_score_scales(). An upstream change to any of those has to be made in the header in the same PR;test_hip_ms_ssim_arithfails otherwise, with or without an AMD device. - The dead
ms_ssim_warp_reduce()helper is gone from the kernel, andinteger_ms_ssim_hip.cno longer hasg_alphas/g_betas/g_gammas. scripts/ci/tidy-baseline-hip.json:integer_ms_ssim_hip.c1 -> 0,test_hip_ms_ssim_parity.c21 -> 0.
docs/rc-phase-map-rc3-rc8 — first-release candidate map RC3 to RC8 (2026-10-01)¶
No upstream source impact: this is fork-only release governance and documentation. ADR-1421 supersedes the candidate mapping of ADR-1352 (ADR-1341's other rules stay). RC3 owns twin exactness, RC4 the first full Rust metric, RC5 deduplication, RC6 the GPU capability table, RC7 benchmarks, profiling and tuning, and RC8 the one-shot real retrain. When rebasing release, roadmap, runbook, tester or ledger documents, never restore "RC3 = benchmarks" or "RC4 = retrain" in forward-looking text; the tools/rc1-tester catalog phases are RC1, RC7 and RC8. tools/rc1-tester/ keeps its name and the backlog IDs T-RC2-BENCH-TUNE and T-RC3-MODEL-RETRAIN stay stable (ADR-1303). docs/state.md keeps its two ADR-1352 disposition labels until T-STATE-LEDGER-RC-RELABEL-2026-10-01 relabels them; the section "First-release phase classification" states how they read meanwhile. The Netflix golden assertions are untouched.
ADR-1416 — adm_cuda runs the CPU's host routines and folds the denominator per row (2026-10-01)¶
fix/cuda-adm-cpu-arithmetic, Research-1416, ADR-1416.
core/src/feature/cuda/integer_adm_cuda.cincludesfeature/integer_adm_kernels.hand has no copy ofdwt_quant_step(),adm_csf_factors(),conclude_adm_cm()orconclude_adm_csf_den(). The CSF weights, the denominator border and shifts and the per-scale scores come from the CPU'sadm_csf_factors(),adm_csf_den_ctx_init()/i4_adm_csf_den_ctx_init(),adm_cm_ctx_init()/i4_adm_cm_ctx_init()and the four*_result()routines. If upstream Netflix changes any of them, the twin follows through the header; do not reintroduce a copy.AdmStateCudalostrfactor[]andcsf_normalization_shift[].core/src/feature/cuda/integer_adm/adm_csf_den.cuis rewritten: kernelsadm_csf_den_scale_row_kernelandadm_csf_den_s123_row_kernel(the_line_kernel_8_128names are gone), one block of 128 threads per row and band, one fold per row, shifts as arguments.core/src/feature/adm_cm_accumulator.h: newadm_csf_den_round_row_total();integer_adm_kernels.h::adm_csf_den_fold()calls it (same arithmetic, CPU scores unchanged). A twin that folds a denominator row must call it on the whole row.scripts/ci/cross_backend_calibration.py:EXACT_TWINS["adm"] = {"cuda"}. On a conflict with another twin's entry keep both.core/test/test_cuda_adm_parity.cis rewritten around exact cases (textured, sparse, 962x13542, 10-bit, options).adm_cm.cu, the HIP, SYCL and Metal twins and every Netflix golden assertion are untouched. No public C API or FFmpeg patch impact.
psnr_hvs_sycl and psnr_hvs_hip compact nonzero terms on the device before readback (2026-10-01)¶
perf/sycl-hip-psnr-hvs-tune, T-SYCL-HIP-PSNR-HVS-EXACT-SUM-THROUGHPUT-2026-10-01, ADR-1397 / ADR-1401.
- Reused
vmaf_psnr_hvs_plane_score_compacted()incore/src/feature/psnr_hvs_score.c/.hacross the twins (HISS-19), summing only nonzero terms in CPU block order ($x + 0.0\\text{f} == x$). - Device prefix scan and compaction kernels added in
core/src/feature/sycl/integer_psnr_hvs_sycl.cppandcore/src/feature/hip/integer_psnr_hvs/psnr_hvs_score.hip(hvs_scan_reduce_hip,hvs_scan_prefix_hip,hvs_compact_hip). - Readback shrinks by 94-97% (from ~198.3 MB to ~11.0 MB at 4K).
- Scratch memory on SYCL remains 0 private memory and 0 register spills (
test_sycl_kernel_scratchpasses). - Bit-identical parity maintained against
--backend cpuof the same binary at--precision max.
ADR-1424 — integer_ssim_cuda adds its terms in the CPU's raster order (2026-10-01)¶
fix/cuda-ssim-cpu-arithmetic, Research-1424, ADR-1424.
core/src/feature/cuda/integer_ssim/integer_ssim_score.cu:integer_ssim_vert_combinetakesdouble *terms(width x height, raster order) where it took per-blockpartials, and stores each pixel's term instead of reducing it. The int64 weight reduction is unchanged. If upstream changesssim_reduce_row_range()(the term) or the order in whichcalc_ssim()adds the terms, changeissim_term()or the host loop in the same PR.core/src/feature/cuda/ssim_cuda.c:rb_ssimis one double per pixel;issim_frame_sum()adds it in index order. No other host arithmetic changed.scripts/ci/cross_backend_parity_gate.py: new gate featuressim(FEATURE_METRICS,FEATURE_TOLERANCE, threeBACKEND_EXTRACTOR_ALIASESentries).scripts/ci/cross_backend_calibration.py:EXACT_TWINS["ssim"] = {"cuda"}. On a conflict with another twin's entry keep both.core/test/test_cuda_ssim_parity.cis rewritten around a case table and asserts equality;core/test/test_cuda_ssim_exact_contract.pyis new.- No Netflix golden-data, public C API or FFmpeg patch impact. The SYCL, HIP and Metal
ssimtwins are untouched (T-GPU-SSIM-FRAME-SUM-ORDER-2026-10-01).
ADR-1419 — float_motion_hip adds its SAD in the CPU's order (2026-10-01)¶
fix/hip-float-motion-cpu-float-sum, T-GPU-FLOAT-MOTION-CPU-FLOAT-SUM-2026-10-01 (HIP part).
- New
core/src/feature/hip/float_motion/float_motion_rows.h(plain C and HIP C++, listed inhip_kernel_shared_headers): the absolute difference, the transposed layout index, the per-row fp32 sum,motion.c's bilinear sample (moved here from the kernel) and, on the host, the plane score. float_motion_score.hip: the blur kernels takediffinstead ofpartialsand store|cur - prev|;float_motion_hip_scale1_sadbecamefloat_motion_hip_scale1_diff;float_motion_hip_row_sumis new. The wave and block reductions (fm_warp_reduce,fm_block_reduce) are gone.float_motion_hip.c: each plane hasdiff[2](scale 0, scale 1); the readback holdsh(+sh) row sums per plane instead of block partials (row_count,rows1,off0,off1);fm_hip_frame_score()callsvmaf_hip_float_motion_plane_score().- It mirrors
float_motion.c::float_sad_line_c()/compute_motion_simd()andmotion.c::vmaf_image_sad_c()/motion_scale_bilinear(). An upstream change to the order or the types of those sums has to be made in the header in the same PR;test_hip_float_motion_rowsfails otherwise, with or without an AMD device. scripts/ci/cross_backend_calibration.py:EXACT_TWINS["float_motion"]gainship;scripts/ci/test_cross_backend_parity_gate.pyexpects it. When another backend joins, keep every listed twin.
ADR-1423 — adm_hip computes with the CPU's routines and folds the denominator per row (2026-10-01)¶
fix/hip-adm-cpu-arithmetic, T-HIP-ADM-NOT-CPU-ARITHMETIC-2026-10-01, T-HIP-ADM-FIRST-FRAME-STALE-ACCUMULATORS-2026-10-01.
core/src/feature/hip/integer_adm_hip.cincludesinteger_adm_kernels.h. Its copies ofdwt_quant_step(),adm_csf_factors(),conclude_adm_cm()andconclude_adm_csf_den()are gone, and so isAdmStateHip::csf_normalization_shift; the scores come fromadm_cm_result()/adm_csf_den_result()and theiri4_forms. A change to those CPU routines or to the context initialisers reaches the twin without an edit here.core/src/feature/hip/integer_adm/adm_csf_den.hip: the two kernels areadm_csf_den_scale_row_kernelandadm_csf_den_s123_row_kernel(were..._line_kernel_8_128), launched as1 x rows x 3blocks of 128 threads with the border and the shifts as arguments. The fold callsadm_csf_den_round_row_total()(adm_cm_accumulator.h, ADR-1416), like the CUDA kernel and the CPU.- The per-frame
hipMemsetAsyncoftmp_rescomes afteradm_hip_stage_luma(), directly ahead of the kernels. Do not move it back ahead of the upload in a rebase: there it is lost in the first context of a process that needs larger planes than the contexts before it. scripts/ci/cross_backend_calibration.py:EXACT_TWINS["adm"]listshipnext tocuda. A textual merge of two branches that each add an"adm"entry leaves two dictionary keys, and Python keeps only the last one; keep one entry with every twin.
vif_sycl rounds its sums as the CPU does (2026-10-01)¶
fix/sycl-vif-cpu-float-sums. No ADR: a bug fix with one way to do it.
core/src/feature/sycl/integer_vif_sycl.cpp:vif_compute_scores()is gone.vif_scale_sums()rounds each scale's numerator and denominator tofloat(the two(float)(...)casts mirrorinteger_vif.c::vif_store_residuals()), andvif_score_set()mirrorsinteger_vif.c::write_scores(): it adds the rounded values and sets.single_precision_ratio = true. On rebase: if the other side still computesdoublesums incollect_fex_sycl(), keep this side; if upstream Netflix changes whereinteger_vif.crounds, mirror it here, incuda/integer_vif_cuda.cand inhip/integer_vif_hip.c, which carry the same tail.- The twin's
debugoption defaults tofalse, as on the CPU. - The kernels are unchanged.
dev_vif_stats_log_domain()still forms the gain in fp32 (T-SYCL-VIF-FP32-GAIN-2026-10-01). - Tests:
core/test/test_sycl_vif_parity.c(device) andcore/test/test_sycl_vif_float_sums_contract.py(device-free, four planted regressions). - No Netflix golden-data, public API or FFmpeg patch impact.
ADR-1426 — ciede_cuda computes the CPU's arithmetic; the gate bounds the math library (2026-10-01)¶
fix/cuda-ciede-cpu-arithmetic, Research-1426, ADR-1426.
core/src/feature/cuda/integer_ciede/ciede_device.h(new):ciede.c'sget_lab_color(),ciede2000()and helpers in plain C and CUDA C++, with the reference's fp64 / float split and every libm promotion written out;CIEDE_POWFis glibc'spowfon the host and the correctly rounded power on the device;ciede_frame_sum()isextract()'s accumulator. If upstream Netflix changes any of those functions, change this header in the same PR;core/test/test_ciede_device_math.c(device-free, replays whole frames against the CPU extractor) fails until it follows.core/src/feature/cuda/integer_ciede/ciede_score.cu: the fp32 formulation and the warp / block reduction are gone. Both kernels callciede_pixel()and store one float per pixel; channel reads keep the ADR-0762__ldg()pattern.core/src/feature/cuda/integer_ciede_cuda.c: the read-back is the term plane (one float per pixel); the score is the reference's expression.scripts/ci/cross_backend_calibration.py: newLIBM_TWINS/libm_pair_tolerance()(ciede:cudaat1e-9), consumed by both gates. NotEXACT_TWINS: the twin is not bit-identical.core/test/test_cuda_ciede_parity.cis rewritten around a case table at1e-8;core/test/test_cuda_ciede_exact_contract.pyis new.- No Netflix golden-data, public C API or FFmpeg patch impact;
ciede.cis not touched. Theciede_sycl,ciede_hipandciede_metaltwins are untouched (T-GPU-CIEDE-CPU-ARITHMETIC-2026-10-01).
ADR-1427 — a HIP frame clears its accumulators after its upload (2026-10-01)¶
fix/hip-first-frame-accumulator-clear, T-HIP-FIRST-FRAME-ASYNC-CLEAR-OTHER-TWINS-2026-10-01.
core/src/feature/hip/float_moment_hip.c(moment_hip_launch()),integer_vif_hip.c(submit_fex_hip()) andfloat_psnr_hip.c(float_psnr_hip_launch()) queue theirhipMemsetAsyncaftervmaf_hip_plane_source_acquire_luma(), directly ahead of the kernels.integer_adm_hip.cdoes since ADR-1423. Keep that order in every conflict resolution: a clear queued ahead of the upload is lost on a gfx1036 in the first context of a process that needs larger planes than the contexts before it, and the twin then returns a wrong first frame without an error.core/test/test_hip_clear_after_upload_contract.pyreads every*.cundercore/src/feature/hip/andcore/src/hip/and fails on a function that clears and uploads afterwards, helpers included. A new upload entry point or clear call belongs in itsUPLOADS/CLEARStuples.core/test/meson.buildbuildstest_hip_first_frame_clear.conce per name inhip_first_frame_twins; the contract fails when that list and the test'scases[]table differ.
ADR-1422 — float_vif_sycl computes the CPU's arithmetic without fp64 (2026-10-01)¶
fix/sycl-float-vif-cpu-arithmetic, ADR-1422 (follows ADR-1412).
core/src/feature/sycl/float_vif_sycl.cpp: the tap table (VifFilterConstants<SCALE>::coeff) is gone;init_vif_taps()callsvif_get_filter()and every launch passes the scale'sVifTapsby value. The filter kernel storessigma1_sq/sigma2_sq/sigma12(store_vif_sigmas());vif_contribution(),reduce_vif_group(), the sub-group accessors and the per-group partial buffers are gone. New:FloatVifStatisticKernel(one work-item per pixel, sub-group size 16) andvif_row_sums()(one work-item per row, sub-group size 8);sum_vif_rows()adds the rows in fp32. Those loop shapes and types are load-bearing: a group, sub-group, strided or atomic reduction, an fp64 fold of the rows, a tap literal orsycl::log2gives a different rounding and the twin stops matching the CPU. On rebase: if the other side still has thecoefftables,sycl::log2orsycl::reduce_over_groupin this TU, keep this side.core/src/feature/sycl/sycl_float_vif_math.h(new) mirrorsvif_tools.c::vif_pixel_statistic_s()andlog2f_approx(). If upstream Netflix changes either,vif_statistic_s(),vif_get_filter()orVIF_OPT_FAST_LOG2, mirror it in this header and incore/src/feature/cuda/float_vif/float_vif_device.hin the same change. The header is fp64-free outsidemake_noise_variance()andmake_statistic_params(), which run on the host; do not replaceone_plus_ratio()by a plain fp32 expression or by the pair alone.- The twin has the CPU's
vif_scale1_min_val/vif_scale2_min_val/vif_scale3_min_valoptions now. scripts/ci/cross_backend_calibration.py:EXACT_TWINS["float_vif"]is{"cuda", "sycl"}. Another branch may add twins to the same table; keep both sides' entries.- Depends on ADR-1367 (the SYCL strict FP line) and ADR-1395: no kernel of the twin may use scratch memory. The statistic must not move back into the filter kernel, and
soft_add()/noise_plus()must keep selecting scalars, not structs. - Tests:
core/test/test_sycl_float_vif_parity.covercore/test/float_vif_twin_parity.h(==, every output, device),core/test/test_sycl_float_vif_math.cwith its probetest_sycl_float_vif_math_probe.cpp(host and device),core/test/test_sycl_float_vif_exact_contract.py(eleven planted regressions),scripts/ci/test_cross_backend_parity_gate.py. - No Netflix golden-data, public API or FFmpeg patch impact.
ADR-1420 — float_adm_cuda computes the CPU's arithmetic and is bit-identical to it (2026-10-01)¶
fix/cuda-float-adm-cpu-arithmetic, Research-1420, ADR-1420.
core/src/feature/adm_tools.c(Netflix file): three changes, none in the arithmetic.rcp_s()reads the processor's estimate throughrcp_estimate_s();adm_decouple_s()readscos^2fromadm_decouple_cos_1deg_sq_s(); the tails ofadm_csf_den_scale_s(),adm_csf_den_scale_s_p3(),adm_cm_s()andadm_cm_s_p3()calladm_pool_bands_s().adm_border_s()andadm_csf_rfactor_s()loststatic, andAdmBorderSmoved to the newcore/src/feature/adm_float_reference.h, which declares all of these plusadm_divs_is_reciprocal_s()/adm_divs_reciprocal_estimate_s(). On an upstream sync that touches these functions keep the exported names and the single pooling routine:float_adm_cuda.ccalls them, andcore/test/test_cuda_float_adm_exact_contract.pycounts the four call sites. CPU scores are unchanged (4 752 outputs over six fixtures and twelve option sets).core/src/feature/adm_reciprocal_model.{c,h}(new): the table model of the host'sRCPSSestimate and its probe. The header'sadm_reciprocal_model_bits()is integer-only and is compiled by the device as well.core/src/feature/cuda/float_adm/float_adm_device.h(new): the kernel argument blocks and the per-sample arithmetic, in plain C and CUDA C++:fadm_divs()=DIVS(),fadm_angle_flag()=adm_angle_flag_s()(theADM_OPT_AVOID_ATANbranch),fadm_decouple_band()=adm_decouple_band_s(),fadm_csf_flt()= thefltstore ofadm_csf_s(),fadm_thresh_band()/fadm_threshold()=adm_cm_thresh3x3_s(),fadm_den_term()/fadm_cm_term()= the terms ofadm_csf_den_scale_s()/adm_cm_s(),fadm_row_sum()/fadm_fold_rows()= their two accumulators. If upstream Netflix changes any of those, change this header in the same PR;core/test/test_float_adm_device_math.c(device-free, bit compare against the CPU routines) and the contract fail until it follows.core/src/feature/cuda/float_adm/float_adm_score.cu: the DWT kernels are unchanged except thatfadm_mirror()stays inside a one-sample input.float_adm_csf_cm,float_adm_csf_randfloat_adm_aim_cmare gone, withFADM_ACCUM_SLOTSand every warp reduction;float_adm_decouple_csfwrites both CSF pairs,float_adm_termsstores nine terms per sample of the reduced region (slot by slot, column by column),float_adm_row_sumsadds one row per thread. Each takes one by-value argument block.core/src/feature/cuda/float_adm_cuda.c: no copy ofdwt_quant_step(); the readback is nine floats per row per scale (rows_host); the frame sums are floored at1e-10, as incompute_adm().scripts/ci/cross_backend_calibration.py:EXACT_TWINS["float_adm"] = {"cuda"}. On a conflict with another twin's entry keep both.core/test/test_cuda_float_adm_parity.cis rewritten around a case table and asserts equality.- No Netflix golden-data, public C API or FFmpeg patch impact. The
float_adm_sycl,float_adm_hipandfloat_adm_metaltwins are untouched (T-GPU-FLOAT-ADM-CPU-ARITHMETIC-2026-10-01).
ADR-1433 — ssimulacra2_cuda returns the sums of the CPU's loops (2026-10-01)¶
fix/cuda-ssimulacra2-cpu-sum-order, Research-1433, ADR-1433.
core/src/feature/ordered_sum.h(new): the bits offor (...) s += x[i];over non-negative doubles from chunk-wise integer increments (plan, increments for an even and an odd start, checked walk). Plain C for the host, device code throughVMAF_ORDSUM_FUNC/VMAF_ORDSUM_BITS/VMAF_ORDSUM_FROM_BITS.core/test/test_ordered_sum.ccompares it with the loop on the host.core/src/feature/cuda/ssimulacra2/ssimulacra2_device.cu:ssimulacra2_combine_partialsandssimulacra2_combine_finalare gone.ssimulacra2_chunk_sums,ssimulacra2_chunk_plan,ssimulacra2_chunk_unitsandssimulacra2_ordered_totalsreplace them;ss2c_terms()holds the per-pixel terms ofssim_map()andedge_diff_map(). If upstream Netflix or libjxl changes those two functions inssimulacra2.c(a term, its clamp, the order of the six sums), changess2c_terms()in the same PR; the terms must stay non-negative or NaN.core/src/feature/cuda/ssimulacra2_cuda.{c,h}: four launches per scale for the sums, three small device buffers (d_chunk_sums,d_plan,d_units) instead ofd_partials; the readback andcollect()are unchanged.scripts/ci/cross_backend_calibration.py:ssimulacra2/cudainEXACT_TWINS.core/test/test_cuda_ssimulacra2_parity.casserts==on three fixtures;core/test/test_cuda_ssimulacra2_exact_contract.pyis new.- No Netflix golden-data, public C API or FFmpeg patch impact;
ssimulacra2.cis not touched. The SYCL, HIP and Metal twins are untouched (T-GPU-SSIMULACRA2-SUM-ORDER-2026-10-01).
ADR-1428 — exact twins are declared by fragment files (2026-10-01)¶
refactor/exact-twins-fragments, ADR-1428.
scripts/ci/exact_twins.d/<feature>.<backend>(new, one per listed twin;adr:andevidence:) replaces theEXACT_TWINSdict literal inscripts/ci/cross_backend_calibration.py, which now loads and validates the directory at import. A pull request that addsEXACT_TWINS[...]entries, per-feature exact tests or enumerating prose conflicts with this once: drop those hunks and add a fragment file instead (steps inscripts/ci/AGENTS.md). Never reintroduce the literal.docs/development/cross-backend-exact-twins.mdis generated (scripts/docs/generate-exact-twins.py,make docs-fragments-write): on a conflict take master's side and regenerate.
float_adm refuses frames below 17x17 (2026-10-01)¶
fix/float-adm-min-frame, T-FLOAT-ADM-TINY-FRAME-BAND-READS-2026-10-01.
core/src/feature/float_adm.c::init()callsadm_frame_size_check()(adm_csf_fixed_point.h) before it allocates. Upstream Netflix has no such check and reads outside the scale-3 bands of smaller frames; on an upstream sync that touchesinit(), keep the check and its place before the allocations.core/src/feature/cuda/float_adm_cuda.c::init_fex_cuda()has the same check, before any device resource is claimed.provided_featuresinfloat_adm.ckeeps upstream'sadm_scale0entry. It looks like a mistake (the extractor emitsadm), and it is why the debug ratio gets no option suffix, but the Netflix golden tests read that unsuffixed key under non-default options (T-FLOAT-ADM-DEBUG-KEY-UNSUFFIXED-2026-10-01). Do not change the list.- No Netflix golden-data, public C API or FFmpeg patch impact. The SYCL, HIP and Metal
float_admtwins are untouched (T-GPU-FLOAT-ADM-TINY-FRAME-FLOOR-2026-10-01).
ADR-1432 — vif_sycl computes the gain terms exactly (2026-10-01)¶
fix/sycl-vif-fp64-gain, ADR-1432 (follows the host-tail fix above and ADR-1422).
core/src/feature/sycl/sycl_integer_vif_math.h(new) mirrors six lines ofinteger_vif.c::vif_accumulate_pixel()(also inx86/vif_avx2.c,x86/vif_avx512.c,arm64/vif_neon.c). If upstream Netflix changes them, changegain_terms_integer()andgain_terms_replayed()in the same change;core/test/test_sycl_vif_exact_gain_contract.pyfails when the lines move.kEpsMant/kEpsExpare the fp64 bits of65536 * 1.0e-10.core/src/feature/sycl/sycl_soft_double.h(new) holds the soft-fp64 primitives;sycl_float_vif_math.hlost its own copies ofSoftDouble,soft_round(),soft_add(),soft_div(),soft_from_float()andsoft_to_float()and includes it. On rebase: if the other side edits those functions insycl_float_vif_math.h, move the edit to the shared header.core/src/feature/sycl/integer_vif_sycl.cpp:dev_vif_stats_log_domain()callsvmaf_sycl_ivif::gain_terms(); the fp32 gain, itssycl::fmaandsycl::fminare gone and must not come back. The gain limit travels asVifGainLimit(built ininit), the per-pixel terms asvif_terms(sevenint32_t), and the fused kernel of scale 0 takes the large register file at SIMD-16 (vif_fused_grf_size()). The last two keep the kernels free of scratch memory (ADR-1395).- Do not use
sycl::mul_hi()on 64-bit operands in a kernel: it returned wrong values on an Arc A380.u128_mul()forms the product in limbs. vifis declared exact forsyclbyscripts/ci/exact_twins.d/vif.sycl(ADR-1397's exact cell in ADR-1428's fragment form); the row indocs/development/cross-backend-exact-twins.mdis generated (make docs-fragments-write).- Tests:
core/test/test_sycl_integer_vif_math.cwith its probetest_sycl_integer_vif_math_probe.cpp(host and device),core/test/test_sycl_vif_parity.c(==, every output),core/test/test_sycl_vif_exact_gain_contract.py(seven planted regressions). - No Netflix golden-data, public API or FFmpeg patch impact.
ADR-1430 — speed_chroma on CUDA joins the parity gate with a log2f bound (2026-10-01)¶
fix/cuda-speed-chroma-libm-bound, Research-1430, ADR-1430.
- No source of
speed_chroma_cudachanges.speed_log2()incore/src/feature/cuda/speed/speed_score.custays correctly rounded; do not replace it by a port of a C library'slog2f. scripts/ci/cross_backend_parity_gate.pyandscripts/ci/cross_backend_vif_diff.py: new featurespeed_chroma(speed_chroma_u,speed_chroma_v,speed_chroma_uv), places=4 for an unlisted twin.scripts/ci/cross_backend_calibration.py:LIBM_TWINS["speed_chroma"] = {"cuda": 5e-6}.core/test/test_cuda_speed_chroma_parity.c: 960x960 textured fixture (regular covariance), all three scores of every frame, relative bound1e-6. If upstream changesspeed.c'slog2fcalls or scoring, re-run it and the gate cell; a new difference is the twin's until a run with a correctly roundedlog2fpreloaded shows otherwise.- No Netflix golden-data, public C API or FFmpeg patch impact;
speed.cand the SYCL, HIP and Metal twins are untouched.
ADR-1437 — HIP twins declared exact after a sweep (2026-10-01)¶
test/hip-exact-twins-declared, T-HIP-EXACT-TWINS-UNDECLARED-2026-10-01.
- Seven fragment files under
scripts/ci/exact_twins.d/(ADR-1428) declare the twins:motion.hip,motion_debug.hip,motion_v2.hip,psnr.hip,cambi.hip,float_ms_ssim.hipandfloat_ms_ssim_lcs.hip. Nothing shared is edited; another backend's declaration for the same feature is another file. core/test/test_hip_exact_twins.c(new) asserts==for the five twins. A rebase that brings a float reduction or a host copy of a CPU routine intointeger_motion_sad_hip.c,integer_psnr_hip.c,integer_cambi_hip.corinteger_ms_ssim/ms_ssim_arith.hfails it on a device; no test without a device covers the listing.- No source of a twin changed. No Netflix golden-data, public API or FFmpeg patch impact.
ADR-1438 — integer_ssim_hip adds every frame in the CPU's order (2026-10-01)¶
fix/hip-ssim-cpu-frame-sum, T-GPU-SSIM-FRAME-SUM-ORDER-2026-10-01 (HIP part).
core/src/feature/hip/integer_ssim/integer_ssim_score.hip:integer_ssim_vert_combineandissim_pixel_term()are gone.integer_ssim_vert_termsis the only pass-2 kernel; its seventh and eighth arguments are nowdouble *terms(one per pixel) andint64_t *block_weights(one per block). A rebase that restores the per-block term tree, or an identical-window shortcut, makes the twin inexact again;test_hip_kernel_source_contract.pyrejects both.core/src/feature/hip/integer_ssim_hip.c:raster,pair_countandfunc_vertare gone;term_countsizesrb_ssim(width x height doubles) andblock_countsizesrb_wgt.ISSIM_HIP_RASTER_MAX_PIXELSis removed frominteger_ssim_hip.h.docs/adr/1400-hip-integer-ssim-raster-sum-small-frames.mdis superseded (status line and index fragment only).scripts/ci/silent-revert-allowlist.jsongains two ADR-1438 entries forinteger_ssim_hip.h: removing ADR-1400's bound returns the header to the blob it had before #1673, which the silent-revert gate reports as a rewind and as a reverse hunk. Both entries can go once this change is on the target.scripts/ci/exact_twins.d/ssim.hip(new) declares the twin exact (ADR-1428); nothing shared is edited.
ADR-1440 — float_psnr_hip adds integer block sums (2026-10-01)¶
fix/hip-float-psnr-exact-block-sums, T-HIP-FLOAT-PSNR-FP32-BLOCK-SUMS-2026-10-01.
core/src/feature/hip/float_psnr/float_psnr_score.hip: both kernels takeuint32_t *partials(wasfloat *) and write two values per block,partials[2 * block]andpartials[2 * block + 1].fpsnr_warp_reduce()reducesuint32_t;fpsnr_square()andfpsnr_block_sum()are new. A rebase that brings afloataccumulator back makes the twin inexact at 10 bits and above;test_hip_float_psnr_exact_contract.pyrejects it.core/src/feature/hip/float_psnr_hip.c: the read-back isfloat_psnr_hip_partials_bytes()(twouint32per block), andcollect()divides the sum byscaler * scaler.scripts/ci/exact_twins.d/float_psnr.hip(new) declares the twin exact (ADR-1428). The CUDA, SYCL and Metal twins still reduce in fp32.- No Netflix golden-data, public API or FFmpeg patch impact. The CPU extractor is untouched.
ADR-1441 — float_ssim_hip uses the shared window arithmetic (2026-10-01)¶
fix/hip-float-ssim-cpu-arithmetic, T-HIP-FLOAT-SSIM-NOT-CPU-ARITHMETIC-2026-10-01.
core/src/feature/hip/float_ssim/ssim_score.hipincludes../integer_ms_ssim/ms_ssim_arith.h.SSIM_G,struct SsimMomentsandssim_lcs()are gone;ssim_horiz()callsvmaf_hip_ms_ssim_horizontal(),ssim_vertical_moments()returnsvmaf_hip_ms_ssim_vertical()andssim_pixel()callsvmaf_hip_ms_ssim_lcs(). A rebase that restores an fp32 running sum or a locall/c/smakes the twin inexact again;test_hip_kernel_source_contract.pyrejects both.- A change to
iqa/convolve.cor tossim_accumulate_default_scalar()now reachesfloat_ssim_hipandfloat_ms_ssim_hipthrough one header. scripts/ci/exact_twins.d/float_ssim.hipandfloat_ssim_lcs.hip(new) declare the twin exact (ADR-1428).- No Netflix golden-data, public API or FFmpeg patch impact. The CPU extractor is untouched.
ADR-1435 — vif_hip reads the CPU's log2 table (2026-10-01)¶
fix/hip-vif-cpu-log2-table, T-HIP-VIF-DEVICE-LOG2-2026-10-01.
core/src/feature/integer_vif.h(upstream-mirror) gainsvif_log2_table_generate()and#include <math.h>;core/src/feature/integer_vif.closes itsstatic inline log_generate()and calls the header's function ininit(). Same expression, moved. If upstream Netflix changeslog_generate(), apply the change tovif_log2_table_generate()and keepinteger_vif.cfree of a second copy (test_hip_vif_log2_table_contract.pyrejects one).core/src/feature/hip/integer_vif/vif_statistics.hip: the five horizontal kernels take one more argument,const uint16_t *log2_table, betweenvif_enhn_gain_limitandaccum_out;log_generate()islog2_lookup().core/src/feature/hip/integer_vif_hip.callocateslog2_table_dev, fills it invif_hip_tables_upload()and passes it in bothargs_hori[]lists. A rebase that keeps one side's kernel signature and the other's argument list launches the kernels with shifted arguments: keep both from this side.scripts/ci/exact_twins.d/vif.hip(new) declares the twin exact (ADR-1428). Another backend'svifdeclaration is another file; nothing shared is edited.- No Netflix golden-data, public API or FFmpeg patch impact. The CPU extractor's table and scores are unchanged.
ADR-1444 — float_vif_hip runs the CUDA twin's arithmetic from a shared header (2026-10-02)¶
fix/hip-float-vif-cpu-arithmetic, T-GPU-FLOAT-VIF-CPU-ARITHMETIC-2026-10-01 (HIP part), T-HIP-FLOAT-VIF-SMALL-FRAME-GPU-FAULT-2026-10-02.
core/src/feature/float_vif_gpu_common.h(new): the arithmetic and the kernel argument blocks that were incore/src/feature/cuda/float_vif/float_vif_device.h, unchanged, with the rounding operators as overridable macros and the blocks namedFloatVifGpu*. The CUDA header keeps theDEVICE_CODEmapping to__fmul_rn()and friends, includes the new header and aliasesFloatVifCuda*. A rebase that brings a change to the old header's arithmetic applies it to the new header instead; the CUDA header must not regain a copy.core/src/feature/hip/float_vif/float_vif_score.hipandcore/src/feature/hip/float_vif_hip.c: rewritten afterfloat_vif_score.cu/float_vif_cuda.c. Three kernels, each taking one argument block by value; the old discrete argument lists are gone. Keep kernel and host from the same side of a conflict.- Mirror list, same PR when the CPU side changes:
vif_get_filter(),VIF_OPT_FAST_LOG2/log2f_approx(),vif_pixel_statistic_s(),vif_statistic_s()invif_tools.cchangefloat_vif_gpu_common.h(test_float_vif_device_mathfails until it follows). scripts/ci/exact_twins.d/float_vif.hip(new) declares the twin exact (ADR-1428).core/test/test_cuda_float_vif_exact_contract.pyreads the shared header for the arithmetic checks and the CUDA header for the device spelling.- No Netflix golden-data, public API or FFmpeg patch impact. The CPU extractor's scores are unchanged;
float_vif_cudameasured bit-identical on an RTX 4090 after the split (48 of 48 and 50 of 50 frames).
ADR-1442 — float ADM divides; no reciprocal estimate (2026-10-02)¶
fix/float-adm-reference-divides, Research-1442, ADR-1442.
This is a deliberate divergence from upstream in two Netflix files. A sync must keep the fork's side of both:
core/src/feature/adm_options.h: upstream has#define ADM_OPT_RECIP_DIVISION. The fork has a comment in its place and no definition. Do not take the upstream line back.core/src/feature/adm_tools.c: upstream has, under__SSE2__and that macro,#include <emmintrin.h>,rcp_s()built on_mm_rcp_ss(), and#define DIVS(n, d) ((n) * rcp_s(d)). The fork has one#define DIVS(n, d) ((n) / (d))and an#errorwhen the macro is defined. Keep that block whole. If upstream changes how the decouple usesDIVS(), port the change with the quotient.core/src/feature/adm_reciprocal_model.{c,h}and the exportsadm_divs_is_reciprocal_s()/adm_divs_reciprocal_estimate_s()(ADR-1420) are deleted. Nothing may bring them back; a twin divides with its device's correctly rounded fp32 division.core/src/feature/cuda/float_adm/float_adm_device.h:fadm_divs(n, d)isFADM_FDIV(n, d),__fdiv_rn()on the device.float_adm_cuda.chas no probe, no table buffer and no upload.- Guards:
core/test/test_float_adm_divides_contract.py(the reference, every*float_adm*file undercore/src/feature/and the CUDA flags incore/src/meson.build) andtest_decouple_dividesincore/test/test_float_adm_device_math.c. Both fail on upstream's code. - Scores: x86
float_admmoves by up to 1.3e-7 against upstream and against earlier fork releases. The Netflix golden assertions hold unchanged (271 passed); no file underpython/test/is touched. - No public C API or FFmpeg patch impact. The SYCL, HIP and Metal twins are untouched (
T-GPU-FLOAT-ADM-CPU-ARITHMETIC-2026-10-01); an in-flight twin that includesadm_reciprocal_model.hno longer builds and has to divide.
test-netflix-golden checks for pytest presence (2026-10-02)¶
fix/golden-gate-pytest-check, closes T-TEST-NETFLIX-GOLDEN-PYTEST-MISSING-HINT-2026-10-02.
Makefile:test-netflix-goldenprobespython3 -m pytest --versionbefore invoking the test suite and fails with an actionable error directing the developer to.venv/bin/pip install pytest (see docs/development/languages.md).- Tests:
scripts/ci/tests/test_golden_gate_makefile_contract.py(test_test_netflix_golden_checks_pytest_presence). - No public API, ABI, SIMD/GPU twin or Netflix golden-data impact.
ADR-1445 — ssimulacra2_hip evaluates fp64 terms and forms the CPU's sums (2026-10-02)¶
fix/hip-ssimulacra2-cpu-sum-order, T-HIP-SSIMULACRA2-NOT-CPU-BITS-2026-10-01.
core/src/feature/hip/ssimulacra2/ssimulacra2_device.hip: the fp32 pair arithmetic (Ff,two_sum,ff_add,ff_div, ...) and the kernelsssimulacra2_combine_partials/ssimulacra2_combine_finalare gone. In their place:ss2h_terms()(the CPU's fp64 expressions) and the four kernelsssimulacra2_chunk_sums,_chunk_plan,_chunk_units,_ordered_totals, ported fromcuda/ssimulacra2/ssimulacra2_device.cu(ADR-1433). A change to one of the two files' sum kernels belongs in the other as well.core/src/feature/hip/ssimulacra2_hip.h:struct Ss2hCombineArgshas the CUDA layout (chunk_sums,plan,units,totals,pixels,chunks);struct Ss2hFinalArgsand theSS2H_REDUCE_WG/SS2H_MAX_GROUPS/SS2H_PAIRconstants are gone. Kernel and host must come from the same side of a conflict.core/src/feature/hip/ssimulacra2_hip.c: four device buffers for the sums (ss2h_alloc_sums()), four launches per scale, a readback of 108 doubles.- A change to
ssim_map()/edge_diff_map()inssimulacra2.cchangesss2h_terms()(andss2c_terms()) in the same PR. scripts/ci/exact_twins.d/ssimulacra2.hip(new) declares the twin exact (ADR-1428).
ADR-1447 — float_moment_hip adds the CPU's float squares (2026-10-02)¶
fix/hip-float-moment-cpu-squares, T-HIP-FLOAT-MOMENT-16BIT-SQUARES-2026-10-01.
core/src/feature/hip/float_moment/moment_score.hip: newmoment_float_square(); the 10/12/16-bit kernel adds it for the second sums instead ofr * r/d * d. The 8-bit kernel is unchanged. If upstream changes howmoment.c::compute_2nd_moment()forms or adds its term, change the kernel with it.core/src/feature/hip/float_moment_hip.c: comment only (the range of exactness and the bound past 2^53).core/test/test_hip_float_moment_parity.cis one table-driven binary now; the_10bitmeson variant is gone and a_largevariant is registered throughhip_parity_large_fixture_tests.scripts/ci/exact_twins.d/float_moment.hip(new) declares the twin exact (ADR-1428).- No Netflix golden-data, public API or FFmpeg patch impact.
vif_log2_table.h — one definition of the VIF log2 table for every backend (2026-10-02)¶
refactor/vif-log2-table-one-definition; follow-up of ADR-1435.
core/src/feature/vif_log2_table.h(new):VIF_LOG2_TABLE_SIZE,VIF_LOG2_TABLE_OFFSETandvif_log2_table_generate(), moved out ofcore/src/feature/integer_vif.h(upstream-mirror), which now includes it. The header is plain C that is also valid C++ and Objective-C++ and includes only<math.h>and<stdint.h>. If upstream Netflix changeslog_generate()or the table size, apply the change to this header.core/src/feature/sycl/integer_vif_sycl.cpp::vif_init_log2_lut()andcore/src/feature/metal/integer_vif_metal.mmcall the generator instead of their own loops; the Metal file's copy ofVIF_LOG2_TABLE_SIZEand itsfill_log2_table()are gone. A rebase that brings either loop back reintroduces a second definition;test_hip_vif_log2_table_contract.pyrejects it.core/test/test_integer_vif_log2.cbuilds its table with the generator.- No Netflix golden-data, public API or FFmpeg patch impact; no score changes.
ADR-1446 — ssimulacra2_sycl forms the CPU's fp64 terms in integers and the CPU's sums (2026-10-02)¶
fix/sycl-ssimulacra2-cpu-bits, the SYCL part of T-GPU-SSIMULACRA2-SUM-ORDER-2026-10-01.
core/src/feature/ordered_sum.h(shared with the CUDA and HIP twins): every function that took or returned adoublehas a_bitsform on the fp64 bit pattern, and the fp64 form is a wrapper around it. WithVMAF_ORDSUM_NO_FP64defined the header names nodouble. A change to an fp64 form belongs in its_bitsform; keep everydoubleinside#ifndef VMAF_ORDSUM_NO_FP64(core/test/test_sycl_ssimulacra2_exact_contract.pychecks it).core/src/feature/sycl/sycl_ssimulacra2_math.h(new): the six per-sample terms ofssim_map()/edge_diff_map()as fp64 bit patterns, computed in 64-bit integers. A change to those two functions inssimulacra2.cchanges this header andreference_terms()incore/test/test_sycl_ssimulacra2_math.cin the same PR (as it changesss2c_terms()andss2h_terms()).core/src/feature/sycl/sycl_ordered_sum.h(new): the plan from fp32 advice, kept terms and runs for chunks the walk adds term by term, and the walk, on bit patterns.core/src/feature/sycl/sycl_soft_signed.h:signed_from_float(),signed_abs()andkQuietNanBitsadded; nothing else changed.core/src/feature/sycl/ssimulacra2_sycl.cpp: stage 3c is rewritten. The kernelslaunch_combine_partials/launch_combine_finaland the bufferd_partialsare gone; in their placelaunch_chunk_sums,launch_chunk_plan,launch_chunk_units(three kernels) andlaunch_ordered_totals, and eight device buffers. The fp32 pair functions (ss2s_ssim_term(),ss2s_edge_terms()) stay, as advice for the plan only. The read-back is 108uint64_t. Take the whole stage from one side of a conflict.core/test/test_sycl_ssimulacra2_parity.cis rewritten as equality cases; newtest_sycl_ssimulacra2_mathandtest_sycl_ordered_sum(each with a SYCL probe TU built by acustom_target, liketest_sycl_integer_ssim_math) andtest_sycl_ssimulacra2_exact_contract.py.scripts/ci/exact_twins.d/ssimulacra2.sycl(new) declares the twin exact (ADR-1428);scripts/ci/gpu_ulp_calibration.yamlloses the Arc A380'sssimulacra2: 5.0e-2rows, which an exact twin never reads.- No Netflix golden-data, public API or FFmpeg patch impact.
ADR-1449 — float_moment_sycl adds the CPU's float squares (2026-10-02)¶
fix/sycl-float-moment-cpu-float-squares, T-SYCL-FLOAT-MOMENT-16BIT-SQUARES-2026-10-02.
core/src/feature/sycl/integer_moment_sycl.cpp: newmoment_float_square(); the kernel adds it for the second sums instead ofr * r/d * d, at every bit depth (it is the integer square up to 12 bits). The collect comment states the range of exactness and the bound past 2^53. If upstream changes howmoment.c::compute_2nd_moment()forms or adds its term, change the kernel with it.core/test/test_sycl_float_moment_parity.cis one binary of equality cases over the newcore/test/float_moment_twin_parity.h; the_10bitmeson variant is gone (the binary covers 8, 10, 12 and 16 bits) and the_largevariant stays registered throughsycl_parity_large_fixture_tests.core/test/test_sycl_float_moment_exact_contract.py(new) is device-free.scripts/ci/exact_twins.d/float_moment.sycl(new) declares the twin exact (ADR-1428).- No Netflix golden-data, public API or FFmpeg patch impact.
ADR-1450 — float_psnr_sycl adds its squared differences as integers (2026-10-02)¶
fix/sycl-float-psnr-exact-block-sums, T-SYCL-FLOAT-PSNR-FP32-BLOCK-SUMS-2026-10-02.
core/src/feature/sycl/float_psnr_sycl.cpp:fpsnr_pixel_noise()returns the fp32 square of the raw sample difference asuint64(fpsnr_inv_scaler()is gone,fpsnr_scaler()is its host counterpart);fpsnr_store_workgroup_sum(), the local accessor,d_partials/h_partialsandFpsnrOutput::partialsareuint64;collect_fex_sycl()adds integers and divides the total by scaler^2 and the pixel count. Take kernel and host from the same side of a conflict. If upstream changes howfloat_psnr.cforms or adds its term, change the kernel with it.core/test/test_sycl_float_psnr_parity.cis one binary of equality cases over the newcore/test/float_psnr_twin_parity.h;core/test/test_sycl_float_psnr_exact_contract.py(new) is device-free.scripts/ci/exact_twins.d/float_psnr.sycl(new) declares the twin exact (ADR-1428).- No Netflix golden-data, public API or FFmpeg patch impact.
ADR-1451 — six SYCL twins declared exact as a group (2026-10-02)¶
test/sycl-exact-twins-declared, T-SYCL-EXACT-TWINS-UNDECLARED-2026-10-02.
scripts/ci/exact_twins.d/gainsadm.sycl,motion.sycl,motion_debug.sycl,motion_v2.sycl,psnr.sycl,float_ssim.sycl,float_ssim_lcs.syclandcambi.sycl(ADR-1428): the gate compares those cells with tolerance 0. No twin's code changes.core/test/test_sycl_exact_twins.c(new) asserts==on every output of the six twins at 8 and 10 bits. A change to one of them, or to its CPU extractor, has to keep it passing.scripts/ci/gpu_ulp_calibration.yaml: notes only; the Arc A380'sfloat_ssim: 5.0e-4stays for cells whose other side is not exact.- No Netflix golden-data, public API or FFmpeg patch impact; no score changes.
ADR-1452 — speed_chroma on HIP gets the CUDA cell's log2f bound (2026-10-02)¶
test/hip-speed-chroma-libm-bound, ADR-1452.
scripts/ci/cross_backend_calibration.py:LIBM_TWINS["speed_chroma"]listshipat5e-6next tocuda. Keep both entries; a rebase that drops one puts that cell back at the general5e-5.core/test/speed_chroma_twin_parity.h(new): the 960x960 textured fixture, the CPU run and the comparison oftest_cuda_speed_chroma_paritymoved out of that test unchanged.test_cuda_speed_chroma_parity.candtest_hip_speed_chroma_parity.cwrap it. A change to the fixture or the bound is made in the header.test_hip_speed_chroma_parityexits 77 when it skips (no device or a scaffold); it passed before.- No source of a twin changes; no score changes. No Netflix golden-data, public API or FFmpeg patch impact.
ADR-1453 — float_moment_cuda adds the CPU's float squares (2026-10-02)¶
fix/cuda-float-moment-cpu-float-squares, T-CUDA-FLOAT-MOMENT-16BIT-SQUARES-2026-10-02.
core/src/feature/cuda/integer_moment/moment_score.cu: newmoment_float_square()andsample_square<T>();thread_sums()addssample_square<T>()for the second sums instead ofr * r/d * d. Auint16_tsample contributes the fp32 square, auint8_tsample the integer square (the same number at 8 bits). If upstream changes howmoment.c::compute_2nd_moment()forms or adds its term, change the kernel with it.core/src/feature/cuda/integer_moment_cuda.c:moment_cuda_scaler()holds the bit-depth scaler and the statement of the exact range and the bound past 2^53; the arithmetic ofcollect_fex_cuda()is unchanged.core/test/test_cuda_float_moment_parity.cis one binary of equality cases overcore/test/float_moment_twin_parity.h; the_10bitmeson variant is gone (the binary covers 8, 10, 12 and 16 bits) and the_largevariant stays registered throughcuda_parity_large_fixture_tests.core/test/test_cuda_float_moment_exact_contract.py(new) is device-free.scripts/ci/exact_twins.d/float_moment.cuda(new) declares the twin exact
ADR-1455 — float_psnr_cuda adds its squared differences as integers (2026-10-02)¶
fix/cuda-float-psnr-exact-block-sums, T-CUDA-FLOAT-PSNR-FP32-BLOCK-SUMS-2026-10-02.
core/src/feature/cuda/float_psnr/float_psnr_score.cu:fpsnr_square()returns the fp32 square of the raw sample difference asunsigned long long;fpsnr_block_sum()and the per-block partials are 64-bit integers; one templatedfpsnr_block<T>()is the body of both kernels, and the 16bpc kernel no longer takesbpc.core/src/feature/cuda/float_psnr_cuda.c: the readback is oneuint64per block (partials_bytes), both kernels are launched with the same seven arguments, andfloat_psnr_noise()adds integers and divides the total by scaler^2 and the pixel count. Take kernel and host from the same side of a conflict. If upstream changes howfloat_psnr.cforms or adds its term, change the kernel with it.core/test/test_cuda_float_psnr_parity.cis one binary of equality cases overcore/test/float_psnr_twin_parity.h;core/test/test_cuda_float_psnr_exact_contract.py(new) is device-free.scripts/ci/exact_twins.d/float_psnr.cuda(new) declares the twin exact (ADR-1428).
ADR-1454 — scripts/ci/AGENTS.md is a generated index over AGENTS.d/ pages (2026-10-02)¶
docs/agents-index-topic-pages, opens T-AGENTS-INDEX-MIGRATION-2026-10-02.
- No rebase impact from upstream: an upstream sync never touches an
AGENTS.mdor anAGENTS.d/; Netflix/vmaf has neither. - Fork branches that append to
scripts/ci/AGENTS.mdconflict once. Take master's side ofscripts/ci/AGENTS.md, put the added text into the page whoseTouchingrow matches the files (or a new page underscripts/ci/AGENTS.d/), thenmake docs-fragments-write. scripts/docs/agents_index.py(new) renders everyAGENTS.mdnext to anAGENTS.d/;make docs-fragments-checkand thecheck-generated-docshook run it.scripts/docs/agents_migration_check.py(new) is the one-off proof for a migration pull request..pre-commit-config.yaml: the hooks triggered byscripts/ci/AGENTS.md(test-codex-hook-config,check-research-digest-ids,test-research-digest-ids) also name the page that now holds their text.- No Netflix golden-data, public API or FFmpeg patch impact.
motion_cuda emits the CPU's SAD score (2026-10-02)¶
fix/cuda-motion-sad-score, T-CUDA-MOTION-SAD-SCORE-NOT-EMITTED-2026-10-02.
core/src/feature/cuda/integer_motion_cuda.c:provided_featureslistsVMAF_integer_feature_motion_sad_scorefirst, asinteger_motion.cdoes, andextract_force_zero(),motion_collect_first_frame()andemit_batch_scores()append it on every frame. The value is the one the debugmotion_scorealready carried. If upstream adds, renames or drops an output ofinteger_motion.c, change the twin's list and its three append sites with it.core/test/test_cuda_motion_sad_score.c(new) compares every output of eleven frames with==under four option sets.- No kernel change, no Netflix golden-data, public API or FFmpeg patch impact. The JSON / XML of a
--backend cuda --feature motionrun gains the key the CPU run has.
ADR-1456 — vif_cuda's device logarithm is pinned to the CPU's table (2026-10-02)¶
fix/cuda-vif-cpu-log2-table, T-CUDA-VIF-DEVICE-LOG2-UNPROBED-2026-10-02.
core/src/feature/cuda/integer_vif/vif_log2_probe.cu(new): the kernelvif_log2_table_probe, its own entry incuda_cu_sources(core/src/meson.build), launched only bycore/test/test_cuda_vif_log2_table.c.filter1d.cuis untouched.core/src/feature/cuda/integer_vif/vif_statistics.cuh:log_generate()is unchanged in code (a dead commented-out range check is gone) and documented as the mirror ofvif_log2_table_generate(). If upstream changes either expression, change the other and re-run the device test. Keeplog_generate()an inline function of this header: the probe includes it.core/src/feature/vif_log2_table.h: comment only.core/test/test_cuda_vif_log2_table.c(device) andcore/test/test_cuda_vif_log2_contract.py(device-free) are new;EXPECTED_CUDA_TARGET_COUNTincore/test/test_device_target_header_dependencies.pyis 22.scripts/ci/exact_twins.d/vif.cuda(new) declares the twin exact (ADR-1428).- No scoring kernel, Netflix golden-data, public API or FFmpeg patch impact.
ADR-1448 — ciede_hip runs the SYCL twin's fp32-pair arithmetic from shared headers (2026-10-02)¶
fix/hip-ciede-cpu-arithmetic, T-GPU-CIEDE-CPU-ARITHMETIC-2026-10-01 (HIP part).
core/src/feature/ff_math.handcore/src/feature/ciede_ff_math.h(new): the pair functions and the ciede2000 statements that werecore/src/feature/sycl/sycl_ff_math.handsycl_ciede_math.h, unchanged apart fromsycl::functions becomingVMAF_FF_*macros, the namespaces becomingvmaf_ffm/vmaf_ciede_ff, andmake_constants()beingconstexpr. The two SYCL headers keep their names, define the SYCL primitives, include the shared headers and alias the old namespaces. A rebase that brings a change to the old SYCL headers' arithmetic applies it to the shared headers; the SYCL headers must not regain function definitions (test_sycl_ciede_exact_contract.pyrejects that).scripts/dev/gen_sycl_ff_math.pywrites the generated block ofcore/src/feature/ff_math.h.core/src/feature/ff_pair.h(new): the exact pair operations for a backend with IEEE fp32 operators;core/src/feature/hip/integer_ciede/ciede_hip_math.h(new): the HIP primitives.core/src/feature/ciede_frame_sum.h(new):ciede_frame_sum(), moved out ofcuda/integer_ciede/ciede_device.handsycl/integer_ciede_sycl.cpp, which include it.core/src/feature/hip/integer_ciede/ciede_score.hipandcore/src/feature/hip/ciede_hip.c: rewritten. The kernels take the planes of each picture as oneCiedeHipPlanesblock by value, atermspointer and the index of the bit depth; the readback is one float per pixel. Keep kernel and host from the same side of a conflict.core/src/meson.build:hip_kernel_extra_argsgivesciede_score-std=c++20; the shared headers are inhip_kernel_shared_headers.scripts/ci/cross_backend_calibration.py:LIBM_TWINS["ciede"]gains"hip": 1e-9;scripts/ci/test_cross_backend_parity_gate.pyfollows. A conflict with another lane's entry in that literal keeps both entries.core/test/test_hip_ciede_parity.cwrapsciede_twin_parity.h; the_oddwmeson variant is gone (the shared cases include 577x325).test_hip_ciede_mathistest_sycl_ciede_math.cbuilt withVMAF_TEST_CIEDE_MATH_HOST_ONLYandtest_hip_ciede_math_probe.cpp.- Mirror list, same PR when the CPU side changes:
get_lab_color(),ciede2000(),get_r_sub_t()and the order ofextract()'s sum inciede.cchangeciede_ff_math.hand the CUDA twin'sciede_device.h. - No Netflix golden-data, public API or FFmpeg patch impact.
ciede_syclmeasured bit-identical before and after on an Arc A380 (178 frames);ciede_cudaunchanged on an RTX 4090.
ADR-1457 — six CUDA twins declared exact as a group (2026-10-02)¶
test/cuda-exact-twins-declared, T-CUDA-EXACT-TWINS-UNDECLARED-2026-10-02.
scripts/ci/exact_twins.d/gainsmotion.cuda,motion_debug.cuda,motion_v2.cuda,psnr.cuda,float_ssim.cuda,float_ssim_lcs.cuda,float_ms_ssim.cuda,float_ms_ssim_lcs.cudaandcambi.cuda(ADR-1428): the gate compares those cells with tolerance 0. No twin's code changes.core/test/test_cuda_exact_twins.c(new) asserts==on every output of the six twins at 8 and 10 bits. A change to one of them, or to its CPU extractor, has to keep it passing.docs/research/1457-cuda-twin-exactness-sweep.mdholds the sweep of all 21 gate features.- No Netflix golden-data, public API or FFmpeg patch impact; no score changes.
speed_chroma SYCL cell is a libm twin at 5e-6 (2026-10-02)¶
test/sycl-speed-chroma-libm-bound, T-SYCL-SPEED-CHROMA-GATE-DEFAULT-TOLERANCE-2026-10-02.
scripts/ci/cross_backend_calibration.py:LIBM_TWINS["speed_chroma"]gains"sycl": 5e-6. When this dictionary conflicts in a rebase, keep every backend of both sides; the entry for a backend is never dropped to resolve a conflict.scripts/ci/test_cross_backend_parity_gate.py:test_speed_chroma_sycl_cell_is_bounded_by_the_cpu_log2f; the CUDA and HIP tests no longer use the SYCL cell as their example of an unlisted twin.- Do not replace the entry by an exact-twin fragment because an icx run shows 0: the twin equals an icx CPU and differs from a glibc CPU.
- No library code, score, Netflix golden-data, public API or FFmpeg patch impact.
ADR-1465 — float_ms_ssim_cuda adds the terms of every scale in the CPU's raster order (2026-10-02)¶
fix/cuda-float-ms-ssim-raster-order-sum, T-CUDA-FLOAT-MS-SSIM-FRAME-SUM-ORDER-2026-10-02.
core/src/feature/cuda/integer_ms_ssim/ms_ssim_score.cu:ms_ssim_vert_lcskeeps its name and parameter list and stores terms instead of reducing them:landcas doubles,sas a float, aty * w_final + x. The shuffle loop and the shared warp arrays are gone: do not take them back from an older branch, per-block sums are the defect.core/src/feature/cuda/integer_ms_ssim_cuda.c: the per-scale buffers are term planes (l_terms,c_terms,s_termsand their pinned host copies,scale_window_count) where they were block partials;ms_ssim_scale_sums()is the only place terms are added.- When
iqa/ssim_tools.c::iqa_ssim()or its accumulate functions change the order or the number of their sums, change these two files in the same PR. core/test/float_ms_ssim_order_frame.his added byte-identically by the HIP, CUDA and SYCL lanes. A rebase that sees it added twice keeps one copy and never merges edits into it;test_cuda_float_ms_ssim_exact_contract.pyholds its sha256. The second frame oftest_cuda_float_ms_ssim_order.cis rebuilt from a formula in that file.core/test/meson.build:test_cuda_float_ms_ssim_order(device) andtest_cuda_float_ms_ssim_exact_contract(device-free), one block aftertest_cuda_float_ms_ssim_parity.scripts/ci/exact_twins.d/float_ms_ssim.cuda,float_ms_ssim_lcs.cuda:adr:names ADR-1465 in place of ADR-1457.- No Netflix golden-data, public API or FFmpeg patch impact. Stored
float_ms_ssim_cudascores can move in their last digits on rare frames.
ADR-1464 — float_ssim_cuda adds its frame sums in the CPU's raster order (2026-10-02)¶
fix/cuda-float-ssim-raster-order-sum, T-CUDA-FLOAT-SSIM-FRAME-SUM-ORDER-2026-10-02.
core/src/feature/cuda/integer_ssim/ssim_score.cu: the two pass-2 kernels keep their names and store terms instead of reducing them.calculate_ssim_vert_combinewrites one double per window aty * w_final + x;calculate_ssim_vert_combine_lcswritesLCS_TERMS(4) per window.block_sum()and the shared warp array are gone: do not take them back from an older branch, a per-block sum is the defect.core/src/feature/cuda/integer_ssim_cuda.c: one read-back (rb) ofn_windows * n_sumsdoubles replaces the partials andrb_lcs;float_ssim_frame_sum()andfloat_ssim_frame_sums_lcs()are the only places terms are added. Both kernels take the same parameter list now.- When
iqa/ssim_tools.c::iqa_ssim()or its accumulate functions change the order or the number of their sums, change these two files in the same PR. core/test/float_ssim_order_frame.his added byte-identically by the CUDA, HIP and SYCL lanes. A rebase that sees it added twice keeps one copy and never merges edits into it;test_cuda_float_ssim_exact_contract.pyholds its sha256.core/test/meson.build:test_cuda_float_ssim_order(device) andtest_cuda_float_ssim_exact_contract(device-free), one block aftertest_cuda_float_ssim_parity.scripts/ci/exact_twins.d/float_ssim.cuda,float_ssim_lcs.cuda:adr:names ADR-1464 in place of ADR-1457, whose "exact up to one rounding" no longer describes the twin.- No Netflix golden-data, public API or FFmpeg patch impact. Stored
float_ssim_cudascores can move by one float step on rare frames.
ADR-1460 — speed_temporal is a parity-gate feature; registry coverage test (2026-10-02)¶
test/gate-speed-temporal, T-GATE-SPEED-TEMPORAL-UNGATED-2026-10-02.
scripts/ci/cross_backend_parity_gate.pyandscripts/ci/cross_backend_vif_diff.py:speed_temporalinFEATURE_METRICS(andFEATURE_TOLERANCE);psnrlistspsnr_y,psnr_cbandpsnr_crin both; the single-feature gate gainsssimand its threeinteger_ssim_<backend>aliases. The two scripts' tables are now equal and a test keeps them so.scripts/ci/cross_backend_calibration.py:LIBM_TWINS["speed_temporal"] = {"cuda": 4e-5, "hip": 4e-5, "sycl": 4e-5}.core/test/test_parity_gate_covers_registered_twins.py(new, in the fast suite) readscore/src/feature/feature_extractor.cpp. A sync or a new backend that registers a twin has to give it a gate feature, or the test fails with the twin's name.- No library code, Netflix golden-data, public API or FFmpeg patch impact.
ADR-1458 — float_adm_hip runs the CUDA twin's arithmetic from a shared header (2026-10-02)¶
fix/hip-float-adm-cpu-arithmetic, T-HIP-FLOAT-ADM-NOT-CPU-ARITHMETIC-2026-10-01.
core/src/feature/float_adm_gpu_common.h(new): the arithmetic and the argument blocks that were incore/src/feature/cuda/float_adm/float_adm_device.h, unchanged, with the rounding macros,FADM_HD, the bit casts andFADM_POWFoverridable and the blocks namedFloatAdmGpu*. The CUDA header keeps theDEVICE_CODEdefinitions, includes the new header and typedefs theFloatAdmCuda*names. A rebase that brings a change to the old header's arithmetic applies it to the new header; the CUDA header must not regain a copy (test_cuda_float_adm_exact_contract.pyrejects that).core/src/feature/hip/float_adm/float_adm_hip_math.h(new): the HIP spelling, plain operators.core/src/feature/hip/float_adm/float_adm_score.hipandcore/src/feature/hip/float_adm_hip.c: rewritten afterfloat_adm_score.cu/float_adm_cuda.c(five kernels, row sums, the reference's routines). Keep kernel and host from the same side of a conflict.float_adm_hipgains the optionadm_skip_aim_scaleand refuses frames below 17x17.scripts/ci/exact_twins.d/float_adm.hip(new) declares the twin exact (ADR-1428).core/test/test_hip_float_adm_parity.cwrapsfloat_adm_twin_parity.h;core/test/test_hip_float_adm_math.cwith its probe kerneltest_hip_float_adm_math_probe.hipandhip_float_adm_math_sample.h(new) compare the device's arithmetic with the host's;core/test/test_hip_float_adm_exact_contract.py(new) is device-free.test_float_adm_divides_contract.pyandtest_cuda_float_adm_exact_contract.pyread the shared header.- Mirror list, same PR when the CPU side changes:
adm_decouple_s(),adm_csf_s(),adm_cm_thresh3x3_s(),adm_csf_den_scale_s()andadm_cm_s()inadm_tools.cchangefloat_adm_gpu_common.h(test_float_adm_device_mathfails until it follows). - No Netflix golden-data, public API or FFmpeg patch impact.
float_adm_cudare-measured unchanged on an RTX 4090 after the split.
ADR-1454 — ai/AGENTS.md is a generated index over AGENTS.d/ pages (2026-10-02)¶
docs/agents-index-ai, T-AGENTS-INDEX-MIGRATION-2026-10-02.
- No rebase impact from upstream: an upstream sync never touches an
AGENTS.mdor anAGENTS.d/. - Fork branches that append to
ai/AGENTS.mdconflict once: take master's side of that file, put the added text into the page whoseTouchingrow matches the files, thenmake docs-fragments-write. - No Netflix golden-data, public API or FFmpeg patch impact.
ADR-1454 — core/src/AGENTS.md is a generated index over AGENTS.d/ pages (2026-10-02)¶
docs/agents-index-core-src, T-AGENTS-INDEX-MIGRATION-2026-10-02.
- No rebase impact from upstream: an upstream sync never touches an
AGENTS.mdor anAGENTS.d/. - Fork branches that append to
core/src/AGENTS.mdconflict once: take master's side of that file, put the added text into the page whoseTouchingrow matches the files, thenmake docs-fragments-write. - No Netflix golden-data, public API or FFmpeg patch impact.
ADR-1462 — vif_cuda reads the CPU's log2 table (2026-10-02)¶
fix/cuda-vif-reads-host-log2-table, T-CUDA-VIF-DEVICE-LOG2-HOST-DEPENDENT-2026-10-02. Supersedes the ADR-1456 entry above.
core/src/feature/cuda/integer_vif/vif_statistics.cuh:log_generate()is gone. The module globalvif_cuda_log2_table,log2_lookup()and the kernelvif_cuda_log2_table_transferreplace it; the statistic's three logarithm sites calllog2_lookup(). When upstream changes this header, keep the lookup: do not take a devicelog2f()back.core/src/feature/cuda/integer_vif/filter1d.cuis untouched (it includes the header);vif_log2_probe.cuand itscuda_cu_sourcesentry are deleted,EXPECTED_CUDA_TARGET_COUNTis 21 again.core/src/feature/cuda/integer_vif_cuda.c/.h:vmaf_cuda_vif_log2_table_transfer()andvmaf_cuda_vif_upload_log2_table();init_fex_cuda()calls the upload right aftervif_init_cuda_context().core/src/feature/vif_log2_table.h: comment only (the rule is "no twin computes the table on its device" again).core/test/test_cuda_vif_log2_table.ctests the upload on a device;core/test/test_cuda_vif_log2_contract.pyis rewritten for the lookup.scripts/ci/silent-revert-allowlist.jsoncarries two ADR-1462 entries for the removal of the probe fatbin: areverse-hunkentry forcore/src/meson.buildand arewindentry forcore/test/test_device_target_header_dependencies.py. They are expiring declarations: once this change is onmasterthe findings are gone and the entries can be removed. Therewindentry names blob ids, so a rebase over a later change to that file makes it unused, not wrong.- No score, Netflix golden-data, public API or FFmpeg patch impact.
ADR-1454 — core/AGENTS.md is a generated index over AGENTS.d/ pages (2026-10-02)¶
docs/agents-index-core, T-AGENTS-INDEX-MIGRATION-2026-10-02.
- No rebase impact from upstream: an upstream sync never touches an
AGENTS.mdor anAGENTS.d/. - Fork branches that append to
core/AGENTS.mdconflict once: take master's side of that file, put the added text into the page whoseTouchingrow matches the files, thenmake docs-fragments-write. - No Netflix golden-data, public API or FFmpeg patch impact.
motion_sycl emits the SAD score and honours motion_force_zero (2026-10-02)¶
fix/sycl-motion-sad-score, T-GPU-MOTION-SAD-SCORE-NOT-EMITTED-2026-10-02 (SYCL part), T-SYCL-MOTION-FORCE-ZERO-IGNORED-2026-10-02.
core/src/feature/sycl/integer_motion_sycl.cpp:provided_featureslistsVMAF_integer_feature_motion_sad_scorefirst, asinteger_motion.cdoes, andmotion_append_sad_score()appends it on every frame at the three collect sites (the debugmotion_scorerepeats it).motion_force_zeromoved out ofextract_fex_sycl(), which libvmaf never calls for a SYCL extractor, intosubmit/collect/flush; under the optioninitallocates nothing on the device and does not register with the combined graph. Keep that pairing on a conflict: a registered extractor has to callvmaf_sycl_graph_submit()every frame, an unregistered one must not. Scores ofmodel/other_models/vmaf_v0.6.1mfz.jsonon--backend syclchange to the CPU's (72.321 instead of 76.668 on the Netflix pair).scripts/ci/cross_backend_parity_gate.pyandscripts/ci/cross_backend_vif_diff.py: themotionandmotion_debugtuples ofFEATURE_METRICSstart withVMAF_integer_feature_motion_sad_score. A twin without the output is a cellERROR(ADR-1418).motion_cuda(#1809) andmotion_hipemit it.core/test/test_sycl_motion_sad_score.c(new) compares every output of eleven frames with==at 8 and 10 bits under four option sets and requires that the twin has no output the CPU lacks;core/test/test_sycl_exact_twins.ccompares the SAD score in its two motion cases.- A new output or emit site in
integer_motion.c::extractchanges the SYCL twin in the same PR. - No Netflix golden-data, public C API or FFmpeg patch impact: the output name exists on the CPU already and the FFmpeg filter reads none of it.
float_ssim_hip adds the frame sum in the CPU's raster order (2026-10-02)¶
fix/hip-float-ssim-cpu-frame-sum, T-GPU-FLOAT-SSIM-FRAME-SUM-ORDER-2026-10-02 (HIP part); construction of ADR-1438, window arithmetic of ADR-1441 unchanged.
core/src/feature/hip/float_ssim/ssim_score.hip:SSIM_BLOCK_SIZE,ssim_block_sum()andssim_store_block_sum()are gone.calculate_ssim_hip_vert_combinestores onedoubleper window intermsaty * w_final + x; its argument list is the five pass-1 planes,terms,w_horiz,w_final,h_final,c1,c2.calculate_ssim_hip_vert_combine_lcstakestermsandlcs_terms(three planes ofw_final * h_finaldoubles,[l | c | s]) and no longer a partial count. A rebase that restores a per-block or per-wave sum makes the twin inexact again;test_hip_kernel_source_contract.pyrejects it.core/src/feature/hip/float_ssim_hip.c:partials_capacityandpartials_countare gone;windowssizesrb(windowsdoubles) andrb_lcs(3 * windows).fssim_hip_frame_sums()adds the windows in ascending order, onedoublechain per sum. Keep kernel and host from the same side of a conflict.core/test/float_ssim_order_frame.h(new) is the constructed 64x64 pair, added byte-identically by the CUDA, HIP and SYCL changes of the same row (sha2566dee502f6f583ea5...). Take either side of an add/add conflict; do not edit the file.core/test/test_hip_float_ssim_parity.cgainstest_float_ssim_frame_sum_orderand anenable_lcscase atscale=1.scripts/ci/exact_twins.d/float_ssim.hipandfloat_ssim_lcs.hip: evidence line only.- No Netflix golden-data, public API, CLI or FFmpeg patch impact. The CPU extractor is not touched.
ADR-1454 — dev/AGENTS.md is a generated index over AGENTS.d/ pages (2026-10-02)¶
docs/agents-index-dev, T-AGENTS-INDEX-MIGRATION-2026-10-02.
- No rebase impact from upstream: an upstream sync never touches an
AGENTS.mdor anAGENTS.d/. - Fork branches that append to
dev/AGENTS.mdconflict once: takeAGENTS.mdfrom master, put the new rule into a new or matching page underdev/AGENTS.d/, runmake docs-fragments-write. - No Netflix golden-data, public API or FFmpeg patch impact.
float_ms_ssim_hip adds the per-scale sums in the CPU's raster order (2026-10-02)¶
fix/hip-float-ms-ssim-cpu-frame-sum, T-GPU-FLOAT-SSIM-FRAME-SUM-ORDER-2026-10-02 (HIP float_ms_ssim part); construction of ADR-1438, sample arithmetic of ADR-1403 unchanged.
core/src/feature/hip/integer_ms_ssim/ms_ssim_score.hip:ms_ssim_vert_lcshas 12 arguments instead of 14: the three partial pointers are onedouble *terms(three planes ofw_final * h_finaldoubles,[l | c | s], raster order). The wave and block reduction,BLOCK_*,MIN_WARP_SIZEandWARPS_PER_BLOCKare gone. A rebase that restores a sum on the device makes the twin inexact again;test_hip_kernel_source_contract.pyrejects it.core/src/feature/hip/integer_ms_ssim_hip.c:scale_block_count,l_partials/c_partials/s_partialsand theirh_copies are gone;scale_windows[i],terms[i]andh_terms[i]replace them (ms_ssim_terms_bytes()),ms_ssim_alloc_partials()/ms_ssim_unwind_partials()arems_ssim_alloc_terms()/ms_ssim_unwind_terms(), and the pinned planes arehipHostMallocDefaultinstead of write-combined (the host reads them).ms_ssim_hip_scale_sums()adds the windows in ascending order. Keep kernel and host from the same side of a conflict.core/test/float_ms_ssim_order_frame.h(new, 389 kB): the luma planes of the constructed 176x176 pair. Data, not code: never edit it; a twin test of another backend includes this file instead of adding a copy.core/test/test_hip_ms_ssim_parity.cgainstest_ms_ssim_frame_sum_order.scripts/ci/exact_twins.d/float_ms_ssim.hipandfloat_ms_ssim_lcs.hip: evidence line and ADR list.- No Netflix golden-data, public API, CLI or FFmpeg patch impact. The CPU extractor is not touched.
ADR-1454 — core/tools/AGENTS.md is a generated index over AGENTS.d/ pages (2026-10-02)¶
docs/agents-index-core-tools, T-AGENTS-INDEX-MIGRATION-2026-10-02.
- No rebase impact from upstream: an upstream sync never touches an
AGENTS.mdor anAGENTS.d/. - Fork branches that append to
core/tools/AGENTS.mdconflict once: take
ADR-1454 — tools/vmaf-tune/AGENTS.md is a generated index over AGENTS.d/ pages (2026-10-02)¶
docs/agents-index-vmaf-tune, T-AGENTS-INDEX-MIGRATION-2026-10-02.
- No rebase impact from upstream: an upstream sync never touches an
AGENTS.mdor anAGENTS.d/. - Fork branches that append to
tools/vmaf-tune/AGENTS.mdconflict once: take
ADR-1454 — core/test/AGENTS.md is a generated index over AGENTS.d/ pages (2026-10-02)¶
docs/agents-index-core-test, T-AGENTS-INDEX-MIGRATION-2026-10-02.
- No rebase impact from upstream: an upstream sync never touches an
AGENTS.mdor anAGENTS.d/. - Fork branches that append to
core/test/AGENTS.mdconflict once: take master's side of that file, put the added text into the page whoseTouchingrow matches the files, thenmake docs-fragments-write. - No Netflix golden-data, public API or FFmpeg patch impact.
SYCL kernels require sub-group size 16 or 32 (ADR-1468, 2026-10-02)¶
fix/sycl-aot-xe2-subgroup-size, T-SYCL-AOT-XE2-SUB-GROUP-SIZE-8-2026-10-02.
core/src/feature/sycl/sycl_compat.h:VmafSyclSubGroupSize<N>static_assertsN == 16 || N == 32;VMAF_SYCL_REQD_SG_SIZE(N)andVmafSyclKernelShape<SG, GRF>go through it. A rebase that brings a kernel with size 8, or a raw sub-group attribute or property, fails to compile or failscore/test/test_sycl_sub_group_size_contract.py: set the kernel to 16 and re-measure it (scratch audit, parity test).float_motion_sycl.cpp,float_adm_sycl.cpp,float_vif_sycl.cpp(row kernels),ssimulacra2_sycl.cpp(SS2S_WALK_SG),core/test/test_sycl_float_adm_math_probe.cppandcore/test/test_sycl_ordered_sum_probe.cpp: 8 became 16; the probe's work-group is 16 items so that the walk stays alone in its sub-group.core/test/sycl_aot_targets.py(new): the default targets and the measured sizes per family. A target added tosycl_icpx_aot_targetsincore/meson_options.txtneeds an entry.core/test/test_sycl_sub_group_size_contract.py(new, suitefast) andcore/test/test_sycl_aot_default_targets.py(new, suitesycl-aot, registered incore/test/meson.buildfor icpx builds). The second reads the SYCL compile commands frombuild.ninja; if the waycore/src/meson.buildspells them changes (-fsycl, the-devicelist,-MD -MF,-o), itsfor_targets()changes with it.- No score, public C API, Netflix golden-data or FFmpeg patch impact.
ADR-1454 — .github/AGENTS.md is a generated index over AGENTS.d/ pages (2026-10-02)¶
docs/agents-index-github, T-AGENTS-INDEX-MIGRATION-2026-10-02.
- No rebase impact from upstream: an upstream sync never touches an
AGENTS.mdor anAGENTS.d/. - Fork branches that append to
.github/AGENTS.mdconflict once: take.github/AGENTS.mdfrom master, put the new rule into a new or matching page under.github/AGENTS.d/, runmake docs-fragments-write.
ADR-1454 — core/src/feature/x86/AGENTS.md is a generated index over AGENTS.d/ pages (2026-10-02)¶
docs/agents-index-feature-x86, T-AGENTS-INDEX-MIGRATION-2026-10-02.
- No rebase impact from upstream: an upstream sync never touches an
AGENTS.mdor anAGENTS.d/. - Fork branches that append to
core/src/feature/x86/AGENTS.mdconflict once: take
ADR-1454 — core/src/dnn/AGENTS.md is a generated index over AGENTS.d/ pages (2026-10-02)¶
docs/agents-index-dnn, T-AGENTS-INDEX-MIGRATION-2026-10-02.
- No rebase impact from upstream: an upstream sync never touches an
AGENTS.mdor anAGENTS.d/. - Fork branches that append to
core/src/dnn/AGENTS.mdconflict once: take master's side of that file, put the added text into the page whoseTouchingrow matches the files, thenmake docs-fragments-write. - No Netflix golden-data, public API or FFmpeg patch impact.
float_ssim_sycl adds the CPU's terms in the CPU's order (ADR-1463, 2026-10-02)¶
fix/sycl-float-ssim-raster-sum, the SYCL float_ssim part of T-GPU-FLOAT-SSIM-FRAME-SUM-ORDER-2026-10-02.
core/src/feature/sycl/sycl_ssim_terms.h:ssim_float_parts()is the fp32 part that was insidessim_terms()(which now calls it and is used byfloat_ms_ssim_syclalone);ssim_double_terms()andssim_product_bits()form the CPU'slv,cvand(lv * cv) * svas fp64 bit patterns onsycl_soft_signed.h;ssim_frame_sums()/accumulate_window()are the host's sums. The header mirrorsiqa/ssim_accumulate_lane.hand the means ofiqa/ssim_tools.c: an upstream change to those lines changes the header in the same PR (core/test/test_sycl_float_ssim_exact_contract.pyfails until it does).core/src/feature/sycl/integer_ssim_sycl.cpp: the float twin's pass 2 has no reduction.FloatSsimTermKernel/FloatSsimLcsKernelstore one term (orlv,cv,sv) per window at its raster position; the work-group partials (d_partials,d_lcs_partials,store_fixed_group(),launch_vert_combine*(),sum_partials()) are gone.collectadds the read-back planes in index order.integer_ssim_frame_sum()becameframe_sum_of_terms(), shared by both twins of the file. Keep kernels and host sums from the same side of a conflict.vmaf_sycl_float_ssim_host_means()takesdouble means[5]now (SSIM as the default kernel forms it, L, C, S, SSIM as theenable_lcspath forms it). It is a test hook, not public API.core/test/float_ssim_order_frame.h(new) is the constructed pair, shared byte for byte with the CUDA and HIP tests (sha2566dee502f6f583ea5...). Never edit it; a lane that lands the same file later takes either side.core/test/test_sycl_float_ssim_parity.ccompares with==and has the order cases;core/test/test_sycl_float_ssim_exact_contract.py(new) is device-free;test_sycl_kernel_source_contract.pyandtest_sycl_ssim_exact_contract.pyread the renamed pieces.scripts/ci/exact_twins.d/float_ssim.syclandfloat_ssim_lcs.syclcite ADR-1463 and name the constructed frame.- No Netflix golden-data, public C API or FFmpeg patch impact. Stored
float_ssim_syclscores change only on frames whose mean lies next to afloatrounding boundary, by onefloatstep.
ADR-1454 — core/src/hip/AGENTS.md is a generated index over AGENTS.d/ pages (2026-10-02)¶
docs/agents-index-hip-runtime, T-AGENTS-INDEX-MIGRATION-2026-10-02.
- No rebase impact from upstream: an upstream sync never touches an
AGENTS.mdor anAGENTS.d/. - Fork branches that append to
core/src/hip/AGENTS.mdconflict once: takecore/src/hip/AGENTS.mdfrom master, put the new rule into a new or matching page undercore/src/hip/AGENTS.d/, runmake docs-fragments-write. - No Netflix golden-data, public API or FFmpeg patch impact.
SYCL runtime lint cleanup: three invariants kept (2026-10-02)¶
refactor/std-sycl-a (PR #1837), T-SYCL-VA-IMPORT-DETILE-EXCEPTION-2026-10-02.
core/src/sycl/common.cpp:VmafSyclStatelost its constructor and is an aggregate;vmaf_sycl_state_init()initialisesqueueandcopy_queuewith designated initialisers in the new-expression. The two queues stay the first two members. A rebase must not turn this into a default construction followed by assignments (queues on the default device), nor bring the constructor back (38 clang-tidy findings).core/src/sycl/picture_sycl.h,common.h,dmabuf_import.h: one definition per type for C and C++ (plain C form, cited NOLINT). Do not re-introduce#ifdef __cpluspluspairs, and never a C++-only enum underlying type.core/src/sycl/dmabuf_import.cpp:vmaf_sycl_import_va_surface()and the readback are split into helpers; the de-tile kernels live indetile_tile4()/detile_y_tiled()with unchanged bodies and captures;dispatch_detile()catches a throwing submit (detile_submit_failed()). The Level Zero descriptors use designated initialisers.core/src/sycl/d3d11_import.cpp: the threegotos became helper functions (Windows only; not compiled on the Linux lanes).core/test/test_sycl_runtime_contract.py(new, suitefast) holds the three invariants without a device.- No score, public C API, Netflix golden-data or FFmpeg patch impact.
One enum definition for C and C++ (ADR-1470, 2026-10-02)¶
fix/c-cxx-enum-one-definition, T-ENUM-CXX-ONLY-UNDERLYING-TYPE-2026-10-02.
core/src/feature/nonfinite_score.h:VmafVifNameSetis one plaintypedef enumfor both languages (it was: unsigned charunder__cplusplus), inside a citedNOLINTBEGIN(modernize-use-using, performance-enum-size)block. A rebase or a lint pass that re-adds a C++-only underlying type failscore/test/test_c_cxx_enum_definition_contract.py.core/src/model.handcore/src/feature/luminance_tools.hare unchanged: their: unsigned intC++ heads are size-compatible and pinned by the..._ABI_UINT_MAX = UINT_MAXenumerators. Keep those enumerators.- No score, output, public C API, Netflix golden-data or FFmpeg patch impact.
CUDA include guards renamed (standards batch B5, 2026-10-02)¶
refactor/b5-cuda-host-standards, ADR-1142.
core/src/cuda/picture_cuda.h: the guard__VMAF_SRC_CUDA_PICTURE_CUDA_H__isVMAF_SRC_CUDA_PICTURE_CUDA_H_;core/src/cuda/cuda_helper.cuh:__CUDA_HELPER_H__isVMAF_SRC_CUDA_HELPER_CUH_. Upstream Netflix/vmaf keeps the reserved spellings inlibvmaf/src/cuda/: a sync that touches the first or last lines of either file conflicts there; keep the fork's guard.core/src/cuda/picture_cuda.c:vmaf_cuda_picture_download_async()andvmaf_cuda_picture_upload_async()initialiseCUDA_MEMCPY2Dwith the two memory types as designators instead of{0}and two assignments. Same descriptor; keep the designators if upstream changes the neighbouring lines.core/src/feature/cuda/integer_psnr_hvs_cuda.c:init_fex_cuda()calls the newpsnr_hvs_load_module();reduce_hvs_planes()loops overpsnr_hvs_plane_count(). Fork-local file, no upstream counterpart.
filter1d.cu: the integer VIF kernels are assembled from stages (standards batch B5, 2026-10-02)¶
refactor/b5-cuda-vif-filter-kernels, ADR-1142 (HISS-04).
core/src/feature/cuda/integer_vif/filter1d.cu: the four kernel bodies (filter1d_8_vertical_kernel,filter1d_8_horizontal_kernel,filter1d_16_vertical_kernel,filter1d_16_horizontal_kernel; 133 to 225 lines each upstream) are short functions that call__forceinline__stages:vif_mirror_index(),vif_vert_load_tiles(),vif_vert8_accumulate()/vif_vert16_accumulate(),vif_vert16_round(),vif_vert_store(), and for the horizontal passvif_hori_load_tile(),vif_hori_center_tap(),vif_hori_tap_pairs(),vif_hori_border(),vif_hori_statistics(),vif_hori_flush_accums(),vif_hori_store_rd(). The two horizontal kernels are one template,vif_hori_kernel<val_per_thread, fwidth, fwidth_rd, filt_row, use_ldg>: the 8-bit kernel is row 0 with a rounding of 2^15, a shift of 16 and__ldg()loads; the 16-bit kernel is the scale's row with itsadd_shift_round_HP/shift_HPand plain loads.- The
__global__entry points, their names, their argument lists and__launch_bounds__(128, 10)are unchanged, and so areinteger_vif_cuda.cand the shared-memory layout. - An upstream change to one of the four bodies no longer applies as a hunk: port it into the stage that holds the statement (the arithmetic lines are upstream's, with
accum_x[off]spelledo.x[off]ors.x[off]), then runtest_cuda_exact_twinsand thevifgate cell on a device. Do not restore a long body:praetorctl auditrecords none for this file any more. core/src/feature/hip/integer_vif/vif_statistics.hip: one comment namesvif_mirror_index()instead of line numbers of the CUDA file.
psnr_hvs NEON masking threshold takes the scalar's double product; ADR-1469 — the SIMD butterfly is two functions (2026-10-02)¶
fix/psnr-hvs-neon-scalar-bits, T-PSNR-HVS-NEON-NOT-SCALAR-BITS-2026-10-02.
core/src/feature/arm64/psnr_hvs_neon.ccompute_masks()now reads(float)(sqrt((double)b->s_mask * b->s_gvar) / 32.0), the expression ofthird_party/xiph/psnr_hvs.c(sqrt((double)s_mask * s_gvar) / 32.f) and ofx86/psnr_hvs_avx2.c. Keep the cast in all three: without it the product is afloatproduct and the threshold is onefloatstep off on about one block in twenty.- An upstream change to
calc_psnrhvs()(Netflix or xiph) is mirrored inx86/psnr_hvs_avx2.candarm64/psnr_hvs_neon.cin the same PR;core/test/test_psnr_hvs_dispatch_invariance.c(new) fails on x86-64 or underqemu-aarch64when one of the three leaves the others. - The test includes
core/test/psnr_hvs_twin_parity.hfor its fixtures; a change toHvsFixtureorhvs_fill_pic()there reaches it. - ADR-1469:
od_bin_fdct8_simd()inx86/psnr_hvs_avx2.candarm64/psnr_hvs_neon.ccalls two stages,od_bin_fdct8_even_simd()(the first 21 statements of the scalarod_bin_fdct8()) andod_bin_fdct8_odd_simd()(the other 13), over anod_fdct8_state. The statements are unchanged. An upstream change to the scalar butterfly goes into the matching stage of both files; a conflict in the old single function is resolved by taking this side and re-applying the statement. testdata/ir-snapshots/od_bin_fdct8x8_avx2.llis regenerated (make ir-diff-update, ADR-0918):assert()line numbers and one commuted integer add. Any later edit tox86/psnr_hvs_avx2.cthat moves itsassert()lines needs the same;make ir-diffshows it.- No Netflix golden-data, public API or FFmpeg patch impact. x86-64 scores and the scalar path are unchanged; aarch64 NEON scores move by at most 5.7e-7 dB onto the scalar's.
ADR-1454 — core/src/feature/cuda/AGENTS.md is a generated index over AGENTS.d/ pages (2026-10-02)¶
docs/agents-index-feature-cuda, T-AGENTS-INDEX-MIGRATION-2026-10-02.
- No rebase impact from upstream: an upstream sync never touches an
AGENTS.mdor anAGENTS.d/. - Fork branches that append to
core/src/feature/cuda/AGENTS.mdconflict once: takecore/src/feature/cuda/AGENTS.mdfrom master, put the new rule into a new or matching page undercore/src/feature/cuda/AGENTS.d/, runmake docs-fragments-write. - Coupled edits:
scripts/ci/check-issue-reference-provenance.pyandscripts/ci/tests/test_issue_reference_provenance.pyupdate contracts forkernel-launch-params.mdandhost-preprocessing-download.md.
ADR-1454 — core/src/feature/hip/AGENTS.md is a generated index over AGENTS.d/ pages (2026-10-02)¶
docs/agents-index-feature-hip, T-AGENTS-INDEX-MIGRATION-2026-10-02.
- No rebase impact from upstream: an upstream sync never touches an
AGENTS.mdor anAGENTS.d/. - Fork branches that append to
core/src/feature/hip/AGENTS.mdconflict once: takecore/src/feature/hip/AGENTS.mdfrom master, put the new rule into a new or matching page undercore/src/feature/hip/AGENTS.d/, runmake docs-fragments-write.
ADR-1454 — core/src/feature/sycl/AGENTS.md is a generated index over AGENTS.d/ pages (2026-10-02)¶
docs/agents-index-feature-sycl, T-AGENTS-INDEX-MIGRATION-2026-10-02.
- No rebase impact from upstream: an upstream sync never touches an
AGENTS.mdor anAGENTS.d/. - Fork branches that append to
core/src/feature/sycl/AGENTS.mdconflict once: take master's side of that file, put the added text into the page whoseTouchingrow matches the files, thenmake docs-fragments-write. - No Netflix golden-data, public API or FFmpeg patch impact.
The x86 float ADM wavelet kernels are split into row helpers (ADR-1142, 2026-10-02)¶
refactor/float-adm-x86-standards. No score impact: no dispatch table calls float_adm_*_avx2 / float_adm_*_avx512 (only their headers name them), and an old-against-new comparison of all eight exported functions is bit-identical on 71,840 inputs.
core/src/feature/x86/float_adm_avx2.c and float_adm_avx512.c are fork files under the Netflix header (the float ADM SIMD port); upstream has no counterpart, so a sync does not touch them. A later rewrite or a re-wiring into adm.c's dispatch meets this layout:
| Statements of the former single function | Now in |
|---|---|
| broadcast of the eight filter taps | Dwt2TapsAvx2 / Dwt2TapsAvx512, filled once in float_adm_dwt2_avx2() / float_adm_dwt2_avx512() |
| vertical pass of a row (vector loop and scalar tail) | dwt2_vertical_row_avx2() / dwt2_vertical_row_avx512() |
| horizontal pass of a row, AVX2 (scalar) | dwt2_horizontal_row_avx2() |
horizontal pass of a row, AVX-512 (j = 0 scalar, 16-wide loop, scalar tail) | dwt2_horizontal_row_avx512() over dwt2_horizontal_scalar() and dwt2_horizontal_16_avx512() |
allocation of tmplo / tmphi, the guard-zone memset (AVX-512), the row loop | the entry point |
Kept: every multiply and add in its place and order (((c0*s0 + c1*s1) + c2*s2) + c3*s3, no FMA intrinsic), the float temporaries, the tail bounds, the aligned_malloc sizes. Changed outside the wavelet: row pointers in float_adm_csf_*, float_adm_csf_den_scale_* and float_adm_sum_cube_* are formed with (ptrdiff_t)i * stride (same address for every valid int product; clang-tidy bugprone-implicit-widening-of-multiplication-result).
The AVX-512 filter tables are file-level dwt2_filter_lo / dwt2_filter_hi (they were function-local). The invariant page is core/src/feature/x86/AGENTS.d/float-adm.md.
The NEON float ADM wavelet kernel is split into row helpers (ADR-1142, 2026-10-02)¶
refactor/float-adm-neon-standards. No score impact: every recorded adm / float_adm output is identical before and after on aarch64 (scalar and NEON dispatch, 1066 cases); the x86 build has no changed object; both golden gates 271 passed, 12 skipped.
core/src/feature/arm64/float_adm_dwt2_neon.c and float_adm_neon.c are fork files under the Netflix header; upstream has no counterpart, so a sync does not touch them. They mirror adm_dwt2_s() and the other scalar references in adm_tools.c: an upstream change to those is ported into these helpers.
Statements of the former single float_adm_dwt2_neon() | Now in |
|---|---|
the four vaddq_f32(acc, vmulq_laneq_f32(sN, f, N)) steps from +0, once for the low-pass and once for the high-pass taps | dwt2_vertical_4_neon(), called with flo, then with fhi |
| vertical pass of a row (4-wide loop, scalar tail) | dwt2_vertical_row_neon() |
horizontal pass of a row (scalar, through ind_x) | dwt2_horizontal_row_neon(); it writes a[j], v[j], h[j], d[j] where the single function wrote dst->band_X[i * dst_px_stride + j] |
allocation of tmplo / tmphi, the two vld1q_f32 of the taps, the row loop | float_adm_dwt2_neon() |
Kept: the +0 start of every sum (signed-zero parity with adm_dwt2_s()), multiply then add with no fused form, the float accum sequences of the tail and of the horizontal pass, the tail bounds. Each function carries the GCC optimize("-ffp-contract=off") attribute, as adm_dwt2_s() and its two pass helpers do; a new helper needs it too. In float_adm_neon.c only declarations and row-pointer casts changed ((ptrdiff_t)i * stride).
core/src/feature/adm_tools.c: the comment above adm_dwt2_vert_pass_s() said the wavelet stays one function; it is three functions, and the comment now says so. The NOLINTNEXTLINE(readability-function-size) on adm_dwt2_s() suppressed nothing and is gone. No code changed in that file, nor in x86/adm_avx2.c / x86/adm_avx512.c (SPDX line only).
SYCL tidy wrapper: -Wno-overriding-option (2026-10-02)¶
fix/sycl-tidy-overriding-option, T-SYCL-TIDY-OVERRIDING-OPTION-2026-10-02.
scripts/ci/clang-tidy-sycl.shpasses-extra-arg-before=-Wno-overriding-option. Keep it when the wrapper is touched: without it every translation unit whose target repeatsvmaf_strict_fp_args(151 on an icx build) is a compile failure in thesycltidy lane.- No source, score, public C API, Netflix golden-data or FFmpeg patch impact.
ADR-1454 — core/src/feature/AGENTS.md is a generated index over AGENTS.d/ pages (2026-10-02)¶
docs/agents-index-feature, T-AGENTS-INDEX-MIGRATION-2026-10-02.
- No rebase impact from upstream: an upstream sync never touches an
AGENTS.mdor anAGENTS.d/. - Fork branches that append to
core/src/feature/AGENTS.mdconflict once: takecore/src/feature/AGENTS.mdfrom master, put the new rule into a new or matching page undercore/src/feature/AGENTS.d/, runmake docs-fragments-write. - No coupled edits.
- No Netflix golden-data, public API or FFmpeg patch impact.
float_ms_ssim_sycl adds the CPU's terms in the CPU's order (ADR-1466, 2026-10-02)¶
fix/sycl-float-ms-ssim-raster-sum, the SYCL float_ms_ssim part of T-GPU-FLOAT-SSIM-FRAME-SUM-ORDER-2026-10-02. Stacked on ADR-1463.
core/src/feature/sycl/integer_ms_ssim_sycl.cpp: the vertical pass isMsSsimLcsKernel, one work-item per window, no reduction.d_partials/h_partials,store_lcs_group(), the work-group counts andLcsFixedare gone;d_terms/h_terms(two fp64 patterns per window) andd_structure/h_structure(onefloat) hold every (plane, scale) atwindow_offset.sum_scale_lcs()callsssim_lcs_sums(). Keep kernel, layout and host sums from the same side of a conflict.core/src/feature/sycl/sycl_ssim_terms.h:ssim_terms(),SsimTerms,ssim_term(),term_fixed(),FixedSumand the two fixed-point constants are removed (no user left);ssim_lcs_sums()is new. A rebase that brings a user of a removed name back converts it tossim_double_terms()and a host sum.core/test/float_ms_ssim_order_frame.h(new) is the pair the HIP lane found, shared byte for byte with the CUDA and HIP tests (sha256be2341f63ce74151...). Never edit it; a lane that lands the same file later takes either side.core/test/ssim_order_noise.h(new) is the seeded noise both SYCL SSIM tests regenerate their search pairs from.core/test/test_sycl_ms_ssim_parity.chas the two order cases;core/test/test_sycl_kernel_source_contract.pypins the new kernel, layout and sums.scripts/ci/exact_twins.d/float_ms_ssim.syclandfloat_ms_ssim_lcs.syclcite ADR-1466 and name the pairs.- No Netflix golden-data, public C API or FFmpeg patch impact. Stored
float_ms_ssim_syclscores change only where a per-scale mean lies next to afloatrounding boundary.
adm.c: the band planes are carved with a typed cursor (cppcheck, 2026-10-02)¶
fix/ci-cppcheck-exhaustive-findings.
core/src/feature/adm.c:init_dwt_band(),init_dwt_band_d()andinit_dwt_band_hvd()take and return afloat *(double *for the_dvariant) cursor and a step in samples, where upstream passes achar *cursor and a step in bytes and casts each plane.adm_alloc_bands()passesbuf_sz_one / sizeof(float). The addresses are upstream's.- An upstream change to these helpers or to the carving in
compute_adm()conflicts here. Keep the typed cursor: a(float *)cast of achar *cursor fails the requiredCppcheckcheck (invalidPointerCast), and the(float *)(void *)form fails clang-tidy (bugprone-casting-through-void).
Integer ADM weight limits follow from the contrast-masking cube (ADR-1472, 2026-10-02)¶
fix/integer-adm-aim-wrap, T-ADM-AIM-BARTEN-SCALE-TERM-WRAP-2026-10-01.
core/src/feature/adm_csf_fixed_point.h(fork file, no upstream counterpart):ADM_CSF_S123_LIMIT(2^30) is gone.adm_csf_fixed_limit(scale, band)returns the limit of one weight, andadm_csf_fixed_scale()normalises against the three limits of a scale. The constantsADM_DWT_BAND_MAX_SCALE0..3,ADM_CM_EXCESS_MAX_SQ29,ADM_CM_EXCESS_MAX_SQ30,ADM_I4_CM_EXCESS_SLACKandADM_I4_CM_WEIGHT_SHIFTdescribe upstream arithmetic ininteger_adm.c:- the wavelet taps
dwt2_db2_coeffs_lo/_hiand the DWT shifts (adm_dwt2_88 and 16;i4_dwt2_round()0 + 15, 16 + 16, 16 + 15) give the band bounds; shift_sq29 / 30 ofADM_CM_ACCUM_ROUND/I4_ADM_CM_ACCUM_ROUNDgives the excess budgets;shift_dst28 ofi4_adm_cm()gives the weight shift. An upstream change to any of these changes the bounds.core/test/test_integer_adm_cm_budget.cderives them again from the taps and fails until the constants follow; it does not read the shifts from the code, so a changed shift needsBAND_FORMATin the test and the constants updated together.- No kernel, SIMD file or device file changes. CUDA, HIP, SYCL and Metal call
adm_csf_fixed_scale()and follow. - Scores: Watson97 (default) and the two blend modes are bit-identical, so no Netflix golden value moves. Integer
admwithadm_csf_mode=1changes: by at most 7.7e-7 where it was right (Netflix pair), and from a failure or a wrong value to the right one where the square wrapped. A snapshot or test that pins a Barten-modeinteger_adm*value at more than six decimals needs regenerating; none exists on master (python/test/feature_extractor_test.pypinsadm_csf_mode=1withadm_csf_scale=0.002893, whose weights need no normalisation and are unchanged). - Upstream Netflix/vmaf has
adm_csf_modeinlibvmaf/src/feature/integer_adm.cwith no weight normalisation; its Barten mode wraps in the weight conversion itself. A port of an upstream fix there must keepadm_csf_fixed_point.has the single place that converts weights.
Integer motion SIMD kernels split into stages (ADR-1142, 2026-10-02)¶
refactor/std-motion-simd.
core/src/feature/x86/motion_avx2.candmotion_avx512.c: the two pipelines per file keep their names and signatures and are row loops over inlined stages (y_conv_row_{8,16}_*for the vertical pass of one row,x_conv_row_sad_*for the horizontal pass, with one vector block and one scalar edge helper each). An upstream change to a pipeline lands in the stage that holds the statement; the arithmetic of every statement is unchanged, including the logical_mm256_srlv_epi64of the AVX2 16-bit path.motion_avx512.c:y_convolution_8_avx512,y_convolution_16_avx512andx_convolution_16_avx512(used bycore/test/test_motion_avx512_parity.conly) sharefilter5_epu32_avx512()andfilter5_scalar(); the three passes ofx_convolution_16_avx512(left edge, interior, right edge) keep their order.core/src/feature/arm64/motion_neon.c:x_convolution_16_neonis three passes overx_conv_edge_cols_neon()/x_conv_row_interior_neon().- Each file includes its own header, so the exported functions are checked against their declarations.
- No score, public C API, Netflix golden-data or FFmpeg patch impact.
ADR-1467 — ciede.c writes its squares as products (2026-10-02)¶
fix/ciede-powf-explicit, T-CIEDE-CLANG-POWF-BUILTIN-2026-10-02.
core/src/feature/ciede.cdiffers from upstream inget_r_sub_t()(exp(-(degrees * degrees))where upstream haspowf(degrees, 2)) and inciede2000()(square(x), astatic double square(const float x), where upstream haspow(x, 2), 13 times). Keep the fork's side on an upstream sync:powf(degrees, 2)makes a GCC build and a clang build disagree again (65 of 180 measured frames, up to 2.0e-11), andtest_ciede_device_mathfails under GCC.- A new square in an upstream change to
ciede2000()is written assquare(x)ifxis afloat(the product is then exact indouble); a square of adoubleexpression is a different case and needs measuring. - Mirrors of those statements, same PR when they change:
feature/cuda/integer_ciede/ciede_device.h(ciede_r_sub_t(),ciede_sq()),feature/ciede_ff_math.h(r_sub_t(),sq()), and the pinned lines incore/test/test_sycl_ciede_exact_contract.pyandcore/test/test_cuda_ciede_exact_contract.py. ciede_device.h:ciede_r_sub_t()computes-(degrees * degrees);CIEDE_POWFis left with one use,powf(x, 7).core/test/meson.build:test_ciede_device_mathis registered on every architecture.- Netflix golden gate unchanged in outcome (271 passed, 12 skipped on x86-64 and aarch64 with GCC and clang); no golden assertion, public API or FFmpeg patch impact.
ciede2000of a GCC build moves by at most 2.0e-11.
integer_ssim.c: calc_ssim() takes its buffers from helpers (ADR-1142, 2026-10-02)¶
refactor/std-cpu-extractors.
core/src/feature/integer_ssim.c(upstream path, Xiph.Org code):calc_ssim()keeps its signature and its row loop. The twogaussian_filter_init()calls and the twomalloc()s moved intossim_work_init()(same order: vertical kernel, row pointers, row storage, horizontal kernel;-ENOMEMwith nothing left allocated) and the fourfree()s intossim_work_free(). An upstream change to the allocation part lands in those helpers; the arithmetic (ssim_accumulate_row*,ssim_reduce_row_range) is untouched. TheNOLINT(readability-function-size)that kept the function whole is gone: the HISS gate has no opt-out, and the function is 37 lines.core/test/test_feature_collector.c: the twoHAVE_CUDAtests are split into helpers (close_retry_fixture,close_retry_first_close,cuda_overwrite_and_release). Same assertions, same calls in the same order, except that the duplicate-owner test now asserts on the overwrite attempt after the four releases have run.- No score, public C API, Netflix golden-data or FFmpeg patch impact.
CLISettings is ordered by alignment (ADR-1142, 2026-10-02)¶
refactor/std-cli-tools.
core/tools/cli_parse.h(upstream path): the members ofCLISettingsare grouped as pointers, the model and feature tables, 4-byte values, flags. The set of members, their names and their types are unchanged; nothing initialises the structure by position (CLISettings c = {};invmaf.cpp, assignment by name incli_parse.cpp). An upstream commit that adds a member puts it into the group of its type, not at the place upstream has it.CLIModelConfigandCLIFeatureConfigkeep their order:cli_parse.cppuses designated initialisers on them, and C++ requires declaration order.core/tools/cli_parse.h,core/tools/vidinput.h: thetypedefs (andvideo_input_pixel_format) stay plain C inside citedNOLINTBEGIN(modernize-use-using[,performance-enum-size])blocks, because C translation units include both headers (ADR-0141, ADR-1470).- No CLI behaviour, score, public C API, Netflix golden-data or FFmpeg patch impact.
Python harness command builders are assembled from helpers (ADR-1142, 2026-10-02)¶
refactor/std-python-harness-init, T-PYTHON-CALL-VMAFEXEC-FORCE-ZERO-SECOND-MODEL-2026-10-02.
compat/python-vmaf/__init__.py(Netflixpython/vmaf/__init__.py):ExternalProgramCaller.call_vmafexec()andcall_vmafexec_multi_features()keep their signatures and call module-level helpers (_vmafexec_base_command,_vmafexec_feature_flags,_vmafexec_model_flags,_vmafexec_model_overloads,_vmafexec_run_flags,_multi_features_run_arguments,_feature_argument). An upstream sync that touches either function will conflict: put the changed flag into the helper that emits it and keep the order of the parts (python/test/python_harness_coverage_test.pypins both command texts).- Deliberate difference from upstream:
_vmafexec_model_overloads()does not overwritemotion_force_zero, so the overload reaches every model. Do not take upstream's in-loop assignment back. - No score, Netflix golden-data, public C API or FFmpeg patch impact.
x86 float ADM: wavelet and CSF kernels are exact and dispatched; the reduction kernels are gone (ADR-1473, 2026-10-02)¶
feat/float-adm-x86-simd-exact, T-FLOAT-ADM-X86-KERNELS-NOT-EXACT-NOT-DISPATCHED-2026-10-02. Maintainer decision (popup, 2026-10-02): make the kernels exact, test them and wire them.
Which files are whose:
core/src/feature/x86/float_adm_avx2.{c,h},float_adm_avx512.{c,h}: fork files under the Netflix header. Upstream Netflix/vmaf has no float ADM SIMD (libvmaf/src/feature/x86/holdsadm_avx2.c/adm_avx512.cfor the fixed-point extractor only). A sync never touches them.core/src/feature/adm_tools.c,adm_tools.h,adm.c: upstream mirrors with fork changes. This change adds to them:adm_tools.c:adm_csf_s()is now a call ofadm_csf_planes_s(..., adm_csf_plane_s). Upstream's element loop isadm_csf_plane_s(), its three statements verbatim (flt_ptr[dst_offset + j] = FLOAT_ONE_BY_30 * fabsf(dst_val);is held by the GPU twins' contract tests); the weights and the border region are inadm_csf_planes_s(). Port an upstream hunk inadm_csf_s()into the function that owns the statement.adm.c:adm_dwt2_dispatch()has x86 branches next to the NEON one;#define adm_csf adm_csf_sbecame#define adm_csf(...) adm_csf_planes_s(__VA_ARGS__, adm_csf_plane_select()). An upstream change to the twoadm_csf(...)calls incompute_adm()keeps working as long as the argument list isadm_csf_s()'s.
Rules a rebase must keep (the test core/test/test_float_adm_x86.c fails otherwise):
- every four-tap sum of the wavelet kernels starts at
+0and adds one product per step in tap order (the scalaraccum = 0; accum += c[0] * s0; ...), multiply then add, no fused form; - the CSF kernels compute
fltas a double product narrowed to float; - the wavelet kernels return
int(-ENOMEMon a failed allocation), asadm_dwt2_s()does.
Removed: float_adm_csf_den_scale_avx2() / _avx512() and float_adm_sum_cube_avx2() / _avx512() with their declarations, and the helpers only they used (hadd_pd4(), hsum_ps_to_double()). Do not bring them back from an older branch: they add the cubes in a lane tree in double, the scalar reference adds float values in column order, and compute_adm() never called a sum of cubes. ADR-0844 described their accumulation; ADR-1473 replaces it for these kernels.
No score, output, public C API, Netflix golden-data or FFmpeg patch impact. float_adm is 13 to 24 % faster with AVX2 or AVX-512.
Ten fork-authored files retagged EUPL-1.2 (ADR-1250, 2026-10-02)¶
chore/relicense-pending-mechanical, T-RELICENSE-CHECK-PENDING-2026-10-02.
- no rebase impact: the ten files exist only in the fork (tests, the
vmaf_close_retryhelper, the golden-build script); one tag line each. scripts/dev/relicense_fork_files.py --checkstill reports 31 entries; do not run--writeover the tree until the state row's decisions are made (it would put a C comment into theexact_twins.d/*.hipfragments, a header into two praetor-managed files, and rewrite a template insidescripts/sync-pelorus-interop.sh).
SPDX lines follow the notice in the file (ADR-1250, 2026-10-02)¶
fix/spdx-tags-match-notices, T-SPDX-TAG-DISAGREES-WITH-NOTICE-2026-10-02.
- Eleven upstream-path or upstream-derived files had their tag corrected to the licence of the notice they carry (
BSD-2-Clausefor the Daala, Xiph.Org, dav1d, LIME and scanf texts;AND BSD-3-Clause/AND MIT/AND BSD-2-Clausewhere Netflix's header sits next to a quoted notice). Upstream has no SPDX lines, so a sync does not touch them; if a sync replaces a header block wholesale, keep the fork's SPDX line. core/src/feature/ciede.c: the SPDX line is in the file's own header, not in the quoted MIT notice. Do not move it back.REUSE.toml:core/src/feature/third_party/xiph/**isBSD-2-Clause.scripts/ci/tests/test_spdx_tag_matches_notice.pyfails when a tag and the text next to it disagree.- No source, score, public C API, Netflix golden-data or FFmpeg patch impact.
The clang-tidy lanes are measured in the dev container (ADR-1471, 2026-10-02)¶
ci/tidy-lanes-dev-container.
- No rebase impact from upstream:
scripts/dev/tidy-lane.sh,scripts/ci/clang-tidy-hip.sh,scripts/ci/gen-gpu-compile-commands.py,build-aux/aarch64-linux-gnu-qemu-user.ini, thetidy-*targets of theMakefileandscripts/ci/tidy-baseline-*.jsonare fork-local. - After an upstream sync or any rebase that changes C, C++, CUDA, HIP or SYCL sources, the five baselines are re-measured with
make tidy-lane-write LANE=all(dev container). A conflict in a baseline JSON is never resolved by hand and never by amake tidy-ratchet-writeon the host: take either side, then re-measure. - A fork branch that tightened a baseline with a scoped write on a host (
tidy-ratchet.py --only ... --write) conflicts with the re-measured files. Take master's baselines and repeat the tightening in the container:scripts/dev/tidy-lane.sh --write --only <file> <lane>. - An upstream change that adds a dependency file to a kernel target or renames a meson custom-command rule must keep
scripts/ci/tests/test_gen_gpu_compile_commands.pygreen: the generator exits 1 when a.cu/.hipbuild statement exists that it cannot read.
Helper headers carry EUPL-1.2 AND the reproduced code's licences (ADR-1474, 2026-10-02)¶
fix/relicense-tool-clean-check, T-RELICENSE-CHECK-PENDING-2026-10-02.
- no rebase impact on upstream files: the three headers (
hip/float_ssim/ssim_decimate.h,metal/float_ms_ssim_option_semantics.h,sycl/sycl_integer_ssim_math.h) exist only in the fork; one notice block and one tag line each. scripts/dev/relicense_provenance.tomlhas five new[ports]entries (ssim_decimate.h,float_ms_ssim_option_semantics.h,sycl_integer_ssim_math.h,sycl_ssim_terms.h,sycl_ssimulacra2_math.h). A port or sync that renames one of these files must move its entry, or the family default returns and--checkasks for notices the file does not owe. The same holds for the four new[not_ports]entries (speed_cuda_params.h,float_adm_hip_math.h,ciede_hip_math.h,sycl_ciede_math.h): they hold no reference code and stayEUPL-1.2.scripts/dev/relicense_fork_files.py:EXCLUDED_PREFIXESgainedscripts/ci/exact_twins.d/,tools/figures/and.config/agent/hooks/block_evasion.py; prose grants are rewritten only when they start in the first 60 lines (header_prose_blocks()).relicense_fork_files.py --checkexits 0 on this tree;--writeis safe to run again.
Required check Licence Provenance reads the recorded upstream head (ADR-1474, 2026-10-02)¶
ci/relicense-check-required, closes T-RELICENSE-CHECK-PENDING-2026-10-02.
- An upstream port or sync moves one heading.
docs/development/known-upstream-bugs.mdhas exactly one heading## Upstream head the fork is at parity with: `<commit id>` (<date>);scripts/ci/upstream_parity_pin.pyreads it and thelicence-provenancejob of.github/workflows/lint-and-format.ymlrunsrelicense_fork_files.py --check --upstream-ref <that commit>. Update the id in the port's own pull request and keep the wording; do not add a second heading of that form (retitle the older section instead). - Before pushing a port, run the check against the new head:
python3 scripts/dev/relicense_fork_files.py --check --upstream-ref <new id>. A fork file whose path or name now exists upstream changes verdict (upstream-path,upstream-name) and keeps its terms from then on. - The job fetches
https://github.com/Netflix/vmaf.git masterand needsfetch-depth: 0;relicense_fork_files.pyrefuses a shallow checkout (require_full_history()). Licence Provenanceis in the aggregator'srequiredandstrictMustReportarrays and inADR_1474_STRICT_CONTEXTSofscripts/ci/tests/test_hiss_replay_contract.py; rename all of them together.
iqa_ssim() counts windows for the SIMD kernels (2026-10-02)¶
fix/float-ssim-8x8-avx-garbage, T-FLOAT-SSIM-SUB-WINDOW-SIMD-COUNT-2026-10-02.
core/src/feature/iqa/ssim_tools.c: the two calls into the SIMD dispatch (g_ssim_variance,g_ssim_accumulate) passssim_window_count(w, h), which is 0 unless both extents are positive, andiqa_convolve_dispatch()uses the SIMD convolve only forw >= k->w && h >= k->h. Both are fork-local: upstream has no SIMD dispatch in this file and its scalar loops need neither guard.- On an upstream sync of
ssim_tools.ckeep the guards and keep the divisor of the four means as upstream writes it (w * h): a frame smaller than the window scores0 / (w * h)in both trees, andcore/test/test_iqa_ssim_sub_window.cpins that value and the equality of every dispatch with the scalar path. - No score of a frame of 11x11 or larger moves; no snapshot or golden value is involved.
Eight deliberate deviations from Netflix's source have their ADR (ADR-1479 to ADR-1486, 2026-10-02)¶
docs/adr-deliberate-upstream-deviations. Documentation only; no code moves.
Each of these fork lines differs from Netflix 9e48141b on purpose. On an upstream sync keep the fork's side until the named upstream pull request is merged, then take upstream's lines (the ADR says where the two forms differ without differing in value) and remove the deviation's entry from the upstream parity guard's allowlist.
| ADR | Fork lines to keep | Upstream form | Ends with |
|---|---|---|---|
| ADR-1479 | core/src/feature/ciede.c: ss_hor for the chroma column index, ss_ver for the row advance | flags swapped (ciede.c:71, :73, :89, :91) | Netflix/vmaf#1611 |
| ADR-1480 | core/src/feature/speed.c: speed_temporal buffers of float_stride * alloc_height | float_stride * h (speed.c:1578) | Netflix/vmaf#1627 |
| ADR-1481 | core/src/thread_pool.c (last_error), core/src/libvmaf.c (threaded_extract_batch_func() returns f->err) | void job function, vmaf_thread_pool_wait() returns 0 | no upstream pull request |
| ADR-1482 | core/src/feature/integer_adm.c::dwt2_src_indices_1d(), adm_half_shift() in adm_csf_fixed_point.h and its callers in x86/adm_avx2.c, x86/adm_avx512.c | dwt2_src_indices_filt() (integer_adm.c:708), pow(2, shift - 1) | Netflix/vmaf#1599, #1600 |
| ADR-1483 | vmaf_chroma_extent() (core/src/picture_geometry.h) and every caller | w >> ss_hor, h >> ss_ver (picture.c:74, :76) | no upstream pull request |
| ADR-1484 | core/src/feature/ms_ssim.c: fabs() on l, c, s before pow() | no fabs() (ms_ssim.c:294) | Netflix/vmaf#1665 (fabs() on s only; equal in value) |
| ADR-1485 | core/src/feature/integer_psnr.c::flush(), vmaf_psnr_aggregate() in psnr_score.h | three planes, ceiling with the factor 2 (integer_psnr.c:226) | Netflix/vmaf#1666 (keeps the factor 2; equal in value) |
| ADR-1486 | core/src/feature/motion.c: img1_stride, img2_stride to the scale-1 scaler | stride recomputed from the width (motion.c:70) | Netflix/vmaf#1667 |
Upstream parity guard and its allowlist (ADR-1487, 2026-10-02)¶
feat/upstream-parity-guard.
- No rebase impact on upstream files:
scripts/dev/upstream_parity*.py,scripts/dev/upstream_parity_harness.c,scripts/ci/upstream_parity.d/,scripts/ci/upstream_parity_allowlist.pyand the generateddocs/development/upstream-parity-allowlist.mdexist only in the fork. - An upstream port or sync moves the recorded head in
docs/development/known-upstream-bugs.md. In the same pull request runmake upstream-parity-full: a difference that appears is the port's to fix or a change upstream made that the port left out; a fragment the port makes stale (upstream took the fork's fix) is removed there. - A sync that takes upstream's side of a line an ADR-recorded deviation covers turns that deviation's fragment stale: either the line is restored (the deviation stands) or the fragment and the deviation go together. The table under "Eight deliberate deviations" above names the lines.
- The harness asks this tree for its feature collector through
vmaf_feature_collector_get()(core/src/libvmaf_priv.h); keep that accessor and the collector'sfeature_vectorandaggregate_vectorfields. scripts/ci/cross_backend_calibration.py: the fragment line parser is nowparse_fragment_fields()andcheck_fragment_adrs(), shared byexact_twins.dandupstream_parity.d. A conflict in the generated allowlist table is resolved like the exact-twin table: master's side, thenmake docs-fragments-write.scripts/ci/setup-golden-build.shhonoursGOLDEN_NINJA_JOBS.testdata/bench_upstream_ab.pyno longer clones upstream itself and has no--max-score-delta; a branch that still passes the option fails at argument parsing. Its--fork-builddefault moved fromcore/build-goldeninto the guard's work directory.- The guard measures in the dev container image only (
--container, whatmake upstream-paritypasses); outside it, it exits 2 unless--unpinnedmarks the verdict advisory. A bound measured on a host is not evidence: a branch that re-sizes a fragment does it from an--containerrun, and says so in the fragment'sevidence. Result documents are schema 2 (they carry the environment); a schema-1 document is refused. - A bound over a value the heap check (
--heap-check) finds undefined upstream must beinf. A dev image rebuilt with another compiler or C library is a new environment: runmake upstream-parity-fullin it before trusting a bound.
SpEED's three fp64 expressions are upstream's again; GPU twins score on the host (ADR-1477, 2026-10-02)¶
fix/speed-upstream-double-math, T-SPEED-UPSTREAM-DOUBLE-MATH-2026-10-02.
core/src/feature/speed.c:create_givens()(1.0 / sqrt(1 + t * t)),update_entropy()(log2(...) + log2(2 * M_PI * M_E)) andget_speed_score()(log2(1 + ...),/ 2.0,0.75 * ...) are upstream's lines again (Netflix9e48141b,libvmaf/src/feature/speed.c418, 423, 802, 897 to 928). A sync takes upstream's side there. The fork adds onlyNOLINT(performance-type-promotion-in-math-fn)comments citing ADR-1477; never resolve that lint by writingsqrtf/log2f.si_create_givens()inspeed_internal.cmirrors the first. Afterget_speed_score()the fork adds three test entries,speed_internal_cpu_create_givens(),speed_internal_cpu_update_entropy()andspeed_internal_cpu_speed_score()(declared inspeed_internal.h), which call the three functions unchanged forcore/test/test_speed_upstream_form.c; keep them when taking upstream's side.speed_internal.cgainedspeed_internal_gpu_tail_scores(): upstream'supdate_entropy(),est_params()steps 8 and 9,get_speed_score()and the one-side-singular rule ofspeed_extract_score(), for the GPU twins. An upstream change to one of those functions is ported into the tail in the same PR.speed_internal_entropy_constant()/speed_internal_base_entropy()andspeed_constants.hare gone.speed_gpu_common.h:SpeedGpuScoringis{sigma_nn, nn_floor, weight_mode}(host only);SpeedGpuTailLayoutdescribes the block a twin reads back.SpeedCudaFrameArgsandSpeedHipParamslostent,contrib,resultandscoring.- New
core/src/feature/speed_givens.h(speed_givens_unit()), included bycuda/speed/speed_score.cu,hip/speed/speed_hip_device.handsycl/speed_sycl_pipeline.cpp; listed incuda_kernel_shared_headersand the HIP kernel header list ofcore/src/meson.build. - Removed:
speed_log2_hard_cases.h, each twin'sspeed_log2()/speed_hd_log2_rn(),speed_score_kernel,speed_hip_score,launch_score(). - Gate:
scripts/ci/exact_twins.d/speed_{chroma,temporal}.{cuda,hip,sycl}added,LIBM_TWINSlost both features. On a conflict indocs/development/cross-backend-exact-twins.mdtake master's side and runmake docs-fragments-write.
float_ms_ssim_cuda and integer_ms_ssim_hip score the chroma planes (2026-10-03)¶
fix/ms-ssim-chroma-cuda-hip, T-MS-SSIM-GPU-CHROMA-OPTION-DRIFT-2026-09-06.
core/src/feature/cuda/integer_ms_ssim_cuda.candcore/src/feature/hip/integer_ms_ssim_hip.ckeep their geometry, pyramid and term buffers per plane (MsSsimPlaneCuda,MsSsimPlaneHip) and run the luma pipeline once per scored plane, asfloat_ms_ssim.cdoes. Both declare the CPU's four options (enable_chromais new on CUDA; on HIP it was accepted and ignored) and providefloat_ms_ssim_cb/float_ms_ssim_cr. Both are fork files with no upstream counterpart; an upstream change tofloat_ms_ssim.c's plane loop or its chroma minimum changes both twins in the same PR.- Both include
core/src/feature/metal/float_ms_ssim_option_semantics.hfor the active plane count and the ceil-subsampled plane size; keep its helper names, which the Metal twin and its device-free test also call. - The HIP twin no longer stages level 0 in
d_ref0/d_cmp0: each plane uploads straight into pyramid level 0. A rebase onto an older layout keeps the direct upload. - Gate: new cell
float_ms_ssim_chroma(float_ms_ssimwithenable_chroma=true) in both gate scripts,exact_twins.dfragments for CUDA, HIP and SYCL, andFEATURE_MIN_CHROMA_DIM, which reports the cell SKIP on a fixture whose chroma is below 176 pixels (the 576x324 4:2:0 pair). On a conflict indocs/development/cross-backend-exact-twins.mdtake master's side and runmake docs-fragments-write. - No CPU extractor, snapshot or golden value moves.
float_adm_sycl's term kernel takes the large register file; its probe's queue is in order (ADR-1501, 2026-10-03)¶
fix/sycl-xe2-float-adm-terms, T-SYCL-FLOAT-ADM-TERMS-XE2-SPILL-2026-10-03, T-SYCL-FLOAT-ADM-PROBE-OUT-OF-ORDER-QUEUE-2026-10-03.
core/src/feature/sycl/sycl_compat.h:VmafSyclKernelShape<0, 256>means no required sub-group size with the large register file (VmafSyclShapeSubGroup<0>; astatic_assertrefuses size 0 with GRF 0). A rebase that changes the shape template keeps both.core/src/feature/sycl/float_adm_sycl.cpp(FadmTermsKernel) andcore/test/test_sycl_float_adm_math_probe.cpp(TermsKernel): the term kernel is a functor in the shapekTermsSubGroup/kTermsGrfofsycl_float_adm_math.h(0 / 256). Do not turn it back into a lambda or give it a fixed sub-group size: either spills on some default AOT target. The probe's queue is created in order.core/test/test_sycl_sub_group_size_contract.pyandcore/test/test_sycl_float_adm_exact_contract.pyguard it.- No score, public C API, Netflix golden-data or FFmpeg patch impact.
Declared ruleset matches the live one¶
.github/rulesets/main.jsonis removed and.standards.yamldeclinesbranch-ruleset; a praetor pin move orsyncmust not bring the template back (test_repository_security.pyfails; ADR-1504). No rebase impact.
Praetor pin 0af07a733e65 (ADR-1506)¶
PRAETOR_REFis0af07a733e6534269b435cea185da4d1df7aba0c; the vendoredtools/markdownlint/lock has nobraces. A rebase keeps the engine's files (documentation gate, DevContainer bundle, compiled context,.standards.lock) and the newtools/apicompat/and.github/workflows/praetor-api.yml; on conflict take master's side and regenerate with the engine, never by hand..standards-baseline.jsonis re-recorded once at the tip with the pinned engine (503), and the README figure follows it.
Tester workflows: version string and release identity¶
.github/workflows/macos-tester-bundle.ymland.github/workflows/docker-publish-tester.ymlusegit describe --tags --match 'v*.*.*'(no--always, no--long), andscripts/ci/check-vcs-version-not-bare-sha.shholds both to it. The macOS workflow'sCreate the tester prereleasestep alone speaks as the release-bot identity (App, elseRELEASE_BOT_TOKEN, else fails), after theChoose the release-bot identityandMint the release-bot installation tokensteps. A rebase keeps all three and the job token on every other step. No score, public API or FFmpeg patch impact.
Tester manifests name tests relative to the package root¶
tools/rc1-tester/image/prepare_build.py stagewrites each test'scmdrelative to the image root (tests/<test>), andvmaf_rc1_tester.hw_suites.command_path()resolves a relativecmdagainst the root the manifest is read from. A rebase keeps both halves: an absolute path of the build machine names nothing on the machine a bundle is unpacked on (T-TESTER-BUNDLE-UNIT-PATHS-ABSOLUTE-2026-10-04). No score, public API or FFmpeg patch impact.
Hardware we need page and its generator¶
docs/usage/hardware-we-need.mdholds a table between thehardware-needs:beginandhardware-needs:endmarkers thatscripts/docs/generate-hardware-reports.pyrewrites fromscripts/docs/hardware-needs.jsonand the reports underdocs/hardware-reports/; never edit the table by hand. On a conflict inside the markers take master's side and runmake docs-fragments-write. A new GPU family in a row map oftools/rc1-tester/image/needs a row inhardware-needs.jsonin the same PR (the generator refuses otherwise)..github/ISSUE_TEMPLATE/hardware_report.ymllists all five tester packages; a new package adds an entry there. No score, public API or FFmpeg patch
Windows tester zip (ADR-1515)¶
.github/workflows/windows-tester-bundle.ymlbuilds the zips onwindows-2025andwindows-11-vs2026-armwithscripts/ci/build-windows-tester-bundle.py;-Db_vscrt=mtis load-bearing (ADR-1503 rule 7: no runtime DLL for VMAFx programs), andscripts/ci/check-windows-bundle-imports.pymust stay between the notices and the pack, as must thewindows-ziplicence check. The interpreter'svcruntime140*.dllcome from the runner'sVCToolsRedistDir, never from python-build-standalone ordebug_nonredist.tools/rc1-tester/src/vmaf_rc1_tester/hw_winfacts.pygives Windows hosts the Linux machine names (x86_64,aarch64);hw_facts.machine_name()andhw_facts.vmaf_binary()are the one place the report andprepare_build.pylearn the architecture and thevmafprogram's name. The report schema keepsschema_version3 and gains the enum valueswindowsandwindows-zip.scripts/ci/check-vcs-version-not-bare-sha.shholds the new workflow'sgit describeto--match 'v*.*.*'. No score, public API or FFmpeg patch impact.
Windows CUDA tester zip (ADR-1516)¶
windows-tester-bundle.ymlgains the matrix legx64-cuda(VMAFX_GPU=cuda): the toolkit fromscripts/ci/install-cuda-toolkit.ps1, nv-codec-headers atNV_CODEC_HEADERS_COMMITread fromdocker/Dockerfile.tester(one pin for both CUDA kits), artifactwindows-cuda-zip. The CUDA EULA check (CUDA_EULA_MARKERSofscripts/ci/build-windows-tester-bundle.py) moves with the Linux image's check on a CUDA bump.hw_cuda.pyreaches the GPU on Windows throughnvcuda.dllin System32 (pathwindows);hw_cudaprobe.DRIVERnames that DLL there.prepare_build.pyleaves a shell test out of a Windows build.generate-hardware-reports.pylets a GPU row name a platform; such a row covers no family in the coverage check. No score, public API or FFmpeg patch impact.
Agent-imported pages regrouped¶
docs/development/rebase-sensitive-invariants.mdis grouped under H2 sections by area (documentation, build and CI, upstream, backends and the gate, floating-point policy, then one section per feature family). A sync that adds an invariant puts it in its area's section and keeps the contents list at the top in step;core/test/test_agent_pages_contract.pyfails when an entry stops at "See", a link or path stops resolving, or retired status text returns. On a conflict in the page, resolve per hunk inside the section of the entry, never take a whole side. No score, public API or FFmpeg patch impact.
Windows tester zip: unimported runtime DLLs and per-leg verify¶
scripts/ci/build-windows-tester-bundle.pyremoves an interpretervcruntime140*.dllthat no program of the interpreter imports before it replaces the others from the runner's redistributable folder (drop_unimported_runtime(), reading imports with the parser ofscripts/ci/check-windows-bundle-imports.py), and prints the output of the unit tests the report counts as failed. The verify job ofwindows-tester-bundle.ymlruns per leg whenvalidatepassed; the publish job still needs every leg. No score, public API or FFmpeg patch impact.
Stale-text fixes and the vmafx-mcp alias (ADR-1521)¶
core/test/test_stale_text_contract.pypins help strings, header comments, option descriptions, CI comments, the fuzz README, the perf-gate page and the Metal gate row to the code. An upstream sync that rewritescli_parse.cpp's usage text,picture.h'svmaf_picture_alloccomment ordispatch_strategy.cppmust keep what those tests assert (HIP / Metal help says opt-in; the picture header says 64 samples and zero-filled; theVMAF_SYCL_NO_GRAPHwarning namesVMAF_SYCL_DISPATCH=<feature>:direct).mcp-server/vmaf-mcp/pyproject.tomlkeepsvmafx-mcppointing atdeprecated_vmafx_mcp_aliasuntil the Python package is removed (ADR-1229); never point it back atmain. No score, public API or FFmpeg patch impact.
MSVC float ADM wavelet: the first +0 + of the vector sums¶
core/src/feature/x86/float_adm_avx2.candfloat_adm_avx512.cstart the four-tap vector sum withdwt2_plus_zero_avx2()/dwt2_plus_zero_avx512()(a compare with_CMP_NEQ_UQand a mask), not with_mm*_add_ps(_mm*_setzero_ps(), p): MSVC 19.51 removes that intrinsic addition under/fp:preciseand the kernel then returns -0 whereadm_dwt2_s()returns +0. A sync or cleanup must not restore the addition; GCC and Clang builds cannot show the defect, the MSVC lane can..github/workflows/build.ymlrunstest_float_adm_x86inWindows MSVC+CUDA (full)for that reason; keep it in the list. No score, public API or FFmpeg patch impact.
Windows zips: the Visual Studio 2026 licence terms¶
tools/rc1-tester/image/licensing.jsonpins the Visual Studio 2026 licence document infetched_textswith"extract": "docx-text";licensing.py(docx_text(),fetch_text()) writes its paragraphs totexts/visual-studio-2026-license-terms.txt. Both Microsoft components ofwindows-zipcarry that text, andwindows-cuda-ziptakes them with{"from": "windows-zip", ...}, so they have one definition. A sync that touches the record keeps the text on both components; a Visual Studio major version change records its new document first. No score, public API or FFmpeg patch impact.
TransNet V2 upstream pin (fix/transnet-exporter-pin)¶
UPSTREAM_COMMITinai/scripts/export_transnet_v2.py,upstream_commitandlicense_urlinmodel/tiny/transnet_v2.json,license_urlinmodel/tiny/registry.jsonand the model page carrya0942ca347ee00aa455631147641954278b1d1a5, the commit that added the weights upstream (its LFS object ids equal the two pinned hashes).77498b8enever existed upstream; a sync must not restore it.ai/tests/test_transnet_pin_consistency.pyguards the four places. No score, public API or FFmpeg patch impact.
Registry schema test dependency (fix/registry-schema-test-deps)¶
jsonschema==4.26.0is a line ofpython/requirements-test.in; regeneratepython/requirements-test-lock.txtwithmake python-locks-writeafter a sync that touches the file, and keepmodel_registry_schema_test.pyimporting it at module level (noimportorskip). No score, public API or FFmpeg patch impact.
PTQ stub scripts removed (fix/quantize-stubs)¶
ai/scripts/gen_calibration.pyandai/scripts/quantize_int8.pyare removed, with the.standards-baseline.jsonrow of the first;vmaf-train quantize-int8is the entry point. A sync must not restore them. No score, public API or FFmpeg patch impact.
Controller client credentials shared by node and operator (ADR-1569)¶
pkg/controllerclientowns the TLS and bearer-token code of every controller client:cmd/vmafx-node/controller_auth.goand the node'stransportCredentialsare gone,controllerConfig.Credsholds acontrollerclient.Credentials, and the operator'sVmafxJobReconciler.ControllerCredentialsfeedsConnFactory.Dial. Both binaries appendcontrollerclient.CompoundKeysto their config options. A sync that touches either binary's dial must keep going through the package; no score, public C API or FFmpeg patch impact.
Sidecar key opset (fix/sidecar-opset-key)¶
core/src/dnn/model_loader.creadsopset(notonnx_opset), asregistry.schema.jsonandai/scripts/validate_model_registry.pydo;vmaf_train.registry.ModelMetadatahas the fieldopset. A sync or a new sidecar writer must not bringonnx_opsetback.ai/tests/test_sidecar_opset_key.pyandtest_model_loaderguard it. No score or FFmpeg patch impact;vmaf_model_meta.opsetis now filled for every sidecar that has the key.
eBPF program licence string (ADR-1559)¶
cmd/vmafx-node/bpf/rclone_bypass.bpf.cdeclares"GPL"in itsSEC("license")section; its SPDX line staysEUPL-1.2. Never restore"Dual BSD/GPL"(a grant the project never made) or put"EUPL-1.2"there (the kernel refuses the GPL-only helpers the program calls). Regenerate the object withgo generate ./cmd/vmafx-node/bpf/after any change to the C file;TestEmbeddedObjectLicencechecks the string. No score, public API or FFmpeg patch impact.
Orphan documentation pages audited¶
- The pages that were outside
mkdocs.ymlare in the navigation (Development, Server, Architecture, Choose a backend) or under Records.docs/api/vulkan-image-import.mdanddocs/superpowers/plans/2026-09-20-pelorus-interop-sync.mdare deleted; archived changelog text that links the first is left as written. On a conflict in a corrected page, keep the verified statement (each names the workflow, script or file it was checked against). No score, public API or FFmpeg patch impact.
GPU device code compressed (build/compress-everything, ADR-1590)¶
core/src/meson.builddefinescuda_compress_args,hip_compress_argsandsycl_compress_argsbetweenBEGIN/END VMAF {CUDA,HIP,SYCL} device code compression policymarkers, gated by the new optioncompress_device_codeincore/meson_options.txt. The nvcc fatbin command, every[hipcc_exe, '--genco']command (the two HIP test probes incore/test/meson.buildtoo), the SYCL AOT compile line (toolchain and per-TU skip path),sycl_link_args(oneif not sycl_msvc_device_linkappend after the existing assignment) and the MSVCsycl_device_link_argstake the list. The literal--offload-compressleftsycl_icpx_aot_base_argsand the MSVC device link list. An upstream sync that rewrites the CUDA gencode or HIP target blocks keeps the lists on the commands; a new device compile site adds its backend's list. The build-time checks (*_device_compression_check,core/src/check_device_compression.py) fail a build that stores raw device code.scripts/ci/gen-sycl-compile-commands.pystrips--offload-compression-level=. Guard:core/test/test_device_code_compression.py. No score, public API or FFmpeg patch impact; builds with the clang CUDA driver or AdaptiveCpp need-Dcompress_device_code=false.
Job cancel reaches the node (ADR-1567)¶
cmd/vmafx-controller/proto/controller.protoaddsrunning_job_ids(4) toHeartbeatRequestandcancel_job_ids(2) toHeartbeatResponse; the bindings undergen/go/controller/are regenerated withcmd/vmafx-controller/proto/generate.shand keep their two// SAFETY:comments (protoc does not emit them).Heartbeatanswers withqueue.CancelledAmong; the node'srunJobuses a per-jobcontext.WithCancelCause. Keep the tenant in the SQLWHEREand the 64-entry bound. No score, public C API or FFmpeg patch impact.
Windows SYCL tester zip (ADR-1566)¶
scripts/ci/build-windows-tester-bundle.pybuilds the SYCL zip withicx-cland-Db_vscrt=md(SYCL_OPTIONS,KITS) and stages everything a program loads beside it (PROGRAM_DIRS): the Visual C++ runtime closure (copy_program_runtime()), Intel's runtime fromtools/rc1-tester/image/sycl-runtime-windows.jsonpruned to what is imported or loaded by name (LOADED_AT_RUN_TIME), and the Level Zero loader.scripts/ci/check-windows-bundle-imports.py --runtime mdholds the layout; the CPU and CUDA zips keep--runtime mt.tools/rc1-tester/image/prepare_build.pyreadscredist_dir(Windows entries are<installdir>/bin/<name>) and a component'sdests;hw_sycl.pyon Windows opens the zip's own loader throughVMAFX_ZE_LOADER(hw_l0probe.py).core/test/test_sycl_kernel_scratch.creadsVMAF_SYCL_SCRATCH_RATCHET_FILEbefore its compiled path; an upstream or fork change to the test keeps that override, which the zip's manifest sets. No score, public API or FFmpeg patch impact.
Scoring roots per tenant (ADR-1577)¶
pkg/scoringscopedecides which inputs a tenant may score; the controller calls it inScore,POST /v1/score(scoring_scope.go,http_server.go::decodeScoreRequestsplit out ofhandleScore) andSubmitJob, and setsJob.scoring_roots(proto field 10) only inPullWorkanswers; the node'sscopedSourcesresolves a job's inputs beforepkg/storageprepares them.TenantSpec.Scoring, the CRD'sspec.scoring.rootsand the chart'sauth.scoringRoots/auth.tenants[].scoringcarry the configuration. Deny by default: a sync must not make an empty root list admit inputs, nor drop the node-side check. No score, public C API or FFmpeg patch impact.
vmaf-tune predictor trainer uses the shared ONNX exporter¶
tools/vmaf-tune/src/vmaftune/predictor_train.py::_export_onnx()delegates toai/src/vmaf_train/models/exports.py::export_to_onnx(); it must not regain atorch.onnx.export(..., dynamo=False)call._ensure_ai_src_importable()is the single place that addsai/srctosys.path. No score, public API or FFmpeg patch impact. Guard:tools/vmaf-tune/tests/test_predictor_train.py.
pkg/codecadapter: one definition per codec¶
softwareAdapters(),acceleratedAdapters()andadditionalAdapters()inpkg/codecadapter/codecadapter.gocall the named constructors (libx265Adapter()and the rest); a sync or a Python-parity update edits the constructor and never adds anAdapterliteral to a list. Guard:pkg/codecadapter/one_definition_test.go. No score, public API or FFmpeg patch impact.
ai/scripts stubs removed, two placeholder generators implemented¶
- Eight
ai/scripts/*.pystubs are deleted (eval_loso_fr_regressor_v2,external_benchmark_pvmaf,fetch_lsvq,gen_ssimulacra2_eotf_lut,hdrsdr_vqa_to_corpus_jsonl,my_corpus_to_corpus_jsonl,train_fr_regressor_v4,train_video_saliency_student); the realscripts/gen_ssimulacra2_eotf_lut.pyis untouched. A sync must not restore them.gen_dists_sq_placeholder_onnx.pyandgen_mobilesal_placeholder_onnx.pyare real and guarded byai/tests/test_no_stub_scripts.py. No score, public API or FFmpeg patch impact.
Go tests resolve the vmaf CLI through internal/vmaftest¶
- Go tests that run the
vmafCLI callvmaftest.Binary(t)(internal/vmaftest/vmaftest.go):VMAF_BIN, elsecore/build-cpu/tools/vmaf, else a failure. Do not add aPATHlookup, a/usr/local/bin/vmafcandidate or a private resolver to a test;libvmaf.FindBinary()(production code, keeps the installed-container candidate) is not a test resolver. No score, public API or FFmpeg patch impact.
Controller node role (ADR-1563)¶
cmd/vmafx-controller/grpc_roles.golistsauth.RoleNodealone for the four node-API methods;auth.IsKnownRoleanddevClaimsinclude it, andtenantRolesrefuses it asdefaultRole. TheVmafxTenantCRD (deploy/helm/vmafx/crds/vmafx.dev_vmafxtenants.yaml) offers it inallowedRolesonly. A sync that touches the role table must not putvmafx:adminback on the node API;TestGRPCRolesEnforcedPerRPCholds the table independently of the code. No score, public C API or FFmpeg patch impact.
Controller workload in the Helm chart (ADR-1589)¶
deploy/helm/vmafx/templates/controller.yamlis the controller workload;vmafx.controllerAuthEnv(_helpers.tpl) is the only place the auth settings are rendered, andtemplates/deployment.yaml(server) carries none.auth-validate.yamlrequirescontroller.enabledwithauth.enabledand the other way round, and refuses avmafx-controllerimage.repository. The tenant-reader Role and the API-server NetworkPolicy selectcomponent: controller.docker/Dockerfile.controllermirrorsDockerfile.go-serverstage for stage; a change to one build recipe changes the other, and licence recordproduction-controller-imagereuses the go-server components throughrewrite.cmd/vmafx-controllerhas--version(pkg/version).e2e-k8s.ymldownloads carry--max-timeand its diagnostics usediag(). No score, public C API or FFmpeg patch impact.
Controller service account (ADR-1592)¶
templates/controller.yamlcreates and usesvmafx.controllerServiceAccountName(<serviceAccount>-controller);controller-tenant-rbac.yamlbinds only it;operator-rbac.yamlhas novmafxtenantsrule. A sync of the operator RBAC must not bring the tenant rule back (ADR-1058's rule is replaced). No score, public C API or FFmpeg patch impact.
Python Package Tests (vmaf-tune) runs the two real-x265 tests (2026-10-04)¶
.github/workflows/tests-and-quality-gates.yml, job vmaf-tune-tests: installs the distribution ffmpeg (libx265 included), sets VMAF_TUNE_INTEGRATION=1 for the suite step, and fails the job on a skip whose reason is ffmpeg not on PATH or libx265 unavailable. A rebase keeps the install step, the variable and the widened skip pattern together. The research digest docs/research/1178-dev-container-image-publish.md no longer calls the dev image published "for transparency" (ADR-1564: the package stays private). No score, public API or FFmpeg patch impact.
Affected-suite runner (run_affected_suites.py)¶
.github/test-suites.jsoncarries the local-run fieldssource_paths,install,pytestandfail_on_skip(or anot_localreason) per suite;scripts/ci/suite_registry.pyparses and checks them, andscripts/ci/run_affected_suites.pyreads them. A sync that adds a suite adds the fields in the same change; a lock rename changes the suite'sinstall. The CI jobs do not readinstallyet. No score, public C API or FFmpeg patch impact.
Release workflows skip tester tags¶
docker-publish-production.yml,docker-publish-operator-node.ymlandsupply-chain.ymlcarry the job-level guardgithub.event_name != 'release' || startsWith(github.event.release.tag_name, 'v')onvalidate-releaseand on anyif: always()summary job. A sync or rebase keeps it, and a newon: releaseworkflow adds it (scripts/ci/tests/test_release_workflows_version_tag_guard.pyfails otherwise). No score, public API or FFmpeg patch impact.
gen-node-bpf prefers the versioned clang; oneAPI installer removal retries¶
scripts/dev/gen-node-bpf.shfind_tool()tries<tool>-<major of the pin>before<tool>; an upstream sync or rebase keeps that order (scripts/dev/tests/test_gen_node_bpf.py::test_versioned_pinned_clang_wins_over_a_newer_default_clang). TheInstall Intel oneAPIstep of.github/workflows/windows-tester-bundle.ymlkeeps its exit-code checks and the boundedRemove-Itemretry (scripts/ci/tests/test_windows_tester_oneapi_install.py). No score, public API or FFmpeg patch impact.
Python tests resolve the vmaf CLI through scripts/lib/vmaftest.py¶
- Python tests that run the
vmafCLI callscripts.lib.vmaftest.find()(VMAF_BIN,VMAF_BIN_FOR_TESTS, thenbuild/,core/build/,core/build-cpu/; a set variable that names no executable raises) and skip or fail withvmaftest.MISSING_MESSAGEwhen it returnsNone. The vmaf-tune suite goes throughtools/vmaf-tune/tests/_vmaf_cli.py(vmaf_under_test(),fork_vmaf_under_test()), and themcpsuite'stests/conftest.pypoints the server'sVMAF_BINat the resolved binary before every test. Do not add aPATHlookup, a/usr/local/bin/vmafcandidate or a private resolver to a test;scripts/ci/tests/test_tests_use_vmaf_under_test.pyfails on either outside itsALLOWEDdata-only files. The shipped tools' runtime discovery (mcp-server/.../server.py::_vmaf_binary,tools/rc1-tester/.../probe.py) andscripts/ci/run_affected_suites.pyare not test resolvers. No score, public API or FFmpeg patch impact.
cpu tidy lane reaches zero findings¶
core/src/mcp/3rdparty/cJSON/cJSON.h(vendored):cJSON_SetNumberValue,cJSON_SetBoolValueandcJSON_ArrayForEachparenthesise their macro arguments; an upstream cJSON sync keeps that form.core/src/dict_internal.hisnumeric()copies itsstring_viewinto astd::stringbeforestrtof(). Shared C headers keep theirNOLINTBEGIN/ENDblocks citing ADR-1138 (a lint cleanup must not turn theirtypedefintousingor give an enum a C++-only base). No score, public API or FFmpeg patch impact.
ADR status sweep 2026-10-05¶
- no rebase impact: docs only. 78 ADR status lines and their index rows changed (
docs/adr/*.md,docs/adr/_index_fragments/*.md, regenerateddocs/adr/README.md); a sync that conflicts on one takes master's status line and keeps the dated### Status update 2026-10-05note at the end of the file. The new gatescripts/ci/check-adr-status-drift.py(exceptions inscripts/ci/adr-status-exceptions.json, expiring 2027-01-05) has its own test (scripts/ci/tests/test_check_adr_status_drift.py).
Rust CI lints the whole workspace¶
.github/workflows/rust-ci.ymlrunscargo fmt --all --checkandcargo clippy --workspace --all-targets -- -D warnings. A sync or rebase keeps both on--workspace/--all: a-p <crate>form would leave the other members unlinted again. No score, public API or FFmpeg patch impact.
compat/python-vmaf helpers keep their iterative form (HISS-01 / HISS-07)¶
compat/python-vmaf/tools/misc.py(_to_ordered_dict,_load_module_from_path,_write_overridden_copy,import_python_file),core/result_store.py(_to_python_nativeswith its work stack) andtools/decorator.py(persist_to_fileraisesPersistCacheErrorinstead of callingsys.exit(1)) differ from Netflix's text. An upstream sync that touches them keeps the fork's side of each hunk;compat/python-vmaf/tests/test_decorator_extended.pyandcompat/python-vmaf/tests/test_result_store.pycover the new shapes. No score, public API or FFmpeg patch impact.- HISS native batch 4 (
refactor/hiss-zero-native-vendor):PELORUS_VENDOR_SHAmoves to aVMAFx/peloruscommit that carries the HISS splits ofinterop.candqp_report_csv.cand Pelorus's UTF-8 CSV path opening (ADR-0149 upstream). The mirror stays verbatim (ADR-1113): an upstream sync that conflicts incore/src/interop/pelorus_*.c,core/include/libvmaf/pelorus/*.horcore/test/test_pelorus_interop.ctakes master's side and re-runsscripts/sync-pelorus-interop.sh --update; never merge a hunk by hand. No score, public API or FFmpeg patch impact.
Composite actions are linted by a script of their own¶
scripts/ci/check_composite_actions.pykeeps the shellcheck ignore list of actionlint v1.7.12'srule_shellcheck.go; when the actionlint pin in.pre-commit-config.yamlmoves, compare the list. A new composite action under.github/actions/is picked up without a config edit. No score, public API or FFmpeg patch impact.
Copyright and SPDX hook reads every language; exceptions are declared¶
check-copyright(.pre-commit-config.yaml) selects files by the extension regex that equals thecaselists ofscripts/ci/check-copyright.sh; a rebase keeps the two lists equal and never re-adds a path-name skip to the script. A file that cannot carry a line goes into.config/lint-exceptions.d/<rule>.tomlwith a reason and an expiry (scripts/ci/lint_exceptions.py). The ten Pelorus mirror files stay byte-identical to the pin (ADR-1113);compat/python-vmaf/core/adm_dwt2_cy.pyxkeeps its first line,# SPDX-License-Identifier: BSD-2-Clause-Patent, when upstream Netflix is synced. No score, public API or FFmpeg patch impact.
Pelorus pin moves to the qp_report_csv initialisation fix¶
PELORUS_VENDOR_SHAmoves to42cb17106a2d(VMAFx/pelorus #79:csv_colsstarts at -1 inx265_csv_read_rows()). The mirror stays verbatim (ADR-1113): a conflict incore/src/interop/pelorus_*.ctakes master's side and re-runsscripts/sync-pelorus-interop.sh --update; never merge a hunk by hand. No score, public API or FFmpeg patch impact.
clang-format covers the HIP and Metal kernels¶
- The second
clang-formatentry of.pre-commit-config.yaml(clang-format-hip-metal) andCLANG_FORMAT_FILESin theMakefileread.hipand.metal; an upstream sync or a rebase of a kernel formats it with the pinned clang-format before committing (pre-commit run clang-format-hip-metal --files <file>). The 18 files formatted here changed line breaks only, so a conflict in one of them is resolved by taking the incoming side and re-running the formatter. No score, public API or FFmpeg patch impact. - Collector owns mounted models (
refactor/hiss-zero-native-rust, ADR-1755):VmafModelgainedstruct VmafRef *owners(core/src/model.h);vmaf_model_ref()andvmaf_model_destroy()now live incore/src/model_lifetime.c, compiled into thepredict_carchive, andcore/src/read_json_model.{c,cpp}create the owner count.vmaf_feature_collector_mount_model()takes an owner and the unmount path drops it (core/src/feature/feature_collector.cpp). An upstream sync that touchesvmaf_model_destroy()inmodel.c, the loaders or the mount / unmount helpers keeps all four together; a model must never be freed with a plainfree(). The RustDropimpls leak instead of aborting. No score or FFmpeg patch impact; the public header only gains documentation.
SYCL psnr_hvs scan helpers and once-read SYCL env switches (ADR-1142)¶
Once-read SYCL env switches (ADR-1142)¶
-
core/src/sycl/common.cppreadsVMAF_SYCL_PROFILE,_TIMING,_IMPORT_DEBUGand_CHECKSUMthroughvmaf_gpu_dispatch_env_get(); a sync keeps that and does not bringgetenv()back.core/src/feature/ssimulacra2_eotf_lut.hkeeps itsNOLINTblock (andscripts/gen_ssimulacra2_eotf_lut.pyemits it). No score or public API impact. -
HISS native batch 2 (
refactor/hiss-zero-native-x86): the x86 SSIMULACRA 2 kernels (core/src/feature/x86/ssimulacra2_avx2.c,ssimulacra2_avx512.c,ssimulacra2_host_avx2.c) are drivers over static helpers (vector block, scalar tail pixel, shared IIR step). An upstream or fork change to one of these functions edits the helper that holds the changed statement; the intrinsics, the FMA pattern and the summation order stay as the ADR-1205 / ADR-1208 contracts fix them. No score, public API or FFmpeg patch impact.
Server contracts name the default model; embedded OpenAPI follows the YAML¶
fix/server-default-model-docs.
scripts/ci/check-default-model-single-source.shreadsproto/*.proto,api/openapi/*.yamlanddocs/server/*.md; a sync that brings back a "defaults tovmaf_v0.6.1" sentence in them fails the gate. Name the library default or drop the sentence.gen/go/oapi/vmafx_server_v1.gen.gois regenerated wheneverapi/openapi/vmafx-server-v1.yamlchanges (header kept, seegen/go/AGENTS.md);TestEmbeddedSpecMatchesContractfails otherwise. No score, public C API or FFmpeg patch impact.
ROCm 10.1.0 and the Renovate base-image coverage¶
- ROCm installs under
/opt/rocm/core-10.1:dev/Containerfile(rocm-srcstage) anddocker/Dockerfile.nodename that directory, andtools/rc1-tester/image/hip-runtime.jsonnames the LLVM 24 sonames. A conflict there takes the side that matchesROCM_VERSIONinbuild-config.env. The prune lists of therocm-srcstage and ofscripts/ci/install-rocm-from-image.shstay one list (RPP joined both). scripts/ci/clang-tidy-hip.shpasses--extra-arg=--cuda-host-onlyfor.hipfiles: ROCm 10.1.0's clang-tidy otherwise analyses the device job and the hip lane's counts move. Keep the flag when the wrapper or the hip lane's Makefile recipe changes.renovate.jsonselectsdocker/Dockerfile.testerin the base-image custom manager and in the built-in manager's disable rule;scripts/ci/tests/test_renovate_file_patterns.pyderives that set from the tree, so a sync that adds a Dockerfile mirroring abuild-config.envimage key adds it to both lists. No score, public API or FFmpeg patch impact.
black and ruff read every Python file¶
- The
blackandruff-checkhooks of.pre-commit-config.yamlhave nofiles:filter; theirexcluderegexes are exactly the files of.config/lint-exceptions.d/{black,ruff}.toml, whichpyproject.toml'sextend-excluderepeats (scripts/ci/tests/test_python_format_scope.py). A rebase that adds an exception edits the three together. An upstream Netflix sync ofcompat/python-vmaf/resource/*.pyformats the incoming file with black before committing (the 30 files here were reformatted; a conflict takes the incoming side and re-runs black).core/test/test_*_contract.pyfiles carry named constants for the counts they assert; a change to a counted construct changes the constant. No score, public API or FFmpeg patch impact.
torch only in the training packages (ADR-1886)¶
vmaftune.predictor_trainis nowai/src/vmaf_train/predictor_train.py, with its tests inai/tests/; an upstream-independent fork file, so no Netflix sync touches it. A change that re-adds the module to vmaf-tune, or torch to any pyproject outsideai/andtools/ensemble-training-kit/, failsscripts/ci/check-torch-scope.py.mcp-server/vmaf-mcp/src/vmaf_mcp/vlm.pyis the only VLM path ofdescribe_worst_frames(ONNX Runtime GenAI, localVMAF_MCP_VLM_MODEL); keep thevlmextra free of torch and transformers. Thevmaf-tune-traintest suite is removed from.github/test-suites.jsonandtests-and-quality-gates.yml; a conflict there takes the side without it. No score, public C API or FFmpeg patch impact.
A feature score of a picture read waits for the worker threads (2026-10-06)¶
fix/feature-score-fed-frame-einval (T-ENGINE-READ-FED-FRAME-EINVAL-2026-10-06).
core/src/libvmaf.c::vmaf_feature_score_at_index()fences on-EINVALas well as-EAGAINwhen the index is at most the last picture read (have_last_index/last_index). Upstream Netflix/vmaf returns the collector's answer with no fence; an upstream sync that touches this function keeps the fork's body. On the RC4 branches the body lives invmaf_engine_feature_score_at_index()with the same condition (rc4/api-motion-incremental); a merge keeps one copy of it.- New test
core/test/test_feature_score_fed_frame.cand its block incore/test/meson.build. No score or golden impact: a call that returned a score before returns the same score; only-EINVALfor a picture still with a worker becomes that picture's score.
A model collection's per-frame score reads its stored values first¶
fix/model-set-score-idempotent (T-MODEL-SET-SCORE-NOT-IDEMPOTENT-2026-10-05).
core/src/libvmaf.cgainsread_predicted_collection_score(), whichvmaf_score_at_index_model_collection()calls beforevmaf_predict_score_at_index_model_collection(): a frame whose four named bootstrap scores are already in the collector returns them. Upstream Netflix/vmaf predicts every time and has the same failure; an upstream sync that touches this function keeps the read.- New test
core/test/test_model_collection_score_repeat.cand its block incore/test/meson.build. No score or golden impact: a first prediction is unchanged and a repeat returns its stored values.
Pelorus re-vendor at the tidy-clean commit (RC3 exit, 2026-10-05)¶
rc3-revendor-pelorus-2. PELORUS_VENDOR_SHA moves to 5f5614b0229d (VMAFx/pelorus #78). The ten vendored files are rendered by scripts/sync-pelorus-interop.sh --update, never edited by hand; a rebase that conflicts in them takes either side and re-runs the script, then the drift check.
govulncheck gate and the Go OpenVEX document (ADR-1899)¶
scripts/ci/govulncheck-gate.pyruns ingo-ci.ymlaftergo vet;GOVULNCHECK_VERSIONlives inbuild-config.envwith its Renovate manager. A finding that is not called needs a statement insecurity/vex/go.openvex.json; a "not present" justification covers module-level findings only. No upstream file is involved; no score, public API or FFmpeg patch impact.
The process log level is atomic¶
fix/log-level-atomic (T-LOG-LEVEL-GLOBAL-DATA-RACE-2026-10-06).
core/src/log.cppkeepsvmaf_log_levelandisttyasstd::atomic<int>:vmaf_set_log_level()stores them relaxed,vmaf_log()loads them relaxed (the tty flag once per line, intotty). An upstream change to the logger keeps the atomics; upstream's plain globals race as soon as two threads create contexts or one logs while another creates one.core/src/log.cis not built (ADR-0708) and is unchanged.- New test
core/test/test_log_level_threads.cand its block incore/test/meson.build. No score, output or golden impact.
golusoris modules composed in bootstrap (ADR-1899)¶
internal/app/bootstrap/bootstrap.godefinesCore(config, log, clock, id, validate) andHTTP(router, server);BaseusesCore, andcmd/vmafx-server/cmd/vmafx-controllertakebootstrap.HTTP. No vmafx file imports the golusoris root package: it imports every golusoris module and broughtx/crypto/md4,x/crypto/argon2and 59 otherwise unused modules into the build. A conflict ingo.mod/go.sumtakes this side and rerunsgo mod tidy. No upstream file is involved; no score, public API or FFmpeg patch impact.- HISS native batch 1 (
refactor/hiss-zero-native-1):niqe_extract_aggd()incore/src/feature/niqe_math.his nowniqe_aggd_moments(),niqe_aggd_gamma_index()and a short tail. An upstream or fork change to the AGGD fit edits the helper that holds the changed line; the float operations and their order are fixed by the NIQE snapshot. No score, public API or FFmpeg patch impact.
The Metal host code at zero clang-tidy findings (RC3 exit, 2026-10-06)¶
rc3-tidy-metal-zero. core/src/metal/objc_handle.h (the vmaf_metal::borrow<>() / retain_to_slot() / transfer() bridges and vmaf_metal_library_load(), implemented in kernel_template.mm) is the only place a handle slot becomes a Metal object or the embedded metallib is read: a Metal twin that comes from upstream or from another branch with its own libvmaf_metallib_start / (__bridge ...)(void *)slot code takes the helper instead. The .mm files keep their file-scope helpers and types in anonymous namespaces. The metal lane reads only .mm / .c units and objc_handle.h (--header-filter in tidy-metal.yml); the other headers are the cpu lane's, read as C. A rebase takes master's side of a conflicting .mm hunk and re-runs the Tidy Metal workflow with fix=true; tidy-baseline-metal.json is generated.
vmaf_cuda_picture_get_pix_fmt() is a fork accessor (PR #1118)¶
vmaf_cuda_picture_get_pix_fmt() (core/src/cuda/picture_cuda.h, defined in picture_cuda.c as return pic->pix_fmt;) sits next to vmaf_cuda_picture_get_stream() and the event accessors. The PR #1067 refactor dropped both the definition and the declaration, which broke the link of any CUDA extractor that calls it; PR #1118 restored them. Upstream Netflix has no such function, so a sync or a rebase of a branch that predates #1118 takes the fork's side of both hunks and keeps the accessor. No score, public API or FFmpeg patch impact.
float_adm debug key and unsuffixed_debug_key (ADR-2056)¶
VmafFeatureExtractor gains unsuffixed_debug_key; float_adm.c and the four twins set it to adm, and the twins list adm_scale0 where they listed adm. float_adm.c is a Netflix file: the fork adds one line to its extractor table. A sync keeps that line and the refuse_debug_key_collision() call in core/src/fex_ctx_vector.cpp. No score, public API or FFmpeg patch impact.
x86 AVX2 level requires FMA (ADR-2055)¶
core/src/x86/cpu.c is dav1d's CPU probe with one fork change: the AVX2 flag is set only when CPUID leaf 1 ECX bit 12 (FMA) is set as well as BMI1, BMI2 and AVX2, because two AVX2 kernels are built with -mfma. A sync that takes upstream's cpu.c keeps the has_fma test before the leaf 7 read. core/test/test_x86_cpu_gate.c compiles the file with a mock CPUID and fails without it. No score, public API or FFmpeg patch impact.
MATLAB MEX sources are linted and edited (ADR-2062)¶
The ten MEX sources of compat/python-vmaf/matlab/ are Netflix training-harness files that the fork now edits for lint (braces, static, const, includes) and one defect (ical_std.c destroyed the data pointer of a matrix instead of the mxArray). An upstream sync takes upstream's text and re-applies clang-tidy -fix through make tidy-ratchet LANE=cpu. edges-orig.c also gets FILTER renamed to REDUCE (the name convolve.h defines). No score, public API or FFmpeg patch impact.
float_ms_ssim_cuda builds level 0 on the device (fix/cuda-ms-ssim-device-level0)¶
float_ms_ssim_cuda converts level 0 of its pyramids on the device (ms_ssim_picture_to_float in core/src/feature/cuda/integer_ms_ssim/ms_ssim_score.cu) instead of copying each plane to pinned host memory, running picture_copy() there and uploading the floats (T-CUDA-MS-SSIM-HOST-STAGING-2026-10-06). An upstream sync of the MS-SSIM CUDA twin keeps the device conversion and must not bring back h_input_uint, the per-plane h_ref / h_cmp staging or the #include "picture_copy.h"; a conflict in ms_ssim_stage_inputs() takes this side. test_cuda_float_ms_ssim_exact_contract.py (device-free) and test_cuda_float_ms_ssim_host_traffic (on a device) fail on the host staging. No score, public API or FFmpeg patch impact.
.gitattributes pins the files praetorctl hashes to LF (fix/gitattributes-praetor-hashed-lf)¶
.gitattributes holds a fork block of eol=lf rules above the praetor managed block: the archetypes .standards.lock pins, .standards.*, AGENTS.md, its six compiled targets, the persona sources under .agents/ and their four projections (T-WINDOWS-CRLF-PRAETOR-HASHED-FILES-2026-10-06). Upstream's .gitattributes has none of these paths; a sync keeps the block where it is (the managed block must stay at the tail). scripts/ci/tests/test_praetor_hashed_files_lf.py fails when a rule is missing. No score, public API or FFmpeg patch impact.
nvcc on Windows uses the build's MSVC (fix/nvcc-ccbin-build-msvc)¶
The Windows discovery block of core/src/meson.build (ported from the unmerged Netflix PR #1472, ADR-0150) now gives nvcc the build's own cl.exe when cxx is MSVC (nvcc_build_msvc, assigned just before the block), otherwise the newest toolset under the latest vswhere install, otherwise cl on PATH (T-WINDOWS-NVCC-CCBIN-OLDEST-TOOLSET-2026-10-06). Upstream master has no such block; a re-port keeps this order and core/test/test_windows_cuda_compiler_discovery.py. No score, public API or FFmpeg patch impact.
Option numbers parse in the C locale (fix/numeric-options-c-locale)¶
core/src/opt.cpp parse_double(), core/src/dict.cpp dict_normalize_numeric() and core/src/feature/feature_name.cpp format_double_c_locale() read and write option numbers inside a thread C-locale scope (vmaf_thread_locale_push_c() / _pop(); CLocaleScope in dict.cpp) (T-OPTION-NUMBERS-CALLER-LOCALE-2026-10-06). Upstream's opt.c, dict.c and feature_name.c call strtod() and snprintf("%g") in the caller's locale; a sync that ports a change to either function keeps the scope. The test programs that compile dict.cpp or opt.cpp on their own (test_dict, test_opt, test_feature) link thread_locale.cpp. test_locale_handling fails when either scope is missing. No score, public API or FFmpeg patch impact.
The Windows hooks job copies origin refs into its scratch clone (fix/windows-hooks-scratch-origin)¶
.github/workflows/standards-gate.yml windows-hooks runs the pre-commit hooks in a git clone --shared scratch copy and then fetches the checkout's refs/remotes/origin/* into it, because hooks such as check-research-digest-ids resolve origin/master and a pull-request checkout has no local master (T-CI-WINDOWS-HOOKS-SCRATCH-NO-ORIGIN-MASTER-2026-10-07). Keep the fetch if the step is reworked; scripts/ci/tests/test_windows_hooks_scratch_origin.py fails without it. No score, public API or build change.
The float extractors refuse depths they do not scale (fix/refuse-unsupported-bit-depths)¶
core/src/feature/feature_extractor.cpp refuse_unscaled_bpc(), called first in vmaf_feature_extractor_context_init(), makes the float_ssim, float_ms_ssim, float_adm, float_vif and float_motion families (every backend's twin, matched by name prefix) return -EINVAL for bpc other than 8, 10, 12 and 16 (T-ODD-BIT-DEPTHS-SILENT-WRONG-FLOAT-SCORES-2026-10-07): picture_copy() and the twins that mirror it scale 10, 12 and 16 bits only. When the odd-depth support of RC4 (#2378) lands, remove the families it fixes from the list in the same PR, with the twin matrix at those depths. test_read_pictures_bpc fails without the guard and checks that psnr_hvs still scores 9 and 11 bits. No score at 8, 10, 12 or 16 bits changes.
The SYCL dma-buf import keeps the caller's descriptor (fix/sycl-dmabuf-fd-ownership)¶
core/src/sycl/dmabuf_import.cpp gives Level Zero a private duplicate of the caller's descriptor (driver_fd() / driver_fd_done()), because compute runtime 26.35 closes the descriptor of a re-import. A rebase keeps the duplicate on every zeMemAllocDevice() import path; the RC4 integration branch carries the same change in vmaf_sycl_dmabuf_import_queue(). core/test/test_sycl_dmabuf_fd_ownership.c (GBM, Intel render node, skipped without them) guards it. No ABI, golden-data or FFmpeg patch impact.
-qpfile on libx264 through quant_offsets (fix/x264-qpfile-quant-offsets)¶
RC4 WP15 (ADR-2167) changes the libx264 hunks of ffmpeg-patches/0007-libvmaf-tune-qpfile-unified.patch, adds ffmpeg-patches/test/qpfile_check.py, and makes pkg/saliency and tools/vmaf-tune pass -qpfile for libx264.
Invariants a rebase keeps:
libavcodec/libx264.chas nox264_param_parse(.., "qpfile", ..): libx264 has no such key.X264_init()loads the file withff_qpfile_load()and refusesaq-mode=0and a block grid other than the video's macroblock grid;setup_frame()callssetup_qpfile()after the ROI side-data block, withqpf_framescounting input frames;X264_close()frees the file. An upstream change tosetup_roi()orsetup_frame()keeps the three.pkg/saliency.ExtraParamsFor("libx264")andvmaftune.saliency.augment_extra_params_with_qpfile()return-qpfile;-x264-params qpfile=comes back only with a libx264 that has the key.
No score or public C API impact.
Tool FFmpeg argv fixes (fix/tune-ffmpeg-argv)¶
RC4 WP15: pkg/ffencode, pkg/corpus, pkg/predictor, pkg/hdr and tools/vmaf-tune (encode.py, executor.py, hdr.py, predictor_features.py) merge repeated encoder-parameter options, convert the shot start to seconds, and drop -master_display / -max_cll for hevc_nvenc; pkg/hdr/testdata/python_hdr.json is the Python dump the Go test replays and lost the two options. Invariants: see docs/development/rebase-sensitive-invariants.md ("Encoder-parameter options are merged"). No score or C API impact.
Cross-device parity report fails closed (2026-10-05)¶
Fork-only: ai/src/vmaf_train/cross_backend.py (CrossBackendReport.ok, unbound), the new scripts/ci/tiny_ai_cross_device_parity_gate.py and its tests. Keep the empty-list guard in ok on a sync; no upstream file is involved and no score, public API or FFmpeg patch changes.
Mini retrain, stage runner and the motion metric alias (2026-10-05)¶
Fork-only: ai/src/aiutils/pipeline.py, mini_corpus.py, retrain_checks.py, ai/scripts/mini_retrain.py, ai/e2e/, .github/workflows/mini-retrain.yml are new. ai/data/feature_extractor.py gains _METRIC_KEY_ALIASES and a branch in _lookup(); ai/scripts/extract_full_features.py gains --assume-dims. On an upstream sync that touches either file keep both additions. .github/test-suites.json gets a mini-retrain suite and the Makefile two targets; on a conflict keep both sides.
VMAFx API prototype: four libvmaf entry points become generated shims¶
rc4/api-generation-prototype, ADR-1852.
core/src/libvmaf.crenames the bodies ofvmaf_init,vmaf_close,vmaf_versionandvmaf_feature_score_at_indextovmaf_engine_init,vmaf_engine_close,vmaf_engine_versionandvmaf_engine_feature_score_at_index(declared incore/src/vmafx/engine.h, not exported) and addsVmafContext.api_ownerplus four small accessors. The public functions are defined in the generatedcore/src/vmafx/compat_libvmaf_gen.con the newvmafx_*API. An upstream change to one of the four bodies is ported into itsvmaf_engine_*function; re-adding the old definition is a duplicate symbol.- Generated files (
core/include/vmafx/*.h,core/src/vmafx/*_gen.*,core/test/test_vmafx_abi_layout.c,bindings/python/vmafx/_api.py,docs/api/vmafx/reference.md) are regenerated, never merged by hand:python3 scripts/codegen/vmafx-api.py --write. core/test/check_exported_symbols.pyacceptsvmafx_symbols declared undercore/include/. No score or FFmpeg patch impact; libvmaf return values are unchanged (the shims return the engine's own errno).
VMAFx API generator: header split and versioned vmafx_ symbols¶
rc4/api-wp1-generator, ADR-1852.
- The generated headers are now
core/include/vmafx/{vmafx,version,types,error,context,device,frame,model,score,provenance,report,dnn,mcp,libvmaf_bridge}.h;core/include/vmafx/meson.build(their install list),core/src/vmafx.map,core/src/vmafx.def,core/src/vmafx_symbols.txtanddocs/api/vmafx/<header>.mdare generated too. Regenerate on a conflict, never merge by hand:python3 scripts/codegen/vmafx-api.py --write. core/src/meson.buildlinkslibvmafwith-Wl,--version-script=core/src/vmafx.map -Wl,--no-undefined-versionon ELF targets (vmaf_link_args,link_depends). A sync that rewrites thelibrary('vmaf', ...)call keeps both; thevmaf_*exports stay unversioned while[api] hide_unlisted = false.core/test/check_exported_symbols.pytakes a third argument (the symbol list) and judgesvmafx_exports by it, not by the header regex.- No score, FFmpeg patch or
libvmaf.himpact.
Tester legs build where their inputs change; the cut checks them (2026-10-07, ADR-2198)¶
Fork-only CI: windows_tester_zip_sycl in .github/ci-impact.json, own_input_lanes in .github/ci-tier.json (read by scripts/ci/ci_tier.py), the light gate of windows-tester-bundle.yml, run-name on the three tester workflows and scripts/release/check-candidate-legs.py. Keep the lane's impact and gate on outputs.light and the SYCL selector a superset of the x64 one on a sync. No upstream file, score, public API or FFmpeg patch is involved.
Pelorus re-vendor at the _wfsopen commit (2026-10-07)¶
refactor/pelorus-revendor-wfsopen, ADR-1113. PELORUS_VENDOR_SHA moves to 4aae30711c65 (VMAFx/pelorus #89). The second local edit of core/src/interop/pelorus_qp_report_csv.c (_wfsopen, added by fix/msvc-zero-warnings-crt) is now pelorus's own code, so the mirror carries only the banner and the include rewrite again and scripts/sync-pelorus-interop.sh reports no drift. A sync takes pelorus's side of every vendored file. no upstream file.
VMAFx core API: engine entry points, per-thread log sink, shared picture helpers¶
rc4/api-wp2-core, ADR-1852, ADR-1906.
core/src/libvmaf.crenames the bodies ofvmaf_use_feature,vmaf_use_features_from_model,vmaf_use_features_from_model_collection,vmaf_import_feature_score,vmaf_set_perceptual_weight_enabled,vmaf_set_perceptual_weight_strength,vmaf_feature_backend_twin,vmaf_registered_feature_extractor,vmaf_read_pictures,vmaf_score_at_index,vmaf_score_at_index_model_collection,vmaf_feature_score_pooled,vmaf_score_pooledandvmaf_score_pooled_model_collectiontovmaf_engine_*(declared incore/src/vmafx/engine.h) and keeps the libvmaf names as one-line forwarders in a block near the end of the file. An upstream change to one of these bodies goes into itsvmaf_engine_*function; engine-internal callers (the pooling loops, the Metal import, the tiny-model registration) call thevmaf_engine_names. New helpers there:vmaf_engine_frame_retention,vmaf_engine_is_flushed,vmaf_engine_extractor_backend.core/src/log.cppandcore/src/log.hgainvmaf_get_log_level()and a per-thread sink (VmafLogSink,vmaf_log_swap_thread_sink(),vmaf_log_thread_sink()): while one is installedvmaf_log()delivers to it, filtered by the sink's level. Keep the sink check invmaf_log()on an upstream sync of the logger.core/src/log.cis not built (ADR-0708 moved the logger tolog.cpp) and is unchanged.- Full log routing (ADR-1906):
struct ThreadDataBatchincore/src/libvmaf.ccarrieslog_sink, set fromvmaf_log_thread_sink()where the job is enqueued, andthreaded_extract_batch_func()installs it around the job and restores the previous sink before it returns. A sync that rewrites the job or adds anothervmaf_thread_pool_enqueue()caller keeps both.core/src/thread_pool.cis unchanged. vmaf_engine_init()(the formervmaf_init()body incore/src/libvmaf.c) no longer callsvmaf_set_log_level();vmafx_context_create()does, for a context without a log callback (everyvmaf_init()context). An upstream sync that touches the init body keeps the call out. The atomicvmaf_log_level/isttyofcore/src/log.cppcome from master (PR #2207, T-LOG-LEVEL-GLOBAL-DATA-RACE-2026-10-06): when this branch rebases onto it,log.cpptakes master's atomics and keeps this branch's sink (thread_sink,log_to_sink(), the sink branch invmaf_log(),vmaf_get_log_level()as a relaxed load).- The error prints of
core/src/feature/adm.c,ssim.c,ms_ssim.c,motion.candvif.c(allocation and stride errors,printfto stdout plusfflush(stdout)) arevmaf_log(VMAF_LOG_LEVEL_ERROR, ...)with the same text, and the files includelog.h. An upstream change to one of these lines keepsvmaf_log();core/test/test_engine_log_routing_contract.pyfails on a direct stdout / stderr write.vifdiff()and theVIF_OPT_DEBUG_DUMPoutput invif.ckeep their prints. core/src/picture.cexportsvmaf_picture_plane_extents()(the plane geometrypicture_compute_geometry()used inline);core/src/model.caddsvmaf_model_builtin_data()(the embedded bytes of a built-in model).core/src/vmafx/gainsdevice.c,frame_host.c,model.c,options.c,register.c,score.c,sha256.c,sized.c,submit.cand the internal headersinternal.h,options_internal.h,sha256.h. Generated files as before: regenerate withpython3 scripts/codegen/vmafx-api.py --write.- No score impact (
test_vmafx_bitexactcompares every score with thelibvmaf.hpath; golden gate green), no FFmpeg patch impact; libvmaf return values are unchanged.
Pelorus re-vendor at the fixture _fsopen commit (2026-10-07)¶
refactor/pelorus-revendor-fsopen-fixture, ADR-1113. PELORUS_VENDOR_SHA moves to 11e183ec0aed (VMAFx/pelorus #91): the conformance fixture body of core/test/test_pelorus_interop.c opens its files for reading through fixture_open_read() (_fsopen on Windows). Rendered by scripts/sync-pelorus-interop.sh --update; a sync takes pelorus's side and re-runs the script, then the drift check. no upstream file.
icx-cl: the CRT's deprecated calls (2026-10-07)¶
fix/icx-cl-crt-residuals. Upstream-mirror files touched: core/src/libvmaf.c (VMAF_STRDUP in the tiny-model attach), core/test/test_model.c and core/test/test_output.c (vmaf_fopen_utf8(), vmaf_tmpfile_portable()). A sync that brings a plain strdup / fopen / tmpfile / getenv back into a file the Windows builds compile keeps the fork's spelling (compat/crt_portable.h, compat/path_utf8.h). Fork files: compat/crt_portable.h gains vmaf_tmpfile_portable(); vmaf_tiny_ai_resolve_model_path() takes a caller-owned buffer for the environment value (VMAF_TINY_AI_ENV_PATH_MAX), and its five callers pass one. core/src/feature/common/macros.h (upstream-mirror) defines UNUSED_FUNCTION as the GNU attribute for every GCC or Clang front end, clang-cl and icx-cl included (they define _MSC_VER); a sync keeps that condition. The SYCL leg's configure step no longer sets /experimental:c11atomics.