Skip to content

Research digests

The site search finds a digest by its title through the title list; digest bodies are not in the search index (ADR-1512). To search their text, use GitHub code search, for example repo:VMAFx/vmafx path:docs/research/ zensical.

Iteration-time research notes for the VMAFx fork. Each digest captures what was investigated and why for a fork-local workstream — source links, alternatives weighed, prior art, dead ends.

These are not ADRs:

  • An ADR records a decision and its alternatives at the moment it was made. The body is frozen once Accepted.
  • A research digest records the learning behind that decision (and the iterations that followed). It can be amended as new evidence arrives, the same way a lab notebook is.

A typical workstream has one ADR (the decision) and one research digest (the supporting investigation). Some PRs reuse an existing digest by linking; that is fine.

When to write one

Required by ADR-0108 on every fork-local PR that makes a non-trivial design choice. PRs without a design choice (e.g., a one-line bug fix in fork-added code) state "no research digest needed: trivial" in the PR description and skip the file. Reuse over duplication: if the workstream already has a digest, link it from the new PR instead of starting a parallel one.

Format

Each file is named NNNN-kebab-case-topic.md with a 4-digit zero-padded ID assigned in commit order. The structure mirrors 0000-template.md:

# Research-NNNN: <short, descriptive title>

- **Status**: Active | Superseded by Research-MMMM | Archived
- **Workstream**: <ADR-NNNN, ADR-MMMM, ...>
- **Last updated**: YYYY-MM-DD

## Question         — what was the unknown going in
## Sources          — papers, upstream docs, Netflix issues, prior PRs
## Findings         — what was learned, with citations
## Alternatives explored — what didn't work and why
## Open questions   — what is still unknown
## Related          — ADRs, PRs, issues

Conventions:

  • IDs are assigned in commit order and never reused.
  • Digests are amendable — update the Last updated date when you add findings. To replace one entirely, add Status: Superseded by Research-MMMM and write a new file.
  • Cite sources inline with [link text](URL) so readers can verify.
  • Keep one digest per workstream, not per PR. Cross-link from the PR description.

Index

ID Title Status Workstream
0001 Cache shape for bisect-model-quality nightly Active ADR-0109
0002 Automating process-ADR enforcement (0100 / 0105 / 0106 / 0108) Active ADR-0124
0003 SSIMULACRA 2 port source selection + upstream-drift strategy Active ADR-0126
0004 Vulkan compute backend — loader, shader language, allocator, DMABUF import Active ADR-0127
0005 Embedded MCP in libvmaf — threading, JSON library, SSE server, Power-of-10 fit Active ADR-0128
0006 Tiny-AI PTQ int8 — accuracy targets, ORT API comparison, calibration sourcing Active ADR-0129
0007 SSIMULACRA 2 scalar port — YUV handling, blur deviation, snapshot tooling Active ADR-0126, ADR-0130
0008 MS-SSIM decimate SIMD — FLOP accounting, summation order, bit-exactness Active ADR-0125
0010 Is Netflix about to ship a SpEED-driven VMAF successor? (informational) Active —
0011 _iqa_convolve AVX2 — bit-exactness via __m256d, kernel invariants, Amdahl Active ADR-0138
0012 SSIM SIMD bit-exactness to scalar — where the ULP drifted Active ADR-0139
0013 SIMD DX framework — audit + NEON bit-exactness port Active ADR-0140
0014 psnr_hvs NEON sister port — half-wide split strategy, aarch64 gotchas, QEMU verification limits Active ADR-0160
0015 SSIMULACRA 2 AVX2 + AVX-512 + NEON — per-lane cbrtf, left-to-right summation, 2×2 downsample deinterleave Active ADR-0161
0016 SSIMULACRA 2 IIR blur SIMD — row-batching with gather (horizontal), column-SIMD (vertical), bit-exact to scalar Active ADR-0162
0017 SSIMULACRA 2 picture_to_linear_rgb SIMD — per-lane scalar reads, SIMD matmul, per-lane scalar powf Active ADR-0163
0018 SSIMULACRA 2 snapshot-JSON regression gate — why fork self-consistency beats libjxl/Pacidus cross-check at this scope Active ADR-0164
0031 Intel AI-PC NPU + EP applicability to tiny-AI / dnn/ — verdict: defer NPU; iGPU already covered by OpenVINO EP Active — (backlog T7-9)
0046 vmaf_tiny_v3 (mlp_medium 6→32→16→1, 769 params) vs v2 (mlp_small 257 params): 4-corpus parquet, identical recipe; Netflix LOSO mean PLCC 0.9986 ± 0.0015 vs v2's 0.9978 ± 0.0021 (+0.0008 mean, -29 % std). Decision matrix + per-fold table; ship-alongside-v2 recommendation. Active ADR-0389
0048 vmaf_tiny_v4 (mlp_large, 3 073 params) — does the architecture ladder saturate? Verdict: yes, +0.0001 mean PLCC vs v3 (below 1 std). Ladder stops at v4. Active ADR-0390
0053 iqa_convolve block-of-N tap widen — failed-attempt post-mortem; per-tap widen is load-bearing for bit-exactness, block-of-4 reorder mismatches scalar on 27.67 % of pixels (10 M Monte Carlo) Active ADR-0138
0054 precise decoration audit on vif.comp + ciede.comp — Step A of the Vulkan 1.4 bump path. ciede improves 19× (42/48 → 5/48 mismatches at NVIDIA driver 595.71); vif decorated correctly but the 1.4 regression is not in the tagged float ops. Step B stays blocked. Active ADR-0269, ADR-0264
0055 Root-causes the residual 5/48 NVIDIA-Vulkan ciede2000 places=4 mismatch (1.78× threshold, max abs 8.9e-05) deferred from PR #346. Triangulates double-CPU vs experimental float-CPU vs NVIDIA-Vulkan: f32-CPU matches NVIDIA-GPU to ~6e-7 on the 5 failing high-ΔE frames. Conclusion: structural f32/f64 colour-space-chain precision gap, not a driver fast-math bug. Mitigations rejected; documented as fork debt. Active ADR-0391
0085 Vendor-neutral VVC (H.266) GPU encode landscape — survey of NVENC (Ada+ silicon only), AMD AMF / Intel QSV (decode-only in 2026), VK_KHR_video_encode_h266 (unratified), HIP / SYCL ports of VVenC (3–6 eng-month effort), NN-VC tools via ONNXRuntime EPs (vendor-neutral today), ZLUDA (rejected). Cost / risk / value matrix + three-tier rollout recommendation feeding ADR-0315. Active ADR-0315
0086 Usage-doc coverage audit against ADR-shipped surfaces — 255 ADRs scanned, 46 GOOD / 31 BACKFILL / 178 N/A; identifies 5 highest-leverage gaps (vmaf-tune codec adapters, --score-backend, --cache, Vulkan image import, HDR + sample-clip) for full backfill in this PR; remaining 26 land as ADR-cited stubs. Active ADR-0100, ADR-0167
0090 Phase-A-promotion audit (2026-05-08) — repo-wide scan for surfaces still flagged "Phase A only / scaffold-only / Phase B pending" whose follow-up wiring hasn't shipped. 5 production-blocking promotions (HDR not actually wired into iter_rows; 15 of 17 codec adapters bypass ffmpeg_codec_args; vmaf-tune fast has no CLI subcommand; embedded MCP and HIP runtimes still -ENOSYS), 12 cosmetic doc-drift items, 9 ADRs ready for Proposed→Accepted. Recommended sprint plan + sibling-agent coordination notes. Active ADR-0237, ADR-0300, ADR-0276, ADR-0209
0091 End-to-end integration audit of every shipped libvmaf feature extractor against an 8-rung ladder (CPU → backends → SIMD → corpus → trainer → predictor → docs → tests). 22 extractors inventoried; 0 score 8/8. Engine rungs (1-3) mostly green; learning rungs (4-6) red across the board because CORPUS_ROW_KEYS captures only vmaf_score and ShotFeatures accepts no libvmaf metric outputs. Surprise findings: vmaf_fex_ssim (integer SSIM) is defined but never registered — dead symbol since CPU registration list ships without it. Top-5 promotions ranked by AI-stack ROI. Active —
0126 vmaf-tune HDR dispatch coverage — widens the central hdr_codec_args() table for AV1 NVENC, HEVC/AV1 QSV, HEVC/AV1 AMF, HEVC VideoToolbox, and libaom while keeping private SEI flags limited to verified families. Active ADR-0300
0135 CHUG/K150K extractor I/O cost breakdown and Win 1 + Win 2 optimisations — per-clip cost audit from perf-audit §6; replaces O(N²) parquet flush with at-end-only write via JSONL staging; adds ffprobe skip from CHUG sidecar geometry; decision matrix for in-memory vs streaming vs DuckDB; projected wall-time savings for 5992-clip CHUG run. Active — (perf-audit-pipeline-2026-05-16.md §6)
0136 HDR/UGC dataset license + access audit (2026-05-15) — evaluates 13 candidate corpora from Audit Slice C.7; 6 datasets classified ACTIONABLE-NOW (Beyond8Bits, HDRSDR-VQA, LIVE HDR Database, IPI-MobileHDRVQA, HDR-VDC, CHUG already active); 5 BLOCKED on access or license; 1 BLOCKED on infrastructure. HDRSDR-VQA's 6-display pairwise design surfaces the new panel/display-aware workstream scoped in ADR-0459. Active ADR-0459
1314 Semgrep registry SARIF authority routing Active ADR-1314
1333 Meson test environment secret credential sanitization Active ADR-1333
1366 vmaf CLI frame read-ahead: where the ~7 ms per 4K frame read floor comes from, reader threads per input, pool sizing and shutdown, bit-exactness and CPU/B580 timings Active ADR-1366
1371 motion_sycl tiny-frame parity: the twin blurred each frame where the CPU (since the PR #532 upstream port) blurs the frame difference; per-size before/after on Arc B580 and UHD 770, the CUDA / HIP / Metal twins, kernel cost, and the motion_add_uv host wait Active ADR-1371
1372 CUDA RC3 parity port: motion_cuda moves to the diff-first kernel motion_v2_cuda already had (plus the weighted debug score, device-side ping-pong ordering, one readback wait), the CPU option tables on the PSNR / SSIM / float-motion twins (why identical SSIM windows needed uncontracted products), and the launch-grid traces behind the ADM tiny-height and VIF 16-pixel guards Active ADR-1372, ADR-1373, ADR-1374
1377 HIP RC3 CPU parity: motion_hip's blur-then-diff order and the shared diff-first kernel, CPU motion / motion_v2 SAD equivalence, host replays of the HIP tile and ADM scale-0 row loads, the vif_hip 16-pixel bound, identical-window SSIM exactness, and what runs without an AMD device Active ADR-1377, ADR-1381, ADR-1382
1378 HIP CAMBI and SpEED on the device: why HIP's __f*_rn intrinsics are not rounding primitives, the SpEED kernel IR check, glibc log2f misrounding and what it means against a glibc-built CPU, and the host replays (lockstep c-values with a histogram canary, planted regressions) that pin both twins without an AMD device Active ADR-1378, ADR-1384
2041 Thread-pool backpressure and shutdown lifetime Active ADR-1243
2072 Pelorus v0.2.2 interop parser safety sync Active ADR-1276
2073 FFmpeg n9.0.2 stable refresh Active ADR-1240
2079 SYCL clang-tidy required-gate live verification and fail-closed contract hardening Active ADR-1297, ADR-0623
2080 Required-workflow path-filter closure Completed ADR-1297, ADR-0313, ADR-1140
2082 SYCL upload host-buffer lifetime Active BUG-040, ADR-0214
2083 dev-MCP smoke-probe contract restoration Active BUG-048 A13; no ADR (bug fix)
2085 Restore strict JSON boundaries for external-bench and vmaf-roi-score Active — (BUG048 A9 restoration)
2086 BUG-048 strict JSON emitter restoration Active BUG-048 A10
2089 Archived issue-reference provenance after the repository migration Active BUG-048 Section E
2094 CodeQL SVM solver lifecycle and parser loop alerts Active ADR-0889, ADR-1039
2096 CodeQL unused-static-function identity audit for repeated test compilations Active ADR-1142
2097 Resolving live CodeQL cpp/equality-on-floats alerts Active ADR-1308
2098 Resolution of CodeQL AVX-512 large-parameter alerts Active — (private ABI cleanup)
2099 Non-finite score laundering in the metric engine Active ADR-1302
2100 Feature-collector single C++ source authority and duplicate elimination Active — (BUG-048 section E)
2101 SYCL extractor init-failure resource cleanup restoration Complete — (BUG-048 section E)
2102 SYCL clang-tidy wrapper executable resolution for safe_subprocess Active ADR-1142, ADR-1270
2103 Required Python Lint fail-closed merge-base gate restoration Complete ADR-1310
2104 GPU twin option aliases and collector-key compatibility Complete ADR-1312
2105 GPU option-value capability inventory and model-dispatch fallback Complete ADR-1316
2106 CUDA and HIP device target header dependency tracking Complete ADR-1320
2108 GPU float-SSIM auto-scale fallback Complete ADR-1324
2109 Integer ADM Barten fixed-point normalization Complete ADR-1325
2110 Metal float_ms_ssim option and score parity Complete ADR-1334, ADR-1221
2111 Integer ADM contrast-masking row-rounding observability Complete ADR-1167
2112 SYCL motion-add-UV fixed-point oracle Complete ADR-1326
2113 Metal float-motion force-zero ownership, debug gating, and flush idempotency Complete — (BUG048 A5 restoration)
2114 Trusted merge-base authority for research-digest identity debt Complete ADR-1335
2115 HIP float-motion force-zero ownership and option-aware flush idempotency Active — (BUG048 A5 restoration)
2116 CodeQL Python alerts triage and exception semantics Active —
2117 CUDA context-owned resource teardown Complete ADR-1336
2118 BUG-048 script environment drift Complete — (BUG-048 restoration)
2119 GCC 16 C++ placement new/delete symbol visibility Complete ADR-1337
2120 SYCL ciede2000 throughput: --feature ciede never reached SYCL; ciede_sycl stages chroma at native size Active — (RC3 performance)
2121 --feature with an explicit GPU --backend: twin pairing through the model-dispatch lookup, option/geometry gates, SYCL pairing table and JSON receipt Active ADR-1359
2122 Device-resident CAMBI on SYCL: c-values and exact top-K pooling on the GPU, parity proof and measurements Active ADR-1357
1362 AIM on the SYCL integer ADM twin: one more stored CSF band, one fused per-row reduction, int64 decouple clamp, CPU float finalisation for every output (bit-exact); parity and timings on Arc B580 and UHD 770 Active ADR-1362
2123 Arc B580 SYCL crashes: IGC SIMD32 crash on the psnr_hvs private-memory DCT, tile-halo loads outside the plane, dropped graph-wait errors, integer VIF minimum size and odd-width stride, area-scaled psnr_hvs gate Active ADR-1361
1397 Making a GPU psnr_hvs twin return the CPU's bits: the four differences between calc_psnrhvs() and the per-block twin (summation order, double masking table, double threshold root, contraction), bit-identity on 576x324 to 3840x2160 at 8 to 12 bits on an RTX 4090, the ablation of each difference, the CPU float sum as a bias against a double sum, the cost (2.4 to 12.2 ms per 3840x2160 frame) and the tuning options (zero compaction, order-preserving device reduction) Active ADR-1397, ADR-1361
1401 Bit-exact psnr_hvs on SYCL and HIP: why the CPU's double product and square root in the masking threshold equal one fp32 rounding of the exact root, sqrt_prod_rn() as an integer square root for the fp64-free SYCL kernel (0 mismatches on 450 million operand pairs), bit-identity on an Arc A380 and a gfx1036 from 576x324 to 3840x2160 at 8 to 12 bits, a scratch-free SYCL kernel, the one-ulp log10 difference between an icx and a gcc build, the ablations and the cost (22.4 to 35.9 ms per 3840x2160 frame on the A380, 18.3 to 37.9 on the gfx1036) Active ADR-1401, ADR-1397
2125 First native Windows run of the SYCL build on an Arc B580 and a UHD 770: link.exe never registered the device images (explicit device link and anchor symbol), the Windows test runner returned before Meson finished, harness failures it hid, CPU / Linux parity of every SYCL twin, per-frame timings Active ADR-1364
2127 CPU options on the SYCL psnr / ssim / float_ssim / float_motion twins: where each option acts, identical-frame dB exactness in fp32, per-option CPU parity on Arc B580 and UHD 770, planted-regression checks Active ADR-1365
2128 oneAPI release image: the 25.18 compute runtime from Intel's 2025 runtime image crashes Arc B580 SYCL runs (NEO swap alone fixes it, loader swap does not); the image rebuilt on Debian 13 with Intel's 2026.1 apt packages, pinned NEO and Level Zero loader, 2.39 GB instead of 5.86 GB, scores within the parity gate on B580 and UHD 770 Active ADR-1368
2130 float_ssim decimation on the SYCL device: exactness of the CPU's double window sum, int64 fixed-point reproduction, byte-identical planes on Arc B580 and UHD 770, score parity and 4K timing, the fp32 frame-mean rounding behind enable_db, the L·C·S headline experiment Active ADR-1370
2131 vmaf CLI option dictionaries: which early exits leaked the --feature / --model overload dictionaries (158 to 329 bytes), the default LeakSanitizer scan hiding the overload leak, telling vmaf_use_feature()'s two -EINVAL cases apart, and where the CLI releases them Active T-CLI-PRE-REGISTRATION-OPTS-DICT-LEAK-2026-09-30 (no ADR)
2132 CAMBI on frames shorter or narrower than its window: the affected sizes per width and height, the in-place decimation that makes the scalar walk read another scale's columns, master-versus-branch scores at three dispatch levels, the unit tests, the GPU twins, and why the c-values scan driver keeps a cited cppcheck suppression Active ADR-1393
2133 Licence audit of the published tester image and macOS bundle: every component with version, licence, binary-redistribution obligation and whether it was met; the copyleft parts (Debian base, sqv crates, LGPL libquadmath grafted into wheels, identified by ELF build ID), the interpreter archive without licence texts, the missing SBOMs, and the sizes of the source companion Active ADR-1503
2137 Documentation site redesign inputs: MkDocs 1.x and 2.0 status, Material for MkDocs end of life (2027-05-05) and Zensical measured on this tree (34 s against 257 s, its unsupported settings and 987 strict-mode link warnings), Starlight and Docusaurus migration cost, chart libraries against the house chart style, diagram tooling, content ready for charts and diagrams, and the site's measured typography (81 to 88 characters per line) with targets Active ADR-1508
2138 NVIDIA GPU tester image: the binaries link no NVIDIA library (readelf -d), the NVIDIA code inside the kernels (CUDA headers, libdevice) and the EULA clauses that allow it, the host driver through the NVIDIA Container Toolkit with a read-only root, which GPU runs which cubin or PTX of the default gencode list, and the RTX 4090 run of the documented command Active ADR-1509
2139 AMD GPU tester image: the ROCm 10.0.0 runtime closure of the HIP build (libamdhip64 to the bundled rocm_sysdeps, kept in its RPATH layout), the licence of each part, the source the bundled LGPL libelf and libnuma owe (upstream archives and the TheRock tree, by ELF build ID), the 18 gfx targets, why WSL2 is out, and the gfx1036 run of the documented command Active ADR-1511
2141 Windows tester zip: what the hosted Windows lanes build and run, the licence of every part (the static Microsoft runtime, the interpreter's Visual C++ runtime DLLs from the runner's redistributable folder, python-build-standalone on Windows), the interpreter's import tables and pruning, host facts from the registry and IsProcessorFeaturePresent, the Wine run of the report from a MinGW build, and what only the hosted runner and a tester can prove Active ADR-1515
1413 Integer ADM enhancement gain limit under a non-integer limit: what the scalar stores (the double product truncated toward zero), per-input distances of the AVX2 / AVX-512 kernels and the SYCL twin before the change, the integer form of the truncated double product (two 64-bit products and a bit-length comparison) and its check on 106.8 million pairs, CUDA / HIP conversions in PTX, the madd_epi16 wrap in the scale-0 angle test, seven planted defects, timings Active ADR-1413
1402 Integer ADM masking centre tap in int32: the golden gate before and after, which inputs move and by how much, what the AVX2 / AVX-512 kernels must compute to equal the scalar on every operand (and why upstream's vector form does not), sixteen planted defects, the UBSan reports of the x86 tails, CUDA / HIP / SYCL before-and-after, stage times per dispatch level Active ADR-1402
1369 SYCL psnr_hvs / motion_v2 / psnr at 4K: per-phase profile (host conversion, private uploads, kernels), shared luma + opt-in shared chroma planes, bit-identical kernel reworks, pinned chroma staging, slot fence, CUDA/HIP review Active ADR-1369
1379 Device-resident CAMBI and SpEED on CUDA: the round-to-nearest intrinsics that reproduce the CPU's fp32 rounding, the CPU reference's dependence on its libm log2f and on FMA contraction, the exact top-K sum, the lanczos4 and CPU speed_temporal prescale findings, the exhaustive check of the device log2, and frame-by-frame parity, transfer counts and timing through a host emulation of the CUDA driver and on an RTX 4090 Active ADR-1379, ADR-1380
1417 Integer AIM above 1 on a flat reference: per-scale numerators and denominators of the integer and float pipelines (equal up to the final clip), the two upstream lines, upstream master's output on the same pictures, the 24x24 scale-3 difference and its cause, adm3 and the default model's score with and without a clip, device twins, golden-gate coverage Active ADR-1417

| 0053 | Post-merge CPU profile 2026-05-03 — perf top-10 after lusoris/vmaf#310 through lusoris/vmaf#321; surfaces 3 new opt targets (convolve widen, SSIM double reduction, VIF gather elimination) | Active | — | | 0081 | Real-corpus retrain methodology for the fr_regressor_v2 deep ensemble — corpus-size sufficiency (9 ref + 70 dis @ .workingdir2/netflix/), 9-fold LOSO sizing inherited from the deterministic ADR-0291 baseline, seed-diversity hyperparameters, and the Seeking_25fps weak-fold diagnostic for HOLD-on-spread cases. | Active | ADR-0309 | | 0089 | CPU double vs Vulkan float stage bisect on the residual NVIDIA-Vulkan integer_vif_scale2 45/48-frame places=4 mismatch at API 1.4 (T-VK-VIF-1.4-RESIDUAL). Static SPIR-V re-verification confirms only 5 FP-arithmetic ops in vif.comp and all 5 are NoContraction-decorated post-PR #346 — SPIR-V mitigation surface is exhausted. SYCL counter-example (same f32 contract, passes the gate) rules out a pure f32-vs-f64 class issue. Localised root cause: NVIDIA shaderFloatControls2-v2 codegen flip at API 1.4 on a non-IEEE-bound default (reciprocal-multiply, fast-rsq) outside the SPIR-V declarable surface. Phase-2 shader fix not warranted; recommends per-stage NVIDIA dynamic dump or places=3 override ADR. | Active | ADR-0264, ADR-0269 | | 0090 | Per-commit triage of the 41 upstream commits binned SKIP-doc-or-format by the /sync-upstream Pass-2 heuristic on 2026-05-08. Splits into 5 PORT_NOW (motion_v2 mirroring bugfix, motion_v2 option cluster + prev_prev_ref API, two cambi internals), 18 PORT_LATER (python/test MyTestCase migration, blocked on agent-E worktree), 4 DEFER_INDEFINITELY, 1 PORTED_SILENTLY (662fb9ce semaphores → fork commit e5a52e74), 12 MERGE_BOUNDARY. Surfaces the riskiest item: the python/test mass-port is +5 600 LOC and crosses Netflix-golden assertions in feature_extractor_test.py. | Active | — (companion to Research-0089) | | 0135 | Vulkan dispatch overhead characterization — T7-18: startup dominated by uncached vkCreateComputePipelines; per-frame fence/submit overhead ruled out; pipeline-cache fix recommended | Active | T7-18 | | 0091 | CAMBI CUDA integration trade-offs (T3-15a): per-thread 49-read vs shared-memory SAT for the spatial-mask kernel; synchronous vs async ring-buffer DtoH for the 5-scale pipeline; host_pinned slot reuse for score storage; two compile-time bugs found and fixed (cuMemcpyDtoH arg order, VMAF_FEATURE_DISPATCH_SEQUENTIAL non-existence). Predecessor: Research-0032 (Vulkan twin). | Active | ADR-0360 |

| 0734 | CUDA VIF filter1d.cu ncu hotpath profile on RTX 4090 (sm_89). Primary bottleneck: launch-width-limited (0.84 waves) + register pressure (56 regs/thread). Top kernel filter1d_8_horizontal_kernel_2_17_9 = 35 % of VIF filter time. Three optimization candidates: increase val_per_thread 2→4, reduce register live range, add __ldg() for smem loads. | Active | — | (Index seeded by ADR-0108's adoption PR; backfilled digests for the existing major workstreams will be added as their authors revisit the corresponding code.) | 0135 | CAMBI CUDA spatial-mask SLM tile -- design analysis: img-tile correctness bug, 26x read reduction via direct zd_tile load, bank-conflict accepted at uint8 row access | Active | ADR-0464 | | 0751 | Cross-backend 4K (3840x2160) baseline (CPU + CUDA) and PR #79 adm_cm_line_kernel_8 A/B at 4K. RTX 4090 medians: vif CUDA 147 fps, adm CUDA 161 fps. filter1d fully saturated at 4K (253 waves, 69.7% occ). adm_cm __launch_bounds__ win is zero at 4K (-0.3%) vs -9.3% at 1080p (register-bound regime only). ms_ssim_decimate scale 0 saturated at 4K (88.1% occ). | Active | — |