Perf gate guide¶
The wall-clock perf regression gate (ADR-0907, status Proposed) compares a fresh multi-resolution benchmark run against a committed baseline. It is not wired into CI, a git hook or a Makefile target. Nothing in the tree calls scripts/perf/check-regression.py or scripts/perf/bench-multi-resolution.sh; you run both by hand, and no job reports a perf regression for you. Wiring the gate in belongs to the benchmarks candidate, RC8 (ADR-1490), because benchmarks and tuning never happen before it.
What the scripts do¶
scripts/perf/bench-multi-resolution.shruns thevmafbinary over resolution, backend and metric combinations and writes the timings to a JSON file you name with--output.scripts/perf/check-regression.pyjoins that file against the committed baseline (testdata/perf_multi_resolution.json) by (resolution, backend, metric) and reports every cell whose median wall time exceeds the baseline by more than--tolerance-pct(default 5). It exits non-zero on a regression;--advisoryprints the report and always exits 0. Cells whose status is notokon either side are reported and skipped, not failed.
The baseline¶
The committed baseline holds 40 ok cells (CPU and CUDA) recorded on 2026-05-29 on one workstation (AMD Ryzen 9 9950X3D, RTX 4090). Timings do not transfer between machines: compare a run only against a baseline recorded on the same hardware and the same build configuration, or refresh the baseline first.
Run it locally¶
# Build first
meson setup core/build -Denable_cuda=false -Denable_sycl=false --buildtype=release
ninja -C core/build
# Benchmark
VMAF_BIN=core/build/tools/vmaf \
scripts/perf/bench-multi-resolution.sh \
--backends cpu \
--runs 3 \
--output /tmp/perf_current.json
# Compare against the committed baseline
python3 scripts/perf/check-regression.py \
--baseline testdata/perf_multi_resolution.json \
--current /tmp/perf_current.json \
--tolerance-pct 5.0 \
--backend cpu
Refresh the baseline¶
Record a new baseline on the machine you want to compare against:
VMAF_BIN=core/build/tools/vmaf \
scripts/perf/bench-multi-resolution.sh \
--backends cpu,cuda \
--runs 5 \
--output testdata/perf_multi_resolution.json
Check that the cells you care about have status: ok, and commit the file only as a deliberate baseline change with its hardware named in the commit message (agent hard rule 13: ad-hoc benchmark output is never committed).
What wiring it in would take (RC8)¶
- A baseline recorded on the runner class the job runs on; the workstation baseline is several times faster than a hosted runner, so a hard 5 % tolerance against it flags noise.
- A job that runs the two scripts and uploads the run JSON, advisory first (
--advisory --skip-if-no-baseline), then blocking once the tolerance has been calibrated against runner variance. - GPU cells on a self-hosted GPU runner, as a second
check-regression.pyinvocation with--backend cuda. - An ADR that moves ADR-0907 out of Proposed and records the tolerance.
Relevant files¶
| File | Purpose |
|---|---|
testdata/perf_multi_resolution.json | Committed baseline (refreshed by hand) |
scripts/perf/bench-multi-resolution.sh | Benchmark harness |
scripts/perf/check-regression.py | Comparison and reporting script |
scripts/perf/test_check_regression.py | Unit tests for the comparison script |
See also¶
- Performance guide: profiling and the same not-wired note.
- ADR-0907: the gate's design.
- ADR-1005: advisory mode and baseline refresh (the CI step it describes does not exist yet).