Self-hosted GPU runner — enrollment guide¶
Follow this runbook to enroll a self-hosted runner for the coverage-gpu CI job (in tests-and-quality-gates.yml). The job needs a runner that exposes both NVIDIA and Intel GPUs alongside an AVX-512-capable CPU; hosted GitHub runners cannot reach those code paths. The Arc-only SYCL parity lane is a separate capability with its own operator guide.
This page is a provisioning runbook, not evidence that the hardware lane has run.
How the job is admitted¶
The coverage-gpu job runs only after a hosted admission job sees both GPU_COVERAGE_ENABLED=true and an online runner carrying every label in its runs-on set.
- A missing, partial or offline match fails the hosted probe and skips self-hosted dispatch, so the hardware job cannot queue forever.
- While the switch is enabled, the required-check aggregator treats that skip as a failure.
This is the ADR-1319 contract.
Note
State on 2026-09-25 (dated snapshot): the repository and organisation runner APIs both returned zero runners, GPU_COVERAGE_ENABLED was absent, and the documented local runner service, environment and state paths were absent. Re-check with gh api /repos/VMAFx/vmafx/actions/runners --jq '.total_count' and gh variable list -R VMAFx/vmafx.
Historical backlog reference: T7-3. The local .workingdir/BACKLOG.md notebook is not part of the published documentation; the linked workflow is the current configuration source.
Required labels¶
The workflows match on a label triple. Match these exactly:
| Label | Meaning |
|---|---|
self-hosted | GitHub default for any non-hosted runner. |
linux | OS family. Bare-metal Linux only — WSL2 and containers can't pass cudaSetDevice to the host driver reliably. |
gpu-full | The fork's combined CUDA + SYCL + AVX-512 tag — at minimum NVIDIA + Intel + AVX-512. It does not claim HIP coverage. Add fine-grained labels (gpu-cuda, gpu-intel, avx512) so future jobs can target a subset without re-tagging. |
Hardware expectations¶
The runner needs to satisfy the union of all jobs that target it. Toolkit versions are the ones pinned in build-config.env (CUDA_VERSION, ONEAPI_VERSION); an older toolkit cannot build the current tree (ADR-1223 raises the CUDA floor to Ampere, compute capability 8.0).
- NVIDIA GPU + driver + the pinned CUDA toolkit — drives the
coverage-gpuCUDA build (-Denable_cuda=true) and the CUDA test suite.nvidia-smimust succeed withoutsudo. - Intel GPU + Level Zero + the pinned oneAPI toolkit — drives the SYCL build (
-Denable_sycl=true).sycl-lsmust list at least one Intel GPU device. - Intel
ocloc— the SYCL build compiles native GPU code ahead of time andmeson setuprefuses withoutocloconPATH(ADR-1360). Runbash scripts/ci/install-intel-ocloc.shonce on the host; it installs the release pinned asINTEL_NEO_VERSIONinbuild-config.env. - AVX-512-capable CPU (Ice Lake or newer / Zen 4 or newer) — the AVX-512 SIMD code paths.
lscpu | grep avx512fshould return a hit. - ≥ 16 GB RAM — coverage + multi-backend builds peak around 8 GB resident; doubling that gives headroom for parallel CPU test processes. Tests tagged
gpurun exclusively because they share one accelerator. - ≥ 60 GB free disk — coverage builds + nightly artifact retention rotate through ~30 GB.
- HIP is not part of this job. An AMD device on the host does not make
Coverage GPUa HIP gate; HIP needs its own executable job and capability label before CI can claim hardware coverage.
A typical workstation that runs the fork's local dev loop already satisfies all of the above. Per maintainer direction (2026-04-25), the primary dev workstation is the first runner.
Enrollment steps¶
These are the canonical GitHub Actions runner setup steps, lightly adapted for this fork's labels.
1. Generate a registration token¶
The token is short-lived (1 hour) and single-use. Generate one in the GitHub UI:
https://github.com/VMAFx/vmafx/settings/actions/runners/new
Or via gh:
2. Install the runner agent¶
Pick a working directory the agent will own (e.g. ~/actions-runner):
mkdir -p ~/actions-runner && cd ~/actions-runner
# Pin to a known release; bump deliberately. 2.319.1 is an example:
# take the current release from https://github.com/actions/runner/releases.
RUNNER_VERSION=2.319.1
curl -O -L \
"https://github.com/actions/runner/releases/download/v${RUNNER_VERSION}/actions-runner-linux-x64-${RUNNER_VERSION}.tar.gz"
echo "<sha256-from-release-page> actions-runner-linux-x64-${RUNNER_VERSION}.tar.gz" | sha256sum -c -
tar xzf "actions-runner-linux-x64-${RUNNER_VERSION}.tar.gz"
3. Configure with the fork's labels¶
./config.sh \
--url https://github.com/VMAFx/vmafx \
--token "<token from step 1>" \
--labels self-hosted,linux,gpu-full,gpu-cuda,gpu-intel,avx512 \
--name "$(hostname)-gpu-full" \
--work _work \
--replace
The fine-grained labels (gpu-cuda, gpu-intel, avx512) are not required by any current workflow but are reserved so future jobs can match a subset of capabilities without forcing a label change.
4. Run as a systemd service (recommended)¶
The service auto-starts on boot and respawns on crash. Logs land in ~/actions-runner/_diag/ and journalctl -u actions.runner.VMAFx-vmafx.*.
5. Verify it's online¶
gh api /repos/VMAFx/vmafx/actions/runners --jq '.runners[] | {name, status, labels: [.labels[].name]}'
The runner should appear with "status": "online" and the full label set. Trigger the coverage-gpu job once with gh workflow run to confirm end-to-end:
gh variable set GPU_COVERAGE_ENABLED -b true -R VMAFx/vmafx
gh workflow run tests-and-quality-gates.yml
Enable the variable last. The hosted probe must first report a complete online self-hosted,linux,gpu-full match. The Coverage GPU check is already required while enabled; a failed probe produces a skipped hardware check that the aggregator rejects.
Operational notes¶
- GPU driver upgrades: stop the agent (
sudo ./svc.sh stop), drain in-flight jobs (the runner finishes the current job before exiting), upgrade, restart. The runner deregisters automatically on prolonged offline. - Disk hygiene: GitHub Actions does not clean
_work/on its own. A nightlyfind _work -mtime +14 -deleteis standard. The fork's coverage artifacts retain for 14 days on the GitHub side, so local rotation at the same cadence is safe. - Secrets: this runner has access to repo + organisation secrets scoped to the
gpu-coverageenvironment if and when one is added. Treat the host as security-sensitive — no third-party SSH keys, no shared user accounts, audit~/.ssh/authorized_keysquarterly. - Fork pull requests: the hosted probe excludes fork heads, so untrusted fork code never reaches this self-hosted runner.
- Concurrency: see GPU test serialisation below.
- Second runner: adding a runner with the same label set (for example a remote Intel-only or NVIDIA-only host) lets
coverage-gpuparallelise with whatever fine-grained-label job comes next without label collisions.
GPU test serialisation¶
A single runner serialises GPU jobs.
- Inside a Meson test job, every test tagged
gpusetsis_parallel : false. Meson drains running tests before each one and does not start another until it finishes. - Keep the normal parallel Meson invocation.
-j1and inflated timeouts hide accelerator contention instead of enforcing the shared-device contract. - The source guard covers dormant backends that cannot be configured on the current host; the metadata guard checks Meson's effective scheduling flags.
Verify both the source registry and the configured metadata:
python3 "$(git rev-parse --show-toplevel)/scripts/ci/run_meson_test.py" -- \
-C build --no-rebuild test_gpu_serialization_contract check_gpu_test_serialization
Decommissioning¶
# Disable admission before taking the runner offline.
gh variable set GPU_COVERAGE_ENABLED -b false -R VMAFx/vmafx
cd ~/actions-runner
sudo ./svc.sh stop && sudo ./svc.sh uninstall
./config.sh remove --token "$(gh api -X POST \
/repos/VMAFx/vmafx/actions/runners/remove-token --jq .token)"
rm -rf ~/actions-runner
References¶
- Historical T7-3 reference:
.workingdir/BACKLOG.md(local notebook, not shipped with this documentation). tests-and-quality-gates.yml§ coverage-gpu — the first consumer of thegpu-fulllabel.- GitHub Actions: self-hosted runners
- Maintainer direction, 2026-04-25 (paraphrased): use the local workstation GPUs for CUDA and Intel testing for now.