Alert runbooks¶
One page per alert of the generated rule file (deploy/prometheus/vmafx-rules.yaml). Every alert carries a runbook_url annotation that links its page here. Each page says what the alert means, what it costs, how to find the cause and how to fix it.
| Alert | Severity | Page |
|---|---|---|
VMAFxComponentDown | critical | Component down |
VMAFxNoLiveNodes | critical | No live nodes |
VMAFxQueueAging | warning | Queue aging |
VMAFxJobErrorBudgetBurn | critical (fast), warning (slow) | Job error budget burn |
VMAFxScoreErrorBudgetBurn | critical (fast), warning (slow) | Score error budget burn |
VMAFxScoreLatencyBudgetBurn | critical (fast), warning (slow) | Score latency budget burn |
VMAFxScoreRegression | warning | Score regression |
VMAFxMetricsReadErrors | warning | Metrics read errors |
The burn-rate alerts follow the multi-window, multi-burn-rate pattern. With the default settings the fast rule (critical) fires when an hour and its last five minutes spend the error budget 14.4 times faster than the 30-day objective allows, which uses 2 % of the budget in that hour; the slow rule (warning) fires at 6 times over six hours and their last thirty minutes, 5 % of the budget. Both windows have to agree, so an alert clears soon after the cause does. The default objectives are 99 % for each of jobs, Score request errors and Score requests within 30 seconds. The objectives, windows, factors and the other alerts' thresholds are settings: chart values on Kubernetes (monitoring on Kubernetes), the same keys for a rule file rendered with go run ./tools/obsgen -render-rules.