Observability¶
VMAFx reports what it does through three signals and ships everything to watch them with. This guide is for operators who run the server, the controller and the nodes; the developer guide covers how the telemetry is wired and generated.
| What you get | Where it comes from |
|---|---|
| Prometheus metrics from every long-running component | /metrics of vmafx-server, vmafx-controller, vmafx-node (and the operator's controller-runtime metrics); metric reference |
| Ten Grafana dashboards | deploy/grafana/dashboards/; dashboard tour |
| Eight alerts with a runbook each, and the recording rules they read | deploy/prometheus/vmafx-rules.yaml; alert runbooks |
| Traces | OpenTelemetry over OTLP/gRPC to your collector; OpenTelemetry |
| Settings: SLO objectives, burn rates, alert thresholds, prices | Helm values monitoring.slo, .burnRates, .alerts, .cost, or the same keys in one file for Compose |
Choose how to install it¶
- Kubernetes with the Prometheus operator: set
monitoring.enabledin the Helm chart. It renders a ServiceMonitor or PodMonitor per component, a PrometheusRule with your settings and a ConfigMap per dashboard for the Grafana sidecar. See monitoring on Kubernetes. - One machine, or to try it out: the Docker Compose stack runs VMAFx with Prometheus, an OpenTelemetry Collector, Tempo, Loki and Grafana, already wired together. See the Compose stack.
- Your own Prometheus and Grafana: scrape the three components'
/metrics, load the rule file (render your own withgo run ./tools/obsgen -render-rules -values <file>), and import the dashboards' JSON. See without the Prometheus operator.
The first hour¶
- Check the scrape. In Prometheus,
vmafx_build_inforeturns one series per server, controller and node pod; Status > Targets shows them up. - Open the Overview. Components up matches your pods, Live nodes matches your nodes, and the queue is empty or draining (tour).
- Check the rules. Alerts in Prometheus lists the eight VMAFx alerts, none of them in error. Each alert's
runbook_urlopens its runbook. - Set your objectives. The defaults are 99 % for job success, Score request success and Score requests within 30 seconds. Change them, the burn-rate windows and the thresholds in
monitoring.slo,.burnRatesand.alerts; the chart's schema refuses a value the rules cannot use (settings). - Set prices, if you account usage.
monitoring.cost.perJobSecondandmonitoring.cost.perJob, inmonitoring.cost.currency, fill the cost panels of the usage and cost dashboard. VMAFx assumes no price: without them the cost panels stay empty. - Point the traces somewhere. Set
VMAFX_OTEL_ENDPOINT=collector:4317(orOTEL_EXPORTER_OTLP_ENDPOINT=http://collector:4317) on the components (OpenTelemetry).
When an alert fires¶
Every alert links its runbook: what it means, what it costs, how to find the cause and how to fix it. The runbook index lists them; the dashboard tour says which dashboard each one starts from.