Alert Runbooks
Alert runbooks are split per alert under this directory so each alert annotation links to a dedicated procedure.
Shared Incident Workflow​
- Confirm blast radius (
single pod,single tenant, orfleet-wide). - Validate health checks:
curl -fsS http://<proxy>:3100/readycurl -fsS http://<victorialogs>:9428/health
- Check request failures and latency from metrics.
/metricsis served on the chart's dedicated metrics port (:9091by default); the binary serves it only with-server.register-instrumentation=true(on-listen, or on-metrics-listenwhen set). - Review proxy logs for translation, backend, timeout, or auth errors.
- Apply mitigation, then verify alert recovery criteria.
Runbooks​
- Deployment And Scaling Best Practices
- LokiVLProxyDown
- LokiVLProxyHighErrorRate
- LokiVLProxyHighLatency
- LokiVLProxyBackendHighLatency
- LokiVLProxyBackendUnreachable
- LokiVLProxyCircuitBreakerOpen
- LokiVLProxyTenantHighErrorRate
- LokiVLProxyRateLimiting
- LokiVLProxyClientBadRequestBurst
- LokiVLProxyUnexpectedTupleMode and LokiVLProxyDefault2TupleMissing
- LokiVLProxySystemMetricsMissing, LokiVLProxySystemMemoryHigh, LokiVLProxySystemCPUPressureHigh and LokiVLProxySystemIOPressureHigh
These names match the Helm PrometheusRule alerts in
alerting/loki-vl-proxy-prometheusrule.yaml, whose runbook_url annotations
link here. The standalone rule file alerting/loki-vl-proxy-alerting-rules.yaml
(for Prometheus, vmalert or Grafana alerting) uses a partly different alert set
and carries no runbook links.