Skip to main content

Operations Guide

Deployment​

Minimum Requirements​

ResourceMinimumRecommended
CPU50m200m
Memory64Mi256Mi
Replicas12+ (with PDB)

The proxy is stateless (except optional disk cache). Scale horizontally without coordination.

Key scaling controls (all tunable via CLI flags):

  • -max-concurrent 100 — per-replica in-flight request cap (excess gets 503); also bounds concurrent backend operations
  • -rate-limit-per-second 50 / -rate-limit-burst 100 — per-client token bucket
  • -cb-fail-threshold 5 / -cb-open-duration 10s — backend circuit breaker
  • use Grafana refresh policy, ingress shaping, HPA, and cache tuning as complementary levers

Helm Deployment​

helm install loki-vl-proxy oci://ghcr.io/reliablyobserve/charts/loki-vl-proxy \
--version <release> \
--set extraArgs.backend=http://victorialogs:9428 \
--set extraArgs.label-style=underscores

# Local chart (development)
helm install loki-vl-proxy ./charts/loki-vl-proxy \
--set extraArgs.backend=http://victorialogs:9428 \
--set extraArgs.label-style=underscores

For multi-replica fleets with HPA, prefer peerCache.enabled=true over static peer lists. The chart creates a headless service and the proxy refreshes DNS-discovered peers automatically, so scaling events do not require manual replica or peer updates.

For Grafana Logs Drilldown pattern discovery, keep the default extraArgs.patterns-enabled=true or set it explicitly during rollout if you need to control the surface area:

extraArgs:
backend: http://victorialogs:9428
label-style: underscores
patterns-enabled: "true"

Required Configuration​

FlagRequiredDescription
-backendYesVictoriaLogs URL
-listenNoListen address (default :3100)
-label-styleNounderscores (default) or passthrough

Backend Auth Forwarding​

If VictoriaLogs authentication is delegated from upstream clients, you can forward client Authorization to backend explicitly:

-forward-authorization=true

Equivalent manual mode:

-forward-headers=Authorization

Use this only in trusted topologies (for example Grafana/auth-proxy -> Loki-VL-proxy -> VictoriaLogs).


Operational Assets​

Treat these as one versioned operational package:

AssetCanonical sourcePurpose
Grafana operations dashboarddashboard/loki-vl-proxy.jsonThree-section layout: Section 1 — SLO/SLI + Health (8-stat top strip: circuit breaker, active requests, QPS, error %, P99 client latency, P95 backend latency, cache hit ratio, uptime; plus SLI time-series rows). Section 2 — Client → Proxy → VL + Resources (client visibility: request rate by route, errors by reason, query length, per-client inflight, latency by route; proxy internals: coalescing, internal ops, response tuple mode, tenant QPS; VL backend: upstream fanout, window count, backend latency, fetch/merge latency, adaptive parallelism; process resources: CPU, memory, goroutines, GC, network, disk I/O, PSI pressure). Section 3 — Deep Proxy Internals (cache tiers: T0/L1/L2/L3 hit/miss, sizes, stale hits, backend fallthrough; peer cache fleet: cluster members, hit/miss, write-through, hot read-ahead, error breakdown; query-range windowing: window cache, prefilter efficiency, retries, partial responses, prefilter duration, adaptive parallelism trace; patterns engine: in-memory count/bytes, mining rate, source line pipeline, snapshot hits/reuse, persistence; HTTP connection lifecycle: states, rotation reasons, transitions; tenant deep dive: per-tenant QPS/P99/errors)
Alert rulesalerting/loki-vl-proxy-prometheusrule.yamlPrometheusRule/vmalert-oriented alert set with standardized labels and annotations
SRE runbooksdocs/runbooks/alerts.mdIndex plus per-alert runbook files referenced directly from alert runbook_url

When using the Helm chart, the runtime templates consume synced copies in charts/loki-vl-proxy/{dashboards,alerting}. Keep canonical and chart copies aligned with:

./scripts/ci/sync_observability_assets.sh sync
./scripts/ci/sync_observability_assets.sh --check

--check is already enforced in CI to prevent drift.


Preventive Scaling And Deployment​

Use the dedicated guide for prevention-oriented operations hardening:

Critical defaults to reduce incident frequency:

  • run at least 2 replicas with PDB enabled
  • enable HPA with conservative downscale
  • tune cache TTLs differently for query paths vs metadata paths
  • monitor backend p95 and proxy p99 histograms, not averages
  • add synthetic in-cluster e2e query probes in addition to /ready

Multi-Tenancy​

Tenant Mapping Strategies​

The proxy maps X-Scope-OrgID headers to VictoriaLogs tenant IDs. Three strategies are available depending on deployment size and dynamism.

1. Inline JSON (-tenant-map)​

Best for small, static tenant maps. The entire map is provided directly as a CLI flag or env var value:

-tenant-map='{"team-a":{"account_id":"1","project_id":"0"},"team-b":{"account_id":"2","project_id":"0"}}'

This requires a proxy restart to update. (A map supplied through the TENANT_MAP environment variable is re-read on SIGHUP when no -tenant-map-file is set.)

2. File-based (-tenant-map-file)​

Best for Kubernetes environments where tenant maps are mounted as ConfigMaps. The proxy hot-reloads the file on SIGHUP and also polls for mtime changes on the configured interval:

-tenant-map-file=/etc/proxy/tenants.yaml
-tenant-map-reload-interval=30s

The default reload interval is 30s. To trigger an immediate reload without restarting the proxy:

kill -HUP <pid>

In Helm, configure a lifecycle hook to send SIGHUP on ConfigMap updates:

lifecycle:
postStart:
exec:
command: ["/bin/sh", "-c", "kill -HUP 1"]

Polling every 30s means changes are picked up automatically even without an explicit signal, which suits ConfigMap-mounted files that are updated by an external controller.

3. Label-based (-tenant-label)​

Scopes each request by a stream field instead of VictoriaLogs AccountID/ProjectID headers. Useful when the VictoriaLogs default tenant (0:0) holds data for several tenants distinguished by a stream label such as tenant:

-tenant-label=tenant

When set, the client still sends X-Scope-OrgID. For an org ID that is not in the tenant map, not a default-tenant alias (0, fake, default) and not *, the proxy adds a VictoriaLogs extra_stream_filters constraint {"tenant":"<orgID>"} to backend queries; request parameters cannot override it. The configured field must be a VictoriaLogs stream field (part of _stream_fields at ingestion). Explicit tenant-map entries take priority, and an unmapped X-Scope-OrgID: * is still rejected with 403 unless -tenant.allow-global=true.


-require-tenant-header Flag​

-require-tenant-header=true enforces that every request carries an X-Scope-OrgID header (returns HTTP 401 if missing) without enabling full auth. This is useful for catching misconfigured clients in multi-tenant setups without a full auth proxy.

-auth.enabled=true has the same effect on requests without the header (401). Neither flag authenticates the header value; put an authenticating proxy in front of the proxy when tenants must not be able to choose their own X-Scope-OrgID.


Health Check Endpoints​

The proxy exposes these operational endpoints:

EndpointPurposeKubernetes probe
/aliveLiveness — confirms the process is runninglivenessProbe
/readyReadiness — confirms the proxy is ready to serve traffic (backend reachable, warm-up complete)readinessProbe
/metricsPrometheus metrics scrape. Off by default since v1.56.0: requires -server.register-instrumentation=true, and is served on --metrics-listen when set (the Helm chart uses :9091), otherwise on the main listenerServiceMonitor / scrape config

Admin and debug routes (/admin/cache/flush, /debug/pprof/*, /debug/queries) are served on the loopback --admin-listen address (default 127.0.0.1:3101) unless -server.admin-auth-token is set, in which case they move to the main listener and require that token. /admin/cache/flush exists only when -server.register-instrumentation=true; POST /admin/cache/flush?peers=1 also purges every peer in the ring through the token-protected POST /_cache/purge peer endpoint.

If /ready stays non-ok immediately after a restart, check whether patterns or indexed label-values startup warm is configured — those persistence restores can intentionally hold readiness at 503 until warm-up completes.


Translation Modes​

Translation guidance moved to dedicated docs:

Operational recommendation:

  • use label-style=underscores (default) when upstream VL stores dotted OTel fields; passthrough when VL already stores underscore names
  • use metadata-field-mode=hybrid for mixed Loki + OTel field workflows
  • use metadata-field-mode=translated (default) for strict Loki-style field surfaces
  • use metadata-field-mode=native for OTel-native field-only surfaces

Capacity Planning​

Memory​

ComponentMemory per Unit
L1 cache~50MB per 10k entries
L2 disk cache (bbolt)~10MB mmap overhead
Per active query~1-5MB (depends on result size)
Singleflight coalescing bufferUp to 256MB per unique query
Base process~20MB

Formula: base(20MB) + cache(entries × 5KB) + concurrent_queries × 3MB

Default -cache-max is 10000 (binary default). The Helm chart ships 50000 to suit light-to-moderate production use. For 50k cache entries and 100 concurrent queries: ~570MB recommended limit.

CPU​

The proxy is CPU-light. Main costs:

  • JSON marshaling/unmarshaling (~70% of CPU)
  • LogQL→LogsQL translation (~10%)
  • Label translation (~5%)
  • HTTP overhead (~15%)

Guideline: 1 CPU core handles ~2000 req/s.

Disk Cache​

L2 disk cache with bbolt:

  • 1 million entries ≈ 2-5GB on disk (gzip compressed)
  • Write amplification: ~2x with bbolt
  • Use fast SSD (NVMe) for the cache volume
  • Set disk-cache-flush-size=500 and disk-cache-flush-interval=10s for batched writes

Performance Tuning​

Cache TTLs​

Default TTLs are conservative. Adjust for your query patterns:

-cache-ttl=120s # Increase for stable label sets
-cache-max=50000 # Increase for high-cardinality environments
EndpointDefault TTLNotes
labels, label_values5m-labels-cache-ttl; scaled up for longer request windows, capped at 1h
detected_fields, detected_field_values, detected_labels90sscaled up for longer request windows, capped at 1h
series30sTier0 compatibility cache
query_range, query5m final-response cacherequests ending within -recent-tail-refresh-window (default 2m) of now refetch once the entry is older than -recent-tail-refresh-max-staleness (default 2s)
index_stats, volume, volume_range10s

The per-endpoint TTLs above are built in; -labels-cache-ttl is the only per-endpoint override. Query-range split windows use -query-range-history-cache-ttl / -query-range-recent-cache-ttl.

Concurrency Limits​

-http-max-header-bytes=1048576 # 1MB default
-http-max-body-bytes=10485760 # 10MB default

The proxy uses singleflight to coalesce identical concurrent queries. N identical requests → 1 backend request.

Built-In Traffic Guards​

All traffic guard controls are tunable via CLI flags (or extraArgs in the Helm chart):

FlagDefaultDescription
-rate-limit-per-second50Per-client request rate (req/s)
-rate-limit-burst100Per-client burst allowance
-max-concurrent100Per-replica in-flight request cap (excess gets 503 with Retry-After: 5); also bounds concurrent backend operations
-cb-fail-threshold5Failures within window to open circuit breaker
-cb-open-duration10sHow long circuit breaker stays open
-cb-window-duration30sFailure counting window
-backend-max-concurrent-heavy-queries2Heavy VictoriaLogs calls running at once per replica (raw-row metric fetches, stats and hits over -backend-heavy-query-min-range); excess waits -backend-heavy-query-queue-wait (20s), then gets 429
-backend-max-concurrent-metadata-scans8Ceiling of the adaptive limit on long-range VictoriaLogs metadata scans per replica (/labels, /label/{name}/values, /series, detected_fields over -backend-heavy-query-min-range, and the day bucket scans of the label inventory); selects VictoriaLogs runs for others count against it, and it follows scan latency (-backend-metadata-scan-latency-tolerance, 1.5), backend failures and VictoriaLogs memory (-backend-metadata-scan-memory-headroom, 0.4, read from its /metrics) between -backend-min-concurrent-metadata-scans (1) and this ceiling; excess waits -backend-heavy-query-queue-wait, then gets 429; background warm-ups skip instead of waiting
-metadata-inventory-parallelism4Bucket listings in flight per label request from the time-bucketed label inventory (day, hour, 5-minute and minute buckets cached and merged; a refresh reads only its edges). 0 turns the inventory off

If defaults are too strict or too loose for your workload, tune at the proxy first, then complement with:

  • reduced Grafana auto-refresh and retry pressure
  • ingress or service-mesh shaping in front of the proxy
  • scale out replicas and raise cache effectiveness before pushing more uncached load

Monitoring​

See the dedicated Observability Guide for the full metrics catalog, JSON log schema, OTLP push configuration, and collector/agent integration examples.

Metrics​

The proxy exposes Prometheus metrics at /metrics:

Use the Observability Guide as the canonical catalog for:

  • every documented loki_vl_proxy_* metric family
  • cardinality level (Low, Medium, High (capped)) for each family
  • scrape versus OTLP field/label mapping
  • the new fanout and proxy-internal operation metrics/log fields
MetricTypePrimary dimensionsDescription
loki_vl_proxy_requests_totalcountersystem, direction, endpoint, route, statusTotal requests by downstream Loki route or upstream backend route
loki_vl_proxy_request_duration_secondshistogramsystem, direction, endpoint, routeEnd-to-end request latency
loki_vl_proxy_backend_duration_secondshistogramsystem, direction, endpoint, routeUpstream-only latency for VictoriaLogs and rules/alerts backends
loki_vl_proxy_cache_hits_by_endpoint / loki_vl_proxy_cache_misses_by_endpointcountersystem, direction, endpoint, routeCache efficiency by normalized route
loki_vl_proxy_tenant_requests_total / loki_vl_proxy_client_requests_totalcountertenant/client plus route dimensionsHot tenants and clients per route
loki_vl_proxy_process_*gauges/countersmetric family specificRuntime, CPU, memory, disk, network, and PSI health

Key Ratios to Monitor​

  • Route cache hit ratio: cache_hits_by_endpoint / (cache_hits_by_endpoint + cache_misses_by_endpoint) by endpoint,route — target >80% on stable metadata paths
  • Downstream error rate: requests_total{system="loki",direction="downstream",status=~"5.."} over total downstream requests — target <1%
  • Upstream latency: backend_duration_seconds by endpoint,route — use this to separate VictoriaLogs slowness from proxy-side work
  • End-to-end latency: request_duration_seconds{system="loki",direction="downstream"} by endpoint,route — compare with upstream latency and request logs

OTLP Push​

Push metrics to an OTLP collector:

-otlp-endpoint=http://otel-collector:4318/v1/metrics
-otlp-interval=30s
-otlp-compression=gzip

The OTLP exporter reuses the same core proxy metric names that /metrics exposes, so dashboards and alert logic can stay aligned across scrape and push modes.

For exact proxy-only overhead on translated paths, use structured request logs with proxy.overhead_ms, proxy.duration_ms, and upstream.duration_ms. The metrics intentionally keep route-aware end-to-end and upstream histograms, while logs carry the per-request decomposition.


Troubleshooting​

No Data in Grafana​

  1. Check proxy health: curl http://proxy:3100/ready
  2. Check VL backend: curl http://vl:9428/health
  3. Check proxy logs for translation errors and execution-limit rejections
  4. Verify label-style matches your VL ingestion format
  5. Check /loki/api/v1/labels for available labels

Label Names Don't Match​

SymptomCauseFix
Dots in Grafana labelslabel-style=passthrough with dotted VL dataSet label-style=underscores (the default)
Empty label_values for service_nameVL stores service.name, query asks service_nameSet label-style=underscores (the default)
Grafana Drilldown "failed to fetch"Volume/stats endpoint issueCheck proxy logs, ensure VL v1.49+

High Memory Usage​

  • Reduce -cache-max (default 10000)
  • Reduce -http-max-body-bytes
  • Add memory limits in Kubernetes
  • Check for singleflight amplification (many unique queries)

High Latency​

  • Keep -response-compression=gzip for broad Loki/Grafana compatibility; auto now behaves the same on the frontend for legacy configs
  • Set -response-compression-min-bytes around 1024 to avoid wasting CPU on small metadata/control responses
  • Increase cache TTLs
  • Check VL backend latency via metrics
  • Rely on built-in singleflight coalescing for identical concurrent reads

502 / 503 Errors On Large Queries​

Not every 502 or 503 means VictoriaLogs is down. Built-in execution limits reject oversized work instead of returning truncated data:

  • 502 with manual range metric row limit exceeded — raise -manual-range-metric-row-limit or narrow the query
  • 400 with maximum number of series (N) reached for a single query — Loki's own answer above max_query_series; narrow the query or raise -max-stats-query-series. The log line for the rejected request names that flag. Grafana Logs Drilldown requests instead receive a partial result with the warning ... returning partial results
  • 503 with too many concurrent queries — the -max-concurrent admission cap was reached
  • 429 with too many outstanding requests: heavy VictoriaLogs queries are limited to -backend-max-concurrent-heavy-queries=N — heavy long-range VictoriaLogs calls waited longer than -backend-heavy-query-queue-wait for a slot. Retry, narrow the time range, or raise -backend-max-concurrent-heavy-queries together with VictoriaLogs memory (see Heavy VictoriaLogs Query Admission)
  • 429 with too many outstanding requests: long-range VictoriaLogs metadata scans are limited to at most -backend-max-concurrent-metadata-scans=N per replica (adaptive limit L, ...) — a long-range /labels, /label/{name}/values, /series or detected_fields request waited longer than -backend-heavy-query-queue-wait for one of its long scans: VictoriaLogs was running as many long selects as the adaptive limit (for this replica or others), concurrent scans had slowed down, a scan failed, VictoriaLogs' memory in use left less than -backend-metadata-scan-memory-headroom free, or a listing no replica has measured yet was waiting for VictoriaLogs to run no other select (the message says which). The buckets the request did scan stay cached, so a retry continues where it stopped. The backend is saturated: give VictoriaLogs memory or cores, or narrow the range; raising the ceiling does not help while the memory gate or VictoriaLogs' own load is what closed. The same 429 ending in the scan was stopped to keep VictoriaLogs' memory inside the headroom means the scan had started and VictoriaLogs' memory in use then passed the brake mark (a third of the way into -backend-metadata-scan-memory-headroom); the proxy stopped it before VictoriaLogs ran out of memory.

line_format and binary-expression evaluation limits return 400. See Fixed Execution Limits.

Circuit Breaker Tripping​

The circuit breaker opens after -cb-fail-threshold backend transport failures (connection errors) within -cb-window-duration. HTTP error responses from VictoriaLogs and timeouts do not count, because they prove the backend is reachable. Check:

  • VL backend health and logs
  • Network connectivity between proxy and VL
  • VL resource usage (CPU/memory/disk)

Backup & Recovery​

The proxy is stateless. Only the optional disk cache needs backup:

  • L1 cache: In-memory, rebuilds on restart
  • L2 disk cache: bbolt file at -disk-cache-path. Can be deleted safely — will be repopulated.
  • Configuration: All config is CLI flags / env vars. Store in Helm values or ConfigMap.

Scaling​

Horizontal Scaling​

horizontalPodAutoscaling:
enabled: true
minReplicas: 2
maxReplicas: 10
metrics:
- type: Resource
resource:
name: cpu
target:
type: Utilization
averageUtilization: 70

Pod Disruption Budget​

podDisruptionBudget:
enabled: true
minAvailable: 1

Multi-Zone Deployment​

topologySpreadConstraints:
- maxSkew: 1
topologyKey: topology.kubernetes.io/zone
whenUnsatisfiable: ScheduleAnyway
labelSelector:
matchLabels:
app.kubernetes.io/name: loki-vl-proxy