Skip to main content

Configuration

All flags follow VictoriaMetrics naming conventions (-flagName=value).

Generated reference​

Three pages are generated from the flags the binary declares and the limits registry in internal/config, so they cannot drift from the code (go run ./cmd/configdoc regenerates them, and a test fails when they differ):

  • Configuration reference β€” every flag with its type, default, Helm value and description, grouped by category.
  • Limits registry β€” every bound on work: what it bounds, the error returned when it is hit, its metric and alert, whether a tenant can override it, Loki parity, and how to size it up or down. Includes worked examples for a small and a large deployment.
  • Errors and alerts index β€” from an error message or a firing alert back to the limit that produced it.

The same run writes conformance/registry/generated/proxy/limits.json, the machine-readable export the conformance registry reads, so the inventory of limits is generated from the flags rather than kept as a second hand-written list.

The sections below stay hand-written: they explain how the pieces fit together.

Server​

FlagEnvDefaultDescription
-listenLISTEN_ADDR:3100Listen address
-backendVL_BACKEND_URLhttp://localhost:9428VictoriaLogs backend URL
-proc-rootPROC_ROOT/procProc filesystem root used for CPU/memory/disk/network/PSI metrics (/proc for container scope, /host/proc for host scope)
-host-proc-rootHOST_PROC_ROOT/procSeparate proc filesystem root used for host-scope reads (stat, meminfo, pressure/{cpu,memory,io}). Set to /host/proc when the chart mounts the five individual host /proc files at that path; keeps self-scope reads on /proc while host-scope reads pull from the dedicated mount. Falls back to -proc-root when unset.
-backend-version-strictβ€”falsePromote the existing soft backend version check to a hard startup failure. When true, a /health failure, non-2xx response, or sub-minimum semver causes the proxy to exit during startup instead of logging a warning. Mirrors the upstream vmselect strict-version pattern.
-metadata-default-lookbackβ€”12hDefault lookback bound applied to /labels, /label/{name}/values, and /series when the client omits both start and end. Bounds the backend time window so an unbounded Grafana label-discovery query cannot fan out across all retention. 0 disables (prior unbounded behavior).
-ruler-backendRULER_BACKEND_URLβ€”Optional rules backend for legacy Loki YAML rules routes and Prometheus-style /prometheus/api/v1/rules passthrough
-alerts-backendALERTS_BACKEND_URLβ€”Optional alerts backend for Loki and Prometheus-style alerts endpoints (defaults to -ruler-backend)
-log-levelβ€”infoLog level: debug, info, warn, error
-tls-cert-fileβ€”β€”TLS certificate file for HTTPS
-tls-key-fileβ€”β€”TLS key file for HTTPS
-tls-client-ca-fileβ€”β€”CA file for verifying HTTPS client certificates
-tls-require-client-certβ€”falseRequire and verify HTTPS client certificates
-log-request-sample-rateβ€”0Static sampling for successful (2xx) access logs: 0 or 1 logs every request, N>1 logs 1 in every N requests; 4xx/5xx are always logged. The proxy always runs adaptive sampling (-log-rate-threshold, -log-stats-interval), which takes precedence; this static rate is only a fallback when the adaptive sampler is not initialized.
-log-bufferedβ€”trueWrite logs in the background to avoid slowing down requests under high load
-log-stats-intervalβ€”10sHow often to print a request statistics summary (total, errors, latency, cache rate)
-log-rate-thresholdβ€”10When traffic exceeds this rate (req/s), replace per-request logs with periodic summaries. Errors are always logged.
-cache-disabledβ€”falseDisable the in-memory L1 cache entirely (useful for benchmarking raw translation overhead)

Label Translation​

See Translation Modes Guide for mode-selection profiles and exact underscore vs dotted exposure behavior across labels, field APIs, and structured metadata.

FlagEnvDefaultDescription
-label-styleβ€”underscorespassthrough or underscores
-metadata-field-modeβ€”translatednative, translated, or hybrid for detected_fields and structured metadata exposure
-translate-otel-attributesTRANSLATE_OTEL_ATTRIBUTEStrueTranslate known OTel semantic-convention labels from underscore to dotted form in upstream queries; set false to preserve client label names (for example when VictoriaLogs stores underscore names)
-backend-default-msg-valueβ€”β€”VictoriaLogs -defaultMsgValue when it is customized. The proxy returns _msg as the log line; rows whose _msg is empty, starts with VictoriaLogs' default missing _msg field text, or equals this value stored no line (for example a Loki push of a JSON line), so the proxy rebuilds the line as a JSON object of the row's non-stream fields. See Logging for compatibility for the limits
-emit-structured-metadataβ€”trueEnable Loki categorize-labels response encoding: requests with X-Loki-Response-Encoding-Flags: categorize-labels emit 3-tuples [timestamp, line, metadata], while default/no-flag requests stay canonical 2-tuples
-logql-dotted-namesβ€”autoauto, reject or accept: dotted names in LogQL (k8s.pod.name in a matcher, label filter, by/without list, keep/drop, label_format, parser parameters, unwrap). reject answers Loki's exact 400 parse error and names dotted JSON keys in detected_fields by Loki's sanitized label (http_method, key in jsonPath); accept translates them to the dotted VictoriaLogs field and keeps dotted JSON keys in detected_fields; auto rejects in the Loki-compatible profile and accepts otherwise. See Compatibility option matrix
-label-browse-extensionsβ€”autoauto, on or off: limit, offset and search/q on /labels and /label/{name}/values, a proxy extension Loki ignores. auto is off in the Loki-compatible profile unless -label-values-indexed-cache=true, on otherwise
-error-response-message-fieldβ€”trueAdd message (the field Grafana's Loki datasource displays) with the error text to JSON error bodies; false restores the previous {status, errorType, error} body
-detected-level-body-scanβ€”trueDerive detected_level like Loki for rows without a stored level field: the log line is read as JSON, then logfmt, then scanned for level keywords, with unknown as the fallback. false uses only stored fields (a stored detected_level, Loki's log_level_fields such as level, severity, lvl, severity_text, then severity_number) and answers unknown otherwise, which skips the per-line scan on large plain-text tenants but differs from Loki for lines that carry their level only in the text. Applies to log query responses, tail, /detected_fields and patterns collected from log queries; the /patterns endpoint's own sample always reads the line, and reads a row whose _msg VictoriaLogs replaced with -backend-default-msg-value as an ordinary line
-patterns-enabledβ€”trueEnable GET /loki/api/v1/patterns (Grafana Logs Drilldown patterns view). When false, the endpoint returns an empty successful response ({"status":"success","data":[]})
-patterns-autodetect-from-queriesβ€”falseWarm /loki/api/v1/patterns cache from successful query and query_range responses (global autodetect mode, opt-in)
-patterns-customβ€”β€”Static custom patterns prepended to every /patterns response. Accepts JSON string array or newline-separated text payload
-patterns-custom-fileβ€”β€”File path for static custom patterns prepended to every /patterns response. Supports JSON string array or newline-separated text with optional # comments
-patterns-persist-pathβ€”β€”Disk path for persisted patterns snapshot JSON file (empty disables persistence)
-patterns-persist-intervalβ€”30sPeriodic flush interval for in-memory patterns snapshot
-patterns-startup-stale-thresholdβ€”60sFreshness threshold used by startup warm logic/peer snapshot cache
-patterns-startup-peer-warm-timeoutβ€”5sStartup timeout for peer warm merge of pattern snapshots
-field-mappingFIELD_MAPPINGβ€”JSON custom field mappings
-stream-fieldsβ€”β€”Comma-separated _stream_fields labels used for stream selector optimization and label-surface hints
-extra-label-fieldsEXTRA_LABEL_FIELDSβ€”Comma-separated additional VL fields to expose on label-facing APIs and alias resolution paths (for example host.id,k8s.cluster.name)

Configure -label-style and -metadata-field-mode with flags (or Helm extraArgs). The LABEL_STYLE and METADATA_FIELD_MODE environment variables do not override the current defaults.

Label Style Modes​

ModeWhen to UseResponseQuery
passthroughVL stores underscore labels (Vector/FluentBit normalize)No translationNo translation
underscores (default)VL stores OTel dotted labels (OTLP direct)service.name β†’ service_name{service_name="x"} β†’ VL "service.name":"x"

Custom Field Mappings​

./loki-vl-proxy -label-style=underscores \
-field-mapping='[{"vl_field":"my_trace_id","loki_label":"traceID"}]'

Custom Drilldown Patterns​

Examples:

# Inline JSON list
./loki-vl-proxy \
-patterns-custom='["time=\"<_>\" level=info msg=\"finished unary call\"","grpc.code=<_> grpc.method=<_>"]'

# File-backed list (JSON array or newline-separated text)
./loki-vl-proxy \
-patterns-custom-file=/etc/loki-vl-proxy/custom-patterns.json

Extra Label Fields (-extra-label-fields)​

Use this when you need explicit label exposure and alias resolution for fields that are not always discoverable from stream_field_names.

What it affects:

  • GET /loki/api/v1/labels: appends configured fields to the label set.
  • GET /loki/api/v1/label/{name}/values: resolves underscore↔dot aliases for custom fields.
  • GET /loki/api/v1/index/volume and GET /loki/api/v1/index/volume_range: resolves targetLabels aliases (for example host_id -> host.id) before backend grouping.

How values are handled:

  • Accepts a comma-separated list.
  • Entries are normalized to canonical VL field names by translator rules.
  • Duplicates are removed.
  • It extends exposure and aliasing only; it does not create or index fields in VictoriaLogs.

Caching and limits:

  • It reuses existing label-path caches (/labels and /label/{name}/values), so repeated lookups do not repeatedly rescan backend metadata.
  • There is no dedicated hard cap on how many entries you can set in -extra-label-fields; behavior is intentionally open to match VictoriaLogs field discovery.
  • Existing proxy safeguards still apply (endpoint TTLs, backend timeouts, and standard response-size protections).
  • Practical guidance: keep this list focused on fields you want Loki users to filter on frequently; very large lists can add UI noise in Grafana.

Examples:

# Conservative Loki-facing UX + explicit custom fields
./loki-vl-proxy \
-label-style=underscores \
-metadata-field-mode=translated \
-emit-structured-metadata=true \
-extra-label-fields='host.id,k8s.cluster.name,custom.pipeline.processing'
# Equivalent via env
export EXTRA_LABEL_FIELDS='host.id,k8s.cluster.name,custom.pipeline.processing'
./loki-vl-proxy -label-style=underscores -metadata-field-mode=translated

When to also use -field-mapping:

  • Use -extra-label-fields to extend discovery and alias resolution.
  • Use -field-mapping when you need a non-default alias name (for example custom.pipeline.processing <-> pipeline_proc).

Indexed Label Values Browse Cache (Optional)​

This mode is designed for very high-cardinality labels (for example k8s_pod_name) where returning every value on first browse is expensive.

FlagEnvDefaultDescription
-label-values-indexed-cacheβ€”falseEnable indexed hotset browsing for GET /loki/api/v1/label/{name}/values
-label-values-hot-limitβ€”200Default values returned for empty-query browse when limit is not specified
-label-values-index-max-entriesβ€”200000Maximum indexed values retained per tenant+label in memory
-label-values-index-persist-pathβ€”β€”Disk path for persisted label-values index snapshot (JSON)
-label-values-index-persist-intervalβ€”30sPeriodic snapshot flush interval
-label-values-index-startup-stale-thresholdβ€”60sDisk snapshot freshness threshold before peer warm fallback
-label-values-index-startup-peer-warm-timeoutβ€”5sStartup timeout for peer warm fallback

Behavior when enabled:

  • Empty-query browse (query omitted or query=*) serves a hot subset first.
  • Supports optional pagination-style parameters on the same endpoint: limit and offset.
  • Supports optional in-proxy value filtering with search (alias q), without forcing a full backend refetch when index is warm.
  • Query-scoped requests (query={...}) keep standard behavior and still update index state.
  • The index is not time-scoped. Once a label has indexed values, empty-query browse returns them for any start/end, including values only seen in other windows. The first browse of a label that is not indexed yet, and every query-scoped request, reads the full requested range from VictoriaLogs.
  • Startup warm order: restore disk snapshot first, then warm from peer cache when disk snapshot is stale/missing.
  • Rolling update safety: graceful shutdown writes a final snapshot before exit.
  • Readiness behavior: /ready stays 503 until label-values startup warm is finished.

Example:

./loki-vl-proxy \
-label-values-indexed-cache=true \
-label-values-hot-limit=200 \
-label-values-index-max-entries=200000 \
-label-values-index-persist-path=/cache/label-values-index.json \
-label-values-index-persist-interval=30s \
-label-values-index-startup-stale-threshold=60s \
-label-values-index-startup-peer-warm-timeout=5s

Sizing guidance:

- RAM estimate per label key: `label-values-index-max-entries * ~96 bytes`.
- Disk snapshot estimate per label key: roughly `~60%` of RAM estimate.
- Total footprint scales with `(tenant_count * indexed_label_count)`.

Enable this to keep detected /loki/api/v1/patterns results warm across rolling restarts and fleet members.

Behavior:

  • Pattern cache entries are retained long-term and updated in-place by cache key.
  • Every detected pattern response is appended to the in-memory snapshot map (no periodic TTL-based pruning in snapshot state).
  • Optional global autodetect (-patterns-autodetect-from-queries=true) passively mines successful query and query_range responses and pre-warms matching /patterns cache keys.
  • Snapshot is persisted to disk periodically and on graceful shutdown.
  • Startup restore order: disk snapshot first, then peer merge (newest entry wins per cache key).
  • Readiness stays 503 during startup warm when patterns persistence is configured.

Operational guidance:

  • For restart-safe persistence, run StatefulSet + PVC and set -patterns-persist-path to a writable mounted path.
  • If -patterns-persist-path is set but not writable, proxy startup fails fast with a clear error.
  • If -patterns-persist-path is not set, patterns still work, but persistence is disabled (cold start after restart).

Sizing guidance:

  • Endpoint clamp: max returned patterns per request is 1000.
  • Approximate persisted bytes:
    • snapshot_size ~= sum(pattern_response_payload_bytes_per_cached_query_key)
    • Rule of thumb per pattern entry: ~(120 bytes base + pattern length + ~24 bytes per sample bucket)
  • Real footprint depends on:
    • number of unique (tenant, rawQuery) keys,
    • sample bucket count per pattern (step, query range),
    • pattern text length distribution.

References:

Metadata Field Modes​

ModeWhen to UseField APIs
nativeYou want only raw VictoriaLogs field namesservice.name, k8s.pod.name
translatedDefault. You want strict Loki-style field names onlyservice_name, k8s_pod_name
hybridYou need Loki compatibility plus OTel-native correlationBoth native dotted names and translated aliases

The default -metadata-field-mode=translated exposes only Loki-style aliases. -metadata-field-mode=hybrid keeps the label surface Loki-compatible while making field-oriented APIs like detected_fields and detected_field/{name}/values expose both service.name and service_name when they differ.

Compatibility Profiles​

The three flags below define the compatibility profile:

  • -label-style
  • -metadata-field-mode
  • -emit-structured-metadata
ProfileSettingsBest For
Loki/Grafana conservative (default)label-style=underscores, metadata-field-mode=translated, emit-structured-metadata=trueStrict Loki-style field naming plus Explore/Drilldown event metadata
Drilldown/OTel mixed modelabel-style=underscores, metadata-field-mode=hybrid, emit-structured-metadata=trueGrafana + OTel correlation where both dotted and translated field names are useful
Native VL field surfacelabel-style=passthrough, metadata-field-mode=native, emit-structured-metadata=trueConsumers that prefer raw VictoriaLogs field names and structured metadata

The Loki/Grafana conservative profile is the Loki-compatible profile: the proxy holds requests to Loki's contract as well as responses. A dotted name in LogQL (| k8s.namespace.name="x", by (k8s.pod.name), {service.name="x"}) is answered with Loki's exact 400 parse error - line, column and expected-token list included - on query, query_range, tail and every selector parameter (labels, label values, series, index, patterns, detected fields), before any VictoriaLogs call; the label endpoints ignore limit, offset and search the way Loki does. Both follow from -logql-dotted-names=auto and -label-browse-extensions=auto, which match the profile; set either explicitly to keep an extension with Loki's names, or Loki's grammar with the hybrid or native names. The hybrid and native modes expose dotted VictoriaLogs names, so with auto they keep accepting dotted names in queries and keep the browse parameters: those are documented extensions, not Loki behaviour.

There is no single profile flag: the flags already default to the Loki-compatible profile, each also changes one surface on its own (for example -emit-structured-metadata=false for clients that reject metadata objects), and a profile flag would need precedence rules against them. Set the flags explicitly to pin a profile; the option matrix lists what each combination does.

For tuple behavior and endpoint-level details, see API Reference. For support scope by product/version track, see Compatibility Matrix.

If you must interoperate with legacy clients that reject metadata objects in categorize-labels mode, explicitly set -emit-structured-metadata=false.

Compatibility option matrix​

Every option below is its own flag, and any combination is supported. The Loki-compatible profile is the default; each extension can be turned on alone without leaving it, and each Loki behaviour can be turned off alone.

OptionValues (default first)Loki behaviourWhat the other values do
-label-styleunderscores, passthroughunderscores: stream labels and label names sanitized (service.name -> service_name)passthrough: VictoriaLogs field names as stored, dots included
-metadata-field-modetranslated, hybrid, nativetranslated: detected_fields and structured metadata under Loki's names onlyhybrid: both the dotted and the Loki name; native: the dotted name only
-emit-structured-metadatatrue, falsetrue: categorize-labels 3-tuples carry structured metadata and parsed fieldsfalse: the metadata object stays empty, for clients that reject it
-logql-dotted-namesauto, reject, acceptreject: a dotted name is Loki's 400 parse error; dotted JSON keys appear in detected_fields as Loki's sanitized label with the key in jsonPathaccept: dotted names are translated to the dotted VictoriaLogs field; detected_fields keeps dotted JSON keys. auto = reject exactly when -label-style=underscores and -metadata-field-mode=translated
-label-browse-extensionsauto, on, offoff: limit, offset, search/q are ignored on the label endpointson: the browse window applies (hot set first with -label-values-indexed-cache). auto = off in the Loki-compatible profile without -label-values-indexed-cache, on otherwise
-error-response-message-fieldtrue, false(Loki answers text/plain) true: Grafana shows the error textfalse: the previous {status, errorType, error} body
-label-values-max-response-bytes67108864label values fail with Loki's ResourceExhausted 500 past the querier's message sizeraise it (per tenant too) to answer larger label value lists; 268435456 is the previous implicit bound

What each combination of the three profile flags does with the auto settings (the unit test TestCompatOptionMatrix checks every combination of all six options: dotted-name handling, browse parameters, structured-metadata keys and detected_fields names):

-label-style-metadata-field-modeLabelsStructured metadata / detected fieldsDotted names in LogQL (auto)Browse params (auto)e2e variant
underscorestranslatedservice_namehttp_targetrejected, Loki's errorignored (on with the indexed cache)loki-vl-proxy, -underscore, -patterns-autodetect, -translated-metadata, -vmauth, -tail*
underscoreshybridservice_namehttp.target and http_targetacceptedhonouredloki-vl-proxy-otel-hybrid
underscoresnativeservice_namehttp.targetacceptedhonouredloki-vl-proxy-native-metadata
passthroughtranslatedservice.namehttp.target (nothing to translate)acceptedhonoured(unit test only)
passthroughhybridservice.namehttp.targetacceptedhonoured(unit test only)
passthroughnativeservice.namehttp.targetacceptedhonoured(unit test only)

-emit-structured-metadata=false empties the metadata object in any row (loki-vl-proxy-no-metadata runs it with the Loki profile). Combinations that mix the rows are valid: -metadata-field-mode=translated -logql-dotted-names=accept keeps Loki's names in every response while still running queries a user wrote with dotted names; -metadata-field-mode=hybrid -logql-dotted-names=reject shows both spellings but holds queries to Loki's grammar (the dotted spelling then cannot be queried, so it is only useful for display). With -label-style=passthrough the labels themselves carry dots, so rejecting dotted names leaves them unqueryable; keep accept there.

Maximum Loki Compatibility β€” Annotated Configuration​

This is the recommended starting point when your priority is full compatibility with Grafana's Loki datasource and Logs Drilldown. Copy, adjust the backend URL, and run.

./loki-vl-proxy \
# ── Backend ──────────────────────────────────────────────────────────────
-backend=http://victorialogs:9428 \ # VictoriaLogs HTTP API address
\
# ── Label translation: maximum Loki compatibility ─────────────────────────
# OTel dotted labels (service.name, k8s.pod.name) are translated to Loki-style
# underscores (service_name, k8s_pod_name) in every response.
# Queries written with underscores are rewritten to dotted before hitting VL.
-label-style=underscores \
\
# ── Metadata field exposure ────────────────────────────────────────────────
# "translated" exposes only underscore-style labels in detected_fields and
# detected_labels. This gives Grafana a clean, Loki-identical field list
# with no dotted duplicates in the fields panel.
-metadata-field-mode=translated \
\
# ── Structured metadata (3-tuple responses) ────────────────────────────────
# Required for Grafana 10+ Loki datasource "categorize-labels" mode.
# Enables per-line label context in log details panel without extra queries.
# Default is true; listed here explicitly so the intent is visible.
-emit-structured-metadata=true \
\
# ── Patterns (Logs Drilldown) ──────────────────────────────────────────────
# Enables GET /loki/api/v1/patterns used by the Patterns tab in Logs Drilldown.
# Default is true. Set false only if you do not use the Drilldown plugin.
-patterns-enabled=true

Environment variables (for Docker Compose or Kubernetes): only the backend address has an environment form here.

VL_BACKEND_URL=http://victorialogs:9428

The other settings in this profile are the binary defaults. Set -label-style, -metadata-field-mode, -emit-structured-metadata and -patterns-enabled as flags or Helm extraArgs; LABEL_STYLE and METADATA_FIELD_MODE do not override the current defaults.

Grafana datasource β€” global single-tenant (no tenant mapping needed):

datasources:
- name: Logs
type: loki
url: http://loki-vl-proxy:3100
jsonData:
httpHeaderName1: X-Scope-OrgID
secureJsonData:
# "0", "fake", and "default" all resolve to VL's built-in default tenant (0:0).
# Pick whichever your Loki configuration already uses; the proxy treats them identically.
httpHeaderValue1: "0"

What each flag does and why it matters for Loki compatibility:

FlagValueEffect
-label-styleunderscoresTranslates service.name β†’ service_name so Grafana label filters work
-metadata-field-modetranslatedNo dotted duplicates in the fields panel; matches real Loki output
-emit-structured-metadatatrueEnables 3-tuple [ts, line, metadata] required by Grafana 10+ log details
-patterns-enabledtrueUnlocks the Patterns tab in Logs Drilldown

Switch -metadata-field-mode to hybrid if you also need OTel correlation (traces/metrics linked by service.name); the proxy then exposes both service.name and service_name in field APIs.

Cache (L1 In-Memory)​

FlagEnvDefaultDescription
-cache-ttlβ€”60sDefault cache TTL
-cache-maxβ€”10000Maximum cache entries
-cache-max-bytesβ€”268435456Maximum in-memory L1 cache size in bytes (256 MiB by default)
-labels-cache-ttlβ€”0 (uses 5m)Cache TTL for /labels and /label/{name}/values responses. 0 uses the built-in 5-minute default. The keep-warm loop runs at 75% of this TTL (the built-in 5 minutes when unset). A cache miss queries VictoriaLogs over the full requested start–end range, so the first response is complete; there is no reduced first scan.
-max-metadata-cache-freshnessβ€”24hLoki's max_metadata_cache_freshness, applied to every tenant (Loki's option is per tenant). Loki caches /labels, /label/{name}/values and /series results per 24h split and only for splits that end before now minus this window, so it reads the last 24h plus the split that straddles the boundary live. A request of the proxy that ends within this window of now refetches a cached answer once it is older than -recent-tail-refresh-max-staleness (default 2s), so streams written in the window appear at once, as in Loki; concurrent identical refetches share one pass. The refetch is answered from the time-bucketed inventory: sealed buckets come from its cache, while the newest minute, the unaligned edges and buckets due for revalidation are read from VictoriaLogs (a window under 10 minutes, a limit, -metadata-inventory-parallelism below 1 and the label-values field-name pass read it directly). Requests ending before the window keep the cached answer, its TTL and its background refresh. Independent of -recent-tail-refresh-enabled. 0 keeps the previous caching for every range. Rows that land more than a minute late in an already sealed, non-empty bucket appear after that bucket's revalidation; for a query of stream and field filters only, empty buckets in the window are confirmed on every request by one row count per contiguous run (cheap for * and stream filters, but with a field filter it reads the filtered column of every row in the run's range, so it is not free there), so rows written into them with old timestamps appear at once; a bucket stored empty with rows that lack the listed field is not part of a run, so rows that gain the field appear up to the negative TTL late.
-labels-cache-warmβ€”trueWarm the labels cache for the 1h, 6h, 24h and 7d time-picker presets at startup and keep them warm in the background (every 75% of -labels-cache-ttl). Each refresh is a label-name scan of up to 7 days in VictoriaLogs, per replica. Set false on replicas that serve no interactive label pickers (batch/API-only replicas, test variants): label requests are still cached and answered, only the proactive scans stop.
-compat-cache-enabledβ€”trueEnable the Tier0 compatibility-edge response cache for safe GET read endpoints
-compat-cache-max-percentβ€”10Percent of -cache-max-bytes reserved for Tier0 (0 disables, max 50)

Tier0 Compatibility-Edge Cache​

Tier0 is a separate in-memory cache instance that reuses the same cache implementation as the deeper L1/L2/L3 stack, but stores only final Loki-shaped response bodies.

  • It runs only after tenant validation and route classification.
  • It only applies to safe GET read endpoints such as query, query_range, series, labels, volume, patterns, and Drilldown metadata endpoints.
  • It does not apply to /tail, websocket upgrades, writes, deletes, admin/debug paths, or non-JSON responses.
  • Its budget is derived from -cache-max-bytes, so -compat-cache-max-percent=10 means Tier0 gets 10% of the primary L1 memory budget.
  • Tenant-map and field-mapping reloads invalidate Tier0 immediately.

Per-Endpoint TTLs​

EndpointTTL
labels, label_values5m (-labels-cache-ttl) for request windows up to 1h; longer windows scale the TTL up (Γ—3 up to 6h, Γ—10 up to 24h, Γ—20 up to 7d), capped at 1h. Background refreshes re-query the same full range and keep the scaled TTL. Requests ending within -max-metadata-cache-freshness (default 24h) of now refetch an entry older than -recent-tail-refresh-max-staleness instead of serving it
detected_fields, detected_field_values, detected_labels90s, scaled for longer request windows the same way, capped at 1h
series30s
patterns100y (effectively persistent; update-on-write)
query_range, query5m final-response cache; requests ending within -recent-tail-refresh-window (default 2m) of now are refetched once the cached entry is older than -recent-tail-refresh-max-staleness (default 2s)
index_stats, volume, volume_range10s

Cache (L2 On-Disk)​

FlagEnvDefaultDescription
-disk-cache-pathβ€”β€”Path to bbolt DB file (empty = disabled)
-disk-cache-compressβ€”trueGzip compression for disk cache
-disk-cache-flush-sizeβ€”100Flush write buffer after N entries
-disk-cache-flush-intervalβ€”5sWrite buffer flush interval
-disk-cache-min-ttlβ€”30sMinimum TTL required before an entry is eligible for L2 disk-cache writes. Empty /labels and /label/{name}/values answers are cached for max(30s, this, -peer-write-through-min-ttl), so they still overwrite an older non-empty copy on disk and peers when either minimum is raised (new labels can then take that long to appear after an empty answer)
-disk-cache-max-bytesβ€”0Maximum on-disk L2 cache size in bytes (0 = unlimited)

Cold Storage Backend​

Route queries for old data to a separate cold backend (e.g. Victoria Lakehouse or a second VictoriaLogs instance). Queries spanning both hot and cold data are split and merged transparently.

FlagEnvDefaultDescription
-cold-enabledβ€”falseEnable cold storage backend routing
-cold-backendβ€”β€”Cold storage backend URL (e.g. http://lakehouse:9428)
-cold-boundaryβ€”168hData older than this age is routed to the cold backend (default 7 days)
-cold-overlapβ€”1hOverlap window around the cold boundary β€” queries spanning the boundary include this overlap on both sides to avoid gaps
-cold-manifest-refreshβ€”5mHow often to refresh the cold backend capability manifest
-cold-timeoutβ€”30sTimeout for cold backend requests
./loki-vl-proxy \
-backend=http://victorialogs:9428 \
-cold-enabled \
-cold-backend=http://lakehouse:9428 \
-cold-boundary=336h # 14 days

Query Length Limits​

The proxy can enforce a maximum query time range, matching Loki's max_query_length behavior.

FlagEnvDefaultDescription
-default-max-query-lengthβ€”0Default maximum query time range enforced for all tenants unless overridden by tenant limits. 0 disables the limit (unlimited, matching Loki's default). Uses Go duration syntax (h, m, s; there is no d unit), for example 720h.

Precedence​

When multiple limits are configured, the most specific takes effect:

  1. Per-tenant β€” max_query_length for the tenant in -tenant-limits (highest priority)
  2. Tenant defaults β€” max_query_length in -tenant-default-limits
  3. Global flag β€” -default-max-query-length
  4. No limit β€” when all of the above are 0 or unset

Behavior​

  • Enforced, as in Loki's limits middleware, on query_range, series, labels, label values, detected_labels and index/stats requests
  • Independently of this flag, query_range requests where (end - start) / step exceeds 11,000 (integer division, as in Loki) return HTTP 400 with exceeded maximum resolution of 11,000 points per time series. Try increasing the value of the step parameter before any backend call, for log and metric queries alike
  • Applied after LogQL offset extraction (the enforced range reflects any offset modifier in the query)
  • Tenant-limit values accept Loki duration units, including d, w and y (for example "90d"); the global flag does not
  • Rejected queries return HTTP 400 with Loki's message: {"status":"error","errorType":"bad_data","error":"the query time range exceeds the limit (query length: 3h0m0s, limit: 1h)"}
  • A multi-tenant request is held to the smallest non-zero limit of its tenants
  • Matches Loki's max_query_length semantics for client compatibility

Examples​

# Limit all queries to 30 days by default
loki-vl-proxy -default-max-query-length=720h

# Allow a specific tenant to query further back (tenant limits accept Loki units such as d)
loki-vl-proxy -default-max-query-length=720h \
-tenant-limits='{"team-a":{"max_query_length":"90d"}}'

# No global limit (default β€” operator manages limits via per-tenant config)
loki-vl-proxy -default-max-query-length=0

Query Range Window Cache​

These flags control Loki-compatible query_range split/merge execution with per-window cache reuse.

FlagEnvDefaultDescription
-query-range-windowingβ€”trueEnable window split/merge path for log query_range
-query-range-split-intervalβ€”1hPer-window time span used for split/merge
-query-range-max-parallelβ€”2Static maximum window fetch parallelism (used when adaptive mode is disabled)
-query-range-adaptive-parallelβ€”trueEnable adaptive window fetch parallelism
-query-range-adaptive-min-parallelβ€”2Lower bound for adaptive parallelism
-query-range-adaptive-max-parallelβ€”8Upper bound for adaptive parallelism
-query-range-latency-targetβ€”1.5sTarget backend window fetch latency for safe adaptive increase
-query-range-latency-backoffβ€”3sBackoff threshold; adaptive mode reduces parallelism when EWMA latency exceeds this
-query-range-adaptive-cooldownβ€”30sMinimum interval between adaptive parallelism changes
-query-range-error-backoff-thresholdβ€”0.02Backoff threshold for EWMA backend error ratio (0-1)
-query-range-freshnessβ€”10mNear-now freshness boundary
-query-range-recent-cache-ttlβ€”0sCache TTL for windows newer than now-freshness (0s disables near-now cache)
-query-range-history-cache-ttlβ€”24hCache TTL for historical windows older than now-freshness
-query-range-prefilter-index-statsβ€”trueUse /select/logsql/hits preflight to skip empty split windows before expensive log fanout
-query-range-prefilter-min-windowsβ€”8Minimum split-window count required before prefilter is enabled
-query-range-stream-aware-batchingβ€”trueReduce parallelism for windows estimated as expensive from prefilter hits
-query-range-expensive-hit-thresholdβ€”2000Prefilter hit threshold above which a window is treated as expensive
-query-range-expensive-max-parallelβ€”1Maximum parallel fetches for expensive windows
-query-range-align-windowsβ€”trueAlign split windows to fixed interval boundaries for overlap-cache reuse
-query-range-window-timeoutβ€”20sPer-window backend timeout budget (0 disables)
-query-range-partial-responsesβ€”falseAllow partial query_range responses on retryable backend failures
-query-range-background-warmβ€”trueContinue warming failed windows in background after a partial response
-query-range-background-warm-max-windowsβ€”24Cap background warm fanout after a partial response
-recent-tail-refresh-enabledβ€”trueEnable near-now stale-cache bypass for query_range, index/volume, and index/volume_range
-recent-tail-refresh-windowβ€”2mTreat requests ending within this window from now as near-now
-recent-tail-refresh-max-stalenessβ€”2sMaximum acceptable cache age for near-now requests before forcing a fresh backend fetch

Drilldown and Metric Stats Controls​

FlagEnvDefaultDescription
-drilldown-scan-timeoutβ€”5sPer-request timeout for the detected_fields / detected_field/{name}/values log-scan path. 0 disables the cap
-exact-parser-series-identityβ€”falseName the series of a metric over | json or | logfmt with the labels those parsers extracted, as Loki does. Loki's series identity is the stream labels plus every extracted label, so count_over_time({app="x"} | logfmt [1m]) returns one series per distinct set of parsed values (450 on a 30-stream minute of the e2e generator) where the proxy returns one per stream (30). Turning this on evaluates every such query from rows (bounded by the manual row budget) instead of pushing it down to VictoriaLogs stats, so it costs far more on wide ranges; leave it off unless a dashboard depends on those labels. Captures of | regexp and | pattern are always part of the identity, because the query names them and VictoriaLogs can group by them
-max-stats-query-seriesβ€”0 (uses 500)Maximum series a metric query may return (count_over_time, rate, bytes_rate, ...), the proxy's equivalent of Loki's max_query_series; 0 uses the built-in default of 500, which is also Loki's default. Above it the query is rejected with HTTP 400 and Loki's message maximum number of series (N) reached for a single query, on range and instant queries alike; the log line for the rejected request names this flag. Grafana Logs Drilldown requests receive a partial result (the busiest series) with Loki's ... returning partial results warning instead, as Loki does for that client (see Fixed Execution Limits)
-stats-query-range-concurrencyβ€”0 (uses 4)Maximum concurrent stats_query_range calls to VictoriaLogs. 0 uses the built-in default of 4
-stats-query-range-inter-query-delay-msβ€”200Minimum pause in ms between consecutive individual stats_query_range calls; the concurrency slot is held for this long after each call. 0 disables
-drilldown-burst-window-msβ€”50Window in ms for coalescing concurrent Drilldown Fields per-field count queries into one fused VictoriaLogs call. 0 disables the coalescer
-drilldown-burst-max-fieldsβ€”30Maximum fields per coalesced burst call; further fields form another call
-drilldown-field-batch-window-msβ€”100Deprecated, no effect: Logs Drilldown field breakdowns are answered exactly, one stats_query_range call each. Accepted so existing command lines keep working
-drilldown-field-batch-max-fieldsβ€”6Deprecated, no effect (see -drilldown-field-batch-window-ms)
-drilldown-max-stats-bucketsβ€”120Deprecated, no effect: Logs Drilldown breakdowns are answered on the requested step, as Loki answers them. Accepted so existing command lines keep working

Loki-Aligned Defaults​

  • split-interval=1h mirrors the common Loki split-by-interval setup.
  • freshness=10m keeps the latest time range uncached by default, matching typical Loki freshness posture.
  • bounded max-parallel=2 protects VictoriaLogs from fanout bursts while still reducing latency for long ranges.
  • history-cache-ttl=24h keeps older windows warm for repeated drilldowns and expanded-range queries.
  • adaptive mode (min=2, max=8) raises parallelism only when EWMA latency/error remain healthy and backs off quickly when backend pressure rises.

Cache Sizing Guidance​

Use this approximation for disk planning:

required_bytes ~= unique_window_keys_per_day * avg_compressed_window_bytes * ttl_days

From in-repo sizing tests (internal/cache/sizing_test.go), a practical compressed average is around 56 KiB per entry, which yields:

  • 1 GiB -> about 18k cached windows
  • 10 GiB -> about 189k cached windows
  • 50 GiB -> about 948k cached windows

Set -disk-cache-max-bytes when you want deterministic upper bounds instead of relying only on PVC limits.

Practical Tuning Profiles​

ProfileSuggested SettingsWhen to Use
Conservativeadaptive-max=4, latency-target=2s, latency-backoff=4s, history-ttl=24hShared VictoriaLogs clusters with strict backend protection goals
Balanced (default)adaptive-max=8, latency-target=1.5s, latency-backoff=3s, history-ttl=24hMost Grafana + Drilldown workloads
Aggressiveadaptive-max=12, latency-target=1s, latency-backoff=2s, history-ttl=72h + larger disk cacheLarge repeated long-range analytics where backend headroom is proven

Tuning Visibility Metrics​

Use these metrics to tune safely:

  • loki_vl_proxy_window_adaptive_parallel_current (gauge)
  • loki_vl_proxy_window_adaptive_latency_ewma_seconds (gauge)
  • loki_vl_proxy_window_adaptive_error_ewma (gauge)
  • loki_vl_proxy_window_cache_hit_total, loki_vl_proxy_window_cache_miss_total (counters)
  • loki_vl_proxy_window_fetch_seconds and loki_vl_proxy_window_merge_seconds (histograms)
  • loki_vl_proxy_window_prefilter_attempt_total, loki_vl_proxy_window_prefilter_error_total (counters)
  • loki_vl_proxy_window_prefilter_kept_total, loki_vl_proxy_window_prefilter_skipped_total (counters)
  • loki_vl_proxy_window_prefilter_duration_seconds (histogram)

Recommended loop:

  1. Increase -query-range-adaptive-max-parallel only if latency EWMA stays below target and error EWMA stays near zero during peak load.
  2. Increase -query-range-history-cache-ttl only if disk utilization and cache hit ratio justify it.
  3. Set -disk-cache-max-bytes before increasing TTL windows beyond 24h.

Metadata vs Live Query Freshness​

The proxy keeps faster-changing paths conservative and slower-changing metadata slightly warmer:

  • query and query_range stay on short TTLs so new log lines remain visible quickly
  • labels, label_values, series, patterns, detected_fields, and detected_labels can cache longer because they change more slowly and are more expensive to rebuild
  • Drilldown now prefers native VictoriaLogs metadata (field_names, field_values, streams) where possible, which reduces the amount of raw log rescanning needed on cache misses

Multitenancy​

FlagEnvDefaultDescription
-tenant-labelTENANT_LABEL""VL field name for label-based tenant routing. When set, an unmapped X-Scope-OrgID (other than the default-tenant aliases and *) is sent as a VictoriaLogs extra_stream_filters constraint {"<tenant-label>":"<orgID>"} instead of AccountID/ProjectID headers; request parameters cannot override it. The field must be a VictoriaLogs stream field (part of _stream_fields). Use when all data is under the VL default tenant (0:0). Explicit tenant-map entries take priority.
-tenant-mapTENANT_MAP—JSON string→int tenant mapping
-tenant-map-fileTENANT_MAP_FILE""Path to a YAML or JSON file mapping Loki X-Scope-OrgID strings to VictoriaLogs AccountID/ProjectID. Reloaded on SIGHUP and automatically when the file changes (see -tenant-map-reload-interval). Suitable for Kubernetes ConfigMap volumes. File entries override -tenant-map inline entries for the same key.
-tenant-map-reload-interval(flag only)30sHow often to poll -tenant-map-file for mtime changes. Set to 0 to disable polling (SIGHUP-only reload).
-tenant-limits-allow-publishTENANT_LIMITS_ALLOW_PUBLISHbuilt-in allowlistComma-separated limit fields exposed on /config/tenant/v1/limits and /loki/api/v1/drilldown-limits
-tenant-default-limitsTENANT_DEFAULT_LIMITSβ€”JSON map of Loki limits for every tenant, enforced and published; see Published tenant limits
-tenant-limitsTENANT_LIMITSβ€”JSON map of per-tenant Loki limits keyed by X-Scope-OrgID, above -tenant-default-limits; enforced and published
-auth.enabledβ€”falseRequire X-Scope-OrgID on query requests; a missing header returns 401. The header value is not authenticated
-require-tenant-headerβ€”falseReject requests missing X-Scope-OrgID with 401, independent of -auth.enabled
-tenant.allow-globalβ€”falseAllow an unmapped X-Scope-OrgID: * to bypass tenant scoping and use the backend default tenant. Without it, unmapped * returns 403; an explicit "*" tenant-map entry always takes precedence
-forward-tenant-headerFORWARD_TENANT_HEADERtrueWhen false, suppresses X-Scope-OrgID header forwarding to the backend

Tenant Resolution Order​

  1. No header β€” when -auth.enabled=false and -require-tenant-header=false, requests without X-Scope-OrgID use VictoriaLogs' backend default tenant, which is AccountID=0 and ProjectID=0; otherwise they return 401
  2. Tenant map lookup β€” if -tenant-map is configured and the org ID matches a key, use the mapped AccountID/ProjectID
  3. Explicit tenant map override β€” if a tenant map contains an exact key such as "0" or "fake", that explicit mapping wins
  4. Default-tenant aliases β€” X-Scope-OrgID values 0, fake, and default map to VictoriaLogs' built-in 0:0 tenant without rewriting headers
  5. Wildcard global bypass β€” X-Scope-OrgID: * is a proxy-specific convenience. Unless the tenant map has an explicit "*" entry, it uses the backend default tenant only when -tenant.allow-global=true and returns 403 otherwise (also in -tenant-label mode)
  6. Label routing β€” with -tenant-label set, any other org ID is accepted and scoped with a stream-field constraint instead of tenant headers
  7. Numeric passthrough β€” if the org ID is a number other than the default-tenant alias case (for example "42"), pass it directly as AccountID with ProjectID: 0
  8. Fail closed β€” unmapped non-numeric org IDs are rejected with 403 Forbidden

Multi-Tenant Query Headers​

The proxy accepts Loki-style multi-tenant query headers on read/query endpoints by separating tenant IDs with |, for example X-Scope-OrgID: team-a|team-b.

  • Supported on query-style endpoints such as /query, /query_range, /labels, /label/.../values, /series, /index/*, and Drilldown metadata endpoints
  • Rejected on /tail, delete, and write paths
  • Query results inject a synthetic __tenant_id__ label per tenant, matching Loki's documented query behavior
  • __tenant_id__ matchers in the leading selector narrow the tenant fanout set before backend requests are sent
  • multi-tenant detected_fields and detected_labels use exact merged value unions, so cardinality does not double-count identical values across tenants
  • like Loki, a multi-tenant request fails as a whole when any tenant's sub-request fails: 400 when a tenant rejects the query, 504 on a timeout, 500 for any other backend failure (a Grafana Drilldown partial-results reply from one tenant counts as that tenant failing); nothing is merged or cached from the other tenants
  • Wildcard * is not allowed inside a multi-tenant header; use explicit tenant IDs
  • fanout is safety-capped to prevent one request from exploding into an unbounded number of backend queries
  • merged multi-tenant response bodies are also size-capped before they are returned to the client

Tenant Map File Format​

The tenant map file (YAML or JSON) maps each Loki X-Scope-OrgID string to a VictoriaLogs AccountID/ProjectID pair. Both fields are required and must be non-negative integers represented as strings.

YAML (recommended β€” easier to comment and diff):

# /etc/loki-vl-proxy/tenant-map.yaml
#
# Each key is the exact string sent in X-Scope-OrgID by Grafana (or Loki clients).
# account_id maps to VictoriaLogs AccountID header (tenant namespace).
# project_id maps to VictoriaLogs ProjectID header (sub-account partition).
# Both fields must be non-negative integers. Use "0" for the VL default.

# Single VictoriaLogs account, separate projects per team:
team-alpha:
account_id: "1"
project_id: "1"

team-beta:
account_id: "1"
project_id: "2"

# Different VL accounts (e.g. separate VL clusters behind a load balancer):
ops-prod:
account_id: "42"
project_id: "0"

ops-staging:
account_id: "43"
project_id: "0"

# Explicit override for the default-tenant alias "0".
# Without this, X-Scope-OrgID: "0" resolves to VL 0:0 automatically.
# Add it only when you need to redirect "0" to a different VL project:
# "0":
# account_id: "0"
# project_id: "99"

JSON (use when injecting from CI/CD tooling):

{
"team-alpha": { "account_id": "1", "project_id": "1" },
"team-beta": { "account_id": "1", "project_id": "2" },
"ops-prod": { "account_id": "42", "project_id": "0" }
}

Validation rules enforced at load time:

  • account_id and project_id must both be present and non-empty.
  • Both must parse as non-negative integers that fit in a uint32 (0–4294967295).
  • Strings that are not valid integers (for example "prod") are rejected immediately β€” the proxy does not start if the tenant map is invalid.
  • The file extension determines the parser: .json β†’ JSON, anything else β†’ YAML.

Configuration Examples​

# Via flag (inline JSON β€” convenient for single-container deployments)
./loki-vl-proxy \
-tenant-map='{"team-alpha":{"account_id":"1","project_id":"1"},"team-beta":{"account_id":"1","project_id":"2"}}'

# Via environment variable (same JSON format)
export TENANT_MAP='{"ops-prod":{"account_id":"42","project_id":"0"}}'
./loki-vl-proxy -backend=http://victorialogs:9428

# Via file with hot-reload β€” ideal for Kubernetes ConfigMap volumes
./loki-vl-proxy \
-backend=http://victorialogs:9428 \
-tenant-map-file=/etc/loki-vl-proxy/tenant-map.yaml \
-tenant-map-reload-interval=30s

Kubernetes ConfigMap: use -tenant-map-file so the proxy can pick up tenant changes without a restart. The proxy polls the file's modification time every -tenant-map-reload-interval (default 30 s); Kubernetes atomically replaces the ConfigMap symlink, so the new mtime is visible immediately. See k8s-tenant-map-configmap.yaml for a complete Deployment snippet.

Global single-tenant (no map needed): if every Grafana datasource sends the same X-Scope-OrgID value ("0", "fake", or "default") and all data lives in VictoriaLogs' default tenant, you do not need -tenant-map at all. The proxy resolves those three aliases to VL's built-in AccountID=0, ProjectID=0 automatically.

Grafana Datasource per Tenant​

datasources:
- name: Logs (team-alpha)
type: loki
url: http://loki-vl-proxy:3100
jsonData:
httpHeaderName1: X-Scope-OrgID
secureJsonData:
httpHeaderValue1: team-alpha

Grafana Datasource for Explicit Multi-Tenant Reads​

datasources:
- name: Logs (team-a + team-b)
type: loki
url: http://loki-vl-proxy:3100
jsonData:
httpHeaderName1: X-Scope-OrgID
secureJsonData:
httpHeaderValue1: team-a|team-b

Use __tenant_id__ in LogQL when the datasource fans out to more than one tenant:

{app="api-gateway", __tenant_id__="team-b"}
{service_name="checkout", __tenant_id__=~"team-(a|b)"}

Grafana Datasource for Single-Tenant VictoriaLogs​

If the backend only uses VictoriaLogs' default tenant, you can keep Grafana simple:

datasources:
- name: Logs (global tenant)
type: loki
url: http://loki-vl-proxy:3100
jsonData:
httpHeaderName1: X-Scope-OrgID
secureJsonData:
httpHeaderValue1: "0"

X-Scope-OrgID: "0", X-Scope-OrgID: "fake", and X-Scope-OrgID: "default" resolve to VL's default 0:0 tenant in Loki-compatible single-tenant mode. X-Scope-OrgID: "*" remains a proxy-specific wildcard convenience and requires explicit -tenant.allow-global=true unless the tenant map contains an explicit "*" entry.

Hot Reload​

Tenant mappings and field mappings can be reloaded without restart via SIGHUP:

kill -HUP $(pidof loki-vl-proxy)

Published Tenant Limits Compatibility Surface​

The proxy also exposes tenant-limit compatibility endpoints for Grafana and Logs Drilldown bootstrap:

  • GET /config/tenant/v1/limits returns YAML
  • GET /loki/api/v1/drilldown-limits returns JSON

Both publish the tenant's own limits, and the Loki query limits among them are the values its queries run under: one resolver serves enforcement and publishing. Each value resolves, highest first:

  1. the tenant's entry in -tenant-limits (keyed by X-Scope-OrgID)
  2. -tenant-default-limits
  3. the proxy flag
  4. Loki's default (v3.7.7)
Loki limitFlagLoki default / proxy defaultEnforcement
max_query_series-max-stats-query-series500 / 500every metric route, range and instant: 400 maximum number of series (N) reached for a single query; ..., or for Logs Drilldown the busiest N series with Loki's partial-result warning
max_entries_limit_per_query-max-entries-limit-per-query5000 / 10000log queries: 400 max entries limit per query exceeded, limit > max_entries_limit_per_query (L > N) (-max-entries-limit-per-query-cap lowers the limit instead); 0 is unlimited
max_query_length-default-max-query-length721h / 0 (unlimited)query_range, series, labels, label values, detected labels, index stats: 400 the query time range exceeds the limit (query length: X, limit: Y)
max_query_lookbackβ€”0 / 0the same requests (an instant query only when it is a metric query, as Loki's frontend sends instant log queries past its limits): a start before now - lookback is moved to it; a request ending before it is answered empty without running
max_query_rangeβ€”0 / 0metric queries: 400 [interval] value exceeds limit: [R] > [M] for a longer range selector
query_timeout-backend-timeout1m / 2mset in a tenant limit, the requests Loki bounds by it (query, query_range, series, labels, label values, index stats, volume) run under it, never a live tail (504 when it expires); it may not exceed -backend-timeout, which still bounds each VictoriaLogs call. Without one, -backend-timeout is the published value

A request for several tenants (X-Scope-OrgID: a|b) is held to them combined as Loki combines them: the smallest max_query_series, and the smallest non-zero value of every other limit. For query_timeout, a tenant without an override counts with -backend-timeout, so a|b runs under the smaller of the two as a whole-request deadline. /loki/api/v1/drilldown-limits publishes those combined limits for such a header (Loki answers it 401, which would leave a multi-tenant Drilldown without its bootstrap); /config/tenant/v1/limits answers it with Loki's 401 multiple org IDs present.

max_query_bytes_read, max_querier_bytes_read and volume_max_series are published at Loki's disabled or default value (0B, 0B, 1000): VictoriaLogs reports no bytes read before a query runs, and the volume endpoints bound their answer by the request limit. Overriding them is rejected at startup, since nothing would apply the published value. The other published fields (discover_log_levels, discover_service_name, log_level_fields, retention_period, otlp_config, ...) are published as configured; they describe the deployment and do not change how the proxy derives service_name or detected_level. Label values requests are capped at the -max-entries-limit-per-query flag value, a proxy protection Loki does not have, whatever a tenant's max_entries_limit_per_query.

One proxy limit is also settable per tenant in the same maps: label_values_max_response_bytes overrides -label-values-max-response-bytes (see Fixed Execution Limits). Loki has no per-tenant equivalent (its bound is the querier's grpc_server_max_send_msg_size), so it is enforced but never published; a multi-tenant label values request fans out per tenant, and each tenant's read is bounded by that tenant's value.

Invalid values are rejected at startup with a message naming the tenant and the field: max_query_series must be a positive integer, max_entries_limit_per_query 0 or positive, durations Loki duration strings ("5m", "30d1h", "0s"), query_timeout positive and not above -backend-timeout, label_values_max_response_bytes a positive number of bytes.

# Every tenant: Loki's defaults for series and entries; team-a gets more series and a 7-day lookback
loki-vl-proxy \
-tenant-default-limits='{"max_entries_limit_per_query":5000,"query_timeout":"1m"}' \
-tenant-limits='{"team-a":{"max_query_series":2000,"max_query_lookback":"7d"}}'

Operational notes:

  • published fields are filtered by -tenant-limits-allow-publish
  • X-Scope-OrgID omitted on these endpoints publishes the limits of the default tenant
  • an empty retention_stream is left out of the payload, as Loki does

OTLP Telemetry​

FlagEnvDefaultDescription
-otlp-endpointOTLP_ENDPOINTβ€”OTLP HTTP endpoint for proxy metrics
-otlp-intervalβ€”30sPush interval
-otlp-compressionOTLP_COMPRESSIONnonenone, gzip, zstd
-otlp-headersOTLP_HEADERSβ€”Comma-separated OTLP HTTP headers in key=value form
-otlp-timeoutβ€”10sHTTP request timeout
-otlp-tls-skip-verifyβ€”falseSkip TLS verification
-otel-service-nameOTEL_SERVICE_NAMEloki-vl-proxyservice.name resource metadata for OTLP metrics/logs (not duplicated per JSON log line)
-otel-service-namespaceOTEL_SERVICE_NAMESPACEβ€”service.namespace resource metadata for OTLP metrics/logs (not duplicated per JSON log line)
-otel-service-instance-idOTEL_SERVICE_INSTANCE_IDβ€”service.instance.id resource metadata for OTLP metrics/logs (not duplicated per JSON log line)
-deployment-environmentDEPLOYMENT_ENVIRONMENTβ€”deployment.environment.name resource metadata for OTLP metrics/logs (not duplicated per JSON log line)

Go Runtime Tuning​

FlagEnvDefaultDescription
-go-mem-limitβ€”0Explicit GOMEMLIMIT in bytes. Overrides -go-mem-limit-percent when set (0 = disabled)
-go-mem-limit-percentβ€”85GOMEMLIMIT as a percentage of the detected container memory limit (from /proc cgroup). Ignored when -go-mem-limit is set
-go-gc-percentβ€”200GOGC target percentage. The proxy default (200) halves GC frequency versus Go's built-in default of 100, trading higher peak RSS for lower GC CPU. Set -1 to disable GC target entirely

See Performance β€” Go Runtime Tuning for guidance on sizing these for your deployment.

HTTP Hardening​

FlagEnvDefaultDescription
-http-read-timeoutβ€”30sServer read timeout
-http-read-header-timeoutβ€”10sServer read header timeout
-http-write-timeoutβ€”120sServer write timeout
-http-idle-timeoutβ€”120sServer idle timeout
-http-max-header-bytesβ€”1MBMaximum header size
-http-max-body-bytesβ€”10MBMaximum request body size
-http-conn-max-ageβ€”10mMaximum lifetime for downstream HTTP/1.x keepalive connections
-http-conn-max-age-jitterβ€”2mJitter applied to downstream connection age rotation
-http-conn-max-requestsβ€”256Maximum requests per keepalive connection
-http-conn-overload-max-ageβ€”90sShorter connection lifetime during backpressure

Grafana Compatibility​

FlagEnvDefaultDescription
-max-linesβ€”1000Default max lines per query
-manual-range-metric-row-limitβ€”1000000Maximum rows fetched per proxy-side range-metric evaluation (rate, count_over_time, etc.): raw log rows, or stats rows (one per non-empty step bucket and series) on the window stats path. Exceeding it rejects the query with HTTP 502 (manual range metric row limit exceeded ... increase -manual-range-metric-row-limit) instead of returning truncated results. On the window stats path the limit bounds each bucket grid separately (at most one row per step and series), not their sum. For pipe-free selectors over at least -backend-heavy-query-min-range, a one-row count runs first and rejects the query before any log line is fetched. count_over_time, rate, bytes_over_time and bytes_rate normally avoid raw rows by summing stats_query_range buckets of gcd(step, range) (one bucket per step when the range is shorter than the step). When that grid would need more than 11,000 buckets, and for every topk/bottomk input, they instead read one window-sized seed bucket and two step-sized bucket grids from a streamed VictoriaLogs stats pipe and derive each window exactly; topk/bottomk then rank every series at each step. Raw rows remain the fallback only when the bucket is below 1 ms or the evaluation grid is not epoch-aligned and VictoriaLogs is known to be older than v1.45 (an undetected version keeps the bucket path: the offset argument is ignored at worst). The encoded range-metric response is capped at 64 MiB (HTTP 503 above it), and Loki's 11,000-points-per-series limit bounds the evaluation steps. Proxy memory on the window stats path is about 100 bytes per stats row, so the default bounds one evaluation to roughly 100 MiB
-ordered-json-metric-max-bytesβ€”1073741824Safety cap, not the fix for slow queries: the most bytes the proxy-side ordered JSON metric evaluator reads from the VictoriaLogs raw rows response, and the largest response it builds. 0 uses the 1 GiB default; there is no upper bound. Exceeding it rejects the query with HTTP 502 (ordered JSON metric response exceeds N bytes; narrow the query or increase -ordered-json-metric-max-bytes) instead of returning partial results. The default admits about one million rows of 1 KiB, the -manual-range-metric-row-limit default; on the e2e generator (about 290k lines and 300 MiB of rows per hour) that is roughly 3.5 hours. Rows are streamed, not buffered, so raising it costs VictoriaLogs scan work and network transfer; proxy memory grows with the retained rows, which -manual-range-metric-row-limit bounds. Lower it to protect VictoriaLogs from long raw scans. Helm: extraArgs.ordered-json-metric-max-bytes. The evaluator answers only | json range metrics that need Loki's per-line semantics (surviving parser errors, _extracted labels, keep, a second parser, filters that accept the empty value without drop __error__). Summed count_over_time, rate, bytes_over_time and bytes_rate over a single | json or | logfmt parser, grouped by parsed labels (including detected_level and labels with underscores such as service_version) or ungrouped, with string label filters (=, !=, =~, !~) after the parser, are computed from VictoriaLogs stats_query_range buckets over unpack_json/unpack_logfmt and filter pipes and do not read raw rows; this includes Grafana's logs volume queries for plain, | json and | logfmt selectors with query-builder filters, and Logs Drilldown field breakdowns. Parser errors must be dropped unless a filter rejects the empty value (pipeline="logs/loki", field!=""), which no unparsed line passes. A label VictoriaLogs stores under another spelling (service_version as service.version) is read from the stored field where the line carries it, as Loki reads structured metadata before a parsed key. Before answering, cheap | limit 1 probes check the range for a line the two parsers would read differently: a key parsed before a syntax error, a repeated key, an array value, a key spelled differently from the label Loki sanitizes it to (service-version, service.version, a nested {"service":{"version":...}}), a logfmt tab, or, when errors are not dropped, a stored value of a filter label on a line whose body yields none. One such line keeps the previous route: this evaluator for | json, the native stats route for plain and | logfmt. loki_vl_proxy_parser_metric_evaluations_total{evaluator,reason} counts which evaluator answered and why. Grafana Logs Drilldown requests without a JSON parser keep their dedicated routes.
-label-filter-refill-max-pagesβ€”8Loki-compatible profile: further VictoriaLogs pages of at most limit rows a log query reads when a Loki label filter on a log line key no earlier stage exposes dropped rows VictoriaLogs matched ({app="x"} | user="u1" without a parser), so the page still holds limit lines. A response reads at most (1 + N) x limit rows. 0 reads no further page: such a page can return fewer lines than the limit.
-backend-timeoutβ€”120sTimeout for non-streaming VL backend requests. The remaining request budget (this timeout or an earlier request deadline) is passed to VictoriaLogs as its timeout query argument, so VictoriaLogs stops executing a query the proxy has given up on; VictoriaLogs caps it at its own -search.maxQueryDuration. A client disconnect cancels the in-flight VictoriaLogs request
-backend-min-versionβ€”v1.40.0Minimum VictoriaLogs version considered fully supported at startup compatibility gate (first minor of the oldest supported line; the proxy refuses to start below it unless -backend-allow-unsupported-version=true)
-backend-allow-unsupported-versionβ€”falseAllow startup when detected backend version is lower than -backend-min-version (unsafe override)
-backend-version-check-timeoutβ€”5sTimeout for startup backend version compatibility check (/health)
-stream-responseβ€”falseStream via chunked transfer
-response-compressionβ€”autoResponse compression codec: auto, gzip, none
-response-compression-min-bytesβ€”1024Minimum response size before frontend compression starts; smaller responses stay identity
-response-gzipβ€”trueDeprecated compatibility switch; false disables response compression when -response-compression is unset
-derived-fieldsβ€”β€”JSON derived fields for trace linking
-forward-headersβ€”β€”HTTP headers to forward to VL
-forward-authorizationβ€”falseForward client Authorization header to VL backend (adds Authorization to forwarded headers list)
-forward-cookiesβ€”β€”Cookie names to forward to VL
-backend-basic-authβ€”β€”user:password for VL basic auth
-backend-compressionβ€”autoUpstream compression preference: auto, gzip, zstd, none. auto detects loopback backends (localhost/127.x/::1) and disables compression for co-located VL, otherwise advertises zstd, gzip.
-backend-tls-skip-verifyβ€”falseSkip TLS on VL connection
-tail.allowed-originsβ€”β€”Comma-separated WebSocket Origin allowlist for /loki/api/v1/tail
-tail.modeβ€”autoauto, native, or synthetic for /tail streaming mode

-tail.mode=auto prefers native backend tailing and falls back to synthetic polling when native streaming is unavailable. native disables the fallback, and synthetic forces the polling bridge even when the backend can stream natively.

/loki/api/v1/tail accepts at most 4 KiB per client WebSocket message and closes the connection with code 1009 when a client sends more. Log frames sent to Grafana are not limited by this.

Compression notes:

  • -response-compression=auto keeps the optimized gzip path enabled for clients that advertise gzip.
  • -response-compression-min-bytes=1024 keeps small control-plane responses uncompressed and the proxy applies higher effective thresholds on metadata-heavy routes.
  • peer-cache /_cache/get fetches follow the same preference order: zstd, then gzip, then identity.
  • -backend-compression=auto auto-detects the backend address at startup: when the host is a loopback address (localhost, 127.x.x.x, ::1) it sends Accept-Encoding: identity, eliminating 25–35% decompression CPU with no bandwidth cost (loopback has no network overhead); when the host is remote it advertises zstd, gzip and the proxy safely decodes either on the way back. Use none to force identity for non-loopback local deployments (e.g. same host via a non-loopback IP).
  • current VictoriaLogs docs describe HTTP response compression in general terms, not a guaranteed zstd select-query path, so in practice many deployments will still observe gzip or identity from stock VictoriaLogs today.
  • Grafana 12.4.2 datasource proxy requests advertised Accept-Encoding: deflate, gzip, not zstd, in local verification, so the proxy keeps the frontend surface on gzip/identity only.
  • The Helm chart now pins -response-compression=gzip and -response-compression-min-bytes=1024 by default.

Backend version gate notes:

  • On startup, proxy probes backend /health and inspects response headers for a VictoriaLogs semver.
  • If detected version is below -backend-min-version, startup is blocked by default.
  • Set -backend-allow-unsupported-version=true to bypass the gate at your own risk.
  • If version cannot be detected from headers, proxy logs a warning and continues startup.
  • Version-gated behaviour keeps VictoriaLogs-native evaluation while the version is unknown: range metrics still use stats_query_range buckets anchored with the offset argument, which a release predating it (before v1.45) ignores, leaving the epoch-aligned buckets such a backend has anyway. Only a version known to be older than v1.45 restricts buckets to epoch-aligned grids. While the version remains unknown, for example when the startup probe ran before VictoriaLogs was ready or /metrics is not routed through a proxy in front of it, the proxy retries the /metrics version probe in the background at most once every 5 minutes, bounded by -backend-version-check-timeout. There is no version override flag; make /metrics or a version response header reachable to enable version-gated paths.

Built-In Protection Defaults​

All protection controls are tunable via CLI flags:

FlagEnvDefaultDescription
-max-concurrentβ€”100Maximum concurrent requests admitted per replica; excess requests get 503 immediately (Retry-After: 5). The same value also bounds concurrent hot/cold/alerting backend operations, including fanout, until each response body is consumed or closed. 0 disables both
-rate-limit-per-secondβ€”50Per-client request rate limit (requests/second)
-rate-limit-burstβ€”100Per-client burst allowance above the rate limit
-cb-fail-thresholdβ€”5Number of backend failures within the sliding window required to open the circuit breaker
-cb-open-durationβ€”10sHow long the circuit breaker stays open before entering half-open state
-cb-window-durationβ€”30sSliding window duration for failure counting; sporadic failures outside the window do not accumulate
-coalescer-disabledβ€”falseDisable request coalescing (singleflight); every concurrent request makes its own backend call β€” useful with -cache-disabled to measure raw translation overhead
-backend-max-concurrent-heavy-queriesβ€”2Maximum concurrent heavy VictoriaLogs calls per replica (see Heavy VictoriaLogs Query Admission). 0 disables the limiter
-backend-heavy-query-queue-waitβ€”20sHow long a heavy call waits for a slot before the request fails with 429 too many outstanding requests. 0 rejects immediately when all slots are busy
-backend-heavy-query-min-rangeβ€”6hTime range from which stats_query_range, stats_query, hits and unbounded raw query calls count as heavy, and metadata listings count as long-range scans. Must be > 0
-backend-max-concurrent-metadata-scansβ€”8Ceiling of the adaptive limit on concurrent long-range VictoriaLogs metadata listings per replica (stream_field_names, stream_field_values, field_names, field_values, streams spanning at least -backend-heavy-query-min-range), on their own limiter (see Long-range metadata scans). 0 disables the limiter
-backend-min-concurrent-metadata-scansβ€”1Floor of that adaptive limit: latency and failure feedback never shrink it below this. VictoriaLogs' select slots and memory still hold scans back when it has no room
-backend-metadata-scan-memory-headroomβ€”0.4Share of the memory available to VictoriaLogs (from its /metrics) that long-range scans must leave free. 0 disables the memory gate
-backend-metadata-scan-latency-toleranceβ€”1.5How many times its no-load duration a long-range scan that ran beside others may take before the adaptive limit shrinks. Must be >= 1
-metadata-inventory-parallelismβ€”4Bucket listings one label or field listing may have in flight while it is assembled from the time-bucketed inventory (see Label inventory). 0 turns the inventory off: every listing is one VictoriaLogs call over the whole range

Per-client rate limits identify clients by connection source IP and return 429 (Retry-After: 1). -rate-limit-per-second=0 disables them; a burst of 0 with a positive rate rejects every request.

Shape per-client and global traffic at Grafana, ingress, or an outer proxy layer for additional control beyond these flags.

Heavy VictoriaLogs Query Admission​

VictoriaLogs lets each stats, sort or uniq pipe of a single query use up to 40% of its allowed memory (-memory.allowedPercent) and does not account for other queries running at the same time. A few long-range stats or raw-row calls running together can therefore exhaust VictoriaLogs memory, however each one is bounded. The proxy admits a bounded number of heavy VictoriaLogs calls per replica and queues the rest.

A VictoriaLogs call is heavy when it is:

  • a raw /select/logsql/query fetch whose row bound (the limit argument or a trailing | limit N pipe) exceeds 10,000 rows, as proxy-side metric evaluation uses. Log queries, bounded by Loki's max_entries_limit_per_query sizes, are not heavy;
  • an unbounded /select/logsql/query call spanning at least -backend-heavy-query-min-range;
  • a stats_query_range, stats_query or hits call spanning at least -backend-heavy-query-min-range, with a bucket grid finer than 11,000 buckets (Loki's resolution limit), or without a resolvable time range at all (VictoriaLogs then scans every stream it stored).

Metadata lookups (stream_field_names, stream_field_values, field_names, field_values, streams) and tail requests are never heavy, so label browsing keeps working while heavy queries queue. Long-range metadata listings have a limiter of their own, -backend-max-concurrent-metadata-scans, described below.

At most -backend-max-concurrent-heavy-queries heavy calls run at once. A heavy call that finds every slot busy waits up to -backend-heavy-query-queue-wait, a budget shared by every heavy call of the same request; a released slot goes to the waiting tenant that holds the fewest slots. A rejection is proxy backpressure, not a backend failure: it never counts towards the circuit breaker, and Grafana-sourced requests receive the same 429 rather than an empty partial result. When the wait expires the request fails like Loki's query scheduler with a full queue: HTTP 429, body too many outstanding requests: heavy VictoriaLogs queries are limited to -backend-max-concurrent-heavy-queries=N per replica and this query waited -backend-heavy-query-queue-wait=D; .... The rejection is remembered for the request, so a fallback path does not queue a second time. Waits and rejections are exported as the internal operation backend_heavy_query_admission with outcome admitted, rejected or canceled.

Sizing:

Deployment-backend-max-concurrent-heavy-queriesNotes
Default, one VictoriaLogs node with 4-8 GiB2Two concurrent heavy stats states stay within 80% of VictoriaLogs allowed memory
Small (one VictoriaLogs with 1-2 GiB, a few Grafana users)1, queue wait 25sSerializes heavy work; dashboards with many 24h panels load panel by panel
Large (VictoriaLogs with 32 GiB or more, or a cluster)4-8, queue wait 20sRaise with VictoriaLogs memory; the fleet-wide total is replicas Γ— this value

The limit applies per proxy replica: with three replicas and the default, up to six heavy calls reach VictoriaLogs at once. Size VictoriaLogs memory for the fleet-wide total, or lower the value as replicas grow. Keep -backend-heavy-query-queue-wait below the Grafana data source timeout (30s by default) so queued panels fail with the documented 429 instead of a client timeout. The 20s default lets a Grafana Logs Drilldown service page over 7d, about 30 parallel heavy calls, finish with the default two slots.

Long-range metadata scans​

Loki answers /labels and /label/{name}/values from its index in milliseconds whatever the range. VictoriaLogs has no such index: a stream_field_names, stream_field_values, field_names, field_values or streams listing reads every row in its range (stream_field_names first lists every distinct stream), and VictoriaLogs admits any number of them (-search.maxConcurrentRequests) without accounting for their combined memory. What one scan costs depends on the data: on v1.52.0 with the e2e stack's 7-day retention, a 7-day stream_field_names scan held about 0.8 GiB and took 9 s one day and 35 s the next (the stack holds more than 100 000 distinct streams per 6 hours); seven at once from seven proxy replicas each ran past VictoriaLogs' 60 s limit and OOM-killed an 8 GiB VictoriaLogs, and eleven replicas warming their caches at a static two scans each did the same within a minute of starting. A large deployment can need many times that per scan; a small one a few megabytes. Two mechanisms keep this bounded without a constant sized to one data set.

Label inventory​

Label names and values are listed from a time-bucketed inventory, so a request scans only what nobody has scanned yet. A listing over a range is the union of the listings of any partition of it, with hits summed (VictoriaLogs treats end as exclusive), so the proxy splits each window into UTC-aligned day, hour, 5-minute and minute buckets (at each point the largest aligned bucket that fits), keeps each sealed bucket in the read cache (memory, disk and the owner peer, so replicas share it), merges a missing hour or day from its cached children, and asks VictoriaLogs only for the unaligned edges, the last minute (never cached) and missing buckets, -metadata-inventory-parallelism (4) at a time per request. The merged listing is ordered as VictoriaLogs orders its own (hits descending, then natural order), so the answer is exactly that of one full-range call. A Grafana refresh that moves a 7-day window by seconds reads its two edges of under a minute each, plus, when an edge crosses a 5-minute, hour or day boundary, the minute and 5-minute buckets that newly fit there.

A sealed bucket changes only when rows arrive with old timestamps. Each bucket is revalidated after a delay that grows with its age: the labels TTL (-labels-cache-ttl, 5 minutes) for buckets from the last hour, three times that up to 6 hours old, ten times up to a day, twenty times beyond, at most an hour; an empty bucket of a countable query that ended within -max-metadata-cache-freshness is confirmed on every listing (one row count covers each contiguous run of empty buckets of any size, shared by concurrent listings; it is cheap for * and stream filters, but with a field filter it reads the filtered column of every row in the run's range, so it is not free there; a run that received rows is halved to find the buckets to rescan, which are rescanned at once); other empty buckets are revalidated after the negative TTL (30 s) within the last hour, and older ones follow the age schedule. Hour and day buckets of a query made only of stream and field filters (no pipes, no word or phrase filters, whose count would read every message) carry their row count, read before the scan: revalidating them is a count (0.03-0.2 s per day for * on the e2e stack, against 5-20 s for the scan), a changed count forces a rescan, and they are revalidated at least every three labels TTLs. Rows in VictoriaLogs are append-only and retention drops whole days, so an unchanged count means an unchanged listing.

A bucket that fails (a 429, a backend error) stops the listing from starting new buckets but lets the scans already running finish and stay cached, so the retry continues where it stopped. A listing with a limit, a query that is not row-local (stats, uniq, sort, limit, subqueries, _time filters, query options), a window under 10 minutes, and a listing whose bucket is larger than the cache keeps are one full-range call, as before. -metadata-inventory-parallelism=0 turns the inventory off. Its work is exported as loki_vl_proxy_internal_operation_total{operation="metadata_inventory"} with outcome cached, revalidated, partial or scanned.

Adaptive admission​

The long scans that remain β€” listings spanning at least -backend-heavy-query-min-range or without a time range that the inventory does not split, and the day bucket scans of long listings β€” go through a limiter of their own, separate from the heavy-query one so a 7-day label browser never waits behind 7-day charts. Each such scan holds a slot only while its own call runs. Each replica learns the limit from what it can see of the backend:

  • Fleet share. VictoriaLogs exports the selects it runs for all its clients (vl_concurrent_select_current) on /metrics. The ones it runs for others count against this replica's limit, so replicas that never talk to each other hold VictoriaLogs near one replica's limit rather than replicas Γ— limit: a long scan already spreads over every core, and more at once only slow each other down. A scan deciding for the first time counts everything running (a scan of a listing nobody has measured first waits a random fraction of a second, so replicas that receive it at the same moment see each other); a scan that has waited counts only what kept running through its wait, since long work stays and short calls leave gaps. Below -backend-min-concurrent-metadata-scans a replica may still start a scan while others' long work uses less than a quarter of VictoriaLogs' select slots, so other clients' charts cannot keep label listings out for ever.
  • Latency. Each scan class (endpoint, range bucket and tenant) tracks its no-load duration the way Vegas tracks the base round-trip time: a faster scan lowers it at once, slower ones raise it slowly, per row when the scan's rows are known (counted buckets). A wave of scans slower than -backend-metadata-scan-latency-tolerance (1.5) times it shrinks the limit by 20%; a scan that used the whole limit within the tolerance grows it by 1/limit. The limit starts at 2 and stays between -backend-min-concurrent-metadata-scans (1) and -backend-max-concurrent-metadata-scans (8).
  • Distress. A transport failure, a timeout or a 5xx on a scan halves the limit, once for the scans that failed together, and so does a rise of VictoriaLogs' queue-timeout counter; a rise of its queueing counter counts as a slow wave. A scan whose client went away, or that VictoriaLogs refused with a 4xx, is neither distress nor a sample: nothing is learned from it. While VictoriaLogs runs all its select slots, or for up to 10 s after it stops answering /metrics having answered it before, nothing new is admitted; a /metrics that keeps failing leaves the latency and distress feedback.
  • Memory headroom. A scan is admitted only while the memory VictoriaLogs has in use, plus what this replica's scans in flight reserved and have not shown yet, plus the selects VictoriaLogs runs for others counted at this scan's cost (another replica's scan that started a moment ago holds none of its memory yet), plus the scan's own cost stays below 1 - -backend-metadata-scan-memory-headroom (60%) of vm_available_memory_bytes. Memory in use is the anonymous resident set (process_resident_memory_anon_bytes, else process_resident_memory_bytes) minus the Go heap VictoriaLogs has freed but not yet returned to the OS (go_memstats_heap_idle_bytes - go_memstats_heap_released_bytes), which the next scans reuse. The cost of each class is learned from the peak growth of memory in use while its scans ran; measured beside other work it is only an upper bound. A listing a tenant has not scanned over this range yet takes the cost of the same listing over another range (scaled up by the ratio of the ranges from a shorter one), else twice another tenant's; a listing nobody has measured runs alone, only while VictoriaLogs runs no other select, and at most one at a time per replica. The only scan on an idle VictoriaLogs may use the headroom up to the brake mark below; one whose cost is only an upper bound, or was measured more than an hour ago, runs alone to be measured again rather than being refused for good. /metrics (about 40 KB) is read at most every 100 ms while scans are admitted and every 500 ms while they run.
  • Busy backends. Where some select always runs, no listing is ever measured alone. A synchronous request whose listing has no cost of its own measured alone therefore goes ahead after waiting 2-4 s (randomised, so replicas that met the listing together see each other), on a fresh /metrics reading, once memory in use has held steady for 5 s (a large scan that started a moment ago keeps it growing) and no select started during its wait, reserving more than half of the room left in the budget: such escapes go one at a time.
  • Brake. Behind every estimate, when memory in use passes a third of the way into the headroom (73% of the available memory by default; the container also needs room for the page cache) a replica stops its youngest scan, one every 500 ms while the memory stays above the mark. Its request gets the 429 with the scan was stopped to keep VictoriaLogs' memory inside the headroom, the stop counts as distress, and what the scan had grown to becomes its class's cost, so the class is not started again for an hour unless it fits below the mark.

A backend whose /metrics has no vl_concurrent_select_* series (a vmauth in front of VictoriaLogs, an older release) turns the fleet and memory signals off by itself β€” that process's memory says nothing about VictoriaLogs β€” and leaves the latency and distress feedback.

A synchronous request (Explore's label browser, /series, a 7-day /labels, Logs Drilldown's detected_fields whose field_names listing covers its whole range) waits up to -backend-heavy-query-queue-wait for each long scan (a label inventory listing's day-bucket scans together at most twice that), re-reading /metrics every 250 ms, and then fails with HTTP 429, too many outstanding requests: long-range VictoriaLogs metadata scans are limited to at most -backend-max-concurrent-metadata-scans=8 per replica (adaptive limit 1, floor -backend-min-concurrent-metadata-scans=1, VictoriaLogs running 7 of 16 selects, VictoriaLogs memory 43% used with 0% reserved by this replica's scans against -backend-metadata-scan-memory-headroom=0.40) and this query waited -backend-heavy-query-queue-wait=20s, so the message says which signal closed the door; the buckets it did scan stay cached, so a retry continues where it stopped. The startup warm-up covers only the presets below -backend-heavy-query-min-range (the 1h preset by default), and only while VictoriaLogs has memory and select slots to spare; the long presets are warmed by the jittered keep-warm loop, or by the first request that needs them. Background inventory work β€” the startup warm-up, the keep-warm loop and the stale-entry refreshes of labels, label values and detected fields β€” never waits and never takes the last slot while a request could use it: a refresh that finds the limiter closed is skipped and retried on its next schedule, the cached answer keeps serving, a preset window whose warm-up failed (other than by being skipped) backs off (one keep-warm interval, doubling to one hour), the warm-up runs as the /labels request of a client without a tenant header so the buckets it fills are the ones such requests read, and the keep-warm interval is jittered by up to a quarter so replicas started together drift apart. Identical concurrent listings and bucket fills within a replica share one backend call; a request that shared a fill with a request that went away fills it itself. Admissions, waits, client rejections (rejected), background skips (skipped) and scans stopped by the brake (shed) are exported as loki_vl_proxy_internal_operation_total{operation="backend_metadata_scan_admission"}. Drilldown's detected_fields keeps its own scan timeout: it answers an empty list when that expires before a slot is granted, as it did before for a slow scan.

Sizing: the defaults need no tuning for the data size, that is their point. Raise the ceiling only for a VictoriaLogs with many cores and memory to spare; lower it to cap how many long scans a replica starts. Raise the headroom (0.5) when VictoriaLogs serves other clients too or runs close to its limit; lower it (0.2) on a dedicated node with a large limit. The headroom also leaves VictoriaLogs' page cache its room: it reads its parts through it. Lower the tolerance (1.2) to keep label browsing snappy at the cost of concurrency; raise it (2-3) when scan durations are noisy for other reasons. Set the ceiling to 0 to disable the limiter, or the headroom to 0 to keep only the fleet, latency and distress feedback. Keep -metadata-default-lookback bounded (12h by default) so a /labels without a range does not become a full-retention scan.

Fixed Execution Limits​

These protective limits bound what one request may do. Rejections are errors, not empty results; see Security hardening migration for rollout notes. Every limit that an operator may need to move is a flag β€” the table below names them β€” and each error names the flag that bounded it, so a rejection tells you what to raise.

AreaLimitWhen exceeded
line_format64 KiB output per line; 16 MiB per response; 1 MiB template input per line; bounded template execution work400
Binary metric expressionsnesting depth 64; 1,024 child evaluations per request; 256 MiB of captured operand bytes; 64 MiB per operand response and encoded result; 1,000,000 output samples400 for evaluation and operand limits; 500 for errors raised while joining operands (including implicit many-to-one matches)
Manual range-metric rows-manual-range-metric-row-limit (default 1,000,000 raw or stats rows); 64 MiB raw backend response; stats rows of at most 64 KiB each502
Heavy VictoriaLogs calls-backend-max-concurrent-heavy-queries (default 2) running, queued for -backend-heavy-query-queue-wait (default 20s)429 (too many outstanding requests)
Ordered JSON metric evaluator-ordered-json-metric-max-bytes (default 1 GiB) on the raw rows response and on the built response; -manual-range-metric-row-limit rows502
Metric query series-max-stats-query-series (default 500)400 with Loki's maximum number of series (N) reached for a single query on every metric path, range and instant; Grafana Logs Drilldown gets the busiest series and the ... returning partial results warning
Label values responses-label-values-max-response-bytes (default 64 MiB) read from one VictoriaLogs response of /loki/api/v1/label/{name}/values; per tenant as label_values_max_response_bytes500 with Loki's rpc error: code = ResourceExhausted desc = grpc: trying to send message larger than max (N vs. LIMIT) and the flag appended; nothing is cached or indexed
Request coalescer256 MiB per shared response body, or a larger configured response cap such as -label-values-max-response-byteserror instead of silent truncation
VictoriaLogs response aborted after its headers (GET reads: label names and values, series, detected fields, hits and other metadata calls)VictoriaLogs ends a started response with a raw abort line when a query outlives its deadline504 with Loki's request timed out, decrease the duration of the request or add more label matchers (prefer exact match over regex match) to reduce the amount of data processed when the request deadline or the timeout sent to VictoriaLogs ran out (or VictoriaLogs reported its deadline), otherwise 502 VictoriaLogs aborted the response after D, before it was complete; ...; the partial body is never decoded or cached
Hot/cold merge64 MiB buffered per hot or cold responseerror
Multi-tenant reads-multi-tenant-max-fanout (default 64) tenants per request; -multi-tenant-max-merged-response-bytes (default 32 MiB) merged response400 / 413, each naming its flag
/tail client messages4 KiB per client messageWebSocket close 1009

Flags for these limits​

Each takes 0 to mean "use the built-in default", so a configuration that sets them all to 0 behaves exactly like one that sets none of them.

FlagDefaultWhat it bounds
-max-entries-limit-per-query0 (uses 10000)Loki's max_entries_limit_per_query (Loki's default is 5000). A log query (range or instant) asking for more lines gets Loki's 400 max entries limit per query exceeded, limit > max_entries_limit_per_query (L > N); metric queries carry no entry limit; a label values request above it is capped. Per tenant through -tenant-limits / -tenant-default-limits, where 0 is unlimited
-max-entries-limit-per-query-capfalseLower a log query limit above max_entries_limit_per_query to that value and answer, instead of Loki's 400 (the proxy's behaviour before per-tenant limits)
-max-query-length-bytes0 (uses 131072)LogQL query string length. The default is Loki's syntax.maxInputSize, so the proxy rejects only what Loki rejects; lower it to reject long queries earlier
-label-values-max-response-bytes0 (uses 64 MiB)Bytes of the VictoriaLogs answer to a /loki/api/v1/label/{name}/values request: one response, or the size the merged listing would have as one response when the metadata inventory lists it from time buckets. Above it the request fails as Loki's querier does above grpc_server_max_send_msg_size (500 rpc error: code = ResourceExhausted desc = grpc: trying to send message larger than max (N vs. LIMIT), with this flag named), and nothing is cached or indexed. The VictoriaLogs response is about 1.7x the Loki JSON (it carries a hit count per value): 1h of a high-churn pod label on the e2e stack is 2.4 MB (54,405 values), 24h about 53 MB. The default is 16x Loki's default message size, so a day of such a label still answers; lower it towards 4 MiB to fail where Loki's defaults would. Per tenant as label_values_max_response_bytes in -tenant-limits / -tenant-default-limits
-backend-max-buffered-response-bytes0 (uses 64 MiB)Bytes read from one VictoriaLogs response the proxy has to evaluate itself (buffered stats, volume and binary-operand responses, and the encoded metric result). Exceeding it returns 502 naming the flag rather than a truncated result. Proxy memory grows with this value times the concurrent requests that buffer a response
-binary-metric-max-operand-bytes0 (uses 256 MiB)Operand-response bytes one binary metric expression may capture
-binary-metric-max-arrays0 (uses 2000000)JSON arrays one binary metric expression may allocate while joining operands
-max-zero-fill-buckets0 (uses 32768)Buckets the proxy zero-fills in a metric response
-detected-fields-max-scan-lines0 (uses 2000)Log lines the /detected_fields and detected-field-values scan reads per request
-patterns-max-backend-rows0 (uses 20000)Log lines /patterns reads from VictoriaLogs for one request
-patterns-second-pass-max-rows0 (uses 8000)Log lines the /patterns second pass reads when the first pass mined too few patterns
-patterns-second-pass-max-windows0 (uses 8)Windows the /patterns second pass re-reads
-multi-tenant-max-fanout0 (uses 64)Tenants one multi-tenant request may fan out to; more returns 400 naming the flag
-multi-tenant-max-merged-response-bytes0 (uses 32 MiB)Bytes of a merged multi-tenant response; more returns 413 naming the flag

Observability and Admin Surfaces​

BREAKING in v1.56.0 β€” admin/debug endpoints and /metrics no longer share the main proxy listener by default.

  • -server.register-instrumentation now defaults to false (was true). The main proxy listener stops serving /metrics unless instrumentation is explicitly re-enabled.
  • Admin and debug endpoints (/admin/*, /debug/pprof/*, /debug/queries) move to a dedicated loopback listener on 127.0.0.1:3101 (configurable via --admin-listen) unless -server.admin-auth-token is set, in which case they stay on the main listener for back-compat.
  • /metrics can be hosted on a dedicated port via --metrics-listen=:9091; the Helm chart sets this by default and points the ServiceMonitor at it.

Binary users who scrape :3100/metrics today must, after upgrade, either:

  • set -server.register-instrumentation=true to keep /metrics on the main listener (legacy behavior), OR
  • set -server.register-instrumentation=true -metrics-listen=:9091 for a dedicated metrics port (recommended).
FlagEnvDefaultDescription
-server.register-instrumentationβ€”falseRegister /metrics and related instrumentation handlers. BREAKING in v1.56.0 β€” default changed from true to false so the main listener no longer exposes /metrics unless opted in. Pair with --metrics-listen to host /metrics on a dedicated port.
--admin-listenβ€”127.0.0.1:3101Address of the dedicated admin/debug listener used when -server.admin-auth-token is empty. Bound to loopback by default so admin surfaces (/admin/*, /debug/pprof/*, /debug/queries) are unreachable from off-host. Set -server.admin-auth-token to keep these surfaces on the main listener instead.
--metrics-listenβ€”β€”Address of the dedicated /metrics listener (e.g. :9091). When set together with -server.register-instrumentation=true, /metrics moves off the main proxy listener so scrape traffic never competes with queries. Empty disables the dedicated listener.
-server.enable-pprofβ€”falseExpose /debug/pprof/*
-server.enable-query-analyticsβ€”falseExpose /debug/queries
-server.admin-auth-tokenβ€”β€”Token required on admin/debug endpoints (X-Admin-Token header or Authorization: Bearer <token>). When set, admin/debug routes are served on the main listener instead of --admin-listen; without it, enabling admin/debug routes on a non-loopback admin address fails startup
-server.metrics-max-concurrencyβ€”1Maximum concurrent /metrics scrapes served at once (0 disables the cap)
-metrics.max-tenantsβ€”256Max unique tenant labels retained in /metrics before using __overflow__
-metrics.max-clientsβ€”256Max unique client labels retained in /metrics before using __overflow__
-metrics.export-sensitive-labelsβ€”falseExport per-tenant and per-client identity metrics on /metrics and OTLP
-metrics.trust-proxy-headersβ€”falseTrust user/proxy headers (X-Grafana-User, X-Forwarded-User, X-Webauth-User, X-Auth-Request-User, X-Forwarded-*) for client metrics/log attribution and backend context forwarding
-debug-log-raw-queriesβ€”falseInclude raw LogQL/LogsQL strings and backend query params in debug-level logs. By default each query is logged as sha256:<8hex>+len=<n> to keep PII out of operator log pipelines. Set to true only for ad-hoc investigation; INFO-level access logs are unaffected.

Peer Cache​

FlagEnvDefaultDescription
-peer-selfβ€”β€”This instance address used for peer-cache ownership and fetches
-peer-self-azβ€”β€”This instance's availability zone (e.g. us-east-1a). When set, same-AZ peers are preferred when fetching keys at startup and warmup to reduce cross-AZ traffic. Falls back to any peer with fresh data when no same-AZ peer has the key. Helm chart can inject this automatically from the pod's topology label via the Kubernetes Downward API (see peerCache.topologyLabel).
-peer-discoveryβ€”β€”Peer discovery mode: dns, srv, http, or static
-peer-dnsβ€”β€”Headless service DNS name used when -peer-discovery=dns (e.g., proxy-headless.ns.svc.cluster.local)
-peer-srvβ€”β€”Full SRV record name used when -peer-discovery=srv (e.g., _loki-vl-proxy._tcp.proxy-headless.ns.svc.cluster.local). Port is read from the SRV record β€” no separate port flag needed.
-peer-http-urlβ€”β€”URL returning a JSON peer list when -peer-discovery=http. Supported formats: simple array, {"peers":[...]}, Prometheus HTTP SD, Consul catalog. The URL is polled every DiscoveryInterval.
-peer-staticβ€”β€”Comma-separated peer list used when -peer-discovery=static (e.g., 10.0.0.1:3100,10.0.0.2:3100)
-peer-timeoutβ€”2sTimeout applied to peer-cache owner fetches (/_cache/get) before falling back locally
-peer-auth-tokenβ€”β€”Shared token used on /_cache/get and /_cache/set peer-cache requests. Strongly recommended for fleets so peer auth does not depend only on transient discovery/IP membership during startup. Required by default in v1.56.0 β€” the proxy refuses to start when peer cache is configured (via -peer-discovery or -peer-static) and this token is empty, unless -peer-insecure-ip-allowlist=true is set.
-peer-insecure-ip-allowlistβ€”falseExplicit opt-in for the legacy IP-allowlist-only peer auth (membership in the discovered peer set is the only check). When true, the proxy boots with -peer-auth-token="" and falls back to source-IP membership for /_cache/get and /_cache/set. BREAKING in v1.56.0 β€” without this flag, a configured peer cache with an empty token now fails startup.
-peer-write-throughβ€”truePush eligible non-owner cache writes to the owner peer (/_cache/set) to keep owner shards warm under skewed traffic
-peer-write-through-min-ttlβ€”30sMinimum TTL required to push a write-through copy to the owner peer. Also raises the TTL of empty label-list answers: max(30s, -disk-cache-min-ttl, this)
-peer-hot-read-ahead-enabledβ€”falseEnable bounded periodic hot-read-ahead prefetch from peer hot indexes
-peer-hot-read-ahead-intervalβ€”30sBase interval for hot-index pull cycles
-peer-hot-read-ahead-jitterβ€”5sRandom jitter added to the read-ahead interval
-peer-hot-read-ahead-top-nβ€”256Number of hot keys requested per peer hot-index pull
-peer-hot-read-ahead-max-keys-per-intervalβ€”64Max prefetched keys per read-ahead cycle
-peer-hot-read-ahead-max-bytes-per-intervalβ€”8388608Max prefetched bytes per read-ahead cycle
-peer-hot-read-ahead-max-concurrencyβ€”4Max concurrent hot-index and prefetch peer requests
-peer-hot-read-ahead-min-ttlβ€”30sMinimum remaining TTL required for a prefetch candidate
-peer-hot-read-ahead-max-object-bytesβ€”262144Max object size eligible for read-ahead prefetch
-peer-hot-read-ahead-tenant-fair-shareβ€”50Per-tenant first-pass selection cap (% of key budget)
-peer-hot-read-ahead-error-backoffβ€”15sBase cooldown after read-ahead/index pull failures
-warmup-max-jitterβ€”0 (disabled)Maximum random delay before label cache warmup starts on startup. Spreads fleet instances across a random window so they don't all hit VL simultaneously. Recommended: 5s for ≀5 pods, 10s for ≀15 pods, 20s for ≀30 pods, 30s for larger fleets.

Peer-cache notes:

  • the Helm chart manages -peer-self, -peer-discovery, and the mode-specific discovery flag automatically when peerCache.enabled=true (peerCache.discovery = dns, srv with peerCache.srvName, http with peerCache.httpURL, or static with peerCache.peers)
  • AZ-aware peer selection (-peer-self-az) prefers same-AZ peers for key fetches; Helm auto-injects it from metadata.labels['topology.kubernetes.io/zone'] via Downward API when peerCache.topologyLabel is set (default). Set podLabels: {topology.kubernetes.io/zone: <zone>} or rely on a platform webhook/node label syncer to populate the pod label.
  • for http discovery with a Prometheus HTTP SD endpoint, per-peer AZ is read automatically from labels.az or labels.availability_zone in the SD response β€” no extra flag required
  • GET /_cache/peers returns the current peer ring as {"peers":[...],"self":"...","count":N} β€” useful to verify discovery is working; like every /_cache/* endpoint it requires the X-Peer-Token header unless -peer-insecure-ip-allowlist=true
  • POST /admin/cache/flush (admin listener, requires -server.register-instrumentation=true) purges this instance's caches; add ?peers=1 to also send POST /_cache/purge to every peer in the ring (each peer purges only itself)
  • peer-cache fetches preserve owner TTL and can compress larger /_cache/get responses with zstd or gzip
  • loki_vl_proxy_peer_cache_error_reason_total{reason=...} breaks opaque peer fetch failures into low-cardinality reasons like timeout, transport, status_502, body_read, and decode
  • with -peer-write-through=true (default), non-owner writes with TTL above threshold are pushed to owners and stored locally as short-lived shadows to reduce hot-pod disk skew
  • set -peer-auth-token fleet-wide whenever peer cache is enabled; it avoids transient startup or discovery-flap 403 responses caused by IP-membership-only peer auth
  • when -peer-auth-token is set, all peers must share the same token or peer-cache reuse will fail closed
  • /_cache/has?keys=k1,k2,... is a batch key-presence endpoint (metadata only, no value data transferred) used internally during startup warmup; callers can also use it to pick the freshest peer for a given set of keys before calling /_cache/get
  • startup warmup uses a two-phase peer-first strategy: one /_cache/has per peer to discover who has which label windows, then one /_cache/get per covered window from the freshest peer β€” only the first instance per window hits VL; see Fleet Cache Architecture

Current Tuning For Higher Fleet Reuse​

Use these knobs first:

  • set -warmup-max-jitter for any fleet with β‰₯3 pods to prevent thundering herd on rolling restart (see sizing table in Fleet Cache Architecture)
  • keep -peer-write-through=true (default) to warm owner shards under skewed traffic
  • tune -peer-write-through-min-ttl so only stable/hot entries are replicated
  • keep -response-compression=gzip for explicit Loki/Grafana-safe frontend behavior, or auto if you want the same gzip behavior through the legacy default
  • keep -response-compression-min-bytes around 1-4KiB to avoid wasting CPU on small metadata/control responses
  • keep -backend-compression=auto for smart upstream negotiation: loopback backends get identity (no decompression overhead), remote backends get zstd/gzip
  • keep query-range-windowing enabled with long history TTL and near-now freshness controls for mixed historical/live workloads

Bounded Hot Read-Ahead​

A bounded peer hot-read-ahead mode is implemented and can be enabled with:

-peer-hot-read-ahead-enabled=true

It keeps traffic bounded via key/byte/concurrency caps, jitter, tenant fairness, and error backoff. See Fleet Cache Architecture.

Peer Cache Token Management​

The chart auto-generates a Secret named <release>-peer-auth on first install when peerCache.enabled=true and neither peerCache.authToken nor peerCache.existingSecret is set. The Secret holds a single key token populated with a random value (randAlphaNum 32) and is reused on upgrade via Helm lookup. It is exposed to the container as the PEER_AUTH_TOKEN environment variable and passed as -peer-auth-token=$(PEER_AUTH_TOKEN). This satisfies the v1.56.0 "token required by default" startup gate without operator action.

To rotate or replace the token:

  • Rotate (chart-managed) β€” delete the Secret and helm upgrade will regenerate it: kubectl delete secret <release>-peer-auth && helm upgrade <release> charts/loki-vl-proxy. All pods must restart to pick up the new value, so do this in a maintenance window or under a RollingUpdate strategy.
  • Externally manage β€” set peerCache.authToken to a value (rendered into <release>-peer-auth-literal) or peerCache.existingSecret pointing at your own Secret (key token by default, override with peerCache.existingSecretKey). The chart will skip auto-generation. Pin the token this way for render-only GitOps flows or restricted RBAC, where lookup cannot read the existing Secret and would generate a new token on every render. Useful for GitOps flows that source secrets from Vault / External Secrets Operator.
  • Disable peer cache β€” set peerCache.enabled=false; no Secret is created and the token gate is skipped because no peer cache is configured.

To restore the pre-v1.56.0 IP-allowlist-only behavior without a token, pass -peer-insecure-ip-allowlist=true via extraArgs. This is not recommended: peer auth then depends only on source-IP membership in the discovered peer set, which is brittle under discovery flap and pod IP churn.

The owner hot-index endpoint (/_cache/hot) is internal peer-cache surface area and should not be exposed publicly.

Grafana Datasource Mapping​

These Grafana Loki datasource settings now have a direct proxy-side mapping:

Grafana SettingProxy Support
URL-listen, optional -tls-cert-file / -tls-key-file
No AuthenticationSupported by default
TLS Client Authentication-tls-client-ca-file, -tls-require-client-cert
Skip TLS certificate validationGrafana-side only
Add self-signed certificateGrafana-side only
HTTP headers (X-Scope-OrgID, auth headers)-forward-headers, -tenant-map, -auth.enabled
Forward Authorization specifically-forward-authorization=true (or -forward-headers=Authorization)
Allowed cookies-forward-cookies
Timeout-backend-timeout for proxy→VL, Grafana datasource timeout for Grafana→proxy
Maximum lines-max-lines

When -metrics.trust-proxy-headers=true, the proxy forwards trusted user headers and proxy-chain headers to the backend, plus derived context headers X-Loki-VL-Client-ID / X-Loki-VL-Client-Source.

Datasource auth credentials are forwarded separately as X-Loki-VL-Auth-User / X-Loki-VL-Auth-Source and are not used as enduser.id client identity.

To attribute requests to real Grafana users instead of datasource service credentials:

  1. Enable trusted proxy headers on Loki-VL-proxy:
-metrics.trust-proxy-headers=true
  1. Configure Grafana server dataproxy to forward the logged-in user header:
[dataproxy]
send_user_header = true
  1. Keep Loki datasource in proxy mode (access: proxy) and point it to Loki-VL-proxy:
datasources:
- name: Loki (VL Proxy)
type: loki
access: proxy
url: http://loki-vl-proxy:3100
  1. If Grafana sits behind an auth/reverse proxy, preserve user/proxy-chain headers end to end (X-Forwarded-User, X-Grafana-User, X-Forwarded-*, Forwarded).

Expected request-log behavior with this setup:

  • enduser.id reflects the trusted Grafana/auth user.
  • enduser.source indicates user-header source (for example grafana_user or forwarded_user).
  • auth.principal reflects datasource auth identity when present (for example basic-auth datasource user).

Alerting datasource integration is still partial: the proxy supports legacy Loki YAML rules reads and Prometheus-style JSON rules/alerts reads against a configured backend, but it does not yet implement the full Loki ruler write API surface.