Skip to main content

Generated by go run ./cmd/configdoc from the flags in cmd/proxy/main.go and the registry in internal/config. Do not edit by hand.

Errors and alerts index

What a client error or a firing alert means, which limit produced it, and how to change that limit.

Errors​

ResponseLimit that produced itWhat to do
400 binary expression exceeds its evaluation budget-binary-metric-max-arraysRaise it together with -binary-metric-max-operand-bytes; both bound one request's peak allocation. Default 2000000.
400 binary expression exceeds its evaluation budget-binary-metric-max-operand-bytesRaise it for wide binary expressions over many series; the bytes are held in proxy memory for the request. Default 268435456.
400 max entries limit per query exceeded, limit > max_entries_limit_per_query (L > N) (Loki's text); with -max-entries-limit-per-query-cap the limit is lowered to N instead-max-entries-limit-per-queryRaise it for clients that ask for more log lines per request; lower it towards Loki's 5000 to bound response size. Grafana asks for its Max lines setting (1000 by default), well below the default. Default 10000.
400 maximum number of series (N) reached for a single query; consider reducing query cardinality ... (Loki's text); a Logs Drilldown request gets 200 with the busiest N series and the warning maximum number of series (N) reached for a single query; returning partial results-max-stats-query-series500 matches Loki's default and the Drilldown series cap. Raise it for dashboards that legitimately render thousands of series, at the cost of response size and browser rendering; topk and bottomk rank every series regardless of this cap. Raise -backend-max-buffered-response-bytes with it: this cap is what usually keeps a result inside the byte budget, so raising it alone converts Loki's series-limit error into manual metric response exceeds N bytes. Budget about 40 bytes per sample plus about 100 bytes per series - series x steps x 40 bytes - so 64 MiB holds roughly 1,500 dense series over a 1,000-step range. On the e2e stack sum by (pod) (rate({service_name="api-gateway"}[5m])) over 7d covers about 240,000 distinct pods: at the 500 default it answers 491 series in 0.5s, and at 1,000,000 it exceeds 64 MiB and returns 502. Default 0.
400 multi-tenant fanout exceeds limit of N tenants; raise -multi-tenant-max-fanout or query fewer tenants-multi-tenant-max-fanoutEach tenant in the list is a separate backend query, so this multiplies every other limit. Raise it only with matching VictoriaLogs capacity. Default 64.
400 query exceeds max length (N > M); raise -max-query-length-bytes-max-query-length-bytesLower it to reject generated queries earlier; raising it above Loki's own limit has no effect, because Loki's parse error fires first. Default 131072.
400 the query time range exceeds the limit (query length: X, limit: Y)-default-max-query-lengthSet it to the retention you want users to query; per-tenant limits in -tenant-limits override it. Default 0.
413 multi-tenant merged response exceeds configured safety limit; raise -multi-tenant-max-merged-response-bytes or query fewer tenants-multi-tenant-max-merged-response-bytesThe merged body is held in proxy memory once per request; size it with -multi-tenant-max-fanout and replica memory. Default 33554432.
429 too many outstanding requests: heavy VictoriaLogs queries are limited to -backend-max-concurrent-heavy-queries=N per replica and this query waited -backend-heavy-query-queue-wait=D-backend-max-concurrent-heavy-queriesVictoriaLogs lets each stats pipe of one query use up to 40% of -memory.allowedPercent, so 2 keeps concurrent stats state inside its budget. Raise it with VictoriaLogs memory (the fleet total is replicas x this value); lower it to 1 for a VictoriaLogs under 2 GiB. Default 2.
429 too many outstanding requests: long-range VictoriaLogs metadata scans are limited to at most -backend-max-concurrent-metadata-scans=N per replica (adaptive limit L, floor -backend-min-concurrent-metadata-scans=F[, VictoriaLogs running C of S selects][, VictoriaLogs /metrics not answering][, VictoriaLogs memory P% used with R% reserved by this replica's scans against -backend-metadata-scan-memory-headroom=H]) and this query waited -backend-heavy-query-queue-wait=D; a background inventory refresh that finds the limiter closed is skipped (outcome skipped) and retried on its next schedule instead-backend-max-concurrent-metadata-scansWhat one scan costs depends on the data, so each replica learns the limit instead of being sized: on the e2e stack (7-day retention, 18 million rows) a 7-day stream_field_names scan holds 0.8 GiB and took 9-35 s alone, and seven at once from seven replicas each ran past VictoriaLogs' 60 s limit and OOM-killed it, while a backend with cheap scans lets the limit reach the ceiling. Because others' selects count against it, the limit bounds the fleet, not one replica: a simulation of eleven replicas warming up together stays under the memory budget where a static 2 each kills the backend. Raise the ceiling only for a VictoriaLogs with many cores and memory to spare; lower it to cap how many long scans any replica starts. 0 disables the limiter. Default 8.
429 too many outstanding requests naming both admission flags-backend-heavy-query-queue-waitKeep it below the Grafana data source timeout (30s by default) so a queued panel fails with the documented 429 instead of a client timeout. Raise it when dashboards have more heavy panels than slots; lower it to shed load earlier. Default 20 * time.Second.
500 rpc error: code = ResourceExhausted desc = grpc: trying to send message larger than max (N vs. LIMIT); raise -label-values-max-response-bytes or narrow the query (Loki's querier text above grpc_server_max_send_msg_size, with the flag appended; N is the bytes read when the proxy stopped, or the size of the merged listing when the metadata inventory lists the values from time buckets: the cap bounds the listing, not each bucket response)-label-values-max-response-bytesThe VictoriaLogs response carries a hit count per value, so it is about 1.7x the Loki JSON answer: on the e2e stack pod values for {env="production"} are 2.4 MB from VictoriaLogs (1.4 MB of Loki JSON, 54,405 values) over 1h and grow with the window, about 53 MB for 24h. The 64 MiB default (the same per-response budget as -backend-max-buffered-response-bytes, 16x Loki's default message size) admits about 1.4 million such values, so a day of a high-churn label still answers, while Loki on the same stack already fails the 1h request. While an admitted response is decoded, indexed and encoded, proxy memory holds a small multiple of it (the body, the decoded values and the Loki JSON) per concurrent request. Lower it towards 4 MiB to fail where Loki's defaults would; raise it only with replica memory, since Grafana's label pickers cannot render lists of that size anyway. Default 67108864.
502 manual metric response exceeds N bytes; narrow the query or increase -backend-max-buffered-response-bytes-backend-max-buffered-response-bytesProxy memory grows with this value times the requests that buffer a response at once, so treat it as per-request memory: 64 MiB x -max-concurrent is the worst case. Raise it for very wide stats responses; lower it on small replicas. It is the second half of -max-stats-query-series: the series cap decides how many series a result may carry, this decides whether the encoded result fits. Raising either one alone moves the failure rather than removing it - a high series cap with the default byte budget fails with manual metric response exceeds N bytes instead of the series-limit error. Size it as series x steps x 40 bytes for the widest chart you intend to serve, then check the replica has that much memory per concurrent request. Default 67108864.
502 manual range metric row limit exceeded (N); narrow the query or increase -manual-range-metric-row-limit-manual-range-metric-row-limitProxy memory is about 100 bytes per stats row and about 32 bytes per raw sample, so the default bounds one evaluation to roughly 100 MiB. Raise it for high-cardinality groupings over days; lower it to reject expensive scans earlier. Default 1000000.
502 ordered JSON metric response exceeds N bytes; narrow the query or increase -ordered-json-metric-max-bytes-ordered-json-metric-max-bytesThe default admits about one million rows of 1 KiB. Rows are streamed, so raising it costs VictoriaLogs scan work and network transfer rather than proxy memory; lower it to protect VictoriaLogs from long raw scans. Default 1073741824.
503 too many concurrent queries with Retry-After: 5-max-concurrentSize it with replica CPU and VictoriaLogs capacity; 0 disables admission entirely, which lets one client drive unbounded fanout. Default 100.
the same 429 as -backend-max-concurrent-metadata-scans, with the memory figures in the message; background refreshes skip-backend-metadata-scan-memory-headroom0.4 leaves 40% of VictoriaLogs' memory to its page cache (VictoriaLogs reads its parts through it), ingestion, the queries the proxy does not gate and scans whose memory is still building up; the memory in use already includes the caches VictoriaLogs is allowed (-memory.allowedPercent). On the e2e stack (8 GiB container, 3 GiB caches) 0.3 let memory in use reach 5.4 GiB and the resident set 5.7 GiB with eleven replicas starting together; the container also holds about a gigabyte of page cache. Raise it (0.5) when VictoriaLogs also serves other clients or runs close to its limit; lower it (0.2) on a dedicated node with a large limit. 0 disables the memory gate and the brake and leaves the latency, failure and select-slot feedback. A backend whose /metrics has no vl_concurrent_select series (a vmauth in front of VictoriaLogs) turns the memory gate off by itself: that process's memory says nothing about VictoriaLogs. Default 0.4.

Alerts​

AlertLimits that can fire it
LokiVLProxyHighClientErrors-max-stats-query-series, -max-entries-limit-per-query, -max-query-length-bytes, -default-max-query-length, -multi-tenant-max-fanout
LokiVLProxyHighErrorRate-max-concurrent
LokiVLProxyHighErrorRate (any sustained 5xx)-manual-range-metric-row-limit

Alerts without a limit in this index (backend reachability, latency, memory, cache hit rate) report the health of the proxy or of VictoriaLogs, not a configured bound. See alerting.