Generated by
go run ./cmd/configdocfrom the flags incmd/proxy/main.goand the registry ininternal/config. Do not edit by hand.
Limits registry
Every operator-tunable bound on the work one request may do. A limit is never a silent truncation: hitting one returns an error that names the flag, so an operator can raise it. Zero selects the built-in default of each flag; startup rejects negative values.
Fleet arithmetic: per-replica limits multiply by the replica count, and every byte limit multiplies by the requests that hold a response at once (-max-concurrent).
| Limit | Flag | Helm value | Default | Unit | Bounds | Error when hit | Metric | Alert | Per tenant | Loki parity |
|---|---|---|---|---|---|---|---|---|---|---|
| backend max concurrent heavy queries | -backend-max-concurrent-heavy-queries | extraArgs.backend-max-concurrent-heavy-queries | 2 | concurrent calls | heavy VictoriaLogs calls running at once per replica: raw-row metric fetches above 10,000 rows, and stats, hits or unbounded query calls spanning at least -backend-heavy-query-min-range, finer than 11,000 buckets, or without a time range | 429 too many outstanding requests: heavy VictoriaLogs queries are limited to -backend-max-concurrent-heavy-queries=N per replica and this query waited -backend-heavy-query-queue-wait=D | loki_vl_proxy_internal_operation_total{operation="backend_heavy_query_admission"} | none | no | querier.max_concurrent / the query scheduler queue (Loki answers a full queue with the same 429 message) |
| backend max concurrent metadata scans | -backend-max-concurrent-metadata-scans | extraArgs.backend-max-concurrent-metadata-scans | 8 | concurrent calls | the ceiling of the adaptive limit on long-range VictoriaLogs metadata scans running at once per replica: stream_field_names, stream_field_values, field_names, field_values and streams calls spanning at least -backend-heavy-query-min-range or without a time range (the /labels, /label/{name}/values, /series and detected_fields scans), and the day bucket scans of such a listing when it is assembled from the metadata inventory (each holds a slot only while its own call runs; hour buckets are short scans like any short listing). Shorter listings, inventory edges, minute buckets and row counts are never queued. The selects VictoriaLogs runs for others (other replicas, other clients, read from its /metrics) count against the limit, so replicas that never talk to each other hold VictoriaLogs near one replica's limit rather than replicas x limit. The limit starts at 2, grows by 1/limit per scan that used the whole limit within -backend-metadata-scan-latency-tolerance of its no-load duration, shrinks by 20% once per wave of slower scans, halves on a transport failure, timeout or 5xx and on a rise of VictoriaLogs' queue-timeout counter, and never exceeds this value or drops below -backend-min-concurrent-metadata-scans. Independently of the limit, no scan starts while VictoriaLogs runs all its select slots, stops answering /metrics, or lacks the memory headroom (-backend-metadata-scan-memory-headroom) | 429 too many outstanding requests: long-range VictoriaLogs metadata scans are limited to at most -backend-max-concurrent-metadata-scans=N per replica (adaptive limit L, floor -backend-min-concurrent-metadata-scans=F[, VictoriaLogs running C of S selects][, VictoriaLogs /metrics not answering][, VictoriaLogs memory P% used with R% reserved by this replica's scans against -backend-metadata-scan-memory-headroom=H]) and this query waited -backend-heavy-query-queue-wait=D; a background inventory refresh that finds the limiter closed is skipped (outcome skipped) and retried on its next schedule instead | loki_vl_proxy_internal_operation_total{operation="backend_metadata_scan_admission"} | none | no | querier.max_concurrent / the query scheduler queue (Loki answers a full queue with the same 429 message); Loki serves /labels from its index in milliseconds, VictoriaLogs scans every row in the range |
| backend min concurrent metadata scans | -backend-min-concurrent-metadata-scans | extraArgs.backend-min-concurrent-metadata-scans | 1 | concurrent calls | the floor of the adaptive limit on long-range metadata scans per replica: latency and failure feedback never shrink the limit below it, and below it a replica may start a scan while other clients' long work uses up to a quarter of VictoriaLogs' select slots (the fleet share otherwise counts it against the limit). VictoriaLogs' full select slots, its memory headroom, a /metrics that stopped answering (for up to 10 s), and for a listing never measured any other select running, still hold scans back | none directly; it bounds how far -backend-max-concurrent-metadata-scans may adapt down | loki_vl_proxy_internal_operation_total{operation="backend_metadata_scan_admission"} | none | no | proxy-specific |
| backend metadata scan memory headroom | -backend-metadata-scan-memory-headroom | extraArgs.backend-metadata-scan-memory-headroom | 0.4 | fraction | the share of the memory available to VictoriaLogs that long-range metadata scans must leave free. The proxy reads VictoriaLogs' /metrics (process_resident_memory_anon_bytes, or process_resident_memory_bytes, vm_available_memory_bytes, vl_concurrent_select_current and _capacity; about 40 KB, at most every 100 ms while scans are admitted and every 500 ms while they run) and admits a scan only while the memory in use (the anonymous resident set minus the Go heap VictoriaLogs has freed but not returned, go_memstats_heap_idle_bytes - go_memstats_heap_released_bytes), plus what this replica's scans in flight reserved and have not yet shown in it, plus the selects VictoriaLogs runs for others counted at this scan's cost (another replica's scan that started a moment ago holds none of its memory yet), plus the scan's own cost stays below 1 minus this fraction of the available memory. The cost is learned per endpoint, range bucket and tenant from the peak growth of the memory in use while a scan ran. A listing a tenant has not scanned over this range borrows the same listing's cost over another range (scaled up from a shorter one) or twice another tenant's; a listing nobody has measured runs alone: only while VictoriaLogs runs no other select and this replica no other unmeasured scan. The only scan on an idle VictoriaLogs may use the headroom up to the brake mark. On a VictoriaLogs where some select always runs, a synchronous request whose listing has no cost measured alone goes ahead after waiting 2-4 s, once memory in use held steady for 5 s and no select started during its wait, reserving more than half the room left, so such escapes go one at a time. Behind the estimates is a brake: when memory in use passes a third of the way into the headroom (73% of the available memory at 0.4) a replica stops its youngest scan, one every 500 ms while it stays above, and that request gets the 429; what the stopped scan had grown to becomes its class's cost, so the class is not started again for an hour unless it fits below the mark | the same 429 as -backend-max-concurrent-metadata-scans, with the memory figures in the message; background refreshes skip | loki_vl_proxy_internal_operation_total{operation="backend_metadata_scan_admission"} | none | no | proxy-specific |
| backend metadata scan latency tolerance | -backend-metadata-scan-latency-tolerance | extraArgs.backend-metadata-scan-latency-tolerance | 1.5 | ratio | how many times its no-load duration a long-range metadata scan that ran beside others may take before the adaptive limit shrinks by 20%. The no-load duration is tracked per endpoint, range bucket and tenant the way Vegas tracks the base round-trip time: a faster scan lowers it at once, slower ones raise it slowly, so it follows data growth; it is kept per row when the scan's rows are known (counted inventory buckets), so a busy day is not compared with a quiet one | none directly; it decides when -backend-max-concurrent-metadata-scans adapts down | loki_vl_proxy_internal_operation_total{operation="backend_metadata_scan_admission"} | none | no | proxy-specific |
| metadata inventory parallelism | -metadata-inventory-parallelism | extraArgs.metadata-inventory-parallelism | 4 | concurrent calls | bucket listings one /labels, /label/{name}/values or field-name listing may have in flight while it is assembled from the time-bucketed metadata inventory (UTC-aligned day, hour, 5-minute and minute buckets cached in memory, on disk and on the owner peer; the last minute is never cached; a missing hour or day is merged from its cached children first; hour and day buckets of a query made of stream and field filters are revalidated by a row count instead of a rescan). The answer is the answer of one full-range call: values are unioned, hits summed and ordered like VictoriaLogs; a listing with a limit, or a query that is not row-local (stats, uniq, limit, sort, subqueries, _time filters), is one full-range call | none; a failed bucket fails the listing (the endpoint then serves its last known-good answer or the error), never a partial list | loki_vl_proxy_internal_operation_total{operation="metadata_inventory"} (outcome cached, revalidated, partial or scanned) | none | no | proxy-specific |
| backend heavy query queue wait | -backend-heavy-query-queue-wait | extraArgs.backend-heavy-query-queue-wait | 20 * time.Second | duration | how long the heavy calls of one request wait, together, for heavy-query slots, and how long each long-range metadata scan of a request (one listing, or one day bucket of the label inventory, all of a request's bucket scans together at most twice this) waits for a metadata-scan slot; the adaptive metadata limiter re-reads VictoriaLogs' /metrics every 250 ms while a scan waits, and the request context bounds the whole | 429 too many outstanding requests naming both admission flags | loki_vl_proxy_internal_operation_duration_seconds{operation="backend_heavy_query_admission"} | none | no | proxy-specific |
| backend heavy query min range | -backend-heavy-query-min-range | extraArgs.backend-heavy-query-min-range | 6 * time.Hour | duration | the time range from which stats, hits and unbounded raw calls count as heavy, and metadata listings count as long-range scans | none directly; it selects which calls the two admission limiters govern | loki_vl_proxy_internal_operation_total{operation="backend_heavy_query_admission"} | none | no | proxy-specific |
| manual range metric row limit | -manual-range-metric-row-limit | extraArgs.manual-range-metric-row-limit | 1000000 | rows | rows one proxy-side range-metric evaluation reads: raw log rows, or stats rows (one per non-empty step bucket and series) per bucket grid on the window stats path | 502 manual range metric row limit exceeded (N); narrow the query or increase -manual-range-metric-row-limit | loki_vl_proxy_requests_total{status="502"} | LokiVLProxyHighErrorRate (any sustained 5xx) | no | proxy-specific |
| max stats query series | -max-stats-query-series | extraArgs.max-stats-query-series | 0 | series | series in the result of a metric query, on every metric route, range and instant; per tenant as max_query_series in -tenant-limits and -tenant-default-limits, published on /loki/api/v1/drilldown-limits and /config/tenant/v1/limits | 400 maximum number of series (N) reached for a single query; consider reducing query cardinality ... (Loki's text); a Logs Drilldown request gets 200 with the busiest N series and the warning maximum number of series (N) reached for a single query; returning partial results | loki_vl_proxy_requests_total{status="400"} | LokiVLProxyHighClientErrors | yes (-tenant-limits) | max_query_series (500); a multi-tenant request is held to the smallest of its tenants |
| ordered json metric max bytes | -ordered-json-metric-max-bytes | extraArgs.ordered-json-metric-max-bytes | 1073741824 | bytes | bytes the ordered JSON metric evaluator reads from the raw rows response, and the largest response it builds | 502 ordered JSON metric response exceeds N bytes; narrow the query or increase -ordered-json-metric-max-bytes | loki_vl_proxy_requests_total{status="502"} | none | no | proxy-specific |
| label filter refill max pages | -label-filter-refill-max-pages | extraArgs.label-filter-refill-max-pages | 8 | requests | further VictoriaLogs pages of at most limit rows a log query reads, in the Loki-compatible profile, when a Loki label filter on a log line key no earlier stage exposes dropped rows VictoriaLogs matched: a response reads at most (1 + N) x limit rows | none: the page returns fewer lines than the limit when the budget is spent; 0 reads no further page | none | none | no | proxy-specific |
| backend max buffered response bytes | -backend-max-buffered-response-bytes | extraArgs.backend-max-buffered-response-bytes | 67108864 | bytes | bytes read from one VictoriaLogs response the proxy evaluates itself: buffered stats, volume and binary-operand responses, and the encoded metric result | 502 manual metric response exceeds N bytes; narrow the query or increase -backend-max-buffered-response-bytes | loki_vl_proxy_requests_total{status="502"} | none | no | proxy-specific |
| max entries limit per query | -max-entries-limit-per-query | extraArgs.max-entries-limit-per-query | 10000 | rows | log lines a log query may ask for (Loki's max_entries_limit_per_query, per tenant in -tenant-limits and -tenant-default-limits, where 0 is unlimited); label values requests above it are capped | 400 max entries limit per query exceeded, limit > max_entries_limit_per_query (L > N) (Loki's text); with -max-entries-limit-per-query-cap the limit is lowered to N instead | loki_vl_proxy_requests_total{status="400"} | LokiVLProxyHighClientErrors | yes (-tenant-limits) | max_entries_limit_per_query (Loki's default is 5000; the proxy keeps 10000); a multi-tenant request is held to the smallest non-zero value of its tenants |
| label values max response bytes | -label-values-max-response-bytes | extraArgs.label-values-max-response-bytes | 67108864 | bytes | bytes of the VictoriaLogs answer to a /loki/api/v1/label/{name}/values request: one response, or the size the merged listing would have as one response when the metadata inventory lists it from time buckets; the read stops one byte past the limit, and a rejected response is neither cached nor indexed; per tenant as label_values_max_response_bytes in -tenant-limits and -tenant-default-limits (not published: Loki has no per-tenant equivalent) | 500 rpc error: code = ResourceExhausted desc = grpc: trying to send message larger than max (N vs. LIMIT); raise -label-values-max-response-bytes or narrow the query (Loki's querier text above grpc_server_max_send_msg_size, with the flag appended; N is the bytes read when the proxy stopped, or the size of the merged listing when the metadata inventory lists the values from time buckets: the cap bounds the listing, not each bucket response) | loki_vl_proxy_internal_operation_total{operation="label_values_response_cap"} | none | yes (-tenant-limits) | grpc_server_max_send_msg_size (dskit default 4 MiB): Loki's querier fails a larger label values response with 500 ResourceExhausted; Loki has no per-tenant limit on label values count |
| max query length bytes | -max-query-length-bytes | extraArgs.max-query-length-bytes | 131072 | bytes | the LogQL query string length | 400 query exceeds max length (N > M); raise -max-query-length-bytes | loki_vl_proxy_requests_total{status="400"} | LokiVLProxyHighClientErrors | no | syntax.maxInputSize (131072): the default matches it, so the proxy rejects only what Loki rejects |
| default max query length | -default-max-query-length | extraArgs.default-max-query-length | 0 | duration | the query time range accepted for every tenant unless a per-tenant limit overrides it | 400 the query time range exceeds the limit (query length: X, limit: Y) | loki_vl_proxy_requests_total{status="400"} | LokiVLProxyHighClientErrors | yes (-tenant-limits) | max_query_length (Loki's default is 721h; the proxy keeps 0, unlimited); a multi-tenant request is held to the smallest non-zero value of its tenants |
| multi tenant max fanout | -multi-tenant-max-fanout | extraArgs.multi-tenant-max-fanout | 64 | tenants | tenants one multi-tenant request may fan out to | 400 multi-tenant fanout exceeds limit of N tenants; raise -multi-tenant-max-fanout or query fewer tenants | loki_vl_proxy_requests_total{status="400"} | LokiVLProxyHighClientErrors | no | proxy-specific |
| multi tenant max merged response bytes | -multi-tenant-max-merged-response-bytes | extraArgs.multi-tenant-max-merged-response-bytes | 33554432 | bytes | the merged response of a multi-tenant request | 413 multi-tenant merged response exceeds configured safety limit; raise -multi-tenant-max-merged-response-bytes or query fewer tenants | loki_vl_proxy_requests_total{status="413"} | none | no | proxy-specific |
| binary metric max operand bytes | -binary-metric-max-operand-bytes | extraArgs.binary-metric-max-operand-bytes | 268435456 | bytes | operand bytes one binary metric expression may capture | 400 binary expression exceeds its evaluation budget | loki_vl_proxy_requests_total{status="400"} | none | no | proxy-specific |
| binary metric max arrays | -binary-metric-max-arrays | extraArgs.binary-metric-max-arrays | 2000000 | arrays | JSON arrays one binary metric expression may allocate while joining operands | 400 binary expression exceeds its evaluation budget | loki_vl_proxy_requests_total{status="400"} | none | no | proxy-specific |
| detected fields max scan lines | -detected-fields-max-scan-lines | extraArgs.detected-fields-max-scan-lines | 2000 | rows | log lines the detected_fields and detected_field values scan reads per request | none: the scan stops at the limit and returns what it found | none | none | no | proxy-specific |
| patterns max backend rows | -patterns-max-backend-rows | extraArgs.patterns-max-backend-rows | 20000 | rows | log lines /patterns reads from VictoriaLogs for one request | none: pattern mining works on the rows it read | loki_vl_proxy_patterns_lines_scanned_total | none | no | proxy-specific |
| patterns second pass max rows | -patterns-second-pass-max-rows | extraArgs.patterns-second-pass-max-rows | 8000 | rows | log lines the /patterns second pass reads when the first pass mined too few patterns | none | loki_vl_proxy_patterns_lines_scanned_total | none | no | proxy-specific |
| patterns second pass max windows | -patterns-second-pass-max-windows | extraArgs.patterns-second-pass-max-windows | 8 | windows | windows the /patterns second pass re-reads | none | none | none | no | proxy-specific |
| max zero fill buckets | -max-zero-fill-buckets | extraArgs.max-zero-fill-buckets | 32768 | buckets | buckets the proxy zero-fills in a metric response | none: a wider request is served without zero-fill | none | none | no | the 11,000-point resolution limit bounds a Loki response anyway |
| stats query range concurrency | -stats-query-range-concurrency | extraArgs.stats-query-range-concurrency | 0 | concurrent calls | concurrent stats_query_range calls the proxy makes to VictoriaLogs | none: calls wait for a slot | loki_vl_proxy_upstream_requests_total{route="/select/logsql/stats_query_range"} | none | no | proxy-specific |
| max concurrent | -max-concurrent | extraArgs.max-concurrent | 100 | requests | requests admitted per replica, and concurrent backend operations | 503 too many concurrent queries with Retry-After: 5 | loki_vl_proxy_requests_total{status="503"} | LokiVLProxyHighErrorRate | no | proxy-specific |
| max lines | -max-lines | extraArgs.max-lines | 1000 | rows | the default number of log lines per query when the client sends no limit | none | none | none | yes (-tenant-limits) | the Loki data source's Max lines (1000) |
Sizing​
-backend-max-concurrent-heavy-queries​
VictoriaLogs lets each stats pipe of one query use up to 40% of -memory.allowedPercent, so 2 keeps concurrent stats state inside its budget. Raise it with VictoriaLogs memory (the fleet total is replicas x this value); lower it to 1 for a VictoriaLogs under 2 GiB.
-backend-max-concurrent-metadata-scans​
What one scan costs depends on the data, so each replica learns the limit instead of being sized: on the e2e stack (7-day retention, 18 million rows) a 7-day stream_field_names scan holds 0.8 GiB and took 9-35 s alone, and seven at once from seven replicas each ran past VictoriaLogs' 60 s limit and OOM-killed it, while a backend with cheap scans lets the limit reach the ceiling. Because others' selects count against it, the limit bounds the fleet, not one replica: a simulation of eleven replicas warming up together stays under the memory budget where a static 2 each kills the backend. Raise the ceiling only for a VictoriaLogs with many cores and memory to spare; lower it to cap how many long scans any replica starts. 0 disables the limiter.
-backend-min-concurrent-metadata-scans​
1 serves a slow backend serially instead of not at all. Raise it only when a backend that is known to have headroom is being held too low by noisy latency (for example 2 on a VictoriaLogs cluster). Must be at most the ceiling; 0 uses the default.
-backend-metadata-scan-memory-headroom​
0.4 leaves 40% of VictoriaLogs' memory to its page cache (VictoriaLogs reads its parts through it), ingestion, the queries the proxy does not gate and scans whose memory is still building up; the memory in use already includes the caches VictoriaLogs is allowed (-memory.allowedPercent). On the e2e stack (8 GiB container, 3 GiB caches) 0.3 let memory in use reach 5.4 GiB and the resident set 5.7 GiB with eleven replicas starting together; the container also holds about a gigabyte of page cache. Raise it (0.5) when VictoriaLogs also serves other clients or runs close to its limit; lower it (0.2) on a dedicated node with a large limit. 0 disables the memory gate and the brake and leaves the latency, failure and select-slot feedback. A backend whose /metrics has no vl_concurrent_select series (a vmauth in front of VictoriaLogs) turns the memory gate off by itself: that process's memory says nothing about VictoriaLogs.
-backend-metadata-scan-latency-tolerance​
1.5 tolerates the slowdown of a second concurrent scan on a CPU-bound VictoriaLogs (measured 1.45x) and reacts to a third (1.9x). Lower it (1.2) to keep label browsing snappy at the cost of concurrency; raise it (2-3) when scan durations are noisy for other reasons (shared disks, ingestion bursts). Must be >= 1; 0 uses the default.
-metadata-inventory-parallelism​
4: VictoriaLogs already spreads one query over its cores, so more buys little on a cold window (three concurrent 7-day scans took 1.8x as long as one on the e2e stack) while a warm window needs only its two edges. Lower it to 1-2 for a small VictoriaLogs; the day bucket scans of a long listing also take -backend-max-concurrent-metadata-scans slots, each only while its own call runs. 0 turns the inventory off: every listing is one VictoriaLogs call over the whole range, as before.
-backend-heavy-query-queue-wait​
Keep it below the Grafana data source timeout (30s by default) so a queued panel fails with the documented 429 instead of a client timeout. Raise it when dashboards have more heavy panels than slots; lower it to shed load earlier.
-backend-heavy-query-min-range​
Lower it (for example 1h) when short ranges already strain VictoriaLogs; raise it when only multi-day queries need queueing.
-manual-range-metric-row-limit​
Proxy memory is about 100 bytes per stats row and about 32 bytes per raw sample, so the default bounds one evaluation to roughly 100 MiB. Raise it for high-cardinality groupings over days; lower it to reject expensive scans earlier.
-max-stats-query-series​
500 matches Loki's default and the Drilldown series cap. Raise it for dashboards that legitimately render thousands of series, at the cost of response size and browser rendering; topk and bottomk rank every series regardless of this cap. Raise -backend-max-buffered-response-bytes with it: this cap is what usually keeps a result inside the byte budget, so raising it alone converts Loki's series-limit error into manual metric response exceeds N bytes. Budget about 40 bytes per sample plus about 100 bytes per series - series x steps x 40 bytes - so 64 MiB holds roughly 1,500 dense series over a 1,000-step range. On the e2e stack sum by (pod) (rate({service_name="api-gateway"}[5m])) over 7d covers about 240,000 distinct pods: at the 500 default it answers 491 series in 0.5s, and at 1,000,000 it exceeds 64 MiB and returns 502.
-ordered-json-metric-max-bytes​
The default admits about one million rows of 1 KiB. Rows are streamed, so raising it costs VictoriaLogs scan work and network transfer rather than proxy memory; lower it to protect VictoriaLogs from long raw scans.
-label-filter-refill-max-pages​
Each page is one more VictoriaLogs request of at most limit rows. The default of 8 fills a page unless the newest 9 x limit rows of the window are almost all JSON lines holding the filtered key, the case Loki answers with fewer lines too. Raise it for streams where structured metadata named like a JSON key is rare; lower it, or set 0, to bound VictoriaLogs reads to one page per request.
-backend-max-buffered-response-bytes​
Proxy memory grows with this value times the requests that buffer a response at once, so treat it as per-request memory: 64 MiB x -max-concurrent is the worst case. Raise it for very wide stats responses; lower it on small replicas. It is the second half of -max-stats-query-series: the series cap decides how many series a result may carry, this decides whether the encoded result fits. Raising either one alone moves the failure rather than removing it - a high series cap with the default byte budget fails with manual metric response exceeds N bytes instead of the series-limit error. Size it as series x steps x 40 bytes for the widest chart you intend to serve, then check the replica has that much memory per concurrent request.
-max-entries-limit-per-query​
Raise it for clients that ask for more log lines per request; lower it towards Loki's 5000 to bound response size. Grafana asks for its Max lines setting (1000 by default), well below the default.
-label-values-max-response-bytes​
The VictoriaLogs response carries a hit count per value, so it is about 1.7x the Loki JSON answer: on the e2e stack pod values for {env="production"} are 2.4 MB from VictoriaLogs (1.4 MB of Loki JSON, 54,405 values) over 1h and grow with the window, about 53 MB for 24h. The 64 MiB default (the same per-response budget as -backend-max-buffered-response-bytes, 16x Loki's default message size) admits about 1.4 million such values, so a day of a high-churn label still answers, while Loki on the same stack already fails the 1h request. While an admitted response is decoded, indexed and encoded, proxy memory holds a small multiple of it (the body, the decoded values and the Loki JSON) per concurrent request. Lower it towards 4 MiB to fail where Loki's defaults would; raise it only with replica memory, since Grafana's label pickers cannot render lists of that size anyway.
-max-query-length-bytes​
Lower it to reject generated queries earlier; raising it above Loki's own limit has no effect, because Loki's parse error fires first.
-default-max-query-length​
Set it to the retention you want users to query; per-tenant limits in -tenant-limits override it.
-multi-tenant-max-fanout​
Each tenant in the list is a separate backend query, so this multiplies every other limit. Raise it only with matching VictoriaLogs capacity.
-multi-tenant-max-merged-response-bytes​
The merged body is held in proxy memory once per request; size it with -multi-tenant-max-fanout and replica memory.
-binary-metric-max-operand-bytes​
Raise it for wide binary expressions over many series; the bytes are held in proxy memory for the request.
-binary-metric-max-arrays​
Raise it together with -binary-metric-max-operand-bytes; both bound one request's peak allocation.
-detected-fields-max-scan-lines​
Raise it when Drilldown misses fields that appear late in a window, at the cost of a longer VictoriaLogs scan; -drilldown-scan-timeout still bounds the time.
-patterns-max-backend-rows​
Raise it for sparse selectors where patterns are mined from too few lines; every extra row is VictoriaLogs scan work and proxy parsing.
-patterns-second-pass-max-rows​
Raise it with -patterns-max-backend-rows; the second pass doubles the scan for hard selectors.
-patterns-second-pass-max-windows​
Each window is one VictoriaLogs call; raise it only when pattern coverage over long ranges matters more than fanout.
-max-zero-fill-buckets​
Rarely needs changing; it only bounds proxy-side allocation for zero-filled charts.
-stats-query-range-concurrency​
Drilldown Fields fires about 30 of these at once; 4 keeps VictoriaLogs CPU steady. Heavy calls are additionally bounded by -backend-max-concurrent-heavy-queries.
-max-concurrent​
Size it with replica CPU and VictoriaLogs capacity; 0 disables admission entirely, which lets one client drive unbounded fanout.
-max-lines​
Raise it only with -max-entries-limit-per-query; every line is transferred and rendered.
Worked examples​
A small deployment​
One proxy replica with 512 MiB, one VictoriaLogs with 2 GiB, a handful of Grafana users:
extraArgs:
backend-max-concurrent-heavy-queries: "1" # one heavy stats state fits in VictoriaLogs memory
backend-heavy-query-queue-wait: "25s" # still below Grafana's 30s data source timeout
backend-max-buffered-response-bytes: "16777216" # 16 MiB per buffered response
manual-range-metric-row-limit: "250000" # about 25 MiB of proxy-side samples
max-concurrent: "16"
A large, high-volume deployment​
Three proxy replicas with 4 GiB each, VictoriaLogs with 64 GiB, dashboards over days:
extraArgs:
backend-max-concurrent-heavy-queries: "6" # 18 heavy calls fleet-wide; VictoriaLogs has the memory for them
backend-heavy-query-queue-wait: "20s"
backend-heavy-query-min-range: "12h" # only long ranges queue
backend-max-buffered-response-bytes: "268435456" # 256 MiB for wide stats responses
manual-range-metric-row-limit: "4000000" # high-cardinality groupings over 7d
max-stats-query-series: "5000"
max-concurrent: "256"