Performance tuning¶
Use the timings and counters in each query response to find the slow stage. Change one setting at a time, then run the same query again to check the result.
Tell a metadata fast path from a Parquet scan¶
Every query response carries its own explanation:
curl -s http://localhost:8089/api/v1/sql \
-H 'Content-Type: application/json' \
-d '{"query": "…"}' | jq '{cost, stats}'
| Reading | Meaning |
|---|---|
stats.served_by starts with tier1_ |
A warm metadata fast path served it. |
stats.served_by == "materialized" |
Exact per-file fallback; potentially slow even when rows_scanned == 0. |
stats.served_by == "scan" |
The full query plan ran. |
cost.files_considered ≫ cost.files_to_scan |
Planning-time pruning is working. |
scan.files_planned ≫ scan.files_read |
Execution-time pruning is working. |
phases.plan_micros ≫ collect_micros |
Planning-bound: usually too many files. |
phases.collect_micros dominant |
Scan-bound: decode or I/O. |
scan.bytes_footer dominates scan.fetched_bytes |
Footer-bound: improve metadata-cache reuse or compact toward fewer, larger files. If scan.bytes_data dominates instead, narrowing the projection or improving pruning can help. The reader also charges bytes between requested data ranges that it coalesces into one request. The scan statistics do not expose the split between requested bytes and coalescing overhead, so bytes_data alone cannot identify projection width. |
scan.ordering != "advertised" |
Ordered early-stop refused. The value is the reason. |
phases.distributed.shard_wall_micros uneven |
A straggler shard. |
By symptom¶
Aggregates are scanning¶
If an aggregate is slow, check whether it reports served_by: "materialized"
or returns rows_scanned > 0.
Check stats.served_by first. rows_scanned can be zero for both a
millisecond metadata answer and a much slower per-file fallback.
Repair incomplete group counts¶
served_by: "materialized" is an exact answer, but it was assembled from each
live file's group-count footer or, where necessary, by decoding raw pages. If a
group-by that used to be fast starts taking this path:
- Check the
SiglakeGroupCountDeltaRetryingalert. The commit path tries a delta write four times, waiting 250, 500, and 750 ms between retries. A sustained retry alert points first to object-store health or credentials on the named pod. An exhausted write incrementssiglake_group_count_delta_write_failures_total, records a durable rebuild marker, and does not fireSiglakeGroupCountDeltaLostby itself. That alert names the table only when the subsequent automatic rebuild fails or remains incomplete. - If neither
SiglakeGroupCountDeltaLostnorSiglakeGroupCountAggregateShortis firing, wait for the maintenance compactor's next aggregate fold. It consumes the rebuild marker, reconstructs the table's exact maps and bounded sketches from committed files, and deletes the covered markers. Confirm a successful automatic repair with an increase insiglake_group_count_auto_rebuilds_total{outcome="success"};siglake_group_count_aggregate_rebuilds_totalalso increases after any completed rebuild and is a secondary signal. A later delta does not heal the gap; the marker-driven rebuild does.
siglake_group_count_deltas_absorbed_total records deltas folded into the
base, and siglake_group_count_deltas_deleted_total records covered delta
objects removed by maintenance.
Only a compactor running --role maintenance or --role combined performs
this fold. A fleet whose compactors all run --role drain has nothing to
consume rebuild markers.
If SiglakeGroupCountAggregateShort is firing instead, check its outcome.
The default detected outcome has no lost-delta marker for the fold to
consume. The backoff and suppression outcomes come from a separate durable
automatic-repair history. See
Rebuild a short group-count aggregate with durable backoff.
If SiglakeGroupCountDeltaLost is firing, waiting is over: check whether
its outcome is failed or incomplete, read the compactor log, and use the
operator fallback:
The alert fills in both names: --namespace takes the iceberg_namespace
label and defaults to siglake.
Automatic and operator-triggered rebuilds are safe to run live and to
repeat. They record a
rebuilt_through watermark so a late delta at or below the scanned snapshot
is not folded twice; later deltas continue to fold normally.
3. Handle pre-existing typed columns separately: they are never admitted
automatically. Lost-delta repair restores only the columns named by the
failed commit. To admit eligible typed (long, double, or bool) columns
from a table created before typed columns joined the side aggregate, run:
Admission is all-or-nothing and subject to
SIGLAKE_TYPED_GROUP_COUNT_CARDINALITY; the command reports a column rather
than writing a partial total when the cap is exceeded or a live file cannot
serve it.
The repair deliberately writes only the wide aggregate object, so a repaired
column reports served_by: "tier1_wide",
even when it would fit the inline cap. No data-file rewrite is needed merely to
admit a typed column, nor to move a repaired column into the inline object. A
compaction rewrite is needed only when the rebuild reports that a live file
cannot supply the column from either its footer or a raw-page decode (for
example, the file predates typed footers or lacks the column); rerun the rebuild
after those files have been rewritten. Answers remain exact before and during
repair.
Diagnose a scan¶
served_by: "scan" means the full query plan ran. For a dimensional count,
the matcher may not have recognized the predicate shape. Known limits:
- Float literals are excluded: float rendering isn't reproducible enough to match footer keys.
- The predicate must reduce to footer equality or an integer range.
- The column must have group-count footers, which older files may lack until compaction rewrites them.
Check whether the same shape without the predicate is zero-scan; if so, the
predicate is the problem. SIGLAKE_LOG_EXEC_PLANS=1 shows what the planner
actually built.
Rebuild a short group-count aggregate with durable backoff¶
In Siglake 0.2.0, automatic short group-count repair retries after 15 minutes, one hour and four hours. The compactor suppresses repair after the fourth unsuccessful scan for a table incarnation. A successful automatic or manual rebuild clears applicable attempt history after publishing the aggregate.
SiglakeGroupCountAggregateShort means a maintained column's count is short of
total-records although every commit is accounted for. No later delta closes
the gap. Answers stay exact because GROUP BY uses the per-file path until the
aggregate is rebuilt.
The census runs every 15 minutes on a --role maintenance or --role combined
compactor. SIGLAKE_AGG_SHORT_SCAN_INTERVAL_SECS controls that interval. The
default installation reports the deficit without repairing it. Rebuild the
table yourself with the namespace from the alert's iceberg_namespace label:
--namespace defaults to siglake, so a single-tenant install can leave it
out. A compactor maintaining tenant_* namespaces cannot: each one has its own
events, and the alert says which is short.
For automatic repair, set compactor.shortAggregateRepair to true. It repairs
one table per census by default. Each repair runs one Tier-2 query per maintained
column. A local 400k-row scan took 0.85 seconds per column; the linear estimate
for 250M rows is about nine minutes per column.
Before it scans, the compactor writes an incarnation-scoped attempt marker. A marker write failure prevents the Tier-2 scan. A fresh marker prevents a second compactor from duplicating an active scan. If the scan disappears, the next census records it as interrupted and applies backoff. Recreating the table starts new history.
The 600-second watchdog cancels a long repair before publication and increments
siglake_compactor_watchdog_trips_total{stage="agg_short_repair"}. The alert's
fixed outcome identifies watchdog, failed or interrupted backoff, suppression,
and marker write failure.
Run rebuild-group-counts to recover a suppressed table. After publication,
the command deletes attempt records at or below the rebuild watermark. Its
first output line reports how many it cleared. A failed command preserves
suppression. A later census retries failed deletions. Rebuild large tables by
hand and leave automatic repair off. See the three environment variables in
Compaction.
Ordered browses are slow¶
If ORDER BY timestamp … LIMIT n is slow, check whether scan.ordering is
advertised.
ordering value |
Cause | Fix |
|---|---|---|
filtered |
The residual filter defeated the ordering gate: a non-timestamp predicate or a prune spec. Expected on filtered browses, where a bounded TopK is the cheaper plan. |
Nothing to fix if the query is fast. Otherwise narrow the time window. |
no_bounds |
Files lack manifest time bounds. | Let compaction rewrite them. |
fan_in |
One cluster of files overlaps in time more deeply than the per-partition merge will open. | Fix the compaction layout as described below. |
global_fan_in |
Overlap fits per partition but not the whole-scan stream budget. | Same layout problem; SIGLAKE_ORDERED_MERGE_GLOBAL_FANIN raises the budget. |
mixed_sort_direction, unknown_sort_order, order_walk_failed |
The scan mixes files written under different sort orders, or their orders could not be attributed. The table is part-way through the DESC to ASC convergence. |
Let re-clustering rewrite the stragglers. |
missing_sort_column |
The layout checks passed, but timestamp was unexpectedly absent from the scan's output schema after the earlier projection check. |
This indicates a Siglake bug, not a layout or query-shape problem. Report it rather than re-clustering. |
not_projected, no_sort_order, non_identity_sort, missing_sort_field, non_timestamp_sort |
The query does not project timestamp, or the table is not declared sorted by it. |
Query or table shape, not layout: see Ordered early-stop. |
Raise SIGLAKE_ORDERED_MERGE_MAX_FANIN if partitions legitimately have more
overlapping files than the merge will open, but treat that as a workaround for
a layout problem. Check what the node already allows before raising it: both
caps default to a node-adaptive value, four open streams per CPU clamped to
16 to 64 per partition and eight per CPU clamped to 32 to 128 across the
scan, so a large worker is already well above the small-pod floors of 16 and
32. An explicit value in either variable is used as given, not clamped.
Decide whether table layout is converging or falling behind ingest¶
If query latency rises with table size, check whether
siglake_table_overlap_depth is high or flat. The
leveled-compaction rules
define the file-count, overlap-depth, cadence, and backpressure rules.
Treat this as a compaction problem. It is the most common cause of query latency that rises as a table grows.
- Check
siglake_table_overlap_depth. A converged 2 B-row table settles in the low twenties. This gauge comes from a budgeted manifest walk and can lag on a large table, so verify thatsiglake_table_gauges_sampled_at_secondsis advancing before treating a flat value as current. The timestamp covers the level, leading-edge, and overlap-depth gauges and advances only after a complete walk. - Check
siglake_table_live_data_files: unbounded growth means compaction is losing. This count is exact from the snapshot summary on every sampler cycle, regardless of whether the manifest walk finishes. - Check
siglake_compactor_maintenance_skipped_backpressure_totaland…_throttled_backpressure_total: compaction may be permanently yielding to the drain.
Fixes, in order:
- Add compactor replicas (with
catalogClaim.enabled) so the drain stops starving maintenance. - Keep the default leveled compaction enabled. If
SIGLAKE_COMPACTOR_LEVELED=0is present, remove the override so the default leveled scheduler runs. - Lower
SIGLAKE_COMPACTOR_MAX_OVERLAP_DEPTHto trigger depth merges sooner. - Raise
SIGLAKE_COMPACTOR_BACKPRESSURE_COMPACT_EVERYfrequency (lower the number) so throttled passes run more often under backlog.
Allow 4 to 5 hours of background compaction for a 2 B-row table to converge. In the 1 TB test, convergence halved browse latency and raised 32-way throughput from 28 QPS to 115 QPS.
Queries return 503¶
If queries return 503 with Retry-After: 5, read the response body's
reason first.
shard_pin_unresolved is a refused distributed query. Diagnose it
below.
The rest of this section covers the capacity case, where
siglake_query_breaker_trips_total{breaker="pool_exhausted"} is rising.
- Check
SiglakeQueryPoolNearLimitfor sustained pool pressure. - Request
GET /debug/memory-poolfrom the affected query pod to identify the largest live consumers. - Compare the pod's
query.resources.limits.memorysizing budget with its query concurrency and workload. - If the pool budget is adequate, inspect the spill volume. Sorts and
aggregates spill after the pool tightens, but a query is refused when the
spill directory reaches
query.spill.maxBytes(SIGLAKE_QUERY_SPILL_MAX_BYTES). Check the pod's ephemeral-storage usage, and raise the cap together withquery.spill.sizeLimitand the ephemeral-storage limit in the order Query spill and ephemeral storage describes.
Distributed workers do not reserve admission shares; see Distributed admission is per-coordinator for the worker-side bounds.
This is transient server capacity, not a client rate-limit error. Honor
Retry-After; the same query can fit after concurrent work finishes.
A 503 with reason: shard_pin_unresolved¶
If distributed queries return 503 with
reason: "shard_pin_unresolved"
in the body, siglake_query_shard_pin_total{outcome="miss"} is rising, and
SiglakeQueryShardPinUnresolved fires. Single-node
/api/v1/sql/local requests are unaffected.
A worker could not resolve the generation the coordinator pinned the fan-out to: either the snapshot or the schema id. It refused its shard instead of contributing rows from a different generation. See One generation per fan-out. The error message names which half failed.
Isolated misses clear as caches converge. Sustained schema misses mean a
migrate-schema commit has not reached every replica's metadata cache; the
same cache lifetime governs both halves. Sustained snapshot misses mean
coordinators are serving snapshots the catalog has already expired. Both
remedies below close that gap:
- Raise
compactor.snapshotExpire.retainLastso expiry keeps more snapshots and widens the margin. That section covers the packaged default, what a wider window costs and the other consumers: time travel, long external reads and the mixed-version rollout guard. These constrain the value too. - Lower
SIGLAKE_ICEBERG_METADATA_CACHE_TTL_SECSso coordinators stop serving a snapshot before the catalog expires it. This also narrows the window in which replicas disagree, at the cost of more catalog metadata loads.
A format: "ndjson" query refused after its stream started keeps the 200 and
reports the refusal as a final {"_meta": "error", "code": 503,
"retry_after_secs": 5} line, so a client that only inspects status codes will
read a truncated result as a complete one; see NDJSON
trailers.
A new tenant or index is refused by ingester.maxLanes¶
In Siglake 0.2.0, ingester.maxLanes is a per-process bound on distinct
(tenant, index) pairs. A new pair past the bound receives HTTP 503 or gRPC
Unavailable. The response has no Retry-After, plain gRPC retry-after
metadata or RetryInfo. Confirm the cause with
siglake_ingest_lane_refused_total or the
SiglakeIngestLanesRefused alert. An existing
pair continues to write at the cap.
Size maxLanes for the distinct pairs that one pod can receive between
restarts, with headroom for expected changes. Do not size it from the number of
non-empty queues. Each admitted pair holds one lane for the process lifetime,
even after its queue empties, and each lane holds an open file per
ingester.backpressure.shards writer. If any pod can receive every pair, each
pod needs capacity for the full set of pairs. Check both X-Scope-OrgID and
x-siglake-index when observed cardinality exceeds the expected set.
An empty queue does not reclaim a lane, so waiting does not make the refused pair admissible on that pod. Use one of these recovery paths:
- Correct a client that varies either routing header, then restart the affected pod to clear the unwanted lanes.
- Route the pair to an ingester with a free slot, or add a pod and route traffic to its fresh per-process lane map.
- Raise
ingester.maxLanesor--ingest-max-lanes, then restart. The cap is read at startup. - Restart with the same cap only to clear the map. Keys race for the finite slots again, so this does not guarantee admission for the refused pair.
Retries alone can stay pinned to the refusing pod until their budget expires. Recovery requires routing or operator action; the response supplies no delay because this lane map has no timed reclamation.
Ingest is rejecting¶
Five settings control ingest throughput: lane capacity
(--ingest-backpressure-capacity), writer shards
(--ingest-backpressure-shards), group commit (--ingest-group-commit-ms),
the rate budget (--ingest-rate-per-sec) and the memory breaker
(--ingest-mem-limit-mib). Tune them before adding CPU or replicas. A spent
rate budget returns 429; a full lane or open memory breaker returns 503;
all three responses include Retry-After. An oversized body returns 413 and
requires a smaller request.
-
Match the response to the limit that refused the export.
Response Confirm with Limit 429siglake_ingest_rate_limit_rejected_totalrisesRate budget 503siglake_ingest_backpressure_rejected_totalrisesLane capacity 503siglake_ingest_mem_breaker_openis1Memory breaker Hintless 503or gRPCUnavailablesiglake_ingest_lane_refused_totalrisesPersistent lane count 413Request size exceeds the configured cap Body cap -
If a full lane caused
503, compare rejections withsiglake_ingest_backpressure_queue_events. If the queue stays shallow and rejections are brief, raise--ingest-backpressure-capacityoringester.backpressure.capacityabove its default of1024. -
If the queue grows, increase
--ingest-backpressure-shardsoringester.backpressure.shardsabove its default of1. You can also set--ingest-group-commit-msoringester.backpressure.groupCommitMsabove its default of0to amortize fsync cost. Group commit adds acknowledgement latency. Both settings require a positive lane capacity; capacity0selects the legacy writer path. If the queue keeps growing, add CPU or replicas instead of more queue capacity. -
If the rate budget caused
429, raise--ingest-rate-per-secoringester.rateLimit.ratePerSec. Set--ingest-rate-burstoringester.rateLimit.bursthigh enough for the expected bursts. -
If the memory breaker caused
503, raise--ingest-mem-limit-miboringester.memBreaker.limitMibonly when the process has memory available. Otherwise, add memory before raising the limit. The breaker is disabled by default with a limit of0. -
If Siglake returns
413, split the export into smaller requests. Do not retry the same oversized body. -
For a transient
429or a503that includesRetry-After, configure the exporter to honor the delay. A lane-cap503has no hint and needs the separate recovery above. Retries do not guarantee delivery: the exporter's retry deadline can expire, and its queue can fill. The OpenTelemetry guide configures both limits explicitly.
Commits are slow¶
If siglake_iceberg_commit_duration_seconds and sealed_pending rise together,
check the catalog first.
- The catalog is usually the bottleneck. Check RDS CPU and connections.
- Raise
SIGLAKE_COMMIT_BATCH_TARGET_MBto amortize over more rows. This adds committed-visibility lag. - Set
SIGLAKE_SIDE_AGG_WRITE_BEHIND=1to move the side-aggregate read-modify-write off the commit path. - Confirm snapshot expiry is running. Every commit reads and rewrites an
unbounded
metadata.json.
Fresh data isn't visible¶
If new rows take tens of seconds to appear, check
query.walBuffer.enabled first.
The chart defaults query.walBuffer.enabled to false. Without the WAL buffer,
you get commit-cycle visibility rather than visibility within seconds.
If it is enabled:
- Check
siglake_query_wal_buffer_rows: zero means the buffer isn't serving. - Check
siglake_query_wal_buffer_segment_errors_total. - Remember the floor is seal age:
--wal-max-age-secs(default 5). - User indexes are served after a commit, not from the buffer.
See Freshness.
Query throughput plateaus¶
If adding query replicas stops helping around 16-way concurrency, check
siglake_query_exec_pool_queue_seconds.
Check siglake_query_exec_pool_queue_seconds. Sub-millisecond queue wait with
a plateau means CPU-bound browse decode, not scheduling: more replicas will
help, more threads per replica won't.
Confirm layout convergence first. Most "query doesn't scale" turns out to be "layout isn't disjoint".
Search acceleration¶
Three independent switches govern raw-text search. They work together, and none replaces another.
| Switch | Set through | Default | Governs |
|---|---|---|---|
index_at_flush |
per-table index config, else SIGLAKE_INDEX_AT_FLUSH |
inline (on) | Whether an ingest-time (generation-0) write builds raw-text indexes. Under the chart, the compactor performs that write; see below. |
SIGLAKE_INVERTED_INDEX |
compactor.invertedIndex.enabled |
on | Whether inverted indexes are in the index set, on any path. Only the literal 0 turns it off. |
SIGLAKE_INDEX_REBUILD |
compactor.indexRebuild |
off | Whether a post-rewrite pass backfills Puffin inverted indexes a rewrite left missing. Only the literal 1 turns it on. Existing indexes are read whatever it is set to. |
index_at_flush¶
Applies to ingest-time writes only. "Ingest-time" is the generation of the
write, not the pod that makes it: under chart defaults the ingester only
appends to the WAL, and the WAL → Parquet flush that counts as generation 0
runs in the compactor process. So SIGLAKE_INDEX_AT_FLUSH belongs on
compactor.extraEnv. Keep compactor.enabled: true, its default. The chart
refuses --with-compactor in ingester.extraArgs because that embedded
compactor has no catalog claim. A rolling update can otherwise run two
compactors against the same table.
Setting it false makes a gen-0 write skip inverted-index tokenization, the file-level trigram bloom, the row-group token blooms and the Puffin sidecar. These steps took about 30% of append time in the measurement. Group-count and time-bucket footers are written either way, so the aggregate fast paths are unaffected.
Compaction rewrites (generation ≥ 1) always build the index set; index_at_flush
is not consulted on the rewrite path. That is what "deferred to compaction"
means, and it is the only part of the deferral that is unconditional.
Inverted-index enablement¶
Inverted indexes are in the index set by default.
Set
SIGLAKE_INVERTED_INDEX=0 to remove them from both the flush and compaction
paths. Helm users set compactor.invertedIndex.enabled: false. Operator users
put SIGLAKE_INVERTED_INDEX=0 in spec.extraEnv.
This has two consequences:
- The chart wires
SIGLAKE_INVERTED_INDEXinto the compactor Deployment only. The separate compactor makes every Parquet write, including the gen-0 flush and later rewrites. What inline indexing buys at the leading edge is trigram and row-group bloom pruning on top. Setting the variable oningester.extraEnvchanges nothing. - The index set is derived from the table.
eventsindexesraw. A user index indexes itstextfields whose tokenizer isdefaultorstem. A table declaring onlyraw-tokenized text gets no inverted index even with the switch on. See the user indexes guide.
Post-compaction index rebuilding¶
The rebuild pass is off by default.
To opt in, set SIGLAKE_INDEX_REBUILD=1, Helm
compactor.indexRebuild: true, or operator
spec.extraEnv: [{name: SIGLAKE_INDEX_REBUILD, value: "1"}]. The pass then
runs after a compaction or delete-task rewrite commits and registers Puffin
inverted-index sidecars for the rewritten files that have none.
Leaving it off changes writes only. Indexes a file already carries are still discovered and used, and the flush path still writes its footer inverted index. What you lose is pruning on rewrite output that has no index: the query scans those files and returns exact rows. The default is off because a compacted file's inverted index costs about 40 bytes per indexed row, roughly 294 MB parsed for a 7.3M-row file, so a text query over a dozen of them needs several gigabytes of parsed index against the query pod's parsed-index cache. Turn it on where the working set fits that cache, or where pruning is worth more than the decode.
A rewrite output needs a sidecar when the writer could not put the index in the Parquet footer:
- A streamed rewrite writes no footer inverted index at all. Every merge past the in-memory caps takes that path.
- An in-memory rewrite whose serialized index for a column exceeds
SIGLAKE_INDEX_FOOTER_MAX_BYTESsends that column to a sidecar, and the rewrite commit itself does not register it.
The pass looks at each file and column the enabled index specifications name, and skips the column when the file's Parquet footer already carries that index or the table metadata already registers a Puffin blob for the pair. So an in-memory rewrite whose indexes fit the footer gets nothing new, and running the pass twice over the same files registers nothing the second time. A registration outlives the snapshot that made it: after snapshot expiry the column still counts as indexed, and the sidecar still counts as reachable for the orphan sweep.
The pass returns immediately when the enabled index set holds no inverted
index, so it does nothing under compactor.invertedIndex.enabled: false. If it
fails, the compactor logs a warning and the rewrite stays committed: you lose
pruning on those files, not data.
Watch siglake_index_rebuild_files_total, siglake_index_rebuild_bytes_total
and siglake_index_rebuild_seconds; see
v1 index rebuild cost for
their labels and export forms.
A minimal deferred-index configuration¶
Defer the inline cost on a firehose stream and get inverted indexes at L1:
compactor:
extraEnv:
- name: SIGLAKE_INDEX_AT_FLUSH
value: "0" # defer the 30% inline append cost to compaction
Inverted indexes need no setting here: they are on unless you turn them off,
so an in-memory merge at L1 writes them into the Parquet footer. A streamed
merge writes none, and nothing adds them later unless you also set
SIGLAKE_INDEX_REBUILD=1. The variable goes on the compactor because the
compactor makes both writes these switches govern.
Watch siglake_index_build_seconds, siglake_index_build_bytes and
siglake_write_index_deferred_total.
Search stays correct without indexes
The query path scans an unindexed file and still returns the correct result. These switches trade CPU and storage for speed. Deferral raises query latency on recent data until files consolidate.
Attribute promotion¶
If you filter constantly on an attribute, promote it:
Promoted columns get Parquet statistics, so queries prune instead of scanning
and parsing JSON again. Queries do not need to change. The query layer
rewrites attr_get() predicates onto the promoted column.
Measured effect: an attribute filter went from 413 ms to 1.9 s to 90 ms across the promotion work, and typed attribute counts reached about 3.5 ms zero-scan.
Run siglake migrate-schema --table events --promote-attr … with the same set
first, to widen an existing warehouse.
Cache tuning¶
| Variable | Chart setting and default | Effect |
|---|---|---|
SIGLAKE_FOOTER_CACHE_CAP |
No chart setting | Footer cache entries. Raise for wide tables. |
SIGLAKE_OBJECT_CACHE_BYTES |
No chart setting | Byte-range object cache. 0 disables. |
SIGLAKE_QUERY_SCAN_FILE_CACHE_MAX_BYTES |
query.scan.fileCacheMaxBytes: 0 |
Experimental decoded source-file batch cache byte limit. Absorbs warm repeat S3 reads for the scan shapes that can fill it. |
SIGLAKE_QUERY_SCAN_FILE_CACHE_MAX_ENTRIES |
query.scan.fileCacheMaxEntries: 0 |
Experimental decoded source-file batch cache entry limit. |
SIGLAKE_AGG_RESULT_CACHE_CAP |
No chart setting | Aggregate result entries. |
The experimental source-file cache is enabled only when both limits are
positive. The chart defaults both settings to 0, while a bare binary leaves
both variables unset; either configuration disables the cache. The chart passes
both zero values through explicitly and does not derive defaults.
Few scan shapes fill an entry. The scan has to run unordered, carry no
predicate the Iceberg converter accepts (a timestamp window counts as one),
and read a whole file task to its end. Table drains and unfiltered aggregates do
that. A newest-first browse never takes the cache path, a filtered browse
bypasses it, and a LIMIT satisfied from the first batches drops the population
before it inserts, so a log-UI workload fills close to nothing while the byte
limit is still subtracted from the query memory pool. A filtered browse prunes
the same pages with the cache on as with it off, and gains nothing from the
cache until an eligible scan has filled its entry. Watch
siglake_query_scan_file_cache_requests_total{outcome} for insert before
raising the limits: see Decoded-file cache
outcomes.
When SIGLAKE_OBJECT_CACHE_BYTES is not set explicitly, the byte-range object
cache derives its size as 25% of the cgroup memory limit, clamped from 64 MiB to
16 GiB, or uses 1 GiB when no cgroup limit is available. A useful starting
budget when explicitly enabling the source-file cache is 12.5% of the memory
limit; that share is a sizing recommendation, not a reservation or derived
default. See the
query memory budget for how the caches affect the
query pool and process headroom.
A cold top_hosts query on a 2 B-row table took 6.7 seconds once per snapshot.
The warm query took 0.49 ms.
Text queries have two caches of their own, with their own limits: see Tune the Puffin and parsed text-index caches.
What the decoded-file cache byte limit covers¶
SIGLAKE_QUERY_SCAN_FILE_CACHE_MAX_BYTES bounds completed entries and the
decoded batches that in-flight populations hold, together. A scan filling an
entry charges each batch it keeps at that batch's decoded size, against the
same total the finished entries are counted in. Size the limit as all the
memory you will give this cache: populations get no allowance on top of it.
A charge that fits is one atomic add. When residents leave too little room, the population takes the cache lock without waiting, drops oldest entries until its batch fits, and charges then, so a changing working set still replaces residents instead of freezing the first files a scan reached. If another thread holds that lock, or the batch will not fit even with the cache emptied, the population is refused: the query returns the same rows, nothing is cached for that file, and no outcome counter records the refusal.
At end of file the accumulated charge becomes the entry's own. Cancellation, a
LIMIT satisfied from the first batches, a read failure, a candidate over the
quarter-budget entry bound, an entry another partition inserted first and a
contended insert each release it.
Defaults are unchanged: the chart still ships 0 and 0, packaged memory
limits are the same, and the byte limit is still subtracted from the query
memory pool. siglake_query_scan_file_cache_bytes reports completed entries
only; the bytes held by live populations are not exported.
Tune the Puffin and parsed text-index caches¶
A text query reads a per-file inverted index. Siglake caches it on both sides
of its decode: the parsed form a warm query is served from, and the serialized
Puffin blob a parsed eviction falls back on. Both budgets derive from the pod's
memory limit, both are subtracted from the DataFusion query pool, and both are
published on siglake_cache_budget_bytes{kind="text_index"}.
They are also the last claim on that limit. They take only what is left once the pool can still reserve one compacted file's decode working set, 1.25 GiB. A 4 GiB pod has nothing left, so it derives no text-index caches at all and deserializes an index on every text query. A 5 GiB pod derives the full 400 MiB, 320 MiB parsed and 80 MiB of blobs; a 16 GiB pod reaches both caps. Without a cgroup limit to read, both fall back to 1 GiB and 256 MiB.
| Variable | Default | Effect |
|---|---|---|
SIGLAKE_PARSED_INDEX_CACHE_MAX_BYTES |
1/16 of the memory limit, capped at 1 GiB | Byte limit on parsed indexes, which cost about 40 bytes per indexed row. Evicts least-recently-used. |
SIGLAKE_PUFFIN_BLOB_CACHE_MAX_BYTES |
1/64 of the memory limit, capped at 256 MiB | Byte limit on the serialized blobs. |
SIGLAKE_PUFFIN_BLOB_CACHE_MAX_ENTRIES |
128 | Entry limit on both caches. The tighter of the entry and byte bounds governs. |
None of the three has a chart setting. Set them through query.extraEnv. The
query server also accepts --query-parsed-index-cache-max-bytes and
--query-puffin-blob-cache-max-bytes.
To turn both caches off, set SIGLAKE_PUFFIN_BLOB_CACHE_MAX_ENTRIES to 0:
every text query then deserializes its index again, and a Puffin one fetches
the blob again first. To drop only the serialized copy, set
SIGLAKE_PUFFIN_BLOB_CACHE_MAX_BYTES to 0: warm queries still reuse parsed
indexes, and a parsed miss re-fetches the Puffin blob. An index or blob larger
than its whole budget is left uncached rather than evicting the entries that
fit.
Both forms of index are held parsed, under the same byte and entry bounds. A
Puffin entry is keyed by (statistics path, blob offset); a footer-KV index,
which is what an in-memory rewrite writes, is keyed by (data file path,
column). Each identity is written once and never rewritten, so no entry can go
stale. The identity says the bytes behind it never change, not that they are
still sound. A footer-KV index written with a sibling CRC-32, which is every
one written from Siglake 0.2.0, is checked against it before every handout,
warm or cold, so a corrupt one is refused rather than served from an earlier
parse. A Puffin entry is not re-checked and answers from its parse until it is
evicted. See What the footer text-index checksum
covers.
A footer-KV index is the one that skips the object-store fetch: its bytes
arrive with the Parquet metadata the scan already read. It still pays the
same deserialization and the same
SIGLAKE_INDEX_LOAD_CONCURRENCY
permit on a cold read, so warming it buys the decode, not a fetch. Holding the
serialized Puffin form is the other way round: it saves the fetch, not the
deserialization.
The packaged 4 GiB query pod caches no text indexes
Give the pod 5 GiB to get the derived 400 MiB. Setting the two byte limits by hand at 4 GiB also works, and costs the query memory budget the bytes you give them: that trades the pool's decode reservation for warm indexes.
What the packaged 4 GiB query pod measured on the 50 GB suite¶
The 2026-09-15 50 GB benchmark ran the query server under the packaged 4 GiB
limit. It reported siglake_cache_budget_bytes{kind="text_index"} as 0, which
is what the cache derivation
gives at that limit, and it finished with no out-of-memory (OOM) kill and no
restart. A later reread of that round against
the recorded unlimited-pod round reported these p50 latencies:
| Query shape | 4 GiB pod | Unlimited pod |
|---|---|---|
match_all |
12.14 ms | 7.32 ms |
label_filter_last25 |
62.26 ms | 10.95 ms |
deep_pagination |
14.25 ms | 9.94 ms |
multi_label_and |
121.32 ms | 14.87 ms |
The two rounds planned different files for every shape in the table, so this
is not an isolated measurement of what the limit costs, and none of the four
shapes is a text query. Raising the pod to 5 GiB, or setting
SIGLAKE_PARSED_INDEX_CACHE_MAX_BYTES and
SIGLAKE_PUFFIN_BLOB_CACHE_MAX_BYTES by hand, is a sizing choice: no round has
measured what either does to these latencies.
Benchmarking correctly¶
Turn result caches off, or you are measuring the cache
A 1 TB benchmark board was rendered meaningless without this: every shape collapsed onto a 13 ms result-cache-hit floor, with five orders of magnitude difference in rows scanned producing latencies within 2 ms of each other, and throughput plateauing on a saturated cache rather than the engine.
Also:
- Use novel literals per iteration, or you re-hit snapshot-keyed caches.
- Report cold and warm as separate numbers, not one blended figure.
- Verify row order on ordered queries: a reversed-read correctness bug once hid behind count-only validation.
- State the layout state. An unconverged table is a different system than a converged one.
See Performance for the published numbers and their methodology.