Skip to content

Performance tuning

Use the timings and counters in each query response to find the slow stage. Change one setting at a time, then run the same query again to check the result.

Tell a metadata fast path from a Parquet scan

Every query response carries its own explanation:

curl -s http://localhost:8089/api/v1/sql \
  -H 'Content-Type: application/json' \
  -d '{"query": "…"}' | jq '{cost, stats}'
Reading Meaning
stats.served_by starts with tier1_ A warm metadata fast path served it.
stats.served_by == "materialized" Exact per-file fallback; potentially slow even when rows_scanned == 0.
stats.served_by == "scan" The full query plan ran.
cost.files_considered ≫ cost.files_to_scan Planning-time pruning is working.
scan.files_planned ≫ scan.files_read Execution-time pruning is working.
phases.plan_micros ≫ collect_micros Planning-bound: usually too many files.
phases.collect_micros dominant Scan-bound: decode or I/O.
scan.bytes_footer dominates scan.fetched_bytes Footer-bound: improve metadata-cache reuse or compact toward fewer, larger files. If scan.bytes_data dominates instead, narrow the projection or improve pruning.
scan.ordering != "advertised" Ordered early-stop refused. The value is the reason.
phases.distributed.shard_wall_micros uneven A straggler shard.

By symptom

Aggregates are scanning

If an aggregate is slow, check whether it reports served_by: "materialized" or returns rows_scanned > 0.

Check stats.served_by first. rows_scanned can be zero for both a millisecond metadata answer and a much slower per-file fallback.

Repair incomplete group counts

served_by: "materialized" is an exact answer, but it was assembled from each live file's group-count footer or, where necessary, by decoding raw pages. If a group-by that used to be fast starts taking this path:

  1. Check the SiglakeGroupCountDeltaRetrying alert. The commit path tries a delta write four times, waiting 250, 500, and 750 ms between retries. A sustained retry alert points first to object-store health or credentials on the named pod. An exhausted write increments siglake_group_count_delta_write_failures_total, records a durable rebuild marker, and does not fire SiglakeGroupCountDeltaLost by itself. That alert names the table only when the subsequent automatic rebuild fails or remains incomplete.
  2. If SiglakeGroupCountDeltaLost is not firing, wait for the maintenance compactor's next aggregate fold. It consumes the rebuild marker, reconstructs the table's exact maps and bounded sketches from committed files, and deletes the covered markers. Confirm a successful automatic repair with an increase in siglake_group_count_auto_rebuilds_total{outcome="success"}; siglake_group_count_aggregate_rebuilds_total also increases after any completed rebuild and is a secondary signal. A later delta does not heal the gap; the marker-driven rebuild does.

siglake_group_count_deltas_absorbed_total records deltas folded into the base, and siglake_group_count_deltas_deleted_total records covered delta objects removed by maintenance.

Only a compactor running --role maintenance or --role combined performs this fold. A fleet whose compactors all run --role drain has nothing to consume rebuild markers.

If SiglakeGroupCountDeltaLost is firing, waiting is over: check whether its outcome is failed or incomplete, read the compactor log, and use the operator fallback:

siglake rebuild-group-counts --table <table>

Automatic and operator-triggered rebuilds are safe to run live and to repeat. They record a rebuilt_through watermark so a late delta at or below the scanned snapshot is not folded twice; later deltas continue to fold normally. 3. Handle pre-existing typed columns separately: they are never admitted automatically. Lost-delta repair restores only the columns named by the failed commit. To admit eligible typed (long, double, or bool) columns from a table created before typed columns joined the side aggregate, run:

siglake rebuild-group-counts --table <table> --admit-typed-columns

Admission is all-or-nothing and subject to SIGLAKE_TYPED_GROUP_COUNT_CARDINALITY; the command reports a column rather than writing a partial total when the cap is exceeded or a live file cannot serve it.

The repair deliberately writes only the wide aggregate object, so a repaired column reports served_by: "tier1_wide", even when it would fit the inline cap. No data-file rewrite is needed merely to admit a typed column, nor to move a repaired column into the inline object. A compaction rewrite is needed only when the rebuild reports that a live file cannot supply the column from either its footer or a raw-page decode (for example, the file predates typed footers or lacks the column); rerun the rebuild after those files have been rewritten. Answers remain exact before and during repair.

Diagnose a scan

served_by: "scan" means the full query plan ran. For a dimensional count, the matcher may not have recognized the predicate shape. Known limits:

  • Float literals are excluded: float rendering isn't reproducible enough to match footer keys.
  • The predicate must reduce to footer equality or an integer range.
  • The column must have group-count footers, which older files may lack until compaction rewrites them.

Check whether the same shape without the predicate is zero-scan; if so, the predicate is the problem. SIGLAKE_LOG_EXEC_PLANS=1 shows what the planner actually built.

Ordered browses are slow

If ORDER BY timestamp … LIMIT n is slow, check whether scan.ordering is advertised.

ordering value Cause Fix
filtered The residual filter defeated the ordering gate: a non-timestamp predicate or a prune spec. Expected on filtered browses, where a bounded TopK is the cheaper plan. Nothing to fix if the query is fast. Otherwise narrow the time window.
no_bounds Files lack manifest time bounds. Let compaction rewrite them.
fan_in One cluster of files overlaps in time more deeply than the per-partition merge will open. Fix the compaction layout as described below.
global_fan_in Overlap fits per partition but not the whole-scan stream budget. Same layout problem; SIGLAKE_ORDERED_MERGE_GLOBAL_FANIN raises the budget.
mixed_sort_direction, unknown_sort_order, order_walk_failed The scan mixes files written under different sort orders, or their orders could not be attributed. The table is part-way through the DESC to ASC convergence. Let re-clustering rewrite the stragglers.
missing_sort_column The layout checks passed, but timestamp was unexpectedly absent from the scan's output schema after the earlier projection check. This indicates a Siglake bug, not a layout or query-shape problem. Report it rather than re-clustering.
not_projected, no_sort_order, non_identity_sort, missing_sort_field, non_timestamp_sort The query does not project timestamp, or the table is not declared sorted by it. Query or table shape, not layout: see Ordered early-stop.

Raise SIGLAKE_ORDERED_MERGE_MAX_FANIN if partitions legitimately have more overlapping files than the merge will open, but treat that as a workaround for a layout problem. Check what the node already allows before raising it: both caps default to a node-adaptive value, four open streams per CPU clamped to 16 to 64 per partition and eight per CPU clamped to 32 to 128 across the scan, so a large worker is already well above the small-pod floors of 16 and 32. An explicit value in either variable is used as given, not clamped.

Decide whether table layout is converging or falling behind ingest

If query latency rises with table size, check whether siglake_table_overlap_depth is high or flat. The leveled-compaction rules define the file-count, overlap-depth, cadence, and backpressure rules.

Treat this as a compaction problem. It is the most common cause of query latency that rises as a table grows.

  1. Check siglake_table_overlap_depth. A converged 2 B-row table settles in the low twenties. This gauge comes from a budgeted manifest walk and can lag on a large table, so verify that siglake_table_gauges_sampled_at_seconds is advancing before treating a flat value as current. The timestamp covers the level, leading-edge, and overlap-depth gauges and advances only after a complete walk.
  2. Check siglake_table_live_data_files: unbounded growth means compaction is losing. This count is exact from the snapshot summary on every sampler cycle, regardless of whether the manifest walk finishes.
  3. Check siglake_compactor_maintenance_skipped_backpressure_total and …_throttled_backpressure_total: compaction may be permanently yielding to the drain.

Fixes, in order:

  • Add compactor replicas (with catalogClaim.enabled) so the drain stops starving maintenance.
  • Keep the default leveled compaction enabled. If SIGLAKE_COMPACTOR_LEVELED=0 is present, remove the override so the default leveled scheduler runs.
  • Lower SIGLAKE_COMPACTOR_MAX_OVERLAP_DEPTH to trigger depth merges sooner.
  • Raise SIGLAKE_COMPACTOR_BACKPRESSURE_COMPACT_EVERY frequency (lower the number) so throttled passes run more often under backlog.

Allow 4 to 5 hours of background compaction for a 2 B-row table to converge. In the 1 TB test, convergence halved browse latency and raised 32-way throughput from 28 QPS to 115 QPS.

Queries return 503

If queries return 503 with Retry-After: 5, read the response body's reason first.

shard_pin_unresolved is a refused distributed query. Diagnose it below. The rest of this section covers the capacity case, where siglake_query_breaker_trips_total{breaker="pool_exhausted"} is rising.

  1. Check SiglakeQueryPoolNearLimit for sustained pool pressure.
  2. Request GET /debug/memory-pool from the affected query pod to identify the largest live consumers.
  3. Compare the pod's query.resources.limits.memory sizing budget with its query concurrency and workload.
  4. If the pool budget is adequate, inspect the spill volume. Sorts and aggregates spill after the pool tightens, but a query is refused when the spill directory reaches query.spill.maxBytes (SIGLAKE_QUERY_SPILL_MAX_BYTES). Check the pod's ephemeral-storage usage, and raise the cap together with query.spill.sizeLimit and the ephemeral-storage limit in the order Query spill and ephemeral storage describes.

Distributed workers do not reserve admission shares; see Distributed admission is per-coordinator for the worker-side bounds.

This is transient server capacity, not a client rate-limit error. Honor Retry-After; the same query can fit after concurrent work finishes.

A 503 with reason: shard_pin_unresolved

If distributed queries return 503 with reason: "shard_pin_unresolved" in the body, siglake_query_shard_pin_total{outcome="miss"} is rising, and SiglakeQueryShardPinUnresolved fires. Single-node /api/v1/sql/local requests are unaffected.

A worker could not resolve the generation the coordinator pinned the fan-out to: either the snapshot or the schema id. It refused its shard instead of contributing rows from a different generation. See One generation per fan-out. The error message names which half failed.

Isolated misses clear as caches converge. Sustained schema misses mean a migrate-schema commit has not reached every replica's metadata cache; the same cache lifetime governs both halves. Sustained snapshot misses mean coordinators are serving snapshots the catalog has already expired. Both remedies below close that gap:

  1. Raise compactor.snapshotExpire.retainLast so expiry keeps more snapshots and widens the margin. That section covers the packaged default, what a wider window costs and the other consumers: time travel, long external reads and the mixed-version rollout guard. These constrain the value too.
  2. Lower SIGLAKE_ICEBERG_METADATA_CACHE_TTL_SECS so coordinators stop serving a snapshot before the catalog expires it. This also narrows the window in which replicas disagree, at the cost of more catalog metadata loads.

A format: "ndjson" query refused after its stream started keeps the 200 and reports the refusal as a final {"_meta": "error", "code": 503, "retry_after_secs": 5} line, so a client that only inspects status codes will read a truncated result as a complete one; see NDJSON trailers.

Ingest is rejecting

Five settings control ingest throughput: lane capacity (--ingest-backpressure-capacity), writer shards (--ingest-backpressure-shards), group commit (--ingest-group-commit-ms), the rate budget (--ingest-rate-per-sec) and the memory breaker (--ingest-mem-limit-mib). Tune them before adding CPU or replicas. A spent rate budget returns 429; a full lane or open memory breaker returns 503; all three responses include Retry-After. An oversized body returns 413 and requires a smaller request.

  1. Match the response to the limit that refused the export.

    Response Confirm with Limit
    429 siglake_ingest_rate_limit_rejected_total rises Rate budget
    503 siglake_ingest_backpressure_rejected_total rises Lane capacity
    503 siglake_ingest_mem_breaker_open is 1 Memory breaker
    413 Request size exceeds the configured cap Body cap
  2. If a full lane caused 503, compare rejections with siglake_ingest_backpressure_queue_events. If the queue stays shallow and rejections are brief, raise --ingest-backpressure-capacity or ingester.backpressure.capacity above its default of 1024.

  3. If the queue grows, increase --ingest-backpressure-shards or ingester.backpressure.shards above its default of 1. You can also set --ingest-group-commit-ms or ingester.backpressure.groupCommitMs above its default of 0 to amortize fsync cost. Group commit adds acknowledgement latency. Both settings require a positive lane capacity; capacity 0 selects the legacy writer path. If the queue keeps growing, add CPU or replicas instead of more queue capacity.

  4. If the rate budget caused 429, raise --ingest-rate-per-sec or ingester.rateLimit.ratePerSec. Set --ingest-rate-burst or ingester.rateLimit.burst high enough for the expected bursts.

  5. If the memory breaker caused 503, raise --ingest-mem-limit-mib or ingester.memBreaker.limitMib only when the process has memory available. Otherwise, add memory before raising the limit. The breaker is disabled by default with a limit of 0.

  6. If Siglake returns 413, split the export into smaller requests. Do not retry the same oversized body.

  7. Configure the exporter to honor Retry-After on 429 and 503. Retries do not guarantee delivery: the exporter's retry deadline can expire, and its queue can fill. The OpenTelemetry guide configures both limits explicitly.

Commits are slow

If siglake_iceberg_commit_duration_seconds and sealed_pending rise together, check the catalog first.

  • The catalog is usually the bottleneck. Check RDS CPU and connections.
  • Raise SIGLAKE_COMMIT_BATCH_TARGET_MB to amortize over more rows. This adds committed-visibility lag.
  • Set SIGLAKE_SIDE_AGG_WRITE_BEHIND=1 to move the side-aggregate read-modify-write off the commit path.
  • Confirm snapshot expiry is running. Every commit reads and rewrites an unbounded metadata.json.

Fresh data isn't visible

If new rows take tens of seconds to appear, check query.walBuffer.enabled first.

The chart defaults query.walBuffer.enabled to false. Without the WAL buffer, you get commit-cycle visibility rather than visibility within seconds.

If it is enabled:

  • Check siglake_query_wal_buffer_rows: zero means the buffer isn't serving.
  • Check siglake_query_wal_buffer_segment_errors_total.
  • Remember the floor is seal age: --wal-max-age-secs (default 5).
  • User indexes are served after a commit, not from the buffer.

See Freshness.

Query throughput plateaus

If adding query replicas stops helping around 16-way concurrency, check siglake_query_exec_pool_queue_seconds.

Check siglake_query_exec_pool_queue_seconds. Sub-millisecond queue wait with a plateau means CPU-bound browse decode, not scheduling: more replicas will help, more threads per replica won't.

Confirm layout convergence first. Most "query doesn't scale" turns out to be "layout isn't disjoint".

Search acceleration

Three independent switches govern raw-text search. They work together, and none replaces another.

Switch Set through Default Governs
index_at_flush per-table index config, else SIGLAKE_INDEX_AT_FLUSH inline (on) Whether an ingest-time (generation-0) write builds raw-text indexes. Under the chart, the compactor performs that write; see below.
SIGLAKE_INVERTED_INDEX compactor.invertedIndex.enabled on Whether inverted indexes are in the index set, on any path. Only the literal 0 turns it off.
SIGLAKE_INDEX_REBUILD compactor.indexRebuild off Whether a post-rewrite pass backfills Puffin inverted indexes a rewrite left missing. Only the literal 1 turns it on. Existing indexes are read whatever it is set to.

index_at_flush

Applies to ingest-time writes only. "Ingest-time" is the generation of the write, not the pod that makes it: under chart defaults the ingester only appends to the WAL, and the WAL → Parquet flush that counts as generation 0 runs in the compactor process. So SIGLAKE_INDEX_AT_FLUSH belongs on compactor.extraEnv. Keep compactor.enabled: true, its default. The chart refuses --with-compactor in ingester.extraArgs because that embedded compactor has no catalog claim. A rolling update can otherwise run two compactors against the same table.

Setting it false makes a gen-0 write skip inverted-index tokenization, the file-level trigram bloom, the row-group token blooms and the Puffin sidecar. These steps took about 30% of append time in the measurement. Group-count and time-bucket footers are written either way, so the aggregate fast paths are unaffected.

Compaction rewrites (generation ≥ 1) always build the index set; index_at_flush is not consulted on the rewrite path. That is what "deferred to compaction" means, and it is the only part of the deferral that is unconditional.

Inverted-index enablement

Inverted indexes are in the index set by default. Set SIGLAKE_INVERTED_INDEX=0 to remove them from both the flush and compaction paths. Helm users set compactor.invertedIndex.enabled: false. Operator users put SIGLAKE_INVERTED_INDEX=0 in spec.extraEnv.

This has two consequences:

  • The chart wires SIGLAKE_INVERTED_INDEX into the compactor Deployment only. The separate compactor makes every Parquet write, including the gen-0 flush and later rewrites. What inline indexing buys at the leading edge is trigram and row-group bloom pruning on top. Setting the variable on ingester.extraEnv changes nothing.
  • The index set is derived from the table. events indexes raw. A user index indexes its text fields whose tokenizer is default or stem. A table declaring only raw-tokenized text gets no inverted index even with the switch on. See the user indexes guide.

Post-compaction index rebuilding

The rebuild pass is off by default. To opt in, set SIGLAKE_INDEX_REBUILD=1, Helm compactor.indexRebuild: true, or operator spec.extraEnv: [{name: SIGLAKE_INDEX_REBUILD, value: "1"}]. The pass then runs after a compaction or delete-task rewrite commits and registers Puffin inverted-index sidecars for the rewritten files that have none.

Leaving it off changes writes only. Indexes a file already carries are still discovered and used, and the flush path still writes its footer inverted index. What you lose is pruning on rewrite output that has no index: the query scans those files and returns exact rows. The default is off because a compacted file's inverted index costs about 40 bytes per indexed row, roughly 294 MB parsed for a 7.3M-row file, so a text query over a dozen of them needs several gigabytes of parsed index against the query pod's parsed-index cache. Turn it on where the working set fits that cache, or where pruning is worth more than the decode.

A rewrite output needs a sidecar when the writer could not put the index in the Parquet footer:

  • A streamed rewrite writes no footer inverted index at all. Every merge past the in-memory caps takes that path.
  • An in-memory rewrite whose serialized index for a column exceeds SIGLAKE_INDEX_FOOTER_MAX_BYTES sends that column to a sidecar, and the rewrite commit itself does not register it.

The pass looks at each file and column the enabled index specifications name, and skips the column when the file's Parquet footer already carries that index or the table metadata already registers a Puffin blob for the pair. So an in-memory rewrite whose indexes fit the footer gets nothing new, and running the pass twice over the same files registers nothing the second time. A registration outlives the snapshot that made it: after snapshot expiry the column still counts as indexed, and the sidecar still counts as reachable for the orphan sweep.

The pass returns immediately when the enabled index set holds no inverted index, so it does nothing under compactor.invertedIndex.enabled: false. If it fails, the compactor logs a warning and the rewrite stays committed: you lose pruning on those files, not data.

Watch siglake_index_rebuild_files_total, siglake_index_rebuild_bytes_total and siglake_index_rebuild_seconds.

A minimal deferred-index configuration

Defer the inline cost on a firehose stream and get inverted indexes at L1:

compactor:
  extraEnv:
    - name: SIGLAKE_INDEX_AT_FLUSH
      value: "0"       # defer the 30% inline append cost to compaction

Inverted indexes need no setting here: they are on unless you turn them off, so an in-memory merge at L1 writes them into the Parquet footer. A streamed merge writes none, and nothing adds them later unless you also set SIGLAKE_INDEX_REBUILD=1. The variable goes on the compactor because the compactor makes both writes these switches govern.

Watch siglake_index_build_seconds, siglake_index_build_bytes and siglake_write_index_deferred_total.

Search stays correct without indexes

The query path scans an unindexed file and still returns the correct result. These switches trade CPU and storage for speed. Deferral raises query latency on recent data until files consolidate.

Attribute promotion

If you filter constantly on an attribute, promote it:

compactor:
  promoteAttrs:
    - http.status_code:int
    - k8s.namespace:string

Promoted columns get Parquet statistics, so queries prune instead of scanning and parsing JSON again. Queries do not need to change. The query layer rewrites attr_get() predicates onto the promoted column.

Measured effect: an attribute filter went from 413 ms to 1.9 s to 90 ms across the promotion work, and typed attribute counts reached about 3.5 ms zero-scan.

Run siglake migrate-schema --table events --promote-attr … with the same set first, to widen an existing warehouse.

Cache tuning

Variable Chart setting and default Effect
SIGLAKE_FOOTER_CACHE_CAP No chart setting Footer cache entries. Raise for wide tables.
SIGLAKE_OBJECT_CACHE_BYTES No chart setting Byte-range object cache. 0 disables.
SIGLAKE_QUERY_SCAN_FILE_CACHE_MAX_BYTES query.scan.fileCacheMaxBytes: 0 Experimental decoded source-file batch cache byte limit. Absorbs warm repeat S3 reads.
SIGLAKE_QUERY_SCAN_FILE_CACHE_MAX_ENTRIES query.scan.fileCacheMaxEntries: 0 Experimental decoded source-file batch cache entry limit.
SIGLAKE_AGG_RESULT_CACHE_CAP No chart setting Aggregate result entries.

The experimental source-file cache is enabled only when both limits are positive. The chart defaults both settings to 0, while a bare binary leaves both variables unset; either configuration disables the cache. The chart passes both zero values through explicitly and does not derive defaults.

When SIGLAKE_OBJECT_CACHE_BYTES is not set explicitly, the byte-range object cache derives its size as 25% of the cgroup memory limit, clamped from 64 MiB to 16 GiB, or uses 1 GiB when no cgroup limit is available. A useful starting budget when explicitly enabling the source-file cache is 12.5% of the memory limit; that share is a sizing recommendation, not a reservation or derived default. See the query memory budget for how the caches affect the query pool and process headroom.

A cold top_hosts query on a 2 B-row table took 6.7 seconds once per snapshot. The warm query took 0.49 ms.

Text queries have two caches of their own, with their own limits: see Tune the Puffin and parsed text-index caches.

Tune the Puffin and parsed text-index caches

A text query reads a per-file inverted index. Siglake caches it on both sides of its decode: the parsed form a warm query is served from, and the serialized Puffin blob a parsed eviction falls back on. Both budgets derive from the pod's memory limit, both are subtracted from the DataFusion query pool, and both are published on siglake_cache_budget_bytes{kind="text_index"}.

They are also the last claim on that limit. They take only what is left once the pool can still reserve one compacted file's decode working set, 1.25 GiB. A 4 GiB pod has nothing left, so it derives no text-index caches at all and deserializes a Puffin index on every text query. A 5 GiB pod derives the full 400 MiB, 320 MiB parsed and 80 MiB of blobs; a 16 GiB pod reaches both caps. Without a cgroup limit to read, both fall back to 1 GiB and 256 MiB.

Variable Default Effect
SIGLAKE_PARSED_INDEX_CACHE_MAX_BYTES 1/16 of the memory limit, capped at 1 GiB Byte limit on parsed indexes, which cost about 40 bytes per indexed row. Evicts least-recently-used.
SIGLAKE_PUFFIN_BLOB_CACHE_MAX_BYTES 1/64 of the memory limit, capped at 256 MiB Byte limit on the serialized blobs.
SIGLAKE_PUFFIN_BLOB_CACHE_MAX_ENTRIES 128 Entry limit on both caches. The tighter of the entry and byte bounds governs.

None of the three has a chart setting. Set them through query.extraEnv. The query server also accepts both byte limits as command-line flags.

To turn both caches off, set SIGLAKE_PUFFIN_BLOB_CACHE_MAX_ENTRIES to 0: every text query then fetches and deserializes again. To drop only the serialized copy, set SIGLAKE_PUFFIN_BLOB_CACHE_MAX_BYTES to 0: warm queries still reuse parsed indexes, and a parsed miss re-fetches the blob. An index or blob larger than its whole budget is left uncached rather than evicting the entries that fit.

Entries are keyed by (statistics path, blob offset). A Puffin blob at an offset is written once and never rewritten, so neither entry can go stale. Holding the serialized form saves the object-store fetch, not the deserialization. Only Puffin-backed indexes are held parsed: a file whose inverted index sits in its Parquet footer, which is what an in-memory rewrite writes, deserializes again on every text query that reads it.

The packaged 4 GiB query pod caches no text indexes

Give the pod 5 GiB to get the derived 400 MiB. Setting the two byte limits by hand at 4 GiB also works, and costs the query memory budget the bytes you give them: that trades the pool's decode reservation for warm indexes.

What the packaged 4 GiB query pod measured on the 50 GB suite

The 2026-09-15 50 GB benchmark ran the query server under the packaged 4 GiB limit. It reported siglake_cache_budget_bytes{kind="text_index"} as 0, which is what the cache derivation gives at that limit, and it finished with no out-of-memory (OOM) kill and no restart. A later reread of that round against the recorded unlimited-pod round reported these p50 latencies:

Query shape 4 GiB pod Unlimited pod
match_all 12.14 ms 7.32 ms
label_filter_last25 62.26 ms 10.95 ms
deep_pagination 14.25 ms 9.94 ms
multi_label_and 121.32 ms 14.87 ms

The two rounds planned different files for every shape in the table, so this is not an isolated measurement of what the limit costs, and none of the four shapes is a text query. Raising the pod to 5 GiB, or setting SIGLAKE_PARSED_INDEX_CACHE_MAX_BYTES and SIGLAKE_PUFFIN_BLOB_CACHE_MAX_BYTES by hand, is a sizing choice: no round has measured what either does to these latencies.

Benchmarking correctly

Turn result caches off, or you are measuring the cache

SIGLAKE_QUERY_RESULT_CACHE=off

A 1 TB benchmark board was rendered meaningless without this: every shape collapsed onto a 13 ms result-cache-hit floor, with five orders of magnitude difference in rows scanned producing latencies within 2 ms of each other, and throughput plateauing on a saturated cache rather than the engine.

Also:

  • Use novel literals per iteration, or you re-hit snapshot-keyed caches.
  • Report cold and warm as separate numbers, not one blended figure.
  • Verify row order on ordered queries: a reversed-read correctness bug once hid behind count-only validation.
  • State the layout state. An unconverged table is a different system than a converged one.

See Performance for the published numbers and their methodology.