Skip to content

Performance

All numbers here come from internal AWS validation rounds on 3-node m6i-class clusters with an S3 warehouse and an RDS catalog. Roughly 100 rounds were run across the 2026-06/07 performance arc. Each round's topology, corpus, and per-shape latencies are recorded in internal reports that are not part of the public source tree. Public supporting material includes the source repository's DESIGN_*.md records and its OTLP ingest performance note. Published engine comparisons and their methodology are available on benchmarks.siglake.dev. Those records are separate from the historical AWS measurements summarized on this page.

Read measurements with their context

Each measurement describes a specific scale, layout, and warm or cold state. Use that context to compare with your workload and plan your own capacity measurements. See Methodology below.

What performance to expect at 200 GB and 1 TB scale

At 200 GB (394 million rows), warm metadata aggregates take about 3 ms, attribute filters about 9 ms, keyword search 8 to 12 ms, and an ordered 100-row browse about 45 ms. At 1 TB (2.02 billion rows) on a converged layout, aggregates take about 13 ms, a 100-row match_all browse 180 ms, and mixed query load reaches 115 queries per second at 32-way concurrency. Both scales reported about 5.5 seconds from ingest to query visibility at full ingest rate, with about 415,000 rows per second across four ingesters and three drain nodes.

For those numbers to hold you need the same conditions: AWS m6i-class nodes, S3 and RDS, a converged layout, result caches disabled or novel literals, cold and warm results kept apart, exact cross-arm row counts, and the same query shapes. To reproduce them, follow the methodology.

Ingest

Metric Result
Fleet throughput ~415 K rows/s sustained end to end (4 ingesters, 3 drain nodes, RDS catalog)
Peak (60 s window) ~537 K rows/s
Per-ingester scaling ~108 K → 200 K → 324 K → ~430 K rows/s at 1/2/3/4 ingesters
Sizing rule ~50 000 EPS per pod-CPU at default WAL thresholds
Correctness Exact row counts held through node crashes, via at-least-once recommit with dedup-by-proof

Ten consecutive 1 TB rounds produced exactly 2,016,590,653 rows, including through an abort-and-resume recovery.

Once drains keep up with accept, wall-clock throughput equals sustained throughput. One round completed in 27.5 minutes with 3 ingesters and 8 drains, where an earlier round with more ingesters and fewer drains took 36.5.

Freshness

Ingest to queryable is ~5.5 s at full ingest rate, 20/20 probes, at both 200 GB and 1 TB.

That figure is dominated by seal age, not by the commit cycle. A segment becomes visible when it seals, and a lone marker record waits out the 5 s age bound. Sustained streams seal by size in well under a second.

Query at 200 GB / 394 M rows

Shape Result
Zero-scan aggregates (whole-table GROUP BY, windowed counts, histogram, negation) ~3 ms warm, exact, rows_scanned: 0
Attribute GROUP BY (promoted) ~3 ms over 394 M rows
Attribute filters ~9 ms
count(distinct) ~3 ms, exact
Keyword search 8 to 12 ms
Label filter ~9 ms
Substring (trigram bloom) Bloom-pruned scans touching 10⁴ to 10⁵ of 394 M rows
Ordered browse (ORDER BY timestamp DESC LIMIT 100) ~45 ms warm via early-stop, no sort
Deep pagination (ORDER BY timestamp DESC LIMIT 100 OFFSET 50000) ~52 ms
Under 32-way mixed load Metadata fast paths hold ~6 ms flat on a dedicated runtime

Query at 1 TB / 2.02 B rows

On a converged layout, the same 21-shape suite runs zero-error with exact counts:

Shape Result
match_all browse 180 ms
Windowed browse 100 ms
Aggregates ~13 ms
Deep pagination (ORDER BY timestamp DESC LIMIT 100 OFFSET 50000) 192 ms
Freshness 20/20 at ~5.5 s

Concurrency: 93 / 117 / 115 QPS at 8 / 16 / 32-way, zero errors.

Both deep-pagination cells are the OFFSET shape, kept because it is what the comparison arms can all run. They do not describe the keyset cursor the cookbook recommends: a page carrying a timestamp_ns predicate gives up the ordered early-stop and runs as a bounded top-n over its window, so it is a different shape with no published latency here.

Layout convergence dominates

Convergence moved the query numbers further than any other single factor in the arc.

Layout state 32-way mixed QPS match_all
Unconverged (overlap depth 59) 28 QPS ~360 ms
Converged (depth ~21 to 23) 115 QPS 180 ms

Browse cells halved on convergence. Full convergence of a 2 B-row table takes roughly 4 to 5 hours of background compaction.

If query performance is disappointing, check siglake_table_overlap_depth before adding replicas. See Compaction.

Selected optimization results

Individual improvements, with their before and after:

Change Effect
Typed-count fast path status = 404 count: 134.5 ms → 1.15 ms; 5xx range 129.9 → 1.14 ms; negation 150.9 → 1.06 ms, all zero-scan
Attribute promotion Attribute filter: 413 ms → 1.9 s → 90 ms across the arc
Compact group-count footer 9.2× smaller, 5.3× faster on a wide footer: 133,506 → 14,548 bytes, 553 → 104 µs
Progressive tail chunks + single-partition coalesce match_all rows scanned 320 K → 8,192 (39×); windowed browse 1.83 M → 131 K (14×)
Depth-trigger fix (time-adjacent bins) Settle depth 105 → 16, converged in under an hour, through a historically stalling band
Warm aggregate caching top_hosts 701 ms → 0.49 ms warm (6.7 s cold, once per snapshot)

Most of these cut the work done rather than doing the same work faster: pruning and metadata answers instead of tighter loops.

Methodology

The result-cache failure

The 2026-07-27 1 TB benchmark board was unusable as a performance comparison, and the failure is easy to repeat.

Symptoms: repeated match_all replayed byte-identical plan and collect timings. All shapes sat within 2 ms of each other despite five orders of magnitude difference in rows scanned. Throughput plateaued at ~580 QPS while p50 doubled as concurrency doubled, which is Little's Law on a saturated cache.

The cause: a result-cache extension had landed after the prior board, so "13× faster / 5× QPS" was measuring a feature the baseline did not have. The suite was measuring memoization rather than the engine.

The fix was a cache-bypass switch, SIGLAKE_QUERY_RESULT_CACHE=off, with a test asserting it actually stops memoization.

Rules for reproducing these numbers

  1. Disable result caches (SIGLAKE_QUERY_RESULT_CACHE=off) or use novel literals per iteration.
  2. Report cold and warm as distinct metrics. A single blended number hides which one you measured. Cold top_hosts was 6.7 s where warm was 0.49 ms.
  3. State the layout convergence. An unconverged table is a different system.
  4. Verify row order on ordered queries. A reversed-read correctness bug once hid behind count-only validation: every reversed read of a multi-batch row group returned wrong rows, invisible to count and latency checks.
  5. Verify exact counts cross-arm, not approximate ones.
  6. Run arms concurrently when comparing systems, so both see the same node and network conditions.

Applying the measurements

  • Match the hardware, data shape, and query mix when comparing results; these measurements are not a performance guarantee for another deployment.
  • Account for compaction state. Results measured on converged layouts do not describe an unconverged backfill.
  • Match the configuration, including any opt-in features. The freshness figure requires the WAL buffer, which the chart ships disabled.

Comparative benchmarks

The comparisons cover Quickwit, Elasticsearch, ClickHouse, and a vanilla-Parquet DuckDB baseline. Visit the benchmarks site for published comparison records and methodology.

Corpora used include the Rally WorldCup98 HTTP-log corpus (247,249,096 documents exact, 45-day span, with realistic skew: 77% of documents in the last quarter).

In those rounds ClickHouse held the raw ordered-browse cell on local disk.

Benchmark publication status

The benchmarks site is live and serves published results and methodology. Comparative results are separate from public-release validation.