Skip to content

Compaction

Ingest creates small, often overlapping files. Compaction rewrites them into larger files with fewer overlapping time ranges while ingest continues.

This work costs object-store bandwidth and CPU. Siglake bounds each pass so the compactor can keep draining new WAL segments instead of waiting for a whole partition rewrite.

Roles

siglake compactor --role separates WAL draining from table maintenance.

Role Work
drain Claims sealed WAL segments and commits them to Iceberg.
maintenance Rewrites files, expires snapshots, samples gauges, folds aggregates, and executes enabled delete tasks.
combined Runs both sets of work. This is the default.

SIGLAKE_COMPACTOR_ROLE sets the same option. A zero SIGLAKE_RECLUSTER_INTERVAL_SECS disables file rewrites, but it does not make the process drain-only. Use the drain role when the process must not run maintenance.

Catalog claims let multiple drain processes share work. Maintenance uses per-table leases, so extra drain replicas do not require duplicate maintenance loops. See Scaling.

The disjointness invariant

File size alone does not determine query cost. If file time ranges overlap, an ordered scan must merge those files. Files with separate ranges can be read in sequence and stopped after a LIMIT is full.

The compactor therefore keeps each cluster of overlapping ranges together when it packs a rewrite. It may produce uneven output sizes to avoid splitting one cluster across bins.

Overlap depth is the largest number of file ranges that contain the same instant. siglake_table_overlap_depth reports this value. A high value means an ordered scan needs more simultaneous input streams.

Two strategies

Leveled compaction (default)

Leveled compaction groups files by byte size. A nonterminal level becomes eligible when it holds eight rewrite-eligible files. The compactor checks L0 every 10 seconds, L1 every 60 seconds, and L2 every 3,600 seconds. It selects the deepest overlapping cluster for rewrite when overlap depth exceeds 12.

When the sealed write-ahead log (WAL) backlog exceeds 256 segments, draining takes priority. Only every fourth backlogged cycle runs a restricted L0 pass. siglake_compactor_maintenance_skipped_backpressure_total counts the skipped cycles. siglake_compactor_maintenance_throttled_backpressure_total counts the restricted cycles.

A normal pass considers every due eligible level, ordered by file pressure. The levels share the pass's file and bin budgets, so eligibility does not guarantee that a level finishes a rewrite in that pass.

Level Default size range
L0 Less than 128 MiB
L1 128 MiB to 1 GiB
L2 1 GiB to 8 GiB
L3 More than 8 GiB

The level comes from the file size in the Iceberg manifest. It is not stored as separate file metadata.

Variable Default Meaning
SIGLAKE_COMPACTOR_LEVEL_CEILINGS_MB 128,1024,8192 Boundaries between levels.
SIGLAKE_COMPACTOR_LEVEL_TRIGGER_FILES 8 Rewrite-eligible files required before a nonterminal level is eligible.
SIGLAKE_COMPACTOR_LEVEL_MAX_FANIN 64 Maximum inputs for one level merge.
SIGLAKE_COMPACTOR_LEVEL_MAX_GEN 4 Rewrite-generation limit. 0 disables it.
SIGLAKE_COMPACTOR_MAX_OVERLAP_DEPTH 12 Overlap-depth trigger. 0 disables it.
SIGLAKE_COMPACTOR_LEVEL_INTERVALS_SECS 10,60,3600 Check interval for each level.

Set SIGLAKE_COMPACTOR_LEVELED=0 to select the legacy strategy. The values off and false have the same effect.

Legacy flat whole-partition pass

The legacy strategy considers a whole partition without the level policy. It is available for compatibility, but its work grows with the partition.

The depth trigger

File counts can stay below the level trigger while several files still cover the same time range. If overlap depth exceeds SIGLAKE_COMPACTOR_MAX_OVERLAP_DEPTH, the compactor selects the deepest overlapping cluster for a rewrite.

Depth-triggered bins group adjacent files in time. They can exceed the normal byte target, but they still obey the fan-in limit. This lets a rewrite reduce the overlap that caused it instead of joining distant ranges.

Graded backpressure

When the sealed WAL backlog exceeds SIGLAKE_COMPACTOR_MAX_SEALED_FOR_RECLUSTER, which defaults to 256, the drain gets priority. The compactor skips most maintenance cycles while the backlog remains high.

Skipping every rewrite would let file counts grow without a bound. Every fourth backlogged cycle therefore runs a restricted compaction pass by default. SIGLAKE_COMPACTOR_BACKPRESSURE_COMPACT_EVERY changes that cadence; 0 disables backlogged compaction. The restricted pass considers L0 files and uses a separate bin-concurrency budget.

These counters distinguish skipped and restricted cycles:

  • siglake_compactor_maintenance_skipped_backpressure_total
  • siglake_compactor_maintenance_throttled_backpressure_total

The live-file and overlap gauges come from table metadata walks. Check siglake_table_gauges_sampled_at_seconds before treating them as current.

Write-amplification bound

Each rewrite increments the generation in the file name. A file at SIGLAKE_COMPACTOR_LEVEL_MAX_GEN leaves normal level selection. The compactor can still rewrite it when it overlaps another file, because overlap would otherwise keep ordered-read cost high.

This policy limits repeated work for disjoint files while allowing extra work to restore the time layout. Set the limit to 0 if you do not want a generation cap.

Bounded merges

The storage crate has three merge paths:

  1. An in-memory merge for inputs that fit its decoded-size guard.
  2. A streaming k-way merge for larger bins within the fan-in limit.
  3. A page-bounded plan merge for bins that need bounded decoding.

The page-bounded path reads timestamps first, builds a run plan, then reads selected rows in chunks. Its decoded memory depends on the chunk size instead of the total input size.

What each merge path writes

Which search metadata a rewrite writes depends on the path it took. The in-memory merge holds the whole output batch, so it writes everything the enabled settings ask for: the group-count and time-bucket footers, the row-group token blooms, the whole-file raw trigram bloom, and an inverted index for each column the enabled specifications name. That index goes in the Parquet footer unless it exceeds SIGLAKE_INDEX_FOOTER_MAX_BYTES, which defaults to 1 MiB per column. Inverted indexes need compactor.invertedIndex.enabled, which is on by default. See Search acceleration.

The two streamed paths write the group-count, time-bucket and raw row-group bloom footers. They write neither a footer inverted index nor the whole-file raw trigram bloom, because both are computed from the whole decoded batch that a streamed path declines to hold. A delete-task rewrite splits the same way: past its own caps, its output is streamed too.

A second pass can run after a rewrite commits, and registers Puffin sidecars for the rewritten files whose columns have neither a footer index nor a registration already. That pass is off by default and opts in under compactor.indexRebuild: true, so streamed output stays unindexed until you turn it on, and a search over it scans those files for exact rows. Its registrations survive expiry of the snapshot they were made on. Reading is unaffected either way: an index a file already carries is still discovered and used.

The rebuild restores inverted indexes and nothing else. Its sidecars do not stand in for the inline footer index the streamed write skipped, and no pass adds the whole-file raw trigram bloom after the fact. Until a later in-memory rewrite covers the file, LIKE '%substr%' on raw prunes against the per-row-group blooms alone.

Turning either switch back on backfills nothing already committed. A rebuild only ever sees the files of the rewrite it follows, and only a rewrite of the file itself writes an inline footer index or the whole-file bloom. Files committed while a switch was off stay as they are until some later pass rewrites them.

index_at_flush controls ingest-time work and does not defer indexing during a compaction rewrite. A delete-task rewrite is the exception: it writes at rewrite generation 0, like an ingest flush, so on a table with index_at_flush: false even its in-memory arm skips the inline index work.

Every gap here costs pruning, not correctness. A file without an index or a bloom is scanned and row-evaluated, and the query returns exact rows.

Maintenance on the same loop

The maintenance role can run file rewrites, snapshot expiry, aggregate folding, retention, delete tasks, and table gauge sampling. Features that can delete data or add index work have separate switches.

The delete-task switch is on: the chart ships compactor.deleteTasks: true, and the sweep runs in the idle cycle unless SIGLAKE_DELETE_TASKS is 0, off, false or no. deleteTasks: false renders SIGLAKE_DELETE_TASKS=0; the operator's CR has no field for it, so the opt-out is spec.extraEnv. Turning the sweep off does not reject deletion requests: they are still recorded, and siglake delete-sweep --apply still executes them by hand.

The Helm chart leaves compactor.catalogClaim off by default. Footer inverted indexes are on unless you set compactor.invertedIndex.enabled: false. The post-rewrite Puffin rebuild is the other way round: it runs only under compactor.indexRebuild: true. See Helm for the deployment settings.

Commit-accumulation batching

The drain batches sealed segments until their combined size reaches commitBatch.targetMb or the oldest segment reaches commitBatch.maxAgeSecs. The code and Helm defaults are 32 MiB and 10 seconds. targetMb: 0 restores one commit per segment.

Batching reduces catalog commits, but it delays committed storage by up to the age setting. The optional WAL buffer can expose built-in events rows and managed user-index rows during that delay. External Iceberg readers wait for the commit.

Snapshot expiry trade-off

The Helm chart expires snapshots every 60 seconds and retains the latest 100. This bounds Iceberg metadata growth at the cost of older time travel. Set snapshotExpire.intervalSecs: 0 to disable expiry.

Snapshot expiry removes metadata references. Orphan garbage collection is the separate step that can delete files no retained snapshot references.

Convergence economics

Compaction exchanges write work for lower read amplification. Small files and overlapping time ranges disappear through repeated bounded passes rather than one global rewrite.

Convergence time depends on ingest rate, file sizes, available compactor resources, and the overlap already present. Use live-file counts, level gauges, and siglake_table_overlap_depth to judge the current table instead of assuming a fixed completion time.