Storage layout¶
Siglake stores committed rows in Apache Iceberg format-version-2 tables backed by Parquet files. Iceberg is the live storage format, so another compatible reader can query the tables without a Siglake export step.
Table layout¶
The built-in events table, each managed user index, and system tables have
separate Iceberg tables. Siglake applies these layout rules when it owns the
table:
day(timestamp)partitions- Zstandard level 3 compression in Parquet
- a row-group target sized in rows from a byte budget and a measured row size
- the table's declared timestamp sort order
Drain appends and compaction rewrites use the same Parquet writer. This keeps their sort metadata, encodings, and optional search metadata consistent.
The events.timestamp field is a required microsecond timestamp. Its required
timestamp_ns sibling preserves the source nanoseconds. timestamp is the
partition and first sort field; timestamp_ns breaks ties.
Iceberg format version 2, Parquet's writer version, and Siglake's own metadata versions are separate version numbers.
Two estimators read one row-group target¶
SIGLAKE_PARQUET_TARGET_ROW_GROUP_BYTES (256 MiB by default) is one number, and
the write paths read it two ways. Each path divides the target by a row size it
measures on the first batch it writes, then clamps the result to 128 Ki through
4 Mi rows. The target sets a row count. It is not a ceiling on resident bytes.
| Write path | Row size measured as | Fixture measurement |
|---|---|---|
| Ingest flush | whole Arrow buffer allocations | 432 B/row |
| Leveled merge, re-cluster, delete rewrite | the extent the rows span | 201 B/row |
Flush prices allocations because it hands the writer a batch it assembled itself, so the allocation is that batch. The merge paths price extent because a page-bounded merge writes zero-copy slices of a decoded input part, and pricing allocations would charge a 2,000-row slice for the whole 8,192-row part behind it. The two measures differ by about 2x on the same data, so one target asks for roughly half as many rows on flush as on merge. The figures above come from a 512 Ki-row half-deleted fixture.
The open row group is buffered decoded, and a buffered batch keeps whole buffers. Where the merge priced those rows by extent, the bytes resident for them run over the target: on that fixture, by about twice. We keep the split deliberately (decided 2026-09-16). Moving flush onto the extent measure would change ingest output layout, and that side is unmeasured.
The catalog¶
Siglake uses an Iceberg SQL catalog. SQLite supports a local single-process catalog. PostgreSQL supports a catalog shared by several processes.
The workspace uses vendored Iceberg crates for catalog, scan, and storage behavior. The in-memory Iceberg catalog is not the local persistence path. See Local development for the repository setup.
The time-ordered invariant¶
Before splitting a batch into partition files, the writer sorts rows by the
table's declared Iceberg sort order. It also writes Parquet SortingColumn
metadata.
This order supports three later decisions. Manifest bounds can remove files
from a time query. A compatible ordered query can stop after its LIMIT.
Compaction can concatenate separate time ranges instead of merging every row.
New tables use ascending timestamp order. Older tables can declare descending order. The query provider reads the table and file sort metadata instead of assuming a direction. It can reverse decoding for a query in the other direction.
The source design record is docs/DESIGN_time_ordered_storage.md in the
Siglake repository.
Search acceleration¶
Siglake writes search metadata with Parquet files or in registered Puffin sidecars. Query readers treat these artifacts as optional. If an artifact is missing, disabled, or unknown, the reader scans the relevant data.
An artifact the reader read but could not trust is handled per storage path. A footer inverted index whose checksum fails is refused and the file is answered by another index or an exact scan. A Puffin blob whose Zstd frame fails its content checksum fails the query instead. See What the footer text-index checksum covers and Search v1 bounds.
Parquet-native per-column blooms¶
SIGLAKE_PARQUET_NATIVE_BLOOMS=on writes Parquet-native bloom filters for
configured columns. This setting is off by default. It changes future writes;
it does not add blooms to existing files on its own.
Trigram and token blooms over raw¶
A file-level trigram bloom can rule out files for substring searches on
raw. Per-row-group token blooms can then rule out row groups inside a file.
The metadata keys are siglake.raw_trigram_bloom.v1 and
siglake.raw_trigram_rowgroup_blooms.v1. Their payloads contain a magic value
and version. A reader that does not recognize the payload ignores it and scans
the data.
Inverted indexes¶
An inverted index maps terms in a configured text column to row positions. Small index payloads can live in Parquet metadata. Larger payloads use Puffin sidecars registered with the Iceberg snapshot.
The metadata name begins with siglake.inverted_index.v1. Siglake builds these
indexes by default. Set SIGLAKE_INVERTED_INDEX=0, Helm
compactor.invertedIndex.enabled: false, or operator
spec.extraEnv: [{name: SIGLAKE_INVERTED_INDEX, value: "0"}] to opt out.
Files without one remain queryable through scans.
From Siglake 0.2.0, a footer payload is written beside a CRC-32 of its bytes,
under siglake.inverted_index.crc32.v1 with the same per-column suffix. A
reader verifies the checksum before it uses the index, cold or warm, and
refuses an index that fails: it tries the column's Puffin index and otherwise
scans the file. A file written earlier has no checksum, and a query prunes with
its index as before. A reader that does not recognize the name ignores it.
Refusals are counted by siglake_index_footer_checksum_refused_total; see
Footer text-index checksum
refusals.
Segmented text indexes (seg2)¶
A segmented index (seg2) holds the same terms and row positions as a v1 inverted index, split into blocks behind a directory, so a query reads the part it needs instead of parsing the file's whole index. Siglake 0.2.0 adds the format behind two switches that move independently:
SIGLAKE_SEGMENTED_INDEX_WRITES=1 # compactor: build seg2 during a re-cluster
SIGLAKE_SEGMENTED_INDEX_READS=1 # query: answer text predicates from seg2
Only the compactor's streaming re-cluster writes seg2, one index group per Parquet row group, registered in the same commit as the rewritten files. The in-memory merge arm, ingest flush and delete-task rewrites keep writing v1.
The metadata name begins with siglake.inverted_index.seg2 and the Puffin
blob type is siglake-inverted-seg-v2. Neither collides with the v1 names, and
the format is per-file metadata rather than table state, so one table can hold
v1 files, seg2 files and files with neither. There is no migration: turning
writes on changes what the next re-cluster produces, and nothing already
committed. A query uses whichever format a file carries and scans a file with
neither, which costs pruning rather than correctness. Metadata from the
retired seg1 prototype counts as neither.
Both switches are off by default in 0.2.0. The published measurements come
from local file:// runs, so treat the format as experimental until it is
qualified against object storage.
Group-count and time-bucket footers¶
Per-file group counts and time buckets support exact aggregate paths. Group counts use a compact binary encoding stored in Parquet metadata. The reader can decode one requested column without decoding other columns in the payload.
The keys are siglake.group_counts.v1 and siglake.time_buckets.v1. These
summaries can replace data-file reads only when their validity checks pass.
Otherwise, the query server uses an exact fallback.
Layout metadata¶
siglake.layout.v1 records layout information used by readers. The file name
records rewrite generation. The scheduler derives compaction levels from file
sizes in the manifest rather than storing a mutable level on each file.
Side-object aggregates¶
Table-wide aggregate metadata lives beside the table instead of inside every
Iceberg snapshot summary. This avoids growing metadata.json with cumulative
group counts that the commit path must read and rewrite.
| Object | Purpose |
|---|---|
metadata/siglake-aggregates.json |
Current inline aggregate values and time aggregates. |
metadata/siglake-agg-deltas/<seq>.json |
One commit's group-count contribution. |
metadata/siglake-agg-wide.json |
Deltas folded by the maintenance compactor. |
metadata/siglake-agg-deltas/<seq>.rebuild.json |
Request to rebuild a contribution after retry exhaustion. |
The delta path adds an object write to a commit without requiring that commit
to merge the wide aggregate. The maintenance compactor folds deltas later.
SIGLAKE_SIDE_AGG_WRITE_BEHIND=1 can also move inline aggregate maintenance
off the foreground commit path.
A reader checks a column's aggregate total against the snapshot record count.
If they differ, it uses per-file metadata or scans the data. The response can
report that fallback as served_by: "materialized".
A publication carries counts that exist nowhere else, so a failed one is
retried on the delta write's budget: four attempts, 250, 500 and 750 ms apart.
The retry is replay-safe, because the merge reads the object first and skips a
publication whose coverage links are already there. A publication that spends
all four attempts increments
siglake_side_aggregate_publish_failures_total{iceberg_namespace,table} and,
where the incremental delta path is active, leaves the same rebuild marker a
lost delta does, so the maintenance compactor restores the wide group counts. Nothing
rebuilds the inline object automatically, so its time aggregates stay short of
the snapshot record count and windowed GROUP BY on that table answers from the
per-file path until it is rebuilt. The chart's
SiglakeSideAggregatePublicationLost alert fires on that counter.
That counter marks an event. The same state also arrives without a failed
publication: retention, a delete task, a foreign overwrite and the residual
windows at snapshot expiry all leave an inline object that cannot prove
coverage of the current snapshot, and no commit repairs it. The maintenance
compactor censuses each maintained table for that state every 15 minutes by
default, reports it on
siglake_inline_coverage_unproven{iceberg_namespace,table}, and pages on it
with the critical SiglakeInlineCoverageUnproven. The census reads; it rebuilds
nothing. The repair is on the operator page, under
Data loss or durable inconsistency.
File naming¶
Compacted files follow this pattern:
<N> is the rewrite generation. The compactor uses it to stop selecting a
disjoint file after the configured generation limit. It can still select the
file to remove a time overlap.
How the storage format is versioned, and what happens on upgrade¶
Three numbers move independently: the Iceberg table format version, which is 2 for every table Siglake creates; the Parquet writer version; and the version of each Siglake metadata artifact.
An artifact carries its version in the name (siglake.group_counts.v1), in a
payload magic, or both. A reader that meets an unknown version ignores the
artifact and falls back to an exact path or a scan. No data file is rewritten
because a format changes; compaction writes the current versions when it
rewrites a file for layout.
What happens on upgrade depends on which version changes:
- A new artifact version: nothing. Old files stay readable and converge as compaction rewrites them.
- A new declared column: run
siglake migrate-schemabefore the new binaries serve traffic. The Helm chart runs it in a pre-upgrade hook. - An incompatible Iceberg type: recreate the table and re-ingest. Iceberg
cannot change a column's precision in place. The 2026-09-06 timestamp
contract is the example: it moved
timestampfrom a nanosecond Iceberg type, which is format-version-3 only and closed to v2 readers, to a microsecondtimestamptzwith the requiredtimestamp_nslong beside it. Warehouses written before that change are format version 3 and have no upgrade path.
The registry of format keys and versions is docs/DESIGN_file_formats.md in
the source tree. If you are changing a persisted format rather than living
with one, the writer rules are in Change a persisted
format.
Schema evolution¶
If code tries to populate a declared column that the stored table lacks, the write fails and names the migration command. This trades an ingest refusal for protection against silently dropping the new value.
siglake migrate-schema adds columns declared by the running build and absent
from the table. It does not drop, rename, reorder, or retype columns. Older
data files remain readable and return null for a newly added nullable field.
The command can inspect all tables and tenant namespaces, and it supports a dry run. The Helm chart runs it in a pre-upgrade hook by default. Operators follow the Helm upgrade procedure.
Metadata hygiene¶
Iceberg keeps snapshots and files until maintenance removes their references or data. Siglake separates these operations:
| Operation | Effect |
|---|---|
| Snapshot expiry | Removes old snapshot metadata while retaining the configured recent snapshots. |
| Orphan garbage collection | Deletes files that no retained snapshot references after a safety age. |
| Retention sweep | Removes data files outside an index's retention period. |
| Delete-task sweep | Rewrites files to apply queued predicate deletes. |
siglake gc-orphans, siglake retention-sweep, and
siglake delete-sweep are dry runs unless --apply is present.
siglake audit-rotate has different defaults and can delete audit history.
Read Retention and deletes before running
these commands.