Storage layout¶
Siglake stores committed rows in Apache Iceberg format-version-2 tables backed by Parquet files. Iceberg is the live storage format, so another compatible reader can query the tables without a Siglake export step.
Table layout¶
The built-in events table, each managed user index, and system tables have
separate Iceberg tables. Siglake applies these layout rules when it owns the
table:
day(timestamp)partitions- Zstandard level 3 compression in Parquet
- a row-group target derived from measured row size and a byte budget
- the table's declared timestamp sort order
Drain appends and compaction rewrites use the same Parquet writer. This keeps their sort metadata, encodings, and optional search metadata consistent.
The events.timestamp field is a required microsecond timestamp. Its required
timestamp_ns sibling preserves the source nanoseconds. timestamp is the
partition and first sort field; timestamp_ns breaks ties.
Iceberg format version 2, Parquet's writer version, and Siglake's own metadata versions are separate version numbers.
The catalog¶
Siglake uses an Iceberg SQL catalog. SQLite supports a local single-process catalog. PostgreSQL supports a catalog shared by several processes.
The workspace uses vendored Iceberg crates for catalog, scan, and storage behavior. The in-memory Iceberg catalog is not the local persistence path. See Local development for the repository setup.
The time-ordered invariant¶
Before splitting a batch into partition files, the writer sorts rows by the
table's declared Iceberg sort order. It also writes Parquet SortingColumn
metadata.
This order supports three later decisions. Manifest bounds can remove files
from a time query. A compatible ordered query can stop after its LIMIT.
Compaction can concatenate separate time ranges instead of merging every row.
New tables use ascending timestamp order. Older tables can declare descending order. The query provider reads the table and file sort metadata instead of assuming a direction. It can reverse decoding for a query in the other direction.
The source design record is docs/DESIGN_time_ordered_storage.md in the
Siglake repository.
Search acceleration¶
Siglake writes search metadata with Parquet files or in registered Puffin sidecars. Query readers treat these artifacts as optional. If an artifact is missing, disabled, or unknown, the reader scans the relevant data.
Parquet-native per-column blooms¶
SIGLAKE_PARQUET_NATIVE_BLOOMS=on writes Parquet-native bloom filters for
configured columns. This setting is off by default. It changes future writes;
it does not add blooms to existing files on its own.
Trigram and token blooms over raw¶
A file-level trigram bloom can rule out files for substring searches on
raw. Per-row-group token blooms can then rule out row groups inside a file.
The metadata keys are siglake.raw_trigram_bloom.v1 and
siglake.raw_trigram_rowgroup_blooms.v1. Their payloads contain a magic value
and version. A reader that does not recognize the payload ignores it and scans
the data.
Inverted indexes¶
An inverted index maps terms in a configured text column to row positions. Small index payloads can live in Parquet metadata. Larger payloads use Puffin sidecars registered with the Iceberg snapshot.
The metadata name begins with siglake.inverted_index.v1. Siglake builds these
indexes by default. Set SIGLAKE_INVERTED_INDEX=0, Helm
compactor.invertedIndex.enabled: false, or operator
spec.extraEnv: [{name: SIGLAKE_INVERTED_INDEX, value: "0"}] to opt out.
Files without one remain queryable through scans.
Group-count and time-bucket footers¶
Per-file group counts and time buckets support exact aggregate paths. Group counts use a compact binary encoding stored in Parquet metadata. The reader can decode one requested column without decoding other columns in the payload.
The keys are siglake.group_counts.v1 and siglake.time_buckets.v1. These
summaries can replace data-file reads only when their validity checks pass.
Otherwise, the query server uses an exact fallback.
Layout metadata¶
siglake.layout.v1 records layout information used by readers. The file name
records rewrite generation. The scheduler derives compaction levels from file
sizes in the manifest rather than storing a mutable level on each file.
Side-object aggregates¶
Table-wide aggregate metadata lives beside the table instead of inside every
Iceberg snapshot summary. This avoids growing metadata.json with cumulative
group counts that the commit path must read and rewrite.
| Object | Purpose |
|---|---|
metadata/siglake-aggregates.json |
Current inline aggregate values and time aggregates. |
metadata/siglake-agg-deltas/<seq>.json |
One commit's group-count contribution. |
metadata/siglake-agg-wide.json |
Deltas folded by the maintenance compactor. |
metadata/siglake-agg-deltas/<seq>.rebuild.json |
Request to rebuild a contribution after retry exhaustion. |
The delta path adds an object write to a commit without requiring that commit
to merge the wide aggregate. The maintenance compactor folds deltas later.
SIGLAKE_SIDE_AGG_WRITE_BEHIND=1 can also move inline aggregate maintenance
off the foreground commit path.
A reader checks a column's aggregate total against the snapshot record count.
If they differ, it uses per-file metadata or scans the data. The response can
report that fallback as served_by: "materialized".
A publication carries counts that exist nowhere else, so a failed one is
retried on the delta write's budget: four attempts, 250, 500 and 750 ms apart.
The retry is replay-safe, because the merge reads the object first and skips a
publication whose coverage links are already there. A publication that spends
all four attempts increments
siglake_side_aggregate_publish_failures_total{table} and, where the
incremental delta path is active, leaves the same rebuild marker a lost delta
does, so the maintenance compactor restores the wide group counts. Nothing
rebuilds the inline object, so its time aggregates stay short of the snapshot
record count and windowed GROUP BY on that table answers from the per-file
path until it is rebuilt. The chart's SiglakeSideAggregatePublicationLost
alert fires on that counter.
File naming¶
Compacted files follow this pattern:
<N> is the rewrite generation. The compactor uses it to stop selecting a
disjoint file after the configured generation limit. It can still select the
file to remove a time overlap.
How the storage format is versioned, and what happens on upgrade¶
Three numbers move independently: the Iceberg table format version, which is 2 for every table Siglake creates; the Parquet writer version; and the version of each Siglake metadata artifact.
An artifact carries its version in the name (siglake.group_counts.v1), in a
payload magic, or both. A reader that meets an unknown version ignores the
artifact and falls back to an exact path or a scan. No data file is rewritten
because a format changes; compaction writes the current versions when it
rewrites a file for layout.
What happens on upgrade depends on which version changes:
- A new artifact version: nothing. Old files stay readable and converge as compaction rewrites them.
- A new declared column: run
siglake migrate-schemabefore the new binaries serve traffic. The Helm chart runs it in a pre-upgrade hook. - An incompatible Iceberg type: recreate the table and re-ingest. Iceberg
cannot change a column's precision in place. The 2026-09-06 timestamp
contract is the example: it moved
timestampfrom a nanosecond Iceberg type, which is format-version-3 only and closed to v2 readers, to a microsecondtimestamptzwith the requiredtimestamp_nslong beside it. Warehouses written before that change are format version 3 and have no upgrade path.
The registry of format keys and versions is docs/DESIGN_file_formats.md in
the source tree. If you are changing a persisted format rather than living
with one, the writer rules are in Change a persisted
format.
Schema evolution¶
If code tries to populate a declared column that the stored table lacks, the write fails and names the migration command. This trades an ingest refusal for protection against silently dropping the new value.
siglake migrate-schema adds columns declared by the running build and absent
from the table. It does not drop, rename, reorder, or retype columns. Older
data files remain readable and return null for a newly added nullable field.
The command can inspect all tables and tenant namespaces, and it supports a dry run. The Helm chart runs it in a pre-upgrade hook by default. Operators follow the Helm upgrade procedure.
Metadata hygiene¶
Iceberg keeps snapshots and files until maintenance removes their references or data. Siglake separates these operations:
| Operation | Effect |
|---|---|
| Snapshot expiry | Removes old snapshot metadata while retaining the configured recent snapshots. |
| Orphan garbage collection | Deletes files that no retained snapshot references after a safety age. |
| Retention sweep | Removes data files outside an index's retention period. |
| Delete-task sweep | Rewrites files to apply queued predicate deletes. |
siglake gc-orphans, siglake retention-sweep, and
siglake delete-sweep are dry runs unless --apply is present.
siglake audit-rotate has different defaults and can delete audit history.
Read Retention and deletes before running
these commands.