Skip to content

Storage layout

Siglake stores committed rows in Apache Iceberg format-version-2 tables backed by Parquet files. Iceberg is the live storage format, so another compatible reader can query the tables without a Siglake export step.

Table layout

The built-in events table, each managed user index, and system tables have separate Iceberg tables. Siglake applies these layout rules when it owns the table:

  • day(timestamp) partitions
  • Zstandard level 3 compression in Parquet
  • a row-group target derived from measured row size and a byte budget
  • the table's declared timestamp sort order

Drain appends and compaction rewrites use the same Parquet writer. This keeps their sort metadata, encodings, and optional search metadata consistent.

The events.timestamp field is a required microsecond timestamp. Its required timestamp_ns sibling preserves the source nanoseconds. timestamp is the partition and first sort field; timestamp_ns breaks ties.

Iceberg format version 2, Parquet's writer version, and Siglake's own metadata versions are separate version numbers.

The catalog

Siglake uses an Iceberg SQL catalog. SQLite supports a local single-process catalog. PostgreSQL supports a catalog shared by several processes.

The workspace uses vendored Iceberg crates for catalog, scan, and storage behavior. The in-memory Iceberg catalog is not the local persistence path. See Local development for the repository setup.

The time-ordered invariant

Before splitting a batch into partition files, the writer sorts rows by the table's declared Iceberg sort order. It also writes Parquet SortingColumn metadata.

This order supports three later decisions. Manifest bounds can remove files from a time query. A compatible ordered query can stop after its LIMIT. Compaction can concatenate separate time ranges instead of merging every row.

New tables use ascending timestamp order. Older tables can declare descending order. The query provider reads the table and file sort metadata instead of assuming a direction. It can reverse decoding for a query in the other direction.

The source design record is docs/DESIGN_time_ordered_storage.md in the Siglake repository.

Search acceleration

Siglake writes search metadata with Parquet files or in registered Puffin sidecars. Query readers treat these artifacts as optional. If an artifact is missing, disabled, or unknown, the reader scans the relevant data.

Parquet-native per-column blooms

SIGLAKE_PARQUET_NATIVE_BLOOMS=on writes Parquet-native bloom filters for configured columns. This setting is off by default. It changes future writes; it does not add blooms to existing files on its own.

Trigram and token blooms over raw

A file-level trigram bloom can rule out files for substring searches on raw. Per-row-group token blooms can then rule out row groups inside a file.

The metadata keys are siglake.raw_trigram_bloom.v1 and siglake.raw_trigram_rowgroup_blooms.v1. Their payloads contain a magic value and version. A reader that does not recognize the payload ignores it and scans the data.

Inverted indexes

An inverted index maps terms in a configured text column to row positions. Small index payloads can live in Parquet metadata. Larger payloads use Puffin sidecars registered with the Iceberg snapshot.

The metadata name begins with siglake.inverted_index.v1. Siglake builds these indexes by default. Set SIGLAKE_INVERTED_INDEX=0, Helm compactor.invertedIndex.enabled: false, or operator spec.extraEnv: [{name: SIGLAKE_INVERTED_INDEX, value: "0"}] to opt out. Files without one remain queryable through scans.

Group-count and time-bucket footers

Per-file group counts and time buckets support exact aggregate paths. Group counts use a compact binary encoding stored in Parquet metadata. The reader can decode one requested column without decoding other columns in the payload.

The keys are siglake.group_counts.v1 and siglake.time_buckets.v1. These summaries can replace data-file reads only when their validity checks pass. Otherwise, the query server uses an exact fallback.

Layout metadata

siglake.layout.v1 records layout information used by readers. The file name records rewrite generation. The scheduler derives compaction levels from file sizes in the manifest rather than storing a mutable level on each file.

Side-object aggregates

Table-wide aggregate metadata lives beside the table instead of inside every Iceberg snapshot summary. This avoids growing metadata.json with cumulative group counts that the commit path must read and rewrite.

Object Purpose
metadata/siglake-aggregates.json Current inline aggregate values and time aggregates.
metadata/siglake-agg-deltas/<seq>.json One commit's group-count contribution.
metadata/siglake-agg-wide.json Deltas folded by the maintenance compactor.
metadata/siglake-agg-deltas/<seq>.rebuild.json Request to rebuild a contribution after retry exhaustion.

The delta path adds an object write to a commit without requiring that commit to merge the wide aggregate. The maintenance compactor folds deltas later. SIGLAKE_SIDE_AGG_WRITE_BEHIND=1 can also move inline aggregate maintenance off the foreground commit path.

A reader checks a column's aggregate total against the snapshot record count. If they differ, it uses per-file metadata or scans the data. The response can report that fallback as served_by: "materialized".

A publication carries counts that exist nowhere else, so a failed one is retried on the delta write's budget: four attempts, 250, 500 and 750 ms apart. The retry is replay-safe, because the merge reads the object first and skips a publication whose coverage links are already there. A publication that spends all four attempts increments siglake_side_aggregate_publish_failures_total{table} and, where the incremental delta path is active, leaves the same rebuild marker a lost delta does, so the maintenance compactor restores the wide group counts. Nothing rebuilds the inline object, so its time aggregates stay short of the snapshot record count and windowed GROUP BY on that table answers from the per-file path until it is rebuilt. The chart's SiglakeSideAggregatePublicationLost alert fires on that counter.

File naming

Compacted files follow this pattern:

siglake-g<N>-<uuid>.parquet

<N> is the rewrite generation. The compactor uses it to stop selecting a disjoint file after the configured generation limit. It can still select the file to remove a time overlap.

How the storage format is versioned, and what happens on upgrade

Three numbers move independently: the Iceberg table format version, which is 2 for every table Siglake creates; the Parquet writer version; and the version of each Siglake metadata artifact.

An artifact carries its version in the name (siglake.group_counts.v1), in a payload magic, or both. A reader that meets an unknown version ignores the artifact and falls back to an exact path or a scan. No data file is rewritten because a format changes; compaction writes the current versions when it rewrites a file for layout.

What happens on upgrade depends on which version changes:

  • A new artifact version: nothing. Old files stay readable and converge as compaction rewrites them.
  • A new declared column: run siglake migrate-schema before the new binaries serve traffic. The Helm chart runs it in a pre-upgrade hook.
  • An incompatible Iceberg type: recreate the table and re-ingest. Iceberg cannot change a column's precision in place. The 2026-09-06 timestamp contract is the example: it moved timestamp from a nanosecond Iceberg type, which is format-version-3 only and closed to v2 readers, to a microsecond timestamptz with the required timestamp_ns long beside it. Warehouses written before that change are format version 3 and have no upgrade path.

The registry of format keys and versions is docs/DESIGN_file_formats.md in the source tree. If you are changing a persisted format rather than living with one, the writer rules are in Change a persisted format.

Schema evolution

If code tries to populate a declared column that the stored table lacks, the write fails and names the migration command. This trades an ingest refusal for protection against silently dropping the new value.

siglake migrate-schema adds columns declared by the running build and absent from the table. It does not drop, rename, reorder, or retype columns. Older data files remain readable and return null for a newly added nullable field.

The command can inspect all tables and tenant namespaces, and it supports a dry run. The Helm chart runs it in a pre-upgrade hook by default. Operators follow the Helm upgrade procedure.

Metadata hygiene

Iceberg keeps snapshots and files until maintenance removes their references or data. Siglake separates these operations:

Operation Effect
Snapshot expiry Removes old snapshot metadata while retaining the configured recent snapshots.
Orphan garbage collection Deletes files that no retained snapshot references after a safety age.
Retention sweep Removes data files outside an index's retention period.
Delete-task sweep Rewrites files to apply queued predicate deletes.

siglake gc-orphans, siglake retention-sweep, and siglake delete-sweep are dry runs unless --apply is present. siglake audit-rotate has different defaults and can delete audit history. Read Retention and deletes before running these commands.