Scope and limitations¶
Siglake provides the telemetry storage and query layer: ingestion, durable WAL storage, open Iceberg tables, distributed SQL, and interfaces for external consumers. This page separates intentional product scope from configuration choices and implementation constraints, so you can plan a deployment around the contracts it supports.
The implementation constraints on this page summarize
docs/LIMITATIONS.md
in the Siglake repository at release commit 65c4eef.
Product scope¶
User interface, alerting interface, dashboards¶
Siglake integrates with external presentation and alerting tools through its
SQL and HTTP APIs, open tables, and WAL consumer interface. Use Grafana over
/api/v1/sql and the Jaeger shim for dashboards and trace views, as described
in Grafana and Jaeger.
Dashboard authoring, alert evaluation and delivery, acknowledgement, and on-call workflows belong in those external tools. They are outside the core storage and query product's scope, rather than unfinished core components.
Relevance scoring (BM25)¶
Log exploration uses time ordering and SQL filters. Siglake does not apply
BM25 relevance scores or scored TopK ranking. Term frequency is often less
useful for one-line templated logs, where the same message repeats many times.
The source design docs/DESIGN_bm25_scoring.md records a possible alternative;
it is not the current search contract.
Deployment planning¶
Account for these operating contracts when choosing a topology:
- Compaction is single-pod unless you enable the catalog claim.
- Multi-tenancy is soft: tenants share compute, the catalog, and the bucket.
- Retention is file- and day-granular, not row-level.
- With the default metadata-cache settings, query replicas can disagree across requests for up to 60 seconds.
- Query spill uses bounded node-local scratch space.
- There is no qualified rollback to an older image.
- A warehouse that uses the pre-0.1.0 v3 and nanosecond timestamp contract must be recreated instead of migrated.
- The Elasticsearch-compatible API is write-only.
Features off by default in Helm¶
These options let you choose freshness, scaling, resource use, and cluster integrations explicitly. Enable the ones your deployment needs.
| Feature | Setting | Default behavior and configuration choice |
|---|---|---|
| Query WAL buffer | query.walBuffer.enabled: false |
Queries read committed data. Enable the buffer for visibility into sealed WAL segments before commit. |
| Catalog claim | compactor.catalogClaim.enabled: false |
Single-pod compaction only. The chart refuses to render a compactor tier that can hold more than one pod. |
| Attribute auto-promotion | SIGLAKE_AUTO_PROMOTE_MIN_PCT=0 |
Automatic promotion widens the schema from the compactor's sample, and that widening cannot be undone. Hot keys stay in the JSON blob unless declared explicitly. |
| KEDA autoscaling | keda.enabled: false |
Replica counts are configured directly. Enable KEDA for metric-driven autoscaling. |
| Prometheus alert rules | prometheusRule.enabled: false |
Enable the bundled rules with the Prometheus Operator and an external alert-delivery workflow. |
| Source-file batch cache | query.scan.fileCacheMaxBytes: 0 and query.scan.fileCacheMaxEntries: 0 |
The byte limit is reserved from the query memory pool, so a non-zero default would reduce headroom on pods that never reread a file. Warm repeat reads do not reuse decoded source batches. Set both limits to positive values to enable this experimental cache. |
| ServiceMonitor, PDBs, NetworkPolicies, anti-affinity | all false |
No Prometheus Operator integration, no disruption budgets, no network isolation. |
A query pod at the 4 GiB memory floor also assigns no cache to parsed or
serialized text indexes by default. Each text query deserializes the index
again unless you set SIGLAKE_PARSED_INDEX_CACHE_MAX_BYTES and
SIGLAKE_PUFFIN_BLOB_CACHE_MAX_BYTES yourself. A 5 GiB limit gives both
derived cache budgets some capacity. A pod in that state shows it: every index
acquisition is a miss on
siglake_iceberg_parsed_index_cache_lookups_total{outcome} with no eviction
beside it. See Parsed text-index cache
outcomes.
WAL mirroring is enabled by default when you configure a warehouse URL. Each
sealed segment uploads asynchronously to <warehouse-url>/wal-mirror/.
An acknowledgement is durable on the local WAL after the default fsync(2);
the rows get a copy in object storage when their Iceberg commit or that upload
completes, whichever comes first. Mirroring shortens the pre-commit window, so
losing the WAL volume costs you the acknowledged rows that have been neither
committed nor uploaded.
The default single-replica compactor does not reclaim mirror objects, so give
the prefix an object-store lifecycle expiry unless you enable
compactor.catalogClaim.enabled. Set wal.mirror.enabled: false to opt out.
Leveled compaction is enabled by default and therefore is not listed above.
Set SIGLAKE_COMPACTOR_LEVELED=0 only to select the legacy flat
whole-partition pass.
Footer inverted indexes are enabled by default. Set
SIGLAKE_INVERTED_INDEX=0, Helm
compactor.invertedIndex.enabled: false, or operator
spec.extraEnv: [{name: SIGLAKE_INVERTED_INDEX, value: "0"}] to opt out.
The post-rewrite Puffin rebuild is disabled by default. A rewrite that
could not put an inverted index in the Parquet footer, which is every streamed
merge and any column whose index exceeds SIGLAKE_INDEX_FOOTER_MAX_BYTES,
leaves that output unindexed, and queries scan it for exact rows. Set
SIGLAKE_INDEX_REBUILD=1, Helm compactor.indexRebuild: true, or operator
spec.extraEnv: [{name: SIGLAKE_INDEX_REBUILD, value: "1"}] to opt in: the
pass then registers a Puffin sidecar after the rewrite commits, skipping files
whose columns already carry a footer index or a registration. With inverted
indexes off it does nothing either way. Reads do not depend on the switch.
Indexes that already exist are still used, and the flush path still writes its
footer inverted index.
The rebuild covers inverted indexes only. A streamed merge also skips the
whole-file raw trigram bloom, and no later pass registers one, so
LIKE '%substr%' on raw prunes against per-row-group blooms alone until an
in-memory rewrite covers the file. A delete-task rewrite writes at rewrite
generation 0, like an ingest flush, so on a table with index_at_flush: false
even its in-memory arm skips the inline index work.
Turning either switch back on backfills nothing already committed: a rebuild only ever sees the files of the rewrite it follows, and only a rewrite of the file itself writes an inline footer index or the whole-file bloom. Files committed while a switch was off keep the metadata they were written with. The cost throughout is pruning rather than correctness: an unindexed file is scanned and row-evaluated, and the answer is exact. See What each merge path writes.
Segmented text indexes (seg2) are off at both ends in Siglake 0.2.0.
SIGLAKE_SEGMENTED_INDEX_WRITES=1 makes the compactor's streaming re-cluster
write one, and SIGLAKE_SEGMENTED_INDEX_READS=1 lets a query read one; with
reads off, a file whose only text index is seg2 is scanned. Neither switch has
a chart value or an operator field, so set them through compactor.extraEnv
and query.extraEnv. The published measurements are local file:// runs, not
an object-store, distributed or HTTP result, which is why both defaults stay
off. See Segmented text
indexes.
Delete-task execution is on by default too. The compactor runs the sweep in its
idle cycle unless SIGLAKE_DELETE_TASKS is 0, off, false or no. The
chart ships compactor.deleteTasks: true and renders the variable either way,
so false renders SIGLAKE_DELETE_TASKS=0. The operator's compactor renders
SIGLAKE_DELETE_TASKS=1 and the CR has no field for it: put
SIGLAKE_DELETE_TASKS=0 in spec.extraEnv to opt out, which is also what
--adopt-values reports when the chart values file it reads set
compactor.deleteTasks: false. With the sweep off, POST /api/v1/delete-tasks
still records the task, and siglake delete-sweep --apply still executes it by
hand.
OTLP/gRPC is on by default as well. The ingester listens for logs and traces on
0.0.0.0:4317 next to OTLP/HTTP, and both transports apply the same
authentication, tenant routing and WAL-fsync acknowledgement contract. The
chart ships ingester.otlpGrpc.enabled: true with
ingester.otlpGrpc.port: 4317; setting enabled: false drops the gRPC
container and Service ports and passes --disable-otlp-grpc to the binary.
Outside Helm, leaving --otlp-grpc-listen unset no longer disables the
listener: pass --disable-otlp-grpc instead. The operator renders the listener at
every replica count, and its CR has no field to turn it off.
Persistent batch jobs are on by default (query.jobs.persistent: true). The
query pods keep batch-job state - submissions, their status and their result
rows - in the Postgres instance that already holds the Iceberg catalog. The
chart renders SIGLAKE_JOBS_POSTGRES_URI from the same Secret, so there is no
second connection to configure, and the query server creates its own tables on
start. The operator points an adopted query tier at its own
spec.catalogUri. One store for the tier is what makes a job_id usable at
the shipped query.replicas: 2: the Service has no session affinity, so a
status, result or cancel request lands on either pod.
Each job row records the replica incarnation executing it, and owners
heartbeat. If a query pod crashes, recovery fails its pending and running
rows once the owner's lease expires (SIGLAKE_JOBS_OWNER_LEASE_SECS, default
120 s); jobs owned by a sibling that is still heartbeating are left alone. A planned SIGINT or SIGTERM shutdown
deletes that pod's registration after the HTTP server drains, so the next
recovery pass can fail its remaining non-terminal rows without waiting out the
lease. Neither path resumes interrupted work: the client sees failed and has
to resubmit.
Setting query.jobs.persistent: false drops the variable and gives each pod
its own in-memory store. It is a supported opt-out for a tier you can hold at
one pod, and both deployment paths refuse the rest. The chart fails the install
when query.replicas, or keda.query.maxReplicas with keda.query.enabled:
true, is above 1 while the store is off. The operator refuses the spec with
QueryJobsStoreRequiredForScaleOut when spec.autoscaling.query.max is above
1 and the effective jobs-store URI is blank. Without those guards, the job
status, result and cancel reads that land on the pod which did not take the
submission answer 404 while the job runs normally elsewhere. The one-pod cost
remains: a pod restart loses every job in flight on it.
The catalog claim stays off by default, and it is a requirement rather than a recommendation once you scale the compactor. The chart fails the install on the three combinations that cannot work.
- More than one compactor pod with the claim off. That is
compactor.replicasabove1, orautoscaling.compactor.maxReplicasabove1whenautoscaling.compactor.enabledistrue. The shipped HPA ceiling is4, so turning the compactor HPA on is refused until the claim is on. Uncoordinated replicas share no claim table: each runs the whole drain and maintenance loop against the same tables, and they lose Iceberg commit races to each other while the layout stops converging. - The claim with
wal.mirror.enabled: false. In claim mode the drain reads thewal_segmentscatalog table and never the localsealed/directory, and rows land there only from an ingester that mirrors its segments. That compactor would claim nothing whilesiglake_compactor_sealed_pendingread zero, because in claim mode that gauge counts the shared queue of sealed, unclaimed catalog rows. It cannot reveal a mirrored object that has no catalog row. - The enabled compactor HPA custom metric with the claim on. The refusal needs
autoscaling.compactor.enabled,autoscaling.compactor.customMetric.enabledandcompactor.catalogClaim.enabledall set totrue. Its error starts withautoscaling.compactor.customMetric.enabled is true with compactor.catalogClaim.enabled. Every compactor reports the whole shared sealed queue rather than its own share, but the chart renders a per-pod average target. Disable the custom metric for a claim-mode CPU-only HPA, or use siglake-operator for backlog-driven scaling. The custom metric scales nothing in filesystem mode because that mode is capped at one pod.
All three refusals name the value to change. The operator needs none of these
chart combinations. It selects the drain from
spec.autoscaling.compactor.max, not the current replica count.
A maximum above 1 keeps the WAL mirror and catalog claim on at every current
replica count, including while the autoscaler holds the compactor at one. A
maximum of 1 keeps the filesystem drain.
Deployment constraints¶
- The operator cannot scale the ingester or query tier to zero. It accepts a
zero compactor floor only with the catalog-claim drain and positive smoothing.
The packaged compactor floor remains
1because no cluster round has yet proved that an ingester-published queue depth wakes a parked compactor. - Helm cannot run the embedded compactor. The chart refuses
ingester.extraArgs: [--with-compactor]even at one replica because a rolling update can overlap two ingesters running maintenance. Use the dedicated compactor tier.
Automatic operator drain-mode handover¶
The operator does not convert WAL between filesystem and catalog-claim drains.
Changing spec.autoscaling.compactor.max across 1 changes the ownership
protocol, so reconciliation reports DrainModeHandoverRequired and stops
before applying workloads. You must finish the current drain, account for local
WAL, mirror objects and catalog rows, then follow the manual handover
procedure. The operator
does not convert WAL or delete those segments and rows for you.
Aggregating merge kinds¶
No rollup or last-write-wins dedup merges exist. The exact rows-conserved commit guard is per-merge-kind-waivable by design, but nothing waives it yet.
Metrics downsampling¶
None. Retention is the only size-control mechanism.
Row-level retention¶
Retention drops whole data files whose manifest maximum timestamp is past the horizon. It is file- and day-granular. Your effective retention is the policy plus the span of the last surviving file.
Implementation constraints¶
Rows sharing a custom event time come back in an unspecified order¶
- A bare browse over an index orders newest-first on the field its mapping
names as
timestamp_field. - Equal custom event times have no implicit tiebreak. Which rows land at a
LIMITboundary can change between executions. - A deterministic tie order needs a second sort column declared on the index and named in the query.
Ingester autoscaling lacks live Prometheus evidence¶
- The operator's ingest-rate query depends on the
podlabel that Prometheus Operator adds during target relabeling. The chart does not add that label. - No retained run has confirmed the label and per-pod average with two active ingesters. Without the label, the query reads the fleet total as one pod's rate and can request too many replicas.
An audit batch abandoned at its deadline is lost¶
- Query responses do not wait for the best-effort audit worker. Its retained rows and conversion memory are bounded.
- Siglake 0.2.0 and later bound one append at 30 s
(
SIGLAKE_QUERY_AUDIT_APPEND_DEADLINE_SECS;0restores 0.1.0's unbounded await, where a stuck append held the whole retained budget until the process restarted). A batch that outlives the deadline is abandoned and never re-appended, so a catalog commit already on its way can land unseen. Those rows are counted bysiglake_query_audit_dropped_total{reason="append_deadline"}.
Concurrent index updates are not merged¶
- Two concurrent mapping updates that add different fields do not merge. The
losing request gets
400and must read the current mapping before it sends the update again. - Two writers updating the same index template are last-write-wins. A template
carries no version and takes no
If-Matchcondition, which 0.2.0 added for an index mapping alone. Deleting a template leaves a tombstone that is not garbage-collected.
External-reader evidence is local and narrow¶
- The Trino 483, Spark 3.5.9, DuckDB 1.5.5 and PyIceberg 0.12.0 demonstration
used a
file://warehouse and a SQLite catalog. It does not cover object storage, Postgres, or reads concurrent with compaction. - A later run checked decoded microsecond timestamps with Spark, DuckDB and PyIceberg once. Its retained record omits two reader versions and does not include Trino.
Query probes do not detect query-path degradation¶
The query server's /healthz is a constant 200 for as long as the process
can serve the probe. /readyz verifies only that the Iceberg catalog can be
reached; it does not execute a query or check the cache-warm loop. A pod can
therefore pass both probes while browse queries time out, so probe success is
not evidence that the query path is healthy.
The shipped signal for the known degradation is
SiglakeQueryWarmCycleStalled, not an automatic readiness failure or liveness
restart. See Monitoring for
the operational consequence. These semantics were confirmed by Siglake's
September 3, 2026 review of query degradation and probe behavior.
Tier-1 group counts require first-commit columns¶
The Tier-1 whole-table group-count aggregate initially covers only columns
present since the table's first commit. A table created before typed
measurements joined side aggregates can therefore keep serving GROUP BY on
those columns through the exact but slower materialized path, even when its
files already contain usable group-count footers.
Two older aggregate cases still need manual repair:
- Lost-delta repair does not admit an eligible typed column that predates the failed commit.
- Pre-coverage side aggregates are not adopted because they cannot prove which
snapshot they describe.
rebuild-group-countsrepairs the wide group-count object. No 0.1.0 command rebuilds the inline time aggregates.
The release after 0.1.0 adds siglake rebuild-time-aggregates, which narrows
the second case without closing it. The command recomputes a pre-coverage
object's time aggregates from the committed files and publishes a coverage edge
naming the snapshot it read. It does not certify the maps that were already
there. You run it by hand; no commit path repairs the object for you. It
repairs an existing inline object rather than creating one, because the column
set it would rebuild is recorded nowhere else. Its
publication drops the unproven whole-table group counts and restores only the
components that account for every row in the snapshot summary, so hourly
buckets short of that count, and any column short of it, are left out instead
of published partial. See Inline aggregate
repair for the counters it
increments.
Use the group-count repair
procedure
to admit eligible typed columns with --admit-typed-columns. This does not
require rewriting data files unless an older live file lacks both a usable
footer and a raw-page representation for the column.
Elasticsearch bulk API support is write-only¶
Siglake ingests over the Elasticsearch bulk API but has no Elasticsearch query
API, and none is planned. _bulk and _cluster/health work. _search,
_msearch, _search/scroll, _field_caps and _cat/* are registered and
answer 501 with a pointer to POST /api/v1/sql.
Existing Kibana or Elasticsearch query clients will not work. Query with SQL over HTTP instead: see the SQL cookbook.
Query replicas can briefly disagree across a commit¶
Each replica serves table metadata from its own cache. The default TTL is 5
seconds, and stale metadata can be served for up to 60 seconds while a
background refresh runs
(SIGLAKE_ICEBERG_METADATA_CACHE_TTL_SECS
controls the TTL). Two replicas behind one Service can therefore see different
snapshots during that window. Clients that need a consistent sequence of
queries must send them to one pod or wait for every replica's cache to
converge.
This is a cross-request limit only. Within one distributed query the fan-out
is pinned to the coordinator's generation, its snapshot id and schema id, and
a worker that cannot resolve that generation refuses its shard (503,
reason: "shard_pin_unresolved")
rather than contributing rows from another one, so a single answer is never
merged across generations. See One generation per
fan-out. The cost of that
guarantee is availability: while a coordinator's cache is stale on a snapshot
the catalog has already expired, its fanned-out queries return 503 until the
cache converges (/api/v1/sql/local still answers). That state alerts as
SiglakeQueryShardPinUnresolved.
Raise compactor.snapshotExpire.retainLast
to widen the margin, or lower SIGLAKE_ICEBERG_METADATA_CACHE_TTL_SECS so
coordinators stop serving a snapshot before the catalog expires it.
Jaeger is HTTP-only¶
The shim covers the HTTP subset Grafana's Jaeger data source renders. There is no gRPC SpanReader, so tools expecting that interface will not work.
Kubernetes operator gaps compared with Helm¶
It renders the core data plane with the same workload kinds as the chart: ingester and compactor as Deployments, query as a StatefulSet. The Helm chart remains the supported, more complete install path.
The operator also renders the WAL PVC, Services, retention CronJobs, and a
one-shot schema-migration Job when spec.schemaVersion is set. It does not
render Ingress, PDBs, NetworkPolicies, ServiceMonitors, HPA/KEDA objects, or
scheduling constraints. Per-tier resources default to the chart's and are
overridable through
spec.resources. The CR cannot
express query bearer tokens, OIDC, TLS, or a WAL-buffer volume.
Catalog credentials are plaintext in the spec.catalogUri custom resource.
The operator copies them into pod environment variables and CronJob arguments.
The chart instead keeps Secret references and an unexpanded URI in pod
specifications.
Adoption of an existing Helm release through --adopt-values is experimental
and offline. It has no live test coverage and leaves several values unmapped.
The ownership handover (annotation flip plus release-secret removal) is a
manual runbook.
Freshness does not cover system tables¶
The WAL buffer serves the events table and the managed user indexes a query
references. It does not serve query_audit or any other system table, which
see commit-cycle visibility. When the buffer is off, or a query pod cannot
read the WAL mount, every table falls back to commit-cycle visibility.
Strict mapping mode does not reject¶
"mode": "strict" document mapping enforces at commit time as
dynamic-plus-a-counter: it counts violations rather than rejecting documents,
and undeclared attributes still land in the residual attributes column. Do
not rely on it as a data-quality gate.
Search v1 bounds¶
- Full-text pruning engages only on columns that have index blobs. Other columns fall back to row evaluation.
- Puffin registrations survive expiry of the snapshot that made them. Snapshot expiry keeps the statistics entry in table metadata and the orphan sweep treats the sidecar as reachable, so a long-lived file keeps its sidecar index. Siglake's own expiry keeps the entry; another engine expiring snapshots on the same table may not.
- The two index storage paths protect their bytes differently, and only one of them answers a failed check with a scan. See What the footer text-index checksum covers.
What the footer text-index checksum covers¶
A footer inverted index's checksum sits in the same Parquet footer as the index
it checks, so footer-wide damage that changes both consistently is outside its
cover. From Siglake 0.2.0, a footer blob is written with a CRC-32 of its bytes,
eight hexadecimal characters under
siglake.inverted_index.crc32.v1[.<column>]. The reader checks it before the
parsed-cache handout and before any decode, so a warm index is covered as well
as a cold one. A malformed or disagreeing value is refused: the query tries the
column's Puffin index, otherwise scans the file exactly, and counts the refusal
in siglake_index_footer_checksum_refused_total{reason}. A checksum inside the
blob would close the shared-footer boundary, at the price of every 0.1.x reader
refusing every newly written index until the fleet finished upgrading.
Which path a column takes is a size decision: a v1 index goes in the Parquet
footer while that column's serialized index fits
SIGLAKE_INDEX_FOOTER_MAX_BYTES (1 MiB by default), and in a Puffin sidecar
above it. The sidecar is written as a Zstd frame with a content checksum, so a
corrupt stored byte fails the query with Restored data doesn't match checksum
instead of falling back to a scan, and it is not re-checked once warm.
Files written before 0.2.0 carry no sibling checksum, and a query prunes with their indexes as before: Parquet checksums data pages and not footer metadata, so a corruption that still decodes and still covers the file's rows prunes with it, and the query succeeds while answering short. Upgrading protects what you write next, not the files you already have; a file gains the checksum when a rewrite replaces it. An 0.1.x reader ignores the new footer entry and reads the unchanged blob, so a mixed-version fleet keeps its index coverage.
A single-bit sweep of each stored form on a 1,000-row, 1,010-term fixture measured the uncovered footer path in September 2026, before the checksum shipped. It is the evidence for what an unverified footer blob does, which is still what a file written before 0.2.0 carries. The last column follows one probe term of the 1,010; the column before it counts any term that lost rows.
| stored form | flips | flips that lost some term's rows | flips that lost the probe term's rows |
|---|---|---|---|
| footer KV, 28 KB of hex as stored | 224,368 | 10,432 | 125 |
| Puffin sidecar, Zstd frame as stored | 19,888 | 0 | 0 |
| the same frame, content checksum off | 19,856 | 5,883 | 20 |
Non-mergeable queries run single-pod¶
The distributed classifier falls back to DistPlan::Local for shapes it cannot
safely split and merge, including non-mergeable aggregates and queries over
more than one relation. See the query engine's
mergeability rules for the
full list.
Fan-out membership converges, it does not switch¶
Both packages render --query-peer-discovery-srv instead of a static
--query-peers list, so the old ceilings are gone: the
chart accepts keda.query.maxReplicas > query.replicas and the operator accepts
a real spec.autoscaling.query range. What remains is convergence lag. A new
pod is eligible for shard work one readiness probe plus one discovery refresh
(--query-peer-discovery-interval-secs, default 5 s, with CoreDNS TTL on top)
after it starts, and each query pins the membership it captured. Scale-out
therefore helps the next query, not the one in flight. A pod whose discovery
has never succeeded serves every query single-pod; one whose refreshes have
started failing serves a stale membership. Both are visible on
siglake_query_peer_discovery_members and
siglake_query_peer_discovery_refresh_total, and alert as
SiglakeQueryPeerDiscoveryStalled.
Cross-replica job cancellation is a poll¶
With the default shared job store and more than one query pod, the replica
serving DELETE /api/v1/jobs/{id} is usually not the one executing the job. The 202
is terminal for the job row immediately, but the executing replica learns of it
by re-reading its own in-flight rows every --jobs-cancel-poll-secs (default
2 seconds); there is no LISTEN/NOTIFY fast path. The 202 therefore
bounds when the admission share and the storage scans are released rather than
reporting that they already are. See Batch
jobs. That watch loop also runs on the
batch runtime, alongside the batch futures, so a job occupying a worker without
yielding delays its own cancellation.
Persistent batch-job recovery evidence is incomplete¶
- The two-store Postgres ownership regression needs a scratch database and is ignored by the normal test suite. One retained container run passed it; recovery through a real Postgres outage remains unmeasured.
- A job verdict that cannot be persisted is remembered only in that query process, for at most 1,024 job IDs. If the pod dies first, lease-expiry recovery later marks the job failed without its computed verdict. No live Postgres outage has measured this path.
Catalog-claim commits and table subscriptions have gaps¶
?commit=forcereturns400when a remote catalog-claim drain consumes the WAL mirror. The server cannot observe the segment leaving local directories; use the default?commit=wait_foracknowledgement and query for visibility.siglake subscribeadvances past every non-append commit. An external writer'sINSERT OVERWRITE,MERGE, or row-level delete can therefore add rows that a subscriber never receives. Query the affected interval to recover them.
No rollback is qualified against an older image¶
Rolling the image back keeps the columns a migration added, and the additive direction is regression tested. Every one of those tests runs the current binary against a table widened past what it declares, so it is evidence about the mechanism and not about a released image. No run has deployed image N, migrated, deployed N-1 and read the result back, and both packages' rollback behavior is reasoned from the chart templates and the reconciler. An image that differs by more than its declared column set is outside even that reasoning. See What a rollback does for Helm and Reverting the image for the operator.
Design decisions with consequences¶
Siglake's metrics do not leave the process as OTLP¶
Since 0.2.0 a Siglake process exports its own log records and spans over OTLP,
but never its metrics. Those are siglake_* Prometheus series on /metrics,
which is what the alert rules, the KEDA scalers and the Grafana dashboard are
written against. An OTLP-only backend needs a collector with a Prometheus
receiver scraping the metrics ports, which Siglake's own
telemetry
configures. Moving the metric call sites onto the OpenTelemetry API would
rename every series those rules depend on, so the collector carries the cost
instead.
Neither the chart nor the CRD has an OTel block¶
OTLP export is configured with environment variables on both deployment paths.
The Helm chart renders no otel: values, so the variables go in per-role or
global extraEnv; a SiglakeCluster has no OTel field either, so they go in
spec.extraEnv, which reaches every container the operator renders. Two
consequences follow from the operator's list being plain name/value and
cluster-wide: the tiers export together or not at all, and a variable that
would need valueFrom, such as POD_NAME from the downward API, cannot be
set there. A first-class block waits until a deployment has run with export on
and shown which settings operators reach for.
Query memory below 4 GiB¶
- Below
4 GiB, the query pool cannot reserve one compacted file's decode estimate. A scan still opens one file, but its decode buffers sit outside the pool's accounting. - The operator reports
QueryMemoryUndersized=TruewithQueryMemoryBelowDecodeFloorand still renders the StatefulSet. Local-disk measurements did not show a latency floor; overlapping object-store reads remain unmeasured.
Query spill is bounded node-local scratch space¶
Query spill is neither durable nor globally observable storage. PVC-backed spill for sorts larger than a pod's ephemeral-storage budget is not implemented, and Siglake exposes no whole-runtime spill-bytes metric. Query failures and pod ephemeral-storage usage are the operational signals.
Distributed admission is per-coordinator¶
A distributed query reserves one admission share on its coordinator, at most a
quarter of the pod budget, held from admission through the merge. Shard work
reserves nothing on workers, whose concurrent heavy shard scans are
bounded by the process memory pool and the per-shard rows-scanned breaker
instead; they can reach replicas × 4 at once. A cluster-scoped budget or a
worker reservation priced at the shard's own share is not implemented.
WAL-append acks, and a shared filesystem¶
An acknowledgement means the batch is in one node's local WAL, not in object
storage. The default wait_for mode fsync(2)s that WAL before it answers, on
OTLP/HTTP logs and traces, the Elasticsearch-compatible bulk endpoints, and
OTLP/gRPC logs and traces. The sync covers segment bytes and directory entries
for the tenant, index, active segment, and sealed segment. Siglake syncs the
sealed name before it unlinks the active copy.
The power-loss guarantee assumes ext4 or xfs on a node-attached volume. With
EFS or another network filesystem, the guarantee is whatever that filesystem's
fsync(2) and rename semantics provide. If you would rather trade durability
for latency, send commit=auto (or the equivalent X-Siglake-Commit: auto
header) to acknowledge once write(2) has reached the kernel page cache. That
write survives a process crash, an OOM kill, and a pod restart, but a node
crash or power failure can lose it before the kernel flushes.
Neither mode waits for the rows to become queryable: the compactor commits to Iceberg on its own cadence. Neither waits for the WAL mirror to upload the segment to object storage. An object-store group-commit "PUT-as-ack" model is designed but not built.
Every WAL consumer that may run on a different node needs a ReadWriteMany
filesystem (EFS on AWS). That includes a single ingester and a single compactor
when they run on separate nodes, plus query pods when the freshness buffer is
enabled. ReadWriteOnce is supported only when every WAL consumer is
explicitly co-located on one node. RWX brings its own cost, throughput
characteristics, and operational surface.
Mirror reconciliation scales with retained objects¶
In a catalog-claim deployment, one elected compactor repairs mirror objects
whose catalog registration was missed. It registers at most 1,024 sealed
objects per pass and stores a shared (last_key, rotation) cursor in the
mirror_sync_cursors catalog table. The cursor advances only after the whole
page registers, survives owner handoff, and clears at end-of-prefix so a later
rotation repairs keys inserted behind it. S3 listings use OpenDAL
start_after; other backends locally filter the listing as a correctness-first
fallback.
The default compactor.committedRetentionSecs: 86400 purges committed objects
and catalog rows after 24 hours; 0 is the explicit never-purge opt-out, and
non-zero values are floored at 901 seconds. Each retention run drains
512-object pages up to a 16,384-object bound and is paced start-to-start.
siglake_compactor_mirror_sync_objects records each bounded page and should
plateau at 1,024 while a rotation is in progress. The whole-rotation metrics
are siglake_compactor_mirror_sync_rotation_duration_seconds,
siglake_compactor_mirror_sync_rotation_objects,
siglake_compactor_mirror_sync_rotations_completed, and
siglake_compactor_mirror_sync_last_completed_timestamp_seconds. The first two
are recorded when the cursor wraps; the last two are durable gauges that
survive owner handoff. Completion age is the backlog or stall signal when every
page stays full; a rising retained-per-rotation count with proportional
duration means retention is backlogged. See Mirror reconciliation and retention
for retention sizing and dashboard guidance.
The Helm value prometheusRule.mirrorRotationStallSecs defaults to 21600
(six hours). It controls how long bounded pages may run without a full rotation
completing before SiglakeMirrorReconciliationStalled can fire; 0 disables
that alert.
Snapshot expiry limits time travel¶
Snapshot expiry is on by default and retains the last 100 snapshots, because
metadata.json is read and rewritten on every commit. Iceberg time travel
reaches no further back than the retained snapshots. Widening the window costs
commit throughput.
In catalog-claim deployments, new writers atomically maintain the bounded
siglake.consumed_proof.v1 table property, so committed-claim evidence survives
snapshot expiry and reclustering. Old writers know only the retained snapshot
summaries. Before a rolling upgrade, set
compactor.snapshotExpire.retainLast
>= 400; keep it there until every old drain/maintenance writer is gone and
for another 1,025 seconds. The dual reader uses both sources during that
interval. See consumed-proof rolling
upgrades for the
deployment procedure.
Corrupt or over-cap durable state fails closed: it refuses writes and makes
uncovered reclaim decisions unprovable rather than silently discarding positive
proof. Reclaim starts after SIGLAKE_CLAIM_RECLAIM_MAX_AGE_SECS (900 seconds by
default); 0 is refused and resolves to 900 with a once-per-process warning.
When neither the durable proof nor retained snapshot history covers an
abandoned segment, the compactor increments
siglake_compactor_reclaim_unprovable_total and requeues it, preferring a
detectable possible duplicate to silent loss.
Commit batching delays external visibility¶
Commit-accumulation batching defers commits by up to maxAgeSecs (10 s
default). With the optional WAL buffer enabled, Siglake queries can see sealed
WAL segments before commit. Acknowledgement does not make a row visible: the
segment must first seal, after 4,096 events or 5 seconds by default. External
Iceberg readers always wait for the Iceberg commit, and so does every Siglake
table when the buffer is off. See OTLP-to-SQL visibility
delay.
Multi-tenancy is soft¶
Tenants share compute, the catalog, the bucket, and object-storage credentials. No routing mode gives a tenant its own compute quota.
Ingest is single-tenant by default. Every request routes to the default
tenant, and an X-Scope-OrgID naming another tenant is refused: 403 on
HTTP, PermissionDenied on the OTLP/gRPC logs and traces exporters. Naming
default is a no-op, so a client that always sends the header keeps working.
Routing tenants is something you turn on, with one of two settings:
--oidc-tenant-claim/ingester.oidc.tenantClaimtakes the tenant from the caller's verified JSON Web Token, on both transports. A header may only agree with the claim, and a token carrying no usable claim is refused.--trust-scope-header/ingester.trustScopeHeadertakes the client's word: the header selects the tenant again. Use it where a gateway in front of the ingester sets the header itself and strips the client's; on anything else, any accepted credential can write as any tenant.
Neither setting makes tenants isolated from each other. For hostile-tenant isolation, run separate deployments. See Multi-tenancy.
Deletes are not immediately permanent, and cover one table¶
Delete tasks rewrite files, but the old files persist in expired snapshots until orphan GC reclaims them, and then as noncurrent S3 versions if versioning is on. Compliance-grade erasure of the table requires the full chain: execute, expire, GC, and expire object versions.
That chain stops at the table. Orphan GC is rooted at the table location, so
the WAL mirror prefix, its _active/ blobs, the WAL volume's poison/
quarantine, replica buckets and catalog backups still hold the same events.
Each needs its own expiry rule, and no Siglake command sweeps them for you. See
Copies the delete chain does not
reach.
Delete-task claims and recovery¶
- A delete-task claim is never released, including after completion. A process that dies after claiming a task can strand it, and deleting old claim objects is unsafe because a delayed executor might still hold the task.
- Each task records the table UUID resolved at submission. If the index name now points to a replacement table, the task fails instead of deleting from the new incarnation.
- A failed task, or a
runningor claimedpendingtask stranded by a crash, cannot return topending. Fix the cause and submit a new task. If the first commit outcome was ambiguous, check the replacement task's deleted-row count before deciding what the first run changed.
Dropping an index does not reclaim its storage¶
DELETE /api/v1/indexes/{id} records the table UUID, committed-file inventory
and UUID-scoped aggregate prefix before removing the catalog entry. The record
is report-only, so the dropped table's committed files and aggregates stay in
the bucket. Retention and orphan GC both resolve a table through its catalog
entry, so neither can reach those files after the drop. Deleting an index frees
no space. An index created later with the same id reuses the table location
with a fresh table UUID and reads only its own rows. The aggregate inventory
stays pinned to the dropped UUID and never resolves the reusable name. See
Index management.
A pre-0.1.0 warehouse has to be recreated, not migrated¶
Tables written before September 6, 2026 are Iceberg format version 3 with a
nanosecond timestamp and no timestamp_ns sibling. Siglake still reads them,
but no v2-only external engine can. Iceberg cannot change a column's precision
and cannot downgrade a v3 table, so there is no in-place upgrade:
migrate-schema refuses such a table instead of migrating it, and the remedy
is to delete the warehouse and re-ingest. A read-old/write-new rewrite tool is
not built. Tables that current builds create are format version 2 with a
microsecond timestamptz timestamp and a timestamp_ns column, so this
applies only to warehouses that predate the contract. The same boundary bounds
a rollback: see What a rollback
does and Reverting the
image.
Prometheus histogram and counter limits¶
- Only
*_secondsmetrics and four per-call counts have Prometheus histogram buckets:siglake_group_count_deltas_folded,siglake_group_count_tier2_files_per_call,siglake_compactor_mirror_sync_objects, andsiglake_compactor_mirror_sync_rotation_objects. Other histogram calls export summaries without_bucketseries. - Alert counters with fixed labels are registered at zero, so
increase()sees their first event. Dynamic label values cannot be registered in advance;siglake_storage_schema_drift_total{column=...}and everysiglake_group_count_delta_write_failures_totalseries outside{iceberg_namespace="siglake",table="events"}, which covers tenant namespaces and index tables, need a second event before anincrease()alert sees a delta.
Scale and performance caveats¶
- Ordinary log search scales through replication, not per-query fan-out. More
replicas add throughput for small-
LIMITbrowse and Tier-1 aggregate queries, but one replica answers each query. Size this tier for concurrency, not lower latency for one search. - Interactive scans slow during a backfill in proportion to compaction lag. The packaged 1 GiB / 2 CPU compactor merges one bin at a time and lags hardest when ingest drives up file-overlap depth. In the 2026-08-27 full-rate 1 TB validation, adequately sizing the compactor reduced the same windowed browse from 40.9 s to 0.13 s without reducing ingest throughput. If interactive search must remain responsive during backfill, follow the compactor sizing recipe.
- Query throughput plateaus around 16-way concurrency, CPU-bound on browse decode.
- Layout convergence takes hours. A 2 B-row table needs roughly 4 to 5 hours of background compaction to converge. Query performance on an unconverged layout is substantially worse: 28 QPS versus 115 QPS at 32-way in one measured case.
- Ingester vertical scaling stops paying past ~2 CPU, because of per-tenant writer-task serialization.
- The catalog is on the commit path. At high commit rates Postgres becomes the bottleneck.
- High tenant cardinality is untested. Per-tenant WAL subdirectories and writer tasks mean tenant count drives resource use.
A rare-term match_terms query with no LIMIT is the case for a larger parsed-index cache¶
One text query shape runs faster with a parsed-index budget above its derived
1 GiB cap: a repeated match_terms query on a
rare term, with no LIMIT, over a fully-compacted table. Without a LIMIT the
server keeps the index instead of declining it
(siglake_query_inverted_index_declined_total does not move), and a 1 GiB
budget then evicts it between lookups. A local paired-budget experiment
raised SIGLAKE_PARSED_INDEX_CACHE_MAX_BYTES and
SIGLAKE_PUFFIN_BLOB_CACHE_MAX_BYTES together, from 1 GiB / 256 MiB to
8 GiB / 2 GiB:
| Measurement | At 1 GiB / 256 MiB |
At 8 GiB / 2 GiB |
|---|---|---|
rare_scan p50 |
22,798.9 ms | 120.2 ms |
rare_scan_last25 p50 |
4,843.4 ms | 38.2 ms |
| Total process RSS | 14.3-15.6 GiB | 20.2-21.3 GiB |
Neither knob was varied alone, so the pair is what those latencies cover and
they carry no guaranteed production speedup. The RSS figures are the whole
process on a local fixture, not the caches' incremental cost and not a pod
limit to copy. The run qualified no HTTP, object storage, distributed execution
or AWS, and the five clipped query shapes beside these two stayed within noise.
Treat the override as a sizing option for a deployment that has this query
shape and the memory to spare. Defaults, the 1 GiB derivation cap and the
packaged 4 GiB query limit are unchanged. For the knobs themselves, see Tune
the Puffin and parsed text-index
caches.
Benchmarking caveat¶
Result caches will silently invalidate a benchmark. A 1 TB round collapsed every query shape onto a ~13 ms result-cache-hit floor, with five orders of magnitude of difference in rows scanned producing latencies within 2 ms of each other.
Set SIGLAKE_QUERY_RESULT_CACHE=off, or use novel literals, and report cold
and warm separately. See Performance.
Reporting¶
Features do land, so parts of this page go out of date. If you find something here that no longer matches the code, open an issue at https://github.com/siglake/siglake/issues.
Public release validation¶
The renamed v0.1.0, v0.2.0 and v0.2.1 public images passed basic smoke checks on October 8, 2026: image identity, binary startup, Compose health, exact committed rows and grouped counts, query-server restart, and cleanup. The v0.1.0 and v0.2.0 checks used pinned public MinIO fixture substitutions for their unavailable historical dependencies.
These runs verify the renamed release artifacts, not a full burn-in, Kubernetes installation, or upgrade qualification. The release-validation ledger records the evidence and remaining coverage separately from historical install failures.