Skip to content

Scope and limitations

Siglake provides the telemetry storage and query layer: ingestion, durable WAL storage, open Iceberg tables, distributed SQL, and interfaces for external consumers. This page separates intentional product scope from configuration choices and implementation constraints, so you can plan a deployment around the contracts it supports.

The implementation constraints on this page summarize docs/LIMITATIONS.md in the Siglake repository at release commit 65c4eef.

Product scope

User interface, alerting interface, dashboards

Siglake integrates with external presentation and alerting tools through its SQL and HTTP APIs, open tables, and WAL consumer interface. Use Grafana over /api/v1/sql and the Jaeger shim for dashboards and trace views, as described in Grafana and Jaeger.

Dashboard authoring, alert evaluation and delivery, acknowledgement, and on-call workflows belong in those external tools. They are outside the core storage and query product's scope, rather than unfinished core components.

Relevance scoring (BM25)

Log exploration uses time ordering and SQL filters. Siglake does not apply BM25 relevance scores or scored TopK ranking. Term frequency is often less useful for one-line templated logs, where the same message repeats many times. The source design docs/DESIGN_bm25_scoring.md records a possible alternative; it is not the current search contract.

Deployment planning

Account for these operating contracts when choosing a topology:

Features off by default in Helm

These options let you choose freshness, scaling, resource use, and cluster integrations explicitly. Enable the ones your deployment needs.

Feature Setting Default behavior and configuration choice
Query WAL buffer query.walBuffer.enabled: false Queries read committed data. Enable the buffer for visibility into sealed WAL segments before commit.
Catalog claim compactor.catalogClaim.enabled: false Single-pod compaction only. The chart refuses to render a compactor tier that can hold more than one pod.
Attribute auto-promotion SIGLAKE_AUTO_PROMOTE_MIN_PCT=0 Automatic promotion widens the schema from the compactor's sample, and that widening cannot be undone. Hot keys stay in the JSON blob unless declared explicitly.
KEDA autoscaling keda.enabled: false Replica counts are configured directly. Enable KEDA for metric-driven autoscaling.
Prometheus alert rules prometheusRule.enabled: false Enable the bundled rules with the Prometheus Operator and an external alert-delivery workflow.
Source-file batch cache query.scan.fileCacheMaxBytes: 0 and query.scan.fileCacheMaxEntries: 0 The byte limit is reserved from the query memory pool, so a non-zero default would reduce headroom on pods that never reread a file. Warm repeat reads do not reuse decoded source batches. Set both limits to positive values to enable this experimental cache.
ServiceMonitor, PDBs, NetworkPolicies, anti-affinity all false No Prometheus Operator integration, no disruption budgets, no network isolation.

A query pod at the 4 GiB memory floor also assigns no cache to parsed or serialized text indexes by default. Each text query deserializes the index again unless you set SIGLAKE_PARSED_INDEX_CACHE_MAX_BYTES and SIGLAKE_PUFFIN_BLOB_CACHE_MAX_BYTES yourself. A 5 GiB limit gives both derived cache budgets some capacity. A pod in that state shows it: every index acquisition is a miss on siglake_iceberg_parsed_index_cache_lookups_total{outcome} with no eviction beside it. See Parsed text-index cache outcomes.

WAL mirroring is enabled by default when you configure a warehouse URL. Each sealed segment uploads asynchronously to <warehouse-url>/wal-mirror/. An acknowledgement is durable on the local WAL after the default fsync(2); the rows get a copy in object storage when their Iceberg commit or that upload completes, whichever comes first. Mirroring shortens the pre-commit window, so losing the WAL volume costs you the acknowledged rows that have been neither committed nor uploaded. The default single-replica compactor does not reclaim mirror objects, so give the prefix an object-store lifecycle expiry unless you enable compactor.catalogClaim.enabled. Set wal.mirror.enabled: false to opt out.

Leveled compaction is enabled by default and therefore is not listed above. Set SIGLAKE_COMPACTOR_LEVELED=0 only to select the legacy flat whole-partition pass.

Footer inverted indexes are enabled by default. Set SIGLAKE_INVERTED_INDEX=0, Helm compactor.invertedIndex.enabled: false, or operator spec.extraEnv: [{name: SIGLAKE_INVERTED_INDEX, value: "0"}] to opt out.

The post-rewrite Puffin rebuild is disabled by default. A rewrite that could not put an inverted index in the Parquet footer, which is every streamed merge and any column whose index exceeds SIGLAKE_INDEX_FOOTER_MAX_BYTES, leaves that output unindexed, and queries scan it for exact rows. Set SIGLAKE_INDEX_REBUILD=1, Helm compactor.indexRebuild: true, or operator spec.extraEnv: [{name: SIGLAKE_INDEX_REBUILD, value: "1"}] to opt in: the pass then registers a Puffin sidecar after the rewrite commits, skipping files whose columns already carry a footer index or a registration. With inverted indexes off it does nothing either way. Reads do not depend on the switch. Indexes that already exist are still used, and the flush path still writes its footer inverted index.

The rebuild covers inverted indexes only. A streamed merge also skips the whole-file raw trigram bloom, and no later pass registers one, so LIKE '%substr%' on raw prunes against per-row-group blooms alone until an in-memory rewrite covers the file. A delete-task rewrite writes at rewrite generation 0, like an ingest flush, so on a table with index_at_flush: false even its in-memory arm skips the inline index work.

Turning either switch back on backfills nothing already committed: a rebuild only ever sees the files of the rewrite it follows, and only a rewrite of the file itself writes an inline footer index or the whole-file bloom. Files committed while a switch was off keep the metadata they were written with. The cost throughout is pruning rather than correctness: an unindexed file is scanned and row-evaluated, and the answer is exact. See What each merge path writes.

Segmented text indexes (seg2) are off at both ends in Siglake 0.2.0. SIGLAKE_SEGMENTED_INDEX_WRITES=1 makes the compactor's streaming re-cluster write one, and SIGLAKE_SEGMENTED_INDEX_READS=1 lets a query read one; with reads off, a file whose only text index is seg2 is scanned. Neither switch has a chart value or an operator field, so set them through compactor.extraEnv and query.extraEnv. The published measurements are local file:// runs, not an object-store, distributed or HTTP result, which is why both defaults stay off. See Segmented text indexes.

Delete-task execution is on by default too. The compactor runs the sweep in its idle cycle unless SIGLAKE_DELETE_TASKS is 0, off, false or no. The chart ships compactor.deleteTasks: true and renders the variable either way, so false renders SIGLAKE_DELETE_TASKS=0. The operator's compactor renders SIGLAKE_DELETE_TASKS=1 and the CR has no field for it: put SIGLAKE_DELETE_TASKS=0 in spec.extraEnv to opt out, which is also what --adopt-values reports when the chart values file it reads set compactor.deleteTasks: false. With the sweep off, POST /api/v1/delete-tasks still records the task, and siglake delete-sweep --apply still executes it by hand.

OTLP/gRPC is on by default as well. The ingester listens for logs and traces on 0.0.0.0:4317 next to OTLP/HTTP, and both transports apply the same authentication, tenant routing and WAL-fsync acknowledgement contract. The chart ships ingester.otlpGrpc.enabled: true with ingester.otlpGrpc.port: 4317; setting enabled: false drops the gRPC container and Service ports and passes --disable-otlp-grpc to the binary. Outside Helm, leaving --otlp-grpc-listen unset no longer disables the listener: pass --disable-otlp-grpc instead. The operator renders the listener at every replica count, and its CR has no field to turn it off.

Persistent batch jobs are on by default (query.jobs.persistent: true). The query pods keep batch-job state - submissions, their status and their result rows - in the Postgres instance that already holds the Iceberg catalog. The chart renders SIGLAKE_JOBS_POSTGRES_URI from the same Secret, so there is no second connection to configure, and the query server creates its own tables on start. The operator points an adopted query tier at its own spec.catalogUri. One store for the tier is what makes a job_id usable at the shipped query.replicas: 2: the Service has no session affinity, so a status, result or cancel request lands on either pod.

Each job row records the replica incarnation executing it, and owners heartbeat. If a query pod crashes, recovery fails its pending and running rows once the owner's lease expires (SIGLAKE_JOBS_OWNER_LEASE_SECS, default 120 s); jobs owned by a sibling that is still heartbeating are left alone. A planned SIGINT or SIGTERM shutdown deletes that pod's registration after the HTTP server drains, so the next recovery pass can fail its remaining non-terminal rows without waiting out the lease. Neither path resumes interrupted work: the client sees failed and has to resubmit.

Setting query.jobs.persistent: false drops the variable and gives each pod its own in-memory store. It is a supported opt-out for a tier you can hold at one pod, and both deployment paths refuse the rest. The chart fails the install when query.replicas, or keda.query.maxReplicas with keda.query.enabled: true, is above 1 while the store is off. The operator refuses the spec with QueryJobsStoreRequiredForScaleOut when spec.autoscaling.query.max is above 1 and the effective jobs-store URI is blank. Without those guards, the job status, result and cancel reads that land on the pod which did not take the submission answer 404 while the job runs normally elsewhere. The one-pod cost remains: a pod restart loses every job in flight on it.

The catalog claim stays off by default, and it is a requirement rather than a recommendation once you scale the compactor. The chart fails the install on the three combinations that cannot work.

  • More than one compactor pod with the claim off. That is compactor.replicas above 1, or autoscaling.compactor.maxReplicas above 1 when autoscaling.compactor.enabled is true. The shipped HPA ceiling is 4, so turning the compactor HPA on is refused until the claim is on. Uncoordinated replicas share no claim table: each runs the whole drain and maintenance loop against the same tables, and they lose Iceberg commit races to each other while the layout stops converging.
  • The claim with wal.mirror.enabled: false. In claim mode the drain reads the wal_segments catalog table and never the local sealed/ directory, and rows land there only from an ingester that mirrors its segments. That compactor would claim nothing while siglake_compactor_sealed_pending read zero, because in claim mode that gauge counts the shared queue of sealed, unclaimed catalog rows. It cannot reveal a mirrored object that has no catalog row.
  • The enabled compactor HPA custom metric with the claim on. The refusal needs autoscaling.compactor.enabled, autoscaling.compactor.customMetric.enabled and compactor.catalogClaim.enabled all set to true. Its error starts with autoscaling.compactor.customMetric.enabled is true with compactor.catalogClaim.enabled. Every compactor reports the whole shared sealed queue rather than its own share, but the chart renders a per-pod average target. Disable the custom metric for a claim-mode CPU-only HPA, or use siglake-operator for backlog-driven scaling. The custom metric scales nothing in filesystem mode because that mode is capped at one pod.

All three refusals name the value to change. The operator needs none of these chart combinations. It selects the drain from spec.autoscaling.compactor.max, not the current replica count. A maximum above 1 keeps the WAL mirror and catalog claim on at every current replica count, including while the autoscaler holds the compactor at one. A maximum of 1 keeps the filesystem drain.

Deployment constraints

  • The operator cannot scale the ingester or query tier to zero. It accepts a zero compactor floor only with the catalog-claim drain and positive smoothing. The packaged compactor floor remains 1 because no cluster round has yet proved that an ingester-published queue depth wakes a parked compactor.
  • Helm cannot run the embedded compactor. The chart refuses ingester.extraArgs: [--with-compactor] even at one replica because a rolling update can overlap two ingesters running maintenance. Use the dedicated compactor tier.

Automatic operator drain-mode handover

The operator does not convert WAL between filesystem and catalog-claim drains. Changing spec.autoscaling.compactor.max across 1 changes the ownership protocol, so reconciliation reports DrainModeHandoverRequired and stops before applying workloads. You must finish the current drain, account for local WAL, mirror objects and catalog rows, then follow the manual handover procedure. The operator does not convert WAL or delete those segments and rows for you.

Aggregating merge kinds

No rollup or last-write-wins dedup merges exist. The exact rows-conserved commit guard is per-merge-kind-waivable by design, but nothing waives it yet.

Metrics downsampling

None. Retention is the only size-control mechanism.

Row-level retention

Retention drops whole data files whose manifest maximum timestamp is past the horizon. It is file- and day-granular. Your effective retention is the policy plus the span of the last surviving file.

Implementation constraints

Rows sharing a custom event time come back in an unspecified order

  • A bare browse over an index orders newest-first on the field its mapping names as timestamp_field.
  • Equal custom event times have no implicit tiebreak. Which rows land at a LIMIT boundary can change between executions.
  • A deterministic tie order needs a second sort column declared on the index and named in the query.

Ingester autoscaling lacks live Prometheus evidence

  • The operator's ingest-rate query depends on the pod label that Prometheus Operator adds during target relabeling. The chart does not add that label.
  • No retained run has confirmed the label and per-pod average with two active ingesters. Without the label, the query reads the fleet total as one pod's rate and can request too many replicas.

An audit batch abandoned at its deadline is lost

  • Query responses do not wait for the best-effort audit worker. Its retained rows and conversion memory are bounded.
  • Siglake 0.2.0 and later bound one append at 30 s (SIGLAKE_QUERY_AUDIT_APPEND_DEADLINE_SECS; 0 restores 0.1.0's unbounded await, where a stuck append held the whole retained budget until the process restarted). A batch that outlives the deadline is abandoned and never re-appended, so a catalog commit already on its way can land unseen. Those rows are counted by siglake_query_audit_dropped_total{reason="append_deadline"}.

Concurrent index updates are not merged

  • Two concurrent mapping updates that add different fields do not merge. The losing request gets 400 and must read the current mapping before it sends the update again.
  • Two writers updating the same index template are last-write-wins. A template carries no version and takes no If-Match condition, which 0.2.0 added for an index mapping alone. Deleting a template leaves a tombstone that is not garbage-collected.

External-reader evidence is local and narrow

  • The Trino 483, Spark 3.5.9, DuckDB 1.5.5 and PyIceberg 0.12.0 demonstration used a file:// warehouse and a SQLite catalog. It does not cover object storage, Postgres, or reads concurrent with compaction.
  • A later run checked decoded microsecond timestamps with Spark, DuckDB and PyIceberg once. Its retained record omits two reader versions and does not include Trino.

Query probes do not detect query-path degradation

The query server's /healthz is a constant 200 for as long as the process can serve the probe. /readyz verifies only that the Iceberg catalog can be reached; it does not execute a query or check the cache-warm loop. A pod can therefore pass both probes while browse queries time out, so probe success is not evidence that the query path is healthy.

The shipped signal for the known degradation is SiglakeQueryWarmCycleStalled, not an automatic readiness failure or liveness restart. See Monitoring for the operational consequence. These semantics were confirmed by Siglake's September 3, 2026 review of query degradation and probe behavior.

Tier-1 group counts require first-commit columns

The Tier-1 whole-table group-count aggregate initially covers only columns present since the table's first commit. A table created before typed measurements joined side aggregates can therefore keep serving GROUP BY on those columns through the exact but slower materialized path, even when its files already contain usable group-count footers.

Two older aggregate cases still need manual repair:

  • Lost-delta repair does not admit an eligible typed column that predates the failed commit.
  • Pre-coverage side aggregates are not adopted because they cannot prove which snapshot they describe. rebuild-group-counts repairs the wide group-count object. No 0.1.0 command rebuilds the inline time aggregates.

The release after 0.1.0 adds siglake rebuild-time-aggregates, which narrows the second case without closing it. The command recomputes a pre-coverage object's time aggregates from the committed files and publishes a coverage edge naming the snapshot it read. It does not certify the maps that were already there. You run it by hand; no commit path repairs the object for you. It repairs an existing inline object rather than creating one, because the column set it would rebuild is recorded nowhere else. Its publication drops the unproven whole-table group counts and restores only the components that account for every row in the snapshot summary, so hourly buckets short of that count, and any column short of it, are left out instead of published partial. See Inline aggregate repair for the counters it increments.

Use the group-count repair procedure to admit eligible typed columns with --admit-typed-columns. This does not require rewriting data files unless an older live file lacks both a usable footer and a raw-page representation for the column.

Elasticsearch bulk API support is write-only

Siglake ingests over the Elasticsearch bulk API but has no Elasticsearch query API, and none is planned. _bulk and _cluster/health work. _search, _msearch, _search/scroll, _field_caps and _cat/* are registered and answer 501 with a pointer to POST /api/v1/sql.

Existing Kibana or Elasticsearch query clients will not work. Query with SQL over HTTP instead: see the SQL cookbook.

Query replicas can briefly disagree across a commit

Each replica serves table metadata from its own cache. The default TTL is 5 seconds, and stale metadata can be served for up to 60 seconds while a background refresh runs (SIGLAKE_ICEBERG_METADATA_CACHE_TTL_SECS controls the TTL). Two replicas behind one Service can therefore see different snapshots during that window. Clients that need a consistent sequence of queries must send them to one pod or wait for every replica's cache to converge.

This is a cross-request limit only. Within one distributed query the fan-out is pinned to the coordinator's generation, its snapshot id and schema id, and a worker that cannot resolve that generation refuses its shard (503, reason: "shard_pin_unresolved") rather than contributing rows from another one, so a single answer is never merged across generations. See One generation per fan-out. The cost of that guarantee is availability: while a coordinator's cache is stale on a snapshot the catalog has already expired, its fanned-out queries return 503 until the cache converges (/api/v1/sql/local still answers). That state alerts as SiglakeQueryShardPinUnresolved. Raise compactor.snapshotExpire.retainLast to widen the margin, or lower SIGLAKE_ICEBERG_METADATA_CACHE_TTL_SECS so coordinators stop serving a snapshot before the catalog expires it.

Jaeger is HTTP-only

The shim covers the HTTP subset Grafana's Jaeger data source renders. There is no gRPC SpanReader, so tools expecting that interface will not work.

Kubernetes operator gaps compared with Helm

It renders the core data plane with the same workload kinds as the chart: ingester and compactor as Deployments, query as a StatefulSet. The Helm chart remains the supported, more complete install path.

The operator also renders the WAL PVC, Services, retention CronJobs, and a one-shot schema-migration Job when spec.schemaVersion is set. It does not render Ingress, PDBs, NetworkPolicies, ServiceMonitors, HPA/KEDA objects, or scheduling constraints. Per-tier resources default to the chart's and are overridable through spec.resources. The CR cannot express query bearer tokens, OIDC, TLS, or a WAL-buffer volume.

Catalog credentials are plaintext in the spec.catalogUri custom resource. The operator copies them into pod environment variables and CronJob arguments. The chart instead keeps Secret references and an unexpanded URI in pod specifications.

Adoption of an existing Helm release through --adopt-values is experimental and offline. It has no live test coverage and leaves several values unmapped. The ownership handover (annotation flip plus release-secret removal) is a manual runbook.

Freshness does not cover system tables

The WAL buffer serves the events table and the managed user indexes a query references. It does not serve query_audit or any other system table, which see commit-cycle visibility. When the buffer is off, or a query pod cannot read the WAL mount, every table falls back to commit-cycle visibility.

Strict mapping mode does not reject

"mode": "strict" document mapping enforces at commit time as dynamic-plus-a-counter: it counts violations rather than rejecting documents, and undeclared attributes still land in the residual attributes column. Do not rely on it as a data-quality gate.

Search v1 bounds

  • Full-text pruning engages only on columns that have index blobs. Other columns fall back to row evaluation.
  • Puffin registrations survive expiry of the snapshot that made them. Snapshot expiry keeps the statistics entry in table metadata and the orphan sweep treats the sidecar as reachable, so a long-lived file keeps its sidecar index. Siglake's own expiry keeps the entry; another engine expiring snapshots on the same table may not.
  • The two index storage paths protect their bytes differently, and only one of them answers a failed check with a scan. See What the footer text-index checksum covers.

A footer inverted index's checksum sits in the same Parquet footer as the index it checks, so footer-wide damage that changes both consistently is outside its cover. From Siglake 0.2.0, a footer blob is written with a CRC-32 of its bytes, eight hexadecimal characters under siglake.inverted_index.crc32.v1[.<column>]. The reader checks it before the parsed-cache handout and before any decode, so a warm index is covered as well as a cold one. A malformed or disagreeing value is refused: the query tries the column's Puffin index, otherwise scans the file exactly, and counts the refusal in siglake_index_footer_checksum_refused_total{reason}. A checksum inside the blob would close the shared-footer boundary, at the price of every 0.1.x reader refusing every newly written index until the fleet finished upgrading.

Which path a column takes is a size decision: a v1 index goes in the Parquet footer while that column's serialized index fits SIGLAKE_INDEX_FOOTER_MAX_BYTES (1 MiB by default), and in a Puffin sidecar above it. The sidecar is written as a Zstd frame with a content checksum, so a corrupt stored byte fails the query with Restored data doesn't match checksum instead of falling back to a scan, and it is not re-checked once warm.

Files written before 0.2.0 carry no sibling checksum, and a query prunes with their indexes as before: Parquet checksums data pages and not footer metadata, so a corruption that still decodes and still covers the file's rows prunes with it, and the query succeeds while answering short. Upgrading protects what you write next, not the files you already have; a file gains the checksum when a rewrite replaces it. An 0.1.x reader ignores the new footer entry and reads the unchanged blob, so a mixed-version fleet keeps its index coverage.

A single-bit sweep of each stored form on a 1,000-row, 1,010-term fixture measured the uncovered footer path in September 2026, before the checksum shipped. It is the evidence for what an unverified footer blob does, which is still what a file written before 0.2.0 carries. The last column follows one probe term of the 1,010; the column before it counts any term that lost rows.

stored form flips flips that lost some term's rows flips that lost the probe term's rows
footer KV, 28 KB of hex as stored 224,368 10,432 125
Puffin sidecar, Zstd frame as stored 19,888 0 0
the same frame, content checksum off 19,856 5,883 20

Non-mergeable queries run single-pod

The distributed classifier falls back to DistPlan::Local for shapes it cannot safely split and merge, including non-mergeable aggregates and queries over more than one relation. See the query engine's mergeability rules for the full list.

Fan-out membership converges, it does not switch

Both packages render --query-peer-discovery-srv instead of a static --query-peers list, so the old ceilings are gone: the chart accepts keda.query.maxReplicas > query.replicas and the operator accepts a real spec.autoscaling.query range. What remains is convergence lag. A new pod is eligible for shard work one readiness probe plus one discovery refresh (--query-peer-discovery-interval-secs, default 5 s, with CoreDNS TTL on top) after it starts, and each query pins the membership it captured. Scale-out therefore helps the next query, not the one in flight. A pod whose discovery has never succeeded serves every query single-pod; one whose refreshes have started failing serves a stale membership. Both are visible on siglake_query_peer_discovery_members and siglake_query_peer_discovery_refresh_total, and alert as SiglakeQueryPeerDiscoveryStalled.

Cross-replica job cancellation is a poll

With the default shared job store and more than one query pod, the replica serving DELETE /api/v1/jobs/{id} is usually not the one executing the job. The 202 is terminal for the job row immediately, but the executing replica learns of it by re-reading its own in-flight rows every --jobs-cancel-poll-secs (default 2 seconds); there is no LISTEN/NOTIFY fast path. The 202 therefore bounds when the admission share and the storage scans are released rather than reporting that they already are. See Batch jobs. That watch loop also runs on the batch runtime, alongside the batch futures, so a job occupying a worker without yielding delays its own cancellation.

Persistent batch-job recovery evidence is incomplete

  • The two-store Postgres ownership regression needs a scratch database and is ignored by the normal test suite. One retained container run passed it; recovery through a real Postgres outage remains unmeasured.
  • A job verdict that cannot be persisted is remembered only in that query process, for at most 1,024 job IDs. If the pod dies first, lease-expiry recovery later marks the job failed without its computed verdict. No live Postgres outage has measured this path.

Catalog-claim commits and table subscriptions have gaps

  • ?commit=force returns 400 when a remote catalog-claim drain consumes the WAL mirror. The server cannot observe the segment leaving local directories; use the default ?commit=wait_for acknowledgement and query for visibility.
  • siglake subscribe advances past every non-append commit. An external writer's INSERT OVERWRITE, MERGE, or row-level delete can therefore add rows that a subscriber never receives. Query the affected interval to recover them.

No rollback is qualified against an older image

Rolling the image back keeps the columns a migration added, and the additive direction is regression tested. Every one of those tests runs the current binary against a table widened past what it declares, so it is evidence about the mechanism and not about a released image. No run has deployed image N, migrated, deployed N-1 and read the result back, and both packages' rollback behavior is reasoned from the chart templates and the reconciler. An image that differs by more than its declared column set is outside even that reasoning. See What a rollback does for Helm and Reverting the image for the operator.

Design decisions with consequences

Siglake's metrics do not leave the process as OTLP

Since 0.2.0 a Siglake process exports its own log records and spans over OTLP, but never its metrics. Those are siglake_* Prometheus series on /metrics, which is what the alert rules, the KEDA scalers and the Grafana dashboard are written against. An OTLP-only backend needs a collector with a Prometheus receiver scraping the metrics ports, which Siglake's own telemetry configures. Moving the metric call sites onto the OpenTelemetry API would rename every series those rules depend on, so the collector carries the cost instead.

Neither the chart nor the CRD has an OTel block

OTLP export is configured with environment variables on both deployment paths. The Helm chart renders no otel: values, so the variables go in per-role or global extraEnv; a SiglakeCluster has no OTel field either, so they go in spec.extraEnv, which reaches every container the operator renders. Two consequences follow from the operator's list being plain name/value and cluster-wide: the tiers export together or not at all, and a variable that would need valueFrom, such as POD_NAME from the downward API, cannot be set there. A first-class block waits until a deployment has run with export on and shown which settings operators reach for.

Query memory below 4 GiB

  • Below 4 GiB, the query pool cannot reserve one compacted file's decode estimate. A scan still opens one file, but its decode buffers sit outside the pool's accounting.
  • The operator reports QueryMemoryUndersized=True with QueryMemoryBelowDecodeFloor and still renders the StatefulSet. Local-disk measurements did not show a latency floor; overlapping object-store reads remain unmeasured.

Query spill is bounded node-local scratch space

Query spill is neither durable nor globally observable storage. PVC-backed spill for sorts larger than a pod's ephemeral-storage budget is not implemented, and Siglake exposes no whole-runtime spill-bytes metric. Query failures and pod ephemeral-storage usage are the operational signals.

Distributed admission is per-coordinator

A distributed query reserves one admission share on its coordinator, at most a quarter of the pod budget, held from admission through the merge. Shard work reserves nothing on workers, whose concurrent heavy shard scans are bounded by the process memory pool and the per-shard rows-scanned breaker instead; they can reach replicas × 4 at once. A cluster-scoped budget or a worker reservation priced at the shard's own share is not implemented.

WAL-append acks, and a shared filesystem

An acknowledgement means the batch is in one node's local WAL, not in object storage. The default wait_for mode fsync(2)s that WAL before it answers, on OTLP/HTTP logs and traces, the Elasticsearch-compatible bulk endpoints, and OTLP/gRPC logs and traces. The sync covers segment bytes and directory entries for the tenant, index, active segment, and sealed segment. Siglake syncs the sealed name before it unlinks the active copy.

The power-loss guarantee assumes ext4 or xfs on a node-attached volume. With EFS or another network filesystem, the guarantee is whatever that filesystem's fsync(2) and rename semantics provide. If you would rather trade durability for latency, send commit=auto (or the equivalent X-Siglake-Commit: auto header) to acknowledge once write(2) has reached the kernel page cache. That write survives a process crash, an OOM kill, and a pod restart, but a node crash or power failure can lose it before the kernel flushes.

Neither mode waits for the rows to become queryable: the compactor commits to Iceberg on its own cadence. Neither waits for the WAL mirror to upload the segment to object storage. An object-store group-commit "PUT-as-ack" model is designed but not built.

Every WAL consumer that may run on a different node needs a ReadWriteMany filesystem (EFS on AWS). That includes a single ingester and a single compactor when they run on separate nodes, plus query pods when the freshness buffer is enabled. ReadWriteOnce is supported only when every WAL consumer is explicitly co-located on one node. RWX brings its own cost, throughput characteristics, and operational surface.

Mirror reconciliation scales with retained objects

In a catalog-claim deployment, one elected compactor repairs mirror objects whose catalog registration was missed. It registers at most 1,024 sealed objects per pass and stores a shared (last_key, rotation) cursor in the mirror_sync_cursors catalog table. The cursor advances only after the whole page registers, survives owner handoff, and clears at end-of-prefix so a later rotation repairs keys inserted behind it. S3 listings use OpenDAL start_after; other backends locally filter the listing as a correctness-first fallback.

The default compactor.committedRetentionSecs: 86400 purges committed objects and catalog rows after 24 hours; 0 is the explicit never-purge opt-out, and non-zero values are floored at 901 seconds. Each retention run drains 512-object pages up to a 16,384-object bound and is paced start-to-start.

siglake_compactor_mirror_sync_objects records each bounded page and should plateau at 1,024 while a rotation is in progress. The whole-rotation metrics are siglake_compactor_mirror_sync_rotation_duration_seconds, siglake_compactor_mirror_sync_rotation_objects, siglake_compactor_mirror_sync_rotations_completed, and siglake_compactor_mirror_sync_last_completed_timestamp_seconds. The first two are recorded when the cursor wraps; the last two are durable gauges that survive owner handoff. Completion age is the backlog or stall signal when every page stays full; a rising retained-per-rotation count with proportional duration means retention is backlogged. See Mirror reconciliation and retention for retention sizing and dashboard guidance.

The Helm value prometheusRule.mirrorRotationStallSecs defaults to 21600 (six hours). It controls how long bounded pages may run without a full rotation completing before SiglakeMirrorReconciliationStalled can fire; 0 disables that alert.

Snapshot expiry limits time travel

Snapshot expiry is on by default and retains the last 100 snapshots, because metadata.json is read and rewritten on every commit. Iceberg time travel reaches no further back than the retained snapshots. Widening the window costs commit throughput.

In catalog-claim deployments, new writers atomically maintain the bounded siglake.consumed_proof.v1 table property, so committed-claim evidence survives snapshot expiry and reclustering. Old writers know only the retained snapshot summaries. Before a rolling upgrade, set compactor.snapshotExpire.retainLast >= 400; keep it there until every old drain/maintenance writer is gone and for another 1,025 seconds. The dual reader uses both sources during that interval. See consumed-proof rolling upgrades for the deployment procedure.

Corrupt or over-cap durable state fails closed: it refuses writes and makes uncovered reclaim decisions unprovable rather than silently discarding positive proof. Reclaim starts after SIGLAKE_CLAIM_RECLAIM_MAX_AGE_SECS (900 seconds by default); 0 is refused and resolves to 900 with a once-per-process warning. When neither the durable proof nor retained snapshot history covers an abandoned segment, the compactor increments siglake_compactor_reclaim_unprovable_total and requeues it, preferring a detectable possible duplicate to silent loss.

Commit batching delays external visibility

Commit-accumulation batching defers commits by up to maxAgeSecs (10 s default). With the optional WAL buffer enabled, Siglake queries can see sealed WAL segments before commit. Acknowledgement does not make a row visible: the segment must first seal, after 4,096 events or 5 seconds by default. External Iceberg readers always wait for the Iceberg commit, and so does every Siglake table when the buffer is off. See OTLP-to-SQL visibility delay.

Multi-tenancy is soft

Tenants share compute, the catalog, the bucket, and object-storage credentials. No routing mode gives a tenant its own compute quota.

Ingest is single-tenant by default. Every request routes to the default tenant, and an X-Scope-OrgID naming another tenant is refused: 403 on HTTP, PermissionDenied on the OTLP/gRPC logs and traces exporters. Naming default is a no-op, so a client that always sends the header keeps working. Routing tenants is something you turn on, with one of two settings:

  • --oidc-tenant-claim / ingester.oidc.tenantClaim takes the tenant from the caller's verified JSON Web Token, on both transports. A header may only agree with the claim, and a token carrying no usable claim is refused.
  • --trust-scope-header / ingester.trustScopeHeader takes the client's word: the header selects the tenant again. Use it where a gateway in front of the ingester sets the header itself and strips the client's; on anything else, any accepted credential can write as any tenant.

Neither setting makes tenants isolated from each other. For hostile-tenant isolation, run separate deployments. See Multi-tenancy.

Deletes are not immediately permanent, and cover one table

Delete tasks rewrite files, but the old files persist in expired snapshots until orphan GC reclaims them, and then as noncurrent S3 versions if versioning is on. Compliance-grade erasure of the table requires the full chain: execute, expire, GC, and expire object versions.

That chain stops at the table. Orphan GC is rooted at the table location, so the WAL mirror prefix, its _active/ blobs, the WAL volume's poison/ quarantine, replica buckets and catalog backups still hold the same events. Each needs its own expiry rule, and no Siglake command sweeps them for you. See Copies the delete chain does not reach.

Delete-task claims and recovery

  • A delete-task claim is never released, including after completion. A process that dies after claiming a task can strand it, and deleting old claim objects is unsafe because a delayed executor might still hold the task.
  • Each task records the table UUID resolved at submission. If the index name now points to a replacement table, the task fails instead of deleting from the new incarnation.
  • A failed task, or a running or claimed pending task stranded by a crash, cannot return to pending. Fix the cause and submit a new task. If the first commit outcome was ambiguous, check the replacement task's deleted-row count before deciding what the first run changed.

Dropping an index does not reclaim its storage

DELETE /api/v1/indexes/{id} records the table UUID, committed-file inventory and UUID-scoped aggregate prefix before removing the catalog entry. The record is report-only, so the dropped table's committed files and aggregates stay in the bucket. Retention and orphan GC both resolve a table through its catalog entry, so neither can reach those files after the drop. Deleting an index frees no space. An index created later with the same id reuses the table location with a fresh table UUID and reads only its own rows. The aggregate inventory stays pinned to the dropped UUID and never resolves the reusable name. See Index management.

A pre-0.1.0 warehouse has to be recreated, not migrated

Tables written before September 6, 2026 are Iceberg format version 3 with a nanosecond timestamp and no timestamp_ns sibling. Siglake still reads them, but no v2-only external engine can. Iceberg cannot change a column's precision and cannot downgrade a v3 table, so there is no in-place upgrade: migrate-schema refuses such a table instead of migrating it, and the remedy is to delete the warehouse and re-ingest. A read-old/write-new rewrite tool is not built. Tables that current builds create are format version 2 with a microsecond timestamptz timestamp and a timestamp_ns column, so this applies only to warehouses that predate the contract. The same boundary bounds a rollback: see What a rollback does and Reverting the image.

Prometheus histogram and counter limits

  • Only *_seconds metrics and four per-call counts have Prometheus histogram buckets: siglake_group_count_deltas_folded, siglake_group_count_tier2_files_per_call, siglake_compactor_mirror_sync_objects, and siglake_compactor_mirror_sync_rotation_objects. Other histogram calls export summaries without _bucket series.
  • Alert counters with fixed labels are registered at zero, so increase() sees their first event. Dynamic label values cannot be registered in advance; siglake_storage_schema_drift_total{column=...} and every siglake_group_count_delta_write_failures_total series outside {iceberg_namespace="siglake",table="events"}, which covers tenant namespaces and index tables, need a second event before an increase() alert sees a delta.

Scale and performance caveats

  • Ordinary log search scales through replication, not per-query fan-out. More replicas add throughput for small-LIMIT browse and Tier-1 aggregate queries, but one replica answers each query. Size this tier for concurrency, not lower latency for one search.
  • Interactive scans slow during a backfill in proportion to compaction lag. The packaged 1 GiB / 2 CPU compactor merges one bin at a time and lags hardest when ingest drives up file-overlap depth. In the 2026-08-27 full-rate 1 TB validation, adequately sizing the compactor reduced the same windowed browse from 40.9 s to 0.13 s without reducing ingest throughput. If interactive search must remain responsive during backfill, follow the compactor sizing recipe.
  • Query throughput plateaus around 16-way concurrency, CPU-bound on browse decode.
  • Layout convergence takes hours. A 2 B-row table needs roughly 4 to 5 hours of background compaction to converge. Query performance on an unconverged layout is substantially worse: 28 QPS versus 115 QPS at 32-way in one measured case.
  • Ingester vertical scaling stops paying past ~2 CPU, because of per-tenant writer-task serialization.
  • The catalog is on the commit path. At high commit rates Postgres becomes the bottleneck.
  • High tenant cardinality is untested. Per-tenant WAL subdirectories and writer tasks mean tenant count drives resource use.

A rare-term match_terms query with no LIMIT is the case for a larger parsed-index cache

One text query shape runs faster with a parsed-index budget above its derived 1 GiB cap: a repeated match_terms query on a rare term, with no LIMIT, over a fully-compacted table. Without a LIMIT the server keeps the index instead of declining it (siglake_query_inverted_index_declined_total does not move), and a 1 GiB budget then evicts it between lookups. A local paired-budget experiment raised SIGLAKE_PARSED_INDEX_CACHE_MAX_BYTES and SIGLAKE_PUFFIN_BLOB_CACHE_MAX_BYTES together, from 1 GiB / 256 MiB to 8 GiB / 2 GiB:

Measurement At 1 GiB / 256 MiB At 8 GiB / 2 GiB
rare_scan p50 22,798.9 ms 120.2 ms
rare_scan_last25 p50 4,843.4 ms 38.2 ms
Total process RSS 14.3-15.6 GiB 20.2-21.3 GiB

Neither knob was varied alone, so the pair is what those latencies cover and they carry no guaranteed production speedup. The RSS figures are the whole process on a local fixture, not the caches' incremental cost and not a pod limit to copy. The run qualified no HTTP, object storage, distributed execution or AWS, and the five clipped query shapes beside these two stayed within noise. Treat the override as a sizing option for a deployment that has this query shape and the memory to spare. Defaults, the 1 GiB derivation cap and the packaged 4 GiB query limit are unchanged. For the knobs themselves, see Tune the Puffin and parsed text-index caches.

Benchmarking caveat

Result caches will silently invalidate a benchmark. A 1 TB round collapsed every query shape onto a ~13 ms result-cache-hit floor, with five orders of magnitude of difference in rows scanned producing latencies within 2 ms of each other.

Set SIGLAKE_QUERY_RESULT_CACHE=off, or use novel literals, and report cold and warm separately. See Performance.

Reporting

Features do land, so parts of this page go out of date. If you find something here that no longer matches the code, open an issue at https://github.com/siglake/siglake/issues.

Public release validation

The renamed v0.1.0, v0.2.0 and v0.2.1 public images passed basic smoke checks on October 8, 2026: image identity, binary startup, Compose health, exact committed rows and grouped counts, query-server restart, and cleanup. The v0.1.0 and v0.2.0 checks used pinned public MinIO fixture substitutions for their unavailable historical dependencies.

These runs verify the renamed release artifacts, not a full burn-in, Kubernetes installation, or upgrade qualification. The release-validation ledger records the evidence and remaining coverage separately from historical install failures.