Skip to content

Limitations

What Siglake does not do, does partially, or does only when you ask. Read this before planning a deployment.

The deferrals on this page follow the Siglake README's "Things deliberately not yet done" at commit 43c1420 from 2026-09-05.

These limits can change a production design:

Features off by default in Helm

These ship disabled. Several of them are features operators assume are on.

Feature Setting Consequence of leaving it off
Query WAL buffer query.walBuffer.enabled: false No seconds-scale freshness. You get commit-cycle visibility.
Catalog claim compactor.catalogClaim.enabled: false Single-pod compaction only. The chart refuses to render a compactor tier that can hold more than one pod.
Attribute auto-promotion SIGLAKE_AUTO_PROMOTE_MIN_PCT=0 Automatic promotion widens the schema from the compactor's sample, and that widening cannot be undone. Hot keys stay in the JSON blob unless declared explicitly.
KEDA autoscaling keda.enabled: false No autoscaling.
Prometheus alert rules prometheusRule.enabled: false No alerts are installed. Enabling the rules still requires the Prometheus Operator and an external alert-delivery workflow.
Source-file batch cache query.scan.fileCacheMaxBytes: 0 and query.scan.fileCacheMaxEntries: 0 The byte limit is reserved from the query memory pool, so a non-zero default would reduce headroom on pods that never reread a file. Warm repeat reads do not reuse decoded source batches. Set both limits to positive values to enable this experimental cache.
ServiceMonitor, PDBs, NetworkPolicies, anti-affinity all false No Prometheus Operator integration, no disruption budgets, no network isolation.

WAL mirroring is enabled by default when you configure a warehouse URL. Each sealed segment uploads asynchronously to <warehouse bucket>/wal-mirror/. An acknowledgement is durable on the local WAL after the default fsync(2); the rows get a copy in object storage when their Iceberg commit or that upload completes, whichever comes first. Mirroring shortens the pre-commit window, so losing the WAL volume costs you the acknowledged rows that have been neither committed nor uploaded. The default single-replica compactor does not reclaim mirror objects, so give the prefix an object-store lifecycle expiry unless you enable compactor.catalogClaim.enabled. Set wal.mirror.enabled: false to opt out.

Leveled compaction is enabled by default and therefore is not listed above. Set SIGLAKE_COMPACTOR_LEVELED=0 only to select the legacy flat whole-partition pass.

Footer inverted indexes are enabled by default. Set SIGLAKE_INVERTED_INDEX=0, Helm compactor.invertedIndex.enabled: false, or operator spec.extraEnv: [{name: SIGLAKE_INVERTED_INDEX, value: "0"}] to opt out.

The post-rewrite Puffin rebuild is disabled by default. A rewrite that could not put an inverted index in the Parquet footer, which is every streamed merge and any column whose index exceeds SIGLAKE_INDEX_FOOTER_MAX_BYTES, leaves that output unindexed, and queries scan it for exact rows. Set SIGLAKE_INDEX_REBUILD=1, Helm compactor.indexRebuild: true, or operator spec.extraEnv: [{name: SIGLAKE_INDEX_REBUILD, value: "1"}] to opt in: the pass then registers a Puffin sidecar after the rewrite commits, skipping files whose columns already carry a footer index or a registration. With inverted indexes off it does nothing either way. Reads do not depend on the switch. Indexes that already exist are still used, and the flush path still writes its footer inverted index.

The rebuild covers inverted indexes only. A streamed merge also skips the whole-file raw trigram bloom, and no later pass registers one, so LIKE '%substr%' on raw prunes against per-row-group blooms alone until an in-memory rewrite covers the file. A delete-task rewrite writes at rewrite generation 0, like an ingest flush, so on a table with index_at_flush: false even its in-memory arm skips the inline index work.

Turning either switch back on backfills nothing already committed: a rebuild only ever sees the files of the rewrite it follows, and only a rewrite of the file itself writes an inline footer index or the whole-file bloom. Files committed while a switch was off keep the metadata they were written with. The cost throughout is pruning rather than correctness: an unindexed file is scanned and row-evaluated, and the answer is exact. See What each merge path writes.

Delete-task execution is on by default too. The compactor runs the sweep in its idle cycle unless SIGLAKE_DELETE_TASKS is 0, off, false or no. The chart ships compactor.deleteTasks: true and renders the variable either way, so false renders SIGLAKE_DELETE_TASKS=0. The operator's compactor renders SIGLAKE_DELETE_TASKS=1 and the CR has no field for it: put SIGLAKE_DELETE_TASKS=0 in spec.extraEnv to opt out, which is also what --adopt-values reports when the chart values file it reads set compactor.deleteTasks: false. With the sweep off, POST /api/v1/delete-tasks still records the task, and siglake delete-sweep --apply still executes it by hand.

OTLP/gRPC is on by default as well. The ingester listens for logs and traces on 0.0.0.0:4317 next to OTLP/HTTP, and both transports apply the same authentication, tenant routing and WAL-fsync acknowledgement contract. The chart ships ingester.otlpGrpc.enabled: true with ingester.otlpGrpc.port: 4317; setting enabled: false drops the gRPC container and Service ports and passes the binary's explicit opt-out flag. Outside Helm, leaving --otlp-grpc-listen unset no longer disables the listener: pass that opt-out flag instead. The operator renders the listener at every replica count, and its CR has no field to turn it off.

Persistent batch jobs are on by default (query.jobs.persistent: true). The query pods keep batch-job state - submissions, their status and their result rows - in the Postgres instance that already holds the Iceberg catalog. The chart renders SIGLAKE_JOBS_POSTGRES_URI from the same Secret, so there is no second connection to configure, and the query server creates its own tables on start. The operator points an adopted query tier at its own spec.catalogUri. One store for the tier is what makes a job_id usable at the shipped query.replicas: 2: the Service has no session affinity, so a status, result or cancel request lands on either pod.

Each job row records the replica incarnation executing it, and owners heartbeat. If a query pod crashes, recovery fails its pending and running rows once the owner's lease expires (SIGLAKE_JOBS_OWNER_LEASE_SECS, default 120 s); jobs owned by a sibling that is still heartbeating are left alone. A planned SIGINT or SIGTERM shutdown deletes that pod's registration after the HTTP server drains, so the next recovery pass can fail its remaining non-terminal rows without waiting out the lease. Neither path resumes interrupted work: the client sees failed and has to resubmit.

Setting query.jobs.persistent: false drops the variable and gives each pod its own in-memory store. It is a supported opt-out for a tier you can hold at one pod, and both deployment paths refuse the rest. The chart fails the install when query.replicas, or keda.query.maxReplicas with keda.query.enabled: true, is above 1 while the store is off. The operator refuses the spec with QueryJobsStoreRequiredForScaleOut when spec.autoscaling.query.max is above 1 and the effective jobs-store URI is blank. Without those guards, the job status, result and cancel reads that land on the pod which did not take the submission answer 404 while the job runs normally elsewhere. The one-pod cost remains: a pod restart loses every job in flight on it.

The catalog claim stays off by default, and it is a requirement rather than a recommendation once you scale the compactor. The chart fails the install on the three combinations that cannot work.

  • More than one compactor pod with the claim off. That is compactor.replicas above 1, or autoscaling.compactor.maxReplicas above 1 when autoscaling.compactor.enabled is true. The shipped HPA ceiling is 4, so turning the compactor HPA on is refused until the claim is on. Uncoordinated replicas share no claim table: each runs the whole drain and maintenance loop against the same tables, and they lose Iceberg commit races to each other while the layout stops converging.
  • The claim with wal.mirror.enabled: false. In claim mode the drain reads the wal_segments catalog table and never the local sealed/ directory, and rows land there only from an ingester that mirrors its segments. That compactor would claim nothing while siglake_compactor_sealed_pending read zero, because in claim mode that gauge counts the shared queue of sealed, unclaimed catalog rows. It cannot reveal a mirrored object that has no catalog row.
  • The enabled compactor HPA custom metric with the claim on. The refusal needs autoscaling.compactor.enabled, autoscaling.compactor.customMetric.enabled and compactor.catalogClaim.enabled all set to true. Its error starts with autoscaling.compactor.customMetric.enabled is true with compactor.catalogClaim.enabled. Every compactor reports the whole shared sealed queue rather than its own share, but the chart renders a per-pod average target. Disable the custom metric for a claim-mode CPU-only HPA, or use siglake-operator for backlog-driven scaling. The custom metric scales nothing in filesystem mode because that mode is capped at one pod.

All three refusals name the value to change. The operator needs none of these chart combinations. It selects the drain from spec.autoscaling.compactor.max, not the current replica count. A maximum above 1 keeps the WAL mirror and catalog claim on at every current replica count, including while the autoscaler holds the compactor at one. A maximum of 1 keeps the filesystem drain.

Not built

Automatic operator drain-mode handover

The operator does not convert WAL between filesystem and catalog-claim drains. Changing spec.autoscaling.compactor.max across 1 changes the ownership protocol, so reconciliation reports DrainModeHandoverRequired and stops before applying workloads. You must finish the current drain, account for local WAL, mirror objects and catalog rows, then follow the manual handover procedure. The operator does not convert WAL or delete those segments and rows for you.

Relevance scoring (BM25)

Search results sort by time. There is no scored TopK.

The reasoning: term frequency degenerates on one-line templated log text, where the same message repeats thousands of times. A full design exists in docs/DESIGN_bm25_scoring.md (per-file df-sketch footers, snapshot-keyed corpus stats, scored TopK) if that calculus changes.

User interface, alerting interface, dashboards

Siglake ships no user interface. Use Grafana over /api/v1/sql and the Jaeger shim, described in Grafana and Jaeger. There is also no built-in alert evaluation, delivery, acknowledgement, or on-call workflow.

Aggregating merge kinds

No rollup or last-write-wins dedup merges exist. The exact rows-conserved commit guard is per-merge-kind-waivable by design, but nothing waives it yet.

Metrics downsampling

None. Retention is the only size-control mechanism.

Row-level retention

Retention drops whole data files whose manifest maximum timestamp is past the horizon. It is file- and day-granular. Your effective retention is the policy plus the span of the last surviving file.

Partial

Query probes do not detect query-path degradation

The query server's /healthz is a constant 200 for as long as the process can serve the probe. /readyz verifies only that the Iceberg catalog can be reached; it does not execute a query or check the cache-warm loop. A pod can therefore pass both probes while browse queries time out, so probe success is not evidence that the query path is healthy.

The shipped signal for the known degradation is SiglakeQueryWarmCycleStalled, not an automatic readiness failure or liveness restart. See Monitoring for the operational consequence. These semantics were confirmed by Siglake's September 3, 2026 review of query degradation and probe behavior.

Tier-1 group counts require first-commit columns

The Tier-1 whole-table group-count aggregate initially covers only columns present since the table's first commit. A table created before typed measurements joined side aggregates can therefore keep serving GROUP BY on those columns through the exact but slower materialized path, even when its files already contain usable group-count footers.

Lost-delta automatic repair restores only the columns named by the failed commit; it does not admit pre-existing typed columns automatically.

Use the group-count repair procedure to admit eligible typed columns with --admit-typed-columns. This does not require rewriting data files unless an older live file lacks both a usable footer and a raw-page representation for the column.

Elasticsearch bulk API support is write-only

Siglake ingests over the Elasticsearch bulk API but has no Elasticsearch query API, and none is planned. _bulk and _cluster/health work. _search, _msearch, _search/scroll, _field_caps and _cat/* are registered and answer 501 with a pointer to POST /api/v1/sql.

Existing Kibana or Elasticsearch query clients will not work. Query with SQL over HTTP instead: see the SQL cookbook.

Query replicas can briefly disagree across a commit

Each replica serves table metadata from its own cache. The default TTL is 5 seconds, and stale metadata can be served for up to 60 seconds while a background refresh runs (SIGLAKE_ICEBERG_METADATA_CACHE_TTL_SECS controls the TTL). Two replicas behind one Service can therefore see different snapshots during that window. Clients that need a consistent sequence of queries must send them to one pod or wait for every replica's cache to converge.

This is a cross-request limit only. Within one distributed query the fan-out is pinned to the coordinator's generation, its snapshot id and schema id, and a worker that cannot resolve that generation refuses its shard (503, reason: "shard_pin_unresolved") rather than contributing rows from another one, so a single answer is never merged across generations. See One generation per fan-out. The cost of that guarantee is availability: while a coordinator's cache is stale on a snapshot the catalog has already expired, its fanned-out queries return 503 until the cache converges (/api/v1/sql/local still answers). That state alerts as SiglakeQueryShardPinUnresolved. Raise compactor.snapshotExpire.retainLast to widen the margin, or lower SIGLAKE_ICEBERG_METADATA_CACHE_TTL_SECS so coordinators stop serving a snapshot before the catalog expires it.

Jaeger is HTTP-only

The shim covers the HTTP subset Grafana's Jaeger data source renders. There is no gRPC SpanReader, so tools expecting that interface will not work.

Kubernetes operator gaps compared with Helm

It renders the core data plane with the same workload kinds as the chart: ingester and compactor as Deployments, query as a StatefulSet. The Helm chart remains the supported, more complete install path.

The operator also renders the WAL PVC, Services, retention CronJobs, and a one-shot schema-migration Job when spec.schemaVersion is set. It does not render Ingress, PDBs, NetworkPolicies, ServiceMonitors, HPA/KEDA objects, or scheduling constraints. Per-tier resources default to the chart's and are overridable through spec.resources. The CR cannot express query bearer tokens, OIDC, TLS, or a WAL-buffer volume.

Catalog credentials are plaintext in the spec.catalogUri custom resource. The operator copies them into pod environment variables and CronJob arguments. The chart instead keeps Secret references and an unexpanded URI in pod specifications.

Adoption of an existing Helm release through --adopt-values is experimental and offline. It has no live test coverage and leaves several values unmapped. The ownership handover (annotation flip plus release-secret removal) is a manual runbook.

Freshness does not cover system tables

The WAL buffer serves the events table and the managed user indexes a query references. It does not serve query_audit or any other system table, which see commit-cycle visibility. When the buffer is off, or a query pod cannot read the WAL mount, every table falls back to commit-cycle visibility.

Strict mapping mode does not reject

"mode": "strict" document mapping enforces at commit time as dynamic-plus-a-counter: it counts violations rather than rejecting documents, and undeclared attributes still land in the residual attributes column. Do not rely on it as a data-quality gate.

Search v1 bounds

  • Full-text pruning engages only on columns that have index blobs. Other columns fall back to row evaluation.
  • Puffin registrations survive expiry of the snapshot that made them. Snapshot expiry keeps the statistics entry in table metadata and the orphan sweep treats the sidecar as reachable, so a long-lived file keeps its sidecar index. Siglake's own expiry keeps the entry; another engine expiring snapshots on the same table may not.

Non-mergeable queries run single-pod

The distributed classifier falls back to DistPlan::Local for shapes it cannot safely split and merge, including non-mergeable aggregates and queries over more than one relation. See the query engine's mergeability rules for the full list.

Fan-out membership converges, it does not switch

Both packages render --query-peer-discovery-srv instead of a static --query-peers list, so the old ceilings are gone: the chart accepts keda.query.maxReplicas > query.replicas and the operator accepts a real spec.autoscaling.query range. What remains is convergence lag. A new pod is eligible for shard work one readiness probe plus one discovery refresh (--query-peer-discovery-interval-secs, default 5 s, with CoreDNS TTL on top) after it starts, and each query pins the membership it captured. Scale-out therefore helps the next query, not the one in flight. A pod whose discovery has never succeeded serves every query single-pod; one whose refreshes have started failing serves a stale membership. Both are visible on siglake_query_peer_discovery_members and siglake_query_peer_discovery_refresh_total, and alert as SiglakeQueryPeerDiscoveryStalled.

Cross-replica job cancellation is a poll

With the default shared job store and more than one query pod, the replica serving DELETE /api/v1/jobs/{id} is usually not the one executing the job. The 202 is terminal for the job row immediately, but the executing replica learns of it by re-reading its own in-flight rows every --jobs-cancel-poll-secs (default 2 seconds); there is no LISTEN/NOTIFY fast path. The 202 therefore bounds when the admission share and the storage scans are released rather than reporting that they already are. See Batch jobs. That watch loop also runs on the batch runtime, alongside the batch futures, so a job occupying a worker without yielding delays its own cancellation.

No rollback is qualified against an older image

Rolling the image back keeps the columns a migration added, and the additive direction is regression tested. Every one of those tests runs the current binary against a table widened past what it declares, so it is evidence about the mechanism and not about a released image. No run has deployed image N, migrated, deployed N-1 and read the result back, and both packages' rollback behavior is reasoned from the chart templates and the reconciler. An image that differs by more than its declared column set is outside even that reasoning. See What a rollback does for Helm and Reverting the image for the operator.

Design decisions with consequences

Query spill is bounded node-local scratch space

Query spill is neither durable nor globally observable storage. PVC-backed spill for sorts larger than a pod's ephemeral-storage budget is not implemented, and Siglake exposes no whole-runtime spill-bytes metric. Query failures and pod ephemeral-storage usage are the operational signals.

Distributed admission is per-coordinator

A distributed query reserves one admission share on its coordinator, at most a quarter of the pod budget, held from admission through the merge. Shard work reserves nothing on workers, whose concurrent heavy shard scans are bounded by the process memory pool and the per-shard rows-scanned breaker instead; they can reach replicas × 4 at once. A cluster-scoped budget or a worker reservation priced at the shard's own share is not implemented.

WAL-append acks, and a shared filesystem

An acknowledgement means the batch is in one node's local WAL, not in object storage. The default wait_for mode fsync(2)s that WAL before it answers, on OTLP/HTTP logs and traces, the Elasticsearch-compatible bulk endpoints, and OTLP/gRPC logs and traces. The sync covers segment bytes and directory entries for the tenant, index, active segment, and sealed segment. Siglake syncs the sealed name before it unlinks the active copy.

The power-loss guarantee assumes ext4 or xfs on a node-attached volume. With EFS or another network filesystem, the guarantee is whatever that filesystem's fsync(2) and rename semantics provide. If you would rather trade durability for latency, send commit=auto (or the equivalent X-Siglake-Commit: auto header) to acknowledge once write(2) has reached the kernel page cache. That write survives a process crash, an OOM kill, and a pod restart, but a node crash or power failure can lose it before the kernel flushes.

Neither mode waits for the rows to become queryable: the compactor commits to Iceberg on its own cadence. Neither waits for the WAL mirror to upload the segment to object storage. An object-store group-commit "PUT-as-ack" model is designed but not built.

Every WAL consumer that may run on a different node needs a ReadWriteMany filesystem (EFS on AWS). That includes a single ingester and a single compactor when they run on separate nodes, plus query pods when the freshness buffer is enabled. ReadWriteOnce is supported only when every WAL consumer is explicitly co-located on one node. RWX brings its own cost, throughput characteristics, and operational surface.

Mirror reconciliation scales with retained objects

In a catalog-claim deployment, one elected compactor repairs mirror objects whose catalog registration was missed. It registers at most 1,024 sealed objects per pass and stores a shared (last_key, rotation) cursor in the mirror_sync_cursors catalog table. The cursor advances only after the whole page registers, survives owner handoff, and clears at end-of-prefix so a later rotation repairs keys inserted behind it. S3 listings use OpenDAL start_after; other backends locally filter the listing as a correctness-first fallback.

The default compactor.committedRetentionSecs: 86400 purges committed objects and catalog rows after 24 hours; 0 is the explicit never-purge opt-out, and non-zero values are floored at 901 seconds. Each retention run drains 512-object pages up to a 16,384-object bound and is paced start-to-start.

siglake_compactor_mirror_sync_objects records each bounded page and should plateau at 1,024 while a rotation is in progress. The whole-rotation metrics are siglake_compactor_mirror_sync_rotation_duration_seconds, siglake_compactor_mirror_sync_rotation_objects, siglake_compactor_mirror_sync_rotations_completed, and siglake_compactor_mirror_sync_last_completed_timestamp_seconds. The first two are recorded when the cursor wraps; the last two are durable gauges that survive owner handoff. Completion age is the backlog or stall signal when every page stays full; a rising retained-per-rotation count with proportional duration means retention is backlogged. See Mirror reconciliation and retention for retention sizing and dashboard guidance.

The Helm value prometheusRule.mirrorRotationStallSecs defaults to 21600 (six hours). It controls how long bounded pages may run without a full rotation completing before SiglakeMirrorReconciliationStalled can fire; 0 disables that alert.

Snapshot expiry limits time travel

Snapshot expiry is on by default and retains the last 100 snapshots, because metadata.json is read and rewritten on every commit. Iceberg time travel reaches no further back than the retained snapshots. Widening the window costs commit throughput.

In catalog-claim deployments, new writers atomically maintain the bounded siglake.consumed_proof.v1 table property, so committed-claim evidence survives snapshot expiry and reclustering. Old writers know only the retained snapshot summaries. Before a rolling upgrade, set compactor.snapshotExpire.retainLast >= 400; keep it there until every old drain/maintenance writer is gone and for another 1,025 seconds. The dual reader uses both sources during that interval. See consumed-proof rolling upgrades for the deployment procedure.

Corrupt or over-cap durable state fails closed: it refuses writes and makes uncovered reclaim decisions unprovable rather than silently discarding positive proof. Reclaim starts after SIGLAKE_CLAIM_RECLAIM_MAX_AGE_SECS (900 seconds by default); 0 is refused and resolves to 900 with a once-per-process warning. When neither the durable proof nor retained snapshot history covers an abandoned segment, the compactor increments siglake_compactor_reclaim_unprovable_total and requeues it, preferring a detectable possible duplicate to silent loss.

Commit batching delays external visibility

Commit-accumulation batching defers commits by up to maxAgeSecs (10 s default). With the optional WAL buffer enabled, Siglake queries can see sealed WAL segments before commit. Acknowledgement does not make a row visible: the segment must first seal, after 4,096 events or 5 seconds by default. External Iceberg readers always wait for the Iceberg commit, and so does every Siglake table when the buffer is off. See OTLP-to-SQL visibility delay.

Multi-tenancy is soft

Tenants share compute, the catalog, the bucket, and object-storage credentials. No routing mode gives a tenant its own compute quota.

Ingest is single-tenant by default. Every request routes to the default tenant, and an X-Scope-OrgID naming another tenant is refused: 403 on HTTP, PermissionDenied on the OTLP/gRPC logs and traces exporters. Naming default is a no-op, so a client that always sends the header keeps working. Routing tenants is something you turn on, with one of two settings:

  • --oidc-tenant-claim / ingester.oidc.tenantClaim takes the tenant from the caller's verified JSON Web Token, on both transports. A header may only agree with the claim, and a token carrying no usable claim is refused.
  • --trust-scope-header / ingester.trustScopeHeader takes the client's word: the header selects the tenant again. Use it where a gateway in front of the ingester sets the header itself and strips the client's; on anything else, any accepted credential can write as any tenant.

Neither setting makes tenants isolated from each other. For hostile-tenant isolation, run separate deployments. See Multi-tenancy.

Deletes are not immediately permanent

Delete tasks rewrite files, but the old files persist in expired snapshots until orphan GC reclaims them, and then as noncurrent S3 versions if versioning is on. Compliance-grade erasure requires the full chain: execute, expire, GC, and expire object versions.

Dropping an index does not reclaim its storage

DELETE /api/v1/indexes/{id} removes the catalog entry; the dropped table's committed files stay in the bucket. Retention and orphan GC both resolve a table through that entry, so once it is gone neither sweep can reach those files. Deleting an index frees no space. An index created later with the same id reuses the table location with a fresh table UUID and reads only its own rows, which is why out-of-band cleanup has to target the dropped table's UUID and file inventory rather than the shared prefix or the index id. See Index management.

A pre-0.1.0 warehouse has to be recreated, not migrated

Tables written before September 6, 2026 are Iceberg format version 3 with a nanosecond timestamp and no timestamp_ns sibling. Siglake still reads them, but no v2-only external engine can. Iceberg cannot change a column's precision and cannot downgrade a v3 table, so there is no in-place upgrade: migrate-schema refuses such a table instead of migrating it, and the remedy is to delete the warehouse and re-ingest. A read-old/write-new rewrite tool is not built. Tables that current builds create are format version 2 with a microsecond timestamptz timestamp and a timestamp_ns column, so this applies only to warehouses that predate the contract. The same boundary bounds a rollback: see What a rollback does and Reverting the image.

Scale and performance caveats

  • Interactive scans slow during a backfill in proportion to compaction lag. The packaged 1 GiB / 2 CPU compactor merges one bin at a time and lags hardest when ingest drives up file-overlap depth. In the 2026-08-27 full-rate 1 TB validation, adequately sizing the compactor reduced the same windowed browse from 40.9 s to 0.13 s without reducing ingest throughput. If interactive search must remain responsive during backfill, follow the compactor sizing recipe.
  • Query throughput plateaus around 16-way concurrency, CPU-bound on browse decode.
  • Layout convergence takes hours. A 2 B-row table needs roughly 4 to 5 hours of background compaction to converge. Query performance on an unconverged layout is substantially worse: 28 QPS versus 115 QPS at 32-way in one measured case.
  • Ingester vertical scaling stops paying past ~2 CPU, because of per-tenant writer-task serialization.
  • The catalog is on the commit path. At high commit rates Postgres becomes the bottleneck.
  • High tenant cardinality is untested. Per-tenant WAL subdirectories and writer tasks mean tenant count drives resource use.

Benchmarking caveat

Result caches will silently invalidate a benchmark. A 1 TB round collapsed every query shape onto a ~13 ms result-cache-hit floor, with five orders of magnitude of difference in rows scanned producing latencies within 2 ms of each other.

Set SIGLAKE_QUERY_RESULT_CACHE=off, or use novel literals, and report cold and warm separately. See Performance.

Reporting

Features do land, so parts of this page go out of date. If you find something here that no longer matches the code, open an issue at https://github.com/siglake/siglake/issues.