Limitations¶
What Siglake does not do, does partially, or does only when you ask. Read this before planning a deployment.
The deferrals on this page follow the Siglake README's "Things deliberately
not yet done" at commit 43c1420 from 2026-09-05.
These limits can change a production design:
- Compaction is single-pod unless you enable the catalog claim.
- Multi-tenancy is soft: tenants share compute, the catalog, and the bucket.
- Retention is file- and day-granular, not row-level.
- Search has no relevance scoring; results sort by time.
- With the default metadata-cache settings, query replicas can disagree across requests for up to 60 seconds.
- Query spill uses bounded node-local scratch space.
- There is no qualified rollback to an older image.
- A warehouse that uses the pre-0.1.0 v3 and nanosecond timestamp contract must be recreated instead of migrated.
- The Elasticsearch-compatible API is write-only.
Features off by default in Helm¶
These ship disabled. Several of them are features operators assume are on.
| Feature | Setting | Consequence of leaving it off |
|---|---|---|
| Query WAL buffer | query.walBuffer.enabled: false |
No seconds-scale freshness. You get commit-cycle visibility. |
| Catalog claim | compactor.catalogClaim.enabled: false |
Single-pod compaction only. The chart refuses to render a compactor tier that can hold more than one pod. |
| Attribute auto-promotion | SIGLAKE_AUTO_PROMOTE_MIN_PCT=0 |
Automatic promotion widens the schema from the compactor's sample, and that widening cannot be undone. Hot keys stay in the JSON blob unless declared explicitly. |
| KEDA autoscaling | keda.enabled: false |
No autoscaling. |
| Prometheus alert rules | prometheusRule.enabled: false |
No alerts are installed. Enabling the rules still requires the Prometheus Operator and an external alert-delivery workflow. |
| Source-file batch cache | query.scan.fileCacheMaxBytes: 0 and query.scan.fileCacheMaxEntries: 0 |
The byte limit is reserved from the query memory pool, so a non-zero default would reduce headroom on pods that never reread a file. Warm repeat reads do not reuse decoded source batches. Set both limits to positive values to enable this experimental cache. |
| ServiceMonitor, PDBs, NetworkPolicies, anti-affinity | all false |
No Prometheus Operator integration, no disruption budgets, no network isolation. |
WAL mirroring is enabled by default when you configure a warehouse URL. Each
sealed segment uploads asynchronously to <warehouse bucket>/wal-mirror/.
An acknowledgement is durable on the local WAL after the default fsync(2);
the rows get a copy in object storage when their Iceberg commit or that upload
completes, whichever comes first. Mirroring shortens the pre-commit window, so
losing the WAL volume costs you the acknowledged rows that have been neither
committed nor uploaded.
The default single-replica compactor does not reclaim mirror objects, so give
the prefix an object-store lifecycle expiry unless you enable
compactor.catalogClaim.enabled. Set wal.mirror.enabled: false to opt out.
Leveled compaction is enabled by default and therefore is not listed above.
Set SIGLAKE_COMPACTOR_LEVELED=0 only to select the legacy flat
whole-partition pass.
Footer inverted indexes are enabled by default. Set
SIGLAKE_INVERTED_INDEX=0, Helm
compactor.invertedIndex.enabled: false, or operator
spec.extraEnv: [{name: SIGLAKE_INVERTED_INDEX, value: "0"}] to opt out.
The post-rewrite Puffin rebuild is disabled by default. A rewrite that
could not put an inverted index in the Parquet footer, which is every streamed
merge and any column whose index exceeds SIGLAKE_INDEX_FOOTER_MAX_BYTES,
leaves that output unindexed, and queries scan it for exact rows. Set
SIGLAKE_INDEX_REBUILD=1, Helm compactor.indexRebuild: true, or operator
spec.extraEnv: [{name: SIGLAKE_INDEX_REBUILD, value: "1"}] to opt in: the
pass then registers a Puffin sidecar after the rewrite commits, skipping files
whose columns already carry a footer index or a registration. With inverted
indexes off it does nothing either way. Reads do not depend on the switch.
Indexes that already exist are still used, and the flush path still writes its
footer inverted index.
The rebuild covers inverted indexes only. A streamed merge also skips the
whole-file raw trigram bloom, and no later pass registers one, so
LIKE '%substr%' on raw prunes against per-row-group blooms alone until an
in-memory rewrite covers the file. A delete-task rewrite writes at rewrite
generation 0, like an ingest flush, so on a table with index_at_flush: false
even its in-memory arm skips the inline index work.
Turning either switch back on backfills nothing already committed: a rebuild only ever sees the files of the rewrite it follows, and only a rewrite of the file itself writes an inline footer index or the whole-file bloom. Files committed while a switch was off keep the metadata they were written with. The cost throughout is pruning rather than correctness: an unindexed file is scanned and row-evaluated, and the answer is exact. See What each merge path writes.
Delete-task execution is on by default too. The compactor runs the sweep in its
idle cycle unless SIGLAKE_DELETE_TASKS is 0, off, false or no. The
chart ships compactor.deleteTasks: true and renders the variable either way,
so false renders SIGLAKE_DELETE_TASKS=0. The operator's compactor renders
SIGLAKE_DELETE_TASKS=1 and the CR has no field for it: put
SIGLAKE_DELETE_TASKS=0 in spec.extraEnv to opt out, which is also what
--adopt-values reports when the chart values file it reads set
compactor.deleteTasks: false. With the sweep off, POST /api/v1/delete-tasks
still records the task, and siglake delete-sweep --apply still executes it by
hand.
OTLP/gRPC is on by default as well. The ingester listens for logs and traces on
0.0.0.0:4317 next to OTLP/HTTP, and both transports apply the same
authentication, tenant routing and WAL-fsync acknowledgement contract. The
chart ships ingester.otlpGrpc.enabled: true with
ingester.otlpGrpc.port: 4317; setting enabled: false drops the gRPC
container and Service ports and passes the binary's explicit opt-out flag.
Outside Helm, leaving --otlp-grpc-listen unset no longer disables the
listener: pass that opt-out flag instead. The operator renders the listener at
every replica count, and its CR has no field to turn it off.
Persistent batch jobs are on by default (query.jobs.persistent: true). The
query pods keep batch-job state - submissions, their status and their result
rows - in the Postgres instance that already holds the Iceberg catalog. The
chart renders SIGLAKE_JOBS_POSTGRES_URI from the same Secret, so there is no
second connection to configure, and the query server creates its own tables on
start. The operator points an adopted query tier at its own
spec.catalogUri. One store for the tier is what makes a job_id usable at
the shipped query.replicas: 2: the Service has no session affinity, so a
status, result or cancel request lands on either pod.
Each job row records the replica incarnation executing it, and owners
heartbeat. If a query pod crashes, recovery fails its pending and running
rows once the owner's lease expires (SIGLAKE_JOBS_OWNER_LEASE_SECS, default
120 s); jobs owned by a sibling that is still heartbeating are left alone. A planned SIGINT or SIGTERM shutdown
deletes that pod's registration after the HTTP server drains, so the next
recovery pass can fail its remaining non-terminal rows without waiting out the
lease. Neither path resumes interrupted work: the client sees failed and has
to resubmit.
Setting query.jobs.persistent: false drops the variable and gives each pod
its own in-memory store. It is a supported opt-out for a tier you can hold at
one pod, and both deployment paths refuse the rest. The chart fails the install
when query.replicas, or keda.query.maxReplicas with keda.query.enabled:
true, is above 1 while the store is off. The operator refuses the spec with
QueryJobsStoreRequiredForScaleOut when spec.autoscaling.query.max is above
1 and the effective jobs-store URI is blank. Without those guards, the job
status, result and cancel reads that land on the pod which did not take the
submission answer 404 while the job runs normally elsewhere. The one-pod cost
remains: a pod restart loses every job in flight on it.
The catalog claim stays off by default, and it is a requirement rather than a recommendation once you scale the compactor. The chart fails the install on the three combinations that cannot work.
- More than one compactor pod with the claim off. That is
compactor.replicasabove1, orautoscaling.compactor.maxReplicasabove1whenautoscaling.compactor.enabledistrue. The shipped HPA ceiling is4, so turning the compactor HPA on is refused until the claim is on. Uncoordinated replicas share no claim table: each runs the whole drain and maintenance loop against the same tables, and they lose Iceberg commit races to each other while the layout stops converging. - The claim with
wal.mirror.enabled: false. In claim mode the drain reads thewal_segmentscatalog table and never the localsealed/directory, and rows land there only from an ingester that mirrors its segments. That compactor would claim nothing whilesiglake_compactor_sealed_pendingread zero, because in claim mode that gauge counts the shared queue of sealed, unclaimed catalog rows. It cannot reveal a mirrored object that has no catalog row. - The enabled compactor HPA custom metric with the claim on. The refusal needs
autoscaling.compactor.enabled,autoscaling.compactor.customMetric.enabledandcompactor.catalogClaim.enabledall set totrue. Its error starts withautoscaling.compactor.customMetric.enabled is true with compactor.catalogClaim.enabled. Every compactor reports the whole shared sealed queue rather than its own share, but the chart renders a per-pod average target. Disable the custom metric for a claim-mode CPU-only HPA, or use siglake-operator for backlog-driven scaling. The custom metric scales nothing in filesystem mode because that mode is capped at one pod.
All three refusals name the value to change. The operator needs none of these
chart combinations. It selects the drain from
spec.autoscaling.compactor.max, not the current replica count.
A maximum above 1 keeps the WAL mirror and catalog claim on at every current
replica count, including while the autoscaler holds the compactor at one. A
maximum of 1 keeps the filesystem drain.
Not built¶
Automatic operator drain-mode handover¶
The operator does not convert WAL between filesystem and catalog-claim drains.
Changing spec.autoscaling.compactor.max across 1 changes the ownership
protocol, so reconciliation reports DrainModeHandoverRequired and stops
before applying workloads. You must finish the current drain, account for local
WAL, mirror objects and catalog rows, then follow the manual handover
procedure. The operator
does not convert WAL or delete those segments and rows for you.
Relevance scoring (BM25)¶
Search results sort by time. There is no scored TopK.
The reasoning: term frequency degenerates on one-line templated log text,
where the same message repeats thousands of times. A full design exists in
docs/DESIGN_bm25_scoring.md (per-file df-sketch footers, snapshot-keyed
corpus stats, scored TopK) if that calculus changes.
User interface, alerting interface, dashboards¶
Siglake ships no user interface. Use Grafana over /api/v1/sql and the Jaeger
shim, described in Grafana and Jaeger. There
is also no built-in alert evaluation, delivery, acknowledgement, or on-call
workflow.
Aggregating merge kinds¶
No rollup or last-write-wins dedup merges exist. The exact rows-conserved commit guard is per-merge-kind-waivable by design, but nothing waives it yet.
Metrics downsampling¶
None. Retention is the only size-control mechanism.
Row-level retention¶
Retention drops whole data files whose manifest maximum timestamp is past the horizon. It is file- and day-granular. Your effective retention is the policy plus the span of the last surviving file.
Partial¶
Query probes do not detect query-path degradation¶
The query server's /healthz is a constant 200 for as long as the process
can serve the probe. /readyz verifies only that the Iceberg catalog can be
reached; it does not execute a query or check the cache-warm loop. A pod can
therefore pass both probes while browse queries time out, so probe success is
not evidence that the query path is healthy.
The shipped signal for the known degradation is
SiglakeQueryWarmCycleStalled, not an automatic readiness failure or liveness
restart. See Monitoring for
the operational consequence. These semantics were confirmed by Siglake's
September 3, 2026 review of query degradation and probe behavior.
Tier-1 group counts require first-commit columns¶
The Tier-1 whole-table group-count aggregate initially covers only columns
present since the table's first commit. A table created before typed
measurements joined side aggregates can therefore keep serving GROUP BY on
those columns through the exact but slower materialized path, even when its
files already contain usable group-count footers.
Lost-delta automatic repair restores only the columns named by the failed commit; it does not admit pre-existing typed columns automatically.
Use the group-count repair
procedure
to admit eligible typed columns with --admit-typed-columns. This does not
require rewriting data files unless an older live file lacks both a usable
footer and a raw-page representation for the column.
Elasticsearch bulk API support is write-only¶
Siglake ingests over the Elasticsearch bulk API but has no Elasticsearch query
API, and none is planned. _bulk and _cluster/health work. _search,
_msearch, _search/scroll, _field_caps and _cat/* are registered and
answer 501 with a pointer to POST /api/v1/sql.
Existing Kibana or Elasticsearch query clients will not work. Query with SQL over HTTP instead: see the SQL cookbook.
Query replicas can briefly disagree across a commit¶
Each replica serves table metadata from its own cache. The default TTL is 5
seconds, and stale metadata can be served for up to 60 seconds while a
background refresh runs
(SIGLAKE_ICEBERG_METADATA_CACHE_TTL_SECS
controls the TTL). Two replicas behind one Service can therefore see different
snapshots during that window. Clients that need a consistent sequence of
queries must send them to one pod or wait for every replica's cache to
converge.
This is a cross-request limit only. Within one distributed query the fan-out
is pinned to the coordinator's generation, its snapshot id and schema id, and
a worker that cannot resolve that generation refuses its shard (503,
reason: "shard_pin_unresolved")
rather than contributing rows from another one, so a single answer is never
merged across generations. See One generation per
fan-out. The cost of that
guarantee is availability: while a coordinator's cache is stale on a snapshot
the catalog has already expired, its fanned-out queries return 503 until the
cache converges (/api/v1/sql/local still answers). That state alerts as
SiglakeQueryShardPinUnresolved.
Raise compactor.snapshotExpire.retainLast
to widen the margin, or lower SIGLAKE_ICEBERG_METADATA_CACHE_TTL_SECS so
coordinators stop serving a snapshot before the catalog expires it.
Jaeger is HTTP-only¶
The shim covers the HTTP subset Grafana's Jaeger data source renders. There is no gRPC SpanReader, so tools expecting that interface will not work.
Kubernetes operator gaps compared with Helm¶
It renders the core data plane with the same workload kinds as the chart: ingester and compactor as Deployments, query as a StatefulSet. The Helm chart remains the supported, more complete install path.
The operator also renders the WAL PVC, Services, retention CronJobs, and a
one-shot schema-migration Job when spec.schemaVersion is set. It does not
render Ingress, PDBs, NetworkPolicies, ServiceMonitors, HPA/KEDA objects, or
scheduling constraints. Per-tier resources default to the chart's and are
overridable through
spec.resources. The CR cannot
express query bearer tokens, OIDC, TLS, or a WAL-buffer volume.
Catalog credentials are plaintext in the spec.catalogUri custom resource.
The operator copies them into pod environment variables and CronJob arguments.
The chart instead keeps Secret references and an unexpanded URI in pod
specifications.
Adoption of an existing Helm release through --adopt-values is experimental
and offline. It has no live test coverage and leaves several values unmapped.
The ownership handover (annotation flip plus release-secret removal) is a
manual runbook.
Freshness does not cover system tables¶
The WAL buffer serves the events table and the managed user indexes a query
references. It does not serve query_audit or any other system table, which
see commit-cycle visibility. When the buffer is off, or a query pod cannot
read the WAL mount, every table falls back to commit-cycle visibility.
Strict mapping mode does not reject¶
"mode": "strict" document mapping enforces at commit time as
dynamic-plus-a-counter: it counts violations rather than rejecting documents,
and undeclared attributes still land in the residual attributes column. Do
not rely on it as a data-quality gate.
Search v1 bounds¶
- Full-text pruning engages only on columns that have index blobs. Other columns fall back to row evaluation.
- Puffin registrations survive expiry of the snapshot that made them. Snapshot expiry keeps the statistics entry in table metadata and the orphan sweep treats the sidecar as reachable, so a long-lived file keeps its sidecar index. Siglake's own expiry keeps the entry; another engine expiring snapshots on the same table may not.
Non-mergeable queries run single-pod¶
The distributed classifier falls back to DistPlan::Local for shapes it cannot
safely split and merge, including non-mergeable aggregates and queries over
more than one relation. See the query engine's
mergeability rules for the
full list.
Fan-out membership converges, it does not switch¶
Both packages render --query-peer-discovery-srv instead of a static
--query-peers list, so the old ceilings are gone: the
chart accepts keda.query.maxReplicas > query.replicas and the operator accepts
a real spec.autoscaling.query range. What remains is convergence lag. A new
pod is eligible for shard work one readiness probe plus one discovery refresh
(--query-peer-discovery-interval-secs, default 5 s, with CoreDNS TTL on top)
after it starts, and each query pins the membership it captured. Scale-out
therefore helps the next query, not the one in flight. A pod whose discovery
has never succeeded serves every query single-pod; one whose refreshes have
started failing serves a stale membership. Both are visible on
siglake_query_peer_discovery_members and
siglake_query_peer_discovery_refresh_total, and alert as
SiglakeQueryPeerDiscoveryStalled.
Cross-replica job cancellation is a poll¶
With the default shared job store and more than one query pod, the replica
serving DELETE /api/v1/jobs/{id} is usually not the one executing the job. The 202
is terminal for the job row immediately, but the executing replica learns of it
by re-reading its own in-flight rows every --jobs-cancel-poll-secs (default
2 seconds); there is no LISTEN/NOTIFY fast path. The 202 therefore
bounds when the admission share and the storage scans are released rather than
reporting that they already are. See Batch
jobs. That watch loop also runs on the
batch runtime, alongside the batch futures, so a job occupying a worker without
yielding delays its own cancellation.
No rollback is qualified against an older image¶
Rolling the image back keeps the columns a migration added, and the additive direction is regression tested. Every one of those tests runs the current binary against a table widened past what it declares, so it is evidence about the mechanism and not about a released image. No run has deployed image N, migrated, deployed N-1 and read the result back, and both packages' rollback behavior is reasoned from the chart templates and the reconciler. An image that differs by more than its declared column set is outside even that reasoning. See What a rollback does for Helm and Reverting the image for the operator.
Design decisions with consequences¶
Query spill is bounded node-local scratch space¶
Query spill is neither durable nor globally observable storage. PVC-backed spill for sorts larger than a pod's ephemeral-storage budget is not implemented, and Siglake exposes no whole-runtime spill-bytes metric. Query failures and pod ephemeral-storage usage are the operational signals.
Distributed admission is per-coordinator¶
A distributed query reserves one admission share on its coordinator, at most a
quarter of the pod budget, held from admission through the merge. Shard work
reserves nothing on workers, whose concurrent heavy shard scans are
bounded by the process memory pool and the per-shard rows-scanned breaker
instead; they can reach replicas × 4 at once. A cluster-scoped budget or a
worker reservation priced at the shard's own share is not implemented.
WAL-append acks, and a shared filesystem¶
An acknowledgement means the batch is in one node's local WAL, not in object
storage. The default wait_for mode fsync(2)s that WAL before it answers, on
OTLP/HTTP logs and traces, the Elasticsearch-compatible bulk endpoints, and
OTLP/gRPC logs and traces. The sync covers segment bytes and directory entries
for the tenant, index, active segment, and sealed segment. Siglake syncs the
sealed name before it unlinks the active copy.
The power-loss guarantee assumes ext4 or xfs on a node-attached volume. With
EFS or another network filesystem, the guarantee is whatever that filesystem's
fsync(2) and rename semantics provide. If you would rather trade durability
for latency, send commit=auto (or the equivalent X-Siglake-Commit: auto
header) to acknowledge once write(2) has reached the kernel page cache. That
write survives a process crash, an OOM kill, and a pod restart, but a node
crash or power failure can lose it before the kernel flushes.
Neither mode waits for the rows to become queryable: the compactor commits to Iceberg on its own cadence. Neither waits for the WAL mirror to upload the segment to object storage. An object-store group-commit "PUT-as-ack" model is designed but not built.
Every WAL consumer that may run on a different node needs a ReadWriteMany
filesystem (EFS on AWS). That includes a single ingester and a single compactor
when they run on separate nodes, plus query pods when the freshness buffer is
enabled. ReadWriteOnce is supported only when every WAL consumer is
explicitly co-located on one node. RWX brings its own cost, throughput
characteristics, and operational surface.
Mirror reconciliation scales with retained objects¶
In a catalog-claim deployment, one elected compactor repairs mirror objects
whose catalog registration was missed. It registers at most 1,024 sealed
objects per pass and stores a shared (last_key, rotation) cursor in the
mirror_sync_cursors catalog table. The cursor advances only after the whole
page registers, survives owner handoff, and clears at end-of-prefix so a later
rotation repairs keys inserted behind it. S3 listings use OpenDAL
start_after; other backends locally filter the listing as a correctness-first
fallback.
The default compactor.committedRetentionSecs: 86400 purges committed objects
and catalog rows after 24 hours; 0 is the explicit never-purge opt-out, and
non-zero values are floored at 901 seconds. Each retention run drains
512-object pages up to a 16,384-object bound and is paced start-to-start.
siglake_compactor_mirror_sync_objects records each bounded page and should
plateau at 1,024 while a rotation is in progress. The whole-rotation metrics
are siglake_compactor_mirror_sync_rotation_duration_seconds,
siglake_compactor_mirror_sync_rotation_objects,
siglake_compactor_mirror_sync_rotations_completed, and
siglake_compactor_mirror_sync_last_completed_timestamp_seconds. The first two
are recorded when the cursor wraps; the last two are durable gauges that
survive owner handoff. Completion age is the backlog or stall signal when every
page stays full; a rising retained-per-rotation count with proportional
duration means retention is backlogged. See Mirror reconciliation and retention
for retention sizing and dashboard guidance.
The Helm value prometheusRule.mirrorRotationStallSecs defaults to 21600
(six hours). It controls how long bounded pages may run without a full rotation
completing before SiglakeMirrorReconciliationStalled can fire; 0 disables
that alert.
Snapshot expiry limits time travel¶
Snapshot expiry is on by default and retains the last 100 snapshots, because
metadata.json is read and rewritten on every commit. Iceberg time travel
reaches no further back than the retained snapshots. Widening the window costs
commit throughput.
In catalog-claim deployments, new writers atomically maintain the bounded
siglake.consumed_proof.v1 table property, so committed-claim evidence survives
snapshot expiry and reclustering. Old writers know only the retained snapshot
summaries. Before a rolling upgrade, set
compactor.snapshotExpire.retainLast
>= 400; keep it there until every old drain/maintenance writer is gone and
for another 1,025 seconds. The dual reader uses both sources during that
interval. See consumed-proof rolling
upgrades for the
deployment procedure.
Corrupt or over-cap durable state fails closed: it refuses writes and makes
uncovered reclaim decisions unprovable rather than silently discarding positive
proof. Reclaim starts after SIGLAKE_CLAIM_RECLAIM_MAX_AGE_SECS (900 seconds by
default); 0 is refused and resolves to 900 with a once-per-process warning.
When neither the durable proof nor retained snapshot history covers an
abandoned segment, the compactor increments
siglake_compactor_reclaim_unprovable_total and requeues it, preferring a
detectable possible duplicate to silent loss.
Commit batching delays external visibility¶
Commit-accumulation batching defers commits by up to maxAgeSecs (10 s
default). With the optional WAL buffer enabled, Siglake queries can see sealed
WAL segments before commit. Acknowledgement does not make a row visible: the
segment must first seal, after 4,096 events or 5 seconds by default. External
Iceberg readers always wait for the Iceberg commit, and so does every Siglake
table when the buffer is off. See OTLP-to-SQL visibility
delay.
Multi-tenancy is soft¶
Tenants share compute, the catalog, the bucket, and object-storage credentials. No routing mode gives a tenant its own compute quota.
Ingest is single-tenant by default. Every request routes to the default
tenant, and an X-Scope-OrgID naming another tenant is refused: 403 on
HTTP, PermissionDenied on the OTLP/gRPC logs and traces exporters. Naming
default is a no-op, so a client that always sends the header keeps working.
Routing tenants is something you turn on, with one of two settings:
--oidc-tenant-claim/ingester.oidc.tenantClaimtakes the tenant from the caller's verified JSON Web Token, on both transports. A header may only agree with the claim, and a token carrying no usable claim is refused.--trust-scope-header/ingester.trustScopeHeadertakes the client's word: the header selects the tenant again. Use it where a gateway in front of the ingester sets the header itself and strips the client's; on anything else, any accepted credential can write as any tenant.
Neither setting makes tenants isolated from each other. For hostile-tenant isolation, run separate deployments. See Multi-tenancy.
Deletes are not immediately permanent¶
Delete tasks rewrite files, but the old files persist in expired snapshots until orphan GC reclaims them, and then as noncurrent S3 versions if versioning is on. Compliance-grade erasure requires the full chain: execute, expire, GC, and expire object versions.
Dropping an index does not reclaim its storage¶
DELETE /api/v1/indexes/{id} removes the catalog entry; the dropped table's
committed files stay in the bucket. Retention and orphan GC both resolve a
table through that entry, so once it is gone neither sweep can reach those
files. Deleting an index frees no space. An index created later with the same
id reuses the table location with a fresh table UUID and reads only its own
rows, which is why out-of-band cleanup has to target the dropped table's UUID
and file inventory rather than the shared prefix or the index id. See Index
management.
A pre-0.1.0 warehouse has to be recreated, not migrated¶
Tables written before September 6, 2026 are Iceberg format version 3 with a
nanosecond timestamp and no timestamp_ns sibling. Siglake still reads them,
but no v2-only external engine can. Iceberg cannot change a column's precision
and cannot downgrade a v3 table, so there is no in-place upgrade:
migrate-schema refuses such a table instead of migrating it, and the remedy
is to delete the warehouse and re-ingest. A read-old/write-new rewrite tool is
not built. Tables that current builds create are format version 2 with a
microsecond timestamptz timestamp and a timestamp_ns column, so this
applies only to warehouses that predate the contract. The same boundary bounds
a rollback: see What a rollback
does and Reverting the
image.
Scale and performance caveats¶
- Interactive scans slow during a backfill in proportion to compaction lag. The packaged 1 GiB / 2 CPU compactor merges one bin at a time and lags hardest when ingest drives up file-overlap depth. In the 2026-08-27 full-rate 1 TB validation, adequately sizing the compactor reduced the same windowed browse from 40.9 s to 0.13 s without reducing ingest throughput. If interactive search must remain responsive during backfill, follow the compactor sizing recipe.
- Query throughput plateaus around 16-way concurrency, CPU-bound on browse decode.
- Layout convergence takes hours. A 2 B-row table needs roughly 4 to 5 hours of background compaction to converge. Query performance on an unconverged layout is substantially worse: 28 QPS versus 115 QPS at 32-way in one measured case.
- Ingester vertical scaling stops paying past ~2 CPU, because of per-tenant writer-task serialization.
- The catalog is on the commit path. At high commit rates Postgres becomes the bottleneck.
- High tenant cardinality is untested. Per-tenant WAL subdirectories and writer tasks mean tenant count drives resource use.
Benchmarking caveat¶
Result caches will silently invalidate a benchmark. A 1 TB round collapsed every query shape onto a ~13 ms result-cache-hit floor, with five orders of magnitude of difference in rows scanned producing latencies within 2 ms of each other.
Set SIGLAKE_QUERY_RESULT_CACHE=off, or use novel literals, and report cold
and warm separately. See Performance.
Reporting¶
Features do land, so parts of this page go out of date. If you find something here that no longer matches the code, open an issue at https://github.com/siglake/siglake/issues.