Helm¶
Use the chart in deploy/helm/siglake to deploy Siglake on Kubernetes. This
guide covers the settings you must choose, the checks that protect an upgrade
and the failure modes that need operator action.
Install¶
helm install siglake ./deploy/helm/siglake \
--namespace siglake-system --create-namespace \
-f my-values.yaml
A minimal my-values.yaml:
tenant:
namespace: siglake
image:
repository: ghcr.io/siglake/siglake
tag: "" # empty = Chart.appVersion
serviceAccount:
create: true
annotations:
eks.amazonaws.com/role-arn: arn:aws:iam::123456789012:role/siglake-warehouse-rw
postgres:
existingSecret: siglake-postgres # keys: host, port, user, password, database
s3:
bucket: my-siglake-warehouse
region: us-east-1
warehousePrefix: warehouse
wal:
size: 200Gi
accessMode: ReadWriteMany
storageClassName: efs-sc
The pod specification holds Secret references for the PG* variables and an
unexpanded SIGLAKE_CATALOG_URI. Kubernetes resolves the references and
substitutes the URI in the container environment, with no shell wrapper. If
your secret uses different field names (AWS Secrets Manager exports host as
endpoint, for instance), remap them under postgres.secretKeys.
Namespaces and tenancy¶
Each Helm release pins itself to one Iceberg namespace (tenant.namespace,
default siglake). Running multiple releases against the same catalog and
warehouse gives schema-level isolation between environments.
This namespace is separate from ingest tenant routing. Ingest uses the
default tenant unless you configure routing under ingester. See
Multi-tenancy.
Per-role shape¶
Every role accepts the same keys: enabled, replicas, resources,
nodeSelector, tolerations, affinity, extraEnv, extraArgs.
These per-role extraEnv values reach only that role; global extraEnv also
reaches the schema-migration hook Job.
| Role | Workload kind | Default replicas |
|---|---|---|
ingester |
Deployment | 1 |
compactor |
Deployment | 1 |
query |
StatefulSet + headless Service | 2 |
Rolling upgrade with Helm: order, consumed-proof upgrade, schema migration¶
The safe rolling-upgrade procedure with Helm covers order, consumed-proof upgrades and schema migrations. You need both chart directories and the release's values file.
- Raise
compactor.snapshotExpire.retainLastto400with the current chart and image. Wait for it before any new compactor starts. - Set
image.tagto the target version and runhelm upgradewith the target chart. - Keep
schemaMigration.enabled: true. Helm runs the schema migration as apre-upgradeJob and holds every workload until it succeeds. Wait for the Job and readmigrate-schema completein its log. - Watch the ingester, compactor and query rollouts. The migration is the only barrier between them.
- Hold
retainLast >= 400until the last old drain exits, then 1,025 seconds more. Confirmsiglake_consumed_proof_bytesper table, then restore normal retention. - If a workload fails after the migration, keep
retainLastat400and see What a rollback does.
Step 1: raise snapshot retention before the image change¶
Before you change the image, raise snapshot retention with the chart and image that the release runs now:
Apply that value with the current chart. Do not change image.tag in this
command:
helm upgrade siglake <current-chart> --namespace siglake-system -f my-values.yaml --wait --wait-for-jobs --timeout 10m
Wait for this command to finish before you start a new binary. The value change restarts the current compactor with the required retention guard.
Step 2: start the upgrade with the target chart¶
In my-values.yaml, set image.tag to the target version and leave
compactor.snapshotExpire.retainLast at 400. If you set query.image.tag,
clear it so the query StatefulSet inherits the target global image. Start the
upgrade with the target chart:
helm upgrade siglake <target-chart> --namespace siglake-system -f my-values.yaml --wait --wait-for-jobs --timeout 10m
Step 3: verify the schema-migration Job¶
Verify the schema-migration Job before you accept the rollout. The
pre-upgrade hook uses the target global image and runs migrate-schema
--all-tables --all-namespaces. Helm does not apply the workload changes until
this Job succeeds. The successful log ends with
migrate-schema complete: <N> column(s) added:
kubectl wait --namespace siglake-system --for=condition=complete job/siglake-migrate-schema-<revision> --timeout=10m
kubectl logs --namespace siglake-system job/siglake-migrate-schema-<revision>
If the Job fails, the old workloads remain active. Read its logs, fix the catalog or object-store access, and retry the upgrade. A failed hook does not need a rollback.
Step 4: verify the three workloads¶
Verify all three workloads after the hook completes:
kubectl rollout status --namespace siglake-system deployment/siglake-ingester --timeout=10m
kubectl rollout status --namespace siglake-system deployment/siglake-compactor --timeout=10m
kubectl rollout status --namespace siglake-system statefulset/siglake-query --timeout=10m
The migration is the only cross-component barrier. After it succeeds,
Kubernetes can reconcile the ingester, compactor and query workloads at the
same time. The chart does not enforce an order among them. The ingester and
query tiers roll while the compactor uses Recreate; see What rolls, and
how.
Step 5: hold retention through the consumed-proof overlap¶
Keep retainLast >= 400 until every old drain or maintenance process has
exited, then keep it for another 1,025 seconds. New compactors write the
durable siglake.consumed_proof.v1 property and the legacy snapshot summary.
They also read both sources. Old compactors continue to use the legacy
summary, so the versions can overlap during this window. If you run a
maintenance process outside the chart, its last old instance also starts this
timer.
For each table receiving writes, confirm that Prometheus reports
siglake_consumed_proof_bytes. Confirm that
increase(siglake_consumed_proof_cap_refusals_total[30m]) remains 0. The
first metric appears after the new compactor updates a table's durable proof.
The second catches updates refused at the property's size limit.
After the interval, restore the retention value you chose for normal operation
and run helm upgrade again with the target chart. The default is 100. This
final upgrade runs the idempotent migration hook again and restarts the
compactor with the lower value.
Step 6: roll back a failed workload¶
If a workload fails after the migration succeeded, keep retainLast at 400
and use What a rollback does. Helm rolls the images
back but leaves additive schema columns in place. Restart the 1,025-second
timer when the last old drain or maintenance process exits again.
The settings that matter most¶
Schema migrations¶
schemaMigration:
enabled: true
backoffLimit: 4
ttlSecondsAfterFinished: 86400
resources:
requests:
memory: 64Mi
cpu: 50m
limits:
memory: 256Mi
cpu: 500m
On helm upgrade, schemaMigration.enabled renders a pre-upgrade Job using
the new image. It runs
siglake migrate-schema --all-tables --all-namespaces, including every tenant
namespace, and Helm waits for it before rolling the workloads. The migration
is additive and idempotent; if the Job fails, the upgrade stops before a newer
binary can serve writes against an older table. Fresh installs create tables
at the current schema and do not need the hook.
A column the Job adds is queryable as soon as the migration commits, on distributed queries as well as single-pod ones: a coordinator pins each fan-out to the schema id it planned against, so a query replica whose metadata cache predates the migration refreshes onto the pinned schema instead of planning its shard against the older column set (One generation per fan-out). Workers running an image older than the one that added the pinned schema id ignore it, so during the rollout that guarantee holds only once every query pod runs the new image.
The job-migrate-schema.yaml
container inherits the chart's s3.* settings, postgres.existingSecret, and
global extraEnv, but not ingester.extraEnv, compactor.extraEnv, or
query.extraEnv. Put anything needed to reach the object store or catalog
during an upgrade in those shared settings; credentials supplied only to a
role can leave the hook retrying until Helm reports BackoffLimitExceeded.
The check_hook_credentials render-time guard catches credentials present on
workloads but missing from the hook.
backoffLimit controls the Job's retry limit, and
ttlSecondsAfterFinished controls how long Kubernetes retains the completed
Job. resources sets the migration container's requests and limits. Leave
enabled on unless another deployment controller guarantees that the same
migration completes before rollout.
A newer binary refuses writes that populate a column the table lacks, naming
the column and migration remedy instead of silently discarding the value. When
prometheusRule.enabled is on, the chart also installs the critical
SiglakeSchemaWritesRefused alert for this condition.
What a rollback does¶
The migration Job does not re-run. It is a pre-upgrade hook, and a rollback
fires only pre-rollback and post-rollback hooks, neither of which the chart
declares.
A rollback rolls back the image, never the schema. The columns the migration
added stay, and that is the direction the storage layer supports: the
rolled-back binary writes the columns it declares, the storage layer fills the
ones it does not with nulls, and rows that already carry values keep them.
Re-running migrate-schema on the old image adds and removes nothing, and it
does not lower the table's recorded schema version, because the migration
stamps the higher of the recorded and declared versions. An image built before
that behavior shipped (anything pre-0.1.0) stamps its own lower constant
instead. That is a reporting artifact: it changes no column and no write
decision, and rolling forward restamps it.
Rolling forward again renders a fresh Job. The Job name carries
.Release.Revision, which a rollback increments, so the next helm upgrade
never reuses a name. A Job's spec.template is immutable, so reuse would fail
with a 422 as soon as the image tag changed. Hook Jobs are not part of the
release manifest, so a rollback neither deletes nor recreates the ones already
in the namespace; schemaMigration.ttlSecondsAfterFinished (default 86400)
reaps them.
A migration that fails needs no rollback. Helm blocks on the pre-upgrade
hook, so the release fails before any workload is applied and the previous
revision is still the live one. Read the Job's logs before you delete it:
The next attempt's Job has a different name either way.
Under the operator, reverting spec.image runs one more migration Job on the
old binary instead. See Reverting the
image.
Two limits bound all of this. It covers additive column changes only: there is
no path across the pre-0.1.0 nanosecond timestamp contract, which
migrate-schema refuses rather than migrating. And no rollback has been
qualified against an actual older image. The regression runs one binary against
a table widened past what it declares, which is evidence about the mechanism,
not about a released image; the rest of this section follows from the chart
templates.
What rolls, and how¶
The ingester Deployment uses RollingUpdate with maxUnavailable: 0.
Upgrades do not reduce ingest capacity. The query StatefulSet also uses RollingUpdate and
replaces one ordinal at a time. If a peer being replaced does not answer an
in-flight query, the coordinator runs that shard itself and increments
siglake_query_coordinator_failover_total.
The chart uses Recreate for the compactor regardless of replica count, so
drain and maintenance pause while the pod restarts. Sealed segments wait on the
WAL PVC; acknowledgement happens in the ingester, so the restart does not lose
acknowledged data. A compactor cannot safely surge against the shared local WAL
because two pods can claim the same segment; multi-pod
compaction instead coordinates claims through the
catalog.
A restart that interrupts a commit needs no operator in the ordinary case. The in-flight commit is all or nothing, and the restarted drain decides what to do with the segment from the table's consumed-segment record. See The compactor dies mid-commit for the disposition rules, the two metrics that count them, and the one case that does ask for a decision.
Freshness¶
Without this, query pods do not mount the WAL and you get commit-cycle
visibility (tens of seconds), not the seconds-scale freshness Siglake is
designed for. The query pods read the same WAL as the ingester and compactor,
so the claim must be RWX when those pods can run on different nodes. The chart
defaults wal.accessMode to ReadWriteMany and wal.storageClassName to the
cluster default; if that StorageClass only supports RWO (as EBS-backed classes
commonly do), the RWX claim remains pending.
query.hotCaches.enabled (the last_values() / distinct_values() UDTFs)
tails the same mount and needs walBuffer too.
Disaster recovery¶
These defaults upload each sealed segment to
<warehouse bucket>/wal-mirror/. Acknowledgements do not wait for the upload:
the default wait_for response means the event is fsynced locally, and it has
no copy in the warehouse bucket until its Iceberg commit or a mirror upload
completes, whichever comes first. A committed row keeps its warehouse copy with
the mirror off, so mirroring is what protects the rows still waiting for a
commit. Set activeIntervalSecs above 0 to upload the open segment on that
cadence.
Successful uploads set that recovery target only when you recover a filesystem
drain. A delayed or failed
upload extends the loss window past the configured interval. If you recover a
catalog-claim drain,
reconciliation recovers sealed objects and gets no active-snapshot target. See
mirror-failure monitoring.
The default single-replica compactor drains the local WAL and does not reclaim
mirror objects. Give wal-mirror/ an object-store lifecycle expiry unless you
enable compactor.catalogClaim.enabled. Set wal.mirror.enabled: false to opt
out. The chart then renders an empty SIGLAKE_WAL_MIRROR_PREFIX; omitting that
variable leaves the binary's warehouse-backed mirror running.
| Value | Default | Purpose |
|---|---|---|
compactor.committedRetentionSecs |
86400 |
Purge successfully drained mirror objects and catalog rows after this many seconds; 0 disables purging. |
prometheusRule.mirrorRotationStallSecs |
21600 |
Alert when bounded reconciliation pages run for this long without a full rotation completing; 0 disables the alert. |
Non-zero committed retention is floored at 901 seconds so it outlives the
default 600-second local-WAL settle delay plus one 300-second sweep cadence.
Raise it beyond the sum if either window is increased. Keep it at 0 when the
local-WAL sweep is disabled but mirror catch-up remains enabled. See
Mirror reconciliation and retention
for the scaling consequence of retaining the full prefix.
Multi-pod compaction¶
compactor:
replicas: 2
catalogClaim:
enabled: true
batch: 16
extraArgs: ["--role", "drain"]
wal:
mirror:
enabled: true # required: the compactor reads segments from the bucket
Catalog claim switches the compactor from listing the local sealed/
directory to the wal_segments SQL claim table. It is required above one pod,
not recommended: the chart fails the install when compactor.replicas is above
1, or when autoscaling.compactor.maxReplicas is above 1 with
autoscaling.compactor.enabled: true, while
compactor.catalogClaim.enabled is false. Replicas that share no claim table
each run the whole drain and maintenance loop against the same tables, lose
Iceberg commit races to each other, and leave the layout unconverged.
The install fails the other way round as well: the claim with
wal.mirror.enabled: false. Only an ingester that mirrors its segments writes
wal_segments rows, so a claim drain without the mirror claims nothing and
compaction stops.
The chart also refuses the compactor's custom backlog metric when all three of
these values are true:
autoscaling:
compactor:
enabled: true
customMetric:
enabled: true
compactor:
catalogClaim:
enabled: true
The error starts with autoscaling.compactor.customMetric.enabled is true with
compactor.catalogClaim.enabled. In claim mode,
siglake_compactor_sealed_pending is the whole shared queue of sealed,
unclaimed segments. It is not one pod's share. The chart renders a Pods metric
with an averageValue target, so the requested replica count grows with the
current replica count instead of tracking the backlog. Disable customMetric
to keep the claim-mode CPU HPA, or use siglake-operator for
backlog-driven scaling. In filesystem mode, the existing guard holds the
compactor at one pod, so the custom metric cannot scale that mode either.
The chart has no first-class compactor.role value yet. Set the binary role
with compactor.extraArgs, as above, or use compactor.extraEnv when
environment variables are easier to manage:
The chart renders one compactor Deployment, so this recipe makes all its
replicas drains. Run the single maintenance process as a separately managed
Deployment with --role maintenance; see Scaling.
Leave the role unset for the default single-process combined behavior.
Snapshot retention¶
Expiry is on by default. retainLast counts snapshots rather than time. Only a
newer commit displaces an older snapshot from the retained window, so a
busy table can cycle through 100 snapshots quickly while an idle one keeps its
oldest retained snapshot indefinitely. intervalSecs is how often the
compactor trims, not a wall-clock erasure deadline. Every drain, compaction,
maintenance and delete commit produces a snapshot, so the window's real
duration is a function of that table's commit rate.
Raising it has two costs:
- Every commit reads and rewrites
metadata.json. Each retained snapshot adds an entry, so a larger window increases work on every commit. - Expiry changes metadata only.
gc-orphanscan delete a data file only after no retained snapshot references it. A wider window therefore holds superseded files, including files rewritten by a delete, in the warehouse for longer. See The full sequence to make a delete permanent.
Two needs can pull the value above the packaged default of 100. Use the higher value while both apply:
- Iceberg time travel and long-running external queries reach only as far back as the retained
window (Snapshot expiry limits time
travel). A
mixed-version rollout needs
>= 400temporarily. See Consumed-proof rolling upgrades below, which is a bounded guard, not a steady-state recommendation. - A coordinator can serve table metadata that is stale by
the cache TTL (5 seconds by default, up to 60 seconds while a background
refresh runs). If the catalog expires a snapshot a coordinator is still
pinning fan-outs to, workers refuse their shards with
503andreason: "shard_pin_unresolved". The margin you need is the number of commits that land in that staleness window, so estimate commits per minute for the busiest table and keepretainLastcomfortably above it. The other half of that remedy is loweringSIGLAKE_ICEBERG_METADATA_CACHE_TTL_SECS. See A 503 withreason: shard_pin_unresolved. A pin also names the schema id, and retention does not widen that half. A schema miss right aftermigrate-schemaclears when the lagging replica refreshes its metadata.
The packaged 100 is a fixed value, not one derived from the workload the way
compactor.binConcurrency is. Pick it from whichever consumer above needs the
longest window, and revisit it when the commit rate changes. intervalSecs: 0
disables trimming entirely, which adds a growing cost to every commit.
Changing retainLast in either direction updates metadata without rewriting
data files. Raising it preserves evidence from snapshots that have not expired,
but cannot restore expired snapshots. Keep an ambiguous filesystem orphan held
until available consumption proof or an operator comparison establishes its
disposition. You can lower retainLast after the incident, subject to the
rollout guard below.
Consumed-proof rolling upgrades¶
Rolling ingesters and drains through a version change safely takes one retention setting, applied before the first new binary starts. Old and new writers record a claim's consumption differently, so the reader needs both sources retained while the two versions overlap.
Versions that write siglake.consumed_proof.v1 can reclaim committed claims
after their source snapshots expire. Older writers update only the legacy
snapshot summary. Before starting the first new binary, set:
Keep compactor.snapshotExpire.retainLast >= 400 throughout the rollout and
for at least 1,025 seconds after the last old drain or maintenance writer
exits. The new reader consults both sources during that interval. After the
wait, retainLast may return to the value Snapshot
retention arrives at for the deployment (the packaged
default is 100); no data-file rewrite is required.
Compactor sizing¶
The packaged compactor merges one bin at a time:
When SIGLAKE_COMPACTOR_BIN_CONCURRENCY is not set, Siglake derives safe bin
concurrency from the cgroup memory limit and available cores. The chart sets it
explicitly through compactor.binConcurrency; raise that value, memory, and
CPU together when backfill makes compaction lag. Budget a 1 GiB process reserve
plus roughly 4 GiB per concurrent bin and at least one CPU per bin. Raising
concurrency without proportional memory leads to OOM kills rather than more
throughput.
Query memory budget¶
Treat query.resources.limits.memory as the query server's sizing budget, not
just its kill ceiling. Use Size query pods and understand distributed
fan-out to set
the pod count after you choose this budget. The chart
values
divide the cgroup limit this way with the chart defaults:
- 25% for the byte-range object cache, clamped from 64 MiB to 16 GiB;
- 12.5% for metadata caches, clamped from 64 MiB to 2 GiB;
- no space for the experimental source-file batch cache unless both
query.scan.fileCacheMaxBytesandquery.scan.fileCacheMaxEntriesare set to positive values; - 1/16 of the limit for parsed text indexes, capped at 1 GiB, and 1/64 for their Puffin blobs, capped at 256 MiB, taken only from what remains once the pool can still reserve one compacted file's 1.25 GiB decode working set;
- 50% of the memory left after the object, metadata, text-index and enabled source-file caches for DataFusion's shared memory pool;
- the rest for the process, in-flight Arrow batches, decode buffers and allocator slack.
For a default 4 GiB pod, that is a 1 GiB object cache, 512 MiB of metadata caches, no text-index caches and a 1.25 GiB DataFusion memory pool. About 1.25 GiB remains for the process and transient allocations. If you enable the source-file cache, its byte limit is also subtracted before applying the pool fraction. A 12.5% share of the pod limit is a reasonable starting budget, not a derived default.
The text-index caches get nothing at 4 GiB because the pool takes half of what
the other caches leave, so holding the 1.25 GiB decode reservation costs
2.5 GiB of the limit and the floor pod has no remainder to divide. Raise the
limit to 5 GiB and they derive 400 MiB; at 16 GiB they reach both caps. Both
budgets are published on siglake_cache_budget_bytes{kind="text_index"}, and
tuning the text-index
caches
covers the two environment variables that override them.
DataFusion sorts and aggregates spill when they reach the pool bound. If it
cannot stay within the memory and spill bounds, Siglake reports
ResourcesExhausted as a retryable 503. The pool reduces OOM risk, but it
does not account for every process allocation.
Raise the memory limit for scan-heavy or high-concurrency workloads. Use
SIGLAKE_QUERY_MEMORY_FRACTION to change the pool's default 0.5 share, and
watch siglake_query_memory_pool_bytes.
siglake_query_scan_decode_reservation_total{outcome="unreserved"} counts
scans that fell to one file at a time and then proceeded without a pool
reservation, so their decode memory goes unaccounted. Read it as a coverage
signal, not a latency one, and do not alert on it as a latency proxy: wider
fan-out asks the pool for more, so in the 2026-09-06 decode sweep the fastest
arm also had the highest unreserved rate.
What the packaged 4 GiB query pod measured on the 50 GB suite¶
The 2026-09-15 50 GB benchmark ran the query server under the packaged 4 GiB limit, and the pod held it: the cgroup kept the ceiling through both suites, Docker reported no out-of-memory (OOM) kill and no restart, and the budgets came out as the arithmetic above predicts: a 1 GiB object cache, 512 MiB of metadata caches, no text-index caches and a 1.25 GiB pool. A later reread of that round against the recorded unlimited-pod round reported these p50 latencies:
| Query shape | 4 GiB pod | Unlimited pod |
|---|---|---|
match_all |
12.14 ms | 7.32 ms |
label_filter_last25 |
62.26 ms | 10.95 ms |
deep_pagination |
14.25 ms | 9.94 ms |
multi_label_and |
121.32 ms | 14.87 ms |
The two rounds planned different files for every shape in the table, so this is not an isolated measurement of what the limit costs. Read it as what one 4 GiB round did. Raising the pod to 5 GiB, or setting the two text-cache byte limits by hand, is a sizing choice: no round has measured what either does to these latencies.
Query spill and ephemeral storage¶
Once a sort or aggregate reaches the pool bound, DataFusion spills it to node-local scratch on the query pod. The chart provisions that scratch with three capacity ceilings:
query:
spill:
directory: /var/lib/siglake/spill
maxBytes: "8589934592" # 8Gi; keep the quotes
sizeLimit: 10Gi
resources:
requests:
ephemeral-storage: 10Gi
limits:
ephemeral-storage: 12Gi
query.spill.directory and query.spill.maxBytes reach the binary as
SIGLAKE_QUERY_SPILL_DIR and SIGLAKE_QUERY_SPILL_MAX_BYTES, and the chart
mounts an emptyDir sized by query.spill.sizeLimit at that directory.
Keep the three ceilings ordered: DataFusion's byte cap (8Gi) below the
emptyDir size limit (10Gi) below the pod's ephemeral-storage limit (12Gi).
In that order DataFusion refuses the allocation with ResourcesExhausted
first, and the client receives a retryable 503 with Retry-After: 5, as
Queries return 503 describes,
instead of the kubelet evicting the pod for overflowing the emptyDir or the
container's ephemeral-storage limit. The 2Gi between the size limit and the
pod limit is the container's writable layer and logs. The 10Gi request
reserves enough node storage for a full emptyDir at scheduling time.
Keep maxBytes quoted. Unquoted, Helm renders a large integer through
scientific notation before it reaches the byte-count parser, which rejects it
and falls back to DataFusion's default cap. When a workload needs more spill
room, raise all three values together and preserve the order. Spill is
bounded node-local scratch
space:
there is no PVC-backed alternative and no whole-runtime spill-bytes metric, so
watch siglake_query_breaker_trips_total{breaker="pool_exhausted"} and the
pod's ephemeral-storage usage instead. The chart's SiglakeQueryPoolRefusing
alert fires on those breaker trips and sends the operator back to these three
values; its row under Suggested alerts →
Saturation carries the full triage steps. Outside
the chart, the binary's defaults differ; see the spill
variables.
Distributed query¶
On by default with 2 replicas:
Query pods form a StatefulSet behind a headless Service so each shard has
stable DNS. /api/v1/sql coordinates transparently; /api/v1/sql/local
forces single-pod.
Four settings decide how a fan-out behaves, and the chart renders three of them from values you may already have set for other reasons:
| What it controls | Setting | Default |
|---|---|---|
| Peer discovery | The chart renders --query-peer-discovery-srv at the headless Service, so membership is every Ready replica rather than a fixed list. query.replicas is the starting count only. |
On, refreshed every 5 s |
| Shard authentication | query.distributed.coordinatorToken - the credential a coordinator presents on /api/v1/sql/shard. Required whenever query auth is on and the tier can hold more than one pod; the chart fails the install otherwise, because every peer would answer a non-retryable 401. |
Unset |
| Spill scratch space | query.spill.* and the matching ephemeral-storage limits. Spill is node-local per pod, so each replica needs its own room. |
8 GiB cap under a 10 GiB emptyDir |
| Per-coordinator admission | query.admission.budgetBytes and query.admission.waitMs. The budget is per pod and only the coordinator reserves against it: a shard request takes no reservation of its own, so adding replicas raises how many distributed queries run at once. A worker stays bounded by the process memory pool and the mid-flight rows-scanned breaker, so heavy shard scans can reach replicas x 4 at once: see Distributed admission is per-coordinator. |
Binary defaults: 4 GiB budget, 2,000 ms wait |
A pod whose discovery refreshes all fail keeps serving queries without fan-out
and reports it in siglake_query_peer_discovery_refresh_total. See Query peer
discovery for
the alert and the triage.
Batch job store¶
The default renders SIGLAKE_JOBS_POSTGRES_URI from the catalog Secret, so the
query pods keep batch-job rows in the Postgres instance that already holds the
Iceberg catalog. There is no second connection to configure, and the query
server creates its own tables on start. One store is what makes a job_id
usable at more than one pod: the query Service has no session affinity, so a
status, result or cancel request reaches any ready pod.
persistent: false drops the variable and gives each pod its own in-memory
store. Batch state then lives and dies with the pod, and the chart refuses to
render the configurations where that is wrong:
query.replicasabove1with the store off.keda.query.maxReplicasabove1with the store off, whenkeda.enabledandkeda.query.enabledare bothtrue. The ceiling counts even though the StatefulSet starts at one replica, because KEDA can reach it.
Either refusal names the values to set: turn the store back on, or hold the
tier at one pod with query.replicas: 1 and keda.query.enabled: false or
keda.query.maxReplicas: 1. A dormant ceiling is inert, so global KEDA with
keda.query.enabled: false renders the one-pod opt-out.
Keep the opt-out for a single-pod tier that can afford to lose in-flight jobs on restart. For anything else, leave the store on.
Query guardrails¶
query:
admission:
budgetBytes: 0 # 0 = binary default (4 GiB, 4x decompression, 2000 ms wait)
waitMs: 0
breaker:
rowsScannedCeiling: 0 # 0 = default 100M rows
bytesScannedCeilingGb: 0 # 0 = default 100 GB
The 100 M row ceiling suits datasets of roughly 50 GB. At larger scale, a
legitimate windowed aggregate can trip it. A 25% time window over 400 M rows
is about 98 M rows and returns 413 instead of running slowly. Set the ceiling
to a few times your largest expected windowed scan. The wall-clock timeout and
admission budget still bound runaway work.
Attribute promotion¶
Each entry is attr_key:type[:column] with type in
string|int|float|bool; the column name defaults to the attribute name with
. and - replaced by _.
The compactor widens the events table additively at startup. For an existing
warehouse, first run siglake migrate-schema --table events --promote-attr …
with the same set.
Leveled compaction¶
Leveled compaction is enabled by default; no chart value or extraEnv entry is
needed. To select the legacy flat whole-partition pass, set the opt-out through
extraEnv:
See Compaction.
Authentication¶
Ingester¶
ingester:
auth:
existingSecret: siglake-ingester-tokens
secretKey: tokens
oidc:
issuer: https://cognito-idp.us-east-1.amazonaws.com/us-east-1_ABC123
audience: my-client-id
tenantClaim: custom:tenant
allowedTenants: [] # set known tenant ids here; [] accepts any
maxTenants: 100 # admitted tenants per pod; 0 = unbounded
maxLanes: 500 # distinct (tenant, index) pairs; 0 = unbounded
If both issuer and audience are non-empty, OIDC takes precedence over
tokens. The selected Secret field must contain comma-separated bearer tokens.
The chart also accepts inline tokens through ingester.auth.list, but a Secret
keeps credentials out of values files.
Ingest is single-tenant by default. Every request routes to default. An
X-Scope-OrgID naming another tenant is refused with 403 over HTTP or
PermissionDenied over OpenTelemetry Protocol (OTLP)/gRPC. Naming default
has no effect.
Set ingester.oidc.tenantClaim to route by a verified JSON Web Token (JWT)
claim. A present header must agree with the claim. Set
ingester.trustScopeHeader to route by X-Scope-OrgID on the client's word.
Use the claim on shared clusters. If you trust the header, put a gateway in
front of the ingester to set it and strip client-supplied values.
When the tenant set is known, allowedTenants is the strongest bound.
Otherwise, maxTenants
caps distinct admitted tenants per ingester pod. Use maxLanes to cap distinct
(tenant, index) writer lanes. All three bounds are unbounded by default.
Query¶
Choose one of three sources for the caller bearer tokens:
- Set
query.tokens.existingSecretto readquery.tokens.secretKeyfrom a Secret you manage. - Set
query.tokens.listto have the chart create a Secret from an inline list. Use this path only for development because the tokens remain in your values. - Set
externalSecrets.queryTokens.remoteKeyto have External Secrets Operator (ESO) fetch the tokens from your secret backend.
If you use existing Secrets, create both before you install:
siglake-query-tokenswith a non-emptytokensfield containing comma-separated caller bearer tokenssiglake-query-coordinator-tokenwith a non-emptycoordinatorTokenfield containing the peer credential
query:
tokens:
existingSecret: siglake-query-tokens
secretKey: tokens
distributed:
coordinatorToken:
existingSecret: siglake-query-coordinator-token
secretKey: coordinatorToken
If you use ESO, install it and create the referenced SecretStore or
ClusterSecretStore before you apply these values:
externalSecrets:
enabled: true
secretStore: {name: siglake, kind: ClusterSecretStore}
queryTokens:
remoteKey: siglake/query-tokens
The chart renders an ExternalSecret targeting
<release>-siglake-query-tokens. It writes the remote value under
query.tokens.secretKey. Leave query.tokens.existingSecret empty because it
renames the target to a Secret you already manage. Leave query.tokens.list
empty because the chart would write a Secret with the same name as the ESO
target.
The default distributed tier runs two query pods. Configure a separate
query.distributed.coordinatorToken when the ESO token source enables
authentication, or Helm refuses the install.
With no query.oidc.issuer or query.oidc.audience, the non-empty tokens
value enables static bearer-token authentication. Every caller must present a
matching Authorization: Bearer <token> header. If both OIDC settings are
non-empty, OIDC takes precedence over static tokens.
The query API is open only when OIDC is unset and the static token list has no non-blank entries. The server trims and discards blank token entries, then logs a warning if the resulting list is empty.
tenantClaim enables per-request multi-tenancy: the claim value selects an
Iceberg namespace. Setting it also makes the claim mandatory. A verified
token whose claim is missing, blank, non-string, longer than 128 characters or
outside [A-Za-z0-9_-] is refused with 403 before routing, not routed to the
default namespace, and the value is validated rather than repaired
(acme.corp is a refusal, not acmecorp). Leave it empty for single-tenant
mode, where every verified caller reads the default namespace. Refusals raise
the chart's SiglakeQueryTenantsDenied alert.
When query authentication is enabled and query.distributed.enabled: true
with query.replicas > 1, Helm refuses to render with query auth is enabled
and distributed query is fanning out across replicas, but
query.distributed.coordinatorToken is unset. Prefer an existing Secret:
query:
distributed:
coordinatorToken:
existingSecret: siglake-query-coordinator-token
secretKey: coordinatorToken
Alternatively, the chart can create the Secret from an inline value:
The token must be identical across the peer set. The coordinator presents it
to peers on /api/v1/sql/shard, allowing them to trust the forwarded tenant.
TLS¶
The chart's ingress.tls block terminates TLS for configured Ingress routes.
It does not encrypt a listener exposed independently through a Service,
LoadBalancer, or NodePort.
For direct query API exposure without an Ingress, enable the query server's TLS listener:
query:
tls:
enabled: true
existingSecret: siglake-query-tls # kubernetes.io/tls with tls.crt + tls.key
query.tls protects the query API listener. If distributed query is enabled,
the chart also sends query-peer requests over HTTPS.
Neither setting adds TLS to the separate ingest HTTP and gRPC listeners. This includes OTLP/gRPC on port 4317, which uses plaintext h2c when independently exposed. See Send over OTLP/gRPC.
NetworkPolicies restrict access to these listeners, but do not encrypt traffic. Put a TLS proxy, Ingress, or service mesh in front of ingest listeners that need transport encryption.
Storage¶
wal:
size: 200Gi
accessMode: ReadWriteMany
storageClassName: efs-sc
existingClaim: "" # non-empty = chart leaves the PVC alone
The ingester and compactor share the WAL. Use RWX whenever WAL consumers may
run on different nodes, including when scaling ingesters or enabling the query
WAL buffer without co-locating their pods. The chart does not co-locate them by
default. To use RWO, set matching nodeSelector mappings under ingester and
compactor to select one uniquely labeled node, add the mapping under query
when query.walBuffer.enabled is true, and leave antiAffinity.enabled false.
On EKS, RWX normally uses the AWS EFS CSI driver and an EFS-backed StorageClass.
The Terraform module installs the driver as a managed addon;
deploy/aws/up.sh creates the StorageClass after apply. On a non-Terraform
cluster, install the driver and create the StorageClass by hand. See
Terraform on EKS: Post-apply.
Autoscaling¶
Two mechanisms; KEDA is the current one.
query:
replicas: 2 # starting count; KEDA owns spec.replicas after that
keda:
enabled: true
prometheusServerAddress: http://prometheus-server.monitoring.svc.cluster.local:80
ingester:
minReplicas: 1
maxReplicas: 10
requestsPerSecondTarget: "800"
backpressureQueueTarget: "256"
query:
minReplicas: 2 # keep >= 2 so fan-out always has peers
maxReplicas: 12
inFlightTarget: "8"
p95QueueWaitSecondsTarget: "0.5"
The contention trigger uses queue wait rather than request latency because KEDA divides the signal by replica count: the signal must fall as pods are added, but request latency does not.
keda.query.maxReplicas may exceed query.replicas, with or without
query.distributed.enabled. The chart renders --query-peer-discovery-srv
instead of a peer list, so every Ready replica joins the shard membership and
query.replicas is only the starting (and non-KEDA) count; the chart omits
spec.replicas entirely on that tier so an upgrade does not snap a scaled tier
back. A pod KEDA adds becomes eligible one readiness probe plus one discovery
refresh later, and a query already in flight keeps the membership it pinned.
Anti-flap is built in: scale-up is immediate (absorb spikes), scale-down is stabilized over 300 s.
The legacy autoscaling.* block renders CPU-based HPAs for the ingester and
compactor only, off by default. KEDA is the chart's only query
autoscaler. The compactor's customMetric block does not provide multi-pod
backlog scaling in either drain mode. Use siglake-operator when
the compactor must scale on siglake_compactor_sealed_pending.
See Scaling.
Graceful shutdown¶
ingester:
preStopSleepSeconds: 5
terminationGracePeriodSeconds: 60
query:
preStopSleepSeconds: 5
terminationGracePeriodSeconds: 60
preStop pauses so the Service stops routing new work before SIGTERM.
terminationGracePeriodSeconds must exceed preStop plus the longest
operation you want to complete, such as a WAL force-seal on the ingester or an
in-flight query on the query pod.
Other objects¶
| Value | Renders |
|---|---|
ingress.* |
Ingress with optional TLS. |
serviceMonitor.enabled |
Prometheus Operator ServiceMonitor. |
prometheusRule.* |
Prometheus Operator PrometheusRule with 33 alerts. Opt-in; use labels for rule discovery. queryWarmIntervalSecs follows query.warmIntervalSecs when unset or 0; override it only when query.extraEnv gives the pods a different warm cadence. |
podDisruptionBudget.* |
PDBs. Opt-in. |
antiAffinity.* |
Pod anti-affinity. Opt-in. |
networkPolicy.* |
NetworkPolicies. Opt-in. |
externalSecrets.* |
External Secrets Operator integration. |
podSecurityContext, securityContext |
Security contexts. |
logLevel |
RUST_LOG. Default info,siglake=info. |
extraEnv |
Global extra env, applied to every role and the schema-migration hook Job; per-role extraEnv does not reach the hook. |
Sizing rule of thumb¶
From the benchmark arc, budget roughly 50,000 events per second (EPS) per
pod-CPU on the ingester at default WAL segment thresholds. A single ingester at
limits.cpu: 1 sustains about 55,000 EPS on log-line events.
Scaling vertically past limits.cpu: 2 gives sub-linear returns because each
tenant's writer tasks serialize work. Add replicas with an RWX WAL instead.
See Scaling.