Skip to content

Helm

Use the chart in deploy/helm/siglake to deploy Siglake on Kubernetes. This guide covers the settings you must choose, the checks that protect an upgrade and the failure modes that need operator action.

Install

helm install siglake ./deploy/helm/siglake \
  --namespace siglake-system --create-namespace \
  -f my-values.yaml

A minimal my-values.yaml:

tenant:
  namespace: siglake

image:
  repository: ghcr.io/siglake/siglake
  tag: ""                       # empty = Chart.appVersion

serviceAccount:
  create: true
  annotations:
    eks.amazonaws.com/role-arn: arn:aws:iam::123456789012:role/siglake-warehouse-rw

postgres:
  existingSecret: siglake-postgres      # keys: host, port, user, password, database

s3:
  bucket: my-siglake-warehouse
  region: us-east-1
  warehousePrefix: warehouse

wal:
  size: 200Gi
  accessMode: ReadWriteMany
  storageClassName: efs-sc

The pod specification holds Secret references for the PG* variables and an unexpanded SIGLAKE_CATALOG_URI. Kubernetes resolves the references and substitutes the URI in the container environment, with no shell wrapper. If your secret uses different field names (AWS Secrets Manager exports host as endpoint, for instance), remap them under postgres.secretKeys.

Namespaces and tenancy

Each Helm release pins itself to one Iceberg namespace (tenant.namespace, default siglake). Running multiple releases against the same catalog and warehouse gives schema-level isolation between environments.

This namespace is separate from ingest tenant routing. Ingest uses the default tenant unless you configure routing under ingester. See Multi-tenancy.

Per-role shape

Every role accepts the same keys: enabled, replicas, resources, nodeSelector, tolerations, affinity, extraEnv, extraArgs. These per-role extraEnv values reach only that role; global extraEnv also reaches the schema-migration hook Job. Entries are {name, value} or {name, valueFrom}, so a downward-API field works here. The chart renders no OTel block, so this is also where Siglake's own OTLP export is turned on: see Siglake's own telemetry.

Role Workload kind Default replicas
ingester Deployment 1
compactor Deployment 1
query StatefulSet + headless Service 2

Rolling upgrade with Helm: order, consumed-proof upgrade, schema migration

The safe rolling-upgrade procedure with Helm covers order, consumed-proof upgrades and schema migrations. You need both chart directories and the release's values file.

  1. Raise compactor.snapshotExpire.retainLast to 400 with the current chart and image. Wait for it before any new compactor starts.
  2. Set image.tag to the target version and run helm upgrade with the target chart.
  3. Keep schemaMigration.enabled: true. Helm runs the schema migration as a pre-upgrade Job and holds every workload until it succeeds. Wait for the Job and read migrate-schema complete in its log.
  4. Watch the ingester, compactor and query rollouts. The migration is the only barrier between them.
  5. Hold retainLast >= 400 until the last old drain exits, then 1,025 seconds more. Confirm siglake_consumed_proof_bytes per table, then restore normal retention.
  6. If a workload fails after the migration, keep retainLast at 400 and see What a rollback does.

Step 1: raise snapshot retention before the image change

Before you change the image, raise snapshot retention with the chart and image that the release runs now:

compactor:
  snapshotExpire:
    retainLast: 400

Apply that value with the current chart. Do not change image.tag in this command:

helm upgrade siglake <current-chart> --namespace siglake-system -f my-values.yaml --wait --wait-for-jobs --timeout 10m

Wait for this command to finish before you start a new binary. The value change restarts the current compactor with the required retention guard.

Step 2: start the upgrade with the target chart

In my-values.yaml, set image.tag to the target version and leave compactor.snapshotExpire.retainLast at 400. If you set query.image.tag, clear it so the query StatefulSet inherits the target global image. Start the upgrade with the target chart:

helm upgrade siglake <target-chart> --namespace siglake-system -f my-values.yaml --wait --wait-for-jobs --timeout 10m

Step 3: verify the schema-migration Job

Verify the schema-migration Job before you accept the rollout. The pre-upgrade hook uses the target global image and runs migrate-schema --all-tables --all-namespaces. Helm does not apply the workload changes until this Job succeeds. The successful log ends with migrate-schema complete: <N> column(s) added:

kubectl wait --namespace siglake-system --for=condition=complete job/siglake-migrate-schema-<revision> --timeout=10m
kubectl logs --namespace siglake-system job/siglake-migrate-schema-<revision>

If the Job fails, the old workloads remain active. Read its logs, fix the catalog or object-store access, and retry the upgrade. A failed hook does not need a rollback.

Step 4: verify the three workloads

Verify all three workloads after the hook completes:

kubectl rollout status --namespace siglake-system deployment/siglake-ingester --timeout=10m
kubectl rollout status --namespace siglake-system deployment/siglake-compactor --timeout=10m
kubectl rollout status --namespace siglake-system statefulset/siglake-query --timeout=10m

The migration is the only cross-component barrier. After it succeeds, Kubernetes can reconcile the ingester, compactor and query workloads at the same time. The chart does not enforce an order among them. The ingester and query tiers roll while the compactor uses Recreate; see What rolls, and how.

Step 5: hold retention through the consumed-proof overlap

Keep retainLast >= 400 until every old drain or maintenance process has exited, then keep it for another 1,025 seconds. New compactors write the durable siglake.consumed_proof.v1 property and the legacy snapshot summary. They also read both sources. Old compactors continue to use the legacy summary, so the versions can overlap during this window. If you run a maintenance process outside the chart, its last old instance also starts this timer.

For each table receiving writes, confirm that Prometheus reports siglake_consumed_proof_bytes. Confirm that increase(siglake_consumed_proof_cap_refusals_total[30m]) remains 0. The first metric appears after the new compactor updates a table's durable proof. The second catches updates refused at the property's size limit.

After the interval, restore the retention value you chose for normal operation and run helm upgrade again with the target chart. The default is 100. This final upgrade runs the idempotent migration hook again and restarts the compactor with the lower value.

Step 6: roll back a failed workload

If a workload fails after the migration succeeded, keep retainLast at 400 and use What a rollback does. Helm rolls the images back but leaves additive schema columns in place. Restart the 1,025-second timer when the last old drain or maintenance process exits again.

The settings that matter most

Schema migrations

schemaMigration:
  enabled: true
  backoffLimit: 4
  ttlSecondsAfterFinished: 86400
  resources:
    requests:
      memory: 64Mi
      cpu: 50m
    limits:
      memory: 256Mi
      cpu: 500m

On helm upgrade, schemaMigration.enabled renders a pre-upgrade Job using the new image. It runs siglake migrate-schema --all-tables --all-namespaces, including every tenant namespace, and Helm waits for it before rolling the workloads. The migration is additive and idempotent; if the Job fails, the upgrade stops before a newer binary can serve writes against an older table. Fresh installs create tables at the current schema and do not need the hook.

A column the Job adds is queryable as soon as the migration commits, on distributed queries as well as single-pod ones: a coordinator pins each fan-out to the schema id it planned against, so a query replica whose metadata cache predates the migration refreshes onto the pinned schema instead of planning its shard against the older column set (One generation per fan-out). Workers running an image older than the one that added the pinned schema id ignore it, so during the rollout that guarantee holds only once every query pod runs the new image.

The job-migrate-schema.yaml container inherits the chart's s3.* settings, postgres.existingSecret, and global extraEnv, but not ingester.extraEnv, compactor.extraEnv, or query.extraEnv. Put anything needed to reach the object store or catalog during an upgrade in those shared settings; credentials supplied only to a role can leave the hook retrying until Helm reports BackoffLimitExceeded. The check_hook_credentials render-time guard catches credentials present on workloads but missing from the hook.

backoffLimit controls the Job's retry limit, and ttlSecondsAfterFinished controls how long Kubernetes retains the completed Job. resources sets the migration container's requests and limits. Leave enabled on unless another deployment controller guarantees that the same migration completes before rollout.

A newer binary refuses writes that populate a column the table lacks, naming the column and migration remedy instead of silently discarding the value. When prometheusRule.enabled is on, the chart also installs the critical SiglakeSchemaWritesRefused alert for this condition.

What a rollback does

helm rollback siglake <revision> -n siglake-system

The migration Job does not re-run. It is a pre-upgrade hook, and a rollback fires only pre-rollback and post-rollback hooks, neither of which the chart declares.

A rollback rolls back the image, never the schema. The columns the migration added stay, and that is the direction the storage layer supports: the rolled-back binary writes the columns it declares, the storage layer fills the ones it does not with nulls, and rows that already carry values keep them. Re-running migrate-schema on the old image adds and removes nothing, and it does not lower the table's recorded schema version, because the migration stamps the higher of the recorded and declared versions. An image built before that behavior shipped (anything pre-0.1.0) stamps its own lower constant instead. That is a reporting artifact: it changes no column and no write decision, and rolling forward restamps it.

Rolling forward again renders a fresh Job. The Job name carries .Release.Revision, which a rollback increments, so the next helm upgrade never reuses a name. A Job's spec.template is immutable, so reuse would fail with a 422 as soon as the image tag changed. Hook Jobs are not part of the release manifest, so a rollback neither deletes nor recreates the ones already in the namespace; schemaMigration.ttlSecondsAfterFinished (default 86400) reaps them.

A migration that fails needs no rollback. Helm blocks on the pre-upgrade hook, so the release fails before any workload is applied and the previous revision is still the live one. Read the Job's logs before you delete it:

kubectl logs -n siglake-system -l app.kubernetes.io/component=migrate-schema

The next attempt's Job has a different name either way.

Under the operator, reverting spec.image runs one more migration Job on the old binary instead. See Reverting the image.

Two limits bound all of this. It covers additive column changes only: there is no path across the pre-0.1.0 nanosecond timestamp contract, which migrate-schema refuses rather than migrating. And no rollback has been qualified against an actual older image. The regression runs one binary against a table widened past what it declares, which is evidence about the mechanism, not about a released image; the rest of this section follows from the chart templates.

What rolls, and how

For a 0.1.0 to 0.2.0 upgrade, upgrade the whole query tier together and plan for query downtime while it changes versions. A 0.2.0 coordinator rejects a 0.1.0 worker's response when a zero-scan fast path omits the scan-attribution header, so a mixed query tier can fail distributed queries; a single query pod executes in process and is unaffected.

The ingester Deployment uses RollingUpdate with maxUnavailable: 0. Upgrades do not reduce ingest capacity. The query StatefulSet also uses RollingUpdate and replaces one ordinal at a time. If a peer being replaced does not answer an in-flight query, the coordinator runs that shard itself and increments siglake_query_coordinator_failover_total.

The chart uses Recreate for the compactor regardless of replica count, so drain and maintenance pause while the pod restarts. Sealed segments wait on the WAL PVC; acknowledgement happens in the ingester, so the restart does not lose acknowledged data. A compactor cannot safely surge against the shared local WAL because two pods can claim the same segment; multi-pod compaction instead coordinates claims through the catalog.

A restart that interrupts a commit needs no operator in the ordinary case. The in-flight commit is all or nothing, and the restarted drain decides what to do with the segment from the table's consumed-segment record. See The compactor dies mid-commit for the disposition rules, the two metrics that count them, and the one case that does ask for a decision.

Freshness

query:
  walBuffer:
    enabled: true       # requires wal.accessMode: ReadWriteMany

Without this, query pods do not mount the WAL and you get commit-cycle visibility (tens of seconds), not the seconds-scale freshness Siglake is designed for. The query pods read the same WAL as the ingester and compactor, so the claim must be RWX when those pods can run on different nodes. The chart defaults wal.accessMode to ReadWriteMany and wal.storageClassName to the cluster default; if that StorageClass only supports RWO (as EBS-backed classes commonly do), the RWX claim remains pending.

query.hotCaches.enabled (the last_values() / distinct_values() UDTFs) tails the same mount and needs walBuffer too.

Disaster recovery

wal:
  mirror:
    enabled: true
    prefix: wal-mirror
    activeIntervalSecs: 0

The prefix is relative to the warehouse URL, so these defaults upload each sealed segment to s3://<s3.bucket>/<s3.warehousePrefix>/wal-mirror/, inside the warehouse prefix rather than beside it. Acknowledgements do not wait for the upload: the default wait_for response means the event is fsynced locally, and it has no copy in the warehouse bucket until its Iceberg commit or a mirror upload completes, whichever comes first. A committed row keeps its warehouse copy with the mirror off, so mirroring is what protects the rows still waiting for a commit. Set activeIntervalSecs above 0 to upload every open segment on that cadence, one object per open writer under _active/<tenant>[/<index>]/. That is one PUT per (tenant, managed index, ingester.backpressure.shards) combination holding new rows, per ingester, per tick. Successful uploads set that recovery target only when you recover a filesystem drain. A delayed or failed upload extends the loss window past the configured interval. If you recover a catalog-claim drain, reconciliation recovers sealed objects and gets no active-snapshot target. See mirror-failure monitoring.

The default single-replica compactor drains the local WAL and does not reclaim mirror objects. Give <s3.warehousePrefix>/wal-mirror/ an object-store lifecycle expiry unless you enable compactor.catalogClaim.enabled. Catalog-claim retention reclaims sealed objects only. From Siglake 0.2.0 the ingester deletes each sealed segment's own _active/ partial after it confirms the sealed upload, so the expiry covers the rest: partials written before that cleanup shipped, and partials whose deletes failed. Set wal.mirror.enabled: false to opt out. The chart then renders an empty SIGLAKE_WAL_MIRROR_PREFIX; omitting that variable leaves the binary's warehouse-backed mirror running.

Value Default Purpose
compactor.committedRetentionSecs 86400 Purge successfully drained mirror objects and catalog rows after this many seconds; 0 disables purging.
prometheusRule.mirrorRotationStallSecs 21600 Alert when bounded reconciliation pages run for this long without a full rotation completing; 0 disables the alert.

Non-zero committed retention is floored at 901 seconds so it outlives the default 600-second local-WAL settle delay plus one 300-second sweep cadence. Raise it beyond the sum if either window is increased. Keep it at 0 when the local-WAL sweep is disabled but mirror catch-up remains enabled. See Mirror reconciliation and retention for the scaling consequence of retaining the full prefix.

Multi-pod compaction

compactor:
  replicas: 2
  catalogClaim:
    enabled: true
    batch: 16
  extraArgs: ["--role", "drain"]
wal:
  mirror:
    enabled: true    # required: the compactor reads segments from the bucket

Catalog claim switches the compactor from listing the local sealed/ directory to the wal_segments SQL claim table. It is required above one pod, not recommended: the chart fails the install when compactor.replicas is above 1, or when autoscaling.compactor.maxReplicas is above 1 with autoscaling.compactor.enabled: true, while compactor.catalogClaim.enabled is false. Replicas that share no claim table each run the whole drain and maintenance loop against the same tables, lose Iceberg commit races to each other, and leave the layout unconverged.

The install fails the other way round as well: the claim with wal.mirror.enabled: false. Only an ingester that mirrors its segments writes wal_segments rows, so a claim drain without the mirror claims nothing and compaction stops.

The chart also refuses the compactor's custom backlog metric when all three of these values are true:

autoscaling:
  compactor:
    enabled: true
    customMetric:
      enabled: true
compactor:
  catalogClaim:
    enabled: true

The error starts with autoscaling.compactor.customMetric.enabled is true with compactor.catalogClaim.enabled. In claim mode, siglake_compactor_sealed_pending is the whole shared queue of sealed, unclaimed segments. It is not one pod's share. The chart renders a Pods metric with an averageValue target, so the requested replica count grows with the current replica count instead of tracking the backlog. Disable customMetric to keep the claim-mode CPU HPA, or use siglake-operator for backlog-driven scaling. In filesystem mode, the existing guard holds the compactor at one pod, so the custom metric cannot scale that mode either.

The chart has no first-class compactor.role value yet. Set the binary role with compactor.extraArgs, as above, or use compactor.extraEnv when environment variables are easier to manage:

compactor:
  extraEnv:
    - name: SIGLAKE_COMPACTOR_ROLE
      value: drain

The chart renders one compactor Deployment, so this recipe makes all its replicas drains. Run the single maintenance process as a separately managed Deployment with --role maintenance; see Scaling. Leave the role unset for the default single-process combined behavior.

Switch an existing filesystem drain to the claim: inventory held orphans first

Before enabling compactor.catalogClaim.enabled, inventory and keep each <wal>/**/orphans/ file on the persistent volume claim (PVC). Use an ingester pod; it mounts the PVC in both drain modes:

kubectl exec deployment/<release>-ingester -n <namespace> -- \
  find /var/lib/siglake/wal -type f -path '*/orphans/*' -print

A held orphan has UNKNOWN commit status. Use retention and ingest history to decide whether its rows committed. Deleting it can lose rows; requeueing it can duplicate them.

siglake_compactor_orphans_held{tenant} and SiglakeCompactorOrphansHeld come only from the filesystem drain's per-cycle orphans/ sweep. With a catalog, the cycle switches to the claim path and visits no write-ahead log (WAL) directory. A missing series does not prove the PVC is clear: its files remain while this monitoring stops.

The chart gives the compactor an emptyDir instead of its WAL claim, so the pod cannot inspect the PVC. The operator keeps the claim mounted in both modes, but refuses the conversion with DrainModeHandoverRequired.

If evidence proves an orphan uncommitted after the switch, moving it to sealed/ does not ingest it. The claim drain never reads that directory. The ingester's mirror catch-up sweep must copy the segment, then the claim drain's mirror sync can add it to the table.

Snapshot retention

compactor:
  snapshotExpire:
    intervalSecs: 60
    retainLast: 100

Expiry is on by default. retainLast counts snapshots rather than time. Only a newer commit displaces an older snapshot from the retained window, so a busy table can cycle through 100 snapshots quickly while an idle one keeps its oldest retained snapshot indefinitely. intervalSecs is how often the compactor trims, not a wall-clock erasure deadline. Every drain, compaction, maintenance and delete commit produces a snapshot, so the window's real duration is a function of that table's commit rate.

Raising it has two costs:

  • Every commit reads and rewrites metadata.json. Each retained snapshot adds an entry, so a larger window increases work on every commit.
  • Expiry changes metadata only. gc-orphans can delete a data file only after no retained snapshot references it. A wider window therefore holds superseded files, including files rewritten by a delete, in the warehouse for longer. See The full sequence to make a delete permanent.

Two needs can pull the value above the packaged default of 100. Use the higher value while both apply:

  • Iceberg time travel and long-running external queries reach only as far back as the retained window (Snapshot expiry limits time travel). A mixed-version rollout needs >= 400 temporarily. See Consumed-proof rolling upgrades below, which is a bounded guard, not a steady-state recommendation.
  • A coordinator can serve table metadata that is stale by the cache TTL (5 seconds by default, up to 60 seconds while a background refresh runs). If the catalog expires a snapshot a coordinator is still pinning fan-outs to, workers refuse their shards with 503 and reason: "shard_pin_unresolved". The margin you need is the number of commits that land in that staleness window, so estimate commits per minute for the busiest table and keep retainLast comfortably above it. The other half of that remedy is lowering SIGLAKE_ICEBERG_METADATA_CACHE_TTL_SECS. See A 503 with reason: shard_pin_unresolved. A pin also names the schema id, and retention does not widen that half. A schema miss right after migrate-schema clears when the lagging replica refreshes its metadata.

The packaged 100 is a fixed value, not one derived from the workload the way compactor.binConcurrency is. Pick it from whichever consumer above needs the longest window, and revisit it when the commit rate changes. intervalSecs: 0 disables trimming entirely, which adds a growing cost to every commit.

Changing retainLast in either direction updates metadata without rewriting data files. Raising it preserves evidence from snapshots that have not expired, but cannot restore expired snapshots. Keep an ambiguous filesystem orphan held until available consumption proof or an operator comparison establishes its disposition. You can lower retainLast after the incident, subject to the rollout guard below.

Consumed-proof rolling upgrades

Rolling ingesters and drains through a version change safely takes one retention setting, applied before the first new binary starts. Old and new writers record a claim's consumption differently, so the reader needs both sources retained while the two versions overlap.

Versions that write siglake.consumed_proof.v1 can reclaim committed claims after their source snapshots expire. Older writers update only the legacy snapshot summary. Before starting the first new binary, set:

compactor:
  snapshotExpire:
    retainLast: 400

Keep compactor.snapshotExpire.retainLast >= 400 throughout the rollout and for at least 1,025 seconds after the last old drain or maintenance writer exits. The new reader consults both sources during that interval. After the wait, retainLast may return to the value Snapshot retention arrives at for the deployment (the packaged default is 100); no data-file rewrite is required.

Compactor sizing

The packaged compactor merges one bin at a time:

compactor:
  resources:
    limits:
      cpu: 2
      memory: 1Gi
  binConcurrency: 1

When SIGLAKE_COMPACTOR_BIN_CONCURRENCY is not set, Siglake derives safe bin concurrency from the cgroup memory limit and available cores. The chart sets it explicitly through compactor.binConcurrency; raise that value, memory, and CPU together when backfill makes compaction lag. Budget a 1 GiB process reserve plus roughly 4 GiB per concurrent bin and at least one CPU per bin. Raising concurrency without proportional memory leads to OOM kills rather than more throughput.

Query memory budget

Treat query.resources.limits.memory as the query server's sizing budget, not just its kill ceiling. Use Size query pods and understand distributed fan-out to set the pod count after you choose this budget. The chart values divide the cgroup limit this way with the chart defaults:

  • 25% for the byte-range object cache, clamped from 64 MiB to 16 GiB;
  • 12.5% for metadata caches, clamped from 64 MiB to 2 GiB;
  • no space for the experimental source-file batch cache unless both query.scan.fileCacheMaxBytes and query.scan.fileCacheMaxEntries are set to positive values;
  • 1/16 of the limit for parsed text indexes, capped at 1 GiB, and 1/64 for their Puffin blobs, capped at 256 MiB, taken only from what remains once the pool can still reserve one compacted file's 1.25 GiB decode working set;
  • 50% of the memory left after the object, metadata, text-index and enabled source-file caches for DataFusion's shared memory pool;
  • the rest for the process, in-flight Arrow batches, decode buffers and allocator slack.

For a default 4 GiB pod, that is a 1 GiB object cache, 512 MiB of metadata caches, no text-index caches and a 1.25 GiB DataFusion memory pool. About 1.25 GiB remains for the process and transient allocations. If you enable the source-file cache, its byte limit is also subtracted before applying the pool fraction. That limit covers cached entries and the decoded batches in-flight populations hold, so it is the whole memory the cache can take: see What the decoded-file cache byte limit covers. A 12.5% share of the pod limit is a reasonable starting budget, not a derived default.

The text-index caches get nothing at 4 GiB because the pool takes half of what the other caches leave, so holding the 1.25 GiB decode reservation costs 2.5 GiB of the limit and the floor pod has no remainder to divide. Raise the limit to 5 GiB and they derive 400 MiB; at 16 GiB they reach both caps. Both budgets are published on siglake_cache_budget_bytes{kind="text_index"}, and tuning the text-index caches covers the two environment variables that override them.

DataFusion sorts and aggregates spill when they reach the pool bound. If it cannot stay within the memory and spill bounds, Siglake reports ResourcesExhausted as a retryable 503. The pool reduces OOM risk, but it does not account for every process allocation.

Raise the memory limit for scan-heavy or high-concurrency workloads. Use SIGLAKE_QUERY_MEMORY_FRACTION to change the pool's default 0.5 share, and watch siglake_query_memory_pool_bytes.

siglake_query_scan_decode_reservation_total{outcome="unreserved"} counts scans that fell to one file at a time and then proceeded without a pool reservation, so their decode memory goes unaccounted. Read it as a coverage signal, not a latency one, and do not alert on it as a latency proxy: wider fan-out asks the pool for more, so in the 2026-09-06 decode sweep the fastest arm also had the highest unreserved rate.

What the packaged 4 GiB query pod measured on the 50 GB suite

The 2026-09-15 50 GB benchmark ran the query server under the packaged 4 GiB limit, and the pod held it: the cgroup kept the ceiling through both suites, Docker reported no out-of-memory (OOM) kill and no restart, and the budgets came out as the arithmetic above predicts: a 1 GiB object cache, 512 MiB of metadata caches, no text-index caches and a 1.25 GiB pool. A later reread of that round against the recorded unlimited-pod round reported these p50 latencies:

Query shape 4 GiB pod Unlimited pod
match_all 12.14 ms 7.32 ms
label_filter_last25 62.26 ms 10.95 ms
deep_pagination 14.25 ms 9.94 ms
multi_label_and 121.32 ms 14.87 ms

The two rounds planned different files for every shape in the table, so this is not an isolated measurement of what the limit costs. Read it as what one 4 GiB round did. Raising the pod to 5 GiB, or setting the two text-cache byte limits by hand, is a sizing choice: no round has measured what either does to these latencies.

Query spill and ephemeral storage

Once a sort or aggregate reaches the pool bound, DataFusion spills it to node-local scratch on the query pod. The chart provisions that scratch with three capacity ceilings:

query:
  spill:
    directory: /var/lib/siglake/spill
    maxBytes: "8589934592"   # 8Gi; keep the quotes
    sizeLimit: 10Gi
  resources:
    requests:
      ephemeral-storage: 10Gi
    limits:
      ephemeral-storage: 12Gi

query.spill.directory and query.spill.maxBytes reach the binary as SIGLAKE_QUERY_SPILL_DIR and SIGLAKE_QUERY_SPILL_MAX_BYTES, and the chart mounts an emptyDir sized by query.spill.sizeLimit at that directory. Keep the three ceilings ordered: DataFusion's byte cap (8Gi) below the emptyDir size limit (10Gi) below the pod's ephemeral-storage limit (12Gi). In that order DataFusion refuses the allocation with ResourcesExhausted first, and the client receives a retryable 503 with Retry-After: 5, as Queries return 503 describes, instead of the kubelet evicting the pod for overflowing the emptyDir or the container's ephemeral-storage limit. The 2Gi between the size limit and the pod limit is the container's writable layer and logs. The 10Gi request reserves enough node storage for a full emptyDir at scheduling time.

Keep maxBytes quoted. Unquoted, Helm renders a large integer through scientific notation before it reaches the byte-count parser, which rejects it and falls back to DataFusion's default cap. When a workload needs more spill room, raise all three values together and preserve the order. Spill is bounded node-local scratch space: there is no PVC-backed alternative and no whole-runtime spill-bytes metric, so watch siglake_query_breaker_trips_total{breaker="pool_exhausted"} and the pod's ephemeral-storage usage instead. The chart's SiglakeQueryPoolRefusing alert fires on those breaker trips and sends the operator back to these three values; its row under Suggested alerts → Saturation carries the full triage steps. Outside the chart, the binary's defaults differ; see the spill variables.

Distributed query

On by default with 2 replicas:

query:
  replicas: 2
  distributed:
    enabled: true

Query pods form a StatefulSet behind a headless Service so each shard has stable DNS. /api/v1/sql coordinates transparently; /api/v1/sql/local forces single-pod.

Four settings decide how a fan-out behaves, and the chart renders three of them from values you may already have set for other reasons:

What it controls Setting Default
Peer discovery The chart renders --query-peer-discovery-srv at the headless Service, so membership is every Ready replica rather than a fixed list. query.replicas is the starting count only. On, refreshed every 5 s
Shard authentication query.distributed.coordinatorToken - the credential a coordinator presents on /api/v1/sql/shard. Required whenever query auth is on and the tier can hold more than one pod; the chart fails the install otherwise, because every peer would answer a non-retryable 401. Unset
Spill scratch space query.spill.* and the matching ephemeral-storage limits. Spill is node-local per pod, so each replica needs its own room. 8 GiB cap under a 10 GiB emptyDir
Per-coordinator admission query.admission.budgetBytes and query.admission.waitMs. The budget is per pod and only the coordinator reserves against it: a shard request takes no reservation of its own, so adding replicas raises how many distributed queries run at once. A worker stays bounded by the process memory pool and the mid-flight rows-scanned breaker, so heavy shard scans can reach replicas x 4 at once: see Distributed admission is per-coordinator. Binary defaults: 4 GiB budget, 2,000 ms wait

A pod whose discovery refreshes all fail keeps serving queries without fan-out and reports it in siglake_query_peer_discovery_refresh_total. See Query peer discovery for the alert and the triage.

Batch job store

query:
  jobs:
    persistent: true     # default: one store for the whole tier

The default renders SIGLAKE_JOBS_POSTGRES_URI from the catalog Secret, so the query pods keep batch-job rows in the Postgres instance that already holds the Iceberg catalog. There is no second connection to configure, and the query server creates its own tables on start. One store is what makes a job_id usable at more than one pod: the query Service has no session affinity, so a status, result or cancel request reaches any ready pod.

persistent: false drops the variable and gives each pod its own in-memory store. Batch state then lives and dies with the pod, and the chart refuses to render the configurations where that is wrong:

  • query.replicas above 1 with the store off.
  • keda.query.maxReplicas above 1 with the store off, when keda.enabled and keda.query.enabled are both true. The ceiling counts even though the StatefulSet starts at one replica, because KEDA can reach it.

Either refusal names the values to set: turn the store back on, or hold the tier at one pod with query.replicas: 1 and keda.query.enabled: false or keda.query.maxReplicas: 1. A dormant ceiling is inert, so global KEDA with keda.query.enabled: false renders the one-pod opt-out.

Keep the opt-out for a single-pod tier that can afford to lose in-flight jobs on restart. For anything else, leave the store on.

Query guardrails

query:
  admission:
    budgetBytes: 0        # 0 = binary default (4 GiB, 4x decompression, 2000 ms wait)
    waitMs: 0
  breaker:
    rowsScannedCeiling: 0        # 0 = default 100M rows
    bytesScannedCeilingGb: 0     # 0 = default 100 GB

The 100 M row ceiling suits datasets of roughly 50 GB. At larger scale, a legitimate windowed aggregate can trip it. A 25% time window over 400 M rows is about 98 M rows and returns 413 instead of running slowly. Set the ceiling to a few times your largest expected windowed scan. The wall-clock timeout and admission budget still bound runaway work.

Attribute promotion

compactor:
  promoteAttrs:
    - http.status_code:int
    - k8s.namespace:string

Each entry is attr_key:type[:column] with type in string|int|float|bool; the column name defaults to the attribute name with . and - replaced by _.

The compactor widens the events table additively at startup. For an existing warehouse, first run siglake migrate-schema --table events --promote-attr … with the same set.

Leveled compaction

Leveled compaction is enabled by default; no chart value or extraEnv entry is needed. To select the legacy flat whole-partition pass, set the opt-out through extraEnv:

compactor:
  extraEnv:
    - name: SIGLAKE_COMPACTOR_LEVELED
      value: "0"

See Compaction.

Short-aggregate repair

Every 15 minutes the maintenance compactor censuses each maintained table for a group-count aggregate short of total-records and reports what it finds on siglake_group_count_short_aggregates_total{iceberg_namespace,table,outcome} and SiglakeGroupCountAggregateShort. Rebuilding is separate, and off by default:

compactor:
  shortAggregateRepair: true

The value renders SIGLAKE_AGG_SHORT_REPAIR=1, and the census then rebuilds one table per pass. Leave compactor.shortAggregateRepair at false where tables are wide or large: the repair is one Tier-2 query per maintained column and a table big enough for that to exceed the compactor's 600-second watchdog has the repair cut. Retries wait 15 minutes, one hour and four hours. The fourth unsuccessful scan suppresses automatic repair for that table incarnation. Rebuild those tables with siglake rebuild-group-counts --namespace <ns> --table <table> instead, with the namespace the alert reports. Answers stay exact either way. See Rebuild a short group-count aggregate with durable backoff.

Inline-coverage census

The same maintenance pass reads each maintained table's inline aggregate object every 15 minutes and records whether its coverage edge still reaches the current snapshot, on siglake_inline_coverage_unproven{iceberg_namespace,table} and SiglakeInlineCoverageUnproven. The census is on by default and has nothing to opt into, because it only reads: one object and the table metadata per pass. No chart value renders its cadence. To change it, or to switch the census off:

compactor:
  extraEnv:
    - name: SIGLAKE_INLINE_COVERAGE_SCAN_INTERVAL_SECS
      value: "off"

Switching it off leaves the repair undetected: nothing else names the table that needs siglake rebuild-time-aggregates. See Data loss or durable inconsistency.

Authentication

Ingester

ingester:
  auth:
    existingSecret: siglake-ingester-tokens
    secretKey: tokens
  oidc:
    issuer: https://cognito-idp.us-east-1.amazonaws.com/us-east-1_ABC123
    audience: my-client-id
    tenantClaim: custom:tenant
  allowedTenants: []       # set known tenant ids here; [] accepts any
  maxTenants: 100          # admitted tenants per pod; 0 = unbounded
  maxLanes: 500            # distinct (tenant, index) pairs; 0 = unbounded

If both issuer and audience are non-empty, OIDC takes precedence over tokens. The selected Secret field must contain comma-separated bearer tokens. The chart also accepts inline tokens through ingester.auth.list, but a Secret keeps credentials out of values files.

Ingest is single-tenant by default. Every request routes to default. An X-Scope-OrgID naming another tenant is refused with 403 over HTTP or PermissionDenied over OpenTelemetry Protocol (OTLP)/gRPC. Naming default has no effect.

Set ingester.oidc.tenantClaim to route by a verified JSON Web Token (JWT) claim. A present header must agree with the claim. Set ingester.trustScopeHeader to route by X-Scope-OrgID on the client's word. Use the claim on shared clusters. If you trust the header, put a gateway in front of the ingester to set it and strip client-supplied values.

When the tenant set is known, allowedTenants is the strongest bound. Otherwise, maxTenants caps distinct admitted tenants per ingester pod. Use maxLanes to cap distinct (tenant, index) writer lanes. All three bounds are unbounded by default.

Query

Choose one of three sources for the caller bearer tokens:

  • Set query.tokens.existingSecret to read query.tokens.secretKey from a Secret you manage.
  • Set query.tokens.list to have the chart create a Secret from an inline list. Use this path only for development because the tokens remain in your values.
  • Set externalSecrets.queryTokens.remoteKey to have External Secrets Operator (ESO) fetch the tokens from your secret backend.

If you use existing Secrets, create both before you install:

  • siglake-query-tokens with a non-empty tokens field containing comma-separated caller bearer tokens
  • siglake-query-coordinator-token with a non-empty coordinatorToken field containing the peer credential
query:
  tokens:
    existingSecret: siglake-query-tokens
    secretKey: tokens
  distributed:
    coordinatorToken:
      existingSecret: siglake-query-coordinator-token
      secretKey: coordinatorToken

If you use ESO, install it and create the referenced SecretStore or ClusterSecretStore before you apply these values:

externalSecrets:
  enabled: true
  secretStore: {name: siglake, kind: ClusterSecretStore}
  queryTokens:
    remoteKey: siglake/query-tokens

The chart renders an ExternalSecret targeting <release>-siglake-query-tokens. It writes the remote value under query.tokens.secretKey. Leave query.tokens.existingSecret empty because it renames the target to a Secret you already manage. Leave query.tokens.list empty because the chart would write a Secret with the same name as the ESO target.

The default distributed tier runs two query pods. Configure a separate query.distributed.coordinatorToken when the ESO token source enables authentication, or Helm refuses the install.

With no query.oidc.issuer or query.oidc.audience, the non-empty tokens value enables static bearer-token authentication. Every caller must present a matching Authorization: Bearer <token> header. If both OIDC settings are non-empty, OIDC takes precedence over static tokens.

The query API is open only when OIDC is unset and the static token list has no non-blank entries. The server trims and discards blank token entries, then logs a warning if the resulting list is empty.

To restrict reads to a known tenant set, configure claim routing and the query allow-list together:

query:
  oidc:
    issuer: https://cognito-idp.us-east-1.amazonaws.com/us-east-1_ABC123
    audience: my-client-id
    tenantClaim: custom:tenant
  allowedTenants:
    - acme
    - widgets

query.allowedTenants: [] is the unrestricted default. A non-empty list does not inherit ingester.allowedTenants and requires query.oidc.tenantClaim. Helm refuses to render a non-empty list without that claim setting. The query-server binary also refuses the equivalent flag or environment configuration at startup.

tenantClaim enables per-request multi-tenancy: the claim value selects an Iceberg namespace. Setting it also makes the claim mandatory. A verified token whose claim is missing, blank, non-string, longer than 128 characters or outside [A-Za-z0-9_-] is refused with 403 before routing, not routed to the default namespace, and the value is validated rather than repaired (acme.corp is a refusal, not acmecorp). Leave it empty for single-tenant mode, where every verified caller reads the default namespace. An unlisted usable claim is also refused with 403 before namespace creation. Refusals raise the chart's SiglakeQueryTenantsDenied alert.

When query authentication is enabled and query.distributed.enabled: true with query.replicas > 1, Helm refuses to render with query auth is enabled and distributed query is fanning out across replicas, but query.distributed.coordinatorToken is unset. Prefer an existing Secret:

query:
  distributed:
    coordinatorToken:
      existingSecret: siglake-query-coordinator-token
      secretKey: coordinatorToken

Alternatively, the chart can create the Secret from an inline value:

query:
  distributed:
    coordinatorToken:
      value: replace-with-a-shared-secret

The token must be identical across the peer set. The coordinator presents it to peers on /api/v1/sql/shard, allowing them to trust the forwarded tenant.

TLS

The chart's ingress.tls block terminates TLS for configured Ingress routes. It does not encrypt a listener exposed independently through a Service, LoadBalancer, or NodePort.

For direct query API exposure without an Ingress, enable the query server's TLS listener:

query:
  tls:
    enabled: true
    existingSecret: siglake-query-tls   # kubernetes.io/tls with tls.crt + tls.key

query.tls protects the query API listener. If distributed query is enabled, the chart also sends query-peer requests over HTTPS.

Neither setting adds TLS to the separate ingest HTTP and gRPC listeners. This includes OTLP/gRPC on port 4317, which uses plaintext h2c when independently exposed. See Send over OTLP/gRPC.

NetworkPolicies restrict access to these listeners, but do not encrypt traffic. Put a TLS proxy, Ingress, or service mesh in front of ingest listeners that need transport encryption.

Storage

wal:
  size: 200Gi
  accessMode: ReadWriteMany
  storageClassName: efs-sc
  existingClaim: ""          # non-empty = chart leaves the PVC alone

The ingester and compactor share the WAL. Use RWX whenever WAL consumers may run on different nodes, including when scaling ingesters or enabling the query WAL buffer without co-locating their pods. The chart does not co-locate them by default. To use RWO, set matching nodeSelector mappings under ingester and compactor to select one uniquely labeled node, add the mapping under query when query.walBuffer.enabled is true, and leave antiAffinity.enabled false. On EKS, RWX normally uses the AWS EFS CSI driver and an EFS-backed StorageClass. The Terraform module installs the driver as a managed addon; deploy/aws/up.sh creates the StorageClass after apply. On a non-Terraform cluster, install the driver and create the StorageClass by hand. See Terraform on EKS: Post-apply.

Autoscaling

Two mechanisms; KEDA is the current one.

query:
  replicas: 2              # starting count; KEDA owns spec.replicas after that
keda:
  enabled: true
  prometheusServerAddress: http://prometheus-server.monitoring.svc.cluster.local:80
  ingester:
    minReplicas: 1
    maxReplicas: 10
    requestsPerSecondTarget: "800"
    backpressureQueueTarget: "256"
  query:
    minReplicas: 2           # keep >= 2 so fan-out always has peers
    maxReplicas: 12
    inFlightTarget: "8"
    p95QueueWaitSecondsTarget: "0.5"

The contention trigger uses queue wait rather than request latency because KEDA divides the signal by replica count: the signal must fall as pods are added, but request latency does not.

keda.query.maxReplicas may exceed query.replicas, with or without query.distributed.enabled. The chart renders --query-peer-discovery-srv instead of a peer list, so every Ready replica joins the shard membership and query.replicas is only the starting (and non-KEDA) count; the chart omits spec.replicas entirely on that tier so an upgrade does not snap a scaled tier back. A pod KEDA adds becomes eligible one readiness probe plus one discovery refresh later, and a query already in flight keeps the membership it pinned.

Anti-flap is built in: scale-up is immediate (absorb spikes), scale-down is stabilized over 300 s.

The legacy autoscaling.* block renders CPU-based HPAs for the ingester and compactor only, off by default. KEDA is the chart's only query autoscaler. The compactor's customMetric block does not provide multi-pod backlog scaling in either drain mode. Use siglake-operator when the compactor must scale on siglake_compactor_sealed_pending.

See Scaling.

Graceful shutdown

ingester:
  preStopSleepSeconds: 5
  terminationGracePeriodSeconds: 60
query:
  preStopSleepSeconds: 5
  terminationGracePeriodSeconds: 60

preStop pauses so the Service stops routing new work before SIGTERM. terminationGracePeriodSeconds must exceed preStop plus the longest operation you want to complete, such as a WAL force-seal on the ingester or an in-flight query on the query pod.

Other objects

Value Renders
ingress.* Ingress with optional TLS.
serviceMonitor.enabled Prometheus Operator ServiceMonitor.
prometheusRule.* Prometheus Operator PrometheusRule with 39 alerts. Opt-in; use labels for rule discovery. queryWarmIntervalSecs follows query.warmIntervalSecs when unset or 0; override it only when query.extraEnv gives the pods a different warm cadence.
podDisruptionBudget.* PDBs. Opt-in.
antiAffinity.* Pod anti-affinity. Opt-in.
networkPolicy.* NetworkPolicies. Opt-in.
externalSecrets.* External Secrets Operator integration.
podSecurityContext, securityContext Security contexts.
logLevel RUST_LOG. Default info,siglake=info.
extraEnv Global extra env, applied to every role and the schema-migration hook Job; per-role extraEnv does not reach the hook.

Sizing rule of thumb

From the benchmark arc, budget roughly 50,000 events per second (EPS) per pod-CPU on the ingester at default WAL segment thresholds. A single ingester at limits.cpu: 1 sustains about 55,000 EPS on log-line events.

Scaling vertically past limits.cpu: 2 gives sub-linear returns because each tenant's writer tasks serialize work. Add replicas with an RWX WAL instead. See Scaling.