Operations¶
Use this page to pick a deployment path, gather what the chart expects you to bring, and change the defaults that matter before you take traffic.
Deployment paths¶
| Path | Use when | Page |
|---|---|---|
| Docker Compose | Local development, evaluation, single-host demos. | Quickstart |
| Helm | Kubernetes. The most complete and current path. | Helm |
| Operator | You want a SiglakeCluster custom resource. Incomplete: see below. |
Operator |
| Terraform and EKS | Greenfield AWS. Provisions EKS, EFS, RDS and S3, and emits matching Helm values. | Terraform on EKS |
The operator renders ingester and compactor as Deployments and query as a StatefulSet, which matches the chart's workload kinds. Everything else the chart renders, the operator does not: Ingress, PodDisruptionBudgets, NetworkPolicies, ServiceMonitors, autoscaling objects, query-tier authentication and TLS. Choose Helm unless you specifically want the custom resource.
Object storage, Iceberg catalog, WAL volume, and Kubernetes prerequisites¶
Before you install Siglake, provide Kubernetes 1.27 or later and the infrastructure below. The chart references these resources rather than provisioning them.
- A Postgres instance reachable from the cluster, for the Iceberg catalog. RDS in the reference deployment.
- A Kubernetes Secret holding its connection details, with the keys
host,port,user,passwordanddatabase. - An S3 bucket for the Iceberg warehouse. This is the object storage every Parquet file lands in.
- An IAM role with read and write on that bucket, trusted by the cluster's OIDC provider and annotated on the ServiceAccount. On EKS this is IRSA, IAM Roles for Service Accounts.
- A StorageClass for the shared write-ahead log (WAL) volume, ReadWriteMany in most deployments.
The Terraform module in deploy/terraform/aws/ provisions items 1, 3 and 4,
and emits Helm values matching the chart's schema.
Choosing the WAL StorageClass¶
Use ReadWriteMany (RWX) whenever WAL consumers may run on different nodes,
which includes scaling ingesters and enabling the query freshness buffer
without co-locating their pods. ReadWriteOnce (RWO) requires every WAL consumer
on one node, and the chart does not co-locate them for you. If you want RWO,
set matching nodeSelector mappings under ingester and compactor to select
one uniquely labelled node, add the same mapping under query when the
freshness buffer is on, and leave antiAffinity.enabled false.
On EKS, RWX means the AWS EFS CSI driver and an EFS-backed StorageClass. The
Terraform module installs the driver as a managed addon, and deploy/aws/up.sh
creates the StorageClass after apply. On a cluster you did not build with
Terraform, install the driver and create the StorageClass yourself. See
Terraform on EKS: Post-apply.
Features to enable for a production Helm deployment¶
The chart ships conservative. These features are off by default and will not turn themselves on.
| Value | Default | Consider enabling when |
|---|---|---|
query.walBuffer.enabled |
false |
You want seconds-fresh queryability. Requires an RWX WAL volume. |
compactor.catalogClaim.enabled |
false |
You run more than one compactor pod. Required, not optional: the chart fails the install when compactor.replicas or an enabled autoscaling.compactor.maxReplicas is above 1 without it, and when it is on with wal.mirror.enabled: false. |
keda.enabled |
false |
You want autoscaling on saturation signals. |
serviceMonitor.enabled |
false |
You run Prometheus Operator. |
compactor.indexRebuild |
false |
Searches over streamed compaction output need index pruning rather than a scan, and the parsed indexes fit the query pods' cache. See Post-compaction index rebuilding. |
WAL mirroring is already on when you configure a warehouse URL. The default
single-replica compactor does not reclaim mirror objects, so give
wal-mirror/ an object-store lifecycle expiry. For production, turn on the
Prometheus integration first. Turn on the WAL buffer when queries need
pre-commit data. The rest depend on your workload.
Footer inverted indexes are already on. Set SIGLAKE_INVERTED_INDEX=0,
compactor.invertedIndex.enabled: false, or operator
spec.extraEnv: [{name: SIGLAKE_INVERTED_INDEX, value: "0"}] to opt out.
The post-rewrite Puffin rebuild is not. A compaction or delete-task rewrite
that could not write a footer index leaves its output unindexed, and a search
over those files scans them and returns exact rows. Set
SIGLAKE_INDEX_REBUILD=1, compactor.indexRebuild: true, or operator
spec.extraEnv: [{name: SIGLAKE_INDEX_REBUILD, value: "1"}] to turn the
backfill on. Indexes a file already carries are read whichever way the switch
is set, and with inverted indexes off the pass does nothing. See
Post-compaction index
rebuilding.
Leveled compaction is already on. Set SIGLAKE_COMPACTOR_LEVELED=0 through
compactor.extraEnv only when you want the legacy flat whole-partition pass.
Delete-task execution is on as well: the chart ships
compactor.deleteTasks: true, and the compactor sweeps pending predicate
deletes in its idle cycle unless SIGLAKE_DELETE_TASKS is 0, off, false
or no. Setting deleteTasks: false renders SIGLAKE_DELETE_TASKS=0; under
the operator, put that variable in spec.extraEnv. With the sweep off,
submissions are still accepted and recorded, and only
siglake delete-sweep --apply executes them.
So is the persistent batch-job store: the chart ships
query.jobs.persistent: true and renders SIGLAKE_JOBS_POSTGRES_URI from the
catalog Secret, so every query pod reads the same batch jobs and a job_id
works whichever pod the Service picks. The false opt-out is for a tier you
can hold at one pod. The chart fails the install when query.replicas, or an
enabled keda.query.maxReplicas, is above 1 with the store off, because
status, result and cancel requests would answer 404 from the pod that did not
take the submission.
Read Helm before your first production install.
Operational reading¶
-
Authentication, TLS, tenancy boundaries, and what the audit table contains.
-
Per-role scaling behaviour, sizing rules from the benchmark arc, and autoscaling configuration.
-
The metrics to watch, the shipped alerts, and the Grafana dashboard.
-
WAL mirroring, the recovery procedure, and what each failure mode costs you.
-
Which knob to reach for when a specific thing is slow.
Day-two operations¶
Routine maintenance runs as CronJobs or on the compactor's own loop:
| Task | Mechanism |
|---|---|
| Snapshot expiry | Compactor loop, on by default: every 60 s, retaining 100 snapshots. |
| Orphan GC | siglake gc-orphans --apply on a schedule. |
| Retention | siglake retention-sweep --apply, or per-index policies. |
| Delete tasks | Compactor idle cycle, on by default (compactor.deleteTasks: true); or siglake delete-sweep --apply. |
| Audit rotation | siglake audit-rotate --max-age-secs N. |
| Schema migration | Helm pre-upgrade hook (default), operator Job, or siglake migrate-schema --all-tables --all-namespaces before rollout. |
gc-orphans, retention-sweep and delete-sweep are dry-run by default, so
they report what they would do until you add --apply. Schema migration is the
exception: it applies by default, and you add --dry-run to inspect the
proposed changes first. See
Retention and deletes.
Upgrade¶
Every table records its schema version. When a newer binary tries to populate a
column the table lacks, the write is refused and names the missing column and
the remedy, rather than accepting the row with the value dropped. With the
chart's Prometheus rules enabled, SiglakeSchemaWritesRefused alerts on that.
Migrate every Siglake-owned table and tenant namespace before you roll the new binary. The command is additive and idempotent:
Add --dry-run to see the table's recorded version and the proposed additions
without applying them.
Which control plane runs it depends on your install:
- Helm:
schemaMigration.enableddefaults totrueand runs the command as apre-upgradehook. Helm stops the rollout if the Job fails. - Operator: raise the monotonic
spec.schemaVersion. The operator applies the same migration Job and holds the workload rollout until it completes. - Anything else: run the command yourself before rolling the deployment.
After migration, query pods can roll independently of ingest through the
per-component query.image override.
Ingesters force-seal their WAL on SIGTERM, so a rolling restart does not lose
acknowledged data, as long as terminationGracePeriodSeconds is long enough
for the seal to finish.