Skip to content

Operations

Use this page to pick a deployment path, gather what the chart expects you to bring, and change the defaults that matter before you take traffic.

Deployment paths

Path Use when Page
Docker Compose Local development, evaluation, single-host demos. Quickstart
Helm Kubernetes. The most complete and current path. Helm
Operator You want to manage workloads through a SiglakeCluster custom resource and configure supporting resources separately. Operator
Terraform and EKS Greenfield AWS. Provisions EKS, EFS, RDS and S3, and emits matching Helm values. Terraform on EKS

The operator renders ingester and compactor as Deployments and query as a StatefulSet, which matches the chart's workload kinds. Everything else the chart renders, the operator does not: Ingress, PodDisruptionBudgets, NetworkPolicies, ServiceMonitors, autoscaling objects, query-tier authentication and TLS. Choose Helm unless you specifically want the custom resource.

Object storage, Iceberg catalog, WAL volume, and Kubernetes prerequisites

Before you install Siglake, provide Kubernetes 1.27 or later and the infrastructure below. The chart references these resources rather than provisioning them.

  1. A Postgres instance reachable from the cluster, for the Iceberg catalog. RDS in the reference deployment.
  2. A Kubernetes Secret holding its connection details, with the keys host, port, user, password and database.
  3. An S3 bucket for the Iceberg warehouse. This is the object storage every Parquet file lands in.
  4. An IAM role with read and write on that bucket, trusted by the cluster's OIDC provider and annotated on the ServiceAccount. On EKS this is IRSA, IAM Roles for Service Accounts.
  5. A StorageClass for the shared write-ahead log (WAL) volume, ReadWriteMany in most deployments.

The Terraform module in deploy/terraform/aws/ provisions items 1, 3 and 4, and emits Helm values matching the chart's schema.

Choosing the WAL StorageClass

Use ReadWriteMany (RWX) whenever WAL consumers may run on different nodes, which includes scaling ingesters and enabling the query freshness buffer without co-locating their pods. ReadWriteOnce (RWO) requires every WAL consumer on one node, and the chart does not co-locate them for you. If you want RWO, set matching nodeSelector mappings under ingester and compactor to select one uniquely labelled node, add the same mapping under query when the freshness buffer is on, and leave antiAffinity.enabled false.

On EKS, RWX means the AWS EFS CSI driver and an EFS-backed StorageClass. The Terraform module installs the driver as a managed addon, and deploy/aws/up.sh creates the StorageClass after apply. On a cluster you did not build with Terraform, install the driver and create the StorageClass yourself. See Terraform on EKS: Post-apply.

Features to enable for a production Helm deployment

The chart ships conservative. These features are off by default and will not turn themselves on.

Value Default Consider enabling when
query.walBuffer.enabled false You want seconds-fresh queryability. Requires an RWX WAL volume.
compactor.catalogClaim.enabled false You run more than one compactor pod. Required, not optional: the chart fails the install when compactor.replicas or an enabled autoscaling.compactor.maxReplicas is above 1 without it, and when it is on with wal.mirror.enabled: false.
keda.enabled false You want autoscaling on saturation signals.
serviceMonitor.enabled false You run Prometheus Operator.
compactor.indexRebuild false Searches over streamed compaction output need index pruning rather than a scan, and the parsed indexes fit the query pods' cache. See Post-compaction index rebuilding.

WAL mirroring is already on when you configure a warehouse URL. The default single-replica compactor does not reclaim mirror objects, so give <s3.warehousePrefix>/wal-mirror/ an object-store lifecycle expiry. For production, turn on the Prometheus integration first. Turn on the WAL buffer when queries need pre-commit data. The rest depend on your workload.

Footer inverted indexes are already on. Set SIGLAKE_INVERTED_INDEX=0, compactor.invertedIndex.enabled: false, or operator spec.extraEnv: [{name: SIGLAKE_INVERTED_INDEX, value: "0"}] to opt out.

The post-rewrite Puffin rebuild is not. A compaction or delete-task rewrite that could not write a footer index leaves its output unindexed, and a search over those files scans them and returns exact rows. Set SIGLAKE_INDEX_REBUILD=1, compactor.indexRebuild: true, or operator spec.extraEnv: [{name: SIGLAKE_INDEX_REBUILD, value: "1"}] to turn the backfill on. Indexes a file already carries are read whichever way the switch is set, and with inverted indexes off the pass does nothing. See Post-compaction index rebuilding.

Leveled compaction is already on. Set SIGLAKE_COMPACTOR_LEVELED=0 through compactor.extraEnv only when you want the legacy flat whole-partition pass.

Delete-task execution is on as well: the chart ships compactor.deleteTasks: true, and the compactor sweeps pending predicate deletes in its idle cycle unless SIGLAKE_DELETE_TASKS is 0, off, false or no. Setting deleteTasks: false renders SIGLAKE_DELETE_TASKS=0; under the operator, put that variable in spec.extraEnv. With the sweep off, submissions are still accepted and recorded, and only siglake delete-sweep --apply executes them.

So is the persistent batch-job store: the chart ships query.jobs.persistent: true and renders SIGLAKE_JOBS_POSTGRES_URI from the catalog Secret, so every query pod reads the same batch jobs and a job_id works whichever pod the Service picks. The false opt-out is for a tier you can hold at one pod. The chart fails the install when query.replicas, or an enabled keda.query.maxReplicas, is above 1 with the store off, because status, result and cancel requests would answer 404 from the pod that did not take the submission.

Read Helm before your first production install.

Operational reading

  • Security

    Authentication, TLS, tenancy boundaries, and what the audit table contains.

  • Scaling

    Per-role scaling behaviour, sizing rules from the benchmark arc, and autoscaling configuration.

  • Monitoring

    The metrics to watch, the shipped alerts, and the Grafana dashboard.

  • Siglake's own telemetry

    Export Siglake's logs and traces over OTLP, and the collector an OTLP-only backend needs for its Prometheus metrics.

  • Disaster recovery

    WAL mirroring, the recovery procedure, and what each failure mode costs you.

  • Performance tuning

    Which knob to reach for when a specific thing is slow.

  • Profiling

    The PROFILING=1 image, the two opt-ins that arm /debug/pprof/*, and how to capture a CPU, heap or runtime profile.

Day-two operations

Routine maintenance runs as CronJobs or on the compactor's own loop:

Task Mechanism
Snapshot expiry Compactor loop, on by default: every 60 s, retaining 100 snapshots.
Orphan GC siglake gc-orphans --apply on a schedule.
Retention siglake retention-sweep --apply, or per-index policies.
Delete tasks Compactor idle cycle, on by default (compactor.deleteTasks: true); or siglake delete-sweep --apply.
Audit rotation siglake audit-rotate --max-age-secs N.
Schema migration Helm pre-upgrade hook (default), operator Job, or siglake migrate-schema --all-tables --all-namespaces before rollout.

gc-orphans, retention-sweep and delete-sweep are dry-run by default, so they report what they would do until you add --apply. Schema migration is the exception: it applies by default, and you add --dry-run to inspect the proposed changes first. See Retention and deletes.

Upgrade

Every table records its schema version. When a newer binary tries to populate a column the table lacks, the write is refused and names the missing column and the remedy, rather than accepting the row with the value dropped. With the chart's Prometheus rules enabled, SiglakeSchemaWritesRefused alerts on that.

Migrate every Siglake-owned table and tenant namespace before you roll the new binary. The command is additive and idempotent:

siglake migrate-schema --all-tables --all-namespaces

Add --dry-run to see the table's recorded version and the proposed additions without applying them.

Which control plane runs it depends on your install:

  • Helm: schemaMigration.enabled defaults to true and runs the command as a pre-upgrade hook. Helm stops the rollout if the Job fails.
  • Operator: raise the monotonic spec.schemaVersion. The operator applies the same migration Job and holds the workload rollout until it completes.
  • Anything else: run the command yourself before rolling the deployment.

After migration, query pods can roll independently of ingest through the per-component query.image override.

Ingesters force-seal their WAL on SIGTERM, so a rolling restart does not lose acknowledged data, as long as terminationGracePeriodSeconds is long enough for the seal to finish.

Validate released artifacts

Run release validation to check clean installation, repeatable baseline correctness, component restarts and 24/72-hour operation. The evidence ledger distinguishes completed tests, failures and remaining qualification gaps for each public release.