Skip to content

Kubernetes operator

Use siglake-operator when you want a SiglakeCluster custom resource and can provide the chart objects it does not render. This guide installs the operator, checks its status, repairs rejected specs and says what moving an existing Helm release onto the operator costs.

The operator reconciles the custom resource into a running deployment. It is leader-elected, and reports observedGeneration and the managed tables' schema versions in status.

Choose Helm or the Kubernetes operator

Use Helm for production unless your control plane requires a SiglakeCluster custom resource definition (CRD). The chart exposes more Siglake features and more Kubernetes policy objects. The operator covers the core data plane, per-tier resources, autoscaling, WAL storage, retention rotation and schema migrations, and it scales the three tiers itself rather than through an autoscaler object.

The operator applies at most nine objects: the WAL PersistentVolumeClaim, the ingester and compactor Deployments, the query StatefulSet, a ClusterIP Service for ingester and one for query, the query headless Service, and, when you ask for them, the schema-migration Job and the audit-rotate CronJob. Workload kinds match the chart. Everything else the chart can render is yours to provide.

Chart capabilities the custom resource cannot express today:

Chart capability What the operator does instead
Ingress (ingress.*) Renders none. Expose the query Service yourself.
PodDisruptionBudget, NetworkPolicy, ServiceMonitor, PrometheusRule Renders none. Apply your own. Both sides label objects with app.kubernetes.io/name, instance and component, so a chart-shaped selector still matches.
HorizontalPodAutoscalers and KEDA ScaledObjects (autoscaling, keda) Renders none, deliberately: the operator reads Prometheus and writes .spec.replicas itself. Point --prometheus-url at the server the HPA used.
Headless compactor metrics Service Renders none. The compactor's metrics port 9101 is reachable on the pod, not through a Service, so a scrape config keyed on that Service finds nothing.
Scheduling: nodeSelector, tolerations, affinity, antiAffinity No field on the custom resource. Rendered pods carry no placement rules.
ServiceAccount creation (serviceAccount.create) and imagePullSecrets Binds an existing account through spec.serviceAccountName and creates nothing. Put pull secrets on that ServiceAccount.
Catalog credentials from a Secret (postgres.existingSecret, externalSecrets) spec.catalogUri is a plain string. It reaches the pods as an environment variable and the audit-rotate CronJob as an argument.
Query authentication: static tokens and the coordinator token (query.tokens, query.distributed.coordinatorToken) Renders neither. spec.authTokensSecretRef covers ingest only, and both query tokens come from a Secret, which spec.extraEnv cannot reference.
Query TLS (query.tls) Renders none, and the custom resource cannot mount a certificate Secret. The operator always passes --query-peer-scheme http.
Query WAL buffer and hot caches (query.walBuffer, query.hotCaches) Renders neither. The query pod mounts only its spill emptyDir, so there is no WAL directory for those features to read.
An existing WAL claim (wal.existingClaim) Applies its own claim, named <cluster>-wal. The --adopt-values preflight reports a claim under any other name as a finding.
Per-tier image and extraArgs, valueFrom environment entries, and <tier>.enabled: false One spec.image for every tier, no argument field, plain name/value spec.extraEnv only, and all three tiers always run.
A tier scaled to zero Ingester and query floors of 0 are refused. A compactor floor of 0 needs the catalog-claim drain and positive EWMA smoothing; the packaged floor remains 1. See Scale the compactor to zero.
Compactor rollout strategy Derived from the compactor policy maximum: 1 renders Recreate, above 1 renders RollingUpdate with maxUnavailable: 0 and maxSurge: 1. Helm always uses Recreate.
Pod security context and migration-Job settings (podSecurityContext, securityContext, schemaMigration.backoffLimit, schemaMigration.ttlSecondsAfterFinished) Renders the chart's default values and accepts no override: UID and GID 65532, no privilege escalation, all capabilities dropped, backoff limit 4, one-day TTL.

Chart settings that are plain environment variables do carry over, through spec.extraEnv: ingest rate limits, backpressure, the memory breaker, OIDC issuer and audience, the Iceberg namespace (SIGLAKE_TENANT_NAMESPACE, which the operator otherwise leaves at the binary's default) and a non-AWS object store (AWS_ENDPOINT_URL). Entries go to every container the operator renders, so a variable both the ingester and the query server read, such as SIGLAKE_OIDC_ISSUER, configures both tiers or neither.

OTLP export of Siglake's own logs and traces is one of those settings. Neither the chart nor the custom resource has an OTel block, so on both sides it is environment variables: OTEL_EXPORTER_OTLP_ENDPOINT and the rest go in spec.extraEnv here, and reach every tier at once. See Siglake's own telemetry.

You cannot migrate an existing Helm release in place. Adoption is an offline preflight plus a manual ownership handover, it has no live-cluster test coverage, and it needs a maintenance window. Keep Helm as the controller until the generated custom resource passes preflight with no findings. Migrate a Helm release to the operator has the steps and the abort path.

The custom resource definition

Property Value
Group siglake.limnion.ai
Version v1alpha1
Kind SiglakeCluster
Short name kcluster
Scope Namespaced

Install

Migrate off the old CRD group before you upgrade the operator

Deployments created before the rename to the siglake.limnion.ai group use the old custom resource definition group. A new operator does not see resources under the old group, so deploying it against an existing cluster can leave two control planes managing overlapping objects. Migrate first.

helm install siglake-operator ./deploy/helm/siglake-operator \
  --namespace siglake-system --create-namespace

Outside Helm, siglake-operator defaults to http://prometheus-server.monitoring.svc.cluster.local:80. The --prometheus-url flag and SIGLAKE_PROMETHEUS_URL environment variable override this address. The chart passes --prometheus-url from prometheus.url. If you use kube-prometheus-stack, set that value to its <release>-prometheus Service on port 9090:

helm install siglake-operator ./deploy/helm/siglake-operator \
  --namespace siglake-system --create-namespace \
  --set-string prometheus.url=http://kube-prometheus-stack-prometheus.monitoring.svc.cluster.local:9090

Then apply a cluster:

apiVersion: siglake.limnion.ai/v1alpha1
kind: SiglakeCluster
metadata:
  name: prod
  namespace: siglake-system
spec:
  image: ghcr.io/siglake/siglake:0.1.0
  warehouseUrl: s3://acme-prod/warehouse
  catalogUri: postgres://siglake:secret@catalog.internal:5432/siglake
  awsRegion: us-west-2
  schemaVersion: 1

  # Bind every rendered Pod to an existing ServiceAccount in this namespace.
  # serviceAccountName: siglake

  # Plain name/value variables appended to every rendered siglake container.
  # extraEnv:
  #   - name: SIGLAKE_QUERY_WARM_INTERVAL_SECS
  #     value: "300"

  authTokensSecretRef:
    name: siglake-auth-tokens
    key: tokens

  storage:
    # WAL knobs; defaults work on any cluster with an RWX StorageClass.

  autoscaling:
    ingester:
      min: 2
      max: 8
      target: 1000.0   # ingest requests / sec / pod
    compactor:
      min: 1
      max: 4
      target: 5.0      # shared sealed segments / worker
    query:
      min: 2
      max: 8
      target: 4.0      # in-flight queries / pod
    ewmaHalfLifeSecs: 0    # use raw signals; non-zero damps flapping

  retention:
    queryAuditRotateIntervalDays: 7

  tenants: []      # advisory only; see below

Check that the operator accepted the spec:

kubectl get siglakecluster prod -n siglake-system -o jsonpath='{.status}'

A rejected spec sets InvalidSpec=True with a reason. The repair table says what to change.

Fields that catch people out

Set awsRegion. Operator-rendered pods, including the retention CronJob, need it in their environment or the OpenDAL S3 client fails with region is missing. The chart path picks the region up from Terraform-emitted Helm values; the operator path bypasses that.

Set spec.serviceAccountName to an existing ServiceAccount in the SiglakeCluster namespace. It binds every rendered Pod, including the retention CronJob. Leave it unset and pods use that namespace's default ServiceAccount, which on EKS usually means S3 403 AccessDenied unless the node IAM role already has warehouse permissions. The Terraform EKS deployment scopes its warehouse IRSA role to siglake/siglake, so use namespace: siglake with serviceAccountName: Siglake, or supply an equivalently authorised ServiceAccount in the custom resource's namespace.

Use spec.extraEnv for tuning knobs that have no dedicated field. Entries are appended to every container the operator renders: watchdog ceilings, the query warm interval, compaction overlap and fan-in limits, the buffer-delta budget, catalog claim batching, the mirror sync interval. A value that needs valueFrom needs a first-class field instead. If you override SIGLAKE_COMPACTOR_BIN_CONCURRENCY, follow the compactor sizing guidance.

For SIGLAKE_WAL_MIRROR_PREFIX, the final matching entry wins, and the operator trims its value. With no entry, the operator uses wal-mirror for both ingester uploads and catalog-claim recovery. A blank value turns mirroring off and is valid only when spec.autoscaling.compactor.max is 1.

The compactor policy maximum selects the drain mode for the life of that policy. A maximum above 1 enables the ingester's remote drain, adds --catalog-claim to the compactor and uses RollingUpdate, even while the autoscaler holds the compactor at one replica. A maximum of 1 uses the filesystem drain and Recreate. Changing the maximum across that boundary needs a manual handover.

With a maximum above 1, autoscaling.compactor.target is sealed segments per worker against one shared catalog-claim queue. The operator divides the queue depth by this target once. A maximum of 1 keeps the filesystem drain's per-pod calculation.

Scale the compactor to zero

Set spec.autoscaling.compactor.min: 0 only when spec.autoscaling.compactor.max is above 1 and spec.autoscaling.ewmaHalfLifeSecs is positive. This selects the catalog-claim drain and keeps the tier from parking after one idle scrape. The packaged floor remains 1, so a default custom resource does not opt in.

The ingesters publish siglake_wal_segments_sealed{tenant} from the shared catalog queue. Its sample-age companion lets the operator discard a publisher whose last successful catalog read is more than two minutes old. The depth counts sealed rows and claims older than SIGLAKE_CLAIM_RECLAIM_MAX_AGE_SECS, so work stranded by the last stopped worker can ask for a new one without counting a live claim twice.

If no fresh ingester reading exists, the operator restores the compactor to one replica. That worker's siglake_compactor_sealed_pending gauge is the fallback until an ingester reading returns. A parked tier also wakes for ten minutes after one hour at zero so retention, delete tasks, claim reclaim and mirror synchronization still run.

No retained cluster round has yet shown a compactor reaching zero and waking from the ingester signal. Keep the packaged floor of 1 unless you accept that qualification gap. A maximum of 1 cannot use this mode because its filesystem drain does not read the shared catalog queue.

Treat spec.tenants as documentation. The list provisions nothing and authorises nobody: the operator renders no tenant argument from it, and namespaces are created lazily on first write. An empty list is the normal case.

Operator-rendered ingest is single-tenant. Every request routes to the default tenant. An X-Scope-OrgID naming another tenant is refused with 403 over HTTP or PermissionDenied over OpenTelemetry Protocol (OTLP)/gRPC; naming default has no effect. If you want more than one tenant, put the opt-in in spec.extraEnv, which the operator appends to every container it renders:

  • SIGLAKE_OIDC_TENANT_CLAIM takes the tenant from a verified JSON Web Token (JWT) claim. Set SIGLAKE_OIDC_ISSUER and SIGLAKE_OIDC_AUDIENCE alongside it. The ingester refuses to start with only one of the two, and ignores the claim with neither.
  • SIGLAKE_TRUST_SCOPE_HEADER=1 routes on X-Scope-OrgID instead. Set it only behind a gateway that supplies the header itself and strips the client's.

Ingest tenant selection has the claim validation rules, the refusal counter and the gateway requirement. Because spec.extraEnv reaches the query containers too, the same three OIDC variables bind query routing to the verified claim. The query tier never routes on X-Scope-OrgID, so SIGLAKE_TRUST_SCOPE_HEADER changes nothing there.

authTokensSecretRef is a reference, not a value. The operator surfaces it to the ingester as SIGLAKE_AUTH_TOKENS through valueFrom and never reads the token bytes. Omit it and ingest is open, which is only safe inside a trusted network.

retention.queryAuditRotateIntervalDays renders a CronJob that runs siglake audit-rotate on that cadence.

autoscaling.query is an ordinary range. The operator scales query on in-flight queries per pod, clamped to [min, max] with the same 10% deadband. It renders --query-peer-discovery-srv pointing at the headless Service's _http._tcp record at every replica count, so a replica the decision adds receives shard work once it is Ready: no rollout, and no ceiling tied to the rendered count. The earlier max == min refusal (QueryAutoscalingRangeUnsupported) is gone. The packaged default is still max: 1, so raise it deliberately.

Raising it above 1 requires a shared batch-job store. The operator renders SIGLAKE_JOBS_POSTGRES_URI from spec.catalogUri when that URI is Postgres, and blank otherwise, and the last matching spec.extraEnv entry overrides either. A blank effective value selects the per-pod in-memory store, which is correct only at one pod, so the operator refuses that spec with QueryJobsStoreRequiredForScaleOut. Three configurations pass: a Postgres spec.catalogUri with no override, a non-blank SIGLAKE_JOBS_POSTGRES_URI in spec.extraEnv, and max: 1 with the in-memory store. The last one loses every in-flight job when the pod restarts.

When one load signal is unusable, because Prometheus is down or its series is empty, that tier holds its current size while tiers with readings continue to scale. Each decision is still clamped into its [min, max] range. A zero-floor compactor with no usable activation reading is restored to one replica.

Per-tier resources

spec.resources.{ingester,compactor,query} overrides each data-plane tier's container requests and limits. Every tier starts from the operator's packaged defaults, which mirror the chart: 200m CPU and 256Mi memory requested, a 2-CPU limit, a 1Gi memory limit on ingester and compactor, and on query a 4Gi memory limit with a 10Gi ephemeral-storage request and a 12Gi limit. Values you supply merge over those defaults field by field:

apiVersion: siglake.limnion.ai/v1alpha1
kind: SiglakeCluster
metadata:
  name: prod
spec:
  resources:
    query:
      limits:
        memory: 8Gi  # CPU limit and all requests keep their packaged values

The query server derives its read caches and its shared memory pool from the container memory limit, so treat spec.resources.query.limits.memory as a sizing budget. See the query memory guidance. The operator also renders query's 10Gi spill emptyDir. If you override query.limits.ephemeral-storage, keep it above that size limit, in the order Query spill and ephemeral storage describes.

If spec.extraEnv raises SIGLAKE_COMPACTOR_BIN_CONCURRENCY above one, the tier's effective memory limit must hold 1Gi plus 4Gi per bin. The operator refuses the spec with InvalidSpec when it cannot.

Schema migrations

Raise the monotonic spec.schemaVersion counter to run a one-shot, additive schema-migration Job before the workloads roll out. The Job runs Siglake migrate-schema --all-tables --all-namespaces. Leaving the field unset, or setting it to 0, renders no Job. The completed version appears in status. On an upgrade, the operator holds the workload rollout while the Job is running or after it fails.

The counter is a trigger, not a target. The Job migrates the warehouse to whatever schema the binary in spec.image declares, so raising the counter means "run a migration", not "reach version N".

The Job's name is <cluster>-migrate-schema-v<N>-<digest>, where the digest covers everything that shapes the pod template: spec.image, spec.warehouseUrl, spec.catalogUri, spec.awsRegion and spec.extraEnv. A Job's spec.template is immutable, so a name keyed on the counter alone would fail with a 422 on an image bump, and it would leave the migration unreachable at that counter value.

Reverting the image

Put the old value back in spec.image and leave spec.schemaVersion where it is. The digest covers the image, so the revert renders a new Job name: the operator applies and waits on one more migration Job, this one running the old binary. Against an already-widened table that run adds nothing and exits 0, the hold clears, and the tiers roll back to the old image.

The migration is not reversed. The added columns stay, the reverted binary writes them as nulls, and rows that already carry values keep them. The table's recorded schema version also stays where the wider binary left it, because the migration stamps the higher of the recorded and declared versions. An image built before that behavior shipped (anything pre-0.1.0) stamps its own lower constant instead, which changes no column and no write decision, and rolling forward restamps it.

Do not lower spec.schemaVersion to match the old image. A lower value is still a trigger, so it renders yet another Job and leaves the reported version naming a schema the warehouse never returned to. To ask the warehouse instead, run:

siglake migrate-schema --dry-run --all-namespaces

That prints each events table's recorded version, and the version the binary you ran declares when the two differ.

The limits are the same as under Helm: additive column changes only, no path across the pre-0.1.0 nanosecond timestamp contract, and no revert qualified against an actual older image. See What a rollback does.

Migrate a Helm release to the operator

Migration is manual, and there is no zero-downtime path. The --adopt-values preflight is offline: it reads a values file, prints a SiglakeCluster with a findings list and a runbook below it as comments, and never contacts a cluster. The ownership handover that follows is yours to run, it has no live-cluster test coverage, and doing it wrong leaves Helm and the operator fighting over the same objects. If that is not a trade you want, keep Helm as the controller.

Before you start: the operator is deployed, the CRD is installed, kubectl is 1.22 or later, the first release whose kubectl patch reads a patch from a file, Python 3 is on the path for steps 7 and 8, and you have a maintenance window. The window is there because the handover has never been exercised against a live cluster, not because any step is known to restart a pod.

  1. Run the whole procedure in one shell. Set the release, its namespace, and a new directory path for the adoption artifacts. Put the directory on storage you control: the effective values and release records can contain credentials. The block creates it with mode 0700 and stops if the path already exists. Use a new path instead of replacing an earlier backup.
RELEASE=siglake
NAMESPACE=siglake
BACKUP_DIR="$PWD/siglake-adoption-backup"

Create the directory before you export any values:

(
set -e
umask 077
mkdir -m 700 -- "$BACKUP_DIR"
)
  1. Export the release's effective values, not only your override file. The export is a values file, not a copy of the release. Helm's own record of the release lives in Secrets, and step 7 saves those.
(
set -eC
umask 077
helm get values "$RELEASE" -n "$NAMESPACE" --all -o yaml \
  > "$BACKUP_DIR/effective-values.yaml"
)
  1. Generate the custom resource and its preflight findings. Name the cluster after the release so names and selectors coincide, name the namespace the release runs in, and pass the catalog URI: the chart builds it from a Postgres Secret, so the preflight cannot infer it.
(
set -eC
umask 077
siglake-operator --adopt-values "$BACKUP_DIR/effective-values.yaml" \
  --adopt-cluster-name "$RELEASE" \
  --adopt-namespace "$NAMESPACE" \
  --adopt-catalog-uri postgres://siglake:secret@catalog.internal:5432/siglake \
  > "$BACKUP_DIR/cluster.yaml"
)

The whole output is one YAML document. The custom resource comes first, then the findings and a handover runbook as comment lines, so the saved file is a manifest kubectl can read: step 8 checks it and step 11 applies it unchanged. --adopt-namespace ships in 0.2.0. It sets metadata.namespace and the -n on every runbook command, and it defaults to the release name, so a release whose namespace differs from its name gets the wrong namespace when you leave the flag out.

Each tier's replicas becomes a fixed band (min equals max), so the handover changes no pod count. Each <tier>.resources block carries onto spec.resources; an omitted block takes the operator's packaged defaults, which are the chart's.

  1. Clear every finding. A finding names a value the custom resource cannot carry, and each one is a blocker. Findings cover an empty image.tag, a missing s3.bucket, a non-AWS s3.endpoint, a tenant.namespace other than siglake, an ingester.auth.secretKey with no existingSecret, per-tier extraEnv folded cluster-wide, a detection tier still enabled in values, and each chart opt-out the operator switches back on: compactor.deleteTasks, compactor.invertedIndex.enabled, wal.mirror.enabled and query.jobs.persistent. Fix each one by adding the equivalent spec.extraEnv entry, or keep Helm.

compactor.indexRebuild: true is reported for the opposite reason. The operator renders SIGLAKE_INDEX_REBUILD=0, so the adopted compactor stops registering Puffin sidecars for rewrite output. Keep the opt-in with SIGLAKE_INDEX_REBUILD=1 in spec.extraEnv. Indexes already registered stay readable either way.

The WAL claim is one of those findings, and the one that costs data if you miss it. The preflight carries wal.size, wal.accessMode and wal.storageClassName onto spec.storage, because the operator applies a claim of the same name and the API server rejects a shrink or a storage-class change. The claim name is separate: the chart mounts wal.existingClaim when you set it, and the operator always renders and mounts <cluster>-wal. When the two differ, the finding begins wal.existingClaim is <name> and names both claims, and says the adopted ingester and compactor would start against a different, probably empty WAL while every segment on the old claim that has not reached Iceberg stays there unread. An absent, empty or already-matching existingClaim is silent, because all three are the same volume. The preflight only reports the mismatch: it moves no data and renames no volume, so drain the old claim to empty before the handover, or keep Helm.

  1. Resolve a drain-mode disagreement. Helm selects the drain with compactor.catalogClaim.enabled and always uses Recreate. The generated custom resource selects it from the fixed compactor policy synthesised from compactor.replicas. When the two disagree, the finding begins compactor.catalogClaim.enabled is <value> and says the synthesised policy selects the other mode. Align the chart with that mode and finish its drain, or set spec.autoscaling.compactor.max above 1 to retain catalog claims and accept a scale-out-capable policy.

  2. In the maintenance window, confirm that the objects the operator expects already exist under the release name.

kubectl get -n "$NAMESPACE" \
  "deploy/$RELEASE-ingester" "deploy/$RELEASE-compactor" "sts/$RELEASE-query"
  1. Back up the release records and the current ownership metadata. Step 10 deletes Helm's record of the release, and the two abort paths below read this backup. Every file has mode 0600. The block refuses existing files and symlinks instead of replacing them. It exits non-zero if either read fails or if the label selector matches no Secret, so a failed backup stops the sequence before step 9 changes anything.
(
set -eC
umask 077
if [ ! -d "$BACKUP_DIR" ] || [ -L "$BACKUP_DIR" ]; then
  echo "unsafe adoption artifact directory: $BACKUP_DIR" >&2
  exit 1
fi
chmod 700 "$BACKUP_DIR"

kubectl get secret -n "$NAMESPACE" -l "owner=helm,name=$RELEASE" -o json \
  > "$BACKUP_DIR/release-secrets.json"
kubectl get -n "$NAMESPACE" -o json \
  "deploy/$RELEASE-ingester" "deploy/$RELEASE-compactor" "sts/$RELEASE-query" \
  > "$BACKUP_DIR/workloads.json"

python3 - "$BACKUP_DIR" "$RELEASE" <<'PY'
import json
import os
import sys

backup, release = sys.argv[1:3]
workload_names = (
    f"{release}-ingester",
    f"{release}-compactor",
    f"{release}-query",
)
OWNERSHIP = (
    ("annotations", "meta.helm.sh/release-name"),
    ("annotations", "meta.helm.sh/release-namespace"),
    ("labels", "app.kubernetes.io/managed-by"),
)
ASSIGNED = (
    "creationTimestamp",
    "generation",
    "managedFields",
    "resourceVersion",
    "selfLink",
    "uid",
)


def load(name):
    with open(os.path.join(backup, name)) as handle:
        return json.load(handle)


def write(name, document):
    path = os.path.join(backup, name)
    handle = os.open(path, os.O_WRONLY | os.O_CREAT | os.O_EXCL, 0o600)
    with os.fdopen(handle, "w") as out:
        json.dump(document, out, indent=2)
        out.write("\n")


records = load("release-secrets.json").get("items") or []
if not records:
    sys.exit(f"no Helm release records for {release}; nothing to back up")

workloads = {
    item["metadata"]["name"]: item["metadata"]
    for item in load("workloads.json").get("items") or []
}
absent = [name for name in workload_names if name not in workloads]
if absent:
    sys.exit("no ownership metadata for: " + ", ".join(absent))

patches = {}
for name in workload_names:
    metadata = workloads[name]
    patch = {}
    for section, key in OWNERSHIP:
        current = metadata.get(section) or {}
        patch.setdefault(section, {})[key] = current.get(key)
    patches[name] = {"metadata": patch}

for record in records:
    for field in ASSIGNED:
        record["metadata"].pop(field, None)

write("release-restore.json", {"apiVersion": "v1", "kind": "List", "items": records})
for name, patch in patches.items():
    write(f"ownership-{name}.json", patch)
print(f"saved {len(records)} release records and {len(patches)} ownership patches")
PY
)

release-restore.json is the same Secrets without the fields the API server assigns, among them resourceVersion and uid. A create that carries a resourceVersion is rejected, so release-secrets.json is the record to keep and the stripped copy is the one to apply. Each ownership-<name>.json is a merge patch that puts one workload's two Helm annotations and its app.kubernetes.io/managed-by label back exactly as they were. An annotation or label that is absent today is recorded as null, which the merge patch removes on restore rather than inventing a value.

  1. Accept the saved manifest for this release and namespace. Run this after the edits in steps 4 and 5, because it reads the file as it stands. The block reads the resource at the top of $BACKUP_DIR/cluster.yaml, ignoring the comment lines, and stops unless it is a SiglakeCluster named $RELEASE in namespace $NAMESPACE. On success it writes $BACKUP_DIR/manifest-accepted with mode 0600, which is the file step 9 looks for. A failure here stops the sequence with the annotations and the release Secrets untouched.
(
set -eC
umask 077
python3 - "$BACKUP_DIR" "$RELEASE" "$NAMESPACE" <<'PY'
import os
import re
import sys

backup, release, namespace = sys.argv[1:4]
saved = os.path.join(backup, "cluster.yaml")

fields, section = {}, None
with open(saved) as handle:
    for line in handle.read().splitlines():
        if not line.strip() or line.lstrip().startswith("#"):
            continue
        if not line[:1].isspace():
            section = line.split(":", 1)[0]
            top = re.match(r"(\w+): *(\S.*)$", line)
            if top:
                fields[top.group(1)] = top.group(2).strip("\"'")
            continue
        if section == "metadata":
            field = re.match(r"  (\w+): *(\S.*)$", line)
            if field:
                fields["metadata." + field.group(1)] = field.group(2).strip("\"'")

kind = fields.get("kind")
if kind != "SiglakeCluster":
    sys.exit(f"{saved} holds kind {kind or '(none)'}, not SiglakeCluster")

name = fields.get("metadata.name")
if name != release:
    sys.exit(f"{saved} names {name or '(none)'}, not the release {release}")

target = fields.get("metadata.namespace")
if target != namespace:
    sys.exit(
        f"{saved} is for namespace {target or '(none)'}, not {namespace}: "
        "rerun step 3 with --adopt-namespace"
    )

path = os.path.join(backup, "manifest-accepted")
handle = os.open(path, os.O_WRONLY | os.O_CREAT | os.O_EXCL, 0o600)
with os.fdopen(handle, "w") as out:
    out.write(f"{kind}/{name} for namespace {namespace}\n")
print(f"accepted {kind}/{name} for namespace {namespace}")
PY
)

The namespace is decided in step 3 and checked here. A preflight run without --adopt-namespace stamps the release name instead, and a resource for the wrong namespace would create the cluster beside the workloads whose Helm ownership steps 9 and 10 remove, with nothing to reconcile them. Step 11 passes -n "$NAMESPACE" as well, and kubectl refuses a file whose metadata.namespace disagrees with it.

  1. Flip ownership. Helm forgets the objects without deleting them. This is the first step that changes the cluster, so the block stops unless steps 7 and 8 left every file the abort paths and step 11 read. A failed annotation removal stops it before it changes a label. It writes ownership-flipped only when both commands succeeded, and step 10 refuses to run without that file.
(
set -e
umask 077
for artifact in "release-restore.json" "manifest-accepted" "cluster.yaml" \
  "ownership-$RELEASE-ingester.json" "ownership-$RELEASE-compactor.json" \
  "ownership-$RELEASE-query.json"; do
  if [ ! -f "$BACKUP_DIR/$artifact" ]; then
    echo "missing $artifact: finish steps 7 and 8 first" >&2
    exit 1
  fi
done

kubectl annotate -n "$NAMESPACE" \
  "deploy/$RELEASE-ingester" "deploy/$RELEASE-compactor" "sts/$RELEASE-query" \
  meta.helm.sh/release-name- meta.helm.sh/release-namespace- || {
  status=$?
  echo "annotation removal failed: Abort before the release records are deleted" >&2
  exit "$status"
}
kubectl label -n "$NAMESPACE" \
  "deploy/$RELEASE-ingester" "deploy/$RELEASE-compactor" "sts/$RELEASE-query" \
  app.kubernetes.io/managed-by=siglake-operator --overwrite || {
  status=$?
  echo "label change failed: Abort before the release records are deleted" >&2
  exit "$status"
}
: > "$BACKUP_DIR/ownership-flipped"
)

Either command can change some of the three objects and fail on the rest, so the block names an abort path instead of undoing what it did. The pre-deletion abort patches all three objects back to the metadata step 7 recorded, whether or not the flip reached them.

  1. Delete the Helm release Secrets, which are Helm's record of the release and its revision history. The block stops unless step 9 wrote ownership-flipped, and it writes records-deleted only when the deletion succeeded.

    (
    set -e
    umask 077
    if [ ! -f "$BACKUP_DIR/ownership-flipped" ]; then
      echo "step 9 did not finish: do not delete Helm's records" >&2
      exit 1
    fi
    kubectl delete secret -n "$NAMESPACE" -l "owner=helm,name=$RELEASE" || {
      status=$?
      echo "deletion failed: Abort after the release records are deleted" >&2
      exit "$status"
    }
    : > "$BACKUP_DIR/records-deleted"
    )
    

    A deletion that removed some Secrets and failed on the rest looks from outside like one that removed none, so a failure here points at the post-deletion abort. That block applies release-restore.json, which recreates the deleted records and rewrites any survivor with the same content.

  2. Apply the manifest step 8 accepted, in the release's namespace. This is the file step 3 saved, comments and all. When the specs match, the first reconcile is a no-op apply. The block stops unless step 10 wrote records-deleted.

    (
    set -e
    if [ ! -f "$BACKUP_DIR/records-deleted" ]; then
      echo "step 10 did not finish: do not apply the custom resource" >&2
      exit 1
    fi
    kubectl apply -n "$NAMESPACE" -f "$BACKUP_DIR/cluster.yaml" || {
      status=$?
      echo "apply failed: Abort after the release records are deleted" >&2
      exit "$status"
    }
    )
    

    If the apply fails, ask the API server whether the resource exists before you decide:

    kubectl get siglakecluster "$RELEASE" -n "$NAMESPACE"
    

    If it reports NotFound, nothing is reconciling the workloads and the post-deletion abort still returns them to Helm.

  3. Verify. Status reaches Ready=True with no pod restarts, and the replica counts in status match what Helm was running.

    kubectl get siglakecluster "$RELEASE" -n "$NAMESPACE" -w
    

Abort before the release records are deleted

Nothing in steps 1 through 8 changes the cluster. After the ownership flip in step 9 and before the deletion in step 10, the workloads keep running with no controller and Helm's records are still in the namespace. A step 9 that stopped part way lands here too. Put the saved ownership metadata back, then confirm Helm sees the release again. Like the rest of the handover, these commands have no live-cluster test coverage.

(
set -e
for workload in "deploy/$RELEASE-ingester" "deploy/$RELEASE-compactor" \
  "sts/$RELEASE-query"; do
  kubectl patch -n "$NAMESPACE" "$workload" --type=merge \
    --patch-file "$BACKUP_DIR/ownership-$(basename "$workload").json"
done
helm status "$RELEASE" -n "$NAMESPACE"
helm history "$RELEASE" -n "$NAMESPACE"
)

Abort after the release records are deleted

After step 10, restoring the annotations and labels alone leaves Helm with no release to manage: helm status reports that the release is not found. Apply the saved release records first, then the ownership patches. A failed apply stops the block before it patches anything. A step 10 that deleted part of the list, and a step 11 that failed before the custom resource existed, both end here.

(
set -e
kubectl apply -n "$NAMESPACE" -f "$BACKUP_DIR/release-restore.json"
for workload in "deploy/$RELEASE-ingester" "deploy/$RELEASE-compactor" \
  "sts/$RELEASE-query"; do
  kubectl patch -n "$NAMESPACE" "$workload" --type=merge \
    --patch-file "$BACKUP_DIR/ownership-$(basename "$workload").json"
done
helm status "$RELEASE" -n "$NAMESPACE"
helm history "$RELEASE" -n "$NAMESPACE"
)

Either block ends with the two commands that prove the recovery: helm status reports the release with its last revision, and helm history lists every revision the backup held. If helm history lists fewer revisions than the release had, the backup missed records, and Helm can upgrade the release but cannot roll it back past what is listed.

Once you apply the custom resource in step 11, this procedure stops making promises. The operator owns the objects from that point, and returning them to Helm is not a path this page covers or anyone has tested.

Design record: docs/DESIGN_operator_adoption.md.

Status

The operator reports:

  • observedGeneration, which tells you whether it has caught up with your last edit;
  • schema versions for the managed tables;
  • managed replica counts (siglake_operator_managed_replicas);
  • Kubernetes conditions. Ready reports whether the managed workloads reached their desired replica counts, Progressing reports a replica-count change, InvalidSpec reports that the requested specification cannot be honoured, and DrainModeCompatible=False reports that an existing workload template uses a different drain from the compactor policy.

Invalid specifications

When the operator cannot honour a SiglakeCluster specification it sets InvalidSpec=True with a specific reason, and Ready=False with reason InvalidSpec. Validation runs before any child resource is rendered or applied, so existing managed resources are left alone and a new invalid cluster gets nothing. The operator retries after five minutes to avoid a hot loop, and an edit to the SiglakeCluster wakes reconciliation immediately.

InvalidSpec reason Repair
ImageRequired Set spec.image to the container image to run.
WarehouseUrlRequired Set spec.warehouseUrl to the Iceberg warehouse URL.
CatalogUriRequired Set spec.catalogUri to the Iceberg catalog URI.
AutoscalingRangeInvalid For every tier, set min to 0 or more and max to at least min.
AutoscalingZeroFloorUnsupported Set ingester and query min to 1 or more. For a compactor floor of 0, set compactor.max above 1 so the ingesters can publish the shared catalog queue depth.
AutoscalingZeroFloorNeedsSmoothing Set spec.autoscaling.ewmaHalfLifeSecs above 0, or restore the compactor floor to 1. Positive smoothing keeps one idle scrape from parking the tier.
AutoscalingTargetNotPositive Give every tier a finite target greater than zero.
AutoscalingEwmaHalfLifeInvalid Set spec.autoscaling.ewmaHalfLifeSecs to a finite, non-negative number; zero disables smoothing.
RetentionIntervalUnsupported Set spec.retention.queryAuditRotateIntervalDays to 1 (daily), 7 (weekly) or 30 (monthly), or omit it to disable rotation.
CompactorBinConcurrencyInvalid Set SIGLAKE_COMPACTOR_BIN_CONCURRENCY in spec.extraEnv to a positive integer, or remove the override.
CompactorBinConcurrencyExceedsMemory Reduce bin concurrency, or raise the effective compactor memory limit to at least 1Gi plus 4Gi per concurrent bin. See Per-tier resources.
QueryJobsStoreRequiredForScaleOut Give the query tier a shared batch-job store, or hold spec.autoscaling.query.max at 1. Set a non-blank SIGLAKE_JOBS_POSTGRES_URI in spec.extraEnv, or use a Postgres spec.catalogUri and drop any override of that variable.
WalMirrorRequiredForCompactorScaleOut Set a non-blank SIGLAKE_WAL_MIRROR_PREFIX in spec.extraEnv, or hold spec.autoscaling.compactor.max at 1.

Handover between drain modes

If an existing ingester or compactor Deployment uses the other drain mode, the operator sets DrainModeCompatible=False and Ready=False. Both conditions use reason DrainModeHandoverRequired. The operator patches status only, then retries after five minutes. It does not query Prometheus or apply child resources while the mismatch remains.

If you did not intend to change drain mode, restore the previous spec.autoscaling.compactor.max. To complete the handover:

  1. Stop ingestion. Confirm that clients and exporters no longer send requests to the ingester.
  2. Let the current drain finish. Watch siglake_compactor_sealed_pending fall to zero, then query the affected tables for the expected rows. A zero gauge alone does not prove that every mirror object has a catalog row.
  3. Account for retained local WAL segments, mirror objects and catalog rows. Preserve anything your recovery plan still needs. The operator does not convert or delete any of them during the handover. Inventory held orphans/ files before you switch.
  4. Delete the ingester and compactor Deployments so the operator can recreate both templates in the selected mode:
kubectl delete deployment <cluster>-ingester <cluster>-compactor -n <namespace>
  1. Wait for the recreated workloads. Confirm that Ready=True and that DrainModeHandoverRequired is absent from status before you resume ingestion.

Operator metrics: siglake_operator_reconciles_total, siglake_operator_reconcile_duration_seconds and siglake_operator_prom_query_errors_total{namespace}.