Kubernetes operator¶
Use siglake-operator when you want a SiglakeCluster custom resource and can
provide the chart objects it does not render. This guide installs the operator,
checks its status, repairs rejected specs and says what moving an existing Helm
release onto the operator costs.
The operator reconciles the custom resource into a running deployment. It is
leader-elected, and reports observedGeneration and the managed tables' schema
versions in status.
Choose Helm or the Kubernetes operator¶
Use Helm for production unless your control plane requires a SiglakeCluster
custom resource definition (CRD). The chart exposes more Siglake features and
more Kubernetes policy objects. The operator covers the core data plane,
per-tier resources, autoscaling, WAL storage, retention rotation and schema
migrations, and it scales the three tiers itself rather than through an
autoscaler object.
The operator applies at most nine objects: the WAL PersistentVolumeClaim, the ingester and compactor Deployments, the query StatefulSet, a ClusterIP Service for ingester and one for query, the query headless Service, and, when you ask for them, the schema-migration Job and the audit-rotate CronJob. Workload kinds match the chart. Everything else the chart can render is yours to provide.
Chart capabilities the custom resource cannot express today:
| Chart capability | What the operator does instead |
|---|---|
Ingress (ingress.*) |
Renders none. Expose the query Service yourself. |
| PodDisruptionBudget, NetworkPolicy, ServiceMonitor, PrometheusRule | Renders none. Apply your own. Both sides label objects with app.kubernetes.io/name, instance and component, so a chart-shaped selector still matches. |
HorizontalPodAutoscalers and KEDA ScaledObjects (autoscaling, keda) |
Renders none, deliberately: the operator reads Prometheus and writes .spec.replicas itself. Point --prometheus-url at the server the HPA used. |
| Headless compactor metrics Service | Renders none. The compactor's metrics port 9101 is reachable on the pod, not through a Service, so a scrape config keyed on that Service finds nothing. |
Scheduling: nodeSelector, tolerations, affinity, antiAffinity |
No field on the custom resource. Rendered pods carry no placement rules. |
ServiceAccount creation (serviceAccount.create) and imagePullSecrets |
Binds an existing account through spec.serviceAccountName and creates nothing. Put pull secrets on that ServiceAccount. |
Catalog credentials from a Secret (postgres.existingSecret, externalSecrets) |
spec.catalogUri is a plain string. It reaches the pods as an environment variable and the audit-rotate CronJob as an argument. |
Query authentication: static tokens and the coordinator token (query.tokens, query.distributed.coordinatorToken) |
Renders neither. spec.authTokensSecretRef covers ingest only, and both query tokens come from a Secret, which spec.extraEnv cannot reference. |
Query TLS (query.tls) |
Renders none, and the custom resource cannot mount a certificate Secret. The operator always passes --query-peer-scheme http. |
Query WAL buffer and hot caches (query.walBuffer, query.hotCaches) |
Renders neither. The query pod mounts only its spill emptyDir, so there is no WAL directory for those features to read. |
An existing WAL claim (wal.existingClaim) |
Applies its own claim, named <cluster>-wal. The --adopt-values preflight reports a claim under any other name as a finding. |
Per-tier image and extraArgs, valueFrom environment entries, and <tier>.enabled: false |
One spec.image for every tier, no argument field, plain name/value spec.extraEnv only, and all three tiers always run. |
| A tier scaled to zero | Refused. Every spec.autoscaling.<tier>.min must be 1 or more, because each tier's scaling signal comes from that tier's own pods. |
| Compactor rollout strategy | Derived from the compactor policy maximum: 1 renders Recreate, above 1 renders RollingUpdate with maxUnavailable: 0 and maxSurge: 1. Helm always uses Recreate. |
Pod security context and migration-Job settings (podSecurityContext, securityContext, schemaMigration.backoffLimit, schemaMigration.ttlSecondsAfterFinished) |
Renders the chart's default values and accepts no override: UID and GID 65532, no privilege escalation, all capabilities dropped, backoff limit 4, one-day TTL. |
Chart settings that are plain environment variables do carry over, through
spec.extraEnv: ingest rate limits, backpressure, the memory breaker, OIDC
issuer and audience, the Iceberg namespace (SIGLAKE_TENANT_NAMESPACE, which
the operator otherwise leaves at the binary's default) and a non-AWS object
store (AWS_ENDPOINT_URL). Entries go to every container the operator renders,
so a variable both the ingester and the query server read, such as
SIGLAKE_OIDC_ISSUER, configures both tiers or neither.
You cannot migrate an existing Helm release in place. Adoption is an offline preflight plus a manual ownership handover, it has no live-cluster test coverage, and it needs a maintenance window. Keep Helm as the controller until the generated custom resource passes preflight with no findings. Migrate a Helm release to the operator has the steps and the abort path.
The custom resource definition¶
| Property | Value |
|---|---|
| Group | siglake.limnion.ai |
| Version | v1alpha1 |
| Kind | SiglakeCluster |
| Short name | kcluster |
| Scope | Namespaced |
Install¶
Migrate off the old CRD group before you upgrade the operator
Deployments created before the rename to the siglake.limnion.ai group use
the old custom resource definition group. A new operator does not see
resources under the old group, so deploying it against an existing cluster
can leave two control planes managing overlapping objects. Migrate first.
helm install siglake-operator ./deploy/helm/siglake-operator \
--namespace siglake-system --create-namespace
Outside Helm, siglake-operator defaults to
http://prometheus-server.monitoring.svc.cluster.local:80. The
--prometheus-url flag and SIGLAKE_PROMETHEUS_URL environment variable
override this address. The chart passes --prometheus-url from
prometheus.url. If you use kube-prometheus-stack, set that value to its
<release>-prometheus Service on port 9090:
helm install siglake-operator ./deploy/helm/siglake-operator \
--namespace siglake-system --create-namespace \
--set-string prometheus.url=http://kube-prometheus-stack-prometheus.monitoring.svc.cluster.local:9090
Then apply a cluster:
apiVersion: siglake.limnion.ai/v1alpha1
kind: SiglakeCluster
metadata:
name: prod
namespace: siglake-system
spec:
image: ghcr.io/siglake/siglake:0.1.0
warehouseUrl: s3://acme-prod/warehouse
catalogUri: postgres://siglake:secret@catalog.internal:5432/siglake
awsRegion: us-west-2
schemaVersion: 1
# Bind every rendered Pod to an existing ServiceAccount in this namespace.
# serviceAccountName: siglake
# Plain name/value variables appended to every rendered siglake container.
# extraEnv:
# - name: SIGLAKE_QUERY_WARM_INTERVAL_SECS
# value: "300"
authTokensSecretRef:
name: siglake-auth-tokens
key: tokens
storage:
# WAL knobs; defaults work on any cluster with an RWX StorageClass.
autoscaling:
ingester:
min: 2
max: 8
target: 1000.0 # ingest requests / sec / pod
compactor:
min: 1
max: 4
target: 5.0 # shared sealed segments / worker
query:
min: 2
max: 8
target: 4.0 # in-flight queries / pod
ewmaHalfLifeSecs: 0 # use raw signals; non-zero damps flapping
retention:
queryAuditRotateIntervalDays: 7
tenants: [] # advisory only; see below
Check that the operator accepted the spec:
A rejected spec sets InvalidSpec=True with a reason. The
repair table says what to change.
Fields that catch people out¶
Set awsRegion. Operator-rendered pods, including the retention CronJob, need
it in their environment or the OpenDAL S3 client fails with region is
missing. The chart path picks the region up from Terraform-emitted Helm
values; the operator path bypasses that.
Set spec.serviceAccountName to an existing ServiceAccount in the
SiglakeCluster namespace. It binds every rendered Pod, including the
retention CronJob. Leave it unset and pods use that namespace's default
ServiceAccount, which on EKS usually means S3 403 AccessDenied unless the
node IAM role already has warehouse permissions. The Terraform EKS
deployment scopes its warehouse IRSA role to
siglake/siglake, so use namespace: siglake with serviceAccountName:
siglake, or supply an equivalently authorised ServiceAccount in the custom
resource's namespace.
Use spec.extraEnv for tuning knobs that have no dedicated field. Entries are
appended to every container the operator renders: watchdog ceilings,
the query warm interval, compaction overlap and fan-in limits, the buffer-delta
budget, catalog claim batching, the mirror sync interval. A value that needs
valueFrom needs a first-class field instead. If you override
SIGLAKE_COMPACTOR_BIN_CONCURRENCY, follow the compactor sizing
guidance.
For SIGLAKE_WAL_MIRROR_PREFIX, the final matching entry wins, and the
operator trims its value. With no entry, the operator uses wal-mirror for
both ingester uploads and catalog-claim recovery. A blank value turns
mirroring off and is valid only when spec.autoscaling.compactor.max is 1.
The compactor policy maximum selects the drain mode for the life of that
policy. A maximum above 1 enables the ingester's remote drain, adds
--catalog-claim to the compactor and uses RollingUpdate, even while the
autoscaler holds the compactor at one replica. A maximum of 1 uses the
filesystem drain and Recreate.
Changing the maximum across that boundary needs a manual
handover.
With a maximum above 1, autoscaling.compactor.target is sealed segments per
worker against one shared catalog-claim queue. The operator divides the queue
depth by this target once. A maximum of 1 keeps the filesystem drain's
per-pod calculation.
Treat spec.tenants as documentation. The list provisions nothing and
authorises nobody: the operator renders no tenant argument from it, and
namespaces are created lazily on first write. An empty list is the normal case.
Operator-rendered ingest is single-tenant. Every request routes to the default
tenant. An X-Scope-OrgID naming another tenant is refused with 403 over HTTP
or PermissionDenied over OpenTelemetry Protocol (OTLP)/gRPC; naming default
has no effect. If you want more than one tenant, put the opt-in in
spec.extraEnv, which the operator appends to every container it renders:
SIGLAKE_OIDC_TENANT_CLAIMtakes the tenant from a verified JSON Web Token (JWT) claim. SetSIGLAKE_OIDC_ISSUERandSIGLAKE_OIDC_AUDIENCEalongside it. The ingester refuses to start with only one of the two, and ignores the claim with neither.SIGLAKE_TRUST_SCOPE_HEADER=1routes onX-Scope-OrgIDinstead. Set it only behind a gateway that supplies the header itself and strips the client's.
Ingest tenant selection has the claim
validation rules, the refusal counter and the gateway requirement. Because
spec.extraEnv reaches the query containers too, the same three OIDC variables
bind query routing to the verified claim. The query tier never routes on
X-Scope-OrgID, so SIGLAKE_TRUST_SCOPE_HEADER changes nothing there.
authTokensSecretRef is a reference, not a value. The operator surfaces it to
the ingester as SIGLAKE_AUTH_TOKENS through valueFrom and never reads the
token bytes. Omit it and ingest is open, which is only safe inside a trusted
network.
retention.queryAuditRotateIntervalDays renders a CronJob that runs
siglake audit-rotate on that cadence.
autoscaling.query is an ordinary range. The operator scales query on
in-flight queries per pod, clamped to [min, max] with the same 10% deadband.
It renders --query-peer-discovery-srv pointing at the headless Service's
_http._tcp record at every replica count, so a replica the decision adds
receives shard work once it is Ready: no rollout, and no ceiling tied to the
rendered count. The earlier max == min refusal
(QueryAutoscalingRangeUnsupported) is gone. The packaged default is still
max: 1, so raise it deliberately.
Raising it above 1 requires a shared batch-job store. The operator renders
SIGLAKE_JOBS_POSTGRES_URI from spec.catalogUri when that URI is Postgres,
and blank otherwise, and the last matching spec.extraEnv entry overrides
either. A blank effective value selects the per-pod in-memory store, which is
correct only at one pod, so the operator refuses that spec with
QueryJobsStoreRequiredForScaleOut. Three
configurations pass: a Postgres spec.catalogUri with no override, a non-blank
SIGLAKE_JOBS_POSTGRES_URI in spec.extraEnv, and max: 1 with the in-memory
store. The last one loses every in-flight job when the pod restarts.
When the load signal is unusable, because Prometheus is down or the scrape came
back empty, all three tiers hold their current size rather than scaling down
blind. Each one is still clamped into its [min, max] range, so a tier below
its floor converges up to min.
Per-tier resources¶
spec.resources.{ingester,compactor,query} overrides each data-plane tier's
container requests and limits. Every tier starts from the operator's packaged
defaults, which mirror the chart: 200m CPU and 256Mi memory requested, a 2-CPU
limit, a 1Gi memory limit on ingester and compactor, and on query a 4Gi memory
limit with a 10Gi ephemeral-storage request and a 12Gi limit. Values you supply
merge over those defaults field by field:
apiVersion: siglake.limnion.ai/v1alpha1
kind: SiglakeCluster
metadata:
name: prod
spec:
resources:
query:
limits:
memory: 8Gi # CPU limit and all requests keep their packaged values
The query server derives its read caches and its shared memory pool from the
container memory limit, so treat spec.resources.query.limits.memory as a
sizing budget. See
the query memory guidance. The operator also
renders query's 10Gi spill emptyDir. If you override
query.limits.ephemeral-storage, keep it above that size limit, in the order
Query spill and ephemeral
storage describes.
If spec.extraEnv raises SIGLAKE_COMPACTOR_BIN_CONCURRENCY above one, the
tier's effective memory limit must hold 1Gi plus 4Gi per bin. The operator
refuses the spec with InvalidSpec when it cannot.
Schema migrations¶
Raise the monotonic spec.schemaVersion counter to run a one-shot, additive
schema-migration Job before the workloads roll out. The Job runs siglake
migrate-schema --all-tables --all-namespaces. Leaving the field unset, or
setting it to 0, renders no Job. The completed version appears in status. On
an upgrade, the operator holds the workload rollout while the Job is running or
after it fails.
The counter is a trigger, not a target. The Job migrates the warehouse to
whatever schema the binary in spec.image declares, so raising the counter
means "run a migration", not "reach version N".
The Job's name is <cluster>-migrate-schema-v<N>-<digest>, where the digest
covers everything that shapes the pod template: spec.image,
spec.warehouseUrl, spec.catalogUri, spec.awsRegion and spec.extraEnv.
A Job's spec.template is immutable, so a name keyed on the counter alone
would fail with a 422 on an image bump, and it would leave the migration
unreachable at that counter value.
Reverting the image¶
Put the old value back in spec.image and leave spec.schemaVersion where it
is. The digest covers the image, so the revert renders a new Job name: the
operator applies and waits on one more migration Job, this one running the old
binary. Against an already-widened table that run adds nothing and exits 0, the
hold clears, and the tiers roll back to the old image.
The migration is not reversed. The added columns stay, the reverted binary writes them as nulls, and rows that already carry values keep them. The table's recorded schema version also stays where the wider binary left it, because the migration stamps the higher of the recorded and declared versions. An image built before that behavior shipped (anything pre-0.1.0) stamps its own lower constant instead, which changes no column and no write decision, and rolling forward restamps it.
Do not lower spec.schemaVersion to match the old image. A lower value is
still a trigger, so it renders yet another Job and leaves the reported version
naming a schema the warehouse never returned to. To ask the warehouse instead,
run:
That prints each events table's recorded version, and the version the binary
you ran declares when the two differ.
The limits are the same as under Helm: additive column changes only, no path
across the pre-0.1.0 nanosecond timestamp contract, and no revert qualified
against an actual older image. See What a rollback
does.
Migrate a Helm release to the operator¶
Migration is manual, and there is no zero-downtime path. The --adopt-values
preflight is offline: it reads a values file, prints a SiglakeCluster, a
findings list and a runbook, and never contacts a cluster. The ownership handover that follows is
yours to run, it has no live-cluster test coverage, and doing it wrong leaves
Helm and the operator fighting over the same objects. If that is not a trade
you want, keep Helm as the controller.
Before you start: the operator is deployed, the CRD is installed, kubectl is
1.22 or later, the first release whose kubectl patch reads a patch from a
file, Python 3 is on the path for steps 7 and 8, and you have a
maintenance window. The window is there because the
handover has never been exercised against a live cluster, not because any step
is known to restart a pod.
- Run the whole procedure in one shell. Set the release, its namespace, and a
new directory path for the adoption artifacts. Put the directory on storage
you control: the effective values and release records can contain
credentials. The block creates it with mode
0700and stops if the path already exists. Use a new path instead of replacing an earlier backup.
Create the directory before you export any values:
- Export the release's effective values, not only your override file. The export is a values file, not a copy of the release. Helm's own record of the release lives in Secrets, and step 7 saves those.
(
set -eC
umask 077
helm get values "$RELEASE" -n "$NAMESPACE" --all -o yaml \
> "$BACKUP_DIR/effective-values.yaml"
)
- Generate the custom resource and the preflight report. Name the cluster after the release so names and selectors coincide, and pass the catalog URI: the chart builds it from a Postgres Secret, so the preflight cannot infer it.
(
set -eC
umask 077
siglake-operator --adopt-values "$BACKUP_DIR/effective-values.yaml" \
--adopt-cluster-name "$RELEASE" \
--adopt-catalog-uri postgres://siglake:secret@catalog.internal:5432/siglake \
> "$BACKUP_DIR/cluster.yaml"
)
The saved report contains the custom resource first, then the findings,
then a handover runbook that mixes comments with commands. The report is
therefore not a file kubectl can read: step 8 extracts the resource from
it, and step 11 applies what step 8 wrote.
Each tier's replicas becomes a fixed band (min equals max), so the
handover changes no pod count. Each <tier>.resources block carries onto
spec.resources; an omitted block takes the operator's packaged defaults,
which are the chart's.
- Clear every finding. A finding names a value the custom resource cannot
carry, and each one is a blocker. Findings cover an empty
image.tag, a missings3.bucket, a non-AWSs3.endpoint, atenant.namespaceother thansiglake, aningester.auth.secretKeywith noexistingSecret, per-tierextraEnvfolded cluster-wide, a detection tier still enabled in values, and each chart opt-out the operator switches back on:compactor.deleteTasks,compactor.invertedIndex.enabled,wal.mirror.enabledandquery.jobs.persistent. Fix each one by adding the equivalentspec.extraEnventry, or keep Helm.
compactor.indexRebuild: true is reported for the opposite reason. The
operator renders SIGLAKE_INDEX_REBUILD=0, so the adopted compactor stops
registering Puffin sidecars for rewrite output. Keep the opt-in with
SIGLAKE_INDEX_REBUILD=1 in spec.extraEnv. Indexes already registered
stay readable either way.
The WAL claim is one of those findings, and the one that costs data if you
miss it. The preflight carries wal.size, wal.accessMode and
wal.storageClassName onto spec.storage, because the operator applies a
claim of the same name and the API server rejects a shrink or a
storage-class change. The claim name is separate: the chart mounts
wal.existingClaim when you set it, and the operator always renders and
mounts <cluster>-wal. When the two differ, the finding begins
wal.existingClaim is <name> and names both claims, and says the adopted
ingester and compactor would start against a different, probably empty WAL
while every segment on the old claim that has not reached Iceberg stays
there unread. An absent, empty or already-matching existingClaim is
silent, because all three are the same volume. The preflight only reports
the mismatch: it moves no data and renames no volume, so drain the old claim
to empty before the handover, or keep Helm.
-
Resolve a drain-mode disagreement. Helm selects the drain with
compactor.catalogClaim.enabledand always usesRecreate. The generated custom resource selects it from the fixed compactor policy synthesised fromcompactor.replicas. When the two disagree, the finding beginscompactor.catalogClaim.enabled is <value>and says the synthesised policy selects the other mode. Align the chart with that mode and finish its drain, or setspec.autoscaling.compactor.maxabove1to retain catalog claims and accept a scale-out-capable policy. -
In the maintenance window, confirm that the objects the operator expects already exist under the release name.
kubectl get -n "$NAMESPACE" \
"deploy/$RELEASE-ingester" "deploy/$RELEASE-compactor" "sts/$RELEASE-query"
- Back up the release records and the current ownership metadata. Step 10
deletes Helm's record of the release, and the two abort paths below read
this backup. Every file has mode
0600. The block refuses existing files and symlinks instead of replacing them. It exits non-zero if either read fails or if the label selector matches no Secret, so a failed backup stops the sequence before step 9 changes anything.
(
set -eC
umask 077
if [ ! -d "$BACKUP_DIR" ] || [ -L "$BACKUP_DIR" ]; then
echo "unsafe adoption artifact directory: $BACKUP_DIR" >&2
exit 1
fi
chmod 700 "$BACKUP_DIR"
kubectl get secret -n "$NAMESPACE" -l "owner=helm,name=$RELEASE" -o json \
> "$BACKUP_DIR/release-secrets.json"
kubectl get -n "$NAMESPACE" -o json \
"deploy/$RELEASE-ingester" "deploy/$RELEASE-compactor" "sts/$RELEASE-query" \
> "$BACKUP_DIR/workloads.json"
python3 - "$BACKUP_DIR" "$RELEASE" <<'PY'
import json
import os
import sys
backup, release = sys.argv[1:3]
workload_names = (
f"{release}-ingester",
f"{release}-compactor",
f"{release}-query",
)
OWNERSHIP = (
("annotations", "meta.helm.sh/release-name"),
("annotations", "meta.helm.sh/release-namespace"),
("labels", "app.kubernetes.io/managed-by"),
)
ASSIGNED = (
"creationTimestamp",
"generation",
"managedFields",
"resourceVersion",
"selfLink",
"uid",
)
def load(name):
with open(os.path.join(backup, name)) as handle:
return json.load(handle)
def write(name, document):
path = os.path.join(backup, name)
handle = os.open(path, os.O_WRONLY | os.O_CREAT | os.O_EXCL, 0o600)
with os.fdopen(handle, "w") as out:
json.dump(document, out, indent=2)
out.write("\n")
records = load("release-secrets.json").get("items") or []
if not records:
sys.exit(f"no Helm release records for {release}; nothing to back up")
workloads = {
item["metadata"]["name"]: item["metadata"]
for item in load("workloads.json").get("items") or []
}
absent = [name for name in workload_names if name not in workloads]
if absent:
sys.exit("no ownership metadata for: " + ", ".join(absent))
patches = {}
for name in workload_names:
metadata = workloads[name]
patch = {}
for section, key in OWNERSHIP:
current = metadata.get(section) or {}
patch.setdefault(section, {})[key] = current.get(key)
patches[name] = {"metadata": patch}
for record in records:
for field in ASSIGNED:
record["metadata"].pop(field, None)
write("release-restore.json", {"apiVersion": "v1", "kind": "List", "items": records})
for name, patch in patches.items():
write(f"ownership-{name}.json", patch)
print(f"saved {len(records)} release records and {len(patches)} ownership patches")
PY
)
release-restore.json is the same Secrets without the fields the API server
assigns, among them resourceVersion and uid. A create that carries a
resourceVersion is rejected, so release-secrets.json is the record to
keep and the stripped copy is the one to apply. Each
ownership-<name>.json is a merge patch that puts one
workload's two Helm annotations and its app.kubernetes.io/managed-by label
back exactly as they were. An annotation or label that is absent today is
recorded as null, which the merge patch removes on restore rather than
inventing a value.
- Accept the saved manifest for this namespace. Run this after the edits in
steps 4 and 5, because it reads the file as it stands. The block pulls the
custom resource out of the report, stops unless the resource is a
SiglakeClusternamed$RELEASE, stops if it carries a namespace other than$NAMESPACE, and writes the accepted copy as$BACKUP_DIR/cluster-accepted.yamlwith mode0600. A failure here stops the sequence with the annotations and the release Secrets untouched.
(
set -eC
umask 077
python3 - "$BACKUP_DIR" "$RELEASE" "$NAMESPACE" <<'PY'
import os
import re
import sys
backup, release, namespace = sys.argv[1:4]
report = os.path.join(backup, "cluster.yaml")
with open(report) as handle:
lines = handle.read().splitlines()
manifest, metadata, section = [], {}, None
for line in lines:
if line.startswith("# preflight"):
break
if line.startswith("#"):
continue
manifest.append(line)
if not line[:1].isspace():
section = line.split(":", 1)[0]
continue
if section == "metadata":
field = re.match(r" (\w+): *(\S.*)$", line)
if field:
metadata[field.group(1)] = field.group(2).strip("\"'")
while manifest and not manifest[-1].strip():
manifest.pop()
kind = next(
(line.split(":", 1)[1].strip() for line in manifest if line.startswith("kind:")),
"",
)
if kind != "SiglakeCluster":
sys.exit(f"{report} holds kind {kind or '(none)'}, not SiglakeCluster")
name = metadata.get("name")
if name != release:
sys.exit(f"{report} names {name or '(none)'}, not the release {release}")
saved = metadata.get("namespace")
if saved is not None and saved != namespace:
sys.exit(f"{report} is for namespace {saved}, not {namespace}")
path = os.path.join(backup, "cluster-accepted.yaml")
handle = os.open(path, os.O_WRONLY | os.O_CREAT | os.O_EXCL, 0o600)
with os.fdopen(handle, "w") as out:
out.write("\n".join(manifest) + "\n")
print(f"accepted {kind}/{name} for namespace {namespace}")
PY
)
The preflight names the cluster and leaves metadata.namespace unset, so
the saved resource carries no namespace unless you added one. An absent
namespace is accepted, and step 11 passes $NAMESPACE on the command line:
that flag is the only place the namespace is decided. A namespace that is
present and different is a conflict. Applying it would create the cluster
beside the workloads whose Helm ownership steps 9 and 10 remove, and nothing
would reconcile them.
- Flip ownership. Helm forgets the objects without deleting them. This is the
first step that changes the cluster, so the block stops unless steps 7 and 8
left every file the abort paths and step 11 read. A failed annotation
removal stops it before it changes a label. It writes
ownership-flippedonly when both commands succeeded, and step 10 refuses to run without that file.
(
set -e
umask 077
for artifact in "release-restore.json" "cluster-accepted.yaml" \
"ownership-$RELEASE-ingester.json" "ownership-$RELEASE-compactor.json" \
"ownership-$RELEASE-query.json"; do
if [ ! -f "$BACKUP_DIR/$artifact" ]; then
echo "missing $artifact: finish steps 7 and 8 first" >&2
exit 1
fi
done
kubectl annotate -n "$NAMESPACE" \
"deploy/$RELEASE-ingester" "deploy/$RELEASE-compactor" "sts/$RELEASE-query" \
meta.helm.sh/release-name- meta.helm.sh/release-namespace- || {
status=$?
echo "annotation removal failed: Abort before the release records are deleted" >&2
exit "$status"
}
kubectl label -n "$NAMESPACE" \
"deploy/$RELEASE-ingester" "deploy/$RELEASE-compactor" "sts/$RELEASE-query" \
app.kubernetes.io/managed-by=siglake-operator --overwrite || {
status=$?
echo "label change failed: Abort before the release records are deleted" >&2
exit "$status"
}
: > "$BACKUP_DIR/ownership-flipped"
)
Either command can change some of the three objects and fail on the rest, so the block names an abort path instead of undoing what it did. The pre-deletion abort patches all three objects back to the metadata step 7 recorded, whether or not the flip reached them.
-
Delete the Helm release Secrets, which are Helm's record of the release and its revision history. The block stops unless step 9 wrote
ownership-flipped, and it writesrecords-deletedonly when the deletion succeeded.( set -e umask 077 if [ ! -f "$BACKUP_DIR/ownership-flipped" ]; then echo "step 9 did not finish: do not delete Helm's records" >&2 exit 1 fi kubectl delete secret -n "$NAMESPACE" -l "owner=helm,name=$RELEASE" || { status=$? echo "deletion failed: Abort after the release records are deleted" >&2 exit "$status" } : > "$BACKUP_DIR/records-deleted" )A deletion that removed some Secrets and failed on the rest looks from outside like one that removed none, so a failure here points at the post-deletion abort. That block applies
release-restore.json, which recreates the deleted records and rewrites any survivor with the same content. -
Apply the manifest step 8 accepted, in the release's namespace. When the specs match, the first reconcile is a no-op apply. The block stops unless step 10 wrote
records-deleted.( set -e if [ ! -f "$BACKUP_DIR/records-deleted" ]; then echo "step 10 did not finish: do not apply the custom resource" >&2 exit 1 fi kubectl apply -n "$NAMESPACE" -f "$BACKUP_DIR/cluster-accepted.yaml" || { status=$? echo "apply failed: Abort after the release records are deleted" >&2 exit "$status" } )If the apply fails, ask the API server whether the resource exists before you decide:
If it reports
NotFound, nothing is reconciling the workloads and the post-deletion abort still returns them to Helm. -
Verify. Status reaches
Ready=Truewith no pod restarts, and the replica counts in status match what Helm was running.
Abort before the release records are deleted¶
Nothing in steps 1 through 8 changes the cluster. After the ownership flip in step 9 and before the deletion in step 10, the workloads keep running with no controller and Helm's records are still in the namespace. A step 9 that stopped part way lands here too. Put the saved ownership metadata back, then confirm Helm sees the release again. Like the rest of the handover, these commands have no live-cluster test coverage.
(
set -e
for workload in "deploy/$RELEASE-ingester" "deploy/$RELEASE-compactor" \
"sts/$RELEASE-query"; do
kubectl patch -n "$NAMESPACE" "$workload" --type=merge \
--patch-file "$BACKUP_DIR/ownership-$(basename "$workload").json"
done
helm status "$RELEASE" -n "$NAMESPACE"
helm history "$RELEASE" -n "$NAMESPACE"
)
Abort after the release records are deleted¶
After step 10, restoring the annotations and labels alone leaves Helm with no
release to manage: helm status reports that the release is not found. Apply
the saved release records first, then the ownership patches. A failed apply
stops the block before it patches anything. A step 10 that deleted part of the
list, and a step 11 that failed before the custom resource existed, both end
here.
(
set -e
kubectl apply -n "$NAMESPACE" -f "$BACKUP_DIR/release-restore.json"
for workload in "deploy/$RELEASE-ingester" "deploy/$RELEASE-compactor" \
"sts/$RELEASE-query"; do
kubectl patch -n "$NAMESPACE" "$workload" --type=merge \
--patch-file "$BACKUP_DIR/ownership-$(basename "$workload").json"
done
helm status "$RELEASE" -n "$NAMESPACE"
helm history "$RELEASE" -n "$NAMESPACE"
)
Either block ends with the two commands that prove the recovery: helm status
reports the release with its last revision, and helm history lists every
revision the backup held. If helm history lists fewer revisions than the
release had, the backup missed records, and Helm can upgrade the release but
cannot roll it back past what is listed.
Once you apply the custom resource in step 11, this procedure stops making promises. The operator owns the objects from that point, and returning them to Helm is not a path this page covers or anyone has tested.
Design record: docs/DESIGN_operator_adoption.md.
Status¶
The operator reports:
observedGeneration, which tells you whether it has caught up with your last edit;- schema versions for the managed tables;
- managed replica counts (
siglake_operator_managed_replicas); - Kubernetes conditions.
Readyreports whether the managed workloads reached their desired replica counts,Progressingreports a replica-count change,InvalidSpecreports that the requested specification cannot be honoured, andDrainModeCompatible=Falsereports that an existing workload template uses a different drain from the compactor policy.
Invalid specifications¶
When the operator cannot honour a SiglakeCluster specification it sets
InvalidSpec=True with a specific reason, and Ready=False with reason
InvalidSpec. Validation runs before any child resource is rendered or
applied, so existing managed resources are left alone and a new invalid cluster
gets nothing. The operator retries after five minutes to avoid a hot loop, and
an edit to the SiglakeCluster wakes reconciliation immediately.
InvalidSpec reason |
Repair |
|---|---|
ImageRequired |
Set spec.image to the container image to run. |
WarehouseUrlRequired |
Set spec.warehouseUrl to the Iceberg warehouse URL. |
CatalogUriRequired |
Set spec.catalogUri to the Iceberg catalog URI. |
AutoscalingRangeInvalid |
For every tier, set min to at least 1 and max to at least min. |
AutoscalingZeroFloorUnsupported |
Set every tier's min to 1 or more. A tier stopped at zero exports no scaling signal, so nothing can request it back, and its missing reading also holds the other two tiers at their current size. |
AutoscalingTargetNotPositive |
Give every tier a finite target greater than zero. |
AutoscalingEwmaHalfLifeInvalid |
Set spec.autoscaling.ewmaHalfLifeSecs to a finite, non-negative number; zero disables smoothing. |
RetentionIntervalUnsupported |
Set spec.retention.queryAuditRotateIntervalDays to 1 (daily), 7 (weekly) or 30 (monthly), or omit it to disable rotation. |
CompactorBinConcurrencyInvalid |
Set SIGLAKE_COMPACTOR_BIN_CONCURRENCY in spec.extraEnv to a positive integer, or remove the override. |
CompactorBinConcurrencyExceedsMemory |
Reduce bin concurrency, or raise the effective compactor memory limit to at least 1Gi plus 4Gi per concurrent bin. See Per-tier resources. |
QueryJobsStoreRequiredForScaleOut |
Give the query tier a shared batch-job store, or hold spec.autoscaling.query.max at 1. Set a non-blank SIGLAKE_JOBS_POSTGRES_URI in spec.extraEnv, or use a Postgres spec.catalogUri and drop any override of that variable. |
WalMirrorRequiredForCompactorScaleOut |
Set a non-blank SIGLAKE_WAL_MIRROR_PREFIX in spec.extraEnv, or hold spec.autoscaling.compactor.max at 1. |
Handover between drain modes¶
If an existing ingester or compactor Deployment uses the other drain mode, the
operator sets DrainModeCompatible=False and Ready=False. Both conditions
use reason DrainModeHandoverRequired. The operator patches status only, then
retries after five minutes. It does not query Prometheus or apply child
resources while the mismatch remains.
If you did not intend to change drain mode, restore the previous
spec.autoscaling.compactor.max. To complete the handover:
- Stop ingestion. Confirm that clients and exporters no longer send requests to the ingester.
- Let the current drain finish. Watch
siglake_compactor_sealed_pendingfall to zero, then query the affected tables for the expected rows. A zero gauge alone does not prove that every mirror object has a catalog row. - Account for retained local WAL segments, mirror objects and catalog rows. Preserve anything your recovery plan still needs. The operator does not convert or delete any of them during the handover.
- Delete the ingester and compactor Deployments so the operator can recreate both templates in the selected mode:
- Wait for the recreated workloads. Confirm that
Ready=Trueand thatDrainModeHandoverRequiredis absent from status before you resume ingestion.
Operator metrics: siglake_operator_reconciles_total,
siglake_operator_reconcile_duration_seconds and
siglake_operator_prom_query_errors_total.