Security¶
Use this guide to require authentication, bind requests to tenants, protect traffic and remove sensitive data. Start with authentication and network access. Then configure audit retention and test the deletion procedure before the cluster takes production traffic.
Authentication¶
Choose one of three modes. Shared clusters should use OpenID Connect (OIDC) with a tenant claim.
Open (the default)¶
With no authentication settings, both servers accept requests from anyone who can reach them. The query server logs a warning at startup; the ingester does not. Use this default only inside a trusted network boundary.
Bearer tokens¶
siglake ingest-server --auth-tokens tok1,tok2 # SIGLAKE_AUTH_TOKENS
siglake-query-server --tokens tok1,tok2 # SIGLAKE_QUERY_TOKENS
Every request must carry Authorization: Bearer <token> matching one of the
listed values. Bearer is the only accepted authorization scheme.
These credentials cover the whole cluster. Tokens do not identify a tenant or a caller. Ingest tenant routing follows the settings described below.
In Helm, source them from a Secret rather than inline values:
ingester:
auth:
existingSecret: siglake-ingester-tokens
secretKey: tokens
query:
tokens:
existingSecret: siglake-query-tokens
secretKey: tokens
See Query authentication for Helm's existing Secret, inline list and External Secrets Operator (ESO) sources for query bearer tokens.
If the Helm deployment also uses distributed query with more than one
replica, bearer-token auth requires
query.distributed.coordinatorToken; client bearer
tokens do not substitute for this peer credential.
OIDC¶
--oidc-issuer https://cognito-idp.us-east-1.amazonaws.com/us-east-1_ABC123
--oidc-audience my-client-id
--oidc-tenant-claim custom:tenant
The server discovers the issuer's signing material through
<issuer>/.well-known/openid-configuration. It caches that material and
verifies every bearer JSON Web Token (JWT). If you configure OIDC and static
bearer tokens, the server uses OIDC.
This is the only mode that gives you identity-bound tenancy. With
--oidc-tenant-claim, the query server routes each query to the namespace in
the verified claim, and the ingester takes the verified claim as the
authoritative tenant. Query OIDC also populates subject and email in the
audit table.
Configuring the claim makes it mandatory at both servers. A verified token is
refused if its claim is missing, blank, not a JSON string, longer than 128
characters or outside [A-Za-z0-9_-]. The query server returns 403 before
routing the request and does not fall back to the default namespace.
Siglake validates tenant identifiers without changing them. It refuses
acme.corp instead of changing it to acmecorp. This prevents two claims from
selecting the same namespace and prevents an invalid claim from selecting the
default namespace. If you leave --oidc-tenant-claim unset, every verified
caller reads the default namespace.
OIDC deployments have the same query.distributed.coordinatorToken
requirement when distributed query uses more than one replica.
With Helm, the corresponding settings are ingester.oidc.issuer,
ingester.oidc.audience, and ingester.oidc.tenantClaim (or the same keys
under query.oidc for the query tier).
Ingest tenant selection¶
Ingest is single-tenant by default. Every request routes to the default
tenant. An X-Scope-OrgID naming another tenant is refused with 403 over
HTTP or PermissionDenied over OpenTelemetry Protocol (OTLP)/gRPC. Naming
default has no effect.
Two Helm values turn on ingest tenant routing:
ingester.oidc.tenantClaimtakes the tenant from the verified JSON Web Token (JWT) claim. When the header is present, it must agree with the claim.ingester.trustScopeHeadertakes the client's word and routes byX-Scope-OrgID. Put a trusted gateway in front of the ingester to set the header and strip client-supplied values.
Use claim-based routing or separate ingesters for mutually untrusting tenants.
Configuring the claim is what makes it mandatory, on this boundary as on the
query boundary. A verified token whose claim is missing, blank, not a JSON
string, longer than 128 characters, or outside [A-Za-z0-9_-] is refused with
403 before the batch is routed. Token validity does not grant access to a
tenant. The value is validated without repair: acme.corp is a refusal
rather than a rewrite to acmecorp, so two claims cannot alias onto one
namespace and an unusable claim never falls back to the default tenant.
The server creates no WAL lane, metric label or Iceberg namespace for a refused
batch.
Refusals are counted by siglake_ingest_tenant_denied_total per reason.
Each label comes from one setting:
reason |
Refused by | What the client sent |
|---|---|---|
header_not_trusted |
Single-tenant routing, the shipped default | X-Scope-OrgID naming a tenant other than default, with no verified claim to route by |
claim_missing |
--oidc-tenant-claim / ingester.oidc.tenantClaim |
A valid token that carries no such claim |
claim_invalid |
--oidc-tenant-claim / ingester.oidc.tenantClaim |
A claim that is blank, not a JSON string, over 128 characters, or outside [A-Za-z0-9_-] |
header_mismatch |
--oidc-tenant-claim / ingester.oidc.tenantClaim |
An X-Scope-OrgID that contradicts the claim |
not_allowed |
--allowed-tenants / ingester.allowedTenants |
A tenant outside the allowed set, after routing resolved it |
at_capacity |
--max-tenants / ingester.maxTenants |
A tenant this ingester has not admitted before, arriving when the cap is full |
On the shipped default, header_not_trusted is the label you see. All six
series exist at zero from startup, so a client that loses its tenant routing
after an upgrade shows up on the first refused request rather than the second.
The chart's SiglakeTenantsDenied alert watches that counter; see
Monitoring.
If you set ingester.trustScopeHeader, bound client-controlled tenant creation
with --allowed-tenants / ingester.allowedTenants when the tenant set is
known, or --max-tenants / ingester.maxTenants as a backstop. The cap counts
the distinct tenants one ingester process has admitted, starts empty on
restart, and is held per pod, so it bounds a pod rather than the fleet; see
What the tenant cap counts.
--ingest-max-lanes / ingester.maxLanes separately caps distinct
(tenant, index) backpressure lanes and counts its refusals in
siglake_ingest_lane_refused_total. An allow-list refusal raises
SiglakeTenantsDenied on the not_allowed series above, and a cap refusal on
at_capacity. See
the
configuration reference for
the corresponding environment variables and defaults.
The query tier ignores X-Scope-OrgID. With --oidc-tenant-claim /
query.oidc.tenantClaim set, a verified token without a usable claim gets a
403 instead of the default namespace. query.allowedTenants can further
restrict usable claims. Its empty default accepts all of them, and it stays
independent from ingester.allowedTenants.
siglake_query_tenant_denied_total records the query-tier refusal:
reason |
Refused by | What the client sent |
|---|---|---|
claim_missing |
--oidc-tenant-claim / query.oidc.tenantClaim |
A valid token without the configured claim |
claim_invalid |
--oidc-tenant-claim / query.oidc.tenantClaim |
A blank, non-string, overlength or invalid claim |
not_allowed |
--allowed-tenants / query.allowedTenants |
A usable claim outside the query allow-list |
All three series exist at zero from startup. The chart's
SiglakeQueryTenantsDenied alert reports each refusal. If the claim setting is
empty, every verified caller reads the default namespace.
See Multi-tenancy for what is and isn't isolated.
TLS¶
Choose where to terminate Transport Layer Security (TLS).
The chart can terminate TLS at an Ingress. Its ingress.tls block protects
traffic that uses configured Ingress routes. It does not encrypt a listener
that you expose independently through a Service, LoadBalancer, or NodePort.
If a Service LoadBalancer or NodePort exposes the query API without an Ingress,
terminate TLS in siglake-query-server with rustls:
Or set the equivalent environment variables:
Both values are required. If you set only one, siglake-query-server exits
during startup. With neither value, the query API listener uses HTTP. See the
authentication configuration
for the environment-variable reference.
query.tls protects the query API listener. If distributed query is enabled,
the chart also sends coordinator requests to query peers over HTTPS. With
query.tls.enabled: false, query-peer traffic uses HTTP.
Query TLS does not protect the separate ingest HTTP and gRPC listeners. This includes OTLP/gRPC on port 4317, which uses plaintext h2c when you expose it directly. See Send over OTLP/gRPC.
A peer accepts a request on /api/v1/sql/shard as a shard request only when it
presents the shared query.distributed.coordinatorToken. Without that token,
the request is treated as an ordinary caller and normal bearer-token or OIDC
authentication applies.
NetworkPolicies restrict which workloads can reach a listener. They do not encrypt traffic. Put a TLS proxy, Ingress, or service mesh in front of ingest listeners that need transport encryption.
Object storage credentials¶
Use standard provider environment variables (AWS_ACCESS_KEY_ID,
AWS_SECRET_ACCESS_KEY, AWS_REGION, AWS_ENDPOINT_URL, and the GCP/Azure
equivalents). On Amazon Elastic Kubernetes Service (EKS), use IAM Roles for
Service Accounts (IRSA). Annotate the ServiceAccount with the role ARN and omit
static credentials.
Every role uses the deployment's credentials. There are no per-tenant credentials; tenant isolation is by namespace prefix within one bucket, not by IAM boundary.
Network policies¶
Set networkPolicy.enabled: true to render NetworkPolicies. The rules must
allow query pods to call one another for peer fan-out. They must also allow
each role to reach the shared WAL and catalog services it uses.
The rendered policy declares policyTypes: [Egress] and nothing else. It
writes no ingress rule, for the metrics ports (9100, 9101, 9105) or for any
other listener, so who may open a connection to a Siglake pod stays whatever
your cluster's default is. The metrics port checks no token, so restrict it
yourself if that default is open. A NetworkPolicy cannot select by
ServiceAccount; select the Prometheus pods by namespace and pod labels:
apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
name: siglake-metrics-from-prometheus
namespace: <siglake namespace>
spec:
podSelector:
matchLabels:
app.kubernetes.io/name: siglake
app.kubernetes.io/instance: <release>
policyTypes:
- Ingress
ingress:
- from:
- namespaceSelector:
matchLabels:
kubernetes.io/metadata.name: <prometheus namespace>
podSelector:
matchLabels:
app.kubernetes.io/name: prometheus
ports:
- port: 9100
protocol: TCP
- port: 9101
protocol: TCP
- port: 9105
protocol: TCP
An ingress policy that selects a pod denies every other inbound connection to it, so this one on its own cuts off ingest on 8088 and 4317 and queries on 8089. Allow those ports in the same policy or in a companion one that selects the same pods.
Secrets handling¶
- In Helm deployments, Postgres credentials come from a Kubernetes Secret you
provide (
postgres.existingSecret). Kubernetes assemblesSIGLAKE_CATALOG_URIin the container environment. - Auth tokens can come from a Secret (
ingester.auth.existingSecretandquery.tokens.existingSecret). - See Query authentication to materialize the query bearer token Secret from a remote store with External Secrets Operator.
- The operator references the auth-token Secret via
valueFromand never reads the bytes itself.
Protect catalog credentials
A Helm pod specification contains Secret references for the PG*
variables and an unexpanded URI for SIGLAKE_CATALOG_URI. Kubernetes
resolves both in the container environment. Reading the pod specification
does not reveal the password.
The operator takes a different path. It stores the full URI in plaintext
as spec.catalogUri, then copies the value into pod environments and
CronJob arguments. Permission to read the custom resource or its managed
workload specifications exposes the password.
Restrict permission to read Secrets, execute commands in pods, and create
workloads that reference Secrets. For operator deployments, also restrict
reads of SiglakeCluster resources and their managed workloads.
Audit logging¶
The query server submits completed queries to a bounded, best-effort writer for
the query_audit Iceberg table. A retained row stores the subject, email,
endpoint, SQL text, priority, duration, status, complexity, estimated bytes and
rows, truncation and any error.
The writer has two process-local charged-retention ceilings: 10,000 rows and 64 MiB across its submit channel, flush buffer and Arrow conversion working set. These ceilings limit charged audit work, not total process memory. If a row exceeds either ceiling, the server drops the whole row rather than truncating it. A full channel or stopped worker also drops the whole row.
Alert on increases in siglake_query_audit_dropped_total{reason}. The reason
label is oversized, row_limit, byte_limit, channel_full or
worker_shutdown when the worker refuses a row before storage. If a storage
append returns an error, the worker abandons the whole batch without retry. It
counts each row in siglake_query_audit_dropped_total{reason="append"} and
increments siglake_query_audit_failures_total{reason="append"} once for the
batch. Audit batch build failures increment the failure counter with
reason="build" without incrementing the drop counter.
Each Iceberg append has a 30-second service deadline. Set
SIGLAKE_QUERY_AUDIT_APPEND_DEADLINE_SECS to change it, or set it to 0 for an
unbounded wait. If the deadline expires, the worker abandons the batch without
retry, counts each row with drop reason append_deadline and increments the
matching failure reason once. The worker then drains retained rows behind the
abandoned batch. The append may have committed before the deadline, so retrying
it could duplicate audit rows.
The audit table contains query text
query holds the literal SQL, including any literal values in WHERE
clauses. These values may include usernames, IDs or other sensitive data
searched for. Treat query_audit as sensitive, restrict who can read it,
and set retention deliberately.
siglake audit-rotate --max-age-secs <seconds> expires old snapshots and keeps
every audit row. The age does not set a sensitive-data retention deadline. The
bare siglake audit-rotate command drops and recreates the table, deleting all
historical audit rows. See
Audit rotation.
Data deletion¶
POST /api/v1/delete-tasks creates GDPR-style predicate deletes, executed by
compactor sweeps that rewrite the affected files.
Configure and rehearse these five parts of deletion:
- Deletes are queued in a ledger and executed on a sweep. The compactor sweeps
by default: the chart ships
compactor.deleteTasks: true, and the binary sweeps unlessSIGLAKE_DELETE_TASKSis0,off,falseorno. SettingdeleteTasks: falserendersSIGLAKE_DELETE_TASKS=0, and under the operator that variable goes inspec.extraEnv. With the sweep off, tasks are still accepted and recorded, and onlysiglake delete-sweep --applyexecutes them. - Deleted rows persist in expired snapshots until orphan GC reclaims the
superseded files. For a compliance-grade deletion you must also run snapshot
expiry and
siglake gc-orphans --apply, and account for S3 bucket versioning if enabled. - That chain erases one table, not the deployment. Orphan GC is rooted at the
table location, so the WAL mirror prefix, its
_active/blobs, the WAL volume, replica buckets and catalog backups keep their own copies of the same events until their own retention rules expire them. Write those windows down with the deletion procedure: see Copies the delete chain does not reach. - A
failedtask is terminal. There is no retry API; recovery is resubmitting the same request fields under the same tenant identity, which answers with a newtask_idwhile the failed record and itserrorstay as they were. Keep both IDs in the deletion audit trail.failedalso does not prove the rewrite never committed, so checkrows_deletedon the new task before recording that the first attempt deleted nothing. - A
pendingtask can be permanently stranded. If an executor crashes after claiming a task but before writing itsrunningrecord, the task is leftpendingunder the executor's unreleased claim and never runs.GET /api/v1/delete-tasks/{id}shows you the claim. Apendingtask carries aclaimobject recording whether the claim was observed and, when its body decodes, which process took it and how long ago. But a stranded claim and one held by a rewrite still in flight look identical, so the observation names the holder rather than proving the task is abandoned. Corroborate with the executor-side signals and resubmit with a newtask_idas described in Execution runs by default. - A queued task is bound to the index incarnation it was accepted against. If
you drop an index and recreate it under the same id, the acknowledged request
goes
failedat execution instead of deleting from the replacement. See Delete tasks are bound to one index incarnation.
See Retention and deletes and
Recovering a failed
task.
Pod security¶
The chart exposes podSecurityContext and securityContext per role. The
containers do not require privileged mode or host mounts. Set a non-root user
and a read-only root filesystem if your policy requires them. The data plane
needs write access to the WAL mount.
Reporting vulnerabilities¶
See SECURITY.md in the source repository.