Skip to content

Security

Use this guide to require authentication, bind requests to tenants, protect traffic and remove sensitive data. Start with authentication and network access. Then configure audit retention and test the deletion procedure before the cluster takes production traffic.

Authentication

Choose one of three modes. Shared clusters should use OpenID Connect (OIDC) with a tenant claim.

Open (the default)

With no authentication settings, both servers accept requests from anyone who can reach them. The query server logs a warning at startup; the ingester does not. Use this default only inside a trusted network boundary.

Bearer tokens

siglake ingest-server --auth-tokens tok1,tok2      # SIGLAKE_AUTH_TOKENS
siglake-query-server  --tokens tok1,tok2           # SIGLAKE_QUERY_TOKENS

Every request must carry Authorization: Bearer <token> matching one of the listed values. Bearer is the only accepted authorization scheme.

These credentials cover the whole cluster. Tokens do not identify a tenant or a caller. Ingest tenant routing follows the settings described below.

In Helm, source them from a Secret rather than inline values:

ingester:
  auth:
    existingSecret: siglake-ingester-tokens
    secretKey: tokens
query:
  tokens:
    existingSecret: siglake-query-tokens
    secretKey: tokens

See Query authentication for Helm's existing Secret, inline list and External Secrets Operator (ESO) sources for query bearer tokens.

If the Helm deployment also uses distributed query with more than one replica, bearer-token auth requires query.distributed.coordinatorToken; client bearer tokens do not substitute for this peer credential.

OIDC

--oidc-issuer   https://cognito-idp.us-east-1.amazonaws.com/us-east-1_ABC123
--oidc-audience my-client-id
--oidc-tenant-claim custom:tenant

The server discovers the issuer's signing material through <issuer>/.well-known/openid-configuration. It caches that material and verifies every bearer JSON Web Token (JWT). If you configure OIDC and static bearer tokens, the server uses OIDC.

This is the only mode that gives you identity-bound tenancy. With --oidc-tenant-claim, the query server routes each query to the namespace in the verified claim, and the ingester takes the verified claim as the authoritative tenant. Query OIDC also populates subject and email in the audit table.

Configuring the claim makes it mandatory at both servers. A verified token is refused if its claim is missing, blank, not a JSON string, longer than 128 characters or outside [A-Za-z0-9_-]. The query server returns 403 before routing the request and does not fall back to the default namespace.

Siglake validates tenant identifiers without changing them. It refuses acme.corp instead of changing it to acmecorp. This prevents two claims from selecting the same namespace and prevents an invalid claim from selecting the default namespace. If you leave --oidc-tenant-claim unset, every verified caller reads the default namespace.

OIDC deployments have the same query.distributed.coordinatorToken requirement when distributed query uses more than one replica.

With Helm, the corresponding settings are ingester.oidc.issuer, ingester.oidc.audience, and ingester.oidc.tenantClaim (or the same keys under query.oidc for the query tier).

Ingest tenant selection

Ingest is single-tenant by default. Every request routes to the default tenant. An X-Scope-OrgID naming another tenant is refused with 403 over HTTP or PermissionDenied over OpenTelemetry Protocol (OTLP)/gRPC. Naming default has no effect.

Two Helm values turn on ingest tenant routing:

  • ingester.oidc.tenantClaim takes the tenant from the verified JSON Web Token (JWT) claim. When the header is present, it must agree with the claim.
  • ingester.trustScopeHeader takes the client's word and routes by X-Scope-OrgID. Put a trusted gateway in front of the ingester to set the header and strip client-supplied values.

Use claim-based routing or separate ingesters for mutually untrusting tenants.

Configuring the claim is what makes it mandatory, on this boundary as on the query boundary. A verified token whose claim is missing, blank, not a JSON string, longer than 128 characters, or outside [A-Za-z0-9_-] is refused with 403 before the batch is routed. Token validity does not grant access to a tenant. The value is validated without repair: acme.corp is a refusal rather than a rewrite to acmecorp, so two claims cannot alias onto one namespace and an unusable claim never falls back to the default tenant. The server creates no WAL lane, metric label or Iceberg namespace for a refused batch.

Refusals are counted by siglake_ingest_tenant_denied_total per reason. Each label comes from one setting:

reason Refused by What the client sent
header_not_trusted Single-tenant routing, the shipped default X-Scope-OrgID naming a tenant other than default, with no verified claim to route by
claim_missing --oidc-tenant-claim / ingester.oidc.tenantClaim A valid token that carries no such claim
claim_invalid --oidc-tenant-claim / ingester.oidc.tenantClaim A claim that is blank, not a JSON string, over 128 characters, or outside [A-Za-z0-9_-]
header_mismatch --oidc-tenant-claim / ingester.oidc.tenantClaim An X-Scope-OrgID that contradicts the claim
not_allowed --allowed-tenants / ingester.allowedTenants A tenant outside the allowed set, after routing resolved it
at_capacity --max-tenants / ingester.maxTenants A tenant this ingester has not admitted before, arriving when the cap is full

On the shipped default, header_not_trusted is the label you see. All six series exist at zero from startup, so a client that loses its tenant routing after an upgrade shows up on the first refused request rather than the second. The chart's SiglakeTenantsDenied alert watches that counter; see Monitoring.

If you set ingester.trustScopeHeader, bound client-controlled tenant creation with --allowed-tenants / ingester.allowedTenants when the tenant set is known, or --max-tenants / ingester.maxTenants as a backstop. The cap counts the distinct tenants one ingester process has admitted, starts empty on restart, and is held per pod, so it bounds a pod rather than the fleet; see What the tenant cap counts. --ingest-max-lanes / ingester.maxLanes separately caps distinct (tenant, index) backpressure lanes and counts its refusals in siglake_ingest_lane_refused_total. An allow-list refusal raises SiglakeTenantsDenied on the not_allowed series above, and a cap refusal on at_capacity. See the configuration reference for the corresponding environment variables and defaults.

The query tier ignores X-Scope-OrgID. With --oidc-tenant-claim / query.oidc.tenantClaim set, a verified token without a usable claim gets a 403 instead of the default namespace. query.allowedTenants can further restrict usable claims. Its empty default accepts all of them, and it stays independent from ingester.allowedTenants.

siglake_query_tenant_denied_total records the query-tier refusal:

reason Refused by What the client sent
claim_missing --oidc-tenant-claim / query.oidc.tenantClaim A valid token without the configured claim
claim_invalid --oidc-tenant-claim / query.oidc.tenantClaim A blank, non-string, overlength or invalid claim
not_allowed --allowed-tenants / query.allowedTenants A usable claim outside the query allow-list

All three series exist at zero from startup. The chart's SiglakeQueryTenantsDenied alert reports each refusal. If the claim setting is empty, every verified caller reads the default namespace.

See Multi-tenancy for what is and isn't isolated.

TLS

Choose where to terminate Transport Layer Security (TLS).

The chart can terminate TLS at an Ingress. Its ingress.tls block protects traffic that uses configured Ingress routes. It does not encrypt a listener that you expose independently through a Service, LoadBalancer, or NodePort.

If a Service LoadBalancer or NodePort exposes the query API without an Ingress, terminate TLS in siglake-query-server with rustls:

siglake-query-server \
  --tls-cert /etc/siglake/tls/tls.crt \
  --tls-key /etc/siglake/tls/tls.key

Or set the equivalent environment variables:

SIGLAKE_QUERY_TLS_CERT=/etc/siglake/tls/tls.crt \
SIGLAKE_QUERY_TLS_KEY=/etc/siglake/tls/tls.key

Both values are required. If you set only one, siglake-query-server exits during startup. With neither value, the query API listener uses HTTP. See the authentication configuration for the environment-variable reference.

query:
  tls:
    enabled: true
    existingSecret: siglake-query-tls    # kubernetes.io/tls

query.tls protects the query API listener. If distributed query is enabled, the chart also sends coordinator requests to query peers over HTTPS. With query.tls.enabled: false, query-peer traffic uses HTTP.

Query TLS does not protect the separate ingest HTTP and gRPC listeners. This includes OTLP/gRPC on port 4317, which uses plaintext h2c when you expose it directly. See Send over OTLP/gRPC.

A peer accepts a request on /api/v1/sql/shard as a shard request only when it presents the shared query.distributed.coordinatorToken. Without that token, the request is treated as an ordinary caller and normal bearer-token or OIDC authentication applies.

NetworkPolicies restrict which workloads can reach a listener. They do not encrypt traffic. Put a TLS proxy, Ingress, or service mesh in front of ingest listeners that need transport encryption.

Object storage credentials

Use standard provider environment variables (AWS_ACCESS_KEY_ID, AWS_SECRET_ACCESS_KEY, AWS_REGION, AWS_ENDPOINT_URL, and the GCP/Azure equivalents). On Amazon Elastic Kubernetes Service (EKS), use IAM Roles for Service Accounts (IRSA). Annotate the ServiceAccount with the role ARN and omit static credentials.

Every role uses the deployment's credentials. There are no per-tenant credentials; tenant isolation is by namespace prefix within one bucket, not by IAM boundary.

Network policies

Set networkPolicy.enabled: true to render NetworkPolicies. The rules must allow query pods to call one another for peer fan-out. They must also allow each role to reach the shared WAL and catalog services it uses.

The rendered policy declares policyTypes: [Egress] and nothing else. It writes no ingress rule, for the metrics ports (9100, 9101, 9105) or for any other listener, so who may open a connection to a Siglake pod stays whatever your cluster's default is. The metrics port checks no token, so restrict it yourself if that default is open. A NetworkPolicy cannot select by ServiceAccount; select the Prometheus pods by namespace and pod labels:

apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
  name: siglake-metrics-from-prometheus
  namespace: <siglake namespace>
spec:
  podSelector:
    matchLabels:
      app.kubernetes.io/name: siglake
      app.kubernetes.io/instance: <release>
  policyTypes:
    - Ingress
  ingress:
    - from:
        - namespaceSelector:
            matchLabels:
              kubernetes.io/metadata.name: <prometheus namespace>
          podSelector:
            matchLabels:
              app.kubernetes.io/name: prometheus
      ports:
        - port: 9100
          protocol: TCP
        - port: 9101
          protocol: TCP
        - port: 9105
          protocol: TCP

An ingress policy that selects a pod denies every other inbound connection to it, so this one on its own cuts off ingest on 8088 and 4317 and queries on 8089. Allow those ports in the same policy or in a companion one that selects the same pods.

Secrets handling

  • In Helm deployments, Postgres credentials come from a Kubernetes Secret you provide (postgres.existingSecret). Kubernetes assembles SIGLAKE_CATALOG_URI in the container environment.
  • Auth tokens can come from a Secret (ingester.auth.existingSecret and query.tokens.existingSecret).
  • See Query authentication to materialize the query bearer token Secret from a remote store with External Secrets Operator.
  • The operator references the auth-token Secret via valueFrom and never reads the bytes itself.

Protect catalog credentials

A Helm pod specification contains Secret references for the PG* variables and an unexpanded URI for SIGLAKE_CATALOG_URI. Kubernetes resolves both in the container environment. Reading the pod specification does not reveal the password.

The operator takes a different path. It stores the full URI in plaintext as spec.catalogUri, then copies the value into pod environments and CronJob arguments. Permission to read the custom resource or its managed workload specifications exposes the password.

Restrict permission to read Secrets, execute commands in pods, and create workloads that reference Secrets. For operator deployments, also restrict reads of SiglakeCluster resources and their managed workloads.

Audit logging

The query server submits completed queries to a bounded, best-effort writer for the query_audit Iceberg table. A retained row stores the subject, email, endpoint, SQL text, priority, duration, status, complexity, estimated bytes and rows, truncation and any error.

The writer has two process-local charged-retention ceilings: 10,000 rows and 64 MiB across its submit channel, flush buffer and Arrow conversion working set. These ceilings limit charged audit work, not total process memory. If a row exceeds either ceiling, the server drops the whole row rather than truncating it. A full channel or stopped worker also drops the whole row.

Alert on increases in siglake_query_audit_dropped_total{reason}. The reason label is oversized, row_limit, byte_limit, channel_full or worker_shutdown when the worker refuses a row before storage. If a storage append returns an error, the worker abandons the whole batch without retry. It counts each row in siglake_query_audit_dropped_total{reason="append"} and increments siglake_query_audit_failures_total{reason="append"} once for the batch. Audit batch build failures increment the failure counter with reason="build" without incrementing the drop counter.

Each Iceberg append has a 30-second service deadline. Set SIGLAKE_QUERY_AUDIT_APPEND_DEADLINE_SECS to change it, or set it to 0 for an unbounded wait. If the deadline expires, the worker abandons the batch without retry, counts each row with drop reason append_deadline and increments the matching failure reason once. The worker then drains retained rows behind the abandoned batch. The append may have committed before the deadline, so retrying it could duplicate audit rows.

The audit table contains query text

query holds the literal SQL, including any literal values in WHERE clauses. These values may include usernames, IDs or other sensitive data searched for. Treat query_audit as sensitive, restrict who can read it, and set retention deliberately.

siglake audit-rotate --max-age-secs <seconds> expires old snapshots and keeps every audit row. The age does not set a sensitive-data retention deadline. The bare siglake audit-rotate command drops and recreates the table, deleting all historical audit rows. See Audit rotation.

Data deletion

POST /api/v1/delete-tasks creates GDPR-style predicate deletes, executed by compactor sweeps that rewrite the affected files.

Configure and rehearse these five parts of deletion:

  • Deletes are queued in a ledger and executed on a sweep. The compactor sweeps by default: the chart ships compactor.deleteTasks: true, and the binary sweeps unless SIGLAKE_DELETE_TASKS is 0, off, false or no. Setting deleteTasks: false renders SIGLAKE_DELETE_TASKS=0, and under the operator that variable goes in spec.extraEnv. With the sweep off, tasks are still accepted and recorded, and only siglake delete-sweep --apply executes them.
  • Deleted rows persist in expired snapshots until orphan GC reclaims the superseded files. For a compliance-grade deletion you must also run snapshot expiry and siglake gc-orphans --apply, and account for S3 bucket versioning if enabled.
  • That chain erases one table, not the deployment. Orphan GC is rooted at the table location, so the WAL mirror prefix, its _active/ blobs, the WAL volume, replica buckets and catalog backups keep their own copies of the same events until their own retention rules expire them. Write those windows down with the deletion procedure: see Copies the delete chain does not reach.
  • A failed task is terminal. There is no retry API; recovery is resubmitting the same request fields under the same tenant identity, which answers with a new task_id while the failed record and its error stay as they were. Keep both IDs in the deletion audit trail. failed also does not prove the rewrite never committed, so check rows_deleted on the new task before recording that the first attempt deleted nothing.
  • A pending task can be permanently stranded. If an executor crashes after claiming a task but before writing its running record, the task is left pending under the executor's unreleased claim and never runs. GET /api/v1/delete-tasks/{id} shows you the claim. A pending task carries a claim object recording whether the claim was observed and, when its body decodes, which process took it and how long ago. But a stranded claim and one held by a rewrite still in flight look identical, so the observation names the holder rather than proving the task is abandoned. Corroborate with the executor-side signals and resubmit with a new task_id as described in Execution runs by default.
  • A queued task is bound to the index incarnation it was accepted against. If you drop an index and recreate it under the same id, the acknowledged request goes failed at execution instead of deleting from the replacement. See Delete tasks are bound to one index incarnation.

See Retention and deletes and Recovering a failed task.

Pod security

The chart exposes podSecurityContext and securityContext per role. The containers do not require privileged mode or host mounts. Set a non-root user and a read-only root filesystem if your policy requires them. The data plane needs write access to the WAL mount.

Reporting vulnerabilities

See SECURITY.md in the source repository.