Skip to content

CLI reference

Siglake rename

Product names, commands and links were normalized during the Siglake migration. Retained generation dates and commit IDs below identify the archived pre-rename source and binaries; they are not new build evidence.

The siglake binary runs the ingest and compactor roles. It also provides maintenance commands, offline warehouse tools and a client for the query server. The query server runs as the separate siglake-query-server binary.

Each binary writes tracing logs to stderr and command results to stdout. You can pipe or redirect machine-readable output without mixing it with logs.

This page is generated

Everything below the horizontal rule comes from --help on built siglake, siglake-query-server, and siglake-operator binaries. Regenerate with ./scripts/gen-reference.sh /path/to/siglake after changing any flag. Generated from Siglake commit 65c4eef on 2026-09-16. Binaries built from 65c4eef (Siglake), 65c4eef (siglake-query-server), 65c4eef (siglake-operator).

Command groups

Long-running server roles

Command Role
ingest-server OTLP + bulk ingest → WAL. --with-compactor runs compaction in-process.
compactor Drain sealed WAL segments into Iceberg and run maintenance. Use --role drain, --role maintenance, or the default --role combined to select the work.

The query server is siglake-query-server, not a siglake subcommand. The Kubernetes operator has its own siglake-operator flags.

Query and inspection commands

Command Purpose
sql Interactive client for a running query server. One-shot or REPL.
sql-direct Run DataFusion directly against a warehouse without a query server.
query Run SQL against locally ingested events.
subscribe Tail an Iceberg table by time column, incrementally.

SQL clients

Run a one-off SQL query from the command line with siglake sql '<query>'. This siglake sql command sends the SQL query to a running query server. Run siglake sql-direct --query '<query>' to query the warehouse directly. siglake sql-direct does not use a query server; siglake sql does. Without a query argument, siglake sql opens a read-eval-print loop (REPL).

siglake sql 'SELECT count(*) FROM events'
siglake sql-direct --query 'SELECT count(*) FROM events'

Maintenance commands

Command Purpose Execution
audit-rotate Bound query_audit growth. The default mode deletes all historical audit rows; use --max-age-secs for the non-destructive mode. Writes by default; pass --dry-run to preview.
gc-orphans Delete files no retained snapshot references. Dry-run by default; pass --apply to delete.
retention-sweep Drop whole data files past the retention horizon. Dry-run by default; pass --apply to rewrite the snapshot.
delete-sweep Execute pending GDPR/predicate delete tasks. Dry-run by default; pass --apply to rewrite files and advance the ledger.
migrate-schema Additively reconcile table schemas after a refused write. Use --all-tables --all-namespaces before a deployment-wide upgrade. Writes by default; pass --dry-run to preview.
wal-recover Pull mirrored WAL segments back from object storage. Writes to the --to destination; no preview flag.

Maintenance sweep differences

The purpose of siglake gc-orphans is storage cleanup: gc-orphans removes files that no retained snapshot references. siglake retention-sweep enforces time retention by removing expired data files. siglake delete-sweep executes queued predicate or General Data Protection Regulation (GDPR) delete tasks. These three sweep commands differ by target. All three are dry runs unless you pass --apply.

Development commands

Command Purpose
gen Emit synthetic NDJSON events.
ingest Ingest an NDJSON file.
iceberg-demo Append synthetic events and run canned queries.

Synthetic event generation

Generate synthetic load with siglake gen --n <count>. The --n parameter controls load volume and defaults to 100 events. siglake gen prints synthetic newline-delimited JSON (NDJSON) to stdout. It has no parameter for cardinality.

Ingest server deployment

Running siglake ingest-server directly starts the ingester process in the current environment. Deploying through the Helm chart runs that same siglake ingest-server subcommand in a managed pod. The operator also deploys the siglake ingest-server subcommand in a managed pod. The Helm chart and operator supply the flags and environment variables. The siglake ingest-server behavior is the same in all three cases.

Audit rotation warning

Destructive by default: bare audit-rotate

Without --max-age-secs, audit-rotate drops and recreates the query_audit table. This coarse row TTL deletes all historical audit rows. Iceberg-rust 0.9 has no public row-level delete. Use --max-age-secs N for the non-destructive snapshot-age sweep that keeps the rows.

Common flags

Most commands share the storage-addressing flags:

Flag Env Meaning
--data-dir SIGLAKE_DATA_DIR (siglake-query-server only) Global siglake local root. Used only when no --warehouse-url.
--warehouse-url SIGLAKE_WAREHOUSE_URL e.g. s3://bucket/prefix. Overrides --warehouse.
--catalog-uri SIGLAKE_CATALOG_URI postgres://… or sqlite://path?mode=rwc. Defaults to SQLite under the warehouse.

Flag and environment precedence

--data-dir is a global siglake flag, so every subcommand accepts it. The other storage flags in the table belong to individual commands. The SIGLAKE_DATA_DIR environment variable configures siglake-query-server, not the siglake binary. When a command accepts both a flag and its environment-variable configuration, the command-line flag takes precedence.

Object-store credentials

Object-store credentials come from the standard Amazon Web Services (AWS), Google Cloud and Azure environment variables. S3-compatible stores use AWS_ACCESS_KEY_ID, AWS_SECRET_ACCESS_KEY, AWS_REGION and AWS_ENDPOINT_URL.


siglake

siglake command-line interface: servers, maintenance jobs and a SQL client

Usage: siglake [OPTIONS] <COMMAND>

Commands:
  sql                   Interactive SQL client for a RUNNING siglake query server (`/api/v1/sql`): one-shot with a query argument, or a REPL without. Table output includes the per-query scan stats + server time; use `--dry-run` for the cost estimate without executing
  ingest                Ingest newline-delimited JSON events
  query                 Run a SQL query against the ingested events. Table name: `events`
  gen                   Generate `n` synthetic events to stdout (NDJSON) for local testing and demos
  iceberg-demo          Iceberg demo: append `n` synthetic events through the SQLite-backed `IcebergContext`, then run a few canned SQL queries against the resulting snapshot
  ingest-server         Run the OTLP ingest server
  wal-recover           Disaster-recovery: pull every WAL segment under `<warehouse-url>/<prefix>/` back onto a local WAL root, rebuilding the `<tenant>[/<index>]/sealed/` layout so the ordinary drain commits each segment to the namespace and table it came from. Skips segments already present locally, and recovers the active mirror too (preferring the sealed copy of any segment present as both)
  audit-rotate          Bound the `query_audit` Iceberg table's growth
  gc-orphans            Reclaim orphan files: physically delete data/manifest/ manifest-list files under a table's location that no retained snapshot references — the storage left behind by re-clustering overwrites + snapshot expiry
  retention-sweep       Enforce per-index retention policies by dropping whole data files whose manifest max timestamp is older than the configured horizon
  delete-sweep          Execute pending GDPR/delete tasks for one managed index
  rebuild-group-counts  Rebuild a table's group-count aggregate from the committed data files
  migrate-schema        Additively reconcile a table's stored schema toward the schema the running build declares for it
  sql-direct            Run a SQL query DIRECTLY against a warehouse via DataFusion — no query server involved (offline/ops tool; `siglake sql` is the client for a running server)
  subscribe             Tail an Iceberg-backed table by time-column. Prints each new row batch as it commits. The tailing primitive for a consumer of a siglake TABLE; to consume the write-ahead log itself — earlier, and with retention that waits for you — see `siglake_wal::consumer` and `docs/CONSUMING_SEGMENTS.md`
  compactor             Run the compactor: drain sealed WAL segments into Iceberg
  help                  Print this message or the help of the given subcommand(s)

Options:
      --data-dir <DATA_DIR>  Root directory used as the local object store [default: ./data]
  -h, --help                 Print help
  -V, --version              Print version

siglake sql

Interactive SQL client for a RUNNING siglake query server (`/api/v1/sql`): one-shot with a query argument, or a REPL without. Table output includes the per-query scan stats + server time; use `--dry-run` for the cost estimate without executing

Usage: siglake sql [OPTIONS] [QUERY]

Arguments:
  [QUERY]
          SQL to run once; omit for an interactive session

Options:
      --data-dir <DATA_DIR>
          Root directory used as the local object store

          [default: ./data]

      --endpoint <ENDPOINT>
          Query server base URL

          [env: SIGLAKE_ENDPOINT=]
          [default: http://localhost:8089]

      --token <TOKEN>
          Bearer token (when the server runs with --auth-tokens / OIDC)

          [env: SIGLAKE_TOKEN=]

      --format <FORMAT>
          Output format

          Possible values:
          - table:  Aligned columns + a stats footer (default)
          - json:   The raw response JSON, pretty-printed
          - ndjson: One JSON object per row (streamed from the server)

          [default: table]

      --dry-run
          Cost estimate only — plan the query without executing it

      --quiet
          Suppress the stats footer

  -h, --help
          Print help (see a summary with '-h')

siglake ingest

Ingest newline-delimited JSON events

Usage: siglake ingest [OPTIONS]

Options:
      --data-dir <DATA_DIR>  Root directory used as the local object store [default: ./data]
      --input <INPUT>        NDJSON input file. If omitted, reads from stdin
  -h, --help                 Print help

siglake query

Run a SQL query against the ingested events. Table name: `events`

Usage: siglake query [OPTIONS] --sql <SQL>

Options:
      --data-dir <DATA_DIR>  Root directory used as the local object store [default: ./data]
      --sql <SQL>            SQL query string
  -h, --help                 Print help

siglake gen

Generate `n` synthetic events to stdout (NDJSON) for local testing and demos

Usage: siglake gen [OPTIONS]

Options:
      --data-dir <DATA_DIR>  Root directory used as the local object store [default: ./data]
      --n <N>                Number of events to generate [default: 100]
  -h, --help                 Print help

siglake iceberg-demo

Iceberg demo: append `n` synthetic events through the SQLite-backed `IcebergContext`, then run a few canned SQL queries against the resulting snapshot.

The catalog (SQLite) and data (Parquet + manifests) persist between runs, so a second invocation appends to the same table rather than recreating it. Use `--reset` to wipe the warehouse first.

Usage: siglake iceberg-demo [OPTIONS]

Options:
      --data-dir <DATA_DIR>
          Root directory used as the local object store

          [default: ./data]

      --n <N>
          Number of events to ingest

          [default: 1000]

      --warehouse <WAREHOUSE>
          Warehouse subdirectory under `--data-dir`

          [default: warehouse]

      --reset
          Wipe the warehouse before running

  -h, --help
          Print help (see a summary with '-h')

siglake ingest-server

Run the OTLP ingest server.

HTTP endpoints: - POST /v1/logs        — OTLP/HTTP logs (JSON or protobuf) - GET  /api/v1/stream  — SSE tail, teed ahead of the WAL append - GET  /healthz        — liveness/readiness

Tenancy is SINGLE-TENANT by default: every request routes to the `default` tenant, and `X-Scope-OrgID` naming another one is refused. Multi-tenant routing is explicit — `--oidc-tenant-claim` takes the tenant from a verified JWT, `--trust-scope-header` takes the client's word. Each tenant gets its own WAL subtree + Iceberg namespace.

Without `--with-compactor`, this only writes WAL segments — you'll need to run `siglake compactor` separately to commit them to Iceberg. With `--with-compactor`, runs an in-process compactor every 1s for single-process demos.

Usage: siglake ingest-server [OPTIONS]

Options:
      --bind <BIND>
          Address to bind. The OTLP/HTTP ingest surface (`POST /v1/logs`) serves here; 8088 is the long-standing default and is kept so existing deployments do not have to move

          [default: 0.0.0.0:8088]

      --data-dir <DATA_DIR>
          Root directory used as the local object store

          [default: ./data]

      --metrics-bind <METRICS_BIND>
          Address to bind for the Prometheus `/metrics` endpoint

          [default: 0.0.0.0:9100]

      --wal <WAL>
          Subdirectory under `--data-dir` (or absolute path) for WAL segments

          [default: wal]

      --warehouse <WAREHOUSE>
          Subdirectory under `--data-dir` for the Iceberg warehouse, when running locally with no `--warehouse-url`

          [default: warehouse]

      --warehouse-url <WAREHOUSE_URL>
          Full warehouse URL (e.g. `s3://bucket/prefix`). Overrides `--warehouse`. Reads `SIGLAKE_WAREHOUSE_URL` if not given

          [env: SIGLAKE_WAREHOUSE_URL=]

      --catalog-uri <CATALOG_URI>
          Iceberg catalog URI (e.g. `postgres://user:pass@host/db`, `sqlite://path?mode=rwc`). Reads `SIGLAKE_CATALOG_URI` if not given. Defaults to a SQLite db under the warehouse dir

          [env: SIGLAKE_CATALOG_URI=]

      --wal-max-events <WAL_MAX_EVENTS>
          Roll the WAL segment after this many events

          [default: 4096]

      --wal-max-age-secs <WAL_MAX_AGE_SECS>
          Roll the WAL segment after this many seconds

          [default: 5]

      --with-compactor
          Also run a polling compactor in-process (every 1s)

      --otlp-grpc-listen <OTLP_GRPC_LISTEN>
          OTLP/gRPC listen address. Use --disable-otlp-grpc to turn it off

          [env: SIGLAKE_OTLP_GRPC_LISTEN=]
          [default: 0.0.0.0:4317]

      --disable-otlp-grpc
          Disable the default OTLP/gRPC logs and traces listener

      --wal-mirror-prefix <WAL_MIRROR_PREFIX>
          WAL → object-store mirror at `<warehouse-url>/<prefix>/`. Unset mirrors to `wal-mirror/` whenever `--warehouse-url` is set (the default since 2026-09-11), and is off without one. Pass an EMPTY value to turn the mirror off

          [env: SIGLAKE_WAL_MIRROR_PREFIX=]

      --wal-active-mirror-interval-secs <WAL_ACTIVE_MIRROR_INTERVAL_SECS>
          When > 0 and `--wal-mirror-prefix` is set, also mirror the currently-active segment every N seconds to `<prefix>/_active/<filename>`. N is the upload window for the in-flight segment, and it becomes an N-second data-loss bound on ONE recovery path: a successful upload is recovered when an operator runs `siglake wal-recover` onto the WAL root and the filesystem drain commits what it finds. Nothing reads an active snapshot on its own, and the catalog-claim drain reconciles sealed objects only — it never reads `_active/` and gets no N-second target

          [env: SIGLAKE_WAL_ACTIVE_MIRROR_INTERVAL_SECS=]
          [default: 0]

      --auth-tokens <AUTH_TOKENS>
          Comma-separated bearer tokens. When set, every request must carry `Authorization: Bearer <token>` where `<token>` matches one of these. These tokens say who may write, not what they may write as: binding tenancy to the caller is `--oidc-tenant-claim`. Empty/unset = no auth (the v0 behavior; only safe inside a trusted network)

          [env: SIGLAKE_AUTH_TOKENS=]

      --oidc-issuer <OIDC_ISSUER>
          OIDC issuer URL for JWT verification (e.g. `https://cognito-idp.us-east-1.amazonaws.com/<pool-id>`). When set together with `--oidc-audience`, every request must carry a valid `Authorization: Bearer <jwt>`. Takes precedence over `--auth-tokens`

          [env: SIGLAKE_OIDC_ISSUER=]

      --oidc-audience <OIDC_AUDIENCE>
          OIDC audience (client ID) to validate in the JWT `aud` claim

          [env: SIGLAKE_OIDC_AUDIENCE=]

      --oidc-tenant-claim <OIDC_TENANT_CLAIM>
          JWT claim that carries the tenant. When set, the tenant comes from the VERIFIED token rather than the `X-Scope-OrgID` header, and a header naming a different tenant is refused.

          Requires `--oidc-issuer` and `--oidc-audience`: the tenant comes from a verified token, so with no verifier there is nothing to take it from. Setting it alone — with static `--auth-tokens`, with open auth, or with `--trust-scope-header` still routing on the header — is refused at startup rather than accepted and ignored.

          Setting it also makes the claim MANDATORY: a verified token whose claim is missing, blank, not a string, longer than 128 chars, or outside `[A-Za-z0-9_-]` is refused with `403` before the batch is routed anywhere — token validity is not tenant authorization. The value is validated, never repaired, so an unusable claim never falls back to the `default` tenant. Refusals are counted by `siglake_ingest_tenant_denied_total` (`reason="claim_missing"` / `"claim_invalid"` / `"header_mismatch"`).

          Set this on any shared deployment: it is the only way to route tenants that does not take a client header at its word. Without it this ingester is single-tenant unless `--trust-scope-header` says otherwise.

          [env: SIGLAKE_OIDC_TENANT_CLAIM=]

      --trust-scope-header [<TRUST_SCOPE_HEADER>]
          Let the unverified `X-Scope-OrgID` header select the tenant.

          OFF by default since 2026-09-11: this ingester is single-tenant unless told otherwise, and a header naming any tenant but `default` is refused with `403` rather than honoured or ignored. Prefer `--oidc-tenant-claim`, which binds the tenant to a verified identity; reach for this only where a gateway in front of the ingester sets the header itself and strips the client's.

          Accepts `1`/`true`/`yes`/`on` (or the bare flag). Anything else, including an unset or empty value, leaves the header untrusted: a typo must not silently open multi-tenant routing.

          [env: SIGLAKE_TRUST_SCOPE_HEADER=]

      --ingest-rate-per-sec <INGEST_RATE_PER_SEC>
          Steady-state request rate per bearer token (or per `X-Forwarded-For` IP in open mode). 0 disables the limiter

          [env: SIGLAKE_INGEST_RATE_PER_SEC=]
          [default: 0]

      --ingest-rate-burst <INGEST_RATE_BURST>
          Burst capacity for the token-bucket rate limiter. Ignored when `--ingest-rate-per-sec=0`

          [env: SIGLAKE_INGEST_RATE_BURST=]
          [default: 0]

      --ingest-rate-redis-url <INGEST_RATE_REDIS_URL>
          Redis URL for a cross-replica shared rate budget. Empty/unset = per-replica in-memory limiter (the default). When set, the rate-per-sec / burst values apply to a single token-bucket shared across every ingester replica via a Redis Lua script. Example: `redis://redis.siglake-system:6379/0`

          [env: SIGLAKE_INGEST_RATE_REDIS_URL=]

      --ingest-rate-redis-prefix <INGEST_RATE_REDIS_PREFIX>
          Hash-key prefix the Redis rate budget uses. Defaults to `siglake:rb`. Use a unique prefix per siglake deployment that shares a Redis with other tenants

          [env: SIGLAKE_INGEST_RATE_REDIS_PREFIX=]

      --ingest-backpressure-capacity <INGEST_BACKPRESSURE_CAPACITY>
          Per-tenant capacity of the mpsc-fed writer queue. Default 1024 enables the BackpressureRouter (non-blocking mpsc path); set 0 to opt out and keep the legacy mutex-serialized path. A full lane returns 503 + `Retry-After` instead of blocking on `Mutex<WalWriter>`. Tune up for higher per-tenant burst tolerance

          [env: SIGLAKE_INGEST_BACKPRESSURE_CAPACITY=]
          [default: 1024]

      --ingest-group-commit-ms <INGEST_GROUP_COMMIT_MS>
          Group commit. When >0, the per-tenant writer task waits this many ms after the first command for more commands to accumulate before flushing. Trades p50 ack latency for amortized fsync cost across many batches. Requires `--ingest-backpressure-capacity > 0`

          [env: SIGLAKE_INGEST_GROUP_COMMIT_MS=]
          [default: 0]

      --ingest-backpressure-shards <INGEST_BACKPRESSURE_SHARDS>
          Per-tenant write parallelism. N writer tasks per tenant; ingest handlers round-robin across them. 1 (default) preserves the single-writer-per-tenant behavior. Requires `--ingest-backpressure-capacity > 0`

          [env: SIGLAKE_INGEST_BACKPRESSURE_SHARDS=]
          [default: 1]

      --ingest-mem-limit-mib <INGEST_MEM_LIMIT_MIB>
          RSS memory circuit breaker (MiB). When >0, a background task samples this process's resident set every `--ingest-mem-sample-secs`; ingest handlers shed load with `503` + `Retry-After` while RSS is at or above this ceiling. 0 (default) disables the breaker

          [env: SIGLAKE_INGEST_MEM_LIMIT_MIB=]
          [default: 0]

      --ingest-mem-sample-secs <INGEST_MEM_SAMPLE_SECS>
          RSS sampling interval (seconds) for the memory circuit breaker. Only meaningful when `--ingest-mem-limit-mib > 0`

          [env: SIGLAKE_INGEST_MEM_SAMPLE_SECS=]
          [default: 60]

      --allowed-tenants <ALLOWED_TENANTS>
          Tenants this ingester accepts, comma-separated. Empty = any. Checked against the tenant actually resolved, so it bounds a JWT claim as well as a trusted header.

          Each novel tenant mints a backpressure lane holding an open file, a fresh set of metric label values, and an Iceberg namespace with seven tables downstream. Set this on any deployment whose tenants are known.

          [env: SIGLAKE_ALLOWED_TENANTS=]

      --max-tenants <MAX_TENANTS>
          Distinct tenants to mint before refusing new ones. 0 = unbounded.

          The backstop for when the tenant set is not known ahead of time. Refusing an unexpected tenant is a bad day; exhausting file descriptors is an outage for every tenant that was real.

          Counts tenants, not `(tenant, index)` lanes — see `--ingest-max-lanes` for those. At the cap the tenants already writing are unaffected and a novel one is refused with `403` (`siglake_ingest_tenant_denied_total{reason="at_capacity"}`). The count is this process's: it starts empty on restart, and each pod holds its own.

          [env: SIGLAKE_MAX_TENANTS=]
          [default: 0]

      --ingest-max-lanes <INGEST_MAX_LANES>
          Distinct `(tenant, index)` backpressure lanes to create before refusing new ones. 0 = unbounded.

          Each lane holds an open file per shard. Both halves of the key are client headers, so the cross product is client-controlled.

          [env: SIGLAKE_INGEST_MAX_LANES=]
          [default: 0]

  -h, --help
          Print help (see a summary with '-h')

siglake wal-recover

Disaster-recovery: pull every WAL segment under `<warehouse-url>/<prefix>/` back onto a local WAL root, rebuilding the `<tenant>[/<index>]/sealed/` layout so the ordinary drain commits each segment to the namespace and table it came from. Skips segments already present locally, and recovers the active mirror too (preferring the sealed copy of any segment present as both)

Usage: siglake wal-recover [OPTIONS] --from <FROM> --to <TO>

Options:
      --data-dir <DATA_DIR>  Root directory used as the local object store [default: ./data]
      --from <FROM>          Full source URL (e.g. `s3://bucket/wal-mirror`) [env: SIGLAKE_WAL_MIRROR_URL=]
      --to <TO>              Local WAL ROOT to populate — the same path the ingester and compactor are pointed at (e.g. `/var/lib/siglake/wal`), NOT a `sealed/` subdirectory. The tenant/index layout is rebuilt beneath it
  -h, --help                 Print help

siglake audit-rotate

Bound the `query_audit` Iceberg table's growth.

Two modes: * default — **drop + recreate** (a coarse row-TTL: deletes ALL historical audit rows; iceberg-rust 0.9 has no public row-level delete). The recreated table keeps the same schema/partition/sort. * `--max-age-secs N` — **non-destructive snapshot-age sweep**: keep the rows, but expire `query_audit` snapshots older than N seconds (the per-flush commit churn) so the metadata.json stays bounded, then reclaim the now-unreferenced files. Uses the in-fork `expire_snapshots` age sweep + the same orphan GC as `gc-orphans`.

Schedule on a CronJob; the running query-server picks up the table on its next audit write (`ensure_query_audit_table` is idempotent). `--dry-run` reports what would happen without touching the catalog.

Usage: siglake audit-rotate [OPTIONS]

Options:
      --data-dir <DATA_DIR>
          Root directory used as the local object store

          [default: ./data]

      --warehouse-url <WAREHOUSE_URL>
          Iceberg warehouse URL (`s3://bucket/prefix`, `file:///abs/path`). Reads `SIGLAKE_WAREHOUSE_URL` if not given

          [env: SIGLAKE_WAREHOUSE_URL=]

      --catalog-uri <CATALOG_URI>
          Iceberg catalog URI (`postgres://...`, `sqlite://...`). Reads `SIGLAKE_CATALOG_URI` if not given

          [env: SIGLAKE_CATALOG_URI=]

      --namespace <NAMESPACE>
          Per-deployment Iceberg namespace. Defaults to the chart's `tenant.namespace` value (`siglake`)

          [env: SIGLAKE_TENANT_NAMESPACE=]
          [default: siglake]

      --table <TABLE>
          Table to rotate. Defaults to `query_audit`. Any table works with `--max-age-secs` (the non-destructive snapshot-age sweep — including any user index, e.g. a consumer's own output tables). The destructive drop-and-recreate path (no `--max-age-secs`) is restricted to `query_audit`

          [default: query_audit]

      --max-age-secs <MAX_AGE_SECS>
          Non-destructive mode: expire snapshots older than this many seconds (keeping rows) + reclaim orphans, instead of drop-and-recreate

      --dry-run
          Print what would happen and exit 0; don't touch the catalog. Useful for verifying the catalog URI + warehouse URL before scheduling the rotate

  -h, --help
          Print help (see a summary with '-h')

siglake gc-orphans

Reclaim orphan files: physically delete data/manifest/ manifest-list files under a table's location that no retained snapshot references — the storage left behind by re-clustering overwrites + snapshot expiry.

**Dry-run by default** — reports the orphan count + bytes and deletes nothing. Pass `--apply` to actually delete. Files modified within `--min-age-secs` are skipped (guards the in-flight-write race). Run on a periodic cadence (CronJob) per table. file/s3 warehouses only.

Usage: siglake gc-orphans [OPTIONS]

Options:
      --data-dir <DATA_DIR>
          Root directory used as the local object store

          [default: ./data]

      --warehouse-url <WAREHOUSE_URL>
          Iceberg warehouse URL. Reads `SIGLAKE_WAREHOUSE_URL` if not given

          [env: SIGLAKE_WAREHOUSE_URL=]

      --catalog-uri <CATALOG_URI>
          Iceberg catalog URI. Reads `SIGLAKE_CATALOG_URI` if not given

          [env: SIGLAKE_CATALOG_URI=]

      --namespace <NAMESPACE>
          Per-deployment Iceberg namespace

          [env: SIGLAKE_TENANT_NAMESPACE=]
          [default: siglake]

      --table <TABLE>
          Table to GC (`events`, `query_audit`, or any index id)

          [default: events]

      --min-age-secs <MIN_AGE_SECS>
          Skip files modified within this many seconds (safety window against deleting a concurrently-written, not-yet-referenced file)

          [default: 86400]

      --apply
          Actually delete. Without this, the command is a dry-run

  -h, --help
          Print help (see a summary with '-h')

siglake retention-sweep

Enforce per-index retention policies by dropping whole data files whose manifest max timestamp is older than the configured horizon.

Dry-run by default: prints the files/bytes/rows that would be removed and leaves the snapshot untouched. This command only performs the file drop rewrite; compose it with the existing snapshot-expiry + orphan-GC sweeps to reclaim superseded files physically. A Helm CronJob would mirror `audit-rotate`: same image/env, different subcommand.

Usage: siglake retention-sweep [OPTIONS]

Options:
      --data-dir <DATA_DIR>
          Root directory used as the local object store

          [default: ./data]

      --warehouse-url <WAREHOUSE_URL>
          Iceberg warehouse URL. Reads `SIGLAKE_WAREHOUSE_URL` if not given

          [env: SIGLAKE_WAREHOUSE_URL=]

      --catalog-uri <CATALOG_URI>
          Iceberg catalog URI. Reads `SIGLAKE_CATALOG_URI` if not given

          [env: SIGLAKE_CATALOG_URI=]

      --namespace <NAMESPACE>
          Per-deployment Iceberg namespace

          [env: SIGLAKE_TENANT_NAMESPACE=]
          [default: siglake]

      --index <INDEX>
          Optional managed index id. Absent => sweep every managed index in the resolved tenant namespace

      --apply
          Actually apply the retention rewrite. Without this, the command is a dry-run

  -h, --help
          Print help (see a summary with '-h')

siglake delete-sweep

Execute pending GDPR/delete tasks for one managed index.

Dry-run by default: evaluates the pending tasks and reports the files and rows they would rewrite without touching the snapshot or ledger.

Usage: siglake delete-sweep [OPTIONS] --index <INDEX>

Options:
      --data-dir <DATA_DIR>
          Root directory used as the local object store

          [default: ./data]

      --warehouse-url <WAREHOUSE_URL>
          Iceberg warehouse URL. Reads `SIGLAKE_WAREHOUSE_URL` if not given

          [env: SIGLAKE_WAREHOUSE_URL=]

      --catalog-uri <CATALOG_URI>
          Iceberg catalog URI. Reads `SIGLAKE_CATALOG_URI` if not given

          [env: SIGLAKE_CATALOG_URI=]

      --namespace <NAMESPACE>
          Per-deployment Iceberg namespace

          [env: SIGLAKE_TENANT_NAMESPACE=]
          [default: siglake]

      --index <INDEX>
          Managed index id whose pending delete tasks should execute

      --apply
          Actually apply the rewrites and advance the ledger. Without this, the command is a dry-run

  -h, --help
          Print help (see a summary with '-h')

siglake rebuild-group-counts

Rebuild a table's group-count aggregate from the committed data files.

The operator fallback for a LOST group-count delta. A delta write that exhausts its retries leaves a durable marker, and the maintenance compactor normally rebuilds the aggregate on its next fold. If that automatic rebuild fails or remains incomplete, the query guard refuses the cheap Tier-1 path for the affected columns — correctly, because a short aggregate answers wrongly. Symptom: a `GROUP BY` on a high-cardinality column that used to answer in milliseconds now takes seconds and reports `served_by: "materialized"`, while `siglake_group_count_delta_write_failures_total` is non-zero.

Reads the same tiers a query would — each file's group-count footer where it has one, a raw-page decode where it does not — so it is correct for columns too wide for the per-file footers, which is exactly the case that needs repairing. Cost is one expensive scan per column, once.

Safe to run live: the object is written under an optimistic lock and the rebuild refuses rather than merging if a concurrent writer wins. Deltas at or below the scanned snapshot become redundant and are cleaned up by the compactor; later ones fold on top as usual.

Usage: siglake rebuild-group-counts [OPTIONS]

Options:
      --data-dir <DATA_DIR>
          Root directory used as the local object store

          [default: ./data]

      --warehouse-url <WAREHOUSE_URL>
          Iceberg warehouse URL. Reads `SIGLAKE_WAREHOUSE_URL` if not given

          [env: SIGLAKE_WAREHOUSE_URL=]

      --catalog-uri <CATALOG_URI>
          Iceberg catalog URI. Reads `SIGLAKE_CATALOG_URI` if not given

          [env: SIGLAKE_CATALOG_URI=]

      --namespace <NAMESPACE>
          Per-deployment Iceberg namespace

          [env: SIGLAKE_TENANT_NAMESPACE=]
          [default: siglake]

      --table <TABLE>
          Table whose aggregate to rebuild (`events`, or a managed index id)

          [default: events]

      --admit-typed-columns
          Also add the typed (long/double/bool) columns the table's schema carries but its aggregate never did.

          For a table created before typed columns joined the side aggregates: `status` is in every file's footer and in no aggregate, so `GROUP BY status` reports `served_by: "materialized"` for the life of the table, and a plain rebuild — which repairs only what the aggregate already holds — cannot change that. This admits such columns, computing each exact full-table total from the files; no rewrite is needed. Each is held to `SIGLAKE_TYPED_GROUP_COUNT_CARDINALITY` on its whole-table distinct count, and is left absent (reported, never partial) if some live file cannot serve it. Without the flag the report still names the columns it would add.

  -h, --help
          Print help (see a summary with '-h')

siglake migrate-schema

Additively reconcile a table's stored schema toward the schema the running build declares for it.

Adds any column the code declares but the table lacks, as an optional (nullable) column — never drops, renames, reorders, or retypes. Existing data files stay readable, with the new columns reading back null for rows written before the migration. Idempotent: re-running once the table is up to date adds nothing. Safe to run live against a table being written — the commit is guarded by the schema-id optimistic lock. Until migration, the write path refuses writes that populate a column absent from the table, naming the column and remedy. Reconcile every declared table in every namespace with `siglake migrate-schema --all-tables --all-namespaces`.

Run as a one-shot Job before/at rollout. `--dry-run` reports the columns that *would* be added without touching the catalog.

Usage: siglake migrate-schema [OPTIONS]

Options:
      --data-dir <DATA_DIR>
          Root directory used as the local object store

          [default: ./data]

      --warehouse-url <WAREHOUSE_URL>
          Iceberg warehouse URL. Reads `SIGLAKE_WAREHOUSE_URL` if not given

          [env: SIGLAKE_WAREHOUSE_URL=]

      --catalog-uri <CATALOG_URI>
          Iceberg catalog URI. Reads `SIGLAKE_CATALOG_URI` if not given

          [env: SIGLAKE_CATALOG_URI=]

      --namespace <NAMESPACE>
          Per-deployment Iceberg namespace

          [env: SIGLAKE_TENANT_NAMESPACE=]
          [default: siglake]

      --table <TABLE>
          Table to migrate (`events` or `query_audit` — the tables siglake declares; a managed index's schema belongs to whoever declared it). Ignored when `--all-tables` is set

          [default: events]

      --all-tables
          Migrate every known table to its declared schema in one pass

      --all-namespaces
          Migrate every NAMESPACE in the warehouse, not just `--namespace`.

          Tenancy is header-based, so `events` exists once per tenant namespace. Without this a migration reports success having left every other tenant's table narrow. Namespaces with no such table are skipped, never created.

      --dry-run
          Report the columns that would be added and exit 0; don't commit

      --promote-attr <PROMOTE_ATTR>
          Include these promoted typed columns in the `events` declared schema so the migration adds them. Repeatable; same `attr_key:type[:column]` format as `compactor --promote-attr`. Pass the same set the compactor runs with

  -h, --help
          Print help (see a summary with '-h')

siglake sql-direct

Run a SQL query DIRECTLY against a warehouse via DataFusion — no query server involved (offline/ops tool; `siglake sql` is the client for a running server).

`events` plus every managed index are pre-registered. Any DataFusion-supported SQL works (joins, window functions, etc.).

Usage: siglake sql-direct [OPTIONS] --query <QUERY>

Options:
      --data-dir <DATA_DIR>
          Root directory used as the local object store

          [default: ./data]

      --query <QUERY>
          SQL query, e.g. `"SELECT count(*) FROM events"`

      --warehouse <WAREHOUSE>
          Subdirectory under `--data-dir` for the Iceberg warehouse, when running locally with no `--warehouse-url`

          [default: warehouse]

      --warehouse-url <WAREHOUSE_URL>
          Full warehouse URL. See `ingest-server --warehouse-url`

          [env: SIGLAKE_WAREHOUSE_URL=]

      --catalog-uri <CATALOG_URI>
          Iceberg catalog URI. See `ingest-server --catalog-uri`

          [env: SIGLAKE_CATALOG_URI=]

  -h, --help
          Print help (see a summary with '-h')

siglake subscribe

Tail an Iceberg-backed table by time-column. Prints each new row batch as it commits. The tailing primitive for a consumer of a siglake TABLE; to consume the write-ahead log itself — earlier, and with retention that waits for you — see `siglake_wal::consumer` and `docs/CONSUMING_SEGMENTS.md`

Usage: siglake subscribe [OPTIONS] --table <TABLE>

Options:
      --data-dir <DATA_DIR>            Root directory used as the local object store [default: ./data]
      --table <TABLE>                  Table to tail: `events`, or any index id
      --time-column <TIME_COLUMN>      Time column to cursor on. Defaults to `events.timestamp`, or an index's own declared `timestamp_field`
      --since <SINCE>                  Subscription cursor start (RFC3339). Defaults to "1 minute ago"
      --interval-secs <INTERVAL_SECS>  Polling interval (seconds) [default: 1]
      --once                           Process current snapshot once and exit
      --warehouse <WAREHOUSE>          Subdirectory under `--data-dir` for the Iceberg warehouse, when running locally with no `--warehouse-url` [default: warehouse]
      --warehouse-url <WAREHOUSE_URL>  Full warehouse URL. See `ingest-server --warehouse-url` [env: SIGLAKE_WAREHOUSE_URL=]
      --catalog-uri <CATALOG_URI>      Iceberg catalog URI. See `ingest-server --catalog-uri` [env: SIGLAKE_CATALOG_URI=]
  -h, --help                           Print help

siglake compactor

Run the compactor: drain sealed WAL segments into Iceberg

Usage: siglake compactor [OPTIONS]

Options:
      --data-dir <DATA_DIR>
          Root directory used as the local object store

          [default: ./data]

      --metrics-bind <METRICS_BIND>
          Address to bind for the Prometheus `/metrics` endpoint

          [default: 0.0.0.0:9101]

      --wal <WAL>
          Subdirectory under `--data-dir` (or absolute path) for WAL segments

          [default: wal]

      --warehouse <WAREHOUSE>
          Subdirectory under `--data-dir` for the Iceberg warehouse, when running locally with no `--warehouse-url`

          [default: warehouse]

      --warehouse-url <WAREHOUSE_URL>
          Full warehouse URL. See `ingest-server --warehouse-url`

          [env: SIGLAKE_WAREHOUSE_URL=]

      --catalog-uri <CATALOG_URI>
          Iceberg catalog URI. See `ingest-server --catalog-uri`

          [env: SIGLAKE_CATALOG_URI=]

      --once
          Process the currently-sealed segments and exit

      --interval-secs <INTERVAL_SECS>
          Polling interval (seconds) when running as a daemon

          [default: 1]

      --catalog-claim
          Switch to multi-pod-safe catalog-claim coordination. When set, the compactor no longer reads `<wal>/sealed/`; it claims segments out of the `wal_segments` SQL table and fetches their bytes from the WAL mirror at `<warehouse-url>/<mirror-prefix>/<id>.arrow`. Requires `--warehouse-url` (object-store mirror root) and `--catalog-uri` (Postgres / SQLite for the claim table)

          [env: SIGLAKE_COMPACTOR_CATALOG_CLAIM=]

      --mirror-prefix <MIRROR_PREFIX>
          Mirror prefix under `--warehouse-url` to scan for catalog-claim mode. Defaults to `wal-mirror` — the same constant the ingester writes under, the chart renders and the operator sets, so the reader and the writer cannot drift apart

          [env: SIGLAKE_WAL_MIRROR_PREFIX=]
          [default: wal-mirror]

      --catalog-claim-batch <CATALOG_CLAIM_BATCH>
          Max segments per `try_claim` cycle in catalog-claim mode. Sized to drain an accumulated commit-batch in one commit; aligns with the local-FS 64-segment cap

          [env: SIGLAKE_COMPACTOR_CATALOG_CLAIM_BATCH=]
          [default: 64]

      --role <ROLE>
          Which half of the work this process performs: `drain` (WAL -> Iceberg only), `maintenance` (reclustering, snapshot expiry, gauge sampling, aggregate folding, delete tasks), or `combined` (both, the default).

          Setting the reclustering interval to 0 does NOT make a process drain-only -- expiry, gauge sampling and folding all still run.

          [env: SIGLAKE_COMPACTOR_ROLE=]
          [default: combined]

      --fs-claim-max-segments <FS_CLAIM_MAX_SEGMENTS>
          Max sealed segments to claim per cycle on the local-FS path. `0` disables the limit

          [env: SIGLAKE_COMPACTOR_FS_CLAIM_MAX_SEGMENTS=]
          [default: 64]

      --fs-claim-max-bytes <FS_CLAIM_MAX_BYTES>
          Max total on-disk bytes to claim per cycle on the local-FS path. `0` disables the limit

          [env: SIGLAKE_COMPACTOR_FS_CLAIM_MAX_BYTES=]
          [default: 67108864]

      --promote-attr <PROMOTE_ATTR>
          Promote a declared OTLP attribute out of the `attributes` JSON into its own typed column on write. Repeatable. Format `attr_key:type[:column]`, type ∈ string|int|float|bool, column defaults to the attr key with `.`/`-`→`_`. E.g. `--promote-attr http.status_code:int --promote-attr k8s.namespace:string`. The events table is widened additively at startup

  -h, --help
          Print help (see a summary with '-h')

siglake-query-server

HTTP query API for the siglake Iceberg warehouse

Usage: siglake-query-server [OPTIONS]

Options:
      --bind <BIND>
          Address to bind the HTTP API

          [env: SIGLAKE_QUERY_BIND=]
          [default: 0.0.0.0:8089]

      --metrics-bind <METRICS_BIND>
          Address to bind the Prometheus `/metrics` endpoint

          [env: SIGLAKE_QUERY_METRICS_BIND=]
          [default: 0.0.0.0:9105]

      --data-dir <DATA_DIR>
          Root directory used as the local data dir. Only consulted when `--warehouse-url` is omitted (local-fs warehouse mode)

          [env: SIGLAKE_DATA_DIR=]
          [default: ./data]

      --warehouse <WAREHOUSE>
          Subdirectory under `--data-dir` for the warehouse, when running locally with no `--warehouse-url`

          [default: warehouse]

      --warehouse-url <WAREHOUSE_URL>
          Full warehouse URL (e.g. `s3://bucket/prefix`)

          [env: SIGLAKE_WAREHOUSE_URL=]

      --catalog-uri <CATALOG_URI>
          Iceberg catalog URI (e.g. `postgres://user:pass@host/db`, `sqlite://path?mode=rwc`)

          [env: SIGLAKE_CATALOG_URI=]

      --tokens <TOKENS>
          Comma-separated bearer tokens accepted on `/api/v1/*` routes. Empty or unset = no bearer mode. Ignored when `--oidc-issuer` is also set (OIDC wins). Reads `SIGLAKE_QUERY_TOKENS` if not given on the command line

          [env: SIGLAKE_QUERY_TOKENS=]

      --oidc-issuer <OIDC_ISSUER>
          OIDC issuer URL. When set, the server discovers the JWKS via `<issuer>/.well-known/openid-configuration` and verifies every `Authorization: Bearer <jwt>` against it. Requires `--oidc-audience`

          [env: SIGLAKE_OIDC_ISSUER=]

      --oidc-audience <OIDC_AUDIENCE>
          Expected `aud` claim. Required when `--oidc-issuer` is set

          [env: SIGLAKE_OIDC_AUDIENCE=]

      --oidc-tenant-claim <OIDC_TENANT_CLAIM>
          Name of the JWT claim that carries the per-request tenant identifier. Setting it enables per-request multi-tenancy and makes the claim MANDATORY: each verified caller is routed to its own Iceberg namespace (`tenant_<claim>`, created on first use), and a token whose claim is missing, blank, not a string, longer than 128 chars, or outside `[A-Za-z0-9_-]` is refused with `403` before the request is routed anywhere. The value is validated, never repaired, so an unusable claim never falls back to the default namespace. When unset, tenancy is not derived from a claim at all and every caller reads the default namespace (the `--tenant-namespace` value). Requires `--oidc-issuer`

          [env: SIGLAKE_OIDC_TENANT_CLAIM=]

      --max-rows <MAX_ROWS>
          Server-side cap on rows returned per query. NDJSON streams stop at the cap and emit a truncation marker; records responses truncate, flip the `truncated` flag, and return 413

          [env: SIGLAKE_QUERY_MAX_ROWS=]
          [default: 1000000]

      --query-scan-partitions <QUERY_SCAN_PARTITIONS>
          Number of source partitions to expose from the custom siglake Iceberg scan. Unset keeps the DataFusion default

          [env: SIGLAKE_QUERY_SCAN_PARTITIONS=]

      --query-peers <QUERY_PEERS>
          Distributed query (#7): comma-separated worker base URLs, one per shard (e.g. `http://q0:8089,http://q1:8089`). Enables `/api/v1/sql/distributed`, which fans a query across these peers and merges. Each node also serves `/api/v1/sql/shard` as a worker.

          A FIXED list: a pod added beyond it receives no shard work. Kubernetes deployments should use `--query-peer-discovery-srv` instead; this remains the non-Kubernetes and test compatibility mode, and its documented contract is that the coordinator is peer zero. Setting both is refused.

          [env: SIGLAKE_QUERY_PEERS=]

      --query-peer-discovery-srv <QUERY_PEER_DISCOVERY_SRV>
          Distributed query (#967): the SRV record naming the query tier's headless Service port, e.g. `_http._tcp.siglake-query-headless.default.svc.cluster.local`. A background task re-resolves it and publishes a membership snapshot, so every Ready replica becomes eligible for shard work without a rollout. Each query pins ONE snapshot, so membership never moves under a running query. Mutually exclusive with `--query-peers`

          [env: SIGLAKE_QUERY_PEER_DISCOVERY_SRV=]

      --query-peer-scheme <QUERY_PEER_SCHEME>
          Scheme discovered peers are addressed with (`http` or `https`). SRV records carry a target and a port but no scheme, so this is explicit. Only meaningful with `--query-peer-discovery-srv`

          [env: SIGLAKE_QUERY_PEER_SCHEME=]

      --query-peer-discovery-interval-secs <QUERY_PEER_DISCOVERY_INTERVAL_SECS>
          How often peer discovery re-resolves the SRV record, in seconds. CoreDNS remains the TTL authority; this bounds how quickly a scale event becomes visible to fan-out

          [env: SIGLAKE_QUERY_PEER_DISCOVERY_INTERVAL_SECS=]
          [default: 5]

      --query-peer-self-name <QUERY_PEER_SELF_NAME>
          This pod's name, matched against the SRV answer to find the coordinator's OWN worker URL (the failover target). Defaults to the `HOSTNAME` the container runtime sets, which for a StatefulSet pod is its pod name

          [env: SIGLAKE_QUERY_PEER_SELF_NAME=]

      --query-coordinator-token <QUERY_COORDINATOR_TOKEN>
          Bearer token the coordinator presents to peer workers (when they enforce auth)

          [env: SIGLAKE_QUERY_COORDINATOR_TOKEN=]

      --query-scan-reader-budget <QUERY_SCAN_READER_BUDGET>
          Aggregate object-store reader budget per query. When set, the custom scan derives per-partition reader concurrency from the planned source partition count so total reader fan-out stays bounded under wider scans

          [env: SIGLAKE_QUERY_SCAN_READER_BUDGET=]

      --query-scan-file-concurrency <QUERY_SCAN_FILE_CONCURRENCY>
          Per-source-partition object-store read concurrency inside the custom siglake Iceberg scan. Unset derives automatically

          [env: SIGLAKE_QUERY_SCAN_FILE_CONCURRENCY=]

      --query-scan-batch-size <QUERY_SCAN_BATCH_SIZE>
          Target Arrow batch size inside the custom siglake Iceberg scan. Unset keeps the iceberg reader default

          [env: SIGLAKE_QUERY_SCAN_BATCH_SIZE=]

      --query-scan-range-coalesce-bytes <QUERY_SCAN_RANGE_COALESCE_BYTES>
          Merge nearby object-store byte ranges into larger reads when the gap is smaller than this many bytes. Unset keeps the iceberg reader default

          [env: SIGLAKE_QUERY_SCAN_RANGE_COALESCE_BYTES=]

      --query-scan-range-fetch-concurrency <QUERY_SCAN_RANGE_FETCH_CONCURRENCY>
          Maximum concurrent merged byte-range fetches per source partition. Unset keeps the iceberg reader default

          [env: SIGLAKE_QUERY_SCAN_RANGE_FETCH_CONCURRENCY=]

      --query-scan-range-adaptive-min-bytes <QUERY_SCAN_RANGE_ADAPTIVE_MIN_BYTES>
          Only enable the merged-range reader path when the planned scan touches at least this many bytes. Unset applies the range settings to every scan

          [env: SIGLAKE_QUERY_SCAN_RANGE_ADAPTIVE_MIN_BYTES=]

      --query-scan-range-adaptive-min-files <QUERY_SCAN_RANGE_ADAPTIVE_MIN_FILES>
          Only enable the merged-range reader path when the planned scan touches at least this many files. Unset applies the range settings to every scan

          [env: SIGLAKE_QUERY_SCAN_RANGE_ADAPTIVE_MIN_FILES=]

      --query-scan-adaptive-partition-min-bytes <QUERY_SCAN_ADAPTIVE_PARTITION_MIN_BYTES>
          Reduce source partition fan-out for smaller planned scans by targeting at least this many bytes per partition. Unset keeps the fixed `--query-scan-partitions` ceiling

          [env: SIGLAKE_QUERY_SCAN_ADAPTIVE_PARTITION_MIN_BYTES=]

      --query-scan-adaptive-partition-min-files <QUERY_SCAN_ADAPTIVE_PARTITION_MIN_FILES>
          Reduce source partition fan-out for smaller planned scans by targeting at least this many files per partition. Unset keeps the fixed `--query-scan-partitions` ceiling

          [env: SIGLAKE_QUERY_SCAN_ADAPTIVE_PARTITION_MIN_FILES=]

      --query-scan-file-cache-max-bytes <QUERY_SCAN_FILE_CACHE_MAX_BYTES>
          Total in-memory byte budget for the experimental source-file batch cache (process-lifetime, LRU). Disabled when unset or 0; enabling it requires positive byte and entry limits

          [env: SIGLAKE_QUERY_SCAN_FILE_CACHE_MAX_BYTES=]

      --query-scan-file-cache-max-entries <QUERY_SCAN_FILE_CACHE_MAX_ENTRIES>
          Maximum number of source files to retain in the experimental source-file batch cache. Disabled when unset or 0; enabling it requires positive byte and entry limits

          [env: SIGLAKE_QUERY_SCAN_FILE_CACHE_MAX_ENTRIES=]

      --query-object-cache-max-bytes <QUERY_OBJECT_CACHE_MAX_BYTES>
          Byte-range object cache budget (bytes) — the cold-S3 hot cache that caches Parquet footers, column chunks, and index sidecars so repeat/warm reads (incl. the FTS/bloom pruning path) don't re-hit S3.

          DEFAULT-ON, sized at **1/4 of the container memory limit** (64 MiB floor, 16 GiB cap); 1 GiB when there is no cgroup limit to read. Set to 0 to disable. A 64 GiB pod derives to the 16 GiB every published board used.

          [env: SIGLAKE_OBJECT_CACHE_BYTES=]

      --query-parsed-index-cache-max-bytes <QUERY_PARSED_INDEX_CACHE_MAX_BYTES>
          Budget (bytes) for parsed per-file inverted indexes, the form a warm text query is served from.

          DEFAULT-ON, sized at **1/16 of the container memory limit** (64 MiB floor, 1 GiB cap); 1 GiB when there is no cgroup limit to read. A parsed index costs about 40 bytes per indexed row. Set to 0 to deserialize per query, as before this cache existed.

          [env: SIGLAKE_PARSED_INDEX_CACHE_MAX_BYTES=]

      --query-puffin-blob-cache-max-bytes <QUERY_PUFFIN_BLOB_CACHE_MAX_BYTES>
          Budget (bytes) for the serialized Puffin blobs those indexes are parsed from, which a parsed eviction falls back on.

          DEFAULT-ON, sized at **1/64 of the container memory limit** (16 MiB floor, 256 MiB cap); 256 MiB when there is no cgroup limit to read — a quarter of the parsed budget, which is about what a blob is of its parsed form, so the two cover the same files. Set to 0 to keep only the parsed form and re-fetch on a miss.

          [env: SIGLAKE_PUFFIN_BLOB_CACHE_MAX_BYTES=]

      --query-wal-buffer-dir <QUERY_WAL_BUFFER_DIR>
          WAL root the query node can read. When set, queries union the Iceberg snapshot with the uncommitted WAL segments (sealed/ + processing/) under this directory, so just-ingested rows are visible before the compaction commit. This covers `events` and every managed user index the query references. Requires this pod to see the ingester's WAL — a shared RWX volume (EFS) in the distributed deployment. Unset ⇒ disabled (historical commit-cycle visibility)

          [env: SIGLAKE_QUERY_WAL_BUFFER_DIR=]

      --query-hot-caches
          Enable query-tier hot caches. When set (and `--query-wal-buffer-dir` is provided), a background task tails the WAL root and serves the `last_values()` and `distinct_values(<dim>)` UDTFs from per-tenant last-value + distinct caches. Unset ⇒ disabled

          [env: SIGLAKE_QUERY_HOT_CACHES=]

      --audit-disabled
          Disable best-effort audit-log emission. By default each completed query is submitted to the bounded `siglake.query_audit` writer; rows are dropped whole if that writer is over budget or unavailable

          [env: SIGLAKE_QUERY_AUDIT_DISABLED=]

      --jobs-postgres-uri <JOBS_POSTGRES_URI>
          Postgres URI for the persistent batch-job store. When set, batch jobs survive pod restart. Every replica shares the store, so recovery is scoped to execution ownership: a starting replica leaves jobs owned by a still-heartbeating sibling alone, and marks a job `failed` only once its owner's lease has expired (see `--jobs-owner-lease-secs`). That includes this pod's own previous incarnation, so a restart surfaces a terminal state within one lease period rather than instantly. When unset or BLANK, batch state is in-memory and private to this process: pod restart drops every queued and running job, and a peer replica answers 404 for them. The chart sets this from the catalog URI by default (`query.jobs.persistent: true`)

          [env: SIGLAKE_JOBS_POSTGRES_URI=]

      --jobs-owner-lease-secs <JOBS_OWNER_LEASE_SECS>
          How stale a query replica's job-owner heartbeat may get before its in-flight batch jobs are considered orphaned and failed. The owner heartbeats at a third of this interval, so three consecutive lost writes are survivable. Postgres-backed store only

          [env: SIGLAKE_JOBS_OWNER_LEASE_SECS=]
          [default: 120]

      --jobs-ownerless-grace-secs <JOBS_OWNERLESS_GRACE_SECS>
          Grace period for batch-job rows that carry no execution owner — rows written by a build older than ownership tracking. Their executor's liveness is unknowable, so they are failed only once they are this old. Postgres-backed store only

          [env: SIGLAKE_JOBS_OWNERLESS_GRACE_SECS=]
          [default: 86400]

      --jobs-cancel-poll-secs <JOBS_CANCEL_POLL_SECS>
          How often an executing replica re-reads the status of the batch jobs it is running, so a cancellation persisted by *another* replica reaches the future that holds the admission reservation and the storage-scan cancel guard. This is the bound the `202` from `DELETE /api/v1/jobs/<id>` promises: after it, the work has stopped. The query is by primary key over this pod's in-flight jobs only. Postgres-backed store only — with the in-memory store the replica that accepts the DELETE is necessarily the executor and the abort is immediate

          [env: SIGLAKE_JOBS_CANCEL_POLL_SECS=]
          [default: 2]

      --tls-cert <TLS_CERT>
          Path to a PEM-encoded TLS cert chain. Pair with `--tls-key` to serve HTTPS instead of HTTP. When either is unset the binary stays HTTP-only and TLS is left to a fronting Ingress

          [env: SIGLAKE_QUERY_TLS_CERT=]

      --tls-key <TLS_KEY>
          Path to the matching PEM-encoded private key

          [env: SIGLAKE_QUERY_TLS_KEY=]

      --tenant-namespace <TENANT_NAMESPACE>
          Iceberg namespace this query-server reads from. Defaults to `siglake`. Set per-tenant to run multiple isolated siglake deployments against one shared Iceberg catalog + warehouse; each tenant gets its own namespace. This is deployment-level tenancy: it is the namespace callers read while `--oidc-tenant-claim` is unset. With a claim configured, every request is routed by its own claim instead

          [env: SIGLAKE_TENANT_NAMESPACE=]
          [default: siglake]

  -h, --help
          Print help (see a summary with '-h')

  -V, --version
          Print version

siglake-operator

Kubernetes operator for SiglakeCluster

Usage: siglake-operator [OPTIONS]

Options:
      --prometheus-url <PROMETHEUS_URL>
          Prometheus base URL the reconciler queries for load signals. Defaults to the `prometheus-server` Service installed by the prometheus-community/prometheus chart: `http://prometheus-server.monitoring.svc.cluster.local:80`. kube-prometheus-stack uses a different Service; override this value [env: SIGLAKE_PROMETHEUS_URL=] [default: http://prometheus-server.monitoring.svc.cluster.local:80]
      --watch-namespace <NS>
          Restrict the watch to one or more namespaces. Repeatable. When empty, the operator watches every namespace
      --lease-namespace <LEASE_NAMESPACE>
          Namespace where the leader-election Lease lives. Defaults to the pod's namespace via `$POD_NAMESPACE`, falling back to `siglake-system` [env: POD_NAMESPACE=] [default: siglake-system]
      --pod-name <POD_NAME>
          Unique identity for this replica (typically the pod name). Reads `$POD_NAME` so the Helm Deployment pods carry their own names via the downward API [env: POD_NAME=]
      --no-leader-election
          Disable leader election. Safe for single-replica deployments; required to be off for `replicas > 1`
      --metrics-bind <METRICS_BIND>
          Bind address for the operator's own `/metrics` endpoint. Set to empty string to disable [env: SIGLAKE_METRICS_BIND=] [default: 0.0.0.0:9190]
      --print-crd
          Print the rendered CRD YAML and exit. Used by the install flow: `siglake-operator --print-crd | kubectl apply -f -`
      --adopt-values <VALUES_YAML>
          Helm-release adoption preflight (offline): synthesize a SiglakeCluster from the given chart VALUES file, run the parity checks, print the CR + handover runbook, and exit. See docs/DESIGN_operator_adoption.md
      --adopt-cluster-name <ADOPT_CLUSTER_NAME>
          CR / release name for adoption (must equal the helm release name for name+selector parity). Default "siglake" [default: siglake]
      --adopt-catalog-uri <ADOPT_CATALOG_URI>
          Catalog URI for adoption (cannot be inferred from chart values — the chart wires Postgres through a Secret)
  -h, --help
          Print help
  -V, --version
          Print version