Multi-tenancy¶
Siglake uses a tenant identifier to route WAL segments and Iceberg tables. Authentication decides whether a caller may connect. Tenant routing decides which namespace that accepted caller uses.
Ingest is single-tenant by default. Every request uses default, and Siglake
refuses a header that names another tenant. The two opt-ins select a tenant
from either a verified JSON Web Token (JWT) claim or a trusted header.
The opt-ins change routing only. Every tenant still shares compute, the SQL catalog, the warehouse bucket, and object-storage credentials. What is not isolated describes the effect of sharing these resources.
How a tenant is selected¶
Tenant selection differs between ingest and query. Ingest routing describes the default tenant and both opt-ins. Query routing uses only a configured JWT claim.
Ingest routing¶
By default, every ingest request uses default. A header naming another tenant
returns 403 over HTTP or PermissionDenied over OpenTelemetry Protocol
(OTLP)/gRPC. A header naming default has no effect.
Two Helm values opt the ingester into tenant routing:
| Value | Ingest routing |
|---|---|
ingester.oidc.tenantClaim |
Takes the tenant from the verified JWT claim. A present header must match it. |
ingester.trustScopeHeader |
Trusts X-Scope-OrgID from the caller. An absent header means default. |
If both values are set, the verified JWT claim takes precedence. The claim
must be a non-empty string of at most 128 characters containing only A-Z,
a-z, 0-9, _, or -. It must not match a reserved WAL directory name.
Siglake rejects an unusable claim with 403; it does not remove invalid
characters or choose default.
If X-Scope-OrgID is present in claim-based mode, it must equal the claim.
The mismatch also returns 403.
Trusted-header routing requires a gateway that sets X-Scope-OrgID and
filters caller-supplied values. Authentication does not bind a cluster-wide
bearer token to one tenant. An invalid HTTP header value returns 400.
--allowed-tenants restricts ingest to a known set. --max-tenants caps how
many distinct tenants one ingester admits when the set is open.
--ingest-max-lanes separately limits the number of tenant and index writer
lanes.
siglake_ingest_tenant_denied_total{reason} counts the 403 refusals above,
one label per setting. On the shipped single-tenant default the label is
header_not_trusted: X-Scope-OrgID named a tenant other than default, and
no setting lets the ingester route by it. Claim-based routing emits
claim_missing for a token without the claim, claim_invalid for a claim that
is not a usable identifier, and header_mismatch for a header that contradicts
the claim. --allowed-tenants emits not_allowed once routing has resolved a
tenant outside its set, and --max-tenants emits at_capacity for a novel
tenant it turns away. All six series exist at zero from startup, so the first
refused request is visible. The request is denied before a WAL lane or Iceberg
namespace is created.
What the tenant cap counts¶
--max-tenants / ingester.maxTenants caps the distinct tenants one ingester
admits. It counts resolved tenants, not (tenant, index) lanes: one tenant
writing to twelve indexes is one tenant and twelve lanes, and
--ingest-max-lanes bounds that cross product. The default is 0, which
leaves the tenant count unbounded; an unbounded ingester keeps no admission set
at all.
At the cap, the tenants already admitted keep writing. Only a tenant the
process has not admitted before is refused, with 403 over HTTP or
PermissionDenied over OTLP/gRPC, before a writer opens, a WAL directory is
made or a row is published. The refusal increments
siglake_ingest_tenant_denied_total{reason="at_capacity"}. The allow-list is
checked first, so a tenant outside --allowed-tenants never spends a slot.
The cap is held per ingester process and per pod, and it starts empty on
restart. A 2-replica ingester with maxTenants: 100 admits up to 100 tenants
on each pod, and the pods decide independently: a novel tenant can be admitted
by one pod and refused by another that is already full. A restarted pod
re-admits the tenants that already have WAL directories, in the order they next
write. Size the value per pod, and use --allowed-tenants when you need a
bound on the tenant set itself.
What is isolated¶
WAL¶
The filesystem WAL has a directory for each tenant. TenantWalRouter creates
writers beneath that directory, including separate lanes for managed indexes.
The default tenant uses the WAL root's default layout. Reserved WAL directory
names cannot be tenant identifiers.
Iceberg namespaces¶
A non-default tenant maps to the Iceberg namespace tenant_<id>. The default
tenant uses the main namespace. Siglake creates a tenant namespace only after
the boundary accepts its identifier.
Tables with the same name in two tenant namespaces have separate snapshots, manifests, and data files. The namespace is the storage boundary.
User indexes¶
Managed user indexes live inside the tenant namespace. Two tenants can use the same index name with different mappings and retention settings.
Rate budgets¶
Ingest rate limiting is a token bucket, and it is off until you configure it. A bucket starts full at the burst value, spends one token per request, and refills at the configured rate. The budget counts requests, not events: a client that batches 10,000 events into one export spends one token.
Set ingester.rateLimit.ratePerSec above 0 to turn the limiter on.
| Helm value | Flag | Default | What it sets |
|---|---|---|---|
ingester.rateLimit.ratePerSec |
--ingest-rate-per-sec |
0, the limiter off |
Steady-state requests per second for one bucket. |
ingester.rateLimit.burst |
--ingest-rate-burst |
0, which takes the rate value, and at least 1 |
Bucket capacity, so the largest burst one bucket can spend at once. |
ingester.rateLimit.redis.url |
--ingest-rate-redis-url |
Empty, so each replica keeps its own buckets | Redis holding one bucket per budget for the whole deployment. |
ingester.rateLimit.redis.prefix |
--ingest-rate-redis-prefix |
siglake:rb |
Prefix on the Redis entries, so two deployments sharing one Redis do not spend the same budget. |
The configuration reference gives the environment variable behind each flag.
Which bucket a request spends¶
The bucket is per tenant only where X-Scope-OrgID selects the tenant. The
limiter runs ahead of authentication, because it is what protects the
authentication path, so it reads the header before anything has verified it.
| Ingester mode | What the bucket is per |
|---|---|
ingester.trustScopeHeader on |
The tenant in X-Scope-OrgID. A request without the header falls through to the next rule. |
ingester.oidc.tenantClaim set |
The bearer token. The claim is read after the limiter, so the budget is per token, not per tenant. |
| Single-tenant, the default | The bearer token, else the first X-Forwarded-For address, else one anonymous bucket shared by every caller. |
With the in-memory limiter each replica holds its own buckets, so a deployment
of N replicas grants N times the configured rate. Setting redis.url replaces
that with one shared bucket per budget across replicas. If a Redis command
fails, the limiter admits the request and counts
siglake_rate_budget_backend_errors_total, so the budget goes unenforced until
Redis recovers. Alert on that counter.
The limiter covers the HTTP ingest surface. OTLP/gRPC exports never reach it, so the lane queue and the memory breaker are the only limits they meet.
What a client sees when a lane fills or memory runs out¶
Each tenant writes through its own bounded queue: one lane per tenant and index
pair, ingester.backpressure.capacity commands deep, drained by a dedicated
writer task. A full lane refuses the request instead of blocking the handler,
and it refuses only that tenant's request.
The memory circuit breaker is per process, not per tenant. It samples the
ingester's resident set size (RSS) every ingester.memBreaker.sampleSecs
seconds, and while RSS is at or above ingester.memBreaker.limitMib the
ingester sheds every request that replica receives, including requests from
tenants whose own lanes are empty. The sampler reads /proc, so the breaker
never trips on a host that is not Linux.
| Condition | What the client gets | Which metric moves |
|---|---|---|
| Rate budget spent | 429, Retry-After set to the seconds until the next token, and a JSON body repeating it as retry_after_secs. HTTP only. |
siglake_ingest_rate_limit_rejected_total, labelled by key_kind |
| Tenant lane full | 503, Retry-After: 1, and the text Service Unavailable: ingester backlog full. |
siglake_ingest_backpressure_rejected_total, labelled by endpoint, tenant and index |
| Memory breaker tripped | 503, Retry-After: 5, and the text Service Unavailable: ingester over memory limit. |
siglake_ingest_mem_breaker_open goes to 1, and siglake_ingest_mem_breaker_rejected_total, labelled by endpoint, counts the refusal once on either transport |
| Lane limit reached | 503, with no Retry-After. This is the --ingest-max-lanes bound refusing a new tenant or index, not backpressure. |
siglake_ingest_lane_refused_total |
What a 503 without Retry-After means: the persistent lane cap¶
A 503 without Retry-After means ingester.maxLanes refused a novel
(tenant, index) lane after the process reached its cap. The body starts
Service Unavailable: ingester lane limit reached (<N> lanes) and directs you
to route traffic elsewhere or raise --ingest-max-lanes. OTLP/gRPC returns
Unavailable without retry-after metadata or RetryInfo. Admitted lanes
still accept writes.
An empty queue does not release its lane. Recovery requires routing or operator
action: send traffic
to a pod with capacity, add a pod, or raise the startup cap and restart. Correct
clients that vary X-Scope-OrgID or x-siglake-index.
The maxTenants cap instead returns HTTP 403
or gRPC PermissionDenied for a novel tenant. Queue backpressure returns
503 with Retry-After: 1. A writer failure remains 500 or
Internal. Siglake's pinned opentelemetry-otlp 0.32.0 exporter makes one HTTP
attempt because experimental retry is off. A retry-enabled build makes three
retries before reporting loss. A bounded retry can exhaust against the full
pod. This result does not describe every client.
Configure transient ingest limits¶
On OTLP/gRPC all three HTTP 503 conditions map to UNAVAILABLE. Queue and
memory refusals include a plain retry-after metadata entry; the lane-cap
refusal does not. A breaker refusal on gRPC counts the same as one on HTTP: one
increment each of
siglake_ingest_mem_breaker_rejected_total and
siglake_ingest_requests_total{status="503"}, with endpoint set to otlp for
the logs exporter and otlp_traces for the traces exporter.
| Helm value | Flag | Default | What it sets |
|---|---|---|---|
ingester.backpressure.capacity |
--ingest-backpressure-capacity |
1024 |
Commands one lane queues before it refuses. 0 selects the older mutex path, which has no backpressure signal. |
ingester.backpressure.shards |
--ingest-backpressure-shards |
1 |
Writer tasks per tenant, each with its own lane of that capacity. |
ingester.memBreaker.limitMib |
--ingest-mem-limit-mib |
0, the breaker off |
RSS ceiling for the ingest process. Set it below the container memory limit. |
ingester.memBreaker.sampleSecs |
--ingest-mem-sample-secs |
60 |
How often RSS is read, so how late the breaker can be. |
Configure your exporter to honor Retry-After when the response includes it.
For the steps to tune these limits after a refusal, see Ingest is
rejecting;
Monitoring covers the queue
depth to compare rejections against.
Query routing¶
Without a configured tenant claim, every accepted query uses the default namespace. This includes open authentication, static bearer tokens, and OIDC authentication without tenant routing.
With query.oidc.tenantClaim configured, the query server applies the ingest
identifier rules and selects the matching IcebergContext. Missing or invalid
claims return 403 before query planning.
query.allowedTenants is an exact allow-list for verified query claims. Its
default, [], accepts every usable claim. A non-empty list requires
query.oidc.tenantClaim. The query list never reads or inherits
ingester.allowedTenants: stopping new writes for a tenant does not prevent
reads during its retention window.
An unlisted tenant gets 403 before the query server creates a namespace or
tenant context. This applies to direct query requests and to
coordinator-authenticated /api/v1/sql/shard requests. If a worker has a more
restrictive list than its coordinator, the coordinator forwards the worker's
403 to the caller.
siglake_query_tenant_denied_total{reason} distinguishes missing, invalid and
unlisted claims. Monitoring lists
the operator action for each reason.
Authentication is separate¶
| Mode | Configuration | Tenant effect |
|---|---|---|
| Open | Default | No identity check. Routing still follows the configured tenant mode. |
| Bearer token | --auth-tokens for ingest, --tokens for query |
Tokens are cluster-wide and do not select a tenant. |
| OpenID Connect (OIDC) | --oidc-issuer and --oidc-audience |
Verifies the JWT. --oidc-tenant-claim can then bind routing to a claim. |
OIDC takes precedence when it is configured. A valid token proves identity, but it provides tenant routing only when the tenant claim option is also set. See Security for configuration.
What is not isolated¶
All tenants share these resources within one deployment:
- Ingest, compactor, and query CPU and memory.
- The SQL catalog.
- The warehouse bucket.
- Object-storage credentials.
Ingest rate budgets and query guardrails limit some work, but there is no general per-tenant compute quota. A busy tenant can affect another tenant's latency, and a tenant that drives one replica's RSS over the ceiling trips the memory breaker for every tenant on it. Namespace prefixes are not an Identity and Access Management (IAM) boundary.
Use separate deployments when tenants need separate credentials, capacity, or failure domains.
Operational notes¶
Header-only mode creates tenant state from accepted header values. A typo can
therefore create a new namespace unless --allowed-tenants or --max-tenants
stops it.
Tenant count also increases WAL directories and writer state. The lane limit bounds the tenant and index combination, but very high tenant counts remain an untested deployment shape.
Schema migration and retention operate per table. A multi-tenant schema
migration needs --all-namespaces so it reaches existing tenant namespaces.
The configuration reference
lists every tenant setting.