HTTP API reference¶
Siglake rename
Product names, commands and links were normalized during the Siglake migration. Retained generation dates and commit IDs below identify the archived pre-rename source and binaries; they are not new build evidence.
The ingester exposes HTTP on port 8088 by default. The query server uses port 8089. This page lists their endpoints, request fields, response fields and status codes.
Machine-readable spec
The checked-in OpenAPI 3.1 specifications for the ingester API
and query API
define the endpoint tables on this page.
The tables were reconciled against Siglake commit 8a517387 on 2026-10-02.
Regenerate the specifications from a Siglake checkout with
cargo run -p siglake-openapi -- --out docs/api. The servers do not expose
these build artifacts at /openapi.json.
Ingester endpoints¶
| Method | Path | Summary |
|---|---|---|
GET |
/ |
Elasticsearch version fingerprint. |
HEAD |
/ |
Reachability probe used by ES clients that HEAD / before sending. |
GET |
/_cluster/health |
Cluster health, always green. |
POST |
/api/v1/_elastic/_bulk |
Elasticsearch _bulk ingest. |
GET |
/api/v1/_elastic/_cluster/health |
/api/v1/_elastic-prefixed alias of GET /_cluster/health. |
POST |
/api/v1/_elastic/{index}/_bulk |
Elasticsearch _bulk ingest with a default index from the path. |
GET |
/api/v1/stream |
GET /api/v1/stream — Server-Sent Events tail of the ingest path. |
GET |
/healthz |
Liveness probe. |
GET |
/readyz |
Readiness probe: can this ingester still durably accept writes? |
POST |
/v1/logs |
Ingest OTLP logs. |
POST |
/v1/traces |
Ingest OTLP traces. |
Commit acknowledgement modes¶
All four ingest routes, POST /v1/logs, POST /v1/traces,
POST /api/v1/_elastic/_bulk, and
POST /api/v1/_elastic/{index}/_bulk, accept either the commit query
parameter (for example, ?commit=wait_for) or the X-Siglake-Commit header.
The query parameter takes precedence when both are present.
| Value | Acknowledgement boundary |
|---|---|
auto |
Returns after the batch has reached the WAL through write(2); the acknowledgement does not require fsync(2). The write survives a process crash, OOM kill, or pod restart, but can be lost in a node crash or power failure before the kernel flushes it. |
wait_for (default) |
Returns after the segment bytes and the directory entry that names the active segment are synced. It does not wait for the rows to become queryable. |
force |
Seals the segment, starts an immediate commit, and waits until this ingester observes it. The server refuses this mode with 400 when a remote catalog-claim drain consumes the mirrored WAL. Otherwise it returns 504 if the commit is not observed within SIGLAKE_COMMIT_FORCE_TIMEOUT_SECS (30 seconds by default); the rows may still commit, so retrying can create duplicates. |
The wait_for power-loss guarantee assumes ext4 or xfs on a node-attached
volume. A network filesystem provides whatever its fsync(2) and rename
semantics guarantee. The WAL also syncs new tenant and index directories and
the sealed segment name before unlinking the active copy.
An unknown value returns 400. A successful wait_for response means the rows
are fsynced on the local WAL. It waits for neither the Iceberg commit that makes
them queryable nor the mirror upload of their segment. With an object-store
warehouse, the rows get their first copy off the WAL volume at whichever of
those two completes first: the commit writes their Parquet files before it
publishes the snapshot, so a committed row has a warehouse copy with the mirror
off or behind.
If a remote catalog-claim drain consumes the mirrored WAL, send
commit=wait_for and then query for the rows to confirm they are visible. The
durability model explains why
that drain cannot support force and which drains can.
Elasticsearch compatibility¶
Siglake ingests over the Elasticsearch bulk API but has no Elasticsearch query
API, and none is planned. Query with SQL over HTTP instead: see the SQL
cookbook. The read routes (_search, _msearch,
_search/scroll, _field_caps and _cat/*) are registered and answer 501
with a pointer to POST /api/v1/sql. The table above omits them, and so does
the OpenAPI specification: a routed 501 is not a surface to build on.
POST /v1/logs¶
Standard OTLP ExportLogsServiceRequest. Content-Type selects the encoding:
application/json or application/x-protobuf.
Headers:
| Header | Meaning |
|---|---|
X-Scope-OrgID |
In the default single-tenant mode, absent or default routes to the default tenant; any other value gets a 403. With --trust-scope-header, the header selects the tenant and absent means default. With --oidc-tenant-claim, the verified JWT claim selects the tenant and a present header must agree. See the ingest routing contract and generated --oidc-tenant-claim help. |
X-Siglake-Commit |
Acknowledgement mode. See Commit acknowledgement modes. |
Authorization: Bearer <token> |
Required when auth is enabled. Bearer is the only accepted authorization scheme. |
Responses:
| Status | Meaning |
|---|---|
200 |
Accepted at the requested commit acknowledgement boundary. This does not by itself mean the rows are queryable. |
400 |
Invalid request or commit mode, or commit=force while a remote catalog-claim drain consumes the mirrored WAL. |
401 |
Auth enabled and the token is missing or invalid. |
403 |
Tenant refused: this ingester is single-tenant and the header named another tenant, or the header disagrees with (or the token lacks) the JWT tenant claim, or the resolved tenant is outside --allowed-tenants, or this ingester is at --max-tenants and the resolved tenant is one it has not admitted before. |
413 |
Body over SIGLAKE_INGEST_MAX_BODY_BYTES (16 MiB default). |
429 |
Rate budget exhausted. Honor Retry-After. |
503 |
Backpressure lane full or memory breaker open. Honor Retry-After. |
504 |
A commit=force commit was not observed before SIGLAKE_COMMIT_FORCE_TIMEOUT_SECS. The rows may still commit. |
POST /api/v1/_elastic/_bulk¶
The Elasticsearch-compatible POST /api/v1/_elastic/_bulk endpoint maps bulk
index and create actions onto Siglake ingest semantics. Each bulk action
selects a target index and supplies one document. Siglake validates each bulk
document, maps accepted bulk actions to the ingest path and applies the same
ingest commit modes as POST /v1/logs. A bulk 200 can contain item failures.
Inspect each bulk item before accepting the bulk request as successful.
Bulk request format¶
NDJSON, alternating action and document lines:
{"index":{"_index":"app-logs"}}
{"@timestamp":"2026-07-28T12:00:00Z","message":"hello","level":"info"}
Documents are validated against the target index's document mapping.
Both this route and POST /api/v1/_elastic/{index}/_bulk accept the
commit query parameter and X-Siglake-Commit header described in Commit
acknowledgement modes. A bulk 200 can still
contain item-level failures; inspect errors and items in the response.
GET /api/v1/stream¶
GET /api/v1/stream delivers live ingest events as Server-Sent Events (SSE).
The stream keeps no result history or resume cursor. After a disconnect, a
client cannot resume the old stream; it must reconnect and receives only new
events. The route has no server-side filter and drops clients that fall behind.
Stream durability¶
The ingest path publishes each event to this route before appending it to the
write-ahead log (WAL). A later 503 refusal or 500 append failure can leave
subscribers with an event that Siglake did not store. A retry publishes that
event again. See Ingest and the WAL
for acknowledgement boundaries.
Query server endpoints¶
| Method | Path | Summary |
|---|---|---|
GET |
/api/v1/delete-tasks |
List delete tasks and their progress. |
POST |
/api/v1/delete-tasks |
Submit a delete task (GDPR / targeted deletion). |
GET |
/api/v1/delete-tasks/{id} |
Fetch one delete task, including rows deleted and files rewritten. |
GET |
/api/v1/index-templates |
List every index template in the caller's tenant namespace. |
DELETE |
/api/v1/index-templates/{id} |
Delete an index template from the caller's tenant namespace. |
PUT |
/api/v1/index-templates/{id} |
Create or replace an index template. |
GET |
/api/v1/indexes |
List every index in the caller's tenant. |
POST |
/api/v1/indexes |
Create an index. |
DELETE |
/api/v1/indexes/{id} |
Delete an index. |
GET |
/api/v1/indexes/{id} |
Fetch one index configuration. |
PUT |
/api/v1/indexes/{id} |
Replace an index configuration. |
GET |
/api/v1/jaeger/{index}/api/services |
Jaeger-compatible service-name list. |
GET |
/api/v1/jaeger/{index}/api/services/{service}/operations |
Jaeger-compatible operation-name list for one service. |
GET |
/api/v1/jaeger/{index}/api/traces |
Jaeger-compatible trace search. |
GET |
/api/v1/jaeger/{index}/api/traces/{trace_id} |
Jaeger-compatible single-trace lookup. |
DELETE |
/api/v1/jobs/{id} |
Cancel a batch job. |
GET |
/api/v1/jobs/{id} |
Poll a batch job's status. |
GET |
/api/v1/jobs/{id}/result |
Fetch a finished batch job's rows. |
POST |
/api/v1/sql |
POST /api/v1/sql entry point. |
POST |
/api/v1/sql/distributed |
Coordinator endpoint: fan req.query across the configured worker peers and merge into the single-pod answer. |
POST |
/api/v1/sql/explain |
Estimate a query's cost without running it. |
POST |
/api/v1/sql/local |
Run a SQL query on this pod only. |
POST |
/api/v1/sql/shard |
Worker endpoint: run the (sharded) query locally and return the result as an Arrow IPC stream (lossless transport). |
GET |
/debug/memory-pool |
Snapshot the process-wide DataFusion memory pool. |
GET |
/healthz |
Liveness probe. |
GET |
/readyz |
Readiness probe: verifies the Iceberg catalog is reachable. |
Every /api/v1/* response carries the x-siglake-server-micros header with
total server-side latency.
POST /api/v1/sql¶
Request body¶
{
"query": "SELECT host, count(*) FROM events GROUP BY host",
"format": "records",
"priority": "interactive",
"default_order": true,
"dry_run": false,
"limits": { "max_rows_returned": 1000 },
"shard": { "index": 0, "count": 4 }
}
| Field | Type | Default | Meaning |
|---|---|---|---|
query |
string | required | The SQL to execute. |
format |
records|ndjson |
records |
Response encoding. ndjson streams and bypasses the coordinator's metadata fast paths. |
priority |
interactive|batch |
interactive |
batch returns a job id. |
default_order |
bool | true |
Order a bare interactive SELECT newest-first. Set it to false to run the query exactly as written. See Implicit newest-first ordering. |
dry_run |
bool | false |
Return the cost estimate; execute nothing. |
limits.max_rows_returned |
int | server default | Per-request row cap, clamped by the server ceiling. |
shard |
object | None | Worker-only. Scan shard index of count. |
Records response¶
{
"columns": ["host", "count(*)"],
"row_count": 3,
"rows": [ ["web-01", 4211], ["web-02", 3980], ["db-01", 122] ],
"truncated": false,
"max_rows": null,
"cost": { },
"stats": { }
}
truncated and max_rows appear only when the row cap was hit. A truncated
records response returns HTTP 413; an ndjson stream stops at the cap and
emits a truncation marker. See NDJSON trailers.
Request and response name the same number differently. You request a row cap
with limits.max_rows_returned. The envelope reports the applied cap
as max_rows. max_rows_returned is the sole request spelling; there is no
alias. On /api/v1/sql, /api/v1/sql/local, /api/v1/sql/explain, and
/api/v1/sql/distributed, an unknown field inside limits is refused with HTTP
422. The most common mistake is max_rows. The JSON extractor refuses it
before planning, execution or batch enqueue, so no work is done and a batch
submission creates no job. The response is the extractor's plain-text message
naming the unknown field, not an ApiErrorBody JSON envelope.
Only the limits object is strict; unknown fields at the top level of these
request bodies are still ignored. /api/v1/sql/shard is not affected because
it takes a shard request and resolves the tier defaults. A known limit above
its tier ceiling is still clamped rather than refused, while an omitted or
empty limits still selects the tier defaults (see Row
caps).
cost contains the preflight estimate:
| Field | Meaning |
|---|---|
files_to_scan |
Files assigned after time-bounds pruning. |
files_considered |
Live files before pruning. The difference is planning's win. |
estimated_bytes_scanned |
Estimated bytes. |
estimated_rows_processed |
Estimated rows. |
estimated_runtime_seconds |
Estimated wall time. |
complexity_class |
small | medium | large | huge. |
exact |
true when the underlying stats were exact, false when heuristic. |
warnings |
Advisory strings. |
stats contains execution measurements:
| Field | Meaning |
|---|---|
rows_scanned |
Leaf rows emitted by a DataFusion data-file scan. A footer-sum or raw-page path bypasses this accounting and can also report 0; inspect served_by. |
bytes_scanned |
Leaf bytes read. |
spill_bytes |
Aggregation/sort spill. |
phases |
Wall-clock breakdown, microseconds. |
scan |
Per-request pruning attribution. Absent when no data-file scan ran. |
served_by |
Execution path that produced the answer. Use this with rows_scanned: metadata, footer-sum, and raw-page paths can all report zero rows scanned. |
stats.served_by values¶
| Value | Meaning |
|---|---|
scan |
The full DataFusion plan ran and scanned data files. |
tier1_inline |
Exact group counts came from the warm inline whole-table aggregate. |
tier1_wide |
Exact group counts came from the warm wide aggregate, including columns repaired by rebuild-group-counts. |
tier1_windowed_agg |
Exact windowed group counts came from the per-snapshot time-by-group aggregate, with at most the boundary ranges materialized. |
materialized |
Exact group counts were assembled per live file from footers, with raw-page decoding where a usable footer was absent. This can be slow even when rows_scanned is 0. |
sketch |
A bounded approximate heavy-hitter summary served the query; inspect the response's approximation object for its bounds. |
stats.phases¶
| Field | Meaning |
|---|---|
plan_micros |
Physical planning (single-pod) or logical planning + classification (coordinator). |
buffer_delta_micros |
WAL-buffer delta load. 0 if no buffer or drained. |
collect_micros |
Plan execution. Single-pod only. |
render_micros |
Arrow → JSON rendering. |
distributed.mode |
scan | ordered_scan | aggregate | ordered_aggregate | local fallback. |
distributed.shard_wall_micros |
Per-shard wall time, in shard order. Fan-out is concurrent, so the cost is the maximum rather than the sum. |
distributed.merge_micros |
Coordinator-side merge. |
stats.scan¶
| Field | Meaning |
|---|---|
files_planned |
Assigned file set after planning. |
files_read |
Files actually opened. An early-stopped browse opens fewer. |
files_pruned_bloom |
Files skipped whole by the trigram/token bloom after the footer read. |
planned_bytes / planned_rows |
Assigned totals. |
row_groups_considered |
Row groups in read files. |
row_groups_pruned_bloom |
Skipped by bloom. |
row_groups_pruned_stats |
Skipped by Parquet statistics. |
row_groups_read |
Actually decoded. |
rows_pruned_selection |
Rows skipped before decode (page index, deletes, inverted index). |
object_store_reads |
Object-store GET count. |
fetched_bytes / decoded_bytes |
Bytes over the wire vs. decoded. |
bytes_footer / bytes_index / bytes_data / bytes_other |
Reader-side fetched bytes for Parquet footers; page, offset, and column indexes plus Siglake index blobs; requested column-chunk data and bytes between requested ranges that the reader coalesces into one request; and reader-side manifest or otherwise unclassified reads, respectively. Together they sum to fetched_bytes; planning-time manifest reads use a separate path and appear in neither. Each class is omitted when zero. |
file_cache_hits |
File tasks served from the decoded file-batch cache. A hit builds no reader, so it counts in none of files_read, row_groups_read, object_store_reads, or the byte fields. Omitted when zero. |
file_cache_misses |
File tasks that consulted the decoded file-batch cache, missed, and read the file while populating the cache. A task carrying a predicate the Iceberg converter accepts, a raw prune spec or a promoted prune spec bypasses the cache: it counts in file_cache_bypasses, not in file_cache_hits or file_cache_misses. Omitted when zero. |
file_cache_bypasses |
File tasks that consulted the decoded file-batch cache and declined to populate it because they carry a converted predicate, a raw prune spec or a promoted prune spec. They read with their filtering intact, as they would on a cache-disabled install. Without this field, a response with no population reads the same whether nothing was decoded or every task was ineligible. Omitted when zero. |
file_cache_populate_rows |
Rows this request's populations were handed before they stopped, summed over the tasks that populated. Counted before the residual filter above the scan, and it keeps rising after a candidate crosses the entry bound, so it measures how deep the read got rather than what the cache kept or the query returned. The process histogram siglake_query_scan_file_cache_populate_rows splits the same quantity by outcome. Omitted when zero. |
file_attribution |
Bounded file-task membership and decoded-cache outcomes. Added in Siglake 0.2.0. |
unsettled_partitions |
Scan partitions still unwinding when the block was built. When non-zero, every count above is missing those partitions' reads. Normally omitted after the server waits for early-stopped partitions; it appears only if that wait reaches its deadline. |
ordering |
advertised when the scan streamed sorted (ordered LIMITs early-stop), otherwise the refusal reason (filtered, fan_in, no_bounds, …). |
Siglake 0.2.0 adds file_attribution. Its files array retains at most 32
entries across the request after coordinator fan-in. The table,
table-relative object_key, start and length fields identify one file task.
The three boolean fields record its cache and reader outcomes.
files_omitted counts identities excluded from the retained list. If
identity_complete is false, a consumer must refuse every file-membership
claim instead of treating the retained list as complete.
stats.scan.file_attribution fields¶
| Field | Meaning |
|---|---|
file_attribution |
Bounded file-task membership and decoded-cache outcomes for this request. |
file_attribution.files |
Retained file-task entries. |
file_attribution.files[].table |
Iceberg table identifier. |
file_attribution.files[].object_key |
Path relative to the Iceberg table location. |
file_attribution.files[].start |
Start of the file task's byte range. |
file_attribution.files[].length |
Length of the file task's byte range. |
file_attribution.files[].cache_candidate |
Whether the task consulted the decoded file-batch cache. |
file_attribution.files[].reader_opened |
Whether the request opened a reader for the task. |
file_attribution.files[].cache_hit |
Whether the decoded cache supplied data for the task. |
file_attribution.files_omitted |
Number of file-task identities omitted from files. |
file_attribution.identity_complete |
Whether files contains the request's complete file-task membership. |
The decoded file-batch cache ships off, with
SIGLAKE_QUERY_SCAN_FILE_CACHE_MAX_BYTES and
SIGLAKE_QUERY_SCAN_FILE_CACHE_MAX_ENTRIES both 0, so all four
file_cache_* fields are absent from a default install's response.
Implicit newest-first ordering¶
An interactive SELECT that names one timestamp-bearing table and asks for no
ordering of its own is given ORDER BY timestamp DESC. That is what a log
reader means by SELECT timestamp, raw FROM events LIMIT 100, and it is what
puts the query on the ordered early-stop path instead of a file-order scan.
default_order: false on the request runs the query exactly as written.
Eligible tables are events, query_audit, and any managed index. A managed
index orders on the field its doc mapping declares as timestamp_field.
Siglake quotes a non-canonical field name in the injected clause, so names that
contain uppercase letters or collide with SQL keywords keep their mapping
identity.
| Input | What the server does |
|---|---|
Bare SELECT over one eligible table |
Appends ORDER BY timestamp DESC. |
Explicit ORDER BY, GROUP BY, an aggregate, a CTE, a join, DISTINCT |
Runs it as written. |
A projection that aliases another column to timestamp, such as SELECT raw AS timestamp |
Runs it as written, in file order. |
EXPLAIN, or priority: "batch" |
Runs it as written. |
Rewritten query with an explicit LIMIT |
Keeps that LIMIT. |
Rewritten query with no LIMIT |
Adds max_rows_returned + 1, so the truncation signal still fires. |
Rewrites are counted by siglake_query_default_order_applied_total (see
metrics). The scan reports whether it streamed sorted in
stats.scan.ordering.
The rewrite asks for the early stop; storage decides whether it happens. A scan advertises an ordering only when the field it was asked to order on is the table's identity sort lead in the current schema and reads as an Arrow timestamp. A mismatched, unknown or otherwise typed field keeps DataFusion's blocking sort, which returns the same rows without the early-stop saving.
Rows that share a custom event time have no tiebreak, so a LIMIT boundary can
cut an equal-time group anywhere. See
Limitations.
NDJSON trailers¶
POST /api/v1/sql emits NDJSON result rows followed by at most one _meta
trailer. The NDJSON trailer fields identify truncation, a row-scan limit or a
query that failed mid-stream. Detect a failed query by reading the final NDJSON
line and checking its _meta value before accepting the streamed results.
NDJSON trailer fields¶
An ndjson stream commits its 200 status with the first bytes. Later errors
therefore appear in the body. The stream ends with at most one trailer: a JSON
object containing a _meta field. Result rows never contain _meta. Inspect
the final line before treating a short stream as complete.
_meta |
Extra fields | Meaning |
|---|---|---|
truncated |
row_count, max_rows |
The per-request row cap was reached. The rows already emitted are complete; the result is a prefix. |
midflight_rows_scanned_exceeded |
rows_scanned, max_rows_scanned |
The mid-flight rows-scanned breaker tripped and the stream was cut. Equivalent to the 413 a records request would have received. |
error |
code, error, and retry_after_secs on a pool refusal |
Execution failed after the response started. |
NDJSON trailer examples¶
One of these lines, never more than one, follows the last row:
{"_meta": "truncated", "row_count": 1000, "max_rows": 1000}
{"_meta": "midflight_rows_scanned_exceeded", "rows_scanned": 51200000, "max_rows_scanned": 50000000}
{"_meta": "error", "code": 503, "error": "query memory pool exhausted; retry after 5s: …", "retry_after_secs": 5}
The error trailer's code is the status the request would have carried had it
failed before streaming. A query memory-pool or spill-cap refusal uses 503,
adds retry_after_secs (5, the same delay the Retry-After header carries),
and increments siglake_query_breaker_trips_total{breaker="pool_exhausted"}.
Handle it like a 503 status. See Queries return
503. Every other
execution failure uses 500 and omits retry_after_secs.
POST /api/v1/sql/explain¶
Use POST /api/v1/sql/explain to inspect a DataFusion physical plan without
executing the SQL query. Send the same request body as POST /api/v1/sql. The
explain response includes the DataFusion plan and preflight cost report, but no
query rows. This endpoint is equivalent to dry_run: true on the main route.
Batch jobs¶
Create a batch job with POST /api/v1/sql and "priority": "batch". Poll the
batch job at GET /api/v1/jobs/{id} while its state is pending or running.
Before results are available, the job must reach succeeded; fetch them from
GET /api/v1/jobs/{id}/result. Other terminal states are failed,
cancelled and timeout.
Batch job admission¶
Submit with "priority": "batch". An admitted submission returns 202 with
the job id. If the query admission budget remains full through
SIGLAKE_QUERY_ADMISSION_WAIT_MS, the server creates no job and returns 429
with Retry-After.
| Status | Meaning |
|---|---|
pending |
Accepted but not yet executing. |
running |
An executor has published ownership. |
succeeded |
The result route returns the stored rows. |
failed |
Execution or result storage failed. |
cancelled |
A cancellation reached the job row before another terminal state. |
timeout |
The query exceeded its wall-clock deadline. |
Each admitted batch job reserves one share of the pod's admission budget: the
budget divided by
SIGLAKE_QUERY_ADMISSION_MAX_SHARE_DIVISOR
(default 4). It holds the share until the job finishes or a cancellation reaches the replica
running it.
After admission:
curl -s http://localhost:8089/api/v1/jobs/$ID # status
curl -s http://localhost:8089/api/v1/jobs/$ID/result # rows
curl -sX DELETE http://localhost:8089/api/v1/jobs/$ID # cancel
Batch job cancellation¶
DELETE cancels a pending or running job and answers 202 with the cancelled
job id. That 202 means the cancellation is persisted and terminal: the row
reads cancelled from every query replica, and a completion the executor
writes afterwards is refused. It does not mean the work has already stopped.
The abort handle that releases the query-admission reservation and cancels the
storage scans is process-local. When the replica serving the DELETE is the
one executing the job, the abort is immediate. Otherwise, the executor polls
for cancellations. With
query.jobs.persistent
and more than one query pod, it re-reads the status of its own
in-flight jobs every --jobs-cancel-poll-secs /
SIGLAKE_JOBS_CANCEL_POLL_SECS (default
2 seconds) and aborts the ones that read cancelled. There is no
LISTEN/NOTIFY fast path, so the 202 bounds when the work stops rather
than reporting that it already has.
Batch job terminal conflicts¶
Whichever terminal state lands first wins. A job that already reached one
(succeeded, failed, cancelled, or timeout) answers 409; a run that
finishes just after a cancellation has its result discarded and reports no
succeeded outcome. Recovery rechecks the owner's lease in the conditional
write that changes a job to failed, so a heartbeat renewed after candidate
selection prevents condemnation. If recovery wins first, that failed state
remains terminal, and what the executor loses depends on where it was when the
store told it so.
The executor publishes running before it does any work, and that write is
conditional on the row still reading pending or running. When the
publication is refused because recovery, cancellation or TTL deletion won,
the run stops before execution. The query is dropped unpolled, so
nothing is planned, no scan starts, no terminal write is attempted, and the
admission reservation is released. There is no computed output to discard and
no message is added to the job, so its client-visible error stays the one the
winner wrote. A job-store error on that publication is not evidence of a
terminal verdict, so the run executes as before and lets its terminal write
discover the authoritative state.
When the refusal instead meets a run that already finished, the completion is
the thing that is refused: a late executor completion cannot reinstate the job,
and its computed output is discarded. The job's client-visible error then
includes a result for this job was computed after recovery declared its
executor gone; it was discarded, so resubmit the query.
Aborts driven by a peer's cancellation count on
siglake_query_jobs_cancel_propagated_total, and both refusal shapes count on
siglake_query_job_terminal_conflict_total{attempted,actual,cause}:
attempted="running" is the refused start that executed nothing, and
attempted="succeeded", "failed" or "timeout" is a refused completion
whose output was discarded. See Query metrics for what each
cause value means and why neither counter takes a plain rate threshold.
An unknown id, or a job owned by another tenant, answers 404 on all three
job routes.
GET /api/v1/jobs/{id}/result answers 404 while the job is still pending or
running and 409 when it ended in a non-success state; poll
GET /api/v1/jobs/{id} to tell those apart.
Batch job store failures¶
A transient failure of the persistent job store, such as a lost connection,
a pool timeout, a serialization failure or deadlock, a Postgres that is
restarting, answers 503 with Retry-After: 5 on all three job routes
instead. The store never answered, so the job's existence and state are
unknown: a 503 is not a 404 and not a 409, and the correct response is to
retry the same request unchanged. Any other store error is a 500.
Batch job terminal-state reconciliation¶
An outage that catches a finished run rather than your request shows up
differently: as a running status that outlives the query. The terminal write
of a finished run makes three attempts within 30 seconds. If none lands,
the run ends anyway, releases its admission share, and hands the job id to its
own executor. That executor retries the same conditional write every five
seconds until it lands. The row reads running the whole time, so a poll loop
can see running long after execution stopped. The five seconds is the delay
between passes, not a bound on how long this lasts: every pass needs a job
store that answers, so the row stays running for as long as the outage does.
Poll to a deadline rather than until a terminal state, and read running as
"no verdict yet", not as "still executing".
Nothing is re-executed and no result body is republished when a pass finally lands, so the state it installs comes from a fixed set:
| Verdict the run reached | Installed status | error |
|---|---|---|
succeeded |
failed |
this job finished but its result could not be stored (the job store was unavailable); the result was discarded, so resubmit the query |
failed |
failed |
this job failed and its error could not be stored (the job store was unavailable); resubmit the query to see why |
timeout |
timeout |
query exceeded the timeout (recorded by its executor after the job store became reachable again) |
The first row is the one to handle: the query did run and did produce rows,
but they were discarded rather than held through the outage, the result route
answers 409 like any other non-success, and the only way to get that answer
is to submit the query again. The second keeps the status and loses only the
run's own error text. Neither is a new execution, so neither costs anything a
resubmit would not.
That reconciling write is conditional on the row still reading pending or
running, exactly like the run's own terminal write. A terminal state that won
in the meantime therefore stands, and its error is what you read. A
cancellation still reads cancelled, and a recovery-condemned row still reads
failed with batch executor gone: …. A row the TTL sweep already removed
answers 404.
One case never reconciles. An executor already tracking 1,024
finished-but-unpersisted ids refuses the next one, and that row reads running
until the pod exits and a peer condemns it on lease expiry as failed with
batch executor gone: the replica running this job stopped heartbeating.
Operators page on that (SiglakeBatchRowStrandedNonTerminal, see Suggested
alerts); a client sees a longer
wait and then that recovery text.
Batch job row caps¶
The per-request row cap (limits.max_rows_returned, clamped by the server
ceiling) also applies to batch execution. A finished batch job still returns
200 from the result route when it hits that cap, but its RecordsResponse
sets truncated to true and reports the cap in max_rows. Check those
fields before treating rows as the whole answer; unlike a synchronous
truncated records response, the batch result does not return 413. A job
that sets no cap of its own gets
the batch tier's default, which is far below the interactive one and is not
raised by SIGLAKE_QUERY_MAX_ROWS. See Row
caps.
Batch queries run on a dedicated runtime, separate from interactive execution, but both tiers share the query admission budget.
Index management¶
# create
curl -sX POST http://localhost:8089/api/v1/indexes \
-H 'Content-Type: application/json' \
-d '{
"index_id": "app-logs",
"doc_mapping": {
"mode": "dynamic",
"timestamp_field": "timestamp",
"field_mappings": [
{"name": "timestamp", "type": "datetime", "required": true},
{"name": "message", "type": "text", "tokenizer": "default"},
{"name": "level", "type": "text", "tokenizer": "raw"},
{"name": "status", "type": "long"}
],
"tag_fields": ["level"],
"default_search_fields": ["message"]
},
"retention": {"period_secs": 2592000},
"index_at_flush": false
}'
See the user indexes guide for field types, tokenizers, and mapping modes.
Managed-index ETag and If-Match preconditions¶
GET /api/v1/indexes/{id} and a successful PUT return the full index
configuration with a strong ETag header. The validator covers the full
configuration and the table incarnation. Data appends, compaction, snapshot
expiry, and other data-only commits keep it unchanged. A configuration update
changes it. Deleting and recreating an index changes it even if the id and
configuration are identical.
PUT /api/v1/indexes/{id} accepts an optional RFC 9110 If-Match header. The
value can be * or a comma-separated list of entity-tags. A list matches when
any strong tag equals the current ETag; weak tags do not strongly match. *
matches any existing index. Malformed syntax returns 400.
If the condition is false, the server returns 412 Precondition Failed. The
response ETag matches current, which is the configuration from the exact
commit base that rejected the update:
{
"error": "managed index `app-logs` changed since the supplied If-Match value",
"code": 412,
"current": {
"index_id": "app-logs",
"doc_mapping": {
"mode": "dynamic",
"timestamp_field": "timestamp",
"field_mappings": [
{"name": "timestamp", "type": "datetime", "required": true}
],
"tag_fields": [],
"default_search_fields": []
},
"retention": null,
"index_at_flush": null
}
}
Another writer can replace that configuration before the response arrives.
Treat the body and ETag as one retry base, not as a lock. A later PUT using
the returned tag can therefore receive another 412.
A PUT without If-Match keeps the existing behavior. Schema changes remain
additive only whether the header is present or absent. A matching tag does not
permit a drop, rename, reorder, retype, or required-field addition; those
invalid mappings return 400.
Delete and recreate a managed index¶
Before DELETE /api/v1/indexes/{id} removes the catalog entry, it writes and
reads back a cleanup record under
_siglake/config/dropped_indexes/<namespace>/. The record pins the loaded
table's UUID, location, retained committed-file inventory and UUID-scoped
aggregate prefix. The index then disappears from GET /api/v1/indexes and
stops answering queries.
Production-created cleanup records are report-only. The aggregate sweep can inventory the recorded prefix, but it cannot delete from it. Retention and orphan GC cannot reclaim the committed files because both resolve the table through the catalog entry that the delete removed. Dropping an index does not free storage. Account for those bytes until Siglake exposes reviewed cleanup authority.
Creating an index with the same id again reuses that table location but writes a fresh table UUID. The replacement exposes only its own rows, while the dropped incarnation's files sit alongside them under the same prefix. Any out-of-band cleanup must therefore be pinned to the dropped table's UUID and file inventory, never to the location or the index id: a prefix-wide delete against a recreated id destroys the replacement's live data.
The write-ahead log (WAL) also remains under the tenant and index name. Each
current segment identifies its Iceberg table by UUID. If an ingester keeps a
writer open across the delete and recreation, all WAL-buffer query paths skip
the old table's segments. This includes local reads, count fast paths, and
distributed fan-out. The filesystem drain moves those segments to
stale/<dropped-uuid>/ instead of committing them to the replacement.
With a catalog-claim drain, the mirror prefix is re-stamped for the replacement and each object is checked against its own frame UUID. Objects from the dropped table stay in the mirror and their claims enter quarantine. After a prefix has named a dropped table, an object without its own UUID is also quarantined.
For compatibility, an unmarked local directory or a segment without a UUID has no opinion and serves as before. The mirror exception above applies after an index recreation because its re-stamped prefix can no longer identify an older, UUID-less object.
An ingest lane checks the table UUID when it opens, then refreshes it every 30
seconds by default. A write accepted after a recreation but before that refresh
returns 200, then goes to quarantine instead of Iceberg. The row does not
become queryable. Set
SIGLAKE_WAL_IDENTITY_REFRESH_SECS to shorten this
window or disable refreshes.
Index templates auto-create indexes whose IDs match a glob pattern, so a
shipper can write to a new index without an explicit create call. PUT
/api/v1/index-templates/{id} takes template_id (which must equal the path
id), a non-empty index_id_patterns, a priority, a doc_mapping, and an
optional retention; the highest-priority match wins, ties breaking on the
lexicographically smallest template_id. See the worked
example for the full request,
where the auto-created index becomes visible, and what happens to writes that
match no template.
Delete tasks¶
curl -sX POST http://localhost:8089/api/v1/delete-tasks \
-H 'Content-Type: application/json' \
-d '{
"index_id": "app-logs",
"predicate_sql": "user_id = 12345",
"start_ts": "2026-01-01T00:00:00Z",
"end_ts": "2026-07-01T00:00:00Z"
}'
| Field | Required | Meaning |
|---|---|---|
index_id |
yes | Target index. |
predicate_sql |
yes | SQL predicate identifying rows to delete. |
start_ts / end_ts |
no | Bound the time range, so fewer files are rewritten. |
POST, the collection list and GET …/delete-tasks/{id} all return the
recorded task. The server sets every field below; none of them is a request
field.
| Field | Type | Always present |
|---|---|---|
created_at |
string (date-time) | yes |
end_ts |
string (date-time) or null | no |
error |
string or null | no |
files_rewritten |
integer | yes |
index_id |
string | yes |
predicate_sql |
string | yes |
rows_deleted |
integer (int64) | yes |
start_ts |
string (date-time) or null | no |
state |
pending, running, done, failed |
yes |
table_uuid |
string or null | no |
task_id |
string (uuid) | yes |
"Always present" is the spec's required set. The rest are null unless
something set them: error until the task fails, start_ts and end_ts unless
the request bounded the range, table_uuid on a task recorded before the
binding described below.
table_uuid is the Iceberg table the submission was accepted against: the index
incarnation the task is authorized to delete from. An index id is reusable, so
DELETE /api/v1/indexes/app-logs followed by a POST of the same id gives a
different table. The executor refuses a task whose index_id now resolves to a
different table: state becomes failed and error names both tables. A task
recorded before this binding existed carries a null table_uuid and is refused
the same way; nothing infers an incarnation from a reused name. A refusal
commits no snapshot, and error says so when the rewrite had already run and
left files no snapshot references. Recovery is the same resubmission as for any
failed task. See Delete tasks are bound to one index
incarnation.
Tasks are recorded in a ledger and executed by compactor sweeps. They are not
immediate. Execution requires compactor.deleteTasks: true, which defaults to off,
or a manual siglake delete-sweep. See Retention and
deletes.
GET …/delete-tasks/{id} reports claim data only for a pending task.
An executor create-only-writes a claim object beside a task before running it.
The single-task GET returns that object in a claim field beside the task
fields while state is pending. A done, failed or running task has no
claim field. Neither POST nor the collection route returns claim data.
| Field | Meaning |
|---|---|
present |
Whether the claim object was observed by this read. |
observed_at |
When this read observed it. Always set. |
claimant |
UUID of the process that wrote the claim, or null. |
claimed_at |
When the claim was created, or null. |
age_seconds |
Whole seconds from claimed_at to observed_at, floored at zero for clock skew, or null. Not execution duration. |
The claim data has these limits:
- A claim whose body does not decode is still
present. An empty, truncated or unparseable body, or one naming a differenttask_id, answerspresent: truewithclaimant,claimed_atandage_secondsallnull. Presence is what excludes the next executor, so the object counts even when its diagnostics do not. A store failure other than "not found" is a500, never apresent: false. present: falsedescribes only that read. An executor may claim the task immediately afterwards.- Presence, the UUID and age do not prove liveness or abandonment. A claim
hours old may belong to a rewrite still running or to
a process long dead; the response cannot tell them apart, and no value of
age_secondsmakes taking the task over safe.
Siglake does not run a terminal or stranded task again. This includes tasks in
failed, tasks stranded in running and pending tasks with an unreleased
claim. The API has no retry endpoint. To recover, post the same request fields
again under the same tenant identity. The
response carries a new task_id in pending, and the failed record and its
error are left unchanged. The endpoint is not idempotent (every POST creates
another task), it promises nothing about exactly-once execution, and each task
reports only what its own run rewrote, never the request's cumulative effect.
See Recovering a failed
task.
Jaeger shim¶
The HTTP subset Grafana's Jaeger data source renders. {index} is a
traces-shaped index.
The two name-list routes, GET …/api/services and
GET …/api/services/{service}/operations, share /api/v1/sql's result cache.
The first poll for an Iceberg snapshot executes the complete distinct-name
query. Repeats for that snapshot are cache hits. A new snapshot gets a new
cache entry and executes the query again. The cache stores a complete name list
only when its retained arena and its lookup string together fit the 512 KiB
per-entry cap. The route's derived name ceiling supplies the row bound; the SQL
cache's 128-row entry cap does not apply to these lists. An overlength entry is
returned whole but not cached, so every poll runs the query.
SIGLAKE_QUERY_RESULT_CACHE controls this cache; there is no Jaeger-specific
cache setting. Cache hits preserve the 413, 429, 503 and 504 behavior
of a miss.
The shared store still holds at most 256 entries and 4 MiB. That 4 MiB, and the
siglake_query_sql_result_cache_bytes gauge that
reports it, count each entry's body, the lookup string each entry is stored
under, and one pointer per recency marker. For /api/v1/sql that lookup string
is the normalized query text, so a long query is charged against the same
budget as the rows it returns. An entry whose body and lookup string together
exceed its caller's per-entry allowance (4 MiB for /api/v1/sql, 512 KiB for a
name list) is refused rather than stored and then evicted.
GET …/api/traces accepts service, operation, start, end, limit,
minDuration, maxDuration, and tags (a JSON object string).
All four routes share /api/v1/sql's per-pod interactive admission budget
(SIGLAKE_QUERY_ADMISSION_BUDGET_BYTES) and
query wall-clock timeout. Each request makes one admission reservation. If the
budget remains full through
SIGLAKE_QUERY_ADMISSION_WAIT_MS, the route
returns 429 with Retry-After; wait for that delay before retrying.
The wall clock starts before tenant resolution and includes index lookup,
table registration, and query execution. A trace search gets one deadline
across trace-id selection and span retrieval; span retrieval does not get a
fresh budget. Expiry during preparation or either query phase returns 504
and cancels the in-flight scan.
The Jaeger ceilings refuse the whole response instead of truncating it. The four routes have
trace, span-row, accumulated Arrow-byte, and distinct-name ceilings derived
from the admission reservation each request already takes. After a 2 MiB
per-request floor, the server conservatively prices each trace at 12 KiB, each
span row at 8.5 KiB, each Arrow byte at 4 bytes of peak render memory, and each
name at 1.4 KiB. The reservation is min(16 MiB, admission budget / share), so
the ceilings follow
SIGLAKE_QUERY_ADMISSION_BUDGET_BYTES, the
share divisor, and pod memory; there is no Jaeger-specific environment knob.
The trace ceiling is also capped at half the span-row ceiling because even a
one-span trace consumes a selection row and a span row. If admission is
disabled, the calculation uses the packaged minimum reservation rather than
becoming unbounded. The server's resolved interactive max_rows_returned,
clamped by SIGLAKE_QUERY_MAX_ROWS, still wins
when tighter.
A trace search's limit cannot exceed the derived trace ceiling. The server
checks it before taking the admission reservation, looking up the index, or
planning, and returns 400 with the requested and allowed values. Span rows
and Arrow bytes are bounded mid-flight at a batch boundary across both phases
of a search; the name-list routes similarly bound names and Arrow bytes. One
request spends one budget: for a search, the spans of every selected trace
together rather than each trace separately. Narrow service, operation,
start/end, tags, or limit when a result is refused.
Every result-ceiling refusal is a whole-response 413: no data, no partial
trace, and no Retry-After. /api/v1/sql can return a truncated body that
declares truncated and max_rows in its records envelope,
but Jaeger's data and total response cannot say that a trace is missing
spans. This surface also has no per-request limits object.
| Status | On these routes |
|---|---|
400 |
A trace search's limit exceeded the derived trace ceiling (or another request parameter was invalid). Checked before admission, index lookup, or planning. |
413 |
A span-row, Arrow-byte, distinct-name, or tighter interactive row ceiling was exceeded. Refused whole: no data and no Retry-After; narrow the request. |
429 |
The shared interactive admission budget stayed full through the admission wait. Honor Retry-After. |
503 |
The query memory pool or the spill byte cap refused an allocation, exactly as on the SQL routes. Honor Retry-After. |
504 |
The one interactive wall-clock budget expired, in preparation or in either query phase. |
There is no gRPC SpanReader. See Grafana and Jaeger.
GET /debug/memory-pool¶
A snapshot of the process-wide DataFusion memory pool on the pod that answers.
Use it to find what holds memory when queries return 503 or when
SiglakeQueryPoolReservedWhileIdle fires; see Queries return
503 and the
saturation alerts. The route is
authenticated like /api/v1/* and answers 401 without valid credentials.
| Field | Type | Meaning |
|---|---|---|
limit_bytes |
integer or null |
Process-wide DataFusion pool limit in bytes. null when the pool is unbounded. |
reserved_bytes |
integer or null |
Bytes currently reserved by DataFusion operators. null when the pool is unbounded. |
top_consumers |
string or null |
DataFusion's own report text for the ten consumers with the largest current reservations, ordered by reservation size. null when consumer tracking is disabled (SIGLAKE_QUERY_MEMORY_POOL_TRACK_CONSUMERS set to 0 or off) or the pool is unbounded. |
Authentication¶
| Mode | Ingester flag | Query flag |
|---|---|---|
| Bearer tokens | --auth-tokens / SIGLAKE_AUTH_TOKENS |
--tokens / SIGLAKE_QUERY_TOKENS |
| OIDC | --oidc-issuer + --oidc-audience |
same |
OIDC takes precedence when both authentication modes are set. On the query server,
--oidc-tenant-claim routes each query by the named verified JWT claim. Without that
option, query callers use the default namespace; X-Scope-OrgID does not enable
query-side header routing.
--allowed-tenants optionally limits those verified query claims. An empty list accepts
every usable claim. A non-empty list requires --oidc-tenant-claim and never inherits
the ingester's allow-list, so a tenant can remain readable after write admission stops.
An unlisted claim gets a 403 before the query server opens its namespace or context.
The ingester uses single-tenant routing by default. An absent X-Scope-OrgID header or
an explicit default value routes to the default tenant. Any other value gets a 403.
Set --trust-scope-header only behind a gateway that sets the header and strips the
client's value. In that mode, the header selects the tenant and an absent header means
default.
With --oidc-tenant-claim, the ingester uses the verified claim for every request,
including when --trust-scope-header is also set. A present X-Scope-OrgID header must
agree with the claim.
The ingester and query server read separate bearer-token settings. A token works on both only when you configure it in both settings. An OIDC deployment can use the same issuer and audience on both processes.
Configuring the claim makes it mandatory on both boundaries, and the value is validated
rather than repaired. A verified token whose claim is missing, blank, not a JSON string,
longer than 128 characters, or outside [A-Za-z0-9_-] gets a 403 before routing. The
query server refuses it in auth middleware, so the refusal reaches every authenticated
operation, including operations that read no tenant data. The ingester refuses it before
the batch creates a write-ahead log (WAL) lane, metric label or namespace. Siglake does
not rewrite the claim or fall back to the default namespace.
siglake_query_tenant_denied_total uses reason="claim_missing" when a verified token lacks the configured tenant claim, reason="claim_invalid" when the tenant claim is unusable, and reason="not_allowed" when the verified tenant claim is outside --allowed-tenants.
siglake_ingest_tenant_denied_total uses reason="header_not_trusted" when a header selects another tenant in single-tenant mode, reason="not_allowed" when the resolved tenant is outside --allowed-tenants, reason="at_capacity" when a novel tenant arrives after --max-tenants is full, reason="claim_missing" when a verified token lacks the configured tenant claim, reason="claim_invalid" when the tenant claim is unusable, and reason="header_mismatch" when the header contradicts the tenant claim.
Health and metrics endpoints are unauthenticated.
Error responses¶
Siglake HTTP APIs use two main error response body shapes. Query errors contain
error and code fields; ingester errors contain text and code. A
Retry-After header distinguishes retryable 429 responses and capacity,
shard-pin or job-store 503 responses. Terminal HTTP failures omit that
header. The exceptions appear below.
HTTP status codes¶
Query-server errors return JSON with an error string and numeric code.
Some errors add fields such as reason, pin or cost. Ingester errors use a
text string and numeric code; Elasticsearch-compatible routes retain the
Elasticsearch error shape. The 422 response below is the JSON extractor's
plain text instead. Common statuses:
| Status | Cause |
|---|---|
400 |
Malformed request or invalid SQL. |
401 |
Missing or invalid credentials. |
403 |
The credential verified but does not authorize the request. On either server, once --oidc-tenant-claim is configured: the tenant claim it requires is missing or is not a usable tenant id ([A-Za-z0-9_-], 1 to 128 characters). Retrying or refreshing the token cannot help; the identity provider has to mint the claim. On the query server, a usable claim outside its --allowed-tenants also returns 403 before namespace or context creation. On the ingester also: an X-Scope-OrgID that names a different tenant than the claim, or a tenant outside its own --allowed-tenants. |
413 |
Body too large; a records response that hit the row cap, with the truncated prefix and its truncated and max_rows fields; or a Jaeger read that exceeded a result ceiling and returned no data. |
422 |
On the four SqlRequest routes, a field inside limits is not a known per-request limit. The JSON extractor refuses the request before planning, execution or batch enqueue. Its plain-text body names the unknown field. Unknown top-level fields are ignored. |
429 |
On the ingester, the rate budget is exhausted; on the query server, the admission budget stayed full through the configured wait. Honor Retry-After. |
500 |
Internal error. |
503 |
Backpressure, an ingest memory breaker, query memory-pool or spill-cap exhaustion, an unresolvable shard pin (reason: "shard_pin_unresolved"), a transient job-store failure on the batch job routes (Retry-After: 5), or a not-ready dependency. |
504 |
Query wall-clock timeout. |
The 429 responses include Retry-After. Capacity, shard-pin and job-store
503 responses include it too. The status code and this header distinguish
them from terminal request errors such as 400, 403 and 413.
In a distributed query, a worker's refusal status is forwarded rather than
collapsed: a mid-flight row-ceiling trip returns 413, a rejected query
400, a query tenant refusal 403, a budget refusal 429, a memory-pool or
spill-cap refusal 503, an
unresolvable shard pin 503, and a timeout 504. For 503, the coordinator
re-attaches Retry-After: 5 and does
not fail over or re-run the query locally. Any other worker 5xx is reported
as 500. A worker fault is a cluster fault. A client
acts on them differently: 413 means "narrow the query", 503 means "retry
after the requested delay", and 500 means "retry, or page someone".
A refusal that happens after an ndjson response has started cannot change the
status: it arrives as an _meta trailer on an otherwise
successful 200.
reason: "shard_pin_unresolved"¶
The error reason shard_pin_unresolved means a worker could not resolve the
table generation pinned by the coordinator. A client should respond to this
retryable 503 error by waiting for Retry-After: 5 and retrying. The
response echoes the pin in pin, as table, snapshot_id and schema_id;
either id is null when the pin did not carry it. If the error persists,
adjust snapshot retention or the metadata-cache lifetime as described below.
Shard pin failure details¶
Not every 503 is capacity. The coordinator pins each shard request to the
generation it planned against: the snapshot id, and the Iceberg schema id in
the optional pin.schema_id field of the /api/v1/sql/shard body. A worker
refuses its fragment if it cannot resolve either half: the snapshot has
expired from the catalog, or the schema id is in no generation the worker's
metadata retains, and neither becomes visible after a metadata refresh. A
worker whose metadata merely runs ahead of the coordinator serves the pinned
historical schema instead of refusing. The worker does not answer from its own
current generation. The body carries reason: "shard_pin_unresolved" and a
pin object naming the table, the snapshot id and the schema id, alongside
Retry-After: 5, so it is distinguishable
from a memory-pool refusal without reading the message. The message names the
half that failed, because the operator response differs: a snapshot refusal
points at snapshot retention, and a schema refusal at a migrate-schema that
has not reached every replica's metadata cache. The coordinator forwards the
refusal rather than re-running the fragment locally, so the whole distributed
query is refused; /api/v1/sql/local still answers. A distributed answer is
from one generation or absent. Siglake does not merge results across
generations.
See One generation per fan-out.
The schema id is optional so that a rolling upgrade works in both directions. An older coordinator omits it and gets snapshot-only pinning. An older worker ignores it, so a mixed-version fleet keeps snapshot-only pinning until every replica runs the new image.
Misses are counted by siglake_query_shard_pin_total{outcome="miss"} and alert
as SiglakeQueryShardPinUnresolved. If they persist rather than clearing as
caches converge, raise
compactor.snapshotExpire.retainLast
or lower
SIGLAKE_ICEBERG_METADATA_CACHE_TTL_SECS; see
Queries return 503.