Skip to content

HTTP API reference

Siglake rename

Product names, commands and links were normalized during the Siglake migration. Retained generation dates and commit IDs below identify the archived pre-rename source and binaries; they are not new build evidence.

The ingester exposes HTTP on port 8088 by default. The query server uses port 8089. This page lists their endpoints, request fields, response fields and status codes.

Machine-readable spec

The checked-in OpenAPI 3.1 specifications for the ingester API and query API define the endpoint tables on this page. The tables were reconciled against Siglake commit 65c4eef on 2026-09-16. Regenerate the specifications from a Siglake checkout with cargo run -p siglake-openapi -- --out docs/api. The servers do not expose these build artifacts at /openapi.json.

Ingester endpoints

Method Path Summary
GET / Elasticsearch version fingerprint.
HEAD / Reachability probe used by ES clients that HEAD / before sending.
GET /_cluster/health Cluster health, always green.
POST /api/v1/_elastic/_bulk Elasticsearch _bulk ingest.
GET /api/v1/_elastic/_cluster/health /api/v1/_elastic-prefixed alias of GET /_cluster/health.
POST /api/v1/_elastic/{index}/_bulk Elasticsearch _bulk ingest with a default index from the path.
GET /api/v1/stream GET /api/v1/stream — Server-Sent Events tail of the ingest path.
GET /healthz Liveness probe.
GET /readyz Readiness probe: can this ingester still durably accept writes?
POST /v1/logs Ingest OTLP logs.
POST /v1/traces Ingest OTLP traces.

Commit acknowledgement modes

All four ingest routes, POST /v1/logs, POST /v1/traces, POST /api/v1/_elastic/_bulk, and POST /api/v1/_elastic/{index}/_bulk, accept either the commit query parameter (for example, ?commit=wait_for) or the X-Siglake-Commit header. The query parameter takes precedence when both are present.

Value Acknowledgement boundary
auto Returns after the batch has reached the WAL through write(2); the acknowledgement does not require fsync(2). The write survives a process crash, OOM kill, or pod restart, but can be lost in a node crash or power failure before the kernel flushes it.
wait_for (default) Returns after the segment bytes and the directory entry that names the active segment are synced. It does not wait for the rows to become queryable.
force Seals the segment, starts an immediate commit, and waits until this ingester observes it. The server refuses this mode with 400 when a remote catalog-claim drain consumes the mirrored WAL. Otherwise it returns 504 if the commit is not observed within SIGLAKE_COMMIT_FORCE_TIMEOUT_SECS (30 seconds by default); the rows may still commit, so retrying can create duplicates.

The wait_for power-loss guarantee assumes ext4 or xfs on a node-attached volume. A network filesystem provides whatever its fsync(2) and rename semantics guarantee. The WAL also syncs new tenant and index directories and the sealed segment name before unlinking the active copy.

An unknown value returns 400. A successful wait_for response means the rows are fsynced on the local WAL. It waits for neither the Iceberg commit that makes them queryable nor the mirror upload of their segment. With an object-store warehouse, the rows get their first copy off the WAL volume at whichever of those two completes first: the commit writes their Parquet files before it publishes the snapshot, so a committed row has a warehouse copy with the mirror off or behind.

If a remote catalog-claim drain consumes the mirrored WAL, send commit=wait_for and then query for the rows to confirm they are visible. The durability model explains why that drain cannot support force and which drains can.

Elasticsearch compatibility

Siglake ingests over the Elasticsearch bulk API but has no Elasticsearch query API, and none is planned. Query with SQL over HTTP instead: see the SQL cookbook. The read routes (_search, _msearch, _search/scroll, _field_caps and _cat/*) are registered and answer 501 with a pointer to POST /api/v1/sql. The table above omits them, and so does the OpenAPI specification: a routed 501 is not a surface to build on.

POST /v1/logs

Standard OTLP ExportLogsServiceRequest. Content-Type selects the encoding: application/json or application/x-protobuf.

Headers:

Header Meaning
X-Scope-OrgID In the default single-tenant mode, absent or default routes to the default tenant; any other value gets a 403. With --trust-scope-header, the header selects the tenant and absent means default. With --oidc-tenant-claim, the verified JWT claim selects the tenant and a present header must agree. See the ingest routing contract and generated --oidc-tenant-claim help.
X-Siglake-Commit Acknowledgement mode. See Commit acknowledgement modes.
Authorization: Bearer <token> Required when auth is enabled. Bearer is the only accepted authorization scheme.

Responses:

Status Meaning
200 Accepted at the requested commit acknowledgement boundary. This does not by itself mean the rows are queryable.
400 Invalid request or commit mode, or commit=force while a remote catalog-claim drain consumes the mirrored WAL.
401 Auth enabled and the token is missing or invalid.
403 Tenant refused: this ingester is single-tenant and the header named another tenant, or the header disagrees with (or the token lacks) the JWT tenant claim, or the resolved tenant is outside --allowed-tenants, or this ingester is at --max-tenants and the resolved tenant is one it has not admitted before.
413 Body over SIGLAKE_INGEST_MAX_BODY_BYTES (16 MiB default).
429 Rate budget exhausted. Honor Retry-After.
503 Backpressure lane full or memory breaker open. Honor Retry-After.
504 A commit=force commit was not observed before SIGLAKE_COMMIT_FORCE_TIMEOUT_SECS. The rows may still commit.

POST /api/v1/_elastic/_bulk

The Elasticsearch-compatible POST /api/v1/_elastic/_bulk endpoint maps bulk index and create actions onto Siglake ingest semantics. Each bulk action selects a target index and supplies one document. Siglake validates each bulk document, maps accepted bulk actions to the ingest path and applies the same ingest commit modes as POST /v1/logs. A bulk 200 can contain item failures. Inspect each bulk item before accepting the bulk request as successful.

Bulk request format

NDJSON, alternating action and document lines:

{"index":{"_index":"app-logs"}}
{"@timestamp":"2026-07-28T12:00:00Z","message":"hello","level":"info"}

Documents are validated against the target index's document mapping.

Both this route and POST /api/v1/_elastic/{index}/_bulk accept the commit query parameter and X-Siglake-Commit header described in Commit acknowledgement modes. A bulk 200 can still contain item-level failures; inspect errors and items in the response.

GET /api/v1/stream

GET /api/v1/stream delivers live ingest events as Server-Sent Events (SSE). The stream keeps no result history or resume cursor. After a disconnect, a client cannot resume the old stream; it must reconnect and receives only new events. The route has no server-side filter and drops clients that fall behind.

Stream durability

The ingest path publishes each event to this route before appending it to the write-ahead log (WAL). A later 503 refusal or 500 append failure can leave subscribers with an event that Siglake did not store. A retry publishes that event again. See Ingest and the WAL for acknowledgement boundaries.

Query server endpoints

Method Path Summary
GET /api/v1/delete-tasks List delete tasks and their progress.
POST /api/v1/delete-tasks Submit a delete task (GDPR / targeted deletion).
GET /api/v1/delete-tasks/{id} Fetch one delete task, including rows deleted and files rewritten.
GET /api/v1/index-templates List every index template in the caller's tenant namespace.
DELETE /api/v1/index-templates/{id} Delete an index template from the caller's tenant namespace.
PUT /api/v1/index-templates/{id} Create or replace an index template.
GET /api/v1/indexes List every index in the caller's tenant.
POST /api/v1/indexes Create an index.
DELETE /api/v1/indexes/{id} Delete an index.
GET /api/v1/indexes/{id} Fetch one index configuration.
PUT /api/v1/indexes/{id} Replace an index configuration.
GET /api/v1/jaeger/{index}/api/services Jaeger-compatible service-name list.
GET /api/v1/jaeger/{index}/api/services/{service}/operations Jaeger-compatible operation-name list for one service.
GET /api/v1/jaeger/{index}/api/traces Jaeger-compatible trace search.
GET /api/v1/jaeger/{index}/api/traces/{trace_id} Jaeger-compatible single-trace lookup.
DELETE /api/v1/jobs/{id} Cancel a batch job.
GET /api/v1/jobs/{id} Poll a batch job's status.
GET /api/v1/jobs/{id}/result Fetch a finished batch job's rows.
POST /api/v1/sql POST /api/v1/sql entry point.
POST /api/v1/sql/distributed Coordinator endpoint: fan req.query across the configured worker peers and merge into the single-pod answer.
POST /api/v1/sql/explain Estimate a query's cost without running it.
POST /api/v1/sql/local Run a SQL query on this pod only.
POST /api/v1/sql/shard Worker endpoint: run the (sharded) query locally and return the result as an Arrow IPC stream (lossless transport).
GET /debug/memory-pool Snapshot the process-wide DataFusion memory pool.
GET /healthz Liveness probe.
GET /readyz Readiness probe: verifies the Iceberg catalog is reachable.

Every /api/v1/* response carries the x-siglake-server-micros header with total server-side latency.

POST /api/v1/sql

Request body

{
  "query": "SELECT host, count(*) FROM events GROUP BY host",
  "format": "records",
  "priority": "interactive",
  "default_order": true,
  "dry_run": false,
  "limits": { "max_rows_returned": 1000 },
  "shard": { "index": 0, "count": 4 }
}
Field Type Default Meaning
query string required The SQL to execute.
format records|ndjson records Response encoding. ndjson streams and bypasses the coordinator's metadata fast paths.
priority interactive|batch interactive batch returns a job id.
default_order bool true Order a bare interactive SELECT newest-first. Set it to false to run the query exactly as written. See Implicit newest-first ordering.
dry_run bool false Return the cost estimate; execute nothing.
limits.max_rows_returned int server default Per-request row cap, clamped by the server ceiling.
shard object None Worker-only. Scan shard index of count.

Records response

{
  "columns": ["host", "count(*)"],
  "row_count": 3,
  "rows": [ ["web-01", 4211], ["web-02", 3980], ["db-01", 122] ],
  "truncated": false,
  "max_rows": null,
  "cost": { },
  "stats": { }
}

truncated and max_rows appear only when the row cap was hit. A truncated records response returns HTTP 413; an ndjson stream stops at the cap and emits a truncation marker. See NDJSON trailers.

Request and response name the same number differently. You request a row cap with limits.max_rows_returned. The envelope reports the applied cap as max_rows. max_rows_returned is the sole request spelling; there is no alias. On /api/v1/sql, /api/v1/sql/local, /api/v1/sql/explain, and /api/v1/sql/distributed, an unknown field inside limits is refused with HTTP 422. The most common mistake is max_rows. The JSON extractor refuses it before planning, execution or batch enqueue, so no work is done and a batch submission creates no job. The response is the extractor's plain-text message naming the unknown field, not an ApiErrorBody JSON envelope.

Only the limits object is strict; unknown fields at the top level of these request bodies are still ignored. /api/v1/sql/shard is not affected because it takes a shard request and resolves the tier defaults. A known limit above its tier ceiling is still clamped rather than refused, while an omitted or empty limits still selects the tier defaults (see Row caps).

cost contains the preflight estimate:

Field Meaning
files_to_scan Files assigned after time-bounds pruning.
files_considered Live files before pruning. The difference is planning's win.
estimated_bytes_scanned Estimated bytes.
estimated_rows_processed Estimated rows.
estimated_runtime_seconds Estimated wall time.
complexity_class small | medium | large | huge.
exact true when the underlying stats were exact, false when heuristic.
warnings Advisory strings.

stats contains execution measurements:

Field Meaning
rows_scanned Leaf rows emitted by a DataFusion data-file scan. A footer-sum or raw-page path bypasses this accounting and can also report 0; inspect served_by.
bytes_scanned Leaf bytes read.
spill_bytes Aggregation/sort spill.
phases Wall-clock breakdown, microseconds.
scan Per-request pruning attribution. Absent when no data-file scan ran.
served_by Execution path that produced the answer. Use this with rows_scanned: metadata, footer-sum, and raw-page paths can all report zero rows scanned.

stats.served_by values

Value Meaning
scan The full DataFusion plan ran and scanned data files.
tier1_inline Exact group counts came from the warm inline whole-table aggregate.
tier1_wide Exact group counts came from the warm wide aggregate, including columns repaired by rebuild-group-counts.
tier1_windowed_agg Exact windowed group counts came from the per-snapshot time-by-group aggregate, with at most the boundary ranges materialized.
materialized Exact group counts were assembled per live file from footers, with raw-page decoding where a usable footer was absent. This can be slow even when rows_scanned is 0.
sketch A bounded approximate heavy-hitter summary served the query; inspect the response's approximation object for its bounds.

stats.phases

Field Meaning
plan_micros Physical planning (single-pod) or logical planning + classification (coordinator).
buffer_delta_micros WAL-buffer delta load. 0 if no buffer or drained.
collect_micros Plan execution. Single-pod only.
render_micros Arrow → JSON rendering.
distributed.mode scan | ordered_scan | aggregate | ordered_aggregate | local fallback.
distributed.shard_wall_micros Per-shard wall time, in shard order. Fan-out is concurrent, so the cost is the maximum rather than the sum.
distributed.merge_micros Coordinator-side merge.

stats.scan

Field Meaning
files_planned Assigned file set after planning.
files_read Files actually opened. An early-stopped browse opens fewer.
files_pruned_bloom Files skipped whole by the trigram/token bloom after the footer read.
planned_bytes / planned_rows Assigned totals.
row_groups_considered Row groups in read files.
row_groups_pruned_bloom Skipped by bloom.
row_groups_pruned_stats Skipped by Parquet statistics.
row_groups_read Actually decoded.
rows_pruned_selection Rows skipped before decode (page index, deletes, inverted index).
object_store_reads Object-store GET count.
fetched_bytes / decoded_bytes Bytes over the wire vs. decoded.
bytes_footer / bytes_index / bytes_data / bytes_other Reader-side fetched bytes for Parquet footers; page, offset, and column indexes plus Siglake index blobs; projected column-chunk data pages; and reader-side manifest or otherwise unclassified reads, respectively. Together they sum to fetched_bytes; planning-time manifest reads use a separate path and appear in neither. Each class is omitted when zero.
file_cache_hits File tasks served from the decoded file-batch cache. A hit builds no reader, so it counts in none of files_read, row_groups_read, object_store_reads, or the byte fields. Omitted when zero.
file_cache_misses File tasks that consulted the decoded file-batch cache, missed, and read the file while populating the cache. Tasks with a raw or promoted prune spec bypass the cache and count in neither cache field. Omitted when zero.
unsettled_partitions Scan partitions still unwinding when the block was built. When non-zero, every count above is missing those partitions' reads. Normally omitted after the server waits for early-stopped partitions; it appears only if that wait reaches its deadline.
ordering advertised when the scan streamed sorted (ordered LIMITs early-stop), otherwise the refusal reason (filtered, fan_in, no_bounds, …).

Implicit newest-first ordering

An interactive SELECT that names one timestamp-bearing table and asks for no ordering of its own is given ORDER BY timestamp DESC. That is what a log reader means by SELECT timestamp, raw FROM events LIMIT 100, and it is what puts the query on the ordered early-stop path instead of a file-order scan. default_order: false on the request runs the query exactly as written.

Eligible tables are events, query_audit, and any managed index whose doc mapping declares timestamp as its timestamp_field.

Input What the server does
Bare SELECT over one eligible table Appends ORDER BY timestamp DESC.
Explicit ORDER BY, GROUP BY, an aggregate, a CTE, a join, DISTINCT Runs it as written.
EXPLAIN, or priority: "batch" Runs it as written.
Rewritten query with an explicit LIMIT Keeps that LIMIT.
Rewritten query with no LIMIT Adds max_rows_returned + 1, so the truncation signal still fires.

Rewrites are counted by siglake_query_default_order_applied_total (see metrics). The scan reports whether it streamed sorted in stats.scan.ordering.

An index whose mapping names some other field as its timestamp_field is left unordered, because a column called timestamp there need not be a time order. Those queries browse in file order, and do not reach the ordered early-stop path, unless they write their own ORDER BY.

NDJSON trailers

POST /api/v1/sql emits NDJSON result rows followed by at most one _meta trailer. The NDJSON trailer fields identify truncation, a row-scan limit or a query that failed mid-stream. Detect a failed query by reading the final NDJSON line and checking its _meta value before accepting the streamed results.

NDJSON trailer fields

An ndjson stream commits its 200 status with the first bytes. Later errors therefore appear in the body. The stream ends with at most one trailer: a JSON object containing a _meta field. Result rows never contain _meta. Inspect the final line before treating a short stream as complete.

_meta Extra fields Meaning
truncated row_count, max_rows The per-request row cap was reached. The rows already emitted are complete; the result is a prefix.
midflight_rows_scanned_exceeded rows_scanned, max_rows_scanned The mid-flight rows-scanned breaker tripped and the stream was cut. Equivalent to the 413 a records request would have received.
error code, error, and retry_after_secs on a pool refusal Execution failed after the response started.

NDJSON trailer examples

One of these lines, never more than one, follows the last row:

{"_meta": "truncated", "row_count": 1000, "max_rows": 1000}
{"_meta": "midflight_rows_scanned_exceeded", "rows_scanned": 51200000, "max_rows_scanned": 50000000}
{"_meta": "error", "code": 503, "error": "query memory pool exhausted; retry after 5s: …", "retry_after_secs": 5}

The error trailer's code is the status the request would have carried had it failed before streaming. A query memory-pool or spill-cap refusal uses 503, adds retry_after_secs (5, the same delay the Retry-After header carries), and increments siglake_query_breaker_trips_total{breaker="pool_exhausted"}. Handle it like a 503 status. See Queries return 503. Every other execution failure uses 500 and omits retry_after_secs.

POST /api/v1/sql/explain

Use POST /api/v1/sql/explain to inspect a DataFusion physical plan without executing the SQL query. Send the same request body as POST /api/v1/sql. The explain response includes the DataFusion plan and preflight cost report, but no query rows. This endpoint is equivalent to dry_run: true on the main route.

Batch jobs

Create a batch job with POST /api/v1/sql and "priority": "batch". Poll the batch job at GET /api/v1/jobs/{id} while its state is pending or running. Before results are available, the job must reach succeeded; fetch them from GET /api/v1/jobs/{id}/result. Other terminal states are failed, cancelled and timeout.

Batch job admission

Submit with "priority": "batch". An admitted submission returns 202 with the job id. If the query admission budget remains full through SIGLAKE_QUERY_ADMISSION_WAIT_MS, the server creates no job and returns 429 with Retry-After.

Status Meaning
pending Accepted but not yet executing.
running An executor has published ownership.
succeeded The result route returns the stored rows.
failed Execution or result storage failed.
cancelled A cancellation reached the job row before another terminal state.
timeout The query exceeded its wall-clock deadline.

Each admitted batch job reserves one share of the pod's admission budget: the budget divided by SIGLAKE_QUERY_ADMISSION_MAX_SHARE_DIVISOR (default 4). It holds the share until the job finishes or a cancellation reaches the replica running it.

After admission:

curl -s http://localhost:8089/api/v1/jobs/$ID          # status
curl -s http://localhost:8089/api/v1/jobs/$ID/result   # rows
curl -sX DELETE http://localhost:8089/api/v1/jobs/$ID  # cancel

Batch job cancellation

DELETE cancels a pending or running job and answers 202 with the cancelled job id. That 202 means the cancellation is persisted and terminal: the row reads cancelled from every query replica, and a completion the executor writes afterwards is refused. It does not mean the work has already stopped.

The abort handle that releases the query-admission reservation and cancels the storage scans is process-local. When the replica serving the DELETE is the one executing the job, the abort is immediate. Otherwise, the executor polls for cancellations. With query.jobs.persistent and more than one query pod, it re-reads the status of its own in-flight jobs every --jobs-cancel-poll-secs / SIGLAKE_JOBS_CANCEL_POLL_SECS (default 2 seconds) and aborts the ones that read cancelled. There is no LISTEN/NOTIFY fast path, so the 202 bounds when the work stops rather than reporting that it already has.

Batch job terminal conflicts

Whichever terminal state lands first wins. A job that already reached one (succeeded, failed, cancelled, or timeout) answers 409; a run that finishes just after a cancellation has its result discarded and reports no succeeded outcome. Recovery rechecks the owner's lease in the conditional write that changes a job to failed, so a heartbeat renewed after candidate selection prevents condemnation. If recovery wins first, that failed state remains terminal, and what the executor loses depends on where it was when the store told it so.

The executor publishes running before it does any work, and that write is conditional on the row still reading pending or running. When the publication is refused because recovery, cancellation or TTL deletion won, the run stops before execution. The query is dropped unpolled, so nothing is planned, no scan starts, no terminal write is attempted, and the admission reservation is released. There is no computed output to discard and no message is added to the job, so its client-visible error stays the one the winner wrote. A job-store error on that publication is not evidence of a terminal verdict, so the run executes as before and lets its terminal write discover the authoritative state.

When the refusal instead meets a run that already finished, the completion is the thing that is refused: a late executor completion cannot reinstate the job, and its computed output is discarded. The job's client-visible error then includes a result for this job was computed after recovery declared its executor gone; it was discarded, so resubmit the query.

Aborts driven by a peer's cancellation count on siglake_query_jobs_cancel_propagated_total, and both refusal shapes count on siglake_query_job_terminal_conflict_total{attempted,actual,cause}: attempted="running" is the refused start that executed nothing, and attempted="succeeded", "failed" or "timeout" is a refused completion whose output was discarded. See Query metrics for what each cause value means and why neither counter takes a plain rate threshold.

An unknown id, or a job owned by another tenant, answers 404 on all three job routes.

GET /api/v1/jobs/{id}/result answers 404 while the job is still pending or running and 409 when it ended in a non-success state; poll GET /api/v1/jobs/{id} to tell those apart.

Batch job store failures

A transient failure of the persistent job store, such as a lost connection, a pool timeout, a serialization failure or deadlock, a Postgres that is restarting, answers 503 with Retry-After: 5 on all three job routes instead. The store never answered, so the job's existence and state are unknown: a 503 is not a 404 and not a 409, and the correct response is to retry the same request unchanged. Any other store error is a 500.

Batch job terminal-state reconciliation

An outage that catches a finished run rather than your request shows up differently: as a running status that outlives the query. The terminal write of a finished run makes three attempts within 30 seconds. If none lands, the run ends anyway, releases its admission share, and hands the job id to its own executor. That executor retries the same conditional write every five seconds until it lands. The row reads running the whole time, so a poll loop can see running long after execution stopped. The five seconds is the delay between passes, not a bound on how long this lasts: every pass needs a job store that answers, so the row stays running for as long as the outage does. Poll to a deadline rather than until a terminal state, and read running as "no verdict yet", not as "still executing".

Nothing is re-executed and no result body is republished when a pass finally lands, so the state it installs comes from a fixed set:

Verdict the run reached Installed status error
succeeded failed this job finished but its result could not be stored (the job store was unavailable); the result was discarded, so resubmit the query
failed failed this job failed and its error could not be stored (the job store was unavailable); resubmit the query to see why
timeout timeout query exceeded the timeout (recorded by its executor after the job store became reachable again)

The first row is the one to handle: the query did run and did produce rows, but they were discarded rather than held through the outage, the result route answers 409 like any other non-success, and the only way to get that answer is to submit the query again. The second keeps the status and loses only the run's own error text. Neither is a new execution, so neither costs anything a resubmit would not.

That reconciling write is conditional on the row still reading pending or running, exactly like the run's own terminal write. A terminal state that won in the meantime therefore stands, and its error is what you read. A cancellation still reads cancelled, and a recovery-condemned row still reads failed with batch executor gone: …. A row the TTL sweep already removed answers 404.

One case never reconciles. An executor already tracking 1,024 finished-but-unpersisted ids refuses the next one, and that row reads running until the pod exits and a peer condemns it on lease expiry as failed with batch executor gone: the replica running this job stopped heartbeating. Operators page on that (SiglakeBatchRowStrandedNonTerminal, see Suggested alerts); a client sees a longer wait and then that recovery text.

Batch job row caps

The per-request row cap (limits.max_rows_returned, clamped by the server ceiling) also applies to batch execution. A finished batch job still returns 200 from the result route when it hits that cap, but its RecordsResponse sets truncated to true and reports the cap in max_rows. Check those fields before treating rows as the whole answer; unlike a synchronous truncated records response, the batch result does not return 413. A job that sets no cap of its own gets the batch tier's default, which is far below the interactive one and is not raised by SIGLAKE_QUERY_MAX_ROWS. See Row caps.

Batch queries run on a dedicated runtime, separate from interactive execution, but both tiers share the query admission budget.

Index management

# create
curl -sX POST http://localhost:8089/api/v1/indexes \
  -H 'Content-Type: application/json' \
  -d '{
    "index_id": "app-logs",
    "doc_mapping": {
      "mode": "dynamic",
      "timestamp_field": "timestamp",
      "field_mappings": [
        {"name": "timestamp", "type": "datetime", "required": true},
        {"name": "message",   "type": "text", "tokenizer": "default"},
        {"name": "level",     "type": "text", "tokenizer": "raw"},
        {"name": "status",    "type": "long"}
      ],
      "tag_fields": ["level"],
      "default_search_fields": ["message"]
    },
    "retention": {"period_secs": 2592000},
    "index_at_flush": false
  }'

See the user indexes guide for field types, tokenizers, and mapping modes.

DELETE /api/v1/indexes/{id} removes only the catalog entry. The index disappears from GET /api/v1/indexes and stops answering queries, but its committed data files remain at the table location, and the current retention and orphan-GC sweeps cannot reclaim them afterward: both resolve a table through the catalog entry that the delete just removed. Dropping an index does not free storage. Account for those bytes until a cleanup path exists.

Creating an index with the same id again reuses that table location but writes a fresh table UUID. The replacement exposes only its own rows, while the dropped incarnation's files sit alongside them under the same prefix. Any out-of-band cleanup must therefore be pinned to the dropped table's UUID and file inventory, never to the location or the index id: a prefix-wide delete against a recreated id destroys the replacement's live data.

The write-ahead log (WAL) also remains under the tenant and index name. Each current segment identifies its Iceberg table by UUID. If an ingester keeps a writer open across the delete and recreation, all WAL-buffer query paths skip the old table's segments. This includes local reads, count fast paths, and distributed fan-out. The filesystem drain moves those segments to stale/<dropped-uuid>/ instead of committing them to the replacement.

With a catalog-claim drain, the mirror prefix is re-stamped for the replacement and each object is checked against its own frame UUID. Objects from the dropped table stay in the mirror and their claims enter quarantine. After a prefix has named a dropped table, an object without its own UUID is also quarantined.

For compatibility, an unmarked local directory or a segment without a UUID has no opinion and serves as before. The mirror exception above applies after an index recreation because its re-stamped prefix can no longer identify an older, UUID-less object.

An ingest lane checks the table UUID when it opens, then refreshes it every 30 seconds by default. A write accepted after a recreation but before that refresh returns 200, then goes to quarantine instead of Iceberg. The row does not become queryable. Set SIGLAKE_WAL_IDENTITY_REFRESH_SECS to shorten this window or disable refreshes.

Index templates auto-create indexes whose IDs match a glob pattern, so a shipper can write to a new index without an explicit create call. PUT /api/v1/index-templates/{id} takes template_id (which must equal the path id), a non-empty index_id_patterns, a priority, a doc_mapping, and an optional retention; the highest-priority match wins, ties breaking on the lexicographically smallest template_id. See the worked example for the full request, where the auto-created index becomes visible, and what happens to writes that match no template.

Delete tasks

curl -sX POST http://localhost:8089/api/v1/delete-tasks \
  -H 'Content-Type: application/json' \
  -d '{
    "index_id": "app-logs",
    "predicate_sql": "user_id = 12345",
    "start_ts": "2026-01-01T00:00:00Z",
    "end_ts": "2026-07-01T00:00:00Z"
  }'
Field Required Meaning
index_id yes Target index.
predicate_sql yes SQL predicate identifying rows to delete.
start_ts / end_ts no Bound the time range, so fewer files are rewritten.

POST, the collection list and GET …/delete-tasks/{id} all return the recorded task. The server sets every field below; none of them is a request field.

Field Type Always present
created_at string (date-time) yes
end_ts string (date-time) or null no
error string or null no
files_rewritten integer yes
index_id string yes
predicate_sql string yes
rows_deleted integer (int64) yes
start_ts string (date-time) or null no
state pending, running, done, failed yes
table_uuid string or null no
task_id string (uuid) yes

"Always present" is the spec's required set. The rest are null unless something set them: error until the task fails, start_ts and end_ts unless the request bounded the range, table_uuid on a task recorded before the binding described below.

table_uuid is the Iceberg table the submission was accepted against: the index incarnation the task is authorized to delete from. An index id is reusable, so DELETE /api/v1/indexes/app-logs followed by a POST of the same id gives a different table. The executor refuses a task whose index_id now resolves to a different table: state becomes failed and error names both tables. A task recorded before this binding existed carries a null table_uuid and is refused the same way; nothing infers an incarnation from a reused name. A refusal commits no snapshot, and error says so when the rewrite had already run and left files no snapshot references. Recovery is the same resubmission as for any failed task. See Delete tasks are bound to one index incarnation.

Tasks are recorded in a ledger and executed by compactor sweeps. They are not immediate. Execution requires compactor.deleteTasks: true, which defaults to off, or a manual siglake delete-sweep. See Retention and deletes.

GET …/delete-tasks/{id} reports claim data only for a pending task. An executor create-only-writes a claim object beside a task before running it. The single-task GET returns that object in a claim field beside the task fields while state is pending. A done, failed or running task has no claim field. Neither POST nor the collection route returns claim data.

Field Meaning
present Whether the claim object was observed by this read.
observed_at When this read observed it. Always set.
claimant UUID of the process that wrote the claim, or null.
claimed_at When the claim was created, or null.
age_seconds Whole seconds from claimed_at to observed_at, floored at zero for clock skew, or null. Not execution duration.

The claim data has these limits:

  • A claim whose body does not decode is still present. An empty, truncated or unparseable body, or one naming a different task_id, answers present: true with claimant, claimed_at and age_seconds all null. Presence is what excludes the next executor, so the object counts even when its diagnostics do not. A store failure other than "not found" is a 500, never a present: false.
  • present: false describes only that read. An executor may claim the task immediately afterwards.
  • Presence, the UUID and age do not prove liveness or abandonment. A claim hours old may belong to a rewrite still running or to a process long dead; the response cannot tell them apart, and no value of age_seconds makes taking the task over safe.

Siglake does not run a terminal or stranded task again. This includes tasks in failed, tasks stranded in running and pending tasks with an unreleased claim. The API has no retry endpoint. To recover, post the same request fields again under the same tenant identity. The response carries a new task_id in pending, and the failed record and its error are left unchanged. The endpoint is not idempotent (every POST creates another task), it promises nothing about exactly-once execution, and each task reports only what its own run rewrote, never the request's cumulative effect. See Recovering a failed task.

Jaeger shim

The HTTP subset Grafana's Jaeger data source renders. {index} is a traces-shaped index.

The two name-list routes, GET …/api/services and GET …/api/services/{service}/operations, share /api/v1/sql's result cache. The first poll for an Iceberg snapshot executes the complete distinct-name query. Repeats for that snapshot are cache hits. A new snapshot gets a new cache entry and executes the query again. The cache stores a complete name list only when its retained arena and its lookup string together fit the 512 KiB per-entry cap. The route's derived name ceiling supplies the row bound; the SQL cache's 128-row entry cap does not apply to these lists. An overlength entry is returned whole but not cached, so every poll runs the query. SIGLAKE_QUERY_RESULT_CACHE controls this cache; there is no Jaeger-specific cache setting. Cache hits preserve the 413, 429, 503 and 504 behavior of a miss.

The shared store still holds at most 256 entries and 4 MiB. That 4 MiB, and the siglake_query_sql_result_cache_bytes gauge that reports it, count each entry's body, the lookup string each entry is stored under, and one pointer per recency marker. For /api/v1/sql that lookup string is the normalized query text, so a long query is charged against the same budget as the rows it returns. An entry whose body and lookup string together exceed its caller's per-entry allowance (4 MiB for /api/v1/sql, 512 KiB for a name list) is refused rather than stored and then evicted.

GET …/api/traces accepts service, operation, start, end, limit, minDuration, maxDuration, and tags (a JSON object string).

All four routes share /api/v1/sql's per-pod interactive admission budget (SIGLAKE_QUERY_ADMISSION_BUDGET_BYTES) and query wall-clock timeout. Each request makes one admission reservation. If the budget remains full through SIGLAKE_QUERY_ADMISSION_WAIT_MS, the route returns 429 with Retry-After; wait for that delay before retrying.

The wall clock starts before tenant resolution and includes index lookup, table registration, and query execution. A trace search gets one deadline across trace-id selection and span retrieval; span retrieval does not get a fresh budget. Expiry during preparation or either query phase returns 504 and cancels the in-flight scan.

The Jaeger ceilings refuse the whole response instead of truncating it. The four routes have trace, span-row, accumulated Arrow-byte, and distinct-name ceilings derived from the admission reservation each request already takes. After a 2 MiB per-request floor, the server conservatively prices each trace at 12 KiB, each span row at 8.5 KiB, each Arrow byte at 4 bytes of peak render memory, and each name at 1.4 KiB. The reservation is min(16 MiB, admission budget / share), so the ceilings follow SIGLAKE_QUERY_ADMISSION_BUDGET_BYTES, the share divisor, and pod memory; there is no Jaeger-specific environment knob. The trace ceiling is also capped at half the span-row ceiling because even a one-span trace consumes a selection row and a span row. If admission is disabled, the calculation uses the packaged minimum reservation rather than becoming unbounded. The server's resolved interactive max_rows_returned, clamped by SIGLAKE_QUERY_MAX_ROWS, still wins when tighter.

A trace search's limit cannot exceed the derived trace ceiling. The server checks it before taking the admission reservation, looking up the index, or planning, and returns 400 with the requested and allowed values. Span rows and Arrow bytes are bounded mid-flight at a batch boundary across both phases of a search; the name-list routes similarly bound names and Arrow bytes. One request spends one budget: for a search, the spans of every selected trace together rather than each trace separately. Narrow service, operation, start/end, tags, or limit when a result is refused.

Every result-ceiling refusal is a whole-response 413: no data, no partial trace, and no Retry-After. /api/v1/sql can return a truncated body that declares truncated and max_rows in its records envelope, but Jaeger's data and total response cannot say that a trace is missing spans. This surface also has no per-request limits object.

Status On these routes
400 A trace search's limit exceeded the derived trace ceiling (or another request parameter was invalid). Checked before admission, index lookup, or planning.
413 A span-row, Arrow-byte, distinct-name, or tighter interactive row ceiling was exceeded. Refused whole: no data and no Retry-After; narrow the request.
429 The shared interactive admission budget stayed full through the admission wait. Honor Retry-After.
503 The query memory pool or the spill byte cap refused an allocation, exactly as on the SQL routes. Honor Retry-After.
504 The one interactive wall-clock budget expired, in preparation or in either query phase.

There is no gRPC SpanReader. See Grafana and Jaeger.

GET /debug/memory-pool

A snapshot of the process-wide DataFusion memory pool on the pod that answers. Use it to find what holds memory when queries return 503 or when SiglakeQueryPoolReservedWhileIdle fires; see Queries return 503 and the saturation alerts. The route is authenticated like /api/v1/* and answers 401 without valid credentials.

{
  "limit_bytes": 4294967296,
  "reserved_bytes": 1073741824,
  "top_consumers": "..."
}
Field Type Meaning
limit_bytes integer or null Process-wide DataFusion pool limit in bytes. null when the pool is unbounded.
reserved_bytes integer or null Bytes currently reserved by DataFusion operators. null when the pool is unbounded.
top_consumers string or null DataFusion's own report text for the ten consumers with the largest current reservations, ordered by reservation size. null when consumer tracking is disabled (SIGLAKE_QUERY_MEMORY_POOL_TRACK_CONSUMERS set to 0 or off) or the pool is unbounded.

Authentication

Mode Ingester flag Query flag
Bearer tokens --auth-tokens / SIGLAKE_AUTH_TOKENS --tokens / SIGLAKE_QUERY_TOKENS
OIDC --oidc-issuer + --oidc-audience same

OIDC takes precedence when both authentication modes are set. On the query server, --oidc-tenant-claim routes each query by the named verified JWT claim. Without that option, query callers use the default namespace; X-Scope-OrgID does not enable query-side header routing.

The ingester uses single-tenant routing by default. An absent X-Scope-OrgID header or an explicit default value routes to the default tenant. Any other value gets a 403. Set --trust-scope-header only behind a gateway that sets the header and strips the client's value. In that mode, the header selects the tenant and an absent header means default.

With --oidc-tenant-claim, the ingester uses the verified claim for every request, including when --trust-scope-header is also set. A present X-Scope-OrgID header must agree with the claim.

The ingester and query server read separate bearer-token settings. A token works on both only when you configure it in both settings. An OIDC deployment can use the same issuer and audience on both processes.

Configuring the claim makes it mandatory on both boundaries, and the value is validated rather than repaired. A verified token whose claim is missing, blank, not a JSON string, longer than 128 characters, or outside [A-Za-z0-9_-] gets a 403 before routing. The query server refuses it in auth middleware, so the refusal reaches every authenticated operation, including operations that read no tenant data. The ingester refuses it before the batch creates a write-ahead log (WAL) lane, metric label or namespace. Siglake does not rewrite the claim or fall back to the default namespace.

siglake_query_tenant_denied_total uses reason="claim_missing" when a verified token lacks the configured tenant claim, and reason="claim_invalid" when the tenant claim is unusable. siglake_ingest_tenant_denied_total uses reason="header_not_trusted" when a header selects another tenant in single-tenant mode, reason="not_allowed" when the resolved tenant is outside --allowed-tenants, reason="at_capacity" when a novel tenant arrives after --max-tenants is full, reason="claim_missing" when a verified token lacks the configured tenant claim, reason="claim_invalid" when the tenant claim is unusable, and reason="header_mismatch" when the header contradicts the tenant claim.

Health and metrics endpoints are unauthenticated.

Error responses

Siglake HTTP APIs use two main error response body shapes. Query errors contain error and code fields; ingester errors contain text and code. A Retry-After header distinguishes retryable 429 responses and capacity, shard-pin or job-store 503 responses. Terminal HTTP failures omit that header. The exceptions appear below.

HTTP status codes

Query-server errors return JSON with an error string and numeric code. Some errors add fields such as reason, pin or cost. Ingester errors use a text string and numeric code; Elasticsearch-compatible routes retain the Elasticsearch error shape. The 422 response below is the JSON extractor's plain text instead. Common statuses:

Status Cause
400 Malformed request or invalid SQL.
401 Missing or invalid credentials.
403 The credential verified but does not authorize the request. On either server, once --oidc-tenant-claim is configured: the tenant claim it requires is missing or is not a usable tenant id ([A-Za-z0-9_-], 1 to 128 characters). Retrying or refreshing the token cannot help; the identity provider has to mint the claim. On the ingester also: an X-Scope-OrgID that names a different tenant than the claim, or a tenant outside --allowed-tenants.
413 Body too large; a records response that hit the row cap, with the truncated prefix and its truncated and max_rows fields; or a Jaeger read that exceeded a result ceiling and returned no data.
422 On the four SqlRequest routes, a field inside limits is not a known per-request limit. The JSON extractor refuses the request before planning, execution or batch enqueue. Its plain-text body names the unknown field. Unknown top-level fields are ignored.
429 On the ingester, the rate budget is exhausted; on the query server, the admission budget stayed full through the configured wait. Honor Retry-After.
500 Internal error.
503 Backpressure, an ingest memory breaker, query memory-pool or spill-cap exhaustion, an unresolvable shard pin (reason: "shard_pin_unresolved"), a transient job-store failure on the batch job routes (Retry-After: 5), or a not-ready dependency.
504 Query wall-clock timeout.

The 429 responses include Retry-After. Capacity, shard-pin and job-store 503 responses include it too. The status code and this header distinguish them from terminal request errors such as 400, 403 and 413.

In a distributed query, a worker's refusal status is forwarded rather than collapsed: a mid-flight row-ceiling trip returns 413, a rejected query 400, a budget refusal 429, a memory-pool or spill-cap refusal 503, an unresolvable shard pin 503, and a timeout 504. For 503, the coordinator re-attaches Retry-After: 5 and does not fail over or re-run the query locally. Any other worker 5xx is reported as 500. A worker fault is a cluster fault. A client acts on them differently: 413 means "narrow the query", 503 means "retry after the requested delay", and 500 means "retry, or page someone".

A refusal that happens after an ndjson response has started cannot change the status: it arrives as an _meta trailer on an otherwise successful 200.

reason: "shard_pin_unresolved"

The error reason shard_pin_unresolved means a worker could not resolve the table generation pinned by the coordinator. A client should respond to this retryable 503 error by waiting for Retry-After: 5 and retrying. The response echoes the pin in pin, as table, snapshot_id and schema_id; either id is null when the pin did not carry it. If the error persists, adjust snapshot retention or the metadata-cache lifetime as described below.

Shard pin failure details

Not every 503 is capacity. The coordinator pins each shard request to the generation it planned against: the snapshot id, and the Iceberg schema id in the optional pin.schema_id field of the /api/v1/sql/shard body. A worker refuses its fragment if it cannot resolve either half: the snapshot has expired from the catalog, or the schema id is in no generation the worker's metadata retains, and neither becomes visible after a metadata refresh. A worker whose metadata merely runs ahead of the coordinator serves the pinned historical schema instead of refusing. The worker does not answer from its own current generation. The body carries reason: "shard_pin_unresolved" and a pin object naming the table, the snapshot id and the schema id, alongside Retry-After: 5, so it is distinguishable from a memory-pool refusal without reading the message. The message names the half that failed, because the operator response differs: a snapshot refusal points at snapshot retention, and a schema refusal at a migrate-schema that has not reached every replica's metadata cache. The coordinator forwards the refusal rather than re-running the fragment locally, so the whole distributed query is refused; /api/v1/sql/local still answers. A distributed answer is from one generation or absent. Siglake does not merge results across generations. See One generation per fan-out.

The schema id is optional so that a rolling upgrade works in both directions. An older coordinator omits it and gets snapshot-only pinning. An older worker ignores it, so a mixed-version fleet keeps snapshot-only pinning until every replica runs the new image.

Misses are counted by siglake_query_shard_pin_total{outcome="miss"} and alert as SiglakeQueryShardPinUnresolved. If they persist rather than clearing as caches converge, raise compactor.snapshotExpire.retainLast or lower SIGLAKE_ICEBERG_METADATA_CACHE_TTL_SECS; see Queries return 503.