Monitoring¶
Configure Prometheus to scrape every role, install the chart's alert rules, then work from the question that matches the incident. Most investigations reduce to four: is ingest healthy, is the drain keeping up, is the layout converging, and why is a query slow. The metrics reference lists every metric and the smaller set you need for routine operations.
Wiring it up¶
serviceMonitor:
enabled: true # Prometheus Operator
prometheusRule:
enabled: true # 39 alert rules; off by default
labels:
release: kube-prometheus-stack
Both objects require the Prometheus Operator. Set prometheusRule.labels to
the labels your Prometheus installation uses for rule discovery. The chart
only installs alert definitions; Prometheus and Alertmanager remain responsible
for evaluation, routing, delivery, acknowledgement, and the on-call workflow.
Or scrape directly:
| Role | Port |
|---|---|
| Ingester | 9100 |
| Compactor | 9101 |
| Query server | 9105 |
The metrics port checks no token of its own. Bearer tokens and OIDC guard the
query API on 8089, not this listener, and --metrics-bind defaults to
0.0.0.0, so each role answers /metrics on every interface its pod has. The
chart's networkPolicy.enabled renders egress rules only, so nothing shipped
restricts who may reach 9100, 9101 or 9105: reachability is whatever your
cluster allows between pods. Keep the port off any Ingress or LoadBalancer, and
if your cluster's default is open, write the ingress restriction yourself. See
Network policies.
Metrics are Prometheus only: no Siglake process exports them as OTLP, whatever its OTel configuration says. An OTLP-only backend gets them from a collector with a Prometheus receiver scraping these ports, which Siglake's own telemetry configures. That page also covers the logs and traces Siglake exports about itself, which are not on this listener.
Starter Grafana dashboard¶
The
deploy/grafana/siglake-overview.json
file is a starter dashboard that you import by hand; the chart does not render
it.
No panel alerts or pages anyone. Only the chart's
PrometheusRule
does that, so use the alert table below to configure what
pages.
The rows and panels worth knowing before you need them:
| Row or panel | What it graphs |
|---|---|
| Durability row | Mirror register and upload abandons, schema drift, and WAL recovery / integrity: CRC mismatches, adopted partials, dropped partial tails and IPC framing refused on one graph. The last series is the five-minute increase in siglake_wal_ipc_framing_refused_total, which counts refused read attempts rather than damaged segments; anything above zero means a WAL segment Siglake cannot read. |
| WAL mirror pin duration, in the Durability row | Average and p99 time that sealing waits for the synchronous mirror pin. Read this beside mirror failures to separate slow pins from failed pins. |
| Group-count delta writes lost / retried | Lost delta writes and retried ones, both by iceberg_namespace and table, plus a third series, side publication lost <namespace>.<table>, for inline aggregate objects that lost a publication outright. |
| Group-count automatic rebuilds | Rebuild attempts, by iceberg_namespace, table and outcome. |
| Drain and compaction row | Reclaims without proof and the mirror-sync panels, including Mirror repairs / rotation completion. Its completion age series is how you spot a backlog or a stalled rotation when every page stays full. |
| Mirror objects left unreclaimed (ledger mark never landed) | Marking failures and mirror objects left behind after the local evidence reaches its retention ceiling. Both series should stay at zero when compactor.mirrorLedgerReclaim is enabled. |
| Segments quarantined, in the Drain and compaction row | Two backlog levels on one graph, each summed over the selected namespaces: siglake_catalog_claim_quarantined (claim quarantine) and siglake_compactor_segments_poisoned (poison / set-aside). Both are gauges, so they fall as you requeue. The event arm, siglake_compactor_segments_poisoned_total, is not plotted: a segment requeued and set aside again counts twice there and once here. The two drains recover differently; see the SiglakeSegmentsQuarantined row. |
| WAL owner mismatches (dropped incarnation held back), in the Drain and compaction row | Hourly refusals of WAL segments written for a deleted and recreated index, by tenant and index, plus both cumulative totals. See WAL segments from a deleted index. |
| Durable reclaim proof row | Consumed-proof size (entries and bytes by table) and Consumed-proof watermark lag / cap refusals. |
| Breaker trips by kind, in the Query row | The five-minute increase in query-breaker trips, by breaker and priority. The pool_exhausted series covers both memory-pool and spill-cap refusals. |
| Incomplete scan attribution rate and Scan settle p99 latency | Requests whose scan counters did not finish settling, and how long requests waited for them. |
| Batch job cancellation across replicas | siglake_query_jobs_cancel_propagated_total, which counts cross-replica aborts and should track client cancellations in a multi-replica tier. Also siglake_query_job_terminal_conflict_total by attempted, actual and cause, and batch outcomes by outcome. |
| Batch jobs finished but not persisted | siglake_query_jobs_unreconciled (finished runs parked for retry) and the siglake_query_jobs_reconciled_total rate by computed and outcome. SiglakeBatchReconciliationBacklogStalled warns per pod when the parked count stays above zero for 15 seconds. The critical signals are siglake_query_jobs_unreconciled_dropped_total and the write_abandoned conflict: those rows have no retry before the pod exits. |
| Table-metadata cache fences by action | Healthy publication contention, separated from reloads that leave a cache entry unpublished. |
| Text-index startup by stage (p50 / p99), in the Fast paths row | Per-file text-index startup grouped by stage: permit_wait, blob_fetch, decode and selection. This is the panel that says whether a slow text query is queueing, fetching, decoding or selecting. See Why is a query slow?. |
| Parsed text-index cache outcomes / evictions, in the Fast paths row | Lookup hit and miss rates beside the byte_bound, entry_bound and oversized drop rates. This is the panel that says whether the parsed cache is being evicted rather than read cold for the first time. |
| Text-index cache resident / budget, in the Fast paths row | The parsed-cache resident and budget gauges beside siglake_iceberg_puffin_blob_cache_bytes and siglake_iceberg_puffin_blob_cache_max_bytes. The Puffin pair is published on every admission attempt, including refusals, so a disabled cache reports a budget of 0. |
| Puffin blob cache fetches / hits / evictions, in the Fast paths row | The full-width panel under the parsed-cache panels: the siglake_iceberg_puffin_blob_fetches_total rate, the lookup hit and miss rates, and the stale, redundant, fifo and oversized drop rates on one graph. The last arm is a refused admission rather than an eviction. Read it beside the parsed-cache panel above it, not alone. See Why is a query slow?. |
| Footer text-index checksum refusals, in the Fast paths row | The five-minute increase in siglake_index_footer_checksum_refused_total, split into malformed and mismatch. Each refusal is a footer-KV inverted index the reader would not trust, so the file was answered from a Puffin index or an exact scan instead. Both reasons are pre-registered at zero on query pods. See Footer text-index checksum refusals. |
| Decoded-file cache populations, in the Fast paths row | All nine outcome arms of siglake_query_scan_file_cache_requests_total, including population_refused. The decoded-file cache is experimental and off by default, so a flat zero on every arm is the right reading for most installs. Once it is on, compare refused and abandoned populations with inserts. See Why is a query slow?. |
| Decoded-file cache bytes (accounted vs entries), in the Fast paths row | siglake_query_scan_file_cache_accounted_bytes against completed-entry bytes. Their gap is decoded data held by admitted populations that have not inserted. |
| Index sidecars written / refused by reason, in the Fast paths row | The hourly increase in siglake_iceberg_segmented_index_writes_total by outcome and reason, which is what the seg2 sidecar writer decided at the close of a compaction rewrite. Read the refusal arms first: column, file_rows and row_domain are the three ways a sidecar would have described a layout the output file does not have, and each one leaves that file on the scan path with no index. A refusal is the safe outcome, so this is where a compactor writing files without indexes is visible rather than a correctness signal. All four arms are pre-registered at 0, so flat zeros are the correct reading of a default install: the writer is opt-in. See Segmented text-index writes. |
| Index sidecar bytes / per-group index bytes (p50 / p99 per pod), in the Fast paths row | siglake_iceberg_segmented_index_written_bytes, one finished sidecar's encoded size, beside siglake_iceberg_segmented_index_group_index_bytes, the heap one row group's postings and dictionary occupy while they are built and release when that group closes. The second is per-row-group allocation, not the writer's peak process heap, and a series that grew with file size would mean the build had stopped being per-row-group. Both export as summaries, so the panel charts each pod's own quantiles rather than a sum or an average, and both are absent until that pod closes its first sidecar. |
| v1 index rebuild: files and source bytes, in the Fast paths row | The hourly increase in siglake_index_rebuild_files_total, with siglake_index_rebuild_bytes_total on the right axis. The bytes are the source Parquet the pass re-read and decoded, not the index it produced. Both counters carry tenant and table, so an empty panel means no rebuild has committed on any pod and not a broken scrape. Files rising while the sidecar panel above stays flat is the compactor paying for whole-file decodes that a seg2 sidecar would have avoided. See v1 index rebuild cost. |
| Auto-promotion | Per-table auto-promotion state: pass outcomes, the last sampled candidate count, the fraction of the configured column cap in use, and whether promotion is enabled at all. The candidate series is hidden when siglake_auto_promotion_candidates_available is 0, because a pass that did not sample has no candidate verdict. The panel carries no attribute names; those are in the compactor's INFO line. See What is attribute auto-promotion doing to my schema?. |
| v1 index rebuild duration (p50 / p99), in the Fast paths row | histogram_quantile() over siglake_index_rebuild_seconds_bucket by table: how long one pass took, from the footer probe over every candidate file through the whole-file decodes to the Puffin registration. These are quantiles over the fleet's aggregated buckets, unlike the per-pod byte panel beside it. The pass holds the compaction cycle it runs in, so a p99 approaching that interval is a compactor whose cycle is the rebuild. A p99 that climbs while the file rate stays flat is per-file decode cost rather than more work. |
The cause label on siglake_query_job_terminal_conflict_total is the one that
needs reading rather than watching:
attempted="running"is a start the store refused, so the query never executed.succeeded,failedandtimeoutare refused completions whose output was discarded.- Cancellation and
goneare expected races. recoverymeans recovery refused executor work of either shape.write_deferredandwrite_abandonedare terminal writes that never landed. The first is being retried by its own executor; the second is not.result_too_largeis an appliedfailedstate.
The four cause="recovery" and three cause="write_abandoned" series are
pre-registered, because the chart's alerts read them. Other causes appear on
first occurrence; see Query metrics.
The local Compose stack ships Prometheus at http://localhost:9090 with
deploy/prometheus.yml already scraping every role.
Query server probes¶
On the query server, /healthz is a process-liveness check that returns a
constant 200, and /readyz only round-trips the Iceberg catalog. Neither
probe executes a query or checks the cache-warm loop, so keep
SiglakeQueryWarmCycleStalled enabled: it is the shipped signal for the known
case where browse requests time out while both probes stay green. Do not treat
successful probes as end-to-end query-path health. See Query probes do not
detect query-path
degradation
for why Siglake does not automatically restart or remove the pod from service.
This behavior comes from Siglake's September 3, 2026 review of query
degradation and probe behavior.
Which build is running?¶
Every long-running service publishes siglake_build_info with version and
commit labels. In Kubernetes, list the build scraped from each pod with:
Each series has a value of 1; its labels are the useful part. Different
version or commit values make a mixed-version fleet visible before any
behavior differs between pods.
Is ingest healthy?¶
rate(siglake_events_accepted_total[5m])
rate(siglake_ingest_backpressure_rejected_total[5m])
siglake_ingest_backpressure_queue_events
Rejections mean Siglake is shedding client requests. Compare them with queue depth. If both rise, the writer cannot keep up, so add ingester CPU or replicas. If the queue stays shallow, a burst filled the lane before it could drain.
siglake_wal_crc_mismatch_total and siglake_wal_ipc_framing_refused_total
should always be zero. The first counts segment bodies that failed their CRC
check; the second counts WAL reads refused because the Arrow IPC length
framing does not describe a message the segment's bytes can back.
Is the drain keeping up, and is the layout converging?¶
siglake_compactor_sealed_pending_oldest_age_seconds and
SiglakeDrainBacklogGrowing tell you whether the WAL drain is keeping up.
siglake_table_overlap_depth and SiglakeTableNotConverging tell you whether
the layout is converging. The supporting metrics identify why either stage is
behind, and the chart ships every alert named here.
Is the WAL drain keeping up?¶
| Metric | One-line reading | Shipped alert | Threshold and direction |
|---|---|---|---|
siglake_compactor_sealed_pending_oldest_age_seconds |
Rising means the drain is falling behind ingest. | SiglakeDrainBacklogGrowing (warning) |
siglake_compactor_sealed_pending_oldest_age_seconds > 900 for 15 minutes. |
siglake_compactor_sealed_pending |
Rising with oldest age means the shared queue backlog is growing. | None | No chart threshold. |
siglake_compactor_drain_inflight |
At SIGLAKE_DRAIN_CONCURRENCY with a growing backlog means the drain is capacity-bound. |
None | No chart threshold. |
siglake_compactor_commit_duration_seconds |
Rising means catalog commits are slowing. | None | No chart threshold. |
siglake_compactor_watchdog_trips_total |
Any increase means a drain cycle or recluster pass missed its deadline. | SiglakeCompactorWatchdogTripping (warning) |
increase(siglake_compactor_watchdog_trips_total[30m]) > 0, with no hold. |
siglake_catalog_claim_quarantined |
Above zero means claimed segments cannot drain without intervention. | SiglakeSegmentsQuarantined (warning) |
siglake_catalog_claim_quarantined > 0 or siglake_compactor_segments_poisoned > 0 for 10 minutes. |
siglake_compactor_segments_poisoned{tenant} |
Above zero means the local filesystem drain is holding unreadable segments under <wal>/poison/ for that tenant. |
SiglakeSegmentsQuarantined (warning) |
siglake_catalog_claim_quarantined > 0 or siglake_compactor_segments_poisoned > 0 for 10 minutes. |
siglake_compactor_segments_poisoned_total{tenant} |
Any increase means another segment was set aside; a requeued segment that fails again counts twice. | None | No chart threshold. |
Is the layout converging?¶
| Metric | One-line reading | Shipped alert | Threshold and direction |
|---|---|---|---|
siglake_table_overlap_depth |
Rising or stuck high means the layout is not converging. | SiglakeTableNotConverging (warning) |
siglake_table_overlap_depth > 60 for 60 minutes. |
siglake_table_gauges_sampled_at_seconds |
Not advancing means the overlap-depth reading is stale. | None | No chart threshold. |
siglake_table_live_data_files |
Rising without bound means compaction is losing to ingest. | None | No chart threshold. |
siglake_compactor_maintenance_throttled_backpressure_total |
Rising means compaction is yielding to the drain. | None | No chart threshold. |
siglake_compactor_watchdog_trips_total |
Any increase can mean a recluster pass missed its deadline. | SiglakeCompactorWatchdogTripping (warning) |
increase(siglake_compactor_watchdog_trips_total[30m]) > 0, with no hold. |
The drain commits sealed WAL segments to Iceberg. Compaction then merges what the drain wrote into a layout queries can scan. If the drain falls behind, rows are durable but not yet queryable. If the layout stops converging, the rows are queryable and scans get slower as depth grows. More compactor replicas clear a drain backlog. They do not help a table whose merge passes keep aborting at their deadline.
Read the drain signals¶
siglake_compactor_sealed_pending_oldest_age_seconds
sum by (namespace) (
max by (namespace, tenant) (
siglake_compactor_sealed_pending
)
)
siglake_compactor_drain_inflight
histogram_quantile(0.95, rate(siglake_compactor_commit_duration_seconds_bucket[5m]))
increase(siglake_compactor_watchdog_trips_total[30m])
siglake_catalog_claim_quarantined
siglake_compactor_segments_poisoned
With two or more catalog-claim workers, every worker publishes the same queue
total.
Deduplicate the copies with max by (namespace, tenant) before you sum tenants.
Dashboard panel 111 does this for both siglake_compactor_sealed_pending and
siglake_compactor_sealed_pending_bytes, then sums each result by namespace.
If the oldest age climbs and sealed_pending grows with it, ingest is running
faster than the drain: rows are durable but not yet queryable, and the WAL
volume is filling. Investigate commit duration or add compactor replicas with
catalogClaim.enabled.
Rising commit duration with flat throughput usually means catalog pressure: check RDS.
A segment the drain cannot commit at all is set aside rather than retried
forever, and SiglakeSegmentsQuarantined reports both drains doing it: the
claim table quarantines the segment, and the local filesystem drain moves the
file under <wal>/poison/. Its row under Stalled carries the
triage for each.
WAL segments from a deleted index¶
Delete an index and recreate it under the same name, and the drain can still find WAL segments written for the old table. It refuses to commit those rows as the replacement's, and counts each refusal:
| Metric | What one increment is |
|---|---|
siglake_compactor_wal_owner_mismatch_total{tenant,index} |
A WAL directory or mirror object prefix stamped for the dropped incarnation. On the filesystem drain the counter moves by the number of segments quarantined under stale/<dropped-uuid>/; on the mirror there is nothing to move, so it moves by one per prefix re-stamp. |
siglake_compactor_wal_stale_segments_total{tenant,index} |
One object refused on its own frame identity: a header naming the dropped table, or, under a prefix that has itself named a dropped table, an object carrying no identity at all. |
Both are refusal events rather than a count of distinct objects held back. A
filesystem segment whose move under stale/ fails stays in sealed/ and is
counted again on the next cycle that examines it, and requeueing a quarantined
mirror claim runs the identity check again, which produces another refusal. So
read the hourly increase as an event rate, and do not read the cumulative total
as a backlog depth. The metrics
reference carries both descriptions.
The WAL owner mismatches (dropped incarnation held back) panel graphs both
series in the starter dashboard's Drain and compaction row. It plots the
hourly increase by tenant and index alongside both cumulative totals,
because a series that appears with its first increment has no rate to read yet.
Neither counter has an alert, and that is deliberate. A refused mirrored object
ends up as a quarantined claim, which SiglakeSegmentsQuarantined and the
Segments quarantined panel already report.
Resolve the identity mismatch before you try to recover the rows. On the mirror
side, requeueing alone does not repair it: follow the
SiglakeSegmentsQuarantined row, which covers what the requeue does
and does not adopt. On the filesystem side the refused segments stay visible
under stale/<dropped-uuid>/ in the index's WAL directory. Nothing on either
path is deleted, so recovering or discarding those rows stays your decision.
Delete tasks that stay non-terminal¶
Delete-task observation runs as part of the compactor's delete sweep:
| Metric | Use |
|---|---|
siglake_compactor_delete_tasks_stalled_total{state} |
Alerting counter. Each increment is one observation of a task whose claim is older than twice the sweep's watchdog ceiling, not one distinct task. |
siglake_compactor_delete_tasks_nonterminal{state} |
Dashboard only. The last complete observation, aggregated across namespaces. A live pod retains it when delete-task execution is disabled, the lease is elsewhere, or the stage watchdog trips, so it is not a liveness or coverage signal. |
The claim object's age is an upper bound on execution duration, not a running
duration. Neither metric says that an executor died or whether its rewrite
committed. The state set is exactly running and pending_claimed.
Read the layout signals¶
siglake_table_overlap_depth
siglake_table_live_data_files
siglake_table_gauges_sampled_at_seconds
rate(siglake_compactor_maintenance_throttled_backpressure_total[5m])
Use overlap_depth as the layout-health metric. A converged 2 B-row table
settles in the low twenties. A flat plateau over hours means compaction is not
converging, and query performance is silently degraded: an unconverged 1 TB
layout collapsed 32-way query throughput from 115 QPS to 28. Because this
gauge comes from the budgeted manifest walk, it can lag; first confirm that
siglake_table_gauges_sampled_at_seconds is advancing. The same freshness
signal covers the level and leading-edge gauges and changes only when a table's
walk completes.
live_data_files growing without bound means compaction is losing to ingest.
Unlike the walked gauges, it is read exactly from the snapshot summary every
sampler cycle, regardless of table size. The throttled-backpressure counter
tells you compaction is deliberately yielding. That is expected under load,
but it is a problem if it never stops.
What is attribute auto-promotion doing to my schema?¶
Auto-promotion is off by default. Once SIGLAKE_AUTO_PROMOTE_MIN_PCT is above
zero, the compactor widens table schemas from a sample with no operator in the
loop, and that widening cannot be undone, so watch the column count and what
each pass decided.
siglake_auto_promotion_enabled
sum by (iceberg_namespace, table, outcome) (increase(siglake_auto_promotion_passes_total[1h]))
siglake_auto_promotion_columns{kind="used"} / ignoring (kind) siglake_auto_promotion_columns{kind="limit"}
siglake_auto_promotion_candidates and on (namespace, pod, iceberg_namespace, table) siglake_auto_promotion_candidates_available == 1
| Reading | What it means |
|---|---|
enabled is 0 |
Promotion is switched off on that compactor, by SIGLAKE_AUTO_PROMOTE_MIN_PCT=0 or SIGLAKE_AUTO_PROMOTE_MAX_COLUMNS=0. Nothing samples and nothing widens. |
enabled is 1 and all four passes_total outcomes are 0 |
Promotion is on, and no pass has completed on that table yet. All four outcomes are pre-registered, so this is a measured "never ran", not a gap in the scrape. |
nothing_cleared rising |
Passes are sampling and adding nothing: no unpromoted attribute reaches the threshold. This is the steady state on a settled table. |
at_ceiling rising |
The table is at SIGLAKE_AUTO_PROMOTE_MAX_COLUMNS, so the pass skips sampling entirely. Hot attributes stay in attributes until you raise the cap. |
failed rising |
The pass could not read the table metadata or could not finish sampling. Read the compactor WARN lines for the error. |
candidates_available is 0 |
The last pass reached no verdict about candidates, so siglake_auto_promotion_candidates is either absent or holding an earlier pass's reading. Gate every candidate query on this gauge. |
Names are in the log: auto-promotion pass finished at INFO carries up to 32
attribute names per disposition, promoted and declined, with a _truncated
count for the rest. The dashboard's Auto-promotion panel graphs these four
families, and SiglakeAutoPromotionNearCeiling warns at 80% of the cap. See
Attribute auto-promotion for
each metric and Declared vs automatic
promotion for what a
promotion costs.
Why is a query slow?¶
Start with the response. Every response carries cost, stats.scan and
stats.phases. See Query
engine.
Then, cluster-wide:
histogram_quantile(0.95, rate(siglake_query_request_duration_seconds_bucket[5m]))
siglake_query_exec_pool_queue_seconds
siglake_query_in_flight
rate(siglake_query_scan_output_ordering_total[5m])
| Symptom | Likely cause |
|---|---|
High exec_pool_queue_seconds |
Pool saturation, not slow scans. Add query replicas. |
scan_output_ordering{no_bounds} rising |
Files lack manifest bounds: ordered early-stop is refusing. |
scan_output_ordering{fan_in} or {global_fan_in} rising |
Layout not disjoint: overlap exceeds the merge budget. Compaction problem. |
scan_output_ordering{filtered} dominating |
Expected: filtered browses take a bounded TopK instead. Only chase it if those browses are slow. |
query_breaker_trips_total rising |
Check the breaker label: SQL row-ceiling trips indicate unbounded queries or a low ceiling; jaeger_trace_limit, jaeger_span_rows, jaeger_render_bytes, and jaeger_name_rows identify the derived Jaeger ceiling that refused a read; pool_exhausted means a memory-pool reservation or spill-cap refusal. |
Aggregates with rows_scanned > 0 |
A fast path didn't engage. |
| A text query is slow and the response does not say where | Read Text-index startup by stage (p50 / p99) in the Fast paths row. permit_wait is files queueing behind SIGLAKE_INDEX_LOAD_CONCURRENCY, blob_fetch is the Puffin read from object storage, decode is a parsed-cache miss, and selection is the postings and row-selection work above the index. |
decode samples climbing toward selection samples |
The plan's indexes are not staying warm: decode counts cold loads, selection counts index-pruned files. Read Parsed text-index cache outcomes / evictions next. |
parsed_index_cache_evictions_total rising with the misses |
The cache is decoding indexes and throwing them away. byte_bound and entry_bound name the bound that bit; oversized means one index exceeds the whole budget and that file decodes on every query. Raise the matching knob in Tune the Puffin and parsed text-index caches. |
| Misses with no evictions beside them | First reads, or a pod that derived no text-index cache at all. See Features off by default in Helm. |
puffin_blob_cache_lookups_total{outcome="hit"} climbing with the parsed misses |
The working set is past the parsed budget but inside the blob budget, so each re-decode is served from bytes already held. Read Puffin blob cache fetches / hits / evictions beside the parsed panel. Raising SIGLAKE_PARSED_INDEX_CACHE_MAX_BYTES is what removes the decode. |
puffin_blob_fetches_total tracking the parsed miss rate |
Every re-decode is also re-reading object storage. Check the eviction arms: a steady fifo rate means neither the stale nor the redundant arm found a victim, and the oldest entry went instead, which can be the blob the next query wants. Raise SIGLAKE_PUFFIN_BLOB_CACHE_MAX_BYTES or SIGLAKE_PUFFIN_BLOB_CACHE_MAX_ENTRIES in Tune the Puffin and parsed text-index caches. |
puffin_blob_cache_evictions_total{reason="oversized"} rising beside fetches |
One blob exceeds the byte budget, so the cache refuses it and reads it from object storage on every decode. This can resemble a cold or disabled cache. Raise SIGLAKE_PUFFIN_BLOB_CACHE_MAX_BYTES in Tune the Puffin and parsed text-index caches. |
Fetches above the lookup miss rate |
Index reading that consulted no cache: a rebuild that bypasses it, or a pod with either blob-cache bound at 0. Confirm the bounds before you read the hit ratio, because a disabled cache leaves both lookup arms flat at zero. |
scan_file_cache_requests_total{outcome="abandoned"} tracking the miss rate with insert at zero |
You turned the decoded-file cache on and it is adding nothing during that window: every population is dropped before its insert, typically a LIMIT served from the first batches or cancelled queries. The budget is still subtracted from the query memory pool. Read hit, siglake_query_scan_file_cache_entries and siglake_query_scan_file_cache_bytes before you conclude the cache is empty; earlier inserts can still be serving hits. If the shape holds for the workload you run, clear both knobs in Cache tuning. |
index_footer_checksum_refused_total rising |
A footer text index failed its checksum and was refused, so those files are read by Puffin index or exact scan. Answers stay correct and the scan cost stays until the file is rewritten. Read Footer text-index checksum refusals. |
Footer text-index checksum refusals¶
siglake_index_footer_checksum_refused_total{reason} above zero means a query
found a footer-KV inverted index that did not match the sibling CRC-32 stored
with it, and refused to prune with it. Answers stay exact: the reader tries the
column's Puffin index and otherwise scans the file. The cost is pruning on that
file, and it stays until something rewrites the file.
Read the two reasons differently. mismatch is the corruption case: the blob
in the Parquet footer no longer hashes to the value written beside it. Any
rewrite of that file, a compactor re-cluster or a delete-task rewrite, replaces
the index it carries, and that is what clears the refusal. malformed is a
value the reader cannot parse at all: a checksum that is not eight hexadecimal
characters, or a blob that is not even-length hexadecimal. The Siglake writer
emits neither shape, so it is damage that landed on the characters themselves,
or a footer some other tool rewrote.
Neither series carries a file label and no log line names the refused path, so the metric says that this is happening and not where. The rate follows query traffic rather than the number of damaged files: one file read by ten queries counts ten times. The Footer text-index checksum refusals panel in the Fast paths row graphs it by reason, and ships no alert.
Query peer discovery and table-cache publication¶
When SRV peer discovery is enabled, each refresh increments
siglake_query_peer_discovery_refresh_total. The changed and unchanged
outcomes mean the pod resolved a usable membership; empty, unmatched, and
error are failures. A failed refresh keeps the last good membership, which
protects in-flight queries from a DNS blip but means a pod broken since startup
can keep answering every query locally without an obvious request failure.
siglake_query_peer_discovery_members shows the published member count, and
siglake_query_peer_discovery_last_success_seconds records the last usable
answer.
SiglakeQueryPeerDiscoveryStalled is the warning for a pod cut off from its
membership. One failed refresh is normal during a rollout. The trigger instead
means that refreshes are being attempted and none are succeeding. The failing
outcomes increase over a ten-minute window while neither changed nor
unchanged increases, then the rule holds for ten minutes. Failures alongside
successes are a cluster in churn or a partially answering resolver; that pod
still has a membership, so the rule deliberately stays quiet. With the default
--query-peer-discovery-interval-secs of 5, ten minutes without one usable
answer is roughly 120 consecutive attempts.
Split the refresh counter by outcome to act on it: for error, check CoreDNS,
NetworkPolicy, and that --query-peer-discovery-srv names the headless
Service's http port; for empty, check that the Service has Ready endpoints;
for unmatched, compare
SIGLAKE_QUERY_PEER_SELF_NAME with the SRV
targets. A pod that cannot find its own name in the answer is also left
without a failover target.
What the stall costs fan-out depends on whether the pod ever published a membership. One that has not answers every query itself, single-pod, using none of the query tier's other replicas, and no query-side counter marks it as degraded. One that published earlier keeps serving its last good membership, so its fan-out is stale. It will not see replicas added since, and can briefly address a member that has gone away. Same-shard failover turns that into a latency event instead of a wrong answer.
The query server and compactor also count table-metadata cache publication
races in siglake_iceberg_table_cache_fenced_total. Its pre-registered
action="superseded" and action="reload" series are healthy contention under
concurrent commits. action="unpublished" means the reload exhausted its
attempts, left that table's cache entry empty, and made later reads pay for a
full load_table.
SiglakeTableCacheUnpublished pages only when unpublished continues, rather
than on the healthy arms. If superseded or reload is moving too, reduce the
commit rate by raising compactor.commitBatch.targetMb or maxAgeSecs. If
metadata loads themselves are slow, shrink metadata.json by lowering
compactor.snapshotExpire.retainLast, or run
snapshot expiry more frequently by lowering its intervalSecs.
External consumer health¶
Processes built on the WAL consumer interface must export their own lag, error, and result telemetry. Their alert names depend on that separate deployment and are not part of Siglake's metric inventory.
Disaster-recovery readiness¶
rate(siglake_wal_mirror_failures_total[5m])
rate(siglake_wal_mirror_segments_total[5m])
siglake_wal_mirror_queue_depth
rate(siglake_wal_mirror_queue_wait_seconds_sum[5m])
/ rate(siglake_wal_mirror_queue_wait_seconds_count[5m])
sum(rate(siglake_wal_mirror_pin_duration_seconds_sum[5m]))
/ sum(rate(siglake_wal_mirror_pin_duration_seconds_count[5m]))
histogram_quantile(0.99,
sum by (le) (rate(siglake_wal_mirror_pin_duration_seconds_bucket[5m])))
Mirror failures move your recovery point further back without affecting ingest. If mirroring is enabled, alert on any failure and restore uploads before the WAL volume becomes the only copy.
siglake_wal_mirror_queue_depth counts in-memory segments waiting for the
uploader. It excludes the active upload and does not inventory durable
mirror-pending/ pins. A stuck active upload can therefore coincide with a
queue depth of zero.
siglake_wal_mirror_queue_wait_seconds records each segment's time from
enqueue to dequeue. It does not measure upload duration or the current oldest
pending segment's age. Keep the failure alert and WAL volume free-space
monitoring alongside both queue signals. A prolonged remote failure retains
successfully pinned bytes across local compaction and retention, which can
exhaust the volume.
siglake_wal_mirror_pin_duration_seconds records each seal's whole synchronous
pin_segment call, whether the call succeeds or fails. Its first four buckets
are below 1 ms, where healthy pins commonly finish. The average and p99 show
how long sealing waits, but they do not attribute time to local lookup retries,
hard-link creation, or the mirror-pending/ directory fsync.
The ingester checks sealed/ and mirror-pending/ at startup and every 300
seconds to resume uploads after an outage or process exit. Set
SIGLAKE_WAL_MIRROR_SWEEP_SECS=0 to disable both checks. A failed pin appears
as siglake_wal_mirror_failures_total{reason="pin"} and does not carry the
retention protection described above.
Two reasons on the same counter report leaked storage rather than a worse
recovery point. reason="active_cleanup" means three deletes of a sealed
segment's _active/ partial failed, and reason="active_cleanup_stat" means
an active upload could not check whether its segment had sealed. The sealed
segment is uploaded and registered in both cases. The next catch-up sweep
retries the delete while the local segment is still on the WAL volume; after
that, only an object-store lifecycle expiry removes the partial. Treat a
rising rate as a storage-growth signal on the mirror prefix, not as a lost
recovery point.
Unreadable objects found by a wal-recover run¶
siglake wal-recover reports mirror damage in its printed output and in log
events, and exports no metric for it. The command is a one-shot subcommand that
installs no metrics recorder and binds no /metrics, so a counter it charged
would be discarded when the process exits, and no row in the metric
inventory covers a recovery run. What you read
instead is the counts on the printed report and one WARN event per refused
object and unrecognised name. Unreadable reporting arrives in Siglake 0.2.0,
with the plan-first wal-recover.
UNREADABLE: `default/siglake-ingester-0-01a0afc1-2901-7cc2-8b41-6f5f0f4fd83e.arrow` is not a WAL segment (WAL frame integrity check: <object-store segment> body length 972 != header 625); it will be left in the mirror
totals: 467 segments, 3.5 GiB (1 unreadable: body does not decode, 17 keys skipped: unrecognised layout)
One UNREADABLE: line prints per refused object, capped at ten, and the count
lands on totals:. Any nonzero count is rows you cannot restore from the
mirror. The refused objects stay where they are and the local drain never sees
them, so the count comes back every time you plan. For the disposition, and for
what separates an unreadable candidate from a skipped name, read What an
UNREADABLE: line in the plan
means.
Each refused object and each unrecognised name also emits a WARN event carrying the object name and, for a body that did not decode, the error. These are the three messages a log rule matches on:
| Message | Emitted when |
|---|---|
wal-recover: unrecognised key, skipped |
A mirror object's name does not fit the layout, so the plan never reads its body. |
wal-recover: candidate body is not a readable WAL segment, refused |
The plan read a candidate's body and could not decode it. |
wal-recover: candidate body is not a readable WAL segment, nothing written |
--apply could not decode a body, so it wrote nothing for that object. |
The events go to stderr, and over OTLP when OTEL_EXPORTER_OTLP_ENDPOINT is
set; the command flushes the exporter before it returns. So alerting on a
damaged mirror is a rule in your log pipeline, not a Prometheus rule, and it
depends on you collecting the run's logs: a restore whose stderr goes nowhere
leaves no record once the terminal scrolls. Setting RUST_LOG replaces the
default filter outright, so keep WARN enabled for siglake_wal if you set it.
Suggested alerts¶
The chart's
PrometheusRule
is the authoritative alert set.
Siglake CI checks that every rule uses a metric exported by the build. The
groups describe the operator action, not the component that emits the metric.
Alert counters whose label sets are known at startup are pre-registered at
zero, so Prometheus increase() can observe the first event after a fresh pod
starts. The exceptions have labels that are discovered only when the event
occurs: siglake_storage_schema_drift_total{column=...} and most series of
siglake_group_count_delta_write_failures_total,
siglake_group_count_delta_write_retries_total,
siglake_side_aggregate_publish_failures_total,
siglake_group_count_auto_rebuilds_total and
siglake_group_count_short_aggregates_total. Those five carry
iceberg_namespace alongside table, and only
{iceberg_namespace="siglake",table="events"} is pre-registered, including the
automatic-rebuild counter's success, incomplete and failed outcomes and
all four census outcomes. A tenant_* namespace, an index table and a base
namespace moved off the default by SIGLAKE_TENANT_NAMESPACE are known only at
the first increment. For a dynamically discovered series, the first increment
is not visible to increase(); the schema-drift alert
therefore fires on the second refusal, which the next drain cycle produces.
SiglakeGroupCountDeltaRetrying reads rate() over 15 minutes and holds for
30, so it needs retries in two consecutive windows whether or not its series
was pre-registered.
Data loss or durable inconsistency¶
Page on the critical rules in this group. The warning rules expose precursors or recoverable inconsistencies that still need investigation.
| Alert | Severity | Operator action |
|---|---|---|
SiglakeBatchCompletionRejectedByRecovery |
Warning | Recovery made a batch job terminal before its executor finished, so a lifecycle transition was refused. Check attempted. For running, read the WARN line batch lifecycle transition was refused for the job ID, owner and superseding status; the query never ran. For succeeded, failed or timeout, read batch completion was refused for the job ID, owner, attempted outcome, and dropped rows and bytes; the computed output was discarded. The failed job tells the client to resubmit. The rule has no hold. Client cancellation, TTL expiry, oversized results and abandoned terminal writes do not fire it. |
SiglakeBatchReconciliationBacklogStalled |
Warning | Restore the shared job store and inspect the affected query-server pod's WARN logs while it keeps retrying the terminal writes. The rule warns per pod after siglake_query_jobs_unreconciled stays above zero for 15 seconds. Its timer starts when the gauge becomes nonzero, not when the store returns. Verified kind run #129 peaked at seven parked rows, split 3/4 between two pods, and drained within 13.19887 seconds after restoration. The 15-second hold rounds that upper bound to the next 5-second reconciliation pass. This warning covers rows still being retried; do not restart the pod for this warning. |
SiglakeBatchRowStrandedNonTerminal |
Critical | A batch run finished without persisting a verdict, and nothing is retrying it. Restore the job store, then restart the pod. Restarting before the store recovers only strands the next runs. Read the query-server ERROR lines batch terminal state could not be persisted and could not be tracked and job-store outage outlasted this replica's reconciliation bookkeeping for the job IDs and owner. The rule groups both abandoned-write counters by pod, has no hold and excludes rows that an executor is still reconciling. |
SiglakeWalMirrorRegisterAbandoned |
Critical | Catalog registration of an uploaded mirror object gave up after its retries, so acknowledged rows are durable in object storage but have no catalog row and no query returns them. Repair the object-store or catalog fault the compactor reported, then let mirror reconciliation register the unregistered objects on a later pass. siglake wal-recover is not the remedy here: it restores a local WAL root, which the catalog-claim drain never reads. Follow Recover a catalog-claim drain, and decide from its row counts rather than from siglake_compactor_sealed_pending. |
SiglakeWalMirrorUploadAbandoned |
Critical | Protect the WAL PVC and restore mirror uploads; it is now the only copy of the affected segments. |
SiglakeWalCrcMismatch |
Critical | Investigate the corrupted, quarantined segment immediately and recover its rows from an intact copy if one exists. |
SiglakeWalIpcFramingRefused |
Critical | A WAL read walked the Arrow IPC length prefixes against the file size and found framing the segment's bytes cannot back, so the segment is unreadable and needs investigation. Read the failing role's error: it names the byte offset where the walk stopped and what the message declared there, and the recovered-partial paths name the segment file too. Keep the file and recover its rows from an intact copy, such as the object-storage mirror, if one exists. What happens to the refused segment depends on the caller, so check rather than assume it was quarantined. The counter counts refused read attempts, not damaged segments: a retried read counts again. A body that fails its CRC fires SiglakeWalCrcMismatch instead, a recovered partial whose complete prefix still decoded fires SiglakeWalPartialTailDropped, and a decode that fails after the framing was accepted fires neither. The rule fires on any increase over 15 minutes and has no hold. |
SiglakeSchemaWritesRefused |
Critical | Run siglake migrate-schema --all-tables --all-namespaces; writes are being refused to prevent silent column loss. |
SiglakeReclaimWithoutProof |
Warning | Neither the durable consumed proof nor retained snapshot history covered the claims, so requeueing may duplicate rows. Inspect durable-proof decode or cap errors; during a mixed-version rollout, keep compactor.snapshotExpire.retainLast >= 400 through 1,025 seconds after the final old writer exits. |
SiglakeConsumedProofAtCapacity |
Critical | Siglake fails closed at the consumed-proof metadata cap. Inspect non-terminal catalog claims holding back the proof watermark before retrying writes. |
SiglakeWalPartialAdopted |
Warning | Treat an isolated alert after a pod death as expected; if it repeats on healthy pods, investigate writer starvation and the adoption threshold. |
SiglakeWalPartialTailDropped |
Warning | A recovered partial segment ended inside an Arrow IPC message. Siglake drained the complete fsynced batches ahead of it and left the source bytes on disk untouched, so no acknowledged row was lost: the discarded tail was never acknowledged. Read the warning log line for the segment, the rows recovered and the bytes dropped, then match the alert against the ingester restart or storage fault that caused it. The rule fires on any increase over 30 minutes and has no hold. Repeats on a healthy fleet mean an unstable WAL volume or a process that keeps dying mid-append. |
SiglakeGroupCountDeltaLost |
Warning | Check the outcome label: failed means the automatic rebuild errored and left its marker for retry; incomplete means the next fold ran but could not restore full coverage. Read the compactor log, then run siglake rebuild-group-counts --namespace <ns> --table <table>, which the alert renders with both labels filled in. Answers remain exact on the per-file path meanwhile. |
SiglakeGroupCountAggregateShort |
Warning | The maintenance census found a maintained column short of total-records with every commit's contribution already accounted for. Answers stay exact: GROUP BY uses the per-file path until the aggregate is rebuilt. Check outcome: detected is the default reporting-only mode; incomplete and failed mean a rebuild could not restore coverage. The backed_off_watchdog, backed_off_failed and backed_off_interrupted outcomes identify the last unsuccessful automatic scan. Retries wait 15 minutes, one hour and four hours. suppressed means the table incarnation spent all four attempts, or its marker is malformed. marker_failed means the compactor skipped the scan because it could not persist the attempt first. The WARN line gives the attempt count, last reason, next eligible time and manual command when they apply. Run siglake rebuild-group-counts --namespace <ns> --table <table> for suppression or repeated failure. A successful command publishes first, clears applicable attempt history and reports the number cleared. The selector excludes repaired; the rule fires on any included increase over an hour with no hold. See Rebuild a short group-count aggregate with durable backoff. |
SiglakeGroupCountDeltaRetrying |
Warning | Check warehouse object-store health and credential refresh on the named pod before retries become a lost delta. Its counter, siglake_group_count_delta_write_retries_total, carries iceberg_namespace and table like the four counters beside it, so the alert names <iceberg_namespace>.<table> and the pod that retried. It renders no repair command: a retry that succeeded lost nothing, and the fix is the pod's warehouse access. The compactor INFO line group-count delta write succeeded on retry names the delta object. |
SiglakeSideAggregatePublicationLost |
Warning | A publication of the table's inline aggregate object spent all four attempts, 250, 500 and 750 ms apart, so that commit's group counts and time aggregates are gone. Answers stay exact: the read guard refuses a short aggregate and the query falls back to the per-file path. Where the incremental delta path is active, the publication leaves the same rebuild marker a lost delta does, so the compactor restores the wide group counts; run siglake rebuild-group-counts --namespace <ns> --table <table>, which the alert renders with both labels filled in, if that rebuild fails. Nothing rebuilds the inline time aggregates, so windowed GROUP BY on that table answers from the per-file path until the object is rebuilt. Check warehouse object-store health and credential refresh on the named pod. The rule fires on any increase over an hour and has no hold. |
SiglakeInlineCoverageUnproven |
Critical | The maintenance census found a maintained table whose inline aggregate object cannot prove coverage of the current snapshot, with no publication in flight. Retention, a delete task, a foreign overwrite and the residual windows at snapshot expiry all leave that state, and no commit repairs it: the table answers windowed GROUP BY, date histograms and windowed counts from the per-file tiers for the rest of its life. Answers stay exact, because the read guard refuses the object. The severity is critical because the state does not heal and the repair is manual: run siglake rebuild-time-aggregates --namespace <ns> --table <table>, which the alert renders with both labels filled in and which is safe to re-run. The compactor WARN line names the same table. The rule pairs siglake_inline_coverage_unproven > 0 with increase(siglake_inline_coverage_census_total[1h]) > 0 on the same pod, so a compactor that stopped censusing leaves the alert instead of paging from a reading nobody is refreshing, and holds for 30 minutes: two censuses at the default cadence. Nothing reports the repair's own cost: the command records how long each component took and how much it decoded, but it starts no metrics listener, so time the run from the outside and read what it did from the report it prints. See Inline aggregate repair. |
SiglakeMirrorReconciliationErrors |
Warning | A mirror listing or catalog-registration pass failed in the last 30 minutes; the rule has no additional hold time. Inspect Cycles by outcome and the compactor warning log, then repair the reported object-store or catalog fault so a later pass can register the uploaded segments. Its catalog_sync_error series is pre-registered at startup. |
Stalled¶
| Alert | Severity | Operator action |
|---|---|---|
SiglakeCompactorWatchdogTripping |
Warning | Inspect compactor deadline failures; repeated trips mean drain or reclustering is not making progress. |
SiglakeDrainBacklogGrowing |
Warning | Add drain capacity and investigate commit throughput before the WAL volume fills. |
SiglakeSegmentsQuarantined |
Warning | Both drains fire this rule, and the recovery differs. On the claim path, read the reason field on the segment QUARANTINED log line. Repeated drain failures quarantine a segment after twelve attempts; restore its missing index or fix the reported fault, then run requeue_quarantined. A mirrored segment with a mismatched table identity, or no identity under a displaced-owner prefix, is refused on its first cycle. Requeueing a terminally refused segment runs the identity check again; it does not adopt the old rows into the replacement table. On the local filesystem drain, siglake_compactor_segments_poisoned{tenant} is above zero because a segment failed to read on every attempt it was given and now sits under <wal>/poison/ beside a .poison.json note; the segment SET ASIDE log line names the file and the read error. Fix what made it unreadable, then run siglake wal-requeue --wal <wal-root> to move every held segment back into sealed/, --segment <filename> to move one, or --dry-run first to see what would move. Nothing under poison/ is deleted or requeued for you, so discarding those rows stays your decision. |
SiglakeCompactorOrphansHeld |
Critical | siglake_compactor_orphans_held > 0 for 15 minutes, per tenant and pod. A compactor killed mid-commit left a segment under <wal>/orphans/ whose name is absent from the table's consumed set, and the snapshot that would prove whether it was committed is no longer retained. No later cycle settles it. Preserve the files: their rows may already be in the table or may exist nowhere else, so deleting one can lose rows and requeueing one can duplicate them. Establish which from your own evidence, the compactor's orphan auto-disposition INFO line names the directory and table, and the segment's rows can be read and compared against the table, before you move anything into sealed/ or delete it. Raising compactor.snapshotExpire.retainLast protects the proof for orphans a future crash creates; it does not restore expired history and will not clear this hold. See When the compactor holds an orphan for you. |
SiglakeMirrorReconciliationStalled |
Warning | Bounded pages ran without a full rotation completing for prometheusRule.mirrorRotationStallSecs (21,600 seconds by default), then remained stalled for the rule's 15-minute hold. Inspect Mirror repairs / rotation completion and Mirror sync objects per pass; raise the threshold or reduce retained-prefix size if a rotation is merely slow, otherwise investigate the stuck walk. |
SiglakeDeleteTaskStalled |
Critical | increase(siglake_compactor_delete_tasks_stalled_total[1h]) > 0 reports observations, not distinct tasks, with no hold. The named state has remained non-terminal longer than twice the delete sweep's watchdog ceiling, but claim age only bounds execution duration from above: this proves neither that the executor is dead nor whether its rewrite committed. Find the task id and index in the compactor WARN line, read GET /api/v1/delete-tasks/{id}, decide from the audit trail whether the original rewrite landed, and resubmit the request under a new task id. A watchdog ceiling of 0 disables stalled classification. |
SiglakeTableNotConverging |
Warning | Restore compaction headroom and investigate the sustained overlap depth; scanning latency degrades as depth grows. |
SiglakeQueryPeerDiscoveryStalled |
Warning | The rule fires when failed outcomes increase over ten minutes while changed and unchanged do not, then holds for ten minutes. Split the counter by outcome. For error, check CoreDNS, NetworkPolicy and the headless Service's http port. For empty, check Ready endpoints. For unmatched, compare SIGLAKE_QUERY_PEER_SELF_NAME with the SRV targets. Until a refresh succeeds, a new pod serves queries without fan-out; a pod with an earlier membership keeps using that stale membership. See Query peer discovery. |
SiglakeTableCacheUnpublished |
Warning | Cache reloads keep exhausting their publication attempts, leaving the affected table entry empty and forcing full metadata loads. Reduce commit frequency with compactor.commitBatch.targetMb / maxAgeSecs, or reduce metadata size with snapshot-expiry retention and interval settings. |
Refusing work¶
| Alert | Severity | Operator action |
|---|---|---|
SiglakeIngestLanesRefused |
Warning | Check whether clients vary X-Scope-OrgID or x-siglake-index; otherwise raise ingester.maxLanes deliberately. Only the lane cap increments siglake_ingest_lane_refused_total: a refusal at ingester.maxTenants lands on SiglakeTenantsDenied with reason="at_capacity" instead. |
SiglakeTenantsDenied |
Warning | Split siglake_ingest_tenant_denied_total by reason. For header_not_trusted, the shipped single-tenant default refused an X-Scope-OrgID naming a tenant other than default; a header naming default is accepted. Stop sending the unintended header, or route tenancy by verified identity with ingester.oidc.tenantClaim, or set ingester.trustScopeHeader if a gateway sets the header and strips caller-supplied values. For not_allowed, fix the client's tenant or add the intended tenant to ingester.allowedTenants. For claim_missing, fix the identity provider or the configured claim name. For claim_invalid, send a value of 1 to 128 characters from [A-Za-z0-9_-]. For header_mismatch, remove the contradictory X-Scope-OrgID header. For at_capacity, the pod is at ingester.maxTenants and refused a tenant it had not admitted before; the tenants already writing on that pod are unaffected. Raise the cap or name the tenants you expect in ingester.allowedTenants, remembering that the count is per pod and starts empty on restart. The rule has no reason selector, so every label feeds it. All six series are pre-registered. Each label and the setting that emits it is in Ingest tenant selection. |
SiglakeQueryShardPinUnresolved |
Warning | Workers cannot resolve the generation pinned by the coordinator. Distributed queries return 503 with reason: "shard_pin_unresolved"; see Error responses. Read the worker log for which half of the pin it refused. For a snapshot, raise compactor.snapshotExpire.retainLast or lower SIGLAKE_ICEBERG_METADATA_CACHE_TTL_SECS. For a schema id, check that every worker reads the same warehouse and that its metadata refreshes are succeeding. The rule fires on a non-zero miss rate for five minutes. Both outcomes are pre-registered. |
SiglakeQueryTenantsDenied |
Warning | Query tenant admission returned 403. Split siglake_query_tenant_denied_total by reason. For claim_missing, fix the identity provider or configured query.oidc.tenantClaim. For claim_invalid, send a value of 1 to 128 characters from [A-Za-z0-9_-]. For not_allowed, fix the caller's claim or add the intended tenant to query.allowedTenants. All three series are pre-registered. |
Saturation¶
| Alert | Severity | Operator action |
|---|---|---|
SiglakeAutoPromotionNearCeiling |
Warning | A table with auto-promotion enabled has used at least 80% of its configured column cap for 10 minutes. The rule divides siglake_auto_promotion_columns{kind="used"} by the kind="limit" series and keeps only tables whose siglake_auto_promotion_enabled is 1, so it reads each table against that compactor's own SIGLAKE_AUTO_PROMOTE_MAX_COLUMNS rather than a universal ceiling. Decide which way to settle it: raising the cap permits more schema additions you cannot undo, and SIGLAKE_AUTO_PROMOTE_MIN_PCT=0 stops future automatic ones while keeping the columns already added. Read the compactor's auto-promotion pass finished INFO lines first for the attribute names the cap refused, and remember that the count includes columns you declared by hand. See What is attribute auto-promotion doing to my schema?. |
SiglakeQueryScanAttributionIncomplete |
Warning | A scan partition was still unwinding two seconds after the root stream ended. The response was returned, but its stats.scan omits that partition's counters. Confirm the affected request has stats.scan.unsettled_partitions > 0 and compare it with the warning scan partitions still unwinding at the settle deadline; the alert recovers on its own when subsequent requests settle completely. |
SiglakeQueryPoolNearLimit |
Warning | Reduce query pressure or add query memory/capacity before spilling and reduced file concurrency degrade service. |
SiglakeQueryPoolRefusing |
Warning | Compare SiglakeQueryPoolNearLimit, then make an authenticated GET /debug/memory-pool request to the affected query pod. Review query.resources.limits.memory and the query.spill.maxBytes / query.spill.sizeLimit sizing. The warning recovers when pressure drops, but sustained trips mean queries are receiving retryable 503 responses. |
SiglakeQueryPoolReservedWhileIdle |
Warning | While the residual is present, make an authenticated GET /debug/memory-pool request to the affected query pod. Compare its ten largest live DataFusion consumers with the reservation-owner warning and siglake_query_memory_pool_idle_residual_reports_total. Consumer tracking is on unless SIGLAKE_QUERY_MEMORY_POOL_TRACK_CONSUMERS is 0 or off; if it is disabled, re-enable it before the next occurrence. Restart a pod with a wedged warm probe, or investigate a named operator as a memory leak. These diagnostics were added by Siglake's September 3, 2026 query-memory observability change. |
SiglakeQueriesAbandoned |
Warning | Investigate client timeouts and disconnects, then check the query pod for the degradation that accumulated abandonments can precede. |
SiglakeQueryWarmCycleStalled |
Warning | Restart the pod, then inspect siglake_query_warm_cycle_in_progress to distinguish a wedged cycle from a dead warm loop. This is the shipped signal for the query degradation that /healthz cannot see: browse requests can time out while health stays green. |
SiglakeQueryWarmCyclesAbandoned |
Warning | Investigate slow storage or cancellation; repeated alerts are an early query-pod wedge signal, so restart the affected pod if degradation follows. |
Logs¶
RUST_LOG controls filtering; the chart's logLevel default is
info,siglake=info. For debugging, info,siglake=debug.
Container runtimes capture the server roles' tracing logs from stderr; the server roles write nothing operational to stdout.
SIGLAKE_LOG_EXEC_PLANS logs physical query plans: very verbose, useful when
a fast path isn't engaging and you need to see what DataFusion actually built.
Query audit as an observability surface¶
query_audit is an ordinary Iceberg table, so your query history is
queryable:
-- most expensive query shapes in the last day
SELECT query, count(*) AS runs, avg(duration_ms) AS avg_ms,
max(estimated_bytes_scanned) AS max_bytes
FROM query_audit
WHERE timestamp >= now() - INTERVAL '24 hours'
AND complexity IN ('large', 'huge')
GROUP BY query
ORDER BY avg_ms DESC
LIMIT 20;
-- rejections and errors
SELECT status, count(*) FROM query_audit
WHERE timestamp >= now() - INTERVAL '1 hour'
GROUP BY status;
The table contains literal SQL text. Treat it as sensitive and rotate it
with siglake audit-rotate --max-age-secs N.