Monitoring¶
Configure Prometheus to scrape every role, install the chart's alert rules, then work from the question that matches the incident. Most investigations reduce to four: is ingest healthy, is the drain keeping up, is the layout converging, and why is a query slow. The metrics reference lists every metric and the smaller set you need for routine operations.
Wiring it up¶
serviceMonitor:
enabled: true # Prometheus Operator
prometheusRule:
enabled: true # 33 alert rules; off by default
labels:
release: kube-prometheus-stack
Both objects require the Prometheus Operator. Set prometheusRule.labels to
the labels your Prometheus installation uses for rule discovery. The chart
only installs alert definitions; Prometheus and Alertmanager remain responsible
for evaluation, routing, delivery, acknowledgement, and the on-call workflow.
Or scrape directly:
| Role | Port |
|---|---|
| Ingester | 9100 |
| Compactor | 9101 |
| Query server | 9105 |
Starter Grafana dashboard¶
The
deploy/grafana/siglake-overview.json
file is a starter dashboard that you import by hand; the chart does not render
it.
No panel alerts or pages anyone. Only the chart's
PrometheusRule
does that, so use the alert table below to configure what
pages.
The rows and panels worth knowing before you need them:
| Row or panel | What it graphs |
|---|---|
| Durability row | Mirror register and upload abandons, schema drift, and WAL recovery / integrity: CRC mismatches, adopted partials and dropped partial tails on one graph. |
| Group-count delta writes lost / retried | Lost and retried delta writes, by table, plus a third series, side publication lost <table>, for inline aggregate objects that lost a publication outright. |
| Group-count automatic rebuilds | Rebuild attempts, by table and outcome. |
| Drain and compaction row | Reclaims without proof and the mirror-sync panels, including Mirror repairs / rotation completion. Its completion age series is how you spot a backlog or a stalled rotation when every page stays full. |
| WAL owner mismatches (dropped incarnation held back), in the Drain and compaction row | Hourly refusals of WAL segments written for a deleted and recreated index, by tenant and index, plus both cumulative totals. See WAL segments from a deleted index. |
| Durable reclaim proof row | Consumed-proof size (entries and bytes by table) and Consumed-proof watermark lag / cap refusals. |
| Breaker trips by kind, in the Query row | The five-minute increase in query-breaker trips, by breaker and priority. The pool_exhausted series covers both memory-pool and spill-cap refusals. |
| Incomplete scan attribution rate and Scan settle p99 latency | Requests whose scan counters did not finish settling, and how long requests waited for them. |
| Batch job cancellation across replicas | siglake_query_jobs_cancel_propagated_total, which counts cross-replica aborts and should track client cancellations in a multi-replica tier. Also siglake_query_job_terminal_conflict_total by attempted, actual and cause, and batch outcomes by outcome. |
| Batch jobs finished but not persisted | siglake_query_jobs_unreconciled (finished runs parked for retry) and the siglake_query_jobs_reconciled_total rate by computed and outcome. Also the two series that page: siglake_query_jobs_unreconciled_dropped_total and the write_abandoned conflict. A parked count that falls back to zero is reconciliation working; the two paging series are rows nothing will resolve before the pod exits. |
| Table-metadata cache fences by action | Healthy publication contention, separated from reloads that leave a cache entry unpublished. |
The cause label on siglake_query_job_terminal_conflict_total is the one that
needs reading rather than watching:
attempted="running"is a start the store refused, so the query never executed.succeeded,failedandtimeoutare refused completions whose output was discarded.- Cancellation and
goneare expected races. recoverymeans recovery refused executor work of either shape.write_deferredandwrite_abandonedare terminal writes that never landed. The first is being retried by its own executor; the second is not.result_too_largeis an appliedfailedstate.
The four cause="recovery" and three cause="write_abandoned" series are
pre-registered, because the chart's alerts read them. Other causes appear on
first occurrence; see Query metrics.
The local Compose stack ships Prometheus at http://localhost:9090 with
deploy/prometheus.yml already scraping every role.
Query server probes¶
On the query server, /healthz is a process-liveness check that returns a
constant 200, and /readyz only round-trips the Iceberg catalog. Neither
probe executes a query or checks the cache-warm loop, so keep
SiglakeQueryWarmCycleStalled enabled: it is the shipped signal for the known
case where browse requests time out while both probes stay green. Do not treat
successful probes as end-to-end query-path health. See Query probes do not
detect query-path
degradation
for why Siglake does not automatically restart or remove the pod from service.
This behavior comes from Siglake's September 3, 2026 review of query
degradation and probe behavior.
Which build is running?¶
Every long-running service publishes siglake_build_info with version and
commit labels. In Kubernetes, list the build scraped from each pod with:
Each series has a value of 1; its labels are the useful part. Different
version or commit values make a mixed-version fleet visible before any
behavior differs between pods.
Is ingest healthy?¶
rate(siglake_events_accepted_total[5m])
rate(siglake_ingest_backpressure_rejected_total[5m])
siglake_ingest_backpressure_queue_events
Rejections mean Siglake is shedding client requests. Compare them with queue depth. If both rise, the writer cannot keep up, so add ingester CPU or replicas. If the queue stays shallow, a burst filled the lane before it could drain.
siglake_wal_crc_mismatch_total should always be zero.
Is the drain keeping up, and is the layout converging?¶
siglake_compactor_sealed_pending_oldest_age_seconds and
SiglakeDrainBacklogGrowing tell you whether the WAL drain is keeping up.
siglake_table_overlap_depth and SiglakeTableNotConverging tell you whether
the layout is converging. The supporting metrics identify why either stage is
behind, and the chart ships every alert named here.
Is the WAL drain keeping up?¶
| Metric | One-line reading | Shipped alert | Threshold and direction |
|---|---|---|---|
siglake_compactor_sealed_pending_oldest_age_seconds |
Rising means the drain is falling behind ingest. | SiglakeDrainBacklogGrowing (warning) |
siglake_compactor_sealed_pending_oldest_age_seconds > 900 for 15 minutes. |
siglake_compactor_sealed_pending |
Rising with oldest age means the shared queue backlog is growing. | None | No chart threshold. |
siglake_compactor_drain_inflight |
At SIGLAKE_DRAIN_CONCURRENCY with a growing backlog means the drain is capacity-bound. |
None | No chart threshold. |
siglake_compactor_commit_duration_seconds |
Rising means catalog commits are slowing. | None | No chart threshold. |
siglake_compactor_watchdog_trips_total |
Any increase means a drain cycle or recluster pass missed its deadline. | SiglakeCompactorWatchdogTripping (warning) |
increase(siglake_compactor_watchdog_trips_total[30m]) > 0, with no hold. |
siglake_catalog_claim_quarantined |
Above zero means segments cannot drain without intervention. | SiglakeSegmentsQuarantined (warning) |
siglake_catalog_claim_quarantined > 0 for 10 minutes. |
Is the layout converging?¶
| Metric | One-line reading | Shipped alert | Threshold and direction |
|---|---|---|---|
siglake_table_overlap_depth |
Rising or stuck high means the layout is not converging. | SiglakeTableNotConverging (warning) |
siglake_table_overlap_depth > 60 for 60 minutes. |
siglake_table_gauges_sampled_at_seconds |
Not advancing means the overlap-depth reading is stale. | None | No chart threshold. |
siglake_table_live_data_files |
Rising without bound means compaction is losing to ingest. | None | No chart threshold. |
siglake_compactor_maintenance_throttled_backpressure_total |
Rising means compaction is yielding to the drain. | None | No chart threshold. |
siglake_compactor_watchdog_trips_total |
Any increase can mean a recluster pass missed its deadline. | SiglakeCompactorWatchdogTripping (warning) |
increase(siglake_compactor_watchdog_trips_total[30m]) > 0, with no hold. |
The drain commits sealed WAL segments to Iceberg. Compaction then merges what the drain wrote into a layout queries can scan. If the drain falls behind, rows are durable but not yet queryable. If the layout stops converging, the rows are queryable and scans get slower as depth grows. More compactor replicas clear a drain backlog. They do not help a table whose merge passes keep aborting at their deadline.
Read the drain signals¶
siglake_compactor_sealed_pending_oldest_age_seconds
sum by (namespace) (
max by (namespace, tenant) (
siglake_compactor_sealed_pending
)
)
siglake_compactor_drain_inflight
histogram_quantile(0.95, rate(siglake_compactor_commit_duration_seconds_bucket[5m]))
increase(siglake_compactor_watchdog_trips_total[30m])
siglake_catalog_claim_quarantined
With two or more catalog-claim workers, every worker publishes the same queue
total.
Deduplicate the copies with max by (namespace, tenant) before you sum tenants.
Dashboard panel 111 does this for both siglake_compactor_sealed_pending and
siglake_compactor_sealed_pending_bytes, then sums each result by namespace.
If the oldest age climbs and sealed_pending grows with it, ingest is running
faster than the drain: rows are durable but not yet queryable, and the WAL
volume is filling. Investigate commit duration or add compactor replicas with
catalogClaim.enabled.
Rising commit duration with flat throughput usually means catalog pressure: check RDS.
A segment the drain cannot commit at all is quarantined rather than retried
forever, which SiglakeSegmentsQuarantined reports; its row under
Stalled carries the triage.
WAL segments from a deleted index¶
Delete an index and recreate it under the same name, and the drain can still find WAL segments written for the old table. It refuses to commit those rows as the replacement's, and counts each refusal:
| Metric | What one increment is |
|---|---|
siglake_compactor_wal_owner_mismatch_total{tenant,index} |
A WAL directory or mirror object prefix stamped for the dropped incarnation. On the filesystem drain the counter moves by the number of segments quarantined under stale/<dropped-uuid>/; on the mirror there is nothing to move, so it moves by one per prefix re-stamp. |
siglake_compactor_wal_stale_segments_total{tenant,index} |
One object refused on its own frame identity: a header naming the dropped table, or, under a prefix that has itself named a dropped table, an object carrying no identity at all. |
Both are refusal events rather than a count of distinct objects held back. A
filesystem segment whose move under stale/ fails stays in sealed/ and is
counted again on the next cycle that examines it, and requeueing a quarantined
mirror claim runs the identity check again, which produces another refusal. So
read the hourly increase as an event rate, and do not read the cumulative total
as a backlog depth. The metrics
reference carries both descriptions.
The WAL owner mismatches (dropped incarnation held back) panel graphs both
series in the starter dashboard's Drain and compaction row. It plots the
hourly increase by tenant and index alongside both cumulative totals,
because a series that appears with its first increment has no rate to read yet.
Neither counter has an alert, and that is deliberate. A refused mirrored object
ends up as a quarantined claim, which SiglakeSegmentsQuarantined and the
Segments quarantined panel already report.
Resolve the identity mismatch before you try to recover the rows. On the mirror
side, requeueing alone does not repair it: follow the
SiglakeSegmentsQuarantined row, which covers what the requeue does
and does not adopt. On the filesystem side the refused segments stay visible
under stale/<dropped-uuid>/ in the index's WAL directory. Nothing on either
path is deleted, so recovering or discarding those rows stays your decision.
Delete tasks that stay non-terminal¶
Delete-task observation runs as part of the compactor's delete sweep:
| Metric | Use |
|---|---|
siglake_compactor_delete_tasks_stalled_total{state} |
Alerting counter. Each increment is one observation of a task whose claim is older than twice the sweep's watchdog ceiling, not one distinct task. |
siglake_compactor_delete_tasks_nonterminal{state} |
Dashboard only. The last complete observation, aggregated across namespaces. A live pod retains it when delete-task execution is disabled, the lease is elsewhere, or the stage watchdog trips, so it is not a liveness or coverage signal. |
The claim object's age is an upper bound on execution duration, not a running
duration. Neither metric says that an executor died or whether its rewrite
committed. The state set is exactly running and pending_claimed.
Read the layout signals¶
siglake_table_overlap_depth
siglake_table_live_data_files
siglake_table_gauges_sampled_at_seconds
rate(siglake_compactor_maintenance_throttled_backpressure_total[5m])
Use overlap_depth as the layout-health metric. A converged 2 B-row table
settles in the low twenties. A flat plateau over hours means compaction is not
converging, and query performance is silently degraded: an unconverged 1 TB
layout collapsed 32-way query throughput from 115 QPS to 28. Because this
gauge comes from the budgeted manifest walk, it can lag; first confirm that
siglake_table_gauges_sampled_at_seconds is advancing. The same freshness
signal covers the level and leading-edge gauges and changes only when a table's
walk completes.
live_data_files growing without bound means compaction is losing to ingest.
Unlike the walked gauges, it is read exactly from the snapshot summary every
sampler cycle, regardless of table size. The throttled-backpressure counter
tells you compaction is deliberately yielding. That is expected under load,
but it is a problem if it never stops.
Why is a query slow?¶
Start with the response. Every response carries cost, stats.scan and
stats.phases. See Query
engine.
Then, cluster-wide:
histogram_quantile(0.95, rate(siglake_query_request_duration_seconds_bucket[5m]))
siglake_query_exec_pool_queue_seconds
siglake_query_in_flight
rate(siglake_query_scan_output_ordering_total[5m])
| Symptom | Likely cause |
|---|---|
High exec_pool_queue_seconds |
Pool saturation, not slow scans. Add query replicas. |
scan_output_ordering{no_bounds} rising |
Files lack manifest bounds: ordered early-stop is refusing. |
scan_output_ordering{fan_in} or {global_fan_in} rising |
Layout not disjoint: overlap exceeds the merge budget. Compaction problem. |
scan_output_ordering{filtered} dominating |
Expected: filtered browses take a bounded TopK instead. Only chase it if those browses are slow. |
query_breaker_trips_total rising |
Check the breaker label: SQL row-ceiling trips indicate unbounded queries or a low ceiling; jaeger_trace_limit, jaeger_span_rows, jaeger_render_bytes, and jaeger_name_rows identify the derived Jaeger ceiling that refused a read; pool_exhausted means a memory-pool reservation or spill-cap refusal. |
Aggregates with rows_scanned > 0 |
A fast path didn't engage. |
Query peer discovery and table-cache publication¶
When SRV peer discovery is enabled, each refresh increments
siglake_query_peer_discovery_refresh_total. The changed and unchanged
outcomes mean the pod resolved a usable membership; empty, unmatched, and
error are failures. A failed refresh keeps the last good membership, which
protects in-flight queries from a DNS blip but means a pod broken since startup
can keep answering every query locally without an obvious request failure.
siglake_query_peer_discovery_members shows the published member count, and
siglake_query_peer_discovery_last_success_seconds records the last usable
answer.
SiglakeQueryPeerDiscoveryStalled is the warning for a pod cut off from its
membership. One failed refresh is normal during a rollout. The trigger instead
means that refreshes are being attempted and none are succeeding. The failing
outcomes increase over a ten-minute window while neither changed nor
unchanged increases, then the rule holds for ten minutes. Failures alongside
successes are a cluster in churn or a partially answering resolver; that pod
still has a membership, so the rule deliberately stays quiet. With the default
--query-peer-discovery-interval-secs of 5, ten minutes without one usable
answer is roughly 120 consecutive attempts.
Split the refresh counter by outcome to act on it: for error, check CoreDNS,
NetworkPolicy, and that --query-peer-discovery-srv names the headless
Service's http port; for empty, check that the Service has Ready endpoints;
for unmatched, compare
SIGLAKE_QUERY_PEER_SELF_NAME with the SRV
targets. A pod that cannot find its own name in the answer is also left
without a failover target.
What the stall costs fan-out depends on whether the pod ever published a membership. One that has not answers every query itself, single-pod, using none of the query tier's other replicas, and no query-side counter marks it as degraded. One that published earlier keeps serving its last good membership, so its fan-out is stale. It will not see replicas added since, and can briefly address a member that has gone away. Same-shard failover turns that into a latency event instead of a wrong answer.
The query server and compactor also count table-metadata cache publication
races in siglake_iceberg_table_cache_fenced_total. Its pre-registered
action="superseded" and action="reload" series are healthy contention under
concurrent commits. action="unpublished" means the reload exhausted its
attempts, left that table's cache entry empty, and made later reads pay for a
full load_table.
SiglakeTableCacheUnpublished pages only when unpublished continues, rather
than on the healthy arms. If superseded or reload is moving too, reduce the
commit rate by raising compactor.commitBatch.targetMb or maxAgeSecs. If
metadata loads themselves are slow, shrink metadata.json by lowering
compactor.snapshotExpire.retainLast, or run
snapshot expiry more frequently by lowering its intervalSecs.
External consumer health¶
Processes built on the WAL consumer interface must export their own lag, error, and result telemetry. Their alert names depend on that separate deployment and are not part of Siglake's metric inventory.
Disaster-recovery readiness¶
rate(siglake_wal_mirror_failures_total[5m])
rate(siglake_wal_mirror_segments_total[5m])
siglake_wal_mirror_queue_depth
rate(siglake_wal_mirror_queue_wait_seconds_sum[5m])
/ rate(siglake_wal_mirror_queue_wait_seconds_count[5m])
Mirror failures move your recovery point further back without affecting ingest. If mirroring is enabled, alert on any failure and restore uploads before the WAL volume becomes the only copy.
siglake_wal_mirror_queue_depth counts in-memory segments waiting for the
uploader. It excludes the active upload and does not inventory durable
mirror-pending/ pins. A stuck active upload can therefore coincide with a
queue depth of zero.
siglake_wal_mirror_queue_wait_seconds records each segment's time from
enqueue to dequeue. It does not measure upload duration or the current oldest
pending segment's age. Keep the failure alert and WAL volume free-space
monitoring alongside both queue signals. A prolonged remote failure retains
successfully pinned bytes across local compaction and retention, which can
exhaust the volume.
The ingester checks sealed/ and mirror-pending/ at startup and every 300
seconds to resume uploads after an outage or process exit. Set
SIGLAKE_WAL_MIRROR_SWEEP_SECS=0 to disable both checks. A failed pin appears
as siglake_wal_mirror_failures_total{reason="pin"} and does not carry the
retention protection described above.
Suggested alerts¶
The chart's
PrometheusRule
is the authoritative alert set.
Siglake CI checks that every rule uses a metric exported by the build. The
groups describe the operator action, not the component that emits the metric.
Alert counters whose label sets are known at startup are pre-registered at
zero, so Prometheus increase() can observe the first event after a fresh pod
starts. The exceptions have labels that are discovered only when the event
occurs: siglake_storage_schema_drift_total{column=...} and the index-table
series for siglake_group_count_delta_write_failures_total,
siglake_side_aggregate_publish_failures_total and
siglake_group_count_auto_rebuilds_total. The events-table series are
pre-registered, including the automatic-rebuild counter's success,
incomplete, and failed outcomes. For a dynamically discovered series, the
first increment is not visible to increase(); the schema-drift alert
therefore fires on the second refusal, which the next drain cycle produces.
Data loss or durable inconsistency¶
Page on the critical rules in this group. The warning rules expose precursors or recoverable inconsistencies that still need investigation.
| Alert | Severity | Operator action |
|---|---|---|
SiglakeBatchCompletionRejectedByRecovery |
Warning | Recovery made a batch job terminal before its executor finished, so a lifecycle transition was refused. Check attempted. For running, read the WARN line batch lifecycle transition was refused for the job ID, owner and superseding status; the query never ran. For succeeded, failed or timeout, read batch completion was refused for the job ID, owner, attempted outcome, and dropped rows and bytes; the computed output was discarded. The failed job tells the client to resubmit. The rule has no hold. Client cancellation, TTL expiry, oversized results and abandoned terminal writes do not fire it. |
SiglakeBatchRowStrandedNonTerminal |
Critical | A batch run finished without persisting a verdict, and nothing is retrying it. Restore the job store, then restart the pod. Restarting before the store recovers only strands the next runs. Read the query-server ERROR lines batch terminal state could not be persisted and could not be tracked and job-store outage outlasted this replica's reconciliation bookkeeping for the job IDs and owner. The rule groups both abandoned-write counters by pod, has no hold and excludes rows that an executor is still reconciling. |
SiglakeWalMirrorRegisterAbandoned |
Critical | Run siglake wal-recover or re-register the affected object keys; acknowledged rows are durable but absent from the catalog. |
SiglakeWalMirrorUploadAbandoned |
Critical | Protect the WAL PVC and restore mirror uploads; it is now the only copy of the affected segments. |
SiglakeWalCrcMismatch |
Critical | Investigate the corrupted, quarantined segment immediately and recover its rows from an intact copy if one exists. |
SiglakeSchemaWritesRefused |
Critical | Run siglake migrate-schema --all-tables --all-namespaces; writes are being refused to prevent silent column loss. |
SiglakeReclaimWithoutProof |
Warning | Neither the durable consumed proof nor retained snapshot history covered the claims, so requeueing may duplicate rows. Inspect durable-proof decode or cap errors; during a mixed-version rollout, keep compactor.snapshotExpire.retainLast >= 400 through 1,025 seconds after the final old writer exits. |
SiglakeConsumedProofAtCapacity |
Critical | Siglake fails closed at the consumed-proof metadata cap. Inspect non-terminal catalog claims holding back the proof watermark before retrying writes. |
SiglakeWalPartialAdopted |
Warning | Treat an isolated alert after a pod death as expected; if it repeats on healthy pods, investigate writer starvation and the adoption threshold. |
SiglakeWalPartialTailDropped |
Warning | A recovered partial segment ended inside an Arrow IPC message. Siglake drained the complete fsynced batches ahead of it and left the source bytes on disk untouched, so no acknowledged row was lost: the discarded tail was never acknowledged. Read the warning log line for the segment, the rows recovered and the bytes dropped, then match the alert against the ingester restart or storage fault that caused it. The rule fires on any increase over 30 minutes and has no hold. Repeats on a healthy fleet mean an unstable WAL volume or a process that keeps dying mid-append. |
SiglakeGroupCountDeltaLost |
Warning | Check the outcome label: failed means the automatic rebuild errored and left its marker for retry; incomplete means the next fold ran but could not restore full coverage. Read the compactor log, then run siglake rebuild-group-counts --table <table>. Answers remain exact on the per-file path meanwhile. |
SiglakeGroupCountDeltaRetrying |
Warning | Check warehouse object-store health and credential refresh on the named pod before retries become a lost delta. |
SiglakeSideAggregatePublicationLost |
Warning | A publication of the table's inline aggregate object spent all four attempts, 250, 500 and 750 ms apart, so that commit's group counts and time aggregates are gone. Answers stay exact: the read guard refuses a short aggregate and the query falls back to the per-file path. Where the incremental delta path is active, the publication leaves the same rebuild marker a lost delta does, so the compactor restores the wide group counts; run siglake rebuild-group-counts --table <table> if that rebuild fails. Nothing rebuilds the inline time aggregates, so windowed GROUP BY on that table answers from the per-file path until the object is rebuilt. Check warehouse object-store health and credential refresh on the named pod. The rule fires on any increase over an hour and has no hold. |
SiglakeMirrorReconciliationErrors |
Warning | A mirror listing or catalog-registration pass failed in the last 30 minutes; the rule has no additional hold time. Inspect Cycles by outcome and the compactor warning log, then repair the reported object-store or catalog fault so a later pass can register the uploaded segments. Its catalog_sync_error series is pre-registered at startup. |
Stalled¶
| Alert | Severity | Operator action |
|---|---|---|
SiglakeCompactorWatchdogTripping |
Warning | Inspect compactor deadline failures; repeated trips mean drain or reclustering is not making progress. |
SiglakeDrainBacklogGrowing |
Warning | Add drain capacity and investigate commit throughput before the WAL volume fills. |
SiglakeSegmentsQuarantined |
Warning | Read the reason field on the segment QUARANTINED log line. Repeated drain failures quarantine a segment after twelve attempts; restore its missing index or fix the reported fault, then run requeue_quarantined. A mirrored segment with a mismatched table identity, or no identity under a displaced-owner prefix, is refused on its first cycle. Requeueing a terminally refused segment runs the identity check again; it does not adopt the old rows into the replacement table. |
SiglakeMirrorReconciliationStalled |
Warning | Bounded pages ran without a full rotation completing for prometheusRule.mirrorRotationStallSecs (21,600 seconds by default), then remained stalled for the rule's 15-minute hold. Inspect Mirror repairs / rotation completion and Mirror sync objects per pass; raise the threshold or reduce retained-prefix size if a rotation is merely slow, otherwise investigate the stuck walk. |
SiglakeDeleteTaskStalled |
Critical | increase(siglake_compactor_delete_tasks_stalled_total[1h]) > 0 reports observations, not distinct tasks, with no hold. The named state has remained non-terminal longer than twice the delete sweep's watchdog ceiling, but claim age only bounds execution duration from above: this proves neither that the executor is dead nor whether its rewrite committed. Find the task id and index in the compactor WARN line, read GET /api/v1/delete-tasks/{id}, decide from the audit trail whether the original rewrite landed, and resubmit the request under a new task id. A watchdog ceiling of 0 disables stalled classification. |
SiglakeTableNotConverging |
Warning | Restore compaction headroom and investigate the sustained overlap depth; scanning latency degrades as depth grows. |
SiglakeQueryPeerDiscoveryStalled |
Warning | The rule fires when failed outcomes increase over ten minutes while changed and unchanged do not, then holds for ten minutes. Split the counter by outcome. For error, check CoreDNS, NetworkPolicy and the headless Service's http port. For empty, check Ready endpoints. For unmatched, compare SIGLAKE_QUERY_PEER_SELF_NAME with the SRV targets. Until a refresh succeeds, a new pod serves queries without fan-out; a pod with an earlier membership keeps using that stale membership. See Query peer discovery. |
SiglakeTableCacheUnpublished |
Warning | Cache reloads keep exhausting their publication attempts, leaving the affected table entry empty and forcing full metadata loads. Reduce commit frequency with compactor.commitBatch.targetMb / maxAgeSecs, or reduce metadata size with snapshot-expiry retention and interval settings. |
Refusing work¶
| Alert | Severity | Operator action |
|---|---|---|
SiglakeIngestLanesRefused |
Warning | Check whether clients vary X-Scope-OrgID or x-siglake-index; otherwise raise ingester.maxLanes deliberately. Only the lane cap increments siglake_ingest_lane_refused_total: a refusal at ingester.maxTenants lands on SiglakeTenantsDenied with reason="at_capacity" instead. |
SiglakeTenantsDenied |
Warning | Split siglake_ingest_tenant_denied_total by reason. For header_not_trusted, the shipped single-tenant default refused an X-Scope-OrgID naming a tenant other than default; a header naming default is accepted. Stop sending the unintended header, or route tenancy by verified identity with ingester.oidc.tenantClaim, or set ingester.trustScopeHeader if a gateway sets the header and strips caller-supplied values. For not_allowed, fix the client's tenant or add the intended tenant to ingester.allowedTenants. For claim_missing, fix the identity provider or the configured claim name. For claim_invalid, send a value of 1 to 128 characters from [A-Za-z0-9_-]. For header_mismatch, remove the contradictory X-Scope-OrgID header. For at_capacity, the pod is at ingester.maxTenants and refused a tenant it had not admitted before; the tenants already writing on that pod are unaffected. Raise the cap or name the tenants you expect in ingester.allowedTenants, remembering that the count is per pod and starts empty on restart. The rule has no reason selector, so every label feeds it. All six series are pre-registered. Each label and the setting that emits it is in Ingest tenant selection. |
SiglakeQueryShardPinUnresolved |
Warning | Workers cannot resolve the generation pinned by the coordinator. Distributed queries return 503 with reason: "shard_pin_unresolved"; see Error responses. Read the worker log for which half of the pin it refused. For a snapshot, raise compactor.snapshotExpire.retainLast or lower SIGLAKE_ICEBERG_METADATA_CACHE_TTL_SECS. For a schema id, check that every worker reads the same warehouse and that its metadata refreshes are succeeding. The rule fires on a non-zero miss rate for five minutes. Both outcomes are pre-registered. |
SiglakeQueryTenantsDenied |
Warning | Verified tokens lack a usable query.oidc.tenantClaim, so queries return 403. Split siglake_query_tenant_denied_total by reason. For claim_missing, fix the identity provider or configured claim name. For claim_invalid, send a value of 1 to 128 characters from [A-Za-z0-9_-]. Both series are pre-registered. |
Saturation¶
| Alert | Severity | Operator action |
|---|---|---|
SiglakeQueryScanAttributionIncomplete |
Warning | A scan partition was still unwinding two seconds after the root stream ended. The response was returned, but its stats.scan omits that partition's counters. Confirm the affected request has stats.scan.unsettled_partitions > 0 and compare it with the warning scan partitions still unwinding at the settle deadline; the alert recovers on its own when subsequent requests settle completely. |
SiglakeQueryPoolNearLimit |
Warning | Reduce query pressure or add query memory/capacity before spilling and reduced file concurrency degrade service. |
SiglakeQueryPoolRefusing |
Warning | Compare SiglakeQueryPoolNearLimit, then make an authenticated GET /debug/memory-pool request to the affected query pod. Review query.resources.limits.memory and the query.spill.maxBytes / query.spill.sizeLimit sizing. The warning recovers when pressure drops, but sustained trips mean queries are receiving retryable 503 responses. |
SiglakeQueryPoolReservedWhileIdle |
Warning | While the residual is present, make an authenticated GET /debug/memory-pool request to the affected query pod. Compare its ten largest live DataFusion consumers with the reservation-owner warning and siglake_query_memory_pool_idle_residual_reports_total. Consumer tracking is on unless SIGLAKE_QUERY_MEMORY_POOL_TRACK_CONSUMERS is 0 or off; if it is disabled, re-enable it before the next occurrence. Restart a pod with a wedged warm probe, or investigate a named operator as a memory leak. These diagnostics were added by Siglake's September 3, 2026 query-memory observability change. |
SiglakeQueriesAbandoned |
Warning | Investigate client timeouts and disconnects, then check the query pod for the degradation that accumulated abandonments can precede. |
SiglakeQueryWarmCycleStalled |
Warning | Restart the pod, then inspect siglake_query_warm_cycle_in_progress to distinguish a wedged cycle from a dead warm loop. This is the shipped signal for the query degradation that /healthz cannot see: browse requests can time out while health stays green. |
SiglakeQueryWarmCyclesAbandoned |
Warning | Investigate slow storage or cancellation; repeated alerts are an early query-pod wedge signal, so restart the affected pod if degradation follows. |
Logs¶
RUST_LOG controls filtering; the chart's logLevel default is
info,siglake=info. For debugging, info,siglake=debug.
Container runtimes capture the server roles' tracing logs from stderr; the server roles write nothing operational to stdout.
SIGLAKE_LOG_EXEC_PLANS logs physical query plans: very verbose, useful when
a fast path isn't engaging and you need to see what DataFusion actually built.
Query audit as an observability surface¶
query_audit is an ordinary Iceberg table, so your query history is
queryable:
-- most expensive query shapes in the last day
SELECT query, count(*) AS runs, avg(duration_ms) AS avg_ms,
max(estimated_bytes_scanned) AS max_bytes
FROM query_audit
WHERE timestamp >= now() - INTERVAL '24 hours'
AND complexity IN ('large', 'huge')
GROUP BY query
ORDER BY avg_ms DESC
LIMIT 20;
-- rejections and errors
SELECT status, count(*) FROM query_audit
WHERE timestamp >= now() - INTERVAL '1 hour'
GROUP BY status;
The table contains literal SQL text. Treat it as sensitive and rotate it
with siglake audit-rotate --max-age-secs N.