Metrics reference¶
Siglake rename
Product names, commands and links were normalized during the Siglake migration. Retained generation dates and commit IDs below identify the archived pre-rename source and binaries; they are not new build evidence.
Each Siglake role exposes Prometheus metrics on a separate /metrics port.
This page lists the main signals, labels and export types.
| Role | Port | Flag |
|---|---|---|
| Ingester | 9100 | --metrics-bind |
| Compactor | 9101 | --metrics-bind |
| Query server | 9105 | --metrics-bind / SIGLAKE_QUERY_METRICS_BIND |
The Helm chart can render a ServiceMonitor with serviceMonitor.enabled and
a 39-alert PrometheusRule with prometheusRule.enabled. Both values default
to off and require the Prometheus Operator. Monitoring
lists the alerts and operator actions. The chart does not render
deploy/grafana/siglake-overview.json; import that dashboard separately. See
Starter Grafana dashboard
for its panels and related alerts.
Selected operational metrics¶
These tables group selected metrics by operational purpose.
Build provenance¶
siglake_build_info exposes the binary's version and source commit labels.
These build provenance labels identify the running binary. Confirm that it
matches a released version by comparing both labels with the intended release
and across every scrape target.
Build information metric¶
| Metric | Watch for |
|---|---|
siglake_build_info{version,commit} |
The version and source commit running in each process. Compare labels with the intended release and across scrape targets to find a mixed-version fleet. |
Ingest health¶
| Metric | Watch for |
|---|---|
siglake_events_accepted_total |
The top-line ingest rate. |
siglake_ingest_backpressure_rejected_total |
Non-zero means Siglake is refusing clients. Compare it with lane depth. |
siglake_ingest_backpressure_queue_events |
Lane depth. Sustained growth means the writer can't keep up. |
siglake_ingest_rate_limit_rejected_total |
Rate budget exhaustion. |
siglake_ingest_mem_breaker_open |
The RSS circuit breaker is shedding load. |
siglake_ingest_tenant_denied_total{reason} |
Batches refused with 403 before Siglake creates tenant state. Reasons cover a header selecting a tenant on a single-tenant ingester (header_not_trusted), tenants outside ingester.allowedTenants (not_allowed), a novel tenant at ingester.maxTenants (at_capacity), missing claims (claim_missing), invalid claims (claim_invalid) and claims contradicted by X-Scope-OrgID (header_mismatch). |
siglake_wal_identity_resolve_failures_total |
Failed index table identity lookups. Each increment is one lookup during a lane's initial binding or refresh. Existing writers keep their binding after a refresh failure; a new writer remains unbound after an initial failure. An increase does not by itself mean a retained identity is stale. |
siglake_wal_writer_rebinds_total |
WAL writer bindings to a different table identity. Each increment is one lane changing from its current binding, including the initial binding at startup. One startup increment per lane is normal; later increments mean the lane learned that the index was recreated. |
siglake_wal_segments_sealed_total |
Segment roll rate, which sets the freshness floor. |
siglake_wal_segments_sealed{tenant} |
Shared catalog-claim queue depth published by the ingester for a zero-floor compactor. It counts sealed rows and claims older than SIGLAKE_CLAIM_RECLAIM_MAX_AGE_SECS; live claims are excluded. |
siglake_wal_segments_sealed_sample_age_seconds |
Age of the last successful shared-queue sample. A failed catalog read holds the previous depth while this age rises. The operator discards a publisher after two minutes. |
siglake_wal_crc_mismatch_total |
Must remain zero. Non-zero means corruption. |
siglake_wal_ipc_framing_refused_total |
Must remain zero. Each increment is one WAL read refused because the Arrow IPC length framing does not describe a message the segment's bytes can back, so that segment is unreadable and needs investigation. Read attempts, not segments: a retried read counts again. |
All six siglake_ingest_tenant_denied_total{reason} series are
pre-registered, so the first denial after enabling ingester.oidc.tenantClaim
is visible rather than hidden behind an empty graph, and the chart's
SiglakeTenantsDenied alert watches the counter. It is the ingest counterpart
of siglake_query_tenant_denied_total; see
Monitoring.
siglake_wal_ipc_framing_refused_total carries no labels and is
pre-registered at 0 on the ingester, the compactor and the query server, so
the first refusal is visible to increase() on whichever role read the
segment. It covers sealed, legacy and recovered partial segments alike, and
sits beside two counters it never overlaps: siglake_wal_crc_mismatch_total
for a body that failed its CRC check before any framing was walked, and
siglake_wal_partial_tail_dropped_total for a recovered partial whose
complete prefix decoded and whose torn tail was discarded. A decode that fails
after the framing walk passed increments none of the three. The chart's
critical SiglakeWalIpcFramingRefused alert fires on any increase over 15
minutes; see
Monitoring.
Drain and commit¶
| Metric | Watch for |
|---|---|
siglake_compactor_sealed_pending |
The shared catalog-claim backlog. Each worker publishes the same queue total, so deduplicate with max by (namespace, tenant) before summing tenants. Sustained growth means the drain is behind accept. |
siglake_compactor_drain_inflight |
In-flight commits. Should sit near SIGLAKE_DRAIN_CONCURRENCY. |
siglake_compactor_rows_committed_total |
Commit throughput. |
siglake_compactor_commit_duration_seconds |
Per-commit cost. Rising means catalog pressure. |
siglake_compactor_watchdog_trips_total |
A drain cycle wedged. |
siglake_compactor_segments_poisoned{tenant} |
Segments the local filesystem drain is holding under <wal>/poison/ for this tenant, summed over its events directory and every managed index. Above zero means those rows are durable and unqueryable until someone requeues them. Each complete sweep republishes the level, and a tenant that drains to empty reports zero; a poison/ the sweep cannot list counts zero for that sweep. SiglakeSegmentsQuarantined fires on this gauge as well as on siglake_catalog_claim_quarantined. |
siglake_compactor_segments_poisoned_total{tenant} |
One set-aside event: a segment that failed to read on SIGLAKE_COMPACTOR_POISON_ATTEMPTS consecutive drain cycles, 3 by default, and was moved under poison/ with a note saying why. A cycle that reads the segment clears its charge, and 0 switches the set-aside off so the drain retries forever. A segment that is requeued and set aside again counts twice here and once on the gauge, so read this counter as an event rate and the gauge as the standing backlog. An unlabelled series is pre-registered at 0, so a fresh compactor's first set-aside is visible to increase(). |
siglake_compactor_wal_owner_mismatch_total{tenant,index} |
An index was deleted and recreated under the same id while WAL segments were still on disk, so the drain refused to commit the dropped incarnation's rows as the replacement's. On the filesystem drain each increment is one segment file moved to stale/<dropped-uuid>/; on the catalog-claim drain it is one mirror prefix re-stamped for the live table. Quarantined segments are held, never deleted, so reclaiming them is a decision you make on visible files. See Index management. |
siglake_compactor_wal_stale_segments_total{tenant,index} |
One object refused on its own frame identity: a header naming the dropped table, or, under a prefix that names a dropped table, an object carrying no identity. A refused mirror claim enters terminal quarantine on its first cycle. This counter records refusal events, so a failed filesystem quarantine or a manually requeued mirror claim can produce another increment. |
Layout quality¶
| Metric | Watch for |
|---|---|
siglake_table_overlap_depth |
Layout-health metric. Converged tables settle low; a plateau means compaction is behind after you rule out a stale sample. |
siglake_table_live_data_files |
Unbounded growth means compaction is losing. Read exactly from the snapshot summary on every 30 s sampler cycle, regardless of table size. |
siglake_table_level_files{level} |
Per-level distribution (leveled mode); like the leading-edge and overlap-depth gauges, this can lag when the manifest walk times out. |
siglake_table_gauges_sampled_at_seconds |
Last successful manifest-walk time for the level, leading-edge, and overlap-depth gauges. Use this to detect stale values. |
siglake_compactor_maintenance_skipped_backpressure_total |
Compaction yielding to the drain. |
siglake_compactor_maintenance_throttled_backpressure_total |
Throttled single-merge passes under backlog. |
The manifest-walk gauges, siglake_table_level_files,
siglake_table_leading_edge_small_files,
siglake_table_leading_edge_small_bytes and siglake_table_overlap_depth,
share a per-table manifest-walk budget set by
SIGLAKE_GAUGE_TABLE_TIMEOUT_SECS (default 20 seconds). The sampler updates
them and siglake_table_gauges_sampled_at_seconds only after a complete walk;
siglake_table_gauge_table_timeouts_total counts walks that did not finish.
siglake_table_live_data_files does not share that staleness: the snapshot
summary's total-data-files value is published before the walk.
siglake_table_live_data_files_sample_failures_total retains its broader
meaning that the sampler sweep did not finish; it does not mean that a
published snapshot-summary count was approximate.
Query¶
| Metric | Watch for |
|---|---|
siglake_query_requests_total |
Completed requests, by bounded endpoint and final HTTP status. All four Jaeger read routes use endpoint="jaeger"; the endpoint label never contains an index, service, trace id, raw URL, or request text. |
siglake_query_request_duration_seconds |
Latency distribution. |
siglake_query_in_flight |
Concurrency. A KEDA scaling signal. |
siglake_query_exec_pool_queue_seconds |
Queue wait. High means pool saturation, not slow scans. |
siglake_query_breaker_trips_total |
Query refusals and aborts, split by breaker and request priority. |
siglake_query_scan_attribution_incomplete_total |
Requests returned before every scan partition folded its counters into stats.scan; inspect stats.scan.unsettled_partitions and the query-server warning log. |
siglake_query_scan_settle_seconds |
Time spent waiting for scan partitions to fold their counters before rendering stats.scan; values at the two-second deadline accompany incomplete attribution. |
siglake_query_scan_file_cache_bytes |
Bytes in completed decoded-file cache entries on this query pod. The series is absent until the pod's first insert. |
siglake_query_scan_file_cache_accounted_bytes |
Completed entries plus admitted in-flight populations, which is the total enforced against SIGLAKE_QUERY_SCAN_FILE_CACHE_MAX_BYTES. It starts at 0; subtract the completed-entry gauge to see bytes held by live populations. |
siglake_query_scan_file_cache_entries |
Completed decoded-file cache entries on this query pod. |
siglake_query_scan_output_ordering_total{outcome} |
Whether ordered early-stop engaged: advertised, or the refusal reason. filtered (a non-timestamp predicate) is the common and expected one; no_bounds, fan_in, and global_fan_in point at layout. Full label list under Ordered early-stop. |
siglake_query_wal_buffer_rows |
Rows served from the freshness buffer. |
siglake_query_wal_buffer_segment_errors_total |
Buffer read failures. Fresh data may be missing. |
siglake_query_wal_buffer_stale_owner_total |
A WAL-buffered read declined a directory whose owner marker names a dropped incarnation of the index, and served the committed view only. Each increment is one such check, from a local buffer read, the count fast path, or a distributed buffer partial. Non-zero means an index was deleted and recreated with WAL segments still on disk; siglake_compactor_wal_owner_mismatch_total names the tenant and index. |
siglake_query_wal_buffer_stale_segments_total |
WAL segments excluded from a buffered read because their frame identity does not match the live table. Each increment is one segment excluded from one read; repeated reads can count the same file again. |
siglake_query_tenant_denied_total{reason} |
Verified tokens refused with 403 before namespace or tenant-context creation. claim_missing means the configured claim is absent, claim_invalid means it is not a usable tenant id, and not_allowed means the usable claim is absent from query.allowedTenants. All three series are pre-registered at zero. |
siglake_query_jobs_cancel_propagated_total |
Batch jobs a replica aborted because a DELETE served by another replica had persisted cancelled; the executor finds them on its --jobs-cancel-poll-secs sweep. A normal-operation counter, not a fault: in a multi-replica tier it should track client cancellations, and sitting at zero while jobs are being cancelled means either every DELETE happened to land on the executor or the poll is not running. |
siglake_query_job_terminal_conflict_total{attempted,actual,cause} |
A batch lifecycle transition the store refused. attempted is the proposed state. actual is the stored state. cause identifies recovery, cancellation, TTL deletion, an oversized result or a terminal write that did not land. Refused starts and discarded results do not increment siglake_query_jobs_total. An oversized result is stored and counted as failed. |
siglake_query_job_terminal_write_failed_total{attempted} |
A failed or timed-out terminal-write attempt. A run makes up to three attempts within 30 seconds. Compare this metric with the query-server warnings persisting a batch job's terminal state failed and persisting a batch job's terminal state ran out of time. |
siglake_query_jobs_unreconciled |
Finished runs whose terminal write did not land and whose IDs this replica retries every five seconds. Clients read running until reconciliation installs the verdict. |
siglake_query_jobs_unreconciled_dropped_total |
A finished run's id that reconciliation bookkeeping refused because it was already holding its cap of 1,024 ids. Each increment is a row stranded non-terminal until this replica exits, and is paged by SiglakeBatchRowStrandedNonTerminal. |
siglake_query_jobs_reconciled_total{computed,outcome} |
Result of owner-local reconciliation. computed is the run's verdict. outcome is the reconciled status, a terminal status that already won or gone after TTL deletion. A lost success becomes failed with a resubmit instruction. Siglake does not run the query again. |
siglake_query_breaker_trips_total has a breaker label with these values:
admission, timeout, shard_timeout, preflight_bytes, pool_exhausted,
midflight_rows_shard, midflight_rows_scanned, and
midflight_rows_scanned_ndjson, plus jaeger_trace_limit,
jaeger_span_rows, jaeger_render_bytes, and jaeger_name_rows. Its
priority label is either interactive or batch. The pool_exhausted
series covers both memory-pool and spill-cap refusals; the three
midflight_rows_* series identify SQL row-ceiling trips, while the four
jaeger_* series identify the Jaeger ceiling that refused a read and always
have priority="interactive".
Read attempted first: running means the store refused the run's start
publication, so the query was dropped unpolled and nothing was executed; any of
succeeded, failed and timeout means a completed run's verdict was refused
and its computed output discarded. The cause label then makes the conflict
actionable:
cancellation: The run finished just after a client cancellation landed.attempted="succeeded", actual="cancelled"is the common shape, and a low rate is expected anywhere clients cancel.recovery: Recovery installed the monotonicfailedstate before the executor finished. If the executor had not started, its refusedrunningpublication prevented the query from running at all; if it had, the computed output was discarded.SiglakeBatchCompletionRejectedByRecoverycovers both shapes.gone: The row was deleted before the run finished, usually by the job TTL sweep.result_too_large: A success body exceeded the job-row limit and was installed asfailed. This is the only conflict that also counts insiglake_query_jobs_total, asfailed.write_deferred: The terminal write never landed, soactual="unknown", but the run's own executor parked the id and retries the same conditional write every five seconds. The row readsrunninguntil that pass resolves it, which needs no operator; watchsiglake_query_jobs_unreconciledfall back to zero.write_abandoned: The terminal write never landed and nothing is retrying it: reconciliation bookkeeping was full, so the row stays non-terminal until this replica exits and lease-expiry recovery on a peer condemns it.actual="unknown"and the answer the run computed is gone. This is the second cause covered by a chart alert,SiglakeBatchRowStrandedNonTerminal.other: Another terminal state won without matching one of the cases above; inspectactualand the query-server warning.
Do not put a rate threshold on either counter as a whole:
siglake_query_jobs_cancel_propagated_total is a normal-operation counter, and
an unlabelled siglake_query_job_terminal_conflict_total rate pages on
ordinary client cancellations and TTL expiry. Watch both on the starter
dashboard's Batch job cancellation across replicas panel. The chart warns
whenever recovery refused executor work, of either shape:
increase(siglake_query_job_terminal_conflict_total{cause="recovery"}[10m]) > 0,
and pages per pod on the abandoned write
(SiglakeBatchRowStrandedNonTerminal, which also reads
siglake_query_jobs_unreconciled_dropped_total). Cancellation, TTL expiry and
write_deferred do not alert. Seven series are pre-registered so the first
conflict of each is visible to increase(): the four cause="recovery" ones
(attempted is running, succeeded, failed, or timeout; actual is
failed) and the three cause="write_abandoned" ones (attempted is
succeeded, failed, or timeout; actual is unknown). Only a finished
run's verdict can be deferred or abandoned: a job-store error on the start
publication is no conflict at all, and the run executes as before. The
unlabelled siglake_query_jobs_unreconciled_dropped_total is pre-registered
too, for the other half of that alert. Other causes appear on first occurrence.
Which WARN line to read follows from attempted. A refused completion logs
batch completion was refused; computed output was discarded with the job id,
owner, attempted outcome, superseding status, cause, and dropped rows and
bytes. A refused start logs batch lifecycle transition was refused with
phase="running", the job id, owner, superseding status, and whether the row
was recovered; it reports no dropped rows, because there were none. A write
that never landed logs neither: write_deferred logs the WARN batch terminal
state could not be persisted; deferred to owner-local reconciliation, and
write_abandoned the ERROR batch terminal state could not be persisted and
could not be tracked, both with the job id, owner and attempted verdict. See
Batch jobs for completion and start-refusal behavior.
Cache effectiveness¶
SIGLAKE_FOOTER_CACHE_CAP, SIGLAKE_AGG_RESULT_CACHE_CAP and
SIGLAKE_DATA_FILE_LIST_CACHE_CAP set the corresponding cache capacities.
SIGLAKE_QUERY_RESULT_CACHE controls whole-response caching. Siglake does not
define a healthy cache hit rate. Interpret the cache effectiveness metrics
against query repetition and snapshot churn; the configuration options change
capacity, not a target hit rate.
Cache counters¶
| Metric | Watch for |
|---|---|
siglake_footer_cache_hits_total / _misses_total |
Footer cache hit rate. |
siglake_agg_result_cache_hits_total / _misses_total |
Aggregate memoization. |
siglake_file_list_cache_hits_total / _misses_total |
File-list cache. |
siglake_query_sql_result_cache_requests_total |
Whole-response cache. Cache hits distort repeated-query benchmarks. See Performance. |
siglake_query_sql_result_cache_bytes |
Heap the whole-response cache retains, against its 4 MiB cap: each entry's body, the lookup string it is stored under, and one pointer per recency marker. Query text is charged here, so long queries leave less room for result bodies. |
Text-index startup stages¶
siglake_iceberg_text_index_startup_seconds{stage,storage} times what a text
query spends on a planned indexed file before that file produces rows. The name
ends in _seconds, so it has _bucket series and histogram_quantile()
applies. storage is puffin or footer_kv.
stage |
What it times |
|---|---|
permit_wait |
Waiting for the global index-load semaphore, SIGLAKE_INDEX_LOAD_CONCURRENCY (4 by default). Both index forms pay it on a cold read. |
blob_fetch |
The Puffin blob read from object storage. A footer-KV index has no such stage: its bytes arrive with the Parquet metadata the scan already read. |
decode |
InvertedIndex::from_bytes, about 30 ns per indexed row. Recorded only when the parsed-index cache missed. |
selection |
The postings lookup plus the row-selection runs. Recorded on every file the index prunes. |
Read the sample counts as well as the quantiles. decode's count is the number
of cold loads, and selection's count is the number of files the index pruned.
So decode climbing toward selection means the plan is not staying warm, and
each stage's share of the total says which of the four to chase.
Parsed text-index cache outcomes¶
| Metric | Watch for |
|---|---|
siglake_iceberg_parsed_index_cache_lookups_total{outcome,storage} |
Exactly one hit or miss per file a text query acquires an index for, so the hit ratio reads directly. The repeated lookups one cold load makes count as a single miss. All four series are pre-registered at 0 on the query server. |
siglake_iceberg_parsed_index_cache_evictions_total{reason} |
An index the cache would not keep. This separates a first read from an entry the cache decoded and threw away, which a hit ratio cannot. All three series are pre-registered at 0. |
siglake_iceberg_parsed_index_cache_bytes |
Parsed bytes resident on this pod. |
siglake_iceberg_parsed_index_cache_max_bytes |
The byte budget in force on this pod. |
reason |
Which bound dropped it |
|---|---|
byte_bound |
The working set went over SIGLAKE_PARSED_INDEX_CACHE_MAX_BYTES, so the least-recently-used entry was evicted. |
entry_bound |
The entry count went over the shared SIGLAKE_PUFFIN_BLOB_CACHE_MAX_ENTRIES. |
oversized |
One index alone exceeds the whole byte budget, so it is never admitted and that file decodes on every query. |
Both gauges are published where the bounds are enforced, after a successful
insert into the cache; an oversized refusal and a zero bound publish nothing.
They are absent until the pod's first indexed text query, and they hold their
last value while no text query runs. Do not read either one as a liveness
signal.
Puffin and footer-KV indexes share this cache under the same bounds, so
storage splits both arms of the lookup counter. The three eviction reasons map
to the three knobs described in Tune the Puffin and parsed text-index
caches.
Puffin blob cache outcomes¶
The Puffin blob cache holds the compressed index bytes a decode reads, one
layer below the parsed cache. A parsed miss is what brings a query here, so
read these three counters beside
siglake_iceberg_parsed_index_cache_lookups_total.
| Metric | Watch for |
|---|---|
siglake_iceberg_puffin_blob_fetches_total |
Index blobs read from object storage. Unlabelled, and charged on every fetch, including the reads no lookup preceded. |
siglake_iceberg_puffin_blob_cache_lookups_total{outcome} |
One hit or miss per file a text query decodes an index for. A lookup happens only while both blob-cache bounds are positive; with either at 0 the cache is off and neither arm moves. |
siglake_iceberg_puffin_blob_cache_evictions_total{reason} |
Evicted entries and refused admissions, grouped by the reason. |
siglake_iceberg_puffin_blob_cache_bytes |
Compressed blob bytes resident on this pod. Published on every admission attempt, including a refused one. |
siglake_iceberg_puffin_blob_cache_max_bytes |
Enforced byte budget. An entry bound of 0 publishes 0 here even when the byte knob is positive, because the cache cannot admit an entry. |
reason |
What happened |
|---|---|
stale |
Nothing read the blob while the cache turned over four times. Tried first. The test is inactivity, not proof that the file left the plan. |
redundant |
The blob's parsed twin is resident, so the blob cannot be read until that twin goes. Tried when nothing is stale; among the candidates, the one whose twin sits furthest from parsed eviction loses. |
fifo |
Neither arm applied, so the oldest entry goes. This is how every eviction worked before the two caches were coupled, and it can drop a blob that is live and recent. |
oversized |
One blob exceeds SIGLAKE_PUFFIN_BLOB_CACHE_MAX_BYTES, so the cache refuses it before copying it and evicts nothing. Each decode fetches the blob again. |
Blob hits climbing with parsed misses is a working set past the parsed budget
but inside the blob budget: a re-decode that needed no re-read. A fetch rate
that tracks the parsed miss rate is the opposite reading: the refetch shape the
coupled eviction rule fixed. A fetch rate above the miss rate is index
reading that consulted no cache: a rebuild that bypasses it, or a pod with
SIGLAKE_PUFFIN_BLOB_CACHE_MAX_BYTES or
SIGLAKE_PUFFIN_BLOB_CACHE_MAX_ENTRIES at 0. All seven series are
pre-registered at 0 on the query server, so a tier serving no text query charts
zero rather than no data. Both knobs are described in Tune the Puffin and
parsed text-index
caches.
Footer text-index checksum refusals¶
siglake_index_footer_checksum_refused_total{reason} counts footer-KV inverted
indexes the reader refused because the sibling CRC-32 stored beside them did
not check out. The check runs before the parsed-cache lookup, so a warm index
is covered as well as a cold one. The reader then tries a Puffin index for that
column, and scans the file exactly when there is none. A rising rate is
correctness holding at the price of scan work.
reason |
What the reader found |
|---|---|
malformed |
The sibling entry is present but unusable: an empty value, a value that is not exactly eight hexadecimal characters, or a blob that is not even-length hexadecimal. |
mismatch |
Both values parsed and the CRC-32 of the stored blob disagrees with the stored checksum. This is the corruption case. |
Both series are pre-registered at 0 on query pods, so the first refusal charts as a step rather than replacing No data. A file written before 0.2.0 carries no sibling checksum, so zero on a fleet mid-upgrade says nothing about the files it has not rewritten. The dashboard graphs both reasons on Footer text-index checksum refusals in the Fast paths row, and ships no alert for them.
Decoded-file cache outcomes¶
siglake_query_scan_file_cache_requests_total{outcome} counts what the
experimental decoded-file cache did with each file task a scan opened, and with
each population that task started, under the nine outcomes below.
outcome |
What it counts |
|---|---|
hit |
A cached entry served the file task, so no reader was built. |
miss |
A file task opened on the cached path with no predicate or prune spec to bypass with, and nothing cached for that file and scan direction, so a population was attempted for it. Charged per opened task, not per query. |
bypass |
The file task carries a predicate the Iceberg converter accepts, a raw-text prune spec or a promoted-column prune spec, so it reads through the pruning reader with its predicate intact and never populates. Charged in place of miss, not alongside it. |
insert |
A population read its file to the end and its batches were kept. |
insert_skipped_contended |
A finished population lost the cache lock and dropped its batches instead of waiting. One skip costs one later miss. |
skip_oversized |
One file's buffered batches crossed a quarter of SIGLAKE_QUERY_SCAN_FILE_CACHE_MAX_BYTES, so they are never admitted. |
abandoned |
A population was dropped before its end-of-stream insert: a LIMIT satisfied from the first batches, or a cancelled query. Charged whether or not the population had buffered anything yet. |
evict |
The oldest entry was dropped, by insertion order. An insert that crosses the byte or entry bound evicts, and so does a live population that needs room for its next batch under the shared byte bound. |
population_refused |
The shared byte bound refused one population because no room could be made or another population held the replacement lock. Counted once for that population and never with abandoned. |
A column comparison, an IN list, IS [NOT] NULL or a prefix LIKE bypasses
whenever the Iceberg predicate converter accepts it, including a timestamp
range. Order-preserving scans never reach this path, so a newest-first browse
records neither outcome. A predicate query is still served by an entry a
predicate-free scan left behind, and DataFusion's residual filter keeps that
answer exact.
The outcomes do not partition miss. miss is charged before the population
stream is built, and hit and bypass never build one, so a stream that fails
to construct leaves a miss with no second outcome. A population that went
oversized charges skip_oversized and never abandoned, and a population that
ended in an error counts as finished rather than abandoned. A population the
shared byte bound refuses records population_refused once. The label does
not distinguish a full budget from a busy replacement lock. See What the
decoded-file cache byte limit
covers.
siglake_query_scan_file_cache_accounted_bytes is the total the shared byte
bound enforces: completed entries plus admitted live populations. It starts at
0 and updates wherever that total moves. siglake_query_scan_file_cache_bytes
reports completed entries only, and siglake_query_scan_file_cache_entries
reports their count.
The cache is on only when both SIGLAKE_QUERY_SCAN_FILE_CACHE_MAX_BYTES and
SIGLAKE_QUERY_SCAN_FILE_CACHE_MAX_ENTRIES are positive, and the chart, the
Compose file and the operator all set them to 0. All nine outcome series are
pre-registered on the query server, so a default install reads 0 on every arm
rather than absent. The knobs, and the narrow set of scan shapes that fill an
entry, are described in Cache
tuning.
Decoded-file cache population depth¶
siglake_query_scan_file_cache_populate_rows{outcome} is a histogram with one
observation per population of the decoded-file cache: the rows the reader handed
that population before it stopped.
outcome |
What it records |
|---|---|
completed |
The population read its file to the end. The entry may or may not have been kept, and the depth is the same either way. |
clipped |
The population was polled and then dropped before end-of-stream. A LIMIT satisfied from the first batches and a cancelled query both land here, because the stream cannot tell them apart. |
unpolled |
The plan opened the task and never polled it, so the observation is 0. The same task is charged miss and abandoned. |
error |
The read failed. |
The count is cumulative over the stream and is taken before the residual filter above the scan drops any rows. It keeps rising after a candidate crosses the entry bound, so it measures how far the read got, not what the cache kept and not what the query returned.
The buckets carry an edge at 131,071, so +Inf minus le="131071" is the exact
count of populations handed 131,072 rows or more, the write path's row-group
floor. Reaching that floor does not prove a row group was closed: a file whose
groups are larger closes none at that depth, and a read that does not start on a
group boundary closes none at any depth.
Unlike the nine counter arms above, this histogram is not pre-registered. It is
absent until the cache is on and a population ends, so a default install exports
no series for it. The same depth per request is
stats.scan.file_cache_populate_rows.
Storage maintenance¶
| Metric | Watch for |
|---|---|
siglake_gc_orphans_deleted_total / siglake_gc_bytes_reclaimed_total |
GC effectiveness. |
siglake_compactor_snapshots_expired_total |
Snapshot trimming. |
siglake_iceberg_commit_duration_seconds |
Catalog commit cost. |
siglake_wal_mirror_failures_total |
Non-zero means disaster-recovery mirroring is failing. |
siglake_wal_mirror_upload_attempts_total{path} |
One application-level object-store write call, including retries made by Siglake. path is sealed or active. Local source failures and retries inside the client or service are excluded, so this is not a billed S3 request count. |
siglake_wal_mirror_active_bytes_uploaded_total |
Bytes written by the active-segment uploader. Subtract this from siglake_wal_mirror_bytes_uploaded_total to obtain sealed-uploader bytes. Active mirroring remains off at its default interval of 0. |
siglake_dropped_index_cleanup_records_total |
Incarnation-bound cleanup records confirmed before an index catalog entry was removed. Production records are report-only. |
siglake_dropped_index_aggregate_objects_deleted_total |
Aggregate objects removed by a separately authorized cleanup record. Production-created records do not grant that authority, so this remains 0 on the shipped path. |
siglake_dropped_index_aggregate_cleanup_failures_total |
Object deletes that failed after an authorized cleanup sweep began. Report-only inventory does not increment it. |
siglake_side_aggregate_publish_failures_total{iceberg_namespace,table} |
A publication of the table's inline aggregate object spent all four attempts, so that commit's group counts and time aggregates are lost. Answers stay exact; windowed GROUP BY on the table moves to the per-file path. SiglakeSideAggregatePublicationLost fires on it. See Data loss or durable inconsistency. |
Attribute auto-promotion¶
Five metrics report what the compactor's attribute auto-promotion pass did to
one table's schema. Every series carries iceberg_namespace and table, and a
compactor publishes one set per table the pass visits: the events table plus
every user index it can list. The gauges refresh on each re-clustering pass;
sampling itself runs at most once per 300 seconds per process, so the pass
counter moves on that cadence. Auto-promotion is off by default
(SIGLAKE_AUTO_PROMOTE_MIN_PCT=0; see
Indexing and promotion), and a
disabled compactor still publishes the gauges and pre-registers all four pass
outcomes at 0.
| Metric | Watch for |
|---|---|
siglake_auto_promotion_enabled{iceberg_namespace,table} |
1 when this compactor would promote: SIGLAKE_AUTO_PROMOTE_MIN_PCT is above zero and SIGLAKE_AUTO_PROMOTE_MAX_COLUMNS is not 0. 0 is a pass that is switched off, which is the shipped default. An absent series is a different reading again: no compactor has visited that table. |
siglake_auto_promotion_passes_total{iceberg_namespace,table,outcome} |
One increment per pass over that table. promoted added at least one column, nothing_cleared sampled and added none, at_ceiling declined to sample because the table already holds its configured cap, and failed could not read the table metadata or could not finish sampling. All four outcomes are pre-registered at 0 for every table a compactor visits, switched-off ones included, so zero is a measurement rather than an absence. |
siglake_auto_promotion_candidates{iceberg_namespace,table} |
Unpromoted attribute names the last sampled pass found in at least SIGLAKE_AUTO_PROMOTE_MIN_PCT of sampled rows. It includes the names that pass then refused for the cap, a schema name collision or mixed sampled types, so it is the size of the queue and not the size of the next widening. It is a last observation and is never cleared, so read it only where the gauge below is 1: on a table that has never sampled it is absent, and on one that stopped sampling it holds the earlier pass's number. |
siglake_auto_promotion_candidates_available{iceberg_namespace,table} |
1 when the reading above comes from a pass that sampled. 0 when the pass deliberately did not sample, because promotion is off or the table is at its cap, and when the pass failed. An unavailable candidate count is not a measured zero: those paths reach no verdict about candidates at all. |
siglake_auto_promotion_columns{iceberg_namespace,table,kind} |
kind="limit" is the effective SIGLAKE_AUTO_PROMOTE_MAX_COLUMNS on that compactor. kind="used" is the length of the table's promoted-column list, which counts the columns you declared by hand as well as the automatic ones. used / limit is the fraction SiglakeAutoPromotionNearCeiling reads. |
The restored cost series price sampling separately from the backfill it can trigger. They do not mean that automatic promotion is qualified for an object-store workload.
| Metric | What it measures |
|---|---|
siglake_auto_promotion_sample_reads_total{iceberg_namespace,table} |
Object-store reads made by successful sampling passes. Pre-registered at 0 for each table the compactor visits. |
siglake_auto_promotion_sample_bytes_total{iceberg_namespace,table,phase} |
Bytes attributed to footer, index or data work during sampling. All three phases are pre-registered at 0. |
siglake_auto_promotion_pass_duration_seconds{iceberg_namespace,table} |
Duration of one successful sampled pass. |
siglake_compactor_promotion_backfill_files_total{table} |
Input files rewritten by committed promotion-backfill bins. |
siglake_compactor_promotion_backfill_bytes_in_total{table} / _bytes_out_total{table} |
Input and output bytes for committed backfill bins. |
siglake_compactor_promotion_backfill_duration_seconds{table} |
Duration of each committed backfill bin. |
The attribute names are in the compactor's log, not in a metric. A pass that
sampled logs auto-promotion pass finished at INFO with its outcome, the
candidate count, the column count before and after, and up to 32 names per
disposition: promoted, declined_ceiling, declined_name_collision and
declined_mixed_type. Each list carries a _truncated count of the names it
left out, so the totals stay exact while one line stays bounded. A table at its
cap logs auto-promotion pass declined sampling with candidates=unavailable
instead, and a failed pass logs a WARN naming whether the metadata read or the
sampling was what broke.
SiglakeAutoPromotionNearCeiling warns when a table with promotion enabled has
used at least 80% of its configured cap for 10 minutes. It reads each table
against that compactor's own cap, not against a universal column ceiling. See
Saturation and
What is attribute auto-promotion doing to my schema?.
Deferred text-index registration¶
siglake_index_registration_deferred_total{reason="snapshot_has_statistics"}
counts registration attempts that published no index. Iceberg allows one
statistics file per snapshot, so a registration that finds the snapshot already
carrying one keeps the registered file and drops the blobs it just built. The
compactor logs the deferred data-file paths beside each increment. Those files
stay readable and answer text predicates by scan, so a rising count costs query
time on them, not correctness; they gain an index from a later rewrite or from a
CLI index rebuild against a later snapshot. The counter is pre-registered with
its one reason, so a healthy process reads 0 rather than absent. The dashboard
graphs it by reason over one hour, and the chart ships no alert for it.
Statistics retirement¶
Two counters report the pass that removes Iceberg statistics entries whose
Puffin index blobs describe no data file a retained snapshot still references.
The elected snapshot-expiry pass runs it, and so does
siglake gc-orphans --apply; a dry run classifies the entries and increments
nothing. Removal only makes the Puffin object unreachable. The orphan sweep's
min_age gate still decides when the object is deleted, and
siglake_gc_bytes_reclaimed_total is where its bytes show up. Both counters are
pre-registered, so a table with nothing to retire reads 0.
| Metric | Watch for |
|---|---|
siglake_iceberg_statistics_removed_total |
Entries removed from table metadata by a committed retirement. No labels. An entry that still names one live data file is kept whole and counted by neither metric; the gc-orphans output line reports the eligible, removed, live-kept and skipped counts for the table it ran on. |
siglake_iceberg_statistics_retirement_skipped_total{reason} |
Entries the pass declined to touch. unowned_blob_type means the entry holds a blob type Siglake does not own, which is the case when another writer registered statistics on the table. missing_data_file means an owned entry has a blob with no data_file property, so the pass cannot prove what it describes. Either way the entry and its object stay in the table, and the bytes stay with them. |
Segmented text-index writes¶
Three metrics report the segmented (seg2) sidecar build a compactor re-cluster
runs when SIGLAKE_SEGMENTED_INDEX_WRITES=1. That switch is off by default in
0.2.0; see
Indexing and promotion. All four
(outcome, reason) series of the counter are pre-registered on every
compactor, so a default install reads 0 on each of them rather than absent.
The two _bytes families export as summaries, which have nothing to
pre-register, so each stays absent until that pod closes its first sidecar.
Neither is among the bucketed families: select a
quantile label directly instead of calling histogram_quantile().
| Metric | Watch for |
|---|---|
siglake_iceberg_segmented_index_writes_total{outcome,reason} |
One increment per (output file, indexed column) at outcome="written" and reason="none", and one per refused output file at outcome="refused", because the writer abandons that file's whole sidecar set at the first column it cannot stand behind. A refusal publishes nothing for that file: reason="column" is a row group whose text column could not be indexed, file_rows is a sidecar whose total rows are not the data file's, and row_domain is per-group rows that are not the footer's row groups. A refused file is read the way an unindexed one is, so refusals cost query time and not correctness. Sustained refusals mean the opt-in is buying less than it appears to. |
siglake_iceberg_segmented_index_written_bytes |
Size of each published sidecar blob, one sample per blob. This is the object-storage cost of the opt-in. Read {quantile="0.5"} and {quantile="0.99"} per pod: a summary quantile covers one exporter's rolling window, and a fleet mean of percentiles is a number no pod measured. |
siglake_iceberg_segmented_index_group_index_bytes |
Postings and dictionary the writer holds while encoding one row group, one sample per row group, dropped when that group closes. This is one row group's parsed-index allocation and not the writer's peak process heap, which is why the number is small; it tracks row-group size rather than file size. Read its quantiles per pod, as above. |
Segmented text-index reads¶
Four histograms price the store traffic one seg2 lookup makes: the rounds of
waits it took, the reads its reader made across them, the ranges the store
served, and the bytes those ranges carried. A lookup runs only
when SIGLAKE_SEGMENTED_INDEX_READS is on, which it is not by default in
0.2.0 (see Segmented text indexes: the seg2 write and read
switches),
and only on a scanned file that carries a seg2 sidecar for the predicate's
column. All four record one sample per (file, column) lookup that opened a
sidecar, whatever that lookup concluded, so a declined file sits in the numbers
beside an answered one: a lookup that finds the blob compressed records four
zeros, and one refused for row_group_order records nothing at all. All four
export as summaries and are not among the bucketed
families: select a quantile label instead of
calling histogram_quantile(), and expect them absent on a pod that has read
no sidecar yet. None has a dashboard row or an alert; they report what the
staged reader costs, not a fault.
| Metric | Watch for |
|---|---|
siglake_iceberg_segmented_index_stages |
Rounds of store waits one lookup took: the trailer, the directory, the dictionary blocks the terms name, then the posting sections. That is four for any shape, plus one per extra term of a conjunction, and two fewer when the directory is already held from an earlier lookup on the same blob. A stage's ranges go out together, so this is the lookup's serial depth rather than its read count. A lookup that runs past its stage budget declines the file and is counted at siglake_iceberg_segmented_index_declined_total{reason="stages"}. |
siglake_iceberg_segmented_index_reader_reads |
Reads the reader made, counted across every stage, including the ranges a later stage walks again. Read it against siglake_iceberg_segmented_index_range_reads, which counts what the object store served: a range asked for twice is fetched and counted once there. The gap between the two is re-decoding of bytes the lookup already holds, which costs query-pod CPU and no store traffic. |
siglake_iceberg_segmented_index_range_reads |
Ranges the object store served, deduplicated: a range two terms share, or one a later stage walks again, is fetched and counted once. This is the round trips a source serving one range at a time would have taken, against the _stages the staged reader waited for. |
siglake_iceberg_segmented_index_fetched_bytes |
Bytes those ranges carried: the sliver of the sidecar this lookup read. Read it against the blob sizes on the write side, siglake_iceberg_segmented_index_written_bytes, for what the segmented layout saves over decoding a whole index. |
Segmented text-index lookups: used and declined¶
Two counters and two histograms say what seg2 lookups concluded. A lookup that
answers returns a row selection for the file and is counted once at
siglake_iceberg_segmented_index_used_total; one that concludes nothing about
the file is counted once at siglake_iceberg_segmented_index_declined_total,
under the reason it stopped for. The ratio between them is what the read switch
buys. No series here is pre-registered, so a pod with
SIGLAKE_SEGMENTED_INDEX_READS off, or one whose files carry no seg2 sidecar,
exports none of them rather than zeros. The two histograms export as summaries,
like the cost families above: select a quantile label rather than calling
histogram_quantile().
| Metric | Watch for |
|---|---|
siglake_iceberg_segmented_index_used_total{source} |
Lookups that produced a row selection, one per (file, column). source="fts_udf" is a match, match_any, match_phrase or match_prefix predicate; source="like_substring" is a LIKE the planner turned into terms. The v1 counterpart is siglake_iceberg_inverted_index_used_total, which also carries storage; a file answered here never reaches the v1 path. |
siglake_iceberg_segmented_index_declined_total{reason} |
Lookups that concluded nothing about the file, by reason: see Why a seg2 lookup declines a file. A decline costs the reads the lookup already made and then the file is answered another way, so a rising rate is query time rather than wrong answers. |
siglake_iceberg_segmented_index_resident_bytes |
Parsed directory the answering lookup held, which is the size the directory cache below charges against its budget. Recorded only on the lookups counted at _used_total. |
siglake_iceberg_segmented_index_selected_rows |
Rows that lookup selected in the file, before the scan reads them. Recorded with _resident_bytes on the same answering lookups. A count close to the file's row count is a predicate the index barely narrowed. |
Why a seg2 lookup declines a file¶
siglake_iceberg_segmented_index_declined_total{reason} carries one of nine
reasons. A decline is the safe outcome: the file falls back to its v1 Puffin or
footer index for that column, and to a scan when it has none. A LIMIT query
that reached seg2 through the clipped path has no v1 fallback and scans the
file.
reason |
What the reader found |
|---|---|
row_group_order |
The scan handed the reader row groups that are not strictly ascending. Refused before the sidecar is opened, so this decline costs no read. |
compressed |
The sidecar blob is compressed, so the reader cannot address byte ranges inside it. Every lookup on that blob declines the same way. |
open |
The bytes are there and do not parse as a seg2 index: a truncated blob, or a range the store refused and the reader filled as unreadable. |
row_domain |
The sidecar states row groups that are not the file's, so it describes a different layout and is not trusted to prune. |
stages |
The lookup asked for new ranges past its stage budget of 8 plus 4 per predicate term, which bounds a reader that is not converging. |
unanswerable |
One term's postings could not be read from the bytes the lookup holds, so the index cannot prove anything about that predicate. |
no_hints |
The predicate carried no term the index can look up for that column. |
clipped_estimate_unavailable |
A LIMIT query whose predicate is a substring match. The reader cannot estimate document frequency for one, so it does not spend the reads. |
clipped_document_frequency |
The terms of a LIMIT query match more rows than the limit's budget, so the selection would not save the scan. |
Segmented directory cache¶
The reader holds a parsed seg2 directory between lookups, keyed by sidecar path
and blob offset, so a repeat lookup on the same blob skips the trailer and
directory stages.
SIGLAKE_SEGMENTED_INDEX_DIRECTORY_CACHE_MAX_BYTES sets the process-wide
budget, 64 MiB by default; at 0 the cache is off, no lookup is counted and
neither gauge is published. This cache is separate from the parsed text-index
cache, holds seg2 directories alone, and
charges each one its _resident_bytes plus the path it is filed under.
| Metric | Watch for |
|---|---|
siglake_iceberg_segmented_index_directory_cache_lookups_total{outcome} |
One hit or miss per lookup that opened a sidecar while the cache is on. A hit is two rounds of store waits that lookup did not take. |
siglake_iceberg_segmented_index_directory_cache_evictions_total{reason} |
byte_bound is a least-recently-used directory dropped to get back inside the budget. oversized is an admission refused because that one directory exceeds the whole budget; it evicts nothing, and every lookup on that file re-parses its directory. |
siglake_iceberg_segmented_index_directory_cache_bytes |
Directory bytes resident on this pod. |
siglake_iceberg_segmented_index_directory_cache_max_bytes |
The budget in force on this pod. |
Both gauges are published on admission, so they stay absent until the cache accepts its first directory and hold their last value while no seg2 lookup runs. A refused oversized directory publishes neither. Do not read either one as a liveness signal.
Group-count aggregate labels¶
Five aggregate-maintenance counters name the table they report on with an
Iceberg namespace as well as a table name. One compactor maintains the base
namespace and every tenant_* namespace, each with its own events, so a bare
table="events" merged every tenant onto one series.
| Metric | Labels |
|---|---|
siglake_group_count_short_aggregates_total |
iceberg_namespace, table, outcome: detected, repaired, incomplete, failed, backed_off_watchdog, backed_off_failed, backed_off_interrupted, suppressed, marker_failed |
siglake_group_count_delta_write_failures_total |
iceberg_namespace, table |
siglake_side_aggregate_publish_failures_total |
iceberg_namespace, table |
siglake_group_count_auto_rebuilds_total |
iceberg_namespace, table, outcome: success, incomplete, failed |
siglake_group_count_delta_write_retries_total |
iceberg_namespace, table |
The namespace label is iceberg_namespace, not namespace, because Prometheus
attaches the Kubernetes namespace under namespace and renames a colliding
metric label to exported_namespace. siglake_inline_coverage_unproven carries
it for the same reason. All four alerts on these counters report
<iceberg_namespace>.<table>. Three of them - SiglakeGroupCountDeltaLost,
SiglakeGroupCountAggregateShort and SiglakeSideAggregatePublicationLost -
render siglake rebuild-group-counts --namespace <ns> --table <table> with both
labels filled in. The fourth, SiglakeGroupCountDeltaRetrying, has no repair
command to render: a retry is a transient object-store or credential failure, so
it names the pair and the pod label Prometheus adds, and sends you to that
pod's warehouse access.
All five counters gained iceberg_namespace in the 0.2.0 line, the retry
counter last; before that they carried table alone, and one table name used
in several namespaces shared a series. A dashboard or alert that selects on
table alone still matches, because a PromQL matcher ignores the labels it
does not name, but it now returns one series per namespace where it used to
return one in total. Series identity changes with the label: the pre-upgrade
series and the per-namespace ones that replace it have separate histories, so
rate() over a window that spans the upgrade reports each of them that still
holds two samples in the window. A recording rule, a join or a by clause that
pairs these counters with another series on an exact label set needs
iceberg_namespace added to it; one that aggregates across namespaces on
purpose keeps the label out.
Inline aggregate repair¶
Two maintenance paths rewrite a table's inline aggregate object outside the
commit path. The compactor re-roots the coverage edge off ancestry that a
snapshot expiry is about to drop. You run
siglake rebuild-time-aggregates --namespace <ns> --table <table> to recompute
the time aggregates from committed files. Both paths are fenced on the object
they read, so each counts the writes it published apart from the writes it
abandoned. A third path only reads: every 15 minutes
(SIGLAKE_INLINE_COVERAGE_SCAN_INTERVAL_SECS) the maintenance compactor
censuses each maintained table for an object whose coverage edge no longer
reaches the current snapshot, and reports the tables that need the rebuild
command. The census is the only signal here with an alert,
SiglakeInlineCoverageUnproven.
| Metric | Watch for |
|---|---|
siglake_inline_coverage_reroots_total |
One published re-root: an expiry moved the coverage edge onto ancestry that survives it, so windowed GROUP BY on the table keeps its fast path. No labels. |
siglake_inline_coverage_reroot_conflicts_total |
One re-root the same pass declined to write, because a publication landed under it and the object no longer carried the edge the expiry proved. Nothing is lost: that newer publication's own edge is current. No labels. A re-root that fails outright increments neither counter and reports itself in the compactor warning log. |
siglake_inline_time_aggregate_rebuilds_total |
One published rebuild, once per table. A run that finds the object already covered, or that can prove no component, publishes nothing and does not increment. No labels. |
siglake_inline_time_aggregate_rebuild_conflicts_total |
One rebuild attempt restarted because the table committed under it. The command retries three times, then fails without writing and asks for a window with no ingest to that table. No labels. |
siglake_inline_time_rebuild_files_total{source} |
Files one time-bucket rebuild read. source="footer" is a file whose minute histogram came out of the Parquet footer; source="decode" is one whose footer was absent or failed its validity guard, so the pass scanned the timestamp column instead. |
siglake_inline_time_group_rebuild_files_total{source} |
The same split for the per-column time group counts a rebuild recomputes. |
siglake_inline_time_rebuild_seconds{component} |
How long one rebuild attempt spent on one component: component="time_buckets" is the hourly time-bucket pass, component="time_group_counts" the per-column group-count pass. Not the command's duration, which covers both components and up to three attempts. A _seconds name, so it has _bucket series (Histogram export). |
siglake_inline_time_rebuild_decoded_bytes{component} |
What the same attempt decoded for that component, for the files whose footer was unusable: the projected Arrow batches' in-memory size, not bytes fetched from the object store. A reading near zero on a table with many files means nearly every file answered from its footer. Exported as a summary, with a quantile label, _sum and _count and no _bucket series, so histogram_quantile() does not apply to it and its quantiles do not average across processes. |
siglake_inline_coverage_unproven{iceberg_namespace,table} |
Current state, 1 or 0, set by the census for every table it reaches a verdict on. 1 means the object's coverage edge does not reach the current snapshot and no publication is in flight, so windowed GROUP BY, date histograms and windowed counts on that table run on the per-file tiers until siglake rebuild-time-aggregates repairs it; answers stay exact throughout. A repaired table reads 0 at the next pass. The namespace label is iceberg_namespace so it does not collide with the namespace label Prometheus attaches. An object the census cannot read leaves the previous reading standing, because a failed GET is not evidence either way; a table the census stops reaching at all, a dropped index, is zeroed. SiglakeInlineCoverageUnproven fires on it. |
siglake_inline_coverage_census_total |
One increment per completed census pass. No labels. The gauge beside it is a last observation, so the alert reads this counter as its liveness arm: a compactor that stopped censusing leaves the alert instead of paging from a reading nobody is refreshing. |
siglake_inline_coverage_census_requests_total{op} |
Object-store calls from the census. A table with an aggregate object uses the shipped two head calls and one get. Attempts count at the call site, including a failed GET. |
siglake_inline_coverage_census_bytes_total{iceberg_namespace,table} |
Bytes from usable GET responses. Failed or unparseable responses add no bytes. |
siglake_inline_coverage_census_pass_duration_seconds |
Duration of each completed census pass. The histogram exists at zero before the maintenance loop starts. |
siglake_inline_coverage_census_tables |
Number of tables seen by the last completed pass. The gauge starts at 0. |
The two source splits say where a rebuild's work went: a decode file reads
column data where a footer file reads metadata, so a decode-heavy run is the
slow one. The two cost families put a number on that per component. All four
rebuild series record as soon as the pass reads, so an attempt the table
committed under records the seconds and bytes it spent before its work was
thrown away.
This output is scheduled for
Siglake 0.3.0
and is not present in 0.2.x. The command's final stdout line starts with
rebuild_cost and ends with one JSON object.
rebuild_cost {"measurement":"complete","publications":1,"conflicts":0,"components":{"time_buckets":{"seconds":0.021,"decoded_bytes":27136,"files":{"footer":0,"decode":16}},"time_group_counts":{"seconds":0.004,"decoded_bytes":196864,"files":{"footer":0,"decode":16}}}}
publications counts writes completed by this invocation. conflicts counts
attempts discarded because the table committed before publication. The command
makes up to three attempts. Every total accumulates across those attempts,
including work later discarded by a conflict.
Each component reports the sum of its attempt durations in seconds.
decoded_bytes is the projected Arrow batches' in-memory size, not bytes read
from object storage. files.footer counts files answered from Parquet footer
metadata. files.decode counts files read from column data because the footer
was missing or invalid.
measurement="complete" means the command finished measuring, not that it
published a write. An already-covered table is a successful no-op with zero
publications, conflicts, seconds, bytes and files. A command error keeps the
observations collected before the error, prints measurement="incomplete",
then retains its nonzero exit status.
The command installs a recorder for its own invocation and writes the captured
totals to stdout. It starts no metrics listener or network exporter, and these
observations do not appear on any role's /metrics endpoint. No flag changes
that boundary.
The re-root pair and the two census series belong to the maintenance compactor, which does export. Read those as the standing signal for this section; the census never rebuilds anything itself.
Table subscriptions¶
| Metric | Watch for |
|---|---|
siglake_subscription_history_gap_total{table} |
A siglake subscribe poll refused because snapshot expiry removed an ancestor between the consumer cursor and current snapshot. The subscription cannot advance until you re-bootstrap or backfill it. The chart has no alert for this metric. See Tailing a table with siglake subscribe. |
siglake_subscription_rewrite_commits_skipped_total{table,origin} |
A skipped non-append commit. origin="siglake" covers compaction, retention and delete tasks, which do not add rows. origin="foreign" means another writer may have added rows that the subscription did not deliver. The chart has no alert for this metric. If another engine writes the table, alert on increase(siglake_subscription_rewrite_commits_skipped_total{origin="foreign"}[10m]) > 0 and recover with a query backfill. See External overwrite semantics are not supported. |
v1 index rebuild cost¶
Three metrics report the post-rewrite v1 inverted-index rebuild, the opt-in
pass that reads a committed data file back whole and indexes the columns a
rewrite left without one (SIGLAKE_INDEX_REBUILD=1 or
compactor.indexRebuild; see
Indexing and promotion and
Post-compaction index rebuilding).
All three carry tenant and table, known only at the increment, so a table
nothing has rebuilt is absent rather than 0. Increments land after the
registration commits: a pass that rebuilt nothing, or whose registration was
deferred, records nothing here.
| Metric | Watch for |
|---|---|
siglake_index_rebuild_files_total{tenant,table} |
Data files one pass rebuilt, charged once per pass. Counts that rise while the seg2 write counter stays flat are whole-file decodes a seg2 sidecar would have avoided. |
siglake_index_rebuild_bytes_total{tenant,table} |
Those files' own size: the source Parquet the pass re-read and decoded, not the index it produced. |
siglake_index_rebuild_seconds{tenant,table} |
Pass duration: the footer probe over every candidate file, the whole-file decodes and the Puffin registration. A _seconds name, so it has _bucket series (Histogram export). |
The pass holds the compaction cycle it runs in, so a p99 near that interval is a compactor whose cycle is the rebuild.
Suggested alerts¶
The chart provides 39 alert rules. Monitoring lists each expression and operator action.
Histogram export¶
The histogram export format is fixed in code. Siglake metrics ending in
*_seconds and five named count families use the Prometheus histogram format
with bucket series. Other histogram metrics use the Prometheus summary format.
You do not configure buckets at scrape time. Scrape the buckets as emitted and
use histogram_quantile() only with metrics that expose _bucket series.
Bucketed metric families¶
The inventory below uses histogram for every metric recorded through the
histogram API, but not every such metric is exported as a Prometheus histogram.
Only names ending in *_seconds and these five count histograms have
bucket series:
siglake_group_count_deltas_foldedsiglake_group_count_tier2_files_per_callsiglake_compactor_mirror_sync_objectssiglake_compactor_mirror_sync_rotation_objectssiglake_query_scan_file_cache_populate_rows
The first four share one per-call count layout.
siglake_query_scan_file_cache_populate_rows has its own row-count layout,
which places an edge at 131,071 next to one at 131,072 so that the fraction at
or above the write path's row-group floor is exact.
Other histogram families, including _bytes, _rows, ratios, and other
counts, are exported as summaries with a quantile label and no _bucket
series. There is no per-metric switch that changes this export type. Use
histogram_quantile() only with the bucketed families above; if a
future cross-pod quantile needs another family, that metric first needs an
explicit bucket layout in Siglake. The monitoring examples follow this rule:
their histogram_quantile() expressions use _seconds_bucket metrics. This
export split comes from Siglake's September 3, 2026 Prometheus histogram-bucket
correction.
Full inventory¶
This section is generated
Names, types, and emitters are extracted from the public Siglake source
tree: workspace crates by name, the vendored Iceberg forks by their
third_party/<fork> path. Regenerate with
./scripts/gen-reference.sh /path/to/siglake.
Generated from Siglake commit 8a517387 on 2026-10-02.
A row means the source tree registers the metric, not that your binary emits
it. Three rows are behind a cargo feature nothing in the workspace turns on,
experimental-exact-point-rollup in siglake-storage:
siglake_exact_point_rollup_build_nanoseconds_total,
siglake_exact_point_rollup_query_total and
siglake_exact_point_boundary_data_files_read_total. A default build never
exports those three series.
| Metric | Type | Emitted by |
|---|---|---|
siglake_agg_result_cache_hits_total |
counter | siglake-storage |
siglake_agg_result_cache_misses_total |
counter | siglake-storage |
siglake_auto_promotion_candidates |
gauge | siglake-compactor |
siglake_auto_promotion_candidates_available |
gauge | siglake-compactor |
siglake_auto_promotion_columns |
gauge | siglake-compactor |
siglake_auto_promotion_enabled |
gauge | siglake-compactor |
siglake_auto_promotion_pass_duration_seconds |
histogram | siglake-storage |
siglake_auto_promotion_passes_total |
counter | siglake-compactor |
siglake_auto_promotion_sample_bytes_total |
counter | siglake-compactor, siglake-storage |
siglake_auto_promotion_sample_reads_total |
counter | siglake-compactor, siglake-storage |
siglake_auto_promotion_sampled_keys_dropped_total |
counter | siglake-storage |
siglake_build_info |
gauge | siglake-cli, siglake-operator, siglake-query-server |
siglake_cache_admission_refused_total |
counter | siglake-storage |
siglake_cache_budget_bytes |
gauge | siglake-storage |
siglake_cache_bytes |
gauge | siglake-storage |
siglake_cache_entries |
gauge | siglake-storage |
siglake_cache_evictions_total |
counter | siglake-storage |
siglake_catalog_cas_total |
counter | third_party/iceberg-catalog-sql |
siglake_catalog_claim_ineligible_total |
counter | siglake-storage |
siglake_catalog_claim_quarantined |
gauge | siglake-compactor |
siglake_catalog_claim_quarantined_total |
counter | siglake-storage |
siglake_catalog_claim_released_foreign_shard_total |
counter | siglake-storage |
siglake_catalog_claims_reclaimed_already_committed_total |
counter | siglake-storage |
siglake_catalog_claims_reclaimed_total |
counter | siglake-storage |
siglake_catalog_rows_purged_total |
counter | siglake-storage |
siglake_catalog_table_lease_total |
counter | siglake-storage |
siglake_compactor_agg_fold_duration_seconds |
histogram | siglake-compactor |
siglake_compactor_auto_promotions_total |
counter | siglake-compactor |
siglake_compactor_bin_concurrency |
gauge | siglake-storage |
siglake_compactor_bin_duration_seconds |
histogram | siglake-storage |
siglake_compactor_bin_files_added_total |
counter | siglake-storage |
siglake_compactor_bin_files_removed_total |
counter | siglake-storage |
siglake_compactor_bin_rows_total |
counter | siglake-storage |
siglake_compactor_bins_committed_total |
counter | siglake-storage |
siglake_compactor_carrier_map_duration_seconds |
histogram | siglake-compactor |
siglake_compactor_claim_duration_seconds |
histogram | siglake-compactor |
siglake_compactor_claimed_bytes |
histogram | siglake-compactor |
siglake_compactor_claimed_last_cycle |
gauge | siglake-compactor |
siglake_compactor_claimed_segments |
histogram | siglake-compactor |
siglake_compactor_commit_deferred_total |
counter | siglake-compactor |
siglake_compactor_commit_duration_seconds |
histogram | siglake-compactor |
siglake_compactor_concat_duration_seconds |
histogram | siglake-compactor |
siglake_compactor_cycle_duration_seconds |
histogram | siglake-compactor |
siglake_compactor_cycles_total |
counter | siglake-compactor |
siglake_compactor_delete_task_files_rewritten_total |
counter | siglake-compactor |
siglake_compactor_delete_task_rows_deleted_total |
counter | siglake-compactor |
siglake_compactor_delete_tasks_completed_total |
counter | siglake-compactor |
siglake_compactor_delete_tasks_duration_seconds |
histogram | siglake-compactor |
siglake_compactor_delete_tasks_nonterminal |
gauge | siglake-compactor |
siglake_compactor_delete_tasks_stalled_total |
counter | siglake-compactor |
siglake_compactor_depth_trigger_bins_total |
counter | siglake-storage |
siglake_compactor_drain_inflight |
gauge | siglake-compactor |
siglake_compactor_drain_schema_aligned_total |
counter | siglake-compactor |
siglake_compactor_effective_catalog_claim_batch |
gauge | siglake-compactor |
siglake_compactor_effective_fs_only_settings_present |
gauge | siglake-compactor |
siglake_compactor_effective_mirror_sync |
gauge | siglake-compactor |
siglake_compactor_effective_reclustering |
gauge | siglake-compactor |
siglake_compactor_effective_role |
gauge | siglake-compactor |
siglake_compactor_expire_attempts_total |
counter | siglake-compactor |
siglake_compactor_expire_duration_seconds |
histogram | siglake-compactor |
siglake_compactor_expire_passes_total |
counter | siglake-compactor |
siglake_compactor_finish_failures_total |
counter | siglake-compactor |
siglake_compactor_gauge_sample_duration_seconds |
histogram | siglake-compactor |
siglake_compactor_iceberg_append_duration_seconds |
histogram | siglake-compactor |
siglake_compactor_idle_sleep_duration_seconds |
histogram | siglake-compactor |
siglake_compactor_index_resolve_duration_seconds |
histogram | siglake-compactor |
siglake_compactor_index_unresolved_total |
counter | siglake-compactor |
siglake_compactor_level_compactions_total |
counter | siglake-storage |
siglake_compactor_loop_iteration_duration_seconds |
histogram | siglake-compactor |
siglake_compactor_maintenance_lease_denied_total |
counter | siglake-compactor |
siglake_compactor_maintenance_lease_errors_total |
counter | siglake-compactor |
siglake_compactor_maintenance_skipped_backpressure_total |
counter | siglake-compactor |
siglake_compactor_maintenance_throttled_backpressure_total |
counter | siglake-compactor |
siglake_compactor_mark_committed_duration_seconds |
histogram | siglake-compactor |
siglake_compactor_merge_fetch_gross_nanos_total |
counter | siglake-storage |
siglake_compactor_merge_fetch_stalled_nanos_total |
counter | siglake-storage |
siglake_compactor_merge_page_bounded_total |
counter | siglake-storage |
siglake_compactor_merge_path_total |
counter | siglake-storage |
siglake_compactor_merge_plan_runs_total |
counter | siglake-storage |
siglake_compactor_merge_rowgroups_copyable_total |
counter | siglake-storage |
siglake_compactor_merge_rowgroups_total |
counter | siglake-storage |
siglake_compactor_merge_rows_copyable_total |
counter | siglake-storage |
siglake_compactor_merge_stage_nanos_total |
counter | siglake-storage |
siglake_compactor_mirror_mark_errors_total |
counter | siglake-compactor |
siglake_compactor_mirror_marked_total |
counter | siglake-compactor |
siglake_compactor_mirror_sync_duration_seconds |
histogram | siglake-compactor |
siglake_compactor_mirror_sync_last_completed_timestamp_seconds |
gauge | siglake-compactor |
siglake_compactor_mirror_sync_objects |
histogram | siglake-compactor |
siglake_compactor_mirror_sync_registered_total |
counter | siglake-compactor |
siglake_compactor_mirror_sync_rotation_duration_seconds |
histogram | siglake-compactor, siglake-core |
siglake_compactor_mirror_sync_rotation_objects |
histogram | siglake-compactor, siglake-core |
siglake_compactor_mirror_sync_rotation_objects_examined |
gauge | siglake-compactor |
siglake_compactor_mirror_sync_rotations_completed |
gauge | siglake-compactor |
siglake_compactor_mirror_sync_total |
counter | siglake-compactor |
siglake_compactor_mirror_unreclaimed_total |
counter | siglake-compactor |
siglake_compactor_orphans_disposed_total |
counter | siglake-compactor |
siglake_compactor_orphans_held |
gauge | siglake-compactor |
siglake_compactor_orphans_quarantined_total |
counter | siglake-compactor |
siglake_compactor_pass_claim_attempts_exhausted_total |
counter | siglake-compactor |
siglake_compactor_peek_duration_seconds |
histogram | siglake-compactor |
siglake_compactor_processing_pending |
gauge | siglake-compactor |
siglake_compactor_promotion_backfill_bins_total |
counter | siglake-storage |
siglake_compactor_promotion_backfill_bytes_in_total |
counter | siglake-storage |
siglake_compactor_promotion_backfill_bytes_out_total |
counter | siglake-storage |
siglake_compactor_promotion_backfill_duration_seconds |
histogram | siglake-storage |
siglake_compactor_promotion_backfill_files_total |
counter | siglake-storage |
siglake_compactor_proof_watermark_skipped_total |
counter | siglake-compactor |
siglake_compactor_reclaim_duration_seconds |
histogram | siglake-compactor |
siglake_compactor_reclaim_unprovable_total |
counter | siglake-compactor |
siglake_compactor_recluster_files_added_total |
counter | siglake-compactor |
siglake_compactor_recluster_files_net_removed_total |
counter | siglake-compactor |
siglake_compactor_recluster_files_removed_total |
counter | siglake-compactor |
siglake_compactor_recluster_merge_tiers_total |
counter | siglake-storage |
siglake_compactor_recluster_pass_duration_seconds |
histogram | siglake-compactor |
siglake_compactor_recluster_pass_net_files |
histogram | siglake-compactor |
siglake_compactor_recluster_passes_noop_total |
counter | siglake-compactor |
siglake_compactor_recluster_passes_total |
counter | siglake-compactor |
siglake_compactor_recluster_preempted_total |
counter | siglake-storage |
siglake_compactor_retention_duration_seconds |
histogram | siglake-compactor |
siglake_compactor_retention_object_delete_errors_total |
counter | siglake-compactor |
siglake_compactor_retention_purged_total |
counter | siglake-compactor |
siglake_compactor_rewrite_bytes_in_total |
counter | siglake-storage |
siglake_compactor_rewrite_bytes_out_total |
counter | siglake-storage |
siglake_compactor_rewrite_commit_duration_seconds |
histogram | siglake-storage |
siglake_compactor_rewrite_rows_total |
counter | siglake-storage |
siglake_compactor_rows_committed_total |
counter | siglake-compactor |
siglake_compactor_sealed_pending |
gauge | siglake-compactor |
siglake_compactor_sealed_pending_bytes |
gauge | siglake-compactor |
siglake_compactor_sealed_pending_oldest_age_seconds |
gauge | siglake-compactor |
siglake_compactor_segment_fetch_duration_seconds |
histogram | siglake-compactor |
siglake_compactor_segment_read_duration_seconds |
histogram | siglake-compactor |
siglake_compactor_segments_committed_total |
counter | siglake-compactor |
siglake_compactor_segments_fetched_total |
counter | siglake-compactor |
siglake_compactor_segments_poisoned |
gauge | siglake-compactor |
siglake_compactor_segments_poisoned_total |
counter | siglake-compactor |
siglake_compactor_segments_swept_total |
counter | siglake-compactor |
siglake_compactor_snapshots_expired_total |
counter | siglake-compactor |
siglake_compactor_strict_residual_rows_total |
counter | siglake-compactor |
siglake_compactor_wal_owner_mismatch_total |
counter | siglake-compactor |
siglake_compactor_wal_stale_segments_total |
counter | siglake-compactor |
siglake_compactor_watchdog_trips_total |
counter | siglake-compactor |
siglake_consumed_proof_bytes |
gauge | siglake-storage |
siglake_consumed_proof_cap_refusals_total |
counter | siglake-storage |
siglake_consumed_proof_current_watermark_lag_seconds |
gauge | siglake-storage |
siglake_consumed_proof_entries |
gauge | siglake-storage |
siglake_consumed_proof_watermark_lag_seconds |
gauge | siglake-storage |
siglake_delete_task_rewrite_path_total |
counter | siglake-storage |
siglake_dropped_index_aggregate_cleanup_failures_total |
counter | siglake-storage |
siglake_dropped_index_aggregate_objects_deleted_total |
counter | siglake-storage |
siglake_dropped_index_cleanup_records_total |
counter | siglake-storage |
siglake_events_accepted_total |
counter | siglake-ingest |
siglake_exact_point_boundary_data_files_read_total |
counter | siglake-storage |
siglake_exact_point_rollup_build_nanoseconds_total |
counter | siglake-storage |
siglake_exact_point_rollup_query_total |
counter | siglake-storage |
siglake_file_list_cache_hits_total |
counter | siglake-storage |
siglake_file_list_cache_misses_total |
counter | siglake-storage |
siglake_file_list_cache_phase_seconds |
histogram | siglake-storage |
siglake_footer_cache_hits_total |
counter | siglake-storage |
siglake_footer_cache_misses_total |
counter | siglake-storage |
siglake_gc_bytes_reclaimed_total |
counter | siglake-storage |
siglake_gc_orphans_deleted_total |
counter | siglake-storage |
siglake_gc_orphans_found_total |
counter | siglake-storage |
siglake_group_count_aggregate_rebuilds_total |
counter | siglake-storage |
siglake_group_count_auto_rebuilds_total |
counter | siglake-storage |
siglake_group_count_column_demoted_total |
counter | siglake-storage |
siglake_group_count_delta_fold_retries_total |
counter | siglake-storage |
siglake_group_count_delta_write_failures_total |
counter | siglake-storage |
siglake_group_count_delta_write_retries_total |
counter | siglake-storage |
siglake_group_count_deltas_absorbed_total |
counter | siglake-compactor |
siglake_group_count_deltas_deleted_total |
counter | siglake-compactor |
siglake_group_count_deltas_folded |
histogram | siglake-storage |
siglake_group_count_short_aggregates_total |
counter | siglake-storage |
siglake_group_count_sketched_columns |
gauge | siglake-storage |
siglake_group_count_tier2_calls_total |
counter | siglake-storage |
siglake_group_count_tier2_files_per_call |
histogram | siglake-storage |
siglake_group_count_tier2_files_total |
counter | siglake-storage |
siglake_iceberg_aggregate_join_duration_seconds |
histogram | siglake-storage |
siglake_iceberg_commit_attempts_total |
counter | third_party/iceberg |
siglake_iceberg_commit_duration_seconds |
histogram | siglake-storage |
siglake_iceberg_commit_stale_base_total |
counter | third_party/iceberg |
siglake_iceberg_data_files_per_append |
histogram | siglake-storage |
siglake_iceberg_data_flush_duration_seconds |
histogram | siglake-storage |
siglake_iceberg_data_write_duration_seconds |
histogram | siglake-storage |
siglake_iceberg_footer_cache_total |
counter | third_party/iceberg |
siglake_iceberg_inverted_index_used_total |
counter | third_party/iceberg |
siglake_iceberg_load_table_duration_seconds |
histogram | siglake-storage |
siglake_iceberg_object_store_bytes_read_total |
counter | third_party/iceberg |
siglake_iceberg_parquet_encode_duration_seconds |
histogram | siglake-storage |
siglake_iceberg_parsed_index_cache_bytes |
gauge | third_party/iceberg |
siglake_iceberg_parsed_index_cache_evictions_total |
counter | third_party/iceberg |
siglake_iceberg_parsed_index_cache_lookups_total |
counter | third_party/iceberg |
siglake_iceberg_parsed_index_cache_max_bytes |
gauge | third_party/iceberg |
siglake_iceberg_puffin_blob_cache_bytes |
gauge | third_party/iceberg |
siglake_iceberg_puffin_blob_cache_evictions_total |
counter | third_party/iceberg |
siglake_iceberg_puffin_blob_cache_lookups_total |
counter | third_party/iceberg |
siglake_iceberg_puffin_blob_cache_max_bytes |
gauge | third_party/iceberg |
siglake_iceberg_puffin_blob_fetches_total |
counter | third_party/iceberg |
siglake_iceberg_raw_bloom_skip_total |
counter | third_party/iceberg |
siglake_iceberg_raw_rowgroup_bloom_skip_total |
counter | third_party/iceberg |
siglake_iceberg_scan_cost_phase_seconds |
histogram | siglake-storage |
siglake_iceberg_scan_cost_requests_total |
counter | siglake-storage |
siglake_iceberg_scan_cost_total_seconds |
histogram | siglake-storage |
siglake_iceberg_schema_columns_added_total |
counter | siglake-storage |
siglake_iceberg_segmented_index_declined_total |
counter | third_party/iceberg |
siglake_iceberg_segmented_index_directory_cache_bytes |
gauge | third_party/iceberg |
siglake_iceberg_segmented_index_directory_cache_evictions_total |
counter | third_party/iceberg |
siglake_iceberg_segmented_index_directory_cache_lookups_total |
counter | third_party/iceberg |
siglake_iceberg_segmented_index_directory_cache_max_bytes |
gauge | third_party/iceberg |
siglake_iceberg_segmented_index_fetched_bytes |
histogram | third_party/iceberg |
siglake_iceberg_segmented_index_group_index_bytes |
histogram | third_party/iceberg |
siglake_iceberg_segmented_index_range_reads |
histogram | third_party/iceberg |
siglake_iceberg_segmented_index_reader_reads |
histogram | third_party/iceberg |
siglake_iceberg_segmented_index_resident_bytes |
histogram | third_party/iceberg |
siglake_iceberg_segmented_index_selected_rows |
histogram | third_party/iceberg |
siglake_iceberg_segmented_index_stages |
histogram | third_party/iceberg |
siglake_iceberg_segmented_index_used_total |
counter | third_party/iceberg |
siglake_iceberg_segmented_index_writes_total |
counter | third_party/iceberg |
siglake_iceberg_segmented_index_written_bytes |
histogram | third_party/iceberg |
siglake_iceberg_snapshots_expired_total |
counter | siglake-storage |
siglake_iceberg_statistics_removed_total |
counter | siglake-storage |
siglake_iceberg_statistics_retirement_skipped_total |
counter | siglake-storage |
siglake_iceberg_table_cache_fenced_total |
counter | siglake-storage |
siglake_iceberg_table_cache_requests_total |
counter | siglake-storage |
siglake_iceberg_text_index_startup_seconds |
histogram | third_party/iceberg |
siglake_index_build_bytes |
histogram | siglake-storage |
siglake_index_build_seconds |
histogram | siglake-storage |
siglake_index_footer_checksum_refused_total |
counter | third_party/iceberg |
siglake_index_rebuild_bytes_total |
counter | siglake-storage |
siglake_index_rebuild_files_total |
counter | siglake-storage |
siglake_index_rebuild_seconds |
histogram | siglake-storage |
siglake_index_registration_deferred_total |
counter | siglake-storage |
siglake_index_row_domain_mismatch_total |
counter | third_party/iceberg |
siglake_index_stamp_mismatch_total |
counter | third_party/iceberg |
siglake_ingest_backpressure_append_duration_seconds |
histogram | siglake-ingest |
siglake_ingest_backpressure_batch_commands |
histogram | siglake-ingest |
siglake_ingest_backpressure_batch_events |
histogram | siglake-ingest |
siglake_ingest_backpressure_batches_total |
counter | siglake-ingest |
siglake_ingest_backpressure_lanes |
gauge | siglake-ingest |
siglake_ingest_backpressure_queue_commands |
gauge | siglake-ingest |
siglake_ingest_backpressure_queue_events |
gauge | siglake-ingest |
siglake_ingest_backpressure_queue_wait_seconds |
histogram | siglake-ingest |
siglake_ingest_backpressure_rejected_total |
counter | siglake-ingest |
siglake_ingest_commit_force_unsupported_total |
counter | siglake-ingest |
siglake_ingest_commit_force_wait_seconds |
histogram | siglake-ingest |
siglake_ingest_commit_total |
counter | siglake-ingest |
siglake_ingest_lane_refused_total |
counter | siglake-ingest |
siglake_ingest_mem_breaker_open |
gauge | siglake-ingest |
siglake_ingest_mem_breaker_rejected_total |
counter | siglake-ingest |
siglake_ingest_metric_label_over_cap_total |
counter | siglake-ingest |
siglake_ingest_not_ready_total |
counter | siglake-ingest |
siglake_ingest_rate_limit_rejected_total |
counter | siglake-ingest |
siglake_ingest_request_duration_seconds |
histogram | siglake-ingest |
siglake_ingest_request_events |
histogram | siglake-ingest |
siglake_ingest_requests_in_flight |
gauge | siglake-ingest |
siglake_ingest_requests_total |
counter | siglake-ingest |
siglake_ingest_rss_bytes |
gauge | siglake-ingest |
siglake_ingest_stream_lagged_total |
counter | siglake-ingest |
siglake_ingest_stream_subscribers_opened_total |
counter | siglake-ingest |
siglake_ingest_tenant_denied_total |
counter | siglake-ingest |
siglake_inline_coverage_census_bytes_total |
counter | siglake-storage |
siglake_inline_coverage_census_pass_duration_seconds |
histogram | siglake-compactor |
siglake_inline_coverage_census_requests_total |
counter | siglake-compactor, siglake-storage |
siglake_inline_coverage_census_tables |
gauge | siglake-compactor |
siglake_inline_coverage_census_total |
counter | siglake-compactor |
siglake_inline_coverage_reroot_conflicts_total |
counter | siglake-storage |
siglake_inline_coverage_reroots_total |
counter | siglake-storage |
siglake_inline_coverage_unproven |
gauge | siglake-storage |
siglake_inline_time_aggregate_rebuild_conflicts_total |
counter | siglake-cli, siglake-storage |
siglake_inline_time_aggregate_rebuilds_total |
counter | siglake-storage |
siglake_inline_time_group_rebuild_files_total |
counter | siglake-cli, siglake-storage |
siglake_inline_time_rebuild_decoded_bytes |
histogram | siglake-cli, siglake-storage |
siglake_inline_time_rebuild_files_total |
counter | siglake-cli, siglake-storage |
siglake_inline_time_rebuild_seconds |
histogram | siglake-cli, siglake-storage |
siglake_object_cache_bytes |
gauge | third_party/iceberg |
siglake_object_cache_requests_total |
counter | third_party/iceberg |
siglake_object_store_debounced_total |
counter | third_party/iceberg |
siglake_object_store_read_bytes_total |
counter | third_party/iceberg |
siglake_object_store_reads_total |
counter | third_party/iceberg |
siglake_object_store_write_chunk_bytes |
gauge | third_party/iceberg-storage-opendal |
siglake_object_store_write_concurrency |
gauge | third_party/iceberg-storage-opendal |
siglake_object_store_writer_opened_total |
counter | third_party/iceberg-storage-opendal |
siglake_operator_managed_replicas |
gauge | siglake-operator |
siglake_operator_prom_query_errors_total |
counter | siglake-operator |
siglake_operator_reconcile_duration_seconds |
histogram | siglake-operator |
siglake_operator_reconcile_errors_total |
counter | siglake-operator |
siglake_operator_reconciles_total |
counter | siglake-operator |
siglake_operator_rollout_held_total |
counter | siglake-operator |
siglake_query_admission_reserved_bytes |
gauge | siglake-query-server |
siglake_query_admission_waiting |
gauge | siglake-query-server |
siglake_query_approximate_group_counts_total |
counter | siglake-query-server |
siglake_query_audit_dropped_total |
counter | siglake-query-server |
siglake_query_audit_failures_total |
counter | siglake-query-server |
siglake_query_audit_retained_bytes |
gauge | siglake-query-server |
siglake_query_audit_retained_rows |
gauge | siglake-query-server |
siglake_query_audit_rows_total |
counter | siglake-query-server |
siglake_query_breaker_trips_total |
counter | siglake-query-server |
siglake_query_buffer_delta_cache_total |
counter | siglake-query-server |
siglake_query_coordinator_failover_skipped_total |
counter | siglake-query-server |
siglake_query_coordinator_failover_total |
counter | siglake-query-server |
siglake_query_coordinator_total |
counter | siglake-query-server |
siglake_query_cost_bytes_estimated |
histogram | siglake-query-server |
siglake_query_count_distinct_fast_path_total |
counter | siglake-query-server |
siglake_query_default_order_applied_total |
counter | siglake-query-server |
siglake_query_exec_pool_abandoned_total |
counter | siglake-query-server |
siglake_query_exec_pool_in_flight |
gauge | siglake-query-server |
siglake_query_exec_pool_queue_seconds |
histogram | siglake-query-server |
siglake_query_exec_route_total |
counter | siglake-query-server |
siglake_query_fast_path_buffer_delta_total |
counter | siglake-query-server |
siglake_query_group_count_sketch_total |
counter | siglake-storage |
siglake_query_histogram_snapshot_agg_total |
counter | siglake-storage |
siglake_query_hot_cache_distinct_dropped_total |
counter | siglake-query-server |
siglake_query_hot_cache_segment_errors_total |
counter | siglake-query-server |
siglake_query_hot_cache_series |
gauge | siglake-query-server |
siglake_query_hot_cache_series_dropped_total |
counter | siglake-query-server |
siglake_query_in_flight |
gauge | siglake-query-server |
siglake_query_inverted_index_declined_total |
counter | siglake-storage |
siglake_query_job_access_denied_total |
counter | siglake-query-server |
siglake_query_job_terminal_conflict_total |
counter | siglake-query-server |
siglake_query_job_terminal_write_failed_total |
counter | siglake-query-server |
siglake_query_jobs_cancel_propagated_total |
counter | siglake-query-server |
siglake_query_jobs_reconciled_total |
counter | siglake-query-server |
siglake_query_jobs_total |
counter | siglake-query-server |
siglake_query_jobs_unreconciled |
gauge | siglake-query-server |
siglake_query_jobs_unreconciled_dropped_total |
counter | siglake-query-server |
siglake_query_memory_pool_available_bytes |
gauge | siglake-storage |
siglake_query_memory_pool_bytes |
gauge | siglake-storage |
siglake_query_memory_pool_idle_residual_reports_total |
counter | siglake-query-server |
siglake_query_memory_pool_reserved_bytes |
gauge | siglake-storage |
siglake_query_memory_pool_used_ratio |
gauge | siglake-storage |
siglake_query_memory_reserved_elsewhere_bytes |
gauge | siglake-storage |
siglake_query_ordered_drain_backpressure_total |
counter | siglake-storage, third_party/iceberg |
siglake_query_ordered_drain_buffered_bytes |
gauge | siglake-storage, third_party/iceberg |
siglake_query_ordered_plan_cache_total |
counter | siglake-storage |
siglake_query_ordered_residual_fallback_total |
counter | siglake-query-server |
siglake_query_ordered_residual_hint_total |
counter | siglake-query-server |
siglake_query_peer_discovery_last_success_seconds |
gauge | siglake-query-server |
siglake_query_peer_discovery_members |
gauge | siglake-query-server |
siglake_query_peer_discovery_refresh_total |
counter | siglake-query-server |
siglake_query_phase_duration_seconds |
histogram | siglake-query-server |
siglake_query_preflight_bytes_skipped_limited_total |
counter | siglake-query-server |
siglake_query_prewarm_files_skipped_total |
counter | siglake-storage |
siglake_query_prewarm_total |
counter | siglake-storage |
siglake_query_registered_tables |
histogram | siglake-query-server |
siglake_query_request_duration_seconds |
histogram | siglake-query-server |
siglake_query_requests_total |
counter | siglake-query-server |
siglake_query_runtime_df_bytes_scanned |
histogram | siglake-query-server |
siglake_query_runtime_df_elapsed_compute_seconds |
histogram | siglake-query-server |
siglake_query_runtime_leaf_rows_scanned |
histogram | siglake-query-server |
siglake_query_scan_attribution_incomplete_total |
counter | siglake-query-server |
siglake_query_scan_clipped_admission_total |
counter | siglake-storage |
siglake_query_scan_clipped_admission_wait_seconds |
histogram | siglake-storage |
siglake_query_scan_decode_reservation_total |
counter | siglake-storage |
siglake_query_scan_file_cache_accounted_bytes |
gauge | siglake-storage |
siglake_query_scan_file_cache_bytes |
gauge | siglake-storage |
siglake_query_scan_file_cache_entries |
gauge | siglake-storage |
siglake_query_scan_file_cache_populate_rows |
histogram | siglake-core, siglake-storage |
siglake_query_scan_file_cache_requests_total |
counter | siglake-storage |
siglake_query_scan_file_cache_row_group_served |
histogram | siglake-storage |
siglake_query_scan_file_cache_row_group_total |
counter | siglake-storage |
siglake_query_scan_ordered_merge_partitions_total |
counter | siglake-storage |
siglake_query_scan_ordered_merge_streams_total |
counter | siglake-storage |
siglake_query_scan_ordered_partition_coalesce_total |
counter | siglake-storage |
siglake_query_scan_ordered_sort_clusters_total |
counter | siglake-storage |
siglake_query_scan_output_ordering_total |
counter | siglake-storage |
siglake_query_scan_partition_decoded_bytes |
histogram | siglake-storage |
siglake_query_scan_partition_fetched_bytes |
histogram | siglake-storage |
siglake_query_scan_partition_first_batch_rows |
histogram | siglake-storage |
siglake_query_scan_partition_first_batch_seconds |
histogram | siglake-storage |
siglake_query_scan_partition_inter_batch_gap_seconds_max |
histogram | siglake-storage |
siglake_query_scan_partition_inter_batch_gap_seconds_mean |
histogram | siglake-storage |
siglake_query_scan_partition_output_batches |
histogram | siglake-storage |
siglake_query_scan_partition_output_rows |
histogram | siglake-storage |
siglake_query_scan_partition_stream_build_seconds |
histogram | siglake-storage |
siglake_query_scan_partition_stream_create_seconds |
histogram | siglake-storage |
siglake_query_scan_partition_wall_seconds |
histogram | siglake-storage |
siglake_query_scan_settle_seconds |
histogram | siglake-query-server |
siglake_query_shard_pin_total |
counter | siglake-query-server |
siglake_query_side_aggs_cache_total |
counter | siglake-storage |
siglake_query_sql_result_cache_bytes |
gauge | siglake-query-server |
siglake_query_sql_result_cache_entries |
gauge | siglake-query-server |
siglake_query_sql_result_cache_requests_total |
counter | siglake-query-server |
siglake_query_tenant_denied_total |
counter | siglake-query-server |
siglake_query_wal_buffer_build_errors_total |
counter | siglake-query-server |
siglake_query_wal_buffer_distributed_used_total |
counter | siglake-query-server |
siglake_query_wal_buffer_refused_bytes_total |
counter | siglake-query-server |
siglake_query_wal_buffer_refused_segment_cap_total |
counter | siglake-query-server |
siglake_query_wal_buffer_rows |
histogram | siglake-query-server |
siglake_query_wal_buffer_segment_errors_total |
counter | siglake-query-server |
siglake_query_wal_buffer_segments_examined |
histogram | siglake-query-server |
siglake_query_wal_buffer_segments_excluded_total |
counter | siglake-query-server |
siglake_query_wal_buffer_segments_kept_total |
counter | siglake-query-server |
siglake_query_wal_buffer_segments_time_pruned_total |
counter | siglake-query-server |
siglake_query_wal_buffer_skipped_no_match_total |
counter | siglake-query-server |
siglake_query_wal_buffer_stale_owner_total |
counter | siglake-query-server |
siglake_query_wal_buffer_stale_segments_total |
counter | siglake-query-server |
siglake_query_wal_buffer_used_total |
counter | siglake-query-server |
siglake_query_warm_cycle_in_progress |
gauge | siglake-query-server |
siglake_query_warm_cycle_timeouts_total |
counter | siglake-query-server |
siglake_query_warm_cycles_abandoned_total |
counter | siglake-query-server |
siglake_query_warm_cycles_completed_total |
counter | siglake-query-server |
siglake_query_warm_cycles_started_total |
counter | siglake-query-server |
siglake_query_warm_last_completed_seconds |
gauge | siglake-query-server |
siglake_query_warm_probe_timeouts_total |
counter | siglake-query-server |
siglake_query_windowed_agg_boundary_ranges |
histogram | siglake-storage |
siglake_query_windowed_agg_boundary_seconds |
histogram | siglake-storage |
siglake_query_windowed_agg_core_buckets |
histogram | siglake-storage |
siglake_query_windowed_agg_core_seconds |
histogram | siglake-storage |
siglake_query_windowed_agg_fallback_total |
counter | siglake-storage |
siglake_query_windowed_count_snapshot_agg_total |
counter | siglake-storage |
siglake_query_windowed_group_snapshot_agg_total |
counter | siglake-storage |
siglake_rate_budget_backend_errors_total |
counter | siglake-ingest |
siglake_side_aggregate_publish_failures_total |
counter | siglake-storage |
siglake_side_aggregates_bytes |
histogram | siglake-storage |
siglake_side_aggregates_load_total |
counter | siglake-storage |
siglake_storage_schema_drift_total |
counter | siglake-storage |
siglake_subscription_bootstrap_total |
counter | siglake-storage |
siglake_subscription_fast_path_total |
counter | siglake-storage |
siglake_subscription_history_gap_total |
counter | siglake-storage |
siglake_subscription_incremental_total |
counter | siglake-storage |
siglake_subscription_polls_total |
counter | siglake-storage |
siglake_subscription_rewrite_commits_skipped_total |
counter | siglake-storage |
siglake_table_gauge_table_timeouts_total |
counter | siglake-storage |
siglake_table_gauges_sampled_at_seconds |
gauge | siglake-storage |
siglake_table_gen_capped_files |
gauge | siglake-storage |
siglake_table_leading_edge_small_bytes |
gauge | siglake-storage |
siglake_table_leading_edge_small_files |
gauge | siglake-storage |
siglake_table_level_files |
gauge | siglake-storage |
siglake_table_live_data_bytes |
gauge | siglake-storage |
siglake_table_live_data_files |
gauge | siglake-storage |
siglake_table_live_data_files_sample_failures_total |
counter | siglake-compactor |
siglake_table_overlap_depth |
gauge | siglake-storage |
siglake_time_bucket_width_normalized_total |
counter | siglake-storage |
siglake_time_group_column_dropped_total |
counter | siglake-storage |
siglake_time_group_delta_built_total |
counter | siglake-storage |
siglake_time_group_delta_columns |
histogram | siglake-storage |
siglake_time_group_delta_skipped_total |
counter | siglake-storage |
siglake_wal_append_batch_rows |
histogram | siglake-wal |
siglake_wal_append_duration_seconds |
histogram | siglake-wal |
siglake_wal_append_events_rows |
histogram | siglake-wal |
siglake_wal_append_flush_duration_seconds |
histogram | siglake-wal |
siglake_wal_append_write_duration_seconds |
histogram | siglake-wal |
siglake_wal_appends_total |
counter | siglake-wal |
siglake_wal_bytes_written_total |
counter | siglake-wal |
siglake_wal_consumer_misdirected |
gauge | siglake-wal |
siglake_wal_crc_mismatch_total |
counter | siglake-wal |
siglake_wal_identity_resolve_failures_total |
counter | siglake-ingest |
siglake_wal_ipc_framing_refused_total |
counter | siglake-wal |
siglake_wal_local_sealed_segments |
gauge | siglake-cli |
siglake_wal_local_sweep_deleted_total |
counter | siglake-cli |
siglake_wal_local_sweep_errors_total |
counter | siglake-cli |
siglake_wal_mirror_active_bytes_uploaded_total |
counter | siglake-wal |
siglake_wal_mirror_active_uploads_total |
counter | siglake-wal |
siglake_wal_mirror_bytes_uploaded_total |
counter | siglake-wal |
siglake_wal_mirror_failures_total |
counter | siglake-wal |
siglake_wal_mirror_pin_duration_seconds |
histogram | siglake-core, siglake-wal |
siglake_wal_mirror_queue_depth |
gauge | siglake-wal |
siglake_wal_mirror_queue_wait_seconds |
histogram | siglake-wal |
siglake_wal_mirror_register_abandoned_total |
counter | siglake-cli |
siglake_wal_mirror_register_lag_seconds |
histogram | siglake-cli |
siglake_wal_mirror_register_total |
counter | siglake-cli |
siglake_wal_mirror_segments_total |
counter | siglake-wal |
siglake_wal_mirror_sweep_committed_skipped_total |
counter | siglake-wal |
siglake_wal_mirror_sweep_unregistered_total |
counter | siglake-cli |
siglake_wal_mirror_sweep_uploads_total |
counter | siglake-wal |
siglake_wal_mirror_upload_abandoned_total |
counter | siglake-wal |
siglake_wal_mirror_upload_attempts_total |
counter | siglake-wal |
siglake_wal_mirror_upload_present_after_error_total |
counter | siglake-wal |
siglake_wal_partial_tail_dropped_total |
counter | siglake-wal |
siglake_wal_partials_adopted_total |
counter | siglake-wal |
siglake_wal_partials_left_to_owner_total |
counter | siglake-wal |
siglake_wal_partials_recovered_total |
counter | siglake-wal |
siglake_wal_record_batch_build_duration_seconds |
histogram | siglake-wal |
siglake_wal_rows_written_total |
counter | siglake-wal |
siglake_wal_seal_bytes |
histogram | siglake-wal |
siglake_wal_seal_compression_ratio |
histogram | siglake-wal |
siglake_wal_seal_duration_seconds |
histogram | siglake-wal |
siglake_wal_seal_finish_duration_seconds |
histogram | siglake-wal |
siglake_wal_seal_rename_duration_seconds |
histogram | siglake-wal |
siglake_wal_seal_rows |
histogram | siglake-wal |
siglake_wal_seal_sync_duration_seconds |
histogram | siglake-wal |
siglake_wal_seals_total |
counter | siglake-wal |
siglake_wal_segments_sealed |
gauge | siglake-cli |
siglake_wal_segments_sealed_sample_age_seconds |
gauge | siglake-cli |
siglake_wal_segments_sealed_total |
counter | siglake-wal |
siglake_wal_writer_rebinds_total |
counter | siglake-wal |
siglake_write_index_build_seconds |
histogram | siglake-storage |
siglake_write_index_deferred_total |
counter | siglake-storage |
siglake_write_inverted_index_build_seconds |
histogram | siglake-storage |
siglake_write_raw_file_bloom_build_seconds |
histogram | siglake-storage |
siglake_write_raw_rowgroup_bloom_build_seconds |
histogram | siglake-storage |