Skip to content

Grafana and Jaeger

Use Grafana to view Siglake metrics, traces, and SQL results. Siglake ships no user interface, and Grafana reaches it three ways: Prometheus metrics for operational dashboards, the Jaeger shim for traces, and the Infinity data source over /api/v1/sql for logs. Each one is wired below. Before you start, give Grafana network access to the Prometheus and Siglake query services.

The Jaeger shim serves only the HTTP subset of the Jaeger query API that Grafana's Jaeger data source renders, with no gRPC SpanReader for tools that expect one. Its two name-list routes are cached only while a list and its lookup string fit a 512 KiB allowance. What Grafana cannot give you either way is in What the Grafana integration does not provide.

Import the operations dashboard

deploy/grafana/siglake-overview.json is a starter dashboard over the Prometheus metrics. Import it manually because the chart does not render it. Point it at the Prometheus that scrapes your Siglake roles. In its Query row, Breaker trips by kind graphs the five-minute increase in query-breaker trips by breaker and priority, making row-ceiling trips and memory-pool or spill-cap refusals visible separately.

With Prometheus Operator:

serviceMonitor:
  enabled: true

See Starter Grafana dashboard for the rest of the panel coverage and the alert rules that page operators.

Connect Grafana to the Jaeger shim

Traces reach Grafana through the Jaeger shim: the query server implements the HTTP subset of the Jaeger query API that Grafana's Jaeger data source renders.

Add a Jaeger data source with this URL:

http://siglake-query.siglake.svc:8089/api/v1/jaeger/<index>

Replace <index> with your traces-shaped index, such as siglake-traces-default.

Endpoints the Jaeger shim supports

Four routes, all GET and all under /api/v1/jaeger/<index>. Each is bounded by ceilings the server derives from the request's admission reservation, and each refuses the whole response rather than returning part of one.

Endpoint Purpose Ceilings that bound it
GET …/api/services Service list. Distinct names, accumulated Arrow bytes.
GET …/api/services/{service}/operations Operations for a service. Distinct names, accumulated Arrow bytes.
GET …/api/traces Find traces. Traces (the limit check), span rows, accumulated Arrow bytes.
GET …/api/traces/{trace_id} Fetch one trace. Span rows, accumulated Arrow bytes.

api/traces accepts service, operation, start, end, limit, minDuration, maxDuration, and tags (a JSON object string).

A refusal is one of five status codes: 400 for a search limit above the trace ceiling, 413 for any other ceiling, 429 when the admission budget stays full, 503 when the memory pool or spill cap refuses an allocation, and 504 when the wall clock expires. Result ceilings and what they refuse gives the derivation, and Status codes the shim returns gives one row per code. No ceiling has a Jaeger-specific setting: they follow the admission budget and pod memory. There is no gRPC SpanReader.

How the name-list cache behaves

The two name-list routes share /api/v1/sql's result cache. The first poll for an Iceberg snapshot executes the complete distinct-name query; repeats for that snapshot are cache hits. A new snapshot is a new lookup, so the query runs again.

The shared store holds at most 256 entries and 4 MiB of retained heap, not encoded JSON size. That 4 MiB, and the siglake_query_sql_result_cache_bytes gauge that reports it, count each entry's body, the lookup string the entry is stored under, and one pointer for every recency marker the store keeps. For /api/v1/sql that lookup string is the normalized query text, so a long query is charged against the same budget as the rows it returns. One entry never retains more than its caller's per-entry allowance: 4 MiB for /api/v1/sql, 512 KiB for a name list. An entry over that allowance is refused rather than stored, so it cannot evict the entries around it.

A name list is stored only when its names and its lookup string together fit that 512 KiB. The shim's derived name ceiling bounds the rows, not the SQL cache's 128-row entry cap. A list over the allowance is still returned whole, not truncated, but is never stored and therefore executes in full on every poll. There is no Jaeger-specific cache knob, and the shared caps are not raised for these routes. Caching does not change error behavior: 413, 429, 503, and 504 behave exactly as they do on a cache miss.

Admission budget and query deadline

All four routes use the same per-pod interactive admission budget (SIGLAKE_QUERY_ADMISSION_BUDGET_BYTES) and query wall-clock timeout as /api/v1/sql. Each request makes one admission reservation. If the budget remains full through SIGLAKE_QUERY_ADMISSION_WAIT_MS, the route returns 429 with Retry-After; wait for that delay before retrying.

The wall clock starts before tenant resolution and includes index lookup, table registration, and query execution. A trace search gets one deadline across both trace-id selection and span retrieval, so the second phase does not get a fresh budget. If the deadline expires during preparation or either query phase, the route returns 504 and cancels the in-flight scan.

Result ceilings and what they refuse

The shim bounds traces, span rows, accumulated Arrow bytes, and distinct names. These ceilings are derived from the admission reservation, after a 2 MiB request floor, using conservative measured costs of 12 KiB per trace, 8.5 KiB per span row, 4x the Arrow bytes, and 1.4 KiB per name. They follow the admission budget, share divisor, and pod memory; there is no Jaeger-specific environment knob. The resolved interactive row cap (the tier default, clamped by SIGLAKE_QUERY_MAX_ROWS) also applies where it is tighter.

A search limit above the trace ceiling is rejected with 400 before admission, index lookup, or planning. Span rows and Arrow bytes are checked mid-flight at batch boundaries across both search phases; the list routes check names and Arrow bytes. The whole request shares each ceiling, so a search counts the spans of all selected traces together. Narrow service, operation, the time range, tags, or limit when a request is refused.

Every result-ceiling refusal is 413 with no data, partial trace, or Retry-After. /api/v1/sql can flag a truncated records envelope, but the Jaeger response has only data and total, so it cannot safely express missing spans. The shim takes no per-request limits object.

Status codes the shim returns

Status On these routes
400 A trace search's limit exceeded the derived trace ceiling (or another request parameter was invalid). Checked before admission, index lookup, or planning.
413 A span-row, Arrow-byte, distinct-name, or tighter interactive row ceiling was exceeded. Refused whole: no data and no Retry-After; narrow the request.
429 The shared interactive admission budget stayed full through the admission wait. Honor Retry-After.
503 The query memory pool or spill byte cap refused an allocation. Honor Retry-After.
504 The interactive wall-clock budget expired during preparation or either search phase.

The shim supports HTTP only

Tools that expect the Jaeger gRPC SpanReader interface will not work. The shim covers what Grafana's Jaeger data source calls over HTTP, and nothing more.

Pointing it at a non-traces index returns a clear error rather than empty results.

Wire Grafana to logs through SQL

Siglake has no purpose-built Grafana data source, so a log panel is a generic JSON data source pointed at /api/v1/sql. Use Infinity unless you already have another one in the cluster.

Infinity is the data source we recommend for SQL panels. The Infinity plugin queries arbitrary JSON HTTP endpoints, which is exactly what /api/v1/sql is.

Configure a POST request to http://siglake-query:8089/api/v1/sql with body:

{"query": "SELECT timestamp, host, raw FROM events WHERE timestamp >= now() - INTERVAL '1 hour' ORDER BY timestamp DESC LIMIT 500"}

Parse the response with root selector rows, and use columns to name the fields.

For dashboard variables, use the dashboard aggregation choice rule. Prefer an eligible metadata shape for a variable that runs on every dashboard load:

{"query": "SELECT host, count(*) AS n FROM events GROUP BY host ORDER BY host"}

Test the query with records format before putting it on an auto-refresh schedule. The choice rule explains how to read stats.served_by and repair a materialized fallback.

NDJSON streaming for large panels

For larger panels, "format": "ndjson" streams newline-delimited JSON (NDJSON), one object per line, which some data sources handle more gracefully than a single large JSON body.

The 200 is committed with the first bytes, so it no longer means the result is complete: a panel that hits the row cap or a pool refusal mid-stream renders a short series with no error unless you inspect the last line for a _meta trailer. See NDJSON trailers.

Bound dashboard queries

Dashboards auto-refresh, and a badly-scoped panel becomes a repeating expensive query. Protect yourself:

  • Always bound the time range in the SQL, not just in Grafana's picker.
  • Always LIMIT.
  • Set query.breaker.rowsScannedCeiling appropriately.
  • Watch query_audit for panels that turn out to be large or huge:
SELECT query, count(*) AS runs, avg(duration_ms) AS avg_ms
FROM query_audit
WHERE timestamp >= now() - INTERVAL '24 hours'
  AND complexity IN ('large','huge')
GROUP BY query ORDER BY runs DESC;

Add a live tail

GET /api/v1/stream is a Server-Sent Events tail of the ingest path. Grafana doesn't consume SSE natively, but it's straightforward for a custom panel or a terminal:

curl -N http://localhost:8088/api/v1/stream

This endpoint is best-effort. It keeps no history, provides no durability, and drops subscribers that fall behind. Filter on the client because the server does not support stream filters. Events are sent before the WAL append, so a tailed event is not yet acknowledged. A batch then refused for backlog (503) or whose append fails (500) has still appeared on the panel. The caller's retry publishes it again. See Ingest and the WAL.

What the Grafana integration does not provide

Three data sources, no Siglake plugin, and no UI of Siglake's own: what you get is metric dashboards, trace views through the Jaeger shim, and log panels that are SQL results. Four things the integration does not give you:

  • Siglake has no alerting UI or built-in alert delivery. Use Grafana alerting or another alerting system when you need notifications.
  • It has no saved searches or query history UI, although you can query query_audit.
  • It has no log-context expansion, field extraction UI, or pattern discovery UI.
  • Search results sort by time rather than relevance. See Limitations.

Two shapes of panel need care rather than being unavailable. A live tail is Server-Sent Events, which Grafana does not consume natively, so it needs a custom panel. A large panel served as NDJSON commits its 200 with the first bytes, so a mid-stream refusal renders a short series unless the panel reads the _meta trailer.

Grafana plus /api/v1/sql covers exploration and dashboards. It is not a Kibana replacement out of the box.