Grafana and Jaeger¶
Use Grafana to view Siglake metrics, traces, and SQL results. The integration
connects three data sources: Prometheus metrics for
operational dashboards, the Jaeger shim for traces, and the Infinity data
source over /api/v1/sql for logs. Each one is wired below. Before you start,
give Grafana network access to the Prometheus and Siglake query services.
The Jaeger shim serves only the HTTP subset of the Jaeger query API that Grafana's Jaeger data source renders, with no gRPC SpanReader for tools that expect one. Its two name-list routes are cached only while a list and its lookup string fit a 512 KiB allowance. See Integration scope for workflow choices and panel-specific constraints.
Import the operations dashboard¶
deploy/grafana/siglake-overview.json
is a starter dashboard over the Prometheus metrics. Import it manually because
the chart does not render it. Point it at the Prometheus that scrapes your
Siglake roles. In its Query row, Breaker trips by kind graphs the
five-minute increase in query-breaker trips by breaker and priority, making
row-ceiling trips and memory-pool or spill-cap refusals visible separately.
With Prometheus Operator:
See Starter Grafana dashboard for the rest of the panel coverage and the alert rules that page operators.
Connect Grafana to the Jaeger shim¶
Traces reach Grafana through the Jaeger shim: the query server implements the HTTP subset of the Jaeger query API that Grafana's Jaeger data source renders.
Add a Jaeger data source with this URL:
Replace <index> with your traces-shaped index, such as
siglake-traces-default.
Endpoints the Jaeger shim supports¶
Four routes, all GET and all under /api/v1/jaeger/<index>. Each is bounded
by ceilings the server derives from the request's admission reservation, and
each refuses the whole response rather than returning part of one.
| Endpoint | Purpose | Ceilings that bound it |
|---|---|---|
GET …/api/services |
Service list. | Distinct names, accumulated Arrow bytes. |
GET …/api/services/{service}/operations |
Operations for a service. | Distinct names, accumulated Arrow bytes. |
GET …/api/traces |
Find traces. | Traces (the limit check), span rows, accumulated Arrow bytes. |
GET …/api/traces/{trace_id} |
Fetch one trace. | Span rows, accumulated Arrow bytes. |
api/traces accepts service, operation, start, end, limit,
minDuration, maxDuration, and tags (a JSON object string).
A refusal is one of five status codes: 400 for a search limit above the
trace ceiling, 413 for any other ceiling, 429 when the admission budget
stays full, 503 when the memory pool or spill cap refuses an allocation, and
504 when the wall clock expires. Result ceilings and what they
refuse gives the derivation, and
Status codes the shim returns gives one row
per code. No ceiling has a Jaeger-specific setting: they follow the admission
budget and pod memory. There is no gRPC SpanReader.
How the name-list cache behaves¶
The two name-list routes share /api/v1/sql's result cache. The first poll for
an Iceberg snapshot executes the complete distinct-name query; repeats for
that snapshot are cache hits. A new snapshot is a new lookup, so the query runs
again.
The shared store holds at most 256 entries and 4 MiB of retained heap, not
encoded JSON size. That 4 MiB, and the
siglake_query_sql_result_cache_bytes
gauge that reports it, count each entry's body, the lookup string the entry is
stored under, and one pointer for every recency marker the store keeps. For
/api/v1/sql that lookup string is the normalized query text, so a long query
is charged against the same budget as the rows it returns. One entry never
retains more than its caller's per-entry allowance: 4 MiB for /api/v1/sql,
512 KiB for a name list. An entry over that allowance is refused rather than
stored, so it cannot evict the entries around it.
A name list is stored only when its names and its lookup string together fit
that 512 KiB. The shim's derived name ceiling bounds the rows, not the SQL
cache's 128-row entry cap. A list over the allowance is still returned whole,
not truncated, but is never stored and therefore executes in full on every
poll. There is no Jaeger-specific cache knob, and the shared caps are not
raised for these routes. Caching does not change error behavior: 413, 429,
503, and 504 behave exactly as they do on a cache miss.
Admission budget and query deadline¶
All four routes use the same per-pod interactive admission budget
(SIGLAKE_QUERY_ADMISSION_BUDGET_BYTES)
and query wall-clock timeout as /api/v1/sql. Each request makes one admission
reservation. If the budget remains full through
SIGLAKE_QUERY_ADMISSION_WAIT_MS,
the route returns 429 with Retry-After; wait for that delay before retrying.
The wall clock starts before tenant resolution and includes index lookup, table
registration, and query execution. A trace search gets one deadline across
both trace-id selection and span retrieval, so the second phase does not get a
fresh budget. If the deadline expires during preparation or either query
phase, the route returns 504 and cancels the in-flight scan.
Result ceilings and what they refuse¶
The shim bounds traces, span rows, accumulated Arrow bytes, and distinct names.
These ceilings are derived from the admission reservation, after a 2 MiB
request floor, using conservative measured costs of 12 KiB per trace, 8.5 KiB
per span row, 4x the Arrow bytes, and 1.4 KiB per name. They follow the
admission budget, share divisor, and pod memory; there is no Jaeger-specific
environment knob. The resolved interactive row cap (the tier default, clamped
by
SIGLAKE_QUERY_MAX_ROWS) also applies
where it is tighter.
A search limit above the trace ceiling is rejected with 400 before
admission, index lookup, or planning. Span rows and Arrow bytes are checked
mid-flight at batch boundaries across both search phases; the list routes
check names and Arrow bytes. The whole request shares each ceiling, so a search
counts the spans of all selected traces together. Narrow service,
operation, the time range, tags, or limit when a request is refused.
Every result-ceiling refusal is 413 with no data, partial trace, or
Retry-After. /api/v1/sql can flag a truncated records envelope, but the
Jaeger response has only data and total, so it cannot safely express
missing spans. The shim takes no per-request limits object.
Status codes the shim returns¶
| Status | On these routes |
|---|---|
400 |
A trace search's limit exceeded the derived trace ceiling (or another request parameter was invalid). Checked before admission, index lookup, or planning. |
413 |
A span-row, Arrow-byte, distinct-name, or tighter interactive row ceiling was exceeded. Refused whole: no data and no Retry-After; narrow the request. |
429 |
The shared interactive admission budget stayed full through the admission wait. Honor Retry-After. |
503 |
The query memory pool or spill byte cap refused an allocation. Honor Retry-After. |
504 |
The interactive wall-clock budget expired during preparation or either search phase. |
The shim supports HTTP only
Tools that expect the Jaeger gRPC SpanReader interface will not work.
The shim covers what Grafana's Jaeger data source calls over HTTP, and
nothing more.
Pointing it at a non-traces index returns a clear error rather than empty results.
Wire Grafana to logs through SQL¶
Siglake has no purpose-built Grafana data source, so a log panel is a generic
JSON data source pointed at /api/v1/sql. Use Infinity unless you already
have another one in the cluster.
Use the Infinity data source (recommended)¶
Infinity is the data source we recommend for SQL panels. The
Infinity
plugin queries arbitrary JSON HTTP endpoints, which is exactly what
/api/v1/sql is.
Configure a POST request to http://siglake-query:8089/api/v1/sql with body:
{"query": "SELECT timestamp, host, raw FROM events WHERE timestamp >= now() - INTERVAL '1 hour' ORDER BY timestamp DESC LIMIT 500"}
Parse the response with root selector rows, and use columns to name the
fields.
For dashboard variables, use the dashboard aggregation choice rule. Prefer an eligible metadata shape for a variable that runs on every dashboard load:
Test the query with records format before putting it on an auto-refresh
schedule. The choice rule explains how to read stats.served_by and repair a
materialized fallback.
NDJSON streaming for large panels¶
For larger panels, "format": "ndjson" streams newline-delimited JSON
(NDJSON), one object per line, which some data sources handle more gracefully
than a single large JSON body.
The 200 is committed with the first bytes, so it no longer means the result is
complete: a panel that hits the row cap or a pool refusal mid-stream renders a
short series with no error unless you inspect the last line for a _meta
trailer. See NDJSON trailers.
Bound dashboard queries¶
Dashboards auto-refresh, and a badly-scoped panel becomes a repeating expensive query. Protect yourself:
- Always bound the time range in the SQL, not just in Grafana's picker.
- Always
LIMIT. - Set
query.breaker.rowsScannedCeilingappropriately. - Watch
query_auditfor panels that turn out to belargeorhuge:
SELECT query, count(*) AS runs, avg(duration_ms) AS avg_ms
FROM query_audit
WHERE timestamp >= now() - INTERVAL '24 hours'
AND complexity IN ('large','huge')
GROUP BY query ORDER BY runs DESC;
Add a live tail¶
GET /api/v1/stream is a Server-Sent Events tail of the ingest path. Grafana
doesn't consume SSE natively, but it's straightforward for a custom panel or a
terminal:
This endpoint is best-effort. It keeps no history, provides no durability, and
drops subscribers that fall behind. Filter on the client because the server
does not support stream filters. Events are sent before the WAL append, so a
tailed event is not yet acknowledged. A batch then refused for backlog (503)
or whose append fails (500) has still appeared on the panel. The caller's
retry publishes it again. See
Ingest and the WAL.
Integration scope¶
The integration provides metric dashboards, trace views through the Jaeger shim, and SQL-backed log panels using Grafana's data sources. Siglake handles storage and query; Grafana or another external tool supplies the presentation and alerting workflow.
- Use an external alerting system for rule evaluation, notifications, and on-call workflows. Alerting is outside Siglake's core product scope.
- Save exploration queries in your client and query
query_auditfor recorded query activity. Siglake exposes these data through APIs rather than a built-in query-history or saved-search UI. - Log-context expansion, visual field extraction, and pattern discovery need client-side tooling; this integration does not supply those UI features.
- Log exploration uses time ordering and SQL filters rather than relevance ranking. See Scope and limitations.
Two shapes of panel need care rather than being unavailable. A live tail is
Server-Sent Events, which Grafana does not consume natively, so it needs a
custom panel. A large panel served as NDJSON commits its 200 with the first
bytes, so a mid-stream refusal renders a short series unless the panel reads
the _meta trailer.
Grafana plus /api/v1/sql provides a starting point for exploration and
dashboards that you can adapt to your team's workflows.