Export Siglake's own logs and traces¶
Point Siglake at an OpenTelemetry Protocol (OTLP) collector and every role
exports its own log records and spans over OTLP/HTTP. This is Siglake as a
client, and it is separate from the ingest endpoints your applications export
to. Export is off until OTEL_EXPORTER_OTLP_ENDPOINT is set, and it never
carries metrics: those stay on /metrics for Prometheus, so an OTLP-only
backend needs a collector with a Prometheus
receiver.
Sending application telemetry to Siglake is a different job, covered by the OpenTelemetry guide.
Before you start¶
- Siglake 0.2.0 or later on every role you want telemetry from. Earlier images ignore these variables.
- A collector with an OTLP/HTTP receiver, reachable from the pods. The examples use port 4318. The exporter speaks OTLP/HTTP with protobuf payloads and nothing else, so an endpoint on a gRPC port does not work.
- Permission to change environment variables on the roles and restart them.
- If
networkPolicy.enabledis set, put the collector in the release namespace. The chart's egress rules allow DNS, the configured S3, Postgres and OIDC destinations, and the release namespace, and nothing else.
Turn on export for a Helm release¶
The chart renders no OTel block. Set the variables through extraEnv, which
every role accepts, and which also accepts a valueFrom entry:
extraEnv:
- name: OTEL_EXPORTER_OTLP_ENDPOINT
value: http://otel-collector.siglake.svc:4318
- name: OTEL_SERVICE_NAMESPACE
value: siglake-prod
- name: POD_NAME
valueFrom:
fieldRef:
fieldPath: metadata.name
- Add the block to your values file. Global
extraEnvreaches every role and the schema-migration hook Job;ingester.extraEnv,compactor.extraEnvandquery.extraEnvreach one role each. - Run
helm upgradeand let the roles restart. Each process reads these variables once, at startup, so the change takes effect on the restart and not before. - Verify the export.
Each process names itself siglake-<component>: siglake-ingest,
siglake-compactor, siglake-query, siglake-operator, and siglake-cli for
the maintenance subcommands. OTEL_SERVICE_NAME overrides that, so set it per
role or not at all. The configuration
reference
lists every variable.
Turn on export for a SiglakeCluster¶
The custom resource has no OTel block either. spec.extraEnv carries the
variables, and the operator appends them to every container it renders:
apiVersion: siglake.limnion.ai/v1alpha1
kind: SiglakeCluster
metadata:
name: prod
spec:
extraEnv:
- name: OTEL_EXPORTER_OTLP_ENDPOINT
value: http://otel-collector.siglake.svc:4318
- name: OTEL_SERVICE_NAMESPACE
value: siglake-prod
Two consequences of that one list. The tiers export together or not at all,
because there is no per-tier extraEnv on the custom resource. And entries are
plain name/value, so POD_NAME from the downward API is not expressible: spans
and log records carry host.name from the container's HOSTNAME, without the
service.instance.id the Helm example above sets. Fields that catch people
out covers the rest of that escape
hatch.
What Siglake exports¶
Log records are the log lines the process already writes. The bridge forwards
whatever RUST_LOG admits, so the filter selects the exported records and the
console output alike, and the console layer keeps writing to stderr either way.
Spans sit at the boundaries that cost something:
| Role | Span | Covers |
|---|---|---|
| Ingester | post_otlp_logs, post_otlp_traces |
One OTLP export request |
| Ingester | ingest_batch |
The write of one decoded batch |
| Compactor | run_once |
One drain cycle |
| Query | http.server |
One API request, with method, route and status |
| Query | distributed_inner |
One distributed SQL query and its fan-out |
A distributed query is one trace. The coordinator injects a W3C traceparent
header into each shard request and the worker's middleware extracts it, so the
worker spans are children of the coordinator's span rather than unrelated
roots. A caller that sends its own traceparent becomes the parent of
http.server the same way.
Siglake entry points that use the shared telemetry initializer flush its batch
processors on the way out. This lets a one-shot command and a graceful
SIGTERM deliver buffered data. siglake-loadgen and siglake-corpus use
console-only subscribers, so they create no OpenTelemetry providers or batches
to flush.
Join query scan log records¶
Every query execution logs query execution start with a process-local
query_execution_id, its endpoint and its SQL. The id is also present on the
buffered path's sql execution profile and terminal sql query profile
records. Use it to join concurrent executions without relying on log order.
Two scan events carry that execution id and a scan_id minted for the scan
node:
| Event | When it is written |
|---|---|
siglake query scan reader tuning |
Once when the scan node is planned. |
siglake query source partition profile |
Once when a partition ends, fails or is dropped. |
An early LIMIT can leave partitions unwinding after the request profile is
written. An NDJSON body can outlive its handler too. Join these events on both
fields instead of assuming adjacent lines belong to one request. A
query_execution_id of 0 marks an internal scan whose session carried no
request id. One execution can have several scan ids.
Execution ids are local to one process. A distributed query's coordinator and workers mint unrelated values, so join those pods with the W3C trace context that the coordinator propagates. That cross-pod join needs OpenTelemetry export configured. Streaming and refused executions do not gain a terminal profile record.
Send the Prometheus metrics through the same collector¶
Metrics do not leave the process as OTLP. Every role exports its counters,
gauges and histograms to /metrics on the metrics port, which is what the
chart's ServiceMonitor, the alert rules, the KEDA scalers and the Grafana
dashboard read. To land them in an OTLP-only backend, give the collector a
Prometheus receiver that discovers pods. Discover pods rather than Services: a
Service hostname is one scrape target however many pods stand behind it, and
the chart runs two query replicas by default. Each scrape of that target reads
one pod, the next can read the other, and both readings arrive under the same
target identity, so counters go backwards and gauges flip between processes.
receivers:
otlp:
protocols:
http:
endpoint: 0.0.0.0:4318
prometheus:
config:
scrape_configs:
- job_name: siglake
scrape_interval: 30s
kubernetes_sd_configs:
- role: pod
namespaces:
names: [<namespace>]
relabel_configs:
- source_labels: [__meta_kubernetes_pod_label_app_kubernetes_io_name]
regex: siglake
action: keep
- source_labels: [__meta_kubernetes_pod_container_port_name]
regex: metrics
action: keep
- source_labels: [__meta_kubernetes_namespace]
target_label: namespace
- source_labels: [__meta_kubernetes_pod_name]
target_label: pod
- source_labels: [__meta_kubernetes_pod_label_app_kubernetes_io_component]
target_label: component
exporters:
otlphttp/backend:
endpoint: https://otlp.example.com
headers:
api-key: ${env:BACKEND_API_KEY}
service:
pipelines:
logs:
receivers: [otlp]
exporters: [otlphttp/backend]
traces:
receivers: [otlp]
exporters: [otlphttp/backend]
metrics:
receivers: [prometheus]
exporters: [otlphttp/backend]
Pod discovery makes one target per declared container port, so the two keep
rules do the narrowing: the first to Siglake's pods, the second to the port
named metrics, which is 9100 on the ingester, 9101 on the compactor and 9105
on the query server. Without the port rule the job also scrapes 8088 and 8089,
which serve no /metrics. The three target_label rules put pod identity on
every series: namespace and pod are what the chart's alert rules group by,
and component names the role. instance stays the pod address and changes
when a pod is replaced, so group by pod.
Two releases in one namespace need a third keep rule, on
__meta_kubernetes_pod_label_app_kubernetes_io_instance against the release
name. A SiglakeCluster needs no change: the operator writes the same three
app.kubernetes.io labels, and it renders no compactor metrics Service at
all, so pod discovery is the only way to reach 9101 there.
The metric names do not change on the way through, so siglake_* in the
backend means the same thing as siglake_* in Prometheus.
Monitoring has the port per role. The metrics
port checks no token, so keep the collector inside the cluster and off any
Ingress.
Let the collector list pods¶
Pod discovery reads the Kubernetes API, so the collector's ServiceAccount
needs get, list and watch on pods in the namespace it scrapes. Without
them the receiver logs a pods is forbidden error from the API server and the
job holds no targets. Nothing reports a scrape failure, because there is
nothing to scrape.
apiVersion: rbac.authorization.k8s.io/v1
kind: Role
metadata:
name: otel-collector-pod-discovery
namespace: <namespace>
rules:
- apiGroups: [""]
resources: ["pods"]
verbs: ["get", "list", "watch"]
---
apiVersion: rbac.authorization.k8s.io/v1
kind: RoleBinding
metadata:
name: otel-collector-pod-discovery
namespace: <namespace>
roleRef:
apiGroup: rbac.authorization.k8s.io
kind: Role
name: otel-collector-pod-discovery
subjects:
- kind: ServiceAccount
name: <collector service account>
namespace: <collector namespace>
A Role covers one namespace, which is enough for a collector next to one
release. For a collector that scrapes several namespaces, use a ClusterRole
and a ClusterRoleBinding with the same rule, and drop the namespaces block
from kubernetes_sd_configs to discover cluster-wide.
Check that both query pods are scraped¶
The chart's default replica counts are one ingester, one compactor and two query servers, so a working job holds four targets and two of them are query pods. Run this check after any change to the scrape config.
- List the pods the job should find. Expect four, two of them
query:
- Count the targets in the backend. The Prometheus receiver emits one
upseries per target and the metrics pipeline forwards it:
Expect query at 2, and ingester and compactor at 1 each.
A query count of 1 means the second replica was never discovered: compare
the pod list from step 1, then read the collector's log for the API error. A
count of 0 while the pods are running means the keep rules matched nothing,
usually a nameOverride that changed app.kubernetes.io/name.
If your backend drops up, count by (component) (siglake_build_info) reads
the same way: every role exports it, one series per process.
Turn export off¶
| What you want | Setting |
|---|---|
| No export at all | Leave OTEL_EXPORTER_OTLP_ENDPOINT unset or empty. No exporter is built. |
| Off for now, endpoint kept | SIGLAKE_OTEL_DISABLED=1 |
| Log records only | OTEL_TRACES_EXPORTER=none |
| Spans only | OTEL_LOGS_EXPORTER=none |
Turning export off costs one restart, the same as turning it on. With the endpoint unset, the exporters and their batch processors are never constructed and the process behaves as it did in 0.1.0.
Verify the export¶
Run a one-shot command from a machine that can reach both the collector and a query server. It initializes telemetry, does its work, and flushes on exit:
export OTEL_EXPORTER_OTLP_ENDPOINT=http://localhost:4318
export OTEL_SERVICE_NAME=siglake-verify
siglake sql --endpoint http://localhost:8089 "SELECT 1"
Then check three things, in this order:
- The backend has a service called
siglake-verify. If it does not, the endpoint is wrong or the collector is unreachable from where you ran the command. - The deployed roles report themselves as
siglake-ingest,siglake-compactorandsiglake-query. A missing role means the variables did not reach its pods, whichkubectl exec <pod> -- env | grep OTEL_confirms. - A SQL query against the query tier produces one trace with an
http.serverspan and, on a distributed deployment, worker spans under it. Query traces are the slowest signal to appear, because the batch span processor holds them briefly.