OpenTelemetry¶
Use the OpenTelemetry Collector or an OpenTelemetry Protocol (OTLP) exporter to send logs and traces to Siglake. The HTTP endpoints accept protobuf or JSON without a Siglake plugin.
Configure an OpenTelemetry Collector exporter¶
Before you start, make the ingest service reachable from the Collector. If
ingest authentication is enabled, set SIGLAKE_TOKEN to a valid bearer token.
Then add this pipeline to the Collector configuration:
receivers:
otlp:
protocols:
grpc:
endpoint: 0.0.0.0:4317
http:
endpoint: 0.0.0.0:4318
filelog:
include: [/var/log/containers/*.log]
operators:
- type: container
processors:
batch:
timeout: 5s
send_batch_size: 8192
send_batch_max_size: 16384 # keep under the 16 MiB body cap
k8sattributes:
resource:
attributes:
- key: host.name
from_attribute: k8s.node.name
action: insert
transform:
log_statements:
- context: log
statements:
- set(attributes["sourcetype"], "nginx:access")
- set(attributes["index"], "web")
exporters:
otlphttp/siglake:
logs_endpoint: http://siglake-ingester.siglake.svc:8088/v1/logs
traces_endpoint: http://siglake-ingester.siglake.svc:8088/v1/traces
encoding: proto
headers:
Authorization: "Bearer ${env:SIGLAKE_TOKEN}"
sending_queue:
enabled: true
queue_size: 10000
retry_on_failure:
enabled: true
initial_interval: 1s
max_interval: 30s
max_elapsed_time: 5m
service:
pipelines:
logs:
receivers: [filelog, otlp]
processors: [k8sattributes, resource, transform, batch]
exporters: [otlphttp/siglake]
traces:
receivers: [otlp]
processors: [k8sattributes, resource, batch]
exporters: [otlphttp/siglake]
Retries do not guarantee delivery
Siglake sheds load with 429 and 503 plus Retry-After under
backpressure. The Collector retries those responses, but retries expire
after five minutes and the bounded queue can fill. This configuration has
no persistent queue storage, so a Collector restart can also lose queued
telemetry.
These limits apply before Siglake acknowledges an export over HTTP or
gRPC. With the default wait_for mode, a Siglake acknowledgement means the
rows are fsynced to its local write-ahead log (WAL). See Ingest and the
WAL
for the later durability boundaries.
Map Collector attributes into the events schema¶
Check how OTLP fields land in the events schema before you send production
traffic:
events column |
Source | Fallback |
|---|---|---|
timestamp |
timeUnixNano |
observedTimeUnixNano, then ingest wall-clock |
timestamp_ns |
timeUnixNano |
observedTimeUnixNano, then ingest wall-clock |
host |
resource attribute host.name |
"unknown" |
source |
resource attribute service.name |
scope name, then "otel" |
sourcetype |
Record attribute sourcetype |
"otel:logs" |
index |
Record attribute index |
"main" |
raw |
log record body |
"" |
attributes |
everything else | null |
Both time columns resolve the same instant: a nonzero timeUnixNano, else a
nonzero observedTimeUnixNano, else the ingester's wall clock. timestamp
stores that instant to microsecond precision and remains the day() partition
and first sort column. timestamp_ns stores its exact Unix nanoseconds and is
the second sort column, breaking ties within a microsecond.
host.name must be a resource attribute. If it is missing, every event
gets host = "unknown". In the configuration above, the resource processor
copies the node name after k8sattributes enriches the resource.
sourcetype and index are record attributes, not resource attributes. The
transform processor in the configuration above sets them per log record.
When you add it to an existing configuration, merge transform into the
existing processors mapping. Keep the other processor definitions. Add it to
the logs pipeline after resource and before batch. Merge the
logs.processors list into the existing service mapping, and keep the traces
pipeline unchanged:
processors:
transform:
log_statements:
- context: log
statements:
- set(attributes["sourcetype"], "nginx:access")
- set(attributes["index"], "web")
service:
pipelines:
logs:
processors: [k8sattributes, resource, transform, batch]
A meaningful sourcetype makes grouping and filtering much more useful.
Attributes are preserved¶
Fields that do not map to a core column go into attributes:
You can choose promoted attributes after you start collecting data. If an
attribute becomes common in filters, promote
it. Existing queries keep
working, because the query layer rewrites attr_get() onto the promoted
column.
Send directly from an SDK¶
You can skip the Collector, but then the SDK must handle batching and retries. It also loses the Collector's enrichment processors. Export these variables:
export OTEL_EXPORTER_OTLP_LOGS_ENDPOINT=http://siglake-ingester:8088/v1/logs
export OTEL_EXPORTER_OTLP_TRACES_ENDPOINT=http://siglake-ingester:8088/v1/traces
export OTEL_EXPORTER_OTLP_PROTOCOL=http/protobuf
export OTEL_RESOURCE_ATTRIBUTES=service.name=checkout,host.name=web-01
Make sure your SDK's exporter retries on 429/503.
The Collector and SDK examples above omit custom tenant routing. They write to
the default tenant, which is also what the query server reads by default.
Custom tenant variant¶
To route this telemetry to team-platform on a shared deployment, configure
OIDC tenant routing on the ingester and query server:
siglake ingest-server \
--oidc-issuer https://id.example.com \
--oidc-audience siglake \
--oidc-tenant-claim tenant
siglake-query-server \
--oidc-issuer https://id.example.com \
--oidc-audience siglake \
--oidc-tenant-claim tenant
Both the exporter and query client need valid bearer JWTs for that issuer and
audience with tenant: team-platform. The Collector variant may confirm the
claim with the ingest header; a different header value is refused:
exporters:
otlphttp/siglake:
headers:
X-Scope-OrgID: team-platform
Authorization: "Bearer ${env:SIGLAKE_TOKEN}"
Pass a JWT with the same claim to queries, for example with
siglake sql --token "$SIGLAKE_TOKEN". Adding X-Scope-OrgID (or
another ingest-style tenant header) to a query is not a substitute for
query-side tenant selection: queries use the verified claim only when the
query server has --oidc-tenant-claim; otherwise they use default.
See Multi-tenancy and Security for routing and credential details.
Configure OTLP over gRPC¶
OTLP over gRPC listens on 0.0.0.0:4317 by default. It accepts logs and traces
with the same authentication, tenant routing and WAL-fsync acknowledgement
contract as OTLP over HTTP.
To turn off the listener in a Helm deployment, set:
For a non-Helm run, pass the disable flag:
Change the default port with ingester.otlpGrpc.port in Helm or
--otlp-grpc-listen on the command line.
The listener serves both the OTLP logs and trace services, so the Collector
side is one otlp exporter rather than the two endpoints the HTTP exporter
takes:
exporters:
otlp/siglake:
endpoint: siglake-ingester.siglake.svc:4317
tls:
insecure: true
headers:
authorization: "Bearer ${env:SIGLAKE_TOKEN}"
Name otlp/siglake instead of otlphttp/siglake in the logs and traces
pipelines. Keep sending_queue and retry_on_failure as configured above.
The same five-minute retry limit and bounded, non-persistent queue apply to
gRPC exports.
HTTP is the better-tested path.
Traces¶
POST /v1/traces accepts OTLP spans, queryable through the Jaeger
shim or directly with SQL against the traces index.
Set the batch size¶
The ingester's body cap is 16 MiB (SIGLAKE_INGEST_MAX_BODY_BYTES).
Oversized batches get 413.
send_batch_max_size: 16384 records is a safe starting point for typical log
lines. If your records are large, lower it.
Larger batches amortize per-request overhead; the trade-off is latency and
memory on both sides. The ingester's own --ingest-group-commit-ms provides
similar amortization server-side.
Verify Collector export and attribute mapping¶
# is anything arriving?
curl -s localhost:9100/metrics | grep siglake_events_accepted_total
# is anything being shed?
curl -s localhost:9100/metrics | grep -E 'backpressure_rejected|rate_limit_rejected'
# did the mapping work?
siglake sql "SELECT host, source, sourcetype, index, count(*)
FROM events
WHERE timestamp >= now() - INTERVAL '5 minutes'
GROUP BY host, source, sourcetype, index"
If the last query returns host = unknown, fix the host.name resource
attribute. If every row has sourcetype = otel:logs, set the record attribute
in the transform processor.