Skip to content

OpenTelemetry

Use the OpenTelemetry Collector or an OpenTelemetry Protocol (OTLP) exporter to send logs and traces to Siglake. The HTTP endpoints accept protobuf or JSON without a Siglake plugin.

This page is about telemetry arriving at Siglake. For Siglake exporting its own logs and traces to your backend, see Siglake's own telemetry.

Configure an OpenTelemetry Collector exporter

Before you start, make the ingest service reachable from the Collector. If ingest authentication is enabled, set SIGLAKE_TOKEN to a valid bearer token. Then add this pipeline to the Collector configuration:

receivers:
  otlp:
    protocols:
      grpc:
        endpoint: 0.0.0.0:4317
      http:
        endpoint: 0.0.0.0:4318
  filelog:
    include: [/var/log/containers/*.log]
    operators:
      - type: container

processors:
  batch:
    timeout: 5s
    send_batch_size: 8192
    send_batch_max_size: 16384      # keep under the 16 MiB body cap
  k8sattributes:
  resource:
    attributes:
      - key: host.name
        from_attribute: k8s.node.name
        action: insert
  transform:
    log_statements:
      - context: log
        statements:
          - set(attributes["sourcetype"], "nginx:access")
          - set(attributes["index"], "web")

exporters:
  otlphttp/siglake:
    logs_endpoint: http://siglake-ingester.siglake.svc:8088/v1/logs
    traces_endpoint: http://siglake-ingester.siglake.svc:8088/v1/traces
    encoding: proto
    headers:
      Authorization: "Bearer ${env:SIGLAKE_TOKEN}"
    sending_queue:
      enabled: true
      queue_size: 10000
    retry_on_failure:
      enabled: true
      initial_interval: 1s
      max_interval: 30s
      max_elapsed_time: 5m

service:
  pipelines:
    logs:
      receivers: [filelog, otlp]
      processors: [k8sattributes, resource, transform, batch]
      exporters: [otlphttp/siglake]
    traces:
      receivers: [otlp]
      processors: [k8sattributes, resource, batch]
      exporters: [otlphttp/siglake]

Retries do not guarantee delivery

Siglake sheds load with 429 and 503 plus Retry-After under backpressure. The Collector retries those responses, but retries expire after five minutes and the bounded queue can fill. This configuration has no persistent queue storage, so a Collector restart can also lose queued telemetry.

These limits apply before Siglake acknowledges an export over HTTP or gRPC. With the default wait_for mode, a Siglake acknowledgement means the rows are fsynced to its local write-ahead log (WAL). See Ingest and the WAL for the later durability boundaries.

Map Collector attributes into the events schema

Check how OTLP fields land in the events schema before you send production traffic:

events column Source Fallback
timestamp timeUnixNano observedTimeUnixNano, then ingest wall-clock
timestamp_ns timeUnixNano observedTimeUnixNano, then ingest wall-clock
host resource attribute host.name "unknown"
source resource attribute service.name scope name, then "otel"
sourcetype Record attribute sourcetype "otel:logs"
index Record attribute index "main"
raw log record body ""
attributes everything else null

Both time columns resolve the same instant: a nonzero timeUnixNano, else a nonzero observedTimeUnixNano, else the ingester's wall clock. timestamp stores that instant to microsecond precision and remains the day() partition and first sort column. timestamp_ns stores its exact Unix nanoseconds and is the second sort column, breaking ties within a microsecond.

host.name must be a resource attribute. If it is missing, every event gets host = "unknown". In the configuration above, the resource processor copies the node name after k8sattributes enriches the resource.

sourcetype and index are record attributes, not resource attributes. The transform processor in the configuration above sets them per log record. When you add it to an existing configuration, merge transform into the existing processors mapping. Keep the other processor definitions. Add it to the logs pipeline after resource and before batch. Merge the logs.processors list into the existing service mapping, and keep the traces pipeline unchanged:

processors:
  transform:
    log_statements:
      - context: log
        statements:
          - set(attributes["sourcetype"], "nginx:access")
          - set(attributes["index"], "web")

service:
  pipelines:
    logs:
      processors: [k8sattributes, resource, transform, batch]

A meaningful sourcetype makes grouping and filtering much more useful.

Attributes are preserved

Fields that do not map to a core column go into attributes:

SELECT attr_get(attributes, 'k8s.namespace') AS ns, count(*)
FROM events GROUP BY ns;

You can choose promoted attributes after you start collecting data. If an attribute becomes common in filters, promote it. Existing queries keep working, because the query layer rewrites attr_get() onto the promoted column.

Send directly from an SDK

You can skip the Collector, but then the SDK must handle batching and retries. It also loses the Collector's enrichment processors. Export these variables:

export OTEL_EXPORTER_OTLP_LOGS_ENDPOINT=http://siglake-ingester:8088/v1/logs
export OTEL_EXPORTER_OTLP_TRACES_ENDPOINT=http://siglake-ingester:8088/v1/traces
export OTEL_EXPORTER_OTLP_PROTOCOL=http/protobuf
export OTEL_RESOURCE_ATTRIBUTES=service.name=checkout,host.name=web-01

Make sure your SDK's exporter retries on 429/503.

The Collector and SDK examples above omit custom tenant routing. They write to the default tenant, which is also what the query server reads by default.

Custom tenant variant

To route this telemetry to team-platform on a shared deployment, configure OIDC tenant routing on the ingester and query server:

siglake ingest-server \
  --oidc-issuer https://id.example.com \
  --oidc-audience siglake \
  --oidc-tenant-claim tenant

siglake-query-server \
  --oidc-issuer https://id.example.com \
  --oidc-audience siglake \
  --oidc-tenant-claim tenant

Both the exporter and query client need valid bearer JWTs for that issuer and audience with tenant: team-platform. The Collector variant may confirm the claim with the ingest header; a different header value is refused:

exporters:
  otlphttp/siglake:
    headers:
      X-Scope-OrgID: team-platform
      Authorization: "Bearer ${env:SIGLAKE_TOKEN}"

Pass a JWT with the same claim to queries, for example with siglake sql --token "$SIGLAKE_TOKEN". Adding X-Scope-OrgID (or another ingest-style tenant header) to a query is not a substitute for query-side tenant selection: queries use the verified claim only when the query server has --oidc-tenant-claim; otherwise they use default.

See Multi-tenancy and Security for routing and credential details.

Configure OTLP over gRPC

OTLP over gRPC listens on 0.0.0.0:4317 by default. It accepts logs and traces with the same authentication, tenant routing and WAL-fsync acknowledgement contract as OTLP over HTTP.

To turn off the listener in a Helm deployment, set:

ingester:
  otlpGrpc:
    enabled: false

For a non-Helm run, pass the disable flag:

siglake ingest-server --disable-otlp-grpc

Change the default port with ingester.otlpGrpc.port in Helm or --otlp-grpc-listen on the command line.

The listener serves both the OTLP logs and trace services, so the Collector side is one otlp exporter rather than the two endpoints the HTTP exporter takes:

exporters:
  otlp/siglake:
    endpoint: siglake-ingester.siglake.svc:4317
    tls:
      insecure: true
    headers:
      authorization: "Bearer ${env:SIGLAKE_TOKEN}"

Name otlp/siglake instead of otlphttp/siglake in the logs and traces pipelines. Keep sending_queue and retry_on_failure as configured above. The same five-minute retry limit and bounded, non-persistent queue apply to gRPC exports.

HTTP is the better-tested path.

Traces

POST /v1/traces accepts OTLP spans, queryable through the Jaeger shim or directly with SQL against the traces index.

Set the batch size

The ingester's body cap is 16 MiB (SIGLAKE_INGEST_MAX_BODY_BYTES). Oversized batches get 413.

send_batch_max_size: 16384 records is a safe starting point for typical log lines. If your records are large, lower it.

Larger batches amortize per-request overhead; the trade-off is latency and memory on both sides. The ingester's own --ingest-group-commit-ms provides similar amortization server-side.

Verify Collector export and attribute mapping

# is anything arriving?
curl -s localhost:9100/metrics | grep siglake_events_accepted_total

# is anything being shed?
curl -s localhost:9100/metrics | grep -E 'backpressure_rejected|rate_limit_rejected'

# did the mapping work?
siglake sql "SELECT host, source, sourcetype, index, count(*)
             FROM events
             WHERE timestamp >= now() - INTERVAL '5 minutes'
             GROUP BY host, source, sourcetype, index"

If the last query returns host = unknown, fix the host.name resource attribute. If every row has sourcetype = otel:logs, set the record attribute in the transform processor.