Skip to content

Sending real data

This guide connects a real source to Siglake: an OpenTelemetry Collector, a language SDK, or an existing Elasticsearch bulk shipper. It assumes you have an ingester reachable on port 8088, from the quickstart or from a deployment of your own.

The ingester serves four write surfaces:

Surface Endpoint Use for
OTLP/HTTP logs POST /v1/logs Anything OpenTelemetry-shaped. The primary path.
OTLP/HTTP traces POST /v1/traces Spans, queryable through the Jaeger shim.
OTLP/gRPC logs and traces port 4317 Exporters that speak gRPC, including the Collector's default OTLP exporter.
Elasticsearch bulk POST /api/v1/_elastic/_bulk Existing Filebeat, Logstash or Fluentd pipelines.

Point an OTLP exporter at the ingester

POST /v1/logs accepts both the JSON and the protobuf encoding, chosen by Content-Type. Any OTLP exporter works: the OpenTelemetry Collector, a language SDK, or an agent. Point it at http://<ingester>:8088.

A Collector configuration that reads local log files and forwards both logs and traces:

receivers:
  filelog:
    include: [/var/log/**/*.log]
  otlp:
    protocols:
      grpc:
      http:

exporters:
  otlphttp/siglake:
    logs_endpoint: http://siglake-ingester:8088/v1/logs
    traces_endpoint: http://siglake-ingester:8088/v1/traces
    encoding: proto      # or json
    headers:
      Authorization: "Bearer ${SIGLAKE_TOKEN}"   # if auth is enabled

service:
  pipelines:
    logs:
      receivers: [filelog, otlp]
      exporters: [otlphttp/siglake]
    traces:
      receivers: [otlp]
      exporters: [otlphttp/siglake]

To verify it worked, query for what you sent:

curl -s http://localhost:8089/api/v1/sql \
  -H 'Content-Type: application/json' \
  -d '{"query": "SELECT source, count(*) AS n FROM events GROUP BY source"}' | jq

Send over OTLP/gRPC

The ingester listens for OTLP/gRPC logs and traces on 0.0.0.0:4317, and that listener is on by default. Point the Collector's OTLP exporter at it. The listener speaks plaintext h2c, so an exporter that does not go through a TLS proxy needs insecure: true:

exporters:
  otlp/siglake:
    endpoint: siglake-ingester:4317
    tls:
      insecure: true
    headers:
      Authorization: "Bearer ${SIGLAKE_TOKEN}"   # if auth is enabled

Name otlp/siglake in the logs and traces pipelines where the configuration above names otlphttp/siglake. Both transports apply the same authentication, tenant routing and WAL-fsync acknowledgement rules, so moving an exporter from HTTP to gRPC changes nothing else about how a batch is accepted.

To turn the listener off under Helm, set the chart's own value for it; the chart drops the gRPC container and Service ports with it and passes the binary's opt-out flag:

ingester:
  otlpGrpc:
    enabled: false

ingester.otlpGrpc.port moves the listener to another port. For a non-Helm run, --otlp-grpc-listen (or SIGLAKE_OTLP_GRPC_LISTEN) chooses the address, and a separate opt-out flag disables the listener; leaving the address unset no longer disables it. Run siglake ingest-server --help for the opt-out flag's name.

How OTLP maps to the events schema

The mapping decides what your columns look like:

events column Sourced from Fallback
timestamp timeUnixNano observedTimeUnixNano, then ingest wall-clock
timestamp_ns timeUnixNano observedTimeUnixNano, then ingest wall-clock
host resource attribute host.name "unknown"
source resource attribute service.name instrumentation scope name, then "otel"
sourcetype record attribute sourcetype "otel:logs"
index record attribute index "main"
raw log record body (string value, else JSON-encoded) ""
attributes everything else, as a JSON object string null

Both time columns resolve the same instant: a nonzero timeUnixNano, else a nonzero observedTimeUnixNano, else the ingester's wall clock. timestamp stores that instant to microsecond precision and is both the day() partition column and the leading sort column. timestamp_ns stores the exact Unix nanoseconds and sorts second, breaking ties within a microsecond.

Set host.name if you care about host

host defaults to the literal string "unknown". Set host.name when host-level grouping and filtering matter to your queries.

Attribute capture is lossless. Any resource or log attribute that is not promoted to a core column is preserved in the attributes JSON string and stays queryable through attr_get().

The rest of the log record is not. Ingest reads four fields off a log record: timeUnixNano, observedTimeUnixNano, body and attributes. severityText, severityNumber, traceId, spanId and the dropped-attribute counts are discarded on both the JSON and the protobuf path. If you query on severity, or join logs to spans by trace id, send those values as record attributes, where they land in attributes like any other.

A complete example:

curl -s http://localhost:8088/v1/logs \
  -H 'Content-Type: application/json' \
  -d '{
    "resourceLogs": [{
      "resource": {
        "attributes": [
          {"key": "host.name",    "value": {"stringValue": "web-01"}},
          {"key": "service.name", "value": {"stringValue": "checkout"}},
          {"key": "k8s.namespace","value": {"stringValue": "prod"}}
        ]
      },
      "scopeLogs": [{
        "logRecords": [{
          "timeUnixNano": "'"$(date +%s)000000000"'",
          "body": {"stringValue": "upstream connect error: connection refused"},
          "attributes": [
            {"key": "sourcetype",       "value": {"stringValue": "nginx:error"}},
            {"key": "severity",         "value": {"stringValue": "ERROR"}},
            {"key": "http.status_code", "value": {"intValue": "502"}}
          ]
        }]
      }]
    }]
  }'

That request produces a row with host = web-01, source = checkout, sourcetype = nginx:error, index = main, and attributes = {"k8s.namespace":"prod","severity":"ERROR","http.status_code":502}, which you read back with attr_get(attributes, 'k8s.namespace').

Elasticsearch bulk API support and query-client limits

For shippers you do not want to replace, the ingester exposes an ES-compatible NDJSON bulk API that writes into user indexes:

curl -s http://localhost:8088/api/v1/_elastic/_bulk \
  -H 'Content-Type: application/x-ndjson' \
  --data-binary $'{"index":{"_index":"app-logs"}}\n{"@timestamp":"2026-07-28T12:00:00Z","message":"hello","level":"info"}\n'

POST /api/v1/_elastic/{index}/_bulk is also accepted, taking the index from the path.

Read-side ES compatibility is not implemented

Elasticsearch read routes return 501 Not Implemented and point you to POST /api/v1/sql. See Elasticsearch compatibility for details. _cluster/health always reports green. An existing Kibana or ES query client will not work against Siglake.

Select a tenant

Ingest is single-tenant by default. Every request routes to default. If X-Scope-OrgID names another tenant, Siglake returns 403 over HTTP and PermissionDenied over OpenTelemetry Protocol (OTLP)/gRPC. The examples above send no tenant header, so their rows are reachable through the default query endpoint once ingestion makes them visible.

For multiple ingest tenants, set ingester.oidc.tenantClaim to take the tenant from a verified JSON Web Token (JWT) claim. Set ingester.trustScopeHeader to trust the caller's X-Scope-OrgID header instead. Each tenant gets its own WAL subtree and Iceberg namespace, tenant_<id>; default maps to the main namespace.

With --oidc-tenant-claim, the verified JWT claim selects the tenant instead, and a missing claim or a conflicting header is refused. On a shared deployment, bind a tenant such as team-platform to that claim. Configure the same issuer, audience and claim name on both tiers:

siglake ingest-server \
  --oidc-issuer https://id.example.com \
  --oidc-audience siglake \
  --oidc-tenant-claim tenant
siglake-query-server \
  --oidc-issuer https://id.example.com \
  --oidc-audience siglake \
  --oidc-tenant-claim tenant

The ingest exporter and the query client must each send a valid bearer JWT whose aud is siglake and whose tenant claim is team-platform. You may keep sending the ingest header as an explicit confirmation of that claim:

headers:
  X-Scope-OrgID: team-platform
  Authorization: "Bearer ${SIGLAKE_TOKEN}"

Query with a JWT carrying the same tenant claim:

siglake sql --endpoint http://query-server:8089 \
  --token "$SIGLAKE_TOKEN" \
  "SELECT count(*) FROM events"

Adding X-Scope-OrgID, or any other ingest-style tenant header, to a query does not select its tenant. Query tenant selection comes from the verified OIDC claim; without --oidc-tenant-claim on the query server, every query reads the default tenant. Security covers the tenant trust boundaries.

Handle backpressure

If Siglake refuses an export, follow Tune rejected ingest to identify the limit and change the matching setting. The procedure also covers exporter retries and their loss boundary.

Tail the ingest path live

GET /api/v1/stream is a Server-Sent Events tail of the ingest path. It is push-only and best-effort: it replays nothing, promises no durability, and drops subscribers that fall behind. Events are teed before the WAL append, so a tailed event is not yet acknowledged. A batch that is then refused for backlog (503), or whose append fails (500), has still been published, and the caller's retry publishes it again. See Ingest and the WAL. For a replayable subscription use siglake subscribe, which tails committed Iceberg tables incrementally.

Next