Skip to content

Siglake

Siglake is a horizontally-scalable, OTLP-native log analytics platform written in Rust, built on Parquet v2, Apache Iceberg and DataFusion.

Logs and traces go in over OTLP or an Elasticsearch-compatible bulk API. Every row lands as open Parquet in an Iceberg catalog on object storage. Siglake's own distributed SQL tier reads it, and so do external Iceberg readers with Siglake out of the path.

  • Quickstart

    The whole stack in one docker compose command. Ingest a record, query it back, in about five minutes.

  • Concepts

    How the ingest path, WAL, Iceberg storage, compaction, and distributed query tier fit together.

  • Reference

    HTTP API, CLI, configuration, table schemas, and Prometheus metrics.

  • Operations

    Helm, the Kubernetes operator, Terraform/EKS, scaling, and monitoring.

Design choices

The write-ahead log is the streaming layer

Siglake needs no Kafka or Flink. The compactor and the query freshness buffer consume sealed write-ahead log (WAL) segments directly, and a supported consumer interface lets separate processes do the same with durable cursors and retention watermarks.

The warehouse is plain Iceberg

Siglake writes ordinary Iceberg tables of ordinary Parquet files. The format-version-2 event-time contract uses a microsecond timestamptz plus an exact timestamp_ns long sibling. Trino 483, Spark 3.5.9, DuckDB 1.5.5 and PyIceberg 0.12.0 each read the measured local fixture without Siglake in the path. See External query engines for the pinned libraries and the catalog caveats.

Search acceleration lives in the files

Trigram and token blooms, inverted-index blobs, group-count footers and time-bucket footers all ride inside the Parquet files and their Puffin sidecars. Any reader that knows how to look can prune with them, and compaction rebuilds them as data consolidates.

Fresh data is queryable before it is committed

The query tier reads sealed but uncommitted WAL segments and unions them with Iceberg under the same table name, de-overlapped by commit-stamped consumed-segment lists. Measured ~5.5 s ingest to queryable at full ingest rate.

The event stream is open to external consumers

Separate processes can read sealed WAL segments before commit through the supported consumer interface. Consume WAL segments explains the read loop, retention window, and recovery after a segment is swept.

Architecture at a glance

                 ┌───────────────────────── ingest tier ─────────────────────────┐
OTLP /v1/logs ──►│  ingester ─► WAL (Arrow IPC segments)                          │
Elastic _bulk ──►│             active/ → sealed/ → processing/ → committed/       │
OTLP /v1/traces ►│             backpressure lanes, per-tenant rate budgets        │
                 └──────────┬──────────────────────────────┬─────────────────────┘
                            │ drain (continuous,           │ SegmentConsumer
                            │ N commits in flight)         ▼
                            ▼                       external consumers
             Iceberg on object storage
             ┌────────────────────────────┐
             │ events + user indexes      │
             │ Parquet v2, time-ordered,  │
             │ day-partitioned, blooms +  │
             │ inverted-index sidecars    │
             │ leveled compaction         │
             └─────────────┬──────────────┘
                           │
             ┌─────────────▼──────────────┐
             │ query tier (2+ replicas)   │  DataFusion SQL, transparent distributed
             │ /api/v1/sql coordinator    │  coordination, aggregate fast paths,
             │ + WAL real-time buffer     │  ordered early-stop, cost guardrails
             └────────────────────────────┘

Read the full walkthrough in Concepts.

Project status

This documentation covers Siglake v0.1.0. See the release list for current releases. The core platform is complete and AWS-validated end to end: ingestion, WAL, Iceberg storage, distributed SQL query, the WAL consumer interface, Helm charts, a Kubernetes operator, and Terraform/EKS deployment.

It has been validated at 200 GB (394 M rows) and 1 TB (2.0 B rows) scales across roughly 100 AWS benchmark rounds. A four-ingester fleet sustained ~415 K rows/s end to end, and exact row counts held through node crashes.

Siglake focuses on telemetry storage and query. Dashboards, alert evaluation, notifications, and on-call workflows integrate through external tools; these are intentional boundaries of the core product. See Grafana and Jaeger for dashboard and trace integrations. Search uses time ordering and SQL filters. The scope and limitations page records configuration defaults and operational constraints.

License

Apache License 2.0. The vendored Iceberg forks under third_party/ retain their upstream Apache-2.0 LICENSE and NOTICE files.