Skip to content

Siglake

Siglake is a horizontally-scalable, OTLP-native log analytics platform written in Rust, built on Parquet v2, Apache Iceberg and DataFusion.

Logs and traces go in over OTLP or an Elasticsearch-compatible bulk API. Every row lands as open Parquet in an Iceberg catalog on object storage. Siglake's own distributed SQL tier reads it, and so do external Iceberg readers with Siglake out of the path.

  • Quickstart

    The whole stack in one docker compose command. Ingest a record, query it back, in about five minutes.

  • Concepts

    How the ingest path, WAL, Iceberg storage, compaction, and distributed query tier fit together.

  • Reference

    HTTP API, CLI, configuration, table schemas, and Prometheus metrics.

  • Operations

    Helm, the Kubernetes operator, Terraform/EKS, scaling, and monitoring.

Design choices

The write-ahead log is the streaming layer

Siglake needs no Kafka or Flink. The compactor and the query freshness buffer consume sealed write-ahead log (WAL) segments directly, and a supported consumer interface lets separate processes do the same with durable cursors and retention watermarks.

The warehouse is plain Iceberg

Siglake writes ordinary Iceberg tables of ordinary Parquet files. The format-version-2 event-time contract uses a microsecond timestamptz plus an exact timestamp_ns long sibling. Trino 483, Spark 3.5.9, DuckDB 1.5.5 and PyIceberg 0.12.0 each read the measured local fixture without Siglake in the path. See External query engines for the pinned libraries and the catalog caveats.

Search acceleration lives in the files

Trigram and token blooms, inverted-index blobs, group-count footers and time-bucket footers all ride inside the Parquet files and their Puffin sidecars. Any reader that knows how to look can prune with them, and compaction rebuilds them as data consolidates.

Fresh data is queryable before it is committed

The query tier reads sealed but uncommitted WAL segments and unions them with Iceberg under the same table name, de-overlapped by commit-stamped consumed-segment lists. Measured ~5.5 s ingest to queryable at full ingest rate.

The event stream is open to external consumers

Separate processes can read sealed WAL segments before commit through the supported consumer interface. Consume WAL segments explains the read loop, retention window, and recovery after a segment is swept.

Architecture at a glance

                 ┌───────────────────────── ingest tier ─────────────────────────┐
OTLP /v1/logs ──►│  ingester ─► WAL (Arrow IPC segments)                          │
Elastic _bulk ──►│             active/ → sealed/ → processing/ → committed/       │
OTLP /v1/traces ►│             backpressure lanes, per-tenant rate budgets        │
                 └──────────┬──────────────────────────────┬─────────────────────┘
                            │ drain (continuous,           │ SegmentConsumer
                            │ N commits in flight)         ▼
                            ▼                       external consumers
             Iceberg on object storage
             ┌────────────────────────────┐
             │ events + user indexes      │
             │ Parquet v2, time-ordered,  │
             │ day-partitioned, blooms +  │
             │ inverted-index sidecars    │
             │ leveled compaction         │
             └─────────────┬──────────────┘
                           │
             ┌─────────────▼──────────────┐
             │ query tier (2+ replicas)   │  DataFusion SQL, transparent distributed
             │ /api/v1/sql coordinator    │  coordination, aggregate fast paths,
             │ + WAL real-time buffer     │  ordered early-stop, cost guardrails
             └────────────────────────────┘

Read the full walkthrough in Concepts.

Project status

The current published release is v0.2.1. These latest docs also cover the 0.3.0 development line. Select 0.2.1 in the version menu for the released contracts, or read the changelog for the differences.

The core platform provides ingestion, a durable WAL, open Iceberg storage, distributed SQL queries, and a WAL consumer interface. Docker Compose, Helm charts, a Kubernetes operator, and Terraform/EKS support local evaluation and cloud deployment.

It has been validated at 200 GB (394 M rows) and 1 TB (2.0 B rows) scales across roughly 100 AWS benchmark rounds. A four-ingester fleet sustained ~415 K rows/s end to end, and exact row counts held through node crashes.

Product scope

Siglake is the telemetry storage and query layer. Its APIs, SQL interface, open tables, and WAL consumer interface let you build on it with your own tools. Use Grafana and Jaeger for dashboards and trace exploration, and an external alerting system for evaluation, notifications, and on-call workflows. These interfaces are intentional integration boundaries of the core product.

Log exploration uses time ordering and SQL filters. Deployment options such as the WAL freshness buffer and autoscaling are configurable opt-ins. Scope and limitations explains those defaults and the operational constraints to account for when sizing and configuring a deployment.

License

Apache License 2.0. The vendored Iceberg forks under third_party/ retain their upstream Apache-2.0 LICENSE and NOTICE files.