Siglake¶
Siglake is a horizontally-scalable, OTLP-native log analytics platform written in Rust, built on Parquet v2, Apache Iceberg and DataFusion.
Logs and traces go in over OTLP or an Elasticsearch-compatible bulk API. Every row lands as open Parquet in an Iceberg catalog on object storage. Siglake's own distributed SQL tier reads it, and so do external Iceberg readers with Siglake out of the path.
-
The whole stack in one
docker composecommand. Ingest a record, query it back, in about five minutes. -
How the ingest path, WAL, Iceberg storage, compaction, and distributed query tier fit together.
-
HTTP API, CLI, configuration, table schemas, and Prometheus metrics.
-
Helm, the Kubernetes operator, Terraform/EKS, scaling, and monitoring.
Design choices¶
The write-ahead log is the streaming layer¶
Siglake needs no Kafka or Flink. The compactor and the query freshness buffer consume sealed write-ahead log (WAL) segments directly, and a supported consumer interface lets separate processes do the same with durable cursors and retention watermarks.
The warehouse is plain Iceberg¶
Siglake writes ordinary Iceberg tables of ordinary Parquet files. The
format-version-2 event-time contract uses a microsecond timestamptz plus an
exact timestamp_ns long sibling. Trino 483, Spark 3.5.9, DuckDB 1.5.5 and
PyIceberg 0.12.0 each read the measured local fixture without Siglake in the
path. See External query engines for the pinned
libraries and the catalog caveats.
Search acceleration lives in the files¶
Trigram and token blooms, inverted-index blobs, group-count footers and time-bucket footers all ride inside the Parquet files and their Puffin sidecars. Any reader that knows how to look can prune with them, and compaction rebuilds them as data consolidates.
Fresh data is queryable before it is committed¶
The query tier reads sealed but uncommitted WAL segments and unions them with Iceberg under the same table name, de-overlapped by commit-stamped consumed-segment lists. Measured ~5.5 s ingest to queryable at full ingest rate.
The event stream is open to external consumers¶
Separate processes can read sealed WAL segments before commit through the supported consumer interface. Consume WAL segments explains the read loop, retention window, and recovery after a segment is swept.
Architecture at a glance¶
┌───────────────────────── ingest tier ─────────────────────────┐
OTLP /v1/logs ──►│ ingester ─► WAL (Arrow IPC segments) │
Elastic _bulk ──►│ active/ → sealed/ → processing/ → committed/ │
OTLP /v1/traces ►│ backpressure lanes, per-tenant rate budgets │
└──────────┬──────────────────────────────┬─────────────────────┘
│ drain (continuous, │ SegmentConsumer
│ N commits in flight) ▼
▼ external consumers
Iceberg on object storage
┌────────────────────────────┐
│ events + user indexes │
│ Parquet v2, time-ordered, │
│ day-partitioned, blooms + │
│ inverted-index sidecars │
│ leveled compaction │
└─────────────┬──────────────┘
│
┌─────────────▼──────────────┐
│ query tier (2+ replicas) │ DataFusion SQL, transparent distributed
│ /api/v1/sql coordinator │ coordination, aggregate fast paths,
│ + WAL real-time buffer │ ordered early-stop, cost guardrails
└────────────────────────────┘
Read the full walkthrough in Concepts.
Project status¶
The current published release is v0.2.1. These latest docs also cover the 0.3.0 development line. Select 0.2.1 in the version menu for the released contracts, or read the changelog for the differences.
The core platform provides ingestion, a durable WAL, open Iceberg storage, distributed SQL queries, and a WAL consumer interface. Docker Compose, Helm charts, a Kubernetes operator, and Terraform/EKS support local evaluation and cloud deployment.
It has been validated at 200 GB (394 M rows) and 1 TB (2.0 B rows) scales across roughly 100 AWS benchmark rounds. A four-ingester fleet sustained ~415 K rows/s end to end, and exact row counts held through node crashes.
Product scope¶
Siglake is the telemetry storage and query layer. Its APIs, SQL interface, open tables, and WAL consumer interface let you build on it with your own tools. Use Grafana and Jaeger for dashboards and trace exploration, and an external alerting system for evaluation, notifications, and on-call workflows. These interfaces are intentional integration boundaries of the core product.
Log exploration uses time ordering and SQL filters. Deployment options such as the WAL freshness buffer and autoscaling are configurable opt-ins. Scope and limitations explains those defaults and the operational constraints to account for when sizing and configuring a deployment.
License¶
Apache License 2.0. The vendored Iceberg forks under third_party/ retain
their upstream Apache-2.0 LICENSE and NOTICE files.