Skip to content

About Siglake

Siglake is a horizontally-scalable, OTLP-native log analytics platform built on Parquet v2, Apache Iceberg, and DataFusion, written in Rust. It is developed by Limnion AI and licensed under the Apache License 2.0.

Page Contents
Performance Published measurements, and the methodology behind them.
Scope and limitations Integration boundaries, configuration choices, and operational constraints.
Contributing Building, testing, and submitting changes.
Changelog Release history.

Status

The current published release is v0.2.1. The latest docs cover ongoing 0.3.0 development as well as released features. Use the version menu to select the 0.2.1 documentation, and the changelog to compare release behavior.

Siglake provides OTLP and bulk ingestion, a durable WAL, Iceberg storage, distributed SQL queries, and an interface for external WAL consumers. Deployment paths include Docker Compose, Helm, the Kubernetes operator, and Terraform/EKS.

It has been validated at 200 GB (394 M rows) and 1 TB (2.02 B rows) across roughly 100 AWS benchmark rounds. A four-ingester fleet sustained ~415 K rows/s end to end, and exact row counts held through node crashes.

The core product focuses on telemetry storage and query. Dashboards, alert evaluation, delivery, and on-call workflows integrate through external tools; see Grafana and Jaeger. Search uses time ordering and SQL filters. Scope and limitations documents these design choices alongside configuration defaults and operational limits.

History

The project was called knulps until 2026-06-12. Pre-rename docs are kept in the internal history, and older git history uses the old name. It is the same system.

The build was phased, each phase shipping compiling, clippy-clean, tested code:

Phase Scope
1 to 4 Storage engine, ingest, query, scale-out, multi-tenancy, hardening: backpressure, catalog claims, audit, retention, GDPR deletes.
5 The former in-tree detector pipeline, retired on 2026-08-29.
2026-06/07 The performance arc: time-ordered storage, distributed query, ordered early-stop, leveled compaction with graded backpressure, deferred indexing, continuous-dispatch drain.

The public design record lives in the source repo under docs/ as DESIGN_*.md. The round-by-round AWS validation reports behind the performance numbers are internal, not part of the public source tree. Public supporting material also includes the OTLP ingest performance note. Published comparisons and their methodology are available on the benchmarks site. Each result should be read with its workload, configuration, and measurement context.

Design principles

Six commitments show up repeatedly in the code.

Storage stays open. The warehouse is plain Iceberg on Parquet, so any Iceberg reader can query it directly. There is no proprietary format and no second copy.

Accelerators degrade to a scan. Every accelerator is optional by construction. A stale side aggregate, an unrecognized footer version, an unreadable bloom: each falls back to a scan. Pruning artifacts fail closed, because a false negative silently drops rows from a result.

Every persisted format carries a version, in the name the artifact is stored under (siglake.group_counts.v1), in the payload bytes, or both. A reader that does not recognize a version ignores the artifact rather than guessing. Files are never rewritten for format reasons alone; tables converge as compaction rewrites them for its own reasons.

Caches are snapshot-keyed, never TTL-expired. A cache entry must be a pure function of (table, snapshot, query) and invalidate on commit. A TTL'd whole-result cache silently serves stale leading-edge answers, and is prohibited.

Measurements carry their context: scale, topology, and warm or cold state. When a benchmark round turned out to be measuring a result cache rather than the engine, that was recorded as a methodology failure and a cache-bypass switch was added. See Performance.

Scope and operating contracts are documented. Scope and limitations distinguishes integration boundaries from implementation constraints and tracks the code that changes them.

License

Apache License 2.0. See LICENSE in the source repository.

The vendored Iceberg forks under third_party/ retain their upstream Apache-2.0 LICENSE and NOTICE files. See third_party/README.md and NOTICE.