Quickstart: a minimal Siglake stack, end to end¶
In this tutorial you get a minimal Siglake stack running locally with one command: a Postgres catalog, a MinIO warehouse, an ingester, a compactor, a query server, and a Prometheus that scrapes every role. You then send your first log record and read it back with SQL.
You need Docker with the Compose plugin, curl, and jq. The first build
compiles the Rust workspace, so give it a few gigabytes of disk and some time.
1. Get the stack running locally¶
git clone https://github.com/siglake/siglake.git
cd siglake
docker compose -f deploy/docker-compose.yml up --build
That command stays in the foreground. Later starts reuse the built image and come up in seconds.
scripts/up.sh runs the same compose file detached, waits for the ingester's
health check, and prints the connection details. scripts/down.sh tears
everything down, volumes included.
In a second terminal, wait until the ingester reports healthy:
2. Send your first log record¶
Siglake speaks OTLP/HTTP. POST /v1/logs accepts both the JSON and the
protobuf encoding, chosen by Content-Type:
curl -s http://localhost:8088/v1/logs \
-H 'Content-Type: application/json' \
-d '{
"resourceLogs": [{
"resource": {
"attributes": [
{"key": "service.name", "value": {"stringValue": "quickstart"}}
]
},
"scopeLogs": [{
"logRecords": [{
"timeUnixNano": "'"$(date +%s)000000000"'",
"body": {"stringValue": "hello siglake"}
}]
}]
}]
}'
What the acknowledgement means¶
With no commit option, the request uses wait_for. Siglake acknowledges the
record after fsync(2) syncs the local WAL. The sync covers segment bytes and
directory entries for the tenant, index, active segment, and sealed segment.
The power-loss guarantee assumes ext4 or xfs on a
node-attached volume. A network filesystem provides whatever its fsync(2)
and rename semantics guarantee.
This acknowledgement does not wait for an object-store mirror, an Iceberg
commit, or query visibility. To acknowledge after write(2) while bytes can
remain in the kernel page cache, send ?commit=auto. Ingest and the
WAL has the full durability model.
3. Query the record end to end¶
curl -s http://localhost:8089/api/v1/sql \
-H 'Content-Type: application/json' \
-d '{"query": "SELECT timestamp, host, raw FROM events ORDER BY timestamp DESC LIMIT 10"}' \
| jq
Your record comes back within a few seconds, before the compactor has committed it to Iceberg. The query server reads sealed-but-uncommitted WAL segments from its real-time buffer and unions them with the committed Iceberg snapshot under the same table name. See Freshness.
host comes back as unknown, because the payload set service.name and no
host.name. Sending real data covers the whole mapping.
If the row is not there yet, wait and retry. A WAL segment seals after 4096 events or 5 seconds, whichever comes first, so a single record can take up to about 5 s to appear. Under sustained load segments seal on size instead, and freshness is dominated by the seal interval rather than the commit cycle.
4. Use the interactive client¶
The binary that runs the servers is also a SQL client. Give it a query for a single answer, or no query for a REPL:
docker compose -f deploy/docker-compose.yml exec ingester \
siglake sql --endpoint http://query-server:8089 \
"SELECT count(*) AS n FROM events"
docker compose -f deploy/docker-compose.yml exec ingester \
siglake sql --endpoint http://query-server:8089
Table output carries the per-query scan statistics and the server-side time.
Add --dry-run to get the pre-flight cost estimate without executing anything.
5. Drive some load¶
To watch the system under sustained ingest, with WAL segments rolling, the drain committing and compaction running, drive load from the image the stack already built:
docker compose -f deploy/docker-compose.yml exec ingester \
siglake-loadgen --target http://localhost:8088 --eps 5000 --duration 60s
The generator prints progress every 5 s and a summary at the end: events sent, achieved rate, error and status-code counts, and latency percentiles from an HDR histogram.
scripts/loadgen.sh wraps the same generator with a concurrent SQL query
worker, a drain wait, and a wider telemetry snapshot: query latencies,
Prometheus counters, and the final Iceberg row count. It runs
target/release/siglake-loadgen on the host rather than in a container, so
build that binary first:
Watch the compactor work:
Prometheus is at http://localhost:9090. The MinIO console is at
http://localhost:9001 (minioadmin / minioadmin), where you can watch
Parquet files land in the warehouse bucket.
6. Tear down¶
The -v removes the volumes: the WAL, the warehouse bucket and the catalog
database. Omit it to keep your data between runs.
Where to go next¶
- Your first queries: the
eventsschema and the SQL surface in detail. - Sending real data: wire up an OpenTelemetry Collector or an Elasticsearch bulk client.
- Concepts: how the pieces work.
- Operations: deploying this for real.