Local development¶
This guide covers the development tasks you repeat: building the workspace,
running the gate, running the stack, and iterating in a single process. You
need rustup, Docker with the Compose plugin, and git.
To reproduce a bug against a source build, four commands take a clean checkout to a running stack:
Run the gate before you suspect your own change: it is the same set of jobs CI
runs, and it keeps every job's full output in a directory it names on the last
line. Run the test gate covers the jobs and the modes,
and Run the stack locally covers load generation and
the smoke check. For a reproduction you intend to file, include the full
siglake --version output, which names the source commit of the build, and
the failing job's log path.
Build from source¶
Siglake is a Rust workspace. rust-toolchain.toml pins the toolchain and
rustup picks it up automatically.
For anything performance-related, build in release mode:
Three of the binaries it produces are the ones you run by hand:
siglake(crates/siglake-cli): every server role, plus the ops commands and the SQL client.siglake-query-server(crates/siglake-query-server): the query tier.siglake-loadgen(crates/siglake-loadgen): the OTLP load generator thatscripts/loadgen.shdrives.
Run the binaries you just built¶
Cargo writes the binaries under target/ and does not install them, so a fresh
checkout leaves you with no siglake on your PATH. The examples on this page
call the binaries by name. From the checkout root, put the profile you built at
the front of PATH:
If you built in release mode, use target/release instead.
Check which binary the shell finds now:
The path it prints is under your checkout, in the profile directory you
exported. If it prints /usr/local/bin/siglake or a path under ~/.cargo/bin,
the export did not take in this shell and you are about to run an older
installed copy: check that you ran it from the checkout root, then run it
again.
The export applies to the current shell only, and it changes nothing outside
it. Each new terminal repeats both steps: change into the checkout, then export
PATH.
Run the test gate¶
Every change is expected to pass the local gate before review:
The default gate runs thirteen jobs: build-env, fmt, shell, claude-md,
set-var, dashboard, test, clippy, helm, public-tree, generated,
deny and fork-tests.
The generated job checks that the committed OpenAPI specifications and both
copies of the CRDs match the code. The fork-tests job runs the vendored
Iceberg forks' own unit tests, which a workspace test run never compiles,
through scripts/check-fork-tests.sh. The dashboard job checks that every
siglake_* series a Grafana panel queries is one the code emits, and that
every deployed SIGLAKE_* environment variable is named by non-test Rust
source. The helm job lints both charts and renders the siglake chart across
its whole value matrix, confirming that each install guard still refuses the
configuration it exists for.
A job whose tool is missing (helm, promtool, cargo-deny) reports skipped and
does not change the exit code.
Every job's full output goes to a directory this run alone owns, under
target/ci-local/, and it survives the exit whether the run was red or green.
The summary line names the directory, a red line prints the failing job's log
path, and target/ci-local/latest points at the newest run. Quote that log in
a bug report rather than the summary line.
Use the extended gate for nightly and pre-release checks, and strict mode when a job this machine cannot run should fail rather than be skipped:
The extended gate adds three heavy jobs: operator-cluster, which needs kind
and kubectl; docker, which builds both images and runs the MinIO and Postgres
integration tests; and external-readers.
During development, three commands cover most of it:
cargo test --workspace
cargo clippy --workspace --all-targets -- -D warnings
cargo fmt --all -- --check
Most tests are hermetic: they reach neither the network nor a cloud account. A
plain workspace test run needs no external services, and tests that want object
storage use in-memory or file:// backends.
One workspace lint is deny, deliberately
clippy::await_holding_lock is denied workspace-wide. A std or
parking_lot guard held across an .await serializes the executor and can
deadlock under cancellation, so it is always a bug in this codebase.
tokio::sync::Mutex is designed to be held across .await and is exempt.
Run the stack locally¶
The Docker Compose stack from the quickstart is the normal development environment:
cargo build --release -p siglake-loadgen
scripts/up.sh
scripts/loadgen.sh --eps 5000 --duration 60s
scripts/smoke.sh
scripts/down.sh
up.sh builds and starts everything, then waits for the ingester's health
check. loadgen.sh drives ingest and concurrent SQL, then prints a telemetry
snapshot; it runs target/release/siglake-loadgen on the host, which is why
the release build comes first. smoke.sh posts a batch, waits for the
compactor, and checks that the row count in Iceberg matches what it sent.
down.sh tears the stack down and removes the volumes. To keep your data
between runs:
To run the chart in a local Kubernetes cluster instead:
scripts/kind-round.sh is the monitoring evidence round. It installs pinned
kube-prometheus-stack and KEDA charts, enables the chart's ServiceMonitors,
PrometheusRules and ScaledObjects, drives five and a half minutes of load
through the benchmark SQL shapes, and prints the alerted counters, dashboard
panel evidence and KEDA scaling state. It deletes the cluster on exit unless
you set KEEP=1.
Iterate in a single process¶
For fast iteration you do not need the whole stack. The ingester can run the compactor in-process, against a local filesystem warehouse and a SQLite catalog.
Run both terminals from the checkout root with PATH set as above. The default
--data-dir is ./data, relative to the working directory, so the shared
working directory is what makes both terminals read and write the same data.
In the first terminal, run ingest and compaction together:
In a second terminal, change into the same checkout, export the same profile, then send something:
cd <siglake-checkout>
export PATH="$PWD/target/debug:$PATH"
curl -s http://localhost:8088/v1/logs \
-H 'Content-Type: application/json' \
-d '{"resourceLogs":[{"scopeLogs":[{"logRecords":[{"body":{"stringValue":"hi"}}]}]}]}'
Once the first terminal reports the Iceberg commit, query the warehouse offline:
With no --warehouse-url, the warehouse is a directory under --data-dir
(default ./data) and the catalog is a SQLite file alongside it.
The HTTP response only acknowledges the WAL write. sql-direct reads the
committed Iceberg snapshot and does not read the query server's WAL buffer, so
wait for the in-process compactor to seal and commit the segment before
expecting the offline query to see the record.
siglake sql-direct runs DataFusion against a warehouse with no server in the
path, which is what makes it useful for debugging and ops. siglake sql is the
client for a running query server, and it is the one that exercises the real
query path: fast paths, caches, WAL buffer, distributed coordination.
Generate test data¶
These utilities are separate from the running ingest server. siglake gen
emits NDJSON, and siglake ingest reads that file into the local
Iceberg-backed data directory without sending the records through the server's
HTTP and WAL path.
siglake gen --n 10000 > events.ndjson
siglake ingest --input events.ndjson
siglake iceberg-demo --n 1000
iceberg-demo appends synthetic events and runs a few canned queries against
the resulting snapshot; it persists between runs unless you pass --reset.
For sustained load, siglake-loadgen generates OTLP traffic at a configurable
rate, worker count and batch size, and reports latency percentiles from an HDR
histogram.
Repo layout¶
crates/
siglake-core events, schemas, doc mappings, sharding
siglake-ingest OTLP/bulk decode → WAL writes, backpressure
siglake-wal segment format (Arrow IPC + CRC framing), lifecycle
siglake-storage Iceberg integration: writes, compaction, aggregates,
indexes, GC, retention, delete tasks
siglake-compactor drain loop + maintenance scheduling
siglake-query-server distributed SQL, fast paths, WAL buffer,
Elastic/Jaeger shims, jobs, audit
siglake-index inverted-index build/serve
siglake-bloom trigram/token blooms, group-count footers
siglake-cli the `siglake` binary (all roles + ops commands)
siglake-operator Kubernetes operator
siglake-loadgen load generator
siglake-openapi emits the committed OpenAPI 3.1 specifications
third_party/
iceberg vendored Apache Iceberg fork
iceberg-catalog-sql vendored catalog fork
iceberg-storage-opendal
vendored object-store writer fork
deploy/ Helm charts, Dockerfile, Terraform, Grafana
docs/ DESIGN_*.md records, CONSUMING_SEGMENTS.md,
SOAK_CORPUS.md, PERF_OTEL_INGEST_2026-06-07.md, api/
scripts/ local + kind dev environments
Crate map covers what each crate owns and how they depend on each other.
The performance numbers come from internal AWS validation rounds whose reports
are not part of this public tree. Public supporting material includes the
source repository's
DESIGN_*.md records, its
OTLP ingest performance note,
and, once public, the
benchmarks repository.
The vendored Iceberg forks¶
third_party/iceberg, third_party/iceberg-catalog-sql and
third_party/iceberg-storage-opendal are first-class forks of Apache Iceberg
0.9.1 crates, wired in through [patch.crates-io]. Siglake owns the Parquet
read path, the commit path and the object-store write path there, and the forks
carry:
rewrite_files(atomic replace) as a transaction action;- count- and age-based
expire_snapshots; - commit-reload elision and commit-attempt observability;
- S3 conditional-put primitives;
- the incremental append scan used by table subscriptions, closed as not planned upstream;
- configurable and observable multipart upload concurrency, chunk sizing, and separate S3 write permits for drain and compaction.
They are rebased against upstream periodically, and feature work does not block on upstream releases.
Contributing¶
See Contributing for the PR process, DCO sign-off, and coding conventions.