Skip to content

Local development

This guide covers the development tasks you repeat: building the workspace, running the gate, running the stack, and iterating in a single process. You need rustup, Docker with the Compose plugin, and git.

To reproduce a bug against a source build, four commands take a clean checkout to a running stack:

cargo build --workspace
export PATH="$PWD/target/debug:$PATH"
scripts/ci-local.sh
scripts/up.sh

Run the gate before you suspect your own change: it is the same set of jobs CI runs, and it keeps every job's full output in a directory it names on the last line. Run the test gate covers the jobs and the modes, and Run the stack locally covers load generation and the smoke check. For a reproduction you intend to file, include the full siglake --version output, which names the source commit of the build, and the failing job's log path.

Build from source

Siglake is a Rust workspace. rust-toolchain.toml pins the toolchain and rustup picks it up automatically.

git clone https://github.com/siglake/siglake.git
cd siglake
cargo build --workspace

For anything performance-related, build in release mode:

cargo build --workspace --release

Three of the binaries it produces are the ones you run by hand:

  • siglake (crates/siglake-cli): every server role, plus the ops commands and the SQL client.
  • siglake-query-server (crates/siglake-query-server): the query tier.
  • siglake-loadgen (crates/siglake-loadgen): the OTLP load generator that scripts/loadgen.sh drives.

Run the binaries you just built

Cargo writes the binaries under target/ and does not install them, so a fresh checkout leaves you with no siglake on your PATH. The examples on this page call the binaries by name. From the checkout root, put the profile you built at the front of PATH:

export PATH="$PWD/target/debug:$PATH"

If you built in release mode, use target/release instead.

Check which binary the shell finds now:

command -v siglake

The path it prints is under your checkout, in the profile directory you exported. If it prints /usr/local/bin/siglake or a path under ~/.cargo/bin, the export did not take in this shell and you are about to run an older installed copy: check that you ran it from the checkout root, then run it again.

The export applies to the current shell only, and it changes nothing outside it. Each new terminal repeats both steps: change into the checkout, then export PATH.

Run the test gate

Every change is expected to pass the local gate before review:

scripts/ci-local.sh

The default gate runs thirteen jobs: build-env, fmt, shell, claude-md, set-var, dashboard, test, clippy, helm, public-tree, generated, deny and fork-tests. The generated job checks that the committed OpenAPI specifications and both copies of the CRDs match the code. The fork-tests job runs the vendored Iceberg forks' own unit tests, which a workspace test run never compiles, through scripts/check-fork-tests.sh. The dashboard job checks that every siglake_* series a Grafana panel queries is one the code emits, and that every deployed SIGLAKE_* environment variable is named by non-test Rust source. The helm job lints both charts and renders the siglake chart across its whole value matrix, confirming that each install guard still refuses the configuration it exists for.

A job whose tool is missing (helm, promtool, cargo-deny) reports skipped and does not change the exit code.

Every job's full output goes to a directory this run alone owns, under target/ci-local/, and it survives the exit whether the run was red or green. The summary line names the directory, a red line prints the failing job's log path, and target/ci-local/latest points at the newest run. Quote that log in a bug report rather than the summary line.

Use the extended gate for nightly and pre-release checks, and strict mode when a job this machine cannot run should fail rather than be skipped:

scripts/ci-local.sh --all
scripts/ci-local.sh --strict

The extended gate adds three heavy jobs: operator-cluster, which needs kind and kubectl; docker, which builds both images and runs the MinIO and Postgres integration tests; and external-readers.

During development, three commands cover most of it:

cargo test --workspace
cargo clippy --workspace --all-targets -- -D warnings
cargo fmt --all -- --check

Most tests are hermetic: they reach neither the network nor a cloud account. A plain workspace test run needs no external services, and tests that want object storage use in-memory or file:// backends.

One workspace lint is deny, deliberately

clippy::await_holding_lock is denied workspace-wide. A std or parking_lot guard held across an .await serializes the executor and can deadlock under cancellation, so it is always a bug in this codebase. tokio::sync::Mutex is designed to be held across .await and is exempt.

Run the stack locally

The Docker Compose stack from the quickstart is the normal development environment:

cargo build --release -p siglake-loadgen
scripts/up.sh
scripts/loadgen.sh --eps 5000 --duration 60s
scripts/smoke.sh
scripts/down.sh

up.sh builds and starts everything, then waits for the ingester's health check. loadgen.sh drives ingest and concurrent SQL, then prints a telemetry snapshot; it runs target/release/siglake-loadgen on the host, which is why the release build comes first. smoke.sh posts a batch, waits for the compactor, and checks that the row count in Iceberg matches what it sent. down.sh tears the stack down and removes the volumes. To keep your data between runs:

scripts/down.sh --keep-volumes

To run the chart in a local Kubernetes cluster instead:

scripts/kind-up.sh
scripts/kind-smoke.sh
scripts/kind-round.sh
scripts/kind-down.sh

scripts/kind-round.sh is the monitoring evidence round. It installs pinned kube-prometheus-stack and KEDA charts, enables the chart's ServiceMonitors, PrometheusRules and ScaledObjects, drives five and a half minutes of load through the benchmark SQL shapes, and prints the alerted counters, dashboard panel evidence and KEDA scaling state. It deletes the cluster on exit unless you set KEEP=1.

Iterate in a single process

For fast iteration you do not need the whole stack. The ingester can run the compactor in-process, against a local filesystem warehouse and a SQLite catalog.

Run both terminals from the checkout root with PATH set as above. The default --data-dir is ./data, relative to the working directory, so the shared working directory is what makes both terminals read and write the same data.

In the first terminal, run ingest and compaction together:

siglake ingest-server --with-compactor

In a second terminal, change into the same checkout, export the same profile, then send something:

cd <siglake-checkout>
export PATH="$PWD/target/debug:$PATH"
curl -s http://localhost:8088/v1/logs \
  -H 'Content-Type: application/json' \
  -d '{"resourceLogs":[{"scopeLogs":[{"logRecords":[{"body":{"stringValue":"hi"}}]}]}]}'

Once the first terminal reports the Iceberg commit, query the warehouse offline:

siglake sql-direct --query "SELECT count(*) FROM events"

With no --warehouse-url, the warehouse is a directory under --data-dir (default ./data) and the catalog is a SQLite file alongside it.

The HTTP response only acknowledges the WAL write. sql-direct reads the committed Iceberg snapshot and does not read the query server's WAL buffer, so wait for the in-process compactor to seal and commit the segment before expecting the offline query to see the record.

siglake sql-direct runs DataFusion against a warehouse with no server in the path, which is what makes it useful for debugging and ops. siglake sql is the client for a running query server, and it is the one that exercises the real query path: fast paths, caches, WAL buffer, distributed coordination.

Generate test data

These utilities are separate from the running ingest server. siglake gen emits NDJSON, and siglake ingest reads that file into the local Iceberg-backed data directory without sending the records through the server's HTTP and WAL path.

siglake gen --n 10000 > events.ndjson
siglake ingest --input events.ndjson
siglake iceberg-demo --n 1000

iceberg-demo appends synthetic events and runs a few canned queries against the resulting snapshot; it persists between runs unless you pass --reset.

For sustained load, siglake-loadgen generates OTLP traffic at a configurable rate, worker count and batch size, and reports latency percentiles from an HDR histogram.

Repo layout

crates/
  siglake-core          events, schemas, doc mappings, sharding
  siglake-ingest        OTLP/bulk decode → WAL writes, backpressure
  siglake-wal           segment format (Arrow IPC + CRC framing), lifecycle
  siglake-storage       Iceberg integration: writes, compaction, aggregates,
                        indexes, GC, retention, delete tasks
  siglake-compactor     drain loop + maintenance scheduling
  siglake-query-server  distributed SQL, fast paths, WAL buffer,
                        Elastic/Jaeger shims, jobs, audit
  siglake-index         inverted-index build/serve
  siglake-bloom         trigram/token blooms, group-count footers
  siglake-cli           the `siglake` binary (all roles + ops commands)
  siglake-operator      Kubernetes operator
  siglake-loadgen       load generator
  siglake-openapi       emits the committed OpenAPI 3.1 specifications
third_party/
  iceberg               vendored Apache Iceberg fork
  iceberg-catalog-sql   vendored catalog fork
  iceberg-storage-opendal
                        vendored object-store writer fork
deploy/                 Helm charts, Dockerfile, Terraform, Grafana
docs/                   DESIGN_*.md records, CONSUMING_SEGMENTS.md,
                        SOAK_CORPUS.md, PERF_OTEL_INGEST_2026-06-07.md, api/
scripts/                local + kind dev environments

Crate map covers what each crate owns and how they depend on each other.

The performance numbers come from internal AWS validation rounds whose reports are not part of this public tree. Public supporting material includes the source repository's DESIGN_*.md records, its OTLP ingest performance note, and, once public, the benchmarks repository.

The vendored Iceberg forks

third_party/iceberg, third_party/iceberg-catalog-sql and third_party/iceberg-storage-opendal are first-class forks of Apache Iceberg 0.9.1 crates, wired in through [patch.crates-io]. Siglake owns the Parquet read path, the commit path and the object-store write path there, and the forks carry:

  • rewrite_files (atomic replace) as a transaction action;
  • count- and age-based expire_snapshots;
  • commit-reload elision and commit-attempt observability;
  • S3 conditional-put primitives;
  • the incremental append scan used by table subscriptions, closed as not planned upstream;
  • configurable and observable multipart upload concurrency, chunk sizing, and separate S3 write permits for drain and compaction.

They are rebased against upstream periodically, and feature work does not block on upstream releases.

Contributing

See Contributing for the PR process, DCO sign-off, and coding conventions.