Skip to content

Changelog

The authoritative changelog is CHANGELOG.md in the source repository. This page covers the 0.3.0 development line, 0.2.1, 0.2.0 and 0.1.0. No 0.1.1 was tagged. The 15 changes written up for it ship inside 0.2.0. The dated entries under 0.1.0 are that release's own history and are kept as they were written.

0.3.0 development line

Siglake commit 8a517387 restores the feature and telemetry work held out of the fixes-only 0.2.1 release. It does not change the packaged defaults or claim that release qualification has passed.

  • A bare interactive query over a managed index now orders by the timestamp_field in that index's mapping. Explicit ordering, batch queries and default_order: false keep their existing behavior.
  • spec.autoscaling.compactor.min: 0 is accepted with a catalog-claim drain and a positive ewmaHalfLifeSecs. The packaged floor remains 1, and a live cluster wake-up has not yet been qualified.
  • Deleting an index writes an incarnation-bound cleanup record. The cleanup sweep reports the dropped UUID's aggregate objects and bytes, but production-created records grant no delete authority.
  • Query scan log records carry process-local query_execution_id and scan_id fields, so tuning and partition events can be joined when they arrive after the request log.
  • Auto-promotion reports sampling reads, bytes and duration, plus committed backfill files, input and output bytes, and duration. Auto-promotion remains off by default.
  • WAL mirror metrics separate application-level write attempts by active and sealed path and report active-uploader bytes. Active mirroring remains off by default.
  • Puffin blob-cache gauges report resident and enforced-budget bytes. The decoded-file cache reports all bytes charged to its shared budget and counts refused populations. Existing cache defaults are unchanged.
  • The inline-coverage census reports object-store requests, usable bytes, table count and completed-pass duration. Its 900-second interval and read-only behavior are unchanged.

The HTTP API, operator guide, telemetry guide and metrics reference describe the restored contracts.

0.2.1

Published as v0.2.1 at Siglake commit 08cbeaa6d508eba1a4a6810de4002e004b8686b9 after the rename. Its corresponding pre-rename source revision was b9f77f86fe03102784e8ff2575cb21ae4bd4eb1a. This fixes-only patch preserves the 0.2.0 schema, storage format, API, configuration options and shipped defaults.

  • Delete-task ownership checks conditional-write behavior at runtime. Stores that ignore the required preconditions are refused before task execution.
  • The Helm chart forwards the WAL mirror prefix to the compactor in both drain modes and respects disabled mirroring.
  • Product prose and labels use Siglake; commands, identifiers and URLs remain lowercase siglake.

The dependency-performance, cache-depth and text-index-policy field measurements carried over from 0.2.0 remain required work for 0.3.0. This release does not claim those measurements passed. See the release-validation ledger for installation evidence and gaps.

The following work is deferred to 0.3.0 and does not ship in the 0.2.1 source:

  • compactor scale-to-zero and its ingester-published wake-up gauges;
  • implicit newest-first ordering on a mapping's non-canonical event-time field;
  • Puffin blob-cache resident and budget gauges, and the decoded-file cache's accounted-bytes gauge and population_refused outcome;
  • query execution and scan correlation records;
  • auto-promotion sampling and backfill cost series;
  • dropped-index cleanup records and the report-only cleanup sweep;
  • per-path write attempts and active-uploader bytes for the WAL mirror; and
  • object-store requests, bytes, table count and pass duration for the inline-coverage census.

Existing cache admission limits, decoded-cache outcomes, WAL mirror result counters and inline-coverage census behavior remain in 0.2.1. Deferring new telemetry does not remove those safeguards or the census itself. After the 0.2.1 tag, Siglake task #6062 must replay the deferred work onto 0.3.0 without dropping any 0.2.1 fix, test or qualification change.

0.2.0

A minor release rather than the patch it was planned as. The vendored storage stack moved up a major version and four contracts changed, so read What to do on upgrade before you roll a 0.1.0 deployment forward. The workspace version, both chart version and appVersion pairs, the pinned image tags under deploy/ and both OpenAPI documents' info.version read 0.2.0, and the v0.2.0 tag publishes image tag 0.2.0.

The vendored storage stack moved up a major version

Arrow and Parquet are at 58, DataFusion at 53.1, OpenDAL at 0.57 and the four Iceberg crates at 0.10.1. The engine reads and writes the same tables: Iceberg format version 2 on Parquet v2, with 0.1.0's timestamp contract intact. A warehouse written by 0.1.0 needs no migration for this change.

siglake wal-recover plans by default

The command prints a plan and writes nothing unless you pass --apply. The plan is the listing the restore already did before its first GET: one line per (tenant, index) with the segment count, the byte total the listing reported, a sample object name and the destination it reconstructs, plus the already-present, skipped and unreadable counts. A script or runbook that calls the old single-command form stops writing and prints a plan instead. No flag skips the plan.

Five recovery fixes travel with it. A --from URL carrying a path no longer lists <path>/<path>/, which used to print pulled 0 segments and exit 0. A candidate that does not decode to at least one row is counted unreadable and left in the mirror rather than published under a sealed name. A restore holding only index segments for a tenant now rebuilds the <tenant>/sealed/ discovery directory the drain enumerates. A run where every candidate was unreadable and nothing was already present exits nonzero. The new --catalog <uri> compares the routing each listed object name implies against the uploader's ledger rows, opening the catalog read-only. See Recover a filesystem drain.

Scan file attribution in query responses

A scanning /api/v1/sql response can carry stats.scan.file_attribution: a request-wide list of at most 32 sorted, table-relative file-task identities and their decoded-cache outcomes. files_omitted counts the identities the bound dropped, and identity_complete is false when any shard or the coordinator omitted one. A distributed worker must send a valid bounded attribution header, null included for a scan-free result. A missing, malformed or oversized header fails the shard instead of producing incomplete scan stats. See stats.scan.file_attribution fields.

Re-cluster bins may not span partition values

IcebergContext::recluster_files{,_with} refuses a bin that spans two partition values on every dispatch, not only on the streaming ones. The refusal runs before the catalog is read and before any output is written, and the error names the partition span and the remedy. Both shipped planners already group their files by partition value, so the compactor and the operator behave as before. A caller of the library API groups by partition value and calls once per group.

Persistent lane-cap refusals

When ingester.maxLanes reaches its per-process limit, a novel tenant and index lane now returns HTTP 503 or gRPC Unavailable. The response has no retry hint because an empty queue does not release an admitted lane. Existing lanes keep accepting writes, and writer failures remain 500 or Internal. Recovery requires routing traffic to a pod with capacity or changing the deployment. See What a 503 without Retry-After means: the persistent lane cap. Shipped in Siglake merge commit e3b4c1a.

Query tenant admission

query.allowedTenants and SIGLAKE_QUERY_ALLOWED_TENANTS add an optional exact allow-list for verified query tenant claims. The empty default accepts every usable claim. A non-empty list requires query.oidc.tenantClaim and stays independent from ingester.allowedTenants, so retained data remains readable after write admission stops.

An unlisted claim returns 403 before namespace or tenant-context creation. Workers apply their own list to coordinator-authenticated shard requests, and the coordinator forwards a worker's 403. The pre-registered siglake_query_tenant_denied_total{reason="not_allowed"} series and SiglakeQueryTenantsDenied alert identify the refusal. Shipped in Siglake merge commit bb48ab4.

Siglake 0.2.0 writes a sibling CRC-32 checksum with each new footer-v1 inverted index. Its metadata name is siglake.inverted_index.crc32.v1[.<column>]. A 0.1.x reader ignores the entry and reads the unchanged index blob, so old and new readers can share files. See What the footer text-index checksum covers.

Group-count counters carry the Iceberg namespace

The five aggregate-maintenance counters now name the Iceberg namespace as well as the table: siglake_group_count_short_aggregates_total, siglake_group_count_delta_write_failures_total, siglake_side_aggregate_publish_failures_total, siglake_group_count_auto_rebuilds_total and siglake_group_count_delta_write_retries_total. One compactor maintains the base namespace and every tenant_* namespace, each with its own events, so a bare table="events" merged every tenant onto one series and the alerts named a table an operator could not locate. The label is iceberg_namespace rather than namespace because Prometheus attaches the Kubernetes namespace under that name. All four alerts on these counters now name <iceberg_namespace>.<table>. See Group-count aggregate labels. Shipped in Siglake merge commits 66e2f11 and 33832f2.

Segmented text indexes, off by default

An opted-in streaming re-cluster builds the compressed segmented inverted index (seg2) as it emits Parquet row groups, then registers the finished Puffin statistics file in the same transaction as the data-file rewrite. A failed transaction leaves no discoverable index. Both SIGLAKE_SEGMENTED_INDEX_WRITES=1 and SIGLAKE_SEGMENTED_INDEX_READS=1 are required to build and to use the format, and both stay off in every packaged release pending AWS qualification. On a 14 x 7.34M-row corpus the writer averaged 282.77 s of rewrite time and 251.4 MiB of peak tracked heap, 12.0% faster and 80.8% smaller than the post-commit v1 rebuild it replaces. Queries recognize seg2 and keep reading whole-file v1 indexes. The unreleased seg1 prototype is no longer discovered. See Segmented text indexes (seg2).

Managed-index mapping ETags and If-Match

GET /api/v1/indexes/{id} and a successful index PUT return a strong ETag over the mapping, and PUT accepts an optional If-Match. A false condition returns 412 Precondition Failed with the exact rejecting base config and its matching ETag, including after a lost catalog compare-and-swap replays the transaction on another writer's mapping. The validator covers the table UUID and the full parsed index config, so a data-only commit preserves the tag and a delete-and-recreate changes it. A PUT without the header keeps its additive-only, idempotent behavior. See Managed-index ETag and If-Match preconditions.

OTLP export of Siglake's own logs and traces, off by default

Setting OTEL_EXPORTER_OTLP_ENDPOINT exports the log lines every binary already writes as OTLP log records, and spans at the ingest handlers, the compactor drain and the query server's per-request path as OTLP traces over HTTP. The coordinator injects W3C traceparent into each shard request and the worker extracts it, so a distributed query is one trace rather than a root span per replica. /metrics does not move: siglake_* stays Prometheus, which is what the alerts, the KEDA scalers and the dashboard read. Neither the chart nor the operator renders these variables, so they reach a pod through <tier>.extraEnv. See Export Siglake's own logs and traces.

Profile endpoints in a profiling build

/debug/pprof/{profile,heap,runtime} serve on-demand CPU, heap and tokio-runtime profiles on the --metrics-bind router, so one mount point covers the ingester, the compactor and the query tier. Reaching them takes two opt-ins no released binary carries: siglake-core's off-by-default profiling cargo feature, which only the PROFILING=1 build of deploy/Dockerfile enables, and SIGLAKE_PPROF_ENABLED=1 in the process. Without both, the routes are absent and answer 404. Nothing on the metrics port is authenticated. See Profiling a running deployment.

WAL mirror reclamation for the filesystem drain, off by default

compactor.mirrorLedgerReclaim (SIGLAKE_MIRROR_LEDGER_RECLAIM) lets the default single-replica compactor delete the mirror objects it has committed. The filesystem drain commits out of local sealed/ and never reads the mirror, so before this the prefix grew for as long as the cluster ingested: 421,632 objects and 37.3 GB a day at 20K events per second. With the knob on, the compactor marks each committed file's wal_segments row without claiming anything, and the unchanged retention pass deletes the object and then the row under committedRetentionSecs, whose 0 still deletes nothing. See Committed retention purges drained mirror objects.

A WAL segment the drain cannot read no longer blocks its batch

A segment the local filesystem drain fails to read three times in a row (SIGLAKE_COMPACTOR_POISON_ATTEMPTS; 0 restores the old behavior) moves to <wal>/poison/ with a .poison.json note holding the read error and the attempts spent. Its batch siblings commit on the next pass. poison/ is excluded from the automatic orphans/ disposition, survives restarts, and is never deleted or rewritten: siglake wal-requeue --wal <wal-root> is the way back once the cause is fixed. Each drain pass now spends at most three claims on a segment and then leaves it in sealed/ for the next cycle, so a pass that keeps failing no longer re-claims the same segments for its whole budget. See siglake wal-requeue.

The ingester's local WAL sweep reaches managed indexes

The sweep listed the WAL root and its tenant directories one level deep, and a managed index's segments sit at <root>/<tenant>/<index>/sealed/. In catalog-claim mode, where this sweep is the only thing that deletes an ingester's local copies, nothing was removed for a managed index and the PVC grew for the life of the pod. siglake_wal_local_sealed_segments counted only the directories the sweep visited, so it did not report the backlog either. The walk is now the drain's own: the root, each tenant, then each tenant's indexes. The deletion gates are unchanged.

wal.mirror.activeIntervalSecs uploads the segments ingest is writing

The active-segment mirror loop was handed the ingester's root WAL writer, and every request lands in a per-tenant writer or a backpressure lane's writer instead. Each tick flushed an empty writer and sent nothing: a process that logged "WAL active-segment mirror enabled" wrote no _active/ object, and the N-second loss bound the flag advertises held nowhere. The loop now asks the writer sets the HTTP handlers resolve against on every tick, so tenants, managed indexes and write shards created by later traffic are covered as they appear. Off by default, unchanged. See Your PVC-loss exposure window with active-segment uploads.

New censuses, alerts and a repair command for stranded aggregates

The maintenance compactor censuses every maintained table every 15 minutes for two states that do not heal on their own. A group-count aggregate short of its row count is reported on siglake_group_count_short_aggregates_total{iceberg_namespace,table,outcome} and warned on by SiglakeGroupCountAggregateShort; automatic repair is opt-in (compactor.shortAggregateRepair, SIGLAKE_AGG_SHORT_REPAIR=1) because it costs one Tier-2 query per maintained column. An inline time aggregate whose coverage chain the read guard refuses sets siglake_inline_coverage_unproven{iceberg_namespace,table} and fires SiglakeInlineCoverageUnproven after two censuses; the new Siglake rebuild-time-aggregates is its repair. Snapshot expiry also re-roots a coverage edge it would otherwise strand, counted on siglake_inline_coverage_reroots_total. Answers were exact throughout both states: what they cost is the fast path.

Six alerts added to the packaged rule file

The chart's PrometheusRule goes from 33 alerts to 39. Nothing was removed or renamed. Each of the six reads a metric the metrics reference lists.

New alert Fires on
SiglakeCompactorOrphansHeld A WAL orphan whose commit status the drain could not establish, held under <wal>/orphans/ for 15 minutes. Critical.
SiglakeWalIpcFramingRefused A sealed, legacy or recovered partial WAL segment an IPC framing walk refused as unreadable. Critical.
SiglakeInlineCoverageUnproven A table whose inline time aggregate the read guard refuses, over two censuses. Critical.
SiglakeGroupCountAggregateShort A group-count aggregate short of its table's row count. Warning.
SiglakeAutoPromotionNearCeiling A table at 80% or more of its effective auto-promotion column cap for 10 minutes. Warning.
SiglakeBatchReconciliationBacklogStalled Finished batch jobs on a query pod without a persisted terminal state past the 15 s recovery bound. Warning.

Query planning and caching fixes

Four fixes change what a query reads and leave every answer as it was. A text query over more indexed files than the text-index caches hold no longer re-reads every index blob from object storage on every execution: eviction now drops the blobs the parsed cache still covers and keeps the ones it has dropped. Whether to use a text index a file already carries is decided per execution, so a text predicate under a bare LIMIT stays on the scan path and increments siglake_query_inverted_index_declined_total{reason}. With the experimental decoded-file cache switched on, a scan whose predicate converts to an Iceberg predicate reads with that predicate intact and no longer populates the cache. A browse written WHERE TIMESTAMP >= ... gets the same scan-order hint and distribution estimate as the lowercase spelling.

An audit append has a deadline

SIGLAKE_QUERY_AUDIT_APPEND_DEADLINE_SECS bounds one query_audit append at 30 s by default; 0 restores 0.1.0's unbounded await. A single append that stopped answering used to hold the whole retention budget, so every later row was refused and nothing reached the table again until the process restarted. A batch that outlives the deadline is abandoned, never re-appended, and its rows are counted by siglake_query_audit_dropped_total{reason="append_deadline"}. Audit rows can be lost whole at the deadline, which Limitations records.

Release images name the commit they were built from

siglake --version and the siglake_build_info metric read a revision stamped in at build time, and the publish workflow passed none, so every published 0.1.0 image said unknown. Both the server and the operator image now carry the commit the release checkout resolved to. Image repositories and the image tag are unchanged: the tag says what the release is called, and the revision says what is in it.

The AWS reference deployment tears down what it provisioned

deploy/aws/down.sh exits with terraform's status when the destroy fails. Before this it ended on a log call, so a failed terraform destroy exited 0 with RDS, the warehouse bucket and the IAM role still running and billing. It also settles its teardown mode before it writes to the cluster: an unrecognised mode exits 1 naming the value, having run no helm, no kubectl delete, no aws s3 and no terraform destroy. deploy/aws/up.sh writes its own mode-0600 kubeconfig, checks that the context it wrote names the cluster it just provisioned, and hands that kubeconfig and context to every kubectl and helm call, leaving your own default kubeconfig and current context alone.

The adoption preflight prints one appliable manifest

siglake-operator --adopt-values prints one YAML document. The synthesized SiglakeCluster comes first, and everything below it is a comment: the findings, and the handover runbook including its commands. A saved copy is therefore a file kubectl apply -f reads, and nothing has to cut the resource out of it. The header line changed with it, so a script that matched the old # --- synthesized SiglakeCluster (save as cluster.yaml) --- line, or stopped at # preflight, now extracts nothing. Apply the saved file as it stands.

A new flag, --adopt-namespace <NS>, names the namespace the release runs in. It sets metadata.namespace on the resource and the -n on every runbook command, both abort paths included. It defaults to the release name, which is what the runbook assumed before the flag existed, so a release installed into a namespace other than its name needs it. See Migrate a Helm release to the operator.

What to do on upgrade

Each change takes effect when the tier that owns it starts: the lane refusal with the new ingester, the group-count labels with the new compactor, the shard header with the new query pods. Upgrading also changes future footer-v1 index writes, but it does not backfill existing files. No siglake migrate-schema run is required by this release: the events schema version does not move. The chart's pre-upgrade hook still runs the command, and it stays idempotent.

Change in the 0.2.0 line What to do on upgrade
A new lane at ingester.maxLanes returns hintless 503 / Unavailable instead of 500 / Internal. Check alerts and clients that classify ingest failures by status. A retry can exhaust against the same full pod, so send traffic to another pod or change the deployment.
A 0.2.0 coordinator refuses a shard response that carries no x-siglake-scan header, and a 0.1.0 worker omits that header when its shard ran no data-file scan. Roll the whole query tier in one go and keep it short. While both versions serve, a distributed query whose shard answered from a zero-scan fast path can fail on a 0.2.0 coordinator paired with a 0.1.0 worker. Single-pod query deployments are unaffected.
siglake wal-recover plans by default and writes only under --apply. Add --apply to any runbook, script or recovery drill that calls the old single-command form. Without it the command prints a plan and creates nothing, --to included. Read Recover a filesystem drain before the next drill.
A query_audit append is bounded at 30 s (SIGLAKE_QUERY_AUDIT_APPEND_DEADLINE_SECS). Nothing, unless you need 0.1.0's unbounded await: set the variable to 0. A batch that outlives the deadline is dropped whole and counted on siglake_query_audit_dropped_total{reason="append_deadline"}, so watch that series for the first days after the upgrade.
The filesystem drain sets an unreadable segment aside under <wal>/poison/ after three failed reads. Include <wal>/poison/ in whatever you use to inventory the WAL volume, and alert on SiglakeSegmentsQuarantined. Return a set-aside segment with siglake wal-requeue --wal <wal-root> once you have fixed the cause; nothing deletes the directory for you.
wal.mirror.activeIntervalSecs now uploads the open segments it always advertised. If you had the interval set on 0.1.0, expect the PUT rate and the mirror object count to rise to what the setting asks for, because the loop was sending nothing. Re-check the mirror's cost against What mirroring costs at the ingester.
The chart ships 39 alerts, six more than 0.1.0's 33. Re-apply the chart's PrometheusRule and check that your receiver routes SiglakeCompactorOrphansHeld, SiglakeWalIpcFramingRefused and SiglakeInlineCoverageUnproven, which are critical.
IcebergContext::recluster_files{,_with} refuses a bin that spans two partition values. Nothing for a chart or operator install: both shipped planners already group by partition value. A caller of the Rust API groups its files by partition value and calls once per group.
Arrow and Parquet 58, DataFusion 53.1, OpenDAL 0.57, Iceberg 0.10.1. Nothing. Tables stay Iceberg format version 2 on Parquet v2 and external readers see no change.
Footer-v1 index writes add a sibling checksum. A malformed or mismatched checksum refuses the index and increments siglake_index_footer_checksum_refused_total{reason}. Upgrade normally. A replacement that writes a footer-v1 index gains the checksum, but re-clustering with segmented writes enabled writes seg2 instead. An absent checksum does not increment the refusal counter, so zero refusals do not establish checksum coverage.
The five group-count aggregate-maintenance counters carry iceberg_namespace alongside table. A selector on table alone still matches, but it returns one series per namespace where it returned one in total, so re-check panels and alerts that expect a single series. Add iceberg_namespace to the recording rules, joins and by clauses that need per-namespace identity, and leave the ones that aggregate across namespaces on purpose as they are. The pre-upgrade series and the per-namespace ones have separate histories: rate() over a window that spans the upgrade reports each of them that still holds two samples in the window.

0.1.0: initial public release

Siglake is a horizontally-scalable, OTLP-native log analytics platform on Parquet v2, Apache Iceberg and DataFusion.

Timestamp handling, Jaeger reads, and what to do on upgrade

0.1.0 is the first public release, so there is no earlier release line to upgrade from. Five changes inside the line need an action if you are moving a pre-release install onto the release binary.

Change in the 0.1.0 line What to do on upgrade
The 2026-09-06 timestamp contract: tables are Iceberg format version 2, events.timestamp is a microsecond timestamptz, and the required timestamp_ns long holds the OTLP nanosecond verbatim. Recreate any warehouse written before that change. Those warehouses are format version 3 and cannot be migrated. Where a reader needs nanosecond exactness, convert timestamp_ns with that engine's own function: see Reconstruct nanosecond timestamps from timestamp_ns.
Additive schema changes are gated: a binary newer than the table refuses writes to columns the table lacks. Run siglake migrate-schema --all-tables --all-namespaces before or with the new binary. Each control plane runs it on a trigger: the chart renders the pre-upgrade hook while schemaMigration.enabled is true, its default, and the operator renders the Job when you raise the monotonic spec.schemaVersion counter; an unset counter, or 0, renders no Job. Run the command yourself wherever neither trigger fires.
The 2026-09-08 bounded Jaeger reads, with the 2026-09-09 name-list cache correction: the shim derives trace, span-row, Arrow-byte and distinct-name ceilings from each request's admission reservation. Re-check every Grafana Jaeger panel and API client. A search limit above the derived trace ceiling now returns 400 before planning, and a result that crosses another ceiling is refused whole with 413. Narrow the time range, service, operation, tags or limit on the panels that refuse.
Acknowledgement calls fsync(2) by default, covering segment bytes and the directory entries that name the segment. Expect acknowledgement latency to follow the WAL volume's flush latency. Send ?commit=auto per request where a client prefers the older write(2) boundary and can lose unflushed rows on a node crash.
The ingester listens for OTLP/gRPC logs and traces on 0.0.0.0:4317 by default. Allow or block port 4317 deliberately in your NetworkPolicy and Service. Set ingester.otlpGrpc.enabled: false if you do not want the listener; see Configure OTLP over gRPC.

Rolling ingesters and drains across a version change also takes one retention setting, applied before the first new binary starts. See Consumed-proof rolling upgrades.

2026-09-09: Jaeger name-list cache correction

Later 0.1.0 changes replaced the cache limits in the 2026-09-08 entry. A complete Jaeger name list is cached when its retained arena and stored lookup string together fit the 512 KiB per-entry cap. The shim's derived name ceiling supplies the row bound. The SQL cache's 128-row entry cap does not apply to these lists. See the Jaeger shim reference.

2026-09-08: bounded Jaeger reads

The Jaeger HTTP shim now derives trace, span-row, accumulated Arrow-byte, and distinct-name ceilings from each request's admission reservation. An excessive search limit is rejected with 400 before admission or planning. A result that crosses another ceiling is refused whole with 413, no data, and no Retry-After. The four refusal causes are exposed as jaeger_* values on siglake_query_breaker_trips_total. See the Jaeger shim reference.

The span plan now fetches at most one row beyond its row ceiling. A refused request no longer materializes the full match before returning 413.

2026-09-08: release hardening

  • An unknown field in a SQL request's limits object now returns 422 before planning, execution or batch enqueue. Unknown top-level keys remain allowed, and known limits above their tier ceiling remain clamped. The row-limit field is max_rows_returned, not the response field max_rows. See the SQL query reference.
  • Delete sweeps now report claimed tasks that remain non-terminal for more than twice SIGLAKE_DRAIN_WATCHDOG_SECS. The siglake_compactor_delete_tasks_stalled_total{state} counter drives the SiglakeDeleteTaskStalled alert. The dashboard-only siglake_compactor_delete_tasks_nonterminal{state} gauge records the last complete observation. See Monitor Siglake.
  • The Jaeger service and operation name routes now cache complete lists by table snapshot. Lists above 128 rows or 4 MiB remain uncached. See the Jaeger shim reference.

2026-09-06: portable timestamp contract

Siglake tables now use Iceberg format version 2. events.timestamp is a microsecond timestamptz (Parquet INT64 TIMESTAMP(MICROS, UTC)), and the required timestamp_ns long sibling preserves the OTLP nanosecond verbatim. Trino 483, Spark 3.5.9 with Iceberg 1.11.0, DuckDB 1.5.5 with core iceberg extension 45163a28, and PyIceberg 0.12.0 all read a fresh 2,000-row fixture and agreed on its exact timestamp_ns bounds. See External query engines.

Release overview

Ingest accepts OTLP/HTTP logs over POST /v1/logs and an Elasticsearch-compatible _bulk surface with per-index document mappings. The ingester also listens for OTLP/gRPC logs and traces on 0.0.0.0:4317 by default. The WAL force-seals on graceful shutdown. Backpressure lanes and per-tenant rate budgets bound the accept path. Auth is Authorization: Bearer <token> or OIDC.

Durability rests on the WAL. A request is acked once its rows are in the WAL. By default, the ingester calls fsync(2) before it acknowledges the request. The sync covers segment bytes and the directory entries that name the segment. A seal publishes the sealed name durably before unlinking the active copy. The power-loss guarantee assumes ext4 or xfs on a node-attached volume. A network filesystem provides whatever its fsync(2) and rename semantics guarantee.

Send ?commit=auto to accept an acknowledgement after write(2) instead. A node-level crash or power loss can discard rows that the kernel has not flushed in that mode. Sealed WAL segments are mirrored to object storage by default, but the asynchronous upload is outside the acknowledgement path. Neither mode waits for the rows to become queryable; that follows the drain's cadence (measured p50 ~5.5 s).

Multi-tenancy is single-tenant by default. Every ingest request routes to default. An X-Scope-OrgID naming another tenant is refused with 403 over HTTP or PermissionDenied over OpenTelemetry Protocol (OTLP)/gRPC. Set ingester.oidc.tenantClaim to route by a verified JSON Web Token (JWT) claim, or ingester.trustScopeHeader to trust the header. The claim option requires a usable claim, and a present header may only agree with it. Optional --allowed-tenants and --max-tenants bound what a client can create. The query server applies the claim rule through its own --oidc-tenant-claim: a claim that is missing, blank, non-string, over 128 characters or outside [A-Za-z0-9_-] is a 403 before routing rather than a fall back to the default namespace, and identifiers are validated rather than repaired, so acme.corp cannot reach acmecorp's data.

Storage is Iceberg tables on any object store: physically time-ordered Parquet with per-file group-count footers (typed columns included), token and trigram bloom filters, and continuous leveled compaction with overlap-depth convergence. The compactor's siglake_table_live_data_files gauge is read from the snapshot summary each cycle, so it stays exact on tables too large to walk inside the sampler's per-table budget. The level and depth gauges can lag, and siglake_table_gauges_sampled_at_seconds says when they were last walked. Group-count delta writes get four attempts with retry and failure metrics. An exhausted write leaves a durable marker for the maintenance compactor to rebuild automatically. siglake_group_count_auto_rebuilds_total{table,outcome} records the result, and SiglakeGroupCountDeltaLost fires only when that rebuild fails or remains incomplete. As the operator fallback, siglake rebuild-group-counts repairs the aggregate from committed files without double-folding late deltas; its --admit-typed-columns adds the typed columns a table created before typed side aggregates never carried, so no rewrite is needed for those.

Query is DataFusion SQL (/api/v1/sql) with a transparent distributed coordinator. Ordered-scan early stop in both time directions (reversed tail-chunk decode). Zero-scan fast paths for counts, group-bys, histograms, distinct counts, and dimensional or typed filters served from footers and snapshot-keyed side aggregates. Snapshot-keyed result and decoded-chunk caches, invalidated by commit and never by TTL. Per-request scan and cost attribution (stats.scan). Bounded query spill uses SIGLAKE_QUERY_SPILL_DIR and SIGLAKE_QUERY_SPILL_MAX_BYTES, the chart's query.spill.* ladder, and a matching query-pod ephemeral-storage limit. Memory-pool or spill-cap refusals return 503 plus Retry-After, including refusals forwarded from workers, and increment siglake_query_breaker_trips_total{breaker="pool_exhausted"}.

Freshness comes from sealed-WAL buffer serving. Records are queryable in seconds, before commit, and are folded exactly into every fast path.

Attribute capture is lossless. OTLP resource and log attributes stay queryable via attr_get(). Hot keys auto-promote to typed columns with backfill and query rewrite (opt-in).

Streaming consumers read the WAL directly. siglake_wal::consumer::SegmentConsumer is a supported interface for reading WAL segments as they seal, with a durable cursor, at-least-once delivery, retention that waits for slow consumers (bounded, so a stuck consumer degrades to "you missed some" rather than filling the disk), and CRC integrity. See docs/CONSUMING_SEGMENTS.md in the source repository. The four-tier semantic detection pipeline that shipped inside Siglake through 2026-08-29 was moved out to run entirely on top of this interface and is maintained as its reference consumer.

Schema evolution is additive and gated. A table records the schema version it is at. When the running binary is newer, the columns the table lacks cannot be written, so the write is refused, naming the column and the remedy, rather than silently dropped. Run siglake migrate-schema --all-tables --all-namespaces; it is additive-only and idempotent, and --all-namespaces matters on any multi-tenant install because events exists once per tenant namespace. Each control plane runs it on a trigger. The chart renders the pre-upgrade hook while schemaMigration.enabled is true, its default. The operator renders the Job when you raise the monotonic spec.schemaVersion counter, and holds the workload rollout until the Job finishes; the counter says "run a migration" rather than naming a target version, and leaving it unset or at 0 renders no Job. See Schema migrations. Anywhere neither trigger fires, run the command yourself before the new binary serves writes.

The operational surface is Helm charts and a Kubernetes operator (SiglakeCluster CRD), with per-tier spec.resources.<tier>, the chart's 4Gi query-pod default, schema-migration jobs, and offline Helm-release adoption. Invalid configurations set InvalidSpec with one of ten reasons rather than being partly honored; the former QueryAutoscalingIgnored condition is removed. A Terraform/EKS reference deployment covers BYOC installs. The release also carries Prometheus metrics throughout, Prometheus alert rules for the silent-loss counters, audit logging, and retention, GC and delete sweeps.

Interoperability rests on the open format. The warehouse is plain Iceberg-on-Parquet at format version 2, with no v3-only type in any schema. Event time is timestamp, a microsecond timestamptz (Parquet INT64 TIMESTAMP(MICROS, UTC)) that every Iceberg reader maps. events also carries timestamp_ns, a required long holding the OTLP time_unix_nano value verbatim, so nanosecond exactness is available externally and Siglake's own scans lose nothing. Each engine reconstructs the nanosecond timestamp with its own conversion, listed under Reconstruct nanosecond timestamps from timestamp_ns. Trino, Spark, DuckDB and PyIceberg all read the warehouse without Siglake in the path; the versions tested and the measured output are in External query engines . Warehouses written before this change are format version 3 and must be recreated, not migrated (docs/DESIGN_time_ordered_storage.md, "Timestamp contract").

See the Performance section of README.md for measured comparisons against Quickwit, Elasticsearch, ClickHouse, and a vanilla-Parquet DuckDB baseline.

Before 0.1.0

The project was named knulps until 2026-06-12. Pre-rename docs are kept in the internal history, and older git history uses the old name.

Development was phased: the storage engine, ingest, and query (phases 1 to 4); scale-out, multi-tenancy, and hardening; the former in-tree detector pipeline (phase 5, retired in source commit e2a6ab2 on 2026-08-29); and the 2026-06/07 performance arc. See About.