Retention and deletes¶
Use this guide to expire old rows, execute predicate deletes, and reclaim their files. These jobs are separate, and a complete deletion usually needs more than one.
| Mechanism | Removes | Granularity |
|---|---|---|
| Retention sweep | Whole data files past a horizon | File / day |
| Delete tasks | Rows matching a predicate | Row (via file rewrite) |
| Snapshot expiry | Old Iceberg snapshots | Snapshot (metadata only) |
| Orphan GC | Unreferenced files on storage | File (physical) |
Retention and delete tasks change what queries see. Only orphan garbage collection (GC) frees storage. A row is gone from results after a delete sweep, but the bytes remain in superseded files until snapshot expiry drops the references and GC reclaims them.
Apply a retention policy¶
Set the per-index policy¶
Send retention with the index, on the create or update request. Its one
field, period_secs, is a horizon in seconds: 2592000 is 30 days. Omit
retention for an index that keeps data forever.
Run a retention sweep¶
A policy is a record, and nothing acts on it until a sweep runs. Only
siglake retention-sweep applies index retention: the compactor's retention
pass covers committed WAL segments, not index data. Schedule the sweep
yourself, as in the complete setup.
# dry-run — prints files/bytes/rows that would be removed
siglake retention-sweep --namespace siglake
# actually apply
siglake retention-sweep --namespace siglake --apply
# one index only
siglake retention-sweep --index app-logs --apply
Without --index, it sweeps every managed index in the namespace.
Retention removes whole files
A file is dropped only when its manifest maximum timestamp is older than the horizon. A file spanning the boundary survives whole.
Your effective retention is therefore the policy plus the span of the last surviving file. There is no row-level retention.
Per-table window sealing during compaction bounds output time spans. This makes retention closer to the configured period but not exact.
Schedule it as a CronJob mirroring the audit-rotate pattern: same image,
same env, different subcommand.
Delete tasks¶
For GDPR-style "delete everything matching this predicate".
curl -sX POST http://localhost:8089/api/v1/delete-tasks \
-H 'Content-Type: application/json' \
-d '{
"index_id": "app-logs",
"predicate_sql": "user_id = 12345",
"start_ts": "2026-01-01T00:00:00Z",
"end_ts": "2026-07-01T00:00:00Z"
}'
A valid request returns 201 with a new task in pending. The request itself
does not execute the task; the compactor's next idle-cycle sweep does, unless
you have turned that sweep off.
Set start_ts and end_ts to limit how many files the task must rewrite.
Check status:
Large candidates are rewritten in a stream¶
Execution rewrites each candidate file without its matching rows. A candidate
that fits the in-memory caps is rewritten from one decoded batch. Past 16 MiB
compressed or 128 Ki rows, the rewrite streams, and streamed output carries no
footer inverted index and no whole-file raw trigram bloom. Under
compactor.indexRebuild: true the post-commit rebuild adds Puffin
inverted-index sidecars for those files; it is off by default, and nothing adds
the bloom.
A delete rewrite also writes at rewrite generation 0, like an ingest flush, so
a table with index_at_flush: false defers the in-memory arm's inline indexes
as well.
Either gap costs pruning, not correctness. A later query scans and row-evaluates that file instead of skipping it, and still returns exact rows. See What each merge path writes for the same split during compaction.
Delete tasks are bound to one index incarnation¶
A queued task belongs to the index incarnation it was accepted against, not to
the index name. Submission records the table_uuid the server resolved
index_id to, and the executor rewrites that Iceberg table or none.
Index ids are reusable. DELETE /api/v1/indexes/app-logs followed by a POST
of the same id gives a different table under the same name. A task queued
against the first table is refused once the second exists: it goes failed,
and error names both tables. The deletion was authorized for rows that left
with their index, so the replacement's rows are never touched. Recover by
resubmitting against the index that exists, as for any failed
task.
Tasks recorded by a build older than this binding carry a null table_uuid.
They stay readable, nothing rewrites or migrates them, and they are refused at
execution with the same instruction, because nothing infers an incarnation from
a reused name. If your warehouse holds requests that were acknowledged before
you upgraded, resubmit them.
A refusal commits no snapshot. The executor checks the table twice, when it
loads the table and again against a fresh catalog read before the commit, so a
task refused after its rewrite has already run leaves output files that no
snapshot references. The error says so, and orphan GC reclaims them. That
guarantee is specific to this refusal; a failed task in general does not prove
that nothing was deleted.
See Delete tasks for table_uuid and
the rest of the response fields.
Execution runs by default¶
Creating a task records it in a ledger, and a sweep executes it. The compactor runs that sweep in its idle cycle unless you turn it off, so a task you submit executes without any further configuration.
The chart ships the switch on, and renders the environment variable either way,
so deleteTasks: false renders SIGLAKE_DELETE_TASKS=0:
Outside Helm, the compactor sweeps unless SIGLAKE_DELETE_TASKS is 0, off,
false or no. Any other value, a typo included, leaves the sweep running. The
operator renders SIGLAKE_DELETE_TASKS=1 and its CR has no field for the
switch, so the opt-out is spec.extraEnv:
siglake-operator --adopt-values reports that same line when the chart values
it reads set compactor.deleteTasks: false.
You can also sweep by hand, with the compactor's sweep on or off:
Turning the sweep off leaves accepted tasks unexecuted
With SIGLAKE_DELETE_TASKS=0, POST /api/v1/delete-tasks still answers
201 and records the task. Nothing then runs it except
siglake delete-sweep --apply, so a compliance workflow that disables the
compactor's sweep has to schedule that command instead.
Applying claims each task, and the claim is never released
A dry-run is always safe to run: without --apply, delete-sweep
evaluates the pending tasks and reports what they would rewrite. It writes
neither a task record nor a snapshot and claims nothing, so a
preview can never poison the apply that follows it. Its
tasks_already_claimed is always 0. Run it against a busy index whenever
you like.
Applying takes ownership of each task first. A task's record is still
last-write-wins, so ownership is a separate object. Before executing a
pending task, an executor create-only-writes a sibling
_siglake/config/delete_tasks/<namespace>/<task_id>.claim next to the
task's record. Of any set of executors racing for that task exactly one
wins the write and runs it; each loser counts it in tasks_already_claimed,
leaves the record alone and rewrites nothing. Two processes can no longer
each rewrite the same task and each write their own terminal state over the
other's.
[APPLY] app-logs: tasks_examined=2 tasks_completed=1 tasks_failed=0 tasks_already_claimed=1 files_rewritten=3 rows_deleted=1188
1 pending task(s) are claimed by another executor and were not re-run
The claim covers every entry point; the compactor's own guard does not.
A maintaining compactor still takes a cluster-wide delete_tasks
maintenance lease before each sweep, but that lease is not per-task, and
siglake delete-sweep --apply neither takes it nor is held off by it. A
manual sweep alongside compactor.deleteTasks: true, or two manual sweeps
of the same index, is therefore no longer a double-execution hazard. The
two executors split the pending set. Overlap still complicates
accounting: each sweep's printed totals cover only the tasks it won, so
read them per executor, never as the index's total. One executor per index
remains the setup that is easiest to reason about.
A crashed executor strands its task under a permanent claim
A claim is never released after completion, failure, or a crash during a rewrite. That is deliberate: a process that died somewhere around its rewrite commit is exactly the case where a second executor must not start, and a takeover on claim age alone, with no fencing token the Iceberg commit could check, would let the dead owner's rewrite land after its successor's.
The cost is a new stranded state. An executor that dies between claiming a
task and writing its running record leaves that task pending forever.
Start with the API. For a task still in pending,
GET /api/v1/delete-tasks/{id} returns a read-only claim object
reporting what that read saw of the claim: present, the observed_at
it applies to, and, when the claim body decodes, the claimant UUID
identifying the process that took it, its claimed_at, and the
age_seconds since then. See Delete
tasks for the field-by-field
shape.
That response does not establish the executor's state:
- A claim whose body does not decode stays visible. An empty,
truncated or unparseable body, or one written against a different
task_id, still answerspresent: true, withclaimant,claimed_atandage_secondsallnull. Presence alone is what excludes the next executor, so the object counts either way. present: falsedescribes that one read. An executor can take the task a millisecond later.- Presence, the UUID, and the age do not tell you whether the claimant is alive. A claim hours old fits a long rewrite and a dead process equally well, and no age makes a takeover safe.
The API observation tells you what holds the task, not that the task is stranded. Corroborate it with the executor-side signals:
tasks_already_claimedin adelete-sweep --applyline, and theN pending task(s) are claimed by another executornote under it.- The executor logs
delete task is already claimed by another executor; skippingwith thetask_id. The compactor emits this on its own sweeps. siglake_compactor_delete_tasks_stalled_total{state="running"|"pending_claimed"}and theSiglakeDeleteTaskStalledalert. The counter increments once per stalled-task observation in a complete sweep, not once per distinct task; the WARN line carries the task id and index. See Monitoring for the exact rule and operator action.
The stall threshold is twice the sweep's resolved watchdog ceiling. Its age
is the claim object's age: an upper bound on execution duration, not a
running duration, and it keeps growing after a task strands. A stalled
report proves neither that the executor is dead nor whether its rewrite
committed. A pending task without a claim is outside this signal, and a
watchdog ceiling of 0 disables stalled classification. The dashboard-only
siglake_compactor_delete_tasks_nonterminal{state} gauge is only the last
complete observation; it can remain unchanged when execution or observation
stops, so it does not establish liveness or current coverage.
A non-zero tasks_already_claimed count also covers the healthy case of a
task another executor is running right now. Treat a claim as stranded only
after reading the task record, the compactor WARN, and your audit trail.
Never decide from claim age or the retained gauge alone.
Siglake has no flag or endpoint that releases a claim. Orphan GC does not
reach _siglake/config/, so one small claim object remains for each task.
Recovery uses the same explicit resubmission as a failed task. The new
task ID gets a new claim object. Deleting a claim
object out of band is not a supported step: it re-arms the double execution
the claim exists to prevent, and it does not recover a task already
stranded in running.
A warehouse without create-only writes fails the sweep
The claim needs a conditional create: the filesystem backend takes
O_EXCL, S3 sends If-None-Match: *. If the warehouse store does not
advertise that capability, or advertises it and then rejects the write,
as an S3-compatible store that ignores If-None-Match does,
the sweep returns an error rather than executing the task unclaimed. It
names the task it refused. There is no unconditional fallback because one
would silently restore double execution.
The check happens as the first pending task is about to run, so a sweep
over an index with nothing pending still succeeds on such a store and the
problem surfaces only once a deletion request exists. The compactor's
multi-index sweep aborts at the first index it cannot claim in, logs
delete-task sweep failed, leaves the remaining indexes unswept, and fails
the same way every later interval. On that store, acknowledged requests
pile up pending with the only evidence in the compactor's log.
Recovering a failed task¶
The lifecycle is one-way. The executor takes pending tasks only, and nothing
returns a terminal task to pending. There is no retry endpoint or automatic
retry. A task that reached failed never runs again, even though
its submission was acknowledged. The same holds for one left in running by a
crash, or by a status write that failed after the rewrite committed, and for
one left pending under a claim its executor never released. Nothing sweeps
any of them up, and they all recover the same way.
Recovery is an explicit resubmission:
-
Read the failure with
GET /api/v1/delete-tasks/{id}. Checkerrorfor the cause where the task has one. -
Remedy that cause, and make sure the executor runs the fixed build.
- POST the request again under the same tenant identity. Send
index_id,predicate_sql,start_ts,end_ts, request fields only, exactly as in the submission above. - Record both IDs. The response carries a new
task_idinpending; the failed record and itserrorare left exactly as they were, so both ids belong in the audit trail for the request. - Confirm execution through the compactor's maintenance sweep
or a manual
delete-sweep --apply. Poll the new id until it reachesdoneorfailed. The new id carries its own claim, so the stranded one's claim does not hold it back.
A resubmission is not a retry
- It is not idempotent at the HTTP layer: every POST creates another task.
- It promises nothing about exactly-once execution.
- Each task reports only what its own run rewrote, not the request's cumulative effect across attempts.
- Over an unchanged dataset, with the same predicate and the same fixed
bounds, the repeat finishes with
rows_deleted: 0,files_rewritten: 0and commits no snapshot. A time-dependent predicate, or rows that arrived since, makes it delete a different set.
failed does not prove nothing was deleted
A task recorded failed because its commit result was ambiguous, or
because the status write failed after the rewrite committed, may have
committed that rewrite. Check rows_deleted on the new task against what
you expected before concluding the first attempt did nothing.
A task that failed before its commit is the easier case: the whole task commits once. It leaves inactive, unreferenced output files for orphan GC rather than a partly deleted snapshot.
A new task reaching done removes the rows from the live snapshot, and no
more: erasure still needs the chain below.
The full sequence to make a delete permanent¶
Deleting rows rewrites the affected files. The old files still exist and are
still referenced by older snapshots. Every step is scoped to the table you
deleted from, app-logs in the examples above, not to events.
- Execute the delete task. The rows leave the current snapshot.
- Let snapshot expiry drop every snapshot taken before the delete.
- Check readiness with a table-scoped
gc-orphansdry run. - Run orphan GC with
--applyagainst that table and namespace. GC is per table and--tabledefaults toevents, so a scheduled events GC never reclaims anything underapp-logs. - Expire the noncurrent object versions. If S3 versioning is on (the Terraform module enables it), deleted objects persist as noncurrent versions until a lifecycle rule expires them.
Step 3 is a dry run:
Step 4 is the same command with --apply:
Its effect on time travel is absolute. Once step 2 finishes you cannot query the table as of any instant before the delete, and there is no way back. That is what the sequence buys you, for the table it is scoped to. It does not reach the copies of the same events held outside that table, which expire on their own schedules.
Skipping step 5 is the most common compliance gap. Running GC only for
events while the deleted rows live in a user index is the next most common.
Copies the delete chain does not reach: WAL mirror, backups and replicas¶
The chain above erases one table's storage, not the deployment. Orphan GC is rooted at the table location and refuses any file outside it, so a WAL copy, a backup or a replica keeps the original event bytes until its own rule expires them.
| Copy | Where it lives | What expires it |
|---|---|---|
| Sealed mirror segments | <warehouse>/wal-mirror/ |
Committed retention, or a lifecycle rule |
| Partial segment snapshots | <warehouse>/wal-mirror/_active/ |
The uploader's delete when the segment seals (Siglake 0.2.0), or a lifecycle rule for older and failed-delete partials |
| Committed local segments | The WAL volume's committed/ |
The drain's sweep, within 3,600 seconds |
| Quarantined segments | The WAL volume's poison/ |
An operator, by hand |
| Noncurrent versions, replicas | The bucket and any replica bucket | A lifecycle rule in each bucket |
| Catalog backups | Postgres snapshots and point-in-time recovery | Your backup retention |
Record each window beside the delete chain: an auditor asking when the data is gone is asking about the longest one. Do not shorten one to hurry an erasure, because a mirror expiry below your worst-case drain backlog deletes segments a recovery needs. Segments the drain has not committed are a copy too, and their rows enter the table on a later commit, so delete after the events you target have committed.
What goes wrong in steps 2 and 3¶
Snapshot expiry is on by default (60 s, retain last 100), but retainLast is
a count of snapshots, not an age: only newer snapshots displace an older one
from the retained window. The 60-second cadence sets no wall-clock erasure
deadline. It bounds how promptly a table with more than 100 snapshots is
trimmed, nothing more. An idle or low-traffic index, which is what most user
indexes are, may sit below the retain count and keep its pre-delete snapshots,
and their files, indefinitely.
Siglake has no subcommand, HTTP endpoint, or sql-direct Iceberg metadata
table that lists a table's retained snapshots or the files they reference. The
intended operator check is the table-scoped gc-orphans dry run in step 3. It
reports the aggregate orphan count and bytes only: it does not name files or
prove that particular deleted data is no longer referenced. It also skips
files modified within --min-age-secs (24 hours by default), so a zero count
may mean the superseded files are still retained, or still inside that safety
window. Re-run the dry run after the table has taken further commits and the
rewritten files have aged past the window. File-level compliance evidence
still requires examining the Iceberg metadata and manifests outside Siglake.
Don't narrow retainLast to hurry erasure mid-upgrade
Lowering compactor.snapshotExpire.retainLast does drop the pre-delete
snapshots sooner, but during a mixed-version rollout the consumed-proof
guard requires retainLast >= 400 until well after the last old writer
exits. Read
Consumed-proof rolling upgrades
before changing it.
Expire snapshots¶
On by default:
It exists for performance, not cleanliness: metadata.json is read and
rewritten on every commit, so unbounded snapshot growth slows the entire
commit path.
You cannot use Iceberg time travel beyond the retained snapshots. Set
intervalSecs: 0 to keep every snapshot, which adds a growing cost to each
commit.
Expiry is metadata-only. Expired snapshots' files become orphans.
Erasure is not the only claim on this value: distributed queries need the
retained window to outlast a coordinator's stale metadata cache, and a
mixed-version rollout needs it at >= 400. Snapshot
retention weighs all three consumers
against the per-commit cost.
Reclaim orphan files¶
# dry-run — reports orphan count and bytes
siglake gc-orphans --table events
# apply
siglake gc-orphans --table events --apply
# tighten the safety window
siglake gc-orphans --table events --min-age-secs 3600 --apply
| Flag | Default | Meaning |
|---|---|---|
--table |
events |
Table to GC. Run per table. |
--min-age-secs |
86400 |
Skip files modified within this window. |
--apply |
off | Actually delete. |
Don't shorten --min-age-secs casually
The 24-hour default guards against deleting a concurrently-written file that isn't yet referenced by a committed snapshot. Shortening it narrows that guard. Only do so if you understand your write cadence, and never below your longest commit latency.
File and S3 warehouses only.
Audit rotation¶
query_audit commits per flush, so it accumulates snapshots fast.
# non-destructive: expire old snapshots, keep the rows
siglake audit-rotate --max-age-secs 604800
# preview
siglake audit-rotate --max-age-secs 604800 --dry-run
The bare form is destructive
Without --max-age-secs, audit-rotate drops and recreates the table
as a coarse row TTL that deletes all historical audit rows. (Iceberg-rust
0.9 has no public row-level delete.) The recreated table keeps the same
schema, partitioning, and sort order.
Use --max-age-secs unless you actually want to discard your audit
history.
The running query server picks the table up on its next audit write because
ensure_query_audit_table is idempotent.
Via the operator, retention.queryAuditRotateIntervalDays renders a CronJob.
Schedule the complete setup¶
The compactor runs two of these jobs for you, both on by default: snapshot expiry on its own cadence, and delete-task execution in its idle cycle. Retention sweeps, orphan GC and audit rotation have no in-process scheduler, so they run as CronJobs.
compactor:
deleteTasks: true # the shipped default; `false` renders SIGLAKE_DELETE_TASKS=0
snapshotExpire:
intervalSecs: 60
retainLast: 100
Plus CronJobs:
| Schedule | Command | Scope |
|---|---|---|
| Daily | siglake retention-sweep --apply |
Every managed index in the namespace |
| Daily | siglake gc-orphans --table events --apply |
events only |
| Daily | siglake gc-orphans --namespace siglake --table app-logs --apply |
One user index; add a row per index |
| Weekly | siglake audit-rotate --max-age-secs 604800 |
query_audit |
retention-sweep is the only one of these that fans out across the namespace.
GC is per table, so the events row covers only events. Each index whose
data must leave storage, including every index named in a delete task, needs
its own row.
And an S3 lifecycle rule expiring noncurrent versions, without which none of the above erases anything on a versioned bucket.
Always dry-run first
retention-sweep, gc-orphans, and delete-sweep are all dry-run by
default, and they print exactly what they would remove. Read that output
before adding --apply to a CronJob.