Skip to content

Retention and deletes

Use this guide to expire old rows, execute predicate deletes, and reclaim their files. These jobs are separate, and a complete deletion usually needs more than one.

Mechanism Removes Granularity
Retention sweep Whole data files past a horizon File / day
Delete tasks Rows matching a predicate Row (via file rewrite)
Snapshot expiry Old Iceberg snapshots Snapshot (metadata only)
Orphan GC Unreferenced files on storage File (physical)

Retention and delete tasks change what queries see. Only orphan garbage collection (GC) frees storage. A row is gone from results after a delete sweep, but the bytes remain in superseded files until snapshot expiry drops the references and GC reclaims them.

Apply a retention policy

Set the per-index policy

Send retention with the index, on the create or update request. Its one field, period_secs, is a horizon in seconds: 2592000 is 30 days. Omit retention for an index that keeps data forever.

{
  "index_id": "app-logs",
  "retention": {"period_secs": 2592000}
}

Run a retention sweep

A policy is a record, and nothing acts on it until a sweep runs. Only siglake retention-sweep applies index retention: the compactor's retention pass covers committed WAL segments, not index data. Schedule the sweep yourself, as in the complete setup.

# dry-run — prints files/bytes/rows that would be removed
siglake retention-sweep --namespace siglake

# actually apply
siglake retention-sweep --namespace siglake --apply

# one index only
siglake retention-sweep --index app-logs --apply

Without --index, it sweeps every managed index in the namespace.

Retention removes whole files

A file is dropped only when its manifest maximum timestamp is older than the horizon. A file spanning the boundary survives whole.

Your effective retention is therefore the policy plus the span of the last surviving file. There is no row-level retention.

Per-table window sealing during compaction bounds output time spans. This makes retention closer to the configured period but not exact.

Schedule it as a CronJob mirroring the audit-rotate pattern: same image, same env, different subcommand.

Delete tasks

For GDPR-style "delete everything matching this predicate".

curl -sX POST http://localhost:8089/api/v1/delete-tasks \
  -H 'Content-Type: application/json' \
  -d '{
    "index_id": "app-logs",
    "predicate_sql": "user_id = 12345",
    "start_ts": "2026-01-01T00:00:00Z",
    "end_ts": "2026-07-01T00:00:00Z"
  }'

A valid request returns 201 with a new task in pending. The request itself does not execute the task; the compactor's next idle-cycle sweep does, unless you have turned that sweep off.

Set start_ts and end_ts to limit how many files the task must rewrite.

Check status:

curl -s localhost:8089/api/v1/delete-tasks
curl -s localhost:8089/api/v1/delete-tasks/$ID

Large candidates are rewritten in a stream

Execution rewrites each candidate file without its matching rows. A candidate that fits the in-memory caps is rewritten from one decoded batch. Past 16 MiB compressed or 128 Ki rows, the rewrite streams, and streamed output carries no footer inverted index and no whole-file raw trigram bloom. Under compactor.indexRebuild: true the post-commit rebuild adds Puffin inverted-index sidecars for those files; it is off by default, and nothing adds the bloom.

A delete rewrite also writes at rewrite generation 0, like an ingest flush, so a table with index_at_flush: false defers the in-memory arm's inline indexes as well.

Either gap costs pruning, not correctness. A later query scans and row-evaluates that file instead of skipping it, and still returns exact rows. See What each merge path writes for the same split during compaction.

Delete tasks are bound to one index incarnation

A queued task belongs to the index incarnation it was accepted against, not to the index name. Submission records the table_uuid the server resolved index_id to, and the executor rewrites that Iceberg table or none.

Index ids are reusable. DELETE /api/v1/indexes/app-logs followed by a POST of the same id gives a different table under the same name. A task queued against the first table is refused once the second exists: it goes failed, and error names both tables. The deletion was authorized for rows that left with their index, so the replacement's rows are never touched. Recover by resubmitting against the index that exists, as for any failed task.

Tasks recorded by a build older than this binding carry a null table_uuid. They stay readable, nothing rewrites or migrates them, and they are refused at execution with the same instruction, because nothing infers an incarnation from a reused name. If your warehouse holds requests that were acknowledged before you upgraded, resubmit them.

A refusal commits no snapshot. The executor checks the table twice, when it loads the table and again against a fresh catalog read before the commit, so a task refused after its rewrite has already run leaves output files that no snapshot references. The error says so, and orphan GC reclaims them. That guarantee is specific to this refusal; a failed task in general does not prove that nothing was deleted.

See Delete tasks for table_uuid and the rest of the response fields.

Execution runs by default

Creating a task records it in a ledger, and a sweep executes it. The compactor runs that sweep in its idle cycle unless you turn it off, so a task you submit executes without any further configuration.

The chart ships the switch on, and renders the environment variable either way, so deleteTasks: false renders SIGLAKE_DELETE_TASKS=0:

compactor:
  deleteTasks: true # the shipped default

Outside Helm, the compactor sweeps unless SIGLAKE_DELETE_TASKS is 0, off, false or no. Any other value, a typo included, leaves the sweep running. The operator renders SIGLAKE_DELETE_TASKS=1 and its CR has no field for the switch, so the opt-out is spec.extraEnv:

spec:
  extraEnv:
    - name: SIGLAKE_DELETE_TASKS
      value: "0"

siglake-operator --adopt-values reports that same line when the chart values it reads set compactor.deleteTasks: false.

You can also sweep by hand, with the compactor's sweep on or off:

siglake delete-sweep --index app-logs           # dry-run
siglake delete-sweep --index app-logs --apply

Turning the sweep off leaves accepted tasks unexecuted

With SIGLAKE_DELETE_TASKS=0, POST /api/v1/delete-tasks still answers 201 and records the task. Nothing then runs it except siglake delete-sweep --apply, so a compliance workflow that disables the compactor's sweep has to schedule that command instead.

Applying claims each task, and the claim is never released

A dry-run is always safe to run: without --apply, delete-sweep evaluates the pending tasks and reports what they would rewrite. It writes neither a task record nor a snapshot and claims nothing, so a preview can never poison the apply that follows it. Its tasks_already_claimed is always 0. Run it against a busy index whenever you like.

Applying takes ownership of each task first. A task's record is still last-write-wins, so ownership is a separate object. Before executing a pending task, an executor create-only-writes a sibling _siglake/config/delete_tasks/<namespace>/<task_id>.claim next to the task's record. Of any set of executors racing for that task exactly one wins the write and runs it; each loser counts it in tasks_already_claimed, leaves the record alone and rewrites nothing. Two processes can no longer each rewrite the same task and each write their own terminal state over the other's.

[APPLY] app-logs: tasks_examined=2 tasks_completed=1 tasks_failed=0 tasks_already_claimed=1 files_rewritten=3 rows_deleted=1188
1 pending task(s) are claimed by another executor and were not re-run

The claim covers every entry point; the compactor's own guard does not. A maintaining compactor still takes a cluster-wide delete_tasks maintenance lease before each sweep, but that lease is not per-task, and siglake delete-sweep --apply neither takes it nor is held off by it. A manual sweep alongside compactor.deleteTasks: true, or two manual sweeps of the same index, is therefore no longer a double-execution hazard. The two executors split the pending set. Overlap still complicates accounting: each sweep's printed totals cover only the tasks it won, so read them per executor, never as the index's total. One executor per index remains the setup that is easiest to reason about.

A crashed executor strands its task under a permanent claim

A claim is never released after completion, failure, or a crash during a rewrite. That is deliberate: a process that died somewhere around its rewrite commit is exactly the case where a second executor must not start, and a takeover on claim age alone, with no fencing token the Iceberg commit could check, would let the dead owner's rewrite land after its successor's.

The cost is a new stranded state. An executor that dies between claiming a task and writing its running record leaves that task pending forever.

Start with the API. For a task still in pending, GET /api/v1/delete-tasks/{id} returns a read-only claim object reporting what that read saw of the claim: present, the observed_at it applies to, and, when the claim body decodes, the claimant UUID identifying the process that took it, its claimed_at, and the age_seconds since then. See Delete tasks for the field-by-field shape.

That response does not establish the executor's state:

  • A claim whose body does not decode stays visible. An empty, truncated or unparseable body, or one written against a different task_id, still answers present: true, with claimant, claimed_at and age_seconds all null. Presence alone is what excludes the next executor, so the object counts either way.
  • present: false describes that one read. An executor can take the task a millisecond later.
  • Presence, the UUID, and the age do not tell you whether the claimant is alive. A claim hours old fits a long rewrite and a dead process equally well, and no age makes a takeover safe.

The API observation tells you what holds the task, not that the task is stranded. Corroborate it with the executor-side signals:

  • tasks_already_claimed in a delete-sweep --apply line, and the N pending task(s) are claimed by another executor note under it.
  • The executor logs delete task is already claimed by another executor; skipping with the task_id. The compactor emits this on its own sweeps.
  • siglake_compactor_delete_tasks_stalled_total{state="running"|"pending_claimed"} and the SiglakeDeleteTaskStalled alert. The counter increments once per stalled-task observation in a complete sweep, not once per distinct task; the WARN line carries the task id and index. See Monitoring for the exact rule and operator action.

The stall threshold is twice the sweep's resolved watchdog ceiling. Its age is the claim object's age: an upper bound on execution duration, not a running duration, and it keeps growing after a task strands. A stalled report proves neither that the executor is dead nor whether its rewrite committed. A pending task without a claim is outside this signal, and a watchdog ceiling of 0 disables stalled classification. The dashboard-only siglake_compactor_delete_tasks_nonterminal{state} gauge is only the last complete observation; it can remain unchanged when execution or observation stops, so it does not establish liveness or current coverage.

A non-zero tasks_already_claimed count also covers the healthy case of a task another executor is running right now. Treat a claim as stranded only after reading the task record, the compactor WARN, and your audit trail. Never decide from claim age or the retained gauge alone.

Siglake has no flag or endpoint that releases a claim. Orphan GC does not reach _siglake/config/, so one small claim object remains for each task. Recovery uses the same explicit resubmission as a failed task. The new task ID gets a new claim object. Deleting a claim object out of band is not a supported step: it re-arms the double execution the claim exists to prevent, and it does not recover a task already stranded in running.

A warehouse without create-only writes fails the sweep

The claim needs a conditional create: the filesystem backend takes O_EXCL, S3 sends If-None-Match: *. If the warehouse store does not advertise that capability, or advertises it and then rejects the write, as an S3-compatible store that ignores If-None-Match does, the sweep returns an error rather than executing the task unclaimed. It names the task it refused. There is no unconditional fallback because one would silently restore double execution.

The check happens as the first pending task is about to run, so a sweep over an index with nothing pending still succeeds on such a store and the problem surfaces only once a deletion request exists. The compactor's multi-index sweep aborts at the first index it cannot claim in, logs delete-task sweep failed, leaves the remaining indexes unswept, and fails the same way every later interval. On that store, acknowledged requests pile up pending with the only evidence in the compactor's log.

Recovering a failed task

The lifecycle is one-way. The executor takes pending tasks only, and nothing returns a terminal task to pending. There is no retry endpoint or automatic retry. A task that reached failed never runs again, even though its submission was acknowledged. The same holds for one left in running by a crash, or by a status write that failed after the rewrite committed, and for one left pending under a claim its executor never released. Nothing sweeps any of them up, and they all recover the same way.

Recovery is an explicit resubmission:

  1. Read the failure with GET /api/v1/delete-tasks/{id}. Check error for the cause where the task has one.

    curl -s localhost:8089/api/v1/delete-tasks/$ID
    
  2. Remedy that cause, and make sure the executor runs the fixed build.

  3. POST the request again under the same tenant identity. Send index_id, predicate_sql, start_ts, end_ts, request fields only, exactly as in the submission above.
  4. Record both IDs. The response carries a new task_id in pending; the failed record and its error are left exactly as they were, so both ids belong in the audit trail for the request.
  5. Confirm execution through the compactor's maintenance sweep or a manual delete-sweep --apply. Poll the new id until it reaches done or failed. The new id carries its own claim, so the stranded one's claim does not hold it back.

A resubmission is not a retry

  • It is not idempotent at the HTTP layer: every POST creates another task.
  • It promises nothing about exactly-once execution.
  • Each task reports only what its own run rewrote, not the request's cumulative effect across attempts.
  • Over an unchanged dataset, with the same predicate and the same fixed bounds, the repeat finishes with rows_deleted: 0, files_rewritten: 0 and commits no snapshot. A time-dependent predicate, or rows that arrived since, makes it delete a different set.

failed does not prove nothing was deleted

A task recorded failed because its commit result was ambiguous, or because the status write failed after the rewrite committed, may have committed that rewrite. Check rows_deleted on the new task against what you expected before concluding the first attempt did nothing.

A task that failed before its commit is the easier case: the whole task commits once. It leaves inactive, unreferenced output files for orphan GC rather than a partly deleted snapshot.

A new task reaching done removes the rows from the live snapshot, and no more: erasure still needs the chain below.

The full sequence to make a delete permanent

Deleting rows rewrites the affected files. The old files still exist and are still referenced by older snapshots. Every step is scoped to the table you deleted from, app-logs in the examples above, not to events.

  1. Execute the delete task. The rows leave the current snapshot.
  2. Let snapshot expiry drop every snapshot taken before the delete.
  3. Check readiness with a table-scoped gc-orphans dry run.
  4. Run orphan GC with --apply against that table and namespace. GC is per table and --table defaults to events, so a scheduled events GC never reclaims anything under app-logs.
  5. Expire the noncurrent object versions. If S3 versioning is on (the Terraform module enables it), deleted objects persist as noncurrent versions until a lifecycle rule expires them.

Step 3 is a dry run:

siglake gc-orphans --namespace siglake --table app-logs

Step 4 is the same command with --apply:

siglake gc-orphans --namespace siglake --table app-logs --apply

Its effect on time travel is absolute. Once step 2 finishes you cannot query the table as of any instant before the delete, and there is no way back. That is what makes the delete permanent, and it is what the full sequence buys you.

Skipping step 5 is the most common compliance gap. Running GC only for events while the deleted rows live in a user index is the next most common.

What goes wrong in steps 2 and 3

Snapshot expiry is on by default (60 s, retain last 100), but retainLast is a count of snapshots, not an age: only newer snapshots displace an older one from the retained window. The 60-second cadence sets no wall-clock erasure deadline. It bounds how promptly a table with more than 100 snapshots is trimmed, nothing more. An idle or low-traffic index, which is what most user indexes are, may sit below the retain count and keep its pre-delete snapshots, and their files, indefinitely.

Siglake has no subcommand, HTTP endpoint, or sql-direct Iceberg metadata table that lists a table's retained snapshots or the files they reference. The intended operator check is the table-scoped gc-orphans dry run in step 3. It reports the aggregate orphan count and bytes only: it does not name files or prove that particular deleted data is no longer referenced. It also skips files modified within --min-age-secs (24 hours by default), so a zero count may mean the superseded files are still retained, or still inside that safety window. Re-run the dry run after the table has taken further commits and the rewritten files have aged past the window. File-level compliance evidence still requires examining the Iceberg metadata and manifests outside Siglake.

Don't narrow retainLast to hurry erasure mid-upgrade

Lowering compactor.snapshotExpire.retainLast does drop the pre-delete snapshots sooner, but during a mixed-version rollout the consumed-proof guard requires retainLast >= 400 until well after the last old writer exits. Read Consumed-proof rolling upgrades before changing it.

Expire snapshots

On by default:

compactor:
  snapshotExpire:
    intervalSecs: 60
    retainLast: 100

It exists for performance, not cleanliness: metadata.json is read and rewritten on every commit, so unbounded snapshot growth slows the entire commit path.

You cannot use Iceberg time travel beyond the retained snapshots. Set intervalSecs: 0 to keep every snapshot, which adds a growing cost to each commit.

Expiry is metadata-only. Expired snapshots' files become orphans.

Erasure is not the only claim on this value: distributed queries need the retained window to outlast a coordinator's stale metadata cache, and a mixed-version rollout needs it at >= 400. Snapshot retention weighs all three consumers against the per-commit cost.

Reclaim orphan files

# dry-run — reports orphan count and bytes
siglake gc-orphans --table events

# apply
siglake gc-orphans --table events --apply

# tighten the safety window
siglake gc-orphans --table events --min-age-secs 3600 --apply
Flag Default Meaning
--table events Table to GC. Run per table.
--min-age-secs 86400 Skip files modified within this window.
--apply off Actually delete.

Don't shorten --min-age-secs casually

The 24-hour default guards against deleting a concurrently-written file that isn't yet referenced by a committed snapshot. Shortening it narrows that guard. Only do so if you understand your write cadence, and never below your longest commit latency.

File and S3 warehouses only.

Audit rotation

query_audit commits per flush, so it accumulates snapshots fast.

# non-destructive: expire old snapshots, keep the rows
siglake audit-rotate --max-age-secs 604800

# preview
siglake audit-rotate --max-age-secs 604800 --dry-run

The bare form is destructive

Without --max-age-secs, audit-rotate drops and recreates the table as a coarse row TTL that deletes all historical audit rows. (Iceberg-rust 0.9 has no public row-level delete.) The recreated table keeps the same schema, partitioning, and sort order.

Use --max-age-secs unless you actually want to discard your audit history.

The running query server picks the table up on its next audit write because ensure_query_audit_table is idempotent.

Via the operator, retention.queryAuditRotateIntervalDays renders a CronJob.

Schedule the complete setup

The compactor runs two of these jobs for you, both on by default: snapshot expiry on its own cadence, and delete-task execution in its idle cycle. Retention sweeps, orphan GC and audit rotation have no in-process scheduler, so they run as CronJobs.

compactor:
  deleteTasks: true # the shipped default; `false` renders SIGLAKE_DELETE_TASKS=0
  snapshotExpire:
    intervalSecs: 60
    retainLast: 100

Plus CronJobs:

Schedule Command Scope
Daily siglake retention-sweep --apply Every managed index in the namespace
Daily siglake gc-orphans --table events --apply events only
Daily siglake gc-orphans --namespace siglake --table app-logs --apply One user index; add a row per index
Weekly siglake audit-rotate --max-age-secs 604800 query_audit

retention-sweep is the only one of these that fans out across the namespace. GC is per table, so the events row covers only events. Each index whose data must leave storage, including every index named in a delete task, needs its own row.

And an S3 lifecycle rule expiring noncurrent versions, without which none of the above erases anything on a versioned bucket.

Always dry-run first

retention-sweep, gc-orphans, and delete-sweep are all dry-run by default, and they print exactly what they would remove. Read that output before adding --apply to a CronJob.