Disaster recovery¶
Use this guide to protect Siglake's three data stores, recover a lost WAL and check that the recovered rows are queryable. Each store fails differently.
| Component | Contents | Loss means |
|---|---|---|
| Object storage | Committed Parquet, manifests, Puffin sidecars | Total data loss. Use bucket versioning and replication. |
| Catalog (Postgres) | Table pointers, snapshots, write-ahead log (WAL) segment claims | The warehouse is unreadable until you restore it. The data is intact but unaddressable. |
| WAL volume | Acknowledged but uncommitted events | Events that have reached neither an Iceberg commit nor a successful mirror upload are lost. |
Ordinary cloud practice covers the first two: S3 versioning and cross-region replication, RDS automated backups and point-in-time recovery. The third is specific to Siglake, and it is the rest of this page.
What WAL mirroring does and does not protect against¶
When you configure a warehouse URL, Siglake copies each sealed WAL segment to
<warehouse bucket>/wal-mirror/ by default. The mirror protects against loss
of the WAL volume: the persistent volume claim (PVC) is deleted, the node
backing it is gone, or the filesystem is corrupt. An event that has reached
neither an Iceberg commit nor a successful mirror upload has no off-volume
copy.
A committed row and a mirrored segment give you different protection. A successful Iceberg commit has already written the rows' Parquet files to the warehouse and published a snapshot that references them, so those rows survive the loss of the WAL volume with no recovery step, whether the mirror is on, off or behind. Their remaining exposure is the object-storage and catalog rows of the table above. Mirrored bytes are acknowledged rows that no commit has reached, and a recovery procedure has to replay them within the limits below. Mirroring covers the window between the acknowledgement and the commit, and nothing after it.
The default wait_for mode syncs the local WAL before acknowledgement. The
sync covers segment bytes and directory entries for the tenant, index, active
segment, and sealed segment. Siglake syncs the sealed name before it unlinks
the active copy.
The power-loss guarantee assumes ext4 or xfs on a node-attached volume. A
network filesystem provides whatever its fsync(2) and rename semantics
guarantee. commit=auto opts into an earlier acknowledgement while bytes can
remain in the kernel page cache. A node crash or power failure can lose those
bytes.
Acknowledgements do not wait for object storage
A successful default acknowledgement means the event is fsynced on the local WAL and stored nowhere else. It gets a copy in object storage at whichever comes first: the Iceberg commit that writes its Parquet file, or the mirror upload of its segment. Losing the WAL volume before either one loses acknowledged data. The commit is also what makes the event visible to queries, unless the query pods mount the WAL buffer.
Configure WAL mirroring¶
wal:
mirror:
enabled: true
prefix: wal-mirror # sibling to s3.warehousePrefix, not underneath
activeIntervalSecs: 0 # only mirror sealed segments
These are the chart defaults. Outside Helm, leaving --wal-mirror-prefix
unset selects wal-mirror whenever you set --warehouse-url.
If you need a shorter loss window for the open segment, set an upload cadence:
To opt out in Helm, set wal.mirror.enabled: false. The chart renders an empty
SIGLAKE_WAL_MIRROR_PREFIX; omitting the variable leaves mirroring on. Outside
Helm, pass an empty prefix:
Sealed segments upload on a best-effort background task as they seal. Active
segments, the ones still open, upload every activeIntervalSecs to
<prefix>/_active/<filename>.
Before queueing a sealed segment, the ingester attempts to make a durable hard
link in the WAL's mirror-pending/ directory. If the link succeeds, local
compaction and retention can remove the segment's other links without removing
the upload source. The ingester removes the pin after it confirms the remote
object. A pin failure increments
siglake_wal_mirror_failures_total{reason="pin"} and leaves the segment without
this extra protection.
For a filesystem drain, successful active uploads set the target for rows that recovery can restore:
activeIntervalSecs |
PVC-loss exposure with successful mirror uploads and filesystem recovery |
|---|---|
0 (default) |
Everything in the open segment since the most recently uploaded sealed segment, so normally up to one full roll interval. |
10 |
About 10 seconds of events, when each scheduled active-segment upload completes. |
Acknowledgements do not wait for mirroring. A delayed or failed upload extends your exposure past the configured interval until a later upload succeeds, so alert on mirror failures as Disaster-recovery readiness describes. The exposure is the acknowledged rows that no commit has reached; rows the drain already committed stay in the warehouse through any length of mirror outage.
For a catalog-claim drain, active uploads
store bytes that reconciliation does not consume. Catalog-claim recovery
reconciles sealed objects only, so activeIntervalSecs does not reduce its
potential row loss.
Each active interval adds one object PUT per ingester. In a loopback test with
a filesystem-backed object store, sealed-segment mirroring raised
acknowledgement p50 by 0.27 ms at 20,000 events per second. Saturation
throughput stayed within run-to-run spread, and uploads stayed within one
sealed segment of the writer. These are local filesystem results, not Amazon
S3 results. They predate the mirror-pending/ pin and do not include its
synchronous hard-link and directory-sync cost. Remote PUT latency and retries
determine the backlog on S3.
A broken mirror is invisible until you need it
Mirroring is best-effort and runs in the background. Ingest can continue
through expired credentials, a bucket policy change or a network policy.
Alert on siglake_wal_mirror_failures_total, and check that
siglake_wal_mirror_segments_total tracks your seal rate. While the remote
failure continues, successful pins retain local bytes and can fill the WAL
volume.
Ingester catch-up¶
Each ingester sweeps its local sealed/ and mirror-pending/ directories once
at startup and every 300 seconds. It uploads any local segment missing from the
mirror, so a retained pin resumes after an outage or process exit. Set
SIGLAKE_WAL_MIRROR_SWEEP_SECS to another interval in seconds. A value of 0
disables both the startup and periodic sweeps. If you disable the sweep, pins
left by an exited uploader remain on disk without automatic retry.
The ingester sweep repairs the local upload backlog. It does not scan the remote prefix for missing catalog rows.
Mirror reconciliation and retention¶
In a catalog-claim deployment, one elected compactor periodically reconciles
the mirror with the shared catalog. This recovery sweep repairs mirrored
objects whose catalog registration was missed. It registers at most 1,024
sealed objects per pass and keeps a shared (last_key, rotation) cursor in the
mirror_sync_cursors catalog table. The cursor advances only after a whole
page registers, survives owner handoff, and clears at end-of-prefix so a later
rotation repairs keys inserted behind it. S3 listings use OpenDAL
start_after; other backends filter the listing locally, which is slower but
correct.
Two warning alerts cover this sweep, and they mean different things.
SiglakeMirrorReconciliationErrors fires with no extra hold when any pass
reports catalog_sync_error in the last 30 minutes: listing the mirror or
registering its objects failed. Inspect the Cycles by outcome panel and the
compactor warning log, fix the object-store or catalog fault it reports, and
let a later pass register the uploaded segments. Until then, segments without
catalog rows stay unqueryable.
SiglakeMirrorReconciliationStalled needs bounded pages to keep running
without one full rotation completing over
prometheusRule.mirrorRotationStallSecs, followed by a 15-minute hold. The
error alert and its outcome series are pre-registered at startup.
| Helm value | Default | Effect |
|---|---|---|
compactor.committedRetentionSecs |
86400 |
Purge successfully drained mirror objects and their catalog rows after 24 hours; 0 disables purging. |
prometheusRule.mirrorRotationStallSecs |
21600 |
Alert after bounded reconciliation pages run for six hours without a full rotation completing; 0 disables the alert. |
This retention runs only in catalog-claim mode. Helm selects that mode with
compactor.catalogClaim.enabled; the operator selects it when
spec.autoscaling.compactor.max is above 1. A filesystem drain never
reclaims mirror objects, so give wal-mirror/ an object-store lifecycle expiry
longer than your worst-case drain backlog. In catalog-claim mode, the 24-hour
default bounds prefix growth. Retention of 0 makes the prefix grow forever.
The compactor floors any non-zero retention at 901 seconds, because it must
outlive the default 600-second local-WAL settle delay plus one 300-second sweep
cadence.
If you raise either local-WAL window, raise retention past their sum. If you
disable the local sweep while mirror catch-up is still enabled, leave committed
retention at 0: otherwise an ingester can re-upload and re-register a segment
after its committed row and mirror object are purged.
Each retention run drains 512-object pages up to a 16,384-object bound. It deletes objects concurrently before batch-deleting their catalog rows, and it is paced start to start, so a long pass does not add another full interval to a backlog.
siglake_compactor_mirror_sync_objects records how many objects each bounded
page examines. It should plateau at 1,024 while a rotation is in progress.
Whole-rotation telemetry is recorded separately:
siglake_compactor_mirror_sync_rotation_duration_secondsandsiglake_compactor_mirror_sync_rotation_objectsare histograms, recorded only when the cursor wraps.siglake_compactor_mirror_sync_rotations_completedandsiglake_compactor_mirror_sync_last_completed_timestamp_secondsare durable gauges that survive owner handoff.
On the dashboard, Mirror repairs / rotation completion shows completion age and completed rotations. When every bounded page stays full, completion age is the backlog or stall signal. A rising retained-per-rotation count in Mirror sync objects per pass, with proportional completed-rotation duration, means retention is backlogged. A flat page count on its own does not prove recovery is converging.
Recovery procedure after the WAL volume is lost¶
This recovery needs a mirror. Determine the drain mode before you copy any
files because each mode has a different recovery path. Under Helm,
compactor.catalogClaim.enabled selects it. Under the operator,
spec.autoscaling.compactor.max above 1 selects the catalog-claim drain and a
maximum of 1 selects the filesystem drain. The current replica count does not
change that choice.
- For the filesystem drain, use the filesystem recovery procedure below.
- For the catalog-claim drain, use the catalog-claim procedure.
siglake wal-recoverwrites local files that this drain never reads.
Recover a filesystem drain¶
This procedure applies to the filesystem drain. Helm selects it with
compactor.catalogClaim.enabled: false; the operator selects it with
spec.autoscaling.compactor.max: 1. It restores every sealed segment the
mirror holds, plus any uploaded active segment snapshots, to the local WAL.
1. Confirm what you lost¶
Check the last successful commit:
Compare that timestamp with the mirror's contents. Record the expected row count or known event identifiers for each affected table and time range.
2. Restore the WAL¶
Provision a new WAL volume, then pull the mirrored segments onto its root. Not
a sealed/ subdirectory: the root.
Recovery rebuilds the <tenant>[/<index>]/sealed/ layout beneath that root, so
the ordinary drain commits each segment to the namespace and table it came
from. --to must therefore be the WAL root the ingester and compactor are
themselves pointed at, --wal /var/lib/siglake/wal on both. Recovering into a
root nothing is serving from leaves the segments outside the drain. Passing
/var/lib/siglake/wal/sealed is refused outright, because the old flattened
spelling would build <wal>/sealed/<tenant>/sealed/, which nothing drains.
The siglake wal-recover command
reference includes the generated help.
The command skips segments already present locally, so it is safe to re-run and safe to interrupt.
--from also reads SIGLAKE_WAL_MIRROR_URL.
3. Let the drain catch up¶
Start the compactor against the restored WAL, with --wal set to the same root
you passed to --to. It claims the recovered segments like any others and
commits them to Iceberg.
Watch siglake_compactor_sealed_pending fall to zero and
siglake_compactor_sealed_pending_oldest_age_seconds fall with it. The same
metrics and alerts that say whether the drain is keeping up in normal operation
say whether this catch-up is progressing: SiglakeDrainBacklogGrowing on the
oldest-segment age, and SiglakeTableNotConverging on the layout the recommit
leaves behind. See Is the WAL drain keeping
up? and Is the layout
converging?. A zero pending gauge
shows that the local queue drained. It does not prove that the recovered rows
are queryable.
4. Verify the recovered rows¶
Query every affected table over the lost time range. For example:
SELECT count(*) AS recovered_rows, min(timestamp), max(timestamp)
FROM <affected_table>
WHERE timestamp >= '<start>' AND timestamp < '<end>';
Compare the result and known event identifiers with the values you recorded in step 1. At-least-once recommit with dedup-by-proof prevents a segment that was partially committed before the loss from double-counting.
Recover a catalog-claim drain¶
This procedure applies to the catalog-claim drain. Helm selects it with
compactor.catalogClaim.enabled: true; the operator selects it when
spec.autoscaling.compactor.max is above 1. Do not restore the mirror to a
local WAL root. The compactor reads catalog rows and mirror objects instead of
local sealed/ files.
- Record the affected tables, lost time range, expected row count and known event identifiers.
- Follow Mirror reconciliation and
retention. Fix the object-store or
catalog fault reported by
SiglakeMirrorReconciliationErrors, then let an elected compactor register sealed mirror objects that have no catalog row. Reconciliation excludes objects under_active/. Treat rows present only in those active snapshots as unrecovered unless you have another copy. - Check that reconciliation rotations resume and that the catalog pending
queue drains. In claim mode,
siglake_compactor_sealed_pendingcounts pending catalog rows. A zero value cannot reveal mirror objects that still lack catalog rows, and it does not prove that recovered rows are queryable. - Run the query from the filesystem procedure for every affected table and lost time range. Compare the result and known event identifiers with the values from step 1. Use these row checks, not restored files or queue gauges, to decide whether recovery succeeded.
What to do in each failure mode¶
An ingester pod dies¶
Nothing. The ingester force-seals its WAL on SIGTERM. After an ungraceful kill,
the partially written active segment is recovered on restart, counted by
siglake_wal_partials_recovered_total.
An ungraceful kill during an append can cut the segment's last Arrow IPC message short. Recovery drains every complete fsynced batch ahead of that message and discards only the incomplete tail. Rows Siglake acknowledged before the crash still reach Iceberg. The segment file keeps its original bytes: nothing truncates or rewrites it, so you can inspect it afterwards.
Each such recovery increments siglake_wal_partial_tail_dropped_total, and a
warning log line names the segment, the rows recovered, and the bytes dropped.
The chart's SiglakeWalPartialTailDropped rule turns that counter into a
warning alert; see Monitoring.
Match it against the restart that caused it.
Everything else about WAL integrity still fails closed. The tolerance covers one case: an active segment recovered as a partial whose final message ends at end of file, after at least one complete batch. Outside that case, damage is still an error:
- A sealed segment whose body fails its CRC check is refused whole and counted
by
siglake_wal_crc_mismatch_total. - A legacy unframed segment stays all-or-nothing.
- A segment written with a frame version this build does not support is refused.
- A partial with no complete batch, or one whose bytes are garbled before end of file, fails to decode instead of recovering a shorter prefix.
The compactor dies mid-commit¶
No acknowledged row is lost and no half-commit lands. Acknowledgement happens in the ingester against the WAL, and an Iceberg commit either becomes a snapshot or does not, so the in-flight commit is all or nothing. What the restart needs from you is usually nothing: both drain shapes recover the in-flight segment from the table's own record of what it has consumed, and ask for an operator only when that record cannot decide.
On the filesystem drain, the segment sits in processing/. Restart moves every
processing/ segment to orphans/ rather than retry it blindly, counting the
move in siglake_compactor_orphans_quarantined_total. The next drain cycle
then disposes of each orphan against the target table's cumulative
consumed-segment set:
| What the table's record shows | What the compactor does |
|---|---|
| The segment's name is in the consumed set. | Its rows are provably committed, so the file is deleted. Recommitting would duplicate them. |
| The name is absent and retained snapshot history reaches back past the seal, or the table has no snapshots at all. | It was provably never committed, so the file is renamed back into sealed/ and committed on the same cycle. |
| The name is absent and snapshot expiry may have dropped the snapshot that would prove it. | Ambiguous, so the orphan is held for you. |
siglake_compactor_orphans_disposed_total{action} counts the deletes and
requeues; siglake_compactor_orphans_held is the gauge that needs you. The
chart has no alert on it, so add your own rule, such as
max(siglake_compactor_orphans_held) > 0 for 15 minutes. Raise
compactor.snapshotExpire.retainLast to preserve
evidence for future incidents. It cannot restore snapshots that have already
expired. Keep an ambiguous segment in orphans/ until available consumption
proof or your comparison against the table establishes its disposition. Delete
it if the comparison proves its rows committed. Move it into sealed/ only if
the comparison proves they did not.
On the catalog-claim drain there is no processing/ directory: the claim row
stays in processing and another drain reclaims it once it is older than
SIGLAKE_CLAIM_RECLAIM_MAX_AGE_SECS (default 900 s), on the
SIGLAKE_CLAIM_RECLAIM_INTERVAL_SECS sweep (default 60 s). The reclaim reads
the same proof first. A claim the proof covers as committed is marked committed
and never redriven, counted by
siglake_catalog_claims_reclaimed_already_committed_total. A claim neither the
durable proof nor retained history covers is still requeued, because duplicate
rows can be found afterwards and unwritten rows cannot, and each one increments
siglake_compactor_reclaim_unprovable_total. That counter is what
SiglakeReclaimWithoutProof fires on, and the remedy is more history rather
than a different guess: see Consumed-proof rolling
upgrades.
A claim the drain keeps failing to commit is a different case. Its row is
quarantined after twelve attempts, which raises
siglake_catalog_claim_quarantined and the SiglakeSegmentsQuarantined
warning. Fix the fault the segment QUARANTINED log line names, then requeue
it.
The catalog is lost¶
The warehouse is intact but unaddressable, because Iceberg metadata pointers
live in the catalog. Restore Postgres from backup. If the restore predates
recent commits, those commits' files become orphans, and siglake gc-orphans
identifies them.
Back the catalog up at least as often as your commit cadence matters.
The WAL volume is lost before upload¶
Any acknowledged event that has reached neither an Iceberg commit nor a successful mirror upload is gone. This includes every uncommitted event when you disable mirroring. The recovery procedure restores only the objects already present in the mirror.
An object-storage object is deleted¶
Enable S3 versioning. Siglake never rewrites an object in place, because compaction writes new files and swaps the manifest, so versioning plus a lifecycle policy gives you a real recovery window.
The backup checklist for Siglake, and the failure modes it does not cover¶
This is what to back up, and what backup cannot save you from.
- S3 bucket versioning enabled
- S3 cross-region replication, if your recovery point objective (RPO) requires it
- RDS automated backups with point-in-time recovery, retention matched to your RPO
-
wal.mirror.enabled: trueconfirmed - For filesystem-drain recovery,
wal.mirror.activeIntervalSecsset to your recovery target - For catalog-claim recovery, your RPO accounts for rows in sealed mirror objects only
- An object-store lifecycle expiry on
wal-mirror/for a filesystem drain - An alert on
siglake_wal_mirror_failures_total - The recovery procedure rehearsed against a non-production cluster
What this does not cover¶
Backup and recovery leave three failure modes to you.
- There is no built-in warehouse backup command. The warehouse is plain Iceberg on object storage, so back it up with your object-storage tooling.
- There is no cross-region failover orchestration. You can replicate the bucket and the catalog, but promoting a standby is a manual procedure you design.
- There is no point-in-time restore of the warehouse beyond retained snapshots.
Snapshot expiry retains the last 100 snapshots by default, and time travel
further back is not available. Widen
snapshotExpire.retainLastif you want more, at the cost of a largermetadata.jsonon every commit.