Disaster recovery¶
Use this guide to protect Siglake's three data stores, recover a lost WAL and check that the recovered rows are queryable. Each store fails differently.
| Component | Contents | Loss means |
|---|---|---|
| Object storage | Committed Parquet, manifests, Puffin sidecars | Total data loss. Use bucket versioning and replication. |
| Catalog (Postgres) | Table pointers, snapshots, write-ahead log (WAL) segment claims | The warehouse is unreadable until you restore it. The data is intact but unaddressable. |
| WAL volume | Acknowledged but uncommitted events | Events that have reached neither an Iceberg commit nor a successful mirror upload are lost. |
Ordinary cloud practice covers the first two: S3 versioning and cross-region replication, RDS automated backups and point-in-time recovery. The third is specific to Siglake, and it is the rest of this page.
What WAL mirroring does and does not protect against¶
When you configure a warehouse URL, Siglake copies each sealed WAL segment to
<warehouse-url>/<wal.mirror.prefix>/ by default, which under the chart is
s3://<s3.bucket>/<s3.warehousePrefix>/wal-mirror/. The mirror protects against
loss of the WAL volume: the persistent volume claim (PVC) is deleted, the node
backing it is gone, or the filesystem is corrupt. An event that has reached
neither an Iceberg commit nor a successful mirror upload has no off-volume
copy.
A committed row and a mirrored segment give you different protection. A successful Iceberg commit has already written the rows' Parquet files to the warehouse and published a snapshot that references them, so those rows survive the loss of the WAL volume with no recovery step, whether the mirror is on, off or behind. Their remaining exposure is the object-storage and catalog rows of the table above. What a recovery replays out of the mirror is narrower: the acknowledged rows no commit has reached, within the limits below. Mirroring covers the window between the acknowledgement and the commit, and nothing after it.
The mirror still holds the segments whose rows did commit. Upload happens at seal, no commit deletes the object, and the copy leaves only when committed retention or a lifecycle rule removes it. See Committed retention purges drained mirror objects for how long those copies live, and Copies the delete chain does not reach for what that means when you delete rows for compliance.
The default wait_for mode syncs the local WAL before acknowledgement. The
sync covers segment bytes and directory entries for the tenant, index, active
segment, and sealed segment. Siglake syncs the sealed name before it unlinks
the active copy.
The power-loss guarantee assumes ext4 or xfs on a node-attached volume. A
network filesystem provides whatever its fsync(2) and rename semantics
guarantee. commit=auto opts into an earlier acknowledgement while bytes can
remain in the kernel page cache. A node crash or power failure can lose those
bytes.
Acknowledgements do not wait for object storage
A successful default acknowledgement means the event is fsynced on the local WAL and stored nowhere else. It gets a copy in object storage at whichever comes first: the Iceberg commit that writes its Parquet file, or the mirror upload of its segment. Losing the WAL volume before either one loses acknowledged data. The commit is also what makes the event visible to queries, unless the query pods mount the WAL buffer.
Configure WAL mirroring¶
Mirroring is on once you set a warehouse URL. These are the chart defaults:
wal:
mirror:
enabled: true
prefix: wal-mirror # relative to the warehouse URL, under
# s3.warehousePrefix, not beside it
activeIntervalSecs: 0 # only mirror sealed segments
The prefix is relative to the warehouse URL, so the mirror root is
s3://<s3.bucket>/<s3.warehousePrefix>/wal-mirror/, one component below the
warehouse URL in your values file. That path is what an object-store lifecycle
rule names and what siglake wal-recover --from takes. Outside Helm, leaving
--wal-mirror-prefix unset selects wal-mirror whenever you set
--warehouse-url.
If you need a shorter loss window for the open segment, set an upload cadence:
To opt out in Helm, set wal.mirror.enabled: false. The chart renders an empty
SIGLAKE_WAL_MIRROR_PREFIX; omitting the variable leaves mirroring on. Outside
Helm, pass an empty prefix:
What the mirror uploads, and when¶
Sealed segments upload on a best-effort background task as they seal. Active
segments, the ones still open, upload every activeIntervalSecs to
<prefix>/_active/<tenant>[/<index>]/<segment>.arrow.partial, one object per
open writer. A writer whose segment has not grown since its last upload is
skipped, so an idle tenant costs nothing on a tick. Active uploads need both
an enabled mirror and activeIntervalSecs above 0; outside Helm the same
interval is SIGLAKE_WAL_ACTIVE_MIRROR_INTERVAL_SECS, and 0 leaves them off.
The order is upload, then delete. From Siglake 0.2.0 the ingester removes a
segment's exact _active/ partial once the sealed object is confirmed. Any of
three paths confirms it: the sealed upload, the check that resolves an
ambiguous upload error, or the catch-up sweep. An active upload that completes
after the seal checks for the sealed object and removes its own partial, so
neither ordering leaves one behind. Each removal gets three attempts. A third
failure increments
siglake_wal_mirror_failures_total{reason="active_cleanup"}, logs a warning,
and leaves the sealed segment registered. Cleanup removes the single object it
derives from the sealed name and never lists _active/.
Before queueing a sealed segment, the ingester attempts to make a durable hard
link in the WAL's mirror-pending/ directory. If the link succeeds, local
compaction and retention can remove the segment's other links without removing
the upload source. The ingester removes the pin after it confirms the remote
object. A pin failure increments
siglake_wal_mirror_failures_total{reason="pin"} and leaves the segment without
this extra protection.
Your PVC-loss exposure window with active-segment uploads¶
activeIntervalSecs decides how much acknowledged data a lost WAL volume takes
with it, as long as the uploads succeed. For a filesystem
drain,
successful active uploads set the target for rows that recovery can restore:
activeIntervalSecs |
PVC-loss exposure with successful mirror uploads and filesystem recovery |
|---|---|
0 (default) |
Everything in the open segment since the most recently uploaded sealed segment, so normally up to one full roll interval. |
10 |
About 10 seconds of events, when each scheduled active-segment upload completes. |
Acknowledgements do not wait for mirroring. A delayed or failed upload extends your exposure past the configured interval until a later upload succeeds, so alert on mirror failures as Disaster-recovery readiness describes. The exposure is the acknowledged rows that no commit has reached; rows the drain already committed stay in the warehouse through any length of mirror outage.
For a catalog-claim drain, active uploads
store bytes that reconciliation does not consume. Catalog-claim recovery
reconciles sealed objects only, so activeIntervalSecs does not reduce its
potential row loss.
A broken mirror is invisible until you need it
Mirroring is best-effort and runs in the background. Ingest can continue
through expired credentials, a bucket policy change or a network policy.
Alert on siglake_wal_mirror_failures_total, and check that
siglake_wal_mirror_segments_total tracks your seal rate. While the remote
failure continues, successful pins retain local bytes and can fill the WAL
volume.
What mirroring costs at the ingester¶
Each active interval adds one object PUT per open writer that took rows since
the last tick, summed over your ingesters. One ingester holds one writer per
(tenant, managed index, ingester.backpressure.shards) combination, so a
busy multi-tenant fleet pays many PUTs a tick, not one. In a loopback test with
a filesystem-backed object store, sealed-segment mirroring raised
acknowledgement p50 by 0.27 ms at 20,000 events per second. Saturation
throughput stayed within run-to-run spread, and uploads stayed within one
sealed segment of the writer. These are local filesystem results, not Amazon
S3 results. They predate the mirror-pending/ pin and do not include its
synchronous hard-link and directory-sync cost. Remote PUT latency and retries
determine the backlog on S3.
Ingester catch-up¶
Each ingester sweeps its local sealed/ and mirror-pending/ directories once
at startup and every 300 seconds. It uploads any local segment missing from the
mirror, so a retained pin resumes after an outage or process exit. Set
SIGLAKE_WAL_MIRROR_SWEEP_SECS to another interval in seconds. A value of 0
disables both the startup and periodic sweeps. If you disable the sweep, pins
left by an exited uploader remain on disk without automatic retry.
The ingester sweep repairs the local upload backlog. It does not scan the remote prefix for missing catalog rows.
Mirror reconciliation and retention¶
In a catalog-claim deployment, one elected compactor periodically reconciles
the mirror with the shared catalog. This recovery sweep repairs mirrored
objects whose catalog registration was missed. It registers at most 1,024
sealed objects per pass and keeps a shared (last_key, rotation) cursor in the
mirror_sync_cursors catalog table. The cursor advances only after a whole
page registers, survives owner handoff, and clears at end-of-prefix so a later
rotation repairs keys inserted behind it. S3 listings use OpenDAL
start_after; other backends filter the listing locally, which is slower but
correct.
Two Helm values set how long the sweep's registered objects are kept and when a sweep that is not completing rotations alerts:
| Helm value | Default | Effect |
|---|---|---|
compactor.committedRetentionSecs |
86400 |
Purge successfully drained mirror objects and their catalog rows after 24 hours; 0 disables purging. |
prometheusRule.mirrorRotationStallSecs |
21600 |
Alert after bounded reconciliation pages run for six hours without a full rotation completing; 0 disables the alert. |
Alerts on failed or stalled mirror reconciliation¶
Two warning alerts cover this sweep, and they mean different things.
SiglakeMirrorReconciliationErrors fires with no extra hold when any pass
reports catalog_sync_error in the last 30 minutes: listing the mirror or
registering its objects failed. Inspect the Cycles by outcome panel and the
compactor warning log, fix the object-store or catalog fault it reports, and
let a later pass register the uploaded segments. Until then, segments without
catalog rows stay unqueryable.
SiglakeMirrorReconciliationStalled needs bounded pages to keep running
without one full rotation completing over
prometheusRule.mirrorRotationStallSecs, followed by a 15-minute hold. The
error alert and its outcome series are pre-registered at startup.
Committed retention purges drained mirror objects¶
Committed retention deletes mirror objects the drain has committed. The
catalog-claim drain runs it as described above: Helm selects that mode with
compactor.catalogClaim.enabled, the operator selects it when
spec.autoscaling.compactor.max is above 1, and there the 24-hour default
bounds prefix growth. Retention of 0 makes the prefix grow forever.
A filesystem drain reclaims nothing by default. It commits out of local
sealed/ and never reads the mirror, so the prefix and its wal_segments rows
grow for as long as the cluster ingests. Setting
compactor.mirrorLedgerReclaim: true closes that for the segments this drain
commits: it marks the ingester's row for each one and lets the same retention
pass delete the object and then the row. Marking needs three inputs:
SIGLAKE_CATALOG_URI, SIGLAKE_WAREHOUSE_URL and a non-empty mirror prefix.
Without all three the compactor warns and keeps draining. The chart renders the
first two on every pod, and both are required for any install: the catalog URI
comes from postgres.existingSecret, read through the field names in
postgres.secretKeys, and the warehouse URL from s3.bucket with
s3.warehousePrefix. So a chart install already has them. The prefix is the
input you can miss. The chart passes wal.mirror.prefix to the compactor only
in catalog-claim mode, and a filesystem drain reclaims under the binary's own
default, wal-mirror. If you changed wal.mirror.prefix, repeat it to the
compactor:
Otherwise the compactor records the default prefix for any segment it commits before the ingester's upload registers, and retention then deletes an object path that nothing wrote and leaves the upload behind. Reclamation never marks an object it did not commit, so a dropped index incarnation's quarantined segments, and the uploads of an ingester whose volume was lost before they drained, stay in the prefix.
Under either drain, give <s3.warehousePrefix>/<wal.mirror.prefix>/ an
object-store lifecycle expiry longer than your worst-case drain backlog. It is
the only thing that removes what the drain cannot.
No drain reclaims _active/. Only a freshly sealed segment gets a
wal_segments row, and reconciliation skips the partial blobs, so retention
never sees them. The uploader collects them instead: from Siglake 0.2.0 each
sealed segment's own partial is deleted once the sealed object is confirmed,
so seals no longer grow the prefix by one blob each. Two populations are left
for the lifecycle rule. Partials written before this cleanup shipped are never
revisited, because cleanup derives one object name rather than listing the
prefix. So are partials whose three deletes all failed, once the matching local
segment is gone and no catch-up sweep can retry them.
Choose a committed retention window¶
The compactor floors any non-zero compactor.committedRetentionSecs at 901
seconds, because it must outlive the default 600-second local-WAL settle delay
plus one 300-second sweep cadence. If you raise either local-WAL window, raise
retention past their sum.
If you disable the local sweep while mirror catch-up is still enabled, leave
committed retention at 0: otherwise an ingester can re-upload and re-register
a segment after its committed row and mirror object are purged.
Each retention run drains 512-object pages up to a 16,384-object bound. It deletes objects concurrently before batch-deleting their catalog rows, and it is paced start to start, so a long pass does not add another full interval to a backlog.
Mirror reconciliation telemetry and dashboard panels¶
siglake_compactor_mirror_sync_objects records how many objects each bounded
page examines. It should plateau at 1,024 while a rotation is in progress.
Whole-rotation telemetry is recorded separately:
siglake_compactor_mirror_sync_rotation_duration_secondsandsiglake_compactor_mirror_sync_rotation_objectsare histograms, recorded only when the cursor wraps.siglake_compactor_mirror_sync_rotations_completedandsiglake_compactor_mirror_sync_last_completed_timestamp_secondsare durable gauges that survive owner handoff.
On the dashboard, Mirror repairs / rotation completion shows completion age and completed rotations. When every bounded page stays full, completion age is the backlog or stall signal. A rising retained-per-rotation count in Mirror sync objects per pass, with proportional completed-rotation duration, means retention is backlogged. A flat page count on its own does not prove recovery is converging.
For a filesystem drain with compactor.mirrorLedgerReclaim enabled,
siglake_compactor_mirror_marked_total counts locally committed segments whose
catalog marks became durable. siglake_compactor_mirror_mark_errors_total
counts failed marking calls; restore the catalog before the 3,600-second local
retention ceiling. If that ceiling removes the local evidence first,
siglake_compactor_mirror_unreclaimed_total counts the mirror objects this
drain can no longer reclaim. The counter does not identify those objects, so
remove them through the mirror prefix's object-store lifecycle rule.
Recovery procedure after the WAL volume is lost¶
This recovery needs a mirror. Determine the drain mode before you copy any
files because each mode has a different recovery path. Under Helm,
compactor.catalogClaim.enabled selects it. Under the operator,
spec.autoscaling.compactor.max above 1 selects the catalog-claim drain and a
maximum of 1 selects the filesystem drain. The current replica count does not
change that choice.
- For the filesystem drain, use Recover a filesystem drain.
- For the catalog-claim drain, use Recover a catalog-claim
drain.
siglake wal-recover --applywrites local files that this drain never reads.
Recover a filesystem drain¶
This procedure applies to the filesystem drain. Helm selects it with
compactor.catalogClaim.enabled: false; the operator selects it with
spec.autoscaling.compactor.max: 1. It restores every sealed segment the
mirror holds, plus any uploaded active segment snapshots, to the local WAL.
Step 1: confirm what you lost¶
Check the last successful commit:
Compare that timestamp with the mirror's contents. Record the expected row count or known event identifiers for each affected table and time range.
Prerequisite: isolate the replacement WAL from every drain¶
Keep the replacement WAL offline from Siglake until you finish the tenant
checks in Step 7. Mount it only in the recovery environment that runs
siglake wal-recover. Stop every standalone compactor and every ingest server
started with --with-compactor before you expose the volume. Keep ordinary
ingesters off it too, so new writes cannot change the directories.
Stopping a process or deleting a pod is not enough when its controller can
start another one. Disable restarts or pause reconciliation before you mount
the replacement WAL. For an operator-managed cluster, pause the operator and
stop the generated ingester and compactor Deployments directly. Do not set
spec.autoscaling.compactor.min or spec.autoscaling.compactor.max to 0.
A supported zero floor is an autoscaling policy, not a controller pause, and
the operator still runs scheduled maintenance wake-ups.
Before you plan, check both sides of the fence. The storage control plane must show the replacement volume attached only to the recovery environment. The workload control plane must show no running or pending pod or process that can mount that volume and run a filesystem drain. Recheck both after any controller or node restart. Keep this isolation in place through Step 7.
Step 2: plan the restore onto the new WAL volume¶
Provision a new WAL volume, then plan the pull onto its root. Not a sealed/
subdirectory: the root. --from must be the mirror root,
<warehouse-url>/<wal.mirror.prefix>/, not the warehouse URL in your values
file. Under the chart that is
s3://<s3.bucket>/<s3.warehousePrefix>/wal-mirror, with wal-mirror as the
default prefix:
if ! wal_recover_help=$(siglake wal-recover --help 2>&1); then
printf '%s\n' "$wal_recover_help" >&2
printf '%s\n' 'Refusing recovery: wal-recover --help failed.' >&2
exit 1
fi
if ! printf '%s\n' "$wal_recover_help" | grep -Eq '^[[:space:]]*--apply([[:space:]]|$)'; then
printf '%s\n' 'Refusing recovery: this wal-recover has no --apply option.' >&2
exit 1
fi
siglake wal-recover \
--from s3://my-siglake-warehouse/warehouse/wal-mirror \
--to /var/lib/siglake/wal
Without --apply the command plans. It lists the mirror, reconstructs the
layout, prints one line per tenant and index, and creates nothing under --to,
--to itself included. The plan costs one listing, plus one read of each
candidate it would write: a body that does not decode as a WAL segment is
refused and named in the plan rather than after the restore.
This two-command form arrives in Siglake 0.2.0, so the planning step needs
0.2.0 or later. In Siglake 0.1.0, the bare siglake wal-recover command
restores files immediately. It does not plan. If your binary has no --apply
option, use the frozen Siglake 0.1.0 recovery
procedure.
If your catalog survived the loss, plan again with --catalog. The flag settles
the same question the root verdict answers from
markers, and it is the check for a default install, which carries no marker
either way. Pass the URI explicitly: --catalog reads no environment variable,
so a SIGLAKE_CATALOG_URI in the shell does not turn it on.
siglake wal-recover \
--from s3://my-siglake-warehouse/warehouse/wal-mirror \
--to /var/lib/siglake/wal \
--catalog sqlite:///var/lib/siglake/catalog.db
Your catalog is eligible if the ingester that uploaded the objects ran with a
catalog URI of its own. That ingester records every object it uploads in
wal_segments, with the tenant and index it wrote the segment for, and those
rows are the evidence --catalog reads. The chart gives every pod
SIGLAKE_CATALOG_URI, so a Helm install qualifies and the URI to pass is its
postgres:// one. The Postgres path has no live test coverage yet: its
statements are parse-checked in the Postgres dialect, and the tested cases are
SQLite. --catalog arrives in Siglake 0.2.0 with the two-command form.
Step 3: check the tenant destinations in the plan¶
The plan prints one line per tenant and index, with the directory each group
restores into. Read those destinations before anything else: every tenant=
value is a directory name the mirror carried, so check it against the tenants
your ingesters route to.
plan for s3://my-siglake-warehouse/warehouse/wal-mirror -> /var/lib/siglake/wal
tenant=acme index=- 412 segments 3.1 GiB -> /var/lib/siglake/wal/acme/sealed/ (e.g. `_active/acme/siglake-ingester-0-01a0afc1-283c-7bf0-b100-ac30fd8d0bd2.arrow.partial`)
tenant=default index=- 18 segments 96.0 MiB -> /var/lib/siglake/wal/default/sealed/ (e.g. `default/siglake-ingester-0-01a0afc1-287c-7113-a219-b89b7422d276.arrow`)
tenant=payments index=http_requests 37 segments 284.0 MiB -> /var/lib/siglake/wal/payments/http_requests/sealed/ (e.g. `payments/http_requests/siglake-ingester-1-01a0afc1-28b9-7622-9b17-9b0d73d24726.arrow`)
UNREADABLE: `default/siglake-ingester-0-01a0afc1-2901-7cc2-8b41-6f5f0f4fd83e.arrow` is not a WAL segment (WAL frame integrity check: <object-store segment> body length 972 != header 625); it will be left in the mirror
totals: 467 segments, 3.5 GiB (1 unreadable: body does not decode, 17 keys skipped: unrecognised layout (e.g. `mirror-inventory.csv`))
root confirmed by `_active/acme/siglake-ingester-0-01a0afc1-283c-7bf0-b100-ac30fd8d0bd2.arrow.partial`, a mirror marker at its own depth
nothing written. Re-run with --apply to restore into /var/lib/siglake/wal
A tenant= value you do not route to stops the restore until you have
accounted for it. The mirror prefix, wal-mirror by default, is the name to
watch for, and the name on its own settles nothing: see Step 4: read the root
verdict. tenant=default is what a
single-tenant ingester produces, and also what a mirrored object whose name
carries no tenant component restores as. An index= value names a managed user
index; index=- is that tenant's events stream.
Ingest takes the tenant from a verified JWT claim where
ingester.oidc.tenantClaim is set, from X-Scope-OrgID where
ingester.trustScopeHeader is set, and otherwise from the single-tenant
default. It does not take it from the Iceberg namespace: tenant.namespace
and SIGLAKE_TENANT_NAMESPACE pin a release to one namespace and select
nothing at ingest. See Namespaces and tenancy
and Ingest tenant selection. Where
ingester.allowedTenants is set, that list is what to compare the plan
against.
Byte totals are whatever the listing carried, so a store that reports no sizes
in a listing prints size unknown on each line, and size unknown (this store
reports no sizes in a listing) on the totals: line.
The (e.g. <key>) sample is the group's first object name in sort order,
spelled relative to --from. A group whose in-flight segment reached the
mirror therefore samples the _active/ object. A listing that caught a
segment between its sealed upload and the delete of its partial retains only
that partial, and the restore still writes the sealed copy: the plan derives
the sealed object name, both the plan's body check and the apply GET prefer
it, and the GET retries it if cleanup removed the listed partial in between.
Segment names are <ingester-id>-<uuid>.arrow, where the ingester id is SIGLAKE_INGESTER_ID or
the ingester's hostname. The totals: line carries the skipped count, and one
refused name, whenever the listing held keys the layout does not have. Plan a
partly restored volume and each line gains [N already present], with an
already present count on totals:. The plan still writes nothing.
An object whose name fits the layout but whose body does not decode gets its
own UNREADABLE: line and an unreadable count on totals:. Read What an
UNREADABLE: line in the plan
means before you apply.
What an UNREADABLE: line in the plan means¶
An UNREADABLE: line names a candidate the plan read and could not decode as a
WAL segment. The plan lists the mirror, then reads the body of every candidate
it would write, so a torn object is named here rather than after the restore.
That object is left in the mirror, it is not counted in N segments, and
--apply does not pull it.
UNREADABLE: `default/siglake-ingester-0-01a0afc1-2901-7cc2-8b41-6f5f0f4fd83e.arrow` is not a WAL segment (WAL frame integrity check: <object-store segment> body length 972 != header 625); it will be left in the mirror
totals: 467 segments, 3.5 GiB (1 unreadable: body does not decode, 17 keys skipped: unrecognised layout (e.g. `mirror-inventory.csv`))
Ten names print, then ... and N more unreadable candidates. The command
publishes no metric for a refusal. The printed lines are the signal, together
with a WARN event per refused object on stderr; see Unreadable objects found by a
wal-recover run.
Tell this apart from a skip. A skip counts a name the mirror layout does not have, and the plan never reads that object's body. An unreadable candidate has a name the layout does have and a body that does not decode.
Nothing reclaims a refused object. Mirror reclamation only reaches segments the local drain commits, and a refused candidate never lands locally, so the count does not fall when you plan again. Copy it somewhere you can inspect it, then delete it from the mirror yourself.
Where the verdict is root CONTRADICTED, no body is read at all. That plan
costs one listing, prints no UNREADABLE: line, and exits nonzero.
When every candidate is unreadable: the run exits nonzero¶
A run that can restore nothing exits nonzero. That is a mirror whose every
candidate is unreadable, with no segment already present under --to. Both the
plan and --apply refuse on the plan, before any write, and the --apply does
not create --to:
plan for s3://my-siglake-warehouse/warehouse/wal-mirror -> /var/lib/siglake/wal
UNREADABLE: `_active/acme/siglake-ingester-0-01a0afc1-283c-7bf0-b100-ac30fd8d0bd2.arrow.partial` is not a WAL segment (StreamReader::try_new: Ipc error: Expected schema message, found empty stream.); it will be left in the mirror
totals: 0 segments, size unknown (this store reports no sizes in a listing) (1 unreadable: body does not decode)
root confirmed by `_active/acme/siglake-ingester-0-01a0afc1-283c-7bf0-b100-ac30fd8d0bd2.arrow.partial`, a mirror marker at its own depth
Error: restored nothing: 1 candidate body does not decode as a WAL segment under --from (e.g. `_active/acme/siglake-ingester-0-01a0afc1-283c-7bf0-b100-ac30fd8d0bd2.arrow.partial`); no recoverable candidate remains. The unreadable objects were left in the mirror, and /var/lib/siglake/wal is unchanged.
Where the same listing also held names the layout does not have, the refusal
adds , and N other keys have an unrecognised layout. It still names the
unreadable bodies as the reason, because those are the objects that would have
restored.
A partial restore is a success. If one candidate decodes, or one segment is already present from an earlier run, the command restores what it can, reports the unreadable count, and exits 0. An empty mirror is a success too: nothing listed means no candidate to refuse.
An --apply can fail this way after its own plan passed, if a body stopped
decoding between the plan's read and the restore's GET. That run prints the
pulled 0 segments summary line first, and its refusal ends no WAL segment
was written under /var/lib/siglake/wal in place of is unchanged.
Step 4: read the root verdict¶
The last line of the plan is the root verdict, read off the same listing. It
says whether --from names the mirror root. With --catalog, a second line
follows it with what wal_segments says, and the marker line stays first
because it refuses on its own: a contradicting marker refuses the restore
whatever the catalog says.
| Verdict | What it means | What to do |
|---|---|---|
root confirmed by <name> |
A marker Siglake writes itself, an _active/ object or a <tenant>/<index>/owner file, sits at the depth the mirror layout puts it. --from is the mirror root. |
Go to Step 5: apply the restore. |
root CONTRADICTED |
The same marker sits one component deeper, so --from is one component above the mirror root. The plan and --apply both exit nonzero, and neither creates --to. |
Re-run with the directory the refusal names. There is no override flag. |
root unverified |
The mirror carries neither marker, so nothing in the listing pins the root. See What root unverified means. |
Plan again with --catalog, as Step 2 shows. With no catalog, decide from the tenant destinations. |
root confirmed by the catalog |
At least one listed object's row routes exactly as its object name does, and the matched rows agree on one mirror prefix. --from is the mirror root. |
Go to Step 5: apply the restore. |
root CONTRADICTED by the catalog |
A matched row disagrees with the object's own name, or two rows claim different mirror prefixes. The plan and --apply both exit nonzero and restore nothing. |
Re-run with the directory the refusal names, or read what --catalog does not certify. |
catalog reachable and silent |
No listed object has a row, so the catalog adds nothing. Retention deletes a row with its object, so your own purged mirror and another deployment's mirror read alike. | Decide from the tenant destinations. |
catalog could not be read |
The catalog is missing, is not a Siglake catalog, or the mount refused it. The run exits nonzero and writes nothing under --to. |
Fix what the message names, or drop --catalog and restore on the listing's evidence. |
On a default install whose catalog survived, the two lines read like this:
root unverified: this mirror carries no `_active/` object and no `<tenant>/<index>/owner` marker, so nothing in the listing pins the root. Read the plan
root confirmed by the catalog: 16 listed segment(s) match a wal_segments row and every one routes as its key does, under the mirror prefix `wal-mirror` (e.g. `default/siglake-ingester-0-01a0afc1-287c-7113-a219-b89b7422d276.arrow`). 2 listed segment(s) have no row and keep the routing their key implies (e.g. `default/siglake-ingester-1-01a0afc1-28b9-7622-9b17-9b0d73d24726.arrow`)
What wal-recover --catalog does not certify¶
wal-recover --catalog certifies the root, not every object. One listed object
whose wal_segments row routes as its object name does settles where --from
points. Objects with no row keep the routing their name implies and count as
uncertified in the plan, the ordinary case: retention deletes a row with its
object. The check is read-only mechanically: SQLite is opened mode=ro,
Postgres runs inside START TRANSACTION READ ONLY, and no schema is created in
the catalog you inspect. What it will not do:
- A disagreement refuses the restore whole. It reports both routings and reroutes nothing.
- A catalog it cannot read fails the run. It never falls back to the marker verdict; drop the flag instead.
- A WAL-journal SQLite catalog on a read-only mount needs
?immutable=1, which reads around the-walsidecar and is exact only for a catalog nothing is writing. Siglake's own catalogs are rollback-journal and need none of it. - One false refusal is known: a
wal.mirror.prefixchanged after some objects were uploaded leaves rows claiming two prefixes for one mirror. The refusal names both; drop--catalogto get past it.
What root unverified means on a default install¶
root unverified means nothing in the listing pins the root. It is the verdict
a default install prints: no managed index, and
wal.mirror.activeIntervalSecs at its default of 0, so the mirror holds no
marker either way. That install is what --catalog
answers: the rows the uploader wrote pin the root where the listing cannot.
Without a catalog, the destinations in the plan are the whole check.
That mirror, listed one component too high, is indistinguishable from a real
mirror whose first tenant happens to be named after a prefix. Recovery reads
<mirror-dir>/<tenant>/<segment>.arrow as a valid two-level name, takes the
mirror directory for the tenant and the tenant for the index, and the counts
and the exit status both look healthy.
unverified is not a statement that the root is right.
To get a verdict instead, give the mirror a marker before you need one. Set
activeIntervalSecs above 0 and the first successful active upload writes
the _active/ object that confirms the root. A managed index writes
<tenant>/<index>/owner whatever the interval is.
When the plan skips every name: --from is above the mirror root¶
A plan in which no name fits the layout exits 1. That is --from two or more
components above the mirror root: every object name carries path components the
layout does not have. The plan prints on stdout with the counts on its
totals: line and one refused name quoted, and the refusal itself prints on
stderr:
plan for s3://my-siglake-warehouse -> /var/lib/siglake/wal
totals: 0 segments, size unknown (this store reports no sizes in a listing) (812 keys skipped: unrecognised layout (e.g. `warehouse/wal-mirror/acme/siglake-ingester-0-01a0afc1-283c-7bf0-b100-ac30fd8d0bd2.arrow`))
root unverified: this mirror carries no `_active/` object and no `<tenant>/<index>/owner` marker, so nothing in the listing pins the root. Read the plan
Error: restored nothing: all 812 keys under --from have a layout recovery will not guess at (e.g. `warehouse/wal-mirror/acme/siglake-ingester-0-01a0afc1-283c-7bf0-b100-ac30fd8d0bd2.arrow`). --from must name the MIRROR ROOT — the directory holding <tenant>[/<index>]/<segment>.arrow and _active/ — not an ancestor of it (…/warehouse rather than …/warehouse/wal-mirror). /var/lib/siglake/wal is unchanged.
The URL to pass is the mirror root itself, the directory that holds
<tenant>[/<index>]/<segment>.arrow and _active/. For a warehouse at
s3://my-siglake-warehouse/warehouse, that is
s3://my-siglake-warehouse/warehouse/wal-mirror. --from also reads
SIGLAKE_WAL_MIRROR_URL.
Step 5: apply the restore¶
Re-run the same command with --apply, and with the same --catalog you
planned with. It takes its own listing and prints the plan again, so a mirror
that changed between the two runs is reported before the first write, and the
catalog check runs again on that fresh listing.
siglake wal-recover \
--from s3://my-siglake-warehouse/warehouse/wal-mirror \
--to /var/lib/siglake/wal \
--catalog sqlite:///var/lib/siglake/catalog.db \
--apply
Recovery rebuilds the <tenant>[/<index>]/sealed/ layout beneath that root, so
the ordinary drain commits each segment to the namespace and table it came
from. --to must therefore be the WAL root the ingester and compactor are
themselves pointed at, --wal /var/lib/siglake/wal on both. Recovering into a
root nothing is serving from leaves the segments outside the drain. Passing
/var/lib/siglake/wal/sealed is refused outright, because the old flattened
spelling would build <wal>/sealed/<tenant>/sealed/, which nothing drains.
The siglake wal-recover command
reference includes the generated help.
The command skips segments already present locally, so it is safe to re-run and safe to interrupt.
Step 6: read the restore summary line and its exit status¶
The restore prints one summary line after the plan. That line and its exit status are what say whether the segments landed:
pulled 464 segments into /var/lib/siglake/wal (3 already present, 17 keys skipped: unrecognised layout)
| Summary line | Exit status | What it means |
|---|---|---|
pulled N segments into <wal-root> with N above zero, with or without counts in parentheses |
0 | Those N segments are on the volume, each one fsynced under its final name before it was counted. Go to Step 7: check the restored tenant directories. |
pulled 0 segments into <wal-root> (K already present) |
0 | A re-run with nothing left to pull. Go to Step 7: check the restored tenant directories. |
pulled 0 segments into <wal-root> with no counts at all |
0 | No file was listed under --from. Check the URL, and check that the compactor was mirroring. |
either of the first two rows with K unreadable: body does not decode as a WAL segment, left in the mirror (e.g. <key>) |
0 | K objects were read and refused, and at least one segment was pulled or is already present. The K are still in the mirror and there is no local copy of them. See What an UNREADABLE: line in the plan means. |
pulled 0 segments into <wal-root> (K unreadable: …) with no segment pulled and none already present |
nonzero | Nothing was restored. Only an --apply whose bodies decoded at plan time and not at GET time prints this line: see When every candidate is unreadable. |
Each count in parentheses appears only when it is nonzero, so a first restore
from a mirror with nothing foreign in it prints the count and the root it wrote
to, and no parentheses at all. A skip is a file under --from whose name does
not fit the mirror layout, whether that is another file kept in the same prefix
or a segment name with more path components than the layout has. An
unreadable count is a different case: those names fit and their bodies did not
decode. The restore leaves them behind, and it still exits 0 as long as it
pulled a segment or found one already present. Plan again and the same count
comes back.
A run whose root is contradicted, whose every name was skipped, or whose every candidate is unreadable with no segment already present stops on the plan and exits nonzero without writing, so none of the three prints this line.
Step 7: check the restored tenant directories¶
Compare the top-level directories under the --to root with the tenants your
ingesters route to, the same set you checked the plan against. This listing
confirms what actually landed, and it is the last check before the drain acts
on it.
The tenant destinations in the
plan are where you catch a
wrong --from before a byte is written.
Recovery creates one directory per tenant in the plan and writes the segments
to <tenant>/sealed/, or to <tenant>/<index>/sealed/ for a managed index. An
index restore also creates that tenant's own sealed/ directory, because the
drain enumerates a tenant by that directory alone. default is a tenant
directory like any other: it holds the single-tenant default's segments, and
anything the mirror carried under a name with no tenant component.
Lifecycle directories are not tenants. Inside a tenant directory, the
ingester's active/ and the drain's processing/, committed/ and
consumers/ appear during normal operation, along with the orphans/,
poison/ and stale/ quarantines. At the top level, a volume that also served
a pre-tenancy build
can carry active/, sealed/, processing/, committed/ and consumers/
from the flat layout. Siglake refuses those five names as tenant ids, and the
drain's tenant walk skips them rather than reading them as tenants.
When a restored top-level directory is not a tenant you route to¶
A top-level name that is neither a tenant nor a lifecycle directory is a reason to stop, not proof of a wrong root. Leave the compactor stopped and confirm all three of these before you remove anything:
- Check the routing. The name is not a tenant any ingester resolves, judged
against
ingester.oidc.tenantClaim,ingester.trustScopeHeaderandingester.allowedTenantsrather than against the shape of the name. - Re-run the plan with the corrected
--from. It accounts for the same segments, under the tenants you expect. - Confirm the mirror still holds those objects. Recovery only reads the mirror, but a lifecycle expiry on the prefix can remove them behind it.
When all three agree, remove the directory this restore created, correct
--from, and run the plan and the restore again. The command skips segments
already present, so a re-run costs only the transfer. If any of the three is
unresolved, or the directory predates this restore, move it aside instead of
deleting it and keep the compactor stopped: an uncommitted segment can exist
nowhere else.
Step 8: let the drain catch up¶
After Step 7 passes, unmount the WAL from the recovery environment and attach
it to the normal workload. Restore the workload controller or restart policy,
then start the compactor with --wal set to the same root you passed to
--to. It claims the recovered segments like any others and commits them to
Iceberg.
Watch siglake_compactor_sealed_pending fall to zero and
siglake_compactor_sealed_pending_oldest_age_seconds fall with it. The same
metrics and alerts that say whether the drain is keeping up in normal operation
say whether this catch-up is progressing: SiglakeDrainBacklogGrowing on the
oldest-segment age, and SiglakeTableNotConverging on the layout the recommit
leaves behind. See Is the WAL drain keeping
up? and Is the layout
converging?. A zero pending gauge
shows that the local queue drained. It does not prove that the recovered rows
are queryable.
Step 9: verify the recovered rows¶
Query every affected table over the lost time range. For example:
SELECT count(*) AS recovered_rows, min(timestamp), max(timestamp)
FROM <affected_table>
WHERE timestamp >= '<start>' AND timestamp < '<end>';
Compare the result and known event identifiers with the values you recorded in Step 1: confirm what you lost. At-least-once recommit with dedup-by-proof prevents a segment that was partially committed before the loss from double-counting.
Recover a catalog-claim drain¶
This procedure applies to the catalog-claim drain. Helm selects it with
compactor.catalogClaim.enabled: true; the operator selects it when
spec.autoscaling.compactor.max is above 1. Do not restore the mirror to a
local WAL root. The compactor reads catalog rows and mirror objects instead of
local sealed/ files.
- Record the affected tables, lost time range, expected row count and known event identifiers.
- Follow Mirror reconciliation and
retention. Fix the object-store or
catalog fault reported by
SiglakeMirrorReconciliationErrors, then let an elected compactor register sealed mirror objects that have no catalog row. Reconciliation excludes objects under_active/. Treat rows present only in those active snapshots as unrecovered unless you have another copy. - Check that reconciliation rotations resume and that the catalog pending
queue drains. In claim mode,
siglake_compactor_sealed_pendingcounts pending catalog rows. A zero value cannot reveal mirror objects that still lack catalog rows, and it does not prove that recovered rows are queryable. - Run the query from Step 9: verify the recovered rows for every affected table and lost time range. Compare the result and known event identifiers with the values from step 1 above. Use these row checks, not restored files or queue gauges, to decide whether recovery succeeded.
What to do in each failure mode¶
An ingester pod dies¶
Nothing. The ingester force-seals its WAL on SIGTERM. After an ungraceful kill,
the partially written active segment is recovered on restart, counted by
siglake_wal_partials_recovered_total.
An ungraceful kill during an append can cut the segment's last Arrow IPC message short. Recovery drains every complete fsynced batch ahead of that message and discards only the incomplete tail. Rows Siglake acknowledged before the crash still reach Iceberg. The segment file keeps its original bytes: nothing truncates or rewrites it, so you can inspect it afterwards.
Each such recovery increments siglake_wal_partial_tail_dropped_total, and a
warning log line names the segment, the rows recovered, and the bytes dropped.
The chart's SiglakeWalPartialTailDropped rule turns that counter into a
warning alert; see Monitoring.
Match it against the restart that caused it.
Everything else about WAL integrity still fails closed. The tolerance covers one case: an active segment recovered as a partial whose final message ends at end of file, after at least one complete batch. Outside that case, damage is still an error:
- A sealed segment whose body fails its CRC check is refused whole and counted
by
siglake_wal_crc_mismatch_total. - A legacy unframed segment stays all-or-nothing.
- A segment written with a frame version this build does not support is refused.
- A partial with no complete batch, or one whose bytes are garbled before end of file, fails to decode instead of recovering a shorter prefix.
Every read walks the Arrow IPC length prefixes against the file size before it
decodes anything. When that walk refuses a sealed, legacy or recovered partial
segment, Siglake increments siglake_wal_ipc_framing_refused_total and the
chart raises the critical SiglakeWalIpcFramingRefused alert on any increase
over 15 minutes: the segment is unreadable and needs investigation. The
counter counts refused read attempts rather than segments, so a retried read
counts again. A CRC failure counts on siglake_wal_crc_mismatch_total
instead, and a partial whose complete prefix decoded counts on
siglake_wal_partial_tail_dropped_total. What becomes of a refused segment
depends on which caller read it; check the file rather than assume it was
quarantined.
The compactor dies mid-commit¶
No acknowledged row is lost and no half-commit lands. Acknowledgement happens in the ingester against the WAL, and an Iceberg commit either becomes a snapshot or does not, so the in-flight commit is all or nothing. What the restart needs from you is usually nothing: both drain shapes recover the in-flight segment from the table's own record of what it has consumed, and ask for an operator only when that record cannot decide.
On the filesystem drain, the segment sits in processing/. Restart moves every
processing/ segment to orphans/ rather than retry it blindly, counting the
move in siglake_compactor_orphans_quarantined_total. The next drain cycle
then disposes of each orphan against the target table's cumulative
consumed-segment set:
| What the table's record shows | What the compactor does |
|---|---|
| The segment's name is in the consumed set. | Its rows are provably committed, so the file is deleted. Recommitting would duplicate them. |
| The name is absent and retained snapshot history reaches back past the seal, or the table has no snapshots at all. | It was provably never committed, so the file is renamed back into sealed/ and committed on the same cycle. |
| The name is absent and snapshot expiry may have dropped the snapshot that would prove it. | Ambiguous, so the orphan is held for you. |
siglake_compactor_orphans_disposed_total{action} counts the deletes and
requeues; siglake_compactor_orphans_held is the gauge that needs you. The
catalog-claim drain reclaims the same in-flight work differently: see The
compactor dies mid-commit on the catalog-claim
drain.
When the compactor holds an orphan for you¶
A held orphan is the one case that asks for a decision: the table's record cannot
say whether the segment's rows were committed, so the compactor leaves the file
in orphans/. The chart alerts on it: SiglakeCompactorOrphansHeld fires at
critical severity on siglake_compactor_orphans_held > 0 held for 15 minutes,
and carries the tenant, pod and namespace. The gauge is the tenant's total over
every WAL directory the sweep visits, republished every cycle, so it reaches 0
on the cycle after the last orphan is settled.
If you switch to the catalog claim, the gauge and alert stop while the files stay on the PVC. Inventory held orphans before the switch.
Keep an ambiguous segment in orphans/ until available consumption proof or
your comparison against the table establishes its disposition. Delete it if the
comparison proves its rows committed. Move it into sealed/ only if the
comparison proves they did not. Raise
compactor.snapshotExpire.retainLast to preserve
evidence for future incidents. It cannot restore snapshots that have already
expired, so it will not clear a hold you already have.
The compactor dies mid-commit on the catalog-claim drain¶
On the catalog-claim drain there is no processing/ directory: the claim row
stays in processing and another drain reclaims it once it is older than
SIGLAKE_CLAIM_RECLAIM_MAX_AGE_SECS (default 900 s), on the
SIGLAKE_CLAIM_RECLAIM_INTERVAL_SECS sweep (default 60 s). The reclaim reads
the table's record of what it has consumed first. A claim the proof covers as committed is marked committed
and never redriven, counted by
siglake_catalog_claims_reclaimed_already_committed_total. A claim neither the
durable proof nor retained history covers is still requeued, because duplicate
rows can be found afterwards and unwritten rows cannot, and each one increments
siglake_compactor_reclaim_unprovable_total. That counter is what
SiglakeReclaimWithoutProof fires on, and the remedy is more history rather
than a different guess: see Consumed-proof rolling
upgrades.
A claim the drain keeps failing to commit is a different case. Its row is
quarantined after twelve attempts, which raises
siglake_catalog_claim_quarantined and the SiglakeSegmentsQuarantined
warning. Fix the fault the segment QUARANTINED log line names, then requeue
it.
The catalog is lost¶
The warehouse is intact but unaddressable, because Iceberg metadata pointers
live in the catalog. Restore Postgres from backup. If the restore predates
recent commits, those commits' files become orphans, and siglake gc-orphans
identifies them.
Back the catalog up at least as often as your commit cadence matters.
The WAL volume is lost before upload¶
Any acknowledged event that has reached neither an Iceberg commit nor a successful mirror upload is gone. This includes every uncommitted event when you disable mirroring. The recovery procedure restores only the objects already present in the mirror.
An object-storage object is deleted¶
Enable S3 versioning. Siglake never rewrites an object in place, because compaction writes new files and swaps the manifest, so versioning plus a lifecycle policy gives you a real recovery window.
The backup checklist for Siglake, and the failure modes it does not cover¶
This is what to back up, and what backup cannot save you from.
- S3 bucket versioning enabled
- S3 cross-region replication, if your recovery point objective (RPO) requires it
- RDS automated backups with point-in-time recovery, retention matched to your RPO
-
wal.mirror.enabled: trueconfirmed - For filesystem-drain recovery,
wal.mirror.activeIntervalSecsset to your recovery target - For catalog-claim recovery, your RPO accounts for rows in sealed mirror objects only
- An object-store lifecycle expiry on
wal-mirror/for a filesystem drain, and for either drain oncewal.mirror.activeIntervalSecsis above0 - An alert on
siglake_wal_mirror_failures_total - The recovery procedure rehearsed against a non-production cluster
What this does not cover¶
Backup and recovery leave three failure modes to you.
- There is no built-in warehouse backup command. The warehouse is plain Iceberg on object storage, so back it up with your object-storage tooling.
- There is no cross-region failover orchestration. You can replicate the bucket and the catalog, but promoting a standby is a manual procedure you design.
- There is no point-in-time restore of the warehouse beyond retained snapshots.
Snapshot expiry retains the last 100 snapshots by default, and time travel
further back is not available. Widen
snapshotExpire.retainLastif you want more, at the cost of a largermetadata.jsonon every commit.