Patroni Cluster Restore
This page is the restore-side companion to Patroni Cluster Backup.
Once a cluster has a stanza with WAL archiving (pg backup setup + pg ha snapshot),
pg ha restore brings the whole cluster back to a point in time — the Patroni
counterpart of the single-instance pg restore. It walks the
mechanism (why a cluster PITR is not just pgbackrest restore), the safety
pre-check on where the leader lives, an end-to-end runbook, and the re-baseline
steps that must follow. Command/flag details also live in
Restore → Patroni cluster; this page strings them
into a procedure you can run.
Prerequisite: a stanza with WAL archiving, on this host
PITR can only replay up to what was archived. Before restoring, the cluster must already have:
- a cluster stanza provisioned against an S3 repo —
pg backup setup --s3-endpoint ...(see Backup → S3 repository and HA backup); - at least one full snapshot, plus the WAL segments covering the target time
(taken with
pg ha snapshot create <scope> --type full).
The target time is bounded by that history: earlier than the newest full backup’s stop time, or later than the last archived WAL, and the recovery cannot land there.
pg backup setup must have run on this host, and the backup container must be
up. This is stricter than it sounds for a cross-host cluster:
- The member that carries the bootstrap restores from S3 using the repo config
mounted into its container —
pgbackrest-archive.confat/etc/pgbackrest.confplus the S3 CA at/etc/pgbackrest/ca.crt. Both are generated bypg backup setup, and the container only mounts them when the files exist. A host that never ran setup has no repo config to mount, so thepgbackrest restorebootstrap cannot reach S3 and fails to start. - The leader-locality and stop-time pre-checks run
pgbackrest ... infoinside the backup container.pg ha restorerefreshes the sharedpgbackrest.confbest-effort but does not create the member archive config and does not start the backup container — so bring it up first (pg backup setup, orpg backup statusto confirm it isUp). If the backup container is down, the pre-checks are skipped and the real failure surfaces only during the bootstrap; the post-restore fresh snapshot needs it running anyway.
In short: run pg ha restore on a host where pg backup setup has provisioned the
S3 repo and the backup container is Up.
Mechanism: custom-bootstrap PITR, not a bare pgbackrest restore
A Patroni cluster cannot be PITR’d by running pgbackrest restore under a live
postmaster — Patroni owns the data directory, the timeline, and the DCS identity.
pg ha restore instead follows Patroni’s custom-bootstrap recipe:
-
Pause the cluster (
patronictl pause). This freezes failover before anything is stopped: stopping the local leader while the cluster is live lets Patroni promote a remote replica this host cannot stop, and that new remote leader then races the local bootstrap to re-claim the DCS on the old timeline — defeating the leader-locality pre-check from inside the destructive path. A restore re-run against a cluster a previous attempt already tore down (DCS cleared, members stopped) has nothing live to pause; that is the already-frozen state, so the pause failure on an empty DCS is tolerated and the run continues. -
Clear the cluster’s DCS identity (
patronictl remove). Patroni’sbootstrap.dcsblock only runs when the DCS has no config key, so removing it is what re-arms the bootstrap path. -
Pick one LOCAL member and rewrite its
patroni.ymlso the default initdb bootstrap is replaced by apgbackrest restoremethod targeting the point in time:no_params: truestops Patroni appending--scope/--datadir(whichpgbackrestrejects with[031] invalid option);keep_existing_recovery_conf: truepreserves therecovery.signal+restore_command+recovery_target_*thatpgbackrest restorewrites itself. Themethod/pgbackrestkeys are siblings ofbootstrap.dcs, never nested inside it — recovery GUCs leaked into the DCS config would apply as live settings on the promoted leader. -
Wipe that member’s PGDATA and recreate its container. Patroni sees an empty data dir with no DCS config, runs the method, recovers to the target on a new timeline, and — via
--target-action=promote— promotes itself to the writable leader. -
The remaining members rejoin the new leader: this host’s other local members are wiped and restarted as standard replicas (they
pg_basebackupthe new leader); cross-host members rejoin through the DCS (or are reinitialized).
The leader-locality pre-check
pg ha restore refuses to run unless the current leader is a local member of this
host. Before touching anything it reads the leader from the DCS and checks it is
one of the members this host can control.
Why. Step 1 clears the DCS and step 3 re-bootstraps a local member on a new timeline. A leader that lives on another host keeps its Patroni daemon running — this host cannot stop it — so after the DCS key is removed that remote leader races to re-claim leadership on the old timeline, colliding with the local bootstrap. Requiring leadership to be local (so it can be stopped) removes the race.
What you see.
Fix. Move leadership onto a member of the host you are standing on, then retry:
Or run pg ha restore on the leader’s own host instead. An empty / unreadable
leader (e.g. the very first bootstrap, before any member has published itself) does
not block the run.
Differences from single-instance pg restore
- It always promotes. There is no read-only “hold, inspect, retry another time”
two-step — the bootstrap recovery ends on a writable leader on a new timeline.
Confirm the target with
--dry-runfirst. (Thepatronictl pausein step 1 of the mechanism is unrelated: it freezes failover during the restore, it is not a read-only inspection window and the cluster is resumed implicitly when the new bootstrap republishes the DCS config.) - There is no
--promoteflag — promotion is automatic (it is baked into the bootstrap command). - It must run on a host that owns a member, and the leader must be local (see the pre-check above).
Runbook: restoring a cluster to a point in time
The full flag set is --time (required, all the formats in Restore),
--member, --dry-run, --tail-logs, --force.
1. Confirm the target is safe — dry-run. Nothing is touched; you see the stanza,
the chosen bootstrap member, the exact pgbackrest command, and (if applicable) the
leader-locality warning.
2. Make sure the leader is local (only if the dry-run warned). Switchover, then re-check status.
3. Restore. The default prompts for confirmation, listing scope, stanza, target,
bootstrap member, and the “PERMANENTLY LOST” warning. Stream the recovery logs with
--tail-logs; skip the prompt with --force for automation.
4. Verify the new cluster comes up. waitForLeader polls the DCS until a leader
appears (15-minute ceiling), then rejoin members start as replicas.
After the restore: re-baseline the new timeline
The restore is not the last step. The data directory was rebuilt, so two follow-ups are required before the cluster is backup-ready again:
-
Take a fresh full snapshot so future PITR has a base on the new timeline:
This usually succeeds directly: a pgBackRest restore of the same stanza preserves the PostgreSQL system-id across the timeline switch (verified), so the post-restore snapshot does not hit
[051] system-id ... do not match stanza. You do not needpg backup stanza-upgradeafter a cluster restore. (A real system-id change — a fresh initdb into the same stanza, e.g. a brand-new cluster reusing an old stanza name — is what[051]and stanza-upgrade are for.) -
Bring back any member that did not auto-rejoin — typically a cross-host member that lost the old leader. Reinitialize it from the new leader (destructive to that replica’s data dir only):
(Full details in HA backup → stuck replica reinit.)
Troubleshooting
- “no local member on this host” — you ran it on a host that owns none of the
scope’s members.
pg ha restorecan only rebuild a local data dir + container. Run it on a host that owns a member, or on the leader’s host. - “target time is before the latest backup stop time” — the earliest usable point
is the newest full backup’s stop time; the error prints it and a suggested
--time. Take a fresher full snapshot if you need a later base. - Recovery exceeds the last archived WAL — you cannot restore past what was
archived. Reduce the target time, or ensure
archive_commandis flowing (pg backup status) before retrying. - First post-restore snapshot fails
[051]— not expected: a same-stanza restore keeps the system-id (see the re-baseline section). If you do hit it, apg backup stanza-upgradefixes it non-destructively. FATAL: recovery ended before configured recovery target was reachedon a repeat PITR into the same repo — the stanza already hosts a previously promoted timeline, and the defaultrecovery_target_timeline=latestjumps onto that branch, which has no commits before your target. Restore again with an explicit older timeline pinned on the command, e.g.--recovery-option=recovery_target_timeline=<old-tl>(added to the bootstrappgbackrest ... restorecommand in the member’spatroni.yml). A first PITR after the target — no branch crossing — never hits this.