Skip to content

Patroni Cluster Restore

Point-in-time recovery (PITR) for a Patroni HA cluster with pg ha restore: the custom-bootstrap mechanism, the leader-locality pre-check, the end-to-end runbook, and the post-restore re-baseline steps

This page is the restore-side companion to Patroni Cluster Backup. Once a cluster has a stanza with WAL archiving (pg backup setup + pg ha snapshot), pg ha restore brings the whole cluster back to a point in time — the Patroni counterpart of the single-instance pg restore. It walks the mechanism (why a cluster PITR is not just pgbackrest restore), the safety pre-check on where the leader lives, an end-to-end runbook, and the re-baseline steps that must follow. Command/flag details also live in Restore → Patroni cluster; this page strings them into a procedure you can run.

Prerequisite: a stanza with WAL archiving, on this host

PITR can only replay up to what was archived. Before restoring, the cluster must already have:

  • a cluster stanza provisioned against an S3 repo — pg backup setup --s3-endpoint ... (see Backup → S3 repository and HA backup);
  • at least one full snapshot, plus the WAL segments covering the target time (taken with pg ha snapshot create <scope> --type full).

The target time is bounded by that history: earlier than the newest full backup’s stop time, or later than the last archived WAL, and the recovery cannot land there.

pg backup setup must have run on this host, and the backup container must be up. This is stricter than it sounds for a cross-host cluster:

  • The member that carries the bootstrap restores from S3 using the repo config mounted into its container — pgbackrest-archive.conf at /etc/pgbackrest.conf plus the S3 CA at /etc/pgbackrest/ca.crt. Both are generated by pg backup setup, and the container only mounts them when the files exist. A host that never ran setup has no repo config to mount, so the pgbackrest restore bootstrap cannot reach S3 and fails to start.
  • The leader-locality and stop-time pre-checks run pgbackrest ... info inside the backup container. pg ha restore refreshes the shared pgbackrest.conf best-effort but does not create the member archive config and does not start the backup container — so bring it up first (pg backup setup, or pg backup status to confirm it is Up). If the backup container is down, the pre-checks are skipped and the real failure surfaces only during the bootstrap; the post-restore fresh snapshot needs it running anyway.

In short: run pg ha restore on a host where pg backup setup has provisioned the S3 repo and the backup container is Up.

Mechanism: custom-bootstrap PITR, not a bare pgbackrest restore

A Patroni cluster cannot be PITR’d by running pgbackrest restore under a live postmaster — Patroni owns the data directory, the timeline, and the DCS identity. pg ha restore instead follows Patroni’s custom-bootstrap recipe:

  1. Pause the cluster (patronictl pause). This freezes failover before anything is stopped: stopping the local leader while the cluster is live lets Patroni promote a remote replica this host cannot stop, and that new remote leader then races the local bootstrap to re-claim the DCS on the old timeline — defeating the leader-locality pre-check from inside the destructive path. A restore re-run against a cluster a previous attempt already tore down (DCS cleared, members stopped) has nothing live to pause; that is the already-frozen state, so the pause failure on an empty DCS is tolerated and the run continues.

  2. Clear the cluster’s DCS identity (patronictl remove). Patroni’s bootstrap.dcs block only runs when the DCS has no config key, so removing it is what re-arms the bootstrap path.

  3. Pick one LOCAL member and rewrite its patroni.yml so the default initdb bootstrap is replaced by a pgbackrest restore method targeting the point in time:

    bootstrap:
      method: pgbackrest
      pgbackrest:
        command: 'pgbackrest --stanza=pgcli_<scope> --type=time --target="<time>" --target-action=promote --delta restore'
        no_params: true
        keep_existing_recovery_conf: true

    no_params: true stops Patroni appending --scope/--datadir (which pgbackrest rejects with [031] invalid option); keep_existing_recovery_conf: true preserves the recovery.signal + restore_command + recovery_target_* that pgbackrest restore writes itself. The method/pgbackrest keys are siblings of bootstrap.dcs, never nested inside it — recovery GUCs leaked into the DCS config would apply as live settings on the promoted leader.

  4. Wipe that member’s PGDATA and recreate its container. Patroni sees an empty data dir with no DCS config, runs the method, recovers to the target on a new timeline, and — via --target-action=promote — promotes itself to the writable leader.

  5. The remaining members rejoin the new leader: this host’s other local members are wiped and restarted as standard replicas (they pg_basebackup the new leader); cross-host members rejoin through the DCS (or are reinitialized).

The leader-locality pre-check

pg ha restore refuses to run unless the current leader is a local member of this host. Before touching anything it reads the leader from the DCS and checks it is one of the members this host can control.

Why. Step 1 clears the DCS and step 3 re-bootstraps a local member on a new timeline. A leader that lives on another host keeps its Patroni daemon running — this host cannot stop it — so after the DCS key is removed that remote leader races to re-claim leadership on the old timeline, colliding with the local bootstrap. Requiring leadership to be local (so it can be stopped) removes the race.

What you see.

# dry-run: a loud warning, no refusal — you can still inspect the plan
pg ha restore app --time "2026-08-26 15:30:00+00" --dry-run
#   [!!] current leader "node3" is NOT a local member — the real run REFUSES until
#        leadership is on this host...

# real run on a host whose leader is remote: hard error, nothing is touched
pg ha restore app --time "2026-08-26 15:30:00+00"
#   current leader "node3" is not a local member of scope "app" ...
#   Move leadership here first (pg ha switchover app ...) or run on the leader's host

Fix. Move leadership onto a member of the host you are standing on, then retry:

pg ha switchover app    # or: pg ha failover app ...
pg ha status app        # confirm the leader is now a local member

Or run pg ha restore on the leader’s own host instead. An empty / unreadable leader (e.g. the very first bootstrap, before any member has published itself) does not block the run.

Differences from single-instance pg restore

  • It always promotes. There is no read-only “hold, inspect, retry another time” two-step — the bootstrap recovery ends on a writable leader on a new timeline. Confirm the target with --dry-run first. (The patronictl pause in step 1 of the mechanism is unrelated: it freezes failover during the restore, it is not a read-only inspection window and the cluster is resumed implicitly when the new bootstrap republishes the DCS config.)
  • There is no --promote flag — promotion is automatic (it is baked into the bootstrap command).
  • It must run on a host that owns a member, and the leader must be local (see the pre-check above).

Runbook: restoring a cluster to a point in time

The full flag set is --time (required, all the formats in Restore), --member, --dry-run, --tail-logs, --force.

1. Confirm the target is safe — dry-run. Nothing is touched; you see the stanza, the chosen bootstrap member, the exact pgbackrest command, and (if applicable) the leader-locality warning.

pg ha restore app --time "2026-08-26 15:30:00+00" --dry-run

2. Make sure the leader is local (only if the dry-run warned). Switchover, then re-check status.

pg ha switchover app
pg ha status app

3. Restore. The default prompts for confirmation, listing scope, stanza, target, bootstrap member, and the “PERMANENTLY LOST” warning. Stream the recovery logs with --tail-logs; skip the prompt with --force for automation.

pg ha restore app --time "2026-08-26 15:30:00+00" --tail-logs
# or, picking the bootstrap member explicitly:
pg ha restore app --time "2026-08-26 15:30:00+00" --member node1 --force

4. Verify the new cluster comes up. waitForLeader polls the DCS until a leader appears (15-minute ceiling), then rejoin members start as replicas.

pg ha status app                     # new leader on a switched timeline, replicas streaming
pg exec --dsn "postgres://<user>@<leader_host>:<port>/postgres" \
  "SELECT ... FROM your_table"       # data stops at the target time; later commits gone

After the restore: re-baseline the new timeline

The restore is not the last step. The data directory was rebuilt, so two follow-ups are required before the cluster is backup-ready again:

  1. Take a fresh full snapshot so future PITR has a base on the new timeline:

    pg ha snapshot create app --type full

    This usually succeeds directly: a pgBackRest restore of the same stanza preserves the PostgreSQL system-id across the timeline switch (verified), so the post-restore snapshot does not hit [051] system-id ... do not match stanza. You do not need pg backup stanza-upgrade after a cluster restore. (A real system-id change — a fresh initdb into the same stanza, e.g. a brand-new cluster reusing an old stanza name — is what [051] and stanza-upgrade are for.)

  2. Bring back any member that did not auto-rejoin — typically a cross-host member that lost the old leader. Reinitialize it from the new leader (destructive to that replica’s data dir only):

    pg ha ctl app -- reinit app-<ns> <member> --force

    (Full details in HA backup → stuck replica reinit.)

Troubleshooting

  • “no local member on this host” — you ran it on a host that owns none of the scope’s members. pg ha restore can only rebuild a local data dir + container. Run it on a host that owns a member, or on the leader’s host.
  • “target time is before the latest backup stop time” — the earliest usable point is the newest full backup’s stop time; the error prints it and a suggested --time. Take a fresher full snapshot if you need a later base.
  • Recovery exceeds the last archived WAL — you cannot restore past what was archived. Reduce the target time, or ensure archive_command is flowing (pg backup status) before retrying.
  • First post-restore snapshot fails [051] — not expected: a same-stanza restore keeps the system-id (see the re-baseline section). If you do hit it, a pg backup stanza-upgrade fixes it non-destructively.
  • FATAL: recovery ended before configured recovery target was reached on a repeat PITR into the same repo — the stanza already hosts a previously promoted timeline, and the default recovery_target_timeline=latest jumps onto that branch, which has no commits before your target. Restore again with an explicit older timeline pinned on the command, e.g. --recovery-option=recovery_target_timeline=<old-tl> (added to the bootstrap pgbackrest ... restore command in the member’s patroni.yml). A first PITR after the target — no branch crossing — never hits this.