# Patroni Cluster Restore

> Point-in-time recovery (PITR) for a Patroni HA cluster with pg ha restore: the custom-bootstrap mechanism, the leader-locality pre-check, the end-to-end runbook, and the post-restore re-baseline steps

---

LLMS index: [llms.txt](/llms.txt)

---

This page is the **restore-side companion to [Patroni Cluster Backup](../ha-backup/)**.
Once a cluster has a stanza with WAL archiving (`pg backup setup` + `pg ha snapshot`),
`pg ha restore` brings the whole cluster back to a point in time — the Patroni
counterpart of the single-instance [`pg restore`](../../restore/). It walks the
mechanism (why a cluster PITR is not just `pgbackrest restore`), the safety
pre-check on where the leader lives, an end-to-end runbook, and the re-baseline
steps that must follow. Command/flag details also live in
[Restore → Patroni cluster](../../restore/#patroni-cluster); this page strings them
into a procedure you can run.

## Prerequisite: a stanza with WAL archiving, on this host

PITR can only replay up to what was archived. Before restoring, the cluster must
already have:

- a cluster stanza provisioned against an S3 repo — `pg backup setup --s3-endpoint ...`
  (see [Backup → S3 repository](../../backup/#s3-object-storage-repository) and
  [HA backup](../ha-backup/#prerequisite-pg-backup-setup));
- at least one **full** snapshot, plus the WAL segments covering the target time
  (taken with `pg ha snapshot create <scope> --type full`).

The target time is bounded by that history: earlier than the newest full backup's
stop time, or later than the last archived WAL, and the recovery cannot land there.

**`pg backup setup` must have run on *this* host, and the backup container must be
up.** This is stricter than it sounds for a cross-host cluster:

- The member that carries the bootstrap restores from S3 using the repo config
  mounted into its container — `pgbackrest-archive.conf` at `/etc/pgbackrest.conf`
  plus the S3 CA at `/etc/pgbackrest/ca.crt`. Both are generated by `pg backup
  setup`, and the container only mounts them when the files exist. A host that never
  ran setup has no repo config to mount, so the `pgbackrest restore` bootstrap
  cannot reach S3 and fails to start.
- The leader-locality and stop-time pre-checks run `pgbackrest ... info` **inside
  the backup container**. `pg ha restore` refreshes the shared `pgbackrest.conf`
  best-effort but does **not** create the member archive config and does **not**
  start the backup container — so bring it up first (`pg backup setup`, or
  `pg backup status` to confirm it is `Up`). If the backup container is down, the
  pre-checks are skipped and the real failure surfaces only during the bootstrap;
  the post-restore fresh snapshot needs it running anyway.

In short: run `pg ha restore` on a host where `pg backup setup` has provisioned the
S3 repo and the backup container is `Up`.

## Mechanism: custom-bootstrap PITR, not a bare `pgbackrest restore`

A Patroni cluster cannot be PITR'd by running `pgbackrest restore` under a live
postmaster — Patroni owns the data directory, the timeline, and the DCS identity.
`pg ha restore` instead follows Patroni's **custom-bootstrap** recipe:

1. **Pause the cluster** (`patronictl pause`). This freezes failover *before*
   anything is stopped: stopping the local leader while the cluster is live lets
   Patroni promote a **remote** replica this host cannot stop, and that new remote
   leader then races the local bootstrap to re-claim the DCS on the old timeline —
   defeating the leader-locality pre-check from inside the destructive path. A
   restore re-run against a cluster a previous attempt already tore down (DCS
   cleared, members stopped) has nothing live to pause; that is the already-frozen
   state, so the pause failure on an empty DCS is tolerated and the run continues.
2. **Clear the cluster's DCS identity** (`patronictl remove`). Patroni's
   `bootstrap.dcs` block only runs when the DCS has no config key, so removing it is
   what re-arms the bootstrap path.
3. **Pick one LOCAL member** and rewrite its `patroni.yml` so the default initdb
   bootstrap is replaced by a `pgbackrest restore` method targeting the point in
   time:

   ```yaml
   bootstrap:
     method: pgbackrest
     pgbackrest:
       command: 'pgbackrest --stanza=pgcli_<scope> --type=time --target="<time>" --target-action=promote --delta restore'
       no_params: true
       keep_existing_recovery_conf: true
   ```

   `no_params: true` stops Patroni appending `--scope`/`--datadir` (which
   `pgbackrest` rejects with `[031] invalid option`); `keep_existing_recovery_conf:
   true` preserves the `recovery.signal` + `restore_command` + `recovery_target_*`
   that `pgbackrest restore` writes itself. The `method`/`pgbackrest` keys are
   **siblings** of `bootstrap.dcs`, never nested inside it — recovery GUCs leaked
   into the DCS config would apply as live settings on the promoted leader.
4. **Wipe that member's PGDATA** and recreate its container. Patroni sees an empty
   data dir with no DCS config, runs the method, recovers to the target on a **new
   timeline**, and — via `--target-action=promote` — promotes itself to the writable
   leader.
5. **The remaining members rejoin** the new leader: this host's other local members
   are wiped and restarted as standard replicas (they `pg_basebackup` the new
   leader); cross-host members rejoin through the DCS (or are reinitialized).

## The leader-locality pre-check

**`pg ha restore` refuses to run unless the current leader is a local member of this
host.** Before touching anything it reads the leader from the DCS and checks it is
one of the members this host can control.

**Why.** Step 1 clears the DCS and step 3 re-bootstraps a local member on a new
timeline. A leader that lives on another host keeps its Patroni daemon running — this
host cannot stop it — so after the DCS key is removed that remote leader **races to
re-claim leadership on the old timeline**, colliding with the local bootstrap. Requiring
leadership to be local (so it can be stopped) removes the race.

**What you see.**

```bash
# dry-run: a loud warning, no refusal — you can still inspect the plan
pg ha restore app --time "2026-08-26 15:30:00+00" --dry-run
#   [!!] current leader "node3" is NOT a local member — the real run REFUSES until
#        leadership is on this host...

# real run on a host whose leader is remote: hard error, nothing is touched
pg ha restore app --time "2026-08-26 15:30:00+00"
#   current leader "node3" is not a local member of scope "app" ...
#   Move leadership here first (pg ha switchover app ...) or run on the leader's host
```

**Fix.** Move leadership onto a member of the host you are standing on, then retry:

```bash
pg ha switchover app    # or: pg ha failover app ...
pg ha status app        # confirm the leader is now a local member
```

Or run `pg ha restore` on the leader's own host instead. An **empty / unreadable**
leader (e.g. the very first bootstrap, before any member has published itself) does
**not** block the run.

## Differences from single-instance `pg restore`

- **It always promotes.** There is no read-only "hold, inspect, retry another time"
  two-step — the bootstrap recovery ends on a writable leader on a new timeline.
  Confirm the target with `--dry-run` first. (The `patronictl pause` in step 1 of
  the mechanism is unrelated: it freezes failover during the restore, it is not a
  read-only inspection window and the cluster is resumed implicitly when the new
  bootstrap republishes the DCS config.)
- **There is no `--promote` flag** — promotion is automatic (it is baked into the
  bootstrap command).
- **It must run on a host that owns a member**, and the leader must be local (see the
  pre-check above).

## Runbook: restoring a cluster to a point in time

The full flag set is `--time` (required, all the [formats in Restore](../../restore/#point-in-time-recovery-pitr)),
`--member`, `--dry-run`, `--tail-logs`, `--force`.

**1. Confirm the target is safe — dry-run.** Nothing is touched; you see the stanza,
the chosen bootstrap member, the exact `pgbackrest` command, and (if applicable) the
leader-locality warning.

```bash
pg ha restore app --time "2026-08-26 15:30:00+00" --dry-run
```

**2. Make sure the leader is local** (only if the dry-run warned). Switchover, then
re-check status.

```bash
pg ha switchover app
pg ha status app
```

**3. Restore.** The default prompts for confirmation, listing scope, stanza, target,
bootstrap member, and the "PERMANENTLY LOST" warning. Stream the recovery logs with
`--tail-logs`; skip the prompt with `--force` for automation.

```bash
pg ha restore app --time "2026-08-26 15:30:00+00" --tail-logs
# or, picking the bootstrap member explicitly:
pg ha restore app --time "2026-08-26 15:30:00+00" --member node1 --force
```

**4. Verify the new cluster comes up.** `waitForLeader` polls the DCS until a leader
appears (15-minute ceiling), then rejoin members start as replicas.

```bash
pg ha status app                     # new leader on a switched timeline, replicas streaming
pg exec --dsn "postgres://<user>@<leader_host>:<port>/postgres" \
  "SELECT ... FROM your_table"       # data stops at the target time; later commits gone
```

## After the restore: re-baseline the new timeline

The restore is not the last step. The data directory was rebuilt, so two follow-ups
are required before the cluster is backup-ready again:

1. **Take a fresh full snapshot** so future PITR has a base on the new timeline:

   ```bash
   pg ha snapshot create app --type full
   ```

   This usually succeeds directly: a pgBackRest restore of the same stanza
   **preserves the PostgreSQL system-id across the timeline switch** (verified),
   so the post-restore snapshot does not hit `[051] system-id ... do not match
   stanza`. You do **not** need `pg backup stanza-upgrade` after a cluster
   restore. (A real system-id change — a fresh initdb into the same stanza, e.g.
   a brand-new cluster reusing an old stanza name — is what `[051]` and
   stanza-upgrade are for.)

2. **Bring back any member that did not auto-rejoin** — typically a cross-host
   member that lost the old leader. Reinitialize it from the new leader (destructive
   to that replica's data dir only):

   ```bash
   pg ha ctl app -- reinit app-<ns> <member> --force
   ```

   (Full details in [HA backup → stuck replica reinit](../ha-backup/#pitfall-a-stuck-replica-lsn--patronictl-reinit).)

## Troubleshooting

- **"no local member on this host"** — you ran it on a host that owns none of the
  scope's members. `pg ha restore` can only rebuild a local data dir + container. Run
  it on a host that owns a member, or on the leader's host.
- **"target time is before the latest backup stop time"** — the earliest usable point
  is the newest full backup's stop time; the error prints it and a suggested
  `--time`. Take a fresher full snapshot if you need a later base.
- **Recovery exceeds the last archived WAL** — you cannot restore past what was
  archived. Reduce the target time, or ensure `archive_command` is flowing
  (`pg backup status`) before retrying.
- **First post-restore snapshot fails `[051]`** — not expected: a same-stanza
  restore keeps the system-id (see the re-baseline section). If you do hit it, a
  `pg backup stanza-upgrade` fixes it non-destructively.
- **`FATAL: recovery ended before configured recovery target was reached`** on a
  *repeat* PITR into the same repo — the stanza already hosts a previously
  promoted timeline, and the default `recovery_target_timeline=latest` jumps onto
  that branch, which has no commits before your target. Restore again with an
  explicit older timeline pinned on the command, e.g.
  `--recovery-option=recovery_target_timeline=<old-tl>` (added to the bootstrap
  `pgbackrest ... restore` command in the member's `patroni.yml`). A first PITR
  after the target — no branch crossing — never hits this.

---

Backlinks:

- [HA Cluster](/docs/ha-cluster/)
