Skip to content

This is the multi-page printable view of this section. .

Return to the regular view of this page.

HA Cluster

Patroni high-availability clusters with pgcli — pg ha commands, dynamic configuration, REST API, extensions, cluster backup and restore, direct SQL

Patroni is the de-facto standard for PostgreSQL high availability: it owns each postmaster’s lifecycle, streams replication between members, and performs automatic failover when the leader is lost. pgcli exposes Patroni as its own top-level command — pg ha — a distinct mode rather than a plain pg instance or an addon installed with pg addon install.

Note on placement. Patroni pages used to live under Addons. They are listed here as a standalone HA Cluster section because pg ha is not installed via the addon system — the addon index only mentions it for discovery. The etcd DCS and the HAProxy load balancer remain addon pages (etcd, HAProxy); this section links to them where relevant.

Pages

Page What it covers
Patroni HA The pg ha command set: create, status, switchover/failover, pause, cross-host members, passwords, namespace & DCS layout
Dynamic Configuration pg ha edit-config / pg ha ctl — the DCS-backed runtime config, and why it never recreates a container
REST API The per-member Patroni REST API: health checks, leader redirects, what HAProxy probes
Extensions in HA Installing/removing PostgreSQL extensions across a cluster (rolling shared_preload_libraries changes)
Cluster Backup pg backup setup + pg ha snapshot — stanza, WAL archiving to S3, cross-host trust
Cluster Restore pg ha restore — cluster PITR via the custom-bootstrap mechanism, the leader-locality pre-check, the post-restore re-baseline
Exec / psql pg ha exec / pg ha psql — run SQL against the leader or any member directly, no dsn, no container exec
Example: HA + self-CA MinIO A verified end-to-end walkthrough: a cluster backed up to a MinIO serving a bring-your-own cert, in a fully isolated second environment
Generating a Certificate pg cert — self-signed cert with DNS and IP SANs for any TLS you need without a CA (dev/test servers, an internal endpoint, or a MinIO serving a bring-your-own cert): its flags, and how the single self-signed leaf acts as its own trust anchor
S3 Storage High Availability Making the MinIO/silo repo itself fault-tolerant: the SNSD and MNSD modes pgcli exposes and why only those two, plus ZFS under the data directory — single-host raidz, per-node pools in a distributed cluster, heterogeneous nodes

The typical path

# one host: bootstrap a cluster (on an installed etcd addon member)
pg ha create app --member node1 --etcd m1
pg ha create app --member node2 --etcd m1

# other hosts register their members with the same password set
pg ha passwords app --file app-passwd.yml
ssh other-host
pg ha create app --member node3 --advertise-host 10.0.0.12 \
    --etcd-endpoints 10.0.0.9:2379 --passwords-file app-passwd.yml

# run SQL directly — the leader is resolved from the DCS
pg ha exec app "SELECT version()"
pg ha psql app

# back it up and be able to go back in time
pg backup setup --s3-endpoint ...
pg ha snapshot create app --type full
pg ha restore app --time "2026-08-26 15:30:00+00"

1 - Patroni HA

Run PostgreSQL high availability with Patroni as a pgcli HA mode — automatic failover, switchover, and a DCS-backed cluster

Patroni is the de-facto standard for PostgreSQL high availability: it manages each postmaster’s lifecycle, streams replication between members, and performs automatic failover when the leader is lost. pgcli exposes Patroni as a distinct mode — pg ha — rather than folding it into the plain pg instance path.

This is a separate mode, not an addon subcommand. Patroni (not pgcli) owns the PostgreSQL process. pg ha members do not appear in cfg.Instances, share no instance lifecycle code, and are not created by pg create. The one thing it borrows from the addon system is the etcd DCS.

Linux only, like the etcd addon: Patroni members rely on podman host networking. Both root and rootless podman are supported. On macOS the commands fail fast with a clear message.

Ownership boundary

The single most important thing to internalize is who owns what:

Responsibility Owner
Patroni container (run/start/stop/rm), image, patroni.yml, port allocation, passwords, DCS wiring pgcli
postmaster lifecycle, initdb, PostgreSQL config rendering, replication slots, failover Patroni
switchover / failover / pause / edit-config pgcli wrapping patronictl in short-lived containers

pgcli never edits a running PostgreSQL’s config directly, and never runs initdb — Patroni does both. This is also why the plain PG image’s docker-entrypoint-initdb.d convention (the admin role / default database) does not apply here: Patroni bootstraps its own cluster, so the role system is postgres (superuser), replicator, and rewind_user instead.

Container lifecycle ≠ safe PG restart

Because Patroni is PID 1 inside its container:

  • pg ha create is re-install semantics (stop + recreate the container). Recreating the leader takes that node fully offline and triggers a failover. Changing a member’s config by re-running create is fine; for dynamic settings prefer pg ha edit-config, which never touches the container.
  • pg ha start / pg ha stop are raw container start/stop. Stopping the leader’s container is exactly as disruptive as the node crashing — Patroni will fail over to a replica. For planned maintenance, pg ha pause first (it disables auto-failover), then stop the container.
  • A paused cluster has no automatic failover until resumed.

How It Works

pg ha runs one Patroni container per member. All members of a scope point at the same DCS (a etcd cluster) which holds the cluster’s dynamic configuration and leader lock:

  1. The first pg ha create for a scope bootstraps: Patroni runs initdb, wins the leader race, and becomes the leader.
  2. Every later member automatically pg_basebackups from the current leader and begins streaming — no flag distinguishes “add” from “join”.
  3. pg ha control commands (switchover, pause, …) run patronictl in an ephemeral container against the DCS, so they work even when no member container is running locally.

The DCS is the source of truth: Patroni re-renders each member’s postgresql.conf/pg_hba.conf from it every loop, so hand-edits on disk are lost — use pg ha edit-config instead.

Scope and the DCS layout

The <scope> in pg ha create <scope> is the Patroni cluster name — the identity of one HA cluster. Every command that takes a scope (status, switchover, failover, pause, ctl, …) names the same cluster; the first create for a scope bootstraps it, later creates with the same scope add members. (--member is the per-host node name inside the cluster — one member per machine.)

In pg.yaml the cluster is keyed by that scope under addons.patroni.<scope>.

What actually lands in etcd

Patroni stores everything under a fixed etcd namespace plus the scope:

  • namespace: /service/ — Patroni’s top-level key, not changed by pgcli.
  • scope: pgcli writes PatroniScope(scope) into patroni.yml, which is your scope plus the pgcli namespace suffix. Patroni has no namespace concept of its own (its etcd prefix is the raw scope), so pgcli bakes the suffix into the scope to keep two pgcli namespaces sharing one etcd from cross-talking.

So the full etcd prefix for a cluster is:

/service/<scope>[-<namespace>]/          # e.g. /service/app/  (no namespace)
                                         #      /service/app-prod/  (namespace: prod)
├── initialize       # bootstrap marker (written once)
├── leader           # current leader; value = member name
├── members/<member> # per-member registration (conn_url, api_url, state)
├── status           # cluster LSN / state
├── config           # dynamic config (pause lives here too)
├── history          # config revision history
├── failover         # manual failover request
└── sync             # synchronous-replication state

The keys below the prefix are Patroni’s own DCS layout. To inspect them, point pg etcdctl at a running etcd member of the same host/config:

ETCDCTL_ENDPOINTS=http://127.0.0.1:2379 pg etcdctl get /service/ -- --prefix --keys-only

Note the scope shown there carries the namespace suffix, even if you configured the cluster with the bare name.

Install

Prerequisite: a DCS. Either reuse local etcd addon members or point at an external etcd.

# 1. a DCS — here, the etcd addon (see the etcd page for multi-node setups)
pg addon install etcd --name m1

# 2. the first member bootstraps the cluster (becomes leader)
pg ha create app --member node1 --etcd m1

# 3. further members join automatically as replicas
pg ha create app --member node2 --etcd m1
pg ha create app --member node3 --etcd m1

# 4. watch it
pg ha status app

pg ha status app renders patronictl list — exactly one Leader row and the rest Replica … streaming:

+ Cluster: app (7683433951661608987) +-----------+----+-------------+-----+------------+-----+
| Member | Host            | Role    | State     | TL | Receive LSN | Lag | Replay LSN | Lag |
+--------+-----------------+---------+-----------+----+-------------+-----+------------+-----+
| node1  | 127.0.0.1:5432  | Leader  | running   |  1 |             |     |            |     |
| node2  | 127.0.0.1:5433  | Replica | streaming |  1 |   0/3000060 |   0 |  0/3000060 |   0 |
| node3  | 127.0.0.1:5434  | Replica | streaming |  1 |   0/3000060 |   0 |  0/3000060 |   0 |
+--------+-----------------+---------+-----------+----+-------------+-----+------------+-----+

With no argument, pg ha status rolls up every cluster: container state, member count, and each member’s pg=/rest= ports.

Choosing a DCS

Patroni only needs a reachable etcd cluster, so there are two layouts:

  • Co-located — run the etcd members as addons on the same hosts as the Patroni members and pass --etcd m1,m2,m3. Simplest: one machine per role, no extra hosts. Good for a starter 3-node HA cluster.

  • Dedicated DCS hosts — run the etcd cluster on its own (virtual) machines and point every Patroni member at it with --etcd-endpoints. Each line below runs on a different machine (pgcli is per-host):

    # host E1 (10.0.0.20) — bootstrap the etcd cluster
    pg addon install etcd --name e1 --cluster prod \
        --advertise-host 10.0.0.20 --client-port 2379 --peer-port 2380
    
    # host E2 (10.0.0.21) — join
    pg addon install etcd --name e2 --cluster prod \
        --advertise-host 10.0.0.21 --client-port 2379 --peer-port 2380 \
        --join http://10.0.0.20:2379
    
    # host E3 (10.0.0.22) — join
    pg addon install etcd --name e3 --cluster prod \
        --advertise-host 10.0.0.22 --client-port 2379 --peer-port 2380 \
        --join http://10.0.0.20:2379
    
    # hosts A / B / C — one Patroni member each; same endpoint list everywhere
    pg ha create app --member node1 --advertise-host 10.0.0.11 \
        --etcd-endpoints 10.0.0.20:2379,10.0.0.21:2379,10.0.0.22:2379

    This is the more available topology: etcd’s quorum survives losing a PG host, and re-imaging a database machine never takes the DCS down with it. The PG hosts only ever talk to the DCS; --etcd-endpoints lists all member client URLs, so Patroni falls through to the next endpoint if one etcd host is down (writes still need the etcd quorum itself — 2 of 3). Keep the endpoint list identical on every PG host.

    A single endpoint (--etcd-endpoints 10.0.0.20:2379) does work — that one etcd member serves the whole cluster, and the other two are hidden behind it. But it re-introduces a single point of failure: if E1 goes down, Patroni cannot reach the DCS even though the etcd quorum is healthy, and a leader that fails to renew its lock demotes itself. List every endpoint you actually have.

A dedicated odd-sized etcd cluster (3 or 5) is the production recommendation; co-locating is fine for dev and small footprints.

Commands

Command What it does
pg ha create <scope> --member <m> … Register + (re)install one member — recreate = node offline
pg ha list Compact table of all HA clusters (scope, members, DCS, status)
pg ha status [scope] All clusters, or one cluster’s patronictl list
pg ha switchover <scope> Planned leader change (patronictl confirms; --yes to script)
pg ha failover <scope> Promote a replica now
pg ha pause / resume <scope> Disable / re-enable automatic failover
pg ha edit-config <scope> -- … View or patch the dynamic config in the DCS (never recreates a container)
pg ha start / stop <scope> --member m | --all Raw container start/stop (see lifecycle caveat)
pg ha remove <scope> --member m | --scope-all [--clean-data] [--force] Remove member(s); --scope-all also clears the DCS
pg ha passwords <scope> [--file F] Export the stored password set (the --passwords-file format, for other hosts)
pg ha remote <scope> [--member m] [--ssh-port P] Register a member living on another host (typically the cross-host leader) for backup SSH
pg ha remote remove <scope> --member m Undo such a registration (the remote container is untouched)
pg ha extension install/remove/list/apply Install, remove, or list PostgreSQL extensions (see Extensions in HA)
pg ha exec <scope> "<sql>" [--member m] [--database db] Run one-shot SQL against the leader (or a named member) — no dsn, no container exec
pg ha psql <scope> [--member m] [--database db] [-- <psql-args>…] Interactive psql against the cluster, same resolved target
pg ha ctl <scope> -- <patronictl args…> Passthrough to any patronictl command

Flags after -- reach patronictl verbatim (cobra strips the --), so pg ha ctl app -- show-config, pg ha ctl app -- topology, and pg ha edit-config app -- -s synchronous_mode=true --force all work. edit-config --show is a convenience alias for ctl … -- show-config.

Cross-host members

Each host runs its own pgcli managing only that host’s members; the cluster reassembles through the shared DCS, so two hosts’ pg.yaml files each hold a partial view of one scope.

# host A (10.0.0.11) — bootstrap (its own etcd, or an external DCS)
pg ha create app --member node1 --advertise-host 10.0.0.11 \
    --etcd-endpoints 10.0.0.9:2379,10.0.0.10:2379

# host A — export the generated password set for other hosts
pg ha passwords app --file app-passwd.yml

# host B (10.0.0.12) — join, sharing the SAME DCS and password set
pg ha create app --member node2 --advertise-host 10.0.0.12 \
    --etcd-endpoints 10.0.0.9:2379,10.0.0.10:2379 \
    --passwords-file app-passwd.yml

Cross-host checklist (all three must line up on every host):

  1. Ports — auto-assignment is per-host, but connect_address is stored in the DCS cluster-wide. Pass --host-port / --restapi-port with the same value on every host, or replicas can’t reach each other.
  2. Passwords — Patroni’s replication / rewind / REST-API auth is cluster-wide. Export the first host’s generated set with pg ha passwords app --file app-passwd.yml and pass --passwords-file app-passwd.yml to every other pg ha create.
  3. --advertise-host — required for cross-host members; it flips the listen address to 0.0.0.0 and puts a reachable IP in connect_address. Left empty, a member is loopback-only. This applies to the first (bootstrap) member too: it becomes the leader, and its loopback connect_address is what every later host would try to pg_basebackup from — cross-host joins fail forever once the cluster was bootstrapped with the default. If you may ever add a member on another host, pass your LAN IPv4 from the very first create.
  4. Firewall — allow the two ports (PG + REST API) pairwise between members.

pgcli does not validate the peers’ config; pg ha status shows each member’s connect_address so you can self-check reachability.

Backing up a cross-host leader

Each host’s pg.yaml records only its own members, so this host cannot see a leader running on another host — yet pgBackRest’s stanza-create and full backups must connect to the leader.

This now happens automatically: after pg ha create installs a member, it publishes that member’s pgcli-private ports (SSH / REST API) into an etcd registry at /pgcli/ha/<scope>/<member> (the topology itself — member name + host:pgport — Patroni already stores in the DCS). The backup config generators then run patronictl list for the cluster-wide topology and look up each remote member’s SSH port in the registry, so pgbackrest.conf / ssh_config include members on every host, with pg1-host and the SSH HostName pointing at the remote IP (local members still go via 127.0.0.1). pg ha start/stop/remove, autostart, and pg ha extension skip members that aren’t local — they belong to their own host’s pgcli.

Just refresh the backup config:

pg backup setup

Cross-host backup trust is now automatic. A cluster-wide stanza lists every member as a pg*-host, so a full backup / check from any host SSH-probes the members on all hosts to find the primary — each of those member sshd processes must therefore accept the initiating host’s backup key, and its --advertise-host already flipped the listener to 0.0.0.0 (see above) so cross-host SSH reaches it. pg backup setup handles this: each host’s create publishes its backup public key into the same /pgcli/ha/<scope>/<member> registry, setup merges every member’s key into a per-cluster authorized_keys file, and each member container bind-mounts that file as an extra AuthorizedKeysFile. Because sshd re-reads it on every login and the merge rewrites the file in place, a host that joins later is trusted by the running members without a restart (the file only needs the one-time recreate to add the mount, which setup does inside the pause window alongside the archive-config recreate). A member removed with pg ha remove drops its registry key, so the next merge stops trusting it.

The S3 repository CA is distributed the same way: the host that configured ca_file publishes the certificate to /pgcli/ha/<scope>/.repo/ca, and a joiner that leaves ca_file empty pulls it and points its own ca_file at the pulled copy. When the store itself lives on yet another machine, the first host gets the CA without any file copy too — pg backup fetch-ca <store>:<port> pulls it out of the endpoint’s TLS chain (see the MinIO doc). Only public material — SSH public keys and the self-signed CA — ever enters the registry; private keys, passwords, and the S3 secret_key never do (the DCS link is unauthenticated).

Fallback: a remote member created on its host before that host upgraded pgcli has no registry entry, so the generators fall back to this scope’s base SSH port for it (every host starts at patroni_ssh_start_port, so the first member is usually right). When it isn’t, register just that member manually to override (creates no container):

pg ha remote app --member node3 --ssh-port 42301   # writes members.<m>.remote_host
pg ha remote remove app --member node3             # undo the registration

S3 backups and WAL archiving

Patroni clusters ship with archiving off — full backups alone can never reach a point-in-time. Enabling the S3 repository (backup docs, English; 中文) flips it on. Each cluster is one stanza — pgcli_<scope>, matching the namespaced DCS scope, so two pgcli namespaces sharing one S3 bucket never collide — whose backup-side config lists every member as a pg*-host, so pgBackRest finds the primary itself and keeps working across failovers. pg backup setup renders the identical archive_command into each member’s local patroni.yml (pgcli never writes archive GUCs into DCS — that is edit-config’s territory, and Patroni applies a local value whenever DCS does not manage the key), recreates stale member containers replicas-first inside a pause window (the recreate is also the postmaster restart archive_mode needs), and runs stanza-create + check for the cluster. From then on WAL streams to S3 continuously, from whichever member holds the leader lock.

Expect a short per-member offline window and a planned leader demotion at the end — the same recreate semantics as pg ha create.

Passwords

The first pg ha create for a scope generates four credentials — superuser, replication, rewind, and the restapi basic-auth pair (restapi_user / restapi_password) — and stores them in pg.yaml under the cluster. You never pass passwords on the command line: the first member generates them, and pg ha passwords exports the stored set:

pg ha passwords app                          # print the stored set (YAML) to stdout
pg ha passwords app --file app-passwd.yml    # write it to a file (mode 0600) instead

(Without --file the YAML goes to stdout — prefer --file so the secrets stay out of shell history and scrollback.)

The four roles and their default usernames (fixed by pgcli — you only ever set the passwords, which default to a random 16-char string per role). Override the length with --password-length on pg ha create / pg create (8–64; it only applies to the generate path, not to a --passwords-file or an already-stored set):

pg.yaml key Role Default username Default password
superuser PostgreSQL superuser postgres random 16-char, generated
replication replication / streaming replicator random 16-char, generated
rewind pg_rewind role rewind_user random 16-char, generated
restapi_user / restapi_password Patroni REST API basic-auth postgres random 16-char, generated

The postgres / replicator / rewind_user usernames are baked into the rendered patroni.yml (postgresql.authentication); only the passwords are the generated / --passwords-file part. restapi_user is the one username you can override, via the passwords file.

For full control (or when you’d rather author the file yourself) use the same format with --passwords-file:

# app-passwd.yml
superuser: <postgres superuser password>
replication: <replicator password>
rewind: <rewind_user password>
restapi_user: postgres          # optional, defaults to postgres
restapi_password: <REST API basic-auth password>
pg ha create app --member node1 --etcd m1 --passwords-file app-passwd.yml

pg ha create prints where each password came from (generated-and-stored vs. a file path). The rendered patroni.yml is written mode 0600 because it embeds all of them.

Auto-start on Boot

Like the other infra addons, a Patroni member’s container can be brought up after a host reboot — but it only starts the existing container, reading the patroni.yml already on disk; it never re-renders config or re-creates data.

pg autostart enable --ha --scope app --name node1

Members are toggled one at a time (--ha --scope <scope> --name <member>). Start order relative to the DCS doesn’t matter: Patroni retries until etcd answers, then re-elects normally. See the autostart page for the boot-service mechanics.

Logs

pg logs addon patroni --scope app --name node1          # last 50 lines
pg logs addon patroni --scope app --name node1 -f       # follow

Container name is pgcli-patroni-<scope>-<member> (namespace-prefixed when a namespace is set).

Connecting

pg ha exec and pg ha psql are the zero-config path: pgcli resolves the leader from the DCS itself, authenticates with the stored superuser password, and runs psql in a throwaway container — no dsn to assemble, no member container to enter, and remote members are reachable (plain TCP + scram, so a replica on another host works too; --member aims at a specific node).

pg ha exec app "SELECT version()"
pg ha exec app --member node2 "SELECT pg_is_in_recovery()"   # read-only, on a replica
pg ha psql app                                               # interactive

For clients outside pgcli, connect to the leader’s PostgreSQL port. Find the leader with pg ha status app (the Leader row’s Host is its connect_address), then point pg psql at it. On a single host with default loopback-only members, that is 127.0.0.1:<host_port>.

pg psql --dsn postgres://postgres@127.0.0.1:<leader_port>/postgres

After a failover the leader changes, so a fixed connection string should be avoided unless fronted by a pooler or the Patroni REST API’s leader redirect.

For a stable endpoint that survives failover — and optional read/write separation — put HAProxy in front of the cluster.

Planned leader changes

Patroni owns failover; pg ha wraps the patronictl verbs that move the leader on purpose. All three take just the scope:

pg ha switchover app                     # patronictl prompts for the candidate
pg ha switchover app --candidate node2   # name it up front
pg ha switchover app --candidate node2 --yes   # skip prompts (scripted)

pg ha failover   app --candidate node2 --yes   # promote now, no handoff
pg ha pause      app                          # stop auto-failover
pg ha resume     app                          # re-enable it
  • switchover is the planned, graceful one: the current leader is demoted first, so nothing is in flight when the candidate is promoted. Both members must be healthy and caught up. The old leader automatically rejoins as a streaming replica a few seconds later (you may briefly see it as stopped while Patroni restarts its postmaster). A new timeline is opened — this is normal, not a split-brain.
  • failover force-promotes a replica immediately without a clean handoff from the leader. Use it only when the leader is already gone or you are deliberately discarding it; on a healthy cluster it just causes an unnecessary blip. Reach for switchover for anything planned.
  • pause turns off automatic failover cluster-wide. Do this before any planned container surgery — pg ha stop --all, a pg ha create recreate, or pg ha extension — so Patroni doesn’t promote a replica out from under you. pg ha resume puts auto-failover back. (A paused cluster has no automatic failover until resumed.)

These are pure DCS operations — they run patronictl in a throwaway container, touch no member’s container or data, and work even when some members live on other hosts. Like every pg ha control command, the scope is resolved to the namespaced DCS key (app → app-default), so a single-host or cross-host leader is switched the same way.

Because clients dial the leader’s port, the connection target moves after a switchover — see Connecting and put HAProxy in front if you need one stable endpoint.

pg ha vs. pg replica — which to pick

pgcli has two ways to get a standby:

pg replica + pg failover pg ha (Patroni)
Model Manual: pg replica builds a standby, pg failover promotes it on demand Automatic: Patroni keeps N members in sync and self-heals
Failover A human runs pg failover; primary stays down until promoted Patroni detects the loss and promotes a replica within ~30s
Ownership pgcli drives postmaster (as with any instance) Patroni drives postmaster; pgcli owns containers only
Best for Simple read-scaling, single planned promotion, staying on pg-instance tooling Zero-RTO availability requirements, unattended failover

If you need a standby you control by hand, use pg replica. If you need the cluster to survive a node crash without a human, use pg ha.

Configuration

The rendered patroni.yml per member is derived entirely from pg.yaml under addons.patroni.<scope>. A typical entry:

addons:
  patroni:
    app:
      name: app
      etcd_members: [m1]                # or etcd_endpoints for an external DCS
      passwords:
        superuser: <superuser-password>
        replication: <replication-password>
        rewind: <rewind-password>
        restapi_user: postgres
        restapi_password: <restapi-password>
      members:
        node1:
          container_name: pgcli-patroni-app-node1
          image_tag: ghcr.io/mars-base/pgcli/pgcli-patroni:18-4.1.5
          host_port: 35590
          restapi_port: 39090
          data_dir: /home/you/.pgcli/addon/patroni/app/node1
          autostart: false

Ports come from two independent pools — patroni_start_port (default 35532) for PostgreSQL and patroni_restapi_start_port (default 39060) for the REST API — so Patroni members never collide with plain-instance or addon ports.

The DCS-scope key is the scope plus the namespace suffix (Patroni has no namespace concept of its own, so the prefix is baked into the scope to keep two pgcli namespaces sharing one etcd from cross-talking).

Default patroni.yml

The file below is what pgcli renders and mounts read-only into each member’s container. Passwords are auto-generated at bootstrap time; --passwords-file lets you supply your own set.

scope: app-default
namespace: /service/
name: node1

etcd3:
    hosts: 10.0.0.11:2379       # from --etcd-endpoints or local etcd member
    protocol: http

restapi:
    listen: 0.0.0.0:8009
    connect_address: 10.0.0.11:8009
    authentication:
        username: postgres
        password: <auto-generated>

bootstrap:
    dcs:
        ttl: 30
        loop_wait: 10
        retry_timeout: 10
        maximum_lag_on_failover: 1048576
        postgresql:
            use_pg_rewind: true
            use_slots: true
            parameters:
                wal_level: replica
                hot_standby: "on"
    initdb:
        - encoding: UTF8
        - data-checksums
    pg_hba:
        - local all all trust
        - local replication all trust
        - host all all all scram-sha-256
        - host replication all all scram-sha-256

postgresql:
    listen: 0.0.0.0:35532
    connect_address: 10.0.0.11:35532
    data_dir: /var/lib/postgresql/data
    bin_dir: /usr/lib/postgresql/18/bin
    pgpass: /patroni/.pgpass
    use_unix_socket: true
    use_unix_socket_repl: true
    authentication:
        superuser:
            username: postgres
            password: <auto-generated>
        replication:
            username: replicator
            password: <auto-generated>
        rewind:
            username: rewind_user
            password: <auto-generated>
    parameters:
        unix_socket_directories: /var/lib/postgresql

Key points:

  • scope = Patroni cluster name. Combined with namespace, it forms the etcd key prefix.
  • etcd3.hosts = host:port only (no http:// scheme). Multiple endpoints are comma-separated.
  • bootstrap.dcs values are defaults only — they take effect on first bootstrap; afterwards use pg ha edit-config.
  • postgresql.listen: 0.0.0.0 is set when --advertise-host is used (cross-host). Without it, listen is 127.0.0.1.
  • Passwords are generated once and stored in pg.yaml; exported via pg ha passwords.

Dynamic Configuration

Patroni’s dynamic configuration lives in the DCS (etcd) under /service/<scope>/config. Every member reads it on each loop (every loop_wait seconds) and applies changes live — so pg ha edit-config is the right way to tune runtime parameters without restarting containers.

bootstrap.dcs is one-time. The bootstrap.dcs block in patroni.yml only takes effect when the first member of a scope runs pg ha create (i.e. the cluster bootstrap). Once Patroni writes the config into the DCS, subsequent changes to bootstrap.dcs in the YAML file are completely ignored — even on re-install (pg ha create). To change dynamic configuration after bootstrap, use pg ha edit-config.

The common pitfall: re-running pg ha create does not re-read bootstrap.dcs from the YAML — Patroni sees the existing config key in the DCS and uses it directly.

Method Description
pg ha edit-config app -- -s key=value Recommended — pgcli’s standard way
pg ha ctl app -- edit-config Passthrough to patronictl, same effect
Patroni REST API (PATCH /config) Requires access to a member’s REST API port

Persistence: changes are written directly to etcd, not to container files. Container restarts, pg ha start/stop, or even pg ha create (re-install) do not lose these settings — new members automatically pick up the latest config from the DCS.

Viewing the current config

pg ha edit-config app --show

This is a convenience alias for pg ha ctl app -- show-config. The output is the full JSON blob stored in /service/<scope>/config.

Modifying parameters

# Set a single parameter
pg ha edit-config app -- -s loop_wait=5

# Set multiple parameters
pg ha edit-config app -- -s loop_wait=5 -s retry_timeout=3

# Apply without confirmation (useful in scripts)
pg ha edit-config app -- -s loop_wait=5 --force

# Interactive edit (opens $EDITOR with the current config)
pg ha edit-config app

After a change, all members apply it on their next loop (within loop_wait seconds). No restart needed.

Common tunable parameters

Parameter Default Description
loop_wait 10 Seconds between leader-loop iterations (lock renewal, DCS updates)
ttl 30 Leader lock TTL. If the leader fails to renew within this window, replicas trigger failover
retry_timeout 10 Timeout for DCS/PostgreSQL operations. Must be < ttl - loop_wait to give the leader at least one retry chance
maximum_lag_on_failover 1048576 Maximum replication lag (bytes) for a replica to be eligible for promotion. Default 1 MB
synchronous_mode false Enable synchronous replication (zero data loss, higher latency)
synchronous_node_count 1 How many synchronous standby nodes (when synchronous_mode=true)
use_pg_rewind true Use pg_rewind to rejoin a failed leader (faster than full pg_basebackup)
use_slots true Use replication slots (prevent WAL loss when a replica disconnects)
failover_timeout 0 How long to wait before failover (0 = immediate when leader is lost)

PostgreSQL runtime parameters can also be set under postgresql.parameters:

pg ha edit-config app -- -s 'postgresql.parameters.max_connections=200'
pg ha edit-config app -- -s 'postgresql.parameters.work_mem=64MB'

These trigger a PostgreSQL reload (or restart, depending on the parameter’s context). Check pg_hba.conf and postgresql.conf parameter documentation for which settings require a restart.

What not to edit

  • Do not edit patroni.yml on disk — it is regenerated from pg.yaml on every pg ha create, and Patroni reads dynamic config from the DCS anyway.
  • Do not edit postgresql.conf directly — Patroni overwrites it each loop from the DCS config.
  • Do not change scope or namespace — these are baked into the DCS key at bootstrap time and cannot be changed without recreating the cluster.

Full parameter reference: Patroni Dynamic Configuration covers every DCS-tunable parameter with defaults, constraints, and examples.

Extensions

pg ha extension install app pg_stat_statements,pg_cron   # install
pg ha extension list app                                  # list
pg ha extension remove app pg_cron                        # remove

Full reference: Extensions in HA Clusters covers the orchestration order, cross-host workflow, shared_preload_libraries ordering, and builtin-only fast path.

Notes

  • Linux only (root or rootless). Rootless members run as the host user via --userns=keep-id; root members chown config and data dirs to postgres (uid 999) so the container can read 0600 config and write its data dir.
  • pg_hba.conf is permissive by design (host all all all scram-sha-256
    • a replication line, plus local trust lines). Rootless podman’s pasta rewrites loopback sources, and Patroni’s own replace_pg_hba step only ever grants its resolved loopback TCP address during custom bootstrap — so postgresql.use_unix_socket/use_unix_socket_repl are set to make Patroni connect to its own instance over the unix socket instead, which is unaffected by the rewrite and keeps bootstrap from deadlocking on its own pg_hba. Tightening the allow-list to a fixed set is a planned future refinement — do not expose these ports to untrusted networks yet.
  • The DCS (etcd) has its own security caveats — see the etcd page: pgcli-managed etcd runs without TLS or auth.
  • No init.sh / docker-entrypoint-initdb.d. Patroni bootstraps the cluster itself, so the admin/default-db convention of plain instances does not exist here; use postgres (superuser) to connect and create roles.
  • use_slots / use_pg_rewind are enabled: Patroni owns replication slots and can rejoin a crashed leader via pg_rewind instead of a full rebase.

2 - Patroni Dynamic Configuration

Reference for all dynamically configurable parameters stored in the DCS

These parameters are stored in the DCS (etcd) under /service/<scope>/config and applied to every member of the cluster. Modify them with pg ha edit-config:

pg ha edit-config app -- -s loop_wait=5
pg ha edit-config app -- -s 'postgresql.parameters.max_connections=200'
pg ha edit-config app --show          # view current config

Reference: Patroni official docs, Pigsty dynamic config.

bootstrap.dcs is one-time. The bootstrap.dcs block in patroni.yml only takes effect when the first member of a scope bootstraps the cluster. Once Patroni writes the config into the DCS, subsequent changes to bootstrap.dcs in the YAML file are completely ignored — even on re-install (pg ha create). Re-running pg ha create does not re-read bootstrap.dcs — Patroni sees the existing config key in the DCS and uses it directly.

Method Description
pg ha edit-config app -- -s key=value Recommended — pgcli’s standard way
pg ha ctl app -- edit-config Passthrough to patronictl, same effect
Patroni REST API (PATCH /config) Requires access to a member’s REST API port

Core timing parameters

Parameter Default Min Description
loop_wait 10 1 Seconds the main loop sleeps between iterations (lock renewal, DCS updates, state refresh)
ttl 30 20 Leader lock TTL in seconds. Effectively the wait time before auto-failover triggers
retry_timeout 10 3 DCS and PostgreSQL operation retry timeout. If the DCS or network is down for less than this, Patroni does not demote the leader

Constraint when changing loop_wait, retry_timeout, or ttl:

loop_wait + 2 * retry_timeout <= ttl

Failover parameters

Parameter Default Description
maximum_lag_on_failover 1048576 Maximum replication lag (bytes) for a replica to be eligible for leader promotion
maximum_lag_on_syncnode -1 Maximum lag (bytes) before a synchronous standby is replaced by a healthy async one. When ≤ 0, Patroni does not replace unhealthy sync standbys. Set high enough to avoid frequent replacement during high-transaction workloads
max_timelines_history 0 Max timeline history entries kept in the DCS. 0 = keep all
primary_start_timeout 300 Seconds the leader has to recover from a failure before failover triggers. 0 = immediate failover on crash detection (may lose transactions with async replication). Max failover time = loop_wait + primary_start_timeout + loop_wait; with 0, just loop_wait
primary_stop_timeout — Seconds to wait when stopping PostgreSQL (only effective with synchronous_mode). If the stop exceeds this timeout, Patroni sends SIGKILL to the postmaster. ≤ 0 or unset = no effect
failover_timeout 0 Seconds to wait before triggering failover after the leader is lost. 0 = immediate

Replication mode

Parameter Default Description
synchronous_mode false Enable synchronous replication (off / on / quorum). The leader manages synchronous_standby_names; only the last-known leader or a sync standby may run for leader. Guarantees zero data loss at the cost of write unavailability when durability cannot be assured
synchronous_mode_strict false When no sync standby is available, refuse to disable sync replication — blocks all client writes to the leader
synchronous_node_count 1 Number of synchronous standby nodes. Dynamically adjusted as members join/leave. Clamped to the number of eligible nodes
failsafe_mode false Enable DCS failsafe mode: when the DCS is unreachable, the leader keeps running instead of demoting itself

PostgreSQL settings

Parameter Default Description
postgresql.use_pg_rewind false Use pg_rewind to rejoin a failed leader (faster than full pg_basebackup). Requires data page checksums (--data-checksums at initdb) or wal_log_hints=on
postgresql.use_slots true Use replication slots (prevents WAL loss when a replica disconnects). Default on PostgreSQL 9.4+
postgresql.pg_hba — Rules for generating pg_hba.conf. Ignored if the PostgreSQL hba_file parameter is set to a non-default value
postgresql.pg_ident — Rules for generating pg_ident.conf. Ignored if ident_file is non-default
postgresql.parameters — PostgreSQL GUCs as key-value pairs, e.g. {max_connections: 100, wal_level: "replica", wal_log_hints: "on"}. Many are required for replication to work
postgresql.recovery_conf — Additional recovery.conf entries for standby configuration (PG12+ handled transparently)

Set PostgreSQL parameters via the postgresql.parameters key:

pg ha edit-config app -- -s 'postgresql.parameters.max_connections=200'
pg ha edit-config app -- -s 'postgresql.parameters.work_mem=64MB'
pg ha edit-config app -- -s 'postgresql.parameters.wal_log_hints=on'

Standby cluster

If defined, the cluster bootstraps as a standby cluster that streams from a remote primary.

Parameter Description
standby_cluster.host Remote primary address
standby_cluster.port Remote primary port
standby_cluster.primary_slot_name Slot name for replication (optional, defaults to member name)
standby_cluster.create_replica_methods Ordered list of methods to bootstrap the standby leader from the remote primary
standby_cluster.restore_command WAL restore command
standby_cluster.archive_cleanup_command Archive cleanup command for the standby leader
standby_cluster.recovery_min_apply_delay Delay before applying WAL records

Replication slots

Parameter Default Description
member_slots_ttl 30min How long a replica’s physical replication slot is retained after it shuts down. 0 = delete immediately when the member key expires from the DCS. Only effective on PostgreSQL 11+
slots — Permanent replication slots (hash map). Preserved across switchover/failover. Physical slots on PG11+ are created on all nodes and advanced every loop_wait seconds. Logical slots are copied from primary to replicas via restart, then advanced every loop_wait seconds. Requires use_slots: true
ignore_slots — Slots managed externally that Patroni should not touch (list of attribute sets). Any subset match causes the slot to be ignored

Permanent slots example

slots:
  permanent_physical_slot:
    type: physical
  permanent_logical_slot:
    type: logical
    database: my_db
    plugin: pgoutput

ignore_slots:
  - name: externally_managed_slot
    type: physical

Node-pinned physical slots

For a fixed cluster topology, define a permanent physical slot per node to prevent slot deletion during temporary outages:

slots:
  node1:
    type: physical
  node2:
    type: physical
  node3:
    type: physical

Warning: Permanent replication slots are synced from the primary/standby_leader to replicas only. Applications should use them on the leader node. Using permanent slots on replicas causes unbounded pg_wal growth across the cluster. Exception: physical slots matching a Patroni member name (created and maintained by Patroni) are synced across all nodes for inter-node replication.

Viewing the current configuration

pg ha edit-config app --show

Or read directly from etcd:

ETCDCTL_ENDPOINTS=http://10.0.0.1:2379 pg etcdctl get /service/app-default/config -- --prefix

3 - Patroni REST API

Patroni REST API endpoints for health checks, monitoring, and cluster management

Patroni exposes an HTTP REST API on each member, providing health check endpoints (for load balancers and Kubernetes probes), monitoring data (including Prometheus metrics), and cluster management operations.

Reference: Patroni official docs, Pigsty REST API.

Finding the REST API

Each member’s REST API port is assigned automatically by pgcli and stored in pg.yaml under addons.patroni.<scope>.members.<name>.restapi_port. The port is also embedded in the rendered patroni.yml as restapi.connect_address.

# Find REST API ports
pg ha status app          # shows pg= and rest= ports for each member

Authentication

pgcli enables basic-auth on the REST API (username postgres, auto-generated password). But authentication is method-based, not blanket:

Method Auth required? Endpoints
GET / HEAD / OPTIONS No All read endpoints — /, /primary, /replica, /health, /cluster, /config, /metrics, /patroni, …
POST / PATCH / PUT / DELETE Yes Write endpoints — /failover, /switchover, /config (modify), /reload, /restart, …

This is by design: read-only health checks stay open so a load balancer or Prometheus can poll them without credentials, while operations that change cluster state (POST /failover, PATCH /config) require authentication. The credentials exist so patronictl and other Patroni members can perform writes.

So a load balancer health check needs no auth:

# No -u needed — returns 200 on leader, 503 on replica
curl -s http://<host>:<port>/primary -w "%{http_code}"

Write operations do require auth. The password is the same restapi_password used in patroni.yml’s restapi.authentication section. Export it with:

# Export passwords, then extract the restapi_password
pg ha passwords app --file app-passwd.yml
grep restapi_password app-passwd.yml
# restapi_password: <password>

# Example: a write operation with auth
curl -u postgres:<restapi-password> -X PATCH http://<host>:<port>/config -d '{"ttl": 60}'

Health Check Endpoints

All health check endpoints respond with GET requests. Patroni returns a JSON document describing the node state, along with an HTTP status code. Use HEAD or OPTIONS instead of GET when only the status code is needed (no response body).

Primary-only endpoints (200 only on leader)

These return HTTP 200 only when the node is the current leader holding the leader lock:

Endpoint Description
GET / Root — primary health check
GET /primary Alias for /
GET /read-write Alias for /
GET /leader Like / but doesn’t distinguish primary vs standby_leader
GET /master Legacy alias for /leader
# Returns 200 on leader, 503 on replica
curl -s -u postgres:<pw> http://<leader-host>:<port>/primary -w "%{http_code}"

These endpoints are useful for load balancer health checks that should route writes only to the current primary.

Replica-only endpoints (200 only on replica)

Endpoint Description
GET /replica Replica health check — 200 when node is running, role is replica, and noloadbalance tag is not set
GET /replica?replication_state=streaming Only 200 when replica is actively streaming (not still catching up via archive recovery)
GET /replica?lag=<max> Only 200 when replication lag is below the threshold (bytes or human-readable: 10MB, 1GB)
# Streaming replica only
curl -s -u postgres:<pw> "http://<replica-host>:<port>/replica?replication_state=streaming"

# Replica with lag < 1 MB
curl -s -u postgres:<pw> "http://<replica-host>:<port>/replica?lag=1048576"

Read-only endpoints (200 on both primary and replica)

Endpoint Description
GET /read-only Any running node (primary or replica)
GET /synchronous / GET /sync Sync standby only
GET /asynchronous / GET /async Async standby only
GET /read-only-sync Primary + sync standby
GET /read-only-quorum Primary + quorum standby
GET /quorum Quorum standby only

PostgreSQL health

Endpoint Description
GET /health 200 when PostgreSQL is running (regardless of role)

Kubernetes Probes

Endpoint Description
GET /liveness 200 if Patroni heartbeat loop is running. 503 if primary’s last heartbeat exceeds ttl seconds, or replica exceeds 2*ttl. Lightweight — no SQL queries. Suitable for livenessProbe.
GET /readiness 200 when node is leader, or when PostgreSQL is running, replicating, and within the allowed lag. Accepts ?lag=<max> (default: maximum_lag_on_failover) and ?mode=apply|write (default: apply). Suitable for readinessProbe.
# Example Kubernetes probes
livenessProbe:
  httpGet:
    scheme: HTTP
    path: /liveness
    port: 8008          # REST API port
  initialDelaySeconds: 3
  periodSeconds: 10
  timeoutSeconds: 5
  failureThreshold: 3

readinessProbe:
  httpGet:
    scheme: HTTP
    path: /readiness
    port: 8008
  initialDelaySeconds: 3
  periodSeconds: 10
  timeoutSeconds: 5
  failureThreshold: 3

Monitoring Endpoints

GET /patroni

Returns detailed node status as JSON. Used internally by Patroni during leader election and also useful for monitoring:

curl -s -u postgres:<pw> http://<host>:<port>/patroni | jq .

Response fields:

Field Description
state Node state: running, stopped, starting, etc.
role primary or replica
server_version PostgreSQL version as integer
xlog.location Current WAL position (primary only)
xlog.received_location WAL received from primary (replica only)
xlog.replayed_location WAL replayed (replica only)
timeline Current timeline number
replication Array of connected replicas (primary only)
cluster_unlocked true if no leader lock is held
pause true if auto-failover is paused
dcs_last_seen Epoch timestamp of last successful DCS contact
patroni.version Patroni version
patroni.scope Cluster scope name
patroni.name Member name

GET /cluster

Returns the full cluster topology — all members with their roles, states, and replication status:

curl -s -u postgres:<pw> http://<host>:<port>/cluster | jq .
{
  "members": [
    {
      "name": "node1",
      "role": "leader",
      "state": "running",
      "api_url": "http://10.0.0.11:8008/patroni",
      "host": "10.0.0.11",
      "port": 35532,
      "timeline": 1
    },
    {
      "name": "node2",
      "role": "replica",
      "state": "streaming",
      "host": "10.0.0.11",
      "port": 35533,
      "timeline": 1,
      "receive_lag": 0,
      "receive_lsn": "0/3000168",
      "replay_lag": 0,
      "replay_lsn": "0/3000168"
    }
  ],
  "scope": "app-default"
}

GET /config

Returns the current dynamic configuration stored in the DCS:

curl -s -u postgres:<pw> http://<host>:<port>/config | jq .

This is equivalent to pg ha edit-config <scope> --show.

GET /history

Returns the timeline history. Empty array [] when no timeline switch has occurred (fresh cluster).

GET /metrics

Returns monitoring data in Prometheus exposition format, suitable for scraping by Prometheus or compatible systems:

curl -s -u postgres:<pw> http://<host>:<port>/metrics

Key metrics:

Metric Description
patroni_primary 1 if this node is the leader
patroni_replica 1 if this node is a replica
patroni_postgres_running 1 if PostgreSQL is running
patroni_postgres_streaming 1 if PostgreSQL is streaming (replica)
patroni_xlog_location Current WAL position (primary only)
patroni_xlog_received_location WAL received (replica only)
patroni_xlog_replayed_location WAL replayed (replica only)
patroni_cluster_unlocked 1 if no leader lock
patroni_is_paused 1 if auto-failover is disabled
patroni_pending_restart 1 if node needs restart
patroni_postgres_timeline Current timeline
patroni_dcs_last_seen Epoch of last DCS contact
patroni_server_version PostgreSQL version

Tag-based Filtering

Health check endpoints accept query parameters to filter by custom tags defined in the member’s patroni.yml tags section:

# Only replicas with tag dc=us-east
curl -u postgres:<pw> "http://<host>:<port>/replica?dc=us-east"

# Only leader with tag region=primary
curl -u postgres:<pw> "http://<host>:<port>/leader?region=primary"

Practical Patterns

Load balancer routing

Because GET health checks need no auth, a load balancer can poll the REST API endpoints directly. The pattern (from Patroni’s example haproxy.cfg) is: the backend TCP port is the PostgreSQL port, but the health check hits the REST API port with GET /, which returns 200 only on the leader.

Single read-write endpoint (all traffic → current leader):

global
    maxconn 100

defaults
    log global
    mode tcp
    retries 2
    timeout client 30m
    timeout connect 4s
    timeout server 30m
    timeout check 5s

listen stats
    mode http
    bind *:7000
    stats enable
    stats uri /

listen app
    bind *:5000
    option httpchk
    http-check expect status 200
    default-server inter 3s fall 3 rise 2 on-marked-down shutdown-sessions
    server node1 <ip>:35532 maxconn 100 check port 8008
    server node2 <ip>:35533 maxconn 100 check port 8009
  • check port 8008 / 8009 — the HTTP health check goes to the REST API port, not the PG port.
  • GET / returns 200 only on the leader → HAProxy marks only the leader as UP.
  • on-marked-down shutdown-sessions — on failover, sessions to the old leader are killed so clients reconnect to the new leader.
  • fall 3 / rise 2 with inter 3s — ~9s to mark down, ~6s to mark up.

Read/write split — add a second listener that uses GET /replica (200 only on replicas) for read traffic:

listen app_rw
    bind *:5000
    mode tcp
    option httpchk GET /
    http-check expect status 200
    default-server inter 3s fall 3 rise 2 on-marked-down shutdown-sessions
    server node1 <ip>:35532 maxconn 100 check port 8008
    server node2 <ip>:35533 maxconn 100 check port 8009

listen app_ro
    bind *:5001
    mode tcp
    option httpchk GET /replica
    http-check expect status 200
    default-server inter 3s fall 3 rise 2
    server node1 <ip>:35532 maxconn 100 check port 8008
    server node2 <ip>:35533 maxconn 100 check port 8009
  • app_rw (:5000) — GET / → only the leader is UP → writes land on the leader.
  • app_ro (:5001) — GET /replica → only replicas are UP → reads spread across replicas.

The same GET /replica?lag=1MB refinement (from the replica endpoints above) keeps a lagging replica out of the read pool.

Monitoring with curl

Quick health check script:

#!/bin/bash
# Check all members — no auth needed for GET
for port in 8008 8009; do
  code=$(curl -s http://10.0.0.11:$port/health -o /dev/null -w "%{http_code}")
  echo "Port $port: HTTP $code"
done

Prometheus scrape config

GET /metrics needs no auth, so the scrape config works without credentials:

scrape_configs:
  - job_name: patroni
    static_configs:
      - targets:
        - '10.0.0.11:8008'   # node1
        - '10.0.0.11:8009'   # node2
    metrics_path: /metrics

4 - Extensions in HA Clusters

Install and manage PostgreSQL extensions in Patroni-managed HA clusters

Installing PostgreSQL extensions in a Patroni-managed HA cluster is fundamentally different from a single-node instance. This page explains why, and walks through the pg ha extension command subtree that handles it.

Why not pg extension install?

The single-node flow (pg extension install <instance> <ext>) works by:

  1. Building a -ext derived image with Pigsty packages
  2. Stopping and recreating the container from the new image
  3. Editing postgresql.conf to set shared_preload_libraries
  4. Running CREATE EXTENSION inside the container

In a Patroni cluster, steps 2-4 all break:

  • Patroni is PID 1. Recreating a container takes that node offline. If it’s the leader, Patroni triggers a failover — the cluster reshuffles while you’re mid-install.
  • Patroni regenerates postgresql.conf every loop from the DCS. Any direct edit to the file is overwritten within seconds. shared_preload_libraries must be set via patronictl edit-config (which writes to the DCS).
  • CREATE EXTENSION must run on the leader, which may be on a remote host — not inside any local container.

pg ha extension orchestrates all of this correctly.

How It Works

The install flow follows a specific order to avoid the pitfalls above:

1. Validate extension names (IsExtensionKnown)
2. Build -ext image with Pigsty packages
3. patronictl pause <scope> --wait            ← disable auto-failover
4. Recreate each member container (one at a time, wait for rejoin)
5. patronictl resume <scope> --wait           ← ⚠ must come BEFORE edit-config
6. patronictl edit-config (set shared_preload_libraries)
   → Patroni triggers a rolling restart (replicas first, leader last)
7. Wait for rolling restart to complete
8. CREATE EXTENSION on leader
9. Save extensions list to pg.yaml

The critical ordering is resume before edit-config: a paused Patroni cluster does not apply edit-config changes. If you edit-config while paused, the shared_preload_libraries update is silently lost.

Builtin-only fast path

If all requested extensions are builtin (contrib extensions like hstore, uuid-ossp that ship with PostgreSQL), steps 2-4 are skipped entirely — no image build, no container recreate, no pause/resume. The flow goes straight to edit-config + CREATE EXTENSION.

Commands

Install

pg ha extension install <scope> <extension>[,<extension>...] [flags]

Builds the -ext image, pauses the cluster, and recreates local member containers.

  • Single-host cluster (all members local): automatically runs apply (resume + edit-config + CREATE EXTENSION)
  • Cross-host cluster: only performs pause + recreate, cluster stays paused. Run install on each host separately, then run pg ha extension apply on any host to complete the installation.

Extensions are passed as a comma-separated list:

# Install a single extension
pg ha extension install app pg_stat_statements

# Install multiple extensions to a specific database
pg ha extension install app pg_cron,pg_stat_statements --database mydb

# Skip the rolling restart confirmation prompt
pg ha extension install app pgvector --auto-restart

Flags:

Flag Default Description
--database postgres Target database for CREATE EXTENSION
--auto-restart false Skip confirmation for the rolling restart triggered by edit-config

⚠ --database affects all extensions. When you specify --database, the apply phase runs CREATE EXTENSION IF NOT EXISTS for all installed extensions (not just the newly added ones) in the target database. For example, if the cluster has [pg_cron, pg_stat_statements, pgvector] installed, running pg ha extension install app hstore --database mydb will create all four extensions in mydb. The extension .so files are already in the image and won’t be reinstalled — only the SQL objects (functions, types, etc.) are registered in the target database.

If you only want to enable an existing extension in a specific database, there’s no need to reinstall. Simply connect to that database with psql and run CREATE EXTENSION IF NOT EXISTS directly.

Remove

pg ha extension remove <scope> <extension>[,<extension>...] [flags]

Runs DROP EXTENSION on the leader, then updates shared_preload_libraries via patronictl edit-config (triggers a rolling restart). Does not rebuild the image or recreate containers — the -ext image only grows; disk reclamation is rare and manual.

Extensions are passed as a comma-separated list:

pg ha extension remove app pg_cron
pg ha extension remove app pg_stat_statements,pg_cron --auto-restart

List

pg ha extension list <scope>

Shows three views of the cluster’s extensions:

  • Config (pg.yaml): the extensions list stored in pg.yaml
  • DCS (preload): the shared_preload_libraries value from patronictl show-config
  • Leader (installed): actual extensions from pg_extension on the leader
$ pg ha extension list app
Cluster "app" extensions:

  Config (pg.yaml):  [pg_stat_statements pg_cron]
  DCS (preload):     shared_preload_libraries: pg_stat_statements,pg_cron
  Leader (installed): [pg_cron pg_stat_statements]

Apply

pg ha extension apply <scope> [flags]

Manually triggers the second half of the install flow: resume the cluster, run patronictl edit-config, and execute CREATE EXTENSION on the leader.

Used in cross-host clusters — after running install on each host (each builds the image and recreates its local members), run apply once on any host to complete the DCS update and extension creation.

pg ha extension apply app
pg ha extension apply app --database mydb --auto-restart

⚠ --database affects all installed extensions. apply runs CREATE EXTENSION IF NOT EXISTS for all extensions in the cluster in the specified database, not just the newly added ones.

Cross-Host Workflow

In a cross-host cluster, each host manages only its own members. The extension installation workflow splits into two phases with global pause/resume — the cluster is paused once and resumed once, not per-host:

Phase 1 — per host (sequentially): Run pg ha extension install on each host. Each host:

  • Builds the -ext image locally (merging DCS’s existing preload list to include all packages)
  • Pauses the cluster (idempotent — second host won’t fail on “already paused”)
  • Recreates its own local members from the new image (replicas first, leader last)
  • Saves config — cluster stays paused, no resume

Recommended order: start from replica hosts. If the host with the leader runs install first, recreating the leader triggers a failover (leader moves to another host). Starting from replica hosts keeps the leader in place until the very end, minimizing failover-related data sync overhead.

# host A (replicas only — run first) — cluster enters paused state
pg ha extension install app pg_stat_statements,pg_cron

# host B (has the leader — run last) — cluster stays paused
pg ha extension install app pg_stat_statements,pg_cron

Phase 2 — once, on any host: Run pg ha extension apply to:

  • Resume the cluster (idempotent — tolerates “not paused”)
  • Update shared_preload_libraries via patronictl edit-config
  • Wait for the rolling restart
  • Run CREATE EXTENSION on the leader
pg ha extension apply app --auto-restart

Note: After install, the cluster is in paused state (auto-failover disabled). If you forget to run apply, run pg ha extension apply or manually patronictl resume <scope> to re-enable failover.

For single-host clusters (all members local), install automatically runs the apply step — no separate command needed.

About leader drift: Leader drift cannot be completely avoided at this time. When the container on the host where the leader resides is recreated, the Patroni process stops. After the leader key in the DCS expires (TTL), even if the cluster is paused, other replicas will still detect the expired key and trigger a new election. The replicas-first recreation order ensures that the host with the leader is the last one to run install, but it cannot prevent the leader drift itself. This is an inherent limitation of Patroni + container recreation.

The extensions Config Field

Extensions are tracked at the cluster level in pg.yaml, under the Patroni cluster config:

addons:
  patroni:
    app:
      name: app
      extensions:
        - pg_stat_statements
        - pg_cron
      passwords:
        superuser: ...
      members:
        node1: { ... }
        node2: { ... }

This list is the source of truth for what pg ha extension list reports and what apply installs. Both install and remove update it automatically.

shared_preload_libraries Ordering

Some extensions must appear at position 0 in shared_preload_libraries (PostgreSQL fatals if they’re not first). pgcli tracks this as a catalog attribute (PreloadFirst) — currently set on:

  • citus — distributed PostgreSQL, must be loaded before anything else
  • timescaledb — time-series engine, same constraint

Additional extensions may require companion DCS parameters:

  • pg_cron needs cron.database_name

pg ha extension handles all of this automatically:

  • Extensions with PreloadFirst are always placed at the start of the CSV, regardless of input order
  • When pg_cron is present, cron.database_name is set to the --database value (default postgres)
  • When pg_cron is removed, cron.database_name is cleared

Examples

Install pg_stat_statements and pg_cron

pg ha extension install app pg_stat_statements,pg_cron --database mydb --auto-restart

This builds an -ext image, recreates all members (with pause/resume), sets shared_preload_libraries=pg_stat_statements,pg_cron and cron.database_name=mydb in the DCS, then creates both extensions in mydb.

Install Citus (must be first in preload)

pg ha extension install app pg_stat_statements,citus

Even though citus is listed second, the preload CSV is generated as citus,pg_stat_statements — extensions with PreloadFirst are always placed at position 0.

Remove an extension

pg ha extension remove app pg_cron

Drops pg_cron from the leader, removes it from shared_preload_libraries, and clears cron.database_name. Triggers a rolling restart.

Check what’s installed

pg ha extension list app

Limitations

  • Image rebuild is additive. Removing an extension does not rebuild the -ext image or shrink it. To reclaim disk, manually prune old images with podman image prune.
  • No per-member extension list. Extensions are cluster-wide — all members share the same -ext image and the same shared_preload_libraries.
  • CREATE EXTENSION targets one database. PostgreSQL extensions are per-database. To install in multiple databases, re-run with --database pointing at each one.
  • Cross-host requires manual coordination. Each host must run install before apply is run once. pgcli does not SSH into remote hosts.

See Also

5 - Patroni Cluster Backup

pgBackRest backups for a Patroni HA cluster: backup setup, S3 repository and WAL archiving, pg ha snapshot operations, and reinit when a replica’s LSN is stuck

This page walks the complete procedure for backing up a Patroni HA cluster: getting the backup infrastructure up with pg backup setup, taking full/incr/diff backups with pg ha snapshot, and one operational pitfall — when a replica’s Replay LSN stalls and pg ha status shows members disagreeing on LSN, how to pull it back with patronictl reinit. Command and mechanism details live in Patroni HA and Backup; this page strings them into a path you can run end to end.

Model: one stanza per cluster, the backup container finds the primary

A Patroni cluster (scope) maps to one pgBackRest stanza: pgcli_<scope> (matching the namespace-qualified DCS scope, so two pgcli namespaces sharing one S3 bucket never collide). The stanza is cluster-wide — the backup-side config lists every member as a pg*-host, and pgBackRest SSH-probes them, locates the current primary itself, and keeps following it across failovers with no config change. Therefore:

  • There is no per-member backup. full / incr / diff all operate on the one cluster stanza; pgBackRest snapshots whatever host is the leader at that moment.
  • Backups run from the shared backup container, reaching the primary over SSH. Cross-host members are reachable too (trust is wired up automatically by setup).
  • pg ha snapshot commands never need you to name the leader — pass the scope; pgBackRest resolves the primary.

Prerequisite: pg backup setup

One idempotent run brings up the whole backup path and gives the cluster archiving:

# simplest (backup data lands under the default base-dir)
pg backup setup

# S3 object repository (the prerequisite for archiving to MinIO / AWS S3).
# When the store is on another host and you have no ca.crt yet, pull it with
# one TLS handshake first — no scp:
pg backup fetch-ca 10.0.0.9:9000          # prints a SHA-256 fingerprint to cross-check
pg backup setup --s3-endpoint 10.0.0.9:9000 --s3-bucket pgbackrest \
    --s3-access-key admin --s3-ca-file ~/.pgcli/backup/repo-ca/ca-10.0.0.9-9000.crt

What setup does (when an S3 repo is configured, it first runs an endpoint preflight — a plain TCP dial that fails the run immediately on a mistyped host or a store that is down, rather than after the image pull / container start buries the real cause):

  1. Build/pull the pgbackrest image, create the network and dirs, generate the backup-container pgbackrest.conf and the member-local archive view pgbackrest-archive.conf.
  2. Start the shared backup container, then verify repository connectivity: a pgbackrest repo-ls proves the whole stack (TLS, credentials, bucket) and maps a failure to an actionable hint — a cert error points at pg backup fetch-ca, access-denied at the credentials, a refused/timeout connection at the endpoint.
  3. Wire up cross-host backup trust: each host publishes its backup public key into the cluster’s etcd registry at pg ha create; setup merges every member’s key into one per-cluster authorized_keys bind-mounted into each member container. sshd re-reads that file on every login and the merge rewrites it in place, so hosts that join later are trusted without a restart. The S3 repo CA follows the same path: a host with ca_file publishes it into the registry, joiners with it empty pull it. (Details: HA → Backing up a cross-host leader.)
  4. Enable WAL archiving: when a member’s archive config is stale, pause the cluster, recreate stale members replicas-first / leader last (a recreate is exactly the postmaster restart archive_mode needs), resume, and render the same archive_command into every member’s patroni.yml.
  5. Run stanza-create + check for every cluster stanza.

HTTPS is mandatory. pgBackRest rejects plaintext S3. A MinIO meant to receive cluster archives must serve TLS — install it with pg addon install minio --tls (pgcli’s self-signed CA; point --s3-ca-file at its ca.crt). When the MinIO is on another machine, pull its CA with one TLS handshake — pg backup fetch-ca <store-host>:<port> — no scp needed (see MinIO → Getting the CA onto a remote host).

Confirm after setup:

pg backup status

Once the backup container is Up and the stanzas are ready, start backing up.

Create members after setup, on every host. A member’s archiving is fixed at creation time: the container only mounts pgbackrest-archive.conf if that file already exists, and the rendered patroni.yml only carries the archive GUCs if an S3 repo is configured. A member pg ha created before pg backup setup on its host therefore starts without WAL archiving — pg ha snapshot/pg ha restore cannot cover it until a later setup flags it stale and recreates it inside a pause window. pg ha create warns about this up front; run pg backup setup first on each host for a backup-ready create.

Snapshot operations: pg ha snapshot

These act on a cluster (scope), running pgBackRest against the current primary from the shared backup container. This is the Patroni-cluster counterpart of pg snapshot (single instance).

# full backup (default --type full); --tail-logs streams pgBackRest output
pg ha snapshot create app --tail-logs

# incremental (changes since the last backup)
pg ha snapshot create app --type incr

# differential (changes since the last full)
pg ha snapshot create app --type diff

# list every snapshot for the cluster (Start / Stop / Name / Type)
pg ha snapshot list app
pg ha snapshot list app --limit 5

# delete a specific snapshot (label from list) — prompts unless --force
pg ha snapshot delete app 20260916-140131F_20260917-015154I
pg ha snapshot delete app <label> --force

Notes:

  • Before each command, pg ha snapshot re-renders the backup-container pgbackrest.conf from the current topology, so member adds/removes and port changes are reflected in the stanza immediately — the file is bind-mounted, no recreate.
  • You can only delete a non-unique full: pgBackRest keeps at least one full, so deleting the only one is refused (take a new full first).
  • create connects to whatever host is leader at the time; after a failover it still succeeds — that is the point of the cluster-wide stanza.

Pitfall: a stuck replica LSN → patronictl reinit

Symptom. In pg ha status app one replica’s Replay LSN trails its own Receive LSN (or visibly differs from another replica) and never converges over time. Note Patroni’s Lag column is relative to the leader, so every number moves when the leader’s LSN advances; to spot a real stall compare a replica’s own recv vs replay, or watch whether its Replay LSN stays frozen across checks.

Confirm. Three signals point to the same conclusion — the replica’s WAL replay is at a hole and streaming has actually stopped:

# 1) leader has no replication clients, slots inactive
pg exec --dsn "postgres://<user>@<leader_host>:<port>/postgres" \
  "SELECT client_addr, state FROM pg_stat_replication"
pg exec --dsn "postgres://<user>@<leader_host>:<port>/postgres" \
  "SELECT slot_name, active, confirmed_flush_lsn FROM pg_replication_slots"

# 2) the stuck replica has no walreceiver
pg exec --dsn "postgres://<user>@<replica_host>:<port>/postgres" \
  "SELECT status FROM pg_stat_wal_receiver"

# 3) the replica container log loops the same two lines (pg logs ha fetches by
#    scope + member name; -f follows, -n takes more lines):
pg logs ha <scope> -m <member> -n 50 | grep -iE "prev-link|waiting for WAL"
#   LOG: record with incorrect prev-link ... at 0/20000060
#   LOG: waiting for WAL to become available at 0/20000078

record with incorrect prev-link + a repeating waiting for WAL to become available means the WAL stream has a hole/discontinuity at a segment boundary and the startup process is dead-waiting for the next segment. Patroni often keeps logging “no action. I am … a secondary, and following a leader” each cycle — it believes it is following while the postmaster’s walreceiver has actually stopped, so it will not self-heal. This is a replication runtime issue, unrelated to backups.

Fix: reinit the replica. Have Patroni take a fresh pg_basebackup from the leader and restart replication. This destructively rebuilds that replica’s data directory (the leader and other replicas are untouched), so Patroni requires --force:

# CLUSTER_NAME is the namespace-qualified scope (the name after "Cluster:" in pg ha status)
pg ha ctl <scope> -- reinit <scope>-<ns-suffix> <member> --force
# example (namespace default -> -default suffix):
pg ha ctl app -- reinit app-default node1 --force

-f does not work: patronictl reinit only accepts the full --force (unlike switchover/failover’s --yes). Success: reinitialize for member node1 means it was dispatched.

Verify recovery. reinit wipes and rebuilds the data dir, then basebackups + replays. Poll until it is recv == replay and status=streaming:

pg exec --dsn "postgres://<user>@<replica_host>:<port>/postgres" \
  "SELECT pg_last_wal_receive_lsn() recv, pg_last_wal_replay_lsn() replay,
          (SELECT status FROM pg_stat_wal_receiver) wr"
#   expected: recv == replay, wr = streaming

pg ha status app    # the replica's Replay LSN catches up, Lag converges

After reinit completes, run pg ha snapshot create app --type full once to realign the snapshot baseline with the healthy topology.

6 - Patroni Cluster Restore

Point-in-time recovery (PITR) for a Patroni HA cluster with pg ha restore: the custom-bootstrap mechanism, the leader-locality pre-check, the end-to-end runbook, and the post-restore re-baseline steps

This page is the restore-side companion to Patroni Cluster Backup. Once a cluster has a stanza with WAL archiving (pg backup setup + pg ha snapshot), pg ha restore brings the whole cluster back to a point in time — the Patroni counterpart of the single-instance pg restore. It walks the mechanism (why a cluster PITR is not just pgbackrest restore), the safety pre-check on where the leader lives, an end-to-end runbook, and the re-baseline steps that must follow. Command/flag details also live in Restore → Patroni cluster; this page strings them into a procedure you can run.

Prerequisite: a stanza with WAL archiving, on this host

PITR can only replay up to what was archived. Before restoring, the cluster must already have:

  • a cluster stanza provisioned against an S3 repo — pg backup setup --s3-endpoint ... (see Backup → S3 repository and HA backup);
  • at least one full snapshot, plus the WAL segments covering the target time (taken with pg ha snapshot create <scope> --type full).

The target time is bounded by that history: earlier than the newest full backup’s stop time, or later than the last archived WAL, and the recovery cannot land there.

pg backup setup must have run on this host, and the backup container must be up. This is stricter than it sounds for a cross-host cluster:

  • The member that carries the bootstrap restores from S3 using the repo config mounted into its container — pgbackrest-archive.conf at /etc/pgbackrest.conf plus the S3 CA at /etc/pgbackrest/ca.crt. Both are generated by pg backup setup, and the container only mounts them when the files exist. A host that never ran setup has no repo config to mount, so the pgbackrest restore bootstrap cannot reach S3 and fails to start.
  • The leader-locality and stop-time pre-checks run pgbackrest ... info inside the backup container. pg ha restore refreshes the shared pgbackrest.conf best-effort but does not create the member archive config and does not start the backup container — so bring it up first (pg backup setup, or pg backup status to confirm it is Up). If the backup container is down, the pre-checks are skipped and the real failure surfaces only during the bootstrap; the post-restore fresh snapshot needs it running anyway.

In short: run pg ha restore on a host where pg backup setup has provisioned the S3 repo and the backup container is Up.

Mechanism: custom-bootstrap PITR, not a bare pgbackrest restore

A Patroni cluster cannot be PITR’d by running pgbackrest restore under a live postmaster — Patroni owns the data directory, the timeline, and the DCS identity. pg ha restore instead follows Patroni’s custom-bootstrap recipe:

  1. Pause the cluster (patronictl pause). This freezes failover before anything is stopped: stopping the local leader while the cluster is live lets Patroni promote a remote replica this host cannot stop, and that new remote leader then races the local bootstrap to re-claim the DCS on the old timeline — defeating the leader-locality pre-check from inside the destructive path. A restore re-run against a cluster a previous attempt already tore down (DCS cleared, members stopped) has nothing live to pause; that is the already-frozen state, so the pause failure on an empty DCS is tolerated and the run continues.

  2. Clear the cluster’s DCS identity (patronictl remove). Patroni’s bootstrap.dcs block only runs when the DCS has no config key, so removing it is what re-arms the bootstrap path.

  3. Pick one LOCAL member and rewrite its patroni.yml so the default initdb bootstrap is replaced by a pgbackrest restore method targeting the point in time:

    bootstrap:
      method: pgbackrest
      pgbackrest:
        command: 'pgbackrest --stanza=pgcli_<scope> --type=time --target="<time>" --target-action=promote --delta restore'
        no_params: true
        keep_existing_recovery_conf: true

    no_params: true stops Patroni appending --scope/--datadir (which pgbackrest rejects with [031] invalid option); keep_existing_recovery_conf: true preserves the recovery.signal + restore_command + recovery_target_* that pgbackrest restore writes itself. The method/pgbackrest keys are siblings of bootstrap.dcs, never nested inside it — recovery GUCs leaked into the DCS config would apply as live settings on the promoted leader.

  4. Wipe that member’s PGDATA and recreate its container. Patroni sees an empty data dir with no DCS config, runs the method, recovers to the target on a new timeline, and — via --target-action=promote — promotes itself to the writable leader.

  5. The remaining members rejoin the new leader: this host’s other local members are wiped and restarted as standard replicas (they pg_basebackup the new leader); cross-host members rejoin through the DCS (or are reinitialized).

The leader-locality pre-check

pg ha restore refuses to run unless the current leader is a local member of this host. Before touching anything it reads the leader from the DCS and checks it is one of the members this host can control.

Why. Step 1 clears the DCS and step 3 re-bootstraps a local member on a new timeline. A leader that lives on another host keeps its Patroni daemon running — this host cannot stop it — so after the DCS key is removed that remote leader races to re-claim leadership on the old timeline, colliding with the local bootstrap. Requiring leadership to be local (so it can be stopped) removes the race.

What you see.

# dry-run: a loud warning, no refusal — you can still inspect the plan
pg ha restore app --time "2026-08-26 15:30:00+00" --dry-run
#   [!!] current leader "node3" is NOT a local member — the real run REFUSES until
#        leadership is on this host...

# real run on a host whose leader is remote: hard error, nothing is touched
pg ha restore app --time "2026-08-26 15:30:00+00"
#   current leader "node3" is not a local member of scope "app" ...
#   Move leadership here first (pg ha switchover app ...) or run on the leader's host

Fix. Move leadership onto a member of the host you are standing on, then retry:

pg ha switchover app    # or: pg ha failover app ...
pg ha status app        # confirm the leader is now a local member

Or run pg ha restore on the leader’s own host instead. An empty / unreadable leader (e.g. the very first bootstrap, before any member has published itself) does not block the run.

Differences from single-instance pg restore

  • It always promotes. There is no read-only “hold, inspect, retry another time” two-step — the bootstrap recovery ends on a writable leader on a new timeline. Confirm the target with --dry-run first. (The patronictl pause in step 1 of the mechanism is unrelated: it freezes failover during the restore, it is not a read-only inspection window and the cluster is resumed implicitly when the new bootstrap republishes the DCS config.)
  • There is no --promote flag — promotion is automatic (it is baked into the bootstrap command).
  • It must run on a host that owns a member, and the leader must be local (see the pre-check above).

Runbook: restoring a cluster to a point in time

The full flag set is --time (required, all the formats in Restore), --member, --dry-run, --tail-logs, --force.

1. Confirm the target is safe — dry-run. Nothing is touched; you see the stanza, the chosen bootstrap member, the exact pgbackrest command, and (if applicable) the leader-locality warning.

pg ha restore app --time "2026-08-26 15:30:00+00" --dry-run

2. Make sure the leader is local (only if the dry-run warned). Switchover, then re-check status.

pg ha switchover app
pg ha status app

3. Restore. The default prompts for confirmation, listing scope, stanza, target, bootstrap member, and the “PERMANENTLY LOST” warning. Stream the recovery logs with --tail-logs; skip the prompt with --force for automation.

pg ha restore app --time "2026-08-26 15:30:00+00" --tail-logs
# or, picking the bootstrap member explicitly:
pg ha restore app --time "2026-08-26 15:30:00+00" --member node1 --force

4. Verify the new cluster comes up. waitForLeader polls the DCS until a leader appears (15-minute ceiling), then rejoin members start as replicas.

pg ha status app                     # new leader on a switched timeline, replicas streaming
pg exec --dsn "postgres://<user>@<leader_host>:<port>/postgres" \
  "SELECT ... FROM your_table"       # data stops at the target time; later commits gone

After the restore: re-baseline the new timeline

The restore is not the last step. The data directory was rebuilt, so two follow-ups are required before the cluster is backup-ready again:

  1. Take a fresh full snapshot so future PITR has a base on the new timeline:

    pg ha snapshot create app --type full

    This usually succeeds directly: a pgBackRest restore of the same stanza preserves the PostgreSQL system-id across the timeline switch (verified), so the post-restore snapshot does not hit [051] system-id ... do not match stanza. You do not need pg backup stanza-upgrade after a cluster restore. (A real system-id change — a fresh initdb into the same stanza, e.g. a brand-new cluster reusing an old stanza name — is what [051] and stanza-upgrade are for.)

  2. Bring back any member that did not auto-rejoin — typically a cross-host member that lost the old leader. Reinitialize it from the new leader (destructive to that replica’s data dir only):

    pg ha ctl app -- reinit app-<ns> <member> --force

    (Full details in HA backup → stuck replica reinit.)

Troubleshooting

  • “no local member on this host” — you ran it on a host that owns none of the scope’s members. pg ha restore can only rebuild a local data dir + container. Run it on a host that owns a member, or on the leader’s host.
  • “target time is before the latest backup stop time” — the earliest usable point is the newest full backup’s stop time; the error prints it and a suggested --time. Take a fresher full snapshot if you need a later base.
  • Recovery exceeds the last archived WAL — you cannot restore past what was archived. Reduce the target time, or ensure archive_command is flowing (pg backup status) before retrying.
  • First post-restore snapshot fails [051] — not expected: a same-stanza restore keeps the system-id (see the re-baseline section). If you do hit it, a pg backup stanza-upgrade fixes it non-destructively.
  • FATAL: recovery ended before configured recovery target was reached on a repeat PITR into the same repo — the stanza already hosts a previously promoted timeline, and the default recovery_target_timeline=latest jumps onto that branch, which has no commits before your target. Restore again with an explicit older timeline pinned on the command, e.g. --recovery-option=recovery_target_timeline=<old-tl> (added to the bootstrap pgbackrest ... restore command in the member’s patroni.yml). A first PITR after the target — no branch crossing — never hits this.

7 - HA Exec / psql

Run SQL against a Patroni HA cluster with pg ha exec and pg ha psql: zero-config leader resolution from the DCS, no hand-assembled dsn, no entering a member container, and read-only or cross-host targets via –member

pg ha exec and pg ha psql are the direct way to run SQL against a Patroni cluster. They are the cluster-side twins of the single-instance pg exec and pg psql: you name a scope, they resolve the target and connect. No --dsn to hand-assemble, no member container to podman exec into, and — because the connection is plain TCP — a leader or replica on another host is just as reachable as a local one.

Why not pg exec --dsn or podman exec

Before these commands, ad-hoc SQL on a cluster meant one of two awkward paths:

  • pg exec --dsn postgres://postgres:<pass>@<leader>:<port>/db — correct, but you assemble it by hand: read the leader’s connect_address out of pg ha status, copy the superuser password out of the config, plug in the port. A fixed dsn also goes stale the moment the leader moves.
  • podman exec -it <member> psql ... — reaches only a local member, so a leader on another host is out of reach; and you are inside a container the tool is meant to abstract away.

pg ha exec/psql fold both away: pgcli resolves the leader from the cluster’s own state, authenticates with the stored password, and runs psql in a short-lived container. You never see it.

How the target is resolved

The DCS roster (patronictl list -f json) is the only membership view that spans the whole cluster. Each host’s pg.yaml registers only its own members, so this host’s config cannot even see a node running elsewhere — but the DCS knows every member and its advertised connect_address. Both commands go through it:

  • Default (no --member) → the current leader. If leadership has moved since you last looked, you still hit the right node — there is no stale dsn.
  • --member <name> → that named member, wherever it lives. Aim it at a replica for a read-only query (SELECT off a standby, pg_is_in_recovery() returns t), or at a specific remote node.

Authentication is the cluster’s superuser over scram; pg_hba (host all all all scram-sha-256) accepts any source, so a remote member is reachable across hosts exactly the way replication traffic already is.

Usage

# One-shot SQL against the leader
pg ha exec app "SELECT version()"
pg ha exec app "SELECT count(*) FROM pg_stat_activity"

# A different database on the leader
pg ha exec app --database mydb "SELECT * FROM t LIMIT 5"

# Read-only: target a replica instead of the leader
pg ha exec app --member node2 "SELECT pg_is_in_recovery()"

# Interactive psql against the leader
pg ha psql app

# Interactive against a specific (even cross-host) member
pg ha psql app --member node3

# Pass extra psql arguments after --
pg ha psql app -- -c "SELECT 1"
pg ha psql app --database mydb -- -x

pg ha exec takes the SQL as trailing words (so you normally quote one string); pg ha psql opens an interactive shell and forwards anything after -- straight to psql. --database selects the target database on both (default postgres). There is no --user — the managed superuser is the only role pgcli holds credentials for, so that is what connects.

Interactive vs. scripted

pg ha psql allocates a TTY when your stdin is one; when it is not (piping a script, running under CI) it turns the pager off so the session terminates instead of blocking in less. That makes both of these behave as expected:

pg ha psql app < migrations.sql      # pipe a file, no pager, exits when done
printf '\conninfo\n' | pg ha psql app   # scripted one-liner, output streams back

Reading the output

pg ha exec streams psql’s own formatting — column headers, alignment, row counts, and errors go straight to your terminal, exactly as psql -c would print them. It is not a machine-parseable dump; if you need one, pipe through your own tooling or use pg ha exec app "..." --csv-style psql flags via pg ha psql app -- -c "...".

Where these fit

For a stable client endpoint that survives failover independently of pgcli, put HAProxy in front of the cluster and connect through it. pg ha exec/psql are for operators and scripts driving a cluster directly, not a substitute for a pooler in front of an application.

The pg ha status / pg ha ctl verbs remain the way to inspect cluster state and reach unwrapped patronictl commands; these two are specifically about running SQL.

8 - Example: HA Cluster with a Self-CA MinIO

A complete, verified end-to-end walkthrough: a single-member Patroni cluster backed up to a MinIO addon serving a bring-your-own (self-signed) domain certificate, run in a fully isolated pgcli environment

This page is a worked example, run end to end and verified: create a Patroni HA cluster, stand up a MinIO addon that serves a bring-your-own domain certificate (a self-signed CA of our own making — the same shape as a public-CA or corporate-CA cert, just without buying one), and point the cluster’s pgBackRest backups and WAL archiving at it. Every command below is a real one from the run; only hostnames, ports, passwords, and the certificate material are genericized.

Two ideas make the example worth reading rather than skimming:

  • BYO TLS on the store. The MinIO addon does not have to serve pgcli’s generated self-signed pair — --tls-cert / --tls-key let it serve a cert you already hold. A client that trusts the issuing CA needs no extra material; with a private (here: self-signed) CA you hand that CA to the backup stack via --s3-ca-file, exactly as with the generated one. See MinIO → Bring your own certificate.
  • Isolation via a second config file. backup.repo.s3 is a single, global repository shared by every stanza in one pgcli environment. Repointing it at a new store would silently redirect the backups of every existing cluster on that host too. The clean way to try something new against a live estate is a separate config file with its own base_dir, namespace, and port ranges — which is what this example does.

The environment we built

Piece Value Notes
Config ~/.pgcli-app1/pg.yaml separate from the production ~/.pgcli/pg.yaml
Base dir /home/fish/bucket/pgcli-data-app1 nothing touches the production data dir
Namespace app1 prefixes every container name — pgcli-minio-app1-store1, pgcli-patroni-app1-app1-nodea, pgcli-backup-app1
DCS existing etcd m1 at 127.0.0.1:2379 created on the production config long before this example — see Prerequisite; reused here as an external endpoint, and cluster membership is keyed by scope, so a new scope is naturally isolated
MinIO store1, ports 9010/9011, listen 0.0.0.0 BYO self-signed cert, SANs minio1.test,127.0.0.1,<host-ip>
Patroni scope app1, member nodea, PG <host-ip>:35632 single member; the leader by definition
Backup container pgcli-backup-app1 coexists with the production pgcli-backup-default

Prerequisite — the etcd DCS

The example reuses the etcd that already serves this host’s production clusters. If you are starting from scratch, that first piece to install is the same addon, one member at a time — a lone member is a healthy one-node etcd cluster, plenty for HA dev/test:

pg addon install etcd --name m1
# pgcli-etcd-default-m1, client 127.0.0.1:2379, peer 2380, cluster "pgcli-etcd"

(For members other hosts will join as a DCS, give it a reachable address at install time — --advertise-host <host-ip> — so its peer/client URLs announce that IP instead of loopback; and grow it to a real 3-node ensemble with --name m2/--name m3 the same way. None of that matters for this example: the cluster lives on the same host and dials 127.0.0.1:2379.)

Step 0 — a domain certificate

Any PEM leaf + key works. For the demo we minted a self-signed one with pgcli’s own pg cert (validity via --valid-duration, SANs covering the name and IPs clients will dial):

mkdir -p ~/.pgcli-app1/certs
pg cert --host "minio1.test,127.0.0.1,<host-ip>" --valid-duration 8760h \
    --cert-file ~/.pgcli-app1/certs/store1.crt --key-file ~/.pgcli-app1/certs/store1.key
chmod 600 ~/.pgcli-app1/certs/store1.key

openssl x509 -in ~/.pgcli-app1/certs/store1.crt -noout -subject -issuer -ext subjectAltName -dates
# subject=...
# issuer=...                         <- self-signed: leaf and root are one and the same
# DNS:minio1.test, IP Address:127.0.0.1, IP Address:<host-ip>

pg cert writes a single self-signed leaf (CA:FALSE, server-auth) that is its own trust anchor — exactly the shape a public- or corporate-CA cert has, minus having bought one. With a public-CA (or corporate-CA) cert this step is simply “you already have the files” — point the flags at your fullchain.pem (leaf first, then intermediates) and key, and the client-side story gets easier: chains rooted in a trusted CA need no --s3-ca-file at all.

Step 1 — an isolated environment

mkdir -p ~/.pgcli-app1
pg config init -o ~/.pgcli-app1/pg.yaml \
    --base-dir /home/fish/bucket/pgcli-data-app1 \
    --namespace app1 \
    --pg-start-port 36500 --pg-ssh-port 43500

--namespace makes every container name this config creates distinct, so two environments live side by side in one podman. The config file is passed with -c on every subsequent command — from here on, every pg in this page means pg -c ~/.pgcli-app1/pg.yaml.

Port ranges. The default port pools (etcd 2379, MinIO 9000, Patroni 35532, …) start at the same values as the production config, and podman publishes the loopback ports of a host-networked container — so pass explicit, disjoint ports everywhere rather than letting both environments auto-assign the same numbers.

Step 2 — MinIO with the BYO cert

pg addon install minio --name store1 \
    --api-port 9010 --console-port 9011 --listen 0.0.0.0 \
    --tls-cert ~/.pgcli-app1/certs/store1.crt \
    --tls-key  ~/.pgcli-app1/certs/store1.key
-> MinIO BYO cert: CN="" issuer="" valid 2026-09-18 → 2027-09-18
  [OK] TLS certs (BYO: /home/…/.pgcli-app1/certs/store1.crt, valid 2026-09-18 → 2027-09-18)
         self-signed: clients pin the issuing CA (pg backup setup --s3-ca-file <ca.pem>)
  [OK] MinIO container started

Install-time validation pairs the cert and key (a mismatched, expired, CA-instead-of-leaf, or non-server-auth pair fails right here), prints the validity window, and — because the cert is self-signed — tells you clients will pin it as their trust anchor. --listen 0.0.0.0 rather than loopback: the backup container reaches the store over host networking via the host IP, and that IP is one of the SANs. A quick proof it serves your certificate:

curl -s --cacert ~/.pgcli-app1/certs/store1.crt \
     https://127.0.0.1:9010/minio/health/live -o /dev/null -w "%{http_code}\n"
# 200  — and curl verified the chain against exactly the cert we supplied

Create the repository bucket (pgBackRest does not create it):

pg mc alias set store1 https://127.0.0.1:9010 admin '<root-password>' -- --insecure
pg mc mb store1/pgbackrest

Step 3 — the Patroni cluster

pg ha create app1 --member nodea \
    --etcd-endpoints 127.0.0.1:2379 \
    --host-port 35632 --restapi-port 8028 \
    --advertise-host <host-ip>
!  No S3 backup repo configured on this host (backup.repo.s3): the member will
   be created WITHOUT WAL archiving ... Configure a repo and run `pg backup
   setup` to wire archiving in (it recreates members as needed).
  [OK] Patroni member nodea started
✓ Patroni member "nodea" installed in scope "app1"
  scope (DCS):  app1-app1        # namespace-qualified: "app1-" + scope "app1"
  pg:           <host-ip>:35632

The warning is the designed order of operations, not a problem: at create time no repo exists yet, so the member comes up without archiving; the backup setup in the next step configures the repo and recreates the member with archiving enabled. Note the DCS scope is app1-app1 (namespace prefix + scope) — pass the plain scope (app1) to every pg ha command and pgcli resolves the qualified form; it also keeps the stanza distinct from any app1 another namespace might run on the same bucket.

Seed something worth backing up (this also proves client auth over scram from the host):

pg ha exec app1 "CREATE TABLE IF NOT EXISTS byo_probe(id serial primary key,
    note text, ts timestamptz default now());
    INSERT INTO byo_probe(note) SELECT 'row-'||g FROM generate_series(1,500) g;"
pg ha exec app1 "SELECT count(*) FROM byo_probe"   # 500

Step 4 — point the backup stack at the BYO store

pg backup setup \
    --s3-endpoint <host-ip>:9010 --s3-bucket pgbackrest \
    --s3-access-key admin --s3-secret-key '<root-password>' \
    --s3-ca-file ~/.pgcli-app1/certs/store1.crt

--s3-ca-file takes our own cert as the trust anchor — for a self-signed cert the leaf embeds its own root, so the file you passed to --tls-cert is the CA file. pgBackRest passes the bundle to its verifier with no content inspection, so a public-CA cert needs no flag at all and a private-CA one just gets its chain here. The run then proves the full stack and rewires the member:

  [OK] <host-ip>:9010 reachable
  [OK] repo CA published to 1 cluster registry/registries (from …/store1.crt)
  [OK] repository reachable                      # a repo-ls over TLS: creds + CA + bucket all good
  [stale] app1/nodea: mounts the backup-side pgbackrest.conf (pg1-host breaks archive-push)
-> Pausing cluster app1 (no auto-failover during recreate)
  -> Recreating member nodea...                  # now with archive-push enabled
  [OK] app1: archive-ready after recreating 1 member(s)
  [OK] stanza pgcli_app1-app1 created
  [OK] check pgcli_app1-app1

The “repo CA published to the cluster registry” line is the etcd-distribution path: every other member of this scope (or future ones, like a second host) picks the CA up from the DCS with no manual scp and no extra flags.

Step 5 — back it up, and prove it landed

pg ha snapshot create app1 --type full
#   Name:    20260918-091519F
#   Type:    full
pg ha snapshot list app1                          # pgBackRest sees the backup in the repo

The independent proof is that the objects are physically in the bucket, written over a TLS session terminated by our certificate:

pg mc ls store1/pgbackrest -- --recursive
# …/backup/pgcli_app1-app1/20260918-091519F/backup.manifest
# …/backup/pgcli_app1-app1/20260918-091519F/pg_data/base/…  (.zst files)
# …/archive/pgcli_app1-app1/18-1/0000000200000000/000000020000000000000005-….zst
# …/archive/pgcli_app1-app1/18-1/0000000200000000/000000020000000000000005.00000028.backup
# …/archive/pgcli_app1-app1/archive.info

archive/ is continuous WAL archiving (the archive-push the recreated member runs), backup/ is the full snapshot. Confirm the archiver is healthy:

pg ha exec app1 "SELECT archived_count, failed_count, last_archived_wal
                 FROM pg_stat_archiver"

One subtlety from our run, worth knowing so you don’t chase a ghost: failed_count was 1, for 00000002.history — the timeline-history file, archived one second later successfully. It raced the brief window in which backup setup was pausing and recreating the member. A lone recent failure during a setup/recreate is self-healing; a growing count with the same WAL name is a real repo problem.

What the example demonstrates

  • A MinIO addon serving a bring-your-own certificate works as a pgBackRest repository end to end: stanza-create, check, full backup, and continuous WAL archiving all flow over TLS verified against the supplied cert.
  • Self-signed is not a lesser path: the same --s3-ca-file slot carries it (leaf==root), a public CA’s chain, or an internal CA bundle — the transport and the trust plumbing are identical.
  • backup.repo.s3 is one global repo per config file. To experiment beside a running estate, don’t repoint it — init a second config (pg config init --namespace … --base-dir …) and run the whole experiment there. Namespace-qualified container names, DCS scopes, and stanzas keep the two environments from ever seeing each other.

Teardown

Everything the example created lives under the app1 config, so teardown is ordered through that same -c (data directories deleted explicitly):

pg -c ~/.pgcli-app1/pg.yaml ha remove app1 --scope-all --clean-data
pg -c ~/.pgcli-app1/pg.yaml backup remove --clean-data
pg -c ~/.pgcli-app1/pg.yaml addon remove minio --name store1 --clean-data
rm -rf ~/.pgcli-app1 /home/fish/bucket/pgcli-data-app1

--scope-all (not --member) is what clears the scope’s DCS keys in the shared etcd — patronictl remove plus the pgcli member registry. Removing members one by one leaves both behind.

The etcd m1 belonged to the production environment — it is untouched by all of the above, which is exactly the point of the isolation.

9 - Generating a Certificate — pg cert

pg cert mints a self-signed certificate with DNS and IP SubjectAltNames for any TLS you need without a CA — dev/test servers, internal endpoints, and a MinIO serving a bring-your-own TLS certificate. Its flags, what it writes (a single self-signed leaf, not a chain), and how it acts as its own trust anchor

pg cert mints a self-signed certificate — its SubjectAltNames covering any mix of DNS names and IP addresses you ask for — for any place you need TLS without going through a CA: a dev or test server, an internal endpoint, a service whose clients you control. The certificate it writes is a normal TLS server leaf, usable anywhere.

The use this docs tree exercises is pgBackRest’s S3 path: a Patroni cluster’s backups push to an S3 store (MinIO or silo, typically) that must serve TLS — pgBackRest refuses plaintext S3 — and pgcli’s MinIO/silo addons can serve a bring-your-own certificate (pg addon install minio --tls-cert ... --tls-key ...). The quickest way to get a cert to try that with is pg cert:

pg cert --host "minio.test,127.0.0.1,10.0.0.9" \
  --cert-file minio.crt --key-file minio.key

Nothing is registered with pgcli — it touches no pg.yaml and starts no container; it just writes the two files you name and prints the SANs. See MinIO → Bring your own certificate for the serving side, and the worked HA + self-CA MinIO example for it end to end.

What it produces

A single self-signed leaf certificate, not a chain:

  • one CERTIFICATE block in --cert-file — the cert signs itself;
  • CA:FALSE, extended key usage serverAuth — a proper TLS server leaf;
  • SANs split automatically: any --host entry that parses as an IP lands in the IP SANs, the rest in the DNS SANs, so "minio.test,10.0.0.9" needs no special syntax (and *.wild.test works as a wildcard DNS entry).

Because it is self-signed, the leaf is its own trust anchor — hand the same .crt to any TLS client that lets you point at a CA file or bundle (curl --cacert, a browser’s import, an app’s SSL_CERT_FILE …). For pgBackRest’s S3 path specifically that means backup.repo.s3.ca_file / pg backup setup --s3-ca-file minio.crt. A remote host that never saw the file needs no scp either: pg backup fetch-ca <endpoint> recognizes a self-signed leaf in the TLS handshake and saves it back as the anchor. This is not a workaround pgBackRest merely tolerates: OpenSSL treats whatever you hand its trust store as an anchor regardless of CA:TRUE, and pgBackRest’s S3 path (curl over OpenSSL) is that mechanism — verified directly, openssl verify -CAfile minio.crt minio.crt returns OK.

With a real public- or private-CA certificate none of the self-anchor machinery applies: a chain rooted in a trusted CA needs no ca_file at all, and you already hold the issuing CA for a private one.

Flags

Flag Default Meaning
--host 127.0.0.1 Comma-separated DNS names and/or IPs for the SAN (repeatable). IPs are auto-detected; *.wild.test works as a wildcard.
--cert-file cert.pem Path to write the PEM certificate.
--key-file key.pem Path to write the PEM private key (PKCS8, mode 0600).
--valid-duration 825 days (19800h) How long the cert stays valid, e.g. --valid-duration 8760h for a year.
--ecdsa P-256 Curve: P-224/P-256/P-384/P-521. Set to "" to disable ECDSA (then pass --rsa).
--rsa (off) RSA key size (e.g. 2048, 4096); set only when you specifically need RSA instead of the default ECDSA key.
--ca false Issue a self-signed CA (CA:TRUE, keyCertSign) instead of a server leaf — see below.

Selecting no key type at all (--rsa 0 --ecdsa "") is an error: exactly one of them must produce the key.

The --ca mode is not for --tls-cert

--ca makes the cert its own Certificate Authority — for when you want a private root to go sign further certs with. It is not something to hand to MinIO’s --tls-cert: ValidateBYOCert inspects the file’s first certificate and rejects one that is a CA, because a CA’s private key is a signing key, not a server key. The default (no --ca) is the server leaf you do want.

Standalone binary: gencert

The same generator is built as a standalone program for use outside pg:

make gencert          # -> bin/gencert
bin/gencert -host "minio.test,10.0.0.9" -cert-file minio.crt -key-file minio.key

Identical flag set, single-dash form (-host, -cert-file, …). Both front the one internal/certgen package, so their output is the same kind of certificate — pg cert is just the path that keeps you inside the CLI.

Example: full BYO chain in one host

# 1. mint a cert the MinIO will serve (SANs cover how clients dial it)
mkdir -p ~/.pgcli/certs
pg cert --host "minio.test,127.0.0.1,<host-ip>" --valid-duration 8760h \
  --cert-file ~/.pgcli/certs/store.crt --key-file ~/.pgcli/certs/store.key
chmod 600 ~/.pgcli/certs/store.key

# 2. MinIO serves it over HTTPS
pg addon install minio --name store --listen 0.0.0.0 \
  --tls-cert ~/.pgcli/certs/store.crt --tls-key ~/.pgcli/certs/store.key

# 3. the cluster's backups point at it; the served leaf is its own CA file
pg backup setup --s3-endpoint <host-ip>:9000 --s3-bucket pgbackrest \
  --s3-access-key admin --s3-ca-file ~/.pgcli/certs/store.crt

Step 3 references the .crt you minted, because on this host you already hold it. A remote host that never saw the file pulls it back with one handshake instead of an scp — fetch-ca recognizes the self-signed leaf MinIO serves and saves it as the anchor:

# on another host, instead of copying ~/.pgcli/certs/store.crt over:
pg backup fetch-ca <host-ip>:9000
#   [OK] CA fetched from <host-ip>:9000
#        saved:    <base-dir>/backup/repo-ca/ca-<host-ip>-9000.crt
#        SHA-256:  d4df…81e2   (cross-check against the store host's sha256sum)
pg backup setup --s3-ca-file <base-dir>/backup/repo-ca/ca-<host-ip>-9000.crt

10 - S3 Storage High Availability

Making the MinIO/silo object store behind pgBackRest highly available: the four deployment modes pgcli exposes (SNSD/SNMD/MNSD/MNMD), how native multi-drive compares to a ZFS layer for disk redundancy, and the two paths for a distributed cluster whose nodes each hold several disks

A Patroni cluster survives a node loss; the S3 repository its WAL and backups stream into must survive one too, or “high availability” quietly ends at the backup path. The store is a MinIO or silo addon (the two are interchangeable — silo is Pigsty’s MinIO fork with the same feature surface). Two fault domains matter, and they are orthogonal:

Fault Who absorbs it
A disk dies inside a node the drive-level EC under SNMD/MNMD, or the storage layer under MinIO’s data directory (ZFS)
A node dies entirely MinIO’s own erasure coding (EC) across hosts

This page records the shapes pgcli supports for combining the two, and how to pick among them.

The deployment modes

MinIO classifies its layouts by node count and drive count per node (SNSD / SNMD / MNSD / MNMD). pgcli exposes all four:

Mode Shape Use it for
SNSD (single-node, single-drive) one node, one data directory — the default dev, test, demos — and, paired with ZFS below, any single-host deployment that needs disk redundancy
SNMD (single-node, multi-drive) one node, several drives — one --drive per drive a single host with N ≥ 4 data disks that must survive a disk loss without a filesystem layer
MNSD (multi-node, single-drive) ≥ 4 nodes, one data directory per node compact high-availability deployments
MNMD (multi-node, multi-drive) ≥ 2 nodes, several drives each — --drive plus the full --endpoint matrix surviving a disk loss and a node loss without a filesystem layer

SNMD is native: pg addon install minio --drive /mnt/disk1 --drive ... (one flag per drive) starts one MinIO process that erasure-codes across the drives. What it buys, measured on a live 4-drive set: 2 parity shards by default, so it tolerates 2 drive failures; with 1 drive down reads and writes continue, with 2 down reads still succeed but writes are refused (the quorum boundary); usable capacity is about half the raw total; a drive that returns is healed by MinIO itself. The same --drive works on silo.

MNMD is native too: keep --drive for this node’s drives and add the whole cluster’s host×drive endpoint matrix with --endpoint — one URL per drive on every node, each naming that node’s /data1../dataN slot. Measured on a live 4-node × 4-drive set (both minio and silo): 16 drives online report EC:4 in a single erasure set of stripe size 16; losing one whole node (12/16) keeps reads and writes working; losing a second node (8/16) refuses writes and fails reads, and the set self-heals to 16/16 on restart. The full walkthrough is Addons → MinIO → Multi-Node Multi-Drive (MNMD).

SNMD vs ZFS: two ways to survive a disk

Both protect a single host’s data against disk loss; they differ in where the redundancy lives and what that costs.

native SNMD/MNMD (--drive) ZFS pool under --data-dir
quorum unit the drive — MinIO counts drives as failure members invisible to MinIO — one big drive, the pool absorbs disk loss
layout changeable later fixed at install (the drive set is the EC set) freely — swap disks, grow, migrate raidz1→raidz2; --data-dir never changes
usable capacity, 4 disks ~half (2 data + 2 parity) depends on vdev: raidz2 = 2×disk, raidz1 = 3×disk (more usable)
rebuild MinIO heals a returned drive ZFS resilvers locally
serves which modes SNMD (single host) and MNMD (across hosts) SNSD and MNSD (same recipe under each node)
heterogeneous nodes every node must contribute the same drive count a 2-disk node and a 6-disk node look identical (one endpoint each)

The honest tradeoff: native multi-drive is simpler — one command, no filesystem to provision, and it gets you disk redundancy with zero ZFS setup. ZFS is more flexible — the layout stays changeable underneath, it saves more capacity on the same disks (raidz1 keeps 3 of 4 usable where SNMD’s EC:2 keeps 2 of 4), and it tolerates nodes that are not identical. So:

  • single host, want disk redundancy without touching ZFS → SNMD
  • single host, want the most usable capacity / a layout you can change later → SNSD + ZFS
  • distributed cluster, disk and node failure → both paths are first-class now: native MNMD, or MNSD + per-node ZFS — see “HA with 4 hosts, each with several disks” below

Neither is wrong on a single box — SNMD’s 50% capacity is the price of not managing a pool, and ZFS’s flexibility is the price of provisioning one. The rest of this page documents the ZFS path and the MNMD alternative for multi-host sets.

ZFS: the flexible disk layer

The recipe is identical under SNSD and MNSD: build a zpool from the host’s data disks, create one dataset for the store, and point --data-dir at its mount point. MinIO sees “one big reliable drive”; it never learns how many physical disks are under it.

# one pool from the host's spare disks (mount point defaults to /minio-pool)
sudo zpool create minio-pool raidz1 /dev/sdb /dev/sdc /dev/sdd

# one dataset for the store: 1M recordsize suits object blobs, no atime churn
sudo zfs create -o recordsize=1M -o atime=off minio-pool/store

pg addon install minio --name store --data-dir /minio-pool/store   # or: silo

The pool must sit on separate devices from the root filesystem — which is also what MinIO’s drive check and pgcli’s install-time advisory want (drive is part of root drive, will not be used). A ZFS mount on its own disks satisfies that trivially.

Layout by disk count

Data disks SNSD (ZFS is the only defense) MNSD node (EC above absorbs node loss)
1 no redundancy — fine for dev/test, a dead disk means re-seeding the repo one big disk per node is the plain MNSD shape; no ZFS needed
2 mirror mirror
3 raidz1 (2D usable) raidz1
4 raidz2 on spinning disks; raidz1 when SSD rebuilds are quick and the capacity matters; 2 × mirror for write-heavy stores raidz1 — one parity is enough to keep disk loss invisible to MinIO
5–8 raidz2 ((N−2)D usable) raidz1, or raidz2 with large HDDs
> 8 prefer two smaller raidz2/raidz1 vdevs striped over the pool — a resilver across 12+ disks is a long exposure window same: keep any single vdev ≤ ~8 disks

The two columns differ in one line of reasoning: under SNSD ZFS must survive the disk and whatever happens during its rebuild, so parity depth buys safety outright; under MNSD, ZFS only has to keep a disk failure from escalating into a node loss, so one parity plus quick local resilver is the job — and beyond that, prefer more nodes over deeper local RAID.

For a 4-disk host under SNSD — the most common single-box question — the usual choices are:

Layout Usable Survives Character
raidz2 (like RAID6) 2 × disk 2 disks the safe default on spinning disks: resilvers on large drives are long, and raidz1’s single parity does not survive one more loss during a rebuild
raidz1 (like RAID5) 3 × disk 1 disk the capacity pick for SSDs / small disks, where a resilver takes minutes, not hours
2 × mirror (striped) 2 × disk 1 disk per mirror (2 if in different mirrors) best small-write performance — wide raidz is the worst shape for it

Under SNSD ZFS is the store’s only line of defense, so the default leans conservative: raidz2 unless the disks are fast enough to rebuild quickly and the capacity is worth the thinner margin.

MinIO’s own guidance to avoid RAID underneath its EC mode targets the double-redundancy of RAID + cross-node EC. Under SNSD there is no EC — ZFS is the only protection the data has — so the guidance does not apply, and following it literally (“single drive, no ZFS”) on a 4-disk host would mean losing the whole store to one dead disk.

Co-locating with PostgreSQL: ZFS caches aggressively (ARC, by default up to half of RAM). On a host that also runs the database, cap it — e.g. options zfs:zfs_arc_max=8589934592 in /etc/modprobe.d/zfs.conf — so the two workloads do not fight over memory.

HA with 4 hosts, each with several disks

This is the hybrid the layout table above points at, and there are two first-class ways to build it: native MNMD — MinIO erasure-codes across every disk of every node — or MNSD + a ZFS pool under each node’s data directory — MinIO sees one drive per node and the pools absorb disk loss. They fail differently, so the choice is real either way.

Native MNMD

One matrix, no filesystem layer: each node passes its own drives with --drive and every node carries the identical host×drive endpoint list. For TLS, generate one shared leaf and install the same pair on every node (see the TLS note after the measurements below).

# once, anywhere — one self-signed leaf whose SANs cover every node address:
pg cert --host 10.0.0.11,10.0.0.12,10.0.0.20,10.0.0.21 \
        --cert-file grid.crt --key-file grid.key
# copy grid.crt + grid.key to every node — byte-identical files everywhere

# node 1 (10.0.0.11) — four data disks mounted; nodes 2–4: same command, the
# SAME grid.crt/grid.key, own --drive paths, SAME 16-endpoint matrix, SAME
# --root-password:
pg addon install minio --name store \
  --tls-cert grid.crt --tls-key grid.key \
  --listen 0.0.0.0 \
  --drive /mnt/minio/disk1 --drive /mnt/minio/disk2 \
  --drive /mnt/minio/disk3 --drive /mnt/minio/disk4 \
  --root-password '<shared-secret>' \
  --endpoint https://10.0.0.11:9000/data1 --endpoint https://10.0.0.11:9000/data2 \
  --endpoint https://10.0.0.11:9000/data3 --endpoint https://10.0.0.11:9000/data4 \
  --endpoint https://10.0.0.12:9000/data1 --endpoint https://10.0.0.12:9000/data2 \
  --endpoint https://10.0.0.12:9000/data3 --endpoint https://10.0.0.12:9000/data4 \
  --endpoint https://10.0.0.20:9000/data1 --endpoint https://10.0.0.20:9000/data2 \
  --endpoint https://10.0.0.20:9000/data3 --endpoint https://10.0.0.20:9000/data4 \
  --endpoint https://10.0.0.21:9000/data1 --endpoint https://10.0.0.21:9000/data2 \
  --endpoint https://10.0.0.21:9000/data3 --endpoint https://10.0.0.21:9000/data4

Measured on a live 4-node × 4-drive set (the same on minio and silo): the 16-drive set reports EC:4 in one erasure set of stripe size 16 — losing 4 drives of 16 stays healthy.

Event Online Effect (measured)
1–4 drives lost ≥ 12/16 reads and writes continue; returned drives are healed by MinIO
1 node down (its 4 drives) 12/16 reads and writes continue — a 64 MiB round-trip stayed byte-identical
2 nodes down 8/16 writes refused (Resource requested is unwritable), reads fail too
nodes restarted 16/16 self-heals

Usable capacity follows the parity ratio: EC:4 over 16 drives keeps 12/16 of raw bytes. The costs: the layout is fixed at install (the matrix is the EC set), every node must contribute the same number of drives, and pgcli cannot detect a matrix typo on another node — keeping the N pg.yaml files identical is the operator’s job.

For TLS across the grid, serve one shared certificate: mint a single self- signed leaf with pg cert --host <every node address, comma-separated> and install it on every node with --tls-cert/--tls-key (the same two files byte-identical everywhere, the endpoint matrix all https://). The grid then forms with no per-node CA to reconcile — verified on a live 4-node set: the ring came up Network: 4/4 OK with an empty CAs/ dir, because that shared leaf is its own trust anchor and every node already holds it. The fallback is the generated mode: if you let each node run a bare --tls, every host mints its own CA and the first cross-node handshake dies with x509: certificate signed by unknown authority until you seed one shared CA into every node’s cert dir by hand — which the shared-pg cert path avoids entirely.

MNSD + per-node ZFS

MNSD across the hosts, ZFS under each node’s data directory. Each node exposes exactly one endpoint (its ZFS-backed /data); disk failures are healed by the local pool and never reach MinIO; node failures are absorbed by EC quorum.

# on EACH of the 4 nodes: build the local pool (layout per node — see below)
sudo zpool create minio-pool raidz1 /dev/sdb /dev/sdc /dev/sdd
sudo zfs create -o recordsize=1M -o atime=off minio-pool/store

# node 1 (10.0.0.11):
pg addon install minio --name store \
  --listen 10.0.0.11 \
  --data-dir /minio-pool/store \
  --root-password '<shared-secret>' --tls \
  --endpoint http://10.0.0.11:9000/data \
  --endpoint http://10.0.0.12:9000/data \
  --endpoint http://10.0.0.20:9000/data \
  --endpoint http://10.0.0.21:9000/data

# nodes 2–4: same command, own --listen, identical endpoint list and password

What each layer then tolerates on a 4-node cluster:

Event Handled by Effect
1 disk in a node’s raidz pool ZFS (resilver) invisible to MinIO, no quorum math
1 node down EC (writes need ⌈4/2⌉+1 = 3 of 4 online) reads and writes continue
2 nodes down EC reads only (⌈4/2⌉ = 2 of 4) readable, writes refused until a node returns
a disk and its node failing together EC (3 of 4) still safe — the surviving pools heal after the node returns

Two properties make the scheme flexible:

  • Endpoint count = node count, not disk count. A node with 2 disks and a node with 6 disks look identical to MinIO. The four hosts may run different layouts — raidz1 ×4 disks here, 2-way mirror there, raidz2 ×6 over there — heterogeneity costs nothing.
  • Pool capacities should be roughly aligned. EC sets the usable capacity of the whole cluster to the weakest member’s free space (3 × 12T pools + 1 × 4T pool → you get 4T × striping, not 40T). If the disks genuinely differ, thin provisioning (zvol-based sparse datasets, or simply zfs set refquota) hides the mismatch from MinIO — the quota caps the big pools at the small one’s size, and nothing wastes a rebuild.

Why not stripe the per-node pool to reclaim capacity. MinIO’s “no RAID under EC” advice is about capacity, and the arithmetic is real: a plain 4-node MNSD EC set already halves the raw total, so an unmirrored raidz_none pool on each node (all disks striped, zero local redundancy) does squeeze out roughly a third more cluster capacity than raidz1 would. But EC counts a failure in nodes, and a striped pool turns a single dead disk into a whole offline node — one ordinary disk failure burns a slot of the budget EC set aside for losing an entire machine, forcing a network-wide rebuild of that node’s whole pool and leaving zero margin until it finishes. The capacity is only “free” because you quietly downgraded disk fault-tolerance to node fault-tolerance. raidz1 per node is the point where the two layers stop stealing from each other: the pool absorbs disk loss invisibly, MinIO’s EC budget stays reserved for node loss. (If the reclaim-everything answer truly fits, the shape that maximizes it is plain MNSD on one big disk per node — no ZFS at all — not a striped pool that hides the same single-point risk one layer down.)

Which of the two

native MNMD MNSD + per-node ZFS
setup one matrix command per node, no filesystem to provision zpool + dataset on every node first, then one endpoint per node
what EC counts drives — one node down is just 4 of 16 members lost nodes — a disk failure never reaches MinIO at all
measured fault margin 4-node × 4-drive set: OK to 12/16 drives (one full node), fails at 8/16 4 nodes: reads at 2/4, writes at 3/4 nodes
usable capacity 12/16 of raw on that set (EC:4) — the parity is MinIO’s choice for the stripe local raidz1 keeps 3/4 per pool, then EC halves across nodes
layout later fixed at install — the matrix is the EC set changeable — disks, vdevs, layouts move without touching MinIO
heterogeneous nodes not possible: every node must contribute the same drive count natural — each node is just one endpoint, any local layout
TLS one shared cert serves the whole grid: pg cert a single leaf covering every node address, --tls-cert/--tls-key it identically everywhere (no CA to reconcile) the same — it’s still a TLS grid between nodes, so the same shared-leaf setup applies

Pick MNMD when the nodes are identical, the disk plan is settled, and you want disk and node redundancy from MinIO alone with nothing to provision underneath. Pick MNSD + ZFS when nodes differ in disk count or size, the layout may change later, or you want to reason about failures in whole nodes instead of in drives.

Choosing a shape

Situation Shape
laptop / demo / CI SNSD, default data dir — nothing to decide
one host with N ≥ 4 data disks, backups must survive a disk SNMD (--drive × N) — zero filesystem setup; or SNSD + ZFS (raidz2 on spinning disks, raidz1 on fast SSDs) for more usable capacity and a changeable layout
a few hosts, one data disk each, must survive a host MNSD plain
identical hosts with several data disks each, must survive a disk and a host MNMD (--drive + full --endpoint matrix) — no filesystem layer; or MNSD + per-node ZFS, see “Which of the two” above
hosts whose disk counts/sizes differ MNSD + per-node ZFS (layouts may differ per node; keep capacities aligned) — MNMD needs equal drive counts

The store’s TLS story is independent of the topology: whichever shape you pick, --tls (or pg cert-minted BYO certificates) works the same — see Generating a Certificate.