This is the multi-page printable view of this section. .
HA Cluster
- 1: Patroni HA
- 2: Patroni Dynamic Configuration
- 3: Patroni REST API
- 4: Extensions in HA Clusters
- 5: Patroni Cluster Backup
- 6: Patroni Cluster Restore
- 7: HA Exec / psql
- 8: Example: HA Cluster with a Self-CA MinIO
- 9: Generating a Certificate — pg cert
- 10: S3 Storage High Availability
Patroni is the de-facto standard for PostgreSQL
high availability: it owns each postmaster’s lifecycle, streams replication
between members, and performs automatic failover when the leader is lost.
pgcli exposes Patroni as its own top-level command — pg ha — a distinct mode
rather than a plain pg instance or an addon installed with pg addon install.
Note on placement. Patroni pages used to live under Addons. They are listed here as a standalone HA Cluster section because
pg hais not installed via the addon system — the addon index only mentions it for discovery. The etcd DCS and the HAProxy load balancer remain addon pages (etcd, HAProxy); this section links to them where relevant.
Pages
| Page | What it covers |
|---|---|
| Patroni HA | The pg ha command set: create, status, switchover/failover, pause, cross-host members, passwords, namespace & DCS layout |
| Dynamic Configuration | pg ha edit-config / pg ha ctl — the DCS-backed runtime config, and why it never recreates a container |
| REST API | The per-member Patroni REST API: health checks, leader redirects, what HAProxy probes |
| Extensions in HA | Installing/removing PostgreSQL extensions across a cluster (rolling shared_preload_libraries changes) |
| Cluster Backup | pg backup setup + pg ha snapshot — stanza, WAL archiving to S3, cross-host trust |
| Cluster Restore | pg ha restore — cluster PITR via the custom-bootstrap mechanism, the leader-locality pre-check, the post-restore re-baseline |
| Exec / psql | pg ha exec / pg ha psql — run SQL against the leader or any member directly, no dsn, no container exec |
| Example: HA + self-CA MinIO | A verified end-to-end walkthrough: a cluster backed up to a MinIO serving a bring-your-own cert, in a fully isolated second environment |
| Generating a Certificate | pg cert — self-signed cert with DNS and IP SANs for any TLS you need without a CA (dev/test servers, an internal endpoint, or a MinIO serving a bring-your-own cert): its flags, and how the single self-signed leaf acts as its own trust anchor |
| S3 Storage High Availability | Making the MinIO/silo repo itself fault-tolerant: the SNSD and MNSD modes pgcli exposes and why only those two, plus ZFS under the data directory — single-host raidz, per-node pools in a distributed cluster, heterogeneous nodes |
The typical path
Related
- Addons: etcd (the DCS), HAProxy (a stable client endpoint with optional read/write split), MinIO (the S3 repo the backups live in)
- Restore → Patroni cluster for PITR basics shared with single-instance restores
- Failover: Replica Promotion for the non-Patroni, single-replica promotion path (
pg replica) - Namespace Isolation — scopes, stanzas, and container names are namespace-prefixed
1 - Patroni HA
Patroni is the de-facto standard for
PostgreSQL high availability: it manages each postmaster’s lifecycle, streams
replication between members, and performs automatic failover when the
leader is lost. pgcli exposes Patroni as a distinct mode — pg ha — rather than
folding it into the plain pg instance path.
This is a separate mode, not an addon subcommand. Patroni (not pgcli) owns the PostgreSQL process.
pg hamembers do not appear incfg.Instances, share no instance lifecycle code, and are not created bypg create. The one thing it borrows from the addon system is the etcd DCS.
Linux only, like the etcd addon: Patroni members rely on podman host networking. Both root and rootless podman are supported. On macOS the commands fail fast with a clear message.
Ownership boundary
The single most important thing to internalize is who owns what:
| Responsibility | Owner |
|---|---|
Patroni container (run/start/stop/rm), image, patroni.yml, port allocation, passwords, DCS wiring |
pgcli |
postmaster lifecycle, initdb, PostgreSQL config rendering, replication slots, failover |
Patroni |
| switchover / failover / pause / edit-config | pgcli wrapping patronictl in short-lived containers |
pgcli never edits a running PostgreSQL’s config directly, and never runs
initdb — Patroni does both. This is also why the plain PG image’s
docker-entrypoint-initdb.d convention (the admin role / default database)
does not apply here: Patroni bootstraps its own cluster, so the role system
is postgres (superuser), replicator, and rewind_user instead.
Container lifecycle ≠ safe PG restart
Because Patroni is PID 1 inside its container:
pg ha createis re-install semantics (stop + recreate the container). Recreating the leader takes that node fully offline and triggers a failover. Changing a member’s config by re-runningcreateis fine; for dynamic settings preferpg ha edit-config, which never touches the container.pg ha start/pg ha stopare raw container start/stop. Stopping the leader’s container is exactly as disruptive as the node crashing — Patroni will fail over to a replica. For planned maintenance,pg ha pausefirst (it disables auto-failover), then stop the container.- A paused cluster has no automatic failover until resumed.
How It Works
pg ha runs one Patroni container per member. All members of a scope point at
the same DCS (a etcd cluster) which holds the cluster’s dynamic
configuration and leader lock:
- The first
pg ha createfor a scope bootstraps: Patroni runsinitdb, wins the leader race, and becomes the leader. - Every later member automatically
pg_basebackups from the current leader and begins streaming — no flag distinguishes “add” from “join”. pg hacontrol commands (switchover,pause, …) runpatronictlin an ephemeral container against the DCS, so they work even when no member container is running locally.
The DCS is the source of truth: Patroni re-renders each member’s
postgresql.conf/pg_hba.conf from it every loop, so hand-edits on disk are
lost — use pg ha edit-config instead.
Scope and the DCS layout
The <scope> in pg ha create <scope> is the Patroni cluster name — the
identity of one HA cluster. Every command that takes a scope (status,
switchover, failover, pause, ctl, …) names the same cluster; the first
create for a scope bootstraps it, later creates with the same scope add
members. (--member is the per-host node name inside the cluster — one
member per machine.)
In pg.yaml the cluster is keyed by that scope under addons.patroni.<scope>.
What actually lands in etcd
Patroni stores everything under a fixed etcd namespace plus the scope:
- namespace:
/service/— Patroni’s top-level key, not changed by pgcli. - scope: pgcli writes
PatroniScope(scope)intopatroni.yml, which is your scope plus the pgcli namespace suffix. Patroni has no namespace concept of its own (its etcd prefix is the raw scope), so pgcli bakes the suffix into the scope to keep two pgcli namespaces sharing one etcd from cross-talking.
So the full etcd prefix for a cluster is:
The keys below the prefix are Patroni’s own DCS layout. To inspect them, point
pg etcdctl at a running etcd member of the same host/config:
Note the scope shown there carries the namespace suffix, even if you configured the cluster with the bare name.
Install
Prerequisite: a DCS. Either reuse local etcd addon members or point at an external etcd.
pg ha status app renders patronictl list — exactly one Leader row and the
rest Replica … streaming:
With no argument, pg ha status rolls up every cluster: container state, member
count, and each member’s pg=/rest= ports.
Choosing a DCS
Patroni only needs a reachable etcd cluster, so there are two layouts:
-
Co-located — run the etcd members as addons on the same hosts as the Patroni members and pass
--etcd m1,m2,m3. Simplest: one machine per role, no extra hosts. Good for a starter 3-node HA cluster. -
Dedicated DCS hosts — run the etcd cluster on its own (virtual) machines and point every Patroni member at it with
--etcd-endpoints. Each line below runs on a different machine (pgcli is per-host):This is the more available topology: etcd’s quorum survives losing a PG host, and re-imaging a database machine never takes the DCS down with it. The PG hosts only ever talk to the DCS;
--etcd-endpointslists all member client URLs, so Patroni falls through to the next endpoint if one etcd host is down (writes still need the etcd quorum itself — 2 of 3). Keep the endpoint list identical on every PG host.A single endpoint (
--etcd-endpoints 10.0.0.20:2379) does work — that one etcd member serves the whole cluster, and the other two are hidden behind it. But it re-introduces a single point of failure: if E1 goes down, Patroni cannot reach the DCS even though the etcd quorum is healthy, and a leader that fails to renew its lock demotes itself. List every endpoint you actually have.
A dedicated odd-sized etcd cluster (3 or 5) is the production recommendation; co-locating is fine for dev and small footprints.
Commands
| Command | What it does |
|---|---|
pg ha create <scope> --member <m> … |
Register + (re)install one member — recreate = node offline |
pg ha list |
Compact table of all HA clusters (scope, members, DCS, status) |
pg ha status [scope] |
All clusters, or one cluster’s patronictl list |
pg ha switchover <scope> |
Planned leader change (patronictl confirms; --yes to script) |
pg ha failover <scope> |
Promote a replica now |
pg ha pause / resume <scope> |
Disable / re-enable automatic failover |
pg ha edit-config <scope> -- … |
View or patch the dynamic config in the DCS (never recreates a container) |
pg ha start / stop <scope> --member m | --all |
Raw container start/stop (see lifecycle caveat) |
pg ha remove <scope> --member m | --scope-all [--clean-data] [--force] |
Remove member(s); --scope-all also clears the DCS |
pg ha passwords <scope> [--file F] |
Export the stored password set (the --passwords-file format, for other hosts) |
pg ha remote <scope> [--member m] [--ssh-port P] |
Register a member living on another host (typically the cross-host leader) for backup SSH |
pg ha remote remove <scope> --member m |
Undo such a registration (the remote container is untouched) |
pg ha extension install/remove/list/apply |
Install, remove, or list PostgreSQL extensions (see Extensions in HA) |
pg ha exec <scope> "<sql>" [--member m] [--database db] |
Run one-shot SQL against the leader (or a named member) — no dsn, no container exec |
pg ha psql <scope> [--member m] [--database db] [-- <psql-args>…] |
Interactive psql against the cluster, same resolved target |
pg ha ctl <scope> -- <patronictl args…> |
Passthrough to any patronictl command |
Flags after -- reach patronictl verbatim (cobra strips the --), so
pg ha ctl app -- show-config, pg ha ctl app -- topology, and
pg ha edit-config app -- -s synchronous_mode=true --force all work. edit-config --show is a convenience alias for ctl … -- show-config.
Cross-host members
Each host runs its own pgcli managing only that host’s members; the cluster
reassembles through the shared DCS, so two hosts’ pg.yaml files each hold a
partial view of one scope.
Cross-host checklist (all three must line up on every host):
- Ports — auto-assignment is per-host, but
connect_addressis stored in the DCS cluster-wide. Pass--host-port/--restapi-portwith the same value on every host, or replicas can’t reach each other. - Passwords — Patroni’s replication / rewind / REST-API auth is cluster-wide.
Export the first host’s generated set with
pg ha passwords app --file app-passwd.ymland pass--passwords-file app-passwd.ymlto every otherpg ha create. --advertise-host— required for cross-host members; it flips the listen address to0.0.0.0and puts a reachable IP inconnect_address. Left empty, a member is loopback-only. This applies to the first (bootstrap) member too: it becomes the leader, and its loopbackconnect_addressis what every later host would try topg_basebackupfrom — cross-host joins fail forever once the cluster was bootstrapped with the default. If you may ever add a member on another host, pass your LAN IPv4 from the very firstcreate.- Firewall — allow the two ports (PG + REST API) pairwise between members.
pgcli does not validate the peers’ config; pg ha status shows each member’s
connect_address so you can self-check reachability.
Backing up a cross-host leader
Each host’s pg.yaml records only its own members, so this host cannot see a
leader running on another host — yet pgBackRest’s stanza-create and full
backups must connect to the leader.
This now happens automatically: after pg ha create installs a member, it
publishes that member’s pgcli-private ports (SSH / REST API) into an etcd
registry at /pgcli/ha/<scope>/<member> (the topology itself — member name +
host:pgport — Patroni already stores in the DCS). The backup config generators
then run patronictl list for the cluster-wide topology and look up each remote
member’s SSH port in the registry, so pgbackrest.conf / ssh_config include
members on every host, with pg1-host and the SSH HostName pointing at the
remote IP (local members still go via 127.0.0.1). pg ha start/stop/remove,
autostart, and pg ha extension skip members that aren’t local — they belong to
their own host’s pgcli.
Just refresh the backup config:
Cross-host backup trust is now automatic. A cluster-wide stanza lists every
member as a pg*-host, so a full backup / check from any host SSH-probes the
members on all hosts to find the primary — each of those member sshd processes
must therefore accept the initiating host’s backup key, and its
--advertise-host already flipped the listener to 0.0.0.0 (see above) so
cross-host SSH reaches it. pg backup setup
handles this: each host’s create publishes its backup public key into the
same /pgcli/ha/<scope>/<member> registry, setup merges every member’s key
into a per-cluster authorized_keys file, and each member container bind-mounts
that file as an extra AuthorizedKeysFile. Because sshd re-reads it on every
login and the merge rewrites the file in place, a host that joins later is
trusted by the running members without a restart (the file only needs the
one-time recreate to add the mount, which setup does inside the pause window
alongside the archive-config recreate). A member removed with pg ha remove
drops its registry key, so the next merge stops trusting it.
The S3 repository CA is distributed the same way: the host that configured
ca_file publishes the certificate to /pgcli/ha/<scope>/.repo/ca, and a
joiner that leaves ca_file empty pulls it and points its own ca_file at the
pulled copy. When the store itself lives on yet another machine, the first host
gets the CA without any file copy too — pg backup fetch-ca <store>:<port>
pulls it out of the endpoint’s TLS chain (see the MinIO doc). Only public
material — SSH public keys and the self-signed CA —
ever enters the registry; private keys, passwords, and the S3 secret_key never
do (the DCS link is unauthenticated).
Fallback: a remote member created on its host before that host upgraded
pgcli has no registry entry, so the generators fall back to this scope’s base
SSH port for it (every host starts at patroni_ssh_start_port, so the first
member is usually right). When it isn’t, register just that member manually to
override (creates no container):
S3 backups and WAL archiving
Patroni clusters ship with archiving off — full backups alone can never
reach a point-in-time. Enabling the S3 repository
(backup docs, English;
中文) flips it on. Each cluster is one
stanza — pgcli_<scope>, matching the namespaced DCS scope, so two pgcli
namespaces sharing one S3 bucket never collide — whose backup-side config lists
every member as a pg*-host, so pgBackRest finds the primary itself and keeps
working across failovers. pg backup setup renders the identical
archive_command into each member’s local patroni.yml (pgcli never writes
archive GUCs into DCS — that is edit-config’s territory, and Patroni applies
a local value whenever DCS does not manage the key), recreates stale member
containers replicas-first inside a pause window (the recreate is also the
postmaster restart archive_mode needs), and runs stanza-create + check
for the cluster. From then on WAL streams to S3 continuously, from whichever
member holds the leader lock.
Expect a short per-member offline window and a planned leader demotion at
the end — the same recreate semantics as pg ha create.
Passwords
The first pg ha create for a scope generates four credentials — superuser,
replication, rewind, and the restapi basic-auth pair (restapi_user /
restapi_password) — and stores them in pg.yaml under the cluster. You never
pass passwords on the command line: the first member generates them, and
pg ha passwords exports the stored set:
(Without --file the YAML goes to stdout — prefer --file so the secrets stay
out of shell history and scrollback.)
The four roles and their default usernames (fixed by pgcli — you only ever set
the passwords, which default to a random 16-char string per role). Override the
length with --password-length on pg ha create / pg create (8–64; it only
applies to the generate path, not to a --passwords-file or an already-stored set):
| pg.yaml key | Role | Default username | Default password |
|---|---|---|---|
superuser |
PostgreSQL superuser | postgres |
random 16-char, generated |
replication |
replication / streaming | replicator |
random 16-char, generated |
rewind |
pg_rewind role |
rewind_user |
random 16-char, generated |
restapi_user / restapi_password |
Patroni REST API basic-auth | postgres |
random 16-char, generated |
The postgres / replicator / rewind_user usernames are baked into the rendered
patroni.yml (postgresql.authentication); only the passwords are the generated /
--passwords-file part. restapi_user is the one username you can override, via the
passwords file.
For full control (or when you’d rather author the file yourself) use the same
format with --passwords-file:
pg ha create prints where each password came from (generated-and-stored vs. a
file path). The rendered patroni.yml is written mode 0600 because it embeds
all of them.
Auto-start on Boot
Like the other infra addons, a Patroni member’s container can be brought up
after a host reboot — but it only starts the existing container, reading the
patroni.yml already on disk; it never re-renders config or re-creates data.
Members are toggled one at a time (--ha --scope <scope> --name <member>).
Start order relative to the DCS doesn’t matter: Patroni retries until etcd
answers, then re-elects normally. See the autostart page for
the boot-service mechanics.
Logs
Container name is pgcli-patroni-<scope>-<member> (namespace-prefixed when a
namespace is set).
Connecting
pg ha exec and pg ha psql are the zero-config path: pgcli resolves the
leader from the DCS itself, authenticates with the stored superuser password,
and runs psql in a throwaway container — no dsn to assemble, no member
container to enter, and remote members are reachable (plain TCP + scram, so a
replica on another host works too; --member aims at a specific node).
For clients outside pgcli, connect to the leader’s PostgreSQL port. Find the
leader with pg ha status app (the Leader row’s Host is its
connect_address), then point pg psql at it. On a single host with default
loopback-only members, that is 127.0.0.1:<host_port>.
After a failover the leader changes, so a fixed connection string should be avoided unless fronted by a pooler or the Patroni REST API’s leader redirect.
For a stable endpoint that survives failover — and optional read/write separation — put HAProxy in front of the cluster.
Planned leader changes
Patroni owns failover; pg ha wraps the patronictl verbs that move the
leader on purpose. All three take just the scope:
switchoveris the planned, graceful one: the current leader is demoted first, so nothing is in flight when the candidate is promoted. Both members must be healthy and caught up. The old leader automatically rejoins as astreamingreplica a few seconds later (you may briefly see it asstoppedwhile Patroni restarts its postmaster). A new timeline is opened — this is normal, not a split-brain.failoverforce-promotes a replica immediately without a clean handoff from the leader. Use it only when the leader is already gone or you are deliberately discarding it; on a healthy cluster it just causes an unnecessary blip. Reach forswitchoverfor anything planned.pauseturns off automatic failover cluster-wide. Do this before any planned container surgery —pg ha stop --all, apg ha createrecreate, orpg ha extension— so Patroni doesn’t promote a replica out from under you.pg ha resumeputs auto-failover back. (A paused cluster has no automatic failover until resumed.)
These are pure DCS operations — they run patronictl in a throwaway container,
touch no member’s container or data, and work even when some members live on
other hosts. Like every pg ha control command, the scope is resolved to the
namespaced DCS key (app → app-default), so a single-host or cross-host
leader is switched the same way.
Because clients dial the leader’s port, the connection target moves after a switchover — see Connecting and put HAProxy in front if you need one stable endpoint.
pg ha vs. pg replica — which to pick
pgcli has two ways to get a standby:
pg replica + pg failover |
pg ha (Patroni) |
|
|---|---|---|
| Model | Manual: pg replica builds a standby, pg failover promotes it on demand |
Automatic: Patroni keeps N members in sync and self-heals |
| Failover | A human runs pg failover; primary stays down until promoted |
Patroni detects the loss and promotes a replica within ~30s |
| Ownership | pgcli drives postmaster (as with any instance) | Patroni drives postmaster; pgcli owns containers only |
| Best for | Simple read-scaling, single planned promotion, staying on pg-instance tooling |
Zero-RTO availability requirements, unattended failover |
If you need a standby you control by hand, use pg replica. If you need the
cluster to survive a node crash without a human, use pg ha.
Configuration
The rendered patroni.yml per member is derived entirely from pg.yaml under
addons.patroni.<scope>. A typical entry:
Ports come from two independent pools — patroni_start_port (default 35532) for
PostgreSQL and patroni_restapi_start_port (default 39060) for the REST API —
so Patroni members never collide with plain-instance or addon ports.
The DCS-scope key is the scope plus the namespace suffix (Patroni has no namespace concept of its own, so the prefix is baked into the scope to keep two pgcli namespaces sharing one etcd from cross-talking).
Default patroni.yml
The file below is what pgcli renders and mounts read-only into each member’s
container. Passwords are auto-generated at bootstrap time; --passwords-file
lets you supply your own set.
Key points:
scope= Patroni cluster name. Combined withnamespace, it forms the etcd key prefix.etcd3.hosts=host:portonly (nohttp://scheme). Multiple endpoints are comma-separated.bootstrap.dcsvalues are defaults only — they take effect on first bootstrap; afterwards usepg ha edit-config.postgresql.listen: 0.0.0.0is set when--advertise-hostis used (cross-host). Without it, listen is127.0.0.1.- Passwords are generated once and stored in
pg.yaml; exported viapg ha passwords.
Dynamic Configuration
Patroni’s dynamic configuration lives in the DCS (etcd) under /service/<scope>/config.
Every member reads it on each loop (every loop_wait seconds) and applies
changes live — so pg ha edit-config is the right way to tune runtime parameters
without restarting containers.
bootstrap.dcsis one-time. Thebootstrap.dcsblock inpatroni.ymlonly takes effect when the first member of a scope runspg ha create(i.e. the cluster bootstrap). Once Patroni writes the config into the DCS, subsequent changes tobootstrap.dcsin the YAML file are completely ignored — even on re-install (pg ha create). To change dynamic configuration after bootstrap, usepg ha edit-config.The common pitfall: re-running
pg ha createdoes not re-readbootstrap.dcsfrom the YAML — Patroni sees the existingconfigkey in the DCS and uses it directly.
Method Description pg ha edit-config app -- -s key=valueRecommended — pgcli’s standard way pg ha ctl app -- edit-configPassthrough to patronictl, same effect Patroni REST API ( PATCH /config)Requires access to a member’s REST API port
Persistence: changes are written directly to etcd, not to container files.
Container restarts, pg ha start/stop, or even pg ha create (re-install)
do not lose these settings — new members automatically pick up the latest config
from the DCS.
Viewing the current config
This is a convenience alias for pg ha ctl app -- show-config. The output is
the full JSON blob stored in /service/<scope>/config.
Modifying parameters
After a change, all members apply it on their next loop (within loop_wait
seconds). No restart needed.
Common tunable parameters
| Parameter | Default | Description |
|---|---|---|
loop_wait |
10 |
Seconds between leader-loop iterations (lock renewal, DCS updates) |
ttl |
30 |
Leader lock TTL. If the leader fails to renew within this window, replicas trigger failover |
retry_timeout |
10 |
Timeout for DCS/PostgreSQL operations. Must be < ttl - loop_wait to give the leader at least one retry chance |
maximum_lag_on_failover |
1048576 |
Maximum replication lag (bytes) for a replica to be eligible for promotion. Default 1 MB |
synchronous_mode |
false |
Enable synchronous replication (zero data loss, higher latency) |
synchronous_node_count |
1 |
How many synchronous standby nodes (when synchronous_mode=true) |
use_pg_rewind |
true |
Use pg_rewind to rejoin a failed leader (faster than full pg_basebackup) |
use_slots |
true |
Use replication slots (prevent WAL loss when a replica disconnects) |
failover_timeout |
0 |
How long to wait before failover (0 = immediate when leader is lost) |
PostgreSQL runtime parameters can also be set under postgresql.parameters:
These trigger a PostgreSQL reload (or restart, depending on the parameter’s
context). Check pg_hba.conf and postgresql.conf parameter documentation for
which settings require a restart.
What not to edit
- Do not edit
patroni.ymlon disk — it is regenerated frompg.yamlon everypg ha create, and Patroni reads dynamic config from the DCS anyway. - Do not edit
postgresql.confdirectly — Patroni overwrites it each loop from the DCS config. - Do not change
scopeornamespace— these are baked into the DCS key at bootstrap time and cannot be changed without recreating the cluster.
Full parameter reference: Patroni Dynamic Configuration covers every DCS-tunable parameter with defaults, constraints, and examples.
Extensions
Full reference: Extensions in HA Clusters covers the orchestration order, cross-host workflow,
shared_preload_librariesordering, and builtin-only fast path.
Notes
- Linux only (root or rootless). Rootless members run as the host user via
--userns=keep-id; root members chown config and data dirs to postgres (uid 999) so the container can read0600config and write its data dir. pg_hba.confis permissive by design (host all all all scram-sha-256- a
replicationline, pluslocaltrust lines). Rootless podman’s pasta rewrites loopback sources, and Patroni’s ownreplace_pg_hbastep only ever grants its resolved loopback TCP address duringcustom bootstrap— sopostgresql.use_unix_socket/use_unix_socket_replare set to make Patroni connect to its own instance over the unix socket instead, which is unaffected by the rewrite and keeps bootstrap from deadlocking on its own pg_hba. Tightening the allow-list to a fixed set is a planned future refinement — do not expose these ports to untrusted networks yet.
- a
- The DCS (etcd) has its own security caveats — see the etcd page: pgcli-managed etcd runs without TLS or auth.
- No
init.sh/docker-entrypoint-initdb.d. Patroni bootstraps the cluster itself, so theadmin/default-db convention of plain instances does not exist here; usepostgres(superuser) to connect and create roles. use_slots/use_pg_rewindare enabled: Patroni owns replication slots and can rejoin a crashed leader viapg_rewindinstead of a full rebase.
2 - Patroni Dynamic Configuration
These parameters are stored in the DCS (etcd) under /service/<scope>/config and
applied to every member of the cluster. Modify them with pg ha edit-config:
Reference: Patroni official docs, Pigsty dynamic config.
bootstrap.dcsis one-time. Thebootstrap.dcsblock inpatroni.ymlonly takes effect when the first member of a scope bootstraps the cluster. Once Patroni writes the config into the DCS, subsequent changes tobootstrap.dcsin the YAML file are completely ignored — even on re-install (pg ha create). Re-runningpg ha createdoes not re-readbootstrap.dcs— Patroni sees the existingconfigkey in the DCS and uses it directly.
| Method | Description |
|---|---|
pg ha edit-config app -- -s key=value |
Recommended — pgcli’s standard way |
pg ha ctl app -- edit-config |
Passthrough to patronictl, same effect |
Patroni REST API (PATCH /config) |
Requires access to a member’s REST API port |
Core timing parameters
| Parameter | Default | Min | Description |
|---|---|---|---|
loop_wait |
10 |
1 | Seconds the main loop sleeps between iterations (lock renewal, DCS updates, state refresh) |
ttl |
30 |
20 | Leader lock TTL in seconds. Effectively the wait time before auto-failover triggers |
retry_timeout |
10 |
3 | DCS and PostgreSQL operation retry timeout. If the DCS or network is down for less than this, Patroni does not demote the leader |
Constraint when changing
loop_wait,retry_timeout, orttl:
Failover parameters
| Parameter | Default | Description |
|---|---|---|
maximum_lag_on_failover |
1048576 |
Maximum replication lag (bytes) for a replica to be eligible for leader promotion |
maximum_lag_on_syncnode |
-1 |
Maximum lag (bytes) before a synchronous standby is replaced by a healthy async one. When ≤ 0, Patroni does not replace unhealthy sync standbys. Set high enough to avoid frequent replacement during high-transaction workloads |
max_timelines_history |
0 |
Max timeline history entries kept in the DCS. 0 = keep all |
primary_start_timeout |
300 |
Seconds the leader has to recover from a failure before failover triggers. 0 = immediate failover on crash detection (may lose transactions with async replication). Max failover time = loop_wait + primary_start_timeout + loop_wait; with 0, just loop_wait |
primary_stop_timeout |
— | Seconds to wait when stopping PostgreSQL (only effective with synchronous_mode). If the stop exceeds this timeout, Patroni sends SIGKILL to the postmaster. ≤ 0 or unset = no effect |
failover_timeout |
0 |
Seconds to wait before triggering failover after the leader is lost. 0 = immediate |
Replication mode
| Parameter | Default | Description |
|---|---|---|
synchronous_mode |
false |
Enable synchronous replication (off / on / quorum). The leader manages synchronous_standby_names; only the last-known leader or a sync standby may run for leader. Guarantees zero data loss at the cost of write unavailability when durability cannot be assured |
synchronous_mode_strict |
false |
When no sync standby is available, refuse to disable sync replication — blocks all client writes to the leader |
synchronous_node_count |
1 |
Number of synchronous standby nodes. Dynamically adjusted as members join/leave. Clamped to the number of eligible nodes |
failsafe_mode |
false |
Enable DCS failsafe mode: when the DCS is unreachable, the leader keeps running instead of demoting itself |
PostgreSQL settings
| Parameter | Default | Description |
|---|---|---|
postgresql.use_pg_rewind |
false |
Use pg_rewind to rejoin a failed leader (faster than full pg_basebackup). Requires data page checksums (--data-checksums at initdb) or wal_log_hints=on |
postgresql.use_slots |
true |
Use replication slots (prevents WAL loss when a replica disconnects). Default on PostgreSQL 9.4+ |
postgresql.pg_hba |
— | Rules for generating pg_hba.conf. Ignored if the PostgreSQL hba_file parameter is set to a non-default value |
postgresql.pg_ident |
— | Rules for generating pg_ident.conf. Ignored if ident_file is non-default |
postgresql.parameters |
— | PostgreSQL GUCs as key-value pairs, e.g. {max_connections: 100, wal_level: "replica", wal_log_hints: "on"}. Many are required for replication to work |
postgresql.recovery_conf |
— | Additional recovery.conf entries for standby configuration (PG12+ handled transparently) |
Set PostgreSQL parameters via the postgresql.parameters key:
Standby cluster
If defined, the cluster bootstraps as a standby cluster that streams from a remote primary.
| Parameter | Description |
|---|---|
standby_cluster.host |
Remote primary address |
standby_cluster.port |
Remote primary port |
standby_cluster.primary_slot_name |
Slot name for replication (optional, defaults to member name) |
standby_cluster.create_replica_methods |
Ordered list of methods to bootstrap the standby leader from the remote primary |
standby_cluster.restore_command |
WAL restore command |
standby_cluster.archive_cleanup_command |
Archive cleanup command for the standby leader |
standby_cluster.recovery_min_apply_delay |
Delay before applying WAL records |
Replication slots
| Parameter | Default | Description |
|---|---|---|
member_slots_ttl |
30min |
How long a replica’s physical replication slot is retained after it shuts down. 0 = delete immediately when the member key expires from the DCS. Only effective on PostgreSQL 11+ |
slots |
— | Permanent replication slots (hash map). Preserved across switchover/failover. Physical slots on PG11+ are created on all nodes and advanced every loop_wait seconds. Logical slots are copied from primary to replicas via restart, then advanced every loop_wait seconds. Requires use_slots: true |
ignore_slots |
— | Slots managed externally that Patroni should not touch (list of attribute sets). Any subset match causes the slot to be ignored |
Permanent slots example
Node-pinned physical slots
For a fixed cluster topology, define a permanent physical slot per node to prevent slot deletion during temporary outages:
Warning: Permanent replication slots are synced from the primary/standby_leader to replicas only. Applications should use them on the leader node. Using permanent slots on replicas causes unbounded
pg_walgrowth across the cluster. Exception: physical slots matching a Patroni member name (created and maintained by Patroni) are synced across all nodes for inter-node replication.
Viewing the current configuration
Or read directly from etcd:
3 - Patroni REST API
Patroni exposes an HTTP REST API on each member, providing health check endpoints (for load balancers and Kubernetes probes), monitoring data (including Prometheus metrics), and cluster management operations.
Reference: Patroni official docs, Pigsty REST API.
Finding the REST API
Each member’s REST API port is assigned automatically by pgcli and stored in
pg.yaml under addons.patroni.<scope>.members.<name>.restapi_port. The port
is also embedded in the rendered patroni.yml as restapi.connect_address.
Authentication
pgcli enables basic-auth on the REST API (username postgres, auto-generated
password). But authentication is method-based, not blanket:
| Method | Auth required? | Endpoints |
|---|---|---|
GET / HEAD / OPTIONS |
No | All read endpoints — /, /primary, /replica, /health, /cluster, /config, /metrics, /patroni, … |
POST / PATCH / PUT / DELETE |
Yes | Write endpoints — /failover, /switchover, /config (modify), /reload, /restart, … |
This is by design: read-only health checks stay open so a load balancer or
Prometheus can poll them without credentials, while operations that change
cluster state (POST /failover, PATCH /config) require authentication. The
credentials exist so patronictl and other Patroni members can perform writes.
So a load balancer health check needs no auth:
Write operations do require auth. The password is the same restapi_password
used in patroni.yml’s restapi.authentication section. Export it with:
Health Check Endpoints
All health check endpoints respond with GET requests. Patroni returns a JSON
document describing the node state, along with an HTTP status code. Use HEAD
or OPTIONS instead of GET when only the status code is needed (no response
body).
Primary-only endpoints (200 only on leader)
These return HTTP 200 only when the node is the current leader holding
the leader lock:
| Endpoint | Description |
|---|---|
GET / |
Root — primary health check |
GET /primary |
Alias for / |
GET /read-write |
Alias for / |
GET /leader |
Like / but doesn’t distinguish primary vs standby_leader |
GET /master |
Legacy alias for /leader |
These endpoints are useful for load balancer health checks that should route writes only to the current primary.
Replica-only endpoints (200 only on replica)
| Endpoint | Description |
|---|---|
GET /replica |
Replica health check — 200 when node is running, role is replica, and noloadbalance tag is not set |
GET /replica?replication_state=streaming |
Only 200 when replica is actively streaming (not still catching up via archive recovery) |
GET /replica?lag=<max> |
Only 200 when replication lag is below the threshold (bytes or human-readable: 10MB, 1GB) |
Read-only endpoints (200 on both primary and replica)
| Endpoint | Description |
|---|---|
GET /read-only |
Any running node (primary or replica) |
GET /synchronous / GET /sync |
Sync standby only |
GET /asynchronous / GET /async |
Async standby only |
GET /read-only-sync |
Primary + sync standby |
GET /read-only-quorum |
Primary + quorum standby |
GET /quorum |
Quorum standby only |
PostgreSQL health
| Endpoint | Description |
|---|---|
GET /health |
200 when PostgreSQL is running (regardless of role) |
Kubernetes Probes
| Endpoint | Description |
|---|---|
GET /liveness |
200 if Patroni heartbeat loop is running. 503 if primary’s last heartbeat exceeds ttl seconds, or replica exceeds 2*ttl. Lightweight — no SQL queries. Suitable for livenessProbe. |
GET /readiness |
200 when node is leader, or when PostgreSQL is running, replicating, and within the allowed lag. Accepts ?lag=<max> (default: maximum_lag_on_failover) and ?mode=apply|write (default: apply). Suitable for readinessProbe. |
Monitoring Endpoints
GET /patroni
Returns detailed node status as JSON. Used internally by Patroni during leader election and also useful for monitoring:
Response fields:
| Field | Description |
|---|---|
state |
Node state: running, stopped, starting, etc. |
role |
primary or replica |
server_version |
PostgreSQL version as integer |
xlog.location |
Current WAL position (primary only) |
xlog.received_location |
WAL received from primary (replica only) |
xlog.replayed_location |
WAL replayed (replica only) |
timeline |
Current timeline number |
replication |
Array of connected replicas (primary only) |
cluster_unlocked |
true if no leader lock is held |
pause |
true if auto-failover is paused |
dcs_last_seen |
Epoch timestamp of last successful DCS contact |
patroni.version |
Patroni version |
patroni.scope |
Cluster scope name |
patroni.name |
Member name |
GET /cluster
Returns the full cluster topology — all members with their roles, states, and replication status:
GET /config
Returns the current dynamic configuration stored in the DCS:
This is equivalent to pg ha edit-config <scope> --show.
GET /history
Returns the timeline history. Empty array [] when no timeline switch has
occurred (fresh cluster).
GET /metrics
Returns monitoring data in Prometheus exposition format, suitable for scraping by Prometheus or compatible systems:
Key metrics:
| Metric | Description |
|---|---|
patroni_primary |
1 if this node is the leader |
patroni_replica |
1 if this node is a replica |
patroni_postgres_running |
1 if PostgreSQL is running |
patroni_postgres_streaming |
1 if PostgreSQL is streaming (replica) |
patroni_xlog_location |
Current WAL position (primary only) |
patroni_xlog_received_location |
WAL received (replica only) |
patroni_xlog_replayed_location |
WAL replayed (replica only) |
patroni_cluster_unlocked |
1 if no leader lock |
patroni_is_paused |
1 if auto-failover is disabled |
patroni_pending_restart |
1 if node needs restart |
patroni_postgres_timeline |
Current timeline |
patroni_dcs_last_seen |
Epoch of last DCS contact |
patroni_server_version |
PostgreSQL version |
Tag-based Filtering
Health check endpoints accept query parameters to filter by custom tags defined
in the member’s patroni.yml tags section:
Practical Patterns
Load balancer routing
Because GET health checks need no auth, a load balancer can poll the REST API
endpoints directly. The pattern (from Patroni’s example haproxy.cfg) is: the backend TCP port
is the PostgreSQL port, but the health check hits the REST API port with
GET /, which returns 200 only on the leader.
Single read-write endpoint (all traffic → current leader):
check port 8008/8009— the HTTP health check goes to the REST API port, not the PG port.GET /returns 200 only on the leader → HAProxy marks only the leader asUP.on-marked-down shutdown-sessions— on failover, sessions to the old leader are killed so clients reconnect to the new leader.fall 3/rise 2withinter 3s— ~9s to mark down, ~6s to mark up.
Read/write split — add a second listener that uses GET /replica (200 only on replicas) for read traffic:
app_rw(:5000) —GET /→ only the leader isUP→ writes land on the leader.app_ro(:5001) —GET /replica→ only replicas areUP→ reads spread across replicas.
The same GET /replica?lag=1MB refinement (from the replica endpoints above)
keeps a lagging replica out of the read pool.
Monitoring with curl
Quick health check script:
Prometheus scrape config
GET /metrics needs no auth, so the scrape config works without credentials:
4 - Extensions in HA Clusters
Installing PostgreSQL extensions in a Patroni-managed HA cluster is
fundamentally different from a single-node instance. This page explains why,
and walks through the pg ha extension command subtree that handles it.
Why not pg extension install?
The single-node flow (pg extension install <instance> <ext>) works by:
- Building a
-extderived image with Pigsty packages - Stopping and recreating the container from the new image
- Editing
postgresql.confto setshared_preload_libraries - Running
CREATE EXTENSIONinside the container
In a Patroni cluster, steps 2-4 all break:
- Patroni is PID 1. Recreating a container takes that node offline. If it’s the leader, Patroni triggers a failover — the cluster reshuffles while you’re mid-install.
- Patroni regenerates
postgresql.confevery loop from the DCS. Any direct edit to the file is overwritten within seconds.shared_preload_librariesmust be set viapatronictl edit-config(which writes to the DCS). CREATE EXTENSIONmust run on the leader, which may be on a remote host — not inside any local container.
pg ha extension orchestrates all of this correctly.
How It Works
The install flow follows a specific order to avoid the pitfalls above:
The critical ordering is resume before edit-config: a paused Patroni cluster
does not apply edit-config changes. If you edit-config while paused, the
shared_preload_libraries update is silently lost.
Builtin-only fast path
If all requested extensions are builtin (contrib extensions like hstore,
uuid-ossp that ship with PostgreSQL), steps 2-4 are skipped entirely — no
image build, no container recreate, no pause/resume. The flow goes straight to
edit-config + CREATE EXTENSION.
Commands
Install
Builds the -ext image, pauses the cluster, and recreates local member containers.
- Single-host cluster (all members local): automatically runs apply (resume + edit-config + CREATE EXTENSION)
- Cross-host cluster: only performs pause + recreate, cluster stays paused. Run install on each host separately, then run
pg ha extension applyon any host to complete the installation.
Extensions are passed as a comma-separated list:
Flags:
| Flag | Default | Description |
|---|---|---|
--database |
postgres |
Target database for CREATE EXTENSION |
--auto-restart |
false |
Skip confirmation for the rolling restart triggered by edit-config |
⚠
--databaseaffects all extensions. When you specify--database, theapplyphase runsCREATE EXTENSION IF NOT EXISTSfor all installed extensions (not just the newly added ones) in the target database. For example, if the cluster has[pg_cron, pg_stat_statements, pgvector]installed, runningpg ha extension install app hstore --database mydbwill create all four extensions inmydb. The extension.sofiles are already in the image and won’t be reinstalled — only the SQL objects (functions, types, etc.) are registered in the target database.If you only want to enable an existing extension in a specific database, there’s no need to reinstall. Simply connect to that database with psql and run
CREATE EXTENSION IF NOT EXISTSdirectly.
Remove
Runs DROP EXTENSION on the leader, then updates shared_preload_libraries
via patronictl edit-config (triggers a rolling restart). Does not rebuild
the image or recreate containers — the -ext image only grows; disk reclamation
is rare and manual.
Extensions are passed as a comma-separated list:
List
Shows three views of the cluster’s extensions:
- Config (pg.yaml): the
extensionslist stored inpg.yaml - DCS (preload): the
shared_preload_librariesvalue frompatronictl show-config - Leader (installed): actual extensions from
pg_extensionon the leader
Apply
Manually triggers the second half of the install flow: resume the cluster,
run patronictl edit-config, and execute CREATE EXTENSION on the leader.
Used in cross-host clusters — after running install on each host (each
builds the image and recreates its local members), run apply once on any host
to complete the DCS update and extension creation.
⚠
--databaseaffects all installed extensions.applyrunsCREATE EXTENSION IF NOT EXISTSfor all extensions in the cluster in the specified database, not just the newly added ones.
Cross-Host Workflow
In a cross-host cluster, each host manages only its own members. The extension installation workflow splits into two phases with global pause/resume — the cluster is paused once and resumed once, not per-host:
Phase 1 — per host (sequentially): Run pg ha extension install on each host.
Each host:
- Builds the
-extimage locally (merging DCS’s existing preload list to include all packages) - Pauses the cluster (idempotent — second host won’t fail on “already paused”)
- Recreates its own local members from the new image (replicas first, leader last)
- Saves config — cluster stays paused, no resume
Recommended order: start from replica hosts. If the host with the leader runs
installfirst, recreating the leader triggers a failover (leader moves to another host). Starting from replica hosts keeps the leader in place until the very end, minimizing failover-related data sync overhead.
Phase 2 — once, on any host: Run pg ha extension apply to:
- Resume the cluster (idempotent — tolerates “not paused”)
- Update
shared_preload_librariesviapatronictl edit-config - Wait for the rolling restart
- Run
CREATE EXTENSIONon the leader
Note: After
install, the cluster is in paused state (auto-failover disabled). If you forget to runapply, runpg ha extension applyor manuallypatronictl resume <scope>to re-enable failover.
For single-host clusters (all members local), install automatically runs
the apply step — no separate command needed.
About leader drift: Leader drift cannot be completely avoided at this time. When the container on the host where the leader resides is recreated, the Patroni process stops. After the leader key in the DCS expires (TTL), even if the cluster is paused, other replicas will still detect the expired key and trigger a new election. The
replicas-firstrecreation order ensures that the host with the leader is the last one to runinstall, but it cannot prevent the leader drift itself. This is an inherent limitation of Patroni + container recreation.
The extensions Config Field
Extensions are tracked at the cluster level in pg.yaml, under the Patroni
cluster config:
This list is the source of truth for what pg ha extension list reports and
what apply installs. Both install and remove update it automatically.
shared_preload_libraries Ordering
Some extensions must appear at position 0 in shared_preload_libraries
(PostgreSQL fatals if they’re not first). pgcli tracks this as a catalog
attribute (PreloadFirst) — currently set on:
citus— distributed PostgreSQL, must be loaded before anything elsetimescaledb— time-series engine, same constraint
Additional extensions may require companion DCS parameters:
pg_cronneedscron.database_name
pg ha extension handles all of this automatically:
- Extensions with
PreloadFirstare always placed at the start of the CSV, regardless of input order - When
pg_cronis present,cron.database_nameis set to the--databasevalue (defaultpostgres) - When
pg_cronis removed,cron.database_nameis cleared
Examples
Install pg_stat_statements and pg_cron
This builds an -ext image, recreates all members (with pause/resume), sets
shared_preload_libraries=pg_stat_statements,pg_cron and
cron.database_name=mydb in the DCS, then creates both extensions in mydb.
Install Citus (must be first in preload)
Even though citus is listed second, the preload CSV is generated as
citus,pg_stat_statements — extensions with PreloadFirst are always
placed at position 0.
Remove an extension
Drops pg_cron from the leader, removes it from shared_preload_libraries,
and clears cron.database_name. Triggers a rolling restart.
Check what’s installed
Limitations
- Image rebuild is additive. Removing an extension does not rebuild the
-extimage or shrink it. To reclaim disk, manually prune old images withpodman image prune. - No per-member extension list. Extensions are cluster-wide — all members
share the same
-extimage and the sameshared_preload_libraries. CREATE EXTENSIONtargets one database. PostgreSQL extensions are per-database. To install in multiple databases, re-run with--databasepointing at each one.- Cross-host requires manual coordination. Each host must run
installbeforeapplyis run once. pgcli does not SSH into remote hosts.
See Also
- Patroni HA — cluster setup, commands, and architecture
- Extensions — single-node extension management and the Pigsty catalog
- Patroni Dynamic Configuration — DCS parameters reference
5 - Patroni Cluster Backup
This page walks the complete procedure for backing up a Patroni HA cluster:
getting the backup infrastructure up with pg backup setup, taking
full/incr/diff backups with pg ha snapshot, and one operational pitfall — when a
replica’s Replay LSN stalls and pg ha status shows members disagreeing on LSN,
how to pull it back with patronictl reinit. Command and mechanism details live in
Patroni HA and Backup;
this page strings them into a path you can run end to end.
Model: one stanza per cluster, the backup container finds the primary
A Patroni cluster (scope) maps to one pgBackRest stanza: pgcli_<scope>
(matching the namespace-qualified DCS scope, so two pgcli namespaces sharing one S3
bucket never collide). The stanza is cluster-wide — the backup-side config lists
every member as a pg*-host, and pgBackRest SSH-probes them, locates the current
primary itself, and keeps following it across failovers with no config change.
Therefore:
- There is no per-member backup.
full/incr/diffall operate on the one cluster stanza; pgBackRest snapshots whatever host is the leader at that moment. - Backups run from the shared backup container, reaching the primary over SSH.
Cross-host members are reachable too (trust is wired up automatically by
setup). pg ha snapshotcommands never need you to name the leader — pass the scope; pgBackRest resolves the primary.
Prerequisite: pg backup setup
One idempotent run brings up the whole backup path and gives the cluster archiving:
What setup does (when an S3 repo is configured, it first runs an endpoint preflight — a plain TCP dial that fails the run immediately on a mistyped host or a store that is down, rather than after the image pull / container start buries the real cause):
- Build/pull the pgbackrest image, create the network and dirs, generate the
backup-container
pgbackrest.confand the member-local archive viewpgbackrest-archive.conf. - Start the shared backup container, then verify repository connectivity: a
pgbackrest repo-lsproves the whole stack (TLS, credentials, bucket) and maps a failure to an actionable hint — a cert error points atpg backup fetch-ca, access-denied at the credentials, a refused/timeout connection at the endpoint. - Wire up cross-host backup trust: each host publishes its backup public key
into the cluster’s etcd registry at
pg ha create;setupmerges every member’s key into one per-clusterauthorized_keysbind-mounted into each member container. sshd re-reads that file on every login and the merge rewrites it in place, so hosts that join later are trusted without a restart. The S3 repo CA follows the same path: a host withca_filepublishes it into the registry, joiners with it empty pull it. (Details: HA → Backing up a cross-host leader.) - Enable WAL archiving: when a member’s archive config is stale, pause the
cluster, recreate stale members replicas-first / leader last (a recreate is exactly
the postmaster restart
archive_modeneeds), resume, and render the samearchive_commandinto every member’spatroni.yml. - Run
stanza-create+checkfor every cluster stanza.
HTTPS is mandatory. pgBackRest rejects plaintext S3. A MinIO meant to receive cluster archives must serve TLS — install it with
pg addon install minio --tls(pgcli’s self-signed CA; point--s3-ca-fileat itsca.crt). When the MinIO is on another machine, pull its CA with one TLS handshake —pg backup fetch-ca <store-host>:<port>— no scp needed (see MinIO → Getting the CA onto a remote host).
Confirm after setup:
Once the backup container is Up and the stanzas are ready, start backing up.
Create members after setup, on every host. A member’s archiving is fixed at creation time: the container only mounts
pgbackrest-archive.confif that file already exists, and the renderedpatroni.ymlonly carries the archive GUCs if an S3 repo is configured. A memberpg ha created beforepg backup setupon its host therefore starts without WAL archiving —pg ha snapshot/pg ha restorecannot cover it until a latersetupflags it stale and recreates it inside a pause window.pg ha createwarns about this up front; runpg backup setupfirst on each host for a backup-ready create.
Snapshot operations: pg ha snapshot
These act on a cluster (scope), running pgBackRest against the current primary from
the shared backup container. This is the Patroni-cluster counterpart of pg snapshot
(single instance).
Notes:
- Before each command,
pg ha snapshotre-renders the backup-containerpgbackrest.conffrom the current topology, so member adds/removes and port changes are reflected in the stanza immediately — the file is bind-mounted, no recreate. - You can only delete a non-unique full: pgBackRest keeps at least one full, so deleting the only one is refused (take a new full first).
- create connects to whatever host is leader at the time; after a failover it still succeeds — that is the point of the cluster-wide stanza.
Pitfall: a stuck replica LSN → patronictl reinit
Symptom. In pg ha status app one replica’s Replay LSN trails its own
Receive LSN (or visibly differs from another replica) and never converges over
time. Note Patroni’s Lag column is relative to the leader, so every number moves
when the leader’s LSN advances; to spot a real stall compare a replica’s own recv vs
replay, or watch whether its Replay LSN stays frozen across checks.
Confirm. Three signals point to the same conclusion — the replica’s WAL replay is at a hole and streaming has actually stopped:
record with incorrect prev-link + a repeating waiting for WAL to become available
means the WAL stream has a hole/discontinuity at a segment boundary and the startup
process is dead-waiting for the next segment. Patroni often keeps logging
“no action. I am … a secondary, and following a leader” each cycle — it believes it
is following while the postmaster’s walreceiver has actually stopped, so it will not
self-heal. This is a replication runtime issue, unrelated to backups.
Fix: reinit the replica. Have Patroni take a fresh pg_basebackup from the leader
and restart replication. This destructively rebuilds that replica’s data directory
(the leader and other replicas are untouched), so Patroni requires --force:
-f does not work: patronictl reinit only accepts the full --force (unlike
switchover/failover’s --yes). Success: reinitialize for member node1 means it
was dispatched.
Verify recovery. reinit wipes and rebuilds the data dir, then basebackups + replays.
Poll until it is recv == replay and status=streaming:
After reinit completes, run pg ha snapshot create app --type full once to realign the
snapshot baseline with the healthy topology.
6 - Patroni Cluster Restore
This page is the restore-side companion to Patroni Cluster Backup.
Once a cluster has a stanza with WAL archiving (pg backup setup + pg ha snapshot),
pg ha restore brings the whole cluster back to a point in time — the Patroni
counterpart of the single-instance pg restore. It walks the
mechanism (why a cluster PITR is not just pgbackrest restore), the safety
pre-check on where the leader lives, an end-to-end runbook, and the re-baseline
steps that must follow. Command/flag details also live in
Restore → Patroni cluster; this page strings them
into a procedure you can run.
Prerequisite: a stanza with WAL archiving, on this host
PITR can only replay up to what was archived. Before restoring, the cluster must already have:
- a cluster stanza provisioned against an S3 repo —
pg backup setup --s3-endpoint ...(see Backup → S3 repository and HA backup); - at least one full snapshot, plus the WAL segments covering the target time
(taken with
pg ha snapshot create <scope> --type full).
The target time is bounded by that history: earlier than the newest full backup’s stop time, or later than the last archived WAL, and the recovery cannot land there.
pg backup setup must have run on this host, and the backup container must be
up. This is stricter than it sounds for a cross-host cluster:
- The member that carries the bootstrap restores from S3 using the repo config
mounted into its container —
pgbackrest-archive.confat/etc/pgbackrest.confplus the S3 CA at/etc/pgbackrest/ca.crt. Both are generated bypg backup setup, and the container only mounts them when the files exist. A host that never ran setup has no repo config to mount, so thepgbackrest restorebootstrap cannot reach S3 and fails to start. - The leader-locality and stop-time pre-checks run
pgbackrest ... infoinside the backup container.pg ha restorerefreshes the sharedpgbackrest.confbest-effort but does not create the member archive config and does not start the backup container — so bring it up first (pg backup setup, orpg backup statusto confirm it isUp). If the backup container is down, the pre-checks are skipped and the real failure surfaces only during the bootstrap; the post-restore fresh snapshot needs it running anyway.
In short: run pg ha restore on a host where pg backup setup has provisioned the
S3 repo and the backup container is Up.
Mechanism: custom-bootstrap PITR, not a bare pgbackrest restore
A Patroni cluster cannot be PITR’d by running pgbackrest restore under a live
postmaster — Patroni owns the data directory, the timeline, and the DCS identity.
pg ha restore instead follows Patroni’s custom-bootstrap recipe:
-
Pause the cluster (
patronictl pause). This freezes failover before anything is stopped: stopping the local leader while the cluster is live lets Patroni promote a remote replica this host cannot stop, and that new remote leader then races the local bootstrap to re-claim the DCS on the old timeline — defeating the leader-locality pre-check from inside the destructive path. A restore re-run against a cluster a previous attempt already tore down (DCS cleared, members stopped) has nothing live to pause; that is the already-frozen state, so the pause failure on an empty DCS is tolerated and the run continues. -
Clear the cluster’s DCS identity (
patronictl remove). Patroni’sbootstrap.dcsblock only runs when the DCS has no config key, so removing it is what re-arms the bootstrap path. -
Pick one LOCAL member and rewrite its
patroni.ymlso the default initdb bootstrap is replaced by apgbackrest restoremethod targeting the point in time:no_params: truestops Patroni appending--scope/--datadir(whichpgbackrestrejects with[031] invalid option);keep_existing_recovery_conf: truepreserves therecovery.signal+restore_command+recovery_target_*thatpgbackrest restorewrites itself. Themethod/pgbackrestkeys are siblings ofbootstrap.dcs, never nested inside it — recovery GUCs leaked into the DCS config would apply as live settings on the promoted leader. -
Wipe that member’s PGDATA and recreate its container. Patroni sees an empty data dir with no DCS config, runs the method, recovers to the target on a new timeline, and — via
--target-action=promote— promotes itself to the writable leader. -
The remaining members rejoin the new leader: this host’s other local members are wiped and restarted as standard replicas (they
pg_basebackupthe new leader); cross-host members rejoin through the DCS (or are reinitialized).
The leader-locality pre-check
pg ha restore refuses to run unless the current leader is a local member of this
host. Before touching anything it reads the leader from the DCS and checks it is
one of the members this host can control.
Why. Step 1 clears the DCS and step 3 re-bootstraps a local member on a new timeline. A leader that lives on another host keeps its Patroni daemon running — this host cannot stop it — so after the DCS key is removed that remote leader races to re-claim leadership on the old timeline, colliding with the local bootstrap. Requiring leadership to be local (so it can be stopped) removes the race.
What you see.
Fix. Move leadership onto a member of the host you are standing on, then retry:
Or run pg ha restore on the leader’s own host instead. An empty / unreadable
leader (e.g. the very first bootstrap, before any member has published itself) does
not block the run.
Differences from single-instance pg restore
- It always promotes. There is no read-only “hold, inspect, retry another time”
two-step — the bootstrap recovery ends on a writable leader on a new timeline.
Confirm the target with
--dry-runfirst. (Thepatronictl pausein step 1 of the mechanism is unrelated: it freezes failover during the restore, it is not a read-only inspection window and the cluster is resumed implicitly when the new bootstrap republishes the DCS config.) - There is no
--promoteflag — promotion is automatic (it is baked into the bootstrap command). - It must run on a host that owns a member, and the leader must be local (see the pre-check above).
Runbook: restoring a cluster to a point in time
The full flag set is --time (required, all the formats in Restore),
--member, --dry-run, --tail-logs, --force.
1. Confirm the target is safe — dry-run. Nothing is touched; you see the stanza,
the chosen bootstrap member, the exact pgbackrest command, and (if applicable) the
leader-locality warning.
2. Make sure the leader is local (only if the dry-run warned). Switchover, then re-check status.
3. Restore. The default prompts for confirmation, listing scope, stanza, target,
bootstrap member, and the “PERMANENTLY LOST” warning. Stream the recovery logs with
--tail-logs; skip the prompt with --force for automation.
4. Verify the new cluster comes up. waitForLeader polls the DCS until a leader
appears (15-minute ceiling), then rejoin members start as replicas.
After the restore: re-baseline the new timeline
The restore is not the last step. The data directory was rebuilt, so two follow-ups are required before the cluster is backup-ready again:
-
Take a fresh full snapshot so future PITR has a base on the new timeline:
This usually succeeds directly: a pgBackRest restore of the same stanza preserves the PostgreSQL system-id across the timeline switch (verified), so the post-restore snapshot does not hit
[051] system-id ... do not match stanza. You do not needpg backup stanza-upgradeafter a cluster restore. (A real system-id change — a fresh initdb into the same stanza, e.g. a brand-new cluster reusing an old stanza name — is what[051]and stanza-upgrade are for.) -
Bring back any member that did not auto-rejoin — typically a cross-host member that lost the old leader. Reinitialize it from the new leader (destructive to that replica’s data dir only):
(Full details in HA backup → stuck replica reinit.)
Troubleshooting
- “no local member on this host” — you ran it on a host that owns none of the
scope’s members.
pg ha restorecan only rebuild a local data dir + container. Run it on a host that owns a member, or on the leader’s host. - “target time is before the latest backup stop time” — the earliest usable point
is the newest full backup’s stop time; the error prints it and a suggested
--time. Take a fresher full snapshot if you need a later base. - Recovery exceeds the last archived WAL — you cannot restore past what was
archived. Reduce the target time, or ensure
archive_commandis flowing (pg backup status) before retrying. - First post-restore snapshot fails
[051]— not expected: a same-stanza restore keeps the system-id (see the re-baseline section). If you do hit it, apg backup stanza-upgradefixes it non-destructively. FATAL: recovery ended before configured recovery target was reachedon a repeat PITR into the same repo — the stanza already hosts a previously promoted timeline, and the defaultrecovery_target_timeline=latestjumps onto that branch, which has no commits before your target. Restore again with an explicit older timeline pinned on the command, e.g.--recovery-option=recovery_target_timeline=<old-tl>(added to the bootstrappgbackrest ... restorecommand in the member’spatroni.yml). A first PITR after the target — no branch crossing — never hits this.
7 - HA Exec / psql
pg ha exec and pg ha psql are the direct way to run SQL against a Patroni
cluster. They are the cluster-side twins of the single-instance
pg exec and pg psql: you name
a scope, they resolve the target and connect. No --dsn to hand-assemble, no
member container to podman exec into, and — because the connection is plain TCP
— a leader or replica on another host is just as reachable as a local one.
Why not pg exec --dsn or podman exec
Before these commands, ad-hoc SQL on a cluster meant one of two awkward paths:
pg exec --dsn postgres://postgres:<pass>@<leader>:<port>/db— correct, but you assemble it by hand: read the leader’sconnect_addressout ofpg ha status, copy the superuser password out of the config, plug in the port. A fixed dsn also goes stale the moment the leader moves.podman exec -it <member> psql ...— reaches only a local member, so a leader on another host is out of reach; and you are inside a container the tool is meant to abstract away.
pg ha exec/psql fold both away: pgcli resolves the leader from the cluster’s
own state, authenticates with the stored password, and runs psql in a short-lived
container. You never see it.
How the target is resolved
The DCS roster (patronictl list -f json) is the only membership view that
spans the whole cluster. Each host’s pg.yaml registers only its own members,
so this host’s config cannot even see a node running elsewhere — but the DCS
knows every member and its advertised connect_address. Both commands go through
it:
- Default (no
--member) → the current leader. If leadership has moved since you last looked, you still hit the right node — there is no stale dsn. --member <name>→ that named member, wherever it lives. Aim it at a replica for a read-only query (SELECToff a standby,pg_is_in_recovery()returnst), or at a specific remote node.
Authentication is the cluster’s superuser over scram; pg_hba
(host all all all scram-sha-256) accepts any source, so a remote member is
reachable across hosts exactly the way replication traffic already is.
Usage
pg ha exec takes the SQL as trailing words (so you normally quote one string);
pg ha psql opens an interactive shell and forwards anything after -- straight
to psql. --database selects the target database on both (default postgres).
There is no --user — the managed superuser is the only role pgcli holds
credentials for, so that is what connects.
Interactive vs. scripted
pg ha psql allocates a TTY when your stdin is one; when it is not (piping a
script, running under CI) it turns the pager off so the session terminates
instead of blocking in less. That makes both of these behave as expected:
Reading the output
pg ha exec streams psql’s own formatting — column headers, alignment, row
counts, and errors go straight to your terminal, exactly as psql -c would print
them. It is not a machine-parseable dump; if you need one, pipe through your own
tooling or use pg ha exec app "..." --csv-style psql flags via pg ha psql app -- -c "...".
Where these fit
For a stable client endpoint that survives failover independently of pgcli,
put HAProxy in front of the cluster and connect through it.
pg ha exec/psql are for operators and scripts driving a cluster directly,
not a substitute for a pooler in front of an application.
The pg ha status / pg ha ctl verbs remain the way to inspect cluster state and
reach unwrapped patronictl commands; these two are specifically about running
SQL.
8 - Example: HA Cluster with a Self-CA MinIO
This page is a worked example, run end to end and verified: create a Patroni HA cluster, stand up a MinIO addon that serves a bring-your-own domain certificate (a self-signed CA of our own making — the same shape as a public-CA or corporate-CA cert, just without buying one), and point the cluster’s pgBackRest backups and WAL archiving at it. Every command below is a real one from the run; only hostnames, ports, passwords, and the certificate material are genericized.
Two ideas make the example worth reading rather than skimming:
- BYO TLS on the store. The MinIO addon does not have to serve pgcli’s
generated self-signed pair —
--tls-cert/--tls-keylet it serve a cert you already hold. A client that trusts the issuing CA needs no extra material; with a private (here: self-signed) CA you hand that CA to the backup stack via--s3-ca-file, exactly as with the generated one. See MinIO → Bring your own certificate. - Isolation via a second config file.
backup.repo.s3is a single, global repository shared by every stanza in one pgcli environment. Repointing it at a new store would silently redirect the backups of every existing cluster on that host too. The clean way to try something new against a live estate is a separate config file with its ownbase_dir,namespace, and port ranges — which is what this example does.
The environment we built
| Piece | Value | Notes |
|---|---|---|
| Config | ~/.pgcli-app1/pg.yaml |
separate from the production ~/.pgcli/pg.yaml |
| Base dir | /home/fish/bucket/pgcli-data-app1 |
nothing touches the production data dir |
| Namespace | app1 |
prefixes every container name — pgcli-minio-app1-store1, pgcli-patroni-app1-app1-nodea, pgcli-backup-app1 |
| DCS | existing etcd m1 at 127.0.0.1:2379 |
created on the production config long before this example — see Prerequisite; reused here as an external endpoint, and cluster membership is keyed by scope, so a new scope is naturally isolated |
| MinIO | store1, ports 9010/9011, listen 0.0.0.0 |
BYO self-signed cert, SANs minio1.test,127.0.0.1,<host-ip> |
| Patroni | scope app1, member nodea, PG <host-ip>:35632 |
single member; the leader by definition |
| Backup container | pgcli-backup-app1 |
coexists with the production pgcli-backup-default |
Prerequisite — the etcd DCS
The example reuses the etcd that already serves this host’s production clusters. If you are starting from scratch, that first piece to install is the same addon, one member at a time — a lone member is a healthy one-node etcd cluster, plenty for HA dev/test:
(For members other hosts will join as a DCS, give it a reachable address at
install time — --advertise-host <host-ip> — so its peer/client URLs announce
that IP instead of loopback; and grow it to a real 3-node ensemble with
--name m2/--name m3 the same way. None of that matters for this example:
the cluster lives on the same host and dials 127.0.0.1:2379.)
Step 0 — a domain certificate
Any PEM leaf + key works. For the demo we minted a self-signed one with pgcli’s
own pg cert (validity via --valid-duration, SANs covering the name and IPs
clients will dial):
pg cert writes a single self-signed leaf (CA:FALSE, server-auth) that is its
own trust anchor — exactly the shape a public- or corporate-CA cert has, minus
having bought one. With a public-CA (or corporate-CA) cert this step is
simply “you already have the files” — point the flags at your fullchain.pem
(leaf first, then intermediates) and key, and the client-side story gets
easier: chains rooted in a trusted CA need no --s3-ca-file at all.
Step 1 — an isolated environment
--namespace makes every container name this config creates distinct, so two
environments live side by side in one podman. The config file is passed with
-c on every subsequent command — from here on, every pg in this page means
pg -c ~/.pgcli-app1/pg.yaml.
Port ranges. The default port pools (etcd 2379, MinIO 9000, Patroni 35532, …) start at the same values as the production config, and podman publishes the loopback ports of a host-networked container — so pass explicit, disjoint ports everywhere rather than letting both environments auto-assign the same numbers.
Step 2 — MinIO with the BYO cert
Install-time validation pairs the cert and key (a mismatched, expired,
CA-instead-of-leaf, or non-server-auth pair fails right here), prints the
validity window, and — because the cert is self-signed — tells you clients will
pin it as their trust anchor. --listen 0.0.0.0 rather than loopback: the
backup container reaches the store over host networking via the host IP, and
that IP is one of the SANs. A quick proof it serves your certificate:
Create the repository bucket (pgBackRest does not create it):
Step 3 — the Patroni cluster
The warning is the designed order of operations, not a problem: at create time
no repo exists yet, so the member comes up without archiving; the backup setup in the next step configures the repo and recreates the member with
archiving enabled. Note the DCS scope is app1-app1 (namespace prefix + scope)
— pass the plain scope (app1) to every pg ha command and pgcli resolves
the qualified form; it also keeps the stanza distinct from any app1 another
namespace might run on the same bucket.
Seed something worth backing up (this also proves client auth over scram from the host):
Step 4 — point the backup stack at the BYO store
--s3-ca-file takes our own cert as the trust anchor — for a self-signed
cert the leaf embeds its own root, so the file you passed to --tls-cert is
the CA file. pgBackRest passes the bundle to its verifier with no content
inspection, so a public-CA cert needs no flag at all and a private-CA one just
gets its chain here. The run then proves the full stack and rewires the member:
The “repo CA published to the cluster registry” line is the etcd-distribution path: every other member of this scope (or future ones, like a second host) picks the CA up from the DCS with no manual scp and no extra flags.
Step 5 — back it up, and prove it landed
The independent proof is that the objects are physically in the bucket, written over a TLS session terminated by our certificate:
archive/ is continuous WAL archiving (the archive-push the recreated
member runs), backup/ is the full snapshot. Confirm the archiver is healthy:
One subtlety from our run, worth knowing so you don’t chase a ghost:
failed_count was 1, for 00000002.history — the timeline-history file,
archived one second later successfully. It raced the brief window in which
backup setup was pausing and recreating the member. A lone recent failure
during a setup/recreate is self-healing; a growing count with the same WAL
name is a real repo problem.
What the example demonstrates
- A MinIO addon serving a bring-your-own certificate works as a pgBackRest
repository end to end: stanza-create,
check, full backup, and continuous WAL archiving all flow over TLS verified against the supplied cert. - Self-signed is not a lesser path: the same
--s3-ca-fileslot carries it (leaf==root), a public CA’s chain, or an internal CA bundle — the transport and the trust plumbing are identical. backup.repo.s3is one global repo per config file. To experiment beside a running estate, don’t repoint it — init a second config (pg config init --namespace … --base-dir …) and run the whole experiment there. Namespace-qualified container names, DCS scopes, and stanzas keep the two environments from ever seeing each other.
Teardown
Everything the example created lives under the app1 config, so teardown is
ordered through that same -c (data directories deleted explicitly):
--scope-all (not --member) is what clears the scope’s DCS keys in the
shared etcd — patronictl remove plus the pgcli member registry. Removing
members one by one leaves both behind.
The etcd m1 belonged to the production environment — it is untouched by all
of the above, which is exactly the point of the isolation.
Related
- Patroni Cluster Backup — the backup machinery in depth
- MinIO — the addon, and its BYO certificate section
- Namespace Isolation — how
--namespacescopes names - Patroni HA — the
pg hacommand set
9 - Generating a Certificate — pg cert
pg cert mints a self-signed certificate — its SubjectAltNames covering any
mix of DNS names and IP addresses you ask for — for any place you need TLS
without going through a CA: a dev or test server, an internal endpoint, a
service whose clients you control. The certificate it writes is a normal TLS
server leaf, usable anywhere.
The use this docs tree exercises is pgBackRest’s S3 path: a Patroni cluster’s
backups push to an S3 store (MinIO or silo, typically) that must serve
TLS — pgBackRest refuses plaintext S3 — and pgcli’s MinIO/silo addons can
serve a bring-your-own certificate (pg addon install minio --tls-cert ... --tls-key ...). The quickest way to get a cert to try that with is pg cert:
Nothing is registered with pgcli — it touches no pg.yaml and starts no
container; it just writes the two files you name and prints the SANs. See
MinIO → Bring your own certificate
for the serving side, and the worked
HA + self-CA MinIO example for it end to end.
What it produces
A single self-signed leaf certificate, not a chain:
- one
CERTIFICATEblock in--cert-file— the cert signs itself; CA:FALSE, extended key usageserverAuth— a proper TLS server leaf;- SANs split automatically: any
--hostentry that parses as an IP lands in the IP SANs, the rest in the DNS SANs, so"minio.test,10.0.0.9"needs no special syntax (and*.wild.testworks as a wildcard DNS entry).
Because it is self-signed, the leaf is its own trust anchor — hand the same
.crt to any TLS client that lets you point at a CA file or bundle (curl
--cacert, a browser’s import, an app’s SSL_CERT_FILE …). For pgBackRest’s
S3 path specifically that means backup.repo.s3.ca_file /
pg backup setup --s3-ca-file minio.crt. A remote host that never saw the
file needs no scp either: pg backup fetch-ca <endpoint> recognizes a
self-signed leaf in the TLS handshake and saves it back as the anchor. This is
not a workaround pgBackRest merely tolerates: OpenSSL treats whatever you hand
its trust store as an anchor regardless of CA:TRUE, and pgBackRest’s S3 path
(curl over OpenSSL) is that mechanism — verified directly,
openssl verify -CAfile minio.crt minio.crt returns OK.
With a real public- or private-CA certificate none of the self-anchor
machinery applies: a chain rooted in a trusted CA needs no ca_file at all,
and you already hold the issuing CA for a private one.
Flags
| Flag | Default | Meaning |
|---|---|---|
--host |
127.0.0.1 |
Comma-separated DNS names and/or IPs for the SAN (repeatable). IPs are auto-detected; *.wild.test works as a wildcard. |
--cert-file |
cert.pem |
Path to write the PEM certificate. |
--key-file |
key.pem |
Path to write the PEM private key (PKCS8, mode 0600). |
--valid-duration |
825 days (19800h) |
How long the cert stays valid, e.g. --valid-duration 8760h for a year. |
--ecdsa |
P-256 |
Curve: P-224/P-256/P-384/P-521. Set to "" to disable ECDSA (then pass --rsa). |
--rsa |
(off) | RSA key size (e.g. 2048, 4096); set only when you specifically need RSA instead of the default ECDSA key. |
--ca |
false |
Issue a self-signed CA (CA:TRUE, keyCertSign) instead of a server leaf — see below. |
Selecting no key type at all (--rsa 0 --ecdsa "") is an error: exactly one
of them must produce the key.
The --ca mode is not for --tls-cert
--ca makes the cert its own Certificate Authority — for when you want a
private root to go sign further certs with. It is not something to hand to
MinIO’s --tls-cert: ValidateBYOCert inspects the file’s first certificate
and rejects one that is a CA, because a CA’s private key is a signing key, not
a server key. The default (no --ca) is the server leaf you do want.
Standalone binary: gencert
The same generator is built as a standalone program for use outside pg:
Identical flag set, single-dash form (-host, -cert-file, …). Both front the
one internal/certgen package, so their output is the same kind of
certificate — pg cert is just the path that keeps you inside the CLI.
Example: full BYO chain in one host
Step 3 references the .crt you minted, because on this host you already hold
it. A remote host that never saw the file pulls it back with one handshake
instead of an scp — fetch-ca recognizes the self-signed leaf MinIO serves and
saves it as the anchor:
Related
- Silo — the same BYO TLS surface on Pigsty’s MinIO fork
- MinIO — the addon itself, and its bring-your-own certificate and generating a test certificate sections
- Example: HA + self-CA MinIO — the whole chain, run end to end in an isolated environment
- Cluster Backup — the S3 repository and WAL archiving that consume this certificate
- Patroni HA — the
pg hacommand set
10 - S3 Storage High Availability
A Patroni cluster survives a node loss; the S3 repository its WAL and backups stream into must survive one too, or “high availability” quietly ends at the backup path. The store is a MinIO or silo addon (the two are interchangeable — silo is Pigsty’s MinIO fork with the same feature surface). Two fault domains matter, and they are orthogonal:
| Fault | Who absorbs it |
|---|---|
| A disk dies inside a node | the drive-level EC under SNMD/MNMD, or the storage layer under MinIO’s data directory (ZFS) |
| A node dies entirely | MinIO’s own erasure coding (EC) across hosts |
This page records the shapes pgcli supports for combining the two, and how to pick among them.
The deployment modes
MinIO classifies its layouts by node count and drive count per node (SNSD / SNMD / MNSD / MNMD). pgcli exposes all four:
| Mode | Shape | Use it for |
|---|---|---|
| SNSD (single-node, single-drive) | one node, one data directory — the default | dev, test, demos — and, paired with ZFS below, any single-host deployment that needs disk redundancy |
| SNMD (single-node, multi-drive) | one node, several drives — one --drive per drive |
a single host with N ≥ 4 data disks that must survive a disk loss without a filesystem layer |
| MNSD (multi-node, single-drive) | ≥ 4 nodes, one data directory per node | compact high-availability deployments |
| MNMD (multi-node, multi-drive) | ≥ 2 nodes, several drives each — --drive plus the full --endpoint matrix |
surviving a disk loss and a node loss without a filesystem layer |
SNMD is native: pg addon install minio --drive /mnt/disk1 --drive ...
(one flag per drive) starts one MinIO process that erasure-codes across the
drives. What it buys, measured on a live 4-drive set: 2 parity shards by
default, so it tolerates 2 drive failures; with 1 drive down reads and
writes continue, with 2 down reads still succeed but writes are refused (the
quorum boundary); usable capacity is about half the raw total; a drive that
returns is healed by MinIO itself. The same --drive works on
silo.
MNMD is native too: keep --drive for this node’s drives and add the whole
cluster’s host×drive endpoint matrix with --endpoint — one URL per drive on
every node, each naming that node’s /data1../dataN slot. Measured on a live
4-node × 4-drive set (both minio and silo): 16 drives online report EC:4
in a single erasure set of stripe size 16; losing one whole node (12/16)
keeps reads and writes working; losing a second node (8/16) refuses writes
and fails reads, and the set self-heals to 16/16 on restart. The full
walkthrough is Addons → MinIO → Multi-Node Multi-Drive
(MNMD).
SNMD vs ZFS: two ways to survive a disk
Both protect a single host’s data against disk loss; they differ in where the redundancy lives and what that costs.
native SNMD/MNMD (--drive) |
ZFS pool under --data-dir |
|
|---|---|---|
| quorum unit | the drive — MinIO counts drives as failure members | invisible to MinIO — one big drive, the pool absorbs disk loss |
| layout changeable later | fixed at install (the drive set is the EC set) | freely — swap disks, grow, migrate raidz1→raidz2; --data-dir never changes |
| usable capacity, 4 disks | ~half (2 data + 2 parity) | depends on vdev: raidz2 = 2×disk, raidz1 = 3×disk (more usable) |
| rebuild | MinIO heals a returned drive | ZFS resilvers locally |
| serves which modes | SNMD (single host) and MNMD (across hosts) | SNSD and MNSD (same recipe under each node) |
| heterogeneous nodes | every node must contribute the same drive count | a 2-disk node and a 6-disk node look identical (one endpoint each) |
The honest tradeoff: native multi-drive is simpler — one command, no
filesystem to provision, and it gets you disk redundancy with zero ZFS setup.
ZFS is more flexible — the layout stays changeable underneath, it saves
more capacity on the same disks (raidz1 keeps 3 of 4 usable where SNMD’s
EC:2 keeps 2 of 4), and it tolerates nodes that are not identical. So:
- single host, want disk redundancy without touching ZFS → SNMD
- single host, want the most usable capacity / a layout you can change later → SNSD + ZFS
- distributed cluster, disk and node failure → both paths are first-class now: native MNMD, or MNSD + per-node ZFS — see “HA with 4 hosts, each with several disks” below
Neither is wrong on a single box — SNMD’s 50% capacity is the price of not managing a pool, and ZFS’s flexibility is the price of provisioning one. The rest of this page documents the ZFS path and the MNMD alternative for multi-host sets.
ZFS: the flexible disk layer
The recipe is identical under SNSD and MNSD: build a zpool from the host’s
data disks, create one dataset for the store, and point --data-dir at its
mount point. MinIO sees “one big reliable drive”; it never learns how many
physical disks are under it.
The pool must sit on separate devices from the root filesystem — which is also
what MinIO’s drive check and
pgcli’s install-time advisory want (drive is part of root drive, will not be used). A ZFS mount on its own disks satisfies that trivially.
Layout by disk count
| Data disks | SNSD (ZFS is the only defense) | MNSD node (EC above absorbs node loss) |
|---|---|---|
| 1 | no redundancy — fine for dev/test, a dead disk means re-seeding the repo | one big disk per node is the plain MNSD shape; no ZFS needed |
| 2 | mirror |
mirror |
| 3 | raidz1 (2D usable) |
raidz1 |
| 4 | raidz2 on spinning disks; raidz1 when SSD rebuilds are quick and the capacity matters; 2 × mirror for write-heavy stores |
raidz1 — one parity is enough to keep disk loss invisible to MinIO |
| 5–8 | raidz2 ((N−2)D usable) |
raidz1, or raidz2 with large HDDs |
| > 8 | prefer two smaller raidz2/raidz1 vdevs striped over the pool — a resilver across 12+ disks is a long exposure window |
same: keep any single vdev ≤ ~8 disks |
The two columns differ in one line of reasoning: under SNSD ZFS must survive the disk and whatever happens during its rebuild, so parity depth buys safety outright; under MNSD, ZFS only has to keep a disk failure from escalating into a node loss, so one parity plus quick local resilver is the job — and beyond that, prefer more nodes over deeper local RAID.
For a 4-disk host under SNSD — the most common single-box question — the usual choices are:
| Layout | Usable | Survives | Character |
|---|---|---|---|
raidz2 (like RAID6) |
2 × disk | 2 disks | the safe default on spinning disks: resilvers on large drives are long, and raidz1’s single parity does not survive one more loss during a rebuild |
raidz1 (like RAID5) |
3 × disk | 1 disk | the capacity pick for SSDs / small disks, where a resilver takes minutes, not hours |
2 × mirror (striped) |
2 × disk | 1 disk per mirror (2 if in different mirrors) | best small-write performance — wide raidz is the worst shape for it |
Under SNSD ZFS is the store’s only line of defense, so the default leans
conservative: raidz2 unless the disks are fast enough to rebuild quickly and
the capacity is worth the thinner margin.
MinIO’s own guidance to avoid RAID underneath its EC mode targets the double-redundancy of RAID + cross-node EC. Under SNSD there is no EC — ZFS is the only protection the data has — so the guidance does not apply, and following it literally (“single drive, no ZFS”) on a 4-disk host would mean losing the whole store to one dead disk.
Co-locating with PostgreSQL: ZFS caches aggressively (ARC, by default up to half of RAM). On a host that also runs the database, cap it — e.g.
options zfs:zfs_arc_max=8589934592in/etc/modprobe.d/zfs.conf— so the two workloads do not fight over memory.
HA with 4 hosts, each with several disks
This is the hybrid the layout table above points at, and there are two first-class ways to build it: native MNMD — MinIO erasure-codes across every disk of every node — or MNSD + a ZFS pool under each node’s data directory — MinIO sees one drive per node and the pools absorb disk loss. They fail differently, so the choice is real either way.
Native MNMD
One matrix, no filesystem layer: each node passes its own drives with
--drive and every node carries the identical host×drive endpoint list. For
TLS, generate one shared leaf and install the same pair on every node (see the
TLS note after the measurements below).
Measured on a live 4-node × 4-drive set (the same on minio and silo): the 16-drive set reports EC:4 in one erasure set of stripe size 16 — losing 4 drives of 16 stays healthy.
| Event | Online | Effect (measured) |
|---|---|---|
| 1–4 drives lost | ≥ 12/16 | reads and writes continue; returned drives are healed by MinIO |
| 1 node down (its 4 drives) | 12/16 | reads and writes continue — a 64 MiB round-trip stayed byte-identical |
| 2 nodes down | 8/16 | writes refused (Resource requested is unwritable), reads fail too |
| nodes restarted | 16/16 | self-heals |
Usable capacity follows the parity ratio: EC:4 over 16 drives keeps 12/16
of raw bytes. The costs: the layout is fixed at install (the matrix is the EC
set), every node must contribute the same number of drives, and pgcli cannot
detect a matrix typo on another node — keeping the N pg.yaml files
identical is the operator’s job.
For TLS across the grid, serve one shared certificate: mint a single self-
signed leaf with pg cert --host <every node address, comma-separated> and
install it on every node with --tls-cert/--tls-key (the same two files
byte-identical everywhere, the endpoint matrix all https://). The grid then
forms with no per-node CA to reconcile — verified on a live 4-node set: the ring
came up Network: 4/4 OK with an empty CAs/ dir, because that shared leaf is
its own trust anchor and every node already holds it. The fallback is the
generated mode: if you let each node run a bare --tls, every host mints its
own CA and the first cross-node handshake dies with x509: certificate signed by unknown authority until you seed one shared CA into every node’s cert dir by
hand — which the shared-pg cert path avoids entirely.
MNSD + per-node ZFS
MNSD across the hosts, ZFS under each node’s data directory. Each node
exposes exactly one endpoint (its ZFS-backed /data); disk failures are
healed by the local pool and never reach MinIO; node failures are absorbed by
EC quorum.
What each layer then tolerates on a 4-node cluster:
| Event | Handled by | Effect |
|---|---|---|
| 1 disk in a node’s raidz pool | ZFS (resilver) | invisible to MinIO, no quorum math |
| 1 node down | EC (writes need ⌈4/2⌉+1 = 3 of 4 online) | reads and writes continue |
| 2 nodes down | EC reads only (⌈4/2⌉ = 2 of 4) | readable, writes refused until a node returns |
| a disk and its node failing together | EC (3 of 4) | still safe — the surviving pools heal after the node returns |
Two properties make the scheme flexible:
- Endpoint count = node count, not disk count. A node with 2 disks and a
node with 6 disks look identical to MinIO. The four hosts may run
different layouts —
raidz1×4 disks here, 2-waymirrorthere,raidz2×6 over there — heterogeneity costs nothing. - Pool capacities should be roughly aligned. EC sets the usable capacity
of the whole cluster to the weakest member’s free space (3 × 12T pools +
1 × 4T pool → you get 4T × striping, not 40T). If the disks genuinely
differ, thin provisioning (
zvol-based sparse datasets, or simplyzfs set refquota) hides the mismatch from MinIO — the quota caps the big pools at the small one’s size, and nothing wastes a rebuild.
Why not stripe the per-node pool to reclaim capacity. MinIO’s “no RAID
under EC” advice is about capacity, and the arithmetic is real: a plain
4-node MNSD EC set already halves the raw total, so an unmirrored raidz_none
pool on each node (all disks striped, zero local redundancy) does squeeze out
roughly a third more cluster capacity than raidz1 would. But EC counts a
failure in nodes, and a striped pool turns a single dead disk into a whole
offline node — one ordinary disk failure burns a slot of the budget EC set
aside for losing an entire machine, forcing a network-wide rebuild of that
node’s whole pool and leaving zero margin until it finishes. The capacity is
only “free” because you quietly downgraded disk fault-tolerance to node
fault-tolerance. raidz1 per node is the point where the two layers stop
stealing from each other: the pool absorbs disk loss invisibly, MinIO’s EC
budget stays reserved for node loss. (If the reclaim-everything answer truly
fits, the shape that maximizes it is plain MNSD on one big disk per node —
no ZFS at all — not a striped pool that hides the same single-point risk one
layer down.)
Which of the two
| native MNMD | MNSD + per-node ZFS | |
|---|---|---|
| setup | one matrix command per node, no filesystem to provision | zpool + dataset on every node first, then one endpoint per node |
| what EC counts | drives — one node down is just 4 of 16 members lost | nodes — a disk failure never reaches MinIO at all |
| measured fault margin | 4-node × 4-drive set: OK to 12/16 drives (one full node), fails at 8/16 | 4 nodes: reads at 2/4, writes at 3/4 nodes |
| usable capacity | 12/16 of raw on that set (EC:4) — the parity is MinIO’s choice for the stripe | local raidz1 keeps 3/4 per pool, then EC halves across nodes |
| layout later | fixed at install — the matrix is the EC set | changeable — disks, vdevs, layouts move without touching MinIO |
| heterogeneous nodes | not possible: every node must contribute the same drive count | natural — each node is just one endpoint, any local layout |
| TLS | one shared cert serves the whole grid: pg cert a single leaf covering every node address, --tls-cert/--tls-key it identically everywhere (no CA to reconcile) |
the same — it’s still a TLS grid between nodes, so the same shared-leaf setup applies |
Pick MNMD when the nodes are identical, the disk plan is settled, and you want disk and node redundancy from MinIO alone with nothing to provision underneath. Pick MNSD + ZFS when nodes differ in disk count or size, the layout may change later, or you want to reason about failures in whole nodes instead of in drives.
Choosing a shape
| Situation | Shape |
|---|---|
| laptop / demo / CI | SNSD, default data dir — nothing to decide |
| one host with N ≥ 4 data disks, backups must survive a disk | SNMD (--drive × N) — zero filesystem setup; or SNSD + ZFS (raidz2 on spinning disks, raidz1 on fast SSDs) for more usable capacity and a changeable layout |
| a few hosts, one data disk each, must survive a host | MNSD plain |
| identical hosts with several data disks each, must survive a disk and a host | MNMD (--drive + full --endpoint matrix) — no filesystem layer; or MNSD + per-node ZFS, see “Which of the two” above |
| hosts whose disk counts/sizes differ | MNSD + per-node ZFS (layouts may differ per node; keep capacities aligned) — MNMD needs equal drive counts |
The store’s TLS story is independent of the topology: whichever shape you
pick, --tls (or pg cert-minted BYO certificates) works the same — see
Generating a Certificate.
Related
- Addons → MinIO — install,
--tls, distributed mode constraints (the hard rules this page builds on), and the MNMD walkthrough - Addons → Silo — the drop-in fork; identical modes
- Cluster Backup — how the repo is wired into Patroni
(
pg backup setup) - Example: HA Cluster with a Self-CA MinIO — an end-to-end verified walkthrough of the single-node variant