Patroni HA
Patroni is the de-facto standard for
PostgreSQL high availability: it manages each postmaster’s lifecycle, streams
replication between members, and performs automatic failover when the
leader is lost. pgcli exposes Patroni as a distinct mode — pg ha — rather than
folding it into the plain pg instance path.
This is a separate mode, not an addon subcommand. Patroni (not pgcli) owns the PostgreSQL process.
pg hamembers do not appear incfg.Instances, share no instance lifecycle code, and are not created bypg create. The one thing it borrows from the addon system is the etcd DCS.
Linux only, like the etcd addon: Patroni members rely on podman host networking. Both root and rootless podman are supported. On macOS the commands fail fast with a clear message.
Ownership boundary
The single most important thing to internalize is who owns what:
| Responsibility | Owner |
|---|---|
Patroni container (run/start/stop/rm), image, patroni.yml, port allocation, passwords, DCS wiring |
pgcli |
postmaster lifecycle, initdb, PostgreSQL config rendering, replication slots, failover |
Patroni |
| switchover / failover / pause / edit-config | pgcli wrapping patronictl in short-lived containers |
pgcli never edits a running PostgreSQL’s config directly, and never runs
initdb — Patroni does both. This is also why the plain PG image’s
docker-entrypoint-initdb.d convention (the admin role / default database)
does not apply here: Patroni bootstraps its own cluster, so the role system
is postgres (superuser), replicator, and rewind_user instead.
Container lifecycle ≠ safe PG restart
Because Patroni is PID 1 inside its container:
pg ha createis re-install semantics (stop + recreate the container). Recreating the leader takes that node fully offline and triggers a failover. Changing a member’s config by re-runningcreateis fine; for dynamic settings preferpg ha edit-config, which never touches the container.pg ha start/pg ha stopare raw container start/stop. Stopping the leader’s container is exactly as disruptive as the node crashing — Patroni will fail over to a replica. For planned maintenance,pg ha pausefirst (it disables auto-failover), then stop the container.- A paused cluster has no automatic failover until resumed.
How It Works
pg ha runs one Patroni container per member. All members of a scope point at
the same DCS (a etcd cluster) which holds the cluster’s dynamic
configuration and leader lock:
- The first
pg ha createfor a scope bootstraps: Patroni runsinitdb, wins the leader race, and becomes the leader. - Every later member automatically
pg_basebackups from the current leader and begins streaming — no flag distinguishes “add” from “join”. pg hacontrol commands (switchover,pause, …) runpatronictlin an ephemeral container against the DCS, so they work even when no member container is running locally.
The DCS is the source of truth: Patroni re-renders each member’s
postgresql.conf/pg_hba.conf from it every loop, so hand-edits on disk are
lost — use pg ha edit-config instead.
Scope and the DCS layout
The <scope> in pg ha create <scope> is the Patroni cluster name — the
identity of one HA cluster. Every command that takes a scope (status,
switchover, failover, pause, ctl, …) names the same cluster; the first
create for a scope bootstraps it, later creates with the same scope add
members. (--member is the per-host node name inside the cluster — one
member per machine.)
In pg.yaml the cluster is keyed by that scope under addons.patroni.<scope>.
What actually lands in etcd
Patroni stores everything under a fixed etcd namespace plus the scope:
- namespace:
/service/— Patroni’s top-level key, not changed by pgcli. - scope: pgcli writes
PatroniScope(scope)intopatroni.yml, which is your scope plus the pgcli namespace suffix. Patroni has no namespace concept of its own (its etcd prefix is the raw scope), so pgcli bakes the suffix into the scope to keep two pgcli namespaces sharing one etcd from cross-talking.
So the full etcd prefix for a cluster is:
The keys below the prefix are Patroni’s own DCS layout. To inspect them, point
pg etcdctl at a running etcd member of the same host/config:
Note the scope shown there carries the namespace suffix, even if you configured the cluster with the bare name.
Install
Prerequisite: a DCS. Either reuse local etcd addon members or point at an external etcd.
pg ha status app renders patronictl list — exactly one Leader row and the
rest Replica … streaming:
With no argument, pg ha status rolls up every cluster: container state, member
count, and each member’s pg=/rest= ports.
Choosing a DCS
Patroni only needs a reachable etcd cluster, so there are two layouts:
-
Co-located — run the etcd members as addons on the same hosts as the Patroni members and pass
--etcd m1,m2,m3. Simplest: one machine per role, no extra hosts. Good for a starter 3-node HA cluster. -
Dedicated DCS hosts — run the etcd cluster on its own (virtual) machines and point every Patroni member at it with
--etcd-endpoints. Each line below runs on a different machine (pgcli is per-host):This is the more available topology: etcd’s quorum survives losing a PG host, and re-imaging a database machine never takes the DCS down with it. The PG hosts only ever talk to the DCS;
--etcd-endpointslists all member client URLs, so Patroni falls through to the next endpoint if one etcd host is down (writes still need the etcd quorum itself — 2 of 3). Keep the endpoint list identical on every PG host.A single endpoint (
--etcd-endpoints 10.0.0.20:2379) does work — that one etcd member serves the whole cluster, and the other two are hidden behind it. But it re-introduces a single point of failure: if E1 goes down, Patroni cannot reach the DCS even though the etcd quorum is healthy, and a leader that fails to renew its lock demotes itself. List every endpoint you actually have.
A dedicated odd-sized etcd cluster (3 or 5) is the production recommendation; co-locating is fine for dev and small footprints.
Commands
| Command | What it does |
|---|---|
pg ha create <scope> --member <m> … |
Register + (re)install one member — recreate = node offline |
pg ha list |
Compact table of all HA clusters (scope, members, DCS, status) |
pg ha status [scope] |
All clusters, or one cluster’s patronictl list |
pg ha switchover <scope> |
Planned leader change (patronictl confirms; --yes to script) |
pg ha failover <scope> |
Promote a replica now |
pg ha pause / resume <scope> |
Disable / re-enable automatic failover |
pg ha edit-config <scope> -- … |
View or patch the dynamic config in the DCS (never recreates a container) |
pg ha start / stop <scope> --member m | --all |
Raw container start/stop (see lifecycle caveat) |
pg ha remove <scope> --member m | --scope-all [--clean-data] [--force] |
Remove member(s); --scope-all also clears the DCS |
pg ha passwords <scope> [--file F] |
Export the stored password set (the --passwords-file format, for other hosts) |
pg ha remote <scope> [--member m] [--ssh-port P] |
Register a member living on another host (typically the cross-host leader) for backup SSH |
pg ha remote remove <scope> --member m |
Undo such a registration (the remote container is untouched) |
pg ha extension install/remove/list/apply |
Install, remove, or list PostgreSQL extensions (see Extensions in HA) |
pg ha exec <scope> "<sql>" [--member m] [--database db] |
Run one-shot SQL against the leader (or a named member) — no dsn, no container exec |
pg ha psql <scope> [--member m] [--database db] [-- <psql-args>…] |
Interactive psql against the cluster, same resolved target |
pg ha ctl <scope> -- <patronictl args…> |
Passthrough to any patronictl command |
Flags after -- reach patronictl verbatim (cobra strips the --), so
pg ha ctl app -- show-config, pg ha ctl app -- topology, and
pg ha edit-config app -- -s synchronous_mode=true --force all work. edit-config --show is a convenience alias for ctl … -- show-config.
Cross-host members
Each host runs its own pgcli managing only that host’s members; the cluster
reassembles through the shared DCS, so two hosts’ pg.yaml files each hold a
partial view of one scope.
Cross-host checklist (all three must line up on every host):
- Ports — auto-assignment is per-host, but
connect_addressis stored in the DCS cluster-wide. Pass--host-port/--restapi-portwith the same value on every host, or replicas can’t reach each other. - Passwords — Patroni’s replication / rewind / REST-API auth is cluster-wide.
Export the first host’s generated set with
pg ha passwords app --file app-passwd.ymland pass--passwords-file app-passwd.ymlto every otherpg ha create. --advertise-host— required for cross-host members; it flips the listen address to0.0.0.0and puts a reachable IP inconnect_address. Left empty, a member is loopback-only. This applies to the first (bootstrap) member too: it becomes the leader, and its loopbackconnect_addressis what every later host would try topg_basebackupfrom — cross-host joins fail forever once the cluster was bootstrapped with the default. If you may ever add a member on another host, pass your LAN IPv4 from the very firstcreate.- Firewall — allow the two ports (PG + REST API) pairwise between members.
pgcli does not validate the peers’ config; pg ha status shows each member’s
connect_address so you can self-check reachability.
Backing up a cross-host leader
Each host’s pg.yaml records only its own members, so this host cannot see a
leader running on another host — yet pgBackRest’s stanza-create and full
backups must connect to the leader.
This now happens automatically: after pg ha create installs a member, it
publishes that member’s pgcli-private ports (SSH / REST API) into an etcd
registry at /pgcli/ha/<scope>/<member> (the topology itself — member name +
host:pgport — Patroni already stores in the DCS). The backup config generators
then run patronictl list for the cluster-wide topology and look up each remote
member’s SSH port in the registry, so pgbackrest.conf / ssh_config include
members on every host, with pg1-host and the SSH HostName pointing at the
remote IP (local members still go via 127.0.0.1). pg ha start/stop/remove,
autostart, and pg ha extension skip members that aren’t local — they belong to
their own host’s pgcli.
Just refresh the backup config:
Cross-host backup trust is now automatic. A cluster-wide stanza lists every
member as a pg*-host, so a full backup / check from any host SSH-probes the
members on all hosts to find the primary — each of those member sshd processes
must therefore accept the initiating host’s backup key, and its
--advertise-host already flipped the listener to 0.0.0.0 (see above) so
cross-host SSH reaches it. pg backup setup
handles this: each host’s create publishes its backup public key into the
same /pgcli/ha/<scope>/<member> registry, setup merges every member’s key
into a per-cluster authorized_keys file, and each member container bind-mounts
that file as an extra AuthorizedKeysFile. Because sshd re-reads it on every
login and the merge rewrites the file in place, a host that joins later is
trusted by the running members without a restart (the file only needs the
one-time recreate to add the mount, which setup does inside the pause window
alongside the archive-config recreate). A member removed with pg ha remove
drops its registry key, so the next merge stops trusting it.
The S3 repository CA is distributed the same way: the host that configured
ca_file publishes the certificate to /pgcli/ha/<scope>/.repo/ca, and a
joiner that leaves ca_file empty pulls it and points its own ca_file at the
pulled copy. When the store itself lives on yet another machine, the first host
gets the CA without any file copy too — pg backup fetch-ca <store>:<port>
pulls it out of the endpoint’s TLS chain (see the MinIO doc). Only public
material — SSH public keys and the self-signed CA —
ever enters the registry; private keys, passwords, and the S3 secret_key never
do (the DCS link is unauthenticated).
Fallback: a remote member created on its host before that host upgraded
pgcli has no registry entry, so the generators fall back to this scope’s base
SSH port for it (every host starts at patroni_ssh_start_port, so the first
member is usually right). When it isn’t, register just that member manually to
override (creates no container):
S3 backups and WAL archiving
Patroni clusters ship with archiving off — full backups alone can never
reach a point-in-time. Enabling the S3 repository
(backup docs, English;
中文) flips it on. Each cluster is one
stanza — pgcli_<scope>, matching the namespaced DCS scope, so two pgcli
namespaces sharing one S3 bucket never collide — whose backup-side config lists
every member as a pg*-host, so pgBackRest finds the primary itself and keeps
working across failovers. pg backup setup renders the identical
archive_command into each member’s local patroni.yml (pgcli never writes
archive GUCs into DCS — that is edit-config’s territory, and Patroni applies
a local value whenever DCS does not manage the key), recreates stale member
containers replicas-first inside a pause window (the recreate is also the
postmaster restart archive_mode needs), and runs stanza-create + check
for the cluster. From then on WAL streams to S3 continuously, from whichever
member holds the leader lock.
Expect a short per-member offline window and a planned leader demotion at
the end — the same recreate semantics as pg ha create.
Passwords
The first pg ha create for a scope generates four credentials — superuser,
replication, rewind, and the restapi basic-auth pair (restapi_user /
restapi_password) — and stores them in pg.yaml under the cluster. You never
pass passwords on the command line: the first member generates them, and
pg ha passwords exports the stored set:
(Without --file the YAML goes to stdout — prefer --file so the secrets stay
out of shell history and scrollback.)
The four roles and their default usernames (fixed by pgcli — you only ever set
the passwords, which default to a random 16-char string per role). Override the
length with --password-length on pg ha create / pg create (8–64; it only
applies to the generate path, not to a --passwords-file or an already-stored set):
| pg.yaml key | Role | Default username | Default password |
|---|---|---|---|
superuser |
PostgreSQL superuser | postgres |
random 16-char, generated |
replication |
replication / streaming | replicator |
random 16-char, generated |
rewind |
pg_rewind role |
rewind_user |
random 16-char, generated |
restapi_user / restapi_password |
Patroni REST API basic-auth | postgres |
random 16-char, generated |
The postgres / replicator / rewind_user usernames are baked into the rendered
patroni.yml (postgresql.authentication); only the passwords are the generated /
--passwords-file part. restapi_user is the one username you can override, via the
passwords file.
For full control (or when you’d rather author the file yourself) use the same
format with --passwords-file:
pg ha create prints where each password came from (generated-and-stored vs. a
file path). The rendered patroni.yml is written mode 0600 because it embeds
all of them.
Auto-start on Boot
Like the other infra addons, a Patroni member’s container can be brought up
after a host reboot — but it only starts the existing container, reading the
patroni.yml already on disk; it never re-renders config or re-creates data.
Members are toggled one at a time (--ha --scope <scope> --name <member>).
Start order relative to the DCS doesn’t matter: Patroni retries until etcd
answers, then re-elects normally. See the autostart page for
the boot-service mechanics.
Logs
Container name is pgcli-patroni-<scope>-<member> (namespace-prefixed when a
namespace is set).
Connecting
pg ha exec and pg ha psql are the zero-config path: pgcli resolves the
leader from the DCS itself, authenticates with the stored superuser password,
and runs psql in a throwaway container — no dsn to assemble, no member
container to enter, and remote members are reachable (plain TCP + scram, so a
replica on another host works too; --member aims at a specific node).
For clients outside pgcli, connect to the leader’s PostgreSQL port. Find the
leader with pg ha status app (the Leader row’s Host is its
connect_address), then point pg psql at it. On a single host with default
loopback-only members, that is 127.0.0.1:<host_port>.
After a failover the leader changes, so a fixed connection string should be avoided unless fronted by a pooler or the Patroni REST API’s leader redirect.
For a stable endpoint that survives failover — and optional read/write separation — put HAProxy in front of the cluster.
Planned leader changes
Patroni owns failover; pg ha wraps the patronictl verbs that move the
leader on purpose. All three take just the scope:
switchoveris the planned, graceful one: the current leader is demoted first, so nothing is in flight when the candidate is promoted. Both members must be healthy and caught up. The old leader automatically rejoins as astreamingreplica a few seconds later (you may briefly see it asstoppedwhile Patroni restarts its postmaster). A new timeline is opened — this is normal, not a split-brain.failoverforce-promotes a replica immediately without a clean handoff from the leader. Use it only when the leader is already gone or you are deliberately discarding it; on a healthy cluster it just causes an unnecessary blip. Reach forswitchoverfor anything planned.pauseturns off automatic failover cluster-wide. Do this before any planned container surgery —pg ha stop --all, apg ha createrecreate, orpg ha extension— so Patroni doesn’t promote a replica out from under you.pg ha resumeputs auto-failover back. (A paused cluster has no automatic failover until resumed.)
These are pure DCS operations — they run patronictl in a throwaway container,
touch no member’s container or data, and work even when some members live on
other hosts. Like every pg ha control command, the scope is resolved to the
namespaced DCS key (app → app-default), so a single-host or cross-host
leader is switched the same way.
Because clients dial the leader’s port, the connection target moves after a switchover — see Connecting and put HAProxy in front if you need one stable endpoint.
pg ha vs. pg replica — which to pick
pgcli has two ways to get a standby:
pg replica + pg failover |
pg ha (Patroni) |
|
|---|---|---|
| Model | Manual: pg replica builds a standby, pg failover promotes it on demand |
Automatic: Patroni keeps N members in sync and self-heals |
| Failover | A human runs pg failover; primary stays down until promoted |
Patroni detects the loss and promotes a replica within ~30s |
| Ownership | pgcli drives postmaster (as with any instance) | Patroni drives postmaster; pgcli owns containers only |
| Best for | Simple read-scaling, single planned promotion, staying on pg-instance tooling |
Zero-RTO availability requirements, unattended failover |
If you need a standby you control by hand, use pg replica. If you need the
cluster to survive a node crash without a human, use pg ha.
Configuration
The rendered patroni.yml per member is derived entirely from pg.yaml under
addons.patroni.<scope>. A typical entry:
Ports come from two independent pools — patroni_start_port (default 35532) for
PostgreSQL and patroni_restapi_start_port (default 39060) for the REST API —
so Patroni members never collide with plain-instance or addon ports.
The DCS-scope key is the scope plus the namespace suffix (Patroni has no namespace concept of its own, so the prefix is baked into the scope to keep two pgcli namespaces sharing one etcd from cross-talking).
Default patroni.yml
The file below is what pgcli renders and mounts read-only into each member’s
container. Passwords are auto-generated at bootstrap time; --passwords-file
lets you supply your own set.
Key points:
scope= Patroni cluster name. Combined withnamespace, it forms the etcd key prefix.etcd3.hosts=host:portonly (nohttp://scheme). Multiple endpoints are comma-separated.bootstrap.dcsvalues are defaults only — they take effect on first bootstrap; afterwards usepg ha edit-config.postgresql.listen: 0.0.0.0is set when--advertise-hostis used (cross-host). Without it, listen is127.0.0.1.- Passwords are generated once and stored in
pg.yaml; exported viapg ha passwords.
Dynamic Configuration
Patroni’s dynamic configuration lives in the DCS (etcd) under /service/<scope>/config.
Every member reads it on each loop (every loop_wait seconds) and applies
changes live — so pg ha edit-config is the right way to tune runtime parameters
without restarting containers.
bootstrap.dcsis one-time. Thebootstrap.dcsblock inpatroni.ymlonly takes effect when the first member of a scope runspg ha create(i.e. the cluster bootstrap). Once Patroni writes the config into the DCS, subsequent changes tobootstrap.dcsin the YAML file are completely ignored — even on re-install (pg ha create). To change dynamic configuration after bootstrap, usepg ha edit-config.The common pitfall: re-running
pg ha createdoes not re-readbootstrap.dcsfrom the YAML — Patroni sees the existingconfigkey in the DCS and uses it directly.
Method Description pg ha edit-config app -- -s key=valueRecommended — pgcli’s standard way pg ha ctl app -- edit-configPassthrough to patronictl, same effect Patroni REST API ( PATCH /config)Requires access to a member’s REST API port
Persistence: changes are written directly to etcd, not to container files.
Container restarts, pg ha start/stop, or even pg ha create (re-install)
do not lose these settings — new members automatically pick up the latest config
from the DCS.
Viewing the current config
This is a convenience alias for pg ha ctl app -- show-config. The output is
the full JSON blob stored in /service/<scope>/config.
Modifying parameters
After a change, all members apply it on their next loop (within loop_wait
seconds). No restart needed.
Common tunable parameters
| Parameter | Default | Description |
|---|---|---|
loop_wait |
10 |
Seconds between leader-loop iterations (lock renewal, DCS updates) |
ttl |
30 |
Leader lock TTL. If the leader fails to renew within this window, replicas trigger failover |
retry_timeout |
10 |
Timeout for DCS/PostgreSQL operations. Must be < ttl - loop_wait to give the leader at least one retry chance |
maximum_lag_on_failover |
1048576 |
Maximum replication lag (bytes) for a replica to be eligible for promotion. Default 1 MB |
synchronous_mode |
false |
Enable synchronous replication (zero data loss, higher latency) |
synchronous_node_count |
1 |
How many synchronous standby nodes (when synchronous_mode=true) |
use_pg_rewind |
true |
Use pg_rewind to rejoin a failed leader (faster than full pg_basebackup) |
use_slots |
true |
Use replication slots (prevent WAL loss when a replica disconnects) |
failover_timeout |
0 |
How long to wait before failover (0 = immediate when leader is lost) |
PostgreSQL runtime parameters can also be set under postgresql.parameters:
These trigger a PostgreSQL reload (or restart, depending on the parameter’s
context). Check pg_hba.conf and postgresql.conf parameter documentation for
which settings require a restart.
What not to edit
- Do not edit
patroni.ymlon disk — it is regenerated frompg.yamlon everypg ha create, and Patroni reads dynamic config from the DCS anyway. - Do not edit
postgresql.confdirectly — Patroni overwrites it each loop from the DCS config. - Do not change
scopeornamespace— these are baked into the DCS key at bootstrap time and cannot be changed without recreating the cluster.
Full parameter reference: Patroni Dynamic Configuration covers every DCS-tunable parameter with defaults, constraints, and examples.
Extensions
Full reference: Extensions in HA Clusters covers the orchestration order, cross-host workflow,
shared_preload_librariesordering, and builtin-only fast path.
Notes
- Linux only (root or rootless). Rootless members run as the host user via
--userns=keep-id; root members chown config and data dirs to postgres (uid 999) so the container can read0600config and write its data dir. pg_hba.confis permissive by design (host all all all scram-sha-256- a
replicationline, pluslocaltrust lines). Rootless podman’s pasta rewrites loopback sources, and Patroni’s ownreplace_pg_hbastep only ever grants its resolved loopback TCP address duringcustom bootstrap— sopostgresql.use_unix_socket/use_unix_socket_replare set to make Patroni connect to its own instance over the unix socket instead, which is unaffected by the rewrite and keeps bootstrap from deadlocking on its own pg_hba. Tightening the allow-list to a fixed set is a planned future refinement — do not expose these ports to untrusted networks yet.
- a
- The DCS (etcd) has its own security caveats — see the etcd page: pgcli-managed etcd runs without TLS or auth.
- No
init.sh/docker-entrypoint-initdb.d. Patroni bootstraps the cluster itself, so theadmin/default-db convention of plain instances does not exist here; usepostgres(superuser) to connect and create roles. use_slots/use_pg_rewindare enabled: Patroni owns replication slots and can rejoin a crashed leader viapg_rewindinstead of a full rebase.