PostgreSQL disaster recovery#
Ametnes PostgreSQL supports cross-site streaming replication with a promotable standby: you run a primary and a standby in different locations and fail over when the primary is lost.
This is disaster recovery, not automatic high availability
Replication is asynchronous and failover is manual. A primary/standby pair protects against losing a location, but it does not give you zero data loss, automatic failover, or multi-region writes. Read Guarantees and limitations before you rely on it.
Model#
| Field | Meaning |
|---|---|
dr.role |
primary (writable, replication source) or standby (read-only replica of a remote primary). |
dr.replication.user / dr.replication.password |
Shared credentials a standby uses to stream from this primary. |
dr.uri |
Connection URI of the upstream primary, e.g. postgresql://repl_user:secret@primary.example.com:5432/postgres. Required when dr.role is standby. |
A standby is a single node (it cannot also run local read replicas). Replication is asynchronous, so a primary stays writable if the standby is unreachable.
Guarantees and limitations#
| Property | What this setup provides |
|---|---|
| Replication | Asynchronous physical streaming replication |
| Writers | One at a time (the current primary); the standby is read-only |
| Recovery point (RPO) | Greater than zero — writes on the primary that have not yet shipped are lost if the primary fails |
| Recovery time (RTO) | Manual — the time to detect the failure, promote the standby, and repoint clients |
| Failover trigger | Manual (dr.role → primary) |
| Automatic failover | No |
| Failback | Incremental after a clean shutdown; full re-seed after an ungraceful failure |
| Split-brain protection | None — the platform does not fence the old primary; you must ensure it never becomes a writer again |
| Multi-region access | A read-only standby in a second location, with a single writer |
Provision a primary and a standby#
Create the standby with the primary's external host and port:
Use the primary's admin password on the standby
A standby is a physical clone (pg_basebackup) of the primary, so it
inherits the primary's superuser password. Set admin.password on the
standby to the same value as the primary; otherwise the secret the
platform advertises for the standby will not authenticate against it.
Verify replication by writing on the primary and reading on the standby.
Fail over#
Promote the standby by changing its dr.role to primary. The platform runs
pg_promote() and the instance becomes writable.
Promotion does not contact the failed primary, so it works even when the primary is completely unreachable — for example after a blackout.
Promote before repointing clients
Promotion is not automatic. Only the standby you promote can accept writes;
repoint your application once it reports ready.
Fail back#
To return to the original primary, make it a standby of the promoted instance by
setting dr.role = standby and dr.uri to the promoted instance's endpoint.
The platform resynchronises the old primary's data and rejoins it as a streaming
replica — data written on the promoted instance is then replicated back.
How the resynchronisation works depends on how the old primary stopped. For a planned, clean stop it is incremental. For an ungraceful failure it is a full re-seed — see the next section.
What happens in an ungraceful failure (blackout)#
A blackout — power loss, hardware failure, or a host that disappears without shutting down — is not the same as a clean stop, and the two legs of a failover behave differently.
Failover works normally. The standby streams until the primary stops, keeps the data it has already received, and you promote it manually. Because replication is asynchronous, you lose whatever the primary committed but had not yet shipped to the standby (the RPO gap described above).
Failback is not incremental. The graceful failback path uses pg_rewind,
which requires the old primary to have been shut down cleanly. After a
blackout the old primary's data directory is left unclean, so an incremental
rewind is not possible and the node is rebuilt by a full re-seed — a fresh
clone from the promoted instance. If the failed location is unreachable, or its
storage is lost, the old primary has to be rebuilt as a new standby once that
location is available again.
Rebuild the old primary before it serves traffic
After a failover the old primary may hold older data. Do not let it run as a writer again until it has rejoined as a standby of the promoted instance — otherwise you risk a split-brain with two writable copies. The platform does not fence it for you.
When you need stronger guarantees#
This pattern is deliberately simple: one writer, asynchronous replication, manual failover, and an explicit data-loss window. If any of the following are hard requirements, it is the wrong tool:
- Automatic failover, with no manual promote.
- A near-zero or zero data-loss guarantee (RPO ≈ 0).
- Active reads and writes in more than one region.
- Horizontal write scaling.
- No manual operational steps during an incident.
For those workloads, use a distributed, PostgreSQL-compatible database instead. Ametnes is adding this as an option, built on YugabyteDB: consensus-based synchronous replication, automatic failover, and multi-region placement, while still speaking the PostgreSQL wire protocol so your application connects to a single endpoint unchanged. If your requirements match the list above, talk to us about it rather than using the primary/standby pattern.
Operational notes#
- Replication is asynchronous; monitor lag before failing over.
- A role change is a config update (
dr.role/dr.uri), not a new resource — the same DNS name and endpoint are reused. - The replication password on the primary must match the password in the
standby's
dr.uri. - Only ever run one writable instance. The platform does not fence the old primary, so reconfigure or rebuild it as a standby before it accepts traffic.
- Test your drill — including a hard stop of the primary — before you rely on it, so the RPO and failback behaviour are not a surprise during a real incident.