Skip to content

PostgreSQL disaster recovery#

Ametnes PostgreSQL supports cross-site streaming replication with a promotable standby: you run a primary and a standby in different locations and fail over when the primary is lost.

This is disaster recovery, not automatic high availability

Replication is asynchronous and failover is manual. A primary/standby pair protects against losing a location, but it does not give you zero data loss, automatic failover, or multi-region writes. Read Guarantees and limitations before you rely on it.

Model#

Field Meaning
dr.role primary (writable, replication source) or standby (read-only replica of a remote primary).
dr.replication.user / dr.replication.password Shared credentials a standby uses to stream from this primary.
dr.uri Connection URI of the upstream primary, e.g. postgresql://repl_user:secret@primary.example.com:5432/postgres. Required when dr.role is standby.

A standby is a single node (it cannot also run local read replicas). Replication is asynchronous, so a primary stays writable if the standby is unreachable.

Guarantees and limitations#

Property What this setup provides
Replication Asynchronous physical streaming replication
Writers One at a time (the current primary); the standby is read-only
Recovery point (RPO) Greater than zero — writes on the primary that have not yet shipped are lost if the primary fails
Recovery time (RTO) Manual — the time to detect the failure, promote the standby, and repoint clients
Failover trigger Manual (dr.role → primary)
Automatic failover No
Failback Incremental after a clean shutdown; full re-seed after an ungraceful failure
Split-brain protection None — the platform does not fence the old primary; you must ensure it never becomes a writer again
Multi-region access A read-only standby in a second location, with a single writer

Provision a primary and a standby#

primary.tf
resource "ametnes_service" "postgres_primary" {
  name    = "PostgresPrimary"
  project = ametnes_project.project.id
  location = data.ametnes_location.location_a.id
  kind    = "service/postgres:17.6"
  nodes   = 1
  capacity { storage = 20 }
  config = {
    "architecture"           = "Small"
    "admin.password"         = "change-me"
    "dr.role"                = "primary"
    "dr.replication.user"    = "repl_user"
    "dr.replication.password" = "shared-repl-secret"
  }
}

Create the standby with the primary's external host and port:

standby.tf
resource "ametnes_service" "postgres_standby" {
  name    = "PostgresStandby"
  project = ametnes_project.project.id
  location = data.ametnes_location.location_b.id
  kind    = "service/postgres:17.6"
  nodes   = 1
  capacity { storage = 20 }
  config = {
    "architecture" = "Small"
    "admin.password" = "change-me"
    "dr.role"      = "standby"
    "dr.uri"       = "postgresql://repl_user:shared-repl-secret@<primary-host>:<primary-port>/postgres"
  }
}

Use the primary's admin password on the standby

A standby is a physical clone (pg_basebackup) of the primary, so it inherits the primary's superuser password. Set admin.password on the standby to the same value as the primary; otherwise the secret the platform advertises for the standby will not authenticate against it.

Verify replication by writing on the primary and reading on the standby.

Fail over#

Promote the standby by changing its dr.role to primary. The platform runs pg_promote() and the instance becomes writable.

Promotion does not contact the failed primary, so it works even when the primary is completely unreachable — for example after a blackout.

Promote before repointing clients

Promotion is not automatic. Only the standby you promote can accept writes; repoint your application once it reports ready.

Fail back#

To return to the original primary, make it a standby of the promoted instance by setting dr.role = standby and dr.uri to the promoted instance's endpoint. The platform resynchronises the old primary's data and rejoins it as a streaming replica — data written on the promoted instance is then replicated back.

How the resynchronisation works depends on how the old primary stopped. For a planned, clean stop it is incremental. For an ungraceful failure it is a full re-seed — see the next section.

What happens in an ungraceful failure (blackout)#

A blackout — power loss, hardware failure, or a host that disappears without shutting down — is not the same as a clean stop, and the two legs of a failover behave differently.

Failover works normally. The standby streams until the primary stops, keeps the data it has already received, and you promote it manually. Because replication is asynchronous, you lose whatever the primary committed but had not yet shipped to the standby (the RPO gap described above).

Failback is not incremental. The graceful failback path uses pg_rewind, which requires the old primary to have been shut down cleanly. After a blackout the old primary's data directory is left unclean, so an incremental rewind is not possible and the node is rebuilt by a full re-seed — a fresh clone from the promoted instance. If the failed location is unreachable, or its storage is lost, the old primary has to be rebuilt as a new standby once that location is available again.

Rebuild the old primary before it serves traffic

After a failover the old primary may hold older data. Do not let it run as a writer again until it has rejoined as a standby of the promoted instance — otherwise you risk a split-brain with two writable copies. The platform does not fence it for you.

When you need stronger guarantees#

This pattern is deliberately simple: one writer, asynchronous replication, manual failover, and an explicit data-loss window. If any of the following are hard requirements, it is the wrong tool:

  • Automatic failover, with no manual promote.
  • A near-zero or zero data-loss guarantee (RPO ≈ 0).
  • Active reads and writes in more than one region.
  • Horizontal write scaling.
  • No manual operational steps during an incident.

For those workloads, use a distributed, PostgreSQL-compatible database instead. Ametnes is adding this as an option, built on YugabyteDB: consensus-based synchronous replication, automatic failover, and multi-region placement, while still speaking the PostgreSQL wire protocol so your application connects to a single endpoint unchanged. If your requirements match the list above, talk to us about it rather than using the primary/standby pattern.

Operational notes#

  • Replication is asynchronous; monitor lag before failing over.
  • A role change is a config update (dr.role/dr.uri), not a new resource — the same DNS name and endpoint are reused.
  • The replication password on the primary must match the password in the standby's dr.uri.
  • Only ever run one writable instance. The platform does not fence the old primary, so reconfigure or rebuild it as a standby before it accepts traffic.
  • Test your drill — including a hard stop of the primary — before you rely on it, so the RPO and failback behaviour are not a surprise during a real incident.