RabbitMQ disaster recovery#
Ametnes RabbitMQ replicates across clusters with the broker's bundled
shovel plugin. One cluster runs as the primary (the active broker) and a
second cluster runs as a standby whose shovel pulls a named queue from the
primary. The standby can be promoted when the primary is lost.
Quorum queues are in-cluster HA, not cross-cluster DR
Queues default to the quorum type (default_queue_type = quorum), which
Raft-replicates them across the nodes of a single cluster - that is the
in-cluster HA mechanism (a 3/5-node tier survives node loss automatically).
Quorum queues cannot span clusters, so replication to a standby in
another location uses the shovel link described here.
Why shovel and not federation#
RabbitMQ ships two cross-cluster plugins, federation and shovel, but only shovel is a fit for a warm standby. The difference is in when messages move:
| Shovel (what Ametnes uses) | Queue federation | |
|---|---|---|
| Message movement | Moves messages unconditionally | Moves messages only when a local consumer is waiting |
| Direction | Unidirectional (primary -> standby) | Needs consumers on both sides |
| Warm standby | Holds a copy with no consumer attached | Stays empty until an application connects |
Queue federation is consumer-driven: it exists to balance a single logical
queue across brokers by moving messages "where the consumers are". RabbitMQ's
own documentation states a federated queue "will only retrieve messages when it
has run out of messages locally, it has consumers that need messages, and
the upstream queue has 'spare' messages". It also cannot serve basic.get
polling when the local queue is empty.
A disaster-recovery standby has no application consumers sitting on it waiting for data, so federation would leave it empty until failover - which defeats the point. A shovel pulls the queue regardless of consumers, so the standby holds a real, current copy of the data. Federation remains a valid RabbitMQ feature for load balancing and blue/green migrations, but it is not exposed as a DR mode on this platform.
Model#
| Field | Meaning |
|---|---|
dr.role |
primary (active) or standby (pulls from the primary). |
dr.uri |
AMQPS URI of the peer, e.g. amqps://repl_user:secret@broker.example.com:5671/. Required for a standby. |
dr.queue |
Queue the shovel replicates. Defaults to ametnes.dr. |
dr.vhost |
Vhost for the shovel parameter (default /). |
dr.replication.user / dr.replication.password |
Credentials used to authenticate to the peer. |
Replication links are AMQPS — always pass an amqps:// peer URI so the
replication traffic is encrypted with the peer's zone certificate.
One queue per standby
A shovel moves a single named queue. To replicate more than one queue,
create the standby with the dr.queue you need, or run additional standbys.
Provision a primary and a standby#
Create the primary with dr.role = primary, then create a second cluster in
another location as dr.role = standby pointing at the primary's amqps
endpoint:
Declare and write to the queue on the primary; the standby's shovel pulls it and declares its own copy of the queue. There is no need to declare queues on the standby, and only the primary is operated by the application.
Fail over#
Promote the standby by setting its dr.role to primary. The shovel is
removed and the cluster becomes the active broker; repoint your clients at the
promoted cluster's amqps endpoint.
Promotion is restart-free. The definitions volume stays mounted in both roles - only its contents change (a standby carries the shovel parameter, a primary does not) - so the broker pod is not rolled and there is no extra outage window on top of the failover itself.
Fail back#
To return to the original primary, set its dr.role to standby, point
dr.uri at the promoted cluster's amqps endpoint, and set dr.queue to
the queue you now want replicated back. It establishes a shovel from the
promoted cluster, so data written there replicates to it.
Operational notes#
- A role change is a config update on the existing service; the endpoint name is reused.
- Replication is asynchronous — expect a small propagation delay.
- The replication credentials in
dr.urimust match those configured on the peer. - The shovel keeps running for the life of the standby; it never self-deletes after a queue is drained.