Skip to content

RabbitMQ disaster recovery#

Ametnes RabbitMQ replicates across clusters with the broker's bundled shovel plugin. One cluster runs as the primary (the active broker) and a second cluster runs as a standby whose shovel pulls a named queue from the primary. The standby can be promoted when the primary is lost.

Quorum queues are in-cluster HA, not cross-cluster DR

Queues default to the quorum type (default_queue_type = quorum), which Raft-replicates them across the nodes of a single cluster - that is the in-cluster HA mechanism (a 3/5-node tier survives node loss automatically). Quorum queues cannot span clusters, so replication to a standby in another location uses the shovel link described here.

Why shovel and not federation#

RabbitMQ ships two cross-cluster plugins, federation and shovel, but only shovel is a fit for a warm standby. The difference is in when messages move:

Shovel (what Ametnes uses) Queue federation
Message movement Moves messages unconditionally Moves messages only when a local consumer is waiting
Direction Unidirectional (primary -> standby) Needs consumers on both sides
Warm standby Holds a copy with no consumer attached Stays empty until an application connects

Queue federation is consumer-driven: it exists to balance a single logical queue across brokers by moving messages "where the consumers are". RabbitMQ's own documentation states a federated queue "will only retrieve messages when it has run out of messages locally, it has consumers that need messages, and the upstream queue has 'spare' messages". It also cannot serve basic.get polling when the local queue is empty.

A disaster-recovery standby has no application consumers sitting on it waiting for data, so federation would leave it empty until failover - which defeats the point. A shovel pulls the queue regardless of consumers, so the standby holds a real, current copy of the data. Federation remains a valid RabbitMQ feature for load balancing and blue/green migrations, but it is not exposed as a DR mode on this platform.

Model#

Field Meaning
dr.role primary (active) or standby (pulls from the primary).
dr.uri AMQPS URI of the peer, e.g. amqps://repl_user:secret@broker.example.com:5671/. Required for a standby.
dr.queue Queue the shovel replicates. Defaults to ametnes.dr.
dr.vhost Vhost for the shovel parameter (default /).
dr.replication.user / dr.replication.password Credentials used to authenticate to the peer.

Replication links are AMQPS — always pass an amqps:// peer URI so the replication traffic is encrypted with the peer's zone certificate.

One queue per standby

A shovel moves a single named queue. To replicate more than one queue, create the standby with the dr.queue you need, or run additional standbys.

Provision a primary and a standby#

Create the primary with dr.role = primary, then create a second cluster in another location as dr.role = standby pointing at the primary's amqps endpoint:

rabbitmq-dr.tf
resource "ametnes_service" "rabbitmq_primary" {
  name     = "RabbitMQPrimary"
  project  = ametnes_project.project.id
  location = data.ametnes_location.location_a.id
  kind     = "service/rabbitmq:4.3"
  nodes    = 3
  capacity { storage = 100 }
  config = {
    "architecture"   = "Medium"
    "admin.user"     = "admin"
    "admin.password" = "change-me"
    "dr.role"        = "primary"
  }
}

resource "ametnes_service" "rabbitmq_standby" {
  name     = "RabbitMQStandby"
  project  = ametnes_project.project.id
  location = data.ametnes_location.location_b.id
  kind     = "service/rabbitmq:4.3"
  nodes    = 3
  capacity { storage = 100 }
  config = {
    "architecture"            = "Medium"
    "admin.user"              = "admin"
    "admin.password"          = "change-me"
    "dr.role"                 = "standby"
    "dr.queue"                = "orders"
    "dr.uri"                  = "amqps://repl_user:secret@<primary-host>:<primary-port>/"
    "dr.replication.user"     = "repl_user"
    "dr.replication.password" = "secret"
  }
}

Declare and write to the queue on the primary; the standby's shovel pulls it and declares its own copy of the queue. There is no need to declare queues on the standby, and only the primary is operated by the application.

Fail over#

Promote the standby by setting its dr.role to primary. The shovel is removed and the cluster becomes the active broker; repoint your clients at the promoted cluster's amqps endpoint.

Promotion is restart-free. The definitions volume stays mounted in both roles - only its contents change (a standby carries the shovel parameter, a primary does not) - so the broker pod is not rolled and there is no extra outage window on top of the failover itself.

Fail back#

To return to the original primary, set its dr.role to standby, point dr.uri at the promoted cluster's amqps endpoint, and set dr.queue to the queue you now want replicated back. It establishes a shovel from the promoted cluster, so data written there replicates to it.

Operational notes#

  • A role change is a config update on the existing service; the endpoint name is reused.
  • Replication is asynchronous — expect a small propagation delay.
  • The replication credentials in dr.uri must match those configured on the peer.
  • The shovel keeps running for the life of the standby; it never self-deletes after a queue is drained.