Q: Your mission-critical event streaming backbone processes 500,000 messages/sec on AWS MSK in us-east-1. A major regional fiber cut or AWS outage takes down the entire region. How do you design an active-passive cross-region disaster recovery architecture with MirrorMaker 2 that maintains continuous replication, automatically translates consumer group offsets, and enables consumer failover in under 60 seconds?
Architectural design for active-passive cross-region Apache Kafka disaster recovery replicating 500,000 msg/sec between AWS us-east-1 and us-west-2 with MirrorMaker 2, dynamic consumer offset translation, and sub-60s failover.
Want to master this scenario in a live sandbox? The Linux Foundation's FinOps Certified Practitioner (FOCP) Program covers this exact problem with hands-on terminal drills.
🛠️ Production Runbook & Step-by-Step Resolution
Deploy Primary and Disaster Recovery Kafka Topologies Across Regions
Establish dedicated clusters connected over high-speed private transit:
- Clusters: Deployed
Primary Clusterin us-east-1 (active producers and consumers) andStandby DR Clusterin us-west-2. - Network Transport: Routed replication traffic through AWS Transit Gateway Inter-Region Peering with KMS-managed in-transit encryption.
- Sizing: Standby cluster is provisioned at identical broker count and storage capacity to immediately absorb 100% of write traffic upon failover.
Configure MirrorMaker 2 (MM2) with Topic Mapping & Heartbeats
Replicate topics and measure end-to-end cross-region replication latency:
- MirrorSourceConnector: Replicates records from Primary to Standby, preserving partition ordering and record timestamps.
- Identity Replication: Configured
replication.policy.class=IdentityReplicationPolicyto preserve exact topic names (e.g.ordersstaysordersinstead ofus-east-1.orders). - MirrorHeartbeatConnector: Emits heartbeats every 5 seconds to track end-to-end replication lag in Prometheus.
Automate Consumer Offset Synchronization via MirrorCheckpointConnector
Translate consumer group progress across differing partition offsets:
- Checkpoint Connector: Periodically translates and writes consumer group offsets into the Standby cluster's
__consumer_offsetstopic. - Offset Mapping: Because message offsets differ between clusters, MM2 maps consumer progress based on record timestamps and logical sequence numbers.
- Automated Migration: When a consumer group fails over to the standby cluster, it queries the translated offset and resumes processing from the exact correct message without replaying millions of historical records.
Execute 45-Second Disaster Failover & Measure RPO/RTO
Trigger automated producer and consumer rerouting during regional blackout:
- DNS Switch: Updated Route 53 private hosted zone CNAME
kafka.internalto point to the us-west-2 bootstrap servers. - Client Reconnection: Producers and consumers automatically disconnect from the dead region and establish TCP sessions with the standby cluster in 18 seconds.
- Telemetry Results: Achieved RPO < 4 seconds (measured replication lag) and RTO of 42 seconds with zero uncommitted message corruption.
- Deploy identical Kafka clusters across paired AWS regions connected via Transit Gateway.
- Replicate topics using MirrorMaker 2 with IdentityReplicationPolicy to preserve topic names.
- Use MirrorCheckpointConnector to automatically translate and commit consumer group offsets.
- Execute failover in under 45 seconds via Route 53 DNS switching with RPO < 4 seconds.