⚡ ~/naveed Interview Prep
⚡ Portfolio Home ✍️ Engineering Blog Deep Dives 🎯 Interview Hub 1,000+ Scenarios ☸️ Kubernetes Mastery Hub 24 Modules 🎮 DevOps Arcade & Quizzes Subnet Blitz ⚡ 🗺️ DevOps Roadmaps PDFs & Guides 🤖 Morpheus Analysis AI Quant ↗ 🛠️ Developer Tools Utilities 🧪 Labs & Experiments 📄 Interactive CV & Certs 🔗 All Links & Socials ⚡ Join The Dispatch (Weekly SRE Newsletter) →
← Back to All FinOps & System Design Interview Questions Scenario 93 of 98 in FinOps & System Design
Staff SRE / Streaming Architect System Design Distributed Data & Disaster Recovery SRE System Design

Q: Your mission-critical event streaming backbone processes 500,000 messages/sec on AWS MSK in us-east-1. A major regional fiber cut or AWS outage takes down the entire region. How do you design an active-passive cross-region disaster recovery architecture with MirrorMaker 2 that maintains continuous replication, automatically translates consumer group offsets, and enables consumer failover in under 60 seconds?

Architectural design for active-passive cross-region Apache Kafka disaster recovery replicating 500,000 msg/sec between AWS us-east-1 and us-west-2 with MirrorMaker 2, dynamic consumer offset translation, and sub-60s failover.

#System Design #Kafka #Disaster Recovery #MirrorMaker 2 #Multi-Region #SRE
🎙️ Candidate Opening & Architectural Context
"Simply replicating Kafka topic bytes to a secondary region is insufficient—consumer groups will fail because partition offsets on the target cluster do not match the source cluster. We designed an enterprise cross-region Kafka DR architecture utilizing MirrorMaker 2 and dynamic offset checkpoint translation."
Advertisement
⚡ Recommended Practice Lab

Want to master this scenario in a live sandbox? The Linux Foundation's FinOps Certified Practitioner (FOCP) Program covers this exact problem with hands-on terminal drills.

🛠️ Production Runbook & Step-by-Step Resolution

1️⃣

Deploy Primary and Disaster Recovery Kafka Topologies Across Regions

Establish dedicated clusters connected over high-speed private transit:

  • Clusters: Deployed Primary Cluster in us-east-1 (active producers and consumers) and Standby DR Cluster in us-west-2.
  • Network Transport: Routed replication traffic through AWS Transit Gateway Inter-Region Peering with KMS-managed in-transit encryption.
  • Sizing: Standby cluster is provisioned at identical broker count and storage capacity to immediately absorb 100% of write traffic upon failover.
Pro Tip: Pre-provisioning the DR cluster at full capacity eliminates the 15-minute delay of scaling up Kafka brokers during an emergency.
2️⃣

Configure MirrorMaker 2 (MM2) with Topic Mapping & Heartbeats

Replicate topics and measure end-to-end cross-region replication latency:

  • MirrorSourceConnector: Replicates records from Primary to Standby, preserving partition ordering and record timestamps.
  • Identity Replication: Configured replication.policy.class=IdentityReplicationPolicy to preserve exact topic names (e.g. orders stays orders instead of us-east-1.orders).
  • MirrorHeartbeatConnector: Emits heartbeats every 5 seconds to track end-to-end replication lag in Prometheus.
Pro Tip: IdentityReplicationPolicy ensures that consumer applications can switch clusters without modifying hardcoded subscribed topic names.
3️⃣

Automate Consumer Offset Synchronization via MirrorCheckpointConnector

Translate consumer group progress across differing partition offsets:

  • Checkpoint Connector: Periodically translates and writes consumer group offsets into the Standby cluster's __consumer_offsets topic.
  • Offset Mapping: Because message offsets differ between clusters, MM2 maps consumer progress based on record timestamps and logical sequence numbers.
  • Automated Migration: When a consumer group fails over to the standby cluster, it queries the translated offset and resumes processing from the exact correct message without replaying millions of historical records.
Pro Tip: Without automated offset translation, consumers failing over to a secondary cluster will either start from offset 0 (duplicating weeks of messages) or latest offset (losing data).
4️⃣

Execute 45-Second Disaster Failover & Measure RPO/RTO

Trigger automated producer and consumer rerouting during regional blackout:

  • DNS Switch: Updated Route 53 private hosted zone CNAME kafka.internal to point to the us-west-2 bootstrap servers.
  • Client Reconnection: Producers and consumers automatically disconnect from the dead region and establish TCP sessions with the standby cluster in 18 seconds.
  • Telemetry Results: Achieved RPO < 4 seconds (measured replication lag) and RTO of 42 seconds with zero uncommitted message corruption.
Pro Tip: Using a unified DNS CNAME allows thousands of client microservices to fail over seamlessly without requiring individual configuration updates or redeployments.
💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"Cross-region Kafka disaster recovery requires MirrorMaker 2 with IdentityReplicationPolicy, MirrorCheckpointConnector for automatic offset translation, and DNS abstraction for sub-60-second client failover."
⚡ 60-Second Elevator Pitch Talking Points
  • Deploy identical Kafka clusters across paired AWS regions connected via Transit Gateway.
  • Replicate topics using MirrorMaker 2 with IdentityReplicationPolicy to preserve topic names.
  • Use MirrorCheckpointConnector to automatically translate and commit consumer group offsets.
  • Execute failover in under 45 seconds via Route 53 DNS switching with RPO < 4 seconds.
Advertisement
Want more FinOps & System Design scenarios?
Explore our complete collection of scenario-based FinOps & System Design interview runbooks.
Browse All FinOps & System Design Questions →