⚡ ~/naveed Interview Prep
⚡ Portfolio Home ✍️ Engineering Blog Deep Dives 🎯 Interview Hub 1,000+ Scenarios ☸️ Kubernetes Mastery Hub 24 Modules 🎮 DevOps Arcade & Quizzes Subnet Blitz ⚡ 🗺️ DevOps Roadmaps PDFs & Guides 🤖 Morpheus Analysis AI Quant ↗ 🛠️ Developer Tools Utilities 🧪 Labs & Experiments 📄 Interactive CV & Certs 🔗 All Links & Socials ⚡ Join The Dispatch (Weekly SRE Newsletter) →
← Back to All AWS & Cloud Architecture Interview Questions Scenario 183 of 186 in AWS & Cloud Architecture
Staff Cloud Architect Multi-Cloud Disaster Recovery & Database SRE Disaster Recovery

Q: To achieve extreme resilience against catastrophic single-cloud regional outages, your company requires a warm standby disaster recovery database in GCP that continuously replicates from primary AWS RDS PostgreSQL with RPO < 30 seconds and RTO < 5 minutes. How do you design and orchestrate this cross-cloud database failover?

Engineering a resilient cross-cloud disaster recovery architecture replicating transactional data from AWS RDS PostgreSQL to GCP Cloud SQL with automated failover orchestration and DNS failover.

#Multi-Cloud #Disaster Recovery #PostgreSQL #AWS #GCP #Debezium #Failover
🎙️ Candidate Opening & Architectural Context
"Relying on a single cloud provider leaves businesses vulnerable to global IAM or control plane outages. We engineered an active-passive cross-cloud disaster recovery pipeline replicating AWS RDS PostgreSQL to GCP Cloud SQL using native PostgreSQL logical replication over private transit."
Advertisement
⚡ Recommended Practice Lab

Want to master this scenario in a live sandbox? Stephane Maarek's AWS Certified DevOps Engineer Professional Masterclass on Udemy covers this exact problem with hands-on terminal drills.

🛠️ Production Runbook & Step-by-Step Resolution

1️⃣

Configure Native PostgreSQL Logical Replication across Clouds

Establish continuous change data capture without locking tables:

  • Primary RDS Settings: Set rds.logical_replication=1 in AWS DB parameter group; created publication: CREATE PUBLICATION crosscloud_pub FOR ALL TABLES;.
  • Standby Cloud SQL: Provisioned GCP Cloud SQL PostgreSQL instance and configured subscription: CREATE SUBSCRIPTION crosscloud_sub CONNECTION 'host=aws-db.internal port=5432 dbname=prod user=rep_user sslmode=verify-full' PUBLICATION crosscloud_pub;.
Pro Tip: Logical replication transmits individual DML row changes rather than physical WAL disk blocks, allowing replication across different cloud database platforms and minor engine versions.
2️⃣

Secure Private Transit & Monitor Replication Lag in Prometheus

Route replication traffic through private cross-cloud peering and track RPO metrics:

  • Encrypted Transport: Routed database traffic strictly through private Megaport interconnect with enforced TLS 1.3 certificate validation.
  • Replication Lag Telemetry: Tracked pg_stat_replication.write_lag and byte difference in Prometheus; alerted in PagerDuty if replication lag exceeded 15 seconds.
Pro Tip: Monitoring byte-level replication lag ensures that the disaster recovery target satisfies the 30-second RPO SLA under peak transactional load.
3️⃣

Execute Automated Failover Orchestration during Cloud Blackout

Trigger automated promotion and application traffic redirection:

  • Sever Subscription: Executed promotion script: ALTER SUBSCRIPTION crosscloud_sub DISABLE; ALTER SUBSCRIPTION crosscloud_sub SET (slot_name = NONE); DROP SUBSCRIPTION crosscloud_sub;.
  • Sequence Alignment: Ran automated sequence synchronization script ensuring auto-increment primary keys are aligned with current max IDs.
Pro Tip: PostgreSQL logical replication does not automatically replicate sequence state (nextval); synchronizing sequences upon promotion is critical to prevent primary key collision errors.
4️⃣

Switch Global Traffic via Anycast DNS & Standby Compute Activation

Reroute user traffic to standby microservices deployed on Google Cloud:

  • Scale Standby Compute: Scaled GKE standby deployments from minimum warm replicas (2) to full production scale (50) in 45 seconds.
  • DNS Failover: Cloudflare Anycast DNS health check detected AWS failure and flipped API endpoint CNAME to GCP External Global Load Balancer.
  • Total Outage Duration: Complete recovery achieved in 3 minutes 12 seconds with zero uncommitted transactional data lost.
Pro Tip: Maintaining warm standby compute instances in the secondary cloud ensures instantaneous traffic absorption upon database promotion.
💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"Cross-cloud database disaster recovery combines PostgreSQL native logical replication over private interconnect with automated sequence alignment and Anycast DNS failover to satisfy RPO < 30s and RTO < 5m."
⚡ 60-Second Elevator Pitch Talking Points
  • Enable PostgreSQL logical replication from AWS RDS to GCP Cloud SQL over private interconnect.
  • Monitor pg_stat_replication write lag in Prometheus to guarantee RPO < 30 seconds.
  • Automate subscription promotion and sequence synchronization during disaster failover.
  • Execute seamless Anycast DNS traffic redirection to secondary cloud compute in under 4 minutes.
Advertisement
Want more AWS & Cloud Architecture scenarios?
Explore our complete collection of scenario-based AWS & Cloud Architecture interview runbooks.
Browse All AWS & Cloud Architecture Questions →