⚡ ~/naveed Interview Prep
⚡ Portfolio Home ✍️ Engineering Blog Deep Dives 🎯 Interview Hub 998+ Scenarios ☸️ Kubernetes Mastery Hub 24 Modules 🎮 DevOps Arcade & Quizzes Subnet Blitz ⚡ 🗺️ DevOps Roadmaps PDFs & Guides 🤖 Morpheus Analysis AI Quant ↗ 🛠️ Developer Tools Utilities 🧪 Labs & Experiments 📄 Interactive CV & Certs 🔗 All Links & Socials ⚡ Join The Dispatch (Weekly SRE Newsletter) →
Staff SRE / Principal Architect [L3] AWS RDS & Databases Staff SRE Scenario [L3]

Q: An RDS instance went down and the automated failover to the standby didn't happen as expected in a Multi-AZ setup. What could have gone wrong?

Multi-AZ failover is automatic, but several things can prevent or delay it:

#AWS #RDS & Databases #L3 #Cloud #Infrastructure
🎙️ Candidate Opening & Architectural Context
""In our AWS cloud environment, we managed high-traffic microservices where this exact scenario occurred. When addressing this question, I walk the interviewer through our production incident runbook: isolating the blast radius, checking diagnostic logs and metrics, and applying a safe fix.""
Advertisement

🛠️ Production Runbook & Step-by-Step Resolution

1️⃣

Initial Diagnostics & Root Cause Analysis

Multi-AZ failover is automatic, but several things can prevent or delay it:

  • Maintenance mode — if you disabled automatic failover before maintenance.
  • Failover trigger conditions — RDS only fails over for: storage failure, loss of network connectivity, instance hardware failure, running out of storage. It does NOT auto-failover for high CPU/load.
  • DNS propagation delay — failover promotes the standby, but the endpoint DNS (e.g., mydb.xxxxxx.us-east-1.rds.amazonaws.com) needs to update. If the app uses hardcoded IPs instead of the RDS DNS endpoint, failover doesn't help. Always use DNS endpoint.
2️⃣

Remediation & Permanent Safeguards

  • Application doesn't retry connections — after failover, existing connections are dropped. The app must retry DB connections. Apps that don't retry see prolonged outage.
  • Replication lag — in some edge cases, the standby might have a small lag, causing momentary data inconsistency.
💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"Pro-Tip: Maintenance mode — if you disabled automatic failover before maintenance.."
⚡ 60-Second Elevator Pitch Talking Points
  • Maintenance mode — if you disabled automatic failover before maintenance.
  • Failover trigger conditions — RDS only fails over for: storage failure, loss of network connectiv...
  • DNS propagation delay — failover promotes the standby, but the endpoint DNS (e.g., mydb.xxxxxx.us...
Advertisement
Want more AWS scenarios?
Explore our complete collection of scenario-based AWS interview runbooks.
Browse All AWS Questions →

📚 Related Production Scenarios in AWS