⚡ ~/naveed Interview Prep
⚡ Portfolio Home ✍️ Engineering Blog Deep Dives 🎯 Interview Hub 998+ Scenarios ☸️ Kubernetes Mastery Hub 24 Modules 🎮 DevOps Arcade & Quizzes Subnet Blitz ⚡ 🗺️ DevOps Roadmaps PDFs & Guides 🤖 Morpheus Analysis AI Quant ↗ 🛠️ Developer Tools Utilities 🧪 Labs & Experiments 📄 Interactive CV & Certs 🔗 All Links & Socials ⚡ Join The Dispatch (Weekly SRE Newsletter) →
Staff SRE / Principal Architect [L3] General DevOps General DevOps — Scenario-Based Interview Questions Staff SRE Scenario [L3]

Q: Your company’s primary region (`us-east-1`) completely goes offline (a major AWS outage affecting the whole region). You are tasked with deciding when to trigger the Disaster Recovery failover to `us-west-2`. What metrics/business factors dictate this decision?

A regional failover is the most destructive action an SRE can take; it risks massive data loss, split-brain scenarios, and hours of compl...

#General DevOps #General DevOps — Scenario-Based Interview Questions #L3 #DevOps #SRE #Architecture
🎙️ Candidate Opening & Architectural Context
""We faced this organizational and technical challenge while scaling our engineering teams. The interviewer is testing: RTO, RPO, split-brain, cost of failover vs cost of downtime.. I structure my answer around systematic triage first, root cause analysis second, and permanent remediation third.""
Advertisement

🛠️ Production Runbook & Step-by-Step Resolution

1️⃣

Initial Diagnostics & Root Cause Analysis

A regional failover is the most destructive action an SRE can take; it risks massive data loss, split-brain scenarios, and hours of complex resolution. It should not be triggered lightly.

  • Is the outage estimated to last longer than our business-defined RTO? (e.g., if RTO is 4 hours, and AWS says it's an hour fix, we wait. If AWS says ETA unknown, we fail over).
  • What is the state of the cross-region database replication? If the replication lag was 2 hours behind when the region died, failing over now guarantees 2 hours of permanent data loss (violating a strict 1-hour RPO).
  • We must ensure the us-east-1 environment is completely fenced off (DNS isolated) to prevent a "split-brain" where systems come back online and start writing conflicting data to the old primary database while the new us-west-2 primary is active.
2️⃣

Remediation & Permanent Safeguards

The decision is governed by RTO (Recovery Time Objective) and RPO (Recovery Point Objective):

💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"Pro-Tip: Is the outage estimated to last longer than our business-defined RTO? (e.g., if RTO is 4 hours, and AWS says it's an hour fix, we ."
⚡ 60-Second Elevator Pitch Talking Points
  • Is the outage estimated to last longer than our business-defined RTO? (e.g., if RTO is 4 hours, a...
  • What is the state of the cross-region database replication? If the replication lag was 2 hours be...
  • We must ensure the us-east-1 environment is completely fenced off (DNS isolated) to prevent a "sp...
Advertisement
Want more General DevOps scenarios?
Explore our complete collection of scenario-based General DevOps interview runbooks.
Browse All General DevOps Questions →

📚 Related Production Scenarios in General DevOps