⚡ ~/naveed Interview Prep
⚡ Portfolio Home ✍️ Engineering Blog Deep Dives 🎯 Interview Hub 998+ Scenarios ☸️ Kubernetes Mastery Hub 24 Modules 🎮 DevOps Arcade & Quizzes Subnet Blitz ⚡ 🗺️ DevOps Roadmaps PDFs & Guides 🤖 Morpheus Analysis AI Quant ↗ 🛠️ Developer Tools Utilities 🧪 Labs & Experiments 📄 Interactive CV & Certs 🔗 All Links & Socials ⚡ Join The Dispatch (Weekly SRE Newsletter) →
Staff SRE / Principal Architect [L3] AWS Cost & Architecture Staff SRE Scenario [L3]

Q: How do you design disaster recovery and balance RTO/RPO targets against AWS costs?

Strategic framework for designing disaster recovery across the 4 AWS DR tiers (Backup & Restore, Pilot Light, Warm Standby, Multi-Site Active/Active) while balancing RTO/RPO requirements against cloud infrastructure costs.

#AWS #Disaster Recovery #RTO #RPO #FinOps #Cost Optimization #Backup
🎙️ Candidate Opening & Architectural Context
"I design disaster recovery by first categorizing systems by business criticality and establishing non-negotiable RTO (Recovery Time Objective) and RPO (Recovery Point Objective) targets. I select the most cost-effective DR pattern: Backup & Restore for non-critical services (hours RTO/RPO), Pilot Light for medium workloads, Warm Standby for mission-critical apps (minutes RTO), and Multi-Site Active/Active only when seconds count. For cost optimization, I ensure we do not over-provision standby infrastructure by leveraging IaC, automated snapshot lifecycles, and auto-scaling."
Advertisement

🛠️ Production Runbook & Step-by-Step Resolution

1️⃣

The 4 AWS Disaster Recovery Tiers & Cost Trade-Offs

Aligning technical architecture with business recovery objectives:

# Inspecting RDS automated cross-region snapshot copy and replication
aws rds describe-db-instances --db-instance-identifier prod-db \
  --query 'DBInstances[*].[DBInstanceIdentifier,ReadReplicaDBInstanceIdentifiers]' --output json

# Verifying S3 cross-region replication configuration
aws s3api get-bucket-replication --bucket prod-primary-backups
  • Backup & Restore (Lowest Cost): Data is continuously backed up to S3 with cross-region replication (CRR). Zero compute is running in the DR region until a disaster strikes (RTO: 4-24h, RPO: 1-24h).
  • Pilot Light (Low Cost): Core data stores are continuously replicated (e.g. RDS Read Replica or DynamoDB Global Tables). Minimal core infra exists in DR; app compute is spun up via Terraform/ASGs upon failover (RTO: 10-60 min, RPO: minutes).
  • Warm Standby (Medium/High Cost): A scaled-down, functional production replica runs 24/7 in the secondary region. During disaster, Route 53 switches traffic and the Auto Scaling Group scales out to 100% capacity (RTO: minutes, RPO: seconds).
  • Multi-Site Active/Active (Highest Cost): Full production capacity operates concurrently across multiple regions with global load balancing (RTO: zero, RPO: zero/sub-second).
2️⃣

FinOps Guardrails & Standby Cost Controls

Prevent secondary DR environments from doubling your cloud bill unnecessarily:

# Inspect AWS Cost Explorer monthly spend and DR bucket lifecycle
aws ce get-cost-and-usage --time-period Start=2026-08-01,End=2026-09-01 \
  --granularity MONTHLY --metrics UnblendedCost \
  --group-by Type=DIMENSION,Key=SERVICE

# Check S3 lifecycle configuration for backup tiering to Glacier
aws s3api get-bucket-lifecycle-configuration --bucket prod-backup-vault
  • Minimal Idle Compute: In Pilot Light / Warm Standby, keep EC2/EKS compute at minimum viable size (e.g., 2 small instances) and use Terraform to scale up only when triggered.
  • Snapshot Lifecycle Policies: Transition older EBS snapshots and S3 backups to S3 Glacier Flexible / Deep Archive after 30 days.
  • Automated Testing (Game Days): Run quarterly DR rehearsals to measure actual recovery time against RTO/RPO baselines and verify runbooks.
💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"Never default to expensive Active-Active DR. Classify workloads by business impact, select the appropriate tier (Backup & Restore vs Pilot Light vs Warm Standby), and keep standby compute minimal until failover."
⚡ 60-Second Elevator Pitch Talking Points
  • Tier workloads by business impact to establish realistic RTO and RPO targets before designing infrastructure.
  • Choose Pilot Light or Warm Standby to achieve sub-hour recovery times without paying for duplicate 24/7 compute fleets.
  • Automate snapshot retention to S3 Glacier and conduct regular Game Day rehearsals to validate failover automation.
Advertisement
Want more AWS scenarios?
Explore our complete collection of scenario-based AWS interview runbooks.
Browse All AWS Questions →

📚 Related Production Scenarios in AWS