⚡ ~/naveed Interview Prep
⚡ Portfolio Home ✍️ Engineering Blog Deep Dives 🎯 Interview Hub 1,000+ Scenarios ☸️ Kubernetes Mastery Hub 24 Modules 🎮 DevOps Arcade & Quizzes Subnet Blitz ⚡ 🗺️ DevOps Roadmaps PDFs & Guides 🤖 Morpheus Analysis AI Quant ↗ 🛠️ Developer Tools Utilities 🧪 Labs & Experiments 📄 Interactive CV & Certs 🔗 All Links & Socials ⚡ Join The Dispatch (Weekly SRE Newsletter) →
← Back to All AWS & Cloud Architecture Interview Questions Scenario 171 of 171 in AWS & Cloud Architecture
Senior DevOps / SRE AWS Disaster Recovery & High Availability Architecture & Design

Q: What are RTO and RPO in AWS cloud architecture? Explain the 4 primary AWS disaster recovery strategies, their cost trade-offs, and how you architect systems to achieve sub-minute recovery.

Architectural breakdown of RTO and RPO in AWS: comparing Backup & Restore, Pilot Light, Warm Standby, and Multi-Region Active-Active across latency, replication, and cost.

#rto and rpo in aws #RTO #RPO #AWS Disaster Recovery #Multi-Region #High Availability #Cloud Architecture #Route 53 #Aurora Global Database
🎙️ Candidate Opening & Architectural Context
"Every disaster recovery (DR) architecture is defined by two mathematical metrics: Recovery Time Objective (RTO) and Recovery Point Objective (RPO). In AWS, as RTO and RPO approach zero, architectural complexity and cloud costs increase exponentially."
Advertisement
⚡ Recommended Practice Lab

Want to master this scenario in a live sandbox? Stephane Maarek's AWS Certified DevOps Engineer Professional Masterclass on Udemy covers this exact problem with hands-on terminal drills.

🛠️ Production Runbook & Step-by-Step Resolution

1️⃣

Defining RTO and RPO

The foundational disaster recovery metrics:

  • RTO (Recovery Time Objective): The maximum acceptable duration of downtime before application service must be restored after a disaster. (How long can the business afford to be offline?).
  • RPO (Recovery Point Objective): The maximum acceptable data loss measured in time between the disaster and the latest recoverable backup. (How much data can the business afford to permanently lose?).
Pro Tip: If your database took a backup at 2:00 AM and a disaster occurs at 2:45 AM, recovering to the 2:00 AM snapshot means an RPO of 45 minutes of lost transactions.
2️⃣

The 4 AWS Disaster Recovery Strategies

Ranging from lowest cost to lowest downtime:

  • 1. Backup and Restore (Hours RTO, Hours RPO): Regular snapshots stored in Amazon S3/AWS Backup with Cross-Region Replication. Cheapest strategy, but spinning up compute from scratch in a secondary region takes hours.
  • 2. Pilot Light (10s of minutes RTO, Minutes RPO): Critical core data is continuously replicated to the secondary region (e.g., Amazon Aurora read replica). Compute templates (AMIs, Terraform) are ready, but EC2 instances/EKS node groups are scaled to 0 until failover.
  • 3. Warm Standby (Minutes RTO, Seconds RPO): A scaled-down, functional duplicate of the system runs continuously in the secondary region. During disaster, Route 53 redirects traffic while Auto Scaling groups rapidly scale compute up to 100%.
  • 4. Multi-Region Active-Active (Near-zero RTO, Near-zero RPO): Full production workloads run simultaneously in two or more AWS regions. Route 53 latency/geo-routing distributes live user requests. Amazon Aurora Global Database or DynamoDB Global Tables provide sub-second cross-region replication.
Advertisement
3️⃣

Key AWS Services Powering Low RTO/RPO

Technologies enabling enterprise DR automation:

  • Amazon Route 53 Application Recovery Controller (ARC): Automated health checks, DNS failover, and routing control cells to isolate regional failures.
  • Amazon Aurora Global Database: Storage-level cross-region physical replication with typical latency < 1 second and zero performance impact on the primary cluster; promote secondary in < 1 minute.
  • DynamoDB Global Tables: Fully managed multi-master active-active database with automatic multi-region replication.
  • AWS Backup: Centrally managed cross-region and cross-account immutable vault backups (protected against ransomware).
4️⃣

Automated Failover Architecture

Failover workflow for Warm Standby or Pilot Light:

  • Conduct quarterly 'GameDay' disaster simulations in staging environments to verify that automation works without manual intervention.
💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"RTO is acceptable downtime; RPO is acceptable data loss. AWS offers 4 disaster recovery tiers: Backup & Restore, Pilot Light, Warm Standby, and Multi-Region Active-Active. Achieving near-zero RTO/RPO requires active data replication (Aurora Global DB) and automated DNS routing."
⚡ 60-Second Elevator Pitch Talking Points
  • RTO defines the maximum acceptable duration of service downtime, while RPO defines the maximum acceptable data loss measured in time.
  • In AWS, we choose between four architectures based on cost and SLA: Backup & Restore for non-critical workloads, Pilot Light or Warm Standby for standard production, and Multi-Region Active-Active for tier-0 revenue-critical services.
  • For sub-minute RTO and RPO, we leverage Amazon Aurora Global Database for storage-layer replication and Route 53 ARC for rapid automated DNS traffic shift.
Advertisement
Want more AWS & Cloud Architecture scenarios?
Explore our complete collection of scenario-based AWS & Cloud Architecture interview runbooks.
Browse All AWS & Cloud Architecture Questions →