Q: How do you design disaster recovery and balance RTO/RPO targets against AWS costs?
Strategic framework for designing disaster recovery across the 4 AWS DR tiers (Backup & Restore, Pilot Light, Warm Standby, Multi-Site Active/Active) while balancing RTO/RPO requirements against cloud infrastructure costs.
#AWS #Disaster Recovery #RTO #RPO #FinOps #Cost Optimization #Backup
🎙️ Candidate Opening & Architectural Context
"I design disaster recovery by first categorizing systems by business criticality and establishing non-negotiable RTO (Recovery Time Objective) and RPO (Recovery Point Objective) targets. I select the most cost-effective DR pattern: Backup & Restore for non-critical services (hours RTO/RPO), Pilot Light for medium workloads, Warm Standby for mission-critical apps (minutes RTO), and Multi-Site Active/Active only when seconds count. For cost optimization, I ensure we do not over-provision standby infrastructure by leveraging IaC, automated snapshot lifecycles, and auto-scaling."
Advertisement
🛠️ Production Runbook & Step-by-Step Resolution
1️⃣
The 4 AWS Disaster Recovery Tiers & Cost Trade-Offs
Aligning technical architecture with business recovery objectives:
# Inspecting RDS automated cross-region snapshot copy and replication
aws rds describe-db-instances --db-instance-identifier prod-db \
--query 'DBInstances[*].[DBInstanceIdentifier,ReadReplicaDBInstanceIdentifiers]' --output json
# Verifying S3 cross-region replication configuration
aws s3api get-bucket-replication --bucket prod-primary-backups
- Backup & Restore (Lowest Cost): Data is continuously backed up to S3 with cross-region replication (CRR). Zero compute is running in the DR region until a disaster strikes (RTO: 4-24h, RPO: 1-24h).
- Pilot Light (Low Cost): Core data stores are continuously replicated (e.g. RDS Read Replica or DynamoDB Global Tables). Minimal core infra exists in DR; app compute is spun up via Terraform/ASGs upon failover (RTO: 10-60 min, RPO: minutes).
- Warm Standby (Medium/High Cost): A scaled-down, functional production replica runs 24/7 in the secondary region. During disaster, Route 53 switches traffic and the Auto Scaling Group scales out to 100% capacity (RTO: minutes, RPO: seconds).
- Multi-Site Active/Active (Highest Cost): Full production capacity operates concurrently across multiple regions with global load balancing (RTO: zero, RPO: zero/sub-second).
2️⃣
FinOps Guardrails & Standby Cost Controls
Prevent secondary DR environments from doubling your cloud bill unnecessarily:
# Inspect AWS Cost Explorer monthly spend and DR bucket lifecycle
aws ce get-cost-and-usage --time-period Start=2026-08-01,End=2026-09-01 \
--granularity MONTHLY --metrics UnblendedCost \
--group-by Type=DIMENSION,Key=SERVICE
# Check S3 lifecycle configuration for backup tiering to Glacier
aws s3api get-bucket-lifecycle-configuration --bucket prod-backup-vault
- Minimal Idle Compute: In Pilot Light / Warm Standby, keep EC2/EKS compute at minimum viable size (e.g., 2 small instances) and use Terraform to scale up only when triggered.
- Snapshot Lifecycle Policies: Transition older EBS snapshots and S3 backups to S3 Glacier Flexible / Deep Archive after 30 days.
- Automated Testing (Game Days): Run quarterly DR rehearsals to measure actual recovery time against RTO/RPO baselines and verify runbooks.
💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"Never default to expensive Active-Active DR. Classify workloads by business impact, select the appropriate tier (Backup & Restore vs Pilot Light vs Warm Standby), and keep standby compute minimal until failover."
⚡ 60-Second Elevator Pitch Talking Points
- Tier workloads by business impact to establish realistic RTO and RPO targets before designing infrastructure.
- Choose Pilot Light or Warm Standby to achieve sub-hour recovery times without paying for duplicate 24/7 compute fleets.
- Automate snapshot retention to S3 Glacier and conduct regular Game Day rehearsals to validate failover automation.
Advertisement