⚡ ~/naveed Interview Prep
⚡ Portfolio Home ✍️ Engineering Blog Deep Dives 🎯 Interview Hub 1,000+ Scenarios ☸️ Kubernetes Mastery Hub 24 Modules 🎮 DevOps Arcade & Quizzes Subnet Blitz ⚡ 🗺️ DevOps Roadmaps PDFs & Guides 🤖 Morpheus Analysis AI Quant ↗ 🛠️ Developer Tools Utilities 🧪 Labs & Experiments 📄 Interactive CV & Certs 🔗 All Links & Socials ⚡ Join The Dispatch (Weekly SRE Newsletter) →
← Back to All AWS & Cloud Architecture Interview Questions Scenario 168 of 186 in AWS & Cloud Architecture
Staff Cloud Architect Azure & Cloud Database Architecture & Disaster Recovery Disaster Recovery

Q: Your Tier-1 banking transactional database requires 100 TB of storage capacity, rapid scaling up to 80 vCores, and a strict DR requirement of RPO < 5 seconds and RTO < 2 minutes across paired Azure regions. How do you design and operate Azure SQL Hyperscale with Auto-Failover Groups?

Architectural design for achieving near-zero RPO and RTO < 5 minutes on Azure SQL Hyperscale using Auto-Failover Groups, geo-replication, and read-scale replica offloading.

#Azure #Azure SQL #Hyperscale #Auto-Failover Groups #Disaster Recovery #High Availability
🎙️ Candidate Opening & Architectural Context
"Standard Azure SQL tiers hit storage limits at 4 TB and suffered multi-hour restore times. We re-platformed our core ledger database to Azure SQL Hyperscale, leveraging tiered page servers and active Auto-Failover Groups."
Advertisement
⚡ Recommended Practice Lab

Want to master this scenario in a live sandbox? Stephane Maarek's AWS Certified DevOps Engineer Professional Masterclass on Udemy covers this exact problem with hands-on terminal drills.

🛠️ Production Runbook & Step-by-Step Resolution

1️⃣

Provision Azure SQL Database Hyperscale with High Availability Replicas

Deploy tiered compute and storage architecture supporting up to 100 TB:

  • Provision Hyperscale: Executed az sql db create -g rg-data -s sql-primary-eastus -n core-ledger --edition Hyperscale --family Gen5 --capacity 32 --ha-replicas 2 --zone-redundant true.
  • Storage Architecture: Hyperscale separates compute nodes from storage page servers, enabling near-instant database scaling and snapshots regardless of data size.
Pro Tip: Configuring 2 High-Availability (HA) replicas and zone redundancy provides 99.99% availability within the primary region.
2️⃣

Configure Cross-Region Auto-Failover Group

Establish continuous asynchronous replication to a paired secondary region:

  • Create Failover Group: Created failover group linking East US (primary) to Central US (secondary): az sql failover-group create -g rg-data --server sql-primary-eastus --name fg-ledger --partner-server sql-dr-centralus --failover-policy Automatic --grace-period 1.
  • Failover Policy & Grace Period: Set grace period to 1 hour for automated failover, preventing premature split-brain during transient network blips.
Pro Tip: Auto-Failover Groups provide read-write and read-only listener FQDNs that remain static during failovers, eliminating client connection string changes.
3️⃣

Offload Heavy Reporting Queries to Read-Scale Replicas

Segregate transactional OLTP traffic from analytical reporting traffic:

  • Read-Only Listener: Configured BI dashboards and analytics services to connect to fg-ledger.secondary.database.windows.net.
  • Zero Resource Contention: Hyperscale secondary replicas utilize dedicated compute nodes, preventing analytical queries from consuming write transaction log throughput.
Pro Tip: Connecting with ApplicationIntent=ReadOnly automatically routes queries to free secondary replicas, improving primary write throughput.
4️⃣

Execute Disaster Recovery Drill via Forced & Graceful Failover

Validate RPO and RTO compliance under simulated primary region outage conditions:

  • Planned Failover Drill: Executed az sql failover-group set-primary -g rg-data --server sql-dr-centralus --name fg-ledger.
  • Metrics Verification: Confirmed zero data loss (RPO = 0s) and listener redirection completed in 34 seconds (RTO well below 2-minute SLA).
Pro Tip: Planned failovers perform a clean synchronization before swapping primary/secondary roles, guaranteeing zero data loss.
💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"Azure SQL Hyperscale decouples compute and storage to scale up to 100 TB seamlessly, while Auto-Failover Groups provide continuous cross-region replication and rapid failover under static listener endpoints."
⚡ 60-Second Elevator Pitch Talking Points
  • Deploy Azure SQL Hyperscale with 2 zone-redundant HA replicas for intra-region resilience.
  • Establish an Auto-Failover Group with Central US for cross-region disaster recovery.
  • Route heavy BI reporting to read-scale listeners to protect primary write performance.
  • Validate RPO < 5s and RTO < 2m through regular automated planned failover drills.
Advertisement
Want more AWS & Cloud Architecture scenarios?
Explore our complete collection of scenario-based AWS & Cloud Architecture interview runbooks.
Browse All AWS & Cloud Architecture Questions →