⚡ ~/naveed Interview Prep
⚡ Portfolio Home ✍️ Engineering Blog Deep Dives 🎯 Interview Hub 1,000+ Scenarios ☸️ Kubernetes Mastery Hub 24 Modules 🎮 DevOps Arcade & Quizzes Subnet Blitz ⚡ 🗺️ DevOps Roadmaps PDFs & Guides 🤖 Morpheus Analysis AI Quant ↗ 🛠️ Developer Tools Utilities 🧪 Labs & Experiments 📄 Interactive CV & Certs 🔗 All Links & Socials ⚡ Join The Dispatch (Weekly SRE Newsletter) →
← Back to All Databases & Storage Interview Questions Scenario 52 of 52 in Databases & Storage
Senior SRE / Infrastructure Engineer Database Redis & In-Memory Caching High Availability

Q: During a transient network partition between AWS Availability Zones, a Redis master node is isolated from its Sentinel monitors and replica instances. The Sentinels detect quorum loss, declare the master dead, and promote a replica in AZ-b to become the new master. However, client applications in AZ-a still have network access to the old partitioned master and continue executing writes for 90 seconds. When the network heals, all writes made to the old master are irrevocably lost. How do you configure Redis min-replicas-to-write and quorum failover semantics to prevent split-brain write loss?

Prevent silent data loss in Redis Sentinel clusters during network partitions by tuning quorum configurations and enforcing min-replicas-to-write rules.

#Redis #Redis Sentinel #Split-Brain #Failover #High Availability #min-replicas-to-write
🎙️ Candidate Opening & Architectural Context
"Redis replication is asynchronous by default. When a network partition divides a Redis cluster, the Sentinel quorum promotes a replica to master in the majority partition. However, if clients in the minority partition continue writing to the partitioned old master, those writes cannot be synchronized. When the partition heals and the old master demotes to a replica, it resynchronizes from the new master by discarding its own local dataset, permanently deleting all unacknowledged minority writes."
Advertisement
⚡ Recommended Practice Lab

Want to master this scenario in a live sandbox? KodeKloud's PostgreSQL Database Administration & High Availability Course covers this exact problem with hands-on terminal drills.

🛠️ Production Runbook & Step-by-Step Resolution

⚡

Situation: Asynchronous Failover Causes 90 Seconds of Lost Writes

⚡

Task: Enforce Strict Write Quorum Guardrails on Master Instances

Advertisement
⚡

Action: Configure min-replicas-to-write & Sentinel Down-After-Milliseconds

⚡

Result: 0 Data Loss During Partition Drills

💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"Redis does not guarantee strong consistency during network partitions. Always configure `min-replicas-to-write 1` and `min-replicas-max-lag` so an isolated master stops accepting writes before Sentinels promote a new master, preventing split-brain data loss."
⚡ 60-Second Elevator Pitch Talking Points
  • Network partitions can cause split-brain where clients write to an isolated master that gets wiped upon healing.
  • Configure min-replicas-to-write 1 and min-replicas-max-lag to reject writes on partitioned masters.
  • Keep down-after-milliseconds low and distribute Sentinel instances across at least 3 availability zones.
Advertisement
Want more Databases & Storage scenarios?
Explore our complete collection of scenario-based Databases & Storage interview runbooks.
Browse All Databases & Storage Questions →