Q: During a transient network partition between AWS Availability Zones, a Redis master node is isolated from its Sentinel monitors and replica instances. The Sentinels detect quorum loss, declare the master dead, and promote a replica in AZ-b to become the new master. However, client applications in AZ-a still have network access to the old partitioned master and continue executing writes for 90 seconds. When the network heals, all writes made to the old master are irrevocably lost. How do you configure Redis min-replicas-to-write and quorum failover semantics to prevent split-brain write loss?
Prevent silent data loss in Redis Sentinel clusters during network partitions by tuning quorum configurations and enforcing min-replicas-to-write rules.
Want to master this scenario in a live sandbox? KodeKloud's PostgreSQL Database Administration & High Availability Course covers this exact problem with hands-on terminal drills.
🛠️ Production Runbook & Step-by-Step Resolution
Situation: Asynchronous Failover Causes 90 Seconds of Lost Writes
Task: Enforce Strict Write Quorum Guardrails on Master Instances
Action: Configure min-replicas-to-write & Sentinel Down-After-Milliseconds
Result: 0 Data Loss During Partition Drills
- Network partitions can cause split-brain where clients write to an isolated master that gets wiped upon healing.
- Configure min-replicas-to-write 1 and min-replicas-max-lag to reject writes on partitioned masters.
- Keep down-after-milliseconds low and distribute Sentinel instances across at least 3 availability zones.