Q: You have an ELK stack. The Elasticsearch cluster status turns Yellow. What does this mean, and what do you do?
Elasticsearch cluster states are:
#Observability #Observability #L2 #Monitoring #Prometheus #SRE
🎙️ Candidate Opening & Architectural Context
""In an interview, I explain how we designed actionable, symptom-based alerting using the Four Golden Signals. The interviewer is testing: Elasticsearch shard mechanics, replica management.. I structure my answer around systematic triage first, root cause analysis second, and permanent remediation third.""
Advertisement
🛠️ Production Runbook & Step-by-Step Resolution
1️⃣
Initial Diagnostics & Root Cause Analysis
Elasticsearch cluster states are:
- Green: All primary and replica shards are allocated.
- Yellow: All primary shards are allocated (data is safe, searching works), but one or more replica shards are unassigned.
- Red: One or more primary shards are missing (data loss or downtime).
2️⃣
Remediation & Permanent Safeguards
A Yellow state usually happens because a node went down or restarted, and ES cannot allocate the replica shard to the same node holding the primary shard. To fix:
- Check
_cat/healthand_cluster/allocation/explainto see *why* and *which* shards aren't allocating. - The usual cause is either an offline node (I need to bring it back up, or wait for ES to timeout and recreate the replica on another node if there's space) or a disk watermark issue (disks are >85% full, preventing new shard allocation, so I need to clear old indices or add disk space).
💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"Pro-Tip: Green: All primary and replica shards are allocated.."
⚡ 60-Second Elevator Pitch Talking Points
- Green: All primary and replica shards are allocated.
- Yellow: All primary shards are allocated (data is safe, searching works), but one or more replica...
- Red: One or more primary shards are missing (data loss or downtime).
Advertisement