Q: Your readiness probe keeps failing even though the app is working fine. What could be wrong?
1. Wrong port or path — probe is checking a different port/endpoint than the app actually serves.
#Kubernetes #Advanced Scenarios #L2 #Container Orchestration #K8s
🎙️ Candidate Opening & Architectural Context
""In one of our high-traffic production clusters, our SRE team handled this incident using a standardized runbook. When addressing this question, I walk the interviewer through our production incident runbook: isolating the blast radius, checking diagnostic logs and metrics, and applying a safe fix.""
Advertisement
🛠️ Production Runbook & Step-by-Step Resolution
1️⃣
Initial Diagnostics & Root Cause Analysis
Debug: kubectl describe pod shows probe failure reason. Also exec into the pod and manually curl the readiness endpoint to see what it returns.
- Wrong port or path — probe is checking a different port/endpoint than the app actually serves.
- Probe timeout too short — if the app takes 2 seconds to respond and
timeoutSeconds: 1, it always times out. - App returns non-200 for the health path under load — the readiness endpoint has a bug.
2️⃣
Remediation & Permanent Safeguards
- initialDelaySeconds too short — app needs more time to start before readiness checks begin.
- Checking the wrong container port name — if using
port: httpand the port isn't named, it fails.
💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"Pro-Tip: Wrong port or path — probe is checking a different port/endpoint than the app actually serves.."
⚡ 60-Second Elevator Pitch Talking Points
- Wrong port or path — probe is checking a different port/endpoint than the app actually serves.
- Probe timeout too short — if the app takes 2 seconds to respond and timeoutSeconds: 1, it always ...
- App returns non-200 for the health path under load — the readiness endpoint has a bug.
Advertisement