⚡ ~/naveed Interview Prep
⚡ Portfolio Home ✍️ Engineering Blog Deep Dives 🎯 Interview Hub 998+ Scenarios ☸️ Kubernetes Mastery Hub 24 Modules 🎮 DevOps Arcade & Quizzes Subnet Blitz ⚡ 🗺️ DevOps Roadmaps PDFs & Guides 🤖 Morpheus Analysis AI Quant ↗ 🛠️ Developer Tools Utilities 🧪 Labs & Experiments 📄 Interactive CV & Certs 🔗 All Links & Socials ⚡ Join The Dispatch (Weekly SRE Newsletter) →
Senior DevOps / SRE Kubernetes Networking & Ingress Production Incident

Q: During high traffic, your app shows intermittent 502 errors through Ingress — how do you debug and fix it?

Root-cause analysis and resolution for intermittent HTTP 502 Bad Gateway errors during traffic peaks: diagnosing upstream timeouts, CPU throttling, readiness probe flapping, tuning Ingress proxy timeouts, and autoscaling HPA.

#Kubernetes #Ingress #NGINX Ingress #HTTP 502 #HPA #Performance #Networking
🎙️ Candidate Opening & Architectural Context
"I debug 502s from the edge inward: Ingress controller logs and metrics, upstream service endpoints, pod readiness, connection saturation, timeouts, and application logs. Under high traffic, common causes are insufficient replicas, slow upstreams, readiness flapping, keepalive/timeout mismatches, or node CPU throttling. I correlate timestamps across Ingress, service, and application telemetry before tuning."
Advertisement

🛠️ Production Runbook & Step-by-Step Resolution

1️⃣

Ingress Controller Logs & Service Endpoints Validation

Differentiate whether the 502 is generated by the Ingress controller or returned by the backend application:

# Filter Ingress controller logs for 502 responses
kubectl logs -n ingress-nginx deploy/ingress-nginx-controller --since=15m | grep ' 502 '
kubectl describe ingress public-api -n app

# Check service endpoints and watch pod restarts
kubectl get svc,endpoints -n app api -o wide
kubectl get pods -n app -l app=api -w
kubectl describe pod -n app <pod-name> | egrep 'Readiness|Liveness|Restart'
  • Ingress Controller Error Logs: Search NGINX Ingress controller logs for upstream timed out (110: Connection timed out) or connect() failed (111: Connection refused).
  • Service Endpoints Check: Verify whether active endpoints are dropping or flapping under load.
  • Pod Readiness Flapping: Check if high CPU causes readiness probes to time out, removing pods from endpoints dynamically.
2️⃣

Tuning Resource Pressure, HPA, and Upstream Proxy Timeouts

Remediate upstream latency bottlenecks and adjust Ingress controller keepalive and timeout limits:

# Inspect resource consumption and autoscaling
kubectl top pods -n app
kubectl top nodes
kubectl get hpa -n app
kubectl describe hpa api -n app

# Extend NGINX Ingress timeout annotations
kubectl annotate ingress public-api -n app nginx.ingress.kubernetes.io/proxy-read-timeout='120' --overwrite
kubectl annotate ingress public-api -n app nginx.ingress.kubernetes.io/proxy-send-timeout='120' --overwrite
  • CPU Throttling & HPA: Inspect kubectl top pods and HPA metrics; increase HPA minReplicas so capacity is pre-warmed for peaks.
  • Tune Proxy Timeouts: If backend queries take longer during traffic spikes, annotate the Ingress to extend proxy read and send timeouts from default 60s to 120s.
  • Database & Query Optimization: Resolve backend bottleneck (e.g. unindexed query or connection pool starvation) that triggered slow upstream processing.
💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"502 Bad Gateway means the Ingress controller failed to get a timely HTTP response from upstream pods. Trace from edge logs to service endpoints, verify readiness probe stability, and tune timeouts alongside HPA scaling."
⚡ 60-Second Elevator Pitch Talking Points
  • Inspect NGINX Ingress controller logs to verify whether errors are upstream connection timeouts or dropped endpoints.
  • Check pod CPU throttling and readiness probe failures that temporarily remove pods from Service endpoints during traffic spikes.
  • Apply permanent fixes: increase HPA minReplicas, optimize upstream connection pools, and tune Ingress proxy-read-timeout annotations.
Advertisement
Want more Kubernetes scenarios?
Explore our complete collection of scenario-based Kubernetes interview runbooks.
Browse All Kubernetes Questions →

📚 Related Production Scenarios in Kubernetes