⚡ ~/naveed Interview Prep
⚡ Portfolio Home ✍️ Engineering Blog Deep Dives 🎯 Interview Hub 1,000+ Scenarios ☸️ Kubernetes Mastery Hub 24 Modules 🎮 DevOps Arcade & Quizzes Subnet Blitz ⚡ 🗺️ DevOps Roadmaps PDFs & Guides 🤖 Morpheus Analysis AI Quant ↗ 🛠️ Developer Tools Utilities 🧪 Labs & Experiments 📄 Interactive CV & Certs 🔗 All Links & Socials ⚡ Join The Dispatch (Weekly SRE Newsletter) →
← Back to All Kubernetes Interview Questions Scenario 181 of 194 in Kubernetes
Senior SRE / Kubernetes Architect Kubernetes Ingress & Service Triage High-Impact Incident

Q: A Kubernetes application is healthy according to 'kubectl get pods' (1/1 Running), but users report 504 Gateway Timeout errors. Walk me through your troubleshooting flow.

Architectural diagnostic workflow for resolving 504 Gateway Timeout errors when pod readiness and liveness probes show 100% green.

#Kubernetes #504 Gateway Timeout #Ingress #ALB #Keepalive #CoreDNS #Troubleshooting
🎙️ Candidate Opening & Architectural Context
"A 504 Gateway Timeout means an edge proxy (ALB, NGINX Ingress, Cloudflare) established a connection to an upstream gateway or backend, but did not receive a timely HTTP response before its idle timeout expired. Because 'kubectl get pods' only reflects liveness/readiness probes (which often query a lightweight '/healthz' endpoint), the application worker threads can be completely deadlocked on downstream database queries while the health endpoint responds instantly."
Advertisement
⚡ Recommended Practice Lab

Want to master this scenario in a live sandbox? KodeKloud's CKA & CKAD Hands-On Certification Track covers this exact problem with hands-on terminal drills.

🛠️ Production Runbook & Step-by-Step Resolution

1

Differentiate Shallow Health Probes from Deep Application Starvation

Verify what the probe actually tests. If '/healthz' simply returns HTTP 200 without testing database pools, external APIs, or message queues, the pod appears 1/1 Running while all actual business logic threads are blocked waiting on locked PostgreSQL rows or downstream microservices.

# Inspect pod liveness/readiness probe definitions
kubectl get pod <pod-name> -o jsonpath='{.spec.containers[*].readinessProbe}'
# Exec into container or port-forward to test realistic requests
kubectl port-forward pod/<pod-name> 8080:8080 &
curl -v http://localhost:8080/api/v1/checkout
2

Diagnose Keepalive & Idle Timeout Mismatches

Check timeout alignment between your Ingress controller / ALB and the backend container runtime. If the Load Balancer has an idle timeout of 60s, but the application server (e.g. Node.js, Gunicorn, Tomcat) has a keep-alive timeout of 50s, the backend server will silently close the TCP connection right as the LB reuses it, resulting in 502/504 errors under high concurrency.

Pro Tip: Architecture Requirement: Application keep-alive timeout must ALWAYS be strictly greater than upstream Ingress/ALB idle timeout (e.g. ALB = 60s, App Server keep-alive = 65s).
Advertisement
3

Inspect Ingress Controller Logs & Upstream Response Times

Review NGINX Ingress or AWS Load Balancer access logs. Compare $upstream_connect_time, $upstream_header_time, and $upstream_response_time. If $upstream_response_time is exactly 60.000s, the container received the request but hung.

kubectl logs -n ingress-nginx -l app.kubernetes.io/name=ingress-nginx --tail=200 | grep ' 504 '
# Output pattern:
# upstream_response_time: 60.004, upstream_status: 504, request_time: 60.005
4

Check Thread Pool & Connection Pool Saturation Inside Pod

Inspect socket queues and thread dumps. Use 'ss -lnt' to see if the Send-Q or Recv-Q on the listen port is backed up, indicating that the kernel accepted TCP SYN packets, but the application process is too slow to call accept().

kubectl exec -it <pod-name> -- ss -lnt '( sport = :8080 )'
# Check thread exhaustion / GC pause:
jstack <pid> | grep -E 'BLOCKED|WAITING' # For Java
kill -USR1 <pid> # For Node.js heap/CPU profile
💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"Pod Running status only proves that container PID 1 is alive and '/healthz' returned 200. 504s are caused by upstream keepalive race conditions, database connection pool exhaustion, or downstream service hangs."
⚡ 60-Second Elevator Pitch Talking Points
  • Recognize that shallow /healthz probes pass even when application thread pools are deadlocked.
  • Check ingress logs for upstream_response_time matching the exact proxy timeout threshold (e.g. 60s).
  • Verify keep-alive timeout hierarchy: backend container keepalive must be 5-10s longer than Ingress/ALB idle timeout.
  • Inspect socket listen queues (ss -lnt) and application connection pools to catch thread starvation.
Advertisement
📥 FREE DOWNLOAD · 101-PAGE COMPANION HANDBOOK
Studying for Kubernetes & SRE Technical Rounds?
Download the complete 100-question PDF field guide covering all 11 core modules with offline diagnostic runbooks.
📥 Download PDF (Free) Read Online Guide →
Want more Kubernetes scenarios?
Explore our complete collection of scenario-based Kubernetes interview runbooks.
Browse All Kubernetes Questions →