Q: A Kubernetes application is healthy according to 'kubectl get pods' (1/1 Running), but users report 504 Gateway Timeout errors. Walk me through your troubleshooting flow.
Architectural diagnostic workflow for resolving 504 Gateway Timeout errors when pod readiness and liveness probes show 100% green.
Want to master this scenario in a live sandbox? KodeKloud's CKA & CKAD Hands-On Certification Track covers this exact problem with hands-on terminal drills.
🛠️ Production Runbook & Step-by-Step Resolution
Differentiate Shallow Health Probes from Deep Application Starvation
Verify what the probe actually tests. If '/healthz' simply returns HTTP 200 without testing database pools, external APIs, or message queues, the pod appears 1/1 Running while all actual business logic threads are blocked waiting on locked PostgreSQL rows or downstream microservices.
# Inspect pod liveness/readiness probe definitions
kubectl get pod <pod-name> -o jsonpath='{.spec.containers[*].readinessProbe}'
# Exec into container or port-forward to test realistic requests
kubectl port-forward pod/<pod-name> 8080:8080 &
curl -v http://localhost:8080/api/v1/checkout
Diagnose Keepalive & Idle Timeout Mismatches
Check timeout alignment between your Ingress controller / ALB and the backend container runtime. If the Load Balancer has an idle timeout of 60s, but the application server (e.g. Node.js, Gunicorn, Tomcat) has a keep-alive timeout of 50s, the backend server will silently close the TCP connection right as the LB reuses it, resulting in 502/504 errors under high concurrency.
Inspect Ingress Controller Logs & Upstream Response Times
Review NGINX Ingress or AWS Load Balancer access logs. Compare $upstream_connect_time, $upstream_header_time, and $upstream_response_time. If $upstream_response_time is exactly 60.000s, the container received the request but hung.
kubectl logs -n ingress-nginx -l app.kubernetes.io/name=ingress-nginx --tail=200 | grep ' 504 '
# Output pattern:
# upstream_response_time: 60.004, upstream_status: 504, request_time: 60.005
Check Thread Pool & Connection Pool Saturation Inside Pod
Inspect socket queues and thread dumps. Use 'ss -lnt' to see if the Send-Q or Recv-Q on the listen port is backed up, indicating that the kernel accepted TCP SYN packets, but the application process is too slow to call accept().
kubectl exec -it <pod-name> -- ss -lnt '( sport = :8080 )'
# Check thread exhaustion / GC pause:
jstack <pid> | grep -E 'BLOCKED|WAITING' # For Java
kill -USR1 <pid> # For Node.js heap/CPU profile
- Recognize that shallow /healthz probes pass even when application thread pools are deadlocked.
- Check ingress logs for upstream_response_time matching the exact proxy timeout threshold (e.g. 60s).
- Verify keep-alive timeout hierarchy: backend container keepalive must be 5-10s longer than Ingress/ALB idle timeout.
- Inspect socket listen queues (ss -lnt) and application connection pools to catch thread starvation.