Q: A customer complains of intermittent 502 Bad Gateway errors from an AWS Application Load Balancer. The backend instances show low CPU. What should you look for?
A 502 Bad Gateway from an ALB means the ALB tried to communicate with the target instance, but the connection dropped or the target retur...
🛠️ Production Runbook & Step-by-Step Resolution
Production Solution & Architecture
A 502 Bad Gateway from an ALB means the ALB tried to communicate with the target instance, but the connection dropped or the target returned an invalid response. The most common cause, aside from the app actually crashing, is a mismatch in Keep-Alive timeouts. If the backend web server (like Nginx or Node.js) has an idle timeout configured shorter than the ALB's idle timeout (default 60 seconds), the backend might close the TCP connection just as the ALB decides to send a new request down that established pipe. The ALB gets a connection reset and throws a 502. The fix is to ensure the backend application's idle timeout is greater than the ALB's idle timeout.
- Immediate Triage: A 502 Bad Gateway from an ALB means the ALB tried to communicate with the target instance, but
- Run targeted verification commands before modifying configuration.
- Automate permanent guardrails (CI check, alerts, IaC policy) to prevent recurrence.