⚡ ~/naveed Interview Prep
⚡ Portfolio Home ✍️ Engineering Blog Deep Dives 🎯 Interview Hub 998+ Scenarios ☸️ Kubernetes Mastery Hub 24 Modules 🎮 DevOps Arcade & Quizzes Subnet Blitz ⚡ 🗺️ DevOps Roadmaps PDFs & Guides 🤖 Morpheus Analysis AI Quant ↗ 🛠️ Developer Tools Utilities 🧪 Labs & Experiments 📄 Interactive CV & Certs 🔗 All Links & Socials ⚡ Join The Dispatch (Weekly SRE Newsletter) →
Senior DevOps / SRE AWS Load Balancing High-Severity Incident

Q: ALB starts returning 5xx errors — how would you identify the root cause?

Production runbook for isolating Application Load Balancer (ALB) 5xx errors: distinguishing ELB-generated vs Target-generated codes, debugging 502/503/504, and querying access logs with Athena.

#AWS #ALB #CloudWatch #Target Groups #HTTP 502 #HTTP 504
🎙️ Candidate Opening & Architectural Context
"When an ALB starts throwing 5xx errors, my immediate first goal is to differentiate between errors generated by the ALB itself versus errors generated by backend application targets."
Advertisement

🛠️ Production Runbook & Step-by-Step Resolution

1️⃣

CloudWatch Metrics Differentiation (ELB vs Target)

Look at CloudWatch metrics for the ALB immediately:

  • HTTPCode_ELB_5XX_Count: The ALB itself generated the error before receiving a response from targets (502 Bad Gateway, 503 Service Unavailable, 504 Gateway Timeout).
  • HTTPCode_Target_5XX_Count: The backend application generated a 5xx response (e.g. unhandled 500 internal server error or database crash), and the ALB simply forwarded it to the client.
  • HealthyHostCount / UnHealthyHostCount: Check if backend targets are failing health checks.
2️⃣

Diagnose Specific 5xx Error Codes

Apply targeted root cause analysis based on status code:

  • HTTP 502 (Bad Gateway): Target closed the TCP connection before sending full HTTP headers. Classic cause: Keep-Alive timeout mismatch (backend web server keep-alive must be set GREATER than ALB idle timeout, e.g. backend 65s vs ALB 60s).
  • HTTP 503 (Service Unavailable): No healthy targets available in the Target Group (all targets failed health checks), or ALB capacity exceeded during extreme traffic spikes.
  • HTTP 504 (Gateway Timeout): Target took longer to respond than the ALB idle timeout. Check for slow database queries, thread pool exhaustion, or locked worker processes.
  • HTTP 500 (Target Generated): Application code threw an uncaught exception, or database/Redis connection failed.
3️⃣

Query ALB Access Logs with Amazon Athena

Analyze ALB S3 access logs to extract exact failing requests, targets, and response times:

  • Check elb_status_code vs target_status_code.
  • Check target_processing_time: If high (>30s), target application is hanging.
  • Identify offending request paths, client IPs, or specific unhealthy backend target IP addresses.
4️⃣

Remediate and Safeguard

Immediate mitigation and permanent safeguards:

  • If 503 (unhealthy targets): Inspect target health check path (e.g. /healthz) in Target Group. Are targets returning 404 or failing DB ping?
  • If 502: Align keep-alive timeout: set Nginx/Node.js/Gunicorn keep-alive to 65s+ and ALB timeout to 60s.
  • If 504: Temporarily bump ALB idle timeout to relieve immediate pressure while developers optimize slow DB queries.
  • Configure CloudWatch Alarms on HTTPCode_ELB_5XX_Count > 10 with SNS to PagerDuty.
💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"Always split ELB 5xx from Target 5xx in CloudWatch first. If it's ELB 502, check Keep-Alive timeout mismatches; if 503, check Target Group healthy host count; if 504, check target processing latency."
⚡ 60-Second Elevator Pitch Talking Points
  • Check CloudWatch: HTTPCode_ELB_5XX (ALB fault) vs HTTPCode_Target_5XX (app fault).
  • If Target 5XX: App is throwing 500s; inspect app logs and database connectivity.
  • If ELB 502: Target closed TCP prematurely; fix Keep-Alive timeout (backend keep-alive must exceed ALB 60s).
  • If ELB 503: Zero healthy targets in Target Group; verify health check path (/healthz) and target capacity.
  • If ELB 504: Gateway timeout; target query took >60s. Check target_processing_time and database locks.
  • Query ALB Access Logs via Athena to isolate failing URLs, client IPs, and specific target instance IDs.
Advertisement
Want more AWS scenarios?
Explore our complete collection of scenario-based AWS interview runbooks.
Browse All AWS Questions →

📚 Related Production Scenarios in AWS