Q: Your application is healthy, but users experience intermittent timeouts. How would you identify whether the issue is in the application, network, infrastructure, or load balancer?
Production diagnostic framework to isolate intermittent timeout complaints across the application layer, network path, underlying compute infrastructure, and cloud load balancer using curl timing telemetry, connection metrics, and packet analysis.
🛠️ Production Runbook & Step-by-Step Resolution
Analyze Edge Timing Telemetry (Isolate Load Balancer vs. Backend)
Examine ALB/ELB access logs and CloudWatch metrics. Specifically compare request_processing_time, target_processing_time, and response_processing_time. If target_processing_time spikes to exactly the timeout threshold (e.g., 60.00s), the issue is downstream in the application or database. If target_processing_time is -1, the connection was dropped before reaching the target.
curl -w "@curl-format.txt" -o /dev/null -s -v https://api.production.internal/v1/health
# Sample curl-format.txt metrics:
# time_namelookup: %{time_namelookup}s
# time_connect: %{time_connect}s
# time_appconnect: %{time_appconnect}s
# time_pretransfer:%{time_pretransfer}s
# time_starttransfer: %{time_starttransfer}s
# time_total: %{time_total}s
Verify Infrastructure & Kernel Connection Queues
Log into host nodes or containers and inspect TCP socket allocation, listen backlog drops, and conntrack table saturation. If the application's listen queue is full (netstat -s | grep -i listen), requests sit in the OS backlog until the client or load balancer times out.
ss -lnt '( sport = :8080 )'
netstat -s | grep -E 'listen|overflow|drop'
cat /proc/sys/net/netfilter/nf_conntrack_count
cat /proc/sys/net/netfilter/nf_conntrack_max
Correlate Network Path & Packet Loss
Check NAT Gateway error metrics (ErrorPortAllocation, PacketsDropCount) and VPC flow logs. Verify whether AWS ALB idle timeout is shorter than the application's keep-alive timeout, which causes the load balancer to send requests down half-closed sockets resulting in 504 timeouts.
Inspect Application Thread & Connection Pool Starvation
If network and load balancer metrics are clean, the application is suffering from thread pool starvation or database connection pool exhaustion where 99% of requests succeed fast, but 1% wait in a blocking queue until the HTTP client times out.
- Check ALB access logs: compare target_processing_time against request_processing_time to see if backend or LB is timing out.
- Verify keep-alive parity: application server keepalive must be 5 seconds longer than the ALB idle timeout.
- Inspect socket listen queues: check ss -lnt and netstat -s for listen queue overflows indicating app thread exhaustion.
- Monitor NAT Gateway port allocation errors and conntrack table saturation for network-level drops.