⚡ ~/naveed Interview Prep
⚡ Portfolio Home ✍️ Engineering Blog Deep Dives 🎯 Interview Hub 998+ Scenarios ☸️ Kubernetes Mastery Hub 24 Modules 🎮 DevOps Arcade & Quizzes Subnet Blitz ⚡ 🗺️ DevOps Roadmaps PDFs & Guides 🤖 Morpheus Analysis AI Quant ↗ 🛠️ Developer Tools Utilities 🧪 Labs & Experiments 📄 Interactive CV & Certs 🔗 All Links & Socials ⚡ Join The Dispatch (Weekly SRE Newsletter) →
Senior DevOps / SRE Networking Load Balancing & Traffic Diagnostics Barclays Classic

Q: Your application is healthy, but users experience intermittent timeouts. How would you identify whether the issue is in the application, network, infrastructure, or load balancer?

Production diagnostic framework to isolate intermittent timeout complaints across the application layer, network path, underlying compute infrastructure, and cloud load balancer using curl timing telemetry, connection metrics, and packet analysis.

#Networking #Load Balancer #Timeouts #Incident Triage #Barclays #SRE #Troubleshooting
🎙️ Candidate Opening & Architectural Context
"When users experience intermittent timeouts while target health checks remain green, the root cause is almost always an edge saturation condition—such as TCP connection queue overflow, ephemeral port exhaustion, upstream keep-alive mismatches, or tail-latency spikes that health probes miss. I isolate the fault layer-by-layer using client-side timing telemetry and load balancer access logs."
Advertisement

🛠️ Production Runbook & Step-by-Step Resolution

1

Analyze Edge Timing Telemetry (Isolate Load Balancer vs. Backend)

Examine ALB/ELB access logs and CloudWatch metrics. Specifically compare request_processing_time, target_processing_time, and response_processing_time. If target_processing_time spikes to exactly the timeout threshold (e.g., 60.00s), the issue is downstream in the application or database. If target_processing_time is -1, the connection was dropped before reaching the target.

curl -w "@curl-format.txt" -o /dev/null -s -v https://api.production.internal/v1/health
# Sample curl-format.txt metrics:
# time_namelookup: %{time_namelookup}s
# time_connect:    %{time_connect}s
# time_appconnect: %{time_appconnect}s
# time_pretransfer:%{time_pretransfer}s
# time_starttransfer: %{time_starttransfer}s
# time_total:      %{time_total}s
2

Verify Infrastructure & Kernel Connection Queues

Log into host nodes or containers and inspect TCP socket allocation, listen backlog drops, and conntrack table saturation. If the application's listen queue is full (netstat -s | grep -i listen), requests sit in the OS backlog until the client or load balancer times out.

ss -lnt '( sport = :8080 )'
netstat -s | grep -E 'listen|overflow|drop'
cat /proc/sys/net/netfilter/nf_conntrack_count
cat /proc/sys/net/netfilter/nf_conntrack_max
3

Correlate Network Path & Packet Loss

Check NAT Gateway error metrics (ErrorPortAllocation, PacketsDropCount) and VPC flow logs. Verify whether AWS ALB idle timeout is shorter than the application's keep-alive timeout, which causes the load balancer to send requests down half-closed sockets resulting in 504 timeouts.

Pro Tip: Golden Rule: Always set application server keep-alive timeout higher than the load balancer's idle timeout (e.g., ALB idle timeout = 60s, App server keep-alive = 65s) to prevent race-condition connection resets.
4

Inspect Application Thread & Connection Pool Starvation

If network and load balancer metrics are clean, the application is suffering from thread pool starvation or database connection pool exhaustion where 99% of requests succeed fast, but 1% wait in a blocking queue until the HTTP client times out.

💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"Intermittent timeouts with healthy probes indicate queue saturation or keep-alive race conditions. Break down latency into DNS, TCP handshake, TLS negotiation, and TTFB to immediately identify the guilty layer."
⚡ 60-Second Elevator Pitch Talking Points
  • Check ALB access logs: compare target_processing_time against request_processing_time to see if backend or LB is timing out.
  • Verify keep-alive parity: application server keepalive must be 5 seconds longer than the ALB idle timeout.
  • Inspect socket listen queues: check ss -lnt and netstat -s for listen queue overflows indicating app thread exhaustion.
  • Monitor NAT Gateway port allocation errors and conntrack table saturation for network-level drops.
Advertisement
Want more Networking scenarios?
Explore our complete collection of scenario-based Networking interview runbooks.
Browse All Networking Questions →

📚 Related Production Scenarios in Networking