⚡ ~/naveed Interview Prep
⚡ Portfolio Home ✍️ Engineering Blog Deep Dives 🎯 Interview Hub 998+ Scenarios ☸️ Kubernetes Mastery Hub 24 Modules 🎮 DevOps Arcade & Quizzes Subnet Blitz ⚡ 🗺️ DevOps Roadmaps PDFs & Guides 🤖 Morpheus Analysis AI Quant ↗ 🛠️ Developer Tools Utilities 🧪 Labs & Experiments 📄 Interactive CV & Certs 🔗 All Links & Socials ⚡ Join The Dispatch (Weekly SRE Newsletter) →
Staff SRE / Principal Architect Observability High-Throughput Streaming & SRE Netflix-Scale Systems

Q: 504 errors on the playback service. Cloud LB shows healthy, mesh sidecars pass, but users can’t stream. Triage.

Master production incident triage for a distributed streaming failure where Cloud Load Balancers report healthy target groups, mesh sidecars show green, yet end users experience 504 Gateway Timeouts.

#AWS #ALB #Envoy #Streaming #504 Gateway Timeout #SRE #Observability #Triage
🎙️ Candidate Opening & Architectural Context
"During the premiere of a major title, users flood customer support: video streams freeze and return HTTP 504 Gateway Timeouts. The Cloud Load Balancer (ALB) dashboard reports 100% healthy backend target instances, and internal Kubernetes Istio/Envoy sidecars report normal HTTP 200 health check responses. Yet, video playback requests fail. We must trace the end-to-end request flow to expose why health checks deceive the load balancer."
Advertisement

🛠️ Production Runbook & Step-by-Step Resolution

1️⃣

The Root Illusion: Health Check vs Video Stream Decoupling

Understand why health check metrics diverge completely from user traffic:

  • ALB target health checks query a lightweight endpoint (e.g. GET /healthz) which returns in 2ms from memory.
  • Actual playback requests execute high-overhead operations: manifest generation (.mpd / .m3u8), DRM license token verification, and CDN origin segment fetches.
  • The container has sufficient thread capacity to answer the 2ms health check, while all streaming worker threads are blocked waiting on slow downstream dependencies.
2️⃣

Inspect ALB and Envoy Idle Timeout Race Conditions

A 504 Gateway Timeout means a proxy in the path gave up waiting for an upstream response:

# 1. Inspect ALB Access Logs for upstream latency and target processing time
# Fields: target_processing_time, request_processing_time, response_processing_time
cat alb-access.log | awk '{print $9, $10, $11, $13}' | grep "504"

# 2. Check Envoy upstream request timeout metrics
kubectl exec -it <playback-pod> -c istio-proxy -- \
  curl -s localhost:15000/stats | grep "upstream_rq_timeout\|upstream_cx_destroy_local"
  • If ALB idle timeout is 60 seconds, and backend playback service manifest generation takes 62 seconds due to DRM database locks, the ALB closes the connection with 504.
  • Conversely, if Envoy's route timeout (default 15s) is shorter than the client request expectation, Envoy returns 504 with response flag UT (Upstream Timeout).
3️⃣

Trace the Blocking Dependency via Distributed Tracing (OpenTelemetry / Jaeger)

Inspect distributed traces for the playback session span:

  • DRM License Token Service: Is the external Widevine/FairPlay key exchange service hitting rate limits?
  • Object Storage / S3 Egress: Are S3 GET requests for video manifest chunks experiencing 503 SlowDown or NAT Gateway bandwidth saturation?
  • Egress Connection Pool Starvation: Did the playback service exhaust outbound HTTP client connection pool sockets connecting to the metadata store?
4️⃣

Immediate Mitigation Actions

Execute tactical containment during live incident response:

  • Graceful Degradation: Shed non-critical calls (e.g. disable real-time viewing history and personalized bitrate recommendations) to free up playback thread pools.
  • Static Fallback Manifests: Serve pre-computed static video manifests from edge CDN cache rather than dynamic computation.
  • Tune Keep-Alive & Timeouts: Ensure backend idle timeouts exceed load balancer idle timeouts to eliminate silent connection drops.
💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"A 504 error with green health checks indicates that the health check endpoint is decoupled from actual application work. The app is alive enough to ping, but deadlocked on a downstream bottleneck."
⚡ 60-Second Elevator Pitch Talking Points
  • A 504 Gateway Timeout while health checks pass means the health endpoint is decoupled from real work: it returns 200 OK while worker threads are saturated on streaming dependencies.
  • First, I inspect the ALB access log fields: if 'target_processing_time' exceeds 60s, the ALB timed out waiting on the backend. If it's 15s, an intermediate Envoy proxy timeout triggered the 504.
  • Second, I pull OpenTelemetry traces for the failing playback endpoint to pinpoint the bottleneck—typically DRM license verification, S3 API throttling, or egress DB pool exhaustion.
  • To mitigate immediately, we activate graceful degradation: shed personalization and analytics dependencies to reclaim worker threads, and serve pre-generated static manifests from CDN edge cache.
Advertisement
Want more Observability scenarios?
Explore our complete collection of scenario-based Observability interview runbooks.
Browse All Observability Questions →

📚 Related Production Scenarios in Observability