⚡ ~/naveed Interview Prep
⚡ Portfolio Home ✍️ Engineering Blog Deep Dives 🎯 Interview Hub 1,000+ Scenarios ☸️ Kubernetes Mastery Hub 24 Modules 🎮 DevOps Arcade & Quizzes Subnet Blitz ⚡ 🗺️ DevOps Roadmaps PDFs & Guides 🤖 Morpheus Analysis AI Quant ↗ 🛠️ Developer Tools Utilities 🧪 Labs & Experiments 📄 Interactive CV & Certs 🔗 All Links & Socials ⚡ Join The Dispatch (Weekly SRE Newsletter) →
← Back to All Networking & Cloud DNS Interview Questions Scenario 48 of 53 in Networking & Cloud DNS
Senior DevOps / SRE Networking Load Balancing & Traffic Diagnostics J.P. Morgan Technical Loop

Q: In a canary deployment to production, half the traffic returns HTTP 502 Bad Gateway, while the other half succeeds. Walk us through your troubleshooting approach.

Diagnostic framework for resolving split-traffic 502 Bad Gateway errors during a canary deployment rollout in Kubernetes.

#Networking #Canary Deployment #502 Bad Gateway #Istio #Ingress #Argo Rollouts #Troubleshooting
🎙️ Candidate Opening & Architectural Context
"When exactly 50% of requests return HTTP 502 during a canary deployment while the remaining 50% succeed, traffic is splitting evenly between two upstream target groups: one healthy (stable version) and one failing (canary version). A 502 indicates that the Ingress controller or Service Mesh reverse proxy received an invalid response or connection reset (RST) from the canary backend."
Advertisement
⚡ Recommended Practice Lab

Want to master this scenario in a live sandbox? KodeKloud's Istio Service Mesh & Advanced Kubernetes Networking Course covers this exact problem with hands-on terminal drills.

🛠️ Production Runbook & Step-by-Step Resolution

1

Confirm Traffic Split Routing & Upstream Targets

Inspect the Ingress, VirtualService, or Load Balancer configuration to confirm how the 50/50 split is allocated. Identify the exact Pod IP addresses belonging to the canary ReplicaSet.

# Identify Pods in canary ReplicaSet
kubectl get pods -l app=payment-service --show-labels
kubectl get endpoints payment-service-canary -o wide
2

Inspect Upstream Connection Resets & Port Mismatches

Common causes of 502 on newly deployed canaries: 1. **Container Port Mismatch**: Service targetPort does not match the actual port exposed by the new container image (e.g. Service routes to 8080, but new Dockerfile exposes 8000). 2. **Immediate Crash on Connection**: The canary pod binds the socket but crashes on the first incoming request due to missing environment variables or secret misconfigurations. 3. **TLS Mismatch**: Ingress expects plaintext HTTP upstream, but canary was configured with HTTPS (or vice versa).

# Test canary pod directly from within the cluster bypassing ingress
kubectl exec -it <debug-pod> -- curl -v http://<canary-pod-ip>:<target-port>/healthz
Advertisement
3

Halt Canary Progression & Route Traffic Back to 100% Stable

Before deep-dive debugging, mitigate user impact immediately. Roll canary traffic back to 0% in Argo Rollouts or patch the Istio VirtualService weights.

# Rollback canary immediately via Argo Rollouts
kubectl argo rollouts abort payment-service
# Or update Istio VirtualService weight to 100% stable
kubectl patch virtualservice payment-service --type merge -p '{"spec":{"http":[{"route":[{"destination":{"host":"payment-stable"},"weight":100},{"destination":{"host":"payment-canary"},"weight":0}]}]}}'
4

Extract Canary Logs & Core Dump

Check the canary pod logs and previous container logs (`kubectl logs --previous`) to analyze application exceptions, uncaught startup errors, or missing runtime configurations.

💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"50% 502 errors indicate that all traffic reaching the canary pool fails. Immediately halt the canary and route 100% of traffic back to stable, then verify targetPort mismatches, missing secrets, or HTTP/HTTPS protocol mismatches."
⚡ 60-Second Elevator Pitch Talking Points
  • Recognize that a 50% failure rate maps directly to the canary replica pool failing completely.
  • Immediately abort the canary rollout to direct 100% of user traffic back to the stable deployment.
  • Verify container targetPort mapping against the Kubernetes Service spec.
  • Curl the canary pod IP directly from a cluster debug pod to isolate whether the app is crashing or misconfigured.
Advertisement
Want more Networking & Cloud DNS scenarios?
Explore our complete collection of scenario-based Networking & Cloud DNS interview runbooks.
Browse All Networking & Cloud DNS Questions →