⚡ ~/naveed Interview Prep
⚡ Portfolio Home ✍️ Engineering Blog Deep Dives 🎯 Interview Hub 1,000+ Scenarios ☸️ Kubernetes Mastery Hub 24 Modules 🎮 DevOps Arcade & Quizzes Subnet Blitz ⚡ 🗺️ DevOps Roadmaps PDFs & Guides 🤖 Morpheus Analysis AI Quant ↗ 🛠️ Developer Tools Utilities 🧪 Labs & Experiments 📄 Interactive CV & Certs 🔗 All Links & Socials ⚡ Join The Dispatch (Weekly SRE Newsletter) →
← Back to All Kubernetes Interview Questions Scenario 183 of 194 in Kubernetes
Senior DevOps / SRE Kubernetes Pod Lifecycle & Health Probes J.P. Morgan Technical Loop

Q: You’ve deployed an application to Azure Kubernetes Service (AKS) and it fails health checks randomly. How do you debug this end-to-end?

End-to-end troubleshooting methodology to diagnose and resolve intermittent liveness and readiness probe failures in Azure Kubernetes Service (AKS).

#Kubernetes #AKS #Azure #Health Checks #Probes #Networking #Troubleshooting
🎙️ Candidate Opening & Architectural Context
"Intermittent health check failures on AKS typically point to four architectural issues: probe timeout thresholds too aggressive under burst traffic, node-level CPU throttling starving kubelet worker threads, Azure CNI IP exhaustion/conntrack packet drops, or application garbage collection (GC) stop-the-world pauses blocking the health check endpoint."
Advertisement
⚡ Recommended Practice Lab

Want to master this scenario in a live sandbox? KodeKloud's CKA & CKAD Hands-On Certification Track covers this exact problem with hands-on terminal drills.

🛠️ Production Runbook & Step-by-Step Resolution

1

Inspect Kubelet Events & Probe Configuration

Check whether failure events report 'Unhealthy' or 'Connection refused'. Examine `timeoutSeconds`, `periodSeconds`, and `failureThreshold` in the pod manifest. If `timeoutSeconds` is set to the default 1s and the application experiences brief GC pauses or database latency, kubelet immediately registers a failure.

kubectl describe pod <pod-name> -n production
# Look for Events:
# Warning  Unhealthy  Liveness probe failed: HTTP probe failed with statuscode: 500 or timeout
2

Check Node & Container CPU Throttling (CFS Quotas)

When container CPU limits (`resources.limits.cpu`) are too low, the Linux Completely Fair Scheduler (CFS) throttles the application threads periodically. When throttled, the application process cannot respond to the kubelet HTTP GET probe within 1 second.

# Check if container is throttled via cgroup
kubectl exec -it <pod-name> -- cat /sys/fs/cgroup/cpu/cpu.stat
# nr_throttled: count of periods where application was throttled
Advertisement
3

Verify Azure CNI Networking & Conntrack Table Limits

In AKS with Azure CNI, verify whether the node's VNet subnet has IP exhaustion or if the node kernel conntrack table (`nf_conntrack_count`) is approaching `nf_conntrack_max`. When conntrack overflows, kubelet packets to pod IPs are silently dropped.

kubectl get nodes -o wide
az aks show -g rg-banking -n aks-prod --query "networkProfile.networkPlugin"
# Check node conntrack count
kubectl get node <node-name> -o jsonpath='{.status.conditions}'
4

Isolate Endpoint Logic from Heavy Business Workloads

Ensure the readiness probe endpoint does not execute deep blocking database queries (`SELECT 1 FROM massive_table`) or external API calls. Move deep dependencies to startup probes and keep readiness probes strictly evaluating local process health and connection pool availability.

Pro Tip: Best Practice: Separate startup probes (for initial bootstrapping) from readiness probes. Increase timeoutSeconds to 3s-5s and failureThreshold to 3.
💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"Intermittent AKS probe failures stem from aggressive 1s probe timeouts, CFS CPU quota throttling, or heavy queries inside health endpoints. Tune timeout thresholds and verify cgroup throttling metrics."
⚡ 60-Second Elevator Pitch Talking Points
  • Check pod describe events to determine whether failures are timeouts or connection resets.
  • Verify container CPU limits to ensure CFS throttling is not pausing application threads.
  • Inspect Azure CNI networking and conntrack table saturation on the host node.
  • Decouple heavy database checks from readiness probes; rely on lightweight process health endpoints.
Advertisement
📥 FREE DOWNLOAD · 101-PAGE COMPANION HANDBOOK
Studying for Kubernetes & SRE Technical Rounds?
Download the complete 100-question PDF field guide covering all 11 core modules with offline diagnostic runbooks.
📥 Download PDF (Free) Read Online Guide →
Want more Kubernetes scenarios?
Explore our complete collection of scenario-based Kubernetes interview runbooks.
Browse All Kubernetes Questions →