Q: You’ve deployed an application to Azure Kubernetes Service (AKS) and it fails health checks randomly. How do you debug this end-to-end?
End-to-end troubleshooting methodology to diagnose and resolve intermittent liveness and readiness probe failures in Azure Kubernetes Service (AKS).
Want to master this scenario in a live sandbox? KodeKloud's CKA & CKAD Hands-On Certification Track covers this exact problem with hands-on terminal drills.
🛠️ Production Runbook & Step-by-Step Resolution
Inspect Kubelet Events & Probe Configuration
Check whether failure events report 'Unhealthy' or 'Connection refused'. Examine `timeoutSeconds`, `periodSeconds`, and `failureThreshold` in the pod manifest. If `timeoutSeconds` is set to the default 1s and the application experiences brief GC pauses or database latency, kubelet immediately registers a failure.
kubectl describe pod <pod-name> -n production
# Look for Events:
# Warning Unhealthy Liveness probe failed: HTTP probe failed with statuscode: 500 or timeout
Check Node & Container CPU Throttling (CFS Quotas)
When container CPU limits (`resources.limits.cpu`) are too low, the Linux Completely Fair Scheduler (CFS) throttles the application threads periodically. When throttled, the application process cannot respond to the kubelet HTTP GET probe within 1 second.
# Check if container is throttled via cgroup
kubectl exec -it <pod-name> -- cat /sys/fs/cgroup/cpu/cpu.stat
# nr_throttled: count of periods where application was throttled
Verify Azure CNI Networking & Conntrack Table Limits
In AKS with Azure CNI, verify whether the node's VNet subnet has IP exhaustion or if the node kernel conntrack table (`nf_conntrack_count`) is approaching `nf_conntrack_max`. When conntrack overflows, kubelet packets to pod IPs are silently dropped.
kubectl get nodes -o wide
az aks show -g rg-banking -n aks-prod --query "networkProfile.networkPlugin"
# Check node conntrack count
kubectl get node <node-name> -o jsonpath='{.status.conditions}'
Isolate Endpoint Logic from Heavy Business Workloads
Ensure the readiness probe endpoint does not execute deep blocking database queries (`SELECT 1 FROM massive_table`) or external API calls. Move deep dependencies to startup probes and keep readiness probes strictly evaluating local process health and connection pool availability.
- Check pod describe events to determine whether failures are timeouts or connection resets.
- Verify container CPU limits to ensure CFS throttling is not pausing application threads.
- Inspect Azure CNI networking and conntrack table saturation on the host node.
- Decouple heavy database checks from readiness probes; rely on lightweight process health endpoints.