Q: A Kubernetes cluster upgrade works perfectly in staging, but when applied to production, it corrupts CoreDNS and breaks internal name resolution. How do you approach patching and restoring service?
Root cause analysis and emergency restoration procedure when a Kubernetes cluster upgrade breaks CoreDNS in production despite passing in staging.
Want to master this scenario in a live sandbox? KodeKloud's CKA & CKAD Hands-On Certification Track covers this exact problem with hands-on terminal drills.
🛠️ Production Runbook & Step-by-Step Resolution
Immediate CoreDNS Pod Logs & Crash Diagnosis
Inspect CoreDNS pod events and container logs in `kube-system`. If CoreDNS pods are in CrashLoopBackOff, check the exit logs. Deprecated plugins in the Corefile will cause CoreDNS to refuse initialization with a syntax validation error.
kubectl get pods -n kube-system -l k8s-app=kube-dns
kubectl logs -n kube-system -l k8s-app=kube-dns --tail=100
# Common failure:
# plugin/ready: this plugin can only be used once
# Corefile:5 - Error during parsing: Unknown directive 'upstream'
Restore Known Good CoreDNS ConfigMap & Plugin Directives
Compare production's `coredns` ConfigMap against staging. Edit or restore the valid Corefile configuration, removing deprecated directives while maintaining internal forward resolvers. Trigger a rollout restart of CoreDNS pods.
kubectl edit configmap coredns -n kube-system
# Remove unsupported plugin lines, verify valid Corefile:
# .:53 {
# errors
# health
# ready
# kubernetes cluster.local in-addr.arpa ip6.arpa {
# pods insecure
# fallthrough in-addr.arpa ip6.arpa
# }
# forward . /etc/resolv.conf
# cache 30
# loop
# reload
# loadbalance
# }
kubectl rollout restart deployment coredns -n kube-system
Bypass CoreDNS via NodeLocal DNSCache or Direct IP Emergency Override
If CoreDNS binary itself is broken, deploy or fallback to NodeLocal DNSCache. In extreme emergencies where pods cannot resolve databases, temporarily patch critical application Service environment variables or /etc/hosts with ClusterIP addresses while fixing CoreDNS.
Verify Post-Fix DNS Resolution Across All Namespaces
Spin up a transient debug container to execute DNS lookups across cluster.local, external domains, and reverse in-addr.arpa entries.
kubectl run dns-test --rm -it --image=busybox:1.36 -- nslookup kubernetes.default
kubectl run dns-test --rm -it --image=busybox:1.36 -- nslookup google.com
- Inspect kube-system CoreDNS pod logs to identify rejected Corefile directives.
- Compare production vs staging coredns ConfigMaps to spot un-migrated plugins.
- Revert the Corefile ConfigMap to clean, forward-compatible syntax and rollout restart coredns.
- Deploy NodeLocal DNSCache as a permanent architectural safeguard against cluster-wide DNS failures.