⚡ ~/naveed Interview Prep
⚡ Portfolio Home ✍️ Engineering Blog Deep Dives 🎯 Interview Hub 1,000+ Scenarios ☸️ Kubernetes Mastery Hub 24 Modules 🎮 DevOps Arcade & Quizzes Subnet Blitz ⚡ 🗺️ DevOps Roadmaps PDFs & Guides 🤖 Morpheus Analysis AI Quant ↗ 🛠️ Developer Tools Utilities 🧪 Labs & Experiments 📄 Interactive CV & Certs 🔗 All Links & Socials ⚡ Join The Dispatch (Weekly SRE Newsletter) →
← Back to All Kubernetes Interview Questions Scenario 182 of 194 in Kubernetes
Staff Kubernetes SRE / Platform Architect Kubernetes Control Plane & CoreDNS Triage Upgrade Disaster Recovery

Q: A Kubernetes cluster upgrade works perfectly in staging, but when applied to production, it corrupts CoreDNS and breaks internal name resolution. How do you approach patching and restoring service?

Root cause analysis and emergency restoration procedure when a Kubernetes cluster upgrade breaks CoreDNS in production despite passing in staging.

#Kubernetes #CoreDNS #Cluster Upgrade #ConfigMap #etcd #Rollback #DNS Outage
🎙️ Candidate Opening & Architectural Context
"CoreDNS breakages during minor or major Kubernetes upgrades almost always occur due to deprecated Corefile plugin syntax (e.g. 'upstream', 'ready', or 'forward' plugin changes), ConfigMap schema discrepancies between staging and production (such as custom stub-domains or rewrite rules in prod that staging lacked), or iptables/IPVS conntrack packet loss. Because all microservices rely on CoreDNS for service discovery, this represents a cluster-wide Sev-1 outage."
Advertisement
⚡ Recommended Practice Lab

Want to master this scenario in a live sandbox? KodeKloud's CKA & CKAD Hands-On Certification Track covers this exact problem with hands-on terminal drills.

🛠️ Production Runbook & Step-by-Step Resolution

1

Immediate CoreDNS Pod Logs & Crash Diagnosis

Inspect CoreDNS pod events and container logs in `kube-system`. If CoreDNS pods are in CrashLoopBackOff, check the exit logs. Deprecated plugins in the Corefile will cause CoreDNS to refuse initialization with a syntax validation error.

kubectl get pods -n kube-system -l k8s-app=kube-dns
kubectl logs -n kube-system -l k8s-app=kube-dns --tail=100
# Common failure: 
# plugin/ready: this plugin can only be used once
# Corefile:5 - Error during parsing: Unknown directive 'upstream'
2

Restore Known Good CoreDNS ConfigMap & Plugin Directives

Compare production's `coredns` ConfigMap against staging. Edit or restore the valid Corefile configuration, removing deprecated directives while maintaining internal forward resolvers. Trigger a rollout restart of CoreDNS pods.

kubectl edit configmap coredns -n kube-system
# Remove unsupported plugin lines, verify valid Corefile:
# .:53 {
#     errors
#     health
#     ready
#     kubernetes cluster.local in-addr.arpa ip6.arpa {
#         pods insecure
#         fallthrough in-addr.arpa ip6.arpa
#     }
#     forward . /etc/resolv.conf
#     cache 30
#     loop
#     reload
#     loadbalance
# }
kubectl rollout restart deployment coredns -n kube-system
Advertisement
3

Bypass CoreDNS via NodeLocal DNSCache or Direct IP Emergency Override

If CoreDNS binary itself is broken, deploy or fallback to NodeLocal DNSCache. In extreme emergencies where pods cannot resolve databases, temporarily patch critical application Service environment variables or /etc/hosts with ClusterIP addresses while fixing CoreDNS.

Pro Tip: Cluster Resilience: Always run NodeLocal DNSCache in enterprise Kubernetes clusters to provide high-speed local caching and isolate applications from CoreDNS control plane upgrade hiccups.
4

Verify Post-Fix DNS Resolution Across All Namespaces

Spin up a transient debug container to execute DNS lookups across cluster.local, external domains, and reverse in-addr.arpa entries.

kubectl run dns-test --rm -it --image=busybox:1.36 -- nslookup kubernetes.default
kubectl run dns-test --rm -it --image=busybox:1.36 -- nslookup google.com
💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"Production CoreDNS failures post-upgrade stem from undocumented ConfigMap differences and deprecated plugins. Quickly inspect pod logs for syntax rejections, roll back the Corefile ConfigMap, restart coredns pods, and enforce NodeLocal DNSCache."
⚡ 60-Second Elevator Pitch Talking Points
  • Inspect kube-system CoreDNS pod logs to identify rejected Corefile directives.
  • Compare production vs staging coredns ConfigMaps to spot un-migrated plugins.
  • Revert the Corefile ConfigMap to clean, forward-compatible syntax and rollout restart coredns.
  • Deploy NodeLocal DNSCache as a permanent architectural safeguard against cluster-wide DNS failures.
Advertisement
📥 FREE DOWNLOAD · 101-PAGE COMPANION HANDBOOK
Studying for Kubernetes & SRE Technical Rounds?
Download the complete 100-question PDF field guide covering all 11 core modules with offline diagnostic runbooks.
📥 Download PDF (Free) Read Online Guide →
Want more Kubernetes scenarios?
Explore our complete collection of scenario-based Kubernetes interview runbooks.
Browse All Kubernetes Questions →