⚡ ~/naveed Interview Prep
⚡ Portfolio Home ✍️ Engineering Blog Deep Dives 🎯 Interview Hub 998+ Scenarios ☸️ Kubernetes Mastery Hub 24 Modules 🎮 DevOps Arcade & Quizzes Subnet Blitz ⚡ 🗺️ DevOps Roadmaps PDFs & Guides 🤖 Morpheus Analysis AI Quant ↗ 🛠️ Developer Tools Utilities 🧪 Labs & Experiments 📄 Interactive CV & Certs 🔗 All Links & Socials ⚡ Join The Dispatch (Weekly SRE Newsletter) →
Senior DevOps / SRE [L2] Networking Production Scenario [L2]

Q: You're designing a microservices architecture with 50 services. Each service is dynamically deployed by Kubernetes, with IPs changing hourly. How do you enable service discovery so one service can reliably reach another without hardcoding IPs or DNS names?

There are three main service discovery approaches, each with tradeoffs:

#Networking #Networking #L2 #VPC #DNS #Security
🎙️ Candidate Opening & Architectural Context
""Isolating network failures requires proving whether packets are dropped by route tables, security groups, or stateless NACLs. The interviewer is testing: Service discovery patterns, DNS vs API-based discovery, microservices networking.. I structure my answer around systematic triage first, root cause analysis second, and permanent remediation third.""
Advertisement

🛠️ Production Runbook & Step-by-Step Resolution

1️⃣

Production Solution & Architecture

There are three main service discovery approaches, each with tradeoffs: 1. DNS-Based Discovery (Traditional): Kubernetes automatically registers each service in DNS: my-service.default.svc.cluster.local resolves to the service's cluster IP (virtual, stable). Pros: Simple, works with existing apps. Cons: DNS caching issues (Java caches DNS indefinitely), DNS TTL can cause stale endpoints, no realtime updates if a pod crashes mid-request. 2. API-Based Discovery (Service Mesh): Istio/Linkerd intercepts all outbound traffic via sidecar proxies. The proxy dynamically queries the service registry (etcd) for live endpoint lists, updating in realtime as pods scale up/down. Pros: Realtime, handles pod failures gracefully, circuit breaking, retries, mTLS. Cons: Complexity, 5-10% CPU overhead per pod, steep learning curve. 3. Hybrid (DNS + API): Use DNS for initial discovery, but rely on service mesh sidecars for active health checking and load balancing. Recommendation for 50 microservices: Start with Kubernetes native DNS (simplest, lowest overhead): Services call each other: curl http://payment-service:8080/charge. Kubernetes DNS resolves and load balances automatically. If you hit scaling issues (DNS TTL, pod crash recovery), adopt Istio for realtime discovery and traffic management.

Client → '`curl http://my-service:8080/api`' → Kubelet's DNS (CoreDNS) → Service ClusterIP → Round-robin load balance to Pod IPs
💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"Pro-Tip: There are three main service discovery approaches, each with tradeoffs:."
⚡ 60-Second Elevator Pitch Talking Points
  • Immediate Triage: There are three main service discovery approaches, each with tradeoffs:
  • Run targeted verification commands before modifying configuration.
  • Automate permanent guardrails (CI check, alerts, IaC policy) to prevent recurrence.
Advertisement
Want more Networking scenarios?
Explore our complete collection of scenario-based Networking interview runbooks.
Browse All Networking Questions →

📚 Related Production Scenarios in Networking