Q: You're designing a microservices architecture with 50 services. Each service is dynamically deployed by Kubernetes, with IPs changing hourly. How do you enable service discovery so one service can reliably reach another without hardcoding IPs or DNS names?
There are three main service discovery approaches, each with tradeoffs:
🛠️ Production Runbook & Step-by-Step Resolution
Production Solution & Architecture
There are three main service discovery approaches, each with tradeoffs: 1. DNS-Based Discovery (Traditional): Kubernetes automatically registers each service in DNS: my-service.default.svc.cluster.local resolves to the service's cluster IP (virtual, stable). Pros: Simple, works with existing apps. Cons: DNS caching issues (Java caches DNS indefinitely), DNS TTL can cause stale endpoints, no realtime updates if a pod crashes mid-request. 2. API-Based Discovery (Service Mesh): Istio/Linkerd intercepts all outbound traffic via sidecar proxies. The proxy dynamically queries the service registry (etcd) for live endpoint lists, updating in realtime as pods scale up/down. Pros: Realtime, handles pod failures gracefully, circuit breaking, retries, mTLS. Cons: Complexity, 5-10% CPU overhead per pod, steep learning curve. 3. Hybrid (DNS + API): Use DNS for initial discovery, but rely on service mesh sidecars for active health checking and load balancing. Recommendation for 50 microservices: Start with Kubernetes native DNS (simplest, lowest overhead): Services call each other: curl http://payment-service:8080/charge. Kubernetes DNS resolves and load balances automatically. If you hit scaling issues (DNS TTL, pod crash recovery), adopt Istio for realtime discovery and traffic management.
Client → '`curl http://my-service:8080/api`' → Kubelet's DNS (CoreDNS) → Service ClusterIP → Round-robin load balance to Pod IPs
- Immediate Triage: There are three main service discovery approaches, each with tradeoffs:
- Run targeted verification commands before modifying configuration.
- Automate permanent guardrails (CI check, alerts, IaC policy) to prevent recurrence.