Q: Pod is running but Service isn't accessible — what's your approach?
Systematic networking approach for when pods report Running/Ready but the Kubernetes ClusterIP/NodePort/LoadBalancer Service cannot be reached.
#Kubernetes #Service #Endpoints #CoreDNS #NetworkPolicy #kube-proxy
🎙️ Candidate Opening & Architectural Context
"When a Pod is Running but its Service isn't accessible, I isolate the failure across 5 discrete layers: Service Selector/Endpoints -> Port mapping -> Readiness probes -> Cluster DNS -> NetworkPolicies."
Advertisement
🛠️ Production Runbook & Step-by-Step Resolution
1️⃣
Step 1: Check Endpoints & EndpointSlices (The #1 Culprit)
Services do not route to Pods directly; they route to Endpoints populated by label matching:
- Run:
kubectl get endpoints <service-name>andkubectl get endpointslices -l kubernetes.io/service-name=<service-name>. - If Endpoints is <none>: The Service's
spec.selectordoes NOT match the Pod's labels! Comparekubectl get svc <svc> -o yamlagainstkubectl get pods --show-labels. - Common typos:
app: webin service vsapp: frontendortier: webon pod.
2️⃣
Step 2: Check Pod Readiness Probes
A pod can be 'Running' but failing its readiness probe:
- Check
kubectl get pods: Is the pod showing0/1 READY? - If a readiness probe fails, Kubernetes removes the pod's IP from the Service Endpoints to prevent traffic from hitting unready pods.
- Check
kubectl describe podfor readiness probe failures.
3️⃣
Step 3: Verify Port & TargetPort Mapping
Confirm the port translation pipeline:
port: 80: The port clients connect to on the Service ClusterIP.targetPort: 8080: The port the container is actually listening on.- Verify the app is listening inside the container:
kubectl exec -it <pod> -- ss -tulpnorcurl localhost:8080. - If
targetPortis a named port (e.g.http), verifycontainerPort: 8080in pod spec matches the name.
4️⃣
Step 4: Test In-Cluster DNS & NetworkPolicies
Spin up a temporary debug pod inside the cluster:
kubectl run curl-test --rm -it --image=curlimages/curl -- sh- Test by IP first:
curl -Iv http://<ClusterIP>:<port>. If IP works, problem is CoreDNS resolution. - Test by FQDN:
curl -Iv http://<service>.<namespace>.svc.cluster.local:<port>. - Check NetworkPolicies: Run
kubectl get netpol. If an ingress default-deny NetworkPolicy exists on the namespace without an allow rule for the client, all traffic is dropped silently at the CNI layer.
💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"Always check 'kubectl get endpoints <svc>' first. If endpoints are empty, it's either a label selector mismatch or a failing readiness probe. Then test port mappings and NetworkPolicies."
⚡ 60-Second Elevator Pitch Talking Points
- Check Endpoints: 'kubectl get endpoints <svc>'. If empty, Service selector doesn't match Pod labels.
- Check Readiness: If pod is 0/1 READY, failing readiness probe stripped pod IP from endpoints.
- Check Port Translation: Verify service port -> targetPort matches the port the container is listening on (ss -tulpn).
- Test via curl container: Test ClusterIP directly, then FQDN (service.ns.svc.cluster.local) to rule out CoreDNS.
- Check NetworkPolicies: 'kubectl get netpol -n <ns>' - verify ingress allow rules exist between namespaces.
Advertisement