Q: What really happens under the hood when you run 'kubectl apply -f deployment.yaml'? Walk me through the entire journey from client CLI to running container.
The definitive lifecycle of a Kubernetes resource: tracing 'kubectl apply' from client OpenAPI validation, API server authentication & admission webhooks, etcd persistence, controller reconciliation, scheduler scoring, kubelet CRI execution, to CNI network namespace plumbing.
🛠️ Production Runbook & Step-by-Step Resolution
Client-Side Processing & Validation (kubectl)
Before any network request reaches the cluster, kubectl performs client-side inspection:
- Client-Side Validation: kubectl verifies YAML syntax and checks resource fields against the locally cached OpenAPI / Swagger schema from the cluster (
~/.kube/cache/discovery). - Three-Way Strategic Merge Patch: Reads the
kubectl.kubernetes.io/last-applied-configurationannotation, compares it with the live cluster state and the new local YAML, and computes a JSON Strategic Merge Patch. - HTTP REST Request: Formats the payload into JSON and dispatches an HTTP POST/PATCH request to the API server:
POST /apis/apps/v1/namespaces/default/deploymentswith TLS client certificates or OIDC Bearer tokens.
API Server: Authentication, Authorization & Admission Control
The kube-apiserver is the single front door to the cluster. The request traverses 4 sequential filters:
- 1. Authentication: Validates TLS cert, ServiceAccount token, or Entra/AWS IAM OIDC identity to establish the caller's username and groups.
- 2. Authorization (RBAC): Evaluates ClusterRoles/RoleBindings to verify if the subject has
create/patchpermissions ondeploymentsin the target namespace. - 3. Mutating Admission Webhooks: Plugins and webhooks (e.g. Istio sidecar injector, Vault agent injector) modify the object or set default values.
- 4. Schema Validation: Enforces API schema rules, required fields, and immutability constraints.
- 5. Validating Admission Webhooks: Webhooks (e.g. Kyverno, OPA Gatekeeper) run final compliance checks (e.g. enforcing non-root execution or image registry whitelisting) and reject invalid manifests.
etcd: State Persistence & Consensus
The single source of truth commits the change:
- Once admission passes, kube-apiserver serializes the Deployment object into protocol buffers and writes it into etcd at key
/registry/deployments/default/my-app. - etcd commits the write across a quorum of nodes using the Raft consensus algorithm.
- The API server returns an HTTP
201 Createdor200 OKresponse back to the client CLI. - Important: At this exact second, NO containers or pods exist yet! Only the desired state is recorded.
Deployment Controller & ReplicaSet Controller (kube-controller-manager)
The control loops take over asynchronously:
- Deployment Controller: Watches the API server for Deployment changes. Detects the new Deployment and creates a child
ReplicaSetobject with pod template hash. - ReplicaSet Controller: Detects the new ReplicaSet desiring e.g. 3 replicas. Compares desired replicas (3) with existing replicas (0).
- It issues 3 API requests to create 3
Podobjects. Crucially, these Pod objects have no assigned node (spec.nodeName: '') and enter the Pending state.
kube-scheduler: Node Filtering & Scoring
Matching unbound pods to healthy worker nodes:
- Watch Loop: kube-scheduler continuously watches the API server for Pods where
spec.nodeName == ''. - Phase 1 (Filtering / Predicates): Eliminates ineligible nodes that lack sufficient CPU/memory requests, have untolerated taints, or fail nodeSelector / affinity constraints.
- Phase 2 (Scoring / Priorities): Scores remaining nodes based on image locality, topology spread, and resource fragmentation (least/most requested).
- Binding: Scheduler picks the highest-scoring node and sends a
BindingAPI call to the API server, settingspec.nodeName: 'worker-node-2'.
Kubelet & Container Runtime (CRI containerd)
The local node agent brings the container to life:
- Kubelet Watch: The kubelet daemon running on
worker-node-2observes that a Pod has been assigned to its node name. - Container Runtime Interface (CRI): Kubelet calls containerd via gRPC.
- Image Pull: containerd checks if the image exists in local cache; if not, pulls it from registry using node IAM credentials or imagePullSecrets.
- Pod Sandbox Creation: Kubelet instructs containerd to create the Pod Sandbox (pause container) establishing Linux namespaces (IPC, UTS, PID, Network).
CNI Plugin: Network Namespace & Pod IP Allocation
Plumbing the container into the cluster network:
- Kubelet invokes the CNI plugin (AWS VPC CNI, Calico, Flannel, Cilium).
- The CNI creates a virtual ethernet pair (
veth), moves one interface into the pod's network namespace aseth0, and connects the other interface to the host bridge or routing table. - Allocates a dedicated Pod IP from the subnet CIDR and configures routing and MTU.
- Once networking is ready, containerd starts the application containers inside the pod sandbox.
Readiness Probes & Service Endpoints Routing
Connecting the running pod to live client traffic:
- Kubelet continuously executes configured
startupProbeandreadinessProbe. - Once probes return HTTP 200 / Success, kubelet reports pod status as
Readyto the API server. - The Endpoints Controller detects the Ready pod and appends the Pod IP to the Service's
Endpoints/EndpointSlices. - kube-proxy (or Cilium eBPF) on every node updates local iptables/IPVS rules, enabling load balancers and clients to route live traffic to the new pod!
- 1. kubectl: Computes strategic 3-way merge patch and sends HTTP POST to kube-apiserver.
- 2. API Server: Authenticates caller, checks RBAC, runs Mutating/Validating admission webhooks, and writes desired state to etcd via Raft consensus.
- 3. Deployment & ReplicaSet Controllers: Detect the change via API watches and create unbound Pod definitions (spec.nodeName empty).
- 4. kube-scheduler: Filters nodes (predicates) and scores nodes (priorities), then writes a Binding object assigning the pod to a node.
- 5. Kubelet: Detects the assigned pod, instructs CRI (containerd) to create the pause container sandbox, and pulls the image.
- 6. CNI: Configures pod network namespace, veth pair, and assigns the Pod IP.
- 7. Service Ingress: Once readiness probes pass, Endpoints Controller adds Pod IP to Service Endpoints, and kube-proxy updates iptables/IPVS.