Q: Have you worked with Azure Kubernetes Service (AKS)? How would you deploy, monitor, and troubleshoot applications on AKS?
End-to-end operational guide for Azure Kubernetes Service (AKS): control plane vs node pools, Azure CNI vs Kubenet, GitOps deployment with Azure DevOps / GitHub Actions, and diagnosing pod failures via Azure Container Insights.
#Azure #AKS #Azure CNI #Workload Identity #Container Insights #Troubleshooting
🎙️ Candidate Opening & Architectural Context
"Yes, extensively. AKS provides a managed control plane while offloading worker node management into System and User node pools. Managing AKS effectively requires understanding Azure CNI networking, Workload Identity, and Azure Container Insights."
Advertisement
🛠️ Production Runbook & Step-by-Step Resolution
1️⃣
AKS Architecture & Networking Fundamentals
Control plane and networking choices:
- Architecture: Free/Standard tier managed control plane (etcd, API server managed by Microsoft); customer pays only for Virtual Machine Scale Set (VMSS) worker nodes grouped into System (CoreDNS, Metrics Server) and User (microservices) node pools.
- Azure CNI vs Kubenet: Kubenet uses internal pod overlay subnets (saves VNet IPs). Azure CNI gives every pod a real, routable IP from the Azure VNet subnet (lower latency, Direct VNet integration, but requires careful VNet CIDR sizing to avoid IP exhaustion).
2️⃣
Deploying Applications to AKS
Modern automated deployment workflow:
Git Push→GitHub Actions / Azure Pipelines→ACR Docker Build→Workload Identity Auth→ArgoCD / Helm Sync→AKS Cluster
- CI pipeline builds image and pushes to Azure Container Registry (ACR).
- Authenticates to AKS using Microsoft Entra Workload Identity (OIDC federation—no long-lived service principal client secrets stored in CI).
- CD pipeline (or ArgoCD GitOps) renders Helm templates and applies manifests to AKS.
3️⃣
Monitoring & Troubleshooting in AKS
Diagnostic workflow when issues occur:
- Monitoring: Enable Azure Monitor Container Insights with Managed Prometheus and Grafana. Run KQL queries in Log Analytics:
ContainerInventory | where ContainerStatus == 'Failed'. - Troubleshooting Pods: Standard
kubectl describe podandkubectl logs --previous. - AKS Diagnose and Solve Problems: Native Azure Portal blade that runs automated diagnostic checks on node readiness, subnet IP allocation, and API server throttles.
- Node Issues: Check VMSS instance health in Azure Portal or run
az aks check-acrto verify network connectivity between AKS nodes and ACR.
💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"AKS architecture relies on System vs User node pools, Azure CNI for routable VNet networking, and Entra Workload Identity for secretless IAM. Monitor via Azure Monitor Container Insights (Prometheus/Grafana) and troubleshoot using kubectl alongside the Azure 'Diagnose and Solve' blade."
⚡ 60-Second Elevator Pitch Talking Points
- Architecture: Microsoft-managed control plane + VMSS worker node pools (System pool for core add-ons, User pool for workloads).
- Networking: Azure CNI gives pods native VNet IPs; requires large subnets to prevent IP exhaustion.
- Security: Entra Workload Identity federates Kubernetes ServiceAccounts with Azure Managed Identities (zero stored keys).
- Deployment: Azure Pipelines / GitHub Actions -> build & scan image -> push to ACR -> deploy via Helm / ArgoCD.
- Troubleshooting: Use 'kubectl describe/logs' for pod issues; 'az aks check-acr' for registry connectivity; and the Azure Portal 'Diagnose and Solve Problems' blade for node and network health.
Advertisement