⚡ ~/naveed Interview Prep
⚡ Portfolio Home ✍️ Engineering Blog Deep Dives 🎯 Interview Hub 998+ Scenarios ☸️ Kubernetes Mastery Hub 24 Modules 🎮 DevOps Arcade & Quizzes Subnet Blitz ⚡ 🗺️ DevOps Roadmaps PDFs & Guides 🤖 Morpheus Analysis AI Quant ↗ 🛠️ Developer Tools Utilities 🧪 Labs & Experiments 📄 Interactive CV & Certs 🔗 All Links & Socials ⚡ Join The Dispatch (Weekly SRE Newsletter) →
⚡

All 998+ Production DevOps & SRE Interview Scenarios

The complete, battle-tested repository of real-world production incident triage, Kubernetes failure debugging, AWS cloud architecture, Terraform state recovery, and high-scale SRE interview scenarios. Every question includes STAR candidate storytelling answers, CLI runbooks, and 60-second elevator pitches.

⚡ 998 Scenarios ☸️ 11 Domain Tracks 🎯 STAR Method Answers 📋 60s Elevator Pitches 🐧 100% Free & Open
🔍 Live Search & Filter 🎮 Practice Simulator 💼 Real Interview Loops 📘 Preparation Guide

Explore by Technical Specialization

Jump directly into dedicated domain hubs or filter the live list below

11 Domain Hubs Available
☸️ 186 Scenarios

Kubernetes

CrashLoopBackOff, etcd quorum loss, CNI packet drops, HPA metrics server, RBAC & EKS zero-downtime upgrades.

☁️ 214 Scenarios

AWS & Cloud

ALB 504 timeouts, VPC peering & transit gateways, IAM permission boundaries, S3 bucket governance & multi-region DR.

🐳 262 Scenarios

Docker & Containers

Multi-stage build optimization, cgroups v2 limits, rootless daemon security, overlay2 exhaustion & container layer caching.

🔄 125 Scenarios

CI/CD & GitOps

ArgoCD drift reconciliation, Helm rollback failure, GitHub Actions runner autoscaling, semantic release & canary deployments.

🏗️ 135 Scenarios

Terraform & IaC

State locking & corruption, drift detection, dynamic blocks, enterprise workspaces & zero-downtime refactors.

🐧 324 Scenarios

Linux & SRE

High CPU wait & load average, socket leaks, unlinked open file cleanup, eBPF tracing & strace kernel debugging.

📊 98 Scenarios

Observability & SRE

Prometheus high cardinality, Loki ingestion lag, OpenTelemetry trace spans, alert fatigue & SLI/SLO error budgets.

🌐 171 Scenarios

Networking & DNS

CoreDNS 5s throttling, TCP handshake timeouts, MTU mismatch, BGP route flapping & Envoy proxy latency.

🔀 87 Scenarios

Git & Workflows

Interactive rebase conflicts, git bisect incident triage, monorepo submodules & trunk-based delivery.

🛡️ 138 Scenarios

Security & DevSecOps

IRSA least privilege, Falco runtime enforcement, leaked credential rotation, KMS envelope encryption.

💰 85 Scenarios

FinOps & System Design

EKS Karpenter spot autoscaling, Graviton migrations, NAT gateway bill reduction & unit economics.

Advertisement

Live Question Explorer & Practice Arena

Search by keyword, tool, or error code. Toggle Practice Mode to test your active recall.

Practiced: 0 / 998
Showing 40 of 998 scenarios ● Live & Ready for Practice
Senior DevOps / SRE Kubernetes Amazon EKS & Upgrades Classic Scenario

Q) Walk me through how you upgraded an Amazon EKS cluster from Kubernetes v1.34 to v1.36 and even after v1.37 without downtime.

Production runbook strategy for upgrading Amazon EKS clusters across minor versions (v1.34 → v1.35 → v1.36 → v1.37) with zero application downtime using sequential control plane upgrades, blue/green managed node groups, PDBs, and health validation.

#Kubernetes #Amazon EKS #Zero Downtime #Cluster Upgrade
💡 Practice reciting your story first
Candidate Opening: "In one of my projects, we had a customer-facing application running as microservices on Amazon EKS, and we had a requirement to upgrade Kubernetes from v1.34 to v1.36. Since it was Production, we couldn't afford application downtime, so we followed a proper upgrade runbook."
1️⃣

Pre-checks & Compatibility Verification

Before touching the cluster, perform a complete pre-flight check of deprecated APIs and dependencies:

2️⃣

Upgrade One Version at a Time

Kubernetes minor version upgrades must be performed sequentially. You cannot skip minor versions:

3️⃣

Replace Worker Nodes (Blue/Green Node Groups)

For worker nodes, we didn't immediately terminate the existing node group:

4️⃣

How Did We Avoid Downtime?

Zero downtime was guaranteed because our critical microservices were engineered for High Availability:

5️⃣

Continuous Telemetry & Monitoring

Using CloudWatch and Prometheus/Grafana, continuously monitored throughout the upgrade:

6️⃣

Validate Before Removing Anything

Once all workloads were running on the new node group, we didn't immediately remove the old one:

💡 The Key Interview Point (The "Gold Nugget")
"Zero downtime wasn't achieved just because we carefully upgraded EKS. It was possible because the application was already designed for high availability with multiple replicas, PDBs, readiness probes, Multi-AZ deployment, and controlled node draining."
📋 Quick 60-Second Talking Points (Elevator Pitch)
  • Conducted rigorous pre-checks: EKS Upgrade Insights, API deprecations (Pluto), add-on compatibility, and lower-environment testing.
  • Upgraded strictly one minor version at a time (v1.34 → v1.35 → v1.36 → v1.37): control plane first, followed by managed add-ons.
  • Implemented blue/green worker node replacement with new EKS-optimized AMI node groups; cordoned and drained nodes one-by-one.
  • Guaranteed zero downtime via HA safeguards: replica counts ≥ 2, PodDisruptionBudgets, readiness probes, preStop hooks, and multi-AZ spread.
  • Monitored CloudWatch/Grafana telemetry (5xx errors, latency, pending pods, ALB target health) throughout.
  • Executed critical business flow smoke tests before safely decommissioning old node groups.
View Complete Standalone Runbook →
Senior DevOps / SRE AWS Compute & EC2 Production Incident

Q) EC2 CPU suddenly reaches 100% — how would you troubleshoot?

Triage runbook for handling sudden 100% CPU utilization on production EC2 instances: telemetry triage, identifying culprit processes, handling legit vs malicious load, and permanent safeguards.

#AWS #EC2 #CloudWatch #Linux
💡 Practice reciting your story first
Candidate Opening: "When a production EC2 instance hits 100% CPU, my priority is two-fold: stop customer impact immediately (mitigation) while capturing telemetry to pinpoint the exact culprit process (root cause analysis)."
1️⃣

Triage From Outside (CloudWatch & Telemetry)

Before logging in, look at multi-dimensional metrics to classify the nature of the spike:

2️⃣

SSH / SSM Session & Process Inspection

Log into the instance (using AWS Systems Manager Session Manager or SSH) and inspect active processes:

3️⃣

Diagnose the Culprit Process

Examine the identified process to determine what it is doing:

4️⃣

Mitigate and Restore Service

Take decisive, safe actions based on the diagnosis:

💡 The Key Interview Point (The "Gold Nugget")
"Never reboot an instance blindly. Capture thread dumps and top process telemetry first so you don't lose the smoking gun, and ensure auto-scaling and CPU alarms prevent single-node exhaustion."
📋 Quick 60-Second Talking Points (Elevator Pitch)
  • Check CloudWatch: correlate CPU with NetworkIn and ALB traffic to differentiate traffic surge from rogue process.
  • Access instance via SSM Session Manager / SSH; run 'top -c', 'uptime', and 'ps aux --sort=-%cpu | head -10'.
  • Check 'vmstat 1 5' for 'us' (app code) vs 'wa' (I/O wait bottleneck) vs 'sy' (kernel context switching).
  • Inspect threads with 'top -H -p <PID>' and capture thread dump (jstack/pprof) or strace before killing.
  • Mitigate: graceful SIGTERM, scale out ASG, or quarantine if compromised.
  • Prevent: CloudWatch Alarm at 75% CPU, ASG auto-scaling, CPU limits (cgroups/Docker), and query optimization.
View Complete Standalone Runbook →
Senior DevOps / SRE AWS Networking & Access Classic Troubleshooting

Q) EC2 is running but SSH isn't working — what would you check?

Systematic OSI and AWS-layer troubleshooting methodology when an EC2 instance shows 'Running' but SSH connection fails, covering network timeouts, connection refused, key issues, and SSM rescue.

#AWS #EC2 #SSH #Security Groups
💡 Practice reciting your story first
Candidate Opening: "When SSH fails while EC2 shows 'Running', the very first step is to observe the exact error message: 'Connection timed out' means a networking/firewall block, whereas 'Connection refused' or 'Permission denied' means you reached the OS but sshd or auth failed."
1️⃣

Differentiate Network Timeout vs Connection Refused

Run verbose SSH: <code>ssh -vvv -i key.pem user@ip</code> to see where the handshake stalls:

2️⃣

AWS Network & VPC Checks (If Timed Out)

Verify the packet path from internet to instance ENI:

3️⃣

OS, Key Pair & Disk Checks (If Refused or Denied)

Verify host-side configuration and authentication:

4️⃣

Rescue Strategies Without SSH

How senior SREs regain access when SSH is dead:

💡 The Key Interview Point (The "Gold Nugget")
"Categorize the error immediately: 'Timed out' is AWS network/SG; 'Connection refused' is sshd/port; 'Permission denied' is key/username. Always have AWS SSM Session Manager enabled as a zero-SSH out-of-band management backdoor."
📋 Quick 60-Second Talking Points (Elevator Pitch)
  • Check error type: 'Timed out' = network/firewall; 'Connection refused' = sshd dead; 'Permission denied' = key/user error.
  • If timed out: Check Security Group IP whitelist, subnet route table (IGW attached), NACLs (ephemeral return ports).
  • If refused/hung: Check EC2 Instance Status Check, EC2 console screenshot, and system log for kernel panic or 100% disk.
  • If permission denied: Verify key permissions (chmod 400), correct OS username (ec2-user vs ubuntu).
  • Rescue path: Use AWS SSM Session Manager (no port 22 needed), or detach EBS root volume to a rescue instance to fix config.
View Complete Standalone Runbook →
Senior DevOps / SRE AWS Load Balancing High-Severity Incident

Q) ALB starts returning 5xx errors — how would you identify the root cause?

Production runbook for isolating Application Load Balancer (ALB) 5xx errors: distinguishing ELB-generated vs Target-generated codes, debugging 502/503/504, and querying access logs with Athena.

#AWS #ALB #CloudWatch #Target Groups
💡 Practice reciting your story first
Candidate Opening: "When an ALB starts throwing 5xx errors, my immediate first goal is to differentiate between errors generated by the ALB itself versus errors generated by backend application targets."
1️⃣

CloudWatch Metrics Differentiation (ELB vs Target)

Look at CloudWatch metrics for the ALB immediately:

2️⃣

Diagnose Specific 5xx Error Codes

Apply targeted root cause analysis based on status code:

3️⃣

Query ALB Access Logs with Amazon Athena

Analyze ALB S3 access logs to extract exact failing requests, targets, and response times:

4️⃣

Remediate and Safeguard

Immediate mitigation and permanent safeguards:

💡 The Key Interview Point (The "Gold Nugget")
"Always split ELB 5xx from Target 5xx in CloudWatch first. If it's ELB 502, check Keep-Alive timeout mismatches; if 503, check Target Group healthy host count; if 504, check target processing latency."
📋 Quick 60-Second Talking Points (Elevator Pitch)
  • Check CloudWatch: HTTPCode_ELB_5XX (ALB fault) vs HTTPCode_Target_5XX (app fault).
  • If Target 5XX: App is throwing 500s; inspect app logs and database connectivity.
  • If ELB 502: Target closed TCP prematurely; fix Keep-Alive timeout (backend keep-alive must exceed ALB 60s).
  • If ELB 503: Zero healthy targets in Target Group; verify health check path (/healthz) and target capacity.
  • If ELB 504: Gateway timeout; target query took >60s. Check target_processing_time and database locks.
  • Query ALB Access Logs via Athena to isolate failing URLs, client IPs, and specific target instance IDs.
View Complete Standalone Runbook →
Senior DevOps / SRE AWS VPC & Networking Networking & Security

Q) Application works internally but not from the internet — how would you troubleshoot?

Outside-in OSI network layer troubleshooting model for resolving applications reachable internally on private VPC IPs but inaccessible over the public internet.

#AWS #VPC #Route 53 #Internet Gateway
💡 Practice reciting your story first
Candidate Opening: "Since the application works internally, the compute instance and software service are healthy. The failure is strictly along the ingress network path between the public internet and the AWS VPC."
1️⃣

DNS & IP Resolution (Outside-in Step 1)

Verify public DNS mapping from an external workstation:

2️⃣

Entry Point Topology & Subnet Routing (Step 2)

Verify the Internet Gateway and subnet routing tables:

3️⃣

Firewalls: Security Groups & NACLs (Step 3)

Inspect AWS stateful and stateless firewall layers:

4️⃣

AWS WAF & Target Group Binding (Step 4)

Check perimeter protection and target routing:

💡 The Key Interview Point (The "Gold Nugget")
"Trace outside-in: DNS -> Internet Gateway & Route Table -> Public Subnet ALB -> Security Group (80/443 from 0.0.0.0/0) -> NACLs (ephemeral return) -> WAF. Internal working proves compute is fine; focus purely on the AWS ingress pipeline."
📋 Quick 60-Second Talking Points (Elevator Pitch)
  • Check DNS: dig app.example.com to verify public resolution to ALB CNAME / Elastic IP.
  • Check Route Table: ensure ALB is in public subnets with route 0.0.0.0/0 -> Internet Gateway (IGW).
  • Check Security Groups: ALB SG must allow 80/443 from 0.0.0.0/0; instance SG must allow traffic from ALB SG.
  • Check NACLs: ensure stateless NACLs permit both inbound 80/443 and outbound ephemeral ports (1024-65535).
  • Check AWS WAF: inspect blocked requests in WAF console for geo-blocking or rate-limit blocks.
  • Use AWS VPC Reachability Analyzer to mathematically prove the network path between IGW and instance.
View Complete Standalone Runbook →
Senior DevOps / SRE Kubernetes Workloads & Scheduling Core K8s Scenario

Q) Pod is stuck in Pending — what would you check?

Step-by-step diagnostic workflow for pods stuck in Pending state: decoding kube-scheduler events, capacity exhaustion, taints/tolerations, PVC binding, and autoscaling response.

#Kubernetes #Scheduler #kubectl describe #Resource Limits
💡 Practice reciting your story first
Candidate Opening: "A Pod stuck in Pending means the kube-scheduler cannot find a node that meets all the pod's constraints, or volume mounting / admission controllers are blocked. My primary tool is immediately 'kubectl describe pod <pod-name>'."
1️⃣

Inspect Scheduler Events First

Run <code>kubectl describe pod &lt;pod-name&gt;</code> and look directly at the <strong>Events</strong> section at the bottom:

2️⃣

Root Cause 1: Insufficient Node Capacity (CPU/Memory)

Scheduler calculates fit based on <strong>requests</strong>, not actual usage:

3️⃣

Root Cause 2: Node Selectors, Affinity & Taints

Filter constraints that eliminate eligible nodes:

4️⃣

Root Cause 3: Unbound PVCs & Missing Config

Check persistent storage and required configuration:

💡 The Key Interview Point (The "Gold Nugget")
"Always run 'kubectl describe pod' and read the Scheduler Events. Differentiate resource request starvation (fit calculation) from affinity/taint rules and AZ-locked EBS volume binding."
📋 Quick 60-Second Talking Points (Elevator Pitch)
  • Run 'kubectl describe pod <name>' and read the Events section: FailedScheduling tells you why.
  • Check Resources: verify node allocated requests ('kubectl describe nodes') vs pod spec.resources.requests.
  • Check Constraints: nodeSelector, nodeAffinity, and taints/tolerations that block placement.
  • Check Storage: verify PVC status ('kubectl get pvc') - check if volume is locked to a different AWS AZ.
  • Verify Cluster Autoscaler / Karpenter logs to ensure nodes are actively spinning up.
  • Fix: Tune requests, correct label selectors, add tolerations, or trigger node autoscaling.
View Complete Standalone Runbook →
Senior DevOps / SRE Kubernetes Troubleshooting & Debugging Core K8s Scenario

Q) Pod is in CrashLoopBackOff — how would you troubleshoot?

Exhaustive triage process for debugging CrashLoopBackOff: decoding container exit codes (137 OOMKill, 1 app error, 143 SIGTERM), fetching previous container logs, and fixing probe failures.

#Kubernetes #CrashLoopBackOff #OOMKilled #kubectl logs
💡 Practice reciting your story first
Candidate Opening: "CrashLoopBackOff means the container started, crashed, and Kubernetes is backing off before restarting it. My troubleshooting sequence always follows: Exit Code analysis -> Previous container logs -> Liveness probe evaluation."
1️⃣

Inspect Exit Code via kubectl describe

Run <code>kubectl describe pod &lt;pod-name&gt;</code> and examine the <code>Last State: Terminated</code> block:

2️⃣

Fetch Previous Container Logs (--previous)

Because the crashing container has already terminated, standard logs may be empty or only show startup lines:

3️⃣

Check Liveness & Startup Probes

A misconfigured liveness probe will actively kill a healthy container during slow startup:

4️⃣

Interactive Debugging with Ephemeral Containers

If logs are silent and container crashes instantly:

💡 The Key Interview Point (The "Gold Nugget")
"Look at the Exit Code first (137 = OOM, 1 = App crash). Use 'kubectl logs --previous' to capture the crash stack trace. Verify startupProbe isn't killing slow-starting applications."
📋 Quick 60-Second Talking Points (Elevator Pitch)
  • Run 'kubectl describe pod <name>' -> inspect Last State: Exit Code (137 = OOMKill, 1 = code exception, 127 = binary missing).
  • Run 'kubectl logs <name> --previous' to retrieve the stack trace from the crashed instance.
  • If Exit Code 137: Increase spec.resources.limits.memory or profile memory leak.
  • Check Probes: Verify livenessProbe isn't timing out during slow boots; add startupProbe.
  • If container crashes instantly: Override command with 'sleep 3600' or attach ephemeral container via 'kubectl debug'.
View Complete Standalone Runbook →
Senior DevOps / SRE Kubernetes Services & Networking Core K8s Scenario

Q) Pod is running but Service isn't accessible — what's your approach?

Systematic networking approach for when pods report Running/Ready but the Kubernetes ClusterIP/NodePort/LoadBalancer Service cannot be reached.

#Kubernetes #Service #Endpoints #CoreDNS
💡 Practice reciting your story first
Candidate Opening: "When a Pod is Running but its Service isn't accessible, I isolate the failure across 5 discrete layers: Service Selector/Endpoints -> Port mapping -> Readiness probes -> Cluster DNS -> NetworkPolicies."
1️⃣

Step 1: Check Endpoints & EndpointSlices (The #1 Culprit)

Services do not route to Pods directly; they route to Endpoints populated by label matching:

2️⃣

Step 2: Check Pod Readiness Probes

A pod can be 'Running' but failing its readiness probe:

3️⃣

Step 3: Verify Port & TargetPort Mapping

Confirm the port translation pipeline:

4️⃣

Step 4: Test In-Cluster DNS & NetworkPolicies

Spin up a temporary debug pod inside the cluster:

💡 The Key Interview Point (The "Gold Nugget")
"Always check 'kubectl get endpoints <svc>' first. If endpoints are empty, it's either a label selector mismatch or a failing readiness probe. Then test port mappings and NetworkPolicies."
📋 Quick 60-Second Talking Points (Elevator Pitch)
  • Check Endpoints: 'kubectl get endpoints <svc>'. If empty, Service selector doesn't match Pod labels.
  • Check Readiness: If pod is 0/1 READY, failing readiness probe stripped pod IP from endpoints.
  • Check Port Translation: Verify service port -> targetPort matches the port the container is listening on (ss -tulpn).
  • Test via curl container: Test ClusterIP directly, then FQDN (service.ns.svc.cluster.local) to rule out CoreDNS.
  • Check NetworkPolicies: 'kubectl get netpol -n <ns>' - verify ingress allow rules exist between namespaces.
View Complete Standalone Runbook →
Senior DevOps / SRE Kubernetes Deployment Strategies & Rollbacks Incident Recovery

Q) New deployment breaks production — how would you rollback?

Fast-response production incident runbook for rolling back broken Kubernetes deployments across native kubectl, GitOps (ArgoCD/Flux), and handling database migration traps.

#Kubernetes #Deployment #Rollback #GitOps
💡 Practice reciting your story first
Candidate Opening: "When a new deployment causes production degradation, the priority is mean time to recovery (MTTR). The rollback path depends on whether you run imperative deployments or declarative GitOps."
1️⃣

Immediate Imperative Rollback (kubectl rollout undo)

If using native Kubernetes deployments:

2️⃣

GitOps Rollback (ArgoCD / Flux Reality)

In GitOps, manual kubectl rollouts will be reverted by self-healing:

3️⃣

The Database Migration Trap

Can code be safely rolled back if database schema migrated?

4️⃣

Post-Incident & Blameless Post-Mortem

Preventing the same failure in future releases:

💡 The Key Interview Point (The "Gold Nugget")
"Know your deployment mechanism: In pure K8s, use 'kubectl rollout undo'; in GitOps (ArgoCD), disable auto-sync or 'git revert' to prevent self-heal fighting. Never do destructive DB migrations in single-step releases."
📋 Quick 60-Second Talking Points (Elevator Pitch)
  • Immediate action: Run 'kubectl rollout undo deployment/<name>' to revert to previous ReplicaSet.
  • If GitOps (ArgoCD/Flux): Disable auto-sync immediately or 'git revert HEAD && git push' so self-heal doesn't re-break it.
  • Database check: Verify if DB migrations ran. If additive, rollback is safe. If destructive, apply compensating migration.
  • Verify recovery: Monitor 'kubectl rollout status', ALB 5xx metrics, and application logs.
  • Post-mortem: Implement progressive delivery (Argo Rollouts/Canary) with automatic metric-based rollback.
View Complete Standalone Runbook →
Senior DevOps / SRE CI/CD Pipelines & Delivery Pipeline Triage

Q) Pipeline succeeds but the new version isn't deployed — how would you debug?

Diagnostic guide for when CI/CD logs show green checkmarks but the target production environment continues running old code: image tag caching, branch rules, and GitOps sync gaps.

#CI/CD #Docker #Image Tagging #GitOps
💡 Practice reciting your story first
Candidate Opening: "When a CI/CD pipeline shows green but production is unchanged, the issue lies in artifact immutability, deployment trigger conditions, or GitOps reconciliation gaps."
1️⃣

Verify What is Actually Running in Production

Start by inspecting the live cluster/server before checking pipeline scripts:

2️⃣

The ':latest' Tag & imagePullPolicy Trap

The single most common root cause in container CI/CD:

3️⃣

Pipeline Step Conditions & Environment Mismatch

Audit pipeline execution steps:

4️⃣

GitOps Manifest Repo & Controller Audit

If using a separate manifest repository (ArgoCD / Flux):

💡 The Key Interview Point (The "Gold Nugget")
"Check the live running image tag first. Avoid mutable ':latest' tags that bypass k8s rollouts. Verify pipeline conditions, approval gates, and GitOps manifest commit chains."
📋 Quick 60-Second Talking Points (Elevator Pitch)
  • Verify running container image: 'kubectl get deploy <app> -o jsonpath={..image}'.
  • Check image tagging: Avoid ':latest' with 'imagePullPolicy: IfNotPresent' which ignores new image pushes.
  • Audit CI logs: Ensure the deploy job actually executed and was not skipped by branch/tag conditions.
  • Check GitOps repo: Verify CI successfully pushed new image tag commit to the manifest repository.
  • Check ArgoCD/Flux: Inspect sync status, controller errors, or paused auto-sync.
View Complete Standalone Runbook →
Senior DevOps / SRE CI/CD Release Engineering Critical Incident

Q) Production deployment fails halfway — what's your rollback strategy?

Engineering strategy for recovering from halfway-failed deployments: rolling update halts, automated canary rollback, and protecting stateful databases.

#CI/CD #Rollback #Database Migrations #Canary
💡 Practice reciting your story first
Candidate Opening: "A deployment failing halfway means your system is in a hybrid state — some nodes or pods are running the new version while others run the old version. The strategy depends on deployment pattern and data compatibility."
1️⃣

Stop Forward Rollout & Contain Traffic

Halt the release immediately before more instances degrade:

2️⃣

Execute Safe Rollback

Restore all instances to the previous known-good state:

3️⃣

Verify Database & Stateful Dependencies

The most dangerous aspect of partial failures:

4️⃣

Automated Guardrails to Prevent Recurrence

Architectural improvements for future releases:

💡 The Key Interview Point (The "Gold Nugget")
"Halt forward rollout immediately. Revert traffic to the stable baseline. Never couple breaking database migrations with application deployments — use Expand/Contract."
📋 Quick 60-Second Talking Points (Elevator Pitch)
  • Halt rollout immediately: Abort canary or rolling update to prevent further degradation.
  • Revert instances: Use 'kubectl rollout undo' or flip traffic router back to 100% stable version.
  • Check Database: Verify if migrations ran; ensure schema changes follow Expand/Contract so old code remains functional.
  • Validate: Confirm healthy targets in ALB and verify application error metrics stabilize.
  • Prevent: Introduce automated Canary analysis (Argo Rollouts/Flagger) with automated rollback thresholds.
View Complete Standalone Runbook →
Senior DevOps / SRE CI/CD Security & Compliance DevSecOps

Q) How would you securely manage secrets?

Enterprise secrets management architecture: eliminating secrets from Git and CI logs, integrating HashiCorp Vault / AWS Secrets Manager with Kubernetes, and using SOPS/External Secrets Operator.

#Security #HashiCorp Vault #AWS Secrets Manager #SOPS
💡 Practice reciting your story first
Candidate Opening: "Secure secrets management requires defense-in-depth across three stages: in Git repositories (at rest), during CI/CD execution (in transit), and inside production runtime environments."
1️⃣

Pillar 1: Never Commit Plaintext Secrets to Git

Enforce strict pre-commit and repository-level prevention:

2️⃣

Pillar 2: Centralized Secrets Store of Record

Use dedicated secret management services:

3️⃣

Pillar 3: Runtime Injection into Kubernetes

How applications consume secrets securely in production:

4️⃣

Pillar 4: CI/CD Pipeline Hygiene

Protecting secrets during automated builds:

💡 The Key Interview Point (The "Gold Nugget")
"Store secrets in AWS Secrets Manager / Vault, inject into K8s via External Secrets Operator, encrypt in Git using SOPS, use OIDC for CI/CD instead of static keys, and encrypt etcd at rest."
📋 Quick 60-Second Talking Points (Elevator Pitch)
  • Repository level: Block secrets with TruffleHog / pre-commit hooks; use SOPS/KMS if storing in GitOps repos.
  • Storage of record: AWS Secrets Manager / HashiCorp Vault with automated rotation and CloudTrail audit logging.
  • Kubernetes integration: Use External Secrets Operator (ESO) or Secrets Store CSI driver; enable etcd encryption-at-rest.
  • CI/CD pipeline: Eliminate static AWS keys using OIDC federated IAM roles; mask all secrets in logs.
  • Access control: Enforce strict RBAC on K8s secrets; disable automountServiceAccountToken where not needed.
View Complete Standalone Runbook →
Senior DevOps / SRE CI/CD Deployment Strategies Architecture & CD

Q) How would you implement Blue-Green/Canary deployment?

Architectural blueprints for Blue-Green and Canary deployments: traffic routing mechanics, automated metric verification, database migration considerations, and rollback triggers.

#CI/CD #Canary #Blue-Green #Argo Rollouts
💡 Practice reciting your story first
Candidate Opening: "Both strategies aim for zero downtime, but they serve different risk profiles: Blue-Green is all-or-nothing environment switching, while Canary is progressive, metric-driven traffic shifting."
1️⃣

Implementing Canary Deployments (Progressive Delivery)

Deploy new version alongside stable, route a small percentage of real traffic:

2️⃣

Implementing Blue-Green Deployments (Environment Isolation)

Maintain two identical production environments (Blue = Live, Green = Idle):

3️⃣

Database Schema Strategy for Both

How to handle databases when two different application versions run simultaneously:

💡 The Key Interview Point (The "Gold Nugget")
"Canary uses Argo Rollouts/Flagger for metric-driven progressive traffic shifting (5% -> 100%). Blue-Green uses dual environments with a router flip. Both require backward-compatible Expand/Contract database schemas."
📋 Quick 60-Second Talking Points (Elevator Pitch)
  • Canary: Progressive rollout using Argo Rollouts or Flagger. Incremental traffic shift (5% -> 20% -> 100%) with Prometheus metric analysis for auto-rollback.
  • Blue-Green: Dual identical environments. Deploy to Green, run smoke tests, flip Service selector or ALB Target Group weight, keep Blue alive for fast rollback.
  • Traffic Routing: Managed at Ingress/Service Mesh layer (Istio, NGINX Ingress, or AWS ALB weighted target groups).
  • Database requirement: Backward-compatible Expand/Contract schema migrations to support both versions running concurrently.
View Complete Standalone Runbook →
Senior DevOps / SRE Terraform State & Drift IaC Troubleshooting

Q) terraform plan shows unexpected changes — what would you investigate?

Deep investigation runbook for unexpected terraform plan diffs: distinguishing configuration drift, provider upgrade defaults, dynamic attributes, and preventing accidental destruction.

#Terraform #State Drift #Plan Diff #AWS Provider
💡 Practice reciting your story first
Candidate Opening: "When 'terraform plan' shows unexpected changes — especially destruction or in-place modifications — my absolute first rule is: STOP. Do not run 'terraform apply'. Investigate the diff systematically."
1️⃣

Analyze the Diff & Action Symbols

Read the exact plan output carefully:

2️⃣

Check for Out-of-Band State Drift

Did someone modify the cloud resource manually in the AWS Console or CLI?

3️⃣

Provider Upgrades & Variable Drift

Inspect underlying dependencies:

4️⃣

Resolve and Prevent Recurrence

How to resolve safely:

💡 The Key Interview Point (The "Gold Nugget")
"Never apply an unexpected plan. Look for '# forces replacement' flags. Run 'terraform plan -refresh-only' to isolate cloud drift from code changes, and check CloudTrail for out-of-band console edits."
📋 Quick 60-Second Talking Points (Elevator Pitch)
  • Freeze: Do not apply. Inspect plan diff for '~' (in-place) vs '-/+' (destroy and recreate - forces replacement).
  • Isolate drift: Run 'terraform plan -refresh-only' to identify out-of-band manual changes in AWS Console.
  • Check CloudTrail: Identify who modified the resource outside Terraform and when.
  • Check provider & variables: Check '.terraform.lock.hcl' for provider upgrades and verify correct .tfvars file.
  • Use moved blocks: If refactoring, use 'moved' blocks to prevent destroy/recreate cycles.
  • Safeguards: Add 'lifecycle { prevent_destroy = true }' on critical production resources.
View Complete Standalone Runbook →
Senior DevOps / SRE Terraform Governance & Drift IaC Governance

Q) Someone manually changes Terraform-managed infrastructure — what happens?

Exhaustive breakdown of configuration drift in Terraform: what happens during refresh, how Terraform resolves divergence, and how to eliminate console access in production.

#Terraform #Drift Detection #CloudTrail #State
💡 Practice reciting your story first
Candidate Opening: "This is known as 'Configuration Drift'. What happens depends on whether the resource was modified, added, or deleted, and when the next Terraform execution runs."
1️⃣

What Happens During Next Terraform Run

During the next <code>terraform plan</code> or <code>apply</code>, Terraform performs a refresh:

2️⃣

Two Remediation Paths

The team must choose between restoring code or adopting the change:

3️⃣

How to Prevent It from Happening Again

Enterprise governance safeguards:

💡 The Key Interview Point (The "Gold Nugget")
"On next run, Terraform detects drift during refresh and proposes reverting the manual change back to match code. Prevent it by removing AWS console write access and running automated daily drift detection in CI."
📋 Quick 60-Second Talking Points (Elevator Pitch)
  • What happens: On next 'terraform plan', refresh detects divergence between cloud reality and state. Terraform proposes reverting manual edits to match code.
  • If resource deleted manually: Terraform proposes recreating it.
  • Remediation: Either 'terraform apply' to overwrite manual changes, or update code to match reality and commit.
  • Root prevention: Restrict IAM - zero human write access in Production AWS Console; only CI/CD IAM role has write access.
  • Detection: Schedule daily automated 'terraform plan -detailed-exitcode' in CI; alert team on exit code 2 (drift).
View Complete Standalone Runbook →
Senior DevOps / SRE Terraform State & Collaboration Architecture & Best Practices

Q) How would you manage state for multiple engineers?

Production-grade remote backend architecture for team collaboration: S3 state storage, DynamoDB state locking, encryption, state isolation, and Atlantis/CI execution.

#Terraform #S3 #DynamoDB #State Locking
💡 Practice reciting your story first
Candidate Opening: "Managing Terraform state across a team requires eliminating local state files completely. We use a secure remote backend with distributed locking, encryption, versioning, and execution isolation."
1️⃣

AWS S3 + DynamoDB Remote Backend

The industry standard remote backend configuration:

2️⃣

State Decomposition & Blast Radius Reduction

Never put all infrastructure into a single monolithic state file:

3️⃣

Execute Through CI/CD (Atlantis / Terraform Cloud)

Engineers should not run 'apply' from local laptops:

4️⃣

State Security & Sensitive Data

Protecting secrets inside state:

💡 The Key Interview Point (The "Gold Nugget")
"Use S3 with versioning and KMS encryption for state storage, DynamoDB for distributed locking, split state by layer and environment, and execute applies exclusively through CI/CD (Atlantis/PR workflow)."
📋 Quick 60-Second Talking Points (Elevator Pitch)
  • Remote Backend: S3 bucket with versioning (enables rollback of corrupt state) and KMS encryption.
  • State Locking: DynamoDB table with LockID key to prevent concurrent apply race conditions.
  • Decomposition: Split state by layer (network, compute, database) and environment (dev, staging, prod) to limit blast radius.
  • Access Control: S3 bucket policy allowing only CI/CD IAM role; block public access (state stores plaintext secrets).
  • Team Workflow: Run plan/apply via Atlantis or CI/CD PR workflow rather than individual local laptops.
View Complete Standalone Runbook →
Senior DevOps / SRE Terraform Repository Architecture Enterprise Architecture

Q) How would you structure dev/staging/prod?

Enterprise architectural comparison: Directory-based separation with reusable modules vs Terraform Workspaces, multi-account AWS strategy, and code reuse.

#Terraform #Architecture #Directory Layout #Workspaces
💡 Practice reciting your story first
Candidate Opening: "For enterprise production infrastructure, the recommended pattern is Directory-Based Separation using Reusable Modules across Multi-Account AWS environments, rather than Terraform Workspaces."
1️⃣

Why Directory-Based Separation Over Workspaces

Understanding the fundamental tradeoff:

2️⃣

Recommended Repository Layout

Standardized modular directory structure:

3️⃣

Multi-Account AWS Strategy

Maximum security and blast radius isolation:

4️⃣

Version Pinning & Promotion Workflow

How code moves from Dev to Prod:

💡 The Key Interview Point (The "Gold Nugget")
"Use directory-based separation with version-pinned reusable modules and separate AWS accounts per environment. Avoid workspaces for multi-environment cloud infrastructure due to shared backend risks."
📋 Quick 60-Second Talking Points (Elevator Pitch)
  • Architecture: Directory-based separation with reusable modules (avoid workspaces for dev/prod due to shared backend blast radius).
  • Multi-Account: Separate AWS accounts for Dev, Staging, and Prod under AWS Organizations.
  • State Isolation: Distinct S3 backend keys and DynamoDB locks per environment.
  • Module Versioning: Environments source modules pinned to immutable Git tags (ref=v1.2.0).
  • Promotion: Changes are applied to Dev first, validated in Staging, and promoted to Prod via CI/CD approval.
View Complete Standalone Runbook →
Senior DevOps / SRE Linux Performance & SRE Core SRE Methodology

Q) Production server becomes slow — how do you identify CPU/memory/disk/network issues?

Mastering the USE Method (Utilization, Saturation, Errors) to isolate system bottlenecks across CPU, Memory, Disk I/O, and Network in under 60 seconds.

#Linux #USE Method #vmstat #iostat
💡 Practice reciting your story first
Candidate Opening: "When a production server becomes slow, I apply Brendan Gregg's USE Method (Utilization, Saturation, Errors) using a systematic 60-second command checklist to isolate which subsystem is bottlenecking."
1️⃣

First 15 Seconds: System Overview

Assess overall system pressure:

2️⃣

CPU vs Disk I/O: vmstat 1 5

The most informative single command in Linux:

3️⃣

Memory Check: free -h & vmstat swap

Inspect real available memory and paging:

4️⃣

Disk I/O & Network Deep Dive

Isolate storage or packet bottlenecks:

💡 The Key Interview Point (The "Gold Nugget")
"Use the USE Method. Run 'uptime' -> 'vmstat 1 5' -> 'free -h' -> 'iostat -xz 1 5'. High 'wa' in vmstat means Disk I/O bottleneck; high 'si/so' means memory swapping; high 'r' means CPU saturation."
📋 Quick 60-Second Talking Points (Elevator Pitch)
  • Step 1: 'uptime' -> compare 1/5/15m load average against 'nproc' core count.
  • Step 2: 'vmstat 1 5' -> check 'r' (CPU saturation), 'b' (I/O blocked), and 'wa' (if iowait is high, bottleneck is disk, not CPU).
  • Step 3: 'free -h' -> check 'available' memory; check 'si/so' in vmstat for swap thrashing.
  • Step 4: 'iostat -xz 1 5' -> check '%util' and 'await' response times for storage bottlenecks.
  • Step 5: 'sar -n DEV 1 5' and 'ss -s' -> check network throughput, dropped packets, and socket saturation.
  • Step 6: 'dmesg -T | tail -30' for OOM kills and hardware errors.
View Complete Standalone Runbook →
Senior DevOps / SRE Linux Storage & Filesystems Production Incident

Q) Disk reaches 100% — how would you find and safely remove the cause?

Production runbook for recovering from 100% disk utilization: inode exhaustion, du vs df discrepancies, the unlinked open file descriptor trap (lsof deleted), and safe log truncation.

#Linux #df #du #lsof
💡 Practice reciting your story first
Candidate Opening: "When a disk hits 100%, services begin dropping connections, crash, and cannot write logs or create PID files. The investigation requires checking both disk blocks and inodes, and avoiding the classic 'deleted file still held open' trap."
1️⃣

Check Mounts and Inodes (df -h & df -i)

Identify which filesystem is full and whether it is bytes or inodes:

2️⃣

Locate Largest Directories and Files (du)

Scan the filesystem safely without crossing mount boundaries:

3️⃣

The 'Deleted File Held Open' Trap (lsof +L1)

The #1 issue senior SREs look for:

4️⃣

Safe Cleanup Best Practices

How to free space immediately without breaking services:

💡 The Key Interview Point (The "Gold Nugget")
"Check 'df -i' for inode exhaustion. Use 'lsof +L1' to find deleted files held open by running processes, and truncate them via /proc/<PID>/fd/<FD> rather than rebooting."
📋 Quick 60-Second Talking Points (Elevator Pitch)
  • Check blocks & inodes: 'df -h' for disk space, 'df -i' for inode exhaustion (millions of tiny files).
  • Find large files: 'du -ahx / | sort -rh | head -20' (use -x to avoid traversing other mounts/NFS).
  • Check deleted files held open: 'lsof +L1' or 'lsof | grep deleted' for deleted files still locked by active processes.
  • Fix open deleted files: Truncate via '> /proc/<PID>/fd/<FD>' to free disk space immediately without restarting process.
  • Clean safely: Truncate logs (truncate -s 0 file.log) rather than 'rm'; run 'journalctl --vacuum-size=500M'.
  • Prevent: Set up logrotate, CloudWatch disk space alarm at 80%, and automated cleanup jobs.
View Complete Standalone Runbook →
Senior DevOps / SRE Linux Performance & Processes Core Linux Skills

Q) Process is consuming high CPU — which commands would you use?

Full command toolkit for diagnosing high CPU processes: thread-level drilldown (top -H), syscall tracing (strace), kernel profiling (perf), and graceful mitigation.

#Linux #top #ps #strace
💡 Practice reciting your story first
Candidate Opening: "When a process consumes high CPU, my goal is to drill down from the system process level down to the exact thread, system call, or application function causing the burn."
1️⃣

Identify Process & System Context

Find the PID and understand its resource footprint:

2️⃣

Thread-Level Inspection (top -H)

Multi-threaded runtimes (Java, Go, C++, Node worker threads) distribute load across threads:

3️⃣

Trace Syscalls & Profile CPU (strace / perf)

Determine if CPU is burned in user code or kernel syscalls:

4️⃣

Mitigate and Control Priority

Managing the process safely:

💡 The Key Interview Point (The "Gold Nugget")
"Use 'top -c' to find PID, 'top -H -p <PID>' to find the specific thread, and 'strace -c' or 'perf top' to pinpoint the looping function or lock contention."
📋 Quick 60-Second Talking Points (Elevator Pitch)
  • Find process: 'top -c' (sort by P) or 'ps aux --sort=-%cpu | head -10' or 'pidstat 1 5 -u'.
  • Inspect threads: 'top -H -p <PID>' to find the exact thread ID (TID) burning CPU.
  • Trace syscalls: 'strace -p <PID> -c' to see time spent in kernel syscalls vs user code.
  • Profile CPU: 'perf top -p <PID>' or capture thread dumps (jstack/pprof) before killing.
  • Control: 'renice +10 <PID>' to deprioritize, or 'kill -15 <PID>' for graceful termination.
  • Prevent: Set cgroup / systemd CPUQuota limits and container resource limits.
View Complete Standalone Runbook →
Senior DevOps / SRE Linux Networking & Daemons Core Linux Networking

Q) Service is running but not listening on the expected port — how would you troubleshoot?

Structured diagnostic sequence for services running in systemd but failing port checks: 127.0.0.1 vs 0.0.0.0 binding, privileged ports, firewall drops, and SELinux blocks.

#Linux #ss #netstat #systemd
💡 Practice reciting your story first
Candidate Opening: "When systemctl reports 'active (running)' but you cannot connect to the expected port, I trace through 5 discrete checkpoints: daemon listening sockets, IP binding address, privileged port permissions, firewall rules, and security modules."
1️⃣

Verify Open Sockets (ss / lsof / netstat)

What ports is the process actually listening on?

2️⃣

The IP Binding Address Trap (127.0.0.1 vs 0.0.0.0)

The #1 configuration trap in service setup:

3️⃣

Privileged Port & Permission Issues (Ports < 1024)

Linux security restrictions on well-known ports:

4️⃣

Firewall (iptables/nftables) & SELinux

Host-level security layers blocking incoming packets:

💡 The Key Interview Point (The "Gold Nugget")
"Check 'ss -tulpn | grep <PID>' first. If listening on 127.0.0.1, change binding to 0.0.0.0. If port < 1024, check CAP_NET_BIND_SERVICE. If bound correctly, check iptables and SELinux port policies."
📋 Quick 60-Second Talking Points (Elevator Pitch)
  • Check open sockets: 'ss -tulpn | grep <process>' or 'lsof -i :<port>' to verify if socket exists.
  • Check binding address: Verify service is bound to 0.0.0.0 (all interfaces), NOT 127.0.0.1 (localhost only).
  • Check logs: 'journalctl -u <service> -e' for bind errors, port conflicts, or permission denied.
  • Check privileged ports (<1024): Non-root users cannot bind to 80/443 without 'AmbientCapabilities=CAP_NET_BIND_SERVICE' in systemd.
  • Check firewall: Inspect 'iptables -L -n -v' or 'ufw status' for incoming port drops.
  • Check SELinux/AppArmor: Look for AVC denials in 'audit.log'; assign port context via 'semanage port'.
View Complete Standalone Runbook →
Staff / Principal SRE System Design Cloud Architecture & Scalability Premium Architecture

Q) Design a highly available, production-grade cloud infrastructure for a microservices application handling millions of requests per day. Explain your choices around networking, load balancing, autoscaling, databases, observability, security, and disaster recovery.

Full architectural blueprint for multi-AZ, production-grade microservices handling millions of daily requests: edge routing, Kubernetes compute with Karpenter, Aurora/Redis persistence, OpenTelemetry observability, and DR.

#System Design #AWS #EKS #Architecture
💡 Practice reciting your story first
Candidate Opening: "To support millions of daily requests with 99.99% availability, the architecture is designed around multi-AZ redundancy, zero single points of failure, decoupling of state, automated progressive scaling, and zero-trust security."
1️⃣

Networking & Edge Load Balancing

Multi-layered ingress and defense-in-depth perimeter:

2️⃣

Compute Layer & Elastic Autoscaling

High-density, cost-effective container orchestration:

3️⃣

Databases, Caching & Data Layer

Decoupling hot reads, writes, and cache tiers:

4️⃣

Zero-Trust Security & Secrets

Hardening at rest, in transit, and across identities:

5️⃣

Full-Stack Observability & Disaster Recovery

Unified telemetry and business continuity:

💡 The Key Interview Point (The "Gold Nugget")
"Achieving scale and 99.99% availability isn't about bigger machines: it's CloudFront/WAF edge caching, 3-AZ VPC with Karpenter + KEDA autoscaling, Aurora Multi-AZ with Redis caching, and zero-trust IAM with OpenTelemetry correlation."
📋 Quick 60-Second Talking Points (Elevator Pitch)
  • Edge & Ingress: CloudFront + WAF + Shield -> ALB with IP Target Mode into multi-AZ EKS cluster.
  • Compute & Autoscaling: EKS with Karpenter for sub-minute node scaling; HPA + KEDA for event-driven pod scaling.
  • Data Tier: Multi-AZ Aurora PostgreSQL with auto-scaling read replicas + ElastiCache Redis cluster for 80%+ cache hit ratio.
  • Security: IRSA for pod IAM, External Secrets Operator, Cilium NetworkPolicies (default-deny), KMS encryption.
  • Observability: OpenTelemetry pipeline -> Prometheus, Loki, Tempo with trace_id correlation across all logs.
  • DR: Pilot light / warm standby in secondary region with Aurora Global Database and Route 53 health-check failover.
View Complete Standalone Runbook →
Staff / Principal SRE SRE & Operations Incident Management SEV-1 Incident

Q) A production deployment causes a sudden 40% increase in error rates. Walk through your incident-response process from detection and diagnosis to rollback, mitigation, root-cause analysis, and post-incident prevention.

End-to-end SEV-1 incident response lifecycle: automated SLO detection, Incident Command System, immediate mitigation (stopping the bleeding before debugging), deep root cause isolation, and blameless post-mortem.

#Incident Response #Rollback #SRE #Post-Mortem
💡 Practice reciting your story first
Candidate Opening: "A 40% error rate spike is a SEV-1 outage. My golden rule as an Incident Commander is: 'Mitigate first to protect customers; debug second.' Never spend hours debugging a broken deployment in production when an immediate rollback restores service."
1️⃣

Detection & Incident Command Mobilization (0–3 Minutes)

Rapid escalation and role assignment:

2️⃣

Immediate Mitigation — Stop the Bleeding (3–8 Minutes)

Restore customer availability before analyzing code:

3️⃣

Validate Recovery & Service Stabilization (8–15 Minutes)

Confirm the rollback successfully eliminated the error spike:

4️⃣

Diagnosis & Deep Root-Cause Analysis (Post-Recovery)

Isolate the root cause in a staging or isolated debug environment:

5️⃣

Blameless Post-Mortem & Preventative Action Items

Institutional learning and automated guardrails:

💡 The Key Interview Point (The "Gold Nugget")
"Mitigate first, debug second. Stop customer bleeding via immediate rollback or feature flag disablement within 5 minutes. Formally separate Incident Commander, Ops Lead, and Comms Lead, followed by a blameless post-mortem."
📋 Quick 60-Second Talking Points (Elevator Pitch)
  • Detection (0-3m): PagerDuty alert on SLO error budget burn rate; establish Incident Commander (IC), Ops Lead, and Comms Lead.
  • Mitigate First (3-8m): Correlate with recent release and rollback immediately (kubectl rollout undo / Argo abort / feature flag toggle).
  • Verify (8-15m): Confirm error rates return to baseline via ALB metrics and synthetic smoke tests; update status page.
  • RCA: Isolate traces/logs in staging environment to identify unhandled exception, DB pool exhaustion, or schema mismatch.
  • Prevention: Conduct blameless post-mortem; implement automated Canary deployment analysis gates (auto-abort if error > 1%).
View Complete Standalone Runbook →
Staff / Principal SRE CI/CD Platform Engineering & Delivery Enterprise Platform

Q) Design an enterprise-grade CI/CD pipeline for multiple teams deploying microservices independently. How would you implement automated testing, artifact management, security scanning, approvals, deployment strategies, and rollback?

Full blueprint for an enterprise CI/CD platform supporting autonomous microservices teams: trunk-based CI, immutable container signing with Cosign, GitOps deployment with ArgoCD, and automated canary verification.

#CI/CD #GitOps #ArgoCD #Security
💡 Practice reciting your story first
Candidate Opening: "An enterprise CI/CD system must provide autonomous developer self-service while enforcing strict organizational security, automated progressive delivery, and zero-downtime rollbacks."
1️⃣

CI Stage: Automated Testing & Security Shift-Left

Triggered on PR and commit using GitHub Actions / GitLab CI:

2️⃣

Artifact Management, SBOM & Cryptographic Signing

Securing the software supply chain from build to registry:

3️⃣

CD Stage: Declarative GitOps with ArgoCD

Decoupling continuous integration from continuous deployment:

4️⃣

Progressive Delivery (Canary) & Automated Rollback

Safe, zero-downtime deployment to Production:

💡 The Key Interview Point (The "Gold Nugget")
"Decouple CI from CD via GitOps (ArgoCD). Enforce shift-left security (Trivy, Cosign, SBOM), use immutable Git SHA tags, and deploy to production via metric-driven automated canary analysis with instant rollback."
📋 Quick 60-Second Talking Points (Elevator Pitch)
  • CI Phase: Parallel unit/contract tests, SAST (SonarQube), SCA (Trivy), secret scanning (TruffleHog) in <10 mins.
  • Artifacts: Immutable image tagged by Git SHA, SBOM generated via Syft, cryptographically signed with Cosign/Sigstore.
  • CD Phase: Two-repo GitOps pattern with ArgoCD; CI commits updated image tag to environment manifest repo.
  • Deployment Strategy: Progressive Canary rollout (5% -> 25% -> 100%) via Argo Rollouts with Prometheus metric gates.
  • Rollback: Automated abort if canary error rate exceeds 0.5%; one-click git revert in GitOps repo.
View Complete Standalone Runbook →
Staff / Principal SRE Terraform IaC Architecture & Governance Enterprise Platform

Q) How would you implement Infrastructure as Code at scale using Terraform? Discuss module design, state management, remote backends, state locking, environment separation, secrets, drift detection, and safe changes across hundreds of resources.

Comprehensive architectural guide for scaling Terraform across enterprise engineering teams: layered modules, remote state locking, multi-account isolation, secrets, and automated drift detection.

#Terraform #State Management #DynamoDB #Drift Detection
💡 Practice reciting your story first
Candidate Opening: "Scaling Terraform to hundreds of resources across multiple teams requires decomposing state, eliminating blast radius, enforcing automated PR-driven workflows, and automating drift detection."
1️⃣

Modular Architecture & Version Pinning

Layered, reusable building blocks with strict semantic versioning:

2️⃣

State Management, Remote Backend & Locking

Zero local state files; enterprise concurrency controls:

3️⃣

Environment Separation & Secrets Hygiene

Account isolation and zero plaintext credentials:

4️⃣

PR-Driven Workflow & Automated Drift Detection

Safe execution and continuous compliance:

💡 The Key Interview Point (The "Gold Nugget")
"Decompose state by layer and environment; lock with DynamoDB; pin module versions via Git tags; enforce changes strictly through PR workflows (Atlantis); and eliminate secrets from state via dynamic Vault lookups."
📋 Quick 60-Second Talking Points (Elevator Pitch)
  • Layered Modules: Foundation (VPC) -> Platform (EKS) -> App (RDS), pinned to immutable Git tags (ref=v2.1.0).
  • State & Locking: S3 backend with bucket versioning and KMS encryption; DynamoDB state locking to prevent race conditions.
  • Decomposition: Split state by environment and layer to prevent monolithic blast radius.
  • Multi-Account: Separate AWS accounts for Dev, Staging, and Prod with strictly isolated IAM credentials.
  • Workflow: Atlantis PR-driven plan/apply with OPA policy checks; daily automated drift detection alerts.
View Complete Standalone Runbook →
Staff / Principal SRE Kubernetes Cluster Reliability & Diagnostics Core Diagnostics

Q) Your Kubernetes cluster is experiencing intermittent pod failures and high latency. How would you troubleshoot it systematically? Explain how you would investigate pods, nodes, networking, resource limits, probes, scheduling, DNS, and application metrics.

Exhaustive diagnostic methodology for intermittent Kubernetes failures: node kernel pressure, CPU throttling, CoreDNS latency (ndots:5), CNI IP exhaustion, and conntrack table drops.

#Kubernetes #CoreDNS #Latency #Troubleshooting
💡 Practice reciting your story first
Candidate Opening: "Intermittent failures and latency spikes in Kubernetes are notoriously elusive because they rarely show up as hard crashes. I isolate them using a layered, full-stack diagnostic model: Nodes -> Pods & Limits -> CoreDNS -> Networking/CNI -> Application Probes."
1️⃣

Layer 1: Node Health, Kernel & CPU Throttling

Inspect underlying host instances and container runtime:

2️⃣

Layer 2: Pod Restarts, OOMKills & Exit Codes

Inspect container states across namespaces:

3️⃣

Layer 3: CoreDNS & The 'ndots:5' DNS Latency Trap

Intermittent 1-second latency spikes are almost always DNS issues:

4️⃣

Layer 4: CNI, IP Exhaustion & Conntrack Table

Subtle networking drops at the host and CNI level:

💡 The Key Interview Point (The "Gold Nugget")
"Intermittent K8s latency is usually not pod crashes: it's Linux CFS CPU throttling, CoreDNS ndots:5 search domain multiplication, or Linux nf_conntrack table exhaustion. Deploy NodeLocal DNSCache and tune CPU limits."
📋 Quick 60-Second Talking Points (Elevator Pitch)
  • Layer 1 (Node): Check node pressure flags (Memory/DiskPressure) and CFS CPU throttling (container_cpu_cfs_throttled_seconds_total).
  • Layer 2 (Pods): Sort pods by restart count; inspect 'kubectl logs --previous' for OOMKilled (Exit Code 137).
  • Layer 3 (DNS): Inspect CoreDNS metrics; mitigate ndots:5 search domain amplification using NodeLocal DNSCache.
  • Layer 4 (Network): Verify VPC CNI subnet IP availability; check 'dmesg' for nf_conntrack table exhaustion drops.
  • Layer 5 (Probes): Ensure liveness probes have sufficient timeout/initialDelay to avoid killing slow-starting pods.
View Complete Standalone Runbook →
Staff / Principal SRE CI/CD Deployment Architecture Deployment Architecture

Q) How would you design a zero-downtime deployment strategy for a critical application? Compare rolling, blue-green, canary, and feature-flag-based deployments, and explain when you would choose each.

Deep architectural tradeoff matrix comparing Rolling, Blue-Green, Canary, and Feature Flag deployments, detailing when to use each and how to execute zero-downtime database schema migrations.

#Zero Downtime #Canary #Blue-Green #Feature Flags
💡 Practice reciting your story first
Candidate Opening: "Zero-downtime deployments require isolating application traffic during transitions and ensuring data stores support multiple software versions concurrently. Each strategy presents distinct tradeoffs between risk, cost, and complexity."
1️⃣

Strategy Comparison & Tradeoff Matrix

The four production deployment patterns evaluated:

2️⃣

The Universal Database Requirement: Expand & Contract

Zero-downtime is impossible without two-phase schema migrations:

3️⃣

Decision Framework: When to Choose What

Choosing the right tool for the job:

💡 The Key Interview Point (The "Gold Nugget")
"Choose Canary (Argo Rollouts) for high-traffic microservices; Blue-Green for instant rollback and clean cutover; Feature Flags to decouple deployment from release. Always use Expand/Contract database migrations."
📋 Quick 60-Second Talking Points (Elevator Pitch)
  • Rolling Update: In-place, 0 extra cost, but mixed versions and slow rollback. Good for routine stateless updates.
  • Blue-Green: Dual environments, instant router flip rollback, but 2x cost. Best for monolithic apps and major releases.
  • Canary: Progressive traffic shifting (5% -> 100%) with automated metric-based rollback. Best for critical high-traffic services.
  • Feature Flags: Decouple deploy from release; instant toggling without redeployment. Best for user-facing features.
  • Database rule: Always use Expand/Contract (multi-phase) schema migrations to support concurrent versions.
View Complete Standalone Runbook →
Staff / Principal SRE SRE & Operations Performance Engineering Incident Triage

Q) A critical service suddenly starts consuming excessive CPU and memory. What would your troubleshooting methodology look like? How would you distinguish between application-level problems, infrastructure issues, resource misconfiguration, and traffic anomalies?

Deep diagnostic framework for distinguishing application memory leaks, infinite loops, CFS throttling, external database deadlocks, and traffic anomalies under pressure.

#Performance #Memory Leak #CPU Burn #Profiling
💡 Practice reciting your story first
Candidate Opening: "When a service suddenly spikes in both CPU and memory, engineers often assume traffic surge or code bug. A rigorous SRE methodology classifies the anomaly across 4 distinct buckets before taking action."
1️⃣

Step 1: Classify the Anomaly Across 4 Buckets

Telemetry triage in the first 2 minutes:

2️⃣

Step 2: CPU Deep-Dive (Thread & Syscall Profiling)

Isolate user code vs kernel syscalls:

3️⃣

Step 3: Memory Deep-Dive (Leak vs Cache)

Distinguish healthy cache from dangerous memory leaks:

4️⃣

Step 4: Mitigate Safely and Protect Customers

Immediate stabilization and safeguards:

💡 The Key Interview Point (The "Gold Nugget")
"Classify first: Flat traffic + pegged CPU = application bug/infinite loop. Flat traffic + rising memory without GC drop = memory leak. High response time + rising memory = downstream database bottleneck piling up threads."
📋 Quick 60-Second Talking Points (Elevator Pitch)
  • Classify into 4 buckets: Traffic surge, application bug (leak/loop), downstream bottleneck (thread pileup), or misconfiguration.
  • CPU drilldown: 'top -H -p <PID>' to find rogue thread; capture thread dump (jstack/pprof) or strace before killing.
  • Memory drilldown: Monitor GC pauses in Prometheus. If full GC doesn't reclaim memory, capture heap dump (jmap/heap profile).
  • Downstream check: Verify if slow database queries are causing worker thread accumulation and buffer exhaustion.
  • Mitigate: Trip circuit breakers, scale out horizontally, or roll restart pods with captured diagnostic dumps.
View Complete Standalone Runbook →
Staff / Principal SRE Security & DevSecOps Software Supply Chain Security DevSecOps & Compliance

Q) How would you secure a DevOps pipeline against supply-chain attacks? Discuss secrets management, IAM/RBAC, dependency scanning, container image security, artifact signing, SBOMs, least privilege, and pipeline isolation.

Comprehensive supply-chain defense architecture based on SLSA framework: source code signing, ephemeral OIDC runners, dependency SCA, SBOM generation, and Cosign admission enforcement.

#Security #Supply Chain #SLSA #SBOM
💡 Practice reciting your story first
Candidate Opening: "Supply-chain attacks target vulnerabilities in third-party dependencies, build systems, or unauthorized artifact tampering. Securing the pipeline requires enforcing the SLSA (Supply-chain Levels for Software Artifacts) framework from source code to production deployment."
1️⃣

Source Integrity & Pipeline Identity (OIDC)

Securing the code repository and build credentials:

2️⃣

Dependency Security (SCA) & Software Bill of Materials (SBOM)

Controlling third-party software risks:

3️⃣

Container Image Hardening & Cryptographic Signing (Cosign)

Building tamper-proof container artifacts:

4️⃣

Production Admission Control Enforcement

The cluster gatekeeper that enforces verification:

💡 The Key Interview Point (The "Gold Nugget")
"Securing the supply chain requires zero static CI keys (use OIDC), minimal distroless containers, generating an SBOM (Syft), signing artifacts with Cosign, and enforcing signature verification at K8s admission time via Kyverno."
📋 Quick 60-Second Talking Points (Elevator Pitch)
  • Pipeline Identity: Replace static AWS secrets with OIDC federated IAM roles; pin CI actions to immutable commit SHAs.
  • Dependencies: Scan dependencies via Snyk/Trivy; pull through a private caching proxy; generate SBOMs using Syft.
  • Build: Use multi-stage distroless/non-root containers; sign images and attestations using Cosign/Sigstore.
  • Cluster Enforcement: Kyverno/OPA admission controllers verify Cosign signatures; reject any unsigned image at deploy time.
  • Runtime Security: Enforce read-only root filesystems and monitor runtime threats with Falco.
View Complete Standalone Runbook →
Staff / Principal SRE Observability & SRE Telemetry & Reliability Enterprise Observability

Q) Design an observability strategy for a distributed production system. Explain how you would use logs, metrics, traces, alerting, SLOs/SLIs, dashboards, correlation IDs, and incident-management practices to identify problems quickly.

Full-stack observability architecture: Google's Four Golden Signals, OpenTelemetry standard, log-metric-trace correlation with exemplars, multi-window error budget burn rate alerting, and incident response.

#Observability #Prometheus #OpenTelemetry #Grafana
💡 Practice reciting your story first
Candidate Opening: "Observability is the ability to infer the internal state of a system based on its external outputs. In distributed microservices, the goal is unified correlation across metrics, logs, and traces to drive MTTR down to minutes."
1️⃣

The Three Telemetry Pillars + Correlation (OpenTelemetry)

Standardized collection via the OpenTelemetry (OTel) framework:

2️⃣

SLIs, SLOs & Multi-Window Burn Rate Alerting

Eliminating alert fatigue through Google SRE error budgets:

3️⃣

Role-Based Dashboards & Incident Integration

Turning raw telemetry into rapid operational decisions:

💡 The Key Interview Point (The "Gold Nugget")
"Standardize on OpenTelemetry. Correlate metrics, logs, and traces using trace_id exemplars so engineers jump from metric spikes to exact traces to logs. Alert exclusively on SLO error budget burn rates to eliminate alert fatigue."
📋 Quick 60-Second Talking Points (Elevator Pitch)
  • OpenTelemetry Collection: Unified OTel agent collecting RED metrics (Prometheus), structured logs (Loki), and traces (Tempo).
  • Correlation: Inject W3C trace_id into all logs and metric exemplars; 1-click jump from Grafana metric graph to trace and logs.
  • SLO Framework: Define Availability/Latency SLIs (99.9%); alert on multi-window error budget burn rates via PagerDuty.
  • Dashboards: Tiered dashboards (High-level business health -> Service RED metrics -> Node/Pod saturation).
  • Incident Integration: PagerDuty alerts link directly to troubleshooting runbooks and contextual dashboards.
View Complete Standalone Runbook →
Staff / Principal SRE Cloud Migration Enterprise Cloud Adoption Enterprise Migration

Q) Scenario: Your organization needs to migrate a large production workload from on-premises infrastructure to AWS/Azure with minimal downtime and no major service disruption. Explain your migration strategy, dependency mapping, networking, data replication, security, IaC, testing, cutover, rollback, and post-migration optimization.

Production playbook for migrating mission-critical enterprise workloads from on-premises to AWS with near-zero downtime: discovery, Direct Connect hybrid networking, DMS continuous CDC sync, cutover, and rollback.

#Cloud Migration #AWS #Direct Connect #DMS
💡 Practice reciting your story first
Candidate Opening: "Migrating enterprise workloads with near-zero downtime requires decoupling the migration into distinct phases: hybrid network foundation, continuous data synchronization with Change Data Capture (CDC), and a low-risk DNS/router cutover with an instant rollback mechanism."
1️⃣

Discovery, Dependency Mapping & Classification (The 7 Rs)

Audit and assess every service before touching infrastructure:

2️⃣

Hybrid Networking & Landing Zone Foundation

Establishing high-throughput, secure communication:

3️⃣

Continuous Data Replication (Change Data Capture - CDC)

Eliminating data transfer downtime during cutover:

4️⃣

Dry-Run Testing, Cutover & Reverse Rollback

Executing the cutover with minimum downtime (<5 minutes):

5️⃣

Post-Migration Optimization

Cost reduction and cloud-native refinement:

💡 The Key Interview Point (The "Gold Nugget")
"Near-zero downtime migration is achieved by pre-syncing data with AWS DMS CDC (Change Data Capture) over Direct Connect, lowering DNS TTL to 60s, and configuring reverse CDC back to on-prem as an instant safety rollback net."
📋 Quick 60-Second Talking Points (Elevator Pitch)
  • Phase 1 (Discovery): Map dependencies via AWS Discovery Service; classify workloads using 7 Rs framework.
  • Phase 2 (Hybrid Network): 10G AWS Direct Connect + Transit Gateway + Route 53 Resolver endpoints for hybrid DNS.
  • Phase 3 (Data Sync): AWS DMS with continuous Change Data Capture (CDC) to keep Aurora in real-time sync with on-prem DB.
  • Phase 4 (Testing): Terraform deploys target EKS/Aurora; execute load and security testing against cloud replica.
  • Phase 5 (Cutover & Rollback): Set DNS TTL to 60s, set on-prem read-only, promote Aurora, flip DNS. Keep reverse CDC active for 72h rollback safety.
View Complete Standalone Runbook →
Senior DevOps / DevSecOps Kubernetes Core Workload Primitives Core Architecture

Q) In Kubernetes, what is the difference between Deployment, StatefulSet, and DaemonSet?

Clear decision matrix distinguishing Kubernetes Deployment, StatefulSet, and DaemonSet controllers, covering pod identity, network identity, persistent volume lifecycle, and per-node scheduling.

#Kubernetes #Workloads #Deployment #StatefulSet
💡 Practice reciting your story first
Candidate Opening: "In production, selecting the wrong workload controller leads to data corruption, deployment deadlocks, or wasted compute. I categorize them by statefulness, pod identity, and placement topology."
1️⃣

Deployment: Stateless, Interchangeable Workloads

Designed for stateless microservices, web apps, and API gateways where pods are ephemeral and interchangeable:

2️⃣

StatefulSet: Ordered, Stateful & Identity-Preserving Workloads

Designed for databases, distributed storage, and clustered systems (Kafka, Elasticsearch, PostgreSQL, Redis Cluster, ZooKeeper):

3️⃣

DaemonSet: Exactly One Pod Per Node Workloads

Ensures that a copy of the pod runs on every matching node in the cluster:

💡 The Key Interview Point (The "Gold Nugget")
"Use Deployments for interchangeable stateless APIs; StatefulSets for quorum-based, disk-backed databases needing deterministic DNS and storage; and DaemonSets for infrastructure agents (CNI, logging, metrics, security) that must live on every physical node."
📋 Quick 60-Second Talking Points (Elevator Pitch)
  • Deployment: Stateless apps (APIs, web apps); random pod hashes; shared or no persistent disk; rolling updates.
  • StatefulSet: Clustered databases (Kafka, Postgres, Redis); deterministic ordinal names (app-0, app-1); dedicated volumeClaimTemplates; headless service DNS.
  • DaemonSet: Infrastructure agents (Fluentbit, Node Exporter, Cilium, Falco); runs exactly one pod per node automatically as nodes scale.
  • Pro-Tip: Never run databases on Deployments with ReadWriteOnce EBS volumes—pods will fail to remount across nodes due to multi-attach errors.
View Complete Standalone Runbook →
Senior DevOps / SRE Kubernetes Networking & Ingress Core Networking

Q) Explain Ingress Controller. What is the difference between Ingress resource, Ingress Controller, and Service?

Demystifying the Kubernetes Ingress architecture: distinguishing the Ingress Resource (manifest), Ingress Controller (reverse proxy daemon), and ClusterIP Services, tracing packet flow from client to container.

#Kubernetes #Ingress #Ingress Controller #NGINX
💡 Practice reciting your story first
Candidate Opening: "Candidates often confuse the declarative Ingress YAML with the actual reverse proxy software. Ingress requires two components: the declarative API rule and an active controller running in the cluster."
1️⃣

The Three Discrete Layers (Resource vs Controller vs Service)

Understanding the separation of concerns:

2️⃣

End-to-End Traffic Flow (Internet to Application)

How an HTTP request traverses the layers:

3️⃣

Production Add-Ons: TLS & Security

Essential components paired with Ingress Controllers in enterprise setups:

💡 The Key Interview Point (The "Gold Nugget")
"An Ingress Resource is just a passive config manifest. Without an active Ingress Controller pod listening to the Kubernetes API, your ingress rules do nothing. High-performance controllers route directly to Pod IPs via EndpointSlices rather than bouncing through kube-proxy."
📋 Quick 60-Second Talking Points (Elevator Pitch)
  • Ingress Resource is the YAML routing specification (hosts, paths, TLS secrets).
  • Ingress Controller is the active reverse proxy daemon (NGINX, Traefik, Envoy, AWS ALB Controller) executing the rules.
  • Service is the backend abstraction; the Ingress Controller watches Service Endpoints to stream traffic directly to container IPs.
  • Client -> Cloud LB -> Ingress Controller Pod -> Evaluates Host/Path Rules -> Direct connection to target Pod IP.
View Complete Standalone Runbook →
Senior DevOps / Platform Engineer Helm & GitOps Package Management & Architecture Enterprise Helm

Q) Suppose we have 20 microservices. Can we deploy all these services using one Helm deployment? How would you structure a Helm chart for multiple microservices?

Architectural trade-off analysis between Helm Umbrella Charts (one parent chart with 20 subcharts) vs Independent Helm Charts per microservice for deploying 20 microservices in enterprise environments.

#Helm #Microservices #Umbrella Chart #GitOps
💡 Practice reciting your story first
Candidate Opening: "Technically, yes—you can deploy 20 microservices using a single Helm Umbrella Chart with dependencies. However, doing so in Production creates severe blast-radius and deployment bottleneck problems."
1️⃣

Approach A: The Umbrella Chart Pattern (Can We?)

How an Umbrella Chart works in Helm:

2️⃣

Approach B: Standard Enterprise Pattern (How We Should Structure)

Independent microservice deployments powered by a shared Common/Library Chart:

3️⃣

Orchestrating 20 Charts with ArgoCD / Helmfile

Managing deployment across all 20 services cleanly without a monolithic umbrella chart:

💡 The Key Interview Point (The "Gold Nugget")
"Can you deploy 20 microservices with one Helm chart? Yes, via an Umbrella chart with 20 dependencies. Should you in Production? No. The enterprise best practice is a shared Library Chart with 20 independent releases orchestrated via ArgoCD ApplicationSets to eliminate blast radius."
📋 Quick 60-Second Talking Points (Elevator Pitch)
  • Yes, technically possible via an Umbrella Chart (Chart.yaml dependencies pointing to 20 subcharts).
  • Why it fails in Production: Massive blast radius, slow releases, merge conflicts in values.yaml, and all-or-nothing rollback failure.
  • Recommended Architecture: Build a standardized 'base-microservice' Library Chart; each service maintains its own lightweight chart in its repo.
  • Deploy independently via GitOps (ArgoCD ApplicationSets or Helmfile) so Team A can ship without risking Team B.
View Complete Standalone Runbook →
Senior DevOps / SRE Helm & GitOps Configuration Management Configuration Architecture

Q) How would you manage different configurations for Dev, QA, UAT, and Production using Helm?

Standard operating framework for managing multi-environment Kubernetes configurations across Dev, QA, UAT, and Production using values hierarchies, Helmfile, and secure parameter overrides.

#Helm #Environments #values.yaml #Helmfile
💡 Practice reciting your story first
Candidate Opening: "The golden rule of enterprise Helm configuration is: 'One Chart, Multiple Values'. Never duplicate template manifests across environments. Separate the chart engine from the environment data."
1️⃣

Layered Values File Strategy

Organize configuration by inheritance and environment overrides:

2️⃣

Deployment Command Chain

Helm merges multiple <code>-f</code> flags in order from left to right, with subsequent files overriding earlier ones:

3️⃣

Enterprise Automation: Helmfile or ArgoCD

How senior teams prevent human error in CI pipelines:

💡 The Key Interview Point (The "Gold Nugget")
"Follow the 'One Chart, Environment-Specific Values' principle. Base defaults in values.yaml, environment overrides in values-<env>.yaml merged via `-f` flags, image tags passed dynamically via CI/CD Git SHA, and secret values injected via External Secrets Operator rather than committed to Git."
📋 Quick 60-Second Talking Points (Elevator Pitch)
  • Maintain a single Helm chart; never duplicate templates per environment.
  • Layered values: Base values.yaml for common config, layered with values-dev.yaml, values-prod.yaml via '-f values.yaml -f values-prod.yaml'.
  • Environment differences: Dev uses 1 replica and spot nodes; Prod uses 3+ replicas, multi-AZ topology spread, strict PDBs, and production secrets.
  • Dynamic parameters: Image tag passed as Git SHA via '--set image.tag=$SHA' in pipeline.
  • GitOps integration: ArgoCD Application manifests define the target cluster, namespace, and values file per environment.
View Complete Standalone Runbook →
Senior DevOps / SRE Helm & GitOps Release Engineering Release Recovery

Q) How do you perform a Helm rollback?

Technical deep dive into 'helm rollback': how Helm stores release history in Kubernetes Secrets, three-way merge patching, and the critical limitations around CRDs and database schema migrations.

#Helm #Rollback #helm rollback #Release Secrets
💡 Practice reciting your story first
Candidate Opening: "A Helm rollback is not just 're-running old YAML'. It is a precise historical reconciliation against Helm's release metadata stored inside cluster Secrets."
1️⃣

Step 1: Inspect History & Trigger Rollback

Standard CLI operational runbook:

2️⃣

Step 2: What Happens Under the Hood?

How Helm executes the rollback internally:

3️⃣

Step 3: Critical Limitations & The Database Trap

What senior engineers know that juniors miss:

💡 The Key Interview Point (The "Gold Nugget")
"Helm rollback creates a brand new revision that matches the target historical revision manifest using a 3-way merge against cluster release Secrets. Be aware that Helm never rolls back CRDs or database state."
📋 Quick 60-Second Talking Points (Elevator Pitch)
  • Check history: 'helm history <release> -n <ns>' to find the last healthy revision number.
  • Roll back: 'helm rollback <release> <revision_number> --wait'.
  • Under the hood: Helm reads the compressed release Secret (sh.helm.release.v1.*), computes a 3-way strategic merge patch, and creates a NEW incremented revision.
  • Limitation 1: Helm does not manage or roll back CRDs.
  • Limitation 2: Database migrations are not rolled back by Helm; migrations must be backward-compatible.
View Complete Standalone Runbook →
Senior DevOps / SRE Kubernetes Troubleshooting & Pod Lifecycle Core Troubleshooting

Q) How do you troubleshoot ImagePullBackOff?

Comprehensive diagnostic runbook for isolating the 4 primary root causes of ImagePullBackOff: image name/tag typos, missing or expired registry credentials, Docker Hub rate limiting (429), and VPC/DNS network egress failures.

#Kubernetes #ImagePullBackOff #Docker #ECR
💡 Practice reciting your story first
Candidate Opening: "ImagePullBackOff means kubelet tried to pull the container image from the registry, failed, and is backing off exponentially. My troubleshooting begins by running 'kubectl describe pod <pod-name>' and reading the exact container runtime error in the Events."
1️⃣

Inspect Describe Events for the Exact Error String

Run <code>kubectl describe pod &lt;pod&gt; -n &lt;ns&gt;</code> and check the bottom Events section:

2️⃣

Root Cause 1: Image Name or Tag Typo / Non-Existent Image

Error: <code>manifest unknown</code> or <code>repository does not exist</code>:

3️⃣

Root Cause 2: Authentication & Missing imagePullSecrets

Error: <code>401 Unauthorized</code> or <code>403 Forbidden / Access Denied</code>:

4️⃣

Root Cause 3 & 4: Rate Limiting (429) & Network / Egress Blocks

Error: <code>toomanyrequests: You have reached your pull rate limit</code> or <code>i/o timeout</code>:

💡 The Key Interview Point (The "Gold Nugget")
"Read the exact error in 'kubectl describe pod': 'manifest unknown' = image tag typo or unpushed image; '401/403' = missing imagePullSecret or IAM/AcrPull role; '429' = Docker Hub rate limit; 'i/o timeout' = NAT Gateway or egress firewall block."
📋 Quick 60-Second Talking Points (Elevator Pitch)
  • Run 'kubectl describe pod <name>' and examine the Events message.
  • Case 1: 'manifest unknown' -> Image tag typo or CI/CD failed to push image.
  • Case 2: '401 Unauthorized' -> Missing imagePullSecret in pod spec or node lacks ECR/ACR IAM pull permissions.
  • Case 3: '429 Too Many Requests' -> Docker Hub rate limit; switch to private registry mirror (ECR/ACR).
  • Case 4: 'Connection timeout' -> Node in private subnet has no egress route to NAT Gateway or Security Group blocks 443 outbound.
View Complete Standalone Runbook →
Senior DevOps / SRE Kubernetes Observability & CLI Tooling Core Diagnostics

Q) How do you check Kubernetes pod logs and events?

Practical command toolkit for inspecting Kubernetes pod logs and cluster events: previous container crashes, multi-container pods, real-time log streaming with Stern, and filtering events with JSONPath.

#Kubernetes #kubectl logs #kubectl events #Debugging
💡 Practice reciting your story first
Candidate Opening: "Checking logs and events is fundamental, but in large-scale production with multi-container pods, crashing containers, and thousands of events, using standard 'kubectl logs' alone is insufficient."
1️⃣

Mastering Pod Logs (Crashing, Multi-Container, Live Stream)

Essential commands for container log inspection:

2️⃣

Mastering Cluster Events (Sorting, Filtering & Warning Triage)

Events explain WHY pods are failing, evicted, or unschedulable:

3️⃣

Deep Debugging: Ephemeral Containers & Node Logs

When container logs are silent or container won't run:

💡 The Key Interview Point (The "Gold Nugget")
"Always use '--previous' to catch the smoking gun of a crashed container. Filter events by 'type=Warning' to eliminate noise, and leverage 'stern' for multi-pod regex log streaming across autoscaling pods."
📋 Quick 60-Second Talking Points (Elevator Pitch)
  • Current logs: 'kubectl logs <pod> -f' (add -c for specific container).
  • Crashed logs: 'kubectl logs <pod> --previous' to see crash stack trace.
  • Multi-pod streaming: Use 'stern <app-pattern>' for live tailing across all replicas.
  • Events: 'kubectl get events -n <ns> --sort-by=.metadata.creationTimestamp' and filter by 'type=Warning'.
  • Zero-downtime debugging: 'kubectl debug' to attach ephemeral netshoot container with diagnostics tools.
View Complete Standalone Runbook →
Senior DevOps / SRE Kubernetes Resource Management & Capacity Resource Architecture

Q) How do you configure CPU and memory requests/limits?

Definitive architectural guide for configuring CPU and memory requests/limits: how kube-scheduler uses requests, how the Linux kernel enforces limits, CFS CPU throttling vs OOMKilled (exit 137), and QoS classes.

#Kubernetes #Requests #Limits #QoS
💡 Practice reciting your story first
Candidate Opening: "Configuring requests and limits is not guess-work. Requests determine where the scheduler places pods; limits determine when the Linux kernel throttles or terminates them. Setting them incorrectly causes either cluster starvation or silent application slowness."
1️⃣

Requests vs Limits Mechanics

Understanding how the control plane and Linux kernel treat them differently:

2️⃣

The 3 Kubernetes Quality of Service (QoS) Classes

Kubernetes automatically assigns QoS based on requests and limits:

3️⃣

Production Right-Sizing Best Practices

How senior SREs avoid common traps:

💡 The Key Interview Point (The "Gold Nugget")
"Requests are for scheduling (reservations); limits are for kernel enforcement. CPU is compressible (hits limit -> throttles); memory is incompressible (hits limit -> OOMKilled Exit 137). Set requests = limits on databases for Guaranteed QoS, and right-size requests to p95 peak usage."
📋 Quick 60-Second Talking Points (Elevator Pitch)
  • Requests: Guaranteed minimum reserved by kube-scheduler; placement depends on it.
  • Limits: Maximum ceiling enforced by cgroups.
  • CPU behavior: Compressible -> throttled via CFS quota (app slows down, no crash).
  • Memory behavior: Incompressible -> killed immediately by OOM Killer (Exit Code 137).
  • QoS Classes: Guaranteed (requests == limits, safest), Burstable (requests < limits), BestEffort (no values, evicted first).
  • Best practice: Right-size with VPA recommendation mode; always set memory limits.
View Complete Standalone Runbook →
Senior DevOps / SRE Kubernetes Autoscaling & Reliability Core Autoscaling

Q) How does HPA work in Kubernetes?

Deep dive into Kubernetes HPA: metrics collection pipeline via Metrics Server and Custom Metrics API, the exact mathematical autoscaling formula, stabilization windows to prevent thrashing, and event-driven scaling with KEDA.

#Kubernetes #HPA #Autoscaling #Metrics Server
💡 Practice reciting your story first
Candidate Opening: "HPA automatically scales the number of pod replicas in a Deployment or StatefulSet based on observed metrics like CPU, memory, or custom business metrics. It operates as a continuous reconciliation control loop inside kube-controller-manager."
1️⃣

Step 1: The Control Loop & Metrics Pipeline

How HPA gathers metrics every 15 seconds (default <code>--horizontal-pod-autoscaler-sync-period</code>):

2️⃣

Step 2: The Exact Mathematical Formula

HPA executes this exact formula on every evaluation loop:

3️⃣

Step 3: Flapping Prevention (Behavior) & KEDA for Events

Advanced production safeguards:

💡 The Key Interview Point (The "Gold Nugget")
"HPA evaluates every 15s using `desiredReplicas = ceil[currentReplicas * (currentMetric / targetMetric)]`. Crucially, target CPU percentage is calculated against the pod's resource REQUESTS, not limits. Use stabilization windows to prevent flapping, and KEDA for event-driven queue scaling."
📋 Quick 60-Second Talking Points (Elevator Pitch)
  • Continuous control loop running in kube-controller-manager (polls every 15s).
  • Formula: desiredReplicas = ceil[ currentReplicas * (currentMetric / targetMetric) ].
  • Golden Rule: CPU percentage is relative to resource REQUESTS (not limits!). If requests aren't set, HPA cannot calculate CPU utilization.
  • Flapping prevention: Uses a 5-minute stabilization window for scale-down to absorb traffic dips.
  • For event queues (Kafka, SQS, RabbitMQ): Use KEDA to autoscale on queue length before CPU spikes.
View Complete Standalone Runbook →
Advertisement
📬

Master Technical Rounds with The Weekly Dispatch

Every Tuesday, get 3 new production incident scenarios with complete root-cause post-mortems and architecture trade-offs delivered directly to your inbox.

⚡ Subscribe Free → 💼 Browse Interview Experiences