โšก ~/naveed Interview Prep
โšก Portfolio Home โœ๏ธ Engineering Blog Deep Dives ๐ŸŽฏ Interview Hub 998+ Scenarios โ˜ธ๏ธ Kubernetes Mastery Hub 24 Modules ๐ŸŽฎ DevOps Arcade & Quizzes Subnet Blitz โšก ๐Ÿ—บ๏ธ DevOps Roadmaps PDFs & Guides ๐Ÿค– Morpheus Analysis AI Quant โ†— ๐Ÿ› ๏ธ Developer Tools Utilities ๐Ÿงช Labs & Experiments ๐Ÿ“„ Interactive CV & Certs ๐Ÿ”— All Links & Socials โšก Join The Dispatch (Weekly SRE Newsletter) →
Staff SRE / Principal Architect [L3] CI/CD ๐Ÿ” Supply Chain Security & Advanced CI/CD Staff SRE Scenario [L3]

Q: Your CI/CD pipeline deploys directly to production on every merge to `main`. One Friday afternoon, a developer merges a PR and goes offline. The deployment silently fails halfway through โ€” 30% of pods are running new code, 70% are running old code. There are no automated rollback hooks. How do you recover, and how do you prevent this situation architecturally?

Immediate recovery:

#CI/CD #๐Ÿ” Supply Chain Security & Advanced CI/CD #L3 #DevOps #Automation #Pipelines
๐ŸŽ™๏ธ Candidate Opening & Architectural Context
""In enterprise CI/CD, you cannot rely on manual interventions; every rollback and promotion must be declarative. The interviewer is testing: Deployment readiness, automated rollback, on-call accountability.. I structure my answer around systematic triage first, root cause analysis second, and permanent remediation third.""
Advertisement

๐Ÿ› ๏ธ Production Runbook & Step-by-Step Resolution

1๏ธโƒฃ

Initial Diagnostics & Root Cause Analysis

Immediate recovery:

  • Identify current state: kubectl rollout status deployment/api โ€” if Progressing is stuck, the rollout is hanging.
  • Manually roll back: kubectl rollout undo deployment/api โ€” Kubernetes reverts to the previous ReplicaSet revision immediately. Verify with kubectl rollout history deployment/api.
  • --atomic on Helm, or rollout analysis on Argo Rollouts โ€” if the deployment does not reach 100% healthy within a timeout, it automatically rolls back.
2๏ธโƒฃ

Remediation & Permanent Safeguards

Architectural prevention:

  • Deployment freeze rules: Block merges to main after 3 PM on Fridays via a GitHub branch protection check that queries the current time. Simple but highly effective.
  • Require on-call acknowledgement: The pipeline sends a Slack message with "Deployment starting, confirm you're watching" and waits 60 seconds for a reaction emoji before proceeding. If no reaction, the pipeline pauses and pages on-call.
๐Ÿ’ก The Senior SRE Gold Nugget (Key Architectural Takeaway)
"Pro-Tip: Identify current state: kubectl rollout status deployment/api โ€” if Progressing is stuck, the rollout is hanging.."
โšก 60-Second Elevator Pitch Talking Points
  • Identify current state: kubectl rollout status deployment/api โ€” if Progressing is stuck, the roll...
  • Manually roll back: kubectl rollout undo deployment/api โ€” Kubernetes reverts to the previous Repl...
  • --atomic on Helm, or rollout analysis on Argo Rollouts โ€” if the deployment does not reach 100% he...
Advertisement
Want more CI/CD scenarios?
Explore our complete collection of scenario-based CI/CD interview runbooks.
Browse All CI/CD Questions →

๐Ÿ“š Related Production Scenarios in CI/CD