⚡ ~/naveed Interview Prep
⚡ Portfolio Home ✍️ Engineering Blog Deep Dives 🎯 Interview Hub 1,000+ Scenarios ☸️ Kubernetes Mastery Hub 24 Modules 🎮 DevOps Arcade & Quizzes Subnet Blitz ⚡ 🗺️ DevOps Roadmaps PDFs & Guides 🤖 Morpheus Analysis AI Quant ↗ 🛠️ Developer Tools Utilities 🧪 Labs & Experiments 📄 Interactive CV & Certs 🔗 All Links & Socials ⚡ Join The Dispatch (Weekly SRE Newsletter) →
← Back to All CI/CD & GitOps Interview Questions Scenario 126 of 126 in CI/CD & GitOps
Senior DevOps / SRE CI/CD Deployment Pipelines & Automation Production Runbook

Q: The Jenkins/GitHub Actions pipeline is GREEN, but the application is DOWN in production. What will you investigate?

The CI/CD deployment pipeline reports a green success status, but users cannot access the application in production. Root cause isolation, silent failure modes, and automated rollback.

#CI/CD #Jenkins #GitHub Actions #Deployment #Rollback #Troubleshooting
🎙️ Candidate Opening & Architectural Context
"A green pipeline only proves that all pipeline stages completed with exit code 0. It does NOT guarantee that the deployed application is functioning or serving traffic. Silent failures frequently pass pipelines while taking production offline."
Advertisement
⚡ Recommended Practice Lab

Want to master this scenario in a live sandbox? KodeKloud's Enterprise GitOps with ArgoCD & Kubernetes Rollouts covers this exact problem with hands-on terminal drills.

🛠️ Production Runbook & Step-by-Step Resolution

1

Audit Pipeline Script Exit Codes & Error Handling

Check if deploy scripts swallowed errors or ignored failed commands. In bash, missing 'set -euo pipefail' allows intermediate failures in a pipeline to be masked by a succeeding downstream command.

# Ensure deploy scripts fail fast:
set -euo pipefail
2

Verify Asynchronous Deployment Rollout Status

Many pipelines run 'kubectl apply -f deployment.yaml' and immediately finish with green, while Kubernetes is still performing the rolling update in the background. If new pods crash, the pipeline has already exited successfully.

# Mandatory in CI/CD pipelines to block until rollout finishes:
kubectl rollout status deployment/<deployment-name> -n production --timeout=300s
3

Verify Artifact & Container Image Digest Parity

Verify whether the pipeline deployed the exact commit SHA image or erroneously defaulted to a cached 'latest' tag or stale artifact registry digest.

kubectl get deployment <deployment-name> -o jsonpath='{.spec.template.spec.containers[*].image}'
# Compare image digest against CI build artifact SHA
4

Inspect Environment Configuration, Secrets, and Ingress

A common cause of green pipelines with broken apps is configuration drift: the application binary succeeded in staging, but in production, database credentials or API keys were missing or expired.

kubectl get secrets -n production
kubectl logs -l app=<app-name> -n production --tail=50
5

Execute Immediate Rollback & Postmortem

Mitigate user impact immediately before conducting root cause debugging. Roll back to the previous healthy revision and verify production traffic recovers.

# Immediate rollback:
kubectl rollout undo deployment/<deployment-name> -n production
kubectl rollout status deployment/<deployment-name> -n production
💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"Never equate pipeline exit code 0 with application health. Enforce 'kubectl rollout status --timeout=300s', automated post-deployment synthetic smoke tests, and automated canary rollback triggers."
⚡ 60-Second Elevator Pitch Talking Points
  • A green pipeline merely means every step exited with 0; it doesn't mean the application is serving user requests.
  • First check if the deployment was asynchronous: did the pipeline run 'kubectl apply' without waiting on 'kubectl rollout status'?
  • Inspect container image tags and digests to verify the correct immutable artifact was deployed, not a stale cache.
  • Check production secrets and database migration logs: code often passes staging but crashes on missing prod credentials.
  • Roll back immediately with 'kubectl rollout undo' to restore customer availability, then triage the logs in a blameless postmortem.
Advertisement
Want more CI/CD & GitOps scenarios?
Explore our complete collection of scenario-based CI/CD & GitOps interview runbooks.
Browse All CI/CD & GitOps Questions →