Q: The Jenkins/GitHub Actions pipeline is GREEN, but the application is DOWN in production. What will you investigate?
The CI/CD deployment pipeline reports a green success status, but users cannot access the application in production. Root cause isolation, silent failure modes, and automated rollback.
Want to master this scenario in a live sandbox? KodeKloud's Enterprise GitOps with ArgoCD & Kubernetes Rollouts covers this exact problem with hands-on terminal drills.
🛠️ Production Runbook & Step-by-Step Resolution
Audit Pipeline Script Exit Codes & Error Handling
Check if deploy scripts swallowed errors or ignored failed commands. In bash, missing 'set -euo pipefail' allows intermediate failures in a pipeline to be masked by a succeeding downstream command.
# Ensure deploy scripts fail fast:
set -euo pipefail
Verify Asynchronous Deployment Rollout Status
Many pipelines run 'kubectl apply -f deployment.yaml' and immediately finish with green, while Kubernetes is still performing the rolling update in the background. If new pods crash, the pipeline has already exited successfully.
# Mandatory in CI/CD pipelines to block until rollout finishes:
kubectl rollout status deployment/<deployment-name> -n production --timeout=300s
Verify Artifact & Container Image Digest Parity
Verify whether the pipeline deployed the exact commit SHA image or erroneously defaulted to a cached 'latest' tag or stale artifact registry digest.
kubectl get deployment <deployment-name> -o jsonpath='{.spec.template.spec.containers[*].image}'
# Compare image digest against CI build artifact SHA
Inspect Environment Configuration, Secrets, and Ingress
A common cause of green pipelines with broken apps is configuration drift: the application binary succeeded in staging, but in production, database credentials or API keys were missing or expired.
kubectl get secrets -n production
kubectl logs -l app=<app-name> -n production --tail=50
Execute Immediate Rollback & Postmortem
Mitigate user impact immediately before conducting root cause debugging. Roll back to the previous healthy revision and verify production traffic recovers.
# Immediate rollback:
kubectl rollout undo deployment/<deployment-name> -n production
kubectl rollout status deployment/<deployment-name> -n production
- A green pipeline merely means every step exited with 0; it doesn't mean the application is serving user requests.
- First check if the deployment was asynchronous: did the pipeline run 'kubectl apply' without waiting on 'kubectl rollout status'?
- Inspect container image tags and digests to verify the correct immutable artifact was deployed, not a stale cache.
- Check production secrets and database migration logs: code often passes staging but crashes on missing prod credentials.
- Roll back immediately with 'kubectl rollout undo' to restore customer availability, then triage the logs in a blameless postmortem.