Q: Your CI/CD pipeline has been stable for months but suddenly starts failing without any pipeline changes. How would you isolate the root cause?
Root-cause analysis methodology when a mission-critical CI/CD pipeline that ran stably for months suddenly fails without any commits to the pipeline configuration.
🛠️ Production Runbook & Step-by-Step Resolution
Diff Build Logs Between Last Successful and First Failed Run
Download the raw console logs from the last green run and the first red run and run a line-by-line diff. Look for differences in resolved dependency versions, runner hostnames, or base image digest hashes.
# Diffing package resolution
diff -u build-success.log build-failure.log | grep -E '^[+-]' | head -n 30
Check Floating Dependencies & Base Images
Check if package managers (npm, pip, maven) or Dockerfiles use unpinned versions (e.g., node:18-alpine or ~1.2.0 in package.json) where an upstream transitive release broke compatibility overnight.
Verify Runner Resources & External API Rate Limits
Inspect runner health: Docker Hub pull rate limit (429 Too Many Requests), NPM/PyPI registry outages, runner disk space full (/var/lib/docker filling up /), or expired credentials (GPG signing keys, AWS IAM roles, NPM tokens).
df -h /var/lib/docker
docker system df
curl -I https://registry.npmjs.org/
- Diff the exact stdout logs between the last green run and the first failed run.
- Audit unpinned dependencies (transitive npm/pip libraries) and floating Docker base tags.
- Verify external registry rate limits (Docker Hub 429) and CI runner disk space (df -h).
- Validate CI service credentials and token expirations (AWS STS, npm publish tokens, GPG keys).