CI/CD Pipeline Succeeds but New Version Isn't Deployed — Debugging
Diagnostic guide for when CI/CD logs show green checkmarks but the target production environment continues running old code: image tag caching, branch rules, an...
Master 118+ battle-tested scenario-based CI/CD & GitOps interview questions for Senior DevOps, Cloud, and SRE engineers. Includes incident runbooks, STAR talking points, and CLI commands.
Diagnostic guide for when CI/CD logs show green checkmarks but the target production environment continues running old code: image tag caching, branch rules, an...
Engineering strategy for recovering from halfway-failed deployments: rolling update halts, automated canary rollback, and protecting stateful databases....
Enterprise secrets management architecture: eliminating secrets from Git and CI logs, integrating HashiCorp Vault / AWS Secrets Manager with Kubernetes, and usi...
Architectural blueprints for Blue-Green and Canary deployments: traffic routing mechanics, automated metric verification, database migration considerations, and...
Full blueprint for an enterprise CI/CD platform supporting autonomous microservices teams: trunk-based CI, immutable container signing with Cosign, GitOps deplo...
Deep architectural tradeoff matrix comparing Rolling, Blue-Green, Canary, and Feature Flag deployments, detailing when to use each and how to execute zero-downt...
Architectural trade-off analysis between Helm Umbrella Charts (one parent chart with 20 subcharts) vs Independent Helm Charts per microservice for deploying 20 ...
Standard operating framework for managing multi-environment Kubernetes configurations across Dev, QA, UAT, and Production using values hierarchies, Helmfile, an...
Technical deep dive into 'helm rollback': how Helm stores release history in Kubernetes Secrets, three-way merge patching, and the critical limitation...
A pipeline that's always failing is worse than no pipeline — it loses trust and gets bypassed....
Use layer caching. Docker builds layers top-to-bottom. Unchanged layers are reused from cache....
This is a partial deployment — dangerous state. The two versions may be incompatible for traffic....
Staging != Production. Common gaps:...
Pre-commit prevention (stop it before it lands in Git):...
Parallelize:...
Blue-Green:...
Use feature flags (feature toggles):...
Structure: one Git repo for application code, one (or same repo separate path) for Kubernetes manifests/Helm values....
Use change detection to only build/test what changed:...
Branch protection rules (GitHub/GitLab):...
Ephemeral environments — create a new environment for each PR/branch, run tests, then tear it down....
The Jenkins plugin for that step isn't installed, or the pipeline is running in a restricted sandbox....
1. Workspace cleanup — add cleanWs() at the end of every pipeline. Removes workspace files after the build....
Use Jenkins input step with a timeout:...
Use conditional triggers and filters:...
1. GitHub Secrets — go to repo Settings → Secrets → add secrets. Reference in workflow as ${{ secrets.MY_SECRET }}. Never printed in logs....
Use reusable workflows (.github/workflows/ in a central repository):...
Use rules with changes:...
1. Runner resources — check CPU/memory on the runner machine. Is it overloaded with too many concurrent jobs?...
Two approaches:...
1. Multi-stage builds — build in one stage, copy only the artifact to a slim final stage:...
Never use latest in production — it makes rollback and debugging impossible....
Use service containers in CI:...
Categorize tests:...
1. Generate coverage report — most test frameworks output coverage: pytest --cov=. --cov-report=xml or jest --coverage....
Run migrations as part of the deployment pipeline, but carefully:...
"Shift-left" means moving security testing earlier in the pipeline (to the left side of the timeline) instead of only checking at the end....
Golden path — create a standardized pipeline template that all services inherit:...
CI = Continuous Integration (code merged, built, tested automatically). CD = Continuous Delivery (code deployable at any time, manual tri......
Anything > 10 minutes delays developer feedback. > 20 minutes kills trunk-based development. Aim for < 10 min for the main feedback loop....
Blue/green: flip traffic back to blue. ArgoCD: argocd app rollback <app> --revision=<prev>. Helm: helm rollback <release>. Feature flags:......
Run the same job with multiple combinations of parameters (OS, Python version, Node version). Useful for testing cross-compatibility....
Use actions/setup-python with cache: pip parameter, or actions/cache with the pip cache directory....
Drain connections, use graceful shutdown, wait for in-flight requests to complete, then switch traffic. Use terminationGracePeriodSeconds......
Store build outputs (JARs, Docker images, npm packages) in a repository (Nexus, Artifactory, ECR) for versioning, sharing, and auditabili......
GitHub Actions: on: push: tags: ['v']. GitLab: rules: - if: $CI_COMMIT_TAG....
Use Conventional Commits format (feat:, fix:, chore:). Tools: semantic-release, release-drafter, conventional-changelog auto-generate not......
Use OPA (Open Policy Agent) or Conftest to evaluate infrastructure code (Terraform, K8s YAML) against policies before deployment. E.g., n......
Authentication expired. Run aws ecr get-login-password | docker login --username AWS --password-stdin <ecr-endpoint> before the push step......
Use environment-scoped secrets in GitHub/GitLab. Define production secrets only in the production environment. The job only has access to......
Multiple teams' features are collected over a sprint and released together on a fixed schedule (e.g., every 2 weeks). Good for coordinati......
Use immutable tags: image:commit-sha or image:v1.2.3. Never overwrite the latest tag. Store the deployed tag in your GitOps repo so you a......
Prevents stuck jobs from consuming runner resources indefinitely. If a stage hasn't completed in 30 minutes, it's likely hung. Timeout ki......
Required controls: no secrets in code (secret scanning), audit trail of all deployments (pipeline logs), code review requirement (branch ......
Software Bill of Materials — list of all dependencies and their versions in your software. Required by many compliance frameworks and exe......
Enable branch protection with required CI checks. Squash merge doesn't bypass CI — the PR's branch must pass checks before the merge butt......
Progressive delivery = controlled rollout with automated analysis. Use Argo Rollouts or Flagger. They send 10% traffic to new version, me......
Excludes files from the Docker build context sent to the daemon. Exclude: node_modules, .git, test files, local configs. Smaller context ......
withCredentials([usernamePassword(credentialsId: 'my-cred', usernameVariable: 'USER', passwordVariable: 'PASS')]) { sh "docke...
Gitflow: long-lived branches (develop, release, hotfix, feature). Complex, merge hell with many developers. Trunk-based: everyone commits......
Increment chart version on every change. In CI: if app version bumps, bump chart appVersion and version. Package chart: helm package. Pus......
Never make real third-party API calls in CI. Use Stripe test mode credentials (not live). Or better: use a mock server that simulates Str......
Backstage.io + GitHub Actions. Developer clicks "Deploy to staging" in Backstage → triggers a GitHub Actions workflow → deploys the servi......
Merge creates a merge commit, preserves history. Rebase replays commits on top of target, creates linear history. Linear history makes CI......
Pin dependency versions. Use lock files (package-lock.json, Pipfile.lock, go.sum). Never use latest or ^ (caret) ranges in production dep......
CI job assumes an IAM role (via OIDC — no stored credentials) with only secretsmanager:PutSecretValue permission on the specific secret A......
DORA metrics: Deployment Frequency (how often), Lead Time for Changes (commit to production), Change Failure Rate (% of deployments causi......
Abstract cloud differences behind a common interface. Use Terraform for infra (supports AWS, GCP, Azure). Helm/K8s for app deployment. CI......
Support both old and new credentials simultaneously during rotation window. Steps: generate new credentials → deploy new version that acc......
You must undeniably use Ephemeral Runners combined with Auto-Scaling (e.g., Actions Runner Controller in Kubernetes)....
Docker Hub limits unauthenticated pulls to 100 per 6 hours per IP. CI NAT gateways run out of this allowance instantly....
This is a systemic failure of the Test Pyramid. ...
The application crashed because the Database schema rolled forward, but the application code rolled backward. ...
You would utilize a Directed Acyclic Graph (DAG) by implementing the needs: keyword....
You must implement OpenID Connect (OIDC) identity federation....
Monorepos require highly intelligent Path Filtering and Dependency Graph Analysis....
Because of Goodhart's Law: "When a measure becomes a target, it ceases to be a good measure."...
Artifact registries must never be treated as infinite black holes. You must ruthlessly separate Snapshots (ephemeral development builds) ......
The team treated feature flags essentially as permanent configuration switches rather than highly Ephemeral Technical Debt....
No. That is a traditional Git-Flow anti-pattern. If you heavily branch off production and deploy directly, you run the immense risk of fo......
By default, helm upgrade is a "fire-and-forget" command. It confidently tells the Kubernetes API to update the deployment and immediately......
You must heavily abstract the execution logic into inherently central Pipeline Templates....
By default in GitHub, Branch Protection Rules implicitly natively exempt Repository Administrators and Organization Owners. Because the l......
Even if you offload physical execution to workers natively, the Jenkins Master JVM is still heavily strictly responsible for dynamically ......
You dynamically eliminate the physical bottleneck by intensely adopting Ephemeral Preview Environments....
You cannot fundamentally cache a massive 5GB file efficiently via standard ephemeral CI caching mechanisms natively (they often forcefull......
You must implement an overarching Cryptographic Supply Chain Security Architecture natively....
This is aggressively commonly intensely caused explicitly by Mutating Admission Webhooks natively operating silently entirely within the ......
Absolutely No. Standard Blue/Green fiercely flips the Load Balancer furiously instantly heavily natively. Active TCP WebSockets physicall......
This is a dependency confusion or typosquatting supply chain attack. Three layers of defence are required:...
Three controls must work together:...
Most CI systems emit webhook events on job start, completion, and failure. The observability stack is built on top:...
Use Renovate Bot configured at the organisation level rather than per-repository:...
Two immediate improvements:...
1. Deployment Locks: Introduce a deployment lock in the state store (DynamoDB table or Redis). Both Jenkins and GitHub Actions jobs must ......
Immediate cleanup:...
Immediate recovery:...
Local enforcement (pre-commit):...
This race condition is prevented by state locking. Terraform backend implementations (S3+DynamoDB, Terraform Cloud, etc.) acquire an excl......
ML pipelines introduce three concerns that don't exist in standard pipelines:...
GitOps does not mean you can't move fast — it means Git is always the source of truth....
The architecture uses build isolation per tenant:...
Root cause: Windows uses CRLF (\r\n) line endings; Linux/macOS use LF (\n). Shell scripts with \r characters fail silently or with crypti......
Add a dedicated manifest validation stage before any kubectl apply:...
Use Testcontainers or CI service containers — not a shared or cloud-hosted database:...
Two layers of defence — one at the Git level, one at the cluster level:...
Feature flag rot is a process problem, not just a tooling one:...
CI pipeline gate:...
A systematic cost reduction framework:...
Multi-stage CI/CD defense-in-depth framework for configuration management: schema validation, Helm linting and templating, strict Kubeconform checks, OPA/Confte...
How to refactor and modularize complex 900+ line GitHub Actions workflows: decomposing into reusable workflows and composite actions, utilizing build matrices, ...
Architectural comparison and security trade-offs between public and private GitHub Actions workflow repositories: internal sharing policies, caller permissions,...
How to configure GitHub Actions concurrency groups: cancelling redundant pull-request builds with cancel-in-progress to save runner minutes, while serializing p...
Systematic workflow failure triage: categorizing failures into code, transient infrastructure, or pipeline design defects, using the GitHub CLI (gh) for rapid i...
Enterprise security blueprint for hardening Jenkins: SAML/OIDC SSO, Matrix/Folder-based RBAC, isolating build execution on ephemeral Kubernetes agents, disablin...
End-to-end integration architecture: orchestrating ephemeral Jenkins agent pods in Kubernetes, building multi-arch Docker images via Buildx, authenticating to A...
Root cause analysis of the classic release paradox: why Kubernetes and CI/CD report green rollout success while end-users experience 500 errors, broken routing,...
Structured SRE decision framework for choosing between an immediate rollback versus a forward hotfix during high-severity production incidents: evaluating blast...