โšก ~/naveed Interview Prep
โšก Portfolio Home โœ๏ธ Engineering Blog Deep Dives ๐ŸŽฏ Interview Hub 998+ Scenarios โ˜ธ๏ธ Kubernetes Mastery Hub 24 Modules ๐ŸŽฎ DevOps Arcade & Quizzes Subnet Blitz โšก ๐Ÿ—บ๏ธ DevOps Roadmaps PDFs & Guides ๐Ÿค– Morpheus Analysis AI Quant โ†— ๐Ÿ› ๏ธ Developer Tools Utilities ๐Ÿงช Labs & Experiments ๐Ÿ“„ Interactive CV & Certs ๐Ÿ”— All Links & Socials โšก Join The Dispatch (Weekly SRE Newsletter) →
Staff SRE / Principal Architect [L3] CI/CD ๐Ÿ” Supply Chain Security & Advanced CI/CD Staff SRE Scenario [L3]

Q: Six months after migrating 80 microservices to GitHub Actions, you realise the total GitHub Actions spend has tripled compared to your old Jenkins setup. An audit shows most workflows run on `ubuntu-latest` GitHub-hosted runners. What strategies do you apply to systematically reduce costs without slowing down developer feedback loops?

A systematic cost reduction framework:

#CI/CD #๐Ÿ” Supply Chain Security & Advanced CI/CD #L3 #DevOps #Automation #Pipelines
๐ŸŽ™๏ธ Candidate Opening & Architectural Context
""In enterprise CI/CD, you cannot rely on manual interventions; every rollback and promotion must be declarative. The interviewer is testing: CI cost optimisation, runner right-sizing, caching strategy, job concurrency tuning.. I structure my answer around systematic triage first, root cause analysis second, and permanent remediation third.""
Advertisement

๐Ÿ› ๏ธ Production Runbook & Step-by-Step Resolution

1๏ธโƒฃ

Initial Diagnostics & Root Cause Analysis

A systematic cost reduction framework:

  • Identify top spenders: GitHub's billing API exposes per-workflow minute usage. Export it and sort by cost. Focus optimisation on the top 10 workflows consuming 80% of spend.
  • Self-hosted runners for heavy jobs: Migrate compilation and Docker build jobs to spot EC2 instances (via Actions Runner Controller). Cost drops by ~70% for compute-heavy tasks.
  • Aggressive dependency caching: Calculate cache hit rates (GitHub shows this in each actions/cache step summary). A 30% cache hit rate means 70% of runs are downloading packages from scratch. Fix the cache key to be more stable.
  • Cancel redundant runs: Use concurrency groups with cancel-in-progress: true. If a developer pushes twice in 30 seconds, the first run is cancelled immediately.
2๏ธโƒฃ

Remediation & Permanent Safeguards

--- *More CI/CD scenarios added periodically. PRs welcome.*

  • Step-level parallelism: Replace sequential test steps with a build matrix, running tests in 4 parallel jobs. Total wall-clock time drops 4ร— but total minutes drop ~20% (faster = less idle time on expensive runners).
  • Right-size runners: Jobs that only run shell scripts don't need a 4-CPU runner. Use runs-on: ubuntu-latest (2 vCPU) for light jobs and runs-on: [self-hosted, large] only for heavy jobs.
๐Ÿ’ก The Senior SRE Gold Nugget (Key Architectural Takeaway)
"Pro-Tip: Identify top spenders: GitHub's billing API exposes per-workflow minute usage. Export it and sort by cost. Focus optimisation on t."
โšก 60-Second Elevator Pitch Talking Points
  • Identify top spenders: GitHub's billing API exposes per-workflow minute usage. Export it and sort...
  • Self-hosted runners for heavy jobs: Migrate compilation and Docker build jobs to spot EC2 instanc...
  • Aggressive dependency caching: Calculate cache hit rates (GitHub shows this in each actions/cache...
Advertisement
Want more CI/CD scenarios?
Explore our complete collection of scenario-based CI/CD interview runbooks.
Browse All CI/CD Questions →

๐Ÿ“š Related Production Scenarios in CI/CD