Q: Six months after migrating 80 microservices to GitHub Actions, you realise the total GitHub Actions spend has tripled compared to your old Jenkins setup. An audit shows most workflows run on `ubuntu-latest` GitHub-hosted runners. What strategies do you apply to systematically reduce costs without slowing down developer feedback loops?
A systematic cost reduction framework:
#CI/CD #๐ Supply Chain Security & Advanced CI/CD #L3 #DevOps #Automation #Pipelines
๐๏ธ Candidate Opening & Architectural Context
""In enterprise CI/CD, you cannot rely on manual interventions; every rollback and promotion must be declarative. The interviewer is testing: CI cost optimisation, runner right-sizing, caching strategy, job concurrency tuning.. I structure my answer around systematic triage first, root cause analysis second, and permanent remediation third.""
Advertisement
๐ ๏ธ Production Runbook & Step-by-Step Resolution
1๏ธโฃ
Initial Diagnostics & Root Cause Analysis
A systematic cost reduction framework:
- Identify top spenders: GitHub's billing API exposes per-workflow minute usage. Export it and sort by cost. Focus optimisation on the top 10 workflows consuming 80% of spend.
- Self-hosted runners for heavy jobs: Migrate compilation and Docker build jobs to spot EC2 instances (via Actions Runner Controller). Cost drops by ~70% for compute-heavy tasks.
- Aggressive dependency caching: Calculate cache hit rates (GitHub shows this in each
actions/cachestep summary). A 30% cache hit rate means 70% of runs are downloading packages from scratch. Fix the cache key to be more stable. - Cancel redundant runs: Use
concurrencygroups withcancel-in-progress: true. If a developer pushes twice in 30 seconds, the first run is cancelled immediately.
2๏ธโฃ
Remediation & Permanent Safeguards
--- *More CI/CD scenarios added periodically. PRs welcome.*
- Step-level parallelism: Replace sequential test steps with a build matrix, running tests in 4 parallel jobs. Total wall-clock time drops 4ร but total minutes drop ~20% (faster = less idle time on expensive runners).
- Right-size runners: Jobs that only run shell scripts don't need a 4-CPU runner. Use
runs-on: ubuntu-latest(2 vCPU) for light jobs andruns-on: [self-hosted, large]only for heavy jobs.
๐ก The Senior SRE Gold Nugget (Key Architectural Takeaway)
"Pro-Tip: Identify top spenders: GitHub's billing API exposes per-workflow minute usage. Export it and sort by cost. Focus optimisation on t."
โก 60-Second Elevator Pitch Talking Points
- Identify top spenders: GitHub's billing API exposes per-workflow minute usage. Export it and sort...
- Self-hosted runners for heavy jobs: Migrate compilation and Docker build jobs to spot EC2 instanc...
- Aggressive dependency caching: Calculate cache hit rates (GitHub shows this in each actions/cache...
Advertisement