Q: Your engineering org has 250 developers running intensive build jobs. GitHub-hosted runners are slow, cost $35,000/month, and cannot access private databases. Static self-hosted VMs accumulate leftover cache files and create security cross-contamination between PRs. How do you design an elastic, ephemeral self-hosted runner architecture using Actions Runner Controller (ARC)?
Engineering a secure, auto-scaling ephemeral GitHub Actions runner fleet on Kubernetes using Actions Runner Controller (ARC), Docker-in-Docker isolation, and Spot instance compute.
Want to master this scenario in a live sandbox? KodeKloud's Enterprise GitOps with ArgoCD & Kubernetes Rollouts covers this exact problem with hands-on terminal drills.
🛠️ Production Runbook & Step-by-Step Resolution
Deploy Actions Runner Controller (ARC) v2 Operator via Helm
Establish the Kubernetes operator managing runner pod lifecycles:
- Operator Installation: Deployed ARC v2 controller in
arc-systemsnamespace using Helm. - GitHub App Authentication: Authenticated ARC against GitHub Organization using a GitHub App with scoped permissions (
Actions: Read/Write,Organization Administration: Read).
Configure AutoscalingRunnerSet with Ephemeral Pod Lifecycles
Define elastic runner pools that scale dynamically based on pending GitHub Actions queue depths:
- AutoscalingRunnerSet CRD: Configured runner set with
minRunners: 2,maxRunners: 150, andscaling: { metric: 'PercentageRunnersBusy' }. - Strict Ephemerality: Configured runner pods to terminate and delete immediately after executing a single workflow job:
ephemeral: true.
Configure Secure Container Mode vs. Docker-in-Docker (DinD)
Enable Docker builds without granting root access to the host Kubernetes node:
- Kubernetes Mode: Utilized ARC Container Mode, executing workflow steps as sibling containers inside the pod's network namespace.
- DinD Security Hardening: For legacy docker-compose workflows, deployed Docker-in-Docker sidecars with isolated tmpfs volume mounts and disabled root node privileges.
Target Kubernetes Spot Compute & Measure Performance Metrics
Slash runner infrastructure expenses while accelerating build speeds:
- Node Affinity & Taints: Scheduled runner pods onto dedicated AWS EC2 Spot / Azure Spot node pools with node taints (
workload=ci-runners:NoSchedule). - Results: Pipeline build execution times dropped by 45% due to high-performance vCPU allocation, while monthly CI costs decreased from $35,000 to $8,400 (76% savings).
- Deploy ARC v2 operator authenticated via a secure GitHub App.
- Configure AutoscalingRunnerSets with ephemeral: true to destroy pods after each job.
- Use ARC Container Mode to execute container steps without exposing host docker.sock.
- Schedule runners onto Spot instances, cutting monthly CI spend by over 70% while boosting speed.