Q: Non-production EKS clusters cost $85,000/month because developers leave hundreds of test pods, load balancers, and PVCs running overnight and over weekends. How do you design automated environment sleep schedules and TTL enforcement without breaking active debugging?
Automating after-hours resource hibernation and aggressive garbage collection across non-production Kubernetes clusters.
Want to master this scenario in a live sandbox? KodeKloud's CKA & CKAD Hands-On Certification Track covers this exact problem with hands-on terminal drills.
🛠️ Production Runbook & Step-by-Step Resolution
Deploy Kube-Downscaler with Organizational Business Hours Schedule
Install `kube-downscaler` configured to scale Deployments, StatefulSets, and HorizontalPodAutoscalers to 0 replicas between 8:00 PM and 7:00 AM weekdays, and all weekend long.
# kube-downscaler Helm values
parameters:
DEFAULT_UPTIME: 'Mon-Fri 08:00-20:00 Europe/London'
EXCLUDE_NAMESPACES: 'kube-system,monitoring,argocd'
DOWNSCALE_PERIOD: 'Mon-Fri 20:00-08:00,Sat-Sun 00:00-24:00'
Provide Developer Self-Service Override Annotations
Empower engineers conducting late-night deployments or overseas testing to temporarily pause hibernation using simple annotations or a Backstage button.
# Developer namespace annotation to pause downscaling for 24h
kubectl annotate namespace team-qa downscaler/exclude-until="2026-10-15T12:00:00Z" --overwrite
Garbage Collect Orphaned Ephemeral PVCs and LoadBalancers
Run a daily Kubernetes CronJob that identifies namespaces labeled `environment=ephemeral` with no updated commits in 48 hours, automatically deleting the namespace to release attached EBS volumes and ALBs.
- Enforce automated night and weekend downscaling across non-prod namespaces using kube-downscaler.
- Provide developer self-service annotations to temporarily exempt namespaces under active testing.
- Aggressively garbage-collect orphaned PVCs and Cloud Load Balancers from stale preview namespaces.