Q: Your Platform team rolled out an ArgoCD ApplicationSet with a Matrix generator combining 500 Git repositories across 8 Kubernetes clusters (4,000 Application CRDs). The ArgoCD repo-server and API server are crashing under OOM. How do you remediate and scale this architecture?
Engineering strategy to prevent ArgoCD controller memory exhaustion and API rate-limiting when scaling ApplicationSets across 500+ microservices and multi-cluster topologies.
Want to master this scenario in a live sandbox? KodeKloud's CKA & CKAD Hands-On Certification Track covers this exact problem with hands-on terminal drills.
🛠️ Production Runbook & Step-by-Step Resolution
Sharding the ArgoCD Application Controller
Enable application controller sharding via ARGOCD_CONTROLLER_REPLICAS and ARGOCD_ENABLE_SHARDING=true. ArgoCD uses consistent hashing across cluster secrets to distribute the 4,000 apps evenly across multiple controller pods.
# argocd-cmd-params-cm
data:
controller.sharding.algorithm: round-robin
application.controller.replicas: '4'
Decouple Git Polling with Webhooks and Redis Caching
Increase timeout.reconciliation from default 180s to 600s, and increase Redis cache memory. Rely strictly on GitHub/GitLab commit webhooks to trigger targeted reconciliation rather than polling 500 repositories concurrently.
# argocd-cm
data:
timeout.reconciliation: '600s'
redis.cache.expiration: '24h'
Partition ApplicationSets by Domain and Environment
Split the monolithic matrix ApplicationSet into regional and domain-specific sets with generator filters, limiting each generator to under 250 targets.
- Enable ArgoCD controller sharding to distribute cluster reconciliation across multiple controller replicas.
- Replace 3-minute Git polling with event-driven repository webhooks to prevent Git API saturation.
- Decompose monolithic matrix ApplicationSets into domain-bounded generators to limit controller blast radius.