⚡ ~/naveed Interview Prep
⚡ Portfolio Home ✍️ Engineering Blog Deep Dives 🎯 Interview Hub 1,000+ Scenarios ☸️ Kubernetes Mastery Hub 24 Modules 🎮 DevOps Arcade & Quizzes Subnet Blitz ⚡ 🗺️ DevOps Roadmaps PDFs & Guides 🤖 Morpheus Analysis AI Quant ↗ 🛠️ Developer Tools Utilities 🧪 Labs & Experiments 📄 Interactive CV & Certs 🔗 All Links & Socials ⚡ Join The Dispatch (Weekly SRE Newsletter) →
← Back to All CI/CD & GitOps Interview Questions Scenario 173 of 176 in CI/CD & GitOps
Staff Infrastructure Architect Helm & GitOps Helm & GitOps Engineering Production Scenario

Q: Your enterprise GitOps infrastructure has grown to 550 EKS/GKE clusters and 12,000 Argo CD Applications. SREs report severe performance degradation: the Argo CD UI takes 45 seconds to load, reconciliation loops lag by over 30 minutes, and the single `argocd-repo-server` frequently encounters OOMKilled crashes during Git manifest generation. You must re-architect Argo CD for high availability and enterprise scale, implementing controller sharding, repo-server horizontal autoscaling with Redis caching, and resource exclusion rules.

Architect and tune an enterprise Argo CD instance managing 500+ Kubernetes clusters and 10,000+ Application resources, addressing repo-server bottlenecks, application-controller sharding, and Redis HA.

#Argo CD #Kubernetes #GitOps #Scalability #High Availability
🎙️ Candidate Opening & Architectural Context
"Architect and tune an enterprise Argo CD instance managing 500+ Kubernetes clusters and 10,000+ Application resources, addressing repo-server bottlenecks, application-controller sharding, and Redis HA."
Advertisement
⚡ Recommended Practice Lab

Want to master this scenario in a live sandbox? KodeKloud's Enterprise GitOps with ArgoCD & Kubernetes Rollouts covers this exact problem with hands-on terminal drills.

🛠️ Production Runbook & Step-by-Step Resolution

Step 1

Shard the Application Controller Across Multiple Replicas

The `argocd-application-controller` is stateful. Scale it horizontally by configuring dynamic cluster sharding using the `ARGOCD_CONTROLLER_REPLICAS` environment variable or Argo CD operator sharding algorithms.

apiVersion: apps/v1
kind: StatefulSet
metadata:
  name: argocd-application-controller
  namespace: argocd
spec:
  replicas: 5
  template:
    spec:
      containers:
        - name: argocd-application-controller
          env:
            - name: ARGOCD_CONTROLLER_REPLICAS
              value: '5'
          args:
            - /usr/local/bin/argocd-application-controller
            - --status-processors
            - '50'
            - --operation-processors
            - '25'
            - --repo-server-timeout-seconds
            - '120'
Pro Tip: Shard the Application Controller Across Multiple Replicas
Step 2

Scale and Cache the Repo Server with Redis HA

Deploy Redis Sentinel HA for manifest and OIDC session caching. Scale `argocd-repo-server` to 10 replicas with horizontal pod autoscaling (HPA) and configure memory limits to prevent fork/exec Helm memory exhaustion.

apiVersion: apps/v1
kind: Deployment
metadata:
  name: argocd-repo-server
  namespace: argocd
spec:
  replicas: 10
  template:
    spec:
      containers:
        - name: argocd-repo-server
          resources:
            requests:
              cpu: 1000m
              memory: 2Gi
            limits:
              cpu: 4000m
              memory: 8Gi
          env:
            - name: ARGOCD_REPO_SERVER_PARALLELISM_LIMIT
              value: '10'
            - name: REDIS_SERVER
              value: 'argocd-redis-ha-haproxy:6379'
Pro Tip: Scale and Cache the Repo Server with Redis HA
Advertisement
Step 3

Filter Ephemeral Cluster Resources with Resource Exclusions

Prevent the application controller from monitoring high-churn Kubernetes resources (e.g., EndpointSlices, PodMetrics, Helm temporary secrets) by tuning `resource.exclusions` in the `argocd-cm` ConfigMap.

apiVersion: v1
kind: ConfigMap
metadata:
  name: argocd-cm
  namespace: argocd
data:
  resource.exclusions: |
    - apiGroups:
        - discovery.k8s.io
      kinds:
        - EndpointSlice
    - apiGroups:
        - metrics.k8s.io
      kinds:
        - PodMetrics
    - apiGroups:
        - events.k8s.io
      kinds:
        - Event
Pro Tip: Filter Ephemeral Cluster Resources with Resource Exclusions
Step 4

Benchmark and Monitor Reconciliation Queues

Expose and monitor Argo CD Prometheus metrics: `argocd_app_reconcile_bucket`, `argocd_cluster_api_resource_objects`, and `argocd_git_request_total`. Configure alerts when workqueue depth exceeds 100 items for > 5 minutes.

Pro Tip: Benchmark and Monitor Reconciliation Queues
💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"Scaling Argo CD to hundreds of clusters requires sharding the application-controller StatefulSet across multiple replicas, scaling the repo-server horizontally behind a clustered Redis cache, and filtering high-churn API resources like EndpointSlices."
⚡ 60-Second Elevator Pitch Talking Points
  • T
  • o
  • s
  • u
  • p
  • p
  • o
  • r
  • t
  • 5
  • 5
  • 0
  • c
  • l
  • u
  • s
  • t
  • e
  • r
  • s
  • a
  • n
  • d
  • 1
  • 2
  • ,
  • 0
  • 0
  • 0
  • a
  • p
  • p
  • s
  • i
  • n
  • A
  • r
  • g
  • o
  • C
  • D
  • ,
  • w
  • e
  • r
  • e
  • s
  • o
  • l
  • v
  • e
  • d
  • r
  • e
  • c
  • o
  • n
  • c
  • i
  • l
  • i
  • a
  • t
  • i
  • o
  • n
  • l
  • a
  • g
  • b
  • y
  • s
  • h
  • a
  • r
  • d
  • i
  • n
  • g
  • t
  • h
  • e
  • a
  • p
  • p
  • l
  • i
  • c
  • a
  • t
  • i
  • o
  • n
  • c
  • o
  • n
  • t
  • r
  • o
  • l
  • l
  • e
  • r
  • a
  • c
  • r
  • o
  • s
  • s
  • 5
  • s
  • t
  • a
  • t
  • e
  • f
  • u
  • l
  • r
  • e
  • p
  • l
  • i
  • c
  • a
  • s
  • ,
  • s
  • c
  • a
  • l
  • i
  • n
  • g
  • t
  • h
  • e
  • r
  • e
  • p
  • o
  • -
  • s
  • e
  • r
  • v
  • e
  • r
  • t
  • o
  • 1
  • 0
  • r
  • e
  • p
  • l
  • i
  • c
  • a
  • s
  • w
  • i
  • t
  • h
  • R
  • e
  • d
  • i
  • s
  • H
  • A
  • c
  • a
  • c
  • h
  • i
  • n
  • g
  • ,
  • a
  • n
  • d
  • e
  • x
  • c
  • l
  • u
  • d
  • i
  • n
  • g
  • e
  • p
  • h
  • e
  • m
  • e
  • r
  • a
  • l
  • r
  • e
  • s
  • o
  • u
  • r
  • c
  • e
  • s
  • l
  • i
  • k
  • e
  • E
  • n
  • d
  • p
  • o
  • i
  • n
  • t
  • S
  • l
  • i
  • c
  • e
  • s
  • f
  • r
  • o
  • m
  • t
  • h
  • e
  • i
  • n
  • f
  • o
  • r
  • m
  • e
  • r
  • c
  • a
  • c
  • h
  • e
  • .
  • T
  • h
  • i
  • s
  • b
  • r
  • o
  • u
  • g
  • h
  • t
  • r
  • e
  • c
  • o
  • n
  • c
  • i
  • l
  • i
  • a
  • t
  • i
  • o
  • n
  • t
  • i
  • m
  • e
  • s
  • d
  • o
  • w
  • n
  • f
  • r
  • o
  • m
  • 3
  • 0
  • m
  • i
  • n
  • u
  • t
  • e
  • s
  • t
  • o
  • u
  • n
  • d
  • e
  • r
  • 1
  • 5
  • s
  • e
  • c
  • o
  • n
  • d
  • s
  • .
Advertisement
Want more CI/CD & GitOps scenarios?
Explore our complete collection of scenario-based CI/CD & GitOps interview runbooks.
Browse All CI/CD & GitOps Questions →