Q: How do you maintain identical security configurations, RBAC roles, and network policies across 40 GKE clusters spread across Google Cloud and on-premises environments without manual drift? Walk me through Anthos Config Sync.
Deploying declarative multi-cluster governance across hybrid GKE and on-premises clusters using Anthos Config Sync (RootSync/RepoSync) and Gatekeeper Policy Controller.
Want to master this scenario in a live sandbox? Stephane Maarek's AWS Certified DevOps Engineer Professional Masterclass on Udemy covers this exact problem with hands-on terminal drills.
🛠️ Production Runbook & Step-by-Step Resolution
Register Clusters to Anthos Fleet (GKE Hub)
Establish a unified administrative plane across all distributed clusters:
- Fleet Membership: Registered all cloud and on-prem clusters to the central Google Cloud Project Fleet using
gcloud container fleet memberships register. - Connect Gateway: Enabled Google Connect Gateway to provide unified IAM-based administrative access to all clusters through Google Cloud.
Configure Config Sync with RootSync and RepoSync CRDs
Implement hierarchical multi-tenant GitOps synchronization:
- RootSync (Platform Team): Configured central
RootSyncpointing to the cluster-admin git repository enforcing cluster-wide CRDs, NetworkPolicies, and RBAC. - RepoSync (Product Squads): Configured namespace-scoped
RepoSyncallowing individual product teams to manage their own application deployments in isolated subdirectories. - Authentication: Authenticated Config Sync to GitHub Enterprise via Workload Identity and GitHub App tokens.
Enforce Declarative Policy Controller (OPA Gatekeeper)
Block non-compliant resources at the admission level:
- ConstraintTemplates: Deployed Google-curated policy templates (e.g.
K8sRequiredLabels,K8sDenyPrivileged,K8sRestrictedPorts). - Enforcement Actions: Configured
enforcementAction: denyon production clusters anddryrunon development clusters for progressive rollout. - Audit Dashboard: Monitored policy compliance metrics directly inside the Google Cloud Console Anthos dashboard.
Validate Automated Drift Correction & Self-Healing
Verify that unauthorized changes are automatically overwritten:
- Drift Simulation: Manually deleted a required NetworkPolicy using
kubectl delete networkpolicy deny-egress. - Self-Healing: Within 15 seconds, Config Sync's in-cluster reconciler detected the divergence and recreated the resource.
- CLI Status: Checked synchronization state with
nomos status.
- Registered distributed GKE clusters into an Anthos Fleet with Connect Gateway.
- Configured hierarchical GitOps using RootSync for cluster-admin rules and RepoSync for tenant squads.
- Enforced OPA Gatekeeper Policy Controller constraints to block non-compliant resources at admission time.
- Verified automated 15-second self-healing drift correction using nomos status.