Q: If your primary management Kubernetes cluster hosting Crossplane, Backstage, and ArgoCD suffers a total etcd data corruption or region failure, how do you restore your platform control plane without triggering cloud resource deletion?
Designing active disaster recovery and state restoration workflows for Kubernetes-based internal developer platform control planes.
Want to master this scenario in a live sandbox? KodeKloud's CKA & CKAD Hands-On Certification Track covers this exact problem with hands-on terminal drills.
🛠️ Production Runbook & Step-by-Step Resolution
Enforce DeletionPolicy: Orphan on All Crossplane Resources
Ensure all Composite Resources and Managed Resources define `spec.deletionPolicy: Orphan`. If the cluster dies or CRDs are deleted during DR testing, Crossplane does not issue cloud API DELETE commands against live production AWS RDS or VPCs.
apiVersion: database.aws.upjet.crossplane.io/v1beta1
kind: Instance
spec:
deletionPolicy: Orphan # Prevents accidental deletion during DR re-sync
Automate Velero CRD and State Backups
Configure Velero with daily scheduled backups covering Crossplane CRDs, Secrets, and ArgoCD applications to an isolated S3 bucket in a secondary cloud region.
velero schedule create platform-dr-daily \
--include-namespaces crossplane-system,argocd,backstage \
--include-resources customresourcedefinitions,secrets,configmaps \
--schedule="0 2 * * *"
Test RTO Recovery Drills in Isolated Sandbox
Execute quarterly restore drills into an empty EKS cluster: restore Velero backup, reconnect external PostgreSQL, and verify Crossplane binds to existing cloud resources via `crossplane.io/external-name` without recreating them.
- Enforce deletionPolicy: Orphan across all Crossplane manifests to prevent accidental cloud deletion during recovery.
- Take scheduled Velero backups of CRDs, secrets, and cluster state to a disaster recovery S3 bucket.
- Store Backstage catalog data in a managed Multi-AZ PostgreSQL instance separate from cluster worker nodes.