⚡ ~/naveed Interview Prep
⚡ Portfolio Home ✍️ Engineering Blog Deep Dives 🎯 Interview Hub 1,000+ Scenarios ☸️ Kubernetes Mastery Hub 24 Modules 🎮 DevOps Arcade & Quizzes Subnet Blitz ⚡ 🗺️ DevOps Roadmaps PDFs & Guides 🤖 Morpheus Analysis AI Quant ↗ 🛠️ Developer Tools Utilities 🧪 Labs & Experiments 📄 Interactive CV & Certs 🔗 All Links & Socials ⚡ Join The Dispatch (Weekly SRE Newsletter) →
← Back to All Platform Engineering & IDP Interview Questions Scenario 26 of 50 in Platform Engineering & IDP
Staff Platform Engineer Platform Engineering Control Planes & Disaster Recovery Platform DR
🎯 Target Role / Context: Staff Platform Engineer Interview · Platform Resilience Track

Q: If your primary management Kubernetes cluster hosting Crossplane, Backstage, and ArgoCD suffers a total etcd data corruption or region failure, how do you restore your platform control plane without triggering cloud resource deletion?

Designing active disaster recovery and state restoration workflows for Kubernetes-based internal developer platform control planes.

#Platform Engineering #Disaster Recovery #Crossplane #Velero #Backstage #Kubernetes
🎙️ Candidate Opening & Architectural Context
"Losing a Crossplane management cluster is dangerous because re-applying Crossplane state can accidentally trigger resource recreation or orphaned cloud resources. Disaster recovery requires consistent Velero CRD backups paired with strict `DeletionPolicy=Orphan` safeguards and automated PostgreSQL database backups."
Advertisement
⚡ Recommended Practice Lab

Want to master this scenario in a live sandbox? KodeKloud's CKA & CKAD Hands-On Certification Track covers this exact problem with hands-on terminal drills.

🛠️ Production Runbook & Step-by-Step Resolution

1

Enforce DeletionPolicy: Orphan on All Crossplane Resources

Ensure all Composite Resources and Managed Resources define `spec.deletionPolicy: Orphan`. If the cluster dies or CRDs are deleted during DR testing, Crossplane does not issue cloud API DELETE commands against live production AWS RDS or VPCs.

apiVersion: database.aws.upjet.crossplane.io/v1beta1
kind: Instance
spec:
  deletionPolicy: Orphan # Prevents accidental deletion during DR re-sync
2

Automate Velero CRD and State Backups

Configure Velero with daily scheduled backups covering Crossplane CRDs, Secrets, and ArgoCD applications to an isolated S3 bucket in a secondary cloud region.

velero schedule create platform-dr-daily \
  --include-namespaces crossplane-system,argocd,backstage \
  --include-resources customresourcedefinitions,secrets,configmaps \
  --schedule="0 2 * * *"
Advertisement
3

Test RTO Recovery Drills in Isolated Sandbox

Execute quarterly restore drills into an empty EKS cluster: restore Velero backup, reconnect external PostgreSQL, and verify Crossplane binds to existing cloud resources via `crossplane.io/external-name` without recreating them.

Pro Tip: RTO Target: A resilient platform control plane must restore within 30 minutes with 0 minutes of data loss on active cloud resources.
💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"Protect Crossplane control planes with DeletionPolicy=Orphan on all resources, automated Velero backups to a secondary region, and externalized Backstage PostgreSQL."
⚡ 60-Second Elevator Pitch Talking Points
  • Enforce deletionPolicy: Orphan across all Crossplane manifests to prevent accidental cloud deletion during recovery.
  • Take scheduled Velero backups of CRDs, secrets, and cluster state to a disaster recovery S3 bucket.
  • Store Backstage catalog data in a managed Multi-AZ PostgreSQL instance separate from cluster worker nodes.
Advertisement
Want more Platform Engineering & IDP scenarios?
Explore our complete collection of scenario-based Platform Engineering & IDP interview runbooks.
Browse All Platform Engineering & IDP Questions →