⚡ ~/naveed Interview Prep
⚡ Portfolio Home ✍️ Engineering Blog Deep Dives 🎯 Interview Hub 1,000+ Scenarios ☸️ Kubernetes Mastery Hub 24 Modules 🎮 DevOps Arcade & Quizzes Subnet Blitz ⚡ 🗺️ DevOps Roadmaps PDFs & Guides 🤖 Morpheus Analysis AI Quant ↗ 🛠️ Developer Tools Utilities 🧪 Labs & Experiments 📄 Interactive CV & Certs 🔗 All Links & Socials ⚡ Join The Dispatch (Weekly SRE Newsletter) →
← Back to All FinOps & System Design Interview Questions Scenario 49 of 98 in FinOps & System Design
Staff Platform Architect System Design Platform Engineering & Multi-Tenancy System Design

Q: You are tasked with designing the global Kubernetes platform architecture for an enterprise with 500 engineering teams and 3,000 microservices. Hard multi-tenancy, zero-trust network boundaries, strict CPU/memory quotas, automated self-service onboarding, and exact cost attribution are non-negotiable. How do you design this platform from the physical node layer up to the developer experience?

End-to-end architectural blueprint for designing a secure, high-scale, multi-tenant Kubernetes platform serving 500+ engineering teams with hard compute isolation, virtual clusters (vcluster), policy enforcement, and chargeback metering.

#System Design #Kubernetes #Multi-Tenancy #vcluster #Kyverno #Platform Engineering
🎙️ Candidate Opening & Architectural Context
"When an enterprise scales to hundreds of teams, giving every team their own physical Kubernetes cluster creates catastrophic compute sprawl, management overhead, and millions in wasted idle node costs. We architected a hardened multi-tenant platform combining Virtual Clusters (vcluster), OPA/Kyverno admission policies, and OpenCost attribution."
Advertisement
⚡ Recommended Practice Lab

Want to master this scenario in a live sandbox? The Linux Foundation's FinOps Certified Practitioner (FOCP) Program covers this exact problem with hands-on terminal drills.

🛠️ Production Runbook & Step-by-Step Resolution

1️⃣

Establish Multi-Tenancy Model: Soft vs Hard Multi-Tenancy with Virtual Clusters

Evaluate tenant boundaries and isolate Kubernetes control planes:

  • Virtual Clusters (vcluster): Deployed lightweight virtual control planes inside tenant namespaces. Each team receives full admin access to their own vcluster control plane (CRDs, namespaces) without access to the underlying host cluster.
  • Kernel Isolation: Configured gVisor (runsc) and Kata Containers runtime classes for untrusted or external tenant workloads to prevent container escape exploits to the host Linux kernel.
Pro Tip: vcluster decouples tenant API server lifecycles from the underlying host cluster, eliminating the noisy-neighbor problem for cluster-scoped resources like CRDs and mutating webhooks.
2️⃣

Implement Admission Control Guardrails with Kyverno / Gatekeeper

Enforce strict organizational security policies declaratively:

  • Mandatory Security Context: Kyverno policies reject pods running as root (runAsNonRoot: true), requiring read-only root filesystems and dropping all Linux capabilities (drop: [ALL]).
  • Resource Quotas & LimitRanges: Automatically injected ResourceQuota and LimitRange manifests into every tenant namespace, setting default CPU/memory limits and blocking overcommit.
Pro Tip: Admission webhooks must be deployed in fail-closed mode for security policies, backed by multiple webhook replicas and automated PDBs to avoid cluster lockouts.
3️⃣

Enforce Zero-Trust Microsegmentation with Cilium eBPF

Isolate tenant network traffic without iptables bottlenecks:

  • Default Deny Ingress/Egress: Applied baseline CiliumClusterwideNetworkPolicy isolating all namespaces by default.
  • Tenant Peer Communication: Allowed cross-tenant API communication strictly through explicit Layer 7 mutual TLS (mTLS) policies authenticated via SPIFFE/SPIRE IDs.
Pro Tip: Cilium eBPF enforces network policies at the Linux socket layer, avoiding the severe CPU penalties of thousands of iptables chains across 3,000 services.
4️⃣

Deploy Real-Time Cost Attribution & Chargeback with OpenCost

Attribute compute, storage, and network egress costs directly to engineering cost centers:

  • OpenCost Deployment: Deployed OpenCost mapped to cloud provider billing APIs (AWS CUR / Azure Cost Management).
  • Tenant Metering: Calculated exact hourly cost per team based on requested vs actual utilized CPU/RAM and cross-AZ egress bytes, exporting monthly chargeback metrics directly to Jira and ERP systems.
Pro Tip: Granular chargeback drives engineering accountability, immediately eliminating idle zombie staging deployments.
💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"Modern enterprise multi-tenant Kubernetes relies on Virtual Clusters (vcluster) for control plane isolation, gVisor for runtime security, Kyverno for guardrails, Cilium eBPF for zero-trust networking, and OpenCost for billing attribution."
⚡ 60-Second Elevator Pitch Talking Points
  • Provide isolated virtual control planes per engineering team using vcluster.
  • Enforce rootless containers and strict resource quotas via Kyverno admission policies.
  • Deploy Cilium eBPF for zero-trust default-deny network microsegmentation and mTLS.
  • Implement OpenCost to deliver transparent, automated multi-tenant cost chargeback.
Advertisement
Want more FinOps & System Design scenarios?
Explore our complete collection of scenario-based FinOps & System Design interview runbooks.
Browse All FinOps & System Design Questions →