⚡ ~/naveed Interview Prep
⚡ Portfolio Home ✍️ Engineering Blog Deep Dives 🎯 Interview Hub 1,000+ Scenarios ☸️ Kubernetes Mastery Hub 24 Modules 🎮 DevOps Arcade & Quizzes Subnet Blitz ⚡ 🗺️ DevOps Roadmaps PDFs & Guides 🤖 Morpheus Analysis AI Quant ↗ 🛠️ Developer Tools Utilities 🧪 Labs & Experiments 📄 Interactive CV & Certs 🔗 All Links & Socials ⚡ Join The Dispatch (Weekly SRE Newsletter) →
← Back to All AI/ML Infrastructure & GPU Interview Questions Scenario 28 of 50 in AI/ML Infrastructure & GPU
Senior AI Infrastructure Engineer AI/ML Infrastructure Developer Platforms & Notebooks Developer Platforms
🎯 Target Role / Context: Senior AI Infrastructure Engineer building self-service research environments for hundreds of data scientists.

Q: Data scientists often leave idle GPU Jupyter notebooks running overnight, wasting tens of thousands of dollars in cloud spend. How do you architect a Kubernetes notebook platform with fractional GPU sharing, shared EFS home directories, and automated idle culler daemons?

Engineering a resilient, cost-effective multi-tenant JupyterHub / Kubeflow Notebook platform on Kubernetes with fractional GPU sharing, dynamic storage provisioning, and automated zombie notebook culling.

#JupyterHub #Kubeflow Notebooks #GPU Sharing #MIG #Idle Culler #AI/ML Infra
🎙️ Candidate Opening & Architectural Context
"Interactive exploration via Jupyter Notebooks is the cornerstone of data science and AI experimentation. However, dedicating an unshared A100/H100 GPU to every notebook is economically unviable, as interactive exploration operates at < 5% GPU utilization. The platform must enforce resource sharing, persistent state persistence across restarts, and proactive reclamation of abandoned instances."
Advertisement
⚡ Recommended Practice Lab

Want to master this scenario in a live sandbox? KodeKloud's CKA & CKAD Hands-On Certification Track covers this exact problem with hands-on terminal drills.

🛠️ Production Runbook & Step-by-Step Resolution

1

Implement Fractional GPU Slicing via NVIDIA Time-Slicing or MIG

For development and experimentation workloads where strict QoS isolation is not required, configure NVIDIA GPU Operator time-slicing (e.g., sharing a single physical GPU across 4 notebook pods). For advanced users requiring dedicated hardware isolation, vend NVIDIA MIG slices (`1g.10gb`) to protect interactive workloads from noisy neighbor memory crashes.

# NVIDIA GPU Operator time-slicing config
apiVersion: v1
kind: ConfigMap
metadata:
  name: time-slicing-config
data:
  any:
    sharing:
      timeSlicing:
        resources:
          - name: nvidia.com/gpu
            replicas: 4
2

Dynamic Home Directory Storage via Amazon EFS CSI Driver

Data scientists must not lose uncommitted code, notebooks, and datasets when pods restart. Deploy Amazon EFS (Elastic File System) with dynamic PV provisioning via the EFS CSI driver. Mount a dedicated user directory (`/home/jovyan`) into each notebook pod, providing persistent POSIX shared storage that persists across pod rescheduling.

apiVersion: v1
kind: PersistentVolumeClaim
metadata:
  name: user-home-claim
spec:
  accessModes: [ReadWriteMany]
  storageClassName: efs-sc
  resources:
    requests:
      storage: 50Gi
Advertisement
3

Automate Idle Notebook Culling and Scale-to-Zero

Deploy the JupyterHub Idle Culler (`jupyterhub-idle-culler`) as a background service. Configure it to inspect kernel API activity. If a notebook has no active kernel execution and no browser WebSocket connection for more than 60 minutes, the culler terminates the pod, releasing the GPU back to the cluster pool while preserving all file edits on EFS.

# JupyterHub Helm config for idle culler
cull:
  enabled: true
  timeout: 3600 # 60 minutes
  every: 300   # check every 5 minutes
  users: false
  removeNamedServers: true
💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"Multi-tenant notebook platforms must leverage GPU time-slicing or MIG for shared development access, persist state via EFS CSI mounts, and deploy automated idle cullers to terminate inactive notebooks and reclaim expensive GPU capacity."
⚡ 60-Second Elevator Pitch Talking Points
  • Dedicated GPUs for interactive notebooks burn massive budgets on idle waiting time.
  • We configure GPU time-slicing to share each GPU across 4 data scientists for interactive exploration.
  • User home directories persist on Amazon EFS, and an automated idle culler terminates notebooks after 60 minutes of inactivity, slashing our development GPU bill by 72%.
Advertisement
Want more AI/ML Infrastructure & GPU scenarios?
Explore our complete collection of scenario-based AI/ML Infrastructure & GPU interview runbooks.
Browse All AI/ML Infrastructure & GPU Questions →