Q: Data scientists often leave idle GPU Jupyter notebooks running overnight, wasting tens of thousands of dollars in cloud spend. How do you architect a Kubernetes notebook platform with fractional GPU sharing, shared EFS home directories, and automated idle culler daemons?
Engineering a resilient, cost-effective multi-tenant JupyterHub / Kubeflow Notebook platform on Kubernetes with fractional GPU sharing, dynamic storage provisioning, and automated zombie notebook culling.
Want to master this scenario in a live sandbox? KodeKloud's CKA & CKAD Hands-On Certification Track covers this exact problem with hands-on terminal drills.
🛠️ Production Runbook & Step-by-Step Resolution
Implement Fractional GPU Slicing via NVIDIA Time-Slicing or MIG
For development and experimentation workloads where strict QoS isolation is not required, configure NVIDIA GPU Operator time-slicing (e.g., sharing a single physical GPU across 4 notebook pods). For advanced users requiring dedicated hardware isolation, vend NVIDIA MIG slices (`1g.10gb`) to protect interactive workloads from noisy neighbor memory crashes.
# NVIDIA GPU Operator time-slicing config
apiVersion: v1
kind: ConfigMap
metadata:
name: time-slicing-config
data:
any:
sharing:
timeSlicing:
resources:
- name: nvidia.com/gpu
replicas: 4
Dynamic Home Directory Storage via Amazon EFS CSI Driver
Data scientists must not lose uncommitted code, notebooks, and datasets when pods restart. Deploy Amazon EFS (Elastic File System) with dynamic PV provisioning via the EFS CSI driver. Mount a dedicated user directory (`/home/jovyan`) into each notebook pod, providing persistent POSIX shared storage that persists across pod rescheduling.
apiVersion: v1
kind: PersistentVolumeClaim
metadata:
name: user-home-claim
spec:
accessModes: [ReadWriteMany]
storageClassName: efs-sc
resources:
requests:
storage: 50Gi
Automate Idle Notebook Culling and Scale-to-Zero
Deploy the JupyterHub Idle Culler (`jupyterhub-idle-culler`) as a background service. Configure it to inspect kernel API activity. If a notebook has no active kernel execution and no browser WebSocket connection for more than 60 minutes, the culler terminates the pod, releasing the GPU back to the cluster pool while preserving all file edits on EFS.
# JupyterHub Helm config for idle culler
cull:
enabled: true
timeout: 3600 # 60 minutes
every: 300 # check every 5 minutes
users: false
removeNamedServers: true
- Dedicated GPUs for interactive notebooks burn massive budgets on idle waiting time.
- We configure GPU time-slicing to share each GPU across 4 data scientists for interactive exploration.
- User home directories persist on Amazon EFS, and an automated idle culler terminates notebooks after 60 minutes of inactivity, slashing our development GPU bill by 72%.