⚡ ~/naveed Interview Prep
⚡ Portfolio Home ✍️ Engineering Blog Deep Dives 🎯 Interview Hub 1,000+ Scenarios ☸️ Kubernetes Mastery Hub 24 Modules 🎮 DevOps Arcade & Quizzes Subnet Blitz ⚡ 🗺️ DevOps Roadmaps PDFs & Guides 🤖 Morpheus Analysis AI Quant ↗ 🛠️ Developer Tools Utilities 🧪 Labs & Experiments 📄 Interactive CV & Certs 🔗 All Links & Socials ⚡ Join The Dispatch (Weekly SRE Newsletter) →
← Back to All AI/ML Infrastructure & GPU Interview Questions Scenario 9 of 50 in AI/ML Infrastructure & GPU
Staff AI Infrastructure Engineer AI/ML Infrastructure Distributed Compute & Ray Ray & Distributed Compute
🎯 Target Role / Context: Staff AI Infrastructure Engineer building multi-tenant distributed Python compute platforms for AI teams.

Q: A distributed data processing pipeline on KubeRay experiences random worker pod OOMKilled events that trigger cascading failures across the entire RayCluster. How do you distinguish between Ray Object Store (Plasma) spills vs Linux cgroup worker process memory limits, and how do you configure resilient actor placement?

Deploying and managing KubeRay on Kubernetes, optimizing the Plasma Shared-Memory Object Store, actor scheduling, and triaging worker pod OOMKills without crashing distributed training runs.

#KubeRay #RayCluster #Plasma Store #OOM #Kubernetes #AI/ML Infra
🎙️ Candidate Opening & Architectural Context
"Ray provides an actor-based distributed execution runtime for Python, popular for reinforcement learning (RLHF), distributed batch inference, and data preprocessing. KubeRay translates Ray clusters into Kubernetes CRDs (RayCluster, RayJob). However, Ray's unique memory architecture—combining worker process heap memory with an in-memory Plasma Shared-Memory Object Store backed by `/dev/shm`—frequently confounds standard Kubernetes cgroup memory management."
Advertisement
⚡ Recommended Practice Lab

Want to master this scenario in a live sandbox? KodeKloud's CKA & CKAD Hands-On Certification Track covers this exact problem with hands-on terminal drills.

🛠️ Production Runbook & Step-by-Step Resolution

1

Understand Ray Memory Split: Worker Heap vs Plasma Object Store

Each Ray node allocates a fraction of its memory to the Plasma Object Store (typically 30% of total node RAM) mounted over shared memory (/dev/shm) for zero-copy deserialization of immutable NumPy/Arrow arrays. The remaining memory is used by Python worker processes. If the sum of worker heap allocations and Plasma usage exceeds the container's Kubernetes memory limits, the Linux kernel OOM Killer abruptly terminates the pod.

# RayCluster CRD workerGroupSpec with explicit shm sizing
spec:
  rayVersion: '2.35.0'
  workerGroupSpecs:
  - groupName: gpu-group
    replicas: 8
    template:
      spec:
        volumes:
        - name: dshm
          emptyDir:
            medium: Memory
            sizeLimit: 64Gi
        containers:
        - name: ray-worker
          image: rayproject/ray-ml:2.35.0-py310-gpu
          volumeMounts:
          - mountPath: /dev/shm
            name: dshm
2

Configure Object Spilling to Local NVMe or S3

When the Plasma store fills up, Ray evicts objects. Without object spilling configured, Ray throws an ObjectStoreFullError or worker processes crash. Configure Ray object spilling to local high-speed NVMe storage or S3 so excess objects stream to disk seamlessly rather than exhausting RAM.

# Ray system config for object spilling
ray start --head --system-config='{
  "object_spilling_config": json.dumps({
    "type": "filesystem",
    "params": {"directory_path": "/tmp/ray/spill_dir"}
  })
}'
Advertisement
3

Implement Fault-Tolerant Actor Placement and Re-execution

Configure Ray actor lifecycle decorators: `@ray.remote(max_restarts=3, max_task_retries=3)`. Use Ray Placement Groups to enforce STRICT_SPREAD across Kubernetes worker nodes for high-availability actors. If an individual worker pod is OOMKilled, the Ray GCS (Global Control Store) on the head node detects the lost heartbeat and restarts the failed actor on an alternate worker without aborting the entire RayJob.

@ray.remote(max_restarts=3, max_task_retries=3)
class WorkerActor:
    def __init__(self):
        pass
    def process_batch(self, batch):
        return compute(batch)
💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"Ray pods require properly sized `/dev/shm` volumes for the Plasma object store alongside worker heap headroom. Configuring Ray object spilling to NVMe and enabling actor auto-restarts isolates worker failures and prevents cluster crashes."
⚡ 60-Second Elevator Pitch Talking Points
  • Ray workers get OOMKilled when shared memory Plasma stores collide with Python process heaps inside the same Kubernetes cgroup.
  • We configure explicit `/dev/shm` `emptyDir` memory mounts and enable Ray object spilling to local NVMe instance storage.
  • By setting `@ray.remote(max_restarts=3)` and strict placement groups, single-worker memory spikes recover automatically without terminating multi-hour distributed pipelines.
Advertisement
Want more AI/ML Infrastructure & GPU scenarios?
Explore our complete collection of scenario-based AI/ML Infrastructure & GPU interview runbooks.
Browse All AI/ML Infrastructure & GPU Questions →