Q: A distributed data processing pipeline on KubeRay experiences random worker pod OOMKilled events that trigger cascading failures across the entire RayCluster. How do you distinguish between Ray Object Store (Plasma) spills vs Linux cgroup worker process memory limits, and how do you configure resilient actor placement?
Deploying and managing KubeRay on Kubernetes, optimizing the Plasma Shared-Memory Object Store, actor scheduling, and triaging worker pod OOMKills without crashing distributed training runs.
Want to master this scenario in a live sandbox? KodeKloud's CKA & CKAD Hands-On Certification Track covers this exact problem with hands-on terminal drills.
🛠️ Production Runbook & Step-by-Step Resolution
Understand Ray Memory Split: Worker Heap vs Plasma Object Store
Each Ray node allocates a fraction of its memory to the Plasma Object Store (typically 30% of total node RAM) mounted over shared memory (/dev/shm) for zero-copy deserialization of immutable NumPy/Arrow arrays. The remaining memory is used by Python worker processes. If the sum of worker heap allocations and Plasma usage exceeds the container's Kubernetes memory limits, the Linux kernel OOM Killer abruptly terminates the pod.
# RayCluster CRD workerGroupSpec with explicit shm sizing
spec:
rayVersion: '2.35.0'
workerGroupSpecs:
- groupName: gpu-group
replicas: 8
template:
spec:
volumes:
- name: dshm
emptyDir:
medium: Memory
sizeLimit: 64Gi
containers:
- name: ray-worker
image: rayproject/ray-ml:2.35.0-py310-gpu
volumeMounts:
- mountPath: /dev/shm
name: dshm
Configure Object Spilling to Local NVMe or S3
When the Plasma store fills up, Ray evicts objects. Without object spilling configured, Ray throws an ObjectStoreFullError or worker processes crash. Configure Ray object spilling to local high-speed NVMe storage or S3 so excess objects stream to disk seamlessly rather than exhausting RAM.
# Ray system config for object spilling
ray start --head --system-config='{
"object_spilling_config": json.dumps({
"type": "filesystem",
"params": {"directory_path": "/tmp/ray/spill_dir"}
})
}'
Implement Fault-Tolerant Actor Placement and Re-execution
Configure Ray actor lifecycle decorators: `@ray.remote(max_restarts=3, max_task_retries=3)`. Use Ray Placement Groups to enforce STRICT_SPREAD across Kubernetes worker nodes for high-availability actors. If an individual worker pod is OOMKilled, the Ray GCS (Global Control Store) on the head node detects the lost heartbeat and restarts the failed actor on an alternate worker without aborting the entire RayJob.
@ray.remote(max_restarts=3, max_task_retries=3)
class WorkerActor:
def __init__(self):
pass
def process_batch(self, batch):
return compute(batch)
- Ray workers get OOMKilled when shared memory Plasma stores collide with Python process heaps inside the same Kubernetes cgroup.
- We configure explicit `/dev/shm` `emptyDir` memory mounts and enable Ray object spilling to local NVMe instance storage.
- By setting `@ray.remote(max_restarts=3)` and strict placement groups, single-worker memory spikes recover automatically without terminating multi-hour distributed pipelines.