Q: Standard Kubernetes Horizontal Pod Autoscaler (HPA) using CPU or GPU duty cycle fails for LLM serving because GPUs remain pinned at 100% even during slow token streaming. How do you design autoscaling for Ray Serve using KEDA and custom Prometheus metrics like queue depth and TTFT?
Engineering autoscaling policies for Ray Serve clusters on Kubernetes, driving horizontal replica scaling using request queue depth, pending actors, and inter-token latency instead of CPU/memory metrics.
Want to master this scenario in a live sandbox? KodeKloud's CKA & CKAD Hands-On Certification Track covers this exact problem with hands-on terminal drills.
🛠️ Production Runbook & Step-by-Step Resolution
Understand Ray Serve Native Autoscaling Architecture
Ray Serve provides an internal application-level autoscaler that runs on the Ray head node. It monitors the Ray Serve HTTP proxy queue depth (`num_ongoing_requests_per_replica`) and calculates desired replica counts dynamically based on target concurrency thresholds, spinning up or tearing down worker actors inside the Ray cluster.
# Ray Serve Deployment autoscaling configuration
@serve.deployment(
autoscaling_config={
"min_replicas": 2,
"max_replicas": 16,
"target_ongoing_requests": 10,
"upscale_delay_s": 15,
"downscale_delay_s": 300,
}
)
class LLMDeployment:
def __init__(self):
pass
Bridge Ray Cluster Worker Pods with Kubernetes KEDA
While Ray Serve scales internal actors, the underlying Kubernetes RayCluster worker pods must also scale to provide physical GPU nodes. Deploy KEDA (Kubernetes Event-driven Autoscaling). Configure a KEDA ScaledObject that queries Prometheus for `ray_serve_deployment_queued_queries` or pending Ray placement groups, scaling KubeRay worker pod replicas dynamically.
apiVersion: keda.sh/v1alpha1
kind: ScaledObject
metadata:
name: ray-worker-gpu-scaler
spec:
scaleTargetRef:
apiVersion: ray.io/v1
kind: RayCluster
name: raycluster-production
minReplicaCount: 2
maxReplicaCount: 20
triggers:
- type: prometheus
metadata:
serverAddress: http://prometheus-k8s.monitoring:9090
metricName: ray_serve_queued_queries
query: sum(ray_serve_deployment_queued_queries{deployment="LLMDeployment"})
threshold: '5'
Configure Fast Upscaling and Cautious Downscaling (Hysteresis)
Because provisioning GPU instances takes 1-2 minutes, configure aggressive upscaling (reacting within 10 seconds of queue growth) paired with substantial downscaling hysteresis (10-15 minute stabilization window). This prevents costly thrashing and premature termination of pods that still have cached model weights in VRAM.
advanced:
horizontalPodAutoscalerConfig:
behavior:
scaleDown:
stabilizationWindowSeconds: 600
policies:
- type: Percent
value: 10
periodSeconds: 60
- Standard HPA scaling on GPU utilization fails because LLM engines run at 95% utilization whether processing 1 request or 100.
- We scale Ray Serve actors based on `target_ongoing_requests` per replica, directly reflecting concurrency pressure.
- KEDA monitors proxy queue depth in Prometheus to dynamically scale KubeRay worker pods and Karpenter GPU nodes, ensuring zero queue buildup during burst traffic.