⚡ ~/naveed Interview Prep
⚡ Portfolio Home ✍️ Engineering Blog Deep Dives 🎯 Interview Hub 1,000+ Scenarios ☸️ Kubernetes Mastery Hub 24 Modules 🎮 DevOps Arcade & Quizzes Subnet Blitz ⚡ 🗺️ DevOps Roadmaps PDFs & Guides 🤖 Morpheus Analysis AI Quant ↗ 🛠️ Developer Tools Utilities 🧪 Labs & Experiments 📄 Interactive CV & Certs 🔗 All Links & Socials ⚡ Join The Dispatch (Weekly SRE Newsletter) →
← Back to All AI/ML Infrastructure & GPU Interview Questions Scenario 23 of 50 in AI/ML Infrastructure & GPU
Staff AI Infrastructure Engineer AI/ML Infrastructure Autoscaling & Node Management Ray & Autoscaling
🎯 Target Role / Context: Staff AI Platform Engineer building auto-scaling multi-tenant Ray Serve infrastructure.

Q: Standard Kubernetes Horizontal Pod Autoscaler (HPA) using CPU or GPU duty cycle fails for LLM serving because GPUs remain pinned at 100% even during slow token streaming. How do you design autoscaling for Ray Serve using KEDA and custom Prometheus metrics like queue depth and TTFT?

Engineering autoscaling policies for Ray Serve clusters on Kubernetes, driving horizontal replica scaling using request queue depth, pending actors, and inter-token latency instead of CPU/memory metrics.

#Ray Serve #KubeRay #KEDA #Autoscaling #Queue Depth #AI/ML Infra
🎙️ Candidate Opening & Architectural Context
"In traditional web applications, CPU utilization reflects load. In LLM and deep learning inference, GPU duty cycle is a lagging or deceptive indicator: a GPU serving a single user with continuous batching can register 95% utilization, while a cluster queuing 100 pending requests might show identical GPU utilization. Autoscaling inference replicas must be governed by request queue length and latency SLO violations."
Advertisement
⚡ Recommended Practice Lab

Want to master this scenario in a live sandbox? KodeKloud's CKA & CKAD Hands-On Certification Track covers this exact problem with hands-on terminal drills.

🛠️ Production Runbook & Step-by-Step Resolution

1

Understand Ray Serve Native Autoscaling Architecture

Ray Serve provides an internal application-level autoscaler that runs on the Ray head node. It monitors the Ray Serve HTTP proxy queue depth (`num_ongoing_requests_per_replica`) and calculates desired replica counts dynamically based on target concurrency thresholds, spinning up or tearing down worker actors inside the Ray cluster.

# Ray Serve Deployment autoscaling configuration
@serve.deployment(
    autoscaling_config={
        "min_replicas": 2,
        "max_replicas": 16,
        "target_ongoing_requests": 10,
        "upscale_delay_s": 15,
        "downscale_delay_s": 300,
    }
)
class LLMDeployment:
    def __init__(self):
        pass
2

Bridge Ray Cluster Worker Pods with Kubernetes KEDA

While Ray Serve scales internal actors, the underlying Kubernetes RayCluster worker pods must also scale to provide physical GPU nodes. Deploy KEDA (Kubernetes Event-driven Autoscaling). Configure a KEDA ScaledObject that queries Prometheus for `ray_serve_deployment_queued_queries` or pending Ray placement groups, scaling KubeRay worker pod replicas dynamically.

apiVersion: keda.sh/v1alpha1
kind: ScaledObject
metadata:
  name: ray-worker-gpu-scaler
spec:
  scaleTargetRef:
    apiVersion: ray.io/v1
    kind: RayCluster
    name: raycluster-production
  minReplicaCount: 2
  maxReplicaCount: 20
  triggers:
  - type: prometheus
    metadata:
      serverAddress: http://prometheus-k8s.monitoring:9090
      metricName: ray_serve_queued_queries
      query: sum(ray_serve_deployment_queued_queries{deployment="LLMDeployment"})
      threshold: '5'
Advertisement
3

Configure Fast Upscaling and Cautious Downscaling (Hysteresis)

Because provisioning GPU instances takes 1-2 minutes, configure aggressive upscaling (reacting within 10 seconds of queue growth) paired with substantial downscaling hysteresis (10-15 minute stabilization window). This prevents costly thrashing and premature termination of pods that still have cached model weights in VRAM.

advanced:
  horizontalPodAutoscalerConfig:
    behavior:
      scaleDown:
        stabilizationWindowSeconds: 600
        policies:
        - type: Percent
          value: 10
          periodSeconds: 60
💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"Autoscaling GPU inference on CPU/GPU metrics fails because GPUs run near 100% regardless of queue pressure. Use Ray Serve target ongoing requests and KEDA Prometheus triggers on queue depth to drive elastic scaling."
⚡ 60-Second Elevator Pitch Talking Points
  • Standard HPA scaling on GPU utilization fails because LLM engines run at 95% utilization whether processing 1 request or 100.
  • We scale Ray Serve actors based on `target_ongoing_requests` per replica, directly reflecting concurrency pressure.
  • KEDA monitors proxy queue depth in Prometheus to dynamically scale KubeRay worker pods and Karpenter GPU nodes, ensuring zero queue buildup during burst traffic.
Advertisement
Want more AI/ML Infrastructure & GPU scenarios?
Explore our complete collection of scenario-based AI/ML Infrastructure & GPU interview runbooks.
Browse All AI/ML Infrastructure & GPU Questions →