⚡ ~/naveed Interview Prep
⚡ Portfolio Home ✍️ Engineering Blog Deep Dives 🎯 Interview Hub 1,000+ Scenarios ☸️ Kubernetes Mastery Hub 24 Modules 🎮 DevOps Arcade & Quizzes Subnet Blitz ⚡ 🗺️ DevOps Roadmaps PDFs & Guides 🤖 Morpheus Analysis AI Quant ↗ 🛠️ Developer Tools Utilities 🧪 Labs & Experiments 📄 Interactive CV & Certs 🔗 All Links & Socials ⚡ Join The Dispatch (Weekly SRE Newsletter) →
← Back to All AI/ML Infrastructure & GPU Interview Questions Scenario 31 of 50 in AI/ML Infrastructure & GPU
Staff AI Infrastructure Engineer AI/ML Infrastructure Vector Search & Retrieval Vector Databases
🎯 Target Role / Context: Staff AI Infrastructure Engineer designing enterprise retrieval-augmented generation (RAG) vector infrastructure.

Q: A production RAG platform needs to search over 200 million 1536-dimensional vectors with sub-25ms p99 latency. How do you architect a distributed Milvus or Qdrant cluster on Kubernetes, and how do you choose between HNSW and IVF-PQ indexing to balance RAM costs and recall accuracy?

Deploying, tuning, and horizontally scaling distributed vector databases (Milvus / Qdrant) on Kubernetes, balancing RAM footprints for HNSW graphs versus quantized IVF-PQ indexes.

#Milvus #Qdrant #Vector Database #HNSW #IVF-PQ #Kubernetes #AI/ML Infra
🎙️ Candidate Opening & Architectural Context
"Vector databases form the retrieval backbone of RAG systems. However, high-dimensional embeddings (e.g., OpenAI 1536-dim or Cohere 1024-dim) have massive memory footprints: 200M vectors at 1536 dimensions in raw FP32 format require over 1.2 Terabytes of RAM before indexing graphs are even constructed. Engineering vector infrastructure requires partitioning query nodes from data nodes and selecting optimal index quantization."
Advertisement
⚡ Recommended Practice Lab

Want to master this scenario in a live sandbox? KodeKloud's CKA & CKAD Hands-On Certification Track covers this exact problem with hands-on terminal drills.

🛠️ Production Runbook & Step-by-Step Resolution

1

Architect Milvus Distributed Component Separation on Kubernetes

Deploy Milvus in distributed microservices mode via Milvus Operator: separate QueryNodes (stateless vector search execution), DataNodes (raw data ingestion and segment writing), IndexNodes (background index computation), and Coordinator nodes (RootCoord, QueryCoord). Use Apache Pulsar / Kafka as the WAL log broker and Amazon S3 as the persistent object storage backend.

apiVersion: milvus.io/v1beta1
kind: Milvus
metadata:
  name: milvus-production
spec:
  mode: cluster
  dependencies:
    etcd: { endpoints: ["etcd-cluster:2379"] }
    storage: { type: "s3", bucket: "milvus-vectors" }
    pulsar: { endpoint: "pulsar-proxy:6650" }
2

Evaluate HNSW vs IVF-PQ Index Memory and Recall Trade-Offs

HNSW (Hierarchical Navigable Small World) provides superior recall (>98%) and sub-10ms search latency, but constructs extensive memory-resident link graphs requiring 1.5x to 2x raw vector size in RAM (costing ~2.5TB RAM for 200M vectors). In contrast, IVF-PQ (Inverted File with Product Quantization) compresses vectors into quantized centroids and codes, reducing memory by 85-90% at the cost of 5-10% lower recall. For high-scale deployments, use Scalar Quantization (SQ8) or HNSW with on-disk vectors (DiskANN / Milvus Knowhere) backed by local NVMe SSDs.

# Index creation parameters for HNSW with SQ8 quantization
index_params = {
    "metric_type": "COSINE",
    "index_type": "HNSW",
    "params": {"M": 16, "efConstruction": 200}
}
Advertisement
3

Tune Segment Sizing and Horizontal Query Scaling

Configure Milvus data segment size to 512MB to balance compaction overhead and search parallelism. Horizontally scale stateless QueryNodes using Kubernetes HPA based on CPU and query latency metrics. QueryCoord automatically distributes segment replicas across QueryNodes, enabling linear search throughput scaling.

# QueryNode autoscaling policy
spec:
  components:
    queryNode:
      replicas: 12
      resources:
        requests: { cpu: "16", memory: "64Gi" }
💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"Distributed vector databases require separating stateless query nodes from background indexing nodes. HNSW provides maximum recall but demands huge RAM; combining HNSW with SQ8/DiskANN or IVF-PQ shrinks memory consumption by up to 90% while meeting tight latency SLAs."
⚡ 60-Second Elevator Pitch Talking Points
  • Storing 200M vectors in uncompressed HNSW graphs requires over 2.5TB of expensive RAM.
  • We deploy Milvus in distributed cluster mode on Kubernetes, decoupling QueryNodes from DataNodes over S3 and Pulsar.
  • By adopting SQ8 scalar quantization with HNSW graphs backed by local NVMe drives, we cut memory costs by 75% while maintaining 97% recall and sub-20ms search latency.
Advertisement
Want more AI/ML Infrastructure & GPU scenarios?
Explore our complete collection of scenario-based AI/ML Infrastructure & GPU interview runbooks.
Browse All AI/ML Infrastructure & GPU Questions →