⚡ ~/naveed Interview Prep
⚡ Portfolio Home ✍️ Engineering Blog Deep Dives 🎯 Interview Hub 1,000+ Scenarios ☸️ Kubernetes Mastery Hub 24 Modules 🎮 DevOps Arcade & Quizzes Subnet Blitz ⚡ 🗺️ DevOps Roadmaps PDFs & Guides 🤖 Morpheus Analysis AI Quant ↗ 🛠️ Developer Tools Utilities 🧪 Labs & Experiments 📄 Interactive CV & Certs 🔗 All Links & Socials ⚡ Join The Dispatch (Weekly SRE Newsletter) →
← Back to All FinOps & System Design Interview Questions Scenario 75 of 98 in FinOps & System Design
Staff Distributed Systems Architect System Design AI Infrastructure & Vector Search System Design

Q: Your enterprise Retrieval-Augmented Generation (RAG) platform must index 100 Million documents with 1536-dimensional embeddings (OpenAI text-embedding-3). The vector database must support 5,000 search queries per second with p99 latency < 15ms and filtered metadata search (e.g., tenant_id, role). How do you architect a scalable, highly available vector search tier without exhausting memory or breaking the bank?

Engineering a high-scale vector search architecture indexing 100 million high-dimensional embeddings with Qdrant / Milvus, HNSW indexing, quantization, and sub-10ms similarity search for real-time RAG pipelines.

#System Design #Vector Database #Qdrant #Milvus #RAG #Embeddings #HNSW
🎙️ Candidate Opening & Architectural Context
"Storing 100 million raw 1536-dimensional float32 vectors in RAM requires over 600 GB of pure memory just for raw numbers, ballooning to multi-terabytes when building HNSW graph indexes. We architected a scalable vector search tier using Qdrant/Milvus, Scalar Quantization, and distributed sharding."
Advertisement
⚡ Recommended Practice Lab

Want to master this scenario in a live sandbox? The Linux Foundation's FinOps Certified Practitioner (FOCP) Program covers this exact problem with hands-on terminal drills.

🛠️ Production Runbook & Step-by-Step Resolution

1️⃣

Select Vector Indexing & Compression: HNSW with Scalar Quantization

Balance recall accuracy against massive memory consumption:

  • HNSW Graph Index: Chose Hierarchical Navigable Small World (HNSW) graphs with m=16 and ef_construct=100 for sub-10ms approximate nearest neighbor (ANN) search.
  • Scalar Quantization (SQ): Quantized 32-bit floating point vectors (float32) down to 8-bit integers (int8), achieving a 4x reduction in RAM requirements (reducing raw vector memory from 614 GB to 153 GB) while preserving 99.2% search recall accuracy.
Pro Tip: Scalar quantization compresses vector memory by 75% with negligible recall degradation, allowing millions of embeddings to fit on affordable commodity nodes.
2️⃣

Architect Distributed Vector Sharding & Replication Topology

Scale horizontally across a multi-node cluster:

  • Distributed Cluster: Deployed a 6-node Qdrant / Milvus cluster distributed across 3 Availability Zones with Raft consensus.
  • Custom Sharding by Tenant: Sharded collections using custom sharding keys (tenant_id), ensuring all embeddings for an individual enterprise customer reside on the same shard for fast localized filtered search.
Pro Tip: Tenant-aware sharding prevents cross-node scatter-gather network queries, keeping filtered query latency sub-10ms.
3️⃣

Implement Pre-Filtering with Dedicated Metadata Payload Indexes

Prevent vector search performance collapse when applying complex metadata filters:

  • The Filtering Problem: Naive post-filtering (searching vectors first, then filtering metadata) fails when filters match only a small percentage of documents.
  • Payload Indexes: Built inverted B-tree indexes on metadata attributes (tenant_id, doc_category, access_group) directly inside the vector engine.
  • Iterative Single-Stage Filtering: The engine traverses the HNSW graph while dynamically evaluating payload index conditions at each hop.
Pro Tip: Single-stage filtered vector search guarantees that strict RBAC security permissions are enforced without destroying vector search recall.
4️⃣

Tier Storage via Memory-Mapped (mmap) Disk-Backed Vectors

Store large vector sets on high-speed NVMe SSDs without keeping everything in RAM:

  • mmap Vector Storage: Kept the HNSW navigational graph in RAM while offloading full raw vector payloads to disk using memory-mapped files (mmap).
  • Performance Telemetry: Tested under 5,000 queries/sec load: sustained p95 latency was 8.2ms, p99 latency was 14.1ms, with 99.4% search recall.
  • Cost Impact: Cut cluster memory footprint by 65%, reducing cloud infrastructure spend from $28,000/month to $8,200/month.
Pro Tip: Disk-backed mmap vectors leverage the Linux page cache to keep frequently accessed vectors warm while reading colder vectors from high-speed NVMe.
💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"Scaling vector search to 100 million embeddings requires HNSW graphs with 8-bit scalar quantization for 75% memory compression, tenant-aware sharding, single-stage payload filtering, and mmap disk-backed vector storage."
⚡ 60-Second Elevator Pitch Talking Points
  • Use HNSW indexing with 8-bit scalar quantization to reduce RAM consumption by 4x.
  • Shard collections by tenant_id across a distributed Raft cluster to eliminate scatter-gather queries.
  • Build payload indexes on metadata attributes to ensure fast, secure RBAC pre-filtering.
  • Store raw vectors on disk using mmap to slash cluster infrastructure costs by 65%.
Advertisement
Want more FinOps & System Design scenarios?
Explore our complete collection of scenario-based FinOps & System Design interview runbooks.
Browse All FinOps & System Design Questions →