⚡ ~/naveed Interview Prep
⚡ Portfolio Home ✍️ Engineering Blog Deep Dives 🎯 Interview Hub 1,000+ Scenarios ☸️ Kubernetes Mastery Hub 24 Modules 🎮 DevOps Arcade & Quizzes Subnet Blitz ⚡ 🗺️ DevOps Roadmaps PDFs & Guides 🤖 Morpheus Analysis AI Quant ↗ 🛠️ Developer Tools Utilities 🧪 Labs & Experiments 📄 Interactive CV & Certs 🔗 All Links & Socials ⚡ Join The Dispatch (Weekly SRE Newsletter) →
← Back to All FinOps & System Design Interview Questions Scenario 96 of 98 in FinOps & System Design
Staff Distributed Systems Architect System Design Distributed Caching & High-Performance SRE System Design

Q: Your high-throughput read API serves 300,000 requests per second. Querying central Redis for every request saturates Redis network bandwidth (10 Gbps limit). Storing a local L1 cache inside each Go/Java application pod delivers 0.01ms latency, but pods serve stale data when entities are updated in the database. How do you design a multi-tier cache hierarchy with sub-millisecond distributed cache invalidation across 200 application pods?

Engineering a high-performance two-tier cache hierarchy (L1 in-memory application cache + L2 shared Redis cluster) handling 300,000 RPS with sub-millisecond invalidation broadcasts via Redis Pub/Sub.

#System Design #Caching #Cache Invalidation #Redis #Pub/Sub #Cache-Aside #Two-Tier Cache
🎙️ Candidate Opening & Architectural Context
"Centralized caches hit network bandwidth scaling limits under massive read loads. We architected a two-tier caching platform combining local L1 in-memory pod caching (Go Ristretto / Java Caffeine), shared L2 Redis Cluster, and distributed invalidation broadcasting over Redis Pub/Sub."
Advertisement
⚡ Recommended Practice Lab

Want to master this scenario in a live sandbox? The Linux Foundation's FinOps Certified Practitioner (FOCP) Program covers this exact problem with hands-on terminal drills.

🛠️ Production Runbook & Step-by-Step Resolution

1️⃣

Architect Two-Tier Cache Topology: L1 In-Memory + L2 Redis Cluster

Achieve nanosecond read latency while offloading central Redis network traffic:

  • Tier 1 (L1 In-Memory): Each application pod allocates 500 MB RAM to a local in-memory cache (Caffeine for JVM / Ristretto for Go) using TinyLFU eviction policies.
  • Tier 2 (L2 Shared Cache): Shared multi-AZ Redis Cluster caching complete serialized entities with a 24-hour TTL.
  • Read Flow: Application checks L1 (0.005ms) -> if miss, checks L2 Redis (1.2ms) -> if miss, queries PostgreSQL database (15ms), populating both L2 and L1.
Pro Tip: L1 cache absorbs 88% of all read requests, slashing Redis network bandwidth consumption from 12 Gbps down to 1.4 Gbps.
2️⃣

Broadcast Sub-Millisecond Invalidation via Redis Pub/Sub

Evict stale local pod memory the instant an update occurs:

  • Write Flow: When an update modifies product:1234, the writing pod updates PostgreSQL, deletes the key from L2 Redis, and publishes an invalidation event: PUBLISH cache:invalidate 'product:1234'.
  • Pod Subscription: All 200 application pods maintain an active subscription to cache:invalidate.
  • Local Eviction: Upon receiving the broadcast message, every pod immediately purges product:1234 from its local L1 memory in < 1 millisecond.
Pro Tip: Broadcasting invalidations via Pub/Sub ensures that local L1 pod caches never serve stale data longer than a few milliseconds.
3️⃣

Mitigate Pub/Sub Packet Drops via Monotonic Versioning & Short TTL

Prevent silent data desynchronization if a pod temporarily loses network connection:

  • Short L1 TTL: All L1 cache entries enforce a maximum background TTL of 60 seconds (even if never invalidated), bounding maximum staleness to 60s under catastrophic network splits.
  • Monotonic Entity Versioning: Objects embed an incremental version counter: { 'id': 1234, 'version': 48 }. If an L1 cache entry has a version lower than an incoming write request header, the local entry is evicted immediately.
Pro Tip: Redis Pub/Sub does not guarantee message delivery to disconnected subscribers; bounding L1 TTLs to 60 seconds guarantees eventual convergence.
4️⃣

Prevent Cache Stampedes via SingleFlight Mutex Suppression

Shield the underlying database during sudden cache eviction spikes:

  • SingleFlight Pattern: When an L1/L2 cache miss occurs for a popular key, the pod uses Go singleflight.Group (or Java synchronized loader) to ensure that only ONE worker thread queries the database.
  • Shared Result: All 500 concurrent in-flight requests wait for the single database query to complete and share the result in memory.
  • Database Impact: Database CPU utilization during cache invalidations dropped from 95% spikes down to a steady 12%.
Pro Tip: SingleFlight mutex suppression prevents the classic 'thundering herd' cache stampede from crashing the database when a popular cache key is invalidated.
💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"A two-tier cache hierarchy combines local pod L1 in-memory caches to absorb 88% of reads, L2 Redis Cluster for shared persistence, Redis Pub/Sub for sub-millisecond invalidation broadcasts, and SingleFlight to prevent database stampedes."
⚡ 60-Second Elevator Pitch Talking Points
  • Deploy L1 in-memory pod caching (TinyLFU) backed by an L2 Redis Cluster to slash network bandwidth.
  • Broadcast real-time cache invalidation messages across all pods in < 1ms via Redis Pub/Sub.
  • Enforce short 60s L1 safety TTLs and monotonic entity versioning to handle network blips.
  • Use SingleFlight request coalescing to eliminate thundering herd database stampedes.
Advertisement
Want more FinOps & System Design scenarios?
Explore our complete collection of scenario-based FinOps & System Design interview runbooks.
Browse All FinOps & System Design Questions →