⚡ ~/naveed Interview Prep
⚡ Portfolio Home ✍️ Engineering Blog Deep Dives 🎯 Interview Hub 1,000+ Scenarios ☸️ Kubernetes Mastery Hub 24 Modules 🎮 DevOps Arcade & Quizzes Subnet Blitz ⚡ 🗺️ DevOps Roadmaps PDFs & Guides 🤖 Morpheus Analysis AI Quant ↗ 🛠️ Developer Tools Utilities 🧪 Labs & Experiments 📄 Interactive CV & Certs 🔗 All Links & Socials ⚡ Join The Dispatch (Weekly SRE Newsletter) →
← Back to All FinOps & System Design Interview Questions Scenario 53 of 98 in FinOps & System Design
Staff SRE / Distributed Systems Architect System Design High-Scale Traffic & Edge SRE System Design

Q: Your API platform receives 100,000 requests per second across 5 global regions. Malicious bots and rogue customer scripts frequently cause cascading outages by exhausting backend database connection pools. How do you design and implement a distributed, low-latency rate limiter that enforces per-user, per-IP, and per-endpoint limits with sub-millisecond evaluation latency?

Engineering a high-performance, fault-tolerant distributed rate limiting tier capable of evaluating 100,000 requests/second with sub-millisecond overhead using Envoy Proxy, Redis cluster, and Lua token-bucket algorithms.

#System Design #Rate Limiter #Envoy #Redis #Token Bucket #Traffic Management #DDoS
🎙️ Candidate Opening & Architectural Context
"In-memory rate limiting within application pods fails because local state is not shared across autoscaled replicas, allowing attackers to bypass limits. We architected a global distributed rate limiting system using Envoy Proxy, Redis Cluster, and Lua script atomic Token Bucket evaluation."
Advertisement
⚡ Recommended Practice Lab

Want to master this scenario in a live sandbox? The Linux Foundation's FinOps Certified Practitioner (FOCP) Program covers this exact problem with hands-on terminal drills.

🛠️ Production Runbook & Step-by-Step Resolution

1️⃣

Select Rate Limiting Algorithm: Token Bucket vs Sliding Window Counter

Evaluate algorithmic trade-offs for burst tolerance and memory efficiency:

  • Algorithm Decision: Chose the Sliding Window Counter algorithm with Token Bucket burst smoothing.
  • Memory Footprint: Sliding Window Counter requires only two integer counters per key (current window and previous window), consuming only 64 bytes per user key compared to unbounded sliding log memory consumption.
Pro Tip: Sliding window counters prevent boundary-burst attacks where an attacker fires 2x the rate limit across the 59th second and 1st second of adjacent minute windows.
2️⃣

Deploy Envoy External Rate Limit Service (RLS) Filter

Intercept requests at the ingress gateway layer before reaching application backends:

  • Envoy Filter: Configured Envoy envoy.filters.http.ratelimit filter calling an external gRPC Rate Limit Service (RLS).
  • Rate Limit Descriptors: Extracted hierarchical tuples: [('tenant_id', 'org_123'), ('endpoint', '/checkout'), ('client_ip', '198.51.100.4')].
  • Failure Mode: Configured failure_mode_deny: false (fail-open) to ensure that if the rate limit infrastructure ever completely crashes, customer traffic is never blocked.
Pro Tip: Evaluating rate limits at the Envoy gateway offloads compute entirely from application backends, intercepting DDoS attacks at the outer perimeter.
3️⃣

Execute Atomic Evaluation with Redis Cluster & Custom Lua Scripts

Guarantee sub-millisecond atomic decrement operations without distributed locks:

  • Atomic Lua Script: Executed single atomic Lua script in Redis that checks existing counter, increments value, sets TTL, and returns remaining tokens in a single round-trip (RTT < 0.8ms).
  • Redis Cluster Topology: Sharded keys across 6 Redis master nodes with multi-AZ read replicas, utilizing key hashing tags {tenant_id}:endpoint to ensure related keys colocate on the same shard.
Pro Tip: Executing evaluations inside a Redis Lua script is atomic and non-blocking, eliminating race conditions between concurrent requests.
4️⃣

Optimize High-Volume Keys via Local In-Memory Token Caching

Prevent Redis network interface saturation during massive 100k+ RPS spikes:

  • Local Memory Batching: Envoy instances maintain small local in-memory token buffers (e.g., claiming 50 tokens at once from Redis), reducing Redis network queries by 90%.
  • HTTP Response Headers: Returned standard RFC 6585 rate limiting headers: X-RateLimit-Limit, X-RateLimit-Remaining, and Retry-After on HTTP 429 Too Many Requests.
Pro Tip: Returning explicit Retry-After headers instructs well-behaved API clients to back off exponentially, naturally quenching traffic storms.
💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"A high-performance 100k RPS rate limiter leverages Envoy gateway filters, sliding window counters, atomic Redis Lua evaluation, and fail-open resilience to protect backend databases with sub-millisecond latency."
⚡ 60-Second Elevator Pitch Talking Points
  • Use the Sliding Window Counter algorithm to block burst attacks with minimal memory.
  • Deploy Envoy External Rate Limit Service (RLS) over high-speed gRPC at the edge.
  • Execute atomic evaluation using Redis Lua scripts across a multi-AZ sharded cluster.
  • Implement local token batching and fail-open defaults to guarantee high availability.
Advertisement
Want more FinOps & System Design scenarios?
Explore our complete collection of scenario-based FinOps & System Design interview runbooks.
Browse All FinOps & System Design Questions →