Q: Your API platform receives 100,000 requests per second across 5 global regions. Malicious bots and rogue customer scripts frequently cause cascading outages by exhausting backend database connection pools. How do you design and implement a distributed, low-latency rate limiter that enforces per-user, per-IP, and per-endpoint limits with sub-millisecond evaluation latency?
Engineering a high-performance, fault-tolerant distributed rate limiting tier capable of evaluating 100,000 requests/second with sub-millisecond overhead using Envoy Proxy, Redis cluster, and Lua token-bucket algorithms.
Want to master this scenario in a live sandbox? The Linux Foundation's FinOps Certified Practitioner (FOCP) Program covers this exact problem with hands-on terminal drills.
🛠️ Production Runbook & Step-by-Step Resolution
Select Rate Limiting Algorithm: Token Bucket vs Sliding Window Counter
Evaluate algorithmic trade-offs for burst tolerance and memory efficiency:
- Algorithm Decision: Chose the Sliding Window Counter algorithm with Token Bucket burst smoothing.
- Memory Footprint: Sliding Window Counter requires only two integer counters per key (current window and previous window), consuming only 64 bytes per user key compared to unbounded sliding log memory consumption.
Deploy Envoy External Rate Limit Service (RLS) Filter
Intercept requests at the ingress gateway layer before reaching application backends:
- Envoy Filter: Configured Envoy
envoy.filters.http.ratelimitfilter calling an external gRPC Rate Limit Service (RLS). - Rate Limit Descriptors: Extracted hierarchical tuples:
[('tenant_id', 'org_123'), ('endpoint', '/checkout'), ('client_ip', '198.51.100.4')]. - Failure Mode: Configured
failure_mode_deny: false(fail-open) to ensure that if the rate limit infrastructure ever completely crashes, customer traffic is never blocked.
Execute Atomic Evaluation with Redis Cluster & Custom Lua Scripts
Guarantee sub-millisecond atomic decrement operations without distributed locks:
- Atomic Lua Script: Executed single atomic Lua script in Redis that checks existing counter, increments value, sets TTL, and returns remaining tokens in a single round-trip (RTT < 0.8ms).
- Redis Cluster Topology: Sharded keys across 6 Redis master nodes with multi-AZ read replicas, utilizing key hashing tags
{tenant_id}:endpointto ensure related keys colocate on the same shard.
Optimize High-Volume Keys via Local In-Memory Token Caching
Prevent Redis network interface saturation during massive 100k+ RPS spikes:
- Local Memory Batching: Envoy instances maintain small local in-memory token buffers (e.g., claiming 50 tokens at once from Redis), reducing Redis network queries by 90%.
- HTTP Response Headers: Returned standard RFC 6585 rate limiting headers:
X-RateLimit-Limit,X-RateLimit-Remaining, andRetry-Afteron HTTP 429 Too Many Requests.
- Use the Sliding Window Counter algorithm to block burst attacks with minimal memory.
- Deploy Envoy External Rate Limit Service (RLS) over high-speed gRPC at the edge.
- Execute atomic evaluation using Redis Lua scripts across a multi-AZ sharded cluster.
- Implement local token batching and fail-open defaults to guarantee high availability.