Q: Your microservices make hundreds of feature flag checks per incoming user request. Querying a central feature flag database over the network would add 15ms latency per request and create a single point of failure that could take down the entire company if the flag service crashes. How do you design a feature flagging platform that evaluates flags in < 0.1ms with zero network calls, while propagating flag updates globally in under 2 seconds?
Architectural design for a distributed, sub-millisecond feature flagging platform evaluating 100,000 evaluations/second with local in-memory evaluation SDKs, SSE rule streaming, and emergency kill-switches.
Want to master this scenario in a live sandbox? The Linux Foundation's FinOps Certified Practitioner (FOCP) Program covers this exact problem with hands-on terminal drills.
🛠️ Production Runbook & Step-by-Step Resolution
Implement Local In-Memory Evaluation Architecture in Client SDKs
Eliminate network calls during runtime flag evaluation:
- Zero Network Calls: Client SDKs embedded in Go, Java, and Node.js microservices maintain an in-memory hash map of all active flag rules.
- Sub-Microsecond Latency: Evaluating
isFeatureEnabled('new-checkout', userContext)executes purely in memory in < 0.05 milliseconds (50 microseconds) with ZERO HTTP requests. - Air-Gapped Resilience: If the central feature flag control plane completely crashes, microservices continue evaluating flags normally from local memory with zero disruption.
Stream Rule Updates Globally via Server-Sent Events (SSE) & Redis Pub/Sub
Propagate flag rule changes to thousands of microservice pods in real time:
- Control Plane: Admin UI allows engineers to toggle flags, saving rules in PostgreSQL.
- Edge Streaming Proxy: Relay edge proxies maintain persistent HTTP Server-Sent Events (SSE) connections to all microservice pods.
- Redis Pub/Sub Sync: When an engineer flips a flag in the UI, an event publishes to Redis Pub/Sub; relay proxies broadcast the update over SSE, updating the in-memory cache of 2,000 pods globally in < 1.4 seconds.
Calculate Percentage Rollouts via Deterministic MurmurHash3 Hashing
Ensure stable, consistent user bucket assignment across distributed services:
- Deterministic Hashing: SDK computes hash:
hash = MurmurHash3(flagKey + ':' + userId) % 100. - Consistent User Experience: If rollout is set to 25%, any user with hash < 25 sees the feature. A user consistently receives the exact same evaluation across all 50 microservices without central state synchronization.
Deploy Automated Emergency Kill-Switches & Flag Evaluation Telemetry
Instantly disable broken features and prune stale flags:
- Automated Kill-Switch: If Datadog / Prometheus alerts detect error rate spike > 1% correlated with a flag release, an automated webhook flips the kill-switch, disabling the flag globally in 1.2 seconds.
- Evaluation Telemetry: SDKs batch evaluation counts in memory and flush metrics every 30 seconds, automatically flagging dead/stale flags that haven't been modified in 90 days.
- Evaluate flags locally in client SDK memory with zero network calls and sub-50 microsecond latency.
- Stream rule modifications to thousands of pods globally in < 2 seconds using Server-Sent Events (SSE).
- Use deterministic MurmurHash3 hashing for sticky, stateless percentage rollouts.
- Deploy automated telemetry-triggered kill-switches to instantly disable broken code in production.