⚡ ~/naveed Interview Prep
⚡ Portfolio Home ✍️ Engineering Blog Deep Dives 🎯 Interview Hub 1,000+ Scenarios ☸️ Kubernetes Mastery Hub 24 Modules 🎮 DevOps Arcade & Quizzes Subnet Blitz ⚡ 🗺️ DevOps Roadmaps PDFs & Guides 🤖 Morpheus Analysis AI Quant ↗ 🛠️ Developer Tools Utilities 🧪 Labs & Experiments 📄 Interactive CV & Certs 🔗 All Links & Socials ⚡ Join The Dispatch (Weekly SRE Newsletter) →
← Back to All FinOps & System Design Interview Questions Scenario 78 of 98 in FinOps & System Design
Staff SRE / Distributed Systems Architect System Design Real-Time Systems & High Concurrency System Design

Q: Your mobile trading and chat application must maintain 5,000,000 simultaneous persistent WebSocket connections to deliver instant price updates and notifications with sub-100ms latency. Edge connections frequently terminate due to mobile network handoffs, triggering reconnect storms. How do you design the networking, gateway tier, pub/sub broker, and Linux kernel to handle 5 million connections reliably?

Engineering a real-time push notification service sustaining 5,000,000 concurrent persistent WebSocket connections with Redis Pub/Sub, horizontal connection gateway pods, and Linux kernel epoll tuning.

#System Design #WebSockets #Push Notifications #Redis #Kafka #Epoll #High Concurrency
🎙️ Candidate Opening & Architectural Context
"Maintaining 5 million concurrent stateful WebSocket connections exhausts Linux file descriptors, memory buffers, and breaks standard stateless load balancers. We engineered a massive real-time WebSocket platform utilizing an epoll-based Go gateway tier, Redis Pub/Sub sharding, and Linux socket tuning."
Advertisement
⚡ Recommended Practice Lab

Want to master this scenario in a live sandbox? The Linux Foundation's FinOps Certified Practitioner (FOCP) Program covers this exact problem with hands-on terminal drills.

🛠️ Production Runbook & Step-by-Step Resolution

1️⃣

Tune Linux Kernel & OS Limits for Millions of Sockets

Prevent TCP socket memory starvation and file descriptor limits on worker nodes:

  • File Descriptors: Increased system limits: fs.file-max = 10000000 and process limits: ulimit -n 1048576.
  • TCP Memory Tuning: Tuned TCP buffer windows: net.ipv4.tcp_rmem = 4096 87380 16777216 and net.ipv4.tcp_wmem = 4096 65536 16777216, setting minimum read/write buffers to 4KB to pack 100,000 connections per 16 GB RAM node.
  • Ephemeral Ports: Configured net.ipv4.ip_local_port_range = 1024 65535 and tuned net.core.somaxconn = 65535 to absorb massive connection surges.
Pro Tip: Tuning minimum TCP buffer sizes from 64KB down to 4KB reduces memory consumption per idle connection by over 90%, allowing massive connection density.
2️⃣

Deploy Horizontal Gateway Cluster using epoll / kqueue (Go / Rust)

Manage millions of concurrent connections with minimal CPU context switching:

  • Non-Blocking I/O: Built gateway servers in Go using gobwas/ws or Rust using tokio-tungstenite with native Linux epoll event loops.
  • Node Capacity: Deployed a cluster of 50 gateway pods (each holding 100,000 connections), distributed behind a Layer 4 Network Load Balancer (AWS NLB) with Maglev consistent hashing.
Pro Tip: Using non-blocking epoll event loops allows a single thread to manage thousands of idle connections without allocating dedicated OS threads per socket.
3️⃣

Route Targeted Messages via Sharded Redis Pub/Sub Mesh

Locate which gateway pod holds the target user's active socket connection:

  • Connection Registry: When user user_4821 connects, gateway registers mapping in Redis: SET connection:user_4821 'gateway-pod-14' EX 300 (refreshed via heartbeat).
  • Sharded Redis Pub/Sub: When a notification arrives, publisher checks the registry and publishes the message to the specific pod's channel: gateway-pod-14.
  • Direct Emission: Gateway pod 14 receives the event from Redis and writes the JSON packet directly to the user's active WebSocket in < 4 milliseconds.
Pro Tip: Sharded pub/sub routing avoids broadcasting notifications to all 50 gateway pods, saving immense CPU and network bandwidth.
4️⃣

Mitigate Reconnect Storms via Exponential Backoff & Jitter

Prevent thundering herd collapse when cell towers or wifi networks drop:

  • Client Reconnection Jitter: Mobile client libraries enforce randomized exponential backoff: sleep = min(cap, base * 2^attempt) + rand(0, 1000ms).
  • Heartbeat Ping/Pong: Gateway sends WebSocket PING every 30 seconds; if client fails to respond with PONG within 10 seconds, the dead socket descriptor is cleaned up immediately.
  • Validation Load Test: Sustained 5.2 million live concurrent connections under load with p99 delivery latency of 18ms.
Pro Tip: Randomized jitter prevents millions of disconnected mobile devices from synchronizing retry requests and crushing gateway ingress.
💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"Sustaining 5 million concurrent WebSockets requires Linux kernel socket buffer minimization, horizontal epoll gateway pods, sharded Redis connection routing, and client reconnection jitter to prevent thundering herds."
⚡ 60-Second Elevator Pitch Talking Points
  • Tune Linux TCP socket memory buffers down to 4KB to pack 100k connections per node.
  • Deploy non-blocking epoll gateway pods behind L4 Network Load Balancers.
  • Route targeted notifications through a sharded Redis Pub/Sub cluster in < 5ms.
  • Enforce randomized client exponential backoff and jitter to survive massive reconnect storms.
Advertisement
Want more FinOps & System Design scenarios?
Explore our complete collection of scenario-based FinOps & System Design interview runbooks.
Browse All FinOps & System Design Questions →