⚡ ~/naveed Interview Prep
⚡ Portfolio Home ✍️ Engineering Blog Deep Dives 🎯 Interview Hub 1,000+ Scenarios ☸️ Kubernetes Mastery Hub 24 Modules 🎮 DevOps Arcade & Quizzes Subnet Blitz ⚡ 🗺️ DevOps Roadmaps PDFs & Guides 🤖 Morpheus Analysis AI Quant ↗ 🛠️ Developer Tools Utilities 🧪 Labs & Experiments 📄 Interactive CV & Certs 🔗 All Links & Socials ⚡ Join The Dispatch (Weekly SRE Newsletter) →
← Back to All Databases & Storage Interview Questions Scenario 49 of 52 in Databases & Storage
Staff SRE / Distributed Systems Database Distributed Streaming & Kafka Distributed Systems

Q: A 30-pod consumer group reading from an order processing Kafka topic experiences a catastrophic 'rebalance storm': pods continuously drop out and rejoin the group every 5 minutes, stopping all message consumption and ballooning consumer lag to millions of events. A slow external payment API is causing message processing batches to exceed the consumer poll timeout. How do you stabilize consumer heartbeats, migrate to cooperative sticky assignors, and eliminate rebalance loops?

Remediate catastrophic Kafka consumer group rebalance loops caused by processing timeouts on slow external APIs, and adopt the Cooperative Sticky Assignor to maintain streaming availability.

#Kafka #Consumer Group #Rebalance Storm #Distributed Streaming #Lag #Cooperative Sticky Assignor
🎙️ Candidate Opening & Architectural Context
"In Apache Kafka, consumer groups dynamically distribute partition ownership across members. A consumer must periodically call `poll()` within `max.poll.interval.ms`. If a consumer thread is blocked (e.g. waiting for a downstream database or third-party HTTP call) and misses the interval, the Kafka coordinator assumes the consumer has died, evicts it, and triggers a group-wide rebalance. Under the legacy 'Eager' assignment protocol, all consumers stop processing and revoke all partitions during every rebalance, creating a vicious rebalance storm."
Advertisement
⚡ Recommended Practice Lab

Want to master this scenario in a live sandbox? KodeKloud's PostgreSQL Database Administration & High Availability Course covers this exact problem with hands-on terminal drills.

🛠️ Production Runbook & Step-by-Step Resolution

⚡

Situation: Rebalance Storm Freezes Event Processing Pipeline

⚡

Task: Decouple Heartbeats from Processing & Eliminate Stop-The-World Rebalances

Advertisement
⚡

Action: Cooperative Sticky Rebalancing & Batch Rightsizing

⚡

Result: 100% Elimination of Rebalance Storms & Recovery of 2M Lag

💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"Rebalance storms happen when message processing times exceed `max.poll.interval.ms`. Always rightsize `max.poll.records`, use the modern `CooperativeStickyAssignor` to avoid stop-the-world pauses, and decouple Kafka poll loops from slow external HTTP calls."
⚡ 60-Second Elevator Pitch Talking Points
  • Kafka evicts consumers when processing batches exceed max.poll.interval.ms.
  • Eager rebalancing revokes all partitions globally, halting all stream consumption.
  • Use CooperativeStickyAssignor and lower max.poll.records to allow incremental rebalancing without stopping consumers.
Advertisement
Want more Databases & Storage scenarios?
Explore our complete collection of scenario-based Databases & Storage interview runbooks.
Browse All Databases & Storage Questions →