Staff SRE / Distributed SystemsDatabaseDistributed Streaming & KafkaDistributed Systems
Q: A 30-pod consumer group reading from an order processing Kafka topic experiences a catastrophic 'rebalance storm': pods continuously drop out and rejoin the group every 5 minutes, stopping all message consumption and ballooning consumer lag to millions of events. A slow external payment API is causing message processing batches to exceed the consumer poll timeout. How do you stabilize consumer heartbeats, migrate to cooperative sticky assignors, and eliminate rebalance loops?
Remediate catastrophic Kafka consumer group rebalance loops caused by processing timeouts on slow external APIs, and adopt the Cooperative Sticky Assignor to maintain streaming availability.
"In Apache Kafka, consumer groups dynamically distribute partition ownership across members. A consumer must periodically call `poll()` within `max.poll.interval.ms`. If a consumer thread is blocked (e.g. waiting for a downstream database or third-party HTTP call) and misses the interval, the Kafka coordinator assumes the consumer has died, evicts it, and triggers a group-wide rebalance. Under the legacy 'Eager' assignment protocol, all consumers stop processing and revoke all partitions during every rebalance, creating a vicious rebalance storm."
Result: 100% Elimination of Rebalance Storms & Recovery of 2M Lag
💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"Rebalance storms happen when message processing times exceed `max.poll.interval.ms`. Always rightsize `max.poll.records`, use the modern `CooperativeStickyAssignor` to avoid stop-the-world pauses, and decouple Kafka poll loops from slow external HTTP calls."
⚡ 60-Second Elevator Pitch Talking Points
Kafka evicts consumers when processing batches exceed max.poll.interval.ms.
Eager rebalancing revokes all partitions globally, halting all stream consumption.
Use CooperativeStickyAssignor and lower max.poll.records to allow incremental rebalancing without stopping consumers.
Advertisement
⚡ Free SRE Study Guide
Get 1 DevOps interview question in your inbox every week
Join 14,000+ engineers leveling up their cloud and platform interview game. Subscribe to get our weekly deep-dive scenario plus instant access to the Top 50 Kubernetes Interview Questions & Incident Runbooks PDF.