Q: Your core financial platform relies on Apache Kafka for event-driven transactions processing 500,000 messages/sec. Disk storage on high-speed NVMe brokers fills up in 3 days, causing massive rebalance storms when expanding cluster capacity. A single AZ datacenter failure must cause zero message loss and seamless consumer failover. How do you design Apache Kafka with Tiered Storage and Multi-AZ rack awareness?
Engineering a high-throughput, multi-AZ Apache Kafka cluster supporting 500,000 msg/sec with KIP-405 Tiered Storage, zero message loss (acks=all), and rack-aware replica distribution.
Want to master this scenario in a live sandbox? The Linux Foundation's FinOps Certified Practitioner (FOCP) Program covers this exact problem with hands-on terminal drills.
🛠️ Production Runbook & Step-by-Step Resolution
Deploy KRaft-Based Cluster with Multi-AZ Rack-Aware Topology
Distribute brokers and partition replicas across independent physical failure zones:
- KRaft Consensus: Eliminated ZooKeeper; deployed 5 dedicated KRaft controller nodes across 3 Availability Zones.
- Broker Rack Awareness: Assigned
broker.rack=us-east-1a,1b,1cin broker configurations. - Topic Configuration: Enforced replication factor
rf=3and minimum in-sync replicasmin.insync.replicas=2, guaranteeing that replicas are physically distributed across 3 independent availability zones.
Configure Kafka Tiered Storage (Remote Storage Manager to S3)
Decouple local broker disk capacity from long-term topic retention:
- Tiered Storage Activation: Enabled KIP-405 remote storage manager backed by AWS S3 / Google Cloud Storage.
- Local vs Remote Retention: Set
local.retention.ms=4h(hot tier on NVMe SSD) andretention.ms=30d(cold tier in S3). - Rebalance Elimination: Because local broker disks only store 4 hours of hot data, adding a new broker or rebalancing partitions takes minutes instead of 18 hours.
Enforce Zero-Data-Loss Producer & Consumer Guarantees
Align client configurations to ensure strict at-least-once or exactly-once delivery:
- Producer Settings: Enforced
acks=all,enable.idempotence=true, andretries=MAX_INTwith exponential backoff. - Consumer Fetch from Follower (KIP-392): Configured consumers to fetch directly from the closest in-rack follower replica (
client.rack=us-east-1a), eliminating expensive cross-AZ network egress bandwidth.
Validate Availability During Simulated Full AZ Blackout
Prove resilience during unannounced broker and datacenter termination:
- Chaos Simulation: Terminated all brokers in us-east-1a simultaneously under peak load (500k msg/sec).
- Telemetry Results: Partition leaders smoothly failed over to surviving in-sync replicas in AZs 1b and 1c within 850ms; producer error rate remained 0.00% with idempotent retries.
- Deploy Kafka with KRaft and multi-AZ rack awareness to survive full datacenter outages.
- Use KIP-405 Tiered Storage to offload data older than 4 hours directly to S3 object storage.
- Enforce acks=all and idempotence on producers; enable consumer fetch-from-follower to save cross-AZ egress.
- Validate sub-second automated partition failover during simulated AZ termination drills.