⚡ ~/naveed Interview Prep
⚡ Portfolio Home ✍️ Engineering Blog Deep Dives 🎯 Interview Hub 1,000+ Scenarios ☸️ Kubernetes Mastery Hub 24 Modules 🎮 DevOps Arcade & Quizzes Subnet Blitz ⚡ 🗺️ DevOps Roadmaps PDFs & Guides 🤖 Morpheus Analysis AI Quant ↗ 🛠️ Developer Tools Utilities 🧪 Labs & Experiments 📄 Interactive CV & Certs 🔗 All Links & Socials ⚡ Join The Dispatch (Weekly SRE Newsletter) →
← Back to All FinOps & System Design Interview Questions Scenario 69 of 98 in FinOps & System Design
Staff Distributed Systems Architect System Design Distributed Streaming & Big Data SRE System Design

Q: Your core financial platform relies on Apache Kafka for event-driven transactions processing 500,000 messages/sec. Disk storage on high-speed NVMe brokers fills up in 3 days, causing massive rebalance storms when expanding cluster capacity. A single AZ datacenter failure must cause zero message loss and seamless consumer failover. How do you design Apache Kafka with Tiered Storage and Multi-AZ rack awareness?

Engineering a high-throughput, multi-AZ Apache Kafka cluster supporting 500,000 msg/sec with KIP-405 Tiered Storage, zero message loss (acks=all), and rack-aware replica distribution.

#System Design #Kafka #Tiered Storage #Multi-AZ #Strimzi #Event Streaming
🎙️ Candidate Opening & Architectural Context
"Retaining weeks of event history on broker-attached SSDs is prohibitively expensive and leads to catastrophic multi-hour broker rebalancing operations. We architected a Kafka platform utilizing KIP-405 Tiered Storage (AWS S3), KRaft consensus, and rack-aware partition allocation."
Advertisement
⚡ Recommended Practice Lab

Want to master this scenario in a live sandbox? The Linux Foundation's FinOps Certified Practitioner (FOCP) Program covers this exact problem with hands-on terminal drills.

🛠️ Production Runbook & Step-by-Step Resolution

1️⃣

Deploy KRaft-Based Cluster with Multi-AZ Rack-Aware Topology

Distribute brokers and partition replicas across independent physical failure zones:

  • KRaft Consensus: Eliminated ZooKeeper; deployed 5 dedicated KRaft controller nodes across 3 Availability Zones.
  • Broker Rack Awareness: Assigned broker.rack=us-east-1a, 1b, 1c in broker configurations.
  • Topic Configuration: Enforced replication factor rf=3 and minimum in-sync replicas min.insync.replicas=2, guaranteeing that replicas are physically distributed across 3 independent availability zones.
Pro Tip: Rack awareness ensures that a complete datacenter or AZ power outage leaves at least two in-sync replicas healthy and ready to serve writes.
2️⃣

Configure Kafka Tiered Storage (Remote Storage Manager to S3)

Decouple local broker disk capacity from long-term topic retention:

  • Tiered Storage Activation: Enabled KIP-405 remote storage manager backed by AWS S3 / Google Cloud Storage.
  • Local vs Remote Retention: Set local.retention.ms=4h (hot tier on NVMe SSD) and retention.ms=30d (cold tier in S3).
  • Rebalance Elimination: Because local broker disks only store 4 hours of hot data, adding a new broker or rebalancing partitions takes minutes instead of 18 hours.
Pro Tip: Tiered storage reduces Kafka storage infrastructure costs by 70% while completely eliminating broker disk-full emergency pages.
3️⃣

Enforce Zero-Data-Loss Producer & Consumer Guarantees

Align client configurations to ensure strict at-least-once or exactly-once delivery:

  • Producer Settings: Enforced acks=all, enable.idempotence=true, and retries=MAX_INT with exponential backoff.
  • Consumer Fetch from Follower (KIP-392): Configured consumers to fetch directly from the closest in-rack follower replica (client.rack=us-east-1a), eliminating expensive cross-AZ network egress bandwidth.
Pro Tip: Enabling consumer fetch-from-follower cuts cross-AZ data transfer fees by over 60% across the Kubernetes cluster.
4️⃣

Validate Availability During Simulated Full AZ Blackout

Prove resilience during unannounced broker and datacenter termination:

  • Chaos Simulation: Terminated all brokers in us-east-1a simultaneously under peak load (500k msg/sec).
  • Telemetry Results: Partition leaders smoothly failed over to surviving in-sync replicas in AZs 1b and 1c within 850ms; producer error rate remained 0.00% with idempotent retries.
Pro Tip: Combining acks=all with min.insync.replicas=2 mathematically guarantees that committed messages can never be lost during a single-AZ failure.
💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"Modern enterprise Apache Kafka pairs KRaft and rack-aware replication with KIP-405 Tiered Storage to S3, delivering zero data loss, sub-second leader failover, and decoupled petabyte retention."
⚡ 60-Second Elevator Pitch Talking Points
  • Deploy Kafka with KRaft and multi-AZ rack awareness to survive full datacenter outages.
  • Use KIP-405 Tiered Storage to offload data older than 4 hours directly to S3 object storage.
  • Enforce acks=all and idempotence on producers; enable consumer fetch-from-follower to save cross-AZ egress.
  • Validate sub-second automated partition failover during simulated AZ termination drills.
Advertisement
Want more FinOps & System Design scenarios?
Explore our complete collection of scenario-based FinOps & System Design interview runbooks.
Browse All FinOps & System Design Questions →