⚡ ~/naveed Interview Prep
⚡ Portfolio Home ✍️ Engineering Blog Deep Dives 🎯 Interview Hub 1,000+ Scenarios ☸️ Kubernetes Mastery Hub 24 Modules 🎮 DevOps Arcade & Quizzes Subnet Blitz ⚡ 🗺️ DevOps Roadmaps PDFs & Guides 🤖 Morpheus Analysis AI Quant ↗ 🛠️ Developer Tools Utilities 🧪 Labs & Experiments 📄 Interactive CV & Certs 🔗 All Links & Socials ⚡ Join The Dispatch (Weekly SRE Newsletter) →
← Back to All Observability & Monitoring Interview Questions Scenario 95 of 96 in Observability & Monitoring
Senior SRE / Distributed Systems Engineer Observability Distributed Systems & Telemetry Triage Tesla Scale Loop

Q: A service depends on Kafka and Redis. Both clusters appear 100% healthy on monitoring dashboards, but service requests are intermittently failing. What is your debug path before escalating to application teams?

Forensic diagnostic path to isolate intermittent request failures in a high-throughput microservice dependent on healthy Kafka and Redis clusters.

#Observability #Kafka #Redis #Connection Pools #TCP Sockets #Tesla #Troubleshooting
🎙️ Candidate Opening & Architectural Context
"When upstream Kafka and Redis clusters report healthy CPU, memory, and cluster health while client requests intermittently drop, the problem almost always resides in the client-side transport layer: connection pool exhaustion, file descriptor limits, socket TIME_WAIT accumulation, or thread starvation on client event loops."
Advertisement
⚡ Recommended Practice Lab

Want to master this scenario in a live sandbox? The Linux Foundation's Prometheus Certified Associate (PCA) & Monitoring Labs covers this exact problem with hands-on terminal drills.

🛠️ Production Runbook & Step-by-Step Resolution

1

Check Socket States & Ephemeral Port / File Descriptor Limits

Log into the application pods or host nodes. Inspect socket connection allocation with `ss -s` and `lsof`. Look for thousands of sockets in `TIME_WAIT` or `CLOSE_WAIT` states. Check if the container process is approaching its max open files limit (`ulimit -n`).

# Inspect active socket connection states and file descriptors
ss -s
cat /proc/<pid>/limits | grep "Max open files"
ls -1 /proc/<pid>/fd | wc -l
2

Inspect Client-Side Connection Pool Starvation & Timeouts

Verify client configuration (e.g. Jedis/Lettuce for Redis, librdkafka/confluent-kafka for Kafka). If the client connection pool has `maxTotal = 50` connections and 500 concurrent worker threads burst, threads block waiting for an idle connection until the client timeout (e.g. 1000ms) triggers, throwing intermittent timeout exceptions.

Pro Tip: Scale Reality: Redis and Kafka cluster metrics only show server-side health. A single saturated client pool will drop requests while the Redis server operates at 5% CPU.
Advertisement
3

Analyze Redis Single-Thread Latency & Slowlog

Run `redis-cli slowlog get 20`. Even if overall Redis CPU is low, Redis runs on a single core for command execution. A single O(N) command (like `KEYS *` or an unbounded `HGETALL`) blocks the event loop for 200ms, dropping incoming client TCP requests intermittently.

redis-cli -h redis.internal.tesla.com slowlog get 25
4

Examine Kafka Producer Buffer Pools & Socket Send Buffers

Check if Kafka producers are hitting `buffer.memory` limits. When downstream network latency fluctuates, `max.block.ms` expires, causing the producer to drop messages before they reach Kafka brokers.

💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"Healthy cluster dashboards don't reflect client-side health. Check file descriptor limits, socket TIME_WAIT exhaustion, client connection pool starvation, and Redis slowlog spikes."
⚡ 60-Second Elevator Pitch Talking Points
  • Inspect socket states (ss -s) and file descriptors (ulimit -n) on client service pods.
  • Verify client connection pool capacity (Jedis/Kafka producers) to detect pool exhaustion under burst traffic.
  • Check redis slowlog for blocking O(N) commands choking the single-threaded event loop.
  • Monitor Kafka client buffer.memory saturation and max.block.ms producer timeouts.
Advertisement
Want more Observability & Monitoring scenarios?
Explore our complete collection of scenario-based Observability & Monitoring interview runbooks.
Browse All Observability & Monitoring Questions →