⚡ ~/naveed Interview Prep
⚡ Portfolio Home ✍️ Engineering Blog Deep Dives 🎯 Interview Hub 998+ Scenarios ☸️ Kubernetes Mastery Hub 24 Modules 🎮 DevOps Arcade & Quizzes Subnet Blitz ⚡ 🗺️ DevOps Roadmaps PDFs & Guides 🤖 Morpheus Analysis AI Quant ↗ 🛠️ Developer Tools Utilities 🧪 Labs & Experiments 📄 Interactive CV & Certs 🔗 All Links & Socials ⚡ Join The Dispatch (Weekly SRE Newsletter) →
Staff SRE / Principal Architect [L3] Linux Linux / SRE — Scenario-Based Interview Questions Staff SRE Scenario [L3]

Q: Your application maintains persistent TCP connections to a backend service. The backend silently crashes and restarts behind a load balancer, but the application continues to hold stale connections and times out. How do you configure TCP keepalive at the kernel level to detect dead connections faster?

TCP keepalive is a kernel-level mechanism that sends periodic probe packets on idle connections to verify the remote end is still alive.

#Linux #Linux / SRE — Scenario-Based Interview Questions #L3 #SRE #Systems #Troubleshooting
🎙️ Candidate Opening & Architectural Context
""During an on-call shift, our alerts triggered when a critical Linux production server exhibited this behavior. The interviewer is testing: TCP keepalive tuning, sysctl networking, connection health.. I structure my answer around systematic triage first, root cause analysis second, and permanent remediation third.""
Advertisement

🛠️ Production Runbook & Step-by-Step Resolution

1️⃣

Initial Diagnostics & Root Cause Analysis

TCP keepalive is a kernel-level mechanism that sends periodic probe packets on idle connections to verify the remote end is still alive.

  • net.ipv4.tcp_keepalive_time = 7200 (Wait 2 hours before sending the first probe)
  • net.ipv4.tcp_keepalive_intvl = 75 (75 seconds between probes)
  • net.ipv4.tcp_keepalive_probes = 9 (Send 9 probes before declaring dead)
2️⃣

Remediation & Permanent Safeguards

The default Linux settings are extremely conservative: With defaults, it takes 2 hours + 9×75s = ~2h11m to detect a dead connection. For production, I would aggressively tune: Now dead connections are detected in 60 + 6×10 = 120 seconds. These settings are persisted in /etc/sysctl.conf. Note: the application must also enable SO_KEEPALIVE on its sockets for these kernel settings to take effect.

sysctl -w net.ipv4.tcp_keepalive_time=60
sysctl -w net.ipv4.tcp_keepalive_intvl=10
sysctl -w net.ipv4.tcp_keepalive_probes=6
💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"Pro-Tip: net.ipv4.tcp_keepalive_time = 7200 (Wait 2 hours before sending the first probe)."
⚡ 60-Second Elevator Pitch Talking Points
  • net.ipv4.tcp_keepalive_time = 7200 (Wait 2 hours before sending the first probe)
  • net.ipv4.tcp_keepalive_intvl = 75 (75 seconds between probes)
  • net.ipv4.tcp_keepalive_probes = 9 (Send 9 probes before declaring dead)
Advertisement
Want more Linux scenarios?
Explore our complete collection of scenario-based Linux interview runbooks.
Browse All Linux Questions →

📚 Related Production Scenarios in Linux