⚡ ~/naveed Interview Prep
⚡ Portfolio Home ✍️ Engineering Blog Deep Dives 🎯 Interview Hub 998+ Scenarios ☸️ Kubernetes Mastery Hub 24 Modules 🎮 DevOps Arcade & Quizzes Subnet Blitz ⚡ 🗺️ DevOps Roadmaps PDFs & Guides 🤖 Morpheus Analysis AI Quant ↗ 🛠️ Developer Tools Utilities 🧪 Labs & Experiments 📄 Interactive CV & Certs 🔗 All Links & Socials ⚡ Join The Dispatch (Weekly SRE Newsletter) →
Staff SRE / Principal Architect [L3] Linux Linux / SRE — Scenario-Based Interview Questions Staff SRE Scenario [L3]

Q: An incident occurred where a server ran out of memory, but the `oom-killer` didn't trigger, causing a complete kernel hang (hard lockup). How do you configure the kernel to automatically panic and reboot if this happens again?

If the system becomes completely unresponsive in a hard lockup or OOM stall, manual intervention is slow. I would configure the kernel to...

#Linux #Linux / SRE — Scenario-Based Interview Questions #L3 #SRE #Systems #Troubleshooting
🎙️ Candidate Opening & Architectural Context
""During an on-call shift, our alerts triggered when a critical Linux production server exhibited this behavior. The interviewer is testing: Sysctl tuning, kernel panic parameters, high availability recovery.. I structure my answer around systematic triage first, root cause analysis second, and permanent remediation third.""
Advertisement

🛠️ Production Runbook & Step-by-Step Resolution

1️⃣

Initial Diagnostics & Root Cause Analysis

If the system becomes completely unresponsive in a hard lockup or OOM stall, manual intervention is slow. I would configure the kernel to panic and automatically reboot. This allows auto-scaling groups or load balancers to immediately replace or re-route traffic from the failed node.

  • kernel.panic = 10 (Reboot 10 seconds after a panic)
  • vm.panic_on_oom = 1 (Panic if the kernel hits an unresolvable OOM state, rather than trying to kill processes if the killer is disabled or failing)
  • kernel.hung_task_panic = 1 (Panic if tasks are hung for too long, like the D state lockups).
2️⃣

Remediation & Permanent Safeguards

Using sysctl: These settings are added to /etc/sysctl.conf to persist across reboots, forcing the system to "fail fast" and rely on infrastructure redundancy.

💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"Pro-Tip: kernel.panic = 10 (Reboot 10 seconds after a panic)."
⚡ 60-Second Elevator Pitch Talking Points
  • kernel.panic = 10 (Reboot 10 seconds after a panic)
  • vm.panic_on_oom = 1 (Panic if the kernel hits an unresolvable OOM state, rather than trying to ki...
  • kernel.hung_task_panic = 1 (Panic if tasks are hung for too long, like the D state lockups).
Advertisement
Want more Linux scenarios?
Explore our complete collection of scenario-based Linux interview runbooks.
Browse All Linux Questions →

📚 Related Production Scenarios in Linux