Q: You need to perform a kernel upgrade on a production server with zero unplanned downtime. The server runs on RHEL/CentOS. Describe the safest approach.
Kernel upgrades inherently require a reboot on standard Linux (unlike kpatch/livepatch for security fixes). The goal is to minimize risk ...
#Linux #group: engineering #L3 #SRE #Systems #Troubleshooting
🎙️ Candidate Opening & Architectural Context
""Never reboot a server blindly; always capture top process telemetry, lsof descriptors, and thread dumps first. The interviewer is testing: Kernel lifecycle management, risk mitigation, rollback strategy.. I structure my answer around systematic triage first, root cause analysis second, and permanent remediation third.""
Advertisement
🛠️ Production Runbook & Step-by-Step Resolution
1️⃣
Initial Diagnostics & Root Cause Analysis
Kernel upgrades inherently require a reboot on standard Linux (unlike kpatch/livepatch for security fixes). The goal is to minimize risk and ensure rollback capability.
- Pre-flight checks:
- Verify current kernel:
uname -r - Check available updates:
yum list kernel - Review changelog for breaking changes:
yum updateinfo info kernel - Ensure GRUB has the old kernel as fallback
- Install (don't update) the new kernel:
- Verify boot configuration:
2️⃣
Remediation & Permanent Safeguards
Safe upgrade workflow: This installs alongside the old kernel rather than replacing it.
yum install kernel-5.14.0-362.el9
- Schedule maintenance window:
- Drain traffic from the server (remove from load balancer).
- Reboot into the new kernel.
- Run validation tests (application health checks, network connectivity, disk mounts).
- Rollback plan (if issues arise):
- Post-validation: Keep the old kernel installed for at least 2 weeks before removing it with
yum remove kernel-.
💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"Pro-Tip: Pre-flight checks:."
⚡ 60-Second Elevator Pitch Talking Points
- Pre-flight checks:
- Verify current kernel: uname -r
- Check available updates: yum list kernel
Advertisement