Q: Explain the difference between SIGTERM and SIGKILL. Why must applications handle SIGTERM gracefully before deployment, and what happens during Kubernetes pod termination?
SIGTERM (signal 15): A "polite request" to terminate. The process can catch it, perform cleanup (flush buffers, close connections), and e...
#Linux #Linux / SRE — Scenario-Based Interview Questions #L1 #SRE #Systems #Troubleshooting
🎙️ Candidate Opening & Architectural Context
""During an on-call shift, our alerts triggered when a critical Linux production server exhibited this behavior. The interviewer is testing: Process signals, graceful shutdown, orchestration lifecycle.. I structure my answer around systematic triage first, root cause analysis second, and permanent remediation third.""
Advertisement
🛠️ Production Runbook & Step-by-Step Resolution
1️⃣
Initial Diagnostics & Root Cause Analysis
SIGTERM (signal 15): A "polite request" to terminate. The process can catch it, perform cleanup (flush buffers, close connections), and exit cleanly.
- When scaling down or rolling out updates, the orchestrator sends SIGTERM to the application (e.g., 30-second grace period).
- The application must catch SIGTERM and gracefully shutdown:
- During shutdown: close client connections, finish in-flight requests, flush logs, release resources.
- If the process doesn't exit cleanly within the grace period, Kubernetes sends SIGKILL, forcefully terminating it (potential data loss or corrupted state).
2️⃣
Remediation & Permanent Safeguards
SIGKILL (signal 9): A "forced kill" that cannot be caught or ignored. The kernel immediately terminates the process, potentially leaving corrupted state. In production deployments (especially Kubernetes): Applications that don't handle SIGTERM risk: Example: A web server catching SIGTERM stops accepting *new* connections but finishes serving existing requests before exiting.
trap 'graceful_shutdown' SIGTERM
- Lost in-flight requests
- Corrupted database transactions
- Connection pool exhaustion (clients hang waiting for responses)
💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"Pro-Tip: When scaling down or rolling out updates, the orchestrator sends SIGTERM to the application (e.g., 30-second grace period).."
⚡ 60-Second Elevator Pitch Talking Points
- When scaling down or rolling out updates, the orchestrator sends SIGTERM to the application (e.g....
- The application must catch SIGTERM and gracefully shutdown:
- During shutdown: close client connections, finish in-flight requests, flush logs, release resources.
Advertisement