Q: A memory leak in a Java Spring Boot application causes the JVM memory to grow indefinitely over several days until it crashes. While the developers investigate the root cause, what immediate SRE mitigation can you apply to keep the service stable for users?
While fixing the root cause is the developer's job, SREs must preserve uptime. If we know the leak takes roughly 48 hours to crash the pr...
🛠️ Production Runbook & Step-by-Step Resolution
Production Solution & Architecture
While fixing the root cause is the developer's job, SREs must preserve uptime. If we know the leak takes roughly 48 hours to crash the process, a safe, immediate mitigation is to enforce automated recycling before the threshold is reached. In Kubernetes, you would configure a strict memory limit combined with a Liveness Probe. If the process locks up, the liveness probe fails, and the Kubelet aggressively restarts the Pod, yielding a fresh, empty JVM. Alternatively, use cron or orchestration to gently drain and recycle the instances every 24 hours during low-traffic periods, completely preventing the runaway limit from being hit while developers buy time to fix the code.
- Immediate Triage: While fixing the root cause is the developer's job, SREs must preserve uptime. If we know the l
- Run targeted verification commands before modifying configuration.
- Automate permanent guardrails (CI check, alerts, IaC policy) to prevent recurrence.