Q: What should a production health check endpoint verify, and what should it avoid?
A health check should be cheap, fast, and designed for the action that will be taken when it fails.
🛠️ Production Runbook & Step-by-Step Resolution
Production Solution & Architecture
A health check should be cheap, fast, and designed for the action that will be taken when it fails. For a liveness endpoint, I keep it shallow: can the process respond, is the main event loop alive, and is the app not deadlocked? It should not call every dependency, because a temporary database issue could cause Kubernetes to restart healthy pods unnecessarily. For a readiness endpoint, I check whether the app can safely receive traffic: required config loaded, database connection pool initialized, cache warmed if required, and migrations compatible. Avoid expensive queries, calls to optional third-party services, or checks that can overload dependencies during an outage. Bad health checks can turn a small dependency issue into a full restart storm.
- Immediate Triage: A health check should be cheap, fast, and designed for the action that will be taken when it fa
- Run targeted verification commands before modifying configuration.
- Automate permanent guardrails (CI check, alerts, IaC policy) to prevent recurrence.