Q: A deployment completed successfully, but response time increased from 200 ms to 3 seconds. Walk me through your troubleshooting approach.
Step-by-step diagnostic workflow for triaging a 15x response time regression (200ms to 3s) immediately following a production deployment without catastrophic error spikes.
🛠️ Production Runbook & Step-by-Step Resolution
Immediate Blast Radius Containment
Check if a canary, blue-green, or rolling deployment is in progress. If traffic is still on a subset of pods, halt rollout immediately or divert traffic back to the known good baseline while preserving 1-2 regressed pods for live forensic debugging.
Distributed Tracing Differential Comparison
Open Datadog, Jaeger, or AWS X-Ray and compare latency flame graphs of the new release versus the previous tag. Pinpoint whether the extra 2.8 seconds is spent on external HTTP calls, database query lock waiting, or CPU execution.
# Inspect container CPU throttling and garbage collection pauses
kubectl top pod -l app=payment-service
cat /sys/fs/cgroup/cpu/cpu.stat | grep nr_throttled
Database Query & Connection Pool Analysis
Check database slow query logs and active lock trees. Frequently, a new feature introduces an unindexed query (causing full table scans) or an N+1 query pattern where a loop issues 50 sequential queries instead of a batch fetch.
-- PostgreSQL lock inspection query
SELECT pid, age(clock_timestamp(), query_start), usename, query, state
FROM pg_stat_activity
WHERE state != 'idle' AND age(clock_timestamp(), query_start) > interval '2 seconds';
Cgroup CPU Throttling & Memory Pressure
Verify whether the new image introduced heavier memory consumption causing frequent JVM/Go GC cycles or exceeded Kubernetes CPU limits resulting in CFS throttling.
- Halt rollout immediately or route traffic back to the prior stable version while retaining one canaried instance for debugging.
- Compare APM distributed traces (flame graphs) between v1 and v2 to isolate the exact function or query responsible.
- Check database slow query logs for missing indexes or sequential N+1 query patterns.
- Examine container cgroup metrics for CFS CPU throttling (nr_throttled) and GC execution duration.