Q: A user reports 10-second delays every 15 minutes in an application running on AKS. No code changes happened. How would you begin root cause analysis?
Forensic root cause analysis workflow to track down predictable, periodic 10-second latency spikes occurring on a strict 15-minute cadence without code deployments.
Want to master this scenario in a live sandbox? The Linux Foundation's Prometheus Certified Associate (PCA) & Monitoring Labs covers this exact problem with hands-on terminal drills.
🛠️ Production Runbook & Step-by-Step Resolution
Correlate Latency Spikes with Cluster CronJobs & DaemonSets
Inspect Kubernetes scheduled workloads. Search for any CronJob or batch task running with the schedule `*/15 * * * *` in any namespace on the cluster. A scheduled database backup or cache warm-up running on the same worker node can saturate node disk I/O or CPU bandwidth.
kubectl get cronjobs -A
# Look for SCHEDULE column matching: */15 * * * * or 0,15,30,45 * * * *
Analyze Internal Application Background Threads & Cache Eviction
If no external CronJob exists, inspect the application's internal scheduler (e.g. `@Scheduled(cron = '0 */15 * * * *')` in Spring Boot, Celery beat, or Go ticker). Applications frequently refresh an in-memory cache every 15 minutes; if the refresh runs synchronously with a global read-write lock, all user threads freeze for 10 seconds.
# Capture thread dump right during the 15-minute spike
kubectl exec -it <pod> -- jstack <pid> | grep -E 'locked|waiting to lock'
Inspect Node-Level I/O Bursting & EBS/Managed Disk Credits
Check Azure Managed Disk throttling metrics in Azure Monitor. If the AKS node uses standard burstable SSDs (P10/P20) and a periodic log rotation or OS indexing job exhausts IOPS burst tokens, disk operations stall until credits replenish.
# Check node disk I/O wait times
kubectl exec -it <debug-node> -- iostat -xz 1 10
# High %iowait and %util indicate disk starvation
Correlate Downstream Database & External API Scheduled Syncs
Query database slow query logs for queries executing every 15 minutes. Heavy analytical batch queries holding table-level shared locks will cause fast transaction queries to queue behind them.
- Search for cluster CronJobs configured with schedule */15 * * * *.
- Inspect application source code for internal scheduled cache-refresh routines holding write locks.
- Monitor Azure Managed Disk IOPS and credit throttling metrics on the node hosting the pod.
- Check database slow query logs for periodic analytical queries locking active transaction tables.