⚡ ~/naveed Interview Prep
⚡ Portfolio Home ✍️ Engineering Blog Deep Dives 🎯 Interview Hub 1,000+ Scenarios ☸️ Kubernetes Mastery Hub 24 Modules 🎮 DevOps Arcade & Quizzes Subnet Blitz ⚡ 🗺️ DevOps Roadmaps PDFs & Guides 🤖 Morpheus Analysis AI Quant ↗ 🛠️ Developer Tools Utilities 🧪 Labs & Experiments 📄 Interactive CV & Certs 🔗 All Links & Socials ⚡ Join The Dispatch (Weekly SRE Newsletter) →
← Back to All Observability & Monitoring Interview Questions Scenario 93 of 96 in Observability & Monitoring
Senior SRE / Performance Engineer Observability Periodic Latency Spikes & RCA J.P. Morgan Technical Loop

Q: A user reports 10-second delays every 15 minutes in an application running on AKS. No code changes happened. How would you begin root cause analysis?

Forensic root cause analysis workflow to track down predictable, periodic 10-second latency spikes occurring on a strict 15-minute cadence without code deployments.

#Observability #AKS #Kubernetes #Latency #CronJob #RCA #Garbage Collection
🎙️ Candidate Opening & Architectural Context
"When latency spikes occur on a strict, mathematical interval (e.g. exactly every 15 minutes), the root cause is virtually guaranteed to be a scheduled periodic background task: a Kubernetes CronJob executing on the same node, an internal application cache flush/warmup routine, periodic JVM major garbage collection, or cloud storage snapshotting."
Advertisement
⚡ Recommended Practice Lab

Want to master this scenario in a live sandbox? The Linux Foundation's Prometheus Certified Associate (PCA) & Monitoring Labs covers this exact problem with hands-on terminal drills.

🛠️ Production Runbook & Step-by-Step Resolution

1

Correlate Latency Spikes with Cluster CronJobs & DaemonSets

Inspect Kubernetes scheduled workloads. Search for any CronJob or batch task running with the schedule `*/15 * * * *` in any namespace on the cluster. A scheduled database backup or cache warm-up running on the same worker node can saturate node disk I/O or CPU bandwidth.

kubectl get cronjobs -A
# Look for SCHEDULE column matching: */15 * * * * or 0,15,30,45 * * * *
2

Analyze Internal Application Background Threads & Cache Eviction

If no external CronJob exists, inspect the application's internal scheduler (e.g. `@Scheduled(cron = '0 */15 * * * *')` in Spring Boot, Celery beat, or Go ticker). Applications frequently refresh an in-memory cache every 15 minutes; if the refresh runs synchronously with a global read-write lock, all user threads freeze for 10 seconds.

# Capture thread dump right during the 15-minute spike
kubectl exec -it <pod> -- jstack <pid> | grep -E 'locked|waiting to lock'
Advertisement
3

Inspect Node-Level I/O Bursting & EBS/Managed Disk Credits

Check Azure Managed Disk throttling metrics in Azure Monitor. If the AKS node uses standard burstable SSDs (P10/P20) and a periodic log rotation or OS indexing job exhausts IOPS burst tokens, disk operations stall until credits replenish.

# Check node disk I/O wait times
kubectl exec -it <debug-node> -- iostat -xz 1 10
# High %iowait and %util indicate disk starvation
4

Correlate Downstream Database & External API Scheduled Syncs

Query database slow query logs for queries executing every 15 minutes. Heavy analytical batch queries holding table-level shared locks will cause fast transaction queries to queue behind them.

💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"A strict 15-minute latency cadence points directly to scheduled jobs: CronJobs on the node, internal cache eviction holding synchronized locks, disk IOPS burst credit exhaustion, or periodic DB batch locks."
⚡ 60-Second Elevator Pitch Talking Points
  • Search for cluster CronJobs configured with schedule */15 * * * *.
  • Inspect application source code for internal scheduled cache-refresh routines holding write locks.
  • Monitor Azure Managed Disk IOPS and credit throttling metrics on the node hosting the pod.
  • Check database slow query logs for periodic analytical queries locking active transaction tables.
Advertisement
Want more Observability & Monitoring scenarios?
Explore our complete collection of scenario-based Observability & Monitoring interview runbooks.
Browse All Observability & Monitoring Questions →