Q: A pipeline that normally takes 12 minutes suddenly takes 40 minutes without major code changes. How would you find the bottleneck?
Step-by-step diagnostic workflow to isolate and remediate sudden pipeline runtime inflation from 12 minutes to 40 minutes using timeline analysis, dependency caching, and parallel execution.
🛠️ Production Runbook & Step-by-Step Resolution
Analyze Task Duration Breakdown in Timeline View
Open the failed run in Azure DevOps and check the 'Timeline' view. Compare each step's duration against the 12-minute baseline run to instantly isolate which specific task swelled in duration.
Check Cache Tasks (Cache@2 Misses)
If npm install, pip install, or Docker layer caching was invalidated by a modified lockfile or cache key miss, the agent downloads gigabytes of packages over the network on every run.
- task: Cache@2
inputs:
key: 'npm | "$(Agent.OS)" | package-lock.json'
restoreKeys: |
npm | "$(Agent.OS)"
path: $(npm_config_cache)
displayName: Cache npm packages
Inspect Agent Host Disk IOPS & CPU Limits
For self-hosted Azure VMs, check Azure Monitor metrics for OS Disk IOPS consumption. If using Standard HDD or bursting Standard SSDs, running out of IOPS credits throttles disk read/write throughput to 500 KB/s.
- Inspect the Azure DevOps Timeline view to identify the exact task consuming the additional 28 minutes.
- Audit Cache@2 tasks for cache misses caused by package-lock.json drift.
- Verify self-hosted agent VM disk metrics: check for Premium SSD IOPS throttling.
- Implement parallel test slicing across multiple jobs to regain 12-minute SLA.