Q: How do you monitor container performance in production?
Layered observability architecture for container performance in production: diagnosing cgroup CPU throttling, memory working sets, cAdvisor and Prometheus metrics, docker stats, and SLO-based alerting.
#Docker #cgroups #cAdvisor #Prometheus #Grafana #Observability #CPU Throttling
🎙️ Candidate Opening & Architectural Context
"I use layered observability: container and host metrics (CPU, memory, filesystem, cgroup throttling), application performance metrics (latency, error rate, throughput), logs, and distributed traces. In Kubernetes, Prometheus, Grafana, and Alertmanager with cAdvisor are standard; for standalone Docker hosts, I use cAdvisor and node-exporter alongside centralized log shipping. I focus alerts on SLO breaches and container throttling rather than noisy raw utilization thresholds."
Advertisement
🛠️ Production Runbook & Step-by-Step Resolution
1️⃣
Fast CLI Diagnostics on Container Hosts
Immediate command-line triage when diagnosing container slowdowns on a live host:
# Real-time resource usage stream across running containers
docker stats --no-stream
# Inspect OOM status and termination details
docker inspect <container_id> --format '{{json .State}}' | jq
# Stream container lifecycle events from the last 30 minutes
docker events --since 30m
- Live Resource Streaming: Run
docker statsfor real-time CPU %, memory usage, limits, and network I/O. - Inspect Container State & OOM: Check container exit codes and OOMKilled state flags with
docker inspect. - Container Events: Stream recent container lifecycle events (die, oom, kill) with
docker events.
2️⃣
cgroup Throttling & Prometheus Metrics Architecture
Monitor kernel-level cgroup metrics to catch subtle CPU starvation and memory saturation:
# Prometheus Query Examples:
# 1. Detect CFS CPU Throttling Rate (indicates undersized CPU limits):
# rate(container_cpu_cfs_throttled_seconds_total[5m]) > 0.2
# 2. Container Memory Working Set vs Limit (OOM risk):
# sum(container_memory_working_set_bytes) by (pod) / sum(container_spec_memory_limit_bytes) by (pod) > 0.85
# Live Kubernetes cluster checks
kubectl top pods -A
kubectl top nodes
- cAdvisor & Prometheus: cAdvisor scrapes container cgroup data directly from
/sys/fs/cgroupand exposes metrics for Prometheus. - CPU CFS Throttling: Monitor
container_cpu_cfs_throttled_seconds_total. When CPU limits are too tight, the CFS scheduler throttles threads, spiking tail latency even if CPU % looks low. - Memory Working Set: Alert on
container_memory_working_set_bytesapproaching container limits, as this is the exact metric the Linux kernel uses to trigger OOMKills. - SLO-Driven Alerts: Alert when p95/p99 latency degrades or when containers restart repeatedly, rather than alerting on arbitrary CPU utilization.
💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"Monitor CFS CPU throttling (container_cpu_cfs_throttled_seconds_total) and memory working set rather than simple CPU averages. CPU throttling causes severe latency spikes long before a container crashes."
⚡ 60-Second Elevator Pitch Talking Points
- Implement layered monitoring: cAdvisor/node-exporter for cgroups, Prometheus/Grafana for metrics, and centralized tracing.
- Track critical cgroup metrics: CFS CPU throttling rate and memory working set to prevent silent latency degradation and OOMKills.
- Configure alerting around user-impacting Golden Signals (latency, errors, saturation) rather than static host thresholds.
Advertisement