Q: A high-throughput microservice in production experiences periodic 200ms latency spikes. Traditional application APM agents and CPU profilers show normal CPU usage and no garbage collection pauses. Profiling with `strace` in production caused a 400% performance degradation and crashed client connections. You must use eBPF and `bpftrace` to trace disk I/O latency and network packet drops directly in the Linux kernel filtered by the container's cgroup ID with less than 1% overhead.
Achieve kernel-level observability into Docker containers without injecting code or restarting processes. Use eBPF and `bpftrace` to trace container filesystem latency, dropped packets, and DNS queries by cgroup.
Want to master this scenario in a live sandbox? KodeKloud's Docker Certified Associate (DCA) Hands-On Lab Course covers this exact problem with hands-on terminal drills.
🛠️ Production Runbook & Step-by-Step Resolution
Understand Why eBPF Outperforms strace for Container Observability
`strace` relies on the `ptrace` system call, which intercepts every system call, pauses the target process twice (on entry and exit), and forces context switches to user space. eBPF executes sandboxed bytecode directly inside the Linux kernel on kprobes and tracepoints with sub-microsecond overhead.
<!-- Performance Impact -->
strace: Pauses container process on EVERY syscall! Overhead: 50% - 400% (Never run in Prod!)
eBPF: JIT-compiled in-kernel verification. Copies only aggregated histograms. Overhead: < 1%
Identify Container Cgroup ID for Scoped Filtering
Obtain the container's 64-bit cgroup ID or namespace inode number to filter eBPF probes specifically to that container, avoiding host-wide trace noise.
# Get container cgroup v2 ID
CONTAINER_ID=$(docker inspect my-app --format '{{.Id}}')
CGROUP_PATH="/sys/fs/cgroup/docker/${CONTAINER_ID}"
CGROUP_ID=$(python3 -c "import os; print(os.stat('$CGROUP_PATH').st_ino)")
echo "Filtering eBPF for cgroup ID: $CGROUP_ID"
Trace Container Disk I/O Latency using bpftrace One-Liner
Attach a `bpftrace` program to the kernel's block I/O tracepoints, generating a power-of-two histogram of disk latency specifically for the container.
# Trace block I/O latency histogram (in microseconds) for container cgroup
sudo bpftrace -e '
tracepoint:block:block_bio_issue {
if (cgroup == '$CGROUP_ID') {
@start[args->dev, args->sector] = nsecs;
}
}
tracepoint:block:block_bio_complete {
if (@start[args->dev, args->sector]) {
@lat_us = hist((nsecs - @start[args->dev, args->sector]) / 1000);
delete(@start[args->dev, args->sector]);
}
}
'
Trace Dropped TCP Packets and TCP Retransmissions
Monitor kernel TCP retransmissions and dropped packets per container using the BCC `tcpretrans` tool or bpftrace kernel tracepoints.
# Monitor TCP retransmits across container processes
sudo /usr/share/bcc/tools/tcpretrans -c
- W
- e
- d
- i
- a
- g
- n
- o
- s
- e
- d
- m
- i
- c
- r
- o
- s
- e
- r
- v
- i
- c
- e
- l
- a
- t
- e
- n
- c
- y
- s
- p
- i
- k
- e
- s
- i
- n
- p
- r
- o
- d
- u
- c
- t
- i
- o
- n
- w
- i
- t
- h
- o
- u
- t
- a
- d
- d
- i
- n
- g
- A
- P
- M
- o
- v
- e
- r
- h
- e
- a
- d
- b
- y
- u
- s
- i
- n
- g
- e
- B
- P
- F
- a
- n
- d
- `
- b
- p
- f
- t
- r
- a
- c
- e
- `
- .
- W
- h
- i
- l
- e
- `
- s
- t
- r
- a
- c
- e
- `
- w
- o
- u
- l
- d
- h
- a
- v
- e
- d
- e
- g
- r
- a
- d
- e
- d
- p
- e
- r
- f
- o
- r
- m
- a
- n
- c
- e
- ,
- o
- u
- r
- e
- B
- P
- F
- t
- r
- a
- c
- e
- p
- o
- i
- n
- t
- f
- i
- l
- t
- e
- r
- e
- d
- b
- y
- t
- h
- e
- c
- o
- n
- t
- a
- i
- n
- e
- r
- '
- s
- c
- g
- r
- o
- u
- p
- I
- D
- i
- d
- e
- n
- t
- i
- f
- i
- e
- d
- t
- h
- a
- t
- s
- l
- o
- w
- N
- V
- M
- e
- b
- l
- o
- c
- k
- w
- r
- i
- t
- e
- s
- w
- e
- r
- e
- b
- l
- o
- c
- k
- i
- n
- g
- a
- p
- p
- l
- i
- c
- a
- t
- i
- o
- n
- w
- o
- r
- k
- e
- r
- t
- h
- r
- e
- a
- d
- s
- ,
- l
- e
- a
- d
- i
- n
- g
- t
- o
- a
- n
- i
- m
- m
- e
- d
- i
- a
- t
- e
- E
- B
- S
- v
- o
- l
- u
- m
- e
- c
- l
- a
- s
- s
- u
- p
- g
- r
- a
- d
- e
- .