Q: The system is experiencing unexplained network latency. `ping` times internally jump from 1ms to 200ms periodically. How do you isolate the problem?
Intermittent high latency is tricky. My workflow would be:
#Linux #Linux / SRE — Scenario-Based Interview Questions #L3 #SRE #Systems #Troubleshooting
🎙️ Candidate Opening & Architectural Context
""Never reboot a server blindly; always capture top process telemetry, lsof descriptors, and thread dumps first. The interviewer is testing: Advanced network troubleshooting, packet capture, kernel networking stack.. I structure my answer around systematic triage first, root cause analysis second, and permanent remediation third.""
Advertisement
🛠️ Production Runbook & Step-by-Step Resolution
1️⃣
Initial Diagnostics & Root Cause Analysis
Intermittent high latency is tricky. My workflow would be:
- Scope: Is it affecting all instances? One AZ? One specific app? I'd ping the default gateway and another instance in the same subnet to isolate whether it's the host network stack or the physical network/hypervisor.
- Host metrics: Check
dmesgorjournalctlfor NIC errors, ring buffer drops, ornf_conntracktable full messages. - Traffic Analysis: Use
mtrto track packet loss at hops. To capture the issue, I'd runtcpdump -i eth0 -w capture.pcapduring the latency spikes.
2️⃣
Remediation & Permanent Safeguards
Based on the finding, I might tune ring buffers (ethtool -G), adjust sysctl buffers, or escalate to the cloud provider if the host metrics check out perfectly.
- Kernel Drops: Use
netstat -suto see if UDP/TCP packets are being dropped by the kernel due to full receive buffers. I'd verifytopto check if a specific CPU core handling NIC interrupts (ksoftirqd) is pinned at 100%, causing a processing bottleneck.
💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"Pro-Tip: Scope: Is it affecting all instances? One AZ? One specific app? I'd ping the default gateway and another instance in the same subn."
⚡ 60-Second Elevator Pitch Talking Points
- Scope: Is it affecting all instances? One AZ? One specific app? I'd ping the default gateway and ...
- Host metrics: Check dmesg or journalctl for NIC errors, ring buffer drops, or nf_conntrack table ...
- Traffic Analysis: Use mtr to track packet loss at hops. To capture the issue, I'd run tcpdump -i ...
Advertisement