⚡ ~/naveed Interview Prep
⚡ Portfolio Home ✍️ Engineering Blog Deep Dives 🎯 Interview Hub 1,000+ Scenarios ☸️ Kubernetes Mastery Hub 24 Modules 🎮 DevOps Arcade & Quizzes Subnet Blitz ⚡ 🗺️ DevOps Roadmaps PDFs & Guides 🤖 Morpheus Analysis AI Quant ↗ 🛠️ Developer Tools Utilities 🧪 Labs & Experiments 📄 Interactive CV & Certs 🔗 All Links & Socials ⚡ Join The Dispatch (Weekly SRE Newsletter) →
← Back to All AI/ML Infrastructure & GPU Interview Questions Scenario 8 of 50 in AI/ML Infrastructure & GPU
Staff AI Infrastructure Engineer AI/ML Infrastructure High-Performance Networking Networking
🎯 Target Role / Context: Staff AI Platform Engineer designing multi-thousand GPU leaf-spine fabrics for foundational model training.

Q: Why does distributed AI training demand a lossless network fabric, and how does RoCE v2 achieve lossless transmission over standard Ethernet? What is a PFC deadlock storm, and how do you configure DSCP, PFC headroom buffers, and ECN (WRED) thresholds to prevent it?

Deep architectural analysis of InfiniBand vs RoCE v2 fabrics, resolving Priority Flow Control (PFC) deadlocks, storm propagation, and ECN buffer marking in large-scale AI training clusters.

#InfiniBand #RoCE v2 #PFC Deadlock #ECN #RDMA #AI/ML Infra
🎙️ Candidate Opening & Architectural Context
"Standard TCP/IP retransmissions introduce tail latency spikes that decimate all-reduce collective throughput during distributed training. RDMA (Remote Direct Memory Access) bypasses the host kernel and CPU, allowing GPUs to stream tensors directly into remote GPU VRAM (GPUDirect RDMA). While InfiniBand provides native credit-based flow control at L2, RoCE v2 runs over Ethernet and relies on Priority Flow Control (PFC) and Explicit Congestion Notification (ECN) to maintain lossless delivery."
Advertisement
⚡ Recommended Practice Lab

Want to master this scenario in a live sandbox? KodeKloud's CKA & CKAD Hands-On Certification Track covers this exact problem with hands-on terminal drills.

🛠️ Production Runbook & Step-by-Step Resolution

1

Mechanism of Lossless RoCE v2: PFC and Pause Frames

RoCE v2 encapsulates InfiniBand transport packets inside standard UDP/IP frames (UDP port 4791). Lossless behavior is enforced using IEEE 802.1Qbb Priority Flow Control (PFC). When a switch ingress buffer fills past a threshold, the switch transmits an 802.1Qbb PAUSE frame upstream for a specific CoS (Class of Service) priority (typically priority 3), halting the sender until buffer capacity recovers.

# Verify PFC configuration on ConnectX-7 NIC
mlnx_qos -i ens1f0np0
# Output shows: Priority 3 -> PFC enabled, Trust DSCP
2

Diagnose and Prevent PFC Deadlock and Pause Storms

A PFC deadlock occurs in cyclic buffer dependency topologies (e.g., ring all-reduce or multi-hop routing loops): switch A pauses switch B, switch B pauses switch C, and switch C pauses switch A. All traffic halts permanently across the fabric. To prevent this, configure PFC Deadlock Detection and Recovery (PFC DLD/DLR) on switches to drop packets and release pause states if a queue stays paused longer than a threshold (e.g., 100ms).

# Arista / SONiC switch PFC watchdog config
pfc-watchdog-interval 100
pfc-watchdog-action drop
Advertisement
3

Tune ECN and WRED Marking to Throttle Traffic Before PFC Triggers

PFC is a blunt emergency brake. To prevent PFC from triggering frequently, configure Explicit Congestion Notification (ECN) with Weighted Random Early Detection (WRED) on switch queues. Set ECN marking thresholds (K_min, K_max) well below the PFC XOFF pause threshold. When buffers reach K_min, switches mark the ECN bits in the IP header. The receiver sends a Congestion Notification Packet (CNP) back to the sender, instructing the NIC hardware to throttle its transmission rate smoothly before pause frames are ever generated.

# RoCE ECN switch thresholds (SONiC example)
queue 3:
  wred_green_min_threshold: 150KB
  wred_green_max_threshold: 1500KB
  wred_green_mark_probability: 20
  pfc_xoff_threshold: 2000KB
  pfc_xon_threshold: 1600KB
💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"RoCE v2 achieves lossless RDMA via PFC, but cyclic pause dependencies can trigger catastrophic PFC deadlocks. Tuning ECN/WRED thresholds below PFC trigger points enables proactive congestion throttling via CNPs, keeping fabrics fast and stable."
⚡ 60-Second Elevator Pitch Talking Points
  • AI training requires lossless networks because a single dropped packet stalls the entire distributed cluster.
  • RoCE v2 uses Priority Flow Control (PFC) to send hardware pause frames, but pause storms can lock up the entire fabric in a cyclic deadlock.
  • We tune ECN/WRED thresholds below the PFC trigger point: switches mark packets early, NICs throttle gracefully via CNPs, and PFC watchdog timers drop poisoned loops if deadlocks ever occur.
Advertisement
Want more AI/ML Infrastructure & GPU scenarios?
Explore our complete collection of scenario-based AI/ML Infrastructure & GPU interview runbooks.
Browse All AI/ML Infrastructure & GPU Questions →