Google Systems SRE (L5) Interview Loop Breakdown: NALSD & Debugging
1. Loop Overview & Candidate Context
Comprehensive debrief of Google's Systems SRE track: Non-Abstract Large System Design (NALSD), live hands-on Linux terminal debugging, concurrent coding in Go/Python, and Googleyness.
2. Detailed Round-by-Round Breakdown
Round 1: Practical Systems Troubleshooting (60 mins)
Candidate given SSH access to a broken VM where an API server was dropping packets intermittently. Successfully identified conntrack table saturation and disk I/O fsync throttling in 15 minutes.
Round 2: Non-Abstract Large System Design - NALSD (60 mins)
Design a globally distributed time-series metrics ingestion pipeline handling 100M events/second. Sized network bandwidth, IOPS, cache invalidation, and failure domains.
Round 3: Software Engineering & Data Structures (60 mins)
Implement a distributed thread-safe rate limiter (token bucket algorithm) with multi-worker concurrency and atomic locks.
Round 4: Linux Systems Internals (60 mins)
Deep dive into memory paging, virtual memory vs RSS, page faults, epoll vs select, and process scheduling latency.
Round 5: Googleyness & Blameless Culture (45 mins)
Evaluated collaboration, navigating ambiguity, ethical tech choices, and fostering blameless engineering culture.
โก Exact Scenarios Asked & Matching Runbooks on This Hub:
The candidate encountered variations of these scenarios. Study the step-by-step diagnostic runbooks below:
3. Candidate Retrospective: What Worked & Advice
- Read Google's SRE Book (especially chapters on Service Level Objectives, Alerting, and Eliminating Toil).
- In the NALSD round, calculate concrete numbers upfront: QPS, RAM requirements, Disk IOPS, and network egress bandwidth.
- In the debugging round, communicate your hypothesis aloud before running any command: 'I suspect connection drops are due to conntrack table overflow; let me verify with conntrack -S'.