⚡ ~/naveed Interview Prep
⚡ Portfolio Home ✍️ Engineering Blog Deep Dives 🎯 Interview Hub 998+ Scenarios ☸️ Kubernetes Mastery Hub 24 Modules 🎮 DevOps Arcade & Quizzes Subnet Blitz ⚡ 🗺️ DevOps Roadmaps PDFs & Guides 🤖 Morpheus Analysis AI Quant ↗ 🛠️ Developer Tools Utilities 🧪 Labs & Experiments 📄 Interactive CV & Certs 🔗 All Links & Socials ⚡ Join The Dispatch (Weekly SRE Newsletter) →
Staff+ / Principal Architect Observability Chaos Engineering & Scalability Netflix-Scale Systems

Q: Describe how you’d test infra chaos and graceful degradation for a Netflix Originals release.

How to design pre-launch chaos experiments, synthetic stress tests, and automated tiered graceful degradation for a global high-concurrency streaming premiere.

#Chaos Engineering #Netflix #Graceful Degradation #Chaos Monkey #SRE #Load Testing
🎙️ Candidate Opening & Architectural Context
"A tier-1 global streaming release concentrates millions of concurrent requests within a 5-minute window. Relying on auto-scaling alone is a guaranteed recipe for outages because cold EC2 instances and EKS node groups take 2 to 5 minutes to provision. To guarantee survival, the system must undergo automated chaos testing and enforce three distinct tiers of automated graceful degradation."
Advertisement

🛠️ Production Runbook & Step-by-Step Resolution

1️⃣

Define the 3 Tiers of Graceful Degradation

Decouple core business transactions from auxiliary personalization services:

  • Tier 1 (Critical Path - Must Never Fail): User Authentication, CDN Playback Token Generation, and Video Stream Manifest Delivery. Must have dedicated thread pools and zero dynamic dependencies.
  • Tier 2 (Degradable Features): Personalized Recommendations, Continue Watching rows, and dynamic search. If latency exceeds 200ms, circuit breakers trip and return pre-computed static Top 10 JSON from CDN edge.
  • Tier 3 (Sacrificial Features): User ratings, bookmark syncing, watch history analytics, and email notifications. Immediately shed via rate limiting and async queues during traffic spikes.
2️⃣

Pre-Release Chaos Drills (Chaos Kong & Fault Injection)

Execute progressive failure drills weeks before the launch date:

  • Availability Zone Termination (Chaos Gorilla): Terminate an entire AWS Availability Zone during peak synthetic load to verify that ALB, EKS, and Aurora continue operating without dropped connections.
  • Dependency Latency Injection: Inject 500ms artificial delay into the user recommendations database using Envoy fault injection. Verify that playback services trip circuit breakers cleanly without timing out.
  • Cache Stampede Simulation: Flush Redis caches under 100,000 rps load to ensure SingleFlight / mutex locking prevents dogpiling the underlying relational database.
3️⃣

Pre-Warming & Dark Canary Traffic

Never allow cold infrastructure into a global release:

  • Pre-scale EKS node groups and database read replicas 2 hours prior to the premiere.
  • Replay recorded production traffic (Shadow / Dark Traffic) at 2x scale against the canary deployment to validate downstream microservice limits.
💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"Extreme scale resilience is built on graceful degradation: when infrastructure is overwhelmed, non-essential personalization sheds automatically so that core video playback remains 100% uninterrupted."
⚡ 60-Second Elevator Pitch Talking Points
  • We design for survivability by categorizing services into 3 tiers: Tier 1 (Playback & Auth) must never fail; Tier 2 (Recommendations) degrades to pre-computed static JSON; Tier 3 (Analytics & Bookmarks) sheds completely under load.
  • We test this using progressive chaos experiments: terminating an entire Availability Zone (Chaos Gorilla) under 2x synthetic load, and injecting latency into dependencies to ensure circuit breakers trip gracefully.
  • We simulate cache stampedes to verify SingleFlight locking prevents database thrashing.
  • Finally, we pre-warm infrastructure and replay dark production traffic to validate capacity before the launch window begins.
Advertisement
Want more Observability scenarios?
Explore our complete collection of scenario-based Observability interview runbooks.
Browse All Observability Questions →

📚 Related Production Scenarios in Observability