Q: Explain the concept of Chaos Engineering and how you would implement a basic chaos experiment on a Linux production system.
Chaos Engineering is the discipline of proactively injecting controlled failures into production systems to discover weaknesses before th...
🛠️ Production Runbook & Step-by-Step Resolution
Initial Diagnostics & Root Cause Analysis
Chaos Engineering is the discipline of proactively injecting controlled failures into production systems to discover weaknesses before they cause real outages.
- Start with a hypothesis: "If one app server fails, the load balancer will route traffic to healthy servers within 10 seconds."
- Minimize blast radius: Start with non-critical environments and gradually move to production.
- Run experiments during business hours: When engineers are available to respond.
- Automate rollback: Every experiment must have an automatic abort condition.
- CPU stress test:
Remediation & Permanent Safeguards
Core principles (from Netflix's Chaos Engineering): Basic chaos experiments on Linux: Validates: Does auto-scaling trigger? Do health checks detect degraded performance? Validates: Do circuit breakers activate? Do timeouts work correctly? Validates: Does systemd restart the process? Does the load balancer detect the failure? Validates: Do disk space alerts fire? Does the application handle disk-full errors gracefully? Tools like Chaos Monkey, Litmus Chaos, and Gremlin automate these experiments at scale with safety controls.
stress-ng --cpu 4 --timeout 300s
- Network latency injection:
- Process kill:
- Disk fill:
- Start with a hypothesis: "If one app server fails, the load balancer will route traffic to health...
- Minimize blast radius: Start with non-critical environments and gradually move to production.
- Run experiments during business hours: When engineers are available to respond.