SE-Radio Episode 325: Tammy Butow on Chaos Engineering

Topics covered
Popular Clips
Episode Highlights
Basics
, a Principal Site Reliability Engineer at Gremlin, introduces the foundational concepts of chaos engineering. She emphasizes the importance of starting with simple experiments, like a "hello, world" of chaos engineering, to understand the impact of failure injection on systems 1. Tammy provides practical insights, such as using bash scripts to simulate infrastructure failures, which can be run on demo machines to observe their effects 1.
Chaos engineering is definitely a journey. You need to get everybody on board and treat it like a marathon that you're all going on together.
---
She stresses the importance of having robust monitoring tools and a high-severity incident management program in place before conducting chaos experiments 2.
Strategies
Tammy discusses various strategies for implementing chaos engineering, highlighting the need for a structured approach. She explains that some organizations have dedicated chaos teams, while others integrate chaos engineering into existing roles like SREs or production engineers 3. The key is to ensure that the entire team is aware of the experiments and their outcomes, often facilitated by tools like Gremlin and Slack integrations for real-time updates 4.
You need to make it very visible. You need to let people know what's happening when.
---
Tammy advises starting with small, controlled experiments and gradually expanding as the organization becomes more comfortable with chaos engineering practices 5.
Techniques
Exploring specific failure injection techniques, Tammy shares examples from her workshops, such as CPU and network attacks, which are available on her GitHub 6. She explains the concept of a "black hole attack," where network traffic is dropped, simulating a severe outage 6. Monitoring tools like Datadog are used to visualize the impact of these attacks, providing immediate feedback on system resilience 7.
If you think about it, you're creating a black hole and you're sending all of the traffic into this black hole.
---
These techniques help teams understand potential vulnerabilities and improve their systems' robustness against unexpected failures 6.
Related Episodes


SE-Radio Episode 288: DevSecOps
Answers 383 questions

SE-Radio Episode 301: Jason Hand Handling Outages
Answers 383 questions

SE-Radio Episode 247: Andrew Phillips on DevOps
Answers 383 questions

Episode 474: Paul Butcher on Fuzz Testing
Answers 383 questions

Episode 112: Roles in Software Engineering II
Answers 383 questions

SE Radio 585: Adam Frank on Continuous Delivery vs Continuous Deployment
Answers 383 questions

SE-Radio Episode 344: Pat Helland on Web Scale
Answers 383 questions

SE-Radio Episode 276: Björn Rabenstein on Site Reliability Engineering
Answers 383 questions

SE Radio 572: Gregory Kapfhammer on Flaky Tests
Answers 383 questions

SE-Radio Episode 256: Jay Fields on Working Effectively with Unit Tests
Answers 383 questions

SE Radio 557: Timothy Beamish on React and Next.js
Answers 383 questions

SE Radio 574: Chad Michel on Software as an Engineering Discipline
Answers 383 questions

SE-Radio Episode 355: Randy Shoup Scaling Technology and Organization
Answers 383 questions














