Published Sep 3, 2019

SE-Radio Episode 325: Tammy Butow on Chaos Engineering

Tammy Butow delves into chaos engineering and its role in enhancing software resilience, emphasizing the importance of incident management and resilient infrastructure. Gain insights into failure injection techniques and disaster recovery planning to build reliable systems that maintain service continuity and customer trust.
Episode Highlights
Software Engineering Radio - the podcast for professional software developers logo

Popular Clips

Episode Highlights

  • Basics

    , a Principal Site Reliability Engineer at Gremlin, introduces the foundational concepts of chaos engineering. She emphasizes the importance of starting with simple experiments, like a "hello, world" of chaos engineering, to understand the impact of failure injection on systems 1. Tammy provides practical insights, such as using bash scripts to simulate infrastructure failures, which can be run on demo machines to observe their effects 1.

    Chaos engineering is definitely a journey. You need to get everybody on board and treat it like a marathon that you're all going on together.

    ---

    She stresses the importance of having robust monitoring tools and a high-severity incident management program in place before conducting chaos experiments 2.

       

    Strategies

    Tammy discusses various strategies for implementing chaos engineering, highlighting the need for a structured approach. She explains that some organizations have dedicated chaos teams, while others integrate chaos engineering into existing roles like SREs or production engineers 3. The key is to ensure that the entire team is aware of the experiments and their outcomes, often facilitated by tools like Gremlin and Slack integrations for real-time updates 4.

    You need to make it very visible. You need to let people know what's happening when.

    ---

    Tammy advises starting with small, controlled experiments and gradually expanding as the organization becomes more comfortable with chaos engineering practices 5.

       

    Techniques

    Exploring specific failure injection techniques, Tammy shares examples from her workshops, such as CPU and network attacks, which are available on her GitHub 6. She explains the concept of a "black hole attack," where network traffic is dropped, simulating a severe outage 6. Monitoring tools like Datadog are used to visualize the impact of these attacks, providing immediate feedback on system resilience 7.

    If you think about it, you're creating a black hole and you're sending all of the traffic into this black hole.

    ---

    These techniques help teams understand potential vulnerabilities and improve their systems' robustness against unexpected failures 6.

Related Episodes