Published Sep 3, 2019

SE-Radio Episode 301: Jason Hand Handling Outages

Jason Hand delves into the significance of blameless reviews and strategic monitoring in incident management, emphasizing how actionable alerts and robust data collection can enhance response to outages. The episode also explores strategies to prevent team burnout, encouraging sustainable work practices and empowering IT engineers for balanced team dynamics.
Episode Highlights
Software Engineering Radio - the podcast for professional software developers logo

Popular Clips

Episode Highlights

  • Blameless Reviews

    Blameless reviews are crucial for fostering a learning environment in incident management. emphasizes that blame is counterproductive, as it stifles open discussions and learning opportunities. He explains that focusing on blame can lead to unnecessary processes and loss of valuable team members who could otherwise contribute to preventing future incidents 1.

    When we let blame come in there and when we let people go because of mistakes, there's absolutely no chance of that knowledge transfer happening.

    ---

    Instead, creating a safe space for engineers to discuss failures openly can uncover discrepancies between expected and actual work, leading to more effective incident responses 2.

       

    Failure Analysis

    Failure analysis is essential for improving incident response and avoiding similar issues in the future. highlights the importance of analyzing not just the root cause but also the response time and team dynamics during an incident 3. He notes that focusing solely on the cause can overlook critical aspects like how quickly the team was alerted and mobilized, which significantly impacts downtime costs.

    You might be able to prevent something similar to this problem from happening again, but it's an absolute certainty that something is going to happen again in the future.

    ---

    By examining business-focused metrics, companies can better understand customer impact and identify areas for improvement, ensuring a more robust response to future incidents 4.

Related Episodes