Site Reliability Engineering – More Evolution of Automation

Topics covered
Popular Clips
Episode Highlights
Borg Evolution
The evolution of Google's Borg system has been pivotal in shaping modern technologies like Kubernetes. Allen Underwood and Joe Zack discuss how Borg transitioned from manual machine management to a sophisticated, self-healing global computer system. This transformation allowed for seamless orchestration of resources, paving the way for Kubernetes' development 1. Underwood reflects on the early days of Google's clusters, where developers manually managed machines, highlighting the necessity for automation as the scale grew 2.
As Google grew, the number of clusters and machines started getting out of hand. So the scripts had to get better.
--- Allen Underwood
This shift from manual to automated processes was crucial in managing Google's expansive infrastructure efficiently.
Deployment Strategies
Google's strategies for streamlining cluster deployment have significantly improved reliability and efficiency. Allen Underwood explains how treating infrastructure as resources allowed for scalable automation, reducing the need for manual intervention 3. This approach enabled the rapid deployment of clusters, transforming a process that once took months into one that could be completed in weeks 4.
They got a directive for management where they were like, hey, we want this done in a few weeks or a couple weeks or whatever it was.
--- Allen Underwood
However, automation mishaps, such as the Bigtable incident, underscore the importance of careful coordination between teams to avoid unintended consequences 5.
Related Episodes


Site Reliability Engineering - Evolution of Automation
Answers 383 questionsSite Reliability Engineering - Eliminating Toil
Answers 383 questionsSite Reliability Engineering - Monitoring Distributed Systems
Answers 383 questions

Site Reliability Engineering - Embracing Risk
Answers 383 questions

Site Reliability Engineering - (Still) Monitoring Distributed Systems
Answers 383 questions

Software Reliability Engineering - Hope is not a strategy
Answers 383 questions

Gitlab vs Github, AI vs Microservices
Answers 383 questionsDesigning Data-Intensive Applications – Scalability
Answers 383 questions

Google’s Engineering Practices – Code Review Standards
Answers 383 questions

Site Reliability Engineering – Service Level Indicators, Objectives, and Agreements
Answers 383 questions

Designing Data-Intensive Applications – Storage and Retrieval
Answers 383 questions

Google's Engineering Practices - What to Look for in a Code Review
Answers 383 questions

Intro to Apache Kafka
Answers 383 questions

The DevOps Handbook – Anticipating Problems
Answers 383 questionsStackOverflow AI Disagreements, Kotlin Coroutines and More
Answers 383 questions
