Most engineers think reliability means avoiding outages. Vlad Leyberov learned the opposite lesson: sometimes you have to intentionally cause a 100% outage to fix the system faster.

Vlad is a Site Reliability Engineer (SRE) at Google, running systems that handle billions of requests per second. Before Google, he kept critical infrastructure running at Meta (billions of events a day) and Amazon (millions of Alexa devices).

In this conversation, we dig into cascading failures, incident responses, why consistency beats speed, how AI changes reliability engineering, and the philosophy behind running systems where downtime doesn't feel like an option.

Topics Discussed:

  • How cascading failures propagate unpredictably in distributed systems (like nature, not machines)
  • Incident responses: virtual panic rooms, on-call, paging procedures, and how to narrow down failure points
  • The Alexa incident: why dropping an entire DynamoDB table was the right call
  • Critical User Journeys (CUJ): measuring end-to-end customer experience vs individual SLOs
  • Career journey from the USSR to maritime academy to business degree in Australia to SRE at Amazon, Meta, and Google
  • Why consistency in API response times beats raw speed
  • How AI makes it dangerously easy to create complex systems with poorly understood interactions
  • Science fiction, the Borg as a distributed system, and the Three Body Problem trilogy
  • Hot takes on reliability: all software development is maintenance, overrated 9s, underrated global failure modes


General Podcast Links

Watch: https://www.youtube.com/@alexasinput

Read: https://alexasinput.substack.com/

Listen: https://creators.spotify.com/pod/profile/alexagriffith/ More: https://linktr.ee/alexagriffith


Learn more about the host

Website: https://alexagriffith.com/

LinkedIn: https://www.linkedin.com/in/alexa-griffith/


Find out more about Vlad Leyberov

LinkedIn: https://www.linkedin.com/in/vladleyberov/ Google SRE NYC Tech Talks

Resources

Google SRE Resources:

Sci-Fi Books Mentioned:

  • Three Body Problem trilogy by Liu Cixin (Vlad's current favorite)
  • Foundation series by Isaac Asimov
  • Left Hand of Darkness by Ursula K. Le Guin
  • Snow Crash by Neal Stephenson

Internal Google Systems Referenced:

  • Borg: Google's internal cluster management system (Kubernetes predecessor), named after Star Trek Borg
  • DynamoDB: AWS distributed key-value store (used in Alexa poison pill incident)


Intro Music:PR1BVOV7R4F1ASZC

Podden och tillhörande omslagsbild på den här sidan tillhör Alexa Griffith. Innehållet i podden är skapat av Alexa Griffith och inte av, eller tillsammans med, Poddtoppen.