Most engineers think reliability means avoiding outages. Vlad Leyberov learned the opposite lesson: sometimes you have to intentionally cause a 100% outage to fix the system faster.
Vlad is a Site Reliability Engineer (SRE) at Google, running systems that handle billions of requests per second. Before Google, he kept critical infrastructure running at Meta (billions of events a day) and Amazon (millions of Alexa devices).
In this conversation, we dig into cascading failures, incident responses, why consistency beats speed, how AI changes reliability engineering, and the philosophy behind running systems where downtime doesn't feel like an option.
Topics Discussed:
How cascading failures propagate unpredictably in distributed systems (like nature, not machines)
Incident responses: virtual panic rooms, on-call, paging procedures, and how to narrow down failure points
The Alexa incident: why dropping an entire DynamoDB table was the right call
Critical User Journeys (CUJ): measuring end-to-end customer experience vs individual SLOs
Career journey from the USSR to maritime academy to business degree in Australia to SRE at Amazon, Meta, and Google
Why consistency in API response times beats raw speed
How AI makes it dangerously easy to create complex systems with poorly understood interactions
Science fiction, the Borg as a distributed system, and the Three Body Problem trilogy
Hot takes on reliability: all software development is maintenance, overrated 9s, underrated global failure modes
Podden och tillhörande omslagsbild på den här sidan tillhör
Alexa Griffith. Innehållet i podden är skapat av Alexa Griffith och inte av,
eller tillsammans med, Poddtoppen.