From alert to root cause in one minute - how PagerDuty built autonomous incident response on Amazon Bedrock, and the future of triage and trust.

Topics Include:

  • PagerDuty's agents must perform during 2am outages — stakes are high
  • Software shipping accelerated dramatically; production environments largely did not
  • A 9:30pm slowdown traced to a race condition solved two years earlier
  • The fix was documented — but the context wasn't at hand
  • PagerDuty Advance ships four agents: SRE, Scribe, Shift, Insights
  • Why four, not one? Focus and predictability in non-deterministic systems
  • Saurabh Shanbhag: Bedrock is far more than a model service
  • Zero data retention, PrivateLink, TLS — why enterprises pick Bedrock
  • Frontier models everywhere burns tokens; classify, route, distill, fine-tune
  • SRE agent triages alerts before you even join the call
  • One minute to root cause — context beat raw intelligence
  • Human surfaces versus machine surfaces: MCP and CLI move fastest
  • "The model eats the harness" — every upgrade invalidates foundational components
  • Feeding agents everything failed; compartmentalised investigation threads work better
  • New York Life's three stages of trust, and the seatbelt override that wasn't


Participants:


See how Amazon Web Services gives you the freedom to migrate, innovate, and scale your software company at https://aws.amazon.com/isv/

Podden och tillhörande omslagsbild på den här sidan tillhör AWS - Amazon Web Services. Innehållet i podden är skapat av AWS - Amazon Web Services och inte av, eller tillsammans med, Poddtoppen.