In our last episode, Stanford NLP researcher Dr. Moritz Sudhof showed that 79% of AI product failures are invisible — they don't fire alerts, don't surface in telemetry, and don't get flagged by users, and they quietly erode trust and accelerate churn. This episode is the operational follow-up: the team that built a method for catching exactly that class of failure.
Microsoft's Copilot UX research team spent a year running evaluations on real conversations — every archetype, every industry, every use case — and found that more than half of their quality failures weren't in any eval they were running. At the world's most widely deployed enterprise AI product, with sophisticated engineering and testing infrastructure, standard evals were still missing the majority of what users actually experienced as failure.
That finding isn't limited to Copilot's scale. It's a structural gap in how the industry evaluates AI quality — and if you're running automated evals and calling that sufficient, the gap in your own product is almost certainly larger than you know.
In this episode we cover:
- Token usage and adoption tell you if your AI is being used — not whether it's actually working for anyone.
- Users bring real prompts, test one model fully, then compare — that's what produces honest signal at scale.
- More than half of Copilot's loss patterns — user-driven gaps in model behavior — weren't in any existing eval.
- LLM judges get you to baseline quality. Users catch what automated testing structurally cannot.
- The flywheel: UXR evals → loss pattern taxonomy → log inspection → prompt changes → retention gains.
- This team started with 10 users and one comparative question. Signal strong enough to scale to an entire org.
"More than half of the loss patterns that we've detected were not things that we were measuring in our evals." — Wendy Wang
About the team: This work was developed by Christopher Monnier, Wendy Wang, and Chuck Kwong, UX researchers on the Microsoft Copilot team. Together, they built and continue to refine an interactive evaluation method that brings real user tasks, side-by-side product comparisons, quantitative results, and qualitative feedback into one process. The team uses this work to identify where Copilot succeeds and where people run into issues, understand the reasons behind user preferences, and turn the findings into clear opportunities for product and prompt teams to improve the experience.
..................
If you found this episode useful, please like, share, and send it to anyone on your team who'd find it helpful.
We built https://productimpactpod.com to be your AI product insights and strategic playbook hub. Check it out.
Hosted by:
➜ Arpy Dragffy Guerrero — https://www.linkedin.com/in/adragffy/
➜ Brittany Hobbs — https://www.linkedin.com/in/brittanyhobbs/
Go to Substack to get AI strategy frameworks, news, and jobs: https://productimpactpod.substack.com
This episode was brought to you by:
➜ PH1 (https://ph1.ca) — a strategy & research consultancy specialized in pinpointing how to best leverage AI and improve the impact of your AI product
➜ AI Value Acceleration (https://aivalueacceleration.com) — The consultancy specialising in enterprise value creation. Make sure that your spending doesn't go to waste. Find out exactly where the value creation of adopting AI products stalls.