This week's biggest AI upgrade wasn't a model. It was someone reviewing code differently.

Grok, Gemini, and GPT all shipped new versions in the same seven days. I skimmed the benchmarks. Forgot most of them within the hour.

What stuck was smaller. Fred Schott released Flue 2, a hook-based agent framework. Simon Willison wrote about how he reviews agent-written code now. Not the same checks he runs on a human's pull request, not the same pace, not the same default trust.

Two different people, two different projects. Same discovery underneath.

Neither of them upgraded a model. They upgraded what happens around it.

I keep watching teams do the opposite. New model, same shell script duct-taped to it. Same review habits built for commits that land twice a day, now pointed at code that shows up ten times faster and doesn't get tired or embarrassed when it's wrong.

That's not a tooling gap. That's a process nobody rebuilt.

If your agent produces a bad tool call right now, what actually catches it? A human watching the terminal? A test suite that was never written for this? Or nothing?

More on this in this week's episode of The Human in the Loop.

#TheHumanInTheLoop #AIAgents #SoftwareEngineering

Podden och tillhörande omslagsbild på den här sidan tillhör Enrique Cordero. Innehållet i podden är skapat av Enrique Cordero och inte av, eller tillsammans med, Poddtoppen.