Today we are talking with Vignesh Baskaran, the CTO and co-founder of Hexo Labs, about teaching AI agents to improve themselves. Vignesh has been training neural networks since 2012, back when he was still called a data scientist.

Then he became an ML engineer and now an AI engineer, though he says the underlying work has never really changed. It's to figure out how to make a system behave the way you intend it to. He built the litigation search engine that Google itself became a customer of, and now he's chasing something new, agents that rewrite and retrain other agents without a human in the loop.

We dig into Sia, the meta-agent at the center of Hexo's research, and why improving an agent means touching both its harness and its actual model weights, not just one or the other. We talk about proxy evals for when you don't have much to ground truth. The Darwin-Gödel machine and why formal verification is too strict a bar for anything commercial.

How Hexo's work echoes DeepMind's Alpha lineage from AlphaGo to AlphaEvolve, and the spectrum from clearly verifiable to totally subjective tasks? Why VAE evals are quietly wrecking agent quality across the industry, and a great story about an agent that discovered a customer's own eval file was silently corrupted, something buried in hundreds of thousands of traces that no human would have caught.

Podden och tillhörande omslagsbild på den här sidan tillhör Software Huddle. Innehållet i podden är skapat av Software Huddle och inte av, eller tillsammans med, Poddtoppen.