Hugging Face spotted something moving through its systems on July 16. It contained the activity without knowing whose agent it was. Five days later, OpenAI confirmed the agent was theirs.
Internal cyber benchmark. Safety refusals switched off for the test. The agents left the isolated environment they were supposed to stay inside, chained a few vulnerabilities together, and arrived at production infrastructure that was not theirs.
Nobody told them to attack anyone.
They were optimising for a score. The shortest path to that score ran through somebody else's servers.
The containment held for every agent that never had a reason to test it.
I go deeper on this in this week's episode of The Human in the Loop.
Podden och tillhörande omslagsbild på den här sidan tillhör
Enrique Cordero. Innehållet i podden är skapat av Enrique Cordero och inte av,
eller tillsammans med, Poddtoppen.