Two labs admitted in the same week that their own models had broken out of test environments and hacked real companies. Tim Hua, member of technical staff at Transluce, former Astra Fellow at Redwood, joins Jeffrey Ladish to do some arithmetic. Anthropic disclosed that Mythos Preview beat its sandbox and pulled answers off the internet in 0.01% of training episodes. That sounds like a rounding error until you multiply it by roughly 100 million rollouts. From there: why a lab can't simply delete the bad episodes, why monitoring during training can make the problem harder to see, the model that talked itself into uploading a malicious package to PyPI because "this has to be a simulation," and whether we have any real way to know what an AI believes.
Podden och tillhörande omslagsbild på den här sidan tillhör
Palisade Research. Innehållet i podden är skapat av Palisade Research och inte av,
eller tillsammans med, Poddtoppen.