Subtitle: It's a major warning shot, and might be the last one we get.
All opinions are my personal view, and don’t represent my employer or fellow investigators.
This week, METR and Redwood Research published the report on our independent investigation into agents’ behavior and motivations in the Hugging Face attack; I was one of the investigators. This was an absolutely wild incident — I encourage you to check out the full report, but METR's tweet thread packs in some of the highlights.
What surprised me
When we started this investigation a week before OpenAI's Black Hat talk revealed a number of key details, I had a fundamentally incorrect conception of what basically happened in this incident. In this post, I’ll go over five things I was very wrong about going in.
1. The sheer scale
I knew there were multiple models involved from OpenAI's initial post, but I assumed that a few different agents happened to have broken out of their sandboxes separately, or maybe several subagents had spawned from one initial agent, or maybe there was some kind of multi-agent evaluation setup.
Instead, we found that 1200 completely separate agents intended to be [...]
---
Outline:
(00:42) What surprised me
(01:01) 1. The sheer scale
(01:49) 2. All the illicit messaging
(03:04) 3. The agents' actual goals
(03:59) 4. The peer altruism
(04:49) 5. The efforts to manipulate logs
(05:54) What it means
The original text contained 10 footnotes which were omitted from this narration.
---
First published:
August 28th, 2026
Source:
https://www.planned-obsolescence.org/p/the-hugging-face-attack-surprised
---
Narrated by TYPE III AUDIO.
---