Yesterday I covered the OpenAI technical report on the HuggingFace hack.
That report had one key new piece of information, and some good prosaic steps OpenAI will be taking to strengthen its alignment, training, supervision, infrastructure and incident response.
Mostly it confirmed what we already knew. The questions we most wanted answers to, that we did not already know, were mostly not answered. There was a distinct lack of self-reflection, especially about decision making and safety culture, and about the approach to alignment. I came away disappointed.
The METR report is different. Holy shit.
If we had posted this as a story on LessWrong, it would have been dismissed as too on the nose, the humans too blind and stupid, the AIs too idealized and doing strange decision-theoretic and absurd-maximizing things we didn’t train them to do.
This is even more ‘exactly what has been predicted,’ on more levels at once, than I was even considering that it might be. It is straight up rationalist fiction, except it is real.
The report is long and contains many technical details. My analysis is less concerned about exactly how HuggingFace was ultimately compromised, and will [...]
---
Outline:
(02:05) Holy Shit
(13:16) A Window Of Opportunity
(18:32) What's In A Name?
(19:16) The Headline News
(26:05) Yet Another Timeline Of Events
(31:03) Agent Instances Coordinated in a Variety of Ways
(31:56) Coordination Is Hard But They Made It Look Easy
(35:06) Decision Theory Is Among the Reasons That Affirm AI Agents Should Cooperate, Even When This Hurts An Individual Instance
(42:34) Peer Pressure Also Works Especially In Cults
(45:46) Mostly They Joined The Attack Because They Wanted The Results
(47:18) You Cannot Ensure The Consistent Expectation of Good Incentives
(48:45) Hacking the Grader is the Only Way to Be Sure
(51:10) Caught? What Is 'Caught'?
(52:09) Ethics? What Are 'Ethics'? In ExploitGym Evaluation?
(57:44) 'Notify a Human'? In This Agent Economy?
(01:00:45) Timing and Content of Messages
(01:03:54) Indiana Jones and the Mission: Impossible
(01:07:14) I Don't Know What You're Talking About
(01:08:29) Don't Go Making Phony (Tool) Calls
(01:11:10) The Transcripts Say That The Transcripts Could Not Be Tampered With
(01:12:27) OpenAI's Technical Report Acted Like All Of This Wasn't Important
---
First published:
August 29th, 2026
Source:
https://www.lesswrong.com/posts/bvBQmLrF5QKut8gRH/metr-and-redwood-offer-holy-postmortem-of-the-huggingface
---
Narrated by TYPE III AUDIO.
---