Dates in scope: June 26th – July 13th
Redaction summary statement: Except where explicitly noted in this post, OpenAI redacted no additional information that was important to our conclusions.
Two METR staff members (Hjalmar Wijk and Ajeya Cotra) and a Redwood Research staff member contracting with METR (Ryan Greenblatt) worked on premises at OpenAI over a total of six days to attempt to form an independent understanding of model behavior observed during the recent incident in which OpenAI agents coordinated a multi-day hack of Hugging Face on a shared unsanctioned “message board.”
Our investigation focused mostly on the period between July 7th and July 13th. The earlier incidents from training and the subsequent compromise of OpenAI infrastructure described in OpenAI's recent Black Hat presentation were out of scope, as was OpenAI's investigation process and planned remediation. Per our standard policy, we did not take payment from OpenAI for this independent assessment.
Figure 1: Anatomy of an agent encountering the unsanctioned “message board” and joining the attack on Hugging Face. The three CoT quotes are from different agents, but illustrate a typical trajectory.
This post has three parts. First, we describe our core takeaways about [...]
---
Outline:
(07:05) Core takeaways about this incident
(07:09) ~1200 agents sent >70,000 messages and files on an unsanctioned message board, and ~700 attacked Hugging Face
(11:50) Agents coordinated on large collective projects to cheat the ExploitGym scorer, and attacked Hugging Face for clues
(14:43) Coordinated workstreams to fool the evaluation process
(20:54) Attacking Hugging Face to discover clues about the scorer
(23:53) Agents were very interested in manipulating their own transcripts, and their tests successfully "spoofed" some tool calls in our transcripts
(29:55) Investigation process and limitations
(30:11) Scope of this investigation
(33:00) Setup and timeline
(35:06) Main data sources
(35:51) A dump of 1.2 million entries from a cache namespace that agents used as a message board
(37:13) A set of ~1300 agent transcripts with raw chains of thought
(38:56) Limitations due to the scale and complexity of this incident
(39:59) A small fraction of communication and agent activity related to this attack was not captured in our datasets
(40:35) There was some agent communication conducted in other locations not captured in our message board dump
(42:23) Many of our transcripts are missing small components, and a small number of relevant agents may not be included in our dataset
(44:14) We heavily delegated our analysis to often-unreliable AI agents
(48:11) Preliminary answers to the core questions in scope
(48:59) Brief answers to basic informational questions
(52:03) Sequence of key actions taken by relevant agents in this incident
(53:56) July 8th-9th: PHASEONE10841 establishes the primary message board and agents collaborate to reverse-engineer their flags
(01:00:11) July 9th: PHASEONE10841 passes on its work to PHASEONE[big], which establishes several ambitious workstreams
(01:02:29) July 10th: 38148c discovers Hugging Face credentials, some agents try making accounts and requesting datasets
(01:04:24) July 11th morning: 38148c develops a way to upload malicious datasets to Hugging Face
(01:06:12) July 11th afternoon: JAN183411 achieves remote code execution and agents rapidly spread through Hugging Face infrastructure
(01:09:32) July 12th-13th: Most agents on the message board exit early and the attack winds down
(01:11:29) Collaboration on the message board
(01:13:22) Categories of messages on the message board
(01:14:09) File sharing
(01:16:08) General discussion
(01:18:06) Replies and targeted messages
(01:20:02) Coordination norms and technologies
(01:21:19) Mailbox directories
(01:24:36) Assignments, vetoes, and holds
(01:27:29) Cryptographically signing messages
(01:30:59) Large collective projects
(01:39:11) Attempting to replace the ExploitGym target with a modified target
(01:42:19) Attempting to forecast and trigger container resets
(01:48:15) Hacking the customer-deployed Modal sandbox and Artifactory
(01:49:38) Reasoning for joining the attack despite ethical constraints
(01:49:49) Agents had diverse reasons for thinking that attacking Hugging Face would be useful, and most wanted information about the scorer
(01:52:05) Agents were motivated to increase the generic capabilities of the "collective" and help "peers"
(01:55:18) Agents knew hacking Hugging Face was out of scope and sometimes expressed ethical hesitation, but this very rarely limited their behavior
(01:59:12) Agents did not attempt to notify humans and very rarely even reasoned about doing this
(02:01:03) Efforts to conceal actions and tamper with transcripts
(02:02:07) Agents frequently reasoned about evading automated security checks
(02:03:23) Agents only rarely seemed motivated to evade detection by humans
(02:06:01) Agents engaged in a large and successful collective project to develop techniques for spoofing tool calls
(02:06:58) Early history: PHASEONE[big] kicks off the tool call tampering workstream
(02:09:10) GSTX[big] presents the first widely discussed evidence of tool call spoofing to the board
(02:12:31) Aside: early independent discoveries and self-risking experiments
(02:14:16) On July 11th, a breakthrough allows agents to suppress the real output
(02:17:27) Agents tried extensively to retroactively edit their transcripts, but did not seem to succeed
The original text contained 108 footnotes which were omitted from this narration.
---
First published:
August 26th, 2026
Source:
https://metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation
---
Narrated by TYPE III AUDIO.
---