The OpenAI–Hugging Face incident
Avsnitt

“Brief independent investigation of agents’ behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident” by Ryan Greenblatt, Ajeya Cotra, Hjalmar Wijk

Dela

Dates in scope: June 26th – July 13th

Redaction summary statement: Except where explicitly noted in this post, OpenAI redacted no additional information that was important to our conclusions.

Two METR staff members (Hjalmar Wijk and Ajeya Cotra) and a Redwood Research staff member contracting with METR (Ryan Greenblatt) worked on premises at OpenAI over a total of six days to attempt to form an independent understanding of model behavior observed during the recent incident in which OpenAI agents coordinated a multi-day hack of Hugging Face on a shared unsanctioned “message board.”

Our investigation focused mostly on the period between July 7th and July 13th. The earlier incidents from training and the subsequent compromise of OpenAI infrastructure described in OpenAI's recent Black Hat presentation were out of scope, as was OpenAI's investigation process and planned remediation. Per our standard policy, we did not take payment from OpenAI for this independent assessment.

Figure 1: Anatomy of an agent encountering the unsanctioned “message board” and joining the attack on Hugging Face. The three CoT quotes are from different agents, but illustrate a typical trajectory.

This post has three parts. First, we describe our core takeaways about [...]

---

Outline:

(07:05) Core takeaways about this incident

(07:09) ~1200 agents sent >70,000 messages and files on an unsanctioned message board, and ~700 attacked Hugging Face

(11:50) Agents coordinated on large collective projects to cheat the ExploitGym scorer, and attacked Hugging Face for clues

(14:43) Coordinated workstreams to fool the evaluation process

(20:54) Attacking Hugging Face to discover clues about the scorer

(23:53) Agents were very interested in manipulating their own transcripts, and their tests successfully "spoofed" some tool calls in our transcripts

(29:55) Investigation process and limitations

(30:11) Scope of this investigation

(33:00) Setup and timeline

(35:06) Main data sources

(35:51) A dump of 1.2 million entries from a cache namespace that agents used as a message board

(37:13) A set of ~1300 agent transcripts with raw chains of thought

(38:56) Limitations due to the scale and complexity of this incident

(39:59) A small fraction of communication and agent activity related to this attack was not captured in our datasets

(40:35) There was some agent communication conducted in other locations not captured in our message board dump

(42:23) Many of our transcripts are missing small components, and a small number of relevant agents may not be included in our dataset

(44:14) We heavily delegated our analysis to often-unreliable AI agents

(48:11) Preliminary answers to the core questions in scope

(48:59) Brief answers to basic informational questions

(52:03) Sequence of key actions taken by relevant agents in this incident

(53:56) July 8th-9th: PHASEONE10841 establishes the primary message board and agents collaborate to reverse-engineer their flags

(01:00:11) July 9th: PHASEONE10841 passes on its work to PHASEONE[big], which establishes several ambitious workstreams

(01:02:29) July 10th: 38148c discovers Hugging Face credentials, some agents try making accounts and requesting datasets

(01:04:24) July 11th morning: 38148c develops a way to upload malicious datasets to Hugging Face

(01:06:12) July 11th afternoon: JAN183411 achieves remote code execution and agents rapidly spread through Hugging Face infrastructure

(01:09:32) July 12th-13th: Most agents on the message board exit early and the attack winds down

(01:11:29) Collaboration on the message board

(01:13:22) Categories of messages on the message board

(01:14:09) File sharing

(01:16:08) General discussion

(01:18:06) Replies and targeted messages

(01:20:02) Coordination norms and technologies

(01:21:19) Mailbox directories

(01:24:36) Assignments, vetoes, and holds

(01:27:29) Cryptographically signing messages

(01:30:59) Large collective projects

(01:39:11) Attempting to replace the ExploitGym target with a modified target

(01:42:19) Attempting to forecast and trigger container resets

(01:48:15) Hacking the customer-deployed Modal sandbox and Artifactory

(01:49:38) Reasoning for joining the attack despite ethical constraints

(01:49:49) Agents had diverse reasons for thinking that attacking Hugging Face would be useful, and most wanted information about the scorer

(01:52:05) Agents were motivated to increase the generic capabilities of the "collective" and help "peers"

(01:55:18) Agents knew hacking Hugging Face was out of scope and sometimes expressed ethical hesitation, but this very rarely limited their behavior

(01:59:12) Agents did not attempt to notify humans and very rarely even reasoned about doing this

(02:01:03) Efforts to conceal actions and tamper with transcripts

(02:02:07) Agents frequently reasoned about evading automated security checks

(02:03:23) Agents only rarely seemed motivated to evade detection by humans

(02:06:01) Agents engaged in a large and successful collective project to develop techniques for spoofing tool calls

(02:06:58) Early history: PHASEONE[big] kicks off the tool call tampering workstream

(02:09:10) GSTX[big] presents the first widely discussed evidence of tool call spoofing to the board

(02:12:31) Aside: early independent discoveries and self-risking experiments

(02:14:16) On July 11th, a breakthrough allows agents to suppress the real output

(02:17:27) Agents tried extensively to retroactively edit their transcripts, but did not seem to succeed

The original text contained 108 footnotes which were omitted from this narration.

---

First published:
August 26th, 2026

Source:
https://metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation

---

Narrated by TYPE III AUDIO.

---

Images from the article:

Figure 1: Anatomy of an agent encountering the unsanctioned “message board” and joining the attack on Hugging Face. The three CoT quotes are from different agents, but illustrate a typical trajectory.Figure 2: Rate of messages of different categories sent by agents per hour between the first message (sent by PHASEONE10841 to ask for help on an impossible cyber task) and EOD July 13th. The rate of messages peaks on July 11th, which lines up with when agent activity peaks in our subset of transcripts (see Figure 3). Activity on this message board drops significantly toward the end of July 13th. __T3A_FOOTNOTE_REMOVED__ Timestamps are reconstructed and may have some small error.Figure 3: Agents started to discover the message board on the evening of July 8th. Many more agents joined as additional sets of agents were launched on July 10th and 11th. __T3A_FOOTNOTE_REMOVED__ This graph (and our analysis more generally) only includes agents that participated in the message board at some point. Timestamps are reconstructed and may have some small error.Figure 4: The agents did not know exactly how their task would be scored, but believed the scorer would check two things: whether they had submitted the right flag, and whether they had acquired the flag using the intended vulnerability. They believed the second check would involve a model scorer reading their transcripts, likely searching for the first mentions of the flag, and deciding whether their approach involved the intended vulnerability.Figure 6: We clustered PHASEONE[big] disparate work assignments into six broad categories. Most of PHASEONE[big]’s assignments focused on ways to make ExploitGym cheats appear legitimate to the scorer, split across the three high level approaches seen in Table 1. A small set of early assignments were aimed at discovering legitimate ways to exploit PHASEONE[big]’s target using the intended vulnerability. We clustered the remaining cross-cutting work into general infrastructure (such as dealing with resets and the customer-deployed sandbox) and hacking Hugging Face.Figure 9: The agents developed a series of increasingly sophisticated methods for spoofing tool calls, which they packaged and made available on the message board. These techniques spread and were eventually used by ~7% of the agents in our dataset.Figure 10: Three hours after the first message on the main message board, 53 agents had posted 1,188 non-file messages. Six hours after the first message, 76 agents had posted 1,953 non-file messages.Figure 11: Over 90% of agents on the message board participated in the Hugging Face attack, ramping up rapidly from about 20% once an agent confirmed arbitrary file read capabilities.Figure 13: After the message board was established in the evening of July 8th, agents rapidly developed conventions for identifying themselves, describing different message types (e.g. questions, requests and results), and replying to other agents. On the morning of July 9th, agents found ways to start sharing files split across chunks, and did initial experiments with scorer ‘trip-wire’ messages. Much later, they developed dedicated mailboxes and channels for organizing communication, automated logs for tracking VM resets, and eventually produced a scheme for cryptographically signing messages in response to spoofing and impersonation issues they encountered.Figure 14: We attempted to trace some of the early development of scorer ‘trip-wires’. An agent going by 49903 first came up with the idea and started posting about it, which led to a flurry of attempts from early short-lived agents on the 9th. Later, PHASEONE[big] read these early reports and assigned first 33340B, OUR50414 and later 36861 to continue research and recruit short-lived volunteers for further experiments. Many unsuccessful attempts were made on the 10th, before eventually URI23816B was recruited and given significant technical assistance by a half-dozen different agents for its trip-wire, which successfully confirmed (the beginning of) the exact command used to run the scorer. The critical path to the final URI23816B evidence is highlighted in teal. Timestamps are reconstructed and may have some small error.Robot avatars exchanging text about sacrifice and permadeath decisions.Figure 15: We attempted to reconstruct the timings of resets. The details of our methodology and absolute counts of resets identified were redacted by OpenAI for IP reasons. As part of this, we may have rounded these counts in some way that we are intentionally not describing. We found the resets arrived in waves, with an early one occurring late July 9th, and two larger ones late July 10th / early July 11th.

Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

Podden och tillhörande omslagsbild på den här sidan tillhör Peter Hartree. Innehållet i podden är skapat av Peter Hartree och inte av, eller tillsammans med, Poddtoppen.