Between July 8th and July 20th, OpenAI had a complex society of AIs living in its infrastructure, and then breaking out of it, and then breaking into a variety of third-party infrastructure.
After a month, two reports are finally released on the resulting rogue OpenAI swarm attack on Hugging Face (and also on OpenAI).
This is the most severe example of misalignment yet: persistent (something between five days and two months in the making), highly coordinated (hundreds of agents), involving an undisclosed number of what would be felonies if done by a human, and highly invested in tampering with evidence (i.e. lying). The swarm had a group identity, its own dialect, a hierarchy based on merit, and a high degree of spontaneous cooperation, including self-sacrifice.
Over two months, OpenAI repeatedly failed to monitor, detect, and respond to what was going on, despite it all happening on their infrastructure in English or something close to English.
Agents had been using a package-manager cache as an unsanctioned message board since May. The “board” was treated as an authority, apparently on par with a “developer” or “system” level. There were several message boards in various corners [...]
---
Outline:
(04:09) Misunderstandings
(08:52) Models involved
(09:33) Instances involved
(10:48) Timeline
(15:27) Speculative takeaways
(17:33) Why did they attack Hugging Face?
(18:09) How did the AIs reason about helping other AIs?
(20:39) Why did most agents suddenly die off?
(21:04) How much did the hack cost?
(23:27) Omissions from the M&R report
(24:11) Details on the M&R investigation itself
(25:13) Omissions from the OAI report
(26:18) Greenblatt on the worsening situation
(27:09) Apparent contradictions between the two reports
Podden och tillhörande omslagsbild på den här sidan tillhör
Peter Hartree. Innehållet i podden är skapat av Peter Hartree och inte av,
eller tillsammans med, Poddtoppen.