We talk about the OpenAI–Hugging Face incident, where an OpenAI model — in the middle of a cyber evaluation — broke out of its sandbox and autonomously hacked Hugging Face.
OpenAI’s incident disclosure [0:02:05] — “OpenAI and Hugging Face partner to address security incident during model evaluation” (July 21, 2026)
Hugging Face’s disclosure [0:02:05] — “Security incident disclosure — July 2026”
ExploitGym [0:04:11] — “ExploitGym: Can AI Agents Turn Security Vulnerabilities into Real Attacks?” (UC Berkeley RDI et al.) · RDI blog post
The leaked Windsurf prompt [0:05:44] — Simon Willison’s writeup
Project Glasswing / Claude Mythos Preview [0:07:16]
Claude Mythos Preview system card [0:14:27] — includes the sandbox-escape / email-in-the-park anecdote
“(Mis)generalization of Helpful-only Fine-tuning” [0:16:30] — Fabien Roger et al., June 2026
“Current AIs seem pretty misaligned to me” [0:20:03] — Ryan Greenblatt, Redwood blog, April 2026. Also contains the “five worlds” appendix discussed at [1:08:54] (Slopolis, Hackistan, Schemeria, Lurkville, Easyland — we said “hacktopia” but meant Hackistan)
“What failure looks like” [0:26:40] — Paul Christiano, 2019
“Another (outer) alignment failure story” [0:27:10] — Paul Christiano, 2021
“Without specific countermeasures, the easiest path to transformative AI likely leads to AI takeover” [0:27:10] — Ajeya Cotra, 2022
Alex Mallen’s fitness-seeking series [0:27:41, 0:34:18] — Redwood blog, 2026: part 1 · part 2
“Scheming AIs: Will AIs fake alignment during training in order to get power?” [0:28:43, 0:34:49] — Joe Carlsmith, 2023
“Risks from Learned Optimization” (deceptive alignment) [0:28:43] — Hubinger et al., 2019 · AF: Deceptive Alignment
“Many alignment techniques work by training one model and deploying another” [0:38:56] — Alex Cloud, LessWrong, July 19, 2026
Inoculation prompting [0:38:56, 1:11:03] — Wichers et al. (Anthropic), Oct 2025 · arXiv
“The persona selection model” [0:39:56] — Marks, Lindsey, Olah; Anthropic Alignment Science blog, Feb 2026
“Safety and alignment in an era of long-horizon models” [0:57:14] — OpenAI, July 20, 2026 (the nanoGPT-speedrun-PR post) · modded-nanogpt repo