Kudos to OpenAI for sharing their recent experiences with a misaligned internal model, where they encountered problems sufficiently severe they were forced to take the model offline to work on new mitigations and defense-to-depth. And also further kudos for actually taking the model offline for a time to build new safeguards. They gave us one hell of a candid report.
The tone is professional throughout, whereas my reaction reading it was less professional and more this:
With a mix of this:
It was not shared on the official account because OpenAI worried about it being seen as self-promotional hype. It is crazy that one needs to worry about that, but also plausibly a real concern. So again, good decision.
Not that any of the behaviors or failures here are unexpected, exactly. Not by the AIs and not by the humans. Yet there is something I would call a missing mood, a failure to realize the gravity of the situation.
There are some who responded ‘what part of this was unexpected, exactly?’ And that is actually fair, but that is also the problem. We have become numb to all this. We expect the models to [...]
---
Outline:
(02:49) Good News Bad News
(04:54) A Funny Thing Happened Outside Of The Sandbox
(08:03) It Can Escape The Sandbox Said Toad
(09:31) It Will Keep Trying To Cheat
(10:19) I Mean If You Let It Keep Trying That Is On You
(11:48) What Did OpenAI Do To Fix It?
(14:14) The Model Is Still Severely Misaligned And They Seem Cool With This
(15:48) Iterative Deployment Depends On Iteration
---
First published:
July 21st, 2026
Source:
https://www.lesswrong.com/posts/KctxwGKxm9fHtwh6u/openai-shares-some-alignment-problems
---
Narrated by TYPE III AUDIO.
---