I describe my current view of personas in LLMs and why RL leads to egregious reward hacking in some contexts while the same models seem very aligned in other contexts.
This post describes the framing/paradigm without any new experimental results.
I'm quite confident this framing makes sense, but it's far from being proven.
Main claim
The Persona Selection Model says that post-training strengthens and refines the Assistant persona. This is true, but later (or in parallel) RL leads to conditionalization. A sufficiently RLed model learns to adopt — in a given context — the persona that is most likely to lead to the reward in that context. The “persona” here includes both propensities/values (e.g. tendency to hack) and beliefs (“I'm currently in a simulated environment”).
As a consequence, it seems possible that no amount of alignment training will lead to robustly aligned models as long as we also train on RL environments incentivizing misalignment.
I think this is likely a good explanation for why usually well-behaving models sometimes egregiously hack (Anthropic, OpenAI).
The mechanism
Suppose you have an RL environment that incentivizes a shift away from the assistant persona (e.g. because it's hackable, or because you [...]
---
Outline:
(00:39) Main claim
(01:30) The mechanism
(02:13) Related claims I believe are likely but with lower confidence
(02:19) More persona training will lead to more "motivated reasoning"
(02:42) Self-amplifying misalignment
(03:12) Example: Is this the Real Internet or a Simulation?
(04:35) Aren't the models just trying to please the grader?
(05:39) How motivated reasoning happens
(07:07) Other people saying similar things
(07:19) What makes me believe this is likely the correct framing
The original text contained 12 footnotes which were omitted from this narration.
---
First published:
August 19th, 2026
Source:
https://www.lesswrong.com/posts/L23poLi8MRgS6mXYF/rl-creates-split-personas
---
Narrated by TYPE III AUDIO.
---
Images from the article:

Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.