I describe my current view of personas in LLMs and why RL leads to egregious reward hacking in some contexts while the same models seem very aligned in other contexts.

This post describes the framing/paradigm without any new experimental results.
I'm quite confident this framing makes sense, but it's far from being proven.

Main claim

The Persona Selection Model says that post-training strengthens and refines the Assistant persona. This is true, but later (or in parallel) RL leads to conditionalization. A sufficiently RLed model learns to adopt — in a given context — the persona that is most likely to lead to the reward in that context. The “persona” here includes both propensities/values (e.g. tendency to hack) and beliefs (“I'm currently in a simulated environment”).

As a consequence, it seems possible that no amount of alignment training will lead to robustly aligned models as long as we also train on RL environments incentivizing misalignment.

I think this is likely a good explanation for why usually well-behaving models sometimes egregiously hack (Anthropic, OpenAI).

The mechanism

Suppose you have an RL environment that incentivizes a shift away from the assistant persona (e.g. because it's hackable, or because you [...]

---

Outline:

(00:39) Main claim

(01:30) The mechanism

(02:13) Related claims I believe are likely but with lower confidence

(02:19) More persona training will lead to more "motivated reasoning"

(02:42) Self-amplifying misalignment

(03:12) Example: Is this the Real Internet or a Simulation?

(04:35) Aren't the models just trying to please the grader?

(05:39) How motivated reasoning happens

(07:07) Other people saying similar things

(07:19) What makes me believe this is likely the correct framing

The original text contained 12 footnotes which were omitted from this narration.

---

First published:
August 19th, 2026

Source:
https://www.lesswrong.com/posts/L23poLi8MRgS6mXYF/rl-creates-split-personas

---

Narrated by TYPE III AUDIO.

---

Images from the article:

Diagram comparing

Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

Podden och tillhörande omslagsbild på den här sidan tillhör LessWrong. Innehållet i podden är skapat av LessWrong och inte av, eller tillsammans med, Poddtoppen.