LessWrong (30+ Karma)
Avsnitt

“Debate Training Reduces Reward Hacking in RLAIF” by zac_kenton, Jonah Brown-Cohen

Dela

Paper: Debate Training Reduces Reward Hacking in RLAIF

Linkpost for GDM Alignment blogpost

Work done by the GDM Amplified Oversight team (we're hiring).

TL;DR: When you RL against an LLM judge, the judge gets hacked i.e. fooled into incorrectly giving high reward; adding a debate opponent reduces this.

Many of the most impressive capabilities of current AI systems are produced by training on crisp tasks, like math and coding, where task success can be automatically verified. However, much of AI behavior that we actually care about is in some sense fuzzy, even for the most classical crisp tasks. For example, a coding agent should produce maintainable code, not just code that passes tests. More crucially, a coding agent should not learn to pass tests at all costs, especially by subverting the original intent of the user. However, using an LLM judge to provide reward for fuzzy tasks introduces its own issues. Convincing an LLM judge to give high rewards is often easier than solving the task correctly. So reward hacking becomes an even bigger problem. We show that training with debate, where two AIs argue against each to convince a judge, can mitigate reward hacking, potentially providing a hopeful [...]

The original text contained 1 footnote which was omitted from this narration.

---

First published:
August 19th, 2026

Source:
https://www.lesswrong.com/posts/BB8o7b8A4Aykeksvw/debate-training-reduces-reward-hacking-in-rlaif

---

Narrated by TYPE III AUDIO.

---

Images from the article:

Three line graphs comparing RLAIF-A, Debate-AB, and RLVR across training steps.Diagram comparing four training protocols: RLAIF-A, Debate-AB, Debate-ABA, RLVR.Diagram showing

Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

Podden och tillhörande omslagsbild på den här sidan tillhör LessWrong. Innehållet i podden är skapat av LessWrong och inte av, eller tillsammans med, Poddtoppen.