This episode examines the "Reversal Curse," a 2023 finding from researchers at Vanderbilt, UK AI Safety Institute, Apollo Research, NYU, Sussex, and Oxford showing that autoregressive language models trained on "A is B" statements fail to infer "B is A," even though the two are logically equivalent. Using the example of Valentina Tereshkova, the discussion shows how a model finetuned on "Tereshkova was the first woman in space" answers forward questions perfectly but performs at chance when the question is reversed, and extends this to real-world GPT-4 results where "who is Tom Cruise's mother" scores 79% accuracy versus just 33% for the reverse query about Mary Lee Pfeiffer. The hosts unpack why this happens mechanically — next-token prediction bakes facts into weights as one-directional associations rather than symmetric relations like a knowledge-graph edge — and contrast this with in-context learning, where the same reversal works flawlessly, pinpointing the failure specifically to generalization from gradient-based training rather than a reasoning limitation. A debate over whether this is just a data-coverage gap versus a deeper meta-learning failure leads into the researchers' attempted fix: training on facts stated in both directions to see if models can pick up the general pattern. It's a compelling listen for anyone curious about the hidden asymmetries in how LLMs actually store knowledge versus how humans intuitively reason about it.
Sources:
1. The Reversal Curse: LLMs trained on "A is B" fail to learn "B is A" — Lukas Berglund, Meg Tong, Max Kaufmann, Mikita Balesni, Asa Cooper Stickland, Tomasz Korbak, Owain Evans, 2023
http://arxiv.org/abs/2309.12288
2. Locating and Editing Factual Associations in GPT (ROME) — Kevin Meng, David Bau, Alex Andonian, Yonatan Belinkov, 2022
https://scholar.google.com/scholar?q=Locating+and+Editing+Factual+Associations+in+GPT+%28ROME%29
3. Language Models as Knowledge Bases? — Fabio Petroni, Tim Rocktäschel, Patrick Lewis, Anton Bakhtin, Yuxiang Wu, Alexander H. Miller, Sebastian Riedel, 2019
https://scholar.google.com/scholar?q=Language+Models+as+Knowledge+Bases%3F
4. Reverse Training to Nurse the Reversal Curse — Olga Golovneva, Zeyuan Allen-Zhu, Jason Weston, Sainbayar Sukhbaatar (Meta AI), 2024
https://scholar.google.com/scholar?q=Reverse+Training+to+Nurse+the+Reversal+Curse
5. Physics of Language Models: Part 3.2, Knowledge Manipulation — Zeyuan Allen-Zhu, Yuanzhi Li, 2024
https://scholar.google.com/scholar?q=Physics+of+Language+Models%3A+Part+3.2%2C+Knowledge+Manipulation
6. Studying Large Language Model Generalization with Influence Functions — Roger Grosse, Juhan Bae, Cem Anil, et al., 2023
https://scholar.google.com/scholar?q=Studying+Large+Language+Model+Generalization+with+Influence+Functions
7. Transformer Feed-Forward Layers Are Key-Value Memories — Mor Geva, Roei Schuster, Jonathan Berant, Omer Levy, 2021
https://scholar.google.com/scholar?q=Transformer+Feed-Forward+Layers+Are+Key-Value+Memories
8. Large Language Models Struggle to Learn Long-Tail Knowledge — Nikhil Kandpal, Haikang Deng, Adam Roberts, Eric Wallace, Colin Raffel, 2023
https://scholar.google.com/scholar?q=Large+Language+Models+Struggle+to+Learn+Long-Tail+Knowledge
9. Taken Out of Context: On Measuring Situational Awareness in LLMs — Lukas Berglund, Asa Cooper Stickland, Mikita Balesni, et al., 2023
https://scholar.google.com/scholar?q=Taken+Out+of+Context%3A+On+Measuring+Situational+Awareness+in+LLMs
Interactive Visualization: The Reversal Curse: When A Is B But Not B Is A