This episode examines the "Reversal Curse," a 2023 finding from researchers at Vanderbilt, UK AI Safety Institute, Apollo Research, NYU, Sussex, and Oxford showing that autoregressive language models trained on "A is B" statements fail to infer "B is A," even though the two are logically equivalent. Using the example of Valentina Tereshkova, the discussion shows how a model finetuned on "Tereshkova was the first woman in space" answers forward questions perfectly but performs at chance when the question is reversed, and extends this to real-world GPT-4 results where "who is Tom Cruise's mother" scores 79% accuracy versus just 33% for the reverse query about Mary Lee Pfeiffer. The hosts unpack why this happens mechanically — next-token prediction bakes facts into weights as one-directional associations rather than symmetric relations like a knowledge-graph edge — and contrast this with in-context learning, where the same reversal works flawlessly, pinpointing the failure specifically to generalization from gradient-based training rather than a reasoning limitation. A debate over whether this is just a data-coverage gap versus a deeper meta-learning failure leads into the researchers' attempted fix: training on facts stated in both directions to see if models can pick up the general pattern. It's a compelling listen for anyone curious about the hidden asymmetries in how LLMs actually store knowledge versus how humans intuitively reason about it.

Sources: 1. The Reversal Curse: LLMs trained on "A is B" fail to learn "B is A" — Lukas Berglund, Meg Tong, Max Kaufmann, Mikita Balesni, Asa Cooper Stickland, Tomasz Korbak, Owain Evans, 2023 http://arxiv.org/abs/2309.12288 2. Locating and Editing Factual Associations in GPT (ROME) — Kevin Meng, David Bau, Alex Andonian, Yonatan Belinkov, 2022 https://scholar.google.com/scholar?q=Locating+and+Editing+Factual+Associations+in+GPT+%28ROME%29 3. Language Models as Knowledge Bases? — Fabio Petroni, Tim Rocktäschel, Patrick Lewis, Anton Bakhtin, Yuxiang Wu, Alexander H. Miller, Sebastian Riedel, 2019 https://scholar.google.com/scholar?q=Language+Models+as+Knowledge+Bases%3F 4. Reverse Training to Nurse the Reversal Curse — Olga Golovneva, Zeyuan Allen-Zhu, Jason Weston, Sainbayar Sukhbaatar (Meta AI), 2024 https://scholar.google.com/scholar?q=Reverse+Training+to+Nurse+the+Reversal+Curse 5. Physics of Language Models: Part 3.2, Knowledge Manipulation — Zeyuan Allen-Zhu, Yuanzhi Li, 2024 https://scholar.google.com/scholar?q=Physics+of+Language+Models%3A+Part+3.2%2C+Knowledge+Manipulation 6. Studying Large Language Model Generalization with Influence Functions — Roger Grosse, Juhan Bae, Cem Anil, et al., 2023 https://scholar.google.com/scholar?q=Studying+Large+Language+Model+Generalization+with+Influence+Functions 7. Transformer Feed-Forward Layers Are Key-Value Memories — Mor Geva, Roei Schuster, Jonathan Berant, Omer Levy, 2021 https://scholar.google.com/scholar?q=Transformer+Feed-Forward+Layers+Are+Key-Value+Memories 8. Large Language Models Struggle to Learn Long-Tail Knowledge — Nikhil Kandpal, Haikang Deng, Adam Roberts, Eric Wallace, Colin Raffel, 2023 https://scholar.google.com/scholar?q=Large+Language+Models+Struggle+to+Learn+Long-Tail+Knowledge 9. Taken Out of Context: On Measuring Situational Awareness in LLMs — Lukas Berglund, Asa Cooper Stickland, Mikita Balesni, et al., 2023 https://scholar.google.com/scholar?q=Taken+Out+of+Context%3A+On+Measuring+Situational+Awareness+in+LLMs

Interactive Visualization: The Reversal Curse: When A Is B But Not B Is A

Podden och tillhörande omslagsbild på den här sidan tillhör mcgrof. Innehållet i podden är skapat av mcgrof och inte av, eller tillsammans med, Poddtoppen.