AI Post Transformers
Avsnitt

Language Model Continual Learning: Do Written Facts Survive Repeated Weight Updates?

Dela

This episode examines a solo-authored arXiv paper by Charles O'Neill (Baseten), "Can a Language Model Learn Facts Continually in Its Weights?", which asks whether facts written into a model's weights via LoRA adapters remain usable after dozens or even a hundred subsequent training updates, rather than simply measuring whether accuracy holds up. The discussion traces the theoretical lineage behind the question, from McCloskey and Cohen's 1989 catastrophic forgetting findings through the reversal curse and Gekhman et al.'s 2024 work showing fine-tuned facts fail at paraphrase and multi-hop reasoning even when the same fact works fine when placed directly in a prompt. To isolate what a written fact actually retains, O'Neill invents fictional entities and facts, writes them into Qwen3-4B via per-fact LoRA adapters, and tests recall, paraphrase, application, composition, and counterfactual reasoning against two benchmarks: an untouched base model and a prompt-injected ceiling. A lenient-versus-strict grading scheme introduces the "entailment gap," a metric for how often a model merely restates a trained premise instead of producing the actual answer. Listeners interested in knowledge editing, continual learning, or the mechanics of what it really means for a model to "know" something will find the paper's more rigorous framing of memory durability a useful corrective to accuracy-only benchmarks in this space.

Sources: 1. Can a Language Model Learn Facts Continually in Its Weights? — Charles O'Neill, 2026 http://arxiv.org/abs/2607.11020 2. Catastrophic Interference in Connectionist Networks: The Sequential Learning Problem — Michael McCloskey, Neal J. Cohen, 1989 https://scholar.google.com/scholar?q=Catastrophic+Interference+in+Connectionist+Networks%3A+The+Sequential+Learning+Problem 3. Overcoming Catastrophic Forgetting in Neural Networks — James Kirkpatrick, Razvan Pascanu, Neil Rabinowitz, Joel Veness, Guillaume Desjardins, Andrei A. Rusu, et al. (DeepMind), 2017 https://scholar.google.com/scholar?q=Overcoming+Catastrophic+Forgetting+in+Neural+Networks 4. An Empirical Investigation of Catastrophic Forgetting in Gradient-Based Neural Networks — Ian J. Goodfellow, Mehdi Mirza, Da Xiao, Aaron Courville, Yoshua Bengio, 2013 https://scholar.google.com/scholar?q=An+Empirical+Investigation+of+Catastrophic+Forgetting+in+Gradient-Based+Neural+Networks 5. An Empirical Study of Catastrophic Forgetting in Large Language Models During Continual Fine-tuning — Yun Luo, Zhen Yang, Fandong Meng, Yafu Li, Jie Zhou, Yue Zhang, 2023 https://scholar.google.com/scholar?q=An+Empirical+Study+of+Catastrophic+Forgetting+in+Large+Language+Models+During+Continual+Fine-tuning 6. Locating and Editing Factual Associations in GPT — Kevin Meng, David Bau, Alex Andonian, Yonatan Belinkov, 2022 https://scholar.google.com/scholar?q=Locating+and+Editing+Factual+Associations+in+GPT 7. Mass-Editing Memory in a Transformer — Kevin Meng, Arnab Sen Sharma, Alex Andonian, Yonatan Belinkov, David Bau, 2023 https://scholar.google.com/scholar?q=Mass-Editing+Memory+in+a+Transformer 8. Fast Model Editing at Scale — Eric Mitchell, Charles Lin, Antoine Bosselut, Chelsea Finn, Christopher D. Manning, 2022 https://scholar.google.com/scholar?q=Fast+Model+Editing+at+Scale 9. MQuAKE: Assessing Knowledge Editing in Language Models via Multi-Hop Questions — Zexuan Zhong, Zhengxuan Wu, Christopher D. Manning, Christopher Potts, Danqi Chen, 2023 https://scholar.google.com/scholar?q=MQuAKE%3A+Assessing+Knowledge+Editing+in+Language+Models+via+Multi-Hop+Questions 10. RL's razor: Why online reinforcement learning forgets less — Idan Shenfeld, Jyothish Pari, Pulkit Agrawal, 2025 https://scholar.google.com/scholar?q=RL%27s+razor%3A+Why+online+reinforcement+learning+forgets+less 11. Does localization inform editing? Surprising differences in causality-based localization vs. knowledge editing — Peter Hase, Mohit Bansal, Been Kim, Asma Ghandeharioun, 2023 https://scholar.google.com/scholar?q=Does+localization+inform+editing%3F+Surprising+differences+in+causality-based+localization+vs.+knowledge+editing 12. Towards mechanistically understanding why memorized knowledge fails to generalize in large language model finetuning — Lu Dai, Ziyang Rao, Yili Wang, Hanqing Wang, Hao Liu, Hui Xiong, 2026 https://scholar.google.com/scholar?q=Towards+mechanistically+understanding+why+memorized+knowledge+fails+to+generalize+in+large+language+model+finetuning 13. Model editing at scale leads to gradual and catastrophic forgetting — Akshat Gupta, Anurag Rao, Gopala Anumanchipalli, 2024 https://scholar.google.com/scholar?q=Model+editing+at+scale+leads+to+gradual+and+catastrophic+forgetting 14. AI models collapse when trained on recursively generated data — Ilia Shumailov, Zakhar Shumaylov, Yiren Zhao, Nicolas Papernot, Ross Anderson, Yarin Gal, 2024 https://scholar.google.com/scholar?q=AI+models+collapse+when+trained+on+recursively+generated+data 15. LoRA vs full fine-tuning: An illusion of equivalence — Reece Shuttleworth, Jacob Andreas, Antonio Torralba, Pratyusha Sharma, 2025 https://scholar.google.com/scholar?q=LoRA+vs+full+fine-tuning%3A+An+illusion+of+equivalence 16. Availability versus accessibility of information in memory for words — Endel Tulving, Zena Pearlstone, 1966 https://scholar.google.com/scholar?q=Availability+versus+accessibility+of+information+in+memory+for+words

Interactive Visualization: Language Model Continual Learning: Do Written Facts Survive Repeated Weight Updates?

Podden och tillhörande omslagsbild på den här sidan tillhör mcgrof. Innehållet i podden är skapat av mcgrof och inte av, eller tillsammans med, Poddtoppen.