This episode traces the "Knowing-Using Gap" — the puzzling phenomenon where fine-tuning a language model on a new fact produces instant, perfect recall of that fact in isolation, yet the model fails when asked to actually reason with it in multi-step tasks. Drawing on a July 2026 arXiv paper from HKUST researchers, the discussion covers how the authors use a novel "self-patching" technique — building on ROME's causal tracing and PatchScope — to trace exactly where a memorized fact sits inside a model's layers and why it's inaccessible to reasoning circuits. The central finding is the "knowledge-circuit misalignment hypothesis": the fact isn't missing from the model at all, it's simply stored in the wrong layers — filed in storage/recall circuits rather than the mid-layer circuits that handle chaining and intersection reasoning. The conversation situates this within the broader landscape of knowledge injection methods (RAG, model editing like ROME/MEMIT, and fine-tuning) and prior benchmarks like MQuAKE and RippleEdits that documented the same failure without explaining it. Listeners interested in interpretability, LLM training dynamics, or why fine-tuned knowledge often doesn't "stick" for reasoning will find the mechanistic account — and the paper's numbers on how much of that lost reasoning is actually recoverable — a compelling departure from purely behavioral benchmarking.
Sources:
1. Towards Mechanistically Understanding Why Memorized Knowledge Fails to Generalize in Large Language Model Finetuning — Lu Dai, Ziyang Rao, Yili Wang, Hanqing Wang, Hao Liu, Hui Xiong, 2026
http://arxiv.org/abs/2607.08393
2. Locating and Editing Factual Associations in GPT — Kevin Meng, David Bau, Alex Andonian, Yonatan Belinkov, 2022
https://scholar.google.com/scholar?q=Locating+and+Editing+Factual+Associations+in+GPT
3. MQuAKE: Assessing Knowledge Editing in Language Models via Multi-Hop Questions — Zexuan Zhong, Zhengxuan Wu, Christopher D. Manning, Christopher Potts, Danqi Chen, 2023
https://scholar.google.com/scholar?q=MQuAKE%3A+Assessing+Knowledge+Editing+in+Language+Models+via+Multi-Hop+Questions
4. Do Large Language Models Latently Perform Multi-Hop Reasoning? — Sohee Yang, Elena Gribovskaya, Nora Kassner, Mor Geva, Sebastian Riedel, 2024
https://scholar.google.com/scholar?q=Do+Large+Language+Models+Latently+Perform+Multi-Hop+Reasoning%3F
5. Progress Measures for Grokking via Mechanistic Interpretability — Neel Nanda, Lawrence Chan, Tom Lieberum, Jess Smith, Jacob Steinhardt, 2023
https://scholar.google.com/scholar?q=Progress+Measures+for+Grokking+via+Mechanistic+Interpretability
6. Physics of Language Models: Part 3.2, Knowledge Manipulation — Zeyuan Allen-Zhu, Yuanzhi Li, 2023
https://scholar.google.com/scholar?q=Physics+of+Language+Models%3A+Part+3.2%2C+Knowledge+Manipulation
7. Hopping too late: Exploring the limitations of large language models on multi-hop queries — Eden Biran, Daniela Gottesman, Sohee Yang, Mor Geva, Amir Globerson, 2024
https://scholar.google.com/scholar?q=Hopping+too+late%3A+Exploring+the+limitations+of+large+language+models+on+multi-hop+queries
8. Cake: Circuit-aware editing enables generalizable knowledge learners — Yunzhi Yao, Jizhan Fang, Jia-Chen Gu, Ningyu Zhang, Shumin Deng, Huajun Chen, Nanyun Peng, 2025
https://scholar.google.com/scholar?q=Cake%3A+Circuit-aware+editing+enables+generalizable+knowledge+learners
9. Model editing at scale leads to gradual and catastrophic forgetting — Akshat Gupta, Anurag Rao, Gopala Anumanchipalli, 2024
https://scholar.google.com/scholar?q=Model+editing+at+scale+leads+to+gradual+and+catastrophic+forgetting