This episode explores "Thought Anchors: Which LLM Reasoning Steps Matter?" by Paul C. Bogdan and Uzay Macar, with senior authors Neel Nanda and Arthur Conmy, examining which specific sentences in a chain-of-thought trace carry disproportionate causal weight over a model's final answer. The discussion covers why standard mechanistic interpretability tools, built for single forward passes, break down for reasoning models that generate thousands of sequentially dependent tokens, and how the authors instead treat the sentence as the right unit of analysis. Three independent methods are unpacked: counterfactual resampling with embedding-based filtering, receiver-head attention analysis, and attention suppression measured via KL divergence, all converging on identifying "thought anchors." A concrete case study on a base-16-to-binary conversion problem shows how a single pivot sentence rescues an otherwise wrong reasoning trace, illustrating the stakes in vivid detail. Listeners interested in interpretability, reasoning-model behavior, and how backtracking and self-correction actually work under the hood will find the mechanistic grounding — and its connection to related work like the s1 paper's "Wait"-token forcing — especially compelling.
Sources:
1. Thought Anchors: Which LLM Reasoning Steps Matter? — Paul C. Bogdan, Uzay Macar, Neel Nanda, Arthur Conmy, 2025
http://arxiv.org/abs/2506.19143
2. s1: Simple test-time scaling — Niklas Muennighoff, Zitong Yang, Weijia Shi, et al., 2025
https://scholar.google.com/scholar?q=s1%3A+Simple+test-time+scaling
3. Understanding reasoning in thinking language models via steering vectors — Constantin Venhoff, Iván Arcuschin, Philip Torr, Arthur Conmy, Neel Nanda, 2025
https://scholar.google.com/scholar?q=Understanding+reasoning+in+thinking+language+models+via+steering+vectors
4. Chain-of-thought reasoning in the wild is not always faithful — Iván Arcuschin, Jett Janiak, Robert Krzyzanowski, Senthooran Rajamanoharan, Neel Nanda, Arthur Conmy, 2025
https://scholar.google.com/scholar?q=Chain-of-thought+reasoning+in+the+wild+is+not+always+faithful
5. Reasoning models don't always say what they think — Yanda Chen, Joe Benton, Ansh Radhakrishnan, et al., 2025
https://scholar.google.com/scholar?q=Reasoning+models+don%27t+always+say+what+they+think
6. Chain of thought monitorability: A new and fragile opportunity for AI safety — Tomek Korbak, Mikita Balesni, Elizabeth Barnes, et al., 2025
https://scholar.google.com/scholar?q=Chain+of+thought+monitorability%3A+A+new+and+fragile+opportunity+for+AI+safety
7. Forking paths in neural text generation — Eric Bigelow, Ari Holtzman, Hidenori Tanaka, Tomer Ullman, 2024
https://scholar.google.com/scholar?q=Forking+paths+in+neural+text+generation
8. Monitoring reasoning models for misbehavior and the risks of promoting obfuscation — Bowen Baker, Joost Huizinga, Leo Gao, et al., 2025
https://scholar.google.com/scholar?q=Monitoring+reasoning+models+for+misbehavior+and+the+risks+of+promoting+obfuscation
Interactive Visualization: Thought Anchors: Which Sentences Really Drive LLM Reasoning