This episode closes a three-part arc on "Understanding and Steering the Cognitive Behaviors of Reasoning Models at Test-Time," examining how researchers identify and manipulate the specific attention heads responsible for triggering non-linear reasoning detours in language models. The discussion covers the technical pipeline: segmenting chains of thought at delimiter tokens, fitting per-head linear probes to find heads that predict reasoning-style shifts, then denoising those signals through shared-subspace PCA across heads in a layer. A live walkthrough shows the payoff — pausing mid-generation and rotating a hidden state to suppress or amplify the "non-linear" direction sends the model down a 12-step versus 45-step path to the same correct answer, making abstract "redundant reasoning" concrete. Benchmark results follow, with the CREST method cutting token usage over 30% while matching or beating baseline accuracy across four architectures (dense and mixture-of-experts) and transferring — without recalibration — from math-only training data to code generation, science QA, and scheduling tasks. The hosts close by flagging an unresolved inconsistency in the paper: three different head-selection ratios appear across sections, with no clear statement of which one produced the headline results.

Sources: 1. Understanding and Steering the Cognitive Behaviors of Reasoning Models at Test-Time — Zhenyu Zhang, Xiaoxia Wu, Zhongzhu Zhou, Qingyang Wu, Yineng Zhang, Pragaash Ponnusamy, Harikaran Subbaraj, Jue Wang, Shuaiwen Leon Song, Ben Athiwaratkun, 2025 http://arxiv.org/abs/2512.24574 2. Cognitive Behaviors that Enable Self-Improving Reasoners, or, Four Habits of Highly Effective STaRs — Kanishk Gandhi, Ayush Chakravarthy, Anikait Singh, Nathan Lile, Noah D. Goodman, 2025 https://scholar.google.com/scholar?q=Cognitive+Behaviors+that+Enable+Self-Improving+Reasoners%2C+or%2C+Four+Habits+of+Highly+Effective+STaRs 3. Retrieval Head Mechanistically Explains Long-Context Factuality — Wenhao Wu, Yizhong Wang, Guangxuan Xiao, Hao Peng, Yao Fu, 2024 https://scholar.google.com/scholar?q=Retrieval+Head+Mechanistically+Explains+Long-Context+Factuality 4. Representation Engineering: A Top-Down Approach to AI Transparency — Andy Zou, Long Phan, Sarah Chen, et al., 2025 https://scholar.google.com/scholar?q=Representation+Engineering%3A+A+Top-Down+Approach+to+AI+Transparency 5. Reasoning Models Can Be Effective Without Thinking — Wenjie Ma, Jingxuan He, Charlie Snell, Tyler Griggs, Sewon Min, Matei Zaharia, 2025 https://scholar.google.com/scholar?q=Reasoning+Models+Can+Be+Effective+Without+Thinking 6. Thoughts Are All Over the Place: On the Underthinking of o1-like LLMs — Yue Wang, Qiuzhi Liu, Jiahao Xu, et al., 2025 https://scholar.google.com/scholar?q=Thoughts+Are+All+Over+the+Place%3A+On+the+Underthinking+of+o1-like+LLMs

Interactive Visualization: Steering Reasoning Models' Cognitive Behaviors at Test-Time

Podden och tillhörande omslagsbild på den här sidan tillhör mcgrof. Innehållet i podden är skapat av mcgrof och inte av, eller tillsammans med, Poddtoppen.