This episode explores why identical reinforcement learning recipes produce wildly different reasoning ability in similarly-sized language models, examining "Cognitive Behaviors that Enable Self-Improving Reasoners" from Stanford and SynthLabs. Training Qwen-2.5-3B and Llama-3.2-3B on the number-puzzle game Countdown with identical PPO settings, Qwen jumps to roughly 60% accuracy while Llama plateaus around 30% — despite matched architecture size, algorithm, and hyperparameters. The discussion traces this gap to four cognitive behaviors already present in Qwen's pretrained weights before any RL begins: verification, backtracking, subgoal setting, and backward chaining (the last borrowed straight from 1970s-80s expert-system logic). Using GPT-4o-mini as an automated classifier across thousands of reasoning traces, the hosts unpack how these behaviors' presence — or absence — in a base model predicts whether reinforcement learning takes off or stalls, reframing the "just scale it up" narrative around what a model already knows how to do before training starts. It sets up the next question the arc will tackle: whether these behaviors can be deliberately installed in a model that lacks them.

Sources: 1. Cognitive Behaviors that Enable Self-Improving Reasoners, or, Four Habits of Highly Effective STaRs — Kanishk Gandhi, Ayush Chakravarthy, Anikait Singh, Nathan Lile, Noah D. Goodman, 2025 http://arxiv.org/abs/2503.01307 2. STaR: Bootstrapping Reasoning With Reasoning — Eric Zelikman, Yuhuai Wu, Jesse Mu, Noah D. Goodman, 2022 https://scholar.google.com/scholar?q=STaR%3A+Bootstrapping+Reasoning+With+Reasoning 3. Let's Verify Step by Step — Hunter Lightman, Vineet Kosaraju, Yura Burda, Harri Edwards, Bowen Baker, Teddy Lee, Jan Leike, John Schulman, Ilya Sutskever, Karl Cobbe, 2023 https://scholar.google.com/scholar?q=Let%27s+Verify+Step+by+Step 4. Tree of Thoughts: Deliberate Problem Solving with Large Language Models — Shunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran, Thomas L. Griffiths, Yuan Cao, Karthik Narasimhan, 2023 https://scholar.google.com/scholar?q=Tree+of+Thoughts%3A+Deliberate+Problem+Solving+with+Large+Language+Models 5. DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning — DeepSeek-AI (Daya Guo et al.), 2025 https://scholar.google.com/scholar?q=DeepSeek-R1%3A+Incentivizing+Reasoning+Capability+in+LLMs+via+Reinforcement+Learning 6. LAMBADA: Backward Chaining for Automated Reasoning in Natural Language — Seyed Mehran Kazemi, Najoung Kim, Deepti Bhatia, Xin Xu, Deepak Ramachandran, 2023 https://scholar.google.com/scholar?q=LAMBADA%3A+Backward+Chaining+for+Automated+Reasoning+in+Natural+Language 7. Selection-Inference: Exploiting Large Language Models for Interpretable Logical Reasoning — Antonia Creswell, Murray Shanahan, Irina Higgins, 2022 https://scholar.google.com/scholar?q=Selection-Inference%3A+Exploiting+Large+Language+Models+for+Interpretable+Logical+Reasoning 8. A Machine-Oriented Logic Based on the Resolution Principle — J. A. Robinson, 1965 https://scholar.google.com/scholar?q=A+Machine-Oriented+Logic+Based+on+the+Resolution+Principle 9. Demystifying Long Chain-of-Thought Reasoning in LLMs — Edward Yeo, Yuxuan Tong, Morry Niu, Graham Neubig, Xiang Yue, 2025 https://scholar.google.com/scholar?q=Demystifying+Long+Chain-of-Thought+Reasoning+in+LLMs 10. Stream of Search (SoS): Learning to Search in Language — Kanishk Gandhi, Denise H.J. Lee, Gabriel Grand, Muxin Liu, Winson Cheng, Archit Sharma, Noah Goodman, 2024 https://scholar.google.com/scholar?q=Stream+of+Search+%28SoS%29%3A+Learning+to+Search+in+Language 11. LLMs Can Easily Learn to Reason from Demonstrations: Structure, not Content, is What Matters! — Dacheng Li, Shiyi Cao, Tyler Griggs, Shu Liu, Xiangxi Mo, Shishir G. Patil, Matei Zaharia, Joseph E. Gonzalez, Ion Stoica, 2025 https://scholar.google.com/scholar?q=LLMs+Can+Easily+Learn+to+Reason+from+Demonstrations%3A+Structure%2C+not+Content%2C+is+What+Matters%21 12. There May Not Be Aha Moment in R1-Zero-like Training — A Pilot Study — Zichen Liu, Changyu Chen, Wenjun Li, Tianyu Pang, Chao Du, Min Lin, 2025 https://scholar.google.com/scholar?q=There+May+Not+Be+Aha+Moment+in+R1-Zero-like+Training+%25E2%2580%2594+A+Pilot+Study 13. Human Problem Solving: The State of the Theory in 1970 — Herbert A. Simon, Allen Newell, 1971 https://scholar.google.com/scholar?q=Human+Problem+Solving%3A+The+State+of+the+Theory+in+1970

Interactive Visualization: Cognitive Behaviors Behind Self-Improving Language Model Reasoners

Podden och tillhörande omslagsbild på den här sidan tillhör mcgrof. Innehållet i podden är skapat av mcgrof och inte av, eller tillsammans med, Poddtoppen.