This episode traces the evolution from classic knowledge distillation to on-policy distillation and finally to Lightning OPD, a technique from NVIDIA researchers for post-training large reasoning models. The hosts unpack why on-policy distillation offers denser training signal than RLVR's sparse end-of-trace rewards, but has historically required an expensive live teacher model running alongside the student throughout training. They explain the paper's key insight — that a student's rollout distribution drifts only modestly from its SFT starting point — which motivates capturing the teacher's judgments once offline rather than serving it continuously. The discussion grounds the work in its lineage, from Hinton et al.'s original 2015 distillation paper to DeepMind's 2024 Generalized Knowledge Distillation formalism, while flagging why the infrastructure savings matter even more for sparse Mixture-of-Experts models. Listeners interested in the theory-first rigor behind cost-cutting techniques in LLM post-training will find the paper's formal proof-before-benchmarks approach a refreshing departure from the field's norm.
Sources:
1. Lightning OPD: Efficient Post-Training for Large Reasoning Models with Offline On-Policy Distillation — Yecheng Wu, Song Han, Hai Cai, 2026
http://arxiv.org/abs/2604.13010
2. On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes — Rishabh Agarwal, Nino Vieillard, Yongchao Zhou, Piotr Stanczyk, Sabela Ramos, Matthieu Geist, Olivier Bachem (Google DeepMind), 2024 (ICLR)
https://scholar.google.com/scholar?q=On-Policy+Distillation+of+Language+Models%3A+Learning+from+Self-Generated+Mistakes
3. On-Policy Distillation — Kevin Lu et al. (Thinking Machines Lab), 2025
https://scholar.google.com/scholar?q=On-Policy+Distillation
4. Distilling the Knowledge in a Neural Network — Geoffrey Hinton, Oriol Vinyals, Jeff Dean, 2015
https://scholar.google.com/scholar?q=Distilling+the+Knowledge+in+a+Neural+Network
5. Tulu 3: Pushing Frontiers in Open Language Model Post-Training — Nathan Lambert, Jacob Morrison, Valentina Pyatkin, et al. (Allen Institute for AI), 2024
https://scholar.google.com/scholar?q=Tulu+3%3A+Pushing+Frontiers+in+Open+Language+Model+Post-Training
6. DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning — DeepSeek-AI (Daya Guo et al.), 2025
https://scholar.google.com/scholar?q=DeepSeek-R1%3A+Incentivizing+Reasoning+Capability+in+LLMs+via+Reinforcement+Learning
7. DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models — Zhihong Shao, Peiyi Wang, Qihao Zhu, et al. (DeepSeek-AI), 2024
https://scholar.google.com/scholar?q=DeepSeekMath%3A+Pushing+the+Limits+of+Mathematical+Reasoning+in+Open+Language+Models
8. Let's Verify Step by Step — Hunter Lightman, Vineet Kosaraju, Yura Burda, et al. (OpenAI), 2023
https://scholar.google.com/scholar?q=Let%27s+Verify+Step+by+Step
9. On-policy distillation (Thinking Machines Lab: Connectionism) — Kevin Lu, Thinking Machines Lab, 2025
https://scholar.google.com/scholar?q=On-policy+distillation+%28Thinking+Machines+Lab%3A+Connectionism%29
10. Does reinforcement learning really incentivize reasoning capacity in LLMs beyond the base model? — Yang Yue, Zhiqi Chen, Rui Lu, Andrew Zhao, Zhaokai Wang, Shiji Song, Gao Huang, 2025
https://scholar.google.com/scholar?q=Does+reinforcement+learning+really+incentivize+reasoning+capacity+in+LLMs+beyond+the+base+model%3F
11. RL's razor: Why online reinforcement learning forgets less — Idan Shenfeld, Jyothish Pari, Pulkit Agrawal, 2025
https://scholar.google.com/scholar?q=RL%27s+razor%3A+Why+online+reinforcement+learning+forgets+less
12. Rethinking on-policy distillation of large language models: Phenomenology, mechanism, and recipe — Siyuan Chen, Daya Guo, Yichun Tan, Xiaohan Liang, Hao Zhou, Bo Zheng, Dejian Yang, 2026
https://scholar.google.com/scholar?q=Rethinking+on-policy+distillation+of+large+language+models%3A+Phenomenology%2C+mechanism%2C+and+recipe
13. Learning beyond teacher: Generalized on-policy distillation with reward extrapolation (ExOPD) — Wenkai Yang, Weijie Liu, Ruobing Xie, Kai Yang, Saiyong Yang, Yankai Lin, 2026
https://scholar.google.com/scholar?q=Learning+beyond+teacher%3A+Generalized+on-policy+distillation+with+reward+extrapolation+%28ExOPD%29
Interactive Visualization: Lightning OPD: Precomputing the Teacher for Cheaper Reasoning Distillation