This episode explores OPSDL (On-Policy Self-Distillation for Long-Context Language Models), a technique out of Baidu that tackles the gap between a model's advertised context window and how much of it the model can actually reason over faithfully. Rather than training on a separate reward model or human-labeled preferences, OPSDL has the same model supervise itself: a version reading a short, evidence-only excerpt acts as teacher for the version reading the full long document, with reverse KL divergence pulling the long-context student toward the short-context teacher's most confident token-by-token predictions. The discussion traces how this improves on prior approaches like LongPO and LongReward, which rely on blunt, sequence-level preference signals, and explains why the short-context teacher's immunity to irrelevant material makes this a direct lever against hallucination in noisy long documents. The hosts also flag a methodological wrinkle worth watching: every experiment runs on a single model family (Qwen2.5-Instruct) at three parameter scales, raising questions about whether the paper's "generalization" claims hold up under scrutiny. Listeners interested in the mechanics of long-context reasoning, self-distillation, and how models can be trained to trust their own better-calibrated judgments will find the framing compelling.

Sources: 1. OPSDL: On-Policy Self-Distillation for Long-Context Language Models — Xinsen Zhang, Zhenkai Ding, Tianjun Pan, Run Yang, Chun Kang, Xue Xiong, Jingnan Gu, 2026 http://arxiv.org/abs/2604.17535 2. Longpo: Long context self-evolution of large language models through short-to-long preference optimization — Guanzheng Chen, Xin Li, Michael Qizhe Shieh, Lidong Bing, 2025 https://scholar.google.com/scholar?q=Longpo%3A+Long+context+self-evolution+of+large+language+models+through+short-to-long+preference+optimization 3. Self-distilled reasoner: On-policy self-distillation for large language models (OPSD) — Siyan Zhao, Zhihui Xie, Mengchen Liu, Jing Huang, Guan Pang, Feiyu Chen, Aditya Grover, 2026 https://scholar.google.com/scholar?q=Self-distilled+reasoner%3A+On-policy+self-distillation+for+large+language+models+%28OPSD%29 4. On-policy distillation of language models: Learning from self-generated mistakes (GKD) — Rishabh Agarwal, Nino Vieillard, Yongchao Zhou, Piotr Stanczyk, Sabela Ramos Garea, Matthieu Geist, Olivier Bachem, 2024 https://scholar.google.com/scholar?q=On-policy+distillation+of+language+models%3A+Learning+from+self-generated+mistakes+%28GKD%29 5. Ruler: What's the real context size of your long-context language models? — Cheng-Ping Hsieh, Simeng Sun, Samuel Kriman, Shantanu Acharya, Dima Rekesh, Fei Jia, Yang Zhang, Boris Ginsburg, 2024 https://scholar.google.com/scholar?q=Ruler%3A+What%27s+the+real+context+size+of+your+long-context+language+models%3F 6. Solopo: Unlocking long-context capabilities in llms via short-to-long preference optimization — Huashan Sun, Shengyi Liao, Yansen Han, Yu Bai, Yang Gao, Cheng Fu, Weizhou Shen, Fanqi Wan, Ming Yan, Ji Zhang, et al., 2025 https://scholar.google.com/scholar?q=Solopo%3A+Unlocking+long-context+capabilities+in+llms+via+short-to-long+preference+optimization

Interactive Visualization: OPSDL: Teaching Long-Context Models to Trust Their Short-Context Selves

Podden och tillhörande omslagsbild på den här sidan tillhör mcgrof. Innehållet i podden är skapat av mcgrof och inte av, eller tillsammans med, Poddtoppen.