This episode explores "Post-Training Science for Supervised Fine-Tuning," a study from Baseten researchers that treats SFT hyperparameter choices — learning rate, batch size, LoRA rank, epochs, and optimizer — as empirical questions rather than inherited folklore. Using controlled, one-variable-at-a-time sweeps across Qwen3 and Llama models ranging from 0.6 billion to 235 billion parameters (including dense and mixture-of-experts architectures), the hosts unpack findings like a surprisingly stable optimal LoRA learning rate that holds flat across two orders of magnitude in model scale, and a batch size that behaves more like a compute-cost tradeoff than a quality lever. They dig into how LoRA stacks up against full fine-tuning, with LoRA recovering a median 98% of full fine-tuning's gains using a fraction of the trainable parameters, and discuss where increasing LoRA rank stops paying off. Along the way, they flag a methodological wrinkle worth scrutinizing: the same evaluator used to construct the training data is also used to judge the fine-tuned model's output quality. Listeners interested in practical, evidence-based guidance for production fine-tuning — rather than another one-off trick — will find concrete, scale-tested defaults here.

Sources: 1. Post-Training Science: Scaling Laws for SFT and LoRA https://www.datocms-assets.com/104802/1781805778-baseten-research-sft.pdf 2. LoRA: Low-Rank Adaptation of Large Language Models — Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, Weizhu Chen (Microsoft Research), 2021 (ICLR 2022) https://scholar.google.com/scholar?q=LoRA%3A+Low-Rank+Adaptation+of+Large+Language+Models 3. QLoRA: Efficient Finetuning of Quantized LLMs — Tim Dettmers, Artidoro Pagnoni, Ari Holtzman, Luke Zettlemoyer (University of Washington), 2023 (NeurIPS 2023) https://scholar.google.com/scholar?q=QLoRA%3A+Efficient+Finetuning+of+Quantized+LLMs 4. LIMA: Less Is More for Alignment — Chunting Zhou, Pengfei Liu, Puxin Xu, Srini Iyer, Jiao Sun, Yuning Mao, Xuezhe Ma, Avia Efrat, Ping Yu, Lili Yu, Susan Zhang, Gargi Ghosh, Mike Lewis, Luke Zettlemoyer, Omer Levy (Meta AI / external collaborators), 2023 (NeurIPS 2023) https://scholar.google.com/scholar?q=LIMA%3A+Less+Is+More+for+Alignment 5. Muon: An optimizer for hidden layers in neural networks / Muon is Scalable for LLM Training — Keller Jordan, Yuchen Jin, Vlado Boza, Jiacheng You, Franz Cesista, Laker Newhouse, Jeremy Bernstein (original Muon, 2024); Moonshot AI / Kimi team, led by Jingyuan Liu and Jianlin Su, et al. (scaling follow-up, 2025), 2024-2025 https://scholar.google.com/scholar?q=Muon%3A+An+optimizer+for+hidden+layers+in+neural+networks+%2F+Muon+is+Scalable+for+LLM+Training 6. Model Soups: Averaging Weights of Multiple Fine-Tuned Models Improves Accuracy Without Increasing Inference Time — Mitchell Wortsman, Gabriel Ilharco, Samir Yitzhak Gadre, et al., 2022 https://scholar.google.com/scholar?q=Model+Soups%3A+Averaging+Weights+of+Multiple+Fine-Tuned+Models+Improves+Accuracy+Without+Increasing+Inference+Time 7. Intrinsic Dimensionality Explains the Effectiveness of Language Model Fine-Tuning — Armen Aghajanyan, Luke Zettlemoyer, Sonal Gupta, 2020 https://scholar.google.com/scholar?q=Intrinsic+Dimensionality+Explains+the+Effectiveness+of+Language+Model+Fine-Tuning 8. S-LoRA: Serving Thousands of Concurrent LoRA Adapters — Ying Sheng, Shiyi Cao, Dacheng Li, et al., 2023 https://scholar.google.com/scholar?q=S-LoRA%3A+Serving+Thousands+of+Concurrent+LoRA+Adapters

Interactive Visualization: Post-Training Science: Scaling Laws for SFT and LoRA

Podden och tillhörande omslagsbild på den här sidan tillhör mcgrof. Innehållet i podden är skapat av mcgrof och inte av, eller tillsammans med, Poddtoppen.