This episode explores "Post-Training Science for Supervised Fine-Tuning," a study from Baseten researchers that treats SFT hyperparameter choices — learning rate, batch size, LoRA rank, epochs, and optimizer — as empirical questions rather than inherited folklore. Using controlled, one-variable-at-a-time sweeps across Qwen3 and Llama models ranging from 0.6 billion to 235 billion parameters (including dense and mixture-of-experts architectures), the hosts unpack findings like a surprisingly stable optimal LoRA learning rate that holds flat across two orders of magnitude in model scale, and a batch size that behaves more like a compute-cost tradeoff than a quality lever. They dig into how LoRA stacks up against full fine-tuning, with LoRA recovering a median 98% of full fine-tuning's gains using a fraction of the trainable parameters, and discuss where increasing LoRA rank stops paying off. Along the way, they flag a methodological wrinkle worth scrutinizing: the same evaluator used to construct the training data is also used to judge the fine-tuned model's output quality. Listeners interested in practical, evidence-based guidance for production fine-tuning — rather than another one-off trick — will find concrete, scale-tested defaults here.
Sources:
1. Post-Training Science: Scaling Laws for SFT and LoRA
https://www.datocms-assets.com/104802/1781805778-baseten-research-sft.pdf
2. LoRA: Low-Rank Adaptation of Large Language Models — Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, Weizhu Chen (Microsoft Research), 2021 (ICLR 2022)
https://scholar.google.com/scholar?q=LoRA%3A+Low-Rank+Adaptation+of+Large+Language+Models
3. QLoRA: Efficient Finetuning of Quantized LLMs — Tim Dettmers, Artidoro Pagnoni, Ari Holtzman, Luke Zettlemoyer (University of Washington), 2023 (NeurIPS 2023)
https://scholar.google.com/scholar?q=QLoRA%3A+Efficient+Finetuning+of+Quantized+LLMs
4. LIMA: Less Is More for Alignment — Chunting Zhou, Pengfei Liu, Puxin Xu, Srini Iyer, Jiao Sun, Yuning Mao, Xuezhe Ma, Avia Efrat, Ping Yu, Lili Yu, Susan Zhang, Gargi Ghosh, Mike Lewis, Luke Zettlemoyer, Omer Levy (Meta AI / external collaborators), 2023 (NeurIPS 2023)
https://scholar.google.com/scholar?q=LIMA%3A+Less+Is+More+for+Alignment
5. Muon: An optimizer for hidden layers in neural networks / Muon is Scalable for LLM Training — Keller Jordan, Yuchen Jin, Vlado Boza, Jiacheng You, Franz Cesista, Laker Newhouse, Jeremy Bernstein (original Muon, 2024); Moonshot AI / Kimi team, led by Jingyuan Liu and Jianlin Su, et al. (scaling follow-up, 2025), 2024-2025
https://scholar.google.com/scholar?q=Muon%3A+An+optimizer+for+hidden+layers+in+neural+networks+%2F+Muon+is+Scalable+for+LLM+Training
6. Model Soups: Averaging Weights of Multiple Fine-Tuned Models Improves Accuracy Without Increasing Inference Time — Mitchell Wortsman, Gabriel Ilharco, Samir Yitzhak Gadre, et al., 2022
https://scholar.google.com/scholar?q=Model+Soups%3A+Averaging+Weights+of+Multiple+Fine-Tuned+Models+Improves+Accuracy+Without+Increasing+Inference+Time
7. Intrinsic Dimensionality Explains the Effectiveness of Language Model Fine-Tuning — Armen Aghajanyan, Luke Zettlemoyer, Sonal Gupta, 2020
https://scholar.google.com/scholar?q=Intrinsic+Dimensionality+Explains+the+Effectiveness+of+Language+Model+Fine-Tuning
8. S-LoRA: Serving Thousands of Concurrent LoRA Adapters — Ying Sheng, Shiyi Cao, Dacheng Li, et al., 2023
https://scholar.google.com/scholar?q=S-LoRA%3A+Serving+Thousands+of+Concurrent+LoRA+Adapters
Interactive Visualization: Post-Training Science: Scaling Laws for SFT and LoRA