This episode explores how inference-time scaling breaks down once large language models shift from short chat responses to long chain-of-thought reasoning, drawing on Micron and Argonne National Laboratory's research spanning models from 8 billion to 671 billion parameters. It explains the divide between compute-bound prefill and bandwidth-bound decode phases, and how reasoning traces exceeding ten thousand tokens push systems into a "capacity-bound" regime where the KV cache — not raw FLOPs — becomes the limiting resource. The discussion contrasts three parallelism strategies (data, tensor, and pipeline) and shows why data parallelism, the industry default, hits a capacity wall under reasoning workloads even though it remains optimal for short prompts. It also covers how architectural choices like Grouped-Query Attention versus DeepSeek-R1's Mixture-of-Experts design and Multi-Head Latent Attention change how much cache pressure a model generates per token. Listeners interested in the practical engineering tradeoffs behind serving reasoning models at scale will find concrete guidance on when each parallelism strategy actually wins, backed by measurements on an 8x H200 NVLink node.
Sources:
1. Understanding Inference Scaling for LLMs: Bottlenecks, Trade-offs, and Performance Principles — Moiz Arif, Avinash Maurya, Sudharshan Vazhkudai, Bogdan Nicolae, 2026
http://arxiv.org/abs/2605.19775
2. PyTorch Distributed: Experiences on Accelerating Data Parallel Training — Shen Li, Yanli Zhao, Rohan Varma, et al. (Meta AI / PyTorch team), 2020
https://scholar.google.com/scholar?q=PyTorch+Distributed%3A+Experiences+on+Accelerating+Data+Parallel+Training
3. ZeRO: Memory Optimizations Toward Training Trillion Parameter Models — Samyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, Yuxiong He (Microsoft), 2020
https://scholar.google.com/scholar?q=ZeRO%3A+Memory+Optimizations+Toward+Training+Trillion+Parameter+Models
4. Efficient Memory Management for Large Language Model Serving with PagedAttention (vLLM) — Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, et al. (UC Berkeley), 2023
https://scholar.google.com/scholar?q=Efficient+Memory+Management+for+Large+Language+Model+Serving+with+PagedAttention+%28vLLM%29
5. Orca: A Distributed Serving System for Transformer-Based Generative Models — Gyeong-In Yu, Joo Seong Jeong, Geon-Woo Kim, Soojeong Kim, Byung-Gon Chun (Seoul National University / FriendliAI), 2022
https://scholar.google.com/scholar?q=Orca%3A+A+Distributed+Serving+System+for+Transformer-Based+Generative+Models
6. DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving — Yinmin Zhong, Shengyu Liu, Junda Chen, Jianbo Hu, Yibo Zhu, Xuanzhe Liu, Xin Jin, Hao Zhang, 2024
https://scholar.google.com/scholar?q=DistServe%3A+Disaggregating+Prefill+and+Decoding+for+Goodput-optimized+Large+Language+Model+Serving
7. Splitwise: Efficient Generative LLM Inference Using Phase Splitting — Pratyush Patel, Esha Choukse, Chaojie Zhang, Aashaka Shah, Íñigo Goiri, Saeed Maleki, Ricardo Bianchini, 2024
https://scholar.google.com/scholar?q=Splitwise%3A+Efficient+Generative+LLM+Inference+Using+Phase+Splitting
8. Llumnix: Dynamic Scheduling for Large Language Model Serving — Biao Sun, Ziming Huang, Hanyu Zhao, Wencong Xiao, Xinyi Zhang, Yong Li, Wei Lin, 2024
https://scholar.google.com/scholar?q=Llumnix%3A+Dynamic+Scheduling+for+Large+Language+Model+Serving
9. vAttention: Dynamic Memory Management for Serving LLMs without PagedAttention — Ramya Prabhu, Ajay Nayak, Jayashree Mohan, Ramachandran Ramjee, Ashish Panwar, 2024
https://scholar.google.com/scholar?q=vAttention%3A+Dynamic+Memory+Management+for+Serving+LLMs+without+PagedAttention
10. Efficient Memory Management for Large Language Model Serving with PagedAttention (already cited [22]) — cross-check against KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache — Zirui Liu, Jiayi Yuan, Hongye Jin, Shaochen Zhong, Zhaozhuo Xu, Vladimir Braverman, Beidi Chen, Xia Hu, 2024
https://scholar.google.com/scholar?q=Efficient+Memory+Management+for+Large+Language+Model+Serving+with+PagedAttention+%28already+cited+%5B22%5D%29+%E2%80%94+cross-check+against+KIVI%3A+A+Tuning-Free+Asymmetric+2bit+Quantization+for+KV+Cache
Interactive Visualization: Understanding Inference Scaling: Prefill, Decode, and Reasoning Bottlenecks