This episode explores a paper introducing Gimbal, a serving system for Mixture-of-Experts LLMs that coordinates two scheduling decisions previously handled separately: which backend engine a request enters through and where individual experts are physically placed across GPUs. The hosts walk through why request-count-based load balancing fails for MoE models — a 200-token and a 2,000-token request look identical by count but differ tenfold in KV-cache pressure and time-to-first-token — and why expert activation is highly uneven and source-dependent, with profiling on Qwen3-80B showing over 83% of one layer's traffic from a given engine routing to remote, non-local experts. The core argument is that dispatch and placement are a coupled optimization problem rather than two problems to solve independently, since offline expert rebalancing is blind to real-time backend pressure like queue depth and KV-cache usage. Reported results back the claim: 42.9% lower time-to-first-token and 33.3% lower time-per-output-token versus vLLM. Listeners interested in LLM inference infrastructure will find a concrete, numbers-driven case for why sparse-model serving needs system-level coordination rather than heuristics borrowed from dense-model deployments.

Sources: 1. Coordinated Scheduling for MoE LLM Serving — Yifan Sun, Zhexiang Zhang, Jiantong Jiang, Gholamreza Haffari, Minxian Xu, Feng Liu, Rajkumar Buyya, Adel N. Toosi, 2026 http://arxiv.org/abs/2606.15177v1 2. Orca: A Distributed Serving System for Transformer-Based Generative Models — Gyeong-In Yu, Joo Seong Jeong, Geon-Woo Kim, Soojeong Kim, Byung-Gon Chun, 2022 https://scholar.google.com/scholar?q=Orca%3A+A+Distributed+Serving+System+for+Transformer-Based+Generative+Models 3. Efficient Memory Management for Large Language Model Serving with PagedAttention — Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, et al., 2023 https://scholar.google.com/scholar?q=Efficient+Memory+Management+for+Large+Language+Model+Serving+with+PagedAttention 4. DeepSeek-V3 Technical Report — DeepSeek-AI (large team; Damai Dai, Wenfeng Liang, et al.), 2024 https://scholar.google.com/scholar?q=DeepSeek-V3+Technical+Report 5. Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving — Ruoyu Qin, Zheming Li, Weiran He, Mingxing Zhang, Yongwei Wu, et al. (Moonshot AI / Kimi), 2024 https://scholar.google.com/scholar?q=Mooncake%3A+A+KVCache-centric+Disaggregated+Architecture+for+LLM+Serving 6. Preble: Efficient Distributed Prompt Scheduling for LLM Serving — Vikranth Srivatsa, Zijian He, Reyna Abhyankar, Dongming Li, Yiying Zhang, 2024/2025 (ICLR 2025) https://scholar.google.com/scholar?q=Preble%3A+Efficient+Distributed+Prompt+Scheduling+for+LLM+Serving 7. MoETuner: Optimized Mixture of Expert Serving with Balanced Expert Placement and Token Routing — Seokjin Go, Divya Mahajan, 2025 https://scholar.google.com/scholar?q=MoETuner%3A+Optimized+Mixture+of+Expert+Serving+with+Balanced+Expert+Placement+and+Token+Routing 8. Semantic Parallelism: Redefining Efficient MoE Inference via Model-Data Co-Scheduling (Sem-MoE) — Yan Li, Zhenyu Zhang, Zhengang Wang, Pengfei Chen, Pengfei Zheng, 2026 (ICLR) https://scholar.google.com/scholar?q=Semantic+Parallelism%3A+Redefining+Efficient+MoE+Inference+via+Model-Data+Co-Scheduling+%28Sem-MoE%29 9. Exploiting Inter-layer Expert Affinity for Accelerating Mixture-of-Experts Model Inference — Jinghan Yao, Quentin Anthony, Aamir Shafi, Hari Subramoni, Dhabaleswar K. Panda, 2024 (IPDPS) https://scholar.google.com/scholar?q=Exploiting+Inter-layer+Expert+Affinity+for+Accelerating+Mixture-of-Experts+Model+Inference 10. JANUS: Disaggregating Attention and Experts for Scalable MoE Inference — Zhexiang Zhang, Ye Wang, Yumiao Zhao, et al., 2025 https://scholar.google.com/scholar?q=JANUS%3A+Disaggregating+Attention+and+Experts+for+Scalable+MoE+Inference

Interactive Visualization: Coordinated MoE Scheduling: How Gimbal Fixes Expert Locality

Podden och tillhörande omslagsbild på den här sidan tillhör mcgrof. Innehållet i podden är skapat av mcgrof och inte av, eller tillsammans med, Poddtoppen.