This episode explores how speculative decoding — the standard trick for speeding up autoregressive LLM inference by having a cheap draft model propose tokens for batch verification — breaks down when applied to Mixture-of-Experts models. The discussion traces why decoding is memory-bandwidth bound rather than compute bound, then shows how MoE routing decouples a draft token's acceptance probability from its actual verification cost: tokens that route to disjoint experts (termed "expert scattering") force costly extra weight fetches even when a confidence-only selector rates them highly. The paper introduces EcoSpec, a cost-aware draft selector that accounts for expert-loading overhead rather than optimizing acceptance length (alpha) alone, and the hosts examine tradeoffs in Table 1 where EcoSpec sacrifices a small amount of acceptance probability on models like Qwen3 and GPT-OSS in exchange for reduced memory traffic. Listeners interested in LLM inference serving, hardware-aware systems design, or the practical limits of applying dense-model optimizations to sparse architectures will find the episode's reframing of a three-year-old assumption particularly compelling.
Sources:
1. Less Experts, Faster Decoding: Cost-Aware Speculative Decoding for Mixture-of-Experts — Jincheng Xie, Runheng Liu, Heyan Huang, Yawen Ling, Hanbin Dai, Yu Zheng, Wen Hu, 2026
http://arxiv.org/abs/2607.12696
2. Fast Inference from Transformers via Speculative Decoding — Yaniv Leviathan, Matan Kalman, Yossi Matias, 2023
https://scholar.google.com/scholar?q=Fast+Inference+from+Transformers+via+Speculative+Decoding
3. EAGLE-2: Faster Inference of Language Models with Dynamic Draft Trees — Yuhui Li, Fangyun Wei, Chao Zhang, Hongyang Zhang, 2024
https://scholar.google.com/scholar?q=EAGLE-2%3A+Faster+Inference+of+Language+Models+with+Dynamic+Draft+Trees
4. Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads — Tianle Cai, Yuhong Li, Zhengyang Geng, Hongwu Peng, Tri Dao, et al., 2024
https://scholar.google.com/scholar?q=Medusa%3A+Simple+LLM+Inference+Acceleration+Framework+with+Multiple+Decoding+Heads
5. Fiddler: CPU-GPU Orchestration for Fast Inference of Mixture-of-Experts Models — Keisuke Kamahori, Yile Gu, Kan Zhu, Baris Kasikci, 2024
https://scholar.google.com/scholar?q=Fiddler%3A+CPU-GPU+Orchestration+for+Fast+Inference+of+Mixture-of-Experts+Models
6. EAGLE-3: Scaling up Inference Acceleration of Large Language Models via Training-Time Test — Li, Y., Wei, F., Zhang, C., Zhang, H., 2026
https://scholar.google.com/scholar?q=EAGLE-3%3A+Scaling+up+Inference+Acceleration+of+Large+Language+Models+via+Training-Time+Test
7. MoE-Spec: Expert Budgeting for Efficient Speculative Decoding — McDanel, B., Li, S., Surineni, S., Khaitan, H., 2026
https://scholar.google.com/scholar?q=MoE-Spec%3A+Expert+Budgeting+for+Efficient+Speculative+Decoding
8. SP-MoE: Speculative Decoding and Prefetching for Accelerating MoE-based Model Inference — Chen, L., Wen, Z., Wu, T., Zhang, X., Wu, C., 2025
https://scholar.google.com/scholar?q=SP-MoE%3A+Speculative+Decoding+and+Prefetching+for+Accelerating+MoE-based+Model+Inference
9. MoE-SpeQ: Speculative Quantized Decoding with Proactive Expert Prefetching and Offloading for Mixture-of-Experts — Wang, W., Liu, J., Hou, X., Xia, X., Tang, P., Zhang, M., Li, C., Guo, M., 2025
https://scholar.google.com/scholar?q=MoE-SpeQ%3A+Speculative+Quantized+Decoding+with+Proactive+Expert+Prefetching+and+Offloading+for+Mixture-of-Experts
10. Bridging Draft Policy Misalignment: Group Tree Optimization for Speculative Decoding (GTO) — Hu, S., Li, J., Lu, Z., Zhou, P., 2026
https://scholar.google.com/scholar?q=Bridging+Draft+Policy+Misalignment%3A+Group+Tree+Optimization+for+Speculative+Decoding+%28GTO%29
11. MoE-Infinity: Efficient MoE Inference on Personal Machines with Sparsity-Aware Expert Cache — Xue, L., Fu, Y., Lu, Z., Mai, L., Marina, M., 2025
https://scholar.google.com/scholar?q=MoE-Infinity%3A+Efficient+MoE+Inference+on+Personal+Machines+with+Sparsity-Aware+Expert+Cache
12. A Survey on Inference Optimization Techniques for Mixture of Experts Models — Liu, J., Tang, P., Wang, W., Ren, Y., Hou, X., Heng, P.A., Guo, M., Li, C., 2026
https://scholar.google.com/scholar?q=A+Survey+on+Inference+Optimization+Techniques+for+Mixture+of+Experts+Models
Interactive Visualization: Cost-Aware Speculative Decoding for Mixture-of-Experts Models