Fast State Restoration in LLM Serving with HCache tackles a hidden cost of running LLM chat services: when GPU memory pressure forces eviction of a conversation's KV cache, restoring that state currently means either recomputing it from scratch (20-26x slower than no restoration) or streaming the full cache back from storage over PCIe (6.5-13x slower). Drawing on traces from ShareGPT4 and L-Eval, the discussion lays out why eviction is the common case rather than an edge case — a single A100-40GB holds only enough KV cache for a handful of live conversations at once. The episode walks through the researchers' proposed middle path: caching the hidden state (one layer upstream of the key/value projection) instead of the KV cache itself, then reconstructing K and V on demand via a cheap matrix multiplication. It's a systems paper grounded in first-principles reasoning about transformer architecture before any benchmarks are run, making the case for why this approach should be faster on theoretical grounds alone. Listeners interested in the practical engineering trade-offs behind serving long, multi-turn LLM conversations at scale will find the framing of "recompute vs. offload vs. something smaller in between" a clear lens on a problem most users never realize is happening.

Sources: 1. Fast State Restoration in LLM Serving with HCache — Shiwei Gao, Youmin Chen, Jiwu Shu, 2024 http://arxiv.org/abs/2410.05004 2. Training Deep Nets with Sublinear Memory Cost — Tianqi Chen, Bin Xu, Chiyuan Zhang, Carlos Guestrin, 2016 https://scholar.google.com/scholar?q=Training+Deep+Nets+with+Sublinear+Memory+Cost 3. Efficient Memory Management for Large Language Model Serving with PagedAttention — Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, et al., 2023 https://scholar.google.com/scholar?q=Efficient+Memory+Management+for+Large+Language+Model+Serving+with+PagedAttention 4. Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving — Ruoyu Qin, Zheming Li, Weiran He, et al. (Moonshot AI and Tsinghua University), 2024/2025 https://scholar.google.com/scholar?q=Mooncake%3A+A+KVCache-centric+Disaggregated+Architecture+for+LLM+Serving 5. CacheGen: KV Cache Compression and Streaming for Fast Large Language Model Serving — Yuhan Liu, Junchen Jiang, et al. (University of Chicago), 2024 https://scholar.google.com/scholar?q=CacheGen%3A+KV+Cache+Compression+and+Streaming+for+Fast+Large+Language+Model+Serving 6. CacheBlend: Fast Large Language Model Serving for RAG with Cached Knowledge Fusion — Jiayi Yao, Hanchen Li, Yuhan Liu, et al., 2025 https://scholar.google.com/scholar?q=CacheBlend%3A+Fast+Large+Language+Model+Serving+for+RAG+with+Cached+Knowledge+Fusion 7. SGLang: Efficient Execution of Structured Language Model Programs — Lianmin Zheng, Liangsheng Yin, Ying Sheng, et al., 2024 https://scholar.google.com/scholar?q=SGLang%3A+Efficient+Execution+of+Structured+Language+Model+Programs 8. Cost-Efficient Large Language Model Serving for Multi-turn Conversations with CachedAttention — Bin Gao, Zhuomin He, Puru Sharma, Qingxuan Kang, Djordje Jevdjic, Junbo Deng, Xingkun Yang, Zhou Yu, Pengfei Zuo, 2024 (USENIX ATC) https://scholar.google.com/scholar?q=Cost-Efficient+Large+Language+Model+Serving+for+Multi-turn+Conversations+with+CachedAttention 9. GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints — Joshua Ainslie, James Lee-Thorp, Michiel de Jong, Yury Zemlyanskiy, Federico Lebron, Sumit Sanghai, 2023 (EMNLP) https://scholar.google.com/scholar?q=GQA%3A+Training+Generalized+Multi-Query+Transformer+Models+from+Multi-Head+Checkpoints 10. Prompt Cache: Modular Attention Reuse for Low-Latency Inference — In Gim, Guojun Chen, Seung-seob Lee, Nikhil Sarda, Anurag Khandelwal, Lin Zhong, 2024 (MLSys) https://scholar.google.com/scholar?q=Prompt+Cache%3A+Modular+Attention+Reuse+for+Low-Latency+Inference 11. Efficiently Programming Large Language Models using SGLang — Lianmin Zheng, Liangsheng Yin, Zhiqiang Xie, Jeff Huang, Chuyue Sun, Cody Hao Yu, Shiyi Cao, Christos Kozyrakis, Ion Stoica, Joseph E. Gonzalez, Clark Barrett, Ying Sheng, 2023/2024 https://scholar.google.com/scholar?q=Efficiently+Programming+Large+Language+Models+using+SGLang 12. FlexGen: High-Throughput Generative Inference of Large Language Models with a Single GPU — Ying Sheng, Lianmin Zheng, Binhang Yuan, Zhuohan Li, Max Ryabinin, Beidi Chen, Percy Liang, Christopher Re, Ion Stoica, Ce Zhang, 2023 (ICML) https://scholar.google.com/scholar?q=FlexGen%3A+High-Throughput+Generative+Inference+of+Large+Language+Models+with+a+Single+GPU

Interactive Visualization: Fast State Restoration for Evicted LLM KV Caches

Podden och tillhörande omslagsbild på den här sidan tillhör mcgrof. Innehållet i podden är skapat av mcgrof och inte av, eller tillsammans med, Poddtoppen.