This episode examines "Rethinking KV Cache Eviction via a Unified Information-Theoretic Objective," a paper that attempts to derive an optimal approach to KV cache eviction from first principles rather than stacking heuristics. The discussion covers the two dominant camps in cache eviction—attention-pattern-based methods like SnapKV and H2O versus structure-aware methods like KeyDiff and Knorm—and how this paper unifies them under the Information Bottleneck principle, treating the surviving cache as a compression of KV history that must stay maximally informative about future queries. The hosts trace how the authors make an intractable nonlinear problem solvable by substituting a linear-Gaussian approximation, then use statistical leverage scores borrowed from classical numerical linear algebra and D-optimal experimental design to cheaply identify which tokens are irreplaceable. The conversation also digs into the paper's validation methodology, questioning whether the Spearman correlation results in Section 4.1 truly confirm the theory or reflect a built-in structural bias in how the proxies were constructed. It's a compelling listen for anyone interested in why long-context inference is so memory-hungry and whether cache eviction can finally be grounded in real mathematics instead of empirical guesswork.
Sources:
1. Rethinking KV Cache Eviction via a Unified Information-Theoretic Objective — Jiaming Yang, Chenwei Tang, Liangli Zhen, Jiancheng Lv, 2026
http://arxiv.org/abs/2604.25975
2. The Information Bottleneck Method — Naftali Tishby, Fernando C. Pereira, William Bialek, 1999
https://scholar.google.com/scholar?q=The+Information+Bottleneck+Method
3. Deep Learning and the Information Bottleneck Principle — Naftali Tishby, Noga Zaslavsky, 2015
https://scholar.google.com/scholar?q=Deep+Learning+and+the+Information+Bottleneck+Principle
4. Opening the Black Box of Deep Neural Networks via Information — Ravid Shwartz-Ziv, Naftali Tishby, 2017
https://scholar.google.com/scholar?q=Opening+the+Black+Box+of+Deep+Neural+Networks+via+Information
5. Deep Variational Information Bottleneck — Alexander A. Alemi, Ian Fischer, Joshua V. Dillon, Kevin Murphy, 2017
https://scholar.google.com/scholar?q=Deep+Variational+Information+Bottleneck
6. Relative-Error CUR Matrix Decompositions — Petros Drineas, Michael W. Mahoney, S. Muthukrishnan, 2008
https://scholar.google.com/scholar?q=Relative-Error+CUR+Matrix+Decompositions
7. CUR Matrix Decompositions for Improved Data Analysis — Michael W. Mahoney, Petros Drineas, 2009
https://scholar.google.com/scholar?q=CUR+Matrix+Decompositions+for+Improved+Data+Analysis
8. Fast Approximation of Matrix Coherence and Statistical Leverage — Petros Drineas, Malik Magdon-Ismail, Michael W. Mahoney, David P. Woodruff, 2012
https://scholar.google.com/scholar?q=Fast+Approximation+of+Matrix+Coherence+and+Statistical+Leverage
9. Randomized Algorithms for Matrices and Data — Michael W. Mahoney, 2011
https://scholar.google.com/scholar?q=Randomized+Algorithms+for+Matrices+and+Data
10. Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention — Angelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, Francois Fleuret, 2020
https://scholar.google.com/scholar?q=Transformers+are+RNNs%3A+Fast+Autoregressive+Transformers+with+Linear+Attention
11. Infinite Attention: NNGP and NTK for Multi-Head Attention Networks — Jiri Hron, Yasaman Bahri, Jascha Sohl-Dickstein, Roman Novak, 2020
https://scholar.google.com/scholar?q=Infinite+Attention%3A+NNGP+and+NTK+for+Multi-Head+Attention+Networks
12. Rethinking Attention with Performers — Krzysztof Choromanski, Valerii Likhosherstov, David Dohan, Xingyou Song, et al., 2021
https://scholar.google.com/scholar?q=Rethinking+Attention+with+Performers
13. Elements of Information Theory (Gaussian channel capacity results) — Thomas M. Cover, Joy A. Thomas (textbook synthesis of Shannon-era results), textbook; foundational results from 1948 onward
https://scholar.google.com/scholar?q=Elements+of+Information+Theory+%28Gaussian+channel+capacity+results%29
14. Efficient Memory Management for Large Language Model Serving with PagedAttention — Kwon, W., Li, Z., Zhuang, S., et al., 2023
https://scholar.google.com/scholar?q=Efficient+Memory+Management+for+Large+Language+Model+Serving+with+PagedAttention
15. H2O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models — Zhang, Z., Sheng, Y., Zhou, T., et al., 2023
https://scholar.google.com/scholar?q=H2O%3A+Heavy-Hitter+Oracle+for+Efficient+Generative+Inference+of+Large+Language+Models
16. Information Bottleneck for Gaussian Variables — Chechik, G., Globerson, A., Tishby, N., Weiss, Y., 2003
https://scholar.google.com/scholar?q=Information+Bottleneck+for+Gaussian+Variables
Interactive Visualization: KV Cache Eviction Through an Information Bottleneck Lens