This episode explores Tensor Cache, a memory architecture for transformer inference that addresses the tradeoff between unbounded KV cache growth and the "amnesia" problem of sliding-window attention. Rather than deleting evicted tokens, the approach folds them into a fixed-size associative memory matrix using an outer-product mechanism rooted in Schmidhuber's 1992 fast-weight memory concept, combined with a linear-attention identity from Schlag et al. that lets a single matrix multiply approximate attention over everything compressed into it. The discussion details the two-tier design—an exact local ring-buffer cache (L1) paired with a compressed overflow matrix (L2)—and clarifies how it differs from related approaches like mLSTM, Infini-attention, RetNet, Mamba, and importance-based eviction schemes such as H2O and SnapKV. Listeners get a clear picture of the learned, per-head gating and decay mechanisms that control how much compressed memory blends into each layer's output, along with a look at the practical challenge of training this eviction-conditioned system efficiently across batches without simulating token-by-token eviction at every gradient step.

Sources: 1. Tensor Cache: Eviction-conditioned Associative Memory for Transformers — Kabir Swain, Sijie Han, Daniel Karl I. Weidele, Mauro Martino, Antonio Torralba, 2026 http://arxiv.org/abs/2605.22884 2. H2O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models — Zhang, Sheng, Zhou, Chen, Zheng, Cai, Song, Tian, Ré, Barrett, Wang, Chen, 2023 https://scholar.google.com/scholar?q=H2O%3A+Heavy-Hitter+Oracle+for+Efficient+Generative+Inference+of+Large+Language+Models 3. SnapKV: LLM Knows What You Are Looking For Before Generation — Li, Huang, Yang, Venkitesh, Locatelli, Ye, Cai, Lewis, Chen, 2024 https://scholar.google.com/scholar?q=SnapKV%3A+LLM+Knows+What+You+Are+Looking+For+Before+Generation 4. CAOTE: KV Caching through Attention Output Error Based Token Eviction — Goel, Park, Gagrani, Jones, Morse, Langston, Lee, Lott, 2025 https://scholar.google.com/scholar?q=CAOTE%3A+KV+Caching+through+Attention+Output+Error+Based+Token+Eviction 5. Get More with LESS: Synthesizing Recurrence with KV Cache Compression for Efficient LLM Inference — Dong, Yang, Zhang, Wang, Chi, Chen, 2024 https://scholar.google.com/scholar?q=Get+More+with+LESS%3A+Synthesizing+Recurrence+with+KV+Cache+Compression+for+Efficient+LLM+Inference 6. Parallelizing Linear Transformers with the Delta Rule over Sequence Length — Yang, Wang, Zhang, Shen, Kim, 2024 https://scholar.google.com/scholar?q=Parallelizing+Linear+Transformers+with+the+Delta+Rule+over+Sequence+Length 7. CacheGen: KV Cache Compression and Streaming for Fast Large Language Model Serving — Liu, Li, Cheng, Ray, Huang, Zhang, Du, Yao, Lu, Ananthanarayanan, Maire, Hoffmann, Holtzman, Jiang, 2023 https://scholar.google.com/scholar?q=CacheGen%3A+KV+Cache+Compression+and+Streaming+for+Fast+Large+Language+Model+Serving

Interactive Visualization: Tensor Cache: Compressing Evicted Tokens into Fixed-Size Memory

Podden och tillhörande omslagsbild på den här sidan tillhör mcgrof. Innehållet i podden är skapat av mcgrof och inte av, eller tillsammans med, Poddtoppen.