This episode explores a technique called "Still," which compresses transformer key-value caches into a fixed-size representation in a single forward pass. The discussion walks through why KV caches balloon with long-context agents, the two-axis taxonomy of compression methods (selection versus synthesis, per-context versus amortized), and how prior work only amortized selection while synthesis remained slow. Using a Perceiver-based module descended from DeepMind's Flamingo resampler, Still cross-attends into each layer's cache and distills it down, requiring a careful workaround for RoPE position rotations so blended content from different token positions doesn't destabilize. The conversation highlights why this fills a genuine gap — manufacturing new compressed representations rather than just picking survivors — and why that matters for memory-constrained, long-horizon agent workloads. Listeners interested in LLM systems efficiency and the mechanics behind emerging cache-compression techniques will find the design-space walkthrough particularly clarifying.

Sources: 1. KV Cache Compaction Beyond Selection: Introducing Still https://arxiv.org/pdf/2606.07878v1 2. H2O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models — Zhenyu Zhang, Ying Sheng, Tianyi Zhou, Tianlong Chen, Lianmin Zheng, Ruisi Cai, Zhao Song, Yuandong Tian, Christopher Ré, Clark Barrett, Zhangyang Wang, Beidi Chen, 2023 https://scholar.google.com/scholar?q=H2O%3A+Heavy-Hitter+Oracle+for+Efficient+Generative+Inference+of+Large+Language+Models 3. Efficient Streaming Language Models with Attention Sinks — Guangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han, Mike Lewis, 2023 https://scholar.google.com/scholar?q=Efficient+Streaming+Language+Models+with+Attention+Sinks 4. Learning to Compress Prompts with Gist Tokens — Jesse Mu, Xiang Lisa Li, Noah Goodman, 2023 https://scholar.google.com/scholar?q=Learning+to+Compress+Prompts+with+Gist+Tokens 5. Cartridges: Lightweight and General-Purpose Long Context Representations via Self-Study — Sabri Eyuboglu et al., 2025 https://scholar.google.com/scholar?q=Cartridges%3A+Lightweight+and+General-Purpose+Long+Context+Representations+via+Self-Study 6. Set Transformer: A Framework for Attention-based Permutation-Invariant Neural Networks — Juho Lee, Yoonho Lee, Jungtaek Kim, Adam R. Kosiorek, Seungjin Choi, Yee Whye Teh, 2019 https://scholar.google.com/scholar?q=Set+Transformer%3A+A+Framework+for+Attention-based+Permutation-Invariant+Neural+Networks 7. Perceiver: General Perception with Iterative Attention — Andrew Jaegle, Felix Gimeno, Andrew Brock, Andrew Zisserman, Oriol Vinyals, Joao Carreira, 2021 https://scholar.google.com/scholar?q=Perceiver%3A+General+Perception+with+Iterative+Attention 8. Flamingo: a Visual Language Model for Few-Shot Learning — Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, et al., 2022 https://scholar.google.com/scholar?q=Flamingo%3A+a+Visual+Language+Model+for+Few-Shot+Learning 9. BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models — Junnan Li, Dongxu Li, Silvio Savarese, Steven Hoi, 2023 https://scholar.google.com/scholar?q=BLIP-2%3A+Bootstrapping+Language-Image+Pre-training+with+Frozen+Image+Encoders+and+Large+Language+Models 10. Learned structure in cartridges: Keys as shareable routers in self-studied representations — Maurizio Diaz, 2025 https://scholar.google.com/scholar?q=Learned+structure+in+cartridges%3A+Keys+as+shareable+routers+in+self-studied+representations 11. Fast KV compaction via attention matching — Adam Zweiger, Xinghong Fu, Han Guo, Yoon Kim, 2026 https://scholar.google.com/scholar?q=Fast+KV+compaction+via+attention+matching 12. KV-Distill: Nearly lossless learnable context compression for LLMs — Vivek Chari, Guanghui Qin, Benjamin Van Durme, 2025 https://scholar.google.com/scholar?q=KV-Distill%3A+Nearly+lossless+learnable+context+compression+for+LLMs 13. Learning to evict from key-value cache (KVP) — Luca Moschella, Laura Manduchi, Ozan Sener, 2026 https://scholar.google.com/scholar?q=Learning+to+evict+from+key-value+cache+%28KVP%29 14. DeepSeek-V4: Toward highly efficient million-token context intelligence — DeepSeek-AI, 2026 https://scholar.google.com/scholar?q=DeepSeek-V4%3A+Toward+highly+efficient+million-token+context+intelligence

Podden och tillhörande omslagsbild på den här sidan tillhör mcgrof. Innehållet i podden är skapat av mcgrof och inte av, eller tillsammans med, Poddtoppen.