On episode 4 of Lab Notes, Amir Zohrenejad speaks with Junchen Jiang about why KVCache may be better understood as reusable, AI-native data rather than a temporary inference optimization. They explore how LMCache and CacheBlend can reduce redundant computation, move context across distributed inference systems, and help support increasingly complex AI agents. The conversation also covers multimodal workloads, open-source infrastructure, and the future of AI systems research.
Podden och tillhörande omslagsbild på den här sidan tillhör
Heavybit. Innehållet i podden är skapat av Heavybit och inte av,
eller tillsammans med, Poddtoppen.