Junchen Jiang is the Co-Founder and CEO of Tensormesh, an AI infrastructure company building the first commercial platform for KV cache-accelerated LLM inference. ~~~~~~~~~~~~~This episode is brought to you by Nebius — the ultimate cloud for AI innovators.Nebius provides AI infrastructure you can count on, combining reliability and speed with flexibility and engineering support unmatched by hyperscalers.AI leaders like Meta, Shopify, and Higgsfield already partner with Nebius to run their AI workloads. Plus, venture-backed startups can save up to $150,000 on compute costs when they apply for access. Visit nebius.com or nebius.com/startups to learn more~~~~~~~~~~~~~Built on years of research at the University of Chicago, UC Berkeley, and Carnegie Mellon, Tensormesh helps organizations reduce AI inference latency and GPU costs by up to 10x while keeping models and data on their own infrastructure. The company emerged from stealth with $4.5 million in seed funding led by Laude Ventures. Junchen is also an Associate Professor of Computer Science at the University of Chicago, where he leads research on large-scale AI systems and co-created LMCache and CacheBlend—breakthrough technologies that have become foundational components of modern LLM infrastructure. His work earned the ACM EuroSys 2025 Best Paper Award and has been adopted by organizations including Bloomberg, Red Hat, Redis, WEKA, and Tencent. He is also a recipient of the Google Faculty Award and the Carnegie Mellon Best Computer Science Ph.D. Dissertation Award.Topics
Why LLM Inference, Not Training, Is Becoming the Biggest Bottleneck in Enterprise AI
How KV Cache Optimization Can Reduce AI Latency and GPU Costs by Up to 10x
Building the Next Generation of Open-Source AI Infrastructure: From LMCache Research to Tensormesh
Podden och tillhörande omslagsbild på den här sidan tillhör
Grace Gong. Innehållet i podden är skapat av Grace Gong och inte av,
eller tillsammans med, Poddtoppen.