Next Level BizTech
Avsnitt

Ep.224- The AI Inference Flip — Why Speed, Cost, and “Everywhere” Just Won- with Josh Lupresto

Dela

The AI race just flipped. For years it was all about building bigger models. Now the center of gravity has moved to serving those models fast, cheap, and everywhere. Josh Lupresto, SVP of Sales Engineering, breaks down the “Inference Flip” in plain English: what training vs. inference really mean, why today’s models “think” before they answer, and how that shift pushes the hard problem—and the money—into real‑time inference.

What you’ll learn

  • Training vs. inference, explained simply (write the textbook once, read it billions of times)
  • Why letting models “think” increases tokens 10–100x at answer time—and shifts cost to inference
  • The memory wall: why moving data, not math, bottlenecks GPUs
  • New hardware built for inference: wafer‑scale systems (e.g., Cerebras), LPUs (e.g., Groq), and system‑level designs
  • What changes in the real world: costs fall, latency drops, and AI moves to the edge (“physical AI”)
  • Where trusted advisors win as customers face more diverse AI deployment choices

Practical takeaways

  • Real-time use cases become feasible: voice agents without awkward lag, live translation, multi‑step agents, instant fraud checks
  • Edge/onsite AI gets viable for latency, cost, or data‑residency needs
  • Reopen “shelved” AI projects—math and performance assumptions may have flipped


Podden och tillhörande omslagsbild på den här sidan tillhör Telarus - Josh Lupresto. Innehållet i podden är skapat av Telarus - Josh Lupresto och inte av, eller tillsammans med, Poddtoppen.