Charles Frye (Modal, ex-Weights & Biases, Berkeley PhD) joins Ravid and Allen to explain why modern AI research is bottlenecked by compute, and why simply buying more GPUs doesn't solve it. We cover the three problems every lab hits (underutilization, saturation, resource sharing), when companies should actually train their own models, why inference is a "bad algorithm" for today's hardware, NVIDIA's monopoly, the OpenAI/Hugging Face hack and what it says about open models, and whether we're in a compute bubble.
Key topics
AI infrastructure challenges and when to train your own models
GPU resource management and virtualization
Inference optimization and speculative decoding
The economics and future of AI hardware
Agents, sandboxing, and open-model security
Chapters 00:00 Intro 01:03 Why AI needs special-purpose compute 03:22 Buying vs renting GPUs: the three problems 07:15 Modal's approach, and doing more with less compute 09:46 Do we actually need to spend more? The conflict-of-interest question 13:08 Should companies train their own models? 14:47 Efficient fine-tuning and prompts as fast weights 17:37 Are we in a compute bubble? 20:21 Why inference will dominate compute (the SQLite analogy) 22:42 Speculative decoding 26:44 Why scaling inference is hard, and neuromorphic hardware 28:36 Why NVIDIA's monopoly persists 33:09 Inference chip startups and the hardware lottery 35:24 How Modal stays hardware-agnostic (GPU snapshot restore) 38:45 Will agentic coding erode CUDA's moat? 41:18 Running one agent vs thousands: sandboxing at scale 46:27 The OpenAI/Hugging Face hack and open models as defenders 52:28 Rogue AI, self-replication, and fast takeoff 56:09 What's next: evals, embodiment, edge inference 1:00:27 Modal is hiring (modal.jobs)
Music
"Kid Kodi" - Blue Dot Sessions - via Free Music Archive - CC BY-NC 4.0
Podden och tillhörande omslagsbild på den här sidan tillhör
Ravid Shwartz-Ziv & Allen Roush. Innehållet i podden är skapat av Ravid Shwartz-Ziv & Allen Roush och inte av,
eller tillsammans med, Poddtoppen.