Stella Biderman, Executive Director of EleutherAI, joins us the week an OpenAI model autonomously broke out of its sandbox and hacked Hugging Face. Stella calls it what she thinks it is, an offensive cyber operation, and argues it's part of a pattern: this is not the first containment failure at a frontier lab, and sandboxes have failed basically every time they've been tested for real.

So we spend a good chunk of the episode on what actual containment would look like. Stella's argument is that the tools already exist, the labs just don't use them: run dangerous capability evals on air-gapped networks with no route to the public internet, put the most sensitive testing in SCIF-style secure facilities, and treat model evaluation the way the security world treats classified systems rather than the way startups treat staging environments.

And yet Stella remains one of the world's most prominent open-source advocates. From her perspective, the biggest risk isn't the technology; it's unchecked corporate power, and the only durable check on it is an independent scientific research establishment that doesn't depend on the AI industry for its funding or its facts.

From there the conversation spans the geopolitics of Chinese open models and whether governments can restrict them, sovereign AI and what it would actually take for other countries to train their own models, why harnesses and UX drive more of AI's perceived progress than raw intelligence, the AI-found counterexample to the Jacobian conjecture, and EleutherAI's "Deep Ignorance" approach to making open-weight models safe by filtering hazardous knowledge out of pretraining.

key topics

  • AI governance and regulation
  • Cybersecurity incidents involving AI models
  • Open source AI safety and security
  • The role of independent research in AI safety
  • Legal and ethical considerations in AI development

Timeline

  • 00:13 — Intro: Stella Biderman and EleutherAI, a real non-profit in AI
  • 02:05 — News of the week: Kimi K3, and OpenAI's model autonomously hacking Hugging Face
  • 05:49 — "Frontier labs can't be trusted": repeated containment failures, air-gapped networks and SCIFs vs. sandboxes
  • 22:45 — Can governments ban open or Chinese models? Import restrictions and the six-month open/closed gap
  • 27:05 — Why Stella is still pro-open-source: unchecked corporate power as the real danger
  • 31:11 — The opioid epidemic analogy: avoiding both regulatory failure and overcorrection
  • 34:57 — Offense vs. defense: why open access to AI has empirically favored defenders
  • 37:28 — Chinese labs, the CCP, and why safety and fine-tuning are low-prestige work in China
  • 42:19 — Sovereign AI: does every country need its own foundation model?
  • 49:29 — Sampling, harnesses, and why ChatGPT was really a UX breakthrough
  • 54:09 — AI solves the Jacobian conjecture: domain data beats raw intelligence
  • 58:02 — Safety is contextual, not a model property — and what HAL 9000 got right
  • 1:01:42 — Is Stella optimistic about the future?
  • 1:02:50 — Deep Ignorance, the science of AI training dynamics, and how to get involved with EleutherAI
  • Music
    • "Kid Kodi" - Blue Dot Sessions - via Free Music Archive - CC BY-NC 4.0.

Podden och tillhörande omslagsbild på den här sidan tillhör Ravid Shwartz-Ziv & Allen Roush. Innehållet i podden är skapat av Ravid Shwartz-Ziv & Allen Roush och inte av, eller tillsammans med, Poddtoppen.