In this episode, Katherine Forrest and Scott Caravello explore how Anthropic researchers built a tool to observe the newly discovered “global workspace” inside its models, the company’s experiments to assess this workspace’s effects on model behavior, and what the discovery could mean for interpretability and AI safety.

For the sources referenced in this episode, please see the links below:

Anthropic: Verbalizable Representations Form a Global Workspace in Language Models

IBM: What Anthropic’s J-space research means for the future of AI

##

Learn More About Paul, Weiss’s Artificial Intelligence practice: https://www.paulweiss.com/industries/artificial-intelligence

Podden och tillhörande omslagsbild på den här sidan tillhör Paul, Weiss. Innehållet i podden är skapat av Paul, Weiss och inte av, eller tillsammans med, Poddtoppen.