Earlier this year, the paper "Emergent Misalignment" made the rounds on AI x-risk social media for seemingly showing LLMs generalizing from 'misaligned' training data of insecure code to acting comically evil in response to innocuous questions. In this episode, I chat with one of the authors of that paper, Owain Evans, about that research as well as other work he's done to understand the psychology of large language models.
Podden och tillhörande omslagsbild på den här sidan tillhör
Daniel Filan. Innehållet i podden är skapat av Daniel Filan och inte av,
eller tillsammans med, Poddtoppen.