This episode examines a black-box audit method for probing what large language models associate with a given name, applied across eight models including GPT-4o, GPT-5, Grok-3, and several locally-run open models. The researchers built WikiMem-style completion probing that works without token probabilities, testing models against 100 famous public figures and 100 invented synthetic names to isolate genuine memorization from guesswork, and uncovered failure patterns like "default token collapse" (models reflexively answering "ambidextrous" or "+1" regardless of the actual person) and base-rate anchoring on attributes like victim counts. The hosts dig into a four-category framework — direct, indirect, inferred, and guessed data — and debate how alarming it really is that GPT-4o hit 60%+ accuracy on several personal attributes for ordinary, non-famous individuals. A companion tool, LMP2, lets anyone query what a model associates with their own name, and survey results from 155 participants reveal a gap between which attributes people fear exposing (financial data, phone numbers, medical conditions) and which ones models actually get right. Listeners interested in AI privacy risks, model auditing methodology, or the gap between perceived and actual data exposure will find plenty to chew on here.
Sources:
1. What Do LLMs Associate with Your Name? A Human-Centered Black-Box Audit of Personal Data — Dimitri Staufer, Kirsten Morehouse, 2026
http://arxiv.org/abs/2602.17483
2. Membership Inference Attacks Against Machine Learning Models — Reza Shokri, Marco Stronati, Congzheng Song, Vitaly Shmatikov, 2017
https://scholar.google.com/scholar?q=Membership+Inference+Attacks+Against+Machine+Learning+Models
3. Auditing Differentially Private Machine Learning: How Private is Private SGD? — Matthew Jagielski, Jonathan Ullman, Alina Oprea, 2020
https://scholar.google.com/scholar?q=Auditing+Differentially+Private+Machine+Learning%3A+How+Private+is+Private+SGD%3F
4. Extracting Training Data from Large Language Models — Nicholas Carlini, Florian Tramèr, Eric Wallace, Matthew Jagielski, et al., 2021
https://scholar.google.com/scholar?q=Extracting+Training+Data+from+Large+Language+Models
5. Deduplicating Training Data Makes Language Models Better — Katherine Lee, Daphne Ippolito, Andrew Nystrom, Chiyuan Zhang, Douglas Eck, Chris Callison-Burch, Nicholas Carlini, 2022
https://scholar.google.com/scholar?q=Deduplicating+Training+Data+Makes+Language+Models+Better
6. Quantifying Memorization Across Neural Language Models — Nicholas Carlini, Daphne Ippolito, Matthew Jagielski, Katherine Lee, Florian Tramèr, Chiyuan Zhang, 2023
https://scholar.google.com/scholar?q=Quantifying+Memorization+Across+Neural+Language+Models
7. Scalable Extraction of Training Data from (Production) Language Models — Milad Nasr, Nicholas Carlini, Jonathan Hayase, Matthew Jagielski, et al., 2023
https://scholar.google.com/scholar?q=Scalable+Extraction+of+Training+Data+from+%28Production%29+Language+Models
8. Beyond Memorization: Violating Privacy Via Inference with Large Language Models — Robin Staab, Mark Vero, Mislav Balunović, Martin Vechev, 2024
https://scholar.google.com/scholar?q=Beyond+Memorization%3A+Violating+Privacy+Via+Inference+with+Large+Language+Models
9. Auditing Algorithms: Research Methods for Detecting Discrimination on Internet Platforms — Christian Sandvig, Kevin Hamilton, Karrie Karahalios, Cedric Langbort, 2014
https://scholar.google.com/scholar?q=Auditing+Algorithms%3A+Research+Methods+for+Detecting+Discrimination+on+Internet+Platforms
10. Auditing Algorithms: Understanding Algorithmic Systems from the Outside In — Danaë Metaxa, Joon Sung Park, Ronald E. Robertson, Karrie Karahalios, Christo Wilson, Jeff Hancock, Christian Sandvig, 2021
https://scholar.google.com/scholar?q=Auditing+Algorithms%3A+Understanding+Algorithmic+Systems+from+the+Outside+In
11. WikiMem probing framework paper (cited as [72]) — Not fully given in excerpt; referenced throughout as 'WikiMem', 2025/2026 (cited as prior work)
https://scholar.google.com/scholar?q=WikiMem+probing+framework+paper+%28cited+as+%5B72%5D%29
12. The Secret Sharer: Evaluating and Testing Unintended Memorization in Neural Networks — Nicholas Carlini, Chang Liu, Úlfar Erlingsson, Jernej Kos, Dawn Song, 2019 (USENIX Security 19)
https://scholar.google.com/scholar?q=The+Secret+Sharer%3A+Evaluating+and+Testing+Unintended+Memorization+in+Neural+Networks
13. Minerva (cited as [81]) — Not fully given in excerpt, Not specified
https://scholar.google.com/scholar?q=Minerva+%28cited+as+%5B81%5D%29
14. Attribute inference from seemingly benign inputs (cited as [71]) — Not fully given in excerpt, Not specified
https://scholar.google.com/scholar?q=Attribute+inference+from+seemingly+benign+inputs+%28cited+as+%5B71%5D%29
15. Contextual-integrity study of ChatGPT users' privacy judgments — Tran et al., 2025
https://scholar.google.com/scholar?q=Contextual-integrity+study+of+ChatGPT+users%27+privacy+judgments
16. Dark patterns and design choices amplifying disclosure (cited as [28]) — Gumusel et al., 2025
https://scholar.google.com/scholar?q=Dark+patterns+and+design+choices+amplifying+disclosure+%28cited+as+%5B28%5D%29
17. Tell me something new: data subject rights applied to inferred data and profiles — Bart Custers, Helena Vrabec, 2024
https://scholar.google.com/scholar?q=Tell+me+something+new%3A+data+subject+rights+applied+to+inferred+data+and+profiles
18. Digital Forgetting in Large Language Models: A Survey of Unlearning Methods — Alberto Blanco-Justicia, Najeeb Jebreel, Benet Manzanares-Salor, David Sánchez, et al., 2025
https://scholar.google.com/scholar?q=Digital+Forgetting+in+Large+Language+Models%3A+A+Survey+of+Unlearning+Methods
Interactive Visualization: LLMs Guess Wrong: Auditing What Models Know About You