This episode explores Patchscopes, a unifying framework from Ghandeharioun et al. (Google Research and Tel Aviv University, ICML 2024) for inspecting hidden representations of language models. Rather than decoding internal states through narrow tools like probing classifiers, logit lens, tuned lens, or activation patching, Patchscopes patches a hidden representation from a source prompt directly into a separate target prompt designed to elicit a plain-language explanation of what it holds. The hosts detail how this single mechanism — defined by prompt, layer, position, and an optional transform — subsumes existing interpretability methods as special-case configurations, including logit lens, tuned lens, causal tracing, and attention knockout. They emphasize that Patchscopes remains strictly read-only inspection, distinct from knowledge editing that alters model weights, and discuss how it overcomes the closed-vocabulary and early-layer failure modes that limit earlier techniques. The conversation makes a compelling case that many interpretability tools researchers already use are really the same underlying operation with different settings, offering listeners a clearer theoretical map of the field.

Sources: 1. Patchscopes: A Unifying Framework for Inspecting Hidden Representations of Language Models — Asma Ghandeharioun, Avi Caciularu, Adam Pearce, Lucas Dixon, Mor Geva, 2024 http://arxiv.org/abs/2401.06102 2. Locating and Editing Factual Associations in GPT — Kevin Meng, David Bau, Alex Andonian, Yonatan Belinkov, 2022 https://scholar.google.com/scholar?q=Locating+and+Editing+Factual+Associations+in+GPT 3. Investigating Gender Bias in Language Models Using Causal Mediation Analysis — Jesse Vig, Sebastian Gehrmann, Yonatan Belinkov, Sharon Qian, Daniel Nevo, Yaron Singer, Stuart Shieber, 2020 https://scholar.google.com/scholar?q=Investigating+Gender+Bias+in+Language+Models+Using+Causal+Mediation+Analysis 4. Interpretability in the Wild: A Circuit for Indirect Object Identification in GPT-2 Small — Kevin Wang, Alexandre Variengien, Arthur Conmy, Buck Shlegeris, Jacob Steinhardt, 2023 https://scholar.google.com/scholar?q=Interpretability+in+the+Wild%3A+A+Circuit+for+Indirect+Object+Identification+in+GPT-2+Small 5. Towards Best Practices of Activation Patching in Language Models: Metrics and Methods — Fred Zhang, Neel Nanda, 2024 https://scholar.google.com/scholar?q=Towards+Best+Practices+of+Activation+Patching+in+Language+Models%3A+Metrics+and+Methods 6. interpreting GPT: the logit lens — nostalgebraist, 2020 https://scholar.google.com/scholar?q=interpreting+GPT%3A+the+logit+lens 7. Eliciting Latent Predictions from Transformers with the Tuned Lens — Nora Belrose, Zach Furman, Logan Smith, Danny Halawi, Igor Ostrovsky, Lev McKinney, Stella Biderman, Jacob Steinhardt, 2023 https://scholar.google.com/scholar?q=Eliciting+Latent+Predictions+from+Transformers+with+the+Tuned+Lens 8. Analyzing Transformers in Embedding Space — Guy Dar, Mor Geva, Ankit Gupta, Jonathan Berant, 2023 https://scholar.google.com/scholar?q=Analyzing+Transformers+in+Embedding+Space 9. Jump to Conclusions: Short-Cutting Transformers With Linear Transformations — Alexander Yom Din, Taelin Karidi, Leshem Choshen, Mor Geva, 2023 https://scholar.google.com/scholar?q=Jump+to+Conclusions%3A+Short-Cutting+Transformers+With+Linear+Transformations 10. Locating and Editing Factual Associations in GPT (ROME) — Meng, Bau, Andonian, Belinkov, 2022 https://scholar.google.com/scholar?q=Locating+and+Editing+Factual+Associations+in+GPT+%28ROME%29 11. Linearity of Relation Decoding in Transformer Language Models (LRE) — Hernandez, Sharma, Haklay, Meng, Wattenberg, Andreas, Belinkov, Bau, 2023 https://scholar.google.com/scholar?q=Linearity+of+Relation+Decoding+in+Transformer+Language+Models+%28LRE%29 12. Does Localization Inform Editing? Surprising Differences in Causality-Based Localization vs. Knowledge Editing in Language Models — Hase, Bansal, Kim, Ghandeharioun, 2023 https://scholar.google.com/scholar?q=Does+Localization+Inform+Editing%3F+Surprising+Differences+in+Causality-Based+Localization+vs.+Knowledge+Editing+in+Language+Models 13. The Expressive Power of Transformers with Chain of Thought — Merrill, Sabharwal, 2024 https://scholar.google.com/scholar?q=The+Expressive+Power+of+Transformers+with+Chain+of+Thought 14. Understanding and Patching Compositional Reasoning in LLMs — Li, Jiang, Xie, Song, Lian, Wei, 2024 https://scholar.google.com/scholar?q=Understanding+and+Patching+Compositional+Reasoning+in+LLMs

Interactive Visualization: How Patchscopes Reveals What Language Models Really Think

Podden och tillhörande omslagsbild på den här sidan tillhör mcgrof. Innehållet i podden är skapat av mcgrof och inte av, eller tillsammans med, Poddtoppen.