This episode explores Patchscopes, a unifying framework from Ghandeharioun et al. (Google Research and Tel Aviv University, ICML 2024) for inspecting hidden representations of language models. Rather than decoding internal states through narrow tools like probing classifiers, logit lens, tuned lens, or activation patching, Patchscopes patches a hidden representation from a source prompt directly into a separate target prompt designed to elicit a plain-language explanation of what it holds. The hosts detail how this single mechanism — defined by prompt, layer, position, and an optional transform — subsumes existing interpretability methods as special-case configurations, including logit lens, tuned lens, causal tracing, and attention knockout. They emphasize that Patchscopes remains strictly read-only inspection, distinct from knowledge editing that alters model weights, and discuss how it overcomes the closed-vocabulary and early-layer failure modes that limit earlier techniques. The conversation makes a compelling case that many interpretability tools researchers already use are really the same underlying operation with different settings, offering listeners a clearer theoretical map of the field.
Sources:
1. Patchscopes: A Unifying Framework for Inspecting Hidden Representations of Language Models — Asma Ghandeharioun, Avi Caciularu, Adam Pearce, Lucas Dixon, Mor Geva, 2024
http://arxiv.org/abs/2401.06102
2. Locating and Editing Factual Associations in GPT — Kevin Meng, David Bau, Alex Andonian, Yonatan Belinkov, 2022
https://scholar.google.com/scholar?q=Locating+and+Editing+Factual+Associations+in+GPT
3. Investigating Gender Bias in Language Models Using Causal Mediation Analysis — Jesse Vig, Sebastian Gehrmann, Yonatan Belinkov, Sharon Qian, Daniel Nevo, Yaron Singer, Stuart Shieber, 2020
https://scholar.google.com/scholar?q=Investigating+Gender+Bias+in+Language+Models+Using+Causal+Mediation+Analysis
4. Interpretability in the Wild: A Circuit for Indirect Object Identification in GPT-2 Small — Kevin Wang, Alexandre Variengien, Arthur Conmy, Buck Shlegeris, Jacob Steinhardt, 2023
https://scholar.google.com/scholar?q=Interpretability+in+the+Wild%3A+A+Circuit+for+Indirect+Object+Identification+in+GPT-2+Small
5. Towards Best Practices of Activation Patching in Language Models: Metrics and Methods — Fred Zhang, Neel Nanda, 2024
https://scholar.google.com/scholar?q=Towards+Best+Practices+of+Activation+Patching+in+Language+Models%3A+Metrics+and+Methods
6. interpreting GPT: the logit lens — nostalgebraist, 2020
https://scholar.google.com/scholar?q=interpreting+GPT%3A+the+logit+lens
7. Eliciting Latent Predictions from Transformers with the Tuned Lens — Nora Belrose, Zach Furman, Logan Smith, Danny Halawi, Igor Ostrovsky, Lev McKinney, Stella Biderman, Jacob Steinhardt, 2023
https://scholar.google.com/scholar?q=Eliciting+Latent+Predictions+from+Transformers+with+the+Tuned+Lens
8. Analyzing Transformers in Embedding Space — Guy Dar, Mor Geva, Ankit Gupta, Jonathan Berant, 2023
https://scholar.google.com/scholar?q=Analyzing+Transformers+in+Embedding+Space
9. Jump to Conclusions: Short-Cutting Transformers With Linear Transformations — Alexander Yom Din, Taelin Karidi, Leshem Choshen, Mor Geva, 2023
https://scholar.google.com/scholar?q=Jump+to+Conclusions%3A+Short-Cutting+Transformers+With+Linear+Transformations
10. Locating and Editing Factual Associations in GPT (ROME) — Meng, Bau, Andonian, Belinkov, 2022
https://scholar.google.com/scholar?q=Locating+and+Editing+Factual+Associations+in+GPT+%28ROME%29
11. Linearity of Relation Decoding in Transformer Language Models (LRE) — Hernandez, Sharma, Haklay, Meng, Wattenberg, Andreas, Belinkov, Bau, 2023
https://scholar.google.com/scholar?q=Linearity+of+Relation+Decoding+in+Transformer+Language+Models+%28LRE%29
12. Does Localization Inform Editing? Surprising Differences in Causality-Based Localization vs. Knowledge Editing in Language Models — Hase, Bansal, Kim, Ghandeharioun, 2023
https://scholar.google.com/scholar?q=Does+Localization+Inform+Editing%3F+Surprising+Differences+in+Causality-Based+Localization+vs.+Knowledge+Editing+in+Language+Models
13. The Expressive Power of Transformers with Chain of Thought — Merrill, Sabharwal, 2024
https://scholar.google.com/scholar?q=The+Expressive+Power+of+Transformers+with+Chain+of+Thought
14. Understanding and Patching Compositional Reasoning in LLMs — Li, Jiang, Xie, Song, Lian, Wei, 2024
https://scholar.google.com/scholar?q=Understanding+and+Patching+Compositional+Reasoning+in+LLMs
Interactive Visualization: How Patchscopes Reveals What Language Models Really Think