Interpreting and Editing Memory in Large Transformer Language Models
Name
meng-mengk-meng-eecs-2024-thesis.pdf
Description
Thesis PDF
Size
12.25 MB
Format
Adobe PDF
Checksum (MD5)
70cfa169305d32865dcf6fecbb852ec6
Author(s)
Meng, Kevin
Advisor(s)
Andreas, Jacob D.
Date Issued
May 2024
Publisher
Massachusetts Institute of Technology
Abstract
This thesis investigates the mechanisms of factual recall in large language models. We first apply causal interventions to identify neuron activations that are decisive in a model’s factual predictions; surprisingly, we find that factual recall corresponds to a sparse, localizable computation in the MLP weights of the GPT models we study. Harnessing this insight, we then develop methods for efficiently and surgically inserting up to 10,000 new memories into a transformer; these methods perform well in terms of both generalization and specificity. We conclude with some directions for future work.
MIT Department
Massachusetts Institute of Technology. Department of Electrical Engineering and Computer Science
Terms of Use
Attribution-NonCommercial-NoDerivatives 4.0 International (CC BY-NC-ND 4.0)
Copyright retained by author(s)
Persistent DSpace Link