Learning to Interpret Language Model Diffs
Name
goel-avichal-meng-eecs-2025-thesis.pdf
Description
Thesis PDF
Size
827.01 KB
Format
Adobe PDF
Checksum (MD5)
c99c3ac1354eeb26ab1ac329b6ee92f5
Author(s)
Goel, Avichal
Advisor(s)
Kim, Yoon
Date Issued
September 2025
Publisher
Massachusetts Institute of Technology
Abstract
Finetuning-induced changes to a model’s weights (a “model diff”) are semantically meaningful but often difficult to interpret. This makes us wonder: can we describe the content of an unknown model diff using natural language? We introduce diff interpretation training, a method that teaches a model describe its own finetuning-induced modifications. Our approach uses synthetic model diffs to train a lightweight adapter, which in turn can be applied to a compatible finetuned model to make it self-describing. Using two simple task settings, we demonstrate that our method can successfully decode model diffs into accurate natural language descriptions.
MIT Department
Massachusetts Institute of Technology. Department of Electrical Engineering and Computer Science
Terms of Use
In Copyright - Educational Use Permitted
Copyright retained by author(s)
Persistent DSpace Link