Automated Mechanistic Interpretability for Neural Networks
Name
liao-iliao-meng-eecs-2024-thesis.pdf
Description
Thesis PDF
Size
4.58 MB
Format
Adobe PDF
Checksum (MD5)
3f26d393ce2a89ea6f21e7b9a7a90505
Author(s)
Liao, Isaac C.
Advisor(s)
Tegmark, Max
Date Issued
May 2024
Publisher
Massachusetts Institute of Technology
Abstract
Mechanistic interpretability research aims to deconstruct the underlying algorithms that neural networks use to perform computations, such that we can modify their components, causing them to change behavior in predictable and positive ways. This thesis details three novel methods for automating the interpretation process for neural networks that are too large to manually interpret. Firstly, we detect inherently multidimensional representations of data; we discover that large language models use circular representations to perform modular addition tasks. Secondly, we introduce methods to penalize complexity in neural circuitry; we discover the automatic emergence of interpretable properties such as sparsity, weight tying, and circuit duplication. Last but not least, we apply neural network symmetries to put networks into a simplified normal form, for conversion into human-readable python; we introduce a program synthesis benchmark with this and successfully convert 32 out of 62 of them.
MIT Department
Massachusetts Institute of Technology. Department of Electrical Engineering and Computer Science
Terms of Use
Attribution-NonCommercial-NoDerivatives 4.0 International (CC BY-NC-ND 4.0)
Copyright retained by author(s)
Persistent DSpace Link