Decomposing Deep Neural Network Minds into Parts
Name
michaud-ericjm-phd-physics-2026-thesis.pdf
Size
32.38 MB
Format
Adobe PDF
Checksum (MD5)
501e065701d33645048ff6f228881832
Author(s)
Michaud, Eric J.
Advisor(s)
Tegmark, Max
Date Issued
February 2026
Publisher
Massachusetts Institute of Technology
Abstract
We are no longer alone in the universe. With the success of deep learning, we have new kinds of minds in the world that we can communicate with and study. This thesis is about understanding the internal workings of these new minds. Our approach is to decompose the computation that neural networks perform into parts—intelligible mechanisms that are more specialized and less complex than the network as a whole. First, we study how the ways that neural networks learn and scale are influenced by the properties of these mechanisms. We study how the statistics of these mechanisms can produce “neural scaling laws”, how the relationship of these mechanisms to each other causes “curricula”, and how differences in the complexity of these mechanisms and the speed at which they are learned cause “grokking”. We then study the challenge of automatically decomposing neural network computation into atomic parts. We first demonstrate that we can convert small networks into standalone Python programs. Turning to larger networks, we then study how sparse autoencoders can be used to decompose language model activations into features. We find multiple types of geometric structure to language model features, raising interesting questions about whether individual sparse autoencoder latents are the only sensible units of decomposition.
MIT Department
Massachusetts Institute of Technology. Department of Physics
Terms of Use
In Copyright - Educational Use Permitted
Copyright retained by author(s)
Persistent DSpace Link