Towards understanding residual neural networks
Name
1127292128-MIT.pdf
Size
3.34 MB
Format
Adobe PDF
Checksum (MD5)
a4921202f7cc3523c8d78e9c884d4c7f
Author(s)
Zeng, Brandon.
Advisor(s)
Aleksander Ma̧dry.
Date Issued
2019
Publisher
Massachusetts Institute of Technology
Abstract
Residual networks (ResNets) are now a prominent architecture in the field of deep learning. However, an explanation for their success remains elusive. The original view is that residual connections allows for the training of deeper networks, but it is not clear that added layers are always useful, or even how they are used. In this work, we find that residual connections distribute learning behavior across layers, allowing resnets to indeed effectively use deeper layers and outperform standard networks. We support this explanation with results for network gradients and representation learning that show that residual connections make the training of individual residual blocks easier.
Description
Thesis: M. Eng., Massachusetts Institute of Technology, Department of Electrical Engineering and Computer Science, 2019
Cataloged from PDF version of thesis.
Includes bibliographical references (page 37).
Subjects
Electrical Engineering and Computer Science.
MIT Department
Massachusetts Institute of Technology. Department of Electrical Engineering and Computer Science
Terms of Use
MIT theses are protected by copyright. They may be viewed, downloaded, or printed from this source but further reproduction or distribution in any format is prohibited without written permission.
Persistent DSpace Link