Complexity control by gradient descent in deep networks
Name
s41467-020-14663-9.pdf
Description
Published version
Size
431.68 KB
Format
Unknown
Checksum (MD5)
cb2bfa9216379d5f6becece6d420b4e7
Author(s) • •
Poggio, Tomaso A
Liao, Qianli
Banburski, Andrzej
Date Issued
2020
Journal
Nature Communications
Publisher
Springer Science and Business Media LLC
Version
Final published version
Abstract
© 2020, The Author(s). Overparametrized deep networks predict well, despite the lack of an explicit complexity control during training, such as an explicit regularization term. For exponential-type loss functions, we solve this puzzle by showing an effective regularization effect of gradient descent in terms of the normalized weights that are relevant for classification.
MIT Department
Center for Brains, Minds, and Machines
Terms of Use
Creative Commons Attribution 4.0 International license
Persistent DSpace Link
DOI of Published Version
https://doi.org/10.1038/S41467-020-14663-9