Multiplicative Regularization Generalizes Better Than Additive Regularization
Name
CBMM Memo 158.pdf
Size
4.8 MB
Format
Adobe PDF
Checksum (MD5)
c24ca329b12192aa7ea33dd62bf369f3
Author(s) • •
Dubach, Rafael
Abdallah, Mohamed S.
Poggio, Tomaso
Date Issued
July 2, 2025
Publisher
Center for Brains, Minds and Machines (CBMM)
Series/Report no.
CBMM Memo;158
Abstract
We investigate the effectiveness of multiplicative versus additive (L2) regularization in deep neural networks, focusing on convolutional neural networks for classification. While additive methods constrain the sum of squared weights, multiplicative regularization directly penalizes the product of layerwise Frobenius norms, a quantity theoretically linked to tighter Rademacher-based generalization bounds. Through experiments on binary classification tasks in a controlled setup, we observe that multiplicative regularization consistently yields wider margin distributions, stronger rank suppression in deeper layers, and improved robustness to label noise. Under 20% label corruption, multiplicative regularization preserves margins that are 5.2% higher and achieves 3.59% higher accuracy compared to additive regularization in our main network architecture. Furthermore, multiplicative regularization achieves a 3.53% boost in test performance for multiclass classification compared to additive regularization. Our analysis of training dynamics shows that directly constraining the global product of norms leads to flatter loss landscapes that correlate with greater resilience to overfitting. These findings highlight the practical benefits of multiplicative penalties for improving generalization and stability in deep models.
Persistent DSpace Link