Stable regression: On the power of optimization over randomization in training regression problems
Name
19-408.pdf
Description
Published version
Size
584.47 KB
Format
Unknown
Checksum (MD5)
97b1a1c2d8c5de7497ea6c8f09a974ca
Author(s) •
Bertsimas, Dimitris J
Paskov, Ivan
Date Issued
November 1, 2020
Journal
Journal of Machine Learning Research
Version
Final published version
Abstract
We investigate and ultimately suggest remediation to the widely held belief that the best way to train regression models is via random assignment of our data to training and validation sets. In particular, we show that taking a robust optimization approach, and optimally selecting such training and validation sets, leads to models that not only perform significantly better than their randomly constructed counterparts in terms of prediction error, but more importantly, are considerably more stable in the sense that the standard deviation of the resulting predictions, as well as of the model coefficients, is greatly reduced. Moreover, we show that this optimization approach to training is far more effective at recovering the true support of a given data set, i.e., correctly identifying important features while simultaneously excluding spurious ones. We further compare the robust optimization approach to cross validation and find that optimization continues to have a performance edge albeit smaller. Finally, we show that this optimization approach to training is equivalent to building models that are robust to all subpopulations in the data, and thus in particular are robust to the hardest subpopulation, which leads to interesting domain specific interpretations through the use of optimal classification trees. The proposed robust optimization algorithm is efficient and scales training to essentially any desired size.
MIT Department
Sloan School of Management
Massachusetts Institute of Technology. Operations Research Center
Terms of Use
Creative Commons Attribution 4.0 International license
Persistent DSpace Link