The Rate of Convergence of AdaBoost
Name
Mukherjee-2013-Rate of convergence.pdf
Size
299.06 KB
Format
Adobe PDF
Checksum (MD5)
8b16d4ba765e746d88bb6e88f59f18cb
Author(s) • •
Mukherjee, Indraneel
Rudin, Cynthia
Schapire, Robert E.
Date Issued
August 2013
Journal
Journal of Machine Learning Research
Publisher
Association for Computing Machinery (ACM)
Citation
Mukherjee, Indraneel, Cynthia Rudin, and Robert E. Schapire. “The Rate of Convergence of AdaBoost.” Journal of Machine Learning Research 14 (2013): 2315–2347.
Version
Final published version
Abstract
The AdaBoost algorithm was designed to combine many “weak” hypotheses that perform slightly better than random guessing into a “strong” hypothesis that has very low error. We study the rate at which AdaBoost iteratively converges to the minimum of the “exponential loss”. Unlike previous work, our proofs do not require a weak-learning assumption, nor do they require that minimizers of the exponential loss are finite. Our first result shows that the exponential loss of AdaBoost's computed parameter vector will be at most ε more than that of any parameter vector of ℓ[subscript 1]-norm bounded by B in a number of rounds that is at most a polynomial in B and 1/ε. We also provide lower bounds showing that a polynomial dependence is necessary. Our second result is that within C/ε iterations, AdaBoost achieves a value of the exponential loss that is at most ε more than the best possible value, where C depends on the data set. We show that this dependence of the rate on ε is optimal up to constant factors, that is, at least Ω(1/ε) rounds are necessary to achieve within ε of the optimal exponential loss.
MIT Department
Sloan School of Management
Terms of Use
Article is made available in accordance with the publisher's policy and may be subject to US copyright law. Please refer to the publisher's site for terms of use.
Persistent DSpace Link
DOI of Published Version
http://jmlr.org/papers/v14/mukherjee13b.html