This is not the latest version of this item. The latest version can be found here.
The lingering of gradients: How to reuse gradients over time
Name
7400-the-lingering-of-gradients-how-to-reuse-gradients-over-time.pdf
Description
Published version
Size
1.35 MB
Format
Adobe PDF
Checksum (MD5)
c573e8a7b56a2fe67954687306a29134
Author(s) •
simchi-levi, David
Wang, Xinshang
Date Issued
2018
Journal
Advances in Neural Information Processing Systems
Citation
simchi-levi, David and Wang, Xinshang. 2018. "The lingering of gradients: How to reuse gradients over time." Advances in Neural Information Processing Systems, 2018-December.
Version
Final published version
Abstract
© 2018 Curran Associates Inc..All rights reserved. Classically, the time complexity of a first-order method is estimated by its number of gradient computations. In this paper, we study a more refined complexity by taking into account the “lingering” of gradients: once a gradient is computed at xk, the additional time to compute gradients at xk+1, xk+2, . . . may be reduced. We show how this improves the running time of gradient descent and SVRG. For instance, if the “additional time” scales linearly with respect to the traveled distance, then the “convergence rate” of gradient descent can be improved from 1/T to exp(−T1/3). On the empirical side, we solve a hypothetical revenue management problem on the Yahoo! Front Page Today Module application with 4.6m users to 10−6 error (or 10−12 dual error) using 6 passes of the dataset.
Terms of Use
Article is made available in accordance with the publisher's policy and may be subject to US copyright law. Please refer to the publisher's site for terms of use.
Persistent DSpace Link
DOI of Published Version
https://papers.nips.cc/paper/2018/file/b4288d9c0ec0a1841b3b3728321e7088-Paper.pdf