Least Squares Temporal Difference Methods: An Analysis under General Conditions
Name
Yu-2012-LEAST SQUARES TEMPORAL DIFFERENCE METHODS.pdf
Size
420.44 KB
Format
Adobe PDF
Checksum (MD5)
e17208832af7cab00ba1ace71ae045c7
Author(s)
Yu, Huizhen
Date Issued
December 2012
Journal
SIAM Journal on Control and Optimization
Publisher
Society for Industrial and Applied Mathematics
Citation
Yu, Huizhen. “Least Squares Temporal Difference Methods: An Analysis Under General Conditions.” SIAM Journal on Control and Optimization 50.6 (2012): 3310–3343. © 2012, Society for Industrial and Applied Mathematics
Version
Final published version
Abstract
We consider approximate policy evaluation for finite state and action Markov decision processes (MDP) with the least squares temporal difference (LSTD) algorithm, LSTD($\lambda$), in an exploration-enhanced learning context, where policy costs are computed from observations of a Markov chain different from the one corresponding to the policy under evaluation. We establish for the discounted cost criterion that LSTD($\lambda$) converges almost surely under mild, minimal conditions. We also analyze other properties of the iterates involved in the algorithm, including convergence in mean and boundedness. Our analysis draws on theories of both finite space Markov chains and weak Feller Markov chains on a topological space. Our results can be applied to other temporal difference algorithms and MDP models. As examples, we give a convergence analysis of a TD($\lambda$) algorithm and extensions to MDP with compact state and action spaces, as well as a convergence proof of a new LSTD algorithm with state-dependent $\lambda$-parameters.
MIT Department
Massachusetts Institute of Technology. Laboratory for Information and Decision Systems
Terms of Use
Article is made available in accordance with the publisher's policy and may be subject to US copyright law. Please refer to the publisher's site for terms of use.
Persistent DSpace Link
DOI of Published Version
https://doi.org/10.1137/100807879