Mean-Variance Optimization in Markov Decision Processes
Name
C-11-mv-MDP-ICML.pdf
Size
376.92 KB
Format
Adobe PDF
Checksum (MD5)
16a9e11aab3d502c302c1af93f978538
Author(s) •
Mannor, Shie
Tsitsiklis, John N.
Date Issued
June 2011
Journal
Proceedings of the Twenty-Eighth International Conference on Machine Learning, ICML 2011
Publisher
International Machine Learning Society
Citation
Mannor, Shie and John Tsitsiklis. "Mean-Variance Optimization in Markov Decision Processes ." in Twenty-Eighth International Conference on Machine Learning, ICML 2011, Jun. 28-Jul.2, Bellevue, Washington. 2011.
Version
Author's final manuscript
Abstract
We consider finite horizon Markov decision processes under performance measures that involve both the mean and the variance of the cumulative reward. We show that either randomized or history-based policies can improve performance. We prove that the complexity of computing a policy that maximizes the mean reward under a variance constraint is NP-hard for some cases, and strongly NP-hard for others. We finally offer pseudo-polynomial exact and approximation algorithms.
MIT Department
Massachusetts Institute of Technology. Department of Electrical Engineering and Computer Science
Terms of Use
Creative Commons Attribution-Noncommercial-Share Alike 3.0
Persistent DSpace Link
DOI of Published Version
http://www.icml-2011.org/papers.php