Algorithmic aspects of mean–variance optimization in Markov decision processes
Name
tsit mv-MDP-EJOR-rev4_SM.pdf
Size
279.71 KB
Format
Adobe PDF
Checksum (MD5)
46339b2d7b3575c04e3012cdea9c4a5d
Author(s) •
Tsitsiklis, John N
Mannor, Shie
Date Issued
June 2013
Journal
European Journal of Operational Research
Publisher
Elsevier
Citation
Mannor, Shie and Tsitsiklis, John N. “Algorithmic Aspects of Mean–variance Optimization in Markov Decision Processes.” European Journal of Operational Research 231, no. 3 (December 2013): 645–653. © 2013 Elsevier B.V.
Version
Author's final manuscript
Abstract
We consider finite horizon Markov decision processes under performance measures that involve both the mean and the variance of the cumulative reward. We show that either randomized or history-based policies can improve performance. We prove that the complexity of computing a policy that maximizes the mean reward under a variance constraint is NP-hard for some cases, and strongly NP-hard for others. We finally offer pseudopolynomial exact and approximation algorithms.
MIT Department
Massachusetts Institute of Technology. Department of Electrical Engineering and Computer Science
Massachusetts Institute of Technology. Laboratory for Information and Decision Systems
Terms of Use
Creative Commons Attribution-NonCommercial-NoDerivs License
Persistent DSpace Link
DOI of Published Version
https://doi.org/10.1016/j.ejor.2013.06.019