Bayesian Policy Search with Policy Priors
Name
Kaelbling_Bayesian policy.pdf
Size
173.76 KB
Format
Adobe PDF
Checksum (MD5)
a50a9391d4e33fe8c73834d2348d2783
Author(s) • • • •
Wingate, David
Goodman, Noah D.
Roy, Daniel M.
Kaelbling, Leslie P.
Tenenbaum, Joshua B.
Date Issued
July 2011
Journal
Proceedings of the Twenty-Second International Joint Conference on Artificial Intelligence
Publisher
International Joint Conference on Artificial Intelligence (IJCAI)
Citation
Wingate, David, Noah D. Goodman, Daniel M. Roy, Leslie P. Kaelbling, and Joshua B. Tenenbaum. "Bayesian Policy Search with Policy Priors." Proceedings of the Twenty-Second International Joint Conference on Artificial Intelligence, July 16-22, 2011, Barcelona, Spain.
Version
Author's final manuscript
Abstract
We consider the problem of learning to act in partially observable, continuous-state-and-action worlds where we have abstract prior knowledge about the structure of the optimal policy in the form of a distribution over policies. Using ideas from planning-as-inference reductions and Bayesian unsupervised learning, we cast Markov Chain Monte Carlo as a stochastic, hill-climbing policy search algorithm. Importantly, this algorithm’s search bias is directly tied to the prior and its MCMC proposal kernels, which means we can draw on the full Bayesian toolbox to express the search bias, including nonparametric priors and structured, recursive processes like grammars over action sequences. Furthermore, we can reason about uncertainty in the search bias itself by constructing a hierarchical prior and reasoning about latent variables that determine the abstract structure of the policy. This yields an adaptive search algorithm—our algorithm learns to learn a structured policy efficiently. We show how inference over the latent variables in these policy priors enables intra- and intertask transfer of abstract knowledge. We demonstrate the flexibility of this approach by learning meta search biases, by constructing a nonparametric finite state controller to model memory, by discovering motor primitives using a simple grammar over primitive actions, and by combining all three.
MIT Department
Massachusetts Institute of Technology. Computer Science and Artificial Intelligence Laboratory
Massachusetts Institute of Technology. Department of Brain and Cognitive Sciences
Massachusetts Institute of Technology. Department of Electrical Engineering and Computer Science
Massachusetts Institute of Technology. Laboratory for Information and Decision Systems
Terms of Use
Creative Commons Attribution-Noncommercial-Share Alike
Persistent DSpace Link
DOI of Published Version
https://doi.org/10.5591/978-1-57735-516-8/IJCAI11-263