Active Reward Learning for Co-Robotic Vision Based Exploration in Bandwidth Limited Environments
Name
2003.05016.pdf
Description
Accepted version
Size
4.8 MB
Format
Unknown
Checksum (MD5)
65adc9315793462b5c145e9feafe1c73
Author(s) • •
Jamieson, Stewart Christopher.
How, Jonathan P
Girdhar, Yogesh
Date Issued
May 2020
Journal
Proceedings - IEEE International Conference on Robotics and Automation
Publisher
IEEE
Citation
2020. "Active Reward Learning for Co-Robotic Vision Based Exploration in Bandwidth Limited Environments." Proceedings - IEEE International Conference on Robotics and Automation.
Version
Author's final manuscript
Abstract
© 2020 IEEE. We present a novel POMDP problem formulation for a robot that must autonomously decide where to go to collect new and scientifically relevant images given a limited ability to communicate with its human operator. From this formulation we derive constraints and design principles for the observation model, reward model, and communication strategy of such a robot, exploring techniques to deal with the very high-dimensional observation space and scarcity of relevant training data. We introduce a novel active reward learning strategy based on making queries to help the robot minimize path regret online, and evaluate it for suitability in autonomous visual exploration through simulations. We demonstrate that, in some bandwidth-limited environments, this novel regret-based criterion enables the robotic explorer to collect up to 17% more reward per mission than the next-best criterion.
MIT Department
Joint Program in Applied Ocean Physics and Engineering
Woods Hole Oceanographic Institution
Terms of Use
Creative Commons Attribution-Noncommercial-Share Alike
Persistent DSpace Link
DOI of Published Version
https://doi.org/10.1109/ICRA40945.2020.9196922