Information-theoretic Algorithms for Model-free
Reinforcement Learning
Name
wu-farrellw-meng-eecs-2023-thesis.pdf
Description
Thesis PDF
Size
423.41 KB
Format
Adobe PDF
Checksum (MD5)
89304e1e860e6db873539876c554f4a7
Author(s)
Wu, Farrell Eldrian S.
Advisor(s)
Farias, Vivek F.
Date Issued
September 2023
Publisher
Massachusetts Institute of Technology
Abstract
In this work, we propose a model-free reinforcement learning algorithm for infinte-horizon, average-reward decision processes where the transition function has a finite yet unknown dependence on history, and where the induced Markov Decision Process is assumed to be weakly communicating. This algorithm combines the Lempel-Ziv (LZ) parsing tree structure for states introduced in [4] together with the optimistic Q-learning approach in [9]. We mathematically analyze the algorithm towards showing sublinear regret, providing major steps towards the proof of such. In doing so, we reduce the proof to showing sub-linearity of a key quantity related to the sum of an uncertainty metric at each step. Simulations of the algorithm will be done in a later work.
MIT Department
Massachusetts Institute of Technology. Department of Electrical Engineering and Computer Science
Terms of Use
Attribution-NonCommercial-NoDerivatives 4.0 International (CC BY-NC-ND 4.0)
Copyright retained by author(s)
Persistent DSpace Link