Undermatching Is a Consequence of Policy Compression
Name
447.full.pdf
Description
Published version
Size
686.68 KB
Format
Adobe PDF
Checksum (MD5)
6d60715af95ead2a5b6363031c277bde
Author(s) •
Bari, Bilal A
Gershman, Samuel J
Date Issued
January 18, 2023
Journal
The Journal of Neuroscience
Publisher
Society for Neuroscience
Citation
Bilal A. Bari, Samuel J. Gershman, Journal of Neuroscience 18 January 2023, 43 (3) 447-457.
Version
Final published version
Abstract
The matching law describes the tendency of agents to match the ratio of choices allocated to the ratio of rewards received when choosing among multiple options (Herrnstein, 1961). Perfect matching, however, is infrequently observed. Instead, agents tend to undermatch or bias choices toward the poorer option. Overmatching, or the tendency to bias choices toward the richer option, is rarely observed. Despite the ubiquity of undermatching, it has received an inadequate normative justification. Here, we assume agents not only seek to maximize reward, but also seek to minimize cognitive cost, which we formalize as policy complexity (the mutual information between actions and states of the environment). Policy complexity measures the extent to which the policy of an agent is state dependent. Our theory states that capacity-constrained agents (i.e., agents that must compress their policies to reduce complexity) can only undermatch or perfectly match, but not overmatch, consistent with the empirical evidence. Moreover, using mouse behavioral data (male), we validate a novel prediction about which task conditions exaggerate undermatching. Finally, in patients with Parkinson's disease (male and female), we argue that a reduction in undermatching with higher dopamine levels is consistent with an increased policy complexity.
Terms of Use
Article is made available in accordance with the publisher's policy and may be subject to US copyright law. Please refer to the publisher's site for terms of use.
Persistent DSpace Link
DOI of Published Version
https://doi.org/10.1523/JNEUROSCI.1003-22.2022