On the Optimal Regret of Locally Private Linear
Contextual Bandit
Name
li-jiach334-sm-idss-26-thesis.pdf
Size
795.49 KB
Format
Adobe PDF
Checksum (MD5)
e917fac8fd199cdd551594ee3c92f376
Author(s)
Li, Jiachun
Advisor(s)
Simchi-Levi, David
Date Issued
February 2026
Publisher
Massachusetts Institute of Technology
Abstract
Contextual bandit with linear reward functions is among one of the most extensively studied models in bandit and online learning research. Recently, there has been increasing interest in designing locally private linear contextual bandit algorithms, where sensitive information contained in contexts and rewards is protected against leakage to the general public. While the classical linear contextual bandit algorithm admits cumulative regret upper bounds of Oe(√ T) via multiple alternative methods, it has remained open whether such regret bounds are attainable in the presence of local privacy constraints, with the state-of-the-art result being Oe(T ³/⁴). In this paper, we show that it is indeed possible to achieve an O(√ T) regret upper bound for locally private linear contextual bandit. Our solution relies on several new algorithmic and analytical ideas, such as the analysis of mean absolute deviation errors and layered principal component regression in order to achieve small mean absolute deviation errors.
MIT Department
Massachusetts Institute of Technology. Institute for Data, Systems, and Society
Terms of Use
In Copyright - Educational Use Permitted
Copyright retained by author(s)
Persistent DSpace Link