Causal Inference Under Privacy Constraints
Name
yao-leonyao-phd-idss-2025.pdf
Description
Thesis PDF
Size
3.95 MB
Format
Adobe PDF
Checksum (MD5)
b4e0dcbb29d6830f24782b3edb7475be
Author(s)
Yao, Leon
Advisor(s)
Eckles, Dean
Date Issued
February 2025
Publisher
Massachusetts Institute of Technology
Abstract
Causal inference is an important tool for learning the effects of interventions in observational or experimental settings. It is widely used in many fields such as epidemiology, economics, and political science to find answers like the average treatment effect of a medical procedure or the individual treatment effect of a personalized ad campaign. In commercial applications, the era of big data allows companies to increase their experiment volume, incentivizing them, in turn, to collect more user data. On one hand, large volumes of data are necessary to train generative models like ChatGPT. At the same time, companies’ increasing use of user data has drawn heavy criticism and consumer backlash, incurring legitimate concerns about privacy and consent. As concerns over user data safety and privacy grow, rules and regulations like GDPR change what kinds of data companies and researchers can acquire and how they can analyze the data. The necessity of now performing causal inference under a range of privacy constrants has carved new spaces for research at the intersection of causal inference and privacy. In my thesis, I will be exploring three paradigms for protecting user data — data minimization, differential privacy and synthetic data — and how to perform causal inference techniques under these new privacy regimes.
MIT Department
Massachusetts Institute of Technology. Institute for Data, Systems, and Society
Terms of Use
Attribution-NonCommercial-NoDerivatives 4.0 International (CC BY-NC-ND 4.0)
Copyright retained by author(s)
Persistent DSpace Link