Testing Population Privacy with Synthetic Data
Name
ouyang-aliciao-sm-tpp-eecs-2026-thesis.pdf
Size
34.61 MB
Format
Adobe PDF
Checksum (MD5)
bc1e6ecd9a305eb1a43759ac0337cbdf
Author(s)
Ouyang, Alicia
Advisor(s)
Raghavan, Manish
Date Issued
February 2026
Publisher
Massachusetts Institute of Technology
Abstract
The U.S. Census decennial population data is utilized for policy decisions such as representation allocation in Congress and education funding. However, the dataset that is publicly available and assumed to be the ground truth for the U.S. Population, has disclosure avoidance techniques applied. Under current U.S. law, there is no way for anyone outside of the Census Bureau to access the original dataset to test the impact of disclosure avoidance techniques on accurate population numbers. To uncover the privacy-accuracy tradeoff, this thesis attempts to create a complete synthetic U.S. Census dataset and notes the computational difficulties in the chosen process. Accuracy analysis is performed on the available synthetic data to uncover trends among disparate impacts among different racial groups and urban level, and explain the trends based on characteristics of that state’s population on county and block level. Finally, this thesis recommends future work on creating synthetic population datasets, other metrics to test, and a workflow for the Census Bureau to allow testing on its disclosure avoidance techniques by researchers.
MIT Department
Massachusetts Institute of Technology. Department of Electrical Engineering and Computer Science
Massachusetts Institute of Technology. Institute for Data, Systems, and Society
Terms of Use
In Copyright - Educational Use Permitted
Copyright retained by author(s)
Persistent DSpace Link