Learning from Weak Supervision: Theory, Methods, and Applications
Name
lang-hjl-phd-eecs-2025-thesis.pdf
Description
Thesis PDF
Size
6.01 MB
Format
Adobe PDF
Checksum (MD5)
dfff349c92a57e9a6a599c77aefc7ad3
Author(s)
Lang, Hunter
Advisor(s)
Sontag, David A.
Date Issued
May 2025
Publisher
Massachusetts Institute of Technology
Abstract
The growing demand for high-quality labeled data to train machine learning models has driven widespread adoption of weak supervision and synthetic data methods, which use automated models instead of humans for annotation. Large language models (LLMs) have further accelerated this trend because their zero- and few-shot classification performance enables them to serve as effective “synthetic annotators” for various tasks. In practice, the data generated by these weak annotators is imperfect, but it enables the training of strong models. However, theoretical understanding of why training one model on the outputs of another leads to strong performance remains limited, especially when the annotator model exhibits suboptimal performance on the target task. In this thesis, I develop a theoretical framework for learning from weak supervision that captures the key aspects of the problem better than existing approaches in the crowdsourcing and learning-with-noisy-label literature. This framework establishes structural conditions that explain when and why weak supervision can reliably train strong models. Building on these theoretical results, the second part of the thesis introduces methods to improve how models learn from weak supervision and applies these methods to low-labeled-data settings.
MIT Department
Massachusetts Institute of Technology. Department of Electrical Engineering and Computer Science
Terms of Use
Attribution-NonCommercial-NoDerivatives 4.0 International (CC BY-NC-ND 4.0)
Copyright retained by author(s)
Persistent DSpace Link