Specialization of Vision Representations with Personalized
Synthetic Data
Name
chae-chaenayo-sm-eecs-2025-thesis.pdf
Description
Thesis PDF
Size
18.96 MB
Format
Adobe PDF
Checksum (MD5)
aaae9487ada0f61effca04a9744257a7
Author(s)
Chae, Nayoung (Julia)
Advisor(s)
Beery, Sara
Date Issued
May 2025
Publisher
Massachusetts Institute of Technology
Abstract
Modern vision models excel at general purpose downstream tasks. It is unclear, however, how they may be used for personalized vision tasks, which are both fine-grained and data-scarce. Recent works have successfully applied synthetic data to general-purpose representation learning, while advances in Text-to-Image (T2I) diffusion models have enabled the generation of personalized images from just a few real examples. Here, we explore a potential connection between these ideas, and formalize the challenge of using personalized synthetic data to learn personalized representations, which encode knowledge about an object of interest and may be flexibly applied to any downstream task relating to the target object. We introduce an evaluation suite for this challenge, including reformulations of two existing datasets and a novel dataset explicitly constructed for this purpose, and propose a contrastive learning approach that makes creative use of image generators. We show that our method improves personalized representation learning for diverse downstream tasks, from recognition to segmentation, and analyze characteristics of image generation approaches that are key to this gain.
MIT Department
Massachusetts Institute of Technology. Department of Electrical Engineering and Computer Science
Terms of Use
In Copyright - Educational Use Permitted
Copyright retained by author(s)
Persistent DSpace Link