Unsupervised Compositional Image Decompositionwith Diffusion Models
Name
su-jocelin-meng-eecs-2023-thesis.pdf
Description
Thesis PDF
Size
1.26 MB
Format
Adobe PDF
Checksum (MD5)
364535ced7f54584d9ab7b99f37363f9
Author(s)
Su, Jocelin
Advisor(s)
Tenenbaum, Joshua B.
Date Issued
June 2023
Publisher
Massachusetts Institute of Technology
Abstract
Our visual understanding of the world is factorized and compositional. With just a single observation, we can ascertain both global and local attributes in a scene, such as lighting, weather, and underlying objects. These attributes are highly compositional and can be combined in various ways to create new representations of the world. This paper introduces Decomp Diffusion, an unsupervised method for decomposing images into a set of underlying compositional factors, each represented by a different diffusion model. We demonstrate how each decomposed diffusion model captures a different factor of the scene, ranging from global scene descriptors, (e.g. shadows, foreground, or facial expression) to local scene descriptors (e.g. constituent objects). Furthermore, we show how these inferred factors can be flexibly composed and recombined both within and across different image datasets.
MIT Department
Massachusetts Institute of Technology. Department of Electrical Engineering and Computer Science
Terms of Use
In Copyright - Educational Use Permitted
Copyright retained by author(s)
Persistent DSpace Link