Speech2Face: Learning the Face Behind a Voice
Name
1905.09773.pdf
Description
Submitted version
Size
5.05 MB
Format
Adobe PDF
Checksum (MD5)
fc093b4102d3f644f18739983b04fed3
Author(s) • • • • • •
Oh, Taehyun
Dekel, Tali
Kim, Changil
Mosseri, Inbar
Freeman, William T
Rubinstein, Michael
Matusik, Wojciech
Date Issued
January 2020
Journal
2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition
Publisher
Institute of Electrical and Electronics Engineers (IEEE)
Citation
Oh, Tae-Hyun et al. "Speech2Face: Learning the Face Behind a Voice." 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition, June 2019, Long Beach, California, Institute of Electrical and Electronics Engineers, January 2020. © 2019 IEEE
Version
Original manuscript
Abstract
How much can we infer about a person's looks from the way they speak? In this paper, we study the task of reconstructing a facial image of a person from a short audio recording of that person speaking. We design and train a deep neural network to perform this task using millions of natural Internet/Youtube videos of people speaking. During training, our model learns voice-face correlations that allow it to produce images that capture various physical attributes of the speakers such as age, gender and ethnicity. This is done in a self-supervised manner, by utilizing the natural co-occurrence of faces and speech in Internet videos, without the need to model attributes explicitly. We evaluate and numerically quantify how-and in what manner-our Speech2Face reconstructions, obtained directly from audio, resemble the true face images of the speakers.
MIT Department
Massachusetts Institute of Technology. Computer Science and Artificial Intelligence Laboratory
Massachusetts Institute of Technology. Department of Electrical Engineering and Computer Science
Terms of Use
Creative Commons Attribution-Noncommercial-Share Alike
Persistent DSpace Link
DOI of Published Version
https://doi.org/10.1109/cvpr.2019.00772