Learning Audio-Video Language Representations
Name
Rouditchenko-roudi-meng-eecs-2021-thesis.pdf
Description
Thesis PDF
Size
13.69 MB
Format
Adobe PDF
Checksum (MD5)
d340e5c82435dc6cc3e6a42b93791d7c
Author(s)
Rouditchenko, Andrew
Advisor(s)
Glass, James
Harwath, David
Date Issued
June 2021
Publisher
Massachusetts Institute of Technology
Abstract
Automatic speech recognition has seen recent advancements powered by machine learning, but it is still only available for a small fraction of the more than 7,000 languages spoken worldwide due to the reliance on manually annotated speech data. Unlabeled multi-modal data, such as videos, are now increasingly available in many different languages and provide opportunities to scale speech technologies. In this thesis, we introduce models and datasets for learning visually grounded spoken language from raw audio in videos. We propose a self-supervised audio-video model that learns from the English narration naturally present in instructional videos to relate spoken words and sounds to visual content. Our model can recognize spoken words and natural sounds in audio queries to retrieve relevant visual clips, supporting its application to video search directly using audio and spoken queries, without needing to transcribe speech to text. We further demonstrate that our model can learn multilingual audiovideo representations and can successfully perform retrieval on Japanese videos. Since our approach only requires audio-visual data without transcripts, we believe it is a promising direction to enable novel speech processing tools.
MIT Department
Massachusetts Institute of Technology. Department of Electrical Engineering and Computer Science
Terms of Use
In Copyright - Educational Use Permitted
Copyright MIT
Persistent DSpace Link