Text-Free Audio Captions of Short Videos from Latent Space Representation
Name
Agarwal-anisha24-meng-eecs-2022-thesis.pdf
Description
Thesis PDF
Size
3.63 MB
Format
Adobe PDF
Checksum (MD5)
4900b4513f1ccea710ff7d5405c2e229
Author(s)
Agarwal, Anisha
Advisor(s)
Oliva, Aude
Date Issued
May 2022
Publisher
Massachusetts Institute of Technology
Abstract
In this thesis, we re-implement previous work exploring image to speech captioning. We expand upon the work to implement video to speech captioning. Specifically, we implement a text-free image to speech captioning pipeline that integrates four distinct machine learning models. We alter the models to process video data rather than image data and analyze the resulting speech captions. We conduct experiments on the Wav2Vec2 and HuBERT Automatic Speech Recognition models, and identify which works best with synthesized speech.
MIT Department
Massachusetts Institute of Technology. Department of Electrical Engineering and Computer Science
Terms of Use
In Copyright - Educational Use Permitted
Copyright MIT
Persistent DSpace Link