Foley Music: Learning to Generate Music from Videos
Name
2007.10984.pdf
Description
Submitted version
Size
6.04 MB
Format
Adobe PDF
Checksum (MD5)
1f024e68dec2c68b44161257c2cad9c8
Author(s) • • • •
Gan, Chuang
Huang, Deng
Chen, Peihao
Tenenbaum, Joshua B
Torralba, Antonio
Date Issued
November 2020
Journal
Lecture Notes in Computer Science
Publisher
Springer International Publishing
Citation
Gan, Chuang et al. "Foley Music: Learning to Generate Music from Videos." ECCV: European Conference on Computer Vision, Lecture Notes in Computer Science, 12356, Springer International Publishing, 2020, 758-775. © 2020 Springer Nature Switzerland AG
Version
Original manuscript
Abstract
In this paper, we introduce Foley Music, a system that can synthesize plausible music for a silent video clip about people playing musical instruments. We first identify two key intermediate representations for a successful video to music generator: body keypoints from videos and MIDI events from audio recordings. We then formulate music generation from videos as a motion-to-MIDI translation problem. We present a Graph−Transformer framework that can accurately predict MIDI event sequences in accordance with the body movements. The MIDI event can then be converted to realistic music using an off-the-shelf music synthesizer tool. We demonstrate the effectiveness of our models on videos containing a variety of music performances. Experimental results show that our model outperforms several existing systems in generating music that is pleasant to listen to. More importantly, the MIDI representations are fully interpretable and transparent, thus enabling us to perform music editing flexibly. We encourage the readers to watch the supplementary video with audio turned on to experience the results.
MIT Department
MIT-IBM Watson AI Lab
Massachusetts Institute of Technology. Department of Electrical Engineering and Computer Science
Massachusetts Institute of Technology. Computer Science and Artificial Intelligence Laboratory
Terms of Use
Creative Commons Attribution-Noncommercial-Share Alike
Persistent DSpace Link
DOI of Published Version
https://doi.org/10.1007/978-3-030-58621-8_44