Do You See What I Mean? Visual Resolution of Linguistic Ambiguities
Name
CBMM-Memo-051.pdf
Size
2.74 MB
Format
Adobe PDF
Checksum (MD5)
a843822f7dd2b279f079726f79b915ac
Author(s) • • • •
Berzak, Yevgeni
Barbu, Andrei
Harari, Daniel
Katz, Boris
Ullman, Shimon
Date Issued
June 10, 2016
Publisher
Center for Brains, Minds and Machines (CBMM), arXiv
Citation
arXiv:1603.08079v1 [cs.CV]
Series/Report no.
CBMM Memo Series;051
Abstract
Understanding language goes hand in hand with the ability to integrate complex contextual information obtained via perception. In this work, we present a novel task for grounded language understanding: disambiguating a sentence given a visual scene which depicts one of the possible interpretations of that sentence. To this end, we introduce a new multimodal corpus containing ambiguous sentences, representing a wide range of syntactic, semantic and discourse ambiguities, coupled with videos that visualize the different interpretations for each sentence. We address this task by extending a vision model which determines if a sentence is depicted by a video. We demonstrate how such a model can be adjusted to recognize different interpretations of the same underlying sentence, allowing to disambiguate sentences in a unified fashion across the different ambiguity types.
Subjects
Computer Language
Language understanding
Computer vision
Terms of Use
Attribution-NonCommercial-ShareAlike 3.0 United States
Persistent DSpace Link