Probing Language Models for Contextual ScaleUnderstanding
Name
vedantam-saakethv-meng-eecs-2023-thesis.pdf
Description
Thesis PDF
Size
1.56 MB
Format
Adobe PDF
Checksum (MD5)
0baf7b8b10cd0de80616ec3021dcb425
Author(s)
Vedantam, Saaketh
Advisor(s)
Kim, Yoon
Date Issued
June 2023
Publisher
Massachusetts Institute of Technology
Abstract
Pretrained language models (LMs) have demonstrated a remarkable ability to emit linguistic and factual knowledge in certain fields. Additionally, they seem to encode relational information about different concepts in a knowledge base. However, since they are trained solely on textual corpora, it is unclear whether these models implicitly understand anything grounded about the real world. This work investigates the extent to which LMs learn the structure of the physical world. By probing the contextualized embeddings of sentences, we examine how well LMs predict the sizes of real-world objects. We further explore the effect of adjectival modifiers on object embeddings. We show that while larger models more accurately convey scalar information through their embeddings, they perform on par with smaller models in the task of contextual prediction. Fortunately, the models are capable of identifying a difference in scale when an adjectival modifier is introduced, implying that the relevant context is successfully incorporated into the object’s embedding through the LM’s attention mechanism.
MIT Department
Massachusetts Institute of Technology. Department of Electrical Engineering and Computer Science
Terms of Use
In Copyright - Educational Use Permitted
Copyright retained by author(s)
Persistent DSpace Link