Long Sequence Transformer Variants on Varying Context Length
Name
sun-mmsun-meng-eecs-2023-thesis.pdf
Description
Thesis PDF
Size
498.06 KB
Format
Adobe PDF
Checksum (MD5)
35d9157ebc30b4642da3fb86cf7ac350
Author(s)
Sun, Melinda
Advisor(s)
Kim, Yoon
Date Issued
September 2023
Publisher
Massachusetts Institute of Technology
Abstract
Transformers are powerful and effective tools in natural language processing, but their scalability is limited by the quadratic complexity of attention. Several transformer variants that address this problem have recently been proposed, including Moving Average Equipped Gated Attention (Mega). In this thesis, we evaluate how effectively Mega uses past context, by comparing the perplexity trend as context length varies with the perplexity trend of a standard transformer. We find that Mega does not show greater benefit from longer context in a Wikipedia or book setting, though it does have a much better ability to extrapolate beyond training context lengths.
MIT Department
Massachusetts Institute of Technology. Department of Electrical Engineering and Computer Science
Terms of Use
In Copyright - Educational Use Permitted
Copyright retained by author(s)
Persistent DSpace Link