The Natural Stories corpus: a reading-time corpus of English texts containing rare syntactic constructions

Futrell, Richard; Gibson, Edward A; Tily, Harry J.; Blank, Idan; Vishnevetsky, Anastasia; Piantadosi, Steven T.; Fedorenko, Evelina

dc.contributor.author	Futrell, Richard
dc.contributor.author	Gibson, Edward A
dc.contributor.author	Tily, Harry J.
dc.contributor.author	Blank, Idan
dc.contributor.author	Vishnevetsky, Anastasia
dc.contributor.author	Piantadosi, Steven T.
dc.contributor.author	Fedorenko, Evelina
dc.date.accessioned	2020-09-15T17:47:22Z
dc.date.available	2020-09-15T17:47:22Z
dc.date.issued	2020-09
dc.identifier.issn	1574-020X
dc.identifier.issn	1574-0218
dc.identifier.uri	https://hdl.handle.net/1721.1/127270
dc.description.abstract	It is now a common practice to compare models of human language processing by comparing how well they predict behavioral and neural measures of processing difficulty, such as reading times, on corpora of rich naturalistic linguistic materials. However, many of these corpora, which are based on naturally-occurring text, do not contain many of the low-frequency syntactic constructions that are often required to distinguish between processing theories. Here we describe a new corpus consisting of English texts edited to contain many low-frequency syntactic constructions while still sounding fluent to native speakers. The corpus is annotated with hand-corrected Penn Treebank-style parse trees and includes self-paced reading time data and aligned audio recordings. We give an overview of the content of the corpus, review recent work using the corpus, and release the data.	en_US
dc.description.sponsorship	National Science Foundation (Grants 0844472 and 1534318)	en_US
dc.publisher	Springer Science and Business Media LLC	en_US
dc.relation.isversionof	http://dx.doi.org/10.1007/s10579-020-09503-7	en_US
dc.rights	Creative Commons Attribution	en_US
dc.rights.uri	https://creativecommons.org/licenses/by/4.0/	en_US
dc.source	Springer Netherlands	en_US
dc.title	The Natural Stories corpus: a reading-time corpus of English texts containing rare syntactic constructions	en_US
dc.type	Article	en_US
dc.identifier.citation	Futrell, Richard et al. "The Natural Stories corpus: a reading-time corpus of English texts containing rare syntactic constructions." Language Resources and Evaluation (September 2020): doi.org/10.1007/s10579-020-09503-7 © 2020 Springer Nature	en_US
dc.contributor.department	Massachusetts Institute of Technology. Department of Brain and Cognitive Sciences	en_US
dc.relation.journal	Language Resources and Evaluation	en_US
dc.eprint.version	Final published version	en_US
dc.type.uri	http://purl.org/eprint/type/JournalArticle	en_US
eprint.status	http://purl.org/eprint/status/PeerReviewed	en_US
dc.date.updated	2020-09-05T03:32:23Z
dc.language.rfc3066	en
dc.rights.holder	The Author(s)
dspace.embargo.terms	N
dspace.date.submission	2020-09-05T03:32:23Z
mit.license	PUBLISHER_CC
mit.metadata.status	Complete

Files in this item

Name:: 10579_2020_Article_9503.pdf
Size:: 524.0Kb
Format:: PDF

View/Open

This item appears in the following Collection(s)

MIT Open Access Articles

Show simple item record