Studying the history of the Arabic language: language technology and a large-scale historical corpus
Name
10579_2019_9460_ReferencePDF.pdf
Size
671.41 KB
Format
Unknown
Checksum (MD5)
333672883706bdbd49e1891e0761e4c0
Author(s) • • • •
Belinkov, Yonatan
Magidow, Alexander
Barrón-Cedeño, Alberto
Shmidman, Avi
Romanov, Maxim
Date Issued
April 2019
Journal
Language Resources and Evaluation
Publisher
Springer Netherlands
Version
Author's final manuscript
Abstract
Abstract
Arabic is a widely-spoken language with a long and rich history, but existing corpora and language technology focus mostly on modern Arabic and its varieties.
Therefore, studying the history of the language has so far been mostly limited to manual analyses on a small scale. In this work, we present a large-scale historical corpus of the written Arabic language, spanning 1400 years. We describe our efforts to clean and process this corpus using Arabic NLP tools, including the identification of reused text.
We study the history of the Arabic language using a novel automatic periodization algorithm, as well as other techniques.
Our findings confirm the established division of written Arabic into Modern Standard and Classical Arabic, and confirm other established periodizations, while suggesting that written Arabic may be divisible into still further periods of development.
MIT Department
Massachusetts Institute of Technology. Computer Science and Artificial Intelligence Laboratory
Terms of Use
Article is made available in accordance with the publisher's policy and may be subject to US copyright law. Please refer to the publisher's site for terms of use.
Persistent DSpace Link
DOI of Published Version
https://doi.org/10.1007/s10579-019-09460-w