Adding More Languages Improves Unsupervised Multilingual Part-of-Speech Tagging: A Bayesian Non-Parametric Approach
Name
Barzilay_Adding more.pdf
Size
544.26 KB
Format
Adobe PDF
Checksum (MD5)
668c13fc2f1f6557b89f56bbe28f7f7c
Author(s) • • •
Snyder, Benjamin
Naseem, Tahira
Eisenstein, Jacob
Barzilay, Regina
Date Issued
June 2009
Journal
Proceedings of Human Language Technologies: The 2009 Annual Conference of the North American Chapter of the Association for Computational Linguistics
Publisher
Association for Computational Linguistics
Citation
Snyder, Benjamin. et al. "Adding More Languages Improves Unsupervised Multilingual Part-of-Speech Tagging: A Bayesian Non-Parametric Approach." Human Language Technologies: The 2009 Annual Conference of the North American Chapter of the ACL, pages 83–91,
Boulder, Colorado, June 2009.
Version
Author's final manuscript
Abstract
We investigate the problem of unsupervised part-of-speech tagging when raw parallel data is available in a large number of languages. Patterns of ambiguity vary greatly across languages and therefore even unannotated multilingual data can serve as a learning signal. We propose a non-parametric Bayesian model that connects related tagging decisions across languages through the use of multilingual latent variables. Our experiments show that performance improves steadily as the number of languages increases.
MIT Department
Massachusetts Institute of Technology. Computer Science and Artificial Intelligence Laboratory
Massachusetts Institute of Technology. Department of Electrical Engineering and Computer Science
Terms of Use
Attribution-Noncommercial-Share Alike 3.0 Unported
Persistent DSpace Link
DOI of Published Version
http://portal.acm.org/citation.cfm?id=1620754.1620767