Multilingual Part-of-Speech Tagging Two Unsupervised Approaches

Naseem, Tahira; Snyder, Benjamin; Eisenstein, Jacob; Barzilay, Regina

dc.contributor.author	Naseem, Tahira
dc.contributor.author	Snyder, Benjamin
dc.contributor.author	Eisenstein, Jacob
dc.contributor.author	Barzilay, Regina
dc.date.accessioned	2011-05-10T18:23:52Z
dc.date.available	2011-05-10T18:23:52Z
dc.date.issued	2009-11
dc.date.submitted	2009-05
dc.identifier.issn	1943-5037
dc.identifier.issn	1076-9757
dc.identifier.uri	http://hdl.handle.net/1721.1/62804
dc.description.abstract	We demonstrate the effectiveness of multilingual learning for unsupervised part-of-speech tagging. The central assumption of our work is that by combining cues from multiple languages, the structure of each becomes more apparent. We consider two ways of applying this intuition to the problem of unsupervised part-of-speech tagging: a model that directly merges tag structures for a pair of languages into a single sequence and a second model which instead incorporates multilingual context using latent variables. Both approaches are formulated as hierarchical Bayesian models, using Markov Chain Monte Carlo sampling techniques for inference. Our results demonstrate that by incorporating multilingual evidence we can achieve impressive performance gains across a range of scenarios. We also found that performance improves steadily as the number of available languages increases.	en_US
dc.description.sponsorship	National Science Foundation (U.S.) (CAREER grant IIS-0448168)	en_US
dc.description.sponsorship	National Science Foundation (U.S.) (grant IIS-0835445)	en_US
dc.description.sponsorship	National Science Foundation (U.S.) (grant IIS-0904684)	en_US
dc.description.sponsorship	Microsoft Research (Faculty Fellowship)	en_US
dc.language.iso	en_US
dc.publisher	AI Access Foundation	en_US
dc.rights	Article is made available in accordance with the publisher's policy and may be subject to US copyright law. Please refer to the publisher's site for terms of use.	en_US
dc.source	JAIR	en_US
dc.title	Multilingual Part-of-Speech Tagging Two Unsupervised Approaches	en_US
dc.type	Article	en_US
dc.identifier.citation	Naseem, Tahira, et al. "Multilingual Part-of-Speech Tagging Two Unsupervised Approaches." Journal of Artificial Intelligence Research 36 (2009) 341-385. © AI Access Foundation.	en_US
dc.contributor.department	Massachusetts Institute of Technology. Computer Science and Artificial Intelligence Laboratory	en_US
dc.contributor.department	Massachusetts Institute of Technology. Department of Electrical Engineering and Computer Science	en_US
dc.contributor.approver	Barzilay, Regina
dc.contributor.mitauthor	Barzilay, Regina
dc.contributor.mitauthor	Naseem, Tahira
dc.contributor.mitauthor	Snyder, Benjamin
dc.contributor.mitauthor	Eisenstein, Jacob
dc.relation.journal	Journal of Artificial Intelligence Research	en_US
dc.eprint.version	Final published version	en_US
dc.type.uri	http://purl.org/eprint/type/JournalArticle	en_US
eprint.status	http://purl.org/eprint/status/PeerReviewed	en_US
dspace.orderedauthors	Naseem, Tahira; Snyder, Benjamin; Eisenstein, Jacob; Barzilay, Regina
dspace.orderedauthors	Naseem, Tahira; Snyder, Benjamin; Eisenstein, Jacob; Barzilay, Regina
dc.identifier.orcid	https://orcid.org/0000-0002-2921-8201
mit.license	PUBLISHER_POLICY	en_US
mit.metadata.status	Complete

Files in this item

Name:: Naseem-2009-Multilingual Part- ...
Size:: 1.340Mb
Format:: PDF

View/Open

This item appears in the following Collection(s)

MIT Open Access Articles

Show simple item record