Are Girls Neko or Shōjo? Cross-Lingual Alignment of Non-Isomorphic Embeddings with Iterative Normalization
Name
P19-1307.pdf
Description
Published version
Size
1.16 MB
Format
Adobe PDF
Checksum (MD5)
07c67eca68c01098a741ed147f7a73fc
Author(s) • • • •
Zhang, Mozhi
Xu, Keyulu
Kawarabayashi, Ken-ichi
Jegelka, Stefanie Sabrina
Boyd-Graber, Jordan
Date Issued
July 2019
Journal
57th Annual Meeting of the Association for Computational Linguistics
Publisher
Association for Computational Linguistics
Citation
Zhang, Mozhi et al. "Are Girls Neko or Shōjo? Cross-Lingual Alignment of Non-Isomorphic Embeddings with Iterative Normalization." 57th Annual Meeting of the Association for Computational Linguistics, July 2019, Florence, Italy, Association for Computational Linguistics, July 2019. © 2019 Association for Computational Linguistics
Version
Final published version
Abstract
Cross-lingual word embeddings (CLWE) underlie many multilingual natural language processing systems, often through orthogonal transformations of pre-trained monolingual embeddings. However, orthogonal mapping only works on language pairs whose embeddings are naturally isomorphic. For non-isomorphic pairs, our method (Iterative Normalization) transforms monolingual embeddings to make orthogonal alignment easier by simultaneously enforcing that (1) individual word vectors are unit length, and (2) each language's average vector is zero. Iterative Normalization consistently improves word translation accuracy of three CLWE methods, with the largest improvement observed on English-Japanese (from 2% to 44% test accuracy).
MIT Department
Massachusetts Institute of Technology. Department of Linguistics and Philosophy
Terms of Use
Creative Commons Attribution 4.0 International license
Persistent DSpace Link
DOI of Published Version
https://doi.org/10.18653/v1/p19-1307