Importance-based Neuron Allocation for Multilingual Neural Machine Translation

Wanying Xie, Yang Feng, Shuhao Gu, Dong Yu


Abstract
Multilingual neural machine translation with a single model has drawn much attention due to its capability to deal with multiple languages. However, the current multilingual translation paradigm often makes the model tend to preserve the general knowledge, but ignore the language-specific knowledge. Some previous works try to solve this problem by adding various kinds of language-specific modules to the model, but they suffer from the parameter explosion problem and require specialized manual design. To solve these problems, we propose to divide the model neurons into general and language-specific parts based on their importance across languages. The general part is responsible for preserving the general knowledge and participating in the translation of all the languages, while the language-specific part is responsible for preserving the language-specific knowledge and participating in the translation of some specific languages. Experimental results on several language pairs, covering IWSLT and Europarl corpus datasets, demonstrate the effectiveness and universality of the proposed method.
Anthology ID:
2021.acl-long.445
Volume:
Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers)
Month:
August
Year:
2021
Address:
Online
Editors:
Chengqing Zong, Fei Xia, Wenjie Li, Roberto Navigli
Venues:
ACL | IJCNLP
SIG:
Publisher:
Association for Computational Linguistics
Note:
Pages:
5725–5737
Language:
URL:
https://aclanthology.org/2021.acl-long.445
DOI:
10.18653/v1/2021.acl-long.445
Bibkey:
Cite (ACL):
Wanying Xie, Yang Feng, Shuhao Gu, and Dong Yu. 2021. Importance-based Neuron Allocation for Multilingual Neural Machine Translation. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 5725–5737, Online. Association for Computational Linguistics.
Cite (Informal):
Importance-based Neuron Allocation for Multilingual Neural Machine Translation (Xie et al., ACL-IJCNLP 2021)
Copy Citation:
PDF:
https://preview.aclanthology.org/nschneid-patch-2/2021.acl-long.445.pdf
Video:
 https://preview.aclanthology.org/nschneid-patch-2/2021.acl-long.445.mp4
Code
 ictnlp/NA-MNMT