FreeTransfer-X: Safe and Label-Free Cross-Lingual Transfer from Off-the-Shelf Models

Yinpeng Guo, Liangyou Li, Xin Jiang, Qun Liu


Abstract
Cross-lingual transfer (CLT) is of various applications. However, labeled cross-lingual corpus is expensive or even inaccessible, especially in the fields where labels are private, such as diagnostic results of symptoms in medicine and user profiles in business. Nevertheless, there are off-the-shelf models in these sensitive fields. Instead of pursuing the original labels, a workaround for CLT is to transfer knowledge from the off-the-shelf models without labels. To this end, we define a novel CLT problem named FreeTransfer-X that aims to achieve knowledge transfer from the off-the-shelf models in rich-resource languages. To address the problem, we propose a 2-step knowledge distillation (KD, Hinton et al., 2015) framework based on multilingual pre-trained language models (mPLM). The significant improvement over strong neural machine translation (NMT) baselines demonstrates the effectiveness of the proposed method. In addition to reducing annotation cost and protecting private labels, the proposed method is compatible with different networks and easy to be deployed. Finally, a range of analyses indicate the great potential of the proposed method.
Anthology ID:
2022.findings-naacl.16
Volume:
Findings of the Association for Computational Linguistics: NAACL 2022
Month:
July
Year:
2022
Address:
Seattle, United States
Editors:
Marine Carpuat, Marie-Catherine de Marneffe, Ivan Vladimir Meza Ruiz
Venue:
Findings
SIG:
Publisher:
Association for Computational Linguistics
Note:
Pages:
217–228
Language:
URL:
https://aclanthology.org/2022.findings-naacl.16
DOI:
10.18653/v1/2022.findings-naacl.16
Bibkey:
Cite (ACL):
Yinpeng Guo, Liangyou Li, Xin Jiang, and Qun Liu. 2022. FreeTransfer-X: Safe and Label-Free Cross-Lingual Transfer from Off-the-Shelf Models. In Findings of the Association for Computational Linguistics: NAACL 2022, pages 217–228, Seattle, United States. Association for Computational Linguistics.
Cite (Informal):
FreeTransfer-X: Safe and Label-Free Cross-Lingual Transfer from Off-the-Shelf Models (Guo et al., Findings 2022)
Copy Citation:
PDF:
https://preview.aclanthology.org/landing_page/2022.findings-naacl.16.pdf
Video:
 https://preview.aclanthology.org/landing_page/2022.findings-naacl.16.mp4
Data
MTOP