An Empirical Study of Leveraging Knowledge Distillation for Compressing Multilingual Neural Machine Translation Models

Varun Gumma; Raj Dabre; Pratyush Kumar

An Empirical Study of Leveraging Knowledge Distillation for Compressing Multilingual Neural Machine Translation Models

Abstract

Knowledge distillation (KD) is a well-known method for compressing neural models. However, works focusing on distilling knowledge from large multilingual neural machine translation (MNMT) models into smaller ones are practically nonexistent, despite the popularity and superiority of MNMT. This paper bridges this gap by presenting an empirical investigation of knowledge distillation for compressing MNMT models. We take Indic to English translation as a case study and demonstrate that commonly used language-agnostic and language-aware KD approaches yield models that are 4-5x smaller but also suffer from performance drops of up to 3.5 BLEU. To mitigate this, we then experiment with design considerations such as shallower versus deeper models, heavy parameter sharing, multistage training, and adapters. We observe that deeper compact models tend to be as good as shallower non-compact ones and that fine-tuning a distilled model on a high-quality subset slightly boosts translation quality. Overall, we conclude that compressing MNMT models via KD is challenging, indicating immense scope for further research.

Anthology ID:: 2023.eamt-1.11
Volume:: Proceedings of the 24th Annual Conference of the European Association for Machine Translation
Month:: June
Year:: 2023
Address:: Tampere, Finland
Editors:: Mary Nurminen, Judith Brenner, Maarit Koponen, Sirkku Latomaa, Mikhail Mikhailov, Frederike Schierl, Tharindu Ranasinghe, Eva Vanmassenhove, Sergi Alvarez Vidal, Nora Aranberri, Mara Nunziatini, Carla Parra Escartín, Mikel Forcada, Maja Popovic, Carolina Scarton, Helena Moniz
Venue:: EAMT
SIG:
Publisher:: European Association for Machine Translation
Note:
Pages:: 103–114
Language:
URL:: https://aclanthology.org/2023.eamt-1.11
DOI:
Bibkey:
Cite (ACL):: Varun Gumma, Raj Dabre, and Pratyush Kumar. 2023. An Empirical Study of Leveraging Knowledge Distillation for Compressing Multilingual Neural Machine Translation Models. In Proceedings of the 24th Annual Conference of the European Association for Machine Translation, pages 103–114, Tampere, Finland. European Association for Machine Translation.
Cite (Informal):: An Empirical Study of Leveraging Knowledge Distillation for Compressing Multilingual Neural Machine Translation Models (Gumma et al., EAMT 2023)
Copy Citation:
PDF:: https://preview.aclanthology.org/nschneid-patch-3/2023.eamt-1.11.pdf

PDF Search