Abstract
In this short overview paper, we describe our system submission for the language pairs Spanish to Aragonese (spa-arg), Spanish to Aranese (spa-arn), and Spanish to Asturian (spa-ast). We train a unified model for all language pairs in the constrained scenario. In addition, we add two language control tokens for Aragonese and Aranese Occitan, as there is already one present for Asturian. We take the distilled NLLB-200 model with 600M parameters and extend special tokens with 2 tokens that denote target languages (arn_Latn, arg_Latn) because Asturian was already presented in NLLB-200 model. We adapt the model by training on a special regime of data augmentation with both monolingual and bilingual training data for the language pairs in this challenge.- Anthology ID:
- 2024.wmt-1.94
- Volume:
- Proceedings of the Ninth Conference on Machine Translation
- Month:
- November
- Year:
- 2024
- Address:
- Miami, Florida, USA
- Editors:
- Barry Haddow, Tom Kocmi, Philipp Koehn, Christof Monz
- Venue:
- WMT
- SIG:
- Publisher:
- Association for Computational Linguistics
- Note:
- Pages:
- 955–959
- Language:
- URL:
- https://aclanthology.org/2024.wmt-1.94
- DOI:
- 10.18653/v1/2024.wmt-1.94
- Cite (ACL):
- Igor Kuzmin, Piotr Przybyła, Euan Mcgill, and Horacio Saggion. 2024. TRIBBLE - TRanslating IBerian languages Based on Limited E-resources. In Proceedings of the Ninth Conference on Machine Translation, pages 955–959, Miami, Florida, USA. Association for Computational Linguistics.
- Cite (Informal):
- TRIBBLE - TRanslating IBerian languages Based on Limited E-resources (Kuzmin et al., WMT 2024)
- PDF:
- https://preview.aclanthology.org/dois-2013-emnlp/2024.wmt-1.94.pdf