Cross-lingual Classification of Crisis-related Tweets Using Machine Translation

Shareefa Al Amer, Mark Lee, Phillip Smith


Abstract
Utilisation of multilingual language models such as mBERT and XLM-RoBERTa has increasingly gained attention in recent work by exploiting the multilingualism of such models in different downstream tasks across different languages. However, performance degradation is expected in transfer learning across languages compared to monolingual performance although it is an acceptable trade-off considering the sparsity of resources and lack of available training data in low-resource languages. In this work, we study the effect of machine translation on the cross-lingual transfer learning in a crisis event classification task. Our experiments include measuring the effect of machine-translating the target data into the source language and vice versa. We evaluated and compared the performance in terms of accuracy and F1-Score. The results show that translating the source data into the target language improves the prediction accuracy by 14.8% and the Weighted Average F1-Score by 19.2% when compared to zero-shot transfer to an unseen language.
Anthology ID:
2023.ranlp-1.3
Volume:
Proceedings of the 14th International Conference on Recent Advances in Natural Language Processing
Month:
September
Year:
2023
Address:
Varna, Bulgaria
Editors:
Ruslan Mitkov, Galia Angelova
Venue:
RANLP
SIG:
Publisher:
INCOMA Ltd., Shoumen, Bulgaria
Note:
Pages:
22–31
Language:
URL:
https://aclanthology.org/2023.ranlp-1.3
DOI:
Bibkey:
Cite (ACL):
Shareefa Al Amer, Mark Lee, and Phillip Smith. 2023. Cross-lingual Classification of Crisis-related Tweets Using Machine Translation. In Proceedings of the 14th International Conference on Recent Advances in Natural Language Processing, pages 22–31, Varna, Bulgaria. INCOMA Ltd., Shoumen, Bulgaria.
Cite (Informal):
Cross-lingual Classification of Crisis-related Tweets Using Machine Translation (Al Amer et al., RANLP 2023)
Copy Citation:
PDF:
https://preview.aclanthology.org/emnlp-22-attachments/2023.ranlp-1.3.pdf