@inproceedings{bagheri-nezhad-agrawal-2024-drives,
    title = "What Drives Performance in Multilingual Language Models?",
    author = "Bagheri Nezhad, Sina  and
      Agrawal, Ameeta",
    editor = {Scherrer, Yves  and
      Jauhiainen, Tommi  and
      Ljube{\v{s}}i{\'c}, Nikola  and
      Zampieri, Marcos  and
      Nakov, Preslav  and
      Tiedemann, J{\"o}rg},
    booktitle = "Proceedings of the Eleventh Workshop on NLP for Similar Languages, Varieties, and Dialects (VarDial 2024)",
    month = jun,
    year = "2024",
    address = "Mexico City, Mexico",
    publisher = "Association for Computational Linguistics",
    url = "https://preview.aclanthology.org/sigedu-bea-out-of-sync-correction/2024.vardial-1.2/",
    doi = "10.18653/v1/2024.vardial-1.2",
    pages = "16--27",
    abstract = "This study investigates the factors influencing the performance of multilingual large language models (MLLMs) across diverse languages. We study 6 MLLMs, including masked language models, autoregressive models, and instruction-tuned LLMs, on the SIB-200 dataset, a topic classification dataset encompassing 204 languages. Our analysis considers three scenarios: ALL languages, SEEN languages (present in the model{'}s pretraining data), and UNSEEN languages (not present or documented in the model{'}s pretraining data in any meaningful way). We examine the impact of factors such as pretraining data size, general resource availability, language family, and script type on model performance. Decision tree analysis reveals that pretraining data size is the most influential factor for SEEN languages. However, interestingly, script type and language family become more crucial for UNSEEN languages, highlighting the importance of cross-lingual transfer learning. Notably, model size and architecture do not significantly alter the most important features identified. Our findings provide valuable insights into the strengths and limitations of current MLLMs and hope to guide the development of more effective and equitable multilingual NLP systems."
}Markdown (Informal)
[What Drives Performance in Multilingual Language Models?](https://preview.aclanthology.org/sigedu-bea-out-of-sync-correction/2024.vardial-1.2/) (Bagheri Nezhad & Agrawal, VarDial 2024)
ACL
- Sina Bagheri Nezhad and Ameeta Agrawal. 2024. What Drives Performance in Multilingual Language Models?. In Proceedings of the Eleventh Workshop on NLP for Similar Languages, Varieties, and Dialects (VarDial 2024), pages 16–27, Mexico City, Mexico. Association for Computational Linguistics.