@inproceedings{jauhiainen-etal-2021-comparing,
    title = "Comparing Approaches to {D}ravidian Language Identification",
    author = "Jauhiainen, Tommi  and
      Ranasinghe, Tharindu  and
      Zampieri, Marcos",
    editor = {Zampieri, Marcos  and
      Nakov, Preslav  and
      Ljube{\v{s}}i{\'c}, Nikola  and
      Tiedemann, J{\"o}rg  and
      Scherrer, Yves  and
      Jauhiainen, Tommi},
    booktitle = "Proceedings of the Eighth Workshop on NLP for Similar Languages, Varieties and Dialects",
    month = apr,
    year = "2021",
    address = "Kiyv, Ukraine",
    publisher = "Association for Computational Linguistics",
    url = "https://preview.aclanthology.org/sigedu-bea-out-of-sync-correction/2021.vardial-1.14/",
    pages = "120--127",
    abstract = "This paper describes the submissions by team HWR to the Dravidian Language Identification (DLI) shared task organized at VarDial 2021 workshop. The DLI training set includes 16,674 YouTube comments written in Roman script containing code-mixed text with English and one of the three South Dravidian languages: Kannada, Malayalam, and Tamil. We submitted results generated using two models, a Naive Bayes classifier with adaptive language models, which has shown to obtain competitive performance in many language and dialect identification tasks, and a transformer-based model which is widely regarded as the state-of-the-art in a number of NLP tasks. Our first submission was sent in the closed submission track using only the training set provided by the shared task organisers, whereas the second submission is considered to be open as it used a pretrained model trained with external data. Our team attained shared second position in the shared task with the submission based on Naive Bayes. Our results reinforce the idea that deep learning methods are not as competitive in language identification related tasks as they are in many other text classification tasks."
}Markdown (Informal)
[Comparing Approaches to Dravidian Language Identification](https://preview.aclanthology.org/sigedu-bea-out-of-sync-correction/2021.vardial-1.14/) (Jauhiainen et al., VarDial 2021)
ACL
- Tommi Jauhiainen, Tharindu Ranasinghe, and Marcos Zampieri. 2021. Comparing Approaches to Dravidian Language Identification. In Proceedings of the Eighth Workshop on NLP for Similar Languages, Varieties and Dialects, pages 120–127, Kiyv, Ukraine. Association for Computational Linguistics.