@inproceedings{zaitsev-minchenko-2022-automatic,
    title = "Automatic Detection of Borrowings in Low-Resource Languages of the {C}aucasus: {A}ndic branch",
    author = "Zaitsev, Konstantin  and
      Minchenko, Anzhelika",
    editor = "Serikov, Oleg  and
      Voloshina, Ekaterina  and
      Postnikova, Anna  and
      Klyachko, Elena  and
      Neminova, Ekaterina  and
      Vylomova, Ekaterina  and
      Shavrina, Tatiana  and
      Ferrand, Eric Le  and
      Malykh, Valentin  and
      Tyers, Francis  and
      Arkhangelskiy, Timofey  and
      Mikhailov, Vladislav  and
      Fenogenova, Alena",
    booktitle = "Proceedings of the First Workshop on NLP applications to field linguistics",
    month = oct,
    year = "2022",
    address = "Gyeongju, Republic of Korea",
    publisher = "International Conference on Computational Linguistics",
    url = "https://preview.aclanthology.org/ingest-emnlp/2022.fieldmatters-1.4/",
    pages = "34--41",
    abstract = "Linguistic borrowings occur in all languages. Andic languages of the Caucasus have borrowings from different donor-languages like Russian, Arabic, Persian. To automatically detect these borrowings, we propose a logistic regression model. The model was trained on the dataset which contains words in IPA from dictionaries of Andic languages. To improve model{'}s quality, we compared TfIdf and Count vectorizers and chose the second one. Besides, we added new features to the model. They were extracted using analysis of vectorizer features and using a language model. The model was evaluated by classification quality metrics (precision, recall and F1-score). The best average F1-score of all languages for words in IPA was about 0.78. Experiments showed that our model reaches good results not only with words in IPA but also with words in Cyrillic."
}Markdown (Informal)
[Automatic Detection of Borrowings in Low-Resource Languages of the Caucasus: Andic branch](https://preview.aclanthology.org/ingest-emnlp/2022.fieldmatters-1.4/) (Zaitsev & Minchenko, FieldMatters 2022)
ACL