Aggression Identification in English, Hindi and Bangla Text using BERT, RoBERTa and SVM

Arup Baruah, Kaushik Das, Ferdous Barbhuiya, Kuntal Dey


Abstract
This paper presents the results of the classifiers we developed for the shared tasks in aggression identification and misogynistic aggression identification. These two shared tasks were held as part of the second workshop on Trolling, Aggression and Cyberbullying (TRAC). Both the subtasks were held for English, Hindi and Bangla language. In our study, we used English BERT (En-BERT), RoBERTa, DistilRoBERTa, and SVM based classifiers for English language. For Hindi and Bangla language, multilingual BERT (M-BERT), XLM-RoBERTa and SVM classifiers were used. Our best performing models are EN-BERT for English Subtask A (Weighted F1 score of 0.73, Rank 5/16), SVM for English Subtask B (Weighted F1 score of 0.87, Rank 2/15), SVM for Hindi Subtask A (Weighted F1 score of 0.79, Rank 2/10), XLMRoBERTa for Hindi Subtask B (Weighted F1 score of 0.87, Rank 2/10), SVM for Bangla Subtask A (Weighted F1 score of 0.81, Rank 2/10), and SVM for Bangla Subtask B (Weighted F1 score of 0.93, Rank 4/8). It is seen that the superior performance of the SVM classifier was achieved mainly because of its better prediction of the majority class. BERT based classifiers were found to predict the minority classes better.
Anthology ID:
2020.trac-1.12
Volume:
Proceedings of the Second Workshop on Trolling, Aggression and Cyberbullying
Month:
May
Year:
2020
Address:
Marseille, France
Editors:
Ritesh Kumar, Atul Kr. Ojha, Bornini Lahiri, Marcos Zampieri, Shervin Malmasi, Vanessa Murdock, Daniel Kadar
Venue:
TRAC
SIG:
Publisher:
European Language Resources Association (ELRA)
Note:
Pages:
76–82
Language:
English
URL:
https://aclanthology.org/2020.trac-1.12
DOI:
Bibkey:
Cite (ACL):
Arup Baruah, Kaushik Das, Ferdous Barbhuiya, and Kuntal Dey. 2020. Aggression Identification in English, Hindi and Bangla Text using BERT, RoBERTa and SVM. In Proceedings of the Second Workshop on Trolling, Aggression and Cyberbullying, pages 76–82, Marseille, France. European Language Resources Association (ELRA).
Cite (Informal):
Aggression Identification in English, Hindi and Bangla Text using BERT, RoBERTa and SVM (Baruah et al., TRAC 2020)
Copy Citation:
PDF:
https://preview.aclanthology.org/nschneid-patch-2/2020.trac-1.12.pdf
Data
Urdu Online Reviews