Mitigating Abusive Comment Detection in Tamil Text: A Data Augmentation Approach with Transformer Model

Reshma Sheik, Raghavan Balanathan, Jaya Nirmala S.


Abstract
With the increasing number of users on social media platforms, the detection and categorization of abusive comments have become crucial, necessitating effective strategies to mitigate their impact on online discussions. However, the intricate and diverse nature of lowresource Indic languages presents a challenge in developing reliable detection methodologies. This research focuses on the task of classifying YouTube comments written in Tamil language into various categories. To achieve this, our research conducted experiments utilizing various multi-lingual transformer-based models along with data augmentation approaches involving back translation approaches and other pre-processing techniques. Our work provides valuable insights into the effectiveness of various preprocessing methods for this classification task. Our experiments showed that the Multilingual Representations for Indian Languages (MURIL) transformer model, coupled with round-trip translation and lexical replacement, yielded the most promising results, showcasing a significant improvement of over 15 units in macro F1-score compared to existing baselines. This contribution adds to the ongoing research to mitigate the adverse impact of abusive content on online platforms, emphasizing the utilization of diverse preprocessing strategies and state-of-the-art language models.
Anthology ID:
2023.icon-1.39
Volume:
Proceedings of the 20th International Conference on Natural Language Processing (ICON)
Month:
December
Year:
2023
Address:
Goa University, Goa, India
Editors:
Jyoti D. Pawar, Sobha Lalitha Devi
Venue:
ICON
SIG:
SIGLEX
Publisher:
NLP Association of India (NLPAI)
Note:
Pages:
460–465
Language:
URL:
https://aclanthology.org/2023.icon-1.39
DOI:
Bibkey:
Cite (ACL):
Reshma Sheik, Raghavan Balanathan, and Jaya Nirmala S.. 2023. Mitigating Abusive Comment Detection in Tamil Text: A Data Augmentation Approach with Transformer Model. In Proceedings of the 20th International Conference on Natural Language Processing (ICON), pages 460–465, Goa University, Goa, India. NLP Association of India (NLPAI).
Cite (Informal):
Mitigating Abusive Comment Detection in Tamil Text: A Data Augmentation Approach with Transformer Model (Sheik et al., ICON 2023)
Copy Citation:
PDF:
https://preview.aclanthology.org/nschneid-patch-4/2023.icon-1.39.pdf