ACTSA: Annotated Corpus for Telugu Sentiment Analysis

Sandeep Sricharan Mukku, Radhika Mamidi


Abstract
Sentiment analysis deals with the task of determining the polarity of a document or sentence and has received a lot of attention in recent years for the English language. With the rapid growth of social media these days, a lot of data is available in regional languages besides English. Telugu is one such regional language with abundant data available in social media, but it’s hard to find a labelled data of sentences for Telugu Sentiment Analysis. In this paper, we describe an effort to build a gold-standard annotated corpus of Telugu sentences to support Telugu Sentiment Analysis. The corpus, named ACTSA (Annotated Corpus for Telugu Sentiment Analysis) has a collection of Telugu sentences taken from different sources which were then pre-processed and manually annotated by native Telugu speakers using our annotation guidelines. In total, we have annotated 5457 sentences, which makes our corpus the largest resource currently available. The corpus and the annotation guidelines are made publicly available.
Anthology ID:
W17-5408
Volume:
Proceedings of the First Workshop on Building Linguistically Generalizable NLP Systems
Month:
September
Year:
2017
Address:
Copenhagen, Denmark
Venue:
WS
SIG:
Publisher:
Association for Computational Linguistics
Note:
Pages:
54–58
Language:
URL:
https://aclanthology.org/W17-5408
DOI:
10.18653/v1/W17-5408
Bibkey:
Cite (ACL):
Sandeep Sricharan Mukku and Radhika Mamidi. 2017. ACTSA: Annotated Corpus for Telugu Sentiment Analysis. In Proceedings of the First Workshop on Building Linguistically Generalizable NLP Systems, pages 54–58, Copenhagen, Denmark. Association for Computational Linguistics.
Cite (Informal):
ACTSA: Annotated Corpus for Telugu Sentiment Analysis (Mukku & Mamidi, 2017)
Copy Citation:
PDF:
https://preview.aclanthology.org/auto-file-uploads/W17-5408.pdf