Developing Politeness Annotated Corpus of Hindi Blogs

Ritesh Kumar


Abstract
In this paper I discuss the creation and annotation of a corpus of Hindi blogs. The corpus consists of a total of over 479,000 blog posts and blog comments. It is annotated with the information about the politeness level of each blog post and blog comment. The annotation is carried out using four levels of politeness ― neutral, appropriate, polite and impolite. For the annotation, three classifiers ― were trained and tested maximum entropy (MaxEnt), Support Vector Machines (SVM) and C4.5 - using around 30,000 manually annotated texts. Among these, C4.5 gave the best accuracy. It achieved an accuracy of around 78% which is within 2% of the human accuracy during annotation. Consequently this classifier is used to annotate the rest of the corpus
Anthology ID:
L14-1480
Volume:
Proceedings of the Ninth International Conference on Language Resources and Evaluation (LREC'14)
Month:
May
Year:
2014
Address:
Reykjavik, Iceland
Venue:
LREC
SIG:
Publisher:
European Language Resources Association (ELRA)
Note:
Pages:
1275–1280
Language:
URL:
http://www.lrec-conf.org/proceedings/lrec2014/pdf/594_Paper.pdf
DOI:
Bibkey:
Cite (ACL):
Ritesh Kumar. 2014. Developing Politeness Annotated Corpus of Hindi Blogs. In Proceedings of the Ninth International Conference on Language Resources and Evaluation (LREC'14), pages 1275–1280, Reykjavik, Iceland. European Language Resources Association (ELRA).
Cite (Informal):
Developing Politeness Annotated Corpus of Hindi Blogs (Kumar, LREC 2014)
Copy Citation:
PDF:
http://www.lrec-conf.org/proceedings/lrec2014/pdf/594_Paper.pdf