Bengali and Magahi PUD Treebank and Parser
Pritha Majumdar, Deepak Alok, Akanksha Bansal, Atul Kr. Ojha, John P. McCrae
Abstract
This paper presents the development of the Parallel Universal Dependency (PUD) Treebank for two Indo-Aryan languages: Bengali and Magahi. A treebank of 1,000 sentences has been created using a parallel corpus of English and the UD framework. A preliminary set of sentences was annotated manually - 600 for Bengali and 200 for Magahi. The rest of the sentences were built using the Bengali and Magahi parser. The sentences have been translated and annotated manually by the authors, some of whom are also native speakers of the languages. The objective behind this work is to build a syntactically-annotated linguistic repository for the aforementioned languages, that can prove to be a useful resource for building further NLP tools. Additionally, Bengali and Magahi parsers were also created which is built on machine learning approach. The accuracy of the Bengali parser is 78.13% in the case of UPOS; 76.99% in the case of XPOS, 56.12% in the case of UAS; and 47.19% in the case of LAS. The accuracy of Magahi parser is 71.53% in the case of UPOS; 66.44% in the case of XPOS, 58.05% in the case of UAS; and 33.07% in the case of LAS. This paper also includes an illustration of the annotation schema followed, the findings of the Parallel Universal Dependency (PUD) treebank, and it’s resulting linguistic analysis- Anthology ID:
- 2022.wildre-1.11
- Volume:
- Proceedings of the WILDRE-6 Workshop within the 13th Language Resources and Evaluation Conference
- Month:
- June
- Year:
- 2022
- Address:
- Marseille, France
- Editors:
- Girish Nath Jha, Sobha L., Kalika Bali, Atul Kr. Ojha
- Venue:
- WILDRE
- SIG:
- Publisher:
- European Language Resources Association
- Note:
- Pages:
- 60–67
- Language:
- URL:
- https://aclanthology.org/2022.wildre-1.11
- DOI:
- Cite (ACL):
- Pritha Majumdar, Deepak Alok, Akanksha Bansal, Atul Kr. Ojha, and John P. McCrae. 2022. Bengali and Magahi PUD Treebank and Parser. In Proceedings of the WILDRE-6 Workshop within the 13th Language Resources and Evaluation Conference, pages 60–67, Marseille, France. European Language Resources Association.
- Cite (Informal):
- Bengali and Magahi PUD Treebank and Parser (Majumdar et al., WILDRE 2022)
- PDF:
- https://preview.aclanthology.org/nschneid-patch-5/2022.wildre-1.11.pdf