Abstract
In this paper, we describe a corpus UD Japanese-BCCWJ that was created by converting the Balanced Corpus of Contemporary Written Japanese (BCCWJ), a Japanese language corpus, to adhere to the UD annotation schema. The BCCWJ already assigns dependency information at the level of the bunsetsu (a Japanese syntactic unit comparable to the phrase). We developed a program to convert the BCCWJ to UD based on this dependency structure, and this corpus is the result of completely automatic conversion using the program. UD Japanese-BCCWJ is the largest-scale UD Japanese corpus and the second-largest of all UD corpora, including 1,980 documents, 57,109 sentences, and 1,273k words across six distinct domains.- Anthology ID:
- W18-6014
- Volume:
- Proceedings of the Second Workshop on Universal Dependencies (UDW 2018)
- Month:
- November
- Year:
- 2018
- Address:
- Brussels, Belgium
- Editors:
- Marie-Catherine de Marneffe, Teresa Lynn, Sebastian Schuster
- Venue:
- UDW
- SIG:
- Publisher:
- Association for Computational Linguistics
- Note:
- Pages:
- 117–125
- Language:
- URL:
- https://aclanthology.org/W18-6014
- DOI:
- 10.18653/v1/W18-6014
- Cite (ACL):
- Mai Omura and Masayuki Asahara. 2018. UD-Japanese BCCWJ: Universal Dependencies Annotation for the Balanced Corpus of Contemporary Written Japanese. In Proceedings of the Second Workshop on Universal Dependencies (UDW 2018), pages 117–125, Brussels, Belgium. Association for Computational Linguistics.
- Cite (Informal):
- UD-Japanese BCCWJ: Universal Dependencies Annotation for the Balanced Corpus of Contemporary Written Japanese (Omura & Asahara, UDW 2018)
- PDF:
- https://preview.aclanthology.org/ml4al-ingestion/W18-6014.pdf