Abstract
The paper describes a rule-based system for tagging clause boundaries, implemented for annotating the Estonian Reference Corpus of the University of Tartu, a collection of written texts containing ca 245 million running words and available for querying via Keeleveeb language portal. The system needs information about parts of speech and grammatical categories coded in the word-forms, i.e. it takes morphologically annotated text as input, but requires no information about the syntactic structure of the sentence. Among the strong points of our system we should mention identifying parenthesis and embedded clauses, i.e. clauses that are inserted into another clause dividing it into two separate parts in the linear text, for example a relative clause following its head noun. That enables a corpus query system to unite the otherwise divided clause, a feature that usually presupposes full parsing. The overall precision of the system is 95% and the recall is 96%. If ordinary clause boundary detection and parenthesis and embedded clause boundary detection are evaluated separately, then one can say that detecting an ordinary clause boundary (recall 98%, precision 96%) is an easier task than detecting an embedded clause (recall 79%, precision 100%).- Anthology ID:
- L12-1083
- Volume:
- Proceedings of the Eighth International Conference on Language Resources and Evaluation (LREC'12)
- Month:
- May
- Year:
- 2012
- Address:
- Istanbul, Turkey
- Editors:
- Nicoletta Calzolari, Khalid Choukri, Thierry Declerck, Mehmet Uğur Doğan, Bente Maegaard, Joseph Mariani, Asuncion Moreno, Jan Odijk, Stelios Piperidis
- Venue:
- LREC
- SIG:
- Publisher:
- European Language Resources Association (ELRA)
- Note:
- Pages:
- 1632–1636
- Language:
- URL:
- http://www.lrec-conf.org/proceedings/lrec2012/pdf/229_Paper.pdf
- DOI:
- Cite (ACL):
- Heiki-Jaan Kaalep and Kadri Muischnek. 2012. Robust clause boundary identification for corpus annotation. In Proceedings of the Eighth International Conference on Language Resources and Evaluation (LREC'12), pages 1632–1636, Istanbul, Turkey. European Language Resources Association (ELRA).
- Cite (Informal):
- Robust clause boundary identification for corpus annotation (Kaalep & Muischnek, LREC 2012)
- PDF:
- http://www.lrec-conf.org/proceedings/lrec2012/pdf/229_Paper.pdf