Coreference in Spoken vs. Written Texts: a Corpus-based Analysis
Marilisa Amoia, Kerstin Kunz, Ekaterina Lapshinova-Koltunski
Abstract
This paper describes an empirical study of coreference in spoken vs. written text. We focus on the comparison of two particular text types, interviews and popular science texts, as instances of spoken and written texts since they display quite different discourse structures. We believe in fact, that the correlation of difficulties in coreference resolution and varying discourse structures requires a deeper analysis that accounts for the diversity of coreference strategies or their sub-phenomena as indicators of text type or genre. In this work, we therefore aim at defining specific parameters that classify differences in genres of spoken and written texts such as the preferred segmentation strategy, the maximal allowed distance in or the length and size of coreference chains as well as the correlation of structural and syntactic features of coreferring expressions. We argue that a characterization of such genre dependent parameters might improve the performance of current state-of-art coreference resolution technology.- Anthology ID:
- L12-1362
- Volume:
- Proceedings of the Eighth International Conference on Language Resources and Evaluation (LREC'12)
- Month:
- May
- Year:
- 2012
- Address:
- Istanbul, Turkey
- Editors:
- Nicoletta Calzolari, Khalid Choukri, Thierry Declerck, Mehmet Uğur Doğan, Bente Maegaard, Joseph Mariani, Asuncion Moreno, Jan Odijk, Stelios Piperidis
- Venue:
- LREC
- SIG:
- Publisher:
- European Language Resources Association (ELRA)
- Note:
- Pages:
- 158–164
- Language:
- URL:
- http://www.lrec-conf.org/proceedings/lrec2012/pdf/629_Paper.pdf
- DOI:
- Cite (ACL):
- Marilisa Amoia, Kerstin Kunz, and Ekaterina Lapshinova-Koltunski. 2012. Coreference in Spoken vs. Written Texts: a Corpus-based Analysis. In Proceedings of the Eighth International Conference on Language Resources and Evaluation (LREC'12), pages 158–164, Istanbul, Turkey. European Language Resources Association (ELRA).
- Cite (Informal):
- Coreference in Spoken vs. Written Texts: a Corpus-based Analysis (Amoia et al., LREC 2012)
- PDF:
- http://www.lrec-conf.org/proceedings/lrec2012/pdf/629_Paper.pdf