Combining Multiple Corpora for Readability Assessment for People with Cognitive Disabilities
Victoria Yaneva, Constantin Orăsan, Richard Evans, Omid Rohanian
Abstract
Given the lack of large user-evaluated corpora in disability-related NLP research (e.g. text simplification or readability assessment for people with cognitive disabilities), the question of choosing suitable training data for NLP models is not straightforward. The use of large generic corpora may be problematic because such data may not reflect the needs of the target population. The use of the available user-evaluated corpora may be problematic because these datasets are not large enough to be used as training data. In this paper we explore a third approach, in which a large generic corpus is combined with a smaller population-specific corpus to train a classifier which is evaluated using two sets of unseen user-evaluated data. One of these sets, the ASD Comprehension corpus, is developed for the purposes of this study and made freely available. We explore the effects of the size and type of the training data used on the performance of the classifiers, and the effects of the type of the unseen test datasets on the classification performance.- Anthology ID:
- W17-5013
- Volume:
- Proceedings of the 12th Workshop on Innovative Use of NLP for Building Educational Applications
- Month:
- September
- Year:
- 2017
- Address:
- Copenhagen, Denmark
- Editors:
- Joel Tetreault, Jill Burstein, Claudia Leacock, Helen Yannakoudakis
- Venue:
- BEA
- SIG:
- SIGEDU
- Publisher:
- Association for Computational Linguistics
- Note:
- Pages:
- 121–132
- Language:
- URL:
- https://aclanthology.org/W17-5013
- DOI:
- 10.18653/v1/W17-5013
- Cite (ACL):
- Victoria Yaneva, Constantin Orăsan, Richard Evans, and Omid Rohanian. 2017. Combining Multiple Corpora for Readability Assessment for People with Cognitive Disabilities. In Proceedings of the 12th Workshop on Innovative Use of NLP for Building Educational Applications, pages 121–132, Copenhagen, Denmark. Association for Computational Linguistics.
- Cite (Informal):
- Combining Multiple Corpora for Readability Assessment for People with Cognitive Disabilities (Yaneva et al., BEA 2017)
- PDF:
- https://preview.aclanthology.org/add_acl24_videos/W17-5013.pdf