Exploring the Influence of Spelling Errors on Lexical Variation Measures

Ryo Nagata, Taisei Sato, Hiroya Takamura


Abstract
This paper explores the influence of spelling errors on lexical variation measures. Lexical richness measures such as Type-Token Ration (TTR) and Yule’s K are often used for learner English analysis and assessment. When applied to learner English, however, they can be unreliable because of the spelling errors appearing in it. Namely, they are, directly or indirectly, based on the counts of distinct word types, and spelling errors undesirably increase the number of distinct words. This paper introduces and examines the hypothesis that lexical richness measures become unstable in learner English because of spelling errors. Specifically, it tests the hypothesis on English learner corpora of three groups (middle school, high school, and college students). To be precise, it estimates the difference in TTR and Yule’s K caused by spelling errors, by calculating their values before and after spelling errors are manually corrected. Furthermore, it examines the results theoretically and empirically to deepen the understanding of the influence of spelling errors on them.
Anthology ID:
C18-1202
Volume:
Proceedings of the 27th International Conference on Computational Linguistics
Month:
August
Year:
2018
Address:
Santa Fe, New Mexico, USA
Venue:
COLING
SIG:
Publisher:
Association for Computational Linguistics
Note:
Pages:
2391–2398
Language:
URL:
https://aclanthology.org/C18-1202
DOI:
Bibkey:
Cite (ACL):
Ryo Nagata, Taisei Sato, and Hiroya Takamura. 2018. Exploring the Influence of Spelling Errors on Lexical Variation Measures. In Proceedings of the 27th International Conference on Computational Linguistics, pages 2391–2398, Santa Fe, New Mexico, USA. Association for Computational Linguistics.
Cite (Informal):
Exploring the Influence of Spelling Errors on Lexical Variation Measures (Nagata et al., COLING 2018)
Copy Citation:
PDF:
https://preview.aclanthology.org/auto-file-uploads/C18-1202.pdf