Multi-Reference Benchmarks for Russian Grammatical Error Correction

Frank Palma Gomez; Alla Rozovskaya

Multi-Reference Benchmarks for Russian Grammatical Error Correction

Abstract

This paper presents multi-reference benchmarks for the Grammatical Error Correction (GEC) of Russian, based on two existing single-reference datasets, for a total of 7,444 learner sentences from a variety of first language backgrounds. Each sentence is corrected independently by two new raters, and their corrections are reviewed by a senior annotator, resulting in a total of three references per sentence. Analysis of the annotations reveals that the new raters tend to make more changes, compared to the original raters, especially at the lexical level. We conduct experiments with two popular GEC approaches and show competitive performance on the original datasets and the new benchmarks. We also compare system scores as evaluated against individual annotators and discuss the effect of using multiple references overall and on specific error types. We find that using the union of the references increases system scores by more than 10 points and decreases the gap between system and human performance, thereby providing a more realistic evaluation of GEC system performance, although the effect is not the same across the error types. The annotations are available for research.

Anthology ID:: 2024.eacl-long.76
Volume:: Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers)
Month:: March
Year:: 2024
Address:: St. Julian’s, Malta
Editors:: Yvette Graham, Matthew Purver
Venue:: EACL
SIG:
Publisher:: Association for Computational Linguistics
Note:
Pages:: 1253–1270
Language:
URL:: https://preview.aclanthology.org/jlcl-multiple-ingestion/2024.eacl-long.76/
DOI:
Bibkey:
Cite (ACL):: Frank Palma Gomez and Alla Rozovskaya. 2024. Multi-Reference Benchmarks for Russian Grammatical Error Correction. In Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1253–1270, St. Julian’s, Malta. Association for Computational Linguistics.
Cite (Informal):: Multi-Reference Benchmarks for Russian Grammatical Error Correction (Palma Gomez & Rozovskaya, EACL 2024)
Copy Citation:
PDF:: https://preview.aclanthology.org/jlcl-multiple-ingestion/2024.eacl-long.76.pdf
Video:: https://preview.aclanthology.org/jlcl-multiple-ingestion/2024.eacl-long.76.mp4

PDF Cite Search Video Fix data