REPA: Russian Error Types Annotation for Evaluating Text Generation and Judgment Capabilities
Alexander Pugachev, Alena Fenogenova, Vladislav Mikhailov, Ekaterina Artemova
Abstract
Recent advances in large language models (LLMs) have introduced the novel paradigm of using LLMs as judges, where an LLM evaluates and scores the outputs of another LLM, often correlating highly with human preferences. However, the use of LLM-as-a-judge has been primarily studied in English. In this paper, we evaluate this framework in Russian by introducing the Russian Error tyPes Annotation dataset (REPA, (eng: turnip)), a dataset of 1,000 user queries and 2,000 LLM-generated responses. Human annotators labeled each response pair, expressing their preferences across ten specific error types, as well as selecting an overall preference. We rank six generative LLMs across the error types using three rating systems based on human preferences. We also evaluate responses using eight LLM judges in zero-shot and few-shot settings. We describe the results of analyzing the judges and position and length biases. Our findings reveal a notable gap between LLM judge performance in Russian and English. However, rankings based on human and LLM preferences show partial alignment, suggesting that while current LLM judges struggle with fine-grained evaluation in Russian, there is potential for improvement.- Anthology ID:
- 2025.bsnlp-1.16
- Volume:
- Proceedings of the 10th Workshop on Slavic Natural Language Processing (Slavic NLP 2025)
- Month:
- July
- Year:
- 2025
- Address:
- Vienna, Austria
- Editors:
- Jakub Piskorski, Pavel Přibáň, Preslav Nakov, Roman Yangarber, Michal Marcinczuk
- Venues:
- BSNLP | WS
- SIG:
- Publisher:
- Association for Computational Linguistics
- Note:
- Pages:
- 136–150
- Language:
- URL:
- https://preview.aclanthology.org/acl25-workshop-ingestion/2025.bsnlp-1.16/
- DOI:
- Cite (ACL):
- Alexander Pugachev, Alena Fenogenova, Vladislav Mikhailov, and Ekaterina Artemova. 2025. REPA: Russian Error Types Annotation for Evaluating Text Generation and Judgment Capabilities. In Proceedings of the 10th Workshop on Slavic Natural Language Processing (Slavic NLP 2025), pages 136–150, Vienna, Austria. Association for Computational Linguistics.
- Cite (Informal):
- REPA: Russian Error Types Annotation for Evaluating Text Generation and Judgment Capabilities (Pugachev et al., BSNLP 2025)
- PDF:
- https://preview.aclanthology.org/acl25-workshop-ingestion/2025.bsnlp-1.16.pdf