Automated Refinement of Essay Scoring Rubrics for Language Models via Reflect-and-Revise
Keno Harada, Lui Yoshida, Takeshi Kojima, Yusuke Iwasawa, Yutaka Matsuo
Abstract
Large Language Models (LLMs) are increasingly used for Automated Essay Scoring (AES), yet the scoring rubrics they rely on are typically designed for human raters and may not be optimal for LLMs. Inspired by the calibration process that human raters undergo before formal scoring, we propose Reflect-and-Revise, an iterative framework that refines scoring rubrics by prompting models to reflect on their own chain-of-thought rationales and score discrepancies with human labels. At each iteration, the model identifies scoring-error patterns from sampled mismatches and revises the rubric accordingly. Experiments on three essay scoring benchmarks (ASAP, ASAP 2.0, and TOEFL11) with three LLMs (GPT-5 mini, Gemini 3 Flash, and Qwen3-Next-80B-A3B-Instruct) demonstrate that our method yields improvements in Quadratic Weighted Kappa (QWK), achieving gains of up to +0.403 over human-authored rubrics. Starting from a minimal seed rubric that specifies only the score scale, our method matches or exceeds expert rubric performance in most dataset-model combinations, indicating that iterative refinement can reduce the manual effort of rubric authoring. Analysis of the refined rubrics reveals that the refinement process introduces explicit procedural structures, such as conditional gating rules and quantitative thresholds, that are absent from human-authored rubrics, highlighting a gap between rubrics designed for human raters and those effective for LLMs.- Anthology ID:
- 2026.conll-main.47
- Volume:
- Proceedings of the 30th Conference on Computational Natural Language Learning
- Month:
- July
- Year:
- 2026
- Address:
- San Diego, California, USA
- Editors:
- Claire Bonial, Yevgeni Berzak
- Venues:
- CoNLL | WS
- SIG:
- Publisher:
- Association for Computational Linguistics
- Note:
- Pages:
- 771–789
- Language:
- URL:
- https://preview.aclanthology.org/ingest-acl-workshops/2026.conll-main.47/
- DOI:
- Cite (ACL):
- Keno Harada, Lui Yoshida, Takeshi Kojima, Yusuke Iwasawa, and Yutaka Matsuo. 2026. Automated Refinement of Essay Scoring Rubrics for Language Models via Reflect-and-Revise. In Proceedings of the 30th Conference on Computational Natural Language Learning, pages 771–789, San Diego, California, USA. Association for Computational Linguistics.
- Cite (Informal):
- Automated Refinement of Essay Scoring Rubrics for Language Models via Reflect-and-Revise (Harada et al., CoNLL 2026)
- PDF:
- https://preview.aclanthology.org/ingest-acl-workshops/2026.conll-main.47.pdf