A Critical Study of Automatic Evaluation in Sign Language Translation
Shakib Yazdani, Yasser HAMIDULLAH, Cristina España-Bonet, Eleftherios Avramidis, Josef van Genabith
Abstract
Automatic evaluation metrics are crucial for advancing sign language translation (SLT). Current SLT evaluation metrics, such as BLEU and ROUGE, are only text-based, and it remains unclear to what extent text-based metrics can reliably capture the quality of SLT outputs. To address this gap, we investigate the limitations of text-based SLT evaluation metrics by analyzing six metrics, including BLEU, chrF, and ROUGE, as well as BLEURT on the one hand, and large language model (LLM)-based evaluators such as G-Eval and GEMBA zero-shot direct assessment on the other hand. Specifically, we assess the consistency and robustness of these metrics under three controlled conditions: paraphrasing, hallucinations in model outputs, and variations in sentence length. Our analysis highlights the limitations of lexical overlap metrics and demonstrates that while LLM-based evaluators better capture semantic equivalence often missed by conventional metrics, they can also exhibit bias toward LLM-paraphrased translations. Moreover, although all metrics are able to detect hallucinations, BLEU tends to be overly sensitive, whereas BLEURT and LLM-based evaluators are comparatively lenient toward subtle cases. This motivates the need for multimodal evaluation frameworks that extend beyond text-based metrics to enable a more holistic assessment of SLT outputs.- Anthology ID:
- 2026.lrec-1.749
- Volume:
- Proceedings of the Fifteenth Language Resources and Evaluation Conference
- Month:
- May
- Year:
- 2026
- Address:
- Palma de Mallorca, Spain
- Editors:
- Stelios Piperidis, Núria Bel, Henk van den Heuvel, Nancy Ide, Simon Krek, Antonio Toral
- Venue:
- LREC
- SIG:
- Publisher:
- ELRA Language Resource Association
- Note:
- Pages:
- 9535–9548
- Language:
- External URL:
- https://lrec.elra.info/lrec2026-main-749
- DOI:
- 10.63317/4n2sooe4fb2i
- Cite (ACL):
- Shakib Yazdani, Yasser HAMIDULLAH, Cristina España-Bonet, Eleftherios Avramidis, and Josef van Genabith. 2026. A Critical Study of Automatic Evaluation in Sign Language Translation. In Proceedings of the Fifteenth Language Resources and Evaluation Conference, pages 9535–9548, Palma de Mallorca, Spain. ELRA Language Resource Association.
- Cite (Informal):
- A Critical Study of Automatic Evaluation in Sign Language Translation (Yazdani et al., LREC 2026)