A Critical Study of Automatic Evaluation in Sign Language Translation

Shakib Yazdani, Yasser HAMIDULLAH, Cristina España-Bonet, Eleftherios Avramidis, Josef van Genabith


Abstract
Automatic evaluation metrics are crucial for advancing sign language translation (SLT). Current SLT evaluation metrics, such as BLEU and ROUGE, are only text-based, and it remains unclear to what extent text-based metrics can reliably capture the quality of SLT outputs. To address this gap, we investigate the limitations of text-based SLT evaluation metrics by analyzing six metrics, including BLEU, chrF, and ROUGE, as well as BLEURT on the one hand, and large language model (LLM)-based evaluators such as G-Eval and GEMBA zero-shot direct assessment on the other hand. Specifically, we assess the consistency and robustness of these metrics under three controlled conditions: paraphrasing, hallucinations in model outputs, and variations in sentence length. Our analysis highlights the limitations of lexical overlap metrics and demonstrates that while LLM-based evaluators better capture semantic equivalence often missed by conventional metrics, they can also exhibit bias toward LLM-paraphrased translations. Moreover, although all metrics are able to detect hallucinations, BLEU tends to be overly sensitive, whereas BLEURT and LLM-based evaluators are comparatively lenient toward subtle cases. This motivates the need for multimodal evaluation frameworks that extend beyond text-based metrics to enable a more holistic assessment of SLT outputs.
Anthology ID:
2026.lrec-1.749
Volume:
Proceedings of the Fifteenth Language Resources and Evaluation Conference
Month:
May
Year:
2026
Address:
Palma de Mallorca, Spain
Editors:
Stelios Piperidis, Núria Bel, Henk van den Heuvel, Nancy Ide, Simon Krek, Antonio Toral
Venue:
LREC
SIG:
Publisher:
ELRA Language Resource Association
Note:
Pages:
9535–9548
Language:
External URL:
https://lrec.elra.info/lrec2026-main-749
DOI:
10.63317/4n2sooe4fb2i
Bibkey:
Cite (ACL):
Shakib Yazdani, Yasser HAMIDULLAH, Cristina España-Bonet, Eleftherios Avramidis, and Josef van Genabith. 2026. A Critical Study of Automatic Evaluation in Sign Language Translation. In Proceedings of the Fifteenth Language Resources and Evaluation Conference, pages 9535–9548, Palma de Mallorca, Spain. ELRA Language Resource Association.
Cite (Informal):
A Critical Study of Automatic Evaluation in Sign Language Translation (Yazdani et al., LREC 2026)
Copy Citation: