GaelEval: Benchmarking LLM Performance for Scottish Gaelic

Peter Devine, William Lamb, Beatrice Alex, Ignatius Ezeani, Dawn Knight, Mícheál J. Ó Meachair, Paul Rayson, Martin Wynne


Abstract
Multilingual large language models (LLMs) often exhibit emergent ‘shadow’ capabilities in languages without official support, yet their performance on these languages remains uneven and under-measured. This is particularly acute for morphosyntactically rich minority languages such as Scottish Gaelic, where translation benchmarks fail to capture structural competence. We introduce GaelEval, the first multi-dimensional benchmark for Gaelic, comprising: (i) an expert-authored morphosyntactic MCQA task; (ii) a culturally-grounded translation benchmark and (iii) a large-scale cultural knowledge Q&A task. Evaluating 19 LLMs against a fluent-speaker human baseline (n = 30), we find that Gemini 3 Pro Preview achieves 83.3% accuracy on the linguistic task, surpassing the human baseline (78.1%). Proprietary models consistently outperform open-weight systems, and in-language (Gaelic) prompting yields a small but stable advantage (+2.4pp). On the cultural task, leading models exceed 90% accuracy, though most systems perform worse under Gaelic prompting and absolute scores are inflated relative to the manual benchmark. Overall, GaelEval reveals that frontier models achieve above-human performance on several dimensions of Gaelic grammar, demonstrates the effect of Gaelic prompting and shows a consistent performance gap favouring proprietary over open-weight models.
Anthology ID:
2026.llms4ssh-1.8
Volume:
Proceedings of Shaping Multilingual, Multimodal AI for the Social Sciences and Humanities (LLMs4SSH) @ LREC 2026
Month:
May
Year:
2026
Address:
Palma de Mallorca (Spain)
Editors:
Arturo Montejo-Raez, Cristina Grisot, Joanna Blochowiak, Nikola Ljubešić, Elena Battaner, German Rigau
Venues:
LLMs4SSH | WS
SIG:
Publisher:
ELRA Language Resources Association (ELRA)
Note:
Pages:
73–85
Language:
External URL:
https://lrec.elra.info/lrec2026-ws-llms4ssh-08
DOI:
10.63317/2yqmx3es6ad5
Bibkey:
Cite (ACL):
Peter Devine, William Lamb, Beatrice Alex, Ignatius Ezeani, Dawn Knight, Mícheál J. Ó Meachair, Paul Rayson, and Martin Wynne. 2026. GaelEval: Benchmarking LLM Performance for Scottish Gaelic. In Proceedings of Shaping Multilingual, Multimodal AI for the Social Sciences and Humanities (LLMs4SSH) @ LREC 2026, pages 73–85, Palma de Mallorca (Spain). ELRA Language Resources Association (ELRA).
Cite (Informal):
GaelEval: Benchmarking LLM Performance for Scottish Gaelic (Devine et al., LLMs4SSH 2026)
Copy Citation: