Martin Wynne
2026
GaelEval: Benchmarking LLM Performance for Scottish Gaelic
Peter Devine | William Lamb | Beatrice Alex | Ignatius Ezeani | Dawn Knight | Mícheál J. Ó Meachair | Paul Rayson | Martin Wynne
Proceedings of Shaping Multilingual, Multimodal AI for the Social Sciences and Humanities (LLMs4SSH) @ LREC 2026
Peter Devine | William Lamb | Beatrice Alex | Ignatius Ezeani | Dawn Knight | Mícheál J. Ó Meachair | Paul Rayson | Martin Wynne
Proceedings of Shaping Multilingual, Multimodal AI for the Social Sciences and Humanities (LLMs4SSH) @ LREC 2026
Multilingual large language models (LLMs) often exhibit emergent ‘shadow’ capabilities in languages without official support, yet their performance on these languages remains uneven and under-measured. This is particularly acute for morphosyntactically rich minority languages such as Scottish Gaelic, where translation benchmarks fail to capture structural competence. We introduce GaelEval, the first multi-dimensional benchmark for Gaelic, comprising: (i) an expert-authored morphosyntactic MCQA task; (ii) a culturally-grounded translation benchmark and (iii) a large-scale cultural knowledge Q&A task. Evaluating 19 LLMs against a fluent-speaker human baseline (n = 30), we find that Gemini 3 Pro Preview achieves 83.3% accuracy on the linguistic task, surpassing the human baseline (78.1%). Proprietary models consistently outperform open-weight systems, and in-language (Gaelic) prompting yields a small but stable advantage (+2.4pp). On the cultural task, leading models exceed 90% accuracy, though most systems perform worse under Gaelic prompting and absolute scores are inflated relative to the manual benchmark. Overall, GaelEval reveals that frontier models achieve above-human performance on several dimensions of Gaelic grammar, demonstrates the effect of Gaelic prompting and shows a consistent performance gap favouring proprietary over open-weight models.
Proceedings of The Second Workshop on Holocaust Testimonies as Language Resources (HTRes)
Isuri Anuradha | Martin Wynne
Proceedings of The Second Workshop on Holocaust Testimonies as Language Resources (HTRes)
Isuri Anuradha | Martin Wynne
Proceedings of The Second Workshop on Holocaust Testimonies as Language Resources (HTRes)
The British National Corpus 1994 to 2026
Martin Wynne | Megan Bushnell
Proceedings of the 12th Workshop on Challenges in the Management of Large Corpora
Martin Wynne | Megan Bushnell
Proceedings of the 12th Workshop on Challenges in the Management of Large Corpora
It was not possible to address the objections of reviewer #1 in an updated version of the paper. They seem to wish for a different paper about AI. Our brief was for "submit a poster on what has changed (and what is possibly about to change)" with respect to the BNC. The BNC is a historical corpus not an ongoing project. Reviewer #2 also notes a lack of novelty, but the same applies as above. We have added a little more on ’lessons learned’. Reviewer #3 asks about the availablility on CD of the corpus. This was phased out more than ten years ago, and it is already stated in the paper that the corpus is available for download from the Oxford Text Archive. Likewise, the TEI annotation of the corpus was frozen on release. This point is emphasised in a revision. The importance of the corpus as a TEI flagship project is noted. The typos reported by reviewer #3 have been corrected. Further footnotes with URLs, references to publications and language resources have been added.
2024
Proceedings of the First Workshop on Holocaust Testimonies as Language Resources (HTRes) @ LREC-COLING 2024
Isuri Anuradha | Martin Wynne | Francesca Frontini | Alistair Plum
Proceedings of the First Workshop on Holocaust Testimonies as Language Resources (HTRes) @ LREC-COLING 2024
Isuri Anuradha | Martin Wynne | Francesca Frontini | Alistair Plum
Proceedings of the First Workshop on Holocaust Testimonies as Language Resources (HTRes) @ LREC-COLING 2024
2010
Resource and Service Centres as the Backbone for a Sustainable Service Infrastructure
Peter Wittenburg | Nuria Bel | Lars Borin | Gerhard Budin | Nicoletta Calzolari | Eva Hajicova | Kimmo Koskenniemi | Lothar Lemnitzer | Bente Maegaard | Maciej Piasecki | Jean-Marie Pierrel | Stelios Piperidis | Inguna Skadina | Dan Tufis | Remco van Veenendaal | Tamas Váradi | Martin Wynne
Proceedings of the Seventh International Conference on Language Resources and Evaluation (LREC'10)
Peter Wittenburg | Nuria Bel | Lars Borin | Gerhard Budin | Nicoletta Calzolari | Eva Hajicova | Kimmo Koskenniemi | Lothar Lemnitzer | Bente Maegaard | Maciej Piasecki | Jean-Marie Pierrel | Stelios Piperidis | Inguna Skadina | Dan Tufis | Remco van Veenendaal | Tamas Váradi | Martin Wynne
Proceedings of the Seventh International Conference on Language Resources and Evaluation (LREC'10)
Currently, research infrastructures are being designed and established in many disciplines since they all suffer from an enormous fragmentation of their resources and tools. In the domain of language resources and tools the CLARIN initiative has been funded since 2008 to overcome many of the integration and interoperability hurdles. CLARIN can build on knowledge and work from many projects that were carried out during the last years and wants to build stable and robust services that can be used by researchers. Here service centres will play an important role that have the potential of being persistent and that adhere to criteria as they have been established by CLARIN. In the last year of the so-called preparatory phase these centres are currently developing four use cases that can demonstrate how the various pillars CLARIN has been working on can be integrated. All four use cases fulfil the criteria of being cross-national.
2008
CLARIN: Common Language Resources and Technology Infrastructure
Tamás Váradi | Steven Krauwer | Peter Wittenburg | Martin Wynne | Kimmo Koskenniemi
Proceedings of the Sixth International Conference on Language Resources and Evaluation (LREC'08)
Tamás Váradi | Steven Krauwer | Peter Wittenburg | Martin Wynne | Kimmo Koskenniemi
Proceedings of the Sixth International Conference on Language Resources and Evaluation (LREC'08)
The paper provides a general introduction to the CLARIN project, a large-scale European research infrastructure project designed to establish an integrated and interoperable infrastructure of language resources and technologies. The goal is to make language resources and technology much more accessible to all researchers working with language material, particularly non-expert users in the Humanities and Social Sciences. CLARIN intends to build a virtual, distributed infrastructure consisting of a federation of trusted digital archives and repositories where language resources and tools are accessible through web services. The CLARIN project consists of 32 partners from 22 countries and is currently engaged in the preparatory phase of developing the infrastructure. The paper describes the objectives of the project in terms of its technical, legal, linguistic and user dimensions.
2002
The Language Resource Archive of the 21st Century
Martin Wynne
Proceedings of the Third International Conference on Language Resources and Evaluation (LREC’02)
Martin Wynne
Proceedings of the Third International Conference on Language Resources and Evaluation (LREC’02)
2001
Search
Fix author
Co-authors
- Isuri Anuradha 2
- Kimmo Koskenniemi 2
- Tamás Váradi 2
- Peter Wittenburg 2
- Beatrice Alex 1
- Núria Bel 1
- Lars Borin 1
- Gerhard Budin 1
- Megan Bushnell 1
- Nicoletta Calzolari 1
- Peter Devine 1
- Ignatius Ezeani 1
- Francesca Frontini 1
- Eva Hajicova 1
- Dawn Knight 1
- Steven Krauwer 1
- William Lamb 1
- Lothar Lemnitzer 1
- Bente Maegaard 1
- Maciej Piasecki 1
- Jean-Marie Pierrel 1
- Stelios Piperidis 1
- Alistair Plum 1
- Paul Rayson 1
- Inguna Skadiņa 1
- Dan Tufiş 1
- Remco van Veenendaal 1
- Mícheál J. Ó Meachair 1