Abstract
Like most other minority languages, Scottish Gaelic has limited tools and resources available for Natural Language Processing research and applications. These limitations restrict the potential of the language to participate in modern speech technology, while also restricting research in fields such as corpus linguistics and the Digital Humanities. At the same time, Gaelic has a long written history, is well-described linguistically, and is unusually well-supported in terms of potential NLP training data. For instance, archives such as the School of Scottish Studies hold thousands of digitised recordings of vernacular speech, many of which have been transcribed as paper-based, handwritten manuscripts. In this paper, we describe a project to digitise and recognise a corpus of handwritten narrative transcriptions, with the intention of re-purposing it to develop a Gaelic speech recognition system.- Anthology ID:
- 2022.cltw-1.9
- Volume:
- Proceedings of the 4th Celtic Language Technology Workshop within LREC2022
- Month:
- June
- Year:
- 2022
- Address:
- Marseille, France
- Editors:
- Theodorus Fransen, William Lamb, Delyth Prys
- Venue:
- CLTW
- SIG:
- Publisher:
- European Language Resources Association
- Note:
- Pages:
- 60–70
- Language:
- URL:
- https://aclanthology.org/2022.cltw-1.9
- DOI:
- Cite (ACL):
- William Lamb, Beatrice Alex, and Mark Sinclair. 2022. Handwriting recognition for Scottish Gaelic. In Proceedings of the 4th Celtic Language Technology Workshop within LREC2022, pages 60–70, Marseille, France. European Language Resources Association.
- Cite (Informal):
- Handwriting recognition for Scottish Gaelic (Lamb et al., CLTW 2022)
- PDF:
- https://preview.aclanthology.org/naacl24-info/2022.cltw-1.9.pdf