Sociolinguistic aspects of crowdsourcing for a vocal corpus of Alsatian
Pascale Erhart, Lucile Hamm, Sam Bigeard, Carole Werner, Malek Yaich, Slim Ouni
Abstract
Alsatian is a regional low-resource language spoken in a majority-language context. In order to create a voice dataset suited for training automatic speech recognition and speech-to-text models, we launched a crowdsourcing campaign on the platform Mozilla Common Voice. We describe sociolinguistic issues we ran into, such as participants’ perception of their own language and its role in the AI landscape, which are vital to address to raise the participation in the crowdsourcing effort. We found that the participants are often confused about NLP and AI tools, and have a strong interested in preserving their language.- Anthology ID:
- 2026.dialres-1.25
- Volume:
- Proceedings of the First Workshop on Dialects in NLP — A Resource Perspective
- Month:
- May
- Year:
- 2026
- Address:
- Palma de Mallorca
- Editors:
- Antonis Anastasopoulos, Stella Markantonatou, Angela Ralli, Marcos Zampieri, Stavros Bompolas, Vivian Stamou
- Venues:
- DialRes | WS
- SIG:
- Publisher:
- Association for Computational Linguistics
- Note:
- Pages:
- 256–264
- Language:
- External URL:
- https://lrec.elra.info/lrec2026-ws-dialres-25
- DOI:
- 10.63317/5ch7cwah438g
- Cite (ACL):
- Pascale Erhart, Lucile Hamm, Sam Bigeard, Carole Werner, Malek Yaich, and Slim Ouni. 2026. Sociolinguistic aspects of crowdsourcing for a vocal corpus of Alsatian. In Proceedings of the First Workshop on Dialects in NLP — A Resource Perspective, pages 256–264, Palma de Mallorca. Association for Computational Linguistics.
- Cite (Informal):
- Sociolinguistic aspects of crowdsourcing for a vocal corpus of Alsatian (Erhart et al., DialRes 2026)