Abstract
Corpus design for speech synthesis is a well-researched topic in languages such as English compared to Modern Standard Arabic, and there is a tendency to focus on methods to automatically generate the orthographic transcript to be recorded (usually greedy methods). In this work, a study of Modern Standard Arabic (MSA) phonetics and phonology is conducted in order to create criteria for a greedy method to create a speech corpus transcript for recording. The size of the dataset is reduced a number of times using these optimisation methods with different parameters to yield a much smaller dataset with identical phonetic coverage than before the reduction, and this output transcript is chosen for recording. This is part of a larger work to create a completely annotated and segmented speech corpus for MSA.- Anthology ID:
- L16-1116
- Volume:
- Proceedings of the Tenth International Conference on Language Resources and Evaluation (LREC'16)
- Month:
- May
- Year:
- 2016
- Address:
- Portorož, Slovenia
- Editors:
- Nicoletta Calzolari, Khalid Choukri, Thierry Declerck, Sara Goggi, Marko Grobelnik, Bente Maegaard, Joseph Mariani, Helene Mazo, Asuncion Moreno, Jan Odijk, Stelios Piperidis
- Venue:
- LREC
- SIG:
- Publisher:
- European Language Resources Association (ELRA)
- Note:
- Pages:
- 734–738
- Language:
- URL:
- https://aclanthology.org/L16-1116
- DOI:
- Cite (ACL):
- Nawar Halabi and Mike Wald. 2016. Phonetic Inventory for an Arabic Speech Corpus. In Proceedings of the Tenth International Conference on Language Resources and Evaluation (LREC'16), pages 734–738, Portorož, Slovenia. European Language Resources Association (ELRA).
- Cite (Informal):
- Phonetic Inventory for an Arabic Speech Corpus (Halabi & Wald, LREC 2016)
- PDF:
- https://preview.aclanthology.org/fix-dup-bibkey/L16-1116.pdf