Development of Text and Speech database for Hindi and Indian English specific to Mobile Communication environment

Shyam Agrawal, Shweta Sinha, Pooja Singh, Jesper Olson


Abstract
This paper describes the method and experiences of text and speech data collection in mobile communication in Indian English Hindi. The primary data collection is done in the form of large number of messages as part of Personal communication among natives of Hindi language and Indian speakers of English. To gather the versatility of mobile communication database among Hindi and English, 12 domains were identified for collection of text corpus from speaking population belonging to deferent age groups, sex and dialects. The text obtained in raw form based on slangs and unconventional grammar were cleaned using on language grammar rules and then tagged and expanded to explain context specific meaning of the words. Texts of 1163 participants from Hindi speaking regions and 1405 English users were taken for creating 13 prompt sheets; containing 630 phonetically rich sentences created using a special software. Each prompt sheet was recorded by at least 7 users simultaneously in three channels and recorded by a total of 100 speakers and annotated. The work is a step forward in the direction of development of standards for mobile text and speech data collection for Indian languages. Keywords - Speech data base, Text analysis, mobile communication, Hindi and Indian English Speech, multi-lingual speech processing.
Anthology ID:
L12-1670
Volume:
Proceedings of the Eighth International Conference on Language Resources and Evaluation (LREC'12)
Month:
May
Year:
2012
Address:
Istanbul, Turkey
Editors:
Nicoletta Calzolari, Khalid Choukri, Thierry Declerck, Mehmet Uğur Doğan, Bente Maegaard, Joseph Mariani, Asuncion Moreno, Jan Odijk, Stelios Piperidis
Venue:
LREC
SIG:
Publisher:
European Language Resources Association (ELRA)
Note:
Pages:
3415–3421
Language:
URL:
http://www.lrec-conf.org/proceedings/lrec2012/pdf/1132_Paper.pdf
DOI:
Bibkey:
Cite (ACL):
Shyam Agrawal, Shweta Sinha, Pooja Singh, and Jesper Olson. 2012. Development of Text and Speech database for Hindi and Indian English specific to Mobile Communication environment. In Proceedings of the Eighth International Conference on Language Resources and Evaluation (LREC'12), pages 3415–3421, Istanbul, Turkey. European Language Resources Association (ELRA).
Cite (Informal):
Development of Text and Speech database for Hindi and Indian English specific to Mobile Communication environment (Agrawal et al., LREC 2012)
Copy Citation:
PDF:
http://www.lrec-conf.org/proceedings/lrec2012/pdf/1132_Paper.pdf