Large Corpus of Czech Parliament Plenary Hearings

Jonáš Kratochvíl; Peter Polák; Ondřej Bojar

Large Corpus of Czech Parliament Plenary Hearings

Jonas Kratochvil, Peter Polak, Ondrej Bojar

Abstract

We present a large corpus of Czech parliament plenary sessions. The corpus consists of approximately 1200 hours of speech data and corresponding text transcriptions. The whole corpus has been segmented to short audio segments making it suitable for both training and evaluation of automatic speech recognition (ASR) systems. The source language of the corpus is Czech, which makes it a valuable resource for future research as only a few public datasets are available in the Czech language. We complement the data release with experiments of two baseline ASR systems trained on the presented data: the more traditional approach implemented in the Kaldi ASRtoolkit which combines hidden Markov models and deep neural networks (NN) and a modern ASR architecture implemented in Jaspertoolkit which uses deep NNs in an end-to-end fashion.

Anthology ID:: 2020.lrec-1.781
Volume:: Proceedings of the 12th Language Resources and Evaluation Conference
Month:: May
Year:: 2020
Address:: Marseille, France
Venue:: LREC
SIG:
Publisher:: European Language Resources Association
Note:
Pages:: 6363–6367
Language:: English
URL:: https://aclanthology.org/2020.lrec-1.781
DOI:
Bibkey:
Cite (ACL):: Jonas Kratochvil, Peter Polak, and Ondrej Bojar. 2020. Large Corpus of Czech Parliament Plenary Hearings. In Proceedings of the 12th Language Resources and Evaluation Conference, pages 6363–6367, Marseille, France. European Language Resources Association.
Cite (Informal):: Large Corpus of Czech Parliament Plenary Hearings (Kratochvil et al., LREC 2020)
Copy Citation:
PDF:: https://preview.aclanthology.org/update-css-js/2020.lrec-1.781.pdf

PDF Cite Search