Hana Skoumalova

Also published as: Hana Skoumalová


2016

pdf
SYN2015: Representative Corpus of Contemporary Written Czech
Michal Křen | Václav Cvrček | Tomáš Čapka | Anna Čermáková | Milena Hnátková | Lucie Chlumská | Tomáš Jelínek | Dominika Kováříková | Vladimír Petkevič | Pavel Procházka | Hana Skoumalová | Michal Škrabal | Petr Truneček | Pavel Vondřička | Adrian Jan Zasina
Proceedings of the Tenth International Conference on Language Resources and Evaluation (LREC'16)

The paper concentrates on the design, composition and annotation of SYN2015, a new 100-million representative corpus of contemporary written Czech. SYN2015 is a sequel of the representative corpora of the SYN series that can be described as traditional (as opposed to the web-crawled corpora), featuring cleared copyright issues, well-defined composition, reliability of annotation and high-quality text processing. At the same time, SYN2015 is designed as a reflection of the variety of written Czech text production with necessary methodological and technological enhancements that include a detailed bibliographic annotation and text classification based on an updated scheme. The corpus has been produced using a completely rebuilt text processing toolchain called SynKorp. SYN2015 is lemmatized, morphologically and syntactically annotated with state-of-the-art tools. It has been published within the framework of the Czech National Corpus and it is available via the standard corpus query interface KonText at http://kontext.korpus.cz as well as a dataset in shuffled format.

2015

pdf bib
Analytic Morphology – Merging the Paradigmatic and Syntagmatic Perspective in a Treebank
Vladimír Petkevič | Alexandr Rosen | Hana Skoumalová | Přemysl Vítovec
The 5th Workshop on Balto-Slavic Natural Language Processing

2014

pdf
The SYN-series corpora of written Czech
Milena Hnátková | Michal Křen | Pavel Procházka | Hana Skoumalová
Proceedings of the Ninth International Conference on Language Resources and Evaluation (LREC'14)

The paper overviews the SYN series of synchronic corpora of written Czech compiled within the framework of the Czech National Corpus project. It describes their design and processing with a focus on the annotation, i.e. lemmatization and morphological tagging. The paper also introduces SYN2013PUB, a new 935-million newspaper corpus of Czech published in 2013 as the most recent addition to the SYN series before planned revision of its architecture. SYN2013PUB can be seen as a completion of the series in terms of titles and publication dates of major Czech newspapers that are now covered by complete volumes in comparable proportions. All SYN-series corpora can be characterized as traditional, with emphasis on cleared copyright issues, well-defined composition, reliable metadata and high-quality data processing; their overall size currently exceeds 2.2 billion running words.

2000

pdf
Resources for Multilingual Text Generation in Three Slavic Languages
John Bateman | Elke Teich | Geert-Jan Kruijff | Ivana Kruijff-Korbayová | Serge Sharoff | Hana Skoumalová
Proceedings of the Second International Conference on Language Resources and Evaluation (LREC’00)

pdf
Multilinguality in a Text Generation System For Three Slavic Languages
Geert-Jan Kruijff | Elke Teich | John Bateman | Ivana Kruijff-Korbayova | Hana Skoumalova | Serge Sharoff | Lena Sokolova | Tony Hartley | Kamenka Staykova | Jiri Hana
COLING 2000 Volume 1: The 18th International Conference on Computational Linguistics

1997

pdf
A Czech Morphological Lexicon
Hana Skoumalova
Computational Phonology: Third Meeting of the ACL Special Interest Group in Computational Phonology

1995

pdf
Identifying Topic and Focus by an Automatic Procedure
Eva Hajicova | Hana Skoumalova | Petr Sgall
Computational Linguistics, Volume 21, Number 1, March 1995

1992

pdf
Surface and Deep Cases
Jarmila Panevová | Hana Skoumalova
COLING 1992 Volume 3: The 14th International Conference on Computational Linguistics