The Corpus of Contemporary Polish: 2011-2020 Decade and Beyond

Witold Kieraś, Małgorzata Marciniak, Katarzyna Krasnowska-Kieraś, Marcin Woliński


Abstract
Several questions and suggestions from the Reviewers concerned providing more details about the corpus design, composition and annotation. We submitted our poster’s accompanying paper as an extended abstract since a paper presenting the corpus was accepted at the main LREC 2026 conference. In the below response, we omit the Reviewers’ remarks that called for more detailed explanations or examples that are provided in the full paper. Once anonymity is not an issue at publication time, we added a reference to the paper in our CMLC abstract. Below we address the remaining points raised by the Reviewers. REVIEWER 1 Q: KWJP limits the diversity of texts to three main categories: fiction, non-fiction, and journalistic writing. This raises the question of whether such a restriction affects comparability with NKJP, which included a wider range of text types, including spoken and internet communication, and in particular those that are less represented in specialized corpora, such as correspondence and poetry. A: Any difference in data selection restricts the comparability of corpora. Spoken language and internet communication registers call for significantly different approaches to data collection, selection and annotation processes. As for smaller and less represented text types: to our knowledge both poetry and correspondence in the balanced NKJP corpus are represented by three books each. REVIEWER 3 Q: For the planned 2026 update, will the new sub-corpus maintain the same annotation layers or are there planned expansions (e.g. more fine-grained NE types)? A: We are not planning any new or modified layers of annotation at this point. The planned update refers to textual data as well as technical efficiency of the web service. Q: Is there consideration for releasing larger open subsets or integrating the data into platforms like Sketch Engine for broader academic access? A: We clarified the statement about copyright issues prohibiting us from releasing larger amounts of data. The same legal considerations do not allow us to upload the data to external services such as Sketch Engine, effectively distributing them to third parties. REVIEWER 4 Q: What types of named entity have been annotated in the corpus? A: As it was mentioned in the abstract, we are using a NE annotation schema developed in the NKJP project. The details may be found in relevant publications concerning this corpus. Q: Is "the development of web-based applications" really necessary? Are you not able to use existing solutions without reinventing the wheel? A: We did not state that we are developing a web-based application, rather we are developing a new solution (within our web service) to provide a more efficient searching mechanism. One of the paths explored in the process is using the BlackLab search engine. It is possible that we will index the corpus in two different search engines and allow the user to choose, but no final decisions have been made yet.
Anthology ID:
2026.cmlc-1.12
Volume:
Proceedings of the 12th Workshop on Challenges in the Management of Large Corpora
Month:
May
Year:
2026
Address:
Palma, Mallorca (Spain)
Editors:
Piotr Bański, Dawn Knight, Marc Kupietz, Andreas Witt, Alina Wróblewska
Venues:
CMLC | WS
SIG:
Publisher:
ELRA Language Resources Association (ELRA)
Note:
Pages:
78–79
Language:
External URL:
https://lrec.elra.info/lrec2026-ws-cmlc-12
DOI:
10.63317/35cjnfgvskp4
Bibkey:
Cite (ACL):
Witold Kieraś, Małgorzata Marciniak, Katarzyna Krasnowska-Kieraś, and Marcin Woliński. 2026. The Corpus of Contemporary Polish: 2011-2020 Decade and Beyond. In Proceedings of the 12th Workshop on Challenges in the Management of Large Corpora, pages 78–79, Palma, Mallorca (Spain). ELRA Language Resources Association (ELRA).
Cite (Informal):
The Corpus of Contemporary Polish: 2011-2020 Decade and Beyond (Kieraś et al., CMLC 2026)
Copy Citation: