Małgorzata Marciniak

Papers on this page may belong to the following people: Malgorzata Marciniak, Małgorzata Marciniak


2026

Several questions and suggestions from the Reviewers concerned providing more details about the corpus design, composition and annotation. We submitted our poster’s accompanying paper as an extended abstract since a paper presenting the corpus was accepted at the main LREC 2026 conference. In the below response, we omit the Reviewers’ remarks that called for more detailed explanations or examples that are provided in the full paper. Once anonymity is not an issue at publication time, we added a reference to the paper in our CMLC abstract. Below we address the remaining points raised by the Reviewers. REVIEWER 1 Q: KWJP limits the diversity of texts to three main categories: fiction, non-fiction, and journalistic writing. This raises the question of whether such a restriction affects comparability with NKJP, which included a wider range of text types, including spoken and internet communication, and in particular those that are less represented in specialized corpora, such as correspondence and poetry. A: Any difference in data selection restricts the comparability of corpora. Spoken language and internet communication registers call for significantly different approaches to data collection, selection and annotation processes. As for smaller and less represented text types: to our knowledge both poetry and correspondence in the balanced NKJP corpus are represented by three books each. REVIEWER 3 Q: For the planned 2026 update, will the new sub-corpus maintain the same annotation layers or are there planned expansions (e.g. more fine-grained NE types)? A: We are not planning any new or modified layers of annotation at this point. The planned update refers to textual data as well as technical efficiency of the web service. Q: Is there consideration for releasing larger open subsets or integrating the data into platforms like Sketch Engine for broader academic access? A: We clarified the statement about copyright issues prohibiting us from releasing larger amounts of data. The same legal considerations do not allow us to upload the data to external services such as Sketch Engine, effectively distributing them to third parties. REVIEWER 4 Q: What types of named entity have been annotated in the corpus? A: As it was mentioned in the abstract, we are using a NE annotation schema developed in the NKJP project. The details may be found in relevant publications concerning this corpus. Q: Is "the development of web-based applications" really necessary? Are you not able to use existing solutions without reinventing the wheel? A: We did not state that we are developing a web-based application, rather we are developing a new solution (within our web service) to provide a more efficient searching mechanism. One of the paths explored in the process is using the BlackLab search engine. It is possible that we will index the corpus in two different search engines and allow the user to choose, but no final decisions have been made yet.