Proceedings of the 12th Workshop on Challenges in the Management of Large Corpora

Piotr Bański, Dawn Knight, Marc Kupietz, Andreas Witt, Alina Wróblewska (Editors)



We revised the manuscript following all the suggestions made by the reviewers.
Many thanks to all reviewers for the detailed comments and suggestions. Misspelling, formatting errors and other minor issues have all been corrected, and are not listed below. 1. (Reviewer 1) Comment: For example, the relation of some described corpora to the Bulgarian National Corpus (like MIC21, Bulgarian MARCELL, General News in Bulgarian, ...) is not clear to me. Are they part of BulNC or the other way around? It’s stated that the BulNC is part of CURLICAT, so I would assume the relation with the other corpora is similar? Response: The conclusion now clarifies the relations between BulNC and IfGPT and its subsets. The paper was also restructured to clarify these issues. 2. (Reviewer 1) Comment: Some sections on the paper reference related work, where again the relation is not clear to me, like on page 3 "Another direction, which still presents significant challenges, is towards multilingual data. O’Keeffe et al. (2024) describe a pipeline for capturing professional video-call interactions, including screen recordings, speaker tracking, and facial expression data. Macaire et al. (2024) present a speech text pictogram corpus for French (230 hours), targeted at augmentative and alternative communication research. Lai and Pustejovsky (2024) develop an annotation scheme for iconic, deictic, and beat gestures anchored in Abstract Meaning Representation." As these publications follow the MIC21 corpus, it’s not clear to me, how they relate. Response: These references are reduced; a citation for the MIC corpus is provided where more details are available on the related works of MIC. 3. (Reviewer 1) Comment: I would be very interested in the benefits of a Graph database for the Metadata and I think it would be very valuable to describe, why the Graph database is more appropriate for Metadata than commonly used alternatives. The described relations are not totally convincing to me. Response: A paragraph is added in the Metadata management section on the justification of the use of a graph database. 4. (Reviewer 1) Comment: The publicly accessible web interface should be linked to. It’s also not clear to me, if it provides a fulltext search or only a keyword search in the Metadata (i.e. in the Graph database). Response: Links are provided to both the search interface of BulNC (full text search, mainly for linguistic research) and to IfGPT metadata search interface (allowing selection of subdatasets for NLP tasks and LLM fine-tuning). 5. (Reviewer 2) Comment: My first doubt is why all the data is presented as part or at least related to the Bulgarian National Corpus - to me, national corpora are reference language corpora with all the characteristics this entails, like a carefully balanced corpus representative of contemporary standard Bulgarian. I would not expect to see in this context mentioned artificially produced language for LLMs, multilingual or foreign language corpora or image corpora. To me it would make a lot more sense the re-cast the paper as presenting (newly) available language resources for Bulgarian (maybe in the context of the Bulgarian CLARIN / CLARIAH, which is not even mentioned!) rather than shoe-horning them to the BulNC. Response: The ties between the BulNC and the large dataset IfGPT has been clarified in the Introduction, in the text and in the Conclusion. 6. (Reviewer 2) Comment: Second, it would help the reader that, rather than just listing all the resources, a table with their key characteristics would be provided first, and then each introduced. Similarly, the endless repetitions of "so-and-so many JSON files" for each domain in Sec. 2.5 are not really helpful, as, first, texts or words are a better metric than files, and, second, this information would also be better presented in a table. Response: A table has now been provided listing the resources and their key properties, including size. 7. (Reviewer 2) Comment: "The BulNC-based dataset is publicly accessible through a dedicated web interface" I don’t see that these datasets are in any way BulNC-based. Response: Clarifications are made on this issue in the Introduction and the Conclusion. 8. (Reviewer 3) Comment: I am missing (at least a sketchy) description of tools used for tagging and/or parsing the corpus data. Response: A reference to the Bulgarian Language Processing Tool Set is now provided. 9. (Reviewer 3) Comment: It is also not clear why the corpus is still using the web interface developed before 2014, and not any newer tool (with a FLOSS license, such as CQPweb or NoSketch Engine). Response: This has been clarified in section 2.1. 10. (Reviewer 3) Comment: I would also expect the text to include summary statistics (preferably in a single table), so the sizes of the respective resources can be compared. Response: Table 1 is now provided for this purpose.
We present Merimënga, a pipeline for reproducible Albanian web-corpus construction from Common Crawl. Rather than distributing a static text dump, we publish versioned manifests and append-only JSONL ledgers that make every retrieval and filtering decision replayable at record level. Records are addressed by (WARC filename, byte offset, byte length) and retrieved via HTTP range requests with checksum validation, enabling selective download, resumability, and exact re-materialization. On top of deterministic cleaning and deduplication, Merimënga supports teacher–student filtering: a large LLM labels a stratified sample; the resulting policy is distilled into a faster student model applied at corpus scale. The paper contributes (i) a reproducibility specification for web-corpus construction based on coordinate-addressed retrieval and decision ledgers, (ii) a concrete instantiation for Albanian with language-specific filtering, and (iii) an evaluation protocol for rerun equivalence and filter-stack ablation. Large-scale download and full-corpus filtering are ongoing; this submission focuses on methodology and auditable artifacts rather than final corpus statistics. Keywords:Common Crawl, Reproducibility, Corpus Construction, Learned Filtering, Albanian
I would like to sincerely thank all reviewers for their time, expertise, and thoughtful feedback on my paper! Please find below my responses to your remarks. ============================================================================ REVIEWER #1 ============================================================================ REMARK 1: Since this paper has been submitted to the CMLC workshop, I would have liked to see more explicit discussions of how the work is linked to challenges in the management of large corpora.... ANSWER: Added a paragraph on this in the conclusion section. REMARK 2: It would also be good to explain how the Songkorpus handle code switching in songs between German, English and other languages... ANSWER: Added a paragraph in Section 2.1. REMARK 3: As new songs are being added, how is IP handled? Does this have to be cleared before lyrics are included in updates to the corpus each year e.g. from the current chart? ANSWER: Added a paragraph in Section 2.1. ============================================================================ REVIEWER #2 ============================================================================ REMARK 1: Section 2.5: The methods used to estimate sentiment intensity and detect MPs are somewhat outdated. However, the human-based evaluation of their quality and reliability is very thorough and convincing. Since manually annotated samples are available, it may be worthwhile to test more modern approaches... ANSWER: Expanded Section 2.4.2 accordingly. REMARK 2: Section 3.2: The conclusions presented in this section are very interesting. It might be useful to examine which pronoun has become more prominent over time... ANSWER: Added a paragraph in Section 3.2. REMARK 3: Examples: I strongly recommend adding concrete examples that illustrate the linguistic features. ANSWER: Added a paragraph in Section 2.3.3. REMARK 4: Figures: The font size in the figures is too small and hard to read. ANSWER: Changed! ============================================================================ REVIEWER #3 ============================================================================ REMARK 1: Could you make more explicit whether the paper’s primary contribution is methodological, substantive, or intentionally balanced between the two? ANSWER: Clarified the contribution in the introduction, then reframed the research questions in Section 1.2. REMARK 2: Could you say a little more about the intended interpretive payoff of features such as pronoun usage and modal particles in this domain? ANSWER: Added a short motivation paragraph in the introduction that directly answers "why these features?", and tried to strengthen interpretive payoff in the Results discussion REMARK 3: How confident can we be in interpreting the results as showing cultural features such as emotional flattening? ANSWER: Please see the new remarks on this in the conclusion.
The rapid advancement of digital humanities and Natural Language Processing (NLP) necessitates centralized access to high-quality, large-scale language resources. This paper presents the technical infrastructure and evolving ecosystem of Korpuss.lv, the central access platform for the Latvian National Corpora Collection (LNCC). The LNCC consolidates 42 corpora developed by 14 institutions, comprising 2.8 billion tokens of written and spoken Latvian across diverse genres and annotation layers. Korpuss.lv has evolved from a simple metadata index into a comprehensive digital infrastructure that enhances corpus discoverability, accessibility, and usability for researchers in linguistics, digital humanities, and natural language processing. The platform integrates noSketchEngine as its primary corpus analysis tool and extends its functionality with custom modules, including a metadata-driven Corpora Explorer, a client-side Federated Content Search system, and precomputed UD-based Word Sketches. The ecosystem is further supported by CLARIN DSpace repositories for persistent storage and citation management, as well as a federated academic authentication architecture built on SATOSA and Keycloak via the CLARIN Service Provider Federation. The paper outlines architectural decisions, integration strategies, and future development plans.
The Icelandic Gigaword Corpus (IGC) is a primary resource for Icelandic NLP, with its current version containing 2.7 billion words of curated text. The IGC is traditionally distributed in a TEI-XML format, a hierarchical structure that allows for rich linguistic annotation and metadata. However, this format introduces significant friction for modern machine learning workflows. Even high-quality curated corpora have been found to contain "unwanted" text sequences – such as fragmented lists or repetitive boilerplate that may trigger instabilities during training of large language models. In this paper, we present a new processing pipeline designed to optimize the IGC for AI development. We describe a filtering approach focusing on training stability, including fuzzy deduplication to reduce the risk of data leakage, with the aim to provide high-quality data for stable model convergence. Furthermore, we introduce a new JSONL distribution format that bridges the gap between TEI-XML and machine-actionable data, facilitating easier access and safer training for models aiming to work with Icelandic.
The Hellenic National Corpus (HNC) is an integrated online environment offering access to standard Modern Greek language material and to related analysis tools. The HNC corpus has been developed in two main phases, and currently comprises over 97 million words exclusively of written language, sourced from printed resources or scraped from the internet. The material has been automatically lemmatized and morphologically annotated, while a subset of 100,000 words has been further manually corrected, in order to produce a freely downloadable error-free corpus. Through the dedicated platform, the users have access to concordances, morphological analysis of words and statistical information (frequency) at word, lemma, part of speech and n-gram levels. Future steps include the expansion of the material in both historical and coverage dimensions: the inclusion of material from older phases of the language is foreseen, as well as the addition of dialectal material besides standards language.
(That item was not included by the Authors)
We thank all four reviewers for their constructive feedback. Below we summarize the changes made in response to their recommendations. Reviewer 1: We added examples illustrating morphological disambiguation. We expanded the semantic annotation section with a detailed paragraph and an example. We added a figure (Appendix A) showing the distribution of functional styles across periods. We clarified our position on corpus balance. Reviewer 2: We added quantitative evaluation metrics for TagText. Regarding the interaction between the existing morphological annotation and the planned UD syntactic layer: we acknowledge this is an important question, but since the syntactic layer has not yet been implemented, we can not describe its interoperability in detail. We corrected the capitalization of "Ukrainian" in the bibliography and improved formatting consistency. Reviewer 3: We clarified the token format example. We added a description of the vertical file format used for NoSketch Engine. Regarding the search engine after syntactic annotation is added, this is an open question we cannot yet answer definitively. We addressed the question of mass-generated text. Reviewer 4: We added a detailed Section 2 (Pipeline and Technical Infrastructure) covering text collection, metadata handling, preprocessing with CleanText, and morphological tagging and semantic annotation with TagText.
1. (Reviewer 1) Comment: Some terminology could do with more explanation, such as MARCELL and CURLICAT, however, the general point is very clear. Response: We have added some more details on the international projects that involved the creation of large datasets included in BulNC. 2. (Reviewer 1) Comment: I am especially interested in the enrichment of the corpus with multimodal data. I think this is something that many large corpus managers would be interested in exploring. Response: We added some more details on the multimodal dataset, the organisation of the multimodal data, the ontology description, and the applications. 3. (Reviewer 2) Comment: It does not, however, give the reader a clear idea of current priorities and future directions, and it largely fails to put the work in context of other research. Response: A clarification has been made in the conclusion that we aim at extending the large dataset with more data and extensive metadata description in order to facilitate development of language technologies and fine-tuning of LLMs. 4. (Reviewer 2) Comment: - "Like many other large reference corpora" - please give reference so that readers can know in what context you see your own work Response: We expanded the Introduction with a paragraph citing other related work on large reference corpora: "There are two main approaches to providing search interfaces for large reference corpora. ..." Also, in the conclusion we included more details on large datasets used for LLMs, which provide context for our future work. 5. (Reviewer 2) Comment: - "linguistic and corpus research" - do you see these as two different (sub)disciplines? Explain or use a different wording Response: Thank you for the remark, it is well founded and we reformulated it as ‘linguistic and NLP research’. 6. (Reviewer 2) Comment:- JSONL and CSV - explain what that is and how you are using it for linguistic data – the formats themselves are just very general specifications for textual data. Are you using any particular linguistic standards? Response: More details are provided in the text with respect to the BulNC processing pipelines, and the handling of different file formats. Due to the limited volume of the paper we have not provided details on the linguistic annotation of the corpus. 7. (Reviewer 2) Comment: - "now called the IfGPT dataset" – is that relevant? What does the acronym stand for? Response: In the fourth paragraph of Section 1, we explain: "These efforts led to the development of the large BulNC-based dataset within the project IfGPT: Infrastructure for Fine-tuning Pre-trained Large Language Models (thus, also called the IfGPT dataset), with a special focus on the efficient management of large text data." 8. (Reviewer 2) Comment: - Table 1 - please explain what the different corpus components actually are Response: We clarified in the following way: "Further extensions of the dataset include newly collected and processed texts from various time periods. Older texts, such as news articles, periodicals, and books published before 1990, are also collected and processed using OCR." 9. (Reviewer 2) Comment:- "25 languages" - which ones? How where they chosen? Response: We clarified in the following way: "The selection of languages was based on the availability of wordnets in various languages in the Extended Open Multilingual Wordnet." 10. (Reviewer 2) Comment: - "therefore has a complex graph-based structure" – I don’t see why this follows from the fact that BulNC is designed to support corpus and language research. Can you explain? How does your graph-based structure relate to other approaches to metadata? Response: Justification for the use of Neo4J database is provided: "The metadata are managed using a graph database, Neo4J, that is designed to handle large volumes of interconnected data efficiently and maintains performance under complex queries using the Cypher query language." 11. (Reviewer 2) Comment: - "Corpus Query Tool specifically developed for the BulNC" – reference? URL? Response: Reference is provided both to the BulNC search interface and the IfGPT metadata web search. 12. (Reviewer 2) Comment: - References – these are exclusively self-references. Do you not want to put your work into the context of other CMLC contributions? Response: More references are provided. See 4. 13. (Reviewer 2) Comment: - General remark: You’re leaving implicit what BulNC is *not* doing. Can you devote at least one sentence to your approach to spoken language? What about CMC, learner language, etc.? Response: Currently, we have not extended our work towards including spoken language data or other specialised datasets (e.g., learner data, etc.). 14. (Reviewer 3) Comment: The difference between "BulNC", "BulNC-based dataset" and "IfGPT dataset" is never really made clear and needs to be inferred by the reader. The presentation would benefit greatly from one or two sentences early on that explicitly spell out the relationship between BulNC, the BulNC-based dataset, and the IfGPT dataset. Response: We thank the reviewers for highlighting the need to clarify the relationship between "BulNC", the "BulNC-based dataset", and the "IfGPT dataset". Answer is as 7. above.
It was not possible to address the objections of reviewer #1 in an updated version of the paper. They seem to wish for a different paper about AI. Our brief was for "submit a poster on what has changed (and what is possibly about to change)" with respect to the BNC. The BNC is a historical corpus not an ongoing project. Reviewer #2 also notes a lack of novelty, but the same applies as above. We have added a little more on ’lessons learned’. Reviewer #3 asks about the availablility on CD of the corpus. This was phased out more than ten years ago, and it is already stated in the paper that the corpus is available for download from the Oxford Text Archive. Likewise, the TEI annotation of the corpus was frozen on release. This point is emphasised in a revision. The importance of the corpus as a TEI flagship project is noted. The typos reported by reviewer #3 have been corrected. Further footnotes with URLs, references to publications and language resources have been added.
Several questions and suggestions from the Reviewers concerned providing more details about the corpus design, composition and annotation. We submitted our poster’s accompanying paper as an extended abstract since a paper presenting the corpus was accepted at the main LREC 2026 conference. In the below response, we omit the Reviewers’ remarks that called for more detailed explanations or examples that are provided in the full paper. Once anonymity is not an issue at publication time, we added a reference to the paper in our CMLC abstract. Below we address the remaining points raised by the Reviewers. REVIEWER 1 Q: KWJP limits the diversity of texts to three main categories: fiction, non-fiction, and journalistic writing. This raises the question of whether such a restriction affects comparability with NKJP, which included a wider range of text types, including spoken and internet communication, and in particular those that are less represented in specialized corpora, such as correspondence and poetry. A: Any difference in data selection restricts the comparability of corpora. Spoken language and internet communication registers call for significantly different approaches to data collection, selection and annotation processes. As for smaller and less represented text types: to our knowledge both poetry and correspondence in the balanced NKJP corpus are represented by three books each. REVIEWER 3 Q: For the planned 2026 update, will the new sub-corpus maintain the same annotation layers or are there planned expansions (e.g. more fine-grained NE types)? A: We are not planning any new or modified layers of annotation at this point. The planned update refers to textual data as well as technical efficiency of the web service. Q: Is there consideration for releasing larger open subsets or integrating the data into platforms like Sketch Engine for broader academic access? A: We clarified the statement about copyright issues prohibiting us from releasing larger amounts of data. The same legal considerations do not allow us to upload the data to external services such as Sketch Engine, effectively distributing them to third parties. REVIEWER 4 Q: What types of named entity have been annotated in the corpus? A: As it was mentioned in the abstract, we are using a NE annotation schema developed in the NKJP project. The details may be found in relevant publications concerning this corpus. Q: Is "the development of web-based applications" really necessary? Are you not able to use existing solutions without reinventing the wheel? A: We did not state that we are developing a web-based application, rather we are developing a new solution (within our web service) to provide a more efficient searching mechanism. One of the paths explored in the process is using the BlackLab search engine. It is possible that we will index the corpus in two different search engines and allow the user to choose, but no final decisions have been made yet.
This field was left empty by the Authors. Filling it for the purpose of manual submission of the PDF by the editor.
The third generation of the Hungarian National Corpus (MNSZ3) aims to provide a large-scale, curated, and well-described corpus resource needed for the sustainable digital presence of Hungarian. Building on the domain structure and proportions of MNSZ2 (v2.0.5; 1.04 billion running words), the project targets a substantial increase in scale while also strengthening the coverage and metadata description of Hungarian language use outside Hungary. MNSZ3 retains the six traditional domains of the earlier corpus—press, fiction, scientific, official, personal, and transcribed spoken language—and is planned to reach approximately 10 billion tokens. This paper presents the motivation and design principles of the project, outlines the practical decisions and procedures used in data collection and cleaning, and discusses the annotation strategy developed for large-scale processing. In planning the linguistic analysis, we build on the complementary strengths of HuSpaCy and e-magyar: HuSpaCy provides the unified and efficient UD-oriented processing backbone, while e-magyar (emMorph) is preserved as an explicit additional layer for morphology and lemmatisation.
We are very grateful to the reviewers for their comments, most of which we have implemented. We were not able to fully address the following points: 1. "Preliminary comparison metrics (e.g. parsing accuracy gain) between the old TTL and the new RODNA syntactic annotations on overlapping data": TTL did not provide syntactic annotations. Preliminary experiments with RODNA are described in the paper, where the tool is compared to STANZA and Trankit. 2. "For ADAMo, what specific features distinguish Moldovan from standard Romanian in the added texts, and how will variety differences be quantified": these experiments are still at a preliminary stage. 3. "Providing example queries/snippets illustrating syntactic annotations": since this annotation layer is not yet indexed in CoRoLa, we cannot provide examples from CoRoLa. However, the syntactic annotation available for German in KorAP (korap.ids-mannheim.de) offers an indication of how such annotations could be queried and displayed in the future.
Clinical text resources are a central component for the study of medical language, as well as the training and evaluation of large language models, chatbots, and artificial intelligence systems supporting clinical routines. With the German Medical Text Corpus (GeMTeX), we are currently working on the largest shareable clinical document dataset in German. The multi-centric project ensures diversity across different university hospitals, clinical domains, and text sorts. After a thorough de-identification process, the clinical texts are semantically annotated using Snomed CT, a language-independent, standardized medical ontology. While the corpus is still under active development, it is accessible upon request under controlled access conditions. As of February 2026, GeMTeX comprises more than 15k documents and 20M tokens. We refer researchers interested in the resource to visit https://kiinformatik.mri.tum.de/en/gemtex or reach out to us via gemtex.mi@mh.tum.de.
Launched in 2020, CorCenCC (Corpws Cenedlaethol Cymraeg Cyfoes – National Corpus of Contemporary Welsh) is the first large-scale corpus of the Welsh language to integrate spoken, written, and electronically mediated data, offering a comprehensive snapshot of contemporary Welsh use. Including contributions from over 2,000 speakers, the 11.2-million-word corpus represents the diversity of Wales’s linguistic landscape. As a national resource, CorCenCC enables users to explore real world Welsh. Several tools and resources were developed through the CorCenCC project, including the CyTag POS tagger and CySemTag (adapted from Lancaster University’s USAS semantic system), to enable the grammatical and semantic categorisation of the dataset. The team also built the pedagogic toolkit Y Tiwtiadur, to allow learners and teachers to access corpus-based examples and tasks. Additionally, Yr Amliadur provides curated frequency-based wordlists across modes and parts of speech, supporting linguistic analysis and vocabulary development. Since completing the corpus, the team has focused on extending its impact and reach, to ensure that the resources are maintained and sustained for future use; a challenge often faced when large-scale projects end. This poster profiles the tools and resources created from and inspired by CorCenCC and its associated tools and resources, as a means of supporting the democratisation of linguistic resources for minoritised language contexts.
This paper introduces Swiss-AL, a language data platform designed for the multilingual, comparative analysis of public discourse in Switzerland. Swiss-AL is an open research data resource providing browser-based access to a variety of corpora in all four of Switzerland’s official languages. Corpora contain journalistic, organisational, and parliamentary discourse. The platform supports research in applied linguistics as well as neighbouring disciplines (e.g., social sciences, communication and media studies).
This paper reports on recent technical developments in the European Reference Corpus EuReCo and its current technical implementation based on the corpus search and analysis platform KorAP. We describe updates to the ingestion pipeline, including extensions to the TEI-to-KorAP-XML converter tei2korapxml and the KorAP tokenizer, as well as the newly introduced korapxmltool for annotation and index conversion. We further present Koral-Mapper, a service that enables cross-schema comparability of annotations and metadata at query time, and report on developments in the backend access control system Kustvakt, the web user interface Kalamar, API client libraries for R and Python that promote reproducibility and methodologically sound AI-assisted analysis, and containerized deployment. The corpora and languages currently represented in EuReCo are outlined, and the role of the German Reference Corpus DeReKo, including its metadata-driven virtual corpus design, predefined useful subcorpora, and TEI encoding, is discussed in detail. We further present the National Libraries as Corpus approach and DeLiKo-2025@DNB as its first full-scale proof of concept, and discuss the potential of this approach for extending EuReCo with comparable contemporary fiction corpora across European countries.