@inproceedings{koeva-stoyanova-2026-recent,
title = "Recent Developments of the {B}ulgarian National Corpus",
author = "Koeva, Svetla Peneva and
Stoyanova, Ivelina",
editor = "Ba{\'n}ski, Piotr and
Knight, Dawn and
Kupietz, Marc and
Witt, Andreas and
Wr{\'o}blewska, Alina",
booktitle = "Proceedings of the 12th Workshop on Challenges in the Management of Large Corpora",
month = may,
year = "2026",
address = "Palma, Mallorca (Spain)",
publisher = "ELRA Language Resources Association (ELRA)",
url = "https://preview.aclanthology.org/ingest-lrec/2026.cmlc-1.10/",
doi = "10.63317/3m95ohtw7mjs",
pages = "71--75",
abstract = "1. (Reviewer 1) Comment: Some terminology could do with more explanation, such as MARCELL and CURLICAT, however, the general point is very clear. Response: We have added some more details on the international projects that involved the creation of large datasets included in BulNC. 2. (Reviewer 1) Comment: I am especially interested in the enrichment of the corpus with multimodal data. I think this is something that many large corpus managers would be interested in exploring. Response: We added some more details on the multimodal dataset, the organisation of the multimodal data, the ontology description, and the applications. 3. (Reviewer 2) Comment: It does not, however, give the reader a clear idea of current priorities and future directions, and it largely fails to put the work in context of other research. Response: A clarification has been made in the conclusion that we aim at extending the large dataset with more data and extensive metadata description in order to facilitate development of language technologies and fine-tuning of LLMs. 4. (Reviewer 2) Comment: - ``Like many other large reference corpora'' - please give reference so that readers can know in what context you see your own work Response: We expanded the Introduction with a paragraph citing other related work on large reference corpora: ``There are two main approaches to providing search interfaces for large reference corpora. ...'' Also, in the conclusion we included more details on large datasets used for LLMs, which provide context for our future work. 5. (Reviewer 2) Comment: - ``linguistic and corpus research'' - do you see these as two different (sub)disciplines? Explain or use a different wording Response: Thank you for the remark, it is well founded and we reformulated it as `linguistic and NLP research'. 6. (Reviewer 2) Comment:- JSONL and CSV - explain what that is and how you are using it for linguistic data {--} the formats themselves are just very general specifications for textual data. Are you using any particular linguistic standards? Response: More details are provided in the text with respect to the BulNC processing pipelines, and the handling of different file formats. Due to the limited volume of the paper we have not provided details on the linguistic annotation of the corpus. 7. (Reviewer 2) Comment: - ``now called the IfGPT dataset'' {--} is that relevant? What does the acronym stand for? Response: In the fourth paragraph of Section 1, we explain: ``These efforts led to the development of the large BulNC-based dataset within the project \textit{IfGPT: Infrastructure for Fine-tuning Pre-trained Large Language Models} (thus, also called the \textbf{IfGPT dataset}), with a special focus on the efficient management of large text data.'' 8. (Reviewer 2) Comment: - Table 1 - please explain what the different corpus components actually are Response: We clarified in the following way: ``Further extensions of the dataset include newly collected and processed texts from various time periods. Older texts, such as news articles, periodicals, and books published before 1990, are also collected and processed using OCR.'' 9. (Reviewer 2) Comment:- ``25 languages'' - which ones? How where they chosen? Response: We clarified in the following way: ``The selection of languages was based on the availability of wordnets in various languages in the Extended Open Multilingual Wordnet.'' 10. (Reviewer 2) Comment: - ``therefore has a complex graph-based structure'' {--} I don{'}t see why this follows from the fact that BulNC is designed to support corpus and language research. Can you explain? How does your graph-based structure relate to other approaches to metadata? Response: Justification for the use of Neo4J database is provided: ``The metadata are managed using a graph database, Neo4J, that is designed to handle large volumes of interconnected data efficiently and maintains performance under complex queries using the Cypher query language.'' 11. (Reviewer 2) Comment: - ``Corpus Query Tool specifically developed for the BulNC'' {--} reference? URL? Response: Reference is provided both to the BulNC search interface and the IfGPT metadata web search. 12. (Reviewer 2) Comment: - References {--} these are exclusively self-references. Do you not want to put your work into the context of other CMLC contributions? Response: More references are provided. See 4. 13. (Reviewer 2) Comment: - General remark: You{'}re leaving implicit what BulNC is *not* doing. Can you devote at least one sentence to your approach to spoken language? What about CMC, learner language, etc.? Response: Currently, we have not extended our work towards including spoken language data or other specialised datasets (e.g., learner data, etc.). 14. (Reviewer 3) Comment: The difference between ``BulNC'', ``BulNC-based dataset'' and ``IfGPT dataset'' is never really made clear and needs to be inferred by the reader. The presentation would benefit greatly from one or two sentences early on that explicitly spell out the relationship between BulNC, the BulNC-based dataset, and the IfGPT dataset. Response: We thank the reviewers for highlighting the need to clarify the relationship between ``BulNC'', the ``BulNC-based dataset'', and the ``IfGPT dataset''. Answer is as 7. above."
}