Johan Heinsen
2026
A Growing Literature of the Public Sphere: Fiction in Danish Newspapers (1666–1850)
Pascale Feldkamp | Alie Lassche | Rie Eriksen | Kit Morgenstjerne | Kristoffer Nielbo | Johan Heinsen | Yuri Bizzoni
Proceedings of the First Workshop on Creating Interoperable Corpora of Historical Newspapers
Pascale Feldkamp | Alie Lassche | Rie Eriksen | Kit Morgenstjerne | Kristoffer Nielbo | Johan Heinsen | Yuri Bizzoni
Proceedings of the First Workshop on Creating Interoperable Corpora of Historical Newspapers
Digitized literary corpora of the 19th century largely focus on standalone volumes, sidelining the broader and more diverse literary production of the period. Fiction published in less enduring formats – such as novellas and serialized pieces in newspapers – remains underexplored, particularly for low-resource languages like Danish, despite the growing availability of digitized newspaper archives. This paper addresses that gap by identifying and tagging fiction in Danish newspapers (1666–1850). We (1) present a manually annotated dataset of 1,831 articles with both binary (fiction/nonfiction) and fine-grained subcategories (travelogue, biography, essay), and (2) evaluate a document-embedding classifier that achieves an F1-score of up to 0.89 for the fiction/nonfiction distinction. Building on this pipeline, we further provide two resources for future research: (a) fiction probability scores for nearly five million newspaper articles (n=4,898,084), and (b) a small, cleaned, and curated subset of newspaper fiction (n=139), intended as a growing resource.
Dynaword: From One-shot to Continuously Developed Datasets
Kenneth Enevoldsen | Kristian Nørgaard Jensen | Jan Kostkan | Balázs Szabó | Márton Kardos | Kirsten Vad | Johan Heinsen | Andrea Blasi Núñez | Gianluca Barmina | Jacob Nielsen | Rasmus Larsen | Rob van der Goot | Peter Vahlstrup | Per Møldrup Dalum | Desmond Elliott | Lukas Galke Poech | Peter Schneider-Kamp | Kristoffer Nielbo
Proceedings of the Fifteenth Language Resources and Evaluation Conference
Kenneth Enevoldsen | Kristian Nørgaard Jensen | Jan Kostkan | Balázs Szabó | Márton Kardos | Kirsten Vad | Johan Heinsen | Andrea Blasi Núñez | Gianluca Barmina | Jacob Nielsen | Rasmus Larsen | Rob van der Goot | Peter Vahlstrup | Per Møldrup Dalum | Desmond Elliott | Lukas Galke Poech | Peter Schneider-Kamp | Kristoffer Nielbo
Proceedings of the Fifteenth Language Resources and Evaluation Conference
Large-scale datasets are foundational for research and development in natural language processing. However, current approaches face three key challenges: (1) reliance on ambiguously licensed sources restricting use, sharing, and derivative works; (2) static dataset releases that prevent community contributions and diminish longevity; and (3) quality assurance processes restricted to publishing teams rather than leveraging community expertise. To address these limitations, we introduce two contributions: the Dynaword approach and Danish Dynaword. The Dynaword approach is a framework for creating large-scale, open datasets that can be continuously updated through community collaboration. Danish Dynaword is a concrete implementation that validates this approach and demonstrates its potential. Danish Dynaword contains over five times as many tokens as comparable releases, is exclusively openly licensed, and has received multiple contributions across industry, the public sector and research institutions. The repository includes light-weight tests to ensure data formatting, quality, and documentation, establishing a sustainable framework for ongoing community contributions and dataset evolution.
Evaluating Embedding Models on Danish Historical Newspapers: A Corpus and Benchmark Resource
Alie Lassche | Pascale Feldkamp | Yuri Bizzoni | Katrine Baunvig | Kristoffer Nielbo | Johan Heinsen
Proceedings of the Fifteenth Language Resources and Evaluation Conference
Alie Lassche | Pascale Feldkamp | Yuri Bizzoni | Katrine Baunvig | Kristoffer Nielbo | Johan Heinsen
Proceedings of the Fifteenth Language Resources and Evaluation Conference
We present an enriched dataset of almost five million Danish historical newspaper articles from the late seventeenth to nineteenth century, augmented with semantic embeddings and an annotated subset, to enable semi-automated classification as well as thematic and linguistic exploration. Through three historical benchmark tasks that evaluate the performance of Danish and multilingual embedding models on this historical Danish corpus, we discuss how the choice for an embedding model depends on the type of task, and enrich our corpus with embeddings from the overall best performing model. As a showcase experiment, we look at the distribution of article categories in the three subgenres that can be observed in the corpus. This experiment highlights the corpus and article-level embeddings’ potential for further exploration and analysis of the Danish historical mediascape. The resource is freely available for research use and aims to foster reproducible, data-driven studies of language and culture in the Danish nineteenth century.
Search
Fix author
Co-authors
- Kristoffer Nielbo 3
- Yuri Bizzoni 2
- Pascale Feldkamp 2
- Alie Lassche 2
- Gianluca Barmina 1
- Katrine Baunvig 1
- Andrea Blasi Núñez 1
- Per Møldrup Dalum 1
- Desmond Elliott 1
- Kenneth Enevoldsen 1
- Rie Eriksen 1
- Lukas Galke Poech 1
- Kristian Nørgaard Jensen 1
- Márton Kardos 1
- Jan Kostkan 1
- Rasmus Larsen 1
- Kit Morgenstjerne 1
- Jacob Nielsen 1
- Peter Schneider-Kamp 1
- Balázs Szabó 1
- Kirsten Vad 1
- Peter Vahlstrup 1
- Rob van der Goot 1