Domain-Specific Considerations in the Preparation of Specialized Corpora: A Case Study on a Corpus of German Sermons

Cora Haiber, Adam Roussel, Stefanie Dipper


Abstract
We present a new corpus of contemporary German sermons and describe the steps taken in its preparation. We apply a semi-automatic approach to sentence segmentation, tokenization, and lemmatization, utilizing annotation guidelines that are specialized to this domain. In the process of preparing these data, we find that state-of-the-art tools for these tasks still make problematic errors, especially with non-standard data, despite apparently very high performance on common benchmarks. We obtain test scores of F1 = 96.69 % for sentence segmentation, F1 = 99.99 % for tokenization, and acc = 64.00 % for lemmatization with our domain-adapted models and show that domain-adaptation improves performance over state-of-the-art models for the token and sentence segmentation tasks.
Anthology ID:
2026.slide-1.19
Volume:
Proceedings of the Workshop on Structured Linguistic Data and Evaluation (SLiDE)
Month:
May
Year:
2026
Address:
Palma de Mallorca, Spain
Editors:
Erhard Hinrichs, Joakim Nivre, Petya Osenova, James Pustejovsky, Claus Zinn
Venues:
SLiDE | WS
SIG:
Publisher:
ELRA Language Resources Association (ELRA)
Note:
Pages:
212–223
Language:
External URL:
https://lrec.elra.info/lrec2026-ws-slide-19
DOI:
10.63317/5nhkpqtxbr6i
Bibkey:
Cite (ACL):
Cora Haiber, Adam Roussel, and Stefanie Dipper. 2026. Domain-Specific Considerations in the Preparation of Specialized Corpora: A Case Study on a Corpus of German Sermons. In Proceedings of the Workshop on Structured Linguistic Data and Evaluation (SLiDE), pages 212–223, Palma de Mallorca, Spain. ELRA Language Resources Association (ELRA).
Cite (Informal):
Domain-Specific Considerations in the Preparation of Specialized Corpora: A Case Study on a Corpus of German Sermons (Haiber et al., SLiDE 2026)
Copy Citation: