Hanna Fischer
2026
Onomasiological Sense Alignment Across Dialect Dictionaries. A Taxonomy-Constrained LLM Classification
Nathalie Mederake | Nico Urbach | Hanna Fischer | Alfred Lameli
Proceedings of the 13th Workshop on NLP for Similar Languages, Varieties and Dialects
Nathalie Mederake | Nico Urbach | Hanna Fischer | Alfred Lameli
Proceedings of the 13th Workshop on NLP for Similar Languages, Varieties and Dialects
We propose a taxonomy-guided approach to semantic alignment that assigns lexicographic senses to an onomasiological taxonomy derived from the Hallig–Wartburg/Post system. Using an LLM under strict taxonomic constraints, short and heterogeneous meaning descriptions are assigned to a common conceptual space. Evaluation against expert annotation shows that run-to-run model agreement (kappa = 0.73) closely matches human agreement (kappa = 0.74), with robustness at coarse taxonomic levels and predictable degradation at finer granularity. A qualitative network analysis demonstrates the resulting potential for cross-dictionary exploration of dialectal variation in semantics.
German Dialects Across Situations, Generations, and Regions: The REDE corpus as an Oral Resource for NLP
Hanna Fischer | Alfred Lameli
Proceedings of the First Workshop on Dialects in NLP — A Resource Perspective
Hanna Fischer | Alfred Lameli
Proceedings of the First Workshop on Dialects in NLP — A Resource Perspective
Recent advances in speech and language technologies increasingly rely on large and diverse corpora that represent linguistic variation across dialect regions, communicative situations, and social speaker characteristics. While substantial resources are available for Standard German, comparable spoken corpora for German dialects have so far been largely lacking, limiting the development and evaluation of dialect-sensitive NLP systems. The REDE corpus addresses this gap by providing a methodologically uniform collection of spoken German for 148 locations that systematically covers all major dialect areas in Germany. It comprises contemporary recordings collected in multiple elicitation and interaction settings, capturing variation across speaking styles, situational contexts, and speaker generations. With more than 1,500 hours of speech and rich metadata on regional and social dimensions, the REDE corpus constitutes a large-scale oral resource suitable for both linguistic research and NLP applications. This paper presents the design, structure, and methodological foundations of the corpus and discusses its relevance for current speech technology requirements.
2023
Reconstructing Language History by Using a Phonological Ontology. An Analysis of German Surnames
Hanna Fischer | Robert Engsterhold
Tenth Workshop on NLP for Similar Languages, Varieties and Dialects (VarDial 2023)
Hanna Fischer | Robert Engsterhold
Tenth Workshop on NLP for Similar Languages, Varieties and Dialects (VarDial 2023)
This paper applies the ontology-baseddialectometric technique of Engsterhold(2020) to surnames. The method wasoriginally developed for phonetic analyses. However, as will be shown, it is also suitedfor the study of graphemic representations. Based on data from the German SurnameAtlas (DFA), the method is optimized forgraphemic analysis and illustrated with anexample case.