Nicolai Hartvig Sørensen

Also published as: Nicolai Hartvig Sørensen

2020

pdf bib abs
An automatically generated Danish Renaissance Dictionary
Mette-Marie Møller Svendsen | Nicolai Hartvig Sørensen | Thomas Troelsgård
Proceedings of the 2020 Globalex Workshop on Linked Lexicography

We present the ongoing work on an automatically generated dictionary describing Danish in the 16th century. A series of relevant dictionaries – from the period as well as more recent ones – are linked together at lemma level, and where possible, definitions or keywords are extracted and presented in the new dictionary.

2016

We launch the SemDaX corpus which is a recently completed Danish human-annotated corpus available through a CLARIN academic license. The corpus includes approx. 90,000 words, comprises six textual domains, and is annotated with sense inventories of different granularity. The aim of the developed corpus is twofold: i) to assess the reliability of the different sense annotation schemes for Danish measured by qualitative analyses and annotation agreement scores, and ii) to serve as training and test data for machine learning algorithms with the practical purpose of developing sense taggers for Danish. To these aims, we take a new approach to human-annotated corpus resources by double annotating a much larger part of the corpus than what is normally seen: for the all-words task we double annotated 60% of the material and for the lexical sample task 100%. We include in the corpus not only the adjucated files, but also the diverging annotations. In other words, we consider not all disagreement to be noise, but rather to contain valuable linguistic information that can help us improve our annotation schemes and our learning algorithms.

2015

Co-authors

Bolette Sandford Pedersen 2

Anders Søgaard 2

Mette-Marie Møller Svendsen 1

Thomas Troelsgård 1

Venues

Fix data