An Extensible Multilingual Open Source Lemmatizer

Ahmet Aker, Johann Petrak, Firas Sabbah


Abstract
We present GATE DictLemmatizer, a multilingual open source lemmatizer for the GATE NLP framework that currently supports English, German, Italian, French, Dutch, and Spanish, and is easily extensible to other languages. The software is freely available under the LGPL license. The lemmatization is based on the Helsinki Finite-State Transducer Technology (HFST) and lemma dictionaries automatically created from Wiktionary. We evaluate the performance of the lemmatizers against TreeTagger, which is only freely available for research purposes. Our evaluation shows that DictLemmatizer achieves similar or even better results than TreeTagger for languages where there is support from HFST. The performance drops when there is no support from HFST and the entire lemmatization process is based on lemma dictionaries. However, the results are still satisfactory given the fact that DictLemmatizer isopen-source and can be easily extended to other languages. The software for extending the lemmatizer by creating word lists from Wiktionary dictionaries is also freely available as open-source software.
Anthology ID:
R17-1006
Volume:
Proceedings of the International Conference Recent Advances in Natural Language Processing, RANLP 2017
Month:
September
Year:
2017
Address:
Varna, Bulgaria
Venue:
RANLP
SIG:
Publisher:
INCOMA Ltd.
Note:
Pages:
40–45
Language:
URL:
https://doi.org/10.26615/978-954-452-049-6_006
DOI:
10.26615/978-954-452-049-6_006
Bibkey:
Cite (ACL):
Ahmet Aker, Johann Petrak, and Firas Sabbah. 2017. An Extensible Multilingual Open Source Lemmatizer. In Proceedings of the International Conference Recent Advances in Natural Language Processing, RANLP 2017, pages 40–45, Varna, Bulgaria. INCOMA Ltd..
Cite (Informal):
An Extensible Multilingual Open Source Lemmatizer (Aker et al., RANLP 2017)
Copy Citation:
PDF:
https://doi.org/10.26615/978-954-452-049-6_006