Anastasiia Sirotina


2019

pdf
Named Entity Recognition in Information Security Domain for Russian
Anastasiia Sirotina | Natalia Loukachevitch
Proceedings of the International Conference on Recent Advances in Natural Language Processing (RANLP 2019)

In this paper we discuss the named entity recognition task for Russian texts related to cybersecurity. First of all, we describe the problems that arise in course of labeling unstructured texts from information security domain. We introduce guidelines for human annotators, according to which a corpus has been marked up. Then, a CRF-based system and different neural architectures have been implemented and applied to the corpus. The named entity recognition systems have been evaluated and compared to determine the most efficient one.