SetGNER: General Named Entity Recognition as Entity Set Generation

Yuxin He, Buzhou Tang


Abstract
Recently, joint recognition of flat, nested and discontinuous entities has received increasing attention. Motivated by the observation that the target output of NER is essentially a set of sequences, we propose a novel entity set generation framework for general NER scenes in this paper. Different from sequence-to-sequence NER methods, our method does not force the entities to be generated in a predefined order and can get rid of the problem of error propagation and inefficient decoding. Distinguished from the set-prediction NER framework, our method treats each entity as a sequence and is capable of recognizing discontinuous mentions. Given an input sentence, the model first encodes the sentence in word-level and detects potential entity mentions based on the encoder’s output, then reconstructs entity mentions from the detected entity heads in parallel. To let the encoder of our model capture better right-to-left semantic structure, we also propose an auxiliary Inverse Generation Training task. Extensive experiments show that our model (w/o. Inverse Generation Training) outperforms state-of-the-art generative NER models by a large margin on two discontinuous NER datasets, two nested NER datasets and one flat NER dataset. Besides, the auxiliary Inverse Generation Training task is found to further improve the model’s performance on the five datasets.
Anthology ID:
2022.emnlp-main.200
Volume:
Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing
Month:
December
Year:
2022
Address:
Abu Dhabi, United Arab Emirates
Editors:
Yoav Goldberg, Zornitsa Kozareva, Yue Zhang
Venue:
EMNLP
SIG:
Publisher:
Association for Computational Linguistics
Note:
Pages:
3074–3085
Language:
URL:
https://aclanthology.org/2022.emnlp-main.200
DOI:
10.18653/v1/2022.emnlp-main.200
Bibkey:
Cite (ACL):
Yuxin He and Buzhou Tang. 2022. SetGNER: General Named Entity Recognition as Entity Set Generation. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 3074–3085, Abu Dhabi, United Arab Emirates. Association for Computational Linguistics.
Cite (Informal):
SetGNER: General Named Entity Recognition as Entity Set Generation (He & Tang, EMNLP 2022)
Copy Citation:
PDF:
https://preview.aclanthology.org/nschneid-patch-2/2022.emnlp-main.200.pdf