Dynamic Ensembles in Named Entity Recognition for Historical Arabic Texts

Muhammad Majadly, Tomer Sagi


Abstract
The use of Named Entity Recognition (NER) over archaic Arabic texts is steadily increasing. However, most tools have been either developed for modern English or trained over English language documents and are limited over historical Arabic text. Even Arabic NER tools are often trained on modern web-sourced text, making their fit for a historical task questionable. To mitigate historic Arabic NER resource scarcity, we propose a dynamic ensemble model utilizing several learners. The dynamic aspect is achieved by utilizing predictors and features over NER algorithm results that identify which have performed better on a specific task in real-time. We evaluate our approach against state-of-the-art Arabic NER and static ensemble methods over a novel historical Arabic NER task we have created. Our results show that our approach improves upon the state-of-the-art and reaches a 0.8 F-score on this challenging task.
Anthology ID:
2021.wanlp-1.12
Volume:
Proceedings of the Sixth Arabic Natural Language Processing Workshop
Month:
April
Year:
2021
Address:
Kyiv, Ukraine (Virtual)
Editors:
Nizar Habash, Houda Bouamor, Hazem Hajj, Walid Magdy, Wajdi Zaghouani, Fethi Bougares, Nadi Tomeh, Ibrahim Abu Farha, Samia Touileb
Venue:
WANLP
SIG:
Publisher:
Association for Computational Linguistics
Note:
Pages:
115–125
Language:
URL:
https://aclanthology.org/2021.wanlp-1.12
DOI:
Bibkey:
Cite (ACL):
Muhammad Majadly and Tomer Sagi. 2021. Dynamic Ensembles in Named Entity Recognition for Historical Arabic Texts. In Proceedings of the Sixth Arabic Natural Language Processing Workshop, pages 115–125, Kyiv, Ukraine (Virtual). Association for Computational Linguistics.
Cite (Informal):
Dynamic Ensembles in Named Entity Recognition for Historical Arabic Texts (Majadly & Sagi, WANLP 2021)
Copy Citation:
PDF:
https://preview.aclanthology.org/naacl24-info/2021.wanlp-1.12.pdf