Abstract
The use of Named Entity Recognition (NER) over archaic Arabic texts is steadily increasing. However, most tools have been either developed for modern English or trained over English language documents and are limited over historical Arabic text. Even Arabic NER tools are often trained on modern web-sourced text, making their fit for a historical task questionable. To mitigate historic Arabic NER resource scarcity, we propose a dynamic ensemble model utilizing several learners. The dynamic aspect is achieved by utilizing predictors and features over NER algorithm results that identify which have performed better on a specific task in real-time. We evaluate our approach against state-of-the-art Arabic NER and static ensemble methods over a novel historical Arabic NER task we have created. Our results show that our approach improves upon the state-of-the-art and reaches a 0.8 F-score on this challenging task.- Anthology ID:
- 2021.wanlp-1.12
- Volume:
- Proceedings of the Sixth Arabic Natural Language Processing Workshop
- Month:
- April
- Year:
- 2021
- Address:
- Kyiv, Ukraine (Virtual)
- Editors:
- Nizar Habash, Houda Bouamor, Hazem Hajj, Walid Magdy, Wajdi Zaghouani, Fethi Bougares, Nadi Tomeh, Ibrahim Abu Farha, Samia Touileb
- Venue:
- WANLP
- SIG:
- Publisher:
- Association for Computational Linguistics
- Note:
- Pages:
- 115–125
- Language:
- URL:
- https://aclanthology.org/2021.wanlp-1.12
- DOI:
- Cite (ACL):
- Muhammad Majadly and Tomer Sagi. 2021. Dynamic Ensembles in Named Entity Recognition for Historical Arabic Texts. In Proceedings of the Sixth Arabic Natural Language Processing Workshop, pages 115–125, Kyiv, Ukraine (Virtual). Association for Computational Linguistics.
- Cite (Informal):
- Dynamic Ensembles in Named Entity Recognition for Historical Arabic Texts (Majadly & Sagi, WANLP 2021)
- PDF:
- https://preview.aclanthology.org/naacl24-info/2021.wanlp-1.12.pdf