Abstract
Mapping user locations to countries can be useful for many applications such as dialect identification, author profiling, recommendation system, etc. Twitter allows users to declare their locations as free text, and these user-declared locations are often noisy and hard to decipher automatically. In this paper, we present the largest manually labeled dataset for mapping user locations on Arabic Twitter to their corresponding countries. We build effective machine learning models that can automate this mapping with significantly better efficiency compared to libraries such as geopy. We also show that our dataset is more effective than data extracted from GeoNames geographical database in this task as the latter covers only locations written in formal ways.- Anthology ID:
- 2021.wanlp-1.15
- Volume:
- Proceedings of the Sixth Arabic Natural Language Processing Workshop
- Month:
- April
- Year:
- 2021
- Address:
- Kyiv, Ukraine (Virtual)
- Venue:
- WANLP
- SIG:
- Publisher:
- Association for Computational Linguistics
- Note:
- Pages:
- 145–153
- Language:
- URL:
- https://aclanthology.org/2021.wanlp-1.15
- DOI:
- Cite (ACL):
- Hamdy Mubarak and Sabit Hassan. 2021. UL2C: Mapping User Locations to Countries on Arabic Twitter. In Proceedings of the Sixth Arabic Natural Language Processing Workshop, pages 145–153, Kyiv, Ukraine (Virtual). Association for Computational Linguistics.
- Cite (Informal):
- UL2C: Mapping User Locations to Countries on Arabic Twitter (Mubarak & Hassan, WANLP 2021)
- PDF:
- https://preview.aclanthology.org/remove-xml-comments/2021.wanlp-1.15.pdf