Mahmoud Fawzi
2026
From Cairo to Cape Town: How African Twitter Shapes the Global Palestine-Israel Narrative
Mahmoud Fawzi | Houda Elmimouni | Walid Magdy
Proceedings of the 3rd Workshop on Natural Language Processing for Political Sciences (PoliticalNLP 2026)
Mahmoud Fawzi | Houda Elmimouni | Walid Magdy
Proceedings of the 3rd Workshop on Natural Language Processing for Political Sciences (PoliticalNLP 2026)
African Twitter users are active shapers of the Palestine-Israel conversation but their contribution remains relatively understudied. Using 132.5K geo-located tweets from 2020 to 2023 and 451-term list of keywords in 33 languages, we identify three patterns in this context: (1) broad participation (Egypt supplies 43% of posts, yet Nigeria, South Africa, Kenya and Ghana contribute more than a third); (2) multilingual predominantly pro-Palestine amplification across Arabic, English, French, Swahili and other tongues, with 8% of tweets left “undetermined” by Twitter’s language detector; and (3) a humanitarian framing that centers civilian harm through hashtags such as #GazaUnderAttack and #PalestenianLivesMatter. We outline design implications for language-agnostic interfaces, low-friction source verification and cross-movement recommendation tools that foreground African epistemologies in global civic-tech systems.
EGCSS at StanceNakba Shared Task: Cross-Topic Arabic Stance Detection for Two Middle East Issues
Asmaa Qindeel | Toka Khaled | Batool Najeh Balah | Eman Elrefai | Mahmoud Fawzi
Proceedings of the 2nd International Workshop on Nakba Narratives as Language Resources @ LREC 2026
Asmaa Qindeel | Toka Khaled | Batool Najeh Balah | Eman Elrefai | Mahmoud Fawzi
Proceedings of the 2nd International Workshop on Nakba Narratives as Language Resources @ LREC 2026
Stance detection continues to be an important task sitting at the intersection of Natural Language Processing (NLP) and Computational Social Science (CSS). In this work, we evaluate how different variations of BERT models perform on the cross-topic form of the task. In particular, we inspect their performance on the second subtask of the shared task StanceNakba 2026, where two topics are included, namely Arab Normalization with Israel and The Presence of Refugees in Arab Countries. We find that the best-performing model was bert-base-arabertv02-twitter, and we further improve its performance by providing context about the topic during the training phase, achieving an F1-score of 0.86 and ranking second among the participating teams.
The NakbaEcho Dataset: From Oral Testimonies to a Transcribed Arabic History Corpus
Batool Najeh Balah | Mahmoud Fawzi | Houda Elmimouni | Walid Magdy
Proceedings of the 2nd International Workshop on Nakba Narratives as Language Resources @ LREC 2026
Batool Najeh Balah | Mahmoud Fawzi | Houda Elmimouni | Walid Magdy
Proceedings of the 2nd International Workshop on Nakba Narratives as Language Resources @ LREC 2026
We present NakbaEcho, a dataset derived from Palestinian testimonies about the 1948 Nakba. The resource is constructed from transcribing over 2,180 hours of recorded interviews gathered through the Palestine Remembered Oral History index and linked to multiple repositories, including the Palestinian Oral History Archive (POHA) and YouTube-hosted interviews. We harmonize interview-level metadata and generate timestamp-aligned transcripts from the original Arabic recordings using an automatic transcription pipeline configured for Palestinian Arabic. The dataset includes speaker-labeled segments and auxiliary annotations designed to support downstream research in Arabic speech processing, natural language processing, digital humanities, and oral-history analysis. NakbaEcho contributes a structured computational resource for studying Palestinian oral testimony while expanding the availability of dialectal Arabic materials for speech, text, and social research.
2025
IslamicEval 2025: The First Shared Task of Capturing LLMs Hallucination in Islamic Content
Hamdy Mubarak | Rana Malhas | Watheq Mansour | Abubakr Mohamed | Mahmoud Fawzi | Majd Hawasly | Tamer Elsayed | Kareem Mohamed Darwish | Walid Magdy
Proceedings of The Third Arabic Natural Language Processing Conference: Shared Tasks
Hamdy Mubarak | Rana Malhas | Watheq Mansour | Abubakr Mohamed | Mahmoud Fawzi | Majd Hawasly | Tamer Elsayed | Kareem Mohamed Darwish | Walid Magdy
Proceedings of The Third Arabic Natural Language Processing Conference: Shared Tasks
2022
Tracing Syntactic Change in the Scientific Genre: Two Universal Dependency-parsed Diachronic Corpora of Scientific English and German
Marie-Pauline Krielke | Luigi Talamo | Mahmoud Fawzi | Jörg Knappen
Proceedings of the Thirteenth Language Resources and Evaluation Conference
Marie-Pauline Krielke | Luigi Talamo | Mahmoud Fawzi | Jörg Knappen
Proceedings of the Thirteenth Language Resources and Evaluation Conference
We present two comparable diachronic corpora of scientific English and German from the Late Modern Period (17th c.–19th c.) annotated with Universal Dependencies. We describe several steps of data pre-processing and evaluate the resulting parsing accuracy showing how our pre-processing steps significantly improve output quality. As a sanity check for the representativity of our data, we conduct a case study comparing previously gained insights on grammatical change in the scientific genre with our data. Our results reflect the often reported trend of English scientific discourse towards heavy noun phrases and a simplification of the sentence structure (Halliday, 1988; Halliday and Martin, 1993; Biber and Gray, 2011; Biber and Gray, 2016). We also show that this trend applies to German scientific discourse as well. The presented corpora are valuable resources suitable for the contrastive analysis of syntactic diachronic change in the scientific genre between 1650 and 1900. The presented pre-processing procedures and their evaluations are applicable to other languages and can be useful for a variety of Natural Language Processing tasks such as syntactic parsing.