Batool Najeh Balah
2026
EGCSS at StanceNakba Shared Task: Cross-Topic Arabic Stance Detection for Two Middle East Issues
Asmaa Qindeel | Toka Khaled | Batool Najeh Balah | Eman Elrefai | Mahmoud Fawzi
Proceedings of the 2nd International Workshop on Nakba Narratives as Language Resources @ LREC 2026
Asmaa Qindeel | Toka Khaled | Batool Najeh Balah | Eman Elrefai | Mahmoud Fawzi
Proceedings of the 2nd International Workshop on Nakba Narratives as Language Resources @ LREC 2026
Stance detection continues to be an important task sitting at the intersection of Natural Language Processing (NLP) and Computational Social Science (CSS). In this work, we evaluate how different variations of BERT models perform on the cross-topic form of the task. In particular, we inspect their performance on the second subtask of the shared task StanceNakba 2026, where two topics are included, namely Arab Normalization with Israel and The Presence of Refugees in Arab Countries. We find that the best-performing model was bert-base-arabertv02-twitter, and we further improve its performance by providing context about the topic during the training phase, achieving an F1-score of 0.86 and ranking second among the participating teams.
The NakbaEcho Dataset: From Oral Testimonies to a Transcribed Arabic History Corpus
Batool Najeh Balah | Mahmoud Fawzi | Houda Elmimouni | Walid Magdy
Proceedings of the 2nd International Workshop on Nakba Narratives as Language Resources @ LREC 2026
Batool Najeh Balah | Mahmoud Fawzi | Houda Elmimouni | Walid Magdy
Proceedings of the 2nd International Workshop on Nakba Narratives as Language Resources @ LREC 2026
We present NakbaEcho, a dataset derived from Palestinian testimonies about the 1948 Nakba. The resource is constructed from transcribing over 2,180 hours of recorded interviews gathered through the Palestine Remembered Oral History index and linked to multiple repositories, including the Palestinian Oral History Archive (POHA) and YouTube-hosted interviews. We harmonize interview-level metadata and generate timestamp-aligned transcripts from the original Arabic recordings using an automatic transcription pipeline configured for Palestinian Arabic. The dataset includes speaker-labeled segments and auxiliary annotations designed to support downstream research in Arabic speech processing, natural language processing, digital humanities, and oral-history analysis. NakbaEcho contributes a structured computational resource for studying Palestinian oral testimony while expanding the availability of dialectal Arabic materials for speech, text, and social research.