Kais Attia
2026
StanceNakba Shared Task: Actor and Topic-Aware Stance Detection in Public Discourse
Kholoud Khalil Aldous | Md. Rafiul Biswas | Mabrouka Bessghaier | Shimaa Amer Ibrahim | Kais Attia | Wajdi Zaghouani
Proceedings of the 2nd International Workshop on Nakba Narratives as Language Resources @ LREC 2026
Kholoud Khalil Aldous | Md. Rafiul Biswas | Mabrouka Bessghaier | Shimaa Amer Ibrahim | Kais Attia | Wajdi Zaghouani
Proceedings of the 2nd International Workshop on Nakba Narratives as Language Resources @ LREC 2026
We present StanceNakba 2026, a shared task on stance detection in polarized social media discourse related to the Palestinian-Israeli conflict, organized as part of Nakba-NLP 2026 at LREC-COLING 2026. The task introduces two subtasks: Subtask A (Actor-Level Stance Detection), which classifies English social media posts as Pro-Palestine, Pro-Israel, or Neutral; and Subtask B (Cross-Topic Stance Detection), which identifies Favor, Against, or Neither stances in Arabic posts toward two conflict-related topics, normalization with Israel and refugee presence in Jordan. The task is grounded in an annotated dataset of 2,606 social media posts. A total of 7 teams participated in Subtask A, and 6 teams in Subtask B. Participating systems primarily fine-tuned Arabic and multilingual transformer-based models, including MARBERT, AraBERT, and DeBERTa-v3 variants, with several teams employing cross-validation, ensemble methods, and topic-conditioned architectures. The best-performing systems achieved a Macro F1 of 0.9620 on Subtask A and 0.8724 on Subtask B, demonstrating that transformer-based approaches are highly effective for conflict-domain stance detection while highlighting persistent challenges in cross-topic generalization and neutral class prediction
Nakba Discourse 2025: A Bilingual Social Media Dataset for Collective Trauma Analysis
Wajdi Zaghouani | Mabrouka Bessghaier | Kais Attia
Proceedings of the 2nd International Workshop on Nakba Narratives as Language Resources @ LREC 2026
Wajdi Zaghouani | Mabrouka Bessghaier | Kais Attia
Proceedings of the 2nd International Workshop on Nakba Narratives as Language Resources @ LREC 2026
We introduce Nakba Discourse 2025, a bilingual full-year social media dataset capturing Arabic and English discourse about the 1948 Palestinian Nakba across Twitter/X and Facebook from January to December 2025. The corpus contains 70,312 unique posts organized into intersecting sub-corpora by language, sentiment, gender, geography, and platform, with engagement metadata and automatically extracted rhetorical features. Analyses reveal systematic variation in engagement and framing across communities. Per-post engagement is highest in Israel and UK subsets (50.62 and 49.08 average likes respectively), while Arabic-language discourse shows markedly lower per-post engagement. Sentiment distribution is strongly skewed, with negative sentiment posts outnumbering positive ones at an 11:1 ratio (54,424 vs. 4,827 posts). Despite dramatic variation in absolute engagement levels, virality rates remain structurally constant at approximately 10% across all Twitter/X sub-corpora, regardless of language, gender, or geography, pointing to platform-level amplification regularities. Gender analysis reveals that women achieve proportional virality equal to men despite producing roughly one-third the volume of posts. Temporal patterns align with cultural calendars, including Thursday peaks associated with Jumu’ah across Arabic and female subsets, and Sunday peaks in English-language subsets reflecting Western media cycles. The dataset will be released for research use and supports multilingual stance detection, virality modeling, rhetorical analysis, and computational studies of digital political memory.
ArPoMeme: An Annotated Arabic Multimodal Dataset for Political Ideology and Polarization
Wajdi Zaghouani | Kais Attia | Md. Rafiul Biswas | Fadhl Eryani
Proceedings of the Fifteenth Language Resources and Evaluation Conference
Wajdi Zaghouani | Kais Attia | Md. Rafiul Biswas | Fadhl Eryani
Proceedings of the Fifteenth Language Resources and Evaluation Conference
Memes have become a prominent medium of political communication in the Arab world, reflecting how humor, imagery, and text interact to express ideological and cultural positions. Despite the centrality of memes to online political discourse, there is a lack of systematically curated resources for analyzing their multimodal and ideological dimensions in Arabic. This paper presents ArPoMeme, a large-scale dataset of approximately 7,300 Arabic political memes categorized by ideological orientation, including Leftist, Islamist, Pan-Arabist, and Satirical perspectives. The dataset captures the diversity of Arabic meme ecosystems by grounding classification in the self-identification of public Facebook pages and groups that produce and disseminate these memes. To ensure both scale and accuracy, we designed a semi-automated data collection pipeline combining Playwright-based Facebook scraping with Google Drive synchronization, followed by text extraction using the Qwen2.5-VL-7B vision–language model. The extracted text was manually verified and annotated for three polarization dimensions: Us vs. Them framing, Hostility toward out-groups, and Calls to action. Annotation was conducted through a custom Streamlit-based interface supporting distributed labeling, real-time tracking, and version control. The resulting dataset links visual content, textual messages, and ideological orientation, enabling fine-grained analysis of political antagonism, mobilization, and humor. Quantitative analysis of the annotated corpus reveals strong asymmetries in antagonistic framing across ideological groups, with Islamist and satirical memes exhibiting the highest levels of hostility and mobilization cues. The dataset and the annotation tool offer a reproducible and publicly available resource for studying Arabic political discourse, multimodal ideology detection, and polarization dynamics.