Maziar Kianimoghadam Jouneghani
2026
MKJ at SemEval-2026 Task 9: A Comparative Study of Generalist, Specialist, and Ensemble Strategies for Multilingual Polarization
Maziar Kianimoghadam Jouneghani
Proceedings of the 20th International Workshop on Semantic Evaluation (2026)
Maziar Kianimoghadam Jouneghani
Proceedings of the 20th International Workshop on Semantic Evaluation (2026)
We present a systematic study of multilingual polarization detection across 22 languages for SemEval-2026 Task 9 (Subtask 1), contrasting multilingual generalists with language-specific specialists and hybrid ensembles. While a standard generalist like XLM-RoBERTa suffices when its tokenizer aligns with the target text, it may struggle with distinct scripts (e.g., Khmer, Odia) where monolingual specialists yield significant gains. Rather than enforcing a single universal architecture, we adopt a language-adaptive selection strategy that chooses among multilingual generalists, language-specific specialists, and hybrid ensembles based on development performance. Additionally, cross-lingual augmentation via NLLB-200 yielded mixed results, often underperforming native architecture selection and degrading morphologically rich tracks. Our final system achieves an overall macro-averaged F1 score of 0.796 and an average accuracy of 0.826 across all 22 tracks. Code and final test predictions are publicly available at: https://github.com/Maziarkiani/SemEval2026-Task9-Subtask1-Polarization.
Beyond Fake News Detection: A Community-based Study of the Multicultural Nature of Information Disorder
Sara Gemelli | Giulia Di Cristina | Yiran Zhang | Md Azizul Hoque | Alberto De La Torre Solís | Mohamad Mojtaba Behboudi Eshkiki | Nikolai Efimov | Mariia Everstova | Caterina Maria Cappello | Maziar Kianimoghadam Jouneghani | Payam Latifi | Yashar Mahboudi | Farzaneh Mohseni | Dario Placenti | Tommaso Caselli | Manuela Sanguinetti | Aurora Scarpellini | Chiara Zanchi | Usman Naseem | Marco Antonio Stranisci | Simona Frenda
Proceedings of the Fifteenth Language Resources and Evaluation Conference
Sara Gemelli | Giulia Di Cristina | Yiran Zhang | Md Azizul Hoque | Alberto De La Torre Solís | Mohamad Mojtaba Behboudi Eshkiki | Nikolai Efimov | Mariia Everstova | Caterina Maria Cappello | Maziar Kianimoghadam Jouneghani | Payam Latifi | Yashar Mahboudi | Farzaneh Mohseni | Dario Placenti | Tommaso Caselli | Manuela Sanguinetti | Aurora Scarpellini | Chiara Zanchi | Usman Naseem | Marco Antonio Stranisci | Simona Frenda
Proceedings of the Fifteenth Language Resources and Evaluation Conference
Recognizing disinformation is a challenging task for humans and AI systems. News can be false, misleading, or harmful, and its interpretation often depends on the cultural context of the audience. However, existing datasets rarely account for these contextual and cultural differences, as they are typically not designed from the perspective of news consumers. To address this gap, in this paper, we present the Information Disorder (InDor) corpus, a multilingual dataset of news articles in English, Farsi, Italian, and Russian, annotated for information disorder detection and explanation. The corpus was developed through a participatory process involving contributors from diverse cultural and professional backgrounds, who engaged in data collection, annotation, and evaluation of Large Language Model (LLM) performance on the task. Our findings highlight that false and manipulated news manifest differently across cultural settings, and that current LLMs fail to adequately capture this complexity. This underscores the need for culturally aware computational approaches in the study of information disorder.