2025
pdf
bib
abs
ImageEval 2025: The First Arabic Image Captioning Shared Task
Ahlam Bashiti
|
Alaa Aljabari
|
Hadi Khaled Hamoud
|
Md. Rafiul Biswas
|
Bilal Mohammed Shalash
|
Mustafa Jarrar
|
Fadi Zaraket
|
George Mikros
|
Ehsaneddin Asgari
|
Wajdi Zaghouani
Proceedings of The Third Arabic Natural Language Processing Conference: Shared Tasks
We present ImageEval 2025, the first shared task dedicated to Arabic image captioning. The task addresses the critical gap in multimodal Arabic NLP by focusing on two complementary subtasks: (1) creating the first open-source, manually-captioned Arabic image dataset through a collaborative datathon, and (2) developing and evaluating Arabic image captioning models. A total of 44 teams registered, of which eight submitted during the test phase, producing 111 valid submissions. Evaluation was conducted using automatic metrics, LLM-based judgment, and human assessment. In Subtask 1, the best-performing system achieved a cosine similarity of 65.5, while in Subtask 2, the top score was 60.0. Although these results show encouraging progress, they also confirm that Arabic image captioning remains a challenging task, particularly due to cultural grounding requirements, morphological richness, and dialectal variation. All datasets, baseline models, and evaluation tools are released publicly to support future research in Arabic multimodal NLP.
pdf
bib
abs
MAHED Shared Task: Multimodal Detection of Hope and Hate Emotions in Arabic Content
Wajdi Zaghouani
|
Md. Rafiul Biswas
|
Mabrouka Bessghaier
|
Shimaa Ibrahim
|
George Mikros
|
Abul Hasnat
|
Firoj Alam
Proceedings of The Third Arabic Natural Language Processing Conference: Shared Tasks
This paper presents the MAHED 2025 Shared Task on Multimodal Detection of Hope and Hate Emotions in Arabic Content, comprising three subtasks: (1) text-based classification of Arabic content into hate and hope, (2) multi-task learning for joint prediction of emotions, offensive content, and hate speech, and (3) multimodal detection of hateful content in Arabic memes. We provide three high-quality datasets totaling over 22,000 instances sourced from social media platforms, annotated by native Arabic speakers with Cohen’s Kappa exceeding 0.85. Our evaluation attracted 46 leaderboard submissions from participants, with systems leveraging Arabic-specific pre-trained language models (AraBERT, MARBERT), large language models (GPT-4, Gemini), and multimodal fusion architectures combining CLIP vision encoders with Arabic text models. The best-performing systems achieved macro F1-scores of 0.723 (Task 1), 0.578 (Task 2), and 0.796 (Task 3), with top teams employing ensemble methods, class-weighted training, and OCR-aware multimodal fusion. Analysis reveals persistent challenges in dialectal robustness, minority class detection for hope speech, and highlights key directions for future Arabic content moderation research.
pdf
bib
Proceedings of the 1stWorkshop on GenAI Content Detection (GenAIDetect)
Firoj Alam
|
Preslav Nakov
|
Nizar Habash
|
Iryna Gurevych
|
Shammur Chowdhury
|
Artem Shelmanov
|
Yuxia Wang
|
Ekaterina Artemova
|
Mucahid Kutlu
|
George Mikros
Proceedings of the 1stWorkshop on GenAI Content Detection (GenAIDetect)
pdf
bib
abs
GenAI Content Detection Task 2: AI vs. Human – Academic Essay Authenticity Challenge
Shammur Absar Chowdhury
|
Hind Almerekhi
|
Mucahid Kutlu
|
Kaan Efe Keleş
|
Fatema Ahmad
|
Tasnim Mohiuddin
|
George Mikros
|
Firoj Alam
Proceedings of the 1stWorkshop on GenAI Content Detection (GenAIDetect)
This paper presents a comprehensive overview of the first edition of the Academic Essay Authenticity Challenge, organized as part of the GenAI Content Detection shared tasks collocated with COLING 2025. This challenge focuses on detecting machine-generated vs human-authored essays for academic purposes. The task is defined as follows: “Given an essay, identify whether it is generated by a machine or authored by a human.” The challenge involves two languages: English and Arabic. During the evaluation phase, 25 teams submitted systems for English and 21 teams for Arabic, reflecting substantial interest in the task. Finally, five teams submitted system description papers. The majority of submissions utilized fine-tuned transformer-based models, with one team employing Large Language Models (LLMs) such as Llama 2 and Llama 3. This paper outlines the task formulation, details the dataset construction process, and explains the evaluation framework. Additionally, we present a summary of the approaches adopted by participating teams. Nearly all submitted systems outperformed the n-gram-based baseline, with the top-performing systems achieving F1 scores exceeding 0.98 for both languages, indicating significant progress in the detection of machine-generated text.
2024
pdf
bib
abs
Establishing Control Corpora for Depression Detection in Modern Greek: Methodological Insights
Vivian Stamou
|
George Mikros
|
George Markopoulos
|
Spyridoula Varlokosta
Proceedings of the Fifth Workshop on Resources and ProcessIng of linguistic, para-linguistic and extra-linguistic Data from people with various forms of cognitive/psychiatric/developmental impairments @LREC-COLING 2024
This paper presents a methodological approach for establishing control corpora in the context of depression detection in the Modern Greek language. We discuss various methods used to create control corpora, focusing on the challenge of selecting representative samples from the general population when the target reference is the depressed population. Our approach includes traditional random selection among Twitter users, as well as an innovative method for creating topic-oriented control corpora. Through this study, we provide insights into the development of control corpora, offering valuable considerations for researchers working on similar projects in linguistic analysis and mental health studies. In addition, we identify several dominant topics in the depressed population such as religion, sentiments, health and digestion, which seem to align with findings consistently reported in the literature
2002
pdf
bib
Quantitative parameters in corpus design: Estimating the optimum text size in Modern Greek language
George Mikros
Proceedings of the Third International Conference on Language Resources and Evaluation (LREC’02)
2000
pdf
bib
Modern Greek Corpus Taxonomy
George Mikros
|
George Carayannis
Proceedings of the Second International Conference on Language Resources and Evaluation (LREC’00)