How Foundation Models Behave for Arabic Image Captioning?
Khaoula Dahimi, Amel Belabbaci, Hadda Cherroun, Abdelhamid Haouhat
Abstract
Image captioning plays a crucial role in numerous applications, including educational systems. However, ensuring caption quality remains a significant challenge, particularly for morphologically rich, low-resource languages such as Arabic. We investigate an evaluation of Arabic image captioning using state-of-the-art multimodal foundation models. We systematically assess the performance of leading models—Gemini, Gemma, LLaMA, and Fanar. Our evaluation framework employs a diverse set of metrics spanning rule-based, learnable, visually-grounded, and LLM-based approaches to capture semantic accuracy, linguistic fluency, and hallucination detection. Experiments are conducted on two benchmark datasets: Flickr8k-Arabic and JEEM. Our findings reveal significant performance variations across models and evaluation metrics, highlighting the need for Arabic-specific optimization in multimodal architectures.- Anthology ID:
- 2026.osact-1.6
- Volume:
- The 7th Workshop on Open-Source Arabic Corpora and Processing Tools (OSACT7) with 5 Shared Tasks
- Month:
- May
- Year:
- 2026
- Address:
- Palma, Mallorca (Spain)
- Editors:
- Hend Al-Khalifa, Mo El-Haj, Saad Ezzini
- Venues:
- OSACT | WS
- SIG:
- Publisher:
- Association for Computational Linguistics
- Note:
- Pages:
- 49–58
- Language:
- External URL:
- https://lrec.elra.info/lrec2026-ws-osact-06
- DOI:
- 10.63317/3bhwcpon3fv5
- Cite (ACL):
- Khaoula Dahimi, Amel Belabbaci, Hadda Cherroun, and Abdelhamid Haouhat. 2026. How Foundation Models Behave for Arabic Image Captioning?. In The 7th Workshop on Open-Source Arabic Corpora and Processing Tools (OSACT7) with 5 Shared Tasks, pages 49–58, Palma, Mallorca (Spain). Association for Computational Linguistics.
- Cite (Informal):
- How Foundation Models Behave for Arabic Image Captioning? (Dahimi et al., OSACT 2026)