ARB: A Comprehensive Arabic Multimodal Reasoning Benchmark

Sara Ghaboura; Shubham Patle; Ketan More; Wafa Hamad Mohamed Alghallabi; Omkar Thawakar; Jorma Laaksonen; Hisham Cholakkal; Salman Khan; Rao Anwer

ARB: A Comprehensive Arabic Multimodal Reasoning Benchmark

Sara Ghaboura, Shubham Patle, Ketan More, Wafa Hamad Mohamed Alghallabi, Omkar Thawakar, Jorma Laaksonen, Hisham Cholakkal, Salman Khan, Rao Anwer

Abstract

As Large Multimodal Models (LMMs) become more capable, there is growing interest in evaluating their reasoning processes alongside their final outputs. However, most existing benchmarks remain focused on English, overlooking languages with rich linguistic and cultural depth such as Arabic. To address this gap, we introduce the Comprehensive Arabic Multimodal Reasoning Benchmark (ARB), the first benchmark designed to evaluate step-by-step reasoning in Arabic across both textual and visual modalities. ARB covers 11 diverse domains and over 40 subfields, including visual reasoning, optical character recognition, scientific analysis, and cultural interpretation. It comprises 2,219 multimodal samples paired with over 8K human-curated reasoning steps and corresponding actions, verified through a human-in-the-loop process. We evaluated 15 state-of-the-art open- and closed-source LMMs and found persistent challenges in coherence, faithfulness, and cultural grounding. ARB provides a structured framework for diagnosing multimodal reasoning in underrepresented languages, marking a critical step toward inclusive, transparent, and culturally aware AI systems. The benchmark, rubric, and evaluation suite are publicly available

Anthology ID:: 2026.lrec-main.723
Volume:: Proceedings of the Fifteenth Language Resources and Evaluation Conference
Month:: May
Year:: 2026
Address:: Palma de Mallorca, Spain
Editors:: Stelios Piperidis, Núria Bel, Henk van den Heuvel, Nancy Ide, Simon Krek, Antonio Toral
Venue:: LREC
SIG:
Publisher:: ELRA Language Resource Association
Note:
Pages:: 9202–9216
Language:
URL:: https://preview.aclanthology.org/ingest-lrec/2026.lrec-main.723/
DOI:
Bibkey:
Cite (ACL):: Sara Ghaboura, Shubham Patle, Ketan More, Wafa Hamad Mohamed Alghallabi, Omkar Thawakar, Jorma Laaksonen, Hisham Cholakkal, Salman Khan, and Rao Anwer. 2026. ARB: A Comprehensive Arabic Multimodal Reasoning Benchmark. International Conference on Language Resources and Evaluation, main:9202–9216.
Cite (Informal):: ARB: A Comprehensive Arabic Multimodal Reasoning Benchmark (Ghaboura et al., LREC 2026)
Copy Citation:
PDF:: https://preview.aclanthology.org/ingest-lrec/2026.lrec-main.723.pdf

PDF Cite Search Fix data