MEUR: A Benchmark for Evaluating Vision-Language Models on Multimodal Event Understanding and Reasoning

Zimu Wang, Yuqi Wang, Tong Chen, Changyu Zeng, Hongbin Na, Nijia Han, Fuyu Xing, Qi Chen, Qiufeng Wang, Anh Nguyen, Shuihua Wang, Ling Chen, Jionglong Su, Haiyang Zhang, Wei Wang


Abstract
Event understanding and reasoning play critical roles in thoroughly evaluating the capabilities of Vision-Language Models (VLMs); however, existing Visual Question Answering (VQA) datasets predominantly focus on entity-centric questions, while event- or action-related questions are limited in scale and suffer from significant shortcut issues. We introduce MEUR, the first Multimodal Event Understanding and Reasoning dataset consisting of 1,200 images and 4,217 questions, necessitating VLMs with a diverse range of multimodal understanding and reasoning capabilities to answer, ranging from basic event recognition to more complex tasks such as counting and comparison. To streamline the annotation process, we propose a novel semi-automated pipeline that combines advanced VLMs with human annotators, achieving high quality and efficiency. We conduct extensive experiments on state-of-the-art non-thinking and thinking VLMs to demonstrate their capabilities and limitations in multimodal event understanding and reasoning. Furthermore, we provide a detailed error analysis that points out promising directions for future research.
Anthology ID:
2026.lrec-1.134
Volume:
Proceedings of the Fifteenth Language Resources and Evaluation Conference
Month:
May
Year:
2026
Address:
Palma de Mallorca, Spain
Editors:
Stelios Piperidis, Núria Bel, Henk van den Heuvel, Nancy Ide, Simon Krek, Antonio Toral
Venue:
LREC
SIG:
Publisher:
ELRA Language Resource Association
Note:
Pages:
1696–1709
Language:
External URL:
https://lrec.elra.info/lrec2026-main-134
DOI:
10.63317/4ftadqtyt374
Bibkey:
Cite (ACL):
Zimu Wang, Yuqi Wang, Tong Chen, Changyu Zeng, Hongbin Na, Nijia Han, Fuyu Xing, Qi Chen, Qiufeng Wang, Anh Nguyen, Shuihua Wang, Ling Chen, Jionglong Su, Haiyang Zhang, and Wei Wang. 2026. MEUR: A Benchmark for Evaluating Vision-Language Models on Multimodal Event Understanding and Reasoning. In Proceedings of the Fifteenth Language Resources and Evaluation Conference, pages 1696–1709, Palma de Mallorca, Spain. ELRA Language Resource Association.
Cite (Informal):
MEUR: A Benchmark for Evaluating Vision-Language Models on Multimodal Event Understanding and Reasoning (Wang et al., LREC 2026)
Copy Citation: