MEUR: A Benchmark for Evaluating Vision-Language Models on Multimodal Event Understanding and Reasoning
Zimu Wang, Yuqi Wang, Tong Chen, Changyu Zeng, Hongbin Na, Nijia Han, Fuyu Xing, Qi Chen, Qiufeng Wang, Anh Nguyen, Shuihua Wang, Ling Chen, Jionglong Su, Haiyang Zhang, Wei Wang
Abstract
Event understanding and reasoning play critical roles in thoroughly evaluating the capabilities of Vision-Language Models (VLMs); however, existing Visual Question Answering (VQA) datasets predominantly focus on entity-centric questions, while event- or action-related questions are limited in scale and suffer from significant shortcut issues. We introduce MEUR, the first Multimodal Event Understanding and Reasoning dataset consisting of 1,200 images and 4,217 questions, necessitating VLMs with a diverse range of multimodal understanding and reasoning capabilities to answer, ranging from basic event recognition to more complex tasks such as counting and comparison. To streamline the annotation process, we propose a novel semi-automated pipeline that combines advanced VLMs with human annotators, achieving high quality and efficiency. We conduct extensive experiments on state-of-the-art non-thinking and thinking VLMs to demonstrate their capabilities and limitations in multimodal event understanding and reasoning. Furthermore, we provide a detailed error analysis that points out promising directions for future research.- Anthology ID:
- 2026.lrec-1.134
- Volume:
- Proceedings of the Fifteenth Language Resources and Evaluation Conference
- Month:
- May
- Year:
- 2026
- Address:
- Palma de Mallorca, Spain
- Editors:
- Stelios Piperidis, Núria Bel, Henk van den Heuvel, Nancy Ide, Simon Krek, Antonio Toral
- Venue:
- LREC
- SIG:
- Publisher:
- ELRA Language Resource Association
- Note:
- Pages:
- 1696–1709
- Language:
- External URL:
- https://lrec.elra.info/lrec2026-main-134
- DOI:
- 10.63317/4ftadqtyt374
- Cite (ACL):
- Zimu Wang, Yuqi Wang, Tong Chen, Changyu Zeng, Hongbin Na, Nijia Han, Fuyu Xing, Qi Chen, Qiufeng Wang, Anh Nguyen, Shuihua Wang, Ling Chen, Jionglong Su, Haiyang Zhang, and Wei Wang. 2026. MEUR: A Benchmark for Evaluating Vision-Language Models on Multimodal Event Understanding and Reasoning. In Proceedings of the Fifteenth Language Resources and Evaluation Conference, pages 1696–1709, Palma de Mallorca, Spain. ELRA Language Resource Association.
- Cite (Informal):
- MEUR: A Benchmark for Evaluating Vision-Language Models on Multimodal Event Understanding and Reasoning (Wang et al., LREC 2026)