OVQA: A Dataset for Visual Question Answering and Multimodal Research in Odia Language

Shantipriya Parida; Shashikanta Sahoo; Sambit Sekhar; Kalyanamalini Sahoo; Ketan Kotwal; Sonal Khosla; Satya Ranjan Dash; Aneesh Bose; Guneet Singh Kohli; Smruti Smita Lenka; Ondřej Bojar

OVQA: A Dataset for Visual Question Answering and Multimodal Research in Odia Language

Shantipriya Parida, Shashikanta Sahoo, Sambit Sekhar, Kalyanamalini Sahoo, Ketan Kotwal, Sonal Khosla, Satya Ranjan Dash, Aneesh Bose, Guneet Singh Kohli, Smruti Smita Lenka, Ondřej Bojar

Abstract

This paper introduces OVQA, the first multimodal dataset designed for visual question-answering (VQA), visual question elicitation (VQE), and multimodal research for the low-resource Odia language. The dataset was created by manually translating 6,149 English question-answer pairs, each associated with 6,149 unique images from the Visual Genome dataset. This effort resulted in 27,809 English-Odia parallel sentences, ensuring a semantic match with the corresponding visual information. Several baseline experiments were conducted on the dataset, including visual question answering and visual question elicitation. The dataset is the first VQA dataset for the low-resource Odia language and will be released for multimodal research purposes and also help researchers extend for other low-resource languages.

Anthology ID:: 2025.indonlp-1.7
Volume:: Proceedings of the First Workshop on Natural Language Processing for Indo-Aryan and Dravidian Languages
Month:: January
Year:: 2025
Address:: Abu Dhabi
Editors:: Ruvan Weerasinghe, Isuri Anuradha, Deshan Sumanathilaka
Venues:: IndoNLP | WS
SIG:
Publisher:: Association for Computational Linguistics
Note:
Pages:: 58–66
Language:
URL:: https://preview.aclanthology.org/fix-sig-urls/2025.indonlp-1.7/
DOI:
Bibkey:
Cite (ACL):: Shantipriya Parida, Shashikanta Sahoo, Sambit Sekhar, Kalyanamalini Sahoo, Ketan Kotwal, Sonal Khosla, Satya Ranjan Dash, Aneesh Bose, Guneet Singh Kohli, Smruti Smita Lenka, and Ondřej Bojar. 2025. OVQA: A Dataset for Visual Question Answering and Multimodal Research in Odia Language. In Proceedings of the First Workshop on Natural Language Processing for Indo-Aryan and Dravidian Languages, pages 58–66, Abu Dhabi. Association for Computational Linguistics.
Cite (Informal):: OVQA: A Dataset for Visual Question Answering and Multimodal Research in Odia Language (Parida et al., IndoNLP 2025)
Copy Citation:
PDF:: https://preview.aclanthology.org/fix-sig-urls/2025.indonlp-1.7.pdf

PDF Cite Search Fix data