From Perception to Reasoning: Enhancing Vision-Language Models for Mobile UI Understanding

Settaluri Lakshmi Sravanthi; Ankit Mishra; Debjyoti Mondal; Subhadarshi Panda; Rituraj Singh; Pushpak Bhattacharyya

From Perception to Reasoning: Enhancing Vision-Language Models for Mobile UI Understanding

Settaluri Lakshmi Sravanthi, Ankit Mishra, Debjyoti Mondal, Subhadarshi Panda, Rituraj Singh, Pushpak Bhattacharyya

Abstract

Accurately grounding visual and textual elements within mobile user interfaces (UIs) remains a significant challenge for Vision-Language Models (VLMs). Visual grounding, a critical task in this domain, involves identifying the most relevant UI element or region based on a natural language query—a process that requires both precise perception and context-aware reasoning. In this work, we present - **MoUI**, a light-weight mobile UI understanding model trained on **MoIT**, an instruction-tuning dataset specifically tailored for mobile screen understanding and grounding, designed to bridge the gap between user intent and visual semantics. Complementing this dataset, we also present a human-annotated reasoning benchmark **MoIQ** that rigorously evaluates complex inference capabilities over mobile UIs. To harness these resources effectively, we propose a two-stage training approach that separately addresses perception and reasoning tasks, leading to stronger perception capabilities and improvement in reasoning abilities. Through extensive experiments, we demonstrate that our MoUI models achieve significant gains in accuracy across all perception tasks and _state-of-the-art_ results on public reasoning benchmark **ComplexQA (78%) and our MoIQ (49%)**. We will be open-sourcing our dataset, code, and models to foster further research and innovation in the field.

Anthology ID:: 2025.findings-acl.1295
Volume:: Findings of the Association for Computational Linguistics: ACL 2025
Month:: July
Year:: 2025
Address:: Vienna, Austria
Editors:: Wanxiang Che, Joyce Nabende, Ekaterina Shutova, Mohammad Taher Pilehvar
Venue:: Findings
SIG:
Publisher:: Association for Computational Linguistics
Note:
Pages:: 25250–25269
Language:
URL:: https://preview.aclanthology.org/display_plenaries/2025.findings-acl.1295/
DOI:
Bibkey:
Cite (ACL):: Settaluri Lakshmi Sravanthi, Ankit Mishra, Debjyoti Mondal, Subhadarshi Panda, Rituraj Singh, and Pushpak Bhattacharyya. 2025. From Perception to Reasoning: Enhancing Vision-Language Models for Mobile UI Understanding. In Findings of the Association for Computational Linguistics: ACL 2025, pages 25250–25269, Vienna, Austria. Association for Computational Linguistics.
Cite (Informal):: From Perception to Reasoning: Enhancing Vision-Language Models for Mobile UI Understanding (Sravanthi et al., Findings 2025)
Copy Citation:
PDF:: https://preview.aclanthology.org/display_plenaries/2025.findings-acl.1295.pdf

PDF Cite Search Fix data