Abstract
Referring Expression Comprehension (ReC) is a task that involves localizing objects in images based on natural language expressions. Most ReC methods typically approach the task as a supervised learning problem. However, the need for costly annotations, such as clear image-text pairs or region-text pairs, hinders the scalability of existing approaches. In this work, we propose a novel scene graph-based framework that automatically generates high-quality pseudo region-query pairs. Our method harnesses scene graphs to capture the relationships between objects in images and generate expressions enriched with relation information. To ensure accurate mapping between visual regions and text, we introduce an external module that employs a calibration algorithm to filter out ambiguous queries. Additionally, we employ a rewriter module to enhance the diversity of our generated pseudo queries through rewriting. Extensive experiments demonstrate that our method outperforms previous pseudo-labeling methods by about 10%, 12%, and 11% on RefCOCO, RefCOCO+, and RefCOCOg, respectively. Furthermore, it surpasses the state-of-the-art unsupervised approach by more than 15% on the RefCOCO dataset.- Anthology ID:
- 2023.findings-emnlp.802
- Volume:
- Findings of the Association for Computational Linguistics: EMNLP 2023
- Month:
- December
- Year:
- 2023
- Address:
- Singapore
- Editors:
- Houda Bouamor, Juan Pino, Kalika Bali
- Venue:
- Findings
- SIG:
- Publisher:
- Association for Computational Linguistics
- Note:
- Pages:
- 11978–11990
- Language:
- URL:
- https://preview.aclanthology.org/icon-24-ingestion/2023.findings-emnlp.802/
- DOI:
- 10.18653/v1/2023.findings-emnlp.802
- Cite (ACL):
- Cantao Wu, Yi Cai, Liuwu Li, and Jiexin Wang. 2023. Scene Graph Enhanced Pseudo-Labeling for Referring Expression Comprehension. In Findings of the Association for Computational Linguistics: EMNLP 2023, pages 11978–11990, Singapore. Association for Computational Linguistics.
- Cite (Informal):
- Scene Graph Enhanced Pseudo-Labeling for Referring Expression Comprehension (Wu et al., Findings 2023)
- PDF:
- https://preview.aclanthology.org/icon-24-ingestion/2023.findings-emnlp.802.pdf