SIMMC 2.0: A Task-oriented Dialog Dataset for Immersive Multimodal Conversations

Satwik Kottur, Seungwhan Moon, Alborz Geramifard, Babak Damavandi


Abstract
Next generation task-oriented dialog systems need to understand conversational contexts with their perceived surroundings, to effectively help users in the real-world multimodal environment. Existing task-oriented dialog datasets aimed towards virtual assistance fall short and do not situate the dialog in the user’s multimodal context. To overcome, we present a new dataset for Situated and Interactive Multimodal Conversations, SIMMC 2.0, which includes 11K task-oriented user<->assistant dialogs (117K utterances) in the shopping domain, grounded in immersive and photo-realistic scenes. The dialogs are collection using a two-phase pipeline: (1) A novel multimodal dialog simulator generates simulated dialog flows, with an emphasis on diversity and richness of interactions, (2) Manual paraphrasing of generating utterances to draw from natural language distribution. We provide an in-depth analysis of the collected dataset, and describe in detail the four main benchmark tasks we propose for SIMMC 2.0. Our baseline model, powered by the state-of-the-art language model, shows promising results, and highlights new challenges and directions for the community to study.
Anthology ID:
2021.emnlp-main.401
Volume:
Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing
Month:
November
Year:
2021
Address:
Online and Punta Cana, Dominican Republic
Editors:
Marie-Francine Moens, Xuanjing Huang, Lucia Specia, Scott Wen-tau Yih
Venue:
EMNLP
SIG:
Publisher:
Association for Computational Linguistics
Note:
Pages:
4903–4912
Language:
URL:
https://aclanthology.org/2021.emnlp-main.401
DOI:
10.18653/v1/2021.emnlp-main.401
Bibkey:
Cite (ACL):
Satwik Kottur, Seungwhan Moon, Alborz Geramifard, and Babak Damavandi. 2021. SIMMC 2.0: A Task-oriented Dialog Dataset for Immersive Multimodal Conversations. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 4903–4912, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
Cite (Informal):
SIMMC 2.0: A Task-oriented Dialog Dataset for Immersive Multimodal Conversations (Kottur et al., EMNLP 2021)
Copy Citation:
PDF:
https://preview.aclanthology.org/naacl24-info/2021.emnlp-main.401.pdf
Software:
 2021.emnlp-main.401.Software.zip
Video:
 https://preview.aclanthology.org/naacl24-info/2021.emnlp-main.401.mp4
Code
 facebookresearch/simmc2
Data
SIMMC2.0SIMMCVisual Question Answering