SagaQA: A Multi-hop Reasoning Benchmark for Long-form Narrative Understanding in TV Series
Galann Pennec, Zhengyuan Liu, Nicholas Asher, Philippe Muller, Nancy Chen
Abstract
We introduce SagaQA, a long-form video benchmark for multi-hop reasoning over full-length TV series. Existing video reasoning benchmarks often emphasize local understanding of adjacent frames or clips. SagaQA addresses this gap by requiring high-level comprehension of extended multimodal narratives in entire TV shows. A distinguishing feature of SagaQA is the granularity of its reasoning steps. Our dataset necessitates long-range reasoning hops to connect information across completely different episodes. This requires models to reason over entire events and actions, demanding a deep understanding of the show’s narration and progression at a multimodal level. Motivated by recent progress in agentic methods, we further study how different planning strategies handle such complex reasoning. We categorize these approaches into three classes—parallel, sequential, and hybrid planners—and evaluate their ability to generate coherent and complete reasoning plans. Our results on SagaQA suggest that hybrid planners consistently produce higher-quality plans and exhibit stronger capabilities for complex, high-level narrative understanding in TV shows.- Anthology ID:
- 2026.sigdial-1.22
- Volume:
- Proceedings of the 27th Annual Meeting of the Special Interest Group on Discourse and Dialogue
- Month:
- August
- Year:
- 2026
- Address:
- Atlanta, Georgia, USA
- Editors:
- Jinho D. Choi, Yun-Nung Chen, Kotaro Funakoshi, Ali Emami
- Venue:
- SIGDIAL
- SIG:
- SIGDIAL
- Publisher:
- Association for Computational Linguistics
- Note:
- Pages:
- 307–326
- Language:
- URL:
- https://preview.aclanthology.org/cawl-year/2026.sigdial-1.22/
- DOI:
- Cite (ACL):
- Galann Pennec, Zhengyuan Liu, Nicholas Asher, Philippe Muller, and Nancy Chen. 2026. SagaQA: A Multi-hop Reasoning Benchmark for Long-form Narrative Understanding in TV Series. In Proceedings of the 27th Annual Meeting of the Special Interest Group on Discourse and Dialogue, pages 307–326, Atlanta, Georgia, USA. Association for Computational Linguistics.
- Cite (Informal):
- SagaQA: A Multi-hop Reasoning Benchmark for Long-form Narrative Understanding in TV Series (Pennec et al., SIGDIAL 2026)
- PDF:
- https://preview.aclanthology.org/cawl-year/2026.sigdial-1.22.pdf