SummN: A Multi-Stage Summarization Framework for Long Input Dialogues and Documents

Yusen Zhang; Ansong Ni; Ziming Mao; Chen Henry Wu; Chenguang Zhu; Budhaditya Deb; Ahmed Awadallah; Dragomir Radev; Rui Zhang

doi:10.18653/v1/2022.acl-long.112

Summ^N: A Multi-Stage Summarization Framework for Long Input Dialogues and Documents

Yusen Zhang, Ansong Ni, Ziming Mao, Chen Henry Wu, Chenguang Zhu, Budhaditya Deb, Ahmed Awadallah, Dragomir Radev, Rui Zhang

Abstract

Text summarization helps readers capture salient information from documents, news, interviews, and meetings. However, most state-of-the-art pretrained language models (LM) are unable to efficiently process long text for many summarization tasks. In this paper, we propose Summ^N, a simple, flexible, and effective multi-stage framework for input texts that are longer than the maximum context length of typical pretrained LMs. Summ^N first splits the data samples and generates a coarse summary in multiple stages and then produces the final fine-grained summary based on it. Our framework can process input text of arbitrary length by adjusting the number of stages while keeping the LM input size fixed. Moreover, it can deal with both single-source documents and dialogues, and it can be used on top of different backbone abstractive summarization models. To the best of our knowledge, Summ^N is the first multi-stage split-then-summarize framework for long input summarization. Our experiments demonstrate that Summ^N outperforms previous state-of-the-art methods by improving ROUGE scores on three long meeting summarization datasets AMI, ICSI, and QMSum, two long TV series datasets from SummScreen, and a long document summarization dataset GovReport. Our data and code are available at https://github.com/psunlpgroup/Summ-N.

Anthology ID:: 2022.acl-long.112
Volume:: Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Month:: May
Year:: 2022
Address:: Dublin, Ireland
Editors:: Smaranda Muresan, Preslav Nakov, Aline Villavicencio
Venue:: ACL
SIG:
Publisher:: Association for Computational Linguistics
Note:
Pages:: 1592–1604
Language:
URL:: https://aclanthology.org/2022.acl-long.112
DOI:: 10.18653/v1/2022.acl-long.112
Bibkey:
Cite (ACL):: Yusen Zhang, Ansong Ni, Ziming Mao, Chen Henry Wu, Chenguang Zhu, Budhaditya Deb, Ahmed Awadallah, Dragomir Radev, and Rui Zhang. 2022. SummN: A Multi-Stage Summarization Framework for Long Input Dialogues and Documents. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1592–1604, Dublin, Ireland. Association for Computational Linguistics.
Cite (Informal):: SummN: A Multi-Stage Summarization Framework for Long Input Dialogues and Documents (Zhang et al., ACL 2022)
Copy Citation:
PDF:: https://preview.aclanthology.org/nschneid-patch-5/2022.acl-long.112.pdf
Code: psunlpgroup/summ-n + additional community code
Data: GovReport, QMSum, SummScreen

PDF Search Code