@inproceedings{yang-wan-2022-investigating,
    title = "Investigating Metric Diversity for Evaluating Long Document Summarisation",
    author = "Yang, Cai  and
      Wan, Stephen",
    editor = "Cohan, Arman  and
      Feigenblat, Guy  and
      Freitag, Dayne  and
      Ghosal, Tirthankar  and
      Herrmannova, Drahomira  and
      Knoth, Petr  and
      Lo, Kyle  and
      Mayr, Philipp  and
      Shmueli-Scheuer, Michal  and
      de Waard, Anita  and
      Wang, Lucy Lu",
    booktitle = "Proceedings of the Third Workshop on Scholarly Document Processing",
    month = oct,
    year = "2022",
    address = "Gyeongju, Republic of Korea",
    publisher = "Association for Computational Linguistics",
    url = "https://preview.aclanthology.org/sigedu-bea-out-of-sync-correction/2022.sdp-1.13/",
    pages = "115--125",
    abstract = "Long document summarisation, a challenging summarisation scenario, is the focus of the recently proposed LongSumm shared task. One of the limitations of this shared task has been its use of a single family of metrics for evaluation (the ROUGE metrics). In contrast, other fields, like text generation, employ multiple metrics. We replicated the LongSumm evaluation using multiple test set samples (vs. the single test set of the official shared task) and investigated how different metrics might complement each other in this evaluation framework. We show that under this more rigorous evaluation, (1) some of the key learnings from Longsumm 2020 and 2021 still hold, but the relative ranking of systems changes, and (2) the use of additional metrics reveals additional high-quality summaries missed by ROUGE, and (3) we show that SPICE is a candidate metric for summarisation evaluation for LongSumm."
}Markdown (Informal)
[Investigating Metric Diversity for Evaluating Long Document Summarisation](https://preview.aclanthology.org/sigedu-bea-out-of-sync-correction/2022.sdp-1.13/) (Yang & Wan, sdp 2022)
ACL