@inproceedings{thomson-reiter-2022-accuracy,
    title = "The Accuracy Evaluation Shared Task as a Retrospective Reproduction Study",
    author = "Thomson, Craig  and
      Reiter, Ehud",
    editor = "Shaikh, Samira  and
      Ferreira, Thiago  and
      Stent, Amanda",
    booktitle = "Proceedings of the 15th International Conference on Natural Language Generation: Generation Challenges",
    month = jul,
    year = "2022",
    address = "Waterville, Maine, USA and virtual meeting",
    publisher = "Association for Computational Linguistics",
    url = "https://preview.aclanthology.org/ingest-emnlp/2022.inlg-genchal.11/",
    pages = "71--79",
    abstract = "We investigate the data collected for the Accuracy Evaluation Shared Task as a retrospective reproduction study. The shared task was based upon errors found by human annotation of computer generated summaries of basketball games. Annotation was performed in three separate stages, with texts taken from the same three systems and checked for errors by the same three annotators. We show that the mean count of errors was consistent at the highest level for each experiment, with increased variance when looking at per-system and/or per-error- type breakdowns."
}Markdown (Informal)
[The Accuracy Evaluation Shared Task as a Retrospective Reproduction Study](https://preview.aclanthology.org/ingest-emnlp/2022.inlg-genchal.11/) (Thomson & Reiter, INLG 2022)
ACL