Comprehensiveness Metrics for Automatic Evaluation of Factual Recall in Text Generation

Adam Dejl; James Barry; Alessandra Pascale; Javier Carnerero-Cano

Comprehensiveness Metrics for Automatic Evaluation of Factual Recall in Text Generation

Adam Dejl, James Barry, Alessandra Pascale, Javier Carnerero-Cano

Abstract

Despite demonstrating remarkable performance across a wide range of tasks, large language models (LLMs) have also been found to frequently produce outputs that are incomplete or selectively omit key information. In sensitive domains, such omissions can result in significant harm comparable to that posed by factual inaccuracies, including hallucinations. In this study, we address the challenge of evaluating the comprehensiveness of LLM-generated texts, focusing on the detection of missing information or underrepresented viewpoints. We investigate three automated evaluation metrics: (1) an NLI-based method that decomposes texts into atomic statements and uses natural language inference (NLI) to identify missing facts, (2) a Q A-based metric that extracts question-answer pairs and compares responses across sources, and (3) an end-to-end approach that directly identifies missing content using LLMs. Our experiments demonstrate the surprising effectiveness of the simple end-to-end metric compared to more complex metrics, though at the cost of reduced robustness, interpretability and result granularity. We further assess the comprehensiveness of responses from several popular open-weight LLMs when answering user queries based on multiple sources.

Anthology ID:: 2026.findings-acl.1744
Volume:: Findings of the Association for Computational Linguistics: ACL 2026
Month:: July
Year:: 2026
Address:: San Diego, California, United States
Editors:: Maria Liakata, Viviane P. Moreira, Jiajun Zhang, David Jurgens
Venue:: Findings
SIG:
Publisher:: Association for Computational Linguistics
Note:
Pages:: 34931–34966
Language:
URL:: https://preview.aclanthology.org/ingest-acl/2026.findings-acl.1744/
DOI:
Bibkey:
Cite (ACL):: Adam Dejl, James Barry, Alessandra Pascale, and Javier Carnerero-Cano. 2026. Comprehensiveness Metrics for Automatic Evaluation of Factual Recall in Text Generation. In Findings of the Association for Computational Linguistics: ACL 2026, pages 34931–34966, San Diego, California, United States. Association for Computational Linguistics.
Cite (Informal):: Comprehensiveness Metrics for Automatic Evaluation of Factual Recall in Text Generation (Dejl et al., Findings 2026)
Copy Citation:
PDF:: https://preview.aclanthology.org/ingest-acl/2026.findings-acl.1744.pdf
Checklist:: 2026.findings-acl.1744.checklist.pdf

PDF Cite Search Checklist Fix data