Leveraging Collection-Wide Similarities for Unsupervised Document Structure Extraction

Gili Lior; Yoav Goldberg; Gabriel Stanovsky

doi:10.18653/v1/2024.findings-acl.568

Leveraging Collection-Wide Similarities for Unsupervised Document Structure Extraction

Gili Lior, Yoav Goldberg, Gabriel Stanovsky

Abstract

Document collections of various domains, e.g., legal, medical, or financial, often share some underlying collection-wide structure, which captures information that can aid both human users and structure-aware models.We propose to identify the typical structure of document within a collection, which requires to capture recurring topics across the collection, while abstracting over arbitrary header paraphrases, and ground each topic to respective document locations. These requirements pose several challenges: headers that mark recurring topics frequently differ in phrasing, certain section headers are unique to individual documents and do not reflect the typical structure, and the order of topics can vary between documents. Subsequently, we develop an unsupervised graph-based method which leverages both inter- and intra-document similarities, to extract the underlying collection-wide structure. Our evaluations on three diverse domains in both English and Hebrew indicate that our method extracts meaningful collection-wide structure, and we hope that future work will leverage our method for multi-document applications and structure-aware models.

Anthology ID:: 2024.findings-acl.568
Volume:: Findings of the Association for Computational Linguistics: ACL 2024
Month:: August
Year:: 2024
Address:: Bangkok, Thailand
Editors:: Lun-Wei Ku, Andre Martins, Vivek Srikumar
Venue:: Findings
SIG:
Publisher:: Association for Computational Linguistics
Note:
Pages:: 9538–9550
Language:
URL:: https://preview.aclanthology.org/build-pipeline-with-new-library/2024.findings-acl.568/
DOI:: 10.18653/v1/2024.findings-acl.568
Bibkey:
Cite (ACL):: Gili Lior, Yoav Goldberg, and Gabriel Stanovsky. 2024. Leveraging Collection-Wide Similarities for Unsupervised Document Structure Extraction. In Findings of the Association for Computational Linguistics: ACL 2024, pages 9538–9550, Bangkok, Thailand. Association for Computational Linguistics.
Cite (Informal):: Leveraging Collection-Wide Similarities for Unsupervised Document Structure Extraction (Lior et al., Findings 2024)
Copy Citation:
PDF:: https://preview.aclanthology.org/build-pipeline-with-new-library/2024.findings-acl.568.pdf

PDF Search