Learning Long-Document Embeddings via Chunk–Context Entailment

Waheed Ahmed Abro, Naïm Es-Sebbani, Zied Bouraoui


Abstract
Learning faithful embeddings for long documents remains challenging, especially in domains like law and medicine where inputs are long, structured, and semantically heterogeneous. We introduce the Chunk Prediction Encoder (CPE), a self-supervised framework that treats chunk–context compatibility as an unsupervised NLI problem. Given a document, CPE masks a chunk and learns (i) a contrastive objective that aligns the masked document with its held-out chunk against in-batch negatives, and (ii) a binary entailment head that predicts whether a candidate chunk belongs to the document. This joint objective encourages both geometric smoothness and directional semantic consistency, yielding robust document-level embeddings. We evaluate CPE with hierarchical and sparse-attention backbones on five benchmarks spanning legal and biomedical domains under frozen-embedding and end-to-end fine-tuning protocols. CPE consistently outperforms baselines, and is more compute-efficient than prompt-only LLM baselines under matched token budgets. Ablations demonstrate the effect of chunk length, the contrastive-vs-entailment balance, and skimming strategies.
Anthology ID:
2026.lrec-1.586
Volume:
Proceedings of the Fifteenth Language Resources and Evaluation Conference
Month:
May
Year:
2026
Address:
Palma de Mallorca, Spain
Editors:
Stelios Piperidis, Núria Bel, Henk van den Heuvel, Nancy Ide, Simon Krek, Antonio Toral
Venue:
LREC
SIG:
Publisher:
ELRA Language Resource Association
Note:
Pages:
7405–7414
Language:
External URL:
https://lrec.elra.info/lrec2026-main-586
DOI:
10.63317/4iz34o26i4tt
Bibkey:
Cite (ACL):
Waheed Ahmed Abro, Naïm Es-Sebbani, and Zied Bouraoui. 2026. Learning Long-Document Embeddings via Chunk–Context Entailment. In Proceedings of the Fifteenth Language Resources and Evaluation Conference, pages 7405–7414, Palma de Mallorca, Spain. ELRA Language Resource Association.
Cite (Informal):
Learning Long-Document Embeddings via Chunk–Context Entailment (Abro et al., LREC 2026)
Copy Citation: