The Extent of Repetition in Contract Language

Dan Simonson, Daniel Broderick, Jonathan Herr


Abstract
Contract language is repetitive (Anderson and Manns, 2017), but so is all language (Zipf, 1949). In this paper, we measure the extent to which contract language in English is repetitive compared with the language of other English language corpora. Contracts have much smaller vocabulary sizes compared with similarly sized non-contract corpora across multiple contract types, contain 1/5th as many hapax legomena, pattern differently on a log-log plot, use fewer pronouns, and contain sentences that are about 20% more similar to one another than in other corpora. These suggest that the study of contracts in natural language processing controls for some linguistic phenomena and allows for more in depth study of others.
Anthology ID:
W19-2203
Volume:
Proceedings of the Natural Legal Language Processing Workshop 2019
Month:
June
Year:
2019
Address:
Minneapolis, Minnesota
Venues:
NAACL | WS
SIG:
Publisher:
Association for Computational Linguistics
Note:
Pages:
21–30
Language:
URL:
https://aclanthology.org/W19-2203
DOI:
10.18653/v1/W19-2203
Bibkey:
Cite (ACL):
Dan Simonson, Daniel Broderick, and Jonathan Herr. 2019. The Extent of Repetition in Contract Language. In Proceedings of the Natural Legal Language Processing Workshop 2019, pages 21–30, Minneapolis, Minnesota. Association for Computational Linguistics.
Cite (Informal):
The Extent of Repetition in Contract Language (Simonson et al., 2019)
Copy Citation:
PDF:
https://preview.aclanthology.org/update-css-js/W19-2203.pdf