Abstract
Contract language is repetitive (Anderson and Manns, 2017), but so is all language (Zipf, 1949). In this paper, we measure the extent to which contract language in English is repetitive compared with the language of other English language corpora. Contracts have much smaller vocabulary sizes compared with similarly sized non-contract corpora across multiple contract types, contain 1/5th as many hapax legomena, pattern differently on a log-log plot, use fewer pronouns, and contain sentences that are about 20% more similar to one another than in other corpora. These suggest that the study of contracts in natural language processing controls for some linguistic phenomena and allows for more in depth study of others.- Anthology ID:
- W19-2203
- Volume:
- Proceedings of the Natural Legal Language Processing Workshop 2019
- Month:
- June
- Year:
- 2019
- Address:
- Minneapolis, Minnesota
- Editors:
- Nikolaos Aletras, Elliott Ash, Leslie Barrett, Daniel Chen, Adam Meyers, Daniel Preotiuc-Pietro, David Rosenberg, Amanda Stent
- Venue:
- NAACL
- SIG:
- Publisher:
- Association for Computational Linguistics
- Note:
- Pages:
- 21–30
- Language:
- URL:
- https://aclanthology.org/W19-2203
- DOI:
- 10.18653/v1/W19-2203
- Cite (ACL):
- Dan Simonson, Daniel Broderick, and Jonathan Herr. 2019. The Extent of Repetition in Contract Language. In Proceedings of the Natural Legal Language Processing Workshop 2019, pages 21–30, Minneapolis, Minnesota. Association for Computational Linguistics.
- Cite (Informal):
- The Extent of Repetition in Contract Language (Simonson et al., NAACL 2019)
- PDF:
- https://preview.aclanthology.org/emnlp22-frontmatter/W19-2203.pdf