Modeling Preconditions in Text with a Crowd-sourced Dataset

Heeyoung Kwon; Mahnaz Koupaee; Pratyush Singh; Gargi Sawhney; Anmol Shukla; Keerthi Kumar Kallur; Nathanael Chambers; Niranjan Balasubramanian

doi:10.18653/v1/2020.findings-emnlp.340

Modeling Preconditions in Text with a Crowd-sourced Dataset

Heeyoung Kwon, Mahnaz Koupaee, Pratyush Singh, Gargi Sawhney, Anmol Shukla, Keerthi Kumar Kallur, Nathanael Chambers, Niranjan Balasubramanian

Abstract

Preconditions provide a form of logical connection between events that explains why some events occur together and information that is complementary to the more widely studied relations such as causation, temporal ordering, entailment, and discourse relations. Modeling preconditions in text has been hampered in part due to the lack of large scale labeled data grounded in text. This paper introduces PeKo, a crowd-sourced annotation of preconditions between event pairs in newswire, an order of magnitude larger than prior text annotations. To complement this new corpus, we also introduce two challenge tasks aimed at modeling preconditions: (i) Precondition Identification – a standard classification task defined over pairs of event mentions, and (ii) Precondition Generation – a generative task aimed at testing a more general ability to reason about a given event. Evaluation on both tasks shows that modeling preconditions is challenging even for today’s large language models (LM). This suggests that precondition knowledge is not easily accessible in LM-derived representations alone. Our generation results show that fine-tuning an LM on PeKo yields better conditional relations than when trained on raw text or temporally-ordered corpora.

Anthology ID:: 2020.findings-emnlp.340
Volume:: Findings of the Association for Computational Linguistics: EMNLP 2020
Month:: November
Year:: 2020
Address:: Online
Editors:: Trevor Cohn, Yulan He, Yang Liu
Venue:: Findings
SIG:
Publisher:: Association for Computational Linguistics
Note:
Pages:: 3818–3828
Language:
URL:: https://aclanthology.org/2020.findings-emnlp.340
DOI:: 10.18653/v1/2020.findings-emnlp.340
Bibkey:
Cite (ACL):: Heeyoung Kwon, Mahnaz Koupaee, Pratyush Singh, Gargi Sawhney, Anmol Shukla, Keerthi Kumar Kallur, Nathanael Chambers, and Niranjan Balasubramanian. 2020. Modeling Preconditions in Text with a Crowd-sourced Dataset. In Findings of the Association for Computational Linguistics: EMNLP 2020, pages 3818–3828, Online. Association for Computational Linguistics.
Cite (Informal):: Modeling Preconditions in Text with a Crowd-sourced Dataset (Kwon et al., Findings 2020)
Copy Citation:
PDF:: https://preview.aclanthology.org/naacl-24-ws-corrections/2020.findings-emnlp.340.pdf
Optional supplementary material:: 2020.findings-emnlp.340.OptionalSupplementaryMaterial.zip

PDF Search Optional supplementary material