Adapting Open Domain Fact Extraction and Verification to COVID-FACT through In-Domain Language Modeling

Zhenghao Liu; Chenyan Xiong; Zhuyun Dai; Si Sun; Maosong Sun; Zhiyuan Liu

doi:10.18653/v1/2020.findings-emnlp.216

Adapting Open Domain Fact Extraction and Verification to COVID-FACT through In-Domain Language Modeling

Zhenghao Liu, Chenyan Xiong, Zhuyun Dai, Si Sun, Maosong Sun, Zhiyuan Liu

Abstract

With the epidemic of COVID-19, verifying the scientifically false online information, such as fake news and maliciously fabricated statements, has become crucial. However, the lack of training data in the scientific domain limits the performance of fact verification models. This paper proposes an in-domain language modeling method for fact extraction and verification systems. We come up with SciKGAT to combine the advantages of open-domain literature search, state-of-the-art fact verification systems and in-domain medical knowledge through language modeling. Our experiments on SCIFACT, a dataset of expert-written scientific fact verification, show that SciKGAT achieves 30% absolute improvement on precision. Our analyses show that such improvement thrives from our in-domain language model by picking up more related evidence pieces and accurate fact verification. Our codes and data are released via Github.

Anthology ID:: 2020.findings-emnlp.216
Volume:: Findings of the Association for Computational Linguistics: EMNLP 2020
Month:: November
Year:: 2020
Address:: Online
Editors:: Trevor Cohn, Yulan He, Yang Liu
Venue:: Findings
SIG:
Publisher:: Association for Computational Linguistics
Note:
Pages:: 2395–2400
Language:
URL:: https://aclanthology.org/2020.findings-emnlp.216
DOI:: 10.18653/v1/2020.findings-emnlp.216
Bibkey:
Cite (ACL):: Zhenghao Liu, Chenyan Xiong, Zhuyun Dai, Si Sun, Maosong Sun, and Zhiyuan Liu. 2020. Adapting Open Domain Fact Extraction and Verification to COVID-FACT through In-Domain Language Modeling. In Findings of the Association for Computational Linguistics: EMNLP 2020, pages 2395–2400, Online. Association for Computational Linguistics.
Cite (Informal):: Adapting Open Domain Fact Extraction and Verification to COVID-FACT through In-Domain Language Modeling (Liu et al., Findings 2020)
Copy Citation:
PDF:: https://preview.aclanthology.org/nschneid-patch-5/2020.findings-emnlp.216.pdf
Code: thunlp/KernelGAT
Data: FEVER, SciFact

PDF Search Code