DecOp: A Multilingual and Multi-domain Corpus For Detecting Deception In Typed Text

Pasquale Capuozzo; Ivano Lauriola; Carlo Strapparava; Fabio Aiolli; Giuseppe Sartori

DecOp: A Multilingual and Multi-domain Corpus For Detecting Deception In Typed Text

Pasquale Capuozzo, Ivano Lauriola, Carlo Strapparava, Fabio Aiolli, Giuseppe Sartori

Abstract

In recent years, the increasing interest in the development of automatic approaches for unmasking deception in online sources led to promising results. Nonetheless, among the others, two major issues remain still unsolved: the stability of classifiers performances across different domains and languages. Tackling these issues is challenging since labelled corpora involving multiple domains and compiled in more than one language are few in the scientific literature. For filling this gap, in this paper we introduce DecOp (Deceptive Opinions), a new language resource developed for automatic deception detection in cross-domain and cross-language scenarios. DecOp is composed of 5000 examples of both truthful and deceitful first-person opinions balanced both across five different domains and two languages and, to the best of our knowledge, is the largest corpus allowing cross-domain and cross-language comparisons in deceit detection tasks. In this paper, we describe the collection procedure of the DecOp corpus and his main characteristics. Moreover, the human performance on the DecOp test-set and preliminary experiments by means of machine learning models based on Transformer architecture are shown.

Anthology ID:: 2020.lrec-1.178
Volume:: Proceedings of the 12th Language Resources and Evaluation Conference
Month:: May
Year:: 2020
Address:: Marseille, France
Venue:: LREC
SIG:
Publisher:: European Language Resources Association
Note:
Pages:: 1423–1430
Language:: English
URL:: https://aclanthology.org/2020.lrec-1.178
DOI:
Bibkey:
Cite (ACL):: Pasquale Capuozzo, Ivano Lauriola, Carlo Strapparava, Fabio Aiolli, and Giuseppe Sartori. 2020. DecOp: A Multilingual and Multi-domain Corpus For Detecting Deception In Typed Text. In Proceedings of the 12th Language Resources and Evaluation Conference, pages 1423–1430, Marseille, France. European Language Resources Association.
Cite (Informal):: DecOp: A Multilingual and Multi-domain Corpus For Detecting Deception In Typed Text (Capuozzo et al., LREC 2020)
Copy Citation:
PDF:: https://preview.aclanthology.org/update-css-js/2020.lrec-1.178.pdf

PDF Cite Search