Søren Fomsgaard
Also published as: Søren Kirkegaard Fomsgaard
2026
Mute Cods: A Multilingual Telegram Dataset with Benchmark Models for Conspiracy Theory Detection
Katarina Laken | Erik Bran Marino | Paloma Piot | Davide Bassi | Søren Kirkegaard Fomsgaard | Michele Joshua Maggini | Renata Vieira | Marcos Garcia | Sara Tonelli
Proceedings of the Fifteenth Language Resources and Evaluation Conference
Katarina Laken | Erik Bran Marino | Paloma Piot | Davide Bassi | Søren Kirkegaard Fomsgaard | Michele Joshua Maggini | Renata Vieira | Marcos Garcia | Sara Tonelli
Proceedings of the Fifteenth Language Resources and Evaluation Conference
The proliferation of conspiracy theories and hateful messages on social media poses significant challenges for content moderation and public discourse. Despite their societal impact, existing datasets for automated conspiracy detection remain limited in scope and language coverage. We present a multilingual dataset of conspiracy content on Telegram comprising 5750 messages across English, Dutch, Italian, Spanish and Portuguese from 87 channels documented as disseminating conspiracist and extremist content. Domain experts annotated messages for conspiracist tone, population replacement conspiracy theories, vaccine conspiracies, and hate speech. We extensively report on difficulties and caveats when creating and annotating this type of dataset. We establish classification baselines by evaluating six models in zero-shot fashion and fine-tuning three encoder models, achieving F1 scores up to 0.800 for conspiracist tone, 0.846 for PRCT, 0.843 for vaccine-related conspiracy theories, and 0.734 for hate speech. Inter-annotator agreement was moderate, consistent with the complexity documented in similar annotation tasks.
Combating Disinformation: Is There No Alternative?
Davide Bassi | Søren Kirkegaard Fomsgaard | Erik Bran Marino | Katarina Laken
Proceedings of the 1st Workshop on Information Disorder (InDor) @ LREC 2026
Davide Bassi | Søren Kirkegaard Fomsgaard | Erik Bran Marino | Katarina Laken
Proceedings of the 1st Workshop on Information Disorder (InDor) @ LREC 2026
This position paper critiques the dominance of detection-centered approaches in misinformation research. We argue that the prevailing paradigm treats information disorders as a content-level anomaly to be identified and suppressed, thereby obscuring the structural conditions under which different forms of information disorders emerge and resonate. Drawing on critical anthropology, we propose an alternative “clinical” model: information disorders should be understood not only as informational distortion, but as a syndrome with complex causes embedded in contexts of economic precarity, institutional distrust, and informational inequality. Treating detection as the ends rather than the means of intervention risks misguiding our efforts. Rather than positioning NLP primarily as a tool for boundary enforcement, we outline a reorientation toward structural diagnosis: diversifying data beyond WEIRD contexts, extracting socioeconomic and trust-related signals from discourse, and integrating computational outputs within interdisciplinary causal frameworks. Under this model, detection becomes a means for an epidemiology of discourse, subordinated to the broader objective of cultivating long-term epistemic resilience in our online environments.
Par-ITA: Benchmarking Seq2Seq and LLMs on a Human-Supervised Parallel Corpus for Italian Hyperpartisan Neutralization
Michele Joshua Maggini | Søren Fomsgaard | Michele Maestroni | Gaël Dias | Pablo Gamallo
Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Michele Joshua Maggini | Søren Fomsgaard | Michele Maestroni | Gaël Dias | Pablo Gamallo
Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Neutralizing hyperpartisan content is essential for mitigating online polarization, yet research has largely focused on English. We present Par-ITA, a curated subset from Semeval 2023 task 3, consisting in the first human-supervised parallel corpus for Italian hyperpartisan neutralization of 2,475 paragraph pairs. The dataset is constructed using a rigorous three-stage pipeline: (1) expert-led preliminary selection of LLMs for high-quality generation, (2) human-supervised data production with high editing rates (32–68%), and (3) post-hoc human validation. We establish extensive benchmarks for this task across seq2seq and decoder-only architectures, evaluating standard fine-tuning, Direct Preference Optimization (DPO), and in-context learning. Our analysis highlights that while DPO effectively maximizes neutrality scores in seq2seq models, automated evaluators like GPT-4o-mini exhibit systematic biases, specifically over-penalizing sensitive political topics compared to human experts. Par-ITA provides a foundational resource for non-English neutralization and a reproducible framework for developing high-quality datasets in subjective domains.
Discourse Realization of Generics in Human and LLM-generated Texts
Søren Kirkegaard Fomsgaard | Martial Pastor | Gaël Dias | Nelleke Oostdijk
Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Søren Kirkegaard Fomsgaard | Martial Pastor | Gaël Dias | Nelleke Oostdijk
Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Large Language Models (LLMs) often produce texts that appear coherent and credible, even when their factual reliability is uncertain. This paper investigates whether such perceived credibility correlates with the pervasive use of generics—generalizations without explicit quantification. We introduce a text-level genericity score derived from clause-level annotations and apply it to argumentative essays produced by humans and LLMs. To analyze how generics are realized in discourse, we employ Rhetorical Structure Theory to examine coherence relations across varying levels of genericity. Results show that according to our genericity metric, human texts are less generic than LLM-produced texts. As regards discourse, higher genericity correlates with less structured, paratactic structures, while for some models coherence is maintained through elaboration relations. Our findings suggest that some LLMs maintain well-structured coherence even in highly generic texts, which might enable them to “camouflage” argumentative texts as informative, enhancing their perceived credibility and persuasiveness.