Hazem Amamou
2026
Uncovering Ideological Bias in RAG with Lexical Multidimensional Analysis: A Case Study on COVID-19
Elmira Salari | Maria Claudia Nunes Delfino | Hazem Amamou | José Victor de Souza | Shruti Kshirsagar | Alan Davoust | Anderson Avila
Proceedings of the 15th Joint Conference on Lexical and Computational Semantics (*SEM 2026)
Elmira Salari | Maria Claudia Nunes Delfino | Hazem Amamou | José Victor de Souza | Shruti Kshirsagar | Alan Davoust | Anderson Avila
Proceedings of the 15th Joint Conference on Lexical and Computational Semantics (*SEM 2026)
This paper studies the impact of retrieved ideologically framed texts on the outputs of large language models (LLMs). While interest in understanding ideological framing in LLMs has recently increased, little attention has been given to this issue in the context of Retrieval-Augmented Generation (RAG). To fill this gap, we design an external knowledge source based on ideologically framed texts about COVID-19 treatments. Our corpus is based on 1,117 academic articles representing discourses about controversial and endorsed treatments for the disease. We propose a corpus linguistics framework, based on Lexical Multidimensional Analysis (LMDA), to identify discourse dimensions within the corpus. LLMs are tasked to answer questions derived from three identified discourse dimensions, and two types of contextual prompts are adopted: the first comprises the user question and ideologically framed texts; and the second contains the question, ideologically framed texts, and LMDA descriptions. Alignment between reference ideologically framed texts and LLMs’ responses is assessed using cosine similarity for lexical and semantic representations. Results demonstrate that retrieved ideologically framed texts influence LLM responses toward the discourse framing represented in the external knowledge, with enhanced prompts further amplifying this effect. Our findings highlight the importance of identifying ideological framings within the RAG framework in order to mitigate not just unintended ideological bias, but also the risks of intentional discourse steering of such models.
Infox-QC: A Quebec-Focused French Corpus for Misinformation Detection and AI Robustness Assessment
Moetaz Doghmane | Hazem Amamou | Thiziri Sefsaf | Alan Davoust | Anderson Raymundo Avila
Proceedings of the Fifteenth Language Resources and Evaluation Conference
Moetaz Doghmane | Hazem Amamou | Thiziri Sefsaf | Alan Davoust | Anderson Raymundo Avila
Proceedings of the Fifteenth Language Resources and Evaluation Conference
The pervasive spread of online misinformation, often through social media and political campaigns, makes detecting false claims a crucial task for mitigating societal risks. While the vast majority of fake news datasets are developed in English, a critical gap remains for low-resource languages, such as French. To address this, we introduce Infox-QC, a novel French-language corpus focused on misinformation relevant to the Quebec region. Beyond containing real true and fake news, Infox-QC includes two unique subsets of AI-generated fake news: one created by prompting an AI to paraphrase existing fake news, and a second generated by prompting an AI to fabricate fake news from real true reports. This innovative approach allows us to verify the robustness of detection systems against fabricated content, which modern LLMs can generate with convincing efficacy. We establish comprehensive baselines using traditional machine learning methods, BERT-based models, and Large Language Models, both with and without Retrieval-Augmented Generation (RAG). Our results demonstrate that RAG-augmented LLMs offer the strongest contextual understanding, while traditional models provide valuable interpretable baselines. We further provide an exploratory human–LLM thematic agreement analysis to assess annotation consistency. The Infox-QC resource fills a critical void in French-language NLP research, supporting future efforts to explore the regional and cultural dimensions of misinformation through cross-linguistic comparison.