Anny D. Alvarez Nogales


2026

This work addresses the need for linguistic resources that enable language models to understand and adapt to subjective and abstract concepts in the domain of moral values within texts. In light of the growing interest in the study of moral values and its limited exploration in Spanish-speaking contexts, this work addresses this gap by developing a novel Spanish-language corpus. Furthermore, the corpus’s development process ensures that the annotations capture a wide range of perspectives, resulting in a resource that reflects the diversity of moral interpretations in real-world contexts. Specifically, there are two main contributions. 1 The creation of the first large-scale Spanish corpus annotated according to Moral Foundations Theory. 2 We introduce an experimental framework that investigates how annotators’ religious orientations could shape moral annotation patterns and propagate to model behavior. To do so, we employ a prompt-based alignment method that improves moral detection regardless of religious alignment for which the model was trained. In this scenario, we explore whether language models can align moral interpretations across divergent belief orientations.

2024

Moral values significantly define decision-making processes, notably on contentious issues like global warming. The Moral Foundations Theory (MFT) delineates morality and aims to reconcile moral expressions across cultures, yet different interpretations arise, posing challenges for computational modeling. This paper addresses the need to incorporate diverse moral perspectives into the learning systems used to estimate morality in text. To do so, it explores how training language models with varied annotator perspectives affects the performance of the learners. Building on top if this, this work also proposes an ensemble method that exploits the diverse perspectives of annotators to construct a more robust moral estimation model. Additionally, we investigate the automated identification of texts that pose annotation challenges, enhancing the understanding of linguistic cues towards annotator disagreement. To evaluate the proposed models we use the Moral Foundations Twitter Corpus (MFTC), a resource that is currently the reference for modeling moral values in computational social sciences. We observe that incorporating the diverse perspectives of annotators into an ensemble model benefits the learning process, showing large improvements in the classification performance. Finally, the results also indicate that instances that convey strong moral meaning are more challenging to annotate.