Claudio Savelli
2026
MINDS at SemEval-2026-Task 13: Robust Detection of Machine-Generated Code under Distribution Shift
Giorgia Rosalia Buccelli | Antonella Coviello | Alexandra Elena Holota | Marco Scaglione | Simone Scalora | Claudio Savelli | Riccardo Coppola | Flavio Giobergia
Proceedings of the 20th International Workshop on Semantic Evaluation (2026)
Giorgia Rosalia Buccelli | Antonella Coviello | Alexandra Elena Holota | Marco Scaglione | Simone Scalora | Claudio Savelli | Riccardo Coppola | Flavio Giobergia
Proceedings of the 20th International Workshop on Semantic Evaluation (2026)
The growing use of large language models for code generation makes distinguishing machine-generated code from human-written code increasingly difficult, especially under distribution shifts in language, domain, and generator family. SemEval-2026 Task 13 targets this challenge through three subtasks: binary detection, multi-class authorship attribution, and hybrid/adversarial code detection.In this paper, we conduct an empirical study across all subtasks, comparing a variety of approaches: frozen encoder representations, feature-based classifiers, fine-tuned transformer models, post-hoc calibration, and probability-level ensembling. Our results show a consistent generalisation gap: strong in-domain validation scores substantially overestimate performance on shifted test conditions.The code is available at https://github.com/AlexandraElena-Holota/SemEval-2026-Task13.git
MINDS at SemEval-2026-Task 1: Enhancing Humor Generation through RAG and Synthetic DPO Alignment
Sina Eskandari | Seyed Amirreza Mousavi | Amirreza Rahimi | Mona Pouresmaeil | Marcello Vitaggio | Claudio Savelli | Riccardo Coppola | Flavio Giobergia
Proceedings of the 20th International Workshop on Semantic Evaluation (2026)
Sina Eskandari | Seyed Amirreza Mousavi | Amirreza Rahimi | Mona Pouresmaeil | Marcello Vitaggio | Claudio Savelli | Riccardo Coppola | Flavio Giobergia
Proceedings of the 20th International Workshop on Semantic Evaluation (2026)
Humor generation presents significant challenges due to subjectivity and the limitations of automatic metrics. In this work, we address Task 1 of SemEval 2026 (Subtask A) by evaluating three instruction-tuned models (Llama 3.1, Gemma 2, and Qwen 2.5) via a round-robin LLM judging framework. We investigate the impact of Retrieval-Augmented Generation and Direct Preference Optimization (DPO) on performance. Our results identify Llama 3.1 as the strongest baseline and demonstrate that DPO consistently improves humor quality across configurations. These findings confirm the efficacy of LLM-based judging as a practical training signal for optimizing subjective generation tasks.
MINDS at SemEval-2026 Task 9: A Multi-Paradigm Approach to Cross-Lingual Polarization Detection
Angelo Iannielli | Samuele Maroli | Marco Roberto | Stefano Sammartino | Valentino Vacirca | Claudio Savelli | Riccardo Coppola | Flavio Giobergia
Proceedings of the 20th International Workshop on Semantic Evaluation (2026)
Angelo Iannielli | Samuele Maroli | Marco Roberto | Stefano Sammartino | Valentino Vacirca | Claudio Savelli | Riccardo Coppola | Flavio Giobergia
Proceedings of the 20th International Workshop on Semantic Evaluation (2026)
Online polarization has become a central challenge in digital discourse, characterized by hostility, identity-based division, and culturally dependent expressions that vary across languages. Automatically detecting such phenomena is particularly difficult in multilingual settings, where semantic nuance and implicit rhetoric complicate cross-lingual generalization.In this context, we participate in POLAR, a shared task at SemEval 2026 on multilingual polarization detection and categorization across 22 languages. We compare three modeling paradigms: multilingual encoder fine-tuning, translation-based transfer learning, and prompting-based generative reasoning. For the multi-label categorization task, we introduce a two-stage cascaded architecture to mitigate false positives under severe class imbalance.Our results show that multilingual encoders achieve the most robust performance for binary detection, whereas reasoning-based prompting is competitive for fine-grained category classification. This comparative study highlights the strengths and limitations of each paradigm for cross-lingual polarization analysis.
MALTO at SemEval-2026 Task 13: Detecting Human, AI, and Hybrid Code via Hard Negative Mining and Curriculum-Driven Ensembles
Hüseyin Arda Arslan | Evren Ayberk Munis | Timofei Khudonogov | Mert Akgün | Murat Beşli | Ayhan Meherrem | Claudio Savelli | Flavio Giobergia
Proceedings of the 20th International Workshop on Semantic Evaluation (2026)
Hüseyin Arda Arslan | Evren Ayberk Munis | Timofei Khudonogov | Mert Akgün | Murat Beşli | Ayhan Meherrem | Claudio Savelli | Flavio Giobergia
Proceedings of the 20th International Workshop on Semantic Evaluation (2026)
The rapid advancement of Large Language Models (LLMs) has significantly impacted software engineering, posing challenges for determining the origin and authenticity of source code. This paper presents the MALTO team’s submission for SemEval-2026 Task 13, explicitly focusing on Subtask B (Authorship Attribution among 11 classes) and Subtask C (Hybrid Code Detection). To address severe class imbalance and the complex boundaries of mixed human-machine code, we propose a unified framework that leverages an ensemble of UniXcoder and CodeT5. Our approach integrates a robust Tree-sitter-based Universal Canonicalization strategy, Data Augmentation, and a novel 3-Phase Curriculum Training schedule enhanced by Hard Negative Mining. Specifically, UniXcoder’s cross-modal representations excel at distinguishing among semantically overlapping LLM families (Subtask B), whereas CodeT5’s identifier-aware architecture is superior at detecting subtle structural anomalies in hybrid and adversarial snippets (Subtask C). By aggregating these complementary strengths, our soft-voting ensemble overcomes the limitations of individual models, demonstrating strong robustness against imbalanced distributions and effectively discriminating between purely human, purely machine, hybrid, and adversarial code snippets.
FAME: Fictional Actors for Multilingual Erasure
Claudio Savelli | Moreno La Quatra | Alkis Koudounas | Flavio Giobergia
Proceedings of the Fifteenth Language Resources and Evaluation Conference
Claudio Savelli | Moreno La Quatra | Alkis Koudounas | Flavio Giobergia
Proceedings of the Fifteenth Language Resources and Evaluation Conference
Large Language Models trained on web-scale data raise concerns about privacy and the right to be forgotten. To address these issues, Machine Unlearning provides techniques to remove specific information from trained models without retraining from scratch. However, existing benchmarks for evaluating unlearning in LLMs face two major limitations: they focus only on English and support only entity-level forgetting (removing all information about a person). We introduce FAME (Fictional Actors for Multilingual Erasure), a synthetic benchmark for evaluating Machine Unlearning across five languages: English, French, German, Italian, and Spanish. FAME contains 1,000 fictional actor biographies and 20,000 question-answer pairs. Each biography includes information on 20 topics organized into structured categories (biography, career, achievements, personal information). This design enables both entity-level unlearning (i.e., forgetting entire identities) and instance-level unlearning (i.e., forgetting specific facts while retaining others). We provide two dataset splits to support these two different unlearning scenarios and enable systematic comparison of unlearning techniques across languages. Since FAME uses entirely fictional data, it ensures that the information was never encountered during model pretraining, allowing for a controlled evaluation of unlearning methods.
Confabulations from ACL Publications (CAP): A Dataset for Scientific Hallucination Detection
Federica Gamba | Aman Sinha | Timothee Mickus | Raul Vazquez | Patanjali Bhamidipati | Claudio Savelli | Ahana Chattopadhyay | Laura A. Zanella | Yash Kankanampati | Binesh Arakkal Remesh | Aryan Ashok Chandramania | Rohit Agarwal | Chuyuan Li | Ioana Buhnila | Radhika Mamidi
Proceedings of the Fifteenth Language Resources and Evaluation Conference
Federica Gamba | Aman Sinha | Timothee Mickus | Raul Vazquez | Patanjali Bhamidipati | Claudio Savelli | Ahana Chattopadhyay | Laura A. Zanella | Yash Kankanampati | Binesh Arakkal Remesh | Aryan Ashok Chandramania | Rohit Agarwal | Chuyuan Li | Ioana Buhnila | Radhika Mamidi
Proceedings of the Fifteenth Language Resources and Evaluation Conference
We introduce the CAP (Confabulations from ACL Publications) dataset, a multilingual resource for studying hallucinations in large language models (LLMs) within scientific text generation. CAP focuses on the scientific domain, where hallucinations can distort factual knowledge, as they frequently do. In this domain, however, the presence of specialized terminology, statistical reasoning, and context-dependent interpretations further exacerbates these distortions, particularly given LLMs’ lack of true comprehension, limited contextual understanding, and bias toward surface-level generalization. CAP operates in a cross-lingual setting covering five high-resource languages (English, French, Hindi, Italian, and Spanish) and four low-resource languages (Bengali, Gujarati, Malayalam, and Telugu). The dataset comprises 900 curated scientific questions and over 7,000 LLM-generated answers from 16 publicly available models, provided as question–answer pairs along with token sequences and corresponding logits. Each instance is annotated with a binary label indicating the presence of a scientific hallucination, denoted as a factuality error, and a fluency label, capturing issues in the linguistic quality or naturalness of the text. CAP is publicly released to facilitate advanced research on hallucination detection, multilingual evaluation of LLMs, and the development of more reliable scientific NLP systems.
2025
MALTO at SemEval-2025 Task 4: Dual Teachers for Unlearning Sensitive Content in LLMs
Claudio Savelli | Evren Munis | Erfan Bayat | Andrea Grieco | Flavio Giobergia
Proceedings of the 19th International Workshop on Semantic Evaluation (SemEval-2025)
Claudio Savelli | Evren Munis | Erfan Bayat | Andrea Grieco | Flavio Giobergia
Proceedings of the 19th International Workshop on Semantic Evaluation (SemEval-2025)
Large language models (LLMs) may retain and reproduce sensitive information learned during training, posing significant privacy and ethical concerns. Once detected, this personal information should be deleted from the model. A naive answer could be to retrain these models from scratch when needed. However, this solution is unfeasible given the immense computational, economic, and environmental costs required to train these models. For this reason, Machine Unlearning (MU) has risen in recent years as an emerging field of research to efficiently delete specific information from a model’s knowledge. This paper presents our solution to the “Unlearning sensitive content from Large Language Models” shared task at SemEval-2025, which challenges researchers to develop effective LLM MU techniques. We adopt a Dual-Teacher framework that leverages a Competent and an Incompetent Teacher to erase unwanted information while selectively preserving model utility. Our approach adapts established computer vision unlearning methods to the sequential nature of language models through KL divergence minimization over next-token prediction probabilities. Our experimental results demonstrate that our method outperforms the state-of-the-art techniques.
MALTO at SemEval-2025 Task 3: Detecting Hallucinations in LLMs via Uncertainty Quantification and Larger Model Validation
Claudio Savelli | Alkis Koudounas | Flavio Giobergia
Proceedings of the 19th International Workshop on Semantic Evaluation (SemEval-2025)
Claudio Savelli | Alkis Koudounas | Flavio Giobergia
Proceedings of the 19th International Workshop on Semantic Evaluation (SemEval-2025)
Large language models (LLMs) often produce {textit{hallucinations}} —factually incorrect statements that appear highly persuasive. These errors pose risks in fields like healthcare, law, and journalism. This paper presents our approach to the Mu-SHROOM shared task at SemEval 2025, which challenges researchers to detect hallucination spans in LLM outputs. We introduce a new method that combines probability-based analysis with Natural Language Inference to evaluate hallucinations at the word level. Our technique aims to better align with human judgments while working independently of the underlying model. Our experimental results demonstrate the effectiveness of this method compared to existing baselines.
2024
MALTO at SemEval-2024 Task 6: Leveraging Synthetic Data for LLM Hallucination Detection
Federico Borra | Claudio Savelli | Giacomo Rosso | Alkis Koudounas | Flavio Giobergia
Proceedings of the 18th International Workshop on Semantic Evaluation (SemEval-2024)
Federico Borra | Claudio Savelli | Giacomo Rosso | Alkis Koudounas | Flavio Giobergia
Proceedings of the 18th International Workshop on Semantic Evaluation (SemEval-2024)
In Natural Language Generation (NLG), contemporary Large Language Models (LLMs) face several challenges, such as generating fluent yet inaccurate outputs and reliance on fluency-centric metrics. This often leads to neural networks exhibiting “hallucinations.” The SHROOM challenge focuses on automatically identifying these hallucinations in the generated text. To tackle these issues, we introduce two key components, a data augmentation pipeline incorporating LLM-assisted pseudo-labelling and sentence rephrasing, and a voting ensemble from three models pre-trained on Natural Language Inference (NLI) tasks and fine-tuned on diverse datasets.
Search
Fix author
Co-authors
- Flavio Giobergia 8
- Riccardo Coppola 3
- Alkis Koudounas 3
- Rohit Agarwal 1
- Mert Akgun 1
- Binesh Arakkal Remesh 1
- Hüseyin Arslan 1
- Erfan Bayat 1
- Murat Besli 1
- Patanjali Bhamidipati 1
- Federico Borra 1
- Giorgia Rosalia Buccelli 1
- Ioana Buhnila 1
- Aryan Ashok Chandramania 1
- Ahana Chattopadhyay 1
- Antonella Coviello 1
- Sina Eskandari 1
- Federica Gamba 1
- Andrea Grieco 1
- Alexandra Elena Holota 1
- Angelo Iannielli 1
- Yash Kankanampati 1
- Timofei Khudonogov 1
- Moreno La Quatra 1
- Chuyuan Li 1
- Radhika Mamidi 1
- Samuele Maroli 1
- Ayhan Meherrem 1
- Timothee Mickus 1
- Seyed Amirreza Mousavi 1
- Evren Munis 1
- Evren Ayberk Munis 1
- Mona Pouresmaeil 1
- Amirreza Rahimi 1
- Marco Roberto 1
- Giacomo Rosso 1
- Stefano Sammartino 1
- Marco Scaglione 1
- Simone Scalora 1
- Aman Sinha 1
- Valentino Vacirca 1
- Marcello Vitaggio 1
- Raúl Vázquez 1
- Laura A. Zanella 1