Marco Scialanga - ACL Anthology

This page is part of a temporary preview of a proposed change that may be incomplete or contain mistakes. It is not official and will be removed when the change is merged or abandoned.

Marco Scialanga

2026

Apertus: Democratizing Open and Compliant LLMs for Global Language Environments
Alejandro Hernández-Cano | Alexander Hägele | Allen Hao Huang | Angelika Romanou | Antoni-Joan Solergibert | Barna Pásztor | Bettina Messmer | Dhia Garbaya | Eduard Frank Ďurech | Ido Hakimi | Juan Garcia Giraldo | Mete Ismayilzada | Negar Foroutan | Skander Moalla | Tiancheng Chen | Vinko Sabolčec | Yixuan Xu | Michael Aerni | Badr AlKhamissi | Inés Altemir Marinas | Mohammad Hossein Amani | Matin Ansaripour | Ilia Badanin | Harold Benoit | Emanuela Boros | Nicholas John Browning | Fabian Bösch | Maximilian Böther | Niklas Canova | Camille Challier | Clément Charmillot | Jonathan Coles | Jan Milan Deriu | Arnout Devos | Lukas Drescher | Daniil Dzenhaliou | Maud Ehrmann | Dongyang Fan | Simin Fan | Silin Gao | Miguel Gila | María Grandury | Diba Hashemi | Alexander Miserlis Hoyle | Jiaming Jiang | Mark Klein | Andrei Kucharavy | Anastasiia Kucherenko | Frederike Lübeck | Roman Machacek | Theofilos Ioannis Manitaras | Andreas Marfurt | Kyle Matoba | Simon Matrenok | Henrique Mendonça | Fawzi Roberto Mohamed | Syrielle Montariol | Luca Mouchel | Sven Najem-Meyer | Jingwei Ni | Gennaro Oliva | Matteo Pagliardini | Elia Palme | Andrei Panferov | Léo Paoletti | Marco Passerini | Ivan Pavlov | Auguste Poiroux | Kaustubh Ponkshe | Nathan Ranchin | Javier Rando | Mathieu Sauser | Jakhongir Saydaliev | Mukhammadali Sayfiddinov | Marian Schneider | Stefano Schuppli | Marco Scialanga | Andrei Semenov | Kumar Shridhar | Raghav Singhal | Anna Sotnikova | Alexander Sternfeld | Ayush Kumar Tarun | Paul Teiletche | Jannis Vamvas | Xiaozhe Yao | Hao Zhao | Alexander Ilic | Ana Klimovic | Andreas Krause | Caglar Gulcehre | David Rosenthal | Elliott Ash | Florian Tramèr | Joost VandeVondele | Livio Veraldi | Martin Rajman | Thomas C. Schulthess | Torsten Hoefler | Antoine Bosselut | Martin Jaggi | Imanol Schlag
Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)

Open LLMs enable AI practitioners to control development costs by building on an existing foundation for downstream applications. While offering substantial promise, current models often fail to meet the needs of users needing open solutions aligned with responsible AI principles, including data compliance, transparency, and inclusivity. In this work, we present Apertus, a fully open suite of large language models (LLMs) designed to address responsibility shortcomings in today’s open model ecosystem, namely data responsibility and global representation. Unlike many prior models that release weights without reproducible data pipelines or regard for content-owner rights, Apertus models are pretrained exclusively on openly available data, retroactively respecting robots.txt exclusions and filtering for non-permissive, toxic, and personally identifiable content. To mitigate risks of data memorization, we also adopt the Goldfish objective during pretraining, strongly suppressing verbatim recall of data while retaining downstream task performance. Apertus also drastically expands multilingual coverage, training on 15T tokens from over approximately 1800 languages, with about 40% of pretraining data allocated to non-English content. Released at 8B and 70B scales, Apertus approaches state-of-the-art results among fully open models on multilingual benchmarks, rivaling or surpassing open-weight counterparts.

2025

SAKE: Steering Activations for Knowledge Editing
Marco Scialanga | Thibault Laugel | Vincent Grari | Marcin Detyniecki
Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)

As Large Langue Models have been shown to memorize real-world facts, the need to update this knowledge in a controlled and efficient manner arises. Designed with these constraints in mind, Knowledge Editing (KE) approaches propose to alter specific facts in pretrained models. However, they have been shown to suffer from several limitations, including their lack of contextual robustness and their failure to generalize to logical implications related to the fact. To overcome these issues, we propose SAKE, a steering activation method that models a fact to be edited as a distribution rather than a single prompt. Leveraging Optimal Transport, SAKE alters the LLM behavior over a whole fact-related distribution, defined as paraphrases and logical implications. Several numerical experiments demonstrate the effectiveness of this method: SAKE is thus able to perform more robust edits than its existing counterparts.

Co-authors

Harold Benoit 1

Emanuela Boroş 1

Antoine Bosselut 1

Nicholas John Browning 1

Fabian Bösch 1

Maximilian Böther 1

Niklas Canova 1

Camille Challier 1

Clément Charmillot 1

Tiancheng Chen 1

Jonathan Coles 1

Jan Milan Deriu 1

Marcin Detyniecki 1

Lukas Drescher 1

Daniil Dzenhaliou 1

Negar Foroutan 1

Juan Garcia Giraldo 1

María Grandury 1

Vincent Grari 1

Çağlar Gu̇lçehre 1

Alejandro Hernández-Cano 1

Torsten Hoefler 1

Alexander Miserlis Hoyle 1

Allen Hao Huang 1

Alexander Hägele 1

Alexander Ilic 1

Mete Ismayilzada 1

Jiaming Jiang 1

Andreas Krause 1

Andrei Kucharavy 1

Anastasiia Kucherenko 1

Thibault Laugel 1

Frederike Lübeck 1

Roman Machacek 1

Theofilos Ioannis Manitaras 1

Andreas Marfurt 1

Inés Altemir Marinas 1

Simon Matrenok 1

Henrique Mendonça 1

Bettina Messmer 1

Skander Moalla 1

Fawzi Roberto Mohamed 1

Syrielle Montariol 1

Sven Najem-Meyer 1

Gennaro Oliva 1

Matteo Pagliardini 1

Andrei Panferov 1

Léo Paoletti 1

Marco Passerini 1

Auguste Poiroux 1

Kaustubh Ponkshe 1

Barna Pásztor 1

Martin Rajman 1

Nathan Ranchin 1

Angelika Romanou 1

David Rosenthal 1

Vinko Sabolčec 1

Mathieu Sauser 1

Jakhongir Saydaliev 1

Mukhammadali Sayfiddinov 1

Imanol Schlag 1

Marian Schneider 1

Thomas C. Schulthess 1

Stefano Schuppli 1

Andrei Semenov 1

Kumar Shridhar 1

Raghav Singhal 1

Antoni-Joan Solergibert 1

Anna Sotnikova 1

Alexander Sternfeld 1

Ayush Kumar Tarun 1

Paul Teiletche 1

Florian Tramèr 1

Jannis Vamvas 1

Joost VandeVondele 1

Livio Veraldi 1

Eduard Frank Ďurech 1

Venues

ACL2