Elvira Mercatanti
2026
Balancing FAIR and GDPR: A Governance Framework for Oral Archives
Elvira Mercatanti | Monica Monachini | Giovanni Abete | Silvia Calamai | Sergio Canazza | Alessandro Casellato | Virginia Niri | Cesarina Vecchia | Giulia Zitelli Conti | Giada Zuccolo
Proceedings of the Joint Workshop on Legal and Ethical Issues in Human Language Technologies and Computational Approaches to Language Data Pseudonymization, Anonymization, De-identification, and Data Privacy (LEGAL2026 and CALD-pseudo 2026) @ LREC 2026
Elvira Mercatanti | Monica Monachini | Giovanni Abete | Silvia Calamai | Sergio Canazza | Alessandro Casellato | Virginia Niri | Cesarina Vecchia | Giulia Zitelli Conti | Giada Zuccolo
Proceedings of the Joint Workshop on Legal and Ethical Issues in Human Language Technologies and Computational Approaches to Language Data Pseudonymization, Anonymization, De-identification, and Data Privacy (LEGAL2026 and CALD-pseudo 2026) @ LREC 2026
This paper presents a governance framework developed within the research project ROADS to support thesustainable management of oral archives, which constitute essential linguistic resources for interdisciplinary research and cultural heritage preservation. Oral archives raise complex ethical and legal challenges due to the hybrid nature of voice data, which function simultaneously as historical documents, scientific sources and biometric identifiers, thereby creating tensions between open science principles and data protection regulations. The proposed framework integrates FAIR principles (Findable, Accessible, Interoperable, Reusable) with Privacy by Design and the GDPR accountability principle through a multilayered approach. It introduces an access model that distinguishes between publicly available metadata and controlled access to identifiable audio materials, following trusted repository standards. The framework also incorporates consent management procedures and safeguards for legacy collections, enabling responsible data sharing while preserving scientific usability. More broadly, ROADS provides a transferable model to guide the transition from project-based archives to FAIR, sustainable and reusable research resources, ensuring compliance with data protection requirements and respect for the sensitivity of the documented contexts.
Integrating TEI Publication, Guided Exploration, and Vector Databases for Semantic Search in the Voci dall’Inferno Project
Angelo Mario Del Grosso | Elvira Mercatanti | Carla Congiu | Marina Riccucci
Proceedings of The Second Workshop on Holocaust Testimonies as Language Resources (HTRes)
Angelo Mario Del Grosso | Elvira Mercatanti | Carla Congiu | Marina Riccucci
Proceedings of The Second Workshop on Holocaust Testimonies as Language Resources (HTRes)
This paper presents recent advances toward an integrated framework that combines TEI-based digital publishing with embedding-based semantic search to support the preservation, exploration and analysis of Holocaust survivor testimonies. The corpus includes written and oral sources and preserves them within a XML-TEI model supported by an ODD customization that preserves provenance, structure and interpretability. A dedicated web application developed within the eXistdb platform provides guided access to the digital corpus and supports the management, visualization, and exploration of the encoded data. The project aims to investigate a specific research goal: to verify the presence of references to the Divine Comedy by Dante within Holocaust testimonies. To this end, we implement a semantic retrieval component based on SentenceTransformers’ embeddings and a vector database, enabling the discovery of both literal and non-literal Dantean passages within the testimonies. The paper presents the advances achieved toward this objective and the ethical constraints shaping access policies, resulting in a sustainable archive and a reproducible methodology for intertextual research in sensitive historical collections.
2024
The Impact of Digital Editing on the Study of Holocaust Survivors’ Testimonies in the context of Voci dall’Inferno Project
Angelo Mario Del Grosso | Marina Riccucci | Elvira Mercatanti
Proceedings of the First Workshop on Holocaust Testimonies as Language Resources (HTRes) @ LREC-COLING 2024
Angelo Mario Del Grosso | Marina Riccucci | Elvira Mercatanti
Proceedings of the First Workshop on Holocaust Testimonies as Language Resources (HTRes) @ LREC-COLING 2024
In Nazi concentration camps, approximately 20 million people perished. This included young and old, men and women, Jews, dissidents, and homosexuals. Only 10% of those deported survived. This paper introduces “Voci dall’Inferno” project, which aims to achieve two key objectives: a) Create a comprehensive digital archive: by encoding a corpus of non-literary testimonies including both written and oral sources. b) Analyze the use of Dante’s language: by identifying the presence of Dante’s lexicon and allusions. Currently, the project holds 47 testimonies, with 29 transcribed in full text and 18 encoded using the XML-TEI format. This project is propelled by a multidisciplinary and educational context with experts in humanities and computer science. The project’s findings will be disseminated through a user-friendly web application built on an XML foundation. Though currently in its prototyping phase, the application boasts several features, including a search engine for testimonies, terms, or phrases within the corpus. Additionally, a browsing interface allows users to read and listen the original testimonies, while a visualization tool enables deeper exploration of the corpus’s content. Adhering to the Text Encoding Initiative (TEI) guidelines, the project ensures a structured digital archive, aligned with the FAIR principles for data accessibility and reusability.