Caterina Mauri
2026
Coconstructions in Spoken Data: UD Annotation Guidelines and First Results
Ludovica Pannitto | Kaja Dobrovoljc Zor | Sylvain Kahane | Elena Battaglia | Bruno Guillaume | Caterina Mauri | Eleonora Zucchini
Proceedings of the Ninth Workshop on Universal Dependencies (UDW 2026)
Ludovica Pannitto | Kaja Dobrovoljc Zor | Sylvain Kahane | Elena Battaglia | Bruno Guillaume | Caterina Mauri | Eleonora Zucchini
Proceedings of the Ninth Workshop on Universal Dependencies (UDW 2026)
The paper proposes annotation guidelines for syntactic dependencies that span across speaker turns — including collaborative coconstructions proper, wh-question answers, and backchannels — in spoken language treebanks within the Universal Dependencies framework. Two representations are proposed: a speaker-based representation following the segmentation into speech turns, and a dependency-based representation with dependencies across speech turns. New propositions are also put forward to distinguish between reformulations and repairs, and to promote elements in unfinished phrases.
Say Again? The Limits of Whisper with Conversation. A Case Study on the KIParla Corpus.
Martina Simonotti | Ludovica Pannitto | Caterina Mauri | Adriano Ferraresi | Gabriele Carioli
Proceedings of Speech Language Models in Low-Resource Settings: Performance, Evaluation, and Bias Analysis (SPEAKABLE) @ LREC 2026
Martina Simonotti | Ludovica Pannitto | Caterina Mauri | Adriano Ferraresi | Gabriele Carioli
Proceedings of Speech Language Models in Low-Resource Settings: Performance, Evaluation, and Bias Analysis (SPEAKABLE) @ LREC 2026
This study investigates how Whisper handles interactional phenomena in spontaneous Italian conversation, focusing on backchannels, repairs, and filled pauses. We compare standard Word Error Rate (WER) optimization with a decoding strategy that explicitly rewards the preservation of interactional events. Results show that decoding choices have limited impact on overall accuracy, while recognition remains strongly phenomenon-dependent, suggesting structural limitations in the handling of interactional phenomena, with systematic linearization of repairs and frequent suppression of short conversational items.
Is Semi-Automatic Transcription Useful in Corpus Creation? Preliminary Considerations on the KIParla Corpus
Martina Simonotti | Ludovica Pannitto | Eleonora Zucchini | Silvia Ballarè | Caterina Mauri
Proceedings of the Fifteenth Language Resources and Evaluation Conference
Martina Simonotti | Ludovica Pannitto | Eleonora Zucchini | Silvia Ballarè | Caterina Mauri
Proceedings of the Fifteenth Language Resources and Evaluation Conference
This paper analyses the implementation of Automatic Speech Recognition (ASR) into the transcription workflow of the KIParla corpus, a resource of spoken Italian. Through a two-phase experiment, 11 expert and novice transcribers produced both manual and ASR-assisted transcriptions of identical audio segments across three different types of conversation, which were subsequently analyzed through a combination of statistical modeling, word-level alignment and a series of annotation-based metrics. Results show that ASR-assisted workflows can increase transcription speed but do not systemically improve accuracy or prosodic annotation quality. Improvements appear to depend on multiple factors, including workflow configuration, conversation type and annotator experience. These findings are therefore yet not generalizable and highlight the complex interplay between transcription expertise, data type and workflow design. Despite current limitations, ASR-assisted transcription, potentially when supported by task-specific fine-tuning, could be integrated into the KIParla transcription workflow to accelerate corpus creation without compromising linguistic and annotation quality. More broadly, this work underscores the potential of semi-automatic transcription for corpus building, especially in complex settings involving multiple speakers and spontaneous, conversational data.
2025
Introducing KIParla Forest: seeds for a UD annotation of interactional syntax
Ludovica Pannitto | Eleonora Zucchini | Silvia Ballarè | Cristina Bosco | Caterina Mauri | Manuela Sanguinetti
Proceedings of the Eighth International Conference on Dependency Linguistics (Depling, SyntaxFest 2025)
Ludovica Pannitto | Eleonora Zucchini | Silvia Ballarè | Cristina Bosco | Caterina Mauri | Manuela Sanguinetti
Proceedings of the Eighth International Conference on Dependency Linguistics (Depling, SyntaxFest 2025)
The present project endeavors to enrich the linguistic resources available for Italian by introducing KIParla Forest, a treebank for the KIParla corpus - an existing and well-known resource for spoken Italian. This article contextualizes the project, describes the treebank creation process and design choices, and highlights future plans for next improvements.
2024
Did Somebody Say ‘Gest-IT’? A Pilot Exploration of Multimodal Data Management
Ludovica Pannitto | Lorenzo Albanesi | Laura Marion | Federica Martines | Carmelo Caruso | Claudia Bianchini | Francesca Masini | Caterina Mauri
Proceedings of the Tenth Italian Conference on Computational Linguistics (CLiC-it 2024)
Ludovica Pannitto | Lorenzo Albanesi | Laura Marion | Federica Martines | Carmelo Caruso | Claudia Bianchini | Francesca Masini | Caterina Mauri
Proceedings of the Tenth Italian Conference on Computational Linguistics (CLiC-it 2024)
The paper presents a pilot exploration of the construction, management and analysis of a multimodal corpus. Through athree-layer annotation that provides orthographic, prosodic, and gestural transcriptions, the gest-IT resource allows oneto investigate the variation of gesture-making patterns in conversations between sighted people and people with visualimpairment. After discussing the transcription methods and technical procedures employed in our study, we will propose aunified CoNLL-U corpus and indicate our future steps.
2019
KIParla Corpus: A New Resource for Spoken Italian
Caterina Mauri | Silvia Ballarè | Eugenio Goria | Massimo Cerruti | Francesco Suriano
Proceedings of the Sixth Italian Conference on Computational Linguistics (CLiC-it 2019)
Caterina Mauri | Silvia Ballarè | Eugenio Goria | Massimo Cerruti | Francesco Suriano
Proceedings of the Sixth Italian Conference on Computational Linguistics (CLiC-it 2019)
2013
Search
Fix author
Co-authors
- Ludovica Pannitto 5
- Silvia Ballarè 3
- Eleonora Zucchini 3
- Martina Simonotti 2
- Lorenzo Albanesi 1
- Elena Battaglia 1
- Claudia Bianchini 1
- Cristina Bosco 1
- Gabriele Carioli 1
- Carmelo Caruso 1
- Massimo Cerruti 1
- Kaja Dobrovoljc 1
- Adriano Ferraresi 1
- Eugenio Goria 1
- Bruno Guillaume 1
- Sylvain Kahane 1
- Laura Marion 1
- Federica Martines 1
- Francesca Masini 1
- Malvina Nissim 1
- Paola Pietrandrea 1
- Manuela Sanguinetti 1
- Andrea Sansò 1
- Francesco Suriano 1