Sergi Àlvarez Vidal
Also published as: Sergi Alvarez-Vidal, Sergi Alvarez Vidal
2026
A Comparative Study in Corpus Linguistics Applied to Automatic Terminology Extraction
Mercè Vàzquez | Sergi Alvarez-Vidal | Antoni Oliver
Proceedings of the 19th Workshop on Building and Using Comparable Corpora (BUCC)
Mercè Vàzquez | Sergi Alvarez-Vidal | Antoni Oliver
Proceedings of the 19th Workshop on Building and Using Comparable Corpora (BUCC)
Parallel and comparable corpora are the main linguistic resources to identify multilingual terminology using automatic term extraction tools. However, parallel corpora are available only for certain languages, domains and genres, and comparable corpora have some limitations when identifying corresponding terms. To implement a more efficient selection of multilingual terminology, we compared the performance of using specialised parallel and comparable corpora applied to languages with various forms of capital in linguistic resources. This paper presents a comparative study in corpus linguistics in which we automatically identify terms in Catalan, Spanish and English in legislation and administrative law using parallel corpora, comparable corpora and a combined methodology based on both typologies of corpora together with word embeddings. We observe that the combined methodology implemented obtains a higher number of term candidates than when working exclusively with parallel or comparable corpora. The evaluation of the results is performed using a terminological thesaurus as a gold standard. The new methodology presented in our study permits us to identify multilingual terminology in an efficient way, especially in Catalan-Spanish languages.
2025
Using Translation Techniques to Characterize MT Outputs
Sergi Alvarez-Vidal | Maria Do Campo | Christian Olalla-Soler | Pilar Sánchez-Gijón
Proceedings of Machine Translation Summit XX: Volume 1
Sergi Alvarez-Vidal | Maria Do Campo | Christian Olalla-Soler | Pilar Sánchez-Gijón
Proceedings of Machine Translation Summit XX: Volume 1
While current NMT and GPT models improve fluency and context awareness, they struggle with creative texts, where figurative language and stylistic choices are crucial. Current evaluation methods fail to capture these nuances, which requires a more descriptive approach. We propose a taxonomy based on translation techniques to assess machine-generated translations more comprehensively. The pilot study we conducted comparing human machine-produced translations reveals that human translations employ a wider range of techniques, enhancing naturalness and cultural adaptation. NMT and GPT models, even with prompting, tend to simplify content and introduce accuracy errors. Our findings highlight the need for refined frameworks that consider stylistic and contextual accuracy, ultimately bridging the gap between human and machine translation performance.
Fine-tuning and evaluation of NMT models for literary texts using RomCro v.2.0
Bojana Mikelenić | Antoni Oliver | Sergi Àlvarez Vidal
Proceedings of the Second Workshop on Creative-text Translation and Technology (CTT)
Bojana Mikelenić | Antoni Oliver | Sergi Àlvarez Vidal
Proceedings of the Second Workshop on Creative-text Translation and Technology (CTT)
This paper explores the fine-tuning and evaluation of neural machine translation (NMT) models for literary texts using RomCro v.2.0, an expanded multilingual and multidirectional parallel corpus. RomCro v.2.0 is based on RomCro v.1.0, but includes additional literary works, as well as texts in Catalan, making it a valuable resource for improving MT in underrepresented language pairs. Given the challenges of literary translation, where style, narrative voice, and cultural nuances must be preserved, fine-tuning on high-quality domain-specific data is essential for enhancing MT performance. We fine-tune existing NMT models with RomCro v.2.0 and evaluate their performance for six different language combinations using automatic metrics and for Spanish-Croatian and French-Catalan using manual evaluation. Results indicate that fine-tuned models outperform general-purpose systems, achieving greater fluency and stylistic coherence. These findings support the effectiveness of corpus-driven fine-tuning for literary translation and highlight the importance of curated high-quality corpus.
2024
Training an NMT system for legal texts of a low-resource language variety South Tyrolean German - Italian
Antoni Oliver | Sergi Alvarez-Vidal | Egon Stemle | Elena Chiocchetti
Proceedings of the 25th Annual Conference of the European Association for Machine Translation (Volume 1)
Antoni Oliver | Sergi Alvarez-Vidal | Egon Stemle | Elena Chiocchetti
Proceedings of the 25th Annual Conference of the European Association for Machine Translation (Volume 1)
This paper illustrates the process of training and evaluating NMT systems for a language pair that includes a low-resource language variety.A parallel corpus of legal texts for Italian and South Tyrolean German has been compiled, with South Tyrolean German being the low-resourced language variety. As the size of the compiled corpus is insufficient for the training, we have combined the corpus with several parallel corpora using data weighting at sentence level. We then performed an evaluation of each combination and of two popular commercial systems.
LitPC: A set of tools for building parallel corporafrom literary works
Antoni Oliver | Sergi Alvarez-Vidal
Proceedings of the 1st Workshop on Creative-text Translation and Technology
Antoni Oliver | Sergi Alvarez-Vidal
Proceedings of the 1st Workshop on Creative-text Translation and Technology
In this paper, we describe the LitPC toolkit, a variety of tools and methods designed for the quick and effective creation of parallel corpora derived from literary works. This toolkit can be a useful resource due to the scarcity of curated parallel texts for this domain. We also feature a case study describing the creation of a Russian-English parallel corpus based on the literary works by Leo Tolstoy. Furthermore, an augmented version of this corpus is used to both train and assess neural machine translation systems specifically adapted to the author’s style.
2023
Proceedings of the 24th Annual Conference of the European Association for Machine Translation
Mary Nurminen | Judith Brenner | Maarit Koponen | Sirkku Latomaa | Mikhail Mikhailov | Frederike Schierl | Tharindu Ranasinghe | Eva Vanmassenhove | Sergi Alvarez Vidal | Nora Aranberri | Mara Nunziatini | Carla Parra Escartín | Mikel Forcada | Maja Popovic | Carolina Scarton | Helena Moniz
Proceedings of the 24th Annual Conference of the European Association for Machine Translation
Mary Nurminen | Judith Brenner | Maarit Koponen | Sirkku Latomaa | Mikhail Mikhailov | Frederike Schierl | Tharindu Ranasinghe | Eva Vanmassenhove | Sergi Alvarez Vidal | Nora Aranberri | Mara Nunziatini | Carla Parra Escartín | Mikel Forcada | Maja Popovic | Carolina Scarton | Helena Moniz
Proceedings of the 24th Annual Conference of the European Association for Machine Translation
Search
Fix author
Co-authors
- Antoni Oliver 4
- Nora Aranberri 1
- Judith Brenner 1
- Maria Do Campo 1
- Elena Chiocchetti 1
- Mikel L. Forcada 1
- Maarit Koponen 1
- Sirkku Latomaa 1
- Bojana Mikelenić 1
- Mikhail Mikhailov 1
- Helena Moniz 1
- Mara Nunziatini 1
- Mary Nurminen 1
- Christian Olalla-Soler 1
- Carla Parra Escartín 1
- Maja Popović 1
- Tharindu Ranasinghe 1
- Carolina Scarton 1
- Frederike Schierl 1
- Egon Stemle 1
- Pilar Sánchez-Gijón 1
- Eva Vanmassenhove 1
- Mercè Vàzquez 1