Marta Vázquez Abuín
2025
WiC Evaluation in Galician and Spanish: Effects of Dataset Quality and Composition
Marta Vázquez Abuín
|
Marcos Garcia
Proceedings of the 14th Joint Conference on Lexical and Computational Semantics (*SEM 2025)
This work explores the impact of dataset quality and composition on Word-in-Context performance for Galician and Spanish. We assess existing datasets, validate their test sets, and create new manually constructed evaluation data. Across five experiments with controlled variations in training and test data, we find that while the validation of test data tends to yield better model performance, evaluations on manually created datasets suggest that contextual embeddings are not sufficient on their own to reliably capture word meaning variation. Regarding training data, our results suggest that performance is influenced not only by size and human validation but also by deeper factors related to the semantic properties of the datasets. All new resources will be freely released.
2024
Nós-TTS: aWeb User Interface for Galician Text-to-Speech
Carmen Magariños
|
Alp Öktem
|
Antonio Moscoso Sánchez
|
Marta Vázquez Abuín
|
Noelia García Díaz
|
Adina Ioana Vladu
|
Elisa Fernández Rei
|
María Baqueiro Vidal
Proceedings of the 16th International Conference on Computational Processing of Portuguese - Vol. 2