Verginica Barbu Mititelu
Other people with similar names: Verginica Barbu Mititelu
Unverified author pages with similar names: Verginica Barbu Mititelu
2026
PARSEME 2.0 Multilingual Corpus of Multiword Expressions
Agata Savary | Manon Scholivet | Carlos Ramisch | Takuya Nakamura | Eric Bilinski | Sara Stymne | Voula Giouli | Stella Markantonatou | Vasile Pais | Maria Mitrofan | Louis Estève | Bruno Guillaume | Verginica Barbu Mititelu | Jaka Čibej | Roberto Díaz Hernández | Victoria Fendel | Polona Gantar | Olha Kanishcheva | Cvetana Krstev | Chaya Liebeskind | Irina Lobzhanidze | Aleksandra M. Marković | Gunta Nešpore-Bērzkalne | Adriana S. Pagano | Mehrnoush Shamsfard | Ranka Stankovic | Vahide Tajalli | Carole Tiberius | Aakanksha Padhye
Proceedings of the Fifteenth Language Resources and Evaluation Conference
Agata Savary | Manon Scholivet | Carlos Ramisch | Takuya Nakamura | Eric Bilinski | Sara Stymne | Voula Giouli | Stella Markantonatou | Vasile Pais | Maria Mitrofan | Louis Estève | Bruno Guillaume | Verginica Barbu Mititelu | Jaka Čibej | Roberto Díaz Hernández | Victoria Fendel | Polona Gantar | Olha Kanishcheva | Cvetana Krstev | Chaya Liebeskind | Irina Lobzhanidze | Aleksandra M. Marković | Gunta Nešpore-Bērzkalne | Adriana S. Pagano | Mehrnoush Shamsfard | Ranka Stankovic | Vahide Tajalli | Carole Tiberius | Aakanksha Padhye
Proceedings of the Fifteenth Language Resources and Evaluation Conference
We present edition 2.0 of the PARSEME multilingual corpus annotated for multiword expressions (MWEs), resulting from efforts of the PARSEME community towards universality-driven modeling of idiomaticity. With respect to previous editions, we extend the annotation scope to all syntactic MWE categories: verbal, nominal, adjectival, adverbial and functional. We cover 17 languages, of which 7 are new. The annotation process is based on cross-lingually unified guidelines, phrased as decision diagrams over linguistic tests, and a typology of 18 MWE categories. The corpus contains almost 5 million tokens, over 250,000 sentences and 140,000 MWE annotations. The applicability of the corpus is tested in baseline experiments with a prompt-based MWE identification system. Results show that generic large language models do not encode sufficient knowledge to solve the MWE identification task.
The Romanian Corpus Annotated with Multiword Expressions. PARSEME-Ro Version 2.0
Verginica Barbu Mititelu | Mihaela Cristescu | Elena Irimia | Carmen Mîrzea Vasile
Proceedings of the Fifteenth Language Resources and Evaluation Conference
Verginica Barbu Mititelu | Mihaela Cristescu | Elena Irimia | Carmen Mîrzea Vasile
Proceedings of the Fifteenth Language Resources and Evaluation Conference
The Romanian journalistic corpus previously annotated with verbal multiword expressions (PARSEME-Ro) has been extended recently with other journalistic texts and annotated with multiword expressions of all parts of speech closely observing version 2.0 of the PARSEME guidelines. The corpus size has been increased by about 40%, it underwent automatic morpho-syntactic annotation following the Universal Dependencies principles, as well as extensive semi-automatic annotation of multiword expressions of all morphological types (nominal, adjectival, adverbial, determiner, pronominal, prepositional, conjunction, interjection, and verbal for the newly added texts). We present here our work methodology, which involves an automatic annotation phase, but the manual work prevails in checking the annotation and its consistency. We also offer quantitative data about the new version of the corpus, the types of multiword expressions existing in Romanian and occurring therein, and characteristics thereof. The new version of the PARSEME-Ro corpus contributes to the field of developing multiword expressions resources per se, i.e. describing this language phenomenon, as well as resources for training, tuning and testing the performance of tools and large language models when dealing with this linguistic phenomenon.The paper also discusses some remarks on the MWE paraphrasing subtask in which a part of the corpus was used. The corpus is released with a permissive license.
Search
Fix author
Co-authors
- Eric Bilinski 1
- Mihaela Cristescu 1
- Roberto Díaz Hernández 1
- Louis Estève 1
- Victoria Fendel 1
- Polona Gantar 1
- Voula Giouli 1
- Bruno Guillaume 1
- Elena Irimia 1
- Olha Kanishcheva 1
- Cvetana Krstev 1
- Chaya Liebeskind 1
- Irina Lobzhanidze 1
- Stella Markantonatou 1
- Aleksandra M. Marković 1
- Maria Mitrofan 1
- Takuya Nakamura 1
- Gunta Nešpore-Bērzkalne 1
- Aakanksha Padhye 1
- Adriana Silvina Pagano 1
- Vasile Pais 1
- Carlos Ramisch 1
- Agata Savary 1
- Manon Scholivet 1
- Mehrnoush Shamsfard 1
- Ranka Stankovic 1
- Sara Stymne 1
- Vahide Tajalli 1
- Carole Tiberius 1
- Carmen Mîrzea Vasile 1
- Jaka Čibej 1
Venues
- LREC2