Aakanksha Padhye
2026
PARSEME 2.0 Multilingual Corpus of Multiword Expressions
Agata Savary | Manon Scholivet | Carlos Ramisch | Takuya Nakamura | Eric Bilinski | Sara Stymne | Voula Giouli | Stella Markantonatou | Vasile Pais | Maria Mitrofan | Louis Estève | Bruno Guillaume | Verginica Barbu Mititelu | Jaka Čibej | Roberto Díaz Hernández | Victoria Fendel | Polona Gantar | Olha Kanishcheva | Cvetana Krstev | Chaya Liebeskind | Irina Lobzhanidze | Aleksandra M. Marković | Gunta Nešpore-Bērzkalne | Adriana S. Pagano | Mehrnoush Shamsfard | Ranka Stankovic | Vahide Tajalli | Carole Tiberius | Aakanksha Padhye
Proceedings of the Fifteenth Language Resources and Evaluation Conference
Agata Savary | Manon Scholivet | Carlos Ramisch | Takuya Nakamura | Eric Bilinski | Sara Stymne | Voula Giouli | Stella Markantonatou | Vasile Pais | Maria Mitrofan | Louis Estève | Bruno Guillaume | Verginica Barbu Mititelu | Jaka Čibej | Roberto Díaz Hernández | Victoria Fendel | Polona Gantar | Olha Kanishcheva | Cvetana Krstev | Chaya Liebeskind | Irina Lobzhanidze | Aleksandra M. Marković | Gunta Nešpore-Bērzkalne | Adriana S. Pagano | Mehrnoush Shamsfard | Ranka Stankovic | Vahide Tajalli | Carole Tiberius | Aakanksha Padhye
Proceedings of the Fifteenth Language Resources and Evaluation Conference
We present edition 2.0 of the PARSEME multilingual corpus annotated for multiword expressions (MWEs), resulting from efforts of the PARSEME community towards universality-driven modeling of idiomaticity. With respect to previous editions, we extend the annotation scope to all syntactic MWE categories: verbal, nominal, adjectival, adverbial and functional. We cover 17 languages, of which 7 are new. The annotation process is based on cross-lingually unified guidelines, phrased as decision diagrams over linguistic tests, and a typology of 18 MWE categories. The corpus contains almost 5 million tokens, over 250,000 sentences and 140,000 MWE annotations. The applicability of the corpus is tested in baseline experiments with a prompt-based MWE identification system. Results show that generic large language models do not encode sufficient knowledge to solve the MWE identification task.
The Lock, Stock, and Barrel of Marathi Multiwords
Aakanksha Padhye | Ashwini Vaidya
Proceedings of the 22nd Workshop on Multiword Expressions (MWE 2026)
Aakanksha Padhye | Ashwini Vaidya
Proceedings of the 22nd Workshop on Multiword Expressions (MWE 2026)
Multiword expressions are an important area of study in linguistics and natural language processing as they represent combination of words that function as a single unit, and display properties that cannot be predicated fully from their individual components. This paper describes annotated corpora of about 3000 multiword expressions across syntactic categories in Marathi. This is the first exhaustive resource for Marathi which includes both verbal and non-verbal multiwords. In order to develop the guidelines for annotation, we have used the existing literature on the identification and classification of these expressions. Following the PARSEME 2.0 guidelines, we discuss the categories of multiwords and their behaviour in the corpus. Throughout the annotation process, we encounter variability in compositionality and syntactic realization and discuss our design decisions during annotation. Such a dataset will further our understanding of how grammatical structure can be integrated with lexically stored multiword units in Marathi.
Search
Fix author
Co-authors
- Verginica Barbu Mititelu 1
- Eric Bilinski 1
- Roberto Díaz Hernández 1
- Louis Estève 1
- Victoria Fendel 1
- Polona Gantar 1
- Voula Giouli 1
- Bruno Guillaume 1
- Olha Kanishcheva 1
- Cvetana Krstev 1
- Chaya Liebeskind 1
- Irina Lobzhanidze 1
- Stella Markantonatou 1
- Aleksandra M. Marković 1
- Maria Mitrofan 1
- Takuya Nakamura 1
- Gunta Nešpore-Bērzkalne 1
- Adriana Silvina Pagano 1
- Vasile Pais 1
- Carlos Ramisch 1
- Agata Savary 1
- Manon Scholivet 1
- Mehrnoush Shamsfard 1
- Ranka Stankovic 1
- Sara Stymne 1
- Vahide Tajalli 1
- Carole Tiberius 1
- Ashwini Vaidya 1
- Jaka Čibej 1