Proceedings of the Ninth Workshop on Universal Dependencies (UDW 2026)
Çağrı Çöltekin, Kaja Dobrovoljc (Editors)
- Anthology ID:
- 2026.udw-1
- Month:
- May
- Year:
- 2026
- Address:
- Palma de Mallorca, Spain
- Venues:
- UDW | WS
- Events:
- Fifteenth Language Resources and Evaluation Conference | Universal Dependencies Workshop (2026) | Other Workshops and Events (2026)
- SIG:
- Publisher:
- ELRA Language Resources Association (ELRA)
- URL:
- https://preview.aclanthology.org/ingest-lrec/2026.udw-1/
- DOI:
- PDF:
- https://preview.aclanthology.org/ingest-lrec/2026.udw-1.pdf
Proceedings of the Ninth Workshop on Universal Dependencies (UDW 2026)
Çağrı Çöltekin | Kaja Dobrovoljc
Çağrı Çöltekin | Kaja Dobrovoljc
Probing the Dynamics of Syntactic Ability Acquisition Throughout LLM Pretraining
Hiroshi Matsuda | Masayuki Asahara
Hiroshi Matsuda | Masayuki Asahara
In this research, we introduce LoRA probing, a lightweight approach for observing how core syntactic abilities emerge during LLM pretraining. Leveraging OLMo-2’s public intermediate checkpoints, we trace learning curves across 24 pretraining stages on 33 Universal Dependencies languages by fine-tuning LoRA with step-by-step parsing instructions and a simple tabular output. To fit the relatively short context length of the OLMo-2, we design a compact 2-step-no-form prompt template and this matches the baseline in average accuracy while halving the context length and substantially increasing throughput, enabling efficient large-scale evaluation. Token Recall surpasses 0.9 within the first 1–2K pretraining steps, indicating that stable output formatting emerges early. Despite OLMo-2-7B’s English-centric pretraining, LAS exceeds 80 points in 29 of 33 languages; however, relations such as iobj and csubj show delayed onset and instability across many languages. LoRA probing thus provides a practical, reproducible lens on the cross-lingual dynamics of syntactic acquisition during LLM pretraining.
Which languages are "hot", and which are "cool"? Using Universal Dependencies for large-scale comparisons of subject expression
Natalia Levshina
Natalia Levshina
This study uses Universal Dependencies to investigate subject omission across fifty-six news corpora and twenty geographic varieties of English. Building on McLuhan’s "hot–cool" distinction, Hall’s LC–HC continuum, and Bisang’s notion of overt vs. hidden complexity, it tests whether subject omission rates reflect degrees of contextual reliance. The results broadly support these theories: low-context, "hot" languages, such as German, Dutch and Swedish, show low omission rates, while high-context, "cool" languages, such as Japanese, Korean and Chinese, show higher rates. English behaves as a "hot" language but exhibits internal variation across varieties, with Southeast Asian varieties exhibiting more omission than African ones. The study provides large-scale quantitative evidence while highlighting the need for further theoretical and methodological refinement, particularly regarding the role of word order and agreement.
This paper aims to give a preliminary analysis of focus marking constructions in Tigrinya within the Universal Dependency (UD) framework. We identify three types of constructions that we consider clefting. All three involve a copula placed after the focus element, the UD root, and a subordinate verb form representing the presupposed information. Tigrinya has three possibilities for this subordinate verb, which we annotate as csubj, advcl, and xcomp. In each, we use the subrelation cleft to indicate the common structure and function of the different cleft types on a par with the use of subrelations pass for passive in different languages.
Coconstructions in Spoken Data: UD Annotation Guidelines and First Results
Ludovica Pannitto | Kaja Dobrovoljc Zor | Sylvain Kahane | Elena Battaglia | Bruno Guillaume | Caterina Mauri | Eleonora Zucchini
Ludovica Pannitto | Kaja Dobrovoljc Zor | Sylvain Kahane | Elena Battaglia | Bruno Guillaume | Caterina Mauri | Eleonora Zucchini
The paper proposes annotation guidelines for syntactic dependencies that span across speaker turns — including collaborative coconstructions proper, wh-question answers, and backchannels — in spoken language treebanks within the Universal Dependencies framework. Two representations are proposed: a speaker-based representation following the segmentation into speech turns, and a dependency-based representation with dependencies across speech turns. New propositions are also put forward to distinguish between reformulations and repairs, and to promote elements in unfinished phrases.
Verifying the Menzerath-Altmann law in the verbal domain in 180 languages
Pegah Faghiri | Kim Gerdes | Sylvain Kahane
Pegah Faghiri | Kim Gerdes | Sylvain Kahane
We present a large-scale evaluation of the Menzerath-Altmann law (MAL) in the verbal domain across 180 languages, using the Universal Dependencies (UD) treebank collection (v2.17). MAL predicts that as the number of constituents of a linguistic unit increases, their average size decreases. We propose a robust metric to estimate the MAL effect across corpora of widely varying sizes and define threshold-based categories to classify languages along a MAL preference cline. Crucially, we analyse the preverbal and postverbal domains separately, in addition to the standard bilateral MAL, and control for potential sampling bias by comparing results across language families (Indo-European vs. non-Indo-European) and syntactic types (VO, OV and no dominant order). Our results confirm MAL as a typologically widespread preference but not an absolute universal: several languages display a trivial or even opposite (anti-MAL) tendency. Furthermore, we uncover a significant asymmetry between the two sides of the verb: the MAL effect is stronger in the postverbal domain, while anti-MAL is stronger in the preverbal domain. VO languages tend to show a stronger MAL preference postverbally, whereas OV languages do so preverbally. These findings challenge the widespread assumption that length-based ordering constraints apply symmetrically on both sides of the verb and contribute new cross-linguistic evidence to the debate on the interaction between dependency length minimization and constituent size.
Comparing Dependency Distances of Esperanto and Other Languages in a Multi-Lingual Parallel Corpus
Masanori Oya
Masanori Oya
This study attempts to integrate Esperanto into the research of dependency distance, based on a small-scale multi-lingual parallel corpus of Manifesto de Prago, annotated with UD relations, to contribute to the development of research on Esperanto as a natural language. The mean dependency distance and the distribution of dependency distances of Esperanto are not significantly different from the majority of the 21 languages in the corpus, thus partially supporting the claim that Esperanto is no less natural than other natural languages.
Negation of Turkic non-verbal clauses: Analysis and Universal Dependencies Implementation
Nikolett Mus | Furkan Akkurt | Bermet Chontaeva | Soudabeh Eslami | Sardana Ivanova | Çağrı Çöltekin | Jonathan N. Washington | Gulnura Dzhumalieva | Aida Kasieva
Nikolett Mus | Furkan Akkurt | Bermet Chontaeva | Soudabeh Eslami | Sardana Ivanova | Çağrı Çöltekin | Jonathan N. Washington | Gulnura Dzhumalieva | Aida Kasieva
The paper examines the grammatical behavior of the negative element used to negate predicates in non-verbal clauses in three Turkic languages: Azerbaijani, Kyrgyz, and Turkish. We focus on its interaction with verbal copulas, subject agreement, and the distribution of agreement suffixes, as well as its position within the predicate phrase. The study draws on both previously described corpus data and newly collected examples. Across all three languages, agreement features are realised on the negative element only in the absence of an overt copula. The agreement morphology involved is identical to that found with nominal, adjectival, and adverbial predicates. In all the languages examined, the negative element remains within the predicate phrase; thus, its position is syntactically constrained. At the same time, we observe differences among the languages in the degree to which the position of the negator is fixed within the predicate. In Turkish and Azerbaijani, regardless of which element of the nominal predicate it negates, the negator invariably follows the predicate. In Kyrgyz, by contrast, it consistently appears immediately after the element it negates within the predicate. These patterns suggest that the negative element behaves syntactically as a phrasal operator associated with non-verbal predicates. For annotation purposes, we therefore propose analysing the negative element as a negation modifier, assigning it the POS tag ADV and the dependency relation advmod:neg.
Towards Universal Dependencies for L2 Learners of Modern Greek: Annotation and Challenges
Christina Klironomou | Thelka Pasparaki | Arianna Masciolini | Alexandros Tantos | Despoina Ourania Touriki | Konstantinos Tsiotskas | Eleni Tsourilla
Christina Klironomou | Thelka Pasparaki | Arianna Masciolini | Alexandros Tantos | Despoina Ourania Touriki | Konstantinos Tsiotskas | Eleni Tsourilla
This paper focuses on annotating the Greek Learner Corpus in Universal Dependencies (UD). It presents the annotation process, development of guidelines and evaluation of the attempted annotation of two annotators. This work is part of a larger annotation project which aims to compile a sizeable learner treebank that can be used to promote research on second language acquisition and its automatic processing.
A Comparative Linguistic Analysis of Ottoman and Modern Turkish through UD Treebanks
Enes Yılandiloğlu
Enes Yılandiloğlu
While the linguistic shifts between Ottoman and modern Turkish are well-documented qualitatively, quantitative analyses remain scarce. This study addresses this by conducting a comparative computational analysis using two Universal Dependencies treebanks: OTA-DUDU for Ottoman Turkish and TR-BOUN for modern Turkish. By employing descriptive statistics and a log-likelihood ratio test, we demonstrate the change and quantify the magnitude of diachronic variation. The analysis yields three primary statistical findings. First, our data reveals a 77% compliance rate with labial vowel harmony for suffixes, while this value is 98% in modern Turkish. This discrepancy can be explained by the presence of rounding in Ottoman Turkish, which disappears in modern Turkish. On the other hand, the compliance rate of palatal vowel harmony is quite high for both languages, 96% for Ottoman Turkish and 99% for modern Turkish. Second, some suffixes, such as the converb -(y)Ip and the dative infinitive -mAyA, changed by reducing their allomorphs in modern Turkish. Third, we demonstrate that Arabic and Persian pluralization rules, which constituted 28% of plural nouns in Ottoman Turkish, lost their pluralizing function in modern Turkish, although the words remain with singular meaning.
The Southern Bantu language family contains languages with so-called conjunctive orthographies and disjunctive orthographies. In languages with conjunctive orthographies, such as isiZulu, orthographic words correspond to linguistic words, whereas in languages with disjunctive orthographies, prefix morphemes of verbs and other predicates are written as disjunct, orthographic words. When developing Universal Dependencies treebanks, the basic principle is to consider syntactic (linguistic) words, but for languages with agglutinating morphology, it has been argued that this reduces the informativeness of the treebank. In this paper we investigate this claim by analysing and measuring the effects of annotating universal dependencies on the basis of orthographic words on two morphosyntactically parallel treebanks for isiZulu and Sepedi.
CoBra: A Compound Branching Resource for Nominal Triconstituent Compounds in English and German
Carmen Schacht | Isabell Landwehr | Diana Davidson | Konrad Grabowski | Magdalena Meiser | Sophia Wiedmann
Carmen Schacht | Isabell Landwehr | Diana Davidson | Konrad Grabowski | Magdalena Meiser | Sophia Wiedmann
We present CoBra, a resource containing triconstituent nominal compounds in English and German. This addresses an understudied aspect of compound processing, since research and resources in psycholinguistics and NLP have mostly focused on two-constituent compounds. In addition, our resource covers both general and scientific language, allowing for a register-informed perspective on compounds. It provides syntactic and semantic annotation of compound structure, in particular of the branching direction (i.e. the internal embedding structure, the Compound Branching) and the semantic relationship between constituents. Annotations are implemented using extensions of Universal Dependencies (UD) labels. To explore applications of our new resource, we also conduct a pilot study investigating the relationship between semantic transparency and branching direction. Our results indicate that there is indeed a correlation. Overall, our resource contributes to gaining a more detailed understanding of the structure and processing of morphologically complex words within the UD framework.
SE Constructions Revisited: Focus on Treebanks for Romance Languages
Verginica Barbu Mititelu | Elena Irimia | Adriana S. Pagano | Roxana Ciolaneanu | Ioana Buhnila
Verginica Barbu Mititelu | Elena Irimia | Adriana S. Pagano | Roxana Ciolaneanu | Ioana Buhnila
We analyze the current annotation of SE constructions, i.e. verbal constructions marked by the clitic se and its cognates across five Romance languages (French, Italian, Portuguese (European and Brazilian), Romanian and Spanish) in several Universal Dependencies treebanks (version 2.17). We discuss the morphologic, syntactic and semantic characteristics of such constructions in each of the languages considered, both from a theoretical perspective and from that of existing annotation. To address inconsistencies in the data and strengthen Universal Dependencies as a scaffold for the automatic conversion of morphosyntactic annotation into semantic representations (Uniform Meaning Representation), we propose a clear distinction between argumental and non-argumental uses of the reflexive clitic, and outline systematic ways to implement this distinction in annotation guidelines. We also examine how some of the reported inconsistencies are handled in the treebanks under study and discuss the extent to which these practices can be extended to other treebanks, within the same or across different languages.
Say "No" to Missing Polarity: A Negation Enrichment of Porttinari UD Treebank
Isaac Souza de Miranda Junior | Oto Araújo Vale | Marie-Catherine de Marneffe
Isaac Souza de Miranda Junior | Oto Araújo Vale | Marie-Catherine de Marneffe
Negation is a central phenomenon in linguistics: every language has some way of expressing the difference between an affirmative sentence and a negative one (Horn and Wansing, 2025). However, the treatment of negation remains uneven in Natural Language Processing (Jimenez-Zafra et al., 2017; Jiménez-Zafra et al., 2020). This paper presents the enrichment of a Brazilian Portuguese corpus with negation-related morphological information within the Universal Dependencies (UD) framework (Nivre et al., 2020; de Marneffe et al., 2021). We enrich the Porttinari-base corpus (Duran et al., 2023) by systematically adding the UD morphological features Polarity=Neg and PronType=Neg for 18 negation-related lexical items. The enrichment only modifies the morphological features, leaving tokenization and dependency structure unchanged. To evaluate the computational results of this enrichment, we present an experiment using the Brazilian Portuguese parser PortParser (Lopes and Pardo, 2024), which we trained both on the original Porttinari-base data (Duran et al., 2023) and on our enriched version. Our results show that after enrichment, the parser’s performance remains stable, and the newly introduced features are being learned.
The Grammar Does the Work: Functional vs. Lexical Dependency Length Minimization Across the UD Languages
Kim Gerdes
Kim Gerdes
Dependency length minimization (DLM) is a well-documented processing universal, but previous studies report a single mean dependency distance (MDD) per language, obscuring variation across syntactic relation types. We analyze 122 languages in UD and SUD (version 2.17), showing that DLM operates on two distinct levels. Grammar-driven optimization targets functional dependencies (det, case, aux), which are universally short (mean 1.71, σ=0.33) and invariant across typologically diverse languages. Processing-driven optimization operates on lexical dependencies (nsubj, obj, obl), which are longer (mean 2.87), highly variable (σ=0.63), and constrained by word-order typology. This asymmetry holds in SUD despite reversed head direction (r=0.92). We conclude that "the grammar does the work" of minimization by scaffolding sentences with local functional attachments, leaving processing pressures to determine the ordering of lexical heads.
Greenberg’s Universal 45 in Universal Dependencies: Gender Distinctions and Annotation Challenges
Antoni Brosa-Rodriguez | M. Dolores Jimenez Lopez
Antoni Brosa-Rodriguez | M. Dolores Jimenez Lopez
This paper revisits and extends Greenberg’s Universal 45 on gender distinctions using Universal Dependencies 2.17, comprising 339 treebanks across 186 languages. A systematic analysis of morphosyntactic patterns confirms the implicational hierarchy (singular > plural gender marking), with 98.6 % conformity in pronominal categories. Only two potential exceptions are detected, both with minimal occurrences and likely attributable to annotation errors. Extending the analysis beyond pronouns to 13 UPOS categories shows that core categories maintain near-perfect compliance, while peripheral categories exhibit higher violation rates, primarily driven by annotation inconsistencies rather than genuine linguistic exceptions. A total of 90 treebanks display gender-number features in traditionally invariable categories (e.g., adpositions, conjunctions, adverbs), indicating annotation issues such as prepositional contraction handling, homophone merging, and erroneous feature assignment. The study establishes a replicable computational methodology for large-scale typological validation, highlighting both the potential of corpus-based approaches and key limitations, including genealogical sampling biases, annotation heterogeneity despite universal schemas, and the false sense of comparability across treebanks.
A Proposal for a More Universal Annotation of Relative Clauses in Universal Dependencies
Santiago Herrera | Sylvain Kahane
Santiago Herrera | Sylvain Kahane
This paper proposes a new encoding of relative clauses compatible with the Universal Dependencies annotation scheme for syntactic treebanks. After showing that the current guidelines are based on the main strategy of relativization in European languages, we show that it cannot be easily extended to other relativization strategies, especially head-internal relatives clauses. The criteria for the POS of relativizers are discussed. We apply our annotation to an important variety of strategies in Mandarin, Japanese, Bambara, German, Basque, Turkish, Latin, Beja, Gbaya, and Wolof, as well as to various strategies in English, including participial clauses
MesoTree: Annotated Linguistic Resources for Quantitative Comparative Linguistic Analysis and NLP in Mesoamerica
Robert Pugh | Francis Tyers | Robert Henderson
Robert Pugh | Francis Tyers | Robert Henderson
One aspect of descriptive and documentary linguistic materials that is becoming increasingly important in the information age is that they be searchable, quantifiable, and comparable. In this paper, we describe an effort to create morphosyntactically-annotated corpora for a number of under-served Mesoamerican languages using Universal Dependencies. We describe the Mesoamerican linguistic area and languages involved in the project, the training and annotation process, and give a status report on the current state of the corpora. Finally, we describe a comparitive syntax experiment and train UD parsing models on the data, demonstrating the usefulness of UD for facilitating quantitative, comparative linguistic research.
Nonprototypical Predication and Nonpredicational Clauses in Universal Dependencies
Joakim Nivre | William Croft | Andre Coneglian
Joakim Nivre | William Croft | Andre Coneglian
To assess whether the framework of Universal Dependencies (UD) is compatible with findings from linguistic typology, we need to systematically review how UD represents linguistic constructions and how it handles the range of morphosyntactic variation attested across languages. In this paper, we present such a review focusing on nonprototypical predication and nonpredicational clauses. We find that, while nonprototypical predication is generally handled well in the UD framework, nonpredicational clauses are not discussed as such in the guidelines and often have to be annotated in a way that does not reflect their special information packaging functions. We briefly discuss ways in which the UD framework could be extended in order to better capture these functions.
To assess whether the framework of Universal Dependencies (UD) is compatible with findings from linguistic typology, we need to systematically review how UD represents linguistic constructions and how it handles the range of morphosyntactic variation attested across languages. In this paper, we present the results of such a review focusing on complex predicates. We arrive at distinct findings regarding the two main types of complex predicates. The UD framework can well accommodate eventive complex predicates, particularly serial verbs, and more grammaticalized forms of complex predicates, such as voice and TAMP auxiliaries, with the exception of incorporating strategies. However, the guidelines for stative complex predicates could be revised based on the typology of morphosyntactic strategies. We briefly discuss possible ways in which UD can be extended to better capture these strategies.
Bringing Information Structure to Universal Dependencies
Nikolett Mus | Andrew Dyer | Claudia Corbetta | Sylvain Kahane
Nikolett Mus | Andrew Dyer | Claudia Corbetta | Sylvain Kahane
The aim of the paper is to present a first attempt at annotating Information Structure roles in syntactic treebanks of the Universal Dependencies collection, discussing theoretical considerations and practical methodological questions while presenting our core annotation principles. We focus on constructions in which Topic or Focus is overtly marked through (morpho)syntactic means. The proposed annotation is illustrated using examples from five languages: Wolof, Japanese, Tundra Nenets, Hungarian, and Italian.
Introducing Universal Dependencies for Sardinian: the UD ContSar Treebank
Nicoletta Puddu | Manuela Sanguinetti | Luigi Talamo
Nicoletta Puddu | Manuela Sanguinetti | Luigi Talamo
This paper introduces the first steps towards the creation of a novel resource for contemporary Sardinian within the Universal Dependencies framework. Sardinian is a Romance language spoken in Sardinia, an island belonging to the Italian Republic and located in the center of the western Mediterranean. It is a minority and endangered language, traditionally transmitted mainly orally, and characterized by a multiplicity of varieties (usually grouped into two macro-varieties Logudorese and Campidanese), all recognized as part of the Sardinian linguistic continuum. These varieties share basic morphosyntactic features, while presenting differences at the lexical level and in the realization of specific constructions. This internal variation can be particularly challenging with regard to the normalization of lemmas and the linguistic characterization of certain phenomena. The development of the treebank therefore aims to provide an annotated resource for contemporary Sardinian that takes into account the specificities of the different varieties, using Universal Dependencies to represent them within a unified theoretical framework, in order to facilitate both linguistic analysis and automatic processing. The present paper thus describes some linguistic characteristics of Sardinian and the attempts to encode them within the UD framework. Finally, we present the results of our evaluation of an NLP pipeline for Sardinian, trained on our corpus, for the Stanford Stanza parser.
In this paper we describe TEITOK’s approach to parallel aligned treebanks: a framework that supports word-level aligned, as well as multiple level text alignment (text, paragraph, sentence). It also provides various ways a visualizing aligned data, including a newly introduced parallel visualisation for dependency trees, with mouse-over aligned highlighting. The also is a parallel search function under development, that allows queries and statistics that are both tree capable and multi-level alignment capable, including word-level alignment.
Cross-Dialectal Transfer for Low-Resource Arabic: The Tunisian Arabic Dependency Treebank
Amal Aissaoui
Amal Aissaoui
This paper presents a small-scale dependency treebank for Tunisian Arabic (TADT) developed within the Universal Dependencies framework, addressing the scarcity of linguistic resources for the Arabic varieties. The approach employs domain adaptation, leveraging a machine learning model (UDPipe 1.0) trained on Algerian Arabic data to annotate 100 Tunisian Arabic social media comments, followed by manual correction. This pilot study evaluates the feasibility of using machine learning-assisted annotation to scale resource development for spoken Arabic and identifies key challenges in cross-dialectal transfer for improving annotation quality and efficiency. This work contributes to more inclusive and fair representation of Arabic linguistic varieties in academic research and NLP applications.
From Treebank Metadata to Sentence-Level Genre in Universal Dependencies: A Reproducible, Versioned Resource
Egon Stemle
Egon Stemle
We release a sentence-level genre layer for Universal Dependencies as a separate, joinable dataset, computed across UD revisions and linked back to the underlying treebanks via a release-aware composite key comprising treebank, split, sent_id, and UD release metadata. The annotations are derived rather than authoritative and are accompanied by provenance and uncertainty indicators, enabling downstream users to choose appropriate precision-coverage trade-offs and to re-run the pipeline as UD evolves. To support both parity tracking and deployment-oriented interpretation, we report results under two complementary regimes: a fixed-partition setting aligned with earlier protocols, and a language-grouped 10-fold generalisation setting that highlights cross-language heterogeneity and anchor sparsity as operational constraints. The resulting resource is intended to make genre a practical control variable for UD-based experimentation, including genre-stratified evaluation and training data selection for POS tagging and parsing, where performance varies substantially across text types. Finally, we note that reduced genre spaces aligned with recurring robustness profiles (e.g. transcribed speech versus interactional web/social text versus edited prose/news) appear pragmatically useful, but should be treated as a community coordination task implemented through explicit, versioned mapping tables.
Towards a Universal Dependency Corpus for Old Saxon (Old Low German)
Christian Chiarcos | Janine Siewert
Christian Chiarcos | Janine Siewert
Among the West Germanic languages of the first millenium C.E. (Old English, Old Low Franconian/Old Dutch, Old High German, and – although much later – Old Frisian), Old Saxon occupies a special role both linguistically – in that it represents a middle ground in the dialect spectrum between Old English at one extreme and Old High German on the other –, and in terms of material quality, in that it is attested with considerable amounts of coherent text (unlike Old Low Franconian) which is not only particularly old (unlike, especially, Old Frisian), but also original (i.e., not translated, a rarity in the attested Old English and Old High German material). It is thus a language central to the understanding of the emergence of several modern major languages, incl. English, Dutch and German, and has been studied intensely, albeit – so far – not in the context of the Universal Dependencies. This paper addresses this gap and describes the introduction of (a) a manually annotated test corpus of Old Saxon, (b) a highly reusable conversion pipeline for converting the Penn bracketing syntax of the Penn Historical Corpora (and the Old Saxon Heliand) to UD, and (c) the evaluation of the latter against the manual annotations.
Syntactic annotation is time- and resource-consuming, especially for historical and heterogeneous data. The Universal Dependencies (UD) framework provides a stable and cross-linguistically consistent annotation scheme, offering a crucial backbone for diachronic corpus studies. However, ensuring internal consistency within historical UD treebanks remains challenging due to syntactic variation and parser errors. We address this issue for Medieval and Classical French by integrating valency information into our corrections to support UD treebank maintenance. Valency frames were extracted from the Profiterole treebank (v. 2.7) and used to enrich OFrLex with structured valency information for Medieval French. Existing lexical resources such as Lefff are also exploited for Contemporary French. These valency frames are used to detect and correct inconsistencies in automatically annotated data through batch operations, thereby reinforcing UD guideline compliance and improving annotation coherence across diachronic stages. Preliminary experiments on Medieval French and exploratory annotation of Classical French data suggest that lexicon-informed error mining can reduce manual revision effort while strengthening the diachronic continuity enabled by the UD framework.
Syntax is the Key to Semantics: Combining Universal Dependencies and Abstract Meaning Representation
Johannes Heinecke
Johannes Heinecke
This paper presents a new Abstract Meaning Representation (AMR) dataset using sentences which are already translated in 21 languages and annotated in dependency syntax (Parallel Universal Dependencies, PUD) and . We annotated the English version of the 1000 sentences available in the PUD dataset. Since PUD provides syntactic annotations, this new AMR dataset (PUD-AMR) allows comparisons of syntactic phenomena and their corresponding semantics. We finally provide a first analysis of parallels between syntactic and semantic structures.
Extending Retag to Conversion Error Detection: A Case Study on SynTagRus Morphology
Andrei Movsesian | Daniil Timchenko
Andrei Movsesian | Daniil Timchenko
Linguistically annotated corpora are often converted between annotation schemes, but errors introduced during conversion can compromise their reliability. While annotation error detection is a well-studied topic, conversion error detection remains largely unexplored. We adapt the Retag method, which is traditionally used for finding annotation errors, to identify conversion errors by comparing model performance on original and converted versions of the same corpus, aligned at the token level. Applying this approach to the SynTagRus corpus converted to Universal Dependencies, we achieve high-precision detection of conversion errors in morphological annotation. Our analysis reveals systematic errors in distinction of auxiliary verbs, pronouns, numerals, and multi-word named entities, and uncovers previously undocumented annotation inconsistencies between different sections of the corpus. The method can be applied to any converted dataset for which an aligned source is available, providing an efficient way to target conversion errors for manual correction without exhaustive inspection.
Is the framework of Universal Dependencies (UD) compatible with findings from linguistic typology about constructions in the world’s languages? To address this question, we need to systematically review how UD represents these constructions, and how it handles the range of morphosyntactic variation attested across languages. In this paper, we present the results of such a review focusing on speech act constructions. We find that UD currently lack mechanisms for systematically capturing speech act constructions and briefly discuss ways in which this can be remedied.