Peter Uhrig
2026
A Linguistic Ontology for Constructicography: The Research Constructicon and its Ontology Modules
Elodie Winckel | Peter Uhrig | Stephanie Evert
Proceedings of 10th Workshop on Linked Data in Linguistics (LDL-2026)
Elodie Winckel | Peter Uhrig | Stephanie Evert
Proceedings of 10th Workshop on Linked Data in Linguistics (LDL-2026)
This paper introduces the Research Constructicon (RCxn), a project developed within the Research Training Group Dimensions of Constructional Space. The training group finances PhD projects in the framework of Construction Grammar (CxG), which views language as a network of form-meaning pairings. The RCxn is designed as a dynamic, community-driven resource that documents linguistic constructions while also capturing the research processes and findings associated with them. The project addresses three core dimensions: (1) the development of a modular ontology to represent constructions, their relationships, and the research surrounding them; (2) the implementation of database populated by researchers’ contributions; and (3) the creation of a web application to visualize and interact with the data. This paper focuses on our work to implement a rich ontology for the RCxn, which has to accommodate diverse research needs, from cross-linguistic comparisons to multimodal analyses, while ensuring flexibility and interoperability. We detail the modular design of the ontology, its alignment with semantic web standards (RDF/OWL), and the integration of existing ontologies (e.g., OLiA, FOAF). The RCxn’s development is iterative, driven by feedback from our diverse group of PhD researchers.
2024
MoCCA: A Model of Comparative Concepts for Aligning Constructicons
Arthur Lorenzi | Peter Ljunglöf | Ben Lyngfelt | Tiago Timponi Torrent | William Croft | Alexander Ziem | Nina Böbel | Linnéa Bäckström | Peter Uhrig | Ely E. Matos
Proceedings of the 20th Joint ACL - ISO Workshop on Interoperable Semantic Annotation @ LREC-COLING 2024
Arthur Lorenzi | Peter Ljunglöf | Ben Lyngfelt | Tiago Timponi Torrent | William Croft | Alexander Ziem | Nina Böbel | Linnéa Bäckström | Peter Uhrig | Ely E. Matos
Proceedings of the 20th Joint ACL - ISO Workshop on Interoperable Semantic Annotation @ LREC-COLING 2024
This paper presents MoCCA, a Model of Comparative Concepts for Aligning Constructicons under development by a consortium of research groups building Constructicons of different languages including Brazilian Portuguese, English, German and Swedish. The Constructicons will be aligned by using comparative concepts (CCs) providing language-neutral definitions of linguistic properties. The CCs are drawn from typological research on grammatical categories and constructions, and from FrameNet frames, organized in a conceptual network. Language-specific constructions are linked to the CCs in accordance with general principles. MoCCA is organized into files of two types: a largely static CC Database file and multiple Linking files containing relations between constructions in a Constructicon and the CCs. Tools are planned to facilitate visualization of the CC network and linking of constructions to the CCs. All files and guidelines will be versioned, and a mechanism is set up to report cases where a language-specific construction cannot be easily linked to existing CCs.
2023
A Pipeline for the Creation of Multimodal Corpora from YouTube Videos
Nathan Dykes | Anna Wilson | Peter Uhrig
Proceedings of the 1st Workshop on Linguistic Insights from and for Multimodal Language Processing
Nathan Dykes | Anna Wilson | Peter Uhrig
Proceedings of the 1st Workshop on Linguistic Insights from and for Multimodal Language Processing
2019
The_Illiterati: Part-of-Speech Tagging for Magahi and Bhojpuri without even knowing the alphabet
Thomas Proisl | Peter Uhrig | Andreas Blombach | Natalie Dykes | Philipp Heinrich | Besim Kabashi | Sefora Mammarella
Proceedings of the First International Workshop on NLP Solutions for Under Resourced Languages (NSURL 2019) co-located with ICNLSP 2019 - Short Papers
Thomas Proisl | Peter Uhrig | Andreas Blombach | Natalie Dykes | Philipp Heinrich | Besim Kabashi | Sefora Mammarella
Proceedings of the First International Workshop on NLP Solutions for Under Resourced Languages (NSURL 2019) co-located with ICNLSP 2019 - Short Papers
2016
SoMaJo: State-of-the-art tokenization for German web and social media texts
Thomas Proisl | Peter Uhrig
Proceedings of the 10th Web as Corpus Workshop
Thomas Proisl | Peter Uhrig
Proceedings of the 10th Web as Corpus Workshop
2012
Efficient Dependency Graph Matching with the IMS Open Corpus Workbench
Thomas Proisl | Peter Uhrig
Proceedings of the Eighth International Conference on Language Resources and Evaluation (LREC'12)
Thomas Proisl | Peter Uhrig
Proceedings of the Eighth International Conference on Language Resources and Evaluation (LREC'12)
State-of-the-art dependency representations such as the Stanford Typed Dependencies may represent the grammatical relations in a sentence as directed, possibly cyclic graphs. Querying a syntactically annotated corpus for grammatical structures that are represented as graphs requires graph matching, which is a non-trivial task. In this paper, we present an algorithm for graph matching that is tailored to the properties of large, syntactically annotated corpora. The implementation of the algorithm is built on top of the popular IMS Open Corpus Workbench, allowing corpus linguists to re-use existing infrastructure. An evaluation of the resulting software, CWB-treebank, shows that its performance in real world applications, such as a web query interface, compares favourably to implementations that rely on a relational database or a dedicated graph database while at the same time offering a greater expressive power for queries. An intuitive graphical interface for building the query graphs is available via the Treebank.info project.