Proceedings of the 22nd Joint ACL - ISO Workshop on Interoperable Semantic Annotation and Representation (ISA-22) @ LREC 2026
Harry Bunt (Editor)
- Anthology ID:
- 2026.isa-1
- Month:
- May
- Year:
- 2026
- Address:
- Palma, Mallorca (Spain)
- Venues:
- ISA | WS
- Events:
- ISO Workshop on Interoperable Semantic Annotation (2026) | Fifteenth Language Resources and Evaluation Conference | Other Workshops and Events (2026)
- SIG:
- Publisher:
- ELRA Language Resources Association (ELRA)
- URL:
- https://preview.aclanthology.org/ingest-lrec/2026.isa-1/
- DOI:
- PDF:
- https://preview.aclanthology.org/ingest-lrec/2026.isa-1.pdf
Proceedings of the 22nd Joint ACL - ISO Workshop on Interoperable Semantic Annotation and Representation (ISA-22) @ LREC 2026
Harry Bunt
Harry Bunt
Med2Story Referential: A Domain-Specific Extension of ISO 24617-9 for Clinical Narratives Annotation
Ana Luisa Fernandes | Purificação Silvano | Nuno Guimarães | Luís Filipe Cunha | Rita Rb-Silva | Alipio Mario Jorge
Ana Luisa Fernandes | Purificação Silvano | Nuno Guimarães | Luís Filipe Cunha | Rita Rb-Silva | Alipio Mario Jorge
The semantic annotation of clinical narratives is particularly challenging due to the complexity of medical discourse and the need to integrate linguistic, semantic, and domain-specific information within a unified framework. Existing schemes tend to fall into two categories: general-purpose frameworks, which offer robust linguistic modelling but lack specialised medical representation, and domain-specific schemes, which capture clinical content yet often fail to distinguish fundamental semantic types, especially eventive expressions and referential entities. To address this gap, this study proposes Med2Story Referential, a new extension of the Text2Story annotation scheme (Silvano et al., 2021; Leal et al., 2022) (based on ISO 24617-9: 2019 ) dedicated to referential entities in clinical narratives. Building on previous work that introduced a specialised branch for eventive entities (Fernandes et al., 2025a), and informed by the UMLS Metathesaurus and expert validation from a consultant haematologist, the extension introduces eight referential categories that refine the representation of clinical actors, substances, biological entities, instruments, and documentation. The results show that ISO 24617-9: 2019 can be applied to this type of text; however, several adaptations are required, particularly with regard to the grammatical domain and the inclusion of specialised domain labels. Nonetheless, the annotation experiment conducted to validate our proposal showed that the annotation scheme and its accompanying guidelines enable a comprehensive and detailed representation of both grammatical and medical aspects. Moreover, the results indicate that the scheme can be applied effectively by annotators without medical expertise.
This article focus on a methodology for representing the semantics of polysemous markers whose meanings cannot (or do not have to) be disambiguated, even in context. We name this task (multi-)sense representation and present here the French modal verb devoir as a case study. Specifically, we reframe this task — traditionally treated as a multi-class problem — as a multi-label classification problem to account for instances that remain ambiguous due to contextual and intentional factors. In order to fine-tune our model (CamemBERT), we implement an active learning loop to enhance the annotation process and we demonstrate that combining global and local features yields the best results (F1-micro = 0.83; F1-macro = 0.79). The model is then applied on two distinct corpora, showing that the automatic analysis of devoir’s modal senses provides deeper insights into modal verb usage and facilitates comparisons across corpora differing in medium (spoken vs. written) or genre (e.g. legal discourse). Furthermore, our multi-label approach enables the detection and analysis of double-labeled instances, offering valuable applications, as for example legal discourse interpretation and second language acquisition.
A Frame and Canvas-Based Perspective-Encoding Methodology for Multimodal Semantic Annotation of Classroom Settings
Claudia Ferraz | Ely E. Matos | Frederico Belcavello | Julia Gasparetto | Juliana de Oliveira | Janina Wildfeuer | Tiago Timponi Torrent
Claudia Ferraz | Ely E. Matos | Frederico Belcavello | Julia Gasparetto | Juliana de Oliveira | Janina Wildfeuer | Tiago Timponi Torrent
We propose a methodology for the multimodal semantic annotation of classroom interactions that takes interactional canvases and semantic frames as its core analytical categories. The approach enables the systematic recording of semantic correlations among interactants, communicative modes, and material supports involved in situated meaning-making processes. The methodology encodes participant perspective by relying on the temporal alignment of multiple video recordings of the same instructional event captured from different viewpoints, allowing for the representation of how meaning construction unfolds across perceptual and interactional positions. To operationalize the proposal, we introduce an annotation tool that implements the scheme and supports the integration of multimodal data streams within a unified semantic representation framework. We conclude by discussing the limitations of the current proposal and the possibilities for extending it to other interactional settings.
Pattern Analysis (CPA) procedure developed by Hanks (2004) to manually extract recurrent language patterns from texts, can be automated using LLMs. Specifically, we examine ChatGPT and Gemini performance in the task of semantic type tagging of arguments in 150 Italian sentences realising 30 verb patterns (5 sentences per pattern). We run two experiments. In the first, we prompt ChatGPT to use the CPA ontology (about 200 hierarchically organized semantic types) in the annotation task; we provide the model with 5 sentences per pattern and ask it to assign the most specific type to the argument(s) of each sentence. In the second, we prompt both ChatGPT and Gemini to perform the task without the ontology, and ask the models to assign a single label to the argument(s) of the 5 sentences. Both experiments are performed in a zero-shot setting. We evaluate the results using the existing Italian T-PAS pattern resource as benchmark. Our results show that LLMs perform comparably well on both concrete and abstract type tagging and can therefore be used in a pilot study to support analysts in acquiring verb patterns from text.
Korean Quantification in Abstract Meaning Representation
Kiyong Lee | Chongwon Park | Younggyun Hahm | Harry Bunt | Byongrae Ryu
Kiyong Lee | Chongwon Park | Younggyun Hahm | Harry Bunt | Byongrae Ryu
This paper explores the meaning of quantification in Korean and how it is encoded in Abstract Meaning Representation (AMRg:2019) and an enriched version AMR+ accommodating Uniform Meaning Representation (UMRg:2022) and some of the contextual constraints proposed by Bos(2020). The extension makes five special references: Bunt et al. (2018), Bunt and Lee (2025), Pustejovsky et al. (2019), Bos(2020), and ISO (2025), the main reference. The aim of this paper is threefold. First, it focuses on implementing Korean AMR with the rich specification of QuantML (ISO, 2025) and its partially DRT-based semantics (Kamp and Reyle, 1993). Second, it supports the AMR multilingual development project by exploring methods for constructing a large-scale Korean AMR-annotated corpus. This line of research is necessary because Korean AMR resources remain severely underdeveloped. In addition, Korean’s agglutinative morphology and head-final syntax challenge AMR frameworks that are largely based on the analytic inflectional language English. Third, it advances the current state of the UMR 2026 multilingual shared task by contributing more fine-grained annotations of quantification specified by ISO QuantML for resource domain, individuation, distributivity, and determinacy, as well as by treating coreference and lexical or scope ambiguities in Korean.
Towards Corpus-Based Population and Visualization of ISO 24617-8 Ontology (Short Paper)
Maciej Ogrodniczuk | Dariusz Czerski
Maciej Ogrodniczuk | Dariusz Czerski
This paper presents an extension of the ISO 24617-8 ontology for discourse relations through the integration of corpus-based examples and the development of a dedicated Ontology Viewer. The goal is to bridge the gap between formal ontological representations and practical corpus-based linguistic analysis, making discourse annotation frameworks more accessible to researchers. The proposed approach introduces a method for populating the ISO ontology with instances derived from three corpora (in Polish and English) compliant with the ISO 24617-8 standard. These instances formally connect discourse relations, argument roles, and explicit connectives within a unified semantic model. The Ontology Viewer enables intuitive browsing, filtering, and full-text searching of examples by language, relation type, and connective, offering both a relation-oriented and connective-oriented perspective. The experiment demonstrates the feasibility and effectiveness of this corpus-driven instantiation method and its visualization. The system provides a foundation for future integration of multilingual discourse corpora and contributes to the development of interoperable language resources for the Semantic Web and Natural Language Processing applications.
Tracing Consensus Formation in Meetings: Annotation and Incremental Decision Modelling in the MEET Corpus
Ghazaleh Esfandiari-Baiat | Jens Edlund
Ghazaleh Esfandiari-Baiat | Jens Edlund
We present an incremental annotation scheme and discourse model designed specifically for the study of consensus formation in collaborative meetings. By grounding the representation in observable contributions and enforcing a strict no-lookahead principle, the model provides a tractable way to analyse how decisions emerge over the course of interaction. The resulting structures are intentionally minimal yet expressive enough to capture the evolving task state and support dynamic visualisation and replay of the decision process. A web-based reference implementation of the model demonstrates how the evolving decision state can be inspected and replayed during analysis. Together with a suitable corpus, this framework provides a practical foundation for investigating the multimodal dynamics of collaborative decision-making in professional meetings.
GeoAffect: A Multi-Layer Annotation Schema and Few-Shot LLM Evaluation for Geoaffective Analysis of Literary Texts
Fotini Koidaki | Stergios Chatzykiriakidis
Fotini Koidaki | Stergios Chatzykiriakidis
GeoAffect is an annotation framework that has been especially developed to capture how places are emotionally framed in literary narrative. The project focuses on nineteenth-century Greek prose fiction and brings together named entity recognition with an affect schema that distinguishes experiential, appraisal, and identity-oriented relations to place. The annotation design linked entities, emotion spans, and rhetorical devices, allowing us to model not only sentiment but also forms of belonging, alienation, and longing. To test the schema, we created a manually annotated gold dataset of approximately 360 sentences and evaluated thirteen Large Language Models in a few-shot setting for both entity recognition and affect classification. The results indicate that, with carefully designed prompts and selection strategies, LLMs can support structured geoaffective annotation even in low-resource historical language contexts.
Evaluating the Impact of LLM-Assisted Annotation in a Perspectivized Setting: The Case of FrameNet Annotation
Frederico Belcavello | Ely E. Matos | Arthur Lorenzi | Lisandra Bonoto | Livia Pádua Ruiz | Luiz Fernando Pereira | Victor Herbst | Yulla Liquer Navarro | Helen de Andrade Abreu | Lívia Vicente Dutra | Tiago Timponi Torrent
Frederico Belcavello | Ely E. Matos | Arthur Lorenzi | Lisandra Bonoto | Livia Pádua Ruiz | Luiz Fernando Pereira | Victor Herbst | Yulla Liquer Navarro | Helen de Andrade Abreu | Lívia Vicente Dutra | Tiago Timponi Torrent
The use of LLM-based applications as a means to accelerate and/or substitute human labor in the creation of language resources and datasets is a reality. Nonetheless, despite the potential of such tools for linguistic research, an evaluation of their performance and impact on the creation of annotated datasets, especially under a perspectivized approach to NLP, is still missing. This paper contributes to the reduction of this gap by reporting on an extensive evaluation of the (semi-)automatization of FrameNet-like semantic annotation by the use of an LLM-based semantic role labeler. The methodology employed compares annotation time, coverage, and diversity in three experimental settings: manual, automatic, and semi-automatic annotation. Results show that the hybrid, semi-automatic annotation setting leads to increased frame diversity and similar annotation coverage, when compared to the human-only setting, while the automatic setting performs considerably worse in all metrics, except for annotation time, which remains similar.
Annotating Word Meanings over Time: The Trade-off between Scalability, Reliability and Expressivity Power
Pierluigi Cassotti | Nina Tahmasebi
Pierluigi Cassotti | Nina Tahmasebi
Annotating the meanings of a word over time in order to document their emergence or disappearance presents substantial implementation challenges. These difficulties arise for several reasons, notably the need for sufficient expressive power in the annotation paradigm to capture unconventional or rare meanings, as well as issues of scalability related to the number of annotations required. The first challenge is particularly acute in the context of historical texts, where modern annotators must interpret word meanings in sources that are temporally distant and often absent from contemporary dictionaries and language use. The second challenge is inherent to the distribution of word meanings, which tend to occur sparsely and intermittently over long time spans. In this paper, we examine several annotation paradigms, discussing their respective advantages and limitations. We also present a pilot study on English and Swedish. Our results indicate that a usage-sense inventory based annotation paradigm can be adopted in place of a usage-pairs-based approach while maintaining expressivity power and reducing the complexity from quadratic to linear.
Dialogue interactions have varied internal structure, with flow varying, inter alia, in face of both difficulty and agreement. This study investigates eye-gaze in the linguistic progression of interactions. We observe the relation of gaze to illocutionary functions of turns through dialogue acts, and to how turns present "new" or "old" content, through lexical entropy and repetition. Results on the HCRC Map Task corpus, enabled by an event alignment annotation method described, show how gaze is related to linguistic progression. A gaze towards the conversation partner at the end of a turn tends to align with complexity and difficulties being expressed in the turn, while keeping gaze down at the map is more typical of obstacle- and disagreement-free interactions. Addressees who look up or off at the start of a turn show evidence of lexicon adaptation to gaze values.
CATS: An Annotation Scheme of Causality and Temporal Structure
Nana Yu | Purificação Silvano | Luís Filipe Cunha | Alípio Jorge
Nana Yu | Purificação Silvano | Luís Filipe Cunha | Alípio Jorge
This paper presents CATS, a causal and temporal annotation scheme designed to jointly represent causal relations and temporal structures in news texts. The proposed framework integrates components of ISO 24617 Semantic Annotation Framework (SemAF), drawing in particular on Part 1 (Time and Events) (ISO 24617-1: 2012iso24617) and Part 8 (Semantic Relations in Discourse) (ISO 24617-8: 2016iso24617-8). Building on the Text2Story annotation framework (CITATION), the scheme adapts and extends its principles for representing temporal information while introducing new entities and links for modeling causal relations. The resulting annotation model enables the integrated representation of causal arguments, events, temporal relations, and causal signals within a unified structure. By jointly capturing causal and temporal dependencies, CATS provides a resource for studying the interaction between causality and temporality in discourse and supports downstream NLP tasks such as event extraction, temporal ordering, and causal reasoning.
ISO-TimeML Semantics for Interlinking Annotations
Harry Bunt | Alex Chengyu Fang | Kiyong Lee | Volha Petukhova | James Pustejovsky | Purificação Silvano
Harry Bunt | Alex Chengyu Fang | Kiyong Lee | Volha Petukhova | James Pustejovsky | Purificação Silvano
This paper describes a step in the development of a methodology for combining annotation made with different annotation schemes. The methodology, called ‘interlinking’, assumes that different annotations of the same data will contain certain elements that refer to the same entities. This can be represented by a set of ‘identity links’. These links are used for constructing a single, integrated annotation structure at the level of abstract syntax with a semantic interpretation. In this paper we focus on the interlinking of annotations of time and events with ISO-TimeML (ISO 24617-1:2012) and quantification with QuantML (ISO 24617-12:2025).interlinking annotations is in practice only feasible if the respective annotation schemes use the same or convertible representation and interpretation formalisms. Since QuantML and ISO-TimeML use different formalisms and QuantML has a more fully developed semantics than ISO-TimeML, we developed a new, DRT–based semantics for ISO-TimeML which is presented and discussed in this paper.
From Categories to Decisions: A Framework for Attitudinal Analysis of Evaluative Language
Jiamei Zeng | Haitao Wang | Harry Bunt | Xinyu Cao | Min Dong | Tianyong Hao | Kiyong Lee | James Pustejovsky | Laurent Romary | Jianfang Zong | François Claude Rey | Sylviane Cardey | Yangli Jia | Shengqing Liao | Alex Chengyu Fang
Jiamei Zeng | Haitao Wang | Harry Bunt | Xinyu Cao | Min Dong | Tianyong Hao | Kiyong Lee | James Pustejovsky | Laurent Romary | Jianfang Zong | François Claude Rey | Sylviane Cardey | Yangli Jia | Shengqing Liao | Alex Chengyu Fang
The study reported in this paper aims to contribute to the development of an annotation scheme for evaluative language, based on Appraisal Theory, that addresses key sources of classification problems. In particular. it aims to develops a unified annotation scheme that proposes (1) a three-component annotation model comprising Appraiser, Appraised and Appraisal Element, (2) the operationalised distinction between Affect and Appreciation governed by a criterion of experiencer salience and a criterion distinguishing personal emotions from evaluations of conduct, and (3) a decision framework for the Judgement-Appreciation distinction structured on the target and lexis types operating through override conditions and substitution tests. The revised framework is illustrated with examples selected from a corpus of news discourse in English and is designed to be replicable across future Appraisal-based studies of evaluative language.