Proceedings of the ParlaCLARIN V Workshop on Interoperability, Multilinguality, and Multimodality in Parliamentary Corpora
Maria Eskevich, Vincent Vandeghinste, David Bodron (Editors)
- Anthology ID:
- 2026.parlaclarin-1
- Month:
- May
- Year:
- 2026
- Address:
- Palma de Mallorca (Spain)
- Venues:
- ParlaCLARIN | WS
- Events:
- Fifteenth Language Resources and Evaluation Conference | Workshop on Creating, Using And Linking Parliamentary Corpora (2026) | Other Workshops and Events (2026)
- SIG:
- Publisher:
- ELRA Language Resources Association (ELRA)
- URL:
- https://preview.aclanthology.org/ingest-lrec/2026.parlaclarin-1/
- DOI:
- PDF:
- https://preview.aclanthology.org/ingest-lrec/2026.parlaclarin-1.pdf
Proceedings of the ParlaCLARIN V Workshop on Interoperability, Multilinguality, and Multimodality in Parliamentary Corpora
Maria Eskevich | Vincent Vandeghinste | David Bodron
Maria Eskevich | Vincent Vandeghinste | David Bodron
ParlaCAP is an OSCARS Open Science cascading grant project aimed at extending the use of the ParlaMint parliamentary corpora beyond corpus linguistics into the wider Social Sciences and Humanities (SSH). While ParlaMint provides a rich, comparable collection of parliamentary debates and accompanying metadata, its broader uptake has been limited. ParlaCAP addresses this by enriching the data with automatically derived political agendas and sentiment, enabling new forms of comparative political analysis. Using recent advances in multilingual transformer models, the project annotates over 8 million speeches from 28 European parliaments in more than 20 languages. By integrating ParlaMint with the Comparative Agendas Project (CAP) coding schema, ParlaCAP produces a FAIR dataset suitable for cross-national research on interaction of policy, sentiment, and political identity. The enrichments rely on two models, XLM-R-ParlaSent and XLM-R-ParlaCAP, both performing comparably to human annotators. The latter is trained using a teacher–student approach, where GPT-4o-generated labels are used to fine-tune a scalable classifier. The dataset is available via the CROSSDA repository and a user-friendly API. The talk concludes with a series of use cases demonstrating how meaningful insights can be obtained with minimal technical effort.
Quantifying Code-Switching in a Ukrainian Parliamentary Dataset 1990-2021
Olha Kanishcheva | Maria Shvedova
Olha Kanishcheva | Maria Shvedova
Analyzing code-switching – the practice of mixing multiple languages in one discourse – remains a significant task in natural language processing (NLP). This study examines the Ukrainian-Russian bilingual context, focusing on quantifying language alternation in a multilingual dataset. We introduce metrics to assess linguistic boundaries and patterns, specifically addressing the complexities of processing texts where Ukrainian and Russian are used interchangeably, including word-level hybridization. Using a corpus of approximately 200,000 tokens derived from parliamentary transcripts (1990-2021), we apply code-switching metrics to identify frequency and patterns of language use. Our findings provide insights into bilingual communication dynamics and can be used to improve language identification models for mixed-language data.
Ours and Yours: A Discourse Analysis of Political Identity Markers in Slovenian Parliamentary Discourse
Meden Katja
Meden Katja
With recent enrichments of the ParlaMint corpora, new opportunities have emerged for examining a range of political and discursive phenomena. This paper utilises the Slovenian ParlaMint corpus to investigate the construction of political identities through the possessive pronouns ’our’ (slv. naš) and ’your’ (slv. vaš) in Slovenian parliamentary discourse. The analysis uses a corpus-assisted approach, combining text-type, collocation, and keyword analyses of the lemmas naš and vaš. The data are drawn from three subcorpora compiled from speeches of Members of Parliament: (1) Our, containing speeches in which naš occurs; (2) Your, containing speeches in which vaš occurs; and (3) Our&Your, containing speeches in which both lemmas co-occur within the same sentence. The results indicate a clear alignment with established patterns of positive self-representation and negative other-representation. Occurrences of ’our’ are predominantly associated with positively evaluative discourse, with Norms & Values emerging as a central category. In contrast, occurrences of ’your’ are more typically linked to negative sentiment and keywords, with Activities/Discourse identified as the most prominent category. These findings suggest that these possessive pronouns function as markers of ideological positioning and discursive polarisation in Slovenian parliamentary debates.
Representations of Europe and the European Union in Parliamentary Discourse from a Corpus-Assisted Perspective
Anna Kryvenko
Anna Kryvenko
This article examines how Europe and the European Union are represented in parliamentary discourse across three contrasting European political trajectories: the United Kingdom, Slovenia, and Ukraine. Using the ParlaMint 5.0 corpora — uniformly encoded, linguistically annotated, and enriched with sentiment and topic metadata — the study applies a longitudinal, cross-linguistic corpus-assisted discourse approach. Mentions of the EU and Europe in English, Slovenian, and Ukrainian were extracted through targeted queries, supplemented by sentiment profiling, topic distribution analysis, and collocational comparison across three built-in subcorpora (Reference, COVID, COVID,War). The findings show that although the two concepts can overlap, their discursive functions diverge systematically: Europe appears as a broader cultural and geopolitical frame, while the EU attracts more policy-oriented uses. These differences intensify at moments of institutional change or crisis, with sentiment around Europe displaying sharper fluctuations than sentiment around the EU. Cross-national patterns align closely with each country’s EU membership status — past, present, or aspirational — shaping how the EU is invoked, assessed, or contested. The study demonstrates the value of multilingual, longitudinal corpus analysis for tracing the evolution of political concepts and for understanding how parliaments discursively negotiate Europe’s shifting institutional and geopolitical landscape.
Towards ParlaMint-DE: Improving the Interoperability of the GermaParl Corpus of Plenary Protocols of the German Bundestag
Christoph Leonhardt | Andreas Blätte
Christoph Leonhardt | Andreas Blätte
With the number of machine-readable corpora of plenary protocols continuously increasing, concerns about the potentials of harmonisation and shared encoding standards gain prominence. Interoperability of corpora can contribute to innovative research, in particular when comparative analyses are concerned. The ParlaMint encoding schema introduced by CLARIN provides comprehensive guidelines towards this goal. This contribution shows how GermaParl, a large corpus of plenary protocols of the German Bundestag, is transformed from a TEI-inspired XML format to the ParlaMint encoding schema. Based on previous work, this paper presents an adjusted preparation pipeline and discusses challenges of advancing an established resource into a new data format. The prospective ParlaMint-DE corpus will make the plenary debates in Germany from 1949 to 2025 available in a highly interoperable data format. Clear documentation and taxonomies increase the usefulness of the resource in comparative analyses, whereas additional metadata and linguistic annotation broaden its general applicability.
From Transcripts to Insights: A Digital Corpus and Interactive Speech Analysis Platform for Turkish Parliamentary Records
Basak Tepe | Irem Nur Yildirim | Onur Gungor | Susan Uskudarli
Basak Tepe | Irem Nur Yildirim | Onur Gungor | Susan Uskudarli
Turkish parliamentary transcripts constitute a unique longitudinal record of the country’s political, institutional, and linguistic evolution starting from 1920. Yet much of this archive has remained computationally inaccessible due to scanned and analog typewritten transcripts, historical orthography, and heterogeneous formats. We present a unified, machine-readable corpus of the Grand National Assembly of Türkiye (TBMM), comprising 26,648 session transcripts and 1.7 million pages encompassing ten diverse parliamentary entities spanning a century of legislative history. In addition, we introduce an open-access web platform for speech-level analysis of parliamentary debates from 1983 to 2024. The platform integrates named entity recognition, topic modeling, and diachronic semantic shift detection, enabling exploration of discourse patterns across time and parties, including the frequency and thematic focus of speech activities of specific Members of Parliament. By bridging the gap between raw archival scans and modern NLP tools, the dataset and platform support reproducible research in NLP, digital humanities, and computational social science.
Transcription and Recognition of Italian Parliamentary Speeches Using Vision-Language Models
Luigi Curini | Alfio Ferrara | Giovanni Pagano | Sergio Picascia
Luigi Curini | Alfio Ferrara | Giovanni Pagano | Sergio Picascia
Parliamentary proceedings represent a rich yet challenging resource for computational analysis, particularly when preserved only as scanned historical documents. Existing efforts to digitise Italian parliamentary speeches have relied on traditional Optical Character Recognition pipelines, resulting in transcription errors and limited semantic annotation. In this paper, we propose a pipeline based on Vision-Language Models for the automatic transcription, semantic segmentation, and entity linking of Italian parliamentary speeches. The pipeline employs a specialised OCR model to extract text while preserving reading order, followed by a large-scale Vision-Language Model that performs transcription refinement, element classification, and speaker identification by jointly reasoning over visual layout and textual content. Extracted speakers are then linked to the Chamber of Deputies knowledge base through SPARQL queries and a multi-strategy fuzzy matching procedure. Evaluation against an established benchmark demonstrates substantial improvements both in transcription quality and speaker tagging.
Beyond OCR: Structural Segmentation and Speaker Attribution in Historical Italian Parliamentary Debates
Claudia Corbetta | Samuele Mazzei | Alessio Palmero Aprosio
Claudia Corbetta | Samuele Mazzei | Alessio Palmero Aprosio
Historical parliamentary debates are essential for longitudinal political and linguistic research, yet much early material remains available only as scanned images. In the Italian context, proceedings from 1848–1996 lack large-scale, structurally annotated, machine-readable representations. This paper addresses the challenge of transforming historical Italian parliamentary debates into structured corpora by moving beyond plain Optical Character Recognition (OCR) toward functional block segmentation and speaker attribution. We present detailed annotation guidelines and a manually annotated dataset of 300 randomly sampled pages. Two approaches are compared: (i) direct multimodal Large Language Model (LLM) annotation and (ii) a modular pipeline combining OCR with LLM-based structural reconstruction under zero-shot and few-shot prompting. Evaluation on a held-out test set shows that separating transcription from structural reasoning improves performance, with few-shot prompting yielding the most reliable results. The study demonstrates the feasibility of integrating LLM-based reasoning into historical parliamentary digitisation workflows.
Computational Political Landscape of the Netherlands and Prime Minister Schoof’s Position
Wessel Ledder | Iris Hendrickx
Wessel Ledder | Iris Hendrickx
This study presents a computational model of the Dutch political landscape during the Schoof government period, constructed using debate speeches from the House of Representatives. We construct a two-dimensional representation of the Dutch political landscape by fine-tuning a BERT model on parliamentary debate speeches and applying dimensionality reduction techniques to the resulting embeddings. We evaluate the validity of this model by comparing it to an independently developed model from an external research institute, finding that both models reveal similar patterns along the socio-economic left–right dimension. We also examine content patterns and word frequency distributions in targeted samples located at distinct regions of the landscape to interpret the model. We further evaluate the stability of the landscape to ensure that the observed patterns are not driven by random variation. Finally, we position Prime Minister Schoof within this computational landscape. Schoof was intended to be a neutral Prime Minister without any party affiliation that would represent the coalition parties of the government equally. Our analysis will show whether Schoof was indeed neutral in his statements or not.