Jürgen Neyer


2026

Political scientists are interested in changes in political discourse over time. However, the topics of interest, such as the changes in support for or understanding of certain narratives, are often ill-defined and require deliberation, which prevents most lexical or metadata-based methods of temporal aggregation. To enable a diachronic analysis, we propose to model such settings as a series of binary document classification tasks – which current reasoning LLMs can adequately solve – and aggregate the decisions into a temporal signal. Specifically, we propose to use LLMs to classify if a parliament speech is in support of either of two narratives, and we use the monthly count of positives per narrative to track the support over time. We show that the classification is sufficiently accurate and use it to create detailed time series data showing support for the selected narratives in speeches given in the European Parliament from 2006 to 2023. The method is developed in close collaboration with political scientists and is considered an ideal starting point for diachronic analyses of political decision-making processes by domain experts.
We present a large annotated corpus of scholarly discourse in the domain of International Relations, a subfield of political science. The corpus comprises 190 articles (over 1500K tokens) annotated at the argumentation, basic rhetorical, and domain level. Five of the included articles (ca. 62K tokens) constitute a Gold-standard, coded by domain experts. The remaining articles were coded by annotators trained on the Gold-standard and monitored for annotation quality. We describe our corpus creation methodology, the annotation process and quality assurance, the corpus itself, and present insights into the data: Most argumentative structures in the data are simple premise-conclusion structures, fewer than half of the claims have explicit supporting evidence. Counter-arguments to claims are rare. The claim-to-support ratio varies widely between articles; possibly to some extent due to the topics covered (with clear common ground) or to the differences between authors’ styles. The distribution of theoretical vs. evaluative statements varies strongly between articles; this can be attributed to such factors as different methodological approaches between the articles and the methodological focus of the publishing journal.

2025

We present the first dataset, an annotation scheme, discourse analysis, and baseline experiments on argumentation and domain content types in scholarly articles on political science, specifically on the theory of International Relations (IR). The dataset comprises over 1 600 sentences stemming from three foundational articles on Neo-Realism, Liberalism, and Constructivism. We show that our annotation scheme enables educationally-relevant insight into the scholarly IR discourse and that state-of-the-art classifiers, while effective in distinguishing basic argumentative elements (Claims and Support/Attack relations) reaching up to 0.97 micro F1 , require domain-specific training and fine-tuning on the more fine-grained tasks of relation and content type prediction.