Julia Krasselt
2026
Do We Still Need Corpora and Corpus Analysis Platforms? Discourse Analysis in Times of LLMs
Julia Krasselt | Dolores Lemmenmeier-Batinić | Philipp Dreesen
Proceedings of Shaping Multilingual, Multimodal AI for the Social Sciences and Humanities (LLMs4SSH) @ LREC 2026
Julia Krasselt | Dolores Lemmenmeier-Batinić | Philipp Dreesen
Proceedings of Shaping Multilingual, Multimodal AI for the Social Sciences and Humanities (LLMs4SSH) @ LREC 2026
Corpus-based discourse analysis investigates the linguistic construction of societally shared knowledge by iterating between quantitative pattern detection and qualitative interpretation in large text collections. Large Language Models (LLMs) promise to lower practical barriers to such work (e.g., natural-language querying, qualitative coding), yet they also introduce risks that are especially consequential in discourse-analytic settings, where fluent summaries can encourage ungrounded interpretation. This position paper argues that integrating LLMs into corpus analysis platforms is appropriate only insofar as it remains compatible with three epistemic premises of corpus research: (1) transparency of the data basis and traceability of analytical operations; (2) interpretability as evidence-constrained sense-making; and (3) seriality and patternedness as distributional structure and variation. In this opinion paper, we contribute a platform-oriented requirements perspective that translates these premises into design constraints for tool-calling/RAG-style integration, and we outline implementation directions that treat LLMs as an interaction layer over inspectable corpus retrieval and platform-based analysis.
Swiss-AL: Language Data Platform for Applied Sciences
Julia Krasselt | Philipp Dreesen | Dolores Lemmenmeier-Batinić | Sooyeon Geckeler | Klaus Rothenhäusler | Matthias Fluor
Proceedings of the 12th Workshop on Challenges in the Management of Large Corpora
Julia Krasselt | Philipp Dreesen | Dolores Lemmenmeier-Batinić | Sooyeon Geckeler | Klaus Rothenhäusler | Matthias Fluor
Proceedings of the 12th Workshop on Challenges in the Management of Large Corpora
This paper introduces Swiss-AL, a language data platform designed for the multilingual, comparative analysis of public discourse in Switzerland. Swiss-AL is an open research data resource providing browser-based access to a variety of corpora in all four of Switzerland’s official languages. Corpora contain journalistic, organisational, and parliamentary discourse. The platform supports research in applied linguistics as well as neighbouring disciplines (e.g., social sciences, communication and media studies).
2020
Swiss-AL: A Multilingual Swiss Web Corpus for Applied Linguistics
Julia Krasselt | Philipp Dressen | Matthias Fluor | Cerstin Mahlow | Klaus Rothenhäusler | Maren Runte
Proceedings of the Twelfth Language Resources and Evaluation Conference
Julia Krasselt | Philipp Dressen | Matthias Fluor | Cerstin Mahlow | Klaus Rothenhäusler | Maren Runte
Proceedings of the Twelfth Language Resources and Evaluation Conference
The Swiss Web Corpus for Applied Linguistics (Swiss-AL) is a multilingual (German, French, Italian) collection of texts from selected web sources. Unlike most other web corpora it is not intended for NLP purposes, but rather designed to support data-based and data-driven research on societal and political discourses in Switzerland. It currently contains 8 million texts (approx. 1.55 billion tokens), including news and specialist publications, governmental opinions, and parliamentary records, web sites of political parties, companies, and universities, statements from industry associations and NGOs, etc. A flexible processing pipeline using state-of-the-art components allows researchers in applied linguistics to create tailor-made subcorpora for studying discourse in a wide range of domains. So far, Swiss-AL has been used successfully in research on Swiss public discourses on energy and on antibiotic resistance.