Vojko Gorjanc
2026
Thematic Landscapes of the Past: Analysing Slovene Historical Periodicals With Topic Modeling
Filip Dobranić | Uroš Šmajdek | Oliver Pejić | Ciril Bohak | Vojko Gorjanc | Tina Munda | Darja Fiser
Proceedings of the First Workshop on Creating Interoperable Corpora of Historical Newspapers
Filip Dobranić | Uroš Šmajdek | Oliver Pejić | Ciril Bohak | Vojko Gorjanc | Tina Munda | Darja Fiser
Proceedings of the First Workshop on Creating Interoperable Corpora of Historical Newspapers
This paper explores the thematic landscapes of three Slovene historical periodicals—Slovenka, Slovenec, and Slovenski narod—from the sPeriodika corpus, a comprehensive collection of Slovene press published between 1771 and 1914. Using BERTopic, we analyse the thematic profiles of these periodicals, enriched with diachronic perspectives. Our study examines the thematic commonalities and specificities of the selected periodicals, highlighting their distinct political orientations, target audiences, and the increasing nationalist polarisation in public discourse. This work contributes to digital humanities by demonstrating the potential of modern topic modelling techniques, such as BERTopic, to advance historical and cultural research.
Benchmarking LLMs for Aspect-Based Sentiment Classification in Slovene Historical Periodicals
Tina Munda | Filip Dobranić | Uroš Šmajdek | Oliver Pejić | Ciril Bohak | Vojko Gorjanc | Darja Fišer
Proceedings of Shaping Multilingual, Multimodal AI for the Social Sciences and Humanities (LLMs4SSH) @ LREC 2026
Tina Munda | Filip Dobranić | Uroš Šmajdek | Oliver Pejić | Ciril Bohak | Vojko Gorjanc | Darja Fišer
Proceedings of Shaping Multilingual, Multimodal AI for the Social Sciences and Humanities (LLMs4SSH) @ LREC 2026
Historical newspapers present substantial challenges for computational sentiment analysis due to OCR noise, archaic linguistic features, and the absence of domain-specific labeled training data. This paper examines whether instruction-following LLMs can support targeted, mention-level sentiment inference in such conditions. We benchmark four instruction-following LLMs on a manually annotated sample of collective-identity mentions drawn from Slovene historical newspapers. The results provide a benchmark for targeted sentiment classification in OCR-degraded historical Slovene and offer an empirically grounded assessment of the capabilities and limitations of an instruction-tuned LLM in digital humanities research.
2008
Slovene Terminology Web Portal and the TBX-Compatible Simplified DTD/schema
Simon Krek | Vojko Gorjanc | Špela Arhar
Proceedings of the Sixth International Conference on Language Resources and Evaluation (LREC'08)
Simon Krek | Vojko Gorjanc | Špela Arhar
Proceedings of the Sixth International Conference on Language Resources and Evaluation (LREC'08)
The paper describes the project whose main purpose is the creation of the Slovene terminology web portal, funded by the Slovene Research Agency and the Amebis software company. It focuses on the DTD/schema used for the unification of different terminology resources in different input formats into one database available on the web. Two projects involving unification DTD/schemas were taken as the model for the resulting DTD/schema: the CONCEDE project and the TMF project. The final DTD/schema was tested on twenty different specialized dictionaries, both monolingual and bilingual, in various formats either without any existing markup or with complex XML structure. The result of the project will be an on-line terminology resource for Slovenian which will also include didactic material on terminology and free tools for uploading domain-specific text collections to be processed with NLP software, including a term extractor.