Martin Johansson
2026
Exploring the similarities and differences between VLM-driven and traditional OCR for Historical Swedish Data
Martin Johansson | Selma Waginder | Dana Dannélls
Proceedings of the Fourth Workshop on the Role of Resources in the Age of Large Language Models (RESOURCEFUL 2026)
Martin Johansson | Selma Waginder | Dana Dannélls
Proceedings of the Fourth Workshop on the Role of Resources in the Age of Large Language Models (RESOURCEFUL 2026)
Recent Swedish OCR efforts rely primarily on traditional OCR methods, including deep CNN–LSTM hybrid neural networks and transformer-based models. Some approaches have also demonstrated the applicability of VLM-driven OCR to historical material. However, to date, no studies have examined in depth the performance of VLM-based OCR on historical Swedish sources. In this paper, we ask: How do transformers and VLMs differ in character- and word-level recognition performance across typefaces, and what qualitative differences can be observed in their error patterns? We show that fine-tuned versions of the Alibaba Cloud Qwen3-VL-8B-Instruct and Qwen3-VL-2B-Instruct, combined with a simple repetition-trimming step, outperform conventional OCR systems. Remaining errors are primarily attributable to challenges associated with the Blackletter typeface and formatting issues, such as missing or extra line breaks, characters, and spaces. Even when characters are correctly recognized, formatting inconsistencies can substantially increase transcription error rates.
2015
Opportunities and Obligations to Take Turns in Collaborative Multi-Party Human-Robot Interaction
Martin Johansson | Gabriel Skantze
Proceedings of the 16th Annual Meeting of the Special Interest Group on Discourse and Dialogue
Martin Johansson | Gabriel Skantze
Proceedings of the 16th Annual Meeting of the Special Interest Group on Discourse and Dialogue
Modelling situated human-robot interaction using IrisTK
Gabriel Skantze | Martin Johansson
Proceedings of the 16th Annual Meeting of the Special Interest Group on Discourse and Dialogue
Gabriel Skantze | Martin Johansson
Proceedings of the 16th Annual Meeting of the Special Interest Group on Discourse and Dialogue
2014
The Tutorbot Corpus — A Corpus for Studying Tutoring Behaviour in Multiparty Face-to-Face Spoken Dialogue
Maria Koutsombogera | Samer Al Moubayed | Bajibabu Bollepalli | Ahmed Hussen Abdelaziz | Martin Johansson | José David Aguas Lopes | Jekaterina Novikova | Catharine Oertel | Kalin Stefanov | Gül Varol
Proceedings of the Ninth International Conference on Language Resources and Evaluation (LREC'14)
Maria Koutsombogera | Samer Al Moubayed | Bajibabu Bollepalli | Ahmed Hussen Abdelaziz | Martin Johansson | José David Aguas Lopes | Jekaterina Novikova | Catharine Oertel | Kalin Stefanov | Gül Varol
Proceedings of the Ninth International Conference on Language Resources and Evaluation (LREC'14)
This paper describes a novel experimental setup exploiting state-of-the-art capture equipment to collect a multimodally rich game-solving collaborative multiparty dialogue corpus. The corpus is targeted and designed towards the development of a dialogue system platform to explore verbal and nonverbal tutoring strategies in multiparty spoken interactions. The dialogue task is centered on two participants involved in a dialogue aiming to solve a card-ordering game. The participants were paired into teams based on their degree of extraversion as resulted from a personality test. With the participants sits a tutor that helps them perform the task, organizes and balances their interaction and whose behavior was assessed by the participants after each interaction. Different multimodal signals captured and auto-synchronized by different audio-visual capture technologies, together with manual annotations of the tutor’s behavior constitute the Tutorbot corpus. This corpus is exploited to build a situated model of the interaction based on the participants’ temporally-changing state of attention, their conversational engagement and verbal dominance, and their correlation with the verbal and visual feedback and conversation regulatory actions generated by the tutor.