Andreas Søeborg Kirkedal
Also published as: Andreas Søeborg Kirkedal
2026
Quantizing Whisper: How Design Choices Affect ASR Performance
Arthur Söhler | Julian Irigoyen | Andreas Søeborg Kirkedal
Proceedings of Speech Language Models in Low-Resource Settings: Performance, Evaluation, and Bias Analysis (SPEAKABLE) @ LREC 2026
Arthur Söhler | Julian Irigoyen | Andreas Søeborg Kirkedal
Proceedings of Speech Language Models in Low-Resource Settings: Performance, Evaluation, and Bias Analysis (SPEAKABLE) @ LREC 2026
Large speech recognition models like OpenAI’s Whisper achieve high accuracy but are difficult to deploy in resource-constrained environments due to their high memory and computational demands. This matters for low-resource and on-device settings, where compute and memory constraints often limit the practical use and evaluation of ASR systems. To address this, we present a unified, cross-library evaluation of post-training quantization (PTQ) on Whisper-small, comparing supported configurations across quantization scheme, method, granularity, and bit-width. Our study is based on four libraries—PyTorch, Optimum-Quanto, HQQ, and bitsandbytes. Experiments on LibriSpeech test-clean and test-other show that dynamic int8 quantization with Optimum-Quanto offers the best trade-off, reducing model size by 57% while lowering Word Error Rate below the baseline. Additional experiments on Whisper-base and Whisper-tiny confirm these trends, though with more pronounced degradation at lower bit-widths. Static quantization performed worse, likely due to the absence of efficient low-bit implementations for operations such as LayerNorm and Softmax. More aggressive formats (e.g., nf4, int3) achieved up to 71% compression at the cost of accuracy in acoustically challenging conditions. Our results demonstrate that carefully chosen PTQ methods can substantially reduce model size and inference cost without retraining, enabling efficient deployment of Whisper on constrained hardware.
2018
Acoustic Word Disambiguation with Phonogical Features in Danish ASR
Andreas Søeborg Kirkedal
Proceedings of the Fifteenth Workshop on Computational Research in Phonetics, Phonology, and Morphology
Andreas Søeborg Kirkedal
Proceedings of the Fifteenth Workshop on Computational Research in Phonetics, Phonology, and Morphology
Phonological features can indicate word class and we can use word class information to disambiguate both homophones and homographs in automatic speech recognition (ASR). We show Danish stød can be predicted from speech and used to improve ASR. We discover which acoustic features contain the signal of stød, how to use these features to predict stød and how we can make use of stød and stødpredictive acoustic features to improve overall ASR accuracy and decoding speed. In the process, we discover acoustic features that are novel to the phonetic characterisation of stød.
2015
Assessing the Performance of Automatic Speech Recognition Systems When Used by Native and Non-Native Speakers of Three Major Languages in Dictation Workflows
Julián Zapata | Andreas Søeborg Kirkedal
Proceedings of the 20th Nordic Conference of Computational Linguistics (NODALIDA 2015)
Julián Zapata | Andreas Søeborg Kirkedal
Proceedings of the 20th Nordic Conference of Computational Linguistics (NODALIDA 2015)
2013
Analysis of Phonetic Transcription for Danish Automatic Speech Recognition
Andreas Søeborg Kirkedal
Proceedings of the 19th Nordic Conference of Computational Linguistics (NODALIDA 2013)
Andreas Søeborg Kirkedal
Proceedings of the 19th Nordic Conference of Computational Linguistics (NODALIDA 2013)