NAVER LABS Europe Submission to the Instruction-following Track

Beomseok Lee, Marcely Zanon Boito, Laurent Besacier, Ioan Calapodescu


Abstract
In this paper we describe NAVER LABS Europe submission to the instruction-following speech processing short track at IWSLT 2025. We participate in the constrained settings, developing systems that can simultaneously perform ASR, ST, and SQA tasks from English speech input into the following target languages: Chinese, Italian, and German. Our solution leverages two pretrained modules: (1) a speech-to-LLM embedding projector trained using representations from the SeamlessM4T-v2-large speech encoder; and (2) LoRA adapters trained on text data on top of Llama-3.1-8B-Instruct. These modules are jointly loaded and further instruction-tuned for 1K steps on multilingual and multimodal data to form our final system submitted for evaluation.
Anthology ID:
2025.iwslt-1.17
Volume:
Proceedings of the 22nd International Conference on Spoken Language Translation (IWSLT 2025)
Month:
July
Year:
2025
Address:
Vienna, Austria (in-person and online)
Editors:
Elizabeth Salesky, Marcello Federico, Antonis Anastasopoulos
Venues:
IWSLT | WS
SIG:
Publisher:
Association for Computational Linguistics
Note:
Pages:
186–200
Language:
URL:
https://preview.aclanthology.org/landing_page/2025.iwslt-1.17/
DOI:
Bibkey:
Cite (ACL):
Beomseok Lee, Marcely Zanon Boito, Laurent Besacier, and Ioan Calapodescu. 2025. NAVER LABS Europe Submission to the Instruction-following Track. In Proceedings of the 22nd International Conference on Spoken Language Translation (IWSLT 2025), pages 186–200, Vienna, Austria (in-person and online). Association for Computational Linguistics.
Cite (Informal):
NAVER LABS Europe Submission to the Instruction-following Track (Lee et al., IWSLT 2025)
Copy Citation:
PDF:
https://preview.aclanthology.org/landing_page/2025.iwslt-1.17.pdf