Steering Large Language Models for Machine Translation with Finetuning and In-Context Learning

Duarte M. Alves; Nuno M. Guerreiro; João Alves; José Pombal; Ricardo Rei; José G. C. de Souza; Pierre Colombo; André F. T. Martins

doi:10.18653/v1/2023.findings-emnlp.744

Steering Large Language Models for Machine Translation with Finetuning and In-Context Learning

Duarte M. Alves, Nuno M. Guerreiro, João Alves, José Pombal, Ricardo Rei, José G. C. de Souza, Pierre Colombo, André F. T. Martins

Abstract

Large language models (LLMs) are a promising avenue for machine translation (MT). However, current LLM-based MT systems are brittle: their effectiveness highly depends on the choice of few-shot examples and they often require extra post-processing due to overgeneration. Alternatives such as finetuning on translation instructions are computationally expensive and may weaken in-context learning capabilities, due to overspecialization. In this paper, we provide a closer look at this problem. We start by showing that adapter-based finetuning with LoRA matches the performance of traditional finetuning while reducing the number of training parameters by a factor of 50. This method also outperforms few-shot prompting and eliminates the need for post-processing or in-context examples. However, we show that finetuning generally degrades few-shot performance, hindering adaptation capabilities. Finally, to obtain the best of both worlds, we propose a simple approach that incorporates few-shot examples during finetuning. Experiments on 10 language pairs show that our proposed approach recovers the original few-shot capabilities while keeping the added benefits of finetuning.

Anthology ID:: 2023.findings-emnlp.744
Volume:: Findings of the Association for Computational Linguistics: EMNLP 2023
Month:: December
Year:: 2023
Address:: Singapore
Editors:: Houda Bouamor, Juan Pino, Kalika Bali
Venue:: Findings
SIG:
Publisher:: Association for Computational Linguistics
Note:
Pages:: 11127–11148
Language:
URL:: https://preview.aclanthology.org/fix___bootstrap-utility-classes/2023.findings-emnlp.744/
DOI:: 10.18653/v1/2023.findings-emnlp.744
Bibkey:
Cite (ACL):: Duarte M. Alves, Nuno M. Guerreiro, João Alves, José Pombal, Ricardo Rei, José G. C. de Souza, Pierre Colombo, and André F. T. Martins. 2023. Steering Large Language Models for Machine Translation with Finetuning and In-Context Learning. In Findings of the Association for Computational Linguistics: EMNLP 2023, pages 11127–11148, Singapore. Association for Computational Linguistics.
Cite (Informal):: Steering Large Language Models for Machine Translation with Finetuning and In-Context Learning (Alves et al., Findings 2023)
Copy Citation:
PDF:: https://preview.aclanthology.org/fix___bootstrap-utility-classes/2023.findings-emnlp.744.pdf

PDF Cite Search Fix data