Itziar Cortes Etxabe


2018

pdf
Neural Machine Translation of Basque
Thierry Etchegoyhen | Eva Martínez Garcia | Andoni Azpeitia | Gorka Labaka | Iñaki Alegria | Itziar Cortes Etxabe | Amaia Jauregi Carrera | Igor Ellakuria Santos | Maite Martin | Eusebi Calonge
Proceedings of the 21st Annual Conference of the European Association for Machine Translation

We describe the first experimental results in neural machine translation for Basque. As a synthetic language featuring agglutinative morphology, an extended case system, complex verbal morphology and relatively free word order, Basque presents a large number of challenging characteristics for machine translation in general, and for data-driven approaches such as attentionbased encoder-decoder models in particular. We present our results on a large range of experiments in Basque-Spanish translation, comparing several neural machine translation system variants with both rule-based and statistical machine translation systems. We demonstrate that significant gains can be obtained with a neural network approach for this challenging language pair, and describe optimal configurations in terms of word segmentation and decoding parameters, measured against test sets that feature multiple references to account for word order variability.