LMU Bilingual Dictionary Induction System with Word Surface Similarity Scores for BUCC 2020
Silvia Severini, Viktor Hangya, Alexander Fraser, Hinrich Schütze
Abstract
The task of Bilingual Dictionary Induction (BDI) consists of generating translations for source language words which is important in the framework of machine translation (MT). The aim of the BUCC 2020 shared task is to perform BDI on various language pairs using comparable corpora. In this paper, we present our approach to the task of English-German and English-Russian language pairs. Our system relies on Bilingual Word Embeddings (BWEs) which are often used for BDI when only a small seed lexicon is available making them particularly effective in a low-resource setting. On the other hand, they perform well on high frequency words only. In order to improve the performance on rare words as well, we combine BWE based word similarity with word surface similarity methods, such as orthography In addition to the often used top-n translation method, we experiment with a margin based approach aiming for dynamic number of translations for each source word. We participate in both the open and closed tracks of the shared task and we show improved results of our method compared to simple vector similarity based approaches. Our system was ranked in the top-3 teams and achieved the best results for English-Russian.- Anthology ID:
- 2020.bucc-1.8
- Volume:
- Proceedings of the 13th Workshop on Building and Using Comparable Corpora
- Month:
- May
- Year:
- 2020
- Address:
- Marseille, France
- Editors:
- Reinhard Rapp, Pierre Zweigenbaum, Serge Sharoff
- Venue:
- BUCC
- SIG:
- Publisher:
- European Language Resources Association
- Note:
- Pages:
- 49–55
- Language:
- English
- URL:
- https://aclanthology.org/2020.bucc-1.8
- DOI:
- Cite (ACL):
- Silvia Severini, Viktor Hangya, Alexander Fraser, and Hinrich Schütze. 2020. LMU Bilingual Dictionary Induction System with Word Surface Similarity Scores for BUCC 2020. In Proceedings of the 13th Workshop on Building and Using Comparable Corpora, pages 49–55, Marseille, France. European Language Resources Association.
- Cite (Informal):
- LMU Bilingual Dictionary Induction System with Word Surface Similarity Scores for BUCC 2020 (Severini et al., BUCC 2020)
- PDF:
- https://preview.aclanthology.org/emnlp22-frontmatter/2020.bucc-1.8.pdf