Peter Nabende

2022

Recent advances in the pre-training for language models leverage large-scale datasets to create multilingual models. However, low-resource languages are mostly left out in these datasets. This is primarily because many widely spoken languages that are not well represented on the web and therefore excluded from the large-scale crawls for datasets. Furthermore, downstream users of these models are restricted to the selection of languages originally chosen for pre-training. This work investigates how to optimally leverage existing pre-trained models to create low-resource translation systems for 16 African languages. We focus on two questions: 1) How can pre-trained models be used for languages not included in the initial pretraining? and 2) How can the resulting translation models effectively transfer to new domains? To answer these questions, we create a novel African news corpus covering 16 languages, of which eight languages are not part of any existing evaluation dataset. We demonstrate that the most effective strategy for transferring both additional languages and additional domains is to leverage small quantities of high-quality translation data to fine-tune large pre-trained models.

2011

pdf
Mining Transliterations from Wikipedia using Dynamic Bayesian Networks
Peter Nabende
Proceedings of the International Conference Recent Advances in Natural Language Processing 2011

2010

pdf
Mining Transliterations from Wikipedia Using Pair HMMs
Peter Nabende
Proceedings of the 2010 Named Entities Workshop

pdf abs
Applying a Dynamic Bayesian Network Framework to Transliteration Identification
Peter Nabende
Proceedings of the Seventh International Conference on Language Resources and Evaluation (LREC'10)

Identification of transliterations is aimed at enriching multilingual lexicons and improving performance in various Natural Language Processing (NLP) applications including Cross Language Information Retrieval (CLIR) and Machine Translation (MT). This paper describes work aimed at using the widely applied graphical models approach of Dynamic Bayesian Networks (DBNs) to transliteration identification. The task of estimating transliteration similarity is not very different from specific identification tasks where DBNs have been successfully applied; it is also possible to adapt DBN models from the other identification domains to the transliteration identification domain. In particular, we investigate the applicability of a DBN framework initially proposed by Filali and Bilmes (2005) to learn edit distance estimation parameters for use in pronunciation classification. The DBN framework enables the specification of a variety of models representing different factors that can affect string similarity estimation. Three DBN models associated with two of the DBN classes originally specified by Filali and Bilmes (2005) have been tested on an experimental set up of Russian-English transliteration identification. Two of the DBN models result in high transliteration identification accuracy and combining the models leads to even much better transliteration identification accuracy.

2009

pdf
Transliteration System Using Pair HMM with Weighted FSTs
Peter Nabende
Proceedings of the 2009 Named Entities Workshop: Shared Task on Transliteration (NEWS 2009)

Peter Nabende

2022

2011

2010

2009

Co-authors

Venues