Davis David
2025
AFRIDOC-MT: Document-level MT Corpus for African Languages
Jesujoba Oluwadara Alabi | Israel Abebe Azime | Miaoran Zhang | Cristina España-Bonet | Rachel Bawden | Dawei Zhu | David Ifeoluwa Adelani | Clement Oyeleke Odoje | Idris Akinade | Iffat Maab | Davis David | Shamsuddeen Hassan Muhammad | Neo Putini | David O. Ademuyiwa | Andrew Caines | Dietrich Klakow
Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing
Jesujoba Oluwadara Alabi | Israel Abebe Azime | Miaoran Zhang | Cristina España-Bonet | Rachel Bawden | Dawei Zhu | David Ifeoluwa Adelani | Clement Oyeleke Odoje | Idris Akinade | Iffat Maab | Davis David | Shamsuddeen Hassan Muhammad | Neo Putini | David O. Ademuyiwa | Andrew Caines | Dietrich Klakow
Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing
This paper introduces AFRIDOC-MT, a document-level multi-parallel translation dataset covering English and five African languages: Amharic, Hausa, Swahili, Yorùbá, and Zulu. The dataset comprises 334 health and 271 information technology news documents, all human-translated from English to these languages. We conduct document-level translation benchmark experiments by evaluating the ability of neural machine translation (NMT) models and large language models (LLMs) to translate between English and these languages, at both the sentence and pseudo-document levels, the outputs being realigned to form complete documents for evaluation. Our results indicate that NLLB-200 achieves the best average performance among the standard NMT models, while GPT-4o outperforms general-purpose LLMs. Fine-tuning selected models leads to substantial performance gains, but models trained on sentences struggle to generalize effectively to longer documents. Furthermore, our analysis reveals that some LLMs exhibit issues such as under-generation, over-generation, repetition of words and phrases, and off-target translations, specifically for translation into African languages.
2023
AfriSenti: A Twitter Sentiment Analysis Benchmark for African Languages
Shamsuddeen Hassan Muhammad | Idris Abdulmumin | Abinew Ali Ayele | Nedjma Ousidhoum | David Ifeoluwa Adelani | Seid Muhie Yimam | Ibrahim Sa'id Ahmad | Meriem Beloucif | Saif M. Mohammad | Sebastian Ruder | Oumaima Hourrane | Pavel Brazdil | Alipio Jorge | Felermino Dário Mário António Ali | Davis David | Salomey Osei | Bello Shehu Bello | Falalu Ibrahim | Tajuddeen Gwadabe | Samuel Rutunda | Tadesse Belay | Wendimu Baye Messelle | Hailu Beshada Balcha | Sisay Adugna Chala | Hagos Tesfahun Gebremichael | Bernard Opoku | Stephen Arthur
Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing
Shamsuddeen Hassan Muhammad | Idris Abdulmumin | Abinew Ali Ayele | Nedjma Ousidhoum | David Ifeoluwa Adelani | Seid Muhie Yimam | Ibrahim Sa'id Ahmad | Meriem Beloucif | Saif M. Mohammad | Sebastian Ruder | Oumaima Hourrane | Pavel Brazdil | Alipio Jorge | Felermino Dário Mário António Ali | Davis David | Salomey Osei | Bello Shehu Bello | Falalu Ibrahim | Tajuddeen Gwadabe | Samuel Rutunda | Tadesse Belay | Wendimu Baye Messelle | Hailu Beshada Balcha | Sisay Adugna Chala | Hagos Tesfahun Gebremichael | Bernard Opoku | Stephen Arthur
Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing
Africa is home to over 2,000 languages from over six language families and has the highest linguistic diversity among all continents. This includes 75 languages with at least one million speakers each. Yet, there is little NLP research conducted on African languages. Crucial in enabling such research is the availability of high-quality annotated datasets. In this paper, we introduce AfriSenti, a sentiment analysis benchmark that contains a total of >110,000 tweets in 14 African languages (Amharic, Algerian Arabic, Hausa, Igbo, Kinyarwanda, Moroccan Arabic, Mozambican Portuguese, Nigerian Pidgin, Oromo, Swahili, Tigrinya, Twi, Xitsonga, and Yoruba) from four language families. The tweets were annotated by native speakers and used in the AfriSenti-SemEval shared task (with over 200 participants, see website: https://afrisenti-semeval.github.io). We describe the data collection methodology, annotation process, and the challenges we dealt with when curating each dataset. We further report baseline experiments conducted on the AfriSenti datasets and discuss their usefulness.
MasakhaNEWS: News Topic Classification for African languages
David Ifeoluwa Adelani | Marek Masiak | Israel Abebe Azime | Jesujoba Alabi | Atnafu Lambebo Tonja | Christine Mwase | Odunayo Ogundepo | Bonaventure F. P. Dossou | Akintunde Oladipo | Doreen Nixdorf | Chris Chinenye Emezue | Sana Al-azzawi | Blessing Sibanda | Davis David | Lolwethu Ndolela | Jonathan Mukiibi | Tunde Ajayi | Tatiana Moteu | Brian Odhiambo | Abraham Owodunni | Nnaemeka Obiefuna | Muhidin Mohamed | Shamsuddeen Hassan Muhammad | Teshome Mulugeta Ababu | Saheed Abdullahi Salahudeen | Mesay Gemeda Yigezu | Tajuddeen Gwadabe | Idris Abdulmumin | Mahlet Taye | Oluwabusayo Awoyomi | Iyanuoluwa Shode | Tolulope Adelani | Habiba Abdulganiyu | Abdul-Hakeem Omotayo | Adetola Adeeko | Abeeb Afolabi | Anuoluwapo Aremu | Olanrewaju Samuel | Clemencia Siro | Wangari Kimotho | Onyekachi Ogbu | Chinedu Mbonu | Chiamaka Chukwuneke | Samuel Fanijo | Jessica Ojo | Oyinkansola Awosan | Tadesse Kebede | Toadoum Sari Sakayo | Pamela Nyatsine | Freedmore Sidume | Oreen Yousuf | Mardiyyah Oduwole | Kanda Tshinu | Ussen Kimanuka | Thina Diko | Siyanda Nxakama | Sinodos Nigusse | Abdulmejid Johar | Shafie Mohamed | Fuad Mire Hassan | Moges Ahmed Mehamed | Evrard Ngabire | Jules Jules | Ivan Ssenkungu | Pontus Stenetorp
Proceedings of the 13th International Joint Conference on Natural Language Processing and the 3rd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics (Volume 1: Long Papers)
David Ifeoluwa Adelani | Marek Masiak | Israel Abebe Azime | Jesujoba Alabi | Atnafu Lambebo Tonja | Christine Mwase | Odunayo Ogundepo | Bonaventure F. P. Dossou | Akintunde Oladipo | Doreen Nixdorf | Chris Chinenye Emezue | Sana Al-azzawi | Blessing Sibanda | Davis David | Lolwethu Ndolela | Jonathan Mukiibi | Tunde Ajayi | Tatiana Moteu | Brian Odhiambo | Abraham Owodunni | Nnaemeka Obiefuna | Muhidin Mohamed | Shamsuddeen Hassan Muhammad | Teshome Mulugeta Ababu | Saheed Abdullahi Salahudeen | Mesay Gemeda Yigezu | Tajuddeen Gwadabe | Idris Abdulmumin | Mahlet Taye | Oluwabusayo Awoyomi | Iyanuoluwa Shode | Tolulope Adelani | Habiba Abdulganiyu | Abdul-Hakeem Omotayo | Adetola Adeeko | Abeeb Afolabi | Anuoluwapo Aremu | Olanrewaju Samuel | Clemencia Siro | Wangari Kimotho | Onyekachi Ogbu | Chinedu Mbonu | Chiamaka Chukwuneke | Samuel Fanijo | Jessica Ojo | Oyinkansola Awosan | Tadesse Kebede | Toadoum Sari Sakayo | Pamela Nyatsine | Freedmore Sidume | Oreen Yousuf | Mardiyyah Oduwole | Kanda Tshinu | Ussen Kimanuka | Thina Diko | Siyanda Nxakama | Sinodos Nigusse | Abdulmejid Johar | Shafie Mohamed | Fuad Mire Hassan | Moges Ahmed Mehamed | Evrard Ngabire | Jules Jules | Ivan Ssenkungu | Pontus Stenetorp
Proceedings of the 13th International Joint Conference on Natural Language Processing and the 3rd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics (Volume 1: Long Papers)
2021
MasakhaNER: Named Entity Recognition for African Languages
David Ifeoluwa Adelani | Jade Abbott | Graham Neubig | Daniel D’souza | Julia Kreutzer | Constantine Lignos | Chester Palen-Michel | Happy Buzaaba | Shruti Rijhwani | Sebastian Ruder | Stephen Mayhew | Israel Abebe Azime | Shamsuddeen H. Muhammad | Chris Chinenye Emezue | Joyce Nakatumba-Nabende | Perez Ogayo | Aremu Anuoluwapo | Catherine Gitau | Derguene Mbaye | Jesujoba Alabi | Seid Muhie Yimam | Tajuddeen Rabiu Gwadabe | Ignatius Ezeani | Rubungo Andre Niyongabo | Jonathan Mukiibi | Verrah Otiende | Iroro Orife | Davis David | Samba Ngom | Tosin Adewumi | Paul Rayson | Mofetoluwa Adeyemi | Gerald Muriuki | Emmanuel Anebi | Chiamaka Chukwuneke | Nkiruka Odu | Eric Peter Wairagala | Samuel Oyerinde | Clemencia Siro | Tobius Saul Bateesa | Temilola Oloyede | Yvonne Wambui | Victor Akinode | Deborah Nabagereka | Maurice Katusiime | Ayodele Awokoya | Mouhamadane MBOUP | Dibora Gebreyohannes | Henok Tilaye | Kelechi Nwaike | Degaga Wolde | Abdoulaye Faye | Blessing Sibanda | Orevaoghene Ahia | Bonaventure F. P. Dossou | Kelechi Ogueji | Thierno Ibrahima DIOP | Abdoulaye Diallo | Adewale Akinfaderin | Tendai Marengereke | Salomey Osei
Transactions of the Association for Computational Linguistics, Volume 9
David Ifeoluwa Adelani | Jade Abbott | Graham Neubig | Daniel D’souza | Julia Kreutzer | Constantine Lignos | Chester Palen-Michel | Happy Buzaaba | Shruti Rijhwani | Sebastian Ruder | Stephen Mayhew | Israel Abebe Azime | Shamsuddeen H. Muhammad | Chris Chinenye Emezue | Joyce Nakatumba-Nabende | Perez Ogayo | Aremu Anuoluwapo | Catherine Gitau | Derguene Mbaye | Jesujoba Alabi | Seid Muhie Yimam | Tajuddeen Rabiu Gwadabe | Ignatius Ezeani | Rubungo Andre Niyongabo | Jonathan Mukiibi | Verrah Otiende | Iroro Orife | Davis David | Samba Ngom | Tosin Adewumi | Paul Rayson | Mofetoluwa Adeyemi | Gerald Muriuki | Emmanuel Anebi | Chiamaka Chukwuneke | Nkiruka Odu | Eric Peter Wairagala | Samuel Oyerinde | Clemencia Siro | Tobius Saul Bateesa | Temilola Oloyede | Yvonne Wambui | Victor Akinode | Deborah Nabagereka | Maurice Katusiime | Ayodele Awokoya | Mouhamadane MBOUP | Dibora Gebreyohannes | Henok Tilaye | Kelechi Nwaike | Degaga Wolde | Abdoulaye Faye | Blessing Sibanda | Orevaoghene Ahia | Bonaventure F. P. Dossou | Kelechi Ogueji | Thierno Ibrahima DIOP | Abdoulaye Diallo | Adewale Akinfaderin | Tendai Marengereke | Salomey Osei
Transactions of the Association for Computational Linguistics, Volume 9
We take a step towards addressing the under- representation of the African continent in NLP research by bringing together different stakeholders to create the first large, publicly available, high-quality dataset for named entity recognition (NER) in ten African languages. We detail the characteristics of these languages to help researchers and practitioners better understand the challenges they pose for NER tasks. We analyze our datasets and conduct an extensive empirical evaluation of state- of-the-art methods across both supervised and transfer learning settings. Finally, we release the data, code, and models to inspire future research on African NLP.1
Search
Fix author
Co-authors
- David Ifeoluwa Adelani 4
- Shamsuddeen Hassan Muhammad 4
- Jesujoba Alabi 3
- Israel Abebe Azime 3
- Idris Abdulmumin 2
- Chiamaka Chukwuneke 2
- Bonaventure F. P. Dossou 2
- Chris Chinenye Emezue 2
- Tajuddeen Gwadabe 2
- Jonathan Mukiibi 2
- Salomey Osei 2
- Sebastian Ruder 2
- Blessing Kudzaishe Sibanda 2
- Clemencia Siro 2
- Seid Muhie Yimam 2
- Teshome Mulugeta Ababu 1
- Jade Abbott 1
- Habiba Abdulganiyu 1
- Adetola Adeeko 1
- Tolulope Adelani 1
- David O. Ademuyiwa 1
- Tosin Adewumi 1
- Mofetoluwa Adeyemi 1
- Abeeb Afolabi 1
- Orevaoghene Ahia 1
- Ibrahim Sa'id Ahmad 1
- Tunde Ajayi 1
- Idris Akinade 1
- Adewale Akinfaderin 1
- Victor Akinode 1
- Sana Al-Azzawi 1
- Felermino Dário Mário António Ali 1
- Emmanuel Anebi 1
- Aremu Anuoluwapo 1
- Anuoluwapo Aremu 1
- Stephen Arthur 1
- Ayodele Awokoya 1
- Oyinkansola Awosan 1
- Oluwabusayo Awoyomi 1
- Abinew Ali Ayele 1
- Hailu Beshada Balcha 1
- Tobius Saul Bateesa 1
- Rachel Bawden 1
- Tadesse Belay 1
- Bello Shehu Bello 1
- Meriem Beloucif 1
- Pavel Brazdil 1
- Happy Buzaaba 1
- Andrew Caines 1
- Sisay Adugna Chala 1
- Thierno Ibrahima DIOP 1
- Abdoulaye Diallo 1
- Thina Diko 1
- Daniel D’souza 1
- Cristina España-Bonet 1
- Ignatius Ezeani 1
- Samuel Fanijo 1
- Abdoulaye Faye 1
- Hagos Tesfahun Gebremichael 1
- Dibora Gebreyohannes 1
- Catherine Gitau 1
- Tajuddeen Rabiu Gwadabe 1
- Fuad Mire Hassan 1
- Oumaima Hourrane 1
- Falalu Ibrahim 1
- Abdulmejid Johar 1
- Alipio Jorge 1
- Jules Jules 1
- Maurice Katusiime 1
- Tadesse Kebede 1
- Ussen Kimanuka 1
- Wangari Kimotho 1
- Dietrich Klakow 1
- Julia Kreutzer 1
- Constantine Lignos 1
- Mouhamadane MBOUP 1
- Iffat Maab 1
- Tendai Marengereke 1
- Marek Masiak 1
- Stephen Mayhew 1
- Derguene Mbaye 1
- Chinedu Mbonu 1
- Moges Ahmed Mehamed 1
- Wendimu Baye Messelle 1
- Muhidin Mohamed 1
- Shafie Mohamed 1
- Saif Mohammad 1
- Tatiana Moteu 1
- Gerald Muriuki 1
- Christine Mwase 1
- Deborah Nabagereka 1
- Joyce Nakatumba-Nabende 1
- Lolwethu Ndolela 1
- Graham Neubig 1
- Evrard Ngabire 1
- Samba Ngom 1
- Sinodos Nigusse 1
- Doreen Nixdorf 1
- Rubungo Andre Niyongabo 1
- Kelechi Nwaike 1
- Siyanda Nxakama 1
- Pamela Nyatsine 1
- Nnaemeka Obiefuna 1
- Brian Odhiambo 1
- Clement Oyeleke Odoje 1
- Nkiruka Odu 1
- Mardiyyah Oduwole 1
- Perez Ogayo 1
- Onyekachi Ogbu 1
- Kelechi Ogueji 1
- Odunayo Ogundepo 1
- Jessica Ojo 1
- Akintunde Oladipo 1
- Temilola Oloyede 1
- Abdul-Hakeem Omotayo 1
- Bernard Opoku 1
- Iroro Orife 1
- Verrah Otiende 1
- Nedjma Ousidhoum 1
- Abraham Toluwase Owodunni 1
- Samuel Oyerinde 1
- Chester Palen-Michel 1
- Neo Putini 1
- Paul Rayson 1
- Shruti Rijhwani 1
- Samuel Rutunda 1
- Toadoum Sari Sakayo 1
- Saheed Abdullahi Salahudeen 1
- Olanrewaju Samuel 1
- Iyanuoluwa Shode 1
- Freedmore Sidume 1
- Ivan Ssenkungu 1
- Pontus Stenetorp 1
- Mahlet Taye 1
- Henok Tilaye 1
- Atnafu Lambebo Tonja 1
- Kanda Tshinu 1
- Eric Peter Wairagala 1
- Yvonne Wambui 1
- Degaga Wolde 1
- Mesay Gemeda Yigezu 1
- Oreen Yousuf 1
- Miaoran Zhang 1
- Dawei Zhu 1