Davis David - ACL Anthology

This is an internal, incomplete preview of a proposed change to the ACL Anthology. For efficiency reasons, we don't generate MODS or Endnote formats, and the preview may be incomplete in other ways, or contain mistakes. Do not treat this content as an official publication.

Davis David

2025

This paper introduces AFRIDOC-MT, a document-level multi-parallel translation dataset covering English and five African languages: Amharic, Hausa, Swahili, Yorùbá, and Zulu. The dataset comprises 334 health and 271 information technology news documents, all human-translated from English to these languages. We conduct document-level translation benchmark experiments by evaluating the ability of neural machine translation (NMT) models and large language models (LLMs) to translate between English and these languages, at both the sentence and pseudo-document levels, the outputs being realigned to form complete documents for evaluation. Our results indicate that NLLB-200 achieves the best average performance among the standard NMT models, while GPT-4o outperforms general-purpose LLMs. Fine-tuning selected models leads to substantial performance gains, but models trained on sentences struggle to generalize effectively to longer documents. Furthermore, our analysis reveals that some LLMs exhibit issues such as under-generation, over-generation, repetition of words and phrases, and off-target translations, specifically for translation into African languages.

2023

Africa is home to over 2,000 languages from over six language families and has the highest linguistic diversity among all continents. This includes 75 languages with at least one million speakers each. Yet, there is little NLP research conducted on African languages. Crucial in enabling such research is the availability of high-quality annotated datasets. In this paper, we introduce AfriSenti, a sentiment analysis benchmark that contains a total of >110,000 tweets in 14 African languages (Amharic, Algerian Arabic, Hausa, Igbo, Kinyarwanda, Moroccan Arabic, Mozambican Portuguese, Nigerian Pidgin, Oromo, Swahili, Tigrinya, Twi, Xitsonga, and Yoruba) from four language families. The tweets were annotated by native speakers and used in the AfriSenti-SemEval shared task (with over 200 participants, see website: https://afrisenti-semeval.github.io). We describe the data collection methodology, annotation process, and the challenges we dealt with when curating each dataset. We further report baseline experiments conducted on the AfriSenti datasets and discuss their usefulness.

MasakhaNEWS: News Topic Classification for African languages
David Ifeoluwa Adelani | Marek Masiak | Israel Abebe Azime | Jesujoba Alabi | Atnafu Lambebo Tonja | Christine Mwase | Odunayo Ogundepo | Bonaventure F. P. Dossou | Akintunde Oladipo | Doreen Nixdorf | Chris Chinenye Emezue | Sana Al-azzawi | Blessing Sibanda | Davis David | Lolwethu Ndolela | Jonathan Mukiibi | Tunde Ajayi | Tatiana Moteu | Brian Odhiambo | Abraham Owodunni | Nnaemeka Obiefuna | Muhidin Mohamed | Shamsuddeen Hassan Muhammad | Teshome Mulugeta Ababu | Saheed Abdullahi Salahudeen | Mesay Gemeda Yigezu | Tajuddeen Gwadabe | Idris Abdulmumin | Mahlet Taye | Oluwabusayo Awoyomi | Iyanuoluwa Shode | Tolulope Adelani | Habiba Abdulganiyu | Abdul-Hakeem Omotayo | Adetola Adeeko | Abeeb Afolabi | Anuoluwapo Aremu | Olanrewaju Samuel | Clemencia Siro | Wangari Kimotho | Onyekachi Ogbu | Chinedu Mbonu | Chiamaka Chukwuneke | Samuel Fanijo | Jessica Ojo | Oyinkansola Awosan | Tadesse Kebede | Toadoum Sari Sakayo | Pamela Nyatsine | Freedmore Sidume | Oreen Yousuf | Mardiyyah Oduwole | Kanda Tshinu | Ussen Kimanuka | Thina Diko | Siyanda Nxakama | Sinodos Nigusse | Abdulmejid Johar | Shafie Mohamed | Fuad Mire Hassan | Moges Ahmed Mehamed | Evrard Ngabire | Jules Jules | Ivan Ssenkungu | Pontus Stenetorp
Proceedings of the 13th International Joint Conference on Natural Language Processing and the 3rd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics (Volume 1: Long Papers)

2021

MasakhaNER: Named Entity Recognition for African Languages
David Ifeoluwa Adelani | Jade Abbott | Graham Neubig | Daniel D’souza | Julia Kreutzer | Constantine Lignos | Chester Palen-Michel | Happy Buzaaba | Shruti Rijhwani | Sebastian Ruder | Stephen Mayhew | Israel Abebe Azime | Shamsuddeen H. Muhammad | Chris Chinenye Emezue | Joyce Nakatumba-Nabende | Perez Ogayo | Aremu Anuoluwapo | Catherine Gitau | Derguene Mbaye | Jesujoba Alabi | Seid Muhie Yimam | Tajuddeen Rabiu Gwadabe | Ignatius Ezeani | Rubungo Andre Niyongabo | Jonathan Mukiibi | Verrah Otiende | Iroro Orife | Davis David | Samba Ngom | Tosin Adewumi | Paul Rayson | Mofetoluwa Adeyemi | Gerald Muriuki | Emmanuel Anebi | Chiamaka Chukwuneke | Nkiruka Odu | Eric Peter Wairagala | Samuel Oyerinde | Clemencia Siro | Tobius Saul Bateesa | Temilola Oloyede | Yvonne Wambui | Victor Akinode | Deborah Nabagereka | Maurice Katusiime | Ayodele Awokoya | Mouhamadane MBOUP | Dibora Gebreyohannes | Henok Tilaye | Kelechi Nwaike | Degaga Wolde | Abdoulaye Faye | Blessing Sibanda | Orevaoghene Ahia | Bonaventure F. P. Dossou | Kelechi Ogueji | Thierno Ibrahima DIOP | Abdoulaye Diallo | Adewale Akinfaderin | Tendai Marengereke | Salomey Osei
Transactions of the Association for Computational Linguistics, Volume 9

We take a step towards addressing the under- representation of the African continent in NLP research by bringing together different stakeholders to create the first large, publicly available, high-quality dataset for named entity recognition (NER) in ten African languages. We detail the characteristics of these languages to help researchers and practitioners better understand the challenges they pose for NER tasks. We analyze our datasets and conduct an extensive empirical evaluation of state- of-the-art methods across both supervised and transfer learning settings. Finally, we release the data, code, and models to inspire future research on African NLP.1

Co-authors

Chiamaka Chukwuneke 2

Bonaventure F. P. Dossou 2

Chris Chinenye Emezue 2

Tajuddeen Gwadabe 2

Jonathan Mukiibi 2

Sebastian Ruder 2

Blessing Kudzaishe Sibanda 2

Clemencia Siro 2

Seid Muhie Yimam 2

Teshome Mulugeta Ababu 1

Habiba Abdulganiyu 1

Adetola Adeeko 1

Tolulope Adelani 1

David O. Ademuyiwa 1

Tosin Adewumi 1

Mofetoluwa Adeyemi 1

Abeeb Afolabi 1

Orevaoghene Ahia 1

Ibrahim Sa'id Ahmad 1

Idris Akinade 1

Adewale Akinfaderin 1

Victor Akinode 1

Sana Al-Azzawi 1

Felermino Dário Mário António Ali 1

Emmanuel Anebi 1

Aremu Anuoluwapo 1

Anuoluwapo Aremu 1

Stephen Arthur 1

Ayodele Awokoya 1

Oyinkansola Awosan 1

Oluwabusayo Awoyomi 1

Abinew Ali Ayele 1

Hailu Beshada Balcha 1

Tobius Saul Bateesa 1

Rachel Bawden 1

Tadesse Belay 1

Bello Shehu Bello 1

Meriem Beloucif 1

Pavel Brazdil 1

Happy Buzaaba 1

Andrew Caines 1

Sisay Adugna Chala 1

Thierno Ibrahima DIOP 1

Abdoulaye Diallo 1

Daniel D’souza 1

Cristina España-Bonet 1

Ignatius Ezeani 1

Samuel Fanijo 1

Abdoulaye Faye 1

Hagos Tesfahun Gebremichael 1

Dibora Gebreyohannes 1

Catherine Gitau 1

Tajuddeen Rabiu Gwadabe 1

Fuad Mire Hassan 1

Oumaima Hourrane 1

Falalu Ibrahim 1

Abdulmejid Johar 1

Maurice Katusiime 1

Tadesse Kebede 1

Ussen Kimanuka 1

Wangari Kimotho 1

Dietrich Klakow 1

Julia Kreutzer 1

Constantine Lignos 1

Mouhamadane MBOUP 1

Tendai Marengereke 1

Stephen Mayhew 1

Derguene Mbaye 1

Chinedu Mbonu 1

Moges Ahmed Mehamed 1

Wendimu Baye Messelle 1

Muhidin Mohamed 1

Shafie Mohamed 1

Saif Mohammad 1

Tatiana Moteu 1

Gerald Muriuki 1

Christine Mwase 1

Deborah Nabagereka 1

Joyce Nakatumba-Nabende 1

Lolwethu Ndolela 1

Graham Neubig 1

Evrard Ngabire 1

Sinodos Nigusse 1

Doreen Nixdorf 1

Rubungo Andre Niyongabo 1

Kelechi Nwaike 1

Siyanda Nxakama 1

Pamela Nyatsine 1

Nnaemeka Obiefuna 1

Brian Odhiambo 1

Clement Oyeleke Odoje 1

Mardiyyah Oduwole 1

Onyekachi Ogbu 1

Kelechi Ogueji 1

Odunayo Ogundepo 1

Akintunde Oladipo 1

Temilola Oloyede 1

Abdul-Hakeem Omotayo 1

Bernard Opoku 1

Verrah Otiende 1

Nedjma Ousidhoum 1

Abraham Toluwase Owodunni 1

Samuel Oyerinde 1

Chester Palen-Michel 1

Shruti Rijhwani 1

Samuel Rutunda 1

Toadoum Sari Sakayo 1

Saheed Abdullahi Salahudeen 1

Olanrewaju Samuel 1

Iyanuoluwa Shode 1

Freedmore Sidume 1

Ivan Ssenkungu 1

Pontus Stenetorp 1

Atnafu Lambebo Tonja 1

Eric Peter Wairagala 1

Yvonne Wambui 1

Mesay Gemeda Yigezu 1

Miaoran Zhang 1

Venues