Automatic Spell Checker and Correction for Under-represented Spoken Languages: Case Study on Wolof

Thierno Ibrahima Cissé, Fatiha Sadat


Abstract
This paper presents a spell checker and correction tool specifically designed for Wolof, an under-represented spoken language in Africa. The proposed spell checker leverages a combination of a trie data structure, dynamic programming, and the weighted Levenshtein distance to generate suggestions for misspelled words. We created novel linguistic resources for Wolof, such as a lexicon and a corpus of misspelled words, using a semi-automatic approach that combines manual and automatic annotation methods. Despite the limited data available for the Wolof language, the spell checker’s performance showed a predictive accuracy of 98.31% and a suggestion accuracy of 93.33%.Our primary focus remains the revitalization and preservation of Wolof as an Indigenous and spoken language in Africa, providing our efforts to develop novel linguistic resources. This work represents a valuable contribution to the growth of computational tools and resources for the Wolof language and provides a strong foundation for future studies in the automatic spell checking and correction field.
Anthology ID:
2023.rail-1.1
Volume:
Proceedings of the Fourth workshop on Resources for African Indigenous Languages (RAIL 2023)
Month:
May
Year:
2023
Address:
Dubrovnik, Croatia
Editors:
Rooweither Mabuya, Don Mthobela, Mmasibidi Setaka, Menno Van Zaanen
Venue:
RAIL
SIG:
Publisher:
Association for Computational Linguistics
Note:
Pages:
1–10
Language:
URL:
https://aclanthology.org/2023.rail-1.1
DOI:
10.18653/v1/2023.rail-1.1
Bibkey:
Cite (ACL):
Thierno Ibrahima Cissé and Fatiha Sadat. 2023. Automatic Spell Checker and Correction for Under-represented Spoken Languages: Case Study on Wolof. In Proceedings of the Fourth workshop on Resources for African Indigenous Languages (RAIL 2023), pages 1–10, Dubrovnik, Croatia. Association for Computational Linguistics.
Cite (Informal):
Automatic Spell Checker and Correction for Under-represented Spoken Languages: Case Study on Wolof (Cissé & Sadat, RAIL 2023)
Copy Citation:
PDF:
https://preview.aclanthology.org/emnlp-22-attachments/2023.rail-1.1.pdf
Video:
 https://preview.aclanthology.org/emnlp-22-attachments/2023.rail-1.1.mp4