Part-of-Speech Tagging of Northern Sotho: Disambiguating Polysemous Function Words
Gertrud Faaß | Ulrich Heid | Elsabé Taljard | Danie Prinsloo
Proceedings of the First Workshop on Language Technologies for African Languages


Grammar-based tools for the creation of tagging resources for an unresourced language: the case of Northern Sotho
Ulrich Heid | Elsabé Taljard | Danie J. Prinsloo
Proceedings of the Fifth International Conference on Language Resources and Evaluation (LREC’06)

We describe an architecture for the parallel construction of a tagger lexicon and an annotated reference corpus for the part-of-speech tagging of Nothern Sotho, a Bantu language of South Africa, for which no tagged resources have been available so far. Our tools make use of grammatical properties (morphological and syntactic) of the language. We use symbolic pretagging, followed by stochastic tagging, an architecture which proves useful not only for the bootstrapping of tagging resources, but also for the tagging of any new text. We discuss the tagset design, the tool architecture and the current state of our ongoing effort.