Tanja Gaustad
2026
Extension of Linguistic Resources for South African Languages: Part-of-Speech Annotated Domain-Specific Data
Tanja Gaustad | Roald Eiselen | Cindy Arlene McKellar
Proceedings of Resources for African Indigenous Languages (RAIL) 2026 @ LREC 2026
Tanja Gaustad | Roald Eiselen | Cindy Arlene McKellar
Proceedings of Resources for African Indigenous Languages (RAIL) 2026 @ LREC 2026
In this paper, we present part-of-speech (POS) annotated domain-specific data for nine South African languages. The data has been sourced from five different domains (two academic domains, Caps and theses, two non-academic domains, news and magazines, and one fiction domain, novels), uniformly pre-processed, automatically POS-tagged and then corrected by linguistic experts. The widely used NCHLT government data sets (Eiselen and Puttkammer, 2014) have also been re-tagged with the current tag sets and manually corrected. Both the new domain-specific data sets and the re-tagged NCHL data sets have been uploaded into a public repository. To illustrate the characteristics of the domain data in comparison to government data, we include and discuss data statistics, namely type-token ration (TTR), tokens per sentence and out-of-vocabulary (OOV) rates, as well as POS tagging results with a baseline tagger trained on NCHLT data and applied to the different domains for all languages. Both the data statistics and the POS results clearly show that the domain data is significantly different to government data: For all domains and languages, the tagging accuracy decreases significantly compared to testing on in-domain government data. Also, POS results for the two domains with the highest OOV rates for all languages (Caps and novels) are much lower than for the other domains. These findings emphasise the need for more diverse data resources which in turn will aid in the development of more domain-independent language technologies.
2024
The First Universal Dependency Treebank for Tswana: Tswana-Popapolelo
Tanja Gaustad | Ansu Berg | Rigardt Pretorius | Roald Eiselen
Proceedings of the Fifth Workshop on Resources for African Indigenous Languages @ LREC-COLING 2024
Tanja Gaustad | Ansu Berg | Rigardt Pretorius | Roald Eiselen
Proceedings of the Fifth Workshop on Resources for African Indigenous Languages @ LREC-COLING 2024
This paper presents the first publicly available UD treebank for Tswana, Tswana-Popapolelo. The data used consists of the 20 Cairo CICLing sentences translated to Tswana. After pre-processing these sentences with detailed POS (XPOS) and converting them to universal POS (UPOS), we proceeded to annotate the data with dependency relations, documenting decisions for the language specific constructions. Linguistic issues encountered are described in detail as this is the first application of the UD framework to produce a dependency treebank for the Bantu language family in general and for Tswana specifically.
2023
Deep learning and low-resource languages: How much data is enough? A case study of three linguistically distinct South African languages
Roald Eiselen | Tanja Gaustad
Proceedings of the Fourth workshop on Resources for African Indigenous Languages (RAIL 2023)
Roald Eiselen | Tanja Gaustad
Proceedings of the Fourth workshop on Resources for African Indigenous Languages (RAIL 2023)
In this paper we present a case study for three under-resourced linguistically distinct South African languages (Afrikaans, isiZulu, and Sesotho sa Leboa) to investigate the influence of data size and linguistic nature of a language on the performance of different embedding types. Our experimental setup consists of training embeddings on increasing amounts of data and then evaluating the impact of data size for the downstream task of part of speech tagging. We find that relatively little data can produce useful representations for this specific task for all three languages. Our analysis also shows that the influence of linguistic and orthographic differences between languages should not be underestimated: morphologically complex, conjunctively written languages (isiZulu in our case) need substantially more data to achieve good results, while disjunctively written languages require substantially less data. This is not only the case with regard to the data for training the embedding model, but also annotated training material for the task at hand. It is therefore imperative to know the characteristics of the language you are working on to make linguistically informed choices about the amount of data and the type of embeddings to use.
2007
TAT: An Author Profiling Tool with Application to Arabic Emails
Dominique Estival | Tanja Gaustad | Son Bao Pham | Will Radford | Ben Hutchinson
Proceedings of the Australasian Language Technology Workshop 2007
Dominique Estival | Tanja Gaustad | Son Bao Pham | Will Radford | Ben Hutchinson
Proceedings of the Australasian Language Technology Workshop 2007
2004
A Lemma-Based Approach to a Maximum Entropy Word Sense Disambiguation System for Dutch
Tanja Gaustad
COLING 2004: Proceedings of the 20th International Conference on Computational Linguistics
Tanja Gaustad
COLING 2004: Proceedings of the 20th International Conference on Computational Linguistics