A Word-and-Paradigm Workflow for Fieldwork Annotation
Maria Copot, Sara Court, Noah Diewald, Stephanie Antetomaso, Micha Elsner
Abstract
There are many challenges in morphological fieldwork annotation, it heavily relies on segmentation and feature labeling (which have both practical and theoretical drawbacks), it’s time-intensive, and the annotator needs to be linguistically trained and may still annotate things inconsistently. We propose a workflow that relies on unsupervised and active learning grounded in Word-and-Paradigm morphology (WP). Machine learning has the potential to greatly accelerate the annotation process and allow a human annotator to focus on problematic cases, while the WP approach makes for an annotation system that is word-based and relational, removing the need to make decisions about feature labeling and segmentation early in the process and allowing speakers of the language of interest to participate more actively, since linguistic training is not necessary. We present a proof-of-concept for the first step of the workflow, in a realistic fieldwork setting, annotators can process hundreds of forms per hour.- Anthology ID:
- 2022.computel-1.20
- Volume:
- Proceedings of the Fifth Workshop on the Use of Computational Methods in the Study of Endangered Languages
- Month:
- May
- Year:
- 2022
- Address:
- Dublin, Ireland
- Venue:
- ComputEL
- SIG:
- Publisher:
- Association for Computational Linguistics
- Note:
- Pages:
- 159–169
- Language:
- URL:
- https://aclanthology.org/2022.computel-1.20
- DOI:
- 10.18653/v1/2022.computel-1.20
- Cite (ACL):
- Maria Copot, Sara Court, Noah Diewald, Stephanie Antetomaso, and Micha Elsner. 2022. A Word-and-Paradigm Workflow for Fieldwork Annotation. In Proceedings of the Fifth Workshop on the Use of Computational Methods in the Study of Endangered Languages, pages 159–169, Dublin, Ireland. Association for Computational Linguistics.
- Cite (Informal):
- A Word-and-Paradigm Workflow for Fieldwork Annotation (Copot et al., ComputEL 2022)
- PDF:
- https://preview.aclanthology.org/paclic-22-ingestion/2022.computel-1.20.pdf