Abstract
This paper describes an ongoing effort to create, from the original hand-written text, a machine-readable, linguistically-annotated, and easily-searchable corpus of the Nahuatl portion of the Florentine Codex, a 16th century Mesoamerican manuscript written in Nahuatl and Spanish. The Codex consists of 12 books and over 300,000 tokens. We describe the process of annotating 3 of these books, the steps of text preprocessing undertaken, our approach to efficient manual processing and annotation, and some of the challenges faced along the way. We also report on a set of experiments evaluating our ability to automate the text processing tasks to aid in the remaining annotation effort, and find the results promising despite the relatively low volume of training data. Finally, we briefly present a real use case from the humanities that would benefit from the searchable, linguistically annotated corpus we describe.- Anthology ID:
- 2023.americasnlp-1.4
- Volume:
- Proceedings of the Workshop on Natural Language Processing for Indigenous Languages of the Americas (AmericasNLP)
- Month:
- July
- Year:
- 2023
- Address:
- Toronto, Canada
- Editors:
- Manuel Mager, Abteen Ebrahimi, Arturo Oncevay, Enora Rice, Shruti Rijhwani, Alexis Palmer, Katharina Kann
- Venue:
- AmericasNLP
- SIG:
- Publisher:
- Association for Computational Linguistics
- Note:
- Pages:
- 19–29
- Language:
- URL:
- https://aclanthology.org/2023.americasnlp-1.4
- DOI:
- 10.18653/v1/2023.americasnlp-1.4
- Cite (ACL):
- Francis Tyers, Robert Pugh, and Valery Berthoud F.. 2023. Codex to corpus: Exploring annotation and processing for an open and extensible machine-readable edition of the Florentine Codex. In Proceedings of the Workshop on Natural Language Processing for Indigenous Languages of the Americas (AmericasNLP), pages 19–29, Toronto, Canada. Association for Computational Linguistics.
- Cite (Informal):
- Codex to corpus: Exploring annotation and processing for an open and extensible machine-readable edition of the Florentine Codex (Tyers et al., AmericasNLP 2023)
- PDF:
- https://preview.aclanthology.org/dois-2013-emnlp/2023.americasnlp-1.4.pdf