Field to Model: Pairing Community Data Collection with Scalable NLP through the LiFE Suite
Karthick Narayanan R, Siddharth Singh, Saurabh Singh, Aryan Mathur, Ritesh Kumar, Shyam Ratan, Bornini Lahiri, Benu Pareek, Neerav Mathur, Amalesh Gope, Meiraba Takhellambam, Yogesh Dawer
Abstract
We present LiFE Suite as a “Field-to-Model” pipeline, designed to bridge community-centred data collection with scalable language model development. This paper describes the various tools integrated into the LiFE Suite that make this unified pipeline possible. Atekho, a mobile-first data collection platform, is designed to empower communities to assert their rights over their data. MATra-Lab, a web-based data processing and annotation tool, supports the management of field data and the creation of NLP-ready datasets with support from existing state-of-the-art NLP models. LiFE Model Studio, built on top of Hugging Face AutoTrain, offers a no-code solution for building scalable language models using the field data. This end-to-end integration ensures that every dataset collected in the field retains its linguistic, cultural, and metadata context, all the way through to deployable AI models and archive-ready datasets.- Anthology ID:
- 2025.fieldmatters-1.7
- Volume:
- Proceedings of the Fourth Workshop on NLP Applications to Field Linguistics
- Month:
- August
- Year:
- 2025
- Address:
- Vienna, Austria
- Editors:
- Éric Le Ferrand, Elena Klyachko, Anna Postnikova, Tatiana Shavrina, Oleg Serikov, Ekaterina Voloshina, Ekaterina Vylomova
- Venues:
- FieldMatters | WS
- SIG:
- Publisher:
- Association for Computational Linguistics
- Note:
- Pages:
- 76–84
- Language:
- URL:
- https://preview.aclanthology.org/transition-to-people-yaml/2025.fieldmatters-1.7/
- DOI:
- Cite (ACL):
- Karthick Narayanan R, Siddharth Singh, Saurabh Singh, Aryan Mathur, Ritesh Kumar, Shyam Ratan, Bornini Lahiri, Benu Pareek, Neerav Mathur, Amalesh Gope, Meiraba Takhellambam, and Yogesh Dawer. 2025. Field to Model: Pairing Community Data Collection with Scalable NLP through the LiFE Suite. In Proceedings of the Fourth Workshop on NLP Applications to Field Linguistics, pages 76–84, Vienna, Austria. Association for Computational Linguistics.
- Cite (Informal):
- Field to Model: Pairing Community Data Collection with Scalable NLP through the LiFE Suite (R et al., FieldMatters 2025)
- PDF:
- https://preview.aclanthology.org/transition-to-people-yaml/2025.fieldmatters-1.7.pdf