Cuk Tho
2023
NusaCrowd: Open Source Initiative for Indonesian NLP Resources
Samuel Cahyawijaya | Holy Lovenia | Alham Fikri Aji | Genta Winata | Bryan Wilie | Fajri Koto | Rahmad Mahendra | Christian Wibisono | Ade Romadhony | Karissa Vincentio | Jennifer Santoso | David Moeljadi | Cahya Wirawan | Frederikus Hudi | Muhammad Satrio Wicaksono | Ivan Parmonangan | Ika Alfina | Ilham Firdausi Putra | Samsul Rahmadani | Yulianti Oenang | Ali Septiandri | James Jaya | Kaustubh Dhole | Arie Suryani | Rifki Afina Putri | Dan Su | Keith Stevens | Made Nindyatama Nityasya | Muhammad Adilazuarda | Ryan Hadiwijaya | Ryandito Diandaru | Tiezheng Yu | Vito Ghifari | Wenliang Dai | Yan Xu | Dyah Damapuspita | Haryo Wibowo | Cuk Tho | Ichwanul Karo Karo | Tirana Fatyanosa | Ziwei Ji | Graham Neubig | Timothy Baldwin | Sebastian Ruder | Pascale Fung | Herry Sujaini | Sakriani Sakti | Ayu Purwarianti
Findings of the Association for Computational Linguistics: ACL 2023
Samuel Cahyawijaya | Holy Lovenia | Alham Fikri Aji | Genta Winata | Bryan Wilie | Fajri Koto | Rahmad Mahendra | Christian Wibisono | Ade Romadhony | Karissa Vincentio | Jennifer Santoso | David Moeljadi | Cahya Wirawan | Frederikus Hudi | Muhammad Satrio Wicaksono | Ivan Parmonangan | Ika Alfina | Ilham Firdausi Putra | Samsul Rahmadani | Yulianti Oenang | Ali Septiandri | James Jaya | Kaustubh Dhole | Arie Suryani | Rifki Afina Putri | Dan Su | Keith Stevens | Made Nindyatama Nityasya | Muhammad Adilazuarda | Ryan Hadiwijaya | Ryandito Diandaru | Tiezheng Yu | Vito Ghifari | Wenliang Dai | Yan Xu | Dyah Damapuspita | Haryo Wibowo | Cuk Tho | Ichwanul Karo Karo | Tirana Fatyanosa | Ziwei Ji | Graham Neubig | Timothy Baldwin | Sebastian Ruder | Pascale Fung | Herry Sujaini | Sakriani Sakti | Ayu Purwarianti
Findings of the Association for Computational Linguistics: ACL 2023
We present NusaCrowd, a collaborative initiative to collect and unify existing resources for Indonesian languages, including opening access to previously non-public resources. Through this initiative, we have brought together 137 datasets and 118 standardized data loaders. The quality of the datasets has been assessed manually and automatically, and their value is demonstrated through multiple experiments.NusaCrowd’s data collection enables the creation of the first zero-shot benchmarks for natural language understanding and generation in Indonesian and the local languages of Indonesia. Furthermore, NusaCrowd brings the creation of the first multilingual automatic speech recognition benchmark in Indonesian and the local languages of Indonesia. Our work strives to advance natural language processing (NLP) research for languages that are under-represented despite being widely spoken.
Search
Fix author
Co-authors
- Muhammad Adilazuarda 1
- Alham Fikri Aji 1
- Ika Alfina 1
- Timothy Baldwin 1
- Samuel Cahyawijaya 1
- Wenliang Dai 1
- Dyah Damapuspita 1
- Kaustubh Dhole 1
- Ryandito Diandaru 1
- Tirana Noor Fatyanosa 1
- Pascale Fung 1
- Vito Ghifari 1
- Ryan Hadiwijaya 1
- Frederikus Hudi 1
- James Jaya 1
- Ziwei Ji 1
- Ichwanul Karo Karo 1
- Fajri Koto 1
- Holy Lovenia 1
- Rahmad Mahendra 1
- David Moeljadi 1
- Graham Neubig 1
- Made Nindyatama Nityasya 1
- Yulianti Oenang 1
- Ivan Parmonangan 1
- Ayu Purwarianti 1
- Ilham Firdausi Putra 1
- Rifki Afina Putri 1
- Samsul Rahmadani 1
- Ade Romadhony 1
- Sebastian Ruder 1
- Sakriani Sakti 1
- Jennifer Santoso 1
- Ali Septiandri 1
- Keith Stevens 1
- Dan Su 1
- Herry Sujaini 1
- Arie Suryani 1
- Karissa Vincentio 1
- Christian Wibisono 1
- Haryo Wibowo 1
- Muhammad Satrio Wicaksono 1
- Bryan Wilie 1
- Genta Indra Winata 1
- Cahya Wirawan 1
- Yan Xu 1
- Tiezheng Yu 1