Börje F. Karlsson - ACL Anthology

This is an internal, temporary preview of a proposed change to the ACL Anthology. It may be incomplete or contain mistakes. Please do not link to this content or treat it as official. It will be removed when the change is merged or abandoned.

Börje F. Karlsson

2024

pdf bib abs
SEACrowd: A Multilingual Multimodal Data Hub and Benchmark Suite for Southeast Asian Languages
Holy Lovenia | Rahmad Mahendra | Salsabil Maulana Akbar | Lester James Validad Miranda | Jennifer Santoso | Elyanah Aco | Akhdan Fadhilah | Jonibek Mansurov | Joseph Marvin Imperial | Onno P. Kampman | Joel Ruben Antony Moniz | Muhammad Ravi Shulthan Habibi | Frederikus Hudi | Railey Montalan | Ryan Ignatius Hadiwijaya | Joanito Agili Lopo | William Nixon | Börje F. Karlsson | James Jaya | Ryandito Diandaru | Yuze Gao | Patrick Amadeus Irawan | Bin Wang | Jan Christian Blaise Cruz | Chenxi Whitehouse | Ivan Halim Parmonangan | Maria Khelli | Wenyu Zhang | Lucky Susanto | Reynard Adha Ryanda | Sonny Lazuardi Hermawan | Dan John Velasco | Muhammad Dehan Al Kautsar | Willy Fitra Hendria | Yasmin Moslem | Noah Flynn | Muhammad Farid Adilazuarda | Haochen Li | Johanes Lee | R. Damanhuri | Shuo Sun | Muhammad Reza Qorib | Amirbek Djanibekov | Wei Qi Leong | Quyet V. Do | Niklas Muennighoff | Tanrada Pansuwan | Ilham Firdausi Putra | Yan Xu | Tai Ngee Chia | Ayu Purwarianti | Sebastian Ruder | William Chandra Tjhi | Peerat Limkonchotiwat | Alham Fikri Aji | Sedrick Keh | Genta Indra Winata | Ruochen Zhang | Fajri Koto | Zheng Xin Yong | Samuel Cahyawijaya
Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing

Southeast Asia (SEA) is a region rich in linguistic diversity and cultural variety, with over 1,300 indigenous languages and a population of 671 million people. However, prevailing AI models suffer from a significant lack of representation of texts, images, and audio datasets from SEA, compromising the quality of AI models for SEA languages. Evaluating models for SEA languages is challenging due to the scarcity of high-quality datasets, compounded by the dominance of English training data, raising concerns about potential cultural misrepresentation. To address these challenges, through a collaborative movement, we introduce SEACrowd, a comprehensive resource center that fills the resource gap by providing standardized corpora in nearly 1,000 SEA languages across three modalities. Through our SEACrowd benchmarks, we assess the quality of AI models on 36 indigenous languages across 13 tasks, offering valuable insights into the current AI landscape in SEA. Furthermore, we propose strategies to facilitate greater AI advancements, maximizing potential utility and resource equity for the future of AI in Southeast Asia.

pdf bib abs
Universal NER: A Gold-Standard Multilingual Named Entity Recognition Benchmark
Stephen Mayhew | Terra Blevins | Shuheng Liu | Marek Šuppa | Hila Gonen | Joseph Marvin Imperial | Börje F. Karlsson | Peiqin Lin | Nikola Ljubešić | LJ Miranda | Barbara Plank | Arij Riabi | Yuval Pinter
Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers)

We introduce Universal NER (UNER), an open, community-driven project to develop gold-standard NER benchmarks in many languages. The overarching goal of UNER is to provide high-quality, cross-lingually consistent annotations to facilitate and standardize multilingual NER research. UNER v1 contains 19 datasets annotated with named entities in a cross-lingual consistent schema across 13 diverse languages. In this paper, we detail the dataset creation and composition of UNER; we also provide initial modeling baselines on both in-language and cross-lingual learning settings. We will release the data, code, and fitted models to the public.

Co-authors

Muhammad Dehan Al Kautsar 1

Terra Blevins 1

Samuel Cahyawijaya 1

Tai Ngee Chia 1

Jan Christian Blaise Cruz 1

Ryandito Diandaru 1

Amirbek Djanibekov 1

Akhdan Fadhilah 1

Muhammad Ravi Shulthan Habibi 1

Ryan Ignatius Hadiwijaya 1

Willy Fitra Hendria 1

Sonny Lazuardi Hermawan 1

Frederikus Hudi 1

Patrick Amadeus Irawan 1

Onno P. Kampman 1

Peerat Limkonchotiwat 1

Nikola Ljubešić 1

Joanito Agili Lopo 1

Rahmad Mahendra 1

Jonibek Mansurov 1

Stephen Mayhew 1

Lester James Validad Miranda 1

Joel Ruben Antony Moniz 1

Jann Railey Montalan 1

Yasmin Moslem 1

Niklas Muennighoff 1

William Nixon 1

Tanrada Pansuwan 1

Ivan Halim Parmonangan 1

Barbara Plank 1

Ayu Purwarianti 1

Ilham Firdausi Putra 1

Muhammad Reza Qorib 1

Sebastian Ruder 1

Reynard Adha Ryanda 1

Jennifer Santoso 1

Lucky Susanto 1

William Chandra Tjhi 1

Dan John Velasco 1

Chenxi Whitehouse 1

Genta Indra Winata 1

Zheng-Xin Yong 1

Ruochen Zhang 1

Venues

emnlp1
naacl1