Abdullah Alfaifi
2026
Saudi ASWAT: A Large-Scale Corpus of Spontaneous Saudi Arabic Speech
Abdullah I. Alharbi | Afrah A. Altamimi | Muneera Alhoshan | Amal Almazrua | Halah Munif Alharbi | Bayan M. Almuqhim | Hawra Aljasim | Abdulrahman Alosaimy | Yahya A. Asiri | Abdullah Alfaifi
Proceedings of the Fifteenth Language Resources and Evaluation Conference
Abdullah I. Alharbi | Afrah A. Altamimi | Muneera Alhoshan | Amal Almazrua | Halah Munif Alharbi | Bayan M. Almuqhim | Hawra Aljasim | Abdulrahman Alosaimy | Yahya A. Asiri | Abdullah Alfaifi
Proceedings of the Fifteenth Language Resources and Evaluation Conference
Spontaneous Arabic speech is scarce in current corpora, and it is not well represented. This poses a limitation invisibility of spontaneous Arabic to automatic speech recognition (ASR), speaker diarization, and sociolinguistic research. The Saudi ASWAT project fills a major gap by creating the first nationwide corpus of natural Saudi speech, where data has been recorded and transcribed under a systematic methodology and ecologically valid conditions. The corpus aims to collect 2,500 hours of natural conversations from a diverse range of participants. These has been selected from five major Saudi regional varieties, Najdi (Central), Eastern, Hijazi (Western), Northern, and Southern, covering more than fifty five local varieties. Speech has been recorded by trained fieldworkers using participants own devices to reflect real-life variation. The annotated data incorporate a variety of speaker demographics, regional vocabularies which differ from the standard lexicon, and structured metadata. TF–IDF profiling shows regional differences in a range of performing words. Data also represent balanced age and gender sampling to support studies of intergenerational and sociophonetic variation. Saudi ASWAT provides the most linguistically diverse resources of Saudi Arabia to date. Additionally, it establishes an ethical governed framework for Arabic speech data creation to enable advances in both computational modeling and linguistic research.
Mu’jam Arriyadh: A Comprehensive Lexicon for Contemporary Arabic Language
Afrah A. Altamimi | Abdulrahman Alosaimy | Halah Munif Alharbi | Hawra Aljasim | Muneera Alhoshan | Amal Almazrua | Hanan Alharbi | Abdulrahman Saeed Alshehri | Bayan M. Almuqhim | Maryam H. Algarny | Yahya A. Asiri | Abdullah I. Alharbi | Saleh Zaidan Albalawi | Fawziah Mohammed Asiri | Sara Ali Alhifthi | Abdullah Alfaifi
Proceedings of the Fifteenth Language Resources and Evaluation Conference
Afrah A. Altamimi | Abdulrahman Alosaimy | Halah Munif Alharbi | Hawra Aljasim | Muneera Alhoshan | Amal Almazrua | Hanan Alharbi | Abdulrahman Saeed Alshehri | Bayan M. Almuqhim | Maryam H. Algarny | Yahya A. Asiri | Abdullah I. Alharbi | Saleh Zaidan Albalawi | Fawziah Mohammed Asiri | Sara Ali Alhifthi | Abdullah Alfaifi
Proceedings of the Fifteenth Language Resources and Evaluation Conference
This paper provides an overview of Contemporary Arabic Lexicon (Mu’jam Arriyadh). It is a contemporary and inclusive Arabic dictionary that has been specifically developed to cater to the needs of both native and non-native Arabic speakers. The corpus utilized in this study is derived from the Arabic Contemporary Corpus for Analysis (ACCA), which encompasses a vast collection of 450 million words of Modern Standard Arabic spanning the previous century. Significantly, the lexicon in question prioritizes lemma-based entries over root forms, hence enhancing its user-friendliness and adaptability across different contexts. The resource offers comprehensive linguistic data pertaining to a wide array of Arabic vocabulary, encompassing morphological, morph-syntactic, and semantic aspects. The Lexicon has been developed in accordance with the ISO 24613 standard, which improves its ability to be processed by machines and facilitates the utilization of natural language processing systems. The database encompasses a range of linguistic aspects, such as synonyms, antonyms, and root forms, offering a comprehensive compilation. Mu’jam Arriyadh is a contemporary Arabic lexicon that is designed to be accessible to users, compatible with machine processing, and highly beneficial for anyone studying the language, conducting research, and utilizing natural language processing technologies.
2025
BALSAM: A Platform for Benchmarking Arabic Large Language Models
Rawan Al-Matham | Kareem Darwish | Raghad Al-Rasheed | Waad Alshammari | Muneera Alhoshan | Amal Almazrua | Asma Al Wazrah | Mais Alheraki | Firoj Alam | Preslav Nakov | Norah Alzahrani | Eman AlBilali | Nizar Habash | Abdelrahman El-Sheikh | Muhammad Elmallah | Haonan Li | Hamdy Mubarak | Mohamed Anwar | Zaid Alyafeai | Ahmed Abdelali | Nora Altwairesh | Maram Hasanain | Abdulmohsen Al Thubaity | Shady Shehata | Bashar Alhafni | Injy Hamed | Go Inoue | Khalid Elmadani | Ossama Obeid | Fatima Haouari | Tamer Elsayed | Emad Alghamdi | Khalid Almubarak | Saied Alshahrani | Ola Aljarrah | Safa Alajlan | Areej Alshaqarawi | Maryam Alshihri | Sultana Alghurabi | Atikah Alzeghayer | Afrah Altamimi | Abdullah Alfaifi | Abdulrahman AlOsaimy
Proceedings of The Third Arabic Natural Language Processing Conference
Rawan Al-Matham | Kareem Darwish | Raghad Al-Rasheed | Waad Alshammari | Muneera Alhoshan | Amal Almazrua | Asma Al Wazrah | Mais Alheraki | Firoj Alam | Preslav Nakov | Norah Alzahrani | Eman AlBilali | Nizar Habash | Abdelrahman El-Sheikh | Muhammad Elmallah | Haonan Li | Hamdy Mubarak | Mohamed Anwar | Zaid Alyafeai | Ahmed Abdelali | Nora Altwairesh | Maram Hasanain | Abdulmohsen Al Thubaity | Shady Shehata | Bashar Alhafni | Injy Hamed | Go Inoue | Khalid Elmadani | Ossama Obeid | Fatima Haouari | Tamer Elsayed | Emad Alghamdi | Khalid Almubarak | Saied Alshahrani | Ola Aljarrah | Safa Alajlan | Areej Alshaqarawi | Maryam Alshihri | Sultana Alghurabi | Atikah Alzeghayer | Afrah Altamimi | Abdullah Alfaifi | Abdulrahman AlOsaimy
Proceedings of The Third Arabic Natural Language Processing Conference
The impressive advancement of Large Language Models (LLMs) in English has not been matched across all languages. In particular, LLM performance in Arabic lags behind, due to data scarcity, linguistic diversity of Arabic and its dialects, morphological complexity, etc. Progress is further hindered by the quality of Arabic benchmarks, which typically rely on static, publicly available data, lack comprehensive task coverage, or do not provide dedicated platforms with blind test sets. This makes it challenging to measure actual progress and to mitigate data contamination. Here, we aim to bridge these gaps. In particular, we introduce BALSAM, a comprehensive, community-driven benchmark aimed at advancing Arabic LLM development and evaluation. It includes 78 NLP tasks from 14 broad categories, with 52K examples divided into 37K test and 15K development, and a centralized, transparent platform for blind evaluation. We envision BALSAM as a unifying platform that sets standards and promotes collaborative research to advance Arabic LLM capabilities.
Search
Fix author
Co-authors
- Abdulrahman AlOsaimy 3
- Muneera Alhoshan 3
- Amal Almazrua 3
- Abdullah I. Alharbi 2
- Halah Munif Alharbi 2
- Hawra Aljasim 2
- Bayan M. Almuqhim 2
- Afrah A. Altamimi 2
- Yahya A. Asiri 2
- Ahmed Abdelali 1
- Asma Al Wazrah 1
- Rawan Al-Matham 1
- Raghad Al-Rasheed 1
- Abdulmohsen Al-Thubaity 1
- Safa Alajlan 1
- Firoj Alam 1
- Saleh Zaidan Albalawi 1
- Eman Albilali 1
- Maryam H. Algarny 1
- Emad Alghamdi 1
- Sultana Alghurabi 1
- Bashar Alhafni 1
- Hanan Alharbi 1
- Mais Alheraki 1
- Sara Ali Alhifthi 1
- Ola Aljarrah 1
- Khalid Almubarak 1
- Saied Alshahrani 1
- Waad Thuwaini Alshammari 1
- Areej Alshaqarawi 1
- Abdulrahman Saeed Alshehri 1
- Maryam Alshihri 1
- Afrah Altamimi 1
- Nora Altwairesh 1
- Zaid Alyafeai 1
- Norah A. Alzahrani 1
- Atikah Alzeghayer 1
- Mohamed Anwar 1
- Fawziah Mohammed Asiri 1
- Kareem Darwish 1
- Abdelrahman El-Sheikh 1
- Khalid Elmadani 1
- Muhammad Elmallah 1
- Tamer Elsayed 1
- Nizar Habash 1
- Injy Hamed 1
- Fatima Haouari 1
- Maram Hasanain 1
- Go Inoue 1
- Haonan Li 1
- Hamdy Mubarak 1
- Preslav Nakov 1
- Ossama Obeid 1
- Shady Shehata 1