Mohammad Mohammadamini
2026
Central Kurdish Text-to-Speech and Its Application in Speech-to-Text Translation
Mohammad Mohammadamini | Meysam Shamsi | Marie Tahon
Proceedings of the Fifteenth Language Resources and Evaluation Conference
Mohammad Mohammadamini | Meysam Shamsi | Marie Tahon
Proceedings of the Fifteenth Language Resources and Evaluation Conference
In this study, we show how from available resources develop high-quality TTS models for low-resource scenarios that according to our extensive evaluation surpass the models trained on dedicated TTS data recorded in the studio. We develop three Text-to-Speech (TTS) models for Central Kurdish as a low-resource language using F5-TTS architecture. The models are trained on Central Kurdish TTS datasets in which two of them are curated from audiobooks during this study and the third one is evaluated for the first time. We also demonstrate the potential of TTS models for developing other speech technologies in low-resource languages by proposing a speech synthesis framework used in a speech-to-text translation application, achieving promising results on standard speech translation benchmarks. The curated TTS resources and models will be publicly available under CC BY-NC-ND 4.0 license
English to Central Kurdish Speech Translation: Corpus Creation, Evaluation, and Orthographic Standardization
Mohammad Mohammadamini | Daban Jaff | Josep Crego | Marie Tahon | Antoine LAURENT
Proceedings of the Fifteenth Language Resources and Evaluation Conference
Mohammad Mohammadamini | Daban Jaff | Josep Crego | Marie Tahon | Antoine LAURENT
Proceedings of the Fifteenth Language Resources and Evaluation Conference
We present KUTED, a speech-to-text translation (S2TT) dataset for Central Kurdish, derived from TED and TEDx talks. The corpus comprises 91,000 sentence pairs, including 170 hours of English audio, 1.65 million English tokens, and 1.40 million Central Kurdish tokens. We evaluate KUTED on the S2TT task and find that orthographic variation significantly degrades Kurdish translation performance, producing nonstandard outputs. To address this, we propose a systematic text standardization approach that yields substantial performance gains and more consistent translations. On a test set separated from TED talks, a fine-tuned Seamless model achieves 15.18 BLEU, and we improve Seamless baseline by 3.0 BLEU on the FLEURS benchmark. We also train a Transformer model from scratch and evaluate a cascaded system that combines Seamless (ASR) with NLLB (MT).
Southern Kurdish Speech Recognition Resources and Benchmarking
Mohammad Mohammadamini | Marie Tahon
Proceedings of the Fifteenth Language Resources and Evaluation Conference
Mohammad Mohammadamini | Marie Tahon
Proceedings of the Fifteenth Language Resources and Evaluation Conference
This article introduces a dedicated speech recognition dataset for Southern Kurdish, which is a threatened variant of Kurdish macrolanguage. We present 30 hours of validated read speech for training and an evaluation benchmark for Southern Kurdish Automatic Speech Recognition (ASR). Both the training data and evaluation benchmark are read speech recorded by crowdsourcing campaigns. Besides a detailed description of the provided resources, we provide the ASR baselines using Whisper-turbo and wav2vec-bert CTC architectures. We achieved a 4.09 CER and 24.26 WER on our benchmark using wav2vec-bert model. We also provide a categorization of errors to support further improvements in future studies.The resources and trained models are released under the CC BY-NC-ND 4.0 license and are publicly available at https://huggingface.co/datasets/aranemini/southern-kurdish-asr
Speech Translation and Metrics in 2026: Findings of the IWSLT Campaign
David Ifeoluwa Adelani | Victor Agostinelli | Antonios Anastasopoulos | Luisa Bentivogli | Ondřej Bojar | Sébastien Bratières | Marine Carpuat | Fabrício Carraro | Roldano Cattoni | Mauro Cettolo | Lizhong Chen | Marcello Federico | Marco Gaido | Mahendra Gupta | HyoJung Han | Ali Hatami | Lewis C. Howe | Dávid Javorský | Yejin Jeon | Marek Kasztelnik | Antoine Laurent | Danni Liu | Nam Luu | Min Ma | Dominik Macháček | Marie Maltais | Evgeny Matusov | John McCrae | Chutong Meng | Chandresh Kumar Maurya | Mohammad Mohammadamini | Yasmin Moslem | Kenton Murray | Satoshi Nakamura | Matteo Negri | Jan Niehues | Atul Kr. Ojha | John E. Ortega | Siqi Ouyang | Sara Papi | Peter Polák | Fabian Retkowski | Stephanny Sánchez | Beatrice Savoldi | Claytone Sikasote | Matthias Sperber | Sebastian Stüker | Katsuhito Sudoh | Marie Tahon | Marco Turchi | Alexander Waibel | Patrick Wilken | Rodolfo Joel Zevallos | Vilem Zouhar | Maike Züfle
Proceedings of the 23rd International Conference on Spoken Language Translation (IWSLT 2026)
David Ifeoluwa Adelani | Victor Agostinelli | Antonios Anastasopoulos | Luisa Bentivogli | Ondřej Bojar | Sébastien Bratières | Marine Carpuat | Fabrício Carraro | Roldano Cattoni | Mauro Cettolo | Lizhong Chen | Marcello Federico | Marco Gaido | Mahendra Gupta | HyoJung Han | Ali Hatami | Lewis C. Howe | Dávid Javorský | Yejin Jeon | Marek Kasztelnik | Antoine Laurent | Danni Liu | Nam Luu | Min Ma | Dominik Macháček | Marie Maltais | Evgeny Matusov | John McCrae | Chutong Meng | Chandresh Kumar Maurya | Mohammad Mohammadamini | Yasmin Moslem | Kenton Murray | Satoshi Nakamura | Matteo Negri | Jan Niehues | Atul Kr. Ojha | John E. Ortega | Siqi Ouyang | Sara Papi | Peter Polák | Fabian Retkowski | Stephanny Sánchez | Beatrice Savoldi | Claytone Sikasote | Matthias Sperber | Sebastian Stüker | Katsuhito Sudoh | Marie Tahon | Marco Turchi | Alexander Waibel | Patrick Wilken | Rodolfo Joel Zevallos | Vilem Zouhar | Maike Züfle
Proceedings of the 23rd International Conference on Spoken Language Translation (IWSLT 2026)
This paper reports on the outcomes of the shared tasks organized as part of the 23rd International Workshop on Spoken Language Translation (IWSLT). The workshop covered ten major challenges in spoken language translation, including speech-to-text translation for both high-resource and low-resource language pairs, customized speech translation, speech generation, instruction-following speech processing, and the evaluation of speech translation systems. The shared tasks received strong participation, with more than 30 teams submitting runs. This year’s edition broadened the range of tasks, placing particular emphasis on speech generation and evaluation metrics.
LIUM Submission for IWSLT 2026 Low-resource Speech Translation Track
Mohammad Mohammadamini | Marie Tahon
Proceedings of the 23rd International Conference on Spoken Language Translation (IWSLT 2026)
Mohammad Mohammadamini | Marie Tahon
Proceedings of the 23rd International Conference on Spoken Language Translation (IWSLT 2026)
This paper describes the LIUM submission to the IWSLT 2026 low-resource speech translation track. It proposes different data augmentation methods for low-resource speech-to-text translation, including two main pipelines: pseudo-labeling and speech synthesis. The goal is to generate parallel speech data in low-resource scenarios without relying on human-annotated speech translation data. Our submission focuses on Central Kurdish–English language pairs. The objective of this work is to explore the advantages and limitations of each data augmentation method. Our best results are obtained using the pseudo-labeling pipeline, achieving a BLEU score of 25.73 on the development set and 21.09 on the test set for Central Kurdish–English translation.
Fleurs-Badini: Translation and Recording Fleurs Dataset for Badini Variant of Northern Kurdish
Mohammad Mohammadamini | Dilgash Mohammed Salih Tayib | Dezheen H. Abdulazeez | Barzan Hussein Mohammed | Imad Saeed Sadeeq | Aveen Jalal Mohammed | Amera Ismail Melhum | Abuobaida Abdullah Dheyab
Proceedings of the 23rd International Conference on Spoken Language Translation (IWSLT 2026)
Mohammad Mohammadamini | Dilgash Mohammed Salih Tayib | Dezheen H. Abdulazeez | Barzan Hussein Mohammed | Imad Saeed Sadeeq | Aveen Jalal Mohammed | Amera Ismail Melhum | Abuobaida Abdullah Dheyab
Proceedings of the 23rd International Conference on Spoken Language Translation (IWSLT 2026)
Multilingual speech benchmarks such as the FLEURS benchmark have significantly advanced research across a wide range of languages. However, important dialects, including Badini Kurdish, remain underrepresented, limiting bechmarking in automatic speech recognition (ASR) and speech-to-text translation (S2TT). To address this limitation, this study introduces FLEURS-Badini, a dialect-focused extension designed to support research on Northern Kurdish (Badini). The dataset is constructed through a structured process of translation, recording, and validation, resulting in 5,224 utterances paired with their corresponding translated text. The data were collected from 45 speakers. To evaluate the dataset, baseline experiments are conducted using state-of-the-art models for both ASR and S2TT. The results indicate that ASR remains challenging, with the best performance achieved by the W2V-BERT CTC model, reaching a Word Error Rate (WER) of approximately 55% on the test set. Similarly, speech-to-text translation performance is limited, with BLEU scores 6.13 and 5.24 on dev and test sets. Overall, FLEURS-Badini expands multilingual coverage and provides a standardized foundation for evaluating ASR and speech translation systems in the Badini dialect.
Exploring the reusability of Northern Kurdish resources for Badini speech recognition
Mohammad Mohammadamini | Aveen Jalal Mohammed | Barzan Hussein Mohammed | Dezheen H. Abdulazeez | Imad Saeed Sadeeq | Dilgash Mohammed Salih | Amera Ismail Melhum | Abuobaida Abdullah Dheyab
Proceedings of the First Workshop on Dialects in NLP — A Resource Perspective
Mohammad Mohammadamini | Aveen Jalal Mohammed | Barzan Hussein Mohammed | Dezheen H. Abdulazeez | Imad Saeed Sadeeq | Dilgash Mohammed Salih | Amera Ismail Melhum | Abuobaida Abdullah Dheyab
Proceedings of the First Workshop on Dialects in NLP — A Resource Perspective
Badini is a variant of the Kurdish language spoken in the Duhok province of the Kurdistan Region of Iraq. It is written mainly in a modified version of the Arabic script. Although it shares the same script as Central Kurdish (CKB), it is linguistically classified under the Northern Kurdish (KMR) branch. In this paper, we explore the potential and limitations of Northern Kurdish ASR resources for the Badini variant. Firstly, we transliterate the Common Voice 18 dataset from the Latin script into the modified Arabic script and revised it to align with the orthographic conventions of Badini variant. Additionally, we introduce the first text collection for the Badini variant, containing 14,22 million tokens, which serves as a source for speech synthesis. A third resource developed in this research is a standard speech recognition benchmark recorded by 5 speakers which includes 2 hours and 46 minutes of multi-domain read speech. Results show that combining transliterated and synthetic data significantly improves recognition accuracy, achieving a 6.8% CER and 34% WER. All three resources curated during this research will be made available under the CC BY-NC-ND 4.0 license.
FLEURS-Kobani: Extending FLEURS dataset for Northern Kurdish
Daban Q. Jaff | Mohammad Mohammadamini
Proceedings of the First Workshop on Dialects in NLP — A Resource Perspective
Daban Q. Jaff | Mohammad Mohammadamini
Proceedings of the First Workshop on Dialects in NLP — A Resource Perspective
We present FLEURS-Kobani, a Northern Kurdish (ISO 639-3 KMR) spoken extension of FLEURS benchmark. Although FLEURS offers n-way parallel speech for 100+ languages, Northern Kurdish is absent, limits benchmarking automatic speech recognition and speech translation tasks for this language. The FLEURS-Kobani dataset consists of 5,162 validated utterances, totaling 18 hours and 24 minutes. As baselines, we fine-tuned Whisper v3-large for ASR and E2E S2TT. A two-stage fine-tuning strategy (Common Voice→FLEURS-Kobani) yields the best ASR performance (WER 28.11, CER 9.84 on test). For end-to-end S2TT (KMR→EN), Whisper achieves 8.68 BLEU on test; we additionally report pivot-derived targets and a cascaded S2TT setup. FLEURS-Kobani provides the first Northern Kurdish public benchmark for evaluation of ASR, S2TT and S2ST tasks. The data can be accessed (LINK IS BLANK DUE TO REVISION RULES) under CC BY 4.0 license.
2025
Kuvost: A Large-Scale Human-Annotated English to Central Kurdish Speech Translation Dataset Driven from English Common Voice
Mohammad Mohammadamini | Daban Jaff | Sara Jamal | Ibrahim Ahmed | Hawkar Omar | Darya Sabr | Marie Tahon | Antoine Laurent
Proceedings of the 22nd International Conference on Spoken Language Translation (IWSLT 2025)
Mohammad Mohammadamini | Daban Jaff | Sara Jamal | Ibrahim Ahmed | Hawkar Omar | Darya Sabr | Marie Tahon | Antoine Laurent
Proceedings of the 22nd International Conference on Spoken Language Translation (IWSLT 2025)
In this paper, we introduce the Kuvost, a large-scale English to Central Kurdish speech-to-text-translation (S2TT) dataset. This dataset includes 786k utterances derived from Common Voice 18, translated and revised by 230 volunteers into Central Kurdish. Encompassing 1,003 hours of translated speech, this dataset can play a groundbreaking role for Central Kurdish, which severely lacks public-domain resources for speech translation. Following the dataset division in Common Voice, there are 298k, 6,226, and 7,253 samples in the train, development, and test sets, respectively. The dataset is evaluated on end-to-end English-to-Kurdish S2TT using Whisper V3 Large and SeamlessM4T V2 Large models. The dataset is available under the Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License https://huggingface.co/datasets/aranemini/kuvost.
2024
RoboVox: A Single/Multi-channel Far-field Speaker Recognition Benchmark for a Mobile Robot
Mohammad Mohammadamini | Driss Matrouf | Michael Rouvier | Jean-Francois Bonastre | Romain Serizel | Theophile Gonos
Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024)
Mohammad Mohammadamini | Driss Matrouf | Michael Rouvier | Jean-Francois Bonastre | Romain Serizel | Theophile Gonos
Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024)
In this paper, we introduce a new far-field speaker recognition benchmark called RoboVox. RoboVox is a French corpus recorded by a mobile robot. The files are recorded from different distances under severe acoustical conditions with the presence of several types of noise and reverberation. In addition to noise and reverberation, the robot’s internal noise acts as an extra additive noise. RoboVox can be used for both single-channel and multi-channel speaker recognition. In the evaluation protocols, we are considering both cases. The obtained results demonstrate a significant decline in performance in far-filed speaker recognition and urge the community to further research in this domain
2022
Far-Field Speaker Recognition Benchmark Derived From The DiPCo Corpus
Mickael Rouvier | Mohammad Mohammadamini
Proceedings of the Thirteenth Language Resources and Evaluation Conference
Mickael Rouvier | Mohammad Mohammadamini
Proceedings of the Thirteenth Language Resources and Evaluation Conference
In this paper, we present a far-field speaker verification benchmark derived from the publicly-available DiPCo corpus. This corpus comprise three different tasks that involve enrollment and test conditions with single- and/or multi-channels recordings. The main goal of this corpus is to foster research in far-field and multi-channel text-independent speaker verification. Also, it can be used for other speaker recognition tasks such as dereverberation, denoising and speech enhancement. In addition, we release a Kaldi and SpeechBrain system to facilitate further research. And we validate the evaluation design with a single-microphone state-of-the-art speaker recognition system (i.e. ResNet-101). The results show that the proposed tasks are very challenging. And we hope these resources will inspire the speech community to develop new methods and systems for this challenging domain.
Search
Fix author
Co-authors
- Marie Tahon 6
- Daban Jaff 3
- Antoine Laurent 3
- Dezheen H. Abdulazeez 2
- Abuobaida Abdullah Dheyab 2
- Amera Ismail Melhum 2
- Aveen Jalal Mohammed 2
- Barzan Hussein Mohammed 2
- Imad Saeed Sadeeq 2
- David Ifeoluwa Adelani 1
- Victor Agostinelli 1
- Ibrahim Ahmed 1
- Antonios Anastasopoulos 1
- Luisa Bentivogli 1
- Ondřej Bojar 1
- Jean-François Bonastre 1
- Sébastien Bratières 1
- Marine Carpuat 1
- Fabrício Carraro 1
- Roldano Cattoni 1
- Mauro Cettolo 1
- Lizhong Chen 1
- Josep M. Crego 1
- Marcello Federico 1
- Marco Gaido 1
- Theophile Gonos 1
- Mahendra Gupta 1
- HyoJung Han 1
- Ali Hatami 1
- Lewis C. Howe 1
- Sara Jamal 1
- Dávid Javorský 1
- Yejin Jeon 1
- Marek Kasztelnik 1
- Danni Liu 1
- Nam Luu 1
- Min Ma 1
- Dominik Macháček 1
- Marie Maltais 1
- Driss Matrouf 1
- Evgeny Matusov 1
- Chandresh Kumar Maurya 1
- John Philip McCrae 1
- Chutong Meng 1
- Yasmin Moslem 1
- Kenton Murray 1
- Satoshi Nakamura 1
- Matteo Negri 1
- Jan Niehues 1
- Atul Kr. Ojha 1
- Hawkar Omar 1
- John E. Ortega 1
- Siqi Ouyang 1
- Sara Papi 1
- Peter Polák 1
- Fabian Retkowski 1
- Michael Rouvier 1
- Mickael Rouvier 1
- Darya Sabr 1
- Dilgash Mohammed Salih 1
- Beatrice Savoldi 1
- Romain Serizel 1
- Meysam Shamsi 1
- Claytone Sikasote 1
- Matthias Sperber 1
- Sebastian Stüker 1
- Katsuhito Sudoh 1
- Stephanny Sánchez 1
- Dilgash Mohammed Salih Tayib 1
- Marco Turchi 1
- Alexander Waibel 1
- Patrick Wilken 1
- Rodolfo Zevallos 1
- Vilém Zouhar 1
- Maike Züfle 1