Zaokere Kadeer
Also published as: 卡德尔 早克热
2026
Synergizing Semantic Anchors and Ordinal Smoothed Cross-Entropy for Speech Fluency Classification
Mulati Kahaer | Sirajahmat Ruzmamat | XuDong Pang | Subinuer Maimaitituerxun | Zaokere Kadeer | Abudurexiti Reheman | Wenwen Lu | Panpan Zheng | Aishan Wumaier
Findings of the Association for Computational Linguistics: ACL 2026
Mulati Kahaer | Sirajahmat Ruzmamat | XuDong Pang | Subinuer Maimaitituerxun | Zaokere Kadeer | Abudurexiti Reheman | Wenwen Lu | Panpan Zheng | Aishan Wumaier
Findings of the Association for Computational Linguistics: ACL 2026
Speech fluency is a core indicator of second language proficiency and a critical component of Computer-Assisted Pronunciation Training (CAPT) systems. Accurate assessment requires models to perceive both macroscopic speech flow trends and microscopic local anomalies. However, existing methods struggle to bridge the semantic gap between static expert priors and dynamic temporal representations, while often overlooking the inherent ordinal nature of fluency scores. To address these challenges, we first construct a set of expert features targeting fluency disruptions and rhythmic regularity to provide explicit linguistic priors. Building on this, we propose the Multimodal Multi-Stream Fusion Classification (MMSFC) network. It employs a Mutual Cross-Attention (MCA) mechanism that leverages these expert features as “semantic anchors” to actively guide Whisper’s temporal representations and integrate decoder contexts, achieving deep interaction between global priors and local dynamics. Furthermore, we propose the Ordinal Smoothed Cross-Entropy (OSCE) loss. By constructing distance-aware soft target distributions coupled with confidence-adaptive smoothing and boundary enhancement, OSCE explicitly models ordinal relationships to resolve boundary ambiguity. Experiments on SpeechOcean762 show MMSFC achieves 83.40% accuracy, significantly outperforming strong baselines. Notably, OSCE also demonstrates superior generalization potential in cross-domain CV and NLP tasks. Our code is available at https://github.com/speech26ai/MMSFCCode.
2021
基于改进Conformer的新闻领域端到端语音识别(End-to-End Speech Recognition in News Field based on Conformer)
Jimin Zhang (张济民) | Zaokere Kadeer (早克热卡德尔) | Yunfei Shen (申云飞) | Aishan Wumaier (艾山吾买尔) | Liejun Wang (汪烈军)
Proceedings of the 20th Chinese National Conference on Computational Linguistics
Jimin Zhang (张济民) | Zaokere Kadeer (早克热卡德尔) | Yunfei Shen (申云飞) | Aishan Wumaier (艾山吾买尔) | Liejun Wang (汪烈军)
Proceedings of the 20th Chinese National Conference on Computational Linguistics
目前,开源的中文语音识别数据集多为面向通用领域,缺少面向新闻领域的开源语音识别语料库,因此本文构建了面向新闻领域的中文语音识别数据集CHNEWSASR并使用ESPNET-0.9.6框架的RNN、Transformer和Conformer等模型对数据集的有效性进行了验证,实验表明本文所构建的语料在最好的模型上CER为4.8%,SER为39.4%。由于新闻联播主持人说话语速相对较快,本文构建的数据集文本平均长度为28个字符是Aishell1数据集文本平均长度的2倍,且以往的研究中训练目标函数通常为基于字或词水平,缺乏明确的句子水平关系,因此本文提出了一个句子层级的一致性模块与Conformer模型结合直接减少源语音和目标文本的表示差异,在开源的Aishell1数据集上其CER降低0.4%,SER降低2%;在CHNEWSASR数据集上其CER降低0.9%,SER降低3%,实验结果表明该方法不提升模型参数量的前提下能有效提升语音识别的质量。