Xianzhi Yu
2026
Analytical FFN-to-MoE Restructuring via Activation Pattern Analysis
Zehua Pei | Hui-Ling Zhen | Lancheng Zou | Xianzhi Yu | Wulong Liu | Sinno Jialin Pan | Mingxuan Yuan | Bei Yu
Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Zehua Pei | Hui-Ling Zhen | Lancheng Zou | Xianzhi Yu | Wulong Liu | Sinno Jialin Pan | Mingxuan Yuan | Bei Yu
Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Scaling large language models (LLMs) improves performance but significantly increases inference costs, with feed-forward networks (FFNs) consuming the majority of computational resources. While Mixture-of-Experts (MoE) architectures can reduce this cost through sparse activation, restructuring existing dense models into MoEs typically requires extensive retraining on hundreds of billions of tokens.We propose an analytical post-training framework that rapidly restructures FFNs into sparse MoE architectures using only a small calibration dataset. The method analyzes neuron activation patterns to partition neurons into always-active shared experts and conditionally activated routed experts, then constructs a router analytically from representative neuron statistics, enabling immediate deployment or optional lightweight fine-tuning. This approach applies both to dense models and recursively to existing MoE models for hierarchical sparsity.Experiments demonstrate up to 1.17× speedup in compute-bound scenarios with only minutes of processing and 2k-sample fine-tuning, outperforming methods requiring orders of magnitude more resources.
Benchmarking Post-Training Quantization of Large Language Models under Microscaling Floating Point Formats
Manyi Zhang | Ji-Fu Li | Zhongao Sun | Haoli Bai | Hui-Ling Zhen | Zhenhua Dong | Xianzhi Yu
Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Manyi Zhang | Ji-Fu Li | Zhongao Sun | Haoli Bai | Hui-Ling Zhen | Zhenhua Dong | Xianzhi Yu
Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Microscaling Floating-Point (MXFP) has emerged as a promising low-precision format for large language models (LLMs). Despite various post-training quantization (PTQ) algorithms being proposed, they mostly focus on integer quantization, while their applicability and behavior under MXFP formats remain largely unexplored. To address this gap, this work conducts a systematic investigation of PTQ under MXFP formats, encompassing over 7 PTQ algorithms, 15 evaluation benchmarks, and 3 LLM families. The key findings include: 1) MXFP8 consistently achieves near-lossless performance, while MXFP4 introduces substantial accuracy degradation and remains challenging; 2) PTQ effectiveness under MXFP depends strongly on format compatibility, with some algorithmic paradigms being consistently more effective than others; 3) PTQ performance exhibits highly consistent trends across model families and modalities, in particular, quantization sensitivity is dominated by the language model rather than the vision encoder in multimodal LLMs; 4) The scaling factor of quantization is a critical error source in MXFP4, and a simple pre-scale optimization strategy can significantly mitigate its impact. Together, these results provide practical guidance on adapting existing PTQ methods to MXFP quantization.
Unleashing Low-Bit Inference on Ascend NPUs: A Comprehensive Evaluation of HiFloat Formats
Pengxiang Zhao | Hui-Ling Zhen | Xing Li | Han Bao | Weizhe Lin | Zhiyuan Yang | Yu Zi Wei | Xin Wang | Mingxuan Yuan | Xianzhi Yu | Zhenhua Dong
Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (ACL 2026)
Pengxiang Zhao | Hui-Ling Zhen | Xing Li | Han Bao | Weizhe Lin | Zhiyuan Yang | Yu Zi Wei | Xin Wang | Mingxuan Yuan | Xianzhi Yu | Zhenhua Dong
Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (ACL 2026)
As LLMs scale, low-bit floating-point formats like MXFP and NVFP4 offer new opportunities for precision and efficiency. In this work, we evaluate HiFloat (HiF8 and HiF4), a family of formats tailored for Ascend NPUs. Through rigorous comparison across weight-activation and KV-cache tasks, we provide three key insights: (1) INT8 suits narrow-range data, while floating-point formats excel with high-variance data; (2) in 4-bit regimes, HiF4’s hierarchical scaling prevents the accuracy collapse seen in integer formats; and (3) HiFloat is fully compatible with state-of-the-art post-training quantization frameworks. Overall, HiFloat provides a solution for high-efficiency LLM inference on NPUs.
2025
Faster and Better LLMs via Latency-Aware Test-Time Scaling
Zili Wang | Tianyu Zhang | Haoli Bai | Lu Hou | Xianzhi Yu | Wulong Liu | Shiming Xiang | Lei Zhu
Findings of the Association for Computational Linguistics: EMNLP 2025
Zili Wang | Tianyu Zhang | Haoli Bai | Lu Hou | Xianzhi Yu | Wulong Liu | Shiming Xiang | Lei Zhu
Findings of the Association for Computational Linguistics: EMNLP 2025
Test-Time Scaling (TTS) has proven effective in improving the performance of Large Language Models (LLMs) during inference. However, existing research has overlooked the efficiency of TTS from a latency-sensitive perspective. Through a latency-aware evaluation of representative TTS methods, we demonstrate that a compute-optimal TTS does not always result in the lowest latency in scenarios where latency is critical. To address this gap and achieve latency-optimal TTS, we propose two key approaches by optimizing the concurrency configurations: (1) branch-wise parallelism, which leverages multiple concurrent inference branches, and (2) sequence-wise parallelism, enabled by speculative decoding. By integrating these two approaches and allocating computational resources properly to each, our latency-optimal TTS enables a 32B model to reach 82.3% accuracy on MATH-500 within 1 minute and a smaller 3B model to achieve 72.4% within 10 seconds. Our work emphasizes the importance of latency-aware TTS and demonstrates its ability to deliver both speed and accuracy in latency-sensitive scenarios.
2022
HW-TSC’s Submission for the WMT22 Efficiency Task
Hengchao Shang | Ting Hu | Daimeng Wei | Zongyao Li | Xianzhi Yu | Jianfei Feng | Ting Zhu | Lizhi Lei | Shimin Tao | Hao Yang | Ying Qin | Jinlong Yang | Zhiqiang Rao | Zhengzhe Yu
Proceedings of the Seventh Conference on Machine Translation (WMT)
Hengchao Shang | Ting Hu | Daimeng Wei | Zongyao Li | Xianzhi Yu | Jianfei Feng | Ting Zhu | Lizhi Lei | Shimin Tao | Hao Yang | Ying Qin | Jinlong Yang | Zhiqiang Rao | Zhengzhe Yu
Proceedings of the Seventh Conference on Machine Translation (WMT)
This paper presents the submission of Huawei Translation Services Center (HW-TSC) to WMT 2022 Efficiency Shared Task. For this year’s task, we still apply sentence-level distillation strategy to train small models with different configurations. Then, we integrate the average attention mechanism into the lightweight RNN model to pursue more efficient decoding. We tried adding a retrain step to our 8-bit and 4-bit models to achieve a balance between model size and quality. We still use Huawei Noah’s Bolt for INT8 inference and 4-bit storage. Coupled with Bolt’s support for batch inference and multi-core parallel computing, we finally submit models with different configurations to the CPU latency and throughput tracks to explore the Pareto frontiers.
Search
Fix author
Co-authors
- Hui-Ling Zhen 3
- Haoli Bai 2
- Zhenhua Dong 2
- Wulong Liu 2
- Mingxuan Yuan 2
- Han Bao 1
- Jianfei Feng 1
- Lu Hou 1
- Ting Hu 1
- Lizhi Lei 1
- Ji-Fu Li 1
- Xing Li 1
- Zongyao Li 1
- Weizhe Lin 1
- Sinno Jialin Pan 1
- Zehua Pei 1
- Ying Qin 1
- Zhiqiang Rao 1
- Hengchao Shang 1
- Zhongao Sun 1
- Shimin Tao 1
- Xin Wang 1
- Zili Wang 1
- Daimeng Wei 1
- Yu Zi Wei 1
- Shiming Xiang 1
- Hao Yang 1
- Jinlong Yang 1
- Zhiyuan Yang 1
- Bei Yu 1
- Zhengzhe Yu 1
- Manyi Zhang 1
- Tianyu Zhang 1
- Pengxiang Zhao 1
- Lei Zhu 1
- Ting Zhu 1
- Lancheng Zou 1