Nana Hou
2026
Chronological Thinking in Full-Duplex Spoken Dialogue Language Models
Donghang Wu | Haoyang Zhang | Chen Chen | Tianyu Zhang | Fei Tian | Xuerui Yang | Gang Yu | Hexin Liu | Nana Hou | Yuchen Hu | Eng Siong Chng
Proceedings of the 27th Annual Meeting of the Special Interest Group on Discourse and Dialogue
Donghang Wu | Haoyang Zhang | Chen Chen | Tianyu Zhang | Fei Tian | Xuerui Yang | Gang Yu | Hexin Liu | Nana Hou | Yuchen Hu | Eng Siong Chng
Proceedings of the 27th Annual Meeting of the Special Interest Group on Discourse and Dialogue
Recent advances in spoken dialogue language models (SDLMs) reflect growing interest in shifting from turn-based to full-duplex systems, where the models continuously perceive user speech streams while generating responses. This simultaneous listening and speaking design enables real-time interaction and the agent can handle dynamic conversational behaviors like user barge-in. However, during the listening phase, existing systems keep the agent idle by repeatedly predicting the silence token, which departs from human behavior: we usually engage in lightweight thinking during conversation rather than remaining absent-minded. Inspired by this, we propose Chronological Thinking, an on-the-fly conversational thinking mechanism that aims to improve response quality in full-duplex SDLMs. Specifically, chronological thinking presents a paradigm shift from conventional LLM thinking approaches, such as Chain-of-Thought, purpose-built for streaming acoustic input. (1) Strictly causal: the agent reasons incrementally while listening, updating internal hypotheses only from past audio with no lookahead. (2) No additional latency: reasoning is amortized during the listening window; once the user stops speaking, the agent halts thinking and begins speaking without further delay. Experiments demonstrate the effectiveness of chronological thinking through both objective metrics and human evaluations show consistent improvements in response quality. Furthermore, chronological thinking robustly handles conversational dynamics and attains competitive performance on full-duplex interaction metrics.
Evaluating the Expressive Appropriateness of Speech in Rich Contexts
Tianrui Wang | Ziyang Ma | Yizhou Peng | Haoyu Wang | Zhikang Niu | Zikang Huang | Yihao Wu | Yi-Wen Chao | Yu Jiang | Yuheng Lu | Guanrou Yang | Xuanchen Li | Hexin Liu | Chunyu Qiang | Cheng Gong | Yifan Yang | Tianchi Liu | Junyu Wang | Nana Hou | Meng Ge | Fuming You | Yang Wei | Zhongqian Sun | Hu Haifeng | Xiaobao Wang | Eng Siong Chng | Xie Chen | Longbiao Wang | Jianwu Dang
Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Tianrui Wang | Ziyang Ma | Yizhou Peng | Haoyu Wang | Zhikang Niu | Zikang Huang | Yihao Wu | Yi-Wen Chao | Yu Jiang | Yuheng Lu | Guanrou Yang | Xuanchen Li | Hexin Liu | Chunyu Qiang | Cheng Gong | Yifan Yang | Tianchi Liu | Junyu Wang | Nana Hou | Meng Ge | Fuming You | Yang Wei | Zhongqian Sun | Hu Haifeng | Xiaobao Wang | Eng Siong Chng | Xie Chen | Longbiao Wang | Jianwu Dang
Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Evaluating expressive speech remains challenging, as existing methods mainly assess emotional intensity and overlook whether a speech sample is expressively appropriate for its contextual setting. This limitation hinders reliable evaluation of speech systems used in narrative-driven and interactive applications, such as audiobooks and conversational agents. We introduce CEAEval, a Context-rich framework for Evaluating Expressive Appropriateness in speech, which assesses whether a speech sample expressively aligns with the underlying communicative intent implied by its discourse-level narrative context. To support this task, we construct CEAEval-D, the first context-rich speech dataset with real human performances in Mandarin conversational speech, providing narrative descriptions together with fifteen dimensions of human annotations covering expressive attributes and expressive appropriateness. We further develop CEAEval-M, a model that integrates knowledge distillation, planner-based multi-model collaboration, adaptive audio attention bias, and reinforcement learning to perform context-rich expressive appropriateness evaluation. Experiments on a human-annotated test set demonstrate that CEAEval-M substantially outperforms existing speech evaluation and analysis systems.
Search
Fix author
Co-authors
- Eng Siong Chng 2
- Hexin Liu 2
- Yi-Wen Chao 1
- Chen Chen 1
- Xie Chen 1
- Jianwu Dang 1
- Meng Ge 1
- Cheng Gong 1
- Hu Haifeng 1
- Yuchen Hu 1
- Zikang Huang 1
- Yu Jiang 1
- Xuanchen Li 1
- Tianchi Liu 1
- Yuheng Lu 1
- Ziyang Ma 1
- Zhikang Niu 1
- Yizhou Peng 1
- Chunyu Qiang 1
- Zhongqian Sun 1
- Fei Tian 1
- Haoyu Wang 1
- Junyu Wang 1
- Longbiao Wang 1
- Tianrui Wang 1
- Xiaobao Wang 1
- Yang Wei 1
- Donghang Wu 1
- Yihao Wu 1
- Guanrou Yang 1
- Xuerui Yang 1
- Yifan Yang 1
- Fuming You 1
- Gang Yu 1
- Haoyang Zhang 1
- Tianyu Zhang 1