Wataru Nakata
2026
DialogueSidon: Recovering Full-Duplex Dialogue Tracks from In-the-Wild Dialogue Audio
Wataru Nakata | Yuki Saito | Kazuki Yamauchi | Emiru Tsunoo | Hiroshi Saruwatari
Proceedings of the 27th Annual Meeting of the Special Interest Group on Discourse and Dialogue
Wataru Nakata | Yuki Saito | Kazuki Yamauchi | Emiru Tsunoo | Hiroshi Saruwatari
Proceedings of the 27th Annual Meeting of the Special Interest Group on Discourse and Dialogue
Full-duplex dialogue audio, in which each speaker is recorded on a separate track, is an important resource for spoken dialogue research, but is difficult to collect at scale. Most in-the-wild two-speaker dialogue is available only as degraded monaural mixtures, making it unsuitable for systems requiring clean speaker-wise signals. We propose DialogueSidon, a model for joint restoration and separation of degraded monaural two-speaker dialogue audio. DialogueSidon combines a variational autoencoder (VAE) operates on the speech self-supervised learning (SSL) model feature, which compresses SSL model features into a compact latent space, with a diffusion-based latent predictor that recovers speaker-wise latent representations from the degraded mixture. Experiments on English, multilingual, and in-the-wild dialogue datasets show that DialogueSidon substantially improves intelligibility and separation quality over a baseline, while also achieving much faster inference.
J-CHAT: Japanese Large-scale Spoken Dialogue Corpus for Spoken Dialogue Language Modeling
Wataru Nakata | Kentaro Seki | Hitomi Yanaka | Yuki Saito | Shinnosuke Takamichi | Hiroshi Saruwatari
Proceedings of the Fifteenth Language Resources and Evaluation Conference
Wataru Nakata | Kentaro Seki | Hitomi Yanaka | Yuki Saito | Shinnosuke Takamichi | Hiroshi Saruwatari
Proceedings of the Fifteenth Language Resources and Evaluation Conference
Spoken dialogue is essential for human-AI interactions, providing expressive capabilities beyond text. Developing effective spoken dialogue systems (SDSs) requires large-scale, high-quality, and diverse spoken dialogue corpora. However, existing datasets are often limited in size, spontaneity, or linguistic coherence. To address these limitations, we introduce J-CHAT, a 76,000-hour open-source Japanese spoken dialogue corpus. Constructed using an automated, language-independent methodology, J-CHAT ensures acoustic cleanliness, diversity, and natural spontaneity. The corpus is built from YouTube and podcast data, with extensive filtering and denoising to enhance quality. Experimental results with generative spoken dialogue language models trained on J-CHAT demonstrate its effectiveness for SDS development. By providing a robust foundation for training advanced dialogue models, we anticipate that J-CHAT will drive progress in human-AI dialogue research and applications.