DialogueSidon: Recovering Full-Duplex Dialogue Tracks from In-the-Wild Dialogue Audio
Wataru Nakata, Yuki Saito, Kazuki Yamauchi, Emiru Tsunoo, Hiroshi Saruwatari
Abstract
Full-duplex dialogue audio, in which each speaker is recorded on a separate track, is an important resource for spoken dialogue research, but is difficult to collect at scale. Most in-the-wild two-speaker dialogue is available only as degraded monaural mixtures, making it unsuitable for systems requiring clean speaker-wise signals. We propose DialogueSidon, a model for joint restoration and separation of degraded monaural two-speaker dialogue audio. DialogueSidon combines a variational autoencoder (VAE) operates on the speech self-supervised learning (SSL) model feature, which compresses SSL model features into a compact latent space, with a diffusion-based latent predictor that recovers speaker-wise latent representations from the degraded mixture. Experiments on English, multilingual, and in-the-wild dialogue datasets show that DialogueSidon substantially improves intelligibility and separation quality over a baseline, while also achieving much faster inference.- Anthology ID:
- 2026.sigdial-1.1
- Volume:
- Proceedings of the 27th Annual Meeting of the Special Interest Group on Discourse and Dialogue
- Month:
- August
- Year:
- 2026
- Address:
- Atlanta, Georgia, USA
- Editors:
- Jinho D. Choi, Yun-Nung Chen, Kotaro Funakoshi, Ali Emami
- Venue:
- SIGDIAL
- SIG:
- SIGDIAL
- Publisher:
- Association for Computational Linguistics
- Note:
- Pages:
- 1–12
- Language:
- URL:
- https://preview.aclanthology.org/paragraph-normalization/2026.sigdial-1.1/
- DOI:
- Cite (ACL):
- Wataru Nakata, Yuki Saito, Kazuki Yamauchi, Emiru Tsunoo, and Hiroshi Saruwatari. 2026. DialogueSidon: Recovering Full-Duplex Dialogue Tracks from In-the-Wild Dialogue Audio. In Proceedings of the 27th Annual Meeting of the Special Interest Group on Discourse and Dialogue, pages 1–12, Atlanta, Georgia, USA. Association for Computational Linguistics.
- Cite (Informal):
- DialogueSidon: Recovering Full-Duplex Dialogue Tracks from In-the-Wild Dialogue Audio (Nakata et al., SIGDIAL 2026)
- PDF:
- https://preview.aclanthology.org/paragraph-normalization/2026.sigdial-1.1.pdf