Yuya Chiba

Other people with similar names: Yuya Chiba

Unverified author pages with similar names: Yuya Chiba


2026

Verbalization, the process of expressing one’s internal states in words, has been shown to deepen self-understanding and improve well-being. However, this can be difficult to achieve for some individuals. One approach to supporting verbalization is Focusing-Oriented Psychotherapy. Focusing facilitates verbalization by guiding a speaker’s attention to their felt sense, an internal state that has not yet been verbalized, and helping them to find the words that appropriately express this felt sense through dialogue. Toward developing Focusing dialogue systems that support user verbalization, we constructed a dataset containing Focusing dialogues between professionally trained listeners and speakers. This dataset consists of 50 dialogues (762 minutes, 17,986 utterances) and contains ratings of verbalization progress, subjective evaluations, and dialogue act (DA) annotations. To clarify the characteristics of Focusing dialogues, we analyzed associations among verbalization progress, subjective evaluations, and DA usage. Based on these analyses, we discussed design implications for Focusing dialogue systems.
Voice Activity Projection (VAP) has been actively studied to enable natural turn-taking in spoken dialogue systems, relying primarily on acoustic features. Visual cues such as head movements are also known to contribute to turn-taking prediction; however, camera-based approaches are affected by placement and lighting conditions and are not always reliably available to dialogue systems. As a camera-independent approach for directly capturing head motion, earable devices offer a promising solution. In this study, we propose Sensor-Augmented VAP, a framework that integrates in-ear inertial measurement unit (IMU) signals with a pre-trained VAP model via a lightweight residual fusion module. To validate our proposed method, we collected a dataset pairing conversational audio with in-ear IMU data, comprising 12 dyadic Japanese dialogues recorded using microphones and earbuds. Experiments in speaker-independent and speaker-dependent settings demonstrate that IMU fusion consistently improves weighted F1 score for shift detection and reduces VAP loss over the audio-only baseline. These results confirm that head-motion cues are effective for enhancing turn-taking prediction.
Dialogic reading, which involves interactive exchanges between a parent and a child during picture book reading, has been shown to effectively promote children’s language development. While many support systems for picture book reading have been developed to reduce the burden on parents, existing systems are not yet capable of handling dialogic reading, which requires dynamic parent-child interaction. To develop conversational agents capable of dialogic reading, we constructed a multimodal corpus of parent-child picture-book reading dialogues. The corpus comprises recordings from 36 Japanese parent-child pairs taken during actual picture book reading sessions. In this study, we annotated the corpus with dialogue acts relevant to parent-child communication and categorized the types of quizzes and questions used in the sessions, analyzing the linguistic aspects of parent-child interaction during dialogic reading. After dividing the dialogues into two groups based on the proportion of the child’s utterances, our analyses revealed that dialogue systems should adapt their interaction strategies according to individual child characteristics.