EDBT 2026 Demo / reviewers in the wild / expert
Yuya Chiba
dblp:120/6490
· DBLP profile ↗
23ranked-venue papers
8as first author
14since 2021 · last 2026
0000-0003-1987-4368ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 19 · 7 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 2 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Construction and Analysis of Japanese Parent-Child Dialogic Reading Corpus for Conversational Agents
Yuko Nakagi, Yuya Chiba, Sanae Fujita, Shoko Araki |
LREC | 2 |
| 2026 | Sensor-Augmented Voice Activity Projection for Enhancing Turn-Taking PredictionabstractVoice Activity Projection (VAP) has been actively studied to enable natural turn-taking in spoken dialogue systems, relying primarily on acoustic features. Visual cues such as head movements are also known to contribute to turn-taking prediction; however, camera-based approaches are affected by placement and lighting conditions and are not always reliably available to dialogue systems. As a camera-independent approach for directly capturing head motion, earable devices offer a promising solution. In this study, we propose Sensor-Augmented VAP, a framework that integrates in-ear inertial measurement unit (IMU) signals with a pre-trained VAP model via a lightweight residual fusion module. To validate our proposed method, we collected a dataset pairing conversational audio with in-ear IMU data, comprising 12 dyadic Japanese dialogues recorded using microphones and earbuds. Experiments in speaker-independent and speaker-dependent settings demonstrate that IMU fusion consistently improves weighted F1 score for shift detection and reduces VAP loss over the audio-only baseline. These results confirm that head-motion cues are effective for enhancing turn-taking prediction. Satoki Hamanaka, Yasue Kishino, Yuiko Tsunomori, Shin Mizutani, Yuya Chiba, Tadashi Okoshi, Jin Nakazawa |
SIGDIAL | 5 |
| 2026 | From Felt Sense to Words: Construction and Analysis of a Focusing Dialogue Dataset for VerbalizationabstractVerbalization, the process of expressing one’s internal states in words, has been shown to deepen self-understanding and improve well-being. However, this can be difficult to achieve for some individuals. One approach to supporting verbalization is Focusing-Oriented Psychotherapy. Focusing facilitates verbalization by guiding a speaker’s attention to their felt sense, an internal state that has not yet been verbalized, and helping them to find the words that appropriately express this felt sense through dialogue. Toward developing Focusing dialogue systems that support user verbalization, we constructed a dataset containing Focusing dialogues between professionally trained listeners and speakers. This dataset consists of 50 dialogues (762 minutes, 17,986 utterances) and contains ratings of verbalization progress, subjective evaluations, and dialogue act (DA) annotations. To clarify the characteristics of Focusing dialogues, we analyzed associations among verbalization progress, subjective evaluations, and DA usage. Based on these analyses, we discussed design implications for Focusing dialogue systems. Yuiko Tsunomori, Yuya Chiba, Yosuke Koshikawa |
SIGDIAL | 2 |
| 2025 | Predictive ASR and Turn-taking Prediction at Once: Towards More Responsive Spoken Dialog SystemabstractSpoken dialog systems usually wait for users to finish speaking before generating responses, resulting in response delays. A possible solution for reducing the response delay is to predict future words and/or turn-ends while the user is speaking. To realize this, we propose a method to jointly perform predictive automatic speech recognition and turn-taking prediction. Our model receives partial utterances as input and performs speech recognition, future word prediction, and turntaking prediction via autoregressive decoding. It enables turntaking prediction based on prosodic and linguistic cues of observed partial utterances and predicted future linguistic cues. We also incorporate dialogue contexts to improve the performance. Experiments on the Switchboard corpus showed that our multi-task model outperforms a single-task model in turn-taking prediction. We found that conditioning turn-taking prediction on predicted words improved performance when words were correctly predicted. Ryo Fukuda, Takatomo Kano, Naohiro Tawara, Marc Delcroix, Atsunori Ogawa, Yuya Chiba, Atsushi Ando |
ASRU | 6 |
| 2025 | Investigating the Impact of Incremental Processing and Voice Activity Projection on Spoken Dialogue SystemsabstractThe naturalness of responses in spoken dialogue systems has been significantly improved by the introduction of large language models (LLMs), although many challenges remain until human-like turn-taking can be achieved. A turn-taking model called Voice Activity Projection (VAP) is gaining attention because it can be trained in an unsupervised manner using the spoken dialogue data between two speakers. For such a turn-taking model to be fully effective, systems must initiate response generation as soon as a turn-shift is detected. This can be achieved by incremental response generation, which reduces the delay before the system responds. Incremental response generation is done using partial speech recognition results while user speech is incrementally processed. Combining incremental response generation with VAP-based turn-taking will enable spoken dialogue systems to achieve faster and more natural turn-taking. However, their effectiveness remains unclear because they have not yet been evaluated in real-world systems. In this study, we developed spoken dialogue systems that incorporate incremental response generation and VAP-based turn-taking and evaluated their impact on task success and dialogue satisfaction through user assessments. Yuya Chiba, Ryuichiro Higashinaka |
COLING | 1 |
| 2025 | Improving User Impression of Spoken Dialogue Systems by Controlling Para-linguistic Expression Based on Intimacy
Shoki Kawanishi, Akinori Ito, Yuya Chiba, Takashi Nose |
INTERSPEECH | 3 |
| 2024 | Travel Agency Task Dialogue Corpus: A Multimodal Dataset with Age-Diverse SpeakersabstractWhen individuals communicate, they use different vocabularies, speaking speeds, facial expressions, and gestural languages, depending on those with whom they are speaking. This study focuses on the age of the speaker as a factor that affects the style of communication. We collected a multimodal dialogue corpus with various speaker ages. We used travel as the topic, as it interests people of all ages, and we set up a task based on a tourism consultation between an operator and a customer at a travel agency. This article presents the details of the dialogue task, collection procedures and annotations, and analysis of the characteristics of the dialogues and facial expressions, focusing on the age of the speakers. The results of the analysis suggest that the adult speakers have more independent opinions, the older speakers express their opinions more frequently than other age groups, and those in the operator role smile more frequently at minors. Michimasa Inaba, Yuya Chiba, Zhiyang Qi, Ryuichiro Higashinaka, Kazunori Komatani, Yusuke Miyao, Takayuki Nagai |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 2 |
| 2023 | Personality-aware Natural Language Generation for Task-oriented Dialogue using Reinforcement LearningabstractA task-oriented dialogue system capable of expressing personality can improve user engagement and satisfaction. To realize such a system, this paper presents a method of building a personality-aware natural language generation (NLG) module in task-oriented dialogue using reinforcement learning (RL). This method handles both the expression of personality and system intent. During the RL process, a positive reward is given when the generated utterance correctly expresses the assigned personality and conveys its system intent simultaneously. In our experiments on the MultiWOZ dataset, we fine-tuned a personality-aware NLG module for two personality traits (extraversion and neural perception sensitivity). We experimented with data from MultiWOZ and a user simulator to confirm its effectiveness in terms of the ability to express personality and task performance. Atsumoto Ohashi, Yuya Chiba, Yuiko Tsunomori, Ryu Hirai, Ryuichiro Higashinaka |
RO-MAN | 3 |
| 2023 | Analyzing Variations of Everyday Japanese Conversations Based on Semantic Labels of Functional ExpressionsabstractTo achieve effective dialogue processing, the kinds of daily conversations people have must be clarified. Unfortunately, the characteristics of everyday conversations remain insufficiently investigated. In recent years, the Corpus of Everyday Japanese Conversation (CEJC) was developed, which is a large-scale corpus constructed by recording everyday Japanese conversations. In this article, we investigate the linguistic variations of everyday conversations in a multitude of situations using CEJC. We conducted factor analysis of it using the semantic categories of functional expressions that represent such subjective information as modality, thoughts, and communicative intention in addition to various tenses and facts. Our analysis identified seven factors that characterize everyday conversations and suggests that they are expressed by a combination of a dialogue’s purpose (e.g., “Explanation” and “Suggestion”) and its manners (e.g., “Politeness” and “Involvement”). We also analyzed the BTSJ–Japanese natural conversation corpus with transcripts and recordings and the Nagoya University conversational corpus and confirmed the generalizability of these factors. Yuya Chiba, Ryuichiro Higashinaka |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 1 |
| 2022 | Collection and Analysis of Travel Agency Task Dialogues with Age-Diverse SpeakersabstractWhen individuals communicate with each other, they use different vocabulary, speaking speed, facial expressions, and body language depending on the people they talk to. This paper focuses on the speaker’s age as a factor that affects the change in communication. We collected a multimodal dialogue corpus with a wide range of speaker ages. As a dialogue task, we focus on travel, which interests people of all ages, and we set up a task based on a tourism consultation between an operator and a customer at a travel agency. This paper provides details of the dialogue task, the collection procedure and annotations, and the analysis on the characteristics of the dialogues and facial expressions focusing on the age of the speakers. Results of the analysis suggest that the adult speakers have more independent opinions, the older speakers more frequently express their opinions frequently compared with other age groups, and the operators expressed a smile more frequently to the minor speakers. Michimasa Inaba, Yuya Chiba, Ryuichiro Higashinaka, Kazunori Komatani, Yusuke Miyao, Takayuki Nagai |
LREC | 2 |
| 2022 | Empirical Analysis of Training Strategies of Transformer-Based Japanese Chit-Chat SystemsabstractIn recent years, several high-performance conversational systems have been proposed based on the Transformer encoder-decoder model. Although previous studies analyzed the effects of the model parameters and the decoding method on subjective dialogue evaluations with overall metrics, it is not analyzed enough how the differences of fine-tuning datasets affect the user's detailed impressions. In addition, the Transformer-based approach has mostly been verified for English, not for such languages as Japanese that have large inter-language distances. In this study, we developed large-scale Transformer-based Japanese dialogue models and Japanese chit-chat datasets and examined their effectiveness. We analyzed the relationships between users' multifaceted impressions and fine-tuning datasets. Hiroaki Sugiyama, Masahiro Mizukami, Tsunehiro Arimoto, Hiromi Narimatsu, Yuya Chiba, Hideharu Nakajima, Toyomi Meguro |
SLT | 5 |
| 2021 | Dialogue Situation Recognition for Everyday Conversation Using Multimodal Information
Yuya Chiba, Ryuichiro Higashinaka |
Interspeech | 1 |
| 2021 | Neural Spoken-Response Generation Using Prosodic and Linguistic Context for Conversational Systems
Yoshihiro Yamazaki, Yuya Chiba, Takashi Nose, Akinori Ito |
Interspeech | 2 |
| 2021 | Variation across Everyday Conversations: Factor Analysis of Conversations using Semantic Categories of Functional Expressions
Yuya Chiba, Ryuichiro Higashinaka |
PACLIC | 1 |
| 2020 | Multi-Stream Attention-Based BLSTM with Feature Segmentation for Speech Emotion Recognition
Yuya Chiba, Takashi Nose, Akinori Ito |
INTERSPEECH | 1 |
| 2020 | Construction and Analysis of a Multimodal Chat-talk Corpus for Dialog Systems Considering Interpersonal ClosenessabstractThere are high expectations for multimodal dialog systems that can make natural small talk with facial expressions, gestures, and gaze actions as next-generation dialog-based systems. Two important roles of the chat-talk system are keeping the user engaged and establishing rapport. Many studies have conducted user evaluations of such systems, some of which reported that considering the relationship with the user is an effective way to improve the subjective evaluation. To facilitate research of such dialog systems, we are currently constructing a large-scale multimodal dialog corpus focusing on the relationship between speakers. In this paper, we describe the data collection and annotation process, and analysis of the corpus collected in the early stage of the project. This corpus contains 19,303 utterances (10 hours) from 19 pairs of participants. A dialog act tag is annotated to each utterance by two annotators. We compare the frequency and the transition probability of the tags between different closeness levels to help construct a dialog system for establishing a relationship with the user. Yoshihiro Yamazaki, Yuya Chiba, Takashi Nose, Akinori Ito |
LREC | 2 |
| 2020 | Automatic assessment of English proficiency for Japanese learners without reference sentences based on deep neural network acoustic models
Jiang Fu, Yuya Chiba, Takashi Nose, Akinori Ito |
Speech Commun. | 2 |
| 2019 | Improving human scoring of prosody using parametric speech synthesis
Hafiyan Prafianto, Takashi Nose, Yuya Chiba, Akinori Ito |
Speech Commun. | 3 |
| 2018 | Analyzing Effect of Physical Expression on English Proficiency for Multimodal Computer-Assisted Language Learning
Yuya Chiba, Takashi Nose, Akinori Ito |
INTERSPEECH | 2 |
| 2018 | An Analysis of the Effect of Emotional Speech Synthesis on Non-Task-Oriented Dialogue SystemabstractThis paper explores the effect of emotional speech synthesis on a spoken dialogue system when the dialogue is non-task-oriented.Although the use of emotional speech responses has been shown to be effective in a limited domain, e.g., scenario-based and counseling dialogue, the effect is still not clear in the non-task-oriented dialogue such as voice chat.For this purpose, we constructed a simple dialogue system with example-and rule-based dialogue management.In the system, two types of emotion labeling with emotion estimation are adopted, i.e., system-driven and user-cooperative emotion labeling.We conducted a dialogue experiment where subjects evaluate the subjective quality of the system and the dialogue from multiple aspects such as richness of the dialogue and impression of the agent.We then analyze and discuss the results and show the advantage of using appropriate emotions for expressive speech responses in the non-task-oriented system. Yuya Chiba, Takashi Nose, Taketo Kase, Mai Yamanaka, Akinori Ito |
SIGDIAL Conference | 1 |
| 2018 | Improving User Impression in Spoken Dialog System with Gradual Speech Form ControlabstractThis paper examines a method to improve the user impression of a spoken dialog system by introducing a mechanism that gradually changes form of utterances every time the user uses the system.In some languages, including Japanese, the form of utterances changes corresponding to social relationship between the talker and the listener.Thus, this mechanism can be effective to express the system's intention to make social distance to the user closer; however, an actual effect of this method is not investigated enough when introduced to the dialog system.In this paper, we conduct dialog experiments and show that controlling the form of system utterances can improve the users' impression. Yukiko Kageyama, Yuya Chiba, Takashi Nose, Akinori Ito |
SIGDIAL Conference | 2 |
| 2014 | User Modeling by Using Bag-of-Behaviors for Building a Dialog System Sensitive to the Interlocutor's Internal StateabstractWhen using spoken dialog systems in ac-tual environments, users sometimes aban-don the dialog without making any in-put utterance. To help these users before they give up, the system should know why they could not make an utterance. Thus, we have examined a method to estimate the state of a dialog user by capturing the user’s non-verbal behavior even when the user’s utterance is not observed. The pro-posed method is based on vector quan-tization of multi-modal features such as non-verbal speech, feature points of the face, and gaze. The histogram of the VQ code is used as a feature for determining the state. We call this feature “the Bag-of-Behaviors. ” According to the experi-mental results, we prove that the proposed method surpassed the results of conven-tional approaches and discriminated the target user’s states with an accuracy of more than 70%. 1 Yuya Chiba, Masashi Ito, Takashi Nose, Akinori Ito |
SIGDIAL Conference | 1 |
| 2012 | Estimation of User's Internal State before the User's First Utterance Using Acoustic Features and Face OrientationabstractIntroduction of user models (e.g. models of a user's belief, skill and familiarity to the system) is believed to increase flexibility of response of a dialogue system. Conventionally, the internal state is estimated based on linguistic information of the previous utterance, but this approach cannot applied to the user who did not make an input utterance in the first place. Thus, we are developing a method to estimate an internal state of a spoken dialogue system's user before his/her input utterance. In a previous report, we used three acoustic features and a visual feature based on manual labels. In this paper, we introduced new features for the estimation: length of filled pause and face orientation angles. Then, we examined effectiveness of the proposed features by experiments. As a result, we obtained a three-class discrimination accuracy of 85.6% in an open test, which was 1.5 point higher than the result obtained using the previous feature set. Yuya Chiba, Masashi Ito, Akinori Ito |
HSI | 1 |