EDBT 2026 Demo / reviewers in the wild / expert
Ryuichiro Higashinaka
dblp:35/4482
· DBLP profile ↗
102ranked-venue papers
25as first author
36since 2021 · last 2025
0000-0002-6994-3977ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 87 · 21 first-author · 33 since 2021Graphics, computer vision, multimedia, augmented reality and games · 26 · 11 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 12 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Universal Post-Processing Networks for Joint Optimization of Modules in Task-Oriented Dialogue SystemsabstractPost-processing networks (PPNs) are components that modify the outputs of arbitrary modules in task-oriented dialogue systems and are optimized using reinforcement learning (RL) to improve the overall task completion capability of the system. However, previous PPN-based approaches have been limited to handling only a subset of modules within a system, which poses a significant limitation in improving the system performance. In this study, we propose a joint optimization method for post-processing the outputs of all modules using universal post-processing networks (UniPPNs), which are language-model-based networks that can modify the outputs of arbitrary modules in a system as a sequence-transformation task. Moreover, our RL algorithm, which employs a module-level Markov decision process, enables fine-grained value and advantage estimation for each module, thereby stabilizing joint learning for post-processing the outputs of all modules. Through both simulation-based and human evaluation experiments using the MultiWOZ dataset, we demonstrated that UniPPN outperforms conventional PPNs in the task completion capability of task-oriented dialogue systems. Atsumoto Ohashi, Ryuichiro Higashinaka |
AAAI | 2 |
| 2025 | Common Ground Building through Generative Cognitive Modules: Examining the Roles of Initial Perception, Imaging and Captioning
Ryunosuke Baba, Junya Morita, Takeru Amaya, Ryuichiro Higashinaka, Yugo Takeuchi |
CogSci | 4 |
| 2025 | Investigating the Impact of Incremental Processing and Voice Activity Projection on Spoken Dialogue SystemsabstractThe naturalness of responses in spoken dialogue systems has been significantly improved by the introduction of large language models (LLMs), although many challenges remain until human-like turn-taking can be achieved. A turn-taking model called Voice Activity Projection (VAP) is gaining attention because it can be trained in an unsupervised manner using the spoken dialogue data between two speakers. For such a turn-taking model to be fully effective, systems must initiate response generation as soon as a turn-shift is detected. This can be achieved by incremental response generation, which reduces the delay before the system responds. Incremental response generation is done using partial speech recognition results while user speech is incrementally processed. Combining incremental response generation with VAP-based turn-taking will enable spoken dialogue systems to achieve faster and more natural turn-taking. However, their effectiveness remains unclear because they have not yet been evaluated in real-world systems. In this study, we developed spoken dialogue systems that incorporate incremental response generation and VAP-based turn-taking and evaluated their impact on task success and dialogue satisfaction through user assessments. Yuya Chiba, Ryuichiro Higashinaka |
COLING | 2 |
| 2025 | Teleoperation System Enabling Operator-Robot Dialogue for Reducing Operator Boredom during Long-Duration TasksabstractTeleoperated customer service robots have attracted attention to improve customer service efficiency. However, operators experience boredom during long-duration operation due to monotony and idle time, leading to decreased task motivation. This study proposes and evaluates a method to reduce operator boredom through dialogue with the robot to be operated. Field experiments demonstrated that operator-robot dialogue significantly reduced boredom and contributed to maintaining task engagement during long-duration operation. Manato Uetake, Tomonori Kubota, Masaya Iwasaki, Shota Mochizuki, Sanae Yamashita, Kenya Hoshimure, Jun Baba, Ryuichiro Higashinaka, Satoshi Sato, Kohei Ogawa |
HAI | 9 |
| 2025 | Towards a Japanese Full-duplex Spoken Dialogue System
Atsumoto Ohashi, Shinya Iizuka, Jingjing Jiang, Ryuichiro Higashinaka |
INTERSPEECH | 4 |
| 2025 | Anomaly Detection in Human-Robot Interaction Using Multimodal Models Constructed from In-the-Wild InteractionsabstractIn recent years, numerous studies have been conducted on dialogue robots powered by large language models,enabling sophisticated interactions such as providing guidance and engaging in small talk. However, the interaction performance remains imperfect, and the robots sometimes cause problems during interactions. In this study, we aim to automatically detect such anomalies in human-robot interactions by creating a dataset and developing anomaly detection models. To this end, we created a dataset by manually annotating videos of in-the-wild interactions collected from our field experiment designed to test a framework of parallel conversations in which a human intervenes when a problem occurs in the interaction. Using this dataset, we trained classification models to construct anomaly detection models. We then conducted another field experiment in which the model’s detection results were presented as alerts to operators within the parallel conversation framework. The results confirmed that providing alerts on the basis of the anomaly detection model was useful for facilitating operator intervention. Shota Mochizuki, Sanae Yamashita, Kenya Hoshimure, Jun Baba, Tomonori Kubota, Kohei Ogawa, Ryuichiro Higashinaka |
IROS | 7 |
| 2025 | Analysis of the Correlation Between Theory of Mind and Dialogue Ability to Identify Essential ToM for Dialogue Systems
Haruhisa Iseno, Atsumoto Ohashi, Tetsuji Ogawa, Shinnosuke Takamichi, Ryuichiro Higashinaka |
PACLIC | 5 |
| 2025 | Exploring Factors Influencing Hospitality in Mobile Robot Guidance: A Wizard-of-Oz Study with a Teleoperated Humanoid RobotabstractDeveloping mobile robots that can provide guidance with high hospitality remains challenging, as it requires the coordination of spoken interaction, physical navigation, and user engagement. To gain insights that contribute to the development of such robots, we conducted a Wizard-of-Oz (WOZ) study using Teleco, a teleoperated humanoid robot, to explore the factors influencing hospitality in mobile robot guidance. Specifically, we enrolled 30 participants as visitors and two trained operators, who teleoperated the Teleco robot to provide mobile guidance to the participants. A total of 120 dialogue sessions were collected, along with evaluations from both the participants and the operators regarding the hospitality of each interaction. To identify the factors that influence hospitality in mobile guidance, we analyzed the collected dialogues from two perspectives: linguistic usage and multimodal robot behaviors. We first clustered system utterances and analyzed the frequency of categories in high- and low-satisfaction dialogues. The results showed that short responses appeared more frequently in high-satisfaction dialogues. Moreover, we observed a general increase in participant satisfaction over successive sessions, along with shifts in linguistic usage, suggesting a mutual adaptation effect between operators and participants. We also conducted a time-series analysis of multimodal robot behaviors to explore behavioral patterns potentially linked to hospitable interactions. Shota Mochizuki, Sanae Yamashita, Saya Nikaido, Tomoko Isomura, Ryuichiro Higashinaka |
SIGDIAL | 6 |
| 2025 | Integrating Physiological, Speech, and Textual Information Toward Real-Time Recognition of Emotional Valence in DialogueabstractAccurately estimating users’ emotional states in real time is crucial for enabling dialogue systems to respond adaptively. While existing approaches primarily rely on verbal information, such as text and speech, these modalities are often unavailable in non-speaking situations. In such cases, non-verbal information, particularly physiological signals, becomes essential for understanding users’ emotional states. In this study, we aimed to develop a model for real-time recognition of users’ binary emotional valence (high-valence vs. low-valence) during conversations. Specifically, we utilized an existing Japanese multimodal dialogue dataset, which includes various physiological signals, namely electrodermal activity (EDA), blood volume pulse (BVP), photoplethysmography (PPG), and pupil diameter, along with speech and textual data. We classify the emotional valence of every 15-second segment of dialogue interaction by integrating such multimodal inputs. To this end, time-series embeddings of physiological signals are extracted using a self-supervised encoder, while speech and textual features are obtained from pre-trained Japanese HuBERT and BERT models, respectively. The modality-specific embeddings are integrated using a feature fusion mechanism for emotional valence recognition. Experimental results show that while each modality individually contributes to emotion recognition, the inclusion of physiological signals leads to a notable performance improvement, particularly in non-speaking or minimally verbal situations. These findings underscore the importance of physiological information for enhancing real-time valence recognition in dialogue systems, especially when verbal information is limited. Jingjing Jiang, Ryuichiro Higashinaka |
SIGDIAL | 3 |
| 2025 | Key Challenges in Multimodal Task-Oriented Dialogue Systems: Insights from a Large Competition-Based DatasetabstractChallenges in multimodal task-oriented dialogue between humans and systems, particularly those involving audio and visual interactions, have not been sufficiently explored or shared, forcing researchers to define improvement directions individually without a clearly shared roadmap. To address these challenges, we organized a competition for multimodal task-oriented dialogue systems and constructed a large competition-based dataset of 1,865 minutes of Japanese task-oriented dialogues. This dataset includes audio and visual interactions between diverse systems and human participants. After analyzing system behaviors identified as problematic by the human participants in questionnaire surveys and notable methods employed by the participating teams, we identified key challenges in multimodal task-oriented dialogue systems and discussed potential directions for overcoming these challenges. Shiki Sato, Shinji Iwata, Asahi Hentona, Yuta Sasaki, Takato Yamazaki, Shoji Moriya, Masaya Ohagi, Hirofumi Kikuchi, Zhiyang Qi, Takashi Kodama, Akinobu Lee, Masato Komuro, Hiroyuki Nishikawa, Ryosaku Makino, Takashi Minato, Kurima Sakai, Tomo Funayama, Kotaro Funakoshi, Mayumi Usami, Michimasa Inaba, Tetsuro Takahashi, Ryuichiro Higashinaka |
SIGDIAL | 23 |
| 2025 | Analyzing Dialogue System Behavior in a Specific Situation Requiring Interpersonal ConsiderationabstractIn human-human conversation, interpersonal consideration for the interlocutor is essential, and similar expectations are increasingly placed on dialogue systems. This study examines the behavior of dialogue systems in a specific interpersonal scenario where a user vents frustrations and seeks emotional support from a long-time friend represented by a dialogue system. We conducted a human evaluation and qualitative analysis of 15 dialogue systems under this setting. These systems implemented diverse strategies, such as structuring dialogue into distinct phases, modeling interpersonal relationships, and incorporating cognitive behavioral therapy techniques. Our analysis reveals that these approaches contributed to improved perceived empathy, coherence, and appropriateness, highlighting the importance of design choices in socially sensitive dialogue. Tetsuro Takahashi, Hirofumi Kikuchi, Hiroyuki Nishikawa, Masato Komuro, Ryosaku Makino, Shiki Sato, Yuta Sasaki, Shinji Iwata, Asahi Hentona, Takato Yamazaki, Shoji Moriya, Masaya Ohagi, Zhiyang Qi, Takashi Kodama, Akinobu Lee, Takashi Minato, Kurima Sakai, Tomo Funayama, Kotaro Funakoshi, Mayumi Usami, Michimasa Inaba, Ryuichiro Higashinaka |
SIGDIAL | 23 |
| 2025 | Optimizing pipeline task-oriented dialogue systems using post-processing networksabstractMany studies have proposed methods for optimizing the dialogue performance of an entire pipeline task-oriented dialogue system by jointly training modules in the system using reinforcement learning . However, these methods are limited in that they can only be applied to modules implemented using trainable neural-based methods. To solve this problem, we propose a method for optimizing the dialogue performance of a pipeline system that consists of modules implemented with arbitrary methods for dialogue. With our method, neural-based components called post-processing networks (PPNs) are installed inside such a system to post-process the output of each module. All PPNs are updated to improve the overall dialogue performance of the system by using reinforcement learning, not necessitating that each module be differentiable. Through dialogue simulations and human evaluations on two well-studied task-oriented dialogue datasets, CamRest676 and MultiWOZ, we show that our method can improve the dialogue performance of pipeline systems consisting of various modules. In addition, a comprehensive analysis of the results of the MultiWOZ experiments reveals the patterns of post-processing by PPNs that contribute to the overall dialogue performance of the system. Atsumoto Ohashi, Ryuichiro Higashinaka |
Comput. Speech Lang. | 2 |
| 2024 | JMultiWOZ: A Large-Scale Japanese Multi-Domain Task-Oriented Dialogue DatasetabstractDialogue datasets are crucial for deep learning-based task-oriented dialogue system research. While numerous English language multi-domain task-oriented dialogue datasets have been developed and contributed to significant advancements in task-oriented dialogue systems, such a dataset does not exist in Japanese, and research in this area is limited compared to that in English. In this study, towards the advancement of research and development of task-oriented dialogue systems in Japanese, we constructed JMultiWOZ, the first Japanese language large-scale multi-domain task-oriented dialogue dataset. Using JMultiWOZ, we evaluated the dialogue state tracking and response generation capabilities of the state-of-the-art methods on the existing major English benchmark dataset MultiWOZ2.2 and the latest large language model (LLM)-based methods. Our evaluation results demonstrated that JMultiWOZ provides a benchmark that is on par with MultiWOZ2.2. In addition, through evaluation experiments of interactive dialogues with the models and human participants, we identified limitations in the task completion capabilities of LLMs in Japanese. Atsumoto Ohashi, Ryu Hirai, Shinya Iizuka, Ryuichiro Higashinaka |
LREC/COLING | 4 |
| 2024 | I Remember You!: SUI Corpus for Remembering and Utilizing Users' Information in Chat-oriented Dialogue SystemsabstractTo construct a chat-oriented dialogue system that will be used for a long time by users, it is important to build a good relationship between the user and the system. To achieve a good relationship, several methods for remembering and utilizing information on users (preferences, experiences, jobs, etc.) in system utterances have been investigated. One way to do this is to utilize user information to fill in utterance templates for use in response generation, but the utterances do not always fit the context. Another way is to use neural-based generation, but in current methods, user information can be incorporated only when the current dialogue topic is similar to that of the user information. This paper tackled these problems by constructing a novel corpus to incorporate arbitrary user information into system utterances regardless of the current dialogue topic while retaining appropriateness for the context. We then fine-tuned a model for generating system utterances using the constructed corpus. The result of a subjective evaluation demonstrated the effectiveness of our model. Furthermore, we incorporated our fine-tuned model into a dialogue system and confirmed the effectiveness of the system through interactive dialogues with users. Yuiko Tsunomori, Ryuichiro Higashinaka |
LREC/COLING | 2 |
| 2024 | Collecting and Analyzing Dialogues in a Tagline Co-Writing TaskabstractThe potential usage scenarios of dialogue systems will be greatly expanded if they are able to collaborate more creatively with humans. Many studies have examined ways of building such systems, but most of them focus on problem-solving dialogues, and relatively little research has been done on systems that can engage in creative collaboration with users. In this study, we designed a tagline co-writing task in which two people collaborate to create taglines via text chat, created an interface for data collection, and collected dialogue logs, editing logs, and questionnaire results. In total, we collected 782 Japanese dialogues. We describe the characteristic interactions comprising the tagline co-writing task and report the results of our analysis, in which we examined the kind of utterances that appear in the dialogues as well as the most frequent expressions found in highly rated dialogues in subjective evaluations. We also analyzed the relationship between subjective evaluations and workflow utilized in the dialogues and the interplay between taglines and utterances. Xulin Zhou, Takuma Ichikawa, Ryuichiro Higashinaka |
LREC/COLING | 3 |
| 2024 | Learning Anomaly Detection Models for Human-Robot InteractionabstractDialogue robots powered by large language models can generate advanced utterances. However, the interaction performance is not yet perfect, and the robots sometimes cause problems during interactions. In this study, to detect anomalies in human-robot interactions, we created a dataset and constructed anomaly detection models. For the dataset creation, we collected videos of human-robot interactions in a framework where humans intervene when a dialogue breakdown occurs and labeled the scenes where humans intervened as anomalies. Using this dataset, we built classification models and deep metric learning models utilizing encoders for video, audio, and multimodal information. The results showed that we could successfully train the models and achieve an accuracy and F1-score of over 80%. The performance of the deep metric learning models surpassed that of the classification models, thus demonstrating the importance of separating the differences between classes. We also clarified the importance of audio information. Shota Mochizuki, Sanae Yamashita, Reiko Yuasa, Tomonori Kubota, Kohei Ogawa, Ryuichiro Higashinaka |
RO-MAN | 6 |
| 2024 | Estimating the Emotional Valence of Interlocutors Using Heterogeneous Sensors in Human-Human DialogueabstractDialogue systems need to accurately understand the user's mental state to generate appropriate responses, but accurately discerning such states solely from text or speech can be challenging.To determine which information is necessary, we first collected human-human multimodal dialogues using heterogeneous sensors, resulting in a dataset containing various types of information including speech, video, physiological signals, gaze, and body movement.Additionally, for each time step of the data, users provided subjective evaluations of their emotional valence while reviewing the dialogue videos.Using this dataset and focusing on physiological signals, we analyzed the relationship between the signals and the subjective evaluations through Granger causality analysis.We also investigated how sensor signals differ depending on the polarity of the valence.Our findings revealed several physiological signals related to the user's emotional valence. Jingjing Jiang, Ryuichiro Higashinaka |
SIGDIAL | 3 |
| 2024 | Travel Agency Task Dialogue Corpus: A Multimodal Dataset with Age-Diverse SpeakersabstractWhen individuals communicate, they use different vocabularies, speaking speeds, facial expressions, and gestural languages, depending on those with whom they are speaking. This study focuses on the age of the speaker as a factor that affects the style of communication. We collected a multimodal dialogue corpus with various speaker ages. We used travel as the topic, as it interests people of all ages, and we set up a task based on a tourism consultation between an operator and a customer at a travel agency. This article presents the details of the dialogue task, collection procedures and annotations, and analysis of the characteristics of the dialogues and facial expressions, focusing on the age of the speakers. The results of the analysis suggest that the adult speakers have more independent opinions, the older speakers express their opinions more frequently than other age groups, and those in the operator role smile more frequently at minors. Michimasa Inaba, Yuya Chiba, Zhiyang Qi, Ryuichiro Higashinaka, Kazunori Komatani, Yusuke Miyao, Takayuki Nagai |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 4 |
| 2023 | Enhancing Task-oriented Dialogue Systems with Generative Post-processing NetworksabstractRecently, post-processing networks (PPNs), which modify the outputs of arbitrary modules including non-differentiable ones in taskoriented dialogue systems, have been proposed.PPNs have successfully improved the dialogue performance by post-processing natural language understanding (NLU), dialogue state tracking (DST), and dialogue policy (Policy) modules with a classification-based approach.However, they cannot be applied to natural language generation (NLG) modules because the post-processing of utterances output by NLG modules requires a generative approach.In this study, we propose a new postprocessing component for NLG, generative post-processing networks (GenPPNs).For optimizing GenPPNs via reinforcement learning, the reward function incorporates dialogue act contribution, a new measure to evaluate the contribution of GenPPN-generated utterances with regard to task completion in dialogue.Through simulation and human evaluation experiments based on the MultiWOZ dataset, we confirmed that GenPPNs improve the task completion performance of task-oriented dialogue systems 1 . Atsumoto Ohashi, Ryuichiro Higashinaka |
EMNLP | 2 |
| 2023 | Investigating the Intervention in Parallel ConversationsabstractIn recent years, a framework of parallel conversations has been proposed to facilitate efficient conversations through cooperation between humans and dialogue systems. This approach aims to enable simultaneous conversations with multiple users by enabling the system to handle basic conversation and human operators to intervene when problems arise in the system’s conversation. Previous studies on parallel conversations have primarily focused on delegating simple exchanges such as greetings and acknowledgments to the system, with humans taking over for more complex interactions like providing guidance. Recent advancements in large language models may change this situation, enabling dialogue systems to engage in more advanced interactions. In this study, to examine which interventions will be made when large language models are utilized, we placed six dialogue robots based on large language models in an actual facility and conducted a field experiment involving parallel conversations for about a month. Our analysis of the collected data on dialogues and interventions showed that the most frequent interventions were made for supporting interactions when the system failed to react to the user utterances, indicating the limitations of using large language models alone and clarifying our next steps for facilitating smoother parallel conversations. Shota Mochizuki, Sanae Yamashita, Kazuyoshi Kawasaki, Reiko Yuasa, Tomonori Kubota, Kohei Ogawa, Jun Baba, Ryuichiro Higashinaka |
HAI | 8 |
| 2023 | Investigating the Effects of Dialogue Summarization on Intervention in Human-System Collaborative DialogueabstractDialogue systems are widely utilized in chatbots and call centers. However, it is often difficult for such systems to deliver fully autonomous dialogue. For users to have a better dialogue experience, a framework for human-system collaborative dialogue is proposed in which a human operator takes over the dialogue when needed, engaging in conversation with the user instead of the system (we call this process intervention). Operators join the dialogue in the middle; therefore, it is believed that dialogue summarization can be helpful for interventions. However, it is currently unclear whether dialogue summarization is actually useful. Therefore, in this study, we aim to investigate the usefulness of dialogue summaries for interventions through a field experiment conducted at an actual facility combining an aquarium and a zoo. The results of the field experiment revealed that dialogue summaries were more useful for intervention than dialogue history. Furthermore, we found no differences in the word categories included in the operator utterances during interventions irrespective of whether the dialogue history or dialogue format summary was presented to the operators, suggesting that dialogue format summary has content similar to that of dialogue history but improves the usefulness in intervention. Sanae Yamashita, Shota Mochizuki, Kazuyoshi Kawasaki, Tomonori Kubota, Kohei Ogawa, Jun Baba, Ryuichiro Higashinaka |
HAI | 7 |
| 2023 | RealPersonaChat: A Realistic Persona Chat Corpus with Interlocutors' Own Personalities
Sanae Yamashita, Koji Inoue, Shota Mochizuki, Tatsuya Kawahara, Ryuichiro Higashinaka |
PACLIC | 6 |
| 2023 | Personality-aware Natural Language Generation for Task-oriented Dialogue using Reinforcement LearningabstractA task-oriented dialogue system capable of expressing personality can improve user engagement and satisfaction. To realize such a system, this paper presents a method of building a personality-aware natural language generation (NLG) module in task-oriented dialogue using reinforcement learning (RL). This method handles both the expression of personality and system intent. During the RL process, a positive reward is given when the generated utterance correctly expresses the assigned personality and conveys its system intent simultaneously. In our experiments on the MultiWOZ dataset, we fine-tuned a personality-aware NLG module for two personality traits (extraversion and neural perception sensitivity). We experimented with data from MultiWOZ and a user simulator to confirm its effectiveness in terms of the ability to express personality and task performance. Atsumoto Ohashi, Yuya Chiba, Yuiko Tsunomori, Ryu Hirai, Ryuichiro Higashinaka |
RO-MAN | 6 |
| 2023 | Applying Item Response Theory to Task-oriented Dialogue Systems for Accurately Determining User's Task Success AbilityabstractWhile task-oriented dialogue systems have improved, not all users can fully accomplish their tasks.Users with limited knowledge about the system may experience dialogue breakdowns or fail to achieve their tasks because they do not know how to interact with the system.For addressing this issue, it would be desirable to construct a system that can estimate the user's task success ability and adapt to that ability.In this study, we propose a method that estimates this ability by applying item response theory (IRT), commonly used in education for estimating examinee abilities, to task-oriented dialogue systems.Through experiments predicting the probability of a correct answer to each slot by using the estimated task success ability, we found that the proposed method significantly outperformed baselines. Ryu Hirai, Ryuichiro Higashinaka |
SIGDIAL | 3 |
| 2023 | Analyzing Variations of Everyday Japanese Conversations Based on Semantic Labels of Functional ExpressionsabstractTo achieve effective dialogue processing, the kinds of daily conversations people have must be clarified. Unfortunately, the characteristics of everyday conversations remain insufficiently investigated. In recent years, the Corpus of Everyday Japanese Conversation (CEJC) was developed, which is a large-scale corpus constructed by recording everyday Japanese conversations. In this article, we investigate the linguistic variations of everyday conversations in a multitude of situations using CEJC. We conducted factor analysis of it using the semantic categories of functional expressions that represent such subjective information as modality, thoughts, and communicative intention in addition to various tenses and facts. Our analysis identified seven factors that characterize everyday conversations and suggests that they are expressed by a combination of a dialogue’s purpose (e.g., “Explanation” and “Suggestion”) and its manners (e.g., “Politeness” and “Involvement”). We also analyzed the BTSJ–Japanese natural conversation corpus with transcripts and recordings and the Nagoya University conversational corpus and confirmed the generalizability of these factors. Yuya Chiba, Ryuichiro Higashinaka |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 2 |
| 2022 | Adaptive Natural Language Generation for Task-oriented Dialogue via Reinforcement LearningabstractWhen a natural language generation (NLG) component is implemented in a real-world task-oriented dialogue system, it is necessary to generate not only natural utterances as learned on training data but also utterances adapted to the dialogue environment (e.g., noise from environmental sounds) and the user (e.g., users with low levels of understanding ability). Inspired by recent advances in reinforcement learning (RL) for language generation tasks, we propose ANTOR, a method for Adaptive Natural language generation for Task-Oriented dialogue via Reinforcement learning. In ANTOR, a natural language understanding (NLU) module, which corresponds to the user’s understanding of system utterances, is incorporated into the objective function of RL. If the NLG’s intentions are correctly conveyed to the NLU, which understands a system’s utterances, the NLG is given a positive reward. We conducted experiments on the MultiWOZ dataset, and we confirmed that ANTOR could generate adaptive utterances against speech recognition errors and the different vocabulary levels of users. Atsumoto Ohashi, Ryuichiro Higashinaka |
COLING | 2 |
| 2022 | Dialogue Corpus Construction Considering Modality and Social Relationships in Building Common GroundabstractBuilding common ground with users is essential for dialogue agent systems and robots to interact naturally with people. While a few previous studies have investigated the process of building common ground in human-human dialogue, most of them have been conducted on the basis of text chat. In this study, we constructed a dialogue corpus to investigate the process of building common ground with a particular focus on the modality of dialogue and the social relationship between the participants in the process of building common ground, which are important but have not been investigated in the previous work. The results of our analysis suggest that adding the modality or developing the relationship between workers speeds up the building of common ground. Specifically, regarding the modality, the presence of video rather than only audio may unconsciously facilitate work, and as for the relationship, it is easier to convey information about emotions and turn-taking among friends than in first meetings. These findings and the corpus should prove useful for developing a system to support remote communication. Yuki Furuya, Koki Saito, Kosuke Ogura, Koh Mitsuda, Ryuichiro Higashinaka, Kazunori Takashio |
LREC | 5 |
| 2022 | Analysis of Dialogue in Human-Human Collaboration in MinecraftabstractRecently, many studies have focused on developing dialogue systems that enable collaborative work; however, they rarely focus on creative tasks. Collaboration for creative work, in which humans and systems collaborate to create new value, will be essential for future dialogue systems. In this study, we collected 500 dialogues of human-human collaboration in Minecraft as a basis for developing a dialogue system that enables creative collaborative work. We conceived the Collaborative Garden Task, where two workers interact and collaborate in Minecraft to create a garden, and we collected dialogue, action logs, and subjective evaluations. We also collected third-person evaluations of the gardens and analyzed the relationship between dialogue and collaborative work that received high scores on the subjective and third-person evaluations in order to identify dialogic factors for high-quality collaborative work. We found that two essential aspects in creative collaborative work are performing more processes to ask for and agree on suggestions between workers and agreeing on a particular image of the final product in the early phase of work and then discussing changes and details. Takuma Ichikawa, Ryuichiro Higashinaka |
LREC | 2 |
| 2022 | Collection and Analysis of Travel Agency Task Dialogues with Age-Diverse SpeakersabstractWhen individuals communicate with each other, they use different vocabulary, speaking speed, facial expressions, and body language depending on the people they talk to. This paper focuses on the speaker’s age as a factor that affects the change in communication. We collected a multimodal dialogue corpus with a wide range of speaker ages. As a dialogue task, we focus on travel, which interests people of all ages, and we set up a task based on a tourism consultation between an operator and a customer at a travel agency. This paper provides details of the dialogue task, the collection procedure and annotations, and the analysis on the characteristics of the dialogues and facial expressions focusing on the age of the speakers. Results of the analysis suggest that the adult speakers have more independent opinions, the older speakers more frequently express their opinions frequently compared with other age groups, and the operators expressed a smile more frequently to the minor speakers. Michimasa Inaba, Yuya Chiba, Ryuichiro Higashinaka, Kazunori Komatani, Yusuke Miyao, Takayuki Nagai |
LREC | 3 |
| 2022 | Dialogue Collection for Recording the Process of Building Common Ground in a Collaborative TaskabstractTo develop a dialogue system that can build common ground with users, the process of building common ground through dialogue needs to be clarified. However, the studies on the process of building common ground have not been well conducted; much work has focused on finding the relationship between a dialogue in which users perform a collaborative task and its task performance represented by the final result of the task. In this study, to clarify the process of building common ground, we propose a data collection method for automatically recording the process of building common ground through a dialogue by using the intermediate result of a task. We collected 984 dialogues, and as a result of investigating the process of building common ground, we found that the process can be classified into several typical patterns and that conveying each worker’s understanding through affirmation of a counterpart’s utterances especially contributes to building common ground. In addition, toward dialogue systems that can build common ground, we conducted an automatic estimation of the degree of built common ground and found that its degree can be estimated quite accurately. Koh Mitsuda, Ryuichiro Higashinaka, Yuhei Oga, Sen Yoshida |
LREC | 2 |
| 2022 | A Speculative and Tentative Common Ground Handling for Efficient Composition of Uncertain DialogueabstractThis study investigates how the grounding process is composed and explores new interaction approaches that adapt to human cognitive processes that have not yet been significantly studied. The results of an experiment indicate that grounding through dialogue is mutually accepted among participants through holistic expressions and suggest that common ground among participants may not necessarily be formed in a bottom-up way through analytic expressions. These findings raise the possibility of a promising new approach to creating a human-like dialogue system that may be more suitable for natural human communication. Saki Sudo, Kyoshiro Asano, Koh Mitsuda, Ryuichiro Higashinaka, Yugo Takeuchi |
LREC | 4 |
| 2022 | Data Collection for Empirically Determining the Necessary Information for Smooth Handover in DialogueabstractDespite recent advances, dialogue systems still struggle to achieve fully autonomous transactions. Therefore, when a system encounters a problem, human operators need to take over the dialogue to complete the transaction. However, it is unclear what information should be presented to the operator when this handover takes place. In this study, we conducted a data collection experiment in which one of two operators talked to a user and switched with the other operator periodically while exchanging notes when the handovers took place. By examining these notes, it is possible to identify the information necessary for handing over the dialogue. We collected 60 dialogues in which two operators switched periodically while performing chat, consultation, and sales tasks in dialogue. We found that adjacency pairs are a useful representation for recording conversation history. In addition, we found that key-value-pair representation is also useful when there are underlying tasks, such as consultation and sales. Sanae Yamashita, Ryuichiro Higashinaka |
LREC | 2 |
| 2022 | Post-processing Networks: Method for Optimizing Pipeline Task-oriented Dialogue Systems using Reinforcement LearningabstractMany studies have proposed methods for optimizing the dialogue performance of an entire pipeline task-oriented dialogue system by jointly training modules in the system using reinforcement learning.However, these methods are limited in that they can only be applied to modules implemented using trainable neuralbased methods.To solve this problem, we propose a method for optimizing a pipeline system composed of modules implemented with arbitrary methods for dialogue performance.With our method, neural-based components called post-processing networks (PPNs) are installed inside such a system to post-process the output of each module.All PPNs are updated to improve the overall dialogue performance of the system by using reinforcement learning, not necessitating each module to be differentiable.Through dialogue simulation and human evaluation on the MultiWOZ dataset, we show that our method can improve the dialogue performance of pipeline systems consisting of various modules 1 . Atsumoto Ohashi, Ryuichiro Higashinaka |
SIGDIAL | 2 |
| 2021 | Dialogue Situation Recognition for Everyday Conversation Using Multimodal Information
Yuya Chiba, Ryuichiro Higashinaka |
Interspeech | 2 |
| 2021 | Variation across Everyday Conversations: Factor Analysis of Conversations using Semantic Categories of Functional Expressions
Yuya Chiba, Ryuichiro Higashinaka |
PACLIC | 2 |
| 2021 | Integrated taxonomy of errors in chat-oriented dialogue systemsabstractThis paper proposes a taxonomy of errors in chat-oriented dialogue systems.Previously, two taxonomies were proposed; one is theorydriven and the other data-driven.The former suffers from the fact that dialogue theories for human conversation are often not appropriate for categorizing errors made by chat-oriented dialogue systems.The latter has limitations in that it can only cope with errors of systems for which we have data.This paper integrates these two taxonomies to create a comprehensive taxonomy of errors in chat-oriented dialogue systems.We found that, with our integrated taxonomy, errors can be reliably annotated with a higher Fleiss' kappa compared with the previously proposed taxonomies. Ryuichiro Higashinaka, Masahiro Araki, Hiroshi Tsukahara, Masahiro Mizukami |
SIGDIAL | 1 |
| 2020 | Generating Responses that Reflect Meta Information in User-Generated Question Answer PairsabstractThis paper concerns the problem of realizing consistent personalities in neural conversational modeling by using user generated question-answer pairs as training data. Using the framework of role play-based question answering, we collected single-turn question-answer pairs for particular characters from online users. Meta information was also collected such as emotion and intimacy related to question-answer pairs. We verified the quality of the collected data and, by subjective evaluation, we also verified their usefulness in training neural conversational models for generating utterances reflecting the meta information, especially emotion. Takashi Kodama, Ryuichiro Higashinaka, Koh Mitsuda, Ryo Masumura, Yushi Aono, Ryuta Nakamura, Noritake Adachi, Hidetoshi Kawabata |
LREC | 2 |
| 2020 | Collection and Analysis of Dialogues Provided by Two Speakers Acting as OneabstractTsunehiro Arimoto, Ryuichiro Higashinaka, Kou Tanaka, Takahito Kawanishi, Hiroaki Sugiyama, Hiroshi Sawada, Hiroshi Ishiguro. Proceedings of the 21th Annual Meeting of the Special Interest Group on Discourse and Dialogue. 2020. Tsunehiro Arimoto, Ryuichiro Higashinaka, Kou Tanaka, Takahito Kawanishi, Hiroaki Sugiyama, Hiroshi Sawada, Hiroshi Ishiguro |
SIGdial | 2 |
| 2019 | Improving Speech-Based End-of-Turn Detection Via Cross-Modal Representation Learning with Punctuated Text DataabstractThis paper presents a novel training method for speech-based end-of-turn detection for which not only manually annotated speech data sets but also punctuated text data sets are utilized. The speech-based end-of-turn detection estimates whether a target speaker's utterance is ended or not using speech information. In previous studies, the speech-based end-of-turn detection models were trained using only speech data sets that contained manually annotated end-of-turn labels. However, since the amounts of annotated speech data sets are often limited, the end-of-turn detection models were unable to correctly handle a wide variety of speech patterns. In order to mitigate the data scarcity problem, our key idea is to leverage punctuated text data sets for building more effective speech-based end-of-turn detection. Therefore, the proposed method introduces cross-modal representation learning to construct a speech encoder and a text encoder that can map speech and text with the same lexical information into similar vector representations. This enables us to train speech-based end-of-turn detection models from the punctuated text data sets by tackling text-based sentence boundary detection. In experiments on contact center calls, we show that speech-based end-of-turn detection models using hierarchical recurrent neural networks can be improved through the use of punctuated text data sets. Ryo Masumura, Mana Ihori, Tomohiro Tanaka, Atsushi Ando, Ryo Ishii, Takanobu Oba, Ryuichiro Higashinaka |
ASRU | 7 |
| 2019 | Determining Iconic Gesture Forms based on Entity Image RepresentationabstractIconic gestures are used to depict physical objects mentioned in speech, and the gesture form is assumed to be based on the image of a given object in the speaker’s mind. Using this idea, this study proposes a model that learns iconic gesture forms from an image representation obtained from pictures of physical entities. First, we collect a set of pictures of each entity from the web, and create an average image representation from them. Subsequently, the average image representation is fed to a fully connected neural network to decide the gesture form. In the model evaluation experiment, our two-step gesture form selection method can classify seven types of gesture forms with over 62% accuracy. Furthermore, we demonstrate an example of gesture generation in a virtual agent system in which our model is used to create a gesture dictionary that assigns a gesture form for each entry word in the dictionary. Fumio Nihei, Yukiko I. Nakano, Ryuichiro Higashinaka, Ryo Ishii |
ICMI | 3 |
| 2019 | Overview of the sixth dialog system technology challenge: DSTC6
Chiori Hori, Julien Perez, Ryuichiro Higashinaka, Takaaki Hori, Y-Lan Boureau, Michimasa Inaba, Yuiko Tsunomori, Tetsuro Takahashi, Koichiro Yoshino, Seokhwan Kim |
Comput. Speech Lang. | 3 |
| 2018 | Multi-task and Multi-lingual Joint Learning of Neural Lexical Utterance Classification based on Partially-shared ModelingabstractThis paper is an initial study on multi-task and multi-lingual joint learning for lexical utterance classification. A major problem in constructing lexical utterance classification modules for spoken dialogue systems is that individual data resources are often limited or unbalanced among tasks and/or languages. Various studies have examined joint learning using neural-network based shared modeling; however, previous joint learning studies focused on either cross-task or cross-lingual knowledge transfer. In order to simultaneously support both multi-task and multi-lingual joint learning, our idea is to explicitly divide state-of-the-art neural lexical utterance classification into language-specific components that can be shared between different tasks and task-specific components that can be shared between different languages. In addition, in order to effectively transfer knowledge between different task data sets and different language data sets, this paper proposes a partially-shared modeling method that possesses both shared components and components specific to individual data sets. We demonstrate the effectiveness of proposed method using Japanese and English data sets with three different lexical utterance classification tasks. Ryo Masumura, Tomohiro Tanaka, Ryuichiro Higashinaka, Hirokazu Masataki, Yushi Aono |
COLING | 3 |
| 2018 | Adversarial Training for Multi-task and Multi-lingual Joint Modeling of Utterance Intent ClassificationabstractThis paper proposes an adversarial training method for the multi-task and multi-lingual joint modeling needed for utterance intent classification.In joint modeling, common knowledge can be efficiently utilized among multiple tasks or multiple languages.This is achieved by introducing both languagespecific networks shared among different tasks and task-specific networks shared among different languages.However, the shared networks are often specialized in majority tasks or languages, so performance degradation must be expected for some minor data sets.In order to improve the invariance of shared networks, the proposed method introduces both language-specific task adversarial networks and task-specific language adversarial networks; both are leveraged for purging the task or language dependencies of the shared networks.The effectiveness of the adversarial training proposal is demonstrated using Japanese and English data sets for three different utterance intent classification tasks. Ryo Masumura, Yusuke Shinohara, Ryuichiro Higashinaka, Yushi Aono |
EMNLP | 3 |
| 2018 | Neural Confnet Classification: Fully Neural Network Based Spoken Utterance Classification Using Word Confusion NetworksabstractThis paper describes neural ConfNet classification, a novel fully neural network based spoken utterance classification method that uses word confusion networks (ConfNets). Our motivation is to establish a spoken utterance classification method that can precisely understand natural language and robustly handle automatic speech recognition (ASR) errors. Remarkable progress has been made in neural networks for accurate modeling, however, most previous methods could not handle ASR errors since they were developed for reference transcriptions. Therefore, in our work we utilized ConfNets, which are compact and efficient graph representations of ASR hypotheses. Our idea is to regard the ConfNet as a sequence of bag-of-weighted-arcs and introduce a mechanism that converts the bag-of-weighted-arcs into a continuous representation called a modified weighted sum representation. This enables us to flexibly connect ConfNets to arbitrary model structures developed for reference transcriptions. We demonstrate the effectiveness of the neural ConfNet classification in dialogue act, extended named entity, and question type classification tasks. Ryo Masumura, Yusuke Ijima, Taichi Asami, Hirokazu Masataki, Ryuichiro Higashinaka |
ICASSP | 5 |
| 2018 | Analyzing Gaze Behavior and Dialogue Act during Turn-taking for Estimating Empathy Skill LevelabstractWe explored the gaze behavior towards the end of utterances and dialogue act (DA), i.e., verbal-behavior information indicating the intension of an utterance, during turn-keeping/changing to estimate empathy skill levels in multiparty discussions. This is the first attempt to explore the relationship between such a combination. First, we collected data on Davis' Interpersonal Reactivity Index (which measures empathy skill level), utterances that include the DA categories of Provision, Self-disclosure, Empathy, Turn-yielding, and Others, and gaze behavior from participants in four-person discussions. The results of analysis indicate that the gaze behavior accompanying utterances that include these DA categories during turn-keeping/changing differs in accordance with people's empathy skill levels. The most noteworthy result was that speakers with low empathy skill levels tend to avoid making eye contact with the listener when the DA category is Self-disclosure during turn-keeping. However, they tend to maintain eye contact when the DA category is Empathy. A listener who has a high empathy skill level often looks away from the speaker during turn-changing when the DA category of a speaker's utterance is Provision or Empathy. There was also no difference in gaze behavior between empathy skill levels when the DA category of the speaker's utterance was turn-yielding. From these findings, we constructed and evaluated models for estimating empathy skill level using gaze behavior and DA information. The evaluation results indicate that using both gaze behavior and DA during turn-keeping/changing is effective for estimating an individual's empathy skill level in multi-party discussions. Ryo Ishii, Kazuhiro Otsuka, Shiro Kumano, Ryuichiro Higashinaka, Junji Tomita |
ICMI | 4 |
| 2018 | Generating Body Motions using Spoken Language in DialogueabstractWe propose a model to automatically generate whole body motions accompanying utterances at appropriate times, similar to humans, by using various types of natural-language-analysis information obtained from spoken language. Specifically, we focus on the co-occurrence relationship between various types of natural-language-analysis information such as words included in the spoken language, parts of speech, a thesaurus, word positions, dialogue acts of the spoken language, and human motions. Our model automatically generates nods, head postures, facial expressions, hand gestures, and upper-body posture using such information. We first recorded a two-person dialogue and constructed a multimodal corpus including utterance and whole body motion information. Next, using the constructed corpus, we constructed our model for generating a motion for each phrase unit using machine learning and using words, parts of speech, a thesaurus, word positions, and speech acts of the entire spoken language as inputs. These types of natural-language-analysis information were useful for motion generation. The effectiveness of our model was verified through a subjective experiment using a virtual conversational agent. As a result, the agent's body motions and impressions regarding naturalness of motion, degree of coincidence between utterance and motion, humanness of the agent, and likability of the agent improved with our model. Ryo Ishii, Taichi Katayama, Ryuichiro Higashinaka, Junji Tomita |
IVA | 3 |
| 2018 | Automatic Generation System of Virtual Agent's Motion using Natural LanguageabstractA virtual agent in a dialogue system should express appropriate body motions according to utterances and effectively communicate with a user. We previously proposed a generation model of whole body motions such as head direction, nodding, facial expressions, hand gestures, and upper-body posture accompanying utterances at appropriate times similar to humans by using various types of natural-language-analysis information obtained from spoken language. As an attempt to promote this model, we constructed an API that can easily generate motions by using the generation model and constructed a demonstration system that can automatically control a virtual agent from only the spoken language. When inputting an arbitrary utterance language, synthesized sound and motion information are acquired from the speech synthesizer and motion-generation API, and the vocalization of the virtual agent and animated motion are generated. A dialog agent that is more attractive and can communicate smoothly by automatically generating natural motions is expected. Ryo Ishii, Taichi Katayama, Ryuichiro Higashinaka, Junji Tomita |
IVA | 3 |
| 2018 | Predicting Nods by using Dialogue Acts in Dialogue
Ryo Ishii, Ryuichiro Higashinaka, Junji Tomita |
LREC | 2 |
| 2018 | Creating Large-Scale Argumentation Structures for Dialogue Systems
Kazuki Sakai, Akari Inago, Ryuichiro Higashinaka, Yuichiro Yoshikawa, Hiroshi Ishiguro, Junji Tomita |
LREC | 3 |
| 2018 | Automatic Generation of Head Nods using Utterance TextsabstractWe propose a model to generate head nods accompanying an utterance from natural language. To the best of our knowledge, previous models generated simple nods from the final words at the end of an utterance, i.e., using bag of words. We focused on various text analyzed using various types of language information such as dialog act, part of speech, a large-scale Japanese thesaurus, and word position in a sentence. We also generated detailed parameters of speaker's nodding presence, frequency, and depth, which was the first attempt to do so. First, we compiled a Japanese corpus of 24 dialogues including utterance and nod information. Next, using the corpus, we constructed our generation model that estimates nodding presence, frequency, and depth, during a phrase by using such various types of language information as well as bag of words. The results indicate that our model outperformed simple automatic nod-generating models using only bag of words and chance level. The results also indicate that dialog act, part of speech, the large-scale Japanese thesaurus, and word position are useful for generating nods. We also evaluated, through subjective evaluation, if our nod-generation model is useful with conversational agents. The results show that the nodding generated with our model improves user impressions of naturalness, humanness, likability, and reliability toward a conversational agent. Ryo Ishii, Taichi Katayama, Ryuichiro Higashinaka, Junji Tomita |
RO-MAN | 3 |
| 2018 | Role play-based question-answering by real users for building chatbots with consistent personalitiesabstractHaving consistent personalities is important for chatbots if we want them to be believable.Typically, many questionanswer pairs are prepared by hand for achieving consistent responses; however, the creation of such pairs is costly.In this study, our goal is to collect a large number of question-answer pairs for a particular character by using role playbased question-answering in which multiple users play the roles of certain characters and respond to questions by online users.Focusing on two famous characters, we conducted a large-scale experiment to collect question-answer pairs by using real users.We evaluated the effectiveness of role play-based questionanswering and found that, by using our proposed method, the collected pairs lead to good-quality chatbots that exhibit consistent personalities. Ryuichiro Higashinaka, Masahiro Mizukami, Hidetoshi Kawabata, Emi Yamaguchi, Noritake Adachi, Junji Tomita |
SIGDIAL Conference | 1 |
| 2018 | Neural Dialogue Context Online End-of-Turn DetectionabstractThis paper proposes a fully neural network based dialogue-context online end-of-turn detection method that can utilize longrange interactive information extracted from both target speaker's and interlocutor's utterances.In the proposed method, we combine multiple time-asynchronous long short-term memory recurrent neural networks, which can capture target speaker's and interlocutor's multiple sequential features, and their interactions.On the assumption of applying the proposed method to spoken dialogue systems, we introduce target speaker's acoustic sequential features and interlocutor's linguistic sequential features, each of which can be extracted in an online manner.Our evaluation confirms the effectiveness of taking dialogue context formed by the target speaker's utterances and interlocutor's utterances into consideration. Ryo Masumura, Tomohiro Tanaka, Atsushi Ando, Ryo Ishii, Ryuichiro Higashinaka, Yushi Aono |
SIGDIAL Conference | 5 |
| 2018 | Introduction method for argumentative dialogue using paired question-answering interchange about personalityabstractTo provide a better discussion experience in current argumentative dialogue systems, it is necessary for the user to feel motivated to participate, even if the system already responds appropriately.In this paper, we propose a method that can smoothly introduce argumentative dialogue by inserting an initial discourse, consisting of question-answer pairs concerning personality.The system can induce interest of the users prior to agreement or disagreement during the main discourse.By disclosing their interests, the users will feel familiarity and motivation to further engage in the argumentative dialogue and understand the system's intent.To verify the effectiveness of a questionanswer dialogue inserted before the argument, a subjective experiment was conducted using a text chat interface.The results suggest that inserting the questionanswer dialogue enhances familiarity and naturalness.Notably, the results suggest that women more than men regard the dialogue as more natural and the argument as deepened, following an exchange concerning personality. Kazuki Sakai, Ryuichiro Higashinaka, Yuichiro Yoshikawa, Hiroshi Ishiguro, Junji Tomita |
SIGDIAL Conference | 2 |
| 2017 | Understanding the Semantic Structures of Tables with a Hybrid Deep Neural Network ArchitectureabstractWe propose a new deep neural network architecture, TabNet, for table type classification. Table type is essential information for exploring the power of Web tables, and it is important to understand the semantic structures of tables in order to classify them correctly. A table is a matrix of texts, analogous to an image, which is a matrix of pixels, and each text consists of a sequence of tokens. Our hybrid architecture mirrors the structure of tables: its recurrent neural network (RNN) encodes a sequence of tokens for each cell to create a 3d table volume like image data, and its convolutional neural network (CNN) captures semantic features, e.g., the existence of rows describing properties, to classify tables. Experiments using Web tables with various structures and topics demonstrated that TabNet achieved considerable improvements over state-of-the-art methods specialized for table classification and other deep neural network architectures. Kyosuke Nishida, Kugatsu Sadamitsu, Ryuichiro Higashinaka, Yoshihiro Matsuo |
AAAI | 3 |
| 2017 | Online End-of-Turn Detection from Speech Based on Stacked Time-Asynchronous Sequential Networks
Ryo Masumura, Taichi Asami, Hirokazu Masataki, Ryo Ishii, Ryuichiro Higashinaka |
INTERSPEECH | 5 |
| 2017 | Zero-Shot Learning for Natural Language Understanding Using Domain-Independent Sequential Structure and Question Types
Kugatsu Sadamitsu, Yukinori Homma, Ryuichiro Higashinaka, Yoshihiro Matsuo |
INTERSPEECH | 3 |
| 2016 | The dialogue breakdown detection challenge: Task description, datasets, and evaluation metrics
Ryuichiro Higashinaka, Kotaro Funakoshi, Yuka Kobayashi, Michimasa Inaba |
LREC | 1 |
| 2016 | Analyzing Post-dialogue Comments by Speakers - How Do Humans Personalize Their Utterances in Dialogue? -abstractWe have been studying methods to personalize system utterances for users in casual conversations.We know that personalization is important, but no wellestablished way to personalize system utterances for users has been proposed.In this paper, we report the results of our experiment that examined how humans personalize utterances when speaking to each other in casual conversations.In particular, we elicited post-dialogue comments from speakers and analyzed the comments to determine what they thought about the dialogues while they engaged in them.In addition, by analyzing the effectiveness of their thoughts, we found that dialogue strategies for personalization related to "topic elaboration", "topic changing" and "tempo" significantly increased the satisfaction with regard to the dialogues. Toru Hirano, Ryuichiro Higashinaka, Yoshihiro Matsuo |
SIGDIAL Conference | 2 |
| 2016 | Towards an Entertaining Natural Language Generation System: Linguistic Peculiarities of Japanese Fictional CharactersabstractOne of the key ways of making dialogue agents more attractive as conversation partners is characterization, as it makes the agents more friendly, humanlike, and entertaining.To build such characters, utterances suitable for the characters are usually manually prepared.However, it is expensive to do this for a large number of utterances.To reduce this cost, we are developing a natural language generator that can express the linguistic styles of particular characters.To this end, we analyze the linguistic peculiarities of Japanese fictional characters (such as those in cartoons or comics and mascots), which have strong characteristics.The contributions of this study are that we (i) present comprehensive categories of linguistic peculiarities of Japanese fictional characters that cover around 90% of such characters' linguistic peculiarities and (ii) reveal the impact of each category on characterizing dialogue system utterances. Chiaki Miyazaki, Toru Hirano, Ryuichiro Higashinaka, Yoshihiro Matsuo |
SIGDIAL Conference | 3 |
| 2015 | Fatal or not? Finding errors that lead to dialogue breakdowns in chat-oriented dialogue systemsabstractRyuichiro Higashinaka, Masahiro Mizukami, Kotaro Funakoshi, Masahiro Araki, Hiroshi Tsukahara, Yuka Kobayashi. Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing. 2015. Ryuichiro Higashinaka, Masahiro Mizukami, Kotaro Funakoshi, Masahiro Araki, Hiroshi Tsukahara, Yuka Kobayashi |
EMNLP | 1 |
| 2015 | Automatic conversion of sentence-end expressions for utterance characterization of dialogue systems
Chiaki Miyazaki, Toru Hirano, Ryuichiro Higashinaka, Toshiro Makino, Yoshihiro Matsuo |
PACLIC | 3 |
| 2015 | Discourse Relation Recognition by Comparing Various Units of Sentence Expression with Recursive Neural Network
Atsushi Otsuka, Toru Hirano, Chiaki Miyazaki, Ryo Masumura, Ryuichiro Higashinaka, Toshiro Makino, Yoshihiro Matsuo |
PACLIC | 5 |
| 2015 | Towards Taxonomy of Errors in Chat-oriented Dialogue SystemsabstractRyuichiro Higashinaka, Kotaro Funakoshi, Masahiro Araki, Hiroshi Tsukahara, Yuka Kobayashi, Masahiro Mizukami. Proceedings of the 16th Annual Meeting of the Special Interest Group on Discourse and Dialogue. 2015. Ryuichiro Higashinaka, Kotaro Funakoshi, Masahiro Araki, Hiroshi Tsukahara, Yuka Kobayashi, Masahiro Mizukami |
SIGDIAL Conference | 1 |
| 2014 | Towards an open-domain conversational system fully based on natural language processing
Ryuichiro Higashinaka, Kenji Imamura, Toyomi Meguro, Chiaki Miyazaki, Nozomi Kobayashi, Hiroaki Sugiyama, Toru Hirano, Toshiro Makino, Yoshihiro Matsuo |
COLING | 1 |
| 2014 | Predicate-Argument Structure Analysis with Zero-Anaphora Resolution for Dialogue Systems
Kenji Imamura, Ryuichiro Higashinaka, Tomoko Izumi |
COLING | 2 |
| 2014 | Evaluating coherence in open domain conversational systems
Ryuichiro Higashinaka, Toyomi Meguro, Kenji Imamura, Hiroaki Sugiyama, Toshiro Makino, Yoshihiro Matsuo |
INTERSPEECH | 1 |
| 2014 | Large-scale Collection and Analysis of Personal Question-Answer Pairs for Conversational Agents
Hiroaki Sugiyama, Toyomi Meguro, Ryuichiro Higashinaka, Yasuhiro Minami |
IVA | 3 |
| 2014 | Extraction of Daily Changing Words for Question Answering
Kugatsu Sadamitsu, Ryuichiro Higashinaka, Yoshihiro Matsuo |
LREC | 2 |
| 2014 | Open-domain utterance generation using phrase pairs based on dependency relationsabstractThe development of open-domain conversational systems remains difficult since user utterances are widely varied for such systems to respond appropriately. To address this is- sue, previous research has retrieved sentences from the web as system utterances by shallow sentence matching with user utterances. However, since the retrieved sentences include the inherent contexts of the document in which the sentences originally appeared, the retrieved sentences have the possibility of containing information that is irrelevant to user utter-ances. We propose combining two strongly related semantic units (phrase pairs with dependency relations) to create a system utterance. Here, the first semantic unit is the one found in the user utterance and the second semantic unit is the one that has a dependency relation with the first one in a large text cor- pus. This way, we can guarantee that the generated utterance is related to the input user utterance. Our experiments, which examine the appropriateness of response sentences, show that our proposed method significantly outperforms other retrieval and rule-based approaches. Hiroaki Sugiyama, Toyomi Meguro, Ryuichiro Higashinaka, Yasuhiro Minami |
SLT | 3 |
| 2013 | Using role play for collecting question-answer pairs for dialogue agents
Ryuichiro Higashinaka, Kohji Dohsaka, Hideki Isozaki |
INTERSPEECH | 1 |
| 2013 | Estimating callers' levels of knowledge in call center dialogues
Chiaki Miyazaki, Ryuichiro Higashinaka, Toshiro Makino, Yoshihiro Matsuo |
INTERSPEECH | 2 |
| 2013 | Open-domain Utterance Generation for Conversational Dialogue Systems using Web-scale Dependency Structures
Hiroaki Sugiyama, Toyomi Meguro, Ryuichiro Higashinaka, Yasuhiro Minami |
SIGDIAL Conference | 3 |
| 2012 | Automatic Detection of Metonymies using Associative Relations between Words
Takehiro Teraoka, Ryuichiro Higashinaka, Jun Okamoto, Shun Ishizaki |
CogSci | 2 |
| 2012 | Creating an Extended Named Entity Dictionary from Wikipedia
Ryuichiro Higashinaka, Kugatsu Sadamitsu, Kuniko Saito, Toshiro Makino, Yoshihiro Matsuo |
COLING | 1 |
| 2011 | Building a conversational model from two-tweetsabstractThe current problem in building a conversational model from Twitter data is the scarcity of long conversations. According to our statistics, more than 90% of conversations in Twitter are composed of just two tweets. Previous work has utilized only conversations lasting longer than three tweets for dialogue modeling so that more than a single interaction can be successfully modeled. This paper verifies, by experiment, that two-tweet exchanges alone can lead to conversational models that are comparable to those made from longer-tweet conversations. This finding leverages the value of Twitter as a dialogue corpus and opens the possibility of better conversational modeling using Twitter data. Ryuichiro Higashinaka, Noriaki Kawamae, Kugatsu Sadamitsu, Yasuhiro Minami, Toyomi Meguro, Kohji Dohsaka, Hirohito Inagaki |
ASRU | 1 |
| 2011 | Wizard of Oz evaluation of listening-oriented dialogue control using POMDPabstractWe have been working on dialogue control for listening agents. In our previous study [1], we proposed a dialogue control method that maximizes user satisfaction using partially observable Markov decision processes (POMDPs) and evaluated it by a dialogue simulation. We found that it significantly outperforms other stochastic dialogue control methods. However, this result does not necessarily mean that our method works as well in real dialogues with human users. Therefore, in this paper, we evaluate our dialogue control method by a Wizard of Oz (WoZ) experiment. The experimental results show that our POMDP-based method achieves significantly higher user satisfaction than other stochastic models, confirming the validity of our approach. This paper is the first to show the usefulness of POMDP-based dialogue control using human users when the target function is to maximize user satisfaction. Toyomi Meguro, Yasuhiro Minami, Ryuichiro Higashinaka, Kohji Dohsaka |
ASRU | 3 |
| 2011 | Unsupervised Clustering of Utterances Using Non-Parametric Bayesian MethodsabstractUnsupervised clustering of utterances can be useful for the modeling of dialogue acts for dialogue applications. Previously, the Chinese restaurant process (CRP), a non-parametric Bayesian method, has been introduced and has shown promising results for the clustering of utterances in dialogue. This paper newly introduces the infinite HMM, which is also a nonparametric Bayesian method, and verifies its effectiveness. Experimental results in two dialogue domains show that the infinite HMM, which takes into account the sequence of utterances in its clustering process, significantly outperforms the CRP. Although the infinite HMM outperformed other methods, we also found that clustering complex dialogue data, such as humanhuman conversations, is still hard when compared to humanmachine dialogues. Index Terms: Unsupervised clustering, Nonparametric Bayesian methods, Chinese restaurant process, Infinite HMM Ryuichiro Higashinaka, Noriaki Kawamae, Kugatsu Sadamitsu, Yasuhiro Minami, Toyomi Meguro, Kohji Dohsaka, Hirohito Inagaki |
INTERSPEECH | 1 |
| 2011 | Evaluation of Listening-Oriented Dialogue Control Rules Based on the Analysis of HMMsabstractWe have been working on listening-oriented dialogues for the purpose of building listening agents. In our previous work [1], we trained hidden Markov models (HMMs) from listeningoriented dialogues (LoDs) between humans, and by analyzing them, discovered a distinguishing dialogue flow of LoD. For example, listeners suppress their information giving and selfdisclosure, and instead, increase acknowledgments and questions to elicit speakers ’ utterances. As an initial step for building listening agents, we decided to create dialogue control rules based on our analysis of the HMMs. We built our rule-based system and compared it with three other systems by a Wizard of Oz (WoZ) experiment. As a result, we found that our rule-based system achieved as much user satisfaction as human listeners. Index Terms: Listening-oriented dialogue, Dialogue system, Wizard of Oz Toyomi Meguro, Yasuhiro Minami, Ryuichiro Higashinaka, Kohji Dohsaka |
INTERSPEECH | 3 |
| 2010 | Controlling Listening-oriented Dialogue using Partially Observable Markov Decision Processes
Toyomi Meguro, Ryuichiro Higashinaka, Yasuhiro Minami, Kohji Dohsaka |
COLING | 2 |
| 2010 | User-adaptive Coordination of Agent Communicative Behavior in Spoken Dialogue
Kohji Dohsaka, Atsushi Kanemoto, Ryuichiro Higashinaka, Yasuhiro Minami, Eisaku Maeda |
SIGDIAL Conference | 3 |
| 2010 | Modeling User Satisfaction Transitions in Dialogues from Overall Ratings
Ryuichiro Higashinaka, Yasuhiro Minami, Kohji Dohsaka, Toyomi Meguro |
SIGDIAL Conference | 1 |
| 2010 | Improving hmm-based extractive summarization for multi-domain contact center dialoguesabstractThis paper reports the improvements we made to our previously proposed hidden Markov model (HMM) based summarization method for multi-domain contact center dialogues. Since the method relied on Viterbi decoding for selecting utterances to include in a summary, it had the inability to control compression rates. We enhance our method by using the forward-backward algorithm together with integer linear programming (ILP) to enable the control of compression rates, realizing summaries that contain as many domain-related utterances and as many important words as possible within a predefined character length. Using call transcripts as input, we verify the effectiveness of our enhancement. Ryuichiro Higashinaka, Yasuhiro Minami, Hitoshi Nishikawa, Kohji Dohsaka, Toyomi Meguro, Satoshi Kobashikawa, Hirokazu Masataki, Osamu Yoshioka, Satoshi Takahashi, Gen-ichiro Kikui |
SLT | 1 |
| 2010 | Trigram dialogue control using POMDPsabstractThis paper proposes hybrid dialogue control of both trigram and POMDP dialogue controls by extending our proposed method that uses two approaches: automatically acquiring POMDP structures and rewards for target dialogues through Dynamic Bayesian Networks (DBNs) with a large amount of dialogue data and reflecting action predictive probabilities into the POMDP structures. In this extension, we modify the action predictive probabilities to treat trigram dialogue controls. Experimental results show that the proposed method can treat a trigram dialogue control with robustness for erroneous conditions and can simultaneously maximize trigram probability and the dialogue evaluations obtained from users. Yasuhiro Minami, Ryuichiro Higashinaka, Kohji Dohsaka, Toyomi Meguro, Eisaku Maeda |
SLT | 2 |
| 2010 | Trend detection modelabstractThis paper presents a topic model that detects topic distributions over time. Our proposed model, Trend Detection Model (TDM) introduces a latent trend class variable into each document. The trend class has a probability distribution over topics and a continuous distribution over time. Experiments using our data set show that TDM is useful as a generative model in the analysis of the evolution of trends. Noriaki Kawamae, Ryuichiro Higashinaka |
WWW | 2 |
| 2009 | Effects of Conversational Agents on Human Communication in Thought-Evoking Multi-Party Dialogues
Kohji Dohsaka, Ryota Asai, Ryuichiro Higashinaka, Yasuhiro Minami, Eisaku Maeda |
SIGDIAL Conference | 3 |
| 2009 | Analysis of Listening-Oriented Dialogue for Building Listening Agents
Toyomi Meguro, Ryuichiro Higashinaka, Kohji Dohsaka, Yasuhiro Minami, Hideki Isozaki |
SIGDIAL Conference | 2 |
| 2008 | Corpus-based Question Answering for why-Questions
Ryuichiro Higashinaka, Hideki Isozaki |
IJCNLP | 1 |
| 2008 | Effects of self-disclosure and empathy in human-computer dialogueabstractTo build trust or cultivate long-term relationships with users, conversational systems need to perform social dialogue. To date, research has primarily focused on the overall effect of social dialogue in human-computer interaction, leading to little work on the effects of individual linguistic phenomena within social dialogue. This paper investigates such individual effects through dialogue experiments. Focusing on self-disclosure and empathic utterances (agreement and disagreement), we empirically calculate their contributions to the dialogue quality. Our analysis shows that (1) empathic utterances by users are strong indicators of increasing closeness and user satisfaction, (2) the system's empathic utterances are effective for inducing empathy from users, and (3) self-disclosure by users increases when users have positive preferences on topics being discussed. Ryuichiro Higashinaka, Kohji Dohsaka, Hideki Isozaki |
SLT | 1 |
| 2008 | "Who is this" quiz dialogue system and users' evaluationabstractIn order to design a dialogue system that users enjoy and want to be near for a long time, it is important to know the effect of the system's action on users. This paper describes ldquoWho is thisrdquo quiz dialogue system and its users' evaluation. Its quiz-style information presentation has been found effective for educational tasks. In our ongoing effort to make it closer to a conversational partner, we implemented the system as a stuffed-toy (or CG equivalent). Quizzes are automatically generated from Wikipedia articles, rather than from hand-crafted sets of biographical facts. Network mining is utilized to prepare adaptive system responses. Experiments showed the effectiveness of person network and the relationship of user attribute and interest level. Minako Sawaki, Yasuhiro Minami, Ryuichiro Higashinaka, Kohji Dohsaka, Eisaku Maeda |
SLT | 3 |
| 2008 | Automatically Acquiring Causal Expression Patterns from Relation-annotated Corpora to Improve Question Answering for why-QuestionsabstractThis article describes our approach for answering why-questions that we initially introduced at NTCIR-6 QAC-4. The approach automatically acquires causal expression patterns from relation-annotated corpora by abstracting text spans annotated with a causal relation and by mining syntactic patterns that are useful for distinguishing sentences annotated with a causal relation from those annotated with other relations. We use these automatically acquired causal expression patterns to create features to represent answer candidates, and use these features together with other possible features related to causality to train an answer candidate ranker that maximizes the QA performance with regards to the corpus of why-questions and answers. NAZEQA, a Japanese why-QA system based on our approach, clearly outperforms baselines with a Mean Reciprocal Rank (top-5) of 0.223 when sentences are used as answers and with a MRR (top-5) of 0.326 when paragraphs are used as answers, making it presumably the best-performing fully implemented why-QA system. Experimental results also verified the usefulness of the automatically acquired causal expression patterns. Ryuichiro Higashinaka, Hideki Isozaki |
ACM Trans. Asian Lang. Inf. Process. | 1 |
| 2007 | Learning to Rank Definitions to Generate Quizzes for Interactive Information Presentation
Ryuichiro Higashinaka, Kohji Dohsaka, Hideki Isozaki |
ACL | 1 |
| 2007 | The world of mushrooms: human-computer interaction prototype systems for ambient intelligenceabstractOur new research project called concentrates on the creation of new lifestyles through research on communication science and intelligence integration. It is premised on the creation of such virtual communication partners as fairies and goblins that can be constantly at our side. We call these virtual communication partners mushrooms.To show the essence of ambient intelligence, we developed two multimodal prototype systems: mushrooms that watch, listen, and answer questions and a Quizmaster Mushroom. These two systems work in real time using speech, sound, dialogue, and vision technologies.We performed preliminary experiments with the Quizmaster Mushroom. The results showed that the system can transmit knowledge to users while they are playing the quizzes.Furthermore, through the two mushrooms, we found policies for design effects in multimodal interface and integration. Yasuhiro Minami, Minako Sawaki, Kohji Dohsaka, Ryuichiro Higashinaka, Kentaro Ishizuka, Hideki Isozaki, Tatsushi Matsubayashi, Masato Miyoshi, Atsushi Nakamura, Takanobu Oba, Hiroshi Sawada, Takeshi Yamada, Eisaku Maeda |
ICMI | 4 |
| 2007 | Effects of quiz-style information presentation on user understandingabstractThis paper proposes quiz-style information presentation for interactive systems as a means to improve user understanding in educational tasks. Since the nature of quizzes can highly motivate users to stay voluntarily engaged in the interaction and keeptheir attention on receiving information, it is expectedthat information presented as quizzes can be better understood by users. To verify the effectiveness of the approach, we implemented read-out and quiz systems and performed comparison experiments using human subjects. In the task of memorizing biographical facts, the results showed that user understanding for the quiz system was significantly better than that for the read-out system, and that the subjects were more willing to use the quiz system despite the long duration of the quizzes. This indicates that quiz-style information presentation promotes engagement in the interaction with the system, leading to the improved user understanding. Index Terms: information presentation, quiz, user understanding Ryuichiro Higashinaka, Kohji Dohsaka, Shigeaki Amano, Hideki Isozaki |
INTERSPEECH | 1 |
| 2006 | Learning to Generate Naturalistic Utterances Using Reviews in Spoken Dialogue SystemsabstractSpoken language generation for dialogue systems requires a dictionary of mappings between semantic representations of concepts the system wants to express and realizations of those concepts. Dictionary creation is a costly process; it is currently done by hand for each dialogue domain. We propose a novel unsupervised method for learning such mappings from user reviews in the target domain, and test it on restaurant reviews. We test the hypothesis that user reviews that provide individual ratings for distinguished attributes of the domain entity make it possible to map review sentences to their semantic representation with high precision. Experimental analyses show that the mappings learned cover most of the domain ontology, and provide good linguistic variation. A subjective user evaluation shows that the consistency between the semantic representations and the learned realizations is high and that the naturalness of the realizations is higher than a hand-crafted baseline. Ryuichiro Higashinaka, Rashmi Prasad, Marilyn A. Walker |
ACL | 1 |
| 2006 | Simulating Cub Reporter Dialogues: The collection of naturalistic human-human dialogues for information access to text archives
Emma Barker, Ryuichiro Higashinaka, François Mairesse, Robert J. Gaizauskas, Marilyn A. Walker, Jonathan Foster |
LREC | 2 |
| 2006 | Incorporating discourse features into confidence scoring of intention recognition results in spoken dialogue systems
Ryuichiro Higashinaka, Katsuhito Sudoh, Mikio Nakano |
Speech Commun. | 1 |
| 2005 | Incorporating Discourse Features into Confidence Scoring of Intention Recognition Results in Spoken Dialogue SystemsabstractThe paper proposes a method for the confidence scoring of intention recognition results in spoken dialogue systems. To achieve tasks, a spoken dialogue system has to recognize user intentions. However, because of speech recognition errors and ambiguity in user utterances, it sometimes has difficulty recognizing them correctly. Confidence scoring allows errors to be detected in intention recognition results and has proved useful for dialogue management. Conventional methods use the features obtained from speech recognition results for single utterances for confidence scoring. However, this may be insufficient since the intention recognition result is a result of discourse processing. We propose incorporating discourse features for a more accurate confidence scoring of intention recognition results. Experimental results show that incorporating discourse features significantly improves the confidence scoring. Ryuichiro Higashinaka, Katsuhito Sudoh, Mikio Nakano |
ICASSP (1) | 1 |
| 2003 | Corpus-Based Discourse Understanding in Spoken Dialogue SystemsabstractThis paper concerns the discourse understanding process in spoken dialogue systems. This process enables the system to understand user utterances based on the context of a dialogue. Since multiple candidates for the understanding result can be obtained for a user utterance due to the ambiguity of speech understanding, it is not appropriate to decide on a single understanding result after each user utterance. By holding multiple candidates for understanding results and resolving the ambiguity as the dialogue progresses, the discourse understanding accuracy can be improved. This paper proposes a method for resolving this ambiguity based on statistical information obtained from dialogue corpora. Unlike conventional methods that use hand-crafted rules, the proposed method enables easy design of the discourse understanding process. Experiment results have shown that a system that exploits the proposed method performs sufficiently and that holding multiple candidates for understanding results is effective. Ryuichiro Higashinaka, Mikio Nakano, Kiyoaki Aikawa |
ACL | 1 |
| 2003 | Evaluating discourse understanding in spoken dialogue systemsabstractThis paper describes a method for creating an evaluation measure for discourse understanding in spoken dialogue systems. Discourse understanding means utterance understanding taking the context into account. Since the measure needs to be determined based on its correlation with the system’s performance, conventional measures, such as the concept error rate, cannot be easily applied. Using the multiple linear regression analysis, we have previously shown that the weighted sum of various metrics concerning dialogue states can be used for the evaluation of discourse understanding in a single domain. This paper reports the progress of our work: verification of our approach by additional experiments in another domain. The support vector regression method performs better than the multiple linear regression method in creating the measure, indicating non-linearity in mapping the metrics to the system’s performance. The results give strong support for our approach and hint at its suitability as a universal evaluation measure for discourse understanding. 1. Ryuichiro Higashinaka, Noboru Miyazaki, Mikio Nakano, Kiyoaki Aikawa |
INTERSPEECH | 1 |
| 2002 | Interactive Paraphrasing Based on Linguistic Annotation
Ryuichiro Higashinaka, Katashi Nagao |
COLING | 1 |
| 2002 | A method for evaluating incremental utterance understanding in spoken dialogue systemsabstractIn single utterance understanding, which does not include discourse understanding, the concept error rate (CER), or the keyword error rate, has been widely used as an evaluation measure for utterance understanding. However, the CER cannot be used for evaluating systems that understand user utterances based on previous user utterances. In this paper, we propose a method for evaluating incremental utterance understanding, which involves speech recognition, language understanding and discourse processing in spoken dialogue systems, by finding a measure that correlates closely with the system’s performance based on dialogue states and their way of update. We defined dialogue performance by task completion time, and performed a multiple linear regression analysis using task completion time as the explained variable and various metrics concerning dialogue states as explaining variables. The obtained multiple regression model fits comparatively well and shows validity as an evaluation measure. 1. Ryuichiro Higashinaka, Noboru Miyazaki, Mikio Nakano, Kiyoaki Aikawa |
INTERSPEECH | 1 |
| 2002 | Learning decision trees to determine turn-taking by spoken dialogue systemsabstractThis paper presents a method for deciding the timing of turn-taking in spoken dialogue systems. This method uses a decision tree learned from the corpus of dialogues between human users and systems in which desirable turn-taking behaviors are annotated by hand. It utilizes a variety of attributes, such as recognition and understanding results and prosodic information. Unlike most of the existing systems it enables spoken dialogue systems to decide the timing of turn-taking based on not only pauses but also other features, so that users can speak to the system even if they put pauses in the middle of their utterances. The result of a preliminary experiment shows that the learned decision tree outperforms the baseline strategy, which takes turn at every user pauses. 1. Ryo Sato, Ryuichiro Higashinaka, Masafumi Tamoto, Mikio Nakano, Kiyoaki Aikawa |
INTERSPEECH | 2 |