Koichiro Yoshino

dblp:59/8883 · DBLP profile ↗
← Back
64ranked-venue papers
13as first author
26since 2021 · last 2025
0000-0001-6579-361XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 54 · 12 first-author · 21 since 2021Graphics, computer vision, multimedia, augmented reality and games · 21 · 1 first-author · 11 since 2021Human-computer interaction and ubiquitous computing · 4 · 3 since 2021Databases, data management, data science and information retrieval · 3 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Systems, architecture and hardware · 2 · 2 since 2021
YearPublicationVenuePosition
2025 Disambiguating Reference in Visually Grounded Dialogues through Joint Modeling of Textual and Multimodal Semantic Structures
abstract
Multimodal reference resolution, including phrase grounding, aims to understand the semantic relations between mentions and realworld objects.Phrase grounding between images and their captions is a well-established task.In contrast, for real-world applications, it is essential to integrate textual and multimodal reference resolution to unravel the reference relations within dialogue, especially in handling ambiguities caused by pronouns and ellipses.This paper presents a framework that unifies textual and multimodal reference resolution by mapping mention embeddings to object embeddings and selecting mentions or objects based on their similarity. 1 Our experiments show that learning textual reference resolution, such as coreference resolution and predicateargument structure analysis, positively affects performance in multimodal reference resolution.In particular, our model with coreference resolution performs better in pronoun phrase grounding than representative models for this task, MDETR and GLIP.Our qualitative analysis demonstrates that incorporating textual reference relations strengthens the confidence scores between mentions, including pronouns and predicates, and objects, which can reduce the ambiguities that arise in visually grounded dialogues.
Shun Inadumi, Nobuhiro Ueda, Koichiro Yoshino
ACL (1)3
2025 Teaching Text Agents to Learn Sequential Decision Making from Failure
abstract
Text-based reinforcement-learning agents improve their policies by interacting with their environments to collect more training data.However, these self-collected data inevitably contain intermediate failed actions caused by attempting physically infeasible behaviors and/or hallucinations.Directly learning a policy from such trajectories can reinforce incorrect behaviors and reduce task success rates.In this paper, we propose a failed action-aware objective that suppresses the negative impact of failed actions during training by assigning zero return based on textual feedback.Building on this objective, we introduce a perturbation method that leverages unsuccessful trajectories to construct new successful ones that share the same goal.This allows agents to benefit from diverse experiences without further interaction with the environment.Experiments in ALFWorld and ScienceWorld demonstrate that our method significantly outperforms strong baselines and generalizes across environments.
Canasai Kruengkrai, Koichiro Yoshino
ACL (1)2
2025 Efficient ASR Domain Adaptation with Long Noun Phrases: Harnessing the Linguistic Characteristics of Japanese
abstract
Adapting ASR systems to the vocabulary used in the application domain is crucial in practical usage. Domain adaptation in ASR is often performed by fine-tuning with sentence-level synthesized speech or methods like logit bias and LM fusion. However, these approaches can be inefficient, costly, and vulnerable to unseen terms. This paper proposes a novel adaptation method that leverages the linguistic property of Japanese: the interchangeability of noun phrases within a sentence. The proposed method connects domain-specific nouns with postpositional particles, enabling adaptation to multiple terms with minimal synthetic data while preserving the natural flow of word sequences. Experimental results demonstrate that models trained with the proposed training data achieve high performance in overall recognition and recognizing domainspecific words with reduced training cost. The results also indicate that the proposed method enhances robustness against speech from noisy environments.
Shusuke Komatsu, Kazuyo Onishi, Koki Tanaka, Koichiro Yoshino
ASRU5
2025 Training Dialogue Systems by AI Feedback for Improving Overall Dialogue Impression
abstract
To improve user engagement during conversations with dialogue systems, we must improve individual dialogue responses and dialogue impressions such as consistency, personality, and empathy throughout the entire dialogue. While such dialogue systems have been developing rapidly with the help of large language models (LLMs), reinforcement learning from AI feedback (RLAIF) has attracted attention to align LLM-based dialogue models for such dialogue impressions. In RLAIF, a reward model based on another LLM is used to create a training signal for an LLM-based dialogue model using zero-shot/few-shot prompting techniques. However, evaluating an entire dialogue only by prompting LLMs is challenging. In this study, the supervised fine-tuning (SFT) of LLMs prepared reward models corresponding to 12 metrics related to the impression of the entire dialogue for evaluating dialogue responses. We tuned our dialogue models using the reward model signals as feedback to improve the impression of the system. The results of automatic and human evaluations showed that tuning the dialogue model using our reward model corresponding to dialogue impression improved the evaluation of individual metrics and the naturalness of the dialogue response.
Kai Yoshida, Masahiro Mizukami, Seiya Kawano, Canasai Kruengkrai, Hiroaki Sugiyama, Koichiro Yoshino
ICASSP6
2025 Rapport-Building Dialogue Strategies for Deeper Connection: Integrating Proactive Behavior, Personalization, and Aizuchi Backchannels
Muhammad Yeza Baihaqi, Angel F. Garcia Contreras, Seiya Kawano, Koichiro Yoshino
INTERSPEECH4
2025 Co-Speech Motion for Virtual Agents in Dialogue Using LLM-Driven Primitive Action Selection
Muhammad Yeza Baihaqi, Angel F. Garcia Contreras, Seiya Kawano, Koichiro Yoshino
INTERSPEECH4
2025 Dialogue Response Prefetching Based on Semantic Similarity and Prediction Confidence of Language Model
Kiyotada Mori, Seiya Kawano, Angel F. Garcia Contreras, Koichiro Yoshino
INTERSPEECH4
2025 What Do Humans Hear When Interacting? Experiments on Selective Listening for Evaluating ASR of Spoken Dialogue Systems
Kiyotada Mori, Seiya Kawano, Carlos Toshinori Ishi, Angel F. Garcia Contreras, Koichiro Yoshino
INTERSPEECH6
2025 J-ORA: A Framework and Multimodal Dataset for Japanese Object Identification, Reference, Action Prediction in Robot Perception
abstract
We introduce J-ORA, a novel multimodal dataset that bridges the gap in robot perception by providing detailed object attribute annotations within Japanese human-robot dialogue scenarios. J-ORA is designed to support three critical perception tasks, object identification, reference resolution, and next-action prediction, by leveraging a comprehensive template of attributes (e.g., category, color, shape, size, material, and spatial relations). Extensive evaluations with both proprietary and open-source Vision Language Models (VLMs) reveal that incorporating detailed object attributes substantially improves multimodal perception performance compared to without object attributes. Despite the improvement, we find that there still exists a gap between proprietary and open-source VLMs. In addition, our analysis of object affordances demonstrates varying abilities in understanding object functionality and contextual relationships across different VLMs. These findings underscore the importance of rich, context-sensitive attribute annotations in advancing robot perception in dynamic environments. Code and data available at https://github.com/jatuhurrra/J-ORA.
Jesse Atuhurra, Hidetaka Kamigaito, Taro Watanabe, Koichiro Yoshino
IROS4
2025 Using Language Models to Generate and Forget the Narrative Memories of an Assistive Robot
Angel F. Garcia Contreras, Wen-Yu Chang, Seiya Kawano, Yun-Nung Chen, Koichiro Yoshino
MMM (5)5
2025 RoboDJ: Live Commentary Robots System Driven by Physical- and Cyber-World Observations
Yasutomo Kawanishi, Yutaka Nakamura, Taiken Shintani, Carlos Toshinori Ishi, Seiya Kawano, Koichiro Yoshino, Takashi Minato, Michihiko Minoh
MMM (5)6
2025 What Should Autonomous Robots Verbalize and What Should They Not?
Daichi Yoshihara, Akishige Yuguchi, Seiya Kawano, Takamasa Iio, Koichiro Yoshino
MMM (5)5
2025 LLM-Driven Approach for Motion Control in Human-Robot Dialogue for Elevating Engagement
abstract
Non-verbal behaviors, such as body movements, play a crucial role in enhancing a robot’s speech to elevate engagement in human-robot dialogue. Many existing approach based on rules offered natural and engaging motions aligned with the robot’s utterances but required significant resources to maintain. Recent methods leveraging large language models (LLMs) offer a promising alternative to reduce these costs. However, there is a trade-off between flexibility and safety when determining whether the language model should generate motions based on joint angle parameters or action primitives. In this study, we evaluated two LLM-based motion control models: one for motion generation based on joint angle parameters (LLM-GJA) and the other for motion generation based on primitive actions (LLM-GPA). Our human evaluations indicated that directly generating joint angles outperformed generating action primitives in naturalness, timing consistency, and overall engagement, even achieving performance comparable to rule-based systems. This work highlights the potential of LLMs in generating expressive and contextually appropriate robot motions at the joint angle level.
Muhammad Yeza Baihaqi, Angel F. Garcia Contreras, Seiya Kawano, Koichiro Yoshino
RO-MAN4
2025 Modeling Turn-Taking Speed and Speaker Characteristics
abstract
Modeling turn-taking speed while considering speaker characteristics and the relationships between speakers is essential for realizing dialogue systems capable of natural interactions. In this study, we focused on dialogue participants’ roles, relationships, and personality, analyzing and modeling turn-taking speeds observed in real conversations. The analysis confirmed that the expression of these attributes—role, relationship, and personality—is closely associated with turn-taking speed. Based on these findings, we constructed a model that predicts the distribution of turn-taking speeds according to each attribute using a gamma distribution. Evaluation results demonstrated that appropriate parameter fitting to the three-parameter gamma distribution enables effective modeling of turn-taking speeds based on participants’ roles, relationships, and characteristics.
Kazuyo Onishi, Hien Ohnaka, Koichiro Yoshino
SIGDIAL3
2025 Multi-step or Direct: A Proactive Home-Assistant System Based on Commonsense Reasoning
abstract
There is a growing expectation for the realization of proactive home-assistant robots that can assist users in their daily lives. It is essential to develop a framework that closely observes the user’s surrounding context, selectively extracts relevant information, and infers the user’s needs to proactively propose appropriate assistance. In this study, we first extend the Do-I-Demand dataset to define expected proactive assistance actions in domestic situations, where users make ambiguous utterances. These behaviors were defined based on common patterns of support that a majority of users would expect from a robot. We subsequently constructed a framework that infers users’ expected assistance actions from ambiguous utterances through commonsense reasoning. We explored two approaches: (1) multi-step reasoning using COMET as a commonsense reasoning engine, and (2) direct reasoning using large language models. Our experimental results suggest that both the multi-step and direct reasoning methods can successfully derive necessary assistance actions even when dealing with ambiguous user utterances.
Konosuke Yamasaki, Shohei Tanaka, Akishige Yuguchi, Seiya Kawano, Koichiro Yoshino
SIGDIAL5
2024 ClaimBrush: A Novel Framework for Automated Patent Claim Refinement Based on Large Language Models
abstract
Automatic refinement of patent claims in patent applications is crucial from the perspective of intellectual property strategy. In this paper, we propose "ClaimBrush," a novel framework for automated patent claim refinement that includes a dataset and a rewriting model. We constructed a dataset for training and evaluating patent claim rewriting models by collecting a large number of actual patent claim rewriting cases from the patent examination process. Using the constructed dataset, we built an automatic patent claim rewriting model by fine-tuning a large language model. Furthermore, we enhanced the performance of the automatic patent claim rewriting model by applying preference optimization based on a prediction model of patent examiners’ Office Actions. The experimental results showed that our proposed rewriting model outperformed heuristic baselines and zero-shot learning in state-of-the-art large language models. Moreover, preference optimization based on patent examiners’ preferences boosted the performance of patent claim refinement.
Seiya Kawano, Hirofumi Nonaka, Koichiro Yoshino
IEEE Big Data3
2024 A Gaze-grounded Visual Question Answering Dataset for Clarifying Ambiguous Japanese Questions
abstract
Situated conversations, which refer to visual information as visual question answering (VQA), often contain ambiguities caused by reliance on directive information. This problem is exacerbated because some languages, such as Japanese, often omit subjective or objective terms. Such ambiguities in questions are often clarified by the contexts in conversational situations, such as joint attention with a user or user gaze information. In this study, we propose the Gaze-grounded VQA dataset (GazeVQA) that clarifies ambiguous questions using gaze information by focusing on a clarification process complemented by gaze information. We also propose a method that utilizes gaze target estimation results to improve the accuracy of GazeVQA tasks. Our experimental results showed that the proposed method improved the performance in some cases of a VQA system on GazeVQA and identified some typical problems of GazeVQA tasks that need to be improved.
Shun Inadumi, Seiya Kawano, Akishige Yuguchi, Yasutomo Kawanishi, Koichiro Yoshino
LREC/COLING5
2024 J-CRe3: A Japanese Conversation Dataset for Real-world Reference Resolution
abstract
Understanding expressions that refer to the physical world is crucial for such human-assisting systems in the real world, as robots that must perform actions that are expected by users. In real-world reference resolution, a system must ground the verbal information that appears in user interactions to the visual information observed in egocentric views. To this end, we propose a multimodal reference resolution task and construct a Japanese Conversation dataset for Real-world Reference Resolution (J-CRe3). Our dataset contains egocentric video and dialogue audio of real-world conversations between two people acting as a master and an assistant robot at home. The dataset is annotated with crossmodal tags between phrases in the utterances and the object bounding boxes in the video frames. These tags include indirect reference relations, such as predicate-argument structures and bridging references as well as direct reference relations. We also constructed an experimental model and clarified the challenges in multimodal reference resolution tasks.
Nobuhiro Ueda, Hideko Habe, Akishige Yuguchi, Seiya Kawano, Yasutomo Kawanishi, Sadao Kurohashi, Koichiro Yoshino
LREC/COLING7
2024 Rapport-Driven Virtual Agent: Rapport Building Dialogue Strategy for Improving User Experience at First Meeting
Muhammad Yeza Baihaqi, Angel F. Garcia Contreras, Seiya Kawano, Koichiro Yoshino
INTERSPEECH4
2024 Capturing Contact Surfaces by a Frustrated Total Internal Reflection System Using a Curved Plate for Comparison of the Beginning of Touching Motions by Humanitude Experts and Novices
abstract
Analyzing time-series changes of the contact surface by touching is important to elucidate the touching skills of Humanitude as one of the pervasive multimodal comprehensive care methodologies. For the analysis, there is a frustrated total internal reflection (FTIR) method to capture contact surfaces on a transparent flat plate by a camera. However, this conventional flat plate is far from the actual surfaces of care receivers because the surfaces of humans consist of curved shapes. In this paper, we propose an FTIR sensing system using a transparent curved plate to capture more ideal contact states with a surface shape more similar to the human body. We collect the contact surface data of the beginning of touching motions by Humanitude experts and novices using the FTIR sensing system with the curved and flat plates. Then, we compare the data by the experts and novices in terms of time-series contact areas and the quantitative indices and subjective evaluation and discuss the analysis results. Through these experiments, we confirm that the proposed system has the potential for the novices to perform more correctly the beginning of Humanitude's touching motions.
Akishige Yuguchi, Mayuki Toyoda, Sung-Gwi Cho, Atsushi Nakazawa, Jun Takamatsu, Koichiro Yoshino, Tsukasa Ogasawara
SMC6
2024 Overview of the Tenth Dialog System Technology Challenge: DSTC10
abstract
This article introduces the Tenth Dialog System Technology Challenge (DSTC-10). This edition of the DSTC focuses on applying end-to-end dialog technologies for five distinct tasks in dialog systems, namely 1. Incorporation of Meme images into open domain dialogs, 2. Knowledge-grounded Task-oriented Dialogue Modeling on Spoken Conversations, 3. Situated Interactive Multimodal dialogs, 4. Reasoning for Audio Visual Scene-Aware Dialog, and 5. Automatic Evaluation and Moderation of Open-domainDialogue Systems. This article describes the task definition, provided datasets, baselines, and evaluation setup for each track. We also summarize the results of the submitted systems to highlight the general trends of the state-of-the-art technologies for the tasks.
Koichiro Yoshino, Yun-Nung Chen, Paul A. Crook, Satwik Kottur, Jinchao Li, Behnam Hedayatnia, Seungwhan Moon, Zhengcong Fei, Zekang Li, Jinchao Zhang 0001, Yang Feng 0004, Jie Zhou 0016, Seokhwan Kim, Yang Liu 0004, Di Jin 0005, Alexandros Papangelis, Karthik Gopalakrishnan 0001, Dilek Hakkani-Tür, Babak Damavandi, Alborz Geramifard, Chiori Hori, Chen Zhang 0020, Haizhou Li 0001, João Sedoc, Luis Fernando D'Haro, Rafael E. Banchs, Alexander I. Rudnicky
IEEE ACM Trans. Audio Speech Lang. Process.1
2023 Operative Action Captioning for Estimating System Actions
abstract
Human-assistive systems, such as robots, need to correctly understand the surrounding situation based on obser-vations and output the required support actions for humans. Language is one of the important channels to communicate with humans, and robots are required to have the ability to express their understanding and action-planning results. In this study, we propose a new task of operative action captioning that estimates and verbalizes the actions to be taken by the system in a human-assisting domain. We constructed a system that outputs a verbal description of a possible operative action that changes the current state to the given target state. We collected a dataset consisting of two images as observations, which express the current state and the state changed by actions and a caption that describes the actions that change the current state to the target state, by crowdsourcing in daily life situations. Then we constructed a system that estimates an operative action by a caption. Since the operative action's caption is expected to contain some state-changing actions, we use scene graph prediction as an auxiliary task because the events written in the scene graphs correspond to the state changes. Experimental results showed that our system successfully described the operative actions that should be conducted between the current and target states. The auxiliary tasks that predict the scene graphs improved the quality of the estimation results.
Taiki Nakamura, Seiya Kawano, Akishige Yuguchi, Yasutomo Kawanishi, Koichiro Yoshino
ICRA5
2023 Reflective action selection based on positive-unlabeled learning and causality detection model
abstract
Task-oriented dialogue systems need to take appropriate actions not only for clear user requests but also for ambiguous and vague ones. In this study, “ambiguous” denotes that although users have potential requests, they failed to clearly define and verbalize their content and conditions which can be associated with system actions. For such ambiguous requests, taking reflective actions is one plausible choice for such systems. In our study, “reflective” denotes taking actions that satisfy user requests before the users themselves clarify their demands. We constructed such a reflective dialogue agent by collecting a corpus that includes pairs of ambiguous user requests and corresponding reflective system actions on sightseeing navigation with a smartphone. Since annotating every possible combination of user requests and system actions is impossible, this study built a corpus where one reflective action is annotated to one ambiguous user request. To train an action selection model on such incomplete training data in which only one action is associated with a request, we applied the positive/unlabeled (PU) learning method, which assumes that only part of the data is labeled with positive examples. In addition, we enhanced the action selection by extracting and distilling knowledge that corresponds to causality from the training data using a causality detection model. The experimental results show that both the PU learning method and the causality detection model improved the performances of the reflective action selection compared to the conventional positive/negative (PN) learning method.
Shohei Tanaka, Koichiro Yoshino, Katsuhito Sudoh, Satoshi Nakamura 0001
Comput. Speech Lang.2
2022 Butsukusa: A Conversational Mobile Robot Describing Its Own Observations and Internal States
abstract
This paper presents an autonomous conversational mobile robot Butsukusa that can describe its own observations and internal states during patrolling tasks. The proposed robot can observe the surrounding environment using the recognition module for objects, humans, environment, localization, and speech and then move autonomously around an indoor living space. Interaction skills via language are required for the robot to perform in such human-centered spaces. To investigate a better communication protocol with users, we evaluate various language generation patterns based on different observations and interaction patterns. The evaluation results indicate that the importance of describing the robot's observation results and internal states, as well as the necessity of an appropriate description, depends on the situation.
Akishige Yuguchi, Seiya Kawano, Koichiro Yoshino, Carlos Toshinori Ishi, Yasutomo Kawanishi, Yutaka Nakamura, Takashi Minato, Yasuki Saito, Michihiko Minoh
HRI3
2022 Multimodal Persuasive Dialogue Corpus using Teleoperated Android
Seiya Kawano, Muteki Arioka, Akishige Yuguchi, Kenta Yamamoto, Koji Inoue, Tatsuya Kawahara, Satoshi Nakamura 0001, Koichiro Yoshino
INTERSPEECH8
2021 ARTA: Collection and Classification of Ambiguous Requests and Thoughtful Actions
abstract
Human-assisting systems such as dialogue systems must take thoughtful, appropriate actions not only for clear and unambiguous user requests, but also for ambiguous user requests, even if the users themselves are not aware of their potential requirements.To construct such a dialogue agent, we collected a corpus and developed a model that classifies ambiguous user requests into corresponding system actions.In order to collect a high-quality corpus, we asked workers to input antecedent user requests whose pre-defined actions could be regarded as thoughtful.Although multiple actions could be identified as thoughtful for a single user request, annotating all combinations of user requests and system actions is impractical.For this reason, we fully annotated only the test data and left the annotation of the training data incomplete.In order to train the classification model on such training data, we applied the positive/unlabeled (PU) learning method, which assumes that only a part of the data is labeled with positive examples.The experimental results show that the PU learning method achieved better performance than the general positive/negative (PN) learning method to classify thoughtful actions given an ambiguous user request.
Shohei Tanaka, Koichiro Yoshino, Katsuhito Sudoh, Satoshi Nakamura 0001
SIGDIAL2
2020 Improving Spoken Language Understanding by Wisdom of Crowds
abstract
Spoken language understanding (SLU), which converts user requests in natural language to machine-interpretable expressions, is becoming an essential task.The lack of training data is an important problem, especially for new system tasks, because existing SLU systems are based on statistical approaches.In this paper, we proposed to use two sources of the "wisdom of crowds," crowdsourcing and knowledge community website, for improving the SLU system.We firstly collected paraphrasing variations for new system tasks through crowdsourcing as seed data, and then augmented them using similar questions from a knowledge community website.We investigated the effects of the proposed data augmentation method in SLU task, even with small seed data.In particular, the proposed architecture augmented more than 120,000 samples to improve SLU accuracies.
Koichiro Yoshino, Kana Ikeuchi, Katsuhito Sudoh, Satoshi Nakamura 0001
COLING1
2020 Emotional Speech Corpus for Persuasive Dialogue System
abstract
Expressing emotion is known as an efficient way to persuade one’s dialogue partner to accept one’s claim or proposal. Emotional expression in speech can express the speaker’s emotion more directly than using only emotion expression in the text, which will lead to a more persuasive dialogue. In this paper, we built a speech dialogue corpus in a persuasive scenario that uses emotional expressions to build a persuasive dialogue system with emotional expressions. We extended an existing text dialogue corpus by adding variations of emotional responses to cover different combinations of broad dialogue context and a variety of emotional states by crowd-sourcing. Then, we recorded emotional speech consisting of of collected emotional expressions spoken by a voice actor. The experimental results indicate that the collected emotional expressions with their speeches have higher emotional expressiveness for expressing the system’s emotion to users.
Sara Asai, Koichiro Yoshino, Seitaro Shinagawa, Sakriani Sakti, Satoshi Nakamura 0001
LREC2
2020 Overview of the seventh Dialog System Technology Challenge: DSTC7
Luis Fernando D'Haro, Koichiro Yoshino, Chiori Hori, Tim K. Marks, Lazaros Polymenakos, Jonathan K. Kummerfeld, Michel Galley, Xiang Gao 0011
Comput. Speech Lang.2
2019 Neural Conversation Model Controllable by Given Dialogue Act Based on Adversarial Learning and Label-aware Objective
abstract
Building a controllable neural conversation model (NCM) is an important task.In this paper, we focus on controlling the responses of NCMs by using dialogue act labels of responses as conditions.We introduce an adversarial learning framework for the task of generating conditional responses with a new objective to a discriminator, which explicitly distinguishes sentences by using labels.This change strongly encourages the generation of label-conditioned sentences.We compared the proposed method with some existing methods for generating conditional responses.The experimental results show that our proposed method has higher controllability for dialogue acts even though it has higher or comparable naturalness to existing methods.
Seiya Kawano, Koichiro Yoshino, Satoshi Nakamura 0001
INLG2
2019 An Incremental Turn-Taking Model for Task-Oriented Dialog Systems
abstract
In a human-machine dialog scenario, deciding the appropriate time for the machine to take the turn is an open research problem. In contrast, humans engaged in conversations are able to timely decide when to interrupt the speaker for competitive or non-competitive reasons. In state-of-the-art turn-by-turn dialog systems the decision on the next dialog action is taken at the end of the utterance. In this paper, we propose a token-by-token prediction of the dialog state from incremental transcriptions of the user utterance. To identify the point of maximal understanding in an ongoing utterance, we a) implement an incremental Dialog State Tracker which is updated on a token basis (iDST) b) re-label the Dialog State Tracking Challenge 2 (DSTC2) dataset and c) adapt it to the incremental turn-taking experimental scenario. The re-labeling consists of assigning a binary value to each token in the user utterance that allows to identify the appropriate point for taking the turn. Finally, we implement an incremental Turn Taking Decider (iTTD) that is trained on these new labels for the turn-taking decision. We show that the proposed model can achieve a better performance compared to a deterministic handcrafted turn-taking algorithm.
Andrei Catalin, Koichiro Yoshino, Yukitoshi Murase, Satoshi Nakamura 0001, Giuseppe Riccardi
INTERSPEECH2
2019 Overview of the sixth dialog system technology challenge: DSTC6
Chiori Hori, Julien Perez, Ryuichiro Higashinaka, Takaaki Hori, Y-Lan Boureau, Michimasa Inaba, Yuiko Tsunomori, Tetsuro Takahashi, Koichiro Yoshino, Seokhwan Kim
Comput. Speech Lang.9
2019 Associative knowledge feature vector inferred on external knowledge base for dialog state tracking
Yukitoshi Murase, Koichiro Yoshino, Satoshi Nakamura 0001
Comput. Speech Lang.2
2019 Positive Emotion Elicitation in Chat-Based Dialogue Systems
abstract
We aim to draw on an important overlooked potential of affective dialogue systems-their application to promote positive emotional states, similar to that of emotional support between humans. This can be achieved by eliciting a more positive emotional valence throughout a dialogue system interaction, i.e., positive emotion elicitation. Existing works on emotion elicitation have not yet paid attention to the emotional benefit for the users. Moreover, a positive emotion elicitation corpus does not yet exist despite the growing number of emotion-rich corpora. Towards this goal, first, we propose a response retrieval approach for positive emotion elicitation by utilizing examples of emotion appraisal from a dialogue corpus. Second, we efficiently construct a corpus using the proposed retrieval method, by replacing responses in a dialogue with those that elicit a more positive emotion. We validate the corpus through crowdsourcing to ensure its quality. Finally, we propose a novel neural network architecture for an emotion-sensitive neural chat-based dialogue system, optimized on the constructed corpus to elicit positive emotion. Objective and subjective evaluations show that the proposed methods result in dialogue responses that are more natural and elicit a more positive emotional response. Further analyses of the results are discussed in this paper.
Nurul Lubis, Sakriani Sakti, Koichiro Yoshino, Satoshi Nakamura 0001
IEEE ACM Trans. Audio Speech Lang. Process.3
2018 Eliciting Positive Emotion through Affect-Sensitive Dialogue Response Generation: A Neural Network Approach
abstract
An emotionally-competent computer agent could be a valuable assistive technology in performing various affective tasks. For example caring for the elderly, low-cost ubiquitous chat therapy, and providing emotional support in general, by promoting a more positive emotional state through dialogue system interaction. However, despite the increase of interest in this task, existing works face a number of shortcomings: system scalability, restrictive modeling, and weak emphasis on maximizing user emotional experience. In this paper, we build a fully data driven chat-oriented dialogue system that can dynamically mimic affective human interactions by utilizing a neural network architecture. In particular, we propose a sequence-to-sequence response generator that considers the emotional context of the dialogue. An emotion encoder is trained jointly with the entire network to encode and maintain the emotional context throughout the dialogue. The encoded emotion information is then incorporated in the response generation process. We train the network with a dialogue corpus that contains positive-emotion eliciting responses, collected through crowd-sourcing. Objective evaluation shows that incorporation of emotion into the training process helps reduce the perplexity of the generated responses, even when a small dataset is used. Subsequent subjective evaluation shows that the proposed method produces responses that are more natural and likely to elicit a more positive emotion.
Nurul Lubis, Sakriani Sakti, Koichiro Yoshino, Satoshi Nakamura 0001
AAAI3
2018 TRANS-AM: Discovery Method of Optimal Input Vectors Corresponding to Objective Variables
Yu Suzuki 0001, Koichiro Yoshino, Satoshi Nakamura 0001
DaWaK3
2018 Dialogue Scenario Collection of Persuasive Dialogue with Emotional Expressions via Crowdsourcing
Koichiro Yoshino, Yoko Ishikawa, Masahiro Mizukami, Yu Suzuki 0001, Sakriani Sakti, Satoshi Nakamura 0001
LREC1
2018 Japanese Dialogue Corpus of Information Navigation and Attentive Listening Annotated with Extended ISO-24617-2 Dialogue Act Tags
Koichiro Yoshino, Hiroki Tanaka, Kyoshiro Sugiyama, Makoto Kondo, Satoshi Nakamura 0001
LREC1
2018 Unsupervised Counselor Dialogue Clustering for Positive Emotion Elicitation in Neural Dialogue System
abstract
Positive emotion elicitation seeks to improve user's emotional state through dialogue system interaction, where a chatbased scenario is layered with an implicit goal to address user's emotional needs.Standard neural dialogue system approaches still fall short in this situation as they tend to generate only short, generic responses.Learning from expert actions is critical, as these potentially differ from standard dialogue acts.In this paper, we propose using a hierarchical neural network for response generation that is conditioned on 1) expert's action, 2) dialogue context, and 3) user emotion, encoded from user input.We construct a corpus of interactions between a counselor and 30 participants following a negative emotional exposure to learn expert actions and responses in a positive emotion elicitation scenario.Instead of relying on the expensive, labor intensive, and often ambiguous human annotations, we unsupervisedly cluster the expert's responses and use the resulting labels to train the network.Our experiments and evaluation show that the proposed approach yields lower perplexity and generates a larger variety of responses.
Nurul Lubis, Sakriani Sakti, Koichiro Yoshino, Satoshi Nakamura 0001
SIGDIAL Conference3
2018 Optimizing Neural Response Generator with Emotional Impact Information
abstract
The potential of dialogue systems to address user's emotional need has steadily grown. In particular, we focus on dialogue systems application to promote positive emotional states, similar to that of emotional support between humans. Positive emotion elicitation takes form as chat-based dialogue interactions that is layered with an implicit goal to improve user's emotional state. To this date, existing approaches have only relied on mimicking the target responses without considering their emotional impact, i.e. the change of emotional state they cause on the listener, in the model itself. In this paper, we propose explicitly utilizing emotional impact information to optimize neural dialogue system towards generating responses that elicit positive emotion. We examine two emotion-rich corpora with different data collection scenarios: Wizard-of-Oz and spontaneous. Evaluation shows that the proposed method yields lower perplexity, as well as produces responses that are perceived as more natural and likely to elicit a more positive emotion.
Nurul Lubis, Sakriani Sakti, Koichiro Yoshino, Satoshi Nakamura 0001
SLT3
2017 Processing negative emotions through social communication: Multimodal database construction and analysis
abstract
Emotion-rich data is pre-requisite in the efforts of transferring emotional aspects of human communication into Human-Computer Interaction (HCI). An important facet of human social-affective interaction is its ability to facilitate social sharing of emotion, a fundamental part of the emotional processes. When conducted properly, such an interaction can give a positive effect to emotion-related problems. However, there is still a lack of resources that are: 1) explicitly designed for studying the emotional problems commonly encountered in everyday life, and 2) involving a professional as an expert in the conversation. In this paper, we present recordings of dyadic social-affective interactions between a professional counselor as an expert and 30 participants, summing up to 23 hours and 41 minutes of material. In each interaction, a negative emotion inducer is shown to the dyad, and the goal of the expert is to aid emotion processing and elicit a positive emotional change through the interaction. Specifically, we aim to observe how an external party can guide and facilitate emotion processing, especially after a negative emotional response in a commonly encountered social situation. The construction, development, and analysis of the database is detailed in this paper.
Nurul Lubis, Michael Heck, Sakriani Sakti, Koichiro Yoshino, Satoshi Nakamura 0001
ACII4
2017 Neural Machine Translation via Binary Code Prediction
abstract
In this paper, we propose a new method for calculating the output layer in neural machine translation systems.The method is based on predicting a binary code for each word and can reduce computation time/memory requirements of the output layer to be logarithmic in vocabulary size in the best case.In addition, we also introduce two advanced approaches to improve the robustness of the proposed model: using error-correcting codes and combining softmax and binary codes.Experiments on two English ↔ Japanese bidirectional translation tasks show proposed models achieve BLEU scores that approach the softmax, while reducing memory usage to the order of less than 1/10 and improving decoding speed on CPUs by x5 to x10.
Yusuke Oda, Philip Arthur, Graham Neubig, Koichiro Yoshino, Satoshi Nakamura 0001
ACL (1)4
2017 Acquisition and Assessment of Semantic Content for the Generation of Elaborateness and Indirectness in Spoken Dialogue Systems
abstract
In a dialogue system, the dialogue manager selects one of several system actions and thereby determines the system’s behaviour. Defining all possible system actions in a dialogue system by hand is a tedious work. While efforts have been made to automatically generate such system actions, those approaches are mostly focused on providing functional system behaviour. Adapting the system behaviour to the user becomes a difficult task due to the limited amount of system actions available. We aim to increase the adaptability of a dialogue system by automatically generating variants of system actions. In this work, we introduce an approach to automatically generate action variants for elaborateness and indirectness. Our proposed algorithm extracts RDF triplets from a knowledge base and rates their relevance to the original system action to find suitable content. We show that the results of our algorithm are mostly perceived similarly to human generated elaborateness and indirectness and can be used to adapt a conversation to the current user and situation. We also discuss where the results of our algorithm are still lacking and how this could be improved: Taking into account the conversation topic as well as the culture of the user is likely to have beneficial effect on the user’s perception.
Louisa Pragst, Koichiro Yoshino, Wolfgang Minker, Satoshi Nakamura 0001, Stefan Ultes
IJCNLP(1)2
2017 Information Navigation System with Discovering User Interests
abstract
We demonstrate an information navigation system for sightseeing domains that has a dialogue interface for discovering user interests for tourist activities.The system discovers interests of a user with focus detection on user utterances, and proactively presents related information to the discovered user interest.A partially observable Markov decision process (POMDP)-based dialogue manager, which is extended with user focus states, controls the behavior of the system to provide information with several dialogue acts for providing information.We transferred the belief-update function and the policy of the manager from other system trained on a di↵erent domain to show the generality of defined dialogue acts for our information navigation system.
Koichiro Yoshino, Yu Suzuki 0001, Satoshi Nakamura 0001
SIGDIAL Conference1
2017 Semantically readable distributed representation learning for social media mining
abstract
The problem with distributed representations generated by neural networks is that the meaning of the features is difficult to understand. We propose a new method that gives a specific meaning to each node of a hidden layer by introducing a manually created word semantic vector dictionary into the initial weights and by using paragraph vector models. Our experimental results demonstrated that weights obtained based on learning and weights based on the dictionary are more strongly correlated in a closed test and more weakly correlated in an open test, compared with the results of a control test. Additionally, we found that the learned vector are better than the performance of the existing paragraph vector in the evaluation of the sentiment analysis task. Finally, we determined the readability of document embedding in a user test. The definition of readability in this paper is that people can understand the meaning of large weighted features of distributed representations. A total of 52.4% of the top five weighted hidden nodes were related to tweets where one of the paragraph vector models learned the document embedding. Because each hidden node maintains a specific meaning, the proposed method succeeds in improving readability.
Ikuo Keshi, Yu Suzuki 0001, Koichiro Yoshino, Satoshi Nakamura 0001
WI3
2016 Unsupervised Joint Estimation of Grapheme-to-Phoneme Conversion Systems and Acoustic Model Adaptation for Non-Native Speech Recognition
Satoshi Tsujioka, Sakriani Sakti, Koichiro Yoshino, Graham Neubig, Satoshi Nakamura 0001
INTERSPEECH3
2016 Construction of Japanese Audio-Visual Emotion Database and Its Application in Emotion Recognition
Nurul Lubis, Randy Gomez, Sakriani Sakti, Keisuke Nakamura, Koichiro Yoshino, Satoshi Nakamura 0001, Kazuhiro Nakadai
LREC5
2016 Parallel Speech Corpora of Japanese Dialects
Koichiro Yoshino, Naoki Hirayama, Shinsuke Mori, Fumihiko Takahashi, Katsutoshi Itoyama, Hiroshi G. Okuno
LREC1
2016 Cultural Communication Idiosyncrasies in Human-Computer Interaction
abstract
In this work, we investigate whether the cultural idiosyncrasies found in humanhuman interaction may be transferred to human-computer interaction.With the aim of designing a culture-sensitive dialogue system, we designed a user study creating a dialogue in a domain that has the potential capacity to reveal cultural differences.The dialogue contains different options for the system output according to cultural differences.We conducted a survey among Germans and Japanese to investigate whether the supposed differences may be applied in human-computer interaction.Our results show that there are indeed differences, but not all results are consistent with the cultural models.
Juliana Miehle, Koichiro Yoshino, Louisa Pragst, Stefan Ultes, Satoshi Nakamura 0001, Wolfgang Minker
SIGDIAL Conference2
2016 Analyzing the Effect of Entrainment on Dialogue Acts
abstract
Entrainment is a factor in dialogue that affects not only human-human but also human-machine interaction. While entrainment on the lexical level is well documented, less is known about how entrainment affects dialogue on a more abstract, structural level. In this paper, we investigate the effect of entrainment on dialogue acts and on lexical choice given dialogue acts, as well as how entrainment changes during a dialogue. We also define a novel measure of entrainment to measure these various types of entrainment. These results may serve as guidelines for dialogue systems that would like to entrain with users in a similar manner.
Masahiro Mizukami, Koichiro Yoshino, Graham Neubig, David R. Traum, Satoshi Nakamura 0001
SIGDIAL Conference2
2016 The fifth dialog state tracking challenge
abstract
Dialog state tracking - the process of updating the dialog state after each interaction with the user - is a key component of most dialog systems. Following a similar scheme to the fourth dialog state tracking challenge, this edition again focused on human-human dialogs, but introduced the task of cross-lingual adaptation of trackers. The challenge received a total of 32 entries from 9 research groups. In addition, several pilot track evaluations were also proposed receiving a total of 16 entries from 4 groups. In both cases, the results show that most of the groups were able to outperform the provided baselines for each task.
Seokhwan Kim, Luis Fernando D'Haro, Rafael E. Banchs, Jason D. Williams, Matthew Henderson, Koichiro Yoshino
SLT6
2016 Deep bottleneck features and sound-dependent i-vectors for simultaneous recognition of speech and environmental sounds
abstract
In speech interfaces, it is often necessary to understand the overall auditory environment, not only recognizing what is being said, but also being aware of the location or actions surrounding the utterance. However, automatic speech recognition (ASR) becomes difficult when recognizing speech with environmental sounds. Standard solutions treat environmental sounds as noise, and remove them to improve ASR performance. On the other hand, most studies on environmental sounds construct classifiers for environmental sounds only, without interference of spoken utterances. But, in reality, such separate situations almost never exist. This study attempts to address the problem of simultaneous recognition of speech and environmental sounds. Particularly, we examine the possibility of using deep neural network (DNN) techniques to recognize speech and environmental sounds simultaneously, and improve the accuracy of both tasks under respective noisy conditions. First, we investigate DNN architectures including two parallel single-task DNNs, and a single multi-task DNN. However, we found direct multi-task learning of simultaneous speech and environmental recognition to be difficult. Therefore, we further propose a method that combines bottleneck features and sound-dependent i-vectors within this framework. Experimental evaluation results reveal that the utilizing bottleneck features and i-vectors as the input of DNNs can help to improve accuracy of each recognition task.
Sakriani Sakti, Seiji Kawanishi, Graham Neubig, Koichiro Yoshino, Satoshi Nakamura 0001
SLT4
2015 A study of social-affective communication: Automatic prediction of emotion triggers and responses in television talk shows
abstract
Advancements in spoken language technologies have allowed users to interact with computers in an increasingly natural manner. However, most conversational agents or dialogue systems are yet to consider emotional awareness in interaction. To consider emotion in these situations, social-affective knowledge in conversational agents is essential. In this paper, we present a study of the social-affective process in natural conversation from television talk shows. We analyze occurrences of emotion (emotional responses), and the events that elicit them (emotional triggers). We then utilize our analysis for prediction to model the ability of a dialogue system to decide an action and response in an affective interaction. This knowledge has great potential to incorporate emotion into human-computer interaction. Experiments in two languages, English and Indonesian, show that automatic prediction performance surpasses random guessing accuracy.
Nurul Lubis, Sakriani Sakti, Graham Neubig, Koichiro Yoshino, Tomoki Toda, Satoshi Nakamura 0001
ASRU4
2015 Adaptive selection from multiple response candidates in example-based dialogue
abstract
In spoken dialogue systems, dialogue modeling is one of the most important factors for contributing to user satisfaction improvement. Especially in Example-Based Dialogue Modeling (EBDM), effective methods to build dialogue example databases and to select response utterances from examples are the keys for improving dialogue quality. In dialogue corpora, it often have plural appropriate responses for one utterance. However, the system merges these plural appropriate responses into the one system response, thus, it does not try to use plural responses properly by user preference. In fact, responses that each user thinks to be preferable are different. In this paper, we propose a framework that select an appropriate response from plural appropriate response candidates satisfies users. It has a multi-response example database, and selects an appropriate response based on collaborative filtering. Experimental results showed that the proposed framework were successfully choosing appropriate responses, considering multi-response candidates improves user satisfaction to 4.1 from 3.7 of single response, and the adaptive response selection method increased user satisfaction from 3.7 to 4.3.
Masahiro Mizukami, Hideaki Kizuki, Toshio Nomura, Graham Neubig, Koichiro Yoshino, Sakriani Sakti, Tomoki Toda, Satoshi Nakamura 0001
ASRU5
2015 Conversational system for information navigation based on POMDP with user focus tracking
Koichiro Yoshino, Tatsuya Kawahara
Comput. Speech Lang.1
2015 Automatic Speech Recognition for Mixed Dialect Utterances by Mixing Dialect Language Models
abstract
This paper presents an automatic speech recognition (ASR) system that accepts a mixture of various kinds of dialects. The system recognizes dialect utterances on the basis of the statistical simulation of vocabulary transformation and combinations of several dialect models. Previous dialect ASR systems were based on handcrafted dictionaries for several dialects, which involved costly processes. The proposed system statistically trains transformation rules between a common language and dialects, and simulates a dialect corpus for ASR on the basis of a machine translation technique. The rules are trained with small sets of parallel corpora to make up for the lack of linguistic resources on dialects. The proposed system also accepts mixed dialect utterances that contain a variety of vocabularies. In fact, spoken language is not a single dialect but a mixed dialect that is affected by the circumstances of speakers’ backgrounds (e.g., native dialects of their parents or where they live). We addressed two methods to combine several dialects appropriately for each speaker. The first was recognition with language models of mixed dialects with automatically estimated weights that maximized the recognition likelihood. This method performed the best, but calculation was very expensive because it conducted grid searches of combinations of dialect mixing proportions that maximized the recognition likelihood. The second was integration of results of recognition from each single dialect language model. The improvements with this model were slightly smaller than those with the first method. Its calculation cost was, however, inexpensive and it worked in real-time on general workstations. Both methods achieved higher recognition accuracies for all speakers than those with the single dialect models and the common language model, and we could choose a suitable model for use in ASR that took into consideration the computational costs and recognition accuracies.
Naoki Hirayama, Koichiro Yoshino, Katsutoshi Itoyama, Shinsuke Mori, Hiroshi G. Okuno
IEEE ACM Trans. Audio Speech Lang. Process.2
2014 FlowGraph2Text: Automatic Sentence Skeleton Compilation for Procedural Text Generation
abstract
In this paper we describe a method for generating a procedural text given its flow graph representation. Our main idea is to automatically collect sen-tence skeletons from real texts by re-placing the important word sequences with their type labels to form a skeleton pool. The experimental results showed that our method is feasible and has a potential to generate natural sentences. 1
Shinsuke Mori, Hirokuni Maeta, Tetsuro Sasada, Koichiro Yoshino, Atsushi Hashimoto 0001, Takuya Funatomi, Yoko Yamakata
INLG4
2014 Information Navigation System Based on POMDP that Tracks User Focus
abstract
We present a spoken dialogue system for navigating information (such as news ar-ticles), and which can engage in small talk. At the core is a partially observ-able Markov decision process (POMDP), which tracks user’s state and focus of at-tention. The input to the POMDP is pro-vided by a spoken language understanding (SLU) component implemented with lo-gistic regression (LR) and conditional ran-dom fields (CRFs). The POMDP selects one of six action classes; each action class is implemented with its own module. 1
Koichiro Yoshino, Tatsuya Kawahara
SIGDIAL Conference1
2013 Incorporating semantic information to selection of web texts for language model of spoken dialogue system
abstract
A novel text selection approach for training a language model (LM) with Web texts is proposed for automatic speech recognition (ASR) of spoken dialogue systems. Compared to the conventional approach based on perplexity criterion, the proposed approach introduces a semantic-level relevance measure with the back-end knowledge base used in the dialogue system. We focus on the predicate-argument (P-A) structure characteristic to the domain in order to filter semantically relevant sentences in the domain. Several choices of statistical models and combination methods with the perplexity measure are investigated in this paper. Experimental evaluations in two different domains demonstrate the effectiveness and generality of the proposed approach. The combination method realizes significant improvement not only in ASR accuracy but also in semantic and dialogue-level accuracy.
Koichiro Yoshino, Shinsuke Mori, Tatsuya Kawahara
ICASSP1
2013 Predicate Argument Structure Analysis using Partially Annotated Corpora
Koichiro Yoshino, Shinsuke Mori, Tatsuya Kawahara
IJCNLP1
2013 Statistical Dialogue Management using Intention Dependency Graph
Koichiro Yoshino, Shinji Watanabe 0001, Jonathan Le Roux, John R. Hershey
IJCNLP1
2013 Automatic estimation of dialect mixing ratio for dialect speech recognition
abstract
This paper proposes methods for determining an appropriate mixing ratio of dialects in automatic speech recognition (ASR) for dialects. To handle ASR for various dialects, it has been reported to be effective to train a language model using a dialectmixed corpus. One reason behind this is geographical continuity of spoken dialect; we regard spoken dialect as a mixture of various dialects. This mixing ratio changes at every moment as well as depends on a speaker. We can improve recognition accuracybygivingan appropriatedialectmixingratio foraspeaker’s dialect. The mixing ratio is generally unknown and requires to be estimated and updated referring to input utterances. We handle two methods for updating it based on recognition results; one is to compute contribution of dialects for each recognized word, and the other is to predict mixture information referring to a whole recognized sentence based on topic modeling. The experimental result shows that the mixing ratio estimated by these methods realized higher recognition accuracy than a fixed mixing ratio. Index Terms: dialect, supervised latent Dirichlet allocation (sLDA), mixing ratio.
Naoki Hirayama, Koichiro Yoshino, Katsutoshi Itoyama, Shinsuke Mori, Hiroshi G. Okuno
INTERSPEECH2
2012 Language Modeling for Spoken Dialogue System based on Filtering using Predicate-Argument Structures
Koichiro Yoshino, Shinsuke Mori, Tatsuya Kawahara
COLING1
2011 Spoken Dialogue System based on Information Extraction using Similarity of Predicate Argument Structures
Koichiro Yoshino, Shinsuke Mori, Tatsuya Kawahara
SIGDIAL Conference1