VLDB 2026 Research / reviewers in the wild / expert
Kazunori Komatani
dblp:21/6468
· DBLP profile ↗
152ranked-venue papers
24as first author
24since 2021 · last 2026
0000-0002-6052-600XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 122 · 19 first-author · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 61 · 10 first-author · 6 since 2021Systems, architecture and hardware · 39Human-computer interaction and ubiquitous computing · 12 · 4 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 3Databases, data management, data science and information retrieval · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Suppressing Unnecessary Clarification Requests for Unknown Word Acquisition in Spoken Dialogue Using Syllable-Based ASR ConfidenceabstractClarification requests are a promising way for spoken dialogue systems to acquire unknown words from users, but asking too often can burden users. Because unknown words are not in the system’s vocabulary, utterances must first be represented as syllable sequences, which are then segmented into words. In this setting, syllable-based automatic speech recognition (S-ASR) errors can cause utterances containing only known words to appear to contain unknown words, leading to unnecessary clarification requests. To address this issue within a stream-based active learning framework, we extend the reinforcement learning policy state with recognition reliability features. Specifically, we incorporate two confidence measures derived from S-ASR to make clarification request selection sensitive to S-ASR errors. We further incorporate segmentation confidence over N-best hypotheses to reduce the impact of minor S-ASR errors. Experiments using pre-recorded speech data showed that the number of clarification requests on utterances affected by S-ASR errors was reduced by 1.34. The area under the learning curve for word segmentation also numerically increased by 0.07. Takumi Furuta, Ryu Takeda, Kazunori Komatani |
SIGDIAL | 3 |
| 2026 | Question Types for Knowledge Acquisition in LLM-based Dialogue Systems: Experiments with Simulated and Human UsersabstractTo acquire knowledge from users through dialogue, systems must decide not only what to ask but also how to ask it. Since users may not always be willing to answer such questions, question formulation affects both user experience and the information obtained. Prior work has examined question types in controlled, template-based settings, but it remains unclear whether similar effects are observed in LLM-based dialogues. We investigated the effects of question type on knowledge acquisition through experiments with both an LLM-based user simulator and crowdsourced human participants. We compared three question types: implicit questions, explicit questions, and wh-questions. The results showed no substantial differences in user annoyance among the question types. In contrast, the question types differed in how well they elicited correct responses: wh-questions were less effective, whereas implicit and explicit questions performed comparably. The findings suggest that candidate-guided question forms are useful when the system has a plausible candidate answer, whereas wh-questions may be appropriate when the system lacks sufficient confidence to ask a more specific question. Kazunori Komatani, Ryu Takeda, Mikio Nakano |
SIGDIAL | 1 |
| 2026 | Rethinking Binary Evaluation of Turn-Taking under Inherent AmbiguityabstractTurn-taking prediction models output probabilities of turn shifts, yet they are typically evaluated by thresholding these probabilities into binary decisions and comparing them against corpus-observed labels. This practice implicitly treats corpus-observed turn shifts as definitive ground truth, even though under inherent turn-taking ambiguity they reflect one realized interactional outcome among multiple plausible outcomes, rather than a uniquely correct binary label. We argue that binary evaluation is a practical simplification rather than a theoretical necessity. Instead, predicted probabilities should be evaluated at the distributional level without being reduced to binary decisions. To this end, we propose a distribution-based evaluation framework that compares model output distributions with reference distributions and measures their divergence using the Wasserstein distance. We further show how discrepancies between model predictions and corpus-observed turn shifts can be used as a basis for training-data refinement. Experiments on Japanese conversational data, using linguistic information alone, showed that the proposed refinement reduced distributional divergence, indicating better alignment between predicted probabilities and the reference distributions. The refinement also improved balanced accuracy in a supplementary binary evaluation. Yunosuke Kubo, Kenta Yamamoto, Ryu Takeda, Kazunori Komatani |
SIGDIAL | 4 |
| 2025 | Generating Diverse Personas for User Simulators to Test Interview Dialogue SystemsabstractThis paper addresses the issue of the significant labor required to test interview dialogue systems. While interview dialogue systems are expected to be useful in various scenarios, like other dialogue systems, testing them with human users requires significant effort and cost. Therefore, testing with user simulators can be beneficial. Since most conventional user simulators have been primarily designed for training task-oriented dialogue systems, little attention has been paid to the personas of the simulated users. During development, testing interview dialogue systems requires simulating a wide range of user behaviors, but manually creating a large number of personas is labor-intensive. We propose a method that automatically generates personas for user simulators using a large language model. Furthermore, by assigning personality traits related to communication styles when generating personas, we aim to increase the diversity of communication styles in the user simulator. Experimental results show that the proposed method enables the user simulator to generate utterances with greater variation. Mikio Nakano, Kazunori Komatani, Hironori Takeuchi |
SIGDIAL | 2 |
| 2025 | Learning to Ask Efficiently in Dialogue: Reinforcement Learning Extensions for Stream-based Active LearningabstractOne essential function of dialogue systems is the ability to ask questions and acquire necessary information from the user through dialogue. To avoid degrading user engagement through repetitive questioning, the number of such questions should be kept low. In this study, we cast knowledge acquisition through dialogue as stream-based active learning, exemplified by the segmentation of user utterances containing novel words. In stream-based active learning, data instances are presented sequentially, and the system selects an action for each instance based on an acquisition function that determines whether to request the correct answer from the oracle (in this case, the user). To improve the efficiency of training the acquisition function via reinforcement learning, we introduce two extensions: (1) a new action that performs semi-supervised learning, and (2) a state representation that takes the remaining budget into account. Our simulation-based experiments showed that these two extensions improved word segmentation performance with fewer questions for the user, compared to a baseline without these extensions. Issei Waki, Ryu Takeda, Kazunori Komatani |
SIGDIAL | 3 |
| 2024 | Incremental Multimodal Sentiment Analysis for HAIs Based on Multitask Active Learning with Interannotator AgreementabstractMultimodal sentiment analysis (MSA) is critical in developing empathetic and adaptive multimodal dialogue systems or conversational agents that can naturally interact with users by recognizing sentiment and engagement. Addressing the challenges of collecting labeled data for MSA in human-agent interaction (HAI), this study introduces an innovative approach that combines active learning and multitask learning. Our efficient sentiment recognition model leverages active learning to select informative data for learning models, significantly reducing the labor-intensive data labeling process. Furthermore, we employ multitask learning to improve annotation (label) quality by evaluating alignment with true labels and interannotator agreement, thus enhancing the reliability of sentiment annotations. We evaluate the proposed multitask and active learning methods via a human-agent multimodal dialogue dataset that includes various types of sentiment annotations, which are publicly available. The experimental results demonstrate that by learning to predict the agreement score, multitask learning becomes better than singletask learning at capturing the uncertainties in the data. This study lays the groundwork for incremental learning strategies in MSA, aiming to adaptively understand user sentiments in human-agent interactions. Thus Karnjanapatchara, Sixia Li, Candy Olivia Mawalim, Kazunori Komatani, Shogo Okada |
ACII | 4 |
| 2024 | Collecting Human-Agent Dialogue Dataset with Frontal Brain Signal toward Capturing Unexpressed SentimentabstractMultimodal information such as text and audiovisual data has been used for emotion/sentiment estimation during human-agent dialogue; however, user sentiments are not necessarily expressed explicitly during dialogues. Biosignals such as brain signals recorded using an electroencephalogram (EEG) sensor have been the subject of focus in affective computing regions to capture unexpressed emotional changes in a controlled experimental environment. In this study, we collect and analyze multimodal data with an EEG during a human-agent dialogue toward capturing unexpressed sentiment. Our contributions are as follows: (1) a new multimodal human-agent dialogue dataset is created, which includes not only text and audiovisual data but also frontal EEGs and physiological signals during the dialogue. In total, about 500-minute chat dialogues were collected from thirty participants aged 20 to 70. (2) We present a novel method for dealing with eye-blink noise for frontal EEGs denoising. This method applies facial landmark tracking to detect and delete eye-blink noise. (3) An experimental evaluation showed the effectiveness of the frontal EEGs. It improved sentiment estimation performance when used with other modalities by multimodal fusion, although it only has three channels. Shun Katada, Ryu Takeda, Kazunori Komatani |
LREC/COLING | 3 |
| 2024 | Personalized Sentiment Estimation Based on Recall and Resting Ratio of Frontal EEG
Shun Katada, Kazunori Komatani |
MMAsia | 2 |
| 2024 | DialBB: A Dialogue System Development Framework as an Educational MaterialabstractWe demonstrate DialBB, a dialogue system development framework, which we have been building as an educational material for dialogue system technology.Building a dialogue system requires the adoption of an appropriate architecture depending on the application and the integration of various technologies.However, this is not easy for those who have just started learning dialogue system technology.Therefore, there is a demand for educational materials that integrate various technologies to build dialogue systems, because traditional dialogue system development frameworks were not designed for educational purposes.DialBB enables the development of dialogue systems by combining modules called building blocks.After understanding sample applications, learners can easily build simple systems using built-in blocks and can build advanced systems using their own developed blocks. Mikio Nakano, Kazunori Komatani |
SIGDIAL | 2 |
| 2024 | Travel Agency Task Dialogue Corpus: A Multimodal Dataset with Age-Diverse SpeakersabstractWhen individuals communicate, they use different vocabularies, speaking speeds, facial expressions, and gestural languages, depending on those with whom they are speaking. This study focuses on the age of the speaker as a factor that affects the style of communication. We collected a multimodal dialogue corpus with various speaker ages. We used travel as the topic, as it interests people of all ages, and we set up a task based on a tourism consultation between an operator and a customer at a travel agency. This article presents the details of the dialogue task, collection procedures and annotations, and analysis of the characteristics of the dialogues and facial expressions, focusing on the age of the speakers. The results of the analysis suggest that the adult speakers have more independent opinions, the older speakers express their opinions more frequently than other age groups, and those in the operator role smile more frequently at minors. Michimasa Inaba, Yuya Chiba, Zhiyang Qi, Ryuichiro Higashinaka, Kazunori Komatani, Yusuke Miyao, Takayuki Nagai |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 5 |
| 2023 | Recursive Sound Source Separation with Deep Learning-based Beamforming for Unknown Number of Sources
Hokuto Munakata, Ryu Takeda, Kazunori Komatani |
INTERSPEECH | 3 |
| 2023 | Analyzing Differences in Subjective Annotations by Participants and Third-party Annotators in Multimodal Dialogue CorpusabstractEstimating the subjective impressions of human users during a dialogue is necessary when constructing a dialogue system that can respond adaptively to their emotional states.However, such subjective impressions (e.g., how much the user enjoys the dialogue) are inherently ambiguous, and the annotation results provided by multiple annotators do not always agree because they depend on the subjectivity of the annotators.In this paper, we analyzed the annotation results using 13,226 exchanges from 155 participants in a multimodal dialogue corpus called Hazumi that we had constructed, where each exchange was annotated by five third-party annotators.We investigated the agreement between the subjective annotations given by the third-party annotators and the participants themselves, on both perexchange annotations (i.e., participant's sentiments) and per-dialogue (-participant) annotations (i.e., questionnaires on rapport and personality traits).We also investigated the conditions under which the annotation results are reliable.Our findings demonstrate that the dispersion of third-party sentiment annotations correlates with agreeableness of the participants, one of the Big Five personality traits. Kazunori Komatani, Ryu Takeda, Shogo Okada |
SIGDIAL | 1 |
| 2023 | Joint Separation and Localization of Moving Sound Sources Based on Neural Full-Rank Spatial Covariance AnalysisabstractThis paper presents an unsupervised multichannel method that can separate moving sound sources based on an amortized variational inference (AVI) of joint separation and localization. A recently proposed blind source separation (BSS) method called neural full-rank spatial covariance analysis (FCA) trains a neural separation model based on a nonlinear generative model of multichannel mixtures and can precisely separate unseen mixture signals. This method, however, assumes that the sound sources hardly move, and thus its performance is easily degraded by the source movements. In this paper, we solve this problem by introducing time-varying spatial covariance matrices and directions of arrival of sources into the nonlinear generative model of the neural FCA. This generative model is used for training a neural network to jointly separate and localize moving sources by using only multichannel mixture signals and array geometries. The training objective is derived as a lower bound on the log-marginal posterior probability in the framework of AVI. Experimental results obtained with mixture signals of moving sources show that our method outperformed an existing joint separation and localization method and standard BSS methods. Hokuto Munakata, Yoshiaki Bando, Ryu Takeda, Kazunori Komatani, Masaki Onishi |
IEEE Signal Process. Lett. | 4 |
| 2023 | Effects of Physiological Signals in Different Types of Multimodal Sentiment EstimationabstractMultimodal sentiment analysis has become a focus of research in recent years. However, most studies of multimodal sentiment analysis have considered only signals that are observable by humans, such as linguistic, audio and visual information, whereas the contribution of the multimodal fusion of such signals with unobservable signals, i.e., physiological signals, has not been comprehensively explored. In this study, we aim to investigate effects of physiological signals in multimodal sentiment analysis by evaluating all of the fusion models for different types of sentiment estimation in naturalistic human-agent interaction settings. Our results suggest that physiological features are effective in the unimodal model and that the fusion of linguistic representations with physiological features provides the best results for estimating self-sentiment labels as annotated by the users themselves. In contrast, the tensor fusion of linguistic representations with audiovisual features is effective for estimating sentiment labels as annotated by a third party in regression tasks, which can be derived from the corresponding signals that are observable by humans. A detailed analysis of the self-sentiment estimation results suggests that different modalities play different roles in sentiment estimation, and corresponding implications are discussed. Shun Katada, Shogo Okada, Kazunori Komatani |
IEEE Trans. Affect. Comput. | 3 |
| 2022 | Transformer-Based Physiological Feature Learning for Multimodal Analysis of Self-Reported SentimentabstractOne of the main challenges in realizing dialog systems is adapting to a user’s sentiment state in real time. Large-scale language models, such as BERT, have achieved excellent performance in sentiment estimation; however, the use of only linguistic information from user utterances in sentiment estimation still has limitations. In fact, self-reported sentiment is not necessarily expressed by user utterances. To mitigate the issue that the true sentiment state is not expressed as observable signals, psychophysiology and affective computing studies have focused on physiological signals that capture involuntary changes related to emotions. We address this problem by efficiently introducing time-series physiological signals into a state-of-the-art language model to develop an adaptive dialog system. Compared with linguistic models based on BERT representations, physiological long short-term memory (LSTM) models based on our proposed physiological signal processing method have competitive performance. Moreover, we extend our physiological signal processing method to the Transformer language model and propose the Time-series Physiological Transformer (TPTr), which captures sentiment changes based on both linguistic and physiological information. In ensemble models, our proposed methods significantly outperform the previous best result (p < 0.05). Shun Katada, Shogo Okada, Kazunori Komatani |
ICMI | 3 |
| 2022 | Training Data Generation with DOA-based Selecting and Remixing for Unsupervised Training of Deep Separation Models
Hokuto Munakata, Ryu Takeda, Kazunori Komatani |
INTERSPEECH | 3 |
| 2022 | Empirical Sampling from Latent Utterance-wise Evidence Model for Missing Data ASR based on Neural Encoder-Decoder Model
Ryu Takeda, Yui Sudo, Kazuhiro Nakadai, Kazunori Komatani |
INTERSPEECH | 4 |
| 2022 | Collection and Analysis of Travel Agency Task Dialogues with Age-Diverse SpeakersabstractWhen individuals communicate with each other, they use different vocabulary, speaking speed, facial expressions, and body language depending on the people they talk to. This paper focuses on the speaker’s age as a factor that affects the change in communication. We collected a multimodal dialogue corpus with a wide range of speaker ages. As a dialogue task, we focus on travel, which interests people of all ages, and we set up a task based on a tourism consultation between an operator and a customer at a travel agency. This paper provides details of the dialogue task, the collection procedure and annotations, and the analysis on the characteristics of the dialogues and facial expressions focusing on the age of the speakers. Results of the analysis suggest that the adult speakers have more independent opinions, the older speakers more frequently express their opinions frequently compared with other age groups, and the operators expressed a smile more frequently to the minor speakers. Michimasa Inaba, Yuya Chiba, Ryuichiro Higashinaka, Kazunori Komatani, Yusuke Miyao, Takayuki Nagai |
LREC | 4 |
| 2022 | Knowledge Graph Augmentation with Entity Identification for Improving Knowledge Graph Completion Performance
Shuichi Chikatsuji, Kenta Yamamoto, Ryu Takeda, Kazunori Komatani |
PRICAI (1) | 4 |
| 2021 | Multimodal Human-Agent Dialogue Corpus with Annotations at Utterance and Dialogue LevelsabstractThe behaviors of general users for a dialogue system differ greatly from those for a human interlocutor. We have been collecting a multimodal dialogue corpus between a human participant and a virtual agent operated by the Wizard-of-Oz method. This paper presents the collected corpus, Hazumi, which was released in August 2020 and March 2021. The corpus consists of three versions: Hazumi1712, Hazumi1902, and Hazumi1911. The version numbers correspond to the periods during which the data were collected. The three versions contain the dialogue data of 29, 30, and 30 participants, respectively, each of whom spoke with the agent for about 15 to 20 minutes. The corpus contains multimodal recordings, along with several annotations given to every exchange, feature files extracted from the recorded data, and the results of questionnaires conducted before and after the dialogues. The third version Hazumi1911 also contains the physiological signals of the participants during the dialogues and additional questionnaire items. We also show several analyses conducted using this corpus. We anticipate that the corpus will be useful for developing user-adaptive multimodal dialogue systems. Kazunori Komatani, Shogo Okada |
ACII | 1 |
| 2021 | Recognizing Social Signals with Weakly Supervised Multitask Learning for Multimodal Dialogue SystemsabstractSocial signal processing is a methodology that is used to infer human inner states, including attitudes, sentiments and impressions, from verbal and nonverbal multimodal information. The difficulty in training a social signal recognition model is that the ground-truth (target) labels given by multiple coders often disagree because the annotation of social signals such as sentiments is a subjective and ambiguous task. We introduce weakly supervised learning (WSL) algorithms to such an inaccurate supervision setting in which the target label is not necessarily accurate. The novel challenge in this paper is to explore an effective WSL strategy for recognizing social signals. The strategy is verified through two multimodal datasets including audio, visual, and linguistic data collected in a human-agent dialogue setting. First, we clarify that the proposed WSL strategy for deep neural networks (DNNs), called tri-teaching works well in almost all classification tasks. Second, we demonstrate the effectiveness of integrating WSL and multitask learning (MTL), which exploits several label types in the datasets. Third, we show that our proposed approach achieves less accuracy degradation than an existing training algorithm for a DNN (curriculum learning) in a cross-corpus setting, with a maximum improvement of 7.2%. Yuki Hirano, Shogo Okada, Kazunori Komatani |
ICMI | 3 |
| 2021 | Multimodal User Satisfaction Recognition for Non-task Oriented Dialogue SystemsabstractMultimodal dialogue systems (MDSs) are needed to allow users to converse with virtual agents that use natural language by sensing the multimodal behavior of users. One crucial step in the development of an MDS is measuring how well the dialogue system performs. Though previous research focused on the user satisfaction modeling from linguistic modality in text-to-text dialogue systems, the user satisfaction is observed by not only spoken dialogue contents but also the acoustic and visual nonverbal behaviors of users. Multimodal social signal sensing provides a solution that automatically measures dialogue systems based on subjective evaluation. With this background, we proposed a multimodal recognition model of the user using sequence modeling algorithms (RNN, LSTM, and GRU). It is a novel challenge to recognize the user satisfaction label at the dialogue level. Each label was annotated by the user based on the overall dialogue. We extracted both verbal features and nonverbal features at the exchange level (the unit is a pair of system and user utterances) and analyzed the contributions of multimodal features and unimodal features to recognize user satisfaction labels at the dialogue level. We used a multimodal user-system dialogue data corpus with user satisfaction labels at the dialogue level. To validate the recognition accuracy of the proposed multimodal modeling approach, we compared the proposed method with two models based on human perception by external human coders and the system operator (called “Wizard”) with whom the user talks. The experimental results showed that the multimodal model achieved a better performance in both classification and regression tasks. The results indicated that the performance of the multimodal model was higher than that of the human models. Wenqing Wei, Sixia Li, Shogo Okada, Kazunori Komatani |
ICMI | 4 |
| 2021 | Age Estimation with Speech-Age Model for Heterogeneous Speech Datasets
Ryu Takeda, Kazunori Komatani |
Interspeech | 2 |
| 2021 | Knowledge Graph Completion-based Question Selection for Acquiring Domain Knowledge through DialoguesabstractBuilding a perfect knowledge base in a certain domain is practically impossible, so it is effective for dialogue systems to acquire knowledge for enhancing an imperfect knowledge base through natural language dialogues with users. This paper proposes a framework for selecting questions for such knowledge acquisition when a knowledge graph is used as the knowledge base. The framework uses knowledge graph completion (KGC) for predicting new links that are likely to be correct and selects questions on the basis of the KGC scores. One of the problems with this framework is that questions with incorrect content might be selected, which often occurs when the link prediction performance is low, and this would reduce the users’ willingness to engage in dialogues. To alleviate this problem, this paper presents two modifications to the KGC training: 1) creating pseudo entities having substrings of the names of the entities in the graph so that the entities whose names share substrings are connected and 2) limiting the range of negative sampling. Cross validation-based experiments we conducted showed that these modifications improved KGC performance. We also conducted a user study with crowdsourcing to investigate the subjective perception of the correctness of the predicted links. The results suggest that the model trained with the modifications is capable of avoiding questions with incorrect content. Kazunori Komatani, Yuma Fujioka, Keisuke Nakashima, Katsuhiko Hayashi 0001, Mikio Nakano |
IUI | 1 |
| 2020 | Is She Truly Enjoying the Conversation?: Analysis of Physiological Signals toward Adaptive Dialogue SystemsabstractIn human-agent interactions, it is necessary for the systems to identify the current emotional state of the user to adapt their dialogue strategies. Nevertheless, this task is challenging because the current emotional states are not always expressed in a natural setting and change dynamically. Recent accumulated evidence has indicated the usefulness of physiological modalities to realize emotion recognition. However, the contribution of the time series physiological signals in human-agent interaction during a dialogue has not been extensively investigated. This paper presents a machine learning model based on physiological signals to estimate a user's sentiment at every exchange during a dialogue. Using a wearable sensing device, the time series physiological data including the electrodermal activity (EDA) and heart rate in addition to acoustic and visual information during a dialogue were collected. The sentiment labels were annotated by the participants themselves and by external human coders for each exchange consisting of a pair of system and participant utterances. The experimental results showed that a multimodal deep neural network (DNN) model combined with the EDA and visual features achieved an accuracy of 63.2%. In general, this task is challenging, as indicated by the accuracy of 63.0% attained by the external coders. The analysis of the sentiment estimation results for each individual indicated that the human coders often wrongly estimated the negative sentiment labels, and in this case, the performance of the DNN model was higher than that of the human coders. These results indicate that physiological signals can help in detecting the implicit aspects of negative sentiments, which are acoustically/visually indistinguishable. Shun Katada, Shogo Okada, Yuki Hirano, Kazunori Komatani |
ICMI | 4 |
| 2020 | Frame-Wise Online Unsupervised Adaptation of DNN-HMM Acoustic Model from Perspective of Robust Adaptive Filtering
Ryu Takeda, Kazunori Komatani |
INTERSPEECH | 2 |
| 2020 | User Impressions of Questions to Acquire Lexical KnowledgeabstractFor the acquisition of knowledge through dialogues, it is crucial for systems to ask questions that do not diminish the user's willingness to talk, i.e., that do not degrade the user's impression.This paper reports the results of our analysis on how user impression changes depending on the types of questions to acquire lexical knowledge, that is, explicit and implicit questions, and the correctness of the content of the questions.We also analyzed how sequences of the same type of questions affect user impression.User impression scores were collected from 104 participants recruited via crowdsourcing and then regression analysis was conducted.The results demonstrate that implicit questions give a good impression when their content is correct, but a bad impression otherwise.We also found that consecutive explicit questions are more annoying than implicit ones when the content of the questions is correct.Our findings reveal helpful insights for creating a strategy to avoid user impression deterioration during knowledge acquisition. Kazunori Komatani, Mikio Nakano |
SIGdial | 1 |
| 2020 | A framework for building closed-domain chat dialogue systemsabstractThis paper presents HRIChat, a framework for developing closed-domain chat dialogue systems. Being able to engage in chat dialogues has been found effective for improving communication between humans and dialogue systems. This paper focuses on closed-domain systems because they would be useful when combined with task-oriented dialogue systems in the same domain. HRIChat enables domain-dependent language understanding so that it can deal well with domain-specific utterances. In addition, HRIChat makes it possible to integrate state transition network-based dialogue management and reaction-based dialogue management. FoodChatbot, which is an application in the food and restaurant domain, has been developed and evaluated through a user study. Its results suggest that reasonably good systems can be developed with HRIChat. This paper also reports lessons learned from the development and evaluation of FoodChatbot. Mikio Nakano, Kazunori Komatani |
Knowl. Based Syst. | 2 |
| 2019 | Binarized Knowledge Graph Embeddings
Koki Kishimoto, Katsuhiko Hayashi 0001, Genki Akai, Masashi Shimbo, Kazunori Komatani |
ECIR (1) | 5 |
| 2019 | Multitask Prediction of Exchange-level Annotations for Multimodal Dialogue SystemsabstractThis paper presents multimodal computational modeling of three labels that are independently annotated per exchange to implement an adaptation mechanism of dialogue strategy in spoken dialogue systems based on recognizing user sentiment by multimodal signal processing. The three labels include (1) user’s interest label pertaining to the current topic, (2) user’s sentiment label, and (3) topic continuance denoting whether the system should continue the current topic or change it. Predicting the three types of labels that capture different aspects of the user’s sentiment level and the system’s next action contribute to adopting a dialogue strategy based on the user’s sentiment. For this purpose, we enhanced shared multimodal dialogue data by annotating impressed sentiment labels and the topic continuance labels. With the corpus, we develop a multimodal prediction model for the three labels. A multitask learning technique is applied for binary classification tasks of the three labels considering the partial similarities among them. The prediction model was efficiently trained even with a small data set (less than 2000 samples) thanks to the multitask learning framework. Experimental results show that the multitask deep neural network (DNN) model trained with multimodal features including linguistics, facial expressions, body and head motions, and acoustic features, outperformed those trained as single-task DNNs by 1.6 points at the maximum. Yuki Hirano, Shogo Okada, Haruto Nishimoto, Kazunori Komatani |
ICMI | 4 |
| 2019 | Clarifying Privacy, Property, and Power: Case Study on Value Conflict Between CommunitiesabstractThis study analyzes the value conflict of a paper on fan fiction writing that used online fan fiction novels as a source to extract and filter sexual expressions from text. The boundaries of public and private information are ambiguous because users are not always aware of or have agreed to the fact that their content is to be used openly. The case was complicated by the fact that the use of these data by researchers violated an unconsciously infringed upon right of a vulnerable community with a weak legal position. This paper describes the debate on this topic among researchers from engineering and humanities fields on whether the purpose of the research was ethically acceptable; how the systems can be embedded in ethical values; and what ethical, legal, social, and educational lessons are appropriate for governance of artificial intelligence (AI). Our analysis aimed not only to clarify the abstract concept of privacy but also to make changes to the submission guidelines for authors. We hope that our analysis contributes to the governance of ethical AIs and AI ethics on handling sensitive aspects of online activities. Arisa Ema, Hirotaka Osawa, Reina Saijo, Akinori Kubo, Takushi Otani, Hiromitsu Hattori, Naonori Akiya, Nobutsugu Kanzaki, Minao Kukita, Kazunori Komatani, Ryutaro Ichise |
Proc. IEEE | 10 |
| 2018 | Unsupervised Adaptation of Neural Networks for Discriminative Sound Source Localization with Eliminative ConstraintabstractThis paper describes an unsupervised adaptation method of deep neural networks (DNNs) regarding discriminative sound source localization (SSL). DNNs-based SSL and its unsupervised adaptation fail under different conditions from those during training. The estimations sometimes include incoherent unpredictable errors due to the NN's non-linearity. We propose an eliminative posterior probability constraint using a model-based SSL for unsupervised DNNs adaptation. This constraint forces the probability of “less possible candidates” to become zero to eliminate incoherent errors. The candidates are indicated by a model-based SSL method because it can estimate the azimuth of the sound source with moderate accuracy and explicit reasoning. As a result, the localization performance of adapted DNNs improved more than that of model-based SSL. Experimental results showed that our method improved localization correctness of 1D azimuth and 3D regions by a maximum of 13.3 and 5.9 points compared with the model-based SSL. Ryu Takeda, Yoshiki Kudo, Kazuki Takashima, Yoshifumi Kitamura, Kazunori Komatani |
ICASSP | 5 |
| 2018 | Investigating Effectiveness of Linguistic Features Based on Speech Recognition for Storytelling Skill Assessment
Shogo Okada, Kazunori Komatani |
IEA/AIE | 2 |
| 2018 | Multi-timescale Feature-extraction Architecture of Deep Neural Networks for Acoustic Model Training from Raw Speech SignalabstractThis paper describes a new architecture of deep neural networks (DNNs) for acoustic models. Training DNNs from raw speech signals will provide 1) novel features of signals, 2) normalization-free processing such as utterance-wise mean subtraction, and 3) low-latency speech recognition for robot audition. Exploiting the longer context of raw speech signals seems useful in improving recognition accuracy. However, naive use of longer contexts results in the loss of short-term patterns; thus, recognition accuracy degrades. We propose a multi-timescale feature-extraction architecture of DNNs with blocks of different time scales, which enable capturing long- and short-term patterns of speech signals. Each block consists of complex-valued networks that correspond to Fourier and filterbank transformations for analysis. Experiments showed that the proposed multi-timescale architecture reduced the word error rate by about 3% compared with those only with the longterm context. Analysis of the extracted features revealed that our architecture efficiently captured the slow and fast changes of speech features. Ryu Takeda, Kazuhiro Nakadai, Kazunori Komatani |
IROS | 3 |
| 2018 | Collection of Multimodal Dialog Data and Analysis of the Result of Annotation of Users' Interest Level
Masahiro Araki, Sayaka Tomimasu, Mikio Nakano, Kazunori Komatani, Shogo Okada, Shinya Fujie, Hiroaki Sugiyama |
LREC | 4 |
| 2018 | Word Segmentation From Phoneme Sequences Based On Pitman-Yor Semi-Markov Model Exploiting Subword InformationabstractWord segmentation from phoneme sequences is essential to identify unknown words -of-vocabulary; OOV) in spoken dialogues. The Pitman-Yor semi-Markov model (PYSMM) is used for word segmentation that handles dynamic increase in vocabularies. The obtained vocabularies, however, still include meaningless entries due to insufficient cues for phoneme sequences. We focus here on using subword information to capture patterns as “words.” We propose 1) a model based on subword N-gram and subword estimation using a vocabulary set, and 2) posterior fusion of the results of a PYSMM and our model to take advantage of both. Our experiments showed 1) the potential of using subword information for OOV acquisition, and 2) that our method outperformed the PYSMM by 1.53 and 1.07 in terms of the F-measure of the obtained OOV set for English and Japanese corpora, respectively. Ryu Takeda, Kazunori Komatani, Alexander I. Rudnicky |
SLT | 2 |
| 2017 | Unsupervised adaptation of deep neural networks for sound source localization using entropy minimizationabstractThis paper describes an unsupervised method of adapting deep neural networks (DNNs) for sound source localization (SSL). DNNs-based SSL achieves high localization accuracy for sound data that are similar to training data. However, the accuracy deteriorates if a sound source is at an unknown position in unknown reverberant environments. We solve the problem by using unsupervised adaption of the DNNs' parameters to the observed sound signals. Entropy is used as the objective function and minimized to optimize the parameters on the basis of the gradient method. Adaptation without overfitting is achieved by using 1) a parameter adaptation layer, such as linear transform network, and 2) early stopping of the parameter updates. Experimental results indicated that our method improved localization accuracy by a maximum of 20 points for unknown positions and reverberant data. Ryu Takeda, Kazunori Komatani |
ICASSP | 2 |
| 2017 | Unsupervised Segmentation of Phoneme Sequences based on Pitman-Yor Semi-Markov Model using Phoneme Length ContextabstractUnsupervised segmentation of phoneme sequences is an essential process to obtain unknown words during spoken dialogues. In this segmentation, an input phoneme sequence without delimiters is converted into segmented sub-sequences corresponding to words. The Pitman-Yor semi-Markov model (PYSMM) is promising for this problem, but its performance degrades when it is applied to phoneme-level word segmentation. This is because of insufficient cues for the segmentation, e.g., homophones are improperly treated as single entries and their different contexts are also confused. We propose a phoneme-length context model for PYSMM to give a helpful cue at the phoneme-level and to predict succeeding segments more accurately. Our experiments showed that the peak performance with our context model outperformed those without such a context model by 0.045 at most in terms of F-measures of estimated segmentation. Ryu Takeda, Kazunori Komatani |
IJCNLP(1) | 2 |
| 2017 | Node Pruning Based on Entropy of Weights and Node Activity for Small-Footprint Acoustic Model Based on Deep Neural Networks
Ryu Takeda, Kazuhiro Nakadai, Kazunori Komatani |
INTERSPEECH | 3 |
| 2017 | Lexical Acquisition through Implicit Confirmations over Multiple DialoguesabstractWe address the problem of acquiring the ontological categories of unknown terms through implicit confirmation in dialogues.We develop an approach that makes implicit confirmation requests with an unknown term's predicted category.Our approach does not degrade user experience with repetitive explicit confirmations, but the system has difficulty determining if information in the confirmation request can be correctly acquired.To overcome this challenge, we propose a method for determining whether or not the predicted category is correct, which is included in an implicit confirmation request.Our method exploits multiple user responses to implicit confirmation requests containing the same ontological category.Experimental results revealed that the proposed method exhibited a higher precision rate for determining the correctly predicted categories than when only single user responses were considered. Kohei Ono, Ryu Takeda, Eric Nichols, Mikio Nakano, Kazunori Komatani |
SIGDIAL Conference | 5 |
| 2017 | Acoustic model training based on node-wise weight boundary model for fast and small-footprint deep neural networks
Ryu Takeda, Kazuhiro Nakadai, Kazunori Komatani |
Comput. Speech Lang. | 3 |
| 2016 | Bayesian Language Model based on Mixture of Segmental Contexts for Spontaneous Utterances with Unexpected WordsabstractThis paper describes a Bayesian language model for predicting spontaneous utterances. People sometimes say unexpected words, such as fillers or hesitations, that cause the miss-prediction of words in normal N-gram models. Our proposed model considers mixtures of possible segmental contexts, that is, a kind of context-word selection. It can reduce negative effects caused by unexpected words because it represents conditional occurrence probabilities of a word as weighted mixtures of possible segmental contexts. The tuning of mixture weights is the key issue in this approach as the segment patterns becomes numerous, thus we resolve it by using Bayesian model. The generative process is achieved by combining the stick-breaking process and the process used in the variable order Pitman-Yor language model. Experimental evaluations revealed that our model outperformed contiguous N-gram models in terms of perplexity for noisy text including hesitations. Ryu Takeda, Kazunori Komatani |
COLING | 2 |
| 2016 | Sound source localization based on deep neural networks with directional activate function exploiting phase informationabstractThis paper describes sound source localization (SSL) based on deep neural networks (DNNs) using discriminative training. A naïve DNNs for SSL can be configured as follows. Input is the frequency-domain feature used in other SSL methods, and the structure of DNNs is a fully-connected network using real numbers. The training fails because its network structure loses two important properties, i.e., the orthogonality of sub-bands and the intensity- and time-information saved in complex numbers. We solved these two problems by 1) integrating directional information at each sub-band hierarchically, and 2) designing a directional activator that could treat the complex numbers at each sub-band. Our experiments indicated that our method outperformed the naive DNN-based SSL by 20 points in terms of the block-level accuracy. Ryu Takeda, Kazunori Komatani |
ICASSP | 2 |
| 2016 | Discriminative multiple sound source localization based on deep neural networks using independent location modelabstractWe propose a training method for multiple sound source localization (SSL) based on deep neural networks (DNNs). Such networks function as posterior probability estimator of sound location in terms of position labels and achieve high localization correctness. Since the previous DNNs' configuration for SSL handles one-sound-source cases, it should be extended to multiple-sound-source cases to apply it to real environments. However, a naïve design causes 1) an increase in the number of labels and training data patterns and 2) a lack of label consistency across different numbers of sound sources, such as one and two-or-more-sound cases. These two problems were solved using our proposed method, which involves an independent location model for the former and an block-wise consistent labeling with ordering for the latter. Our experiments indicated that the SSL based on DNNs trained by our proposed training method out-performed a conventional SSL method by a maximum of 18 points in terms of block-level correctness. Ryu Takeda, Kazunori Komatani |
SLT | 2 |
| 2015 | Acoustic model training based on node-wise weight boundary model increasing speed of discrete neural networksabstractOur purpose is to realize discrete neural networks (NNs), whose some parameters are discretized, as a low-resource and fast NNs for acoustic models. Two essential problems should be tackled for its realization; 1) the reduction of discretization errors and 2) the implementation method for fast processing. We propose a new parameter training algorithm for 1) and an implementation using look-up table (LUT) on general-purpose CPUs for 2), respectively. The former can set proper boundaries of discretization at each node of NNs, resulting in the reduction of discretization error. The latter can reduce the memory usage of NNs within the cache size of CPU by encoding parameters of NNs. Experiments with 2-bit discrete NNs showed that our algorithm maintained almost the same word accuracy as 8-bit discrete NNs and achieved a 40% increase in speed of the NN's forward calculation. Ryu Takeda, Kazunori Komatani, Kazuhiro Nakadai |
ASRU | 2 |
| 2015 | User Adaptive Restoration for Incorrectly-Segmented Utterances in Spoken Dialogue SystemsabstractIdeally, the users of spoken dialogue systems should be able to speak at their own tempo.The systems thus need to correctly interpret utterances from various users, even when these utterances contain disfluency.In response to this issue, we propose an approach based on a posteriori restoration for incorrectly segmented utterances.A crucial part of this approach is to classify whether restoration is required or not.We improve the accuracy by adapting the classifier to each user.We focus on the dialogue tempo of each user, which can be obtained during dialogues, and determine the correlation between each user's tempo and the appropriate thresholds for the classification.A linear regression function used to convert the tempos into thresholds is also derived.Experimental results showed that the proposed user adaptation for two classifiers, thresholding and decision tree, improved the classification accuracies by 3.0% and 7.4%, respectively, in ten-fold cross validation. Kazunori Komatani, Naoki Hotta, Satoshi Sato, Mikio Nakano |
SIGDIAL Conference | 1 |
| 2015 | Introduction for Speech and language for interactive robots
Heriberto Cuayáhuitl, Kazunori Komatani, Gabriel Skantze |
Comput. Speech Lang. | 2 |
| 2014 | Detecting incorrectly-segmented utterances for posteriori restoration of turn-taking and ASR resultsabstractAppropriate turn-taking is important in spoken dialogue sys-tems as well as generating correct responses. We have devel-oped a method that performs a posteriori restoration of incor-rectly segmented utterances caused by erroneous voice activity detection (VAD), which result in automatic speech recognition (ASR) errors and inappropriate turn-taking. A crucial part of the method is to classify whether the restoration is required or not. We cast it as a binary classification problem detecting originally single utterances from pairs of utterance fragments. Various features are used representing timing, prosody, and ASR result information to improve its accuracy. Furthermore, two kinds of feature selection are performed to obtain effective and domain-independent features. The experimental results showed that the proposed method outperformed a baseline with manually-selected features by 4.8 % and 3.9 % in cross-domain evalua-tions with two domains. More detailed analysis revealed that the dominant and domain-independent features were utterance intervals and results from the Gaussian mixture model (GMM). Index Terms: spoken dialogue system, VAD error, turn taking, a posteriori restoration Naoki Hotta, Kazunori Komatani, Satoshi Sato, Mikio Nakano |
INTERSPEECH | 2 |
| 2013 | Generating More Specific Questions for Acquiring Attributes of Unknown Concepts from Users
Tsugumi Otsuka, Kazunori Komatani, Satoshi Sato, Mikio Nakano |
SIGDIAL Conference | 2 |
| 2012 | Detecting System-directed Utterances using Dialogue-level FeaturesabstractWe have developed a method to determine whether a user utterance is directed at the system or not. A spoken dialogue system should not respond to audio inputs that are not directed at it (i.e., a user’s mutter), and it therefore needs to detect such inputs to avoid unsuitable responses. We classify the two cases by logistic regression based on a feature set including utterance timing, utterance length, and dialogue status. We conducted experiments using 5395 user utterances for both transcription and automatic speech recognition results. Results showed that the classification accuracy improved by 11.0 and 4.1 points, respectively. We also discuss which features are effective in the classification. Index Terms: spoken dialogue system, system-directed utterance, utterance timing Kazunori Komatani, Akira Hirano, Mikio Nakano |
INTERSPEECH | 1 |
| 2012 | Efficient Blind Dereverberation and Echo Cancellation Based on Independent Component Analysis for Actual Acoustic SignalsabstractThis letter presents a new algorithm for blind dereverberation and echo cancellation based on independent component analysis (ICA) for actual acoustic signals. We focus on frequency domain ICA (FD-ICA) because its computational cost and speed of learning convergence are sufficiently reasonable for practical applications such as hands-free speech recognition. In applying conventional FD-ICA as a preprocessing of automatic speech recognition in noisy environments, one of the most critical problems is how to cope with reverberations. To extract a clean signal from the reverberant observation, we model the separation process in the short-time Fourier transform domain and apply the multiple input/output inverse-filtering theorem (MINT) to the FD-ICA separation model. A naive implementation of this method is computationally expensive, because its time complexity is the second order of reverberation time. Therefore, the main issue in dereverberation is to reduce the high computational cost of ICA. In this letter, we reduce the computational complexity to the linear order of the reverberation time by using two techniques: (1) a separation model based on the independence of delayed observed signals with MINT and (2) spatial sphering for preprocessing. Experiments show that the computational cost grows in proportion to the linear order of the reverberation time and that our method improves the word correctness of automatic speech recognition by 10 to 20 points in a RT₂₀= 670 ms reverberant environment. Ryu Takeda, Kazuhiro Nakadai, Toru Takahashi 0001, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno |
Neural Comput. | 4 |
| 2011 | Simultaneous processing of sound source separation and musical instrument identification using Bayesian spectral modelingabstractThis paper presents a method of both separating audio mixtures into sound sources and identifying the musical instruments of the sources. A statistical tone model of the power spectrogram, called an integrated model, is defined and source separation and instrument identification are carried out on the basis of Bayesian inference. Since, the parameter distributions of the integrated model depend on each instrument, the instrument name is identified by selecting the one that has the maximum relative instrument weight. Experimental results showed correct instrument identification enables precise source separation even when many overtones overlap. Katsutoshi Itoyama, Masataka Goto, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno |
ICASSP | 3 |
| 2011 | A Two-Stage Domain Selection Framework for Extensible Multi-Domain Spoken Dialogue Systems
Mikio Nakano, Kazunori Komatani, Kyoko Matsuyama, Kotaro Funakoshi, Hiroshi G. Okuno |
SIGDIAL Conference | 3 |
| 2011 | A multi-expert model for dialogue and behavior control of conversational robots and agents
Mikio Nakano, Yuji Hasegawa, Kotaro Funakoshi, Johane Takeuchi, Toyotaka Torii, Kazuhiro Nakadai, Naoyuki Kanda, Kazunori Komatani, Hiroshi G. Okuno, Hiroshi Tsujino |
Knowl. Based Syst. | 8 |
| 2010 | Design and Implementation of Two-level Synchronization for Interactive Music RobotabstractOur goal is to develop an interactive music robot, i.e., a robot that presents a musical expression together with humans. A music interaction requires two important functions: synchronization with the music and musical expression, such as singing and dancing. Many instrument-performing robots are only capable of the latter function, they may have difficulty in playing live with human performers. The synchronization function is critical for the interaction. We classify synchronization and musical expression into two levels: (1) the rhythm level and (2) the melody level. Two issues in achieving two-layer synchronization and musical expression are: (1) simultaneous estimation of the rhythm structure and the current part of the music and (2) derivation of the estimation confidence to switch behavior between the rhythm level and the melody level. This paper presents a score following algorithm, incremental audio to score alignment, that conforms to the two-level synchronization design using a particle filter. Our method estimates the score position for the melody level and the tempo for the rhythm level. The reliability of the score position estimation is extracted from the probability distribution of the score position. Experiments are carried out using polyphonic jazz songs. The results confirm that our method switches levels in accordance with the difficulty of the score estimation. When the tempo of the music is less than 120 (beats per minute; bpm), the estimated score positions are accurate and reported; when the tempo is over 120 (bpm), the system tends to report only the tempo to suppress the error in the reported score position predictions. Takuma Otsuka, Kazuhiro Nakadai, Toru Takahashi 0001, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno |
AAAI | 4 |
| 2010 | Improvement in listening capability for humanoid robot HRP-2abstractThis paper describes improvement of sound source separation for a simultaneous automatic speech recognition (ASR) system of a humanoid robot. A recognition error in the system is caused by a separation error and interferences of other sources. In separability, an original geometric source separation (GSS) is improved. Our GSS uses a measured robot's head related transfer function (HRTF) to estimate a separation matrix. As an original GSS uses a simulated HRTF calculated based on a distance between microphone and sound source, there is a large mismatch between the simulated and the measured transfer functions. The mismatch causes a severe degradation of recognition performance. Faster convergence speed of separation matrix reduces separation error. Our approach gives a nearer initial separation matrix based on a measured transfer function from an optimal separation matrix than a simulated one. As a result, we expect that our GSS improves the convergence speed. Our GSS is also able to handle an adaptive step-size parameter. These new features are added into open source robot audition software (OSS) called "HARK" which is newly updated as version 1.0.0. The HARK has been installed on a HRP-2 humanoid with an 8-element microphone array. The listening capability of HRP-2 is evaluated by recognizing a target speech signal which is separated from a simultaneous speech signal by three talkers. The word correct rate (WCR) of ASR improves by 5 points under normal acoustic environments and by 10 points under noisy environments. Experimental results show that HARK 1.0.0 improves the robustness against noises. Toru Takahashi 0001, Kazuhiro Nakadai, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno |
ICRA | 3 |
| 2010 | Upper-limit evaluation of robot audition based on ICA-BSS in multi-source, barge-in and highly reverberant conditionsabstractThis paper presents the upper-limit evaluation of robot audition based on ICA-BSS in multi-source, barge-in and highly reverberant conditions. The goal is that the robot can automatically distinguish a target speech from its own speech and other sound sources in a reverberant environment. We focus on the multi-channel semi-blind ICA (MCSB-ICA), which is one of the sound source separation methods with a microphone array, to achieve such an audition system because it can separate sound source signals including reverberations with few assumptions on environments. The evaluation of MCSB-ICA has been limited to robot's speech separation and reverberation separation. In this paper, we evaluate MCSB-ICA extensively by applying it to multi-source separation problems under common reverberant environments. Experimental results prove that MCSB-ICA outperforms conventional ICA by 30 points in automatic speech recognition performance. Ryu Takeda, Kazuhiro Nakadai, Toru Takahashi 0001, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno |
ICRA | 4 |
| 2010 | Violin Fingering Estimation Based on Violin Pedagogical Fingering Model Constrained by Bowed Sequence Estimation from Audio Input
Akira Maezawa, Katsutoshi Itoyama, Toru Takahashi 0001, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno |
IEA/AIE (3) | 4 |
| 2010 | Improving Identification Accuracy by Extending Acceptable Utterances in Spoken Dialogue System Using Barge-in Timing
Kyoko Matsuyama, Kazunori Komatani, Toru Takahashi 0001, Tetsuya Ogata, Hiroshi G. Okuno |
IEA/AIE (2) | 2 |
| 2010 | Music-Ensemble Robot That Is Capable of Playing the Theremin While Listening to the Accompanied Music
Takuma Otsuka, Takeshi Mizumoto, Kazuhiro Nakadai, Toru Takahashi 0001, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno |
IEA/AIE (1) | 5 |
| 2010 | Analyzing user utterances in barge-in-able spoken dialogue system for improving identification accuracyabstractIn our barge-in-able spoken dialogue system, the user’s behaviors such as barge-in timing and utterance expressions vary according to his/her characteristics and situations. The system adapts to the behaviors by modeling them. We analyzed 1584 utterances collected by our systems of quiz and news-listing tasks and showed that ratio of using referential expressions depends on individual users and average lengths of listed items. This tendency was incorporated as a prior probability into our method and improved the identification accuracy of the user’s intended items. Index Terms: barge-in, spoken dialogue systems, utterance timing, user characteristics Kyoko Matsuyama, Kazunori Komatani, Ryu Takeda, Toru Takahashi 0001, Tetsuya Ogata, Hiroshi G. Okuno |
INTERSPEECH | 2 |
| 2010 | Effects of modelling within- and between-frame temporal variations in power spectra on non-verbal sound recognitionabstractResearch on environmental sound recognition has not shown great development in comparison with that on speech and musical signals. One of the reasons is that the sound category of environmental sounds covers a broad range of acoustical natures. We classified them in order to explore suitable recognition techniques for each characteristic. We focus on impulsive sounds and their non-stationary feature within and between analytic frames. We used matching-pursuit as a framework to use wavelet analysis for extracting temporal variation of audio features inside a frame. We also investigated the validity of modeling decaying patterns of sounds using Hidden markov models. Experimental results indicate that sounds with multiple impulsive signals are recognized better by using time-frequency analyzing bases than by frequency domain analysis. Classification of sound classes with a long and clear decaying pattern improves when HMMs with multiple number of hidden states are applied. Index Terms: audio signal classification, non-speech sound recognition, environmental sound recognition, time-frequency analysis, Matching-Pursuit Nobuhide Yamakawa, Tetsuro Kitahara, Toru Takahashi 0001, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno |
INTERSPEECH | 4 |
| 2010 | Exploiting harmonic structures to improve separating simultaneous speech in under-determined conditionsabstractIn real-world situations, a robot may often encounter “under-determined” situation, where there are more sound sources than microphones. This paper presents a speech separation method using a new constraint on the harmonic structure for a simultaneous speech-recognition system in under-determined conditions. The requirements for a speech separation method in a simultaneous speech-recognition system are (1) ability to handle a large number of talkers, and (2) reduction of distortion in acoustic features. Conventional methods use a maximum likelihood estimation in sound source separation, which fulfills requirement (1). Since it is a general approach, the performance is limited when separating speech. This paper presents a two-stage method to improve the separation. The first stage uses maximum likelihood estimation and extracts the harmonic structure, and the second stage exploits the harmonic structure as a new constraint to achieve requirement (2). We carried out an experiment that simulated three simultaneous utterances using impulse responses recorded by two microphones in an anechoic chamber. The experimental results revealed that our method could improve speech recognition correctness by about four points. Yasuharu Hirasawa, Toru Takahashi 0001, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno |
IROS | 3 |
| 2010 | Robot musical accompaniment: integrating audio and visual cues for real-time synchronization with a human flutistabstractMusicians often have the following problem: they have a music score that requires 2 or more players, but they have no one with whom to practice. So far, score-playing music robots exist, but they lack adaptive abilities to synchronize with fellow players' tempo variations. In other words, if the human speeds up their play, the robot should also increase its speed. However, computer accompaniment systems allow exactly this kind of adaptive ability. We present a first step towards giving these accompaniment abilities to a music robot. We introduce a new paradigm of beat tracking using 2 types of sensory input - visual and audio - using our own visual cue recognition system and state-of-the-art acoustic onset detection techniques. Preliminary experiments suggest that by coupling these two modalities, a robot accompanist can start and stop a performance in synchrony with a flutist, and detect tempo changes within half a second. Angelica Lim, Takeshi Mizumoto, Louis-Kenzo Cahier, Takuma Otsuka, Toru Takahashi 0001, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno |
IROS | 6 |
| 2010 | Human-robot ensemble between robot thereminist and human percussionist using coupled oscillator modelabstractThis paper presents a novel synchronizing method for a human-robot ensemble using coupled oscillators. We define an ensemble as a synchronized performance produced through interactions between independent players. To attain better synchronized performance, the robot should predict the human's behavior to reduce the difference between the human's and robot's onset timings. Existing studies in such synchronization only adapts to onset intervals, thus, need a considerable time to synchronize. We use a coupled oscillator model to predict the human's behavior. Experimental results show that our method reduces the average of onset time errors; when we use a metronome, a tempo-varying metronome or a human drummer, errors are reduced by 38%, 10% or 14% on the average, respectively. These results mean that the prediction of human's behaviors is effective for the synchronized performance. Takeshi Mizumoto, Takuma Otsuka, Kazuhiro Nakadai, Toru Takahashi 0001, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno |
IROS | 5 |
| 2010 | Motion generation based on reliable predictability using self-organized object featuresabstractPredictability is an important factor for determining robot motions. This paper presents a model to generate robot motions based on reliable predictability evaluated through a dynamics learning model which self-organizes object features. The model is composed of a dynamics learning module, namely Recurrent Neural Network with Parametric Bias (RNNPB), and a hierarchical neural network as a feature extraction module. The model inputs raw object images and robot motions. Through bi-directional training of the two models, object features which describe the object motion are self-organized in the output of the hierarchical neural network, which is linked to the input of RNNPB. After training, the model searches for the robot motion with high reliable predictability of object motion. Experiments were performed with the robot's pushing motion with a variety of objects to generate sliding, falling over, bouncing, and rolling motions. For objects with single motion possibility, the robot tended to generate motions that induce the object motion. For objects with two motion possibilities, the robot evenly generated motions that induce the two object motions. Shun Nishide, Tetsuya Ogata, Jun Tani, Toru Takahashi 0001, Kazunori Komatani, Hiroshi G. Okuno |
IROS | 5 |
| 2010 | An improvement in automatic speech recognition using soft missing feature masks for robot auditionabstractWe describe integration of preprocessing and automatic speech recognition based on Missing-Feature-Theory (MFT) to recognize a highly interfered speech signal, such as the signal in a narrow angle between a desired and interfered speakers. As a speech signal separated from a mixture of speech signals includes the leakage from other speech signals, recognition performance of the separated speech degrades. An important problem is estimating the leakage in time-frequency components. Once the leakage is estimated, we can generate missing feature masks (MFM) automatically by using our method. A new weighted sigmoid function is introduced for our MFM generation method. An experiment shows that a word correct rate improves from 66 % to 74 % by using our MFM generation method tuned by a search base approach in the parameter space. Toru Takahashi 0001, Kazuhiro Nakadai, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno |
IROS | 3 |
| 2010 | Speedup and performance improvement of ICA-based robot audition by parallel and resampling-based block-wise processingabstractThis paper describes a speedup and performance improvement of multi-channel semi-blind ICA (MCSB-ICA) with parallel and resampling-based block-wise processing. MCSB-ICA is an integrated method of sound source separation that accomplishes blind source separation, blind dereverberation, and echo cancellation. This method enables robots to separate user's speech signals from observed signals including the robot's own speech, other speech and their reverberations without a priori information. The main problem when MCSB-ICA is applied to robot audition is its high computational cost. We tackle this by multi-threading programming, and the two main issues are 1) the design of parallel processing and 2) incremental implementation. These are solved by a) multiple-stack-based parallel implementation, and b) resampling-based overlaps and block-wise separation. The experimental results proved that our method reduced the real-time factor to less than 0.5 with an eight-core CPU, and it improves the performance of automatic speech recognition by 2-10 points compared with the single-stack-based parallel implementation without the resampling technique. Ryu Takeda, Kazuhiro Nakadai, Toru Takahashi 0001, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno |
IROS | 4 |
| 2010 | Online Error Detection of Barge-In Utterances by Using Individual Users' Utterance Histories in Spoken Dialogue System
Kazunori Komatani, Hiroshi G. Okuno |
SIGDIAL Conference | 1 |
| 2010 | Human-robot cooperation in arrangement of objects using confidence measure of neuro-dynamical systemabstractThe objective of our study was to develop dynamic collaboration between a human and a robot. Most conventional studies have created pre-designed rule-based collaboration systems to determine the timing and behavior of robots to participate in tasks. Our aim is to introduce the confidence of the task as a criterion for robots to determine their timing and behavior. In this paper, we report the effectiveness of applying reproduction accuracy as a measure for quantitatively evaluating confidence in an object arrangement task. Our method is comprised of three phases. First, we obtain human-robot interaction data through the Wizard of OZ method. Second, the obtained data are trained using a neuro-dynamical system, namely, the Multiple Time-scales Recurrent Neural Network (MTRNN). Finally, the prediction error in MTRNN is applied as a confidence measure to determine the robot's behavior. The robot participated in the task when its confidence was high, while it just observed when its confidence was low. Training data were acquired using an actual robot platform, Hiro. The method was evaluated using a robot simulator. The results revealed that motion trajectories could be precisely reproduced with a high degree of confidence, demonstrating the effectiveness of the method. Hiromitsu Awano, Tetsuya Ogata, Shun Nishide, Toru Takahashi 0001, Kazunori Komatani, Hiroshi G. Okuno |
SMC | 5 |
| 2010 | Inter-modality mapping in robot with recurrent neural network
Tetsuya Ogata, Shun Nishide, Hideki Kozima, Kazunori Komatani, Hiroshi G. Okuno |
Pattern Recognit. Lett. | 4 |
| 2009 | ICA-based efficient blind dereverberation and echo cancellation method for barge-in-able robot auditionabstractThis paper describes a new method that allows ldquoBarge-Inrdquo in various environments for robot audition. ldquoBarge-inrdquo means that a user begins to speak simultaneously while a robot is speaking. To achieve the function, we must deal with problems on blind dereverberation and echo cancellation at the same time. We adopt Independent Component Analysis (ICA) because it essentially provides a natural framework for these two problems. To deal with reverberation, we apply a Multiple Input/Output INverse-filtering Theorem-based model of observation to the frequency domain ICA. The main problem is its high-computational cost of ICA. We reduce the computational complexity to the linear order of reverberation time by using two techniques: 1) a separation modelbased on observed signal independence, and 2) enforced spatial sphering for preprocessing. The experimental results revealed that our method improved word correctness of reverberant speech by 10-20 points. Ryu Takeda, Kazuhiro Nakadai, Toru Takahashi 0001, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno |
ICASSP | 4 |
| 2009 | Continuous vocal imitation with self-organized vowel spaces in Recurrent Neural NetworkabstractA continuous vocal imitation system was developed using a computational model that explains the process of phoneme acquisition by infants. Human infants perceive speech sounds not as discrete phoneme sequences but as continuous acoustic signals. One of critical problems in phoneme acquisition is the design for segmenting these continuous speech sounds. The key idea to solve this problem is that articulatory mechanisms such as the vocal tract help human beings to perceive speech sound units corresponding to phonemes. To segment acoustic signal with articulatory movement, we apply the segmenting method to our system by Recurrent Neural Network with Parametric Bias (RNNPB). This method determines the multiple segmentation boundaries in a temporal sequence using the prediction error of the RNNPB model, and the PB values obtained by the method can be encoded as kind of phonemes. Our system was implemented by using a physical vocal tract model, called the Maeda model. Experimental results demonstrated that our system can self-organize the same phonemes in different continuous sounds, and can imitate vocal sound involving arbitrary numbers of vowels using the vowel space in the RNNPB. This suggests that our model reflects the process of phoneme acquisition. Hisashi Kanda, Tetsuya Ogata, Toru Takahashi 0001, Kazunori Komatani, Hiroshi G. Okuno |
ICRA | 4 |
| 2009 | Prediction and imitation of other's motions by reusing own forward-inverse model in robotsabstractThis paper proposes a model that enables a robot to predict and imitate the motions of another by reusing its body forward-inverse model. Our model includes three approaches: (i) projection of a self-forward model for predicting phenomena in the external environment (other individuals), (ii) ldquotriadic relationrdquo that is mediation by a physical object between self and others, (iii) introduction of infant imitation by a parent. The recurrent neural network with parametric bias (RNNPB) model is used as the robot's self forward-inverse model. A group of hierarchical neural networks are attached to the RNNPB model as ldquoconversion modulesrdquo. Experiments demonstrated that a robot with our model could imitate a human's motions by translating the viewpoint. It could also discriminate known/unknown motions appropriately, and associate whole motion dynamics from only one motion snap image. Tetsuya Ogata, Ryunosuke Yokoya, Jun Tani, Kazunori Komatani, Hiroshi G. Okuno |
ICRA | 4 |
| 2009 | Adjusting Occurrence Probabilities of Automatically-Generated Abbreviated Words in Spoken Dialogue Systems
Masaki Katsumaru, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno |
IEA/AIE | 2 |
| 2009 | Improving speech understanding accuracy with limited training data using multiple language models and multiple understanding modelsabstractWe aim to improve a speech understanding module with a small amount of training data. A speech understanding module uses a language model (LM) and a language understanding model (LUM). A lot of training data are needed to improve the models. Such data collection is, however, difficult in an actual process of development. We therefore design and develop a new framework that uses multiple LMs and LUMs to improve speech understanding accuracy under various amounts of training data. Even if the amount of available training data is small, each LM and each LUM can deal well with different types of utterances and more utterances are understood by using multiple LM and LUM. As one implementation of the framework, we develop a method for selecting the most appropriate speech understanding result from several candidates. The selection is based on probabilities of correctness calculated by logistic regressions. We evaluate our framework with various amounts of training data. Index Terms: speech understanding, multiple language models and language understanding models, limited training data Masaki Katsumaru, Mikio Nakano, Kazunori Komatani, Kotaro Funakoshi, Tetsuya Ogata, Hiroshi G. Okuno |
INTERSPEECH | 3 |
| 2009 | Enabling a user to specify an item at any time during system enumeration - item identification for barge-in-able conversational dialogue systemsabstractIn conversational dialogue systems, users prefer to speak at any time and to use natural expressions. We have developed an Independent Component Analysis (ICA) based semi-blind source separation method, which allows users to barge-in over system utterances at any time. We created a novel method from timing information derived from barge-in utterances to identify one item that a user indicates during system enumeration. First, we determine the timing distribution of user utterances containing referential expressions and then approximate it using a gamma distribution. Second, we represent both the utterance timing and automatic speech recognition (ASR) results as probabilities of the desired selection from the system’s enumeration. We then integrate these two probabilities to identify the item having the maximum likelihood of selection. Experimental results using 400 utterances indicated that our method outperformed two methods used as a baseline (one of ASR results only and one of utterance timing only) in identification accuracy. Index Terms: spoken dialogue system, conversational interaction, barge-in, utterance timing Kyoko Matsuyama, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno |
INTERSPEECH | 2 |
| 2009 | Phoneme acquisition model based on vowel imitation using Recurrent Neural NetworkabstractA phoneme-acquisition system was developed using a computational model that explains the developmental process of human infants in the early period of acquiring language. There are two important findings in constructing an infant's acquisition of phonemes: (1) an infant's vowel like cooing tends to invoke utterances that are imitated by its caregiver, and (2) maternal imitation effectively reinforces infant vocalization. Therefore, we hypothesized that infants can acquire phonemes to imitate their caregivers' voices by trial and error, i. e., infants use self-vocalization experience to search for imitable and unimitable elements in their caregivers' voices. On the basis of this hypothesis, we constructed a phoneme acquisition process using interaction involving vowel imitation between a human and an infant model. Our infant model had a vocal tract system, called the Maeda model, and an auditory system implemented by using mel-frequency cepstral coefficients (MFCCs) through STRAIGHT analysis. We applied recurrent neural network with parametric bias (RNNPB) to learn the experience of self-vocalization, to recognize the human voice, and to produce the sound imitated by the infant model. To evaluate imitable and unimitable sounds, we used the prediction error of the RNNPB model. The experimental results revealed that as imitation interactions were repeated, the formants of sounds imitated by our system moved closer to those of human voices, and our system could self-organize the same vowels in different continuous sounds. This suggests that our system can reflect the process of phoneme acquisition. Hisashi Kanda, Tetsuya Ogata, Toru Takahashi 0001, Kazunori Komatani, Hiroshi G. Okuno |
IROS | 4 |
| 2009 | Incremental polyphonic audio to score alignment using beat tracking for singer robotsabstractWe aim at developing a singer robot capable of listening to music with its own ¿ears¿ and interacting with a human's musical performance. Such a singer robot requires at least three functions: listening to the music, understanding what position in the music is being performed, and generating a singing voice. In this paper, we focus on the second function, that is, the capability to align an audio signal to its musical score represented symbolically. Issues underlying the score alignment problem are: (1) diversity in the sounds of various musical instruments, (2) difference between the audio signal and the musical score, (3) fluctuation in tempo of the musical performance. Our solutions to these issues are as follows: (1) the design of features based on a chroma vector in the 12-tone model and onset of the sound, (2) defining the rareness for each tone based on the idea that scarcely used tone is salient in the audio signal, and (3) the use of a switching Kalman filter for robust tempo estimation. The experimental result shows that our score alignment method improves the average of cumulative absolute errors in score alignment by 29% using 100 popular music tunes compared to the beat tracking without score alignment. Takuma Otsuka, Toru Takahashi 0001, Hiroshi G. Okuno, Kazunori Komatani, Tetsuya Ogata, Kazumasa Murata, Kazuhiro Nakadai |
IROS | 4 |
| 2009 | Missing-feature-theory-based robust simultaneous speech recognition system with non-clean speech acoustic modelabstractA humanoid robot must recognize a target speech signal while people around the robot chat with them in real-world. To recognize the target speech signal, robot has to separate the target speech signal among other speech signals and recognize the separated speech signal. As separated signal includes distortion, automatic speech recognition (ASR) performance degrades. To avoid the degradation, we trained an acoustic model from non-clean speech signals to adapt acoustic feature of distorted signal and adding white noise to separated speech signal before extracting acoustic feature. The issues are (1) To determine optimal noise level to add the training speech signals, and (2) To determine optimal noise level to add the separated signal. In this paper, we investigate how much noises should be added to clean speech data for training and how speech recognition performance improves for different positions of three talkers with soft masking. Experimental results show that the best performance is obtained by adding white noises of 30 dB. The ASR with the acoustic model outperforms with ASR with the clean acoustic model by 4 points. Toru Takahashi 0001, Kazuhiro Nakadai, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno |
IROS | 3 |
| 2009 | Step-size parameter adaptation of multi-channel semi-blind ICA with piecewise linear model for barge-in-able robot auditionabstractThis paper describes a step-size parameter adaptation technique of multi-channel semi-blind independent component analysis (MCSB-ICA) for a ¿barge-in-able¿ robot audition system. By ¿barge-in¿, we mean that the user can speak simultaneously when the robot is speaking.We focused on MCSB-ICA to achieve such an audition system because it can separate a user's and a robot's speech under reverberant environments. The problem with MCSB-ICA for robot audition is the slow speed of convergence in estimating a separation filter due to its step-size parameters. Many optimization methods cannot be adopted because their computational costs are proportional to the 2nd order of the reverberation time. Our method yields adaptive step-size parameters with MCSB-ICA at low computational costs. It is based on three techniques; (1) recursive expression of the separation process, (2) a piecewise linear model of the step-size of the separation filter, and (3) adaptive step-size parameters with a sub-ICA-filter. Experimental results show that our approach attains faster convergence speed and lower computational costs than those with a fixed step-size parameter. Ryu Takeda, Kazuhiro Nakadai, Toru Takahashi 0001, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno |
IROS | 4 |
| 2009 | Ranking Help Message Candidates Based on Robust Grammar Verification Results and Utterance History in Spoken Dialogue Systems
Kazunori Komatani, Satoshi Ikeda, Yuichiro Fukubayashi, Tetsuya Ogata, Hiroshi G. Okuno |
SIGDIAL Conference | 1 |
| 2009 | A Model of Temporally Changing User Behaviors in a Deployed Spoken Dialogue System
Kazunori Komatani, Tatsuya Kawahara, Hiroshi G. Okuno |
UMAP | 1 |
| 2008 | Two-channel-based voice activity detection for humanoid robots in noisy home environmentsabstractThe purpose of this research is to accurately classify the speech signals originating from the front even in noisy home environments. This ability can help robots to improve speech recognition and to spot keywords. We therefore developed a new voice activity detection (VAD) based on the complex spectrum circle centroid (CSCC) method. It can classify the speech signals that are received at the front of two microphones by comparing the spectral energy of observed signals with that of target signals estimated by CSCC. Also, it can work in real time without training filter coefficients beforehand even in noisy environments (SNR ≫ 0 dB) and can cope with speech noises generated by audio-visual equipments such as televisions and audio devices. Since the CSCC method requires the directions of the noise signals, we also developed a sound source localization system integrated with cross-power spectrum phase (CSP) analysis and an expectation-maximization (EM) algorithm. This system was demonstrated to enable a robot to cope with multiple sound sources using two microphones. Hyun-Don Kim, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno |
ICRA | 2 |
| 2008 | Object dynamics prediction and motion generation based on reliable predictabilityabstractConsistency of object dynamics, which is related to reliable predictability, is an important factor for generating object manipulation motions. This paper proposes a technique to generate autonomous motions based on consistency of object dynamics. The technique resolves two issues: construction of an object dynamics prediction model and evaluation of consistency. The authors utilize Recurrent Neural Network with Parametric Bias to self-organize the dynamics, and link static images to the self-organized dynamics using a hierarchical neural network to deal with the first issue. For evaluation of consistency, the authors have set an evaluation function based on object dynamics relative to robot motor dynamics. Experiments have shown that the method is capable of predicting 90% of unknown object dynamics. Motion generation experiments have proved that the technique is capable of generating autonomous pushing motions that generate consistent rolling motions. Shun Nishide, Tetsuya Ogata, Ryunosuke Yokoya, Jun Tani, Kazunori Komatani, Hiroshi G. Okuno |
ICRA | 5 |
| 2008 | Integrating Topic Estimation and Dialogue History for Domain Selection in Multi-domain Spoken Dialogue Systems
Satoshi Ikeda, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno |
IEA/AIE | 2 |
| 2008 | Rapid Prototyping of Robust Language Understanding Modules for Spoken Dialogue Systems
Yuichiro Fukubayashi, Kazunori Komatani, Mikio Nakano, Kotaro Funakoshi, Hiroshi Tsujino, Tetsuya Ogata, Hiroshi G. Okuno |
IJCNLP | 2 |
| 2008 | Extensibility verification of robust domain selection against out-of-grammar utterances in multi-domain spoken dialogue systemabstractWe developed a robust domain selection method and verified its extensibility. An issue in domain selection is its robustness against out-of-grammar utterances. It is essential to generate correct system responses because such utterances often cause domain selection errors. We therefore integrated the topic estimation results and the dialogue history to construct a robust domain classifier. Another issue is that domain selection should be performed within an extensible framework, because the system is often modified and extended. That is, the classifier should still have high performance without reconstructing it after adding new domains. The extensibility of our method was not experimentally verified yet, because it requires a lot of effort to collect new dialogue data after extending the system. Therefore, we verified extensibility without collecting new data. We constructed the classifier by leaving out some domains in the dialogue data and then evaluated its accuracy as the classifier for the data where the left-out domains were virtually added. Index Terms: multi-domain spoken dialogue system, domain selection, out-of-grammar utterance Satoshi Ikeda, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno |
INTERSPEECH | 2 |
| 2008 | Expanding vocabulary for recognizing user's abbreviations of proper nouns without increasing ASR error rates in spoken dialogue systemsabstractUsers often abbreviate long words when using spoken dialogue systems, which results in automatic speech recognition (ASR) errors. We define abbreviated words as sub-words of the original word, and add them into an ASR dictionary. The first problem is that proper nouns cannot be correctly segmented by general morphological analyzers, although long and compounded words need to be segmented in agglutinative languages such as Japanese. The second is that, as vocabulary increases, adding many abbreviated words degrades the ASR accuracy. We develop two methods, (1) to segment words by using conjunction probabilities between characters, and (2) to manipulate occurrence probabilities of generated abbreviated words on the basis of the phonological similarities between abbreviated and original words. By our method, the ASR accuracy is improved by 24.2 points for utterances containing abbreviated words, and degraded by only a 0.1 point for those containing original words. Index Terms: spoken dialogue systems, abbreviated words, proper nouns, vocabulary expansion Masaki Katsumaru, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno |
INTERSPEECH | 2 |
| 2008 | Predicting ASR errors by exploiting barge-in rate of individual users for spoken dialogue systemsabstractWe exploit the barge-in rate of individual users to predict automatic speech recognition (ASR) errors. A barge-in is a situation in which a user starts speaking during a system prompt, and it can be detected even when ASR results are not reliable. Such features not using ASR results can be a clue for managing a situation in which user utterances cannot be successfully recognized. Since individual users in our system can be identified by their phone numbers, we accumulate how often each user barges in and use this rate as a user profile for determining whether a current “barge-in ” utterance should be accepted or not. We furthermore set a window that reflects the temporal transition of the user’s behavior as they get accustomed to the system. Experimental results show that setting the window improves the prediction accuracy of whether the utterance should be accepted or not. The experiments also clarify the minimum window width for improving accuracy. Index Terms: spoken dialogue system, user modeling, barge-in 1. Kazunori Komatani, Tatsuya Kawahara, Hiroshi G. Okuno |
INTERSPEECH | 1 |
| 2008 | Soft missing-feature mask generation for simultaneous speech recognition system in robots
Toru Takahashi 0001, Shun'ichi Yamamoto, Kazuhiro Nakadai, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno |
INTERSPEECH | 4 |
| 2008 | Segmenting acoustic signal with articulatory movement using Recurrent Neural Network for phoneme acquisitionabstractThis paper proposes a computational model for phoneme acquisition by infants. Human infants perceive speech sounds not as discrete phoneme sequences but as continuous acoustic signals. One of critical problems in phoneme acquisition is the design for segmenting these continuous speech sounds. The key idea to solve this problem is that articulatory mechanisms such as the vocal tract help human beings to perceive speech sound units corresponding to phonemes. That is, the ability to distinguish phonemes is learned by recognizing unstable points in the dynamics of continuous sound with articulatory movement. We have developed a vocal imitation system embodying the relationship between articulatory movements and sounds produced by the movements. To segment acoustic signal with articulatory movement, we apply the segmenting method to our system by recurrent neural network with parametric bias (RNNPB). This method determines the multiple segmentation boundaries in a temporal sequence using the prediction error of the RNNPB model, and the PB values obtained by the method can be encoded as kind of phonemes. Our system was implemented by using a physical vocal tract model, called the Maeda model. Experimental results demonstrated that our system can self-organize the same phonemes in different continuous sounds. This suggests that our model reflects the process of phoneme acquisition. Hisashi Kanda, Tetsuya Ogata, Kazunori Komatani, Hiroshi G. Okuno |
IROS | 3 |
| 2008 | Target speech detection and separation for humanoid robots in sparse dialogue with noisy home environmentsabstractIn normal human communication, people face the speaker when listening and usually pay attention to the speaker’ face. Therefore, in robot audition, the recognition of the front talker is critical for smooth interactions. This paper presents an enhanced speech detection method for a humanoid robot that can separate and recognize speech signals originating from the front even in noisy home environments. The robot audition system consists of a new type of voice activity detection (VAD) based on the complex spectrum circle centroid (CSCC) method and a maximum signal-to-noise (Max-SNR) beamformer. This VAD based on CSCC can classify speech signals that are retrieved at the frontal region of two microphones embedded on the robot. The system works in real-time without needing training filter coefficients given in advance even in a noisy environment (SNR ≫ 0 dB). It can cope with speech noise generated from televisions and audio devices that does not originate from the center. Experiments using a humanoid robot, SIG2, with two microphones showed that our system enhanced extracted target speech signals more than 12 dB (SNR) and the success rate of automatic speech recognition for Japanese words was increased about 17 points. Hyun-Don Kim, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno |
IROS | 3 |
| 2008 | Design and evaluation of two-channel-based sound source localization over entire azimuth range for moving talkersabstractWe propose a way to evaluate various sound localization systems for moving sounds under the same conditions. To construct a database for moving sounds, we developed a moving sound creation tool using the API library developed by the ARINIS Company. We developed a two-channel-based sound source localization system integrated with a cross-power spectrum phase (CSP) analysis and EM algorithm. The CSP of sound signals obtained with only two microphones is used to localize the sound source without having to use prior information such as impulse response data. The EM algorithm helps the system cope with several moving sound sources and reduce localization error. We evaluated our sound localization method using artificial moving sounds and confirmed that it can well localize moving sounds slower than 1.125 rad/sec. Finally, we solve the problem of distinguishing whether sounds are coming from the front or back by rotating a robotpsilas head equipped with only two microphones. Our system was applied to a humanoid robot called SIG2, and we confirmed its ability to localize sounds over the entire azimuth range. Hyun-Don Kim, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno |
IROS | 2 |
| 2008 | A robot listens to music and counts its beats aloud by separating music from counting voiceabstractThis paper presents a beat-counting robot that can count musical beats aloud, i.e., speak ldquoone, two, three, four, one, two, ...rdquo along music, while listening to music by using its own ears. Music-understanding robots that interact with humans should be able not only to recognize music internally, but also to express their own internal states. To develop our beat-counting robot, we have tackled three issues: (1) recognition of hierarchical beat structures, (2) expression of these structures by counting beats, and (3) suppression of counting voice (self-generated sound) in sound mixtures recorded by ears. The main issue is (3) because the interference of counting voice in music causes the decrease of the beat recognition accuracy. So we designed the architecture for music-understanding robot that is capable of dealing with the issue of self-generated sounds. To solve these issues, we took the following approaches: (1) beat structure prediction based on musical knowledge on chords and drums, (2) speed control of counting voice according to music tempo via a vocoder called STRAIGHT, and (3) semi-blind separation of sound mixtures into music and counting voice via an adaptive filter based on ICA (independent component analysis) that uses the waveform of the counting voice as a prior knowledge. Experimental result showed that suppressing robotpsilas own voice improved music recognition capability. Takeshi Mizumoto, Ryu Takeda, Kazuyoshi Yoshii, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno |
IROS | 4 |
| 2008 | Active sensing based dynamical object feature extractionabstractThis paper presents a method to autonomously extract object features that describe their dynamics from active sensing experiences. The model is composed of a dynamics learning module and a feature extraction module. Recurrent Neural Network with Parametric Bias (RNNPB) is utilized for the dynamics learning module, learning and self-organizing the sequences of robot and object motions. A hierarchical neural network is linked to the input of RNNPB as the feature extraction module for extracting object features that describe the object motions. The two modules are simultaneously trained using image and motion sequences acquired from the robotpsilas active sensing with objects. Experiments are performed with the robotpsilas pushing motion with a variety of objects to generate sliding, falling over, bouncing, and rolling motions. The results have shown that the model is capable of extracting features that distinguish the characteristics of object dynamics. Shun Nishide, Tetsuya Ogata, Ryunosuke Yokoya, Jun Tani, Kazunori Komatani, Hiroshi G. Okuno |
IROS | 5 |
| 2008 | Barge-in-able robot audition based on ICA and missing feature theory under semi-blind situationabstractThis paper describes a robot audition system that allows the user to barge-in; that is, the user can speak simultaneously when the robot is speaking. Our ldquobarge-in-ablerdquo system consists of two stages: (1) cancellation of robot speech and (2) recognition of the separated user speech under the ldquosemi-blind situationrdquo. The semi-blind situation is where a robotpsilas speech signal is known but a userpsilas speech signal is not. The first stage is achieved by using an adaptive filter based on time-frequency domain Independent Component Analysis, because that can separate robot speech more robustly against noise than conventional echo cancellers. To improve performance in online processing, we utilized known source normalization and the exponentially weighted stepsize method. The second stage is achieved by automatic speech recognition (ASR) based on the missing feature theory which provides robust recognition by exploiting the reliability of speech features distorted due to noise and/or separation. The semi-blind situation simplifies the estimation of such reliabilities. Experiments demonstrated that our system improved word correctness of ASR by 10.0%. Ryu Takeda, Kazuhiro Nakadai, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno |
IROS | 3 |
| 2008 | Design and Implementation of 3D Auditory Scene Visualizer towards Auditory Awareness with Face TrackingabstractIf machine audition can recognize an auditory scene containing simultaneous and moving talkers, what kinds of awareness will people gain from an auditory scene visualizer? This paper presents the design and implementation of 3D Auditory Scene Visualizer based on the visual information seeking mantra, i.e., ldquooverview first, zoom and filter, then details on demandrdquo. The machine audition system called HARK captures 3D sounds with a microphone array, localizes and separates sounds, and recognizes separated sounds by automatic speech recognition (ASR). The 3D visualizer implemented in Java 3D displays each sound stream as a beam originating from the center of the microphones (overview mode), shows temporal snapshots with/without specifying focusing areas (zoom and filter mode), and shows detailed information about a particular sound stream (details on demand). In the details-ondemand mode, ASR results are displayed in a ldquokaraokerdquo manner, i.e., character-by-character. This three-mode visualization will give the user auditory awareness enhanced by HARK. In addition, a face-tracking system automatically changes the focus of attention by tracking the userpsilas face. The resulting system is portable and can be deployed in any place, so it is expected to give more vivid awareness than expensive high-fidelity auditory scene reproduction systems. Yuji Kubota, Masatoshi Yoshida, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno |
ISM | 3 |
| 2008 | SalienceGraph: Visualizing Salience Dynamics of Written Discourse by Using Reference Probability and PLSA
Shun Shiramatsu, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno |
PRICAI | 2 |
| 2008 | Managing out-of-grammar utterances by topic estimation with domain extensibility in multi-domain spoken dialogue systems
Kazunori Komatani, Satoshi Ikeda, Tetsuya Ogata, Hiroshi G. Okuno |
Speech Commun. | 1 |
| 2008 | An Efficient Hybrid Music Recommender System Using an Incrementally Trainable Probabilistic Generative ModelabstractThis paper presents a hybrid music recommender system that ranks musical pieces while efficiently maintaining collaborative and content-based data, i.e., rating scores given by users and acoustic features of audio signals. This hybrid approach overcomes the conventional tradeoff between recommendation accuracy and variety of recommended artists. Collaborative filtering, which is used on e-commerce sites, cannot recommend nonbrated pieces and provides a narrow variety of artists. Content-based filtering does not have satisfactory accuracy because it is based on the heuristics that the user's favorite pieces will have similar musical content despite there being exceptions. To attain a higher recommendation accuracy along with a wider variety of artists, we use a probabilistic generative model that unifies the collaborative and content-based data in a principled way. This model can explain the generative mechanism of the observed data in the probability theory. The probability distribution over users, pieces, and features is decomposed into three conditionally independent ones by introducing latent variables. This decomposition enables us to efficiently and incrementally adapt the model for increasing numbers of users and rating scores. We evaluated our system by using audio signals of commercial CDs and their corresponding rating scores obtained from an e-commerce site. The results revealed that our system accurately recommended pieces including nonrated ones from a wide variety of artists and maintained a high degree of accuracy even when new users and rating scores were added. Kazuyoshi Yoshii, Masataka Goto, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno |
IEEE Trans. Speech Audio Process. | 3 |
| 2007 | Design and implementation of a robot audition system for automatic speech recognition of simultaneous speechabstractThis paper addresses robot audition that can cope with speech that has a low signal-to-noise ratio (SNR) in real time by using robot-embedded microphones. To cope with such a noise, we exploited two key ideas; Preprocessing consisting of sound source localization and separation with a microphone array, and system integration based on missing feature theory (MFT). Preprocessing improves the SNR of a target sound signal using geometric source separation with multichannel post-filter. MFT uses only reliable acoustic features in speech recognition and masks unreliable parts caused by errors in preprocessing. MFT thus provides smooth integration between preprocessing and automatic speech recognition. A real-time robot audition system based on these two key ideas is constructed for Honda ASIMO and Humanoid SIG2 with 8-ch microphone arrays. The paper also reports the improvement of ASR performance by using two and three simultaneous speech signals. Shun'ichi Yamamoto, Kazuhiro Nakadai, Mikio Nakano, Hiroshi Tsujino, Jean-Marc Valin, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno |
ASRU | 6 |
| 2007 | Integration and Adaptation of Harmonic and Inharmonic Models for Separating Polyphonic Musical SignalsabstractThis paper describes a sound source separation method for polyphonic sound mixtures of music to build an instrument equalizer for remixing multiple tracks separated from compact-disc recordings by changing the volume level of each track. Although such mixtures usually include both harmonic and inharmonic sounds, the difficulties in dealing with both types of sounds together have not been addressed in most previous methods that have focused on either of the two types separately. We therefore developed an integrated weighted-mixture model consisting of both harmonic-structure and inharmonic-structure tone models (generative models for the power spectrogram). On the basis of the MAP estimation using the EM algorithm, we estimated all model parameters of this integrated model under several original constraints for preventing over-training and maintaining intra-instrument consistency. Using standard MIDI files as prior information of the model parameters, we applied this model to compact-disc recordings and achieved the instrument equalizer. Katsutoshi Itoyama, Masataka Goto, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno |
ICASSP (1) | 3 |
| 2007 | Vowel Imitation Using Vocal Tract Model and Recurrent Neural Network
Hisashi Kanda, Tetsuya Ogata, Kazunori Komatani, Hiroshi G. Okuno |
ICONIP (2) | 3 |
| 2007 | Predicting Object Dynamics from Visual Images through Active Sensing ExperiencesabstractPrediction of dynamic features is an important task for determining the manipulation strategies of an object. This paper presents a technique for predicting dynamics of objects relative to the robot's motion from visual images. During the learning phase, the authors use recurrent neural network with parametric bias (RNNPB) to self-organize the dynamics of objects manipulated by the robot into the PB space. The acquired PB values, static images of objects, and robot motor values are input into a hierarchical neural network to link the static images to dynamic features (PB values). The neural network extracts prominent features that induce each object dynamics. For prediction of the motion sequence of an unknown object, the static image of the object and robot motor value are input into the neural network to calculate the PB values. By inputting the PB values into the closed loop RNNPB, the predicted movements of the object relative to the robot motion are calculated sequentially. Experiments were conducted with the humanoid robot Robovie-IIs pushing objects at different heights. Reducted grayscale images and shoulder pitch angles were input into the neural network to predict the dynamics of target objects. The results of the experiment proved that the technique is efficient for predicting the dynamics of the objects. Shun Nishide, Tetsuya Ogata, Jun Tani, Kazunori Komatani, Hiroshi G. Okuno |
ICRA | 4 |
| 2007 | Distance Estimation of Hidden Objects Based on Acoustical Holography by applying Acoustic Diffraction of Audible SoundabstractOcclusion is a problem for range finders; ranging systems using cameras or lasers cannot be used to estimate distance to an object (hidden object) that is occluded by another (obstacle). We developed a method to estimate the distance to the hidden object by applying acoustic diffraction of audible sound. Our method is based on time-of-flight (TOF), which has been used in ultrasound ranging systems. We determined the best frequency of audible sound and designed its optimal modulated signal for our system. We determined that the system estimates the distance to the hidden object as well as the obstacle. However, the measurement signal obtained from the hidden object was weak. Thus, interference from sound signals reflected from other objects or walls was not negligible. Therefore, we combined acoustical holography (AH) and TOF, which enabled a partial analysis of the reflection sound intensity field around the obstacle and hidden object. Our method was effective for ranging two objects of the same size within a 1.2 m depth range. The accuracy of our method was 3 cm for the obstacle, and 6 cm for the hidden object. Haruhiko Niwa, Tetsuya Ogata, Kazunori Komatani, Hiroshi G. Okuno |
ICRA | 3 |
| 2007 | Human-Robot Cooperation using Quasi-symbols Generated by RNNPB ModelabstractWe describe a means of human robot interaction based not on natural language but on "quasi symbols," which represent sensory-motor dynamics in the task and/or environment. It thus overcomes a key problem of using natural language for human-robot interaction - the need to understand the dynamic context. The quasi-symbols used are motion primitives corresponding to the attractor dynamics of the sensory-motor flow. These primitives are extracted from the observed data using the recurrent neural network with parametric bias (RNNPB) model. Binary representations based on the model parameters were implemented as quasi symbols in a humanoid robot, Robovie. The experiment task was robot-arm operation on a table. The quasi-symbols acquired by learning enabled the robot to perform novel motions. A person was able to control the arm through speech interaction using these quasi-symbols. These quasi symbols formed a hierarchical structure corresponding to the number of nodes in the model. The meaning of some of the quasi-symbols depended on the context, indicating that they are useful for human-robot interaction. Tetsuya Ogata, Shohei Matsumoto, Jun Tani, Kazunori Komatani, Hiroshi G. Okuno |
ICRA | 4 |
| 2007 | Real-Time Auditory and Visual Talker Tracking Through Integrating EM Algorithm and Particle Filter
Hyun-Don Kim, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno |
IEA/AIE | 2 |
| 2007 | Evaluation of Two Simultaneous Continuous Speech Recognition with ICA BSS and MFT-Based ASR
Ryu Takeda, Shun'ichi Yamamoto, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno |
IEA/AIE | 3 |
| 2007 | Topic estimation with domain extensibility for guiding user's out-of-grammar utterances in multi-domain spoken dialogue systemsabstractIn a multi-domain spoken dialogue system, a user’s utterances are more prone to be out-of-grammar, because this kind of system deals with more tasks than a single-domain system. We defined a topic as a domain about which users want to find more information, and we developed a method of recovering out-ofgrammar utterances based on topic estimation, i.e., by providing a help message in the estimated domain. Moreover, the domain extensibility, that is, to facilitate adding new domains, should be inherently retained in multi-domain systems. We therefore collected documents from the Web as training data for topic estimation. Because the data contained not a few noises, we used Latent Semantic Mapping (LSM), which enables robust topic estimation by removing the effect of noise from the data. The experimental results based on using 272 utterances collected with a Woz-like method showed that our method increased the topic estimation accuracy by 23.1 points from the baseline. Index Terms: multi-domain spoken dialogue system, topic estimation, out-of-grammar utterance 1. Satoshi Ikeda, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno |
INTERSPEECH | 2 |
| 2007 | Analyzing temporal transition of real user's behaviors in a spoken dialogue systemabstractManaging various behaviors of real users is indispensable for spoken dialogue systems to operate adequately in real environments. We have analyzed various users ’ behaviors using data collected over 34 months from the Kyoto City Bus Information System. We focused on “barge-in ” and added barge-in rates to our analysis. Temporal transitions of users ’ behaviors, such as automatic speech recognition (ASR) accuracy, task success rates and barge-in rates, were initially investigated. We then examined the relationship between ASR accuracy and barge-in rates. Analysis revealed that the ASR accuracy of utterances inputted with barge-ins was lower because many novices, who were not accustomed to the timing when to utter, used the system. We also observed that the ASR accuracy of utterances with barge-ins differed based on the barge-in rates of individual users. The results indicate that the barge-in rate can be used as a novel user profile for detecting ASR errors. Index Terms: spoken dialogue system, real user behavior, barge-in Kazunori Komatani, Tatsuya Kawahara, Hiroshi G. Okuno |
INTERSPEECH | 1 |
| 2007 | Vocal imitation using physical vocal tract modelabstractA vocal imitation system was developed using a computational model that supports the motor theory of speech perception. A critical problem in vocal imitation is how to generate speech sounds produced by adults, whose vocal tracts have physical properties (i.e., articulatory motions) differing from those of infants’ vocal tracts. To solve this problem, a model based on the motor theory of speech perception, was constructed. This model suggests that infants simulate the speech generation by estimating their own articulatory motions in order to interpret the speech sounds of adults. Applying this model enables the vocal imitation system to estimate articulatory motions for unexperienced speech sounds that have not actually been generated by the system. The system was implemented by using Recurrent Neural Network with Parametric Bias (RNNPB) and a physical vocal tract model, called the Maeda model. Experimental results demonstrated that the system was sufficiently robust with respect to individual differences in speech sounds and could imitate unexperienced vowel sounds. Hisashi Kanda, Tetsuya Ogata, Kazunori Komatani, Hiroshi G. Okuno |
IROS | 3 |
| 2007 | Auditory and visual integration based localization and tracking of humans in daily-life environmentsabstractThe purpose of this research is to develop techniques that enable robots to choose and track a desired person for interaction in daily-life environments. Therefore, localizing multiple moving sounds and human faces is necessary so that robots can locate a desired person. For sound source localization, we used a cross-power spectrum phase analysis (CSP) method and showed that CSP can localize sound sources only using two microphones and does not need impulse response data. An expectation-maximization (EM) algorithm was shown to enable a robot to cope with multiple moving sound sources. For face localization, we developed a method that can reliably detect several faces using the skin color classification obtained by using the EM algorithm. To deal with a change in color state according to illumination condition and various skin colors, the robot can obtain new skin color features of faces detected by OpenCV, an open vision library, for detecting human faces. Finally, we developed a probability based method to integrate auditory and visual information and to produce a reliable tracking path in real time. Furthermore, the developed system chose and tracked people while dealing with various background noises that are considered loud, even in the daily-life environments. Hyun-Don Kim, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno |
IROS | 2 |
| 2007 | Two-way translation of compound sentences and arm motions by recurrent neural networksabstractWe present a connectionist model that combines motions and language based on the behavioral experiences of a real robot. Two models of recurrent neural network with parametric bias (RNNPB) were trained using motion sequences and linguistic sequences. These sequences were combined using their respective parameters so that the robot could handle many-to-many relationships between motion sequences and linguistic sequences. Motion sequences were articulated into some primitives corresponding to given linguistic sequences using the prediction error of the RNNPB model. The experimental task in which a humanoid robot moved its arm on a table demonstrated that the robot could generate a motion sequence corresponding to given linguistic sequence even if the motions or sequences were not included in the training data, and vice versa. Tetsuya Ogata, Masamitsu Murase, Jun Tani, Kazunori Komatani, Hiroshi G. Okuno |
IROS | 4 |
| 2007 | Exploiting known sound source signals to improve ICA-based robot audition in speech separation and recognitionabstractThis paper describes a new semi-blind source separation (semi-BSS) technique with independent component analysis (ICA) for enhancing a target source of interest and for suppressing other known interference sources. The semi BSS technique is necessary for double-talk free robot audition systems in order to utilize known sound source signals such as self speech, music, or TV-sound, through a line-in or ubiquitous network. Unlike the conventional semi-BSS with ICA, we use the time-frequency domain convolution model to describe the reflection of the sound and a new mixing process of sounds for ICA. In other words, we consider that reflected sounds during some delay time are different from the original. ICA then separates the reflections as other interference sources. The model enables us to eliminate the frame size limitations of the frequency-domain ICA, and ICA can separate the known sources under a highly reverberative environment. Experimental results show that our method outperformed the conventional semi-BSS using ICA under simulated normal and highly reverberative environments. Ryu Takeda, Kazuhiro Nakadai, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno |
IROS | 3 |
| 2007 | Discovery of other individuals by projecting a self-model through imitationabstractThis paper proposes a novel model which enables a humanoid robot infant to discover other individual (e.g. human parent). In this work, the authors define “other individual” as an actor which can be predicted by a self-model. For modeling the developmental process of discovering ability, the following three approaches are employed. (i) Projection of a selfmodel for predicting other individual’s actions. (ii) Mediation by a physical object between self and other individual. (iii) Introduction of infant imitation by parent. For creating the self-model of a robot, we apply Recurrent Neural Network with Parametric Bias (RNNPB) model which can learn the robot’s body dynamics. For the other-model of a human, conventional hierarchical neural networks are attached to the RNNPB model as “conversion modules”. Our target task is a moving an object. For evaluation of our model, human discovery experiments by the robot projecting its self-model were conducted. The results demonstrated that our method enabled the robot to predict the human’s motions, and to estimate the human’s position fairly accurately, which proved its adequacy. Ryunosuke Yokoya, Tetsuya Ogata, Jun Tani, Kazunori Komatani, Hiroshi G. Okuno |
IROS | 4 |
| 2007 | A biped robot that keeps steps in time with musical beats while listening to music with its own earsabstractWe aim at enabling a biped robot to interact with humans through real-world music in daily-life environments, e.g., to autonomously keep its steps (stamps) in time with musical beats. To achieve this, the robot should be able to robustly predict the beat times in real time while listening to musical performance with its own ears (head-embedded microphones). However, this has not previously been addressed in most studies on music-synchronized robots due to the difficulty in predicting the beat times in real-world music. To solve this problem, we implemented a beat-tracking method developed in the field of music information processing. The predicted beat times are then used by a feedback-control method that adjusts the robot's step intervals to synchronize its steps in time with the beats. The experimental results show that the robot can adjust its steps in time with the beat times as the tempo changes. The resulting robot needed about 25 [s] to recognize the tempo change after it and then synchronize its steps. Kazuyoshi Yoshii, Kazuhiro Nakadai, Toyotaka Torii, Yuji Hasegawa, Hiroshi Tsujino, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno |
IROS | 6 |
| 2007 | Auditory and Visual Integration based Localization and Tracking of Multiple Moving Sounds in Daily-life EnvironmentsabstractThis paper presents techniques that enable talker tracking for effective human-robot interaction. To track moving people in daily-life environments, localizing multiple moving sounds is necessary so that robots can locate talkers. However, the conventional method requires an array of microphones and impulse response data. Therefore, we propose a way to integrate a cross-power spectrum phase analysis (CSP) method and an expectation-maximization (EM) algorithm. The CSP can localize sound sources using only two microphones and does not need impulse response data. Moreover, the EM algorithm increases the system's effectiveness and allows it to cope with multiple sound sources. We confirmed that the proposed method performs better than the conventional method. In addition, we added a particle filter to the tracking process to produce a reliable tracking path and the particle filter is able to integrate audio-visual information effectively. Furthermore, the applied particle filter is able to track people while dealing with various noises that are even loud sounds in the daily-life environments. Hyun-Don Kim, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno |
RO-MAN | 2 |
| 2006 | F0 Estimation Method for Singing Voice in Polyphonic Audio Signal Based on Statistical Vocal Model and Viterbi SearchabstractThis paper describes a method for estimating F0s of vocal from polyphonic audio signals. Because melody is sung by a singer in many musical pieces, the estimation of F0s of the vocal part is useful for many applications. Based on existing multiple-F0 estimation method, we evaluate the vocal probabilities of the harmonic structure of each F0 candidate. In order to calculate the vocal probabilities of the harmonic structure, we extract and resynthesize the harmonic structure by using a sinusoidal model and extract feature vectors. Then, we evaluate the vocal probability by using vocal and non-vocal Gaussian mixture models (GMMs). Finally, we track F0 trajectories using these probabilities based on Viterbi search. Experimental results show that our method improves estimation accuracy from 78.1% to 84.3%, which is 28.3% reduction of misestimation Hiromasa Fujihara, Tetsuro Kitahara, Masataka Goto, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno |
ICASSP (5) | 4 |
| 2006 | Instrogram: A New Musical Instrument Recognition Technique Without Using Onset Detection NOR F0 EstimationabstractThis paper describes a new technique for recognizing musical instruments in polyphonic music. Because the conventional framework for musical instrument recognition in polyphonic music had to estimate the onset time and fundamental frequency (F0) of each note, instrument recognition strictly suffered from errors of onset detection and F0 estimation. Unlike such a note-based processing framework, our technique calculates the temporal trajectory of instrument existence probabilities for every possible F0, and the results are visualized with a spectrogram-like graphical representation called instrogram. The instrument existence probability is defined as the product of a nonspecific instrument existence probability calculated using PreFEst and a conditional instrument existence probability calculated using the hidden Markov model. Experimental results show that the obtained instrograms reflect the actual instrumentations and facilitate instrument recognition Tetsuro Kitahara, Masataka Goto, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno |
ICASSP (5) | 3 |
| 2006 | An Error Correction Framework Based on Drum Pattern Periodicity for Improving Drum Sound DetectionabstractThis paper presents a framework for correcting errors of automatic drum sound detection focusing on the periodicity of drum patterns. We define drum patterns as periodic structures found in onset sequences of bass and snare drum sounds. Our framework extracts periodic drum patterns from imperfect onset sequences of detected drum sounds (bottom-up processing) and corrects errors using the periodicity of the drum patterns (top-down processing). We implemented this framework on our drum-sound detection system. We first obtained onset sequences of the drum sounds with our system and extracted drum patterns. On the basis of our observation that the same drum patterns tend to be repeated, we detected time points which deviate from the periodicity as error candidates. Finally, we verified each error candidate to judge whether it is an actual onset or not. Experiments of drum sound detection for polyphonic audio signals of popular CD recordings showed that our correction framework improved the average detection accuracy from 77.4% to 80.7% Kazuyoshi Yoshii, Masataka Goto, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno |
ICASSP (5) | 3 |
| 2006 | Genetic Algorithm-Based Improvement of Robot Hearing Capabilities in Separating and Recognizing Simultaneous Speech Signals
Shun'ichi Yamamoto, Kazuhiro Nakadai, Mikio Nakano, Hiroshi Tsujino, Jean-Marc Valin, Ryu Takeda, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno |
IEA/AIE | 7 |
| 2006 | Speaker identification under noisy environments by using harmonic structure extraction and reliable frame weightingabstractWe present methods for automatic speaker identification in noisy environments. To improve noise robustness of speaker identification, we developed two methods, theharmonic structure extraction method and the reliable frame weighting method. The harmonic structure extraction method enables the speaker of input speech signals to be identified after environmental noise has been reduced. This method first extracts harmonic components of the speech from the sound mixtures and then resynthesizes a clean speech signal by using a sinusoidal model driven by harmonic components. The reliable frame weighting method then determines how each frame of the resynthesized speech is reliable (i.e. little influenced by environmental noises) by using two Gaussian mixture models for the speech and noise. The speaker can be robustly identified by attaching importance to reliable frames. Experimental results with thirty speakers showed that our method was able to reduce the influences of environmental noise and achieved an error rate of 10.7%, while the error rate for a conventional method was 18.9%. Index Terms: speaker identification, noise robustness, voice extraction, voice reliability, Gaussian mixture model. Hiromasa Fujihara, Tetsuro Kitahara, Masataka Goto, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno |
INTERSPEECH | 4 |
| 2006 | Dynamic help generation by estimating user²s mental model in spoken dialogue systemsabstractIn a speech interface, a gap between a user’s mental model and actual structures of systems tends to be large because the amount of information conveyed by speech is limited. We address dynamic help generation adapted to users, to decrease the gap between them. We defined a domain concept tree as an expression of a system’s actual structure. We estimated and maintained user’s knowledge about the system on the tree. Every node in the tree has values representing the degree to which a user understands the concepts corresponding to the nodes. The values are updated based on the content of user’s utterances and help messages the system gives. Help messages provided for users are determined by referring to the domain concept tree and identifying concepts the user does not understand. We evaluated our method by testing twelve novice subjects. Both the average time to complete tasks and the number of utterances significantly decreased because of the help messages provided by our method. Index Terms: spoken dialogue system, adaptive help generation, novice user Yuichiro Fukubayashi, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno |
INTERSPEECH | 2 |
| 2006 | Improving speech recognition of two simultaneous speech signals by integrating ICA BSS and automatic missing feature mask generation
Ryu Takeda, Shun'ichi Yamamoto, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno |
INTERSPEECH | 3 |
| 2006 | Multiple Acoustical Holography Method for Localization of Objects in Broad Range using Audible SoundabstractThis paper describes a new acoustic localization method using audible sound, which can be applied over a broader range of search directions. In the field of robotics, most conventional indoor localization systems based on sonar range finders use ultrasound to obtain a highly accurate distance. Because ultrasound has high directivity, many measurements are required to localize objects in a large space. To achieve localization with one-time measurement, we use audible sounds. We then calculate an intensity field of the reflection sound to estimate object positions. Although acoustical holography (AH) is a well-known technique to do this, it has problems in that it generates false images. We propose multiple AH (MAH) to solve this problem. The method is used to divide a measurement plane into sub-planes and to apply AH to each sub-plane. By integrating the results of applying AH to the sub-planes, false images can be suppressed because the positions of the false images differ depending on the position of the sub-plane. In addition, we use multiple frequencies to advance an accuracy of the localization based on MAH in a real environment. We constructed a localization system with only one speaker and a microphone array. In both simulation and actual experiments, we confirmed that MAH was effective method for the suppression of false image and could be used within the range of an angle view of 120 deg Haruhiko Niwa, Tetsuya Ogata, Kazunori Komatani, Hiroshi G. Okuno |
IROS | 3 |
| 2006 | Missing-Feature based Speech Recognition for Two Simultaneous Speech Signals Separated by ICA with a pair of Humanoid EarsabstractRobot audition is a critical technology in making robots symbiosis with people. Since we hear a mixture of sounds in our daily lives, sound source localization and separation, and recognition of separated sounds are three essential capabilities. Sound source localization has been recently studied well for robots, while the other capabilities still need extensive studies. This paper reports the robot audition system with a pair of omni-directional microphones embedded in a humanoid to recognize two simultaneous talkers. It first separates sound sources by independent component analysis (ICA) with single-input multiple-output (SIMO) model. Then, spectral distortion for separated sounds is estimated to identify reliable and unreliable components of the spectrogram. This estimation generates the missing feature masks as spectrographic masks. These masks are then used to avoid influences caused by spectral distortion in automatic speech recognition based on missing-feature method. The novel ideas of our system reside in estimates of spectral distortion of temporal-frequency domain in terms of feature vectors. In addition, we point out that the voice-activity detection (VAD) is effective to overcome the weak point of ICA against the changing number of talkers. The resulting system outperformed the baseline robot audition system by 15% Ryu Takeda, Shun'ichi Yamamoto, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno |
IROS | 3 |
| 2006 | Real-Time Robot Audition System That Recognizes Simultaneous Speech in The Real WorldabstractThis paper presents a robot audition system that recognizes simultaneous speech in the real world by using robot-embedded microphones. We have previously reported missing feature theory (MFT) based integration of sound source separation (SSS) and automatic speech recognition (ASR) for building robust robot audition. We demonstrated that a MFT-based prototype system drastically improved the performance of speech recognition even when three speakers talked to a robot simultaneously. However, the prototype system had three problems; being offline, hand-tuning of system parameters, and failure in voice activity detection (VAD). To attain online processing, we introduced FlowDesigner-based architecture to integrate sound source localization (SSL), SSS and ASR. This architecture brings fast processing and easy implementation because it provides a simple framework of shared-object-based integration. To optimize the parameters, we developed genetic algorithm (GA) based parameter optimization, because it is difficult to build an analytical optimization model for mutually dependent system parameters. To improve VAD, we integrated new VAD based on a power spectrum and location of a sound source into the system, since conventional VAD relying only on power often fails due to low signal-to-noise ratio of simultaneous speech. We, then, constructed a robot audition system for Honda ASIMO. As a result, we showed that the system worked online and fast, and had a better performance in robustness and accuracy through experiments on recognition of simultaneous speech in a noisy and echoic environment Shun'ichi Yamamoto, Kazuhiro Nakadai, Mikio Nakano, Hiroshi Tsujino, Jean-Marc Valin, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno |
IROS | 6 |
| 2006 | Experience Based Imitation Using RNNPBabstractRobot imitation is a useful and promising alternative to robot programming. Robot imitation involves two crucial issues. The first is how a robot can imitate a human whose physical structure and properties differ greatly from its own. The second is how the robot can generate various motions from finite programmable patterns (generalization). This paper describes a novel approach to robot imitation based on its own physical experiences. Let us consider a target task of moving an object on a table. For imitation, we focused on an active sensing process in which the robot acquires the relation between the object's motion and its own arm motion. For generalization, we applied a recurrent neural network with parametric bias (RNNPB) model to enable recognition/generation of imitation motions. The robot associates the arm motion which reproduces the observed object's motion presented by a human operator. Experimental results demonstrated that our method enabled the robot to imitate not only motion it has experienced but also unknown motion, which proved its capability for generalization Ryunosuke Yokoya, Tetsuya Ogata, Jun Tani, Kazunori Komatani, Hiroshi G. Okuno |
IROS | 4 |
| 2006 | Automatic Synchronization between Lyrics and Music CD Recordings Based on Viterbi Alignment of Segregated Vocal SignalsabstractThis paper describes a system that can automatically synchronize between polyphonic musical audio signals and corresponding lyrics. Although there were methods that can synchronize between monophonic speech signals and corresponding text transcriptions by using Viterbi alignment techniques, they cannot be applied to vocals in CD recordings because accompaniment sounds often overlap with vocals. To align lyrics with such vocals, we therefore developed three methods: a method for segregating vocals from polyphonic sound mixtures, a method for detecting vocal sections, and a method for adapting a speech-recognizer phone model to segregated vocal signals. Experimental results for 10 Japanese popular-music songs showed that our system can synchronize between music and lyrics with satisfactory accuracy for 8 songs Hiromasa Fujihara, Masataka Goto, Jun Ogata, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno |
ISM | 4 |
| 2006 | Musical Instrument Recognizer "Instrogram" and Its Application to Music Retrieval Based on Instrumentation SimilarityabstractInstrumentation is an important cue in retrieving musical content. Conventional methods for instrument recognition performing notewise require accurate estimation of the onset time and fundamental frequency (FO) for each note, which is not easy in polyphonic music. This paper presents a non-notewise method for instrument recognition in polyphonic musical audio signals. Instead of such note-wise estimation, our method calculates the temporal trajectory of instrument existence probabilities for every FO and visualizes it as a spectrogram-like graphical representation, called an instrogram. This method can avoid the influence by errors of onset detection and FO estimation because it does not use them. We also present methods for MPEG-7-based instrument annotation and music information retrieval based on the similarity between instrograms. Experimental results with realistic music show the average accuracy of 76.2% for the instrument annotation and that the instrogram-based similarity measure represents the actual instrumentation similarity better than an MFCC-based one Tetsuro Kitahara, Masataka Goto, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno |
ISM | 3 |
| 2006 | Recognition of Simultaneous Speech by Estimating Reliability of Separated Signals for Robot Audition
Shun'ichi Yamamoto, Ryu Takeda, Kazuhiro Nakadai, Mikio Nakano, Hiroshi Tsujino, Jean-Marc Valin, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno |
PRICAI | 7 |
| 2005 | Distance-Based Dynamic Interaction of Humanoid Robot with Multiple People
Tsuyoshi Tasaki, Shohei Matsumoto, Hayato Ohba, Mitsuhiko Toda, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno |
IEA/AIE | 5 |
| 2005 | Contextual constraints based on dialogue models in database search task for spoken dialogue systemsabstractThis paper describes the incorporation of contextual information into spoken dialogue systems in the database search task. Appropriatedialoguemodeling is requiredto manageautomatic speech recognition (ASR) errors using dialogue-level information. We define two dialogue models: a model for dialogue flow and a model of structured dialogue history. The model for dialogueflowassumesdialoguesin the databasesearchtaskconsist of only two modes. In the structured dialogue history model, query conditions are maintained as a tree structure, taking into consideration their inputted order. The constraints derived from these models are integrated by using a decision tree learning, so that the system candeterminea dialogueact of the utteranceand whether each content word should be accepted or rejected, even when it contains ASR errors. The experimental result showed that our method could interpret content words better than conventional one without the contextual information. Furthermore, it was also shown that our method was domain-independent because it achieved equivalent accuracy in another domain without any more training. Kazunori Komatani, Naoyuki Kanda, Tetsuya Ogata, Hiroshi G. Okuno |
INTERSPEECH | 1 |
| 2005 | Multiple moving speaker tracking by microphone array on mobile robotabstractReal-world applications often require tracking multiple moving speakers for improving human-robot interactions and/or sound source separation. This paper presents multiple moving speaker tracking using an 8ch microphone array system installed on a mobile robot. This problem is difficult because the system does not assume that sound sources and/or the microphone array are fixed. Our solutions consist of two key ideas – time delay of arrival estimation, and multiple Kalman filters. The former localizes multiple sound sources based on beamforming in real time. Non-linear movements are tracked by using a set of Kalman filters with different history lengths in order to reduce errors in tracking multiple moving speakers under noisy and echoic environments. For quantitative evaluation of the tracking, motion references of sound sources and a mobile robot, called SIG2, were measured accurately by ultrasonic 3D tag sensors. As a result, we showed that the system tracked three simultaneous sound sources even when SIG2 moved in a room with large reverberation due to glass walls. 1. Masamitsu Murase, Shun'ichi Yamamoto, Jean-Marc Valin, Kazuhiro Nakadai, Kentaro Yamada, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno |
INTERSPEECH | 6 |
| 2005 | Extracting multi-modal dynamics of objects using RNNPBabstractDynamic features play an important role in recognizing objects that have similar static features in colors and or shapes. This paper focuses on active sensing that exploits dynamic feature of an object. An extended version of the robot, Robovie-IIs, moves an object by its arm to obtain its dynamic features. Its issue is how to extract symbols from various kinds of temporal states of the object. We use the recurrent neural network with parametric bias (RNNPB) that generates self-organized nodes in the parametric bias space. The RNNPB with 42 neurons was trained with the data of sounds, trajectories, and tactile sensors generated while the robot was moving/hitting an object with its own arm. The clusters of 20 kinds of objects were successfully self-organized. The experiments with unknown (not trained) objects demonstrated that our method configured them in the PB space appropriately, which proves its generalization capability. Tetsuya Ogata, Hayato Ohba, Jun Tani, Kazunori Komatani, Hiroshi G. Okuno |
IROS | 4 |
| 2005 | Spatially mapping of friendliness for human-robot interactionabstractIt is important that robots interact with multiple people. However, most research has dealt with only interaction between one robot and one person and assumed that the distance between them does not change. This paper focuses on the spatial relationships between a robot and multiple people during interaction. Based on the distance between them, our robot selects appropriate functions to use. It does this using a method we developed for spatially mapping the friendliness of each space around the robot. The robot interacts with the highest friendliness spaces (people) selectively, thereby enabling interaction between the robot and multiple people. Our humanoid robot, SIG2 which the proposed method was implemented into, interacted with about 30 visitors, at the Kyoto University Museum. The results obtained using questionnaires after interaction showed that the actions of SIG2 were easy to understand even when it interacted with multiple people at the same time and that SIG2 behaved in a friendly manner. Tsuyoshi Tasaki, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno |
IROS | 2 |
| 2005 | Making a robot recognize three simultaneous sentences in real-timeabstractA humanoid robot under real-world environments usually hears mixtures of sounds, and thus three capabilities are essential for robot audition; sound source localization, separation, and recognition of separated sounds. We have adopted the missing feature theory (MFT) for automatic recognition of separated speech, and developed the robot audition system. A microphone array is used along with a real-time dedicated implementation of geometric source separation (GSS) and a multi-channel post-filter that gives us a further reduction of interferences from other sources. The automatic speech recognition based on MFT recognizes separated sounds by generating missing feature masks automatically from the post-filtering step. The main advantage of this approach for humanoid robots resides in the fact that the ASR with a clean acoustic model can adapt the distortion of separated sound by consulting the post-filter feature masks. In this paper, we used the improved Julius as an MFT-based automatic speech recognizer (ASR). The Julius is a real-time large vocabulary continuous speech recognition (LVCSR) system. We performed the experiment to evaluate our robot audition system. In this experiment, the system recognizes a sentence, not an isolated word. We showed the improvement in the system performance through three simultaneous speech recognition on the humanoid SIG2. Shun'ichi Yamamoto, Kazuhiro Nakadai, Jean-Marc Valin, Jean Rouat, François Michaud, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno |
IROS | 6 |
| 2005 | Empirical Verification of Meaning-Game-based Generalization of Centering Theory with Large Japanese Corpus
Shun Shiramatsu, Kazunori Komatani, Takashi Miyata, Koichi Hashida, Hiroshi G. Okuno |
PACLIC | 2 |
| 2005 | User Modeling in Spoken Dialogue Systems to Generate Flexible Guidance
Kazunori Komatani, Shinichi Ueno, Tatsuya Kawahara, Hiroshi G. Okuno |
User Model. User Adapt. Interact. | 1 |
| 2004 | Efficient Confirmation Strategy for Large-scale Text Retrieval Systems with Spoken Dialogue Interface
Kazunori Komatani, Teruhisa Misu, Tatsuya Kawahara, Hiroshi G. Okuno |
COLING | 1 |
| 2004 | Recognition of Emotional States in Spoken Dialogue with a Robot
Kazunori Komatani, Ryosuke Ito, Tatsuya Kawahara, Hiroshi G. Okuno |
IEA/AIE | 1 |
| 2004 | Disambiguation in determining phonemes of sound-imitation words for environmental sound recognitionabstractOnomatopoeia, or sound-imitation words (SIWs) are important in informing sound events in human-computer communication. One problem is listener-dependency in recognizing environmental sounds by means of SIWs, that is, different listener hears the same environmental sound as a different SIW even under the same condition. Therefore, the use of usual Japanese phonemes is not adequate to express SIWs. To cope with this ambiguity problem of phoneme determination, we designed a set of new phonemes, referred to as the basic phoneme-groups, to represent environmental sounds. The basic phonemegroup consists of one or more Japanese phonemes, and thus the ambiguity problem is resolved based on it by generating one or more SIWs for a sound event. An HMM-based scheme is adopted to recognize SIWs using the phoneme-groups. Listening experiments with seven subjects showed that automatic SIW recognition based on the basic phoneme-groups outperformed ones based on the other types of phonemes. The recall and precision rate were 56.4% and 72.2%, respectively. Kazushi Ishihara, Yuya Hattori, Tomohiro Nakatani, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno |
INTERSPEECH | 4 |
| 2004 | Robot motion control using listener's back-channels and head gesture informationabstractA novel method is described for robot gestures and utterances during a dialogue based on the listener’s understanding and interest, which are recognized from back-channels and head gestures. “Back-channels” are defined as sounds like ‘uhhuh’ uttered by a listener during a dialogue, and “head gestures” are defined as nod and tilt motions of the listener’s head. The back-channels are recognized using sound features such as power and fundamental frequency. The head gestures are recognized using the movement of the skin-color area and the optical flow data. Based on the estimated understanding and interest of the listener, the speed and size of robot motions are changed. This method was implemented in a humanoid robot called SIG2. Experiments with six participants demonstrated that the proposed method enabled the robot to increase the listener’s level of interest against the dialogue. Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno, Tsuyoshi Tasaki, Takeshi Yamaguchi |
INTERSPEECH | 1 |
| 2004 | Confirmation strategy for document retrieval systems with spoken dialog interfaceabstractAdequate confirmation is indispensable in spoken dialog systems to eliminate misunderstandings caused by speech recognition errors. Spoken language also inherently includes redundant expressions such as disfluency and out-of-domain phrases, which do not contribute to task achievement. It is easy to define a set of keywords to be confirmed for conventional database query tasks, but not straightforward in general document retrieval tasks. In this paper, we propose two statistical measures for identifying portions to be confirmed. A relevance score (RS) represents matching degree with the document set. A significance score (SS) detects portions that consequently affect the retrieval results. With these measures, the system can generate confirmation prior to and posterior to the retrieval, respectively. The strategy is implemented and evaluated with retrieval from software support knowledge base of 40K entries. It is shown that the proposed strategy using the two measures is more efficient than using the conventional confidence measure. Teruhisa Misu, Tatsuya Kawahara, Kazunori Komatani |
INTERSPEECH | 3 |
| 2003 | Flexible Guidance Generation Using User Model in Spoken Dialogue SystemsabstractWe address appropriate user modeling in order to generate cooperative responses to each user in spoken dialogue systems. Unlike previous studies that focus on user's knowledge or typical kinds of users, the user model we propose is more comprehensive. Specifically, we set up three dimensions of user models: skill level to the system, knowledge level on the target domain and the degree of hastiness. Moreover, the models are automatically derived by decision tree learning using real dialogue data collected by the system. We obtained reasonable classification accuracy for all dimensions. Dialogue strategies based on the user modeling are implemented in Kyoto city bus information system that has been developed at our laboratory. Experimental evaluation shows that the cooperative responses adaptive to individual users serve as good guidance for novice users without increasing the dialogue duration for skilled users. Kazunori Komatani, Shinichi Ueno, Tatsuya Kawahara, Hiroshi G. Okuno |
ACL | 1 |
| 2003 | Spoken dialogue system for queries on appliance manuals using hierarchical confirmation strategyabstractWe address a dialogue framework for queries on manuals of electric appliances with a speech interface. Users can makequeriesbyunconstrainedspeech, fromwhichkeywords are extracted and matched to the items in the manual. As a result, so many items are usually obtained. Thus, we introduce an effective dialogue strategy which narrows down the items using a tree structure extracted from the manual. Three cost functions are presented and compared to minimize the number of dialogue turns. We have evaluated the system performance on VTR manual query task. The numberof averagedialogueturnsis reduced to 71% using our strategy compared with a conventional method that makes confirmation in turn according to the matching likelihood. Thus, the proposed system helps users find their intended items more efficiently. Tatsuya Kawahara, Ryosuke Ito, Kazunori Komatani |
INTERSPEECH | 3 |
| 2003 | User modeling in spoken dialogue systems for flexible guidance generationabstractWe address appropriate user modeling in order to generate cooperative responses to each user in spoken dialogue systems. Unlike previous studies that focus on users’ knowledge or typical kinds of users, the proposed user model is more comprehensive. Specifically, we set up three dimensions of user models: skill level to the system, knowledge level on the target domain and degree of hastiness. Moreover, the models are automatically derived by decision tree learning using real dialogue data. We obtained reasonable classification accuracy for all dimensions. Dialogue strategies based on the user modeling are implemented in Kyoto city bus information system that has been developed at our laboratory. Experimental evaluation shows that the cooperative responses adaptive to individual users serve as good guidance for novice users without increasing the dialogue duration for skilled users. Kazunori Komatani, Shinichi Ueno, Tatsuya Kawahara, Hiroshi G. Okuno |
INTERSPEECH | 1 |
| 2002 | Efficient Dialogue Strategy to Find Users' Intended Items from Information Query Results
Kazunori Komatani, Tatsuya Kawahara, Ryosuke Ito, Hiroshi G. Okuno |
COLING | 1 |
| 2001 | Domain-independent spoken dialogue platform using key-phrase spotting based on combined language modelabstractWe present a portable platform for spoken dialogue systems and its experimental evaluation. Conventional development of speech interfaces involves much labor cost in either describing a task grammar or collecting a task corpus. Our platform automatically generates a lexicon and a language model of keyphrases based on task description and structure of the domain database. By spotting key-phrases using both the generated grammar and word 2-gram model trained with dialogue corpora of similar domains, we realize flexible speech understanding on a variety of utterances. Furthermore, adopting a GUI that explicitly displays acceptable utterance patterns is effective in guiding user utterances within the system’s capability. We evaluate the generated spoken dialogue system using 24 novice users. The number of unacceptable utterances are significantly reduced with the simple phrase grammar and GUI. And the phrase spotter using the combined language model improves the semantic accuracy by 15.5% compared with the conventional method decoding the whole sentence with a fixed grammar. Kazunori Komatani, Katsuaki Tanaka, Hiroaki Kashima, Tatsuya Kawahara |
INTERSPEECH | 1 |
| 2000 | Flexible Mixed-Initiative Dialogue Management using Concept-Level Confidence Measures of Speech Recognizer Output
Kazunori Komatani, Tatsuya Kawahara |
COLING | 1 |
| 2000 | Generating effective confirmation and guidance using two-level confidence measures for dialogue systemsabstractWe present a method to generate effective confirmation and guidance using concept-level confidence measures (CM) derived from speech recognizer output in order to handle speech recognition errors. We define two conceptlevel CM, which are on content-words and on semanticattributes, using 10-best outputs of the speech recognizer and parsing with phrase-level grammars. Content-word CM is useful for selecting plausible interpretations. Less confident interpretations are given to confirmation process, and non-confident ones are rejected. The strategy improved the interpretation accuracy by 11.5%. Moreover, the semantic-attribute CM is used to estimate user’s intention and generates system-initiative guidances even when successful interpretation is not obtained. Kazunori Komatani, Tatsuya Kawahara |
INTERSPEECH | 1 |