Kazuhiro Nakadai

dblp:38/545 · DBLP profile ↗
← Back
189ranked-venue papers
26as first author
26since 2021 · last 2026
0000-0002-6134-4558ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 169 · 22 first-author · 22 since 2021Systems, architecture and hardware · 88 · 15 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 58 · 10 first-author · 12 since 2021Human-computer interaction and ubiquitous computing · 16 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 13 · 5 since 2021Databases, data management, data science and information retrieval · 2Theory of computation · 1
YearPublicationVenuePosition
2026 Unsupervised Single-Channel Audio Separation with Diffusion Source Priors
abstract
Single-channel audio separation aims to separate individual sources from a single-channel mixture. Most existing methods rely on supervised learning with synthetically generated paired data. However, obtaining high-quality paired data in real-world scenarios is often difficult. This data scarcity can degrade model performance under unseen conditions and limit generalization ability. To this end, in this work, we approach this problem from an unsupervised perspective, framing it as a probabilistic inverse problem. Our method requires only diffusion priors trained on individual sources. Separation is then achieved by iteratively guiding an initial state toward the solution through reconstruction guidance. Importantly, we introduce an advanced inverse problem solver specifically designed for separation, which mitigates gradient conflicts caused by interference between the diffusion prior and reconstruction guidance during inverse denoising. This design ensures high-quality and balanced separation performance across individual sources. Additionally, we find that initializing the denoising process with an augmented mixture instead of pure Gaussian noise provides an informative starting point that significantly improves the final performance. To further enhance audio prior modeling, we design a novel time–frequency attention-based network architecture that demonstrates strong audio modeling capability. Collectively, these improvements lead to significant performance gains, as validated across speech–sound event, sound event, and speech separation tasks.
Runwu Shi, Khan Nabeela Khanum, Benjamin Yen 0001, Takeshi Ashizawa, Kazuhiro Nakadai
AAAI8
2025 Improvement in Sign Language Translation Using Text CTC Alignment
abstract
Current sign language translation (SLT) approaches often rely on gloss-based supervision with Connectionist Temporal Classification (CTC), limiting their ability to handle non-monotonic alignments between sign language video and spoken text. In this work, we propose a novel method combining joint CTC/Attention and transfer learning. The joint CTC/Attention introduces hierarchical encoding and integrates CTC with the attention mechanism during decoding, effectively managing both monotonic and non-monotonic alignments. Meanwhile, transfer learning helps bridge the modality gap between vision and language in SLT. Experimental results on two widely adopted benchmarks, RWTH-PHOENIX-Weather 2014 T and CSL-Daily, show that our method achieves results comparable to state-of-the-art and outperforms the pure-attention baseline. Additionally, this work opens a new door for future research into gloss-free SLT using text-based CTC alignment.
Sihan Tan, Taro Miyazaki, Khan Nabeela Khanum, Kazuhiro Nakadai
COLING4
2025 Distance Based Single-Channel Target Speech Extraction
abstract
This paper aims to achieve single-channel target speech extraction (TSE) in enclosures by solely utilizing distance information. This is the first work that utilizes only distance cues without using speaker physiological information for single-channel TSE. Inspired by recent single-channel Distance-based separation and extraction methods, we introduce a novel model that efficiently fuses distance information with time-frequency (TF) bins for TSE. Experimental results in both single-room and multi-room scenarios demonstrate the feasibility and effectiveness of our approach. This method can also be employed to estimate the distances of different speakers in mixed speech. Online demos are available at https://runwushi.github.io/distance-demo-page/.
Runwu Shi, Benjamin Yen 0001, Kazuhiro Nakadai
ICASSP3
2025 SignFlow: End-to-End Sign Language Generation for One-to-Many Modeling using Conditional Flow Matching
Khan Nabeela Khanum, Bowen Wu 0002, Sihan Tan, Carlos Toshinori Ishi, Kazuhiro Nakadai
ICMI5
2025 MultiGAU: Real Time Sign Language Generation Using Multimodal Gated Attention
Khan Nabeela Khanum, Bowen Wu 0002, Carlos Toshinori Ishi, Kazuhiro Nakadai
IEA/AIE (1)4
2025 Swarm Active Audition with Robots and Drones: Real-World Performance Validation
abstract
Search and rescue (SAR) operations in large-scale disaster sites, such as areas affected by earthquakes, require rapid victim detection. While drones equipped with cameras are commonly used for SAR, their effectiveness is limited in visually obstructed environments, because of debris, smoke, or fog. Under such situations, auditory information can play a crucial role in locating victims who are not visible. Existing drone audition research has demonstrated the feasibility of detecting sound sources using onboard microphone arrays. However, most studies focus on single-drone systems, which face limitations in coverage and accessibility, particularly in complex environments such as collapsed buildings or urban canyons. Additionally, real-world validation of multi-drone audition systems remains limited, with prior studies relying primarily on simulations or controlled environments. To address these challenges, we propose and evaluate a Multi-Drone and Robot-Based Active Audition System (SAAS-RD: Swarm Active Audition System with Robots and Drones) that integrates multiple drones and ground robots to enhance acoustic search capabilities. Our work focuses on real-world performance validation, conducting field experiments in outdoor environments and analyzing system feasibility through case studies. The results demonstrate the potential of SAAS-RD as a practical solution for large-scale SAR operations.
Kazuhiro Nakadai, Kotaro Hoshiba, Benjamin Yen 0001, Makoto Kumon, Yoko Sasaki
IROS1
2025 Single-Microphone-Based Sound Source Localization for Mobile Robots in Reverberant Environments
abstract
Accurately estimating sound source positions is crucial for robot audition. However, existing sound source localization methods typically rely on a microphone array with at least two spatially preconfigured microphones. This requirement hinders the applicability of microphone-based robot audition systems and technologies. To alleviate these challenges, we propose an online sound source localization method that uses a single microphone mounted on a mobile robot in reverberant environments. Specifically, we develop a lightweight neural network model with only 43k parameters to perform real-time distance estimation by extracting temporal information from reverberant signals. The estimated distances are then processed using an extended Kalman filter to achieve online sound source localization. To the best of our knowledge, this is the first work to achieve online sound source localization using a single microphone on a moving robot, a gap that we aim to fill in this work. Extensive experiments demonstrate the effectiveness and merits of our approach. To benefit the broader research community, we have open-sourced our code at https://github.com/JiangWAV/single-mic-SSL.
Runwu Shi, Benjamin Yen 0001, He Kong 0001, Kazuhiro Nakadai
IROS5
2025 Towards Online Sign Language Expression for Real-Time Human-Robot Interaction
abstract
Sign language (SL) is the primary mode of communication for Deaf and Hard-of-Hearing (DHH) individuals and differs fundamentally from spoken languages. While Sign Language Expression (SLE) systems have made significant progress in generating gestures from text using deep learning, their integration into assistive Human-Robot Interaction (HRI) remains limited. This position paper introduces online SLE as a novel paradigm for enabling responsive, real-time SL communication on robotic platforms. We analyze the technical, dataset, and evaluation challenges in deploying SLE models on robots and present preliminary experiments illustrating the trade-offs between efficiency and expressiveness. We further propose design considerations for online model architectures, identify key gaps in current datasets, and call for interdisciplinary collaboration with the Deaf community. Our goal is to pave the way toward inclusive, socially-aware robotic agents capable of natural SL communication.
Khan Nabeela Khanum, Sihan Tan, Kazuhiro Nakadai
RO-MAN3
2024 UAV-Enhanced Combination to Application: Comprehensive Analysis and Benchmarking of a Human Detection Dataset for Disaster Scenarios
Ragib Amin Nihal, Benjamin Yen 0001, Katsutoshi Itoyama, Kazuhiro Nakadai
ICPR (14)4
2024 Improving Noise Robustness of Automatic Speech Recognition Based on a Parallel Adapter Model with Near-Identity Initialization
Takahiro Osaki, Yui Sudo, Katsutoshi Itoyama, Kenji Nishida, Kazuhiro Nakadai
IEA/AIE5
2024 Improving Impressions of Response Delay in AI-based Spoken Dialogue Systems
abstract
This paper aims to mitigate impression loss due to response delays in spoken dialogue systems. The system addressed in this paper consists of connected ASR, NLP, and TTS AI services, and large response delays due to processing time and network latency in each module may give a bad impression to users. Therefore, we propose three techniques to deal with the inevitable delays: filler insertion, parallel speech processing, and prompt editing. To evaluate the effectiveness of the proposed method, measurements and subject experiments were conducted. The results showed that filler insertion had statistically significant improvements of 1.1 points on a Likert scale in the impression of response delay. Parallel speech processing reduced delay by 0.6 seconds. Prompt editing was also effective in reducing delay through response length suppression.
Shuhei Asaka, Katsutoshi Itoyama, Kazuhiro Nakadai
RO-MAN3
2024 SLAM-Based Joint Calibration of Multiple Asynchronous Microphone Arrays and Sound Source Localization
abstract
Robot audition systems with multiple microphone arrays have many applications in practice. However, the accurate calibration of multiple microphone arrays remains challenging because there are many unknown parameters to be identified, including the relative transforms (i.e., orientation and translation) and asynchronous factors (i.e., initial time offset and sampling clock difference) between microphone arrays. To tackle these challenges, in this article, we adopt batch simultaneous localization and mapping (SLAM) for joint calibration of multiple asynchronous microphone arrays and sound source localization. Using the Fisher information matrix (FIM) approach, we first conduct the observability analysis (i.e., parameter identifiability) of the abovementioned calibration problem and establish necessary/sufficient conditions under which the FIM and the Jacobian matrix have full column rank, which implies the identifiability of the unknown parameters. We also discover several scenarios where the unknown parameters are not uniquely identifiable. Subsequently, we propose an effective framework to initialize the unknown parameters, which is used as the initial guess in batch SLAM for multiple microphone array calibration, aiming to further enhance optimization accuracy and convergence. Extensive numerical simulations and real experiments have been conducted to verify the performance of the proposed method. The experimental results show that the proposed pipeline achieves higher accuracy with fast convergence in comparison to methods that use the noise-corrupted ground truth of the unknown parameters as the initial guess in the optimization and other existing frameworks.
Yuanzheng He, Daobilige Su, Katsutoshi Itoyama, Kazuhiro Nakadai, Junfeng Wu 0001, Shoudong Huang, Youfu Li 0001, He Kong 0001
IEEE Trans. Robotics5
2023 miniStreamer: Enhancing Small Conformer with Chunked-Context Masking for Streaming ASR Applications on the Edge
Haris Gulzar, Monikka Roslianna Busto, Takeharu Eda, Katsutoshi Itoyama, Kazuhiro Nakadai
INTERSPEECH5
2023 Retraining-free Customized ASR for Enharmonic Words Based on a Named-Entity-Aware Model and Phoneme Similarity Estimation
Yui Sudo, Kazuya Hata, Kazuhiro Nakadai
INTERSPEECH3
2023 Improving Sign Language Understanding Introducing Label Smoothing
abstract
Sign language is one of the most important communication methods when considering equality, diversity, and inclusion. Sign language understanding implies understanding sign language using machines, and it involves mainly two functions; sign language recognition and sign language translation. To improve sign language understanding performance, this paper proposes to use label smoothing with CTC (Connectionist Temporal Classification) loss as training criteria for the sign language understanding neural network. Experimental results showed the effectiveness of the proposed method in both sign language recognition and translation.
Sihan Tan, Khan Nabeela Khanum, Katsutoshi Itoyama, Kazuhiro Nakadai
RO-MAN4
2023 Online Adaptation of Fourier Series Based Acoustic Transfer Function Model to Improve Sound Source Localization and Separation
abstract
This paper proposes an online adaptation method for Fourier series based acoustic transfer function (TF) models for robot audition systems based on microphone array signal processing. The TF represents the signal propagation characteristics from a sound source to a microphone, which is an essential component for real-world auditory scene analysis, including sound source localization and separation. The real-world applications of TF-based array signal processing requires two characteristics: 1) adaptability to changes in the acoustic environment (changes in the signal propagation characteristics between the sound source and the microphone), and 2) a lightweight TF set for use in embedded systems such as robots with limited memory and computational resources. This paper proposes an online adaptation method for lightweight TF models using the Fourier series expansion. This method has both above two characteristics. Experimental results showed that the use of TF set adapted online using the proposed method performs better sound source localization and separation performance than existing online TF adaptation methods.
Yui Sudo, Masayuki Takigahira, Hideo Tsuru, Kazuhiro Nakadai, Hirofumi Nakajima
RO-MAN4
2022 Weakly-Supervised Neural Full-Rank Spatial Covariance Analysis for a Front-End System of Distant Speech Recognition
Yoshiaki Bando, Takahiro Aizawa, Katsutoshi Itoyama, Kazuhiro Nakadai
INTERSPEECH4
2022 Streaming Automatic Speech Recognition with Re-blocking Processing Based on Integrated Voice Activity Detection
Yui Sudo, Muhammad Shakeel 0001, Kazuhiro Nakadai, Jiatong Shi, Shinji Watanabe 0001
INTERSPEECH3
2022 Empirical Sampling from Latent Utterance-wise Evidence Model for Missing Data ASR based on Neural Encoder-Decoder Model
Ryu Takeda, Yui Sudo, Kazuhiro Nakadai, Kazunori Komatani
INTERSPEECH3
2022 Spotforming by NMF Using Multiple Microphone Arrays
abstract
Sound source separation is a method to extract a target sound source from a mixture of various sound sources and noises. One of the typical sound source separation methods is beamforming, which can separate sound sources by direction based on the phase difference between channels from the recorded signal of a microphone array, a multi-channel recording system. However, beamforming is a direction-based method and cannot separate multiple sources in the same direction. In this paper, we propose a method for separating sources in the same direction using multiple microphone arrays. The proposed method performs beamforming using multiple microphone arrays and extracts only the target sound source from the separated sound by the Non-negative Matrix Factorization (NMF), thus reducing the influence of other sources in the same direction. In this paper, to investigate the effectiveness of the proposed method, experiments were conducted assuming the presence of another sound source in the same direction from an arbitrary microphone array. The results show that the proposed method outperforms the delay-sum method in a simulation environment. In addition, experiments were conducted in a real environment to verify the effect of reverberation.
Yasuhiro Kagimoto, Katsutoshi Itoyama, Kenji Nishida, Kazuhiro Nakadai
IROS4
2022 Outdoor evaluation of sound source localization for drone groups using microphone arrays
abstract
For robot and drone auditions, microphone arrays have been used for estimating sound source directions and sound source locations. By using sound source localization techniques, for example, drones can detect people calling for help even if the target person is not visible. Most sound source localization methods are based on estimated sound source directions and triangulation. However, when it comes to situations using drones, severe drone noise distorts direction estimation results which could worsen the localization results badly due to the discreteness of direction estimation. In this perspective, the authors have proposed a sound source localization method that can omit outlying triangulation points, which could improve its localization performance. In this paper, an outdoor experiment has been held, and the proposed method is evaluated whether it can localize a sound source even if real drone noise is added to the recordings. Experiment results show that the proposed method can localize with 4.15 m of estimation error for a sound source up to 50 m away, suppress the impact of outliers, and use only plausible triangulation points.
Taiki Yamada, Katsutoshi Itoyama, Kenji Nishida, Kazuhiro Nakadai
IROS4
2021 Assessment of von Mises-Bernoulli Deep Neural Network in Sound Source Localization
Katsutoshi Itoyama, Yoshiya Morimoto, Shungo Masaki, Ryosuke Kojima, Kenji Nishida, Kazuhiro Nakadai
Interspeech6
2021 Fully-Online Always-Adaptation of Transfer Functions and Its Application to Sound Source Localization and Separation
abstract
This paper addresses fully-online always-adaptation of a transfer function for robot audition systems based on microphone array processing. The transfer function represents signal propagation characteristics between a microphone and a sound source, which provides essential information for real-world scene analysis, such as sound source localization and separation for robots. Although it is commonly defined as a stationary function, it should be considered together with room acoustics and their environmental changes for practical use, that is, it should be defined as a dynamically-changing function. To fulfill this requirement, we propose a fully-online always-adaptation method for a transfer function, by continuously estimating the transfer function from the observed signals in a passive manner, while performing sound source localization and separation. The proposed method was implemented on open source robot audition software HARK as modules which works online. These modules are applied to sound source localization and separation which are primary functions in robot audition. Experimental results showed that the proposed method successfully adapted to an office environment and improved the performance of sound source localization and separation at a close level to the transfer function recorded in the room.
Kazuhiro Nakadai, Masayuki Takigahira, Yusuke Kawai, Hirofumi Nakajima
IROS1
2021 Investigation of Node Pruning Criteria for Neural Networks Model Compression with Non-Linear Function and Non-Uniform Network Topology
abstract
This paper investigates node-pruning-based compression for non-uniform deep learning models such as acoustic models in automatic speech recognition (ASR). Node pruning for small footprint ASR has been well studied, but most studies assumed a sigmoid as an activation function and uniform or simple fully-connected neural networks without bypass connections. We propose a node pruning method that can be applied to non-sigmoid functions such as ReLU and that can deal with network topology related issues such as bypass connections. To deal with non-sigmoid functions, we extend a node entropy technique to estimate node activities. To cope with non-uniform network topology, we propose three criteria; inter-layer pairing, no bypass connection pruning, and layer-based pruning rate configuration. The proposed method as a combination of these four techniques and criteria was applied to compress a Kaldi's acoustic model with ReLU as a non-linear function, time delay neural networks (TDNN) and bypass connections inspired by residual networks. Experimental results showed that the proposed method achieved a 31% speed increase while maintaining the ASR accuracy to be comparable by taking network topology into consideration.
Kazuhiro Nakadai, Yosuke Fukumoto, Ryu Takeda
SLT1
2021 Detecting earthquakes: a novel deep learning-based approach for effective disaster response
Muhammad Shakeel 0001, Katsutoshi Itoyama, Kenji Nishida, Kazuhiro Nakadai
Appl. Intell.4
2021 Multichannel environmental sound segmentation
abstract
Abstract This paper proposes a multichannel environmental sound segmentation method. Environmental sound segmentation is an integrated method to achieve sound source localization, sound source separation and classification, simultaneously. When multiple microphones are available, spatial features can be used to improve the localization and separation accuracy of sounds from different directions; however, conventional methods have three drawbacks: (a) Sound source localization and sound source separation methods using spatial features and classification using spectral features trained in the same neural network, may overfit to the relationship between the direction of arrival and the class of a sound, thereby reducing their reliability to deal with novel events. (b) Although permutation invariant training used in autonomous speech recognition could be extended, it is impractical for environmental sounds that include an unlimited number of sound sources. (c) Various features, such as complex values of short time Fourier transform and interchannel phase differences have been used as spatial features, but no study has compared them. This paper proposes a multichannel environmental sound segmentation method comprising two discrete blocks, a sound source localization and separation block and a sound source separation and classification block. By separating the blocks, overfitting to the relationship between the direction of arrival and the class is avoided. Simulation experiments using created datasets including 75-class environmental sounds showed the root mean squared error of the proposed method was lower than that of conventional methods.
Yui Sudo, Katsutoshi Itoyama, Kenji Nishida, Kazuhiro Nakadai
Appl. Intell.4
2020 Calibration of a Microphone Array Based on a Probabilistic Model of Microphone Positions
Katsuhiro Dan, Katsutoshi Itoyama, Kenji Nishida, Kazuhiro Nakadai
IEA/AIE4
2020 Synchronization of Microphones Based on Rank Minimization of Warped Spectrum for Asynchronous Distributed Recording
abstract
This paper describes a new method for synchronizing microphones based on spectral warping in an asynchronous microphone array. In an audio signal observed by an asynchronous microphone array, two factors are involved: the time lag caused by a mismatch of the sampling rate and offset between microphones, and the modulation caused by differences in spatial transfer function between the sound source and each microphone. A spectrum warping matrix representing a resampling effect in the frequency domain is formulated and an observation model of audio (spectrum) mixture in an asynchronous microphone array is constructed. The proposed synchronization method uses an iterative optimization algorithm based on gradient descent of a new objective function. The function is formulated as a logarithmic determinant of a spectrum correlation matrix that is derived from relaxation of a rank minimization problem. Experimental results showed that the proposed method effectively estimates modulated sampling rate and that the proposed method outperforms an existing synchronization method.
Katsutoshi Itoyama, Kazuhiro Nakadai
IROS2
2020 Learning Three-dimensional Skeleton Data from Sign Language Video
abstract
Data for sign language research is often difficult and costly to acquire. We therefore present a novel pipeline able to generate motion three-dimensional (3D) skeleton data from single-camera sign language videos only. First, three recurrent neural networks are learned to infer the three-dimensional position data of body, face, and finger joints for a high resolution of the signer’s skeleton. Subsequently, the angular displacements of all joints over time are estimated using inverse kinematics and mapped to a virtual sign avatar for animation. Last, the generated data are evaluated in detail, including a sign language recognition and sign language synthesis scenario. Utilizing a neural word classifier trained on real motion capture data, we reliably classify word segments built from our newly generated position data with similar accuracy as motion capture data (absolute difference 3.8%). Furthermore, qualitative evaluation of sign animations shows that the avatar performs natural movements that are comprehensible and resemble animations created with original motion capture data.
Heike Brock, Felix Law, Kazuhiro Nakadai, Yuji Nagashima
ACM Trans. Intell. Syst. Technol.3
2019 An Integrated Framework for Field Recording, Localization, Classification and Annotation of Birdsongs Using Robot Audition Techniques - Harkbird 2.0
abstract
Bird vocalizations are one of the important subjects in ecoacoustics because birds communicate diversely using various vocalizations such as songs and calls. We have developed a portable system, HARKBird to provide a basic function, i.e., birdsong localization, which automatically extracts sound sources and their direction of arrivals (DOA) using robot audition techniques based on HARK. In this paper, we introduce HARKBird 2.0 which is empowered for higher understanding of birdsongs. A new soundscape annotation tool for localization results is enhanced by an interactive interface for song classification based on an unsupervised feature mapping t-SNE. We show that HARKBird 2.0 provides bird researchers with an integrated framework to analyze spatio-spectro-temporal dynamics of birdsongs using the song analysis of Japanese bush warbler (Horornis diphone).
Shinji Sumitani, Reiji Suzuki, Naoaki Chiba, Shiho Matsubayashi, Takaya Arita, Kazuhiro Nakadai, Hiroshi G. Okuno
ICASSP6
2019 Weakly-Supervised Deep Recurrent Neural Networks for Basic Dance Step Generation
abstract
Synthesizing human's movements such as dancing is a flourishing research field which has several applications in computer graphics. Recent studies have demonstrated the advantages of deep neural networks (DNNs) for achieving remarkable performance in motion and music tasks with little effort for feature pre-processing. However, applying DNNs for generating dance to a piece of music is nevertheless challenging, because of 1) DNNs need to generate large sequences while mapping the music input, 2) the DNN needs to constraint the motion beat to the music, and 3) DNNs require a considerable amount of hand-crafted data. In this study, we propose a weakly supervised deep recurrent method for real-time basic dance generation with audio power spectrum as input. The proposed model employs convolutional layers and a multilayered Long Short-Term memory (LSTM) to process the audio input. Then, another deep LSTM layer decodes the target dance sequence. Notably, this end-to-end approach has 1) an auto-conditioned decode configuration that reduces accumulation of feedback error of large dance sequence, 2) uses a contrastive cost function to regulate the mapping between the music and motion beat, and 3) trains with weak labels generated from the motion beat, reducing the amount of hand-crafted data. We evaluate the proposed network based on i) the similarities between generated and the baseline dancer motion with a cross entropy measure for large dance sequences, and ii) accurate timing between the music and motion beat with an F-measure. Experimental results revealed that, after training using a small dataset, the model generates basic dance steps with low cross entropy and maintains an F-measure score similar to that of a baseline dancer.
Nelson Enrique Yalta Soplin, Shinji Watanabe 0001, Kazuhiro Nakadai, Tetsuya Ogata
IJCNN3
2019 Environmental sound segmentation utilizing Mask U-Net
abstract
This paper proposes an environmental sound segmentation method using Mask U-Net. Recent research in robot audition has analyzed noise reduction, section detection, and sound source separation for use in a real-world environment with many noises and overlaps. However, conventional methods apply respective functions in cascades. The biggest problem of cascade systems is the accumulation of errors generated at each function block. Although many methods of human voice separation have been proposed, robots operating in a real-world environment must be able to separate not only human voices but other environmental sounds. Unlike traditional sound source separation using spatial information, environmental sound segmentation must simultaneously detect sections and separate sound sources based on pre-trained features. One such method, U-Net, which was proposed for semantic segmentation of images, has been applied to the separation of singing voices. However, this method deals only with limited classes of sounds. The current study proposes an environmental sound segmentation method using Mask U-Net, which combines segmentation using U-Net with sound event detection using CNN to 75-classes of environmental sounds. Experimental application confirmed that this method improved learning speed and sound source separation compared with the conventional method.
Yui Sudo, Katsutoshi Itoyama, Kenji Nishida, Kazuhiro Nakadai
IROS4
2018 HARK-Bird-Box: A Portable Real-time Bird Song Scene Analysis System
abstract
This paper addresses real-time bird song scene analysis. Observation of animal behavior such as communication of wild birds would be aided by a portable device implementing a real-time system that can localize sound sources, measure their timing, classify their sources, and visualize these factors of sources. The difficulty of such a system is an integration of these functions considering the real-time requirement. To realize such a system, we propose a cascaded approach, cascading sound source detection, localization, separation, feature extraction, classification, and visualization for bird song analysis. Our system is constructed by combining an open source software for robot audition called HARK and a deep learning library to implement a bird song classifier based on a convolutional neural network (CNN). Considering portability, we implemented this system on a single-board computer, Jetson TX2, with a microphone array and developed a prototype device for bird song scene analysis. A preliminary experiment confirms a computational time for the whole system to realize a real-time system. Also, an additional experiment with a bird song dataset revealed a trade-off relationship between classification accuracy and time consuming and the effectiveness of our classifier.
Ryosuke Kojima, Osamu Sugiyama, Kotaro Hoshiba, Reiji Suzuki, Kazuhiro Nakadai
IROS5
2018 Extracting the Relationship between the Spatial Distribution and Types of Bird Vocalizations Using Robot Audition System HARK
abstract
For a deeper understanding of ecological functions and semantics of wild bird vocalizations (i.e., songs and calls), it is important to clarify the fine-scaled and detailed relationships among their characteristics of vocalizations and their behavioral contexts. However, it takes a lot of time and effort to obtain such data using conventional recordings or by human observation. Bringing out a robot to a field is our approach to solve this problem. We are developing a portable observation system called HARKBird using a robot audition HARK and microphone arrays to understand temporal patterns of vocalizations characteristics and their behavioral contexts. In this paper, we introduce a prototype system to 2D localize vocalizations of wild birds in real-time, and to classify their song types after recording. We show that the system can estimate the position of songs of a target individual and classify their songs with a reasonable quality to discuss their song - behavior relationships.
Shinji Sumitani, Reiji Suzuki, Shiho Matsubayashi, Takaya Arita, Kazuhiro Nakadai, Hiroshi G. Okuno
IROS5
2018 Multi-timescale Feature-extraction Architecture of Deep Neural Networks for Acoustic Model Training from Raw Speech Signal
abstract
This paper describes a new architecture of deep neural networks (DNNs) for acoustic models. Training DNNs from raw speech signals will provide 1) novel features of signals, 2) normalization-free processing such as utterance-wise mean subtraction, and 3) low-latency speech recognition for robot audition. Exploiting the longer context of raw speech signals seems useful in improving recognition accuracy. However, naive use of longer contexts results in the loss of short-term patterns; thus, recognition accuracy degrades. We propose a multi-timescale feature-extraction architecture of DNNs with blocks of different time scales, which enable capturing long- and short-term patterns of speech signals. Each block consists of complex-valued networks that correspond to Fourier and filterbank transformations for analysis. Experiments showed that the proposed multi-timescale architecture reduced the word error rate by about 3% compared with those only with the longterm context. Analysis of the extracted features revealed that our architecture efficiently captured the slow and fast changes of speech features.
Ryu Takeda, Kazuhiro Nakadai, Kazunori Komatani
IROS2
2018 To animate or anime-te?: Investigating sign avatar comprehensibility
abstract
In this study, we investigated the effect of avatar design on the perception of sign language animation display. Signed sentence expressions were comparatively evaluated using a natural and an anime-style avatar model of three different clothing colors and patterns to determine each design parameter's influence on comprehensibility, naturalness and user affinity. Results show that plain settled color clothing was perceived as most favorable, and that the natural avatar was clearly preferred over the anime-style avatar as an informant of the signed sentence content.
Heike Brock, Shigeaki Nishina, Kazuhiro Nakadai
IVA3
2018 Deep JSLC: A Multimodal Corpus Collection for Data-driven Generation of Japanese Sign Language Expressions
Heike Brock, Kazuhiro Nakadai
LREC2
2018 Data-driven development of Virtual Sign Language Communication Agents
abstract
Engaging deaf and hearing people in common discussions requires interfaces to help them understand each other, such as robot agents that translate spoken language into Sign Language (SL) expressions and vice-versa. However, the recognition and generation of signed sentences is a complex task of high dimensionality that cannot be solved in sufficient quality yet. Thus, it is necessary to develop new technologies of improved performances. The sequence to sequence neural network model, traditionally used for machine translation, is adapted to the above two tasks by treating a SL sequence as a multi-dimensional sentence. We defined an encoding of the SL annotations and conducted experiments on the network structure to define a most accurate translation model. This study proves the network trainable and possibly applicable in real-life with an extended dataset, which shall be tested for deployment in virtual translation assistants in the following.
Agathe Balayn, Heike Brock, Kazuhiro Nakadai
RO-MAN3
2018 Signal Restoration based on Bi-directional LSTM with Spectral Filtering for Robot Audition
abstract
This paper addresses restoration of acoustic signals for robot audition. A robot usually listens to target acoustic signals such as speech and music in noisy conditions. Acoustic information on such signals inevitably contaminated with noise. Even when noise reduction techniques such as sound source separation are performed, the noise-reduced acoustic signals contain distortion and/or residual noise after the noise reduction to some extent. The distortion and residual noise basically degrade the performance of recognition processes such as automatic speech recognition (ASR). We decided to use bidirectional long short-term memory (Bi-LSTM) for acoustic signal restoration since it can represent dynamic behaviors well for a temporal sequence in the forward and backward directions. When applying Bi-LSTM to recover acoustic signals, there is an issue, that is, acoustic signals tend to be sparse in high frequencies, and thus Bi-LSTM training becomes insufficient in such high frequencies due to a lack of training data. Therefore, we propose a new restoration method based on Bi-LSTM with spectral filtering. The spectral filter and the corresponding inverse filter are introduced to a Bi-LSTM framework to accelerate training in high frequencies. Preliminary results showed that the proposed Bi-LSTM with spectral filtering can perform signal restoration even when a small amount of training data is available.
Ryosuke Taniguchi, Kotaro Hoshiba, Katsutoshi Itoyama, Kenji Nishida, Kazuhiro Nakadai
RO-MAN5
2018 Speech Enhancement Based on Bayesian Low-Rank and Sparse Decomposition of Multichannel Magnitude Spectrograms
abstract
This paper presents a blind multichannel speech enhancement method that can deal with the time-varying layout of microphones and sound sources. Since nonnegative tensor factorization (NTF) separates a multichannel magnitude (or power) spectrogram into source spectrograms without phase information, it is robust against the time-varying mixing system. This method, however, requires prior information such as the spectral bases (templates) of each source spectrogram in advance. To solve this problem, we develop a Bayesian model called robust NTF (Bayesian RNTF) that decomposes a multichannel magnitude spectrogram into target speech and noise spectrograms based on their sparseness and low rankness. Bayesian RNTF is applied to the challenging task of speech enhancement for a microphone array distributed on a hose-shaped rescue robot. When the robot searches for victims under collapsed buildings, the layout of the microphones changes over time and some of them often fail to capture target speech. Our method robustly works under such situations, thanks to its characteristic of time-varying mixing system. Experiments using a 3-m hose-shaped rescue robot with eight microphones show that the proposed method outperforms conventional blind methods in enhancement performance by the signal-to-noise ratio of 1.03 dB.
Yoshiaki Bando, Katsutoshi Itoyama, Masashi Konyo, Satoshi Tadokoro, Kazuhiro Nakadai, Kazuyoshi Yoshii, Tatsuya Kawahara, Hiroshi G. Okuno
IEEE ACM Trans. Audio Speech Lang. Process.5
2017 A Spatial-Cue-Based Probabilistic Model for Bird Song Scene Analysis
abstract
This paper addresses bird song scene analysis focusing on location of birds and acoustic features of bird songs. Such a research area usually requires manual annotation related to positions and/or vocalization types of the target animals for a large amount of observed data. However, this manual annotation has two problems. One is that it is tough to annotate data observed in real environments because environmental noise exist and sound is reflected by trees and the ground, and also several birds at different locations may sing at the same time. The other is that it is inevitable that manual annotation produces inaccurate and inconsistent labels due to human errors and annotators' individual differences. For the first problem, we propose a Spatial-Cue-Based Probabilistic Model (SCBPM), which is a probabilistic model to estimate the maximum likelihood result for a bird song scene analysis by integrating sound source detection, localization, separation and identification based on spatial information of sound sources. For the second problem, we employ a semiautomatic annotation approach, in which a semi-supervised training method is deduced for SCBPM. This method decreases the amount of manual annotation. Preliminary experiments using recorded bird song data from the wild revealed that our system outperformed a conventional bird song scene analysis system by simply connecting sound source detection, localization, separation and identification in a cascade way in terms of identification accuracy.
Ryosuke Kojima, Osamu Sugiyama, Kotaro Hoshiba, Reiji Suzuki, Kazuhiro Nakadai
DSAA5
2017 Node Pruning Based on Entropy of Weights and Node Activity for Small-Footprint Acoustic Model Based on Deep Neural Networks
Ryu Takeda, Kazuhiro Nakadai, Kazunori Komatani
INTERSPEECH2
2017 Development of microphone-array-embedded UAV for search and rescue task
abstract
This paper addresses online outdoor sound source localization using a microphone array embedded in an unmanned aerial vehicle (UAV). In addition to sound source localization, sound source enhancement and robust communication method are also described. This system is one instance of deployment of our continuously developing open source software for robot audition called HARK (Honda Research Institute Japan Audition for Robots with Kyoto University). To improve the robustness against outdoor acoustic noise, we propose to combine two sound source localization methods based on MUSIC (multiple signal classification) to cope with trade-off between latency and noise robustness. The standard Eigenvalue decomposition based MUSIC (SEVD-MUSIC) has smaller latency but less noise robustness, whereas the incremental generalized singular value decomposition based MUSIC (iGSVD-MUSIC) has higher noise robustness but larger latency. A UAV operator can use an appropriate method according to the situation. A sound enhancement method called online robust principal component analysis (ORPCA) enables the operator to detect a target sound source more easily. To improve the stability of wireless communication, and robustness of the UAV system against weather changes, we developed data compression based on free lossless audio codec (FLAC) extended to support a 16 ch audio data stream via UDP, and developed a water-resistant microphone array. The resulting system successfully worked in an outdoor search and rescue task in ImPACT Tough Robotics Challenge in November 2016.
Kazuhiro Nakadai, Makoto Kumon, Hiroshi G. Okuno, Kotaro Hoshiba, Mizuho Wakabayashi, Kai Washizaki, Takahiro Ishiki, Daniel Gabriel, Yoshiaki Bando, Takayuki Morito, Ryosuke Kojima, Osamu Sugiyama
IROS1
2017 Acoustic model training based on node-wise weight boundary model for fast and small-footprint deep neural networks
Ryu Takeda, Kazuhiro Nakadai, Kazunori Komatani
Comput. Speech Lang.2
2016 Reduction of Computational Cost Using Two-Stage Deep Neural Network for Training for Denoising and Sound Source Identification
Takayuki Morito, Osamu Sugiyama, Satoshi Uemura, Ryosuke Kojima, Kazuhiro Nakadai
IEA/AIE5
2016 Localizing Bird Songs Using an Open Source Robot Audition System with a Microphone Array
Reiji Suzuki, Shiho Matsubayashi, Kazuhiro Nakadai, Hiroshi G. Okuno
INTERSPEECH3
2016 Semi-automatic bird song analysis by spatial-cue-based integration of sound source detection, localization, separation, and identification
abstract
This paper addresses bird song analysis based on semi-automatic annotation. Research in animal behavior, especially with birds, would be aided by automated (or semiautomated) systems that can localize sounds, measure their timing, and identify their source. This is difficult to achieve in real environments where several birds may be singing from different locations and at the same time. Analysis of recordings from the wild has in the past typically required manual annotation. Such annotation is not always accurate or even consistent, as it may vary both within or between observers. Here we propose a system that uses automated methods from robot audition, including sound source detection, localization, separation and identification. In robot audition these technologies have typically been studied separately; combining them often leads to poor performance in real-time application from the wild. We suggest that integration is aided by placing a primary focus on spatial cues, then combining other features within a Bayesian framework. A second problem has been that supervised machine learning methods typically requires a pre-trained model that may require a large training set of annotated labels. We have employed a semi-automatic annotation approach that requires much less pre-annotation. Preliminary experiments with recordings of bird songs from the wild revealed that for identification accuracy our system outperformed a method based on conventional robot audition.
Ryosuke Kojima, Osamu Sugiyama, Reiji Suzuki, Kazuhiro Nakadai, Charles E. Taylor
IROS4
2016 Partially Shared Deep Neural Network in sound source separation and identification using a UAV-embedded microphone array
abstract
This paper addresses sound source separation and identification for noise-contaminated acoustic signals recorded with a microphone array embedded in an Unmanned Aerial Vehicle (UAV), aiming at people's voice detection quickly and widely in a disaster situation. The key approach to achieve this is Deep Neural Network (DNN), but it is well known that training a DNN needs a huge dataset to improve its performance. In a practical application, building such a dataset is not often realistic owing to the cost of manual data annotation. Therefore, we propose a Partially-Shared Deep Neural Network (PS-DNN) which can learn multiple tasks at the same time with a small amount of annotated data. Preliminary results show that the PS-DNN outperforms conventional DNN-based approaches which require fully-annotated data in training in terms of identification accuracy. In addition, it maintains performance even when noise-suppressed signals are used for sound source separation training, and partially annotated data is used for sound source identification training.
Takayuki Morito, Osamu Sugiyama, Ryosuke Kojima, Kazuhiro Nakadai
IROS4
2016 Online simultaneous localization and mapping of multiple sound sources and asynchronous microphone arrays
abstract
This paper presents an online method of simultaneous localization and mapping (SLAM) for estimating the positions of multiple moving sound sources and stationary robots and synchronizing microphone arrays attached to those robots. Since each robot with a microphone array can solely estimate the directions of sound sources, the two-dimensional source positions can be estimated from the source directions estimated by multiple robots using a triangulation method. In addition, sound mixtures can be separated accurately by regarding distributed microphone arrays as one big array. To perform these tasks, some methods have been proposed for localizing and synchronizing microphone arrays. These methods, however, can be used only if a single sound source exists because the time differences of arrival (TDOAs) between microphones are assumed to be directly observed. To overcome this limitation, we propose a unified state-space model that encodes the source and robot positions and the time offsets between microphone arrays in a latent space. Given the TDOAs and directions of arrival (DOAs) estimated by separating observed mixture sounds into source sounds, the latent variables are estimated jointly in an online manner using a FastSLAM2.0 algorithm that can deal with an unknown time-varying number of moving sound sources.
Kouhei Sekiguchi, Yoshiaki Bando, Keisuke Nakamura, Kazuhiro Nakadai, Katsutoshi Itoyama, Kazuyoshi Yoshii
IROS4
2016 Robust sound source mapping using three-layered selective audio rays for mobile robots
abstract
This paper investigates sound source mapping in a real environment using a mobile robot. Our approach is based on audio ray tracing which integrates occupancy grids and sound source localization using a laser range finder and a microphone array. Previous audio ray tracing approaches rely on all observed rays and grids. As such observation errors caused by sound reflection, sound occlusion, wall occlusion, sounds at misdetected grids, etc. can significantly degrade the ability to locate sound sources in a map. A three-layered selective audio ray tracing mechanism is proposed in this work. The first layer conducts frame-based unreliable ray rejection (sensory rejection) considering sound reflection and wall occlusion. The second layer introduces triangulation and audio tracing to detect falsely detected sound sources, rejecting audio rays associated to these misdetected sounds sources (short-term rejection). A third layer is tasked with rejecting rays using the whole history (long-term rejection) to disambiguate sound occlusion. Experimental results under various situations are presented, which proves the effectiveness of our method.
Daobilige Su, Keisuke Nakamura, Kazuhiro Nakadai, Jaime Valls Miró
IROS3
2016 Construction of Japanese Audio-Visual Emotion Database and Its Application in Emotion Recognition
Nurul Lubis, Randy Gomez, Sakriani Sakti, Keisuke Nakamura, Koichiro Yoshino, Satoshi Nakamura 0001, Kazuhiro Nakadai
LREC7
2016 Leveraging phantom signals for improved voice-based human-robot interaction
abstract
Voice-based system used in human-robot interaction is susceptible to challenging environment conditions. In an enclosed environment, the speech signal is often reflected which causes smearing as it is observed in the microphone. This phenomenon creates mismatch with the acoustic model, degrading the recognition performance and the robot's ability to understand and execute commands. Moreover, phantoms increase false-alarm in robot's attention system. To address these issues, environment-matched training and model adaptation may be used. The former requires enormous amount of training data to exhaustively cover different matched conditions whereas the latter needs several adaptation data to be collected at runtime. It is important to stress that data collection and the wait time are luxuries in a robot setup. In this paper, we extend our previous work that mitigates these problem by combining environment-adaptive training, speech enhancement with phantom awareness and fast model update, respectively. As a result, we achieve a robust voice-based system that enhances the observed speech, rejects phantoms and automatically updates the model at runtime to minimize the mismatch. Results show that the proposed method significantly outperforms our previous work.
Randy Gomez, Yurii Vasylkiv, Keisuke Nakamura, Takeshi Mizumoto, Kazuhiro Nakadai
RO-MAN5
2015 Acoustic model training based on node-wise weight boundary model increasing speed of discrete neural networks
abstract
Our purpose is to realize discrete neural networks (NNs), whose some parameters are discretized, as a low-resource and fast NNs for acoustic models. Two essential problems should be tackled for its realization; 1) the reduction of discretization errors and 2) the implementation method for fast processing. We propose a new parameter training algorithm for 1) and an implementation using look-up table (LUT) on general-purpose CPUs for 2), respectively. The former can set proper boundaries of discretization at each node of NNs, resulting in the reduction of discretization error. The latter can reduce the memory usage of NNs within the cache size of CPU by encoding parameters of NNs. Experiments with 2-bit discrete NNs showed that our algorithm maintained almost the same word accuracy as 8-bit discrete NNs and achieved a 40% increase in speed of the NN's forward calculation.
Ryu Takeda, Kazunori Komatani, Kazuhiro Nakadai
ASRU3
2015 Robot audition: Its rise and perspectives
abstract
The ability of robots to listen to several things at once with their own “ears”, that is, robot audition, is an important factor in improving interaction and symbiosis between humans and robots. The critical issue in robot audition is real-time processing and robustness against noisy environments with high flexibility to support various kinds of robots and hardware configurations. This paper first overviews activities and issues related to robot audition. Then, it presents the “HARK” robot audition software, which provides three primary functions for robot audition, sound source localization, sound source separation, and separated sound recognition, and then reports their performance. Finally, it discusses future directions in new promising areas as well as robotics.
Hiroshi G. Okuno, Kazuhiro Nakadai
ICASSP2
2015 Temporal smearing compensation in reverberant environment for speech-based human-robot interaction
abstract
Speech-based human-robot interaction is often plagued with issues such as reverberation and changes in speaker position that impacts overall performance. In this paper, we show a method in compensating the joint effects of reverberation and the change in speaker position. The acoustic perturbation caused by these two takes its toll on the Automatic Speech Recognition (ASR) and then the Spoken Language Understanding (SLU). Consequently, these will lead to a failure in the human-robot interaction experience. The proposed method is specifically designed to address the challenging environment condition in which robots are deployed. First, we analyze the impact of reverberation in the form of temporal smearing per change in speaker position. Then, we extract the smearing coefficients that capture the joint dynamics between the speech signal at current position and the room acoustics as observed by the robot. These coefficients are utilized to update the room transfer function (RTF) and the suppression parameters are stored offline. Moreover, all of these processes are optimized in the context of the ASR system for robot application. In the online mode, the reverberant data at an arbitrary position is processed using the parameters pre-computed offline. This effectively compensates the joint effects of reverberation at the arbitrary speaker position. Experimental results using real data gathered in a human-robot communication setting show that the proposed method outperforms existing methods.
Randy Gomez, Keisuke Nakamura, Takeshi Mizumoto, Kazuhiro Nakadai
ICRA4
2015 On-the-spot calibration of microphone array Transfer Functions for robot audition
abstract
This paper investigates the calibration of a microphone array based robot audition system, namely calibration of microphone array Transfer Functions (TFs). There are mainly two methods to obtain TFs: geometrical calculation and measurement. The geometrical calculation has difficulty in simulating robot- and room-acoustics such as diffraction and reflection properties of a robot's body and a room, and the measurement is accurate but time-consuming and requires expertise on acoustics. Thus, we propose fast and simple on-the-spot calibration of TFs including robot- and room-acoustics. The proposed approach first estimates microphone location and clock-difference by Simultaneous Localization And Mapping (SLAM) using hand clap acoustic signals while a human is walking around a robot. Second, TFs with robot- and room-acoustics are estimated by hand clap acoustic signals and interpolated so that the TFs can be roundly arranged at regular intervals in an online manner. In the evaluation, we calibrated TFs only by 20 hand claps (took only 20 seconds), and the TFs showed considerable improvements in sound source localization and separation compared to geometrically calculated TFs and achieved comparable performance towards the measured TFs which are calibrated by approximately 60 minute recordings.
Keisuke Nakamura, Surya Ambrose, Kazuhiro Nakadai
ICRA3
2015 Scene Understanding Based on Sound and Text Information for a Cooking Support Robot
Ryosuke Kojima, Osamu Sugiyama, Kazuhiro Nakadai
IEA/AIE3
2015 Interactive Interface to Optimize Sound Source Localization with HARK
Osamu Sugiyama, Ryosuke Kojima, Kazuhiro Nakadai
IEA/AIE3
2015 Dereverberation for active human-robot communication robust to speaker's face orientation
Randy Gomez, Levko Ivanchuk, Keisuke Nakamura, Takeshi Mizumoto, Kazuhiro Nakadai
INTERSPEECH5
2015 Microphone-accelerometer based 3D posture estimation for a hose-shaped rescue robot
abstract
3D posture estimation for a hose-shaped robot is critical in rescue activities due to complex physical environments. Conventional sound-based posture estimation assumes rather flat physical environments and focuses only on 2D, resulting in poor performance in real world environments with rubble. This paper presents novel 3D posture estimation by exploiting microphones and accelerometers. The idea of our method is to compensate the lack of posture information obtained by sound-based time-difference-of arrival (TDOA) with the tilt information obtained from accelerometers. This compensation is formulated as a nonlinear state-space model and solved by the unscented Kalman filter. Experiments are conducted by using a 3m hose-shaped robot with eight units of a microphone and an accelerometer and seven units of a loudspeaker and a vibration motor deployed in a simple 3D structure. Experimental results demonstrate that our method reduces the errors of initial states to about 20 cm in the 3D space. If the initial errors of initial states are less than 20 %, our method can estimate the correct 3D posture in real-time.
Yoshiaki Bando, Katsutoshi Itoyama, Masashi Konyo, Satoshi Tadokoro, Kazuhiro Nakadai, Kazuyoshi Yoshii, Hiroshi G. Okuno
IROS5
2015 Utilizing visual cues in robot audition for sound source discrimination in speech-based human-robot communication
abstract
It is easy for human beings to discern whether an observed acoustic signal is a direct speech, reflected speech or noise through simple listening. Relying purely on acoustic cues is enough for human beings to discriminate between the different kinds of sound sources which is not straightforward for machines. A robot equipped with the current robot audition mechanism in most cases, will fail to differentiate a direct speech from the other sound sources because acoustic information alone is insufficient for effective discrimination. Robot audition is an important topic in speech-based human-robot communication. It enables the robot to associate the incoming speech signal to the user for an effective human-robot communication. In challenging environments, this task becomes difficult due to reflections of the direct speech signal and background noise sources. To counter this problem, a robot needs to have a minimum amount of prior information to discriminate the valid speech signal (direct speech) from the contaminants (i.e., speech reflections and background noise sources). Failure to do so would lead to false speech-to-speaker association in robot audition and will gravely impact human-robot communication experience. In this paper we propose to using visual cues to augment the traditional robot audition which relies solely on acoustic information. The proposed method significantly improves accuracy of speech-to-speaker association and machine understanding performance in real environment situation. Experimental results show that our expanded system is robust in discriminating direct speech from speech reflections and background noise sources.
Randy Gomez, Levko Ivanchuk, Keisuke Nakamura, Takeshi Mizumoto, Kazuhiro Nakadai
IROS5
2015 Audio-visual scene understanding utilizing text information for a cooking support robot
abstract
This paper addresses multimodal “scene understanding” for a robot using audio-visual and text information. Scene understanding is defined by extracting six-W information such as What, When, Where, Who, Why, and hoW on the surrounding environment. Although scene understanding for a robot has been studied in the fields of robot vision and audition, only the first four Ws except for why and how information were considered. We, thus, focus on extracting how information, in particular, on cooking scenes. In cooking scenes, we define how information as a cooking procedure, and it is useful that a robot gives appropriate advice for cooking. To realize such cooking support, we propose a multi-modal cooking procedure recognition framework consisting of Convolutional Neural Network (CNN), and Hierarchical Hidden Markov Model (HHMM). CNN is knows as one of the most advanced classifiers, and it is applied to recognize a cooking events from audio and visual information. HHMM models a cooking procedure represented by a sequence of cooking events, which is defined as a relationship between cooking events using text data obtained from web, and the cooking events classified with CNN. Therefore, our proposed framework integrates these three types of modalities. We constructed an interactive cooking support system based on the proposed framework, which advice a next step in the current cooking procedure through human-robot communication. Preliminary results with simulated and real recorded multi-modal scenes showed the robustness of the proposed framework in a noisy and/or occluded situation.
Ryosuke Kojima, Osamu Sugiyama, Kazuhiro Nakadai
IROS3
2015 Robot-Audition-based Human-Machine Interface for a Car
abstract
This paper describes a Robot-Audition based Car Human Machine Interface (RA-CHMI). A RA-CHMI, like a car navigation system, has difficulty dealing with voice commands, since there are many noise sources in a car, including road noise, air-conditioner, music, and passengers. Microphone array processing developed in robot audition, may overcome this problem. Robot audition techniques, including sound source localization, Voice Activity Detection (VAD), sound source separation, and barge-in-able processing, were introduced by considering the characteristics of RA-CHMI. Automatic Speech Recognition (ASR), based on a Deep Neural Network (DNN), improved recognition performance and robustness in a noisy environment. In addition, as an integrated framework, HARK-Dialog was developed to build a multi-party and multi-modal dialog system, enabling the seamless use of cloud and local services with pluggable modular architecture. The constructed multi-party and multimodal RA-CHMI system did not require a push-to-talk button, nor did it require reducing the audio volume or air-conditioner when issuing speech commands. It could also control a four-DOF robot agent to make the system's responses more understandable. The proposed RA-CHMI was validated by evaluating essential techniques in the system, such as VAD and DNN-ASR, using real speech data recorded during driving. The entire design of the RA-CHMI system, including the system response time and the proper use of cloud/local services, are also discussed.
Kazuhiro Nakadai, Takeshi Mizumoto, Keisuke Nakamura
IROS1
2015 Robot audition based Acoustic Event Identification using a Bayesian model considering spectral and temporal uncertainties
abstract
To analyze auditory scenes of robots' surrounding environments, not only speeches but also non-speech sounds are important, which are spatially distributed and have different spectral and temporal characteristics. Thus, this paper investigates Acoustic Event Identification (AEI) which includes problems of localization, detection, and identification of sound sources. To achieve AEI by a robot in a real environment, we first propose to use a robot audition framework including sound source localization and separation to localize, detect, and separate acoustic events. For the identification, we propose two Bayesian models, iterative Latent Dirichlet Allocation (it-LDA) and Nested Pitman-Yor process with Uncertainty Compensation (NPY-UC). it-LDA and NPY-UC extract noise-robust sound-units and sound-words, respectively, and they consider probabilistic spectral and temporal uncertainties to robustify AEI against harsh environments such as noise and reverberation, etc. We have implemented these proposed methods using a robot-embedded microphone array. The preliminary results showed 5–18 pts improvement compared to a conventional GMM method in noisy environments thanks to the Bayesian framework in consideration of uncertainties.
Keisuke Nakamura, Kazuhiro Nakadai
IROS2
2015 Interactive sound source localization using robot audition for tablet devices
abstract
This paper investigates localization of sound sources in a real environment using a tablet device. For the localization, we use build-in sensors on a tablet device and additionally mount a cover with a microphone array. Because of the flat shape and limited sensor performance, the localization has mainly the following three issues; 1) the flat microphone array allows only azimuth estimation but elevation estimation; 2) the measurement noise of built-in sensors degrades the direction of arrival estimation performance; 3) the direction of arrival is not sufficient information to localize sound sources in three dimensional space to explore. To solve these issues, we propose interactive sound source localization using robot audition. For 1), we propose elevation estimation by integrating sound azimuth estimation and device motion information under an active audition framework in robot audition. For 2), we introduced a constrained optimization to robustify the direction of arrival estimation. For 3), we propose an interactive interface based on augmented reality for tablet devices, which navigates the device camera view to sound sources to explore. The proposed system was evaluated both subjectively and objectively and showed validity in real cases of sound source localization.
Keisuke Nakamura, Lana Sinapayen, Kazuhiro Nakadai
IROS3
2015 A case study of an automatic volume control interface for a telepresence system
abstract
The study of the telepresence robot as a tool for telecommunication from a remote location is attracting a considerable amount of attention. However, the problem arises that a telepresence robot system does not allow the volume of the user's utterance to be adjusted precisely, because it does not consider varying conditions in the sound environment, such as noise. In addition, when talking with several people in remote location, the user would like to be able to change the speaker volume freely according to the situation. In a previous study, a telepresence robot was proposed that has a function that automatically regulates the volume of the user's utterance. However, the manner in which the user exploits this function in a practical situation needs to be investigated. We propose a telepresence conversation robot system called “TeleCoBot.” TeleCoBot includes an operator's user interface, through which the volume of the user's utterance can be automatically regulated according to the distance between the robot and the conversation partner and the noise level in the robot's environment. We conducted a case study, in which the participants played a game using TeleCoBot's interface. The results of the study reveal the manner in which the participants used TeleCoBot and the additional factors that the system requires.
Masaaki Takahashi, Masa Ogata, Michita Imai, Keisuke Nakamura, Kazuhiro Nakadai
RO-MAN5
2015 Improved sound source localization in horizontal plane for binaural robot audition
Ui-Hyun Kim, Kazuhiro Nakadai, Hiroshi G. Okuno
Appl. Intell.2
2015 Audio-visual speech recognition using deep learning
abstract
Audio-visual speech recognition (AVSR) system is thought to be one of the most promising solutions for reliable speech recognition, particularly when the audio is corrupted by noise. However, cautious selection of sensory features is crucial for attaining high recognition performance. In the machine-learning community, deep learning approaches have recently attracted increasing attention because deep neural networks can effectively extract robust latent features that enable various recognition algorithms to demonstrate revolutionary generalization capabilities under diverse application conditions. This study introduces a connectionist-hidden Markov model (HMM) system for noise-robust AVSR. First, a deep denoising autoencoder is utilized for acquiring noise-robust audio features. By preparing the training data for the network with pairs of consecutive multiple steps of deteriorated audio features and the corresponding clean features, the network is trained to output denoised audio features from the corresponding features deteriorated by noise. Second, a convolutional neural network (CNN) is utilized to extract visual features from raw mouth area images. By preparing the training data for the CNN as pairs of raw images and the corresponding phoneme label outputs, the network is trained to predict phoneme labels from the corresponding mouth area input images. Finally, a multi-stream HMM (MSHMM) is applied for integrating the acquired audio and visual HMMs independently trained with the respective features. By comparing the cases when normal and denoised mel-frequency cepstral coefficients (MFCCs) are utilized as audio features to the HMM, our unimodal isolated word recognition results demonstrate that approximately 65 % word recognition rate gain is attained with denoised MFCCs under 10 dB signal-to-noise-ratio (SNR) for the audio signal input. Moreover, our multimodal isolated word recognition results utilizing MSHMM with denoised MFCCs and acquired visual features demonstrate that an additional word recognition rate gain is attained for the SNR conditions below 10 dB.
Kuniaki Noda, Yuki Yamaguchi, Kazuhiro Nakadai, Hiroshi G. Okuno, Tetsuya Ogata
Appl. Intell.3
2014 Volume adaptation and visualization by modeling the volume level in noisy environments for telepresence system
abstract
The Lombard effect is the involuntary tendency of speakers to increase their vocal effort when speaking in a loud noise to enhance the audibility of their voice. There is a problem in telecommunication due to the Lombard effect. A speaker talks at a louder volume than necessary for the conversation partner at a remote location. This paper proposes a volume model that is required in order to automatically adjust the volume of an operator's voice at a remote communication via a telepresence robot, and develops an optimal volume control system LombaBot equipped on a telepresence robot with the model. The volume model measures the level of noise around the robot and the distance between a conversation partner and the robot to adjust the volume of the operator's voice. It has two types of volume adjustments. Those are called comfortable volume and secret talk volume. LombaBot enables people at a remote site to listen comfortably to the voice of a robot operator. Moreover, the operator is able to talk in low voices when s/he wants to talk in secret with nearby people. We confirmed that LombaBot adjusted the volume of an operator's voice properly in the noisy remote location.
Akira Hayamizu, Michita Imai, Keisuke Nakamura, Kazuhiro Nakadai
HAI4
2014 Ego-motion noise suppression for robots based on Semi-Blind Infinite Non-negative Matrix Factorization
abstract
This paper addresses ego-motion noise suppression for a robot. Many methods use motion information such as position, velocity and acceleration of each joint to infer ego-motion noise. However, such inference is not reliable since motion information and ego-motion noise are not at all times correlated. We propose a new framework for ego-motion noise suppression based on single channel processing without using any explicit motion information. In the proposed framework, ego-motion noise features are estimated in advance from an ego-motion noise input with Infinite Non-negative Matrix Factorization (INMF) which is a non-parametric Bayesian model. After that, the proposed Semi-Blind INMF(SB-INMF) is applied to an input signal consisting of both the target and egomotion noise signals. The ego-motion noise features which are obtained with INMF are used as input to the SB-INMF and treated as the fixed features to extract the target signal. Finally, the target signal is extracted using newly-estimated features with SB-INMF. The proposed framework was applied to ego-motion noise suppression on two types of humanoid robots. Experimental results showed that ego-motion noise was suppressed well compared to a conventional template-based egomotion noise suppression method using motion information, and thus it worked properly on a robot which does not have an interface to provide the robot's motion information.
Taiki Tezuka, Takami Yoshida, Kazuhiro Nakadai
ICRA3
2014 Lipreading using convolutional neural network
Kuniaki Noda, Yuki Yamaguchi, Kazuhiro Nakadai, Hiroshi G. Okuno, Tetsuya Ogata
INTERSPEECH3
2014 Speech-based human-robot interaction robust to acoustic reflections in real environment
abstract
Acoustic reflection inside an enclosed environment is detrimental to human-robot interaction. Reflection may manifest as phantom sources emanating from unknown directions. In effect, a single speaker may falsely manifest as multiple speakers to the robot audition system, impeding the robot's ability to correctly associate the speech command to the actual speaker. Moreover, speech reflection smears the original speech signal due to reverberation. This degrades speech recognition and understanding performance. Conventional robot audition schemes that rely purely on acoustics and spatial information are very sensitive to acoustic reflection which ultimately leads to the failure in human-robot interaction. We propose a method for human-robot interaction robust to the effect of acoustic reflection. First, visual information is utilized and head tracking scheme is employed to reinforce the acoustic information with the visual presence of a prospect user. Second, we employ a model-based sound event identification scheme and scrutinize whether the acoustic information is likely to be speech or non-speech. Using all the information we have gathered, we create a simple rule construct to effectively discriminate the original source (actual speaker) from phantom sources (reflection). Consequently, the corresponding source identified as phantom (reflection) is used to estimate the unwanted smearing for effective suppression via speech enhancement. Experiments are conducted in human-robot interaction setting in which the proposed method outperforms the conventional method.
Randy Gomez, Koji Inoue, Keisuke Nakamura, Takeshi Mizumoto, Kazuhiro Nakadai
IROS5
2014 Improvement in outdoor sound source detection using a quadrotor-embedded microphone array
abstract
This paper addresses sound source detection in an outdoor environment using a quadrotor with a microphone array. Since the previously reported method has a high computational cost, we proposed a sound source detection algorithm called MUltiple SIgnal Classification based on incremental Generalized Singular Value Decomposition (iGSVD-MUSIC), which detects sound source location and temporal activity with low computational cost. In addition, to relax an over-esitimation problem of noise correlation matrix which is used in iGSVD-MUSIC, we proposed Correlation Matrix Scaling (CMS), which realizes soft whitening of noise. The protptype system based on the proposed methods were evaluated with two types of microphone arrays in an outdoor environment. Experimental results showed that the combination of iGSVD-MUSIC and CMS improves sound source detection performance drastically and achieves real-time processing.
Takuma Ohata, Keisuke Nakamura, Takeshi Mizumoto, Taiki Tezuka, Kazuhiro Nakadai
IROS5
2014 Making a robot dance to diverse musical genre in noisy environments
abstract
In this paper we address the problem of musical genre recognition for a dancing robot with embedded microphones capable of distinguishing the genre of a musical piece while moving in a real-world scenario. For this purpose, we assess and compare two state-of-the-art musical genre recognition systems, based on Support Vector Machines and Markov Models, in the context of different real-world acoustic environments. In addition, we compare different preprocessing robot audition variants (single channel and separated signal from multiple channels) and test different acoustic models, learned a priori, to tackle multiple noise conditions of increasing complexity in the presence of noises of different natures (e.g., robot motion, speech). The results with six different musical genres suggest improved results, in the order of 43.6pp for the most complex conditions, when recurring to Sound Source Separation and acoustic models trained in similar conditions to the testing scenarios. A robot dance demonstration session confirms the applicability of the proposed integration for genre-adaptive dancing robots in real-world noisy environments.
João Lobato Oliveira, Keisuke Nakamura, Thibault Langlois, Fabien Gouyon, Kazuhiro Nakadai, Angelica Lim, Luís Paulo Reis, Hiroshi G. Okuno
IROS5
2014 Auditory-aware navigation for mobile robots based on reflection-robust sound source localization and visual SLAM
abstract
Autonomous robot navigation using Simultaneous localization and mapping (SLAM) is essential for scene understanding by robots. Most existing systems use visual information, and even though such visual-based technologies are robust and useful for many situations, they have difficulty dealing with certain scenarios such as occluded goal or when the goal is out of frame. Introducing audio information to the navigation system solves these issues effectively. Several audio-based methods have been developed in the past to deal with these issues. However, these existing audio-based methods have been developed with the assumption that the space around the robot is open i.e. no sound reflection occurs. Hence, the invisible goals where the sound reflection can be localized have not been fully considered. This paper proposes a reflection-robust sound source localization method using visual SLAM. This method can deal with sound sources whose direct paths are not available, and using localization, we can set a goal only for the actual sound source. Also, to correct the drift present in the local estimates of Visual Odometry (VO), SLAM was integrated with the system, thus increasing the accuracy and robustness of mapping and navigation of our proposed method. The performance of the proposed system is compared to conventional methods, and it proves to be very efficient and robust especially in extremely reverberant situations. It was found that the integration of VO and SLAM improved the average error of a map by approximately 50 pts, and the accuracy of SSL for direct path of sounds was improved by approximately 8 pts. With the online implementation of these methods we successfully achieved audio visual navigation for the actual sound sources.
Gautam Narang, Keisuke Nakamura, Kazuhiro Nakadai
SMC3
2014 Sound annotation tool for multidirectional sounds based on spatial information extracted by HARK robot audition software
abstract
With the rise of inexpensive microphone array products and the robot audition software called HARK, we can record and analyze multidirectional sound sources easily. The combination of microphone array and the software enables us to separate, localize, and track multidirectional sound sources. Most of the solutions for accessing these separated sound source information provide clients for interpreting simplified information about the separated sources, but not to directly execute the semantic annotations. Since the multidirectional sound annotation requires simultaneous labeling of separated sound sources and a multidirectional overview of the sources, it is essential to have an efficient way of annotation and an intuitive view of multidirectional sounds. Our proposed sound annotation tool provides drag & drop operation of annotation with a 3D sound source view and also provides annotation autocompletion with a SVM trained with the user's annotation history. The proposed features enable users to do the annotation task intuitively and confirm its result. We also conducted an evaluation demonstrating the efficiency of annotation done using the tool.
Osamu Sugiyama, Katsutoshi Itoyama, Kazuhiro Nakadai, Hiroshi G. Okuno
SMC3
2013 Robustness to speaker position in distant-talking automatic speech recognition
abstract
In this paper, we show a method that significantly improved our previous work in single-channel dereverberation. The proposed method is more robust to changes in speaker position in distant talking ASR. First, we update the room transfer function (RTF) and weighting parameters for dereverberation to the target speaker position. This scheme corrects speech power variation as a function of position in the waveform level. Consequently, its impact to the acoustic model is verified. Then, we implement a fast acoustic model update reflective of the speech power level of the target speaker position. Furthermore, the scheme in updating the model is simple and precludes time-consuming model re-estimation. As a result, the proposed method can be executed online. The synergy of these corrective measures significantly minimizes the mismatch between training and testing conditions. We test our method using real reverberant data with different locations inside the room. Experimental results show that the proposed method outperforms the conventional methods in terms of ASR performance. Moreover, our fast acoustic model update scheme is at par in terms of recognition performance against time-consuming model re-estimation.
Randy Gomez, Keisuke Nakamura, Kazuhiro Nakadai
ICASSP3
2013 Hands-free human-robot communication robust to speaker's radial position
abstract
In this paper we present a method in room transfer function (RTF) estimation, employed specifically for dereverberation in hands-free human-robot communication.We introduce a radial distance compensation scheme which significantly improved the RTF estimate robust to the speech power variation due to changes in speaker's radial position. The proposed method is implemented in two levels; first, waveform-level compensation is executed to reflect the change in power caused by the change of radial position to the RTF. We generated possible RTF estimates within a close neighbourhood based on curve fitting. Then, we select among these estimates the optimal RTF based on acoustic model likelihood criterion, the same criterion employed in automatic speech recognition (ASR) systems. The latter is referred to as acoustic model-level compensation, which links the generated RTF to the ASR. We note that in ASR application, both waveform and acoustic models play an important role in achieving optimal performance. Thus, the synergistic effect of the two processes guarantee ASR performance improvement when used in conjunction with our ASR-based dereverberation scheme. Experimental evaluation show robustness in recognition performance when used in hands-free human-robot communication environment.
Randy Gomez, Keisuke Nakamura, Kazuhiro Nakadai, Ui-Hyun Kim, Hiroshi G. Okuno, Tatsuya Kawahara
ICRA3
2013 Improved Sound Source Localization and Front-Back Disambiguation for Humanoid Robots with Two Ears
Ui-Hyun Kim, Kazuhiro Nakadai, Hiroshi G. Okuno
IEA/AIE2
2013 Posture estimation of hose-shaped robot using microphone array localization
abstract
This paper presents a posture estimation of hose-shaped robot using microphone array localization. The hose-shaped robots, one of major rescue robots, have problems with navigation because their posture is too flexible for a remote operator to control to go as far as desired. For navigational and mission usability, the posture estimation of the hose-shaped robot is essential. We developed a posture estimation method with a microphone array and small loudspeakers equipped on the hose-shaped robot. Our method consists of two steps: (1) playing a known sound from the loudspeaker one-by-one, and (2) estimating the microphone positions on the hose-shaped robot instead of estimating the posture directly. We designed a time difference of arrival (TDOA) estimation method to be robust against directional noise and implemented a prototype system using a posture model of the hose-shaped robot and an Extended Kalman Filter (EKF). The validity of our approach is evaluated by the experiments with both signals recorded in an anechoic chamber and simulated data.
Yoshiaki Bando, Takeshi Mizumoto, Katsutoshi Itoyama, Kazuhiro Nakadai, Hiroshi G. Okuno
IROS4
2013 Noise correlation matrix estimation for improving sound source localization by multirotor UAV
abstract
A method has been developed for improving sound source localization (SSL) using a microphone array from an unmanned aerial vehicle with multiple rotors, a “multirotor UAV”. One of the main problems in SSL from a multirotor UAV is that the ego noise of the rotors on the UAV interferes with the audio observation and degrades the SSL performance. We employ a generalized eigenvalue decomposition-based multiple signal classification (GEVD-MUSIC) algorithm to reduce the effect of ego noise. While GEVD-MUSIC algorithm requires a noise correlation matrix corresponding to the auto-correlation of the multichannel observation of the rotor noise, the noise correlation is nonstationary due to the aerodynamic control of the UAV. Therefore, we need an adaptive estimation method of the noise correlation matrix for a robust SSL using GEVD-MUSIC algorithm. Our method uses a Gaussian process regression to estimate the noise correlation matrix in each time period from the measurements of self-monitoring sensors attached to the UAV such as the pitch-roll-yaw tilt angles, xyz speeds, and motor control values. Experiments compare our method with existing SSL methods in terms of precision and recall rates of SSL. The results demonstrate that our method outperforms existing methods, especially under high signal-to-noise-ratio conditions.
Koutarou Furukawa, Keita Okutani, Kohei Nagira, Takuma Otsuka, Katsutoshi Itoyama, Kazuhiro Nakadai, Hiroshi G. Okuno
IROS6
2013 Dereverberation robust to speaker's azimuthal orientation in multi-channel human-robot communication
abstract
The acoustical dynamics of reverberation in an enclosed environment poses a problem to human-robot communication. Any change in the azimuthal orientation of the speaker contributes to unpredictable acoustical activity resulting in a degradation in the performance of the automatic speech recognition (ASR) system. Thus, dereverberation techniques need to address this issue prior to ASR. Dereverberation in multi-channel applications primarily evolves in the adoption of a suitable reverberant model that results to a computationally feasible solution and at the same time yields an accurate estimate of the harmful reflections (i.e., late reflection) for effective suppression. In this paper we address this problem by introducing a hybrid method based on multi-channel processing on a singlechannel reverberant model platform. The proposed method is capable of accurate signal estimation, a property inherent to a multi-channel system, and at the same time bears the computational efficiency derived from single-channel reverberant model approach. The proposed method is summarized as follows; First, multi-channel sound-source processing is employed to obtain the full reverberant and the late reflection signal estimates. Then, equalization is employed to update the late reflection estimate reflective of the change in azimuth prior to dereverberation. The equalization parameters for azimuthal change are obtained through an offline optimization procedure. Experimental evaluation in an actual human-robot communication environment shows that the proposed method outperforms existing methods in terms of robustness in the ASR performance.
Randy Gomez, Keisuke Nakamura, Kazuhiro Nakadai
IROS3
2013 Real-time super-resolution three-dimensional sound source localization for robots
abstract
This paper investigates Sound Source Localization (SSL) for a robot in a real world. Previously, we focused on one-dimensional SSL for azimuth and assumed that target sources are distributed close to a horizontal plane. Without this assumption, the SSL performance is drastically degraded. Thus, three-dimensional SSL is essential to improve the localization for sound sources distributed in a three-dimensional space. Compared to one-dimensional SSL, three-dimensional SSL mainly has the following problems: 1) a massive number of Transfer Function (TF) measurements for microphone array calibration are required for three dimensions to maintain the spatial resolution of SSL sufficiently-high, 2) the computational cost for searching for sound sources drastically increases in high-dimensional spaces. For the first issue, we extend the previously-proposed one-dimensional TF interpolation method, integrating time-domain-based and frequency-domain-based interpolation, to three dimensions. The interpolation achieves three-dimensional super-resolution SSL and reduction of the number of TF measurements while maintaining the spatial resolution of SSL. For the second issue, we propose optimal hierarchical SSL, which reduces computational cost for searching for sound sources by introducing a hierarchical search algorithm instead of using greedy search in localization. We previously proposed the concept of the algorithm. This paper additionally discusses theoretical optimality in hierarchization to minimize the total computational cost of SSL. The method determines the number of hierarchies and the resolution of each hierarchy depending on desired spatial resolution. These techniques are integrated into an SSL system using a robot. The experimental result showed: 1) the proposed interpolation method achieved super-resolution SSL working better than that with pre-measured TFs, 2) the optimal hierarchical SSL drastically reduced computational cost by approximately 97%.
Keisuke Nakamura, Randy Gomez, Kazuhiro Nakadai
IROS3
2013 Sound Source Localization Using Joint Bayesian Estimation With a Hierarchical Noise Model
abstract
The performance of sound source localization is often reduced by the presence of colored noise in the environment, such as room reverberation. In this study, a method for estimating the noise spatial covariance using a hierarchical model is proposed and its performance is evaluated. By employing the hierarchical model in joint Bayesian estimation, robust estimation of the covariance is expected with a relatively small amount of data. Moreover, a method of jointly estimating the number of sources is introduced so that it can be used for cases in which the number of active sources dynamically changes, for example, speech signals. The results of the experiments performed using actual room reverberation show the effectiveness of the proposed method.
Futoshi Asano, Hideki Asoh, Kazuhiro Nakadai
IEEE Trans. Speech Audio Process.3
2012 Multi-party human-robot interaction with distant-talking speech recognition
abstract
Speech is one of the most natural medium for human communication, which makes it vital to human-robot interaction. In real environments where robots are deployed, distant-talking speech recognition is difficult to realize due to the effects of reverberation. This leads to the degradation of speech recognition and understanding, and hinders a seamless human-robot interaction. To minimize this problem, traditional speech enhancement techniques optimized for human perception are adopted to achieve robustness in human-robot interaction. However, human and machine perceive speech differently: an improvement in speech recognition performance may not automatically translate to an improvement in human-robot interaction experience (as perceived by the users). In this paper, we propose a method in optimizing speech enhancement techniques specifically to improve automatic speech recognition (ASR) with emphasis on the human-robot interaction experience. Experimental results using real reverberant data in a multi-party conversation, show that the proposed method improved human-robot interaction experience in severe reverberant conditions compared to the traditional techniques.
Randy Gomez, Tatsuya Kawahara, Keisuke Nakamura, Kazuhiro Nakadai
HRI4
2012 Sound source localization in spatially colored noise using a hierarchical Bayesian model
abstract
In this paper, source localization in spatially colored noise is addressed. The covariance of colored noise is estimated using a hierarchical model in the joint Bayesian estimation. The results of the experiment show that the spatial resolution was improved compared with the approach without hierarchical modeling.
Futoshi Asano, Hideki Asoh, Kazuhiro Nakadai
ICASSP3
2012 Online audio beat tracking for a dancing robot in the presence of ego-motion noise in a real environment
abstract
This paper presents the design and implementation of a real-time real-world beat tracking system which runs on a dancing robot. The main problem of such a robot is that, while it is moving, ego noise is generated due to its motors, and this directly degrades the quality of the audio signal features used for beat tracking. Therefore, we propose to incorporate ego noise reduction as a pre-processing stage prior to our tempo induction and beat tracking system. The beat tracking algorithm is based on an online strategy of competing agents sequentially processing a continuous musical input, while considering parallel hypotheses regarding tempo and beats. This system is applied to a humanoid robot processing the audio from its embedded microphones on-the-fly, while performing simplistic dancing motions. A detailed and multi-criteria based evaluation of the system across different music genres and varying stationary/non-stationary noise conditions is presented. It shows improved performance and noise robustness, outperforming our conventional beat tracker (i.e., without ego noise suppression) by 15.2 points in tempo estimation and 15.0 points in beat-times prediction.
João Lobato Oliveira, Gökhan Ince, Keisuke Nakamura, Kazuhiro Nakadai
ICRA4
2012 Online learning for template-based multi-channel ego noise estimation
abstract
This paper presents a system that gives a robot the ability to diminish its own disturbing noise (i.e., ego noise) by utilizing template-based ego noise estimation, an algorithm previously developed by the authors. In pursuit of an autonomous, online and adaptive template learning system in this work, we specifically focus on eliminating the requirement of an offline training session performed in advance to build the essential templates, which represent the ego noise. The idea of discriminating ego noise from all other sound sources in the environment enables the robot to learn the templates online without requiring any prior information. Based on the directionality/diffuseness of the sound sources, the robot can easily decide whether the template should be discarded because it is corrupted by external noises, or it should be inserted into the database because the template consists of pure ego noise only. Furthermore, we aim to update the template database optimally by introducing an additional time-variant forgetting factor parameter, which provides a balance between adaptivity and stability of the learning process automatically. Moreover, we enhanced the single-channel noise estimation system to be compatible with the multi-channel robot audition framework so that ego noise can be eliminated from all signals stemming from multiple sound sources respectively. We demonstrate that the proposed system allows the robot to have the ability of online template learning as well as a high performance of noise estimation and suppression for multiple sound sources.
Gökhan Ince, Kazuhiro Nakadai, Keisuke Nakamura
IROS2
2012 Real-time super-resolution Sound Source Localization for robots
abstract
Sound Source Localization (SSL) is an essential function for robot audition and yields the location and number of sound sources, which are utilized for post-processes such as sound source separation. SSL for a robot in a real environment mainly requires noise-robustness, high resolution and real-time processing. A technique using microphone array processing, that is, Multiple Signal Classification based on Standard EigenValue Decomposition (SEVD-MUSIC) is commonly used for localization. We improved its robustness against noise with high power by incorporating Generalized EigenValue Decomposition (GEVD). However, GEVD-based MUSIC (GEVD-MUSIC) has mainly two issues: 1) the resolution of pre-measured Transfer Functions (TFs) determines the resolution of SSL, 2) its computational cost is expensive for real-time processing. For the first issue, we propose a TF interpolation method integrating time-domain-based and frequency-domain-based interpolation. The interpolation achieves super-resolution SSL, whose resolution is higher than that of the pre-measured TFs. For the second issue, we propose two methods, MUSIC based on Generalized Singular Value Decomposition (GSVD-MUSIC), and Hierarchical SSL (H-SSL). GSVD-MUSIC drastically reduces the computational cost while maintaining noise-robustness in localization. H-SSL also reduces the computational cost by introducing a hierarchical search algorithm instead of using greedy search in localization. These techniques are integrated into an SSL system using a robot embedded microphone array. The experimental result showed: the proposed interpolation achieved approximately 1 degree resolution although we have only TFs at 30 degree intervals, GSVD-MUSIC attained 46.4% and 40.6% of the computational cost compared to SEVD-MUSIC and GEVD-MUSIC, respectively, H-SSL reduces 59.2% computational cost in localization of a single sound source.
Keisuke Nakamura, Kazuhiro Nakadai, Gökhan Ince
IROS2
2012 Outdoor auditory scene analysis using a moving microphone array embedded in a quadrocopter
abstract
This paper addresses auditory scene analysis, especially, sound source localization using an aerial vehicle with a microphone array in an outdoor environment. Since such a vehicle is able to search sound sources quickly and widely, it is useful to detect outdoor sound sources, for instance, to find distressed people in a disaster situation. In such an environment, noise is quite loud and dynamically-changing, and conventional microphone array techniques studied in the field of indoor robot audition are of less use. We, thus, proposed MUltiple SIgnal Classification based on incremental Generalized EigenValue Decomposition (iGEVD-MUSIC). It can deal with high power noise by introducing a noise correlation matrix and GEVD even when the signal-to-noise ratio is less than 0 dB. In addition, the noise correlation matrix is incrementally estimated to adapt to dynamic changes in noise. We developed a prototype system for auditory scene analysis based on the proposed method using the Parrot AR.Drone with a microphone array and a Kinect device. Experimental results using the prototype system showed that dynamically-changing noise is properly suppressed with the proposed method and multiple human voice sources are able to be localized even when the AR.Drone is moving in an outdoor environment.
Keita Okutani, Takami Yoshida, Keisuke Nakamura, Kazuhiro Nakadai
IROS4
2012 Live assessment of beat tracking for robot audition
abstract
In this paper we propose the integration of an online audio beat tracking system into the general framework of robot audition, to enable its application in musically-interactive robotic scenarios. To this purpose, we introduced a staterecovery mechanism into our beat tracking algorithm, for handling continuous musical stimuli, and applied different multi-channel preprocessing algorithms (e.g., beamforming, ego noise suppression) to enhance noisy auditory signals lively captured in a real environment. We assessed and compared the robustness of our audio beat tracker through a set of experimental setups, under different live acoustic conditions of incremental complexity. These included the presence of continuous musical stimuli, built of a set of concatenated musical pieces; the presence of noises of different natures (e.g., robot motion, speech); and the simultaneous processing of different audio sources on-the-fly, for music and speech. We successfully tackled all these challenging acoustic conditions and improved the beat tracking accuracy and reaction time to music transitions while simultaneously achieving robust automatic speech recognition.
João Lobato Oliveira, Gökhan Ince, Keisuke Nakamura, Kazuhiro Nakadai, Hiroshi G. Okuno, Luís Paulo Reis, Fabien Gouyon
IROS4
2012 An active audition framework for auditory-driven HRI: Application to interactive robot dancing
abstract
In this paper we propose a general active audition framework for auditory-driven Human-Robot Interaction (HRI). The proposed framework simultaneously processes speech and music on-the-fly, integrates perceptual models for robot audition, and supports verbal and non-verbal interactive communication by means of (pro)active behaviors. To ensure a reliable interaction, on top of the framework a behavior decision mechanism based on active audition policies the robot's actions according to the reliability of the acoustic signals for auditory processing. To validate the framework's application to general auditory-driven HRI, we propose the implementation of an interactive robot dancing system. This system integrates three preprocessing robot audition modules: sound source localization, sound source separation, and ego noise suppression; two modules for auditory perception: live audio beat tracking and automatic speech recognition; and multi-modal behaviors for verbal and non-verbal interaction: music-driven dancing and speech-driven dialoguing. To fully assess the system, we set up experimental and interactive real-world scenarios with highly dynamic acoustic conditions, and defined a set of evaluation criteria. The experimental tests revealed accurate and robust beat tracking and speech recognition, and convincing dance beat-synchrony. The interactive sessions confirmed the fundamental role of the behavior decision mechanism for actively maintaining a robust and natural human-robot interaction.
João Lobato Oliveira, Gökhan Ince, Keisuke Nakamura, Kazuhiro Nakadai, Hiroshi G. Okuno, Luís Paulo Reis, Fabien Gouyon
RO-MAN4
2012 Efficient Blind Dereverberation and Echo Cancellation Based on Independent Component Analysis for Actual Acoustic Signals
abstract
This letter presents a new algorithm for blind dereverberation and echo cancellation based on independent component analysis (ICA) for actual acoustic signals. We focus on frequency domain ICA (FD-ICA) because its computational cost and speed of learning convergence are sufficiently reasonable for practical applications such as hands-free speech recognition. In applying conventional FD-ICA as a preprocessing of automatic speech recognition in noisy environments, one of the most critical problems is how to cope with reverberations. To extract a clean signal from the reverberant observation, we model the separation process in the short-time Fourier transform domain and apply the multiple input/output inverse-filtering theorem (MINT) to the FD-ICA separation model. A naive implementation of this method is computationally expensive, because its time complexity is the second order of reverberation time. Therefore, the main issue in dereverberation is to reduce the high computational cost of ICA. In this letter, we reduce the computational complexity to the linear order of the reverberation time by using two techniques: (1) a separation model based on the independence of delayed observed signals with MINT and (2) spatial sphering for preprocessing. Experiments show that the computational cost grows in proportion to the linear order of the reverberation time and that our method improves the word correctness of automatic speech recognition by 10 to 20 points in a RT₂₀= 670 ms reverberant environment.
Ryu Takeda, Kazuhiro Nakadai, Toru Takahashi 0001, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno
Neural Comput.2
2011 Rhythmic reference of a human while a rope turning task
abstract
This paper addresses the rhythmic reference in physical human-robot interaction. Human refers to a rhythm from multiple sensing modalities when turning a rope with another human synchronously. This study verifies a hypothesis that some humans mix several rhythms of the modalities into a rhythm (rhythmic reference). Six participants, four males and two females, 21-23 years old, took part in eight experiments which examined the hypothesis. In each experiment, we masked the perception of each participant using eight combination of three kinds of masks, an eye-mask, headphones, and a force mask. Each participant interacted with an operator that turned a rope with a constant frequency. As a result of the experiments, a participant increased the controlling error as the number of masks was increased regardless the types of masked modalities. The result strongly supported our hypothesis.
Kenta Yonekura, Chyon Hae Kim, Kazuhiro Nakadai, Hiroshi Tsujino, Shigeki Sugano
HRI3
2011 Correlation matrix interpolation in Sound Source Localization for a robot
abstract
In microphone array processing, a Correlation Matrix (CM) between multiple channel input signals is widely utilized for various purposes such as Sound Source Localization (SSL), etc. The CM corresponds to spatial information between a microphone and a sound source, and it dynamically changes as they move in a dynamic environment. Since pre-measured CMs have difficulties in dealing with dynamically changing environments due to their discreteness, this paper addresses a CM interpolation utilizing a few discrete known CMs. The proposed method is based on an integration of eigen-value-scaling and linear interpolation, which requires a few measurements and low computational costs. Apart from conventional methods, this paper deals with an interpolation based on a correlation-matrix not based on a transfer-function, which achieves the integration approach. The evaluation shows better estimation performance compared to conventional transfer-function-based methods. We also applied the interpolation to SSL and confirmed its validity.
Keisuke Nakamura, Kazuhiro Nakadai, Hirofumi Nakajima, Gökhan Ince
ICASSP2
2011 Assessment of general applicability of ego noise estimation
abstract
Noise generated due to the motion of a robot deteriorates the quality of the desired sounds recorded by robot-embedded microphones. On top of that, a moving robot is also vulnerable to its loud fan noise that changes its orientation relative to the moving limbs where the microphones are mounted on. To tackle the non-stationary ego-motion noise and the direction changes of fan noise, we propose an estimation method based on instantaneous prediction of ego noise using parameterized templates. We verify the ego noise suppression capability of the proposed estimation method on a humanoid robot by evaluating it on two important applications in the framework of robot audition: (1) automatic speech recognition and (2) sound source localization. We demonstrate that our method improves recognition and localization performance during both head and arm motions considerably.
Gökhan Ince, Keisuke Nakamura, Futoshi Asano, Hirofumi Nakajima, Kazuhiro Nakadai
ICRA5
2011 Design and implementation of selectable sound separation on the Texai telepresence system using HARK
abstract
This paper presents the design and implementation of selectable sound separation functions on the telepresence system "Texai" using the robot audition software "HARK." An operator of Texai can "walk" around a faraway office to attend a meeting or talk with people through video-conference instead of meeting in person. With a normal microphone, the operator has difficulty recognizing the auditory scene of the Texai, e.g., he/she cannot know the number and the locations of sounds. To solve this problem, we design selectable sound separation functions with 8 microphones in two modes, overview and filter modes, and implement them using HARK's sound source localization and separation. The overview mode visualizes the direction-of-arrival of surrounding sounds, while the filter mode provides sounds that originate from the range of directions he/she specifies. The functions enable the operator to be aware of a sound even if it comes from behind the Texai, and to concentrate on a particular sound. The design and implementation was completed in five days due to the portability of HARK. Experimental evaluations with actual and simulated data show that the resulting system localizes sound sources with a tolerance of 5 degrees.
Takeshi Mizumoto, Kazuhiro Nakadai, Takami Yoshida, Ryu Takeda, Takuma Otsuka, Toru Takahashi 0001, Hiroshi G. Okuno
ICRA2
2011 Robust Intonation Pattern Classification in Human Robot Interaction
abstract
We present a system for the classification of intonation patterns in human robot interaction. The system distinguishes questions from other types of utterances and can deal with additional re-verberations, background noise, as well as music interfering with the speech signal. The main building blocks of our system are a multi channel source separation, robust fundamental fre-quency extraction and tracking, segmentation of the speech sig-nal, and classification of the fundamental frequency pattern of the last speech segment. We evaluate the system with Japanese sentences which are ambiguous without intonation information in a realistic human robot interaction scenario. Despite the chal-lenging task our system is able to classify the intonation pattern with good accuracy. With several experiments we evaluate the contribution of the different aspects of our system. Index Terms: Intonation pattern classification, human robot in-teraction, source separation, pitch tracking
Martin Heckmann, Kazuhiro Nakadai, Hirofumi Nakajima
INTERSPEECH2
2011 Bayesian Extension of MUSIC for Sound Source Localization and Tracking
abstract
This paper presents a Bayesian extension of MUSIC-based sound source localization (SSL) and tracking method. SSL is important for distant speech enhancement and simultaneous speech separation for improving speech recognition, as well as for auditory scene analysis by mobile robots. One of the drawbacks of existing SSL methods is the necessity of careful parameter tunings, e.g., the sound source detection threshold depending on the reverberation time and the number of sources. Our contribution consists of (1) automatic parameter estimation in the variational Bayesian framework and (2) tracking of sound sources with reliability. Experimental results demonstrate our method robustly tracks multiple sound sources in a reverberant environment with RT20 = 840 (ms). Index Terms: simultaneous sound source localization, MUSIC algorithm, variational Bayes, particle filter
Takuma Otsuka, Kazuhiro Nakadai, Tetsuya Ogata, Hiroshi G. Okuno
INTERSPEECH2
2011 HARK based real-time single pane 3D auditory scene visualizer empowered by Speech Arrow
abstract
Robot audition, which aims at multiple simultaneous human-robot communications in noisy environments, has undergone rapid development since last decade. For a common scenario, robot audition systems usually generate a huge amount of chronological data. These data fall into several categories: sound source locations and orientations, separated sound sources and recognized speeches. In order to efficiently assist people in understanding and analyzing the scenario, it is necessary to organize the data in a delicate way. To achieve this goal, we suggest Speech Arrow which conveys recognized speeches in animated arrow shaped frames. Speech Arrow is highly integrated, intuitive and vivid. In this paper, a single pane auditory scene visualizer is implemented to support Speech Arrow. The visualizer is 3D real-time virtual reality system based on the open source robot audition system HARK. It includes two versions, one of which restores original auditory scenes with offline raw data. The other analyzes and displays online instant data without any delay to help hearing impaired people. We demonstrate that our visualizer improves auditory awareness considerably with experiments.
Kazuhiro Nakadai, Hirofumi Nakajima, Ichiro Hagiwara
IROS2
2011 Assessment of single-channel ego noise estimation methods
abstract
While a robot is moving, ego noise is generated due to the fans and motors of the robot. Furthermore, a robot is not only subject to the ego noise, but also to the ambient noise of the environment, both having different short-term signal characteristics. Because ego-motion noise generated by the motors is non-stationary, and the BackGround Noise (BGN) is stationary, one single noise estimation method is unable to track the changes in both noise spectra rapidly and accurately. Therefore, we propose to use the combination of two different noise estimation methods adequate for each one of co-existing noise types in a unified framework: 1) a stationary noise estimation method called Histogram-based Recursive Level Estimation (HRLE) and 2) a non-stationary noise estimation method called Template-based Estimation (TE). In this paper, we evaluate the performance of several single-channel based noise estimation techniques in terms of their prediction accuracy and quality of the speech signals enhanced by spectral subtraction methods. The experimental results show that our system, compared to the conventional single-stage noise estimation methods, achieves better performance in attaining signal quality and improving word correct rates.
Gökhan Ince, Kazuhiro Nakadai, Tobias Rodemann, Jun-ichi Imura, Keisuke Nakamura, Hirofumi Nakajima
IROS2
2011 Incremental learning for ego noise estimation of a robot
abstract
Using pre-recorded templates to estimate and suppress the ego noise of a robot is advantageous because this method is able to cope with the non-stationarity of this particular type of noise. However, standard template-based estimation requires human intervention in the offline training sessions, storage of large amounts of data and does not adapt to the dynamical changes in the environmental conditions. In this paper we investigate the feasibility of an incremental template learning system to tackle these drawbacks. Incremental learning enables the system to acquire new templates on the fly and update the older ones appropriately. Whilst allowing the system to continually increase its knowledge and enhancing its estimation performance, this learning scheme also reduces the size of the database. We evaluate the performance of the proposed noise estimation method in terms of its estimation accuracy, quality of speech signals enhanced by spectral subtraction method, and size of database. The experimental results show that our system compared to conventional single-channel noise estimation methods achieves better performance in attaining signal quality and improving word correct rates.
Gökhan Ince, Kazuhiro Nakadai, Tobias Rodemann, Jun-ichi Imura, Keisuke Nakamura, Hirofumi Nakajima
IROS2
2011 SLAM-based online calibration of asynchronous microphone array for robot audition
abstract
This paper addresses the online calibration of an asynchronous microphone array for robots. Conventional microphone array technologies require a lot of measurements of transfer functions to calibrate microphone locations, and a multi-channel A/D converter for inter-microphone synchronization. We solve these two problems using a framework combining Simultaneous Localization and Mapping (SLAM) and beamforming in an online manner. To do this, we assume that estimations of microphone locations, a sound source location, and microphone clock difference correspond to mapping, self-localization, observation errors in SLAM, respectively. In our framework, the SLAM process calibrates locations and clock differences of microphones every time a microphone array observes a sound like a human's clapping, and a beamforming process works as a cost function to decide the convergence of calibration by localizing the sound with the estimated locations and clock differences. After calibration, beamforming is used for sound source localization. We implemented a prototype system using Extended Kalman Filter (EKF) based SLAM and Delay-and-Sum Beamforming (DS-BF). The experimental results showed that microphone locations and clock differences were estimated properly with 10–15 sound events (handclaps), and the error of sound source localization with the estimated information was less than the grid size of beamforming, that is, the lowest error was theoretically attained.
Hiroaki Miura, Takami Yoshida, Keisuke Nakamura, Kazuhiro Nakadai
IROS4
2011 Intelligent sound source localization and its application to multimodal human tracking
abstract
We have assessed robust tracking of humans based on intelligent Sound Source Localization (SSL) for a robot in a real environment. SSL is fundamental for robot audition, but has three issues in a real environment: robustness against noise with high power, lack of a general framework for selective listening to sound sources, and tracking of inactive and/or noisy sound sources. To address the first issue, we extended Multiple SIgnal Classification by incorporating Generalized EigenValue Decomposition (GEVD-MUSIC) so that it can deal with high power noise and can select target sound sources. To address the second issue, we proposed Sound Source Identification (SSI) based on hierarchical gaussian mixture models and integrated it with GEVD-MUSIC to realize a selective listening function. To address the third issue, we integrated audio-visual human tracking using particle filtering. Integration of these three techniques into an intelligent human tracking system showed: 1) GEVD-MUSIC improved the noise-robustness of SSL by a signal-to-noise ratio of 5-6 dB; 2) SSI performed more than 70% in F-measure even in a noisy environment; and 3) audio-visual integration improved the average tracking error by approximately 50%.
Keisuke Nakamura, Kazuhiro Nakadai, Futoshi Asano, Gökhan Ince
IROS2
2011 Ego noise cancellation of a robot using missing feature masks
Gökhan Ince, Kazuhiro Nakadai, Tobias Rodemann, Hiroshi Tsujino, Jun-ichi Imura
Appl. Intell.2
2011 A multi-expert model for dialogue and behavior control of conversational robots and agents
Mikio Nakano, Yuji Hasegawa, Kotaro Funakoshi, Johane Takeuchi, Toyotaka Torii, Kazuhiro Nakadai, Naoyuki Kanda, Kazunori Komatani, Hiroshi G. Okuno, Hiroshi Tsujino
Knowl. Based Syst.6
2010 Design and Implementation of Two-level Synchronization for Interactive Music Robot
abstract
Our goal is to develop an interactive music robot, i.e., a robot that presents a musical expression together with humans. A music interaction requires two important functions: synchronization with the music and musical expression, such as singing and dancing. Many instrument-performing robots are only capable of the latter function, they may have difficulty in playing live with human performers. The synchronization function is critical for the interaction. We classify synchronization and musical expression into two levels: (1) the rhythm level and (2) the melody level. Two issues in achieving two-layer synchronization and musical expression are: (1) simultaneous estimation of the rhythm structure and the current part of the music and (2) derivation of the estimation confidence to switch behavior between the rhythm level and the melody level. This paper presents a score following algorithm, incremental audio to score alignment, that conforms to the two-level synchronization design using a particle filter. Our method estimates the score position for the melody level and the tempo for the rhythm level. The reliability of the score position estimation is extracted from the probability distribution of the score position. Experiments are carried out using polyphonic jazz songs. The results confirm that our method switches levels in accordance with the difficulty of the score estimation. When the tempo of the music is less than 120 (beats per minute; bpm), the estimated score positions are accurate and reported; when the tempo is over 120 (bpm), the system tends to report only the tempo to suppress the error in the reported score position predictions.
Takuma Otsuka, Kazuhiro Nakadai, Toru Takahashi 0001, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno
AAAI2
2010 A hybrid framework for ego noise cancellation of a robot
abstract
Noise generated due to the motion of a robot is not desired, because it deteriorates the quality and intelligibility of the sounds recorded by robot-embedded microphones. It must be reduced or cancelled to achieve automatic speech recognition with a high performance. In this work, we divide ego-motion noise problem into three subdomains of arm, leg and head motion noise, depending on their complexity and intensity levels. We investigate methods that make use of single-channel and multi-channel processing in order to suppress ego noise separately. For this purpose, a framework consisting of a microphone-array-based geometric source separation, a consequent post filtering process and a parallel module for template subtraction is used. Furthermore, a control mechanism is proposed, which is based on signal-to-noise ratio and instantaneously detected motions, to switch to the most suitable method to deal with the current type of noise. We evaluate the proposed techniques on a humanoid robot using automatic speech recognition (ASR). The preliminary results of isolated word recognition show the effectiveness of our methods by increasing the word correct rates up to 50% compared to the single channel recognition in arm and leg motion noises and up to 25% in very strong head motion noises.
Gökhan Ince, Kazuhiro Nakadai, Tobias Rodemann, Yuji Hasegawa, Hiroshi Tsujino, Jun-ichi Imura
ICRA2
2010 Improvement in listening capability for humanoid robot HRP-2
abstract
This paper describes improvement of sound source separation for a simultaneous automatic speech recognition (ASR) system of a humanoid robot. A recognition error in the system is caused by a separation error and interferences of other sources. In separability, an original geometric source separation (GSS) is improved. Our GSS uses a measured robot's head related transfer function (HRTF) to estimate a separation matrix. As an original GSS uses a simulated HRTF calculated based on a distance between microphone and sound source, there is a large mismatch between the simulated and the measured transfer functions. The mismatch causes a severe degradation of recognition performance. Faster convergence speed of separation matrix reduces separation error. Our approach gives a nearer initial separation matrix based on a measured transfer function from an optimal separation matrix than a simulated one. As a result, we expect that our GSS improves the convergence speed. Our GSS is also able to handle an adaptive step-size parameter. These new features are added into open source robot audition software (OSS) called "HARK" which is newly updated as version 1.0.0. The HARK has been installed on a HRP-2 humanoid with an 8-element microphone array. The listening capability of HRP-2 is evaluated by recognizing a target speech signal which is separated from a simultaneous speech signal by three talkers. The word correct rate (WCR) of ASR improves by 5 points under normal acoustic environments and by 10 points under noisy environments. Experimental results show that HARK 1.0.0 improves the robustness against noises.
Toru Takahashi 0001, Kazuhiro Nakadai, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno
ICRA2
2010 Upper-limit evaluation of robot audition based on ICA-BSS in multi-source, barge-in and highly reverberant conditions
abstract
This paper presents the upper-limit evaluation of robot audition based on ICA-BSS in multi-source, barge-in and highly reverberant conditions. The goal is that the robot can automatically distinguish a target speech from its own speech and other sound sources in a reverberant environment. We focus on the multi-channel semi-blind ICA (MCSB-ICA), which is one of the sound source separation methods with a microphone array, to achieve such an audition system because it can separate sound source signals including reverberations with few assumptions on environments. The evaluation of MCSB-ICA has been limited to robot's speech separation and reverberation separation. In this paper, we evaluate MCSB-ICA extensively by applying it to multi-source separation problems under common reverberant environments. Experimental results prove that MCSB-ICA outperforms conventional ICA by 30 points in automatic speech recognition performance.
Ryu Takeda, Kazuhiro Nakadai, Toru Takahashi 0001, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno
ICRA2
2010 Robust Ego Noise Suppression of a Robot
Gökhan Ince, Kazuhiro Nakadai, Tobias Rodemann, Hiroshi Tsujino, Jun-ichi Imura
IEA/AIE (1)2
2010 Music-Ensemble Robot That Is Capable of Playing the Theremin While Listening to the Accompanied Music
Takuma Otsuka, Takeshi Mizumoto, Kazuhiro Nakadai, Toru Takahashi 0001, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno
IEA/AIE (1)3
2010 An Improvement in Audio-Visual Voice Activity Detection for Automatic Speech Recognition
Takami Yoshida, Kazuhiro Nakadai, Hiroshi G. Okuno
IEA/AIE (1)2
2010 Applying geometric source separation for improved pitch extraction in human-robot interaction
Martin Heckmann, Claudius Gläser, Frank Joublin, Kazuhiro Nakadai
INTERSPEECH4
2010 A robust speech recognition system against the ego noise of a robot
Gökhan Ince, Kazuhiro Nakadai, Tobias Rodemann, Hiroshi Tsujino, Jun-ichi Imura
INTERSPEECH2
2010 Two-layered audio-visual integration in voice activity detection and automatic speech recognition for robots
Takami Yoshida, Kazuhiro Nakadai
INTERSPEECH2
2010 PROT - An embodied agent for intelligible and user-friendly human-robot interaction
abstract
A system has been developed that can project an embodied agent's image and sound anywhere in a room. It can thus overcome the problems inherent to other embodied agents in dealing with 2D on-screen information and 3D physical information simultaneously. An experiment demonstrated that this `PROT' agent can effectively present both on-screen information and real-world physical information. Because the PROT agent combines the advantages of a robot agent with those of an on-screen agent, it should improve human-robot interaction.
Ryota Fujimura, Kazuhiro Nakadai, Michita Imai, Ren Ohmura
IROS2
2010 Pitch extraction in Human-Robot interaction
abstract
We present a system for real-time fundamental frequency, i. e. pitch, extraction on a humanoid robot. The system extracts pitch using an 8 channel microphone array mounted on the Honda humanoid robot in a realistic Human-Robot interaction scenario. The main building blocks of the system are a multi-channel signal enhancement followed by robust pitch extraction and tracking. The signal enhancement is based on 8 channel Geometric Source Separation. For the pitch extraction the signal is first transformed with a Gammatone filter bank into the frequency domain. Next a histogram of zero crossing distances is calculated from all filter bank signals. During the calculation of the histogram spurious side peaks resulting from harmonics and sub-harmonics of the true fundamental frequency are inhibited. The resulting histogram then serves as input to a grid based Bayesian tracker which deploys Bayesian filtering in a forward step and Bayesian smoothing in a backward step on a 100ms time window. We demonstrate the performance of the system in a scenario where male and female speakers utter different phrases while standing at a normal interaction distance to the robot. For the evaluation we compare the pitch tracking results once obtained from a clean headset signal and once from the signals obtained from the robot. The results show that the tracking performance only degrades to a small extent in the realistic interaction scenario compared to the headset recordings.
Martin Heckmann, Frank Joublin, Kazuhiro Nakadai
IROS3
2010 Multi-talker speech recognition under ego-motion noise using Missing Feature Theory
abstract
This paper presents a system that gives a mobile robot the ability to recognize target speaker's speech, even if the robot performs an action and there are multiple speakers talking in the room. Associated problems to this system are twofold: (1) While the robot is moving, the joints inevitably generate ego-motion noise due to its motors. (2) Recognizing target speech against other interfering speech signals is a difficult task. Since typical solutions to (1) and (2), motor noise suppression and sound source separation, both introduce distortion to the processed signals, the performance of automatic speech recognition (ASR) deteriorates. Instead of removing the ego-motion noise with conventional noise suppression methods, in this work, we investigate methods to eliminate the unreliable parts of the audio features that are contaminated by the ego-motion noise. For this purpose, we model masks that filter unreliable speech features based on the ratio of speech and motor noise energies. We analyze the performance of the proposed technique under various test conditions by comparing it to the performance of existing Missing Feature Theory-based ASR implementations. Finally, we propose an integration framework for two different masks that are designed to eliminate ego noise and to filter the leakage energy of interfering sound sources. We demonstrate that the proposed methods achieve a high ASR accuracy.
Gökhan Ince, Kazuhiro Nakadai, Tobias Rodemann, Hiroshi Tsujino, Jun-ichi Imura
IROS2
2010 Human-robot ensemble between robot thereminist and human percussionist using coupled oscillator model
abstract
This paper presents a novel synchronizing method for a human-robot ensemble using coupled oscillators. We define an ensemble as a synchronized performance produced through interactions between independent players. To attain better synchronized performance, the robot should predict the human's behavior to reduce the difference between the human's and robot's onset timings. Existing studies in such synchronization only adapts to onset intervals, thus, need a considerable time to synchronize. We use a coupled oscillator model to predict the human's behavior. Experimental results show that our method reduces the average of onset time errors; when we use a metronome, a tempo-varying metronome or a human drummer, errors are reduced by 38%, 10% or 14% on the average, respectively. These results mean that the prediction of human's behaviors is effective for the synchronized performance.
Takeshi Mizumoto, Takuma Otsuka, Kazuhiro Nakadai, Toru Takahashi 0001, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno
IROS3
2010 Sound source separation and automatic speech recognition for moving sources
abstract
This paper addresses sound source separation and speech recognition for moving sound sources. Real-world applications such as robots should cope with both moving and stationary sound sources. However, most studies assume only stationary sound sources. We introduce three key techniques to cope with moving sources, that is, Adaptive Step-size control (AS), Optima Controlled Recursive Average (OCRA), and Separation Parameter Switching (SPS). We implemented a real-time robot audition system with these techniques for our humanoid robot with an 8ch microphone array by using HARK which is our open-source software for robot audition. Preliminary results show that the performance of recognition of moving sound sources improved drastically, and also the performance of the system is shown through two speech dialog scenarios which requires sound source separation and automatic speech recognition for moving sources.
Kazuhiro Nakadai, Hirofumi Nakajima, Gökhan Ince, Yuji Hasegawa
IROS1
2010 An easily-configurable robot audition system using Histogram-based Recursive Level Estimation
abstract
This paper presents an easily-configurable robot audition system using the Histogram-based Recursive Level Estimation (HRLE) method. In order to achieve natural human-robot interaction, a robot should recognize human speeches even if there are some noises and reverberations. Since the precision of automatic speech recognizers (ASR) have been degraded by such interference, many systems applying speech enhancement processes have been reported. However, performance of most reported systems suffer from acoustical environmental changes. For example, an enhancement process optimized for steady-state noise, such as fan noise, yields low performance when the process is used for non-steady-state noises, such as background music. The primary reason is mismatches of parameters because the appropriate parameters change according to the acoustical environments. To solve this problem, we propose a robot audition system that optimizes parameters adaptively and automatically. Our system applies and non-linear enhancement sub-processes. For the linear sub-process, we used Geometric Source Separation with the Adaptive Step-size method (GSS-AS). This adjusts the parameters adaptively and does not have any manual parameters. For the non-linear sub-process, we applied a spectral subtraction-based enhancement method with the HRLE method that is newly introduced in this paper. Since HRLE controls the threshold level parameter implicitly based on the statistical characteristics of noise and speech levels, our system has high robustness against acoustical environmental changes. For robot audition systems, all processes should be performed in real-time. We also propose implementation techniques to make HRLE run in real-time and show the effectiveness. We evaluate performance of our system and compare it to conventional systems based on the Minima Controlled Recursive Average (MCRA) method and Minimum Mean Square Error (MMSE) method. The experimental results show that our system achieves better performance than the conventional systems.
Hirofumi Nakajima, Gökhan Ince, Kazuhiro Nakadai, Yuji Hasegawa
IROS3
2010 An improvement in automatic speech recognition using soft missing feature masks for robot audition
abstract
We describe integration of preprocessing and automatic speech recognition based on Missing-Feature-Theory (MFT) to recognize a highly interfered speech signal, such as the signal in a narrow angle between a desired and interfered speakers. As a speech signal separated from a mixture of speech signals includes the leakage from other speech signals, recognition performance of the separated speech degrades. An important problem is estimating the leakage in time-frequency components. Once the leakage is estimated, we can generate missing feature masks (MFM) automatically by using our method. A new weighted sigmoid function is introduced for our MFM generation method. An experiment shows that a word correct rate improves from 66 % to 74 % by using our MFM generation method tuned by a search base approach in the parameter space.
Toru Takahashi 0001, Kazuhiro Nakadai, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno
IROS2
2010 Speedup and performance improvement of ICA-based robot audition by parallel and resampling-based block-wise processing
abstract
This paper describes a speedup and performance improvement of multi-channel semi-blind ICA (MCSB-ICA) with parallel and resampling-based block-wise processing. MCSB-ICA is an integrated method of sound source separation that accomplishes blind source separation, blind dereverberation, and echo cancellation. This method enables robots to separate user's speech signals from observed signals including the robot's own speech, other speech and their reverberations without a priori information. The main problem when MCSB-ICA is applied to robot audition is its high computational cost. We tackle this by multi-threading programming, and the two main issues are 1) the design of parallel processing and 2) incremental implementation. These are solved by a) multiple-stack-based parallel implementation, and b) resampling-based overlaps and block-wise separation. The experimental results proved that our method reduced the real-time factor to less than 0.5 with an eight-core CPU, and it improves the performance of automatic speech recognition by 2-10 points compared with the single-stack-based parallel implementation without the resampling technique.
Ryu Takeda, Kazuhiro Nakadai, Toru Takahashi 0001, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno
IROS2
2010 Two-layered audio-visual speech recognition for robots in noisy environments
abstract
Audio-visual (AV) integration is one of the key ideas to improve perception in noisy real-world environments. This paper describes automatic speech recognition (ASR) to improve human-robot interaction based on AV integration. We developed AV-integrated ASR, which has two AV integration layers, that is, voice activity detection (VAD) and ASR. However, the system has three difficulties: 1) VAD and ASR have been separately studied although these processes are mutually dependent, 2) VAD and ASR assumed that high resolution images are available although this assumption never holds in the real world, and 3) an optimal weight between audio and visual stream was fixed while their reliabilities change according to environmental changes. To solve these problems, we propose a new VAD algorithm taking ASR characteristics into account, and a linear-regression-based optimal weight estimation method. We evaluate the algorithm for auditory-and/or visually-contaminated data. Preliminary results show that the robustness of VAD improved even when the resolution of the images is low, and the AVSR using estimated stream weight shows the effectiveness of AV integration.
Takami Yoshida, Kazuhiro Nakadai, Hiroshi G. Okuno
IROS2
2010 Blind Source Separation With Parameter-Free Adaptive Step-Size Method for Robot Audition
abstract
This paper proposes an adaptive step-size method for blind source separation (BSS) suitable for robot audition systems. The design of the step-size parameter is a critical consideration when we apply BSS to real-world applications such as robot audition systems, because the surrounding environment dynamically changes in the real world. It is common to use a fixed step-size parameter that was obtained empirically. However, because of environmental changes and noise, the performance of BSS with a fixed step-size parameter deteriorates and the separation matrix sometimes diverges. Several adaptive step-size methods for BSS have been proposed. However, there are difficulties when applying them to robot audition systems for example, low-computational cost requirements, being free from manual parameter adjustment and so on. We propose an adaptive step-size method suitable for robot audition systems. The proposed method has the following merits: 1) low computational cost; 2) no parameters to be adjusted manually; and 3) no additional preprocessing requirements. We applied our method to six different BSS algorithms for an eight-channel microphone array embedded in Honda's ASIMO robot. The method improved the performance of all six algorithms in experiments on separation and recognition of simultaneous speech. Moreover, the method increased the amount of calculation by less than 10% compared with the original calculation used in most BSS algorithms.
Hirofumi Nakajima, Kazuhiro Nakadai, Yuji Hasegawa, Hiroshi Tsujino
IEEE Trans. Speech Audio Process.2
2009 Sound source separation of moving speakers for robot audition
abstract
This paper addresses sound source separation and speech recognition for moving sound sources. Real-world applications such as robots should cope with both moving and stationary sound sources. However, most studies assume only stationary sound sources. We introduce two key techniques to cope with moving sources, that is, Adaptive Step-size control (AS) and Optima Controlled Recursive Average (OCRA) to improve blind source separation. We implemented a real-time robot audition system with these techniques for our humanoid robot ASIMO with an 8ch microphone array by using HARK which is our open-source software for robot audition. The performance of the system will be shown through sound source separation for moving sources and automatic speech recognition of separated speeches.
Kazuhiro Nakadai, Hirofumi Nakajima, Yuji Hasegawa, Hiroshi Tsujino
ICASSP1
2009 ICA-based efficient blind dereverberation and echo cancellation method for barge-in-able robot audition
abstract
This paper describes a new method that allows ldquoBarge-Inrdquo in various environments for robot audition. ldquoBarge-inrdquo means that a user begins to speak simultaneously while a robot is speaking. To achieve the function, we must deal with problems on blind dereverberation and echo cancellation at the same time. We adopt Independent Component Analysis (ICA) because it essentially provides a natural framework for these two problems. To deal with reverberation, we apply a Multiple Input/Output INverse-filtering Theorem-based model of observation to the frequency domain ICA. The main problem is its high-computational cost of ICA. We reduce the computational complexity to the linear order of reverberation time by using two techniques: 1) a separation modelbased on observed signal independence, and 2) enforced spatial sphering for preprocessing. The experimental results revealed that our method improved word correctness of reverberant speech by 10-20 points.
Ryu Takeda, Kazuhiro Nakadai, Toru Takahashi 0001, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno
ICASSP2
2009 Ego noise suppression of a robot using template subtraction
abstract
While a robot is moving, the joints inevitably generate noise due to its motors, i.e. ego-motion noise. This problem is very crucial, especially in humanoid robots, because it tends to have a lot of joints and the motors are located closer to the microphones than the sound sources. In this work, we investigate methods for the prediction and suppression of the ego-motion noise. In the first part, we analyze the performance of different noise subtraction strategies, assuming that the noise prediction problem has been solved. In the second part, we present some results for a noise prediction scheme based on the current robot joint status. Performance is evaluated for a number of criteria, including Automatic Speech Recognition (ASR). We demonstrate that our method improves recognition performance during ego-motion considerably.
Gökhan Ince, Kazuhiro Nakadai, Tobias Rodemann, Yuji Hasegawa, Hiroshi Tsujino, Jun-ichi Imura
IROS2
2009 Real-time sound source orientation estimation using a 96 channel microphone array
abstract
This paper proposes real-time sound source orientation estimation based on orientation-extended amplitude beamforming (OE-ABF). To recognize a sound source orientation (such as face orientation) is an important function for a robot who can achieve natural human-robot interaction because the function is required to distinguish the human target from a robot or another person. We developed a sound source orientation system using orientation-extended beamforming (OE-BF) and showed the system worked properly at least under a specific controlled environment. However, in practical use, this system does not work properly because the system doesn't take into account the differences between the supposed model in OE-BF and in practical situations. For example, the system model supposes that there is neither noise nor reverberation, however, this is not a realistic assumption. To solve this assumption mismatch problem, we propose sound source orientation estimation based on OE-ABF, and constructed a real-time sound source orientation estimation system with the proposed method using a 96ch microphone array. Evaluation results of our proposed system show that the average error of estimated angles is lower than 5°, while the error of our previously reported system was greater than 20°. With this system, the robot is able to distinguish that the utterance target of a person standing 1m in front is itself or another person standing 0.2m to the left of the robot. This is valuable for human-robot interaction.
Hirofumi Nakajima, Keiko Kikuchi, Touru Daigo, Yutaka Kaneda, Kazuhiro Nakadai, Yuji Hasegawa
IROS5
2009 Intelligent sound source localization for dynamic environments
abstract
As robotic technology plays an increasing role in human lives, ¿robot audition¿, human-robot communication, is of great interest, and robot audition needs to be robust and adaptable for dynamic environments. This paper addresses sound source localization working in dynamic environments for robots. Previously, noise robustness and dynamic localized sound selection have been enormous issues for practical use. To correct the issues, a new localization system ¿Selective Attention System¿ is proposed. The system has four new functions: localization with Generalized EigenValue Decomposition of correlation matrices for noise robustness(¿Localization with GEVD¿), sound source cancellation and focus (¿Target Source Selection¿), human-like dynamic Focus of Attention (¿Dynamic FoA¿), and correlation matrix estimation for robotic head rotation (¿Correlation Matrix Estimation¿). All are achieved by the dynamic design of correlation matrices. The system is implemented into a humanoid robot, and the experimental validation is successfully verified even when the robot microphones move dynamically.
Keisuke Nakamura, Kazuhiro Nakadai, Futoshi Asano, Yuji Hasegawa, Hiroshi Tsujino
IROS2
2009 Incremental polyphonic audio to score alignment using beat tracking for singer robots
abstract
We aim at developing a singer robot capable of listening to music with its own ¿ears¿ and interacting with a human's musical performance. Such a singer robot requires at least three functions: listening to the music, understanding what position in the music is being performed, and generating a singing voice. In this paper, we focus on the second function, that is, the capability to align an audio signal to its musical score represented symbolically. Issues underlying the score alignment problem are: (1) diversity in the sounds of various musical instruments, (2) difference between the audio signal and the musical score, (3) fluctuation in tempo of the musical performance. Our solutions to these issues are as follows: (1) the design of features based on a chroma vector in the 12-tone model and onset of the sound, (2) defining the rareness for each tone based on the idea that scarcely used tone is salient in the audio signal, and (3) the use of a switching Kalman filter for robust tempo estimation. The experimental result shows that our score alignment method improves the average of cumulative absolute errors in score alignment by 29% using 100 popular music tunes compared to the beat tracking without score alignment.
Takuma Otsuka, Toru Takahashi 0001, Hiroshi G. Okuno, Kazunori Komatani, Tetsuya Ogata, Kazumasa Murata, Kazuhiro Nakadai
IROS7
2009 Missing-feature-theory-based robust simultaneous speech recognition system with non-clean speech acoustic model
abstract
A humanoid robot must recognize a target speech signal while people around the robot chat with them in real-world. To recognize the target speech signal, robot has to separate the target speech signal among other speech signals and recognize the separated speech signal. As separated signal includes distortion, automatic speech recognition (ASR) performance degrades. To avoid the degradation, we trained an acoustic model from non-clean speech signals to adapt acoustic feature of distorted signal and adding white noise to separated speech signal before extracting acoustic feature. The issues are (1) To determine optimal noise level to add the training speech signals, and (2) To determine optimal noise level to add the separated signal. In this paper, we investigate how much noises should be added to clean speech data for training and how speech recognition performance improves for different positions of three talkers with soft masking. Experimental results show that the best performance is obtained by adding white noises of 30 dB. The ASR with the acoustic model outperforms with ASR with the clean acoustic model by 4 points.
Toru Takahashi 0001, Kazuhiro Nakadai, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno
IROS2
2009 Step-size parameter adaptation of multi-channel semi-blind ICA with piecewise linear model for barge-in-able robot audition
abstract
This paper describes a step-size parameter adaptation technique of multi-channel semi-blind independent component analysis (MCSB-ICA) for a ¿barge-in-able¿ robot audition system. By ¿barge-in¿, we mean that the user can speak simultaneously when the robot is speaking.We focused on MCSB-ICA to achieve such an audition system because it can separate a user's and a robot's speech under reverberant environments. The problem with MCSB-ICA for robot audition is the slow speed of convergence in estimating a separation filter due to its step-size parameters. Many optimization methods cannot be adopted because their computational costs are proportional to the 2nd order of the reverberation time. Our method yields adaptive step-size parameters with MCSB-ICA at low computational costs. It is based on three techniques; (1) recursive expression of the separation process, (2) a piecewise linear model of the step-size of the separation filter, and (3) adaptive step-size parameters with a sub-ICA-filter. Experimental results show that our approach attains faster convergence speed and lower computational costs than those with a fixed step-size parameter.
Ryu Takeda, Kazuhiro Nakadai, Toru Takahashi 0001, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno
IROS2
2009 Robot Audition: Missing Feature Theory Approach and Active Audition
Hiroshi G. Okuno, Kazuhiro Nakadai, Hyun-Don Kim
ISRR2
2008 Adaptive step-size parameter control for real-world blind source separation
abstract
This paper describes a method to adaptively control a step-size parameter which is used for updating a separation matrix to extract a target sound source accurately in blind source separation (BSS). The design of the step-size parameter is essential when we apply BSS to real-world applications such as robot audition systems, because the surrounding environment dynamically changes in the real world. It is common to use a fixed step-size parameter that is obtained empirically. However, due to environmental changes and noises, the performance of BSS with the fixed step-size parameter deteriorates and the separation matrix sometimes diverges. We propose a general method that allows adaptive step-size control. The proposed method is an extension of Newton’s method utilizing a complex gradient theory and is applicable to any BSS algorithm. Actually, we applied it to six types of BSS algorithms for an 8 ch microphone array embedded in Honda ASIMO. Experimental results show that the proposed method improves the performance of these six BSS algorithms through experiments of separation and recognition for two simultaneous speeches.
Hirofumi Nakajima, Kazuhiro Nakadai, Yuji Hasegawa, Hiroshi Tsujino
ICASSP2
2008 A robot referee for rock-paper-scissors sound games
abstract
This paper describes a robot referee for “rockpaper-scissors (RPS)” sound games; the robot decides the winner from a combination of rock, paper and scissors uttered by two or three people simultaneously without using any visual information. In this referee task, the robot has to cope with speech with low signal-to-noise ratio (SNR) due to a mixture of speeches, robot motor noises, and ambient noises. Our robot referee system, thus, consists of two subsystems - a real-time robot audition subsystem and a dialog subsystem focusing on RPS sound games. The robot audition subsystem can recognize simultaneous speeches by exploiting two key ideas; preprocessing consisting of sound source localization and separation with a microphone array, and system integration based on missing feature theory (MFT). Preprocessing improves the SNR of a target sound signal using geometric source separation with a multi-channel post-filter. MFT uses only reliable acoustic features in speech recognition and masks out unreliable parts caused by interfering sounds and preprocessing. MFT thus provides smooth integration between preprocessing and automatic speech recognition. The dialog subsystem is implemented as a system-initiative dialog system for multiple players based on deterministic finite automata. It first waits for a trigger command to start an RPS sound game, controls the dialog with players in the game, and finally decides the winner of the game. The referee system is constructed for Honda ASIMO with an 8-ch microphone array. In the case with two players, we attained a 70% task completion rate for the games on average.
Kazuhiro Nakadai, Shun'ichi Yamamoto, Hiroshi G. Okuno, Hirofumi Nakajima, Yuji Hasegawa, Hiroshi Tsujino
ICRA1
2008 Soft missing-feature mask generation for simultaneous speech recognition system in robots
Toru Takahashi 0001, Shun'ichi Yamamoto, Kazuhiro Nakadai, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno
INTERSPEECH3
2008 A robot uses its own microphone to synchronize its steps to musical beats while scatting and singing
abstract
Musical beat tracking is one of the effective technologies for human-robot interaction such as musical sessions. Since such interaction should be performed in various environments in a natural way, musical beat tracking for a robot should cope with noise sources such as environmental noise, its own motor noises, and self voices, by using its own microphone. This paper addresses a musical beat tracking robot which can step, scat and sing according to musical beats by using its own microphone. To realize such a robot, we propose a robust beat tracking method by introducing two key techniques, that is, spectro-temporal pattern matching and echo cancellation. The former realizes robust tempo estimation with a shorter window length, thus, it can quickly adapt to tempo changes. The latter is effective to cancel self noises such as stepping, scatting, and singing. We implemented the proposed beat tracking method for Honda ASIMO. Experimental results showed ten times faster adaptation to tempo changes and high robustness in beat tracking for stepping, scatting and singing noises. We also demonstrated the robot times its steps while scatting or singing to musical beats.
Kazumasa Murata, Kazuhiro Nakadai, Kazuyoshi Yoshii, Ryu Takeda, Toyotaka Torii, Hiroshi G. Okuno, Yuji Hasegawa, Hiroshi Tsujino
IROS2
2008 High performance sound source separation adaptable to environmental changes for robot audition
abstract
This paper describes a novel sound source separation method for a robot that needs to cope with dynamically changing noises in the real world. The sound source separation method, Geometric Source Separation (GSS), is promising because it has high separation performance and requires low computational cost. One of the most important factors in GSS performance is a step-size parameter to update a separation matrix which is generally used for extracting a target sound source. A fixed value that was obtained empirically is commonly used as the step-size parameter. However, in the real world, the surrounding environment changes dynamically. Thus, conventional GSS with a fixed step-size parameter sometimes results in poor separation results, or divergence of the separation matrix. Another important factor is the weight parameter, which adjusts the balance between geometric errors and separation errors and also affects performance. If this parameter is set to a small value, GSS becomes similar to a Blind Source Separation method, by which the output signal may contain errors based on indefinite source amplitudes and orders. In contrast, if this parameter is set to a large value, GSS becomes similar to a delay-and-sum Beamforming method, which does not have high separation performance. GSS gives good performance when the parameters are tuned to an optimum value, which changes according to the environment. We propose two effective methods that can be used for general BSS’s. One is an adaptive step-size parameter control method. By using this method, the step-size and the weight parameters are automatically set to optimum values and are able to adapt to environmental changes. The other is an optima controlled recursive average method for correlation matrix estimation. This method can improve the estimation of a separation matrix, and achieve high separation performance. We evaluated the proposed GSS algorithm with an 8ch microphone array embedded in Honda ASIMO. Experimental results showed that the proposed method improved sound source separation even in dynamically changing environments.
Hirofumi Nakajima, Kazuhiro Nakadai, Yuji Hasegawa, Hiroshi Tsujino
IROS2
2008 Barge-in-able robot audition based on ICA and missing feature theory under semi-blind situation
abstract
This paper describes a robot audition system that allows the user to barge-in; that is, the user can speak simultaneously when the robot is speaking. Our ldquobarge-in-ablerdquo system consists of two stages: (1) cancellation of robot speech and (2) recognition of the separated user speech under the ldquosemi-blind situationrdquo. The semi-blind situation is where a robotpsilas speech signal is known but a userpsilas speech signal is not. The first stage is achieved by using an adaptive filter based on time-frequency domain Independent Component Analysis, because that can separate robot speech more robustly against noise than conventional echo cancellers. To improve performance in online processing, we utilized known source normalization and the exponentially weighted stepsize method. The second stage is achieved by automatic speech recognition (ASR) based on the missing feature theory which provides robust recognition by exploiting the reliability of speech features distorted due to noise and/or separation. The semi-blind situation simplifies the estimation of such reliabilities. Experiments demonstrated that our system improved word correctness of ASR by 10.0%.
Ryu Takeda, Kazuhiro Nakadai, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno
IROS2
2007 Design and implementation of a robot audition system for automatic speech recognition of simultaneous speech
abstract
This paper addresses robot audition that can cope with speech that has a low signal-to-noise ratio (SNR) in real time by using robot-embedded microphones. To cope with such a noise, we exploited two key ideas; Preprocessing consisting of sound source localization and separation with a microphone array, and system integration based on missing feature theory (MFT). Preprocessing improves the SNR of a target sound signal using geometric source separation with multichannel post-filter. MFT uses only reliable acoustic features in speech recognition and masks unreliable parts caused by errors in preprocessing. MFT thus provides smooth integration between preprocessing and automatic speech recognition. A real-time robot audition system based on these two key ideas is constructed for Honda ASIMO and Humanoid SIG2 with 8-ch microphone arrays. The paper also reports the improvement of ASR performance by using two and three simultaneous speech signals.
Shun'ichi Yamamoto, Kazuhiro Nakadai, Mikio Nakano, Hiroshi Tsujino, Jean-Marc Valin, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno
ASRU2
2007 The Design of Phoneme Grouping for Coarse Phoneme Recognition
Kazuhiro Nakadai, Ryota Sumiya, Mikio Nakano, Koichi Ichige, Yasuo Hirose, Hiroshi Tsujino
IEA/AIE1
2007 Coarse speech recognition by audio-visual integration based on missing feature theory
abstract
Audio-visual speech recognition (AVSR) is a promising approach to improve noise robustness of speech recognition in the real world. A phoneme and a viseme are used as an auditory and visual unit for AVSR, respectively. However, in the real world, they are often misclassified due to additional input noises. To solve this problem, we propose two approaches. One is audio-visual integration based on missing feature theory to cope with missing or unreliable audio and visual features for recognition. The other is a biologically-inspired approach, that is, phoneme and viseme grouping based on coarse-to-fine recognition. Preliminary experiments show that audio-visual speech recognition based on these approaches improves the noise robustness of AVSR drastically.
Tomoaki Koiwa, Kazuhiro Nakadai, Jun-ichi Imura
IROS2
2007 Exploiting known sound source signals to improve ICA-based robot audition in speech separation and recognition
abstract
This paper describes a new semi-blind source separation (semi-BSS) technique with independent component analysis (ICA) for enhancing a target source of interest and for suppressing other known interference sources. The semi BSS technique is necessary for double-talk free robot audition systems in order to utilize known sound source signals such as self speech, music, or TV-sound, through a line-in or ubiquitous network. Unlike the conventional semi-BSS with ICA, we use the time-frequency domain convolution model to describe the reflection of the sound and a new mixing process of sounds for ICA. In other words, we consider that reflected sounds during some delay time are different from the original. ICA then separates the reflections as other interference sources. The model enables us to eliminate the frame size limitations of the frequency-domain ICA, and ICA can separate the known sources under a highly reverberative environment. Experimental results show that our method outperformed the conventional semi-BSS using ICA under simulated normal and highly reverberative environments.
Ryu Takeda, Kazuhiro Nakadai, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno
IROS2
2007 A biped robot that keeps steps in time with musical beats while listening to music with its own ears
abstract
We aim at enabling a biped robot to interact with humans through real-world music in daily-life environments, e.g., to autonomously keep its steps (stamps) in time with musical beats. To achieve this, the robot should be able to robustly predict the beat times in real time while listening to musical performance with its own ears (head-embedded microphones). However, this has not previously been addressed in most studies on music-synchronized robots due to the difficulty in predicting the beat times in real-world music. To solve this problem, we implemented a beat-tracking method developed in the field of music information processing. The predicted beat times are then used by a feedback-control method that adjusts the robot's step intervals to synchronize its steps in time with the beats. The experimental results show that the robot can adjust its steps in time with the beat times as the tempo changes. The resulting robot needed about 25 [s] to recognize the tempo change after it and then synchronize its steps.
Kazuyoshi Yoshii, Kazuhiro Nakadai, Toyotaka Torii, Yuji Hasegawa, Hiroshi Tsujino, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno
IROS2
2007 Robust Recognition of Simultaneous Speech by a Mobile Robot
abstract
This paper describes a system that gives a mobile robot the ability to perform automatic speech recognition with simultaneous speakers. A microphone array is used along with a real-time implementation of geometric source separation (GSS) and a postfilter that gives a further reduction of interference from other sources. The postfllter is also used to estimate the reliability of spectral features and compute a missing feature mask. The mask is used in a missing feature theory-based speech recognition system to recognize the speech from simultaneous Japanese speakers in the context of a humanoid robot. Recognition rates are presented for three simultaneous speakers located at 2 m from the robot. The system was evaluated on a 200-word vocabulary at different azimuths between sources, ranging from 10deg to 90deg. Compared to the use of the microphone array source separation alone, we demonstrate an average reduction in relative recognition error rate of 24% with the postfllter and of 42% when the missing features approach is combined with the postfllter. We demonstrate the effectiveness of our multisource microphone array postfilter and the improvement it provides when used in conjunction with the missing features theory.
Jean-Marc Valin, Seiichi Yamamoto, Jean Rouat, François Michaud, Kazuhiro Nakadai, Hiroshi G. Okuno
IEEE Trans. Robotics5
2006 Robust Tracking of Multiple Sound Sources by Spatial Integration of Room And Robot Microphone Arrays
abstract
Sound source tracking is an important function for a robot operating in a daily environment, because the robot should recognize where a sound event such as speech, music and other environmental sounds originates from. This paper addresses sound source tracking by integrating a room and a robot microphone array. The room microphone array consists of 64 microphones attached to the walls. It provides 2D (x-y) sound source localization based on a weighted delay-and-sum beamforming method. The robot microphone array consists of eight microphones installed on a robot head, and localizes multiple sound sources in azimuth. The localization results are integrated to track sound sources by using a particle filter for multiple sound sources. The experimental results show that particle filter based integration reduces localization errors and provides accurate and robust 2D sound source tracking.
Kazuhiro Nakadai, Hirofumi Nakajima, Masamitsu Murase, Satoshi Kaijiri, Kentaro Yamada, Takahiro Nakamura, Yuji Hasegawa, Hiroshi G. Okuno, Hiroshi Tsujino
ICASSP (4)1
2006 Genetic Algorithm-Based Improvement of Robot Hearing Capabilities in Separating and Recognizing Simultaneous Speech Signals
Shun'ichi Yamamoto, Kazuhiro Nakadai, Mikio Nakano, Hiroshi Tsujino, Jean-Marc Valin, Ryu Takeda, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno
IEA/AIE2
2006 Real-Time Tracking of Multiple Sound Sources by Integration of In-Room and Robot-Embedded Microphone Arrays
abstract
Real-time and robust sound source tracking is an important function for a robot operating in a daily environment, because the robot should recognize where a sound event such as speech, music and other environmental sounds originate from. This paper addresses real-time sound source tracking by real-time integration of an in-room microphone array (IRMA) and a robot-embedded microphone array (REMA). The IRMA system consists of 64 ch microphones attached to the walls. It localizes multiple sound sources based on weighted delay-and-sum beam-forming on a 2D plane. The REMA system localizes multiple sound sources in azimuth using eight microphones attached to a robot's head on a rotational table. The localization results are integrated to track multiple sound sources by using a particle filter in real-time. The experimental results show that particle filter based integration improved accuracy and robustness in multiple sound source tracking even when the robot's head was in rotation
Kazuhiro Nakadai, Hirofumi Nakajima, Masamitsu Murase, Hiroshi G. Okuno, Yuji Hasegawa, Hiroshi Tsujino
IROS1
2006 Real-Time Robot Audition System That Recognizes Simultaneous Speech in The Real World
abstract
This paper presents a robot audition system that recognizes simultaneous speech in the real world by using robot-embedded microphones. We have previously reported missing feature theory (MFT) based integration of sound source separation (SSS) and automatic speech recognition (ASR) for building robust robot audition. We demonstrated that a MFT-based prototype system drastically improved the performance of speech recognition even when three speakers talked to a robot simultaneously. However, the prototype system had three problems; being offline, hand-tuning of system parameters, and failure in voice activity detection (VAD). To attain online processing, we introduced FlowDesigner-based architecture to integrate sound source localization (SSL), SSS and ASR. This architecture brings fast processing and easy implementation because it provides a simple framework of shared-object-based integration. To optimize the parameters, we developed genetic algorithm (GA) based parameter optimization, because it is difficult to build an analytical optimization model for mutually dependent system parameters. To improve VAD, we integrated new VAD based on a power spectrum and location of a sound source into the system, since conventional VAD relying only on power often fails due to low signal-to-noise ratio of simultaneous speech. We, then, constructed a robot audition system for Honda ASIMO. As a result, we showed that the system worked online and fast, and had a better performance in robustness and accuracy through experiments on recognition of simultaneous speech in a noisy and echoic environment
Shun'ichi Yamamoto, Kazuhiro Nakadai, Mikio Nakano, Hiroshi Tsujino, Jean-Marc Valin, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno
IROS2
2006 Recognition of Simultaneous Speech by Estimating Reliability of Separated Signals for Robot Audition
Shun'ichi Yamamoto, Ryu Takeda, Kazuhiro Nakadai, Mikio Nakano, Hiroshi Tsujino, Jean-Marc Valin, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno
PRICAI3
2005 Towards New Human-Humanoid Communication: Listening During Speaking by Using Ultrasonic Directional Speaker
abstract
This paper presents a new human-humanoid communication system by using a directional speaker. The directional speaker produces directional sound beams by using intermodulation of ultrasonic sound beams and non linearity in air. This technology solves problems and brings new ways of human-humanoid communication as follows: 1) A humanoid can recognize human speeches during speaking because a microphone installed in the humanoid does not capture self voices played by the directional speaker. 2) A humanoid can speak to a specific person as if it whispers. The directional speaker is installed at the position of the mouth of the humanoid. Preliminary experiments show the efficiency of the directional speaker to realize the above two functions.
Kazuhiro Nakadai, Hiroshi Tsujino
ICRA1
2005 Enhanced Robot Speech Recognition Based on Microphone Array Source Separation and Missing Feature Theory
abstract
A humanoid robot under real-world environments usually hears mixtures of sounds, and thus three capabilities are essential for robot audition; sound source localization, separation, and recognition of separated sounds. While the first two are frequently addressed, the last one has not been studied so much. We present a system that gives a humanoid robot the ability to localize, separate and recognize simultaneous sound sources. A microphone array is used along with a real-time dedicated implementation of Geometric Source Separation (GSS) and a multi-channel post-filter that gives us a further reduction of interferences from other sources. An automatic speech recognizer (ASR) based on the Missing Feature Theory (MFT) recognizes separated sounds in real-time by generating missing feature masks automatically from the post-filtering step. The main advantage of this approach for humanoid robots resides in the fact that the ASR with a clean acoustic model can adapt the distortion of separated sound by consulting the post-filter feature masks. Recognition rates are presented for three simultaneous speakers located at 2m from the robot. Use of both the post-filter and the missing feature mask results in an average reduction in error rate of 42% (relative).
Shun'ichi Yamamoto, Jean-Marc Valin, Kazuhiro Nakadai, Jean Rouat, François Michaud, Tetsuya Ogata, Hiroshi G. Okuno
ICRA3
2005 Multiple moving speaker tracking by microphone array on mobile robot
abstract
Real-world applications often require tracking multiple moving speakers for improving human-robot interactions and/or sound source separation. This paper presents multiple moving speaker tracking using an 8ch microphone array system installed on a mobile robot. This problem is difficult because the system does not assume that sound sources and/or the microphone array are fixed. Our solutions consist of two key ideas – time delay of arrival estimation, and multiple Kalman filters. The former localizes multiple sound sources based on beamforming in real time. Non-linear movements are tracked by using a set of Kalman filters with different history lengths in order to reduce errors in tracking multiple moving speakers under noisy and echoic environments. For quantitative evaluation of the tracking, motion references of sound sources and a mobile robot, called SIG2, were measured accurately by ultrasonic 3D tag sensors. As a result, we showed that the system tracked three simultaneous sound sources even when SIG2 moved in a room with large reverberation due to glass walls. 1.
Masamitsu Murase, Shun'ichi Yamamoto, Jean-Marc Valin, Kazuhiro Nakadai, Kentaro Yamada, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno
INTERSPEECH4
2005 Implementation of active direction-pass filter on dynamically reconfigurable processor
abstract
In this paper, we report the design and implementation of a sound source separation system using a dynamically reconfigurable device. A robot in real-world environments should have an ability to treat a mixture of multiple sound signals. Active direction-pass filter (ADPF) which extracts sound from a specific direction by using a pair of microphones has been developed as such a method of sound source separation. The ADPF was used as a front-end for an automatic speech recognition system, and recognition of three simultaneous speech signals has been reported. The ADPF, however, requires a lot of computational power, while the battery capacity and the physical size of the robot are limited. To reduce the power consumption and the size of the system, we adopted the dynamically reconfigurable device, DRP developed by NEC Electronics. We implemented the ADPF on DRP, and investigated the effectiveness of dynamically reconfigurable device for these applications. The preliminary experiment shows that ADPF on DRP separates a mixture of sound sources in real-time with practical accuracy.
Shunsuke Kurotaki, Noriaki Suzuki, Kazuhiro Nakadai, Hiroshi G. Okuno, Hideharu Amano
IROS3
2005 Sound source tracking with directivity pattern estimation using a 64 ch microphone array
abstract
In human-robot communication, a robot should distinguish between voices uttered by a human and those played by a loudspeaker such as on a TV or a radio. This paper addresses detection of actual human voices by using a microphone array as an extension of auditory function of the robot to support environmental understanding by the robot. We introduce a 64 ch microphone array system in a room and propose a new method based on weighted delay-and-sum beamforming to estimate a directivity pattern of a sound source. The microphone array system localizes a sound source and estimates its directivity pattern. The directivity pattern estimation has two advantages as follows: One is that the system can detect whether the sound source is an actual human voice or not by comparing the estimated directivity pattern with prerecorded directivity patterns. The other is that the heading of the sound source is estimated by detecting the angle with the highest power in the directivity pattern. As a result, we proved the effectiveness of our microphone array through sound source tracking with orientation and detection of actual human voices based on directivity pattern estimation.
Kazuhiro Nakadai, Hirofumi Nakajima, Kentaro Yamada, Yuji Hasegawa, Takahiro Nakamura, Hiroshi Tsujino
IROS1
2005 A two-layer model for behavior and dialogue planning in conversational service robots
abstract
This paper presents a model for the behavior and dialogue planning module of conversational service robots. Most of the previously built conversational robots cannot perform dialogue management necessary for accurately recognizing human intentions and providing information to humans. This model integrates robot behavior planning models with spoken dialogue management that is robust enough to engage in mixed-initiative dialogues in specific domains. It has two layers; the upper layer is responsible for global task planning using hierarchical planning and the lower layer engages in local planning by utilizing modules called experts, which are specialized for performing certain kind of tasks by performing physical actions and engaging in dialogues. This model enables switching and canceling tasks based on recognized human intentions. A preliminary implementation of the model, which has been integrated with Honda ASIMO, has shown its effectiveness.
Mikio Nakano, Yuji Hasegawa, Kazuhiro Nakadai, Takahiro Nakamura, Johane Takeuchi, Toyotaka Torii, Hiroshi Tsujino, Naoyuki Kanda, Hiroshi G. Okuno
IROS3
2005 Making a robot recognize three simultaneous sentences in real-time
abstract
A humanoid robot under real-world environments usually hears mixtures of sounds, and thus three capabilities are essential for robot audition; sound source localization, separation, and recognition of separated sounds. We have adopted the missing feature theory (MFT) for automatic recognition of separated speech, and developed the robot audition system. A microphone array is used along with a real-time dedicated implementation of geometric source separation (GSS) and a multi-channel post-filter that gives us a further reduction of interferences from other sources. The automatic speech recognition based on MFT recognizes separated sounds by generating missing feature masks automatically from the post-filtering step. The main advantage of this approach for humanoid robots resides in the fact that the ASR with a clean acoustic model can adapt the distortion of separated sound by consulting the post-filter feature masks. In this paper, we used the improved Julius as an MFT-based automatic speech recognizer (ASR). The Julius is a real-time large vocabulary continuous speech recognition (LVCSR) system. We performed the experiment to evaluate our robot audition system. In this experiment, the system recognizes a sentence, not an isolated word. We showed the improvement in the system performance through three simultaneous speech recognition on the humanoid SIG2.
Shun'ichi Yamamoto, Kazuhiro Nakadai, Jean-Marc Valin, Jean Rouat, François Michaud, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno
IROS2
2004 Improvement of Robot Audition by Interfacing Sound Source Separation and Automatic Speech Recognition with Missing Feature Theory
abstract
We have been developed robot audition system using the active direction-pass filter (ADPF) with the Scattering Theory, and demonstrated that the humanoid SIG could separate and recognize three simultaneous speeches originating from different directions. This is the first result that a robot can listen to several things simultaneously. However, its general applicability to other robots is not yet confirmed. Since automatic speech recognition (ASR) requires direction- and speaker-dependent acoustic models, it is difficult to adapt various kinds of environments. In addition ASR with lots of acoustic models causes slow processing. In this paper, these three problems are resolved. First, we confirmed the generality of the ADPF by applying it to two humanoids, SIG2 and Replie, under different environments. Next, we present the new interface between ADPF and ASR based on the Missing Feature Theory, which masks broken features of separated sound to make them unavailable to ASR. This new interface improved the recognition performance of three simultaneous speeches up to about 90%. Finally, since the ASR uses only a single acoustic model that is direction- and speaker-independent and created under clean environments, the processing of the whole system was made very light and fast.
Shun'ichi Yamamoto, Kazuhiro Nakadai, Hiroshi Tsujino, Toshio Yokoyama, Hiroshi G. Okuno
ICRA2
2004 Multimodal expression for humanoid robots by integration of human speech mimicking and facial color
Tokitomo Ariyoshi, Kazuhiro Nakadai, Hiroshi Tsujino
INTERSPEECH2
2004 Assessment of general applicability of robot audition system by recognizing three simultaneous speeches
abstract
Robot audition is a critical technology in creating an intelligent robot operating in daily environments. We have developed such a robot audition system by using a new interface between sound source separation and automatic speech recognition (ASR). A mixture of speeches captured with a pair of microphones installed in the ear positions of a humanoid is separated into each speech by using active direction-pass filter (ADPF). The ADPF extracts a sound source originating from a specific direction in real-time by using interaural phase and intensity differences. The separated speech is recognized by a speech recognizer based on the missing feature theory (MFT). By using a missing feature mask, the MFT based ASR neglects distorted and missing features caused during the speech separation. A missing feature mask for each separated speech is generated in speech separation and is sent to the ASR with the separated speech. Thus, this new integration improves the performance of ASR. However, the generality of this robot audition system has not been assessed so far. In this paper, we assess its general applicability by implementing it on the three humanoids, i.e., ASIMO of Honda, SIG2, and Replie of Kyoto University. By using three simultaneous speeches as benchmarks, the robot audition system improved the performance of ASR over 50% in every humanoid, and thus its general applicability was confirmed.
Shun'ichi Yamamoto, Kazuhiro Nakadai, Hiroshi Tsujino, Hiroshi G. Okuno
IROS2
2004 Sound and Visual Tracking for Humanoid Robot
Hiroshi G. Okuno, Kazuhiro Nakadai, Tino Lourens, Hiroaki Kitano
Appl. Intell.2
2004 Improvement of recognition of simultaneous speech signals using AV integration and scattering theory for humanoid robots
Kazuhiro Nakadai, Daisuke Matsuura, Hiroshi G. Okuno, Hiroshi Tsujino
Speech Commun.1
2004 Effects of increasing modalities in recognizing three simultaneous speeches
Hiroshi G. Okuno, Kazuhiro Nakadai, Hiroaki Kitano
Speech Commun.2
2003 Robot recognizes three simultaneous speech by active audition
abstract
Robots should listen to and recognize speeches with their own ears under noisy environments and simultaneous speeches to attain smooth communications with people in a real world. This paper presents three simultaneous speech recognition based on active audition which integrates audition with motion. Our robot audition system consists of three modules - a real-time human tracking system, an active direction-pass filter (ADPF) and a speech recognition system using multiple acoustic models. The real-time human tracking realizes robust and accurate sound source localization and tracking by audio-visual integration. The performance of localization shows that the resolution of the center of the robot is much higher than that of the peripheral. We call this phenomenon "auditory fovea" because it is similar to visual fovea (high resolution in the center of the human eye). Active motions such as being directed at the sound source improve localization because of making the best use if the auditory fovea. The ADPF realizes accurate and fast sound separation by using a pair of microphones. The ADPF separates sounds originating from the specified direction obtained by the real-time human tracking system. Because the performance of separation depends on the accuracy of localization, the extraction of sound from the front direction is more accurate than that of sound from the periphery. This means that the pass range of ADPF should be narrower in the front direction than in periphery. In other words, such active pass range control improves sound separation. The separated speech is recognized by the speech recognition using multiple acoustic models that integrates multiple results to output the result with the maximum likelihood. Active motions such as being directed at a sound source improve speech recognition because it realizes not only improvement of sound extraction but also easier integration of the results using face ID by face recognition. The robot audition system improved by active audition is implemented on an upper-torso humanoid. The system attains localization, separation and recognition of three simultaneous speeches and the results proves the efficiency of active audition.
Kazuhiro Nakadai, Hiroshi G. Okuno, Hiroaki Kitano
ICRA1
2003 Realizing personality in audio-visually triggered non-verbal behaviors
abstract
Controlling robot behaviors becomes more important recently as active perception for robot, in particular active audition in addition to active vision, has made remarkable progress. We are studying how to create social humanoids that perform actions empowered by real-time audio-visual tracking of multiple talkers. In this paper, we present personality as means of controlling on-verbal behaviors. It consists of two dimensions, dominance vs. submissiveness and friendliness vs. hostility, based on the interpersonal theory in psychology. The upper-torso humanoid SIG equipped with real-time audio-visual multiple-talker tracking system is used as a testbed for social interaction. As a companion robot, with friendly personality, it turns toward new sound source in order to show its attention, while with hostile personality, it turns away a new sound source. As a receptionist robot with dominant personality, it focuses its attention on the current customer, while with submissive personality, its attention to the current customer is interrupted by a new one.
Hiroshi G. Okuno, Kazuhiro Nakadai, Hiroaki Kitano
ICRA2
2003 Design and Implementation of Personality of Humanoids in Human Humanoid Non-verbal Interaction
Hiroshi G. Okuno, Kazuhiro Nakadai, Hiroaki Kitano
IEA/AIE2
2003 Three simultaneous speech recognition by integration of active audition and face recognition for humanoid
Kazuhiro Nakadai, Daisuke Matsuura, Hiroshi G. Okuno, Hiroshi Tsujino
INTERSPEECH1
2003 Applying scattering theory to robot audition system: robust sound source localization and extraction
abstract
Robot audition by its own ears (microphones) is essential for natural human-robot communication and interface. Since a microphone is embedded in the head of a robot, the head-related transfer function (HRTF) plays an important role in sound source localization and extraction. Usually, from binaural input, the interaural phase difference (IPD) and interaural intensity difference (IID) are calculated, and then the direction is determined by using IPD and IID with HRTF. The problem of HRTF-based sound source localization is that a HRTF should be measured for each robot in an anechoic chamber, because it depends on the shape of robot's head; HRTF should be interpolated to manipulate a moving talker, because it is available only for discrete azimuth and elevation. To cope with these problems of HRTF, we proposed the auditory epipolar geometry as a continuous function of IPD and IID to dispense with HRTF and have developed a real-time multiple-talker tracking system. This auditory epipolar geometry, however, does not give a good approximation to IID of all range and IPD of peripheral areas. In this paper, the scattering theory in physics is employed to take into consideration the diffraction of sounds around robot's head for better approximation of IID and IPD. The resulting system shows that it is efficient for localization and extraction of sound at higher frequency and from side directions.
Kazuhiro Nakadai, Daisuke Matsuura, Hiroshi G. Okuno, Hiroaki Kitano
IROS1
2002 Real-Time Speaker Localization and Speech Separation by Audio-Visual Integration
abstract
Robot audition in real-world should cope with motor and other noises caused by the robot's own movements in addition to environmental noises and reverberation. This paper reports how auditory processing is improved by audio-visual integration with active movements. The key idea resides in hierarchical integration of auditory and visual streams to disambiguate auditory or visual processing. The system runs in real-time by using distributed processing on 4 PCs connected by a Gigabit Ethernet. The system implemented in a upper-torso humanoid tracks multiple talkers and extracts speech from a mixture of sounds. The performance of epipolar geometry based sound source localization and sound source separation by active and adaptive direction-pass filtering is also reported.
Kazuhiro Nakadai, Ken-ichi Hidai, Hiroshi G. Okuno, Hiroaki Kitano
ICRA1
2002 Social Interaction of Humanoid RobotBased on Audio-Visual Tracking
Hiroshi G. Okuno, Kazuhiro Nakadai, Hiroaki Kitano
IEA/AIE2
2002 Real-time sound source localization and separation for robot audition
abstract
Robot audition in the real world should cope with environment noises and reverberation and motor noises caused by the robot's own movements. This paper presents the active direction-pass filter (ADPF) to separate sounds originating from the specified direction with a pair of microphones. The ADPF is implemented by hierarchical integration of visual and auditory processing with hypothetical reasoning on interaural phase difference (IPD) and interaural intensity difference (IID) for each subband. In creating hypotheses, the reference data of IPD and IID is calculated by the auditory epipolar geometry on demand. Since the performance of the ADPF depends on the direction, the ADPF controls the direction by motor movement. The human tracking and sound source separation based on the ADPF is implemented on an upper-torso humanoid and runs in real-time with 4 PCs connected over Gigabit ethernet. The signal-to-noise ratio (SNR) of each sound separated by the ADPF from a mixture of two speeches with the same loudness is improved to about 10 dB from 0 dB.
Kazuhiro Nakadai, Hiroshi G. Okuno, Hiroaki Kitano
INTERSPEECH1
2002 Auditory fovea based speech enhancement and its application to human-robot dialog system
Kazuhiro Nakadai, Hiroshi G. Okuno, Hiroaki Kitano
INTERSPEECH1
2002 Auditory fovea based speech separation and its application to dialog system
abstract
This paper presents an active direction-pass filter (ADPF) that separates sounds originating from the specified direction by using a pair of microphones. Its application to front-end processing for speech recognition is also reported. Since the performance of sound source separation by the ADPF depends on the accuracy of sound source localization (direction), various localization modules including the interaural phase difference, interaural intensity difference for each sub-band, and other visual and auditory processing are integrated hierarchically. The resulting performance of auditory localization varies according to the relative position of the sound source. The resolution of the center of the robot is much higher than that of peripherals, indicating similar property of visual fovea. To make the best use of this property, the ADPF controls the direction of a head by motor movement. In order to recognize sound streams separated by the ADPF, a hidden Markov model based automatic speech recognition is built with multiple acoustic models trained by the output of the ADPF under different conditions. A preliminary dialog system is thus implemented on an upper-torso humanoid. The experimental results prove that it works well even when two speakers speak simultaneously.
Kazuhiro Nakadai, Hiroshi G. Okuno, Hiroaki Kitano
IROS1
2002 Realizing Audio-Visually Triggered ELIZA-Like Non-verbal Behaviors
Hiroshi G. Okuno, Kazuhiro Nakadai, Hiroaki Kitano
PRICAI2
2001 A computational model of monkey grating cells for oriented repetitive alternating patterns
Tino Lourens, Kazuhiro Nakadai, Hiroshi G. Okuno, Hiroaki Kitano
ESANN2
2001 Graph extraction from color images
Tino Lourens, Kazuhiro Nakadai, Hiroshi G. Okuno, Hiroaki Kitano
ESANN2
2001 Sound and Visual Tracking for Humanoid Robot
Hiroshi G. Okuno, Kazuhiro Nakadai, Tino Lourens, Hiroaki Kitano
IEA/AIE2
2001 Real-Time Auditory and Visual Multiple-Object Tracking for Humanoids
Kazuhiro Nakadai, Ken-ichi Hidai, Hiroshi Mizoguchi, Hiroshi G. Okuno, Hiroaki Kitano
IJCAI1
2001 Real-time multiple speaker tracking by multi-modal integration for mobile robots
abstract
In this paper, real-time multiple speaker tracking is addressed, because it is essential in robot perception and humanrobot social interaction. The difficulty lies in treating a mixture of sounds, occlusion (some talkers are hidden) and real-time processing. Our approach consists of three components; (1) the extraction of the direction of each speaker by using interaural phase difference and interaural intensity difference, (2) the resolution of each speaker's direction by multi-modal integration of audition, vision and motion with canceling inevitable motor noises in motion in case of an unseen or silent speaker, and (3) the distributed implementation to three PCs connected by TCP/IP network to attain real-time processing. As a result, we...
Kazuhiro Nakadai, Ken-ichi Hidai, Hiroshi G. Okuno, Hiroaki Kitano
INTERSPEECH1
2001 Separating three simultaneous speeches with two microphones by integrating auditory and visual processing
abstract
This paper addresses the problem of automatic recognition of three simultaneous speeches with two microphones, that is, that of sound source separation where the number of sound sources is greater than that of microphones. The approach used is the direction-pass filter, which is implemented by hypothetical reasoning on the interaural phase difference (IPD) and interaural intensity difference (IID). Auditory processing calculates IPD and IID for each subband, and generates hypotheses for precalculated IPD and IID for every direction including one obtained by visual processing. Then the system calculates the belief factor of hypothesis by Dempster-Shafer theory and determines the direction of each subband. Subbands of the specific direction are collected and then converted to a wave form by inverse FFT. With 200 benchmarks of three simultaneous utterances of Japanese words, the average 1-best and 10-best recognition rates of extracted speeches are 60% and 81%, respectively.
Hiroshi G. Okuno, Kazuhiro Nakadai, Tino Lourens, Hiroaki Kitano
INTERSPEECH2
2001 Epipolar geometry based sound localization and extraction for humanoid audition
abstract
Sound localization for a robot or an embedded system is usually solved by using inter-aural phase difference (IPD) and inter-aural intensity difference (IID). These values are calculated by using head-related transfer function (HRTF). However, the HRTF depends on the shape of the head and also on changes of the environments. Therefore, sound localization without HRTF is needed for real-world applications. In this paper, we present a new sound localization method based on auditory epipolar geometry with motion control. The auditory epipolar geometry is an extension of an epipolar geometry in stereo vision to audition, and auditory and visual epipolar geometries can share the sound source direction. The key idea is to exploit additional inputs obtained by the motor control in order to compensate damages in the IPD and IID caused by reverberation of the room and the body of the robot. The proposed system can localize and extract simultaneously two sound sources in a real-world room.
Kazuhiro Nakadai, Hiroshi G. Okuno, Hiroaki Kitano
IROS1
2001 Human-robot interaction through real-time auditory and visual multiple-talker tracking
abstract
Nakadai et al. (2001) have developed a real-time auditory and visual multiple-talker tracking technique. In this paper, this technique is applied to human-robot interaction including a receptionist robot and a companion robot at a party. The system includes face identification, speech recognition, focus-of-attention control, and sensorimotor task in tracking multiple talkers. The system is implemented on a upper-torso humanoid and the talker tracking is attained by distributed processing on three nodes connected by 100Base-TX network. The delay of tracking is 200 msec. Focus-of-attention is controlled by associating auditory and visual streams by using the sound source direction and talker position as a clue. Once an association is established, the humanoid keeps its face to the direction of the associated talker.
Hiroshi G. Okuno, Kazuhiro Nakadai, Ken-ichi Hidai, Hiroshi Mizoguchi, Hiroaki Kitano
IROS2
2000 Design and architecture of SIG the humanoid: an experimental platform for integrated perception in RoboCup humanoid challenge
abstract
In this paper, we report an initial design of humanoid head platform for RoboCup humanoid challenge. While many researches in RoboCup humanoid challenge naturally focus on walking and running behaviors, we focus on perception and high-level behavior issues using an upper-torso humanoid. We have designed a head/neck part of the humanoid with aesthetically designed appearance, and various processing for early perception and associated reflex. The goal of this paper is to illustrate issues in designing humanoid head platform for high-level cognition research, and provide some of initial insights obtained during the early stage of implementations. With regards to perception, the emphasis will be placed on how interaction between different perception channel and motor control can complement each other to improve processing.
Hiroaki Kitano, Hiroshi G. Okuno, Kazuhiro Nakadai, Theo Sabisch, Tatsuya Matsui
IROS3
2000 Active audition system and humanoid exterior design
abstract
We present a humanoid active audition system with improved noise cancellation using humanoid cover acoustics and the cover design for an industrial exterior design. For active audition, it is an important problem to distinguish between the internal sounds like motor noises and sounds which originate from the outer world. It is important that inevitable motor noise is cancelled while the humanoid is in motion. Therefore, humanoids require the exterior to cancel such internal noises because it separates humanoid inner world from the outer world. We report an active audition system focused on the sound source tracking by integrating audition, vision and motor movements using sound separation of the humanoid cover. The experiments show that the humanoid can track and localize sound sources more accurately, the noise cannot always be cancelled out optimally and further improvements are required.
Kazuhiro Nakadai, Tatsuya Matsui, Hiroshi G. Okuno, Hiroaki Kitano
IROS1
2000 Humanoid Active Audition System Improved by the Cover Acoustics
Kazuhiro Nakadai, Hiroshi G. Okuno, Hiroaki Kitano
PRICAI1
2000 And the Fans Are Going Wild! SIG plus MIKE
Ian Frank, Kumiko Tanaka-Ishii, Hiroshi G. Okuno, Junichi Akita, Yukiko Nakagawa, Kazuaki Maeda, Kazuhiro Nakadai, Hiroaki Kitano
RoboCup7
1995 Organization of Hierarchical Perceptual Sounds: Music Scene Analysis with Autonomous Processing Modules and a Quantitative Information Integration Mechanism
Kunio Kashino, Kazuhiro Nakadai, Tomoyoshi Kinoshita, Hidehiko Tanaka
IJCAI2